We have built these for a handful of accounting and tax firms. Skipping any one of the three is usually why the project quietly dies within six months.
Written in response to a recurring question in public communities: has anyone built an internal knowledge base for a small accounting firm?
Tax rules change constantly, and old versions pile up on the server. Dense retrieval has no concept of time — it ranks on semantic similarity. An obsolete treatment can match the question better than the current one and get served with total confidence. People usually find out when a client asks about it.
Put an effective date and an expiry date on every document, and archive the old version out of the index when a new one lands. Do not leave them side by side and hope ranking sorts it out. This is baseline, not an optimization.
The real question is rarely "find me notice 2026-XX." It is "how do we treat this situation for this client." That answer is spread across three places: the regulation, last year's working papers, and somebody senior's judgment. The first two are on disk. The third isn't, and no retrieval system will ever reach it.
Organize entries by scenario, each with a named owner and a review date. Entries without an owner rot silently — that is the single most common reason these systems stop being trusted.
Working papers, schedules, filings. If chunking separates the header row from the data rows, "what was this line in Q2" becomes unanswerable — and the system will not say it does not know, it will offer a plausible number from somewhere else.
Parse tables structurally before they go in, preserving the header-to-row relationship. Skipping this step is how firms end up concluding "the system is wrong."
Working papers and client identifiers are highly sensitive. Tag everything with client ownership and confidentiality level at ingest, and filter the candidate set before retrieval — not after. If you filter what is displayed once the text already went into the model context, you are not really controlling anything.
Is there a question your team answers more than ten times a month?
If yes, it is worth building — start with that one scenario and make it work before expanding.
If not, organize your files first and revisit later. Bulk-importing everything just relocates the mess.
We are in an initial pilot phase. Pricing is confirmed per engagement based on material volume and scenario complexity; we do not publish figures before they are confirmed.
Tell us the one query your team answers most often. We will tell you what can and cannot be done with the material you actually have.