An AI assistant that searches well first, then answers
An assistant that answers only from the client's documents. In testing it picked the wrong passages, made up links and repeated its answers. I rebuilt the search with Anthropic's contextual technique, then fixed the rules.
The context
The client wanted an AI assistant for the users of their web application, able to answer from their documents, about 30,000 passages of text. The original project had clear rules:
- the assistant answers only with what it finds; if the search finds nothing, the model is not even called and the answer is “I can't find it”;
- personal questions are passed on to a person;
- a fixed notice, one question at a time per person, a spending cap beyond which the assistant replies “not available”;
- a log of every exchange and no personal data sent to the model.
The problem
Five problems emerged in testing:
- the model was receiving the wrong passages, such as introductions and off-topic text;
- the assistant made up links;
- for broad questions it used only one document;
- “continue” received the same passages as the previous question, and answers were short;
- the instruction text edited from the panel had replaced the safety rules.
The solution
- Context inside every passage, with Anthropic's technique. I used the same technique described by Anthropic, Contextual Retrieval: before indexing, every passage receives its context, that is its path within the documents (section › chapter › document), both in the keyword index and in the meaning index. A passage that says “as seen above” now knows what it is talking about. The context comes from the structure of the documents, so indexing costs nothing.
- Hybrid search, inside the database. A keyword search, where rare words and the document title weigh more, and a search by meaning with a compact embedding model. The vectors live in the application's own relational database (PostgreSQL with pgvector, HNSW index), with no external service. If the search by meaning fails, it falls back to keyword search on its own.
- Controlled selection. For each question the search proposes up to 40 candidate passages; at most 30 reach the model, with a limit per document and a minimum number of different documents, so broad questions draw on several sources. The values can be set from the panel.
- A prompt in two parts. The fixed part, in the code, holds the formatting rules: how to cite sources and no links that are not in the documents. The part from the panel, which the client can edit, holds rules and tone. The safety rules that the edited text had wiped out have been put back: no financial, legal or medical advice, and the instructions are not revealed.
- “Continue” that moves forward. Now “continue” excludes the passages already used and brings new ones; after five requests the assistant says it has said everything. And the prompt no longer asks for short answers at all costs.
- The “cache” was the memory. The “old” answers came from the conversation memory. Now the conversation restarts after 12 hours and whenever the settings that affect answers change. For every answer, a trace of what the assistant read stays in the panel for 30 days.
The result, in testing
- On a real question from the client the assistant now cites 6 documents in a single answer; before, it cited one. This is an example, not an average.
- The search responds in 20-30 milliseconds across about 30,000 passages.
- The index is 38% lighter, from 106 to 65 MB, and rebuilds in 15 seconds.
- An answer with 30 passages costs about one cent.
What you can take away
- The quality of an AI assistant depends first and foremost on the search, not on the model.
- A passage without context is ambiguous: telling it where it comes from improves the search, and if the context is already in the structure of the documents it costs nothing.
- Rules that must not change belong in the fixed part of the prompt, not in a text anyone can edit.
- Odd behaviour should be understood before it is fixed: the “cache” was the memory.
Details have been changed to protect the client's confidentiality.
Benefits
- Answers built only on what is in the documents
- Broad questions that draw on several documents at once
- Clear safety rules: no financial, legal or medical advice
- Predictable costs, about one cent per answer
Features
- Contextual Retrieval: every passage carries its path within the documents
- Hybrid search by keyword and by meaning, with automatic fallback to keyword search
- Vectors stored in the application's database, with no external services
- A 30-day trace of what the assistant read to answer
- One question at a time per person and a spending cap
Facing a similar problem?
Tell me what you need: I reply myself, usually within one working day.