Chatbot on your knowledge base. It answers citing the document, or it doesn't answer.
The same fifteen questions, every day, about the same product sheets and the same policies. This piece answers by drawing on your documents, puts the document it comes from next to the answer, and when the answer isn't there it says so.
Every answer carries its source. Next to it is the document and the point it was taken from, so that checking costs a second instead of a phone call.
“I don't know” is a feature. Outside the agreed boundary the assistant stops and passes to a person, because an assistant that always answers will at some point invent.
Retrieval reduces invention, it doesn't eliminate it. Three legal tools built exactly this way, and sold as hallucination-free, get it wrong between 17% and 33% of the time according to Stanford.
This is the page for a single piece of the system. The other pieces, and the criterion for choosing which to start from, are on the services page.
The problem is the fifteen questions that come back every day
In every business there is a small set of questions that arrives constantly, always the same, and the answer is already written somewhere. The cost isn't the single answer, it's that a person looks for it every time in a different file and in the meantime has stopped doing something else.
The second cost only shows later: whoever answers goes from memory, and the memory of two different people gives two different answers. The trivial question becomes a contradiction in front of a customer, and it is the kind of error nobody records anywhere.
A question repeated fifteen times a day isn't a problem of patience. It's a document nobody has yet put where it's needed.
What changes in practice
The recurring questions stop reaching a person. Whoever asks gets the answer straight away, with the document it comes from next to it, and whoever used to answer finds their day free of those fifteen interruptions.
The second change is that there is only one answer. Everyone receives what is in the approved documents, not what a person remembers reading, and contradictions in front of the customer disappear because there is a single source.
The third is information nobody has ever had: the list of questions the assistant couldn't answer. It is the most honest way to know which documents are missing, and it arrives without asking anyone.
Retrieval from documents reduces invention, and doesn't eliminate it
It is the point where people buy badly, and there is a measurement. Stanford's RegLab group carried out the first preregistered evaluation of three legal research tools built precisely on document retrieval, and sold as hallucination-free.
The result is stated word for word in the paper: «each hallucinate between 17% and 33% of the time». The same authors state that hallucinations are reduced compared with a general-purpose assistant, which is exactly the point: reduced, not removed.
On a simpler task the numbers improve a lot. The Vectara leaderboard, updated to 11 May 2026 on more than 7,700 documents, measures the best model at 1.8% of answers not faithful to the text provided, with 99.5% of questions answered.
Together the two measurements say the useful thing. Summarising a document you have in front of you is reliable; choosing on your own which document is needed and then answering is much less so. Almost all of the building work lies in that second step.
You decide the boundary, beforehand
During analysis you decide what the assistant can answer, what it must refuse and what it must pass to a person. It is a commercial choice, not a technical one, and it has to be made before writing a line: an assistant without a declared boundary will answer everything anyway.
| Type of question | What the assistant does | What you get |
|---|---|---|
| Approved informationit is in the documents | It answers straight away, with the document and the point it took the answer from next to it. |
An identical answer every time, verifiable in a second by whoever receives it. |
| Missing informationit isn't anywhere | It says it doesn't know, passes the question to a person and writes it on the list of things to add. |
No invented answers, and a list that says which documents are really missing. |
| Commitment for the businessprices, discounts, confirmations | It never answers on its own. It collects the request with its context and queues it for whoever decides. |
No figure promised by a machine, and whoever decides receives the request already prepared. |
| Off-topicnothing to do with it | It refuses and brings the conversation back within the boundary, without chasing the question. |
The assistant stays yours, and doesn't become a toy for those who try to derail it. |
What the system does, step by step
Every question goes through the same five steps, and each one leaves a log entry with the question, the documents retrieved and what was answered. That log is why an error gets found, instead of being recounted from memory.
| Step | What happens | What you get |
|---|---|---|
| The baseapproved documents | Sheets, terms, policies and past frequently asked questions go into the base, each with the date of its last check. |
Answers come from documents someone has approved, and you know when. |
| Retrievalwhich pieces are needed | When the question comes, the system retrieves the relevant passages and puts them in front of the model, which answers only on those. |
The answer rests on a text of yours, instead of on what the model remembers about the world. |
| Citationwhere it comes from | Next to the answer appear the document and the exact point, clickable for anyone who wants to check. |
Checking costs a second, and an error shows before it becomes a problem. |
| Refusalwhen it doesn't know | If the retrieved documents don't contain the answer, the assistant says so and passes the question to a person. |
The worst case becomes a delay, instead of wrong information given with confidence. |
| Verificationfixed questions | A set of real questions with the right answer next to them is rerun at every change to the system. |
A deterioration shows immediately, instead of being discovered by a customer three months later. |
The knowledge base is the product; the model is a spare part
An assistant is worth as much as the documents underneath it. If the terms of sale exist in three versions and nobody knows which is right, the assistant will repeat the wrong one with the same confidence it would have repeated the right one.
That is why the work starts from the documents and not from the technology: you choose a good version for each topic, put a date next to it, and decide who updates it. It is the same principle as the base of facts that sales copy and materials works on, and it is best to keep just one for both.
The choice between writing the information directly into the instructions and building a retrieval system is made on volume and update frequency. A few stable pages need no infrastructure, and adding it would be a cost that buys nothing.
The assistant states that it is a system
From 2 August 2026 Article 50 of Regulation (EU) 2024/1689 applies, which requires providers to ensure «that AI systems intended to interact directly with natural persons are designed and developed in such a way that the natural persons concerned are informed that they are interacting with an AI system».
The statement comes before the conversation and takes up one line. What it really means for a small business, with the scope and the four things to put in order, is on the page about Article 50 of the AI Act for Italian SMEs.
In practice it works as a sales argument. Whoever discovers on their own that they were talking to a machine stops trusting even what was true, and the opening line removes that risk instead of adding one. The full reasoning is on the page about the principles we build with.
What goes out on its own and what waits for a person
Everything that repeats already approved information goes out on its own: opening hours, availability, the content of a sheet, written terms, internal procedures. They are things decided beforehand by a person, and the system simply reports them, citing where they are.
Whatever commits the business never goes out on its own. A price, a discount, a confirmation and a delivery promise are collected with their context and queued for whoever decides. Requests that arrive from several channels and need sorting are a different piece, handled by Inbox AI, the first reply to every request.
The same thing changes name with the trade
The mechanism is identical; the knowledge base isn't. It is worth looking at your own case, because that is where you see which document should be put in order first.
| Sector | What is in the base | Where you see it |
|---|---|---|
| Food and agricultureand export | Technical sheets, allergens, formats and pieces per carton, delivery terms. Questions a buyer asks before asking for a price. |
|
| Restaurantsand bars | Updated menu, allergens, opening hours, rules for groups and pets. Questions that arrive in bursts in the hour when nobody can answer. |
The service of a dining room where the phone rings unanswered |
| Hospitalityaccommodation and events | House rules, what is included, hall capacities, arrival and departure times, nearby services. |
What this piece doesn't do
It doesn't promise zero errors, and be wary of anyone who does: the Stanford measurement exists precisely because three large providers had promised it. What gets built is a system in which the error shows, is cited and gets corrected, with a set of fixed questions that keeps it under control over time.
It doesn't replace whoever answers. It removes the repetitive questions and lets those that need judgement reach a person, with more context and sooner.
It doesn't learn on its own from conversations. What goes into the base goes in because someone approved it, and a wrong answer given yesterday doesn't become tomorrow's truth.
Questions and answers
If it draws on our documents, can it still make things up?
Yes, less than before but still yes, and whoever promises otherwise is promising something nobody has demonstrated. Stanford's RegLab group evaluated three legal research tools built on exactly this mechanism, sold as hallucination-free, and measured that they get it wrong between 17% and 33% of the time.
That is why the system cites next to every answer the document it comes from, so that checking costs a second, and why the boundary of admissible questions is decided beforehand.
What happens when the answer isn't in the documents?
The assistant says it doesn't know and passes the question to a person, with a note of what was asked. It is a feature and not a fault: an assistant that always answers is an assistant that at some point invents, because it has no other way of answering.
Questions that fell outside the boundary are collected in a list, and that list is the most honest way to understand which documents are really missing from the base.
Do you need a document retrieval system or is a prompt enough?
It depends on how much knowledge there is to handle, and it is decided during analysis with a simple criterion. If the information fits in a few stable pages, it is written directly into the assistant's instructions and no retrieval infrastructure is needed.
If there are dozens of documents that change, you need a system that indexes them and retrieves them at the moment of the question. The choice is made on volume and update frequency, never to have one more technology in the quote.
Does the assistant have to say it is a system?
Yes, and from 2 August 2026 Article 50 of Regulation (EU) 2024/1689 applies, which requires providers to ensure that natural persons are informed that they are interacting with an AI system. The statement comes before the conversation, not after, and it is a single line.
In practice it also works as a sales argument, because whoever discovers on their own that they were talking to a machine stops trusting even what was true.
How do you measure whether it's working?
With the accuracy of the answer against the real knowledge base, measured on a fixed set of questions built before starting. The questions that really come in are written down, with the right answer taken from the documents next to them, and the same set is rechecked at every change to the system.
The second number is the share of questions to which the assistant replied that it didn't know. It has to stay stable, because if it suddenly drops it means it has started answering things it doesn't know.
Notes on sources
- The 17% and 33% come from the first preregistered evaluation of legal research tools built on document retrieval, by Stanford's RegLab group, with the full text of the paper Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. It measures three US legal tools on questions of law, a domain much harder than questions about a product sheet: the number isn't a forecast about your case, it shows that the mechanism alone isn't enough.
- The 1.8% and 99.5% come from Vectara's public leaderboard, updated to 11 May 2026 on more than 7,700 documents. It measures the faithfulness of a summary to a text already provided, so the easier of the two tasks, and the leaderboard is published by the same company that sells the evaluation model: we cite it because method and data are public and can be rechecked.
- The obligation to inform those who talk to a system is in Article 50 of Regulation (EU) 2024/1689, published in the Official Journal of the European Union, applicable from 2 August 2026. The scope for a small business is reconstructed in our page on the AI Act for Italian SMEs. This page doesn't provide legal advice: the exact scope of an obligation depends on the case and should be taken to whoever assists the business.
- We don't publish any forecast of the share of requests an assistant will end up handling. The estimates going around on this point are projections by analyst firms on markets and future years, and they say little about a single business with its own knowledge base. The share is measured on your own case, by counting the real questions of a month.
- This page doesn't report results obtained for a client, because this piece hasn't yet been delivered to a client. The tests cited are functional checks carried out in testing.
The other pieces in this group
Customers who are out there, and don't reach youFifteen minutes, with your case in front of us.
For a week, note down the questions that reach you more than three times, and next to each write which file holds the answer. If for half of those questions the file doesn't exist, that is the real work. In fifteen minutes on the phone we look at it together and tell you where it makes sense to start, even if we don't end up working together.
You get Mattia Esposito, who then builds the system: there's no salesperson in between. If you'd rather measure on your own before talking, the Diagnostico (in Italian) is twenty questions and five minutes.