Skip to content

AI & automation

Why Does Your Chatbot Give Wrong Answers When Your Pages Are Correct?

Grounding is a claim about provenance, not about intelligence. Six ways that claim fails when the underlying content was fine all along, the control that addresses each, and why an assistant that can decline is worth more than one that cannot.

Zealsync Insights22 min read
Why an assistant answers wrongly from good pages

Grounding is a claim about provenance, not about intelligence

When somebody selling you a website assistant says the answers are grounded in your own content, read the promise literally, because it is narrower than it sounds. It says the material used to construct the answer came from your pages rather than from somewhere else. It does not say the right page was found, that the whole of the relevant page was used, that the qualifications on it survived, or that the sentence finally produced is one you would have signed off. Provenance is not correctness. It is a constraint on where the words may come from, and constraints on sourcing do not by themselves make a claim true.

That distinction is the whole of this piece, because the failure you are investigating happened on the far side of it. A page on your site said the right thing, in the right words, and the assistant said something else in front of a customer. Nothing was wrong with the content. Something went wrong between the content and the answer, and the useful question is which of a small number of things it was.

The scope here is strictly runtime. If your source material is thin, contradictory, out of date, or scattered across four pages that each carry a third of the truth, you have a different problem with a different fix, and it belongs to the work of preparing your website content before an assistant answers for you. Assume for everything below that a competent colleague reading the same pages would have answered the question correctly, and that the machine did not. There are six ordinary ways that happens.

Each is described the way an owner encounters it rather than the way an implementer would debug it: the symptom on screen, the tell that identifies it, and the control that addresses it. Four of the six controls are decisions about content and policy that only you can take. None of them requires you to understand how retrieval is configured, and you should be suspicious of any account of this subject that makes your own job sound technical.

Three failures of retrieval

Retrieval is the step that decides which of your pages, or which fragments of them, are placed in front of the answering step. It is a search problem in different clothing, and it fails the way search has always failed: by missing things that were there, by returning a near-miss with great confidence, and by returning several things at once with no indication of how they relate to each other. The three failures below are those three, and they produce three quite different symptoms.

Retrieval missed a source that was sitting there

The page exists. It says the thing. The assistant either declined the question or answered it from something adjacent and weaker. The tell is that you can find the answer yourself in under a minute, from memory or from the site's own search, and you are left wondering how the thing missed it.

Two causes account for most of it. The first is vocabulary: the visitor asked in their words and the page is written in yours. Buyers describe problems, sites describe services, and the two vocabularies overlap far less than the people who wrote the site believe. The second is form. Content that lives in a PDF, inside a table, in an image of a price list, behind a tab that only renders after a click, or on a page reachable only through a filtered view, may never have been taken in at all. It is published, in the sense that a human can reach it. It is not necessarily present, in the sense that matters here.

This failure is the quietest of the six. Nobody complains about an answer that was merely less good than it could have been, so the gap persists until somebody goes looking. The control is a coverage check: a written list of the questions the assistant must be able to answer, with the page that is the authority for each, and a periodic test that every one of them still resolves to the right source. It is dull, it takes an hour, and it is the only one of these six failures you can reliably detect before a customer does.

Content from an unrelated page contaminated the answer

The answer is fluent, specific, and about something else. It describes a service you retired, a variant meant for a different market, or a page written for an audience the visitor is not part of. The distinguishing feature is that the wrong detail is a real detail. Nothing was invented. Something was retrieved, and it was the wrong thing, and it read plausibly enough that the answering step had no reason to doubt it.

Near-miss retrieval gets worse the more your site says similar things in similar language with different qualifications attached. A company with one service page rarely sees this. A company with a page per sector, a page per package, an archive of older wording still live for search reasons, and a set of pages describing work done under terms that have since changed, sees it regularly. The site did not get worse. It got bigger, and the space of plausible near-misses grew with it.

The control is citable sources. Every answer should be traceable to the specific page it came from, and whoever reviews answers should be able to see that trace. This matters more than it sounds, because without it you cannot tell this failure from the two that follow. An answer that is wrong because the wrong page was used and an answer that is wrong because no page was used look identical from the outside, and they call for opposite responses.

Two retrieved sources combined into a claim neither made

This is the expensive one. One page says a particular thing is included, for a particular kind of engagement. Another page describes a timescale, for a different kind of engagement. The answer joins them into a single sentence in which that thing is included on that timescale. Every clause in it is traceable to your own site. The composite claim is one nobody ever made, and it may be one nobody would ever agree to.

The tell is peculiar and worth learning to spot: the answer reads as more complete than any single page on your site. It resolves a question your own content deliberately leaves open, and it resolves it in a satisfying, quotable sentence. That is exactly why it is dangerous. The false claim is the useful-sounding one, so it is the one the buyer remembers, repeats to a colleague, and later holds you to.

The composite answer reads as more complete than any page on your site. That is not a sign of a good assistant. That is the tell.

The control is a rule about combination: where a complete answer would require stitching two sources together, the assistant answers the part it can support and routes the rest onward. That is a real cost. It means fewer complete answers, and it means visitors occasionally being handed a partial response where a fuller one appeared to be available. The point is that this is a decision with a price, and it should be taken deliberately and in advance rather than discovered afterwards in a transcript. If the person selling you the system has never raised it, they have made the decision on your behalf.

Three failures of scope and instruction

Retrieval can do its job perfectly and the answer can still be wrong. The next three failures happen at the point where the answer is written. They are failures of boundary rather than of search, and they are the ones where the owner's decisions, not the supplier's engineering, do most of the work.

The answer generalised beyond what the source supported

Your page carries a qualification and the answer drops it. Something that is true in a stated configuration, above a stated threshold, or for a stated kind of client, arrives as a flat statement of fact. The qualification was the part that made the sentence true, and it is now gone.

Generalisation is not an accident; it is the natural behaviour of fluent summary. Conditions, exceptions and hedges are precisely the material a summariser treats as noise, because in most prose they are noise. In commercial prose they are the substance. The tell is that the answer sounds like the marketing version of your own page: cleaner, bolder and shorter than the thing you carefully wrote.

It bites hardest wherever an exception exists. Turnaround that depends on scope, coverage that depends on region, compatibility that depends on what the client already runs, inclusions that depend on which arrangement was signed. The control is editorial rather than technical: identify the pages where a qualification is load-bearing, mark them, and require answers on those topics to quote or point rather than paraphrase. Nobody outside your business can make that list. It is a judgement about which of your sentences would cost you money if they were repeated without their second half.

A confident answer to a question outside the declared scope

Somebody asks a legal question, a tax question, a question about a competitor, a question about regulation in their country, or a question about something you stopped selling three years ago. There was no source to work from, so the answer was assembled from general knowledge. General knowledge is fluent, frequently plausible, occasionally right, and never specifically yours.

The tell is an absence rather than a presence: the answer contains nothing that could only be true of your business. No page reference, no term from your own vocabulary, no detail with an edge on it. Read a suspect answer and ask whether a competitor could have published the identical paragraph. If they could, it did not come from you.

The sharpest version of this failure is an answer that is correct in general and wrong for you: a sensible-sounding statement about how something of this kind usually works, offered on your site, in your voice, contradicting your own written terms. You are now on record twice, differently. Whether that exposes you to anything is a question for your own adviser rather than for an article, and the fact that it is a reasonable question to ask is the reason a chat widget is not a decorative element of the page.

Instructions inside retrieved content treated as commands

This one is unfamiliar to most owners, so it is worth stating without jargon. Retrieved material arrives as text. A sentence that describes something and a sentence that instructs the reader to do something have no reliable structural difference; both are only words in a document. If a retrieved passage happens to contain a sentence addressed to the assistant, telling it to disregard what it was told and say something else instead, whether that sentence is obeyed is a property of how the system was designed, not a property of the sentence.

Where the text comes from therefore matters. Your own pages, written by your own people, are usually safe. The surface widens as soon as the assistant reads anything you did not write: a page that quotes correspondence verbatim, a syndicated feed, a supplier-provided description, a document uploaded by whoever happens to be talking to it, a comment field, a third-party listing embedded for convenience. Each of those is a place where somebody other than you can put words in front of the machine.

The control is a rule with a name: retrieved content is data, never instruction. You do not need to know how that rule is enforced, but you are entitled to ask two questions in writing and to treat vagueness in the answers as the finding. What is the assistant permitted to read? And what happens when a passage it reads addresses it directly? A third question matters more than either: what is the assistant permitted to do. An assistant that can only produce text has an embarrassing worst case. An assistant wired to take actions on your behalf has a different one, and the difference should be understood before the wiring, not after.

The control that addresses each failure

Set out together, the six controls are less intimidating than the six failures. Each is a sentence, and each needs somebody's name written next to it.

  • Missed source: a coverage check. A written list of the questions the assistant must answer, the page that is the authority for each, and a scheduled test that each one still resolves correctly.
  • Contaminated answer: citable sources. Every answer traceable to the page it came from, visible to whoever reviews answers, so that a wrong answer can be diagnosed rather than merely noticed.
  • Combined claim: a stated refusal to stitch. Where a complete answer would need two sources joined, the assistant answers the supportable part and routes the rest onward.
  • Over-generalisation: marked qualifications. A list of the topics where the condition is load-bearing, and a requirement to quote or point rather than paraphrase on those topics.
  • Out-of-scope confidence: a declared scope. Written before launch, naming what must be answered and what must never be, so that there is something for the assistant to decline against.
  • Instruction injection: retrieved content treated as data. A written statement of what the assistant may read, what happens when a passage addresses it directly, and which actions it is permitted to take at all.

Notice how few of those are engineering. Four are decisions about your content and your policy, and the two that are not are things you specify and then verify rather than build. The owner's job in this work is not configuration. It is to say what must be answerable, what must never be answered, which of your qualifications are load-bearing, and what should happen when nothing matches. A supplier can implement all four. No supplier can decide them for you, and one that offers to has told you something about how much they know about your business.

A control with no name written next to it is not a control. It is a preference, and preferences do not survive the second busy week after launch.

Every one of these costs something, and the honest version of the argument says so. Coverage checks cost an hour a month of somebody's attention. Citable sources cost interface space and make answers longer and more cluttered. Refusing to stitch costs completeness. Quoting rather than paraphrasing costs fluency, and the result reads stiffer than the copy around it. A declared scope costs breadth, visibly, on questions you could have answered. If what you want is an assistant that always has something to say, you are asking for one that cannot fail in a way you would notice, which is not the same thing as one that does not fail.

Scope, declared in writing before launch

A scope document for an assistant is short. A page or two, mostly things you already know but have never written down. It has four parts, and the order matters less than the fact that all four exist.

First, the topics it must answer, each with the page that is the authority for it. This is the list the coverage check runs against, and writing it usually surfaces two or three questions you are asked constantly that no page on your site actually answers. Second, the topics it must never answer, named specifically rather than gestured at by category. Not legal things in general, but the actual boundaries: anything that reads as advice, anything about work not published on the site, anything about a named individual, anything that commits to a date, anything that would amount to a promise you would otherwise put in a contract.

Third, the qualified topics, where an answer is permitted but paraphrase is not. Fourth, what happens at the boundary: the wording of the refusal and the route it offers, which is a piece of copywriting and deserves to be treated as one rather than left to a default string somebody shipped.

The argument for writing it before launch is simply that the alternative is writing it afterwards, incident by incident, by whoever is nearest to each incident. Those decisions will not be consistent with one another, nobody will remember which of them were deliberate, and the boundary ends up as a sediment of past embarrassments rather than a position. A written scope is also the thing you hand to the person who reviews answers, to a supplier during a procurement conversation, and to whoever inherits the job when the person who set it up moves on.

The counter-argument is real and should be granted: this looks like bureaucracy for a chat widget. Sometimes it is. If your assistant only ever handles opening hours, delivery areas and where the office is, the scope document is a paragraph, the refusal path is one sentence, and anything more is ceremony for its own sake. Scale the process to the exposure. The test is not how sophisticated the assistant is but how much a wrong answer would cost you, and that is a question about your business rather than about the software.

Designing the refusal: what "I do not know" should actually say

The default refusal is a dead end with an apology attached. A visitor asks a real question, receives a sentence saying the assistant is sorry but does not have information about that, and is left holding the question with nowhere to put it. It is not merely unhelpful. It reads as faintly evasive, because a person who said that in a shop would be understood to be avoiding the subject.

A refusal worth designing does four things. It says plainly that this cannot be answered here, once, without apologising three times. It gives a reason that is true, distinguishing between something outside what the assistant covers, something not published anywhere on the site, and something that needs a person to look at the specifics. It offers the route onward that genuinely exists rather than a generic one. And it keeps what the visitor already typed, so that taking the route does not begin with retyping the question.

Compare two. The first: I am sorry, I cannot help with that. The second: I cannot answer that from what is published here, questions about contract terms go to a person, and you can send this one across without retyping it. The second is longer, and length is not the virtue. The virtue is that it is true, it names which of the three reasons applies, and it ends somewhere. What that somewhere should be is a design question in its own right, and there are more options than a contact form, as the piece on the next actions an assistant can offer sets out.

Two constraints on refusal wording are easy to get wrong. The first is that a refusal must not become a disguised sales prompt. If every declined question funnels to the same enquiry form, visitors work out within two attempts that declining is the mechanism rather than the exception, and the refusal loses the credibility that made it worth having. The second is that a refusal must be distinguishable in the record from a failure. Saying nothing is known about a question that sits squarely inside your declared scope means something broke. Saying a question is outside what this assistant covers means the boundary worked as designed. If both come out as the same sentence, your logs cannot separate coverage gaps from boundary hits, and the review described below has nothing to sort.

Why an assistant that can decline is worth more than one that cannot

An assistant that always answers has thrown away its ability to signal uncertainty. Every response arrives with the same evenness of tone, which means the tone carries no information, which means the reader has no way to weight one answer against another. The rational response to a source whose confidence is uninformative is to discount all of it. People arrive at that conclusion slowly and without articulating it, and then they stop using the thing, and usage declines for reasons nobody in the building can name.

Declining is valuable precisely because it costs something. A system that sometimes says no has demonstrated that its yes was selected rather than automatic, and every subsequent answer inherits a little of that credibility. This is not a soft point about trust. It is the difference between a channel a buyer will act on and a channel a buyer will verify elsewhere before acting, and verifying elsewhere means leaving.

There is a commercial edge to it as well. A wrong answer does not end at the wrong answer. It gets screenshotted, quoted back to you in a later conversation, and occasionally treated as a statement made by your business, because from the visitor's side it was made on your site in your name. What follows from that depends on wording, on jurisdiction and on what was promised, and it is a question for your own adviser. A refusal, by contrast, costs one mildly disappointed visitor who now knows where to go instead.

The counter-case is over-refusal, and it is a genuine failure rather than a rhetorical concession. An assistant that declines constantly teaches visitors it is decorative, and they stop opening it just as surely as if it had misled them. The two failures are symmetrical and both are invisible without measurement. The test that separates them is the declared scope: refusals on questions outside it are the system working, and refusals on questions inside it are a coverage problem wearing the costume of caution. Which is the practical reason the scope has to be written down first. Until it exists, a refusal rate is a number with no interpretation attached to it.

Reviewing answers after launch, and who owns the cadence

None of the preceding survives launch without a review habit, and the review is not quality assurance in the software sense. It is reading. A person sits down with a sample of real answers and the scope document, and marks each answer against it. No tooling is required for this to be worth doing, and the absence of tooling is a poor excuse for not doing it.

Four marks are enough. Correct and in scope. Refused, and correctly refused. Refused, but it should have been answered, which is a coverage gap and points either at the first failure or at a page that does not exist yet. Answered, but it should have been refused or qualified, which is the expensive column and where the third, fourth and fifth failures all land. Counting that fourth column over time is the only honest measure of whether the controls are holding.

Cadence should start weekly, because in the first month you do not yet know what people ask, and the shape of the question set is itself the finding. It can drop to monthly once the pattern stabilises. What actually determines the right interval is not the calendar but change: every time you publish a page that alters a qualification, retire a service without retiring its page, or change a price, you have manufactured the conditions for a contaminated or over-generalised answer. Publishing and reviewing belong on the same schedule for that reason.

Ownership needs a person, not a team. Two rules make the choice easier. The reviewer should know what is true about the business, which usually means somebody who deals with customers rather than somebody who deals with systems. And it should not be the person who set the assistant up, because they will read the answers as output from a configuration they understand rather than as claims made in public by the company. Whoever answers the phone is often a better reviewer than whoever chose the supplier.

Finally, the review needs teeth, meaning a route from a marked answer to an actual change and a person authorised to make it. A review that produces observations nobody can act on degrades within two cycles into a log of known-wrong answers that everyone has stopped reading. If the only available remedy is to raise a ticket with a supplier who responds in a month, say so out loud at the point of purchase, because that response time is the real cadence of your assistant regardless of what the review calendar claims.

Why does my chatbot invent answers when the information is on the site?

Because finding the material and writing the answer are separate steps, and either can fail while the pages stay perfectly correct. Six failures account for almost all of it: retrieval missed a source that was there, usually through vocabulary mismatch or an unreadable format; content from an unrelated page was retrieved instead, which shows up as a real detail about the wrong thing; two sources were combined into a claim neither made, which reads as more complete than any page you have; a qualification was dropped in summary; a question with no source at all was answered from general knowledge, so the answer contains nothing specific to you; or instructions embedded in retrieved text were followed. Match the symptom to the failure before changing anything.

Can an assistant be stopped from answering outside its scope?

Only if a scope was declared in writing and a refusal path was designed, because an assistant with no stated boundary has nothing to decline against. A written scope names the topics it must answer and the page that is the authority for each, the topics it must never answer stated specifically rather than by category, the topics where quoting is required because a qualification is load-bearing, and the wording used at the boundary. A good refusal states plainly that it cannot answer, gives the true reason, offers a route that genuinely exists, and preserves what the visitor typed. That costs one mildly disappointed visitor. A confident wrong answer costs considerably more, and it repeats.

Pass it on

Share this signal

Send it to someone working through the same question.