AI & automation
Preparing Your Website Content Before an AI Assistant Answers for You
The setup instruction is usually one line: point it at the site. The governing rule is stricter. An assistant should never be able to say something the site does not say, and that is a readiness audit for the content underneath.

The instruction that ships with most website assistants is a single line: point it at your site. Enter a domain, wait for a crawl to finish, confirm the pages that came back. That instruction is accurate as a description of the setup work and misleading as a description of the decision, because it quietly assumes the site is already fit to be answered from. Most sites are not, and the gap is not a technical one.
What follows is a readiness audit for the content underneath: a sequence of conditions the source material has to meet, each stated with the specific failure it prevents. It is deliberately confined to the content. If the pages are correct and the assistant still answers badly, the fault lies in how the answer was produced rather than in what it was produced from, and that is a separate diagnosis.
The governing rule: it may not say what the site does not say
One rule governs everything below, and it is worth stating before any of the detail. An assistant should never be able to say something the site does not say. Not something the site contradicts, and not merely something the site fails to mention: anything the assistant asserts should be traceable back to a sentence somebody wrote, reviewed and published.
The rule is stricter than it first sounds, and it cuts in two directions. Read forwards, it constrains output: every answer has a source, and where there is no source there is no answer. Read backwards, it becomes a specification for the content: if there is a question the assistant must be able to handle, a page has to exist that settles it. Most readiness work turns out to be the backwards reading. You are not preparing content for a machine so much as discovering which of your own positions were never written down.
The reason to hold the rule this tightly is that an assistant is a publishing surface. Whatever it says arrives in the visitor's browser inside your site, in your typography, under your name. There is no useful sense in which a statement produced by an assistant is less yours than a statement in a paragraph on a services page. It is faster, more specific to the person reading it, and delivered at the moment they asked, which makes it more consequential rather than less. Treating it as a lower-stakes channel than the pages around it is the mistake underneath most of the failures in this piece.
An assistant is a publishing surface. Whatever it says arrives in your typography, under your name, at the moment somebody asked for it.
The obvious objection is that the rule makes the assistant less useful. It does. Held properly, it produces something that declines more often than the demonstration you were shown, because the demonstration was built on a site whose gaps had already been filled in. An assistant that says it does not have that information and offers a route to somebody who does is a worse demonstration and a better business decision. The cost is real, though, and worth admitting: you will find yourself writing pages you had not planned to write, and some of them will be pages you had been avoiding for reasons.
For the rule to be wrong, you would have to be willing to accept an assistant that composes positions on your behalf, and to own whatever it composes as a statement of what your business does. Some organisations are genuinely relaxed about that, usually because the subject matter is low-stakes and the cost of an invented answer is an apology. If that describes you, the audit below is optional. If a wrong answer about scope, process or obligation would cost you money or credibility, it is not, and the readiness question is worth settling before a tool is chosen rather than after.
One boundary, because the two families of failure look identical from the outside. This piece owns every failure that originates in the source material: the answer nobody ever settled, the page that is correct and stale, the pages that disagree with one another, the sentence that cannot survive being quoted alone, the content that could never be reached. Failures that originate at answer time, such as a retrieval step selecting the wrong passage or a copy of the site that is older than the site, belong to the companion piece on why a chatbot gives wrong answers when the pages are correct.
Is the answer written down anywhere at all
The first condition is the one most often skipped, because it feels too obvious to check. For each question you expect the assistant to handle, is there a page that answers it? Not a page adjacent to it, and not a page from which the answer could be inferred by a sympathetic reader who already knows the business. A page that answers it.
The test is mechanical and slightly humbling. Take the questions you actually receive, from the inbox, the contact form, the phone and the first ten minutes of every sales conversation, and against each one write a URL and the sentence on that page that answers it. Not the page you think ought to answer it. The sentence. Questions where you can answer fluently out loud but cannot produce a sentence are the readiness backlog, and there are usually more of them than anyone expects.
Three distinct situations hide inside a failed pointing test, and they need different work. The first is the answer written for a different question. A services page describes what the work involves and reads as though it covers scope, but the question people ask is whether a particular category of problem is in or out. The page asserts a shape; the question needs a boundary. Adding the boundary is a small edit and a genuine change of content.
The second is the answer nobody has ever settled. Somebody asks whether you take on work of a certain size, or what happens to a project when the person leading it is unavailable, and the honest position is that it depends and has never been decided. Writing it down forces a decision, which is precisely why it never got written. This is the most valuable part of the audit and the least comfortable, because the assistant is not creating the ambiguity. It is only the first thing that will state the ambiguity to a stranger, in a sentence, without anybody in the room to soften it.
The third is the answer that is deliberately unwritten. Commercial terms are the usual case: figures are absent from the site because publishing them is a strategic choice you have not made. Readiness here does not mean publishing them. It means deciding what should be said instead and writing that down as its own settled answer, because an unanswered question is not a neutral state. Leave nothing on the subject and an automated reader will reach for whatever is nearest to it, which might be a value claim, a phrase from a capability page, or a line from a comparison table written for a different purpose entirely.
It helps to be systematic rather than exhaustive. Twenty questions is enough to begin with; the tail is long and mostly harmless. Rank them by the cost of getting them wrong rather than by how often they arrive, because the question asked rarely about what you refuse to do is more dangerous unanswered than the question asked constantly about where you are based.
One source of truth per fact, or three pages that disagree
Once the answers exist, the next condition is that each fact exists once. Duplication is the normal state of a site that has been added to steadily: the same claim about what the business does appears on the home page, on a services page, in a footer block, inside an old editorial piece, and in the copy of a landing page built for a campaign that has ended. Each of those carries its own decay schedule. Update the services page and four other statements quietly become historical without anybody touching them.
For a reader, disagreement across pages is a mild irritant, usually resolved by trusting whichever page looks most current. For an automated reader it is a different problem. Given several passages that disagree, something has to choose, and you have no reliable way to predict which one wins or to notice that a choice was made. The answer will be delivered with exactly the composure of a correct one. Confidence is a property of the format, not of the evidence behind it.
The audit is a search exercise. List the ten facts that would embarrass you if stated wrongly: what the business does, what it explicitly does not do, where it is, how to reach it, what happens after somebody makes contact, what is included and what is extra, what the process actually consists of. Then search the site for each one, using a site search, a repository search, or the crudest available approach of a search-engine site query, and count the statements. Anything above one is a liability with a timer on it. A single source of truth in website content is not an architectural aspiration; it is the only version of the site that can be kept true.
- Make one page canonical and have every other mention link to it rather than restate it. Cheapest to maintain, and it costs you some copy that people were fond of.
- Hold the fact once in the publishing system and render it wherever it appears, so that a single edit propagates. Only available if your setup genuinely supports it, which is worth confirming rather than assuming.
- Accept the duplication and record it, so that changing the fact becomes a checklist rather than a memory exercise. This is the weakest option and the one most sites end up with, and it fails silently the moment the person holding the checklist moves on.
Not every overlap is a contradiction waiting to happen, and this should not turn into an argument against summaries. A short statement on a landing page and a detailed treatment on a deeper page can coexist indefinitely, provided they cannot drift into disagreement. The test is drift, not overlap. If the detailed page changes, does the summary become false? If so, it is a duplicate wearing a summary's clothes. If the summary sits a level of abstraction above the detail, describing a category rather than a content, it is a summary and it is fine.
Contradictions also live outside body copy, where nobody thinks to look. A navigation label naming a service the services page no longer describes. A dropdown on the contact form offering categories of work that do not match the categories on the site. A footer address updated in one place only. Structured data describing something the visible page has since stopped saying. Each of those is a statement, and an automated reader will treat it as one.
The page that is correct and stale
A stale page is not a wrong page in the ordinary sense. Every sentence on it was true when it was written and nobody has edited it since. It reads well, it is internally consistent, and it passes proofreading without a flicker, because proofreading checks whether the prose is sound rather than whether the world still matches it.
Staleness is harder to find than error for a structural reason: errors get reported and staleness does not. Somebody eventually notices a broken claim and writes in about it. Nobody writes in about a page describing a process you abandoned, because the page is coherent and the reader has no way of knowing it is out of date. It simply goes on quietly answering a question with an old answer, to whoever happens to find it.
An assistant changes the consequence without changing the cause. A stale page deep in the site may attract almost no traffic; it sits in the sitemap and is read by very few people, which is why nobody has felt the cost of it. Point an assistant at the site and that page's stale sentence acquires an audience, because it will be surfaced to whoever asks the matching question regardless of where the page sits or how rarely it is visited. Low-traffic pages stop being low-consequence. This is the most under-appreciated effect of adding an assistant to an established site.
So when the complaint is that the chatbot is showing old information, there are two candidates, and they are diagnosed in a fixed order. Open the page yourself and read it. If the page says the old thing, the fault is here and it is a content problem, fixed by editing the page. If the page says the current thing and the assistant says the old one, the fault is at answer time, in whatever copy of the site the answer was produced from, and it is covered in the companion piece. Running that diagnosis in the wrong order wastes a great deal of effort in the wrong system.
The recurring shapes are worth naming, because they are much the same everywhere. Descriptions of a way of working that has since changed. Lists presented as complete that quietly stopped being complete. Contact routes that still function but are no longer monitored by anybody. Pages describing what happens next after an enquiry, written for a process that had a different number of steps. And anything written to be current, in language that does not date itself, so that no reader can tell how old it is.
Dates and owners, so staleness becomes detectable
Staleness cannot be found by reading, because stale pages read correctly. It can only be found through metadata, which means two fields per page: when this was last reviewed, and who is responsible for reviewing it. Neither needs to be sophisticated, and both need to exist.
The distinction between last modified and last reviewed matters more than it appears. Last modified is a fact about a file, and it moves when somebody fixes a typo, swaps an image or reformats a list; a page modified recently can be substantively ancient. Last reviewed is a human assertion, meaning somebody read the page in full and confirmed it is still true. Only the second is evidence. If your system records only the first, treat it as a timestamp rather than as a claim, and do not let it reassure you.
Ownership should be a role rather than only a name, because names leave and roles persist. An unowned page has a theoretical review interval, which is to say none at all. The interval itself should follow volatility rather than a single site-wide policy: what the business does not do is likely to change slowly, whereas what happens after somebody makes contact changes whenever the process changes, which is more often than anyone plans for. Applying one review cadence to every page produces a schedule that is simultaneously too frequent for most of the site and too slow for the pages that matter.
Whether to display the date on the page is a genuine trade-off with reasonable arguments either way. A visible date is honest, helps a reader judge the material, and is close to obligatory on anything dated by nature, such as editorial pieces, notices and anything describing a state of affairs. On an evergreen page it can misfire, because a visible date invites a reader to discount a page that is entirely current merely because it has not needed to change. A defensible split is to publish dates on anything time-bound and hold review dates internally on the rest, provided internal means recorded somewhere rather than assumed.
If none of this exists today, the first pass requires no tooling at all. A list of URLs with two extra columns, filled in by the people who own the material, converts an invisible problem into a visible one. It is unglamorous, it will start going out of date the week after it is finished, and it is still the whole of the fix. The alternative is a site whose accuracy depends on somebody happening to remember.
Prose that asserts, and prose that settles
The next condition concerns how sentences are built rather than whether they are present. A great deal of website copy asserts a quality without settling a question. It says the team is responsive, that the approach is collaborative, that the work is thorough. None of those sentences answers anything. They are claims about character, written to create an impression rather than to resolve an enquiry, and they were perfectly good at the job they were given.
Prose that settles is different in a specific and testable way: it answers a question a reader could have asked, in a sentence that remains true and complete when lifted out of the page. Responsive settles nothing. A sentence describing what actually happens to an enquiry when it arrives settles a great deal, and it can be quoted on its own without becoming misleading. That property, survivability outside its own context, is what makes a sentence usable as an answer.
The reason this matters for an assistant rather than only for the reader is that assertion leaves nothing to stand on. Faced with a question about turnaround and a page offering only a claim of responsiveness, an automated answer has two available moves. It can decline, which is correct and unsatisfying. Or it can convert the evocation into something concrete-sounding, at which point a mood word chosen by a copywriter has become a commitment that appears to have been made by the business. The second failure is the expensive one, and its root cause sits in the source copy rather than anywhere downstream of it.
Evocative prose gives an assistant nothing to stand on, so it stands on whatever is nearest.
Two habits of ordinary web writing break the quotable-sentence test almost invariably. The first is reference back: it, this, the above, as mentioned, which is why. These work on a page and fail the moment a sentence travels alone, and there is no way to know in advance which sentence will travel. The second is meaning carried by structure rather than by the sentence, such as a heading reading Not included followed by a list of items that are individually harmless and collectively catastrophic once the heading has been left behind. Wherever a negation, an exception or a condition is expressed by layout, it should also be expressed inside the sentence.
- Could this sentence be read aloud to somebody who has not seen the page, and remain both accurate and complete?
- Does it depend on the sentence before it, the heading above it, or the list it sits inside?
- Does it settle a question somebody actually asks, or assert a quality nobody asked about?
- If it states an exception, does the exception survive being separated from the thing it is an exception to?
- If it were quoted back to you by a prospective client who had read nothing else, would you stand behind it as written?
This is also the honest version of the advice to structure FAQ content for AI. Question-shaped headings with self-contained answers underneath are not a trick played on a machine. They are simply the format that survives extraction, because each unit was written to stand alone in the first place. The common mistake is to write a sequence of answers that depend on one another, so that the fourth is unintelligible without the second. That is a document with question-shaped headings, not a set of answers, and it will come apart exactly where it is cut.
The limit of the argument should be stated plainly, because it is easy to over-apply. This is not a case for rewriting a site as a manual, or for flattening voice into specification. A brand page is doing work a fact sheet cannot do, and a site composed entirely of settled statements would be accurate and dead. The claim is narrower than that: anything you want answered has to be settled somewhere, once, in prose that can be quoted. Everything else can carry on being persuasive, and should.
Content the assistant cannot reach
The last content condition is reachability, and it produces the most confusing symptoms of any of them. The material can be written, single-sourced, current, owned, dated and beautifully quotable, and still be invisible to anything reading the site automatically. Whatever the mechanism doing the reading, the constraint belongs to a familiar class: something requests a document over HTTP and works with what comes back. If the sentence is not in what came back, it is not available, however plainly it appears on screen.
Content that only exists after JavaScript has run
The common case is content the browser assembles rather than receives. Interactive components are the usual culprits, and they are not all equivalent. An accordion whose panels are present in the served markup and merely hidden by styling is generally fine, because the text arrived with the document. An accordion that requests its panel content when somebody clicks the heading is not, because until that click the text was never sent. To a person using the page, the two are indistinguishable, which is why this survives review by everybody including the people who built it.
The test costs a minute and settles the question. Fetch the URL as plain HTTP, using view source rather than inspect element, or a command-line request if you have one to hand, and search the response for the sentence you care about. Inspect element shows the document after scripts have run, which is precisely the thing that will mislead you here. If the sentence is not in the raw response, treat it as unreachable until somebody demonstrates otherwise.
The same test catches the wider family: pages rendered entirely on the client, where the served document is little more than a shell; content behind a login; content revealed only after a consent interaction; content that varies by region or by session. There is a useful overlap with accessibility here, and it is not a coincidence. Content that exists only after a script has run tends to be the same content that behaves poorly for assistive technology and for keyboard users. WCAG 2.2 sets out success criteria across exactly this territory, and 2.1.1 Keyboard requires that functionality be operable through a keyboard interface. Anything failing the raw-response test is worth checking against the guidelines while you are already looking at it, because you are usually seeing one defect wearing two costumes.
Facts trapped in PDFs and attachments
A surprising proportion of the most carefully settled content on a site is not on the site. It is in attachments: capability documents, specification sheets, policy statements, terms, the considered document somebody wrote once and has maintained ever since. These are often the best-written and most precise material an organisation has, and they are among the least usable as a source of answers.
Three problems compound. Extraction is unreliable, because a PDF describes where marks go on a page rather than how a document is structured, so multi-column layouts, tables and footnotes come apart in ways that are hard to predict, and a scanned document with no text layer is an image of words rather than words. Versioning is usually worse than it is for pages, since a PDF tends to be replaced wholesale, rarely carries a review record, and is frequently the oldest artefact on the entire site. And even a clean extraction cannot be cited usefully, because the provenance of the answer is a long file rather than a paragraph a reader can be sent to and shown.
Google Search Central documents that PDF files can be indexed, and that is worth separating from the question here. Being indexable as a document is not the same as being usable as the source of a single sentence. The practical rule is to promote any fact you want stated into an HTML page, and to keep the PDF for the reasons you had it in the first place: something to send, something to print, something to sign. The two are not in competition, but only one of them is a source.
One quick check before assuming a file is readable at all. Open it and try to select a paragraph of text. If nothing selects, it is an image, and every fact inside it is invisible to everything except a person's eyes.
Scope: the pages to include, and the pages to exclude explicitly
Everything above assumes a decision most deployments never make explicitly: which pages are in scope. The default of the whole site is a decision arrived at by not deciding, and it is almost never the right one, because a site contains several categories of page that should not be answered from at all.
Editorial is the first of them. An article arguing a position is not a statement of current policy, and an old article may not even be a statement of a current position. It was written to be read as an argument, by a particular author, at a particular time, with the reader able to see all of those things at once. Extracted as a sentence and delivered as an answer, it becomes what the business says. That applies to the piece you are reading now as much as to anything else.
Legal pages are the most argued-over case, and the argument is a real one. Excluding terms and privacy means the assistant cannot help with the questions people ask most literally. Including them means a system will paraphrase the one category of text on the site that was written to be read exactly as written, where a compression is a change of meaning rather than a convenience. The resolution most people land on is to let the assistant identify the relevant page and send the visitor to it without restating what it says, so the answer is a pointer rather than a summary. That is a weaker answer, and it is defensible in a way the alternative is not.
- Job posts, notices and anything carrying an expiry date that nobody actually set.
- Campaign landing pages, which carry claims tuned for one segment and were never meant to describe the business in general.
- Staging, preview and forgotten subdomains that are reachable only because nobody ever checked whether they were.
- Content you did not write, such as submissions, comments or imported listings, which becomes content you appear to endorse the moment it is quoted back as an answer.
- Near-duplicates produced by the system rather than by an author: paginated archives, tag pages, print variants and parameter versions of the same page, all of which multiply the weight of a single phrasing without anybody deciding that it should carry more weight.
The exclusion list has to be written down and owned, and this is the part that decays fastest. A list agreed once at deployment describes the site as it stood that week. Every page added afterwards is included by default unless somebody decides otherwise, which means the decision quietly reverts to the whole site. The fix is procedural rather than technical: whoever publishes a page states whether it is in scope, in the same breath as everything else they state about it, and the statement is recorded where the next person will find it.
Scope is worth revisiting from the other direction too, at least once. Ask which questions the assistant cannot answer because you excluded the only page that settles them, and whether each exclusion was a judgement or an accident. Both mistakes are common, and only one of them ever gets reported by a visitor.
Preparing content for an assistant is mostly just fixing the site
There is a conclusion available here that is more useful than the obvious one. Very little of this audit is work you are doing for an assistant. It is work you already owed the site, and the assistant is simply the thing that made the bill visible.
Take the conditions one at a time. A question that had never been settled anywhere was already costing you: it was being answered inconsistently over email, differently by different people, and not at all to the visitors who never bothered to ask. Facts held once rather than five times are cheaper to change and stop contradicting each other for readers as well as for machines. Review dates and owners are the only mechanism by which anybody finds a stale page, whatever is doing the reading. Sentences that survive being quoted alone are the sentences that work in a search result, in a link preview, and for the substantial proportion of readers who skim rather than read. Content present in the served document rather than assembled afterwards behaves better for assistive technology and for crawlers. And an explicit scope list is a content inventory, which most organisations do not have and every one of them benefits from.
None of that means doing all of it before doing anything. An audit like this expands to fill whatever time it is given, and a complete content review of a mature site is a project rather than a task. The version that works is bounded: the twenty questions you actually receive, checked against the conditions above, with the rest of the site left alone until something makes it urgent. A narrow assistant answering ten questions from settled, current, reachable content is worth considerably more than a broad one answering fifty from material nobody has read since it was published.
There is one case where the verdict changes, and it is worth stating because it is the reasonable objection to the whole argument. If you are deploying an assistant that answers from a separately maintained knowledge base rather than from the site, then the site's condition matters less directly, and chatbot knowledge base preparation becomes a distinct exercise with its own schedule. It is worth noticing what has been bought, though. Two sources of truth, maintained by different people at different times, describing the same business, is the contradiction problem from earlier in this piece with an extra system inside it and no obvious owner for the disagreement when it arrives.
A disclosure is owed here rather than buried at the end. Zealsync develops Flidu, an AI assistant and contact widget for websites, and the product page sets out what it is. Nothing above describes how Flidu or any other assistant ingests content, because that is not the argument being made. The argument is that the readiness of the source material is the buyer's problem regardless of which assistant is eventually chosen, and it is cheaper to establish before a decision than to discover after one.
It is worth being clear about what happens if the answers stay unwritten, because it is not dramatic. Visitors will go on asking the questions, staff will go on answering them individually, and the assistant will do what it can with prose that was never meant to bear the weight. The failure will not present as an incident. It will present as an assistant that is somehow disappointing, and as a series of small corrections that nobody connects to one another. The alternative starts with a list of questions and a column for URLs, and it is available to anyone willing to be honest about which of those columns is empty. What happens after the answer, once the visitor has one, is a separate design problem, and it only becomes worth solving once the answers themselves are sound.
Relevant Zealsync pages


