AI & automation
How to Review AI-Generated Work Before It Reaches a Client, and Where to Spend the Effort
Always verify is advice, not a procedure. The failures specific to machine drafting, the check that catches each one, and why review effort should scale with severity and irreversibility rather than spread evenly.

A review instruction is only worth having if it tells you what to do differently on a Tuesday afternoon with forty drafted findings in front of you and two hours before the deliverable goes out. Always verify does not do that. It states a standard nobody can fail to endorse and nobody can act on, because it never says what verification consists of, which findings receive it, or what to do when it comes back ambiguous.
Taken literally, it means re-deriving every finding from the original evidence, which costs roughly what drafting it from scratch would have cost and removes most of the reason for working this way. Taken as it is usually taken, it means reading each finding once, nodding at the ones that sound right, and sending the file.
Always verify is advice, not a procedure
The second reading is the dangerous one, because machine-drafted findings are written to sound right. Fluency is what the drafting process is good at, and fluency is exactly the signal a tired reviewer reaches for as a proxy for correctness. This is the specific thing that has changed. A finding written by a junior analyst who half-understood the evidence usually reads like the work of somebody who half-understood the evidence: the hedges sit in the wrong places, the sentence wanders, the recommendation does not quite follow from what precedes it. That texture is a tell, and reviewers have spent careers learning to read it. The same error, drafted by a model, reads like the confident summary of somebody who understood the material completely. The tell is gone, and nothing has replaced it.
So a procedure has to supply structurally what the reviewer used to get for free. Three parts: a named list of the failure modes specific to this way of working, a check for each one that a reviewer can actually run against the evidence, and an allocation rule that decides where a finite review budget goes. The rest of this is those three parts.
It assumes you already know what a finding is made of, and does not re-derive it. The criterion, the evidence, the location, the severity and the recommendation are treated here as fields that exist and can be checked one at a time; the case for defining them that way is made in the anatomy of a finding.
Fluency is what the drafting process is good at, and fluency is what a tired reviewer reaches for as a proxy for correctness.
Three failures specific to machine drafting
Ordinary quality review catches ordinary problems: typos, broken structure, a recommendation that contradicts the finding two paragraphs above it. Keep doing all of that. What follows are three failures that ordinary review does not catch, because in each case the output is well-formed, internally consistent, and wrong only against the source. You cannot find any of them by reading the deliverable carefully. You find them by reading the deliverable against something else.
A superseded standard applied confidently
Standards move, and a draft can cite them with more confidence than the citation deserves. The clearest published example is WCAG. The 2.2 recommendation adds success criteria that did not exist in 2.1, among them 2.4.11 Focus Not Obscured (Minimum) and 2.5.8 Target Size (Minimum), and it removes 4.1.1 Parsing, which was present in 2.0 and 2.1. A finding that raises a parsing failure as a WCAG 2.2 issue is citing a criterion the current recommendation no longer contains. A finding that names one of the new criteria but attaches the wrong conformance level is wrong in a quieter way that survives a proofread completely intact. Search documentation behaves the same way: guidance that was accurate two revisions ago is still fluent, still specific, and no longer what the published documentation says.
The check is unglamorous and takes seconds. Open the published standard, find the criterion by number and by name, and confirm three things: that it exists in the version cited, that the level matches, and that the wording the finding leans on is the wording actually there. The tell for this class is confident specificity with nothing quoted. A criterion number and a criterion name, no source text, is the shape a superseded citation usually takes, because the number and the name are the parts that are easiest to reproduce from memory of an older version.
An inference from an image presented as an observation
The evidence set for an audit is heterogeneous. Screenshots, live URLs, exported spreadsheets, PDFs, design files, campaign exports. Each artefact type can show some things and cannot show others, and the boundary is sharper than it looks. A static screenshot shows what was painted at one moment. It does not show focus order, keyboard reachability, what a screen reader announces, what happens on submit, what happens at another viewport, or whether the state shown is the default state or one the person capturing it produced by accident.
A drafting process handed a screenshot and asked for findings will produce sentences of the form the control cannot be reached by keyboard. That is a reasonable inference from the visual evidence. It is not an observation, and the deliverable presents it as one. Nothing in the wording flags the difference, because the wording is equally fluent in both cases, and the reviewer who reads only the deliverable has no way to tell which of the two they are looking at.
The check is a two-column pass. For each finding, name the artefact it rests on, then ask whether that artefact type is capable of showing what is claimed. Where the answer is no, the finding does not get deleted. It gets demoted to a hypothesis with the test that would settle it written next to it. Perhaps half of those tests are quick enough to run during the review itself, and running one converts a liability into the strongest finding in the report, because it now has a named test and a result behind it. The other half tell the client honestly what was not testable with what they supplied, which is a more useful sentence than a confident guess and a considerably safer one.
True, verifiable, and irrelevant to the objective
The third is the hardest to catch, because nothing in the finding is false. Drafting fills the space it is given. Ask for findings and you receive findings, including ones that are correct, properly evidenced, and have no bearing whatsoever on the question the work was commissioned to answer. Hypothetically: an audit commissioned to explain why a checkout is abandoned comes back with a dozen well-evidenced observations about heading hierarchy on marketing pages. Every one of them is true. Every one is verifiable against the evidence supplied. None of them is the answer, and a reviewer working finding by finding will pass all twelve, because finding by finding is the wrong altitude for this failure.
The check has to happen before you read the draft, not during it. Write the objective in one sentence and put it beside the screen. Then for each finding ask a single question: if this is fixed and nothing else changes, what happens to that objective? Answers that cannot be given without a hedge belong in an appendix or out of the report altogether. This is the only one of the three checks that cannot be handed to whoever is fastest, because it requires knowing why the client commissioned the work, which is rarely written down anywhere in the brief and is often not quite what the brief says.
Severity assigned by pattern match, and re-deriving it properly
Severity is the field most often wrong and least often checked, and the reason is that it is short. A single word carries no visible working, so there is nothing for a reviewer to disagree with. Critical sits on the page looking exactly as authoritative as moderate, and both look like conclusions rather than claims. Drafting assigns severity by resemblance: findings that look like other findings labelled high acquire the label high, and the resemblance doing the work is textual rather than causal.
Re-deriving severity properly means throwing the word away and answering the underlying questions from the evidence, in order. What is the consequence if this is never fixed. Who does that consequence fall on, and are they the client or the client's users, because those are different arguments with different remedies. How many of them, and is it everyone or a subset defined by device, assistive technology or route. Is there a workaround, and does using it require knowing something the affected person has no way to know. Does the consequence compound over time or stay flat. Only then read off the band. If the band that comes back differs from the drafted one, you have learned something about the whole batch, not only about the finding in front of you.
A separate discipline is not to let the severity band decide the running order of the report. Severity describes consequence; sequencing is a different judgement involving effort, dependency and what the client can realistically schedule, and the case for holding the two apart is made in severity is not priority. This piece takes that separation as settled and uses severity only as one input to how much review a finding earns.
The honest limit on this argument: pattern-matched severity is right more often than not, and a reviewer who re-derives every band will find most of them already correct. That is not a defence of the method, because its errors are not distributed randomly. They cluster on the findings that do not resemble anything, which are precisely the unusual, context-dependent ones the client is paying for. The median finding being fine is no comfort when the failures sit in the tail, and the tail is the only part of the deliverable that could not have been written without seeing the evidence.
Sampling: which findings get re-derived and which get a spot check
Two different activities are both called reviewing, and conflating them is why review budgets disappear without anyone noticing. Re-derivation means ignoring the finding's own text, returning to the evidence, independently answering whether the criterion is breached and at what consequence, and only then comparing your answer to the draft. It costs a meaningful fraction of what drafting cost. A spot check means reading the finding, confirming the cited evidence exists, confirming it says what the finding says it says, and confirming the criterion is named correctly. It costs a minute or two. Both are legitimate. Performing the second while describing it as the first is where reports go wrong.
These always get full re-derivation, without exception and without negotiation when the deadline tightens:
- Every finding in the top severity band, on the reasoning that the band itself is one of the things being checked.
- Every finding whose evidence is an inference rather than a direct observation, identified by the two-column pass above.
- Every finding that will be quoted in the executive summary, in the covering note, or read aloud in a meeting.
- Every finding recommending something irreversible or expensive: a migration, a rewrite, a contract change, a public correction.
- Every finding that other findings depend on, because an error there propagates silently into everything built on top of it.
Everything else receives a spot check, plus a sample drawn for full re-derivation. Two rules make the sample worth drawing at all. Choose it before reading the batch, because a sample chosen afterwards is a sample of what already worried you and confirms only what you already suspected. And fix the proportion in advance, honestly against your real capacity, rather than deciding at the end that whatever you managed to get through was the plan all along.
Then the escalation rule, which is what gives the sample teeth. A sampled finding that fails re-derivation is not one bad finding to be quietly corrected in place. It is evidence about the batch it was drawn from. One failure in the sample means every finding of that class gets re-derived, all of them, whatever that does to the schedule. This is deliberately expensive, and the expense is the mechanism: a sample with no consequence attached becomes a ritual, and reviewers unconsciously draw the findings they expect to pass. Knowing that a single failure costs the afternoon is the only thing that keeps the draw honest.
A sample chosen after reading the batch is a sample of what already worried you.
Scaling review effort by severity and irreversibility
Review effort spread evenly across a deliverable is review effort spent wrongly, because the cost of an error is not evenly distributed across it. Two axes decide the allocation. The first is severity in the sense above: how bad it is if this particular sentence turns out to be wrong. The second is irreversibility: what it costs to correct the error at the moment it is discovered, given how far the sentence has travelled by then.
Irreversibility is the axis people forget, and it moves much faster than severity does. A wrong finding caught in the working draft costs an edit. The same finding caught after the file has been sent costs a correction the client will remember longer than they remember the twenty findings that were right. The same finding caught after the client has repeated it to their own board costs a relationship, and no amount of being correct about the rest of the report buys that back. The finding did not change between those three moments. Only its position did, and its position is the thing you control by choosing when to review it.
Worked through, this produces allocations that look wrong against a severity column alone. A low-severity finding on the cover page earns more review than a moderate one on page forty, because the cover page will be read aloud, screenshotted and forwarded, while the moderate one will be read once by somebody who already agrees with it. A finding that recommends leaving something alone earns more review than one recommending a change, because a client will scrutinise the recommendation to change and will simply believe the recommendation to leave, which means the error in the second one never surfaces until it has compounded. Effort follows exposure and consequence together. The severity label is an input to that, not the answer to it.
The counter-case is real and worth stating plainly. If you produce three findings a month, all of this is overhead, and a uniform standard applied to everything is both cheaper and better than a protocol with moving parts. The apparatus earns its keep at volume, which is precisely where a uniform standard quietly becomes a uniformly low one, because the only way to hold a constant level of attention against a rising count is to lower it everywhere at once and not notice you have done it.
The same rule applied to the process, not just the deliverable
The allocation rule generalises beyond findings. Which tasks should not be automated has no answer in general, because the answer depends entirely on what the task touches when it goes wrong. It has three specific answers, and they are the same two axes lifted off the deliverable and applied to the workflow that produces it. A human in the loop review process that puts a person at every step is a process nobody will follow by the third week; the useful version names the steps where the person is load-bearing and lets the rest run.
Steps that are irreversible
Sending, publishing, deleting, merging to a default branch, triggering a deployment, invoicing, anything that puts a file in a client's hands. Automate the preparation of these steps as far as you like, and keep the commit itself with a person. The asymmetry is not close enough to be worth arguing about. A held step costs a delay measured in hours. A sent one costs a retraction, and those two are not comparable quantities on any axis a client cares about.
Reversibility is a property of the world, not of the system. Version control genuinely reverses a change to files: a revert restores the previous content and the history records both states, which is a real guarantee and worth having. It does not reverse the deployment that already served the wrong page, the notification that already went out, or the fact that somebody read it and formed a view. A step whose technical undo is trivial can be entirely irreversible in the only sense that matters to the engagement, and that is the sense this rule is about.
Steps that are externally exposed
Anything that leaves the building carrying your name. Exposure is independent of severity, which is what makes it so easy to under-resource. A trivially wrong sentence in a covering email does more damage than a genuine analytical error on page forty, because the covering email is read by the person who signs off the engagement and page forty is read by nobody until something breaks. The generated summary, the automated status note, the templated follow-up: these are low-effort artefacts with high exposure, and they attract the least review precisely because producing them took no effort. Effort spent producing is a terrible predictor of effort that should be spent checking, and it is the predictor almost every team uses by default.
Judgement a rule cannot capture
Which three findings to lead with. What the client can realistically act on this quarter, given what you know about their year that is written down nowhere. Whether the tone is right for the person who commissioned the work and also for the person who will be blamed by it, who are frequently not the same person and frequently sitting in the same room when it is presented. None of these has a stable rule behind it, and a process that produces a confident answer to any of them has produced a guess written in the register of a decision.
The sharpest case in this class is absence. Rules operate on what is present in the evidence; noticing that something is missing requires a model of what should have been there. A reviewer who knows the client's estate looks at eleven supplied artefacts and registers that the twelfth, the one that would have explained the whole pattern, was never supplied and was probably never mentioned. That noticing is often the entire value of the review, and there is no field in the draft where it could have appeared, because the draft can only be about what it was given.
When the reviewer disagrees and cannot prove it
This case is common and almost never has a procedure attached to it. The reviewer reads a finding. The evidence supports it. The criterion is named correctly. The severity survives re-derivation. And the reviewer still believes it is wrong, without being able to say why in a sentence that would survive being challenged in front of the client. Usually the reason is that the reviewer holds context the evidence set does not contain, and has not yet worked out which piece of that context is doing the work.
Three moves are legitimate. The first is to get the evidence: name the specific test that would settle it, then run it or ask for it. A surprising number of unprovable disagreements become provable within twenty minutes once somebody is forced to write down what test would decide the question. The second is to demote and disclose: restate the finding at the confidence the evidence actually supports, with its limits written into the report where the client can read them rather than into a private note. A stated uncertainty is a professional act and clients respond to it better than suppliers expect; a hidden one is the thing they do not forgive. The third is to hold the finding out of this delivery and say plainly that it is unresolved, which is always available when the disagreement will not close before the deadline.
Two moves are not legitimate, and both are the natural thing to do under time pressure. Deleting the finding quietly because you do not like it, which destroys information and leaves no trace that a decision was taken at all. And shipping it unchanged because you could not articulate the objection, which treats your inability to phrase something as evidence that the thing was not true. Both convert a disagreement into a decision with no record that anybody decided anything, and both are invisible in the finished report, which is why neither gets caught.
Whichever move you take, write down the disagreement and its resolution somewhere the next reviewer will find it. Over a few months that log becomes the only real evidence you have about where drafting fails on your particular kind of work, with your particular clients and your particular evidence sets. The three failure classes above are a starting set, drawn from the mechanics of how drafting works rather than from your material. Your log is what turns them into a list that actually fits what you produce, and once it exists it is worth more than any general advice, this included.
Who signs the report, and what that signature means
A signature on a professional report has never been a claim to have personally typed it. Reports have been drafted by juniors, assembled from templates and composed out of other people's sections for as long as reports have existed. The signature is a claim to have checked the work and to stand behind it when it is challenged, and nothing about how the draft came into being changes what is being claimed. This is exactly why the drafting method is not a defence. Nobody has ever accepted it was in the template, and nobody will accept the newer version of that sentence either.
The practical test is blunt. Pick any finding in a signed report at random and ask which check it received, re-derivation or spot check, and who ran it. If nobody can answer for a given finding, that finding was not reviewed. It was read, which is a different activity that produces a very similar feeling of having done something careful. A protocol that cannot answer this question about its own output is not a protocol, it is a habit, and habits degrade under deadline pressure in exactly the way written rules are meant to prevent.
Accountability also does not distribute. Two reviewers who each assumed the other had covered the high-severity section produce a report that nobody reviewed, and both of them will be genuinely surprised by that in the same meeting. Name one reviewer per section, in writing, before the review starts. Shared responsibility for a section is the standard mechanism by which a section ends up with none, and it fails silently, which is the worst property a control can have.
None of this is an argument against tooling, and it would be a strange argument for us to make. Zealsync builds Prooflin, a product in the AI audits and professional reports category, which takes supplied evidence and produces findings and reviewable reports. A product of that kind produces a report. It does not produce a signature, and the protocol above is written so that it can be run against any such output, ours included, by somebody with no relationship to whoever built the tool.
Disclosure: what the reader should verify for themselves
The interest is declared plainly. Zealsync develops products in this space, described on the Prooflin page, so read the argument above as coming from an interested party and weigh it accordingly. Nothing here rests on a claim about what our products do, and nothing here should be read as one.
The check you can run without us takes about an hour. Take the last report you delivered. Count the findings. For each one, mark whether a specific artefact was named as its evidence, whether it was re-derived or spot-checked, and who did it. Then mark which of them would have changed the client's decision if they had been wrong. If you cannot reconstruct those three columns from what you kept, the first thing to fix is not the protocol but the record, because a review you cannot describe afterwards is a review you cannot defend, and the moment you need to defend it is the moment the record does not exist.
And the condition under which this argument is wrong: if the three failure classes do not appear in your own log after a fair sample of work, then drafting behaves differently on your material than it does on ours, and you should keep your own list rather than borrow this one. The method is the part worth taking. The list is only the starting content for it. If the review problem is one you would rather work through against a specific piece of work, start a conversation.
Relevant Zealsync pages


