Skip to content

Growth

How to Tell If the Website Audit You Were Sold Is Any Good: Six Tests

Every list of audit red flags is published by an agency that would like to sell you a different audit. Six tests you can run yourself, on any report, without technical expertise.

Zealsync Insights21 min read
How to read an audit somebody sold you

A report lands in your inbox. There is a score out of a hundred on the second page, a red band headed critical issues on the fourth, and thirty pages behind that. You cannot evaluate most of the technical claims, which is the reason you commissioned somebody else to make them. What you can evaluate is the document. Whether a finding is arguable, whether it is attributed to something you could look at yourself, whether the report is honest about the limits of its own looking: those are questions about the writing, and a careful non-specialist can answer every one of them without opening a browser console.

That distinction carries the whole piece. None of the six tests below asks you to judge whether a render-blocking script is genuinely blocking the render, or whether a heading level is genuinely out of order. They ask whether you have been given enough to check the claim, or to hand it to somebody who can. A report that fails on those grounds may still be full of true findings. It is simply not a report you can act on, because you have no way of telling the true parts from the filler.

The obvious objection to a piece like this is that lists of red flags in a website audit report are almost always published by somebody who would like to sell you a different audit. The objection is fair, it applies here, and it is dealt with openly further down rather than left implicit. The tests are written so that you can turn them on this article's own author as readily as on anybody else's report.

What an audit is for, and what it is frequently sold for

An audit exists to reduce uncertainty about a decision. Something is going to be done differently, or not done, on the strength of what the report says: a budget released, a migration delayed, a supplier changed, a rebuild avoided. The report earns its fee by making that decision better informed than it would otherwise have been. Everything in the document that does not serve the decision is decoration, however accurate it happens to be.

The second thing an audit is frequently sold for is alarm. Not fraud, and usually not even dishonesty, because a sales audit can be assembled entirely from true observations. It is a document whose organising purpose is to establish that the situation is worse than you thought and that the author is the natural person to put it right. Its findings are real. Its severity labels are inflated but defensible. Its recommendations, taken together, describe a scope of work that happens to be about the size of a retainer.

The two documents are hard to separate on a first read because they are made of the same raw material. The difference is what each one does with uncertainty. Diligence declares uncertainty, because the reader's decision depends on knowing how firm the ground is. A sales instrument resolves uncertainty in whichever direction makes the problem look larger, and does it silently. Every test below is, at bottom, a way of establishing which of those two things the document in your hands actually did.

Tests of the finding itself

The unit of an audit is the finding, not the section and not the score. A report is roughly as good as its median finding, because the sections are only findings grouped and the score is only findings weighted. Read three findings properly, chosen at random rather than from the executive summary, and you will know more about the report than you would from skimming all of it.

Does each finding state a criterion, and name where it came from

A finding without a criterion is an opinion carrying a severity label. The criterion is the rule the observed thing deviates from: a published standard, a documented behaviour, a measurable threshold, or a goal you yourself stated. It answers the question according to what, and without an answer to that question there is nothing in the sentence for anybody to agree or disagree with.

Compare two versions of the same claim. Your images are not optimised names no criterion at all. The product listing template serves hero images at 2400 pixels wide and displays them at a maximum of 600 css pixels names an observation and implies a criterion, that intrinsic width should not greatly exceed rendered width, and you can now argue with it. Perhaps you have a reason for the excess. Perhaps the template is about to be replaced anyway. The point is not that the second version is right; it is that the second version can be shown to be wrong, and the first cannot.

Published standards make the strongest criteria available, because they are external to the author and you can consult them without asking permission. WCAG 2.2 success criterion 1.4.3, Contrast (Minimum), sets out contrast ratios for text along with its own exceptions for incidental and decorative text, so a finding that cites it can be checked by anybody holding the same document. Google Search Central's documentation on canonicalisation works the same way, as do the documented semantics of HTTP status codes. Where no external standard exists, the criterion has to be a goal of yours, and an honest finding says so rather than borrowing the authority of a rule that does not exist.

The full structure of a well-formed finding, and what a criterion has to do in order to be one, is set out in the anatomy of a finding. For reading a report you have already been handed, the first two elements are enough: something observed, and the thing it is being measured against.

Is every finding attributed to something you can re-check yourself

Attribution means the finding points at a specific, locatable thing: a URL, a template name, an exact string, a request, or a screenshot carrying the date it was taken and the viewport it was taken at. The test is mechanical. Could a second party, holding only this finding and with no access to its author, arrive at the same observation? If not, you are being asked to trust rather than to verify, and the gap between those two matters most exactly where the stakes are highest.

The common failure is the site-wide claim with no instance attached. Meta descriptions are duplicated across the site may be perfectly true, but it does not tell you across how many pages, which template generates them, or whether the duplication is a content management default that one change would resolve. Attach two URLs and the same sentence becomes checkable in half a minute, and its remediation becomes estimable. The sentence has not improved because it got longer. It improved because it became falsifiable.

This test is deliberately indifferent to how the finding was produced. A finding drafted by a tool that cites a URL, an observed value and a criterion is more useful to you than one written by a person that cites none of those things. Any standard that turns on authorship rather than on evidence is unenforceable from the outside anyway, because you cannot see who or what wrote a sentence. You can only see whether the sentence carries what it needs in order to be checked, which is the reason that is the thing to test.

Is anything at all marked inconclusive

A report in which every finding is resolved has either examined a very small surface or has quietly settled its uncertainties in whichever direction suited the author. Looking at a site from outside leaves real unknowns. Whether a redirect chain is deliberate. Whether a duplicated page is a campaign variant somebody meant to publish. Whether a low-contrast element is exempt as decorative or disabled under the very standard being cited. Whether a caching behaviour belongs to the origin or to the edge sitting in front of it. None of these is a failure of diligence. They are the honest edge of an outside-in examination, and a good report marks them as open and says what access would close them.

The absence of any inconclusive item is the tell. It suggests either that the awkward questions were never asked, or that they were asked and then answered in the direction that made the finding count. Look also for the softer version of the same move: a finding hedged carefully in its body, this may indicate, this could suggest, which then appears at full weight in the summary table and in whatever score the report produces. Hedging that does not survive into the count is not hedging. It is cover.

Weight this test by scope, because it can be run unfairly. A five-page brochure site examined properly may legitimately produce a report with nothing left open, since there is very little to be unsure about. On an estate with a dozen templates, an authenticated area and a third-party checkout, a clean sweep of certainty is not a sign of thoroughness. It is a sign that the boundaries of what could be seen from outside were never acknowledged.

Tests of the report as a whole

Read the scope statement for what was never examined

Read the scope statement before the findings, and read it for its silences. It should tell you which URLs or templates were examined, at which viewports, on which date, by what method, and, the part that matters most, what was deliberately left out. Authenticated areas, checkout flows, native applications, third-party embeds, non-primary locales and anything behind a login are the usual exclusions. There is nothing whatever wrong with excluding them. There is a great deal wrong with not saying that you did.

Scope also supplies your denominator. Forty-seven issues is not a quantity until you know across how many pages and how many templates they were found, and whether a single template-level defect was counted once or once for every page it renders on. Without a denominator the count is not a measurement. It is a mood, and it is a mood that has been engineered to feel large.

If there is no scope statement at all, you have already found the most serious defect in the document, and it is worth being precise about why. A report with no declared scope cannot be wrong. Anything you later discover it missed becomes, retrospectively, something that was out of scope all along. That is why this is the test to run first, and it is the one that short reports and free reports fail more often than any other.

Check whether the recommendations name an owner

A recommendation that names no owner has not been costed, and has probably not been thought about beyond the point of being written down. Improve page speed assigns work to nobody. The hero component in the listing template needs its image sizing attribute corrected, which is a change in the repository, and the six existing landing page images then need re-uploading through the content management system, which is an editorial task, assigns work to two different people with two different calendars. You can now check whether you have either of them, and you can see that the second half cannot begin before the first.

The useful owner classes are few. Somebody with repository access. Somebody with content management access. Somebody with hosting or DNS access. Somebody with the commercial authority to change the words on a page. And a third party you do not control. That last class is the one to hunt for, because a recommendation that depends on a supplier's roadmap or a platform's feature set is not really a recommendation to you at all, and a report that never once mentions a dependency it cannot see is a report that never asked.

The cost of this test, taken the other way, is real and worth stating. Requiring an owner and a route to delivery for every item in a two-hundred-item list produces a document nobody finishes writing and nobody reads, and it quietly pushes authors towards vaguer recommendations that are easier to assign. The reasonable expectation is that ownership is named for the recommendations the report itself claims matter most. If the top tier is unowned, the tiering was decorative.

Look for what happens if you do nothing, and over what horizon

Every finding carries an implied counterfactual: this is what continues if the work is not done. A report should be willing to write that sentence out, including in the cases where the honest version is nothing much, and not soon. An author who cannot bring themselves to write that sentence for any item in the whole list has told you something about how the list was assembled.

The horizon separates three categories that usually get merged. Some defects are static: a canonical tag pointing at the wrong URL is exactly as wrong in a year as it is today, and the cost of leaving it alone is flat. Some compound: a redirect chain lengthens with every migration that routes through it, and each addition makes the next one harder to unpick. Some have already done whatever they were going to do, so remediating them recovers nothing and only prevents recurrence. Sorting your list by which of those three a finding belongs to will change the order of work more than any severity column will.

This is also where website audit scare tactics are easiest to identify, because they share one shape: unbounded consequence with no mechanism. You are losing customers every day. Your rankings are at risk. This exposes you. Ask of any such sentence, by what path. A mechanism can be examined and can be wrong, which is what makes it useful. The enquiry form's error messages are not programmatically associated with their fields, so somebody using a screen reader who submits an invalid entry is not told which entry was invalid, names a specific person having a specific problem, and you can go and check whether it is true. A threat cannot be examined, which is precisely why it was chosen. The honest form of a consequence gives the mechanism and then admits that the size of the effect cannot be established from outside.

The padding patterns, and why each one is cheap to generate

Padding is not simply filler. It is a small set of specific moves, each of which converts something the author already possesses into pages, without requiring a judgement that could later be shown to be wrong. Once you can name them they become difficult to miss.

  • The per-page repetition. One template-level defect restated once for every URL on which it appears. Four problems become four hundred findings, the headline count becomes impressive, and the actual remediation effort is unchanged.
  • The screenshot without an inference. An image of a tool's output, captioned with a description of what the image shows. Nothing has been concluded, and the reader is left to perform the analysis the report was commissioned to perform.
  • The educational preamble. Several pages explaining what a title element is, or why loading time matters in general. This is tutorial copy. It is not about your site, and it could be printed unchanged in a report about somebody else's.
  • The tool-output transcription. Raw exports pasted in as sections, complete with the tool's own labels and default thresholds, unexamined, unreconciled with each other and unconnected to anything you said you were trying to decide.
  • The universal recommendation. Publish consistently. Improve internal linking. Review your metadata quarterly. All true of nearly every site and therefore about none of them, and none of them requiring anybody to have looked at yours.
  • The unpublished score. A composite number out of a hundred whose formula is not shown. It cannot be reproduced, it cannot be contested, and it moves whenever the author needs it to move.

What these six share is more useful than the list itself, because the patterns change over time and the underlying property does not. Producing any of them required no decision that could turn out to be a mistake. Diligence is expensive precisely because it commits: a finding with a criterion and an instance can be refuted, and the author knew that while writing it. That is the cost being avoided, and it is the cost you were paying for.

Padding is anything that could not have been wrong. Ask of each page what judgement was made here, and what would have shown it to be a mistake.

The legitimate reasons a report is long

Length is not evidence of padding, and brevity is not evidence of discipline. Apply a page-count heuristic in either direction and you will reject careful work while accepting thin work, so it is worth knowing the reasons a report legitimately runs long.

The estate may simply be large. A site with fifteen templates, three locales and a recent migration presents far more surface than a five-page site, and a proportionate report about it is longer. Length that tracks the size of the thing examined is not a warning sign. Length that does not track anything is.

Several surfaces may have been examined. Accessibility conformance, search visibility, delivery performance and content quality are four different criteria sets applied to the same pages, and a report covering all four is close to four reports bound together. This is also the place to check that they have not been silently merged: a finding sitting under a search heading but resting on an accessibility criterion suggests the sections were arranged after the findings were written, which tells you something about how the work was planned.

Evidence carried inline lengthens everything. A finding that states its observation, its criterion, its instance and the step you would take to re-check it runs to several times the length of the same claim asserted in a single sentence. That extra length is the attribution you were told to look for two sections ago, so you cannot consistently ask for brevity and checkability at once. Genuine remediation detail is long for the same reason, and appendices are long because raw exports, crawl data and full URL lists have to live somewhere, correctly separated from the findings rather than counted as findings.

The workable instrument is density rather than page count. Pick three pages at random from anywhere in the document, including the middle where nobody looks, and count the claims on each that could be argued with: claims that name something specific and state what it should have been instead. Two or three such claims to a page describes a working report. A random page carrying none is a page that cost nothing to produce, and pages like that are rarely alone.

A report that passes every test and still recommends the wrong work

The six tests examine the document's honesty. They do not examine its judgement, and the distinction is worth holding onto, because a report can be scrupulously attributed, candid about its scope, careful with its inconclusives and precise about ownership, and still hand you a programme of work that will not move the thing you care about. Honesty and usefulness are separate properties, and only one of them can be checked from the page.

There are two ordinary routes to that outcome. The first is a criteria set that does not match your problem. A conformance audit answers one question, where does this deviate from the standard, and it can answer it very well. It does not answer why enquiries stall on the second step of the form, and it was never going to; if that was the question you had in mind, you have bought a good report written to the wrong brief, and the fault for that is usually shared.

The second route is severity read as an order of work. Severity is a property of the deviation: how far the observed thing sits from the criterion. Priority is a property of your situation: your consequence, your cost, your dependencies, your sequence, your appetite. The two are related and are not the same, and a report can be entirely correct in its severity column while being useless as a plan. The modelling that separates them is set out in severity is not priority.

The protection here is not a further test of the document. It is to write down, before commissioning anything, the decision the audit is meant to inform and what you would do differently under each plausible outcome. If you cannot name a decision that changes, you are not buying diligence. You are buying reassurance, or a document to show somebody else, and both of those are legitimate purchases as long as you know which one you are making and price it accordingly.

Unlike the other tests, this one cannot be run on the document alone. It requires you to know your own question, and no report can supply that, however carefully it has been written.

The questions to put to the author

Whatever the report says, the author's answers to a short list of questions will tell you more than another pass through the pages, and the manner of the answering will tell you most of all.

  • What did you look at, and what did you deliberately not look at?
  • Which of these findings would you keep if you were allowed only ten?
  • Which of them are you least confident about, and what would settle it?
  • For the top item, who has to do the work, and what do they need from us before they can start?
  • What would you expect to be observable once it is fixed, and roughly when?
  • What did you find that you decided not to put in?
  • If we did nothing for six months, which of these gets worse, and which stays exactly where it is?

The last two are the most revealing. A report is a set of inclusion decisions, and an author who cannot describe anything that was considered and then dropped was probably not making decisions at all. An author who answers the confidence question with an unqualified I am not sure, and here is what would settle it, has demonstrated the one habit that all six tests are indirectly looking for.

If the report came from a prospective supplier, that conversation doubles as the interview, and it is a better test of judgement than any sample deliverable, because it costs nobody unpaid work and takes half an hour. Where you would rather interrogate the report you are already holding than commission another one, start with the brief rather than the document.

How long should a website audit be?

There is no page count that signals quality in either direction. A report runs legitimately long when the estate is large, when several distinct criteria sets were examined, or when each finding carries its evidence, its criterion and its re-check step inline instead of asserting a claim in one line. It pads cheaply through repeated boilerplate, tutorial preambles, screenshots captioned with a description of themselves, and a single template-level defect restated once for every page it appears on. Judge density instead of length: pick three pages at random and count the claims that name something specific and state what it should have been. Two or three to a page describes a working report.

Is a free audit worth anything?

Sometimes, and you judge it with exactly the same tests you would apply to a paid one, because price does not change what a finding has to state in order to be arguable. A free report with a criterion, an instance you can open yourself and an honest exclusions list is more useful than an expensive one carrying none of those. The specific caution is that short and free reports are the ones most often silent on scope, which leaves you with no way to know what was never examined and no denominator for the issue count. Read the scope statement first. If there is not one, that absence is the most important thing the document has told you.

Should I trust the severity labels in the report?

Treat a severity label as a claim about how far something deviates from a stated criterion, assigned by whoever or whatever wrote the report, and check that the criterion is actually stated. What the label cannot encode is your consequence, your remediation cost, your dependencies or your delivery sequence, because none of those was visible to the author at the moment of assignment. A severity column is therefore legitimate evidence and is not an order of work, and a report presenting it as a work programme has quietly substituted its own situation for yours. Converting severity into a sequence is a separate question, and it deserves its own treatment.

Pass it on

Share this signal

Send it to someone working through the same question.