Growth
What Can You Honestly Measure in the First Ninety Days of an Engagement?
Reporting cadence and maturation time are not the same clock. What class of evidence each discipline can actually produce, why the horizon comes from your own sales cycle rather than a published average, and what it is dishonest to ask for sooner.

Ninety days is a contract length. It is not a unit of evidence. It became the default review point because quarters exist, and because a quarter is long enough that a buyer feels patient and short enough that they do not feel abandoned. Nothing about the underlying work observes that boundary, and no mechanism in marketing, engineering or brand has ever agreed to finish maturing on the last Friday of a quarter.
The question underneath it is still a fair one, and it deserves better than a number. What you can honestly measure depends on the class of evidence the work is capable of producing. Different disciplines produce different classes, on different clocks, for reasons that have nothing to do with effort. Some of it is checkable the afternoon it ships. Some of it cannot exist yet at any price, because the mechanism that would produce it has not had time to run.
So the useful conversation before an engagement starts is not how long. It is what kind of thing each workstream will be able to show by the review, what each number will be compared against, and what it would mean if the number did not move. Three questions, all answerable in advance, all much harder to argue about later once money has been spent and someone needs the report to say something.
Reporting cadence and maturation time are different clocks
Reporting cadence is how often you look. Maturation time is how long the thing takes to become true. They are set by completely different considerations. Cadence is set by governance, by how often the people paying want reassurance, and by how quickly you would want to catch a mistake. Maturation is set by a kind of physics: how long a mechanism takes to run from cause to observable effect, given your particular traffic, your particular sales cycle, and your particular starting position.
When the two clocks are treated as one, two predictable failures follow. The first is impatience misread as diagnosis. A monthly report on a mechanism with a six-month maturation produces five reports of noise before it produces one report of signal, and if nobody said so at the start, the third noisy report gets read as evidence the work is not working. Direction gets changed at exactly the point where the original direction had not yet had a chance to be wrong.
The second failure is the mirror image, and it is more flattering, so it survives longer. An early number moves, and the movement gets reported as proof. Early movement in a small sample is mostly variance. It is the same variance that would have produced a decline had the coin landed the other way, and the honest version of the report says so at the time rather than quietly hoping the trend continues.
A monthly cadence over a six-month mechanism produces five reports of noise and one report of signal. Say which one you are reading.
The fix is unglamorous and takes about an hour. For each workstream, write the cadence and the maturation horizon as two separate lines in the same document, and write the derivation of the horizon next to it. A monthly report on a workstream with a five-month horizon is then explicitly a progress report, not a results report. Nobody is misled, and the report at month five carries the weight it should carry, because everyone has been waiting for it rather than assuming it already arrived.
What class of evidence each discipline can produce
There are broadly three classes of evidence available in the first quarter, and one thing that is often called evidence but is not. Conformance evidence is judged against a stated criterion, so it is available as soon as the work is done. Volume-dependent evidence needs enough observations to separate a difference from noise, so its clock is measured in events rather than in weeks. Lag-dependent evidence has a mechanism with a built-in delay, so it arrives when the delay expires and not before. The fourth thing, the one to be careful with, is a number that moved for reasons nobody has established.
Technical work: assessed against a published criterion, so checkable immediately
Technical work has the shortest honest horizon of anything in an engagement, and the reason is worth stating precisely, because it is a reason rather than a lucky accident. Technical conformance is judged against a stated criterion rather than against an outcome. WCAG 2.2 defines success criteria that a page either meets or does not meet, and the assessment is an inspection, not a wait. Google Search Central documents what structured data must contain to be valid, and validity is checkable the moment the markup is deployed. A redirect either preserves the request or it does not. A canonical either resolves to an absolute address on the production origin or it does not. A page either returns the status code it claims to return or it does not.
That immediacy is real, and it is also narrow, which is the part that gets skipped. Conformance is not effectiveness. A page can satisfy every criterion you can name and still fail to do the job it exists to do, because the criteria describe correctness rather than persuasion or relevance. When a ninety-day report leans heavily on technical conformance, read it as evidence that the foundations are now defensible, not as evidence that the commercial question has been answered. The two get conflated most often when the conformance numbers are the only ones ready in time, which is precisely when the conflation is least likely to be challenged.
The related trap is counting resolved items as though the count were the result. A list of forty closed findings tells you how much work happened, not whether the closed items were the ones that mattered. That is a separate judgement with separate inputs, and it is the whole distinction between severity and priority. A report that presents a count of closures as an outcome has quietly substituted an activity measure for a result measure, and it will keep doing so for as long as the count keeps rising.
Conversion work: evidence arrives with volume, not with time
Conversion evidence is governed by how many observations you collect, not by how many weeks have passed. Two sites can run the identical change for the identical ninety days and end with completely different levels of confidence, because one collected several thousand completions and the other collected forty. Elapsed time is a proxy for volume, and it is a poor one whenever traffic is low, seasonal, or concentrated in a handful of days in the month.
The practical consequence is that a low-traffic site waits longer for the same certainty, and often waits longer than the engagement lasts. That is not a failure of the work. It is a property of the sample. What it changes is the honest framing: on a low-volume site, conversion work is better described as removing known defects and reducing known friction than as a testing programme, because a testing programme cannot finish inside a horizon anyone is willing to fund.
The tell to watch for is a percentage without a denominator. A reported forty per cent improvement on a base of twelve completions is four extra completions, and four is a number that appears and disappears on its own. Ask for the count alongside every rate, every time, and ask which window the count was drawn from. If neither is offered, the number is decoration rather than evidence.
There is a genuine counter-case, and it cuts the other way. Some conversion changes need no test at all, because they are not experiments. A form that rejects a valid phone number, a submit control that fails on a common browser, a required field that a legitimate user cannot satisfy, an error message that names no remedy: these are defects, judged against a criterion, and they belong in the immediate class rather than the volume-dependent one. Fixing a defect does not require first proving that the broken version was worse. Insisting on a test before repairing something demonstrably broken is not rigour, it is delay wearing the costume of rigour.
Demand work: the horizon is your sales cycle and your crawl cadence
The question of how long before agency work shows results has a genuine answer for demand work, but it is a derivation rather than a duration, and it uses two variables you can measure yourself without asking anyone's permission.
- Your own sales cycle. Read it from your own records: the elapsed time between a first identifiable touch and a closed decision, taken across enough recent deals to see the spread rather than only the average. If half your deals close in three weeks and half take seven months, you have two horizons and you should plan for both, rather than splitting the difference into a single number that describes neither.
- Your own crawl and index cadence. How quickly are your pages actually revisited? That is observable in your own server logs and in your search console reporting, per section of the site, and it varies enormously between a frequently updated section and a page nobody has touched in two years. A change to a page that gets revisited every few weeks cannot influence anything before it has been revisited.
The horizon is roughly the sum of three intervals: the time for a change to be discovered, the time for discovery to affect visibility, and the time for a visitor arriving through that visibility to travel your sales cycle to a decision. None of the three is a constant, and all three are measurable in data you already hold. Add them and you have a defensible expectation with its working shown, which is a very different object from a number somebody quoted at you across a table.
Brand and positioning work: what evidence is even available, and what is not
Brand work is where measurement conversations get evasive, in both directions. One side claims it is unmeasurable and therefore beyond scrutiny. The other side demands a metric and gets handed one that has been reverse-engineered to look reassuring. Neither is honest, and the way out is to be specific about what evidence genuinely exists at ninety days and what does not.
Available early: internal consistency, which is simply whether the same claim appears in the same form across the site, the deck, the proposal template and the outbound messages, and which anyone can audit in an afternoon. Comprehension, which is whether a person outside the company can state what you do and who it is for after reading one page, unprompted, in their own words. The composition of the questions your sales conversations open with, which changes when positioning changes, and which the people having those conversations can report on directly. And whether a salesperson still has to explain the category before they can explain the offer.
Not available early, at any price: preference shift, unaided recall, share of anything, or a movement in how a market as a whole regards you. Those require either instruments you have not deployed or a passage of time that ninety days does not contain. Measuring brand work honestly in the first quarter means accepting that most of the available evidence is qualitative and structural, and saying so plainly rather than dressing it in a chart that implies a precision nobody has.
The limit of that argument matters, because it is the argument most easily abused. If brand work is exempt from measurement, it is also unfalsifiable, and unfalsifiable work is unmanageable. So require it to make a prediction anyway, of the kind that could turn out wrong. After this repositioning, the discovery call should no longer need to establish which category you are in. After this messaging work, the same three objections should stop arriving in the first ten minutes. Those are checkable by asking the people on the calls, they are checkable inside a quarter, and they cost nothing to collect.
Leading indicators stop being indicators once they become targets
A leading indicator earns its place by correlating with something you cannot see yet. Pages published, findings resolved, impressions served, keywords appearing anywhere at all, tickets closed. Each is a reasonable proxy while it is being observed passively. Each degrades the moment it becomes the thing being paid for, because the cheapest route to moving the number is almost never the route that would have moved the outcome the number was standing in for.
The degradation is not usually cynical, which is why it is hard to catch in a report. If the agreed measure is pages published, thin pages are not a betrayal, they are compliance. If the measure is findings resolved, the easy findings get resolved first and the structural one that would take three weeks stays open, and every individual decision along the way was locally sensible. Nobody chose to game anything. The measure selected for a behaviour, and the behaviour duly arrived.
The correction is to pair every leading indicator with a constraint that makes the cheap route unavailable. Publish counts get paired with a quality gate that somebody other than the author applies. Resolution counts get paired with the severity distribution of what was resolved, so that clearing twenty trivial items does not read the same as clearing three structural ones. Volume measures get reported as ratios with their denominators visible. And the lagging measure stays named in the document even during the months when it cannot yet be read, so nobody forgets which question the proxies were standing in for.
The counter-case deserves stating rather than waving away: you cannot simply abandon leading indicators and wait. That would mean flying blind for a quarter, discovering at the end that the direction was wrong, and having no intermediate record with which to work out where it went wrong. Leading indicators are how you steer. The rule is only that steering instruments should not be scored, and that the person doing the work should not be paid against a number they can move directly without touching the outcome.
Where attribution breaks between disciplines, and why it is usually instrumentation
Attribution arguments between disciplines are absorbing and almost always premature. Brand claims the demand, demand claims the conversions, conversion claims the improvement, and each party produces a model that supports its own claim. The models are usually not the problem. The problem sits one layer below, in whether the events being argued over were recorded correctly in the first place.
The recurring instrumentation faults are dull and easily checked. A form submission that fires its event on render rather than on successful submission, so failures are counted as successes. A redirect that drops query parameters, so the origin of a session disappears between the click and the landing. A filtered view that excludes internal traffic in one system and not in another. Consent choices that legitimately change what may be collected, applied differently by two tools. Deduplication windows that differ, so the same person is one visitor here and three there. Any one of these produces two systems that disagree, and no attribution model reconciles a disagreement about what happened.
So the test to run before any attribution debate is deliberately blunt. Take one event that both systems claim to record, fix a window, and compare the counts. If they agree closely, the disagreement is genuinely about modelling and worth having. If they differ by a wide margin, stop: you are arguing about the interpretation of numbers that do not describe the same events. This costs an afternoon, and it settles a surprising share of the arguments a ninety-day review would otherwise contain.
Which makes instrumentation verification one of the most honest deliverables available in the first fortnight of an engagement. It is conformance work, judged against a criterion, so it belongs in the immediate class and can be shown early. It has to happen before the baseline is captured, because a baseline recorded through faulty instrumentation is not a baseline, it is a second unknown layered on the first. And a baseline captured after work has already started is not a baseline at all, whatever it is labelled in the report.
What a good agency refuses to promise, and why that is the useful signal
Refusals carry more information than commitments, because commitments are cheap to make in a pitch and expensive to check afterwards. What an agency declines to promise tells you whether they understand the mechanism they are selling, and it tells you before you have paid anything at all.
The refusals worth hearing are specific ones rather than general modesty.
- A named ranking position by a named date. The result depends on the behaviour of a search system and of competitors, neither of which the agency controls, so a promise here is a promise about other people's decisions.
- A conversion rate percentage quoted before anyone has seen your traffic volumes. Without the volume there is no way to know whether the difference could even be detected inside the engagement, let alone achieved.
- A time-to-result taken from an average. If they cannot show the derivation from your own sales cycle and your own crawl cadence, the figure is borrowed from a business that is not yours.
- An unqualified guarantee of an outcome that depends on your own execution. If converting the demand requires your sales team to respond within a day, the outcome is jointly owned and should be described that way in writing.
- Conformance presented as commercial result. Meeting a published criterion is a real achievement with a real ceiling, and the ceiling should be stated in the same breath as the achievement.
What can be committed to is everything inside the agency's own control: the work delivered, the criterion each piece will be judged against, the reporting cadence, the instrumentation verified before the baseline is captured, and what will be reported versus what will be deliberately withheld until it means something. That is a substantial list, and every item on it is checkable. Intense Path is the Brand, Growth and Technology agency operated by Zealsync Private Limited, and this is the shape of pre-agreement conversation it is written to have: the measure, the baseline and the horizon settled in writing before the first invoice rather than negotiated at the review.
Refusals are also the part of an evaluation you can assess without asking anyone to work for free, which matters if you are trying to judge an agency without commissioning spec work. A conversation about what cannot be promised is free to have, hard to fake, and considerably more diagnostic than a sample deliverable produced without access to your data.
Agreeing the measure, the baseline and the horizon before work starts
Three things get written down per workstream, before anything is built, and each has a failure mode that only shows up if it is left vague.
The measure has to be defined precisely enough that two people reading the definition would extract the same number from the same data, and the source system has to be named. Not conversions, but completed submissions recorded server-side in the named system, excluding the named internal sources, counted once per session. Not visibility, but impressions for the named property in the named report over the named window. Vagueness here is not a small problem: it is the mechanism by which a disappointing quarter gets reinterpreted at the review, because a loosely defined measure can always be recomputed until it flatters somebody.
The baseline has to be captured before work begins, with its window stated and its known gaps written down. Every baseline has gaps. A period with a broken tag, a fortnight of unusual traffic, a section only instrumented halfway through the window. Record them at capture time, because after the fact they become convenient. And say explicitly whether the comparison will run against the immediately preceding period or the same period a year earlier, since a seasonal business can produce two opposite conclusions from identical data depending on which is chosen after the numbers are in.
The horizon has to be derived, and the derivation has to be visible. Written out, it is a short paragraph: this workstream depends first on discovery, which our logs show takes about this long for this section of the site; then on visibility, which we expect to become observable within this range; then on a sales cycle that our own records show runs this long at the median and this long at the upper quartile. Anyone can then argue with a step, which is exactly the point. A horizon nobody can argue with is a horizon nobody has checked.
A fourth item is optional but worth the ten minutes: name who reads the report and what decision it is meant to inform. A report written for a board that wants reassurance and a report written for an operator who has to choose next month's work are different documents, and trying to serve both usually produces one that serves neither. Deciding this in advance also removes a common source of late-quarter friction, where the report is judged for failing to answer a question it was never built to answer.
Naming in advance what would count as the work having failed
The last item is the one most often skipped, and it is the one that makes the rest enforceable. Before work starts, write down what result would mean the approach was wrong. Not what would be disappointing. What would be disconfirming.
It looks different per discipline, and in each case it is a sentence with a date attached. For technical work: the criteria we said we would meet are still unmet at the review, or they were met and the underlying condition recurred within a month because the cause was never addressed. For demand work: at the point where the derived horizon expires, the discovery step has not happened at all, which is a different failure from having been discovered and disregarded, and the two need different responses. For conversion work: we did not accumulate enough observations to distinguish anything, which is a failure of the plan rather than of the change, and it was foreseeable from the traffic on day one. For brand work: the specific prediction we made about sales conversations did not come true, and the people on those calls say so.
If no result could have disconfirmed the approach, then nothing arriving at the review can confirm it either.
Naming failure in advance does something no amount of good intent achieves on its own. It removes the incentive to renegotiate the measure on day eighty-nine. When the definition of failure was written by both parties while both were still optimistic, the review becomes an act of reading rather than an act of persuasion, and the conversation that follows is about what to do next rather than about whose framing wins the room.
There is a legitimate objection to all of this. Circumstances genuinely change inside a quarter. A competitor moves, a platform changes its behaviour, a product decision alters what is being sold. Refusing to revise anything is its own kind of dishonesty, and a plan that cannot absorb new information is not a plan but a script. So the discipline is not that nothing may change; it is that a change is recorded as a change, with its date and its reason, and the original stays visible next to it. A revised measure with its history attached is legitimate. A revised measure that has quietly replaced the original, with no trace that an original ever existed, is the thing to refuse.
Everything described here is a paperwork exercise, which is exactly why it gets skipped in favour of starting the work. It takes a few hours at the beginning and it decides whether the review at the end is a shared reading of evidence or an argument about interpretation. If you want to have that conversation about a specific engagement, start with the diagnosis rather than the deliverable, because the measure you should be agreeing to depends entirely on which problem you actually have.
How long before agency work shows results?
There is no single number, and the mechanism is more useful than one. Technical conformance is checkable immediately, because it is judged against a stated criterion rather than against an outcome: a page either meets a WCAG 2.2 success criterion or it does not, and that is an inspection rather than a wait. Conversion evidence arrives with volume rather than with elapsed time, so a low-traffic site waits considerably longer for the same confidence, sometimes longer than the engagement lasts. Demand horizons are set by two variables you can measure in your own data: your sales cycle, read from your own closed decisions, and your crawl and index cadence, read from your own logs and search console reporting. A published average time-to-result is describing somebody else's business, with their cycle, their site and their starting position, and it should not be treated as a forecast of yours.
What should be in a ninety-day report?
Four things. The measure and the baseline exactly as they were agreed before work started, unrevised, or revised with the change and its reason recorded beside the original. The class of evidence each workstream can produce by now, stated plainly, so that conformance results are not presented as commercial outcomes and volume-dependent results are not presented before there is volume. What has been ruled out, which is real progress even when it is not the progress anyone hoped for. And what remains unknown, with the date at which it should become knowable. That last section is the working test: a report with no unknowns in it is answering a shorter question than the one you asked, because a genuine quarter of work on a real business always leaves something still maturing.
Relevant Zealsync pages


