Skip to content

Growth

What to Ask an Agency Before You Hire One, Without Requesting Free Work

A pitch measures pitch production. It selects for idle capacity and rewards the agency willing to propose before it understands. What to evaluate instead, how to decide whether one supplier or several, and a process that requires no deliverable.

Zealsync Insights29 min read
Judging an agency without asking for free work

You are about to spend money on judgement you cannot inspect. That is the real difficulty in hiring an agency, and no amount of process removes it. A portfolio shows finished work without the constraints that shaped it. References were chosen by the agency. Awards describe the work of the people who entered them. The industry's usual answer to this is the pitch, which asks several suppliers to produce a document about a situation none of them has yet been allowed to examine, and then treats the best document as evidence of the best judgement. There is a better use of the same two hours.

A disclosure before anything else: Zealsync operates an agency brand

Zealsync operates Intense Path, its Brand, Growth and Technology agency brand. This is therefore advice from an interested party about how to interrogate parties like it, and you should read it on that basis. The only defensible test is whether each criterion below would be uncomfortable to sit on the receiving end of, including for the party publishing it. A question that is easy for the author to answer and awkward for a competitor to answer is a sales device wearing the clothes of a criterion, and it has no place in a list like this.

Two things follow. First, no engagement, client, result or example from Zealsync's own work appears anywhere in this article, because none of it would be verifiable by you, and unverifiable evidence is worth less than none at all. Second, the section that matters most is the one on whether to divide the work between several suppliers, because a single practice has an obvious commercial interest in how you answer that. The argument is set out below with its counter-case attached, for exactly that reason.

What a pitch measures, and what it quietly selects for

A pitch measures pitch production. It measures the ability to assemble a persuasive document at short notice, at the agency's own cost, about a problem the agency has been permitted to ask perhaps four questions about. That ability is real, and some agencies have a great deal of it. It is not the ability you are buying, unless the thing you need made is a pitch.

The more serious issue is what a pitch selects for rather than what it measures. Three effects run underneath it, and all three run against you. The first is capacity. Unpaid speculative work is expensive, and the agencies with the most room to do it are the ones with the least booked work; the agency whose senior people are currently doing the work you admired is the agency least able to give you a week of their time for nothing. The second is temperament. A pitch rewards the supplier most comfortable proposing before it understands, because the process has a deadline, and the confident document beats the honest one that says the cause is not yet established. The third is the sample. The people who assemble pitches are, at least in part, the people who are good at assembling pitches. You evaluate them, and then you buy somebody else.

A pitch rewards the supplier most comfortable proposing before it understands, which is close to the opposite of what you are trying to buy.

There is also a cost you do not see. Speculative work is not free to the agency, so it is priced somewhere: into the fees of clients who did not ask for it, into the margin, or into the seniority of the people who eventually get assigned. Asking four agencies to produce free work does not make the work free. It makes three of those bills somebody else's problem and one of them, in some form, part of what you pay.

The case against pitching is not universal, and it is worth stating where it fails. When the deliverable is genuinely specified, bounded and comparable — a defined production job, a fixed media buy, a build against a written specification that will not move — asking several suppliers to quote an approach is a reasonable way to compare price and method, and it costs them little, because the hard thinking has already been done by you. The pitch becomes destructive when the specification is the thing in dispute. If what you need is a view on what the problem actually is, a process that awards the work to whoever describes the solution most attractively has selected on the wrong axis, and the appointment is made before anybody has looked.

So the useful question is not how to run a better pitch. It is what you can observe about judgement in a conversation, given that judgement is invisible and everything else on offer is a proxy for it.

What to evaluate instead

Three things are observable without a single deliverable being produced. Each is available in ordinary conversation, each is difficult to fake for two hours, and each has a specific tell. Taken together they will not tell you whether an agency is good in the abstract. They will tell you whether it is trying to understand your situation or trying to sell you its standard offer, and that distinction predicts more about the next six months than anything in a credentials deck.

The questions they ask before they propose anything

The most informative hour of an agency relationship is usually the first, and the informative part of it is not what they say about themselves. It is the shape of their questions. Count them if it helps. Then sort them, because they are not all the same kind of question, and only one kind tells you anything about how this supplier thinks.

Qualifying questions establish whether you are worth pursuing: budget, timeline, who signs, whether there is an incumbent. They are legitimate, and you should not penalise them — an agency that fails to qualify will waste your time and its own. But they say nothing about judgement. Diagnostic questions have a different property. A diagnostic question has at least two plausible answers, and the answers point at different work. If you can imagine both answers leading to the same recommendation, the question was decoration.

  • What have you already tried, and what made you stop. The answer is where most of the available evidence already lives, and it is the question that most reliably separates suppliers who intend to look from suppliers who intend to present.
  • Who inside the business disagrees with the brief, and on what grounds. A brief is usually a negotiated settlement between people who wanted different things, and the unresolved disagreement resurfaces later as an approval that will not come.
  • What is already decided and not open to revisiting. The constraint nobody mentions in the first meeting is the one that kills the work in month four.
  • What would have to be true for this to be the wrong project. A supplier willing to ask that in a first meeting is telling you it can afford to lose the work, which is the precondition for every piece of honest advice it will give you afterwards.
  • How would you know, without asking anybody, that this had worked. A vague answer here is yours to fix rather than theirs, but the question needs asking before anything is scoped.

The tell to watch for is the question that changes nothing. Some suppliers ask a great many questions and then present the recommendation they would have made regardless. You can detect it at the second session by checking whether anything you said in the first has altered the shape of what is proposed. If nothing moved, the questions were rapport.

Whether they can restate your problem better than you stated it

Give them your problem in your own words, once, without your theory of what is causing it. Then, at the second session, ask them to give it back. You are not testing memory or paraphrase. You are testing whether the version that comes back is more specific than the version that went out.

A good restatement has recognisable properties. It names a particular group of people, at a particular moment, doing or failing to do a particular thing, for a stated reason. It separates the symptom you noticed from the mechanism that would produce it. It can be wrong, and it says how you would find out. Working out which problem you actually have is a piece of work in its own right, and it is worth having attempted it yourself before the meeting, so that you have something of your own to compare the restatement against.

  • It is narrower than what you said, not broader. A restatement that widens the problem is usually widening the scope.
  • It introduces a distinction you had not made, rather than vocabulary you had not used. New words for the same lump are not progress.
  • It is falsifiable. If nothing could show it to be wrong, it is a mood rather than a diagnosis, and it will still be unfalsifiable in month six when you are trying to decide whether the work is going anywhere.
  • It sometimes disagrees with you. The most valuable restatement you receive may be the one telling you that the thing you came in asking for is not the thing that would help, and that sentence is expensive for a supplier to say.

The failure mode is fluency. A restatement can be elegant, well phrased and pitched at exactly the level of abstraction you handed over, adding polish and no resolution. Test it by asking what it rules out. A diagnosis that rules nothing out has narrowed nothing, and narrowing is the entire function.

Whether the people in the room are the people who do the work

Ask plainly who would be assigned, by name and role, and what proportion of their week the account would take. Then ask the second-order questions, which are the ones that actually resolve it: who writes the first draft, who reviews it before it reaches you, who you speak to in week six, and what happens to your account when one of those people is pulled onto something urgent elsewhere.

This matters for a reason that has nothing to do with seniority theatre. Everything else you are evaluating — the questions, the restatement, the willingness to disagree with you — is a property of specific individuals rather than of a company. If the person conducting the conversation will not touch the work, you have carefully assessed one set of people and are about to buy a different set. That can still be an acceptable arrangement, and the section below sets out when it is. But you should know that it is what you are doing.

Whatever you are told, ask for it in the proposal in writing: names, roles, and the review points at which each person is involved. A staffing promise made in a room costs nothing to make and nothing to break. The same promise written into a scope document is something you can point at in week six, which is when you will need it.

One supplier or several, and what has to be true for a split to work

This is a hiring decision rather than a diagnostic one, and it is usually made backwards: the number of suppliers is settled first, out of habit or procurement policy, and the problem is then divided to fit. Decide it the other way round. Neither answer is correct by default, and the honest position is that a split is correct under three conditions and damaging without them.

Splitting works when the workstreams are genuinely independent — when a decision taken inside one of them does not change what is right inside another. Where the problem itself spans disciplines, splitting does not divide the problem. It divides the part of the problem each supplier can see. Each then solves a smaller problem correctly, and what remains is the seams. Three things sit in a seam, and none of them belongs to a supplier by default.

Decisions that would otherwise belong to nobody

Some decisions fall between workstreams and set the ceiling for both at once. Whether a new claim changes what the product actually promises. Whether a technical constraint is worth designing around or worth removing. Whether a launch waits for the mechanism that would make it measurable, or ships without it. Each of these can be answered sensibly from inside either workstream, answered differently from inside the other, and neither supplier has the standing to overrule its counterpart.

The condition, then, is that one person inside your business is named to own the between-decisions before anything starts, with the authority to settle them and the availability to do it within days rather than at the next steering meeting. A committee does not satisfy this; a committee is where between-decisions go to wait. If you cannot name the person, you have learned something useful about your own readiness rather than about any supplier: you are not currently in a position to run a split, and adding suppliers will add seams you have nobody to hold.

Handovers nobody is paid to own

A handover that arrives as a presentation is not a handover. It is a description of one.

Agree in advance what each side hands the other, in what form, and what counts as complete. The claim written out, including the words that may not be used. The page inventory with the decision attached to each entry. The measurement definitions stated precisely enough that two people working separately would compute the same number. None of that survives being summarised into slides, because the value is in the specificity that slides remove.

Then notice the incentive. Whoever's scope ends first has the least reason to make the handover good: they are already finished by the time it matters, and the cost of a poor one lands on somebody else. The remedy is unglamorous. Price the handover into the scope that ends first, name it as a deliverable with acceptance criteria, and have the receiving supplier confirm it is usable before the sending supplier is signed off. Nothing about this is expensive. It is simply nobody's job unless you make it somebody's.

Answerability for the whole, when the whole underperforms

The real test of a supplier structure is not what happens when the work goes well. It is what happens when the result is disappointing and everybody involved is individually correct. Under a split, that outcome is entirely available: each supplier can demonstrate that its part performed against its brief, and the brief was the thing that was wrong. With a single supplier there is nobody to point at, and no version of the conversation in which the answer is that the joins were somebody else's responsibility. That is a genuine advantage for you.

It is also, transparently, the argument a single practice would like you to accept, so weigh it and then discount it. Zealsync's own agency brand, Intense Path, covers brand, growth and technology work, which places the party writing this on one side of the question. The counter-case therefore deserves to be made properly rather than gestured at.

Concentrating the work concentrates the risk. You lose the ability to replace one part without disturbing the rest. You lose the price tension that comes from holding comparable suppliers side by side. Most importantly, you risk buying a discipline you would never have selected them for, because it arrived in the bundle. The test for that is direct: for each discipline in the bundle, ask whether you would have hired this supplier for that discipline alone, on the evidence in front of you. Where the answer is no, either take it out of scope or accept openly that you have chosen convenience over quality on that part, and say so out loud now rather than discovering it in month five.

The summary is unromantic. Split when the workstreams are independent, when the between-decisions have a named owner inside your business, and when the handover form is agreed before anything begins. Consolidate when the problem spans disciplines and any of those three conditions is missing. Do not decide it on supplier count, and do not let a procurement preference for three quotes determine the shape of work whose difficulty lies precisely in the joins between them.

What happens when the person who ran the diagnosis is not assigned to the account

This is the most common substitution in the industry, and it is not automatically a scandal. The activity you evaluated — asking the right questions, restating the problem, disagreeing with you where it counted — is diagnostic work, and there is a legitimate model in which a small number of senior people do that across many accounts while delivery teams execute. What determines whether the substitution is survivable is whether the diagnosis exists anywhere other than inside the diagnostician's head.

It survives when the diagnosis is written down in enough detail to be argued with: the problem statement, the hypotheses, what would confirm or kill each one, and the decisions already taken with the reasons attached. It survives when the person who wrote it returns at named review points rather than at a vague invitation to escalate. And it survives when the delivery lead is present in the second session, so that you have at least evaluated one of the people who will actually be there.

It fails in a recognisable pattern. The diagnosis is never written. The delivery team receives a brief instead, which is a set of instructions with the reasoning removed. The first time they meet you is at a kick-off where you repeat everything you said in the first session to people who cannot tell which parts were load-bearing. Six weeks later the work is defensible against the brief and unrelated to the problem, and nobody involved has done anything wrong.

There is a mirror-image failure worth naming, because the obvious fix creates it. Insisting that the senior person stays on the account can produce a senior person who attends, approves and is billed, while contributing nothing that a written diagnosis would not have contributed for less. What you want is not presence. It is that the reasoning is transferable and that somebody remains answerable for it. Ask which review points the diagnostician attends, what they are there to decide, and what happens if they disagree with the work at one of them.

A two-session evaluation that needs no deliverable

Two conversations, roughly a week apart, with nothing produced in between. The structure is deliberate: in the first you talk and they ask, in the second they talk and you assess. Run it identically with each supplier, in the same order, with the same material, or the comparison is worthless before it starts.

What to bring to the first session:

  • The problem in your own words, stated once, at the start, in under two minutes.
  • What you have already tried, when, and what made you stop. Include the attempts that embarrass you; those carry the most information.
  • The constraints that cannot move: regulatory, contractual, technical, political.
  • The decisions already taken that are not open to revisiting, and who took them.
  • Whatever evidence you already hold, offered rather than withheld. You are not testing whether they can extract data from you; you are testing what they do with it once they have it.

What to withhold, and the reason in each case:

  • Your theory of the cause. This is the important one. Hand over your diagnosis and you can no longer evaluate theirs, and most suppliers will adopt the diagnosis you brought, because agreeing with the buyer is the path of least resistance.
  • Your preferred solution. The same reason, one step further downstream.
  • The names of the other suppliers you are speaking to. It changes what they propose rather than how well they think.
  • Your budget, until you are asked for it, and then answer with a range. Refusing outright wastes everyone's time and produces proposals scoped to a fantasy, but volunteering the number unprompted removes a question whose asking you wanted to observe.

During the first session, your notes should be about their questions rather than their answers. Write each question down as it comes. Afterwards, mark the ones whose answer could have changed what they would recommend. That count, more than anything else in the process, is the thing you are buying.

Between the sessions, nothing is produced. Do not request a document, a deck or a route. If a supplier volunteers one anyway, treat it as data rather than generosity: it tells you the proposal was available before the diagnosis was.

In the second session, ask for four things, none of which is a deliverable:

  • The restatement. Their version of your problem, more specific than yours, together with what it rules out.
  • The questions they would still want answered, and how they would answer them. This is where competence shows most plainly, because it requires admitting in front of a prospective client what is not yet known.
  • Two or three hypotheses, each with what would confirm it and what would kill it. A hypothesis without a kill criterion is a position, not a hypothesis.
  • The shape of a first phase — its sequence, its decision points, who is on it — rather than its content. The content is the work, and you have not bought it yet.

The distinction governing the whole exercise is between a judgement and an artefact. Asking a supplier what they think, how they would find out, and what they would do first is asking for judgement, and it is entirely reasonable to ask for that before signing anything. Asking for a creative route, a sample audit, a keyword plan, a wireframe or a strategy document is asking for an artefact you could take away and use unchanged, which is work. The test is simple: if you cancelled the process tomorrow and could still use what they gave you, you asked for work.

Where you genuinely do need a deliverable before committing, because the scope cannot be settled without one, the honest resolution is to pay for it. Scope it as a small engagement with its own fee, its own acceptance criteria and its own ownership of the output, and make explicit that it obliges neither party to continue. That is a paid discovery rather than a pitch, and it inverts the incentive: the supplier is being paid to find out instead of paid nothing to convince.

Scoring that resists presentation quality

Presentation quality contaminates every judgement made in a room. The difficulty is not that you are impressed by polish; it is that polish raises your assessment of things it has no bearing on, and you will experience that as having found the reasoning more persuasive. The only reliable defence is to commit to criteria before you meet anybody.

  • Write the criteria and their weights before the first session, and do not revise them once you have met a supplier. Revising criteria after a meeting is how criteria come to describe whoever you liked.
  • Score immediately after each session, alone, before discussing it with colleagues. Then compare. Divergence between two people who both attended is more informative than the average of their two scores.
  • For the first session, score the questions rather than the answers: how many were diagnostic, how many were uncomfortable to answer, how many could have changed the recommendation.
  • For the second, score the restatement against the written problem statement you brought in. This is why you write that statement down before you start rather than after.
  • Give rapport its own line with a small weight. It matters, because you will be working with these people for months. Given an explicit home, it stops leaking into your assessment of everything else.
  • Keep disqualifiers separate from scores. Some findings should end a conversation rather than cost it four points: a proposal produced before any diagnostic question, a refusal to name what failure would look like, a refusal to name the delivery team.

None of this is arithmetic, and treating it as arithmetic is a failure of its own. The purpose of a rubric is to make you notice the moment you are about to override it. Overriding it is allowed. Overriding it without writing down why is not, because the written reason is exactly what you will want to re-examine in six months, when you are trying to establish whether the process was sound or merely comfortable.

The approach would be wrong in one case, and it is worth being clear about it. If the thing you are buying genuinely is presentation — a keynote, an investor narrative, a launch film — then presentation quality is the deliverable and belongs at the top of the weighting. The error is not weighting craft highly. It is importing the criteria for one purchase into a different purchase. Decide which one you are making before you decide what to score.

Red flags, with the reasoning behind each one

A warning sign is only useful if you know what it is a sign of. Each of the following is a flag because of a specific mechanism, and it is the mechanism that tells you how much weight it deserves in your particular case.

  • A proposal arrives before a single diagnostic question. The mechanism is straightforward: the shape of the work was decided before your situation was known, which means it came from a template, from the last client, or from whatever the supplier currently has capacity to sell.
  • The room is staffed by people who will not do the work. You are assessing a sample that will not be delivered. The flag weakens if the delivery lead is present and the diagnosis is written down; it does not disappear.
  • They will not name what would count as failure. An approach that cannot fail cannot be evaluated at any later point either, and you discover this at the review where everybody agrees things are progressing and nobody can say progressing against what.
  • Scope grows without any decision being recorded. Each increment is individually reasonable and the aggregate is unaccountable, because there is no artefact to point back at when the budget has gone.
  • Every answer returns to the same offer. A supplier who sells one thing will find that one thing, usually honestly, because it is what they have learned to see. This is a reason to weigh their diagnosis against a second opinion rather than necessarily a reason to walk away.
  • Confident numbers about your outcome before they have seen your data. The number came from another account or from nowhere, and neither origin survives contact with your situation.
  • Results presented without a denominator. A percentage with no baseline, no timeframe and no statement of what else changed during the period is not a result. It is a shape.
  • The incumbent is disparaged early. Criticism of prior work before the constraints that produced it are understood tells you how your own work will be characterised to the next supplier, and how little curiosity gets applied before a verdict.
  • They cannot say what they are not good at, or what they would decline. Every practice has a boundary, and a supplier who will not describe theirs has either not examined it or decided you should not know where it is.
  • Urgency attaches to your decision date rather than to your problem. A deadline that exists because a rate is expiring or a slot is closing is about the supplier's utilisation, and their utilisation is not your constraint.

One caution about the list as a whole. These are signals rather than proofs, and a single flag in an otherwise careful conversation is worth raising with the supplier rather than acting on silently. How they receive the challenge is itself the most useful test in this article. A supplier who takes an awkward observation well, examines it, and adjusts has just shown you what disagreements will look like once money is involved.

The questions, numbered, in the order to ask them

Take these into the room. The order matters more than it looks. The early ones establish whether they intend to look before proposing, the middle ones test the diagnosis they have formed, and the last ones settle the terms on which you would actually work together.

  • 1. Before we go any further, what would you want to ask me?
  • 2. What have you seen so far that makes this worth your time, and what would make you decline it?
  • 3. What would have to be true for this to be the wrong project to run at all?
  • 4. How would you describe our problem back to me, in one or two sentences, more precisely than I described it?
  • 5. What does that description rule out?
  • 6. What would show us that description is wrong, and how quickly would we see it?
  • 7. What do you still not know, and how would you find out?
  • 8. Which two or three hypotheses would you test first, and what would kill each one?
  • 9. Who, by name, would be on this account, what would they do, and how much of their week would it take?
  • 10. Which of those people are here today, and who am I speaking to in week six?
  • 11. If you were only allowed to do one part of this, which part, and why that one?
  • 12. What would you not do, that a supplier in your position would usually propose?
  • 13. What would count as this having failed, and when would we know?
  • 14. What do you need from us for this to work, and what happens if we do not provide it?

The last two questions are the bridge to everything this article deliberately leaves alone. What can honestly be measured in the first ninety days, what the terms of an engagement should say, and what a reasonable horizon looks like are separate questions with separate answers, and they are settled after appointment rather than during selection. What ninety days can honestly show takes that up where this piece stops.

One closing note on how to use the process. It should be capable of returning the answer that you do not need an agency at all, or not yet, or not for the part you were about to buy. If it cannot return that answer, it is not an evaluation, it is a procurement ritual with a conclusion already attached. If you would rather work through the selection with somebody than alone, bring the problem rather than the brief. Intense Path is Zealsync's Brand, Growth and Technology agency brand, and every question above is one it should be asked first.

Is it better to hire one agency or several specialists?

Neither is correct by default. Splitting is right when the workstreams are genuinely independent, when one named person inside your business owns the decisions that fall between them, and when the form each handover must arrive in is agreed before work starts. Miss any of those three and the split creates seams that belong to nobody. It is wrong when the problem itself spans disciplines, because each supplier then solves a smaller problem correctly and the joins are what remain unowned. Consolidating carries its own cost: concentrated risk, less price tension, and the chance of buying a discipline you would never have chosen on its own merits. For each discipline in a bundle, ask whether you would have hired that supplier for it alone.

Is it reasonable to ask an agency for ideas before signing?

Yes, provided you are asking for a judgement rather than an artefact. Asking what they think the problem is, what they would still want to find out, which hypotheses they would test and what would kill each one is reasonable, costs them an hour of thought rather than a week of production, and tells you more than a document would. Asking for a creative route, a sample audit, a keyword plan or a wireframe is asking for work. The test is whether you could cancel the process tomorrow and still use what they gave you unchanged. If you genuinely need a deliverable before committing, pay for it as a small scoped engagement with its own acceptance criteria, which changes the incentive from convincing to finding out.

Should I run a competitive pitch?

Rarely, and not for work whose difficulty is diagnosis. A pitch selects for the suppliers with the most idle capacity, because unpaid speculative work is expensive and the busiest teams cannot afford it, and for the temperament most comfortable proposing before it understands. It also assesses the people who assemble pitches rather than the people who would do the work. The substitute is two conversations a week apart with nothing produced in between. In the first you state the problem and they ask, and you record their questions rather than their answers. In the second they restate the problem, name what they still do not know, and offer hypotheses with kill criteria. Where the deliverable is genuinely specified and comparable, asking for quotes remains reasonable.

What are the strongest warning signs in an agency conversation?

A proposal that arrives before any diagnostic question, because the shape of the work was decided before your situation was known. A room staffed by people who will not do the work, because you are assessing a sample that will not be delivered. A refusal to name what would count as failure, because an approach that cannot fail cannot be reviewed later either. Scope that grows without any decision being recorded, because there is nothing to point back at when the budget has gone. Results quoted without a baseline, a timeframe or a statement of what else changed. Treat each as a signal worth raising with the supplier rather than acting on silently, because how they take the challenge is itself a test.

Pass it on

Share this signal

Send it to someone working through the same question.