Commissioning Teams: 4 Ad Jobs and the Copy Testing Metrics to Use
Discover the key copy testing metrics, from attention and brand recall to persuasion, and learn how to turn creative testing results into better decisions.
Discover the key copy testing metrics, from attention and brand recall to persuasion, and learn how to turn creative testing results into better decisions.
The copy testing metrics worth commissioning fall into six or seven constructs: attention, comprehension or main-point communication, brand linkage and memorability, relevance and interest, liking, and persuasion or purchase intent. The organising rule that matters more than any single number is this: choose the battery that matches the ad’s intended job, whether that’s stopping power, message transmission, brand building or driving action.
Each construct answers a different commissioning question, and conflating them is the single most common mistake we see in creative testing briefs.
Attention measures whether the ad earns a look in the first place. It’s typically captured through eye-tracking, click-through on a thumbnail stimulus, or self-reported stopping power (“How likely would you be to stop and look at this?”). Attention predicts nothing about message quality on its own.
Comprehension and main-point communication test whether the audience extracted the intended message, usually through open-ended “what is this ad telling you?” prompts coded against the brief’s key message.
Brand linkage and memorability separate prompted recognition (“have you seen this before?”) from unaided recall (“what ads have you seen for this category recently?”). Unaided recall is the harder, more diagnostic test, because a memorable ad that nobody attributes to the right brand has a rework problem, not a creative problem.
Relevance and interest gauge whether the audience feels the message applies to them. These are useful early filters but weak standalone predictors of sales effect.
Liking is the most seductive and least reliable metric in the battery. An ad can be warmly liked and commercially inert, or mildly disliked and highly persuasive. Treat liking as a diagnostic input, never a decision criterion.
Persuasion and purchase intent come closest to a behavioural proxy, shifting stated likelihood to buy or consider pre-versus post-exposure. Even here, predictive validity research from Wharton found that copy-testing purchase-intent predictions selected the higher-recall ad correctly only around 57% of the time across large samples, useful above chance, far from a guarantee.
This is precisely why composite single scores are risky: they collapse constructs that move independently into one number, hiding exactly the diagnostic detail a commissioning team needs to decide what to fix, as discussed in frameworks like those outlined in LLM evaluation frameworks for marketing leaders | AuthorityLayer Insights.

Before specifying a questionnaire, define the decision the test needs to support. We find four recurring ad jobs, each pointing to a different compact battery.
Each battery carries different trade-offs. Forced-exposure online surveys are cheaper and faster but less naturalistic; simulated or in-context exposure costs more and takes longer but produces findings closer to real media conditions. Sample type matters too: category buyers versus general population will shift persuasion scores meaningfully.
Before fieldwork begins, pre-specify four things: the criterion outcome the business will act on, the benchmark or threshold for success, the exposure assumptions (single or repeated viewing), and the sample definition.
Pro Tip: Write the decision rule before you see any data: “If brand linkage falls below our benchmark, we rework the opening five seconds” is far more useful after the fact than after the result.
Question wording should map directly to the construct you’re testing, not to a generic “rate this ad” scale. Attention uses stopping-power stems; comprehension uses open-ended main-message capture, coded rather than scaled; brand linkage separates prompted and unaided recall explicitly; persuasion uses a pre/post purchase-intent shift on a standard intent scale.
One in roughly 1.75 predictions is wrong even in validated models, since large-sample predictive-validity research put copy-testing purchase-intent accuracy at about 57%, a meaningful improvement on chance but no substitute for judgement.
The recurring pitfalls we see: treating liking as a proxy for persuasion when the two routinely diverge; leaning on a single metric to carry a decision that needs a battery; and running samples too small to detect a meaningful shift in intent, which makes a “no difference” result impossible to interpret with confidence.
Set thresholds before the data arrives, and separate statistical significance (is the shift real) from practical significance (is the shift big enough to matter commercially). A one-point intent shift that’s statistically significant on a huge sample may still be too small to justify a media change.
A useful diagnostic flow runs in sequence: attention first, then comprehension, then brand linkage, then persuasion. A failure early in that chain usually explains a failure later in it.
That last row matters: when every copy metric looks healthy but business results don’t follow, the fix usually sits in media and targeting decisions, not the creative itself.
Having commissioned and analysed creative tests across financial services, consumer brands and B2B categories, the clearest lesson is that commissioning teams get better decisions by specifying the ad’s job before writing a single question. We’ve applied this metric-battery approach within our ad and creative testing work, pairing it with brand tracking and segmentation so test results connect to longer-term brand measures rather than sitting as one-off scores.
Commissioning a copy test well means specifying the decision first, which is where a senior-led consultancy earns its place over a self-serve panel tool. Our Ad & Creative Testing work builds the metric battery around your specific go/no-go or rework decision, backed by our broader concept and product testing and customer segmentation capabilities when the question extends beyond a single ad.
Pro Tip: Bring your media plan to the discovery call: exposure frequency assumptions change which metrics matter most.
For organisations building a continuous testing capability rather than a one-off study, our Foundation, Growth and Strategic insight partner plans cover ongoing creative and brand research; current prices are on the pricing page.
The core constructs are attention, comprehension, brand linkage, relevance, liking and persuasion or purchase intent. Which ones matter most depends entirely on the ad’s job: a launch ad needs comprehension and persuasion, while a brand ad in a crowded category needs attention and linkage.
Not reliably. Liking and persuasion are related but distinct constructs, and an ad can score well on one while scoring poorly on the other, so liking should be read as a diagnostic input rather than a decision metric.
Predictive-validity research found that purchase-intent predictions selected the higher-performing ad correctly around 57% of the time in large samples, better than chance but not a guaranteed forecast. Combining intent with structured, evidence-based persuasion scoring improved accuracy to roughly 75% in matched comparisons.
Yes. Exposure-frequency modelling shows that repeated exposure changes attitudinal and behavioural response, so a single-exposure test can understate a campaign’s real effect if the media plan involves repetition.
Our Ad & Creative Testing service builds a metric battery around your specific commissioning decision, and our broader service list covers related capabilities such as brand tracking and concept testing for teams building a continuous measurement programme.