Testing · 05
Testing and measurement
A century before A/B testing had a name, mail-order houses were splitting mailings between issues, sorting reply cards by date, and retiring campaigns that did not pay their way.
Direct-response marketing is the discipline that made measurement mandatory. Because every campaign asks for a countable reply, every campaign can be judged. And once judgment is possible, testing becomes obvious: run two versions, count the replies, keep the one that produced more. The mail-order operators did this by hand, patiently, for decades, and the habit — not the software — is what matters.
What is actually worth testing
Not all tests are equal. Some elements of a campaign move the needle by ten or twenty per cent when changed; others move it by fractions. A working priority order, sharpened over decades of practice:
- The audience. Who the message is shown to. Almost always the largest lever.
- The offer. Price, terms, guarantee, bonus. The second-largest lever.
- The headline and opening. Whether the message is read at all.
- The proof. The evidence that makes the promise believable.
- The medium and format. Where and how the message reaches the reader.
- The small stuff. Button colours, punctuation, subject-line emojis. Real, but modest.
Marketers who spend their testing budget from the bottom of that list get a career of small, unremarkable improvements. Marketers who spend it from the top get results that show up in the accounts.
The numbers that matter
A useful campaign can be described with a handful of numbers, and a harmful one can be hidden behind a hundred. The direct-response tradition keeps its dashboards short:
- Response rate. Actions per thousand impressions, or per hundred visitors.
- Cost per response. What each qualified reply costs to acquire.
- Cost per customer. What each paying customer costs, after refunds.
- Average order value. What a customer is worth on the first transaction.
- Lifetime value. What a customer is worth over the relationship — the number that decides how much can be spent to acquire them.
Statistical honesty
The temptation of testing is to declare victory early. A version that pulls ahead in the first hundred replies is not necessarily better; it is often merely lucky. The old operators had a straightforward discipline: never call a test until the sample is large enough that the difference could not reasonably be chance. In practice, this means running tests until at least several hundred conversions have accumulated in each group, and being cautious about small percentage differences on small samples.
The most common testing mistake is not the maths. It is running so many tiny tests, so quickly, that noise is mistaken for signal, and a year is spent chasing improvements that would not survive being re-run.
Test to learn, not to prove
Every test is a small experiment about the world. The winner matters, but the reason it won matters more, because the reason travels to the next campaign. A team that tests to prove a preferred idea learns very little. A team that tests to learnwhatever the data has to say builds, over time, a private working theory of its own market. That theory is the compounding asset of the discipline — worth more than any single winning ad.