The Benchmark Advertising Still Needs

I recently came across a study, Meta-Analysis of Advertising Effectiveness: New Insights from Improved Bias Corrections, by Joseph Korkames, et al. It arrives at an average short-term advertising-to-sales elasticity of 0.0008, a figure statistically indistinguishable from zero.

I found it through a LinkedIn post by a marketing consultant who summarized the finding this way: How much can we expect sales to increase if we significantly increase advertising? His answer was “close to zero.”

That is a fair shorthand for the headline estimate, but it leaves unclear what the number represents. Advertising elasticity measures the marginal effect of changing spending, typically within an existing advertising program. An elasticity of 0.0008 means that a 10% increase in advertising spending is associated with a 0.008% increase in sales. For a company with $10 billion in sales and a $1 billion advertising budget, another $100 million in advertising would produce just $800,000 in additional short-term revenue.

Without that context, the finding can easily be mistaken for evidence that advertising itself has virtually no effect. But it concerns dollars added to an existing program, where diminishing returns may already have set in. It does not tell us advertising’s total contribution to sales, the value of the first dollars spent, what would happen if an established brand stopped advertising, or whether advertising “works.”

The study effectively updates research summarized in Dominique Hanssens’s influential Empirical Generalizations About Marketing Impact. That work included the well-known study How Well Does Advertising Work? Generalizations from Meta-Analysis of Brand Advertising Elasticities, which produced an average elasticity of approximately 0.12. Applied to the same hypothetical company, a 10% increase in spending would generate $120 million in additional sales, rather than $800,000.

The newer study uses more recent methods to correct for publication bias, a real and important problem. But its estimate is not a direct observation of advertising’s effect. It is one model applied to estimates produced largely by other models, spanning different products, media, markets, periods and definitions. It requires assumptions about which findings are comparable, how they should be weighted and how publication bias should be corrected.

The enormous difference between 0.12 and 0.0008 may tell us less about advertising than about the difficulty of constructing universal benchmarks from an inconsistent body of research.

I was reminded of a lecture I recently attended by Aaron Brown, a financial and risk-management expert, about his book Wrong Number. Brown examines how the media, politicians and others become captivated by surprising statistics even when the underlying science is dubious. A striking number travels much farther than the qualifications needed to interpret it.

As Twyman’s Law puts it, any figure that looks particularly interesting or different is usually wrong.

Marketing has no shortage of such numbers. A campaign produced a 500% return. Long-term advertising is many times more powerful than short-term advertising. Advertising’s short-term effect rounds to zero.

The trouble begins when the number becomes detached from the method that produced it.

Science, to its credit, has begun some serious house-cleaning. The replication crisis has exposed p-hacking, HARKing, selective reporting, publication bias and irreproducible findings. Journals have retracted prominent studies. Researchers increasingly preregister hypotheses, share data, report null findings and distinguish exploratory analysis from confirmatory testing.

Marketing should pay attention.

About a decade ago, I watched a celebrated marketing thought leader present general principles derived from a database of advertising-effectiveness award winners. He was a gifted showman, thanks in part to freely swearing like a sailor, and the audience loved him.

I could not get past the evidence.

These were campaigns considered unusually successful by the organizations entering them, supported by whichever measurement methods were available, and then selected by judges as worthy of recognition. The cases might offer useful lessons about successful campaigns. Treating them as representative norms for advertising generally did not pass the sniff test.

Around the same time, I was invited to outline a research roadmap for one of the advertising industry’s leading trade organizations. My central recommendation was a living database of advertising-effectiveness results.

Advertisers, publishers, agencies and measurement vendors would contribute anonymized campaign results on an ongoing basis, with participation tied to membership benefits. The database would include winners, losers and inconclusive findings, reducing the publication bias created when only impressive results are shared.

Ideally, APIs into measurement platforms would automate submissions, regardless of whether the outcome was high, low or nonexistent. Identifying information would be removed, and findings would be released only in rule-based aggregate form.

Every result would carry the equivalent of a methodological nutrition label describing:

  • Product category, purchase cycle, brand size and maturity

  • Campaign objective, media, spending, reach and duration

  • Creative characteristics

  • Outcome measured and measurement period

  • Research design, model assumptions, statistical power and uncertainty

  • Independence of the measurement

Most importantly, dissimilar forms of evidence would not be silently blended. Randomized experiments, quasi-experiments, marketing mix models, attribution studies and brand-lift surveys would be classified separately. Incremental sales would be distinguished from attributed sales, causal effects from associations, observed outcomes from modeled estimates, and immediate response from assumed long-term carryover.

Had the industry built this resource a decade ago, we might know considerably more today about what advertising accomplishes. We could examine not only how effectiveness varies across categories, brands and campaigns, but how reported effectiveness varies according to the method used to measure it.

Several standardization efforts are now advancing across advertising, including AdCP, ARTF, IAB taxonomies, DASH, W3C Attribution and initiatives to validate and benchmark MMM. Each addresses a worthwhile part of the problem. A shared framework for classifying evidence would make them more useful by providing a common language for describing what was measured, how it was measured and how much confidence the result deserves.

It would also address the problem I raised in my essay on enterprise AI and marketing effectiveness: garbage in, garbage out. AI systems will absorb decades of MMM outputs, attribution reports, brand-lift studies, award cases and experimental results as though they form a coherent body of knowledge. Without methodological context, AI cannot reliably distinguish causal evidence from favorable correlation. The next generation of automated, AI-powered media planning and buying systems will then use that deeply flawed body of evidence to make optimization decisions, perpetuating past measurement errors at machine speed and scale.

Advertising does not need another universal number. Cookies and cars should not share one response benchmark. Neither should new brands and household names, strong creative and weak creative, or a short promotion and decades of sustained brand building.

We need the infrastructure to answer a more useful question: What kinds of advertising work, for which brands and products, under what conditions, over what period, according to which methods, and with what degree of confidence?

We have accumulated decades of findings. We still have not organized them well enough to know which ones deserve our trust.

Previous
Previous

Bringing Experimentation to the Ad Context Protocol

Next
Next

Timeless Truths About Advertising, Part III