Timeless Truths About Advertising, Part III

How We Know Advertising Worked

In the first two essays in this series, I argued that advertising follows remarkably consistent patterns. Reach usually beats frequency. Advertising builds memory that fades over time. Brands grow by reaching more buyers, and every buying occasion represents a new opportunity to compete.

Understanding those patterns is only half the challenge. Marketing doesn't have a measurement problem. It has a decision problem. The industry has no shortage of metrics, models, dashboards, and reports. What marketers, CFOs, and CEOs need is evidence they can trust when deciding where to invest the next advertising dollar. The ultimate purpose of measurement isn't to produce reports. It's to improve decisions.

That raises perhaps the most important question in advertising: How do we know whether advertising actually caused incremental sales? The enduring truths that follow are about separating evidence that guides better decisions from evidence that merely describes what happened.

1. Correlation does not imply causation

The purpose of advertising is simple: to cause more people to buy your product than otherwise would have.

That makes advertising measurement fundamentally a question of causality.

Yet many of the industry's most common measurement techniques remain grounded primarily in correlation. They identify patterns associated with sales rather than demonstrating that advertising caused them.

Correlation is often useful. It can improve forecasts, guide optimization, and generate hypotheses. But it cannot distinguish whether advertising created demand or merely found consumers who were already likely to purchase.

That distinction matters. The purpose of advertising is not to predict who is about to buy. It is to persuade people to buy who otherwise would not have. Prediction and persuasion are fundamentally different. Confusing the two can lead marketers to overstate advertising's contribution. Marketing decisions should therefore be based on incremental sales, not coincidental ones.

2. Quasi-experiments are not experiments

Many modern measurement methods attempt to approximate randomized experiments by carefully constructing comparison groups. Matched markets, synthetic controls, Bayesian Structural Time Series (BSTS), propensity score matching, and numerous other techniques all belong to a family of methods known in science as quasi-experiments. The defining characteristic of a quasi-experiment is that the treatment and control groups were not assigned randomly.

Randomization is what makes randomized controlled trials the scientific gold standard. By giving every unit an equal chance of receiving the treatment, it neutralizes both known and unknown sources of bias. Rather than trying to model away differences between test and control groups after the fact, randomization makes those differences equally likely to exist on either side in the first place.

The first rule of quasi-experimental research is to use it when a randomized experiment would be unethical or infeasible. In medicine, it would be unethical to randomly assign people to smoke cigarettes or drink heavily for decades. In advertising, ethical constraints are rarely the issue. Practical constraints sometimes are. National sponsorships, creative already in market, contractual obligations, or channels that cannot be geographically varied may make experimentation difficult or impossible.

But many advertising decisions are amenable to randomized controlled trials, particularly through cluster randomized "geo experiments" that randomize geographic markets rather than individuals. Yet many organizations begin with quasi-experimental methods without seriously considering whether a randomized design is feasible. Familiarity, organizational inertia, perceived operational complexity, or existing investments in modeling tools are understandable reasons to favor quasi-experimental methods. They are not, however, scientific reasons to prefer them over a feasible randomized trial.

When millions of advertising dollars, future revenues, and competitive market share are at stake, the first question should never be, "Which quasi-experimental method should we use?" It should be, "Can we randomize?" Only when the answer is no should quasi-experimental methods become the preferred alternative.

3. Good measurement answers the CFO's question

Not every experiment answers the same question.

Many platform experiments measure how one creative performs against another, or how advertising influences a narrowly defined audience already active on that platform. Those are useful questions.

They are not usually the question the CFO is asking.

The CFO wants to know something much simpler: If I move my next advertising dollar from Channel A to Channel B, which investment will generate more incremental sales?

That requires measuring business outcomes at the level where business decisions are made, not simply measuring behavior within a platform or among a selected audience.

Consider three experiments. Facebook reports a 10% sales lift. TikTok reports a 5% lift. Linear TV reports a 2% lift. At first glance, Facebook appears to be the clear winner.

But those percentages are not directly comparable. Facebook may have deliberately targeted a relatively small audience already exhibiting strong purchase intent, while the TV campaign reached a broad audience with much lower baseline purchase probabilities. A 10% lift among consumers already close to buying is not necessarily more valuable than a 2% lift across a much larger population. Add differences in media costs, and the comparison becomes even more misleading. Each experiment answers the question, "How did this campaign perform within its own context?" None answers the CFO's question: "Which investment generated the greatest incremental return for the business on a comparable basis?"

This is also the question Marketing Mix Models are intended to answer. An MMM estimates the incremental contribution of each marketing channel so budgets can be allocated efficiently across the portfolio. If experiments are used to calibrate or validate an MMM, they should answer that same enterprise-level question. Platform-specific experiments can be excellent tools for optimizing campaigns within a channel, but they are poor calibration data for models designed to compare investments across channels.

The distinction is one of efficiency versus effectiveness. Platform experiments are excellent at improving the efficiency of a channel, helping marketers target the right people, optimize creative, and reduce the cost of acquiring customers within Facebook, TikTok, or another platform. The CFO's problem is different. It is one of effectiveness: deciding which channel deserves the next advertising dollar. Optimizing execution within a channel is not the same as optimizing the allocation of the marketing budget across channels.

The most valuable measurement systems therefore focus on incremental business outcomes at the level where marketing decisions are made, rather than campaign diagnostics at the level where media happens to be delivered.

4. Identity is not necessary for effective measurement

Much of digital advertising is premised on the idea that tracking people is the key to measuring advertising performance.

Often it isn't.

Precision and accuracy are not the same thing. User-level data appears more precise because it records individual impressions, clicks, and conversions. But greater granularity does not necessarily produce more accurate estimates of advertising's incremental effect. In practice, user-level measurement often introduces additional noise through incomplete identity graphs, imperfect match rates, attribution rules, and other assumptions that obscure rather than clarify cause and effect.

The objective is not to reconstruct every consumer's journey. The objective is to determine whether advertising increased sales.

Large-scale geographic experiments can answer that question without identity graphs, clean rooms, household matching, or complex data joins. They also avoid much of the cost, computational overhead, and privacy burden that accompany those technologies. If media can be targeted geographically, and sales can be measured geographically, then advertising effectiveness can often be measured at exactly that same level.

Geo experiments work because they measure outcomes directly rather than attempting to connect millions of individual exposures to millions of individual purchases. The required ingredients are remarkably simple: media that can be geographically targeted, and sales already summarized by postal code, DMA, or another geographic unit. By avoiding the lossy process of matching identities across disparate datasets, geo experiments often produce cleaner estimates with fewer assumptions and less noise.

Sometimes the clearest picture comes from stepping back, not from inspecting every brushstroke. 

5. Advertising's sales lift is usually smaller than marketers imagine

One of the most persistent misconceptions in marketing is that successful advertising should produce dramatic increases in sales.

It rarely does.

Across decades of research, advertising has consistently been found to have relatively small sales elasticities. As Donald Lehmann observed in the foreword to Empirical Generalizations About Marketing Impact, a 20% increase in advertising spending typically produces less than a 1% increase in sales

That doesn't mean advertising is ineffective.

A one-percent increase in sales for a large national brand can represent tens or even hundreds of millions of dollars in incremental revenue. Because advertising budgets are usually only a small fraction of sales, even modest sales gains can generate exceptional returns on investment.

The mistake is to confuse sales lift with ROI. A campaign can produce only a one-percent increase in sales and still be enormously profitable.

These small effects also have profound implications for measurement. Detecting a one-percent lift requires experiments with sufficient statistical power, careful design, and enough observations to separate genuine signal from ordinary business noise. Methods that rely on noisy observational data, small matched samples, or heavily modeled counterfactuals often lack the sensitivity to measure effects of this magnitude reliably. When the expected lift is only one or two percent, weak measurement methods are more likely to produce unstable estimates than trustworthy answers.

The best experiments are designed from the outset to detect exactly these modest, incremental changes, because those are often where advertising creates its greatest economic value.

Advertising has become extraordinarily sophisticated. Measurement has become even more sophisticated. But the central question has never changed.

Did advertising cause more sales than would otherwise have occurred?

Every methodology should ultimately be judged by how confidently it answers that question.

The better our evidence, the better our decisions.

Next
Next

Timeless Truths About Advertising, Part II