Your Ad Reporting Is Taking Credit for Sales You Already Had
The short answer. When eBay stopped bidding on its own brand name on Yahoo and Microsoft's search engines, 99.5% of those clicks came straight back through organic results on the same platform. When researchers compared dashboard-style analysis against controlled experiments on the same Facebook campaigns, the dashboard-style number was too high in all 15 cases. This is our own industry and we sell paid media, so read the section on what we are not saying before you cancel anything. Ads do work. The reporting is what is broken.
Every agency report has a number on it that says what the ads did. Most of them are not measuring what the ads did. They are measuring which ad happened to be nearby when someone bought something, which is a different question with a much more flattering answer.
This is the fourth post in our evidence series. Same rule as the others: every study named, checked against the primary source, with the limits stated, including the limits that cut against us.
The experiment that should have ended the argument
In 2012, eBay's own economists ran the test almost nobody runs. They turned the ads off and watched what happened to sales. The results were published in Econometrica, one of the most demanding journals in economics (Blake, Nosko and Tadelis, 2015).
First they stopped bidding on their own brand name, queries containing the word "eBay", on Yahoo and Microsoft's search engines, while deliberately continuing to pay for those terms on Google so Google could act as a control.
99.5%
Of the clicks lost by halting brand-name paid search on Yahoo and Microsoft, recaptured by organic results on the same platform (Blake, Nosko and Tadelis, Econometrica, 2015).
The paid clicks had been almost entirely substituting for free ones. People searching for eBay by name were going to reach eBay. The ad was a toll booth on a road the traffic was already taking.
Then they ran a larger test on non-brand keywords, switching off bidding across about 30% of US media markets for 60 days. Attributed sales in the switched-off markets fell by more than 72%, exactly the collapse a dashboard would scream about. Total sales barely moved: the whole paid search program added 0.66% to sales, with a confidence interval running from slightly negative to slightly positive.
Two things that matter before anyone quotes this at their agency. Substitution was nowhere near total here: paid clicks fell 41% while total clicks fell only 2%, meaning roughly 42% of those paid non-brand clicks were traffic eBay would not otherwise have received. The 99.5% figure is a branded-search result and does not transfer. And the assignment was not a clean coin flip, markets were matched algorithmically on historical sales patterns.
Still, the gap between the two ways of measuring the same campaign is the point. Measured naively, the way a report does it, that campaign returned over 4,100%. With statistical controls, over 1,400%. Measured by the experiment, negative 63%, with a 95% confidence interval running from minus 124% to minus 3%, and computed from imputed public revenue and spend figures rather than eBay's internal books. The direction is solid. The decimal place is not.
The dashboard is the thing being measured wrong
The obvious objection is that eBay is one company in one year. So here is a different platform and a different failure mode, landing in the same place. Researchers ran 15 large advertising experiments on Facebook, then calculated on the same people what standard non-experimental methods would have reported (Gordon, Zettelmeyer, Bhargava and Chapsky, Marketing Science, 2019).
Comparing people who saw the ad against people who did not, the closest analogue to what a platform reports back to you, overstated the effect in every one of the 15 studies. Not most. All of them. In several cases the two numbers were not in the same postal code: 377% against a real 1.3%, 392% against 8.6%, 233% against 1.2%.
Those 15 studies were hand-selected by two of the authors rather than drawn at random, which is exactly why the same lead author went back with a representative sample: 663 experiments and 7.9 billion observations, using debiased machine learning and stratified propensity score matching (Gordon, Moakler and Zettelmeyer, Marketing Science, 2023). It still could not reliably recover the true causal effect.
Disclosure, because it matters in both directions: two of the four authors on the 2019 paper worked at Facebook, a co-author on the 2023 paper is from Meta Ads Research, and both run on Meta's own experiments. They had access no outsider gets, and their employer sells the experiment-based product this research recommends.
The natural response from anyone in our industry is that better modelling fixes it. Not with the data you have. The 2023 paper does point to a method that recovers the experimental answer, but it requires bid-request-level logs that no major platform exposes. Note also what this research does not cover: it tests statistical adjustment methods, not the multi-touch attribution products sold to advertisers, which nobody has tested at this scale.
Everyone is paid on the flattering number
A randomized experiment with 208,538 consumers at an online retailer looked at what the advertising platform was actually being paid for (Frick, Belo and Telang, Management Science, 2023).
46%
Share of the fees the ad platform collected that corresponded to purchases which would have happened anyway (Frick, Belo and Telang, Management Science, 2023).
Read that correctly, because we are not going to do to this study what we are accusing dashboards of doing. The ads worked: treated shoppers were more likely to purchase, by 0.24 percentage points, and considerably more likely to return to the site, by 1.78 points. The finding is not that the ads did nothing. It is that the fee structure billed for a lot of buyers the ads did not create. The researchers also found no relationship between how likely a shopper already was to buy and how much the ads moved them, which is what you would expect from a system optimizing for credit rather than persuasion.
Caveats, stated: one campaign, one European retailer, one contract type. Its control group used charity ads, a design other researchers in this field argue is unreliable when a platform optimizes delivery, because the platform will not serve two ad types to the same kinds of people. And the purchase effect underneath that 46% sits right at the edge of statistical significance. We cite it for the incentive it documents, not because one campaign settles anything.
That incentive does not stop at the platform. Any agency fee that rises with ad spend, or with the count of conversions a platform attributes, pays more when the flattering number goes up. Ours included. Changing which number sits at the top of a report does not fix that, and we are not going to pretend it does. If you want to know whether your agency's incentives point at your results or at your spend, ask what happens to their invoice when your spend drops 30%. Ask us the same question.
What we are not saying
This is where contrarian marketing posts usually overreach, so here is the ledger against us.
Ads work on people who do not already know you. In the eBay experiment the largest effect, roughly 10%, was on users who had never purchased before. For frequent recent buyers it was near zero. The authors' conclusion was not that search advertising fails, it was that it works when the customer does not already know you have the thing. If you are not a household name, you are in the group the ads helped.
Do not switch off your brand search. eBay had a rare luxury: no competitor bidding on its brand terms was observed. Later work found that when a brand stops bidding and rivals do not, competitors capture roughly 18% to 42% of those clicks, and defensive brand bidding shows strongly positive returns (Simonov, Nosko and Rao). Copying eBay's move without eBay's conditions will cost you money.
Retargeting produces real lift under clean measurement. A ghost-ads design, where the control group is identified without being shown a substitute ad, found retargeting lifted website visits 17.2% and purchases 10.5% (Johnson, Lewis and Nubbemeyer, Journal of Marketing Research, 2017). That was one advertiser, a two-week campaign in winter 2014, on roughly $30,500 of spend, and the design is Google-authored and runs as Google platform infrastructure. Later work reports retargeting returns have fallen since app-tracking and privacy changes (Larson and Dotson, 2026), and at least one experiment with a true no-ad control found no incremental purchase lift. Note the shape too: the visit number is far larger than the purchase number, and quoting a revisit figure as a conversion figure is the most common exaggeration in this category.
The evidence is weakest exactly where our clients live. eBay is a giant with dominant brand recall and the top organic result for its own name. The measurement-difficulty research studies firms with huge, noisy sales volumes. Those authors explicitly carve out small and new businesses. So be precise about what transfers to a smaller advertiser: the mechanism does, credit for buyers you already had, view-through windows that flatter, brand terms where you already rank first. The specific magnitudes do not. Nobody has run these experiments on a business your size, us included.
Much of this data is from an older internet. The Lewis and Rao display experiments ran 2007 to 2011, eBay's tests in 2012, the retargeting experiment in winter 2014, the Facebook comparison on 2015 campaigns. All before app tracking changes and cookie deprecation reshaped measurement. The 2023 work is more recent and finds the same problem, which is the strongest answer to the argument that this is all an old-internet artifact.
So what should you actually do?
The real answer is less satisfying than the pitch you are used to.
Experiments are the only clean answer, and they are often out of reach. Across 25 display advertising experiments run on Yahoo between 2007 and 2011, with 19 major retailers and 6 financial services firms, every one with more than half a million users and most with over a million, the median confidence interval on return was more than 100 percentage points wide (Lewis and Rao, Quarterly Journal of Economics, 2015). Experiments remove the bias. They do not deliver a tidy number. One update in fairness: Lewis later co-authored the ghost-ads work arguing better holdout design substantially improves those economics, letting an advertiser measure the same lift for roughly an order of magnitude less spend. The catch is that ghost ads are platform infrastructure you opt into, not something you or we can build.
Use coarse tests instead of precise fictions. Long on and off windows on one channel. Blended spend against total revenue. Whether the phone rings more. These are crude, and crude and true beats precise and invented. At small budgets a test will be underpowered, and an underpowered result showing nothing is not proof the channel does not work.
Watch the shape of the claim. Ask which conversions were view-through rather than click-through, and how long the attribution window is. A 30-day view-through window credits the ad for nearly everything that moves.
Judge us on the boring number. Total revenue against total marketing spend, over quarters, next to what happened before we started. It is a worse story to tell in a meeting and a better one to run a business on.
What this changes about our own reporting
Three specific changes. Platform-reported return on ad spend stops being the headline number on our reports and moves to a supporting line, labelled as what the platform claims. Blended spend against total revenue goes on top. And where a client's scale allows a real holdout test, we run it, including when it risks making our own work look smaller.
What that does not do is resolve the fee question above, and no report layout can. If that costs us a client who preferred the flattering dashboard, that is the trade.
Frequently asked questions
Does paid search advertising actually work?
Paid search works on people who do not already know you, and works much less on people who were going to buy anyway. In eBay's controlled experiment the largest sales effect, roughly 10%, was among users who had never purchased before, while the effect on recent frequent buyers was near zero. The problem is not that ads do nothing, it is that spend concentrates on the audience least moved by it.
Should I stop bidding on my own brand name?
Probably not. eBay recovered 99.5% of those clicks organically when it stopped bidding on Yahoo and Microsoft, but no competitor bidding on eBay's brand terms was observed. Later work found rivals capture roughly 18% to 42% of the clicks when a brand stops bidding and competitors continue, and defensive brand bidding shows strongly positive returns. Check whether anyone is bidding on your name before changing anything.
Why does my ad dashboard show a better return than reality?
Your dashboard credits conversions that happened near an ad rather than conversions the ad caused, and the platform deliberately targets people most likely to convert anyway. When researchers compared the closest observational analogue, exposed users against unexposed users, against experimental truth on the same 15 Facebook campaigns, it overstated the effect in all 15.
Can better attribution modelling fix this?
Not with the data available to advertisers. A 2023 study applied debiased machine learning and stratified propensity score matching across 663 advertising experiments and 7.9 billion observations and still could not reliably recover the true causal effect. One method does recover it, but requires bid-request-level logs no major platform exposes. That study was run on Meta's own experiments with a Meta co-author, and it tests statistical adjustment methods rather than the multi-touch attribution products sold to advertisers.
What is an incrementality test, and can a small business run one?
An incrementality test measures sales against a group deliberately held back from seeing your ads, so the comparison is causal rather than correlational. At small budgets it is usually underpowered: even experiments at major retailers reaching hundreds of thousands of users produced confidence intervals over 100 percentage points wide. Smaller advertisers are better served by long on and off periods on one channel and by tracking blended spend against total revenue.
Related reading
- Five marketing myths peer-reviewed research says are wrong, the first post in this series.
- Colour psychology, fact-checked, where the famous statistics turned out to have no study behind them.
- Stop sending ad clicks to your homepage, the cheapest fix in most ad accounts.
- How we run performance ads, including what we report and what we refuse to claim.
Measurement, honestly
Ask us what your ads actually did.
We will read your ad account and tell you which results look incremental and which look like credit for sales you already had. If the answer is that your spend is buying customers you would have kept anyway, we would rather say so than bill against it.
Talk to us