Multi-touch attribution isn't a causation engine. It's a correlation engine — and mistaking one for the other can cost a marketing org millions.
A quick scope note: this piece is about standard multi-touch attribution, the rule-based or ML models most teams run day to day to credit touchpoints along a journey. It's not about newer incremental attribution approaches, which build causal experimentation directly into the model. That's a different, and improving, category worth its own piece.
Every marketing team I've collaborated with has had a channel that they take pride in. An impressive campaign. Strong numbers with a high-impact presentation for the quarterly business review. Clear attribution with high ROAS that makes reporting to the executives easy.
For our team, that channel was Google brand campaigns.
From our viewpoint, that was solid MTA data. Brand campaigns attributed to a large proportion of conversions. There was a lot of sense in that. Someone views a brand campaign and then goes on to convert, the model assigns attribution. Easy to understand. We were satisfied with that.
Then came the geo holdout.
In a number of test markets, brand spend was turned off and everything else was kept the same. The results of the holdout contradicted everything the MTA stated. Brand was attributed to a positive lift of conversions, but in reality, those conversions would have taken place anyway. We were not driving conversions, just showing up next to them and taking the attribution.
Here is what no one — or sometimes vendors — forget to tell you when building an MTA: it is a channel tracker. It is not a causation engine.
Multi-touch attribution is a correlation engine. It tracks the customer journey map, notes the conversion touchpoints, and assigns attribution based on the model you have built (e.g. linear, time decay, data driven, etc.). More advanced models use machine learning to provide better accuracy around touchpoints — I've stood one up in the past.
This is not a correlation/causation issue.
Brand campaigns reach people already in the market who will convert. When customers see an ad, the model captures it, credits it, and shows a large positive brand ROAS. The model, however, cannot show whether, if that ad was not shown, the customer would still convert.
This is not an issue with the way you set it up. It is a structural issue with what attribution can show.
Conceptually, a geo holdout test is quite simple. Take two sets of markets that are comparable, run your campaign as you normally would in one set, go dark in the other set, then analyze the results. What you're left with is incrementality — the lift that happened because of your marketing.
It answers a question MTA simply cannot: would you get this conversion if you did not go to market?
When we executed a geo holdout for our brand campaigns at StockX, we found that most of the conversions attributed to brand campaigns were also happening in the holdout markets — with no brand campaigns run at all. The incrementality was significantly less than what was being attributed.
Nobody was happy with the results of the geo holdout. However, it was the most accurate assessment we could get on this marketing channel.
The answer isn't to throw out your MTA. It's to stop asking it to do something it wasn't designed to do.
MTA is useful for directional channel comparisons, optimizing within a channel, and understanding the customer journey at a tactical level. It's fast, granular, and gives you something to work with day to day — optimizing campaigns or generating hypotheses for testing.
Asking MTA to determine whether a channel drives incremental business outcomes is outside its potential, regardless of whether it's ML-based or rule-based. For that, incrementality testing should be used alongside MMM (media mix modeling).
MTA describes what happens along the journey. Incrementality testing describes what causes that to happen. MMM describes what will happen in the future.
The best way to look at them is as a triangulated approach. When the three operate in conjunction, you can be very confident in a decision. When they contradict each other, that's exactly where the most important analysis should happen.
Some orgs avoid running geo holdouts on certain channels because they don't want to discover the truth, or risk losing sales during the test window. Some campaigns — brand, in this case — are costly, highly visible, have internal champions, and are generally perceived as untouchable.
But the alternative is making budget decisions worth millions based on a model that has never been tested, offering a plausible — but potentially wrong — narrative.
This is, in a way, a question of marketing measurement maturity: being willing to go through some short-term pain to gain insight that actually optimizes your budget and generates new demand, instead of just capturing what was already there.
Your MTA is not lying. It is showing you correlation, and calling it causation.