Marketing Analytics & Attribution
Know what you can know, and stop arguing about the rest
Perfect attribution does not exist. What does exist is a measurement setup you trust, a written account of where it is blind, and a way to test the questions attribution cannot answer. That is a better position than most businesses are in.
The starting position
Two things are true at the same time
Most businesses measure far less well than they could. And no business, at any budget, gets a complete picture of what caused a sale. Both of these are true, and a lot of wasted money sits in the gap between them.
The incompleteness is structural rather than a configuration error. Someone sees an ad on a phone, searches on a laptop three days later, asks a colleague, then converts from a link in an email. Consent choices remove part of that path entirely. Browsers cap how long identifiers survive. Platforms report on their own performance using their own definitions. Analytics products fill some of the gaps with modelling, which is reasonable and is still an estimate.
The wrong responses to this are equally common. One is to give up and decide marketing cannot be measured, which usually means budget goes to whoever argues best. The other is to buy a tool that claims to have solved attribution, which relocates the uncertainty into a black box and adds a licence fee to it.
What works is narrower and duller. Instrument the things you can instrument, properly and once. Write down where the data is blind so nobody has to rediscover it in a meeting. Triangulate between analytics, platform reporting, the CRM and what customers tell you directly. And when a question genuinely matters, test it rather than modelling it.
Two different questions
Attribution and incrementality are not competing methods
They answer different questions, and most disputes about marketing performance are caused by asking one of them to do the other one’s job.
| Dimension | Attribution | Incrementality testing |
|---|---|---|
| The question it answers | Which touchpoints appeared on the path to a conversion we managed to record | What would have happened if we had not run this at all |
| How it works | Joins recorded events to a conversion using identifiers, platform data and modelling | Withholds activity from a matched group or region and compares outcomes against a control |
| What it is good for | Daily optimisation, spotting breakage, understanding the sequence of a buying journey | Budget decisions, channel disputes, and anything where the size of the number matters |
| Where it misleads | Rewards whatever appeared last, which is usually the channel that harvested demand rather than created it | Needs enough volume and enough patience. An underpowered test produces confident nonsense |
| What it costs | Setup time, and ongoing maintenance as platforms and privacy rules change | Real money, in the form of demand you deliberately choose not to serve for a period |
| Minimum viable version | Correct GA4 events, offline conversion import, and one agreed definition of a lead | Turn one channel off in one region for a defined period and watch the whole business, not the channel |
The blind spots
Where the numbers stop being reliable
These are not faults to be fixed so much as properties to be known. A team that knows them makes better decisions than a team with a cleaner-looking dashboard and no idea what is missing from it.
- The report is full of "(other)" and hidden rows.
- GA4 applies cardinality limits, which bundle high-variety dimension values into an "(other)" row, and applies data thresholding that withholds rows when a report could identify individuals, which happens more often once Google signals or demographics are enabled. Neither is a bug. Both mean the number on screen is not the number you asked for, and a team that does not know this will chase differences that are artefacts.
- The comparison you need is older than the retention window.
- User and event level data in a standard GA4 property is retained for a limited period, with a default that is shorter than most people assume and a maximum that is shorter than a two-year comparison requires. Aggregate reports go back further, but exploration-style analysis does not. If nobody exported the detail at the time, the analysis is simply unavailable, which is one of the strongest arguments for a warehouse export.
- Consent choices removed part of the picture and nothing accounted for it.
- Where consent is declined, measurement is limited by design, and consent mode fills some of the gap with modelled conversions rather than observed ones. That is a legitimate approach and it is not the same as counting. Any report that mixes observed and modelled conversions without saying so will eventually be challenged by someone in finance, and they will be right to challenge it.
- Brand search looks like the best channel you have.
- Under last-click, the channel that captures demand takes credit for demand something else created. Brand search is the clearest example: it is cheap, it converts well, and cutting the activity that made people search for you in the first place will look free for about a quarter. This is the specific failure mode that costs the most money, because it looks like a good decision on the way in.
- Analytics and the CRM disagree, so nobody trusts either.
- Usually they are counting different things and no one has reconciled the definitions. Analytics counts a form submission; the CRM counts a lead that passed qualification. Both can be correct and they will never match. Until a single definition exists and both systems are instrumented against it, every performance conversation restarts from first principles.
The work
Four layers, built in this order
A measurement plan, before a single tag
Which decisions this data has to support, what a lead and a customer are, which events matter and which are noise. Written down and agreed with finance as well as marketing. Most failed implementations skipped this step and started with tags.
A collection layer that survives privacy changes
GA4 event design, consent mode implemented properly, server-side collection where the accuracy gain justifies the cost, and offline conversion import so closed revenue gets back to the platforms doing the bidding.
Reporting built around decisions
A report should answer what changed, why, and what we are doing differently. Where volume justifies it, a warehouse export so the detail still exists in two years. If nobody opens it, it is not reporting.
A testing cadence
Geo holdouts, platform lift tests where they are available and honest, and a self-reported attribution question on your forms. Planned as a recurring habit, because a test designed after the argument has started is already compromised.
Our position
What we will not tell you
We will never hand you a number that says exactly what caused a sale, because that number does not exist. We will tell you what we measured, how confident we are in it, and what we would test next.
That is a less comfortable pitch than a dashboard promising unified attribution across every channel. It is also the only version that survives contact with a finance director who asks how the number was produced.
In practice the businesses that make the best marketing decisions are not the ones with the most sophisticated models. They are the ones that agreed their definitions, know exactly where their data is blind, and run a small test before making a large change.
Questions
What finance and marketing both ask
Is GA4 good enough on its own?
For most businesses it is a reasonable foundation and a poor single source of truth. It is a modelled, privacy-constrained product: it applies thresholds that hide small segments when certain features are enabled, it groups high-cardinality values into an "(other)" row, and user and event level data are retained for a limited window by default.
Configured well and understood honestly, it answers a lot. The mistake is treating it as a ledger. It is an estimate, and the estimate is better for trends than for absolute numbers.
Will server-side tracking fix our data loss?
Partly, and not in the way it is usually sold. It gives you control over what is collected and sent, it is more resilient to browser restrictions on cookie lifetime, and it improves the quality of the events you forward to advertising platforms.
What it does not do is recover data from people who declined consent, and anyone presenting it as a way around consent requirements is describing a compliance problem rather than a measurement solution. It also costs money to run, so we recommend it when the accuracy gain justifies that, not by default.
Should we add "How did you hear about us?" to our forms?
Yes, in most cases. Self-reported attribution is imprecise, people misremember, and it will not reconcile neatly with your analytics. It is also the only signal that can see the podcast, the recommendation and the conversation at a conference, none of which will ever appear in a tag.
Use it as one input among several rather than as the answer. Where it consistently disagrees with your analytics, that disagreement is itself useful information about where your measurement is blind.
How do we decide budget if attribution is unreliable?
By testing rather than by reading. Turn a channel down or off in one matched region for a defined period, watch total business outcomes rather than that channel’s reported numbers, and compare against a control. That answers the budget question directly in a way no attribution model can.
Between tests, use attribution for what it is good at: spotting breakage, understanding sequence, and day-to-day optimisation. The mistake is asking a model built for the first job to settle an argument that belongs to the second.
We do not have much volume. Does any of this apply?
The instrumentation absolutely does. Correct event design, agreed lead definitions and offline conversion import matter more at low volume, not less, because every misattributed lead is a larger share of the total.
Formal statistical testing usually does not. At low volume a geo holdout will not reach significance in a sensible timeframe, so the honest approach is careful instrumentation, self-reported attribution and larger, slower changes that are easier to observe.
Find out how far your current numbers can be trusted
We will review your GA4 configuration, your consent setup and how your reported conversions compare with your CRM, then tell you which of your numbers are safe to make decisions on and which are not.
If we don't deliver the work we agreed to deliver for reasons within our control, you don't pay for the undelivered work. Read our guarantee
Related
Where to go next
- growth and lifecycle services
- paid media budgets that depend on itBidding is only as good as the signal underneath it.
- Google Ads management
- organic search performanceWhere attribution gaps are widest and least discussed.
- testing changes on siteTesting needs a measurement layer that can be trusted.
- diagnosing a traffic drop
Last updated · Reviewed by Zubair Afzal