Skip to content
Skayle Marketing

measurement · 12 min read

There is no good conversion rate. There is only yours, measured properly

Every benchmark table you have been shown is built on a population you cannot see, a denominator defined differently from yours, and a period nobody stated. Here is why the comparison cannot work, the method for deriving a baseline from your own data, and the arithmetic for telling a real change from noise.

Written by , FounderUpdated

The question has no honest general answer, because a conversion rate is not one quantity. It is a ratio between two things that every business defines for itself.

That is not a dodge, and it is not the preamble to a benchmark table further down the page. There is no table on this page. What follows is why the comparison you were about to make cannot work, and then the method for producing the number that would actually help you, which is a baseline derived from your own data.

The good news is that the derived version is more useful than the borrowed one would have been. A benchmark tells you where you sit against businesses you know nothing about. A baseline tells you when something in your own business changed, which is the only conversion-rate question that has ever led to a decision worth making.

The problem with the tables

Three defects, any one of which is fatal

The first defect is the population. Nearly every published conversion benchmark is computed from the customers of the company publishing it — an analytics vendor, a platform, an agency. That is a self-selected group, not a sample: businesses that bought a particular tool differ systematically from those that did not, in size, sophistication, sector and traffic mix. Averaging them produces a real number about a group you are not in and cannot inspect.

The second is the denominator. A conversion rate divides conversions by something, and that something is usually sessions, which is a configurable rule rather than a natural unit. Google Analytics ends a session after thirty minutes of inactivity by default, and that timeout can be changed. Cross-domain journeys, consent choices, bot filtering and internal traffic exclusions all change the count. Two businesses with identical visitor behaviour can report different session totals because their configurations differ, and the ratio moves accordingly.

The third is the numerator, and it is the one nobody mentions. What counts as a conversion is a decision. One business counts every form submission, including the recruiters and the suppliers. Another counts only enquiries its sales team qualified. A third counts newsletter signups. Advertising platforms add their own layer, since a conversion can be counted once per click or every time it occurs. These are not variations in accuracy. They are different quantities with the same name, and a table averaging them is averaging apples with the concept of fruit.

A filter

Six questions to put to any benchmark before using it

If a source cannot answer all six, the number cannot be compared with yours. This is not a high bar; it is the minimum required for a comparison to mean anything.

Apply these to the next benchmark you are shown, including any we might ever publish ourselves.

  • Who is in the population, and how did they get there? Customers of the publisher is a different thing from a sample of the sector, and only one of them supports a general claim.
  • How many businesses and how many sessions does the figure rest on? A median across eleven accounts and a median across eleven thousand are both quotable and only one is informative.
  • What exactly is the denominator? Sessions, users, visits or something else, and under whose configuration. Without this the ratio cannot be reconstructed.
  • What exactly is the numerator? Which events counted as a conversion, and were they counted once per visitor or every time they occurred.
  • What period does it cover, and does it include a peak trading season? A figure spanning a promotional quarter describes that quarter.
  • Can it be filtered to a segment that matches yours? Device, channel, country and page type all move the rate more than sector does, so a sector average without those splits is too coarse to act on.

The method

Deriving a baseline from your own data

Five steps. None of them needs a tool you do not already have, and the whole thing takes an afternoon plus a year of patience.

  1. Write the two definitions down

    One sentence for the numerator: exactly which events count, and whether they count once per person or every time. One sentence for the denominator: which sessions are included, which are excluded, and what the session timeout is set to. Until both are written, two people in the business will calculate this differently and both will be right.

    You get: A one-paragraph definition, stored where the report is built

  2. Choose the segments before you look at any numbers

    Decide in advance which splits you will hold constant: channel, device, country and page type are the four that move a rate most. Choosing after seeing the data is how a business ends up with a segment that exists because it happened to look good. Pick the splits that correspond to decisions you might actually take.

    You get: A fixed segment list, agreed before the first report

  3. Take a window long enough to contain your cycle

    Long enough to cover a full purchasing rhythm and to accumulate a reasonable number of conversions. For most businesses that is a rolling ninety days rather than a calendar month, because a month is short enough that the day of the week a bank holiday falls on can move the figure visibly.

    You get: A rolling window, stated on the report

  4. Express the result as a band, not a point

    Calculate the rate for each of the last several windows and record the range it moved within while nothing important changed. That range is your baseline. A single point invites people to react to movement inside it, which is the most common way a conversion number causes harm rather than preventing it.

    You get: A band per segment, with the observation period noted

  5. Set the trigger that starts an investigation

    Decide in advance what would make you look: a movement outside the band that persists for a stated number of windows, in a segment that matters. Write it down. Without a pre-agreed trigger, every review meeting relitigates whether this month is worth worrying about, and the answer is decided by whoever speaks first.

    You get: A written trigger, per segment, with a named owner

The arithmetic

How wide the uncertainty around your own rate actually is

A conversion rate estimated from a sample carries a standard error of the square root of p times one minus p, over n. These are worked examples of that formula, not measurements of anybody.

Approximate 95% interval around a 2% rate observed over 1,000 sessions

±0.9 points

Approximate 95% interval around a 2% rate observed over 1,000 sessions

Source: NIST/SEMATECH e-Handbook of Statistical Methods, §7.2.4

The same interval once the window contains 10,000 sessions

±0.3 points

The same interval once the window contains 10,000 sessions

Source: NIST/SEMATECH e-Handbook of Statistical Methods, §7.2.4

Default inactivity timeout after which Google Analytics starts a new session

30 minutes

Default inactivity timeout after which Google Analytics starts a new session

Source: Analytics Help: about Analytics sessions

Take the first line seriously, because it changes what a dashboard means. At a two per cent rate measured over a thousand sessions, the true underlying rate could plausibly be anywhere from about 1.1 to about 2.9 per cent. A month reporting 2.0 and a month reporting 2.7 are entirely consistent with nothing at all having changed, and any meeting that treats the difference as a result is reacting to arithmetic noise.

The second line is why patience beats instrumentation. Ten times the sessions narrows the interval by roughly a factor of three, because the error falls with the square root of the sample. There is no dashboard configuration, tool or model that shortcuts this, which is also why so much conversion reporting at small volumes is discussion rather than evidence.

The third line is a reminder that the denominator is a setting. If a business changes its session timeout, or fixes cross-domain tracking, or starts excluding internal traffic, the rate moves without a single visitor behaving differently. Any baseline has to be re-established after a change like that, and the old band should be marked as belonging to a different measurement.

Comparisons

Four things you could compare your rate against

Only two of these support a decision. The other two are widely used, which is why they are listed rather than omitted.

Your own history, your own segments, a published benchmark and a self-measured competitor set compared on validity, cost and the decision each can support
DimensionIs the comparison validWhat it costs youWhat it can support
Your own past, same segmentYes, provided the definitions and the tracking did not change in between.Nothing beyond the discipline of holding the segment and window fixed.Nearly every decision worth taking: whether something changed, and when it started.
Your own segments against each otherYes for finding where the problem lives; not for judging whether a level is acceptable.A little analysis time. The data is already collected.Prioritisation. Which channel, device or template is dragging, and which is fine.
A published industry benchmarkAlmost never, because the population, the denominator and the numerator are all different or unstated.Free, and the cost arrives later when somebody turns it into a target.Nothing safely. At best a very rough sense of the order of magnitude in a category.
A competitor set you measured yourselfYes in principle, since you control the definitions. Almost nobody can obtain the data.High. It requires access to figures competitors have no reason to give you.Genuine positioning claims, in the rare cases where the data exists, such as regulated disclosures.

What goes wrong

Four ways a benchmark actively costs money

These are not hypothetical failure modes. They are the reasons this page exists in the form it does.

A number from an article became a target in a board paper.
Once a figure is written into a plan it acquires authority nobody intended it to have, and the sourcing conversation is over. Work then gets prioritised to move towards a number computed from a population the business is not in, and activity that would have produced more revenue at a lower rate gets rejected for making the metric worse.
The tracking changed and the baseline did not.
A consent banner is added, cross-domain tracking is repaired, a new event is counted, or the session timeout is adjusted. Each of these moves the rate without any change in visitor behaviour. If the historical band is left in place, the next few months look like a performance story, and somebody will be held responsible for a configuration change.
Two periods with different traffic mixes were compared as if they were alike.
A quarter with a brand campaign running produces different visitors from a quarter without one. Comparing the blended rate across the two measures the mix, not the site. The fix is to compare within a fixed segment, which is why the segments have to be chosen before anybody looks at the numbers.
A supplier was judged on a figure that was never comparable.
An agency or an internal team gets measured against a published average built on different definitions. This punishes work that improved lead quality while reducing volume, and rewards anybody willing to count something easier as a conversion. If a rate is going to be used in an evaluation, the definitions have to be fixed in writing first, by both parties.

Questions

The questions this always produces

So what is a good conversion rate?

There is no honest universal answer, and this is the most useful thing on this page rather than an evasion. A rate is a ratio between two quantities that every business defines differently, measured on traffic every business acquires differently, over periods nobody states.

The answerable version is: is this rate better or worse than the same segment of your own traffic was last quarter, and is the difference larger than the noise? That question has a real answer, and it is the one that should drive any decision.

Why can I not just use an industry benchmark?

Because you cannot see any of the three things that would make the comparison valid. The population is normally the publisher’s own customers, which is a self-selected group rather than a sample of your sector. The denominator depends on how sessions are configured, which varies between businesses. And the numerator depends on what each business decided to count as a conversion.

If a source states its population, its sampling method, its exact denominator and numerator definitions, and the period, and you can filter it to a segment matching yours, it becomes usable. Very few published tables state even two of those.

How much traffic do I need before the number means anything?

Enough that the interval around the estimate is narrower than the difference you care about. That is computable rather than a matter of opinion: the standard error of a proportion is the square root of p times one minus p, divided by the number of sessions, and the ninety-five per cent interval is about twice that either side.

At a two per cent rate over a thousand sessions, that interval is roughly nine tenths of a percentage point wide either way. If you want to detect changes smaller than that, you need either more sessions or a longer window, and there is no arrangement of the dashboard that avoids it.

Our conversion rate went up after traffic fell. Is that good?

It might be excellent or completely meaningless, and the rate alone cannot distinguish them. If the traffic you lost was low-intent, the same number of enquiries now arrives from fewer visits and nothing has been lost. If the traffic you lost was buyers, the rate rose while the business got worse.

This is why enquiries and revenue belong on the same chart as the rate. A ratio can improve because the numerator rose or because the denominator fell, and only one of those is worth celebrating.

What should we report to the board instead of a benchmark?

A band rather than a point, with the segment, the window and the definition attached, plus the count of enquiries or orders underneath it. For example: paid search, mobile, form submissions, rolling ninety days, currently within the band established over the previous year.

Then add one line stating what would trigger an investigation — a movement outside the band that persists for a stated period. That converts the number from a scoreboard into an instrument, which is the only thing it can honestly be.

Do you publish conversion rate benchmarks?

No. We have no dataset we could stand behind, and our editorial policy requires every figure to carry a real primary source with its population and period stated. Republishing somebody else’s unsourced table under our name would breach that.

If we ever run a study large enough to be worth publishing, it will arrive with its methodology, its sample and its limitations attached, and it will still not be a substitute for your own baseline.

Establish the baseline before anyone sets a target against it

Half an hour with your analytics is usually enough to settle the definitions, choose the segments and see how wide the noise band actually is. After that the number starts telling you things instead of starting arguments.

Last updated

We use analytics to understand which pages are useful. Nothing runs until you choose, and we do not sell or share what we collect. What we would set.