Conversion rate optimisation checklist
Twenty-three conversion checks in four groups, tickable on this page, plus an Excel test log that ICE-scores your ideas and flags any result with too few conversions to trust. The obvious fixes come before any split test, because most sites do not have the traffic to test their way to them — and that is where most of the available gain is anyway.
- CHECKLIST
- XLSX
- FREE
- Best for
- Anyone whose traffic is fine and whose enquiries are not
- Includes
- 23 tickable checks + Excel test log with ICE + printable PDF
- Time to use
- 90 minutes, then ongoing
Free to download and use in your own client work. No email address required.
Problems this solves
Recurring complaints about this task, and what this resource does about each one. If it does not solve your version of the problem, that is worth knowing before you download anything.
- A calculator says 95% significant on a few hundred visits, so you ship it
- The test log returns "Not enough data" below 100 conversions per arm regardless of what a significance calculator says, and "Inconclusive" for anything within 5%. Reaching 95% on a tiny sample is a false positive, not a result.
- You do not have the traffic to test anything
- Group two is the whole programme for you: the fixes that need no sample size at all. On most sites that is where nearly all the available gain sits, and it is cheaper than buying more visits.
- Ideas get argued about by opinion and the loudest voice wins
- ICE scoring, with confidence defined as evidence rather than enthusiasm — a 5 requires having seen the problem in recordings, in the data and in what people say.
Most conversion advice assumes you can split test. Below roughly a hundred conversions a month you cannot — a test will not reach significance in a useful timeframe, so you end up acting on noise.
So the obvious fixes come first here. For most sites they are the whole programme, and they are where almost all the gain is.
The checklist
Before testing anything 5
The obvious things, before any test 8
Running a test properly 6
After the test 4
Do these before you test anything
Group two of the checklist is the part worth reading twice, because none of it needs a tool, a developer or a sample size.
The two that recover the most: the form asks only for what you need to respond to someone — every extra field costs completions, and “how did you hear about us?” costs more than it teaches — and errors are shown inline, in words, next to the field that caused them. A form that rejects a submission without saying why is indistinguishable from a broken one.
Submit your own form, on a phone, this month
I have found more silently broken forms than I would like to admit. A form that stopped delivering three weeks ago looks exactly like a quiet month. It costs two minutes to check and it is the only item on this page that can be catastrophic on its own.
Prioritising with ICE
When you do have a list of ideas, score each one and work down. The workbook does the arithmetic and ranks them.
| Axis | The question it asks | The scale |
|---|---|---|
| Impact | If this works, how much would it move the primary metric? | 1 = barely measurable · 5 = would change the quarter |
| Confidence | What evidence do you have that this is a real problem? | 1 = someone’s opinion · 5 = seen in recordings, data and enquiries |
| Ease | How much work is it to build and ship? | 1 = needs a rebuild · 5 = a copy change this afternoon |
Confidence means evidence, not enthusiasm. A 5 requires having seen the problem in session recordings, in the data, and in what people actually say. An idea scoring high on impact and low on confidence is not a test — it is a research task, and treating the two as the same is how testing programmes fill up with inconclusive results.
The test log, and the number that stops you fooling yourself
The log calculates each variant’s rate and the uplift, then returns a verdict. It is deliberately conservative:
- Fewer than 100 conversions per arm returns “Not enough data” rather than a result.
- A difference within 5% returns “Inconclusive” rather than a small win.
That is a sanity check, not a significance calculator — use a proper significance test before acting on a close result. But it exists because the most common CRO failure is stopping a test early, reporting noise as a win, and shipping a change that does nothing.
Record the losses and the inconclusive results
A log containing only wins is a marketing document. It also loses you the most valuable thing a testing programme produces: knowing what does not work on your audience, so you stop proposing it every quarter.
“The calculator says 95% significant” — the trap
This is the most expensive misunderstanding in conversion testing, and it catches careful people.
A significance calculator will happily return 95% on a few hundred visits and a handful of conversions. The number is arithmetically correct and practically meaningless, for two reasons:
Small samples produce large apparent differences by chance. With 12 conversions against 18, the observed uplift is 50% and the true difference may be nothing at all. The calculator is answering “how surprising is this gap, given these totals?” — it is not answering “will this hold”.
Stopping when it turns green guarantees a false positive. Significance fluctuates day to day as data arrives. If you check daily and stop the first time it crosses 95%, you have selected for a lucky moment rather than measured an effect. This is why the stopping rule has to be agreed in advance, and why the checklist asks for it before the test starts.
That is why the test log ignores the calculator and applies a blunter rule: under 100 conversions per arm it says “Not enough data”, and within 5% it says “Inconclusive”. Blunt beats precise-and-wrong here.
What to do when you genuinely cannot get the traffic
Three legitimate options, in order. Test bigger changes — a whole page rather than a headline, because a large effect needs less data to detect. Use fewer variants, since each one splits your traffic again. And accept a judgement call: shipping an obvious improvement without a test is more honest than dressing a coin flip up as evidence.
Common mistakes
- Split testing on traffic too low to reach significance, then acting on the result.
- Stopping a test the moment it looks positive.
- Changing two things in one variant, so a win teaches you nothing.
- Testing a button colour while the page does not say what you sell.
- Recording only the wins, so the same losing idea returns every quarter.
- Optimising a page whose conversion is not tracked correctly in the first place.