Tests read correctly:
not called the moment one bar looks taller
Most A/B tests get called the moment one variant looks ahead, often before there is enough data to say anything. We run a disciplined testing program: a ranked backlog of hypotheses, a proper significance check before any test is called, and a log so a team running many tests never loses track of what was actually learned.
Where the money leaks today
A test dashboard that shows one variant ahead by a visible margin gets called a winner within a day or two in most teams, long before the sample size or test duration actually supports that conclusion. A result that looks decisive on day two can flip by day seven as more data arrives, but by then the “winning” variant is often already live everywhere.
The opposite problem is just as common and less visible: a test that genuinely shows no real difference keeps running for months because nobody formally calls it, quietly consuming traffic that could go toward a hypothesis with a real effect to find.
Running tests without a disciplined read is close to running no tests at all, except it feels like progress, which is worse, because it crowds out the fixes that would actually move the number.
What we do
We build a ranked backlog of hypotheses from the actual funnel data, not from a generic list of “proven” tactics, prioritized by expected impact and how much traffic each step gets, so the first test is the one most likely to matter rather than the easiest to set up. One or two tests run at a time, on your current landing page builder or ad platform’s native testing tool where possible, so nothing new gets bolted onto your stack unnecessarily.
Every test gets a proper significance check: sample size, duration and variance factored in, with an early-call warning shown clearly rather than a confident-looking but premature winner. Where sample size allows, we break results down by device, channel or new versus returning visitor, since an overall winner can lose for a specific segment, a detail easy to miss in an aggregate read.
Every result, win, loss or genuine tie, gets logged with the reasoning behind the call, so a program running tests for a year does not quietly repeat the same question it already answered.
What we need from you
Access to the page or tool where tests will run, and whoever owns the final call on priority within the ranked backlog, since we surface the data but business context on what matters most this quarter is yours. If you already have a testing tool installed and unused, tell us; we will build on what exists rather than adding a new one.
How we measure
Each test’s result against a proper significance check, with the sample size and duration it was based on logged for audit. Over a quarter, the number of real, confirmed wins against the number of tests run, and honesty about how many came back inconclusive.
Price and timeline
| Option | Price | What it covers | Timeline |
|---|---|---|---|
| Launch or audit | from $500 | Single test setup on your current tool, one hypothesis | 1 week |
| Monthly management | from $1,200 / month | Ranked backlog, one to two tests running, monthly log | monthly, no lock-in |
| Full control, handover to your team | from $3,000 | Full testing framework and backlog handed to your team, training | 3 to 4 weeks |
Related
This program pairs directly with conversion rate optimisation and landing page design and optimisation. For the statistics engine behind the significance checks, see automated A/B test analysis. The full build is on the analytics service page. Real examples: the trading app funnel audit and the Bali lead-routing project, where a quiz funnel was measured step by step from entry to buyer.
Calling tests by eyeballing a dashboard? Get in touch and we will look at your current testing setup in the first call.
FAQ
How much does an A/B testing program cost?
From $1,200 a month for a ranked backlog and one to two tests running at a time, read correctly for significance, no lock-in. A single test setup without the ongoing program is available from $500.
How long does a single test need to run?
It depends on your traffic and current conversion rate, typically 2 to 4 weeks for a meaningful sample on a mid-traffic page. We tell you the expected duration before the test starts, not after it has already run too short.
What traffic do we need for testing to make sense?
A few hundred conversions a month on the page or step being tested is a reasonable minimum; below that, tests take too long to read reliably and we recommend bigger, obvious fixes instead of formal testing.
What do you need from us?
Access to the page or platform where the test runs, and sign-off on which hypothesis to test first from the ranked backlog, since we build the list but the business context on priority is yours.
How do you report results?
Every test gets a written result: winner, confidence level, and what it means for the next step, logged centrally so the program compounds instead of repeating the same question six months later.