Skip to content
E-commerce / Home & Furniture10 months · 2026

Two years of winning tests and flat revenue

Their experimentation programme reported a 71% win rate. Revenue had not moved. Almost every winner had been called on a sample too small to mean anything.

+38%

Conversion rate

71% → 22%

Reported win rate

+44%

Contribution per session

4 → 38

Tests reaching significance

The challenge

Bramwell sells furniture at an average order value around $900 and had run an in-house testing programme for two years. The reporting was impressive — 71% of tests declared winners, each with a percentage lift attached — and site conversion rate was almost exactly where it had started. The programme had become a source of internal confidence and no revenue, which is a harder problem to raise than an obvious failure.

What we did

01

Recalculated every historical test

We took two years of test results and computed, for each, the sample size that would have been needed to detect the observed lift. The majority had been stopped well before it, most often when a variant went green in the dashboard. The 71% win rate was almost entirely a measure of how early they had looked.

02

Set a stopping rule before any new test ran

Required sample calculated in advance from baseline conversion rate and the smallest lift worth acting on, with the test not read until it arrived. This is the whole discipline, it is unglamorous, and it immediately cut the number of tests they could run per quarter — which was the difficult conversation.

03

Moved from button tests to argument tests

Most of the historical programme tested surface details, which is where small effects live and small effects need enormous samples. We redirected toward changes big enough to be detectable at their traffic: delivery expectations stated up front, room context rather than cut-out product shots, and finance terms surfaced before the basket.

04

Fixed the measurement underneath it

Their test tool and their order data disagreed by a meaningful margin, so even correctly-powered results were being read against an unreliable baseline. Reconciling the two was a prerequisite nobody had checked in two years of testing.

The result

Conversion rate rose 38% over three quarters against a two-year flat line, and contribution per session rose 44% because the winning changes affected order composition as well as conversion. The reported win rate fell from 71% to 22%, which is the number the team is proudest of — 38 tests reached significance in ten months against four in the preceding two years. Fewer declared winners, all of them real.

The hardest part was telling the board our old reporting had been optimistic, and that we would now be running fewer tests and winning less often. Watching the conversion rate move for the first time in two years made that conversation retrospectively easy.
Rosalind AchterbergHead of Ecommerce, Bramwell Home
Same playbook

Let's run this on your business.

Start with a free teardown. We'll show you where your funnel leaks and what we'd do about it — before you spend anything.