Why Average Order Value Is Lying to You in A/B Tests, And What to Do About It
However, we have a slight problem: AOV, taken in isolation,
is potentially one of the most misleading KPIs you can use to call a test
winner. It can ‘identify’ a clear winner when the reality is far more
complicated. And it can show no meaningful difference between two recipes while
something fundamentally important is happening beneath the surface.
Why? AOV is a mean average. And mean averages, as any
statistician will remind you (and as we all learned and then forgot from
school), are dangerously susceptible to skew, a problem that becomes especially
treacherous when you are making binary implement-or-abandon decisions based on
test results.
The Problem with a Single Number
Let’s take a simple example. Imagine you are running an A/B
test on your product detail page. After two weeks, you pull the results and
compare order values across both recipes.
Control (Recipe A) recorded the following ten orders:
$15, $18, $22, $25, $30, $32, $35, $40, $45, $50
AOV for Recipe A: $31.20.
Recipe B recorded these ten orders:
$8, $10, $12, $14, $16, $80, $85, $90, $95, $100
AOV for Recipe B: $51.00.
Here's how it might look with 30 orders per recipe:
Scale up accordingly, and think about 100 orders, or
150. The volumes become more ‘real’ but
the underlying problem stays the same (yes, even larger sample sizes can suffer
this kind of systemic skew). On the
surface, this looks decisive. Recipe B delivers a 63% uplift in AOV. If this is
your headline metric, you are shipping Recipe B today and writing up a case
study by Friday. And we all love those
case studies (although 63% uplift isn’t as good as the 197% conversion uplift that a promotional case study will tell you is possible with A/B testing!).
And let’s not forget to look at what the
test was trying to achieve. Were you
actively promoting your mid-range products?
Were you offering a promotion on the higher-price products? Did your test just become a spectacular
winner or a tragic loser?
Or did you make a design mistake? Was Recipe B's design inadvertently confusing
for some segments while upselling effectively to others? Did it suppress
mid-tier product discovery? These are critical questions, and AOV alone will
never surface them. It compresses a complex distribution into a single data
point and, in doing so, strips away the nuance that should inform your decision.
How Skew Distorts Test Results
There are several real-world scenarios where AOV can mislead
your A/B test analysis.
A few high-value outliers inflate the winning recipe.
Perhaps Recipe B happened to capture three bulk-buying customers or a couple of corporate purchasers. Those handful of transactions dragged the average up, creating the appearance of a meaningful uplift. Meanwhile, the typical customer in Recipe B spent less than their counterpart in Recipe A. You are about to ship an experience optimized for edge cases. If you’re going to do that, then perhaps it’s time to consider targeting (or personalization – it seems to me sometimes that the correct term to use depends on who’s doing it and how targeted it is).Both recipes perform similarly in the middle, but diverge at the extremes.
Recipe B might be driving more entry-level impulse purchases while simultaneously encouraging a small group of high-intent buyers to add more to their baskets. AOV looks higher, but the majority of your customers are spending less. The "improvement" is an illusion created by a handful of power users subsidizing the average. Again, is it time to segment your visitors based on the products they’re viewing, and serve the winning recipe to the users who are making those entry-level purchases? And let’s not forget the high-intent buyers. They represent the kind of customers that AOV was created for – we assume AOV means ‘customer added more to their order’ and in this case, that’s exactly what’s happening.In each of these cases, AOV gives you a number. It does not
give you a reliable basis for an implementation decision. We need to unlearn that assumption that AOV
means ‘customer added more to their order’ and look more closely at the data.
The Case for Price Band Analysis in A/B Testing
This is where price band analysis becomes essential. Rather
than calculating a single average and then comparing between Recipe A and
Recipe B, price band analysis involves segmenting orders from each recipe into
defined value brackets — for example, $0–$20, $20–$50, $50–$100, $100–$200, and
$200-plus — and then comparing the volume and revenue contribution of each
bracket across both experiences.
This approach gives you something AOV never can: distribution
visibility.
When you evaluate your test through price bands, you can
answer questions and decide whether a result is genuinely actionable:
- Is Recipe B's AOV uplift broad-based across all brackets,
or is it being driven by a single band?
- Did Recipe B improve high-value orders at the expense of
your mid-range core? Or did you
genuinely shift mid-range customers into the high-price band?
- Are you seeing a volume shift into lower bands that is
being masked by a few large transactions?
- Does Recipe B's distribution align with the customer
segments that your test was intended to target?
- Is the "winning" recipe creating a more fragile
revenue profile by concentrating performance in fewer, larger orders?
These are the questions that separate a genuine optimization
win from a statistical mirage. AOV alone cannot answer any of them.
Practical Implementation
For most experimentation teams, layering price band analysis
into test evaluation is not technically difficult. It requires defining
sensible brackets that reflect your catalogue's price architecture, and then
building a view, which cold be in your testing platform's reporting, your BI
tool or your spreadsheet, that breaks down order count, revenue, and share of
total by band for each recipe.
The brackets themselves matter, because they should be
meaningful to your business rather than arbitrary round numbers. Look at your
product pricing structure and your historical order distribution to define
bands that separate genuinely different purchase behaviours. A luxury fashion
retailer and a toy retailer will need entirely different frameworks.
This is especially important when you consider that we run
our A/B tests for a limited period of time. The smaller your sample, the more susceptible
AOV is to skew from a few extreme orders, so again, consider this when you ask
the question ‘how long should I run my test for?’. Price band analysis acts as
a critical sanity check, helping you distinguish between a robust, replicable
improvement and a result that is being propped up by noise.
Making Better Test Implementation Decisions
There is a broader principle at play here, and it is one
that experienced experimenters already know instinctively: any metric that
compresses complexity into a single number should be treated with healthy
skepticism when making shipping decisions. AOV is not unique in this regard; conversion
rate and revenue per visitor (or per visit) all suffer from similar issues.
They are useful as signals, but dangerous as conclusions.
The best experimentation teams treat AOV as a tier-one
indicator, which is worth monitoring for directional movement, but you need to
build out your recommendation and rationale on the richer, more detailed data
sitting underneath it. When you present test results to stakeholders, showing a
price band comparison between recipes tells a far more compelling and
trustworthy story than a single percentage uplift ever could.
Price band analysis is not a radical new technique. It is
straightforward, practical, and surprisingly underused in A/B test evaluation,
given how much it reveals. If AOV is currently your primary metric for judging
test winners, consider what would happen if you supplemented it with a price
band distribution comparison. The decisions it informs will be more confident,
more nuanced, and ultimately more valuable to your business, because in experimentation the most dangerous result is not
the one that is wrong. It's the one that looks right but leads you to implement
an experience you don’t fully understand.
Other posts I've written about online metrics and testing:
Web Analytics - Requirements Gathering
One KPI Too Many - why you need a critical few
Why too many cooks spoil the A/B testing roadmap

No comments:
Post a Comment