Header tag

Wednesday, 12 August 2026

Why Average Order Value Misleads in A/B Tests

 Why Average Order Value Is Lying to You in A/B Tests, And What to Do About It

 Working in ecommerce, Average Order Value is almost certainly one of the key metrics we use to evaluate A/B tests. When we’re deciding whether Recipe B outperforms Control, AOV offers a clean, definitive answer. One number. Higher is better. We sold more items per order, revenue went up, AOV went up.  After all, Average Order Value is defined as Total Revenue divided by Total Orders.  If AOV went up, then Revenue went up too.  Implement Recipe B.

However, we have a slight problem: AOV, taken in isolation, is potentially one of the most misleading KPIs you can use to call a test winner. It can ‘identify’ a clear winner when the reality is far more complicated. And it can show no meaningful difference between two recipes while something fundamentally important is happening beneath the surface. 

Why? AOV is a mean average. And mean averages, as any statistician will remind you (and as we all learned and then forgot from school), are dangerously susceptible to skew, a problem that becomes especially treacherous when you are making binary implement-or-abandon decisions based on test results.

The Problem with a Single Number

Let’s take a simple example. Imagine you are running an A/B test on your product detail page. After two weeks, you pull the results and compare order values across both recipes.

Control (Recipe A) recorded the following ten orders:

$15, $18, $22, $25, $30, $32, $35, $40, $45, $50

AOV for Recipe A: $31.20.

 

Recipe B recorded these ten orders:

$8, $10, $12, $14, $16, $80, $85, $90, $95, $100

AOV for Recipe B: $51.00.

Here's how it might look with 30 orders per recipe:

Scale up accordingly, and think about 100 orders, or 150.  The volumes become more ‘real’ but the underlying problem stays the same (yes, even larger sample sizes can suffer this kind of systemic skew).  On the surface, this looks decisive. Recipe B delivers a 63% uplift in AOV. If this is your headline metric, you are shipping Recipe B today and writing up a case study by Friday.  And we all love those case studies (although 63% uplift isn’t as good as the 197% conversion uplift that a promotional case study will tell you is possible with A/B testing!).

 But look at what  happened. Recipe B produced a polarized order distribution. Many of its customers spent considerably less than Recipe A's lowest order, while the other half spent considerably more. The mid-range products, which could represent the the core of your catalogue have seemingly disappeared from the basket entirely. The average went up, but the underlying customer behavior shifted in a way that deserves scrutiny, not celebration.

And let’s not forget to  look at what the test was trying to achieve.  Were you actively promoting your mid-range products?  Were you offering a promotion on the higher-price products?  Did your test just become a spectacular winner or a tragic loser?

Or did you make a design mistake?  Was Recipe B's design inadvertently confusing for some segments while upselling effectively to others? Did it suppress mid-tier product discovery? These are critical questions, and AOV alone will never surface them. It compresses a complex distribution into a single data point and, in doing so, strips away the nuance that should  inform your decision.

How Skew Distorts Test Results

There are several real-world scenarios where AOV can mislead your A/B test analysis.

A few high-value outliers inflate the winning recipe.

Perhaps Recipe B happened to capture three bulk-buying customers or a couple of corporate purchasers. Those handful of transactions dragged the average up, creating the appearance of a meaningful uplift. Meanwhile, the typical customer in Recipe B  spent less than their counterpart in Recipe A. You are about to ship an experience optimized for edge cases.  If you’re going to do that, then perhaps it’s time to consider targeting (or personalization – it seems to me sometimes that the correct term to use depends on who’s doing it and how targeted it is).

Both recipes perform similarly in the middle, but diverge at the extremes.

Recipe B might be driving more entry-level impulse purchases while simultaneously encouraging a small group of high-intent buyers to add more to their baskets. AOV looks higher, but the majority of your customers are spending less. The "improvement" is an illusion created by a handful of power users subsidizing the average.  Again, is it time to segment your visitors based on the products they’re viewing, and serve the winning recipe to the users who are making those entry-level purchases?  And let’s not forget the high-intent buyers.  They represent the kind of customers that AOV was created for – we assume AOV means ‘customer added more to their order’ and in this case, that’s exactly what’s happening.

In each of these cases, AOV gives you a number. It does not give you a reliable basis for an implementation decision.  We need to unlearn that assumption that AOV means ‘customer added more to their order’ and look more closely at the data.

The Case for Price Band Analysis in A/B Testing

This is where price band analysis becomes essential. Rather than calculating a single average and then comparing between Recipe A and Recipe B, price band analysis involves segmenting orders from each recipe into defined value brackets — for example, $0–$20, $20–$50, $50–$100, $100–$200, and $200-plus — and then comparing the volume and revenue contribution of each bracket across both experiences.

This approach gives you something AOV never can: distribution visibility.

When you evaluate your test through price bands, you can answer questions and decide whether a result is genuinely actionable:

- Is Recipe B's AOV uplift broad-based across all brackets, or is it being driven by a single band?

- Did Recipe B improve high-value orders at the expense of your mid-range core?  Or did you genuinely shift mid-range customers into the high-price band?

- Are you seeing a volume shift into lower bands that is being masked by a few large transactions?

- Does Recipe B's distribution align with the customer segments that your test was intended to target?

- Is the "winning" recipe creating a more fragile revenue profile by concentrating performance in fewer, larger orders?

These are the questions that separate a genuine optimization win from a statistical mirage. AOV alone cannot answer any of them.

Practical Implementation

For most experimentation teams, layering price band analysis into test evaluation is not technically difficult. It requires defining sensible brackets that reflect your catalogue's price architecture, and then building a view, which cold be in your testing platform's reporting, your BI tool or your spreadsheet, that breaks down order count, revenue, and share of total by band for each recipe.

The brackets themselves matter, because they should be meaningful to your business rather than arbitrary round numbers. Look at your product pricing structure and your historical order distribution to define bands that separate genuinely different purchase behaviours. A luxury fashion retailer and a toy retailer will need entirely different frameworks.

 Once established, review the bands alongside AOV. Don’t abandon AOV, because it still has a role as a quick directional signal during a test. But it should be the starting point for investigation, not the final number. When AOV is different between recipes, your first response should be to open the price band comparison and understand why it differs. Is Recipe B lifting performance across the board, or is the uplift concentrated in a single bracket? Is the change driven by many incremental improvements or by a handful of outlier transactions that may not replicate at scale?

This is especially important when you consider that we run our A/B tests for a limited period of time. The smaller your sample, the more susceptible AOV is to skew from a few extreme orders, so again, consider this when you ask the question ‘how long should I run my test for?’. Price band analysis acts as a critical sanity check, helping you distinguish between a robust, replicable improvement and a result that is being propped up by noise.

Making Better Test Implementation Decisions

There is a broader principle at play here, and it is one that experienced experimenters already know instinctively: any metric that compresses complexity into a single number should be treated with healthy skepticism when making shipping decisions.  AOV is not unique in this regard; conversion rate and revenue per visitor (or per visit) all suffer from similar issues. They are useful as signals, but dangerous as conclusions.

The best experimentation teams treat AOV as a tier-one indicator, which is worth monitoring for directional movement, but you need to build out your recommendation and rationale on the richer, more detailed data sitting underneath it. When you present test results to stakeholders, showing a price band comparison between recipes tells a far more compelling and trustworthy story than a single percentage uplift ever could.

Price band analysis is not a radical new technique. It is straightforward, practical, and surprisingly underused in A/B test evaluation, given how much it reveals. If AOV is currently your primary metric for judging test winners, consider what would happen if you supplemented it with a price band distribution comparison. The decisions it informs will be more confident, more nuanced, and ultimately more valuable to your business, because in experimentation the most dangerous result is not the one that is wrong. It's the one that looks right but leads you to implement an experience you don’t fully understand.

Other posts I've written about online metrics and testing:

Web Analytics - Requirements Gathering
One KPI Too Many - why you need a critical few
Why too many cooks spoil the A/B testing roadmap





No comments:

Post a Comment