Header tag

Wednesday, 12 August 2026

Why Average Order Value Misleads in A/B Tests

 Why Average Order Value Is Lying to You in A/B Tests, And What to Do About It

 Working in ecommerce, Average Order Value is almost certainly one of the key metrics we use to evaluate A/B tests. When we’re deciding whether Recipe B outperforms Control, AOV offers a clean, definitive answer. One number. Higher is better. We sold more items per order, revenue went up, AOV went up.  After all, Average Order Value is defined as Total Revenue divided by Total Orders.  If AOV went up, then Revenue went up too.  Implement Recipe B.

However, we have a slight problem: AOV, taken in isolation, is potentially one of the most misleading KPIs you can use to call a test winner. It can ‘identify’ a clear winner when the reality is far more complicated. And it can show no meaningful difference between two recipes while something fundamentally important is happening beneath the surface. 

Why? AOV is a mean average. And mean averages, as any statistician will remind you (and as we all learned and then forgot from school), are dangerously susceptible to skew, a problem that becomes especially treacherous when you are making binary implement-or-abandon decisions based on test results.

The Problem with a Single Number

Let’s take a simple example. Imagine you are running an A/B test on your product detail page. After two weeks, you pull the results and compare order values across both recipes.

Control (Recipe A) recorded the following ten orders:

$15, $18, $22, $25, $30, $32, $35, $40, $45, $50

AOV for Recipe A: $31.20.

 

Recipe B recorded these ten orders:

$8, $10, $12, $14, $16, $80, $85, $90, $95, $100

AOV for Recipe B: $51.00.

Here's how it might look with 30 orders per recipe:

Scale up accordingly, and think about 100 orders, or 150.  The volumes become more ‘real’ but the underlying problem stays the same (yes, even larger sample sizes can suffer this kind of systemic skew).  On the surface, this looks decisive. Recipe B delivers a 63% uplift in AOV. If this is your headline metric, you are shipping Recipe B today and writing up a case study by Friday.  And we all love those case studies (although 63% uplift isn’t as good as the 197% conversion uplift that a promotional case study will tell you is possible with A/B testing!).

 But look at what  happened. Recipe B produced a polarized order distribution. Many of its customers spent considerably less than Recipe A's lowest order, while the other half spent considerably more. The mid-range products, which could represent the the core of your catalogue have seemingly disappeared from the basket entirely. The average went up, but the underlying customer behavior shifted in a way that deserves scrutiny, not celebration.

And let’s not forget to  look at what the test was trying to achieve.  Were you actively promoting your mid-range products?  Were you offering a promotion on the higher-price products?  Did your test just become a spectacular winner or a tragic loser?

Or did you make a design mistake?  Was Recipe B's design inadvertently confusing for some segments while upselling effectively to others? Did it suppress mid-tier product discovery? These are critical questions, and AOV alone will never surface them. It compresses a complex distribution into a single data point and, in doing so, strips away the nuance that should  inform your decision.

How Skew Distorts Test Results

There are several real-world scenarios where AOV can mislead your A/B test analysis.

A few high-value outliers inflate the winning recipe.

Perhaps Recipe B happened to capture three bulk-buying customers or a couple of corporate purchasers. Those handful of transactions dragged the average up, creating the appearance of a meaningful uplift. Meanwhile, the typical customer in Recipe B  spent less than their counterpart in Recipe A. You are about to ship an experience optimized for edge cases.  If you’re going to do that, then perhaps it’s time to consider targeting (or personalization – it seems to me sometimes that the correct term to use depends on who’s doing it and how targeted it is).

Both recipes perform similarly in the middle, but diverge at the extremes.

Recipe B might be driving more entry-level impulse purchases while simultaneously encouraging a small group of high-intent buyers to add more to their baskets. AOV looks higher, but the majority of your customers are spending less. The "improvement" is an illusion created by a handful of power users subsidizing the average.  Again, is it time to segment your visitors based on the products they’re viewing, and serve the winning recipe to the users who are making those entry-level purchases?  And let’s not forget the high-intent buyers.  They represent the kind of customers that AOV was created for – we assume AOV means ‘customer added more to their order’ and in this case, that’s exactly what’s happening.

In each of these cases, AOV gives you a number. It does not give you a reliable basis for an implementation decision.  We need to unlearn that assumption that AOV means ‘customer added more to their order’ and look more closely at the data.

The Case for Price Band Analysis in A/B Testing

This is where price band analysis becomes essential. Rather than calculating a single average and then comparing between Recipe A and Recipe B, price band analysis involves segmenting orders from each recipe into defined value brackets — for example, $0–$20, $20–$50, $50–$100, $100–$200, and $200-plus — and then comparing the volume and revenue contribution of each bracket across both experiences.

This approach gives you something AOV never can: distribution visibility.

When you evaluate your test through price bands, you can answer questions and decide whether a result is genuinely actionable:

- Is Recipe B's AOV uplift broad-based across all brackets, or is it being driven by a single band?

- Did Recipe B improve high-value orders at the expense of your mid-range core?  Or did you genuinely shift mid-range customers into the high-price band?

- Are you seeing a volume shift into lower bands that is being masked by a few large transactions?

- Does Recipe B's distribution align with the customer segments that your test was intended to target?

- Is the "winning" recipe creating a more fragile revenue profile by concentrating performance in fewer, larger orders?

These are the questions that separate a genuine optimization win from a statistical mirage. AOV alone cannot answer any of them.

Practical Implementation

For most experimentation teams, layering price band analysis into test evaluation is not technically difficult. It requires defining sensible brackets that reflect your catalogue's price architecture, and then building a view, which cold be in your testing platform's reporting, your BI tool or your spreadsheet, that breaks down order count, revenue, and share of total by band for each recipe.

The brackets themselves matter, because they should be meaningful to your business rather than arbitrary round numbers. Look at your product pricing structure and your historical order distribution to define bands that separate genuinely different purchase behaviours. A luxury fashion retailer and a toy retailer will need entirely different frameworks.

 Once established, review the bands alongside AOV. Don’t abandon AOV, because it still has a role as a quick directional signal during a test. But it should be the starting point for investigation, not the final number. When AOV is different between recipes, your first response should be to open the price band comparison and understand why it differs. Is Recipe B lifting performance across the board, or is the uplift concentrated in a single bracket? Is the change driven by many incremental improvements or by a handful of outlier transactions that may not replicate at scale?

This is especially important when you consider that we run our A/B tests for a limited period of time. The smaller your sample, the more susceptible AOV is to skew from a few extreme orders, so again, consider this when you ask the question ‘how long should I run my test for?’. Price band analysis acts as a critical sanity check, helping you distinguish between a robust, replicable improvement and a result that is being propped up by noise.

Making Better Test Implementation Decisions

There is a broader principle at play here, and it is one that experienced experimenters already know instinctively: any metric that compresses complexity into a single number should be treated with healthy skepticism when making shipping decisions.  AOV is not unique in this regard; conversion rate and revenue per visitor (or per visit) all suffer from similar issues. They are useful as signals, but dangerous as conclusions.

The best experimentation teams treat AOV as a tier-one indicator, which is worth monitoring for directional movement, but you need to build out your recommendation and rationale on the richer, more detailed data sitting underneath it. When you present test results to stakeholders, showing a price band comparison between recipes tells a far more compelling and trustworthy story than a single percentage uplift ever could.

Price band analysis is not a radical new technique. It is straightforward, practical, and surprisingly underused in A/B test evaluation, given how much it reveals. If AOV is currently your primary metric for judging test winners, consider what would happen if you supplemented it with a price band distribution comparison. The decisions it informs will be more confident, more nuanced, and ultimately more valuable to your business, because in experimentation the most dangerous result is not the one that is wrong. It's the one that looks right but leads you to implement an experience you don’t fully understand.

Other posts I've written about online metrics and testing:

Web Analytics - Requirements Gathering
One KPI Too Many - why you need a critical few
Why too many cooks spoil the A/B testing roadmap





Thursday, 23 July 2026

How Valuable Are Your Meetings?

Work means meetings

Work means meetings.  It means meeting with any combination of your team-mates; your manager and colleagues from other departments (developers, designers, project managers, strategists, architects, builders... you name them, they want to talk to you). 

If I had a dollar (or a pound, or a Euro) for each time I've heard "Can I get 30 minutes of your time?" then I'd be a very wealthy man.  If I included times when I've been given a cold call or an unsolicited email from a supplier asking to get on my calendar, I'd be extremely wealthy!

Which brings me to the point:  how much do you and your leadership value your working time?

How good are your meetings?

Let's suppose I'm paid $60 per hour.  Nice, right?  But it will work for the numbers - I get paid $1 per minute to do my job.  To sell more stuff.  To design more stuff, or to develop more stuff.  My job description says sell/design/develop/build stuff, and I get paid $1 per minute to do that.

But if I get dragged away from my role to attend a meeting, then I have to hope that the meeting I'm attending will be worth the time I'm spending not doing my normal role (unless that meeting is actually part of the design/build/sell activity I normally do).

So I get pulled onto a call for all the managers in my department.  It's a 30-minute call.  Sadly, it's not actually relevant to my role, so the business is paying me to do something that's not part of my actual role.  In fact, it doesn't apply to about 20 of the 40 managers in the meeting, and they're all getting paid the same as me.

20 managers.  $1 per minute, per manager.  30 minutes?  Total cost: $600.

For a meeting that lasts 30 minutes.  Now imagine that it lasts an hour; now it's costing the business $1200.  It's an hour that could have been spent building/selling or whatever.  But why did it happen?  Because leadership determined that it was more important for you to hear what they said instead of doing what you normally do.


Now imagine it's an 'all hands on deck' meeting.  Or a 'strategic communications' presentation, and you invite 100 managers, and it lasts 30 minutes.  It needs to span the entire business strategy to be relevant to everybody; for each manager that it doesn't apply to, that's $30 at risk, and if it isn't relevant to any of them (or if it's a repeat of last month's meeting) then that's $3000 at risk.

So what do you do?

When you call a meeting with an individual, you're pulling them away from their normal role to meet with you.  You know the job you hired them to do?  You're now preventing them from doing it. 

Or perhaps you're arranging a meeting with your full team - of managers, who get paid $60 per hour - you need to make absolutely certain that the meeting you're hosting is going to be at least as valuable to them (and to the business) as their normal work.

And if it isn't?  In that case, there's definitely a strong argument for an asynchronous communication method.  Email... Teams chat... Video plus transcript... it's your choice, but it gives your team the choice of when to consume the content.

"But," you say, "The content is so important I need to be sure they've consumed it!"  

If you can say that for sure, then it's definitely time to hold a meeting.  However, as a gentle suggestion, please make sure that you're not confusing 'the content' with 'my opinion'.


But why aren't your attendees participating?  Are you just talking for 30 (or 60) minutes?  You've been hired to lead, support and manage - not to perform a 30-minute monologue to a tough audience of colleagues who are finding your material irrelevant and uninteresting.  If you have eight people on your team, and you pull them out of their work for an hour to attend your meeting, you have cost the business the equivalent of one full day of work.

Or imagine you call a meeting for 40 people, for an hour.  You've just taken the equivalent of one week's work out of the business, no matter how much that costs.

Consider this:  if your organization is so large that you invite 400 people, and they all turn up at 8:00 am sharp for your meeting, and you don't start until 8:05 am, then that's the equivalent of over four days of work that you've wasted.  

And if your content doesn't apply to a quarter of that audience, and your meeting lasts for an hour, then 100 hours of work got stuck in a meeting it didn't want to be in, when it could have been doing its job.

So, what should you do?

Make your agenda more inclusive

If you're inviting everybody, then make sure your message applies to all of them.  Talk to the builders, the designers, the developers, the sellers.  Talk to the team that sells over the phone.  Include the team that manages the stores.  

If you can't be more inclusive with your agenda, then be more selective with your audience

Do you typically invite all the managers in your department?  Look at the agenda, look at who you're inviting:  do they match?

If you have a guest speaker, ask them to fully introduce themselves; to explain what they do, and to explain why it's relevant to the audience.  Context over content, this time.

Think about it: the more people you invite to a meeting, the more you're costing the business and the harder you're going to have to work to be relevant to everybody.  It's worth the effort, but make sure you're actually putting the effort in.

Other topics I've posted on that may be of interest to you:

Project Management: A Trip To the Moon (when meetings go wrong)
Team-building exercises with Remote Workers

Monday, 6 July 2026

Arsenal's Disciplinary Record

Here's an interesting fact, brought to you by the wonders of social media:

Arsenal, Premier League Champions for 2025-26, had no red cards and no penalties awarded against them throughout the season.


And immediately, social media jumped on the bandwagon with the long-standing VARsenal - the claim that the Video Assistant Referees for the Premier League were in some way biased in favour of Arsenal, and hence any debatable or controversial decision would go in their favour.

So:  how unusual is it for a team to have no red cards, and to have no penalties against them in a season?

To answer this one, we're going to need some data, and fortunately the internet is full of it.  I've looked at football data before, and for some reason, gambling sites tend to keep the best records.  Here's the data on red cards first, taken from MyFootballFacts.com.


2025-26 Red Cards by Clubs2025-26 Yellow Cards by Clubs
ClubNo.ClubNo.
Chelsea8 Tottenham Hotspur98
Everton 4 Chelsea 90
Tottenham Hotspur 4 AFC  Bournemouth88
Burnley3Brighton & Hove Albion86
Manchester United 3 Sunderland82
Newcastle United 3 Wolverhampton Wanderers 77
Sunderland3 Fulham75
West Ham United 3 Crystal Palace74
Wolverhampton Wanderers 3 Everton 72
AFC Bournemouth2 Brentford69
Crystal Palace2 West Ham United 68
Aston  Villa1 Manchester City 67
Brentford1 Burnley64
Fulham1 Newcastle United 63
Leeds United1 Leeds United62
Liverpool 1 Manchester United 62
Nottingham Forest1 Nottingham Forest60
Arsenal 0 Aston Villa58
Brighton & Hove   Albion0 Liverpool 57
Manchester City 0 Arsenal 51
Total  Red Cards44 Total  Yellow Cards1423


Summary of initial findings:

Red cards

It's not impossible for a team to complete the season with zero red cards:  Manchester City and Brighton also achieved this result.  It's rare, with only three out of the 20 teams having a clean sweep, but it's not unheard of.

Average per team is 2.2 red cards per season.  Standard deviation = 1.88
Arsenal are only just outside one standard deviation of the mean, so it's certainly not a statistical outlier - it's still possible.

Chelsea fans might have more of a case of bias against them, with 8 red cards against the average of 2.2, as they are just outside three standard deviations from the mean, making them a definite outlier.

Yellow Cards

Arsenal had the fewest yellow cards (or "bookings"), with 51 against a Premier League average of  71.15.  Spurs had 98, and Chelsea accumulated 90.  It seems that their performance in the Red Card Table is consistent with their overall behaviour.

Penalties Conceded

CLUB PKC
Brentford 8
Burnley 8
Crystal Palace 7
West Ham 7
Bournemouth 6
Brighton 6
Leeds 6
Nottm Forest 6
Newcastle 5
Wolves 5
Everton 4
Fulham 4
Liverpool 4
Man United 4
Sunderland 3
Tottenham 3
Aston Villa 2
Chelsea 2
Man City 2
Arsenal  0

Now this is a more interesting data point.  Arsenal were the only team in the Premier League not to concede a penalty kick in the season.  Whatever Chelsea were doing to get their cards, they were doing it outside the penalty area and hence only conceded two penalties throughout the season. 

The average was 4.6 penalties conceded per team for the season, so 0 is a noticeable outlier.  A follow-up question is - how many times did Arsenal's opponents have possession in their penalty area?  The metric which is measured here is how many times did an opposition player touch the ball in Arsenal's penalty area?  
Arsenal had 1228 touches conceded in the season  (32 per game, and third lowest in the League), so it makes sense that they conceded fewer penalties.

Conclusion

So, Arsenal won the league by keeping their opponents out of their penalty area - leading to no penalties conceded - and by playing a careful game with few reckless challenges and no sendings off.  There's no statistical indication of bias or unfair treatment towards them by the referees; the average for the Premier League was 2.2 red cards in the season, and zero is well within an expected distribution (unlike Chelsea, who have a stronger case for harsh treatment, based only the numbers).  Protecting their penalty area enabled them to concede fewer penalties in the season, and they had no sendings off.

The follow-up question (for another time) is how often do teams win when they've had a man sent off, compared to keeping all 11 men on the pitch?  I'll look into it!

More posts on football data:

Port Vale 0 Arsenal 2 match report