Header tag

Wednesday, 7 October 2026

Battleships and A/B Testing

When you picture a strategic board game, the first thing that often comes to mind is the timeless pen‑and‑paper showdown known as Battleships.  Two players plot secret fleets on a hidden grid, then take turns calling out coordinates in hopes of sinking their opponent’s fleet. Oddly enough, the same logic that underpins this nostalgic pastime also powers the modern marketer’s most valuable tool: A/B testing.  But you knew I was going to say that in this blog, right? ;-)


There are some obvious and some obscure similarities between a round of Battleships and the iterative process of optimizing complex elements of a website, such as a navigation menu or a landing page, or a series of products on a category page. I've observed that it can take a blend of hypothesis, random probing (or sometimes just gut feel) and data‑driven refinement can turn a handful of menu items into a high‑performing traffic‑masterpiece.  You aren't going to hit the perfect combination first time, but you aren't going to have to work out all the combinations (some of them just don't make sense) to get to the optimum answer fairly quickly.

 

The Playing Field: The Grid and The Page Layout 

 

Both scenarios start with an unknown distribution. In Battleships you lack any clue where the opponent's fleet sits; on a fresh website (or on a site component that you've not tested before) you have limited insight into which menu label will attract the most clicks.  Or, you may have extensive data on engagement with your menu options, but you don't know what the most effective sequence might be.  

The first step in both games is to gather information, and the most straightforward way to do that is to make a series of probes. 

  


Random or “Best‑Guess” First Moves 

 

A seasoned Battleship player knows that a completely random first shot is often a good baseline. It gives an unbiased sample of the board and avoids early bias toward a particular quadrant. Likewise, when devising an A/B test on a navigation menu, we have to start with a baseline version (Recipe A, the current state).  This is where we start, based on the original agreed sequence for the menu (or the one that was agreed or imposed most recently).   Your first test recipe - your first guess, or challenger - will be the best guess from existing UX guidelines, past experience, engagement data, clickstream analysis or competitor analysis. Please, please, let it have some form of data backing: it's not essential, but it will certainly help.  At this stage, it's perfectly fine (in fact, it's better) if you have multiple hypotheses on the go... everybody devises a recipe, and you select the best for testing.  The number of recipes you can test will depend on your traffic, but at this stage you want to maximize your options.

I mean, wouldn't Battleships be quicker (and more exciting) if you got to pick three or four co-ordinates per turn, instead of just one?

 

Does the the randomness matter? 

 

At this point, you're focusing more on exploration, and less on exploitation.  Random probes let you explore the unknown space before focusing on promising areas.  More importantly, you'll avoid premature convergence: if you jump straight to an “obvious” menu order, you may lock yourself into a sub‑optimal layout before you ever see the data.  One things that's very important at this stage is to test more than one new idea - in Battleships, this would be like shouting out four or five guesses each turn.  A scattershot approach in your testing will yield the fastest results - you're not necessarily aiming to find a winner immediately, this is about testing a range of different hypotheses.  Will colour be helpful? Which item should go first in the list?  Which should go in the middle?  


The First Hit: Discovering a “Ship” 

 

When a player announces “B5: hit!” they’ve uncovered part of a ship. This moment gives them very specific information: a ship occupies that cell and, by the rules, must extend either horizontally or vertically. 

 

In A/B testing, the equivalent is a statistically significant lift in a key metric, for example, a 6 % increase in overall click‑through rate (CTR) for the new menu order. That lift tells you three things: 

 

1. Presence – The altered ordering resonates with users. 

2. Direction – The change (e.g., moving “Pricing” closer to the left) is beneficial. 

3. Potential shape – The magnitude of uplift hints at how many “cells” (menu items) may still be optimized. 

 

Just as a Battleship player deduces the ship’s orientation after the first round of hits (if they're lucky, or clever), the marketer can infer which type of change matters most: position, wording, visual emphasis, or grouping. 

 

Targeted Follow‑ups: “Scanning” Around a Hit 

 

In the game, once you have a hit you start probing the adjacent squares (B4, B6, A5, C5) to determine the ship’s shape. Each subsequent coordinate is a targeted test rather than a blind guess.  You shift from random guessing to focused fire, and A/B testing follows the same pattern: 

 

Recipe B in the first test showed an initial improvement (e.g., swapping “Products” and “Solutions”). 

In the second test, you can start tweaking other items in the list.  Remember that the number of possible permutations is far more than you would expect (and with just a few items in your list, the total will exceed a 10x10 Battleships board) so you're not going to hit on the perfect winner immediately (unless you are spectacularly lucky).
 

These follow‑up experiments are nested or sequential, concentrating resources on the most promising hypotheses. The data you collect after each iteration narrows the “search space,” just as each new coordinate reduces the number of possible ship layouts, and continues to point you in the right direction.  This is where multiple recipes are important in each test - you'll need to triangulate your results, not just make one-off guesses.  Multiple recipes with one or two differences between them will help you understand what matters, what helps and what makes things worse!

 

 

Mapping the “Board”: Heatmaps, Click‑maps, and Probability Grids 

 

Advanced Battleship players may use a probability heat‑map that shows where ships are *most likely* to be hidden, based on prior hits and the rules of ship lengths.  This takes a simple game and makes it remarkably complicated, but if you're up for a challenge, it's worth a try.

 

Similarly, modern website analytics provide heatmaps and click‑maps that visualize where visitors tend to gravitate. Tools such as Hotjar, Crazy Egg, or built‑in Google Analytics “Behavior Flow” generate a probability surface over the navigation bar: 

 

When you overlay this visual data onto your A/B test results, you get a dual feedback loop: the statistical lift from the test and the qualitative ink‑blot from the heatmap. Together they guide you toward the next set of coordinates (menu tweaks) with a higher chance of “sinking” the under‑performing items. 

 

 

Defining Victory: From “Sinking a Ship” to “Winning the Funnel” 

 

In Battleships, you win by sinking all opponent ships. In navigation optimization, the win condition is more nuanced, and long-time readers will immediately think of my articles on KPIs - do you have enough KPIs; has everybody agreed on the KPIs, and do you have too many? 

 

Primary goal: Increase CTR for a strategic link (e.g., “Contact”). 

Secondary goals: Reduce bounce rate, shorten time‑to‑conversion, improve overall conversion. 

 

When a series of tests successively “sink” the low‑performing placements and elevate the high‑value ones, you achieve a balanced funnel: the menu now routes users efficiently, minimizing friction. 

 

 The Role of Iteration: The Game Never Truly Ends 

 

Even after every ship is sunk, seasoned players often start a new round, perhaps with a different grid size or ship configuration. Likewise, once a navigation menu reaches a satisfactory conversion rate, the digital landscape continues to shift: new products launch, branding evolves, and user expectations change.  Marketing introduce a new product; a new menu item is added as new features, products or services are promoted.

 

Continuous A/B testing becomes a maintenance routine, akin to a never‑ending Battleships campaign. Each new feature or redesign is a fresh fleet to place, and every subsequent test is another salvo.  Sometimes, it can feel like you're hitting moving targets, which certainly makes things more challenging, and you need to think about the principles you've learned, not just the specifics.

  

Takeaways 

 

* Uncertainty is a feature, not a flaw
Both Battleships and A/B testing thrive on systematically reducing unknowns. 


* Start simple, then target
A modest random or best
‑guess change gives you a baseline, after which data‑driven refinements pay off. 


* Visual probability aids decision
‑making
Heatmaps for navigation function the same way as ship‑probability grids in the classic game. 


* Iterative loops lead to mastery
Each successful test recipe brings you closer to an optimal menu, just as each hit brings you closer to sinking a ship. 

* Never stop playing
The digital environment evolves; continuous testing keeps you ahead, much like an endless series of Battleship rounds. 

 

Final Thought 

 

Next time you stare at a spreadsheet of A/B test results, imagine yourself peering over a Battleships board. Each data point is a coordinate, each statistically significant lift a “hit,” and each subsequent experiment a strategic scan for the rest of the vessel. By embracing the game‑like mindset of exploration, hypothesis, and disciplined iteration, you’ll transform a simple website feature into a high‑performing conduit, efficiently guiding visitors precisely where they want to go, and where you'd like them to go, too.

 

(Happy testing, and may your next “shot” be a clean hit!)

Similar posts I've written about online testing

No comments:

Post a Comment