How to read your results
Your experiment is done and the results page is open. What the winner card, the conviction rating, % of Winner and the heatmap are telling you, how to turn the page into a decision, and how to take it further in chat, as a spreadsheet or as a deck.
You open the results page with a decision waiting. Which message goes in the campaign, which benefit leads the launch, which product gets built. The page gives you two things: an answer, and how far to trust it. Each part of the page serves one or the other, and they are meant to be read together.
- →Two colors, two questions. Pink says how well a variant performed against the winner. Green, orange and red say how firmly you can act on it.
- →The winner is always 100%. Every other variant is read as a share of the winner: 93% of the winner, 81% of the winner.
- →The conviction rating answers two questions: how much of the best result you keep by backing the winner, and how often that pick is safe. Ready to act needs both.
- →The heatmap repeats the same read for every segment: where the winner holds, where it flips, and which rows are too thin to trust.
- →A simulation lands on the same page with the same ratings. Its rating describes the simulated panel's evidence. The live run is the ground truth.
Read it in this order
Four steps, then the rest of this guide explains each part.
- The pill first. It tells you what kind of decision the evidence supports today.
- Then the headline and the bars. How big is the gap, and is there a real second option?
- Then the heatmap. Does the winner hold with the people you care most about? Is a flip worth its own test?
- Then the footer and the next move. Heatseeker AI proposes what to do: a follow-up that isolates why the winner won, a live run to confirm a simulation, or the next question on your roadmap.
If the pill says Don't act yet, ask Heatseeker AI on the results page what it would take to get an answer. Usually it is more budget, a longer run, or a sharper set of variants, and it will tell you which.
Two questions, two colors
Every result answers two separate questions, and the page keeps them on separate colors so one is never mistaken for the other.
- Which option won, and by how much? This is performance, and it is always pink. The deeper the pink, the closer a variant came to the winner.
- Can I act on this? This is conviction, and it uses status colors: green when the evidence supports acting, orange when a leader is emerging, red when the evidence cannot yet tell the variants apart.
A pink bar never tells you how sure the result is. A green pill never tells you how big the gap is. You need both, and the page always shows both.
Your answer: the winner card
Results open on the winner card. The winning ad sits on the left, so you see exactly what people responded to. On the right is the answer.
Lead with “Speed over setup”.
Founders respond to being up and running the same afternoon, not to what it costs them later.
- Speed over setup100%
- Proof from peers93%
- Price that scales81%
- Security first64%
Build the launch around time to value. Keep the peer-proof angle as the second message in retargeting. Price is not what moves this audience, so take it out of the headline.
The headline is the decision in one line, written by Heatseeker AI from the goal you set. On a positioning question it names what to lead with. On a pricing or feature question it names the price or the feature. On an exploratory question it reports what the audience responded to.
The line beneath it is Heatseeker AI's read of why the winner won, in plain words: what the winning message does for your audience.
The pill in the top right is the conviction rating. There are four ratings, from Ready to act to Don't act yet, and the next section explains how each is earned.
Variant performance lists every variant as a share of the winner. The winner is 100%. A variant at 93% is read as 93% of the winner, so it is a real alternative. A variant at 64% is not in the race.
The footer answers the question you came with: what this means for the campaign, the launch or the roadmap. Heatseeker AI writes it from your goal and the evidence above, so it reads as a recommendation, not a readout.
How % of Winner is calculated
Every variant is scored on what people did with it, per person who saw it. Clicks on most experiments, and leads when the experiment collects them. Because the score is per person reached, the winner is the message, not the budget behind it.
Heatseeker does not just rank the raw rates. It asks how much of the result you would keep by backing each variant, allowing for how many people actually saw it: a lead on a few clicks is thinner evidence than a lead on a few hundred. The winner is the variant that keeps the most. Every other variant is read as a share of the winner, so 93% means it keeps about 93% of what the winner keeps.
The percentages fall into four brackets, and the brackets set the pink.
A runner-up at 93% is a real alternative. A competitive variant at 84% is a credible second message. A trailing variant is not in the race.
Conviction: can you act on it?
The pill on the winner card is the conviction rating. It answers the only question left once you have a winner: if I act on this, how likely am I to regret it?
- Ready to act
Backing the winner keeps 95% or more of the best result, and the pick is safe at least nine times in ten.
Enough to commit, within what you tested and who you tested it with. Brief the campaign, pick the message, make the call.
- Strong signal
Backing the winner keeps 90% or more of the best result. The leader is clear but has not cleared the decisive bar.
Enough to back the leader when the decision is reversible or you are narrowing a field. A budget or rollout call usually waits for one more rung.
- Early Signal
Backing the winner keeps 80% or more of the best result. A leader is emerging and could still change.
A directional read. Often enough to narrow a field or shape the next test. For a bigger call, give the experiment more time or budget, or if this was a simulation, run it live.
- Don’t act yet
Backing the winner keeps less than 80% of the best result. The evidence cannot yet tell the variants apart.
No decisive winner yet. Ask Heatseeker AI what to test or gather next to get a read you can act on.
Behind the rating sit two questions, answered from the same evidence as the percentages.
How much would I keep? If you back the winner and it turns out not to be the true best, how much of the best result do you keep on average? Keep 95% or more and you are in Ready to act range. 90% or more is Strong signal. 80% or more is Early Signal. Below that, the evidence cannot yet tell the variants apart.
How often is the pick safe? This is decision confidence: the chance that backing the winner keeps at least 90% of the best result. Ready to act requires it to be nine times in ten or better.
Why both? An average can hide a bad tail, usually when the winner leads on a modest number of clicks. Decision confidence catches that, which is why a result can clear the first bar and still read Strong signal.
While the experiment is running
Until a leader clears the Early Signal line, the pill reads Collecting Data and the bars are gray with no numbers, so nothing on the card suggests a winner before the evidence does. Once a leader clears it, the pill switches to an orange Early Read and the headline reads "Leading so far". The ranking can still change until the experiment settles.
On a simulation
Simulations get the same results page and the same ratings. Two things are different. The simulated audience improves with training: every live experiment you run teaches it more about your customers, and a better-trained panel earns a higher conviction rating on the same result. And the panel grades its own confidence. When it is working from little live evidence, the rating is held down and the card tells you why. A simulation points you to the live experiment worth running. The live run is the ground truth.
Who responds, and how: the heatmap
The winner card tells you what won across the whole audience. The heatmap tells you with whom, and where it flips.
Who responds, and how
NumbersRows are segments, grouped by attribute: age, gender and region on Meta; industry, job function, seniority, company size and country on LinkedIn. A segment appears once every variant has reached at least 100 people in it. That is a display minimum. The Signal column says how strong the evidence in the row is.
Columns are your variants. The overall winner wears the crown in its column header.
Each cell is a variant's % of Winner inside that segment, measured against that segment's own winner. The segment winner wears the crown and reads 100%. The colors are the same four brackets as the winner card.
The Signal column is the conviction rating for that row, on the same scale as the pill: Ready to Act, Strong, Early or None. It is judged on the segment alone, so a row can read Ready to Act while the overall experiment reads Strong signal, and the other way around.
Low-signal rows are segments where the evidence cannot yet tell the variants apart. Their cells are outlined instead of colored, no variant is crowned, and the row sits behind Show low-signal segments at the bottom of the map. A grayed row is not a loss for any variant. It is a segment where the experiment did not collect enough to call.
Under the map, Heatseeker AI answers "What does this mean for your targeting?": where the winner holds, which segments flip to a different variant, and what that means for who you talk to with which message.
Which creative, which audience
If you ran the experiment across several audiences, every creative was shown to every audience, and all of those pairings sit on one ranking. Two panels under the winner card answer the two questions that ranking raises: which creative holds up across audiences, and which audience responds across creatives. Each is averaged evenly across the other, so an audience that was cheaper to reach does not win on volume.
Heatseeker AI writes the line under each panel and says plainly whether one message carries every audience or whether the audiences want different things. That is the decision these panels exist for: one campaign, or one per audience.
Go further: ask, download, present
Once you have the read, four things take it further.
Ask Heatseeker AI. The chat beside the results knows this experiment: the goal you set, every variant, every segment and what the page already says. Ask why a variant won with a segment, whether the winner holds for the people you care most about, or what would firm up a Strong signal. Ask how results vary by age, region or industry and it adds a breakdown card to the page: the variants compared inside each age band, or one variant traced across every band.
Download Data. Every number behind the page as a spreadsheet: the experiment context, every variant overall, and one sheet per segment attribute, for your analysts to take into their own tools.
Create Slide Deck. It opens Heatseeker AI with this experiment loaded and asks two things first: who the deck is for, and what decision it supports. It proposes the takeaway and the closing ask, and on your yes it hands you the file in the chat. The deck speaks the same language as the page, and a deck built from a simulation says so on its slides.
You can also read the same result from your own assistant through Connect to Heatseeker with MCP. There the four ratings appear as Ready to Act, Strong Signal, Early Signal and Gathering Data, and every figure on this page is available by name.
The rest of the page
Smaller parts, for when you need them.
- The experiment card at the top shows what ran: type, product, dates, variants and three figures. Reach is the people who saw at least one variant. Engaged is the number who did something with one. Spent is the media cost in your ad account's currency. On a managed experiment, where Heatseeker supplies the media, it reads Credits.
- Click any bar on the winner card to see that variant's ad.
- Numbers on the heatmap shows the percentage in every cell instead of color alone.
- Delivery by audience shows how the platform actually split your spend. Platforms spend where they convert, so an even split of variants is rarely an even split of budget.
- Share Results invites a colleague into the workspace. Copy as image on the winner card or the heatmap gives you a still for a message or a slide.
For your data team: the method
The statistics behind the pageOpenClose
Heatseeker reads experiments with Bayesian decision analysis, the method modern experimentation platforms have converged on because it answers the question a decision maker actually asks: how much do I stand to lose by acting now?
-
Every variant is a distribution, not a single number. Each variant's click-through rate, and on lead experiments its conversion rate, is modeled as a Beta posterior from what it achieved and how many people it reached. A variant with few clicks is treated as uncertain rather than as precisely measured.
-
Ten thousand simulations. Those distributions are sampled together ten thousand times. In each run the best variant is identified and every variant's shortfall against it is recorded as a share of the best rate.
-
Expected loss decides. The average shortfall is a variant's expected loss, also called expected regret, and the winner is the variant with the lowest. "How much you keep" in the main text is one minus that figure; % of Winner is a variant's figure as a share of the winner's. This is the criterion VWO's SmartStats brought to mainstream A/B testing. It is well defined at any sample size and reads correctly whenever you look, which is what a short experiment with a decision waiting on it needs.
-
Thresholds set by the decision, not by convention. The 5%, 10% and 20% lines between the conviction ratings are tolerances on expected loss, tuned to the cost of acting on a wrong call. They are the same tolerances whether the experiment ran for three days or three weeks.
-
Decision confidence guards the top rating. An average cannot see its own tail: a winner can keep 96% on average and still carry a real chance of a much worse outcome. So Ready to act also requires that at least 90% of the simulations keep the winner's shortfall at or below 10%. On lead experiments the conversion rates share a prior fitted from the experiment itself, with the strength of that sharing treated as uncertain, and the top rating must hold under every version of it.
-
Where Heatseeker takes it further. The engine and everything built on it are Heatseeker's own, designed so that a decision maker reads a decision rather than a statistic. % of Winner puts every variant on one scale anchored at the winner, so the size of the gap is readable at a glance without a control arm or an uplift figure. The conviction rating folds expected loss and decision confidence into four action words, so the first thing on the card is what you can do. The Signal column grades every segment on the same scale as the overall result. Panel conviction caps a simulated result by the training evidence behind it. On a multi-audience experiment, the creative and audience rankings are weighted evenly across the other axis, so a cheaper audience cannot win on delivery volume.
-
One estimator everywhere. Segments on the heatmap, simulated experiments, the spreadsheet, the deck and Heatseeker AI all read the same computation, seeded per experiment so every surface shows the same numbers.