Statistics and reading results
Last updated: September 12, 2026
Pick one primary metric when you create a test. The app judges the test on that metric, labels the result, and tells you whether to keep waiting.
What you see per variant
The results table shows one row per variant:
- Visitors - visitors bucketed into this variant who passed the audience rules. Every rate in the row divides by this number.
- Sessions - visits those people made. Shown beside visitors because most merchants think in sessions, but it is not the denominator.
- Your primary metric - the number the verdict comes from.
- Orders and Revenue - net of refunds.
Above the table sits the verdict, the change against control, a confidence range, and a “Which version wins” chart.
Conversions are counted per visitor. Three orders from the same person count as one converting visitor, so a rate can never pass 100%. The one exception is a custom event judged per order, where both sides of the rate are orders: orders carrying the event over orders placed.
Primary metrics you can choose
| Metric | What it measures |
|---|---|
| Revenue per visitor (default) | Revenue divided by bucketed visitors. Conversion rate and order value in one number, so a variant cannot win by selling to more people at a smaller basket. |
| Conversion rate | Share of bucketed visitors who placed an order. |
| Average order value | Revenue divided by orders. |
| Add to cart rate | Share of bucketed visitors who added to cart at least once. |
| A custom event | Share of bucketed visitors who fired that event at least once. Every custom event you have defined appears in the picker. |
Checkout started is not offered as a primary metric. Shop Pay and Apple Pay never load a checkout page the pixel can see, so the number would move whenever a variant changes how many buyers take the wallet path.
Sessions cannot produce a verdict. Session count per variant is a function of the traffic split, so it measures whether bucketing is even, not whether a variant is better. Older tests set to sessions keep working and report that no winner can be declared.
Which test runs
| Primary metric | Test |
|---|---|
| Conversion rate, add to cart rate, any custom event | Two-proportion z-test |
| Revenue per visitor | Welch’s t-test on per-visitor revenue, counting non-buyers as zero |
| Average order value | Welch’s t-test on order values |
Revenue per visitor counts every bucketed visitor, buyers and non-buyers alike, which is why it needs more traffic than conversion rate to resolve. That is honest rather than pessimistic: per-visitor revenue is a genuinely noisy number.
Order value is compared using an assumed spread of order sizes typical for ecommerce rather than a measured one. If your store mixes small accessories with furniture-scale orders, read order-value verdicts with extra caution.
Both revenue metrics need at least two orders in each variant before a verdict is possible. Below that you get “Not enough orders yet to compare revenue per visitor” or “Not enough orders to compare order values yet.”
Minimum data
Each variant needs at least 30 visitors before any significance is calculated. Below that, the result reads:
Too little data. Need at least 30 sessions per variant.
Separately, the page will not name a leader at all until every variant has 2,000 visitors. Under that, it reports what happened (visitors, orders, revenue) and stays quiet about direction. An early lead built on three orders flips on the fourth, and a “leading” badge invites you to act on it.
Confidence level and multiple variants
Pick a confidence level when you create the test:
| Confidence | Use for |
|---|---|
| 80% | Low-traffic shops that want “probably better” rather than proof. One in five clear winners at this level is noise |
| 90% | Cheap, reversible changes |
| 95% (default) | Most tests |
| 99% | Expensive or hard to reverse changes |
With more than two variants, each challenger is compared against the control and the bar is tightened by the number of comparisons. Four challengers means the threshold is four times stricter. This is the right call, and it means a multi-variant test needs more traffic per arm before anything reads as a clear winner.
The verdicts
| Verdict | What it means | Message |
|---|---|---|
| Clear winner | The difference clears your confidence level. | ”Clear winner. Enough data.” |
| Looks promising | Close to the bar but not over it. | ”Looks promising. About N more sessions recommended.” |
| Unclear | Nowhere near the bar, and there is still room for the answer to change. | ”Unclear. Need about N more sessions.” |
| No difference | The whole confidence range sits inside your tie band. The test has answered its question with a no. | ”No difference. Both versions are within 10% of each other on this metric.” |
| Too little data | Under 30 visitors per variant, or under two orders per variant on a revenue metric. | ”Too little data. Need at least 30 sessions per variant.” |
Once a test has enough traffic that no meaningful difference could be hiding, “Unclear” becomes “No clear winner. Any difference so far is too small to act on.”
The tie band
A p-value cannot say “these are the same”. It can only fail to find a difference, and it fails the same way whether the test has fifty visitors or a hundred thousand. So the app judges the confidence range instead: once the whole range sits inside a band you would not act on, the result reads No difference.
The band defaults to 10% relative and you can set it per test under Call it a tie within, anywhere from 1% to 50%. A clear winner always outranks a tie: a large test can resolve a real 3% gain to a range of 1% to 5%, which is inside a 10% band and still worth shipping.
Which version wins
The chart under the verdict gives each variant a share of the wins, plus a Too close to call bar using the same tie band. The three bars add up to 100%. Two versions that behave identically land almost entirely in “Too close to call” rather than splitting 50/50 and making one bar look meaningfully taller.
It needs the same 30 visitors per variant, and two orders per variant on a revenue metric. It is there to describe the picture. The verdict above it is what decides the test.
How many more visitors
The “about N more” figure is how much traffic the test needs before it can answer, sized at the confidence level you picked for the test and at 80% power. A lower level needs less traffic: 90% needs about a fifth less than 95%, 80% a little over half. By default it plans for a 20% relative change on your current rate. Set Minimum detectable effect in percentage points if you only care about a larger change and want a smaller target.
Every chart on the results page carries its own confidence signal beside its title, not just the metric the test is judged on: four small bars lit from one to four, red to green, for how confident the reading is. Four green bars appear only once the reading is decided. Conversion rate, order value and revenue per visitor are each read with the same test and the same confidence level, and each says one of: likely better, likely worse, no real difference, no sign of harm (the change cannot be worse than your equivalence band) or too early to tell. Hover the signal for how confident the reading is (None, Low, Moderate or High), the change, its range, the chance each version is ahead and the call. While a reading is undecided, the call follows the bars: one bar is too early to tell, two is an early signal that the variant is ahead or behind, and three is leaning better or worse. None of those is a result yet. Use them as guardrails: a variant that wins its metric while conversion or order value reads likely worse is not a win.
Stopping automatically
Turn on Auto-stop on significance (Growth and above) and the test marks itself complete the moment its primary metric reaches a clear winner. It is checked hourly, and it always reads the test’s whole run rather than the date range you happen to be looking at.
Losing tests are never stopped for you. A test that has run past its recommended traffic without a winner stays running until you end it.
When to call a test
- Don’t peek and stop. Checking every hour and stopping the moment something looks like a winner inflates your chance of shipping noise. The math assumes you let the test run.
- Don’t run forever either. Once the verdict has settled and the range is inside your tie band, more traffic will not change the answer.
- A “no difference” is a result. It tells you the change is not worth the build cost. Archive it and test something else.
- Several tests at once is fine. Visitors get an independent assignment for each. The app corrects for the variants inside one test, not across separate tests, so the more you run at once, the more likely one of them looks like a winner by chance.
What gets counted
A visitor counts toward a variant when they passed the audience rules and were bucketed into it.
A visitor with no assignment counts for no variant. If the app could not bucket someone (an audience rule excluded them, the shop was over its plan’s session cap, or the storefront could not reach us), they see the control markup but they are recorded against neither arm. Folding them into control would mix a biased group into one side of the experiment. Their events are still recorded, they just do not count for the test.
Sessions are counted as the distinct sessions behind a variant’s events, not as a separate marker, so the same visitor returning across three days is one visitor and three sessions.
Refunds subtract from revenue, and a fully refunded order stops counting as an order. Refunds do not remove add-to-cart or custom-event counts, because those things still happened.
Date ranges
Results default to the whole test. A verdict computed over the last 30 days of a 60-day test is a verdict on half the evidence. You can switch to today, yesterday, the last 7, 30, 90 or 365 days, and the verdict, the chart and the CSV export all follow the window you picked.
Revenue attribution mechanics
Orders attribute through the visitor id, not through the checkout.
When a visitor is bucketed, the embed stores a visitor id in the splt_usr_id cookie (90 days) and writes the same value onto the cart as the splt_user_id cart attribute. Shopify carries cart attributes into the order, so when the order webhook arrives, the order names the visitor, and the visitor names the variant.
This is why attribution survives the paths a storefront script cannot see:
- Express checkout. Shop Pay hands the buyer to a different domain and Apple Pay opens a native sheet. Neither loads a checkout page the pixel runs on. The cart attribute travels with the order regardless.
- Late orders. Abandoned cart recovery and slow shoppers can order days after the session that bucketed them. The 90-day cookie is why those still land on the right variant.
- Refunds. A refund arrives with its order id and nets that exact order’s revenue, however long after the sale it lands.
Two things to know:
- Cross-device orders (start on mobile, finish on desktop) attribute only when Shopify’s cart merge carries the attribute across. If your theme or a custom checkout clears cart attributes on merge, those orders do not attribute. Rare on a vanilla theme, possible on a heavily customised one.
- An order links only to an assignment the visitor already had when they bought. An order that arrives during an outage and is processed later will not be back-dated onto a variant the buyer was bucketed into afterwards.