The numbers you see
How a result is called
At each scheduled check, after the shared rules for data and runtime pass. The numbers are the Balanced level:- Practical equivalence. If ROPE is on and there is at least a 95 % chance that the true lift is inside the band (±7.5 %), the test stops as inconclusive.
- Winner. If the chance to beat control reaches the threshold (98.5 %), the variation wins.
- Losing variation. If the chance falls to the stop losing threshold (5 %) or below, the test stops as Stopped: variation was losing. This protects your conversions but is not counted as a result.
- Expected loss (Careful and Strict). If the expected loss check is on, the test only stops when the expected loss of the leading side is below the expected loss threshold.
- Otherwise the test keeps running.
Priors
Before a test has data, pagent has to start from some assumption. That starting point is the prior.Skeptical prior (default)
Most changes to a website move conversion by a few percent at most. A +40 % lift on day two is almost always noise. The skeptical prior builds this in: it starts from “this change probably does little” and pulls early, extreme lifts toward zero.- With little data, the pull is strong. A big early jump counts for less.
- With more data, the pull fades. The data decides.
Flat prior
The flat prior assumes nothing. Every conversion rate between 0 % and 100 % is equally likely before the test starts, so the chance to beat control follows the raw data from the first visitor on. Websites that existed before October 2026 were kept on the flat prior when the skeptical prior was introduced. Choosing a level switches them to the skeptical prior.How often a false winner is called
pagent checks a running test many times, by default once a day. Every check is another chance for noise to cross the threshold, so a single threshold does not tell you how often a whole test ends with a false winner. That is why the levels use high thresholds: each was set with simulations of tests where the variation does nothing, checked daily, at 100 to 50,000 visitors per variation per day, so that even the worst traffic stays within the level’s promise. Results of 20,000 simulated tests per cell with each level’s Bayesian settings, a 3 % conversion rate and one check a day:
Read the found columns left to right as 300, 3,000 and 30,000 visitors per variation per day.
- Every level keeps its promise at every traffic level.
- Low traffic finds little. At 300 visitors a day, a 5 % lift is rarely found within the runtime at any level. Test bolder changes, or pages with more traffic.
- Fast levels skip small lifts at high traffic. Their wide no-difference band (±15 % for Explore, ±10 % for Fast) ends tests with a real but small lift as “no meaningful difference”. If lifts around 5 % matter, use Balanced or stricter.
- About 1 in 7 tests that change nothing stop as losing at Balanced and 3,000 visitors a day. These are not results, so they add no false losers.
Expected loss
The chance to beat control tells you how likely the variation is better. It does not tell you how much is at stake if it is not. Expected loss covers that: it is the average conversion rate you give up by shipping the variation, counting only the cases where it is actually worse. Example: control converts at 3.0 %. An expected loss of 0.02 pp means that shipping the variation costs, on average, 3.0 % → 2.98 % in the unlucky cases. That is a small risk. At Explore, Fast and Balanced, expected loss is shown but does not stop tests. Careful and Strict require it to be small, 0.1 pp and 0.05 pp, before a test stops. You can turn the check on for any setup with Enable expected loss threshold.Sequential correction (deprecated)
Older websites may still have the Bayesian Sequential correction on. It raises the threshold at early checks, following the same O’Brien-Fleming schedule as the frequentist correction, and the results page then shows the raised threshold. It is being retired. You can switch it off, but not back on, and tests are two-sided while it is on. Choosing a level switches it off.What the Bayesian engine does not do
- It has no p-value. The chance to beat control is not 1 minus a p-value.
- It has no exact false-winner rate for a custom threshold. The levels’ promises come from simulation; for an exact rate, use the frequentist engine.
- It has no power setting. The maximum runtime sets the end of a test.