Statistics settings
The five levels, every setting, its default, and when to change it.
Bayesian statistics
Chance to beat control, priors and expected loss.
Frequentist statistics
p-values, significance level and sequential correction.
Variations and goals
Tests with several variations, goals, revenue and traffic checks.
Choose your setup
Pick a level and match your company’s standard.
Method and level
Under Settings → Tests → Statistics, you make two choices:- Method: Bayesian or frequentist. Both use the same data and the same rules for runtime and data. They differ in the question they ask.
- Level: from Explore (fastest, at most 1 in 7 false winners) to Strict (at most 1 in 100). The default is Balanced: at most 1 in 20 false winners among changes that do nothing.
Each website has a default method. A test takes it when it launches. You can switch a running or paused test to the other engine; see Statistics settings.
When pagent checks a test
pagent updates a test’s numbers every hour. Updating the numbers never ends a test. A test can only be ended at a scheduled check. By default there is one check a day, counted from the moment the test started. You can choose 1, 2, 4 or 24 checks a day with Automatic checks per day. A check runs at the first hourly update after it is due. The results page shows when the next check is due, under Automatic decisions and in the side panel under Analysis → Next check.The order of the checks
At every scheduled check, pagent goes through these rules in order. The first rule that applies decides what happens. The numbers are the Balanced defaults.1
Enough data?
Control and the variation together need at least the minimum total conversions (100). If not, the test keeps running.
2
Ran long enough?
The test must have run for the minimum runtime (7 days). If not, the test keeps running.
3
Out of time?
Once the maximum runtime (21 days) is reached, the test stops. It ends with a result if the final evidence passes the engine’s bar, and as inconclusive otherwise. If steps 1 or 2 still fail at that point, it ends as inconclusive.
4
No meaningful difference?
If practical equivalence is on and the data shows the variation is almost certainly within ±7.5 % of control, the test stops as inconclusive. Further testing would not find a difference worth acting on.
5
Clear winner?
If the evidence passes the bar, the variation wins and the test stops.
6
Clearly losing?
Under the one-sided decision policy (Bayesian default), the test stops when the variation’s chance to beat control falls to 5 % or less, recorded as Stopped: variation was losing. Under the two-sided policy (frequentist default), a clear loser is a Loss.
7
Small enough risk? (Bayesian, Careful and Strict)
If the expected loss check is on, a Bayesian result also needs a small expected loss before the test stops.
Possible results
The test list has tabs for Lost, Stopped: losing, Inconclusive and Failed (tests that errored or were aborted).
Who decides the result
Every finished test shows who decided its result.- Decided by pagent. pagent ended the test at a scheduled check, at the maximum runtime, or because no visitors arrived (see Maximum idle days).
- Decided by you. You stopped the test yourself, or changed the result afterwards.
How runtime is counted
- Runtime starts when the test launches and is counted in days, including parts of a day.
- Paused time counts. Pausing a test does not stop its clock.
- If you resume a test after its maximum runtime has passed, it ends right away.