The numbers you see
How a result is called
At each scheduled check, after the shared rules for data and runtime pass:- Equivalence check. If practical equivalence is on and the variation is clearly within the band around zero, the test stops as inconclusive. pagent uses two one-sided tests: the confidence interval at 1 − 2 × the significance level (80 % at Balanced) must lie fully inside the band.
- Significance check. pagent runs a two-proportion z-test on the primary goal. If the p-value is below the threshold for this check:
- and the variation is better, the test stops with the variation as the winner,
- and the variation is worse, the test stops as a Loss.
- Otherwise the test keeps running.
One-sided frequentist tests
Under Advanced settings → Frequentist, you can set the decision policy to One-sided. Then:- the whole significance level goes to the win side, so 10 % means about 1 in 10 false winners,
- the p-value shown is one-sided,
- a variation that is significantly worse still stops the test, as Stopped: variation was losing, without counting as a result.
Settings
The frequentist engine has three settings of its own. All shared settings apply as well.Significance level
The accepted chance of a false result when the variation makes no difference. Default 10 % (Balanced), allowed above 0 % and up to 50 %. Each level sets the significance level to twice its promise, because two-sided tests split it:
A classic “95 % confidence, two-sided” standard is the Careful level. A lower level means fewer false results, but tests need more data to call a real effect.
Sequential correction
On by default and in every level. We recommend keeping it on. pagent looks at a running test many times, once per check. Each look is another chance for random noise to cross the line. Without a correction, a test checked every day for three weeks calls far more false results than the significance level promises. The sequential correction spreads the significance level over the life of the test. It uses O’Brien-Fleming alpha spending:- Early checks need very strong evidence. Only large, clear effects stop a test early.
- The bar drops as data comes in.
- The final check, at the maximum runtime, sits close to the plain significance level.
Decision policy
Two-sided by default, one-sided as an option. See One-sided frequentist tests.How often a false winner is called
Results of 20,000 simulated tests per cell with each level’s frequentist settings, a 3 % conversion rate and one check a day:
False losers are about as frequent as false winners. At high traffic, the wide no-difference bands of Explore and Fast end most tests early as “no meaningful difference”, including tests with a real lift smaller than the band.
Try your own scenario. Pick a level, set the true lift to 0 % to count false results, or to a real lift to see how often it is found and how fast:
What the frequentist engine does not do
- It does not give a “chance to beat control”. A p-value is not the probability that the variation is better.
- It has no expected-loss check. That is a Bayesian setting.
- It has no power setting. pagent does not plan a sample size up front; the maximum runtime sets the end.