Skip to main content
The frequentist engine answers one question: if the variation really made no difference, how surprising would this data be? If the data would be very surprising, pagent calls a result. Choose this method when your team thinks in p-values and confidence levels, or when you need an exact false-winner rate for settings you choose yourself.

The numbers you see

How a result is called

At each scheduled check, after the shared rules for data and runtime pass:
  1. Equivalence check. If practical equivalence is on and the variation is clearly within the band around zero, the test stops as inconclusive. pagent uses two one-sided tests: the confidence interval at 1 − 2 × the significance level (80 % at Balanced) must lie fully inside the band.
  2. Significance check. pagent runs a two-proportion z-test on the primary goal. If the p-value is below the threshold for this check:
    • and the variation is better, the test stops with the variation as the winner,
    • and the variation is worse, the test stops as a Loss.
  3. Otherwise the test keeps running.
By default the test is two-sided: the significance level is split between a false win and a false loss. At Balanced (10 %), about 1 in 20 tests where nothing changed ends as a false winner and 1 in 20 as a false loser.

One-sided frequentist tests

Under Advanced settings → Frequentist, you can set the decision policy to One-sided. Then:
  • the whole significance level goes to the win side, so 10 % means about 1 in 10 false winners,
  • the p-value shown is one-sided,
  • a variation that is significantly worse still stops the test, as Stopped: variation was losing, without counting as a result.
The levels always use two-sided frequentist tests. A one-sided policy shows as Custom settings.

Settings

The frequentist engine has three settings of its own. All shared settings apply as well.

Significance level

The accepted chance of a false result when the variation makes no difference. Default 10 % (Balanced), allowed above 0 % and up to 50 %. Each level sets the significance level to twice its promise, because two-sided tests split it: A classic “95 % confidence, two-sided” standard is the Careful level. A lower level means fewer false results, but tests need more data to call a real effect.

Sequential correction

On by default and in every level. We recommend keeping it on. pagent looks at a running test many times, once per check. Each look is another chance for random noise to cross the line. Without a correction, a test checked every day for three weeks calls far more false results than the significance level promises. The sequential correction spreads the significance level over the life of the test. It uses O’Brien-Fleming alpha spending:
  • Early checks need very strong evidence. Only large, clear effects stop a test early.
  • The bar drops as data comes in.
  • The final check, at the maximum runtime, sits close to the plain significance level.
See the p-value a variation needs at each daily check: pagent measures progress by visitors. It estimates how many visitors the test will collect by its maximum runtime from the traffic of the first check. To keep a decision possible at every check, the progress never runs ahead of the share of runtime that has passed. With the correction off, every check uses the plain threshold, and false results are several times more frequent than the significance level. Only turn it off if you look at the result once, at a fixed end date. In that case, also turn automatic stopping off and stop the test yourself on that date.

Decision policy

Two-sided by default, one-sided as an option. See One-sided frequentist tests.

How often a false winner is called

Results of 20,000 simulated tests per cell with each level’s frequentist settings, a 3 % conversion rate and one check a day: False losers are about as frequent as false winners. At high traffic, the wide no-difference bands of Explore and Fast end most tests early as “no meaningful difference”, including tests with a real lift smaller than the band. Try your own scenario. Pick a level, set the true lift to 0 % to count false results, or to a real lift to see how often it is found and how fast:

What the frequentist engine does not do

  • It does not give a “chance to beat control”. A p-value is not the probability that the variation is better.
  • It has no expected-loss check. That is a Bayesian setting.
  • It has no power setting. pagent does not plan a sample size up front; the maximum runtime sets the end.