Skip to main content
A test shows each visitor either the original page (control) or a variation. pagent counts how many visitors of each group complete the test’s primary goal. From those counts, it decides whether the variation is better, worse or no different. This page explains how that decision is made. The other pages in this section go deeper:

Statistics settings

The five levels, every setting, its default, and when to change it.

Bayesian statistics

Chance to beat control, priors and expected loss.

Frequentist statistics

p-values, significance level and sequential correction.

Variations and goals

Tests with several variations, goals, revenue and traffic checks.

Choose your setup

Pick a level and match your company’s standard.

Method and level

Under Settings → Tests → Statistics, you make two choices:
  1. Method: Bayesian or frequentist. Both use the same data and the same rules for runtime and data. They differ in the question they ask.
  2. Level: from Explore (fastest, at most 1 in 7 false winners) to Strict (at most 1 in 100). The default is Balanced: at most 1 in 20 false winners among changes that do nothing.
See Statistics levels for what each level sets. Each website has a default method. A test takes it when it launches. You can switch a running or paused test to the other engine; see Statistics settings.

When pagent checks a test

pagent updates a test’s numbers every hour. Updating the numbers never ends a test. A test can only be ended at a scheduled check. By default there is one check a day, counted from the moment the test started. You can choose 1, 2, 4 or 24 checks a day with Automatic checks per day. A check runs at the first hourly update after it is due. The results page shows when the next check is due, under Automatic decisions and in the side panel under Analysis → Next check.

The order of the checks

At every scheduled check, pagent goes through these rules in order. The first rule that applies decides what happens. The numbers are the Balanced defaults.
1

Enough data?

Control and the variation together need at least the minimum total conversions (100). If not, the test keeps running.
2

Ran long enough?

The test must have run for the minimum runtime (7 days). If not, the test keeps running.
3

Out of time?

Once the maximum runtime (21 days) is reached, the test stops. It ends with a result if the final evidence passes the engine’s bar, and as inconclusive otherwise. If steps 1 or 2 still fail at that point, it ends as inconclusive.
4

No meaningful difference?

If practical equivalence is on and the data shows the variation is almost certainly within ±7.5 % of control, the test stops as inconclusive. Further testing would not find a difference worth acting on.
5

Clear winner?

If the evidence passes the bar, the variation wins and the test stops.
6

Clearly losing?

Under the one-sided decision policy (Bayesian default), the test stops when the variation’s chance to beat control falls to 5 % or less, recorded as Stopped: variation was losing. Under the two-sided policy (frequentist default), a clear loser is a Loss.
7

Small enough risk? (Bayesian, Careful and Strict)

If the expected loss check is on, a Bayesian result also needs a small expected loss before the test stops.
If no rule ends the test, it keeps running until the next check.

Possible results

The test list has tabs for Lost, Stopped: losing, Inconclusive and Failed (tests that errored or were aborted).

Who decides the result

Every finished test shows who decided its result.
  • Decided by pagent. pagent ended the test at a scheduled check, at the maximum runtime, or because no visitors arrived (see Maximum idle days).
  • Decided by you. You stopped the test yourself, or changed the result afterwards.
Stopping a test yourself. Admins can open a running test and choose Stop test. The dialog shows pagent’s current reading and asks for Your result: Win, Inconclusive or Loss. If you don’t pick one, it is recorded as inconclusive. pagent does not decide a result when you stop a test, because it only decides at its scheduled checks. Changing a result. On a finished test, admins can choose Change result… in the test’s menu. Your result then counts instead of pagent’s, also for a test that was stopped because the variation was losing. Reports and learnings show that you decided it. Use pagent’s verdict switches back. Shipping is a separate decision. A result never ships anything by itself. Use Mark as shipped… when you have built a variation into your site.

How runtime is counted

  • Runtime starts when the test launches and is counted in days, including parts of a day.
  • Paused time counts. Pausing a test does not stop its clock.
  • If you resume a test after its maximum runtime has passed, it ends right away.

Tests with more than one variation

Each variation is compared with control on its own. pagent makes the bar for a winner stricter for each comparison, so that more variations do not mean more false winners. When the best variation’s comparison ends the test, the whole test stops for all variations. See Variations and goals.