Statistics levels
Most teams only need two choices. Go to Settings → Tests and pick them under Statistics:- Method: Bayesian (how likely is the variation better?) or Frequentist (would this result be rare if nothing changed?).
- Level: how much proof a winner needs.
What a level sets for each method:
Every level uses the skeptical prior and the one-sided decision policy for Bayesian tests, and the two-sided policy with sequential correction for frequentist tests. A Bayesian test also stops when the variation is clearly losing; that stop is not counted as a result.
The band hides small lifts. A level ends a test as “no meaningful difference” once the effect is almost certainly inside its band. That is how fast levels stay fast. It also means Explore and Fast rarely call a real lift smaller than their band, especially with a lot of traffic. If lifts of 5 % matter to you, use Balanced or stricter.
Advanced settings and custom settings
Below the picker, Advanced settings holds every rule on three tabs: Shared, Bayesian and Frequentist. Changing any value that belongs to a level turns the level into Custom settings. Choosing a level again replaces those values. Rules that are not part of a level stay as you set them: automatic stopping, checks per day, maximum idle days and the absolute ROPE band. Click Save changes to apply. Reset discards unsaved edits. Only admins can change these settings.Where settings live
Settings come from three levels. The most specific one wins:- System default: pagent’s defaults, the Balanced level.
- Website setting: your choices under Settings → Tests.
- Test override: a value set for one test only.
See the rules of one test
Open a test, click the test’s actions menu (⋮) and choose View test rules. The dialog shows:- the test’s Level, or Custom settings if its rules match none of the five levels,
- the Statistics engine,
- every rule the test runs by, marked Test override, Website setting or System default.
Override rules for one test
You can override rules for a single test before it launches:- In the pagent Chrome extension, open the settings menu in the toolbar and choose Experiment rules. Change a value, or use Reset on a field to inherit it again.
- When you review a test, you can ask for different rules in your feedback, for example “Let it run for at least 14 days”. pagent applies them before launch.
- In View test rules, a draft can pick its own engine and decision policy.
Statistics engine
Which method decides the test. See Bayesian statistics and Frequentist statistics.
- Draft tests follow the website default unless you pick an engine. The engine is fixed when the test launches.
- Running and paused tests can switch. pagent recalculates the results with the new engine. Collected data and earlier decisions are kept, and later checks use the new engine’s rules.
- Finished tests keep the engine that decided them.
Decision policy
Which directions count as a result. On the Bayesian and Frequentist tabs of Advanced settings.
- One-sided: only a win counts as a result. A variation that is clearly losing still stops the test, to protect your conversions, but it is recorded as Stopped: variation was losing, not as a loss. It does not count as a learning, a loss or a hypothesis outcome. You can still record your own result for the test.
- Two-sided: a clear win and a clear loss both count as results.
Settings for both engines
These are on the Shared tab. They control timing, data requirements and practical equivalence.Automatic stopping
When on, pagent ends the test at a scheduled check once the rules are met.
When off, pagent never ends the test. It still analyses the data every hour and shows its reading and recommendation, but you stop the test yourself. The test keeps running past its maximum runtime, and Maximum idle days does not apply.
Turn it off when you have to end tests on a fixed date, for example for a campaign.
The setting is stored as Require manual stop. Require manual stop = on means automatic stopping is off.
Automatic checks per day
How often pagent may end a test. Checks are spaced evenly from the moment the test started, and each runs at the first hourly update after it is due. A missed check is skipped, not repeated.
More checks let a clear result end the test sooner. With the Bayesian engine, they also give random noise more chances to look like a winner: the levels are calibrated for one check a day. The frequentist sequential correction accounts for the extra checks.
Require minimum data
When on, the test cannot end with a result before it has the minimum total conversions. We recommend keeping it on: with very few conversions, a couple of lucky visitors can swing the result.
Minimum total conversions
The number of conversions control and the variation need together before pagent may call a result. Conversions are counted once per visitor on the primary goal. In a test with several variations, each variation counts with control on its own.
The same number drives the data warnings on the results page:
- Insufficient data: no conversions, or fewer than half the minimum.
- Low confidence: below the minimum, or control or the variation still has no conversions.
Minimum runtime days
The number of days a test must run before it can end with a result. Days are counted from the launch and include paused time.
Visitors behave differently on Monday than on Saturday. Keep at least 7 days so every test covers a full week. Use 14 if your business has a two-week rhythm, such as paydays.
Maximum runtime days
The latest day a test can run. At the first check after this, the test stops even if the evidence is not clear:
- If the evidence passes the engine’s bar, the result is a win, a loss, or Stopped: variation was losing under the one-sided policy.
- If not, or if there is still not enough data, the result is inconclusive.
Maximum idle days
If a test has had no visitors at all since it launched, pagent stops it after this many days as inconclusive: “No visitors reached the test in time.” This usually means the test does not run on your site, for example because the page or audience never matches.
A test that had visitors and then stops getting them is not stopped by this rule. The rule is off when automatic stopping is off.
Practical equivalence (ROPE)
Some tests will never find a difference worth acting on. Practical equivalence stops these tests early, so your traffic goes to the next idea. ROPE stands for “region of practical equivalence”: a band around zero where you treat a change as “no real difference”.Enable ROPE check
When on, the test stops as inconclusive once the data shows the true effect is almost certainly inside the band:
- Bayesian: at least a 95 % chance that the effect is inside the band.
- Frequentist: the confidence interval of the effect lies completely inside the band (two one-sided tests at the 1 − 2 × significance level).
ROPE mode
How the band is measured:
- Relative: as a share of control’s conversion rate. A 7.5 % band at a 4 % conversion rate covers 3.7 % to 4.3 %.
- Absolute: in percentage points of conversion rate, the same whatever the conversion rate.
Relative ROPE band
The band when ROPE mode is relative: effects between −7.5 % and +7.5 % count as no real difference. Set it below the smallest lift you would still act on.
Absolute ROPE band
The band when ROPE mode is absolute, in percentage points. The default of 1 pp is wide for most sites: at a 3 % conversion rate it treats anything from 2 % to 4 % as no difference. If you use absolute mode, set a band that fits your conversion rate.
Bayesian settings
These are on the Bayesian tab and apply to tests that use the Bayesian engine. Bayesian statistics explains the ideas behind them.Chance to beat control threshold
How sure pagent must be before it calls a winner. The variation wins when its chance to beat control reaches the threshold.
Under the two-sided policy, the variation also loses when its chance falls to 100 % minus the threshold. Under the one-sided policy, losers are handled by Stop losing variations at.
With more than one variation, pagent raises the threshold for each comparison automatically; see Variations and goals.
Stop losing variations at
One-sided tests only. pagent stops the test when the variation’s chance to beat control falls to this value or below. The stop protects your conversions and is not recorded as a loss.
A higher value stops losing variations sooner. It does not add false winners: it only ends tests that were heading nowhere. It is not adjusted for several variations.
Prior
What pagent assumes before it sees any data.
- Skeptical starts from “this change probably does little”. Large lifts are treated as rare, so a big jump in the first days counts for less. As data comes in, the data takes over.
- Flat assumes nothing and reacts fully to early data. Flat is not part of any level.
Websites that existed before October 2026 were kept on Flat when the skeptical prior was introduced, so their results did not change. They show Custom settings until you pick a level.
Prior width
Only with the skeptical prior. How large a lift the prior still considers normal. At 30 %, about two thirds of true lifts are expected to fall between −30 % and +30 %. A smaller width is more skeptical and holds back early results longer.
Sequential correction (deprecated)
Raises the threshold at early checks to make up for repeated checking. It is being retired for Bayesian tests. It is only shown where it is still on, and once you switch it off, it cannot be switched back on. While it is on, Bayesian tests are two-sided. Choosing a level switches it off.
Enable expected loss threshold
When on, a result also needs a small expected loss before the test stops. The chance to beat control tells you how likely the variation is better. Expected loss tells you how much you would lose, on average, if you picked the wrong side. Careful and Strict turn it on.
Expected loss threshold
The largest expected loss you accept, in percentage points of conversion rate. At 0.1 pp, choosing the variation may cost on average at most 0.1 percentage points, for example from 3.0 % to 2.9 %. Only used when the expected loss check is on.
Frequentist settings
These are on the Frequentist tab and apply to tests that use the frequentist engine. Frequentist statistics explains them in detail.Significance level
The accepted chance of a false result when the variation makes no difference. Two-sided tests split it between a false win and a false loss: at 10 %, about 1 in 20 such tests ends as a false winner and 1 in 20 as a false loser. One-sided tests spend all of it on the win side. On the settings page, enter it as a fraction: 0.1 for 10 %.
Sequential correction
Keeps the false-winner rate at the significance level although pagent checks the test many times. Early checks need very strong evidence; the bar relaxes as data comes in. Keep it on unless you look at the result only once, at a fixed end date.
What you cannot set
- Power or sample size. pagent does not plan a sample size up front. The maximum runtime sets how long a test may run.
- Credible and confidence interval levels. Bayesian intervals are always 95 %. Frequentist intervals follow the significance level.
- The traffic split check. It is always on at a fixed level; see Variations and goals.