Skip to main content

Sequential Inference

Safe Testing

Safe testing uses e-values for fixed analyses and e-processes for anytime validity. Reverse information projection and growth criteria guide composite-hypothesis constructions; FormalSLT covers the finite-horizon supermartingale case.

AdvancedResearchTier 1FrontierCore spine~50 min
For:MLStatsResearch

Learning position

Place this page in a reading path.

sequential-inference | layer 3 | tier 1. This page has 6 direct prerequisites and 0 published dependents.

What next

Anytime-Valid Inference

This is the first curated or graph-derived continuation from the current page.

Evidence badge

Source-grounded page

This page has no public Lean mapping yet. Use the evidence page to inspect how claim status labels work.

Show the backing system

Why This Matters

Classical Neyman-Pearson testing fixes a sampling plan, a null, and an alternative. In the simple-versus-simple case it constructs a most-powerful level-α\alpha test for that fixed experiment. Reusing a fixed-sample calibration after unplanned monitoring or data-dependent continuation generally invalidates its Type I error guarantee. Sequential boundaries and alpha-spending repair specific monitoring plans, but they must be built into the analysis.

Safe testing, formalized by Grünwald, de Heide, and Koolen (2024, Journal of the Royal Statistical Society Series B, 86(5)), starts with a nonnegative statistic whose expectation is at most one under every distribution in the null. Such a statistic is an e-value and gives a valid threshold test at one analysis time. Sequential validity needs an additional construction: an e-process, or a running product of conditional e-values whose next factor is chosen predictably from past information.

The common currency matters because e-values can be multiplied across conditionally valid studies, mixed using weights fixed before the data used by the mixture, and monitored when assembled into an e-process. Likelihood ratios and some Bayes factors supply important examples. This does not make every classical test a safe test at the same rejection threshold, and it does not imply a universal fixed-sample power penalty. Power depends on the model, alternative, e-value construction, and stopping design.

Formal Setup

Definition

Safe test

A safe test at level 0<α≤10<\alpha\leq1 for a null hypothesis H0H_0 is a function S(X)S(X) of the data such that S(X)≥1/αS(X) \geq 1/\alpha implies rejection, and SS is an e-value: EP[S]≤1\mathbb{E}_{P}[S] \leq 1 for every P∈H0P \in H_0. The Type I error of the rejection rule is at most α\alpha by Markov's inequality.

Definition

Safe anytime-valid test

A safe anytime-valid test at level 0<α≤10<\alpha\leq1 is a sequential procedure with rejection rule "stop and reject the first time St≥1/αS_t \geq 1/\alpha" where (St)(S_t) is a nonnegative process adapted to the stated filtration and satisfies

sup⁡P∈H0sup⁡τEP[Sτ]≤1,\sup_{P\in H_0}\sup_{\tau}\mathbb{E}_P[S_\tau]\leq 1,

with the second supremum over stopping times in the stated filtration. An e-process need not itself be a nonnegative supermartingale. Every test supermartingale is an e-process, and useful e-processes exist that are not test supermartingales.

Definition

Reverse information projection (RIPr)

Let QQ be an alternative distribution and let P0\mathcal{P}_0 denote the convex set of Bayes mixtures of distributions in the composite null. Under the support and finite-KL conditions in Grünwald, de Heide, and Koolen's Theorem 1, the RIPr P0∗P_0^* minimizes D(Q∥P)D(Q\|P) over P∈P0P\in\mathcal{P}_0, or is the limiting sub-probability distribution when the infimum is not attained by a proper mixture. Writing qq and p0∗p_0^* for the corresponding densities, the ratio q(X)/p0∗(X)q(X)/p_0^*(X) is an e-value for the full null and maximizes expected log growth under QQ. In regular examples P0∗P_0^* is a proper null mixture; this is not part of the definition in complete generality.

The Core Result

The construction that makes safe testing work is the e-value itself.

Theorem

E-Value Construction Yields a Safe Test

Statement

Fix 0<α≤10<\alpha\leq1. Let S(X)S(X) be a nonnegative function of the data XX with sup⁡P∈H0EP[S]≤1\sup_{P \in H_0} \mathbb{E}_P[S] \leq 1. The test that rejects H0H_0 when S(X)≥1/αS(X) \geq 1/\alpha is a safe test at level α\alpha: its Type I error is at most α\alpha, uniformly over H0H_0.

If additionally (St)(S_t) is an e-process for H0H_0, the sequential rule "stop and reject at the first tt with St≥1/αS_t \geq 1/\alpha" is a safe anytime-valid test at level α\alpha: Type I error at most α\alpha for every stopping rule.

Intuition

Markov's inequality gives the single-shot guarantee. The stopped-value definition gives the sequential guarantee directly; for a nonnegative test supermartingale, Ville's inequality gives the familiar running-maximum form. The construction does not optimize for power; it guarantees validity. Power comes from the choice of e-value or e-process.

Proof Sketch

Single-shot: Pr⁡P(S≥1/α)≤α⋅EP[S]≤α\Pr_P(S \geq 1/\alpha) \leq \alpha \cdot \mathbb{E}_P[S] \leq \alpha by Markov.

Sequential: apply the stopped-value property to the bounded hitting times τα∧T\tau_\alpha\wedge T, use Markov's inequality, and let T→∞T\to\infty. This gives Pr⁡P(sup⁡tSt≥1/α)≤α\Pr_P(\sup_t S_t\geq1/\alpha)\leq\alpha. Equivalently, for each P∈H0P\in H_0 one may use a nonnegative PP-supermartingale with initial value at most one that dominates the e-process, and apply Ville's inequality under that PP. A single test supermartingale for the whole composite null need not exist.

Validity holds under any P∈H0P \in H_0, simultaneously over the entire null family, because the e-value/e-process bound is uniform in PP.

FormalSLT status: PROVED CONDITIONAL ON a narrower process interface and a finite horizon or bounded stopping time. FormalSLT's eProcess_typeI_control proves the finite-horizon crossing bound for its EProcess, defined as a nonnegative normalized supermartingale. The same module's eProcess_optionalContinuation proves bounded optional continuation. The all-time limit passage used here is not packaged as a Lean theorem in that module. It remains OPEN in FormalSLT, along with the more general stopped-e-value interface that allows e-processes that are not supermartingales, RIPr, and the safe tt-test construction.

Why It Matters

Safe tests make the evidence object explicit. For a simple null and simple alternative, the likelihood ratio is an e-value. Thresholding it at 1/α1/\alpha is safe, while the exact Neyman-Pearson critical value is chosen from its null distribution and is generally different. For composite hypotheses, a valid e-value may require a mixture, universal-inference denominator, RIPr, or another uniform construction.

Failure Mode

The e-value must satisfy EP[S]≤1\mathbb{E}_P[S] \leq 1 for every P∈H0P \in H_0, not just for a chosen point null. Composite nulls require constructions that handle the entire null family at once: universal inference, RIPr, or numerical worst-case search.

Construction Methods

A useful grouping by hypothesis structure is:

Simple-versus-simple (likelihood ratio). When both H0={P0}H_0 = \{P_0\} and H1={P1}H_1 = \{P_1\} are simple, p1/p0p_1/p_0 is an e-value. For a data stream, the running likelihood ratio is a test martingale under P0P_0 when the conditional likelihood factors are correctly specified. The safe threshold 1/α1/\alpha and the exact Neyman-Pearson threshold solve different calibration problems; growth-rate optimality is not the same property as uniformly most-powerful fixed-sample testing.

Simple null, composite alternative. Choose a fixed alternative distribution QQ, possibly a Bayes mixture over the alternative. The likelihood ratio q/p0q/p_0 is an e-value under the simple null. For a composite alternative, the GROW criterion maximizes the worst-case expected log payoff; under the paper's assumptions, solving that criterion can select a least-favorable alternative mixture. It is not the same as choosing an arbitrary point alternative.

Composite null, composite alternative. Two main constructions:

  1. Universal inference (Wasserman-Ramdas-Balakrishnan 2020): use held-out data to evaluate a likelihood fitted without those observations, with a denominator that dominates the null likelihood. Its broad applicability comes with sample-splitting and estimation costs, but there is no universal exact 2\sqrt{2} power loss.
  2. Reverse information projection (Grünwald-de Heide-Koolen 2024): project an alternative distribution onto the convex set of null mixtures, allowing a limiting subdistribution in the general theorem. The resulting ratio is an e-value and is optimal for the paper's expected-log-growth criterion under its assumptions.

Sample-mean tests on bounded outcomes. The betting-strategy construction (Waudby-Smith-Ramdas 2024) builds e-processes for conditional-mean hypotheses such as H0:μ≤μ0H_0: \mu \leq \mu_0 via predictable bets, yielding finite-sample confidence sequences under the paper's boundedness assumptions.

Comparison: Neyman-Pearson vs Safe Tests

PropertyNeyman-Pearson testSafe test
Decision ruleReject if T≥tαT \geq t_\alpha where tαt_\alpha is the 1−α1 - \alpha quantile of TT under H0H_0Reject if S≥1/αS \geq 1/\alpha
Sample sizeFixed or handled by a specified sequential designAny stopping rule only for an e-process or another anytime-valid construction
Type I errorExactly α\alpha (continuous case)At most α\alpha (often strictly less)
Power at fixed nnMost powerful for simple-versus-simple testing at the chosen levelConstruction-dependent; no generic constant-factor guarantee
Composite nullNeeds uniform tail bound (often hard)Handled by universal inference or RIPr
Multiple testingBonferroni / BH adjustments after the facte-BH integrates into the framework natively
Combination across studiesRequires a valid combination rule and its assumptionsPredictable products of conditional e-values are sequentially valid; fixed mixtures preserve e-value validity

The trade-off is design-specific. The universal threshold 1/α1/\alpha can be conservative at a fixed sample size, while optional stopping can reduce expected sample use under alternatives. Worst-case sample requirements, calibration, and power must be reported for the chosen e-value and stopping rule rather than assigned a universal factor.

Canonical Example: Safe t-Testing with Unknown Scale

Consider iid Gaussian observations under H0:μ=0H_0:\mu=0 with unknown scale σ>0\sigma>0. Grünwald, de Heide, and Koolen construct safe tt-test e-values by using the same right-Haar measure on the nuisance scale under the null and alternative. After reducing to a scale-invariant statistic, the resulting Bayes-factor form is an e-value for every σ\sigma in the null. For priors on the standardized effect size with a finite 2+ε2+\varepsilon moment, the paper also proves an expected-log-growth optimality statement for a specified class of alternatives. A Cauchy-prior version remains an e-value, but that growth-optimality result is not established for the momentless Cauchy prior.

Sequential use requires the paper's sequentially decomposable construction and its stated filtration. It is not valid to take an arbitrary fixed-nn Bayes factor, inspect it at every nn, and call the resulting sequence an e-process. Exact power and sample-use comparisons depend on the effect-size prior, monitoring rule, and maximum sample size; they must be computed for that design.

Worked Exercise

ExerciseAdvanced

Problem

For H0:θ=θ0H_0: \theta = \theta_0 versus H1:θ=θ1H_1: \theta = \theta_1 on iid data, let Λn=∏i=1npθ1(Xi)/pθ0(Xi)\Lambda_n=\prod_{i=1}^n p_{\theta_1}(X_i)/p_{\theta_0}(X_i). Show that Λn\Lambda_n is an e-value and that thresholding at 1/α1/\alpha gives a safe single-analysis test. Then compare that threshold with the exact Neyman-Pearson threshold for N(0,1)\mathcal{N}(0, 1) versus N(1,1)\mathcal{N}(1, 1) at n=16n = 16 and α=0.05\alpha=0.05.

Implementation Note

The safestats R package implements safe tests and confidence sequences for several parametric settings, including tt-tests, two-proportion tests, zz-tests, and logrank tests. Package support and the assumptions of each routine should be checked separately; the general safe-testing theory does not turn an arbitrary classical test into an e-process.

For ML-style applications with bounded outcomes, the confseq Python package implements confidence sequences, time-uniform boundaries, and betting methods. Its repository describes the software as early-stage, so production use should pin a version and reproduce the chosen routine's assumptions and tests.

A common implementation pitfall is using the current observation to choose the bet placed on that same observation. Adaptation is allowed, but it must be predictable: the round-tt e-variable and any mixture weights applied to it may depend on information through time t−1t-1, while still satisfying the conditional expectation bound under every null distribution. Pre-registration is one simple way to meet that rule, but it is not required by the mathematics.

Practical Example: Multi-Site Clinical Trial Meta-Analysis

A pharmaceutical sponsor coordinates a multi-center trial of a new treatment against placebo. Centers complete enrollment at different times. The classical approach pre-specifies a meta-analysis at the planned end date and combines pp-values via Fisher's combination. Peeking at interim results is forbidden.

A safe-testing approach has two distinct valid routes:

  1. If each center supplies a full e-process for the common null relative to the joint filtration, choose fixed nonnegative weights summing to one and monitor their weighted mixture.
  2. If centers instead finish in a sequence, require each new site-level factor to be a conditional e-value given all earlier site results, and multiply those factors.

For either route, stop when the resulting e-process crosses 1/α1/\alpha.

Averaging a fixed collection of e-values gives another e-value, even when the inputs are dependent. It does not follow that repeatedly changing the collection or its data-dependent weights while monitoring creates an e-process. Unequal enrollment and interleaving can be handled, but only after the site-level filtrations and conditional validity statements are specified.

References

Canonical:

  • Grünwald, P., de Heide, R., and Koolen, W. (2024). "Safe testing." Journal of the Royal Statistical Society, Series B 86(5), pp. 1091-1128. Defines safe testing, GRO criteria, and the RIPr construction.
  • Ramdas, A., Grünwald, P., Vovk, V., and Shafer, G. (2023). "Game-theoretic statistics and safe anytime-valid inference." Statistical Science 38(4), pp. 576-601. Surveys e-values, e-processes, predictable betting, and sequential combination.
  • Ramdas, A., Ruf, J., Larsson, M., and Koolen, W. (2022). "Testing exchangeability: fork-convexity, supermartingales, and e-processes." International Journal of Approximate Reasoning 141, pp. 83-109. Gives a useful e-process that is not itself a nonnegative supermartingale.
  • Vovk, V. and Wang, R. (2021). "E-values: Calibration, combination and applications." Annals of Statistics 49(3), pp. 1736-1754. The e-value foundations on which safe testing is built.

Current:

  • Wasserman, L., Ramdas, A., and Balakrishnan, S. (2020). "Universal inference." Proceedings of the National Academy of Sciences 117(29), pp. 16880-16890. Split-sample safe tests for composite nulls.
  • Henzi, A. and Ziegel, J. F. (2022). "Valid sequential inference on probability forecast performance." Biometrika 109(3), pp. 647-663. Safe testing for forecast evaluation, central to modern ML model comparison.
  • Waudby-Smith, I. and Ramdas, A. (2024). "Estimating means of bounded random variables by betting." Journal of the Royal Statistical Society, Series B 86(1), pp. 1-27. Bounded-outcome safe tests and finite-sample confidence sequences.

Critique and context:

  • Pawel, S. and Held, L. (2022). "The sceptical Bayes factor for the assessment of replication success." JRSSB (Statistical Methodology) 84(3), pp. 879-911, DOI 10.1111/rssb.12491. Compares safe testing to Bayes factor for replication.
  • Berger, J. O. and Sellke, T. (1987). "Testing a point null hypothesis: the irreconcilability of p values and evidence." Journal of the American Statistical Association 82(397), pp. 112-122. The historical critique of pp-values that safe testing addresses.

Next Topics

Last reviewed: August 8, 2026

Cite this page

Sneiderman, Robby. "Safe Testing." TheoremPath, reviewed 2026-08-08. https://theorempath.com/topics/safe-testing

Canonical URL
https://theorempath.com/topics/safe-testing
Author
Robby Sneiderman, TheoremPath
Last reviewed
2026-08-08
What this is
A reference page on TheoremPath. Written and maintained by the named author. Not peer reviewed and not refereed by any venue. Each claim below carries its own verification status.
Terms
All rights reserved. Non-commercial quotation with attribution permitted.

Citing one claim

One statement on this page is separately addressable. Each carries an identifier of the form claim:<topic>::<statement> printed beside it, and resolves to that identifier's anchor on this URL. Quote the statement, its assumptions, and its failure modes together.

Recorded verification on this page: 1 informal_only.

  • This page is a secondary source. Where a claim cites literature, the result is due to that source and not to this site.
  • 1 of 1 claims on this page name no source. Their attribution to the literature is not recorded here, whatever their verification status says.
  • 1 of 1 carry neither a source nor a proof check. Their provenance is the author's reading of the literature; check them against a primary source before relying on them.
  • Statements are given in LaTeX exactly as authored. The rendered page passes the same LaTeX through KaTeX, so text scraped from the HTML repeats expressions and drops radicals; quote from statementTex.

/claims.json lists 1,178 of the 1,231 statements this site records, with the anchor and the recorded verification of each. What earns a place in it: a theorem, lemma, corollary, or proposition block on a published topic page. Definitions are not listed, and neither are the 27 statements on comparison pages. A missing entry means the rule above did not select the statement, not that this site is silent on the result.

Canonical graph

Required before and derived from this topic

These links come from prerequisite edges in the curriculum graph. Editorial suggestions are shown here only when the target page also cites this page as a prerequisite.

Required prerequisites

6

Derived topics

0

No published topic currently declares this as a prerequisite.