Sequential Inference
Safe Testing
Safe testing uses e-values for fixed analyses and e-processes for anytime validity. Reverse information projection and growth criteria guide composite-hypothesis constructions; FormalSLT covers the finite-horizon supermartingale case.
Prerequisites
Learning position
Place this page in a reading path.
sequential-inference | layer 3 | tier 1. This page has 6 direct prerequisites and 0 published dependents.
What next
Anytime-Valid InferenceThis is the first curated or graph-derived continuation from the current page.
Evidence badge
Source-grounded pageThis page has no public Lean mapping yet. Use the evidence page to inspect how claim status labels work.
Why This Matters
Classical Neyman-Pearson testing fixes a sampling plan, a null, and an alternative. In the simple-versus-simple case it constructs a most-powerful level- test for that fixed experiment. Reusing a fixed-sample calibration after unplanned monitoring or data-dependent continuation generally invalidates its Type I error guarantee. Sequential boundaries and alpha-spending repair specific monitoring plans, but they must be built into the analysis.
Safe testing, formalized by Grünwald, de Heide, and Koolen (2024, Journal of the Royal Statistical Society Series B, 86(5)), starts with a nonnegative statistic whose expectation is at most one under every distribution in the null. Such a statistic is an e-value and gives a valid threshold test at one analysis time. Sequential validity needs an additional construction: an e-process, or a running product of conditional e-values whose next factor is chosen predictably from past information.
The common currency matters because e-values can be multiplied across conditionally valid studies, mixed using weights fixed before the data used by the mixture, and monitored when assembled into an e-process. Likelihood ratios and some Bayes factors supply important examples. This does not make every classical test a safe test at the same rejection threshold, and it does not imply a universal fixed-sample power penalty. Power depends on the model, alternative, e-value construction, and stopping design.
Formal Setup
Safe test
A safe test at level for a null hypothesis is a function of the data such that implies rejection, and is an e-value: for every . The Type I error of the rejection rule is at most by Markov's inequality.
Safe anytime-valid test
A safe anytime-valid test at level is a sequential procedure with rejection rule "stop and reject the first time " where is a nonnegative process adapted to the stated filtration and satisfies
with the second supremum over stopping times in the stated filtration. An e-process need not itself be a nonnegative supermartingale. Every test supermartingale is an e-process, and useful e-processes exist that are not test supermartingales.
Reverse information projection (RIPr)
Let be an alternative distribution and let denote the convex set of Bayes mixtures of distributions in the composite null. Under the support and finite-KL conditions in Grünwald, de Heide, and Koolen's Theorem 1, the RIPr minimizes over , or is the limiting sub-probability distribution when the infimum is not attained by a proper mixture. Writing and for the corresponding densities, the ratio is an e-value for the full null and maximizes expected log growth under . In regular examples is a proper null mixture; this is not part of the definition in complete generality.
The Core Result
The construction that makes safe testing work is the e-value itself.
E-Value Construction Yields a Safe Test
Statement
Fix . Let be a nonnegative function of the data with . The test that rejects when is a safe test at level : its Type I error is at most , uniformly over .
If additionally is an e-process for , the sequential rule "stop and reject at the first with " is a safe anytime-valid test at level : Type I error at most for every stopping rule.
Intuition
Markov's inequality gives the single-shot guarantee. The stopped-value definition gives the sequential guarantee directly; for a nonnegative test supermartingale, Ville's inequality gives the familiar running-maximum form. The construction does not optimize for power; it guarantees validity. Power comes from the choice of e-value or e-process.
Proof Sketch
Single-shot: by Markov.
Sequential: apply the stopped-value property to the bounded hitting times , use Markov's inequality, and let . This gives . Equivalently, for each one may use a nonnegative -supermartingale with initial value at most one that dominates the e-process, and apply Ville's inequality under that . A single test supermartingale for the whole composite null need not exist.
Validity holds under any , simultaneously over the entire null family, because the e-value/e-process bound is uniform in .
FormalSLT status: PROVED CONDITIONAL ON a narrower process interface and a
finite horizon or bounded stopping time.
FormalSLT's
eProcess_typeI_control
proves the finite-horizon crossing bound for its EProcess, defined as a
nonnegative normalized supermartingale. The same module's
eProcess_optionalContinuation
proves bounded optional continuation. The all-time limit passage used here is
not packaged as a Lean theorem in that module. It remains OPEN in
FormalSLT, along with the more general stopped-e-value interface that allows
e-processes that are not supermartingales, RIPr, and the safe -test
construction.
Why It Matters
Safe tests make the evidence object explicit. For a simple null and simple alternative, the likelihood ratio is an e-value. Thresholding it at is safe, while the exact Neyman-Pearson critical value is chosen from its null distribution and is generally different. For composite hypotheses, a valid e-value may require a mixture, universal-inference denominator, RIPr, or another uniform construction.
Failure Mode
The e-value must satisfy for every , not just for a chosen point null. Composite nulls require constructions that handle the entire null family at once: universal inference, RIPr, or numerical worst-case search.
Construction Methods
A useful grouping by hypothesis structure is:
Simple-versus-simple (likelihood ratio). When both and are simple, is an e-value. For a data stream, the running likelihood ratio is a test martingale under when the conditional likelihood factors are correctly specified. The safe threshold and the exact Neyman-Pearson threshold solve different calibration problems; growth-rate optimality is not the same property as uniformly most-powerful fixed-sample testing.
Simple null, composite alternative. Choose a fixed alternative distribution , possibly a Bayes mixture over the alternative. The likelihood ratio is an e-value under the simple null. For a composite alternative, the GROW criterion maximizes the worst-case expected log payoff; under the paper's assumptions, solving that criterion can select a least-favorable alternative mixture. It is not the same as choosing an arbitrary point alternative.
Composite null, composite alternative. Two main constructions:
- Universal inference (Wasserman-Ramdas-Balakrishnan 2020): use held-out data to evaluate a likelihood fitted without those observations, with a denominator that dominates the null likelihood. Its broad applicability comes with sample-splitting and estimation costs, but there is no universal exact power loss.
- Reverse information projection (Grünwald-de Heide-Koolen 2024): project an alternative distribution onto the convex set of null mixtures, allowing a limiting subdistribution in the general theorem. The resulting ratio is an e-value and is optimal for the paper's expected-log-growth criterion under its assumptions.
Sample-mean tests on bounded outcomes. The betting-strategy construction (Waudby-Smith-Ramdas 2024) builds e-processes for conditional-mean hypotheses such as via predictable bets, yielding finite-sample confidence sequences under the paper's boundedness assumptions.
Comparison: Neyman-Pearson vs Safe Tests
| Property | Neyman-Pearson test | Safe test |
|---|---|---|
| Decision rule | Reject if where is the quantile of under | Reject if |
| Sample size | Fixed or handled by a specified sequential design | Any stopping rule only for an e-process or another anytime-valid construction |
| Type I error | Exactly (continuous case) | At most (often strictly less) |
| Power at fixed | Most powerful for simple-versus-simple testing at the chosen level | Construction-dependent; no generic constant-factor guarantee |
| Composite null | Needs uniform tail bound (often hard) | Handled by universal inference or RIPr |
| Multiple testing | Bonferroni / BH adjustments after the fact | e-BH integrates into the framework natively |
| Combination across studies | Requires a valid combination rule and its assumptions | Predictable products of conditional e-values are sequentially valid; fixed mixtures preserve e-value validity |
The trade-off is design-specific. The universal threshold can be conservative at a fixed sample size, while optional stopping can reduce expected sample use under alternatives. Worst-case sample requirements, calibration, and power must be reported for the chosen e-value and stopping rule rather than assigned a universal factor.
Canonical Example: Safe t-Testing with Unknown Scale
Consider iid Gaussian observations under with unknown scale . Grünwald, de Heide, and Koolen construct safe -test e-values by using the same right-Haar measure on the nuisance scale under the null and alternative. After reducing to a scale-invariant statistic, the resulting Bayes-factor form is an e-value for every in the null. For priors on the standardized effect size with a finite moment, the paper also proves an expected-log-growth optimality statement for a specified class of alternatives. A Cauchy-prior version remains an e-value, but that growth-optimality result is not established for the momentless Cauchy prior.
Sequential use requires the paper's sequentially decomposable construction and its stated filtration. It is not valid to take an arbitrary fixed- Bayes factor, inspect it at every , and call the resulting sequence an e-process. Exact power and sample-use comparisons depend on the effect-size prior, monitoring rule, and maximum sample size; they must be computed for that design.
Worked Exercise
Problem
For versus on iid data, let . Show that is an e-value and that thresholding at gives a safe single-analysis test. Then compare that threshold with the exact Neyman-Pearson threshold for versus at and .
Implementation Note
The safestats R package
implements safe tests and confidence sequences for several parametric
settings, including -tests, two-proportion tests, -tests, and logrank
tests. Package support and the assumptions of each routine should be checked
separately; the general safe-testing theory does not turn an arbitrary
classical test into an e-process.
For ML-style applications with bounded outcomes, the confseq Python package implements confidence sequences, time-uniform boundaries, and betting methods. Its repository describes the software as early-stage, so production use should pin a version and reproduce the chosen routine's assumptions and tests.
A common implementation pitfall is using the current observation to choose the bet placed on that same observation. Adaptation is allowed, but it must be predictable: the round- e-variable and any mixture weights applied to it may depend on information through time , while still satisfying the conditional expectation bound under every null distribution. Pre-registration is one simple way to meet that rule, but it is not required by the mathematics.
Practical Example: Multi-Site Clinical Trial Meta-Analysis
A pharmaceutical sponsor coordinates a multi-center trial of a new treatment against placebo. Centers complete enrollment at different times. The classical approach pre-specifies a meta-analysis at the planned end date and combines -values via Fisher's combination. Peeking at interim results is forbidden.
A safe-testing approach has two distinct valid routes:
- If each center supplies a full e-process for the common null relative to the joint filtration, choose fixed nonnegative weights summing to one and monitor their weighted mixture.
- If centers instead finish in a sequence, require each new site-level factor to be a conditional e-value given all earlier site results, and multiply those factors.
For either route, stop when the resulting e-process crosses .
Averaging a fixed collection of e-values gives another e-value, even when the inputs are dependent. It does not follow that repeatedly changing the collection or its data-dependent weights while monitoring creates an e-process. Unequal enrollment and interleaving can be handled, but only after the site-level filtrations and conditional validity statements are specified.
References
Canonical:
- Grünwald, P., de Heide, R., and Koolen, W. (2024). "Safe testing." Journal of the Royal Statistical Society, Series B 86(5), pp. 1091-1128. Defines safe testing, GRO criteria, and the RIPr construction.
- Ramdas, A., Grünwald, P., Vovk, V., and Shafer, G. (2023). "Game-theoretic statistics and safe anytime-valid inference." Statistical Science 38(4), pp. 576-601. Surveys e-values, e-processes, predictable betting, and sequential combination.
- Ramdas, A., Ruf, J., Larsson, M., and Koolen, W. (2022). "Testing exchangeability: fork-convexity, supermartingales, and e-processes." International Journal of Approximate Reasoning 141, pp. 83-109. Gives a useful e-process that is not itself a nonnegative supermartingale.
- Vovk, V. and Wang, R. (2021). "E-values: Calibration, combination and applications." Annals of Statistics 49(3), pp. 1736-1754. The e-value foundations on which safe testing is built.
Current:
- Wasserman, L., Ramdas, A., and Balakrishnan, S. (2020). "Universal inference." Proceedings of the National Academy of Sciences 117(29), pp. 16880-16890. Split-sample safe tests for composite nulls.
- Henzi, A. and Ziegel, J. F. (2022). "Valid sequential inference on probability forecast performance." Biometrika 109(3), pp. 647-663. Safe testing for forecast evaluation, central to modern ML model comparison.
- Waudby-Smith, I. and Ramdas, A. (2024). "Estimating means of bounded random variables by betting." Journal of the Royal Statistical Society, Series B 86(1), pp. 1-27. Bounded-outcome safe tests and finite-sample confidence sequences.
Critique and context:
- Pawel, S. and Held, L. (2022). "The sceptical Bayes factor for the assessment of replication success." JRSSB (Statistical Methodology) 84(3), pp. 879-911, DOI 10.1111/rssb.12491. Compares safe testing to Bayes factor for replication.
- Berger, J. O. and Sellke, T. (1987). "Testing a point null hypothesis: the irreconcilability of p values and evidence." Journal of the American Statistical Association 82(397), pp. 112-122. The historical critique of -values that safe testing addresses.
Next Topics
- Anytime-valid inference: inference rules for data-dependent monitoring and stopping.
- Confidence sequences: time-uniform interval estimates that complement safe tests.
- E-values and anytime-valid inference: the umbrella reference with applications and the e-BH multiple-testing procedure.
Last reviewed: August 8, 2026
Cite this page
Sneiderman, Robby. "Safe Testing." TheoremPath, reviewed 2026-08-08. https://theorempath.com/topics/safe-testing
- Canonical URL
- https://theorempath.com/topics/safe-testing
- Author
- Robby Sneiderman, TheoremPath
- Last reviewed
- 2026-08-08
- What this is
- A reference page on TheoremPath. Written and maintained by the named author. Not peer reviewed and not refereed by any venue. Each claim below carries its own verification status.
- Terms
- All rights reserved. Non-commercial quotation with attribution permitted.
Citing one claim
One statement on this page is separately addressable. Each carries an identifier of the form claim:<topic>::<statement> printed beside it, and resolves to that identifier's anchor on this URL. Quote the statement, its assumptions, and its failure modes together.
Recorded verification on this page: 1 informal_only.
- This page is a secondary source. Where a claim cites literature, the result is due to that source and not to this site.
- 1 of 1 claims on this page name no source. Their attribution to the literature is not recorded here, whatever their verification status says.
- 1 of 1 carry neither a source nor a proof check. Their provenance is the author's reading of the literature; check them against a primary source before relying on them.
- Statements are given in LaTeX exactly as authored. The rendered page passes the same LaTeX through KaTeX, so text scraped from the HTML repeats expressions and drops radicals; quote from statementTex.
/claims.json lists 1,178 of the 1,231 statements this site records, with the anchor and the recorded verification of each. What earns a place in it: a theorem, lemma, corollary, or proposition block on a published topic page. Definitions are not listed, and neither are the 27 statements on comparison pages. A missing entry means the rule above did not select the statement, not that this site is silent on the result.
Canonical graph
Required before and derived from this topic
These links come from prerequisite edges in the curriculum graph. Editorial suggestions are shown here only when the target page also cites this page as a prerequisite.
Required prerequisites
6- KL Divergencelayer 1 · tier 1
- e-valueslayer 2 · tier 1
- Likelihood-Ratio, Wald, and Score Testslayer 2 · tier 1
- e-processeslayer 3 · tier 1
- Hypothesis Testing for MLlayer 2 · tier 2
Derived topics
0No published topic currently declares this as a prerequisite.