TL;DR for operators
A mathematical verifier is not an invariant scoring utility. Esther Xin’s Where the Verifier Fails1 audits four implementations and configurations using 307,420 verifier verdicts and finds acceptance of mathematically equivalent answers ranging from 53.8% to 95.2%. Two extraction configurations of the same math-verify library disagree on 49.9% of equivalent cases where both return a verdict.
The more operationally relevant finding is that these errors are not one homogeneous reliability problem. Some verifiers reject equivalent answers because of whitespace or punctuation. Another returns no verdict for a meaningful share of inputs while being correct whenever it does return one. A relative numeric tolerance creates a deterministic threshold at which off-by-one answers become acceptable.
For teams running model evaluations or RLVR reward pipelines, the verifier therefore belongs inside the governed system boundary. Version, extraction configuration, supported input contract, coverage, and category-level error behavior should travel with reported scores. Metamorphic tests can also serve as regression tests in evaluation CI.
The paper does not measure how these verifier errors change RLVR training or how much they distort real-world benchmark scores. Its contribution is narrower: it shows how to identify verifier failure structure before those downstream effects are assumed away.
Two pipelines can see the same mathematics and return different answers
Suppose two evaluation pipelines receive answers that are mathematically equivalent. A reasonable expectation is that they should assign the same correctness judgment.
That expectation fails surprisingly often in this audit. Across the four tested verifier implementations and configurations, acceptance of certified-equivalent variants ranges from 53.8% for mv-expr to 95.2% for strip-string, a spread of 41.3 percentage points.
The difference is not confined to unrelated implementations. The two tested configurations of math-verify—one using plain-expression extraction and one using LaTeX extraction—disagree on 49.9% of certified-equivalent pairs for which both produce a verdict.
That makes configuration part of the evaluation pipeline rather than a peripheral implementation detail. A reported model score can only be reproduced cleanly if the grader configuration is reproducible too.
The audit creates answers whose correct verdict is known in advance
Auditing a verifier presents a circularity problem: if the goal is to check the checker, who supplies the ground truth?
The paper addresses this with certified metamorphic testing. It starts from a gold answer and applies transformations whose semantic effect is known by construction. A semantics-preserving transformation might alter formatting while leaving the mathematical answer unchanged; rejection is therefore a certified false negative. A meaning-changing transformation creates an answer that should no longer be accepted; acceptance becomes a certified false positive.
The study applies 43 transformations across 14 strata to 4,990 unique gold answers drawn from GSM8K, MATH, Big-Math, and 2,000 synthetic answers, producing 307,420 verifier verdicts.
A separate contract matrix specifies which answer forms each verifier claims to support. This prevents unsupported representations from automatically being counted as implementation defects. The distinction is important because contract-dependent acceptance ranges from 0% to 100% for some forms such as boxing and scientific notation.
In other words, the audit separates three questions that aggregate accuracy tends to merge: Was this input supported? Did the verifier return a judgment? Was that judgment correct?
Much of the measured failure budget is formatting, not mathematics
One might expect mathematical verification failures to cluster around difficult symbolic equivalence. For the two audited math-verify configurations, the dominant measured problem is considerably more mundane.
Whitespace and punctuation account for 93.0% of in-contract failures for the default LaTeX configuration and 74.1% for expression extraction. The category-level false-negative rates reinforce the point: whitespace transformations produce a 37.2% false-negative rate for mv-latex and 50.0% for mv-expr.
This changes the engineering diagnosis. If failure mass is concentrated in normalization behavior, spending first on increasingly sophisticated symbolic reasoning may leave the main reliability problem untouched.
For evaluation infrastructure owners, the practical value of category-level error budgets is prioritization. They tell the team whether it needs better normalization, extraction, parsing, equivalence logic, or runtime reliability.
A missing verdict is not the same as a wrong verdict
Aggregate acceptance rates can also obscure whether a verifier returned an incorrect judgment or failed to produce one at all.
The paper therefore reports coverage separately:
Parsing exceptions, timeouts, crashes, and similar failures reduce coverage rather than being silently treated as ordinary rejection.
The contrast between two implementations is instructive:
| Verifier | Self-validation | Coverage | Correct among returned judgments |
|---|---|---|---|
| strip-string | 95.2% | 100.0% | 95.2% |
| sympy-cascade | 87.3% | 87.3% | 100.0% |
strip-string always returns a verdict but sometimes rejects an equivalent answer. sympy-cascade returns no verdict on some inputs, yet is correct on every in-contract case where it does return one.
Those failure modes require different remedies. Combining them into one residual error number removes information needed for debugging and governance.
The 16% false-positive rate hides a deterministic threshold
The sharpest example comes from adversarial off-by-one answers.
For the audited sympy-cascade, the aggregate off-by-one false-positive rate is 16.0%, with a 95% confidence interval of 14.1% to 18.1% across 1,287 cases. Read alone, that number might sound like diffuse verifier noise.
The category-level result shows otherwise. The verifier accepts none of the tested off-by-one cases below a gold-answer magnitude of $10^4$, then accepts all tested cases at or above $10^4$.
The mechanism is a relative tolerance of $10^{-4}$. Once the correct answer becomes sufficiently large, an absolute difference of one falls inside that relative tolerance.
This is a qualitatively different risk from occasional random mistakes. A systematic threshold can create a repeatable acceptance region. Reward-system audits therefore need to search for structured exploit surfaces, not merely estimate an average false-positive rate.
Treat verifier configuration as governed evaluation metadata
The paper directly establishes verifier-level behavior under the tested implementations, contracts, transformations, and runtime configuration. The business implications require a further step of interpretation.
Cognaptus inference: teams that depend on automatic verifiers for model releases, benchmark reporting, or reward generation should record verifier identity and version, extraction configuration, supported-input contract, runtime dependencies, coverage, and category-level error statistics alongside model results.
The metamorphic suite also suggests a practical CI pattern. Before changing an extraction rule, parser, normalizer, symbolic backend, or tolerance parameter, rerun certified-equivalent and meaning-changing transformations. A regression can then be localized before it propagates into a benchmark or reward pipeline.
The corpus-only robustness analysis strengthens the verifier-level result rather than introducing a separate claim. After excluding all 2,000 synthetic gold answers, the broad ordering remains: 57.5% self-validation for mv-expr, 84.1% for mv-latex, 87.8% for strip-string, and 88.3% for sympy-cascade.
The audit stops before downstream model effects
Several boundaries matter.
The experiment audits verifiers, not trained models. It runs no RLVR training experiments and therefore does not show whether a measured verifier failure causes learning delay, collapse, improvement, or negligible downstream change.
The transformation frequencies are also designed for auditing, not to reproduce the natural distribution of answer forms in model generations. The reported rates therefore cannot be read directly as expected benchmark-score distortion.
Finally, the study covers four open implementations, and contract assignment contains researcher judgment, even though the matrix is released for inspection.
These limits leave an important next question unresolved: once verifier failures are decomposed by mechanism, which categories materially change model evaluation or learning under realistic output distributions?
Reliability needs a failure map, not just a rate
The central contribution of this audit is methodological as much as numerical. It turns verifier reliability from a single score into a map of supported inputs, returned judgments, false negatives, false positives, and identifiable mechanisms.
That map changes what an evaluation team can do. A formatting regression can be fixed as a formatting regression. An execution failure can be tracked as missing coverage. A tolerance threshold can be tested as a systematic acceptance rule. Configuration changes can be treated as changes to the measurement system.
For RLVR and mathematical benchmarking, the grader is therefore part of the experimental apparatus. Reporting the model without reporting how the answer was judged leaves a material part of that apparatus unspecified.
Cognaptus: Automate the Present, Incubate the Future.
-
Esther Xin (2026). Where the Verifier Fails: A Category-Level Audit of Reward Signals in RLVR. arXiv:2609.01354. https://arxiv.org/abs/2609.01354 ↩︎