TL;DR for operators

A fixed seed is not a complete reproducibility guarantee. More surprisingly, neither is restoring the complete state of the random-number generator.

In Reproducible AI Requires Reproducible Randomness, Anthony Bertrand, Tom Schmitt, Engelbert Mephu Nguifo, and David Hill compare library implementations of Mersenne Twister and Philox against canonical reference generators.1 Their deterministic tests show that the same nominal generator, initialized from the corresponding full state, can still emit a different sequence because libraries change what happens between stored state and returned output.

For operators, this moves randomness from a configuration detail into the software dependency graph. Cross-framework validation may need to record the generator implementation, library version, state representation, initialization path, and output procedure—not merely seed=42.

The distinction matters because the failures are not all alike. Several mismatches can be repaired without modifying library source code. Others cannot. Under the tested interface, PyTorch’s Philox path remains unable to reproduce the complete canonical stream.

A restored state can still emit a different stream

Consider a familiar migration problem. A training workflow is moved from one framework or environment to another. The model code and data are held constant. The random seed is restored. The run diverges.

The conventional next step is to save more randomness information. A seed is only an input used to construct a generator’s working configuration; the internal state is the fuller hidden configuration that determines where the generator currently sits in its sequence. If different libraries map one seed into different states, transferring the complete state appears to remove that ambiguity.

The paper shows that this stronger requirement is still insufficient.

Using canonical Mersenne Twister and Philox implementations as references, the authors compare exact output streams from Python Random, NumPy, PyTorch, and TensorFlow. These are deterministic conformance tests rather than statistical estimates: the relevant outcome is whether the emitted sequence matches the reference sequence exactly.

Table 7 makes the initial problem visible. With Mersenne Twister, legacy NumPy naturally matches the reference for both seed- and state-based tests, but several other paths do not. With Philox, none of the tested library implementations naturally reproduces the complete canonical reference stream in the reported state-based comparisons.

That result changes the reproducibility question. Matching the stored state is necessary in some cases, but the receiving library can still interpret or transform that state differently before returning a number.

The mismatch happens below the algorithm name

The paper traces divergence to several layers that are normally hidden behind an API call.

Two libraries can both say they use Mersenne Twister while differing in how a seed becomes state, how integers are assembled, or how raw generator bits are normalized into floating-point values. Philox introduces another issue because it is counter-based: outputs are calculated from counters rather than produced by advancing one conventional sequential state. Counter advancement, partitioning into parallel subsequences, and selection from each generated block therefore become part of observable behavior.

PyTorch’s Mersenne Twister results illustrate the distinction. Its integer stream can be reconstructed by drawing 64-bit integers and decomposing them into two 32-bit unsigned values. Its standard floating-point path, however, uses a different normalization procedure with a 24-bit-style construction. Merely copying the canonical state therefore does not restore canonical floating-point behavior.

This is what the paper calls implementation fidelity: using the same algorithm name is not enough if the surrounding implementation changes the algorithm’s observable stream.

That distinction is operationally useful because a PRNG can still be statistically reputable while being non-portable. Statistical randomness quality and exact stream reproducibility answer different questions.

Some failures are repairable; some are structural

The corrective experiments in Section 8 are more useful than a simple compatibility matrix because they identify why a mismatch occurs.

Failure class What it means Evidence from the tests Operational implication
User-induced The generator can behave as intended, but seeding or interface conventions can lead users onto a different stream Some NumPy Generator paths become portable once initialization behavior is handled correctly Check API semantics before treating divergence as a library defect
User-resolvable The library representation differs, but the canonical stream can be reconstructed externally PyTorch Mersenne Twister integers can be recovered by splitting 64-bit values; TensorFlow Philox integers can be reconstructed similarly A compatibility layer can sometimes repair portability without changing library source
Fundamental under the tested interface Publicly accessible state and outputs are insufficient to reconstruct the canonical behavior PyTorch Philox exposes only part of the canonical counter state and only part of each Philox output block through the examined generation path Exact portability may require implementation changes rather than additional configuration

Table 11 shows the practical effect of these corrections. State-based Mersenne Twister portability is restored for Random and NumPy paths represented in the comparison, while PyTorch integer generation can be aligned through bit manipulation. NumPy Philox can be aligned by accounting for its counter behavior, and TensorFlow Philox by reconstructing 32-bit outputs from generated 64-bit integers.

PyTorch Philox remains the critical counterexample. Its interface constrains access to the canonical counter structure, organizes generation around parallel subsequences, and exposes only one of four 32-bit outputs from each Philox operation in the tested path. The authors therefore cannot reconstruct the full reference stream using their user-level methods.

The point is not that every mismatch is a bug. The more useful classification is whether the mismatch comes from usage, representation, or an implementation architecture that prevents equivalent behavior.

Randomness belongs in the reproducibility contract

The paper directly establishes exact stream compatibility and incompatibility for the tested configurations. The infrastructure implications go one step beyond that evidence.

For an ML platform team moving workloads between frameworks, a stored seed should therefore be treated as metadata rather than proof of reproducibility. A stronger experiment record would preserve the PRNG name, implementation and library version, complete state where accessible, initialization route, and any output transformation required for compatibility.

For model validation and audit teams, the same information can narrow incident diagnosis. If two apparently equivalent executions diverge, randomness implementation becomes an explicit hypothesis alongside data ordering, kernels, hardware, preprocessing, and training code.

Library maintainers have an even more direct use: canonical reference vectors can serve as regression tests for implementation behavior. A library can intentionally deviate from a reference algorithm, but that deviation can then be documented and tested rather than remaining an implicit property of an API.

The business value here is cheaper diagnosis. Without implementation-level checks, teams can spend debugging effort higher in the stack while the divergence originates in the generator underneath it.

The benchmark does not establish downstream model effects

The evidence has a precise boundary.

The experiments cover Mersenne Twister and Philox across selected implementations in Random, NumPy, PyTorch, and TensorFlow, using the reported 2026 software versions and hardware environment. Other generators, future releases, and other execution paths may behave differently.

The study also establishes stream divergence, not downstream harm. It does not estimate how often these PRNG differences materially change model accuracy, convergence, scientific conclusions, or production metrics. A different sequence can alter a stochastic execution trace; the paper does not quantify the resulting application-level effect.

Several fixes are also implementation-specific and require bit or tensor manipulation. Their existence demonstrates that some incompatibilities are recoverable, but not that such reconstruction is an attractive long-term production practice.

Reproducibility needs an implementation check, not just a seed

The common reproducibility progression is seed first, full state if necessary. This paper adds another layer: even state equality can fail when two libraries convert that state into outputs differently.

That makes the random-number implementation itself part of the reproducibility boundary.

For teams that migrate experiments, validate regulated workflows, or investigate unexplained training divergence, the practical check is no longer only whether the same randomness settings were recorded. It is whether the receiving environment has been shown to emit the same stream under those settings—and, when it does not, whether the difference is a usage problem, a reconstructable representation mismatch, or a structural incompatibility.

Cognaptus: Automate the Present, Incubate the Future.


  1. Anthony Bertrand and Tom Schmitt and Engelbert Mephu Nguifo and David Hill (2026). Reproducible AI Requires Reproducible Randomness. arXiv:2609.26461. https://arxiv.org/abs/2609.26461 ↩︎