TL;DR for operators
A team choosing between two tiny models for an edge product could reasonably treat the same nominal parameter budget as a fair baseline. In this controlled study at roughly 60K parameters, however, five approximately parameter-matched full-precision transformer shapes span validation losses from 2.835 to 3.477 nats per byte at the 16M-byte training budget—a 22.6% difference produced by the tested shape package alone. Against the weakest transformer shape, the model that dynamically distributes computation across several internal pathways looks roughly 19–21% better; against the strongest full-precision transformer, it is only 0.95% better, inside the study’s seed-noise threshold.
The larger-budget result still matters. At 130M bytes, that routed model beats all three transformers tested there, but a simpler recurrent model that carries forward a learned state performs better in full precision and statistically matches it when many weights are restricted to three discrete levels. The evidence therefore supports a broader recurrent or state-space advantage more clearly than it supports routing itself as the decisive mechanism.
Training recipe can reverse conclusions as well. For the routed model, a 90/10 full-precision-to-ternary schedule is 15.3% worse than ternary-from-scratch training when the second stage uses a learning rate of $10^{-4}$, yet becomes 4.7% better at $3\times10^{-3}$.
For an edge-AI team, baseline shape, training budget, seed variation, quantization coverage, and transition-stage optimization should be treated as decision variables rather than background settings. The study does not establish latency, energy, SRAM, or deployment superiority on actual microcontrollers.
The baseline can move more than the architecture
Suppose a product team has two models near the same parameter budget. One benchmark shows a clear winner. If memory capacity is approximately matched, it is tempting to interpret the remaining loss gap as evidence about architecture.
The controlled re-examination by Gautam Veldanda shows why that inference can fail at very small scales.1 At a 16M-byte training budget, five approximately parameter-matched full-precision transformers produce validation losses ranging from 2.835 to 3.477 nats per byte. The worst is 22.6% worse than the best despite belonging to the same architecture family and staying within roughly the same parameter envelope.
That spread is large enough to change the story told about the routed model. Against the weakest transformer configuration, the routed model appears roughly 19–21% better. Against the strongest full-precision transformer, it is only 0.95% better. Under the paper’s predefined comparison rule, that smaller difference sits inside estimated seed noise.
Parameter count is therefore not functioning as a sufficient control variable here. Depth, width, feed-forward allocation, and positional-embedding share move together as transformer shape changes. The comparison is between architecture packages, not interchangeable implementations of a fixed baseline.
There is another complication: the preferred shape itself moves with training budget. The shallow-wide 1×52 transformer is strongest at 16M bytes, but becomes the weakest of the three transformer shapes evaluated at 130M bytes. The 3×32 configuration becomes strongest at the larger budget.
For evaluation teams, this makes baseline selection conditional on the optimization regime. Choosing the “best transformer baseline” once and carrying it across budgets can still produce a distorted comparison.
The larger-budget gain is real, but routing is not required
Controlling the short-budget baseline does not erase the routed model’s entire advantage.
At 130M training bytes, the routed model reaches 1.668 nats per byte in full precision, compared with 2.143–2.195 for the three transformers evaluated at that budget. Its advantage is 22.2–24.0%, comfortably beyond the study’s estimated noise bands. In ternary precision—where many weights are restricted to three discrete levels—the routed model also leads those transformers by 13.7–14.4%.
The more revealing comparison is with a simpler gated diagonal state-space model. A state-space model carries forward a learned recurrent state rather than relying entirely on attention over prior positions. It has no dynamic router distributing computation across convolution, state-space, and sparse-attention pathways.
In full precision, that simpler model reaches 1.516 nats per byte, 9.1% better than the routed model. In ternary precision, its 1.942 loss is statistically indistinguishable from the routed model’s 1.993 under the paper’s decision rule.
This comparison functions as a mechanism test: it provides counterevidence to routing being necessary for the observed gain. Router diagnostics point in the same direction—the routed model assigns most of its average pathway weight to its state-space branch—but those diagnostics do not establish causality. The gated model also differs in feed-forward and channel-mixing structure, so the experiment does not isolate recurrence itself as the causal ingredient.
For model-selection work, the narrower conclusion is more useful: a complex routing mechanism should be compared against a simpler recurrent or state-space alternative before its extra machinery is credited with the performance difference.
Quantization results are suggestive, not architecture constants
The study also reports sharply different penalties from ternary weights at 130M bytes: +5.3% for the best reported transformer shape, +19.5% for the routed model, and +28.1% for the gated state-space model.
Those numbers should not be interpreted as intrinsic quantization tolerance for each architecture.
The transformer configurations leave roughly 11–22% of their parameters in full precision because positional embeddings remain unquantized. The recurrent models leave only about 1.6–2.3% in full precision. Their nominally “ternary” systems therefore contain materially different amounts of full-precision state.
The from-scratch ternary learning rate is also inherited from the full-precision recipe rather than tuned separately. Different architecture families could consequently be paying different optimization penalties in addition to different quantization penalties.
For hardware planning, reporting bit width alone is insufficient. Teams also need the fraction of parameters actually quantized and the optimization recipe used to obtain the reported quality.
The learning rate reverses another headline result
A separate experiment tests a 90/10 schedule: train the model for 90% of its budget in full precision, switch to ternary weights, then use the final 10% for adaptation.
At a second-stage learning rate of $10^{-4}$, this looks like a failed strategy. For the routed model, it finishes 15.3% worse than ternary training from scratch.
Increase that rate to $3\times10^{-3}$, however, and the same schedule becomes 4.7% better than the from-scratch baseline. The transformer comparison moves in the same direction.
This learning-rate sweep is a sensitivity test, not evidence that $3\times10^{-3}$ is optimal. Performance is still improving at the highest tested value, so the optimum and instability boundary are not bracketed. The from-scratch ternary baseline is also untuned.
Still, the reversal changes how the training transition should be conceptualized. Moving from full precision to ternary weights may require substantial weight reconstruction rather than the small updates normally associated with fine-tuning. Applying a conventional low continuation rate can make a viable strategy appear defective.
What an architecture review should control before committing
For teams selecting compact models for edge or embedded products, the paper supports a stricter experimental workflow:
| Decision variable | What this study shows | Operational treatment |
|---|---|---|
| Baseline shape | Same-size transformers differ by 22.6% at 16M bytes | Sweep multiple plausible shapes |
| Training budget | Transformer shape ordering changes by 130M bytes | Re-evaluate baselines across budgets |
| Random seed | Some small gaps fall inside seed variation | Use multiple runs and explicit noise thresholds |
| Architectural mechanism | Simpler gated SSM matches or beats routed model | Test simpler structural alternatives |
| Quantization | Nominal ternary models retain different FP shares | Report actual quantized fraction |
| Post-switch optimization | Learning rate reverses FP-to-ternary verdict | Tune the precision-transition phase separately |
The business inference is not that one architecture family has won. It is that architecture selection at this scale is unusually exposed to experimental configuration. A product team that locks in a more complex model after one parameter-matched benchmark may be optimizing around baseline choice rather than a persistent model advantage.
The evidence stops well before deployment
The controlled evidence is narrow by design. All experiments are around 60K parameters, on byte-level TinyStories V2, with only two principal training budgets. TinyStories contains highly local and repetitive structure that may favor convolutional or recurrent processing, while byte-level modeling creates longer effective sequences than subword tokenization.
The study also evaluates validation loss rather than downstream language quality. It does not measure latency, energy use, SRAM consumption, or execution behavior on an actual microcontroller. The gated state-space model has about 5.6% more parameters than the routed model, and the larger-budget transformer sweep evaluates only three of the five short-budget shapes.
So this work changes experimental governance more convincingly than it settles architecture choice. It shows that at tiny scales, baseline shape and optimizer recipe are large enough variables to reverse conclusions that might otherwise be attributed to architectural innovation.
For an engineering team, that makes the benchmark itself part of the system under evaluation.
Cognaptus: Automate the Present, Incubate the Future.
-
Gautam Veldanda (2026). Baseline Shape Decides the Verdict: A Controlled Re-Examination of Ternary Language Models at 60K Parameters. arXiv:2609.29397. https://arxiv.org/abs/2609.29397 ↩︎