Cover image

Check What Breaks: Turning Reasoning Errors into Inference-Time Controls

TL;DR for operators A generic instruction such as “pay more attention” is not a reliable reasoning control. In this study, that prompt barely changes performance for one model, substantially improves another, and makes a third slightly worse. A prompt that instead tells the model which failure points to verify reduces incorrect answers across all three tested model families. ...

October 2, 2026 · 7 min · Zelina
Cover image

Put the LLM Upstream, Not in the Loop

TL;DR for operators The most consequential choice in this paper is where the LLM is allowed to act. It does not decide whether simulated farms adopt solar panels. Instead, it works upstream: drafting behavioural rubrics and techno-economic scenarios that humans inspect, validate against explicit rules, freeze, and then feed into an existing calibrated simulation. ...

September 21, 2026 · 7 min · Zelina
Cover image

One Step Is Not a Workflow: Where LLM Rule Following Starts to Break

TL;DR for operators A model that is highly reliable at applying one explicit rule transition is not necessarily reliable at executing an entire procedure built from those transitions. In Reasoning Capabilities of Large Language Models. Lessons Learned from General Game Playing1, the strongest evaluated model, Gemini 2.5 Pro, achieves 95.6% exact success on one-step next-state generation. At five dependent state transitions, exact success falls to 73.4%. When the model must also choose actions during those five steps, it falls again to 65.3%. ...

September 17, 2026 · 7 min · Zelina
Cover image

Catch the Drift Before the Answer: Reasoning Trajectories as a Runtime Control Surface

TL;DR for operators When a model is midway through a difficult task, the expensive choice is often whether to leave it alone, verify it, or spend more inference on correction. Doing that for every request wastes compute and can damage answers that were already on track. Lihao Sun and colleagues show that the model’s internal state changes systematically as reasoning progresses, and that late-stage changes contain a useful warning signal for eventual failure.1 Their trajectory-based correctness features reach a seed-averaged best-layer ROC-AUC of about 0.85, versus 0.765 for the strongest reported LogitLens baseline and 0.649 for reasoning step count alone. ...

September 16, 2026 · 8 min · Zelina
Cover image

An 8/10 Is Not a Probability: Validating LLM Confidence Before It Controls Workflow

TL;DR for operators A model saying “8/10 confident” does not mean its underlying uncertainty is approximately 20%. Across the evaluated settings, the average instance-level correlation between reported confidence and logits-based confidence is only 0.135. The more useful operating rule is narrower. First test whether reported scores vary enough to distinguish cases. Then measure whether those scores rank examples meaningfully on held-out data. Separately test whether their numerical scale agrees with the comparison signal and whether they are calibrated against correctness. Do not substitute one test for another. ...

September 10, 2026 · 7 min · Zelina
Cover image

Spend the Next Byte Where It Repairs the Model

TL;DR for operators A deployment team can reach an awkward point after quantization: the checkpoint is small enough to ship, but quality is below target, and replacing the compression pipeline would mean another round of calibration, validation, packaging, and operational risk. The paper on Activation-Weighted Seeded Residual Coding, or AWSRC, asks whether some of that lost quality can instead be bought back with a small additional payload while leaving the existing low-bit reconstruction untouched.1 In its cleanest comparison, all tested residual codecs receive exactly 49,245,876 extra serialized bytes on the same RTN-SDQ parent. AWSRC reaches perplexity 7.04 and mean zero-shot accuracy 0.69—the best perplexity and a top accuracy result among the tested sparse, low-rank, learned-vector-quantized, and AWSRC repairs. Learned vector quantization has the best unrounded KL result, so the evidence supports a quality-per-byte advantage rather than dominance on every metric. ...

September 9, 2026 · 7 min · Zelina
Cover image

When Hallucination Is More Than a Wrong Fact: Measuring Reliability Through the User

TL;DR for operators A model change can improve an automated hallucination benchmark while leaving users dissatisfied for a different reason: sources are hard to verify, reasoning appears unsupported, false claims are stated with confidence, or corrections are ignored. The System Hallucination Scale (SHS) gives teams a structured way to measure those experiences across five dimensions rather than reducing reliability to a binary factual-error judgment.1 ...

September 8, 2026 · 7 min · Zelina
Cover image

The Optimizer Comes Before the Quantizer

TL;DR for operators A small model can perform acceptably after task specialization and still lose more accuracy than expected when compressed. The paper examined here suggests that part of that loss may be shaped earlier, during fine-tuning, rather than being determined by the quantizer alone. Paper evidence: Muon-optimized models lost less accuracy after quantization than Adam-optimized models on six of eight benchmarks. On ARC-e, for example, accuracy fell by 0.55 percentage points after quantization with Muon versus 3.16 points with Adam. The integrated pipeline also produced a 2.86 GB quantized model from a 6.01 GB pre-quantized model while improving reported serving throughput and roughly halving per-token latency. ...

September 6, 2026 · 8 min · Zelina
Cover image

When Tables Learn the Meaning Behind Their Columns

TL;DR for operators Many enterprise prediction systems already perform well on structured tables, but those tables often contain information that traditional pipelines treat as symbols rather than meaning: product categories, descriptions, labels, and domain-specific terminology. The CASE framework explores whether language-model-derived representations can add this missing semantic layer without replacing the tabular models already used in production. ...

August 31, 2026 · 6 min · Zelina
Cover image

When the Gazetteer Goes Blank, the Records Still Know Where the Place Is

TL;DR for operators When a historical place name is missing from a gazetteer, the usual workflow treats it as a lookup failure. This paper shows another option: use multiple records that mention the same place from already known coordinates, convert descriptions such as “1 km east of X” into constraints on where X must be, and aggregate those clues. ...

August 26, 2026 · 7 min · Zelina