Cover image

One Forecast, Many Explanations: Why Time-Series Attribution Needs a Horizon Axis

TL;DR for operators A multi-step forecast may produce one trajectory, but the model does not necessarily use the same historical evidence for every point on that trajectory. In the real-world trained-model experiments studied here, an explanation assigned to its own forecast step outperformed explanations borrowed from other steps: the median own-versus-mismatched margin was +0.1418, and 80.6% of runs were positive. ...

September 28, 2026 · 7 min · Zelina
Cover image

The 99% Problem: When a Stroke Benchmark Looks Ready Before It Is

TL;DR for operators A model reporting 99% accuracy creates an obvious decision pressure: should the team fund integration, start clinical workflow design, or treat the experiment as essentially solved? This paper is a useful example of why that decision cannot be made from the headline number alone. Its main results table reports a Stacking Classifier at 99.81% accuracy, Random Forest at 99.52%, and Bagging at 99.45%.1 But those results come from one public Kaggle dataset after the original 5,110 records—249 stroke-positive and 4,861 stroke-negative—were reduced to a balanced working dataset containing 249 observations in each class. ...

September 27, 2026 · 7 min · Zelina
Cover image

Forgotten Until Asked Differently: Unlearning Needs an Adversarial Sign-Off

TL;DR for operators A model can pass an ordinary unlearning evaluation while still yielding supposedly forgotten information when the request is reformulated strategically. Gupta and colleagues demonstrate this gap in a controlled benchmark of LLM unlearning: on the 1% forget split, four fine-tuning-based methods produced average adversarial recovery rates between 72.8% and 84.3%, compared with 87.5% for the unprotected model.1 ...

September 26, 2026 · 7 min · Zelina
Cover image

Put the LLM Upstream, Not in the Loop

TL;DR for operators The most consequential choice in this paper is where the LLM is allowed to act. It does not decide whether simulated farms adopt solar panels. Instead, it works upstream: drafting behavioural rubrics and techno-economic scenarios that humans inspect, validate against explicit rules, freeze, and then feed into an existing calibrated simulation. ...

September 21, 2026 · 7 min · Zelina
Cover image

Sparse Routing May Buy You Inspectability, Not Just Efficiency

TL;DR for operators Herbst, Wermter, and Lee find that the analyzed Mixture-of-Experts models often represent tested concepts in far fewer neurons than comparable dense transformers.1 The difference is largest under the hardest probe constraint: when only one neuron is available, MoE experts often approach their own best probe performance while dense feed-forward layers need more dimensions. Models with sparser routing also tend to show cleaner representations. ...

September 11, 2026 · 8 min · Zelina
Cover image

An 8/10 Is Not a Probability: Validating LLM Confidence Before It Controls Workflow

TL;DR for operators A model saying “8/10 confident” does not mean its underlying uncertainty is approximately 20%. Across the evaluated settings, the average instance-level correlation between reported confidence and logits-based confidence is only 0.135. The more useful operating rule is narrower. First test whether reported scores vary enough to distinguish cases. Then measure whether those scores rank examples meaningfully on held-out data. Separately test whether their numerical scale agrees with the comparison signal and whether they are calibrated against correctness. Do not substitute one test for another. ...

September 10, 2026 · 7 min · Zelina
Cover image

Confidence Needs a Difficulty Check Before It Routes Work

TL;DR for operators A workflow that uses model confidence to auto-accept an answer, escalate it, or send it to human review depends on more than whether the underlying model is accurate. The confidence signal itself has to distinguish cases the model should find easy from cases it should find difficult. Chen et al. test this distinction in Latent Confidence Alignment for LLM Self-Assessment.1 Across 20 LLMs and 100 text-only MedXpertQA questions, supplying an external difficulty signal significantly improved the alignment between models’ stated error probabilities and their expected error probabilities. Structured reflection alone did not significantly improve that alignment. At the same time, latent task ability showed no significant differences across the four evaluated conditions. ...

September 10, 2026 · 7 min · Zelina
Cover image

Decontamination Is a Dial, Not a Delete Key

TL;DR for operators A benchmark score can become unreliable when test items, close paraphrases, or related material have entered training. Removing suspicious questions sounds straightforward, but any detector that misses contaminated items leaves score inflation behind, while filtering also changes the evaluation set. Chai, Zhe, and Sakuma propose DeconIEP,1 a white-box inference-time intervention that keeps the benchmark text and model weights fixed. Instead, it learns small, input-specific changes to the model’s embeddings so contaminated behavior moves closer to a comparatively less-contaminated reference model. ...

September 10, 2026 · 7 min · Zelina
Cover image

Fine-Tuning Changes What Your Model’s Errors Reveal

TL;DR for operators A fine-tuned model can become only slightly more accurate while its remaining errors become substantially easier to distinguish from correct answers. That matters when uncertainty scores feed operational controls. If a production workflow accepts an answer, abstains, calls another model, or sends a case to human review according to a detector threshold, fine-tuning changes more than the benchmark score. It can change the detector itself as an operating signal. ...

September 10, 2026 · 7 min · Zelina
Cover image

Privacy Starts Before the First Gradient

TL;DR for operators Federated fine-tuning keeps raw examples on the client, but that does not mean the client begins from a neutral model state. A malicious coordinating server can send an adapter deliberately structured so that private examples produce recoverable traces during training. Privacy risk can therefore enter through what the client downloads, not only through what it later uploads. ...

August 28, 2026 · 6 min · Zelina