Cover image

When AUROC Agrees and the Decision Still Changes

TL;DR for operators CRS-Bench1 tests a decision that clean-test AUROC does not fully answer: which pretrained medical image encoder deserves the next round of engineering, labeling, adaptation, and validation investment. Across 15 encoder families, AUROC and the benchmark’s broader reliability score are positively associated, yet 21 of 105 pairwise model choices reverse. The disagreement is not evidence that AUROC is useless. It shows that discrimination can preserve the broad ordering while missing differences in calibration, label efficiency, and robustness that change actual selection decisions. ...

September 22, 2026 · 7 min · Zelina
Cover image

Rerank the Regime, Not the Corpus

TL;DR for operators A RAG system can fail even when retrieval looks superficially healthy: several top-ranked documents repeat the query vocabulary, while the document that actually contains the answer sits lower in the list. The paper studies a reranking rule designed specifically for that situation.1 Instead of rewarding more query-document overlap, it removes words that exactly or semantically echo the query and scores what remains. In an eight-theme controlled keyword-stuffing diagnostic, this semantic variant improves mean target rank from 2.88 to 1.25 and pushes stuffed distractors from mean rank 2.00 to 4.50. ...

September 19, 2026 · 7 min · Zelina
Cover image

An 8/10 Is Not a Probability: Validating LLM Confidence Before It Controls Workflow

TL;DR for operators A model saying “8/10 confident” does not mean its underlying uncertainty is approximately 20%. Across the evaluated settings, the average instance-level correlation between reported confidence and logits-based confidence is only 0.135. The more useful operating rule is narrower. First test whether reported scores vary enough to distinguish cases. Then measure whether those scores rank examples meaningfully on held-out data. Separately test whether their numerical scale agrees with the comparison signal and whether they are calibrated against correctness. Do not substitute one test for another. ...

September 10, 2026 · 7 min · Zelina
Cover image

Confidence Is Not a Stop Signal: Test Whether the Model Knows When Information Is Missing

TL;DR for operators A model can be given an explicit way to say “the available information is insufficient” and still choose an unsupported answer most of the time. Tahermazandarani, Mahmood, Islam, and Sheng test this directly across five LLMs.1 They remove the correct answer from medical multiple-choice questions, replace it with an insufficient-information option, and observe abstention rates ranging from just 0.156 to 0.382. Reported unsafe rates range from 0.186 to 0.828. In a separate experiment, progressively stronger warnings that the clinical information may be incomplete or ambiguous also produce little reduction in model confidence. ...

September 10, 2026 · 7 min · Zelina
Cover image

Predicting the Experiment Is Easier Than Knowing When to Trust the Prediction

TL;DR for operators SciPredict finds that frontier LLMs predict outcomes of recent natural-science experiments with roughly 14-26% accuracy, compared with about 20% for domain experts. That headline can make model performance look surprisingly competitive. It should not be read as evidence that these systems are ready to decide which experiments can safely be skipped. ...

September 2, 2026 · 8 min · Zelina
Cover image

Before You Retrain the Guardrail, Ask What It Already Knows

TL;DR for operators When a deployed safety classifier begins making the wrong decisions, retraining the classifier is not necessarily the first intervention to test. Sandoval and Topcu’s Regime-Conditional Verification (RCV)1 adds a small correctness layer around a frozen classifier. It asks a narrower question than the classifier itself: given that the classifier just said “safe” or “unsafe,” how likely is that verdict to agree with the deployer’s policy? ...

August 29, 2026 · 9 min · Zelina
Cover image

Put the Error in Its Place: Why Reliable AI Is a Layering Problem

TL;DR for operators The usual response to an unreliable AI system is to ask for a larger model. That is frequently an expensive way to avoid diagnosing the actual error. Three recent papers point to a more disciplined alternative: Rules that must never be violated should be enforced at generation time, not merely suggested in a prompt. Stable patterns in the problem domain should be built into the model architecture, so the model does not have to rediscover them from every dataset. Residual temporal, class, and modality errors may be better handled through calibration, smoothing, routing, and fusion than through another round of full-model training. These interventions provide different kinds of assurance. A grammar mask can make certain outputs unreachable. An architectural prior can make desirable patterns more likely. Calibration can improve observed performance but usually cannot guarantee behavior. Bigger models still matter where genuine semantic reasoning is required. The lesson is not “small beats large.” It is “do not pay a large model to solve a problem that a rule, prior, or threshold can solve more reliably.” For business leaders, this changes the architecture question from “Which model should we buy?” to “Which layer should own each requirement?” ...

July 21, 2026 · 20 min · Zelina
Cover image

Veto Later, Repair First

TL;DR for operators Most decision systems treat hard constraints like a trapdoor. Candidate violates requirement, candidate disappears. Efficient, clean, and occasionally absurd. The paper behind Repair-Augmented Constraint Learning, or RACL, argues that this is the wrong semantics for systems that already know how to modify an option before showing it to the user.1 A flight missing a checked bag, a hotel missing breakfast, a product bundle missing an accessory, or a schedule slot needing a resource adjustment may not be a bad option. It may be a good option one repair away from being acceptable. ...

June 26, 2026 · 20 min · Zelina
Cover image

Judge, Jury, and Calibration: Why AI Evaluation Needs Anchors

TL;DR for operators AI is becoming very good at producing judgement-shaped output. That is not the same thing as judgement. Two recent papers make the same operational point from different sides: one shows how AI can estimate educational item difficulty before response data are available; the other shows how LLM-generated peer reviews can look serious while diverging from human reviewing behaviour.12 ...

June 15, 2026 · 14 min · Zelina
Cover image

Trust Me, I’m Benchmarked: Why Enterprise AI Needs Two Audits

Enterprise AI has developed two favorite comfort blankets: the model’s confident explanation and the benchmark score. The first says, “Relax, I reasoned through this.” The second says, “Relax, I scored well on a public test.” Both are useful. Neither is a warranty. And when business teams treat either as proof of reliability, the result is not governance. It is theatre with better typography. ...

June 10, 2026 · 14 min · Zelina