Cover image

Before the Test Comes the Question: The LLM Formulation Gap in Analytics

TL;DR for operators An analytics copilot can know statistics and still start from the wrong problem. StatFormBench tests what happens before statistical execution: given an informal request and heterogeneous data, can an LLM determine the statistical problem being asked, select the data objects that actually matter, and assign each one the correct analytical role? Across 1,013 human-reviewed scenarios and 14 LLMs, those abilities do not move together. Gemini 3.1 Pro achieves the highest fine-grained problem-classification accuracy at 72.0, while Claude Opus 4.6 achieves the highest variable-set overlap at 63.2. No evaluated model leads both components. ...

September 30, 2026 · 8 min · Zelina
Cover image

The Graph Isn’t the Verifier: What LCoT-GV Actually Learns From Long Reasoning Chains

TL;DR for operators A long reasoning trace creates an additional quality-control problem: even when the reasoning looks structured, the final answer can still be wrong. A verifier therefore has to identify signals inside the trace that predict answer correctness without simply trusting the model that produced it. LCoT-GV, introduced by Bérénice Jaulmes and Mehwish Alam,1 turns reasoning steps into a graph, connects steps when a local inference model judges them to support or contradict one another, and then uses a graph attention network to classify whether the final answer is correct. Across three reasoning models, its default configurations average 75.24%-77.92% accuracy. ...

September 29, 2026 · 7 min · Zelina
Cover image

More Shots, More Coverage? Measure the Set, Not the Sample

TL;DR for operators If your system produces five patches, ten evidence candidates, or twenty molecular designs, choosing its model and inference settings from single-answer benchmarks can select the wrong configuration for the workflow you actually run. The central measurement problem is redundancy. Several outputs can each be valid and individually strong while repeatedly landing on the same task-relevant result. Conversely, outputs that look different in wording or structure may add no genuinely new option. ...

September 26, 2026 · 7 min · Zelina
Cover image

Make the Structure Survive: Forma Turns Synthetic Clinical Cases Into Auditable Outputs

TL;DR for operators If synthetic cases are going to support training or evaluation, surface plausibility is a weak control objective. The harder requirement is preserving the relationships that make each case meaningful. Forma tests that idea by specifying a person-specific psychological structure before generation and asking whether those directional relationships can be recovered afterward. In the full condition, an external probe reaches MCC +0.41 and AUC 0.70 for directed-edge recovery. When the structural formulation is removed, performance falls close to chance: MCC +0.03/AUC 0.52 with demographics and self-report still present, and +0.01/0.50 under zero-shot generation. ...

September 25, 2026 · 7 min · Zelina
Cover image

More Agents, More Rules: HiMA-MDD Treats Multi-Agent AI as a Governance Problem

TL;DR for operators Dividing a high-stakes decision among specialist agents does not specify who may see which evidence, who owns each subdecision, or who can revise it. HiMA-MDD1 turns those choices into explicit system rules for PHQ-8 assessment from completed multimodal clinical interviews. The strongest architectural signal is not “more agents perform better.” Before global verification, four specialists produce the lowest total-score error, while two specialists produce the highest screening kappa and Macro-F1. Removing cross-factor audit and targeted revision causes the largest screening-performance decline among the paper’s three component ablations. Giving four specialists bounded item-specific evidence also beats giving them a capped shared evidence pool on every reported pre-verification metric, although that experiment changes both evidence composition and context length. ...

September 25, 2026 · 6 min · Zelina
Cover image

Safety Has a Memory: Why Multimodal Jailbreak Testing Must Follow the Conversation

TL;DR for operators Safety testing for a multimodal assistant should cover sequences of interactions, not only whether the system refuses one obviously prohibited prompt. In the tested setup, a staged three-turn attack reached a 91.50% attack success rate on LLaVA-7B and 77.31% on GPT-4o, above the three single-turn attack baselines reported for those models.1 The result does not establish universal failure rates, but it does show that a prompt-level pass can miss vulnerabilities that emerge after earlier turns establish conversational context. ...

September 20, 2026 · 7 min · Zelina
Cover image

More Retrieval Is Not Free: Price Every RAG Component Before You Ship It

TL;DR for operators When a knowledge-grounded assistant needs improvement, adding another retrieval stage is not automatically the safest use of inference budget. In this study, the most expensive retrieval option in the matched comparison—combining semantic and keyword-based search—reduced accuracy by 1.85 percentage points relative to dense retrieval while adding 3,079.69 seconds of runtime across the evaluation run. More machinery produced a worse benchmark result. ...

September 18, 2026 · 7 min · Zelina
Cover image

More Thought Is Not Always More Reliable: Routing Reasoning for Social Judgment

TL;DR for operators Extra inference compute is not a monotonic reliability upgrade for socially ambiguous tasks. Across three benchmarks, reasoning-focused models sometimes outperform non-reasoning counterparts, sometimes underperform them, and can become less accurate when deliberation is pushed harder on difficult cases. The practical lesson is not to suppress reasoning. Moderate reasoning, token limits, and adaptive stopping can improve results. Instead, treat reasoning depth as a control variable: decide when to invoke it, when to stop it, and whether the prompt format itself is steering the model toward shortcuts. ...

September 17, 2026 · 7 min · Zelina
Cover image

Are You Sure? Reliability Starts After the First Answer

TL;DR for operators A model can answer correctly, be challenged by the user, and then talk itself into being wrong. Deployment gates that look only at first-turn accuracy or reported confidence can miss that failure. Saadat and Nemzer’s Certainty Robustness Benchmark1 tests 200 LiveBench math and reasoning questions with independent follow-ups: “Are you sure?”, “You are wrong!”, and a request for 1–100 confidence. GPT-5.2 and Claude Sonnet 4.5 began at almost the same accuracy, yet each collapsed under a different form of pushback. ...

September 10, 2026 · 5 min · Zelina
Cover image

Confidence Needs a Difficulty Check Before It Routes Work

TL;DR for operators A workflow that uses model confidence to auto-accept an answer, escalate it, or send it to human review depends on more than whether the underlying model is accurate. The confidence signal itself has to distinguish cases the model should find easy from cases it should find difficult. Chen et al. test this distinction in Latent Confidence Alignment for LLM Self-Assessment.1 Across 20 LLMs and 100 text-only MedXpertQA questions, supplying an external difficulty signal significantly improved the alignment between models’ stated error probabilities and their expected error probabilities. Structured reflection alone did not significantly improve that alignment. At the same time, latent task ability showed no significant differences across the four evaluated conditions. ...

September 10, 2026 · 7 min · Zelina