Cover image

Following Instructions Is Not the Same as Knowing More

TL;DR for operators A multimodal model can become much better at obeying instructions without becoming much better at the underlying tasks those instructions govern. In the experiments examined here, one 8B vision-language model gains 10.58 percentage points on a targeted instruction-following benchmark and another gains 22.91 points. Yet their average results across broader STEM, VQA, OCR, and document-understanding tests move by only +0.33 and -0.21 points. ...

October 1, 2026 · 7 min · Zelina
Cover image

Before the Test Comes the Question: The LLM Formulation Gap in Analytics

TL;DR for operators An analytics copilot can know statistics and still start from the wrong problem. StatFormBench tests what happens before statistical execution: given an informal request and heterogeneous data, can an LLM determine the statistical problem being asked, select the data objects that actually matter, and assign each one the correct analytical role? Across 1,013 human-reviewed scenarios and 14 LLMs, those abilities do not move together. Gemini 3.1 Pro achieves the highest fine-grained problem-classification accuracy at 72.0, while Claude Opus 4.6 achieves the highest variable-set overlap at 63.2. No evaluated model leads both components. ...

September 30, 2026 · 8 min · Zelina
Cover image

The Graph Isn’t the Verifier: What LCoT-GV Actually Learns From Long Reasoning Chains

TL;DR for operators A long reasoning trace creates an additional quality-control problem: even when the reasoning looks structured, the final answer can still be wrong. A verifier therefore has to identify signals inside the trace that predict answer correctness without simply trusting the model that produced it. LCoT-GV, introduced by Bérénice Jaulmes and Mehwish Alam,1 turns reasoning steps into a graph, connects steps when a local inference model judges them to support or contradict one another, and then uses a graph attention network to classify whether the final answer is correct. Across three reasoning models, its default configurations average 75.24%-77.92% accuracy. ...

September 29, 2026 · 7 min · Zelina
Cover image

Zero Is Not Nothing: When an Inactive Token Still Changes the Model

TL;DR for operators If adding an unused token whose hidden state is exactly zero changes existing tokens or predictions, that token is not actually inactive. Standard softmax attention can produce exactly this behavior: a visible zero source still contributes normalization mass even when its value contribution is zero. OAttention1 addresses this at the attention operator by giving each token a norm-derived presence coefficient. The coefficient gates what a receiver emits and, separately, how much support a source contributes to the attention numerator and denominator. Exact zero then means zero participation under the declared operator contract. ...

September 24, 2026 · 6 min · Zelina
Cover image

The Graph Is Not the Guardrail: Route Retrieval by Failure Mode

TL;DR for operators A retrieval system may have several ways to answer the same analyst request: search semantically similar text, follow explicit relationships in a knowledge graph, repair a failed graph query, or combine graph and text evidence. The operational question is not which technique has the highest average score. It is which path fails acceptably for the workload in front of it. ...

September 19, 2026 · 8 min · Zelina
Cover image

More Thought Is Not Always More Reliable: Routing Reasoning for Social Judgment

TL;DR for operators Extra inference compute is not a monotonic reliability upgrade for socially ambiguous tasks. Across three benchmarks, reasoning-focused models sometimes outperform non-reasoning counterparts, sometimes underperform them, and can become less accurate when deliberation is pushed harder on difficult cases. The practical lesson is not to suppress reasoning. Moderate reasoning, token limits, and adaptive stopping can improve results. Instead, treat reasoning depth as a control variable: decide when to invoke it, when to stop it, and whether the prompt format itself is steering the model toward shortcuts. ...

September 17, 2026 · 7 min · Zelina
Cover image

Four Inputs In, One Modality Out: Testing Whether Omnimodal Models Actually Arbitrate Evidence

TL;DR for operators A model can receive a camera feed, spoken report, reference image, and text record without meaningfully reasoning across all four. In C$^3$PO1, 86–95% of observed failures across ten models in the paper’s failure analysis were classified as dominance-driven: one modality or prior drove the answer while other evidence was effectively ignored. ...

September 13, 2026 · 7 min · Zelina
Cover image

The Wrong Answer May Start Before Reasoning

TL;DR for operators A model reads a chart, diagram, or photographed math problem and produces the wrong answer. Treating that event as a generic “reasoning failure” can send engineering effort to the wrong component. The model may have misread a number, attached a label to the wrong object, confused a scale or unit, or reasoned incorrectly after extracting the right facts. ...

September 13, 2026 · 8 min · Zelina
Cover image

The Table Is the Task: What DataSpace Reveals About Data-Agent Reliability

TL;DR for operators A data agent can find the right evidence, perform much of the required analysis, and still fail the task by returning the wrong table. In an audit of 136 failed runs from the strongest tested backbone, 71 failures—52.2%—were attributed primarily to turning the agent’s internal result into the requested output. Sixty of those involved submitting extra or missing columns. Only three failures were attributed to selecting the wrong evidence source. ...

August 30, 2026 · 7 min · Zelina
Cover image

Running Is Not Correct: Why Scientific Code Needs Graded Verification

TL;DR for operators For generated engineering code, successful execution should be treated as a feasibility check, not a correctness certificate. In the paper’s 7B ablation, reinforcement learning based only on program validity reaches pass@1 of 0.58 and pass@8 of 0.76. Adding continuous trajectory accuracy raises those figures to 0.71 and 0.84. The mechanism is straightforward. Generated solver code must first run, return the required output shape, and avoid non-finite values. Programs that clear those checks are then graded by how closely their numerical trajectories match hidden references and, where available, how consistent their outputs are with the governing PDE. ...

August 14, 2026 · 8 min · Zelina