Cover image

Forgotten Until Asked Differently: Unlearning Needs an Adversarial Sign-Off

TL;DR for operators A model can pass an ordinary unlearning evaluation while still yielding supposedly forgotten information when the request is reformulated strategically. Gupta and colleagues demonstrate this gap in a controlled benchmark of LLM unlearning: on the 1% forget split, four fine-tuning-based methods produced average adversarial recovery rates between 72.8% and 84.3%, compared with 87.5% for the unprotected model.1 ...

September 26, 2026 · 7 min · Zelina
Cover image

More Shots, More Coverage? Measure the Set, Not the Sample

TL;DR for operators If your system produces five patches, ten evidence candidates, or twenty molecular designs, choosing its model and inference settings from single-answer benchmarks can select the wrong configuration for the workflow you actually run. The central measurement problem is redundancy. Several outputs can each be valid and individually strong while repeatedly landing on the same task-relevant result. Conversely, outputs that look different in wording or structure may add no genuinely new option. ...

September 26, 2026 · 7 min · Zelina
Cover image

The Benchmark Is in the Trace: Reusing Agent Trajectories to Shrink SWE Evaluation

TL;DR for operators Repository-level agent benchmarks are expensive enough that teams have a strong incentive to run only part of them. The risk is not merely estimating the wrong average score; a poorly chosen subset can also change which agent appears better. PTA-IRT uses historical execution traces to make that subset more informative.1 At a 10% calibration budget, it records the best reported MAE, Kendall’s tau, and Spearman’s rho in every metric column across the four evaluated SWE-bench variants. Its averages are 0.041 MAE, 0.888 tau, and 0.973 rho. ...

September 22, 2026 · 8 min · Zelina
Cover image

More Retrieval Is Not Free: Price Every RAG Component Before You Ship It

TL;DR for operators When a knowledge-grounded assistant needs improvement, adding another retrieval stage is not automatically the safest use of inference budget. In this study, the most expensive retrieval option in the matched comparison—combining semantic and keyword-based search—reduced accuracy by 1.85 percentage points relative to dense retrieval while adding 3,079.69 seconds of runtime across the evaluation run. More machinery produced a worse benchmark result. ...

September 18, 2026 · 7 min · Zelina
Cover image

One Step Is Not a Workflow: Where LLM Rule Following Starts to Break

TL;DR for operators A model that is highly reliable at applying one explicit rule transition is not necessarily reliable at executing an entire procedure built from those transitions. In Reasoning Capabilities of Large Language Models. Lessons Learned from General Game Playing1, the strongest evaluated model, Gemini 2.5 Pro, achieves 95.6% exact success on one-step next-state generation. At five dependent state transitions, exact success falls to 73.4%. When the model must also choose actions during those five steps, it falls again to 65.3%. ...

September 17, 2026 · 7 min · Zelina
Cover image

When Vision Fails in Both Directions

TL;DR for operators A model can correctly recognize that one object is left of another and still fail when asked to generate that same relationship. More importantly, some visual weaknesses recur in both directions. AMVICC maps visual-language understanding and image generation onto visual concepts derived from the same underlying benchmark. Across the tested systems, weaknesses repeatedly appear in Quantity and Count, Positional and Relational Context, Orientation and Direction, and State and Condition. Text behaves differently: most tested multimodal language models avoid the paper’s Text failure threshold, while all three tested image generators fall below it for explicit generation. ...

September 13, 2026 · 7 min · Zelina
Cover image

The Best Channel Model Depends on the Channel

TL;DR for operators A wireless team choosing a model to standardize has a tempting shortcut: benchmark several candidates, find the highest score, and carry that winner into deployment-specific adaptation. The problem is that wireless-channel performance depends heavily on what is being predicted and on the propagation environment in which the model is evaluated. ...

September 12, 2026 · 7 min · Zelina
Cover image

Are You Sure? Reliability Starts After the First Answer

TL;DR for operators A model can answer correctly, be challenged by the user, and then talk itself into being wrong. Deployment gates that look only at first-turn accuracy or reported confidence can miss that failure. Saadat and Nemzer’s Certainty Robustness Benchmark1 tests 200 LiveBench math and reasoning questions with independent follow-ups: “Are you sure?”, “You are wrong!”, and a request for 1–100 confidence. GPT-5.2 and Claude Sonnet 4.5 began at almost the same accuracy, yet each collapsed under a different form of pushback. ...

September 10, 2026 · 5 min · Zelina
Cover image

Who Did What, When—and From Which Camera? The Perception Gap Behind Video Agents

TL;DR for operators A model can correctly recognize what is visible in a video and still lose track of who performed an action, how many times an event occurred, or which viewpoint observed it. GameplayQA1 makes those failures separately measurable. Across 16 evaluated multimodal models, average accuracy falls from 61.2% on single-reference tasks to 56.0% on temporal tasks and 49.4% on synchronized cross-video tasks. Occurrence counting averages just 36.5%, while cross-video ordering reaches 38.8%. Other-agent actions and states are also harder than questions about world objects. ...

September 8, 2026 · 7 min · Zelina
Cover image

Hide the Worker, Keep the Geometry: What SynthSite Changes About Privacy-Aware Safety Video

TL;DR for operators A safety team may need to conceal workers before video leaves a trusted environment, yet hiding appearance can also remove the geometry needed to detect whether someone is beneath a suspended load. SynthSite tests that conflict directly.1 Across 55 synthetic clips, cartooning produced the highest F2 against human safe/unsafe labels at 0.767, while Canny-edge produced the highest F2 for reproducing the raw-video pipeline at 0.964. ...

August 18, 2026 · 7 min · Zelina