Cover image

Progress Is Not Completion: What CAP Reveals About Browser-Agent Readiness

TL;DR for operators A browser automation can search several sites, manipulate interfaces, gather useful information, and return a polished response while still missing one requirement that makes the workflow unusable. It may apply the wrong filter, misread a value in a chart, or fail to notice that a panel is collapsed. For production decisions, visible progress is not the same thing as reliable completion. ...

August 30, 2026 · 8 min · Zelina
Cover image

The Table Is the Task: What DataSpace Reveals About Data-Agent Reliability

TL;DR for operators A data agent can find the right evidence, perform much of the required analysis, and still fail the task by returning the wrong table. In an audit of 136 failed runs from the strongest tested backbone, 71 failures—52.2%—were attributed primarily to turning the agent’s internal result into the requested output. Sixty of those involved submitting extra or missing columns. Only three failures were attributed to selecting the wrong evidence source. ...

August 30, 2026 · 7 min · Zelina
Cover image

Higher Pass Rate, More Broken Tasks: The Regression Tax in Agent Skill Libraries

TL;DR for operators Adding reusable instructions to an agent creates a release-management problem that average accuracy does not fully expose. A library can solve tasks the baseline missed while simultaneously breaking tasks the baseline handled correctly. Darshan Tank and Baran Nama measure that trade-off directly in The Regression Tax: Decomposing Why Skills Help and Hurt LLM Agents.1 Across 18 skill-library conditions, they observe 553 baseline-fail-to-skill-pass transitions but also 324 baseline-pass-to-skill-fail transitions. The new failures offset 59% of the gross gains. ...

August 19, 2026 · 7 min · Zelina
Cover image

English Looks Ready. Amharic Says Otherwise: What ADAGE Exposes in Multilingual Evaluation

TL;DR for operators A multilingual model can look ready on an English reasoning benchmark and still perform close to chance in a strategically important native language. In the reported zero-shot evaluation, Gemma 3 27B scores 83.0% on English ePiC, 70.1% on Arabic CAPR, 41.3% on Amharic CAPR, and 86.0% on Japanese CAPR. ...

August 16, 2026 · 7 min · Zelina
Cover image

Safe at the Finish, Unsafe on the Way: What SafeRelBench Exposes in Embodied AI

TL;DR for operators A manipulation agent can reach the requested final state and still have executed the task unsafely. That creates a measurement problem for teams using task-completion rates to decide whether an embodied model, prompt, or policy update is ready for deployment. SafeRelBench tests this gap directly. Across seven evaluated VLM-driven agents, the spatial-relation cases produced task Success Rates (SR) of 0.52–0.73 but Safety Success Rates (SSR) of only 0.16–0.40. In matched non-spatial settings, SR rose to 0.83–0.94 and SSR reached as high as 0.91.1 The benchmark therefore measures something final-state success can miss: whether the agent satisfied the relevant safety prerequisite before taking the risky action. ...

August 16, 2026 · 8 min · Zelina
Cover image

The Leaderboard Is Not a Clinical Clearance

TL;DR for operators A healthcare team choosing a model, configuration, and safeguards for diagnostic support should treat leaderboard leadership as evidence of capability—not proof of clinical readiness. GPT-5-medium led ClinMM-Bench, yet produced a completely correct diagnosis in only 33.88% of cases. The benchmark tests a difficult but essential requirement: combining clinical details and images as evidence unfolds, revising earlier conclusions, and explaining a diagnosis without omitting decisive information or introducing unsupported claims. Medical specialization and explicit reasoning settings do not improve these abilities consistently across model scales, metrics, or specialties. ...

August 6, 2026 · 8 min · Zelina
Cover image

Stocked but Not Synthesizable: URSA Tests the Chemistry Inside the Route

TL;DR for operators A team choosing a synthesis-planning system may reasonably treat a route that ends in purchasable starting materials as nearly ready for chemical review. URSA shows why that assumption is risky: reaching stocked compounds proves that a route graph can terminate in available inputs, not that the proposed reactions are chemically plausible. ...

July 31, 2026 · 8 min · Zelina
Cover image

Agree Once, Remember Later: The Commit Boundary in Personal Agents

TL;DR for operators A personal assistant can hear a confident user claim, store it as a preference or rule, and rely on it during a later task after the original conversation is gone. The safety problem is therefore not only the agreeable reply. It is the write that lets the claim survive. ...

July 24, 2026 · 8 min · Zelina
Cover image

The Room Remembers, the Model Forgets

TL;DR for operators A room-tour video is a deceptively simple test for a video model. The objects do not explode, the camera does not enter a car chase, and nobody asks the model to perform cinematic philosophy. The hard part is duller and therefore more operationally relevant: the model must remember where things were, how rooms connected, what changed, and which earlier view matters now. ...

July 2, 2026 · 17 min · Zelina
Cover image

The Solver Isn’t the Strategy: FrontierOR’s Reality Check for AI Optimisation Agents

Scheduling a factory, routing a fleet, pricing airline seats, allocating scarce capacity: these are not “write me a Python script” problems with nicer stationery. In real operations research, the useful answer is not merely a correct mathematical model. It is a method that stays feasible, keeps solution quality high, and finishes before the business context has expired. ...

June 14, 2026 · 15 min · Zelina