Cover image

The Agent Benchmark Without the Agent Bill

Pace shows how carefully selected static tests can screen models for expensive agentic evaluations—provided the proxy remains a filter rather than a substitute for reality.

July 16, 2026 · 16 min · Zelina
Cover image

The Refusal Rate That Refuses to Reassure

A large automated red-team study shows why high aggregate refusal rates can conceal concentrated, inexpensive, and operationally significant jailbreak exposure.

July 16, 2026 · 18 min · Zelina
Cover image

Don’t Retrain the Whole Map When One Neighborhood Moves

A controlled drift study shows when cluster-local monitoring can preserve model performance without paying the full cost of continuous retraining.

July 15, 2026 · 20 min · Zelina
Cover image

Say Less: A Child-Speech Screener Designed to Stop Before Diagnosis

A Polish child-speech pipeline shows that useful AI screening depends less on grand diagnostic claims than on preserving errors, controlling false alarms, and knowing when to remain silent.

July 15, 2026 · 19 min · Zelina
Cover image

The Skill Library That Could Read but Couldn’t Run

Trajectory mining can produce readable agent skills, but this paper shows why readability is not evidence of reusable automation.

July 15, 2026 · 21 min · Zelina
Cover image

Mind the Trigger: When AI Should Read the Room

A causal model reframes theory of mind as an expensive reasoning mode that AI should invoke selectively, not a social-intelligence feature left permanently switched on.

July 14, 2026 · 23 min · Zelina
Cover image

Pick the Mistake Before You Pick the Metric

A practical guide to choosing clustering evaluation metrics according to the errors, entities, and business priorities that should actually count.

July 14, 2026 · 15 min · Zelina
Cover image

Two Heads, One Error Budget

Multi-agent reasoning can rescue a weak model or corrupt a strong one; the operational challenge is deciding when communication deserves to happen.

July 14, 2026 · 17 min · Zelina
Cover image

Architecture Search Has a Training Problem

Neural architecture search works only when the search process respects how each candidate model must be trained.

July 13, 2026 · 18 min · Zelina
Cover image

Count the Missing, Weight the Rare: A Better Bargain for Cardiac Phenotyping

A class-weighted XGBoost pipeline shows why clinical tabular AI improves when imbalance, missingness, and error priorities are designed together rather than patched separately.

July 13, 2026 · 17 min · Zelina