The Agent Benchmark Without the Agent Bill
Pace shows how carefully selected static tests can screen models for expensive agentic evaluations—provided the proxy remains a filter rather than a substitute for reality.
Pace shows how carefully selected static tests can screen models for expensive agentic evaluations—provided the proxy remains a filter rather than a substitute for reality.
A large automated red-team study shows why high aggregate refusal rates can conceal concentrated, inexpensive, and operationally significant jailbreak exposure.
A controlled drift study shows when cluster-local monitoring can preserve model performance without paying the full cost of continuous retraining.
A Polish child-speech pipeline shows that useful AI screening depends less on grand diagnostic claims than on preserving errors, controlling false alarms, and knowing when to remain silent.
Trajectory mining can produce readable agent skills, but this paper shows why readability is not evidence of reusable automation.
A causal model reframes theory of mind as an expensive reasoning mode that AI should invoke selectively, not a social-intelligence feature left permanently switched on.
A practical guide to choosing clustering evaluation metrics according to the errors, entities, and business priorities that should actually count.
Multi-agent reasoning can rescue a weak model or corrupt a strong one; the operational challenge is deciding when communication deserves to happen.
Neural architecture search works only when the search process respects how each candidate model must be trained.
A class-weighted XGBoost pipeline shows why clinical tabular AI improves when imbalance, missingness, and error priorities are designed together rather than patched separately.