Cover image

A Research Agent Should Leave a Paper Trail

TL;DR for operators A long-running research agent can produce an impressive manuscript while still leaving a manager unable to reconstruct what evidence was gathered, what failed, which claims were checked, or where a human should intervene. pAI/MSc1 is most useful as a response to that problem: although its fixed workflow uses 23 specialist agents across 30 graph nodes, its more consequential design choice is to preserve discovery, planning, theory, experimentation, synthesis, review, checkpoints, and budget accounting as named artifacts that can be inspected, resumed, audited, and structurally validated. ...

September 2, 2026 · 6 min · Zelina
Cover image

Benchmarks with Benefits: What DeepScholar-Bench Really Measures

TL;DR for operators DeepScholar-Bench is useful because it turns “deep research” from a demo category into a measurable workflow: retrieve the right sources, synthesize the right facts, and attach citations that actually support the claims.1 The headline result is not flattering. No evaluated system exceeds a 31% geometric mean across all metrics. OpenAI DeepResearch leads overall with a 0.309 geometric mean, but its best-looking strengths hide serious gaps: 0.857 on organization, 0.392 on nugget coverage, 0.187 on reference coverage, and 0.124 on document importance. Translation: the report may read well while still missing the intellectual furniture. ...

August 30, 2025 · 14 min · Zelina