Cover image

Conducting the Agents: Why AORCHESTRA Treats Sub-Agents as Recipes, Not Roles

AOrchestra shows that the practical edge in multi-agent systems may come less from adding more agents and more from dynamically composing the right instruction, context, tools, and model for each subtask.

February 4, 2026 · 14 min · Zelina
Cover image

Conformal Thinking: Teaching LLMs When to Stop Thinking

A mechanism-first reading of Conformal Thinking, showing how risk-controlled early stopping turns reasoning budgets from guesswork into an operational error-budget decision.

February 4, 2026 · 17 min · Zelina
Cover image

More Isn’t Smarter: Why Agent Diversity Beats Agent Count

A mechanism-first reading of why multi-agent LLM systems saturate when agents repeat each other, and why useful diversity beats raw agent count.

February 4, 2026 · 16 min · Zelina
Cover image

Search-R2: When Retrieval Learns to Admit It Was Wrong

Search-R2 shows why reliable retrieval agents need local error repair, not just more search calls or larger rollout budgets.

February 4, 2026 · 16 min · Zelina
Cover image

When Agents Stop Talking to the Wrong People

TodyComm shows why multi-agent AI systems need learned communication governance, not just more agents talking more often.

February 4, 2026 · 15 min · Zelina
Cover image

When Papers Learn to Draw: AutoFigure and the End of Ugly Science Diagrams

AutoFigure shows why publication-ready scientific diagrams need reasoning-first visual pipelines, not prettier text-to-image prompts.

February 4, 2026 · 15 min · Zelina
Cover image

When Your Agent Starts Copying Itself: Breaking Conversational Inertia

A mechanism-first reading of conversational inertia: why long context can make agents imitate their own mistakes, and why strategic forgetting may beat bigger memory.

February 4, 2026 · 17 min · Zelina
Cover image

Click Like a Human: Why Avenir-Web Is a Quiet Breakthrough in Web Agents

Avenir-Web shows why reliable web agents need procedural experience, hybrid grounding, explicit progress tracking, and compressed memory—not just bigger multimodal models.

February 3, 2026 · 16 min · Zelina
Cover image

Click with Confidence: Teaching GUI Agents When *Not* to Click

SafeGround shows how uncertainty calibration can turn GUI agents from reckless clickers into risk-budgeted automation systems.

February 3, 2026 · 17 min · Zelina
Cover image

Coaching the Swarm: Why Multi‑Agent RL Finally Scales

A mechanism-first reading of MAPPA, a process-reward method for turning multiagent LLM workflows from prompted collaboration into trainable systems.

February 3, 2026 · 17 min · Zelina