Cover image

Measure the Chain Before You Target the Reward

TL;DR for operators A tool agent may need to read records, retrieve identifiers, inspect intermediate state, and only then issue the write that completes a task. If the verifier scores only that final write, a seemingly precise per-turn reward can assign learning credit to one visible action while leaving the prerequisite chain unsupervised. ...

October 1, 2026 · 8 min · Zelina

From Fragmented Reports to a Match-Ready Operating Picture

A composite professional football club redesigns its between-match preparation cycle around six coordinated agents that retrieve, reconcile, and refresh evidence while coaches and clinicians retain all consequential decisions.

September 30, 2026 · 8 min · Vox
Cover image

Keep the Rollout Honest: Miles Treats Throughput and Fidelity as One System

TL;DR for operators If rollout GPUs keep generating while training GPUs update the policy, higher utilization creates an accounting problem: some trajectories come from older weights, and serving and training stacks can assign different probabilities even when they nominally use the same model. The operating decision is therefore not simply how to remove idle time, but how much staleness and numerical mismatch the loop can tolerate. ...

September 27, 2026 · 6 min · Zelina
Cover image

Parallel by Default: Manufacturing Agents Need Routing, Not One Reasoning Mode

TL;DR for operators A manufacturing process planner has to preserve relationships across geometry, drawing requirements, material constraints, manufacturing rules, process order, and tooling. If a tolerance or surface-finish requirement is attached to the wrong CAD feature, later reasoning can be internally coherent and still produce the wrong plan. Design-to-Plan addresses that problem with a hybrid architecture: deterministic components handle precision-sensitive perception and machining calculations, while LLM agents reason through structured tools and retrieved manufacturing knowledge.1 Across three 100-case downstream benchmarks, parallel configurations reached 100% execution success and more consistently invoked the tools expected for each case. They also cut average token use from 22,481 to 8,896 for process sequencing and from 36,659 to 11,868 for tool selection. Knowledge retrieval was the exception, where parallel coordination slightly increased token use. ...

September 23, 2026 · 7 min · Zelina
Cover image

State Before Action: OODA-Tool Puts a Control Layer Between Context and Execution

TL;DR for operators A tool-using agent can remember the right customer, constraint, or prior result and still make the wrong call. From State to Action: OODA-Tool for Reliable Multi-Turn Tool Use1 treats that gap as an architectural problem rather than only a prompting problem. Its strongest configuration separates four jobs: reconstruct the active task state, decide whether execution is actually warranted, choose a permitted action structure, and only then bind concrete arguments. On ToolDial, Specialized OODA beats Direct-LoRA at every tested Qwen3 scale, by 4.48 to 6.99 percentage points in Task Success. The gains are largest where state must survive long histories, missing information, changed values, constraints, or sequential dependencies. ...

September 22, 2026 · 8 min · Zelina
Cover image

State Is the Workflow: AstronOS Moves Long-Horizon Agents Beyond Transcript Replay

TL;DR for operators A multi-step agent workflow eventually faces a problem that a larger context window does not solve. One model makes a decision, another step runs later with new evidence, and the system must know which earlier facts and decisions still count as official. Replaying the transcript gives the next executor more text; it does not necessarily tell it what has been accepted, superseded, or rejected. ...

September 22, 2026 · 7 min · Zelina
Cover image

The Graph Is Not the Guardrail: Route Retrieval by Failure Mode

TL;DR for operators A retrieval system may have several ways to answer the same analyst request: search semantically similar text, follow explicit relationships in a knowledge graph, repair a failed graph query, or combine graph and text evidence. The operational question is not which technique has the highest average score. It is which path fails acceptably for the workload in front of it. ...

September 19, 2026 · 8 min · Zelina
Cover image

Guide, Don’t Guess: Reallocating Reasoning Compute with Small-Model Hints

TL;DR for operators When a compact model fails on a difficult multi-step problem, the default remedies are expensive: use a larger model, generate many independent attempts, or accept lower accuracy. This paper tests a different allocation of compute—keep the solver small, but give it localized guidance at difficult intermediate steps. HintMR1 separates those jobs. A hinter tells the solver what to consider next without supplying the full solution; the solver then advances its reasoning one step at a time. On AIME-2024, DeepSeek-R1-Distill-Qwen-7B rises from 20.69% accuracy without hints to 68.97% with GPT-5.2-generated hints. The result is not evidence that any second model helps: non-fine-tuned small-model hints are inconsistent and sometimes reduce accuracy below the no-hint baseline. ...

September 18, 2026 · 6 min · Zelina
Cover image

When Reasoning Leaves the Prompt: Designing the Agentic Control Loop

TL;DR for operators When an LLM has to plan, call a tool, inspect the result, remember what happened, and decide what to do next, the system has changed in a more fundamental way than “adding more reasoning steps.” Agentic Reasoning for Large Language Models frames that change as a move from mostly static generation toward an interactive reasoning-and-control loop.1 ...

September 18, 2026 · 7 min · Zelina
Cover image

One Step Is Not a Workflow: Where LLM Rule Following Starts to Break

TL;DR for operators A model that is highly reliable at applying one explicit rule transition is not necessarily reliable at executing an entire procedure built from those transitions. In Reasoning Capabilities of Large Language Models. Lessons Learned from General Game Playing1, the strongest evaluated model, Gemini 2.5 Pro, achieves 95.6% exact success on one-step next-state generation. At five dependent state transitions, exact success falls to 73.4%. When the model must also choose actions during those five steps, it falls again to 65.3%. ...

September 17, 2026 · 7 min · Zelina