Cover image

Measure the Chain Before You Target the Reward

TL;DR for operators A tool agent may need to read records, retrieve identifiers, inspect intermediate state, and only then issue the write that completes a task. If the verifier scores only that final write, a seemingly precise per-turn reward can assign learning credit to one visible action while leaving the prerequisite chain unsupervised. ...

October 1, 2026 · 8 min · Zelina
Cover image

The Verifier Already Knows: Turn Pass/Fail Checks Into Training Credit

TL;DR for operators A long-running agent may execute dozens of actions before receiving one final pass/fail result. Applying that same terminal signal across the whole trajectory leaves training with little information about which earlier actions actually contributed to satisfying the task. VICT1 treats an existing programmatic verifier as a source of selective training credit. It decomposes the verifier into explicit checks, determines which checks are relevant to the rollout-level preference, and assigns additional credit only when trajectory evidence links a particular action to one of those checks. When the required evidence is missing or the verifier reconstruction is unreliable, the method falls back to the original outcome-based advantage. ...

September 26, 2026 · 7 min · Zelina
Cover image

From Chains to Trees: Why LLM Agents Need Structural Memory

Logs are useful. They are also lazy. A business agent that fails halfway through a product search, customer-support flow, compliance checklist, or research workflow will usually leave behind a long trace: thought, action, observation, thought, action, observation. The standard instinct is to read the failed trace as a chain. This step followed that step; the final reward was bad; therefore the chain was bad. Very tidy. Also very wasteful. ...

April 9, 2026 · 18 min · Zelina
Cover image

MatchTIR: Stop Paying Every Token the Same Salary

Payroll is a useful metaphor for agent training because it makes the absurdity obvious. Imagine a project team where one employee finds the right database, another enters the correct query, a third repeatedly calls the wrong API, and a fourth finally writes the report. If the report is accepted, everyone receives the same bonus. If it fails, everyone receives the same blame. Very democratic. Also very stupid. ...

January 17, 2026 · 16 min · Zelina
Cover image

Credit Where It's Due: How CAPO Brings Verifiable Precision to LLM Reasoning

TL;DR for operators CAPO is not mainly a paper about “making models reason better” in the usual fog-machine sense. It is about fixing a specific training failure: outcome-only reinforcement learning tells a model whether the final answer was right, but not which part of the reasoning earned or destroyed that outcome. The method uses a stronger off-the-shelf LLM as a generative process reward model, or GenPRM, to inspect a rollout and identify wrong reasoning steps in one pass. Those step-level critiques are then converted into token-level penalties, so the policy update can suppress flawed reasoning segments instead of treating the whole answer as one indivisible blob. The authors test this across Llama-3-1B/3B and Qwen2.5-1.5B/7B backbones, with results showing consistent average gains over SFT, GRPO with rule-based verification, and GRPO with generative outcome reward modelling.1 ...

August 5, 2025 · 14 min · Zelina