TL;DR for operators
A tool agent may need to read records, retrieve identifiers, inspect intermediate state, and only then issue the write that completes a task. If the verifier scores only that final write, a seemingly precise per-turn reward can assign learning credit to one visible action while leaving the prerequisite chain unsupervised.
The paper studied here finds that this is not a minor implementation detail. In low-observability settings, spreading a dense terminal advantage broadly across the trajectory can outperform concentrating the same credit on turns identified as making progress. A randomized control that preserves the same concentration but destroys the targeting information performs about as well as the targeted version, pointing to coverage breadth rather than targeting accuracy as the central problem.
For teams training tool-using agents, the sequence of decisions should change: measure how much of the necessary workflow the verifier can independently observe before paying for more sophisticated turn-level targeting. When reads and exploratory calls remain invisible, broader credit propagation is a defensible baseline. Richer process verification is more compelling when it actually exposes previously unseen prerequisite steps rather than merely localizing an already narrow signal.
The final write can hide most of the successful workflow
Consider an agent completing a business workflow against an API or database. Before it changes anything, it may need to find the right account, inspect its current state, recover identifiers, check constraints, and combine information from several calls. Only the final operation produces the state change that an automated evaluator can easily verify.
That final operation is highly visible. It is not necessarily the only action that caused success.
This is the problem examined by Zhou, Jiang, Wu, and Zhou in Coverage, Not Targeting: A Structural Regime in Multi-Turn Agent Credit Assignment.1 Their real-rollout diagnostics make the mismatch concrete. On successful tau^2-bench retail trajectories, the progress-derived credit signal has median effective breadth $k=1$ and collapses onto one terminal batched database write in 98% of trajectories. Successful trajectories, however, require roughly five to eight causally necessary tool calls.
The prerequisite steps are not merely conversational overhead. In the analyzed successful rollouts, arguments used by the terminal write depend on information obtained from precursor reads. Concentrating learning credit on the visible write therefore removes direct signal from actions on which that write depends.
Verifier information density measures the missing supervision
The paper formalizes this mismatch with verifier information density:
Here, $C$ is the number of causally necessary turns in the prerequisite chain, while $k$ is the effective number whose correctness the verifier independently exposes.
This is not a trajectory-length statistic. A ten-turn dialogue may contain only a few causally necessary actions, while a shorter tool sequence may require every call. The relevant denominator is the causal chain.
For tau^2-bench, alternative definitions of $C$ produce median values around six to eight while effective breadth remains about one, placing measured $V_d$ around 0.12-0.20. BFCL V3 exposes roughly two steps in chains of about five to six, giving $V_d$ around 0.4.
The structural constraint is especially severe for state-transition verifiers. Without oracle labels or a learned semantic judge, they can independently check actions that change observable state but usually cannot validate prerequisite reads. The paper expresses this as:
The verifier can therefore become more precise about the write without learning much more about the chain that made the write possible.
Dense reward helps, but concentrating it can still hurt
The first empirical distinction is between reward density and credit geometry.
On the main tau^2 retail experiment with Qwen3-14B, dense uniform credit improves held-out assertion support by 0.079 relative to binary outcome reward. Official task success improves by 5.2 percentage points. A three-level dense reward retains roughly 96% of the dense-versus-binary improvement, suggesting that escaping a nearly degenerate binary signal matters more here than extracting very fine reward resolution.
But density alone does not solve the problem.
When the same dense signal is concentrated onto progress-associated turns, performance falls. Progress-targeted per-turn credit trails uniform credit by 0.053 in mean assertion support and by 0.063 in official task success.
That result challenges a common intuition: if the system can identify the turn where measurable progress occurred, assigning more credit there should be more informative. That reasoning works only if the identified turns cover enough of the actions that success actually depends on.
The shuffled control separates targeting from concentration
A weaker experiment could leave an obvious explanation: perhaps the progress heuristic simply picked the wrong turns.
The paper addresses this with a matched-concentration shuffled control. It preserves the same narrow allocation pattern but randomly relocates the weights. If targeting information were doing substantial work, destroying that information should create an additional performance loss.
Instead, shuffled concentrated credit performs within noise of progress-targeted credit on tau^2, while both remain below uniform allocation. On BFCL, progress-targeted minus uniform official success is -0.036, shuffled minus uniform is -0.034, and shuffled minus targeted is approximately +0.002.
This does not prove that the two concentrated schemes are exactly equivalent. It does provide counterevidence to the claim that poor localization is the main reason concentrated credit loses.
The experiment changes the question from “Did we find the right turn?” to “How much of the causal chain receives any learning signal at all?”
The breadth sweep directly tests the coverage mechanism
The strongest intervention on the proposed mechanism is the pre-registered ToolACE breadth sweep.
The experiment holds total absolute credit fixed while widening its distribution. Relative to normalized full-chain coverage, a single-turn spike is down 0.048. Checker-visible steps reduce the deficit to 0.017. Covering all tool calls reaches +0.010 with a confidence interval spanning zero, effectively eliminating the deficit at broad coverage.
That monotone pattern matters more mechanistically than another uniform-versus-targeted comparison. The total credit budget is held constant while its breadth changes.
Synthetic experiments provide a complementary regime test. Concentrated targeting overtakes uniform allocation only when verifier information density becomes very high, with a crossover around $V_d=0.81$ in the controlled environment. Crossovers remain roughly 0.75-0.91 as causal-chain length varies from four to ten.
The paper does not establish 0.81 as a universal threshold for production agents. The synthetic environment is deliberately simplified. Its role is to demonstrate that targeting can become beneficial once the verifier independently exposes most of the causal chain, not to calibrate an industry-wide cutoff.
Reward-system design should start with observability
The business implication is primarily about sequencing investment.
| Decision | What the paper shows | Cognaptus inference | Boundary |
|---|---|---|---|
| Build finer turn-level targeting | Narrow targeted credit can underperform broad allocation at low $V_d$ | Measure verifier coverage before investing in localization | Learned process rewards may add genuinely new information |
| Improve the verifier | State-transition checking leaves reads and exploration largely invisible | Fund richer verification when it exposes previously unobserved prerequisites | More precise scoring of the same visible writes may not increase coverage |
| Compare reward schemes | Targeting and concentration can move together | Include a matched-concentration shuffled control | A targeting null is not proof of exact equivalence |
| Choose a baseline | Uniform redistribution performs strongly in tested low-$V_d$ regimes | Treat broad redistribution as a zero-information baseline when causal coverage is poor | This is not a claim that uniform credit is universally optimal |
For an agent-platform team, this means the relevant audit is not simply “Do we have per-turn rewards?” It is “Which causally necessary actions can our verifier independently judge?”
A workflow dominated by read-before-write dependencies may require better process supervision more than more elaborate redistribution of a terminal-state score.
The evidence is strongest for immediate updates
The cleanest identification in the paper concerns the first shared-rollout policy update. Competing arms start from the same policy, train on the same rollout batch with matched optimization conditions, and differ in how the same advantage is distributed.
That is strong evidence about immediate credit geometry. It is not equivalent to observing unrestricted multi-round RL, where updated policies begin collecting different trajectories and reward allocation becomes entangled with changing data.
The paper also does not show that learned process reward models are harmful. Oracle or learned supervision can expose prerequisite steps that state-transition verifiers cannot observe, increasing $k$ rather than merely redistributing credit among an already narrow set of turns.
Nor does the real-model evidence establish what happens in naturally high-$V_d$ operational systems. The targeting crossover is demonstrated cleanly in the synthetic environment, while the measured real benchmarks remain below that regime.
Measure coverage before optimizing precision
Credit assignment is often framed as a localization problem: determine which action deserves the reward.
For multi-turn tool agents, that framing can begin one step too late.
If the verifier sees only one or two actions in a workflow whose success depends on six, better localization within those visible actions does not recover the missing supervision. The first design question is how much of the causal chain the verifier can observe. Only after that does precision of targeting become interpretable.
The paper’s contribution is therefore less a new reward heuristic than a diagnostic order of operations: measure verifier information density, test coverage breadth, and then decide whether finer targeting is adding information or merely concentrating it.
Cognaptus: Automate the Present, Incubate the Future.
-
Chenyu Zhou and Qiliang Jiang and Shuning Wu and Xu Zhou (2026). Coverage, Not Targeting: A Structural Regime in Multi-Turn Agent Credit Assignment. arXiv:2609.02417. https://arxiv.org/abs/2609.02417 ↩︎