Measure the Chain Before You Target the Reward
TL;DR for operators A tool agent may need to read records, retrieve identifiers, inspect intermediate state, and only then issue the write that completes a task. If the verifier scores only that final write, a seemingly precise per-turn reward can assign learning credit to one visible action while leaving the prerequisite chain unsupervised. ...