Cover image

Reward the Right Thing: GUI Agents Need Better Success Criteria, Not Just Better Judges

TL;DR for operators After a GUI agent sends a message, edits a document, moves a file, changes a setting, or searches for information, someone—or something—has to decide whether the instruction was actually completed correctly. It is tempting to treat that decision as mainly a model-capability problem: use a stronger vision-language model, show it more screenshots, or improve the prompt. The evidence in Task-Adaptive Rubrics for GUI Reward Modeling suggests another failure point comes earlier. The verifier first needs an adequately specified definition of success.1 ...

September 29, 2026 · 7 min · Zelina