TL;DR for operators

After a GUI agent sends a message, edits a document, moves a file, changes a setting, or searches for information, someone—or something—has to decide whether the instruction was actually completed correctly.

It is tempting to treat that decision as mainly a model-capability problem: use a stronger vision-language model, show it more screenshots, or improve the prompt. The evidence in Task-Adaptive Rubrics for GUI Reward Modeling suggests another failure point comes earlier. The verifier first needs an adequately specified definition of success.1

AdaptRubric separates that definition into two layers. A reusable task-family rubric establishes stable checks such as destination, recipient, scope, final state, or completion. A second stage can add at most two checks grounded in the particular instruction. The resulting criterion is then given to an otherwise unchanged verifier.

The distinction is measurable. In the paper’s Qwen3.5-122B-A10B ablation, the full coarse-plus-fine setup reaches 86.6 F1. Removing either rubric reduces F1 to 84.6; removing both lowers it to 80.7. The same design also performs better in the paper’s controlled online-RL experiment and its reward-guided test-time trajectory-selection study.

For deployment teams, the useful design principle is to make success-definition an explicit, auditable layer before success-judgment. The evidence does not show that the paper’s eight-category rubric bank will generalize unchanged to every GUI product, but it does show that leaving evaluation criteria implicit can materially weaken the reward signal.

A verifier can see the trajectory and still judge the wrong thing

Consider a GUI agent asked to move a file into a particular directory, send a message to a named recipient, or change only one field in a form.

The screenshots may clearly show what happened. Yet judging success still requires knowing what counts as completion. Was the destination correct? Was the original removed or copied? Was the message sent to the right recipient through the requested channel? Did the agent modify only the requested field?

These requirements are not interchangeable across tasks.

This is where AdaptRubric changes the usual framing. Rather than treating outcome reward modeling as simply:

instruction + trajectory → success or failure

the paper makes the judging criterion an explicit input:

$$ \hat{r}=f_{\theta}(\mathcal{I},\tau^{\prime},\mathcal{R})\in\{0,1\} $$

Here, $\mathcal{R}$ is the separately constructed task-adaptive criterion. The verifier still examines the user instruction and selected trajectory evidence, but it no longer has to reconstruct all success conditions implicitly while making the final judgment.

That decomposition matters because errors in the criterion can survive even when perception is adequate.

AdaptRubric separates reusable rules from instruction-specific checks

The method first routes an instruction into one of eight GUI task families, including information queries, creation or modification, deletion, communication, transfer, state changes, composite workflows, and a general fallback.

Each family has a reusable coarse rubric. Its job is to preserve verification boundaries that tend to recur within that class of task.

A communication task, for example, may require attention to source, recipient, channel, and confirmation. A transfer task instead emphasizes source, action, destination, and fidelity. A deletion task needs correct target and scope.

That still does not capture every instruction.

An individual request can specify an exact filename, formatting constraint, value, location, sequence, or other condition that a category-level checklist cannot know in advance. AdaptRubric therefore generates a fine rubric containing no more than two instruction-grounded cues. It can also abstain if no additional cue is justified.

The final criterion preserves the coarse rubric and appends these narrowly scoped additions rather than rewriting the task-family rules.

This is a useful architectural choice: instance adaptation is constrained to explicit instruction content instead of becoming an invitation for the verifier to invent additional requirements.

The ablation shows that both levels are doing work

The clearest evidence for the mechanism comes from the component ablation.

Configuration Accuracy F1 Likely role of test
Coarse + fine 86.9 86.6 Full method
Without fine 84.9 84.6 Tests instance-specific adaptation
Without coarse 85.2 84.6 Tests task-family verification boundaries
Without both 81.5 80.7 Image-only judging baseline within the ablation

Removing either rubric costs 2.0 F1 points. Removing both costs 5.9.

This does not prove that every GUI verifier needs exactly this taxonomy or a two-item fine rubric. It does support the narrower interpretation that reusable category rules and instruction-specific constraints are complementary in the tested setup.

The result also addresses a common alternative explanation: perhaps AdaptRubric simply benefits from seeing more trajectory evidence.

The paper explicitly controls for this in its main offline comparison. Under a matched ten-screenshot budget on OGRBench, AdaptRubric averages 86.7% accuracy and 86.6 F1 across eight judge backbones. Its F1 is 3.6 points above the average of ZeroGUI and OS-Themis.

An appendix experiment further shows why that control is necessary. Increasing ZeroGUI from its last two screenshots to its last ten raises average F1 from 76.8 to 84.9. Image budget clearly matters. The main AdaptRubric comparison therefore becomes more informative because the methods are evaluated with the same ten-screenshot allowance.

Better criteria matter because rewards are used for decisions

Reward classification would be less consequential if it ended as a benchmark score. In agent systems, reward judgments often feed subsequent decisions.

The paper tests two such cases.

First, it holds the MAI-UI-8B policy and ClawGUI-RL/GRPO training protocol fixed while changing the reward agent. The no-external-reward result is 19.70% task success. Training with AdaptRubric reaches 23.93%, an absolute improvement of 4.23 percentage points and the highest result among the compared reward agents in that experiment.

That is evidence within one controlled training configuration, not a general claim that task-adaptive rubrics improve every GUI policy. Still, it demonstrates the relevant pathway: a different success judgment can change which trajectories generate reinforcement signals, and those signals can affect the resulting policy.

Second, the paper tests reward-guided selection from a fixed pool of ten AndroidWorld candidate trajectories per task. On 113 tasks, AdaptRubric improves task success over random selection by 11.88 points for EarlyStop@7 and 13.28 points for BestOfN@8.

Here the verifier is not training the agent. It is deciding which completed trajectory to trust.

The common dependency is the same: reward quality becomes operational once another system acts on the judgment.

For deployment, make success criteria a governed component

Cognaptus’ inference from these results is organizational as much as technical.

A production GUI-agent stack can separate three responsibilities:

  1. Define stable verification boundaries. Maintain reusable rules for recurring task families or product workflows.
  2. Add only instruction-grounded exceptions or specifics. Extract values, scopes, destinations, placements, ordering constraints, or formats that are actually stated by the user.
  3. Judge evidence against the resulting criterion. Let the verifier focus on whether the observed trajectory satisfies the defined requirements.

This makes the reward layer easier to inspect than a system in which all criteria remain latent inside a VLM prompt or reasoning process.

It also creates a practical QA surface. Teams can review whether a category rubric is wrong, whether an instruction-specific cue was unsupported, whether trajectory evidence was insufficient, or whether the verifier misapplied an otherwise sound criterion. Those are different failure modes and can require different fixes.

The paper’s efficiency results also suggest that this decomposition need not require the most expensive verifier pipeline. Using Qwen3-VL-8B-Instruct, AdaptRubric records 84.2 F1 with 4,220 LLM calls and 26.98 million tokens. OS-Themis and ZeroGUI consume substantially more calls or tokens in the reported comparison, while DigiRL is cheaper but lower-performing.

The relevant deployment question is therefore not verifier accuracy in isolation. It is whether the additional criterion-construction work improves the decisions downstream enough to justify its inference cost.

The fixed rubric bank is also the main generalization boundary

The coarse rubric bank is built offline and remains fixed during evaluation. That improves reproducibility and makes its logic inspectable, but it also constrains what the results establish.

The experiments span mobile, desktop, and web environments, yet they cannot cover every application, interface design, enterprise workflow, or instruction distribution. A taxonomy that captures messaging, file transfer, deletion, navigation, and composite workflows in these benchmarks may need extension in a specialized product.

The test-time-scaling experiment also uses a fixed candidate pool rather than a live execution environment. Its results support better ranking and stopping within that setup, not every possible dynamic rollout strategy.

For an operator, then, AdaptRubric is better read as evidence for an architectural separation than as a universal rubric library.

Define success explicitly. Preserve reusable verification boundaries where they exist. Add instruction-specific requirements only when the instruction supports them. Then ask the judge whether the evidence satisfies that definition.

The paper’s central contribution is showing that this intermediate layer is not administrative decoration. Within the tested settings, weakening it measurably weakens the reward.

Cognaptus: Automate the Present, Incubate the Future.


  1. Tao Xiong and Xavier Hu and Wenkai Wang and Qinzhuo Wu and Changqiao Wu and Pengzhi Gao and Wei Liu and Jian Luan and Shengyu Zhang (2026). Task-Adaptive Rubrics for GUI Reward Modeling. arXiv:2608.24174. https://arxiv.org/abs/2608.24174 ↩︎