TL;DR for operators
In World Agent: Can Language Models Keep a World Running?1, the central problem is not whether an agent can produce the expected event. It is whether those events remain correctly organized as the world evolves.
Across the benchmark, models often complete more required events than they correctly satisfy the relationships among those events. Causal-relation checking is especially fragile: its cross-model mean falls from 0.683 in the Easy maintenance tier to 0.356 in Extreme, where substantially more of the world representation must be constructed by the agent.
For persistent AI products, this changes what should be tested. A simulation agent that reaches the requested state after violating prerequisites, executing events in the wrong order, or mishandling concurrency has not maintained the system correctly. WorldAgent-Benchmark therefore separates completion from relation correctness and separately tests prediction with explicit rules, without those rules, with state feedback, and across longer horizons.
The practical implication is to test three dependencies independently: does the agent preserve event relationships, how much scaffolding does it require, and how frequently does it need authoritative state feedback?
Reaching the right state is not enough
Consider an agent operating a persistent simulation. It completes a repair, opens a route, moves a resource, and eventually produces the expected final state. A conventional task-success metric might mark the run as successful.
But suppose the repair happened before its prerequisite inspection, two activities that should have been sequential were performed concurrently, or a state change failed to constrain an action several steps later. The final state may look plausible while the path that produced it violates the world’s rules.
That is the capability gap WorldAgent-Benchmark targets. Rather than evaluating generation alone, the benchmark asks whether a language model can continue operating an already delivered explicit-state world over multiple rounds.
The maintenance track evaluates a timed execution trace against contracts covering required events, structures, prohibitions, causal dependencies, temporal ordering, and concurrency. Crucially, completion and relations are scored separately. An event can therefore count as completed even when the relationships that should organize it are wrong.
This separation exposes failures that a single outcome score would hide.
Event completion and event organization diverge
On the nine-case common maintenance subset, reported completion scores range from 0.692 to 0.813 across model configurations. Relation scores range more widely, from 0.512 to 0.738.
The difference matters because the relation metric is not another measure of whether individual actions happened. It checks whether causal, temporal, concurrency, and chain-order requirements connecting those actions were satisfied.
Causal relations are the weakest component for nearly all models in Easy and for every model in Extreme. Their cross-model mean declines from 0.683 to 0.356 as the benchmark removes more pre-built structure.
The benchmark’s “causal” checks need careful interpretation. They do not identify causality statistically or through interventions. They verify authored requirements: specified causes must occur, and each must finish before the required effect begins.
That still captures an operationally consequential failure mode. A persistent system can execute many locally correct operations while losing the dependency structure that makes the overall process valid.
Removing scaffolding exposes another source of fragility
The maintenance cases keep the narrative and map fixed while varying how much world structure is supplied in advance.
In Easy, types, entities, and capabilities are pre-built. In Hard, capability construction is delegated to the agent. In Extreme, only the map remains pre-built; the agent must construct types, entities, and capabilities itself.
Across the common comparison, every evaluated model shows lower completion from Easy to Extreme, although the size of the decline differs.
This is more than a difficulty adjustment. It identifies how much observed performance depends on representation being prepared for the model.
An agent may appear reliable when it operates inside a carefully structured environment yet degrade once it must construct or maintain more of that structure itself. For a product team, those are different deployment assumptions.
Cognaptus inference: scaffolding should therefore be treated as a measurable system dependency. If production requires fixed schemas, validated entity definitions, capability registries, or other machine-readable structure, reliability tests should preserve those assumptions. If the product is expected to create such structures dynamically, evaluation should remove them deliberately rather than extrapolating from the scaffolded setting.
Explicit rules change what deduction measures
The deduction track isolates a related question: can a model predict how a deterministic world will evolve?
The benchmark compares an Open condition, where detailed machine-readable dynamics are available, with a Naming condition that withholds numeric rule details.
Under pure open-loop M1 deduction, the difference is large. Among stronger models, Open-condition field-level F1 is roughly 0.56–0.72. Naming-condition scores across all eight evaluated models fall to 0.03–0.15.
This result should not be read as a clean test of whether models have internally learned the world’s dynamics. The paper notes that Naming withholds parameters that cannot be recovered from true intermediate-state feedback in M1, so the condition mixes deduction with prior knowledge and guessing.
The cleaner operational conclusion is narrower: providing explicit machine-readable dynamics materially changes predictive performance.
That distinction matters for systems such as digital twins or operational simulators. If the production agent can query authoritative rules, testing it without those rules measures a different system.
Feedback helps tracking, not distant prediction
The benchmark also compares open-loop deduction with M2, where the model receives true state feedback after actions.
That closed loop recovers substantial action-interval prediction performance relative to the severely degraded M1 Naming condition. The result is consistent with a system that can correct its internal state estimate when authoritative observations arrive.
But the recovery does not extend cleanly to distant prediction. Every reported lookahead F1 value remains below 0.25, including the benchmark’s 48-time-unit horizon.
This produces a practical separation between tracking an evolving system with feedback and predicting that system far ahead without intermediate correction.
For persistent agents, those capabilities should not be assumed interchangeable. A workflow can be designed around frequent state refreshes even when long-range internal simulation remains unreliable.
What this changes for persistent AI evaluation
For games, simulations, digital twins, and multi-step operational agents, Cognaptus would translate the benchmark into three separate evaluation questions:
| Evaluation question | Benchmark contrast | Operational interpretation |
|---|---|---|
| Did the required outcomes occur? | Completion score | Measures local task realization |
| Did they occur through valid dependencies and timing? | Relation score | Detects invalid process execution hidden by successful outcomes |
| How much environment support is necessary? | Easy → Hard → Extreme | Reveals dependence on prepared representation and capabilities |
| Can the system operate without explicit dynamics? | Open → Naming | Tests reliance on machine-readable rules |
| Does observation compensate for imperfect internal prediction? | M1 → M2 | Tests whether closed-loop feedback stabilizes state tracking |
| Can the model project far ahead? | Lookahead | Separates feedback-assisted tracking from long-horizon prediction |
The broader design principle is straightforward: evaluate the system under the same information and feedback structure it will receive in production, then remove those supports deliberately to learn where reliability breaks.
The benchmark is auditable, but its boundary matters
WorldAgent-Benchmark uses 31 maintenance cases and six deduction case families. Eight distinct models are evaluated, but only three complete the full 31-case maintenance grid; the broader maintenance comparison relies on a preselected nine-case common subset.
Maintenance also relies partly on an LLM judge for semantic binding of events, structures, and prohibitions. The benchmark saves evidence so judgments can be audited and scores recomputed, while causal, temporal, and concurrency relations are checked programmatically. That improves traceability without guaranteeing that every semantic judgment is correct.
Deduction has a stronger deterministic reference mechanism: reference rollouts are executed programmatically and predictions are compared field by field. Even there, deterministic replay validates tested trajectories rather than proving that every possible world branch perfectly implements the authored specification.
Most importantly, these are synthetic explicit-state worlds. The benchmark demonstrates a measurement problem and documents current model behavior inside that setting. It does not establish that the same failure rates will appear in open-ended business environments.
Persistent agents need process-validity metrics
WorldAgent-Benchmark shifts attention from whether an agent can produce a plausible outcome to whether it can preserve the structure that makes successive outcomes valid.
That distinction becomes more consequential as systems remain active longer. A one-step mistake can become persistent state; an invalid ordering can alter later possibilities; missing feedback can compound prediction error.
For operators, the resulting evaluation agenda is more demanding than task success but also more diagnostic: measure completion, validate relationships, vary scaffolding, expose or hide rules intentionally, and test how much state feedback the agent needs.
An agent that reaches the expected destination through an invalid world trajectory has passed the wrong test.
Cognaptus: Automate the Present, Incubate the Future.
-
Weixing Chen and Weipeng Zhang and Nan An and Yang Liu and Liang Lin (2026). World Agent: Can Language Models Keep a World Running?. arXiv:2609.32692. https://arxiv.org/abs/2609.32692 ↩︎