TL;DR for operators
A world model can be reasonably accurate when predicting the next state and still become unreliable once its own predictions are fed back repeatedly. That difference is operationally relevant because planning and control systems often depend on trajectories, not isolated one-step estimates.
Semigroup-JEPA: Latent Dynamics Consistency for Zero-Shot Physics Generalization1 tests this problem by training models on a narrow range of gravity values and evaluating them across much wider regimes. SG-JEPA jointly trains its visual representation and temporal predictor using several recursively generated future latents. Its largest advantages generally appear where errors have more opportunity to compound: longer forecast horizons and gravity values outside the training distribution.
The paper therefore supports a specific development principle. If a representation will support recursive prediction or downstream control, one-step validation error is an incomplete selection criterion. Evaluate what happens after repeated composition, and resolve results across the operating parameter being shifted. The evidence supports that principle inside controlled MuJoCo experiments; it does not establish general real-world robustness.
Small local errors can become large rollout differences
The paper’s most consequential observation is not simply that SG-JEPA produces lower benchmark error. It is that the size of a model difference depends strongly on how the model is evaluated.
Under teacher forcing, the predictor repeatedly receives representations derived from real observations. Under free rollout, its own predicted latent becomes part of the input for subsequent predictions. The latter creates an error-propagation problem.
On the square task, matched fresh-predictor experiments find lower teacher-forced local-error point estimates for SG-JEPA across all 25 tested gravity values, with roughly 32% lower local error than DINO-WM at far-out-of-distribution gravities. But once predictions are recursively fed back, the gap becomes several times larger from the second forecast step onward and peaks around horizon 20 before narrowing later.
The mechanism is straightforward: a local defect is not merely added to future errors. It is transformed by every later learned transition. The paper formalizes this as
Two representations with similar next-step errors can therefore diverge substantially after repeated application of their learned dynamics.
That changes what should be measured. For a system that will recursively simulate, plan, or supply features to a controller, short-horizon error can conceal instability that becomes visible only along a trajectory.
SG-JEPA trains for repeated composition
SG-JEPA extends LeWorldModel by making two changes central to the experiment: gravity is supplied explicitly to the dynamics model, and training optimizes several autoregressively generated latent predictions rather than only a teacher-forced next step.
Its rollout loss is
The main configuration uses 20 history frames, five recursively predicted training steps, and a discount factor of 0.95. The encoder and predictor are trained jointly, with SIGReg regularization on the latent representation.
The distinction matters because multi-step training changes which representation errors are costly. A feature that looks adequate for the immediate next state may be poor if discarding it creates defects that repeatedly propagate through later transitions.
The headline prediction results are consistent with that objective. On the square task at 44 forecast steps, SG-JEPA GRU records position error of 1.034 m versus 2.006 m for DINO-WM, velocity error of 1.915 m/s versus 3.082 m/s, and cumulative-rotation error of 0.236 turns versus 0.341. On the 3D Approach Ball task, mean position error falls from 0.0705 m with DINO-WM to 0.0491 m with SG-JEPA GRU.
The advantage is not uniform. DINO-WM remains stronger on triangle cumulative rotation and on final-step Approach Ball velocity. Those reversals matter because they show that the result is about error behavior under particular dynamics and horizons, not general dominance of one representation.
The representation carries much of the advantage
A natural explanation would be that SG-JEPA’s recurrent predictor is simply better at long rollouts. The paper tests that possibility by freezing encoders and training fresh temporal predictors.
The ranking largely survives. With either a new GRU or a new Transformer predictor, the encoder originally learned through GRU-based SG-JEPA training produces about 12% lower mean rollout error than the encoder learned with the Transformer-based training setup.
This is a mechanism experiment rather than a second benchmark claim. It localizes part of the observed advantage to what the training process caused the representation to retain.
The paper calls one relevant property predictive closure. In plain terms, a representation is better closed when it preserves enough of the underlying physical state that the next relevant state can be predicted without depending on information the encoder threw away.
Its theoretical decomposition separates local error into two sources:
The first term captures information discarded by the representation that still matters for future dynamics. The second captures mistakes in modeling the dynamics of the information that remains.
For representation selection, that distinction is valuable. Improving the temporal predictor cannot fully compensate for a latent state that has already removed variables required for stable future prediction.
Supplying gravity is not the same as learning transferable physics
The phrase “zero-shot physics generalization” needs a narrow reading here. The model does not infer gravity from video, and the experiments do not show discovery of an unknown physical law. Gravity is given to the predictor as an explicit input.
Even with that information, extrapolation is not guaranteed. The paper’s linear analysis gives an unseen-gravity error bound of the form
The operational interpretation is more useful than the notation. Transfer depends on three things: whether training covered enough variation in the governing parameter, whether the predictor learned the retained dynamics accurately, and whether the representation kept the physical information needed for prediction.
Giving a model the correct environment parameter supplies context. It does not guarantee that the learned state supports reliable behavior at a parameter value far from training.
The control experiments test representation reuse
The three robotic tasks provide a second reason to care about representation quality. Their downstream diffusion policies consume frozen visual features, but they do not query SG-JEPA’s temporal predictor during online control.
On Arm Catcher Ball, average capture success rises from 9.5% with DINO-WM features to 23.3% with SG-JEPA GRU features. On Arm Paddle Ball, success rises from 17.7% to 23.8%. On Franka Paddle-to-Basket, strict basket entry increases from 27.6% for DINO-WM to 30.36% for SG-JEPA GRU, although DINO-WM achieves higher paddle-contact rates.
These results do not demonstrate that an SG-JEPA rollout engine should directly control a robot. They indicate that training a representation around multi-step dynamics can change the information available to a separately trained controller.
That is a more portable architectural lesson: the value of a world-model objective may survive even when its transition model is absent from the production inference path.
Stress-test the operating regime, not just the validation set
For robotics and embodied-AI teams, Cognaptus infers a concrete evaluation change from these results: model selection should include recursive rollouts across the physical parameters expected to vary in deployment.
Aggregate validation error can hide two different risks. Horizon aggregation can hide errors that grow only after repeated composition, while parameter aggregation can hide sharp reversals at particular operating conditions. The paper itself shows both.
Its sparse-gravity adaptation experiment points to a related maintenance strategy. Post-training on four new gravity values improves prediction at unseen intermediate values, and mixing those target examples with replay from the original source distribution produces an average 14.05% improvement versus 6.85% for target-only adaptation across the tested model-shape combinations. This is exploratory adaptation evidence, not proof of a universal replay recipe, but it suggests a lower-cost response to a shifted operating regime than rebuilding the representation from scratch.
The evidence supports an evaluation principle, not a robustness guarantee
The study controls gravity unusually well, but that control also defines its boundary. Gravity is the principal changing physical parameter, it is known to the model, and all headline experiments occur in MuJoCo. Transfer across geometry is uneven, particularly for rotation. Several world-model comparisons use one training seed, so large evaluation cohorts do not measure variation from retraining the underlying models.
The theory is also explanatory rather than literal: its linear feature dynamics simplify neural representations, recurrent predictors, collisions, and contact-driven regime changes.
What survives those boundaries is narrower and more actionable. When a learned representation will participate in repeated dynamics, test the repeated dynamics. When deployment conditions can move along known physical coordinates, resolve performance along those coordinates. And when two models look similar one step ahead, do not assume they will remain similar after their errors have been composed dozens of times.
Cognaptus: Automate the Present, Incubate the Future.
-
Andy Zeyi Liu and Haoran Sun and Lucas Baker and Randall Balestriero and John Sous (2026). Semigroup-JEPA: Latent Dynamics Consistency for Zero-Shot Physics Generalization. arXiv:2609.10464. https://arxiv.org/abs/2609.10464 ↩︎