Cover image

The Navigator Is Not the Motor: What VISOR Gets Right About Embodied AI Architecture

TL;DR for operators A familiar robotics design question is where to draw the boundary between understanding and control: should one learned system interpret the scene, choose where to go next, and explain that choice, or should those functions remain distributed across separate perception and planning modules? VISOR1 tests a middle position. Its 3B model combines language understanding, visual-spatial reasoning, target recognition, and high-level destination selection, while a conventional planner handles the low-level movement. It is not the strongest navigator on the leaderboard, but on OVON its post-trained model records a 21.70% success rate on Val Seen and 22.00% on Val Unseen, an unusually small seen-to-unseen change even though stronger baselines achieve much higher absolute success rates. ...

September 8, 2026 · 9 min · Zelina