Cover image

Left of Whom? Spatial Agents Need More Than an Explicit Viewpoint

TL;DR for operators An embodied assistant may already have seen a room and still fail when asked to place something “to the left of the chair” from a person’s point of view. Supplying more information about where that person is helps less than many teams might expect, because identifying the observer is only one part of the reference-frame problem. ...

September 27, 2026 · 8 min · Zelina
Cover image

Control the Caption by Training What to Omit

TL;DR for operators A generative system can produce a fluent, factually plausible output and still fail because it focuses on the wrong information. Controllable Image Captioning with Prompt-Conditioned Scene Rewards by Jongyeop Hyun, Taeyoung Kim, and Hyounghun Kim1 tests a stricter approach: train the model not only to reward requested content, but also to penalize complementary content that falls outside the requested focus. ...

September 26, 2026 · 7 min · Zelina
Cover image

Reward the Claim, Not Just the Answer: What V-Rubrics Changes in Multimodal RL

TL;DR for operators If a multimodal model reaches the correct final answer after misreading part of an image or making an invalid intermediate inference, a binary success reward gives the optimizer little information about what should actually be reinforced. The paper studied here turns those failure classes into separate training signals and, where possible, assigns their reinforcement-learning credit only to the relevant portion of the response. ...

September 21, 2026 · 7 min · Zelina
Cover image

Same VQA Score, Different Eyes: What Fine-Grained Tests Reveal About VLMs

TL;DR for operators A vision-language model can rank well on broad question-answering benchmarks and still be materially weaker at distinguishing visually similar objects. In Understanding the Fine-Grained Knowledge Capabilities of Vision-Language Models, Ghosh, Zhang, and Schmidt show that CogVLM-Chat and LLaVA-NeXT-Vicuna-13B both score around 48% on aggregated general VQA, yet reach 77.3% and 58.1% respectively on their fine-grained evaluation—a gap of roughly 19 percentage points.1 ...

September 14, 2026 · 8 min · Zelina
Cover image

When Vision Fails in Both Directions

TL;DR for operators A model can correctly recognize that one object is left of another and still fail when asked to generate that same relationship. More importantly, some visual weaknesses recur in both directions. AMVICC maps visual-language understanding and image generation onto visual concepts derived from the same underlying benchmark. Across the tested systems, weaknesses repeatedly appear in Quantity and Count, Positional and Relational Context, Orientation and Direction, and State and Condition. Text behaves differently: most tested multimodal language models avoid the paper’s Text failure threshold, while all three tested image generators fall below it for explicit generation. ...

September 13, 2026 · 7 min · Zelina
Cover image

Plan, Predict, Then Move: What World Action Planner Changes About Robot Generalization

TL;DR for operators A robot action can look reasonable from the current camera image and still fail because its physical consequence is wrong, its coordinates are slightly off, or familiar skills must be composed in an unfamiliar order. World Action Planner (WAP)1 addresses that gap by putting prediction and search between action proposal and execution: a vision-language model proposes an action, a world model imagines its consequences, semantic feedback can revise the proposal, and local search resolves finer manipulation choices. ...

September 7, 2026 · 8 min · Zelina
Cover image

The Model Saw Every Scene. The System Had to Remember the Story.

TL;DR for operators A scene-by-scene assistant can describe every clip fluently while quietly forgetting who the characters are, how they relate, or why an earlier event matters now. StoryTeller addresses that continuity problem without task-specific training by keeping a persistent record of recurring characters and carrying forward only narrative facts that have been checked against the video.1 ...

August 2, 2026 · 8 min · Zelina
Cover image

Same Answer, Different Risk: Visual Semantic Entropy for VLM Review Routing

TL;DR for operators A visual assistant gives the same confident answer several times. That consistency may seem sufficient for automatic acceptance, but it shows only that repeated decoding produced the same response—not that the underlying visual interpretation is stable. Variability can also come from the wrong place. When both the image and question wording are changed, paraphrase choice may drive the resulting answer clusters more strongly than the visual changes. In the paper’s joint-perturbation analysis, text purity exceeds image purity for every reported model and split. A high uncertainty score may therefore indicate prompt sensitivity rather than visual ambiguity. ...

July 26, 2026 · 10 min · Zelina
Cover image

Picture This: When AI Reasoning Leaves the Text Box

Reasoning usually arrives as text. A model explains itself in sentences, equations, bullet points, and the occasional theatrical “therefore.” We have learned to call this chain-of-thought, or CoT, because “the model wrote a long scratchpad and we hope it helped” sounded insufficiently scientific. The paper Optical Reasoning: Rethinking Images as an Expressive Reasoning Medium Beyond Text asks a sharper question: what if the intermediate reasoning medium does not have to be text at all?1 ...

June 9, 2026 · 17 min · Zelina
Cover image

Pixels to Purchase Orders: A Business Map for Choosing Vision-Language Models

Pixels to Purchase Orders: A Business Map for Choosing Vision-Language Models Receipts are a good way to ruin an AI demo. A clean product photo is polite. A scanned receipt is not. It has shadows, folds, strange fonts, tiny numbers, merchant abbreviations, table-like structure, and one suspiciously important total amount hiding near the bottom. Ask a generic multimodal assistant what it sees, and it may produce an answer that sounds fluent enough to make everyone in the meeting relax. That is usually the dangerous part. ...

June 8, 2026 · 19 min · Zelina