Cover image

Look Again Before You Answer: Visual RAG Needs a Search Policy

TL;DR for operators A visual support assistant shown an unfamiliar machine, product, bird, or venue cannot answer by retrieval alone. It must first determine what the image depicts, then locate the missing fact, while deciding whether another search is worth the delay. A wrong first match can redirect every later step toward the wrong entity. ...

July 23, 2026 · 8 min · Zelina
Cover image

CAPTION THIS: Why Multimodal RAG Is Finally Growing Up

Captioning looks easy until the caption has to be true. A consumer image captioning model can say, “a man standing at a podium,” and most people will nod. A newsroom cannot stop there. It needs to know whether the man is a senator, a witness, a CEO, a defendant, or simply someone unlucky enough to stand near a microphone. It may need the committee name, the location, the event, the year, the organization behind the banner, and the person half-visible at the edge of the frame. Journalism, as usual, ruins the demo. ...

November 30, 2025 · 18 min · Zelina