Cover image

Let the Model Design the Poster—Not the Evidence

TL;DR for operators A scientific poster can look polished and still fabricate the plots or diagrams readers interpret as evidence. PosterHarness separates those responsibilities: the image model designs the layout and decides where evidence should appear, but it must leave those regions blank for source-paper figures to be inserted later by deterministic code. ...

August 3, 2026 · 8 min · Zelina
Cover image

Skill Issue or System Design? How LLMs Actually Follow Instructions

The checklist problem that exposes the model Checklist tasks look boring. That is exactly why they are useful. Ask an LLM to write a formal email under 50 words, include one required term, avoid another term, and return the result as JSON. None of this sounds intellectually difficult. No theorem proving. No multimodal reasoning. No dramatic benchmark leaderboard screenshot. Just instructions. ...

April 8, 2026 · 18 min · Zelina
Cover image

Replay the Losses, Win the Game: When Failed Instructions Become Your Best Training Data

Failure logs are usually treated as evidence for the prosecution. A model is asked to produce a concise compliance summary with three bullet points, mention two risks, avoid prohibited claims, and end with a recommendation. It produces three bullets, correctly identifies the risks, avoids the prohibited claims—and forgets the recommendation. Under a strict binary reward, the response receives a zero. Under a partial-credit reward, it might receive 0.75. The first signal says nothing useful happened. The second says something useful happened, but not precisely what. ...

December 30, 2025 · 18 min · Zelina