Cover image

When One Patch Rules Them All: Teaching MLLMs to See What Isn’t There

Image security has an awkward habit of sounding theoretical until the image is inside a business workflow. A product team adds an image-upload feature. A compliance team uses multimodal models to inspect screenshots. A support bot reads photos from customers. A research assistant summarizes figures from PDFs. Everyone understands that the model may occasionally misread an image. That is ordinary error. Annoying, but ordinary. ...

February 3, 2026 · 15 min · Zelina
Cover image

Same Moves, Different Minds: Rashomon Comes to Sequential Decision-Making

A taxi is a useful little trap. It looks harmless: pick up passengers, drive them to destinations, do not run out of fuel. A small grid-world taxi environment is not exactly the sort of thing that makes executives whisper “agentic transformation” over terrible conference coffee. But that is precisely why it works. Strip away the enterprise theatre, and sequential decision-making becomes easier to see. An agent observes a state, chooses an action, receives the next state, and repeats. If two agents always make the same moves and achieve the same objective, most organizations would treat them as equivalent. Same behavior, same operational meaning. Audit passed. Ship it. ...

December 22, 2025 · 18 min · Zelina
Cover image

How to Make Neural Networks Talk: Register Automata as Their Unexpected Interpreters

How to Make Neural Networks Talk: Register Automata as Their Unexpected Interpreters Prices move. Sensors drift. Users click, pause, return, disappear, and sometimes behave exactly like a Markov chain with a caffeine problem. Modern sequence models are good at turning such streams into decisions. A recurrent network or transformer can look at a run of numbers and say: buy, flag, reject, approve, alert. What it usually cannot do is explain the rule it has learned in a form that a risk team, engineer, or auditor can actually inspect. ...

November 25, 2025 · 18 min · Zelina
Cover image

Pop-Ups, Pitfalls, and Planning: Why GUI Agents Break in the Real World

Pop-up. That tiny word hides a surprisingly large operational problem. A human sees a battery warning, an update prompt, a permission dialog, or a frozen app and does something boringly competent: dismiss it, recover context, re-check the screen, and continue. A GUI agent, meanwhile, may confidently continue a plan that no longer matches reality. The machine has not “failed” in the theatrical sense. It has simply treated a live workflow like a polite screenshot sequence. Very enterprise. Very doomed. ...

November 22, 2025 · 13 min · Zelina
Cover image

Don’t Self-Sabotage Me Now: Rational Policy Gradients for Sane Multi-Agent Learning

Kitchen work is not hard because chopping onions is metaphysically difficult. It is hard because two people must agree, implicitly and quickly, who gets the onion, who holds the plate, who waits by the pot, and who moves out of the corridor before everyone performs a small culinary traffic accident. That is why Overcooked remains such a useful multi-agent benchmark. It turns coordination into something visible. Agents do not merely need to “perform a task”; they need to infer what another agent is about to do and avoid becoming a sentient obstacle. ...

November 13, 2025 · 14 min · Zelina
Cover image

Noisy but Wise: How Simple Noise Injection Beats Shortcut Learning in Medical AI

X-rays look clinical. To a neural network, they can also look like stationery. A hospital name in the corner. A scanner signature. A compression pattern. A familiar positioning marker. A slightly different way of cropping the lung field. None of these is pneumonia. None of these is COVID-19. Yet a deep learning model trained on small medical datasets can treat them as wonderfully convenient diagnostic evidence, because machines are very good at passing exams and less naturally committed to understanding what the exam is about. ...

November 9, 2025 · 15 min · Zelina