Cover image

Hard Problems Pay Better: Why Difficulty-Aware DPO Fixes Multimodal Hallucinations

Training data has a bad habit: the easiest examples talk the loudest. Anyone who has trained a model on preference pairs knows the scene. One answer is clearly grounded in the image; the other confidently invents an object, a color, or an action that is not there. The model learns the contrast quickly. Everyone applauds. The loss goes down. The dashboard looks obedient. ...

January 5, 2026 · 15 min · Zelina
Cover image

When Models Teach Themselves: Inside the Rise of SuperIntelliAgent

Image generators fail in very ordinary ways. A prompt asks for a green banana and a blue vase. The model gives you something banana-adjacent, vase-adjacent, and chromatically negotiable. A designer asks for a bowl containing a pizza. The model places the pizza beside the bowl, halfway inside the bowl, or in a bowl-like universe where geometry has apparently resigned. A product team then does the usual dance: collect bad outputs, ask users what they preferred, curate examples, fine-tune later, and call the whole thing “continuous improvement” because the spreadsheet had a date column. ...

December 1, 2025 · 16 min · Zelina
Cover image

Value Collision Course: When LLM Alignment Plays Favorites

A support chatbot does not wake up one morning with a worldview. It gets one, slowly, through the dull machinery of product decisions: who labels the data, how many options they can choose from, whether disagreement is kept or ironed flat, and which optimization method gets the privilege of turning messy human judgement into model behaviour. ...

November 20, 2025 · 14 min · Zelina
Cover image

The Sentiment Edge: How FinDPO Trains LLMs to Think Like Traders

TL;DR for operators News is only useful when it survives the journey from headline to position sizing. FinDPO, proposed by Giorgos Iacovides, Wuyang Zhou, and Danilo Mandic, is a finance-specific Llama-3-8B-Instruct sentiment model trained with Direct Preference Optimization rather than ordinary supervised fine-tuning.1 The paper’s headline result is not merely that FinDPO scores well on sentiment benchmarks. Plenty of models win benchmarks, then politely disappear when transaction costs arrive. ...

July 27, 2025 · 14 min · Zelina
Cover image

Delta Force: How Weak Models are Secretly the Best Teachers

TL;DR for operators Training budget is usually where elegant AI strategy goes to die. The paper behind this article argues that preference tuning does not always need a superior teacher response. It may only need a useful contrast. A model can improve by learning that one weak answer is better than an even weaker one, even when neither answer is as good as what the model can already produce.1 ...

July 9, 2025 · 17 min · Zelina
Cover image

Guardians of the Chain: How Smart-LLaMA-DPO Turns Code into Clarity

TL;DR for operators Smart-LLaMA-DPO is not interesting because it puts another LLM badge on smart contract auditing. We have enough badges. It is interesting because it shows a credible mechanism for making an LLM behave more like a useful junior security analyst: read the contract, identify whether the vulnerability is real, locate the issue, and explain the reasoning in a way a developer can act on. ...

June 24, 2025 · 16 min · Zelina