Longer Yet Dumber: Why LLMs Fail at Catching Their Own Coding Mistakes
FPBench shows that many LLM code assistants can detect faulty requirements when prompted, but often fail to question bad premises on their own.
FPBench shows that many LLM code assistants can detect faulty requirements when prompted, but often fail to question bad premises on their own.
A mechanism-first reading of OpenAI’s malicious fine-tuning study and what it implies for evaluating open-weight model releases.
Multimodal chain-of-thought looks impressive until visual evidence must be used repeatedly, not merely mentioned.
A mechanism-first reading of SQLM, a self-play post-training method where language models generate their own questions, solve them, and learn from proxy rewards without curated training data.
AI shopping agents may reduce search friction, but this paper shows they also create model-specific demand shocks, position bias, and a new market for AI-facing listing optimisation.
CAPO shows how stronger verifier models can turn blunt outcome rewards into step-localised training signals for smaller reasoning models.
FinKario shows that financial AI performance may depend less on bigger language models and more on converting research reports into dynamic, event-aware retrieval infrastructure.
Cupid shows that LLM personalization fails less because models cannot write tailored answers and more because they retrieve, infer, and apply the wrong contextual preference.
A practical reading of multimodal hallucination research: why visual models appear grounded, where that grounding fails, and how operators should diagnose before they deploy.
A practical reading of Multi-Band Variable-Lag Granger Causality and what frequency-specific delays can reveal in complex time series.