Cover image

When Alignment Is Not Enough: Reading Between the Lines of Modern LLM Safety

A practical reading of modern LLM safety research, showing why alignment should be treated as an operational control system rather than a one-time model property.

January 26, 2026 · 15 min · Zelina
Cover image

When Models Listen but Stop Thinking: Teaching Audio Models to Reason Like They Read

CORD shows that audio-language models may fail not because they cannot hear, but because their audio-conditioned reasoning drifts away from their own text pathway.

January 26, 2026 · 16 min · Zelina
Cover image

When SGD Remembers: The Hidden Memory Inside Training Dynamics

A mechanism-first reading of how process-tensor diagnostics turn SGD memory from training folklore into something measurable, testable, and operationally useful.

January 26, 2026 · 15 min · Zelina
Cover image

When Trains Meet Snowstorms: Turning Weather Chaos into Predictable Rail Operations

A new Finnish railway-delay dataset shows that predictive rail AI begins with spatial-temporal data engineering, not with a glamorous model leaderboard.

January 26, 2026 · 20 min · Zelina
Cover image

Gated Sparse Attention: Speed Without the Sink

A mechanism-first reading of Gated Sparse Attention, showing how sparsity, gating, and adaptive token selection jointly target long-context cost, attention sinks, and training instability.

January 24, 2026 · 17 min · Zelina
Cover image

Learning to Discover at Test Time: When Search Learns Back

A mechanism-first reading of TTT-Discover, where test-time search becomes test-time learning for verifiable discovery problems.

January 24, 2026 · 18 min · Zelina
Cover image

PyraTok: When Video Tokens Finally Learn to Speak Human

A mechanism-first reading of PyraTok, showing why language-aligned multi-scale video tokenization matters for generation, understanding, and enterprise video AI.

January 24, 2026 · 15 min · Zelina
Cover image

Training Models to Explain Themselves: Counterfactuals as a First-Class Objective

A mechanism-first reading of counterfactual training: why better recourse may require changing the model, not just improving the explanation generator.

January 24, 2026 · 16 min · Zelina
Cover image

Triage by Token: When Context Clues Quietly Override Clinical Judgment

How proxy-variable testing exposes a quiet failure mode in LLM-based emergency triage: models can change acuity judgments when non-clinical context enters the prompt.

January 24, 2026 · 13 min · Zelina
Cover image

When LLMs Get a Laptop: Why Sandboxes Might Be the Real AGI Benchmark

A mechanism-first reading of LLM-in-Sandbox, showing why giving models a minimal computer environment may matter more than adding another clever prompt.

January 24, 2026 · 16 min · Zelina