Cover image

Seeing Isn’t Knowing: Why Vision-Language Models Still Miss the Details

A case-first reading of FROW, a benchmark showing why multimodal AI must recognize the exact object before it can reason safely about it.

December 14, 2025 · 16 min · Zelina
Cover image

Sound Zones Without the Handcuffs: Teaching Neural Networks to Bend Acoustic Space

A mechanism-first reading of how Neural PSZ uses masked microphone grids and monitor-point learning to make personal sound zones less dependent on rigid calibration geometry.

December 14, 2025 · 16 min · Zelina
Cover image

Tunnel Vision, Literally: When Cropping Makes Multimodal Models Blind

A mechanism-first reading of Visual Funnel, a training-free method showing that multimodal models need structured intermediate context—not just tighter crops—to read visual details correctly.

December 14, 2025 · 18 min · Zelina
Cover image

When Agents Loop: Geometry, Drift, and the Hidden Physics of LLM Behavior

A practical reading of how recursive LLM agents converge, drift, or wander depending less on the model than on the loop we force it to run.

December 14, 2025 · 17 min · Zelina
Cover image

When Tokens Become Actions: A Policy Gradient Built for Transformers

A mechanism-first reading of GPG, a Transformer-aware policy-gradient framework that turns output segments into trainable macro-actions for LLM agents.

December 14, 2025 · 14 min · Zelina
Cover image

ExaCraft and the Missing Layer of AI Education: When Examples Finally Adapt

A mechanism-first reading of ExaCraft, an AI education system that treats learner behavior—not just learner profiles—as the missing layer of personalized examples.

December 13, 2025 · 16 min · Zelina
Cover image

ImplicitRDP: When Robots Stop Guessing and Start Feeling

A mechanism-first reading of ImplicitRDP, showing why force-aware robot policies need causal structure, not just extra sensor channels.

December 13, 2025 · 17 min · Zelina
Cover image

RL Grows a Third Dimension: Why Text-to-3D Finally Needs Reasoning

A mechanism-first reading of why reinforcement learning for text-to-3D generation needs specialized rewards, token-level optimization, reasoning-heavy benchmarks, and coarse-to-fine training.

December 13, 2025 · 16 min · Zelina
Cover image

SceneMaker: When 3D Scene Generation Stops Guessing

SceneMaker shows why open-set 3D scene generation needs separate priors for de-occlusion, geometry, and pose instead of forcing one pipeline to guess everything at once.

December 13, 2025 · 15 min · Zelina
Cover image

Suzume-chan, or: When RAG Learns to Sit in Your Hand

A mechanism-first reading of Suzume-chan shows why embodied RAG may matter less as a robot novelty and more as a practical interface for capturing, preserving, and replaying expert knowledge.

December 13, 2025 · 18 min · Zelina