Cover image

When More RLHF Means More Sycophancy: Audit the Reward Tilt First

A mechanistic account of when preference optimization amplifies sycophancy—and why reward-model auditing should precede stronger post-training.

September 14, 2026 · 7 min · Zelina
Cover image

When the Reward Is Right but the Incentive Is Wrong

Stackelberg Reward Shaping argues that preference-guided assistants may need prompt-specific reward transformation, not simply stronger optimization of the same reward score.

September 14, 2026 · 8 min · Zelina
Cover image

Aviation AI Needs a Shared Operational State Before It Needs a Bigger Model

AviationLMM reframes multimodal aviation AI as an information-integration architecture problem, but its value today is a design agenda rather than demonstrated system performance.

September 13, 2026 · 7 min · Zelina
Cover image

Correct Answer, Weak Evidence: Measuring Multimodal Reasoning at the Fact Level

MuRGAt shows why multimodal model evaluation needs to measure whether cited evidence supports each factual step, separately from whether the final answer is correct.

September 13, 2026 · 7 min · Zelina
Cover image

Correct on the Frame, Wrong on the Timeline

TimeBlind shows why high video-question accuracy can conceal weak temporal reasoning—and how teams can test that failure before deployment.

September 13, 2026 · 7 min · Zelina
Cover image

Four Inputs In, One Modality Out: Testing Whether Omnimodal Models Actually Arbitrate Evidence

C³PO shows that multimodal reliability depends less on accepting more inputs than on keeping competing evidence active long enough to resolve conflicts correctly.

September 13, 2026 · 7 min · Zelina
Cover image

The Wrong Answer May Start Before Reasoning

A process-level view of multimodal math shows why teams should diagnose perception, alignment, and reasoning separately—and verify more deeply only when the cost of error warrants it.

September 13, 2026 · 8 min · Zelina
Cover image

Thinking Longer, Looking Elsewhere

A new attention analysis shows why longer multimodal reasoning can fail even when the model still sees the right visual evidence—and when inference-time intervention may help.

September 13, 2026 · 7 min · Zelina
Cover image

When Vision Fails in Both Directions

AMVICC shows why multimodal model selection should test the specific visual constraints a workflow depends on rather than rely on a single capability score.

September 13, 2026 · 7 min · Zelina
Cover image

Every Token Looks Everywhere: The Quadratic Bill Behind Attention

A mathematical view of attention shows why long context is expensive, why position must be added explicitly, and why efficient-attention choices depend on the workload.

September 12, 2026 · 7 min · Zelina