When More RLHF Means More Sycophancy: Audit the Reward Tilt First
A mechanistic account of when preference optimization amplifies sycophancy—and why reward-model auditing should precede stronger post-training.
A mechanistic account of when preference optimization amplifies sycophancy—and why reward-model auditing should precede stronger post-training.
Stackelberg Reward Shaping argues that preference-guided assistants may need prompt-specific reward transformation, not simply stronger optimization of the same reward score.
AviationLMM reframes multimodal aviation AI as an information-integration architecture problem, but its value today is a design agenda rather than demonstrated system performance.
MuRGAt shows why multimodal model evaluation needs to measure whether cited evidence supports each factual step, separately from whether the final answer is correct.
TimeBlind shows why high video-question accuracy can conceal weak temporal reasoning—and how teams can test that failure before deployment.
C³PO shows that multimodal reliability depends less on accepting more inputs than on keeping competing evidence active long enough to resolve conflicts correctly.
A process-level view of multimodal math shows why teams should diagnose perception, alignment, and reasoning separately—and verify more deeply only when the cost of error warrants it.
A new attention analysis shows why longer multimodal reasoning can fail even when the model still sees the right visual evidence—and when inference-time intervention may help.
AMVICC shows why multimodal model selection should test the specific visual constraints a workflow depends on rather than rely on a single capability score.
A mathematical view of attention shows why long context is expensive, why position must be added explicitly, and why efficient-attention choices depend on the workload.