Cover image

More Critics, Less Gain: Self-Questioning Has a Stability Limit

Counterfactual self-critique can improve compact reasoning models, but the evidence suggests critic count is a stability parameter rather than a quantity to maximize.

September 15, 2026 · 6 min · Zelina
Cover image

No Click Is Not a No: Turning Passive Feedback Into Reward Signals

ImplicitRM shows how reward models can separate user preference from the likelihood of acting on it, making passive interaction logs more informative without pretending that every non-action is a negative label.

September 15, 2026 · 8 min · Zelina
Cover image

One Score, More Signals: Making Reward Models Easier to Rank and Audit

A feature-augmented reward model improves preference ranking on HH-RLHF while exposing measurable signals that product, safety, and governance teams can inspect.

September 15, 2026 · 6 min · Zelina
Cover image

Personalization Starts Before the New User Arrives

A meta-learning approach suggests that cold-start personalization may depend more on learning where user adaptation begins than on adding more user-specific model capacity.

September 15, 2026 · 7 min · Zelina
Cover image

The Reward Model Has to Move Too: R2M Tracks the Policy During RLHF

R2M shows how lightweight reward-model adaptation can track a changing policy without retraining the evaluator’s full backbone.

September 15, 2026 · 7 min · Zelina
Cover image

Buy Fewer Labels, Ask Better Questions: RLHF as an Allocation Problem

A Gemma 9B experiment suggests RLHF teams can extract substantially more value from preference budgets by adapting what they label, when they update, and which comparisons receive judgment.

September 14, 2026 · 7 min · Zelina
Cover image

One Model, Two Routes: Why Audio-Omni Unifies Audio by Splitting Its Controls

Audio-Omni suggests that consolidating multimodal audio systems works best when semantic intent and temporally precise controls remain architecturally distinct.

September 14, 2026 · 9 min · Zelina
Cover image

Same VQA Score, Different Eyes: What Fine-Grained Tests Reveal About VLMs

Broad VLM benchmarks can hide large differences in precise visual discrimination, changing how teams should evaluate models and allocate multimodal training effort.

September 14, 2026 · 8 min · Zelina
Cover image

Search With Vision, Answer With Text? What IRPAPERS Changes About Scientific RAG

IRPAPERS shows why scientific RAG systems may benefit from visual retrieval without making images the default representation for answer generation.

September 14, 2026 · 7 min · Zelina
Cover image

Think Harder, See Less? What Visual Illusions Reveal About Multimodal Reasoning

IllusionReasoning shows that extra multimodal deliberation helps explanation but can hurt simpler perception and constrained-choice tasks, making reasoning depth a routing decision rather than a default upgrade.

September 14, 2026 · 7 min · Zelina