More Critics, Less Gain: Self-Questioning Has a Stability Limit
Counterfactual self-critique can improve compact reasoning models, but the evidence suggests critic count is a stability parameter rather than a quantity to maximize.
Counterfactual self-critique can improve compact reasoning models, but the evidence suggests critic count is a stability parameter rather than a quantity to maximize.
ImplicitRM shows how reward models can separate user preference from the likelihood of acting on it, making passive interaction logs more informative without pretending that every non-action is a negative label.
A feature-augmented reward model improves preference ranking on HH-RLHF while exposing measurable signals that product, safety, and governance teams can inspect.
A meta-learning approach suggests that cold-start personalization may depend more on learning where user adaptation begins than on adding more user-specific model capacity.
R2M shows how lightweight reward-model adaptation can track a changing policy without retraining the evaluator’s full backbone.
A Gemma 9B experiment suggests RLHF teams can extract substantially more value from preference budgets by adapting what they label, when they update, and which comparisons receive judgment.
Audio-Omni suggests that consolidating multimodal audio systems works best when semantic intent and temporally precise controls remain architecturally distinct.
Broad VLM benchmarks can hide large differences in precise visual discrimination, changing how teams should evaluate models and allocate multimodal training effort.
IRPAPERS shows why scientific RAG systems may benefit from visual retrieval without making images the default representation for answer generation.
IllusionReasoning shows that extra multimodal deliberation helps explanation but can hurt simpler perception and constrained-choice tasks, making reasoning depth a routing decision rather than a default upgrade.