The Right Answer Is Not Enough: Verify the Reasoning Before You Train on It
ORACLE shows how selective step-level verification can improve synthetic reasoning data without requiring every reasoning step to be formally executable.
ORACLE shows how selective step-level verification can improve synthetic reasoning data without requiring every reasoning step to be formally executable.
Hidden-state trajectories may let AI systems detect failing reasoning before the answer appears, then spend verification or corrective compute selectively.
KG-Reasoner tests whether graph-backed assistants become more reliable when retrieval, reasoning, and error recovery remain inside one persistent trajectory.
Preference packing removes duplicated prompt processing from preference optimization, but its payoff depends on the balance between prompt and response lengths.
A two-stage training approach suggests that reasoning systems should learn which self-corrections are worth making before operators spend more inference compute on reflection.
Ordinal reward modeling learns preference-strength boundaries directly from graded human feedback, improving fine-grained calibration while reducing severe reward-ranking errors.
A noise-corrected preference loss offers alignment teams a way to reduce the damage from mislabeled comparisons without requiring perfectly cleaned data or an exact error-rate estimate.
A social-choice view of RLHF shows why preprocessing, candidate representation, reward-model constraints, and policy optimization can change whose preferences the model ultimately serves.
A theoretical survey reframes PPO, DPO, SimPO, and related alignment methods as choices about preference assumptions, policy drift, and data coverage rather than as interchangeable losses.
GANPO shows that representation-level regularization may matter more for preserving preference-tuned behavior under noisy decoding than for improving nominal alignment scores.