Cover image

The Right Answer Is Not Enough: Verify the Reasoning Before You Train on It

ORACLE shows how selective step-level verification can improve synthetic reasoning data without requiring every reasoning step to be formally executable.

September 17, 2026 · 7 min · Zelina
Cover image

Catch the Drift Before the Answer: Reasoning Trajectories as a Runtime Control Surface

Hidden-state trajectories may let AI systems detect failing reasoning before the answer appears, then spend verification or corrective compute selectively.

September 16, 2026 · 8 min · Zelina
Cover image

One Trajectory, More Recovery: KG-Reasoner Reworks Multi-Hop Graph Reasoning

KG-Reasoner tests whether graph-backed assistants become more reliable when retrieval, reasoning, and error recovery remain inside one persistent trajectory.

September 16, 2026 · 7 min · Zelina
Cover image

Stop Paying Twice for the Prompt: Preference Packing Reworks DPO’s Execution Layout

Preference packing removes duplicated prompt processing from preference optimization, but its payoff depends on the balance between prompt and response lengths.

September 16, 2026 · 7 min · Zelina
Cover image

Think Again, but Make It Count: Train Reflection Before You Spend More Tokens on It

A two-stage training approach suggests that reasoning systems should learn which self-corrections are worth making before operators spend more inference compute on reflection.

September 16, 2026 · 7 min · Zelina
Cover image

When Preference Strength Becomes Part of the Model

Ordinal reward modeling learns preference-strength boundaries directly from graded human feedback, improving fine-grained calibration while reducing severe reward-ranking errors.

September 16, 2026 · 7 min · Zelina
Cover image

When the Wrong Label Shouts Loudest: Correcting Preference Noise in RLHF and DPO

A noise-corrected preference loss offers alignment teams a way to reduce the damage from mislabeled comparisons without requiring perfectly cleaned data or an exact error-rate estimate.

September 16, 2026 · 8 min · Zelina
Cover image

Your Reward Model Is Also a Voting Rule: What Social Choice Changes About RLHF

A social-choice view of RLHF shows why preprocessing, candidate representation, reward-model constraints, and policy optimization can change whose preferences the model ultimately serves.

September 16, 2026 · 8 min · Zelina
Cover image

Alignment Is a Coverage Problem Before It Is a Loss-Function Problem

A theoretical survey reframes PPO, DPO, SimPO, and related alignment methods as choices about preference assumptions, policy drift, and data coverage rather than as interchangeable losses.

September 15, 2026 · 8 min · Zelina
Cover image

Alignment Under Heat: GANPO Targets the Fragility That Benchmarks Miss

GANPO shows that representation-level regularization may matter more for preserving preference-tuned behavior under noisy decoding than for improving nominal alignment scores.

September 15, 2026 · 7 min · Zelina