Cover image

Shape the Signal, Keep the Objective: MeRLa’s Bet on Reusable RLHF Rewards

TL;DR for operators RLHF teams usually have two obvious levers: improve the reward model or improve the policy optimizer. Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback introduces a third.1 MeRLa learns an additional task-aware reward signal across auxiliary tasks, freezes it, and adds it to the existing reward during subsequent policy optimization. ...

August 23, 2026 · 7 min · Zelina