One Step Forward, One Step Back: OSOL Treats Multi-Domain RL as a Sequential Control Problem
TL;DR for operators Training one model across mathematics, code, instruction following, and chat creates a sequencing problem as well as an allocation problem. An update can improve the capability currently receiving optimization pressure while partially reversing output changes produced by the preceding update. A training run can therefore look relatively benign when gradients are inspected at one checkpoint yet still accumulate damage through the order in which updates are realized. ...