Cover image

More Data, Better Learner: What Children Reveal About Learning Efficiency

TL;DR for operators Adding more data and becoming better at learning from data are different objectives. Across five longitudinal vocabulary datasets covering American English, Norwegian, and Japanese, young children show strongly increasing returns to developmental experience. Using a matched per-word estimator, the language models examined in the study produce a median acceleration estimate of 1.16, with an interquartile range of 0.93–1.46. Corrected child estimates fall around 10.4–13.8. ...

September 24, 2026 · 7 min · Zelina
Cover image

Personalization Starts Before the New User Arrives

TL;DR for operators A new user with only a few preference signals does not necessarily need a richer model built from scratch. The stronger design question may be where personalization starts. The paper studies a reward-modeling system that learns from previous users how a new user’s reward weights should be initialized, then adapts only those lightweight weights from limited feedback. Its average accuracy gains are modest but consistent, while the more informative evidence comes from component ablations, few-shot unseen-user tests, worst-user analysis, and parameter scaling. Removing the learned adaptation mechanism causes the largest ablation drop. ...

September 15, 2026 · 7 min · Zelina
Cover image

Shape the Signal, Keep the Objective: MeRLa’s Bet on Reusable RLHF Rewards

TL;DR for operators RLHF teams usually have two obvious levers: improve the reward model or improve the policy optimizer. Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback introduces a third.1 MeRLa learns an additional task-aware reward signal across auxiliary tasks, freezes it, and adds it to the existing reward during subsequent policy optimization. ...

August 23, 2026 · 7 min · Zelina
Cover image

Textual Gradients and Workflow Evolution: How AdaptFlow Reinvents Meta-Learning for AI Agents

TL;DR for operators Most agent teams eventually discover that “the workflow” is not one thing. A customer-support agent, a coding agent, and a mathematical reasoning agent may all use decomposition, verification, consensus, and answer extraction—but not in the same order, not with the same emphasis, and definitely not with the same failure modes. Static agent templates look tidy in architecture diagrams. Then the first heterogeneous workload arrives, and the diagram starts quietly sweating. ...

August 12, 2025 · 21 min · Zelina