More Data, Better Learner: What Children Reveal About Learning Efficiency
TL;DR for operators Adding more data and becoming better at learning from data are different objectives. Across five longitudinal vocabulary datasets covering American English, Norwegian, and Japanese, young children show strongly increasing returns to developmental experience. Using a matched per-word estimator, the language models examined in the study produce a median acceleration estimate of 1.16, with an interquartile range of 0.93–1.46. Corrected child estimates fall around 10.4–13.8. ...