Cover image

More Hospitals, Different Data: OmniMed-FL Finds Heterogeneity Before Scale

TL;DR for operators OmniMed-FL1 asks a practical systems question: when several institutions train one image-and-text model without pooling raw records, what creates more difficulty—the number of participating sites or differences in what those sites actually contain? Within its controlled benchmark, data composition mattered more. Across 3, 5, 10, and 20 simulated clients, moving between label-skew conditions changed macro-F1 more than the nearly sevenfold increase in client count. Multimodal modeling also produced the highest proxy F1 in a modality comparison, but at 2.3 times the text-only model state, 3.1 times its measured wall time, and 2.1 times its peak memory. ...

October 5, 2026 · 7 min · Zelina