TL;DR for operators

OmniMed-FL1 asks a practical systems question: when several institutions train one image-and-text model without pooling raw records, what creates more difficulty—the number of participating sites or differences in what those sites actually contain?

Within its controlled benchmark, data composition mattered more. Across 3, 5, 10, and 20 simulated clients, moving between label-skew conditions changed macro-F1 more than the nearly sevenfold increase in client count. Multimodal modeling also produced the highest proxy F1 in a modality comparison, but at 2.3 times the text-only model state, 3.1 times its measured wall time, and 2.1 times its peak memory.

For teams evaluating cross-institution multimodal AI, this points to two separate engineering problems. One is statistical: local datasets can pull training in different directions. The other is operational: richer models make every federated round heavier. The benchmark helps characterize those trade-offs, but its synthetic notes, class-level image-text pairing, sequential single-GPU simulation, and absence of a formal privacy audit keep the conclusions firmly on the systems side of the deployment boundary.

The number of hospitals is not the only scaling variable

Suppose five hospitals can already train a shared model while keeping their raw records local. Adding fifteen more appears, at first, to be the obvious scaling challenge. More participants mean more communication, more local updates, and more coordination.

But those institutions may also contain very different proportions of the conditions the model is supposed to recognize. One site may see many examples of one class and few of another. A second site may have the opposite distribution. Their local updates then reflect different training environments even when they are optimizing the same global model.

OmniMed-FL makes that distinction visible. Its main scalability grid varies both the nominal number of clients, $K\in{3,5,10,20}$, and the concentration parameter controlling label skew, $\alpha\in{0.1,1,5}$. Lower $\alpha$ means more uneven local class mixtures.

The effect of skew is larger than the effect of client count in the tested grid. With mild skew at $\alpha=5$, macro-F1 ranges from 0.849 to 0.916. At $\alpha=1$, it ranges from 0.818 to 0.862. Under severe skew at $\alpha=0.1$, the range falls to 0.643–0.738. Within any one skew row, changing the client count from 3 to 20 shifts the mean by no more than 0.10 F1.

That does not establish a universal law for federated learning. It does change which variable deserves early measurement in a deployment study. Counting institutions is straightforward; characterizing how their data distributions differ may be more consequential for optimization.

Under severe skew, close method rankings become unreliable

The benchmark also compares several federation strategies under a matched severe-skew setting with five clients.

Local-only training reaches macro-F1 0.297. FedAvg reaches $0.662\pm0.074$, FedProx $0.737\pm0.085$, and a matched 24-epoch FedMME-style one-shot ensemble $0.647\pm0.080$. Extending that one-shot local budget to 100 epochs does not improve the result: its mean falls to 0.609.

The tempting reading is that FedProx wins. The paper does not support that conclusion strongly. The FedProx–FedAvg difference sits inside the wider two-seed variation, so the benchmark cannot reliably order the two methods.

That distinction matters because many federated-learning design decisions involve small differences between means measured under noisy optimization. Here, the clearer result is not which iterative aggregator is definitively best. It is that iterative aggregation remains much more viable than local-only training under this proxy setup, while strong label heterogeneity makes close comparisons unstable.

The very low SCAFFOLD result also needs a narrow interpretation. OmniMed-FL tests an AdamW-based control-variate adaptation, not classical SGD-based SCAFFOLD. Its poor outcome therefore diagnoses that evaluated implementation and setting, not the entire SCAFFOLD family.

Multimodal improvement comes with a larger operating footprint

The modality comparison gives the benchmark its second major systems result.

At $\alpha=1$, text-only modeling reaches macro-F1 0.880, image-only modeling 0.737, and the multimodal image-text branch 0.906. Within this proxy task, combining modalities improves on either branch alone.

The extra 0.026 F1 over text-only is not free. Relative to the text branch, the multimodal system uses 2.3 times the model state, 3.1 times the measured wall time, and 2.1 times the peak GPU memory.

Federation amplifies that cost because model state moves repeatedly between server and clients. Under the paper’s fixed-state calculation,

$$ V_{\mathrm{nom}}=2KT|\theta|b, $$

so communication grows linearly with client count $K$, number of rounds $T$, parameter count $|\theta|$, and bytes per parameter $b$. With the reported parameter count, FP32 precision, eight rounds, and 20 nominal clients, the analytic bidirectional volume reaches 183.5 GiB.

This is not a measured network bill. It excludes serialization overhead, secure aggregation, compression, networking effects, and algorithm-specific auxiliary state. But it identifies a concrete infrastructure consequence: multimodal gains enlarge the object that federation must repeatedly distribute and collect.

Initialization matters more clearly than many architectural tweaks

Some ablations produce cleaner design signals than others.

At $\alpha=1$, initializing the public DistilBERT and ViT encoders from pretrained weights while keeping task heads random reaches $0.831\pm0.001$ macro-F1. Fully random initialization reaches $0.647\pm0.022$. Under the benchmark’s limited local training budget, starting from pretrained representations therefore has a large association with final performance.

By comparison, several later-stage design choices remain harder to rank. Eight fusion rules span mean F1 values from 0.636 to 0.759, but seed variation is large enough that the authors decline to identify a definitive winner.

The missing-text stress test has a similar shape. Zero filling, deterministic feature imputation, Gaussian beta-NLL imputation, and uncertainty-weighted aggregation all finish within 0.018 F1 of one another. The two probabilistic variants both average 0.706, while the uncertainty-weighted version has the largest two-seed spread.

The anti-collapse ablation reveals a trade-off rather than a simple improvement. Combining class-balanced sampling with entropy-diversity regularization increases minimum predicted-class diversity under severe skew, but lowers mean F1 relative to configurations removing one or both controls. At moderate skew, all four configurations recover full class diversity and sit within 0.809–0.846 F1.

For model-development teams, the practical ordering of evidence is therefore uneven: initialization shows a large separation; fusion and missing-modality alternatives do not; collapse prevention can improve coverage while sacrificing aggregate F1.

What a cross-institution AI team can take from this benchmark

The paper directly shows how a controlled proxy system behaves. Cognaptus’ business inference is narrower: before treating institutional expansion as mainly a client-count problem, teams should characterize the distributions each site contributes.

That changes several early decisions. Data teams may need site-level class and modality audits before architecture tuning. Model teams may get more from robust pretrained initialization than from aggressively optimizing fusion operators that are not yet statistically distinguishable. Infrastructure teams should model communication from parameter count, rounds, and participating clients before assuming multimodal federation has an acceptable operating envelope.

Governance requires a separate track. Keeping raw records at each institution provides data locality, but the study does not establish a formal privacy guarantee. Any real deployment would still need explicit threat models, privacy mechanisms, security testing, and evaluation of what model updates may leak.

This is a systems benchmark, not a clinical validation study

The benchmark deliberately does not reproduce several properties of real clinical deployment.

Its clinical notes are synthetic. Images and notes are paired by class rather than by patient, so the experiment does not test authentic image-report correspondence, contradiction, or missingness. Different image classes partly come from different public source collections, creating potential class-source confounding. Most model comparisons use only two seeds.

The scaling measurements are also simulations: clients execute sequentially on one NVIDIA H100 NVL rather than concurrently across heterogeneous institutional infrastructure. Network latency, stragglers, client dropout, secure aggregation, and real distributed traffic are outside the measurement.

These boundaries determine how the findings should be used. OmniMed-FL provides a benchmark for deciding what to stress-test next: heterogeneity, initialization, multimodal cost, communication, and collapse behavior. It does not provide evidence that the resulting classifier is diagnostically safe or operationally ready for hospitals.

The most consequential result is therefore not a single F1 score. It is a change in the scaling question. In this benchmark, adding institutions was less damaging to model performance than changing the composition of what those institutions held. Meanwhile, adding modalities improved the proxy task while making the federated system materially heavier. Teams exploring multimodal federation should measure both effects before assuming that more sites and richer inputs move the system in the same direction.

Cognaptus: Automate the Present, Incubate the Future.


  1. Ayush Debnath and Ruelia Saha and Sudip Misra (2026). OmniMed-FL: A Robust Multimodal Federated Learning Framework for Clinical Diagnosis. arXiv:2609.10364. https://arxiv.org/abs/2609.10364 ↩︎