TL;DR for operators
An organization that wants several hospitals, branches, or edge devices to train a shared multimodal model faces more than a data-locality problem. Each participating site also needs enough memory and network capacity to take part in training. Full-model federated learning keeps raw data local, but it requires every client to host and exchange the complete model.
USplit-VQA1 tests a different allocation. Clients retain the beginning and end of the model—including the components that touch raw inputs and labels—while the heavier middle runs on a server. In the evaluated configurations, that reduces estimated client memory by about 5.7–5.8× and communication per round by 6.1–10.8× relative to federated learning.
The trade-off is that partitioning is not operationally neutral. With the paper’s lightweight Custom backbone, USplit-VQA beats federated learning on all four evaluated datasets. With BiomedCLIP under one fixed split, it loses to federated learning on all four. Its contribution-aware defense also performs strongly against one malicious client but degrades sharply when two or more of five clients are malicious. For deployment teams, the split point belongs in model validation alongside architecture, accuracy, latency, bandwidth, and threat assumptions.
Distributed learning is also a model-partitioning problem
Federated learning solves one constraint cleanly: raw training examples do not need to be pooled on a central server. It does much less for a second constraint—the endpoint still carries the full model.
USplit-VQA changes that division of labor. The client runs the initial visual and text processing, sends intermediate representations to a server, receives the processed multimodal representation back, and computes the final prediction and loss locally. Because the classification head remains on the client, ground-truth labels also remain there.
This is the “U” in U-shaped split learning: computation starts on the client, moves through the server, and returns to the client before the training objective is evaluated.
For the evaluated BiomedCLIP configuration, federated learning places 226.7 million parameters on each client; USplit-VQA places 38.8 million there. Estimated client memory falls from 2,720 MB to 465 MB. With the Custom model, client parameters fall from 8.5 million to 1.5 million, and estimated memory falls from 102 MB to 18 MB.
Communication falls as well. The reported per-round volume decreases from 907 MB to 84 MB for BiomedCLIP and from 34 MB to 5.6 MB for the Custom model.
For teams whose endpoints are the binding constraint, those are large changes in deployment economics rather than marginal engineering savings.
Offloading computation does not preserve accuracy automatically
The more consequential result appears when resource savings are compared with predictive performance.
Using the Custom model, USplit-VQA exceeds federated learning on every evaluated dataset: VQA-RAD, SLAKE, PathVQA, and VizWiz. On SLAKE, for example, accuracy is 67.77% under USplit-VQA versus 65.88% under federated learning. On VQA-RAD, the corresponding figures are 43.90% and 36.36%.
BiomedCLIP reverses the pattern. Under the paper’s fixed split, federated learning outperforms USplit-VQA on all four datasets. SLAKE drops from 83.51% to 74.18%; PathVQA from 69.18% to 58.25%.
That comparison changes how the architecture should be evaluated. Moving layers to a server is not merely an infrastructure decision made after model selection. The partition can alter what representations are learned and how gradients propagate through the distributed system.
The BiomedCLIP result does not establish that large pretrained models are poor candidates for split learning. The paper evaluates only one BiomedCLIP split. It does establish something narrower and more operationally relevant: a split that saves substantial client resources can still be a poor split for model quality.
There may be no single best split point
The paper’s Custom-model split analysis makes the trade-off more explicit.
Four variants progressively move more computation onto the client. The client share ranges from 2% of model parameters in V1 to 81.4% in V4. V3 produces the highest reported accuracy, 69.75%, while V4 reaches 69.18% but records the lowest training time per round.
The experiment is best read as a sensitivity test rather than a second headline result. Its purpose is to show that partition geometry changes several objectives at once.
For an engineering team, this turns split selection into a constrained optimization problem. A hospital workstation with abundant compute but expensive network connectivity may justify a different boundary from a small edge device connected to fast infrastructure. Neither boundary should be selected solely from parameter count.
CAWA works best when malicious participation is limited
Distributed training creates another allocation problem: how much influence should each participant have?
USplit-VQA adds Contribution-Aware Weighted Aggregation, or CAWA. It compares each client’s gradient direction with those of its peers, gives greater importance to agreement with clients that already have higher reputation, and updates reputation using adaptive thresholds and repeated positive or negative behavior. The resulting trust score scales the client’s training contribution.
The one-attacker experiment is strong. With one malicious participant among five, the attacker’s trust weight is reduced by 98.2%. USplit-VQA retains 65.2% clean accuracy with an attack success rate of 0.6%. Removing CAWA produces 61.2% accuracy and a 44.1% attack success rate.
The defense becomes much less persuasive as the attacker share rises. With two malicious clients, USplit-VQA’s attack success rate reaches 83.5%; with three, 94.2%.
CAWA therefore provides evidence for contribution-aware screening under limited corruption. It does not support treating reputation-weighted aggregation as sufficient protection when coordinated malicious participants may form a substantial part of the training population.
Keeping raw data local reduces exposure, but does not prove privacy
Split learning still transmits intermediate activations and gradients. Those signals can contain information about the original inputs, so retaining raw images on the client does not eliminate reconstruction risk.
The paper tests model inversion and Deep Leakage from Gradients attacks. Under model inversion, USplit-VQA produces markedly poorer reconstructions than centralized or federated training: MSE rises to 0.0175, compared with 0.0013 for federated learning, while PSNR falls to 17.56 from 28.8. Its SSIM is also lower, at 0.796 versus 0.988.
Gradient-inversion differences are smaller but point in the same direction under the tested setup.
These experiments support a specific claim: reconstruction quality was lower under the attacks the authors evaluated. They do not establish formal privacy, immunity to reconstruction, or resistance to future attack methods.
What operators should validate before deployment
Cognaptus inference from these results is that organizations evaluating split learning should add partition geometry to the model-validation process rather than treating it as infrastructure configuration.
For teams operating multimodal systems across resource-constrained institutions, four measurements should move together: endpoint memory, network traffic, predictive performance, and training time. Where adversarial participants are plausible, the assumed fraction and coordination of malicious clients should be tested separately rather than inferred from a single-attacker result.
Several evidence boundaries remain. The primary distributed experiments use IID client partitions, whereas real hospitals and institutions are likely to differ systematically in cases, devices, and populations. Security and split analyses are concentrated on the Custom model with SLAKE. BiomedCLIP receives only one evaluated split. Scalability is tested from five to fifteen clients, not across large production networks.
Those limits do not erase the resource result. They specify what still has to be measured before transferring it into another deployment.
Treat the split as part of the model
USplit-VQA demonstrates a credible architectural route for moving most multimodal computation away from constrained endpoints without moving raw inputs and labels with it. The measured reductions in client memory and communication are substantial.
Its more durable contribution may be the qualification attached to those gains. Once a model is divided across client and server, the location of that boundary affects more than hardware utilization. It can change accuracy, communication, execution time, privacy exposure, and the behavior of security mechanisms.
For operators, the relevant question is therefore not simply whether split learning is cheaper than federated learning. It is which split preserves acceptable model behavior under the resource constraints and threat conditions of the actual deployment.
Cognaptus: Automate the Present, Incubate the Future.
-
Md Khalid Syfullah and Alvi Ataur Khalil (2026). USPLIT-VQA: U-Shaped Split Learning for Visual Question Answering with Contribution-Aware Weighted Aggregation. arXiv:2609.12168. https://arxiv.org/abs/2609.12168 ↩︎