TL;DR for operators

A multimodal dataset can contain the same amount of image, text, audio, or other modality-specific data yet produce different model performance depending on how often those modalities occur together. Marcus Ma and Shrikanth Narayanan isolate that distinction in Paired Multimodal Scaling Laws1, varying pairing while holding modality-specific data budgets fixed across three supervised environments.

The main practical finding is not that every example should be paired. Unpaired examples can still reduce loss associated with information available from either modality alone. What they cannot do by themselves, in the tested settings, is teach information that exists only in the relationship between modalities. Some of that synergistic information also appears only after enough paired examples and paired training exposures have accumulated.

For organizations paying a premium to align modalities, pairing therefore becomes a budget variable. The relevant question is how much joint observation is required to learn the interaction structure of the task, after which cheaper unpaired data may still contribute substantially. That conclusion is supported for the paper’s supervised classification settings, not yet for self-supervised generative pretraining or much larger multimodal systems.

The amount of each modality can stay fixed while performance moves

Suppose a team has collected 100,000 images and 100,000 associated pieces of text. One training set could contain many aligned image-text examples. Another could expose the model to roughly the same quantity of image and text data while presenting much of it separately.

A quantity-only view treats those datasets as broadly similar because the modality-specific data budgets have not changed. The paper tests that assumption directly. Its controlled experiments vary the number of paired examples while holding the amount of data from each modality fixed, along with architecture and parameter count within each environment and gradient-step compute within the relevant comparisons.

Across a constructed image-audio Digits task, VQA-v2, and NLVR2, changing the amount of pairing changes paired-test loss. Pairing is therefore not merely a side effect of how much modality data has been collected. It behaves as an independent data-composition variable.

The authors formalize this with a pairing fraction

$$ p=\frac{2n_p}{D_1+D_2}, $$

where $n_p$ is the number of paired examples and $D_1,D_2$ are modality-specific data counts. Two training regimes can therefore have the same $D_1$ and $D_2$ but different $p$.

For data planning, that distinction matters because creating a pair can require synchronization, matching, annotation, or additional collection effort that a standalone observation does not.

Unpaired data can teach content, but not information that exists only jointly

The paper explains the pairing effect by separating target-relevant information into four channels:

$$ I(X_1,X_2;Y)=R+U_1+U_2+S. $$

$R$ is redundant information available from either modality. $U_1$ and $U_2$ are information unique to each modality. $S$ is synergistic information that becomes available only when the modalities are considered jointly.

This decomposition changes how paired and unpaired data should be interpreted. In the experiments, additional unpaired data can reduce loss associated with redundant and modality-unique information. It does not reduce the synergy channel on its own. Learning $S$ requires observing the modalities together.

That distinction is especially visible because the three environments have different estimated information structures. The constructed Digits task assigns substantial information to all four channels. VQA-v2 is dominated by question-unique information and has relatively little estimated synergy. NLVR2 has near-zero estimated unimodal channels and substantially more dependence on joint information.

The proposed scaling law mirrors that structure:

$$ L(N,D_1,D_2,n_p)=E_N+L_R+L_{U_1}+L_{U_2}+L_S. $$

Instead of forcing all data into a single pooled curve, it allows each type of information to respond to the data that can actually teach it.

A further result complicates any simple choice between paired and unpaired collection. Unpaired observations become more effective for paired deployment after enough paired interaction has been learned. At low pair counts, paired and unpaired data can therefore behave as complements: some pairing appears to help the model transfer what it learns from isolated modalities into the setting where both are present.

Synergy can remain near chance until pairing clears a threshold

More pairing does not always produce a smooth improvement from the first paired example.

In the Digits environment, the paper estimates a distinct-pair floor of roughly 2,900 examples and a paired-exposure requirement of roughly 40,000 under interleaved training. Its empirical threshold is represented as

$$ n^\ast=\max\left(n_{\mathrm{floor}},\frac{r}{k}\right), $$

so synergy begins improving only after both sufficient pair diversity and sufficient repeated exposure have effectively been obtained.

NLVR2 also shows threshold-like behavior, but its escape from chance is stochastic rather than a deterministic hard gate. VQA-v2 shows much less pronounced gating, consistent with its much smaller estimated synergy component.

Presentation order provides an additional sensitivity test rather than a separate thesis. Interleaving paired and unpaired data works differently from concentrating pairs at the beginning or end: frontloaded pairing can be forgotten later, while backloaded pairing can acquire synergy less effectively.

For a training team, the risk is a discontinuous-looking low-pair regime. A pilot with too few paired observations may not merely underestimate a gradual benefit; it may fail to enter the regime in which the cross-modal interaction becomes learnable at all.

Four information channels predict pairing behavior better than mixture terms

The paper compares six pairing-aware extensions of published scaling laws with sixteen variants of its four-channel family. Evaluation uses held-out tests for scale extrapolation, full pairing, and the shape of the pairing curve, supplemented by interpolation and robustness checks.

Every four-channel variant beats the published-law extensions on the scale test. Fourteen of sixteen do so on the full-pairing test, and thirteen of sixteen perform better on pairing-shape $R^2$. The abstract reports 3.2% fit-test error for the proposed law versus 10.4% for the best pairing extension of the published alternatives.

The result supports the paper’s mechanism more strongly than an in-sample fit alone would. In particular, variants containing a synergy gate and an effective-count mechanism help capture behaviors that pooled-data or constant-interaction formulations struggle to express.

It does not establish that the same functional form will remain accurate after orders-of-magnitude increases in model or dataset scale. The main empirical fits hold architecture and parameter count fixed within each environment.

Pairing should be priced against how much the task depends on interaction

The paper’s budget analysis turns the fitted curves into a data-acquisition decision. When paired observations cost more than unpaired ones, the loss-minimizing allocation depends on the task’s information structure.

NLVR2 remains at full pairing in the reported analysis. Digits and VQA-v2 can admit interior pairing optima when paired examples cost roughly two to three times as much as unpaired observations.

The paper directly shows that these fitted optima vary by environment and pair price. Cognaptus’ inference is narrower: before committing to fully paired collection, a team whose production task is supervised and multimodal should estimate whether performance is primarily limited by joint information or by information that cheaper single-modality examples can still supply.

That requires measuring the actual downstream task rather than assigning a fixed economic value to pairing. The PID components here are label-dependent; they describe information about a particular target, not an intrinsic property of “image plus text” or any other modality combination.

Where this evidence stops

Three boundaries affect practical use.

First, all three environments are supervised and classification-oriented. The paper does not test self-supervised generative pretraining, where paired observations may themselves define part of the learning signal.

Second, substantially larger data and model scales are not evaluated. The reported experiments include 5,078 main-sweep training runs, but controlled coverage is not the same as evidence that the same scaling behavior persists into frontier-scale training.

Third, the synergy gate should not be treated as a universal hard threshold. It is pronounced in Digits, stochastic in NLVR2, and relatively weak in VQA-v2.

These constraints still leave a useful decision principle within the tested regime: count co-occurrence separately from modality volume.

The scarce resource may be joint observation

Multimodal data strategy is often expressed as a question of how much image, text, audio, or sensor data to acquire. This paper shows why that accounting can be incomplete. The same modality-specific budgets can support different amounts of joint experience, and that difference can materially change what the model learns.

For operators, the resulting question is more specific than “Do we need paired data?” Some information can be learned from isolated modalities. Some requires co-occurrence. Some unpaired data becomes more valuable only after enough co-occurrence has established the interaction structure.

Where pairing is expensive, those distinctions turn alignment from a default preprocessing assumption into an allocation problem.

Cognaptus: Automate the Present, Incubate the Future.


  1. Marcus Ma and Shrikanth Narayanan (2026). Paired Multimodal Scaling Laws. arXiv:2609.36263. https://arxiv.org/abs/2609.36263 ↩︎