TL;DR for operators

A forecasting team choosing between modeling each signal largely from its own history and building a system that mixes information across signals can reach very different conclusions depending on the benchmark it uses.

In this paper, channel-independent models win all 10 standard multivariate time-series datasets by MSE. On a newly assembled benchmark containing more exploitable cross-channel structure, they win only 3 of 10.1

The operational lesson is narrower than “use channel mixing.” Before paying the engineering and tuning cost of a more complex multivariate architecture, measure whether other signals contain delayed information that can actually improve prediction. Even then, retain simple baselines: DLinear still delivers the lowest MSE on several of the new benchmark datasets.

The benchmark can determine what a model gets to demonstrate

Established benchmarks often look like neutral arenas. Compare several architectures under the same metrics, tune them carefully, and the ranking appears to tell you which modeling strategy works.

That interpretation becomes fragile when the capability being tested depends on a property the datasets barely contain.

For multivariate forecasting, the relevant architectural choice is straightforward. A channel-independent model relies primarily on each signal’s own history. A channel-dependent model can also use the histories of other signals. The latter only has an advantage to exploit when those other histories contain information about the target’s future.

The paper’s starting result is therefore uncomfortable for benchmark interpretation: across 10 standard datasets, the channel-independent strategy wins every dataset by MSE. That could mean cross-channel modeling is generally unhelpful. Or it could mean the datasets offer little useful cross-channel information in the first place.

The authors investigate the second possibility.

Delayed dependence matters more than correlation at the same moment

Two signals can move together without one providing usable forecasting information about the other.

If temperature and electricity demand are strongly related at 3 p.m., that contemporaneous relationship does not automatically help a model forecasting demand at 3 p.m., because the 3 p.m. temperature may also be unavailable when the forecast is issued. A delayed relationship is more operationally relevant: information already observed in one channel may help predict a future value in another.

The paper therefore profiles the fraction of ordered channel pairs whose strongest measured mutual information occurs at a positive lag rather than at the same timestep. It calls this quantity $MI_{\mathrm{lag}}$.

The second measure asks a different question. Statistical dependence can exist without a forecasting architecture being able to exploit it. The paper defines channel-dependency gain as

$$ \mathrm{CD}_{g}=100\cdot\frac{\mathrm{MSE}_{\mathrm{CI}}-\mathrm{MSE}_{\mathrm{CD}}}{\mathrm{MSE}_{\mathrm{CI}}}. $$

Here, the comparison is specifically between TSMixer and its channel-independent counterpart. Positive $\mathrm{CD}_{g}$ means the channel-mixing version achieved lower test error.

That distinction is central: $MI_{\mathrm{lag}}$ describes where dependence appears, while $\mathrm{CD}_{g}$ asks whether one model family can turn cross-channel information into better forecasts.

The synthetic tests validate measurement, not forecasting superiority

The authors first test four dependency measures on six four-channel synthetic datasets with known planted structures: independent, contemporaneous linear, lagged linear, lagged nonlinear, multiplicative joint, and additive mixed dependencies.

Lagged mutual information and CD gain recover all six planted configurations. Transfer Entropy recovers five; Granger Causality three.

This experiment is best read as measurement validation. It supports the use of the two selected indicators across a wider range of planted dependency patterns. It does not establish that either statistic is a complete description of real-world multivariate structure.

That boundary matters because the later benchmark argument depends on measuring the property that channel-mixing architectures are supposed to exploit.

Standard datasets contain much less exploitable coupling

The contrast across dataset groups is substantial.

Dataset group Median $MI_{\mathrm{lag}}$ Median CD gain
10 standard MTSF datasets 23% -4.9%
6 chaotic ODE systems 78% +4.6%
10 MixBench-TS datasets 55.5% +1.7%

On every standard dataset, the channel-independent TSMixer variant beats the channel-mixing version. By contrast, all six chaotic ODE reference systems produce positive CD gain.

Across the full set of 26 profiled standard, coupled, and ODE datasets, $MI_{\mathrm{lag}}$ has a Spearman correlation of 0.78 with CD gain. The corresponding correlations for Granger Causality and Transfer Entropy are 0.04 and 0.08.

This is associative evidence, not causal identification. Still, it connects the dataset property of interest to a measurable change in the value obtained from channel mixing.

MixBench-TS changes the ranking, but does not crown a universal architecture

The authors then screen real-world datasets for positive CD gain and retain 10 as MixBench-TS. Seven also have $MI_{\mathrm{lag}}>50%$ and form MixBench-TS-Core.

Six forecasting models are compared under a common protocol: DLinear, TSMixer_CI, and CycleNet as channel-independent models; TSMixer, iTransformer, and SimpleTM as channel-dependent models. Each model-dataset-horizon combination receives 20 Optuna trials, including lookback-window tuning, with test results averaged across horizons and three seeds.

The ranking changes sharply. Channel-independent models win 10 of 10 standard datasets by MSE but only 3 of 10 MixBench-TS datasets. By MAE, the corresponding counts are 8 of 10 and 2 of 10.

The strongest examples also line up with measured coupling. Bavaria-Donau, Bavaria-Iller, Current-Velocity, and CzeLan have 67–83% lagged-coupled channel pairs, and a channel-dependent model achieves the lowest MSE on each.

But positive CD gain is not equivalent to “complex multivariate model wins.” DLinear has the lowest MSE on ZafNoo, AirQuality, and HouseholdPower. Even on a benchmark deliberately assembled around exploitable coupling, a simple channel-independent baseline remains competitive.

For forecasting teams, profile first and then pay for mixing

Cognaptus inference from these results is operational rather than architectural.

A forecasting platform evaluating whether to add cross-channel modeling can insert a dataset-profiling step before model selection. The affected user is the team choosing the forecasting architecture; the decision is whether added channel-mixing complexity is justified; the relevant condition is the presence of delayed, exploitable information across signals.

That does not require treating the paper’s two measurements as permanent routing rules. A reasonable workflow is to use coupling diagnostics as evidence that determines which model families deserve serious testing.

For a weakly coupled dataset, extensive tuning of complex channel-mixing architectures may have low expected value. For a strongly lag-coupled dataset, excluding those architectures could leave predictive information unused. In both cases, simple baselines remain necessary because measured coupling does not determine which particular model will exploit it best.

Benchmark owners face the same issue at a different level. If a benchmark is intended to evaluate a capability, its datasets need to contain enough of the corresponding structure for that capability to matter.

Where the evidence stops

MixBench-TS is not an independent sample showing that channel mixing generally works better. Candidate datasets are retained precisely when TSMixer produces positive CD gain. Its positive CD-gain distribution is therefore partly built into the benchmark construction.

CD gain is also model-dependent. It compares one channel-dependent TSMixer configuration with its channel-independent variant, so it combines a property of the dataset with the ability of that architecture to exploit the property.

$MI_{\mathrm{lag}}$ has different limits. It is pairwise, records the fraction rather than the strength of lagged dependencies, and captures genuinely joint multichannel structure only indirectly. Wider datasets are profiled using sampled subsets of at most 40 channels.

The paper consequently supports a bounded conclusion: benchmark composition can materially change the apparent competitiveness of channel mixing, and measuring relevant dataset structure can make model evaluation more informative. It does not establish that cross-channel dependence guarantees gains from channel-dependent models on unseen production data.

For operators, that boundary is sufficient to change the workflow. Architecture selection should begin one step earlier—with evidence about whether the data contains the information the architecture was built to use.

Cognaptus: Automate the Present, Incubate the Future.


  1. Ibram Abdelmalak and Mischa Putzke and Jungmin Choi and Tom Hanika and Vijaya Krishna Yalavarthi and Lars Schmidt-Thieme (2026). MixBench-TS: A Multivariate Time Series Forecasting Benchmark Where Channel Mixing Pays Off. arXiv:2609.32656. https://arxiv.org/abs/2609.32656 ↩︎