TL;DR for operators
A wireless team choosing a model to standardize has a tempting shortcut: benchmark several candidates, find the highest score, and carry that winner into deployment-specific adaptation. The problem is that wireless-channel performance depends heavily on what is being predicted and on the propagation environment in which the model is evaluated.
CFM-Bench1 makes those conditions more comparable across six channel domains and six task groups. Its representative experiments still produce repeated ranking reversals. WiFo leads most frequency-extrapolation tests but not all of them; current-beam prediction has different winners across the five evaluated domains; localization winners change again. Higher computational cost is not a reliable substitute for this evidence.
For model owners, network teams, and procurement functions, the decision should therefore be framed around the intended domain-task pair under a controlled evaluation protocol. Treat cross-domain averages as secondary summaries, document pretraining and adaptation exposure, and isolate correlated physical units before trusting validation results.
Standardizing one model requires a harder test than finding one winner
The appeal of a reusable pretrained channel representation is straightforward: learn structure from wireless-channel data once, then adapt that representation across downstream tasks rather than building every task-specific system from scratch. That is the practical promise behind a channel foundation model.
Comparing that promise has been difficult because models have often arrived with their own data sources, partitions, radio configurations, adaptation procedures, targets, and metrics. If the model changes at the same time as the evaluation pipeline, a score difference mixes representation quality with protocol differences.
CFM-Bench addresses that comparison problem by fixing downstream conditions where possible. It contains 157,900 official single-frame examples drawn from six heterogeneous sources: a 3GPP statistical urban-microcell simulation, two independent ray-tracing pipelines, measured industrial channels, measured UAV air-to-ground channels, and a synchronized vehicular channel-plus-sensor simulation.
The benchmark then defines task eligibility rather than pretending every source supports every target. CSI feedback is available across all six domains, for example, while complex temporal extrapolation is evaluated only on R2 and M1. That distinction matters because a superficially uniform benchmark can become physically meaningless if unavailable labels, incompatible codebooks, or differing coordinate systems are forced into one task definition.
Matched evaluation reveals ranking reversals rather than a universal model hierarchy
The representative experiments are main benchmark evidence: their role is to show what happens when established task-specific models and pretrained channel models are evaluated under the common substrate.
Frequency-domain extrapolation initially looks closest to producing a stable leader. Here the benchmark measures error between predicted and target complex channel values using normalized mean-squared error in decibels; lower values mean better reconstruction fidelity.
WiFo records the lowest NMSE on four of five evaluated domains: -10.08 dB on S1, -2.60 dB on R1, -18.93 dB on R2, and -3.86 dB on E1. But E2 breaks the pattern. CSI-MAE reaches -0.04 dB there, WiFo is at 0.00 dB, and LLM4CP is at 1.32 dB. The benchmark also notes that E2 can be fitted on observed training flights while remaining difficult on held-out flights, with validation and test performance near 0 dB.
The ranking becomes less stable on current-beam prediction. PFNet leads S1, R1, and R2 with Top-1 accuracies of 78.87%, 71.10%, and 97.61%. BPNN leads E1 at 49.49%. CSI-MAE leads E2 at 18.56%.
There is an important attribution boundary here: PFNet receives additional position-label supervision during training. Its advantage cannot be interpreted as an architecture-only effect. By making such supervision differences explicit, the benchmark helps separate model design from the information supplied during adaptation.
Localization produces another reshuffling. ABPN gives the lowest mean 3D error on E1 and M1; DyLoc leads E2 and R1; AAResCNN leads R2 and S1. A team looking only at the best result from one environment would therefore carry very little evidence about which architecture should be preferred in another.
The paper directly supports a conditional conclusion: within these benchmark configurations, transfer performance depends on the combination of model, domain, and task. It does not establish a generally superior channel foundation model.
Leakage control changes what a good validation score means
Wireless observations can be strongly correlated across nearby frames. Consecutive samples from the same UAV flight, pedestrian trajectory, measurement session, or vehicle link may contain much of the same physical information. Randomly assigning such frames to training and test sets can make a model appear to generalize when it is partly encountering close relatives of its training observations.
CFM-Bench instead partitions at the largest available independent physical unit. S1 separates simulation realizations and UE routes. R1 uses spatial regions with one-metre guard bands. R2 separates pedestrian trajectories. E1 separates measurement sessions, E2 UAV flights, and M1 vehicle links.
It also separates benchmark exposure from downstream adaptation: benchmark splits are excluded from foundation-model pretraining, parameter updates are restricted to the official training split, validation is reserved for selection, and the test units are used for final scoring.
For an operator, the significance is procedural. A model-selection process should not accept “held-out test data” as sufficient documentation. The relevant question is what physical unit was held out. For spatially and temporally correlated channel data, that choice can determine whether the score measures useful transfer or familiarity with nearly duplicated conditions.
FLOPs are a cost input, not a quality ranking
The benchmark reports computational complexity alongside performance, and the relationship is not monotonic.
On S1 frequency extrapolation, WiFo achieves -10.08 dB at 27.20G FLOPs, while CSI-MAE reaches -9.69 dB at 326.29G FLOPs. On E1, however, WiFo is also the most computationally expensive of the three reported models at 518.28G FLOPs and obtains the best NMSE, -3.86 dB. The direction of the performance-compute relationship changes with the case.
Localization shows the same problem with treating compute as a proxy. On E2, DyLoc produces the lowest positioning error, 7.51 metres, at 4.97G FLOPs. On R2, the much cheaper DyLoc configuration uses 801M FLOPs yet performs far worse than AAResCNN and ABPN, with a mean error of 343.95 metres.
Cognaptus therefore interprets FLOPs as one axis of a deployment decision, not a model-quality score. A wireless platform team should compare the accuracy-cost frontier within the target domain and task, then include the actual serving environment, latency requirements, memory limits, and adaptation burden before standardizing an architecture. CFM-Bench supplies part of that decision record; it does not substitute for deployment profiling.
Keep the domains separate when making the final decision
The natural temptation after building a multi-domain benchmark is to collapse it into one headline number. CFM-Bench explicitly gives reasons not to treat such an aggregate as the primary decision signal.
Its six domains differ in calibration, propagation assumptions, hardware effects, label density, sample counts, and split difficulty. They are not statistically exchangeable samples of one common wireless environment. A macro-average can summarize results, but it cannot establish that a one-point improvement on one domain means the same thing as a one-point improvement on another.
There are broader boundaries as well. The benchmark selects one fixed configuration from each source rather than covering the full range of carriers, antenna configurations, deployments, weather, or out-of-distribution environments. Some tests isolate flights, sessions, links, or trajectories while remaining inside a shared broader environment, so leakage-resistant splitting should not be confused with unseen-scene generalization. CSI-feedback evaluation uses reference CSI and does not reproduce the complete hardware and control chain of deployed FDD feedback.
Those limits narrow the claim without weakening the benchmark’s main use. CFM-Bench provides a more credible way to compare candidate models under controlled, heterogeneous conditions. It shows why the question for a wireless team is not “Which channel model wins the benchmark?” but “Which candidate survives a credible test of the domain, task, supervision, and cost profile we actually intend to deploy?”
For platform standardization, procurement, or model-selection governance, that is a more defensible decision rule than headline accuracy, model scale, or FLOPs alone.
Cognaptus: Automate the Present, Incubate the Future.
-
Yuan Gao and Wenjun Yu and Jun Jiang and Yunfan Li and Xinyu Guo and Shugong Xu (2026). CFM-Bench: A Unified Multi-Domain, Multi-Task Benchmark for Channel Foundation Models. arXiv:2607.14975. https://arxiv.org/abs/2607.14975 ↩︎