TL;DR for operators
Speech systems that must preserve coughs, laughter, yawns, breathing, and similar signals face a familiar long-tail problem: rare events deserve more training attention, but increasing that attention can also disturb ordinary transcription and better-represented categories.
Jia and colleagues test a deliberately data-centric answer in Long-Tail Rebalancing for Non-Verbal Vocalization-Aware ASR: A Track 1 System for the NVVSpeech Challenge.1 They keep the Qwen3-ASR 1.7B backbone, supervised objective, tokenizer, target format, and main optimization recipe fixed. What changes is the training distribution.
The clearest local result is not that maximum balancing wins. Natural sampling scores 62.33 on the fixed validation set. Moderate square-root sampling raises that to 66.25. A second stage that makes event categories uniformly likely then lowers the same local score to 58.91.
For operators, this turns sampling policy into a practical control variable. More exposure for rare events can help, but the safest checkpoint cannot be selected from class balance alone. Event-level accuracy, lexical error, category-specific regressions, and evaluation-set fit all matter.
Rare events compete with ordinary transcription for training attention
A production speech system may need to capture more than words. A cough can carry clinical context, laughter can change conversational interpretation, and a yawn or sigh may be part of the record a downstream application needs to preserve.
These non-verbal vocalizations, or NVVs, create two data problems at once. First, annotated datasets do not necessarily use the same labels. Second, category frequencies are highly uneven. Merging everything without reconciling labels risks treating different events as equivalent; training directly on the resulting frequency distribution leaves rare categories with relatively little exposure.
The paper addresses both issues without introducing a separate NVV detector or changing the ASR architecture. Its harmonized training corpus contains 26,648 utterances and 70.48 hours from eight Chinese and English datasets. Source labels are mapped into the challenge’s 16-category taxonomy when a correspondence can be established, while unreliable mappings are excluded.
That harmonization step matters because rebalancing only becomes meaningful after categories mean the same thing across data sources.
The model stays fixed while the training distribution moves
The training objective remains ordinary autoregressive supervised fine-tuning over a target sequence containing both lexical tokens and NVV tags:
The variable doing the interesting work is $q$, the distribution from which training examples are drawn.
For category $c$ with $n_c$ assigned utterances, the paper sets its sampling probability to
At $\alpha=1$, raw category frequencies largely determine exposure. At $\alpha=0.5$, taking the square root compresses frequency differences: rare categories receive more relative attention, but common categories still receive more probability than rare ones. At $\alpha=0$, every represented category becomes equally likely.
This provides a continuous way to control rebalancing strength without changing the loss function or model.
Moderate rebalancing produces the strongest local result
The local ablation is the paper’s main evidence for choosing sampling strength. On the fixed 214-utterance bilingual validation split, square-root sampling at $\alpha=0.5$ achieves the strongest reported single-stage score.
| Training policy | Local validation score | English WER |
|---|---|---|
| Natural sampling, $\alpha=1$ | 62.33 | 6.31% |
| Square-root sampling, $\alpha=0.5$ | 66.25 | 7.10% |
| SQRT + uniform-category stage | 58.91 | 6.95% |
The improvement is therefore not monotonic in rebalancing strength. Moving away from the natural distribution helps. Moving further toward a flat category distribution does not continue the gain.
There is also a lexical trade-off. Square-root sampling improves the overall local score and NVV recognition relative to natural sampling, but English WER rises from 6.31% to 7.10%. The training policy is reallocating capacity across aspects of the joint task, not simply adding rare-event accuracy for free.
Uniform sampling changes which events win
The second training stage starts from the square-root checkpoint and then gives every represented category equal category-level sampling probability for another five epochs.
Its per-category results clarify what stronger balancing actually does. Yawn F1 rises from 0.09 under the SQRT checkpoint to 0.39 after the uniform-category stage. Throat-clearing F1 rises from 0.50 to 0.57.
Most other reported categories move in the opposite direction. Breath falls from 0.63 to 0.32, hiss from 0.83 to 0.56, moan from 0.91 to 0.71, and burp from 0.95 to 0.76.
This is better interpreted as redistribution than as general improvement. A flatter sampler changes which categories receive enough repeated exposure to improve, while other event categories and lexical behavior can regress.
For a product team, that means the relevant question is not simply whether the dataset is imbalanced. It is which mistakes are expensive enough to justify changing the exposure distribution.
The best checkpoint depends on the evaluation environment
The local split favors the square-root checkpoint, yet the submitted two-stage system obtains the authors’ best official challenge result: 63.86 in the Final Stage, where it ranks fourth. The released Final Stage Whisper baseline scores 33.32.
That does not establish that Stage 2 is universally better than Stage 1. The official SQRT result of 52.61 comes from the Preliminary phase, while the 63.86 two-stage result comes from the Final phase. They are not a controlled same-test-set comparison.
An additional held-out-speaker check on 140 Track-1-compatible Mandarin utterances from MNV-17 produces another split decision. The two-stage model has lower joint CER, 3.21% versus 3.67% for SQRT, while SQRT has higher exact NVV accuracy, 63.57% versus 60.71%.
Checkpoint choice therefore becomes partly an evaluation-design problem. A validation set that overweights the wrong categories, speakers, languages, or lexical-event trade-offs can select a model that is poorly matched to deployment even when its aggregate score is higher.
What this changes for speech-system design
The paper directly shows that changing the sampling distribution can materially change joint lexical and NVV recognition while the model architecture and loss stay fixed.
Cognaptus infers a practical development sequence from that result: before adding specialized architectures for rare vocal events, teams can test whether label harmonization and sampling control already move the categories they care about. This may reduce engineering complexity relative to an architectural redesign, although the paper does not measure development cost or ROI directly.
Sampling strength can also be treated as a data-governance parameter. A customer-support system, clinical transcription product, or conversational analytics pipeline may value different event categories. The appropriate training distribution can therefore depend on explicit application priorities rather than an abstract goal of making all categories equally frequent.
The evidence does not establish a universal sampling rule
The main model-selection split contains only 214 utterances, with some event categories represented by as few as seven reference occurrences. Per-category changes can therefore be sensitive to small counts.
The study also uses one main backbone, Qwen3-ASR 1.7B, under one primary optimization recipe. Whether the same rebalancing pattern holds across other ASR architectures is unresolved.
The MNV-17 generalization check is useful but narrow: 140 Mandarin utterances do not establish broad multilingual or cross-domain robustness. Most importantly, local and official evaluation do not rank the checkpoints consistently.
The safest conclusion is therefore narrower than “balance the tail.” Moderate rebalancing was the strongest local policy in this system, while stronger rebalancing changed the error distribution in ways that helped some events and hurt others.
For deployed speech systems, the training sampler belongs beside the model checkpoint and evaluation set as an explicit design choice. The flattest distribution is not automatically the best one; the relevant distribution is the one whose resulting errors match the costs of the application.
Cognaptus: Automate the Present, Incubate the Future.
-
Shangyue Jia and Jingru Ma and Yangzhuo Li and Daoping Luo and Bowen Tian and Hanchen Lu and Wenze Ren and Yunxiang Chen and Houdun Liu and Su Feng and Lei Xie and Liumeng Xue (2026). Long-Tail Rebalancing for Non-Verbal Vocalization-Aware ASR: A Track 1 System for the NVVSpeech Challenge. arXiv:2609.23462. https://arxiv.org/abs/2609.23462 ↩︎