TL;DR for operators
A public model can come with weights, a dataset, configuration files, inference code, and evaluation tooling while still leaving another team unable to reconstruct what was trained, what was held out, how the released checkpoint was selected, or how broadly its performance claims apply.
Hiwa Asadpour’s audit of a Central Kurdish text-to-speech release1 makes that problem concrete. The strongest evidence is not that the underlying research is invalid; the paper being audited is described as comparatively cautious and well documented. The problem appears in the handoff from publication to public artifacts. Configuration settings conflict with the reported run, evaluation assets are missing, the released dataset does not mark the original test utterances, and the model card states a stronger performance conclusion than the source paper supports.
For model-governance, product, QA, and procurement teams, release readiness should therefore be checked at the artifact level: claims, configurations, data splits, preprocessing, evaluators, linguistic scope, checkpoints, dependencies, and licenses all need to agree closely enough for the intended reuse.
A complete-looking release can still lose the training record
Suppose a team is considering a public speech model for a product prototype or benchmark. The expected due-diligence path is familiar: read the paper, inspect the model card, download the weights and dataset, run the evaluation code, and decide whether the system is mature enough to reuse.
The Central Kurdish case shows where that procedure can fail.
The released configuration contains an eight-GPU comment and an eleven-epoch setting. The associated research paper, however, reports training on one RTX8000 GPU and presents learning curves extending to roughly 275,000 steps. The audit interprets the inconsistent settings as likely inherited upstream defaults rather than evidence that the published training description is false.
That distinction is significant. A configuration file that executes successfully is not necessarily a record of the configuration that produced the released model.
The same problem appears in the data. The publication says that 500 utterances per speaker were held out for testing, yet the public corpus does not identify those utterances. All 17,671 reported utterances appear in a single released training split. A downstream researcher can therefore train on the original evaluation material without realizing it.
Other missing pieces include the exact training command, parts of the data-manifest conversion and audio preparation process, the rule used to select the released checkpoints, and the seven sets of sentences used in the subjective listening evaluation. The public evaluation package itself is functional—the audit reproduced its bundled four-utterance smoke test—but the Kurdish recognition checkpoint required to recompute the reported intelligibility results was not openly recoverable from the examined materials.
Reproducibility here is not primarily a question of whether files exist. It is whether the operational decisions connecting those files survived publication.
The headline ranking is much weaker than the model card implies
The public model card describes the female audiobook system as having the highest overall subjective score and presents it as suitable for general-purpose applications. The source paper is more restrained, describing the audiobook-trained systems as competitive with studio-trained data and its best subjective result as only marginally better.
The numbers explain the difference in tone.
| System | Natural speech MOS | Synthetic MOS | Synthetic loss |
|---|---|---|---|
| audiobook-F | 4.42 | 4.08 | 0.34 |
| audiobook-M | 4.34 | 4.01 | 0.33 |
| studio-M | 4.40 | 4.06 | 0.34 |
The gap between the highest two synthetic scores is 0.02 MOS: 4.08 versus 4.06. Reported confidence intervals for per-category MOS are roughly ±0.14 to ±0.19. Meanwhile, every synthetic system sits about 0.33–0.34 points below its corresponding natural recording.
The audit does not establish that the female audiobook system is worse. It shows that the available evidence does not support turning a 0.02 descriptive difference into a broad superiority claim.
There is another identification problem. Each training condition uses a different individual speaker. Differences between audiobook and studio systems therefore also combine speaker identity, recording conditions, register, and other condition-specific factors. The comparison cannot isolate an independent “audiobook versus studio” effect.
The evaluator has its own domain preference
An automatic speech recognizer can behave differently across recording conditions even before synthetic speech enters the comparison. That makes its baseline behavior part of the interpretation of any intelligibility score.
In this release, a fine-tuned Seamless-family recognizer contributed to transcript preparation, while the same recognizer family was also used to evaluate synthesis through character error rate, or CER.
The natural-speech controls are revealing:
| Source condition | CER on genuine human speech |
|---|---|
| audiobook-F | 0.01 ± 0.01 |
| audiobook-M | 0.02 ± 0.01 |
| studio-M | 0.09 ± 0.08 |
The studio recordings therefore begin with a substantially higher evaluator error rate before the TTS model is assessed.
This does not make CER unusable, nor does the paper estimate how much of each synthetic score is caused by evaluator bias. It establishes a narrower point: the metric is partly conditioned by the measurement instrument and its domain fit.
For QA teams, the practical inference is to evaluate the evaluator. Natural-speech baselines and, where feasible, independently trained measurement systems can reveal whether an apparently portable metric changes with recording domain.
Preprocessing determines what linguistic variation survives
The audit’s third contribution is less visible in benchmark tables. Before speech synthesis, text must be transformed into a pronunciation representation. Every transformation can retain linguistic alternatives or remove them.
In this pipeline, the narrowing happens repeatedly. The source material is prepared read speech. Transcripts pass through automatic and manual processing. Fixed normalization rules standardize orthography. Grapheme-to-phoneme conversion resolves pronunciation choices. A singleOutputPerWord setting keeps the first candidate pronunciation. Syllable markers are discarded. The resulting phoneme sequence must then fit a fixed 2,545-entry vocabulary inherited from pretrained non-Kurdish resources.
None of these choices is presented as inherently unreasonable. Low-resource systems often need strong reuse of pretrained infrastructure. The governance issue is that these choices jointly define the model’s linguistic scope while remaining much less visible than the model weights.
The audit also identifies a deterministic inference defect: uninterrupted digit strings of nine or more digits are replaced with a NUM placeholder, but the transformation is not reversed, and the placeholder’s opening angle bracket is absent from the released vocabulary. The paper does not claim that all numerical input is broken; the verified problem is this specific code path.
Release readiness needs claims, scope, and rights to travel together
Most of the identified defects are repairable without retraining: mark the held-out test split, publish evaluation sentences, reconcile configuration defaults, document preprocessing and checkpoint selection, expose vocabulary failures, and record speaker and listener variety metadata.
Some constraints are harder to fix after adoption.
The speakers’ Central Kurdish varieties and regions are not reported, nor are the dialect backgrounds of the 88 listeners who produced 3,101 subjective ratings. The experiment can therefore compare the three released read-speech systems, but it cannot establish that their voices—or a 0.02 MOS ranking—generalize across Central Kurdish varieties.
Reuse rights also need to be traced before integration. Restrictions on the release are substantially inherited from upstream sources: non-commercial constraints come from pretrained weights, while audiobook permissions contribute a no-derivatives condition under CC BY-NC-ND. The audit does not offer a legal opinion on which model transformations legally count as derivatives. It does show why a team should resolve whether corrected, fine-tuned, dialect-adapted, or deployment-converted models can be redistributed before building a workflow around them.
What this case supports—and what it does not
The evidence is strongest where the audit can directly inspect artifacts: configuration inconsistencies, missing split identifiers, unavailable evaluation dependencies, code paths, documentation gaps, and discrepancies between paper language and model-card claims.
The broader implications require more restraint. This is one Central Kurdish release, not an estimate of how frequently comparable problems occur across low-resource speech projects. The study does not independently evaluate perceptual audio quality. It also cannot measure dialect effects because speaker and listener variety metadata are missing, and several concerns about transfer to other Kurdish varieties remain reasoned reuse risks rather than tested outcomes.
That boundary does not weaken the operational value of the case. It defines it. A model release should be reviewed as a connected system of evidence and artifacts rather than as a folder containing enough files to run inference. For teams deciding whether to adopt public AI resources, the relevant question is not only whether the model works. It is whether its claims, data lineage, evaluation machinery, linguistic assumptions, training record, and redistribution rights remain recoverable when the research leaves the paper.
Cognaptus: Automate the Present, Incubate the Future.
-
Hiwa Asadpour (2026). Assessing the Reusability of Public Speech Resources for Low-Resource Languages: A Central Kurdish Case Study. arXiv:2609.11246. https://arxiv.org/abs/2609.11246 ↩︎