When Open Artifacts Still Hide the Workflow
A Central Kurdish TTS audit shows why downloadable weights, data, and evaluation code are not enough to establish reproducibility, measurement validity, linguistic scope, or reuse rights.
A Central Kurdish TTS audit shows why downloadable weights, data, and evaluation code are not enough to establish reproducibility, measurement validity, linguistic scope, or reuse rights.
SURE-Voice shows why voice-agent reliability may depend as much on deciding whether audio should reach the model as on what the model generates afterward.
A workflow-generation study suggests that when LLMs understand the requested actions but struggle to serialize dense graph structure, deterministic compilation can improve reliability without immediately requiring a stronger model.
Imperfect Restoration Poisoning shows that data protection can recover image quality without fully restoring learnability, changing how teams should think about the trade-off between perceptual usability and resistance to model training.
AdaptRubric shows that GUI reward quality depends not only on what a verifier can see, but on whether the system has defined the right task-specific success criteria before asking for a judgment.
SpatialBlock suggests that deliberately simplified synthetic tasks can transfer better than more realistic spatial labels when they isolate reusable spatial operations.
A category-level audit shows why mathematical reward verifiers should be governed as configurable evaluation infrastructure rather than summarized by one accuracy number.
LCoT-GV shows that graph-based reasoning verification depends far more on semantic step content than on structural metadata alone—and that verifier performance remains sharply domain-dependent.
A mechanistic study shows that LLMs can use a causal phonological feature to choose allomorphs before the conditioning word is generated—and that direct questioning may miss the mechanism.
A new time-series explanation framework shows that different forecast steps often rely on different parts of history, making one importance map an incomplete account of a multi-step forecast.