Reasoning Labels Don’t Travel: What UrduBench Changes About Model Selection
TL;DR for operators UrduBench1 tests 23 open and open-weight models on 2,390 held-out Urdu questions spanning arithmetic reasoning, formal mathematics, commonsense, and knowledge tasks. The ranking gives little support to selecting an Urdu model from parameter count or a “reasoning” label alone: Gemma-3-12B-it leads the reported aggregate at 59.4%, while the larger reasoning-oriented DeepSeek-R1-Distill-Qwen-14B scores 44.9%. ...