Cover image

The 70B Model May Belong Upstream

TL;DR for operators If a team has a small labelled seed set and a large volume of multilingual text to classify, keeping the strongest LLM in every inference request may not be the best allocation of compute. Pecher et al. find that smaller models using examples generated by LLaMA-3 70B can exceed that same 70B model used directly as a zero-shot classifier with roughly 50 synthetic examples in aggregated language groups.1 ...

September 2, 2026 · 8 min · Zelina
Cover image

English Looks Ready. Amharic Says Otherwise: What ADAGE Exposes in Multilingual Evaluation

TL;DR for operators A multilingual model can look ready on an English reasoning benchmark and still perform close to chance in a strategically important native language. In the reported zero-shot evaluation, Gemma 3 27B scores 83.0% on English ePiC, 70.1% on Arabic CAPR, 41.3% on Amharic CAPR, and 86.0% on Japanese CAPR. ...

August 16, 2026 · 7 min · Zelina