The 70B Model May Belong Upstream
TL;DR for operators If a team has a small labelled seed set and a large volume of multilingual text to classify, keeping the strongest LLM in every inference request may not be the best allocation of compute. Pecher et al. find that smaller models using examples generated by LLaMA-3 70B can exceed that same 70B model used directly as a zero-shot classifier with roughly 50 synthetic examples in aggregated language groups.1 ...