Before the Test Comes the Question: The LLM Formulation Gap in Analytics
TL;DR for operators An analytics copilot can know statistics and still start from the wrong problem. StatFormBench tests what happens before statistical execution: given an informal request and heterogeneous data, can an LLM determine the statistical problem being asked, select the data objects that actually matter, and assign each one the correct analytical role? Across 1,013 human-reviewed scenarios and 14 LLMs, those abilities do not move together. Gemini 3.1 Pro achieves the highest fine-grained problem-classification accuracy at 72.0, while Claude Opus 4.6 achieves the highest variable-set overlap at 63.2. No evaluated model leads both components. ...