TL;DR for operators

An analytics copilot can know statistics and still start from the wrong problem.

StatFormBench tests what happens before statistical execution: given an informal request and heterogeneous data, can an LLM determine the statistical problem being asked, select the data objects that actually matter, and assign each one the correct analytical role? Across 1,013 human-reviewed scenarios and 14 LLMs, those abilities do not move together. Gemini 3.1 Pro achieves the highest fine-grained problem-classification accuracy at 72.0, while Claude Opus 4.6 achieves the highest variable-set overlap at 63.2. No evaluated model leads both components.

For analytics systems, the practical implication is architectural. Do not treat “understands statistics” as one gate before automated analysis. Validate the inferred problem type, selected variables, and their roles separately before allowing the system to choose methods or generate code. The benchmark supports that separation strongly within its test setting, but it does not measure live consulting conversations, where requests may be more incomplete and ambiguous.

The first analytics failure can happen before any analysis runs

Consider a business user who uploads a dataset and asks, “Which factors are associated with customers leaving, and does the pattern differ across regions?”

A capable analytics system still has several decisions to make before it can run anything. Is the request primarily predictive, inferential, comparative, or something else? Which columns are relevant? Which variable is the response? Which are predictors? Should identifiers, derived fields, or operational metadata be excluded?

If those decisions are wrong, flawless downstream execution only makes the wrong analysis more reproducible.

This upstream problem is what Wang and colleagues call statistical problem formulation in Benchmarking Language Models for Statistical Problem Formulation.1 The paper separates formulation into two components: identifying the statistical problem type, and identifying the required data objects together with the role each plays in the requested analysis.

That distinction moves the evaluation target. Many statistical or data-science benchmarks begin after the analytical objective has already been specified. StatFormBench asks whether the model can infer that objective in the first place.

StatFormBench makes formulation measurable

The benchmark contains 1,013 human-reviewed scenarios: 629 derived from five statistics textbooks and 384 from a case-based data-science library. It organizes statistical problems into 20 coarse categories and 85 fine-grained categories.

Each example contains background information, an informal request, available data, a reference problem category, and a reference set of variables required to solve the request.

The variable task is narrower than simply naming columns. A relevant data object is represented by its identity, description, values, and its question-dependent role: predictor, response, both, or not applicable. The same field can therefore play different analytical roles under different questions.

This matters because formulation errors are not interchangeable. A system may recognize that the user wants a regression-style analysis while still choosing the wrong response variable. Or it may identify the correct variables while confusing two statistically adjacent problem subtypes.

The benchmark is designed to expose those differences rather than collapsing them into one score.

The strongest classifier is not the strongest variable selector

The clearest result is the lack of a single formulation leader.

Gemini 3.1 Pro records the highest fine-grained classification accuracy, at 72.0. Claude Opus 4.6 leads variable-set recovery, with a Jaccard coefficient of 63.2.

The Jaccard measure asks how much the model-selected variable set overlaps the human reference set, penalizing both missing required variables and adding unnecessary ones. It therefore captures a different failure surface from problem classification.

Formulation component Best reported result What it measures
Fine-grained problem classification Gemini 3.1 Pro — 72.0 Whether the model identifies the specific statistical problem subtype
Variable-set overlap Claude Opus 4.6 — 63.2 Whether the selected data objects match the reference set

The divergence is more informative than either leaderboard value alone. It shows that statistical formulation is not behaving like one latent capability that simply rises with general model strength.

A second result reinforces that interpretation: every evaluated model performs better on coarse-grained classification than on fine-grained classification. Models more reliably recover broad analytical intent than distinguish among closely related statistical subtypes.

That is a plausible production failure mode. A copilot may correctly understand that a user is asking about association, prediction, comparison, or estimation while still selecting an analysis whose assumptions or target differ from what the user actually intended.

Data structure changes which part of formulation looks easy

Performance also shifts with benchmark source.

Representative models classify textbook-derived problems substantially better than case-library problems, but variable identification shows the opposite pattern: the case-library scenarios are easier for extracting relevant variables.

The paper links this asymmetry to differences in how the data are presented. Textbook problems can provide clearer cues about the intended statistical task while embedding quantities in prose or representations that make variable extraction harder. Structured case-library data can make candidate variables more explicit even when the analytical objective is less neatly signposted.

For an operator, this means a single benchmark average can hide deployment-relevant variation. A model evaluated on clean analytical prompts may look strong at recognizing problem types while behaving differently when it receives business tables containing many directly accessible but irrelevant fields.

The input representation is part of the formulation problem.

Better instructions help classification more than they help formulation as a whole

The paper also tests whether prompting can close the gap.

Providing definitions of the benchmark categories improves fine-grained classification accuracy by 1.9 to 3.8 points across the three tested models. That is evidence that some classification errors come from unclear or closely spaced label semantics.

The improvement does not transfer cleanly to the rest of the pipeline. Effects on variable-set recovery and role identification are small and mixed. Few-shot prompting is likewise modest or inconsistent.

This is not evidence that prompting is ineffective in general. It is evidence that better descriptions of the label space mainly address the part of the task that depends on interpreting that label space.

Selecting the correct data and assigning their analytical roles requires something else: connecting the request to the structure and meaning of the supplied data.

A scoring check suggests variable errors are not mainly an evaluation artifact

Variable matching in StatFormBench uses exact equality of serialized values, which creates a legitimate measurement concern. A model can represent an equivalent variable differently and receive no match.

The authors test this on 60 difficult textbook cases for Claude Opus 4.6, GPT-5.5, and Gemini 3.1 Pro. Expert review attributes 11.7%, 10.0%, and 10.0% of those difficult cases, respectively, to matching-induced errors.

This is best read as a robustness check on the metric rather than a second substantive finding. Exact matching does underestimate some semantically correct outputs, but the validation indicates that the matching rule explains only a minority of the examined failures. Low variable-selection performance cannot simply be dismissed as a scoring artifact.

Analytics copilots need formulation checkpoints before execution

The paper directly establishes benchmark-level capability differences. The workflow implication is a Cognaptus inference.

For an analytics copilot, formulation should be represented as an explicit sequence of decisions rather than hidden inside one prompt-response step:

User request + data → problem classification → variable selection → role validation → method selection → execution

Each transition creates a checkpoint.

A system can ask the user to confirm an inferred target when confidence is low. It can test whether required variables are missing. It can flag selected fields that appear irrelevant to the stated question. It can separately verify predictor and response assignments before generating statistical procedures.

The same decomposition also affects model procurement and routing. A model chosen because it performs well on statistical reasoning or code generation has not thereby demonstrated reliable formulation. Organizations evaluating analytics models can test classification, variable recovery, and role assignment separately, then route different stages if the performance tradeoff justifies the added complexity.

Prompt engineering still has a place, particularly where a company maintains its own taxonomy of analytical requests. StatFormBench suggests that supplying those definitions may improve classification. It does not support assuming that the same intervention will resolve data relevance or role errors.

The benchmark supports workflow separation, not a production failure rate

StatFormBench is comparatively strong as a controlled benchmark: all 1,013 final samples were manually reviewed, 14 models were evaluated across the full set, and the paper examines category, source, and prompting differences.

Its main deployment boundary is external validity.

The scenarios come from solved textbooks and a curated case library, not transcripts of live statistical consultations. Real users may leave goals unstated, revise requirements during conversation, misuse statistical language, provide undocumented columns, or ask questions that do not map cleanly to the benchmark taxonomy.

The category distribution is also imbalanced, so results for rare problem types are less precise. The textbook material originated in English while the case-library material originated in Chinese before translation, leaving possible cross-lingual effects unresolved.

Those limits make the benchmark unsuitable for estimating how often a production analytics assistant will formulate a problem correctly. They do not remove the central engineering signal: formulation contains multiple separable decisions, and current models show materially different behavior across them.

Statistical automation should begin with formulation, not code generation

The common picture of an AI analytics system starts too late. It asks whether the model can select a method, write the code, run the analysis, and explain the result.

StatFormBench moves the starting point upstream.

Before method selection, the system has to infer what the user is actually asking. Before execution, it has to determine which data belong in the problem and what those data mean within that question. The benchmark shows that current LLMs are uneven across precisely those steps, and that modest prompting improvements do not collapse them into one solved capability.

For analytics automation, reliability therefore begins before the first line of statistical code is generated.

Cognaptus: Automate the Present, Incubate the Future.


  1. Chen Wang and Junzhe Zhao and Xin Cong and Wanlu Deng and Ke Deng (2026). Benchmarking Language Models for Statistical Problem Formulation. arXiv:2609.01982. https://arxiv.org/abs/2609.01982 ↩︎