TL;DR for operators

When an anomaly detector misses events that matter, teams often look first at the model, the feature selector, or the score-combination rule. The results in LLM-Generated Feature Pools for Time Series Anomaly Detection1 point to an earlier constraint: what the detector is allowed to measure.

In matched experiments on the univariate TSB-AD-U benchmark, changing the candidate feature pool spans 0.226 VUS-PR, compared with 0.031 across feature-selection strategies even after adding a hindsight oracle, and 0.096 across the full score-aggregation grid. A 10-feature hand-crafted pool reaches 0.529, while a stronger LLM-generated pool reaches 0.569 on average. Combining generated and hand-crafted features raises the corresponding result to 0.588.

The LLM is not the anomaly detector. It is used offline to inspect labeled example windows and write candidate NumPy feature functions. Production scoring remains a simple median/MAD statistical procedure. The more defensible deployment pattern is therefore hybrid: let generation expand the vocabulary, retain expert features as prior knowledge, and let the existing selector choose among both.

The largest movement comes before feature selection

Suppose a detector already has a reasonable scoring rule and a feature-selection procedure. There are still two distinct failure modes. It may be searching the available features poorly, or the available features may simply fail to describe the anomaly pattern.

The paper separates these possibilities unusually cleanly. Under a fixed detector, benchmark split, windowing method, and metric, the authors vary three design axes independently. Changing the candidate pool produces a 0.226 VUS-PR range, from 0.362 to 0.588. Changing score aggregation produces a 0.096 range. Changing among greedy forward selection, top-k, mRMR-style selection, and an unattainable per-domain oracle produces only 0.031.

That comparison does not show that selection is unnecessary. Using the entire hand-crafted pool without selection scores 0.435, versus 0.529 after per-domain selection. Selection materially helps. What appears second-order is the choice among several reasonable selectors once a useful pool already exists.

For an engineering team, this distinction matters. If repeated selector refinements deliver small gains, the next investigation may need to move upstream: which anomaly signatures are missing from the representation?

More features are not the same as better coverage

A natural response is to enlarge the library. The benchmark result gives a reason to be more specific.

The generic catch22 library contains 22 features but scores 0.362 under the matched pipeline. The paper’s smaller, 10-feature hand-crafted pool reaches 0.529. For the strongest tested generator, a 10-feature task-conditioned LLM pool reaches 0.569 ± 0.025 across four generation seeds.

The relevant distinction is therefore not feature count. It is whether the pool contains transformations that expose useful domain-specific differences between ordinary and anomalous windows.

The generation procedure is deliberately bounded. For each domain, a multimodal LLM receives five normal and five anomalous tuning-window plots and is instructed to emit exactly ten deterministic NumPy feature functions. Generated code is sandbox-validated, with invalid runs retried. Evaluation labels and evaluation-series examples are not supplied to the generator.

After generation, the LLM leaves the system. Sliding windows are transformed by the selected features, and each feature is scored relative to the median and median absolute deviation observed within the same target series. The reported pipeline then averages the resulting feature scores and maps overlapping window scores back to points.

This architecture separates representation discovery from runtime detection. The paper reports average domain-level generation-call costs ranging from roughly $0.004 to $0.070 for the tested providers, but those calls are design-time operations rather than per-event inference.

Generated features add coverage, but they also add new blind spots

The generated pools are not uniformly superior. That is one of the more useful results for deployment design.

The two Gemini generators average 0.569 and 0.540 respectively, while the tested Claude generator averages 0.497. The 10-feature hand-crafted pool remains at 0.529. Generator choice therefore changes whether generated-only features beat the expert baseline.

Domain results make the reason visible. With the strongest generator, the generated pool scores 0.513 on human activity versus 0.238 for the hand-crafted pool. But the hand-crafted pool remains stronger in finance, medical, and traffic. Neither source consistently contains the best representation.

That makes replacement a weaker design than combination.

Evidence What it supports Operational interpretation Boundary
Hand-crafted selected pool: 0.529 A small expert pool can support a strong simple detector Existing domain features remain valuable One univariate benchmark
Best generated-only pool: 0.569 ± 0.025 Task-conditioned generation can discover competitive features LLMs can expand the representation at design time Performance varies across generators and seeds
Best union: 0.588 ± 0.016 Joint selection can exploit complementary pools Preserve expert features while adding generated candidates Does not prove every domain improves
Union beats generated-only in 12/12 matched generator-seed pairs Complementarity is consistent across the tested runs Expert features provide protection against generator blind spots Only three generators and four seeds each

Across those 12 exact generator-seed comparisons, adding the hand-crafted features improves VUS-PR by between 0.006 and 0.061, with a mean gain of 0.033. Both the reported sign test and Wilcoxon signed-rank test give $p=4.9\times10^{-4}$.

The union is not best in every individual domain. For example, generated-only remains higher on facility and web services for the strongest generator, while hand-crafted remains higher on finance and traffic. The consistency appears at the full evaluation level: selecting from both pools repeatedly reduces the cost of relying on either source alone.

For operators, representation becomes an auditable design decision

The paper directly shows a benchmark result. Cognaptus’ business inference is narrower: teams operating domain-specific sensor or operational time series may want to treat feature coverage as an explicit engineering object rather than a fixed preprocessing choice.

The affected decision is where to spend the next unit of model-development effort. When scoring and subset selection are already competent, another selection heuristic may have less upside than identifying anomaly types for which the current feature vocabulary has no informative statistic.

The paper also suggests a relatively controlled way to test that hypothesis. Generate candidate features from a small labeled tuning set, validate their code, retain the established expert library, and evaluate selection over the combined pool on held-out series. Because the generator is removed from the production scoring loop, deployment does not inherit the same runtime latency and provider dependency that an LLM-in-the-loop detector would.

This is not evidence that feature engineering can now be delegated wholesale to a language model. The generator results themselves argue against that interpretation. The stronger operational use is as a proposal mechanism whose outputs compete with, rather than erase, established features.

The boundary is univariate benchmark evidence

The empirical design is strong for the comparison it makes: the study uses a fixed 48-series tuning split and 350-series evaluation split across nine TSB-AD-U domains, holds the downstream detector constant, evaluates three generators over four seeds, and pairs union-versus-generated comparisons within each seed.

Its external boundary is equally clear. The study covers univariate anomaly detection only. Traffic has five evaluation series and finance eight, limiting domain-specific inference. The generator study is not exhaustive. Published leaderboard systems were taken from reported results rather than rerun through the same implementation, so the union’s 0.588 is best read as reaching the roughly 0.59 level of the strongest pretrained leaderboard entry, not as establishing superiority over it.

The next useful question is therefore not whether LLM-generated features have solved anomaly detection. It is whether representation coverage can be measured and improved systematically across other datasets, multivariate settings, and iterative generation procedures.

Within the evidence available here, one conclusion is already operationally consequential: once scoring and selection are reasonable, a detector may be constrained less by how intelligently it searches its features than by whether those features describe the anomaly at all.

Cognaptus: Automate the Present, Incubate the Future.


  1. Youssef Attia El Hili and Malik Tiomoko and Corinne Ancourt (2026). LLM-Generated Feature Pools for Time Series Anomaly Detection. arXiv:2609.21801. https://arxiv.org/abs/2609.21801 ↩︎