The Detector Can’t See What Its Features Don’t Describe
TL;DR for operators When an anomaly detector misses events that matter, teams often look first at the model, the feature selector, or the score-combination rule. The results in LLM-Generated Feature Pools for Time Series Anomaly Detection1 point to an earlier constraint: what the detector is allowed to measure. In matched experiments on the univariate TSB-AD-U benchmark, changing the candidate feature pool spans 0.226 VUS-PR, compared with 0.031 across feature-selection strategies even after adding a hindsight oracle, and 0.096 across the full score-aggregation grid. A 10-feature hand-crafted pool reaches 0.529, while a stronger LLM-generated pool reaches 0.569 on average. Combining generated and hand-crafted features raises the corresponding result to 0.588. ...