Cover image

The Generalization Stack: Why HAR Robustness Is a Pipeline Property

TL;DR for operators A smartphone activity-recognition model can perform well during development and then degrade when deployed on a different dataset, user population, or sensor position. The natural response is to search for a domain-generalization technique that performs best across these changes. A 410,400-experiment benchmark suggests that this is the wrong unit of comparison. No individual objective, initialization strategy, or architectural modification wins consistently. Yet compatible combinations can produce materially larger gains: the best aggregate joint configurations improve accuracy by about 2.9 percentage points under cross-dataset shift and 4.9 points under cross-position shift. ...

September 28, 2026 · 8 min · Zelina