Cover image

The Generalization Stack: Why HAR Robustness Is a Pipeline Property

TL;DR for operators A smartphone activity-recognition model can perform well during development and then degrade when deployed on a different dataset, user population, or sensor position. The natural response is to search for a domain-generalization technique that performs best across these changes. A 410,400-experiment benchmark suggests that this is the wrong unit of comparison. No individual objective, initialization strategy, or architectural modification wins consistently. Yet compatible combinations can produce materially larger gains: the best aggregate joint configurations improve accuracy by about 2.9 percentage points under cross-dataset shift and 4.9 points under cross-position shift. ...

September 28, 2026 · 8 min · Zelina
Cover image

HAROOD: When Benchmarks Grow Up and Models Stop Cheating

A wearable model can look brilliant in the lab and embarrass itself on Monday morning. The user changes. The watch slides down the wrist. A sensor is mounted on the chest instead of the pocket. The same person walks differently after fatigue, injury, aging, or simply because life has the terrible habit of not matching the training set. Human Activity Recognition, or HAR, has always lived with this problem. It turns sensor streams from accelerometers, gyroscopes, EMG, ECG, and other wearable or ambient devices into labels such as walking, running, sitting, cycling, or stress state. It is useful precisely because it moves into the real world. That is also where benchmark accuracy goes to die. ...

December 12, 2025 · 20 min · Zelina