Who Sets the Score? H-Bench Reframes AI Benchmarking as a Sociotechnical System
TL;DR for operators Two teams can test the same model, observe the same metric values, and still reach different deployment decisions because they assign different costs to reliability, latency, interpretability, fairness, or other constraints. Most benchmarks handle that difference outside the scoring system: someone chooses the metrics and weights, publishes the resulting scorecard, and the benchmark remains comparatively fixed. ...