The Safety Leaderboard Has Conditions: Choose Moderators by Harm, Context, and Cost
TL;DR for operators A moderation team rarely needs a model that is merely “best at safety.” It needs a model that catches the relevant harms at the point where moderation occurs, without creating unacceptable false positives or latency. A large benchmark by Afshin Orojlooyjadid and Hitesh Patel compares 53 specialized moderators and general-purpose language models across 11 public safety datasets.1 Its strongest operational finding is not a new overall winner. It is that the ranking changes with the safety problem and with what the moderator is allowed to inspect. ...