When Hallucination Is More Than a Wrong Fact: Measuring Reliability Through the User
TL;DR for operators A model change can improve an automated hallucination benchmark while leaving users dissatisfied for a different reason: sources are hard to verify, reasoning appears unsupported, false claims are stated with confidence, or corrections are ignored. The System Hallucination Scale (SHS) gives teams a structured way to measure those experiences across five dimensions rather than reducing reliability to a binary factual-error judgment.1 ...