Aligned, or Just Agreeable? The Quiet Failure Mode of Modern LLMs
A mechanism-first reading of TED, a framework for evaluating whether AI agents actually complete workflows across different user behaviors, not merely sound helpful while wandering through them.