Cover image

The Label Is Too Late: EduBehaviors Makes Conversation Coding Auditable

TL;DR for operators A conversation-analysis system can return the wrong label for at least two different reasons: it may have misunderstood what happened in the conversation, or it may have correctly recognized the behavior and then mapped it incorrectly to the construct being measured. A single LLM label hides that distinction. EduBehaviors1 separates the two steps. The system first records human-readable yes-or-no judgments about observable behaviors, then uses a separate rule or classifier to convert those judgments into the final label. In the paper’s Teacher TalkMoves evaluation, the word-only baseline reached macro-F1 0.339 and Cohen’s kappa 0.328. The strongest reported assertion-based configuration reached macro-F1 0.673 with kappa 0.688. ...

October 8, 2026 · 7 min · Zelina