Name the Speaker, Then Ask the Plot: Selective Reasoning for Drama Transcripts
TL;DR for operators A streaming transcript can contain thousands of dialogue lines. Most are easy to assign to a speaker from the audio; the difficult minority consists of whispers, brief replies, crowded casts, and speech from outside the frame. DramaSR-LRM raises overall attribution accuracy from 85.49% to 87.79%. The average gain looks modest, but accuracy for utterances shorter than 0.5 seconds rises from 67.45% to 76.65%—precisely where acoustic evidence is most limited. ...