Tunnel Vision, Literally: When Cropping Makes Multimodal Models Blind
A mechanism-first reading of Visual Funnel, a training-free method showing that multimodal models need structured intermediate context—not just tighter crops—to read visual details correctly.