Evidence Before Elaboration: Why Product-Video Extraction Gains Start With Retrieval
TL;DR for operators When a multimodal model cannot reliably recover product details from video, the first engineering question should not automatically be how to make it reason harder. In a controlled seven-category ablation, retrieving visually similar products and their known attributes raises average micro-F1 from 33.4 to 48.3. Adding interleaved chain-of-thought without retrieval moves the score only to 33.5. The strongest standalone lever in this experiment is better evidence. ...