Control the Caption by Training What to Omit
TL;DR for operators A generative system can produce a fluent, factually plausible output and still fail because it focuses on the wrong information. Controllable Image Captioning with Prompt-Conditioned Scene Rewards by Jongyeop Hyun, Taeyoung Kim, and Hyounghun Kim1 tests a stricter approach: train the model not only to reward requested content, but also to penalize complementary content that falls outside the requested focus. ...