When Seeing Overrides Knowing: Visual Dominance and Deferral-Based Method for Personalized Safety in VLMs
Studies personalized safety in vision-language models, where a generally reasonable response may be unsafe for a specific user whose medical, emotional, or situational context is hidden. Introduces MPS-Bench, a benchmark of 5,181 scenarios from 584 real-world images across 12 high-risk domains, and finds frontier VLMs almost always answer directly (86–99%) instead of seeking missing context. Identifies visual dominance—visual affect enters the text stream in early layers and suppresses textual risk signals during fusion—and proposes PRISM, a lightweight input monitor using bidirectional cross-modal modulation to predict when a query requires deferral, achieving 0.978 AUC and dominating the safety-utility Pareto frontier.