1 briefs tagged #mllms.
The benchmark measures how multimodal models respond when scene text conflicts with visual evidence.