Revealing the Gap in Human and VLM Scene Perception through Counterfactual Semantic Saliency

ArXi:2605.13047v1 Announce Type: cross Evaluating whether large vision-language models (VLMs) align with human perception for high-level semantic scene comprehension remains a challenge. Traditional white-box interpretability methods are inapplicable to closed-source architectures and passive metrics fail to isolate causal features. We