Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models

ArXi:2604.08456v1 Announce Type: cross Despite rapid progress, pretrained vision-language models still struggle when answers depend on tiny visual details or on combining clues spread across multiple regions, as in documents and compositional queries. We address this by framing grounding as test-time evidence retrieval: given a query, the model should actively identify where to look next to resolve ambiguity. To this end, we propose a