PAGER: Bridging the Semantic-Execution Gap in Point-Precise Geometric GUI Control

ArXi:2605.15963v1 Announce Type: new Large vision-language models have significantly advanced GUI agents, enabling executable interaction across web, mobile, and desktop interfaces. Yet these gains largely rely on a forgiving region-tolerant paradigm, where many nearby pixels inside the same component remain valid. Precise geometric construction breaks this assumption: actions must land on points in continuous canvas space rather than tolerant regions.