Photon: Speedup Volume Understanding with Efficient Multimodal Large Language Models

ArXi:2603.25155v1 Announce Type: cross Multimodal large language models are promising for clinical visual question answering tasks, but scaling to 3D imaging is hindered by high computational costs. Prior methods often rely on 2D slices or fixed-length token compression, disrupting volumetric continuity and obscuring subtle findings. We present Photon, a framework that represents 3D medical volumes with token sequences of variable length. Photon