Q-CueGraph: Query-Conditioned Visual Evidence Graphs for Multimodal Reasoning

arXiv:2608.04452v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) can miss fine details in a full image that they recognize in a closer view. Recovering this evidence requires deciding where to look and how much surrounding context to retain. We present Q-CueGraph…

aiscience

Sources

Q-CueGraph: Query-Conditioned Visual Evidence Graphs for Multimodal Reasoning · TechNews