Reasoning Instructions Can Break Answer Decoding in Vision--Language Models

arXiv:2609.29278v1 Announce Type: cross Abstract: Chain-of-thought (CoT) instructions can distort multiple-choice VLM evaluation when a scorer appends a reasoning cue but reads answer-label logits before the model generates any rationale. We call this CoT-prefix scoring. On ScienceQA, Qwen2.5-VL-7B…

science

Sources

Reasoning Instructions Can Break Answer Decoding in Vision--Language Models · TechNews