Prefilling the Reasoning Channel: Output-Prefix Attacks on Reasoning LLMs

arXiv:2609.29775v1 Announce Type: cross Abstract: Large Language Models (LLMs) consume and produce a single sequence of text; hence, if text can be added to the beginning of the LLM's response, i.e., an output prefix, then all subsequent tokens will be conditioned on it. This output-prefix attack…

aiscience

Sources