Minimally Invasive Steering of Language Models

arXiv:2609.30218v1 Announce Type: cross Abstract: Pre-logit steering adapts a frozen language model to a test-time reward by adding vectors to its final hidden states. Unregularized reward optimization can substantially alter the output distribution and degrade generation quality. We propose…

science

Sources