Minimally Invasive Steering of Language Models
arXiv:2609.30218v1 Announce Type: cross Abstract: Pre-logit steering adapts a frozen language model to a test-time reward by adding vectors to its final hidden states. Unregularized reward optimization can substantially alter the output distribution and degrade generation quality. We propose…