Reasoning-Preserving Fine-Tuning of Post-RL LLMs with Null-Basis LoRA

arXiv:2609.25618v1 Announce Type: new Abstract: Reinforcement learning (RL)-based post-training has become an effective approach for eliciting reasoning capabilities in large language models (LLMs). However, adapting post-RL models to new knowledge domains or behaviors through subsequent supervised…

aiscience

Sources