PoEM: Predicting RL Outcomes from Existing Policies
arXiv:2609.30226v1 Announce Type: cross Abstract: Foundation models are post-trained with reinforcement learning (RL) to maximize specific rewards, such as human alignment, correctness, or instruction following. This post-training process is computationally intensive, sometimes unstable, and has to…