Beyond Repeated Sampling: Learning Search Policies for LLM Reasoning
arXiv:2609.26704v1 Announce Type: cross Abstract: Large language models increasingly tackle hard reasoning problems by spending more test-time compute, yet the dominant strategy remains naive repeated sampling: draw many independent solutions and hope one is correct. Because such sampling explores…