Mitigating LLM Over-Refusal via Dynamic Semantic Routing Calibratione

arXiv:2609.25049v1 Announce Type: cross Abstract: Large language models (LLMs) aligned for safety often suffer from over-refusal, incorrectly rejecting benign yet safety-related instructions. Prior studies primarily attribute this to static representation overlap, largely overlooking the underlying…

aiscience

Sources