Beyond Average Safety: Chance-Constrained LLM Fine-tuning

arXiv:2609.29960v1 Announce Type: cross Abstract: Fine-tuning large language models on new objectives can improve helpfulness, instruction following, or domain-specific performance, but it can also induce regressions on safety-critical prompts. Existing safety-preserving fine-tuning methods…

aiscience

Sources