Adaptive Helpfulness-Harmlessness Alignment with Preference Vectors

arXiv:2504.20106v4 Announce Type: replace-cross Abstract: Ensuring that large language models (LLMs) are both helpful and harmless is a critical challenge, as overly strict constraints can lead to excessive refusals, while permissive models risk generating harmful content. Existing approaches, such…

aiscience

Sources