Safety Nudges: User-Facing Interventions for Real-Time AI Risk Awareness

arXiv:2609.26865v2 Announce Type: replace-cross Abstract: Conversational AI systems can pose safety risks to their users such as hallucination, sycophancy, overconfidence, and anthropomorphism, but these risks are difficult for users to detect during everyday use. We introduce Safety Nudges, a…

science

Sources