Refusal without Discrimination: What Encoded Prompts Do to Safety-Trained Models
arXiv:2609.26176v1 Announce Type: cross Abstract: Encoded-prompt attacks are evaluated almost entirely on their harmful arm: a benchmark sends obfuscated harmful requests and reports how often the model complied. We show that this arm carries almost no information about the model under test. Across…