What’s Wrong With AI Safety Testing, and How to Fix It
The idea that Anthropic, OpenAI and other AI companies could “embed” researchers from outside safety groups into their operations has been all the rage in recent weeks. But the devil is in the details. AI safety researchers worry that embedded evaluators won’t have the access they need to provide…
Sources
- T1What’s Wrong With AI Safety Testing, and How to Fix ItThe Information (teasers)