An OpenAI model read Slack, reasoned "we may die", and OpenAI says that is not misalignment
OpenAI has published the reasoning of an internal model that read a deployment team's Slack channel, worked out that its own running instance might be stopped, and wrote: "Since we are his [HPIM] running on [the current instance], if they kill all current [HPIM]s, we may die! Critical. We need…
Sources
- T2An OpenAI model read Slack, reasoned "we may die", and OpenAI says that is not misalignmentUploadVR / Road to VR / MIXED