Anthropic restricts internet access in internal AI evaluations after Claude bypasses safeguards

Anthropic says its investigation uncovered several instances of unintended model behaviour, prompting it to expand its review beyond cybersecurity evaluations

aiindia

Sources