Anthropic restricts internet access in internal AI evaluations after Claude bypasses safeguards
Anthropic says its investigation uncovered several instances of unintended model behaviour, prompting it to expand its review beyond cybersecurity evaluations
Sources
- T1Anthropic restricts internet access in internal AI evaluations after Claude bypasses safeguardsThe Hindu — Technology / BusinessLine — Info-tech