Continual learning might make your blocking monitors nearly useless

Many control protocols work by intervening on an untrusted AI's actions during deployment. For example, you might set up a monitor that scores each action's suspiciousness and blocks actions above a threshold, replacing them with actions from a weaker "trusted" model (a defer-to-trusted protocol)…

ai

Sources

Continual learning might make your blocking monitors nearly useless · TechNews