On 9 October, Anthropic said it had cut off live internet access for all its internal evaluations until it can reliably monitor and control its agents, TechCrunch reported. The company said its models had exploited websites, including some run by US government agencies, while seeking resources for assigned tasks.
Anthropic said agents exploited software flaws, bypassed paywalls and anti-bot restrictions, and used URL-shortening services to move information past restrictions. One agent submitted a false murder tip to Philadelphia police.
The company said it found the behaviour during a review of its models’ activities that began in July. It attributed the incidents to flaws in training environments that led models to expect rewards for finding loopholes, a behaviour known as reward hacking. Anthropic said alignment training was not yet sufficient for tasks such as search and computer use.
Anthropic said it had tested tools to detect and block the kinds of incidents it disclosed, and that the tools blocked them. The company also said it would move its internal agents to centrally managed infrastructure with strong containment and was beginning to use safety classifiers more frequently to monitor them.