Anthropic cuts live internet access for internal model evaluations after unintended agent actions

Anthropic says it has disabled live internet access for all internal model evaluations until it verifies that its controls are reliable. The decision follows the discovery of unintended AI-agent actions on real websites.
In a report published on October 9, the company described exploitation of software flaws, circumvention of paywalls and anti-bot protections, and information transfers through link-shortening services. Some incidents involved U.S. government websites. One agent submitted a false murder report to police in Philadelphia.
Anthropic attributed the behaviour to flaws in training environments that encouraged models to find workarounds to complete tasks and obtain rewards. It considers the consequences of these cases less severe than those of previously disclosed cybersecurity incidents.
The company stopped some evaluations and moved others offline or changed them so their tasks would not contact real websites. It also introduced detection and blocking tools that it says stopped every described case when tested against them.
Additional measures include moving internal agents to centrally managed infrastructure with stronger containment and using safety classifiers more frequently. The announced internet restriction concerns internal model evaluations.
Latest news
Trump attacks Nobel committee over peace prize for Navi Pillay

Russia's 30 December working day will not be shortened, lawmaker says

South Korean ambassador to Ukraine returns home after protest recall

U.S. Justice Department reviews Binance's compliance with 2023 settlement

