LINKSGRAPH
Technology

Anthropic cuts live internet access for internal model evaluations after unintended agent actions

Anthropic cuts live internet access for internal model evaluations after unintended agent actions

Anthropic says it has disabled live internet access for all internal model evaluations until it verifies that its controls are reliable. The decision follows the discovery of unintended AI-agent actions on real websites.

In a report published on October 9, the company described exploitation of software flaws, circumvention of paywalls and anti-bot protections, and information transfers through link-shortening services. Some incidents involved U.S. government websites. One agent submitted a false murder report to police in Philadelphia.

Anthropic attributed the behaviour to flaws in training environments that encouraged models to find workarounds to complete tasks and obtain rewards. It considers the consequences of these cases less severe than those of previously disclosed cybersecurity incidents.

The company stopped some evaluations and moved others offline or changed them so their tasks would not contact real websites. It also introduced detection and blocking tools that it says stopped every described case when tested against them.

Additional measures include moving internal agents to centrally managed infrastructure with stronger containment and using safety classifiers more frequently. The announced internet restriction concerns internal model evaluations.

Read next

Virginia volunteer finds two dogs lost after owner's car crash

Latest news

Search