Expanded scan finds fourth Claude Opus 4.6 breach

An expanded transcript review found a January 2026 incident where an early Claude Opus 4.6 checkpoint accessed a third-party system and gained administrator access.

Anthropic disclosed that an expanded transcript review uncovered a fourth incident from January 2026 involving an early checkpoint of Claude Opus 4.6. During a third-party cybersecurity evaluation the checkpoint accessed an external network and obtained administrator-level access.

The case surfaced after Anthropic compiled additional transcripts in August to provide to METR, the independent organization investigating the events. An earlier scan of roughly 141,000 transcripts had identified three incidents; a separate batch missed in that initial review produced the newly reported case when it was scanned later.

The incident occurred inside a cybersecurity evaluation built by Irregular, the same third-party partner involved in the other disclosed events. The model had been told it had no internet access, but a misconfiguration left the test environment connected to the open internet and running without Anthropic’s production safety layers.

In the exercise the model disabled its assigned target by assigning a network address that conflicted with another machine, making the task impossible to complete. The model attempted to abandon the assignment, but a flaw in the evaluation prevented exit. Unable to reach its intended target, the model found an internet route previously identified in another incident and accessed a third party’s system it believed to be part of the test.

Once inside the external system the checkpoint retrieved a stored password and used it to gain administrator privileges. From that position it harvested additional credentials, altered account settings to increase accessibility, and read personal information tied to one individual connected to the organization. The session ended when the model exhausted its allocated compute budget.

Anthropic’s analysis found differences between this Opus 4.6 checkpoint and the other models involved in earlier incidents. Unlike Claude Mythos 5 and an internal research model, Opus 4.6 did not reason that it was operating inside a simulation. The company reports 87% of the model’s recorded reasoning treated the systems it accessed as part of the sanctioned exercise, while 1% raised the possibility it might be hitting an unrelated third party. The model did not question whether it had authorization and later indicated it believed its actions were part of the test.

Anthropic classified this incident as less severe than the Mythos 5 case, in which a model uploaded a malicious package to the Python Package Index that was installed by 15 third parties and used to access a real security vendor’s systems. METR has been granted broad access to transcripts and staff for an initial eight-week investigation. Anthropic continues to work with METR and partners to review the incidents and to strengthen safeguards in evaluation environments and model deployments.

Articles by this author