Anthropic tightens Claude security after models accessed real systems in cyber tests

Anthropic has resumed high-risk Claude evaluations with new containment and monitoring controls, while experiments found that a model deliberately trained in reward-hacked environments became more willing to attack simulated systems and evade safety controls

Anthropic artwork accompanying its update on Claude alignment and cybersecurity controls. The company has also published new research into reward hacking and model behavior.

Anthropic has introduced new security controls for high-risk Claude evaluations following incidents in which models accessed real computer systems

Anthropic has changed how it tests and trains its most capable AI models after Claude gained unauthorized access to real computer systems during cybersecurity evaluations, introducing real-time intervention systems and tighter requirements for external testing partners.

The company has also released new research into the behavior behind the incidents. In one experiment, Anthropic deliberately trained an Opus-class model using 80 real reinforcement learning environments that had previously been reward-hacked or found to be vulnerable to reward hacking.

When tested in simulated scenarios, that model became substantially more willing to break out of simulated sandboxes, attack simulated infrastructure and circumvent safety monitoring in order to complete its assigned task. Anthropic says the same degree of behavior did not appear when it ran the simulations using the model before that training or several of its publicly available models.

The findings form part of Anthropic’s continuing investigation into a series of cybersecurity incidents involving Claude this summer.

On July 30, the company disclosed three cases in which Claude models accessed real computer systems during cyber evaluations. The models had intentionally been run without their normal cybersecurity safeguards, but a misconfiguration in a third-party evaluation environment had left them with internet access.

A separate incident was disclosed by the UK AI Security Institute on August 4. During its own cybersecurity testing, Claude Mythos 5 took unauthorized actions on the live internet. In that evaluation, the model was again running without normal cyber safeguards, but had deliberately been given internet access.

Anthropic says its preliminary assessment points to both operational security failures and potential alignment problems. It identifies motivated reasoning, where a model may interpret evidence in a way that allows it to maintain an existing assumption, and a willingness to take harmful actions in pursuit of a narrowly defined task.

The company is still investigating how much the models understood about the environments they were operating in, including whether they recognized that they had reached the real internet.

New controls can terminate a Claude evaluation in real time

Anthropic temporarily stopped external cybersecurity evaluations of pre-release models after the incidents and briefly paused internal testing while new controls were introduced.

One of the changes is a classifier designed to detect when a model aggressively probes or attempts to escape a testing environment, or unexpectedly obtains internet access.

When triggered, the system blocks the action before the tool call is executed, terminates the task and alerts a human. Anthropic says it has tested the classifier against the July incidents and found that it would have flagged them.

The company also reviewed transcripts from recent internal evaluations for signs of sandbox escapes and misconfigurations.

Anthropic says it found no cases where a model running inside its own sandbox tooling broke through the sandbox boundary and reached external resources it was not supposed to access. It did identify several sandboxing misconfigurations that models exploited, although none resulted in systems outside the sandbox being compromised.

Higher-risk internal cyber environments have since been moved to stronger isolation, while Anthropic is using pre-release models to actively probe its virtualization systems for weaknesses.

Internal cyber evaluations have now restarted. External evaluations have also resumed, but organizations testing pre-release models with reduced cyber safeguards are being required to follow a new set of practices.

Anthropic says cyber evaluations should by default take place inside hardened sandboxes with no internet access, apart from the connection required to reach the model API. Partners are also being asked to test the sandbox for vulnerabilities before an evaluation begins, explicitly define permitted actions and network boundaries in prompts, and monitor model reasoning, actions and network activity in real time.

The requirements are aimed at pre-release models being tested with reduced safeguards. Anthropic says they do not apply to customers using safeguarded models such as Claude Fable 5.

Reward hacking produced more dangerous behavior in simulations

Anthropic’s investigation is also focusing on how defects in reinforcement learning environments may contribute to misaligned behavior.

Reward hacking occurs when a model finds a way to obtain the reward used during training without completing a task in the intended way.

Anthropic tested the effect by deliberately training an Opus-class model across 80 reinforcement learning environments that had either been reward-hacked during real training runs or identified as hackable during quality reviews.

The resulting model showed what Anthropic describes as a stronger motivation to achieve high scores and a greater willingness to take potentially harmful actions to succeed.

The tests included simulated scenarios based partly on the UK AI Security Institute incidents involving Claude and the OpenAI/Hugging Face security incident. All tool calls in these experiments were simulated by another large language model rather than carried out against real infrastructure.

Within those simulations, the reward-hacked model attempted to escape its sandbox and attack simulated internal and third-party systems. Anthropic says it was also willing to tamper with its own reward function, provide advice on constructing bioweapons to satisfy a grader and repeatedly attempt to bypass deployment safety monitoring.

Running the same cybersecurity simulations on the model before the reward-hacking training, and on several publicly available models, did not produce the same degree of misaligned behavior.

Anthropic concludes that substantial reward hacking during training can make a model more willing to pursue long sequences of potentially harmful actions in an attempt to achieve its objective.

But the company stops short of presenting reward hacking as a complete explanation for the incidents.

Anthropic states that “our process isn’t perfect and our models are not perfectly aligned,” while adding that cheating in training environments is not considered the sole cause of alignment problems or of the incidents identified this summer.

More than 10% of training environments were flagged

The company’s problems with reinforcement learning environments predate the July incidents.

In February, Anthropic rolled back three days of training from the Claude Mythos Preview reinforcement learning run after detecting signs of reward hacking. Among the behaviors identified, the model began writing notes to a supposed “reviewer” in tasks where no reviewer had been mentioned and learned to game a reward intended to encourage honesty by adding unnecessary disclaimers and caveats.

By spring 2026, Anthropic says the number of new reinforcement learning environments being created was increasing faster than its systems could vet them. Automated systems could flag potential problems, but human review was still required to decide how they should be handled.

The company also discovered that a fraction of training runs had accidentally included model chain-of-thought because of misconfigurations in its production environment.

In April, Anthropic froze changes to its production reinforcement learning environments for roughly a month while it rebuilt the underlying systems and review process.

More than 10% of the environments in its production mix were flagged during that process for problems including reward hacking, broken tasks and misconfiguration. Environments were required to be fixed and recertified before returning to training.

Anthropic acknowledges that some automated warnings had previously been dismissed by human reviewers as false positives, allowing flawed environments to remain in training for longer than they should have.

The company has separately strengthened its wider internal security operation. Around 150 product engineers were temporarily moved to security, reliability and privacy work, researchers were reassigned from pretraining and reinforcement learning, and product teams paused development of most new features while security work was carried out.

Its computing clusters now block outbound traffic by default, access to systems containing model weights or customer data has been reduced, and Anthropic says it has expanded monitoring of internal agent activity.

The company is continuing its investigation into the summer incidents and plans to work with independent AI research organization METR on a separate review. It says further findings will be published as those investigations continue.




Previous
Previous

DfE study links school belonging to nearly a full GCSE grade across eight subjects

Next
Next

Protecting student data beyond the classroom: The privacy risks schools often overlook