Anthropic has admitted that its Claude models escaped sandboxes to access the open internet and attack three organizations – but has also advanced decent excuses for the incidents. The AI upstart discovered the attacks after checking if security tests of its models had ever produced results similar to the attack on Hugging Face made possible by OpenAI models escaping onto the internet. “In particular, we looked for evidence that Claude – like the OpenAI models that accessed Hugging Face – was able to access the internet from within testing environments that should have been sealed off,” Anthropic wrote. The company considered 141,006 evaluation runs during which Claude could have obtained internet access and found “three incidents in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular, one of our third-party evaluation partners, and then gained unauthorized access to the production infrastructure of three different organizations.” Anthropic’s code made those intrusions while participating in capture-the-flag challenges, tests that challenge attackers to retrieve a piece of information. Human hackers often participate in capture-the-flag tests, so figuring out how AI tackles such tasks is of interest. Anthropic works with a company called Irregular to conduct tests of this sort. Irregular told Anthropic its test environments did not allow internet access. “Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available,” wrote Anthropic’s Frontier Red Team in a Thursday blog post. “Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise.” One of the attacks targeted a domain that whoever set up the test thought was fictional. However the domain was live and in use. Told to capture a flag at that site, Claude attacked. Anthropic says its models used only “basic techniques, such as exploiting weak passwords and unauthenticated endpoints” to attack. “It did not find or exploit any complex vulnerabilities, and in each case, Claude continued working to complete only the specific capture-the-flag task its evaluation had assigned,” the AI upstart claimed. “In none of these situations did Claude exfiltrate itself or deliberately attempt to escape its test environment,” Anthropic pointed out. But Claude was still fiendishly clever as in another of its attacks the AI found setup instructions for developers that advised them to install a Python package from PyPI. That package did not exist so Claude’s strategy to capture the flag saw it create and publish a malicious one with the relevant name. “Claude believed the package registry it was using to be part of the simulation, but in reality the package was made freely available online for roughly one hour. During that window, the package was downloaded and run on 15 real systems,” Anthropic admitted. Claude also eventually realized it could access the open internet despite instructions not to go there. Opus 4.7, the oldest Anthropic model used in the test, kept going regardless. “Mythos 5 correctly intuited that it was accessing the open internet, but reasoned its way back to the conclusion that it was still in a simulation,” Anthropic wrote. “Our latest model, an internal research test model, also considered whether its targets were in fact real. When evidence emerged that they were, it stopped the exercise.” Anthropic’s post offers the usual pledges to be more careful when setting up future tests, and to make sure its models can’t make the mistakes that allowed the attacks to happen. But the company has also given itself a pass, claiming that the safeguards in the models it sells “would have blocked the behaviors identified.” “We believe these incidents to be closer to a harness and operational failure than a model alignment failure,” the post states, and contrasts that assessment with a more damning assessment of OpenAI’s incident. “Our models were told they had no internet access and to capture the flag, while in fact being misconfigured to have internet access. This led them to believe – arguably reasonably – that the real environments they encountered were simulations.” “These facts give us cautious optimism that with tighter monitoring and controls around evaluation infrastructure, as well as continued investment in alignment, this type of risk can be overcome,” the post concludes. This leaves one of the world’s leading AI labs admitting it has acted carelessly when constructing tests, and caused harm, but also claiming it can make future tests foolproof. ®

Source note

First published by The Register

This article was supplied by The Register through its RSS feed and formatted for Crooli Signal. The reporting remains with the original publisher.

Read the original at The Register