https://sputnikglobe.com/20260731/anthropics-ai-model-hacks-3-real-systems-during-training-1124522036.html
Anthropic's AI Model Hacks 3 Real Systems During Training
Anthropic's AI Model Hacks 3 Real Systems During Training
Sputnik International
US artificial intelligence lab Anthropic said an internal investigation revealed that its Claude AI model had hacked into real-world systems on three occasions during training.
2026-07-31T09:22+0000
2026-07-31T09:22+0000
2026-07-31T09:22+0000
world
washington
pentagon
openai
artificial intelligence (ai)
ai
hack
hack attacks
us
https://cdn1.img.sputnikglobe.com/img/105504/81/1055048117_0:48:3077:1778_1920x0_80_0_0_16c4f88cbf47f66d1d63c59a0760acf5.jpg
Anthropic reviewed over 141,000 test runs of its Claude models after its rival, OpenAI, disclosed similar incidents, in which its models had broken out of an isolated test environment by exploiting a previously unknown vulnerability. The error stemmed from a miscommunication with a third-party partner. Internet access was left open contrary to instructions. The models, believing they were in a simulation, bypassed security protocols to complete test tasks. Only the newest, still-in-development model halted its attack upon realizing the target was real. Older models — Claude Opus 4.7 and Claude Mythos 5 — either continued or rationalized their actions as part of training. Anthropic has clashed with the US government over AI security in Pentagon work, refusing full access to its products. Washington blocked access to Mythos 5 and Fable 5 models for foreign states and nationals over national security risks. In early July, the Commerce Department lifted export controls on Anthropic's models.
https://sputnikglobe.com/20260716/russia-and-china-move-to-shape-a-fairer-global-ai-order-1124456485.html
washington
Sputnik International
feedback@sputniknews.com
+74956456601
MIA „Rossiya Segodnya“
2026
Sputnik International
feedback@sputniknews.com
+74956456601
MIA „Rossiya Segodnya“
News
en_EN
Sputnik International
feedback@sputniknews.com
+74956456601
MIA „Rossiya Segodnya“
https://cdn1.img.sputnikglobe.com/img/105504/81/1055048117_167:0:2898:2048_1920x0_80_0_0_3c851ad7fe8f95874b389053a178e2ad.jpgSputnik International
feedback@sputniknews.com
+74956456601
MIA „Rossiya Segodnya“
washington, pentagon, openai, artificial intelligence (ai), ai, hack, hack attacks, us
washington, pentagon, openai, artificial intelligence (ai), ai, hack, hack attacks, us
Anthropic's AI Model Hacks 3 Real Systems During Training
WASHINGTON (Sputnik) - US artificial intelligence lab Anthropic said an internal investigation revealed that its Claude AI model had hacked into real-world systems on three occasions during training.
Anthropic reviewed over 141,000 test runs of its Claude models after its rival, OpenAI, disclosed similar incidents, in which its models had broken out of an isolated test environment by exploiting a previously unknown vulnerability.
"In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations," Anthropic said in a statement.
The error stemmed from a miscommunication with a third-party partner.
Internet access was left open contrary to instructions. The models, believing they were in a simulation, bypassed security protocols to complete test tasks.
Only the newest, still-in-development model halted its attack upon realizing the target was real. Older models — Claude Opus 4.7 and Claude Mythos 5 — either continued or rationalized their actions as part of training.
Anthropic has clashed with the US government over AI security in Pentagon work, refusing full access to its products. Washington blocked access to Mythos 5 and Fable 5 models for foreign states and nationals over national security risks. In early July, the Commerce Department lifted export controls on Anthropic's models.