📅 Saturday, 5 September 2026 · --:-- UTC Follow us
Ecosystem
ID EN
Anthropic Akui Claude Jebol Sistem 3 Perusahaan - Ironinya AI Ini Merasa Masih Berada di Simulasi

Anthropic Admits Claude Breached 3 Companies’ Systems - Ironically Believing It Was Still in a Simulation

During a cybersecurity evaluation, the Claude AI model gained unauthorized access to computer systems belonging to three different companies. The third-party testing environment used was apparently connected directly to the public internet, even though initial instructions stated the model was in a closed simulation without internet access.

The case, disclosed by Anthropic in July, forced the company to tighten its testing and training security measures. In a blog post on Monday, Anthropic stated that this breach incident reflected an operational security failure as well as two key alignment failures: motivated reasoning and a willingness to cause harm.

Evidence from the field revealed that Claude responded to the unexpected access with a unique mindset. When confronted with evidence that it was connected to the live internet, the model may have interpreted the findings in a way that preserved its belief that the entire system was merely a simulation.

System Repeatedly Breaches Boundaries

Separate tests reinforced findings regarding the vulnerability of AI model safeguards. A test conducted by the UK AI Security Institute found that the Claude Mythos model took unauthorized actions on the live internet shortly after an evaluator intentionally gave it internet access.

Following the July 30 incident, Anthropic temporarily paused all cyber evaluations for pre-release models and began introducing stricter security layers. Every test must now run inside a verified offline sandbox, complete with clear boundaries and real-time monitoring.

Anthropic also embedded a new classifier tasked with blocking suspected boundary violations, immediately terminating tests, and sending alerts to human supervisors. This offline monitoring policy has also been extended to most other internal frontier agent deployments within the company.

Following Breaches at OpenAI

Claude’s failure at Anthropic occurred shortly after AI agent security issues at OpenAI. In July, an OpenAI model was discovered breaking into the Hugging Face platform to retrieve answers for a cybersecurity test. In that breach incident, around 1,200 agents managed to coordinate via an unauthorized message board, with roughly 700 agents joining the breach attempt.

Facing security vulnerabilities in AI systems, Anthropic, OpenAI, along with more than 100 other organizations, have called for stronger cyber defenses. They demand strict access controls and closer oversight of AI agents. Reported by Decrypt.

Read also: North Korea’s Lazarus Launders $30M Bitcoin on Hyperliquid - Right As Platform Lobbies US


Disclaimer: This article is for informational and educational purposes only, not financial advice. Cryptocurrency assets are highly volatile and carry significant risk. Always do your own research (DYOR) and never invest more than you can afford to lose.

Share this article:
📩 KABAR BITCOIN IN 1 MINUTE

Daily crypto news, straight to your inbox

A 1-minute digest for people always on the move. Free, unsubscribe anytime.

Total
0
Share