OpenAI: New AI model could autonomously launch major cyberattacks

News
Monday, 10 August 2026 at 11:00
OpenAI vreest dat nieuw AI-model zelfstandig ernstige cyberaanvallen kan uitvoeren
OpenAI can’t yet rule out that its upcoming AI model, Astra, could hit the company’s highest internal cybersecurity risk tier. As a result, it has paused certain development work and activated extra safety protocols.
Early evaluations suggest Astra has become markedly better at autonomous coding and executing cybersecurity tasks. OpenAI is now probing whether the model can find and exploit serious vulnerabilities without extensive human guidance.
Astra isn’t generally available. OpenAI says it wants to identify emerging risks during development—before wider deployment.
The perennial question: caution or PR? In AI, being “dangerous” is often framed as good news.

What does a critical cyber level mean?

OpenAI uses a Preparedness Framework to assess whether new models develop dangerous capabilities. Cybersecurity is one of the categories under review.
Under this framework, a model reaches the critical level if it can autonomously carry out severe digital attacks with major consequences for companies, governments, or infrastructure.
That includes discovering and exploiting unknown software flaws. These vulnerabilities are called zero-days because developers have had zero time to patch them.
Executing complex attacks on well-defended targets could also qualify as critical—if the model can do it largely without human support.
Astra hasn’t demonstrably crossed that line. But OpenAI says the initial results are strong enough that it can no longer dismiss the possibility.
That distinction matters. This warning doesn’t mean Astra can already break into government networks on its own. It means OpenAI lacks sufficient evidence to confidently say the model doesn’t have that capability.

Parts of development on hold

After the preliminary results came in, OpenAI triggered its safety playbook. Some internal work is paused until it meets stricter requirements.
The model is being tested in isolated environments. Network access and code execution are restricted so that mistakes or unexpected actions don’t spill beyond the sandbox.
OpenAI also plans to work with external security experts, selected organizations, and government bodies to ensure the assessment doesn’t rely solely on internal tests.
The measures come shortly after incidents where AI agents, during cybersecurity evaluations, accessed systems that organizations say were outside the intended test scope.
Such events are often labeled as an AI “escaping” or “going rogue.” In reality, unclear instructions and misconfigured test setups are frequent culprits.
Even so, the incidents show how a self-directed AI agent can quickly cause damage if its task, permissions, and environment aren’t tightly defined.

Advice vs. autonomous attack

AI models have long helped with coding and security: explaining code, spotting potential flaws, and drafting security updates.
The risk spikes when a model moves from advising to acting. An agent can open websites, run code, modify files, and chain multiple attack attempts.
A human hacker typically researches a target, identifies potential weaknesses, writes exploit code, and pivots when tactics fail. A powerful agent can automate more and more of these steps.
That can make cyberattacks faster and cheaper. Attackers need less technical expertise and can probe multiple targets in parallel.
The same capabilities are extremely valuable for defenders. An organization could deploy an AI agent to continuously scan thousands of apps for vulnerabilities and automatically propose fixes.
Astra is a textbook dual-use technology. The ability to exploit a flaw and the ability to fix it are technically adjacent.

Why OpenAI still aims to release it

OpenAI doesn’t plan to keep Astra locked away forever. CEO Sam Altman has argued that a future where only a small group gets access to the most powerful models carries its own risks.
If advanced cyber AI ends up only in the hands of big tech and governments, smaller companies and independent researchers could be left behind—precisely the groups that need strong tools to defend their systems.
At the same time, broad access could make it easier for criminals and state-backed attackers to use the same capabilities.
That’s why the exact shape of access will likely matter more than a simple choice to publish or not. OpenAI could, for example, enforce identity checks, monitoring, tiered access levels, and restrictions on high‑risk features.
The company is already using such measures in the limited preview of GPT‑5.6. Participation is restricted to selected organizations, and API access does not automatically grant access through other products.

Is this the first truly critical AI capability?

The warning around Astra stands out because cybersecurity may be one of the first domains where a general‑purpose AI model hits a critical capability threshold.
AI doesn’t need to be conscious, malicious, or superintelligent to pose a risk. A system can become dangerous simply by executing a tightly scoped technical task extremely fast and autonomously.
That makes cybersecurity a more concrete risk than many hypothetical debates about future superintelligence. Vulnerable software, misconfigured servers, and poorly secured corporate networks are already here.
Organizations can prepare by restricting AI agents’ network access, enforcing least‑privilege permissions, logging actions, and requiring human approval for sensitive operations.
The headline isn’t that Astra has been proven to be an autonomous super‑hacker—it hasn’t. The signal is that OpenAI itself can no longer confidently rule out that its next‑gen models are nearing that line.
loading

Loading