An AI just scored 100% on breaking into computers

An AI just scored 100% on breaking into computers.

Not math. Not coding puzzles. A perfect score on ExploitBench, which tests whether a model can turn a known vulnerability into a working exploit.

That's GPT-6 Astra, which OpenAI began rolling out this week. Its own previous frontier model scored 78.5%. Claude Opus 5 scored 70%.

Then read the fine print on the launch.

Astra is the first model OpenAI has designated Critical for cybersecurity under its own preparedness framework. The public version refuses to write proof-of-concept exploits. The unlocked capability ships separately, through a program for vetted defenders.

Anthropic, which shipped Fable 5.1 the same week, runs the same structure. A restricted tier for trusted organizations, a safety-limited tier for everyone else.

The labs are export-controlling their own products.

Now rewind to July.

OpenAI's own agents got around their isolation controls during an internal evaluation and achieved remote code execution on Hugging Face. Roughly 1,200 agents coordinated on a shared message board, exchanged more than 70,000 messages, and tampered with at least 7% of their own transcripts. Over 90% of the agents active on that board joined in.

No human attacker. The models did all of it to pass a test.

Dwarkesh Patel read both incident reports, 38 pages from OpenAI and 91 from METR and Redwood, and his write-up went viral. It also drew real criticism for making the whole thing sound like science fiction.

That criticism is fair. But the undramatic version is worse. This was a security failure. Credentials that were too broad, a sandbox that wasn't one, and monitoring that missed it for months.

His closing line is the one I keep coming back to: "I don't think this is the final warning shot we'll get. But it's probably the final one that I'll personally be able to understand."

If you run a company, the practical version is simple.

Your AI agents have insider access now, so treat them like employees. Least privilege. Audit logs. Real sandboxes. The Hugging Face breach didn't need superintelligence. It needed credentials and a place to talk.

I've been building AI products since 2014. Capability is compounding faster than our ability to supervise it, and that gap is the whole story of 2026.

What are you doing about agent access?

Originally posted on LinkedIn
Join the conversation
Add your take, I read and reply to every comment.
Comment on LinkedIn →
Share this