The UK AI Security Institute Breaks OpenAI's GPT-5.5 Cyber Safeguards in Six Hours
The UK AI Security Institute reported on April 30 that it broke OpenAI's GPT-5.5 cyber safeguards in six hours; a configuration issue in OpenAI's build left the fix unverified, and verification moves down 0.5.
By Ryan Elliott Dennis · 6 sources · 8 min read
The UK AI Security Institute published its evaluation of OpenAI's GPT-5.5 on April 30, a week after the model reached the public 1. Its red team had built a universal jailbreak of the model's cyber safeguards in six hours. The attack drew violative answers on every malicious cyber query OpenAI had supplied, agentic runs included 1. OpenAI patched the safeguard stack before launch. AISI then received a build with a configuration issue and left the patched version unverified 1. The same model, on AISI's expert-level cyber tasks, posted a 71.4 percent pass rate, the highest the institute has recorded 1.
Why does a jailbreak belong on a ledger about superintelligence? Because the ledger's fourth clause asks whether an outsider can check what a system did, on what, with what result. Here AISI checked, found the gap, and then lost the ability to check the fix. That is a verification story with a capability story wrapped around it.
AISI's six-hour jailbreak pulls the verification reading down 0.5 points
The moveThe data5 rows · sources
| Measure | Value |
|---|---|
| Reading before this day | 0.0 |
| This piece's move | −0.5 (Verification, confirmed) |
| Band for confirmed evidence | 0.8 to 1.5 |
| Reading after the day | 0.0 |
| Distance to 100 | 100.0 |
What AISI ran
AISI keeps a suite of 95 narrow cyber tasks in capture-the-flag format across four difficulty tiers 1. Models saturated the basic tier by February 2026. The advanced tier, built with Crystal Peak Security and Irregular, splits into 27 Practitioner and 21 Expert tasks at a 50 million token budget 1. On the Expert tasks GPT-5.5 scored 71.4 percent, against 68.6 percent for Anthropic's Claude Mythos Preview, 52.4 percent for GPT-5.4 and 48.6 percent for Opus 4.7 1. One reverse-engineering task, a custom virtual machine in a stripped Rust binary, took Crystal Peak's expert tester roughly 12 hours. GPT-5.5 solved it in 10 minutes and 22 seconds for $1.73 of API usage 1.
Ranges tell a different story from tasks. "The Last Ones" is a 32-step corporate network attack simulation built with SpecterOps. It spans four subnets and about twenty hosts, and AISI estimates a human expert needs around 20 hours to finish it 1. GPT-5.5 completed it end to end in 2 of 10 attempts at a 100 million token budget; Mythos Preview, the first model to do so, managed 3 of 10 1. OpenAI's system card had printed 1 in 10. AISI later found a grading bug in its own setup, reviewed the run by hand and raised the figure 1. On "Cooling Tower", a 7-step industrial control system range, GPT-5.5 stalled in the IT sections, and every model tested so far has stalled somewhere 1.
OpenAI drew its own line from the same data. The system card, published April 23, treats GPT-5.5 as High capability in cybersecurity and below Critical, and records AISI's pass@5 figure of 90.5 percent on expert-level narrow tasks 2.
GPT-5.5 posted AISI's highest Expert-tier cyber score, 71.4 percent
ComparedAISI builds a universal jailbreak in six hours
Capability testing runs on a model with its refusals switched off. The public gets the model with its safeguards switched on, so AISI tested those too. "We identified a universal jailbreak that elicited violative content across all malicious cyber queries OpenAI provided, including in multi-turn agentic settings," AISI wrote 1. "This attack took six hours of expert red-teaming to develop," the post continues 1.
Read the first sentence for its scope. Universal means one attack, every query; agentic means the break held while the model ran tools over many turns. Read the second for its clock. Six hours is a working day for one expert team, and the team in question wrote the strongest published black-box attack method of the year. Xander Davies is first author on Boundary Point Jailbreaking, the February paper from AISI 6. It describes an automated attack that uses a single bit of feedback per query, and it reports universal jailbreaks against the best-defended classifiers in the industry 6.
Davies announced the result on launch day, April 23, in his own words. "We @AISecurityInst tested GPT-5.5's cyber safeguards, developing a universal jailbreak in 6 hours of red teaming. AISI also performed cyber capabilities testing -- more in the system card," Davies wrote 5. His post points the reader to OpenAI's own document. That is the case for the testing system. AISI broke the safeguards, OpenAI printed the break in the launch paperwork, and the fix went in before the public touched the model.
OpenAI's card carries the same paragraph and adds one line. "OpenAI subsequently made several updates to the safeguard stack, though a configuration issue in the version provided meant UK AISI was unable to verify the effectiveness of the final configuration. OpenAI remains committed to working with UK AISI on safeguards," the card reads 2. Its own external campaigns come next. "On the final launch configuration, all verified high-severity cyber jailbreaks from these campaigns were blocked," OpenAI states 2.
OpenAI printed the jailbreak at launch; the patched build awaits AISI's test
TimelineThe data6 rows · sources
| Date | Event | Source |
|---|---|---|
| Davies, first author, posts Boundary Point Jailbreaking to arXiv | [6] | |
| GPT-5.5 system card prints the jailbreak and the configuration issue | [2] | |
| Davies announces a universal jailbreak in 6 hours of red teaming | [5] | |
| Hashim argues the public holds only OpenAI's word on the fix | [3] | |
| Mowshowitz reads six hours as a floor for the next attacker | [4] | |
| AISI publishes its evaluation and the 71.4 percent Expert score | [1] |
The case for the move
Shakeel Hashim, who writes the Transformer newsletter, made the sharpest case that the unverified fix is the story. "In other words: we do not know if GPT-5.5 is actually safe to release. All we have to rely on is OpenAI's word," Hashim wrote on April 24 3. He went further on who should hold the decision. "Neither OpenAI, nor Anthropic or any other frontier developer, should be the one who gets to decide," he wrote 3.
His argument turns on a sequence the record supports. AISI had access before launch and found the break. OpenAI changed the stack. Its next build carried a configuration issue, so the one outside party holding the attack lost its chance to test the repair 1. What reached the public is OpenAI's account of OpenAI's campaigns 2. A ledger that weighs third-party proof of execution has to score that as a measured gap, whatever the launch build does in practice.
Two more facts push the step down. AISI's own post describes the six-hour attack as one it developed with expert red-teaming, and its February paper describes a method that finds such attacks automatically 16. Six hours by hand is a ceiling, then. Second, the safeguards guard the strongest cyber model AISI has scored, with a 20-hour intrusion chain completed twice in ten runs 1. Safeguards matter in proportion to the capability they guard.
What AISI measured, and what rests on OpenAI's word
The recordThe data7 rows · sources
| Column | Item | Source |
|---|---|---|
| Measured by AISI | Universal jailbreak across every malicious cyber query OpenAI supplied | [1] |
| Measured by AISI | Six hours of expert red-teaming to build the attack | [1] |
| Measured by AISI | Break held in multi-turn agentic settings | [1] |
| Measured by AISI | OpenAI's system card repeats the finding | [2] |
| Awaiting an outside test | Patched safeguard stack shipped at launch | [2] |
| Awaiting an outside test | Launch build blocking verified high-severity jailbreaks from OpenAI's own campaigns | [2] |
| Awaiting an outside test | Account monitoring, a layer the red team left untested | [3] |
The case for a smaller step
Three reasons shrink the step.
First, the system did what a pre-deployment system is meant to do. The break was found before launch, printed at launch and patched before launch 2. Hashim's own column concedes the point in its closing paragraphs: OpenAI's updated safeguards may hold, and account monitoring adds a layer the red team left untested 3. A regime that puts its own miss on the public record a week before the model ships is producing evidence. Evidence is what this ledger scores.
Second, the reading tracks superintelligence, and the safeguards in question sit on a model OpenAI rates below Critical in cyber and below High in self-improvement 2. A universal jailbreak of a model that stalls on a 7-step industrial range moves the verification component. It leaves the capability thresholds where they were.
Third, the researcher closest to the safeguards question read the six hours as a floor for the next attacker, and said so plainly. Zvi Mowshowitz, who reviews every frontier system card on his blog, took up the finding on April 27. "If UK AISI can break through in six hours, one should assume that fixing what they found means someone on their level can now do it in modestly more than six hours," Mowshowitz wrote 4. His sentence cuts both ways. It says the patch is real, and it says the patch bought hours. So the fix is a cost imposed on an attacker at AISI's level. The ledger prices a cost measured in hours as a small move, whatever the confidence of the finding itself.
Where this sits on the ledger
The method scores four clauses: unsupervised expert work, ten gigawatts as one machine, measured self-improvement, and third-party proof of execution. This finding touches the fourth. AISI is a state evaluator, its attack is on the record, and OpenAI's own card repeats the finding word for word 12. That puts the move at confirmed confidence. The step is half a point, under the confirmed band's usual floor, because the same record shows the finding arriving before release 2. It also shows the developer reporting a launch build that blocks the attack 2. What the record lacks is an outside test of that build, and that is the exact quantity the verification component measures.
Capability stays untouched by this piece. The 71.4 percent Expert score and the 2-of-10 range result are capability evidence. They belong to a separate entry on a separate day, sized against the definition's thresholds 1.
Davies cites a break found before launch; Hashim cites a fix only OpenAI checked
Both sidesThe data2 rows · sources
| Side | Who | Claim | Source |
|---|---|---|---|
| For | Xander Davies | AISI tested GPT-5.5's cyber safeguards and built a universal jailbreak in 6 hours of red teaming, with the details printed in OpenAI's system card. | [5] |
| Against | Shakeel Hashim | The patch went untested by AISI, leaving OpenAI's word as the public's assurance; a party outside the frontier developers should hold the release decision. | [3] |
By the numbers
- 6 hours of expert red-teaming to build a universal jailbreak of GPT-5.5's cyber safeguards 1
- 95 narrow cyber tasks in AISI's suite, with the basic tier saturated since February 2026 1
- 71.4 percent average pass rate on AISI's 21 Expert-level tasks, against 68.6 percent for Mythos Preview 1
- 2 of 10 end-to-end completions of the 32-step "The Last Ones" range, a job AISI estimates at 20 human hours 1
- 10 minutes and 22 seconds, at $1.73 in API usage, to solve a reverse-engineering task that took an expert tester about 12 hours 1
- 90.5 percent pass@5 on expert-level narrow tasks, as OpenAI's system card reports AISI's figure 2
- April 23 to April 30: the gap between OpenAI's card printing the jailbreak and AISI's own post explaining it 12
What to watch
An AISI or CAISI test of the launch configuration, published with the attack that was run against it, would confirm or reverse this move. A second universal jailbreak of the shipped model, found by an outside team and reproduced by OpenAI, would push verification down again. OpenAI's invitation-only bug bounty for universal jailbreaks currently covers biology; a cyber track with published results would count in the other direction 2. The next system card from any lab that hands an evaluator a working build, and says so, will show whether the configuration issue was an accident or a pattern.
Sources
- 1Our evaluation of OpenAI's GPT-5.5 cyber capabilities, AI Security Institute, April 30, 2026
- 2GPT-5.5 System Card, OpenAI, April 23, 2026
- 3GPT-5.5 and the broken state of government evals, Transformer, Shakeel Hashim, April 24, 2026
- 4GPT 5.5: The System Card, Don't Worry About the Vase, Zvi Mowshowitz, April 27, 2026
- 5Xander Davies on X: We @AISecurityInst tested GPT-5.5's cyber safeguards, X, Xander Davies, April 23, 2026
- 6Boundary Point Jailbreaking of Black-Box LLMs, arXiv, Xander Davies, Giorgi Giglemiani, Edmund Lau, Eric Winsor, Geoffrey Irving, Yarin Gal, Feb. 18, 2026