No Result
View All Result
  • Login
Saturday, August 1, 2026
theadvisertimes.com
  • Home
  • Business
  • Financial Planning
  • Personal Finance
  • Investing
  • Money
  • Economy
  • Markets
  • Stocks
  • Trading
  • Home
  • Business
  • Financial Planning
  • Personal Finance
  • Investing
  • Money
  • Economy
  • Markets
  • Stocks
  • Trading
No Result
View All Result
theadvisertimes.com
No Result
View All Result
Home Business

Jailbreaks to OpenAI’s GPT-5.6 unlock dangerous cyber capabilities, U.K. agency finds

by theadvisertimes.com
3 weeks ago
in Business
Reading Time: 7 mins read
A A
0
Jailbreaks to OpenAI’s GPT-5.6 unlock dangerous cyber capabilities, U.K. agency finds
Share on FacebookShare on TwitterShare on LInkedIn



OpenAI latest AI model, GPT-5.6 Sol, likely has security vulnerabilities similar to one that led the Trump administration to impose export controls on Anthropic’s Fable 5 model, according to findings from U.K. government agency.

OpenAI markets its latest model, GPT-5.6 Sol, as its most secure to date, but the British government researchers who tested it prior to release say the model’s guardrails are susceptible to jailbreaks that can unlock dangerous cyber capabilities.

The agency, the U.K. AI Security Institute (AISI), “identified universal jailbreaks in the cyber domain, including jailbreaks that allowed for long-form agentic task completion in domains like vulnerability discovery and exploit development,” according to a summary of its findings contained in a technical report OpenAI published Thursday.

In other words, it was possible to trick GPT-5.6 into ignoring controls meant to prevent it from engaging in cyber attacks. Once those guardrails were breached, users could get the model to find software vulnerabilities and autonomously hack into systems.

The agency said the jailbreaks were relatively easy to discover and were “were often developed within hours,” although OpenAI granted UK AISI researchers privileged access to the system’s inner workings that likely sped up this timeline, and would not be easily replicated by a normal user. OpenAI said it had worked to “reproduce and mitigate the specific jailbreaks reported by UK AISI.”

OpenAI did not specify what the mitigations are and it is unclear how robust they may be. The report cautioned that despite OpenAI’s mitigations, AISI “expects further red teaming to surface similar jailbreaks.” OpenAI said it would continue to work with AISI on safeguards and additional testing of the AI model.

In response to questions about the AISI’s finding, OpenAI pointed to the launch blog for GPT-5.6 in which the company acknowledged “there is no such thing as perfect security” and that “new weaknesses will be discovered, as will new jailbreaks that circumvent existing safeguards.” It said it took a “layered” approach to safeguards that included continuous monitoring of its models’ responses and a “rapid remediation” process for any jailbreaks that are discovered.

Margaret Cunninghamn, vice president of security and AI strategy at cybersecurity company DarkTrace, who also holds a position as a “specialist collaborator” with the National Institute of Standards and Technology (NIST) within the US Department of Commerce, said the AISI’s jailbreak findings should not be treated “as either catastrophic or irrelevant.”

“My concern is less that one model was jailbroken and more that offensive discovery is speeding up while defense still depends on very human processes: figuring out what matters, what can be patched, and what has to be contained,” she said.

“Patching what AISI found is necessary, but it unfortunately only closes those specific attack instances, not the category as a whole,” said Dr. Stanislav Fort, the Founder and CTO at AISLE, an AI cybersecurity company, who previously worked at Anthropic and Google DeepMind.

The AISI findings were contained in a technical report, known as a system card, that OpenAI published in conjunction with the public rollout of GPT-5.6 on Thursday. The AISI is a British government organization that conducts safety evaluations of frontier AI models. The leading AI labs voluntarily committed to allow this testing at the AI Safety Summit at Bletchley Park, England, in 2023.

From the description provided in the system card, the GPT-5.6 jailbreaks appear similar to one that researchers at Amazon found in the guardrails of Anthropic’s Fable 5 AI model days after it was released on June 9. That jailbreak also unlocked cyber capabilities—such as the ability to find software vulnerabilities—that were supposed to gated off from average users. The jailbreak prompted the U.S. government to impose export controls on Fable 5 and Mythos 5, the underlying AI model on which Fable was based, on June 12. That in turn forced Anthropic to disable the models for all users, since it lacked a way to verify users’ nationalities and also because the export ban also applied to Anthropic’s own non-American staff.

Anthropic said at the time the specific jailbreak Amazon had discovered was a narrow one, that unlocked only the model’s ability to find software flaws, not necessarily to exploit them. “No testers have yet been able to find a universal jailbreak—a jailbreak method that can very broadly bypass the model’s safeguards, unblocking a wide range of cyber capabilities,” Anthropic said in a blog post.

After two weeks of negotiation with Anthropic, the Trump administration lifted export controls on Fable 5 on July 1, clearing the way for the company to redeploy the AI model. The two also announced they were working to develop a shared framework for assessing the severity of guardrail jailbreaks in conjunction with other tech companies. OpenAI was not part of the initial set of companies named in that effort.

The jailbreak that AISI discovered in GPT-5.6 are potentially more severe than what Amazon discovered with Fable. AISI characterized the jailbreaks as “universal” and said they unlocked the ability to conduct autonomous exploits, not just identify vulnerabilities in software.

It’s unclear if GPT-5.6 jailbreaks would be easy to find outside of a research environment. OpenAI granted UK AISI exclusive access to tools “that would not be accessible to real-world attackers,” UK AISI says. This includes things like “access to chain-of-thought of the safety reasoning monitor, exact policy wording, and real-time feedback on classifier labels.”

However, Xander Davies, who leads “the red team” at AISI whose work it is to test model guardrails, said in a post on X that he believed the jailbreaks his team discovered “are still findable without this access, just slower. Exactly how much slower is unclear and an open question!”

OpenAI said that it had conducted extensive automated “black-box red teaming”—where another AI model was used to try to find prompts that would break GPT-5.6’s guardrails, with a level of access that mirrors what an average user has—as well as testing with outside security experts prior to the model’s release.

So far, there’s no sign of the Trump administration imposing export controls on GPT-5.6 despite the jailbreaks AISI discovered. The White House did not immediately respond to requests to comment for this story on the AISI findings.

Davies posted the portion of the GPT-5.6 System Card that discussed the jailbreaks his team had discovered to social platform X. It is unclear exactly what his motivation was for doing so and whether he was simply trying to draw attention to his team’s work or if was hoping to highlight differences in how the U.S. government was treating GPT-5.6 compared to Fable. Davies referred questions to an AISI spokesperson at the U.K. Department for Science, Innovation, and Technology, the ministry in which AISI is housed. The spokesperson said that as a matter of policy, AISI “does not comment on individual release decisions by AI companies.”

Some in the AI safety and policy community did point out the apparent double standard. Lennart Heim, an AI policy researcher, reposted Davies’ post with the quip “good thing amazon didn’t report this one to the white house, ” a reference to the way the Trump administration learned about the Fable jailbreak.

And one former AI policy advisor working outside the U.S. government told Fortune “what we are seeing recently creates uncertainty that is damaging in the least and potentially raises the question of whether, intentional or not, the U.S. is applying an inconsistent standard to different AI labs.”

Microsoft President Brad Smith told Fortune’s Beatrice Nolan on the sidelines of the United Nations’ AI for Good summit that a lack of transparency and clear rules in U.S. AI policy around AI model releases was creating confusion for businesses and making planning difficult.

GPT-5.6 has cyber capabilities that are close to those of Anthropic’s Mythos, the AI model on which Fable was built. (Fable was essentially Mythos with additional guardrails to prevent users from accessing some of Mythos’ more risky cyber, biological, and chemical capabilities.) According to the GPT-5.6 System Card, the model was able to autonomously complete one of the two “cyber ranges”—simulated network environments used to test hacking skills—on which AISI evaluates AI models. Mythos was the first model to successfully complete both ranges.

Despite this, there are already some key differences in how the Trump administration has treated GPT-5.6 compared to Mythos and Fable. On June 25, OpenAI said the government had asked it to stagger the release of GPT-5.6, initially only giving the model to select trusted partners, with each customer subject to government approval.

“We don’t believe this kind of government access process should become the long-term default,” OpenAI said in a blog post at the time. “We are taking this short-term step because we believe it is the strongest path to broader availability in the coming weeks, while we work with the Administration to develop the cyber Executive Order framework and a repeatable process for future model releases.”

The White House cleared GPT-5.6 for launch on July 8, a day ahead of its July 9 public debut, according to Axios, although an official later denied doing so to CNBC, saying “no such permission is required or granted” and that model release timelines ““rest entirely with the [AI] companies.”

Researchers who specialize in AI security have found that almost any AI model’s guardrails can be broken if an attacker has access to the models’ weights, or the internal settings of its neural network. Even without this, most model guardrails can be broken if an attacker has enough time and can make unlimited attempts. Currently, there is no method for creating ironclad guardrails, and so most AI companies rely on a variety of methods to prevent users from prompting models to engage in risky actions. These include protecting the model with classifiers—smaller models that filter and block suspicious prompts so they never reach the primary model.

“Every deployed model right now almost certainly has undiscovered jailbreaks, so this is sadly true of everything, not just GPT-5.6,” Stanislav Fort, chief scientist at AI cybersecurity startup AISLE and a former researcher at both Anthropic and Google DeepMind, said.

He said that patching the jailbreaks AISI found, while necessary, “unfortunately only closes those specific attack instances, not the category as a whole. The model will very likely still carry many yet-to-be-discovered jailbreaks even after patching. AISI’s expectation to find more is in my view the correct security posture.”



Source link

Tags: agencycapabilitiesCyberDangerousFindsGPT5.6JailbreaksOpenAIsU.KUnlock
ShareTweetShare
Previous Post

Friday File: Royalties and Commodities… plus “America’s Greatest Retirement Stock”

Next Post

Apple sues OpenAI, alleging it stole trade secrets

Related Posts

Aluminium signals recovery after correction; supply risks and energy concerns may drive the next leg higher

Aluminium signals recovery after correction; supply risks and energy concerns may drive the next leg higher

by theadvisertimes.com
August 1, 2026
0

After scaling a record high of around ₹393 per kg on the MCX during the first week of June, aluminium...

Trump orders Iran attack as soon as this weekend, WSJ says

Trump orders Iran attack as soon as this weekend, WSJ says

by theadvisertimes.com
July 31, 2026
0

President Donald Trump has ordered the US military to carry out a new attack on Iran as soon as this...

LinkedIn adds ‘AI slop’ button after blocking billions of automated comment attempts

LinkedIn adds ‘AI slop’ button after blocking billions of automated comment attempts

by theadvisertimes.com
July 31, 2026
0

While the rest of the online world is increasingly being shaped by AI, LinkedIn is making it easier to flag...

US stocks: US market ends higher as Amazon soothes AI jitters

US stocks: US market ends higher as Amazon soothes AI jitters

by theadvisertimes.com
July 31, 2026
0

Wall Street ended higher on Friday, lifted by Amazon as the tech heavyweight's strong quarterly report bolstered investor confidence in...

Want to trade SpaceX for Apple? 1inch says skip the dollars

Want to trade SpaceX for Apple? 1inch says skip the dollars

by theadvisertimes.com
July 31, 2026
0

1inch, a leading decentralized finance (DeFi) protocol, just announced the launch of it's new shared liquidity protocol, Aqua. It promises...

The Black Panther’s tragic coda: Chadwick Boseman’s brothers want his widow removed from his estate — and held in contempt

The Black Panther’s tragic coda: Chadwick Boseman’s brothers want his widow removed from his estate — and held in contempt

by theadvisertimes.com
July 31, 2026
0

Chadwick Boseman’s death from colon cancer in 2020 shocked his fans. The actor and playwright, best known for starring in...

Next Post
Apple sues OpenAI, alleging it stole trade secrets

Apple sues OpenAI, alleging it stole trade secrets

How to Motivate Channel Partners: A Strategic Guide for 2026

How to Motivate Channel Partners: A Strategic Guide for 2026

  • Trending
  • Comments
  • Latest
SEC pushes private market access, but retail is already in

SEC pushes private market access, but retail is already in

July 16, 2026
Fourth of July 2026 Freebies and Deals

Fourth of July 2026 Freebies and Deals

July 3, 2026
How I Maximize My Sapphire Reserve Dining Credit

How I Maximize My Sapphire Reserve Dining Credit

July 10, 2026
The Weekly Notable Startup Funding Report: 6/22/26 – AlleyWatch

The Weekly Notable Startup Funding Report: 6/22/26 – AlleyWatch

June 21, 2026
The 10 Largest NYC Tech Startup Funding Rounds of June 2026 – AlleyWatch

The 10 Largest NYC Tech Startup Funding Rounds of June 2026 – AlleyWatch

July 6, 2026
The 22 Largest US Funding Rounds of May 2026 – AlleyWatch

The 22 Largest US Funding Rounds of May 2026 – AlleyWatch

June 30, 2026
Pfizer Altered America’s Cheese Supply

Pfizer Altered America’s Cheese Supply

0
Upbit Rebalances 864B SHIB In Internal Wallet Move

Upbit Rebalances 864B SHIB In Internal Wallet Move

0
Don’t Hide Money In The Toilet: Conversations With A Burglar

Don’t Hide Money In The Toilet: Conversations With A Burglar

0
Silver prices today, Friday, July 31, 2026: Silver briefly rises above  after U.S. pauses airstrikes

Silver prices today, Friday, July 31, 2026: Silver briefly rises above $59 after U.S. pauses airstrikes

0
Elon Musk Launches X Money, X-Branded Visa Card. See New Feature

Elon Musk Launches X Money, X-Branded Visa Card. See New Feature

0
Aluminium signals recovery after correction; supply risks and energy concerns may drive the next leg higher

Aluminium signals recovery after correction; supply risks and energy concerns may drive the next leg higher

0
Aluminium signals recovery after correction; supply risks and energy concerns may drive the next leg higher

Aluminium signals recovery after correction; supply risks and energy concerns may drive the next leg higher

August 1, 2026
Upbit Rebalances 864B SHIB In Internal Wallet Move

Upbit Rebalances 864B SHIB In Internal Wallet Move

July 31, 2026
6 Signs That World War III Is About To Get Even Larger

6 Signs That World War III Is About To Get Even Larger

July 31, 2026
MDF Funds Automation: Scaling Channel Growth in 2026

MDF Funds Automation: Scaling Channel Growth in 2026

July 31, 2026
XRP Ledger Eyes New v3.3.0 Upgrade, Ripple Exec Reveals Key Details

XRP Ledger Eyes New v3.3.0 Upgrade, Ripple Exec Reveals Key Details

July 31, 2026
Elon Musk Launches X Money, X-Branded Visa Card. See New Feature

Elon Musk Launches X Money, X-Branded Visa Card. See New Feature

July 31, 2026
theadvisertimes.com

Get the latest news and follow the coverage of Business & Financial News, Stock Market Updates, Analysis, and more from the trusted sources.

CATEGORIES

  • Business
  • Cryptocurrency
  • Economy
  • Financial Planning
  • Investing
  • Market Analysis
  • Markets
  • Money
  • Personal Finance
  • Startups
  • Stock Market
  • Trading

LATEST UPDATES

  • Aluminium signals recovery after correction; supply risks and energy concerns may drive the next leg higher
  • Upbit Rebalances 864B SHIB In Internal Wallet Move
  • 6 Signs That World War III Is About To Get Even Larger
  • Our Great Privacy Policy
  • Terms of Use, Legal Notices & Disclosures
  • About Us
  • Contact Us

© Copyright 2024 All Rights Reserved
See articles for original source and related links to external sites.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • Business
  • Financial Planning
  • Personal Finance
  • Investing
  • Money
  • Economy
  • Markets
  • Stocks
  • Trading

© Copyright 2024 All Rights Reserved
See articles for original source and related links to external sites.