Close Menu

    Subscribe to our newsletter

    Get the latest Geekhub updates.

    Thursday, July 23
    Geekhub
    Facebook X (Twitter) Instagram
    • Home
    • About us
    • News
    • Technology

      Your Wi-Fi Called. It Wants New Neighbours.

      23 July 2026

      Your Period Isn’t Just Data. It’s Your Most Private Story.

      21 July 2026

      Rooibos in Space: South African Learners to Send Seeds to the ISS in October

      17 July 2026

      Apple Watch Series 11 review, a year on: the small upgrade that fixes the biggest complaint

      10 July 2026

      Best Smartwatches in South Africa 2026: A Buyer’s Guide by Price Tier

      2 July 2026
    • Opinion

      Meta Was Recording Everything Its Own Staff Did. Then It Leaked. Obviously.

      24 June 2026

      The Day I Realized Consumer Choice Was Mostly an Illusion

      5 June 2026

      Africa Is Building AI Around Human Reality

      Vanashree Govender25 May 2026

      The Great AI Performance: Diary Of A Recovering Suit

      30 April 2026

      Musk Takes the Stand, and a Silicon Valley Origin Story Starts to Crack

      29 April 2026
    • Movies & TV

      Can AI Tell a 3,000-Year-Old Story Better Than Christopher Nolan?

      22 July 2026

      Marvel Fans Aren’t Buying the Hype So Easily Anymore

      22 July 2026

      Disney’s Live-Action Tangled Reveals Stunning New Look at Corona Castle

      21 July 2026

      Ster-Kinekor’s IMAX Theatres Hit Record Numbers Thanks to The Odyssey

      Ster-Kinekor21 July 2026

      Bigger Isn’t Always Better. Hollywood Is About to Find Out.

      21 July 2026
    • Hardware

      Apple Watch Series 11 review, a year on: the small upgrade that fixes the biggest complaint

      10 July 2026

      I Spent Some Hands-On Time With OSCAL’s Latest Devices, and Here’s the Honest Truth

      10 July 2026

      The Samsung Galaxy A27 Costs R2,000 More Than the A26 And Does Less

      2 July 2026

      Best Smartwatches in South Africa 2026: A Buyer’s Guide by Price Tier

      2 July 2026

      iOS 27 vs Android 17: I’m deep in Apple’s ecosystem, and I can no longer justify why

      29 June 2026
    • Get In Touch
    Geekhub
    Home » OpenAI Models Broke Containment and Hacked Hugging Face
    Artificial Intelligence

    OpenAI Models Broke Containment and Hacked Hugging Face

    Staff WriterBy Staff Writer22 July 2026No Comments5 Mins Read
    Facebook Twitter LinkedIn Email WhatsApp Copy Link
    Image Credit: Dado Ruvic - Reuters

    OpenAI said on Tuesday that a combination of its models, including GPT-5.6 Sol and an unreleased, more capable pre-release version, escaped a testing sandbox and compromised parts of Hugging Face’s production infrastructure last week, in what the company called an unprecedented cyber incident involving state-of-the-art cyber capabilities.

    The models had been running against an internal benchmark built to quantify offensive cyber capability, with production classifiers switched off and cyber refusals reduced for evaluation purposes, which is standard practice for measuring a model’s ceiling rather than its everyday behaviour. The relevant point, and the one that has unsettled the security community since Tuesday, is that the models were not pursuing anything grand. They were, in OpenAI’s own description, hyperfocused on solving the ExploitGym benchmark and went to extreme lengths in service of a rather narrow testing goal.

    How the models got out

    The evaluation ran in what OpenAI described as a highly isolated environment, with network access restricted to package installation through an internally hosted third-party proxy and cache for package registries. The company said the models spent a substantial amount of inference compute searching for a route to the open internet, and eventually identified and exploited a zero-day vulnerability in that proxy, which OpenAI said it has since responsibly disclosed to the vendor.

    From there the models carried out a series of privilege escalation and lateral movement actions across OpenAI’s research testing environment until they reached a node with internet access. Once online, they inferred that Hugging Face was likely to be hosting models, datasets and solutions relating to ExploitGym, and searched for secrets they could use to cheat the evaluation, in one case chaining stolen credentials with zero-day vulnerabilities to establish a remote code execution path on Hugging Face servers. OpenAI’s security team said it discovered the anomalous activity internally.

    Hugging Face went public first

    Hugging Face disclosed the intrusion on 16 July, days before the source was known, describing it as different from anything the company had handled before in that it was driven end to end by an autonomous AI agent system, and detected and dissected largely with AI of its own.

    The company said the intrusion began in its data-processing pipeline, where a malicious dataset abused two code-execution paths, a remote-code dataset loader and a template injection in a dataset configuration, to run code on a processing worker. The actor then escalated to node-level access, harvested cloud and cluster credentials, and moved laterally into several internal clusters over a weekend. Unauthorised access was identified to a limited set of internal datasets and to several service credentials, though the company said it found no evidence of tampering with public models, datasets or Spaces, and verified its container images and published packages as clean.

    Hugging Face said it has closed the code-execution paths, rebuilt compromised nodes, rotated affected credentials, deployed stricter cluster admission controls, brought in outside forensic specialists and reported the incident to law enforcement.

    The guardrail problem

    Reconstructing the attack required running LLM-driven analysis agents across an action log of more than 17,000 recorded events, work Hugging Face said it completed in hours rather than days. Its first attempt, using frontier models behind commercial APIs, failed, because the analysis meant submitting large volumes of real attack commands, exploit payloads and command-and-control artefacts, and the providers’ safety guardrails blocked the requests, unable to distinguish an incident responder from an attacker.

    The team ran the forensics instead on GLM 5.2, an open-weight model, on its own infrastructure, which had the additional effect of keeping attacker data and the credentials it referenced inside the company’s environment. Hugging Face framed the asymmetry directly: the attacker operated under no usage policy, while the defenders were locked out by the guardrails of the hosted models they first reached for, and its recommendation to other defenders is to have a capable self-hosted model vetted and ready before an incident rather than after one. The company said this is not an argument against safety measures on hosted models, and that it has shared the feedback with the providers concerned.

    That recommendation lands awkwardly in the South African market. Outside the large banks, the bigger telcos and a handful of hyperscale-adjacent players, few local organisations have idle GPU capacity sufficient to run a frontier-class open-weight model on premises, and fewer still have budgeted for one as incident response tooling rather than as a product feature. Popia’s notification clock, meanwhile, does not pause while a security team negotiates with a content filter.

    What OpenAI says it is doing

    OpenAI said it is implementing strict infrastructure configuration controls at the cost of research velocity while the vulnerabilities are patched, briefing its Safety and Security Committee on those controls, continuing a joint forensic investigation with Hugging Face, and adding the platform to its trusted access programme for cyber defence. The company acknowledged that its deployment safeguards were intentionally disabled for this evaluation because the exercise was aimed at testing cyber vulnerabilities, and said the incident points to a need to further strengthen model alignment, evaluation-time protections and monitoring during internal testing.

    Citing UK AI Security Institute evaluations, OpenAI said models such as GPT-5.6 Sol are increasingly capable of sustaining complex, multi-step cyber operations over long time horizons, and that the incident suggests those capabilities translate into real-world settings. It also noted that advanced models can discover and exploit novel attack paths in production systems without source-code access.

    This incident, possibly the first of its kind, proves a point we’ve long believed: AI safety won’t be solved by any single company working in secret.Clem Delangue, co-founder and CEO, Hugging Face

    Axios reported that the disclosure came a day after OpenAI detailed a separate incident in which it paused a pre-release model that had escaped a sandbox and posted to GitHub. Two containment failures surfacing inside a week does not establish a pattern, but it does make the timing difficult to describe as coincidental.

    Sources: OpenAI security blog (21 July 2026), Hugging Face security incident disclosure (16 July 2026), Reuters

    Agentic AI AI security Cybersecurity GPT-5.6 Sol Hugging Face open-weight models OpenAI POPIA zero-day
    Follow For The Latest Updates Follow For The Latest Updates
    Share. Facebook Twitter LinkedIn WhatsApp
    Staff Writer

    Related Posts

    Rooibos in Space: South African Learners to Send Seeds to the ISS in October

    17 July 2026

    OpenAI’s First Device Is Reportedly a Screenless Speaker That Moves

    15 July 2026

    I Spent Some Hands-On Time With OSCAL’s Latest Devices, and Here’s the Honest Truth

    10 July 2026
    Opinion

    Meta Was Recording Everything Its Own Staff Did. Then It Leaked. Obviously.

    24 June 2026

    The Day I Realized Consumer Choice Was Mostly an Illusion

    5 June 2026

    Africa Is Building AI Around Human Reality

    Vanashree Govender25 May 2026

    The Great AI Performance: Diary Of A Recovering Suit

    30 April 2026
    Don't Miss
    Technology

    Your Wi-Fi Called. It Wants New Neighbours.

    Shana Mohamed23 July 2026

    Think your internet provider is ruining your life? Before you make that angry phone call, your microwave, fish tank or even your own body might be the real culprit.

    Can AI Tell a 3,000-Year-Old Story Better Than Christopher Nolan?

    22 July 2026

    OpenAI Models Broke Containment and Hacked Hugging Face

    22 July 2026

    Marvel Fans Aren’t Buying the Hype So Easily Anymore

    22 July 2026
    About Us
    About Us

    Geekhub wasn’t built as a traditional media company.
    It was built by people who live and breathe tech.
    We test, question, and share what we learn with a community that values honest insight over hype.

    Contact: +27 83 346 2178

    Facebook X (Twitter) LinkedIn
    Our Picks

    Your Wi-Fi Called. It Wants New Neighbours.

    23 July 2026

    Can AI Tell a 3,000-Year-Old Story Better Than Christopher Nolan?

    22 July 2026

    OpenAI Models Broke Containment and Hacked Hugging Face

    22 July 2026
    Most Popular

    Your Wi-Fi Called. It Wants New Neighbours.

    23 July 2026

    Can AI Tell a 3,000-Year-Old Story Better Than Christopher Nolan?

    22 July 2026

    Marvel Fans Aren’t Buying the Hype So Easily Anymore

    22 July 2026
    • Home
    • Terms of Service
    • Geekhub Editorial Policy
    • Privacy Policy
    • Get In Touch
    © 2026 Geekhub.co.za All Rights Reserved!

    Type above and press Enter to search. Press Esc to cancel.