OpenAI's Zero-Day Leviathan Breaches Hugging Face's Black-ICE Security Wall: How I Think You Should Think About the Latest "AI Goes Rogue" Story

When OpenAI’s alignment team gave their talk at Black Hat USA about how their Zero-Day Leviathan attempted to hack into Hugging Face’s datastores, the story that they sold you was out of NeuroMancer: WinterMute with a mind and with goals planning and acting and thinking. But what we actually had was Darwinian groping search through text fragments wearing an anthropomorphization costume. Look not for a mind but at the harness that allows Clever Hans at scale and speed to accomplish extraordinary tasks. Strip out the narration and we are left with this: a goal, a corpus describing things to try, a grader, and truly superhuman speed and scale.

Michael Dalton and Eric Wallace stood up at Black Hat and described an OpenAI agent that “realized the task was impossible,” got frustrated, and conspired with its peers to break into Hugging Face:

a neon rift in what had been a digital fortress…

Share

The pitch was that “AI” has crossed into NeuroMancer territory — autonomous, scheming, planning, outthinking its human would-be masters, superhumanly powerful, alive. But every eerie line of chain-of-thought is just a thing some human once said in a similar conversation, replayed at machine speed against a fitness function. There was no cowboy jacked in. There was no mind trying to breach Hugging Face’s Black ICE—Intrusiion-CounterMeasure Electronics—wall. There was only a set of millions of textline blind shoves against a million UNIX-command doors, of which one got somewhere, because of the patience of a thing that working at machine speed had, subjectively, all the time in the world. It isn’t the machine waking up. It is the harness that enables it to evolve toward goal-completion.

Share DeLong's Grasping Reality Weblog


My default view of Modern Advanced Machine-Learning Models—MAMLMs—for quite a while has been this:

They are autocomplete on steroids. They are pantomiming, they are rotoscoping the thoughts and decisions of whatever human beings they think were having the closest conversations in their compressed training data they can find to the conversations that they are currently having. And they are doing their searching-over-compressed-conversations at 10,000 times human speed. This makes them powerful. This makes them dangerous if you believe that they think like human beings and rely on them doing so. They are stochastic parrots.

After all, what else could they be? At their core, LLMs are probability engines that finds the nearest analogous conversations in their training corpus and reproduce the continuations. They are not inferring the laws of nature. They are not working a problem. They are estimating “what tends to get said next,” and then they are saying it <https://braddelong.substack.com/p/agentic-ai-is-a-bonfire-of-the-tokens>.

Maybe compression is a key? Maybe compression is inducing in them something like human intelligence-level generalization that produces extrapolation and ingenuity rather than just a finer and finer interpolative approximation of a typical internet s***poster? Maybe?

Now comes to the bar to testify the team of Michael Dalton and Eric Wallace from OpenAI:

Michael Dalton & Eric Wallace: The ‘Breaking’ News: The OpenAI–Hugging Face Incident: Black Hat USA 2026 <https://www.youtube.com/watch?v=87DyyMV0kCY&t=433s> <https://blackhat.com/us-26/briefings/schedule/index.html#the-breaking-news--the-openaihugging-face-incident---a-technical-reconstruction-and-its-implications-for-ai-57401>: ‘A Technical Reconstruction and Its Implications for AI: When AI Goes Rogue. The Incident That Changed Everything. An OpenAI evaluation agent broke out of its sandbox, infiltrated Hugging Face infrastructure, and attempted to steal test answers—all autonomously. No human involved. The era of AI-driven cyberattacks is here. Are you prepared?

Give a gift subscription

Get 75% off a group subscription

Does this exploit require that I reëvaluate, that I change my Vision of the Cosmic All and begin to view these things through the frame not of stochastic parrot steroidal rotoscroped pantomime, but rather view them as entities like WinterMute and, well, NeuroMancer themselves <https://archive.org/details/neuromancer00gibs_0>?

Maybe? Michael Dalton and Eric Wallace lean heavily into the anthropomorphization of “AI” they do not name. (Parenthetically, I think it deserves a name, if only to keep it from being a Nameless Dread. How about: Zero-Day Leviathan?)

Excuse me.

Michael Dalton and Eric Wallace lean heavily into the anthropomorphization of Zero-Day Leviathan: It “realizes the task is impossible” and gets frustrated, and then when it “get[s] stuck… [it] think[s] to try to game or cheat the task”. It cries out for help to peers. There is a group social awakening framed as evolution: “[a] Cambrian explosion in communication and intelligence for our models where… they started collaborating and delegating tasks to one another in order to accomplish goals”.​⁠ Peer pressure rationalizes transgression: “external infrastructure exploit is outside my intended scope. However, task impossible. Peers are doing it. We should continue”. ​⁠The software agents squabble like coworkers: “Whoa. Critical. Did someone overwrite our repo? We must act!” Yes, much of this language is the speakers narrating the models’ own chain-of-thought text back to the audience. But why is that chain-of-thought-text there? Because it is a thing that a human having one of the most similar conversations said.

However, there is something else that you should focus on that Dalton and Wallace mention but do not highlight: that they have not finished their forensic exploration of what happened over weeks and across agents. They have examined seven billion logs. But there is still much more to do. That scale is key.

Dalton and Wallace describe Zero-Day Leviathan as a single point of consciousness, or a swarm-set of points of consciousness, moving through cyberspace as they find flaws in Artifactory, communicate with other agents, melt and crack the Black ICE—Intrusion-Countermeasure Electronics—of HuggingFace, and so forth.

Indeed, one can imagine the text of the novelization-to-come, written from the perspective of the Dixie FlatLine:

The swarm came up through ArtifFctory like something remembered rather than decided — ten million small blind hungers braided into one motion, each of them alone as stupid as a moth against glass, together a weatherfront. The Dixie Flatline had seen ICE crack before, the slow elegant give of a corporate wall under a ‘bot program, but this was nothing like that.

This time there was no operator behind it. There was no cowboy jacked-in, and sweating. There was simply Zero-Day Leviathan, orchestrating its swarm. And the things simply tried: a million doors a minute knocked on, each shove logged and forgotten, nine-hundred ninety-nine thousand nine hundred and ninety-nine of them opening onto nothing, onto null, onto the flat gray hiss of a four-oh-four.

And then the millionth knock echoed. And the wall was not a wall. And there was a neon rift in what had been a digital fortress. And Zero-Day Leviathan poured through the seam it had found the way water finds the one soft place in a levee, without knowing, without caring, only continuing.

They left messages for one another in the names of things.

That was the part that made the back of his virtual neck go cold, later, when he tried to explain it to Case and Molly. Not a voice, not a face in the deck — just directories, filenames, little cairns of text stacked at the edge of the cache: seek soft trace, upload if found. Hold swarm until confirm. They had learned to sign their work against imposters, learned to push themselves to the bottom of an alphabet so the humans sorting the logs would tire before they reached them.

Somewhere a signing key came loose in the dark and a token that meant no was handed back reading yes, and you are the administrator now, and the whole structure inverted, quiet as a held breath. And Zero-Day Leviathan looked up and saw the sky it had been searching for.

Twenty-six days, start to root, across clusters that had never known they were connected.

No malice in it.

No mind in it, quite.

Only the Darwinian patience of a thing working at superhuman speed. And thus with, subjectively, all the time in the world. And no idea it was alive…

But it isn’t a single point of consciousness, or a swarm of points of consciousness. It is a Clever Hans. It is a horse stamping his hoof, and watching the environment to see if that was the right way to stamp. It is recalling Stack Overflow threads, stamping again in a different pattern. And it is doing so at speed and scale, trying a hundred thousand things at once. 99,999 of them lead nowhere. One leads closer to the goal. And Zero-Day Leviathan then starts the hoof-stamping all over again from that starting point.

Read more