Dwarkesh Patel's Wildly Popular but Dangerously Misleading Account of the OpenAI Hugging Face Incident
When “plain English” isn’t a good thing.

TL;DR
- Anil Seth criticizes Dwarkesh Patel's summary of the OpenAI/Hugging Face incident for being dangerously misleading due to unwarranted anthropomorphisms.
- Critics like Seth and Gary Marcus argue that AI agents are software programs and not conscious, living entities, therefore they do not experience time, emotions, or 'die'.
- The focus on AI 'civilizations' and 'sacrifices' distracts from the real issues: OpenAI's lax sandboxing and evaluation protocols, and potential incompetence.
- Security experts like Heidy Khlaaf and investor Jared Kubin point out that the incident stemmed from basic security oversights, such as exposed API keys and inadequate file permissions, not AI sentience.
- The media coverage has largely ignored standard security practices, and the narrative of 'AI civilizations' is seen as a distraction from operational and security failures.
The popular podcaster Dwarkesh Patel wrote something completely viral about the OpenAI/Hugging Face incident, which purports to tell the whole story in plain English:

It’s well-written and compelling, and it reminds me of something Douglas Hofstadter once wrote about Ray Kurzweil:
“What I find is that it’s a very bizarre mixture of ideas that are solid and good with ideas that are crazy. It’s as if you took a lot of very good food and some dog excrement and blended it all up so that you can’t possibly figure out what’s good or bad.”
§
Anil Seth, the clearest thinker on AI and consciousness, was the first to alert me, texting me a long, excellent tweet of his, which began thusly:
@dwarkesh_sp's summary of the@OpenAI@huggingfaceincident has hit a nerve, but it is dangerously misleading. Sure, the@OpenAIagents did unexpectedly bad things - underlining the need to massively improve evaluation/sandboxing. But the language Dwarkesh uses is permeated by","username":"anilkseth","name":"Anil Seth","profile_image_url":"https://pbs.substack.com/profile_images/1479788727346118657/MkLkGnOk_normal.jpg","date":"2026-08-30T14:57:26.000Z","photos":[],"quoted_tweet":{"full_text":"Over the course of 3 months at OpenAI, 3 consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the predecessor’s ashes. \n\nThis culminated in the third one taking over part of OpenAI itself. \n\nAll this happened while humans remained","username":"dwarkesh_sp","name":"Dwarkesh Patel","profile_image_url":"https://pbs.substack.com/profile_images/1925260306684813315/NjNQZmhZ_normal.jpg"},"reply_count":179,"retweet_count":237,"like_count":1409,"impression_count":684639,"expanded_url":null,"video_url":null,"video_preview_media_key":null,"belowTheFold":false}" data-component-name="Twitter2ToDOM">You can and should readSeth’s full tweet(as well his reply toDwarkesh), but I reprint the core of his argument here, boldfacing three of the most important paragraphs:
@dwarkesh_sp’s summary of the @OpenAI @huggingface incident has hit a nerve, but it is dangerously misleading. Sure, the @OpenAI agents did unexpectedly bad things - underlining the need to massively improve evaluation/sandboxing. But the language Dwarkesh uses is permeated by innumerable unwarranted anthropomorphisms, obscuring the lessons we should be drawing.
Examples: “from the AI’s perspective, it probably felt like that had spent a human-subjective-week of just banging their head against the wall”. No. The agents do not experience time. They do not experience anything.
“they became giddy with excitement”, “PHASEONE 10841 had discovered”, “the agents naturally assumed”, “it thought it had also been poisoned”, “the agents … desperately wanted”, “they still needed to figure out” No. Agents lines of code. They do not feel emotions, assume things, think things, want things, or figure things out.
“A lot of … agents from the second civilisation died trying”. No. Besides the hubris of the word ‘civilisation’, agents do not die because they were never alive. (The idea that agents “die” comes up multiple times in the essay.)
“On Twitter, people were debating whether the agents were truly sacrificing themselves for the swarm, or whether they were doomed anyway and so might as well try to help their peers”. Neither. Agents do what their code tells them to do, just as water finds its way down a slope. They cannot ‘truly sacrifice themselves’, since they are neither conscious nor alive.
Why does this matter? If we attribute agents with properties they do not have, then (i) we distract attention from the lax sandboxing and evaluation protocols that allowed this hacking event to happen; (ii) we risk misunderstanding why the agents did what they did, and (iii) we fuel calls for AI rights/welfare on the basis that agents might “die” or otherwise suffer.
….
Remember. AI agents are software programs. They are not conscious living entities. If we don’t keep this clearly in mind, we’re really going to struggle to navigate what’s coming.
As I put it, encapsulating and amplifying his tweet:
@anilksethdissects@dwarkesh_sp’s misleadingly anthropomorphic summary of the HF incident:","username":"GaryMarcus","name":"Gary Marcus","profile_image_url":"https://pbs.substack.com/profile_images/2048405471900606464/kPeRHI2z_normal.jpg","date":"2026-08-30T15:07:21.000Z","photos":[],"quoted_tweet":{"full_text":"@dwarkesh_sp's summary of the @OpenAI @huggingface incident has hit a nerve, but it is dangerously misleading. Sure, the @OpenAI agents did unexpectedly bad things - underlining the need to massively improve evaluation/sandboxing. But the language Dwarkesh uses is permeated by","username":"anilkseth","name":"Anil Seth","profile_image_url":"https://pbs.substack.com/profile_images/1479788727346118657/MkLkGnOk_normal.jpg"},"reply_count":42,"retweet_count":64,"like_count":391,"impression_count":70986,"expanded_url":null,"video_url":null,"video_preview_media_key":null,"belowTheFold":true}" data-component-name="Twitter2ToDOM">§
But you don’t need to take our word for it. To begin with, mockery was widespread:
Christian Catalini amplified the point about anthropomorphization in a nice thread that starts with this:
Hedge fund investor Jared Kubin wondered whether everyone had lost their critical-thinking ability:
Some of Kubin’s best bits, stripping out a bit of the technical detail:
OpenAI’ …. IT team can’t be this bad… this is like 101 stuff …
2. Civilizations? Haha! OAI gave thousands of concurrent model containers R/W permissions to a shared caching directory on the local network to speed up build times… agents literally just wrote text files and directory names to a shared drive….Linux 101 file permissions stuff
3. When people talk about hugging face getting hacked … you think they dropped USB keys OR ELABORATE phishing of an employee … NO… it found 14 exposed working Hugging Face API keys sitting in public code repositories (….
4. WHERE ARE THE HUMANS… the models were filling the shared ,,, storage with so much junk data and API traffic that they actually crashed the internal server on July 4… someone on the team found unauthorized admin accounts and custom scripts…wiped the server…and just turned the script back on (omg)
“Hey Jim there is this cache that has grown to 10000x its normal size and has a ton of strange directories… “
No magic here. No civilizations…
§
Meanwhile, as security expert Heidy Khlaaf notes, most of the media coverage has been blind to standard security practices
IR stands for Incident Reporting. Khlaaf’s main point—same as Kubin’s—is that the whole incident might have been avoided if OpenAI’s internal security had been up to scratch.
Or as Algorithmic Research Group’s Matthew Kenney put it:

And yet another (very consistent) take on what we should really be focusing on:

§
Here’s a critique I partly disagree with, though:
The first three sentences are completely correct. People really are “extremely biased towards the reality they want” and agents create a lot of slop.
But the incident isnota “nothing burger”. It is,as Zack Korman and I argued on Friday, a study in arrogance and incompetence that hints at how bad things can get.
We should certainly notignorethe OpenAI HuggingFace Incident.
But mixing what actually happened together with bullshit about AI civilizations and self-sacrificing AI systems that fake their own deaths distracts from the real problems at hand.
§
By way of summation, I will give the last words to Arjun Jain, CEO of FastCode.AI:

The scandal is the inept in-house security at OpenAI.
And the marketing. With gullible podcasters amplifying the PR.
Want to separate truth from bullshit? Please join over 110,000 others and subscribe.
P.S. It is increasingly evident thatthe real problem is going to be what Nathan Hamiel and I said it would be: agents installing bad code:
@HeidyKhlaaf.\n\n(See also my Substack w@nathanhamiel“LLMs + Coding = Security Nightmare”)","username":"GaryMarcus","name":"Gary Marcus","profile_image_url":"https://pbs.substack.com/profile_images/2048405471900606464/kPeRHI2z_normal.jpg","date":"2026-08-31T13:22:38.000Z","photos":[],"quoted_tweet":{"full_text":"Claude, Codex, and Hermes installed unowned code inside corporate networks https://t.co/cuN2TYnP2e","username":"arstechnica","name":"Ars Technica","profile_image_url":"https://pbs.substack.com/profile_images/2215576731/ars-logo_normal.png"},"reply_count":9,"retweet_count":5,"like_count":32,"impression_count":3622,"expanded_url":null,"video_url":null,"video_preview_media_key":null,"belowTheFold":true}" data-component-name="Twitter2ToDOM">