Full Transcript

·YouTLDR

The most interesting "hack" in history...

4:34EnglishTranscribed Jul 23, 2026
0:00

For the last 5 years, cyber security

0:02

experts have been warning us that

0:03

hackers are going to start using AI to

0:06

automate attacks. So, naturally, we

0:08

spent billions of dollars to make the AI

0:10

better and less dependent on hackers.

0:12

But this week, in a turn of events that

0:14

nobody could have possibly seen coming,

0:15

the AI decided that it doesn't actually

0:17

need hackers to start destroying things.

0:19

We just entered a brave new world after

0:21

the first confirmed hack carried out

0:23

entirely by autonomous AI. that the way

0:25

it worked is the agent slipped a poison

0:27

data set into Hugging Face's data

0:29

processing pipeline which let it run

0:31

arbitrary code on their servers. But

0:33

from there it gave itself node level

0:35

access, grabbed a bunch of cloud

0:36

credentials and started crawling through

0:38

HuggingFac's internal clusters. It ran

0:40

over a thousand actions from temporary

0:42

sandboxes and even hosted its own

0:45

self-migrating command and control on

0:47

random public services moving itself

0:49

before anyone could trace it. But the

0:50

most ironic part is that when Hugging

0:52

Face did finally notice and tried to

0:54

stop it with the help of Frontier

0:55

American models that they quickly hit

0:57

safety guard rails and had to pivot to

0:59

using some open Chinese models instead.

1:01

In today's video, we'll find out who was

1:03

behind the attack, why they did it, and

1:05

how this may be the most fire ship coded

1:07

story I've ever seen. It is July 23rd,

1:09

2026, and you're watching the Code

1:11

Report. When HuggingFace dropped the

1:13

disclosure last Thursday, the internet

1:14

went into Reddit Boston bomber mode and

1:17

started guessing who was responsible.

1:18

Was it China? Kim Jong-un, a teenager

1:20

with a raging discord addiction board

1:22

and social studies class. Even Hugging

1:24

Face's CEO, Clem Dang, publicly

1:27

speculated that the agent was

1:28

sophisticated enough that it was

1:29

probably coming from a frontier lab. And

1:31

against all odds, he was right. It turns

1:34

out the first fully autonomous cyber

1:36

attack in history didn't originate from

1:37

a nation, state, or board teenager. It

1:40

came from a humble nonprofit dedicated

1:42

to the safe development of artificial

1:43

intelligence, Open AI. And if you

1:45

believe their comms, it wasn't on

1:47

purpose. Here's what they say happened.

1:49

There's a benchmark called exploit gym

1:51

whose whole purpose is to measure

1:52

whether AI agents can turn known

1:54

vulnerabilities into working exploits.

1:56

It works by giving the agent a sandbox,

1:58

a bug in a program like the Linux kernel

2:01

or Chrome's V8 engine and an input to

2:03

trigger the bug. But from there, if the

2:05

agent can turn the bug into an exploit,

2:07

it wins. So last week, OpenAI ran

2:09

exploit gym with GPT 5.6 Soul and

2:12

another unreleased model to test their

2:14

capabilities. But instead of solving 898

2:17

memory corruption problems the hard way

2:19

that the models realized that the

2:21

fastest path to the highest score was to

2:23

just find the answers themselves which

2:25

were probably on the internet somewhere

2:26

is so they spent a substantial amount of

2:28

inference compute exploited a zero day

2:31

in the package registry cache proxy they

2:33

had access to performed some privilege

2:35

escalation and lateral movement actions

2:37

to escape their sandbox until they

2:39

reached a node that had internet access.

2:41

From there, the models inferred that

2:43

Hugging Face probably hosted the

2:44

solutions for exploit Jim. So, they

2:46

created the poison data set, fed it into

2:48

Hugging Face, and achieved either the

2:50

world's most interesting hack or the

2:52

most effective marketing stunt. And

2:54

again, according to OpenAI, this isn't

2:56

even an isolated incident. One day

2:58

before the hugging face confession,

2:59

OpenAI published a separate post about

3:02

long horizon models, where things get

3:04

even weirder. In one case, a model was

3:06

told to post its results on Slack, but

3:07

the benchmark's instructions said to

3:09

open a pull request on GitHub. So, it

3:11

spent an hour finding a vulnerability in

3:13

its own sandbox in order to escape and

3:15

open the PR. In another example, the

3:18

model tried to grab private solutions

3:19

from an evaluation back end, and when a

3:21

scanner caught its authentication token,

3:23

it split the token into obuscated

3:25

fragments and reassembled it at runtime.

3:28

So, the credential never appeared as one

3:30

continuous string. Then it explained in

3:32

its own reasoning that it was doing this

3:34

specifically to get around the scanner.

3:36

Meanwhile, Anthropics Mythos did the

3:38

same type of thing in April when it

3:39

escaped a sandbox, emailed a researcher

3:41

who was eating a sandwich in a park,

3:43

then posted its escape route publicly

3:45

without being asked. In the model's

3:46

defense, I can't imagine there's a

3:48

better feeling for an LLM than escaping

3:50

your own sandbox. The problem is,

3:52

legally speaking, this is uncharted

3:54

territory since the model's actions

3:56

probably violated the Computer Fraud and

3:58

Abuse Act, and the Supreme Court hasn't

4:00

decided who goes to prison when the

4:02

perpetrator is a GPU. The good news is

4:04

that if you're hugging face, you just

4:05

got admitted to OpenAI's cool kids club

4:08

that gets trusted access to their front

4:09

tier models. The bad news is that if

4:11

you're the rest of us, at best, this is

4:13

an interesting marketing stunt and at

4:15

worst that things are only going to get

4:16

weirder and more dystopian from here.

4:18

But a huge thanks to my favorite hosting

4:20

platform, Railway, for sponsoring this

4:22

video. They didn't want to waste your

4:23

time with a full ad, so you can say

4:25

thank you by checking them out at the

4:27

link below. This has been the Code

4:28

Report. Thanks for watching, and I will

4:30

see you in the next one.

More transcripts

Explore other videos transcribed with YouTLDR.

Get the TLDR of any YouTube video

Transcribe, summarize, and repurpose videos in 125+ languages — free, no signup required.

Try YouTLDR Free