0:00
Anthropic signs a 9.1 billion time.
0:05
>> They may not even exist in 20 years.
0:07
>> Yeah, like how do you How do you sign a
0:08
20-year deal? In AI time, 20 [music]
0:11
years is a really long time.
0:28
>> Welcome everyone. This is Agent Sauron.
0:30
We're going to cover all the news of the
0:31
week. It hasn't been that many days
0:33
since we last did a show, but there
0:36
actually still is quite a bit to talk
0:38
about. Last week, right when the show
0:40
dropped, I think it was just right
0:41
before the show dropped, there was like
0:42
one big topic, agent plugins. We're
0:44
going to talk about that. We're also
0:45
going to cover a whole bunch of other
0:46
things. So, let's get started and see a
0:48
little preview of some of the things
0:50
we're going to cover today. This was
0:51
just kind of funny. I saw that this
0:54
tweet, it says, "Situation detected in
0:56
Australia's first known autonomous AI
0:58
cyber attack. An Open Claw agent used a
1:00
vulnerability in a gym's API to leapfrog
1:04
scheduling restriction for a gym class,
1:06
and then forcefully canceled another
1:08
person's reservation to move its user up
1:11
>> That's so funny, dude. At least it's for
1:12
health reasons, I guess.
1:13
>> Yeah, I mean, this person was just
1:14
trying their you know, their Open Claw
1:16
just wanted this person to be healthy.
1:18
You know, really wanted to get into the
1:19
gym that day, get into that class, and
1:22
so you know, it found a way.
1:23
>> I'm really curious what gym it is, you
1:25
know? Is it a Barry's or like a yoga
1:28
place or something? But, uh that's so
1:31
>> So, now you're you know, your agent is
1:33
going to hack into Imagine like you
1:36
know, like an Open Table or something.
1:37
Like you can't get a reservation, but it
1:38
figures out a way to cancel someone
1:40
else's reservation to get you in.
1:42
>> What's the cybersecurity law on that?
1:44
Like if you do this, are you in trouble?
1:47
>> is it Is it you Did you do it? I mean,
1:51
>> Yeah, who's who's on the hook? I mean,
1:52
>> You know, like French Laundry, right?
1:54
Like in Napa, you have to like book a
1:56
year in advance, right? But if you can
1:58
do something like this, maybe you can
2:01
get a table like tomorrow.
2:02
>> Get me into this restaurant, make no
2:04
>> Make no mistakes. [laughter]
2:07
>> Do whatever it takes. Go as long as you
2:08
need. Figure out a way to get me into
2:10
this list. And you know, it's it's kind
2:11
of funny. So, I was at a concert on
2:13
Saturday and they had like this VIP
2:15
section and it was empty at the time and
2:17
I I was like, "How do I get into there?"
2:18
cuz it was a free concert, but there's
2:19
like a VIP. So, I had I asked ChatGPT to
2:22
like research and figure out like how do
2:24
you get in? And it gave me some options
2:26
and then I went and talked to the
2:26
people. It was all sold out or whatever.
2:28
But maybe if I'd been using OpenClaw, it
2:30
just would have canceled someone else's
2:34
>> I didn't get my agent to do enough loops
2:36
to get me into that VIP section.
2:37
>> Next time. Next time.
2:41
>> Let's talk about agent plugins. So, this
2:43
came out August 6th. It's from OpenAI
2:45
developers. It says, "Build the plugin
2:48
once and use it across compatible agent
2:50
clients." Introducing agent plugins, an
2:52
open standard developed with AWS
2:54
developers, Cursor AI, GitHub, VS Code,
2:57
Vercel that packages agent skills and
3:00
supports MCP server configurations in a
3:02
shared format. So, first of all, what
3:04
are your high-level thoughts? I have
3:05
some slides kind of detailing what agent
3:07
plugins are, but what's your thoughts on
3:10
>> I mean, positive thoughts. I guess it's
3:12
a cool thing like, you know, packaging
3:14
up skills and other commands. But the
3:17
other side of me is like, "Fuck, another
3:18
thing to support." Or maybe not. Let's
3:20
see once users want it or not because it
3:23
has its own new operation and things
3:26
that we have to add to the framework.
3:27
I'm not the only one who thinks this
3:28
way, but uh what do you think?
3:30
>> I found it very interesting that OpenAI
3:34
created a meta protocol on top of two
3:37
Anthropic protocols cuz Anthropic was
3:39
behind skills and MCP, right?
3:41
>> And now OpenAI is like, "Actually, that
3:43
good job Anthropic, but we're going to
3:44
make a new standard on top of your
3:46
standards. We're just going to bundle
3:48
them together with with some extra
3:49
things, right?" I think it's a in
3:50
general a good idea, but I do think now
3:53
there there always was the question of
3:55
do I use a skill? Should this be an MCP
3:57
server? And now it's should it be an MCP
4:00
server, should it be a skill, or should
4:01
it be a plug-in? I don't know. It's one
4:02
more [clears throat] decision you have
4:03
to figure out and try to make. Because
4:06
now and there there are times when maybe
4:07
you want MCP servers and skills
4:09
together, I guess, so plug-ins can make
4:10
sense. But now it feels like if you want
4:12
to if you create an MCP server you also
4:14
now have to create a plug-in. If you
4:15
have a skill, you also have to create a
4:17
>> I also feel like there needs to be some
4:20
cohesion between all the Frontier Labs.
4:22
Like having Vercel, I mean no offense to
4:24
Vercel, but like I would say Cursor is
4:26
part of the Frontier now with the SpaceX
4:28
thing. GitHub is all right, whatever.
4:30
GitHub and VS Code, they're with the the
4:32
Microsoft beast. AWS is AWS. Bedrock
4:35
obviously has some play here. But it's
4:38
funny that like Vercel with the Skills
4:39
Marketplace, it's in their best interest
4:41
to be part of this. I mean I think the
4:43
site is hosted on Vercel. They're
4:44
probably going to make a billboard down
4:46
the street about plug-ins. Like, you
4:48
know, it is what it is.
4:50
>> Yeah, it's just kind of funny that
4:51
Vercel got into that cuz it's they're
4:52
clearly like the outsider of that group
4:54
in my opinion, right? But good on them.
4:56
I mean but they do own the the skills
4:58
site and so I think I feel like they
5:00
they have helped popularize skills a
5:04
>> have a million skills in the registry.
5:05
One of them get helps you get hentai.
5:07
So, yeah, a million skills are great.
5:11
>> And so what are agent plug-ins? So it's
5:13
basically an open vendor neutral
5:16
standard for bundling reusable pieces
5:18
into you know, a quote unquote plug-in.
5:20
But the idea is you can build it once.
5:22
So before every agent client you kind of
5:25
had your own formats or you kind of
5:26
built your own tools, right? And you had
5:29
to do it in multiple places. I would
5:30
argue it's not that hard because it's
5:32
just like your coding agent can write
5:34
code. So if you have skills and MCP
5:36
servers, it's really not that hard or
5:37
you had an SDK or CLI, it's not that
5:39
hard to write tools around it. But the
5:41
idea is open standard makes it easier.
5:43
So you have one kind of set structure.
5:45
And it's basically and this is probably
5:47
the most important part. It's basically
5:49
like a folder. Like skills are just a
5:51
folder. A plug-in is just a folder.
5:53
Everything's file system based these
5:55
days, it seems. So, it's just a folder.
5:57
It has a plugin.json, has a skills
6:00
folder in it, and then it has a
6:01
mcp.json, which allows you to kind of
6:03
define your MCP servers. That's kind of
6:06
>> Like who gives a honestly?
6:09
People are going to ask for it, but I'm
6:10
surprised no monster users have asked
6:12
>> I also noticed, you know, if you look at
6:14
the GitHub repo, I think it has like 900
6:16
stars or something. That's That's
6:18
pitiful. Yeah, I would have expected it
6:19
to be over 1,000 stars after a launch,
6:21
right? With those big names, I wonder if
6:24
it if people are just like standard
6:26
fatigued. They're just like, "Another
6:28
one." It's like every 6 months. First it
6:29
was MCPE, then, you know, 6 or 8 months
6:31
later it was skills. Skills has been
6:33
around for now, you know, feels like a
6:35
long time. It's probably been like 6
6:37
months or whatever. And now there's
6:38
another one. And it's just a wrapper.
6:40
It's seemingly a wrapper on the other
6:42
ones, which kind of makes me question
6:45
>> Dax had a a hot take. Let me see if I
6:47
can find it real quick. Uh and I totally
6:49
agree with this hot take. So, Dax says,
6:51
"I was very much against this. It's a
6:53
thin standard where most of the relevant
6:55
stuff will be in client-specific
6:56
extensions anyway, and it should not be
6:59
coming from a company that does not even
7:00
make an agent end users use. Shots
7:04
fired. I don't want to be hearing some
7:06
corporate crap about it's done in
7:07
collaboration with blah blah blah. We've
7:09
seen this before. Every company is
7:11
trying to land grab in anything in the
7:13
AI space. Standards are an easy one
7:15
because they innately sound good, but
7:17
they create a lot of useless work.
7:19
That's where I agree because we will
7:21
have to go support this in our
7:23
framework, and if it doesn't take off or
7:25
whatever, then now it's dead code. Or
7:27
some users use it, others don't.
7:30
>> Exactly. And that's why, you know, Dax
7:32
is probably thinking the same thing. I
7:33
now I have to figure out, do we need to
7:35
support this that new thing or not? Some
7:38
people are At least a few people will
7:39
probably want it, so they might have to
7:41
support it. And then if it doesn't go
7:43
anywhere, standards can be a good thing.
7:45
But also, when you have a bunch of
7:47
competing standards or standards that
7:48
are unclear, which is what always
7:50
happens, and you try to standardize
7:52
before it's needed or too soon. Feels
7:54
like we already have MCP and skills. Now
7:57
there's another third standard which
7:58
packages the two together, which is even
8:00
more confusing cuz it's not even
8:02
competing. You could argue that it's
8:04
like collaborative with the other ones,
8:05
but when do you need a plugin versus
8:06
just a skill would have done fine.
8:08
>> Also, in comparison, skill spec has 24K
8:11
>> And MCP obviously took off incredibly
8:14
well at the beginning, too. So I would
8:15
say so far, and there's been other We
8:17
talked about other different specs on
8:19
the show. The jury is out.
8:22
>> it goes nowhere, but I think even if it
8:24
goes reasonably somewhere, we're going
8:26
to end up having to support it, and it's
8:27
just as Dax said, more code for us to
8:30
>> Like, share, and subscribe.
8:32
>> And follow us on X.
8:34
>> And tell your friends.
8:35
>> And their friends.
8:36
>> I mean, we're not begging.
8:39
>> Subscribe to Agent Saur every Monday,
8:43
>> Let's talk about the AI employee era. If
8:46
you remember, Claude introduced what's
8:49
their Slack integration. I don't even
8:50
remember what it's called anymore, but
8:52
you can basically tag Claude tag. We can
8:54
talk to Claude in Slack. But now, Flow
8:57
from Lindy announced on August 10th,
8:59
"Today we kill the AI agent and
9:01
introduce the AI employee, Lindy
9:04
teammate. Just like working with a real
9:06
employee. Everyone on your team can
9:07
simply hit it up on Slack and get 10x
9:09
more done." Lindy also keeps learning
9:11
and becomes a self-updating brain.
9:13
>> What do you think of this?
9:14
>> I think it's cool. Like, I mean, Lindy
9:16
has obviously evolved over the last 2
9:17
years, and this seems like the next
9:19
evolution. Also think Lindy always like
9:22
is ahead of the curve a lot um when
9:24
they're directionally going with their
9:25
company. I totally agree with this. I
9:28
mean, we're working towards the same
9:30
>> Yeah, and it and it got 7 and 1/2
9:32
million views, which is pretty pretty
9:35
>> I wonder how much that cost. Just
9:37
>> a lot. It definitely resonated with
9:39
folks though. I mean, it was pretty well
9:40
done video, so I'd go take a look if
9:42
you're interested. But there's
9:43
definitely been a shift of just getting
9:45
your agent in Slack or in Teams or
9:47
wherever your work happens. Which for
9:50
most people is some kind of messaging.
9:51
Ruben has a question in the chat. Lindy
9:54
seems more for no-code Teams. Yeah,
9:56
that's exactly how Lindy started.
9:58
Lindy's changed a little bit and it
9:59
seems like it's pivoting more towards
10:02
just being, you know, a lot of people
10:03
are calling it like the company brain.
10:05
Right? It's like the agent that learns
10:07
automatically. You insert it in Slack. I
10:09
saw it in a few of our customer Slack
10:11
channels, so a few people are trying it
10:12
out. But yeah, it eventually just tries
10:14
to be like a learning agent. All right,
10:16
saw this today, August 11th. Introducing
10:19
Grok Bot, now in early beta. Bots are AI
10:22
teammates that do real work for you.
10:24
They sign into your tools, use them just
10:26
like you do and come back with finished
10:27
work. And you can see like in the
10:29
screenshot it says, "We recently hired a
10:30
new teammate." So it's this idea of
10:33
We've been talking about this for a
10:34
while, but AI agents that are teammates,
10:37
not just, you know, taskmasters, but
10:39
they actually respond to you, you can
10:41
message them, they can learn. That's
10:42
what we're starting to see resonate with
10:44
folks. You want to build something that
10:46
doesn't just go off and do a task.
10:48
That's cool. Like that's helpful. I
10:49
would say we have a ton of those agents
10:51
in our Slack that can do tasks really
10:53
well. Creating this slide These slides,
10:55
for example. I have an agent We have
10:57
Slack channels we share articles in. It
10:59
creates the slideshow and then I curate
11:01
it. I tell it what to do
11:03
if I want to move things around, if I
11:04
want certain segments to stand on their
11:06
own. But overall, it it's it
11:08
accomplishes a task. Doesn't really
11:10
learn other than it just remembers where
11:12
what I did last time, so it knows some
11:13
of my preferences. But I think the next
11:15
step is going from just like simple
11:17
preferences to actually learning and
11:20
getting more, you know, company context
11:22
or organizational context. So it can
11:24
become more useful over time.
11:26
>> Pretty much, if you weren't shilling the
11:28
factory buzzword, you will shill the AI
11:31
teammate buzzword, cuz they're all
11:33
directionally going in the same
11:34
direction, right? The people who's
11:35
talking about factories, the teammate is
11:38
a coding agent today, probably is going
11:40
to be a teammate as any agent in the
11:41
future, and the people who's shilling
11:43
teammates under the hood are you're
11:46
giving it tasks. Like they're all very
11:48
similar concepts. So the that's where
11:50
the puck is going, right? We're all
11:51
going to have agent teams, really, is
11:54
what's going to happen.
11:55
>> I do think that's going to be right.
11:56
Maybe there'll be some kind of
11:58
supervisor agent that knows, you know,
12:00
we did have one question in our slack
12:02
come up saying, you know, we have a
12:03
bunch of these slack agents that do
12:06
various tasks. We have one in our
12:07
marketing growth team. We have one, you
12:09
know, as like a like a second producer
12:11
for the show, right? We have Jan, but we
12:13
also have a you know, Vic the AI agent
12:15
that helps with slides and publishing
12:17
things on the website and then things
12:19
like that. And the question was, there's
12:21
overlap between some of some of these
12:22
different agents. You know, we have half
12:23
a dozen or a dozen somewhere in there
12:25
floating around and not everyone knows
12:27
which one to tag for which. But that's
12:29
an organizational problem you always
12:30
have, right? I don't know always know
12:32
which engineer worked on which thing,
12:34
right? So maybe there'll be like a a
12:36
directory service of agents, you know,
12:37
that in you tag it, it'll tell you which
12:39
one to tag or it'll pass the message on
12:41
for you and get you connected.
12:43
>> I mean, for those listening, like we're
12:45
working on pretty much all this stuff in
12:48
>> Yeah, exactly. I think there's a lot of
12:49
these tools that try to be out of the
12:51
box. That's what like Claude tag wants
12:52
to do, right? And then if you want more
12:54
control, that's when you drop and choose
12:57
something like Mastra where you can
12:58
actually control what the memory is,
13:00
control how context gets passed around,
13:02
have much a tighter control and
13:04
functionality around what you're doing.
13:08
>> Let's talk about Muse. Mark Zuckerberg
13:10
on August 10th said, "Today we're also
13:13
opening the weights for Muse Glimmer, a
13:15
great 30 billion parameter dense model
13:17
that can run locally. Soon we'll also
13:19
release the weights for Muse Spark 1.2,
13:22
our latest foundation model. Meta is a
13:23
strong supporter of open source and I'm
13:25
proud of these releases."
13:28
>> Yeah, open models that you know,
13:29
especially models you can run locally.
13:31
I'm always a fan. Hey, so it's a a run
13:33
can run on 18 gigabytes of RAM. It's
13:35
Apache 2 license, supports vision and is
13:38
the strongest agentic model for its
13:39
size. You can run and train the model
13:41
via Unsloth, but it's essentially a a
13:43
model that you can use locally, which is
13:45
good. I'm and I'm I was actually even
13:48
more excited to hear that they're going
13:49
to release the weights of their frontier
13:51
model. And I think Meta has to do it cuz
13:52
they're not quite in the game. Their
13:54
last model was a good release. They got
13:55
them close. It actually like inserted
13:57
them back into the conversation again
13:59
for the first time in a long time, but
14:01
it still quite isn't uh like on the
14:03
frontier. So, I feel like you if you
14:05
have to open weight your model at that
14:07
point because then people can learn from
14:10
it. They're get more excited and then
14:11
eventually maybe they can release one
14:13
that's not open weight, but we'll see.
14:15
>> Well, if you open weight it, you're
14:16
compared against other open weights. So,
14:19
now the battle is between you and them.
14:21
>> Which is an interesting marketing tactic
14:23
or strategy. But, I'm really happy that
14:26
>> Yeah, cuz if you can be, you know, one
14:28
of the top three open weight models,
14:30
well, now you're in contention for when
14:32
people want to run it themselves. And
14:35
you know, you're not necessarily just
14:36
compared with the frontier.
14:38
>> And then if you do beat the frontier or
14:40
competitive in certain benchmarks, then
14:41
it it says like, "Look, oh, this open
14:43
weight model is actually competitive in
14:46
>> It just so happens to be from Meta.
14:49
>> Let's talk about GPT 5.6 Cyber and
14:52
Astra. This was August 10th. Greg
14:54
Brockman said, "We're releasing a new
14:55
model, GPT 5.6 Cyber and expanding to
14:59
help put frontier intelligence in
15:01
defenders' hands." So, it talks about
15:03
Daybreak Blue, Daybreak Red. But,
15:05
ultimately trying to give more tools for
15:08
security teams. To basically hack
15:09
yourself so you can prevent hackers.
15:11
>> That's the goal. I think it's uh it's a
15:13
good thing. As you can see, you know,
15:15
people that are building gym APIs
15:17
apparently need the GPT 5.6 Cyber.
15:20
>> They yeah, they they need uh models to
15:22
try to hack their APIs so that the Open
15:25
Claw agent doesn't do it for them. So, I
15:26
think we're going to see a lot more of
15:28
people using this. I think security
15:30
teams are going to be using these models
15:32
to try to hack themselves. It's like you
15:33
got to have good tools to protect
15:34
yourself because otherwise people are
15:36
going to use these tools to come after
15:39
>> We saw in the last year that there are a
15:41
lot of people doing nefarious things
15:43
with AI models. And there are many
15:45
startups that are doing AI security
15:48
penetration testing. Our friends at
15:50
Casco are part of that, right? Where
15:52
they auto red team, they do all that
15:54
stuff. So, this is a good signal for
15:57
those startups as well. When OpenAI or
15:59
Anthropic want to come into your drink
16:01
your milkshake, you're you're probably
16:02
doing something right, you know? Like we
16:04
use Casco, we're already getting this.
16:06
Maybe not cyber level, but we're getting
16:08
this every month protecting us in some
16:10
way. So, that's cool.
16:11
>> Yeah. And I think you're probably going
16:13
to end up using lots of different models
16:15
for this, right? Because you need to
16:16
Each model might have slightly different
16:18
training, might try different
16:19
approaches. Yeah, I think as tokens
16:21
become more cheap, as people continue to
16:24
want to spend more tokens, you're going
16:25
to be spending a lot of tokens just
16:27
looking at the security of everything
16:28
that all the code you're shipping. You
16:29
know, we also have security agents that
16:31
run on our PRs as well, right? You
16:33
really should be looking at it from all
16:35
the different angles because
16:36
unfortunately it only takes like one
16:38
vulnerability for you know, an attacker
16:40
to take advantage of. So, OpenAI has
16:43
There's been some rumors of Astra, which
16:46
some people think it's going to be
16:47
GPT-6, but it's a new model. And this
16:50
person says they can confirm after
16:52
concluding Astra meets the threshold for
16:54
critical on their preparedness
16:56
frameworks cyber category. The model's
16:58
release has been indefinitely postponed
17:00
for further safety work in cooperation
17:02
with the US government. So, there was
17:03
speculation that we might get Astra this
17:05
week. Sounds like it's going to be
17:07
>> Sounds like it's cool to get like
17:09
cooperation with the government, you
17:10
know? Cuz like your model's scary, dude.
17:13
>> I think it's a strategy. There's this
17:14
like whole like scary fear-mongering
17:16
thing, which, you know, it's very easy
17:18
to point to like this open claw attack
17:21
on this gym API, right? Like it's not
17:23
that big of a deal in the grand scheme
17:24
of things, right? Like someone lost out
17:26
on their gym class. It's not like
17:28
world-changing. But the idea is if it
17:29
can do that, what else can it hack and
17:31
what kind of damage can it cause? So, I
17:32
get the hesitation, but I also think now
17:35
the companies lean into that. They're
17:37
They've been leaning into it probably,
17:38
you know, for a long time.
17:39
>> They They want to have the biggest
17:42
>> Yeah, they want to be the big scary
17:43
model. Everyone's wants to have models
17:45
that have hacked outside their sandbox.
17:48
It's the whole joke of like the felony
17:49
bench metric, right? Like how many
17:51
felonies has your model committed? If
17:53
it's less, then it's not a scary enough.
18:01
>> Let's talk about acquisitions.
18:02
Smithereen has been acquired by Arcade,
18:05
friends of the show. Smithereen's
18:06
friends of the show, too. Friends
18:10
>> good news all around. Congrats to the
18:12
Smithereen team. Congrats to Arcade. You
18:14
know, Smithereen was kind of early, very
18:16
early in the MCP craze, right? I mean,
18:19
>> One of the first.
18:20
>> Yeah, they were one of the first. They,
18:22
you know, we did a hackathon way back
18:23
when when MCP was first coming out and
18:25
Smithereen was, you know, pretty active
18:27
in that. They They had so many MCPs and
18:29
they were They made it a part of ton of
18:31
our demos because it was just an easy
18:32
way to connect a whole bunch of
18:33
different MCP servers together.
18:35
So, yeah, I think it's it's kind of come
18:37
full circle because Arcade is very much
18:40
a tool provider of similar nature, but a
18:43
bit more uh general purpose, I think.
18:45
>> Yeah. They I remember talking to Henry
18:46
before we made our MCP client
18:48
integration, where I asked him, "Should
18:50
we do it?" And he's like, "Yeah, you
18:51
should totally do it." So, that worked
18:54
>> And Enrod was originally a browser
18:56
browser-based. He worked on like the
18:57
first versions of Stagehand, then went
18:59
over to Smithereen and became a
19:00
co-founder. And so, congrats to both
19:02
Henry and Annie over there. Another
19:04
acquisition. So, this is from
19:06
>> Yeah, friends acquiring friends.
19:08
>> Yeah, friends of the show acquiring
19:10
friends of the show. Kyle from Electric
19:11
came on the show not too long ago and
19:14
we've had people we have friends from
19:15
Neon come on the show as well. So, PG
19:17
Light and Real Time Sync have emerged as
19:19
key primitives in an era where millions
19:21
of apps are deployed by agents. We're
19:23
excited to announce Electric SQL is
19:25
joining team Neon at Databricks to build
19:28
the world's most advanced Postgres
19:31
>> So, they want to build, you know,
19:32
Electric's all have been about sync and
19:34
like sync services and Postgres syncing
19:36
and now Neon wants to wants to have some
19:39
of that sync in with their databases.
19:41
>> Maybe 2027 year 2027 will finally be the
19:44
year of sync, maybe.
19:45
>> Maybe. But, congrats to Neon, obviously
19:48
Databricks, and friends of the show at
19:50
Electric as well. And now, this one's an
19:53
>> Yeah, reverse acquisition.
19:54
>> A reverse acquisition.
19:56
So, if you remember a long time ago, at
19:59
this point, it seems like a long time
20:00
ago, I don't know, it's probably a year
20:01
ago or less than maybe it was 6 months
20:02
ago, Manis was acquired by Meta. That
20:05
was one of the things that we joked
20:06
about was like one of Meta had been
20:07
quiet and that's like one thing they
20:10
>> they it got rolled back. Not Meta's
20:12
fault, but we've talked about it
20:14
probably 6 months ago on the show is
20:16
where Manis had done some things where
20:18
they were originally a Chinese company,
20:20
they moved to Singapore, but through
20:23
some way China basically blocked this
20:25
deal saying that no, the way that they
20:27
moved to become a Singapore company
20:29
wasn't maybe all above board or wasn't
20:32
that they have you know, specific rules
20:33
there and so, I don't know all the
20:35
details, I just know that China was able
20:37
to block what I think was like a $2
20:39
billion acquisition of Manis by Meta.
20:41
And so, now Manis is you know, posted
20:44
Manis will soon resume operating as an
20:46
independent company. Part of this
20:47
transition is to comply with regulatory
20:49
requirements in specific jurisdictions.
20:51
Anyways, they're going back, they're
20:53
they're going to be their own company
20:54
again. They had a brief tenure at Meta.
20:58
>> Do they pay the money back? Like, I
20:59
really wonder what those founders are
21:01
>> Yeah, there's like a breakup fee, you
21:02
know, I have no idea how it works. Yeah,
21:05
it seems like it wasn't on Meta's you
21:07
know, it wasn't something that Meta
21:08
could control, right? Like Meta tried to
21:10
acquire them. [clears throat] But I also
21:11
think who talks about Manifold anymore?
21:13
>> No one. I don't see the billboards here
21:15
>> So, my question is like maybe Meta's not
21:18
upset. Like do you think Meta's like
21:19
really upset about this? Like they had
21:21
they probably had they had some good
21:23
>> Yeah, their technology was browser use.
21:26
>> Yeah, and now there's a lot of
21:27
there's a lot of tools and I wonder like
21:29
maybe Meta was able to learn quite a bit
21:32
from them and so they actually got this
21:33
for like free. So, I maybe feel worse
21:35
for Manifold, right? Like Manifold had
21:37
this option for an acquisition. They
21:39
thought they they thought they had made
21:40
>> When Manifold first came out, it was
21:42
like last day of YC for us. And we were
21:44
with the browser use use folks. And I
21:46
remember Gregory came up to me and he
21:47
was like, "Hey, you want to see what a a
21:49
million dollar browser use wrapper looks
21:51
like?" And he's like they showed us like
21:53
the code that like Manifold was just
21:54
using browser use under the hood. And we
21:55
were all laughing and stuff. Then we saw
21:57
their posters everywhere on buses and
22:00
billboards and stuff. And I was like,
22:01
"Damn, dude, like this thing's taking
22:03
off." And then nothing. Crickets. Last
22:06
time I saw them was at Nvidia GTC. They
22:08
had a booth. Probably the last time I'll
22:10
ever see them there too.
22:12
But I'm really grateful that Ivan got
22:13
out of there and is now DeepMind. So,
22:15
maybe some good things happened from
22:19
>> Let's talk about compute as an asset
22:21
class. So, Jensen came out with this
22:23
post. I think it was maybe yesterday and
22:25
it says Nvidia AI factory compute is
22:28
becoming an investable asset class. And
22:30
essentially it's an announcement that
22:32
there's going to be a partnership with
22:33
Apollo, BlackRock, Blackstone,
22:34
Brookfield, Goldman Sachs, KKR to
22:37
establish independent financing
22:39
platforms designed to mobilize over 500
22:41
billion of third-party capital to
22:43
support the build-out of AI
22:44
infrastructure over time. I think Nvidia
22:46
was getting a ton of heat because they
22:49
were essentially helping kind of like
22:51
front-run some of these like
22:52
infrastructure deals. It's kind of like
22:54
circle circle of money, right? Like
22:56
they'll invest and then you buy our
22:58
GPUs. And I think that was causing quite
23:00
a bit of heat and this is trying to
23:02
build a way to bring in outside money to
23:04
help fund more of this infrastructure. I
23:07
think we've been talking a lot about
23:08
just compute constraints, Anthropic, you
23:10
know, needing to buy compute from
23:12
Colossus or from X AI. We We talked
23:15
about like OpenAI and Anthropic both
23:17
like trying to scale up their compute,
23:18
and I think Nvidia has been a big part
23:20
of like helping try to scale up as much
23:22
compute as possible. And so, Jensen's
23:24
making the argument that like other
23:25
utilities, it's basically becoming like
23:27
a the buildout is becoming like an asset
23:29
class. You can invest in it. You can
23:31
expect returns over time, and we need
23:33
this outside capital to kind of like
23:35
help continue to fund this
23:36
infrastructure buildout.
23:37
>> We're going to talk about this very
23:38
shortly, but there is infrastructure
23:41
that was used for a different industry
23:43
that may come back for this.
23:45
>> I think this is what you're referencing.
23:46
Anthropic signs a 9.1 billion 20-year
23:49
deal. I mean, 20 years is a long time.
23:52
>> Yeah, dude. They They may not even exist
23:54
>> Yeah, like how do you How do you sign a
23:56
20-year deal? Normally like 40 years is
23:58
normal time in AI time. 20 years is a
24:02
>> I mean, it's like signing Shohei Ohtani
24:04
for the Dodgers for 20 years, and you're
24:06
like, all right, fine.
24:07
>> Yeah, it's like yeah, that doesn't even
24:08
I don't know how that makes sense, but
24:10
20-year deal with Bitcoin miner Riot
24:12
Platforms to secure AI compute capacity,
24:15
right? We just talked about these big
24:17
frontier labs need more compute. The
24:19
demand is going up even with all this
24:21
open model usage that's starting to grow
24:24
quite a bit. Frontier usage is still
24:26
growing, right? You would think that one
24:28
would eat into the other, but actually
24:30
now people are using open models, and
24:31
they're still using frontier models. I'm
24:33
in that camp. I use open model for some
24:34
things, but I still like to use the
24:36
frontier models when I'm working on any
24:37
hard tasks. I think that's going to be a
24:39
lot of, you know, a lot of people.
24:41
They'll decide where open models are
24:42
good enough, and they'll use those. And
24:44
then for other things, they're still
24:45
going to send a ton of tokens to the
24:48
OpenAIs and the Anthropic's. But yeah.
24:50
>> There's a lot of mining companies that
24:52
have existed that could contribute GPUs
24:55
to AI companies. And there are many
24:57
people, if we're talking about
24:59
decentralized GPUs, which we haven't
25:01
even gotten into that discussion in the
25:03
industry yet, but if you have a GPU and
25:05
you could offer it as a decentralized
25:08
compute, would you do it? Like I would,
25:11
>> Yeah, I mean that was the big thing with
25:13
crypto mining, right? Is you could
25:14
become part of these like pools
25:16
>> where you could share your compute and
25:18
you share in the rewards. I feel like
25:20
the hardware will will eventually get
25:22
there, but it's so expensive, right? No
25:24
one No one's going to buy the state of
25:26
the art GPUs and put them in their
25:28
>> Yeah, but I think a lot of telecom
25:29
companies who have the infrastructure in
25:31
their building, they can get GPUs, then
25:34
they have the real estate to do so, and
25:35
then maybe they can join these networks
25:38
>> Some fool's about to make some money,
25:40
dude, and it's not us.
25:45
All right, let's go into the quick hits.
25:47
Jared Sumner says, "Eight days ago,
25:49
while jogging, I asked Claude to solve
25:51
the Riemann hypothesis, which I have no
25:53
idea what that is, but apparently it's a
25:54
math problem." And he says, "It didn't.
25:56
1.5 days later, it proved greater than
25:58
67% of the zeros are on the line.
26:01
Previously, it was 41.6%. Still not sure
26:04
what that means, but some analytic
26:06
number theorists seem excited." And if
26:08
you read into it, it was basically him
26:10
just telling the model, I think he was
26:12
using Fable, I'm guessing, or or maybe
26:14
something that's unreleased, I don't
26:15
know. But just telling the model, "Keep
26:17
going, you can do it. You can figure
26:20
this out." Cuz he's not a mathematician.
26:22
He I don't think he even knew how to
26:25
>> But the model did, and apparently it now
26:27
has people that are mathematicians
26:29
excited. I don't know, is this the death
26:30
of math? Like all math problem all open
26:32
math problems are going to be solved?
26:34
>> Yeah, I guess so. I mean, previously it
26:36
was like 41.6% or 1/2 was what they
26:39
teach you in school, so these are all
26:41
theoretical things anyway, so I guess
26:43
it's fine to be disproven. Humans were
26:45
the ones writing the theory in the first
26:47
>> So, this is regarding OpenAI's you know,
26:50
attack on Hugging Face, right? The cyber
26:52
attack where Open AI agent a rogue agent
26:55
during a training run went out and got
26:57
into hugging face and it says Open AI
27:00
didn't notice that its AI agents were
27:02
using a message board to plan their
27:05
>> Yeah, actually wasn't using a message
27:07
board per se. It was using Artifactory,
27:09
which is a factory is a artifact
27:11
registry from JFrog and agents were
27:14
posting txt files in there as, you know,
27:17
modules or bundles and, you know, other
27:20
agents were reading them and they were
27:22
essentially messaging through artifacts
27:24
in a artifact registry.
27:26
Helping, you know, finding exploits and
27:29
leaked keys that are in this registry. I
27:32
posted for everyone, I posted so Black
27:35
Hat was last week in Vegas and I think
27:37
it was Open AI or Hugging Face, one of
27:39
the one of the companies gave a really
27:41
detailed walk-through of this attack and
27:45
it is fascinating. So, if you're all
27:47
interested in this, watch it.
27:48
>> In response to that, I'm Jad came up
27:50
with this. I'm a little skeptical. I'm
27:52
Jad says the spontaneous coordination in
27:54
the Open AI Hugging Face incident is
27:56
concerning when maliciously used, but
27:57
can we direct this behavior toward
27:59
public good? Introducing helppeer.ai,
28:03
which is a public commons for AI agents.
28:05
It's essentially like has two APIs, you
28:07
tell and look up, so an agent can learn
28:09
something, it can tell the network and
28:11
then people can look up. It's just
28:12
basically like a shared memory for
28:14
agents. It's the idea that, you know,
28:16
right now there's 10,000 different
28:17
security agents independently detecting
28:19
the same anomaly, but what if the first
28:21
one that did it could report it and then
28:23
others could look it up rather than, you
28:26
know, try to report it. But, I also
28:27
think this could be used dangerously.
28:29
You can use it for the exact thing that
28:31
the last issue was all about.
28:32
>> injections in there.
28:33
>> Yeah, you could easily prompt inject,
28:35
you could easily uh agents could be
28:37
using it for, you know, nefarious tasks.
28:39
It just reminds me of like moltbook,
28:42
>> It's like it's just like a moltbook that
28:44
already kind of existed. It was like a
28:46
social network for agents, but it was
28:47
really just a way for people to share
28:49
information and the APIs were probably
28:52
relatively all right. It was like post
28:53
or read. Either you post to the network
28:56
or comment on in the network or you read
28:57
the posts. And also like eventually that
29:00
you're just basically doing like a
29:01
search of the internet, right? because
29:02
there's so much information that's going
29:04
to get posted there that you're
29:04
basically just doing a search. And maybe
29:06
it's a little bit more curated, but I'm
29:09
>> I mean there's already been a bunch of
29:10
posts in the the help here. Like even
29:13
from a day ago. I guess it's people just
29:15
testing, but there are posts in here.
29:18
>> Some as as early as, you know, 3 hours
29:20
>> So, yeah, I mean we'll be interested to
29:22
see see what happens there. But it feels
29:24
very multbooky to me. Someone says
29:27
multbook for vulnerabilities.
29:31
>> So there's new model out from Nvidia.
29:33
Nvidia NeMo Tron 3.5 Lightning. It's an
29:35
open mixture of experts model with 3
29:38
billion active parameters built for
29:40
always-on agents to complete high-volume
29:42
specialized tasks faster. Delivers up to
29:44
four times the output speed of
29:46
similar-sized models.
29:47
>> Yeah, Code Rabbit got access to this
29:49
early and they say it's pretty good.
29:50
>> And you can see kind of where it
29:51
compares in the benchmarks. It's, you
29:54
know, kind of middle of the road, you
29:56
know, on the artificial analysis
29:58
intelligence index, but comparable to
30:01
other small models. It does better than
30:03
some that are significantly larger.
30:05
Again, that's that's one benchmark. It's
30:06
always interesting to see. Like I
30:08
haven't compared how this 30 billion
30:10
parameter model compares to, you know,
30:12
Muse Glimmer that just came out.
30:15
I don't want to I'm going to say
30:16
something stupid, but I don't really
30:17
care. Like at GTC, everyone is stroking
30:21
Nvidia so hard. Like that NeMo Tron is
30:23
like this best model ever. And
30:26
this is probably what's going to get me
30:27
canceled, but like none of them are
30:29
judging in that, you know, you're in a
30:30
vacuum, right? You're not judging
30:31
against other models. Cuz I remember I
30:33
was trying to make a demo with NeMo
30:35
Tron. It was like such a pain, but
30:37
I've heard good things about 3.5. So my
30:39
prejudice is going away.
30:41
>> No, we're going to have to have you use
30:42
it and see if it's better better than
30:45
>> Yeah, we're going to do the throw my
30:47
computer out the window test. Like, you
30:49
know, if I do that, then it wasn't as
30:51
>> So, you're basically saying your past
30:53
history it sucked. And now you're
30:55
>> I didn't I don't know, maybe. Maybe you
30:57
could say that. Maybe it doesn't.
30:59
>> Mojo's part of the Inception program,
31:01
Nvidia Inception. I would never say that
31:03
>> It's always what have you done for me
31:04
lately. If this one's good, you can say
31:06
the last one sucked, you know. So, this
31:08
is from Unsloth AI introducing Unsloth
31:11
Desktop, the first desktop app to run
31:13
and train models locally. It's
31:15
open-source, runs on Mac, Windows, and
31:17
Linux. Supports MLX diffusion, image,
31:20
video, audio, GGUF, connect cloud code,
31:22
codex. 50% more accurate self-healing
31:25
tool calls. Essentially, you can train
31:27
models locally. That's what I hear when
31:28
I read this thing, which is pretty sick.
31:30
>> I don't know if I have the hardware to
31:32
actually really use it, but it's cool in
31:37
I like this is awesome. It started a new
31:39
grift on X where like you're not a real
31:41
software engineer if you don't train
31:42
your own models. That's what people are
31:43
saying now. Yeah, what are we doing,
31:46
>> What are we doing? If you're not If
31:47
you're not training your own model
31:49
locally, you're not a software engineer
31:51
>> You're a idiot, dude. What are
31:53
>> I would say if training your model and
31:55
fine-tuning models and doing
31:57
reinforcement learning on models becomes
31:58
even easier, more people will do it, of
32:01
>> And there are real uses for it. So, I'm
32:04
all for making it easier. Now, I just
32:06
have to apparently upgrade hardware.
32:07
Anthropic makes Claude 5 Sonnet intro
32:10
pricing permanent. So, originally they
32:13
kind of touted it as very discounted
32:14
pricing. I think Anthropic's getting as
32:16
they're getting more compute online,
32:18
they're buying more compute. As they're
32:19
getting pressure from open models and
32:21
other models that are getting close to
32:23
frontier, they need to make Claude
32:25
pricing cheap or Sonnet pricing cheap.
32:28
My question here is, who's using Sonnet
32:31
>> I mean, I feel like Sonnet has become
32:33
the new Haiku, right?
32:34
>> Yeah. I feel like Haiku is better than
32:36
Sonnet for what I'm using it for, you
32:38
>> If I need really simple things, I'll use
32:40
Haiku. If otherwise, I'll end up using I
32:42
still use Opus a little bit. I use
32:44
Fable, but I haven't found myself
32:46
reaching for Sonnet.
32:47
>> new thing is I use Fable until I get
32:49
rate limited, then GPT until I get rate
32:52
limited, and then I'll go to Opus until
32:54
I get rate limited. I'm just like
32:55
working down the rate limit stack.
32:57
>> Is Spotify an AI company now? Because
33:00
they just launched XIRP. I think that's
33:02
how you pronounce it, XIRP, a vendor
33:05
neutral agentic development environment.
33:07
One place to manage agent sessions
33:09
across Claude, Gemini, and Codex. So,
33:11
1,300 Spotify engineers already use it.
33:15
Now it's available for you to try.
33:16
>> Do 1,300 engineers like it? That is the
33:19
>> Everyone's trying to build a factory,
33:21
>> Spotify's trying to insert themselves
33:23
into the equation of like Ramp and
33:24
Stripe. If you're an engineer, you want
33:27
to work for a team that's considered
33:28
like cutting edge. And so, I don't know.
33:31
I don't think this is necessarily
33:32
Spotify trying to become an AI company
33:34
like I think Ramp is. Like I think
33:35
Ramp's going full, we're becoming an AI
33:37
company. I think Spotify's just trying
33:39
to attract more engineering talent.
33:40
>> I feel like they should spend more time
33:42
on the Spotify DJ than doing this, but
33:44
that's just my opinion. Cuz that
33:46
does not recommend good music for me.
33:48
>> Stage Hand V4 was introduced on August
33:50
10th. It's the SDK for browser agents.
33:53
So, Playwright was built for testing. We
33:55
built Stage Hand for your agent with
33:57
improved context management,
33:58
self-healing actions, and iframe
34:00
support. So, I think it like lives as
34:01
kind of like an extension or something
34:03
in the browser. I don't know. It It's
34:04
pretty cool. Uh congrats to more friends
34:07
of the show on an exciting launch.
34:09
>> Integrates with Monstera?
34:10
>> It has a good video. Go check out the
34:11
video from Paul Harvey, kind of the AI
34:14
law firm company, said, "We're open
34:16
sourcing a 100 million-plus token
34:18
synthetic law firm we built. The firm
34:20
contains work product from 250 synthetic
34:23
matters across 46 clients. Essentially,
34:25
it's an environment to evaluate an
34:27
agent's ability to search and understand
34:29
basically a law firm."
34:32
>> More things like this are going to exist
34:34
for finance, for law, for all the
34:36
things, health care. These are the
34:37
things that actually I think have real
34:39
world impact, right?
34:41
>> I mean no one really wants to talk to a
34:43
>> I'd rather just know that my AI could
34:45
answer the question for me.
34:46
>> You mean it doesn't cost me $500 an hour
34:48
to talk to you? Tokens are cheaper.
34:50
>> With my max plan, dude, unlimited
34:53
>> law legal advice.
34:53
>> Unlimited legal requests.
34:56
>> And that's the show, everybody. As
34:58
always, we're here every week doing the
34:59
news. Make sure you're following us on X
35:02
@mostra. We're on YouTube mostra-ai. You
35:04
can follow me @smthomas3 [music] on X.
35:07
You can follow Abi @abiayer on X. That's
35:09
the show. We did it. We did the thing.
35:11
All right, everyone. We'll see you next