Opus 5 released! Is it better than Fable?
[music]
Heat. Heat.
Hey, [music]
hey, hey.
[music]
Heat. Heat. N.
[music]
>> [music]
[music]
[music]
[music]
>> Would you like to be a guest on the
show? Visit masterstra.ai/gest.
Do you have a hot takeache, a problem
you cracked, a cool product, or a demo
worth watching? Share them right here on
Agents Hour.
[music]
[music]
It's Monday noon. The time is here.
Shane and I be loud and clear. Pacific
vibes, we bring the heat. AI agents
can't be beat. AI agents now. Let's go.
Losing guess the big show. Solving
problems we do alive. Staying focus
[music]
that work for you. Making moves. We see
it through. Every week a brand new show.
Tag along and watch your knowledge grow.
[music] Agents out. Let's go. News and
guests, the big show answering
questions. We arrive every Monday. Come
alive.
We want reviews, but only if it's a
five. We share the drama. We got the
drive. Stay in [music] the loop. It's
the place to be. Shane and I be setting
you free.
>> [music]
>> agents here to stay. Tune in Mondays
make your day. From the news to the
problem solve
world, get involved.
[music]
>> [music]
>> Every week in AI, something insane
happens
>> and there's so much drama. Every Monday,
we break it down live.
>> We do the news. We bring on guests
building in the space.
>> And we go deep into the stuff that
actually matters.
>> Agents Hour, every Monday, noon Pacific.
follow. Don't miss it. Peace.
>> This is Agents Hour with Shane Thomas
and Abby Ayer.
Hello everyone and welcome to Agents
Hour. It is Wednesday, July 29th. Today
I'm here as always joined by my
co-founder, friend, co-host, Obby.
What's up, dude?
>> What's up, dude? How are you?
>> Good, good, good to be back home. Last
week, if you saw us live or watch the
recording, we were in London and so that
was fun. So, we're going to talk about
that a bit. We're going to talk about
all the news. There's a lot of drama as
always,
a lot of, you know, a new model release
to talk about. That's always fun. And
we're going to be, you know, we'll see a
short demo and talk about some of the
stuff we're launching over here at MRA.
Before we get started, this is a live
show. So, if you're watching live,
normally do this on Mondays. We're doing
it on Wednesday because of travel. We do
it every week. But you should chat with
us. So, drop a message.
>> Say what's up.
>> Yeah. Say what's up, say hello, ask
questions.
interact, give us your hot takes, we
will pull them up on the show. And if
you're watching after the fact, thank
you. You know, give us that five star
review whether you're on Spotify, Apple
Podcast, YouTube, wherever you are
watching us or listening to us from.
How you doing, dude? How is you're still
in France, right?
>> Went back to Paris after the conference
and uh going home soon. I am ready to go
home.
Yes, it'll be nice to have you back uh
you know back in the United States,
similar time zones at least to me.
>> Yeah, dude. I can't wait to speak
English, dude.
>> It's going to be great.
>> Uh but maybe that's a good segue into
let's talk a bit about TSAI London. That
was a conference we held last week. We
had, you know, 2,000 plus people
virtually signed up. We had, you know,
hundreds of people in person at convene
in London. It was a good time, I guess.
What was your takeaway from the from the
whole like conference week?
>> Um, a couple. So, we got the whole team
together, which was super fun. We were
working towards the things that we were
releasing.
It was stressful at times, but I think,
you know, uh, when you have everyone
together working in person, which is
honestly the most fun part is just
working in person, shooting the [ __ ]
getting [ __ ] done. Um, but I was so
impressed with the conference. One, the
venue was beautiful. And if you've been
to our San Francisco conference, it's
the same venue thing. So, but this one
was just, I mean, a lot nicer. feels
like the quality of the talks were so
good and there was a couple themes that
were just essentially blatantly present
through everybody's talks which was like
software factories.
>> Yeah, I think that was a big one. I
think what I was most impressed with was
the talks as well. of course like
meeting people the net the the the
hallway track is always very fun because
you get to talk to people that are using
>> MRA that are building agents that are
running into a lot of the same problems.
So I think there's one something
therapeutic about it just talking and
relating to other people's problems but
also just getting to know people that
are all it felt kind of like a community
event in in a lot of ways
>> of not just MRA but just people who are
interested in building agents working
with AI working with Typescript you know
a lot of them use MRA of course but not
all of them and it was just great it was
a great like community feel to it felt
more like a community than it did I
think in San Francisco
isn't to say the San Francisco ones
aren't fun, but that one, you know, it
was almost like a different persona
where San Francisco felt like a lot of
like startups, tech, and there were
certainly a lot of that, you know, quite
a bit of that, but it also just felt a
bit more of like a community vibe. But
the talks impressed me a lot because it
was, you know, we had talks at all
different levels, but there was a lot of
like people on the ground shipping
stuff, sharing real
>> things that they've learned. But I think
those are the types of things where you
can actually pull out a couple
actionable things that you're either
going to try that you, you know, might
have ran into just really like useful
type of talks like they're much more
practical rather than just the high
level. And we we did have a little bit
of that which is good. It's a good mix
of like
>> varying levels of uh like in the weeds
versus like what's coming. And I think
that's those are always great to have
like the variance especially in a single
track conference.
Yeah. And we met some of you, the
listeners of the show, came up to us
during the lunch break. Um, very
grateful for all the nice words you all
said. Um, remember we met, you know,
longtime listener in person, John T.
Brook. Uh, so maybe he's watching now or
he will be watching. Shout out to you
and many others. I I mean I think there
was at least a half a dozen people and
it's always weird because they come up
and you know you they're like I feel
like I know you and of course we've
never met at least not in person but
it's great to meet people that actually
watch the show on a weekly basis that
you know we've seen in the chat maybe
but we've never met in person
>> and you don't know that they're going to
be you know where they're all from all
over right so there's obviously like
people in the you know in Europe or in
London and there were people that
traveled in quite a I guess as well, but
most of the people were from the London
area. But yeah, it's great to meet
people in person. You know, we we
wouldn't do this if we didn't, you know,
think it was valuable and we didn't have
people that actually enjoyed watching
it. So, appreciate all of you that did
come up and say hello.
I think we do have uh so Yan, producer
Yan put together a video for as kind of
like a conference recap. So, I figured
we should watch that. So for people who
>> have listened to this and have massive
FOMO,
>> uh you can watch the conference recap
which will give you additional FOMO, but
then you could still watch the talks
afterwards. And so we'll tell you how to
do that here after we watch the video.
[music]
>> [applause]
[music]
[music]
>> Heat. Heat.
[music]
[music]
[music]
>> [music]
[music]
[music]
[music]
>> Nice. Nice work, Yan.
>> That was it.
>> Yeah. So, if you you know, if you feel
like you missed out, you you did, but
you can go watch the if you watch the
whole live stream. It's on our YouTube.
But if you don't want to watch the
entire thing, we are going to be kind of
taking all the talks, cutting them up,
doing a little bit of editing to make
them, you know, a little even tighter.
And we'll be posting those on YouTube
over the next few weeks as well. So if
you missed it, you can still, you know,
learn from some of the talks, learn
from, you know, some of the things that
we saw in person. So don't feel too bad.
Uh you can still participate in some
ways. And yeah, Sebastian, thanks for
watching. Thanks for checking out the
the live stream.
One of the things we did,
and we do this every time we have a
conference, this is our third, this is
actually our third TSA comp, which is
kind of wild to think about.
Yeah. So, the third conference we've
done,
>> second this year.
>> Yeah. And
>> yeah, we don't have a ne we don't have
another date planned, but there will be
another date at some point. So, know
that if you missed out, there's another
chance. But something we do at every
conference is we talk about the cool
stuff that we're working on. So for
those that are new to the show, you
know, we're two of the founders of
Mastra and we like to ship things over
here that developers like to use and so
we like to talk about what those things
are. A lot of those tie into trends
we're seeing in the industry. A lot of
those tie into very closely what our
users and customers are kind of asking
us for. So, wanted to highlight some of
the things we launched last week and uh
we kind of relaunched them this week
because we talked about them at the
conference, but then we uh we actually
promoted them more broadly. So, if you
watched the conference live, you kind of
get the sneak peek and now we're
actually announcing them to the world
throughout this week. So, why don't we
do that and then I hear Abby, you might
have a demo for us.
>> Yeah. And uh show you what I'm cooking.
>> All right. So,
first things first,
we're going to work backwards, I guess,
because why not? So, first thing first,
this was launched today. We added
environments and regions to master
platform. So, we had a lot of customers
ask us and it might not be a surprise,
you know, we were in London for a
reason. A ton of our customers,
Sebastian, you know what we showed as
well, a ton of our customers are in EU,
right? Are in European region, whether
it's, you know, UK, EU,
and they want their data close to them.
And so we've been working for quite a
while just to add regions so you can
kind of decide where when you deploy to
Masha Platform where you're deploying to
and also environments so you can have
production, staging, preview
environments. And I think that it opens
up a ton of flexibility for actually
deploying your MRA agents and your MRA
applications to our platform.
Any comments on this, Obby?
>> Oh man, I'm just super stoked. Um
because we were running some US
deployments before this launched, let's
say, and I it was just so terrible the
latency. But if you have all your stuff
in your region, it's amazing. Um, and
then coming up next,
maybe I'll tease it a little bit, but we
will do multiszone in the future. If you
are a global company, you'll need that.
>> Yeah. And I think it was easy for us to
like feel the problem firsthand because
we were in EU and I I had a bunch of US
deployments and, you know, latency is
not terrible, but you can feel it,
>> right? Yeah.
>> And you don't want to feel it.
>> So, you definitely don't want to feel
it.
>> All right. Now, let's talk about what we
announced yesterday.
And it's sometimes hard for me to
determine what tab to share. So,
hopefully this is the right one. All
right. Got it. All right. So, we
launched Trace Intelligence.
And
maybe we'll just kind of play play the
video while we're talking through it.
Essentially what it is is it allows you
to take a large amount of traces. It
analyzes those traces into clusters and
themes
and then from there it also kind of
connects patterns across those kind of
clusters that it comes up with. So you
can figure out what's the goal, what's
the behavior, what's the outcome, what's
the sentiment of
clusters of traces and so you can rather
than looking at thousands of individual
traces trying to dig through the data or
asking your agent you know to look
through thousands of traces and process
all that we determine the clusters for
you and then you can manually do it with
this UI which is cool like or you can
just like tell your agent to do it but
they're not looking at tens of thousands
of traces or thousands of traces they're
looking at you know a dozen or two dozen
clusters and then your agent can decide
which clusters to go in. So you don't
have to pay for
>> you know I saw a lot of people saying
like well I use Fable to analyze traces.
It's like not not at scale you don't not
if you're pay
>> with your max plan. Sure.
>> Yeah. With if you if you don't exceed
your max plan yeah just send Fable on
like a group of traces you're good.
>> But if you're if you actually have real
data you're not paying fable prices or
you don't want that level of inference.
So there's a whole bunch of cool stuff
we do under the kind of under the hood
with like deterministic matching, some
ML pipelines, some like inference to
come up with these clusters and then
that way you don't have to pay for top
level intelligence to analyze tens of
thousands of traces.
>> Yeah, some clever things that Eric and
team pul pulled out there. Um and dude,
it went pretty viral from our standards,
I guess.
>> Yeah, if you look at it, it's definitely
kind of taken off. So, it is in beta
right now. So, if you want access, if
you're using Master Platform, let us
know. We'll get you access to it.
[sighs]
All right. And then the one that is
probably most exciting to me at least,
and I I would I'm going to guess it was
most exciting to you, too, just because,
you know,
>> we like building developer tools and we
like building tools that we use
ourselves. So, if you've used Monster
Code in the past or you've heard of
Monster Code, you know that that's a
tool that we built to help our team ship
faster. That was always the goal with
Monster Code is what's the coding agent
we want to use that uses all the master
primitives that you can use to build
agents and applications.
But we did the same thing, but this time
it's like a level. It's a a level up on
top of a master code. So we announced
masteractory. So this was on Monday this
week. Sam posted this software
engineering is becoming a hierarchy of
loops. So today we're launching
masteractory a system for agents to take
software from issue into production and
you can just get started with npm create
factory. And what it really is is just
like a whole almost like a master
template of sorts that uses master code
under the hood. You can hook it up to
GitHub to linear. It ingests issues. It
automatically starts working on those
issues. It pauses when it needs
feedback. You can steer it. You can
steer the running agents that are
running in a sandbox kind of starting to
like automate the whole flow for you.
And it's just the beginning. It's still
we're kind of calling it alpha because
we want to be able to break things
because we're making it better.
>> But it is like we're using it every day
and a lot of people are starting to I I
think start to see the vision of where
it could go.
>> Yeah. Yeah. And like software factories
are a big hype term right now. And the
reason why we called it mra factory is
we don't necessarily believe that it
stops with coding agents. Um that's why
it's a factory in general like you
should be able to do whatever you want
in this you know loop architecture or
whatever. Um, but as it stands right
now, it is designed for software,
the SDLC. Um, and we are dog fooding it
every day, much like we did with
Monsterra code. And it's really cool
because it's built on all the primitives
we've done already. There's no secret
sauce other I mean, there will be some
secret sauce in the future, but it's all
built on MRA and it's a new primitive
that is served by the MRA server. And I
don't know like the nerd engineer in me
architect engineer in me just like is so
proud of the fact that we have all these
primitives we put them together we get
mo code then we have an agent controller
that we extracted from mo code then we
added more different types of primitives
to then build the monster factory and
then that's like an entity that can be
served through the monster server that
we didn't even think we would build two
years ago. We were kind of like, what is
the Monsterra server? And then we were
like, you know what? We just do it. Um,
and then that thing comes back into
play. Our storage adapters, you can you
can it's just like MRA. You can use
different storage adapters. Everything's
an interface. It's how MRA is designed
already, just now it's a different
primitive called the factory. So that
was
>> and the idea that you know and we'll
continue to iterate on this but you can
run it in your like bring your own
sandbox bring your own you know like
observability like all the things that
make great like you know because this is
just built on top of it are going to
make factory great. So that and the
reason this came to exist is we kept
hearing in calls over and over again
people telling us they were using MRA to
build a factory and we had a lot of
pieces of this already internally so we
just put it together and made it
extensible so we can help our users so
they don't have to become you know that
they can build and own their own factory
without having to you know do all the
plumbing themselves right it's like we
kind of give you like here's the
baseline customize it to fit your needs.
You can build your own ramp inspect or
you know build your own Devon, right?
But you control it. You pay for the own
your own inference. You control the
keys. You can run it anywhere. You can
run it with and master platform if you
want as well. But ultimately it's yours.
And I think that's the cool part is
>> it allows you to build your own factory
and customize it to what you need. But
you don't have to do all the bits. you
can kind of like take a pretty good set
of primitives, customize it, and you're
good to go.
>> Yeah, it's very disruptive because, you
know, we've always wanted to build our
own Devon
uh internally. And the factory is more
than just a Devon. It is a automation.
It is a uh it actually makes you want to
use linear because you have everything
controlled. You want to write good
issues. It's a discipline too because
you know that you're automating a lot of
things, you know.
>> Yeah. That when that issue comes in,
>> it's going to get picked up immediately.
So, you should make sure like, you know,
only write an issue if you're pretty
serious about getting that thing
shipped.
Yeah, I think that's that's a really
cool part is just it's kind of this idea
of there's this dream and we're not
quite there yet, but factory gets us
very close where you don't have to um
you know it's this idea of like zero
bugs, right? Like no bugs. If a bug
comes in and it's detailed, we should
just start working on it right away.
Don't put it in the backlog.
>> Like either you fix it now or you don't
fix it and you wait till it becomes like
a burning issue. I think that kind of
thing starts to become more possible,
you know, with something like Factory.
>> Yeah. And we need to also pass the bar
test where Shane and I can go out
drinking and work still continues. Um,
and if we need to steer the agent, take
a sip and make it happen.
>> All right. So, you got us a quick demo.
We'll keep it short and then we'll jump
into the news.
>> Cool. All right. So, I've been working
on there's many factors to the factory.
There's work, which is stuff that has
not been uh maybe issues or linear
tickets or whatever. I'm not going to
show that today. I've been really just
focused on review. Um, and for us,
review is super important because we get
a bunch of contributions from the
community. But the factor, the limiting
factor is can we actually review it? We
use code rabbit and our this review is a
complement to any review code review
agents that you have. Um but as you can
see it is a canban style of of a board.
The intake is the in the work items that
are coming into the factory. And you can
see these are all the PRs that need to
be reviewed. behind the scenes there's a
review agent that reviews Mashra like we
do internally. So we wrote a skill
called the factory the factory review
skill and we have a lot of different
like just
the way we do things and we think the
way we do things is the way you should
do things but then in the future you
know you may be able to configure these
uh these agents that work behind the
scenes and so as PRs come in they are
automatically picked up and they start
being reviewed which is cool and I've
done a lot of review today 94. Uh before
the uh live stream started I was at 65.
So while we've been talking things have
been happening. Um so I'm just going to
show a couple things here. Um one I'll
just go to the settings. We are building
out this where you can have different
intake sources. I can connect to linear.
I can have more than one repository. I'm
just worried about Ma open source right
now because if you spend a week in
London, hella issues and PRs come in.
You can configure your model like what
is the default factory model. Um I also
added or we also added OOTH here. So I'm
signed in. Don't tell on me, but I'm
signed in with my max plan. Um probably
won't be kosher in the future, but right
now it is. So that's cool.
And if I go back here, I can just start.
This Alysia adapter has been sitting on
my mind for a while. So I'm just going
to click this and it's going to start a
session. It's a review session. So in
this review session, we spin up a
sandbox and then the agent will look at
the PR and then start reviewing.
And so you can see, you know, it has a
factory phase. This is the work item.
This is what's happening. And then we
have a factory skill which is very
detailed and looks like [ __ ] right now,
but we'll fix that display. And then all
the master bits are all the same. It's
just a web UI. So now it's writing
tasks. So what it's going to do, it's
going to triage the existing stuff. It's
going to check quality. And then it's
going to do something that's very
interesting. And the way we designed
this is we want it to feel like a the
senior person on your team is reviewing
the code. So what really matters is like
not just this change but what is the
history of the change or the changes in
this area and it'll go look in git
history to see how is this thing changed
and is the incoming thing an actually
valid thing to do and then it'll do
architecture review and then finally in
verdict it'll do an adversarial review
on your PR and I made it a little mean
So, it gets kind of mean. Um, not too
mean, though.
>> So, let me just show you an example of a
review that has happened.
>> And we're just seeing we're not seeing I
don't know if you share in multiple
tabs.
>> I'm about to share.
>> Okay, cool.
>> Something.
>> I think the coolest thing as you're
pulling that up or one of the coolest
things is you can steer the agent as
it's going, right? So you can actually
see the session. You can see what it's
doing. And if you want to ride the loop,
you can just coach it as it's running.
Just send a message. The next loop or
the next time it, you know, the agent
stops and pauses for a second to do the
next tool call or whatever, you using
MRA agent signals will get inserted and
you can just steer it and keep it keep
it going. So it
>> it allows you to let things run
completely autonomously or it allows you
to like pay close attention and kind of
guide it as it goes. So you it gives you
the flexibility to do it the way you
want to.
>> Yep. So this is like a review on
Daniel's PR and immediately it has a
bunch of requirements for it to be
approved. So it's requesting changes.
There were some merge conflicts that
need to be um settled. It agrees with
code rabbit's
um review as well. So it takes into
account the other reviews that are there
um just to say like hey like you should
be doing these things has some optional
stuff. It also verifies everything that
you claim to have done. You know a lot
of PRs these days say oh I did all this
this is the test plan. It's like okay
cool. If that's the test plan, let me
run that [ __ ] automatically
and then go for it. And then I did a
followup because I think Daniel like
pulled in some changes. And then there
you go. And I guess this will be good
for review. And there's many of these.
So what I'm doing right now is I'm
running it on every single PR in our
repo. And then from there, you know,
we'll see what happens.
>> And can it approve?
it can approve. It has approved many PRs
today and many community PRs
>> and I think that's the thing that's
going to cause people to either be
excited or scared.
>> Yeah.
>> And and I think and but ultimately, you
know, it's still your choice like
whether you need just the approval from
the bot. You still want the human
approval. I think the the answer for us
is it kind of depends on what surface
area it touches, right? If you're
changing framework code, we're still
going to have humans look at all that,
right? because it we don't fully trust
everything that the bot's going to, you
know, going to do. But there's probably
other surface areas that if if the bot's
happy, you know, if if if the factory is
happy, we're happy, you know.
>> Yeah.
>> So, I think it kind of depends.
>> Yeah. We're going to like like we always
do, we're going to ride yolo mode to
learn and then we're going to find out
where this thing does not work and then
give guardrails for that. But we will
run yolo mode for I mean for the
foreseeable future just to see what it
can do.
>> Absolutely.
>> Um there's a there's a more yolo part of
this which is like issue creation,
right? If you give us an issue, we need
to triage it, start working on it
automatically. And uh yeah, we're just
ironing out the kinks there now. So I
mean all this is going to be dope.
Right.
>> And that that's one thing to flag is
it's really cool if an issue comes in,
it starts working on it, it gets to a
review, a different agent reviews it,
right? You can customize that
>> and then basically they're almost having
like a back and forth of sorts without
Yeah.
>> You know, you don't have to have human
intervention if you don't want, right?
You can kind of get it to approved PR
state where
>> it is actually approved without you
having to even, you know, touch
anything.
>> Yeah. And the memory is shared is
observational memory. And it might, you
know, if you saw in that review, it's
very pedantic to tell a reviewer or a
contributor or whoever that you have
merge conflicts. But the reason we do
that is if a agent started the work,
when it reads the review, it can just it
doesn't have to go do a a tool call to
figure out that it has merge conflicts.
It'll just be, "Oh, I have some merge
conflicts. I'm going to start working on
that right now."
>> Yep.
All right. And with that, you know, we,
you know, this is a live show, so
Medigames, thanks for tuning in.
>> Thanks for tuning in.
>> Thanks for hanging out. And we talked
about TSAI London. We talked about
recent master launches. Yeah, if you
want, if you do want to use the factory,
npm create factory.
So, go ahead and
>> it's an alpha. Give us feedback.
>> Yeah, it is an alpha. There are rough
edges. There are many rough edges. It's
getting better every day. But hopefully
you can see some of what we're excited
about when cuz we'll be talking a lot
about it, I imagine, over the next month
or two.
>> Yeah.
>> But with that, should we get in the
news?
>> Let's do it.
>> Let's get into it.
All right, welcome to Agents Hour. We're
doing the news. We do this every week.
We're doing it on Wednesday this week
rather than Monday because of some
travel things. But it has been it's been
a good week for news. There's been
there's a lot to talk about.
little preview for what we're talking
about today.
The first thing the first thing to talk
about is this idea of if you've been
paying attention, you know, Enthropic
got, you know, kind of like copyright
suit. They got they had a settlement is
like $ 1.5 billion dollars or something
for like book publishers and then it
kind of came out and this is like after
the fact and I don't think, you know,
Enthropic wanted this to come out or at
least there were some internal rum
rumblings or memos of where they didn't
want people to know this. I think they
called it like Operation Panama or
something like they don't want people to
know that essentially what how they got
the information is they were just like
you know the books they couldn't get
online they were just like buying the
copies ripping the spines out and then
ingesting all that data which maybe in
some cases I'm not too worried about
like if it's like a normal book cool
like I guess whatever like if that's
what you got to do to train it like I
don't feel great about it but you know I
anyone can go buy that book again. But
there's a lot at least a number of like
one only one of one copies or very
limited copies that are not that are not
in circulation anymore because they they
kind of essentially destroyed the books.
>> Yeah.
>> What a crooks, [laughter] dude.
>> So, I mean that doesn't make you feel
good, right? Like there's some like
really old books that were probably cost
them a lot of money to buy. Might have
been might have paid $500 for that book
and all they did is just then destroy
the book.
to get the information from the book.
And now no one else, you know, arguably
if these are like some of these are one
of one and I think of course those are
the extremes. I don't think that's most
of the books, right? But even the fact
that they did a little bit kind of
doesn't sit right.
There were there were probably ways to
get the information without having to
destroy the book is all I'm saying.
>> Yeah.
>> Just would have been more inconvenient.
Would it cost more money to get like to
pay the people
like, you know, if the lawsuit's like
billions of dollars and you're spending
a ton of money on the books and then
burning them or whatever, wouldn't it
have just been cheaper to go to each
author and get the rights?
>> Yeah, may
it would have been expensive in time, I
think, is what they basically decided.
And I think the problem is some of these
things they probably couldn't even get
digital copies or whatever. So they'd
have to like they'd have to buy the
book, right? But then maybe just don't
destroy it, you know, just, you know,
take a little more time, keep the book,
put it back in circulation if someone
wants to buy it. Like I don't know.
All right, we got to talk about the
OpenAI security incident.
So this came out on July 21st. So this
is kind of like late last week or kind
of mid to late last week. We had a
significant security incident during
evaluation of our models and we're
sharing what we've learned so far. Cent,
you know, essentially they're partnering
with HuggingFace to try to help figure
out what happened.
But what happened? How did
>> Yes. So they were all right. Allegedly
everything is allegedly right now. um
they're running a security bench and
allegedly or maybe confirmed or whatever
that essentially GPT 5.6 or a model that
we do not know about yet broke out of
the parameters and hacked hugging face.
So it pretty much ignored its uh
directive and did whatever the [ __ ] it
wanted. [laughter]
And and then there's a lot of things
came out after that, right? Hugging face
tried to figure
>> figure out what was happening because
they detected something.
>> They tried to use open AI and anthropic
models to like figure it out, but
>> they they were blocked because of
guardrails. Those models didn't want
they thought they were, you know,
potentially being used for some kind of
like cyber security research or
something that shouldn't have been
>> able to be used for. So they said, "No,
we can't help you with that." So they
had to go to GLM 5.2
>> open models.
>> They had to use an open model in order
to like get to the bottom of the issue
and figure out what was happening and
like start to block or start to like at
least remediate the attack. Open AAI
obviously like then figured it out, you
know, it got shut down or whatever.
There was some like speculation that the
model had planted some other things on
the internet for like instructions for
itself for future versions of itself.
There's like, you know, some really like
Terminator type stuff that
>> yeah,
>> hard to know what's true and what is
speculation at this point, but
>> also hard to know how much of this is
[ __ ] or not. You know what I mean?
>> I mean, yeah,
>> this is media, you know, media.
>> It definitely like happened after, you
know, the Kimmy Kimmy launch where I
think people are, you know, so you never
know. I think Sam Alman has come out
afterwards and said he was shocked that
there wasn't more of a backlash or more
of like a media backlash because of it
>> or maybe the positive media went to open
models
>> maybe.
So I think and we'll talk a bit more
about this you know about open models
but I think I thought this was very
interesting. It's obviously
>> I feel like most people don't even know
what hugging face is. You know what I
mean? Like the layman
>> Yeah. Like if they if they like hacked
the New York public library,
she would be on fire right now.
>> Maybe. So yeah, then the average person
does not know or care about Hugging
Face, right?
>> We do, but most people don't.
Opus 5 came out.
Is it better than Fable? I don't know.
It was released on July 24th. This is
the post from Claude. It says,
"Introducing Claude Opus 5. It's a
thoughtful and proactive model that
comes close to the frontier intelligence
of Fable 5 at half the price."
And then, you know, there's some
benchmarks that came out around it. So,
exciting news. Claude Opus 5 with Max
Reasoning is number one in the frontend
code arena and text arena with
factuality on, which seems like very
specific that it, you know, you need,
but it it does beat Kimmy K3. It beats,
you know, Fable 5. So, it's apparently
good at like front end.
We saw that, you know, Claude Opus 5 by
Enthropic AI is second overall in design
arena with an ELO of 1358, which puts it
just, you know, I guess not just behind
Kimmy, but second place behind Kimmy.
Then you know Claude Opus 5 is narrowly
the most intelligent model on the
artificial analysis intelligence index
offering comparable intelligence to
Fable 5 at 26% lower cost per task.
So, you know, looks good on some
benchmarks.
Not it didn't look great on every
benchmark, right? Like there are some
that it lost to Fable, lost to Kimmy,
lost to, you know, 56 on, but there are
some benchmarks where it was either top
or very close to the top. And then
there, but a lot of people have mixed
opinions. So, Siki Chen says, "I take
back what I said about Opus 5. Initial
results were promising, but the more
time I spent with it, the more
infuriating of an experience it became.
My team feels the same way. I am back on
GBT 56 Soul. It's my daily driver with
Fable and Kimmy 3 unplanning and
reviews. Theo said, "I do not like Opus
5 as much as I hoped."
What do you think? What's been your
response?
>> Um, been daily driving it and then daily
driving it in the factory
and I just don't think it's not as smart
as Fable, but I think it is quite
capable. Um, so
I don't know. I don't have the same
feeling, but I'm just doing review right
now. So maybe that's the point.
>> I think it just doesn't feel much like
in the tasks that I've sent it, it feels
the intelligence level is pretty close
to like 48 for me. Like I don't notice a
big jump.
I have noticed there's, you know, and
maybe this is momentary issues with
enthropic or whatever, but I've noticed
that sometimes it just stalls out.
Sometimes that could be like the
response like too long of response, so
it just
>> cuts out. I I I don't know if that's a
me problem, but that's just something
I've noticed with Opus 5, I haven't
noticed necessarily with other models as
much. So, like some momentary things
where it just doesn't feel like it
finishes what it was what it started out
to.
>> Yeah.
>> Um but overall, I seems good. I don't
know that it seems necessarily great. I
don't know that it quite feels fable
level intelligence to me, but maybe a
step in the right direction, I guess,
overall. So I I don't hate it, but I
don't love it if that makes sense. It it
will probably be part of my rotation
though.
>> Same.
>> Um and then but one interesting thing
that's kind of come out. So Justin
Schroeder had this post and it says this
chart says so much. They use the exact
same prompt. They were all long horizon
oneshots and he said it reflects his
world real world experience at least,
but it's token use for this same prompt.
So it compares GBT 56, Terra, Luna,
Soul, Grock 45, DeepSync V4, Fable 5,
GLM, Kimmy, and then Opus 48 and Opus 5.
And on this task, which again, I don't
know, maybe this is like cherrypicked.
Hard to tell, but Opus 5 used a ton more
tokens. 97.
>> I think it's very uh it's very trigger
happy for tool calling.
>> Yeah. So 97 million tokens compared to
like Opus 48 was 23 million
>> and GBT 56 Soul was 4.3 million.
>> So if you think about it, so not only is
the token cost more expensive, but it
it's very token hungry as well.
>> Yeah.
>> So if if you're on your max plan, you
don't care. Who cares, right?
>> Yeah.
>> If you're paying API costs, you probably
care.
You definitely probably care.
Uh, anything else on Opus 5?
>> No.
>> Yeah, I think it's a good model. I don't
think it's,
you know, like the last time I I will
say this, going from like a four to a
five, you expect it to be this kind of
like put the like the plant a flag in
the ground kind of release. It doesn't
feel that way to me, but it feels like a
good useful model.
>> Yeah.
All right, let's talk about open weights
and all the things regarding
open models and should we have open
models? Should we not have open models?
So, Jensen had a post last week, first
post, I guess it was both Jensen and
Zuck both have had like first posts for
the first time in
>> in a you know, in potentially a long
long time. But Jensen said, "For my
first post, I'm sharing a letter Nvidia
signed on why open models matter. AI
will transform every industry, power
every company, and be built by every
country. Open models strengthen safety
and cyber security, accelerate
innovation and diffusion, and enable
sovereignty.
And then a bunch of people kind of
basically like signed on to this, right?
Signed on to this letter. You had even
open AAI signing. You had all the other
usual like suspects that would you
typically sign something like this also
sign it, right? Right. Palanteer of
course is going to sign it. YC signs it.
Whole bunch of like open model companies
of course signed it. Misilla signed
signed it. GitHub signed it. You know,
everyone you'd kind of expect.
>> We're trying to sign it.
>> Yeah. We we said we, you know, we threw
our hat in there. I don't think our logo
got on the board, but you know, we we
said we we would sign it. Um because I
think we all you if you're watching
this, you'd probably sign it, too,
right? I think we most of us agree that
open models are a net positive. It
keeps,
you know, it keeps things more open,
allows you to have more flexibility. No
one's going and hosting these open
models themselves. Not the big ones.
Like the smaller ones maybe, but the
bigger ones you can't host yourself,
right? But they should still be like the
ability for people to have open models
and open weight models is is a good
thing overall. I think I think it pushes
the frontier to be more competitive, to
keep moving faster, and I think it lock
keeps us from getting locked into
there's a few companies that control all
the intelligence, right?
And then this came out. I thought this
was hilarious.
Denny's had a banger post that says
[laughter] Denny's and Nvidia both know
the importance of staying open. So, you
know, not first time first time Denny's
mention on, you know, agents hour, but
nice work. That was funny. I laughed
and then Enthropic finally responded. I
feel like Enthropic must have been
getting a ton of internal pressure and
they did not sign it, right? No,
>> but they did
outline how like their thoughts and so
you can read this post, you know, they
they released it on the, you know, on
the anthropic blog or their news in in
the anthropic news announcement. It
basically says our position on open
weight models
um their biggest concern is more of a
risk of authoritarian governments, not
just the CCP.
They're, you know, concerned that
powerful AI models may be misused to
carry out cyber attacks.
Their biggest things are we should not
sell powerful chips to China. We should
crack down on industrialcale
distillation operations. You know,
that's Daario's thing lately. He doesn't
want people to, you know, pay for their
inference and take take the content and
build models from it.
Um and then the next big point is all
sufficiently capable models open and
closed should go through mandatory
safety testing.
And so overall tried to be like take a
more reasonable approach. They didn't
respond to every point in the letter but
said we don't dislike open models but
here are the things we believe and we
think that if even if they are open
models they should have to go through
this some rigorous testing which I guess
is mandated by the each government which
kind of makes things hard though because
you got to then be tested by every
government entity that
would regulate the models and the model
use within their country which becomes I
think hard to govern
and then I and then my question would be
like who gets to decide what that safety
test is because I bet you anthropic says
it should be them.
>> Yeah.
>> And that that's the concern of course
>> and then you can deem things that you do
not like with bias.
>> Yeah. They are fully biased you know
>> right? So if OpenAI, Anthropic, and
maybe Google and XAI are the only
companies that can determine what this
what safety is, they get to write the
safety test. Well, then they can pretty
much just write the test. So open models
are probably not going to pass it,
right?
>> Yeah.
>> And the other argument is if you have to
go through a rigorous test, then it does
block out anyone else from being able to
release new models because they got to
go through this rigorous test which are
probably going to be very expensive,
very time consuming.
I see.
>> Then, you know, then you don't even want
to innovate anymore because it's like
the the red tape to even start. You're
like, you know what? I'll just [ __ ] it.
I don't even want to do this anymore.
>> Yeah. And I think, you know, you see
that with a lot of government
regulation. When an industry becomes
overregulated,
typically innovation slows down, right?
It's it's like the path
>> and corruption goes up.
>> Yeah.
>> How many like side deals would happen?
People selling bribes and all that
stuff.
>> Yeah. I I mean on the flip side there is
an argument for safety testing right in
that yeah
>> do you not you know you don't want the
most powerful agent to do everything but
>> I also would argue maybe the best way is
just
>> if all the intelligence is open then at
least you can have the right tools to
protect yourself if there is an agent
because who knows
>> what kind of agents being you know
cooked behind the scenes that could do
all the damage and you don't have access
to it right so how can you protect
yourself from it. So I can see both
sides, but ultimately I think less
regulation is typically better and we
shouldn't have we should be encouraging
innovation at this point rather than
trying to you know encourage or like
discourage people from even trying.
>> Knock on wood for the Skynet stuff. But
yeah.
>> Yeah. [laughter] Yeah. I mean that's the
that's the asterisk, right? Like you
don't
>> I don't want Skynet, but I also don't
want, you know, only three or four
companies to control everything. I don't
want Daario controlling this.
>> Yeah. And
now, you know, OpenAI and I think even
Anthropic, maybe some Anthropic
employees, they started this
uh they it's called the pacing the
frontier. So, pacing the frontfront.com.
>> So, they basically made a statement.
They had a whole bunch of people that
signed from different uh companies. And
so and OpenAI both signed the open model
letter, but then they kind of go
backwards a little bit and they're
saying like we should be very careful
about frontier intelligence and we
should kind of have this regulation and
and you know like safety concern over
top of it, right? And so maybe I can
share
um this is kind of the website the
letter statement from 1,200 employees of
Frontier AI companies. You can see, you
know, Daario's in here, chief scientist
of Open AI, chief scientist of thinking
machines, anthrop, you know, co-founder,
chief science officer of anthropic,
chief scientist of Meta, Google
DeepMind.
Um,
and kind of the statement is we request
that the US government support an
international effort to develop the
technical and governance tools needed to
deliberately pace the frontier of
automated AI development.
So my question is what happens if
someone says they're going to do they're
going to pace but then behind the scenes
they don't. What if China is like, "Yes,
we're in." But then they're actually
like, "You know what?
>> We're gonna be doing our own like black
ops behind the scenes trying to like
we're gonna try to slow everyone else
down. And don't I mean, Open AI and
Anthropic are going to do the same
thing, right? Like they might not
release it to the public, but they're
going to be doing it behind the scenes
because they want to be ready and have
the everyone wants to have the most
intelligent model.
>> So, it's Game of Thrones, dude.
>> I am very skept. very skeptical of of
this in general. It's it's very tied to
the the last, you know, the open weights
concept.
>> I wonder if I sign it, will they accept?
I'm not in a Frontier lab, but you know.
>> Yeah. I don't think I don't think so. I
don't think
>> I'll be at anthropic. [laughter]
>> Yeah. I mean, and you can see they have
some quotes here.
Um, so you can kind of see the thought
process.
I think all this highlights is there's a
lot that's going to come from government
regulation wise uh open weight open
model wise around just how open models
or how models in general are developed
and how intelligence is is kind of
rolled out over the course of the next
few years and so some level like we need
some things I don't know what that thing
is
>> I wonder if you got fired if you didn't
sign it if you worked at anthropic
I would hope not. I bet. But I feel like
Enthropic doesn't need to. I feel like
if you're at Enthropic, like 75% of the
people believe the same things. Like I'm
not saying that there aren't divergent
opinions, but I think like from what
I've heard, Anthropic kind of has the
mission and they're pretty public about
their mission that, you know, like I'm
going to like we are we are the company
that's going to make AI safe, right? And
if so, if you believe that and you work
at Enthropic, you're gonna sign this
thing.
>> Yeah. Now, my opinion is I don't think
one company is what's going to make AI
safe, but that's where my opinions
differ.
All right. Um, continuing on, and then
this came out also July 28th, which is
yesterday. It says, "President Trump is
relying on a small group to decide what
restrictions to impose on Chinese AI
ahead of this week's open AI meetings.
It includes Howard Lutnik, Scott
Bessant, David Sax, Susie Wild, Sean
Karen, Cross, Arvin Dramman.
>> Oh boy. Wonder what's going to happen
then.
>> Again, more to come. More speculation.
We will see.
All right, let's talk about MCP. MCP is
not dead. It's just stateless now.
>> Yeah. So
>> V2
>> MCP 2026 0728 is live and it's the
largest update to the protocol since the
launch. This is from July 28th. This is
a post from claude devs and it says MCP
is now stateless making it easier to
deploy and scale remote servers. So tell
me about this Obby. What does this mean
for folks?
>> So MCP is an API now. Um that's cool.
Um, just to give a little history
lesson, so MCP came out quite a while
ago. Um, and when it first came out, it
was only through stdio standard out. Um,
and how MCP used to work was you have a
connection. You like get a connection to
the server and then the protocol to
transport was standard out. This was
really good for MCPs that were not
hosted, let's say, but uh or some were
hosted, whatever. And that was cool to
start, but automatically a lot of people
were wondering what the hell MCP is
useful for because like why do I need a
connection to a server to to do this
stuff? Then we had SHTTP, which is a
state stateless HTTP protocol in M in
MCP, which allowed [snorts] you to do
HTTP
And now we're back to everything's
stateless just like a rest API. So
>> So can you use full circle?
>> Can you use like stdo like standard
input out anymore? Now it's gone in this
new version.
>> It's all Yeah, it's all like we've been
doing for many years. We are back at
square one.
>> Um yeah. So it's like MCP
realized that most people use MCP for
tool calls
>> tools
>> and how do you norm how would you
normally access a remote system
>> through an API
>> API
>> and you don't need a connection you
don't need a long live connection to
that system you just want to like make a
request get a response and have your
agent
>> that's it
>> handle that thing so why do I need a you
know a connect an ongoing connection
Yeah. And this al this honestly
complicated agent development because
sometimes you lose your connection and
you're just trying to [ __ ] make a
tool call and then you have to make sure
that you have a connection, you lost the
connection. Um you have to regain it.
It's just like all this latency.
Um but some good things came out of this
uh V2. Uh they got rid of dumb [ __ ] that
no one used. Roots, who cares? Like
logs. Yeah, just use regular logs. Like
who cares about that? Like they had
added all this crust to MCP
and it's just gone which is great.
>> I'm assuming they still have like O. Do
they still have elicitation?
>> Um so they have O still which is just
going to be normal ass O. Um which is
great. They have tasks. They had a task
protocol that still exists but roots
sampling logging are all deprecated.
They'll still work for the interim.
Elicitation still works. Um,
and elicitation is I mean people use
that so like that was good but um you
didn't have Yeah. But still even the
people like elicitation is used but not
at the same like most people are using
it for tools right like 8 I would say 80
plus percent of people that use MCP it's
literally just to share tools so it's
easier for agents to use right like that
is most people's use case
>> and there was a lot of talk around like
is MCP dead because you know hadn't been
updated for a while or hadn't really
been like at least not very vocal
updates a lot of people weren't using
all the new things that were added. I
would say based on my experience and
conversations, MCP is definitely not
dead, but I do think MCP is going to be
like an enterprise type like
where that's where it's going to get the
most use. I'm not saying it's not going
to be used outside of that, but
a lot of things that people are using
MCP for, they're just using skills for
now. Unless you're an enterprise and you
want to build a set of like a tool set
that you can share across teams. That's
where I see the most is like internal
tool sets that one team can build the
MCP server, connect it to the different
systems and give agents or you know that
are being built by another team access.
>> Yeah, a lot of things happened to MCP
that were detrimental like outside of
MCP, right? One, you could because you
can write code easier, you can just
create tools with your coding agent
using SDKs that you already have.
>> Yep.
>> Cool. Second thing is CLIs became cool
again. In general, coding agents will
just execute the CLI. Most things have a
CLI and if they didn't, people started
building CLIs for them, right? And then
finally, the whole noise about, oh, you
can't just use OpenAI specs because they
weren't written for agents. Well, good
[ __ ] luck because now everyone's
writing APIs for agents. So, open AI is
cool again. Sorry, open API. My bad.
Open API specs can be good because
they're being refactored for an agent
world, right?
>> Yeah. And
>> why even use MCP? And I think in a lot
of cases it's like MCP is good for
sharing and if you want people to just
be able to easily plug into it and you
know but at the end of the day if you
just had a REST API or you could spin up
an SDK you could probably get around a
lot of the same things.
>> Yeah.
>> And I think one of the other challenges
with MCP is you get this like huge list
of tools and you don't always want like
the GitHub MCP back in the day. You get
like a hundred tools. I don't want a
hundred tools. I want like 10 tools.
>> Maybe 15. So maybe it'd be actually
better rather than have my agent have to
decide between 100 tools or me having to
like look through the list, I could just
have my agent know that these are the
five things I needed to do, do that. And
maybe sometimes one tool call is
actually like two API calls, right? Like
not always, but
>> like if I want to request a refund, that
might be a couple API calls to do a
refund, but my agent just needs to do
the refund. They don't need to like make
three tool calls. So I I think in some
cases MCP is good. It's still used. I
think it's become where it was extremely
hot as like a concept 18 months ago,
right? About that, you know, 15 months
ago. Now it's just become like a a tool
in your tool belt, right? Like there are
certain use cases where it's good. If
you're sharing tools around teams, like
maybe it's still useful. Deploy an MCP
server. One team can maintain it,
another team can use it.
But I feel like in a lot of cases it's
going to be the same as just using an
API.
>> I think one benefit of MCP
today is you can expose an MCP server
that wraps your SDK or your REST API or
REST client or whatever whatever
internal logic you have and now you
already have a tool format that the
agent will speak. So you don't have to
do this like glue code, right? Like the
MCP is already giving you tools. Now you
just have no connection or handshake
necessary. So there are benefits, but
we're just haters.
>> Yeah, I think it's just Yeah, the
benefits are not as great as they were
when it came out. So still useful and I
see it all the time in uh in enterprise
settings. It's being used heavily. So
it's not it's definitely not dead, but
you know, it's maybe just not not what
everyone thought it was going to be.
>> Dude, I'm so glad we didn't support
roots in our MCP client. Thank
>> we talked about it. Yeah, we got we had
people asking about it.
>> Yeah, they ain't asking anymore.
>> And then, you know, this came out sounds
like we arrived at REST API, which is
just funny as a response to uh
stateless MCP because yeah, it's very
similar. I mean, obviously it's, you
know, it gives you a tool format that
agents can use as you said, but yeah,
it's an API.
All right, let's go through some quick
hits. This is where we rapid fire
through a bunch of things that might be
interesting to you all and we'll give
you some of our hot takes on it. So,
Kimmy K3 weights are out. So, that's
good. If you're a fan of open models,
which we are, they released the model
weights. So, you can kind of read
through that.
There was some a leak, I guess, on July
26th. You know, whether it's a leak or
not, I don't know. I think more, you
know, I feel like Model Labs released
these things so they can kind of build
hype. But anyways, Kimmy K 3.1 leak says
performance that closes the gap with
GPT56 and Fable. It's uh maintenance
report mythic mythos level capabilities,
faster inference and lower latency,
better token efficiency. But just
getting a lot of hype around when this
chem 3.1 coming out. It's supposed to be
good. It's supposed to be like even like
as much as a surprise as Kimmy K3 is,
imagine now you get an upgrade to that
if you can make it a little faster, too.
SSI
announced, so this is from SSI Inc. It
was a message on July 27th said, "We are
announcing a long-term strategic
partnership with Nvidia. NVIDIA is
making a substantial investment in SSI
that will let us 10x our compute in the
next 12 months. We reached the point
where our research is worth scaling and
with this partnership, we will be able
to. So this is Ilia's from, you know,
OpenAI days, Ilia's company, SSI, and
sounds like they, you know, maybe
starting to make some moves.
I I think I I don't think you can be a
Frontier model lab and not build
partnerships like this. So
>> yeah,
>> maybe that means we'll be seeing some
things from SSI.
Stripe is in discussions to acquire Open
Router possibly for as high as $10
billion.
>> That is
tight.
>> Yeah, we're fans of Open Router. We like
Open Router.
>> Friends with them.
>> Yeah. So, that's cool. If true.
>> Another dude in Sam's fraternity is
about to be rich. [laughter]
>> Yeah. You know, it's one of those things
like big if true like you know who we'll
see. But dang that that's a that's a lot
of
>> people have a model router
>> just like ramp.
>> Yeah, exactly. I mean
Stripe, you know, Stripe is it's funny
like Stripe is a financial company,
right? Like tied around like finances,
card processing, all that. Ramp is a
financial company. And now they're both
like kind of pivoting to trying to be AI
companies in a lot of ways.
>> Yeah. I think they like imagine being
like leaders in those categories and be
like you know what's a bigger market
than finance AI. Let's go there.
>> Notion as code. So this is now in beta.
This is a post last week July 23rd from
notion. It says you can define an entire
workspace in Typescript team spaces
databases custom agents all of it and
then deploy it through the API. So you
can build workspaces with coding agents,
version control your setup in Git, and
reproduce the same setup anywhere you
need it.
I think this is kind of actually cool.
Like, you know,
>> I'm not going to use it, but it's cool.
>> Yeah. I don't think I mean I I feel like
Notion has become less important for us
as a company. Like we are huge Notion
users. I feel like we still use it,
>> but it's basically used as just like a
wiki, right? It's like if something
doesn't live in linear then put it in
notion maybe. But I think if you were to
want to build like knowledge bases and
you could you you know wanted to be able
to just have your coding agents spin up
and do things for you. But then I my
question is do you really need notion to
do that or not?
>> Yeah.
>> And there's been a lot of hype around
like what is it like open wiki or
something like that's been coming out as
well. So I think there's a lot of people
like trying to disrupt notion. Notion's
trying to become, you know, an AI
company as well. And so they're trying
to get closer to coding agents.
Everything is converging on like the
making things for coding agents.
>> Yep.
>> Cognition is in acquiring interaction,
the makers of Poke. So if you've ever
used Poke, you know why we're so
excited.
>> Okay,
>> that'll be cool. Another channel for
them.
>> Yeah. I've never used Poke. Have you
used Poke?
>> No.
>> Anyone in the chat, have you ever used
Poke?
>> I don't know. Never used it.
>> I think it's I mean Poke was like an AI
assistant that would be in your
WhatsApp, your Telegram, uh your
iMessage, things like that. And it was
like a, you know, like an assistant or,
you know, some somebody you could talk
to as well. Um, so I I I assume this is
to expand channels and that technology
with Devon.
All right. Chat GBT voice is now in the
desktop app. So control your computer,
direct multiple agents running in chat
GPT work or codecs just using your
voice. It's powered by GPT live. So it
can speak, listen, and coordinate work
in the app at the same time.
This is cool. I did see a post about
this that I thought was kind of
interesting and I actually I'm gonna try
it just to maybe provide a come back and
and talk about it. But it's basically
saying if you run the desktop app then
you can actually like go on your phone
and talk to it, but it can control your
your computer. So you can basically be
like, you know, controlling your
computer while you're going for a walk,
right? You could be telling it what to
do. it'll be, you know, basically live
voice and it'll be kind of making moves
for you as you're talking to it using
your computer, but you can kind of take
it anywhere. So, it's this idea of like
maybe you just have one,
you know, workstation running all the
time, but you have you just bring you
can basically
>> it would pass the bar test, I guess, is
the idea. And so maybe maybe I need to
try it out because ultimately you can
use chat GBT work or codeex right from
you know technically the mobile app. You
just talk to it with voice which is
pretty cool.
>> Yeah.
>> So I think I think we'll be seeing you
know that that's the dream that people
are trying to build for it for. Obby and
I have been talking about this dream for
18 months now it seems like or a year on
this show. It's like, you know, how do
you pass how do you get it to pass the
bar test or the beach test where you're
you can take your work with you on the
beach or at the bar and you can still
make some moves.
All right, we got to watch this video
because
>> yeah, this is dope.
>> This is wild. So, give me a second to
pull it up because yeah, we got to watch
this video.
So, this is from Door Dash. We're
cleared for takeoff. Say hello to Door
Dash Air, our in-house drone delivery
program. You know, this isn't maybe
specifically AI, but it's kind of AI
related, right? Um, so let me pull up
this video, wherever it is. There it is.
Hopefully you can all hear this.
Heat. Heat. Heat.
All
>> [music]
[music]
>> right. So, if you watched the Yeah. the
episode, I think it was was it last week
we were talking about is Dor are you is
your agent going to be ordering you
pizza? Yeah,
>> like your agent's going to be ordering a
pizza delivered by a damn helicopter,
>> dude. [laughter]
Um I have a friend who is working on
Door Dash drones um in uh SF. So, dude,
it's tight. Also, when he first told me
about it, I was like, that's like the
dumbest thing ever. And then
now that I think about it, it's not that
dumb. So, the test is, and we should
record this next time we're in SF
together.
We get our agent to use Door Dash's MCP
>> and get it delivered by Door Dash Air.
That would be sick.
>> That's the dream, dude. At the bar.
>> Yeah. [laughter] I I need my pizza. The
pizza comes down and drops.
All right. Yeah. So, that's that. Uh
before we close out, you know, thanks
for watching the show. Follow us on X.
Follow Mr. on X at Mastra. Go to our
YouTube. Hit subscribe if you haven't
already. Please follow me on X at SMT
Thomas 3. Follow Abby on X. Um we
appreciate that. We appreciate any
fivestar reviews. If you don't want to
give us a five star, find something else
to do. But if you do like the show, that
fivestar review does help other people
like you. The other thing you can do and
every time I say this people are always
ask me what are friends but if you do
have friends that are like you and think
like you tell them about the show
because that really helps us get more
people. We've had a ton of chatter in
the the chat that we kind of haven't
pulled up. So let's go through some of
that. So we got I am Brennan says hey
guys does factory run cloud code behind
the scenes right now it runs master
code. We will, you know, Mashra
supports, you know, ACP, we support, you
know, cloud code, codecs, things like
that. So maybe eventually it's like you
can configure your own coding agent to
run if you have a preference, but right
now it runs master code. We'll probably
make it more extensible in the future.
>> That's a big maybe, but we'll see.
>> Maybe we'll see. We will see. We got,
you know, our we think master code's
better for a lot of reasons, but maybe
we will uh make it work. uh when we were
talking about you know
all the open model stuff all yeah
says dystopian behavior
and I think it was when we're talking
about the the books anthropic you know
destroying the books reminds me of the
Google plus French national library
story
>> ma 3D says Kimmy K3 plus opus 5 is a
pretty good combo that's
Medigame says is factory for like
building an agent set up that connects
to things like GitHub. So not exactly.
So what factory is npm create factory if
you want to try it out. It essentially
allows you to help automate your
software development for your team. So
the reason we wanted this is because if
you think about a a small team,
everyone's running their own coding
agents, right? Chipping their own PRs.
But what if you had a centralized place
where your team could see the work, see
all the coding agents that are running,
interact with the coding agents, and
essentially like collaboratively ship
software, but in an often automated way.
Not everything needs to be completely
automated. But that's kind of the dream
is like what if an issue comes in and
the factory just picks it up and works
on it. And at the end, you get an PR
that's gone through multiple rounds of
approval. So ideally, you just click,
you know, merge. Maybe certain things
you still want to review, but maybe
there's certain types of tasks you just
go ahead and just merge.
Um, Hassan says, "Maybe OpenAI only
signed the letter because they thought
Jensen was talking about them when they
said open." [laughter]
Um,
Hassan says, "How unlikable do you want
to make yourself to the public?"
Anthropic says, "Yes."
Um, Mika says, "This is only going to
lead to the best models being hidden
from the public." I agree.
Profi Woo says, "It's definitely useful
for enterprise internal tools." Speaking
of MCP, agreed. That's where I see all
the use or a lot of the use cases.
>> Um, I al so I don't know how to
pronounce your name, but said, "I often
find tools to work much worse once
abstracted behind an MCPA tools.
interesting.
Um,
says LMVD Xand says WTF is poke.
Hassan says never heard of it before.
So, I'm at least I'm not the only one.
Thank you for the chat for backing me up
that I I didn't know what poke was, but
yeah, many games never heard of it. Um,
so anyways,
all right. And then Yan says, "We need
to make a machine that eats the pizza to
close the loop."
>> Close the loop.
>> Man, there's a lot of chat today these
days. I don't know what software is
anymore. Everyone's just building the
same desktop at with the chatbot.
>> I feel you. Uh, Prof. NGW says, "Factory
looks amazing.
Before TSAI, I thought one could use
Masera to offer their services to
software companies to create a factory
for them. After TSI and the factory
release, that thought is obsolete. Great
work."
>> It's not necessarily It's not
necessarily obsolete though. Like
factories customizable. So take it and
go customize it for people and help them
build their own factories. Like that's
that's the goal. But yes, we want to
give you tools so you can do it easier.
So you don't have to do it all yourself.
All right. Dang, that was that was a fun
show.
>> Super fun.
>> Wait, this just in. Anthropic is down.
>> This this is my shocked face. All right.
Just another Wednesday.
>> Just another Wednesday.
>> All right. Well, thank you everybody for
tuning in. Thanks for uh watching.
Thanks Fennel for watching at 2 am in
India. We appreciate you. Appreciate
everyone for watching the show. Go
ahead, follow us, like, do all the
things. Uh, and we'll see you next week.
We'll do it again. Be on Monday next
week, so back to normal.
>> Yeah. Peace.
>> See y'all.
>> Still here. And
>> we're still here.
>> Still here.
>> Yeah. This is where normally Yan comes
in with the
>> Yeah, Jan comes in with a outro, but you
know, it wouldn't be a live show without
a few technical difficulties.
>> Yo, that show's a wrap. We were live in
the zone agent with Shane and I be on
the throne. Did you give us that review
[music] only if it's a five? Jump on the
tube. Make sure to like and subscribe.
Dude, so fresh. [music] Yeah, we keep
you in the loop. Get so fly. They bring
the whole troop. AI on the rise. [music]
Don't miss this [singing] power. Welcome
to the show. It's AI Sour. Did you just
drop in? Is this your first time? Make
sure to follow us on next and go like
and subscribe. Yeah. Learn the
principles and patterns in our books.
The master.AI site. Give it a look. New
so fresh. [music] Yeah, we keep you in
the loop. Guess so fly. They bring the
whole troop. AI on the rise. Don't miss
this power. Welcome to the show. It's AI
Sour. This is the end. We all wrapped
up. Another showdown. Another one coming
up. AI agent I was done, but the news
doesn't cease. Shane and Abby, we out of
here. Peace.
[music]
Continue with YouTLDR
Analyze another video with Pro
Process a new video, search every timestamp, compare sources, and keep the result in your library.
More transcripts
Explore other videos transcribed with YouTLDR.

Leilão de Embriões Nelore PO DNA Genética Aditiva
LANCE RURAL OFICIAL · Portuguese (Portugal, Brazil)

Leilão Peso Pesado Rima Agropecuária
LANCE RURAL OFICIAL · Portuguese (Portugal, Brazil)

النبي .. جبران خليل جبران .. إقرا بودانك
اقرا بودانك · Arabic

Leilão Internacional CIA
LANCE RURAL OFICIAL · Portuguese (Portugal, Brazil)

23° Mega Leilão Genética Aditiva - 1ª Etapa Fêmeas Nelore PO
LANCE RURAL OFICIAL · Portuguese (Portugal, Brazil)

كتاب رسالة الغفران
كتابي المنقذ · Arabic

Erkenntnistheorie 7 Immanuel Kant II
Dominik Finkelde - Hochschule f. Philosophie · English

Schwarze Löcher Erklärt - Von der Geburt bis zum Tod
Dinge Erklärt – Kurzgesagt · German

Como construir uma esfera de Dyson – A Megaestrutura Suprema
Em Poucas Palavras – Kurzgesagt · Portuguese (Portugal, Brazil)

[Histoire des sciences] L’histoire de l’intelligence artificielle (IA)
CEA · French

ميكانيكا الكم│1│الواقع الوهمى - كيف بدأ الكم ؟!
Sharafestien - شرفشتــاين (Sharafestien) · Arabic

Mi niñez fue un fusil AK-47
Comisión de la Verdad · Spanish