How to sell RL envs and data to AI labs: Interview with Sean Cai
What areas of data do you feel are
underserved by now um that labs want to
buy?
>> Uh certainly we're still in a huge darth
of long and unverifiable. Uh the domains
that matured the quickest
matured because they were much much more
easily verifiable with web 2.0
instruments. Coding had GitHub. We don't
have a GitHub for all other domains but
we ventured into finance healthcare and
law afterwards. Nowadays, biological and
cyber security is is is a craze but only
I think by the most sophisticated labs
namely anthropic and some of OEI.
>> You said bio and cyber.
>> Yes.
>> But in general I think incredibly long
horizon realistic is still very much in
demand. A lot of benchmarks are built as
this but not actually that. Does that
mean things that would like basically
serve making co-work and codeex better
for non-technical work?
>> Uh potentially. Yeah. I would classify
that as like back office ERP type tasks.
>> Yeah.
>> Which largely have to do with very
complicated search and retrieval
functions across quite convoluted data
links uh and environments.
>> Okay. Um what kinds of software programs
would people be using?
>> Uh even so Excel file systems.
um just like think very convoluted file
systems and applications with data
across multiple formats sometimes
tabular sometimes graphical even as well
>> um and it's an exercise in model tool
calling as well
>> cool also stuff like SAP or or not
[snorts] necessarily
>> uh yes but I think for maybe some of the
computer use circles computer use
continues to be sort of like a smaller
market but dominated by a few top RLM
companies there uh like a a version of
SAP would have gone for like 500K uh in
the computer use craze of like late 25
probably. Um I suspect a version of SAP
has been created by one of the RL
computer use companies at this point or
maybe by an internal anthropic and OI
but that's speculation. Yeah. The um bio
and cyber for the bio stuff is it uh bio
real world things or like purely digital
you know stuff that's pure computer?
>> Yeah it started from bioinformatics but
nowadays we're trying to make a lot of
these processes with one step in the
digital and one step in the physical
areable.
Uh naturally this gets really hard
because you think about the verification
mechanism for a lot of things in
biology. Anthropic just put out this
benchmark called mystery biobench which
I think enumerates a lot of the problems
pretty succinctly there. We don't even
know even amongst the top experts how to
verify something in biology because
they're literally denovo experiments,
right? Like we combine certain chemicals
like what happens? Um
so uh they're almost veering into
physics based models. Of course if you
get to physics based models then you
have like
uh semreal and robotics uh to to which
you could do RL and robotics and that's
an entirely different domain. But I I'd
say it's like on our slow march to make
sim to real uh and and generally model
much more things in the physical world
more accurately as verifiers and RL um
>> certain domains
>> fall in the middle of purely software
based work and like robotics
>> like biological workflows, chemistry
workflows, even scientific discovery
that make that pretty useful. What are
some of the pieces of software that uh
bioinformaticists might be using that
like they would have you know ends for?
>> Oh, I think I think like I'm not a
biologist but uh there's so many bespoke
tools and like bespoke processes within
a lab itself. I I I'm helping a wet lab
sort of digitize our processes here
doing this and a lot of this stuff
doesn't have purpose-built software for
it. It's
and it does just adds to the environment
complexity and bespoke tool build out
for those environments complexities as
well.
>> Okay. So it's more like general computer
use both both guey and bash.
>> Yeah, I would I I would say so that's
the primitive in which like kind of
everything is based right.
>> Yeah. the um
cyber stuff any particular subsets like
uh uh you know program analysis or um
pentesting or um you know what's the
what's the most underserved cyber
subsets do you think?
Yeah, for sure. Uh, I think cyber is
mostly being bought by Anthropic and
maybe some of OEI right now and then the
rest of the labs are are following
whatever they do. Uh, naturally
irregular security is probably the
company to fall into space who does a
lot of this type of stuff. Um,
I would say a lot of offensive cyber is
uh being modeled right now in sort of
interactive environments. So a lot of
stuff in code level security uh web app
exploitation CTF challenges and like
remediation is already covered by
existing benchmarks like um cyber gym
cybench these are pretty saturated
nowadays so long horizon wise as all
domains are tending towards you're
looking towards um
stuff like infrastructure exploits uh
and agent layer attacks. So there
there's a lot of stuff you can model out
there. Um because there are new zero
days every single day. It's like one of
the most dynamically changing fields. So
naturally you're going to expect there
to need to be uh real-time data streams
to translate these things into model
actionable formats.
>> Yep. Awesome. The um
the is the the sales process is roughly
um either do something in public such
that researchers reach out to you
already know researchers or kind of get
intros or do cold emails to researchers
to get a you know a pilot. Is that
roughly the first step?
Yeah, but you know the our our appetite
for data is voracious and expanding, but
there's still only 24 hours in a day and
a researcher's job is not to talk to
data vendors all the time. So the bar to
get I think a researcher's attention is
getting higher and higher. So those
without research sophistication, it's
just like that's simply not the best
move anymore. the human data supply
chain is expanding such that one can
meaningfully participate in it in it um
without interacting with an end
researcher
>> cross-selling partnerships you mean or
>> yeah protege is an example uh I think
there are companies out there which
almost train companies to produce good
post training data and then sell their
data to researchers themselves using
themselves as a stamp of approval
Yep. Yep. The um
the uh what are what are the most
important things that labs look for?
What are the main reasons that a lab
might uh not kind of renew or increase
their purchase volume? You know, after
they get the uh the first set of data.
>> Uh each lab has their own QC processes
and they run them internally and they
see. But I I you know there there are so
many reasons why
some data might be quite quite poor. Um
so
I think there are just a lot of very
small things that a researcher can look
at a data set and like maybe a data set
is a million lines and they notice
something that is off about one line and
they really start to question whether a
startup even has a QC process at all in
the first place or not. Um,
subsequently,
uh, it's it's it's I think the most
common case is just your tasks are
incredibly incredibly poorly designed,
um, in terms of they're easily reward
hackable. The props are vague, they're
not emblematic of real world tasks from
an obvious setting.
There's a lot of other smaller things
afterwards too like the model failures
you identify and the reasons for why the
model fails at them are not actually
genuine capability failures. So because
you designed the harness quite poorly.
Um it's because you ran it in a very
specific environment that isn't actually
how most users would have ran this task
or model in. Uh you don't do cross
harness testing. You don't do engram
contamination testing. So which is to
say you don't test whether a data set is
already in the pre-trained corpus
literature or not. Um there there's a
lot of uh QC checks that one can run
before delivering say OTS RL data that
would make it just uh that would make
DOM a great partner that researchers
would want to work with and iterating
the shape of posting data.
>> Yeah. the um
the and the QC is both uh researcher
flagged issues and then also just the
data teams as well reviewing things
themselves.
Yeah, I think it's uh
c certainly researcher needs are bespoke
as well for certain projects, but if you
think about
uh researchers exploring a net new
research question like how do we improve
taste and models and they're exploring
different OTS RL data sets to uh along
with maybe a data company provided
benchmark to exploit this question.
There there are just so many things that
they look for in a typical data delivery
that are really is really difficult for
you to know unless you worked in the in
the data industry yourself or you've
been a researcher and you understand
what good RL data is, right?
>> Good data taste, good research taste,
and perceived ability to scale quality
with quantity are the three things that
are necessary to building a a good human
data company out.
>> Good data taste, good research taste,
and what was the third one?
perceived ability to scale quality with
quantity.
>> Yeah. Cool. The um uh when does a
researcher like in a question like that
they're doing some new kind of area um
that's vague. How do they get the model
to do a certain thing?
Um
what do researchers do to first explore
off-the-shelf offerings before um will
they just hit up all their existing
vendors from their team or they go to
their you know data team people
internally and ask them to go and get a
a uh a set of options for them? What's
the what's the first steps that the
researcher takes?
Yeah, I'd say they do some bespoke reach
out themselves, especially for new
research team directions like OpenAI's
newest robotics VA direction which spun
out of Sora or not I wouldn't say Denovo
spun out of Sora but was sort of
combined with the remnants of Sora. they
will go out to real world data vendors
and reach out to get samples uh because
you got to remember their jobs are to
improve model capabilities and if that's
the bottleneck they won't go and solve
the bottleneck themselves but that's why
these labs have human data teams. It's
like one to procure the data necessary
and manage vendor relations but two is
like negotiate on price
>> and and all that all those things
associated. So it's a collaboration
between those two entities.
>> Cool. The um in the robotics data space
is there anything that your viewers
underserved
uh there? Obviously there's a lot of
things that are well served but what are
the ones you view as underserved in
robotics?
>> Data vendors who are genuinely research
first and running post training
experiments on their own data. if
they're trying to sell things like ego
and data up the training mix pyramid to
if you like companies
uh or
uh ju just running a lot of training
experiments on their data that match the
research direction of companies they're
trying to sell to in order to be a bit
more
>> selling egocentric data to VA companies
is underserved.
It is underserved in a sense of there
are not many
>> my naive view is like everyone is
selling egocentric data.
>> Well, there are very few egocentric data
vendors that actually know how the
downstream training is done.
>> So it's egocentric data vendors who are
doing their own postraining and
therefore like research have good
research taste.
>> Yeah. But um you you got to remember
what is being sold when you sell data in
the first place is just model capability
improvement, right? And data is just the
medium to do that. So if you're trying
to sell data, but you're not actually
cognizant of how model improvement is
achieved or you don't have an opinion
there and can't really help the
researcher with that, you are in a
losing battle and losing market uh and
you are going to be commoditized.
>> Yep.
>> Yep. The um
what is the uh what does the initial
meeting look like? the um someone talks
to researcher, the researcher maybe
requests some samples and the founder
sends it in Google Drive. Um how does
that uh what does that typically look
like? What's the formats people are
expecting?
>> For RL data, for the longest time, it's
literally just been a Docker container.
>> Yeah.
>> Uh a Docker container isolated
environment, all the tools on there, all
the verification mechanisms and rubrics.
One simply simply has to plug and play
their agent. Uh and then you get an eval
score and then you can use a multitude
of these software containers to run
rollouts for GRPO RL.
>> Yep.
>> Whatever other training me mechanisms
you employ
>> and labs do they have more kind of
sophisticated internal kind of you know
uh
setups for um running environments now
that need different formats.
>> Yes, they do. Anthropic notably has one
whose name I can't uh disclose but the
most sophisticated labs I would say are
like Anthropic, Open AI, Deep Mind and
then everybody else and then Chinese
labs in that order.
>> Yep. Yep. the um
the um and then in terms of the data
that the non
you know three frontier labs are buying
um are they buying more off-the-shelf
data that you know companies have
already sold to anthropic and open AI on
like a non-exclusive thing
>> uh I think OTS is a relatively new
phenomenon it is the mechanism with
which Serge has done business for a long
long time. Uh but that's because Serge
is a very fundamentally different
company than all the other ventureback
players. Um I I believe
>> how are by the way on that?
>> Oh, they genuinely started off as just
model capability caring about model
capability improvement, right? Not not
as a sort of data company and for the
longest time like mostly SFD data as
well. Um,
>> so
you're talking about exclusivity.
Certainly exclusivity reflects different
labs philosophies towards data vendors.
Anthropic is the only one who I think
really pushes for exclusivity. Open AI
at different points throughout its human
data turnover uh human data teams
turnovers because a lot of people shift
around in OAI a lot. Uh but Enthropic
genuinely views their data vendors as
research partners and if you think your
research partner is genuinely novel
research you probably want to get
exclusivity on that. Um
>> which is the approach that they've
employed with many of the data companies
they've worked with.
>> Yep. Do they have expiry clauses on the
exclusivity like 12 months or 24 or
something like this? I'd imagine they
are starting to think about that pretty
closely now. But I am aware of many many
companies who have recently just ended
anthropic exclusivity. Some of them had
the agreement that they would only have
it for a year. Some of them
for other strategic reasons they've
stopped exclusivity with. So
>> yeah.
>> Yep. the um
the
uh
almost all purchase decisions researcher
led at this point as opposed to you know
like the researcher pulls it in and then
the the data team kind of uh is effect
is like a form of pro procurement or um
is it different? You can imagine it's a
partnership of sorts,
but if you want to think about it from
from a economic buyer perspective, you
always want to be just in general B2B
sales dealing with the economic buyer
because if you can convince the economic
sorry, not the economic buyer, the the
end user, right? If you can convince the
end user of your product that there's
substantial value, the question is not
whether the org is going to buy it or
not. It's just how much are they going
to buy it for. Yep.
>> Um, and so if you're going to the guy
who's pricing it first, who doesn't know
how available it is, doubtlessly it's
going to be a harder sell than if you
had convinced the end user that it's
available first, right?
>> Um, Decagon, Sierra, and Ramp. Um, what
kinds of uh data are they buying
relative to the Frontier Labs?
>> Voice data. Uh, RAMP is not so much
buying data. Actually, the Ramp Labs
report came out um the other day, and I
was surprised at a couple things. One, I
really love the fact that you've got
really sophisticated elite engineering
or applier companies out there
post-raining their own small models for
their own use cases. But I was surprised
that they used a synthetic data set to
inform some of the environments in which
they were training like accounting level
transactions if they're an app layer
company that should have access to that
data themselves
>> which is
which one suggests that one could
sell data to them if they can't use
their own app layer data uh for for
these training environments. But uh two
uh also suggests that Apple companies
may be feasible buyers in the future if
there's a substantial systematic issue
that prevents them from using their own
users data. Certainly doesn't look like
it's been a problem with cursor though.
So I'm sure this is just a small uh
quirk.
>> Yeah. Do you think it's a it's a privacy
thing that they'll just figure out?
>> I think so. Privacy is not like data
privacy is very easy to figure out
nowadays for all these companies.
>> Yep. the um
uh do you have a certain view on you
know long-term
um
the
labs have their applications those
applications give them you know traces
that they can train on um
how it evolves where they still need to
buy data externally versus training on
the data from their users
Um
yeah, one would have thought that
Enthropic has so much data from claude
code and
work
>> right
>> that maybe they would not have needed to
procure from external vendors
but they still do. Um and and and this
reflects the fact that most external
data vendors that are succeeding with
sophisticated research labs and data
markets, they're mostly selling
capabilities that are N plus one of
current tier models, right?
>> Y
>> um Andon Labs, by the way, Andon is a
fantastic company in this regard in
terms of producing really hard realistic
benchmarks, but a bit too ahead of its
time, I think.
>> Yeah. Um, and on labs is a good example
of the fact that we're we we're going to
produce these really real world long
horizon benchmarks that are not going to
be saturated for a long time and that is
quite available to us.
>> Yep. So if it's already within the
capabilities of the model then they can
train on it from their traces but if
it's not and no user is going to attempt
it in the model then they don't have any
traces to train on. And this is from a
purely single axis performance-based
perspective, right? Whereas it's like
there's only one thing to help climb and
it's is perceived performance. Um cost
and latency are also big questions too.
An anthropic researcher I think told me
at some point our benchmarks are really
not going to index on performance and
that we'll have prohibitively expensive
AGI in some sense but like how much does
it cost and how fast does it take to do
something is going to be new
>> new new dimensions of benchmarks. So
then you expect that um end vendors will
uh start to do benchmarks that are
basically performance divided by price
rather than just performance essentially
>> perhaps. Yeah. And this expands greatly
the aperture of different niches that RN
companies can play in because if you
think about the enterprise world, right?
There are many use cases where I just
want a much much cheaper model at a
fixed level of intelligence
>> that is satisfactory for certain like
job functions, right? And then even in
ramp lab's recent implementation on
their Twitter post they showed that they
use a above head frontier model for
planning but they they collapse the
search and retrieval function to a small
model that they post trained just for
that purpose.
>> The um how many labs are spending at the
you know billion dollar plus per year
data level?
>> Seven or eight.
>> Mhm. the how much more than
like you know Anthropic talked about
their billion dollar number. Do you
think it's going to end up being like
closer to like you know three to four
kind of this year?
>> Yeah. I mean I'd say like honestly each
Frontier Lab if you're loose with your
definition of data like they spend
between 10 to 20 billion a year. I think
I posted about this a while back too.
>> 10 to 20 if you're loose with your
definition of data. Um yeah. Can you say
how so? Uh this is this shouldn't be a
surprise to anybody, right? Like three
things Hill climb model capabilities,
compute, data, and talent. And data
spend is still a drop in a bucket
compared to compute costs, right? Um I
I'd say we're generally still supply
constrained in that if you think about
RL data or just data in general, that
means the quality bar for these labs,
we're still very much still in demand of
that data.
>> Yep. So you're saying 10 to 20 billion
in aggregate?
No, per lab.
>> Per lab
with eight labs spending that much
>> uh sevenish.
uh I think for some labs
>> including this isn't like salaries of
data team people is included like how
does it get to
>> like literal data from external vendors
and and and by the way most of this
spent does not actually get satisfied
like I'm sure that there is a data
budget set aside whose upper limit is
not actually met because there's simp
just simply not enough good quality data
vendings. I've still seen I have still
never seen a data contract get turned
down by a top lab if it's good quality
data for budget reasons.
>> Yeah. What's the delta between the
billion dollar number versus the you
know 10 to 20 billion like what's
included in the latter that's not
included in the former?
>> Uh I would say body shop type data
labeling that's very emblematic of scale
type what what scale used to do uh and
what many people still think the data
industry is which is just manual manual
data labeling for pre-training data. Um
>> so then that would be like you know 70
billion plus in aggregate. Um
what's what's what's like the ballpark
of like surge scale
annual revenue
>> surge is between two to three bill
runway rate I'm pretty sure
>> what's the yeah where's the where's the
gap come from like if Serge is you know
leading provider they're doing two to
three 70 billion aggregate spend
>> um there are so many companies that
participate in data markets that you
would have never even expected just a
big massive long tail basically.
>> Yeah, it's an it's a very massive long
tail. Yes. Uh also
>> yeah staffing agencies as well. It's
like uh uh and this encompasses a lot of
the spend that OEI and anthropic
directly have like acquiring companies
from the real world too just for data
assets.
>> Yeah.
>> Uh which certainly I don't know why like
is happening a lot more and more and
people are not discussing this very
closely. M
>> um
>> this is like acquiring little like
little wet labs and that kind of stuff
>> like app layer companies in certain
domains that they're they're interested
in building products in right
>> uh I I I can't name them specifically.
Um
>> enterprise software type small app
player companies.
>> Yeah. Yeah, you could say that with like
network effects from like having I don't
know 10 to 15 years worth of user
activity like a stack overflow type type
thing. Uh so the the data markets as
exemplified by like Merur and these
companies they represent like the tip of
the iceberg in terms of like the the
entire long tale of companies where data
procured actually comes from.
>> Cool. As a last question, um the
what makes inference providers and
neoclouds a good fit uh to acquire RLM
codes is that they basically act as
implementers to the enterprise partners
that are their customers.
>> They they are the compute they are the
compute providers for our labs as well.
It naturally makes sense that they want
to do horizontal product expansion and
bring post-training infrastructure.
>> Yeah.
>> Uh and tooling alongside their product
offering to labs. B 10 actually I think
it was B 10 made an RLM's acquisition
>> uh like in December January time that
very few people are talking about so
there's precedent and I think um some of
the sophisticated RLM targets are very
good acquisition targets for this
>> both help themselves to enterprises and
to labs
>> uh to build out their uh post training
infrastructure product suite
>> the the and the end customer the post
training infra is um mostly like non-top
three frontier labs just like other
enterprises.
>> Yeah. Yeah. Like app layer companies
too. Y
>> um like for a while while you know
Perplexity and Cursor were more than 50%
of fireworks revenue for example.
Continue with YouTLDR
Analyze another video with Pro
Process a new video, search every timestamp, compare sources, and keep the result in your library.
More transcripts
Explore other videos transcribed with YouTLDR.

The Truth Behind Liquidity
Inter Equity Trading · English

Kant: KrV B 46 – 59 – D. Hattrup liest
Dieter Hattrup · German

Kant: KrV B 31 – 45 – D. Hattrup liest
Dieter Hattrup · German

Secrets To Identifying Correct Liquidity
Inter Equity Trading · English

Finding the Daily Bias ONLY Using Liquidity
Inter Equity Trading · English

Kant: Kritik der reinen Vernunft 1787 (Vorrede B VII) – Dieter Hattrup liest
Dieter Hattrup · German

D. Hattrup liest – C.F. von Weizsäcker: Wahrnehmung der Neuzeit: Einstein
Dieter Hattrup · German

هل يمكن الوثوق بعقلك؟ كيف غيّر هيوم وكانط فهمنا للحقيقة
الفلسفة للنوم · Arabic

Opus 5 released! Is it better than Fable?
Mastra · English

Leilão de Embriões Nelore PO DNA Genética Aditiva
LANCE RURAL OFICIAL · Portuguese (Portugal, Brazil)

Leilão Peso Pesado Rima Agropecuária
LANCE RURAL OFICIAL · Portuguese (Portugal, Brazil)

النبي .. جبران خليل جبران .. إقرا بودانك
اقرا بودانك · Arabic