Full Transcript

·YouTLDR

The Prime Intellect Stack — Will Brown, Prime Intellect

46:52EnglishBy AI EngineerTranscribed Jul 15, 2026
Analyze another video with Pro30-day money-back guarantee
0:12

Hey guys, how's it going? Thanks for

0:14

showing up. This was a little bit of a

0:15

last-minute assembly. I

0:18

know a few days ago I was like talking

0:20

to Swix. I was like, "Hey, can I still

0:21

do a workshop?" And he was like, "We

0:23

have one slot left. It's Monday at

0:24

4:30." And I was like, "I'll take it."

0:26

Um and uh

0:27

then yeah, um I wanted to kind of just

0:29

do a bit of an update on uh

0:33

some of the stuff we've been building at

0:34

Primed and Loaded. So, if you don't uh

0:35

know me, hi. I'm Will Brown. I lead

0:37

applied research at Primed and Loaded.

0:39

Uh we do a lot of stuff around uh every

0:42

part of the kind of AI research

0:43

infrastructure stack. Uh today is going

0:45

to be about post-training, which is

0:46

where I spend a lot of my time thinking

0:48

and building. Um and especially want to

0:50

be talking about uh the post-training

0:51

tools that we build uh that are fully

0:53

open source, uh the verifiers and Primed

0:55

RL libraries, uh which kind of go hand

0:57

in hand um

0:58

both on the environment side and the

0:59

training infra side. Um

1:01

and show [clears throat] off some things

1:02

we've been cooking over the past few

1:04

months that I think is uh kind of the

1:06

way that things have evolved as

1:09

the agent use cases have gotten more

1:11

complex, but also kind of clearer in

1:13

terms of what people want out of agents

1:14

and the sorts of things that are needed

1:16

to like do the sort of post-training

1:17

that is needed to uh power like the

1:20

real-world applications people are

1:21

building nowadays. And so, broadly at

1:23

Primed and Loaded, we are

1:25

our goal is to make doing large-scale

1:29

open-source AI research easier and to

1:31

enable companies to train their own

1:32

models and deploy them and have them

1:35

improve based on the scenarios that they

1:37

actually see in production in terms of

1:40

use cases for applications and products

1:42

and internal tasks and workflows. Um and

1:45

to give people an option to not just use

1:47

the open-source models that are getting

1:49

quite good, but to take them and make

1:50

them even better on their own use cases.

1:52

Um and so, we use the phrase the open

1:54

superintelligence stack to describe what

1:56

we mean by this. And I think when we

1:58

said this phrase like a year ago, it

2:00

felt a little more like marketing and

2:03

now it feels a little bit more like, oh,

2:05

yeah, that's that's kind of what it is.

2:06

Like the models are getting very, very

2:08

good. They are superhuman in many ways

2:10

at lots of things.

2:11

Um

2:12

and what we want to do is give people an

2:14

open toolkit that they can use to do

2:17

real training with them.

2:19

>> [clears throat]

2:19

>> And to have the control that they need

2:21

to deploy it where they need to deploy

2:22

it and customize it as much as they need

2:24

to to kind of get the job done. And so

2:26

this is the stack that we built. And it

2:27

all kind of sits on top of compute. So

2:30

we operate a global marketplace of data

2:33

centers around the world. A lot of these

2:34

are like quite large data centers. We

2:36

currently operate over 10,000 GPUs.

2:40

Many in like hundreds or thousands

2:42

within a cluster.

2:43

We have our primary training framework.

2:45

We have environments built with the

2:47

Verifiers library and our environments

2:48

hub platform. We have our

2:51

platform for research workflows that

2:53

we're now calling lab, which is an

2:55

assembly of many pieces including the

2:56

environments hub, hosted training

2:58

evaluations, as well as inference and

2:59

sandboxes. And all of this is in service

3:01

of empowering and unlocking frontier

3:04

model training. And so we do this

3:06

ourselves. We have our intellect model

3:07

series with some exciting things there

3:09

coming soon.

3:11

And we also train models with our

3:12

customers where we have lots of people

3:13

we work with who their goal is to do

3:16

large scale model training on their own

3:18

workflows.

3:19

And so to do all of this, we need to

3:21

give people the tools they can assemble

3:23

into the the the pipelines, the

3:25

workflows, the research that allows them

3:27

to actually get the results that they

3:28

need at scale with everything they need

3:30

to do it.

3:31

And so this talk is going to be about

3:33

going deep into Verifiers and Primary L

3:35

and showing off some of these new

3:37

things, but all under the umbrella of

3:39

what does modern post training look

3:40

like? What does it mean to kind of take

3:42

a model and train it to be better at

3:44

your task? What are all the parts? What

3:46

are all the kind of gotchas?

3:48

And how do you orchestrate this into a

3:50

system that is actually easy for people

3:52

to use without needing to go build a

3:54

massive research team and to be able to

3:56

kind of have it be accessible in the

3:57

sorts of things that anyone who's an AI

4:00

engineer at

4:01

any like startup or enterprise that

4:03

wants to invest in post-training can

4:04

actually do.

4:05

And so there's a cookbook repo that is

4:07

that's kind of like an alpha release

4:09

right now. It's still changing a bit,

4:10

but it's a preview of kind of all the

4:12

stuff we've been building over the past

4:13

several months. And so today we'll be

4:15

kind of following along

4:17

that framing a good bit. And so

4:20

I think the first thing we'll talk about

4:21

is just kind of what is an environment.

4:23

People talk about environment in the

4:24

context of RL and they think of like RL

4:26

environments, but environments are more

4:27

than just for RL. They're for all sorts

4:29

of things in post-training and

4:31

evaluation.

4:32

We're going to talk about what we're

4:33

going to call the the V1 version of the

4:34

Verifiers library, which is a full

4:36

overhaul. Everything else still from

4:38

before still works, but we're kind of we

4:40

kind of wanted to redo it all. And so we

4:41

have a new way of doing everything that

4:42

we think is going to make a lot more

4:44

sense, be a lot more powerful for what

4:45

people are looking to do going forward.

4:47

As well as kind of talk about how Primer

4:49

RL has evolved as a library. And so

4:51

Primer RL is our like full-stack

4:53

open-source training framework to

4:55

support asynchronous reinforcement

4:56

learning.

4:57

And we've got a lot of fun new bells and

4:59

whistles to show off in terms of both

5:01

scale and features. A lot of this is in

5:03

service of custom algorithms. So making

5:05

it much easier to

5:07

do the kinds of things that people are

5:09

interested in for modern post-training.

5:10

If you have been following the news on

5:12

on policy distillation or

5:13

self-distillation or all these other fun

5:15

new algorithms that people are coming

5:16

out with. It is the the age of research

5:18

indeed.

5:19

And we don't want to just like train

5:21

small models. We want to train big

5:23

models. We want to train them really

5:24

efficiently because as models get

5:26

bigger, the compute starts adding up.

5:29

And if you want to make this accessible

5:31

to people, especially if you want to be

5:33

able to iterate on it, it has to be

5:34

fast. It has to be cheap. It has to be

5:36

affordable and reliable.

5:39

And all of these kind of funnel into our

5:41

lab platform. We'll talk about both some

5:43

of the things that we've already

5:43

released there as well as some things

5:44

that are coming soon.

5:46

And so the post-training loop in my mind

5:50

kind of revolves around environments in

5:52

the sense of environments are a language

5:54

for specifying what you want your model

5:56

to do.

5:57

Um, they are an encapsulation of the

6:00

data you might have, the scenario you

6:02

might want your agent to be in, uh, the

6:04

way it'll interact with that

6:05

environment, uh, as well as how to to

6:08

score what good looks like, to determine

6:10

what was good and bad. Um,

6:11

and

6:13

often this is the first thing you want

6:15

to do with an environment is just evals.

6:17

And so I think a lot of people are maybe

6:20

nervous about getting into

6:20

post-training. They're like, "Oh, it

6:22

seems like a lot of work. There's a

6:23

whole new tool chain. Um,

6:25

what if I'm already using like the

6:27

frontier models and I want good results

6:28

out of them, uh, or I'm getting good

6:30

results out of them, or I want to like

6:31

see what I can do at the harness level

6:32

first or prompt optimization." And

6:34

that's all good. Like we're not

6:36

necessarily asking people to just like

6:37

throw everything away. I think in many

6:38

cases what people will find and what we

6:40

see with our customers is that um, the

6:42

systems that work best for them involve

6:44

using both. And you kind of want to be

6:46

able to make these decisions about what

6:48

is the right where is the right place to

6:49

train, uh, where's the right place to

6:51

use a frontier model that is available

6:53

via some API. Um, and so evaluations are

6:56

kind of very key to this. And so like

6:57

evals are the thing that opens the door

7:00

to post-training. And so environments

7:01

and evals are essentially the same

7:03

thing. Um,

7:04

but once you have evals, now this is the

7:06

same kind of unit of, uh, like logic

7:10

that you actually need to do

7:11

post-training anyways. And so, uh,

7:13

building evals is like just good for

7:15

your product hygiene no matter what

7:17

you're doing. If you want to kind of

7:18

decide whether to use GPT or Claude, or

7:20

decide do you need Opus or uh, Sonnet or

7:23

Mythos for a task. Like if you want to

7:24

min-max on like intelligence versus

7:26

dollars, um, evals are a very good way

7:29

to do this. But evals also then unlock

7:31

this this flywheel. And in terms of

7:33

modern post-training, I think

7:34

historically people have done SFT then

7:36

RL as like the main uh, frontier model

7:38

recipe, although on policy distillation

7:40

has certainly found its way into a lot

7:42

of workflows. Um, and I think some

7:43

people are also very eager about

7:45

algorithms like self-distillation. We

7:46

can talk a bit about that and when it

7:48

makes sense and when it doesn't, but

7:50

uh, in particular one area where it does

7:51

make sense to do like the whole

7:53

on-policy distillation thing is when

7:55

you're training experts where you have

7:57

multiple different things you want your

7:58

model to be good at, and people have

7:59

found that if you have a bunch of

8:01

different environments that are all

8:02

different things, and you want to have

8:04

one model be really good at them, a nice

8:06

way to do this is train individual RL

8:08

experts on top of the same base model

8:10

and then do distillation from those

8:11

teachers into the same uh, checkpoint.

8:14

Uh, that just generally ends up being

8:15

more reliable.

8:16

Um, and then once you have this, you

8:18

want to deploy the trained model, which

8:20

could be a full uh, base model with a

8:22

full weight training, or it could be a

8:23

Lora adapter, and you want to serve this

8:24

at scale, and ultimately what is useful

8:28

about this whole process is it's not

8:29

just a thing you do once. Like I think

8:31

some people also say like, "Oh, why

8:33

should I do post-training if the the

8:35

frontier models are going to get

8:36

better?" Well, your model should get

8:37

better, too. Like it's not it like

8:39

everything's going to get better. Uh,

8:40

the point of this is to have flywheels

8:41

that make everything get better. And so

8:43

what you really want is to be able to

8:44

not just post-train like today, but to

8:46

be able to uh, have this iterative

8:49

process of model refinement, and the

8:51

sort of thing where you can kind of have

8:53

the training compute end up be a pretty

8:54

small fraction of your overall inference

8:56

budget that you amortize out such that

8:58

your model is always getting better and

8:59

better um, as you are getting more

9:01

signal from the real world. And getting

9:03

this signal from the real world isn't

9:05

trivial. Like that's kind of largely an

9:07

open question as to like how you go

9:08

about um,

9:10

getting information from real world

9:12

feedback into your environments. It's an

9:14

engineering problem, it's a research

9:15

problem, but it's the sort of thing

9:17

we're all here at this conference to

9:18

kind of think about and learn about. And

9:20

so I'll I'll touch on that a little bit

9:22

uh, in the talk as to kind of how we we

9:23

think about this. But but really the

9:25

goal is going to be thinking about what

9:27

do these tools look like? How do you

9:28

actually do this? What are the parts?

9:30

Um, and how do we build it?

9:31

Um, and so environments as evals, uh,

9:35

Uh, what is an environment? Um,

9:37

I think it's useful to decompose

9:40

environments into tasks and a harness.

9:42

Um, and this is foreshadowing some of

9:44

the the refactoring we've done in

9:46

verifiers over the past months. Uh, if

9:48

any of you have used the verifiers

9:49

library before, you may be familiar with

9:51

the the multi-turn environment pattern

9:53

or tool environment pattern where

9:55

there's kind of one loop that is owned

9:57

by the environment that you can plug in

9:58

various tools into and that was really

10:00

great for a very long time for getting

10:02

started for people, especially back in

10:04

the day when people were mostly just

10:05

trying to graduate from single turn into

10:07

multi-turn tool tool use. Uh, but what

10:09

we found and as we kind of iterated on

10:12

different patterns and extended it in

10:14

various different ways, we found

10:15

ourselves repeating a lot of work of

10:16

like adding patterns for a CLI agent or

10:19

adding patterns for MCP. Um, and we

10:22

wanted to be able to like

10:23

step back and rethink like how should an

10:25

environment work. Um, and what it really

10:27

is there's a notion of a harness. And so

10:30

and you there's also a notion of a task.

10:31

And I think one of the reasons that this

10:33

is kind of subtle and tricky and was a a

10:35

design problem that we went over many

10:37

iterations over the past 6 months really

10:39

is um,

10:40

certain things it's not clear where they

10:42

live. Like there's certain things that

10:43

might belong to the harness and might

10:44

belong to the task. There's certain

10:46

tools that uh, in some cases it's I want

10:49

this uh, task I'm going to do to use a

10:51

certain tool. In some case the harness

10:53

has certain tools. Same with skills or

10:55

system prompts or many other pieces of

10:57

the puzzle in assembling like the full

10:59

world that your agent is going to be

11:00

operating in or your model is going to

11:01

be operating in. Um, but ultimately the

11:03

we're going to call all of these parts

11:05

of the environment. Um, and the goal of

11:07

this is to have some notion of

11:08

verification where you plug in a model

11:12

uh, into the environment which includes

11:13

a harness. You give it a task. It does a

11:16

rollout and then you verify what it did.

11:18

Um, and this same process works both for

11:21

evaluation offline just understanding

11:23

which model is better as well as for

11:25

doing reinforcement learning RL uh, as

11:27

well as for generating data for SFT. Uh,

11:30

I think in many cases people think of

11:32

SFT is this thing where they want to

11:33

like upload a data set, but really often

11:35

what they're doing there is they're

11:37

essentially cobbling together something

11:39

that's essentially an environment and

11:40

they're doing rollouts in it and saving

11:43

it offline and putting it in one format

11:45

and uploading it and then changing it to

11:46

another format and then plugging it into

11:48

a trainer. And the way we've kind of

11:50

approached this is saying, "Well, you

11:51

can just kind of cut all that out and

11:52

just like treat it like a a a problem

11:55

where you're doing rollouts in an

11:57

environment. Just in this case there's a

11:58

teacher." And the teacher could be

12:00

another it could be replaying from

12:02

another data set.

12:03

But ultimately it's about collecting

12:05

rollouts and training on those rollouts.

12:07

And then

12:08

on policy distillation again takes the

12:10

same form where you are doing rollouts

12:12

in an environment just as you would for

12:13

RL, but the scoring is from a teacher

12:16

and the the log probs of the teacher the

12:18

likelihood of the teacher rather than

12:20

the the reward signal itself. And so

12:23

these all in our framework are Python

12:25

packages. So you can have any task as

12:27

you want. You can pull data from

12:30

anywhere you want.

12:31

There's a lot of flexibility that we've

12:33

unlocked in terms of what these tasks

12:34

can look like, what these harnesses can

12:35

look like.

12:37

And our goal is to just make this a

12:38

really flexible toolkit for all the

12:39

kinds of evaluation things people want

12:41

to do both for API models as well as for

12:43

post training.

12:45

And so Verifiers V1 is what we're

12:48

calling it, which is it's not actually

12:49

released as V1 yet, but we took

12:51

inspiration from VLUM doing this

12:54

and decided that we were going to kind

12:56

of have this be the new pattern that we

12:59

want to have everything kind of

13:01

be centered around. And the key pieces

13:04

we broke things down into were a task

13:06

set, a harness, and a runtime. And so

13:09

these are all composable. You can mix

13:11

and match them

13:12

and they're all individually loadable in

13:15

different ways.

13:16

But the way to think about it is that

13:19

task sets are the data and the rules of

13:22

what should be done that are agent

13:25

agnostic. So they're the sort of thing

13:26

you could plug an agent into.

13:28

Um and we wanted to take a very general

13:30

approach in supporting a lot of the

13:31

great work being done throughout the

13:33

ecosystem. So, we integrate natively

13:35

with Hugging Face datasets, with Harbor,

13:37

uh with NeMo Gym and Open Ended. And

13:40

most other tools that you see out in the

13:42

wild that are kind of under the umbrella

13:44

of uh an RL environment, we would call

13:47

these a task set. Um

13:48

we generally have found that it's useful

13:50

to have these be harness agnostic where

13:53

they represent the the back end of the

13:54

server or some state that you're

13:56

interacting with, but they don't own

13:59

everything about what the model is

14:01

doing. And so, it doesn't In some cases,

14:02

it doesn't make sense to plug a model

14:04

into a task set. Especially because

14:06

we're kind of gravitating towards an

14:07

agent world where everything is running

14:09

in a terminal or it has skills or it uh

14:12

is using CLI tools. Um and these things

14:15

like often look more complex than just

14:18

basic loops. Um

14:20

but we also want to support basic loops.

14:21

So, we want to kind of allow both the

14:23

old way of doing things and the new way

14:25

of doing things. And so, uh everything

14:27

that was the old way is now the default

14:28

harness where it's system prompt and

14:30

tools in a loop. Um but the harness

14:32

pattern

14:33

also supports much more flexible

14:35

execution of things like recursive

14:37

language models or CLI agents like

14:39

Codex, Cloud Code, Open Code, or

14:41

classics from the research literature

14:42

like Mini Sweep Agent or building your

14:44

own with arbitrary Python libraries like

14:46

LangChain or DSPy. Um and so, we've been

14:48

able to decouple these into a pattern

14:50

where you get to write your harness

14:52

independently of your task set. Uh there

14:54

are kind of some basic sanity checks

14:56

about properties that like harnesses

14:58

either do or don't support and task sets

15:00

do or don't require. Uh and these kind

15:02

of click together. And the runtime is

15:04

where this executes. And so, we've uh

15:07

still been embracing a lot of the async

15:09

IO patterns from before, but we've

15:11

leaned a little more into having things

15:13

be subprocesses uh where you can still

15:16

run everything locally. You don't have

15:17

to use sandboxes, but you can use local

15:19

Docker, or you can use our own Prime

15:21

sandboxes layer. You can use any other

15:23

sandbox layer you'd like or kind of

15:25

build from scratch. And so the harness

15:28

the runtime back end

15:30

just is a place where the harness can

15:31

run its code. And so the harness just

15:32

needs to be able to run code somewhere

15:35

as a script essentially.

15:37

We've used a lot of the UV tooling where

15:38

UV script is a very powerful pattern to

15:41

be able to kind of mix and match and

15:42

kind of contain dependencies.

15:44

But what happens is once you plug these

15:45

together, you run a rollout on a task

15:48

from a task set and you get a trace.

15:51

This is live on the verifiers main

15:53

branch for prim and elect AI / verifiers

15:55

on GitHub as well as it's released as a

15:57

dev release. The stable main release

15:59

will be kind of coming to PyPI any

16:01

minute now, but you can install the dev

16:03

and play around with it if you want.

16:05

And so what do these look like? So tasks

16:07

are just like a row of a data set.

16:09

And the very basic version of it is

16:11

you just start loading a data set from

16:13

hugging face or anywhere else.

16:16

And so the new pattern here is

16:18

from verifiers V1 I just to keep the old

16:21

stuff separate. The old stuff still

16:22

works just fine, but this is how we have

16:24

been able to kind of decouple and

16:25

iterate on the new version.

16:28

As well as we've really embraced this

16:29

decorator pattern.

16:30

We found it to be very useful. We also

16:32

if you were a rubric fan, we killed

16:33

rubric.

16:35

Didn't make sense anymore if you were

16:36

using old verifiers rubric patterns.

16:39

But still it's you have functions and

16:41

loaders.

16:43

We are very heavy on PyDantic so

16:45

everything is super typed. We have lots

16:47

of powerful config features where you

16:48

can have everything in a toml file. You

16:50

can override it in the CLI

16:53

and everything is kind of clean and

16:54

guaranteed to kind of type check at like

16:57

validation time rather than waiting for

16:59

something to fail later down the road.

17:00

And so examples of this are things like

17:02

sweet wrapper where you can do a genetic

17:03

code search. You can do the classic

17:05

games like Wordle. You can do

17:07

search over documents with judges. You

17:09

can do complex things like harbor that

17:11

support a lot of popular benchmarks now

17:14

that need agents running in a terminal.

17:16

And all of these are going to be

17:17

combinations of the the the task that

17:19

pattern with uh pick your own runtime

17:22

and pick your own harness.

17:24

Um and so rewards and metrics I think

17:25

are also kind of uh just functions that

17:27

take in the kind of records uh of what's

17:30

happening in a rollout and return

17:32

numbers. Uh so rewards are the main

17:34

thing that'll drive progress in RL. Um

17:36

metrics are just kind of like logging uh

17:39

what has happened. So counting tool use

17:40

and counting errors. Um these sorts of

17:42

things are very useful to be able to

17:43

expose in your dashboards. Um and then

17:46

group rewards. I think this is something

17:47

that we have fought hard to kind of make

17:50

sure still is first class because we see

17:52

it as very important to um a lot of the

17:55

research pattern people want to do, but

17:56

I think it's also ignored in a lot of

17:58

like uh tooling out there. Where in in

18:00

many RL frameworks it's actually quite

18:02

hard to do group rewards because things

18:04

are very decoupled and things kind of

18:05

assume that all rollouts are going to

18:07

live independently and that they don't

18:08

need to talk to each other. But there's

18:10

a lot of things where you really want to

18:11

do pairwise judging or you want to do

18:14

ranking or you want to give a bonus to

18:16

the uh the shortest correct answer uh in

18:19

terms of tokens used. Um

18:21

and so these sorts of things are really

18:22

flexible uh in terms of the we really

18:25

design for flexibility in supporting the

18:27

the things that we see as like the most

18:29

exciting papers we've read or all the

18:30

algorithms that we think people may want

18:31

to innovate on um while still allowing

18:34

people to have like the core primitives

18:35

that they kind of expect out of an RL

18:36

framework.

18:38

Um and so like in group rewards, I think

18:40

this concise this pattern is one that I

18:41

find very useful a lot. I think um like

18:44

a big pattern that comes up a lot are

18:46

doing post-training is uh models will

18:49

love to like think and think and think

18:50

if you let them. Um and if you don't

18:53

give them some kind of pressure to like

18:55

be more efficient, uh I think a lot of

18:58

people will notice that like open models

18:59

often have really really long chains of

19:01

thought. Because on one hand it's like

19:04

this is a useful strategy for a model,

19:06

but it's also the sort of thing that

19:07

will grow like out of control if you

19:09

don't counteract it. Um and so in reward

19:12

design like one of the big things people

19:13

will want to do is uh something like a

19:16

length penalty um or a conciseness

19:18

bonus. Um, and so one of the reasons

19:20

this is tricky is because you don't know

19:21

the optimal length for a problem. Like

19:23

if I give you a math problem, I could

19:25

say, "Oh, solve it in less than N

19:27

tokens." But also like who knows what

19:29

the right N is. It's also going to

19:30

change as the model gets smarter over

19:32

time. It's going to be different for

19:33

every problem. And so you kind of can't

19:35

know this up front. And the only way you

19:37

can do it is take advantage of variance.

19:39

So one of the nice things about RL is

19:42

you have multiple samples typically. Um,

19:45

and this allows you to use the fact that

19:46

you have multiple samples to shape the

19:48

reward. Um, and so if you have multiple

19:52

rewards in a group, what you could do is

19:54

uh look at all the ones that were the

19:56

correct answer or just all the ones in

19:58

general and give a bonus to the ones

19:59

that are the most concise. Where if you

20:01

also have a uh a correctness reward like

20:03

these are going to uh ensure that um

20:07

you're both incentivizing correctness as

20:08

well as incentivizing efficiency. And so

20:12

juggling multiple objectives

20:13

simultaneously is kind of one of the

20:14

hard challenges in RL in reward design.

20:17

Um, but doing things like group level

20:19

comparisons and kind of these sorts of

20:20

bonuses are are quite useful in many

20:22

cases.

20:24

Um, I also want to talk about like tools

20:26

and user simulators, which I think have

20:27

been uh

20:28

becoming more important uh in a lot of

20:31

complex applications where you have

20:32

models that are In many cases there's

20:34

like a core agent harness, but there's

20:35

also In many cases you are putting a

20:38

model in a setting where it's going to

20:39

be in some task where a user is giving

20:41

it additional tools. Whether these are

20:43

In some cases you might want to model

20:44

these as skills. In some cases you might

20:46

want to model them as MCP servers. We

20:47

use MCP as a kind of uh a back-end

20:50

framework um that can interact with the

20:52

runtime uh both for tools and for user

20:54

simulators. So user simulators,

20:55

especially if you want to do training

20:56

where there's a a user in the loop, um

20:59

you don't want just your agent to go do

21:00

some task, but you want to be able to do

21:02

some task that involves understanding

21:04

how a user will interact. You can

21:06

essentially have this user be an MCP as

21:08

well where we make it so the model sees

21:10

it as a user, not as a tool. Um but

21:13

behind the scenes, it is a server that

21:15

has some script or something and it has

21:17

some LLM that is going to get some

21:19

context and it's going to be able to be

21:21

like a user in the context of a rollout.

21:24

Uh in many cases, benchmarks have found

21:26

that this is very useful for simulating

21:27

the realism of like having a multi-turn

21:29

setting where there are users in the

21:31

loop, especially given that people are

21:32

building products now where there are

21:34

these users in the loop. And so you want

21:35

kind of a first-class way to incorporate

21:37

this into um your uh your RL

21:40

environments. And so we've found that

21:43

it's useful to have all of these things

21:44

be kind of modular and pluggable. And so

21:48

the harness can connect to each of these

21:49

which run as a as a UV script. Um we

21:52

also have UV script support for grading

21:54

in addition to the basic reward function

21:56

patterns. Um and we also are using this

21:58

idea called an interception server. So

22:00

the harness, we want people to be able

22:02

to use real harnesses and not have to

22:04

like break the harness and like retrofit

22:07

it into like an RL harness. And so

22:09

ideally, we don't know anything about

22:11

the harness code. And so the pattern

22:12

that we use with the interception server

22:14

is that um

22:16

we are responsible for giving each

22:17

harness rollout a fake base URL, which

22:20

could be OpenAI compatible or Anthropic

22:22

compatible. And the harness just thinks

22:24

it's talking to some endpoint. So any

22:25

harness that can just talk to some

22:26

endpoint, we're good to go. And then we

22:28

intercept each request. Uh we can do

22:31

some some back-end maneuvering to make

22:33

sure that we're getting the log probs

22:35

and setting the right temperature. Um

22:37

and then we send this to our uh

22:39

inference server for with the RL

22:40

trainer. Uh and then as it completes, we

22:43

send back the request. And so the the

22:44

harness doesn't know that it's doing RL.

22:46

The harness just is a harness running as

22:48

if it would be running in a real-world

22:50

environment. And so you can kind of very

22:52

easily move between the RL setting and

22:54

the deployment setting where your

22:56

harness is just code. Doesn't need to be

22:57

anything specialized to verifiers or RL.

23:00

We also have found it really easy useful

23:02

to be able to kind of go from this like

23:04

local to global and like have this hot

23:07

swap pattern. Um so we have this new

23:09

like uh eval CLI where um

23:12

you can just like choose the harness you

23:13

want. You can have a task set where

23:15

you're going to say I'm going to run an

23:16

eval on this set of tasks. Okay, I want

23:18

to use recursive language models and

23:19

RLMs. I want to use Codex. I want to run

23:22

it locally. I want to run it in

23:23

sandboxes. Uh I want to run it in

23:24

Docker. These are all just like

23:26

interchangeable. Um and so we found that

23:27

this is super useful in the iteration

23:30

loop as you go from testing something

23:32

out at a small scale towards scaling it

23:34

up towards being able to understand

23:36

questions like what's the best harness

23:38

for this model? Um does this harness

23:40

generalize across tasks? Um as well as

23:43

being able to both have the convenience

23:45

of like local prototyping where you can

23:47

kind of run things fast on your MacBook

23:49

um without having to like wait for a

23:50

cloud job to finish, but also you can

23:52

like go right to the cloud when you need

23:53

to. And so there's been a lot of fun

23:55

patterns we've had to kind of innovate

23:57

on um

23:58

as we've done this overhaul. And we

24:00

we're quite happy with how it's turned

24:01

out. It's made our lives a lot easier

24:03

for both client projects and research

24:05

and just uh being able to have have a

24:07

lot more flexibility and power and

24:08

control over the the kinds of agents we

24:10

want to be training.

24:12

Um and so one of the fun things behind

24:13

the scenes uh is what we call the trace

24:15

graph. And so we had kind of been having

24:19

this grow out of control in terms of the

24:20

old way of doing things and we decided

24:22

this was another opportunity to like

24:23

really overhaul our system to like have

24:26

really good support for sub agents and

24:28

parallel branching trees while also

24:31

still preserving the kind of linear

24:33

sequential dependencies that you need

24:34

for RL with uh careful token control. Um

24:37

and so here uh there's a a notion of a

24:41

of branches that are kind of like at the

24:43

message level. So conceptually um

24:46

the things that matter logically in

24:48

environment space and in harness space

24:49

are messages which are just text. Uh the

24:52

harnesses don't think about tokens, uh

24:54

but if you've done any RL

24:55

experimentation you may have uh

24:57

encountered issues where uh

24:59

re-tokenization or like some messages if

25:01

a model will say something and you turn

25:02

it into text and you put it back through

25:04

tokenizer, it can change a little bit.

25:05

That because tokenization isn't is many

25:07

to one.

25:09

And so this causes lots of very subtle

25:11

numerical problems, especially late in

25:14

large scale training runs. And so you

25:15

want a really nice back and forth

25:17

between uh messages and tokens. And so

25:21

the trace data structure that we created

25:23

here partly is to enable this where we

25:26

can

25:27

store things both at trace level and

25:29

then map them back into token level in

25:31

the right sequences as needed. And we

25:33

also released a library called renderers

25:36

recently, which is a standalone toolkit

25:38

that anyone can use that we have found

25:41

the sort of thing that we're working

25:42

with some of the inference tooling to

25:45

support. So renderers are really all

25:47

about like essentially rethinking

25:49

tokenizers and chat templates where

25:51

behind the scenes it's just making calls

25:53

to the tokenizer,

25:54

but chat templates if people have

25:57

spent time debugging with them, it

25:59

sucks. Ginger is awful.

26:01

It's very very painful and there's so

26:03

many subtle things that we kept running

26:04

into where like

26:05

a model would sometimes have an extra

26:07

new line and the chat template would

26:09

strip it out and this would like cause a

26:11

mismatch in your trainer and inference

26:13

that would either force you to go off

26:15

policy because you now have a trainer

26:16

and inference mismatch or it would cause

26:18

a logical branch where a thing that is a

26:21

branch in like uh

26:24

it becomes a branch in token space even

26:25

though it shouldn't be in logic space

26:28

because of tokenizer subtleties.

26:30

And so renderers as as an abstraction it

26:32

was kind of pioneered by

26:34

OpenAI's harmony with the GPT- OSS

26:36

release and used prominently in thinking

26:38

machines cookbooks as well for tinker,

26:41

but we found it was useful to just kind

26:42

of make it a standalone thing. And so

26:44

this is just a Python library that

26:45

doesn't depend on any other prime stuff.

26:47

You could use it with any inference

26:49

engine you want

26:51

just as a standalone thing that is

26:53

really designed for being able to manage

26:55

this token in token out concatenation

26:59

without thinking about it too much

27:00

yourself because we kind of turn each of

27:03

these chat templates for these the

27:04

popular models into programmable

27:07

artifacts

27:08

where you can do things like

27:10

look up a history of secret you can use

27:12

the kind of history of a trace to be

27:15

able to understand like what is the

27:17

right tokenization? Like do I

27:18

essentially have like a logical prefix

27:20

hit in message space even if I don't in

27:23

tokenization space after re-tokenizing?

27:25

And so this is the sort of thing where I

27:26

think people have gone back and forth on

27:28

like whether they want LM APIs to be

27:30

stateful in general. I think a lot of

27:32

people were hoping that like we could

27:34

just have every model API be stateless.

27:37

I think

27:38

maybe people are less concerned about

27:40

this now because we're moving towards

27:41

this agent world where agents themselves

27:43

are going to be stateful APIs.

27:45

But I think this has revealed to us like

27:47

going through all of the the things here

27:49

like why opening eye responses decided

27:51

to be stateful. There are some kind of

27:53

like unavoidable issues that kind of

27:55

come up when you're doing large-scale

27:58

agentic rollouts where you you do need

28:01

to kind of manage this very carefully

28:03

and it's kind of unavoidable just

28:04

because of how tokenizers work.

28:07

And so you want to be able to maintain

28:08

these dual streams of the logical text

28:11

and the the tokens and you kind of want

28:13

these to be cleanly interoperable where

28:15

users don't have to think about the

28:16

tokens very much but the trainer gets to

28:18

see everything nicely in token space

28:20

as well as the inference engine.

28:22

Um, and so from the harness interception

28:25

server we have clients that can be used

28:28

both for training and inference. And so

28:29

you can kind of like

28:30

swap between these modes without

28:32

thinking about it because

28:34

certain models like don't need to in in

28:37

a training setting you need to be able

28:39

to get log probs and set the

28:40

temperature. Some model APIs won't let

28:42

you do this. They won't return log probs

28:43

because like open eye models with

28:45

reasoning like won't show you the

28:46

reasoning trace so there's no way they

28:47

give you the log probs for everything.

28:48

And so like that's fine, it's just eval

28:50

only and so we have this this client

28:52

layer where you can go between eval and

28:53

train to be able to support all these

28:55

models, but we still use the

28:56

interception server pattern either way

28:58

because it allows us to have like this

29:00

notion of a dialect where like you can

29:02

choose

29:03

OpenAI chat completions or responses or

29:05

Anthropic and all of these are kind of

29:08

easily supportable as just like

29:09

translation layers between a raw request

29:12

into something that'll get passed

29:13

through a renderer potentially if you're

29:14

on the train client side and formatted

29:16

into a a message via tokens.

29:18

Um

29:20

And so this brings us to Primer RL. So

29:21

Primer RL is our training framework that

29:23

is consumes the environment. So once you

29:25

have an environment with your task set

29:26

and your harness and your interception

29:28

server and your runtime and your

29:29

renderer and all those things, this

29:31

plugs into what we call the

29:32

orchestrator.

29:34

And so Primer RL has been async from the

29:36

ground up. Uh so I think

29:39

async RL is one of those things that I

29:40

think people were kind of one foot in

29:42

and one foot out and a lot of training

29:43

frameworks you see them uh will still

29:45

kind of support synchronous training. Um

29:48

some people I think have their reasons

29:49

for wanting to do synchronous training.

29:50

I don't agree with them.

29:52

Um

29:53

I think

29:54

kind of want to bite the bullet of the

29:56

off-policyness anyways for reasons that

29:58

come up with agents

29:59

um in terms of you want to be able to

30:02

overlap long rollouts and not always be

30:04

waiting on your slowest rollout. And

30:05

this kind of means you can't be fully on

30:07

policy unless you want to kind of accept

30:09

always waiting on your slowest rollout.

30:11

Um and so this is really why we went all

30:13

in on async. And so the orchestrator's

30:15

job is to allow the inference and

30:17

trainer to just be separate processes,

30:19

separate servers. Uh they don't share

30:21

GPUs. They don't really know about each

30:23

other all that much. They just consume

30:25

from each other. Um the but the

30:26

orchestrator job is to really like

30:27

manage the run. And so the orchestrator

30:29

uh will

30:31

make sure that the environment is

30:32

running with the endpoint mapping to the

30:33

inference server. It'll do rollouts. Uh

30:37

it'll package these up into a batch.

30:39

It'll send this batch to the trainer and

30:40

it'll be up to the trainer to figure out

30:42

what to do with the batch, uh which will

30:44

be kind of printing some uh sequence to

30:47

feed into a loss function um based on

30:49

the the specification. Um and so the the

30:52

server pattern we use for environments

30:53

is just an engine that can like send

30:55

requests to inference and like send

30:56

batches back to your trainers. It's very

30:57

client-server. Um and we found that this

31:00

is just a really useful way to uh allow

31:02

scaling concerns to be decoupled as

31:04

well. And so like for example, you can

31:06

have a lot of environments running or

31:07

you can have one environment running.

31:08

You can have um a bunch of infra

31:11

replicas, you can have one infra

31:12

replica. Um you can have

31:14

sandboxes or no sandboxes. Uh and the

31:16

trainer doesn't care about this, the

31:18

inference doesn't care about this. It's

31:19

just separation of concerns at a system

31:21

level uh allows you to kind of not

31:23

really worry about these things as uh

31:26

like combined units versus like in some

31:28

cases uh people will want to like have

31:30

training and inference on the same stack

31:31

where it's like especially the the

31:33

closer that you kind of fold in your

31:34

logical in one then it's like sometimes

31:36

you can't even run your environments if

31:38

you're not doing RL, but that means then

31:39

you can't really experiment. You also

31:40

can't use them as evals. There's a lot

31:42

of reasons why just like pulling

31:43

everything apart into these pieces and

31:44

just like having nice APIs for them to

31:46

talk to each other uh makes your life

31:48

way easier.

31:49

Um and we've also just been like really

31:51

scaling it. Um and so we've been doing a

31:53

lot of work on

31:55

uh like GLM-5 series and Kimi K2.5 2.6

31:58

series just to make sure that we can do

32:00

like really good large-scale RL

32:03

efficiently. And so some results we

32:04

found we have recently. This was

32:06

uh the run was on GLM-5 before the 521

32:09

came out. It supports 52 as well, but um

32:12

on the latest primary RL version, we can

32:13

do

32:15

um

32:15

a GLM-5 step on 28 nodes in less than 5

32:20

minutes for long horizon coding tasks

32:22

with 131K context, um

32:25

which means you can do a 1,000-step run

32:27

in 3 days. And that costs for rental

32:31

prices about 50K. And so 50K is not

32:33

cheap, but it's like if you're doing a

32:35

full run on a frontier size model

32:38

on like a proper real world agent

32:40

environment,

32:42

like it's a lot cheaper than what

32:44

OpenAI's raising for it. Um it's a lot

32:46

cheaper than like some of the clusters

32:47

people are are uh selling, and it's the

32:50

sort of thing that like just starts

32:51

making sense for a lot more enterprises

32:53

if you can like actually do this. It

32:54

becomes pretty justifiable if you can

32:56

kind of get to the point where uh you

32:58

have the the tooling chain to be able to

33:00

like build ready valves and like uh do

33:02

all this stuff. Um

33:04

and so this is like the sort of thing

33:05

where it's like let's say you do you

33:06

wanted to like you have a bunch of tasks

33:07

that are like representative of like

33:08

your coding workflows, and you want to

33:10

like do a big RL run. Like this is a

33:12

pretty big RL run, but it's also the

33:13

sort of thing that like you could people

33:15

spend this much on tokens in a month

33:17

sometimes. Um

33:18

and so you can uh

33:20

start finding a lot of savings if you

33:22

kind of think about doing large scale

33:24

post training, and that's kind of what

33:25

we're we're here to help people do. Um

33:28

I guess more on the async side uh

33:30

one of the reasons why you really want

33:32

to do async is that um

33:34

there's a long tail of how long your

33:36

coding agents take. Like if you fire up

33:37

a coding agent task, just like think

33:39

your Claude coder Codex tasks, like how

33:41

many minutes is it going to take you?

33:42

Sometimes it'll be 30 seconds. Sometimes

33:44

it'll be like 2 minutes. Sometimes it'll

33:46

be a goal that goes for like 3 hours. Um

33:49

and these can all be rollouts. And so

33:52

one of the goals of async RL is to have

33:54

your like forward progress speed not be

33:57

tied to the speed of your individual

33:59

rollout. And so this means you can kind

34:00

of allow your rollouts to finish long

34:03

after they start and just go into the

34:05

first batch that they can accept them

34:06

even if that is from a much different

34:08

copy of the model. Um

34:10

and so then the the inference server is

34:13

just always taking the latest version of

34:14

the model. And so um what people have

34:16

generally found and what we've done with

34:17

our experimentation and found as well is

34:19

that you can go reasonably far off

34:21

policy. Like I think 16 is where we

34:22

typically are often operating is like an

34:24

average. Um

34:26

but this means that you just also can

34:27

have a lot more room to not worry about

34:30

certain things about speed. Like you

34:32

don't need to worry about the boot up

34:34

time of your sandboxes as much or the

34:35

weight sync time or like the time of any

34:37

environment or your or your grading.

34:38

Like, it's it's fine if you have these

34:40

things that take time cuz you can't If

34:42

you're doing like grading with a judge,

34:43

you can't like force your judge to be

34:45

super fast all the time. And so, you

34:47

want to have a system where it's okay if

34:49

there are pockets of your life cycle

34:51

that don't use GPU time but do use time

34:54

um and you can kind of overlap these

34:56

without kind of wasting GPU cycles. And

34:58

so, that's really the the key benefit of

35:00

of async RL in our eyes. Um and we've

35:03

done a lot of work on the loss function

35:04

side. The DPPO paper is one that I think

35:06

has gotten popular that we've been using

35:08

a lot as well of like how do you make

35:10

sure that this like stuff stays stable.

35:12

Um we found that it's very stable up to

35:13

thousands of uh steps. Um we're pushing

35:16

towards 10,000 in kind of current

35:17

experiments um in terms of the the scale

35:20

that we're trying to kind of get the

35:21

stuff to reliably. Um

35:23

as well as doing this at big batch sizes

35:26

where you kind of need to pull out all

35:27

the bells and whistles on parallelism.

35:30

And so, what we found is that I guess I

35:31

can go back to

35:33

uh like doing like this uh expert

35:36

parallel on the trainer as well as kind

35:37

of parallel. Um and then on inference

35:39

doing uh wide expert parallel for the

35:42

big MOEs uh meaning multi-node uh

35:44

experts across multiple nodes as well as

35:46

disintegrated prefill. Um and you can

35:49

kind of throw all the inference bells

35:50

and whistles that people would do for

35:51

normal serving into your RL stack and

35:53

get same wins there as well.

35:55

Um

35:55

and so, here to kind of fully kind of

35:57

recap all the the advancements we've

35:58

been pushing into the the stack. We've

36:00

been leaning in towards FP8 um wide

36:04

expert YDP um disintegrated prefill lots

36:07

of stuff at the routing and KV

36:09

offloading and management of just kind

36:10

of where things live. Especially things

36:12

like router re-place. The router

36:13

re-place is actually a really nasty

36:15

systems problem because it requires

36:17

tracking a lot of metadata per rollout

36:19

because you have this for every single

36:20

layer. Um so, it's like it's a big

36:23

multiple over just the like tokens and

36:25

log probs uh that you actually have to

36:26

store um because especially if you're

36:28

multi routing to multiple experts for

36:29

layer, um and so the storage concerns

36:33

for these, as well as like for

36:34

multi-modal, um if you have like images

36:36

that you need to store, like there's a

36:37

lot of stuff where you want to kind of

36:38

offload the heavier artifacts onto some

36:41

like object storage system or other file

36:43

system, um and not just have it floating

36:45

around in memory. Um and so we've

36:46

rebuilt a lot of our systems to support

36:48

this, as well as just really pushing uh

36:51

there's a lot of kind of like

36:53

special cases that you need to worry

36:54

about where it's like certain models

36:55

will want certain types of context

36:56

parallelism,

36:57

which have different considerations

36:59

about what you can do uh in terms of

37:01

like other aspects of the system.

37:03

Um and so we've just been trying to

37:04

really refine the recipes for making

37:07

sure that this works really well,

37:08

especially for models like JLM.

37:10

Um and we've done this all on top of a

37:11

Torch Titan base. I think a lot of

37:12

people are like, "Why do you use Torch

37:14

Titan and not Megatron?" Um and it's

37:16

because Torch Titan is just like really

37:17

easy to like hack, and Megatron is kind

37:19

of this monolith um that I think some

37:21

people will hack it, but it's I think uh

37:24

we just started with Torch Titan a long

37:25

time ago and kind of find the pieces we

37:28

want to bring in, and especially when a

37:30

new model comes out it's like, "Well,

37:30

you want to train on it." Or you want to

37:32

let's say you want to there's some new

37:33

paper that you read that like has some

37:34

new idea, you want to be able to like

37:35

make everything really hackable and

37:37

modular. And so like our our team is not

37:39

that big, like our whole research team

37:41

that maintains Prim Rels like less than

37:43

10 people,

37:44

um and the company is less than 40

37:46

people. Um and so there's a lot of work

37:49

that we want to make sure people can

37:50

parallelize, but also move quickly. Um

37:53

and so uh

37:54

this is all stuff that we've kind of

37:55

been able to to figure out over the past

37:57

few months as we've really been pushing

37:59

it for scale.

38:01

Um another thing that's I think fun is

38:03

algorithms. So I think a lot of what

38:05

researchers care about I think

38:06

researchers some people will love

38:08

thinking about the system problems of

38:09

scale. Um if you do talk to us, um but

38:11

if you don't, I think what a lot of

38:13

people want to spend their time thinking

38:14

about is algorithms in terms of the

38:16

on-policy distillation stuff or like OP

38:18

SD or uh if you saw the Echo paper, I

38:21

think that got a lot of people excited

38:22

of like thinking about how do you do

38:23

world modeling with RL um and folding

38:25

these in as well as just basic stuff

38:27

like SFT. Um there's also this Max RL

38:29

paper from a while back that was super

38:30

cool. And so we were just we were seeing

38:32

all these papers and we were like we

38:33

just want to do all of these. We want it

38:34

to be much easier to kind of like mix

38:36

and match these and to not have to like

38:38

add another if statement buried super

38:40

deep in the code um and like pipe a

38:42

bunch of stuff all the way through. And

38:43

so we kind of decompose things into the

38:45

loss which is like the thing that is

38:47

taking the gradient as well as the

38:49

algorithm which we say is the thing

38:50

that's kind of like preparing the data.

38:52

Um

38:53

and so we have different losses we can

38:54

kind of like pipe things to um in terms

38:56

of like the signal and the masking. Uh

38:58

and you can use these to kind of

38:59

assemble uh different algorithms. Um and

39:02

so now algorithm is like a class where

39:04

you have you can have a class that does

39:06

different things in terms of scoring or

39:07

uh the groups.

39:08

Um and and you'd kind of pick which loss

39:10

you want to target with it. Um

39:12

and so we support all of these all the

39:14

popular ones and you can kind of add

39:15

your own that look like adding in a

39:17

function to like assign advantages

39:19

uh to a rollout. Um

39:25

Uh

39:26

yeah. And so then then from your

39:28

algorithm then from your configs you can

39:29

just kind of say hey I want like this

39:30

algorithm and it's just going to pick

39:32

one from the registry. You can add your

39:33

own to the registry if you want. Um and

39:35

then this will be the algorithm used for

39:37

your training run. Um and you can do

39:39

this on a per environment basis if you

39:41

want. Um

39:43

but all of these algorithms that people

39:44

are looking at kind of fall into this

39:46

like table where there's questions about

39:48

like

39:49

what are your where are your rollouts

39:50

coming from? Are they coming from your

39:52

current policy model or are they coming

39:53

from some other source like a teacher?

39:55

Um and so any algorithm people call like

39:58

on policy or like slightly off policy in

40:00

terms of the async RL stuff. This is one

40:02

where like your your actor in RL sense

40:04

is going to be the model you're training

40:05

your policy. Your policy is your actor.

40:07

Um in other cases you're doing stuff

40:09

where your your actor is some other

40:11

model. So if you're doing context

40:13

distillation or you're doing SFT, like

40:15

these are ones where you are uh going to

40:18

have some other like model or

40:21

prompt potentially be the teacher that

40:23

is generating the data that you're going

40:24

to be training on. And so all of these

40:26

kind of fit within this family. Then the

40:27

other thing is the advantage. And so the

40:28

advantage is just like if you generalize

40:30

it to like a score, like you can now

40:32

call like just cross entropy loss or

40:34

negative log likelihood is like

40:35

everything has like advantage one.

40:37

Um there's a lot of all of the like OPD

40:40

algorithms can kind of you can kind of

40:42

think of it as the log probability is

40:43

being the advantage.

40:45

And with RL like your

40:47

your reward minus some baseline usually

40:50

group mean or something like it is your

40:51

advantage. And so we just kind of like

40:54

took a step back and looked at all these

40:55

things and we're like, "Oh, we can just

40:56

kind of like factor this all out pretty

40:57

nicely." And have almost everything else

41:00

be shared but these things get kind of

41:02

swapped and so your infra doesn't need

41:03

to change just because your loss

41:05

function needs to change.

41:09

And so for on policy distillation like

41:10

it's just plugging in with a different

41:12

loss target and it's like talking to the

41:14

teacher and being like, "Okay, the thing

41:16

I need to get is reference log probs

41:17

from some teacher and I already have my

41:19

like sequences. So I just need to send

41:21

these to a teacher as prefill in terms

41:24

of like a like you can get prefilled by

41:25

just like asking for like a one token

41:26

response. Now I I get my sequences and I

41:28

can just like stick these in as my

41:30

reference log probs." For self

41:31

distillation you can have like a hint

41:33

that you're kind of putting in before

41:34

your

41:36

your

41:37

like

41:38

this teacher. So you're not saying the

41:39

same prompt you're saying a different

41:40

prompt using renders to kind of pack

41:42

these together

41:44

and then you get back the same like

41:45

sequence that you can slice out into the

41:47

original form

41:49

so that you can kind of add these as

41:50

your reference log probs there.

41:51

You could do echo where you have like

41:53

two different algorithm components. You

41:55

could have one that is targeting the uh

41:58

the doing cross entropy on your

41:59

environment tokens while doing like an

42:01

RL objective on your your action tokens.

42:04

And you can mix and match these. So you

42:05

can have

42:07

like different

42:08

like you can decide which teacher is

42:10

going to be done your you're going to be

42:12

sampling from your student or your

42:13

teacher.

42:14

Um you can like decide which algorithm

42:16

you're going to be using on a per

42:17

environment basis.

42:18

Um

42:19

and you can have like this one we have

42:20

like both um a normal like OPD as well

42:24

as GRPO. Um

42:26

and then

42:28

Ah yes. So like this whole family we now

42:30

support within Primerl natively. Um and

42:34

one thing you may also have poked around

42:35

at or seen or tried is uh we have our

42:38

hosted training platform uh which means

42:40

you don't have to worry about GPUs at

42:41

all. Um And so this is just hosted

42:43

Primerl. The version we have that is

42:45

kind of the the broad uh self-serve

42:47

version today is multi-tenant Laura um

42:50

which is focused mostly on RL. Um what

42:52

we have coming quite soon that we'll be

42:54

rolling out is full fine-tuning um which

42:57

supports changing as much as you want in

42:59

Primerl in terms of the model and

43:00

everything else where we still give you

43:02

all the same abstractions for uh not

43:05

needing to think about the GPUs and kind

43:07

of auto-scaling and uh magic restarts

43:10

and uh having a dashboard where you can

43:11

log everything and have unified billing

43:13

for sandboxes and judges and all these

43:14

things. Um but also you get to develop

43:17

your environments on CPU on your laptop,

43:19

push them to the platform as environment

43:21

packages, and specify them in your

43:23

configs. Um and then this allows you to

43:27

um have a lot of flexibility in like

43:30

deciding when you want to go to

43:32

different levels of the stack. So in

43:34

many cases you don't actually want to

43:35

change anything in the trainer. You just

43:37

want to change your reward function. And

43:39

you can do this in environment space.

43:40

Maybe you want to configure the trainer

43:42

but not change it in which case you want

43:44

to kind of poke into like slightly more

43:46

complex knobs for uh things like

43:48

different loss functions or learning

43:49

rates. Um additionally you may want to

43:51

like actually go deeper into the trainer

43:53

and like uh do new algorithms at the

43:54

environment level at the algorithm class

43:56

level or at the loss function level or

43:58

maybe something else entirely that needs

43:59

going even deeper. And so we we support

44:02

all of this with the the um the full

44:03

fine-tuning as well. And so uh

44:05

multi-tenant Laura, if you're

44:06

unfamiliar, is a a very useful pattern

44:08

for allowing multiple people to do

44:10

training runs on the same architecture.

44:12

Um on the same uh model weight copy.

44:13

It's so you have a base model that's

44:15

kind of This is how all inference for

44:17

like token-based pricing usually works

44:19

as you're doing multi-tenant inference

44:20

where you have one big copy of Claude uh

44:22

that is just getting everyone's requests

44:23

and hitting a shared uh like KV pool uh

44:27

that is managed um

44:28

with This is really nice with Laura

44:30

because you can just have multiple

44:32

Lauras available that you can hot swap

44:33

and each person can have their own Laura

44:35

without needing to replace the base

44:36

model. Uh and you get one kind of

44:38

inference pool that serves everybody at

44:39

once even if they're using different

44:40

models. This allows you to do things

44:42

like token-based pricing uh and not need

44:43

to reserve GPUs. Um

44:46

For full fine-tuning, it kind of does

44:47

have to be GPU-based. Um but we've been

44:50

getting a lot more GPUs and so we have

44:51

the ability to let people run stuff on

44:54

their GPUs. Um and uh this allows going

44:57

like pretty like going to full parameter

44:59

training where especially for algorithms

45:01

like large-scale SFT or mid-training, um

45:04

you will want to kind of have more than

45:05

just a Laura adapter um as well as the

45:08

ability to kind of really customize at

45:10

at every layer you might want um like

45:11

how your algorithm is going to work. Um

45:14

but so we The multi-tenant Laura one is

45:16

already live and you can go use it

45:17

today.

45:18

Um full fine-tuning is coming out in the

45:19

next coming next couple weeks. Um the V1

45:22

stuff from before is already uh out uh

45:24

as a kind of a alpha feature that is a

45:27

stable release coming the next I don't

45:29

know, quite soon hopefully. Um and then

45:31

the cookbook I mentioned earlier is also

45:33

out. Um

45:35

And that's mostly what I want to talk

45:36

about. Um bit informal and I can People

45:39

have questions, I can

45:40

uh dig into any individual parts people

45:42

think are curious. We're also uh we're

45:43

hiring. Uh we are small team based in

45:46

mostly San Francisco um

45:48

becoming a much larger team quickly.

45:50

Um and uh we'll have some other exciting

45:53

news about that coming later this week

45:56

uh to

45:57

I don't know.

45:58

Uh demonstrate our commitment to scaling

46:00

the team. Um but yeah, we are we're

46:04

growing and we're

46:06

I think in a unique position of like we

46:08

do very open research work like all this

46:11

code is on GitHub if you want to go play

46:12

with it but we also like our a real

46:14

company that trains big models and makes

46:16

money.

46:17

And so yeah, we think we figured out a

46:19

good business model to do real open

46:21

research and it's been an incredible

46:23

journey thus far and would love to have

46:26

excited passionate talented people join

46:29

the team.

46:32

>> [applause]

46:49

[music]

Continue with YouTLDR

Analyze another video with Pro

Process a new video, search every timestamp, compare sources, and keep the result in your library.

Get Pro — $12/month30-day money-back guarantee

More transcripts

Explore other videos transcribed with YouTLDR.