Full Transcript

·YouTLDR

Ex-NVIDIA Engineer: Why AI Is About to Get 1000x Cheaper

1:23:371,433 summary words · ~7 min readEnglishBy Invest Like The BestTranscribed Aug 26, 2026
Analyze another video with Pro30-day money-back guarantee
Summary

The economics of AI inference will collapse by orders of magnitude as workloads shift from low-latency interactive chatbots to asynchronous, long-horizon background agents that prioritize throughput over speed.

Decoupling inference from real-time human attention allows operators to utilize cheap, non-NVIDIA silicon and intermittent power grids across decentralized 1MW data centers, transforming intelligence into a universally accessible commodity.

Section summaries

0:00-4:12

The Commodity Intelligence Thesis & Asynchronous Agents

watch

Neil introduces Sal Research as an ultra-low-cost token factory and agent sandbox platform designed for long-running virtual machines. He explains that while the existing market optimized for low-latency coding assistants (e.g., Cursor), the real frontier is long-horizon tasks running for hours or days where latency is irrelevant. By decoupling user interaction from token generation, token consumption becomes an autonomous background operation.

  • Reducing token costs by 10x unlocks entirely new product categories.
  • Agentic inference will transition from low-latency interactive queries to long-horizon background workloads.
  • Open-source model weights provide sovereign ownership and flexibility to run customized inference pipelines.

Establishes the core macro thesis of asynchronous background intelligence versus synchronous interactive chat.

4:12-12:36

Unbounded Background Workloads & Test-Time Compute

watch

The discussion covers test-time compute scaling, tracing capabilities from Claude 3.5 Opus to contemporary autonomous models. Neil highlights how background intelligence transforms deep research and cybersecurity into an automated 'proof of work' model, where spending compute to fuzz software finds jagged vulnerabilities. He argues that verifiable domains (formal math, code, scientific simulations) can be solved economically if intelligence is made abundant enough to waste tokens without guaranteed returns.

  • Background agent token volume will scale to 90% of all inference workloads.
  • Cybersecurity is evolving into automated 'proof of work' via heterogeneous model fuzzing.
  • Verifiable tasks (software, math, scientific modeling) scale directly with available token budgets.

Provides concrete operational examples of high-token background agent use cases in software and research.

12:36-23:06

NVIDIA Culture, Kernel Optimization & Throughput Trade-offs

watch

Drawing on his early engineering background at NVIDIA, Neil details the evolution of Tensor Cores since 2016 and the company's cultural obsession with 'speed of light' hardware utilization. He explains the fundamental computer science trade-off between latency and throughput: interactive systems require small batches with idle headroom, whereas background workloads permit massive batching that saturates GPU compute cores. NVLink enables low latency via tensor parallelism, but alternative schemes (pipeline and expert parallelism) offer superior FLOPs-per-dollar economics on non-NVIDIA chips.

  • NVIDIA's engineering ethos revolves around reaching 100% 'speed of light' hardware limits.
  • Latency optimization forces sublinear scaling and underutilizes GPU silicon.
  • High-throughput batching on GPUs functions like a bus rather than a private taxi, maximizing compute density per dollar.

Explains the foundational hardware mechanics and mathematical trade-offs between batch throughput and interactive latency.

23:06-35:42

Silicon Architectures: SRAM vs. DRAM and Transformer Bottlenecks

watch

Neil breaks down the memory hierarchy differences between on-chip SRAM (used by Cerebras and Groq for extreme memory bandwidth) and stacked DRAM/HBM (used by NVIDIA). He identifies the 'original sin' of transformer architectures: pairing compute-bound multi-layer perceptron (MLP) layers directly with memory-bound attention layers. While wafer-scale SRAM chips excel at MLP matrix multiplication, dynamic and unpredictable KV cache expansion necessitates large-capacity DRAM, pointing toward hybrid chip architectures.

  • SRAM offers petabytes per second of bandwidth but suffers from low physical die storage density.
  • DRAM/HBM provides high gigabyte capacity required for dynamic context windows (KV cache) at lower relative bandwidth.
  • Future inference clusters will likely hybridize wafer-scale SRAM engines with traditional high-capacity GPU memory.

Delivers an exceptional technical breakdown of memory bandwidth constraints, KV cache dynamics, and transformer hardware limits.

35:42-44:06

Post-Internet Data, RL Gyms & Rack-Scale Programming

watch

Addressing data scaling, Neil states that the public internet's 30–300 trillion tokens of high-quality text have been exhausted. The next paradigm is recursive self-improvement inside automated reinforcement learning (RL) gyms on verifiable tasks. He returns to infrastructure scaling, explaining how modern GPU programming is moving from individual accelerator kernels toward full rack-scale orchestration (such as NVIDIA's NVL72 liquid-cooled Grace Blackwell systems).

  • Unconditioned human internet feedback has declining utility for frontier model post-training.
  • Verifiable RL environments with automated self-grading provide the engine for continuous synthetic data scaling.
  • Hardware orchestration has shifted from single-GPU kernel optimization to whole-rack (NVL72) cluster programming.

Connects post-training data dynamics directly to physical rack-level computing shifts.

44:06-52:30

Hardware Arbitrage, Alternative Silicon & Market Hype

watch

The discussion covers the market dynamics of top-tier GPUs versus secondary accelerators from AMD, Tenstorrent, or startup ASICs. Neil describes an arbitrage model: while competitors chase scarce Blackwell allocations, huge value exists in programming alternative chips that vendors under-optimize. Comparing today's CapEx boom to the 2000 telecom fiber bubble, he argues current token spend is non-speculative because inference tokens are consumed immediately rather than hoarded.

  • There are no bad chips, only bad pricing; unoptimized silicon provides software alpha.
  • Inference spend is structurally non-speculative because tokens are consumed in real time.
  • NVIDIA strategically allocates high-end chips to prevent single-customer power concentration.

Analyzes the semiconductor market, supply allocation dynamics, and CapEx sustainability.

52:30-1:03:00

The Scavenger Strategy: 1MW Micro Data Centers & Intermittent Power

watch

Neil introduces the physical infrastructure strategy of running distributed 1MW micro-data centers instead of fighting over scarce 100MW+ or gigawatt-scale sites. Liquid cooling allows a megawatt of compute to fit in eight refrigerator-sized racks. By removing diesel backup generators, redundant fiber links, and strict SLAs, Sal can buy cheap intermittent solar/wind power, tolerating 80–95% uptime by shifting asynchronous agent workloads dynamically across facilities when outages occur.

  • A 1MW data center can be compressed into approximately eight high-density liquid-cooled racks.
  • Eliminating power and network redundancy slashes facility CapEx and OpEx.
  • Asynchronous workloads easily tolerate 95% data center uptime via distributed software failover.

Presents a novel blueprint for reducing data center energy and construction costs.

1:03:00-1:13:30

KV Cache Inefficiencies, Fab Realities & Open vs. Closed Source

watch

Neil reviews global compute inefficiencies, highlighting that attention KV caches remain vastly uncompressed. On chip manufacturing, he argues fab yield margins and process corners can be relaxed for inference workloads. Regarding frontier labs, he views open-source distillation as inevitable through public artifact generation (e.g., GitHub repos created with AI tools), predicting that closed-source proprietary leads will remain short-lived 3-to-6 month windows.

  • KV cache entropy is uncompressed by one to two orders of magnitude in current architectures.
  • Process corner variances at fabs could be relaxed to utilize discounted, sub-tier silicon.
  • Model distillation is impossible to halt due to widespread public diffusion of AI-generated artifacts.

Covers high-level fab optimization, memory entropy, and the structural dynamics between open and closed models.

1:13:30-1:21:54

Trillion-Token Days, Supply Chain Bottlenecks & NVIDIA's Moat

watch

Neil envisions a world where users consume a trillion tokens per day for complex proactive automation. He outlines the critical bottlenecks hardware entrepreneurs must navigate: TSMC wafer supply, HBM packaging capacity, advanced packaging, and energy access. Finally, he explains why NVIDIA focuses on building an ecosystem of profitable infrastructure partners (like CoreWeave) rather than selling direct end-user tokens, concluding with reflections on mastering the entire engineering stack.

  • Long-term agent adoption requires dropping trillion-token job costs to tens of thousands of dollars.
  • The primary hard bottleneck in AI silicon scaling is High Bandwidth Memory (HBM) packaging capacity.
  • NVIDIA avoids competing with its infrastructure buyers to preserve a broad, competing ecosystem of customers.

Synthesizes long-term cost targets, semiconductor supply choke points, and ecosystem strategy.

Key points

  • The Latency vs. Throughput Dichotomy in Agentic Inference — Chatbots enforce low latency at the expense of GPU compute efficiency, but autonomous background agents running for hours only need 10 tokens per second. Running wide and slow allows full hardware batching and vastly better token economics without needing low-latency interconnects like NVLink.
  • The Architectural Split: On-Chip SRAM vs. Off-Chip DRAM — Transformers contain compute-bound feed-forward networks (MLP layers) and memory-bound attention layers. Ultra-fast wafer-scale SRAM architectures (e.g., Cerebras) excel at raw compute bandwidth for weights but struggle with dynamically expanding KV cache capacities, necessitating heterogeneous hybrid clusters with traditional high-capacity DRAM/HBM systems.
  • The Scavenger Strategy for Compute and Energy Infrastructure — Instead of competing for scarce 100MW+ tier-4 data centers with triple redundancy, inference workloads can be distributed across decentralized 1MW micro-sites that tolerate 80–95% uptime and run on cheap, intermittent renewable power via resilient control planes.
  • Shift from Internet Data Ingestion to RL Verification Gyms — Having exhausted the ~30–300 trillion tokens of public internet text, frontier model advancement relies on recursive self-improvement inside verifiable reinforcement learning environments (e.g., formal code execution and mathematics) where agents generate and grade their own synthetic task data.
My job is to make the tokens as cheap as humanly possible. I will achieve that and I will do it through every layer in the stack available to me. Neil
We think that whenever you make something 10 times cheaper, it's a new product category and uh we aspire to do that for tokens. Neil

AI-generated from the transcript. May contain errors.

0:00

My job is to make the tokens as cheap as

0:01

humanly possible. I will achieve that

0:03

and I will do it through every layer in

0:05

the stack available to me. I love the

0:06

supply side levers. I will use every

0:07

chip. I'll use every source of power and

0:09

I will use every piece of land in the

0:10

United States that's, you know, suitable

0:12

for this. We still treat the agent as a

0:15

person that is expensive to consult and

0:16

you should ask them when you have a hard

0:18

question. That's not the way to think

0:20

about intelligence. It's incredible that

0:21

the machine can think and we should

0:22

[music] try to get that into as many

0:24

hands as as many people as possible.

0:32

>> [music]

0:37

[music]

0:38

>> I think it's important early in these

0:39

conversations to just say the thing like

0:42

literally what you're building and what

0:43

it does today. So maybe just orient us

0:45

there with a with a brief description

0:47

like literally what the system is that

0:49

you're building and why it should exist.

0:51

>> Sal research is a token factory. We have

0:53

an API where anyone can send us requests

0:55

where they can use large language

0:57

models, open source large language

0:58

models for any task they want. Um, we

1:01

will serve those tokens to them at a

1:02

price that is unbeatable in the market.

1:05

We also support their ability to build

1:07

agents on top of this. We host what we

1:09

call sandboxes, which are longunning

1:11

agent virtual machines hosted in the

1:13

cloud that are designed for agents that

1:15

run for hours, days, or weeks. And so

1:17

you should think about you as a peer

1:20

company to others that serve different

1:22

kinds of inference. You're serving one

1:23

specific kind of inference and your goal

1:25

is to be the absolute cheapest provider

1:28

and enabler of a certain kind of use of

1:30

intelligence.

1:30

>> Exactly. The theme of our company is

1:32

abundance. We want to deliver this new

1:34

commodity of intelligence to as many

1:37

people as possible at a at a cost that

1:39

is sustainable for almost every

1:41

industry. We think that whenever you

1:42

make something 10 times cheaper, it's a

1:43

new product category and uh we aspire to

1:45

do that for tokens. We think it's so

1:47

profound that the machine can think and

1:49

now our job is to make as many machines

1:50

as possible in the world work towards

1:52

thinking.

1:53

>> So if you think about uh the theme of

1:54

the day being token costs is token cost

1:57

the right way to think about this like

1:59

is there some other way you'd put it

2:01

>> to start with? Absolutely. Token cost

2:02

today my north star is I want to have

2:04

the lowest cost per token in the

2:05

industry and do that by a mile. I don't

2:07

think tokens are the final unit of uh of

2:09

work or intelligence but they are what

2:11

we use today and so it's very

2:13

straightforward. I think after tokens

2:14

you start to move more towards um more

2:17

outcomes which is like a vague

2:18

direction. Uh you can imagine for

2:20

example today when you consume tokens

2:22

through an agent you don't actually

2:24

control how many tokens the agent

2:26

reasons for. It can reason for a certain

2:28

amount of time or it can call a certain

2:29

number of tools and increasingly I think

2:31

we will have agents do some unit of work

2:34

take as many shots on goal as they can

2:36

and however many tokens they use to get

2:37

there is going to be kind of a dependent

2:40

variable depending on the task.

2:42

So you think about like agents that

2:44

selfadminister a token budget as opposed

2:46

to a company setting a budget for how

2:48

many tokens engineers can spend per

2:49

month.

2:50

>> Why is there an opportunity that you can

2:53

tackle it? It seems like the entire

2:54

world is oriented around more better

2:57

faster cheaper tokens. Right now it

2:59

seems like the world is trying to solve

3:01

this problem very aggressively.

3:03

>> What was the unique opening that that

3:05

you saw that's maybe the market's not

3:06

being efficient in its attempt to tackle

3:08

this? So I think there's two things that

3:10

are tailwinds for our company. One is

3:12

got to be the rise of open source. I had

3:14

to talk about that first. I think we are

3:16

starting to see an increasing number of

3:18

our customers and the broader market

3:20

care about owning intelligence. They

3:21

they want to have control sovereignty

3:23

over the thing that they depend on. Uh

3:25

and so that created a much more robust

3:27

market for customized models or even

3:28

just like these vanilla open source

3:30

models that no one can ever take away

3:31

from you. You always have the weights.

3:32

You always have the right to deploy them

3:34

however you like. uh in that world

3:36

there's been a reasonably robust market

3:38

for the past couple years uh serving

3:40

these models at large scale. The

3:42

challenge is all those companies, you

3:44

could take your pick, base 10, fireworks

3:46

together, they all focus on low latency

3:48

inference, and they were pulled in that

3:50

direction by one very important

3:51

customer, uh, cursor. And

3:53

[clears throat]

3:54

I think that that was the right choice

3:56

about a year ago, and as of six months

3:58

ago, it started to look like maybe low

4:00

latency wasn't the only thing you wanted

4:01

from an agent. You wanted more

4:02

persistence, more long horizon tasks.

4:05

And now, it's to me very obvious that

4:07

the future of agentic inference is long

4:10

horizon tasks. You're going to run the

4:12

machine for hours or days at a time. It

4:14

doesn't matter if it spits out tokens at

4:16

100 tokens per second. Maybe 10 is just

4:18

fine. That comes with corresponding

4:20

advantages and efficiency.

4:22

>> Why are you so confident in that? It's

4:24

it like to me it seems like I want

4:26

everything as fast as possible.

4:27

>> When you're waiting on it, you

4:28

absolutely deserve the fastest answer

4:30

possible. My trick is I don't want you

4:31

to be waiting on it. I want it to be

4:33

proactive. I want it to be in the

4:34

background. One way to say it is like

4:35

the best latency is no latency at all.

4:37

When you wake up in the morning, the

4:38

work's already been done overnight. you

4:40

didn't even have to ask for it. Uh

4:41

that's the dream. We're not quite there

4:43

yet. But more importantly, I think the

4:45

more you're in the loop as you prompt

4:47

agents and wait for a response, in fact,

4:49

you're the bottleneck uh in helping the

4:51

in having the agent do more or less

4:52

work. What we'd like is the agent to

4:54

operate on more human time scales. You

4:56

don't manage your colleagues every 5

4:58

minutes. You ask them to do a high level

4:59

task and you come back and check in

5:01

maybe every day, but more likely once a

5:03

week. And that to me is the future of

5:05

human agent collaboration, more like

5:07

human time skills. say more about the

5:09

early indications that this is happening

5:11

and therefore you should be building

5:12

this company.

5:13

>> Well, so the first and most important

5:14

thing is the idea of test time compute

5:16

scaling. Uh the idea that you can give

5:17

an agent more time and it will give you

5:19

a better answer. So that was theorized

5:21

about 2 years ago now and uh but it

5:23

wasn't really something that we could

5:25

actually bet on until I would say late

5:27

last year with Opus 45. Opus45 was the

5:29

first agent that was at all suitable for

5:31

longer horizon tasks and you know it was

5:34

pretty mediocre and when it first came

5:35

out but you look at the more recent

5:36

models and what we've done on open

5:38

source as well and you see that agents

5:40

are capable of running for an hour at a

5:43

time. I wouldn't say it's days but

5:44

definitely an hour is quite suitable

5:46

today. And so just seeing that like

5:48

average turn or task length get longer

5:50

and longer uh it doesn't take many

5:52

points to have you kind of draw out the

5:54

exponential and see that agents are

5:56

worth running for longer periods. What

5:58

what do you think will be the market

5:59

share of longrunning agents in 3 years

6:02

or something like this?

6:02

>> You know, I love this market because

6:04

it's unbounded. There's no human in the

6:05

loop. So, you can consume as many tokens

6:06

as you like in the background. Uh versus

6:09

human attention span. If you tell me to

6:10

consume 10x as many tokens at codeex or

6:13

at cloud code, I'm actually not sure if

6:14

I can anymore. I'm already in a loop and

6:17

locked in coding for most of the the day

6:18

that I'm at the laptop. What is undowned

6:21

is how many tokens can be consumed in

6:23

the background or proactively. So

6:24

longterm, I think, you know, we're going

6:26

to end this year at maybe 50/50

6:27

background and uh and real-time

6:29

workloads, but I see this going to 9010

6:31

in favor of background.

6:32

>> What are the sorts of things like what

6:34

are your favorite examples of something

6:36

that gets accomplished much better as a

6:38

background task than as a human in the

6:40

lip task?

6:40

>> Most deep research, most questions where

6:43

you want to have a definitive answer

6:45

over not 100 sources, not a thousand

6:48

sources, but 10,000 sources or more. If

6:50

you want to build an authoritative index

6:52

of information like for example one of

6:54

our customers parallel web systems seeks

6:55

to do. They want to build an index over

6:57

the whole internet and they want to

6:58

monitor the internet in real time for

7:00

changes. That is the kind of crazy

7:03

exabyte scale task that you need a very

7:05

different kind of intelligence or scale

7:07

of intelligence to achieve. Deep

7:08

research is a top category for us and

7:10

then increasingly we see cyber security

7:12

following this direction. If you think

7:14

about, yes, there's so much code you can

7:16

generate, but there's uh exponentially

7:18

more ways to break that same code than

7:20

it is to generate that code. And there

7:22

are some great customers out there who

7:23

are working very hard to uh find agents

7:25

that can break any piece of software and

7:27

proactively patch them. So when Fable

7:29

first came out, for example, or Mythos

7:31

first came out, basically there was this

7:33

push in the cyber security community to

7:34

run Fable against every line of code

7:37

we've ever written and look for bugs in

7:39

20 different ways. uh meaning you're

7:41

looking for both memory errors, you're

7:43

looking for business logic errors and

7:45

looking for like network

7:45

vulnerabilities, all these things. And

7:47

these are all actually things that you

7:49

would write specialized agents for. You

7:50

wouldn't just have Fable look at the

7:51

source code once, you'd have it actually

7:53

set up environments where you can pen

7:54

pentest these applications. And at some

7:58

point, people started to make this joke

7:59

that security has become proof of work.

8:01

When you want secure software, it's

8:03

really a question of how many dollars

8:05

did you spend on anthropics APIs trying

8:07

to break into your software. uh that is

8:09

the best indication for how secure it is

8:11

because that's the best tool in the

8:13

world. And increasingly we found that

8:14

the frontier of intelligence here is

8:15

quite jagged. It's not the case that

8:16

Fable finds a supererset of all bugs in

8:19

software. You would find some bugs with

8:21

a very small model that you don't find

8:22

with the large model. Uh you'd find some

8:24

bugs with Haiku that you would find with

8:26

Fable and vice versa. So it encouraged

8:28

this very diverse approach to sampling

8:30

and trying to build cyber security

8:32

agents that break software autonomously

8:34

such that you can patch them. If you

8:35

were to get sort of like speculative and

8:37

imaginative about the sorts of things

8:40

that longunning very cheap very

8:42

longunning agents can enable. We talked

8:44

about some very practical examples deep

8:46

research um cyber security etc. But if

8:48

you if you get a little bit dreamier

8:50

about the use cases new product category

8:54

that this sort of inference will unlock

8:56

and I guess the question is just like so

8:57

what like what if you're maximally

9:00

successful dream a little bit about what

9:02

that might enable.

9:03

>> Yeah absolutely. So I think for

9:05

individual users what I'm excited about

9:07

most is this idea of proactive

9:09

intelligent agents. Um you can imagine

9:11

Siri that is running in the background

9:12

all the time to understand what's all

9:14

the emails you received in a day, all

9:16

the text messages you receive in a day

9:17

and and has a much more encyclopedic

9:19

view of your life and how to be helpful

9:21

in that life. Right now there's still

9:22

point solutions and so you have to you

9:24

end up doing a lot of prompting. Siri is

9:26

not very proactive. That's something we

9:27

can fix with abundant abundant

9:29

inference. If you trust uh the machine

9:31

enough that it's reliable and also

9:33

trustworthy as in private um you might

9:35

even imagine the machine can understand

9:37

how you interact with it and proactively

9:39

surface your next action whenever you

9:41

open your phone. Can we build a good

9:42

model of what you're going to do next?

9:44

>> My estimation is yes, we totally can.

9:46

>> And the key to that is incredibly cheap

9:48

intelligence.

9:48

>> You have to be willing to spend tokens

9:50

without any promise of return. That's

9:52

the unlock. The long lens view to take

9:55

on this is that we have we have a form

9:58

of intelligence that can tackle any

10:00

verifiable problem. Any verifiable

10:03

problem means most software. It means a

10:05

lot of formal like math proofs and

10:08

similar. And it could also mean

10:10

scientific discovery. These are all

10:13

relatively verifiable problems. And all

10:15

those things currently have a dollar

10:17

cost attached to them essentially.

10:18

That's a hidden one. It's like how many

10:20

tokens could you possibly harness to

10:21

make this work? And we have actually

10:25

started to bring it within view a dollar

10:28

cost for these long horizon tasks that

10:30

is reasonable. It's not millions, it's

10:32

thousands and maybe it could be hundreds

10:34

or even tens of dollars in the near

10:35

future to have a definitive answer to

10:38

any scientific question to any research

10:40

problem. RAMP is the only platform built

10:43

to make your finance team leaner,

10:44

faster, and better, saving businesses 5%

10:46

annually on average, so you can stay

10:48

focused on growth. RAM customers grow

10:50

revenue 3.2 two times faster than the

10:52

average American business. Visa,

10:54

Verscell, Kerser, Stripe, Notion, 11

10:56

Lab, Shopify, and 70,000 other

10:58

businesses all now run on RAMP. Mine

11:01

does too, and so should yours. Learn

11:02

more at ramp.com/invest.

11:06

OpenAI, Cursor, Anthropic, Perplexity,

11:08

and Verscell all have something in

11:09

common. They all use work OS. To achieve

11:12

enterprise adoption at scale, you have

11:14

to deliver on core capabilities like

11:16

SSO, skim, arbback, and audit logs.

11:19

Instead of spending months building

11:20

these missionritical capabilities

11:21

yourself, you can just use Work OS APIs

11:24

to gain all of them on day zero. That's

11:26

why so many of the top AI teams you hear

11:28

about already run on Work OS. Work OS is

11:31

the fastest way to become enterprise

11:32

ready and stay focused on what matters

11:34

most, your product. Visit works.com to

11:37

get started. Felix by Rogo is a personal

11:39

finance agent that turns a single prompt

11:41

into [music] finished client ready work

11:43

using your firm's own templates,

11:45

context, and standards. Send Felix an

11:47

email like, "Take [music] these comments

11:48

and turn them for me." Or, "Udate my

11:50

tracker with the context of these

11:52

emails." And Felix sends back finished

11:53

[music] PowerPoint decks, Excel models,

11:55

and sourced research. Felix works the

11:57

way your team already does, delivering

11:59

[music] work quickly and accurately

12:00

around the clock. Learn more at

12:02

robo.ai/felix.

12:04

[music]

12:04

>> And so, if we dream about that future,

12:06

we're we then become limited just by the

12:09

questions that people can ask.

12:11

Basically,

12:12

>> pretty much the questions we can ask. uh

12:14

the models are on the cusp of basically

12:16

taking even a high level question and

12:18

chasing it down every possible follow-up

12:20

you can have the model essentially take

12:22

that on its own

12:23

>> and the question is what is your token

12:24

budget

12:25

>> and we will solve the token budget

12:26

problem

12:27

>> what about non-verifiable

12:29

tasks

12:29

>> I put basically the entire category of

12:31

human taste into that category we have

12:34

not solved human taste yet and I don't

12:35

know that it fundamentally can be I'm

12:37

excited to be surprised here but um we

12:40

are focused on very quantitative uh

12:42

problems we leave the quality of writing

12:43

we the uh the beauty of art to to

12:46

people.

12:47

>> All right. Now, let's talk about the uh

12:48

the very clever stack of solutions that

12:51

you hope to build

12:52

>> ultimately to have this giant token

12:54

factory, extremely lowcost intelligence

12:56

supplier of extremely lowcost

12:58

intelligence.

12:59

>> I think you think about this in terms of

13:00

level software, hardware, uh and power.

13:03

>> Talk through what your master plan is to

13:05

approach this challenge that's so

13:07

different from what others are thinking

13:09

about doing.

13:09

>> You know, we always have to start with

13:10

software. you know where is the

13:12

opportunity on today's chips with

13:14

today's data centers to improve

13:15

efficiency and the first thing we did

13:17

was we tried to build the entire LM

13:19

software stack around peak GPU

13:21

efficiency meaning we're using Nvidia

13:23

GPUs we wanted to squeeze out more

13:25

tokens from the same chip than anyone

13:26

else in the world and that starts with

13:28

the lowest level of programming kernels

13:31

it's actually my background I spent my

13:32

whole life actually uh my whole

13:34

professional life working on GPUs and

13:35

kernels in Nvidia was my first job while

13:37

I was in college and uh I got to see how

13:39

the tensor cores got to earn their right

13:42

to be on the chip. This is back in 2016.

13:44

>> Just describe what that means for for

13:46

the lay person.

13:47

>> So, okay, tensor core is a specialized

13:48

unit on the GPU that accelerates matrix

13:51

multiplication.

13:52

>> Simple as that. There's been a long

13:53

history of how we evolved at tensor core

13:55

over time that we'll get into.

13:56

>> And why is matrix multiplication so

13:58

important?

13:58

>> That's a great question. I actually I

14:00

cannot say that there is a divine truth

14:01

of the inverse that explains why matrix

14:03

multiplies seem to be the atomic unit of

14:05

computation. But uh one way I've heard

14:07

it described to me is well it's a really

14:09

succinct way to mix two blocks of

14:11

numbers together and have them interact

14:13

in some interesting way. That's as much

14:15

as I can say about it. It is really

14:16

convenient that linear algebra turns out

14:18

to be a very compact representation of

14:19

arbitrary relationships in data. So

14:21

Nvidia great graphics company obviously

14:24

has had market share dominance in GPUs

14:26

and and gaming graphics for quite some

14:28

time. And then starting in like the

14:31

mid2010s they started to actually start

14:33

these like skunk works projects to make

14:35

the graphics processor more suitable for

14:37

machine learning tasks that they were

14:38

tracking. I remember actually reading

14:40

some of the like lab notebooks of some

14:41

of my managers when I was at Nvidia.

14:43

they would visit these small ML

14:45

conferences like ICML or NURPS at the

14:47

time and they would just take note of

14:49

these papers like oh this deep learning

14:51

thing seems to be catching on and what's

14:53

really interesting is that these grad

14:54

students are using gaming Nvidia GPUs in

14:57

order to train their large models we

14:59

should double click on this and figure

15:00

out what's going on here and by 2015

15:03

2016 at least Jensen had the conviction

15:06

to to kind of double down on hey this

15:08

usage of our models is only of our chips

15:10

is only going to grow let's start

15:12

allocating more and more precious

15:13

silicon die area to this capability that

15:16

seems to be emerging. Let's put the

15:18

first version of tensor cores on the

15:19

chip. So, we're talking about, you know,

15:21

taking this gaming chip which is

15:23

designed for painting pixels on a screen

15:25

and adapting it to do metric multiplies

15:28

and it was early and you would be

15:31

competing against the graphics teams

15:32

essentially when you ask for more

15:34

silicon area and any chip company.

15:36

There's always competition for that. It

15:37

is something that that the designers

15:38

guard so carefully. you don't ever want

15:41

to invest in the wrong technology

15:43

because that's opportunity cost that you

15:44

could have allocated to some other

15:46

functionality. And so we we kind of like

15:48

fought and tooth and nail and got just a

15:51

tiny bit of diary maybe like 5 10%

15:53

something like that for the first

15:54

generation of these chips uh to get some

15:57

some amount of acceleration for basic

16:01

convolutions which were the fundamental

16:03

operation for computer vision models in

16:04

the day. Uh, and then we had a software

16:06

team that was trying to squeeze all the

16:08

performance we could out of the chip.

16:11

And I think on that software team, which

16:12

is where I work, that's what actually

16:13

taught me the most about um, just the

16:15

ethos that Nvidia has around they have

16:17

this term called speed of light. They

16:19

always chase the speed of light for any

16:20

piece of hardware that they make. It is

16:22

so ingrained in every engineer's mind

16:24

that if the machine can do it, we're

16:26

going to push the machine to the

16:27

frontier until it does what we think is.

16:29

>> And the speed of light is the edge of

16:30

what's possible.

16:31

>> The speed of light is the edge of what's

16:32

possible. Exactly. uh if we think the

16:33

chip can run at this frequency and

16:35

produce this many multipliers per cycle,

16:36

we're going to get there. We're going to

16:38

break every bottleneck and get to that

16:40

peak level of performance. And so to

16:41

this day, I tell all my engineers like

16:43

we're chasing 100% speed of light. I

16:45

don't care about relative numbers versus

16:46

the competition. I only care about

16:47

absolute numbers. Uh what are we able to

16:49

do on the chip and how do we achieve

16:51

that?

16:51

>> Before we leave that chapter of your

16:53

time at NVIDIA, anything else beyond

16:54

that cultural touch point that really

16:56

like changed the way you think about

16:58

things or that stood out the most about

17:00

how the business ran back then or its

17:02

culture? I have a ton of stories about

17:03

Nvidia. We can I can tell you a few of

17:04

them. Um, one of my favorites is that on

17:06

the tenure side, a lot of people I

17:08

worked with in Nvidia in 2015, 2016 are

17:10

still there today. That company has

17:12

incredible retention and these are the

17:13

best engineers uh, frankly on the

17:15

silicon side at least I've worked with

17:17

in my whole career. They're extremely

17:19

extremely motivated and passionate.

17:20

They've believed in parallel computing

17:22

as a concept through its various

17:23

incarnations and have loved seeing the

17:25

chip evolve. This is their life's work

17:27

and they're extremely extremely

17:29

competent in that direction. They're

17:30

also a very frugal company. Nvidia and

17:32

all, I guess all the Silicon Valley

17:34

companies after 2008, they had some

17:36

cutbacks and like perks. So, no free

17:38

lunch. Uh, for example, Nvidia took it

17:40

one step further. There was no free milk

17:42

in the fridge. So, if you wanted to

17:43

drink coffee at Nvidia and you wanted

17:45

some milk, you actually had to chip in a

17:47

dollar every month to the milk club and

17:49

the milk club would stock Costco milk in

17:51

the fridge. And I remember that

17:53

distinctly. We don't do that at sale,

17:55

but uh

17:56

>> it's a it's a frugality that permeates

17:58

the company. And so coming out of this

17:59

time there, you get this experience of

18:02

what it's like to develop more efficient

18:04

usage of the underlying hardware through

18:05

software.

18:06

>> Yes.

18:06

>> And so so link that to, you know,

18:08

today's environment.

18:10

>> Yeah, absolutely. So, so I think um the

18:12

GPU is fundamentally a throughput

18:13

machine. The GPU is happiest when you

18:14

give it a lot of work to do and let it

18:16

chew through that work at peak

18:18

utilization of its compute units. But

18:19

that's actually not the way that we've

18:21

taken AI in the last couple years. We've

18:24

really pushed AI to be an interactive

18:25

chatbot tool is the most common form of

18:28

AI usage today. And in that world, you

18:31

care a lot about actually spitting

18:32

answers out to the to the person at the

18:34

keyboard as quickly as possible. To your

18:35

point about don't make the user wait, I

18:37

want things as fast as possible. And so

18:38

that's actually quite interesting for

18:40

the GPU. It's very difficult to put the

18:42

GPU in its happy path of being fully

18:45

compute utilized when you're trying to

18:47

spit out tokens quickly. There's a

18:48

fundamental trade-off on the GPU between

18:50

being uh throughput oriented or latency

18:53

optimized and everyone has chosen

18:54

latency optimization because the shape

18:57

of usage was chatbot oriented. I believe

19:00

that's the most profound change we're

19:01

going to see in the next year. We're

19:02

going to move away from chatbots to more

19:04

proactive or background agents. And in

19:06

that world, it makes a lot more sense to

19:08

build a stack around throughput.

19:10

>> Can you explain technically why the

19:12

trade-off between throughput and latency

19:14

is unbreakable? Why can't we have both

19:16

from the same hardware? It's quite

19:17

foundational in almost every system that

19:19

you could ever possibly look at. There's

19:21

always a trade-off between getting a

19:23

small amount of data through the system

19:24

as quickly as possible and leaving a lot

19:26

of buffer uh room for that or trying to

19:30

run wide and slow like narrow and fast

19:32

or wide and slow is like a classic

19:33

trade-off in all computer science. But

19:35

for GPU specifically, I think there's

19:36

one thing to focus on which is there's

19:38

this concept of like batching on the

19:39

GPU. We want to group many users work

19:42

together into a batch that we can uh run

19:45

all at once on the GPU. That's the

19:47

parallel processing of the GPU. We'd

19:48

like to have a lot of parallel work to

19:50

do. The thing is though, you're doing

19:52

net more work when you run a large batch

19:54

of compute together. And so you might be

19:56

filling all the units, but every step

19:59

along the way as you carry a a batch of

20:01

work through the GPU, there's more work

20:04

to be done. And so any individual token

20:06

or any individual user's request in that

20:08

batch, it's going to spend a longer time

20:10

on the GPU being carried with other

20:12

people's traffic. Maybe the way to say

20:13

it is um you know, if you want to get

20:15

downtown and SF, you can take the bus or

20:17

you can take a private transit. And the

20:19

private transit is going to have its own

20:21

direct path as the crow flies or you

20:23

know, using exactly the roads that you

20:24

want from point A to point B. a bus,

20:26

it's going to have to serve many more

20:27

people and it it has to fundamentally uh

20:30

do something that works for everyone

20:32

>> and so it takes a slower path and it

20:33

stops and and waits for other people to

20:35

get on and off. I think the bus versus

20:37

car analogy is pretty accurate

20:38

>> and it's a great analogy and so step one

20:41

for what you're trying to do is like

20:43

create the best possible bus on top of

20:45

Nvidia GPUs. Like that's step one of

20:47

your optimization.

20:49

>> That's exactly right. It means we

20:50

explore things like different

20:51

parallelism schemes. Maybe that's

20:52

another example I can give you is um

20:54

with Nvidia GPUs, one of the things that

20:56

they've really innovated on and done a

20:57

great job with is the NVLink uh

20:59

interconnect between GPUs. And in fact,

21:01

that NVLink system is so good that you

21:04

can if you have a large matrix multiply

21:06

that you want to perform faster. You can

21:09

actually cut that matrix multiply in

21:10

half and shard it [clears throat] across

21:12

two or more up to eight, let's say,

21:14

Nvidia GPUs and have them all work on

21:16

pieces of that larger matrix multiply

21:19

>> and have them connect their results

21:21

together at the end. Reduce their

21:22

results back together at the end. And

21:24

this is a great great way to cut the

21:27

minimum latency of of an operation.

21:29

You're each GPU is now doing 1/8 as much

21:31

work, let's say, uh, and therefore it

21:33

can finish faster but not eight times

21:35

faster. It's sublinear scaling. You'll

21:37

use eight times more hardware, but you

21:39

won't get eight times the speed. You

21:41

might get like four to fivex the speed.

21:43

You're not going to get strong scaling.

21:44

And this is because of communication

21:45

overhead. It's because every GPU is

21:47

going to be a little bit less efficient

21:48

working on a smaller tile of work than a

21:50

larger tile of work. And so, it's the

21:52

only way to speed up if you want the

21:54

minimum latency possible. You can do

21:55

that, but it is not the choice I would

21:57

make. For example, I would prefer to use

21:59

a different parallelism scheme like

22:01

expert parallelism or pipeline

22:02

parallelism. And we may do interesting

22:05

things to overlap and hide the

22:07

communication latency in a way that you

22:09

would have less ability to do that for a

22:11

low latency server.

22:12

>> So is the right way to think about

22:12

NVLink as a technology which improves

22:15

latency performance? Yes.

22:16

>> And only latency performance

22:18

>> which will segue into the next segment

22:20

of what we you know do differently as a

22:22

company. But yes, NVLink is mandatory I

22:24

would say for low latency inference.

22:26

>> So Nvidia is excellent at low latency

22:28

inference. And I'm telling you that we

22:29

don't really care that much about low

22:30

latency inference. So where does that

22:32

leave us? Well, I think I'm not holding

22:34

my breath for other companies broadly to

22:36

figure out NVLink quickly. It's a

22:39

challenging technology to figure out.

22:40

It's hard to scale. It's hard to

22:41

productionize. And so, if I do have some

22:44

other vendors chip and it is good at the

22:46

foundational compute components, it can

22:48

still do metric multiplies really well.

22:49

It just can't communicate those results

22:51

across its peers quickly. Well, maybe

22:53

there's a room for that other chip in my

22:56

stack as a really really good compute

22:59

per dollar option. And that's what I

23:01

actually optimize for in most cases is

23:03

how many flops does this chip have and

23:05

how much is it going to cost me per hour

23:06

to operate to own and operate. Uh and so

23:09

there are other chips that definitely

23:10

rank higher than Nvidia on flops per

23:12

dollar, but they may not have as much

23:14

interconnect. And so it's my job to

23:15

figure out what parallelism scheme am I

23:17

going to use that's going to make this

23:19

chip suitable for inference. It's not

23:21

going to be tensor parallelism. Nvidia

23:23

is basically mandatory for that. But

23:25

other techniques may work well for me.

23:27

So before we leave the latency part of

23:28

the story, can you comment on companies

23:31

like Cerebras or others that can perform

23:33

incredibly fast operations? I'm curious

23:36

like what you think about those

23:37

approaches, those companies, what might

23:38

happen in the future. What is your

23:40

prediction for the future of very low

23:42

latency focused hardware? Cerebrus Grock

23:45

uh and a couple others that are coming

23:46

out of stealth now I think have made a

23:48

very interesting bet on not just

23:51

building another GPU but actually

23:53

building a different kind of accelerator

23:54

that focuses on a different memory

23:56

hierarchy. Uh they want to maximize the

23:58

amount of SRAMM on the chip and use that

24:01

as very very fast memory for for weights

24:03

and KV cache. So SRAM versus DRAM

24:06

there's two ways to make memory for a

24:07

chip. One is to integrate the memory on

24:09

the logic die itself. Like meaning you

24:11

tell TSMC, I want this many megabytes of

24:14

of storage on my chip. Uh and there's a

24:17

way to build that. TSMC has a standard

24:19

cell library you can use and you can

24:20

just print out a bunch of cells of SRAM.

24:23

The problem with SRAMM is it takes a lot

24:25

of area on the silicon die. Um, so if

24:28

you want to build a large die like let's

24:30

say the Nvidia Blackwell at 800 mm

24:32

square. If you made that whole DS RAM,

24:34

it would be in the maybe like

24:36

singledigit gigabytes, it's not a crazy

24:38

amount of of data storage. Compare that

24:41

to if you're willing to take a different

24:43

process technology entirely. So not TSMC

24:45

anymore, but now Micron SKH Highix

24:47

Samsung. They build DRAM, which is a

24:49

whole different way to build memory

24:51

that's more focused on capacitors than

24:53

transistor cells. So, SRAM, the standard

24:55

way to build SRAMM is what's called the

24:57

6T transistor cell. It's a stable

24:59

transistor arrangement that allows you

25:01

to write a bit to it and then it holds

25:03

that state in that bit regardless of

25:05

whether you keep applying. Well, you had

25:06

to apply some power, but uh it it's

25:08

holding that bit without any sort of

25:10

like active management. It's static.

25:12

Now, dynamic RAM, DRAM, it's dynamic

25:16

because what you do to write some data

25:17

is you write a charge onto a capacitor

25:20

and as soon as you write that charge

25:21

into that capacitor, the charge is

25:23

dissipating. it's been leaking. And so

25:25

the dynamic part of DRAM is that you

25:27

must every 50 milliseconds or so refresh

25:29

every bit you've written. So you're

25:31

constantly juggling billions of balls in

25:33

the air essentially billions of bits

25:34

have to be managed by a memory

25:35

controller which is reading and

25:36

refreshing every bit on the DRM. Now the

25:39

benefit of that is you can get much much

25:40

higher density and it's a whole

25:42

different process technology. There's a

25:43

ton of different trade-offs. Hence why

25:45

we split the DM manufacturing into an

25:47

entirely different company like Micron

25:48

SKX and Samsung. These are the best

25:50

companies in the world to do this. They

25:51

build DM. And if you take DM from those

25:54

companies and you stack it uh into many

25:56

layers and you kind of print them or or

25:59

solder them around the main logic die

26:02

that you get from Nvidia, you can now

26:04

get hundreds of gigabytes uh like

26:06

Blackwell has 288 GB of HPM capacity

26:09

around the logic die. And the logic die

26:12

itself maybe only has like 500 megabytes

26:14

of of SRAM. So it's possibly multiple

26:17

orders of magnitude, three orders of

26:19

magnitude difference in density for DRAM

26:21

versus SRAMM. Okay, so let's go back to

26:23

Cerebras. What are they doing? Well,

26:25

they see this problem, there's not

26:26

really an obvious way to increase SRAM

26:28

density on the chip. But thing with

26:30

SRAMM is because it's so physically

26:32

close to the logic gates that actually

26:35

do the computation, the arithmetic logic

26:37

units are right next to the SRAMM that

26:39

they're going to pull from, the compute

26:42

units that are doing the matrix

26:43

multiplies can pull data from SRAMM at

26:45

just mind-boggling speeds. You know,

26:47

Serbis quits pabytes per second, 21

26:49

pabytes per second further away for

26:50

scale engine 3. And so compare that to

26:53

HBM on an Nvidia black wall is u you

26:55

know 10 terabytes per second or so in

26:58

that range. So once again, many orders

26:59

of magnitude difference, more capacity,

27:02

but proportionally less bandwidth

27:04

essentially.

27:04

>> And so what Cerebrus does is they say

27:06

that we're going to take as many of

27:08

these dies as we can. We're not going to

27:10

limit ourselves to the 800 millimeter u

27:12

reticle limit, the TSMC 800 square

27:14

millimeter limit that TSMC imposes on

27:16

us. We're going to take the entire wafer

27:18

and have actually every die connect to

27:21

every other die over scribe lines. And

27:23

we're just going to try to get as much

27:26

SRAM as we can on the whole wafer. and

27:28

we can get to like let's say 50

27:29

gigabytes of SRAM per wafer and then

27:31

we're going to stack many wafers

27:32

together in a pipeline or similar and

27:35

now we can have you know up to a

27:37

terabyte of memory very very fast memory

27:40

and you do all that work just to get to

27:42

the ability to read data from SRAMM at

27:45

yeah 21 pabytes per second per wafer

27:48

therefore you can now serve these

27:49

language models at extremely high tokens

27:51

per second because you can move the

27:53

entire parameter count of a large model

27:55

like Kimmy uh you can move all that data

27:58

in and off the chip or sorry in and off

28:00

the logic cores in about a millisecond

28:02

or something like that.

28:03

>> So there you go you have a path to a

28:05

thousand tokens per second

28:06

>> and so what is your prediction for like

28:08

that segment of the market? Okay. So I

28:10

think what happens to them is some

28:13

hybrid sort of outcome like we we had to

28:15

pair the Cerebras chip where it's very

28:17

strong. It's very very good at fast

28:19

access to memory with something that has

28:21

more capacity for memory because it's

28:24

true that you can take a one trillion

28:26

parameter model like Kimmy and fit it on

28:28

a large number of cerebrus wafers. But

28:31

you can't do something about the KB

28:33

cache very easily. The KV cache is

28:35

something that grows as people use the

28:37

model more and that is always dynamic.

28:39

You don't even know how much KV cache

28:41

you're going to need. It depends on what

28:42

your users how many users you have and

28:43

how many users you want to serve.

28:45

>> Can you explain KV cache just like in

28:46

>> basic? Yeah. So KB cache whenever you

28:49

use a language model every token you

28:51

send through the language model actually

28:55

uh stays in the context window of the

28:57

language model for as long as you're

28:59

having a conversation. So if I we talk

29:01

for 100,000 tokens, the 100,000th and

29:04

oneth token is still in the conversation

29:07

uh behind us and the model is

29:09

referencing all the past conversation

29:10

history in order to make better

29:12

predictions about what the next thing

29:13

we're going to say is. And so that KV

29:16

cache is a bunch of memory. Um you have

29:18

to store a representation for every

29:20

token that you send through the language

29:21

model. And it frequently gets to be

29:24

larger than the weights of the model

29:26

themselves. You have this like

29:27

crystallized knowledge in the model

29:29

weights and you have the dynamic

29:30

knowledge of the exact conversation

29:32

we're having in the KB cache is the way

29:34

I like to think about it.

29:35

>> Yep. And and this is why sometimes

29:36

people would observe like deep in a

29:38

conversation things start to degrade

29:39

because there's some sort of like

29:40

technical problem.

29:42

>> Yeah. So the KB cache is quite

29:43

interesting in that regard. The KB cache

29:44

is an exact representation of everything

29:46

that came before. We we store all the

29:48

information that we've seen in the

29:51

conversation. However, during training,

29:53

the model did not get trained primarily

29:56

on very long context conversations. It

29:58

got trained primarily on, let's say,

30:00

8,000 token conversations or 16,000

30:03

token conversations. So, if you take the

30:05

model to 200,000 tokens, there was some

30:07

training that happened at that context

30:09

length, but it's not the model's like

30:10

core strength. And so, there's there's

30:11

always been a challenge for the Frontier

30:13

Labs to figure out how do we make the

30:14

model exactly as intelligent at 10,000

30:16

tokens as we expect them to be at

30:17

200,000 tokens. And it's going to be a

30:19

perennial battle for us. We've had 1

30:20

million context windows as a concept for

30:22

for years now. Enthropic was I think the

30:24

first to hit the 1 million context

30:25

window length. I still, you know, use

30:27

/compact in my cloud code uh well before

30:30

1 million context length. I don't think

30:32

it's actually great to hit the full

30:33

length.

30:34

>> And so these extremely fast, extremely

30:36

low latency approaches ultimately are

30:38

limited by by this factor.

30:40

>> Yes, you can do whatever you want for

30:42

the weights. It's very possible to have

30:44

just unbeatable performance on weight

30:46

storage. However, the KB cache is going

30:48

to be a big thorn on your side. And so

30:50

three years from now, five years from

30:51

now, what role do you think these kinds

30:53

of chips play? Like what sort of market

30:54

share do they have in the heterogeneous

30:56

chip market?

30:57

>> Crisis and Grock and maybe a couple

30:58

others, you should think of them as

31:00

accelerators. What they are really good

31:01

at is being used in conjunction with an

31:05

more traditional GPU like device that

31:08

critically has this offchip memory built

31:10

in. You want offchip memory for capacity

31:12

and onchip memory for speed. We want to

31:14

hybridize these two things. So if you

31:16

take uh transformers in the limit, you

31:18

take a transformer to a million context

31:20

length. What ends up happening is you

31:22

have this you know computebound stage

31:24

which is the actual matrix multiplies

31:26

for the uh what we call the MLP which is

31:28

where most of the model's knowledge

31:29

world knowledge is encoded and then you

31:31

have the attention layer which is where

31:32

we're kind of dynamically adapting to

31:34

the current conversation. Attention in

31:36

the limit is usually memory bound and

31:39

the MLP in the limit is computebound at

31:41

large enough batch size. And I would say

31:44

the original sin of transformers is that

31:46

you've taken this extremely

31:48

fundamentally memory bound layer and

31:51

juxtaposed it right next to a

31:52

computebound layer. It is very difficult

31:54

to have a single chip that is good at

31:56

both compute operations and memory

31:58

operations. The GPU is quite balanced in

32:00

this regard, but you have to choose one

32:02

or the other. Cerebrus has a very fast

32:04

memory access for something like a

32:06

matrix multiply and it's really good to

32:09

host the the MLP the the weights

32:12

essentially on the Cerebrus chip but the

32:14

GPU has the capacity to scale to really

32:16

long context lengths and so you would

32:18

like to put the uh attention possibly on

32:20

the GPU and the MLP on the cerebrus chip

32:23

>> and I believe this is what's happening

32:24

with Nvidia and Grock. Can you riff for

32:26

a minute just on transformers and uh

32:29

>> yeah, you've been so good at explaining

32:31

some of the core concepts just for

32:33

people that again aren't aren't deeply

32:34

familiar with what this innovation was

32:36

in 2017

32:37

>> like what its strengths and weaknesses

32:39

are and whether or not you think it will

32:41

remain the dominant architecture or a

32:43

dominant architecture for the future of

32:45

AI. What it did was it it allowed us to

32:47

learn an unsupervised data really

32:49

effectively because transformers what

32:51

they're all about at the end of the day

32:52

is taking any sequence any arbitrary

32:55

sequence of data and trying to find

32:57

patterns in that data and they

32:59

critically the attention operation which

33:01

is the headline uh component of

33:03

transformers. It allows the model to

33:05

dynamically adapt to what it thinks is

33:08

the most relevant component of the

33:10

sequence. every step you take through a

33:13

transformer, you are essentially like

33:15

reweing the input that you looked at

33:17

before and figuring out which is most

33:19

relevant for your next prediction. And

33:22

so it it's extremely amenable to

33:25

>> uh learning arbitrary sequence data. And

33:27

the most interesting sequences of data

33:29

that we produce on a regular basis is

33:31

language

33:32

>> and that's how we got to dominance in

33:34

the language regime. But uh to zoom out

33:36

even further, I think what transformers

33:37

really did well is that they scaled.

33:39

Transformers make no such human prior.

33:41

Transformers just say, "Well, there's

33:44

going to be a pattern in the sequence of

33:45

data, and if there is a pattern, I'm

33:46

going to find it. I'm going to throw

33:48

more and more parameters at this problem

33:49

until it works."

33:50

>> Uh, and transformers benefit from a lot

33:52

of the computer vision work, too. For

33:53

example, one of [clears throat] the

33:54

challenges in computer vision was we had

33:56

a hard time going from hundreds of

33:57

thousands of parameters, which you get

33:59

for like linear models like support

34:01

vector machines or other legacy machine

34:04

learning models. Those had, you know, on

34:06

the thousands of parameters. Then we got

34:08

to deep learning and got to tens of

34:09

millions of parameters with computer

34:10

vision. The biggest models were you know

34:12

around like 150 million parameters was a

34:14

huge model for computer vision. And now

34:16

we routinely talk about trillions of

34:17

parameters and transformers are the link

34:20

to go from millions to trillions of

34:21

parameters.

34:22

>> And so if I think about the important

34:24

units of scaling being data and compute

34:26

>> does it stand a reason then that you

34:27

think transformers will just stick

34:28

around because that's the thing that

34:30

we're good at getting more of those two

34:32

things. Well, data is an open question,

34:34

but comput. Yeah, transformers are so

34:37

they're just such great sponges, you

34:39

know, like you you you increase the

34:41

compute available to a transformer by

34:42

10x and you'll get you'll get some log

34:45

improvement somewhere. Uh and and so far

34:47

the scaling laws really work. They're

34:49

really quite beautiful. And to

34:51

[clears throat] the point about I guess

34:53

what do transformers do really well?

34:55

They extend to almost any data set you

34:56

can throw at them. They're extremely

34:58

powerful general learners. And I think

34:59

what's especially useful about

35:01

transformers over other techniques that

35:03

we've tried to replace attention is

35:05

transformers represent any pair wise

35:08

relationship that you want. Any token in

35:11

the sequence can attend to any other

35:12

token in the sequence. So if there's any

35:14

relationship that's in the sequence at

35:15

all, you're going to find it with

35:16

transformer. Now it may be the case that

35:18

you don't need all toall modeling. You

35:21

don't need every token to look at every

35:22

other token. But if you need to,

35:25

transformers give you that option. And

35:28

until we know a better way to kind of

35:30

prune that space down uh a better way to

35:32

kind of have information modeling be

35:35

more selective attention is a very very

35:37

good operation. This is another kind of

35:39

trick that we learned in the computer

35:40

vision days. Like one of the old

35:42

Karpathy sayings is that you know if you

35:44

have a new data set that you want to

35:45

train a model for. Your first goal

35:47

should be to overparameterize the the

35:49

model and try to overfitit the data that

35:51

you have to prove that there is a

35:53

relationship that you can model or

35:54

memorize that your learning algorithm

35:56

works uh that you can instill knowledge

35:58

into the model. Once you can overfit

35:59

then you can compress and the

36:01

compression is how you get

36:02

generalization. You don't want to

36:03

actually memorize the data that you have

36:04

in front of you. you want to generalize

36:06

and therefore once you overfit the data

36:08

set then you can kind of work backwards

36:10

and try to find the general patterns

36:12

that fit into the smallest parameter

36:13

count possible.

36:14

>> What's your prediction for the future of

36:15

data and riff on the importance of data

36:17

in this whole story? I like the phrase

36:19

that internet was a onetime subsidy on

36:21

data. We got it for free. Uh it's

36:23

extremely high quality about 30 trillion

36:25

tokens of high quality text. Uh 300

36:27

trillion tokens if you take a wider view

36:29

on what qualifies as good text and we've

36:32

basically looked at it all already.

36:34

Models have seen the entire internet

36:36

many times over at this point. And there

36:38

is not a whole lot more to be done on

36:39

human data from the internet. The next

36:41

phase of data in my mind is model

36:44

self-improvement through RL environment

36:47

gyms. Basically, in fact, we don't even

36:49

benefit from getting more like random

36:52

user interactions with AI. It used to be

36:53

that, you know, the the new type of data

36:55

that we cared about a lot was the

36:57

interaction data from people using

36:58

chatbt and giving TetBT signals on what

37:01

they liked and didn't like. I like the

37:03

argument now that the median model that

37:05

we serve is so much more advanced than

37:07

the kind of un

37:10

than like a random human uh giving

37:11

feedback that the signal you get from

37:14

random human preference or I guess

37:15

unconditioned human preference is not

37:17

actually worth anything anymore. You

37:19

want expert human preference at this

37:20

point. The model has outgrown everyday

37:22

>> generic Yeah. Everyday Joe. Exactly. So

37:24

the feature of data to me is giving the

37:26

model a hard verifiable task and letting

37:29

it run in this gym where it's kind of

37:32

isolated and it just has a a problem

37:33

that it can make progress on and get

37:35

measurement of whether it made progress

37:36

on that problem or not. You can imagine

37:38

coding problems are in this category.

37:40

math problems are also in this category

37:42

and um increasingly more and more we

37:44

have we can just give the agent a

37:46

computer essentially and have it act

37:48

like it's a human worker and just give

37:51

it feedback on whether it's making

37:53

progress towards the target outcome that

37:55

environment becomes the data. I think

37:57

this is not a super differentiated take

37:58

but uh it's been really really

38:00

productive from what I've seen so far.

38:01

And you think that just goes on for a

38:03

really long period of time or is that

38:05

another like if I think about the

38:06

internet as this one big block like this

38:08

is another big block that will have its

38:10

you know day in the sun and we'll kind

38:12

of get it all and and then we'll have to

38:14

move on to something else.

38:16

>> I think it's actually more profound than

38:17

that. Basically the idea is that if you

38:19

want artificial general intelligence the

38:21

best way to get there is to just keep

38:23

stacking specialized intelligences until

38:24

you have no more gaps to fill. And the

38:27

test here, the only thing you need to

38:29

make sure you do to make this work is

38:32

you must make sure that your task is

38:33

verifiable. You need to give the model a

38:35

self-grading system. If you have that,

38:37

you have the recipe for self-improvement

38:39

on any task you like. And I think you've

38:42

seen this held up by the way frontier

38:44

labs spend. They used to spend much that

38:46

much on data. Now they spend a lot more

38:47

on RL environments. And these

38:49

environments absolutely capture that

38:51

relationship of recursive

38:53

self-improvement on a verifiable task.

38:55

>> Okay. Okay. Now, so I like that we've

38:57

veered off in different little side cars

38:59

here, but coming back to your initial

39:01

task of making existing hardware more

39:05

efficient.

39:05

>> Yes.

39:06

>> By being more in control of what's going

39:08

on at the hardware level through

39:09

software.

39:10

>> Um so, so yeah, just keep going on what

39:13

you've done so far and what you want to

39:15

do and then we're going to jump to

39:16

hardware and then jump to energy

39:18

finally.

39:18

>> Sounds good. So, yeah, I mentioned

39:20

kernels. It's surprising people think

39:21

kernels are done. There are great people

39:23

like Triau who write excellent kernels

39:25

and they're they form the bedrock of all

39:27

of our um modern deep learning is built

39:30

on flash attention. Modern transformers

39:31

are built on flash attention. But if you

39:33

deviate from the happy path at all, if

39:35

there's a new model that comes out that

39:36

has a slightly different way to embed

39:39

positional information like the change

39:40

of the rope system. Suddenly the kernel

39:42

that we had is not suitable for this new

39:45

model and we may have to make a a patch

39:47

to this kernel. I wouldn't say we're in

39:48

the phase where we had to invent new

39:50

kernels from scratch, but having the

39:52

ability to quickly modify existing GPU

39:55

kernel, sorry, a kernel, by the way, is

39:56

a it's a general term for any program

39:58

you run on the GPU. And so,

40:00

historically, kernels tend to be put

40:02

into a library where every kernel has a

40:05

very very scoped purpose. Typically, you

40:07

have a kernel for a matrix multiply. You

40:10

have another kernel for even something

40:11

as simple as addition. You want to add

40:13

two tensors together, that's another

40:14

kernel.

40:15

>> [clears throat]

40:15

>> And then increasingly we've started to

40:17

fuse those kernels together. So if I do

40:19

a matrix multiply and then I want to add

40:21

it to another matrix that I've also

40:22

multiplied maybe those two become one

40:24

kernel and I just fuse the operations

40:26

where instead of writing the data out to

40:28

DRAM and then reading it back in just to

40:30

do the addition maybe I can just do this

40:32

uh easily.

40:33

>> Why are humans still doing this? Like it

40:35

seems like the sort of thing that AIs

40:37

would be exceptionally good at

40:39

engineering more efficient kernels.

40:41

Maybe that's where we're going and we're

40:42

just not quite there yet. But if if we

40:44

aren't there yet, is that where we're

40:45

going? If we're not there yet, why why

40:47

humans still doing this? Why why is Tree

40:49

out so wellknown? You know, it's a name

40:51

I know.

40:51

>> I don't want to speak for Tree, but what

40:52

he taught me was uh you shouldn't write

40:55

kernels by hand anymore necessarily. I

40:58

like to say we write kernels in the

40:59

whiteboard. We go to the whiteboard, we

41:01

describe what we think the machine

41:02

should be doing, then we succinctly

41:04

describe that in in natural language to

41:06

the a model. And then the model is able

41:08

to do the execution of okay, here is my

41:11

input and output. here is the strategy

41:13

of how we want to dispatch this work

41:15

onto the GPU. I'm gonna go implement

41:17

this.

41:17

>> So, we're doing the conceptual design.

41:18

>> Exactly. And that I think I'm not sure

41:21

exactly why models are not superb at

41:23

doing this. I don't think this is like

41:24

our remote or anything like that. I'm

41:25

sure in 6 months time we'll have much

41:27

better models uh on kernel engineering

41:29

and I'm sure the labs would tell you

41:30

that they already do a lot of their

41:32

kernel engineering uh in a fully

41:33

automated way. And so software as an

41:35

edge,

41:36

>> yeah,

41:36

>> if I think about software as maximally

41:39

near speed of light, efficient usage of

41:40

of an underlying piece of hardware

41:42

>> is is going to trend towards not being

41:45

an advantage for a company like yours

41:46

over time.

41:47

>> That's right. The rising tide of

41:49

something like Mythos or GBT 5.6 Soul

41:51

that lifts all boats. It really does. Um

41:54

I actually don't think there's a point

41:55

in specializing to say we work on making

41:59

the model better for just kernel

42:00

engineering. I think that's actually not

42:02

not the most meaningful subset of of

42:05

like coding in general,

42:06

>> uh, kernel engineering in particular.

42:08

Maybe there's some like privilege

42:10

information you inject into the prompt

42:11

that's like a useful way to steer the

42:13

model to be better at writing kernels,

42:15

but broadly speaking, yes, we're all

42:18

we're all downstream of the frontier in

42:20

terms [clears throat] of this

42:21

capability. I I always love this uh this

42:23

from the history of energy there there's

42:25

always this like pendulum between the

42:27

raw source let's say coal

42:29

>> and then if there's a certain amount of

42:31

energy available inside of a chunk hunk

42:33

of coal like what percent of it we can

42:35

harness and use

42:36

>> and a big part of the history of energy

42:39

was getting that number from 10% to 95%

42:42

or whatever

42:42

>> right

42:43

>> where are we in that's like if I just

42:44

think about it at Blackwell or something

42:46

and Blackwell is the piece of coal

42:48

>> like what percent do you think we're at

42:50

like How how efficiently can we use an

42:54

existing piece today?

42:55

>> There's a lot of different ways to

42:56

analyze that. I think in some level we

42:58

are really efficient at optimizing the

43:01

performance when the GPU is doing the

43:03

thing that it's most happy doing which

43:05

is a large dimension matrix will apply

43:08

that operation runs at you know 70 80%

43:10

of peak utilization and it's limited not

43:12

by software but by power. The way Nvidia

43:14

quotes peak flops is a little

43:16

optimistic. You never hit that because

43:17

of power throttling but um

43:19

>> because of heat.

43:20

>> Yeah, exactly. thermals let's say 70 80%

43:22

it's saturated it's pretty good

43:24

>> but in practice you don't spend the

43:26

majority of your time in a transformer

43:28

in that happy path where you're doing a

43:30

large batch m matrix multiply

43:32

>> and so

43:33

>> uh our job is to basically build the

43:35

engine around the chip such that we are

43:37

feeding the GPU these large batches of

43:39

work at all times

43:40

>> and one of the most profound transitions

43:42

we've had in the GPU world in the last

43:44

year has been this moving of you know

43:46

you don't program one GPU at a time

43:48

anymore you should think about the whole

43:49

rack and maybe you should think about

43:51

the whole cluster, the entire data

43:52

center at a time. And with Nvidia again,

43:54

they've started shipping not just a

43:56

single GPU or a single motherboard, but

43:58

actually the the whole rack system is

44:00

something that they prescribe. They call

44:01

it NVL 72. Uh their latest chip, the

44:04

Grace Blackwell 300, um that ships as a

44:07

rack of 72 units. And it is it's an open

44:11

race to figure out who can program the

44:12

whole rack scale computer as efficiently

44:14

as possible. And my belief is that that

44:17

shape of compute is the future of both

44:19

efficiency and speed. In fact, Nvidia

44:21

does a great job of if you want the

44:22

lowest possible latency, you should be

44:24

using that chip. And if you want the

44:25

highest possible throughput, you should

44:26

probably also be using that chip as of

44:28

right now.

44:29

>> And it's all comes down to like this is

44:30

a very new paradigm of programming.

44:32

>> One of the things you hear is that the

44:33

market for the best chips, blackwells,

44:36

let's say,

44:36

>> is like a drug market or something right

44:38

now. Like there's all sorts of

44:40

fascinating things happening to get as

44:42

many of them as possible because

44:43

everyone's so short. Y

44:44

>> I'd love you to react to that analogy

44:46

like is that what it feels like

44:48

>> but then also to talk about what the

44:49

market is like for like not the bleeding

44:51

edge chips like if I if I

44:53

>> am willing to accept a slightly or or

44:55

moderately inferior chip

44:57

>> what's that market like let us into that

45:00

world

45:00

>> yeah okay a couple things number one uh

45:03

yes basically has a long-term view on uh

45:06

on all their chips they they see this

45:08

immense demand for the black hole chips

45:09

and they they can do what other

45:11

suppliers have done in the past which is

45:12

like just crank prices and made the

45:14

market you know supply and demand curves

45:16

will correct they'll intersect at some

45:17

point and everyone will be technically

45:19

happier but Nvidia sees the if they just

45:22

let the most deep pockets buy all the

45:24

chips that maybe hurts them in the long

45:26

term if that customer ends up acrewing a

45:28

lot of more power they understand that

45:30

compute is power today and so uh they're

45:33

quite strategic about how they allocate

45:35

compute that's the first thought the

45:37

second thought is that relationships

45:38

matter a lot nobody wants to have a huge

45:41

order of of a chip rental come in from

45:43

this new startup that says, "Oh yeah,

45:46

I'm going to rent 10,000 black wells for

45:47

for three years or 5 years." The startup

45:49

has only been operating for months

45:51

typically. Who knows they're good for

45:53

the money. The the way you convince

45:55

someone to give you access to compute is

45:57

is quite challenging uh these days and

46:00

requires some pretty either great

46:01

relationships or uh just incredible

46:04

financial backing to make this happen

46:06

>> on the Nvidia side. And it's all because

46:08

the scarcity is so high and demand is

46:11

just off the charts. Now, for other

46:13

chips, I wouldn't even call them

46:14

inferior. I I like to say there's no bad

46:16

chips. There's really bad pricing. And

46:17

uh I will make any chip work at the

46:19

right price. That's like kind of one of

46:20

the ethoses of the company. And let's

46:22

talk about AMD. AMD, I think great chips

46:25

overall. The challenge is that people

46:27

don't um understand how to program them

46:29

very well. So, you know, I've been

46:32

talking to you about how we have such a

46:33

great kernel team. We're so serious

46:34

about squeezing the performance out of

46:35

the hardware. Nvidia is pretty good at

46:37

doing that for their own chips. Frankly,

46:39

there's some alpha that we can squeeze

46:40

out, but actually there's a lot more to

46:41

be done on other chips because the

46:44

vendor does a little bit less work than

46:45

Nvidia does to make the best kernels out

46:47

of the box or or even better for me

46:50

there is alpha and just like other

46:51

people have this perception that AMD is

46:53

not as as good as Nvidia. That's music

46:55

to my ears. I'm very happy for them to

46:57

sleep on this chip and for me to buy as

46:58

much as I can.

46:59

>> Now, I think that that's not actually

47:01

super true anymore. I think AMD is

47:02

actually uh somewhat popular amongst the

47:05

some large buyers. Um you know I think

47:07

publicly Meta and OpenAI have bought a

47:09

ton of AMD chips and so we're

47:11

increasingly seeing that uh all the AMD

47:13

supply is also being allocated but

47:14

there's a long tale of other companies

47:16

that are popping up yet net new

47:18

companies are great uh such as etched or

47:20

or senova or dmatrix all these companies

47:24

are popping up and I think the main

47:25

challenge for them is scale can they

47:28

actually get enough wafer allocation

47:29

from TSMC to pump out chips to make it

47:32

into the market but certainly if there's

47:34

a new chip on the

47:36

I'd like to know about it as quickly as

47:37

possible and evaluate whether we can buy

47:39

a good fraction of that supply.

47:41

>> And so it's fundamentally an arbitrage

47:42

for you. Like if you can be much better

47:44

at eking out performance from chips that

47:46

have received less attention, you can

47:48

then resell that at a margin and it

47:51

could be a great business.

47:52

>> Exactly. Exactly. And I think that it's

47:53

not the case that everyone else is just,

47:55

you know, has a skill issue that they

47:57

can't uh, you know, make these chips

47:59

work as well. I think we're quite

48:00

competent in this. I think we're

48:02

probably one of the best teams in the

48:03

world to use multiple silicon

48:05

architectures and and be quite

48:06

aggressive in chasing down performance

48:08

in unlikely places. But um yeah, I think

48:10

it's the speed at which we we're willing

48:12

to kind of build our stack around a new

48:14

chip. We don't have a huge amount of

48:16

incumbency around well our data center

48:18

providers are only stuck with with this

48:20

class of chip and it's going to be a

48:22

huge pain for us to to deploy uh these

48:25

net new chips. We have some very

48:26

creative data center partners who are

48:28

willing to move very quickly and there's

48:29

a new class of those that we can talk

48:30

about and most importantly we don't shy

48:32

away from the challenge. Uh that's

48:33

frankly a big part of this is just

48:35

saying yes we love TPUs we're going to

48:37

make TPUs work. Yes we love tranium

48:39

we're going to make tranium work and if

48:40

it doesn't work um easily we're going to

48:43

find a way to fit it in with the

48:45

heterogeneous serving system it will

48:47

have a place every chip has a

48:48

comparative advantage we have to find

48:50

that advantage and then squeeze it in

48:51

that direction. Ju just as an interlude

48:53

before we get to hardware, data centers,

48:56

energy, etc. which will be really fun

48:58

part of the conversation. I'd love you

49:00

to talk about your perception of the

49:03

investor classes worry. Yeah.

49:05

>> Like you look at memory stocks.

49:07

>> Um or my current favorite is you look at

49:09

the chart that plots the percent of the

49:11

S&P 500 that's semiconductors.

49:13

>> Historically it was like 2 3 4%. Now

49:15

it's 19 20 21%. And it just sort of

49:18

looks like if you're a student of market

49:19

history, you get all these things

49:21

through time that are sort of reached

49:23

some crazy near-term peak and then and

49:25

then collapsed back to long-term norms.

49:28

>> Um, and I'm curious how that has all

49:30

investors worried.

49:30

>> Yeah. Uh so you know a lot of people

49:33

made a lot of money in Micron and Skhinx

49:35

and companies like this but everyone

49:37

feels like ah these you know on the long

49:39

term like compute's a commodity and uh

49:42

it will not represent a quarter or fifth

49:45

of the entire market capitalization of

49:46

the world and and so they're scared and

49:48

that's the setup. Um everyone

49:50

acknowledges that like there's a huge

49:52

shortage but everyone sort of feels like

49:53

ah we'll figure it out and these things

49:55

will revert back down to their their

49:57

normal place in capital markets. I'm

49:59

curious what you think about about that

50:01

narrative.

50:02

>> One thing, I'm less of a student of

50:03

history as more of a member of history.

50:05

I was I was born in 1997 and uh my mom

50:08

worked at Intel in the 2000 in the

50:10

run-up to the year 2000 and the the do

50:12

boom and crash and you know I remember

50:14

the time where Cisco was the most

50:16

valuable company in the world and and

50:17

Intel was close behind. I mean I mostly

50:19

draw parallels to that period of history

50:21

from 25 years ago to today. And I think

50:24

the main difference is that a lot of the

50:26

investment in networking equipment

50:28

historically was speculative. We

50:30

anticipated this future demand for users

50:33

that never came. And I think what's

50:35

interesting about token consumption or

50:36

AI consumption broadly is that it's no

50:38

longer speculative. People buy tokens

50:40

because they're immediately valuable to

50:42

them. You don't hoard tokens, you use

50:44

them immediately. This is also even

50:46

different from what we had 2 years ago

50:48

where there was a supply crunch for

50:49

hopper generation chips in 2023 2024. Uh

50:53

in that period it was all training

50:54

oriented spend and training is

50:56

inherently speculative. Now it's

50:58

everyone is instituting caps on how much

50:59

you can spend on cloud code. It's a very

51:01

very different world to be talking about

51:03

inference spend and predicting inference

51:04

spend to go up. I do think inference

51:06

spend monotonically increases. Uh

51:08

there's no speculation on inference

51:10

spend. Vanta automates security and

51:13

compliance for over 16,000 fast-moving

51:15

companies like Ramp, Cursor, and Harvey,

51:17

keeping them audit ready around the

51:18

clock. It's the number one Agentic Trust

51:20

platform, and it now helps companies

51:22

like yours watch for the risks that show

51:24

up between audits across your vendors,

51:26

your AI tools, and your whole [music]

51:27

environment. Every new tool your team

51:30

signs up for, every vendor that turns on

51:32

AI features, [music] is an opportunity

51:33

for something to go wrong. And most

51:35

security programs weren't built for AI's

51:37

pace of growth. The Vant agent works

51:39

like a 24/7 GRC engineer in the

51:42

background, finding issues, drafting

51:43

fixes for you, and cutting vendor

51:45

assessment time by up to 50%. Whether

51:47

you're a fast growing startup or a

51:49

global enterprise, [music] Vanta helps

51:50

you earn and prove trust. Invest like

51:52

the best listeners. Get a special offer

51:54

of $1,000 off Vanta at vanta.com/invest.

51:58

Ridgeline is the first endto-end system

52:00

of record with embedded AI for

52:02

investment [music] management firms

52:04

running portfolio accounting,

52:05

reconciliation, reporting, trading, and

52:07

compliance, all on one unified platform.

52:10

Firms are moving off legacy technology

52:12

and onto Ridgeline because of how far

52:14

ahead Ridgeline's AI features are

52:16

compared to anything else in investment

52:17

[music] management software. I've been

52:19

hearing from a lot of investment

52:20

managers about AI, and they fall roughly

52:22

into two camps, with some unsure where

52:24

to even start, [music] and others

52:25

convinced they can build their own order

52:26

management system over just a weekend.

52:28

The reality is that running an

52:29

investment firm will always require

52:31

governance, [music] controls, and a

52:32

single source of truth for your data.

52:34

And no amount of AI enthusiasm changes

52:36

that requirement. If you're serious

52:37

about your firm's AI strategy, Ridgeline

52:39

[music] should be part of that

52:40

conversation. And you can request a demo

52:42

at ridgeline.ai.

52:44

>> Coming back now to your take on

52:47

hardware. And so the unit level is

52:49

interesting to me like talked about

52:51

chips, talked about racks, talked about,

52:53

you know, clusters. I'd love to talk

52:55

about data centers

52:56

>> and you said you've had some interesting

52:58

partners doing some cool things. Talk us

53:00

about the the present and future of data

53:03

centers as you see it.

53:04

>> Y

53:04

>> because this seems like you know

53:06

obviously a critical thing for being

53:08

able to serve all this inference is like

53:09

lots of innovation in in this part of

53:11

the world and obviously you're focused

53:12

on it.

53:12

>> I think one of the themes in our

53:13

conversation has come back to what is

53:15

training versus inference like what is

53:16

the difference between these two?

53:17

categories and you know what was

53:18

different about two years ago being

53:19

training oriented and today being

53:21

inferenceoriented and I think the most

53:23

conservative players in the entire AI

53:24

stack have got to be the infra players

53:27

whether that's data centers or even more

53:28

conservative is TSMC the chip infra

53:30

people uh and so data centers

53:32

historically were built like AI data

53:35

centers they were built for training uh

53:37

and training is the superset workload

53:39

over inference you can make any training

53:41

cluster work for inference but maybe not

53:42

vice versa and what the difference there

53:44

is networking uh how much do you invest

53:47

in bandwidth between chips and how large

53:49

of a cluster do you need? There's

53:51

actually a diseconomy of scale to to

53:53

data centers in some way. Like it's way

53:55

more expensive and difficult to build a

53:57

um you know 100,000 GPUs in one data

53:59

center than it is to build 10,000 than

54:00

it is to build 1,000. And and we now we

54:02

just talk about you know how many

54:03

megawatts or gigawatts do you have? And

54:05

basically there's no way to build a

54:07

gigawatt data center in the United

54:09

States easily anymore. Even 100

54:12

megawatts is is increasingly hard. It's

54:14

basically impossible unless you're a

54:16

very special set of customers. Uh 10

54:18

megawatts is probably on the edge of

54:19

what's possible today and 1 megawatt I

54:22

would argue is plentiful. So there's

54:24

this incredible lore on the market where

54:26

you can find lots of aggregate power but

54:29

it will not be concentrated and that was

54:31

not interesting to anyone who's building

54:33

training uh data centers because you

54:35

just assume all be in one spot for no

54:38

one wants to deal with cross data center

54:39

training.

54:40

>> So the market has some lag in it. I

54:42

think that the market still assumes that

54:44

we have to go shake down those 100

54:46

megawatt and 10 megawatt data centers

54:48

wherever we can find them is still the

54:49

attitude I hear from a lot of data

54:51

center developers

54:52

>> but increasingly we're seeing a few new

54:54

thinkers realize that inference is going

54:56

to be suitable for these distributed 1

54:58

megawatt data centers and uh we're we're

55:01

quite in agreement with that and we are

55:03

very happy to buy small pools of compute

55:05

across the United States and use that as

55:07

our inference fleet. give us a sense of

55:09

literal physical size of uh 1 megawatt

55:12

versus 10.

55:12

>> Yeah. Well, so this got really wonky

55:15

with the advent of liquid cooling. Now

55:16

you can pack insane levels of power

55:18

density into a single physical rack.

55:20

Like a megawatt of compute, you you'd

55:23

imagine this like massive data hall,

55:24

like a huge warehouse basically. And now

55:27

you can actually pack that into Yeah.

55:29

around like around like eight racks

55:32

worth of compute. Each rack is about the

55:34

size of a refrigerator. You can just

55:35

imagine eight of them lined up. Um,

55:37

yeah, that's a megawatt. And so your

55:39

view would be that the future that you

55:41

want to help build is a whole bunch of

55:43

different chips that can be used

55:45

together. Yes.

55:46

>> That you can buy, you know, you're a

55:49

buyer to ek out the most per chip.

55:52

>> And that those chips can then be coupled

55:54

in very small

55:55

>> data centers

55:56

>> to just do inference. And that those two

55:59

steps of a whole bunch of random

56:01

compute, some of which is cheaper than

56:03

it should be,

56:04

>> your ability to eat more out of it, and

56:06

then small units of expression in a data

56:09

center

56:09

>> equals way cheaper intelligence. I

56:12

certainly think so. Yes, there's a lot

56:14

of ways to access cheaper flops if

56:15

you're able to be creative with what you

56:17

take. And so, one of the ways that I

56:19

describe what we do is we will buy any

56:22

chip anywhere in the world for any

56:24

duration of time. That is a level of

56:25

flexibility and liquidity that I think

56:27

no one else has right now. Uh we're very

56:29

aggressive about putting our money where

56:31

our mouth is and we will we will really

56:33

take any capacity uh and find a way to

56:35

make it work in our fleet. And that is a

56:37

big part of our advantage today and long

56:39

term we had to create more of that

56:41

advantage by investing in these data

56:43

centers that other people are going to

56:45

be skeptical of because you know what's

56:46

going to happen when you set up these

56:48

like this army of a thousand small data

56:50

centers versus the one the one big

56:51

gigawatt data center. Well, few things.

56:54

You're not going to have power

56:54

redundancy frank quite often. You're not

56:56

going to have backup diesel generators

56:57

on site. Those are all very expensive.

56:59

We cut all that overhead. We're not even

57:00

going to have redundant networking in a

57:02

lot of cases. We're going to put these

57:03

in facilities where we have good access

57:05

to power, a single source of power and

57:07

we're going to trench one line of fiber

57:08

to these data centers, but we're not

57:10

going to have like three lines of fiber

57:12

with redundancy and failover and SLAs's.

57:14

It's just going to go down sometimes. In

57:16

fact, I won't be surprised if some of

57:17

them get down to like 95% uptime,

57:19

>> which is bad.

57:20

>> Very bad. That's fatal, atrocious for

57:22

anyone else. survive in a in a big a big

57:24

gig

57:25

>> you'd have basically zero buyers for a

57:27

data center that has 95% up time

57:29

>> I'm that first buyer I will buy 95%

57:31

>> up time

57:32

>> and the reason for that is because of

57:33

this background engine thing that if

57:35

there's things running in the background

57:36

you don't care

57:37

>> partially uh it's actually two things

57:39

one is that we have a really robust

57:40

control plane that is going to be fine

57:43

handling any single failure in any

57:44

single data center as long as it's not

57:46

correlated with other data centers and I

57:47

can just move the workload somewhere

57:48

else I'm cool with that um the failures

57:50

happen at some rate and I am basically

57:53

linearly happy with a data center that's

57:55

95% uptime versus 98% versus 99%. It's

57:58

just linearly good or bad for me.

58:01

>> Mhm.

58:01

>> Now, you do need that async piece that I

58:05

mentioned of, you know, we serve these

58:06

long horizon agents because what happens

58:08

when a request fails is that I'm going

58:09

to have to go find a new GPU to put that

58:12

request on. And that means that for that

58:14

single turn of the agent's work, you

58:16

know, it's working for an hour, but then

58:18

it hits a roadblock because it's GPU got

58:20

pulled away. At that moment in time,

58:22

that agent is going to experience maybe

58:24

like an extra minute or two or three,

58:26

maybe even 10 of latency. But my

58:28

argument is that my customers don't care

58:30

because their agent was running for

58:32

hours.

58:32

>> They're sleeping.

58:33

>> Doesn't matter. [laughter] It doesn't

58:34

matter if like a single turn

58:35

occasionally becomes uh a little bit

58:38

longer. Yeah. So we tell our customers,

58:40

look, our average throughput is going to

58:41

be very competitive, but our P99, our

58:44

99th percentile latency, it's not going

58:46

to be controlled. It cannot be. And in

58:49

return, I'll give you unbeatable

58:50

economics.

58:51

>> And I think that's the right fit for

58:52

background agents.

58:52

>> Talk about power as a category. What are

58:54

you seeing that's interesting,

58:55

innovative, where do you think this

58:56

goes?

58:57

>> Okay, so I said I want 95% uptime on my

59:00

on my data centers. Could I even take

59:02

80% up time at the right price?

59:04

Probably. Um, and what does that mean?

59:06

Well, I'm a son of California. I love

59:08

solar and wind. I think solar and wind

59:09

power is way undertapped in the United

59:11

States. And the challenge has always

59:13

been this intermittency. You would even

59:14

consider solar and wind unsuitable for

59:16

data centers because you have a

59:18

persistent base load and an intermittent

59:20

power source. What are you going to do?

59:21

Well, I think we're actually not that

59:23

far from solving that problem. I am

59:25

totally capable of tolerating a outage

59:28

for my data center that's measured in in

59:30

even days or weeks which is like the

59:32

worst case nightmare scenario for a data

59:33

center is that we're going to have a

59:34

long-term outage because the wind is in

59:37

blow and the clouds are in the sky fog

59:39

is hanging over the valley for some

59:40

time. That's the worst case scenario.

59:41

It's in fact highly predictable and I

59:43

can just call in capacity in some other

59:45

place of the world whenever that

59:46

happens. uh I'll just model the weather

59:48

and figure out when my data center is

59:50

going to be offline, move my data my

59:52

workload somewhere else and it's fine.

59:54

The trick is that it's going to give me

59:55

better access to power that no one else

59:57

is going to touch because it is so

59:58

annoying to deal with that kind of

1:00:00

outage.

1:00:01

>> And if my chips are cheap enough,

1:00:03

they're probably not going to be Nvidia

1:00:04

racks. And if my chips are cheap enough,

1:00:06

I don't mind the capital cost of having

1:00:09

idle chips.

1:00:09

>> Yeah. So I've heard you describe this

1:00:11

entire system as like a scavenger

1:00:13

strategy.

1:00:13

>> That's right.

1:00:14

>> Is is that Yeah. Unpack that analogy a

1:00:16

little Well, first we scavenge chips and

1:00:18

then we scavenge power for those chips.

1:00:21

The idea is in both cases I do not want

1:00:23

to be in bidding against Anthropic or

1:00:25

Open AI for compute capacity. I'm not

1:00:26

going to win against them and I don't

1:00:27

want to. I want to be more creative and

1:00:29

use the supply that they don't find

1:00:30

legible today. And over time I amass

1:00:33

enough aggregate supply. I'm never going

1:00:35

to get concentrated supply. I will only

1:00:37

get aggregate supply. And over time I

1:00:39

build my aggregate factory that is

1:00:41

unbeatable in economics.

1:00:43

>> We are building a factory. We're trying

1:00:44

to build the best steel factory in the

1:00:46

world. Uh but it will come through mini

1:00:48

mills not through large

1:00:50

monolithic steel plants.

1:00:52

>> And and if I imagine the different

1:00:53

versions of this like how vertically

1:00:55

integrated you can be.

1:00:57

>> Yeah.

1:00:57

>> One version would be the extreme would

1:00:59

be you own everything. So that it's a

1:01:01

very capital inensive business. You own

1:01:04

the power you know source. You build the

1:01:06

data centers. You design your own chips.

1:01:09

You control the software that eats the

1:01:11

most out of those chips. And you sell

1:01:13

the end finish token

1:01:15

>> to your user like your user is me

1:01:18

>> and you just own the whole stack.

1:01:20

>> But you can imagine many other

1:01:21

permutations of the business

1:01:23

>> where you know you could whatever you

1:01:26

draw the line anywhere you could be

1:01:27

incredibly capital light own nothing and

1:01:29

just be like the coordination plane

1:01:32

>> across all this stuff the virtual

1:01:33

scavenger

1:01:34

>> right

1:01:35

>> how do you think about that question of

1:01:37

like which which type of these

1:01:39

businesses to be

1:01:39

>> you know there's actually two parts of

1:01:40

me to receive that question. One is the

1:01:42

CEO of a company that needs to work

1:01:44

every day and and grow as sustainable

1:01:46

and and as quickly as it possibly can.

1:01:49

The other is the founder and the founder

1:01:50

is much more imaginative and and just

1:01:52

loves this stuff. The founder in me

1:01:55

wants to do everything. This is my

1:01:56

entire life. I spent my entire life

1:01:58

thinking about chips, power, energy.

1:01:59

Like all I care about is this stuff. So

1:02:01

of course I want to be maximally

1:02:02

ambitious. I don't want to stop ever. I

1:02:04

will never stop until I have built the

1:02:06

most efficient system from soup to nuts.

1:02:08

>> You're doing real life factorial

1:02:10

basically.

1:02:10

>> Very much so.

1:02:11

very much so. So that's like the

1:02:12

emotional from the hard answer. On the

1:02:14

CEO side, I think

1:02:16

>> we have to be more pragmatic. I think

1:02:17

that the capital we're we're looking at

1:02:20

for owning everything is like you said,

1:02:22

it's insane. Yeah, software has high

1:02:24

leverage, so we have to start with

1:02:25

software, but ultimately, you know, do

1:02:27

we own power generation or can we get

1:02:29

great power purchase agreements with uh

1:02:31

utilities? I'm more inclined to pursue

1:02:33

like letting other people specialize in

1:02:36

the things that they're historically

1:02:37

good at and then see if we can get to

1:02:39

the scale. I think of it as like I want

1:02:41

to get to the scale where I earn the

1:02:42

right to take this under our wing. I

1:02:44

absolutely think that there's

1:02:46

efficiencies to be gained everywhere in

1:02:47

the stack. If you can break the

1:02:49

assumption that people I would be buying

1:02:50

from, they made assumptions about who

1:02:52

their customers would be. And I maybe

1:02:55

break those assumptions. It's a pretty

1:02:56

optimistic view. Uh I think it's only

1:02:58

possible because we're actually trying

1:03:00

to underwrite the largest market for

1:03:02

compute in the history of computing.

1:03:04

we're actually going to build so many

1:03:07

billions, trillions of dollars of

1:03:08

investment into inference. Uh, and

1:03:11

because of that focus, it makes sense to

1:03:13

build a lot of things that are custom

1:03:14

for inference. And it's my job to seek

1:03:16

all the places where that's possible.

1:03:18

And then as as they become obvious to me

1:03:21

and my and my partners, I will get my

1:03:23

partners to build custom things for me.

1:03:25

And if they can't do it for me, I will

1:03:27

do it myself.

1:03:27

>> If you had to just zoom out on this

1:03:29

entire system, software, hardware,

1:03:31

energy, etc., and stackrank the places

1:03:33

that you think that we are the most

1:03:35

inefficient today at producing useful

1:03:38

intelligent tokens.

1:03:40

>> Yeah.

1:03:40

>> What does that list look like?

1:03:42

>> I think compute scaling is actually like

1:03:44

very efficient. Uh as in like you give

1:03:46

me more flops and I will use more flops.

1:03:48

And I would say we're actually fairly

1:03:50

judicious already with our use of flops.

1:03:52

Uh if you look at a modern model, there

1:03:54

are very few models that are more than

1:03:55

10% dense, meaning 10% of the possible

1:03:58

number of experts you can activate are

1:03:59

activated. And I think the frontier

1:04:01

models are closer to like 1%. So fairly

1:04:04

sparse already. I don't think that we're

1:04:07

wasting too much on the MOE side. People

1:04:09

have been working with for quite some

1:04:10

time. They're pretty good at squeezing.

1:04:12

Where we are not good is attention and

1:04:15

its use of memory. Specifically, the KV

1:04:17

cache is quite uncompressed right now. I

1:04:21

think if you look at the entropy in a KV

1:04:22

cache, it's nowhere near it's not

1:04:25

earning its keep. Like we're storing

1:04:26

many kilobytes of data in the KV cache

1:04:28

per token. Um, and that's probably off

1:04:31

by an order of magnitude or two. And I I

1:04:35

don't know what the Frontier Labs do,

1:04:36

but Deep Seek certainly publishes really

1:04:38

interesting work to compress that

1:04:40

further and further. And they're making

1:04:42

good progress. And I think the fact that

1:04:43

they're able to make order magnitude

1:04:44

progress here every year or so signals

1:04:47

that there's a lot more room to go. I

1:04:49

guess this all on the micro scale. If

1:04:51

you zoom out further, I think that we

1:04:53

actually don't marshall our compute

1:04:54

effectively at all. Like we have all

1:04:55

this compute in the world. Nvidia is

1:04:57

pumping out 5 million Blackwell chips

1:04:58

this year. Where are they all going? Are

1:05:00

they all being used at all all the time?

1:05:01

I certainly doubt it. I think that at

1:05:03

some level we just need better

1:05:04

orchestration of compute across the

1:05:05

world. Uh this is very difficult to do

1:05:07

because a lot of the comput disappears

1:05:09

into private pools of compute that will

1:05:11

never see the light of day and those

1:05:12

GPUs sit very sadly idle. Uh it's

1:05:15

actually it pains me physically to see

1:05:17

that those GPUs are just you know

1:05:18

silicon and power went into that and

1:05:21

it's just sitting idle and I want to fix

1:05:22

that. how we or organize and orchestrate

1:05:25

the world's compute as a shared resource

1:05:27

and and pack it more efficiently. I

1:05:29

would estimate that, you know, we all

1:05:31

make fun of XAI for having, you know,

1:05:33

some challenges with total flop

1:05:36

utilization on its clusters, but um the

1:05:38

reality for the rest of the world is

1:05:39

it's far far worse. A ton of GPUs just

1:05:41

sit in warehouses or sit in private

1:05:43

pools allocated to a specific customer

1:05:46

um just don't get utilized.

1:05:47

>> You're attacking the efficiency of that

1:05:49

very directly.

1:05:49

>> That's way more effective. Yeah. Um,

1:05:51

what about fabs? Like what do you think

1:05:54

is the future of fabs themselves? Like I

1:05:57

think everyone is wondering

1:05:59

>> will the memory companies will TSMC will

1:06:01

Intel and others

1:06:04

>> be able to how will they expand capacity

1:06:06

basically?

1:06:07

>> Yeah.

1:06:07

>> Will we do it here in the US?

1:06:09

>> Um, yeah. Riffon fabrication of chips

1:06:12

themselves. Like if we could just snap

1:06:13

our fingers and have 100 times the

1:06:15

chips, you know, in the stock today, uh,

1:06:18

we'd probably have way cheaper way

1:06:19

cheaper tokens. So yeah, that that seems

1:06:22

like an important part of the universe

1:06:24

to hear your view on.

1:06:25

>> Yeah. Well, it's interesting. Everything

1:06:27

grows in balance with each other, right?

1:06:28

If we snap our fingers and double all

1:06:30

those things, you might fix a TSMC

1:06:32

bottleneck that there you're just going

1:06:33

to run into another bottleneck. You make

1:06:34

20% more chips, then you have another

1:06:35

bottleneck immediately. I will say

1:06:37

though, it is interesting what they

1:06:40

consider to be a mustd deliver. uh like

1:06:42

what they consider to be like an

1:06:44

invariant that their customers me are

1:06:46

always going to want versus what I think

1:06:48

of as like a more fluid relationship. I

1:06:50

think that if a fab exposes more of

1:06:52

their trade-offs to me, I'm able to make

1:06:56

more intelligent decisions about what I

1:06:58

think I can I can do.

1:07:00

>> One of the most interesting examples

1:07:01

here is that any fab has a lot of spread

1:07:04

in their like worst chip that comes out

1:07:06

of the production line and the best chip

1:07:07

that comes out of the production line.

1:07:08

There's a lot of variance in how chips

1:07:09

are made. Uh and then the question is

1:07:11

like you know if you have a company like

1:07:12

TSMC they work very very hard to tighten

1:07:15

what we call these process corners. We

1:07:17

want to keep the worst chip as close in

1:07:20

characterization to the best chip and

1:07:21

they get a great lens to make that

1:07:23

possible. But that means that they are

1:07:25

adding a lot of controls in the process

1:07:27

that maybe I don't need. Maybe I'm

1:07:29

actually willing to find a place for

1:07:30

that worst chip. You don't need to

1:07:32

tighten the process control as much

1:07:34

which takes more time and cost. Uh maybe

1:07:36

I'm willing to take a lot more rejects.

1:07:38

And I think for us it's like a more

1:07:40

holistic optimization around there's you

1:07:42

know cost of the dies, supply of the

1:07:44

dies and then the cost of power and

1:07:45

places we can put them. And my whole

1:07:47

goal is to actually so dramatically

1:07:50

expand the supply of of power uh across

1:07:53

the United States that I have a home for

1:07:55

a lot of chips that otherwise would not

1:07:57

have earned earned their place in a data

1:07:59

center.

1:07:59

>> Can we talk about how you designed the

1:08:01

system of your own business?

1:08:03

>> What lessons have you learned? You

1:08:04

talked about some of interesting Nvidia

1:08:05

lessons. Yeah.

1:08:06

>> But like bring me into the culture and

1:08:08

how you structure a team and a business

1:08:10

where this is the northstar.

1:08:12

>> I think there's a lot of um you know in

1:08:14

the limit thinking we don't worry about

1:08:16

the immediate nature of like when we

1:08:20

start working on a model the efficiency

1:08:21

is not going to be very good. Uh but we

1:08:24

we we think about like where we could

1:08:26

end up in in like a month or or six

1:08:28

months or a year's time. We don't accept

1:08:30

the state of the of the machines we work

1:08:32

on as fixed like even something like the

1:08:35

blackwell chip if we think that there's

1:08:37

some bottleneck that is holding us back

1:08:40

from achieving this performance. I mean

1:08:41

it's very important to me that we we

1:08:43

understand and characterize that very

1:08:44

well and write it down so we can both a

1:08:46

tell Nvidia about it or friends and also

1:08:49

to basically keep this in mind for

1:08:51

future chips that we buy. We want to

1:08:53

learn things that are what we think are

1:08:56

um essentially like invariant for us or

1:08:57

the company uh long term and and kind of

1:09:01

fold that into future decisions that we

1:09:03

make. We're very collaborative. I think

1:09:04

one of the most important traits that we

1:09:06

look for are people who either who are

1:09:08

both good students and great teachers.

1:09:09

Um a lot of our people on the team were

1:09:12

TAs in college and and loved the

1:09:14

experience of of sharing knowledge in

1:09:16

this way. uh we we do whiteboard

1:09:18

sessions all the time and I think the

1:09:20

collegial environment where everyone has

1:09:21

something to teach and something to

1:09:22

learn is is extremely important for us.

1:09:25

What are the attributes of people that

1:09:26

you would want to hire that you think

1:09:28

will be resilient to you know the work

1:09:31

environment 3 years from now when more

1:09:33

stuff is handled by machines?

1:09:35

>> Curiosity. It's 100% curiosity. You know

1:09:37

the one thing I cannot teach is love for

1:09:39

performance, love for uh digging into

1:09:43

every microscond that the machine is

1:09:45

working and understanding what's

1:09:46

happening on the machine at that time.

1:09:49

That to me is the most important trait

1:09:50

for a performance engineer and it's what

1:09:53

I look for. I don't look for lots of AI

1:09:54

experience. I don't look for, you know,

1:09:56

CUDA experience at all. That's actually

1:09:58

a huge red herring. I mean, CUDA as a

1:10:00

concept or GP as a concept have evolved

1:10:02

so much in the last 5 years. There's no

1:10:03

point asking for 10 years of experience.

1:10:05

I want to teach that, but I cannot teach

1:10:07

the love for performance engineering.

1:10:09

That is what I seek.

1:10:10

>> Can you give your assessment of the

1:10:11

major labs

1:10:13

one by one, but also then the

1:10:16

relationship of like closed source as a

1:10:17

category to open source and like what

1:10:19

you think is happening and will happen

1:10:21

>> in a line. I would say the labs pay an

1:10:23

immense premium to be 3 to 6 months

1:10:25

ahead of of everything else. Uh and I

1:10:28

think that's probably still worth it. I

1:10:30

think it makes perfect sense for open

1:10:32

and anthropic to do what they do. You

1:10:33

know there's a sensitive topic around

1:10:35

distillation which I think is part a

1:10:37

very core piece of the relationship

1:10:38

between closed and open frontier. And

1:10:40

you know I'd like to offer an

1:10:41

alternative view on that which is there

1:10:43

is the sense that distillation is theft

1:10:45

that you are taking something from the

1:10:47

frontier models when you distill on

1:10:49

their outputs. And in fact, even if

1:10:52

that's not your intent, even if you

1:10:53

don't ever try to, you know, scrape data

1:10:55

from anthropic, one thing I'll offer is

1:10:57

that an increasingly large percentage of

1:10:58

the artifacts we put out on the internet

1:11:01

are AI generated. Even if you just look

1:11:03

at GitHub alone, you know, what

1:11:04

percentage of repos created in the last

1:11:06

year do we think were created by cloud

1:11:08

code? Um, do we consider that to be

1:11:10

distillation? Because that's probably

1:11:11

all we need. I would not be surprised if

1:11:12

you could train a fable glass model only

1:11:14

on the outputs of code you consider good

1:11:16

on GitHub that's open source. And

1:11:18

certainly if we take the position that

1:11:19

users own the outputs of their

1:11:22

interaction with AI and they choose to

1:11:24

put that up on GitHub, which a lot of

1:11:25

them do, we're going to have latent

1:11:27

distillation for a long time. It seems

1:11:29

fundamentally impossible for me. Like I

1:11:31

I don't think it's fundamentally

1:11:33

possible to prevent the diffusion of of

1:11:37

information or model capabilities. It

1:11:38

will happen. The question is just how

1:11:40

fast. And so then the question becomes,

1:11:42

do scaling and improvement laws hold

1:11:44

forever or for a really long period of

1:11:46

time? And if they do, then there's value

1:11:49

to being three and six months ahead and

1:11:51

that will just last as long as it lasts

1:11:52

and they can charge a huge premium for

1:11:54

those tokens relative to a very cheap

1:11:56

open source token. Is that the right way

1:11:57

to think about it?

1:11:58

>> I think it's possible. I don't know that

1:11:59

the premium for being 3 to six months

1:12:01

ahead is going to last that long. I

1:12:04

mean, if you look at like enterprise

1:12:05

deployments, uh, they don't move at 3 to

1:12:08

six month speed. A lot of enterprises

1:12:09

are probably still on like 46, Opus 46

1:12:12

or Opus 47. They don't they don't adopt

1:12:14

the bleeding edge rapidly. There's a lot

1:12:16

of questions that people have around

1:12:18

rolling out any change at all. And I

1:12:20

think we're just so early in scratching

1:12:21

the surface that um I don't think

1:12:24

there's any way to call a winner in this

1:12:25

race and certainly I don't even think

1:12:27

this is a race that can be decided ever.

1:12:29

There's always it's a continual process

1:12:31

and fundamentally I don't think open

1:12:33

source ever goes away. If there's a

1:12:34

vacuum because one leader steps out, a

1:12:36

new leader will step in. There's too

1:12:37

much incentive and too much there's a

1:12:39

lot of tailwinds too. It's just it gets

1:12:41

easier every day to treat to train a

1:12:43

frontier class model.

1:12:44

>> And so your hope of what the future

1:12:45

looks like is what like what balance

1:12:48

between closed and open, you know, what

1:12:50

balance between model companies doing

1:12:52

everything because they have the

1:12:52

advantage of owning the stack or

1:12:54

whatever. You know, Enthropic can do

1:12:56

that, you know, is like the new Google

1:12:57

could Google just do that or something.

1:12:59

What do you hope the future looks like?

1:13:01

>> I want abundant tokens and diverse

1:13:04

harnesses. I want everyone to build

1:13:05

their own harness and and

1:13:07

>> every company

1:13:07

>> every company every user even make the

1:13:10

agent your your own. Uh I think we're

1:13:12

we're very not that far away from that

1:13:14

level of customization and capability. I

1:13:16

want people to own their intelligence

1:13:18

and I want that intelligence to be

1:13:19

customized probably not through weight

1:13:21

fine-tuning but probably through more in

1:13:23

context learning. That's a more

1:13:24

technical detail. But the underlying

1:13:27

input to this abundance future is about

1:13:30

is basically cheap tokens. My job is to

1:13:32

make the tokens as cheap as humanly

1:13:33

possible. I will achieve that and I will

1:13:35

do it through every layer in the stack

1:13:37

available to me. I love the supply side

1:13:38

levers. I will use every chip. I'll use

1:13:40

every source of power and I will use

1:13:41

every piece of land in the United States

1:13:43

that's you know suitable for this. And

1:13:45

in return, people will have the

1:13:47

incentive to explore what it's like to

1:13:49

have abundant intelligence. We still

1:13:51

treat the agent as a person that is

1:13:53

expensive to consult and you should ask

1:13:55

them when you have a hard question.

1:13:57

That's not the way to think about

1:13:58

intelligence. It's incredible that the

1:14:00

machine can think and we should try to

1:14:01

get that into as many hands as as many

1:14:03

people as possible.

1:14:04

>> You sit in such a unique seat and you

1:14:05

have such a unique perspective on like

1:14:07

what you're trying to do to make this

1:14:09

feature a reality. What do you think are

1:14:11

your most like divergent views of the

1:14:14

world versus your friends who are really

1:14:15

well informed and interested in this

1:14:17

stuff? Like what what make your what

1:14:18

ideas of yours make your friends look at

1:14:20

you like you have three heads?

1:14:21

>> Most of the ideas on chips, I would say.

1:14:23

You know, when I talk about building

1:14:24

custom chips and they ask me, "Oh, so

1:14:26

what's different?" Basically, it's it's

1:14:28

about sidestepping the HPM shortage and

1:14:30

focusing on more extreme offload to

1:14:34

other forms of memory such as flash. Um,

1:14:36

I'm quite passionate about that idea.

1:14:38

Everyone on my team knows that I keep

1:14:40

banging the drum around like what would

1:14:41

we have to change about the model

1:14:42

architecture to make offloading KB cache

1:14:45

to flash work at a much greater level.

1:14:47

And um I'm whiteboarding that all the

1:14:49

time. That's like in the community of

1:14:50

like inference people. you know we have

1:14:52

some divergent views on what you can do

1:14:54

if you design a system around serving at

1:14:57

you know one to 10 tokens per second

1:14:59

which is our whole north star more

1:15:01

broadly I think there is this like

1:15:04

larger sense around you know what do you

1:15:05

do how do people consume a trillion

1:15:07

tokens per day like that's the world we

1:15:09

want to create the capability for them

1:15:10

to do that

1:15:11

>> what's a trillion tokens like ground us

1:15:12

in how much that is

1:15:13

>> a trillion tokens well okay at openi

1:15:15

pricing that's at least $5 million at

1:15:18

the very least for 5.5 or 5.6 six. Yeah,

1:15:20

I think the dollars was probably the

1:15:21

most good metric. Yeah.

1:15:23

>> Yeah. It's millions of dollars.

1:15:24

>> Yeah. So, what's the world in which we

1:15:25

consume what currently costs $5 million

1:15:27

per person per day?

1:15:29

>> Yeah. I mean, we were asking for at

1:15:31

least at least um you know, three to six

1:15:34

orders of magnitude improvement in cost

1:15:36

per token. Uh get that into 5,000. You

1:15:38

probably have some customers. And in

1:15:40

fact, I would argue that we're for some

1:15:42

size of model, we are approaching a

1:15:44

trillion tokens being measured in, you

1:15:47

know, tens of thousands of dollars. And

1:15:49

that's something that you can imagine

1:15:50

running for a single job.

1:15:51

>> Are you at all worried that just like

1:15:52

the average person just can't and won't

1:15:54

do that like doesn't do that now with

1:15:56

their own brain? Like there actually

1:15:58

isn't that much demand for intelligence

1:16:00

in the world.

1:16:01

>> I never will believe in that. There is

1:16:02

always demand for intelligence in the

1:16:04

world. I think that the way in the

1:16:06

on-ramps to that intelligence are our

1:16:07

challenge as a product u you know

1:16:09

community. I'm not a product person so I

1:16:11

cannot say I had the best vision.

1:16:13

>> You want to enable those people.

1:16:14

>> I want to enable those people. I want

1:16:15

them to never be held back by the sense

1:16:17

that, oh, I my free tier users cannot

1:16:19

use or I can't afford to give them this

1:16:21

many tokens. And I hear that from my

1:16:23

customers all the time. Um, we want to

1:16:25

fix that.

1:16:25

>> What about the inverse question? Not

1:16:27

what you think is craziest, but like

1:16:28

what consensus thing you think is wrong?

1:16:31

>> One of the things I keep coming back to

1:16:32

is this question of Nvidia. I am bullish

1:16:35

on Nvidia in the short term. And you

1:16:38

know, Nvidia, you should never bet

1:16:39

against them. They're always going to

1:16:40

reinvent themselves. But like

1:16:41

fundamentally I think one thing that

1:16:43

surprises people is when I tell them

1:16:44

that hey if you look at you know Hopper

1:16:46

to Blackwell to Reuben and you compare

1:16:49

like for like like what is the

1:16:51

performance per watt of Bloat 16

1:16:53

multiply it hasn't improved all that

1:16:55

much or or even you take that one step

1:16:57

further go to TSMC if you look at TSMC 5

1:16:59

nanometer versus four versus three

1:17:01

versus two the performance per watt on

1:17:03

these chips doesn't change like a

1:17:04

dramatic amount

1:17:06

>> so the consequence of this is people

1:17:08

lose their minds over geopolitics like

1:17:10

what what happen if we lost access to

1:17:11

DMC for any reason. And um my contrarian

1:17:14

take is that it wouldn't be that bad.

1:17:17

Supply would take a shock for sure, but

1:17:19

the best processes that we have in the

1:17:20

west uh like Intel not that far behind

1:17:23

at worst like maybe 2x uh worse

1:17:26

performance per watt and the gap is just

1:17:28

far smaller than than you would make it

1:17:30

out to be if you talk if you follow like

1:17:31

the chipboard dialogue. What else is

1:17:33

happening in the AI world that is not in

1:17:36

your path? Meaning it's not like a

1:17:37

component of this whole system that you

1:17:39

would end up doing something in that

1:17:41

interests you most.

1:17:42

>> Well, we're fully downstream of models,

1:17:44

right? So the model people get to decide

1:17:46

how to design their their architectures

1:17:49

and I have only like very light I mean I

1:17:52

don't have any input to open AAI or

1:17:53

anthropic but um I can only pray that

1:17:56

they go in the direction that is a

1:17:57

minimal to me and the ch like or I have

1:18:00

to like do my best to predict where I

1:18:01

think they're going to go and build my

1:18:02

serving architecture accordingly. Both

1:18:04

software and hardware choices they have.

1:18:06

I think the most interesting game in

1:18:08

some ways to play like once again this

1:18:09

is going back to like the profoundity of

1:18:11

the machine thinking and how

1:18:13

consequential it is to decide to use

1:18:15

something like sparse attention versus

1:18:16

dense attention or um how consequential

1:18:19

it is to like use a different data type

1:18:22

like we were training in B16 but now we

1:18:24

can train in FP8 or FP4 lower precision

1:18:26

data types that is just an arbitrary

1:18:28

choice it feels like but it has profound

1:18:30

implications for what chips I can use

1:18:31

and and you know how I should build my

1:18:33

hardware think about the future of

1:18:34

compute

1:18:35

>> if you had a 100 entrepreneurs in a

1:18:37

room, all of whom wanted to create some

1:18:39

new compute startup.

1:18:40

>> Y

1:18:41

>> um and let's say they were specifically

1:18:42

wanted to make hardware chips or systems

1:18:44

or racks or whatever.

1:18:45

>> What advice would you give them on like

1:18:47

how to orient their companies or like

1:18:49

the type of company, not the specific

1:18:51

choice they're making on a tech tech bed

1:18:52

or something like this

1:18:53

>> because it seems like we're going to try

1:18:55

everything and that will be great for

1:18:56

the world. You know, some stuff will

1:18:57

work. But if you had to give them advice

1:18:59

on how to orient their business to be

1:19:01

successful in this coming world, what

1:19:03

advice would you give them? It's all

1:19:05

about the bottlenecks on supply chain.

1:19:07

So, you need to first convince me or

1:19:08

convince an investor that you understand

1:19:10

the like three to five bottlenecks that

1:19:12

dictate modern chip supply. There's TSMC

1:19:14

wafer capacity, there's HPM capacity,

1:19:16

and there's um like advanced packaging,

1:19:18

and maybe a fourth one would be power.

1:19:20

Like, where will you get the power? How

1:19:22

will you build these racks? Uh and I I

1:19:24

want to hear like you should have a

1:19:26

great answer to each of those four

1:19:27

bottlenecks and how you're going to work

1:19:29

around them because it's all arbitrage

1:19:31

at the end of the day. You're building a

1:19:32

chip because you think that Nvidia has

1:19:34

made some choices that are difficult for

1:19:36

them to change, which is true. Nvidia

1:19:38

makes a lot of choices that are

1:19:38

difficult for them to change. They're

1:19:40

not perfect. They're just really well

1:19:41

balanced. And so, you want to be spiky.

1:19:43

You want to pick something and say, I

1:19:46

think they've underpriced the impact of

1:19:47

how short we're going to be on HBM.

1:19:49

We're going to push really hard in this

1:19:50

other direction instead. Which, you

1:19:52

know, as a as an aside, I do think is

1:19:54

probably the thing to attack most.

1:19:55

>> Why? There's no easy way to bring on a

1:19:58

lot more fabs of memory and those guys

1:20:01

have been

1:20:02

>> so it's going to be a while until we

1:20:04

>> Yeah. Yeah. The boys in Boise don't uh

1:20:06

don't love huge capex for for cyclical.

1:20:10

>> They they've been burned on that many

1:20:11

times.

1:20:12

>> But conceivably like because of that

1:20:13

shortage, the world is just going to

1:20:14

route around it by making everything

1:20:17

else in the system more efficient.

1:20:18

>> I think they're gonna make everything

1:20:19

else more expensive. Think that iPhones

1:20:21

will cut their memory. iPhones are going

1:20:22

to go up in price and um we're just

1:20:24

going to deal with it. Why doesn't

1:20:25

Nvidia go all the way to the end and

1:20:28

sell tokens? Do you think

1:20:30

>> Nvidia is really smart about this? They

1:20:31

don't compete with their customers.

1:20:33

Nvidia takes the long view on

1:20:34

everything. Um, why don't they even

1:20:36

start with the Neocloud? Why don't they

1:20:38

just sell computer out the back door?

1:20:40

Well, Nvidia is really good. Jensen is

1:20:42

really good at making his friends

1:20:43

billionaires. He's made Cororeweave a

1:20:44

billion dollar company, many billion

1:20:46

dollar company. And there's no need for

1:20:48

him to kind of uh destroy that goodwill.

1:20:50

like he wants to create a diverse

1:20:52

community of NeoClouds and inference

1:20:54

providers who are all jockeying to

1:20:55

create demand for Nvidia such that if

1:20:58

any one of them decides to I don't know

1:21:00

vertically integrate or go with AMD or

1:21:02

any other option he's got three more

1:21:04

people ready to hungry to fill that

1:21:06

position.

1:21:07

>> It's great to have competition amongst

1:21:08

his buyers.

1:21:09

>> My favorite closing question for

1:21:10

everyone is what is the kindest thing

1:21:12

that anyone's ever done for you?

1:21:13

>> The kindest thing I mean I my immediate

1:21:15

first thought is like all the mentors

1:21:17

that I've had over the years. It's a

1:21:18

rare person who takes a lot of time out

1:21:20

of their their schedule and um and you

1:21:23

know makes it like their personal

1:21:24

interest essentially to to make sure

1:21:26

that you understand something that uh or

1:21:27

teach you something or or like ingrain

1:21:30

some value in you that they think that

1:21:32

you're on the cusp of understanding but

1:21:33

just push you over the line for

1:21:35

understanding. uh a lot of the people in

1:21:36

Nvidia that I mentioned earlier who

1:21:38

instilled that like love of performance

1:21:39

engineering in me but also my professors

1:21:41

in college who I remember like my

1:21:44

adviser in like sophomore year I was

1:21:46

very impatient student so I show up at

1:21:47

his office hours and say like I I want

1:21:50

to build AI chips I know what I want to

1:21:51

do why am I wasting time taking all

1:21:53

these like other basic classes and

1:21:54

networking and you know operating

1:21:56

systems and he just looked at me and

1:21:57

said like you know he laid out basically

1:21:59

like the whole stack and showed me the

1:22:01

depth of or the beauty of like

1:22:04

understanding every piece in the puzzle

1:22:05

like he he took my entire path of like

1:22:07

trying to focus on one piece of the the

1:22:09

system and said that you know it's so

1:22:11

rare that someone can actually

1:22:12

understand the entire stack from the

1:22:14

gate level silicon all the way to

1:22:17

building a great internet scale service

1:22:19

and you know you should aspire to be

1:22:21

someone who over the course of your

1:22:23

lifetime achieves that level of

1:22:24

understanding.

1:22:25

>> It is such a rare rare trait and um you

1:22:28

know it that level of expertise is so

1:22:30

noble to chase and I think and that

1:22:31

stays with me quite a bit. Not a common

1:22:33

but an advisory. Neil, amazing

1:22:35

conversation. Thanks so [music] much for

1:22:36

your time.

1:22:37

>> Thank you so much for having me.

1:22:42

>> You know how small advantages compound

1:22:44

over time? That's true in investing and

1:22:46

just as true in how you run your

1:22:47

company. [music] Your spending system is

1:22:49

your capital allocation strategy. Ramp

1:22:51

makes it smarter by default. Better

1:22:53

data, better decisions, better economics

1:22:55

over time. See how at ramp.com/invest.

1:22:59

As your business grows, Vanta scales

1:23:01

with you, automating compliance and

1:23:02

giving you a single source of truth for

1:23:04

security and risk. Learn more at

1:23:06

vanta.com/invest. [music]

1:23:08

The best AI and software companies from

1:23:10

OpenAI to cursor to perplexity. Use work

1:23:12

OS to become enterprise ready overnight,

1:23:14

not in months. Visit works.com [music]

1:23:17

to skip the unglamorous infrastructure

1:23:18

work and focus on your product.

1:23:21

Ridgeline is redefining asset management

1:23:23

technology as a true partner, not just a

1:23:25

software vendor. They've helped firms 5x

1:23:27

and scale, enabling faster growth,

1:23:29

smarter operations, and [music] a

1:23:30

competitive edge. Visit ridgeland.ai to

1:23:32

see what they can unlock for you.

Continue with YouTLDR

Analyze another video with Pro

Process a new video, search every timestamp, compare sources, and keep the result in your library.

Get Pro — $12/month30-day money-back guarantee

More transcripts

Explore other videos transcribed with YouTLDR.