Full Transcript

·YouTLDR

DuckDB-Quack announcement at AI Council

31:351,084 summary words · ~5 min readEnglishBy DuckDBTranscribed Aug 26, 2026
Analyze another video with Pro30-day money-back guarantee
Summary

DuckDB Labs has unveiled Quack, a native extension that gives DuckDB a lightweight, high-performance client-server protocol over HTTP to let DuckDB instances query and replicate across each other.

It breaks DuckDB out of its single-node, in-process isolation without sacrificing analytical speed, positioning it as a general-purpose database capable of handling real-time ingest and multi-tool workflows.

Section summaries

0:05-2:50

DuckDB Milestone Recap & Adoption Metrics

optional

The speaker opens by humorously reflecting on DuckDB's growth from a university research project into a widely adopted tool. He highlights monthly download stats exceeding 40 million for the Python package alone, alongside over 160 million monthly extension downloads across edge devices, embedded environments, and local data workflows.

  • DuckDB Python downloads exceed 40 million per month, averaging over 1 million downloads daily.
  • Extensions are downloaded more than 160 million times monthly across global installations.

Provides community context and adoption statistics before the main technical announcement.

2:50-7:06

DuckLake & Ecosystem Extension Growth

optional

The talk reviews milestones from the past year, including the introduction of DuckLake as an alternative to complex data lake formats like Apache Iceberg and Delta Lake. Download numbers show DuckLake's extension holding 2.5 million monthly downloads alongside Iceberg (2.9M) and Delta (2.3M) shortly after the DuckDB 1.0 release.

  • DuckLake simplifies lakehouse architectures by maintaining metadata inside a relational database rather than complex file manifests.
  • DuckDB's data lake extensions compete evenly in volume with dedicated Apache Iceberg and Delta Lake integrations.

Details previous lakehouse efforts and extension statistics leading up to the new reveal.

7:06-11:58

The In-Process Limitation & Real-Time Ingestion Problem

watch

The speaker addresses DuckDB's primary architectural bottleneck: while it connects seamlessly to Postgres, SQLite, and cloud object stores, it could not communicate with other DuckDB instances. He details common pain points, including simultaneous multi-tool access on localhost and real-time streaming telemetry from multiple fleet nodes.

  • Pure in-process databases struggle with real-time analytics scenarios involving frequent small inserts from external agents.
  • Local developers frequently encounter concurrency locks when attempting to connect GUIs (e.g., DBeaver) and CLIs to the same DuckDB file.

Crucial context explaining the technical motivation and architectural trade-offs behind building a client-server extension.

11:58-15:12

Introducing Quack: Architecture & Client-Server Mechanics

watch

Quack is introduced as a native extension that implements client-server communication between any two DuckDB nodes using standard SQL syntax. The speaker shows how a server instance exposes local tables via HTTP and how remote nodes connect using ATTACH statements or explicit remote queries across desktop and WebAssembly runtimes.

  • Both client and server ends run standard DuckDB instances without requiring dedicated middleware.
  • Supports query execution directly on remote nodes or local catalog execution through familiar SQL syntax.
  • Available in beta starting with DuckDB 1.5, targeting a stable production release in DuckDB 2.0.

Covers the core reveal, syntax, and fundamental operational model of the Quack extension.

15:12-20:04

Protocol Internals: HTTP, WAL Serialization & Security

watch

The underlying architecture of Quack relies on HTTP/TLS over TCP/IP to ensure firewall friendliness and direct browser/WASM compatibility. Queries and data chunks are serialized losslessly using DuckDB's internal write-ahead log (WAL) structures, while authentication and authorization are managed through pluggable extension callbacks and SQL-defined rules.

  • HTTP was chosen over proprietary protocols due to WebAssembly browser support, firewall traversal, and hardware acceleration.
  • Reuses DuckDB's internal WAL serialization engine for zero-data-loss transmission of data types and vector chunks.
  • Security, token handling, and authorization rules are completely pluggable via extensions or custom SQL callbacks.

Explains critical design choices around networking protocols, data serialization, and security integration.

20:04-24:30

Performance Benchmarks: Bulk Transfers & Small Inserts

watch

The speaker presents AWS benchmark tests pitting Quack against PostgreSQL's wire protocol and Arrow Flight SQL across identical EC2 instances. Quack transferred 60 million records in ~5 seconds compared to 20 seconds for Arrow Flight and 3 minutes for Postgres, while achieving 5,000 transactions per second under multi-threaded single-row insert workloads.

  • Quack delivers an 36x speedup over PostgreSQL wire protocol and 4x over Arrow Flight SQL for bulk analytical row transfers.
  • Achieves 5,000 single-insert transactions per second across 8 concurrent client threads.
  • Enables advanced distributed topologies like proxy coordination, database sharding, and planned WAL-based read replication.

Provides essential empirical data proving performance gains across analytical reads and concurrent writes.

24:30-27:38

The OLAP vs OLTP Paradigm Shift

watch

The presentation explores database classification, challenging the standard spectrum that places Postgres exclusively in OLTP and DuckDB in OLAP. By integrating concurrent transactions, background checkpointing, and Quack networking, DuckDB aims to become a general-purpose database system originating from the analytical domain.

  • PostgreSQL operates as a general-purpose compromise rather than an extreme OLTP-optimized engine like TigerBeetle.
  • DuckDB is expanding toward general-purpose utility by adding transactional concurrency without degrading raw analytical performance.

Delivers the strategic thesis for DuckDB's long-term roadmap and positioning in the broader database landscape.

27:38-31:32

Live Demo: In-Browser WASM Querying EC2 & Conclusion

watch

A live demonstration displays a DuckDB instance running inside a web browser connecting directly to a remote EC2 micro-instance over Quack. The demo highlights the latency difference between pulling raw remote tables versus pushing queries directly to the server, concluding with the public release announcement.

  • Browser-based DuckDB WASM instances can attach and query remote cloud databases directly via HTTP secret tokens.
  • Pushing execution via remote query functions drastically minimizes bandwidth compared to full catalog table pulls.

Demonstrates practical real-world usage and query optimization techniques live.

Key points

  • Quack Enables Native DuckDB-to-DuckDB Networking — Quack is an extension that acts as a client-server protocol where both ends run DuckDB directly, allowing remote attaching and query execution without separate server daemons.
  • HTTP and Write-Ahead Log Serialization for Universal Connectivity — Built atop HTTP with TLS, Quack utilizes DuckDB's native internal WAL serializer to pass types, vectors, and chunks losslessly across networks and WebAssembly (WASM) browser clients.
  • Superior Throughput for Bulk Transfers and Concurrent Writes — Benchmarks demonstrate Quack transferring 60 million rows across AWS instances in 5 seconds (compared to 3 minutes for Postgres wire protocol and 20 seconds for Arrow Flight SQL) and reaching 5,000 insert transactions per second with 8 clients.
  • Pushing DuckDB from Pure OLAP to General-Purpose Database — By pairing Quack with engine improvements like concurrent transactions and background checkpointing, DuckDB is transitioning toward a general-purpose database model spanning analytical and transactional workloads.
DuckDB is a universal data wrangling tool. That's how we like to call it because you can kind of use it for everything under the sun as long it is like vaguely tabular. Hannes Mühleisen
At some point it's not sort of I have to sort of take off my academic hat of being right, which of course I love being right. But at some point it's about solving people's problems. Hannes Mühleisen

AI-generated from the transcript. May contain errors.

0:05

All right. Hello everybody. This is

0:08

quite interesting.

0:11

Thank you so much for the invitation.

0:13

And I'm still amazed I got away with

0:16

this talk proposal, right? Because I

0:18

just basically said this.

0:20

And

0:22

that's what we wanted to surprise you

0:23

and it wouldn't have been a surprise if

0:25

it would have been in the program half a

0:26

year ago. You probably understand this.

0:29

Let me start with

0:31

you know the story so far.

0:33

And you know a small recap like in a in

0:36

a good in a good TNG episode two parter,

0:39

you know, we start with a previously.

0:41

I see my people are here. We're good.

0:43

We're going to have fun today. This is

0:44

this is great.

0:46

So DuckDB, right?

0:48

I've spent the last decade or so working

0:50

on DuckDB together with many other

0:52

wonderful people.

0:53

And

0:54

since we're always looking for

0:55

superlatives, you know, it's the

0:57

friendliest SQL database. I think that's

0:59

fair, right? friendliest.

1:01

We're not going to you know get into

1:02

this whole like you know

1:04

game of who has the fastest and

1:06

whatever. You know, cuz it's the

1:07

friendliest.

1:08

DuckDB is a universal data wrangling

1:11

tool. That's how we like to call it

1:12

because you can kind of use it for

1:13

everything

1:15

under the sun as long it is like vaguely

1:17

tabular, right?

1:18

It's actually used quite everywhere.

1:21

It's kind of wild to see that. I'm still

1:24

totally amazed that our like little

1:25

project from from university became such

1:27

a big thing.

1:29

It runs in space. It runs in a battery.

1:32

It can run fully local in a process.

1:34

It's kind of ideal for tasks that you

1:37

don't you know trust your intern with,

1:39

right? Like we have a lot of these tasks

1:41

where we kind of send the intern to do

1:43

it and then we hope that they don't blow

1:46

up the database. Well, DuckDB is

1:47

actually quite useful for this because

1:49

you can just run all of that locally.

1:52

And you don't end up with a massive

1:53

Snowflake bill. It's great.

1:56

Sorry. I mean I I I like the guys a

1:58

snowflake, but I mean, you know what I

1:59

mean. Um

2:01

DuckDB has actually quite crazy adoption

2:04

at the moment. So, we are at the moment

2:07

running at about 40 million downloads

2:09

per month of DuckDB itself, and that's

2:10

just a Python package. Like, that

2:12

doesn't cover anything else. It's just a

2:14

Python package. It's just the one that

2:16

we have the best metrics on.

2:18

Um and uh that's that's kind of crazy.

2:21

That's like more than 1 million each

2:23

day. And we have

2:25

Also, what's kind of crazy to see is we

2:27

have these extensions, which are plugins

2:28

for DuckDB.

2:30

And those are downloaded like 160

2:32

million times each month, which is also

2:35

totally wild. Like, it's like I'm

2:37

clearly all these downloads must be

2:38

humans, right?

2:40

We're absolutely sure.

2:42

We run out of Americans in about 10

2:44

months, right?

2:46

>> [laughter]

2:47

>> Anyways.

2:48

Um

2:50

But, we went beyond DuckDB. Actually,

2:51

this was from almost exactly 1 year ago.

2:54

I stood on the stage of you know, the

2:57

this conference and hinted in my last

2:59

slide at something new for DuckDB. Um

3:02

also, I only have one set of clothes,

3:04

right?

3:05

Um

3:08

And so, from this logo, um obviously,

3:10

the logo has slightly changed, but we

3:12

announced actually

3:13

DuckLake

3:15

um about a week after the conference.

3:17

And now you know this is DuckLake.

3:18

DuckLake is a

3:20

is of course a lakehouse format that

3:21

kind of admits that metadata is best

3:24

stored in a database.

3:26

And it was released

3:28

um the preview in May, like about a year

3:31

ago.

3:32

And the logo has slightly changed since

3:34

the preview, but that's fine.

3:36

Um

3:37

DuckLake

3:39

is Yeah, it's I mentioned it's the

3:40

simplest way to a data lake. And this is

3:42

like a bit of a you know, aggressive, if

3:44

you want, slide from us, where we just

3:46

show the sort of the bits and pieces of

3:49

Iceberg, the you know, that lakehouse

3:51

format lots of people use. On the left

3:53

side where you see, oh, we need to have

3:54

this catalog and then we have metadata

3:56

files and manifest lists and manifest

3:58

files and data files and it's all a bit

4:01

complex.

4:02

Um, and then on the right side we have

4:04

DuckDB which again admits that metadata

4:06

probably should be stored in database

4:08

and then like you have a much more

4:10

simpler structure.

4:11

And

4:13

this has been hitting quite a nerve. We

4:15

actually expected criticism

4:17

but we got an overwhelmingly positive

4:19

response. Like lots of praise. So things

4:21

like this where somebody said, "DuckDB

4:23

and DuckDB Lake, why we bet the company

4:25

on the Duck stack?"

4:27

Kind of like Duck stack. It's uh

4:29

And this was on a preview. Can you

4:31

imagine, right? Like we we bang

4:32

something out which is a preview and

4:35

then people say, "Now we'll let's bet a

4:36

company on it. It's great." Right? And

4:38

this is only the kind of things that you

4:39

can I can talk about in public, right?

4:41

There's loads of loads of

4:42

insane insanely positive response to to

4:45

DuckDB. Clearly hit a nerve. Great.

4:48

Um, just a month couple of

4:50

weeks ago, a month ago we released

4:52

DuckDB 1.0.

4:54

Um, so the first version of DuckDB that

4:56

we think is

4:57

production ready. Um, this time we

4:59

actually got press coverage which was

5:00

nice.

5:01

Um, and you know, you people have then

5:04

said, "Okay,

5:06

uh, DuckDB is this thing that only the

5:07

DuckDB people love and and you know, the

5:10

the real people use use something else."

5:12

But we actually mentioned earlier we we

5:15

can kind of count how many times people

5:17

download the extensions.

5:18

And DuckDB has extensions for Iceberg,

5:21

for DuckDB and for Delta Lake, right?

5:23

And so, you know, in in in your mind

5:26

imagine how these kind of different

5:27

systems stack up in terms of downloads,

5:29

right? Just make a mental picture.

5:31

And if you look at this, you can

5:33

actually see that within less than a

5:36

year actually and only 4 weeks after

5:37

declaring 1.0, there's already like

5:40

DuckDB is 2.5 million downloads a month

5:43

at the moment for the extension and the

5:45

Iceberg extension is 2.9 and the Delta

5:47

is 2.3, but it's to us totally amazed

5:50

amazing to see how many people are

5:52

really

5:53

happy uh with with this DuckDB thing.

5:56

And yeah, so this is again, this is of

5:58

course only the respective downloads

6:00

within uh the DuckDB ecosystem.

6:03

Uh yeah, within a year. But

6:06

that's not what I want to talk about

6:08

today.

6:09

And now

6:11

um the clicker stopped working.

6:15

Isn't it in a very dramatic pause,

6:17

right?

6:18

Clicker stops working in the in the in

6:20

our the continuation of the of the of

6:21

the Star Trek episode, obviously. Um

6:23

today is really not about DuckDB, and I

6:25

have to admit that we also worked on

6:27

other things this last year. Uh only

6:29

maybe less public than uh DuckDB. And if

6:32

you think about it, it's kind of an

6:34

insane speed for any sort of database

6:36

project to sort of run these gigantic

6:38

initiatives and then kind of finish them

6:40

within a year and do something else in

6:42

the meantime. I I'm I'm still amazed and

6:44

actually really proud of our team to be

6:46

able to do that. And actually Pedro, one

6:48

of our guys, ended up taking over DuckDB

6:51

so that I could do something else. And

6:52

then uh other people have thankfully uh

6:55

taken over some project management, so I

6:57

could do something else. So this is

6:58

something that I have actually worked on

6:59

personally, and so I'm kind of nervous

7:02

to talk to you about it today, okay?

7:04

Let's look at this. Okay.

7:06

New thing. What's the new thing?

7:09

Okay, of course

7:11

single node DuckDB works great, right? I

7:13

would I would even say single node

7:15

DuckDB works extremely great, right? You

7:17

have a single computer.

7:19

We have done experiments. You can go to

7:21

gigantic data sizes. Basically, it will

7:23

just continue crunching stuff until you

7:25

run out of disk space, which is pretty

7:26

good um

7:28

sort of um boundary to have.

7:31

Okay, so this is great.

7:33

Little check mark.

7:35

DuckDB can also talk to almost

7:37

everything under the sun, right? As it

7:38

can talk to Postgres, for example. Oh,

7:40

we just you have a extension that is the

7:42

post scanner. There's my sequel one,

7:44

there's a sequel light one.

7:45

We have run we can talk to random ODBC

7:48

drivers and integrate with and just pull

7:51

data back and forth between these

7:52

things. Works really well. Okay.

7:55

Uh DuckDB can also talk to object

7:57

stores, right? We have integrations for

7:59

all the wonderful object stores. You can

8:01

all the data formats we have support for

8:03

all the you know, parquet and its

8:06

various competitors. All works really

8:07

well. You can

8:08

read and write from the object store. We

8:10

have integrations with catalogs like S3

8:12

tables. It's all it's all wonderful.

8:15

Little check mark. Works well.

8:18

But

8:19

somehow

8:21

DuckDB cannot talk to DuckDB.

8:23

Well, that's a bit annoying, isn't it?

8:25

And we are not the first people

8:27

to notice this. There's actually been a

8:30

quite

8:31

uh

8:32

large amount of GitHub repos and I think

8:34

what about we are running at about one

8:35

per week that's popping up where people

8:38

are um just adding functionality for

8:40

DuckDB to talk to DuckDB.

8:42

And here are just four of them. There's

8:43

many more.

8:45

And

8:46

maybe it's also, you know, something

8:47

that is maybe quite easy to to wipe code

8:49

in an afternoon. That's why they're

8:50

popping up so many. But it's it's it's

8:52

just one of these things where

8:54

yeah, we we really we

8:57

we didn't have a strong sort of feelings

8:58

about it or maybe we did. And then all

9:01

these people out there do things that

9:03

that kind of solve this problem of

9:05

DuckDB to talk to DuckDB. And at some

9:06

point it's not sort of I have to sort of

9:08

take off my academic hat

9:10

of being right, which of course I love

9:12

being right.

9:13

But at some point it's about solving

9:15

people's problems, right?

9:17

And clearly there is a need out there

9:19

because people are people are building

9:21

all these things.

9:22

Um at this point I also have to admit

9:25

that 1 year ago I was standing here

9:28

and talked about the benefits of single

9:30

node in-process databases, right?

9:33

Yeah, yeah.

9:35

And while I fully admit that it is still

9:37

a great idea, that I mean I will still I

9:39

will still think that making a single

9:41

node database is a great idea.

9:43

There are just lots of use cases out

9:45

there

9:46

where a in-process database cannot

9:48

cover, all right? Like if you're an

9:50

in-process database, there's just some

9:51

things that you cannot do.

9:53

And I'm going to show you one example,

9:56

um

9:57

which is uh this sort of the real-time

9:59

analytics use case, right? You have you

10:00

got you have a fleet of nodes. They they

10:03

do whatever, and they try to store some

10:05

for example some telemetry in a central

10:07

sort of authoritative location.

10:10

And it's a flood of sort of fairly small

10:12

inserts.

10:13

So, what do you do now with your

10:14

in-process database, right?

10:16

It's really difficult to to map this

10:18

um to an in-process database, and it's

10:20

unfortunate because it excludes

10:23

a lot of really exciting use cases,

10:25

right?

10:26

And

10:28

also maybe from the academic side that

10:29

we have to kind of admit that this is

10:31

also part of analytics. Like I think for

10:33

the longest time

10:35

uh the academic world has kind of

10:36

treated change in analytical databases

10:38

as like a you know a

10:40

something for other people to fill to

10:42

solve, but we have to kind of we have to

10:44

kind of get around to that.

10:46

And the crazy thing is that this even

10:47

happens

10:48

on localhost. Like people have multiple

10:51

tools they want to talk they want to

10:53

have talk to the same database. Like you

10:55

have For example, you have like a

10:57

DBeaver or DataGrip running on a DuckDB

10:59

database, and then you want to use the

11:00

CLI to do something else. You want to

11:02

start up the DuckDB UI, or I don't know.

11:04

Like this has been a a sort of

11:06

long-standing customer complaint, or I

11:08

mean feature request.

11:10

Sorry.

11:11

Um I should I should get my corp speak

11:13

better.

11:14

Uh but you can kind of do this already

11:16

with DuckLake, but this is quite a lot

11:18

of additional operational complexity.

11:20

And performance for small inserts is

11:22

really like these these lakehouse

11:23

formats are really not made for tons of

11:26

tiny inserts. So, that's kind of

11:27

annoying.

11:29

So, it says breeze on my presenter notes

11:31

here.

11:33

So, that was the motivation. And now I'm

11:35

can I can show you something that makes

11:37

me really happy to show you.

11:39

So,

11:40

I mean, this is just the happiest duck,

11:42

right?

11:43

I was I found this picture and I was I

11:45

was so happy that I get to show you this

11:46

little happy ducky. But, so

11:49

let's say we want to have ducks talk to

11:50

each other.

11:52

What do they do?

11:54

They quack.

11:56

Exactly. Today I I as I'm very very

11:58

happy to to show you Quack. And Quack is

12:01

a DuckDB extension that basically

12:04

extends DuckDB with capabilities with

12:06

quite powerful capabilities and I'm

12:07

going to show you all about it um to

12:10

communicate with other Duck DBs.

12:13

Um and maybe we can do like a a favor

12:15

and we can quack together. So, we do

12:16

this in our company quite a lot. So, can

12:18

do quack quack quack quack quack. Come

12:21

on. Quack quack quack quack quack quack

12:24

quack. Excellent. You made me very

12:25

happy. Uh

12:27

So, what is Quack anyway?

12:29

Um

12:30

So, this is how it looks like. So, Quack

12:32

is kind of a client-server protocol in

12:34

the in a sort of traditional sort of

12:35

sense for DuckDB. So, it is as I

12:37

mentioned implemented as a DuckDB

12:39

extension.

12:40

But, the cool thing is that both sides

12:43

are just DuckDB. So, there's not like a

12:45

separate there's not a separate server.

12:48

There's not a separate client. It's just

12:50

DuckDB. So, you have two DuckDBs. Here

12:52

we have two. Like one is green and one

12:53

is blue.

12:55

Um and on left you can say, "Okay,

12:57

please serve this local database that

12:59

I'm currently in

13:01

on this uh local host and then maybe

13:03

you're creating a table there." And on

13:05

the other side

13:06

you say attach, which is the DuckDB way

13:08

of

13:09

connecting to other things. Like we had

13:11

this for a while. For example, you can

13:12

connect attach a Postgres or you can

13:14

attach a SQLite or a Duck Lake or

13:15

whatever.

13:16

And then you can just type for example

13:18

from remote foo and then magically

13:22

the query will be run on the other side.

13:24

The result will be shipped back

13:25

and be displayed on the blue side,

13:27

obviously.

13:28

The same you can also if you want you

13:29

can also have a sort of explicit uh

13:32

query there where you say from

13:34

remote.query from foo and we actually

13:36

working on

13:38

something really cool there with uh

13:40

DuckDB 2.0 uh to um make this really

13:43

nice in a user interface.

13:46

So, yeah, this is this is really a

13:47

minimum example, but you can do anything

13:50

reachable from SQL, right? Everything.

13:52

All the extensions.

13:54

All the data types, all the all the

13:56

everything it just works, right?

13:59

Here is actually a slightly bigger uh

14:01

example because this actually works

14:04

already.

14:05

So, if you today if you have DuckDB 152,

14:08

you can just run this script and it will

14:10

already do all of this. So, you have

14:12

client server um right right off the box

14:15

in DuckDB 1.5. Um this works for all

14:18

DuckDB distributions, for Python, for

14:19

shell, for web assembly.

14:21

Yeah, you name it, right?

14:23

Like Windows, OSX, Linux, we don't care.

14:25

So, this is being released today as a

14:27

sort of beta.

14:29

And we expect the production release in

14:31

a couple of months uh with DuckDB 2.0 uh

14:35

which are which coming in fall, by the

14:36

way, in case you haven't heard yet.

14:38

Um and yeah, as I mentioned, we're going

14:40

to do um some new simplification some

14:42

simplification on the way that you

14:44

specify queries that should be run

14:46

remotely which I should give you a

14:47

really um great sort of experience in

14:50

writing queries for other servers. It's

14:52

just that it's not there yet and I

14:53

didn't want to show you syntax that

14:54

doesn't work today. So,

14:56

here you go.

14:57

So, let me talk a bit about how Quack

14:59

works internally.

15:00

So, this is like, you know, this is kind

15:03

of how it looks like for you or for your

15:05

agent.

15:06

Um but how does it actually designed?

15:08

And this is something that

15:09

uh we did kind of ourselves. So,

15:12

on the bottom is TCP/IP. Obviously, we

15:14

have no choice but to use TCP/IP to talk

15:16

to other servers, right?

15:17

Uh we're not using UDP. Uh it's TCP/IP.

15:21

And then what is what we have on top?

15:22

Any guesses?

15:25

Hm?

15:27

Any Any guesses what we put on TCP IP as

15:29

a protocol? HTTP.

15:32

Excellent. So, this is um we need a

15:34

protocol on top of HTTP because of WASM.

15:36

Like, in case you don't know, but DuckDB

15:38

runs in the browser, right? And we do

15:40

want to talk to a server from the

15:43

browser. And so, we really don't have a

15:44

lot of choices

15:46

uh to to not use anything else. And this

15:48

is also really great to use HTTP for

15:50

database protocol in that you build in

15:52

2026

15:54

because

15:55

all the firewalls like it, the cyber

15:57

people don't freak out. Um you can just

15:59

slap uh TLS on it and it's going to be

16:01

encrypted. Um and it's actually quite

16:04

impressive. We've done some done some

16:05

benchmarks

16:07

because HTTP is so ubiquitous, all the

16:10

hardware actually is optimized for it.

16:12

So, if you if you if you use a different

16:14

it's it's really counterintuitive, but

16:16

because of you know, I guess because of

16:18

YouTube,

16:20

HTTP is faster than anything else. So,

16:22

it's actually by just because the

16:23

hardware prioritizes and just knows

16:25

better how to deal with it. It's kind of

16:26

funny. So, if you have an an another

16:29

protocol, it's it's actually going to be

16:30

slower. Um and uh yeah, it's kind of

16:33

it's kind of a weird counterintuitive

16:34

thing.

16:36

Okay, so what do we do on top of HTTP?

16:39

We actually use our We need a

16:40

serialization format. We need something

16:42

to basically encode queries and results

16:44

and all that stuff. And so, what we we

16:47

have that already in DuckDB, right? We

16:49

have a serializer because we need to

16:50

serialize things to the write-ahead log.

16:52

And let's just use that. So, we can use

16:55

lossless serialization

16:57

of all the internal structures like

16:58

types, columns,

17:01

data, chunks, all that stuff. It's

17:03

really lossless serialization. And on

17:05

top of that, we just put uh we need some

17:08

interaction protocol.

17:10

All right? And it's really

17:11

straightforward. It's just messages with

17:13

different types like every other

17:14

database RPC out there. So, it's a

17:16

request-response pattern. So, you can

17:18

for example imagine that you can execute

17:20

a statement, fetch more results, and so

17:22

on and so forth, right?

17:24

So, we have a

17:26

RPC, kind of a client-server protocol,

17:28

all these things, and now we have some

17:30

sort of side requirements. For example,

17:34

everybody's favorite. Who here loves

17:35

this?

17:37

No one. Okay, me neither.

17:40

Um

17:41

So, it's really it's really interesting

17:43

because once you start I'm going to come

17:44

over here because I see you people, you

17:45

know, like I love you, too. I just I'm

17:47

not always over there. It's just my my

17:48

notes are over there, so I'm

17:50

um

17:51

This is uh one of these things where now

17:54

that we have a client-server, we cannot

17:55

really hide behind our in in-process

17:58

authentication model anymore, right? We

18:00

have to do authentication, which is

18:01

terrifying,

18:02

right? Um but we also realized that we

18:05

cannot solve this for everybody. So,

18:07

what we have done is we have

18:09

basically made this pluggable through

18:11

DuckDB extensions. So, basically you,

18:13

anyone, can

18:15

override the authentication method that

18:17

that uh Quack has with whatever you

18:19

think is best. There is a default one,

18:22

which is based on tokens. I'll show you

18:23

in a bit,

18:24

but it's basically something that you

18:26

can decide what you want. If you want to

18:28

glue this to your

18:29

um you know, uh

18:31

LDAP server or so, you can do that,

18:34

right? Um I think it's really exciting

18:36

when you have such a community like

18:37

DuckDB has. Um there is this community

18:41

of DuckDB extension writers that have

18:43

already written many extensions.

18:46

We can kind of rely on them to to um to

18:48

just cover some more use cases, and of

18:50

course, if you're a big corp, you can

18:52

just make your own.

18:53

Um you can also just write a SQL

18:55

function, by the way, if you want. You

18:56

can if you can you express your

18:58

authentication in SQL, you can just

18:59

write a SQL function. It's fine.

19:01

So, that's cool.

19:03

And then the other thing, the even more

19:04

fun thing,

19:06

is the authorization, right? This is

19:07

even harder. It's like now we've

19:08

authenticated somebody, how do we

19:11

decide

19:12

what the permissions are going to be,

19:13

right? Like you have table level, column

19:15

level, row level, blah blah blah. It's a

19:17

huge feature set. Uh people expect a ton

19:19

of stuff there. And uh we again we have

19:22

managed to avoid it so far. Again, we

19:24

don't really

19:27

want to solve we're not don't want to

19:28

solve this once and for all. We're going

19:29

to give you the flexibility to do your

19:31

own thing. And you there's going to

19:33

there's also a callback for that where

19:34

basically

19:35

you have a callback that sees what the

19:37

user wants to do. You can decide whether

19:39

you want to allow it or not. We can even

19:40

modify the queries that the user is

19:42

running. There's lots of interest.

19:43

There's some really powerful things

19:44

there. And this is I think interesting

19:46

in Quack is that it's

19:48

it's kind of giving you the tools to

19:50

build something exciting. It's not

19:51

trying to solve everything from the

19:53

get-go, right?

19:56

All right. But

19:58

experiments. Because I'm a database

20:00

person and I cannot live without showing

20:01

a performance experiments, right? It's

20:03

very important.

20:04

So in case you don't know um

20:07

many years ago in 2017 we wrote a paper

20:09

how terrible database client-server

20:11

protocols are.

20:12

Um here you see the plot on this on this

20:14

paper where basically every database

20:16

protocol was worse than netcat by a

20:18

factor of 10.

20:20

It's not so great. Um and actually the

20:22

insights from this paper um

20:25

made us really strong believers in in

20:26

process. But now you know we had to

20:28

reconsider. But it is of course extra

20:30

terrifying to build a client-server

20:32

protocol when you have been the one

20:34

that's been banging out like has been

20:36

complaining about everybody else's

20:37

protocol, right? So we better get this

20:39

right. Okay.

20:40

Okay, so this is the 2017 um

20:43

2017 state of the world.

20:46

Um so let's do some experiments. And the

20:48

experiment is really simple. We have a

20:49

client and server.

20:50

And these are

20:52

uh running in AWS. These are virtual

20:54

machines. Um they are like fairly small.

20:56

They have

20:57

32 GB of RAM and eight CPUs, right? But

21:00

they do have fairly fast network. They

21:02

have 15 GB GB per second networking. And

21:06

they are in the same availability zone,

21:08

right? Like what you would have in your

21:09

in your like own

21:11

sort of setup.

21:12

These are different, and we tried three

21:14

different things.

21:15

So, we tried Postgres, because everybody

21:17

likes Postgres.

21:19

We tried Quack, of course, because of

21:21

course we tried Quack. And we tried

21:23

Arrow Flight SQL. There is So, the Arrow

21:25

project has also created something

21:28

to basically deal with, you know,

21:30

database protocol interactions. It's

21:31

called Flight SQL, and we actually use

21:33

something called Gizmo data, which is

21:34

one of the projects that I've mentioned

21:36

earlier that has shown us that this is

21:37

really something people want.

21:39

And we have two experiments, bulk

21:41

transfer

21:42

and small inserts. So, let's start with

21:44

bulk transfer. Bulk transfer.

21:47

So, here is the result.

21:50

We've transferred millions of rows of a

21:52

database table. So, it's 100,000 rows, 1

21:55

million rows, 10 million rows, 60

21:57

million rows, and we measured wall clock

21:58

time. Lower is better, obviously.

22:01

So, Quack manages to transfer 60 million

22:03

rows in this is scenario in around 5

22:05

seconds, right?

22:07

And Postgres took 3 minutes.

22:09

So, the Postgres protocol took 3

22:11

minutes, and just hold on to that. It's

22:12

like

22:13

it's uh 5 seconds versus 3 minutes.

22:16

Pretty good. Arrow Flight took 20

22:18

seconds. It's better, but it's of course

22:20

still much slower.

22:22

But now, really the

22:24

the other So, I mean, kind of you kind

22:26

of trusted us to get bulk transfers

22:27

right, right? Like I mean, we have been

22:28

working on this for ages.

22:30

It's

22:32

it's okay.

22:33

Um but something that was that we were

22:35

then we we said, "Okay, we also need to

22:37

do transactions." And especially for

22:38

this use case, where you have many small

22:40

inserts coming from all over the place.

22:42

So, here we have a second experiment,

22:45

where we run single insert transactions

22:47

with an incoming increasing number of

22:49

client threads, right?

22:51

So, 1 2 4 8, and so. And we just

22:54

measured the completed transactions per

22:55

second.

22:56

And we do What we see? Well, we see

22:57

Arrow Flight not doing so well. Um it's

23:00

designed for bulk, and kind of only for

23:02

bulk.

23:03

Um Postgres is doing okay. It's kind of

23:04

was designed for this use case. But

23:07

Quack surprisingly did really well. Like

23:09

we finished something like 5,000

23:10

transactions per second on DuckDB uh for

23:13

eight clients here. So, that's that's

23:14

pretty good. We have optimized the

23:16

protocol a little bit

23:18

to make this work. But pretty amazing to

23:20

see those results.

23:21

So, it's fast both for bulk and for

23:24

um

23:25

for small transactions.

23:28

So, what can you do with this now? Well,

23:30

you can do lots of things. Here's the

23:31

classic case. You can use Quack to glue

23:33

two DuckDBs together. Great.

23:35

Um

23:36

But what cool What's really nice about

23:37

this is you have DuckDB on both sides.

23:40

So, you can do really cool things like

23:41

post-process a result coming from the

23:43

server, or you can aggregate your local

23:45

data before you insert it into the into

23:47

the into the central server, right? It's

23:49

really cool.

23:50

You can also do crazier things, right?

23:52

Why don't, you know, proxy a bunch of

23:53

shards that living in separate databases

23:56

through one sort of coordinator node and

23:58

then send that

24:00

to a client and do that through Quack.

24:02

You know, could be shards, could be

24:03

replicas, could be both.

24:05

It's like really the Duck Stack in

24:06

action. Like you really the your own

24:07

imagination is only the the only

24:09

limiting factor here.

24:11

Um By the way, we're also

24:13

planning to um to add valve replication

24:16

to Quack so we can basically have a read

24:17

replica for DuckDB. That's

24:19

um that's automatically kept up to date.

24:22

And again, we've been surprised so many

24:23

times with what people build uh with

24:26

DuckDB. So, surprises again.

24:30

And now I'm going to zoom out a little

24:31

bit more. And this is the part where,

24:33

you know, like uh

24:34

maybe I have to run away quickly after

24:36

the talk, or maybe I need security

24:37

during the

24:39

during the um

24:40

uh office hours. So,

24:44

I would really kind of want to talk

24:45

about this endless OLTP versus OLAP

24:47

debate. You know, let's see where DuckDB

24:49

is on the spectrum here.

24:50

It's on I mean, you would argue that

24:52

it's it's pretty far on the on the like

24:54

side of the analytics, right?

24:57

So, let's look at this scenario in a

24:59

sort of yesterday, today, and tomorrow,

25:00

okay? Let's start with yesterday. The

25:02

common wisdom says that OLTP is sort of

25:05

Postgres land, and DuckDB is OLAP. And

25:07

somewhere in the middle is this elusive

25:09

HTAP, maybe, you know? Who knows?

25:11

Haven't seen it yet.

25:13

However, this is wrong. This is just not

25:16

true.

25:17

To yesterday, actually, is that Postgres

25:19

is actually the middle as a sort of

25:21

general-purpose system, and there are

25:23

hardcore OLTP systems out there like um

25:25

TigerBeetle. Anybody here knows

25:27

TigerBeetle?

25:28

Yes, great system. Um

25:30

and

25:32

TigerBeetle can run like so many circles

25:33

around Postgres in transactions. You

25:35

would You would It's insane, right?

25:37

Um so, Postgres is actually the

25:39

general-purpose system here.

25:41

It's not really great at any single

25:42

task, but good enough for a lot of use

25:44

cases.

25:46

Okay, that was yesterday.

25:48

So, we think that with Quack,

25:50

we've actually been working also behind

25:52

the scenes on concurrent transactions,

25:54

concurrent checkpointing, commits while

25:55

checkpointing.

25:57

There's still work to do, of course.

25:59

Um and we have announced Quack today, so

26:01

we just did that.

26:02

But, we think that actually moves DuckDB

26:05

much closer to the middle without

26:07

compromising anything on the analytics,

26:08

right? We haven't made anything slower.

26:11

So, this is kind of a bigger blob like

26:12

the We're moving closer to the middle.

26:14

So, DuckDB is

26:16

we think with Quack on the way of

26:18

becoming general-purpose

26:20

uh

26:20

as as well, just from a different sort

26:22

of side, right? It's coming from this

26:25

from this analytic analytics side.

26:27

So, this is today with Quack as of

26:29

today. And you know, who knows what

26:30

happens tomorrow.

26:33

Um maybe it's time to use a database

26:36

built after 2000 at some point. I don't

26:38

know.

26:39

>> [sighs]

26:42

>> But, I want to come back to our mission

26:43

uh here at DuckDB Labs. Um our mission

26:46

is to empower people to deal with data

26:48

and to build in amazing sort of systems

26:50

to deal with data um with confidence and

26:53

Quack is really an important ingredient

26:55

here because it really greatly expands

26:59

where DuckDB can be useful, right? Like

27:00

we've had this sort of

27:02

restriction that DuckDB was this really

27:04

in-process only system and it couldn't

27:06

really deal with this

27:07

more complex deployment models, but

27:09

that's really changing and I'm I'm

27:11

really excited that we've built

27:12

something that's, you know, also fast,

27:14

of course.

27:16

Um and now I think I have time for a

27:18

demo. You want to see a demo?

27:20

Yes. Okay, great.

27:22

>> [snorts]

27:22

>> Um this is medium terrifying as you can

27:25

probably imagine.

27:27

Um

27:28

I'm going to have to do the mirroring.

27:32

Mirroring.

27:35

All right, can you see this?

27:37

Actually, I shouldn't show you show you

27:38

this. Okay. So here you have a browser.

27:41

Um and here I have a

27:44

magic URL and I'm just going to put that

27:46

here in my browser and hope that it

27:47

works.

27:48

Um

27:50

Yes, great, it works.

27:52

Um so what happened here? So I loaded up

27:54

a DuckDB instance in the browser which

27:56

is already kind of mind-blowing that

27:57

this is possible in the first place,

27:58

right? So this is just loaded up DuckDB

28:00

in the browser.

28:01

Um

28:02

and then I said, "Hey, install this

28:04

Quack extension, install Quack from

28:06

Cornell nightly." It's currently sitting

28:08

in a separate repository because um it's

28:11

kind of still moving.

28:13

Um then we create a secret

28:16

which is a DuckDB's way of doing

28:18

handling authentication tokens. Also, we

28:20

use secrets for everything like S3

28:22

credentials and things like that.

28:24

So we create a credential um

28:27

and of course you can now copy paste my

28:29

token if you if you want if you're

28:30

quick. Um

28:32

and actually you yeah, you can. Um

28:35

and then we can just run a query here.

28:37

So for example here we just run this

28:39

query and you can see hey, this is a

28:41

node that um

28:43

select meta. You can actually see it.

28:45

This is like something that runs in a

28:47

cloud formation on Amazon where we just

28:49

click a button to spin this up. Okay,

28:51

cool. And in this um

28:54

it's also a small node. It's a T3 micro.

28:56

It's really small. And on this node, I

28:58

have a

29:01

I have a uh

29:03

I cannot type, of course.

29:05

Of course.

29:06

Oh, no. See?

29:08

This is This is you should try live

29:10

demos. This is This is medium

29:11

terrifying. I have to attach.

29:14

So,

29:16

so I have to attach. Now, I'm attaching

29:17

this uh this remote database as a sort

29:20

of as an alias remote. It's just called

29:23

remote. You don't have to call it it.

29:24

You can call it whatever. But now, I can

29:26

basically just look at this table that

29:27

lives on this other server and hopefully

29:29

the Wi-Fi doesn't let me down.

29:32

Yes.

29:34

>> [laughter]

29:35

>> Thank you.

29:36

Um I was I was kind of terrified about

29:38

this. So, what happened is basically

29:40

this used Quack to talk to the server

29:42

that sits in EC2 from the from this

29:43

laptop because this runs in the browser

29:45

and it's all good. Um and of course, you

29:48

can also say something like remote query

29:51

select star

29:54

from line item limit 10.

29:56

So, this is sort of equivalent, but you

29:59

can see that the second one was faster.

30:02

Why was the second one faster? Well,

30:04

because the first one actually

30:05

transferred the whole table. That's why

30:07

we had to wait a little bit.

30:08

Um and this is exactly what what Quack

30:11

allows you to kind of choose where

30:12

computation happens by either shipping

30:14

the query remote side or you're just

30:15

running it locally with this virtual

30:17

uh catalog attachment. So, this is like

30:19

yeah, and this again this works as of

30:21

today. You can you can start building

30:22

stuff with Quack.

30:25

Um how are we on time? Pretty good.

30:28

All right. Um

30:30

now, I can't see my notes. Uh

30:33

it's okay. I can manage this. So, so

30:35

this is I think um we are really excited

30:37

about Quack because it's it's solving

30:39

sort of this long-standing issue and

30:41

it's just really expanding sort of the

30:44

the use cases in which DuckDB can be

30:46

useful, right? Um and if everything goes

30:48

well, then this URL is live now

30:50

um because I Gabo, our DevRel, has

30:52

heroically stayed up uh in Amsterdam to

30:56

click the publish button on on this

30:58

website during the talk. So, let's see

31:01

whether it works. But yeah, so this is

31:03

really that everything the next sort of

31:04

frontier for DuckDB where it becomes

31:06

more a general-purpose sort of system

31:09

without really compromising on its

31:11

really strong in-process roots and also

31:14

of course without losing them. You can

31:15

still use all of these use cases that

31:17

DuckDB has been solving before still all

31:20

work perfectly well. It's just we're

31:21

adding to that.

31:23

Yeah. So, that's it. Um thank you so

31:25

much and I'm happy to answer some

31:26

questions.

31:28

>> [applause]

31:32

[applause]

Continue with YouTLDR

Analyze another video with Pro

Process a new video, search every timestamp, compare sources, and keep the result in your library.

Get Pro — $12/month30-day money-back guarantee

More transcripts

Explore other videos transcribed with YouTLDR.