0:05
All right. Hello everybody. This is
0:11
Thank you so much for the invitation.
0:13
And I'm still amazed I got away with
0:16
this talk proposal, right? Because I
0:18
just basically said this.
0:22
that's what we wanted to surprise you
0:23
and it wouldn't have been a surprise if
0:25
it would have been in the program half a
0:26
year ago. You probably understand this.
0:31
you know the story so far.
0:33
And you know a small recap like in a in
0:36
a good in a good TNG episode two parter,
0:39
you know, we start with a previously.
0:41
I see my people are here. We're good.
0:43
We're going to have fun today. This is
0:48
I've spent the last decade or so working
0:50
on DuckDB together with many other
0:54
since we're always looking for
0:55
superlatives, you know, it's the
0:57
friendliest SQL database. I think that's
0:59
fair, right? friendliest.
1:01
We're not going to you know get into
1:02
this whole like you know
1:04
game of who has the fastest and
1:06
whatever. You know, cuz it's the
1:08
DuckDB is a universal data wrangling
1:11
tool. That's how we like to call it
1:12
because you can kind of use it for
1:15
under the sun as long it is like vaguely
1:18
It's actually used quite everywhere.
1:21
It's kind of wild to see that. I'm still
1:24
totally amazed that our like little
1:25
project from from university became such
1:29
It runs in space. It runs in a battery.
1:32
It can run fully local in a process.
1:34
It's kind of ideal for tasks that you
1:37
don't you know trust your intern with,
1:39
right? Like we have a lot of these tasks
1:41
where we kind of send the intern to do
1:43
it and then we hope that they don't blow
1:46
up the database. Well, DuckDB is
1:47
actually quite useful for this because
1:49
you can just run all of that locally.
1:52
And you don't end up with a massive
1:53
Snowflake bill. It's great.
1:56
Sorry. I mean I I I like the guys a
1:58
snowflake, but I mean, you know what I
2:01
DuckDB has actually quite crazy adoption
2:04
at the moment. So, we are at the moment
2:07
running at about 40 million downloads
2:09
per month of DuckDB itself, and that's
2:10
just a Python package. Like, that
2:12
doesn't cover anything else. It's just a
2:14
Python package. It's just the one that
2:16
we have the best metrics on.
2:18
Um and uh that's that's kind of crazy.
2:21
That's like more than 1 million each
2:25
Also, what's kind of crazy to see is we
2:27
have these extensions, which are plugins
2:30
And those are downloaded like 160
2:32
million times each month, which is also
2:35
totally wild. Like, it's like I'm
2:37
clearly all these downloads must be
2:40
We're absolutely sure.
2:42
We run out of Americans in about 10
2:50
But, we went beyond DuckDB. Actually,
2:51
this was from almost exactly 1 year ago.
2:54
I stood on the stage of you know, the
2:57
this conference and hinted in my last
2:59
slide at something new for DuckDB. Um
3:02
also, I only have one set of clothes,
3:08
And so, from this logo, um obviously,
3:10
the logo has slightly changed, but we
3:15
um about a week after the conference.
3:17
And now you know this is DuckLake.
3:20
is of course a lakehouse format that
3:21
kind of admits that metadata is best
3:24
stored in a database.
3:28
um the preview in May, like about a year
3:32
And the logo has slightly changed since
3:34
the preview, but that's fine.
3:39
is Yeah, it's I mentioned it's the
3:40
simplest way to a data lake. And this is
3:42
like a bit of a you know, aggressive, if
3:44
you want, slide from us, where we just
3:46
show the sort of the bits and pieces of
3:49
Iceberg, the you know, that lakehouse
3:51
format lots of people use. On the left
3:53
side where you see, oh, we need to have
3:54
this catalog and then we have metadata
3:56
files and manifest lists and manifest
3:58
files and data files and it's all a bit
4:02
Um, and then on the right side we have
4:04
DuckDB which again admits that metadata
4:06
probably should be stored in database
4:08
and then like you have a much more
4:13
this has been hitting quite a nerve. We
4:15
actually expected criticism
4:17
but we got an overwhelmingly positive
4:19
response. Like lots of praise. So things
4:21
like this where somebody said, "DuckDB
4:23
and DuckDB Lake, why we bet the company
4:27
Kind of like Duck stack. It's uh
4:29
And this was on a preview. Can you
4:31
imagine, right? Like we we bang
4:32
something out which is a preview and
4:35
then people say, "Now we'll let's bet a
4:36
company on it. It's great." Right? And
4:38
this is only the kind of things that you
4:39
can I can talk about in public, right?
4:41
There's loads of loads of
4:42
insane insanely positive response to to
4:45
DuckDB. Clearly hit a nerve. Great.
4:48
Um, just a month couple of
4:50
weeks ago, a month ago we released
4:54
Um, so the first version of DuckDB that
4:57
production ready. Um, this time we
4:59
actually got press coverage which was
5:01
Um, and you know, you people have then
5:06
uh, DuckDB is this thing that only the
5:07
DuckDB people love and and you know, the
5:10
the real people use use something else."
5:12
But we actually mentioned earlier we we
5:15
can kind of count how many times people
5:17
download the extensions.
5:18
And DuckDB has extensions for Iceberg,
5:21
for DuckDB and for Delta Lake, right?
5:23
And so, you know, in in in your mind
5:26
imagine how these kind of different
5:27
systems stack up in terms of downloads,
5:29
right? Just make a mental picture.
5:31
And if you look at this, you can
5:33
actually see that within less than a
5:36
year actually and only 4 weeks after
5:37
declaring 1.0, there's already like
5:40
DuckDB is 2.5 million downloads a month
5:43
at the moment for the extension and the
5:45
Iceberg extension is 2.9 and the Delta
5:47
is 2.3, but it's to us totally amazed
5:50
amazing to see how many people are
5:53
happy uh with with this DuckDB thing.
5:56
And yeah, so this is again, this is of
5:58
course only the respective downloads
6:00
within uh the DuckDB ecosystem.
6:03
Uh yeah, within a year. But
6:06
that's not what I want to talk about
6:11
um the clicker stopped working.
6:15
Isn't it in a very dramatic pause,
6:18
Clicker stops working in the in the in
6:20
our the continuation of the of the of
6:21
the Star Trek episode, obviously. Um
6:23
today is really not about DuckDB, and I
6:25
have to admit that we also worked on
6:27
other things this last year. Uh only
6:29
maybe less public than uh DuckDB. And if
6:32
you think about it, it's kind of an
6:34
insane speed for any sort of database
6:36
project to sort of run these gigantic
6:38
initiatives and then kind of finish them
6:40
within a year and do something else in
6:42
the meantime. I I'm I'm still amazed and
6:44
actually really proud of our team to be
6:46
able to do that. And actually Pedro, one
6:48
of our guys, ended up taking over DuckDB
6:51
so that I could do something else. And
6:52
then uh other people have thankfully uh
6:55
taken over some project management, so I
6:57
could do something else. So this is
6:58
something that I have actually worked on
6:59
personally, and so I'm kind of nervous
7:02
to talk to you about it today, okay?
7:04
Let's look at this. Okay.
7:06
New thing. What's the new thing?
7:11
single node DuckDB works great, right? I
7:13
would I would even say single node
7:15
DuckDB works extremely great, right? You
7:17
have a single computer.
7:19
We have done experiments. You can go to
7:21
gigantic data sizes. Basically, it will
7:23
just continue crunching stuff until you
7:25
run out of disk space, which is pretty
7:28
sort of um boundary to have.
7:31
Okay, so this is great.
7:35
DuckDB can also talk to almost
7:37
everything under the sun, right? As it
7:38
can talk to Postgres, for example. Oh,
7:40
we just you have a extension that is the
7:42
post scanner. There's my sequel one,
7:44
there's a sequel light one.
7:45
We have run we can talk to random ODBC
7:48
drivers and integrate with and just pull
7:51
data back and forth between these
7:52
things. Works really well. Okay.
7:55
Uh DuckDB can also talk to object
7:57
stores, right? We have integrations for
7:59
all the wonderful object stores. You can
8:01
all the data formats we have support for
8:03
all the you know, parquet and its
8:06
various competitors. All works really
8:08
read and write from the object store. We
8:10
have integrations with catalogs like S3
8:12
tables. It's all it's all wonderful.
8:15
Little check mark. Works well.
8:21
DuckDB cannot talk to DuckDB.
8:23
Well, that's a bit annoying, isn't it?
8:25
And we are not the first people
8:27
to notice this. There's actually been a
8:32
large amount of GitHub repos and I think
8:34
what about we are running at about one
8:35
per week that's popping up where people
8:38
are um just adding functionality for
8:40
DuckDB to talk to DuckDB.
8:42
And here are just four of them. There's
8:46
maybe it's also, you know, something
8:47
that is maybe quite easy to to wipe code
8:49
in an afternoon. That's why they're
8:50
popping up so many. But it's it's it's
8:52
just one of these things where
8:54
yeah, we we really we
8:57
we didn't have a strong sort of feelings
8:58
about it or maybe we did. And then all
9:01
these people out there do things that
9:03
that kind of solve this problem of
9:05
DuckDB to talk to DuckDB. And at some
9:06
point it's not sort of I have to sort of
9:08
take off my academic hat
9:10
of being right, which of course I love
9:13
But at some point it's about solving
9:15
people's problems, right?
9:17
And clearly there is a need out there
9:19
because people are people are building
9:22
Um at this point I also have to admit
9:25
that 1 year ago I was standing here
9:28
and talked about the benefits of single
9:30
node in-process databases, right?
9:35
And while I fully admit that it is still
9:37
a great idea, that I mean I will still I
9:39
will still think that making a single
9:41
node database is a great idea.
9:43
There are just lots of use cases out
9:46
where a in-process database cannot
9:48
cover, all right? Like if you're an
9:50
in-process database, there's just some
9:51
things that you cannot do.
9:53
And I'm going to show you one example,
9:57
which is uh this sort of the real-time
9:59
analytics use case, right? You have you
10:00
got you have a fleet of nodes. They they
10:03
do whatever, and they try to store some
10:05
for example some telemetry in a central
10:07
sort of authoritative location.
10:10
And it's a flood of sort of fairly small
10:13
So, what do you do now with your
10:14
in-process database, right?
10:16
It's really difficult to to map this
10:18
um to an in-process database, and it's
10:20
unfortunate because it excludes
10:23
a lot of really exciting use cases,
10:28
also maybe from the academic side that
10:29
we have to kind of admit that this is
10:31
also part of analytics. Like I think for
10:35
uh the academic world has kind of
10:36
treated change in analytical databases
10:38
as like a you know a
10:40
something for other people to fill to
10:42
solve, but we have to kind of we have to
10:44
kind of get around to that.
10:46
And the crazy thing is that this even
10:48
on localhost. Like people have multiple
10:51
tools they want to talk they want to
10:53
have talk to the same database. Like you
10:55
have For example, you have like a
10:57
DBeaver or DataGrip running on a DuckDB
10:59
database, and then you want to use the
11:00
CLI to do something else. You want to
11:02
start up the DuckDB UI, or I don't know.
11:04
Like this has been a a sort of
11:06
long-standing customer complaint, or I
11:08
mean feature request.
11:11
Um I should I should get my corp speak
11:14
Uh but you can kind of do this already
11:16
with DuckLake, but this is quite a lot
11:18
of additional operational complexity.
11:20
And performance for small inserts is
11:22
really like these these lakehouse
11:23
formats are really not made for tons of
11:26
tiny inserts. So, that's kind of
11:29
So, it says breeze on my presenter notes
11:33
So, that was the motivation. And now I'm
11:35
can I can show you something that makes
11:37
me really happy to show you.
11:40
I mean, this is just the happiest duck,
11:43
I was I found this picture and I was I
11:45
was so happy that I get to show you this
11:46
little happy ducky. But, so
11:49
let's say we want to have ducks talk to
11:56
Exactly. Today I I as I'm very very
11:58
happy to to show you Quack. And Quack is
12:01
a DuckDB extension that basically
12:04
extends DuckDB with capabilities with
12:06
quite powerful capabilities and I'm
12:07
going to show you all about it um to
12:10
communicate with other Duck DBs.
12:13
Um and maybe we can do like a a favor
12:15
and we can quack together. So, we do
12:16
this in our company quite a lot. So, can
12:18
do quack quack quack quack quack. Come
12:21
on. Quack quack quack quack quack quack
12:24
quack. Excellent. You made me very
12:27
So, what is Quack anyway?
12:30
So, this is how it looks like. So, Quack
12:32
is kind of a client-server protocol in
12:34
the in a sort of traditional sort of
12:35
sense for DuckDB. So, it is as I
12:37
mentioned implemented as a DuckDB
12:40
But, the cool thing is that both sides
12:43
are just DuckDB. So, there's not like a
12:45
separate there's not a separate server.
12:48
There's not a separate client. It's just
12:50
DuckDB. So, you have two DuckDBs. Here
12:52
we have two. Like one is green and one
12:55
Um and on left you can say, "Okay,
12:57
please serve this local database that
13:01
on this uh local host and then maybe
13:03
you're creating a table there." And on
13:06
you say attach, which is the DuckDB way
13:09
connecting to other things. Like we had
13:11
this for a while. For example, you can
13:12
connect attach a Postgres or you can
13:14
attach a SQLite or a Duck Lake or
13:16
And then you can just type for example
13:18
from remote foo and then magically
13:22
the query will be run on the other side.
13:24
The result will be shipped back
13:25
and be displayed on the blue side,
13:28
The same you can also if you want you
13:29
can also have a sort of explicit uh
13:32
query there where you say from
13:34
remote.query from foo and we actually
13:38
something really cool there with uh
13:40
DuckDB 2.0 uh to um make this really
13:43
nice in a user interface.
13:46
So, yeah, this is this is really a
13:47
minimum example, but you can do anything
13:50
reachable from SQL, right? Everything.
13:54
All the data types, all the all the
13:56
everything it just works, right?
13:59
Here is actually a slightly bigger uh
14:01
example because this actually works
14:05
So, if you today if you have DuckDB 152,
14:08
you can just run this script and it will
14:10
already do all of this. So, you have
14:12
client server um right right off the box
14:15
in DuckDB 1.5. Um this works for all
14:18
DuckDB distributions, for Python, for
14:19
shell, for web assembly.
14:21
Yeah, you name it, right?
14:23
Like Windows, OSX, Linux, we don't care.
14:25
So, this is being released today as a
14:29
And we expect the production release in
14:31
a couple of months uh with DuckDB 2.0 uh
14:35
which are which coming in fall, by the
14:36
way, in case you haven't heard yet.
14:38
Um and yeah, as I mentioned, we're going
14:40
to do um some new simplification some
14:42
simplification on the way that you
14:44
specify queries that should be run
14:46
remotely which I should give you a
14:47
really um great sort of experience in
14:50
writing queries for other servers. It's
14:52
just that it's not there yet and I
14:53
didn't want to show you syntax that
14:54
doesn't work today. So,
14:57
So, let me talk a bit about how Quack
15:00
So, this is like, you know, this is kind
15:03
of how it looks like for you or for your
15:06
Um but how does it actually designed?
15:08
And this is something that
15:09
uh we did kind of ourselves. So,
15:12
on the bottom is TCP/IP. Obviously, we
15:14
have no choice but to use TCP/IP to talk
15:16
to other servers, right?
15:17
Uh we're not using UDP. Uh it's TCP/IP.
15:21
And then what is what we have on top?
15:27
Any Any guesses what we put on TCP IP as
15:32
Excellent. So, this is um we need a
15:34
protocol on top of HTTP because of WASM.
15:36
Like, in case you don't know, but DuckDB
15:38
runs in the browser, right? And we do
15:40
want to talk to a server from the
15:43
browser. And so, we really don't have a
15:46
uh to to not use anything else. And this
15:48
is also really great to use HTTP for
15:50
database protocol in that you build in
15:55
all the firewalls like it, the cyber
15:57
people don't freak out. Um you can just
15:59
slap uh TLS on it and it's going to be
16:01
encrypted. Um and it's actually quite
16:04
impressive. We've done some done some
16:07
because HTTP is so ubiquitous, all the
16:10
hardware actually is optimized for it.
16:12
So, if you if you if you use a different
16:14
it's it's really counterintuitive, but
16:16
because of you know, I guess because of
16:20
HTTP is faster than anything else. So,
16:22
it's actually by just because the
16:23
hardware prioritizes and just knows
16:25
better how to deal with it. It's kind of
16:26
funny. So, if you have an an another
16:29
protocol, it's it's actually going to be
16:30
slower. Um and uh yeah, it's kind of
16:33
it's kind of a weird counterintuitive
16:36
Okay, so what do we do on top of HTTP?
16:39
We actually use our We need a
16:40
serialization format. We need something
16:42
to basically encode queries and results
16:44
and all that stuff. And so, what we we
16:47
have that already in DuckDB, right? We
16:49
have a serializer because we need to
16:50
serialize things to the write-ahead log.
16:52
And let's just use that. So, we can use
16:55
lossless serialization
16:57
of all the internal structures like
17:01
data, chunks, all that stuff. It's
17:03
really lossless serialization. And on
17:05
top of that, we just put uh we need some
17:08
interaction protocol.
17:10
All right? And it's really
17:11
straightforward. It's just messages with
17:13
different types like every other
17:14
database RPC out there. So, it's a
17:16
request-response pattern. So, you can
17:18
for example imagine that you can execute
17:20
a statement, fetch more results, and so
17:22
on and so forth, right?
17:26
RPC, kind of a client-server protocol,
17:28
all these things, and now we have some
17:30
sort of side requirements. For example,
17:34
everybody's favorite. Who here loves
17:37
No one. Okay, me neither.
17:41
So, it's really it's really interesting
17:43
because once you start I'm going to come
17:44
over here because I see you people, you
17:45
know, like I love you, too. I just I'm
17:47
not always over there. It's just my my
17:48
notes are over there, so I'm
17:51
This is uh one of these things where now
17:54
that we have a client-server, we cannot
17:55
really hide behind our in in-process
17:58
authentication model anymore, right? We
18:00
have to do authentication, which is
18:02
right? Um but we also realized that we
18:05
cannot solve this for everybody. So,
18:07
what we have done is we have
18:09
basically made this pluggable through
18:11
DuckDB extensions. So, basically you,
18:15
override the authentication method that
18:17
that uh Quack has with whatever you
18:19
think is best. There is a default one,
18:22
which is based on tokens. I'll show you
18:24
but it's basically something that you
18:26
can decide what you want. If you want to
18:31
LDAP server or so, you can do that,
18:34
right? Um I think it's really exciting
18:36
when you have such a community like
18:37
DuckDB has. Um there is this community
18:41
of DuckDB extension writers that have
18:43
already written many extensions.
18:46
We can kind of rely on them to to um to
18:48
just cover some more use cases, and of
18:50
course, if you're a big corp, you can
18:53
Um you can also just write a SQL
18:55
function, by the way, if you want. You
18:56
can if you can you express your
18:58
authentication in SQL, you can just
18:59
write a SQL function. It's fine.
19:03
And then the other thing, the even more
19:06
is the authorization, right? This is
19:07
even harder. It's like now we've
19:08
authenticated somebody, how do we
19:12
what the permissions are going to be,
19:13
right? Like you have table level, column
19:15
level, row level, blah blah blah. It's a
19:17
huge feature set. Uh people expect a ton
19:19
of stuff there. And uh we again we have
19:22
managed to avoid it so far. Again, we
19:27
want to solve we're not don't want to
19:28
solve this once and for all. We're going
19:29
to give you the flexibility to do your
19:31
own thing. And you there's going to
19:33
there's also a callback for that where
19:35
you have a callback that sees what the
19:37
user wants to do. You can decide whether
19:39
you want to allow it or not. We can even
19:40
modify the queries that the user is
19:42
running. There's lots of interest.
19:43
There's some really powerful things
19:44
there. And this is I think interesting
19:46
in Quack is that it's
19:48
it's kind of giving you the tools to
19:50
build something exciting. It's not
19:51
trying to solve everything from the
19:58
experiments. Because I'm a database
20:00
person and I cannot live without showing
20:01
a performance experiments, right? It's
20:04
So in case you don't know um
20:07
many years ago in 2017 we wrote a paper
20:09
how terrible database client-server
20:12
Um here you see the plot on this on this
20:14
paper where basically every database
20:16
protocol was worse than netcat by a
20:20
It's not so great. Um and actually the
20:22
insights from this paper um
20:25
made us really strong believers in in
20:26
process. But now you know we had to
20:28
reconsider. But it is of course extra
20:30
terrifying to build a client-server
20:32
protocol when you have been the one
20:34
that's been banging out like has been
20:36
complaining about everybody else's
20:37
protocol, right? So we better get this
20:40
Okay, so this is the 2017 um
20:43
2017 state of the world.
20:46
Um so let's do some experiments. And the
20:48
experiment is really simple. We have a
20:52
uh running in AWS. These are virtual
20:54
machines. Um they are like fairly small.
20:57
32 GB of RAM and eight CPUs, right? But
21:00
they do have fairly fast network. They
21:02
have 15 GB GB per second networking. And
21:06
they are in the same availability zone,
21:08
right? Like what you would have in your
21:12
These are different, and we tried three
21:15
So, we tried Postgres, because everybody
21:19
We tried Quack, of course, because of
21:21
course we tried Quack. And we tried
21:23
Arrow Flight SQL. There is So, the Arrow
21:25
project has also created something
21:28
to basically deal with, you know,
21:30
database protocol interactions. It's
21:31
called Flight SQL, and we actually use
21:33
something called Gizmo data, which is
21:34
one of the projects that I've mentioned
21:36
earlier that has shown us that this is
21:37
really something people want.
21:39
And we have two experiments, bulk
21:42
and small inserts. So, let's start with
21:44
bulk transfer. Bulk transfer.
21:47
So, here is the result.
21:50
We've transferred millions of rows of a
21:52
database table. So, it's 100,000 rows, 1
21:55
million rows, 10 million rows, 60
21:57
million rows, and we measured wall clock
21:58
time. Lower is better, obviously.
22:01
So, Quack manages to transfer 60 million
22:03
rows in this is scenario in around 5
22:07
And Postgres took 3 minutes.
22:09
So, the Postgres protocol took 3
22:11
minutes, and just hold on to that. It's
22:13
it's uh 5 seconds versus 3 minutes.
22:16
Pretty good. Arrow Flight took 20
22:18
seconds. It's better, but it's of course
22:24
the other So, I mean, kind of you kind
22:26
of trusted us to get bulk transfers
22:27
right, right? Like I mean, we have been
22:28
working on this for ages.
22:33
Um but something that was that we were
22:35
then we we said, "Okay, we also need to
22:37
do transactions." And especially for
22:38
this use case, where you have many small
22:40
inserts coming from all over the place.
22:42
So, here we have a second experiment,
22:45
where we run single insert transactions
22:47
with an incoming increasing number of
22:49
client threads, right?
22:51
So, 1 2 4 8, and so. And we just
22:54
measured the completed transactions per
22:56
And we do What we see? Well, we see
22:57
Arrow Flight not doing so well. Um it's
23:00
designed for bulk, and kind of only for
23:03
Um Postgres is doing okay. It's kind of
23:04
was designed for this use case. But
23:07
Quack surprisingly did really well. Like
23:09
we finished something like 5,000
23:10
transactions per second on DuckDB uh for
23:13
eight clients here. So, that's that's
23:14
pretty good. We have optimized the
23:16
protocol a little bit
23:18
to make this work. But pretty amazing to
23:21
So, it's fast both for bulk and for
23:25
for small transactions.
23:28
So, what can you do with this now? Well,
23:30
you can do lots of things. Here's the
23:31
classic case. You can use Quack to glue
23:33
two DuckDBs together. Great.
23:36
But what cool What's really nice about
23:37
this is you have DuckDB on both sides.
23:40
So, you can do really cool things like
23:41
post-process a result coming from the
23:43
server, or you can aggregate your local
23:45
data before you insert it into the into
23:47
the into the central server, right? It's
23:50
You can also do crazier things, right?
23:52
Why don't, you know, proxy a bunch of
23:53
shards that living in separate databases
23:56
through one sort of coordinator node and
24:00
to a client and do that through Quack.
24:02
You know, could be shards, could be
24:03
replicas, could be both.
24:05
It's like really the Duck Stack in
24:06
action. Like you really the your own
24:07
imagination is only the the only
24:09
limiting factor here.
24:11
Um By the way, we're also
24:13
planning to um to add valve replication
24:16
to Quack so we can basically have a read
24:17
replica for DuckDB. That's
24:19
um that's automatically kept up to date.
24:22
And again, we've been surprised so many
24:23
times with what people build uh with
24:26
DuckDB. So, surprises again.
24:30
And now I'm going to zoom out a little
24:31
bit more. And this is the part where,
24:34
maybe I have to run away quickly after
24:36
the talk, or maybe I need security
24:40
uh office hours. So,
24:44
I would really kind of want to talk
24:45
about this endless OLTP versus OLAP
24:47
debate. You know, let's see where DuckDB
24:49
is on the spectrum here.
24:50
It's on I mean, you would argue that
24:52
it's it's pretty far on the on the like
24:54
side of the analytics, right?
24:57
So, let's look at this scenario in a
24:59
sort of yesterday, today, and tomorrow,
25:00
okay? Let's start with yesterday. The
25:02
common wisdom says that OLTP is sort of
25:05
Postgres land, and DuckDB is OLAP. And
25:07
somewhere in the middle is this elusive
25:09
HTAP, maybe, you know? Who knows?
25:11
Haven't seen it yet.
25:13
However, this is wrong. This is just not
25:17
To yesterday, actually, is that Postgres
25:19
is actually the middle as a sort of
25:21
general-purpose system, and there are
25:23
hardcore OLTP systems out there like um
25:25
TigerBeetle. Anybody here knows
25:28
Yes, great system. Um
25:32
TigerBeetle can run like so many circles
25:33
around Postgres in transactions. You
25:35
would You would It's insane, right?
25:37
Um so, Postgres is actually the
25:39
general-purpose system here.
25:41
It's not really great at any single
25:42
task, but good enough for a lot of use
25:46
Okay, that was yesterday.
25:48
So, we think that with Quack,
25:50
we've actually been working also behind
25:52
the scenes on concurrent transactions,
25:54
concurrent checkpointing, commits while
25:57
There's still work to do, of course.
25:59
Um and we have announced Quack today, so
26:02
But, we think that actually moves DuckDB
26:05
much closer to the middle without
26:07
compromising anything on the analytics,
26:08
right? We haven't made anything slower.
26:11
So, this is kind of a bigger blob like
26:12
the We're moving closer to the middle.
26:16
we think with Quack on the way of
26:18
becoming general-purpose
26:20
as as well, just from a different sort
26:22
of side, right? It's coming from this
26:25
from this analytic analytics side.
26:27
So, this is today with Quack as of
26:29
today. And you know, who knows what
26:33
Um maybe it's time to use a database
26:36
built after 2000 at some point. I don't
26:42
>> But, I want to come back to our mission
26:43
uh here at DuckDB Labs. Um our mission
26:46
is to empower people to deal with data
26:48
and to build in amazing sort of systems
26:50
to deal with data um with confidence and
26:53
Quack is really an important ingredient
26:55
here because it really greatly expands
26:59
where DuckDB can be useful, right? Like
27:00
we've had this sort of
27:02
restriction that DuckDB was this really
27:04
in-process only system and it couldn't
27:06
really deal with this
27:07
more complex deployment models, but
27:09
that's really changing and I'm I'm
27:11
really excited that we've built
27:12
something that's, you know, also fast,
27:16
Um and now I think I have time for a
27:18
demo. You want to see a demo?
27:22
>> Um this is medium terrifying as you can
27:28
I'm going to have to do the mirroring.
27:35
All right, can you see this?
27:37
Actually, I shouldn't show you show you
27:38
this. Okay. So here you have a browser.
27:41
Um and here I have a
27:44
magic URL and I'm just going to put that
27:46
here in my browser and hope that it
27:50
Yes, great, it works.
27:52
Um so what happened here? So I loaded up
27:54
a DuckDB instance in the browser which
27:56
is already kind of mind-blowing that
27:57
this is possible in the first place,
27:58
right? So this is just loaded up DuckDB
28:02
and then I said, "Hey, install this
28:04
Quack extension, install Quack from
28:06
Cornell nightly." It's currently sitting
28:08
in a separate repository because um it's
28:11
kind of still moving.
28:13
Um then we create a secret
28:16
which is a DuckDB's way of doing
28:18
handling authentication tokens. Also, we
28:20
use secrets for everything like S3
28:22
credentials and things like that.
28:24
So we create a credential um
28:27
and of course you can now copy paste my
28:29
token if you if you want if you're
28:32
and actually you yeah, you can. Um
28:35
and then we can just run a query here.
28:37
So for example here we just run this
28:39
query and you can see hey, this is a
28:43
select meta. You can actually see it.
28:45
This is like something that runs in a
28:47
cloud formation on Amazon where we just
28:49
click a button to spin this up. Okay,
28:51
cool. And in this um
28:54
it's also a small node. It's a T3 micro.
28:56
It's really small. And on this node, I
29:03
I cannot type, of course.
29:08
This is This is you should try live
29:10
demos. This is This is medium
29:11
terrifying. I have to attach.
29:16
so I have to attach. Now, I'm attaching
29:17
this uh this remote database as a sort
29:20
of as an alias remote. It's just called
29:23
remote. You don't have to call it it.
29:24
You can call it whatever. But now, I can
29:26
basically just look at this table that
29:27
lives on this other server and hopefully
29:29
the Wi-Fi doesn't let me down.
29:36
Um I was I was kind of terrified about
29:38
this. So, what happened is basically
29:40
this used Quack to talk to the server
29:42
that sits in EC2 from the from this
29:43
laptop because this runs in the browser
29:45
and it's all good. Um and of course, you
29:48
can also say something like remote query
29:54
from line item limit 10.
29:56
So, this is sort of equivalent, but you
29:59
can see that the second one was faster.
30:02
Why was the second one faster? Well,
30:04
because the first one actually
30:05
transferred the whole table. That's why
30:07
we had to wait a little bit.
30:08
Um and this is exactly what what Quack
30:11
allows you to kind of choose where
30:12
computation happens by either shipping
30:14
the query remote side or you're just
30:15
running it locally with this virtual
30:17
uh catalog attachment. So, this is like
30:19
yeah, and this again this works as of
30:21
today. You can you can start building
30:25
Um how are we on time? Pretty good.
30:30
now, I can't see my notes. Uh
30:33
it's okay. I can manage this. So, so
30:35
this is I think um we are really excited
30:37
about Quack because it's it's solving
30:39
sort of this long-standing issue and
30:41
it's just really expanding sort of the
30:44
the use cases in which DuckDB can be
30:46
useful, right? Um and if everything goes
30:48
well, then this URL is live now
30:50
um because I Gabo, our DevRel, has
30:52
heroically stayed up uh in Amsterdam to
30:56
click the publish button on on this
30:58
website during the talk. So, let's see
31:01
whether it works. But yeah, so this is
31:03
really that everything the next sort of
31:04
frontier for DuckDB where it becomes
31:06
more a general-purpose sort of system
31:09
without really compromising on its
31:11
really strong in-process roots and also
31:14
of course without losing them. You can
31:15
still use all of these use cases that
31:17
DuckDB has been solving before still all
31:20
work perfectly well. It's just we're
31:23
Yeah. So, that's it. Um thank you so
31:25
much and I'm happy to answer some