WWC26 - The Retrieval Layer for Edge AI
AI agents as well as Sasha who is the
CTO and co-founder of Brain AI Brainform
AI right and has 20 plus years of act
architecturing scalable enterprise
systems across front end backend cloud
mobile and AI from generative AI to the
to on device. Okay, without further ado,
I will hand over to you. But quickly
reminder, if you have questions, please
submit them via the app and we will
check them after. All right, you guys on
track.
>> Thank you.
>> Yeah, thank you. And hello everyone.
Thank you very much for being here. Uh
my name is Shada. I do developer
relations at Quadrant.
>> Yeah, I'm Sasha. uh so I was introduced
as CTO co-ounder of brand form but uh
for quudrant I'm playing a role of
research engineer so uh here in the
stage I will represent myself a research
engineer of quadrant that I'm
responsible for demo part
>> yes and in the next 20 minutes we will
be diving into edge AI and how you can
use quadrant for building semantic
memory for your ondevice edge AI
applications
but uh before we start with the how.
Let's uh talk about why and when we need
edge AI. So there are three real world
constraints uh that make it a must and
maybe you're working under one of these
constraints and you're interested in
using edge AI. So this is the right
place to be. So the first one is uh
connectivity. So there are environments
where we don't have reliable access to
the cloud or no access at all for
extended periods of times such as in
pharmaceutical clean rooms. The second
is latency. In some applications like
autonomous driving, we cannot tolerate
the delay of the round trip to the cloud
and um we need to run that on device
because like even a 100 millisecond
delay can mean the vehicle travels a few
meters. uh which which can be very
critical in some situations. And lastly,
of course, we care a lot about data
privacy. So, if you think about how you
use Face ID or Touch ID to unlock your
phone, your computer every day, it's
great. It's seamless and convenient, but
you don't want your biometric data or
any type of your sensitive data leaving
your device and being uh processed in a
cloud data center.
And uh of course edge AI also has uh
cost reduction benefits. It cuts your uh
cloud processing bills. But here we uh
are talking about situations where it's
really a must and not just an
optimization.
So um edge AI actually has been there
for years. It's not new. We've been
already using it uh when for example
applying uh social media filters or
extracting the foreground of photos. But
what's changing is that now we are
bringing general intelligence. The kind
of reasoning capabilities that not long
ago used to only exist behind an API in
the cloud data center you'd never see
directly on device. And that is uh being
made possible thanks to first small
language models. So think about small
language models as lightweight versions
of LLMs with a much smaller uh parameter
count. So as you can see here on this
timeline and they are designed for
faster inference, lower compute
requirements and lower energy
consumption. So they can actually fit
and work on your computer, on your
phone, on your IoT systems, etc. And
actually we are using Google Gemma for
E2B in um a quick demo that Sasha is
going to present uh in um a minute. So
bear with us.
And the second thing is that open weight
models are becoming very very capable.
So
recent rec recently uh you don't need to
use proprietary models that sit behind
an API to have state-of-the-art
performance or a near state-of-the-art
performance because these openweight
models uh are accessible to you. you can
deploy them on your edge devices and
they are catching up and very close to u
proprietary
frontier models. So they are not the
absolute frontier but surprisingly
close. basically uh on a machine that
costs under $2,500,
you can basically run uh a model that is
uh last year's frontier and the gap uh
is shrinking every quarter.
Okay, great. We have really capable
models uh that we can deploy on the
edge. But real intelligence is not just
the model, it's also the memory. Because
even if you have the smartest model
without memory, it can feel like a
complete stranger that does not know
you, does not know your preferences,
your history, and the context that you
would want to build over time based on
your interaction.
So
memory is a really important piece and
it makes all the difference. For
example, in personal assistance, you
move from a generic chatbot that
basically doesn't remember anything
about you to something very personalized
that knows your preferences, that knows
you and that understands you.
And the second point is uh that
models
are always evolving. So eventually at
some point you will swap the model
because you will move to a more capable
one. Even the hardware you're running
the model on can be switched as devices
evolve. But memory is what you carry
forward. It's the operational history
that for example you learned in whatever
uh industry application you are building
and um it's what transforms an LLM to a
domain expert
and let's see a
how uh you can bring that memory on
device. So this is exactly what Quadrant
edge is designed for. It brings semantic
memory directly to your device
application, to your robot, to your
phone, etc. So before we dive into
Quadrant Edge, let me uh introduce
Quadrant before to those of you who
don't know it. So Quadrant is a vector
search engine and this enables you to
build semantic memory and do semantic
search. So you have different kinds of
data whatever data you are working with
in your day-to-day uh job. It could be
text, it could be images, video, audio,
whatever. The first step is that you
transform this data into a
representation that captures the meaning
and the complex relationships within
this data. And this is done with
embedding models. And after that,
Quadrant comes in as a vector search
engine that is going to index your
vectors and allows you to search them
very very quickly and efficiently. It's
based on REST uh which allows it to be
extremely fast uh and memory efficient.
And Quadrant Edge is the edge deployment
of Quadrant that is designed to run on
your low CPU devices. It runs um
in the background and you can use it for
like whatever uh edge application you
can imagine. And it has the same
features as Quadrant. And also with the
same API so you can synchronize your
data between uh ondevice and your cloud
cluster if you need for example to share
uh within different devices or like to
have a uh snapshot on the cloud.
Now let me give it to uh Sasha who's
going to present uh our first example.
>> Yeah. Hello again. Yeah. My turn. So uh
first please raise your hand if you have
no mobile phone with you.
No one. Oh yeah.
>> But I can see it in your hands.
>> I think you probably for forgot it or
something. Yeah. Because it's
complicated to imagine our life without
mobile phone, right? It's practically
part of you. Yeah. It knows uh it knows
everything about you like all your
secrets, all your conversations,
private, public, whatever. All your
photos, uh your invoices, screenshots
with interesting information,
everything. Yeah. Uh and there's a huge
amount of information. Uh so you have it
there and sometimes it's even hard to
remember uh when did I make this
screenshot or in uh which chat we
discussed this trip or how to find Wi-Fi
password that I uh sent to my friend and
so on. So I prepared demo u with it's
application on flatter impetus flatter
just to be able to run on Android and
iOS. Uh and there is two use cases. One
use case it work with chat. So actually
I generated a fake chat with a tons of
messages uh few friends that discussed
one trip uh to Amsterdam as I remember.
Uh so it's fake information I don't
remember. So the second example with
images. So I also generated a lot of
fakes invoices uh airplane tickets and
so on. Uh uh so
uh I named this gives gave the name of
this demo could run uh on device memory.
Uh so this uh this demo you can with
natural language ask uh like tell me
when we discuss this thing and uh this
demo will find this message in your chat
answer your natural language and give
you opportunity to scroll up to this
message. Uh and the second one you also
can ask natural language like show me
like invoices of how many times I went
in grocery shop or uh where is my uh
ticket to show me ticket to Amsterdam.
Yeah. Uh
and yeah everything happens on device
and no one bite leaves device. So you
get privacy by default. Uh you don't
share your memory with someone else. uh
even if you sure that there is a safe
but anyway all data on device you have
offline functionality you're in plane
you're on mountains I don't know
underground so you can work with it and
you don't have to pay because you
utilize your device power so let's take
a look the first one
so there is a chat I generated like
friends Anna Mark uh Sasha Lena and then
I'm asking so what time are we meeting
and we are meeting at 6 p.m. uh and you
press button and shows a message in chat
where we meet and Wi-Fi password the
Wi-Fi password is sunflower uh 2023 and
show a message when we discuss this so
uh it's talk with you and show you the
message so because all messages are
embedded and stored in coder edge and
you can uh search there so and second
one
uh so there are the invoices I plan
tickets. Uh so I'm asking
grocery shopping
at is open all invoices from grocery
shopping. Yika uh whatever I don't
remember and boarding pass to Amsterdam
and it found boarding pass to Amsterdam.
How does it work?
So there are two part memorize part and
recall part. memorize how to store the
data and recall how to get the data. Uh
so there are two ways. The chat part is
simple. We just
take chat history uh embed them. I use
embedding gema model to generate
embedding on device as well without
sending to cloud and then send them to
vector store. And uh another one with
photos it's a little bit more
complicated. First I use GMA 42B uh to
get information about what do we have on
this screenshot uh uh and embed uh this
uh description uh to to store it as a
text description of this image and be
able to search by description of what
we're looking for. Everything stored in
vector store and next step recall. So in
case of chat and in case of question in
the beginning almost the same first what
do we have query parsing with gema for
model I just took information for
request uh to have uh smart filters uh
for example how to separate uh flight to
Amsterdam and flight from Amsterdam from
vector similarity it's almost the same
but uh I uh got the filters and separate
from and to apply izes filters and uh
then look using semantic search with
embeddings and took only relevant
answers. Uh and in case of chat I there
is one more step answer generation it
generate for your answer in natural
language but then uh show your exact
message. In case of photo just shows the
photo. Uh this is how does it work. So I
just added video to the slides not to
run it here but if you would like to try
it yourself just come to our booth and
uh you can try it in mobile phone. Uh so
what what else? So if you're interested
how uh edge AI works nowadays and
capabilities of AI models on mobile
devices you can come to my talk
tomorrow. I will talk about what does
mean uh edge AI uh edge u intelligence
for mobile developers uh exactly as a
mobile developer. Uh so that's
everything about this demo but then
let's imagine so we were talking about
mobile phones that we have right now.
Let's imagine what we will have tomorrow
in the future. So uh like little robots
will running around us. Yeah. Do
something. Yeah. Uh what what what what
are they robots? So it's actually kind
of mobile devices, right? Uh but mobile
devices on like near future. Uh is it
possible to run edge uh quadrant edge on
robots?
>> This is a very good question and the
answer is yes, absolutely.
>> Okay. Uh some technical problem maybe.
What happened?
Okay. What's Wi-Fi
issue?
What does it happen?
Okay.
>> Yes. Finally, we have our home robot
demo on the screen after a small
technical problem. So, think uh about
having this home robot that you bring
into your home to help you with the
chores uh etc. you enter, you ask it
where did I leave my keys or can you
change the pillowcase for um the pillows
in uh the children's bedroom. So the
first thing what that happens when you
bring this robot into your home is it
has to get familiar with its
environment. It will scan your house for
the objects. And this uh recording here
is a demo that we built and I will very
quickly walk you through how this is
made possible um also of course with
quadrant edge. So we're using YOLO E for
uh object detection and we're embedding
the uh pictures the frames with um an
embedding model and also captioning each
frame and then uh we're doing hybrid
search. So we're combining both the
embeddings of the images with uh BM 25
sparse vectors for the captions uh with
reciprocal rank fusion all with quadrant
but most importantly it's all done on uh
your robot so your house basically data
does not leave uh your home and it can
answer your questions and perform the
search very quickly in sub uh millisec
seconds.
And um this is not just also about
robots. So think about uh the
experiences that you live every day and
that you can index and make searchable
with smart glasses. And this is uh what
Sasha is going to show us live now.
>> Yeah, thank you Sha. So
how many of you for example in the
morning oh where where where I drop my
keys? Yeah I might have to run quickly
to to developers people waiting me on my
boo where is key or where is my mobile
phone actually. Yeah. So [laughter]
uh I created a demo that helps exactly
with this situation. So uh I gave it
name quadrant edge object memory because
uh how does it work? It detects objects
like on the previous demo uh store uh
them exactly on glasses uh and then
using voice recognition you can ask
where is my mobile phone and it will
find it for you. Uh so uh let's try
that. We'll have no technical issues
because Wi-Fi dropped uh and uh they
actually connected by Wi-Fi. Uh so let's
try.
So let's Can you see that? Let's switch
it on.
Okay.
I can try to disconnect it. It still
works.
Okay.
I have to be sure that mobile phone is
okay. Uh so let's try to find it.
Uh let's start with laptop.
Uh where is my laptop?
laptop is there
in
many different ah yeah it's finally
so uh today we figured out then when a
lot of people around uh and Wi-Fi is
very heavy loaded uh sometimes needs to
too many energy to send signal from
glasses to laptop and on the peak of
this energy it's switched off but uh we
demonstrated hated everything I guess.
So now they are not working anymore but
without translation to laptop [laughter]
it won't happen. So actually if you will
try would like to try yourself welcome
to our booth. Uh we will be there today
so you can uh play with this glasses and
with demo uh there. So how does it work?
So uh we have
two parts as well memorize and recall as
previous demo. So we have camera
streaming and there we have a yola model
that recognize and detect the objects.
It runs exactly. So everything runs on
the glasses uh glasses reo x3 pro
is powered by qualcom
and there's a npu and jpu inside. uh so
recognition of object it takes 9
milliseconds there then I have embedding
models tiny clip it's a model small
enough to be run on glasses it's a like
to transform this object to vectors and
store to quant database then during
recall we have Google ISR that recognize
your voice then we use the same
embedding model to trans transform this
voice to the embeddings and organize the
search in quadrant antage
uh some numbers uh so object detection 9
millconds it utilize NPU and works very
fast uh image embedding it's half second
it's a maybe the longest part yeah
because tiny clip models is designed the
way that it can't utilize NPU so that's
why it takes a little bit longer
if you find the model that will utilize
NPU it will be faster of course so
vector search 15 milliseconds and vector
upsert 60 milliseconds. It's longer than
vector search because I do this flash
after every object to be safe if
something will will be wrong with
with power or something to to be sure
that I stored everything. Uh that's why
I don't use batch saving. In case of
batch saving, it would work much faster
than 60 milliseconds. uh uh so query
embedding half a second because it's the
same embedding model and uh so vector
search of top five models
less than 80 milliseconds and take a
look to this uh I just did some
experiments what if we will save object
like during one year I don't know
[laughter]
one year maybe too too much but I
checked on 100,000
vectors uh to check the how fast it will
be and result vector search 78
milliseconds. It's very extremely fast.
So quadr is really good for speed. So uh
the the
longest part is not search part here is
embeddings. So now when we add the the
embedded model that will be details npu
it will be much faster as well.
Uh so let's go further.
>> Yeah. So one last idea about these uh
smart glasses. In this demo we showed
that you can search for objects. But
this can go much further than that.
Imagine uh indexing basically your
day-to-day memories and your experiences
while wearing these uh smart glasses.
And because owning your memories does
not mean not sharing them, you can
actually create a hive mind where uh if
you imagine a family where each member
is wearing smart glasses and uh
recording their best moments uh then
this can be shared and synced from uh
your glasses from Quadrant Edge to the
cloud to Quadrant cloud and creating uh
this uh shared pool where you can search
through your memory. memories but also
uh your loved ones memories all in one
place.
So uh this is it for today's
presentation. Thank you very much for
attending and uh please come join us at
uh the booth and uh yeah we're happy to
hear your questions. Thank you.
Otherwise the question okay so I don't
have a lot of questions because I'm also
experiencing some tech technical
difficulties so let's see the questions
that I have which are on my card is
which edge case use case surprised you
the most in the wild
edge case use case so today I was
actually surprised uh with this Wi-Fi
dancing so it worked uh really good. But
then people start coming and it start to
switch off. What happened? And I then I
find out if I go outside a little bit
there are less people. It works fine.
When more people coming on the pses
needs more energy to send a signal and
on peak switch off. I didn't expect
this. Uh and so what what else
interesting?
Actually, I was surprised how fast
glasses work with NPU with the object.
So, 9 millconds. It's extremely fast. Uh
what what else?
H cases. Um
do I remember something else? So uh in
general I I'm really surprised how good
it works on devices. So and how much
opportunities it opens uh for us. Uh
tomorrow uh I will talk about it also a
little bit how
which approaches we can get. So not only
cloud or on device we also can utilize
like hybrid approach that have partially
on device partially in cloud it's
related to memory it's related to LLM
themselves and you can play with it. So
I think it's kind of future of uh
utilizing AI. It's hybrid approach.
>> Okay. One more question before I let you
go is how do you sync on devices indexes
when connectivity returns?
>> Could you repeat please?
>> How do you sync on device indexes when
connectivity returns?
Uh
so uh it's actually everything works on
device you don't have to sync so uh
everything works here so uh if you would
like uh to sync so it's not exactly edge
approach it's hybrid approach that I
mentioned so you can have uh partially
on device partially on cloud uh and then
sync it but it's not necessary it's
works absolutely independently
All right. Um, that's it for now because
my device is not working either. So, um,
we are going to take a break from the
stage because the next talk will be at
12:50. So, thank you so much Chander and
Sasha and looking forward to having your
talk tomorrow. Okay, everybody clap some
hands
>> and yeah, if you have more questions,
uh, yeah, we are here. Uh, please catch
us here or on our booth and ask all your
questions you have.
Continue with YouTLDR
Analyze another video with Pro
Process a new video, search every timestamp, compare sources, and keep the result in your library.
More transcripts
Explore other videos transcribed with YouTLDR.

Finding the Daily Bias ONLY Using Liquidity
Inter Equity Trading · English

Kant: Kritik der reinen Vernunft 1787 (Vorrede B VII) – Dieter Hattrup liest
Dieter Hattrup · German

D. Hattrup liest – C.F. von Weizsäcker: Wahrnehmung der Neuzeit: Einstein
Dieter Hattrup · German

هل يمكن الوثوق بعقلك؟ كيف غيّر هيوم وكانط فهمنا للحقيقة
الفلسفة للنوم · Arabic

Opus 5 released! Is it better than Fable?
Mastra · English

Leilão de Embriões Nelore PO DNA Genética Aditiva
LANCE RURAL OFICIAL · Portuguese (Portugal, Brazil)

Leilão Peso Pesado Rima Agropecuária
LANCE RURAL OFICIAL · Portuguese (Portugal, Brazil)

النبي .. جبران خليل جبران .. إقرا بودانك
اقرا بودانك · Arabic

Leilão Internacional CIA
LANCE RURAL OFICIAL · Portuguese (Portugal, Brazil)

23° Mega Leilão Genética Aditiva - 1ª Etapa Fêmeas Nelore PO
LANCE RURAL OFICIAL · Portuguese (Portugal, Brazil)

كتاب رسالة الغفران
كتابي المنقذ · Arabic

Erkenntnistheorie 7 Immanuel Kant II
Dominik Finkelde - Hochschule f. Philosophie · English