Connect AI to Billions of Legal Documents — Simon Eskildsen, turbopuffer & Jacob Lauritzen, Legora
Read full transcript 16 segments
-
Okay. Hi everyone. Welcome um to this 20 Okay. Hi everyone. Welcome um to this 20 minute talk about connecting AI to loads minute talk about connecting AI to loads minute talk about connecting AI to loads of legal documents. My name is Jacob. of legal documents. My name is Jacob. of legal documents. My name is Jacob. I'm an engineer at Legora. I'm an engineer at Legora. I'm an engineer at Legora. >> Yeah. And I got to step into frame here. >> Yeah. And I got to step into frame here. >> Yeah. And I got to step into frame here. Uh I'm Simon. I'm the CEO and uh CEO and Uh I'm Simon. I'm the CEO and uh CEO and Uh I'm Simon. I'm the CEO and uh CEO and co-founder of Turopuffer, a search co-founder of Turopuffer, a search co-founder of Turopuffer, a search engine that we work with Legora and engine that we work with Legora and engine that we work with Legora and others on. others on. others on. >> Neat. Um super quickly >> Neat. Um super quickly >> Neat. Um super quickly um introduction to Legora. We're a um introduction to Legora. We're a um introduction to Legora. We're a collaborative AI platform for legal collaborative AI platform for legal collaborative AI platform for legal work. And so that means we have law work. And so that means we have law work. And so that means we have law firms that are clients and we have uh firms that are clients and we have uh firms that are clients and we have uh in-house legal teams that are clients. in-house legal teams that are clients. in-house legal teams that are clients. And they use Legora to do reviews of And they use Legora to do reviews of And they use Legora to do reviews of contracts. They use it to go through an contracts. They use it to go through an contracts. They use it to go through an absurd amount of contracts and make sure absurd amount of contracts and make sure absurd amount of contracts and make sure that they all look good. They use them that they all look good. They use them that they all look good. They use them to create new contracts. They do legal to create new contracts. They do legal to create new contracts. They do legal research which means um looking over all research which means um looking over all research which means um looking over all potential law and they collaborate potential law and they collaborate potential law and they collaborate inside Lor. So you can think of Lorra inside Lor. So you can think of Lorra inside Lor. So you can think of Lorra sort of as a linearfigmas sort of as a linearfigmas sort of as a linearfigmas notion GitHub for legal work. It's a notion GitHub for legal work. It's a notion GitHub for legal work. It's a lot. lot. lot. We um are one of the fastest growing We um are one of the fastest growing We um are one of the fastest growing companies right now. We grow extremely companies right now. We grow extremely companies right now. We grow extremely fast. Yeah, tons of numbers on the fast. Yeah, tons of numbers on the fast. Yeah, tons of numbers on the screen. I'll just skip through that.
-
screen. I'll just skip through that. screen. I'll just skip through that. What we really want to talk about is What we really want to talk about is What we really want to talk about is search today. So at Lor there's two search today. So at Lor there's two search today. So at Lor there's two types of search that we do. There is types of search that we do. There is types of search that we do. There is project search and legal research. project search and legal research. project search and legal research. Project search is um basically projects Project search is um basically projects Project search is um basically projects in Lor is like the the unit of work that in Lor is like the the unit of work that in Lor is like the the unit of work that you have. So if um let's say you are you have. So if um let's say you are you have. So if um let's say you are SpaceX and you want to acquire Cursor, SpaceX and you want to acquire Cursor, SpaceX and you want to acquire Cursor, then that would be one project with your then that would be one project with your then that would be one project with your law firm. H and so they would go into law firm. H and so they would go into law firm. H and so they would go into Lor and they would upload all these Lor and they would upload all these Lor and they would upload all these documents and the law firm that helped documents and the law firm that helped documents and the law firm that helped you would go through all of the you would go through all of the you would go through all of the employment agreements, all of the employment agreements, all of the employment agreements, all of the contracts with suppliers. Um I know contracts with suppliers. Um I know contracts with suppliers. Um I know cursor is using Turbo Puffer so maybe cursor is using Turbo Puffer so maybe cursor is using Turbo Puffer so maybe there's a contract there they'd look at. there's a contract there they'd look at. there's a contract there they'd look at. Um Um Um but basically you do the search confined but basically you do the search confined but basically you do the search confined to a project and projects can be tens of to a project and projects can be tens of to a project and projects can be tens of documents to millions of documents. The documents to millions of documents. The documents to millions of documents. The other use case is legal research. And other use case is legal research. And other use case is legal research. And legal research is sort of a deep legal research is sort of a deep legal research is sort of a deep research style workload where we'll research style workload where we'll research style workload where we'll search across tons of laws, previous search across tons of laws, previous search across tons of laws, previous cases, regulations, etc., etc. And cases, regulations, etc., etc. And cases, regulations, etc., etc. And people use this to answer questions such people use this to answer questions such people use this to answer questions such as like how do I handle this specific as like how do I handle this specific as like how do I handle this specific thing? And they'll also use it for thing? And they'll also use it for thing? And they'll also use it for litigation. Maybe they want to sue litigation. Maybe they want to sue litigation. Maybe they want to sue someone or maybe they are getting sued someone or maybe they are getting sued someone or maybe they are getting sued and they'll use legal research to and they'll use legal research to and they'll use legal research to support and help their case.
-
support and help their case. support and help their case. So if we start at number one, project So if we start at number one, project So if we start at number one, project search, we've been through a little bit search, we've been through a little bit search, we've been through a little bit of a ride here on how we do search. Um, of a ride here on how we do search. Um, of a ride here on how we do search. Um, starting at, you know, hundreds of starting at, you know, hundreds of starting at, you know, hundreds of thousands of documents all the way into thousands of documents all the way into thousands of documents all the way into two billions of documents. And we've two billions of documents. And we've two billions of documents. And we've tried a lot of different things. tried a lot of different things. tried a lot of different things. So first we start started with the very So first we start started with the very So first we start started with the very very simple one, which is just a single very simple one, which is just a single very simple one, which is just a single elastic search cluster for all of our elastic search cluster for all of our elastic search cluster for all of our search workloads. Um, that worked search workloads. Um, that worked search workloads. Um, that worked relatively well. It was sort of a simple relatively well. It was sort of a simple relatively well. It was sort of a simple setup. All of the tenants, you know, our setup. All of the tenants, you know, our setup. All of the tenants, you know, our clients, our users would be on one big clients, our users would be on one big clients, our users would be on one big blob storage where we store the raw blob storage where we store the raw blob storage where we store the raw documents and on one big elastic search documents and on one big elastic search documents and on one big elastic search where we would do all of the searching, where we would do all of the searching, where we would do all of the searching, the indexing and the searching. the indexing and the searching. the indexing and the searching. Super simple, worked relatively well Super simple, worked relatively well Super simple, worked relatively well initially. initially. initially. Then we wanted to enter the land of the Then we wanted to enter the land of the Then we wanted to enter the land of the free and uh we got some new free and uh we got some new free and uh we got some new requirements. They you know Americans requirements. They you know Americans requirements. They you know Americans only want processing to happen within only want processing to happen within only want processing to happen within the US and Europeans only want it to the US and Europeans only want it to the US and Europeans only want it to happen within the EU and Australians happen within the EU and Australians happen within the EU and Australians only want it to happen within Australia. only want it to happen within Australia. only want it to happen within Australia. And so we had to basically um well And so we had to basically um well And so we had to basically um well here's a scaling graph. We had to move here's a scaling graph. We had to move here's a scaling graph. We had to move to multiple elastic searches. And what to multiple elastic searches. And what to multiple elastic searches. And what we actually did was we took the entire we actually did was we took the entire we actually did was we took the entire setup and we just basically iterated setup and we just basically iterated setup and we just basically iterated over the set that is EU, US and Asia over the set that is EU, US and Asia over the set that is EU, US and Asia Pacific. And so we just had this like Pacific. And so we just had this like Pacific. And so we just had this like multiplied by three or four. Kind of multiplied by three or four. Kind of multiplied by three or four. Kind of annoying, lots of overhead, but it uh annoying, lots of overhead, but it uh annoying, lots of overhead, but it uh got us to where we needed to be.
-
Then the next iteration of the story is Then the next iteration of the story is enterprise. So really big banks, the enterprise. So really big banks, the enterprise. So really big banks, the biggest law firms in the world, they biggest law firms in the world, they biggest law firms in the world, they have really annoying requirements. And have really annoying requirements. And have really annoying requirements. And number one they have is they'll ask for number one they have is they'll ask for number one they have is they'll ask for full physical isolation of all of their full physical isolation of all of their full physical isolation of all of their data. data. data. There's probably a little bit of a war There's probably a little bit of a war There's probably a little bit of a war on what physical, you know, isolation on what physical, you know, isolation on what physical, you know, isolation actually means, but essentially means actually means, but essentially means actually means, but essentially means they want their own database. They also they want their own database. They also they want their own database. They also want customer managed encryption keys. want customer managed encryption keys. want customer managed encryption keys. And what that means is they basically And what that means is they basically And what that means is they basically have a key volt thing where they have an have a key volt thing where they have an have a key volt thing where they have an encryption key and they give us access encryption key and they give us access encryption key and they give us access to read the key and we then use that key to read the key and we then use that key to read the key and we then use that key to encrypt and decrypt all of their data to encrypt and decrypt all of their data to encrypt and decrypt all of their data at rest. And what that gives them is at rest. And what that gives them is at rest. And what that gives them is they can just revoke our access to their they can just revoke our access to their they can just revoke our access to their key and then we can't decrypt their data key and then we can't decrypt their data key and then we can't decrypt their data anymore and so it's safe. And so in a anymore and so it's safe. And so in a anymore and so it's safe. And so in a way that gives enterprises a lot of way that gives enterprises a lot of way that gives enterprises a lot of control over all of their data because control over all of their data because control over all of their data because they control you know the key to to they control you know the key to to they control you know the key to to reading it. reading it. reading it. So we moved from elastic search to So we moved from elastic search to So we moved from elastic search to Postgress and I imagine a bunch of you Postgress and I imagine a bunch of you Postgress and I imagine a bunch of you guys are like why would you ever put guys are like why would you ever put guys are like why would you ever put your vectors into Postgress. Um it your vectors into Postgress. Um it your vectors into Postgress. Um it actually worked surprisingly well and actually worked surprisingly well and actually worked surprisingly well and the reason that we did this was we were the reason that we did this was we were the reason that we did this was we were already using Postgress for OOLTP already using Postgress for OOLTP already using Postgress for OOLTP workloads and uh so we sort of already workloads and uh so we sort of already workloads and uh so we sort of already had to do this split of multiple had to do this split of multiple had to do this split of multiple Postgresses and multiple blobs and so it Postgresses and multiple blobs and so it Postgresses and multiple blobs and so it was really easy for us to try to shift was really easy for us to try to shift was really easy for us to try to shift all of our search into Postgress as well all of our search into Postgress as well all of our search into Postgress as well because then we only have one system. So because then we only have one system. So because then we only have one system. So the setup here was PG vector uh the setup here was PG vector uh the setup here was PG vector uh specifically disk NNN TS vector for the specifically disk NNN TS vector for the specifically disk NNN TS vector for the search. So not BM25 which was you know search. So not BM25 which was you know search. So not BM25 which was you know we lost a little bit of retrieval we lost a little bit of retrieval we lost a little bit of retrieval performance there. And what we do is we performance there. And what we do is we performance there. And what we do is we would partition the table where we would would partition the table where we would would partition the table where we would store all the document chunks. We'd store all the document chunks. We'd store all the document chunks. We'd partition it super aggressively like
-
partition it super aggressively like partition it super aggressively like 4,000 partitions and then each project 4,000 partitions and then each project 4,000 partitions and then each project we'd basically have the project key and we'd basically have the project key and we'd basically have the project key and we'd binack them into the partitions. we'd binack them into the partitions. we'd binack them into the partitions. that actually worked relatively well, that actually worked relatively well, that actually worked relatively well, but it was expensive and search but it was expensive and search but it was expensive and search performance wasn't super super good. And performance wasn't super super good. And performance wasn't super super good. And what happened was when we scaled a lot, what happened was when we scaled a lot, what happened was when we scaled a lot, uh everything just broke and exploded. uh everything just broke and exploded. uh everything just broke and exploded. And so what happened was the basically And so what happened was the basically And so what happened was the basically you can imagine like you have a bunch of you can imagine like you have a bunch of you can imagine like you have a bunch of projects and some of them you spin up a projects and some of them you spin up a projects and some of them you spin up a project, you work on it and then you project, you work on it and then you project, you work on it and then you close it and you basically never go back close it and you basically never go back close it and you basically never go back to it again. And we have a bunch of to it again. And we have a bunch of to it again. And we have a bunch of those where like they never get queried those where like they never get queried those where like they never get queried and we have a bunch that get queried all and we have a bunch that get queried all and we have a bunch that get queried all the time because they're super active the time because they're super active the time because they're super active projects. And when we pack them into projects. And when we pack them into projects. And when we pack them into partitions, the cold ones and the hot partitions, the cold ones and the hot partitions, the cold ones and the hot ones would land on the same ones and the ones would land on the same ones and the ones would land on the same ones and the partitions would get really really big. partitions would get really really big. partitions would get really really big. And so when we query them, they would And so when we query them, they would And so when we query them, they would Postgress would pull the partition, put Postgress would pull the partition, put Postgress would pull the partition, put it into memory, we do the stuff and then it into memory, we do the stuff and then it into memory, we do the stuff and then we create another partition and another we create another partition and another we create another partition and another partition and it would essentially partition and it would essentially partition and it would essentially thrash the cache all the time. And what thrash the cache all the time. And what thrash the cache all the time. And what that meant was our latencies would that meant was our latencies would that meant was our latencies would spike. So we went from like search and spike. So we went from like search and spike. So we went from like search and ingestion P99 of 100 milliseconds into ingestion P99 of 100 milliseconds into ingestion P99 of 100 milliseconds into 20 seconds, which you can imagine is a 20 seconds, which you can imagine is a 20 seconds, which you can imagine is a really bad user experience.
-
really bad user experience. really bad user experience. So then we went to turbopuffer in about So then we went to turbopuffer in about So then we went to turbopuffer in about you know when we're about I think 400 you know when we're about I think 400 you know when we're about I think 400 million documents something like that million documents something like that million documents something like that and what we did with turbopuffer was we and what we did with turbopuffer was we and what we did with turbopuffer was we did one name space per project and the did one name space per project and the did one name space per project and the advantages of moving to turbuffer is we advantages of moving to turbuffer is we advantages of moving to turbuffer is we got BM25 real BM25 much better uh got BM25 real BM25 much better uh got BM25 real BM25 much better uh relevancy much better latencies and it relevancy much better latencies and it relevancy much better latencies and it was extremely simple to operate because was extremely simple to operate because was extremely simple to operate because we could just have a single turbo puffer we could just have a single turbo puffer we could just have a single turbo puffer cluster you know we didn't have to have cluster you know we didn't have to have cluster you know we didn't have to have a bunch of different ones like with a bunch of different ones like with a bunch of different ones like with Postgress and it would since it's blob Postgress and it would since it's blob Postgress and it would since it's blob based it could just query the blobs that based it could just query the blobs that based it could just query the blobs that we had anyway and so much lower cost and we had anyway and so much lower cost and we had anyway and so much lower cost and it was extremely simple to operate and it was extremely simple to operate and it was extremely simple to operate and we didn't have this problem with the we didn't have this problem with the we didn't have this problem with the partitions because if a project's not partitions because if a project's not partitions because if a project's not used it's just in blob and so it's used it's just in blob and so it's used it's just in blob and so it's really easy and Simon can talk a bit really easy and Simon can talk a bit really easy and Simon can talk a bit more about why that works so well. Yeah. So, Legora has some of and and Yeah. So, Legora has some of and and legal in general has, by the way, if if legal in general has, by the way, if if legal in general has, by the way, if if uh Jake and I have similar accents and uh Jake and I have similar accents and uh Jake and I have similar accents and maybe even look a bit similar, it's maybe even look a bit similar, it's maybe even look a bit similar, it's because we're both Danish. Um, because we're both Danish. Um, because we're both Danish. Um, Turboper has a very has a particular Turboper has a very has a particular Turboper has a very has a particular architecture that supports these kinds architecture that supports these kinds architecture that supports these kinds of very regulated environments really of very regulated environments really of very regulated environments really really well. But in order to understand really well. But in order to understand really well. But in order to understand that, we have to understand what kind of that, we have to understand what kind of that, we have to understand what kind of search engine is Turboper. Why is it search engine is Turboper. Why is it search engine is Turboper. Why is it different than the ones that they used different than the ones that they used different than the ones that they used in the past?
-
in the past? in the past? Since the very beginning of Turbopuffer, Since the very beginning of Turbopuffer, Since the very beginning of Turbopuffer, um the design has more or less been the um the design has more or less been the um the design has more or less been the same. Uh there may be changes in the same. Uh there may be changes in the same. Uh there may be changes in the future, but the design has stood the future, but the design has stood the future, but the design has stood the test of time. When you do a write to test of time. When you do a write to test of time. When you do a write to Turbopuffer, we write directly to obic Turbopuffer, we write directly to obic Turbopuffer, we write directly to obic storage. There is no like disk storage. There is no like disk storage. There is no like disk replication. There's no paxas. There's replication. There's no paxas. There's replication. There's no paxas. There's there's none of that. Direct to S3. there's none of that. Direct to S3. there's none of that. Direct to S3. That's the fundamental trade-off in That's the fundamental trade-off in That's the fundamental trade-off in Turbopuffer, right? Hundreds of Turbopuffer, right? Hundreds of Turbopuffer, right? Hundreds of milliseconds. If you're like Shopify and milliseconds. If you're like Shopify and milliseconds. If you're like Shopify and doing inventory uh reservations for a doing inventory uh reservations for a doing inventory uh reservations for a Kylie Jenner flash sale not going to Kylie Jenner flash sale not going to Kylie Jenner flash sale not going to work very very good for search because work very very good for search because work very very good for search because generally when you're doing search doing generally when you're doing search doing generally when you're doing search doing a slow write is fine um as long as the a slow write is fine um as long as the a slow write is fine um as long as the read performance is is adaptable and read performance is is adaptable and read performance is is adaptable and good. So that's what happens on write it good. So that's what happens on write it good. So that's what happens on write it just goes into write ahead log you can just goes into write ahead log you can just goes into write ahead log you can imagine you write 1.json 2.json 3.json imagine you write 1.json 2.json 3.json imagine you write 1.json 2.json 3.json JSON. Obviously, it's a it's a database, JSON. Obviously, it's a it's a database, JSON. Obviously, it's a it's a database, so it's not JSON, but for illustrative so it's not JSON, but for illustrative so it's not JSON, but for illustrative purposes, that's what happens. And in purposes, that's what happens. And in purposes, that's what happens. And in the background, we build the vector the background, we build the vector the background, we build the vector indexes, the text indexes, the colmer indexes, the text indexes, the colmer indexes, the text indexes, the colmer indexes, and so on to satisfy the indexes, and so on to satisfy the indexes, and so on to satisfy the queries that that Jake and other queries that that Jake and other queries that that Jake and other customers have. Um, so then at query customers have. Um, so then at query customers have. Um, so then at query time, we can go in and then um the query time, we can go in and then um the query time, we can go in and then um the query reaches some namespace. The namespace is reaches some namespace. The namespace is reaches some namespace. The namespace is kind of our concept of a of a table. Um, kind of our concept of a of a table. Um, kind of our concept of a of a table. Um, you can think of it as a directory on S3 you can think of it as a directory on S3 you can think of it as a directory on S3 that's isolated from everything else. We that's isolated from everything else. We that's isolated from everything else. We go to the node that is most likely to go to the node that is most likely to go to the node that is most likely to have it. It could go to any node, right?
-
have it. It could go to any node, right? have it. It could go to any node, right? It could go to every single node and It could go to every single node and It could go to every single node and they're all read replicas, but we go they're all read replicas, but we go they're all read replicas, but we go with some affinity to the node that has with some affinity to the node that has with some affinity to the node that has the highest probability of having it in the highest probability of having it in the highest probability of having it in cache. We check the memory cache for any cache. We check the memory cache for any cache. We check the memory cache for any objects, NVME SSD cache, and then objects, NVME SSD cache, and then objects, NVME SSD cache, and then finally to object storage. Everything in finally to object storage. Everything in finally to object storage. Everything in Turbopuffer is optimized around doing as Turbopuffer is optimized around doing as Turbopuffer is optimized around doing as much work in as few round trips as much work in as few round trips as much work in as few round trips as possible. Right? S3 has a P99 um on a possible. Right? S3 has a P99 um on a possible. Right? S3 has a P99 um on a like 1 megaby blob size of around uh 200 like 1 megaby blob size of around uh 200 like 1 megaby blob size of around uh 200 milliseconds. So, you want to do as few milliseconds. So, you want to do as few milliseconds. So, you want to do as few round trips as possible, right? Ideally, round trips as possible, right? Ideally, round trips as possible, right? Ideally, you do around three. And everything in you do around three. And everything in you do around three. And everything in Turbopuffer, the database is Oh, I'm Turbopuffer, the database is Oh, I'm Turbopuffer, the database is Oh, I'm going to need your fingerprint. going to need your fingerprint. going to need your fingerprint. >> Got it. >> Got it. >> Got it. >> Um, everything in in Turbopuffer is >> Um, everything in in Turbopuffer is >> Um, everything in in Turbopuffer is designed around minimizing the number of designed around minimizing the number of designed around minimizing the number of round trips. This is also amazing for round trips. This is also amazing for round trips. This is also amazing for modern D discs. If you do a lot of modern D discs. If you do a lot of modern D discs. If you do a lot of concurrency and few round trips, you concurrency and few round trips, you concurrency and few round trips, you utilize them optimally. And everything utilize them optimally. And everything utilize them optimally. And everything in Turbopuffer is designed around this. in Turbopuffer is designed around this. in Turbopuffer is designed around this. So, why is this so good for a company So, why is this so good for a company So, why is this so good for a company like Lora? Well, object storage native like Lora? Well, object storage native like Lora? Well, object storage native if you design it around the atomic unit if you design it around the atomic unit if you design it around the atomic unit of separation being the namespace or the of separation being the namespace or the of separation being the namespace or the table, every single table could be table, every single table could be table, every single table could be encrypted with a different key. Every encrypted with a different key. Every encrypted with a different key. Every single namespace could be in a different single namespace could be in a different single namespace could be in a different bucket. We have customers that have um bucket. We have customers that have um bucket. We have customers that have um thousands of buckets that they have thousands of buckets that they have thousands of buckets that they have namespaces in so that their customers namespaces in so that their customers namespaces in so that their customers get the warm IT fuzzies of having the get the warm IT fuzzies of having the get the warm IT fuzzies of having the bucket in their own cloud account. They bucket in their own cloud account. They bucket in their own cloud account. They can also be encrypted with their own can also be encrypted with their own can also be encrypted with their own keys. You can share buckets. You can you keys. You can share buckets. You can you keys. You can share buckets. You can you can do whatever configuration that you can do whatever configuration that you can do whatever configuration that you need at the namespace level. You can need at the namespace level. You can need at the namespace level. You can re-encrypt with different keys. You can re-encrypt with different keys. You can re-encrypt with different keys. You can move them around. Um and you can move them around. Um and you can move them around. Um and you can re-encrypt with other keys. Um for Lora re-encrypt with other keys. Um for Lora re-encrypt with other keys. Um for Lora in particular, this was really important in particular, this was really important in particular, this was really important for this full physical separation,
-
for this full physical separation, for this full physical separation, right? An encryption separation. All of right? An encryption separation. All of right? An encryption separation. All of the namespaces needed to be physically the namespaces needed to be physically the namespaces needed to be physically at rest with different keys and as at rest with different keys and as at rest with different keys and as separated as possible. S3, GCS, Azure separated as possible. S3, GCS, Azure separated as possible. S3, GCS, Azure BOP stories, they pass that um and the BOP stories, they pass that um and the BOP stories, they pass that um and the other parts of the hierarchy also except other parts of the hierarchy also except other parts of the hierarchy also except the NVME SSD cache because in the SSD the NVME SSD cache because in the SSD the NVME SSD cache because in the SSD cache we consider that to be volatile cache we consider that to be volatile cache we consider that to be volatile like memory um but your customers did like memory um but your customers did like memory um but your customers did not. So in the uh so what we did was not. So in the uh so what we did was not. So in the uh so what we did was that we thought we were going to that we thought we were going to that we thought we were going to implement encryption into the disc cache implement encryption into the disc cache implement encryption into the disc cache but instead we just disabled the disc but instead we just disabled the disc but instead we just disabled the disc cache and saw how it fared and the cache and saw how it fared and the cache and saw how it fared and the performance of turbopuffer even without performance of turbopuffer even without performance of turbopuffer even without the disc cache with just the memory the disc cache with just the memory the disc cache with just the memory cache was so good that we just kept it cache was so good that we just kept it cache was so good that we just kept it that way for some of the lora workloads that way for some of the lora workloads that way for some of the lora workloads where we couldn't have the disc cache where we couldn't have the disc cache where we couldn't have the disc cache for multi-tenency turbuffer will support for multi-tenency turbuffer will support for multi-tenency turbuffer will support that in the future but it just goes to that in the future but it just goes to that in the future but it just goes to show the um natural point where show the um natural point where show the um natural point where turbopuffer allows this encryption and turbopuffer allows this encryption and turbopuffer allows this encryption and storage and separation to become fully storage and separation to become fully storage and separation to become fully multi-tenency uh native. I'll hand it multi-tenency uh native. I'll hand it multi-tenency uh native. I'll hand it back to you on what happened then. back to you on what happened then. back to you on what happened then. >> And then, drum roll please. Latencies >> And then, drum roll please. Latencies >> And then, drum roll please. Latencies look like this. Um, is my mic working?
-
look like this. Um, is my mic working? look like this. Um, is my mic working? No. Could I or I'll start screaming No. Could I or I'll start screaming No. Could I or I'll start screaming really loudly. Um, it speaks for itself really loudly. Um, it speaks for itself really loudly. Um, it speaks for itself if you can't hear me. Okay. Um, latency if you can't hear me. Okay. Um, latency if you can't hear me. Okay. Um, latency is improved in order magnitude is improved in order magnitude is improved in order magnitude basically. And these are median basically. And these are median basically. And these are median latencies. So, P99 were even better. Um, latencies. So, P99 were even better. Um, latencies. So, P99 were even better. Um, so obviously this is a huge thing when so obviously this is a huge thing when so obviously this is a huge thing when you're doing I mean one thing is if you're doing I mean one thing is if you're doing I mean one thing is if you're doing a single sort of rag style you're doing a single sort of rag style you're doing a single sort of rag style thing but if you have an agent that does thing but if you have an agent that does thing but if you have an agent that does 20 queries 100 queries these really 20 queries 100 queries these really 20 queries 100 queries these really really add up. really add up. really add up. So that was on the project side and then So that was on the project side and then So that was on the project side and then a more recent thing is legal research. a more recent thing is legal research. a more recent thing is legal research. So legal research um is a kind of a So legal research um is a kind of a So legal research um is a kind of a difficult problem and the reason it's difficult problem and the reason it's difficult problem and the reason it's difficult is that um the corpus is difficult is that um the corpus is difficult is that um the corpus is extremely big. So we're racing towards extremely big. So we're racing towards extremely big. So we're racing towards 10 billion vectors and we're growing 10 billion vectors and we're growing 10 billion vectors and we're growing extremely fast. We also have quite high extremely fast. We also have quite high extremely fast. We also have quite high read. Um so QPS can spike a lot because read. Um so QPS can spike a lot because read. Um so QPS can spike a lot because we do a lot of fan out. Like if you do a we do a lot of fan out. Like if you do a we do a lot of fan out. Like if you do a a sort of legal research query, we will a sort of legal research query, we will a sort of legal research query, we will fan it out to a bunch of different fan it out to a bunch of different fan it out to a bunch of different queries and we'll keep going. And the queries and we'll keep going. And the queries and we'll keep going. And the reason we do that is we need this um reason we do that is we need this um reason we do that is we need this um heavy filtering because essentially it's heavy filtering because essentially it's heavy filtering because essentially it's a kind of like a graph for a few a kind of like a graph for a few a kind of like a graph for a few different reasons. Firstly, it's um different reasons. Firstly, it's um different reasons. Firstly, it's um hierarchical. you know, you have cities hierarchical. you know, you have cities hierarchical. you know, you have cities and you have counties and you have and you have counties and you have and you have counties and you have states and you have federal law and it's states and you have federal law and it's states and you have federal law and it's the same all around the world. Um, and the same all around the world. Um, and the same all around the world. Um, and so you need to respect that so you need to respect that so you need to respect that authoritative sort of hierarchy. There's authoritative sort of hierarchy. There's authoritative sort of hierarchy. There's also some temporal validity. So, one also some temporal validity. So, one also some temporal validity. So, one judge might overrule a decision that's judge might overrule a decision that's judge might overrule a decision that's been made somewhere else and you need to been made somewhere else and you need to been made somewhere else and you need to also respect that and figure that out.
-
also respect that and figure that out. also respect that and figure that out. And then sometimes there's even like a And then sometimes there's even like a And then sometimes there's even like a new regulation that has exemptions or new regulation that has exemptions or new regulation that has exemptions or special cases of an old regulation. And special cases of an old regulation. And special cases of an old regulation. And so if you're finding this one, you need so if you're finding this one, you need so if you're finding this one, you need to find all the other ones as well. So to find all the other ones as well. So to find all the other ones as well. So you can imagine that it sort of explodes you can imagine that it sort of explodes you can imagine that it sort of explodes the search the search the search and so we started on elastic search for and so we started on elastic search for and so we started on elastic search for this but also moving to turbuffer um this but also moving to turbuffer um this but also moving to turbuffer um elastic search just got extremely elastic search just got extremely elastic search just got extremely expensive because we have to have expensive because we have to have expensive because we have to have everything there but um with turbopuffer everything there but um with turbopuffer everything there but um with turbopuffer we can basically uh take different we can basically uh take different we can basically uh take different jurisdictions and we can make them name jurisdictions and we can make them name jurisdictions and we can make them name spaces in turbo buffer and that means spaces in turbo buffer and that means spaces in turbo buffer and that means some of them here's an example where some of them here's an example where some of them here's an example where like you have the EU that gets quered like you have the EU that gets quered like you have the EU that gets quered all the time that's super hot and some all the time that's super hot and some all the time that's super hot and some of them let's say Danish law cuz we're of them let's say Danish law cuz we're of them let's say Danish law cuz we're Danish, no one cares really. It's such a Danish, no one cares really. It's such a Danish, no one cares really. It's such a small country, so like it doesn't really small country, so like it doesn't really small country, so like it doesn't really get queried and so that can just stay on get queried and so that can just stay on get queried and so that can just stay on blob and that's fine. Um, and because blob and that's fine. Um, and because blob and that's fine. Um, and because it's sort of a deep research style it's sort of a deep research style it's sort of a deep research style workload, if there's 500 milliseconds workload, if there's 500 milliseconds workload, if there's 500 milliseconds latency to fetch that cold blob, that's latency to fetch that cold blob, that's latency to fetch that cold blob, that's okay. That's fine. It's not really a big okay. That's fine. It's not really a big okay. That's fine. It's not really a big problem. So, the way that Turbo is problem. So, the way that Turbo is problem. So, the way that Turbo is designed lends itself super well to this designed lends itself super well to this designed lends itself super well to this super long scale of like cold weird name super long scale of like cold weird name super long scale of like cold weird name spaces and a few that are really, really spaces and a few that are really, really spaces and a few that are really, really hot. Um, yeah. And Simon wants to talk hot. Um, yeah. And Simon wants to talk hot. Um, yeah. And Simon wants to talk more about that.
-
more about that. more about that. Yeah. So, um um I was I was talking Yeah. So, um um I was I was talking Yeah. So, um um I was I was talking about why the company is called about why the company is called about why the company is called Turbopuffer at another uh talk here Turbopuffer at another uh talk here Turbopuffer at another uh talk here earlier today, but uh one of the other earlier today, but uh one of the other earlier today, but uh one of the other explanations of the name of Turbopuffer explanations of the name of Turbopuffer explanations of the name of Turbopuffer is that it's about puffing into the is that it's about puffing into the is that it's about puffing into the different memory hierarchies and really different memory hierarchies and really different memory hierarchies and really mastering when data should be in mastering when data should be in mastering when data should be in particular memory hierarchies. So, you particular memory hierarchies. So, you particular memory hierarchies. So, you can think about it here, right, of can think about it here, right, of can think about it here, right, of something like the EU law might be more something like the EU law might be more something like the EU law might be more or less part of almost every one of the or less part of almost every one of the or less part of almost every one of the legal research uh queries, right? So legal research uh queries, right? So legal research uh queries, right? So that probably sits closer to NVME SSDs that probably sits closer to NVME SSDs that probably sits closer to NVME SSDs in memory, right? The economics kind of in memory, right? The economics kind of in memory, right? The economics kind of change as you move up and down this change as you move up and down this change as you move up and down this hierarchy. Um in in memory, you want hierarchy. Um in in memory, you want hierarchy. Um in in memory, you want things that are queried a lot, right? things that are queried a lot, right? things that are queried a lot, right? Then the economics of memory are great. Then the economics of memory are great. Then the economics of memory are great. NVME SSDs can you can do a lot of things NVME SSDs can you can do a lot of things NVME SSDs can you can do a lot of things directly on them, but the economics directly on them, but the economics directly on them, but the economics change as you move up and down this change as you move up and down this change as you move up and down this boundary, the latency changes and the boundary, the latency changes and the boundary, the latency changes and the way that the database is architected to way that the database is architected to way that the database is architected to take advantage of it in terms of round take advantage of it in terms of round take advantage of it in terms of round trips versus random versus sequential trips versus random versus sequential trips versus random versus sequential all changes as you navigate this all changes as you navigate this all changes as you navigate this hierarchy. Turbopuffer is a database hierarchy. Turbopuffer is a database hierarchy. Turbopuffer is a database that is really designed around the that is really designed around the that is really designed around the memory hierarchy and all of the smarts memory hierarchy and all of the smarts memory hierarchy and all of the smarts in Turbopuffer is that all of these in Turbopuffer is that all of these in Turbopuffer is that all of these namespaces are puffed in and out um of namespaces are puffed in and out um of namespaces are puffed in and out um of the cache. You can think of this as we the cache. You can think of this as we the cache. You can think of this as we want to spend as much time have as much want to spend as much time have as much want to spend as much time have as much data pushed as far down in this data pushed as far down in this data pushed as far down in this hierarchy as possible to get the the hierarchy as possible to get the the hierarchy as possible to get the the best um performance cost ratios. So how best um performance cost ratios. So how best um performance cost ratios. So how does that apply to search? Well, for does that apply to search? Well, for does that apply to search? Well, for something like vector search for something like vector search for something like vector search for example, there's two fundamental ways to example, there's two fundamental ways to example, there's two fundamental ways to do vector search. One is to navigate it do vector search. One is to navigate it do vector search. One is to navigate it basically design a graph. The problem basically design a graph. The problem basically design a graph. The problem with a graph on something like object with a graph on something like object with a graph on something like object storage or disk, again, we want to have
-
storage or disk, again, we want to have storage or disk, again, we want to have things as far down that memory hierarchy things as far down that memory hierarchy things as far down that memory hierarchy as possible. The problem with a graph, as possible. The problem with a graph, as possible. The problem with a graph, this is not a graph, this is a tree. Um, this is not a graph, this is a tree. Um, this is not a graph, this is a tree. Um, but in a graph, you have to navigate but in a graph, you have to navigate but in a graph, you have to navigate from the center of the graph. So, and from the center of the graph. So, and from the center of the graph. So, and then every time you navigate through then every time you navigate through then every time you navigate through these nodes, you're doing 200 these nodes, you're doing 200 these nodes, you're doing 200 millisecond P99 to S3, right? And so, millisecond P99 to S3, right? And so, millisecond P99 to S3, right? And so, you're trying to shrink the diameter of you're trying to shrink the diameter of you're trying to shrink the diameter of the graph. You're trying to do all these the graph. You're trying to do all these the graph. You're trying to do all these tricks to make the graph, but tricks to make the graph, but tricks to make the graph, but fundamentally you're at odds with the fundamentally you're at odds with the fundamentally you're at odds with the fact that graph is about a random fact that graph is about a random fact that graph is about a random sequential trade-off that you have in sequential trade-off that you have in sequential trade-off that you have in memory and in registers, but not further memory and in registers, but not further memory and in registers, but not further down the memory hierarchy. The way down the memory hierarchy. The way down the memory hierarchy. The way Turbopuffer does it is organize it into Turbopuffer does it is organize it into Turbopuffer does it is organize it into clusters, right? Vectors you can think clusters, right? Vectors you can think clusters, right? Vectors you can think of in two dimensions just as point is a of in two dimensions just as point is a of in two dimensions just as point is a massive coordinate system and we can massive coordinate system and we can massive coordinate system and we can organize them into clusters. Turbopuffer organize them into clusters. Turbopuffer organize them into clusters. Turbopuffer then creates clusters of clusters and then creates clusters of clusters and then creates clusters of clusters and clusters of clusters of clusters to clusters of clusters of clusters to clusters of clusters of clusters to essentially organize all of the vector essentially organize all of the vector essentially organize all of the vector data in a tree. You can basically think data in a tree. You can basically think data in a tree. You can basically think of turbopuffer as a very very of turbopuffer as a very very of turbopuffer as a very very complicated B tree, right? Because it's complicated B tree, right? Because it's complicated B tree, right? Because it's a tree on this geometry of this entire a tree on this geometry of this entire a tree on this geometry of this entire space and the clustering of it in an space and the clustering of it in an space and the clustering of it in an approximate way. Now the root approximate way. Now the root approximate way. Now the root centroidids further up the tree you can centroidids further up the tree you can centroidids further up the tree you can imagine are part of every single time imagine are part of every single time imagine are part of every single time you search right there. We're always you search right there. We're always you search right there. We're always trying to figure out which clusters that trying to figure out which clusters that trying to figure out which clusters that we're in and we're always looking at the we're in and we're always looking at the we're in and we're always looking at the upper levels of the tree. So they're upper levels of the tree. So they're upper levels of the tree. So they're going to be further of the memory going to be further of the memory going to be further of the memory hierarchy, right? Closer to the hierarchy, right? Closer to the hierarchy, right? Closer to the registers almost all in DRAM. Now the registers almost all in DRAM. Now the registers almost all in DRAM. Now the leaves that have all of the actual legal leaves that have all of the actual legal leaves that have all of the actual legal cases or whatever long document it could cases or whatever long document it could cases or whatever long document it could be could be images all of that is be could be images all of that is be could be images all of that is probably going to be on SSDs with that probably going to be on SSDs with that probably going to be on SSDs with that single 1 millisecond roundtrip at the single 1 millisecond roundtrip at the single 1 millisecond roundtrip at the end. It doesn't make sense to have all end. It doesn't make sense to have all end. It doesn't make sense to have all that puffed into DRAM. This is that puffed into DRAM. This is that puffed into DRAM. This is fundamentally the cheapest way that you
-
fundamentally the cheapest way that you fundamentally the cheapest way that you can run a database period. So for can run a database period. So for can run a database period. So for something like Lora or even web search something like Lora or even web search something like Lora or even web search which is in hundreds of billion or tens which is in hundreds of billion or tens which is in hundreds of billion or tens of billions depending on how much of the of billions depending on how much of the of billions depending on how much of the web you've scraped this is fundamentally web you've scraped this is fundamentally web you've scraped this is fundamentally the cheapest way to do it. And we have the cheapest way to do it. And we have the cheapest way to do it. And we have customers that are indexing massive customers that are indexing massive customers that are indexing massive parts of the entire web into turbopuffer parts of the entire web into turbopuffer parts of the entire web into turbopuffer which is really also a part of what which is really also a part of what which is really also a part of what legal research is. Full text is also legal research is. Full text is also legal research is. Full text is also also really respectful of the memory also really respectful of the memory also really respectful of the memory hierarchies. The way the text search hierarchies. The way the text search hierarchies. The way the text search works is essentially you can think of it works is essentially you can think of it works is essentially you can think of it as a hashmap. You have a big document as a hashmap. You have a big document as a hashmap. You have a big document and then you take every single one of and then you take every single one of and then you take every single one of the tokens and you put them into the key the tokens and you put them into the key the tokens and you put them into the key in the hashmap. The value in the hashmap in the hashmap. The value in the hashmap in the hashmap. The value in the hashmap is some set with all of the document IDs is some set with all of the document IDs is some set with all of the document IDs that has that term. So then if you that has that term. So then if you that has that term. So then if you search for New York population, you're search for New York population, you're search for New York population, you're finding those three places in the finding those three places in the finding those three places in the hashmap and then you're taking the three hashmap and then you're taking the three hashmap and then you're taking the three sets and doing an intersect on the sets and doing an intersect on the sets and doing an intersect on the intersect on the sets. While you're intersect on the sets. While you're intersect on the sets. While you're intersecting, you're also trying to do intersecting, you're also trying to do intersecting, you're also trying to do some kind of scoring, right? A document some kind of scoring, right? A document some kind of scoring, right? A document that has York in it is probably more that has York in it is probably more that has York in it is probably more valuable than a document that has new in valuable than a document that has new in valuable than a document that has new in it because York is a more rare word. it because York is a more rare word. it because York is a more rare word. When people say BM25, this is the When people say BM25, this is the When people say BM25, this is the scoring that they're referring to.
-
scoring that they're referring to. scoring that they're referring to. The art of fulltech search is one, we The art of fulltech search is one, we The art of fulltech search is one, we want to minimize the number of round want to minimize the number of round want to minimize the number of round trips. So first you download the parts trips. So first you download the parts trips. So first you download the parts of the dictionary that are relevant. of the dictionary that are relevant. of the dictionary that are relevant. Round trip one maybe a round trip one Round trip one maybe a round trip one Round trip one maybe a round trip one before that to index into where the before that to index into where the before that to index into where the parts of the term terms are. And then parts of the term terms are. And then parts of the term terms are. And then the second round trip is to get these the second round trip is to get these the second round trip is to get these massive lists. Try to make the list as massive lists. Try to make the list as massive lists. Try to make the list as small as possible by compressing them. small as possible by compressing them. small as possible by compressing them. But also while you're doing the text But also while you're doing the text But also while you're doing the text search you're trying to minimize the search you're trying to minimize the search you're trying to minimize the amount again of memory bandwidth that amount again of memory bandwidth that amount again of memory bandwidth that you want to intersect these list. You you want to intersect these list. You you want to intersect these list. You can probably imagine that at some point can probably imagine that at some point can probably imagine that at some point there's a point where you've seen so there's a point where you've seen so there's a point where you've seen so many documents with population in York many documents with population in York many documents with population in York that have much higher scores that that have much higher scores that that have much higher scores that documents that just have new in it are documents that just have new in it are documents that just have new in it are irrelevant anymore. This is like a mega irrelevant anymore. This is like a mega irrelevant anymore. This is like a mega crash course in how text search works. crash course in how text search works. crash course in how text search works. And counterintuitively to most people, And counterintuitively to most people, And counterintuitively to most people, text search at web scale is more text search at web scale is more text search at web scale is more difficult and more computationally difficult and more computationally difficult and more computationally expensive than doing vector search. I'll expensive than doing vector search. I'll expensive than doing vector search. I'll hand it over to you. hand it over to you. hand it over to you. Cool. So key learnings from um what you Cool. So key learnings from um what you Cool. So key learnings from um what you heard today, retrieval is extremely heard today, retrieval is extremely heard today, retrieval is extremely important to legora. Um it's key to important to legora. Um it's key to important to legora. Um it's key to legal reasoning. Turbopuffer really legal reasoning. Turbopuffer really legal reasoning. Turbopuffer really excels when uh for us because it makes excels when uh for us because it makes excels when uh for us because it makes it extremely easy to operate. We have 70 it extremely easy to operate. We have 70 it extremely easy to operate. We have 70 plus tenants. We have 100, we have 200 plus tenants. We have 100, we have 200 plus tenants. We have 100, we have 200 tenants. You know, if we had to have tenants. You know, if we had to have tenants. You know, if we had to have separate elastic search databases for separate elastic search databases for separate elastic search databases for each of these, it would be just hell. Um each of these, it would be just hell. Um each of these, it would be just hell. Um but we can do this natively with but we can do this natively with but we can do this natively with turboper with data residency and cmech, turboper with data residency and cmech, turboper with data residency and cmech, etc., etc. And then it's extremely cost etc., etc. And then it's extremely cost etc., etc. And then it's extremely cost efficient generally when you have these efficient generally when you have these efficient generally when you have these types of workflows um or workloads that types of workflows um or workloads that types of workflows um or workloads that we do where there's a long tail of cold
-
we do where there's a long tail of cold we do where there's a long tail of cold indices basically that you don't need to indices basically that you don't need to indices basically that you don't need to query so much and you don't you're sort query so much and you don't you're sort query so much and you don't you're sort of you're okay paying the small latency of you're okay paying the small latency of you're okay paying the small latency cost for it. So now with um turbo buffer cost for it. So now with um turbo buffer cost for it. So now with um turbo buffer in 4 seconds to go now we can focus on in 4 seconds to go now we can focus on in 4 seconds to go now we can focus on making lora we can focus on the product making lora we can focus on the product making lora we can focus on the product making it really really great and not on making it really really great and not on making it really really great and not on scalability and infra and also David our scalability and infra and also David our scalability and infra and also David our CFO is really happy about the cost. So, CFO is really happy about the cost. So, CFO is really happy about the cost. So, it's great. Thanks everyone. it's great. Thanks everyone. it's great. Thanks everyone. [applause]
No summary available yet.
View original episode ↗