← Back
AI Engineer September 16, 2026 18m

Your Agreements Are a Database You Can't Query — Hiral Shah, Docusign & Sean Sodha, NVIDIA

Read full transcript 13 segments
  1. Uh, welcome everyone to our session. Um, Uh, welcome everyone to our session. Um, I'm Haral. I'm a senior director of I'm Haral. I'm a senior director of I'm Haral. I'm a senior director of products um at DocYsine and I'm joined products um at DocYsine and I'm joined products um at DocYsine and I'm joined by Sean. by Sean. by Sean. >> Hello everyone. I'm a product manager at >> Hello everyone. I'm a product manager at >> Hello everyone. I'm a product manager at NVIDIA. So today Sean and I are going to NVIDIA. So today Sean and I are going to NVIDIA. So today Sean and I are going to talk about a massive problem that every talk about a massive problem that every talk about a massive problem that every enterprise faces which is agreement data enterprise faces which is agreement data enterprise faces which is agreement data largecale agreement data. Agreements are largecale agreement data. Agreements are largecale agreement data. Agreements are a big part of any relationship any B-2B a big part of any relationship any B-2B a big part of any relationship any B-2B organization kind of goes through day in organization kind of goes through day in organization kind of goes through day in day out and a lot of that data is day out and a lot of that data is day out and a lot of that data is captured inside that agreements and it's captured inside that agreements and it's captured inside that agreements and it's very critical whether it's pricing very critical whether it's pricing very critical whether it's pricing tables a lot of it and it's all in a lot tables a lot of it and it's all in a lot tables a lot of it and it's all in a lot of different unstructured format and of different unstructured format and of different unstructured format and that's kind of what we're going to show that's kind of what we're going to show that's kind of what we're going to show is how docysine partner with Nvidia are is how docysine partner with Nvidia are is how docysine partner with Nvidia are fixing that on making that data fixing that on making that data fixing that on making that data available readable usable for a lot of available readable usable for a lot of available readable usable for a lot of our kind of organizations. our kind of organizations. our kind of organizations. So here is kind of just a quick map of So here is kind of just a quick map of So here is kind of just a quick map of our talk today. We'll start with just our talk today. We'll start with just our talk today. We'll start with just the stakes. Why does this matter? Why the stakes. Why does this matter? Why the stakes. Why does this matter? Why the scale is so large? And then we'll the scale is so large? And then we'll the scale is so large? And then we'll dive deep into the technical dive deep into the technical dive deep into the technical architecture of how we are approaching architecture of how we are approaching architecture of how we are approaching it, how we have tackled this thing, it, how we have tackled this thing, it, how we have tackled this thing, especially this document processing at especially this document processing at especially this document processing at at scale. And then finally, we'll cover at scale. And then finally, we'll cover at scale. And then finally, we'll cover what we have learned from our evaluation what we have learned from our evaluation what we have learned from our evaluation of all of the different models we've of all of the different models we've of all of the different models we've tried for different purposes. and share tried for different purposes. and share tried for different purposes. and share our learnings with you. So why like you our learnings with you. So why like you our learnings with you. So why like you know just to understand why this is a know just to understand why this is a know just to understand why this is a big problem when you think about big problem when you think about big problem when you think about docysine raise your hands how many of docysine raise your hands how many of docysine raise your hands how many of you have used docuign anyone who's you have used docuign anyone who's you have used docuign anyone who's employed probably the HR docs right so

  2. employed probably the HR docs right so employed probably the HR docs right so it's a massive scale everyone uses docy it's a massive scale everyone uses docy it's a massive scale everyone uses docy for us it's a massive engineering for us it's a massive engineering for us it's a massive engineering problem as well because just look at the problem as well because just look at the problem as well because just look at the scale we have 1.9 million customers who scale we have 1.9 million customers who scale we have 1.9 million customers who are paying us and a billion users what are paying us and a billion users what are paying us and a billion users what does that imply we process us a million does that imply we process us a million does that imply we process us a million agreements a day that needs to now agreements a day that needs to now agreements a day that needs to now structurize make it readable make it structurize make it readable make it structurize make it readable make it querable usable and you know in the past querable usable and you know in the past querable usable and you know in the past we've worked with Deote on a study and we've worked with Deote on a study and we've worked with Deote on a study and it says that there's $2 trillion it says that there's $2 trillion it says that there's $2 trillion captured in this agreement negotiated captured in this agreement negotiated captured in this agreement negotiated value that no one capitalizes no one value that no one capitalizes no one value that no one capitalizes no one goes back and gets that um data back goes back and gets that um data back goes back and gets that um data back right and why it's because they have to right and why it's because they have to right and why it's because they have to do a lot of human reading human reviews do a lot of human reading human reviews do a lot of human reading human reviews there's disconnected systems lot of there's disconnected systems lot of there's disconnected systems lot of manual workflow flows that are there. So manual workflow flows that are there. So manual workflow flows that are there. So that's kind of why docysine built AM an that's kind of why docysine built AM an that's kind of why docysine built AM an intelligent agreement management intelligent agreement management intelligent agreement management platform that takes the entire like platform that takes the entire like platform that takes the entire like applying in an AI first way the entire applying in an AI first way the entire applying in an AI first way the entire agreement life cycle whether you are agreement life cycle whether you are agreement life cycle whether you are creating agreements genai helps a lot creating agreements genai helps a lot creating agreements genai helps a lot with that whether you're negotiating to with that whether you're negotiating to with that whether you're negotiating to understanding and redlinining all the understanding and redlinining all the understanding and redlinining all the way to after signing storing and making way to after signing storing and making way to after signing storing and making a lot of insights from this data. So a lot of insights from this data. So a lot of insights from this data. So when you think about the challenges when you think about the challenges when you think about the challenges involved right in an agreement there is involved right in an agreement there is involved right in an agreement there is the unstructured data it could be a PDF the unstructured data it could be a PDF the unstructured data it could be a PDF it could be a PNG and what are people it could be a PNG and what are people it could be a PNG and what are people wanting to do is simple questions they wanting to do is simple questions they wanting to do is simple questions they can't get that that is the data that is can't get that that is the data that is can't get that that is the data that is trapped inside one agreement but also

  3. trapped inside one agreement but also trapped inside one agreement but also the whole corpus of millions of the whole corpus of millions of the whole corpus of millions of agreement 10 years 20 years of business agreement 10 years 20 years of business agreement 10 years 20 years of business has kind of put into that. So for has kind of put into that. So for has kind of put into that. So for agreements aren't flat they're also agreements aren't flat they're also agreements aren't flat they're also hierarchical in nature they are like you hierarchical in nature they are like you hierarchical in nature they are like you know one agreement governs the other the know one agreement governs the other the know one agreement governs the other the other kind of does something else so you other kind of does something else so you other kind of does something else so you always are needing lot of things to always are needing lot of things to always are needing lot of things to answer this question simple thing I'm answer this question simple thing I'm answer this question simple thing I'm sure you all are using a lot of right sure you all are using a lot of right sure you all are using a lot of right claude and all especially at your claude and all especially at your claude and all especially at your company's organization simple thing what company's organization simple thing what company's organization simple thing what did we contract for the total tokens did we contract for the total tokens did we contract for the total tokens right with claude no one knows that's right with claude no one knows that's right with claude no one knows that's captured inside this agreement in captured inside this agreement in captured inside this agreement in different forms and fashion so you need different forms and fashion so you need different forms and fashion so you need to be extracting this data to find to be extracting this data to find to be extracting this data to find things but also insights and push it things but also insights and push it things but also insights and push it downstream where you're tracking doing downstream where you're tracking doing downstream where you're tracking doing more things and when you analyze an more things and when you analyze an more things and when you analyze an enterprise contract a big set of things enterprise contract a big set of things enterprise contract a big set of things are captured what I call vital terms are captured what I call vital terms are captured what I call vital terms like pricing tiers the skues the like pricing tiers the skues the like pricing tiers the skues the information SLAs's rate cards they're information SLAs's rate cards they're information SLAs's rate cards they're all in table format now traditional all in table format now traditional all in table format now traditional document extraction tools or a generic document extraction tools or a generic document extraction tools or a generic VM BLM completely fail here like you VM BLM completely fail here like you VM BLM completely fail here like you know we've tried we've definitely done know we've tried we've definitely done know we've tried we've definitely done this because they're reading text line this because they're reading text line this because they're reading text line by line which breaks a lot of that by line which breaks a lot of that by line which breaks a lot of that concept within the table. A merge sells concept within the table. A merge sells concept within the table. A merge sells some things are not boundaried. So this some things are not boundaried. So this some things are not boundaried. So this makes a massive operational overhead for makes a massive operational overhead for makes a massive operational overhead for the downstream legal teams, procurement the downstream legal teams, procurement the downstream legal teams, procurement team, sales teams to get that queries team, sales teams to get that queries team, sales teams to get that queries and get that answers done and they spend and get that answers done and they spend and get that answers done and they spend hours and hours digging through this um hours and hours digging through this um hours and hours digging through this um to even just locate a basic thing. And to even just locate a basic thing. And to even just locate a basic thing. And that's kind of where we partnered with that's kind of where we partnered with that's kind of where we partnered with Nvidia and leveraged a purpose-built

  4. Nvidia and leveraged a purpose-built Nvidia and leveraged a purpose-built model like you know tool which is model like you know tool which is model like you know tool which is literally like you know making our literally like you know making our literally like you know making our architecture for table extraction take architecture for table extraction take architecture for table extraction take us to make and solve these complex use us to make and solve these complex use us to make and solve these complex use cases. We're trying with this we're cases. We're trying with this we're cases. We're trying with this we're making things scalable. We can making things scalable. We can making things scalable. We can understand it with the layout but also understand it with the layout but also understand it with the layout but also deliver really accurate results. And to deliver really accurate results. And to deliver really accurate results. And to share more about how we're leveraging share more about how we're leveraging share more about how we're leveraging the Neotron, I'm going to hand it to the Neotron, I'm going to hand it to the Neotron, I'm going to hand it to Sean. All right. Hello everyone. Uh so real All right. Hello everyone. Uh so real quick on uh the Neotron retriever quick on uh the Neotron retriever quick on uh the Neotron retriever initiative. So for those who here knows initiative. So for those who here knows initiative. So for those who here knows about Neotron, you raise your hand real about Neotron, you raise your hand real about Neotron, you raise your hand real quick. Awesome. Uh so Neimatron is all quick. Awesome. Uh so Neimatron is all quick. Awesome. Uh so Neimatron is all about building worldclass open-source about building worldclass open-source about building worldclass open-source models and publishing the data sets, the models and publishing the data sets, the models and publishing the data sets, the techniques, uh the quantization techniques, uh the quantization techniques, uh the quantization approaches, distillation approaches, approaches, distillation approaches, approaches, distillation approaches, pruning approaches, every technique pruning approaches, every technique pruning approaches, every technique possible, blueprints to go with that, possible, blueprints to go with that, possible, blueprints to go with that, you name it. Um throughout the Neimatron you name it. Um throughout the Neimatron you name it. Um throughout the Neimatron portfolio, we have specifically Neatron portfolio, we have specifically Neatron portfolio, we have specifically Neatron Retriever, which is building embedding Retriever, which is building embedding Retriever, which is building embedding models, reranking models, and document models, reranking models, and document models, reranking models, and document extraction models. Um so real quick here extraction models. Um so real quick here extraction models. Um so real quick here we sort of our first initiative is if we sort of our first initiative is if we sort of our first initiative is if you're a large scale enterprise that you're a large scale enterprise that you're a large scale enterprise that deals with pabytes scale data our first deals with pabytes scale data our first deals with pabytes scale data our first initiative is how do we make sure that initiative is how do we make sure that initiative is how do we make sure that you find the right document given a you find the right document given a you find the right document given a certain query your agent sends you know certain query your agent sends you know certain query your agent sends you know set of queries to the corpus afterwards.

  5. set of queries to the corpus afterwards. set of queries to the corpus afterwards. Once you find those top five quer Once you find those top five quer Once you find those top five quer documents whatever it may be then we say documents whatever it may be then we say documents whatever it may be then we say okay you found the right document now okay you found the right document now okay you found the right document now how do you then find the right how do you then find the right how do you then find the right information within the document and this information within the document and this information within the document and this is where the work with the docysteine is where the work with the docysteine is where the work with the docysteine team has gotten really great where we've team has gotten really great where we've team has gotten really great where we've uh worked with them to build the neatron uh worked with them to build the neatron uh worked with them to build the neatron parse model to focus specifically on parse model to focus specifically on parse model to focus specifically on table extraction which is a really table extraction which is a really table extraction which is a really complicated technique um if you think complicated technique um if you think complicated technique um if you think about it the number of permutations of about it the number of permutations of about it the number of permutations of tables are quite vast when you think tables are quite vast when you think tables are quite vast when you think about nested tables merged cells merged about nested tables merged cells merged about nested tables merged cells merged columns merged rows whatever it may be columns merged rows whatever it may be columns merged rows whatever it may be and that can get really really complex and that can get really really complex and that can get really really complex and really hairy of a and really hairy of a and really hairy of a Uh so real quick as I mentioned before Uh so real quick as I mentioned before Uh so real quick as I mentioned before right our team is responsible for uh right our team is responsible for uh right our team is responsible for uh leading a lot of the leaderboards in the leading a lot of the leaderboards in the leading a lot of the leaderboards in the retrieval space. So Vidori V1, V2, V3, retrieval space. So Vidori V1, V2, V3, retrieval space. So Vidori V1, V2, V3, MTB, MMT. Um so our team knows how to MTB, MMT. Um so our team knows how to MTB, MMT. Um so our team knows how to build world-class retrieval models given build world-class retrieval models given build world-class retrieval models given a lot of leadership uh given a lot of a lot of leadership uh given a lot of a lot of leadership uh given a lot of leaderboard winnings that we've had in leaderboard winnings that we've had in leaderboard winnings that we've had in the last year or so and then of course the last year or so and then of course the last year or so and then of course as I mentioned before we open source as I mentioned before we open source as I mentioned before we open source everything right so we share the open everything right so we share the open everything right so we share the open source model weights the techniques and source model weights the techniques and source model weights the techniques and then we release with those blueprints then we release with those blueprints then we release with those blueprints and skills that agents can use then and skills that agents can use then and skills that agents can use then afterwards um so to touch a little bit afterwards um so to touch a little bit afterwards um so to touch a little bit on the actual model that we are working on the actual model that we are working on the actual model that we are working with with docuign was the neatron parse with with docuign was the neatron parse with with docuign was the neatron parse model so when you think VLM you model so when you think VLM you model so when you think VLM you generally think a multi-billion generally think a multi-billion generally think a multi-billion parameter model it's very heavy It's parameter model it's very heavy It's parameter model it's very heavy It's high latency. Um, this is a very small high latency. Um, this is a very small high latency. Um, this is a very small tiny C radio VLM. It's about 850 900 tiny C radio VLM. It's about 850 900 tiny C radio VLM. It's about 850 900 million parameter model. Uh, designed to million parameter model. Uh, designed to million parameter model. Uh, designed to kind of be that all-in-one package sort kind of be that all-in-one package sort kind of be that all-in-one package sort of model where you deploy it and instead of model where you deploy it and instead of model where you deploy it and instead of having small, let's say, YOLO X of having small, let's say, YOLO X of having small, let's say, YOLO X models that do table extraction or page models that do table extraction or page models that do table extraction or page element extraction, whatever it may be.

  6. element extraction, whatever it may be. element extraction, whatever it may be. This is a singleshot model that you can This is a singleshot model that you can This is a singleshot model that you can feed a document in and out comes the feed a document in and out comes the feed a document in and out comes the semantic formatting layouts the text uh semantic formatting layouts the text uh semantic formatting layouts the text uh the reading order uh the the preserved the reading order uh the the preserved the reading order uh the the preserved structure of the table etc. Um this can structure of the table etc. Um this can structure of the table etc. Um this can be served via the NVIDIA NIM or via the be served via the NVIDIA NIM or via the be served via the NVIDIA NIM or via the LLM as well too. Um and so it's a tiny LLM as well too. Um and so it's a tiny LLM as well too. Um and so it's a tiny small model that you can use. It's not a small model that you can use. It's not a small model that you can use. It's not a generator. It's more of an extractor at generator. It's more of an extractor at generator. It's more of an extractor at the end of the day. the end of the day. the end of the day. Uh so real quick as well too um we Uh so real quick as well too um we Uh so real quick as well too um we always want to make sure that we're always want to make sure that we're always want to make sure that we're building towards benchmarks that matter building towards benchmarks that matter building towards benchmarks that matter most to the enterprise space. So we want most to the enterprise space. So we want most to the enterprise space. So we want to make sure that both on the paro curve to make sure that both on the paro curve to make sure that both on the paro curve of accuracy versus performance, we're of accuracy versus performance, we're of accuracy versus performance, we're make sure that we're going to be make sure that we're going to be make sure that we're going to be releasing a world-class models to the releasing a world-class models to the releasing a world-class models to the ecosystem too. So you'll see here ecosystem too. So you'll see here ecosystem too. So you'll see here generally is just a very standard generally is just a very standard generally is just a very standard benchmark of table extraction. I believe benchmark of table extraction. I believe benchmark of table extraction. I believe this one was RD table bench and we this one was RD table bench and we this one was RD table bench and we compare some popular open source models compare some popular open source models compare some popular open source models here and then we compare how our neatron here and then we compare how our neatron here and then we compare how our neatron parse model does compared to that parse model does compared to that parse model does compared to that industry and we continue to kind of industry and we continue to kind of industry and we continue to kind of strive to improve this as time goes on. strive to improve this as time goes on. strive to improve this as time goes on. >> So that I think we believe we have a >> So that I think we believe we have a >> So that I think we believe we have a demo as well. Yeah, demo as well. Yeah, demo as well. Yeah, >> just press one. Yeah. Okay, there we go. >> just press one. Yeah. Okay, there we go. >> just press one. Yeah. Okay, there we go. >> Let me show you how easy it is to turn >> Let me show you how easy it is to turn >> Let me show you how easy it is to turn any agreement into structured usable any agreement into structured usable any agreement into structured usable data with agreement manager, which is a data with agreement manager, which is a data with agreement manager, which is a central repository of every agreement an central repository of every agreement an central repository of every agreement an organization has ever signed. Let's look organization has ever signed. Let's look organization has ever signed. Let's look at this. So when we look at the at this. So when we look at the at this. So when we look at the agreement manager view here, you know, agreement manager view here, you know, agreement manager view here, you know, we have an ability to see the entire we have an ability to see the entire we have an ability to see the entire list of agreements, but also go and list of agreements, but also go and list of agreements, but also go and upload a new agreement. So I'm uploading upload a new agreement. So I'm uploading upload a new agreement. So I'm uploading a new order form into agreement manager.

  7. a new order form into agreement manager. a new order form into agreement manager. As you can see, I can select from my As you can see, I can select from my As you can see, I can select from my computer, import from other places. The computer, import from other places. The computer, import from other places. The moment I select the agreement, it starts moment I select the agreement, it starts moment I select the agreement, it starts uploading and starts processing with AI. uploading and starts processing with AI. uploading and starts processing with AI. And just like that, you can see that the And just like that, you can see that the And just like that, you can see that the jobs engine has processed it. Let's take jobs engine has processed it. Let's take jobs engine has processed it. Let's take a closer look at this agreement. So when a closer look at this agreement. So when a closer look at this agreement. So when you go into the action, you can go and you go into the action, you can go and you go into the action, you can go and browse the file. Within seconds, browse the file. Within seconds, browse the file. Within seconds, agreement manager has extracted a rich agreement manager has extracted a rich agreement manager has extracted a rich set of metadata. Everything from key set of metadata. Everything from key set of metadata. Everything from key terms to commercial details are terms to commercial details are terms to commercial details are automatically structured, highlighted, automatically structured, highlighted, automatically structured, highlighted, and immediately you can jump to that and immediately you can jump to that and immediately you can jump to that section where the details are found. section where the details are found. section where the details are found. Builtin goes deeper. This is where the Builtin goes deeper. This is where the Builtin goes deeper. This is where the Nvidia's model comes in that it's Nvidia's model comes in that it's Nvidia's model comes in that it's extracted all the structured pricing extracted all the structured pricing extracted all the structured pricing data around this agreement. It goes in data around this agreement. It goes in data around this agreement. It goes in breaks down these complex t tables into breaks down these complex t tables into breaks down these complex t tables into order details. order details. order details. And as you can see, we can break it And as you can see, we can break it And as you can see, we can break it down. We can download all of this data. down. We can download all of this data. down. We can download all of this data. This is powered by the advanced parsing This is powered by the advanced parsing This is powered by the advanced parsing leveraging Nvidia's Neotron model, leveraging Nvidia's Neotron model, leveraging Nvidia's Neotron model, turning every even dense tables into turning every even dense tables into turning every even dense tables into something that is instantly usable. And something that is instantly usable. And something that is instantly usable. And of course, you can take this data with of course, you can take this data with of course, you can take this data with you. You can see when we've downloaded you. You can see when we've downloaded you. You can see when we've downloaded into CSV how we've structured all of it into CSV how we've structured all of it into CSV how we've structured all of it for your finance team, procurement team, for your finance team, procurement team, for your finance team, procurement team, even further analysis. All of it is also even further analysis. All of it is also even further analysis. All of it is also available through API. And that's how available through API. And that's how available through API. And that's how agreement manager has transformed agreement manager has transformed agreement manager has transformed agreements into actionable insights in agreements into actionable insights in agreements into actionable insights in seconds leveraging Nvidia.

  8. seconds leveraging Nvidia. seconds leveraging Nvidia. So I think you know what you saw there So I think you know what you saw there So I think you know what you saw there from um a demo perspective we've tried from um a demo perspective we've tried from um a demo perspective we've tried to shortened it. It's like we have a to shortened it. It's like we have a to shortened it. It's like we have a whole repository what you see a list we whole repository what you see a list we whole repository what you see a list we get customers which has thousand get customers which has thousand get customers which has thousand agreements to all the way millions of agreements to all the way millions of agreements to all the way millions of agreements within but the big piece is agreements within but the big piece is agreements within but the big piece is how do we understand and get that data how do we understand and get that data how do we understand and get that data that makes it very valuable to an end that makes it very valuable to an end that makes it very valuable to an end business user right a legal person a business user right a legal person a business user right a legal person a procurement person a salesperson who's procurement person a salesperson who's procurement person a salesperson who's doing a lot of the deals or even a doing a lot of the deals or even a doing a lot of the deals or even a leader right like a business unit the leader right like a business unit the leader right like a business unit the CTO goes and asks what did we do this is CTO goes and asks what did we do this is CTO goes and asks what did we do this is how we have making each of the things a how we have making each of the things a how we have making each of the things a lot more structured so we have our own lot more structured so we have our own lot more structured so we have our own proprietary agreement data model which proprietary agreement data model which proprietary agreement data model which we are structurizing each agreement but we are structurizing each agreement but we are structurizing each agreement but also at a whole organization level and also at a whole organization level and also at a whole organization level and leveraging a lot of the NVIDIA things leveraging a lot of the NVIDIA things leveraging a lot of the NVIDIA things we've been able to do a really good job we've been able to do a really good job we've been able to do a really good job especially with all of those tables like especially with all of those tables like especially with all of those tables like pricing SLAs's and then make that pricing SLAs's and then make that pricing SLAs's and then make that available and then we also have a like available and then we also have a like available and then we also have a like robust kind of search that is um on top robust kind of search that is um on top robust kind of search that is um on top of it so when we think about what have of it so when we think about what have of it so when we think about what have we learned right when you think from a we learned right when you think from a we learned right when you think from a neotron plus docyign we one of the neotron plus docyign we one of the neotron plus docyign we one of the biggest things for us we definitely have biggest things for us we definitely have biggest things for us we definitely have done lot of different models mod for done lot of different models mod for done lot of different models mod for different purposes. So a purpose-built different purposes. So a purpose-built different purposes. So a purpose-built model for the job you're trying to do is model for the job you're trying to do is model for the job you're trying to do is a big big part of how we've been a big big part of how we've been a big big part of how we've been thinking about and that's kind of where thinking about and that's kind of where thinking about and that's kind of where we've been able to accelerate bring we've been able to accelerate bring we've been able to accelerate bring things to market much faster. The second things to market much faster. The second things to market much faster. The second big piece around like the model big piece around like the model big piece around like the model efficiency. So for you know as Sean was efficiency. So for you know as Sean was efficiency. So for you know as Sean was talking about the number of parameters talking about the number of parameters talking about the number of parameters yes context and stuff matters in the you yes context and stuff matters in the you yes context and stuff matters in the you know in a different environment for know in a different environment for know in a different environment for different things. For us the lower kind

  9. different things. For us the lower kind different things. For us the lower kind of context basically also meant lower of context basically also meant lower of context basically also meant lower latency lower cost to deliver the scale latency lower cost to deliver the scale latency lower cost to deliver the scale that we are talking about last around that we are talking about last around that we are talking about last around the faster extraction. So um we we ran the faster extraction. So um we we ran the faster extraction. So um we we ran this against a lot of the other open this against a lot of the other open this against a lot of the other open source models when we think about how source models when we think about how source models when we think about how many tables can it extract per second many tables can it extract per second many tables can it extract per second Neotron was 20x faster which helps us Neotron was 20x faster which helps us Neotron was 20x faster which helps us when we're talking about the millions when we're talking about the millions when we're talking about the millions and billions of scale that we're kind of and billions of scale that we're kind of and billions of scale that we're kind of serving for all of our customers. So a serving for all of our customers. So a serving for all of our customers. So a lot of it is like having that smaller lot of it is like having that smaller lot of it is like having that smaller purpose-built things is the way for an purpose-built things is the way for an purpose-built things is the way for an enterprise as an organization to go and enterprise as an organization to go and enterprise as an organization to go and leverage and then serve that from an leverage and then serve that from an leverage and then serve that from an enduser perspective. enduser perspective. enduser perspective. Um and then what's next? So I'll let Um and then what's next? So I'll let Um and then what's next? So I'll let Sean talk through those. Sean talk through those. Sean talk through those. >> Yeah. So working with the docuign team >> Yeah. So working with the docuign team >> Yeah. So working with the docuign team uh has been awesome so far. Uh and we're uh has been awesome so far. Uh and we're uh has been awesome so far. Uh and we're going to continue to deepen that going to continue to deepen that going to continue to deepen that partnership as well over the next few partnership as well over the next few partnership as well over the next few months. So uh with them we started with months. So uh with them we started with months. So uh with them we started with the how do I extract as much possible the how do I extract as much possible the how do I extract as much possible information from a page and now we'll information from a page and now we'll information from a page and now we'll scale to how do I now find that page to scale to how do I now find that page to scale to how do I now find that page to begin with. Um so we'll start a little begin with. Um so we'll start a little begin with. Um so we'll start a little bit with the neatron retriever uh effort bit with the neatron retriever uh effort bit with the neatron retriever uh effort and then of course we'll talk a little and then of course we'll talk a little and then of course we'll talk a little bit about the NVIDIA agent toolkit with bit about the NVIDIA agent toolkit with bit about the NVIDIA agent toolkit with them over the next few months um and them over the next few months um and them over the next few months um and then actually start scaling into into then actually start scaling into into then actually start scaling into into more production scale agents then.

  10. more production scale agents then. more production scale agents then. >> Perfect. I think that's what we had. We >> Perfect. I think that's what we had. We >> Perfect. I think that's what we had. We have time for a couple questions. I went have time for a couple questions. I went have time for a couple questions. I went in the room. So just to recap for everybody if you So just to recap for everybody if you didn't hear it was the question is right didn't hear it was the question is right didn't hear it was the question is right like OCR is always a thorn in the whole like OCR is always a thorn in the whole like OCR is always a thorn in the whole process so are we thinking about letting process so are we thinking about letting process so are we thinking about letting the go of that and starting from agentic the go of that and starting from agentic the go of that and starting from agentic from the get-go I can talk from my from the get-go I can talk from my from the get-go I can talk from my perspective but so I think for us right perspective but so I think for us right perspective but so I think for us right like there are different use cases at like there are different use cases at like there are different use cases at different points in time many times if different points in time many times if different points in time many times if you are reactive you have a question and you are reactive you have a question and you are reactive you have a question and you're coming some of that can can work you're coming some of that can can work you're coming some of that can can work dynamically at a smaller scale. The dynamically at a smaller scale. The dynamically at a smaller scale. The question is the latency when I am question is the latency when I am question is the latency when I am quering at that scale of thousands I do quering at that scale of thousands I do quering at that scale of thousands I do need to have pre-processed have need to have pre-processed have need to have pre-processed have identified. So that's one. I think the identified. So that's one. I think the identified. So that's one. I think the second big part of the use case for us a second big part of the use case for us a second big part of the use case for us a lot of times businesses want to use this lot of times businesses want to use this lot of times businesses want to use this data to do a lot of downstream work. So data to do a lot of downstream work. So data to do a lot of downstream work. So an example is a procurement team. This an example is a procurement team. This an example is a procurement team. This is my pricing table. I want to put it is my pricing table. I want to put it is my pricing table. I want to put it into Koopa to make sure when I'm paying into Koopa to make sure when I'm paying into Koopa to make sure when I'm paying that works at that time there like you that works at that time there like you that works at that time there like you know the agent is kind of helping but I know the agent is kind of helping but I know the agent is kind of helping but I can't do that on a one document by can't do that on a one document by can't do that on a one document by document. That said, there is ways that document. That said, there is ways that document. That said, there is ways that we are compressing. That's kind of why we are compressing. That's kind of why we are compressing. That's kind of why Neotron worked for us is like how do you

  11. Neotron worked for us is like how do you Neotron worked for us is like how do you do it from a layout understanding just do it from a layout understanding just do it from a layout understanding just for that purpose but I would like let for that purpose but I would like let for that purpose but I would like let you add. you add. you add. >> Yeah, I think it depends on the use case >> Yeah, I think it depends on the use case >> Yeah, I think it depends on the use case a little bit. Um I think for this a little bit. Um I think for this a little bit. Um I think for this specific instance, right, you have specific instance, right, you have specific instance, right, you have pabytes of documents that you want to be pabytes of documents that you want to be pabytes of documents that you want to be queriable at some point, right? So you queriable at some point, right? So you queriable at some point, right? So you are heavy on the compute at the upfront are heavy on the compute at the upfront are heavy on the compute at the upfront side with all the OCR so you don't have side with all the OCR so you don't have side with all the OCR so you don't have to worry about it later on, right? I to worry about it later on, right? I to worry about it later on, right? I think there's some instances where think there's some instances where think there's some instances where people may upload a contract to begin people may upload a contract to begin people may upload a contract to begin with for Q&A and that's a very high with for Q&A and that's a very high with for Q&A and that's a very high that's a very low latency use case, that's a very low latency use case, that's a very low latency use case, right? So you have a high throughput right? So you have a high throughput right? So you have a high throughput versus low latency use case and in that versus low latency use case and in that versus low latency use case and in that scenario your different batch sizes, scenario your different batch sizes, scenario your different batch sizes, your concurrencies, your different your concurrencies, your different your concurrencies, your different techniques on how you process the techniques on how you process the techniques on how you process the document will be different and where you document will be different and where you document will be different and where you spend that compute in that cycle will be spend that compute in that cycle will be spend that compute in that cycle will be changing between the different use changing between the different use changing between the different use cases. cases. cases. >> Okay, one more there. We we do a lot of like more of what I We we do a lot of like more of what I call hybrid approach and a purpose-built call hybrid approach and a purpose-built call hybrid approach and a purpose-built for like the needs and the use cases. So for like the needs and the use cases. So for like the needs and the use cases. So from a table piece it does kind of you from a table piece it does kind of you from a table piece it does kind of you know do the whole layout along with know do the whole layout along with know do the whole layout along with extracting we still do OCR from a lot of extracting we still do OCR from a lot of extracting we still do OCR from a lot of other fields and metadata and a closet other fields and metadata and a closet other fields and metadata and a closet like all of the text kind of thing. So like all of the text kind of thing. So like all of the text kind of thing. So that I think we had a architecture where that I think we had a architecture where that I think we had a architecture where we have our pipeline going through two we have our pipeline going through two we have our pipeline going through two different routes for that. Um as a different routes for that. Um as a different routes for that. Um as a follow we have a blog out there how are follow we have a blog out there how are follow we have a blog out there how are we really solving this at scale across we really solving this at scale across we really solving this at scale across and if you look at that there's a lot of and if you look at that there's a lot of and if you look at that there's a lot of different peacemail modules and stuff different peacemail modules and stuff different peacemail modules and stuff together.

  12. together. together. >> Yeah one moreization >> so we this model is currently on FP16 >> so we this model is currently on FP16 but there are paths towards going on to but there are paths towards going on to but there are paths towards going on to FP8 and VFP4 in the next few months as FP8 and VFP4 in the next few months as FP8 and VFP4 in the next few months as well too. What about center? Are you well too. What about center? Are you well too. What about center? Are you using the 16 or using the 16 or using the 16 or >> we do use that and then we're also kind >> we do use that and then we're also kind >> we do use that and then we're also kind of using some of the older ones and of using some of the older ones and of using some of the older ones and that's the journey as a partnership is that's the journey as a partnership is that's the journey as a partnership is to kind of go tweak as you get more of to kind of go tweak as you get more of to kind of go tweak as you get more of the customers. the customers. the customers. >> Yeah. So for this there are many >> Yeah. So for this there are many >> Yeah. So for this there are many techniques on how to improve the techniques on how to improve the techniques on how to improve the performance side right so quantization performance side right so quantization performance side right so quantization right so we're trying to move everyone right so we're trying to move everyone right so we're trying to move everyone to blackwell right so that's why NVFP4 to blackwell right so that's why NVFP4 to blackwell right so that's why NVFP4 is the big thing now um as well as is the big thing now um as well as is the big thing now um as well as multi-token generation for this it's a multi-token generation for this it's a multi-token generation for this it's a VLM architecture right so your encoder VLM architecture right so your encoder VLM architecture right so your encoder decoder techniques can definitely be decoder techniques can definitely be decoder techniques can definitely be further optimized so not right now this further optimized so not right now this further optimized so not right now this model just generates one token at a time model just generates one token at a time model just generates one token at a time you can do multi token generation of you can do multi token generation of you can do multi token generation of course too so there's plenty of course too so there's plenty of course too so there's plenty of performance things right now we're performance things right now we're performance things right now we're focusing on the accuracy side like are focusing on the accuracy side like are focusing on the accuracy side like are we adding value to the system and then we adding value to the system and then we adding value to the system and then from there we'll then push out that paro from there we'll then push out that paro from there we'll then push out that paro curve on the performance curve on the performance curve on the performance Are you guys using >> I believe they just deploy via directly >> I believe they just deploy via directly recall recall recall when you do that.

  13. >> Oh, this is just an extraction. This is >> Oh, this is just an extraction. This is now a retrieval. now a retrieval. now a retrieval. >> Yeah. >> Yeah. >> Yeah. >> Yeah. Yeah, you're asking question. >> Yeah. Yeah, you're asking question. >> Yeah. Yeah, you're asking question. >> It's coming from that agreement data >> It's coming from that agreement data >> It's coming from that agreement data that we've kind of extracted and stored. that we've kind of extracted and stored. that we've kind of extracted and stored. >> So yeah, maybe I can chat with you >> So yeah, maybe I can chat with you >> So yeah, maybe I can chat with you offline and how like we our architecture offline and how like we our architecture offline and how like we our architecture kind of works fully as well. kind of works fully as well. kind of works fully as well. No, we're almost coming up on time No, we're almost coming up on time No, we're almost coming up on time there. Um, but I think that's kind of there. Um, but I think that's kind of there. Um, but I think that's kind of all we have. I'm happy to hang around uh all we have. I'm happy to hang around uh all we have. I'm happy to hang around uh in the back with more questions. Um, and in the back with more questions. Um, and in the back with more questions. Um, and good luck with a lot of your uh good luck with a lot of your uh good luck with a lot of your uh challenges with AI. So, thank you. challenges with AI. So, thank you. challenges with AI. So, thank you. >> Thank you. >> Thank you. >> Thank you. >> [applause]

Summary

This session discusses the massive problem of unstructured agreement data in enterprises, referencing DocuSign and NVIDIA's partnership. The key takeaway is how their AI-first Intelligent Agreement Management platform makes this critical data readable, usable, and queryable at scale, unlocking significant untapped value.

View original episode ↗