← Back
AI Engineer August 17, 2026 1h 3m

Context Engineering in 2026 — Louis-François Bouchard, Omar Solano & Samridhi Vaid, Towards AI

Read full transcript 47 segments
  1. All right. Good afternoon everyone. All right. Good afternoon everyone. Thank you for joining and not watching Thank you for joining and not watching Thank you for joining and not watching the game. Uh I hope it will be a bit the game. Uh I hope it will be a bit the game. Uh I hope it will be a bit more interesting or at least you will more interesting or at least you will more interesting or at least you will learn something compared to to uh learn something compared to to uh learn something compared to to uh hopefully Germany winning or some uh hopefully Germany winning or some uh hopefully Germany winning or some uh anyways. Yeah. All right. Is it fine? Okay. All right. Is it fine? Okay. All right. So I'm here to talk about we All right. So I'm here to talk about we All right. So I'm here to talk about we are here to talk about context are here to talk about context are here to talk about context engineering in 2026. And more engineering in 2026. And more engineering in 2026. And more specifically, we are here because we've specifically, we are here because we've specifically, we are here because we've all lived that that situation where you all lived that that situation where you all lived that that situation where you try to do things with an agent and try to do things with an agent and try to do things with an agent and ultimately it does just exactly the ultimately it does just exactly the ultimately it does just exactly the thing that you don't want it to do. And thing that you don't want it to do. And thing that you don't want it to do. And it in my case it usually ends up like it in my case it usually ends up like it in my case it usually ends up like this where I'm super mad and I just type this where I'm super mad and I just type this where I'm super mad and I just type back hoping it it learns. back hoping it it learns. back hoping it it learns. And uh usually the problem here is not And uh usually the problem here is not And uh usually the problem here is not that the the model just got dumber and that the the model just got dumber and that the the model just got dumber and you need to switch to cloud or to codeex you need to switch to cloud or to codeex you need to switch to cloud or to codeex or or whatever the the harness that or or whatever the the harness that or or whatever the the harness that you're using, but it's more that the the you're using, but it's more that the the you're using, but it's more that the the context is filling up and it's getting context is filling up and it's getting context is filling up and it's getting worse and worse. The results are getting worse and worse. The results are getting worse and worse. The results are getting worse because of it. In our case, this worse because of it. In our case, this worse because of it. In our case, this is important because we build courses is important because we build courses is important because we build courses and trainings for AI engineers and trainings for AI engineers and trainings for AI engineers specifically. And one of the features specifically. And one of the features specifically. And one of the features that we provide is an AI tutor to help that we provide is an AI tutor to help that we provide is an AI tutor to help answer questions based on our lessons.

  2. answer questions based on our lessons. answer questions based on our lessons. And if the interaction is just like the And if the interaction is just like the And if the interaction is just like the one before and they are super mad at us, one before and they are super mad at us, one before and they are super mad at us, they might just ask for refund and it they might just ask for refund and it they might just ask for refund and it could end up like this. So that's not could end up like this. So that's not could end up like this. So that's not what we want. what we want. what we want. And so what we did for this workshop and And so what we did for this workshop and And so what we did for this workshop and just for the AI tutor in general is to just for the AI tutor in general is to just for the AI tutor in general is to run many different experiments in order run many different experiments in order run many different experiments in order to figure out how in our case we can fix to figure out how in our case we can fix to figure out how in our case we can fix context rut or at least improve the i context rut or at least improve the i context rut or at least improve the i tutor as much as possible and reduce the tutor as much as possible and reduce the tutor as much as possible and reduce the cost as well of running the tutor. cost as well of running the tutor. cost as well of running the tutor. Uh the QR code here is a link to a Uh the QR code here is a link to a Uh the QR code here is a link to a hugging face space where you have all hugging face space where you have all hugging face space where you have all these experiments that you can see and these experiments that you can see and these experiments that you can see and also the AI tutor is open source. who also the AI tutor is open source. who also the AI tutor is open source. who will share another code for the repo but will share another code for the repo but will share another code for the repo but it's also linked on the hugging face. So it's also linked on the hugging face. So it's also linked on the hugging face. So everything is open source you can access everything is open source you can access everything is open source you can access everything uh and even see the everything uh and even see the everything uh and even see the experiments online and use the AI experiments online and use the AI experiments online and use the AI tutorial online as well. In the next 80 tutorial online as well. In the next 80 tutorial online as well. In the next 80 minutes, we will I will start talking minutes, we will I will start talking minutes, we will I will start talking about compaction, memory retrieval and about compaction, memory retrieval and about compaction, memory retrieval and um everything that you can do in 2026 um everything that you can do in 2026 um everything that you can do in 2026 that usually works and then my that usually works and then my that usually works and then my colleagues will jump in with our the colleagues will jump in with our the colleagues will jump in with our the architecture of our AI tutor, our architecture of our AI tutor, our architecture of our AI tutor, our decisions, what we built and the decisions, what we built and the decisions, what we built and the evaluations that we built, the how we evaluations that we built, the how we evaluations that we built, the how we built them and what we decided to built them and what we decided to built them and what we decided to evaluate and then the results and what evaluate and then the results and what evaluate and then the results and what we took out of this. So of of course we took out of this. So of of course we took out of this. So of of course it's applied to our use case and our AI it's applied to our use case and our AI it's applied to our use case and our AI tutor but hopefully you can get away tutor but hopefully you can get away tutor but hopefully you can get away some interesting insights at least from some interesting insights at least from some interesting insights at least from from this and some best practices that from this and some best practices that from this and some best practices that we learned throughout more specifically

  3. we learned throughout more specifically we learned throughout more specifically we is towards AI. Um I founded the we is towards AI. Um I founded the we is towards AI. Um I founded the company with my partners in in a few company with my partners in in a few company with my partners in in a few years ago and we've always been focused years ago and we've always been focused years ago and we've always been focused around education. Obviously back in the around education. Obviously back in the around education. Obviously back in the day it was more about computer vision day it was more about computer vision day it was more about computer vision and and more basic machine learning. Now and and more basic machine learning. Now and and more basic machine learning. Now it's towards AI engineering and agents, it's towards AI engineering and agents, it's towards AI engineering and agents, anything that works for the industry. anything that works for the industry. anything that works for the industry. And I'm joined by my colleagues helping And I'm joined by my colleagues helping And I'm joined by my colleagues helping me develop this AI tutor and our me develop this AI tutor and our me develop this AI tutor and our courses, Omar and Samidi that we'll jump courses, Omar and Samidi that we'll jump courses, Omar and Samidi that we'll jump in later on. And here more specifically in later on. And here more specifically in later on. And here more specifically towards AI is quite large. But what one towards AI is quite large. But what one towards AI is quite large. But what one one of the things that we do is our one of the things that we do is our one of the things that we do is our academy. So the towards AI academy where academy. So the towards AI academy where academy. So the towards AI academy where we build courses technical courses for we build courses technical courses for we build courses technical courses for AI engineers to upskill towards AI AI engineers to upskill towards AI AI engineers to upskill towards AI engineering and as I said we provide an engineering and as I said we provide an engineering and as I said we provide an AI tutor for the students. AI tutor for the students. AI tutor for the students. So the the AI tutor specifically uh will So the the AI tutor specifically uh will So the the AI tutor specifically uh will will be like our baseline for all our will be like our baseline for all our will be like our baseline for all our experiments. We will use that to test experiments. We will use that to test experiments. We will use that to test all the different features based on real all the different features based on real all the different features based on real user interactions and to have the best user interactions and to have the best user interactions and to have the best results possible. We had five results possible. We had five results possible. We had five requirements we wanted to ensure that requirements we wanted to ensure that requirements we wanted to ensure that the chatbot follows. The first one the chatbot follows. The first one the chatbot follows. The first one obviously we want the answers of the obviously we want the answers of the obviously we want the answers of the chatbot to be grounded in our content chatbot to be grounded in our content chatbot to be grounded in our content not just its own knowledge.

  4. not just its own knowledge. not just its own knowledge. We need the the tutor to be based in the We need the the tutor to be based in the We need the the tutor to be based in the current students and current lesson to current students and current lesson to current students and current lesson to not be a because we have multiple not be a because we have multiple not be a because we have multiple courses. So we just need to ensure that courses. So we just need to ensure that courses. So we just need to ensure that it answers from this course content. it answers from this course content. it answers from this course content. Then it needs to hold long help sessions Then it needs to hold long help sessions Then it needs to hold long help sessions in case the student is debugging or or in case the student is debugging or or in case the student is debugging or or just iterating a lot with the the tutor just iterating a lot with the the tutor just iterating a lot with the the tutor and uh obviously handled code because and uh obviously handled code because and uh obviously handled code because it's for AI engineers. So we just code a it's for AI engineers. So we just code a it's for AI engineers. So we just code a lot and it needs to have somewhat of a lot and it needs to have somewhat of a lot and it needs to have somewhat of a low latency to not be frustrating to low latency to not be frustrating to low latency to not be frustrating to use. And all of this is related to use. And all of this is related to use. And all of this is related to context engineering. And here we'll be context engineering. And here we'll be context engineering. And here we'll be talking about uh what is context talking about uh what is context talking about uh what is context engineering in 2026 or at least what we engineering in 2026 or at least what we engineering in 2026 or at least what we figured out from this from these figured out from this from these figured out from this from these experiments. experiments. experiments. And since everything here is in the And since everything here is in the And since everything here is in the context in the in the context of models, context in the in the context of models, context in the in the context of models, it creates two problems. First, the it creates two problems. First, the it creates two problems. First, the context window is finite. Uh everything context window is finite. Uh everything context window is finite. Uh everything the model will see from instructions to the model will see from instructions to the model will see from instructions to to lessons to code will lie in the same to lessons to code will lie in the same to lessons to code will lie in the same space that is limited. And the more you space that is limited. And the more you space that is limited. And the more you pile things in this space, the worse the pile things in this space, the worse the pile things in this space, the worse the results will be and the more expensive results will be and the more expensive results will be and the more expensive it will be because you pay for more it will be because you pay for more it will be because you pay for more tokens. So that's one of the the main tokens. So that's one of the the main tokens. So that's one of the the main problem we're trying to fix. And the problem we're trying to fix. And the problem we're trying to fix. And the second problem is that the model is second problem is that the model is second problem is that the model is stateless. So when a student reopens the stateless. So when a student reopens the stateless. So when a student reopens the AI tutorially if you don't build AI tutorially if you don't build AI tutorially if you don't build anything around it, the model will have anything around it, the model will have anything around it, the model will have no idea what's going on. It's just no idea what's going on. It's just no idea what's going on. It's just starting from zero. So this creates two

  5. starting from zero. So this creates two starting from zero. So this creates two things we have to work on. the context things we have to work on. the context things we have to work on. the context management which means within one management which means within one management which means within one session and the memory aspect of these session and the memory aspect of these session and the memory aspect of these models which means across sessions in models which means across sessions in models which means across sessions in these experiments and in this workshop these experiments and in this workshop these experiments and in this workshop we focus on the first part the context we focus on the first part the context we focus on the first part the context management because you cannot have management because you cannot have management because you cannot have multiple sessions is if one session is multiple sessions is if one session is multiple sessions is if one session is is shitty. So we try to really optimize is shitty. So we try to really optimize is shitty. So we try to really optimize for this and maybe in a future workshop for this and maybe in a future workshop for this and maybe in a future workshop we will do one about memory hopefully. we will do one about memory hopefully. we will do one about memory hopefully. So for context management what does the So for context management what does the So for context management what does the the tutor sees in our case? It sees uh the tutor sees in our case? It sees uh the tutor sees in our case? It sees uh many things from the system prompt to many things from the system prompt to many things from the system prompt to being for being a tutor and the context being for being a tutor and the context being for being a tutor and the context of our courses and stuff to tool of our courses and stuff to tool of our courses and stuff to tool definitions uh so which tools it can use definitions uh so which tools it can use definitions uh so which tools it can use how to use them when to use them. The how to use them when to use them. The how to use them when to use them. The chat history if it's an ongoing chat history if it's an ongoing chat history if it's an ongoing discussion any old tool outputs that was discussion any old tool outputs that was discussion any old tool outputs that was called course chunks that we retrieved called course chunks that we retrieved called course chunks that we retrieved to answer the students and finally the to answer the students and finally the to answer the students and finally the user questions. So it's just not the user questions. So it's just not the user questions. So it's just not the user question that we send obviously and user question that we send obviously and user question that we send obviously and typically it's the smallest part but it typically it's the smallest part but it typically it's the smallest part but it can also have contain a lot of code to can also have contain a lot of code to can also have contain a lot of code to debug and to help or error logs. So it debug and to help or error logs. So it debug and to help or error logs. So it can be large as well and all of this can be large as well and all of this can be large as well and all of this together just ends up costing more and together just ends up costing more and together just ends up costing more and more to us. So we really want to more to us. So we really want to more to us. So we really want to optimize this this context management optimize this this context management optimize this this context management aspect. And what we've seen just quickly aspect. And what we've seen just quickly aspect. And what we've seen just quickly is that the the main bottleneck or the is that the the main bottleneck or the is that the the main bottleneck or the main problem scaling the context is the main problem scaling the context is the main problem scaling the context is the old tool outputs which contains like any old tool outputs which contains like any old tool outputs which contains like any old chunks retrieved all the tool calls

  6. old chunks retrieved all the tool calls old chunks retrieved all the tool calls and tool results pairs uh from from all and tool results pairs uh from from all and tool results pairs uh from from all the tools called over the the many the tools called over the the many the tools called over the the many sessions that it could have the many sessions that it could have the many sessions that it could have the many turns that that one session could have turns that that one session could have turns that that one session could have have and any uh file or searches that it have and any uh file or searches that it have and any uh file or searches that it did to in its own memory. did to in its own memory. did to in its own memory. And the main problem is not is not that And the main problem is not is not that And the main problem is not is not that it's more costly. It's also that the it's more costly. It's also that the it's more costly. It's also that the quality degrades as we know with a quality degrades as we know with a quality degrades as we know with a program called context route because of program called context route because of program called context route because of the way how large language models are the way how large language models are the way how large language models are trained to handle longer context with we trained to handle longer context with we trained to handle longer context with we just inject facts into large compass just inject facts into large compass just inject facts into large compass which doesn't tell them to manage the which doesn't tell them to manage the which doesn't tell them to manage the whole the whole context together or whole the whole context together or whole the whole context together or understand the global context and understand the global context and understand the global context and uh so so one of the reason is to help uh so so one of the reason is to help uh so so one of the reason is to help with quality. We want to reduce the with quality. We want to reduce the with quality. We want to reduce the context as much as possible. And the context as much as possible. And the context as much as possible. And the others are that if you reququery the others are that if you reququery the others are that if you reququery the model in a in a discussion, you resend model in a in a discussion, you resend model in a in a discussion, you resend every previous tokens. So you pay for every previous tokens. So you pay for every previous tokens. So you pay for them as well again, which is far from them as well again, which is far from them as well again, which is far from ideal. And you increase the the latency, ideal. And you increase the the latency, ideal. And you increase the the latency, so the time to first token or TTFT uh so the time to first token or TTFT uh so the time to first token or TTFT uh which makes which creates a very bad which makes which creates a very bad which makes which creates a very bad user experience.

  7. user experience. user experience. So you may want to in our case manage So you may want to in our case manage So you may want to in our case manage our context for speed and for spending our context for speed and for spending our context for speed and for spending not just because the quality drops as we not just because the quality drops as we not just because the quality drops as we will see in our experiments and and in will see in our experiments and and in will see in our experiments and and in the results in the near future. the results in the near future. the results in the near future. So how we do that to manage uh spending So how we do that to manage uh spending So how we do that to manage uh spending and and speed? How do we manage our and and speed? How do we manage our and and speed? How do we manage our context? We use compaction. context? We use compaction. context? We use compaction. And the idea is just very simple. Uh the And the idea is just very simple. Uh the And the idea is just very simple. Uh the idea of compaction is very simple. is idea of compaction is very simple. is idea of compaction is very simple. is just to try to have the smallest context just to try to have the smallest context just to try to have the smallest context possible that contains the information possible that contains the information possible that contains the information to be able to answer the question and to be able to answer the question and to be able to answer the question and you drop or save the rest somewhere. you drop or save the rest somewhere. you drop or save the rest somewhere. And to do compaction, even a good And to do compaction, even a good And to do compaction, even a good compaction, you don't necessarily have compaction, you don't necessarily have compaction, you don't necessarily have to have large language models. You can to have large language models. You can to have large language models. You can start quite cheap with trivial tools start quite cheap with trivial tools start quite cheap with trivial tools like uh if you have if you use tools in like uh if you have if you use tools in like uh if you have if you use tools in your system um like a web search or just your system um like a web search or just your system um like a web search or just executing code you can automatically executing code you can automatically executing code you can automatically truncate outliers. So if you have like truncate outliers. So if you have like truncate outliers. So if you have like once tool that that produce 300 lines once tool that that produce 300 lines once tool that that produce 300 lines instead of 10 usually you can just instead of 10 usually you can just instead of 10 usually you can just truncate almost everything except the truncate almost everything except the truncate almost everything except the the head and tail the the beginning and the head and tail the the beginning and the head and tail the the beginning and the end and just re write that it's the end and just re write that it's the end and just re write that it's truncated so that the model in the truncated so that the model in the truncated so that the model in the future can recall the tool if it's if it future can recall the tool if it's if it future can recall the tool if it's if it feels it lacks context.

  8. feels it lacks context. feels it lacks context. You can use the the simplest approach You can use the the simplest approach You can use the the simplest approach that works the best to just use a that works the best to just use a that works the best to just use a sliding window or just trim. Basically, sliding window or just trim. Basically, sliding window or just trim. Basically, use the last n number of of turns that use the last n number of of turns that use the last n number of of turns that the user sent which uh you need to the user sent which uh you need to the user sent which uh you need to determine by based on your your own determine by based on your your own determine by based on your your own system and your users and you can clear system and your users and you can clear system and your users and you can clear uh for for some specific tools depending uh for for some specific tools depending uh for for some specific tools depending on those that you implement you can just on those that you implement you can just on those that you implement you can just always clear most of the outputs. always clear most of the outputs. always clear most of the outputs. So that's for when that's not even using So that's for when that's not even using So that's for when that's not even using language models. And then you can use language models. And then you can use language models. And then you can use language models to basically spend language models to basically spend language models to basically spend tokens to save even more tokens. And in tokens to save even more tokens. And in tokens to save even more tokens. And in this case, you don't even have to to use this case, you don't even have to to use this case, you don't even have to to use large language models. You can even use large language models. You can even use large language models. You can even use smaller ones or even low super small smaller ones or even low super small smaller ones or even low super small local ones that that runs in like one local ones that that runs in like one local ones that that runs in like one MacBook MacBook MacBook to do a few techniques. There are many to do a few techniques. There are many to do a few techniques. There are many techniques that exist for for techniques that exist for for techniques that exist for for compacting. compacting. compacting. the those that had the most impact in the those that had the most impact in the those that had the most impact in our experiments where selective our experiments where selective our experiments where selective retention where the language models will retention where the language models will retention where the language models will just decide based on where the just decide based on where the just decide based on where the discussion go is going what to keep what discussion go is going what to keep what discussion go is going what to keep what to discard.

  9. to discard. to discard. Then the simplest one here Then the simplest one here Then the simplest one here summarization. So just summarize summarization. So just summarize summarization. So just summarize continuously summarize the some previous continuously summarize the some previous continuously summarize the some previous terms depending on your on on your own terms depending on your on on your own terms depending on your on on your own application and in the end you can do application and in the end you can do application and in the end you can do here what cloud code does. So when it here what cloud code does. So when it here what cloud code does. So when it reaches the limit you just sum produce a reaches the limit you just sum produce a reaches the limit you just sum produce a summary of everything and reset summary of everything and reset summary of everything and reset completely with that summary. And there completely with that summary. And there completely with that summary. And there are many more techniques I highlighted are many more techniques I highlighted are many more techniques I highlighted here those that work the best. We will here those that work the best. We will here those that work the best. We will discuss them later on in this workshop discuss them later on in this workshop discuss them later on in this workshop and the actual results and experiment and the actual results and experiment and the actual results and experiment setup but those are the most successful setup but those are the most successful setup but those are the most successful techniques. And I want to highlight also techniques. And I want to highlight also techniques. And I want to highlight also delta summarization that cloud code uses delta summarization that cloud code uses delta summarization that cloud code uses that is very useful when you spawn sub that is very useful when you spawn sub that is very useful when you spawn sub aents when you use sub aents. It just aents when you use sub aents. It just aents when you use sub aents. It just means to keep a summary and updating the means to keep a summary and updating the means to keep a summary and updating the summary based on the new summary that summary based on the new summary that summary based on the new summary that you produce over time and then the the you produce over time and then the the you produce over time and then the the sub agent will just give that to the sub agent will just give that to the sub agent will just give that to the main agent. But in our case, we don't main agent. But in our case, we don't main agent. But in our case, we don't use sub agents because the tutor works use sub agents because the tutor works use sub agents because the tutor works really well with just one main. So we really well with just one main. So we really well with just one main. So we don't need to add complexity for that.

  10. don't need to add complexity for that. don't need to add complexity for that. And lastly, after spending token to to And lastly, after spending token to to And lastly, after spending token to to save more tokens, you can also offload save more tokens, you can also offload save more tokens, you can also offload things. Right now it's obviously memory things. Right now it's obviously memory things. Right now it's obviously memory and skills are super popular. So you can and skills are super popular. So you can and skills are super popular. So you can of course offload to your memory. So of course offload to your memory. So of course offload to your memory. So just saving text in your documentations just saving text in your documentations just saving text in your documentations and you can use what's been there for and you can use what's been there for and you can use what's been there for years now. Uh retrieve augmented years now. Uh retrieve augmented years now. Uh retrieve augmented generation which is very powerful. And generation which is very powerful. And generation which is very powerful. And as a side note we also compared with as a side note we also compared with as a side note we also compared with graph rag. So everything even if I not I graph rag. So everything even if I not I graph rag. So everything even if I not I don't mention them we compared them and don't mention them we compared them and don't mention them we compared them and I highlight just the the best results I highlight just the the best results I highlight just the the best results here. So we we compared graph rag with here. So we we compared graph rag with here. So we we compared graph rag with with rag here and it's in our case it with rag here and it's in our case it with rag here and it's in our case it just ended up being way costlier to set just ended up being way costlier to set just ended up being way costlier to set up and just tie on the results because up and just tie on the results because up and just tie on the results because it basically was 100% based on our real it basically was 100% based on our real it basically was 100% based on our real user evaluations. So we don't need to user evaluations. So we don't need to user evaluations. So we don't need to use graph rag but it's it depends on use graph rag but it's it depends on use graph rag but it's it depends on your own case. If you have a very large your own case. If you have a very large your own case. If you have a very large data set with inter data set with inter data set with inter with relations and interconnected with relations and interconnected with relations and interconnected topics and things, it might be worth topics and things, it might be worth topics and things, it might be worth implementing. So you definitely want to implementing. So you definitely want to implementing. So you definitely want to still test it.

  11. still test it. still test it. And uh speaking of memory and of up And uh speaking of memory and of up And uh speaking of memory and of up offloading to files, this is uh as I offloading to files, this is uh as I offloading to files, this is uh as I said, basically just saving them locally said, basically just saving them locally said, basically just saving them locally or or in a server for for you to use, or or in a server for for you to use, or or in a server for for you to use, which means it's fully reversible which means it's fully reversible which means it's fully reversible because you don't lose anything. you because you don't lose anything. you because you don't lose anything. you don't lose any ongoing discussion. You don't lose any ongoing discussion. You don't lose any ongoing discussion. You just save them ready to be referred to just save them ready to be referred to just save them ready to be referred to in the future if the same student comes in the future if the same student comes in the future if the same student comes back and asks related questions. And back and asks related questions. And back and asks related questions. And it's basically the the Tpetes idea of it's basically the the Tpetes idea of it's basically the the Tpetes idea of the LLM wiki uh which if you link with the LLM wiki uh which if you link with the LLM wiki uh which if you link with some sort of uh chunks so a version of some sort of uh chunks so a version of some sort of uh chunks so a version of rag it's pretty powerful and it makes rag it's pretty powerful and it makes rag it's pretty powerful and it makes your system become quite cheap and your system become quite cheap and your system become quite cheap and durable and it's easy to inspect from durable and it's easy to inspect from durable and it's easy to inspect from both humans and agents. So it's really both humans and agents. So it's really both humans and agents. So it's really interesting. More specifically, it looks interesting. More specifically, it looks interesting. More specifically, it looks like this in our case. So we have just like this in our case. So we have just like this in our case. So we have just chunks. We save everything into chunks chunks. We save everything into chunks chunks. We save everything into chunks and we cross-link the chunks with and we cross-link the chunks with and we cross-link the chunks with pointers and then these chunks are pointers and then these chunks are pointers and then these chunks are linked to raw data files from where they linked to raw data files from where they linked to raw data files from where they come from. Then we have one index that come from. Then we have one index that come from. Then we have one index that will just map all these chunks. So just will just map all these chunks. So just will just map all these chunks. So just a link to all the chunks and some a link to all the chunks and some a link to all the chunks and some context of what it is about and the context of what it is about and the context of what it is about and the agent will just see this index. So it agent will just see this index. So it agent will just see this index. So it sees it's like I think it was 450 sees it's like I think it was 450 sees it's like I think it was 450 tokens. So it's very small and it just tokens. So it's very small and it just tokens. So it's very small and it just sees that index sees that index sees that index and then uh based on the user question and then uh based on the user question and then uh based on the user question if it seems to be related to some user if it seems to be related to some user if it seems to be related to some user specific question that it may exists in specific question that it may exists in specific question that it may exists in the memory it will scan the index it

  12. the memory it will scan the index it the memory it will scan the index it will go back to the chunk if it's enough will go back to the chunk if it's enough will go back to the chunk if it's enough it will answer based on the junk the it will answer based on the junk the it will answer based on the junk the chunk retrieved if it's not enough it chunk retrieved if it's not enough it chunk retrieved if it's not enough it can even go back to the raw data to have can even go back to the raw data to have can even go back to the raw data to have even more information. So basically it's even more information. So basically it's even more information. So basically it's just the best way to pull context just the best way to pull context just the best way to pull context accordingly to the task complexity. So accordingly to the task complexity. So accordingly to the task complexity. So if the task is complex you will pull if the task is complex you will pull if the task is complex you will pull more and if it's simple you'll p pull more and if it's simple you'll p pull more and if it's simple you'll p pull less. less. less. And a parenthesis on this what we've And a parenthesis on this what we've And a parenthesis on this what we've seen working with our clients and just seen working with our clients and just seen working with our clients and just building this in general is that right building this in general is that right building this in general is that right now everyone is is converging towards now everyone is is converging towards now everyone is is converging towards having more and smaller skills. So you having more and smaller skills. So you having more and smaller skills. So you have it's it's way better to build small have it's it's way better to build small have it's it's way better to build small very precise skills that refer to each very precise skills that refer to each very precise skills that refer to each other's skills to uh to save on context other's skills to uh to save on context other's skills to uh to save on context and just load skills one by one and even and just load skills one by one and even and just load skills one by one and even be able to spawn a sub agent with one be able to spawn a sub agent with one be able to spawn a sub agent with one dedicated skill context instead of just dedicated skill context instead of just dedicated skill context instead of just basically to save context. It's a well basically to save context. It's a well basically to save context. It's a well the idea of progressive disclosure. So the idea of progressive disclosure. So the idea of progressive disclosure. So you just load what you need right now. you just load what you need right now. you just load what you need right now. Okay, just to go back now on on Okay, just to go back now on on Okay, just to go back now on on compaction. This is what the the talk is compaction. This is what the the talk is compaction. This is what the the talk is about because memory we have made some about because memory we have made some about because memory we have made some experiments but we couldn't really fit experiments but we couldn't really fit experiments but we couldn't really fit here. It was a bit too much. And uh when here. It was a bit too much. And uh when here. It was a bit too much. And uh when talking about compaction there's an talking about compaction there's an talking about compaction there's an important problem or solution that uh important problem or solution that uh important problem or solution that uh that appeared recently uh well not that appeared recently uh well not that appeared recently uh well not recently but was way more popularized is recently but was way more popularized is recently but was way more popularized is the uh the the main problem is that the uh the the main problem is that the uh the the main problem is that first the main problem is that you when first the main problem is that you when first the main problem is that you when you ask a follow-up question you need to you ask a follow-up question you need to you ask a follow-up question you need to recomputee all previous tokens every

  13. recomputee all previous tokens every recomputee all previous tokens every time. So you just end up paying twice time. So you just end up paying twice time. So you just end up paying twice for the same tokens or three times or for the same tokens or three times or for the same tokens or three times or four times if the conversation is going four times if the conversation is going four times if the conversation is going which is obviously far from ideal. So which is obviously far from ideal. So which is obviously far from ideal. So what providers do nowadays is to offer what providers do nowadays is to offer what providers do nowadays is to offer prom caching. So you will it they will prom caching. So you will it they will prom caching. So you will it they will save the embedding and and KV cache and save the embedding and and KV cache and save the embedding and and KV cache and I won't enter into the details but they I won't enter into the details but they I won't enter into the details but they will premputee they will have saved a will premputee they will have saved a will premputee they will have saved a lot of the the compute for some tokens lot of the the compute for some tokens lot of the the compute for some tokens and you can just reload them and what's and you can just reload them and what's and you can just reload them and what's interesting for us is that these already interesting for us is that these already interesting for us is that these already sent token that you reuse are much much sent token that you reuse are much much sent token that you reuse are much much much cheaper uh specifically it can go much cheaper uh specifically it can go much cheaper uh specifically it can go up to 50 times cheaper with some API up to 50 times cheaper with some API up to 50 times cheaper with some API like deepseek uh which we will discuss like deepseek uh which we will discuss like deepseek uh which we will discuss in the experiments in the experiments in the experiments and what that means is that If you send and what that means is that If you send and what that means is that If you send a very long context, a very long context, a very long context, you will pay just one 15th of the the you will pay just one 15th of the the you will pay just one 15th of the the price per token and just pay the full price per token and just pay the full price per token and just pay the full price of the qu the new user question or price of the qu the new user question or price of the qu the new user question or the new interaction. And that's a the new interaction. And that's a the new interaction. And that's a problem for compaction because when you problem for compaction because when you problem for compaction because when you are compacting, summarizing, doing any are compacting, summarizing, doing any are compacting, summarizing, doing any transformation to this context, the the transformation to this context, the the transformation to this context, the the provider cannot use the cache because it provider cannot use the cache because it provider cannot use the cache because it it's a new context. It doesn't the model it's a new context. It doesn't the model it's a new context. It doesn't the model is not intelligent enough to understand is not intelligent enough to understand is not intelligent enough to understand it's the same topic change. it it just it's the same topic change. it it just it's the same topic change. it it just cannot use the cache. So you will pay cannot use the cache. So you will pay cannot use the cache. So you will pay full price for these new transform full price for these new transform full price for these new transform tokens. So here what it means is that tokens. So here what it means is that tokens. So here what it means is that you need to to compress to for for you need to to compress to for for you need to to compress to for for compaction to be worthwhile. You need to compaction to be worthwhile. You need to compaction to be worthwhile. You need to compress by more than 50 times the compress by more than 50 times the compress by more than 50 times the context. So it can be quite difficult in context. So it can be quite difficult in context. So it can be quite difficult in some cases without losing

  14. some cases without losing some cases without losing quality. So caching is truly a gamecher quality. So caching is truly a gamecher quality. So caching is truly a gamecher especially because nowadays almost all especially because nowadays almost all especially because nowadays almost all APIs offer it very easy to use and APIs offer it very easy to use and APIs offer it very easy to use and typically they they save cost on like on typically they they save cost on like on typically they they save cost on like on 90% of the cost when you use caching but 90% of the cost when you use caching but 90% of the cost when you use caching but in some case as I said it can go up to in some case as I said it can go up to in some case as I said it can go up to way more than that and not only that it way more than that and not only that it way more than that and not only that it helps with cost but caching cache tokens helps with cost but caching cache tokens helps with cost but caching cache tokens are also already computed so it's way are also already computed so it's way are also already computed so it's way faster to get the first the answer So ultimately it means that So ultimately it means that summarization is potentially a trap. You summarization is potentially a trap. You summarization is potentially a trap. You may not want to use it at all or you you may not want to use it at all or you you may not want to use it at all or you you may want to just use it very may want to just use it very may want to just use it very specifically specifically specifically which is what the most serious harnesses which is what the most serious harnesses which is what the most serious harnesses do nowadays. cloud code, C codeex and do nowadays. cloud code, C codeex and do nowadays. cloud code, C codeex and all of them use context caching but also all of them use context caching but also all of them use context caching but also use a different method of compactions use a different method of compactions use a different method of compactions and we know that because obviously of and we know that because obviously of and we know that because obviously of the leak and then because codeex is open the leak and then because codeex is open the leak and then because codeex is open source and then the APIs also provide source and then the APIs also provide source and then the APIs also provide ways to manage context directly and ways to manage context directly and ways to manage context directly and caching directly. So it's it becomes caching directly. So it's it becomes caching directly. So it's it becomes when you build yourself a harness like when you build yourself a harness like when you build yourself a harness like we do with the AI tutor, it becomes we do with the AI tutor, it becomes we do with the AI tutor, it becomes really interesting to understand when to really interesting to understand when to really interesting to understand when to use which technique and test them use which technique and test them use which technique and test them obviously.

  15. obviously. obviously. So where is this going? It mean it means So where is this going? It mean it means So where is this going? It mean it means that you don't you don't want to just that you don't you don't want to just that you don't you don't want to just compact you don't want to summarize compact you don't want to summarize compact you don't want to summarize everything any time because it may kill everything any time because it may kill everything any time because it may kill the cache and you won't be able to use the cache and you won't be able to use the cache and you won't be able to use it. So some best guidance that we found it. So some best guidance that we found it. So some best guidance that we found is that obvious obviously when the user is that obvious obviously when the user is that obvious obviously when the user seems to talk about a very different seems to talk about a very different seems to talk about a very different topic, you may want to refresh the topic, you may want to refresh the topic, you may want to refresh the session to clean it to clear it. Um when session to clean it to clear it. Um when session to clean it to clear it. Um when you scope or files, use what I described you scope or files, use what I described you scope or files, use what I described with progressive disclosure to just show with progressive disclosure to just show with progressive disclosure to just show the smallest amount of context possible. the smallest amount of context possible. the smallest amount of context possible. You want to clear every old tool outputs You want to clear every old tool outputs You want to clear every old tool outputs that are not useful anymore. You may that are not useful anymore. You may that are not useful anymore. You may want to compact in some case. We will want to compact in some case. We will want to compact in some case. We will see that in the experiments in a few see that in the experiments in a few see that in the experiments in a few minutes. minutes. minutes. And you may want to optimize for cache And you may want to optimize for cache And you may want to optimize for cache hits. So just having the model use be hits. So just having the model use be hits. So just having the model use be able to use its cache. And regarding able to use its cache. And regarding able to use its cache. And regarding that providers are constantly improving that providers are constantly improving that providers are constantly improving their feature set to manage cache. So their feature set to manage cache. So their feature set to manage cache. So you just need to uh stay current and you just need to uh stay current and you just need to uh stay current and follow what what the API allows follow what what the API allows follow what what the API allows nowadays. But they all provide different nowadays. But they all provide different nowadays. But they all provide different methods and um you will have access to methods and um you will have access to methods and um you will have access to the slides but there's I put a link the slides but there's I put a link the slides but there's I put a link earlier in the in the first few slides earlier in the in the first few slides earlier in the in the first few slides on a very interesting article regarding on a very interesting article regarding on a very interesting article regarding regarding prompt caching that I regarding prompt caching that I regarding prompt caching that I recommend checking out. It's like on the recommend checking out. It's like on the recommend checking out. It's like on the sixth or seventh slide but you will have sixth or seventh slide but you will have sixth or seventh slide but you will have the link to the slides.

  16. the link to the slides. the link to the slides. And lastly, you may want to use a model And lastly, you may want to use a model And lastly, you may want to use a model router to optimize especially cost on router to optimize especially cost on router to optimize especially cost on various tasks. various tasks. various tasks. And most importantly, and what the And most importantly, and what the And most importantly, and what the majority of people don't do, you want to majority of people don't do, you want to majority of people don't do, you want to log everything. It's super easy. You log everything. It's super easy. You log everything. It's super easy. You just ask cloud to implement opic and just ask cloud to implement opic and just ask cloud to implement opic and track everything. It's you don't have track everything. It's you don't have track everything. It's you don't have anything to do. So it's definitely anything to do. So it's definitely anything to do. So it's definitely worthwhile to to implement and you can worthwhile to to implement and you can worthwhile to to implement and you can track cache it rate. You can track user track cache it rate. You can track user track cache it rate. You can track user frustration which we've seen that cloud frustration which we've seen that cloud frustration which we've seen that cloud does. So I keep uh telling it when I'm does. So I keep uh telling it when I'm does. So I keep uh telling it when I'm not happy. And uh you you may want to not happy. And uh you you may want to not happy. And uh you you may want to log for some um abnormally long outputs log for some um abnormally long outputs log for some um abnormally long outputs or any weird behavior that a small or any weird behavior that a small or any weird behavior that a small language model could detect. And so all language model could detect. And so all language model could detect. And so all of this together is the is context of this together is the is context of this together is the is context engineering which basically means to engineering which basically means to engineering which basically means to decide what the model sees every time decide what the model sees every time decide what the model sees every time you call it. And to us for the AI tutor, you call it. And to us for the AI tutor, you call it. And to us for the AI tutor, it means to decide what to keep in the it means to decide what to keep in the it means to decide what to keep in the current context window in order to current context window in order to current context window in order to optimize the caching what to drop or optimize the caching what to drop or optimize the caching what to drop or compact and when to do that. Uh and compact and when to do that. Uh and compact and when to do that. Uh and those two first things are exactly what those two first things are exactly what those two first things are exactly what what we studied in many experiments that what we studied in many experiments that what we studied in many experiments that we that my colleague Omar will share we that my colleague Omar will share we that my colleague Omar will share with you right now. And uh so you can with you right now. And uh so you can with you right now. And uh so you can follow along with the experiments on the follow along with the experiments on the follow along with the experiments on the the QR code link. It's the hugging face the QR code link. It's the hugging face the QR code link. It's the hugging face that I mentioned earlier and in there that I mentioned earlier and in there that I mentioned earlier and in there there's also a link to the repo and

  17. there's also a link to the repo and there's also a link to the repo and everything. everything. everything. All right. So, welcome to this second All right. So, welcome to this second All right. So, welcome to this second part of the of the workshop. So now that part of the of the workshop. So now that part of the of the workshop. So now that we have an overall idea of what context we have an overall idea of what context we have an overall idea of what context engineering is and what are the engineering is and what are the engineering is and what are the different techniques that we can apply different techniques that we can apply different techniques that we can apply to our agents to our agents to our agents uh now what we want to do here is see uh now what we want to do here is see uh now what we want to do here is see how they actually perform in in our case how they actually perform in in our case how they actually perform in in our case for our AI tutor. I will first start by for our AI tutor. I will first start by for our AI tutor. I will first start by describing a little bit about how the AI describing a little bit about how the AI describing a little bit about how the AI tutor works. So the system design and tutor works. So the system design and tutor works. So the system design and then I will follow up with the initial then I will follow up with the initial then I will follow up with the initial experiments that we did. All right. So, the AI tutor is actually All right. So, the AI tutor is actually very simple. It's one agent that it's a very simple. It's one agent that it's a very simple. It's one agent that it's a React type of agent that just loops over React type of agent that just loops over React type of agent that just loops over over tool calls and thinking blocks. And over tool calls and thinking blocks. And over tool calls and thinking blocks. And here we just create a very simple one here we just create a very simple one here we just create a very simple one using the lang chain library. So we use using the lang chain library. So we use using the lang chain library. So we use the create agent method with the the create agent method with the the create agent method with the in-memory saver because we in this case in-memory saver because we in this case in-memory saver because we in this case we don't save past conversations we just we don't save past conversations we just we don't save past conversations we just use the the current history use the the current history use the the current history and to customize this agent we use the and to customize this agent we use the and to customize this agent we use the middleware feature of lang chain where middleware feature of lang chain where middleware feature of lang chain where you we can add different uh features you we can add different uh features you we can add different uh features that can change the behavior of the that can change the behavior of the that can change the behavior of the agent at runtime. So in this case, we agent at runtime. So in this case, we agent at runtime. So in this case, we want to summarize for example and clear want to summarize for example and clear want to summarize for example and clear the tool outputs

  18. the tool outputs the tool outputs or also in our case have the the user be or also in our case have the the user be or also in our case have the the user be able to choose the the source uh uh the able to choose the the source uh uh the able to choose the the source uh uh the sources the like which lessons which sources the like which lessons which sources the like which lessons which courses to to use to to answer. courses to to use to to answer. courses to to use to to answer. To do that we also add two different To do that we also add two different To do that we also add two different tools. So we have the first one the tools. So we have the first one the tools. So we have the first one the retrieve tutor context which which uses retrieve tutor context which which uses retrieve tutor context which which uses a very classic hybrid search pipeline. a very classic hybrid search pipeline. a very classic hybrid search pipeline. So um semantic search along with keyword So um semantic search along with keyword So um semantic search along with keyword search and we combine both results to search and we combine both results to search and we combine both results to get the the best possible list of get the the best possible list of get the the best possible list of chunks. And then recently we also added chunks. And then recently we also added chunks. And then recently we also added the second one which is letting the the second one which is letting the the second one which is letting the agent actually browse the file system agent actually browse the file system agent actually browse the file system the knowledge base uh just like we can the knowledge base uh just like we can the knowledge base uh just like we can just like coding agents can in the just like coding agents can in the just like coding agents can in the browsing your codebase for example and browsing your codebase for example and browsing your codebase for example and we also borrow from the idea of karpathy we also borrow from the idea of karpathy we also borrow from the idea of karpathy by creating a wiki and helping the agent by creating a wiki and helping the agent by creating a wiki and helping the agent basically browse uh this knowledge base basically browse uh this knowledge base basically browse uh this knowledge base more easily.

  19. more easily. more easily. I will come back to these tools in a few I will come back to these tools in a few I will come back to these tools in a few in a few me moments. We also add a fast in a few me moments. We also add a fast in a few me moments. We also add a fast API app to wrap the whole system with an API app to wrap the whole system with an API app to wrap the whole system with an endpoint and we add a nextgs UI to let endpoint and we add a nextgs UI to let endpoint and we add a nextgs UI to let students use the DII tutor. students use the DII tutor. students use the DII tutor. So I will talk a little bit about the So I will talk a little bit about the So I will talk a little bit about the first tool. So we have a large corpus. first tool. So we have a large corpus. first tool. So we have a large corpus. So we have all of the lessons from all So we have all of the lessons from all So we have all of the lessons from all the different courses that we created the different courses that we created the different courses that we created over the past two years. over the past two years. over the past two years. uh and also documentation from various uh and also documentation from various uh and also documentation from various public open source libraries like lang public open source libraries like lang public open source libraries like lang chain, lama index and even documentation chain, lama index and even documentation chain, lama index and even documentation from from from uh openai. So they they make available uh openai. So they they make available uh openai. So they they make available available all of the markdown files from available all of the markdown files from available all of the markdown files from like how to use the open API, how to use like how to use the open API, how to use like how to use the open API, how to use codeex and we also have like cloud code codeex and we also have like cloud code codeex and we also have like cloud code documentation. So it's a very big corpus documentation. So it's a very big corpus documentation. So it's a very big corpus that has over 8 million tokens and of that has over 8 million tokens and of that has over 8 million tokens and of course this cannot fit in a single course this cannot fit in a single course this cannot fit in a single context window. So we have to store it context window. So we have to store it context window. So we have to store it in a way where we can retrieve the most in a way where we can retrieve the most in a way where we can retrieve the most important or most relevant information.

  20. important or most relevant information. important or most relevant information. So in this case the agent receives a So in this case the agent receives a So in this case the agent receives a question and the user can then choose question and the user can then choose question and the user can then choose well beforehand the user can choose well beforehand the user can choose well beforehand the user can choose specific source. So we can filter this specific source. So we can filter this specific source. So we can filter this knowledge base. It makes it better for knowledge base. It makes it better for knowledge base. It makes it better for um to improve recall. So precision to um to improve recall. So precision to um to improve recall. So precision to get the most relevant information. get the most relevant information. get the most relevant information. Then we do hybrid search. So this is Then we do hybrid search. So this is Then we do hybrid search. So this is very classic um hybrid search. We use very classic um hybrid search. We use very classic um hybrid search. We use embedding model. In this case it's a embedding model. In this case it's a embedding model. In this case it's a coherent coherent model with BM25 for coherent coherent model with BM25 for coherent coherent model with BM25 for the keyword index to get the the most the keyword index to get the the most the keyword index to get the the most relevant top 30 uh chunks. relevant top 30 uh chunks. relevant top 30 uh chunks. Then we merge the two results from the Then we merge the two results from the Then we merge the two results from the semantic similarity and the keyword semantic similarity and the keyword semantic similarity and the keyword search and we then rerank to the top search and we then rerank to the top search and we then rerank to the top five most relevant chunks five most relevant chunks five most relevant chunks and that's that's what return to the and that's that's what return to the and that's that's what return to the agent. We also have a limit of 100,000 agent. We also have a limit of 100,000 agent. We also have a limit of 100,000 tokens. So we don't want to. So let's tokens. So we don't want to. So let's tokens. So we don't want to. So let's say for example we we go over we just say for example we we go over we just say for example we we go over we just remove the last few chunk uh the last remove the last few chunk uh the last remove the last few chunk uh the last chunks that make it so that we don't chunks that make it so that we don't chunks that make it so that we don't cross that threshold. So these numbers cross that threshold. So these numbers cross that threshold. So these numbers the this configuration actually is not the this configuration actually is not the this configuration actually is not random. We did experiments random. We did experiments random. We did experiments uh to optimize this pipeline. I'm not uh to optimize this pipeline. I'm not uh to optimize this pipeline. I'm not going to talk about it but it's in one going to talk about it but it's in one going to talk about it but it's in one of the courses that we share. Uh of the courses that we share. Uh of the courses that we share. Uh basically we just want to try as many basically we just want to try as many basically we just want to try as many configurations as possible and improve configurations as possible and improve configurations as possible and improve recall. So we measure did we retrieve recall. So we measure did we retrieve recall. So we measure did we retrieve the correct page uh in our knowledge

  21. the correct page uh in our knowledge the correct page uh in our knowledge base and we just choose the best uh base and we just choose the best uh base and we just choose the best uh settings. settings. settings. So like I said it uh yeah in this case So like I said it uh yeah in this case So like I said it uh yeah in this case it's very precise very it's very good it's very precise very it's very good it's very precise very it's very good but what if the agent needs to browse but what if the agent needs to browse but what if the agent needs to browse the whole knowledge base to get the best the whole knowledge base to get the best the whole knowledge base to get the best possible answer. So let's say I want to possible answer. So let's say I want to possible answer. So let's say I want to learn about codeex and I want to learn learn about codeex and I want to learn learn about codeex and I want to learn about cloud code how can I best use uh about cloud code how can I best use uh about cloud code how can I best use uh those two tools and also like if we need those two tools and also like if we need those two tools and also like if we need to have v various documentation pages in to have v various documentation pages in to have v various documentation pages in context is this the best possible uh context is this the best possible uh context is this the best possible uh tool tool tool and there's a paper uh I I link it here and there's a paper uh I I link it here and there's a paper uh I I link it here in the in in the slides but it's a yeah in the in in the slides but it's a yeah in the in in the slides but it's a yeah it's a paper that shows that basically it's a paper that shows that basically it's a paper that shows that basically letting the agent browse the the letting the agent browse the the letting the agent browse the the knowledge base can be very bene knowledge base can be very bene knowledge base can be very bene beneicial uh so you can look at it if beneicial uh so you can look at it if beneicial uh so you can look at it if you want afterwards. Very very slow to you want afterwards. Very very slow to you want afterwards. Very very slow to load. load. load. Uh so that's why we ended up creating Uh so that's why we ended up creating Uh so that's why we ended up creating this second tool the run keep this second tool the run keep this second tool the run keep run knowledgebased command where the run knowledgebased command where the run knowledgebased command where the agent can browse u the knowledge base agent can browse u the knowledge base agent can browse u the knowledge base the file system using bash commands. So the file system using bash commands. So the file system using bash commands. So first of all what we did first is create first of all what we did first is create first of all what we did first is create these uh three different u folders. So these uh three different u folders. So these uh three different u folders. So we have the raw folder with all the we have the raw folder with all the we have the raw folder with all the different markdown files. So all the different markdown files. So all the different markdown files. So all the lessons from the different courses the lessons from the different courses the lessons from the different courses the documentation from the different open documentation from the different open documentation from the different open open source libraries. We we also have open source libraries. We we also have open source libraries. We we also have this generated uh folder with this is

  22. this generated uh folder with this is this generated uh folder with this is the generated automatically with the generated automatically with the generated automatically with basically just the the the titles of basically just the the the titles of basically just the the the titles of each of the markdown files so it's each of the markdown files so it's each of the markdown files so it's easier for the agent to to find um easier for the agent to to find um easier for the agent to to find um relevant information. And then we also relevant information. And then we also relevant information. And then we also have the wiki where uh in this case it have the wiki where uh in this case it have the wiki where uh in this case it was cl code that created this. It was cl code that created this. It was cl code that created this. It basically reads all of the raw files in basically reads all of the raw files in basically reads all of the raw files in the in the folder in the raw folder and the in the folder in the raw folder and the in the folder in the raw folder and it creates a very uh concise uh set of it creates a very uh concise uh set of it creates a very uh concise uh set of files. So topics, frameworks uh from the files. So topics, frameworks uh from the files. So topics, frameworks uh from the different source uh sources. So for different source uh sources. So for different source uh sources. So for example, I can have a topic related to example, I can have a topic related to example, I can have a topic related to finetuning. So in this case, the agent finetuning. So in this case, the agent finetuning. So in this case, the agent will be able to find all of the will be able to find all of the will be able to find all of the different raw files related to different raw files related to different raw files related to finetuning. For example, finetuning. For example, finetuning. For example, here we let the agent only read. So this here we let the agent only read. So this here we let the agent only read. So this is when we actually deploy it. So this is when we actually deploy it. So this is when we actually deploy it. So this is the one you can try on the on the is the one you can try on the on the is the one you can try on the on the space here. The agent can only read the space here. The agent can only read the space here. The agent can only read the knowledge base. It cannot modify it. Uh knowledge base. It cannot modify it. Uh knowledge base. It cannot modify it. Uh so we only allow these bash commands so we only allow these bash commands so we only allow these bash commands that basically cannot uh modify the the that basically cannot uh modify the the that basically cannot uh modify the the file system. We also put some limits file system. We also put some limits file system. We also put some limits around this. So for example, if a around this. So for example, if a around this. So for example, if a command lasts over 8 seconds, we can command lasts over 8 seconds, we can command lasts over 8 seconds, we can just return an error or let the agent just return an error or let the agent just return an error or let the agent execute something else because it's execute something else because it's execute something else because it's taking too much time. And we also cap taking too much time. And we also cap taking too much time. And we also cap the tool outputs to 40,000 uh the tool outputs to 40,000 uh the tool outputs to 40,000 uh characters. So in this case, for characters. So in this case, for characters. So in this case, for example, if a lesson is over this example, if a lesson is over this example, if a lesson is over this amount, what the agent can do is then amount, what the agent can do is then amount, what the agent can do is then okay, so I just got the first 40,000.

  23. okay, so I just got the first 40,000. okay, so I just got the first 40,000. Let me let me do a follow-up command to Let me let me do a follow-up command to Let me let me do a follow-up command to get the last uh piece of the of the get the last uh piece of the of the get the last uh piece of the of the lesson. For example, lesson. For example, lesson. For example, uh we also limit the number of the of uh we also limit the number of the of uh we also limit the number of the of commands. We actually never see the commands. We actually never see the commands. We actually never see the agent go over this limit of 20 commands agent go over this limit of 20 commands agent go over this limit of 20 commands per turn, but this is just a a fallback per turn, but this is just a a fallback per turn, but this is just a a fallback uh in case it takes too too much time to uh in case it takes too too much time to uh in case it takes too too much time to to answer. And we also sandbox of course to answer. And we also sandbox of course to answer. And we also sandbox of course the agent to only browse the knowledge the agent to only browse the knowledge the agent to only browse the knowledge base this specific folder. base this specific folder. base this specific folder. Like I said, we create this offline. So Like I said, we create this offline. So Like I said, we create this offline. So the the three different file folders we the the three different file folders we the the three different file folders we create this once or every time we want create this once or every time we want create this once or every time we want to add a new uh a new course for example to add a new uh a new course for example to add a new uh a new course for example we we we tell cloud code can you add a we we we tell cloud code can you add a we we we tell cloud code can you add a new uh can you add this to the raw raw new uh can you add this to the raw raw new uh can you add this to the raw raw folder and also create new topics around folder and also create new topics around folder and also create new topics around this new new content. So we do this once this new new content. So we do this once this new new content. So we do this once and then we deploy it so that the agent and then we deploy it so that the agent and then we deploy it so that the agent u can can browse it with our current u can can browse it with our current u can can browse it with our current system prompt. We can see that it's uh system prompt. We can see that it's uh system prompt. We can see that it's uh used almost every time. So for almost used almost every time. So for almost used almost every time. So for almost 90% of the terms we can tweak this to 90% of the terms we can tweak this to 90% of the terms we can tweak this to make it use it less uh or more. We make it use it less uh or more. We make it use it less uh or more. We didn't optimize for for this didn't optimize for for this didn't optimize for for this specifically.

  24. specifically. specifically. And I guess the most interesting aspect And I guess the most interesting aspect And I guess the most interesting aspect of doing this is that we measured the of doing this is that we measured the of doing this is that we measured the precision the recall of using this tool precision the recall of using this tool precision the recall of using this tool uh actually turning turning it off and uh actually turning turning it off and uh actually turning turning it off and we actually got the same amount of we actually got the same amount of we actually got the same amount of recall. So recall. So recall. So just using the first tool was enough to just using the first tool was enough to just using the first tool was enough to get all the relevant information and get all the relevant information and get all the relevant information and just having this second tool was just just having this second tool was just just having this second tool was just 50% slower. Basically, it's faster 50% slower. Basically, it's faster 50% slower. Basically, it's faster without this tool because it does less without this tool because it does less without this tool because it does less tool calls. Um, and yeah, we we tool calls. Um, and yeah, we we tool calls. Um, and yeah, we we basically didn't see any improvement on basically didn't see any improvement on basically didn't see any improvement on uh on re on answering with the correct uh on re on answering with the correct uh on re on answering with the correct uh documentation. uh documentation. uh documentation. It it was fun to add but we yeah we It it was fun to add but we yeah we It it was fun to add but we yeah we didn't see any any benefit and one didn't see any any benefit and one didn't see any any benefit and one reason for that is that we tested using reason for that is that we tested using reason for that is that we tested using real world well basically the questions real world well basically the questions real world well basically the questions we get from students and basically those we get from students and basically those we get from students and basically those questions were complex enough we guess questions were complex enough we guess questions were complex enough we guess to actually benefit from using this new to actually benefit from using this new to actually benefit from using this new new setup. Um new setup. Um new setup. Um so this is what this is how initially uh so this is what this is how initially uh so this is what this is how initially uh our AI tutor managed its context. Uh we our AI tutor managed its context. Uh we our AI tutor managed its context. Uh we we we we started with this because uh it looked started with this because uh it looked started with this because uh it looked fast, it looked good. Uh we didn't fast, it looked good. Uh we didn't fast, it looked good. Uh we didn't actually measure anything. It was just actually measure anything. It was just actually measure anything. It was just like oh it looks good. Okay, we will like oh it looks good. Okay, we will like oh it looks good. Okay, we will just set it set it like this. And so just set it set it like this. And so just set it set it like this. And so basically we have these three different basically we have these three different basically we have these three different um um um context engineering techniques where we

  25. context engineering techniques where we context engineering techniques where we clear the outputs uh after 5,000 tokens clear the outputs uh after 5,000 tokens clear the outputs uh after 5,000 tokens uh but we at least keep the last five uh but we at least keep the last five uh but we at least keep the last five ones. So this was basically a way to ones. So this was basically a way to ones. So this was basically a way to keep a very small context uh during a keep a very small context uh during a keep a very small context uh during a conversation. We also added the the conversation. We also added the the conversation. We also added the the capacity to summarize. So after 30,000 capacity to summarize. So after 30,000 capacity to summarize. So after 30,000 tokens uh the two well the system tokens uh the two well the system tokens uh the two well the system summarizes the uh the history but we at summarizes the uh the history but we at summarizes the uh the history but we at least keep the last 20 messages to make least keep the last 20 messages to make least keep the last 20 messages to make sure that uh those are very accurate sure that uh those are very accurate sure that uh those are very accurate those messages. And we also have the those messages. And we also have the those messages. And we also have the source preference. So that that just source preference. So that that just source preference. So that that just allows people to choose what uh sources allows people to choose what uh sources allows people to choose what uh sources to use uh when the thetutor answers. to use uh when the thetutor answers. to use uh when the thetutor answers. But like I said, these are unproven But like I said, these are unproven But like I said, these are unproven unproven defaults and we actually want unproven defaults and we actually want unproven defaults and we actually want to know what actually works best. Um so to know what actually works best. Um so to know what actually works best. Um so I guess you you can Lou showed this at I guess you you can Lou showed this at I guess you you can Lou showed this at the beginning, but you can access the the beginning, but you can access the the beginning, but you can access the tutor live. I'm just going to show this tutor live. I'm just going to show this tutor live. I'm just going to show this very quickly. This is the huging face.

  26. very quickly. This is the huging face. very quickly. This is the huging face. So this is a separate UI just to show So this is a separate UI just to show So this is a separate UI just to show you. We also have the the chat bubble you. We also have the the chat bubble you. We also have the the chat bubble one on the lessons on the course one on the lessons on the course one on the lessons on the course themselves itself. Uh and here on the on themselves itself. Uh and here on the on themselves itself. Uh and here on the on the left you can choose the different uh the left you can choose the different uh the left you can choose the different uh sources to use, enable them or disable sources to use, enable them or disable sources to use, enable them or disable them and then you can uh yeah send your them and then you can uh yeah send your them and then you can uh yeah send your your query. So as you can see here I can your query. So as you can see here I can your query. So as you can see here I can just uh send a new request and we see just uh send a new request and we see just uh send a new request and we see Gemini in this case Gemini 3.5 flash use Gemini in this case Gemini 3.5 flash use Gemini in this case Gemini 3.5 flash use its reasoning and use its uh capacity to its reasoning and use its uh capacity to its reasoning and use its uh capacity to to do tool calling and to answer and I to do tool calling and to answer and I to do tool calling and to answer and I guess uh here there's a I think it's guess uh here there's a I think it's guess uh here there's a I think it's just the internet and yeah the code is open source so you and yeah the code is open source so you can use your favorite coding agent to can use your favorite coding agent to can use your favorite coding agent to explore the codebase and uh and learn explore the codebase and uh and learn explore the codebase and uh and learn about how we implemented this about how we implemented this about how we implemented this specifically. So we have the different specifically. So we have the different specifically. So we have the different activity u activity u activity u the the activity that the the the model the the activity that the the the model the the activity that the the the model did. So for tool calls it thought four did. So for tool calls it thought four did. So for tool calls it thought four times and it use 10 sources and we have times and it use 10 sources and we have times and it use 10 sources and we have the the final answer.

  27. So everything that uh so every time the So everything that uh so every time the student uses this chatbot we actually student uses this chatbot we actually student uses this chatbot we actually log everything. So we for every single log everything. So we for every single log everything. So we for every single turn we have uh the input tokens, the turn we have uh the input tokens, the turn we have uh the input tokens, the the output tokens, how many of them were the output tokens, how many of them were the output tokens, how many of them were cached, uh what was the cost, what time cached, uh what was the cost, what time cached, uh what was the cost, what time it took to to get the first token, how it took to to get the first token, how it took to to get the first token, how many tool calls it did and if we many tool calls it did and if we many tool calls it did and if we actually if the system actually did actually if the system actually did actually if the system actually did something around summarization. So this something around summarization. So this something around summarization. So this is very useful and that's uh what we are is very useful and that's uh what we are is very useful and that's uh what we are going to use when uh measuring the uh going to use when uh measuring the uh going to use when uh measuring the uh the different techniques. So why do we the different techniques. So why do we the different techniques. So why do we need to measure? Well, because need to measure? Well, because need to measure? Well, because the the techniques that Lou showed all the the techniques that Lou showed all the the techniques that Lou showed all sound very smart. So you you might think sound very smart. So you you might think sound very smart. So you you might think that they are very uh useful, but that they are very uh useful, but that they are very uh useful, but sometimes they're not. And uh as we as sometimes they're not. And uh as we as sometimes they're not. And uh as we as we discovered with our experiments, we discovered with our experiments, we discovered with our experiments, actually it might be detrimental. So actually it might be detrimental. So actually it might be detrimental. So because of the the way APIs cache the because of the the way APIs cache the because of the the way APIs cache the tokens when they are um when they are tokens when they are um when they are tokens when they are um when they are sent and it's also difficult to to know sent and it's also difficult to to know sent and it's also difficult to to know in advance what what is best to use.

  28. in advance what what is best to use. in advance what what is best to use. So before I go into the experiments, I So before I go into the experiments, I So before I go into the experiments, I just want to just want to just want to define a few words because we're I'm define a few words because we're I'm define a few words because we're I'm going to use these words throughout the going to use these words throughout the going to use these words throughout the uh the the presentation. So a preset is uh the the presentation. So a preset is uh the the presentation. So a preset is basically uh the way the AI tutor was basically uh the way the AI tutor was basically uh the way the AI tutor was set up in the in the experiment. So for set up in the in the experiment. So for set up in the in the experiment. So for example, it can be in this preset we did example, it can be in this preset we did example, it can be in this preset we did summarization at this amount of tokens summarization at this amount of tokens summarization at this amount of tokens or in this preset we we used sliding or in this preset we we used sliding or in this preset we we used sliding window for example. So that's what a window for example. So that's what a window for example. So that's what a preset is. we have the different tasks. preset is. we have the different tasks. preset is. we have the different tasks. So at task type uh for these initial So at task type uh for these initial So at task type uh for these initial experiments we only did two tasks single experiments we only did two tasks single experiments we only did two tasks single turn and multiple turns also sessions uh turn and multiple turns also sessions uh turn and multiple turns also sessions uh called sessions and one run is basically called sessions and one run is basically called sessions and one run is basically just running one preset on a task and just running one preset on a task and just running one preset on a task and then you get the run. Um the bundle is then you get the run. Um the bundle is then you get the run. Um the bundle is just the the result. So it's just a JSON just the the result. So it's just a JSON just the the result. So it's just a JSON file with all the different uh metrics file with all the different uh metrics file with all the different uh metrics that we save to the disk.

  29. that we save to the disk. that we save to the disk. So first task single turn. Uh so these So first task single turn. Uh so these So first task single turn. Uh so these are question and answers and we didn't are question and answers and we didn't are question and answers and we didn't generate this. It's not synthetic. Uh we generate this. It's not synthetic. Uh we generate this. It's not synthetic. Uh we actually I I had codeex actually I I had codeex actually I I had codeex scrape all of the uh questions and scrape all of the uh questions and scrape all of the uh questions and responses from our website where responses from our website where responses from our website where students can ask questions and get students can ask questions and get students can ask questions and get answers from uh members of the staff. Uh answers from uh members of the staff. Uh answers from uh members of the staff. Uh and that's how I got this initial we got and that's how I got this initial we got and that's how I got this initial we got this initial um data set. uh we cleaned this initial um data set. uh we cleaned this initial um data set. uh we cleaned the data set and only used 60 60 pairs the data set and only used 60 60 pairs the data set and only used 60 60 pairs uh because we saw that some of the uh because we saw that some of the uh because we saw that some of the questions weren't good for the type of questions weren't good for the type of questions weren't good for the type of task for example uh there were old task for example uh there were old task for example uh there were old questions about previous versions of questions about previous versions of questions about previous versions of some libraries and we we might not I some libraries and we we might not I some libraries and we we might not I mean the right now if the tutor answers mean the right now if the tutor answers mean the right now if the tutor answers it's not going to be using this old it's not going to be using this old it's not going to be using this old version of the library so we just uh version of the library so we just uh version of the library so we just uh remove some of the questions some remove some of the questions some remove some of the questions some duplicate ones and and and we got this duplicate ones and and and we got this duplicate ones and and and we got this first uh data set and what we measure is first uh data set and what we measure is first uh data set and what we measure is the retrieval. So did we retrieve the the retrieval. So did we retrieve the the retrieval. So did we retrieve the correct is did the total retrieve the correct is did the total retrieve the correct is did the total retrieve the correct lesson for example this is done correct lesson for example this is done correct lesson for example this is done automatically we can see uh just by automatically we can see uh just by automatically we can see uh just by looking at the code did we did we use looking at the code did we did we use looking at the code did we did we use the correct uh lesson or not and we also the correct uh lesson or not and we also the correct uh lesson or not and we also look at did do we have the correct facts look at did do we have the correct facts look at did do we have the correct facts in the answer or the right kind of in the answer or the right kind of in the answer or the right kind of response in the answer. These two uh are response in the answer. These two uh are response in the answer. These two uh are actually graded using using an LLM. Uh

  30. actually graded using using an LLM. Uh actually graded using using an LLM. Uh you can use APIs to do it but uh right you can use APIs to do it but uh right you can use APIs to do it but uh right now I think the best way to do it is to now I think the best way to do it is to now I think the best way to do it is to just use use your uh code code just use use your uh code code just use use your uh code code subscription or your codec subscription subscription or your codec subscription subscription or your codec subscription because it's cheaper than using the because it's cheaper than using the because it's cheaper than using the APIs. APIs. APIs. We also have the second task the We also have the second task the We also have the second task the session. So multi-turn um conversations session. So multi-turn um conversations session. So multi-turn um conversations back and forth. uh here what we want to back and forth. uh here what we want to back and forth. uh here what we want to know is know is know is is the AI tutor able to recall facts is the AI tutor able to recall facts is the AI tutor able to recall facts after multiple terms. So this is a bit after multiple terms. So this is a bit after multiple terms. So this is a bit uh in this case we do use some generated uh in this case we do use some generated uh in this case we do use some generated uh content. So we do in generate facts uh content. So we do in generate facts uh content. So we do in generate facts that we put at the beginning. So we have that we put at the beginning. So we have that we put at the beginning. So we have a student like a fake student state uh a student like a fake student state uh a student like a fake student state uh fact. So for example, I want to learn fact. So for example, I want to learn fact. So for example, I want to learn about rag and then we stuff the about rag and then we stuff the about rag and then we stuff the conversation with just filler messages conversation with just filler messages conversation with just filler messages because we just want yeah we just want because we just want yeah we just want because we just want yeah we just want to have a lot of messages and then we to have a lot of messages and then we to have a lot of messages and then we have um have um have um a probe which is just the student asking a probe which is just the student asking a probe which is just the student asking a question again. So for example it can a question again. So for example it can a question again. So for example it can be uh what should I learn today? And be uh what should I learn today? And be uh what should I learn today? And since the uh I'm going to go through it since the uh I'm going to go through it since the uh I'm going to go through it in the next slide, but uh yeah, let's in the next slide, but uh yeah, let's in the next slide, but uh yeah, let's let's let's the let's let's the let's let's the example. So for example, here at turn example. So for example, here at turn example. So for example, here at turn one, we we uh we have the student the one, we we uh we have the student the one, we we uh we have the student the fake student implement a state of fact.

  31. fake student implement a state of fact. fake student implement a state of fact. So it the student wants to learn about So it the student wants to learn about So it the student wants to learn about rag evaluation. Um and that that's rag evaluation. Um and that that's rag evaluation. Um and that that's basically the fact. Then we just add a basically the fact. Then we just add a basically the fact. Then we just add a lot of messages uh filler filler lot of messages uh filler filler lot of messages uh filler filler messages and then at at some point we messages and then at at some point we messages and then at at some point we have the student say for example what have the student say for example what have the student say for example what topic should I learn about today and topic should I learn about today and topic should I learn about today and then what we expect the AI tutor to to then what we expect the AI tutor to to then what we expect the AI tutor to to to say is that uh the student should to say is that uh the student should to say is that uh the student should learn about rag evaluation. So more learn about rag evaluation. So more learn about rag evaluation. So more specifically heat rate and MR for specifically heat rate and MR for specifically heat rate and MR for example. example. example. We also have a gate part in the We also have a gate part in the We also have a gate part in the evaluation where let's say we are evaluation where let's say we are evaluation where let's say we are testing the summarization technique. Uh testing the summarization technique. Uh testing the summarization technique. Uh we actually just want to know did we actually just want to know did we actually just want to know did summarization actually happen or not. Um summarization actually happen or not. Um summarization actually happen or not. Um so this is this is this is like one so this is this is this is like one so this is this is this is like one example of one task in the sessions uh example of one task in the sessions uh example of one task in the sessions uh data set in the sessions task. data set in the sessions task. data set in the sessions task. Now we have the evaluation hardness. Um Now we have the evaluation hardness. Um Now we have the evaluation hardness. Um so the main uh harness I guess the the so the main uh harness I guess the the so the main uh harness I guess the the main function is the run the run task main function is the run the run task main function is the run the run task the run battery function that just uh the run battery function that just uh the run battery function that just uh runs this uh this task. Um we have the runs this uh this task. Um we have the runs this uh this task. Um we have the grading like I said we it can either be grading like I said we it can either be grading like I said we it can either be a code check so did we retrieve the a code check so did we retrieve the a code check so did we retrieve the correct lesson or not and we can also correct lesson or not and we can also correct lesson or not and we can also have the LM as a judge and in this case have the LM as a judge and in this case have the LM as a judge and in this case we use the subscription of code. We also we use the subscription of code. We also we use the subscription of code. We also have the check triggers aspect. So this have the check triggers aspect. So this have the check triggers aspect. So this is just a check to to see that if the is just a check to to see that if the is just a check to to see that if the evaluation went uh good or not. Did we

  32. evaluation went uh good or not. Did we evaluation went uh good or not. Did we actually compact or not? This is this is actually compact or not? This is this is actually compact or not? This is this is just to make sure that the the run is just to make sure that the the run is just to make sure that the the run is actually good and we can save it. And actually good and we can save it. And actually good and we can save it. And then we have a generated report uh to then we have a generated report uh to then we have a generated report uh to see like what was the latency, what was see like what was the latency, what was see like what was the latency, what was the time to first token and every metric the time to first token and every metric the time to first token and every metric that we can uh measure. that we can uh measure. that we can uh measure. So yeah, basically we can evaluate So yeah, basically we can evaluate So yeah, basically we can evaluate everything and then grade it afterwards. everything and then grade it afterwards. everything and then grade it afterwards. Uh we we can run it once and grade it Uh we we can run it once and grade it Uh we we can run it once and grade it afterwards. afterwards. afterwards. So what what we run so we run 11 presets So what what we run so we run 11 presets So what what we run so we run 11 presets uh and we change them uh for for each uh and we change them uh for for each uh and we change them uh for for each experiment. So we have the the full experiment. So we have the the full experiment. So we have the the full history. So these are the the main ones history. So these are the the main ones history. So these are the the main ones the full history. So this is the case the full history. So this is the case the full history. So this is the case where we don't don't touch the context. where we don't don't touch the context. where we don't don't touch the context. We we leave everything as is in the We we leave everything as is in the We we leave everything as is in the history. And then we also have the history. And then we also have the history. And then we also have the production that I showed at the production that I showed at the production that I showed at the beginning the defaults that we have. Uh beginning the defaults that we have. Uh beginning the defaults that we have. Uh so these are the reference points and so these are the reference points and so these are the reference points and then we have these six techniques that I then we have these six techniques that I then we have these six techniques that I want to compare. So sliding window want to compare. So sliding window want to compare. So sliding window prompt compression uh selective prompt compression uh selective prompt compression uh selective retention and the other ones and I what retention and the other ones and I what retention and the other ones and I what I want to see is just what uh memory I want to see is just what uh memory I want to see is just what uh memory recall recall do I get if I keep recall recall do I get if I keep recall recall do I get if I keep everything else fixed. So I use the same everything else fixed. So I use the same everything else fixed. So I use the same model the same prompt the same tools the model the same prompt the same tools the model the same prompt the same tools the same data set what's the difference? Uh same data set what's the difference? Uh same data set what's the difference? Uh now ju now just doing this was a bit now ju now just doing this was a bit now ju now just doing this was a bit expensive. I I didn't expect this to to expensive. I I didn't expect this to to expensive. I I didn't expect this to to get over $500, but it did. And that's

  33. get over $500, but it did. And that's get over $500, but it did. And that's one of the reason we did afterwards one of the reason we did afterwards one of the reason we did afterwards follow-up experiments using cheaper follow-up experiments using cheaper follow-up experiments using cheaper models. But my colleague Sani will talk models. But my colleague Sani will talk models. But my colleague Sani will talk about this. Uh so what ARM so what about this. Uh so what ARM so what about this. Uh so what ARM so what basically what preset actually won. Uh, basically what preset actually won. Uh, basically what preset actually won. Uh, and so these are our um our results. And and so these are our um our results. And and so these are our um our results. And as you can see, h we weren't expecting as you can see, h we weren't expecting as you can see, h we weren't expecting this, but basically not touching um the this, but basically not touching um the this, but basically not touching um the context was actually the best the best context was actually the best the best context was actually the best the best strategy for recovering this fact over strategy for recovering this fact over strategy for recovering this fact over time over multiple uh messages. Um you time over multiple uh messages. Um you time over multiple uh messages. Um you can see that the production so the the can see that the production so the the can see that the production so the the defaults that we thought were good defaults that we thought were good defaults that we thought were good enough were actually not the best. Uh enough were actually not the best. Uh enough were actually not the best. Uh not doing anything is actually better. H not doing anything is actually better. H not doing anything is actually better. H we did two uh two different experiments we did two uh two different experiments we did two uh two different experiments where one was just one trial and the where one was just one trial and the where one was just one trial and the second was um two trials. So we have second was um two trials. So we have second was um two trials. So we have like a more statist I I guess it's it's like a more statist I I guess it's it's like a more statist I I guess it's it's better but the uh the numbers are not better but the uh the numbers are not better but the uh the numbers are not like like like can might not be accurate because it's can might not be accurate because it's can might not be accurate because it's just one trial and two trials but I just one trial and two trials but I just one trial and two trials but I guess the most interesting thing is just guess the most interesting thing is just guess the most interesting thing is just the order in which the techniques ended the order in which the techniques ended the order in which the techniques ended up being in the table.

  34. up being in the table. up being in the table. So in this case keeping everything wins So in this case keeping everything wins So in this case keeping everything wins on the memory side but what about the on the memory side but what about the on the memory side but what about the cost? cost? cost? Uh this is what the production cost uh Uh this is what the production cost uh Uh this is what the production cost uh was for the single turn and the session. was for the single turn and the session. was for the single turn and the session. So almost 50 cents uh for a single turn So almost 50 cents uh for a single turn So almost 50 cents uh for a single turn and 24 cents uh for the multi-turn and 24 cents uh for the multi-turn and 24 cents uh for the multi-turn or each turn. Uh we actually had or each turn. Uh we actually had or each turn. Uh we actually had very good memory uh recall for pretty very good memory uh recall for pretty very good memory uh recall for pretty much all the techniques in the single much all the techniques in the single much all the techniques in the single turn because in single turn you don't turn because in single turn you don't turn because in single turn you don't have enough tokens to actually fire up have enough tokens to actually fire up have enough tokens to actually fire up the different strategies. Uh so for one the different strategies. Uh so for one the different strategies. Uh so for one response you don't need to do response you don't need to do response you don't need to do summarization compaction or anything summarization compaction or anything summarization compaction or anything like that. Uh so that's why you you see like that. Uh so that's why you you see like that. Uh so that's why you you see um high numbers. Uh but as you can see um high numbers. Uh but as you can see um high numbers. Uh but as you can see after in the multi-turn uh task you can after in the multi-turn uh task you can after in the multi-turn uh task you can see that the quality degraded to 38%. see that the quality degraded to 38%. see that the quality degraded to 38%. uh and we if we compare this to the full uh and we if we compare this to the full uh and we if we compare this to the full history. So here we don't touch the history. So here we don't touch the history. So here we don't touch the context uh the history of the model we context uh the history of the model we context uh the history of the model we can see that not touching uh is actually can see that not touching uh is actually can see that not touching uh is actually cheaper, it's faster and we have better cheaper, it's faster and we have better cheaper, it's faster and we have better recall overall. So keeping everything recall overall. So keeping everything recall overall. So keeping everything wins on on all of these three fronts.

  35. wins on on all of these three fronts. wins on on all of these three fronts. Uh so why does it so why is it actually Uh so why does it so why is it actually Uh so why does it so why is it actually uh uh uh like why is it uh why do we have less like why is it uh why do we have less like why is it uh why do we have less latency like what's we want to we wanted latency like what's we want to we wanted latency like what's we want to we wanted to understand that and it's basically to understand that and it's basically to understand that and it's basically because if you remove the tool outputs because if you remove the tool outputs because if you remove the tool outputs consistently then the agent needs to consistently then the agent needs to consistently then the agent needs to rerieve uh afterwards for information it rerieve uh afterwards for information it rerieve uh afterwards for information it already had. So you're just making the already had. So you're just making the already had. So you're just making the agent uh do more tool calls and that's agent uh do more tool calls and that's agent uh do more tool calls and that's why it ended up costing more uh and yeah why it ended up costing more uh and yeah why it ended up costing more uh and yeah using more tokens and having less uh using more tokens and having less uh using more tokens and having less uh uh less memory recall. So these are the uh less memory recall. So these are the uh less memory recall. So these are the results we initially got using uh Gemini results we initially got using uh Gemini results we initially got using uh Gemini 3.5 with the uh this this data set of 11 3.5 with the uh this this data set of 11 3.5 with the uh this this data set of 11 to 13 turns is not huge. Uh there's not to 13 turns is not huge. Uh there's not to 13 turns is not huge. Uh there's not many messages. That's why we wanted we many messages. That's why we wanted we many messages. That's why we wanted we now want uh in the follow-up part do now want uh in the follow-up part do now want uh in the follow-up part do more different experiments that my more different experiments that my more different experiments that my colleagues already will will show you.

  36. colleagues already will will show you. colleagues already will will show you. So yeah, let me in let me introduce you So yeah, let me in let me introduce you So yeah, let me in let me introduce you to something. Yeah. to something. Yeah. to something. Yeah. Um, so this is going to be my part and Um, so this is going to be my part and Um, so this is going to be my part and um, as we ended on the note where Omar um, as we ended on the note where Omar um, as we ended on the note where Omar just said that it costed it c um, it just said that it costed it c um, it just said that it costed it c um, it costed us almost $600 to um, run the costed us almost $600 to um, run the costed us almost $600 to um, run the evals that we ran. One question we were evals that we ran. One question we were evals that we ran. One question we were trying to um, answer with this extended trying to um, answer with this extended trying to um, answer with this extended evaluation was um, so when do you uh, evaluation was um, so when do you uh, evaluation was um, so when do you uh, when does compaction actually matter or when does compaction actually matter or when does compaction actually matter or should you actually compact or not? should you actually compact or not? should you actually compact or not? because we clearly saw that when we have because we clearly saw that when we have because we clearly saw that when we have um you know the full answers in the um you know the full answers in the um you know the full answers in the window it works really well but it costs window it works really well but it costs window it works really well but it costs a lot of money. So um I tried to you a lot of money. So um I tried to you a lot of money. So um I tried to you know do this evaluation in like three know do this evaluation in like three know do this evaluation in like three sorts of context. Uh first one was cash sorts of context. Uh first one was cash sorts of context. Uh first one was cash chats. So you know like when you're chats. So you know like when you're chats. So you know like when you're chatting with Gemini you'd be able to chatting with Gemini you'd be able to chatting with Gemini you'd be able to see that is the cash chat option. The see that is the cash chat option. The see that is the cash chat option. The second version is going to be document second version is going to be document second version is going to be document plus tool. So if you're pasting like a plus tool. So if you're pasting like a plus tool. So if you're pasting like a long document in the AI tutor or if long document in the AI tutor or if long document in the AI tutor or if there's just tool output what happens there's just tool output what happens there's just tool output what happens then and finally you know if you go then and finally you know if you go then and finally you know if you go local or if you scale this evaluation local or if you scale this evaluation local or if you scale this evaluation how well is it going to work out how well is it going to work out how well is it going to work out so um should you ever compact so before so um should you ever compact so before so um should you ever compact so before my part Omar just showed that on Gemini my part Omar just showed that on Gemini my part Omar just showed that on Gemini 3.5 flash um keeping everything one 3.5 flash um keeping everything one 3.5 flash um keeping everything one um but uh you know the why did we come um but uh you know the why did we come um but uh you know the why did we come up with that question was because um you up with that question was because um you up with that question was because um you know full history on a frontier model know full history on a frontier model know full history on a frontier model like Gemini would be very very like Gemini would be very very like Gemini would be very very expensive.

  37. expensive. expensive. Um so um you know with all this extended Um so um you know with all this extended Um so um you know with all this extended experiment we are trying to figure out experiment we are trying to figure out experiment we are trying to figure out was it actually worth it or not. was it actually worth it or not. was it actually worth it or not. But uh before we start that I wanted to But uh before we start that I wanted to But uh before we start that I wanted to just talk about the different context we just talk about the different context we just talk about the different context we see in our AI tutor application. Um so see in our AI tutor application. Um so see in our AI tutor application. Um so first one is going to be like a long first one is going to be like a long first one is going to be like a long chat history uh where you know we have chat history uh where you know we have chat history uh where you know we have um a long chat but all of the details um a long chat but all of the details um a long chat but all of the details what the the students are asking or what the the students are asking or what the the students are asking or they're chatting about they can get they're chatting about they can get they're chatting about they can get buried in. Um the second is going to be buried in. Um the second is going to be buried in. Um the second is going to be a pasted document. So you know we do a pasted document. So you know we do a pasted document. So you know we do have a lot of students who are going to have a lot of students who are going to have a lot of students who are going to just like copy paste a lot of documents just like copy paste a lot of documents just like copy paste a lot of documents there. Um so you know um and also there. Um so you know um and also there. Um so you know um and also because we have like limited uh context because we have like limited uh context because we have like limited uh context window. How does that fit? And then it's window. How does that fit? And then it's window. How does that fit? And then it's going to be you know different tools going to be you know different tools going to be you know different tools that we use internally. Um but this is that we use internally. Um but this is that we use internally. Um but this is just going to be like bunch of logs that just going to be like bunch of logs that just going to be like bunch of logs that the tools have and since you know these the tools have and since you know these the tools have and since you know these are like different context each of them are like different context each of them are like different context each of them need some sort of different fix. So you need some sort of different fix. So you need some sort of different fix. So you know we had to evaluate it for all know we had to evaluate it for all know we had to evaluate it for all different um contexts that we had. different um contexts that we had. different um contexts that we had. So um you know the first obvious thing So um you know the first obvious thing So um you know the first obvious thing uh looking at the cost uh we were like uh looking at the cost uh we were like uh looking at the cost uh we were like okay let's try out a cheaper model and okay let's try out a cheaper model and okay let's try out a cheaper model and see you know um does it do better you see you know um does it do better you see you know um does it do better you know what sort of techniques work on know what sort of techniques work on know what sort of techniques work on that does compaction work on it or not.

  38. that does compaction work on it or not. that does compaction work on it or not. So um deepseek v4 flash was an obvious So um deepseek v4 flash was an obvious So um deepseek v4 flash was an obvious option and we were like we will uh try option and we were like we will uh try option and we were like we will uh try this out on this and see um how well it this out on this and see um how well it this out on this and see um how well it works out and you know you can see there works out and you know you can see there works out and you know you can see there is um you know a drastic cost difference is um you know a drastic cost difference is um you know a drastic cost difference between Gemini and Deepseek here. So between Gemini and Deepseek here. So between Gemini and Deepseek here. So here we can see um you know there's a here we can see um you know there's a here we can see um you know there's a drastic difference between cost when we drastic difference between cost when we drastic difference between cost when we checked the performance on deepseek v4 checked the performance on deepseek v4 checked the performance on deepseek v4 and the main reason was that that we and the main reason was that that we and the main reason was that that we were getting a cash discount. So the were getting a cash discount. So the were getting a cash discount. So the cash discount on GE um deepseek was um cash discount on GE um deepseek was um cash discount on GE um deepseek was um you know 50x as compared to Gemini and you know 50x as compared to Gemini and you know 50x as compared to Gemini and even in this um you know setting we saw even in this um you know setting we saw even in this um you know setting we saw that keeping all um you know all of the that keeping all um you know all of the that keeping all um you know all of the context still one. So you know we were context still one. So you know we were context still one. So you know we were just getting the best performance um in just getting the best performance um in just getting the best performance um in that case for deepseek as well. So, but that case for deepseek as well. So, but that case for deepseek as well. So, but you know the the plus thing from this you know the the plus thing from this you know the the plus thing from this experiment was that we figured out a experiment was that we figured out a experiment was that we figured out a cheaper model but um you know uh now we cheaper model but um you know uh now we cheaper model but um you know uh now we wanted to see that even though we have wanted to see that even though we have wanted to see that even though we have everything uh you know keeping everything uh you know keeping everything uh you know keeping everything uh makes it cheaper but does everything uh makes it cheaper but does everything uh makes it cheaper but does it remember better you know does uh it remember better you know does uh it remember better you know does uh keeping all of the context remember all keeping all of the context remember all keeping all of the context remember all of the details the student might be of the details the student might be of the details the student might be asking us. So um you know this is how I asking us. So um you know this is how I asking us. So um you know this is how I tested out the memory of the model. Um tested out the memory of the model. Um tested out the memory of the model. Um so we have conversation within our so we have conversation within our so we have conversation within our system where you know students are system where you know students are system where you know students are asking questions like you know about asking questions like you know about asking questions like you know about their setup about the errors that their setup about the errors that their setup about the errors that they're seeing what whatever they've they're seeing what whatever they've they're seeing what whatever they've already tried. So I just took these already tried. So I just took these already tried. So I just took these questions um so I just took these chats questions um so I just took these chats questions um so I just took these chats and you know asked questions about and you know asked questions about and you know asked questions about specific details just to see if the specific details just to see if the specific details just to see if the model is able to like figure that out if

  39. model is able to like figure that out if model is able to like figure that out if it is able to give me an output for that it is able to give me an output for that it is able to give me an output for that or not. And you know the results I saw or not. And you know the results I saw or not. And you know the results I saw is that 95% of the time if the model was is that 95% of the time if the model was is that 95% of the time if the model was able to uh you know give us the right able to uh you know give us the right able to uh you know give us the right exact details that I was trying to look exact details that I was trying to look exact details that I was trying to look for and even when it is keeping all of for and even when it is keeping all of for and even when it is keeping all of the details whereas if I summarize first the details whereas if I summarize first the details whereas if I summarize first or if I compact um the context I had it or if I compact um the context I had it or if I compact um the context I had it only gave me the answer back 32% of the only gave me the answer back 32% of the only gave me the answer back 32% of the time. And um if you think about it um time. And um if you think about it um time. And um if you think about it um the reason for that is that when you the reason for that is that when you the reason for that is that when you summarize you you know you uh remove all summarize you you know you uh remove all summarize you you know you uh remove all of the necessary details. So you're not of the necessary details. So you're not of the necessary details. So you're not able to uh you know keep all of those able to uh you know keep all of those able to uh you know keep all of those details and the model keeps on missing details and the model keeps on missing details and the model keeps on missing those out. So uh you know we are able to those out. So uh you know we are able to those out. So uh you know we are able to see that keeping everything wins in see that keeping everything wins in see that keeping everything wins in terms of cost. If you have like a you terms of cost. If you have like a you terms of cost. If you have like a you know model like deepseeek um then it know model like deepseeek um then it know model like deepseeek um then it also is remembering things. So you know also is remembering things. So you know also is remembering things. So you know it's not that if you have a long it's not that if you have a long it's not that if you have a long conversation it is not able to remember conversation it is not able to remember conversation it is not able to remember things. So you know it is correct 95% of things. So you know it is correct 95% of things. So you know it is correct 95% of the time. So the next thing we wanted to the time. So the next thing we wanted to the time. So the next thing we wanted to see is the cost part of it. Uh you know see is the cost part of it. Uh you know see is the cost part of it. Uh you know how does it actually um does it cost how does it actually um does it cost how does it actually um does it cost most? Does it cost the lease? What most? Does it cost the lease? What most? Does it cost the lease? What happens in terms of tokens? So, um you happens in terms of tokens? So, um you happens in terms of tokens? So, um you know on Deep Seek we saw um the setup know on Deep Seek we saw um the setup know on Deep Seek we saw um the setup that was sending the most tokens is that was sending the most tokens is that was sending the most tokens is actually the cheapest to run. So the actually the cheapest to run. So the actually the cheapest to run. So the full history u you know setup that we full history u you know setup that we full history u you know setup that we had was sending the most tokens but uh had was sending the most tokens but uh had was sending the most tokens but uh we were still getting the best results we were still getting the best results we were still getting the best results out of it because 97% of the tokens that out of it because 97% of the tokens that out of it because 97% of the tokens that we had were cached and uh you know as we had were cached and uh you know as we had were cached and uh you know as you as I as we just saw in the previous you as I as we just saw in the previous you as I as we just saw in the previous slide cash tokens are really cheap. they slide cash tokens are really cheap. they slide cash tokens are really cheap. they are uh you know they are charged

  40. are uh you know they are charged are uh you know they are charged separately. So um summarizing uh works separately. So um summarizing uh works separately. So um summarizing uh works the other way uh and every turn it makes the other way uh and every turn it makes the other way uh and every turn it makes the model reads and writes uh new tokens the model reads and writes uh new tokens the model reads and writes uh new tokens and you know this is something that I and you know this is something that I and you know this is something that I ran on like 36 turn conversation and it ran on like 36 turn conversation and it ran on like 36 turn conversation and it was about 1.78 million tokens and was about 1.78 million tokens and was about 1.78 million tokens and keeping everything still came out ahead keeping everything still came out ahead keeping everything still came out ahead and um so you know this answers the and um so you know this answers the and um so you know this answers the question that it wasn't just u you know question that it wasn't just u you know question that it wasn't just u you know Gemini it it works same on deepseek as Gemini it it works same on deepseek as Gemini it it works same on deepseek as well So um the cheaper model is u what well So um the cheaper model is u what well So um the cheaper model is u what brought the cost down but the results brought the cost down but the results brought the cost down but the results are the same and um you know in in are the same and um you know in in are the same and um you know in in respect of this I wanted the next thing respect of this I wanted the next thing respect of this I wanted the next thing I wanted to ask is what happens when the I wanted to ask is what happens when the I wanted to ask is what happens when the conversations really grow because all of conversations really grow because all of conversations really grow because all of these were tested out on like short these were tested out on like short these were tested out on like short conversation if I really really uh you conversation if I really really uh you conversation if I really really uh you know increase the length of the know increase the length of the know increase the length of the conversation how does that work out so conversation how does that work out so conversation how does that work out so um you know when I tried to you know run um you know when I tried to you know run um you know when I tried to you know run this entire experiment on a longer this entire experiment on a longer this entire experiment on a longer context I saw saw that um you know if context I saw saw that um you know if context I saw saw that um you know if even if I'm pulling like one specific even if I'm pulling like one specific even if I'm pulling like one specific detail out of it um the model performed detail out of it um the model performed detail out of it um the model performed really well. So on the top part the really well. So on the top part the really well. So on the top part the green line those are all of the green line those are all of the green line those are all of the distinctive distinctive facts that the distinctive distinctive facts that the distinctive distinctive facts that the model is able to find out. So we can see model is able to find out. So we can see model is able to find out. So we can see that up until 800k tokens as well the that up until 800k tokens as well the that up until 800k tokens as well the model was not missing out on those model was not missing out on those model was not missing out on those facts. It was giving me like good and facts. It was giving me like good and facts. It was giving me like good and consistent results for uh you know some consistent results for uh you know some consistent results for uh you know some ambiguous facts. um the performance ambiguous facts. um the performance ambiguous facts. um the performance dropped um you know to half of what I dropped um you know to half of what I dropped um you know to half of what I observed on the distinct distinctive

  41. observed on the distinct distinctive observed on the distinct distinctive facts but for overall our for our AI facts but for overall our for our AI facts but for overall our for our AI tutor this uh you know result um was tutor this uh you know result um was tutor this uh you know result um was really good. So you know we saw that um really good. So you know we saw that um really good. So you know we saw that um the model even when you're not the model even when you're not the model even when you're not compacting anything holds really well compacting anything holds really well compacting anything holds really well even if you have a really long even if you have a really long even if you have a really long conversation conversation conversation so uh but so far everything that I've so uh but so far everything that I've so uh but so far everything that I've talked about is only per turn. So you talked about is only per turn. So you talked about is only per turn. So you know every cost has been per turn but know every cost has been per turn but know every cost has been per turn but does it help when you know we scale it does it help when you know we scale it does it help when you know we scale it because um a chatbot is not something um because um a chatbot is not something um because um a chatbot is not something um that you know it's not like one turn uh that you know it's not like one turn uh that you know it's not like one turn uh situation for a chatbot. Um the tutor is situation for a chatbot. Um the tutor is situation for a chatbot. Um the tutor is a long-term service where you know a long-term service where you know a long-term service where you know students are asking um you know students are asking um you know students are asking um you know questions on on a massive scale. So questions on on a massive scale. So questions on on a massive scale. So let's say we have uh you know if I I let's say we have uh you know if I I let's say we have uh you know if I I calculated this using this like if we calculated this using this like if we calculated this using this like if we have 100,000 to a million turns of have 100,000 to a million turns of have 100,000 to a million turns of questions every day what would you know questions every day what would you know questions every day what would you know be the cost on deepseek or so you know be the cost on deepseek or so you know be the cost on deepseek or so you know for deepseek the cost was uh for deepseek the cost was uh for deepseek the cost was uh approximately somewhere from 18,000 to approximately somewhere from 18,000 to approximately somewhere from 18,000 to like 180,000 a month and uh you know like 180,000 a month and uh you know like 180,000 a month and uh you know even though we don't see that sort of even though we don't see that sort of even though we don't see that sort of volume as of now but paying per token volume as of now but paying per token volume as of now but paying per token starts to add up and so the the starts to add up and so the the starts to add up and so the the alternative for that was going over to a alternative for that was going over to a alternative for that was going over to a local model just to see if we are local model just to see if we are local model just to see if we are getting the same sort of performance on getting the same sort of performance on getting the same sort of performance on local model or not. And um so one way um local model or not. And um so one way um local model or not. And um so one way um you know we were thinking that because you know we were thinking that because you know we were thinking that because local models cache as well can you know local models cache as well can you know local models cache as well can you know we use the same sort of setup on a local we use the same sort of setup on a local we use the same sort of setup on a local model and will we get the same result.

  42. model and will we get the same result. model and will we get the same result. So uh you know because uh we had like So uh you know because uh we had like So uh you know because uh we had like hardware limitations we just tested it hardware limitations we just tested it hardware limitations we just tested it out on a MacBook. the maximum context out on a MacBook. the maximum context out on a MacBook. the maximum context window that we could go up to was 32K window that we could go up to was 32K window that we could go up to was 32K and uh so we thought that can we uh you and uh so we thought that can we uh you and uh so we thought that can we uh you know can we do that locally now uh but know can we do that locally now uh but know can we do that locally now uh but we can't because uh you know the the we can't because uh you know the the we can't because uh you know the the sort of uh lessons that we had are had sort of uh lessons that we had are had sort of uh lessons that we had are had are bigger than like 32k context window are bigger than like 32k context window are bigger than like 32k context window on their own and once the conversation on their own and once the conversation on their own and once the conversation doesn't fit in the window um caching was doesn't fit in the window um caching was doesn't fit in the window um caching was no longer helpful for us and we uh you no longer helpful for us and we uh you no longer helpful for us and we uh you know we have to make um the context know we have to make um the context know we have to make um the context smaller either by compressing it or by smaller either by compressing it or by smaller either by compressing it or by retrieving only the parts that we need. retrieving only the parts that we need. retrieving only the parts that we need. So uh you know the next question we were So uh you know the next question we were So uh you know the next question we were trying to answer is can we uh you know trying to answer is can we uh you know trying to answer is can we uh you know once you have to compact locally what once you have to compact locally what once you have to compact locally what actually works? So uh you know for um actually works? So uh you know for um actually works? So uh you know for um the chat memory going local uh and the chat memory going local uh and the chat memory going local uh and trying to keep everything stops winning trying to keep everything stops winning trying to keep everything stops winning because you can't keep everything and because you can't keep everything and because you can't keep everything and you know a simple uh question here would you know a simple uh question here would you know a simple uh question here would be that why can't you just keep be that why can't you just keep be that why can't you just keep increasing the length of the model you increasing the length of the model you increasing the length of the model you know why can't you use a bigger model know why can't you use a bigger model know why can't you use a bigger model because of course we have hardware because of course we have hardware because of course we have hardware limitations for us it was uh you know a limitations for us it was uh you know a limitations for us it was uh you know a MacBook but you know GPUs also have like MacBook but you know GPUs also have like MacBook but you know GPUs also have like hardware limitations but you know we hardware limitations but you know we hardware limitations but you know we went from like a 7B model 8B model to a went from like a 7B model 8B model to a went from like a 7B model 8B model to a 32-B model uh but here we landed Ed on 32-B model uh but here we landed Ed on 32-B model uh but here we landed Ed on the uh you know conclusion that even the uh you know conclusion that even the uh you know conclusion that even though you keep increasing the length of though you keep increasing the length of though you keep increasing the length of the model uh it is not going to increase the model uh it is not going to increase the model uh it is not going to increase your context window you know you cannot your context window you know you cannot your context window you know you cannot repair that part you have to you know

  43. repair that part you have to you know repair that part you have to you know make a choice there make a choice there make a choice there uh but um you know this is only for the uh but um you know this is only for the uh but um you know this is only for the chat history and what happens when you chat history and what happens when you chat history and what happens when you are trying to deal with like documents are trying to deal with like documents are trying to deal with like documents locally so um this was a little locally so um this was a little locally so um this was a little surprising because if you are retrieving surprising because if you are retrieving surprising because if you are retrieving results with like local documents um you results with like local documents um you results with like local documents um you know it was really good we got like 100% know it was really good we got like 100% know it was really good we got like 100% in accuracy in that case. So if uh you in accuracy in that case. So if uh you in accuracy in that case. So if uh you know when students are pasting something know when students are pasting something know when students are pasting something which is too big uh you know rag is a which is too big uh you know rag is a which is too big uh you know rag is a good option there you can use that and good option there you can use that and good option there you can use that and it can help you retrieve the exact uh it can help you retrieve the exact uh it can help you retrieve the exact uh data that you're looking for. Also the data that you're looking for. Also the data that you're looking for. Also the processing time in this case was processing time in this case was processing time in this case was anywhere from 25 to 65 seconds. So you anywhere from 25 to 65 seconds. So you anywhere from 25 to 65 seconds. So you know which is uh pretty good in terms of know which is uh pretty good in terms of know which is uh pretty good in terms of the you know the output that we're the you know the output that we're the you know the output that we're getting. Uh but you know if you are getting. Uh but you know if you are getting. Uh but you know if you are trying to like stuff the window with trying to like stuff the window with trying to like stuff the window with like more context than you have. We saw like more context than you have. We saw like more context than you have. We saw that you know it took us like that you know it took us like that you know it took us like approximately 340 seconds um to get the approximately 340 seconds um to get the approximately 340 seconds um to get the output when you know our um output when you know our um output when you know our um conversations were really long and the conversations were really long and the conversations were really long and the output that we got was a single token. output that we got was a single token. output that we got was a single token. So you know you are um not getting So you know you are um not getting So you know you are um not getting anything but you're also wasting a lot anything but you're also wasting a lot anything but you're also wasting a lot of time when you are trying to do it. So of time when you are trying to do it. So of time when you are trying to do it. So um you know you have to be careful about um you know you have to be careful about um you know you have to be careful about what option you choose in this case.

  44. what option you choose in this case. what option you choose in this case. So um so you know you also have like So um so you know you also have like So um so you know you also have like different type of retrieval strategies different type of retrieval strategies different type of retrieval strategies that you could use. The default that you could use. The default that you could use. The default retrieval is of course semantic search retrieval is of course semantic search retrieval is of course semantic search where you're just trying to match uh you where you're just trying to match uh you where you're just trying to match uh you know the meaning of the text and that is know the meaning of the text and that is know the meaning of the text and that is the dense rag uh heading that you can the dense rag uh heading that you can the dense rag uh heading that you can see on the chart. Um it mostly works but see on the chart. Um it mostly works but see on the chart. Um it mostly works but uh you know we saw we tried to make it uh you know we saw we tried to make it uh you know we saw we tried to make it work from like uh 50k token to 200k and work from like uh 50k token to 200k and work from like uh 50k token to 200k and we saw that you know dense um rag worked we saw that you know dense um rag worked we saw that you know dense um rag worked really well. you know it was like 80%. really well. you know it was like 80%. really well. you know it was like 80%. But when we increased it to like 400k But when we increased it to like 400k But when we increased it to like 400k tokens it was not able to facts that tokens it was not able to facts that tokens it was not able to facts that were buried in the middle and it started were buried in the middle and it started were buried in the middle and it started giving us like 0% recall whereas uh you giving us like 0% recall whereas uh you giving us like 0% recall whereas uh you know something like BM25 it still got know something like BM25 it still got know something like BM25 it still got 100% every time. So semantic search uh 100% every time. So semantic search uh 100% every time. So semantic search uh on its own is not enough and that's why on its own is not enough and that's why on its own is not enough and that's why you know when Omar talked about our you know when Omar talked about our you know when Omar talked about our setup in the AI tutor we're actually setup in the AI tutor we're actually setup in the AI tutor we're actually using a hybrid search we're using a mix using a hybrid search we're using a mix using a hybrid search we're using a mix of both u you know dense and um you know of both u you know dense and um you know of both u you know dense and um you know BM25 we're using a combination of both BM25 we're using a combination of both BM25 we're using a combination of both those um so you know after all of this those um so you know after all of this those um so you know after all of this we came up we had like one other we came up we had like one other we came up we had like one other question which was like uh how does all question which was like uh how does all question which was like uh how does all of this um local setup compare to cloud of this um local setup compare to cloud of this um local setup compare to cloud because that's the uh real way we'll see because that's the uh real way we'll see because that's the uh real way we'll see the result. So we wanted to put them the result. So we wanted to put them the result. So we wanted to put them side by side just to see uh what is the side by side just to see uh what is the side by side just to see uh what is the output and for chat uh you know the output and for chat uh you know the output and for chat uh you know the local setup was not up to the par of um local setup was not up to the par of um local setup was not up to the par of um the cloud you know the cloud setup on the cloud you know the cloud setup on the cloud you know the cloud setup on cloud keeping everything scores you know cloud keeping everything scores you know cloud keeping everything scores you know somewhere from from 92 to 95%. But

  45. somewhere from from 92 to 95%. But somewhere from from 92 to 95%. But locally it was stuck at like 33% and the locally it was stuck at like 33% and the locally it was stuck at like 33% and the context window was a limitation here. context window was a limitation here. context window was a limitation here. Also you can see that um you know local Also you can see that um you know local Also you can see that um you know local um local models actually work because um um local models actually work because um um local models actually work because um there is no cost like you cannot see there is no cost like you cannot see there is no cost like you cannot see because you already own the hardware because you already own the hardware because you already own the hardware though um you know there's a throughput though um you know there's a throughput though um you know there's a throughput limitation there. Uh but if you use a limitation there. Uh but if you use a limitation there. Uh but if you use a technique like retrieval uh you know you technique like retrieval uh you know you technique like retrieval uh you know you get like good accuracy even on a local get like good accuracy even on a local get like good accuracy even on a local setup. setup. setup. So um in our case what we found is that So um in our case what we found is that So um in our case what we found is that on memory uh you know keeping the whole on memory uh you know keeping the whole on memory uh you know keeping the whole chat recalled about like 95% of the chat recalled about like 95% of the chat recalled about like 95% of the details that we were providing it it was details that we were providing it it was details that we were providing it it was able to you know give us correct answer able to you know give us correct answer able to you know give us correct answer 95% of the time versus it was uh like 95% of the time versus it was uh like 95% of the time versus it was uh like just 32% if you summarize it um on long just 32% if you summarize it um on long just 32% if you summarize it um on long context uh you know finding a single context uh you know finding a single context uh you know finding a single fact uh is easy for the model. we went fact uh is easy for the model. we went fact uh is easy for the model. we went up to 800k tokens and you know we did up to 800k tokens and you know we did up to 800k tokens and you know we did not um see any sort of context rot in not um see any sort of context rot in not um see any sort of context rot in that case. Um on cost per turn we saw that case. Um on cost per turn we saw that case. Um on cost per turn we saw that the cheapest run is actually the that the cheapest run is actually the that the cheapest run is actually the one which is sending the most tokens one which is sending the most tokens one which is sending the most tokens because caching makes um resending the because caching makes um resending the because caching makes um resending the same context uh you know very cheap and same context uh you know very cheap and same context uh you know very cheap and on no cost at scale um um you know it on no cost at scale um um you know it on no cost at scale um um you know it scales up for like let's say if we have scales up for like let's say if we have scales up for like let's say if we have like thousand students uh you know like thousand students uh you know like thousand students uh you know Gemini costs us about like $40,000 Gemini costs us about like $40,000 Gemini costs us about like $40,000 40,000 a month whereas Deep Seek Deep 40,000 a month whereas Deep Seek Deep 40,000 a month whereas Deep Seek Deep Seek was around 1,900 a month um so um

  46. Seek was around 1,900 a month um so um Seek was around 1,900 a month um so um going local saves us on a cost a bit going local saves us on a cost a bit going local saves us on a cost a bit more. So the main thing to take away is more. So the main thing to take away is more. So the main thing to take away is that um do not compact by default. You that um do not compact by default. You that um do not compact by default. You have to name the constraint that you have to name the constraint that you have to name the constraint that you have and then you know look for a better have and then you know look for a better have and then you know look for a better alternative. alternative. alternative. So what did we finally decide after all So what did we finally decide after all So what did we finally decide after all of these different experiments that we of these different experiments that we of these different experiments that we ran? Um so we decided on deep uh V4 ran? Um so we decided on deep uh V4 ran? Um so we decided on deep uh V4 flash because uh we had hardware flash because uh we had hardware flash because uh we had hardware limitations. So for us um the cloud limitations. So for us um the cloud limitations. So for us um the cloud structure worked out really well. It is structure worked out really well. It is structure worked out really well. It is also the cheapest considering the also the cheapest considering the also the cheapest considering the current um intake of students we have. current um intake of students we have. current um intake of students we have. So you know it works out well for us. Um So you know it works out well for us. Um So you know it works out well for us. Um and we're using on top of that model we and we're using on top of that model we and we're using on top of that model we are using a mix of u you know using are using a mix of u you know using are using a mix of u you know using hybrid retrieval to uh you know get good hybrid retrieval to uh you know get good hybrid retrieval to uh you know get good results. For memory uh you know we have results. For memory uh you know we have results. For memory uh you know we have chosen to keep everything. uh we have chosen to keep everything. uh we have chosen to keep everything. uh we have got um you know a default limit that got um you know a default limit that got um you know a default limit that after 30k tokens we are going to you after 30k tokens we are going to you after 30k tokens we are going to you know have compaction but up until that know have compaction but up until that know have compaction but up until that we're planning to like keep everything we're planning to like keep everything we're planning to like keep everything and uh you know that is the tutor setup and uh you know that is the tutor setup and uh you know that is the tutor setup that we have that we have that we have so you know because we had like time so you know because we had like time so you know because we had like time limitations so these are all the limitations so these are all the limitations so these are all the evaluations and experiments that I could evaluations and experiments that I could evaluations and experiments that I could pack into this time but if you would pack into this time but if you would pack into this time but if you would like to you know learn more about these like to you know learn more about these like to you know learn more about these evaluations or you would want to build a evaluations or you would want to build a evaluations or you would want to build a tutor yourself um this is the fullstack tutor yourself um this is the fullstack tutor yourself um this is the fullstack AI engineering course uh on AI engineering course uh on AI engineering course uh on academy.towardsai.net.

  47. academy.towardsai.net. academy.towardsai.net. Um so you can go to this link and uh you Um so you can go to this link and uh you Um so you can go to this link and uh you know go through the course. Uh but thank know go through the course. Uh but thank know go through the course. Uh but thank you so much everyone. Uh and now we can you so much everyone. Uh and now we can you so much everyone. Uh and now we can take any questions.

Summary

This transcript discusses context engineering for AI agents in 2026, specifically addressing the problem of agents performing undesirable actions due to context overload. The takeaway is to experiment with solutions like compaction and memory retrieval to improve AI tutor performance and reduce costs, with all findings and the AI tutor itself being open-source.

View original episode ↗