← Back
AI Engineer July 31, 2026 16m

Emulated: The Data for Fully Autonomous Software Engineers and Companies — Joseph Wang

Read full transcript 14 segments
  1. So, I appreciate the intro. Uh my name So, I appreciate the intro. Uh my name is Joseph and this is my co-founder Sid. is Joseph and this is my co-founder Sid. is Joseph and this is my co-founder Sid. Emulated is a data lab focused on Emulated is a data lab focused on Emulated is a data lab focused on increasing the reliability and autonomy increasing the reliability and autonomy increasing the reliability and autonomy of AI agents. of AI agents. of AI agents. And if you've been a AI engineer and And if you've been a AI engineer and And if you've been a AI engineer and you've watched the talks, seen the you've watched the talks, seen the you've watched the talks, seen the tracks, then there's probably one tracks, then there's probably one tracks, then there's probably one takeaway that all the talks have in takeaway that all the talks have in takeaway that all the talks have in common. And it's that we're headed common. And it's that we're headed common. And it's that we're headed towards a future where agents are able towards a future where agents are able towards a future where agents are able to perform useful work over longer and to perform useful work over longer and to perform useful work over longer and longer horizons with little to no longer horizons with little to no longer horizons with little to no supervision. supervision. supervision. So, today we're going to answer some of So, today we're going to answer some of So, today we're going to answer some of the questions of what this means for the the questions of what this means for the the questions of what this means for the data and model layers. Uh we're going to data and model layers. Uh we're going to data and model layers. Uh we're going to touch on some pretty cool things. Uh so, touch on some pretty cool things. Uh so, touch on some pretty cool things. Uh so, look out for them. Um like how to look out for them. Um like how to look out for them. Um like how to simulate a company within a sandbox or simulate a company within a sandbox or simulate a company within a sandbox or sandboxes for multi-node systems and sandboxes for multi-node systems and sandboxes for multi-node systems and distributed clusters. distributed clusters. distributed clusters. Um and if we have a little bit of time, Um and if we have a little bit of time, Um and if we have a little bit of time, we'll also go into some of the work that we'll also go into some of the work that we'll also go into some of the work that we're doing with post-training pipelines we're doing with post-training pipelines we're doing with post-training pipelines and how these new types of sandboxes are and how these new types of sandboxes are and how these new types of sandboxes are affecting post-training infra as well.

  2. So, where Sid and I come from, um our So, where Sid and I come from, um our backgrounds are in network infra, backgrounds are in network infra, backgrounds are in network infra, distributed databases, and sandbox distributed databases, and sandbox distributed databases, and sandbox infra. And these are all areas where the infra. And these are all areas where the infra. And these are all areas where the workloads are mission critical. workloads are mission critical. workloads are mission critical. We all saw a couple months ago that uh We all saw a couple months ago that uh We all saw a couple months ago that uh when something like DynamoDB goes down, when something like DynamoDB goes down, when something like DynamoDB goes down, so does US East 1 and half the internet. so does US East 1 and half the internet. so does US East 1 and half the internet. Um and working on these systems, we saw Um and working on these systems, we saw Um and working on these systems, we saw model capability gap when it came to model capability gap when it came to model capability gap when it came to operating and building these systems at operating and building these systems at operating and building these systems at scale at scale and thinking about uh the scale at scale and thinking about uh the scale at scale and thinking about uh the consequences of architecture and system consequences of architecture and system consequences of architecture and system system design over the course of years. system design over the course of years. system design over the course of years. >> Yeah, so it led to a pretty uh natural >> Yeah, so it led to a pretty uh natural >> Yeah, so it led to a pretty uh natural question, right? For such question, right? For such question, right? For such mission-critical services, why is it mission-critical services, why is it mission-critical services, why is it that my model or my agent is so that my model or my agent is so that my model or my agent is so proficient at handling the application proficient at handling the application proficient at handling the application layer, but is struggles when it comes to layer, but is struggles when it comes to layer, but is struggles when it comes to reasoning through infrastructure reasoning through infrastructure reasoning through infrastructure complexities? For example, things like complexities? For example, things like complexities? For example, things like MVCC on a database engine, which can MVCC on a database engine, which can MVCC on a database engine, which can lead to corruption issues, which is one lead to corruption issues, which is one lead to corruption issues, which is one of which was one of the roots of the of which was one of the roots of the of which was one of the roots of the DynamoDB failure a few months ago.

  3. >> So, like with everything in NML, uh the >> So, like with everything in NML, uh the gap in models is usually a gap in data. gap in models is usually a gap in data. gap in models is usually a gap in data. Models typically are only as good at as Models typically are only as good at as Models typically are only as good at as data is. data is. data is. Um and to really highlight this point, Um and to really highlight this point, Um and to really highlight this point, right? Model capability has never uh right? Model capability has never uh right? Model capability has never uh regressed whenever you introduce more regressed whenever you introduce more regressed whenever you introduce more high-quality data. high-quality data. high-quality data. Um so, with that being said, what is the Um so, with that being said, what is the Um so, with that being said, what is the data gap then? What does data look like data gap then? What does data look like data gap then? What does data look like right now? And how is this influencing right now? And how is this influencing right now? And how is this influencing the model capability gap here? the model capability gap here? the model capability gap here? So, if you look at any of the frontier So, if you look at any of the frontier So, if you look at any of the frontier or recent benchmarks, like SweBench Pro, or recent benchmarks, like SweBench Pro, or recent benchmarks, like SweBench Pro, Terminal Bench, or something like Terminal Bench, or something like Terminal Bench, or something like Frontier Code and Deep Sweep, Frontier Code and Deep Sweep, Frontier Code and Deep Sweep, um the tasks only operate within the um the tasks only operate within the um the tasks only operate within the code base. Uh code base. Uh code base. Uh the agent is given a the agent is given a the agent is given a pretty large uh task uh and over the pretty large uh task uh and over the pretty large uh task uh and over the course of 50 to 100 turns produces a course of 50 to 100 turns produces a course of 50 to 100 turns produces a couple thousand-line PR. couple thousand-line PR. couple thousand-line PR. Um but it doesn't do all of the work Um but it doesn't do all of the work Um but it doesn't do all of the work that a human does. It doesn't do uh what that a human does. It doesn't do uh what that a human does. It doesn't do uh what a PM does with talking to customers, a PM does with talking to customers, a PM does with talking to customers, understanding their problems, what an understanding their problems, what an understanding their problems, what an engineer does with trying out different engineer does with trying out different engineer does with trying out different approaches, performing performance approaches, performing performance approaches, performing performance testing them, um and owning the testing them, um and owning the testing them, um and owning the underlying infra underlying infra underlying infra for the code base over the course of not for the code base over the course of not for the code base over the course of not just months, but years.

  4. just months, but years. just months, but years. >> And this is really the gap that we're >> And this is really the gap that we're >> And this is really the gap that we're closing. We've taken software closing. We've taken software closing. We've taken software engineering companies and we've put them engineering companies and we've put them engineering companies and we've put them into containerized environments. so this into containerized environments. so this into containerized environments. so this includes uh include like organizational includes uh include like organizational includes uh include like organizational contexts like projects, incidents, contexts like projects, incidents, contexts like projects, incidents, customer conversations. Uh the agent customer conversations. Uh the agent customer conversations. Uh the agent also has to deal with issues that only also has to deal with issues that only also has to deal with issues that only appear at scale like network failures appear at scale like network failures appear at scale like network failures between distributed nodes, data between distributed nodes, data between distributed nodes, data corruption, and clock skew. corruption, and clock skew. corruption, and clock skew. And through all this, we also want the And through all this, we also want the And through all this, we also want the agents to reason about orchestrating agents to reason about orchestrating agents to reason about orchestrating through distributed clusters and also through distributed clusters and also through distributed clusters and also thinking about things like operational thinking about things like operational thinking about things like operational blast radius while solving live traffic. blast radius while solving live traffic. blast radius while solving live traffic. And the result is that the task that And the result is that the task that And the result is that the task that these agents have to complete or we want these agents have to complete or we want these agents have to complete or we want the agents to learn is that environments the agents to learn is that environments the agents to learn is that environments are far more complex and long horizon are far more complex and long horizon are far more complex and long horizon than a simple code diff. So, let's just let's bring a picture So, let's just let's bring a picture into the mix because it tends to make into the mix because it tends to make into the mix because it tends to make things more interesting. Uh here's an things more interesting. Uh here's an things more interesting. Uh here's an example we've built of an SCD consensus example we've built of an SCD consensus example we've built of an SCD consensus cluster that a typical production cluster that a typical production cluster that a typical production service might rely on. So, an old service might rely on. So, an old service might rely on. So, an old environment uh might environment uh might environment uh might tended to operate and work primarily on tended to operate and work primarily on tended to operate and work primarily on that little blue square entitled SCD that little blue square entitled SCD that little blue square entitled SCD source code in the bottom right there.

  5. source code in the bottom right there. source code in the bottom right there. But, a lot of the fun and the model But, a lot of the fun and the model But, a lot of the fun and the model capability gap that results from it is capability gap that results from it is capability gap that results from it is really in everything that surrounds it. really in everything that surrounds it. really in everything that surrounds it. So, you you start with the tickets, So, you you start with the tickets, So, you you start with the tickets, projects, postmortems. What are the projects, postmortems. What are the projects, postmortems. What are the train wrecks? Why did they happen? How train wrecks? Why did they happen? How train wrecks? Why did they happen? How did customers feel about them? And often did customers feel about them? And often did customers feel about them? And often times those aren't necessarily up to times those aren't necessarily up to times those aren't necessarily up to date. date. date. Um the agent has to incorporate all that Um the agent has to incorporate all that Um the agent has to incorporate all that when it's reasoning through the actual when it's reasoning through the actual when it's reasoning through the actual change that change that change that current environments have it make. After current environments have it make. After current environments have it make. After it makes that change, uh you need a kick it makes that change, uh you need a kick it makes that change, uh you need a kick kick off rolling deployments. Those kick off rolling deployments. Those kick off rolling deployments. Those deployment systems can often times be deployment systems can often times be deployment systems can often times be complicated, have conflicts, may not complicated, have conflicts, may not complicated, have conflicts, may not work. work. work. Um and all through that, when you're Um and all through that, when you're Um and all through that, when you're finally migrating off of from old hard finally migrating off of from old hard finally migrating off of from old hard onto new hardware, onto new hardware, onto new hardware, um you run into unforeseen problems um you run into unforeseen problems um you run into unforeseen problems which you which you which you did not sort of the the the the the the did not sort of the the the the the the did not sort of the the the the the the agent has to reason through in real agent has to reason through in real agent has to reason through in real time, just like a human would, right? time, just like a human would, right? time, just like a human would, right? You have um You have um You have um failing nodes. You have stale deprecated failing nodes. You have stale deprecated failing nodes. You have stale deprecated nodes.

  6. nodes. nodes. And while all of this is happening, the And while all of this is happening, the And while all of this is happening, the service can't go down because there is a service can't go down because there is a service can't go down because there is a blast radius to serving live traffic. blast radius to serving live traffic. blast radius to serving live traffic. You have to observe and monitor your You have to observe and monitor your You have to observe and monitor your service. All of these components in in service. All of these components in in service. All of these components in in in the system is really uh what sort of in the system is really uh what sort of in the system is really uh what sort of exemplifies like a full end-to-end exemplifies like a full end-to-end exemplifies like a full end-to-end infrastructure task. infrastructure task. infrastructure task. >> So, what Sid is describing here is an >> So, what Sid is describing here is an >> So, what Sid is describing here is an environment in a single node sandbox environment in a single node sandbox environment in a single node sandbox where we're simulating uh distributed where we're simulating uh distributed where we're simulating uh distributed cluster with multiple nodes, flapping cluster with multiple nodes, flapping cluster with multiple nodes, flapping nodes, lagging learners um in a single nodes, lagging learners um in a single nodes, lagging learners um in a single sandbox. And you can get pretty far with sandbox. And you can get pretty far with sandbox. And you can get pretty far with this, right? Like you can see that this, right? Like you can see that this, right? Like you can see that there's live traffic, there's a lot of there's live traffic, there's a lot of there's live traffic, there's a lot of operational issues that a real engineer operational issues that a real engineer operational issues that a real engineer would have to deal with. And you can would have to deal with. And you can would have to deal with. And you can make this pretty long horizon by just make this pretty long horizon by just make this pretty long horizon by just say say say doing multiple deployments instead of doing multiple deployments instead of doing multiple deployments instead of just one. just one. just one. But really what we're seeing is that But really what we're seeing is that But really what we're seeing is that this is not enough. Uh this fits into this is not enough. Uh this fits into this is not enough. Uh this fits into standard post-training pipelines in the standard post-training pipelines in the standard post-training pipelines in the sense that a standard post-training sense that a standard post-training sense that a standard post-training pipeline is kind of boring. Uh it's kind pipeline is kind of boring. Uh it's kind pipeline is kind of boring. Uh it's kind of homogeneous. You know, everything of homogeneous. You know, everything of homogeneous. You know, everything just runs harbor, everything is a single just runs harbor, everything is a single just runs harbor, everything is a single sandbox, containerized. But real sandbox, containerized. But real sandbox, containerized. But real infrastructure infrastructure infrastructure uh doesn't work like this. Uh this isn't uh doesn't work like this. Uh this isn't uh doesn't work like this. Uh this isn't how real companies run. And how real companies run. And how real companies run. And uh even though you can use something uh even though you can use something uh even though you can use something like deterministic simulation to like deterministic simulation to like deterministic simulation to simulate network failures, it doesn't simulate network failures, it doesn't simulate network failures, it doesn't represent what you might run into if represent what you might run into if represent what you might run into if you're building an AWS-scale service.

  7. you're building an AWS-scale service. you're building an AWS-scale service. So, I did see I think a couple people at So, I did see I think a couple people at So, I did see I think a couple people at AWS. Somebody had Viceroy open on their AWS. Somebody had Viceroy open on their AWS. Somebody had Viceroy open on their laptop. Um fun times. Um but let's laptop. Um fun times. Um but let's laptop. Um fun times. Um but let's imagine here that we are all AWS imagine here that we are all AWS imagine here that we are all AWS engineers or GCP engineers. Azure, too. engineers or GCP engineers. Azure, too. engineers or GCP engineers. Azure, too. No shade, right? Um, and we are building No shade, right? Um, and we are building No shade, right? Um, and we are building a cloud service. Um, it can also be some a cloud service. Um, it can also be some a cloud service. Um, it can also be some infrastructure service like DataDog, infrastructure service like DataDog, infrastructure service like DataDog, Vercel, Superbase. Vercel, Superbase. Vercel, Superbase. Uh, all of these services run into the Uh, all of these services run into the Uh, all of these services run into the same problems. You start off with a same problems. You start off with a same problems. You start off with a shiny piece of software. And this piece shiny piece of software. And this piece shiny piece of software. And this piece of software can service a single of software can service a single of software can service a single customer pretty well. Um, maybe it's customer pretty well. Um, maybe it's customer pretty well. Um, maybe it's running on your machine. If you're running on your machine. If you're running on your machine. If you're working for NLB, it should be a load working for NLB, it should be a load working for NLB, it should be a load balancer, right? If you're working for balancer, right? If you're working for balancer, right? If you're working for AWS Lambda, it'd be some sort of AWS Lambda, it'd be some sort of AWS Lambda, it'd be some sort of serverless runtime. But, uh, it needs to serverless runtime. But, uh, it needs to serverless runtime. But, uh, it needs to actually run somewhere. So, if you're actually run somewhere. So, if you're actually run somewhere. So, if you're infrastructure engineer, next step infrastructure engineer, next step infrastructure engineer, next step is you get into resource provisioning. is you get into resource provisioning. is you get into resource provisioning. Um, and this is already where the single Um, and this is already where the single Um, and this is already where the single node sandbox starts breaking down. How node sandbox starts breaking down. How node sandbox starts breaking down. How do you provision resources within a do you provision resources within a do you provision resources within a single sandbox? You can't exactly single sandbox? You can't exactly single sandbox? You can't exactly simulate something like EC2 or Cloud simulate something like EC2 or Cloud simulate something like EC2 or Cloud Run, right?

  8. Run, right? Run, right? Um, so you get into this host Um, so you get into this host Um, so you get into this host provisioning. Uh, it also includes provisioning. Uh, it also includes provisioning. Uh, it also includes provisioning of other resources like provisioning of other resources like provisioning of other resources like VPCs, subnets, security groups. Um, and VPCs, subnets, security groups. Um, and VPCs, subnets, security groups. Um, and you need to expose this through some you need to expose this through some you need to expose this through some sort of API because your customers are sort of API because your customers are sort of API because your customers are going to want to do things like, "Oh, going to want to do things like, "Oh, going to want to do things like, "Oh, give me this shiny piece of software." give me this shiny piece of software." give me this shiny piece of software." Or, "I don't want it anymore. It cost Or, "I don't want it anymore. It cost Or, "I don't want it anymore. It cost too much. I'm going bankrupt. Delete it, too much. I'm going bankrupt. Delete it, too much. I'm going bankrupt. Delete it, please." please." please." Um, Um, Um, and so you're going to need some sort of and so you're going to need some sort of and so you're going to need some sort of front-end API. And if you have front-end API. And if you have front-end API. And if you have enterprise grade customers who really enterprise grade customers who really enterprise grade customers who really care about quality, then you're going to care about quality, then you're going to care about quality, then you're going to have to meet certain bars like have to meet certain bars like have to meet certain bars like throttling, authentication, throttling, authentication, throttling, authentication, authorization. You can't really like go authorization. You can't really like go authorization. You can't really like go without these things, right? Uh, if without these things, right? Uh, if without these things, right? Uh, if you're AWS, then that's CloudTrail, too. you're AWS, then that's CloudTrail, too. you're AWS, then that's CloudTrail, too. Um, and then beyond this, uh, software Um, and then beyond this, uh, software Um, and then beyond this, uh, software is living. People forget this all the is living. People forget this all the is living. People forget this all the time, especially like investors, right? time, especially like investors, right? time, especially like investors, right? Like they'll be like, "Oh, you wrote it. Like they'll be like, "Oh, you wrote it. Like they'll be like, "Oh, you wrote it. You're done." Um, You're done." Um, You're done." Um, but but but software is living and you probably need software is living and you probably need software is living and you probably need some sort of software deployment some sort of software deployment some sort of software deployment component as well.

  9. component as well. component as well. Uh something whenever you have an update Uh something whenever you have an update Uh something whenever you have an update to roll out roll it out. Um and God to roll out roll it out. Um and God to roll out roll it out. Um and God forbid something goes wrong, roll it forbid something goes wrong, roll it forbid something goes wrong, roll it back. back. back. Uh you need to manage all the different Uh you need to manage all the different Uh you need to manage all the different versions and make sure your deployments versions and make sure your deployments versions and make sure your deployments are gradual to limit your blast radius. are gradual to limit your blast radius. are gradual to limit your blast radius. And we're just kind of getting started And we're just kind of getting started And we're just kind of getting started with this. There's all sorts of things with this. There's all sorts of things with this. There's all sorts of things that you need to think about like health that you need to think about like health that you need to think about like health monitoring with awareness for network monitoring with awareness for network monitoring with awareness for network partitions. partitions. partitions. Uh and then how do you communicate with Uh and then how do you communicate with Uh and then how do you communicate with your host so you can change configs on your host so you can change configs on your host so you can change configs on the fly. the fly. the fly. Um maybe your customer actually wants to Um maybe your customer actually wants to Um maybe your customer actually wants to call your endpoint, so you need DNS and call your endpoint, so you need DNS and call your endpoint, so you need DNS and cert management. And then, you know, cert management. And then, you know, cert management. And then, you know, your our service grows a bunch, you need your our service grows a bunch, you need your our service grows a bunch, you need to keep track of all your resources, to keep track of all your resources, to keep track of all your resources, what's going on, fraud and stuff, then what's going on, fraud and stuff, then what's going on, fraud and stuff, then you need admin consoles, uh telemetry, you need admin consoles, uh telemetry, you need admin consoles, uh telemetry, billing if you're making money, billing if you're making money, billing if you're making money, um all sorts of things. And with all of um all sorts of things. And with all of um all sorts of things. And with all of this, I think there's like one more this, I think there's like one more this, I think there's like one more slide for is it scheduling? Yeah. Um slide for is it scheduling? Yeah. Um slide for is it scheduling? Yeah. Um >> I think the point is fairly clear at >> I think the point is fairly clear at >> I think the point is fairly clear at this point. Beyond this point. Beyond this point. Beyond beyond a certain threshold, there is a beyond a certain threshold, there is a beyond a certain threshold, there is a critical mass at which sandboxing on a critical mass at which sandboxing on a critical mass at which sandboxing on a single node uh can only get you so far.

  10. single node uh can only get you so far. single node uh can only get you so far. And that's why we envision the the And that's why we envision the the And that's why we envision the the future being going towards a world where future being going towards a world where future being going towards a world where environments do provision real environments do provision real environments do provision real infrastructure. infrastructure. infrastructure. >> Yeah, so what this is is um >> Yeah, so what this is is um >> Yeah, so what this is is um a multi-node sandbox a multi-node sandbox a multi-node sandbox with with with access to real infra, real cloud access to real infra, real cloud access to real infra, real cloud resources. Uh we kind of put a cloud in resources. Uh we kind of put a cloud in resources. Uh we kind of put a cloud in box, so cloud box could be another name box, so cloud box could be another name box, so cloud box could be another name for this. Um and as you can imagine, for this. Um and as you can imagine, for this. Um and as you can imagine, changing the sandbox type so drastically changing the sandbox type so drastically changing the sandbox type so drastically here affects post-training pipelines as here affects post-training pipelines as here affects post-training pipelines as well, which um I think we might be well, which um I think we might be well, which um I think we might be running a little bit low on time, so we running a little bit low on time, so we running a little bit low on time, so we won't get like too much into it. Um but won't get like too much into it. Um but won't get like too much into it. Um but yeah, like uh one really cool thing, yeah, like uh one really cool thing, yeah, like uh one really cool thing, too, is like you can put a post-training too, is like you can put a post-training too, is like you can put a post-training pipeline in the sandbox, pipeline in the sandbox, pipeline in the sandbox, um and there's some cool stuff with um and there's some cool stuff with um and there's some cool stuff with model training and RSI that you can get model training and RSI that you can get model training and RSI that you can get into there. into there. into there. Um so, you know, then this begs a Um so, you know, then this begs a Um so, you know, then this begs a question, uh this is all cool stuff, question, uh this is all cool stuff, question, uh this is all cool stuff, Joseph. Uh thank you, Sid, for speaking. Joseph. Uh thank you, Sid, for speaking. Joseph. Uh thank you, Sid, for speaking. Why are you leaking all of this alpha, Why are you leaking all of this alpha, Why are you leaking all of this alpha, right? Why are you like telling all your right? Why are you like telling all your right? Why are you like telling all your organizational secrets and telling organizational secrets and telling organizational secrets and telling everybody like, "Oh, okay, you know, how everybody like, "Oh, okay, you know, how everybody like, "Oh, okay, you know, how do you build a system like this?" Um do you build a system like this?" Um do you build a system like this?" Um it's because uh we're really interested it's because uh we're really interested it's because uh we're really interested in these challenges here. We think in these challenges here. We think in these challenges here. We think they're very fun.

  11. they're very fun. they're very fun. Uh you know, we think they're really Uh you know, we think they're really Uh you know, we think they're really cool. We think that you guys are cool cool. We think that you guys are cool cool. We think that you guys are cool people, uh or maybe I'm just lying, who people, uh or maybe I'm just lying, who people, uh or maybe I'm just lying, who knows. Uh and we want to share these knows. Uh and we want to share these knows. Uh and we want to share these challenges with you uh in case you're challenges with you uh in case you're challenges with you uh in case you're interested in working on them as well. interested in working on them as well. interested in working on them as well. As you can imagine, there's a lot of As you can imagine, there's a lot of As you can imagine, there's a lot of different problems that we haven't different problems that we haven't different problems that we haven't touched on here. Like, for example, touched on here. Like, for example, touched on here. Like, for example, spinning up the entire stack for spinning up the entire stack for spinning up the entire stack for something like AWS Lambda takes hours. something like AWS Lambda takes hours. something like AWS Lambda takes hours. Um how do you fit that into a post Um how do you fit that into a post Um how do you fit that into a post training rollout? Uh and then there's training rollout? Uh and then there's training rollout? Uh and then there's cost as well. How do you efficiently cost as well. How do you efficiently cost as well. How do you efficiently manage this? How do you make sure the manage this? How do you make sure the manage this? How do you make sure the sim-to-real gap, even with real sim-to-real gap, even with real sim-to-real gap, even with real resources, it still exists, right? You resources, it still exists, right? You resources, it still exists, right? You still have to have live customer still have to have live customer still have to have live customer traffic. You still have to have uh traffic. You still have to have uh traffic. You still have to have uh problems that only appear at a certain problems that only appear at a certain problems that only appear at a certain scale. scale. scale. So, you know, if you're a distributed So, you know, if you're a distributed So, you know, if you're a distributed systems engineer, systems engineer, systems engineer, um um um you know, if you know, this stuff if you you know, if you know, this stuff if you you know, if you know, this stuff if you train models before, um if you think train models before, um if you think train models before, um if you think that this stuff is cool, that this stuff is cool, that this stuff is cool, uh then you know, we'd love to talk. Uh uh then you know, we'd love to talk. Uh uh then you know, we'd love to talk. Uh we'd love to talk, uh kind of like see we'd love to talk, uh kind of like see we'd love to talk, uh kind of like see where your opinions are, uh hear what where your opinions are, uh hear what where your opinions are, uh hear what you've worked on, maybe that's like, you've worked on, maybe that's like, you've worked on, maybe that's like, "Oh, Kubernetes." And you have like "Oh, Kubernetes." And you have like "Oh, Kubernetes." And you have like opinions. Well, everybody has like opinions. Well, everybody has like opinions. Well, everybody has like opinions on like auto scaling and opinions on like auto scaling and opinions on like auto scaling and rolling deployments and whatever, but rolling deployments and whatever, but rolling deployments and whatever, but like really niche opinions, right? Like, like really niche opinions, right? Like, like really niche opinions, right? Like, at CDO.

  12. at CDO. at CDO. Um Um Um yeah, we'd love to talk to you and hear yeah, we'd love to talk to you and hear yeah, we'd love to talk to you and hear what you have. what you have. what you have. Thank you. Thank you. Thank you. >> Yep. >> Yep. >> Yep. >> Um what's your like primary goal with >> Um what's your like primary goal with >> Um what's your like primary goal with emulator? emulator? emulator? >> Yeah, um that touches into why it's >> Yeah, um that touches into why it's >> Yeah, um that touches into why it's called emulate in the first place, called emulate in the first place, called emulate in the first place, right? Uh, the real world is very, very right? Uh, the real world is very, very right? Uh, the real world is very, very complex, um, and complex, um, and complex, um, and how we as a industry emulate the real how we as a industry emulate the real how we as a industry emulate the real world is incredibly contrived and low world is incredibly contrived and low world is incredibly contrived and low fidelity. So, emulated goal is really fidelity. So, emulated goal is really fidelity. So, emulated goal is really how do you make these agents own systems how do you make these agents own systems how do you make these agents own systems like this, uh, maybe beyond systems, like this, uh, maybe beyond systems, like this, uh, maybe beyond systems, entire companies, by emulating the real entire companies, by emulating the real entire companies, by emulating the real world with full fidelity. world with full fidelity. world with full fidelity. Yeah, go ahead. Yeah, go ahead. Yeah, go ahead. >> Next question is, >> Next question is, >> Next question is, are you are you are you predominantly focused on like predominantly focused on like predominantly focused on like infra and infra and infra and you have a lot of containers, actually hardware kind of related stuff, actually hardware kind of related stuff, or are you also or are you also or are you also like full RL environments like full RL environments like full RL environments like a digital twin?

  13. like a digital twin? like a digital twin? >> Yeah, of course. Like, the question is >> Yeah, of course. Like, the question is >> Yeah, of course. Like, the question is like, you know, infra is really cool. like, you know, infra is really cool. like, you know, infra is really cool. Um, in 2026, all of us are on an infra's Um, in 2026, all of us are on an infra's Um, in 2026, all of us are on an infra's comeback, and it's very sexy. Everybody comeback, and it's very sexy. Everybody comeback, and it's very sexy. Everybody wants to work on it, right? But, there's wants to work on it, right? But, there's wants to work on it, right? But, there's other types of RL environments as well. other types of RL environments as well. other types of RL environments as well. Um, there is, uh, other workflows that Um, there is, uh, other workflows that Um, there is, uh, other workflows that you really want to capture that aren't you really want to capture that aren't you really want to capture that aren't necessarily infra related. Uh, so the necessarily infra related. Uh, so the necessarily infra related. Uh, so the reason why we're starting with infra is, reason why we're starting with infra is, reason why we're starting with infra is, um, there's a couple. Well, the first um, there's a couple. Well, the first um, there's a couple. Well, the first most important one is it speaks to our most important one is it speaks to our most important one is it speaks to our background the most. Um, we think that background the most. Um, we think that background the most. Um, we think that domain expertise is something that domain expertise is something that domain expertise is something that informs how high quality your data can informs how high quality your data can informs how high quality your data can be. Uh, especially with boutique nature be. Uh, especially with boutique nature be. Uh, especially with boutique nature of data nowadays. Um, and the second is of data nowadays. Um, and the second is of data nowadays. Um, and the second is that when we're simulating full that when we're simulating full that when we're simulating full companies, infra is the easiest to companies, infra is the easiest to companies, infra is the easiest to approach. Uh, if you think about like approach. Uh, if you think about like approach. Uh, if you think about like any infra company out there, whether it any infra company out there, whether it any infra company out there, whether it be Superbase or Modal, uh, or any dev be Superbase or Modal, uh, or any dev be Superbase or Modal, uh, or any dev tools company, tools company, tools company, the problem statement is pretty clear. the problem statement is pretty clear. the problem statement is pretty clear. Uh, engineers kind of know what they Uh, engineers kind of know what they Uh, engineers kind of know what they want. If you are working for Modal, you want. If you are working for Modal, you want. If you are working for Modal, you know that your users want GPU sandbox, know that your users want GPU sandbox, know that your users want GPU sandbox, very low latency, very low cost. You very low latency, very low cost. You very low latency, very low cost. You don't want to fail halfway through your don't want to fail halfway through your don't want to fail halfway through your training run. Um training run. Um training run. Um that's what you care about. So, the that's what you care about. So, the that's what you care about. So, the problem statement becomes much easier.

  14. problem statement becomes much easier. problem statement becomes much easier. Whereas, if you are, you know, a a Whereas, if you are, you know, a a Whereas, if you are, you know, a a company in the YC summer 2026 batch, company in the YC summer 2026 batch, company in the YC summer 2026 batch, you're probably still trying to find you're probably still trying to find you're probably still trying to find product market fit, right? product market fit, right? product market fit, right? >> Yeah. At the same time, there's also >> Yeah. At the same time, there's also >> Yeah. At the same time, there's also lessons learned that going really lessons learned that going really lessons learned that going really vertical on a single domain like vertical on a single domain like vertical on a single domain like infrastructure do translate into other infrastructure do translate into other infrastructure do translate into other horizontal domains. So, we're also um horizontal domains. So, we're also um horizontal domains. So, we're also um exploring going deep into one and exploring going deep into one and exploring going deep into one and scaling out that way. >> All right. Uh >> All right. Uh really appreciate it. Uh appreciate the really appreciate it. Uh appreciate the really appreciate it. Uh appreciate the questions. Um we'll probably step out questions. Um we'll probably step out questions. Um we'll probably step out and we can like take a couple more uh and we can like take a couple more uh and we can like take a couple more uh outside just to make sure the next outside just to make sure the next outside just to make sure the next speaker has room. speaker has room. speaker has room. Um yeah. Uh I thank you guys for Um yeah. Uh I thank you guys for Um yeah. Uh I thank you guys for listening. listening. listening. >> Appreciate it.

Summary

This discussion centers on enhancing the reliability and autonomy of AI agents, particularly in performing useful work with minimal supervision. Key subjects include simulating companies in sandboxes, multi-node systems, distributed clusters, and post-training pipelines, with a focus on how data quality directly impacts AI model capabilities, drawing parallels to mission-critical infrastructure failures. The practical takeaway is that the current model capability gap in handling complex infrastructure issues stems from a deficit in relevant, high-quality data, emphasizing the need for better data to improve AI performance in these areas.

View original episode ↗