← Back
Scott Hanselman December 18, 2025 35m

Trusting Agentic AI with Dr. Dawn Song

Read full transcript 25 segments
  1. As I mentioned, even though LMS are so As I mentioned, even though LMS are so powerful, but it's amazing that none of powerful, but it's amazing that none of powerful, but it's amazing that none of us really understands how it works. us really understands how it works. us really understands how it works. [laughter] [laughter] [laughter] >> That seems concerning like someone ought >> That seems concerning like someone ought >> That seems concerning like someone ought to figure that out, to figure that out, to figure that out, >> right? Like hallucinates, right? Has >> right? Like hallucinates, right? Has >> right? Like hallucinates, right? Has right can be jailbreed and various right can be jailbreed and various right can be jailbreed and various issues and and think about it that issues and and think about it that issues and and think about it that provides the intelligence of our agent provides the intelligence of our agent provides the intelligence of our agent systems. systems. systems. >> Hey friends, it's Scott. I want to thank >> Hey friends, it's Scott. I want to thank >> Hey friends, it's Scott. I want to thank our new sponsor, Mail Trap. Modern email our new sponsor, Mail Trap. Modern email our new sponsor, Mail Trap. Modern email delivery for developers. They integrate delivery for developers. They integrate delivery for developers. They integrate straight into your code with their SDKs. straight into your code with their SDKs. straight into your code with their SDKs. You get unified transactional and You get unified transactional and You get unified transactional and promotional email delivery, 247 support. promotional email delivery, 247 support. promotional email delivery, 247 support. You contact humans, not AI chatbots. You contact humans, not AI chatbots. You contact humans, not AI chatbots. We'll give you 3,500 emails monthly in We'll give you 3,500 emails monthly in We'll give you 3,500 emails monthly in the free tier. And you can try them out the free tier. And you can try them out the free tier. And you can try them out at mailtrap.io at mailtrap.io at mailtrap.io today. That's m a lap.io today. That's m a lap.io today. That's m a lap.io today. today. today. Hi, I'm Scott Hansselman and this Hi, I'm Scott Hansselman and this Hi, I'm Scott Hansselman and this episode of Hansel Minutes is in episode of Hansel Minutes is in episode of Hansel Minutes is in association with the ACM Bitecast. Today association with the ACM Bitecast. Today association with the ACM Bitecast. Today I have the distinct pleasure of speaking I have the distinct pleasure of speaking I have the distinct pleasure of speaking with Dr. Don Song. She's a professor in with Dr. Don Song. She's a professor in with Dr. Don Song. She's a professor in computer science at UC Berkeley and the computer science at UC Berkeley and the computer science at UC Berkeley and the co-director of the Berkeley Center on co-director of the Berkeley Center on co-director of the Berkeley Center on Responsible Decentralized Intelligence.

  2. Responsible Decentralized Intelligence. Responsible Decentralized Intelligence. She's also the recipient of various She's also the recipient of various She's also the recipient of various awards including the MacArthur awards including the MacArthur awards including the MacArthur Fellowship, the Guggenheim Fellowship Fellowship, the Guggenheim Fellowship Fellowship, the Guggenheim Fellowship and more and more. And I'm just thrilled and more and more. And I'm just thrilled and more and more. And I'm just thrilled to be chatting with you today. Thank you to be chatting with you today. Thank you to be chatting with you today. Thank you so much Dr. Song for spending time with so much Dr. Song for spending time with so much Dr. Song for spending time with us. us. us. >> Great. Thanks a lot for having me. So >> Great. Thanks a lot for having me. So >> Great. Thanks a lot for having me. So you're uh you're you have got such an you're uh you're you have got such an you're uh you're you have got such an impressive background. I'm just curious impressive background. I'm just curious impressive background. I'm just curious uh you know when you started your uh you know when you started your uh you know when you started your journey in your academic journey in journey in your academic journey in journey in your academic journey in security did you think that the work you security did you think that the work you security did you think that the work you would be doing would be so recognized it would be doing would be so recognized it would be doing would be so recognized it would be such a big fun long career. would be such a big fun long career. would be such a big fun long career. >> Oh I see. Okay. Um yeah thank you. >> Oh I see. Okay. Um yeah thank you. >> Oh I see. Okay. Um yeah thank you. Thanks for the question. Um actually Thanks for the question. Um actually Thanks for the question. Um actually when I started working in cyber security when I started working in cyber security when I started working in cyber security so first of all uh the field was really so first of all uh the field was really so first of all uh the field was really really small really small really small >> I uh I mean the conference that you go >> I uh I mean the conference that you go >> I uh I mean the conference that you go to is only like u maybe a hundred couple to is only like u maybe a hundred couple to is only like u maybe a hundred couple hundred people. Um and also when I hundred people. Um and also when I hundred people. Um and also when I started I just actually transitioned um started I just actually transitioned um started I just actually transitioned um switch from being a physics major to switch from being a physics major to switch from being a physics major to computer science. So, [laughter] so computer science. So, [laughter] so computer science. So, [laughter] so yeah, so I actually did my, you know, yeah, so I actually did my, you know, yeah, so I actually did my, you know, undergrads in physics and I only undergrads in physics and I only undergrads in physics and I only switched to computer science in grad switched to computer science in grad switched to computer science in grad school and when I first switched, I was school and when I first switched, I was school and when I first switched, I was trying to figure out what I want to trying to figure out what I want to trying to figure out what I want to focus on like you know the domain and I focus on like you know the domain and I focus on like you know the domain and I actually found security really actually found security really actually found security really interesting and also I like the interesting and also I like the interesting and also I like the combination of theory and practice. So combination of theory and practice. So combination of theory and practice. So that's why I chose it and uh you know

  3. that's why I chose it and uh you know that's why I chose it and uh you know given given given um right given the fresh transition and um right given the fresh transition and um right given the fresh transition and also the field was very small. So it was also the field was very small. So it was also the field was very small. So it was I think it was difficult to predict I think it was difficult to predict I think it was difficult to predict what's going to happen in the future. what's going to happen in the future. what's going to happen in the future. But I do know that security um you know But I do know that security um you know But I do know that security um you know was important and was going to be a lot was important and was going to be a lot was important and was going to be a lot more important. So I'm very happy that more important. So I'm very happy that more important. So I'm very happy that uh I chose the path. uh I chose the path. uh I chose the path. >> Yeah. It's funny sometimes people ask me >> Yeah. It's funny sometimes people ask me >> Yeah. It's funny sometimes people ask me like my career like did you plan all of like my career like did you plan all of like my career like did you plan all of this and it's easy to say looking back this and it's easy to say looking back this and it's easy to say looking back oh yeah it was all a plan but it you oh yeah it was all a plan but it you oh yeah it was all a plan but it you just work hard you did your best and just work hard you did your best and just work hard you did your best and people people recognize it and you people people recognize it and you people people recognize it and you follow your your sense of smell to like follow your your sense of smell to like follow your your sense of smell to like what is the next thing? what is the next thing? what is the next thing? >> Yes that's actually a really good way of >> Yes that's actually a really good way of >> Yes that's actually a really good way of putting Yeah. putting Yeah. putting Yeah. >> Yeah. Now, the the MacArthur Fellowship >> Yeah. Now, the the MacArthur Fellowship >> Yeah. Now, the the MacArthur Fellowship and some of the other recognitions that and some of the other recognitions that and some of the other recognitions that you've had like you're an ACM fellow, you've had like you're an ACM fellow, you've had like you're an ACM fellow, you're an ITE E fellow, these are rare you're an ITE E fellow, these are rare you're an ITE E fellow, these are rare and you're stacking them up. I'm and you're stacking them up. I'm and you're stacking them up. I'm curious, when you got something like a curious, when you got something like a curious, when you got something like a genius grant like the MacArthur genius grant like the MacArthur genius grant like the MacArthur Fellowship, did that change what you Fellowship, did that change what you Fellowship, did that change what you chose to follow or do you still plan chose to follow or do you still plan chose to follow or do you still plan your agenda based on your on your gut your agenda based on your on your gut your agenda based on your on your gut and where the research takes you?

  4. and where the research takes you? and where the research takes you? >> Uh, thanks. That's a very good question. >> Uh, thanks. That's a very good question. >> Uh, thanks. That's a very good question. I think in some sense it does give me I think in some sense it does give me I think in some sense it does give me maybe like more freedom more a sense of maybe like more freedom more a sense of maybe like more freedom more a sense of courage is to really [clears throat] courage is to really [clears throat] courage is to really [clears throat] explore things that right I find explore things that right I find explore things that right I find interesting that um I feel that can be interesting that um I feel that can be interesting that um I feel that can be impactful in the future. Uh so yes impactful in the future. Uh so yes impactful in the future. Uh so yes actually my trajectory after the actually my trajectory after the actually my trajectory after the MacArthur fellowship has really even MacArthur fellowship has really even MacArthur fellowship has really even further broadened my research domain and further broadened my research domain and further broadened my research domain and I think um I actually have been taking a I think um I actually have been taking a I think um I actually have been taking a quite unusual path than I think a lot of quite unusual path than I think a lot of quite unusual path than I think a lot of uh a lot of people. uh a lot of people. uh a lot of people. >> I like that idea that it gave you >> I like that idea that it gave you >> I like that idea that it gave you courage in the sense of like it's a very courage in the sense of like it's a very courage in the sense of like it's a very big validation and it's also like the big validation and it's also like the big validation and it's also like the direction we're headed is a good one. direction we're headed is a good one. direction we're headed is a good one. I'm going to now take some risks, make I'm going to now take some risks, make I'm going to now take some risks, make some make some strong decisions. Did it some make some strong decisions. Did it some make some strong decisions. Did it change how you formed your team? Did it change how you formed your team? Did it change how you formed your team? Did it change your your feelings about taking change your your feelings about taking change your your feelings about taking risks? risks? risks? Yes. Yeah, that's that's a very good Yes. Yeah, that's that's a very good Yes. Yeah, that's that's a very good question. So I would say like in my question. So I would say like in my question. So I would say like in my research career um has been quite research career um has been quite research career um has been quite different from a lot of people from most different from a lot of people from most different from a lot of people from most people and that I have as you mentioned people and that I have as you mentioned people and that I have as you mentioned at the beginning I actually have at the beginning I actually have at the beginning I actually have explored uh fairly broad broadly and explored uh fairly broad broadly and explored uh fairly broad broadly and also at the same time you know deeply in also at the same time you know deeply in also at the same time you know deeply in a number of different domains. Um so yes a number of different domains. Um so yes a number of different domains. Um so yes uh so after my cathther you know uh so after my cathther you know uh so after my cathther you know fellowship I right so as you mentioned fellowship I right so as you mentioned fellowship I right so as you mentioned initially my career started in security initially my career started in security initially my career started in security security and privacy and uh and also

  5. security and privacy and uh and also security and privacy and uh and also always been interested in you know how always been interested in you know how always been interested in you know how the brain works how and want to build the brain works how and want to build the brain works how and want to build intelligent machines. So, so yeah, so intelligent machines. So, so yeah, so intelligent machines. So, so yeah, so after the math fellowship, I actually after the math fellowship, I actually after the math fellowship, I actually and then I also did some startups and my and then I also did some startups and my and then I also did some startups and my startup was acquired and I was asking startup was acquired and I was asking startup was acquired and I was asking myself what I want to do if I um you myself what I want to do if I um you myself what I want to do if I um you know retired. I had retired. Then the know retired. I had retired. Then the know retired. I had retired. Then the conclusion was that I want to build conclusion was that I want to build conclusion was that I want to build intelligent machines. So then actually intelligent machines. So then actually intelligent machines. So then actually you know switched my whole group and you know switched my whole group and you know switched my whole group and then actually focused on deep learning then actually focused on deep learning then actually focused on deep learning before deep learning was actually hot. before deep learning was actually hot. before deep learning was actually hot. this was you know even before uh alpha this was you know even before uh alpha this was you know even before uh alpha ago and you know the the last wave uh ago and you know the the last wave uh ago and you know the the last wave uh and so on. So yeah so I would I would and so on. So yeah so I would I would and so on. So yeah so I would I would say it's um you know for most people say it's um you know for most people say it's um you know for most people that would be pretty big change. that would be pretty big change. that would be pretty big change. I was in a meeting recently that I felt I was in a meeting recently that I felt I was in a meeting recently that I felt maybe had too many managers and a person maybe had too many managers and a person maybe had too many managers and a person texted me in the meeting privately and texted me in the meeting privately and texted me in the meeting privately and they said there's a lot of talkers in they said there's a lot of talkers in they said there's a lot of talkers in this meeting and not a lot of doers and this meeting and not a lot of doers and this meeting and not a lot of doers and one of the things that I would give you one of the things that I would give you one of the things that I would give you a compliment about about your career is a compliment about about your career is a compliment about about your career is it seems like as academics go you're a it seems like as academics go you're a it seems like as academics go you're a doer like you create centers you make doer like you create centers you make doer like you create centers you make conferences you you are outward facing conferences you you are outward facing conferences you you are outward facing you're talking to people you're creating you're talking to people you're creating you're talking to people you're creating you know massively online courses is you know massively online courses is you know massively online courses is while other academics tend to kind of while other academics tend to kind of while other academics tend to kind of fold within themselves and they just fold within themselves and they just fold within themselves and they just kind of like disappear and write a paper kind of like disappear and write a paper kind of like disappear and write a paper for a year and then they kind of pop for a year and then they kind of pop for a year and then they kind of pop back up occasionally and you said you back up occasionally and you said you back up occasionally and you said you did startups as well. How do you find

  6. did startups as well. How do you find did startups as well. How do you find that balance between what academia that balance between what academia that balance between what academia expects from a people talking in a room expects from a people talking in a room expects from a people talking in a room perspective and a let's do things, let's perspective and a let's do things, let's perspective and a let's do things, let's ship products, let's make lives better. ship products, let's make lives better. ship products, let's make lives better. You're you seem to be a doer, not a You're you seem to be a doer, not a You're you seem to be a doer, not a talker. talker. talker. >> I see. Okay. So first of all I think um >> I see. Okay. So first of all I think um >> I see. Okay. So first of all I think um I mean everybody has their own path I mean everybody has their own path I mean everybody has their own path everybody has their own preferences and everybody has their own preferences and everybody has their own preferences and people make contributions in their own people make contributions in their own people make contributions in their own ways ways ways >> and I wouldn't say people who are just >> and I wouldn't say people who are just >> and I wouldn't say people who are just working on their papers and maybe you working on their papers and maybe you working on their papers and maybe you know in their offices are talkers I know in their offices are talkers I know in their offices are talkers I think they are I mean some of you know think they are I mean some of you know think they are I mean some of you know some great work actually came out of some great work actually came out of some great work actually came out of that you know that kind of settings as that you know that kind of settings as that you know that kind of settings as well and so yes I wouldn't say well and so yes I wouldn't say well and so yes I wouldn't say necessarily necessarily necessarily right? One, right? One, right? One, you know, one approach is necessary you know, one approach is necessary you know, one approach is necessary better than others. But I think people better than others. But I think people better than others. But I think people they people have different aspirations, they people have different aspirations, they people have different aspirations, people like to do different things. Uh people like to do different things. Uh people like to do different things. Uh I'm glad that um the path that I chose I'm glad that um the path that I chose I'm glad that um the path that I chose the type of works that uh I have been the type of works that uh I have been the type of works that uh I have been doing uh have impacted a lot of people doing uh have impacted a lot of people doing uh have impacted a lot of people and uh you know with the massive open and uh you know with the massive open and uh you know with the massive open online course for example helped you online course for example helped you online course for example helped you know like tens of thousands or even know like tens of thousands or even know like tens of thousands or even hundreds of thousands of people to hundreds of thousands of people to hundreds of thousands of people to actually learn about cutting edge new actually learn about cutting edge new actually learn about cutting edge new topics and and so on and also you know topics and and so on and also you know topics and and so on and also you know the startups help transition research the startups help transition research the startups help transition research technologies into the real world and all technologies into the real world and all technologies into the real world and all these things.

  7. these things. these things. Yes. So I think I'm very happy that my Yes. So I think I'm very happy that my Yes. So I think I'm very happy that my work has been able to help a lot of work has been able to help a lot of work has been able to help a lot of people but I think people also right people but I think people also right people but I think people also right they contribute in different ways. they contribute in different ways. they contribute in different ways. >> I appreciate that. I apologize if that >> I appreciate that. I apologize if that >> I appreciate that. I apologize if that was an indelicate question. It was in it was an indelicate question. It was in it was an indelicate question. It was in it was just meant to show the difference was just meant to show the difference was just meant to show the difference between you know really making things between you know really making things between you know really making things happen in a very physical impactful way. happen in a very physical impactful way. happen in a very physical impactful way. But you're right impact comes in But you're right impact comes in But you're right impact comes in different flavors including our friends different flavors including our friends different flavors including our friends that are maybe more quiet in their that are maybe more quiet in their that are maybe more quiet in their writing. Now you co-direct the Berkeley writing. Now you co-direct the Berkeley writing. Now you co-direct the Berkeley Center for Responsible Decentralized Center for Responsible Decentralized Center for Responsible Decentralized Intelligence RDI. Can you explain that Intelligence RDI. Can you explain that Intelligence RDI. Can you explain that mission and what that means and then how mission and what that means and then how mission and what that means and then how do you select the areas that the center do you select the areas that the center do you select the areas that the center focuses on? focuses on? focuses on? >> Yeah, thanks. That's a very good >> Yeah, thanks. That's a very good >> Yeah, thanks. That's a very good question. So the bricky center RDI question. So the bricky center RDI question. So the bricky center RDI responsible decentralized intelligence responsible decentralized intelligence responsible decentralized intelligence works at an intersection of responsible works at an intersection of responsible works at an intersection of responsible innovation, decentralization innovation, decentralization innovation, decentralization and intelligence as AI for example and I and intelligence as AI for example and I and intelligence as AI for example and I would say agentic AI is actually a very would say agentic AI is actually a very would say agentic AI is actually a very good example of the kind of work that we good example of the kind of work that we good example of the kind of work that we focus on. If you look at agentic AI we focus on. If you look at agentic AI we focus on. If you look at agentic AI we want it to be it's really important that want it to be it's really important that want it to be it's really important that it's safe and secure and responsible. So it's safe and secure and responsible. So it's safe and secure and responsible. So we need to we want aentic AI to help we need to we want aentic AI to help we need to we want aentic AI to help with responsible innovation with responsible innovation with responsible innovation and also right intelligence is a key and also right intelligence is a key and also right intelligence is a key part of uh aentic AI and also we hope part of uh aentic AI and also we hope part of uh aentic AI and also we hope that the agent AI future that we build that the agent AI future that we build that the agent AI future that we build is now centralized it's decentralized is now centralized it's decentralized is now centralized it's decentralized you know each of us we may have our own you know each of us we may have our own you know each of us we may have our own personal um assistant personal uh agent personal um assistant personal uh agent personal um assistant personal uh agent that represent us or help us to interact

  8. that represent us or help us to interact that represent us or help us to interact with others with other agents and so on with others with other agents and so on with others with other agents and so on and we'll have lots and lots of and we'll have lots and lots of and we'll have lots and lots of different uh agents that actually different uh agents that actually different uh agents that actually perform different tasks, have different perform different tasks, have different perform different tasks, have different capabilities to right to help make a capabilities to right to help make a capabilities to right to help make a better world for all of us and for better world for all of us and for better world for all of us and for society and also as SMT it's safe and society and also as SMT it's safe and society and also as SMT it's safe and secure and responsible. secure and responsible. secure and responsible. >> It it feels like for the people out in >> It it feels like for the people out in >> It it feels like for the people out in the community like the non-technical the community like the non-technical the community like the non-technical people that AI is having a moment people that AI is having a moment people that AI is having a moment because it's being well branded. We're because it's being well branded. We're because it's being well branded. We're hearing the word agentic just in the hearing the word agentic just in the hearing the word agentic just in the last year or two, but this has been last year or two, but this has been last year or two, but this has been something you've been thinking about for something you've been thinking about for something you've been thinking about for six, seven, eight years. Like what does six, seven, eight years. Like what does six, seven, eight years. Like what does it feel like to hear things you've been it feel like to hear things you've been it feel like to hear things you've been working on for six or seven years now working on for six or seven years now working on for six or seven years now start to break out into the the start to break out into the the start to break out into the the mainstream? Because I think even now mainstream? Because I think even now mainstream? Because I think even now regular people struggle to understand regular people struggle to understand regular people struggle to understand what what is an agent and what is a what what is an agent and what is a what what is an agent and what is a gentic? Is it just an LLM that has the gentic? Is it just an LLM that has the gentic? Is it just an LLM that has the ability to call a tool or is it is there ability to call a tool or is it is there ability to call a tool or is it is there something more there? something more there? something more there? >> I see. Okay. Yes, that's a very good >> I see. Okay. Yes, that's a very good >> I see. Okay. Yes, that's a very good question. So, first of all, I think it's question. So, first of all, I think it's question. So, first of all, I think it's not even it's not just six, seven years.

  9. not even it's not just six, seven years. not even it's not just six, seven years. Actually, it's been much longer than Actually, it's been much longer than Actually, it's been much longer than that, right? I mean, AI has been in the that, right? I mean, AI has been in the that, right? I mean, AI has been in the making for for many decades. And even making for for many decades. And even making for for many decades. And even for my own transition into deep for my own transition into deep for my own transition into deep learning, as I mentioned, I started learning, as I mentioned, I started learning, as I mentioned, I started working in the field of deep learning. working in the field of deep learning. working in the field of deep learning. Even before actually the term really Even before actually the term really Even before actually the term really became became popular and before became became popular and before became became popular and before no most people actually started working no most people actually started working no most people actually started working in the area. in the area. in the area. >> Mhm. >> Mhm. >> Mhm. >> But even then I think yes >> But even then I think yes >> But even then I think yes I would say almost all of us have been I would say almost all of us have been I would say almost all of us have been really surprised at the speed of really surprised at the speed of really surprised at the speed of advancement for you know frontier AI and advancement for you know frontier AI and advancement for you know frontier AI and so on. Um there has been you know polls so on. Um there has been you know polls so on. Um there has been you know polls and uh also if you just ask most AI and uh also if you just ask most AI and uh also if you just ask most AI researchers who right are working in AI researchers who right are working in AI researchers who right are working in AI today back then before like right like today back then before like right like today back then before like right like CH GPT came out before GPT 3.5 or GP4 CH GPT came out before GPT 3.5 or GP4 CH GPT came out before GPT 3.5 or GP4 came out like what people expected for a came out like what people expected for a came out like what people expected for a lot of the tasks people would expect lot of the tasks people would expect lot of the tasks people would expect that still it would take decades for that still it would take decades for that still it would take decades for those tasks you know to be able to be those tasks you know to be able to be those tasks you know to be able to be accomplished by AI.

  10. accomplished by AI. accomplished by AI. But today, you know, here is where we But today, you know, here is where we But today, you know, here is where we are and I think most people, almost are and I think most people, almost are and I think most people, almost everyone um has been very surprised. everyone um has been very surprised. everyone um has been very surprised. >> Yeah. Yeah. Certainly the math, the >> Yeah. Yeah. Certainly the math, the >> Yeah. Yeah. Certainly the math, the work, the deep learning, the sub, you work, the deep learning, the sub, you work, the deep learning, the sub, you know, the subset of machine learning, know, the subset of machine learning, know, the subset of machine learning, the multi-layered neural networks. This the multi-layered neural networks. This the multi-layered neural networks. This is something that you said has been been is something that you said has been been is something that you said has been been worked on for decades. It popped when worked on for decades. It popped when worked on for decades. It popped when GPT started when the transformer GPT started when the transformer GPT started when the transformer architecture was introduced. Do you architecture was introduced. Do you architecture was introduced. Do you think that there's an overemphasis on think that there's an overemphasis on think that there's an overemphasis on next token prediction on transformer next token prediction on transformer next token prediction on transformer architecture when there's so much other architecture when there's so much other architecture when there's so much other really interesting work happening in really interesting work happening in really interesting work happening in deep learning and in machine learning? deep learning and in machine learning? deep learning and in machine learning? >> Yeah, that's a great question. I think >> Yeah, that's a great question. I think >> Yeah, that's a great question. I think so. So of course what has been shown now so. So of course what has been shown now so. So of course what has been shown now is this next token prediction uh is this next token prediction uh is this next token prediction uh paradigm has been very powerful and also paradigm has been very powerful and also paradigm has been very powerful and also recently the reinforcement learning um recently the reinforcement learning um recently the reinforcement learning um based approaches also have been shown um based approaches also have been shown um based approaches also have been shown um to be really helpful effective at to be really helpful effective at to be really helpful effective at improving right the model capabilities improving right the model capabilities improving right the model capabilities and also in particular for agent and also in particular for agent and also in particular for agent capabilities and and so on. Um and of capabilities and and so on. Um and of capabilities and and so on. Um and of course I think now this is a big course I think now this is a big course I think now this is a big question is is the this transformer with question is is the this transformer with question is is the this transformer with um right the current uh you know um right the current uh you know um right the current uh you know training paradigm with RL and so on will training paradigm with RL and so on will training paradigm with RL and so on will this path be sufficient for us to get to this path be sufficient for us to get to this path be sufficient for us to get to where we want where we want where we want and I mean the truth of the matter is and I mean the truth of the matter is and I mean the truth of the matter is nobody really knows but so far we are

  11. nobody really knows but so far we are nobody really knows but so far we are continuing to see and still the fast continuing to see and still the fast continuing to see and still the fast progress progress progress of model capabilities and also the you of model capabilities and also the you of model capabilities and also the you know the aging development and uh and so know the aging development and uh and so know the aging development and uh and so on. on. on. >> So I mean of course I think we would >> So I mean of course I think we would >> So I mean of course I think we would love to see u more exploration on d more love to see u more exploration on d more love to see u more exploration on d more diverse ideas and so on and even the diverse ideas and so on and even the diverse ideas and so on and even the current paradigm still there are many u current paradigm still there are many u current paradigm still there are many u limitations shortcomings you know not limitations shortcomings you know not limitations shortcomings you know not very data efficient and um very data efficient and um very data efficient and um [clears throat] and so on. So, so we do [clears throat] and so on. So, so we do [clears throat] and so on. So, so we do hope that we can continue to make hope that we can continue to make hope that we can continue to make further progress and identify new ideas, further progress and identify new ideas, further progress and identify new ideas, new breakthroughs and so on. Uh, and in new breakthroughs and so on. Uh, and in new breakthroughs and so on. Uh, and in the meantime, I do foresee that we'll the meantime, I do foresee that we'll the meantime, I do foresee that we'll continue to continue to continue to see the improvements on the model see the improvements on the model see the improvements on the model capabilities and so on. So at at a very capabilities and so on. So at at a very capabilities and so on. So at at a very simp as a very simplistic example, if I simp as a very simplistic example, if I simp as a very simplistic example, if I take a small GPT on my computer and I take a small GPT on my computer and I take a small GPT on my computer and I give it access to tools and I let it run give it access to tools and I let it run give it access to tools and I let it run around on my file system and edit files around on my file system and edit files around on my file system and edit files and do things, I have the basics of an and do things, I have the basics of an and do things, I have the basics of an agentic AI but agentic AI but agentic AI but >> coding agent in that case >> coding agent in that case >> coding agent in that case >> in an agent. Yeah, a small agent making >> in an agent. Yeah, a small agent making >> in an agent. Yeah, a small agent making a small basic agent. I'm basically a small basic agent. I'm basically a small basic agent. I'm basically letting next token prediction run shell letting next token prediction run shell letting next token prediction run shell scripts on my machine and maybe scripts on my machine and maybe scripts on my machine and maybe productivity comes out of it. But one of productivity comes out of it. But one of productivity comes out of it. But one of the themes in your research bio is the the themes in your research bio is the the themes in your research bio is the intersection of deep learning and intersection of deep learning and intersection of deep learning and security. I think I think about this security. I think I think about this security. I think I think about this little agent that runs on my machine and little agent that runs on my machine and little agent that runs on my machine and then maybe a robot in my house that has then maybe a robot in my house that has then maybe a robot in my house that has arms and legs and a model behind it. In

  12. arms and legs and a model behind it. In arms and legs and a model behind it. In both of those instances, do no harm has both of those instances, do no harm has both of those instances, do no harm has always been one of the kind of ideas always been one of the kind of ideas always been one of the kind of ideas around robotics. The first rule is like around robotics. The first rule is like around robotics. The first rule is like do no harm. Is that something that is do no harm. Is that something that is do no harm. Is that something that is possible for an Aentic AI to be both possible for an Aentic AI to be both possible for an Aentic AI to be both secure and helpful or are we always secure and helpful or are we always secure and helpful or are we always going to have that tension? going to have that tension? going to have that tension? >> Oh, that's that's a very good question. >> Oh, that's that's a very good question. >> Oh, that's that's a very good question. So, um, okay. So, also first when you So, um, okay. So, also first when you So, um, okay. So, also first when you know earlier you also asked about what know earlier you also asked about what know earlier you also asked about what is agentic AI. is agentic AI. is agentic AI. >> Thank you. >> Thank you. >> Thank you. >> And so, right, so the examples that you >> And so, right, so the examples that you >> And so, right, so the examples that you mentioned these are very good examples mentioned these are very good examples mentioned these are very good examples of some of the things aentic AI can do. of some of the things aentic AI can do. of some of the things aentic AI can do. But when we talk about gentic AI in But when we talk about gentic AI in But when we talk about gentic AI in general, [clears throat] it's not just general, [clears throat] it's not just general, [clears throat] it's not just about, you know, one type of agent and about, you know, one type of agent and about, you know, one type of agent and so on. It's actually in fact it's a very so on. It's actually in fact it's a very so on. It's actually in fact it's a very broad spectrum. In our recent u overview broad spectrum. In our recent u overview broad spectrum. In our recent u overview paper, we actually lay out um a general paper, we actually lay out um a general paper, we actually lay out um a general landscape for a gentic AI along a number landscape for a gentic AI along a number landscape for a gentic AI along a number of different dimensions. Along each of of different dimensions. Along each of of different dimensions. Along each of these dimensions, essentially the aentic these dimensions, essentially the aentic these dimensions, essentially the aentic systems can be, you know, less flexible systems can be, you know, less flexible systems can be, you know, less flexible versus more flexible. Uh so for example versus more flexible. Uh so for example versus more flexible. Uh so for example you know the kind of tools that they use you know the kind of tools that they use you know the kind of tools that they use whether the tools um pre-specified whether the tools um pre-specified whether the tools um pre-specified in a static set or they can even use in a static set or they can even use in a static set or they can even use dynamic dynamically selected tools dynamic dynamically selected tools dynamic dynamically selected tools during runtime that they didn't even you during runtime that they didn't even you during runtime that they didn't even you know the developer didn't even know that know the developer didn't even know that know the developer didn't even know that or didn't even specify ahead of time and or didn't even specify ahead of time and or didn't even specify ahead of time and so on. and you know the level of so on. and you know the level of so on. and you know the level of autonomy the level of u how um how

  13. autonomy the level of u how um how autonomy the level of u how um how flexible the flow the control flow and flexible the flow the control flow and flexible the flow the control flow and the workflow of the agent is so it's a the workflow of the agent is so it's a the workflow of the agent is so it's a very broad spectrum and given that so very broad spectrum and given that so very broad spectrum and given that so what we also have shown is along with what we also have shown is along with what we also have shown is along with each dimension as a system becomes more each dimension as a system becomes more each dimension as a system becomes more and more flexible and um more and more and more flexible and um more and more and more flexible and um more and more dynamic and so on it also increases the dynamic and so on it also increases the dynamic and so on it also increases the attack surface attack surface attack surface and and also So when we talk about you and and also So when we talk about you and and also So when we talk about you know safety and security of Aentic AI know safety and security of Aentic AI know safety and security of Aentic AI there are actually two main difference there are actually two main difference there are actually two main difference uh uh different aspects. So one is uh uh different aspects. So one is uh uh different aspects. So one is whether the aentic AI system itself is whether the aentic AI system itself is whether the aentic AI system itself is secure whether it can be you know secure secure whether it can be you know secure secure whether it can be you know secure against malicious attacks on the agentic against malicious attacks on the agentic against malicious attacks on the agentic AI system itself. So for example in the AI system itself. So for example in the AI system itself. So for example in the example that you mentioned you have a example that you mentioned you have a example that you mentioned you have a little coding agent that works on your little coding agent that works on your little coding agent that works on your you know uh on your files and so on. You you know uh on your files and so on. You you know uh on your files and so on. You want to be careful that there's no want to be careful that there's no want to be careful that there's no malicious attacks malicious attacks malicious attacks attacking the coding agent so that the attacking the coding agent so that the attacking the coding agent so that the coding agent somehow misbehaved, delete coding agent somehow misbehaved, delete coding agent somehow misbehaved, delete your database and then send out and also your database and then send out and also your database and then send out and also send out like sensitive data right from send out like sensitive data right from send out like sensitive data right from your files to the attacker and and so your files to the attacker and and so your files to the attacker and and so on. So this is one type of concern. And on. So this is one type of concern. And on. So this is one type of concern. And then another type of concern is these uh then another type of concern is these uh then another type of concern is these uh agents as they become powerful agents as they become powerful agents as they become powerful attackers may misuse them as well to attackers may misuse them as well to attackers may misuse them as well to launch attacks you know to other systems launch attacks you know to other systems launch attacks you know to other systems to the internet to the rest of the world to the internet to the rest of the world to the internet to the rest of the world and so on. So that's also um a

  14. and so on. So that's also um a and so on. So that's also um a responsibility that we have uh as we responsibility that we have uh as we responsibility that we have uh as we build these a uh agentic AI is you know build these a uh agentic AI is you know build these a uh agentic AI is you know what people say is with the strong um what people say is with the strong um what people say is with the strong um capabilities also there's the uh strong capabilities also there's the uh strong capabilities also there's the uh strong responsibilities uh as well right responsibilities uh as well right responsibilities uh as well right >> so right so it's both sides and both >> so right so it's both sides and both >> so right so it's both sides and both sides has it own set of challenges and I sides has it own set of challenges and I sides has it own set of challenges and I would say um cyber security has always would say um cyber security has always would say um cyber security has always been challenging a challenging domain. been challenging a challenging domain. been challenging a challenging domain. We are seeing you know attacks every day We are seeing you know attacks every day We are seeing you know attacks every day today already and uh several attacks are today already and uh several attacks are today already and uh several attacks are called causing like you know billions called causing like you know billions called causing like you know billions and billions of dollars of um financial and billions of dollars of um financial and billions of dollars of um financial loss and damages every year loss and damages every year loss and damages every year and now when we add aentic AI I think on and now when we add aentic AI I think on and now when we add aentic AI I think on both sides actually things get a lot both sides actually things get a lot both sides actually things get a lot worse. So for the agent AI systems first worse. So for the agent AI systems first worse. So for the agent AI systems first of all bec because it's much more of all bec because it's much more of all bec because it's much more complex and much more dynamic and so and complex and much more dynamic and so and complex and much more dynamic and so and also we don't actually understand how also we don't actually understand how also we don't actually understand how these large language models work they these large language models work they these large language models work they have intrinsic vulnerabilities issues uh have intrinsic vulnerabilities issues uh have intrinsic vulnerabilities issues uh like jailbreak you know prompt injection like jailbreak you know prompt injection like jailbreak you know prompt injection and so on. So a GTS system itself is and so on. So a GTS system itself is and so on. So a GTS system itself is actually much harder to secure to actually much harder to secure to actually much harder to secure to protect against malicious attacks um on protect against malicious attacks um on protect against malicious attacks um on its own. And then on the other hand when its own. And then on the other hand when its own. And then on the other hand when a Genti systems becomes more powerful a Genti systems becomes more powerful a Genti systems becomes more powerful uh they uh when attackers misuse them uh they uh when attackers misuse them uh they uh when attackers misuse them the consequence can be much worse as

  15. the consequence can be much worse as the consequence can be much worse as well. And this also has been illustrated well. And this also has been illustrated well. And this also has been illustrated with some of our uh our own recent work with some of our uh our own recent work with some of our uh our own recent work in actually evaluating what AI can do in in actually evaluating what AI can do in in actually evaluating what AI can do in cyber security like cyber gym and so on. cyber security like cyber gym and so on. cyber security like cyber gym and so on. >> Yeah. >> Yeah. >> Yeah. A little bit of a of a side rant. I A little bit of a of a side rant. I A little bit of a of a side rant. I remember in the early 90s when they told remember in the early 90s when they told remember in the early 90s when they told us never to trust user input, right? And us never to trust user input, right? And us never to trust user input, right? And you always have your little text boxes you always have your little text boxes you always have your little text boxes and you always put validation on each and you always put validation on each and you always put validation on each text box and you're so careful to not text box and you're so careful to not text box and you're so careful to not trust user input. And now the internet trust user input. And now the internet trust user input. And now the internet is just one giant text box where we type is just one giant text box where we type is just one giant text box where we type pros and we're expected to trust user pros and we're expected to trust user pros and we're expected to trust user input. But that's the now the attack input. But that's the now the attack input. But that's the now the attack vector while the stack is so deep like I vector while the stack is so deep like I vector while the stack is so deep like I have this altter behind me and I have a have this altter behind me and I have a have this altter behind me and I have a PDP11 over there that those are PDP11 over there that those are PDP11 over there that those are computers where you can hold in your computers where you can hold in your computers where you can hold in your brain almost the entire computer. But brain almost the entire computer. But brain almost the entire computer. But now there's no such thing as a full now there's no such thing as a full now there's no such thing as a full stack engineer because a chatbot on the stack engineer because a chatbot on the stack engineer because a chatbot on the internet is a distributed system within internet is a distributed system within internet is a distributed system within another distributed system within another distributed system within another distributed system within virtual machines and you know it's virtual machines and you know it's virtual machines and you know it's complexity all the way down. Is is it complexity all the way down. Is is it complexity all the way down. Is is it problematic that none of us can hold the problematic that none of us can hold the problematic that none of us can hold the full stack in our in our brain anymore?

  16. full stack in our in our brain anymore? full stack in our in our brain anymore? >> That is yes I I think that's a very good >> That is yes I I think that's a very good >> That is yes I I think that's a very good question. It is actually a huge issue as question. It is actually a huge issue as question. It is actually a huge issue as I mentioned. So first of all I mean it's I mentioned. So first of all I mean it's I mentioned. So first of all I mean it's not just about whether we can hold it in not just about whether we can hold it in not just about whether we can hold it in our brain like the as I mentioned even our brain like the as I mentioned even our brain like the as I mentioned even though LMS are so powerful but it's though LMS are so powerful but it's though LMS are so powerful but it's amazing that none of us really amazing that none of us really amazing that none of us really understands how it works understands how it works understands how it works >> that seems concerning like someone ought >> that seems concerning like someone ought >> that seems concerning like someone ought to figure that out [laughter] to figure that out [laughter] to figure that out [laughter] >> right like a hallucin right has right >> right like a hallucin right has right >> right like a hallucin right has right can be jailbreaked and various issues can be jailbreaked and various issues can be jailbreaked and various issues and and think about it that provides the and and think about it that provides the and and think about it that provides the intelligence of our agentic AI systems. intelligence of our agentic AI systems. intelligence of our agentic AI systems. So, so we have this really powerful So, so we have this really powerful So, so we have this really powerful system and uh and also we give it all system and uh and also we give it all system and uh and also we give it all sorts of um you know privileges so that sorts of um you know privileges so that sorts of um you know privileges so that it can do things on our behalf. In the it can do things on our behalf. In the it can do things on our behalf. In the future we may give it our credit card future we may give it our credit card future we may give it our credit card number so it can shop for us right and number so it can shop for us right and number so it can shop for us right and we give it um privilege in our systems we give it um privilege in our systems we give it um privilege in our systems so right it can right to take actions on so right it can right to take actions on so right it can right to take actions on our systems and so on. So it's so our systems and so on. So it's so our systems and so on. So it's so powerful and with all these privileges powerful and with all these privileges powerful and with all these privileges that we gave it but at the same time we that we gave it but at the same time we that we gave it but at the same time we have no idea how it works. We don't know have no idea how it works. We don't know have no idea how it works. We don't know when it can break down. We don't know when it can break down. We don't know when it can break down. We don't know right how it's going to behave under right how it's going to behave under right how it's going to behave under different situations.

  17. different situations. different situations. >> So I think this really causes a huge >> So I think this really causes a huge >> So I think this really causes a huge concerns. So that's why also some of you concerns. So that's why also some of you concerns. So that's why also some of you know my work has been focused on what know my work has been focused on what know my work has been focused on what can we do how we can build more secure can we do how we can build more secure can we do how we can build more secure solutions for these um for these type of solutions for these um for these type of solutions for these um for these type of systems and ideally we want to also systems and ideally we want to also systems and ideally we want to also develop new approaches to still to act develop new approaches to still to act develop new approaches to still to act to to even have probable guarantees of to to even have probable guarantees of to to even have probable guarantees of certain security properties even for certain security properties even for certain security properties even for these agent AI systems. I think that's these agent AI systems. I think that's these agent AI systems. I think that's something that we really need in order something that we really need in order something that we really need in order to actually have a systems to take to actually have a systems to take to actually have a systems to take critical actions for us. critical actions for us. critical actions for us. >> That's a great point. Like here we are >> That's a great point. Like here we are >> That's a great point. Like here we are making these giant distributed programs making these giant distributed programs making these giant distributed programs where the fundamental for loop in the where the fundamental for loop in the where the fundamental for loop in the middle is a black box that's middle is a black box that's middle is a black box that's non-deterministic and we can't trust it non-deterministic and we can't trust it non-deterministic and we can't trust it because it could suddenly decide to be because it could suddenly decide to be because it could suddenly decide to be angry and and cause cause problems. How angry and and cause cause problems. How angry and and cause cause problems. How do you design a system around that? How do you design a system around that? How do you design a system around that? How do you make it so the light switch that do you make it so the light switch that do you make it so the light switch that that can flip off doesn't hurt someone? that can flip off doesn't hurt someone? that can flip off doesn't hurt someone? I'm curious, what do you think about the I'm curious, what do you think about the I'm curious, what do you think about the stochastic parrot analogy? I think stochastic parrot analogy? I think stochastic parrot analogy? I think there's arguments that it's there's arguments that it's there's arguments that it's probabilistic probabilistic probabilistic mimicry and that and the LLM is a kind mimicry and that and the LLM is a kind mimicry and that and the LLM is a kind of a parrot, but then there's also maybe of a parrot, but then there's also maybe of a parrot, but then there's also maybe that that's a simplistic analogy and it that that's a simplistic analogy and it that that's a simplistic analogy and it it under unersells the emergent it under unersells the emergent it under unersells the emergent capabilities of LLMs. I'm curious which capabilities of LLMs. I'm curious which capabilities of LLMs. I'm curious which side you're on.

  18. side you're on. side you're on. >> I see. Yeah, that's a very good >> I see. Yeah, that's a very good >> I see. Yeah, that's a very good question. I mean again to question. I mean again to question. I mean again to [clears throat] at home as I mentioned [clears throat] at home as I mentioned [clears throat] at home as I mentioned we really don't have very good we really don't have very good we really don't have very good understanding of how these LMS work at understanding of how these LMS work at understanding of how these LMS work at all and uh and we do see very all and uh and we do see very all and uh and we do see very interesting phenomenas right on one hand interesting phenomenas right on one hand interesting phenomenas right on one hand these LMS they can win the gold medal in these LMS they can win the gold medal in these LMS they can win the gold medal in these Olympia you know math competitions these Olympia you know math competitions these Olympia you know math competitions right programming contest and so on right programming contest and so on right programming contest and so on >> and uh and they can in certain cases >> and uh and they can in certain cases >> and uh and they can in certain cases solve very hard math problems. solve very hard math problems. solve very hard math problems. But on the other hand, you can easily But on the other hand, you can easily But on the other hand, you can easily see it actually makes very silly see it actually makes very silly see it actually makes very silly mistakes uh very simple right problems. mistakes uh very simple right problems. mistakes uh very simple right problems. Uh and also you know we call the LM has Uh and also you know we call the LM has Uh and also you know we call the LM has this jagged intelligence on certain this jagged intelligence on certain this jagged intelligence on certain things right it does really really well things right it does really really well things right it does really really well and on other things right it does very and on other things right it does very and on other things right it does very poorly and also we have done um some poorly and also we have done um some poorly and also we have done um some recent work also trying to understand recent work also trying to understand recent work also trying to understand better um better um better um whether what is what LM is learning whether what is what LM is learning whether what is what LM is learning whether it how well can it generalize we whether it how well can it generalize we whether it how well can it generalize we actually develop some new benchmarks actually develop some new benchmarks actually develop some new benchmarks Omega, delta and so on to try to u Omega, delta and so on to try to u Omega, delta and so on to try to u develop this controlled experiments to develop this controlled experiments to develop this controlled experiments to really understand how um is doing really understand how um is doing really understand how um is doing generalization both with um supervised generalization both with um supervised generalization both with um supervised funing as well as reinforcement learning funing as well as reinforcement learning funing as well as reinforcement learning and so on. So what our work has shown is

  19. and so on. So what our work has shown is and so on. So what our work has shown is that I mean so first of all yes I mean that I mean so first of all yes I mean that I mean so first of all yes I mean in certain cases our M's capabilities it in certain cases our M's capabilities it in certain cases our M's capabilities it is amazing but then on the other hand uh is amazing but then on the other hand uh is amazing but then on the other hand uh our work does show that um there's still our work does show that um there's still our work does show that um there's still u significant limitations in terms of u significant limitations in terms of u significant limitations in terms of how these LMS actually how well can how these LMS actually how well can how these LMS actually how well can generalize in particular as we increase generalize in particular as we increase generalize in particular as we increase both the difficulty level of the both the difficulty level of the both the difficulty level of the problems and also the compositional problems and also the compositional problems and also the compositional complexity complexity complexity that um doesn't generate and gener that um doesn't generate and gener that um doesn't generate and gener generalize that well and also still I generalize that well and also still I generalize that well and also still I think it can come up with some you know think it can come up with some you know think it can come up with some you know new ideas and so but in general like new ideas and so but in general like new ideas and so but in general like with our benchmark evaluation shows that with our benchmark evaluation shows that with our benchmark evaluation shows that still when some problems that require still when some problems that require still when some problems that require really new solutions new type of really new solutions new type of really new solutions new type of solutions is still not very good uh at solutions is still not very good uh at solutions is still not very good uh at those. You know, I like that term jagged those. You know, I like that term jagged those. You know, I like that term jagged intelligence. Like to to assume that one intelligence. Like to to assume that one intelligence. Like to to assume that one individual is uniquely smart in all individual is uniquely smart in all individual is uniquely smart in all things is to oversimplify. You know, I things is to oversimplify. You know, I things is to oversimplify. You know, I could be a poor driver and I could be a could be a poor driver and I could be a could be a poor driver and I could be a genius in math. I could, you know, genius in math. I could, you know, genius in math. I could, you know, there's a number of things that I am not there's a number of things that I am not there's a number of things that I am not single- dimensional. So, neither are the single- dimensional. So, neither are the single- dimensional. So, neither are the LLMs. Now, you've been teaching for so LLMs. Now, you've been teaching for so LLMs. Now, you've been teaching for so long uh and now you're teaching long uh and now you're teaching long uh and now you're teaching massively open online courses. There's a massively open online courses. There's a massively open online courses. There's a really exciting one that you've been really exciting one that you've been really exciting one that you've been doing, the Agentic AI. This is this has doing, the Agentic AI. This is this has doing, the Agentic AI. This is this has blown up. Did you how many people did blown up. Did you how many people did blown up. Did you how many people did you expect would come to the agentic AI you expect would come to the agentic AI you expect would come to the agentic AI MOO and how many are coming now?

  20. MOO and how many are coming now? MOO and how many are coming now? >> Yes. Yes. Yeah. Thanks. So, right. So, I >> Yes. Yes. Yeah. Thanks. So, right. So, I >> Yes. Yes. Yeah. Thanks. So, right. So, I first started this massive open online first started this massive open online first started this massive open online course MOO on aentic AI actually fall of course MOO on aentic AI actually fall of course MOO on aentic AI actually fall of last year 2024. last year 2024. last year 2024. >> Mh. >> Mh. >> Mh. >> And when I started actually back then um >> And when I started actually back then um >> And when I started actually back then um the agentic AI or agents wasn't quite a the agentic AI or agents wasn't quite a the agentic AI or agents wasn't quite a thing yet. um not many people were thing yet. um not many people were thing yet. um not many people were talking about it but however I could talking about it but however I could talking about it but however I could foresee that this is this is the future foresee that this is this is the future foresee that this is this is the future this is the next frontier so that's how this is the next frontier so that's how this is the next frontier so that's how I actually you know started teaching the I actually you know started teaching the I actually you know started teaching the class and I think because it was the class and I think because it was the class and I think because it was the first it was literally I think the first first it was literally I think the first first it was literally I think the first class on agents uh agentic AI and it's class on agents uh agentic AI and it's class on agents uh agentic AI and it's the first MOO uh on the on the topic as the first MOO uh on the on the topic as the first MOO uh on the on the topic as well so I think it really caught well so I think it really caught well so I think it really caught people's attention people's attention people's attention And now we're actually running the third And now we're actually running the third And now we're actually running the third edition for the class and we have edition for the class and we have edition for the class and we have overall uh over 32,000 enrolled overall uh over 32,000 enrolled overall uh over 32,000 enrolled globally. So that's been really exciting globally. So that's been really exciting globally. So that's been really exciting and also suddenly you know this year and also suddenly you know this year and also suddenly you know this year it's now called the year of agents and it's now called the year of agents and it's now called the year of agents and uh even though as I said when I started uh even though as I said when I started uh even though as I said when I started the class I you know it's because I the class I you know it's because I the class I you know it's because I could see foresee that this is the next could see foresee that this is the next could see foresee that this is the next frontier but even I did not expect frontier but even I did not expect frontier but even I did not expect things to explode so fast like this year things to explode so fast like this year things to explode so fast like this year uh you know especially after the uh you know especially after the uh you know especially after the reasoning models um came out that really reasoning models um came out that really reasoning models um came out that really helped with the reasoning uh capab

  21. helped with the reasoning uh capab helped with the reasoning uh capab capability of agents overall and we are capability of agents overall and we are capability of agents overall and we are really seeing the field uh exploded. So really seeing the field uh exploded. So really seeing the field uh exploded. So that's been really exciting to see. that's been really exciting to see. that's been really exciting to see. >> At what level should a person feel >> At what level should a person feel >> At what level should a person feel comfort around their level of computer comfort around their level of computer comfort around their level of computer science and AI before they join a MOO science and AI before they join a MOO science and AI before they join a MOO like this? Like is this for high school like this? Like is this for high school like this? Like is this for high school students? Is this for graduate students? students? Is this for graduate students? students? Is this for graduate students? Is this for practitioners like myself? Is this for practitioners like myself? Is this for practitioners like myself? what should I come into a course like what should I come into a course like what should I come into a course like this knowing and being prepared for? this knowing and being prepared for? this knowing and being prepared for? >> Yeah, that's that's a great question. >> Yeah, that's that's a great question. >> Yeah, that's that's a great question. So, I would say actually the the course So, I would say actually the the course So, I would say actually the the course is designed to have something to offer is designed to have something to offer is designed to have something to offer for for people all people at different for for people all people at different for for people all people at different levels. levels. levels. >> Mh. >> Mh. >> Mh. >> So, of course the course is mainly >> So, of course the course is mainly >> So, of course the course is mainly designed for people with technical designed for people with technical designed for people with technical backgrounds and in computer science uh backgrounds and in computer science uh backgrounds and in computer science uh and so on. And we actually do and so on. And we actually do and so on. And we actually do systematically cover you know both in systematically cover you know both in systematically cover you know both in terms of different layers of the agent terms of different layers of the agent terms of different layers of the agent AI stack uh all the way you know from AI stack uh all the way you know from AI stack uh all the way you know from the foundation the foundations the model the foundation the foundations the model the foundation the foundations the model um development capabilities uh to a um development capabilities uh to a um development capabilities uh to a genti framework uh all the way to genti framework uh all the way to genti framework uh all the way to applications um both horizontal and applications um both horizontal and applications um both horizontal and vertical applications and so on. Um so vertical applications and so on. Um so vertical applications and so on. Um so and the class is technical but on the and the class is technical but on the and the class is technical but on the other hand um also I think even just other hand um also I think even just other hand um also I think even just from the lectures even for people who from the lectures even for people who from the lectures even for people who don't have too much background there's don't have too much background there's don't have too much background there's still I think quite a bit that they can still I think quite a bit that they can still I think quite a bit that they can learn about just a general uh overall learn about just a general uh overall learn about just a general uh overall development and in the in the space.

  22. development and in the in the space. development and in the in the space. >> Yeah it's really a huge source of of >> Yeah it's really a huge source of of >> Yeah it's really a huge source of of material to explore to look back at material to explore to look back at material to explore to look back at spring and fall of last year. It's worth spring and fall of last year. It's worth spring and fall of last year. It's worth noting that the supplemental readings, noting that the supplemental readings, noting that the supplemental readings, the links to all of the quizzes, the the links to all of the quizzes, the the links to all of the quizzes, the slides, the videos are all available slides, the videos are all available slides, the videos are all available online. So, people should go back and online. So, people should go back and online. So, people should go back and explore. This is really very formal and explore. This is really very formal and explore. This is really very formal and structured and deep with a lot of really structured and deep with a lot of really structured and deep with a lot of really great guest speakers that you've put great guest speakers that you've put great guest speakers that you've put together. And now you've even got a together. And now you've even got a together. And now you've even got a contest for the greater good that you're contest for the greater good that you're contest for the greater good that you're doing, an open competition for agents doing, an open competition for agents doing, an open competition for agents that hopefully make people's lives that hopefully make people's lives that hopefully make people's lives better. better. better. >> Yes. Yeah. Thanks. Yes. So we are >> Yes. Yeah. Thanks. Yes. So we are >> Yes. Yeah. Thanks. Yes. So we are running a competition. Uh so for each running a competition. Uh so for each running a competition. Uh so for each edition of the MOO we actually have uh edition of the MOO we actually have uh edition of the MOO we actually have uh organized the competition. So for organized the competition. So for organized the competition. So for example the last one in this uh past example the last one in this uh past example the last one in this uh past spring we had close to a thousand teams spring we had close to a thousand teams spring we had close to a thousand teams that participated globally and uh for that participated globally and uh for that participated globally and uh for this semester for this addition we have this semester for this addition we have this semester for this addition we have a new competition which actually focuses a new competition which actually focuses a new competition which actually focuses on agent evaluation. So as we develop on agent evaluation. So as we develop on agent evaluation. So as we develop agents actually it's really important to agents actually it's really important to agents actually it's really important to have uh good agent evaluation to have have uh good agent evaluation to have have uh good agent evaluation to have good methodologies for aging evaluation good methodologies for aging evaluation good methodologies for aging evaluation because you know we I mean they're just because you know we I mean they're just because you know we I mean they're just saying that we can only improve what we saying that we can only improve what we saying that we can only improve what we can measure and also in general can measure and also in general can measure and also in general evaluations um essentially the goalpost evaluations um essentially the goalpost evaluations um essentially the goalpost uh for right for development for the uh for right for development for the uh for right for development for the community. uh so you know my group uh community. uh so you know my group uh community. uh so you know my group uh we've had earlier work in a space like

  23. we've had earlier work in a space like we've had earlier work in a space like uh MMLU and some of these other uh MMLU and some of these other uh MMLU and some of these other benchmarks have been widely actually benchmarks have been widely actually benchmarks have been widely actually uh adopted uh in the community and uh uh adopted uh in the community and uh uh adopted uh in the community and uh but a lot of those were focused on I but a lot of those were focused on I but a lot of those were focused on I would say evaluation at the model level would say evaluation at the model level would say evaluation at the model level but for a evaluation actually it but for a evaluation actually it but for a evaluation actually it requires um different um essentially requires um different um essentially requires um different um essentially different focus given that the agent is different focus given that the agent is different focus given that the agent is not just a model. Uh agent evaluation is not just a model. Uh agent evaluation is not just a model. Uh agent evaluation is not just a model evaluation. You not just a model evaluation. You not just a model evaluation. You actually have the model and also you actually have the model and also you actually have the model and also you have the agent itself like also called have the agent itself like also called have the agent itself like also called the harness that actually you know uses the harness that actually you know uses the harness that actually you know uses the model to the model to the model to to essentially perform tasks uh and so to essentially perform tasks uh and so to essentially perform tasks uh and so on. So the hing evaluation essentially on. So the hing evaluation essentially on. So the hing evaluation essentially has more has more components for the has more has more components for the has more has more components for the evaluation and it's very important to evaluation and it's very important to evaluation and it's very important to have um open standardized reproducible have um open standardized reproducible have um open standardized reproducible evaluations for agents and so far this evaluations for agents and so far this evaluations for agents and so far this has been lacking. So we actually have uh has been lacking. So we actually have uh has been lacking. So we actually have uh developed a new paradigm. We call it a a developed a new paradigm. We call it a a developed a new paradigm. We call it a a gentified agent assessment AAA that gentified agent assessment AAA that gentified agent assessment AAA that actually helps to meet this need uh to actually helps to meet this need uh to actually helps to meet this need uh to enable a new [clears throat] paradigm of enable a new [clears throat] paradigm of enable a new [clears throat] paradigm of this open reproducible standard this open reproducible standard this open reproducible standard standardized agent evaluation and also standardized agent evaluation and also standardized agent evaluation and also we are developing a platform for this as we are developing a platform for this as we are developing a platform for this as well. And this agent evaluation well. And this agent evaluation well. And this agent evaluation competition is really to help bring the competition is really to help bring the competition is really to help bring the communities together to develop the best communities together to develop the best communities together to develop the best benchmark evaluation that's

  24. benchmark evaluation that's benchmark evaluation that's standardized, reproducible, and has standardized, reproducible, and has standardized, reproducible, and has broad coverage in diverse domains that broad coverage in diverse domains that broad coverage in diverse domains that helps guides the community development. helps guides the community development. helps guides the community development. And we really hope that u more people And we really hope that u more people And we really hope that u more people can join the competition. We have can join the competition. We have can join the competition. We have actually over $1 million in prizes and actually over $1 million in prizes and actually over $1 million in prizes and resources resources resources um provided by sponsors and partners of um provided by sponsors and partners of um provided by sponsors and partners of the competition including Google Deep the competition including Google Deep the competition including Google Deep Mind and many others and so on. And so I Mind and many others and so on. And so I Mind and many others and so on. And so I think this will be really fun and also think this will be really fun and also think this will be really fun and also it's a great opportunity for the it's a great opportunity for the it's a great opportunity for the community to come together to develop community to come together to develop community to come together to develop public good. So so we hope that more public good. So so we hope that more public good. So so we hope that more people can join us in this competition. people can join us in this competition. people can join us in this competition. Yeah, this is very exciting and folks Yeah, this is very exciting and folks Yeah, this is very exciting and folks can explore all this stuff. There's can explore all this stuff. There's can explore all this stuff. There's agentbeats.org. agentbeats.org. agentbeats.org. They can see the code. They can take a They can see the code. They can take a They can see the code. They can take a look at rdi.berkley.edu look at rdi.berkley.edu look at rdi.berkley.edu to learn about agent X and about agent to learn about agent X and about agent to learn about agent X and about agent beats. And they can learn about the MOO beats. And they can learn about the MOO beats. And they can learn about the MOO at agenticai-learning.org. at agenticai-learning.org. at agenticai-learning.org. I'll put links in the show notes for all I'll put links in the show notes for all I'll put links in the show notes for all of this stuff. As we get ready to close, of this stuff. As we get ready to close, of this stuff. As we get ready to close, I want to ask you as a person on the I want to ask you as a person on the I want to ask you as a person on the forefront of this technology forefront of this technology forefront of this technology in your day-today, what agents are you in your day-today, what agents are you in your day-today, what agents are you using that are helping you be a better using that are helping you be a better using that are helping you be a better professor, be a better thinker, be a professor, be a better thinker, be a professor, be a better thinker, be a better teacher? Are you using commercial better teacher? Are you using commercial better teacher? Are you using commercial products that you just have a products that you just have a products that you just have a subscription to, or are you writing subscription to, or are you writing subscription to, or are you writing these things custom? Are you using these things custom? Are you using these things custom? Are you using cutting edge things? What's an agent cutting edge things? What's an agent cutting edge things? What's an agent expert using for their own agents?

  25. expert using for their own agents? expert using for their own agents? >> Um, that's a very good question. So, so >> Um, that's a very good question. So, so >> Um, that's a very good question. So, so I'm actually trying to develop some of I'm actually trying to develop some of I'm actually trying to develop some of my own agents to to better automate some my own agents to to better automate some my own agents to to better automate some of my own workflows. As you know, you of my own workflows. As you know, you of my own workflows. As you know, you mentioned like I do a lot of different mentioned like I do a lot of different mentioned like I do a lot of different things actually. I mean, they do take a things actually. I mean, they do take a things actually. I mean, they do take a lot of time and even when I have lot of time and even when I have lot of time and even when I have assistance and so on, it can still take assistance and so on, it can still take assistance and so on, it can still take a lot of time and so on and a lot of a lot of time and so on and a lot of a lot of time and so on and a lot of these things now can really be automated these things now can really be automated these things now can really be automated or like large like hugely helped with or like large like hugely helped with or like large like hugely helped with agents and so on. So, that's some of the agents and so on. So, that's some of the agents and so on. So, that's some of the things that I'm doing. things that I'm doing. things that I'm doing. as well. as well. as well. >> So I I like to I've heard in the space >> So I I like to I've heard in the space >> So I I like to I've heard in the space of robotics someone said that a robot of robotics someone said that a robot of robotics someone said that a robot should do things that are dull, dirty or should do things that are dull, dirty or should do things that are dull, dirty or dangerous and we use the term toil. So I dangerous and we use the term toil. So I dangerous and we use the term toil. So I assume you're trying to automate the assume you're trying to automate the assume you're trying to automate the boring stuff so that you can do the boring stuff so that you can do the boring stuff so that you can do the interesting fun thinking. interesting fun thinking. interesting fun thinking. >> Yes, absolutely. >> Yes, absolutely. >> Yes, absolutely. >> That's fantastic. Well, thank you so >> That's fantastic. Well, thank you so >> That's fantastic. Well, thank you so much Dr. Song for spending time with us much Dr. Song for spending time with us much Dr. Song for spending time with us today. today. today. >> Great. Thank you. Thank you so much for >> Great. Thank you. Thank you so much for >> Great. Thank you. Thank you so much for having me. We have been chatting with having me. We have been chatting with having me. We have been chatting with Don Song in association with the ACM Don Song in association with the ACM Don Song in association with the ACM Bitecast and this has been another Bitecast and this has been another Bitecast and this has been another episode of Hansel Minutes and we'll see episode of Hansel Minutes and we'll see episode of Hansel Minutes and we'll see you again next week.

Summary

The discussion centers on the opaque nature of Large Language Models (LLMs) and their role in agent systems, despite their power. It touches on issues like hallucinations and potential for jailbreaking. The practical takeaway is the importance of understanding and addressing these complexities in responsible AI development.

View original episode ↗