← Back
AI Engineer July 29, 2026 16m

We Vetted 2000 AI Skills Before They Reached Developers — Lucas Palma, Nubank

Read full transcript 12 segments
  1. Hello everyone. Good afternoon. Hello everyone. Good afternoon. Today I'm going to talk about how we Today I'm going to talk about how we Today I'm going to talk about how we vetted 2,000 AI skills before they reach vetted 2,000 AI skills before they reach vetted 2,000 AI skills before they reach a developers. a developers. a developers. But before I before of that, I'm Lucas But before I before of that, I'm Lucas But before I before of that, I'm Lucas Palma, but many people call me LP. I'm Palma, but many people call me LP. I'm Palma, but many people call me LP. I'm the product security manager at New the product security manager at New the product security manager at New Bank, the product security structures, Bank, the product security structures, Bank, the product security structures, uh structure that's within security, uh structure that's within security, uh structure that's within security, looking upon how we make code safe and looking upon how we make code safe and looking upon how we make code safe and supporting engineers, product managers supporting engineers, product managers supporting engineers, product managers and everybody to making our products and everybody to making our products and everybody to making our products safer. I have uh over a decade of safer. I have uh over a decade of safer. I have uh over a decade of experience in financial services experience in financial services experience in financial services engineering background also a lot of engineering background also a lot of engineering background also a lot of years working here at security and a years working here at security and a years working here at security and a close relationship with the part that I close relationship with the part that I close relationship with the part that I love which is innovation. love which is innovation. love which is innovation. So before beginning I believe I want to So before beginning I believe I want to So before beginning I believe I want to bring to you uh why are we here. So one bring to you uh why are we here. So one bring to you uh why are we here. So one thing that's important for all of you to thing that's important for all of you to thing that's important for all of you to understand that the understand that the understand that the now that we are using AI everywhere even now that we are using AI everywhere even now that we are using AI everywhere even though even with uh coding one thing though even with uh coding one thing though even with uh coding one thing that uh is important that the AI skills that uh is important that the AI skills that uh is important that the AI skills are being part of the developer workflow are being part of the developer workflow are being part of the developer workflow and that's this might bring some risks and that's this might bring some risks and that's this might bring some risks because because because although they look like configuration although they look like configuration although they look like configuration they behave like supply chain dependence they behave like supply chain dependence they behave like supply chain dependence like uh for example libraries and like uh for example libraries and like uh for example libraries and others. So what we made here was to

  2. others. So what we made here was to others. So what we made here was to build a security review system in order build a security review system in order build a security review system in order to check if these skills were safe or to check if these skills were safe or to check if these skills were safe or not to be used before deploying them. So not to be used before deploying them. So not to be used before deploying them. So the lesson that I want to bring you here the lesson that I want to bring you here the lesson that I want to bring you here by the end of this presentation is that by the end of this presentation is that by the end of this presentation is that we should be protecting the whole we should be protecting the whole we should be protecting the whole workflow not only the code that's being workflow not only the code that's being workflow not only the code that's being generated. generated. generated. All right. All right. All right. So what I mean about the supply chain So what I mean about the supply chain So what I mean about the supply chain part is that uh traditionally the supply part is that uh traditionally the supply part is that uh traditionally the supply chain has uh package containers, models chain has uh package containers, models chain has uh package containers, models and so on. But now in the AI era, it and so on. But now in the AI era, it and so on. But now in the AI era, it doesn't have only that. It still have doesn't have only that. It still have doesn't have only that. It still have the traditional part, but it will it the traditional part, but it will it the traditional part, but it will it also includes skills, plugins, MCP also includes skills, plugins, MCP also includes skills, plugins, MCP servers, agent rules and much more servers, agent rules and much more servers, agent rules and much more things to be acting as supply chain things to be acting as supply chain things to be acting as supply chain and where AI skill fits into this. Uh and where AI skill fits into this. Uh and where AI skill fits into this. Uh I believe that before I go into that I believe that before I go into that I believe that before I go into that it's important for everybody be on the it's important for everybody be on the it's important for everybody be on the same page on what is an AI skill. So an same page on what is an AI skill. So an same page on what is an AI skill. So an AI skill has AI skill has AI skill has there's normally the developer is using there's normally the developer is using there's normally the developer is using AI tools in order to generate an output AI tools in order to generate an output AI tools in order to generate an output which will be code most of the case and which will be code most of the case and which will be code most of the case and within this AI tool there are a bunch of within this AI tool there are a bunch of within this AI tool there are a bunch of things that can be embedded. One of them things that can be embedded. One of them things that can be embedded. One of them are the AI skill. So with this skill we are the AI skill. So with this skill we are the AI skill. So with this skill we can have a capability to a model or to

  3. can have a capability to a model or to can have a capability to a model or to an agent uh bundling some instructions an agent uh bundling some instructions an agent uh bundling some instructions some context in order to have better some context in order to have better some context in order to have better guidance over what it can be done. But guidance over what it can be done. But guidance over what it can be done. But there is also an impact over that there is also an impact over that there is also an impact over that because somebody can create their own because somebody can create their own because somebody can create their own skill and share with others. So when we skill and share with others. So when we skill and share with others. So when we do that this first person is guiding do that this first person is guiding do that this first person is guiding over the code that's being generated by over the code that's being generated by over the code that's being generated by the other person and then that can be the other person and then that can be the other person and then that can be dangerous dangerous dangerous and since we are here talking in the AI and since we are here talking in the AI and since we are here talking in the AI in finance track it's also important for in finance track it's also important for in finance track it's also important for us to understand that we are in a us to understand that we are in a us to understand that we are in a regulated environment. So from one side regulated environment. So from one side regulated environment. So from one side there are are developers wanting better there are are developers wanting better there are are developers wanting better faster coding more context to have less faster coding more context to have less faster coding more context to have less repetitive work but but from the other repetitive work but but from the other repetitive work but but from the other side even more because of the regulate side even more because of the regulate side even more because of the regulate part we need to be aware of the part we need to be aware of the part we need to be aware of the auditability of looking upon credentials auditability of looking upon credentials auditability of looking upon credentials safety by default and many other safety by default and many other safety by default and many other security aspects security aspects security aspects and keeping that balances is hard, and keeping that balances is hard, and keeping that balances is hard, right? So, some people might say like right? So, some people might say like right? So, some people might say like are AI skills dangerous?

  4. are AI skills dangerous? are AI skills dangerous? So, So, So, I brought here a few examples of what do I brought here a few examples of what do I brought here a few examples of what do I mean by AI skills being dangerous? So I mean by AI skills being dangerous? So I mean by AI skills being dangerous? So first uh one thing that can happen is first uh one thing that can happen is first uh one thing that can happen is that when people are describing what that when people are describing what that when people are describing what they skill can or cannot do it can it they skill can or cannot do it can it they skill can or cannot do it can it can ask for it to retrieve a token or can ask for it to retrieve a token or can ask for it to retrieve a token or something and it will begin using that something and it will begin using that something and it will begin using that token hardcoded which will go to logs token hardcoded which will go to logs token hardcoded which will go to logs and so on and it can generate a data and so on and it can generate a data and so on and it can generate a data leak in the future. Another thing that leak in the future. Another thing that leak in the future. Another thing that can happen is also the person to h can happen is also the person to h can happen is also the person to h instruct the AI to use shell comments instruct the AI to use shell comments instruct the AI to use shell comments and then this skill will be used by and then this skill will be used by and then this skill will be used by another person and when they use on another person and when they use on another person and when they use on their shell a lot of dangerous things their shell a lot of dangerous things their shell a lot of dangerous things that that can happen and a lot of files that that can happen and a lot of files that that can happen and a lot of files be modified and so on and there's also be modified and so on and there's also be modified and so on and there's also permissions. So depending on how the permissions. So depending on how the permissions. So depending on how the skill was configured, it might have skill was configured, it might have skill was configured, it might have excessive permissions much more than excessive permissions much more than excessive permissions much more than what was needed and what was needed and what was needed and even a typo can make some dangerous even a typo can make some dangerous even a typo can make some dangerous stuff depending on who is using that stuff depending on who is using that stuff depending on who is using that such skill.

  5. So first thing first what we did So first thing first what we did initially is that how do we share skills initially is that how do we share skills initially is that how do we share skills among ourselves how the engineers would among ourselves how the engineers would among ourselves how the engineers would be sharing the skills. So uh we went be sharing the skills. So uh we went be sharing the skills. So uh we went through the marketplace solution. So the through the marketplace solution. So the through the marketplace solution. So the skills are being canonically shared skills are being canonically shared skills are being canonically shared among marketplace with the plugins among marketplace with the plugins among marketplace with the plugins included the skills among them. So it's included the skills among them. So it's included the skills among them. So it's a internal marketplace where people can a internal marketplace where people can a internal marketplace where people can discover new skills and that's our discover new skills and that's our discover new skills and that's our boundary where we are trying to make it boundary where we are trying to make it boundary where we are trying to make it safer. So what happens is that when safer. So what happens is that when safer. So what happens is that when someone creates an skill it uh will open someone creates an skill it uh will open someone creates an skill it uh will open the pull request and normally it will go the pull request and normally it will go the pull request and normally it will go to the marketplace but we made a step to the marketplace but we made a step to the marketplace but we made a step before that like a CI step where we before that like a CI step where we before that like a CI step where we created a tool that's called skill created a tool that's called skill created a tool that's called skill vector and this is this what this tool vector and this is this what this tool vector and this is this what this tool does is to check if this skill is safe does is to check if this skill is safe does is to check if this skill is safe or not to be used uh using a lot of or not to be used uh using a lot of or not to be used uh using a lot of assessments that I will bring it here assessments that I will bring it here assessments that I will bring it here and also classify those risks and and also classify those risks and and also classify those risks and request remediation and so on.

  6. request remediation and so on. request remediation and so on. So what skill vector does in So what skill vector does in So what skill vector does in a single page is that when a skill is a single page is that when a skill is a single page is that when a skill is created or changed not on the during the created or changed not on the during the created or changed not on the during the creation phase one thing that's creation phase one thing that's creation phase one thing that's important is that the engineers are able important is that the engineers are able important is that the engineers are able to use it locally and also be iterating to use it locally and also be iterating to use it locally and also be iterating until the skill is being considered safe until the skill is being considered safe until the skill is being considered safe before uploaded it and after them upload before uploaded it and after them upload before uploaded it and after them upload the skill. We also the skill. We also the skill. We also runs it again because we can ensure that runs it again because we can ensure that runs it again because we can ensure that the engineer has run locally or has run the engineer has run locally or has run the engineer has run locally or has run the most updated version. So we also be the most updated version. So we also be the most updated version. So we also be scanning that after the upload. And then scanning that after the upload. And then scanning that after the upload. And then we have some determinate checks for the we have some determinate checks for the we have some determinate checks for the uh easiest parts to check some uh easy uh easiest parts to check some uh easy uh easiest parts to check some uh easy risks using regular regular expressions risks using regular regular expressions risks using regular regular expressions and so on. and so on. and so on. After that when we check that we need After that when we check that we need After that when we check that we need better context we then use LLM. Uh it's better context we then use LLM. Uh it's better context we then use LLM. Uh it's important to have this hybrid approach important to have this hybrid approach important to have this hybrid approach with LLM checking the the context but with LLM checking the the context but with LLM checking the the context but also with the determinist because you also with the determinist because you also with the determinist because you know how LLM is depending on the know how LLM is depending on the know how LLM is depending on the temperature that was set. Sometimes it temperature that was set. Sometimes it temperature that was set. Sometimes it will check that it's a risk sometimes it will check that it's a risk sometimes it will check that it's a risk sometimes it might not. And then all of these might not. And then all of these might not. And then all of these findings are reporting the PR that was findings are reporting the PR that was findings are reporting the PR that was open to upload the skill. So it will open to upload the skill. So it will open to upload the skill. So it will improve the usability since the engineer improve the usability since the engineer improve the usability since the engineer will have the in the same PR what has to will have the in the same PR what has to will have the in the same PR what has to be changed before uploading the skill.

  7. be changed before uploading the skill. be changed before uploading the skill. And another good thing that we made And another good thing that we made And another good thing that we made that's important is to have a serif with that's important is to have a serif with that's important is to have a serif with all of these so it can be consumed by all of these so it can be consumed by all of these so it can be consumed by our security tools as well and generate our security tools as well and generate our security tools as well and generate a report on the risks and be part of our a report on the risks and be part of our a report on the risks and be part of our vulnerability management program. vulnerability management program. vulnerability management program. So depending on the severity, depending So depending on the severity, depending So depending on the severity, depending on the policy, on the policy, on the policy, uh the skill can require some uh the skill can require some uh the skill can require some remediation, can be blocked, remediation, can be blocked, remediation, can be blocked, all of this before the marketplace all of this before the marketplace all of this before the marketplace distribution. distribution. distribution. So there is the local scan, the p So there is the local scan, the p So there is the local scan, the p request, the determinist scanner, then request, the determinist scanner, then request, the determinist scanner, then there is the LLM review, PR feedback, there is the LLM review, PR feedback, there is the LLM review, PR feedback, serif, and then the decision. Will we serif, and then the decision. Will we serif, and then the decision. Will we use it? We will allow it, will we allow use it? We will allow it, will we allow use it? We will allow it, will we allow it? but it requires remediation and so it? but it requires remediation and so it? but it requires remediation and so on. on. on. A few examples of what we have ex A few examples of what we have ex A few examples of what we have ex scanned here. It's uh a non-exhaustive scanned here. It's uh a non-exhaustive scanned here. It's uh a non-exhaustive list. So we are looking upon if there list. So we are looking upon if there list. So we are looking upon if there are some unsafe instructions. If there are some unsafe instructions. If there are some unsafe instructions. If there are some drift be within the behavior are some drift be within the behavior are some drift be within the behavior that the agent has if there are some that the agent has if there are some that the agent has if there are some destructive shell comments that I destructive shell comments that I destructive shell comments that I commented earlier if there are some file commented earlier if there are some file commented earlier if there are some file modifications that shouldn't be there.

  8. modifications that shouldn't be there. modifications that shouldn't be there. uh credential requests, how are they uh credential requests, how are they uh credential requests, how are they being done? Some data being exposed uh being done? Some data being exposed uh being done? Some data being exposed uh unintentionally, unintentionally, unintentionally, if there are permissions that are over if there are permissions that are over if there are permissions that are over broad, risky, MCP usage and much more. broad, risky, MCP usage and much more. broad, risky, MCP usage and much more. These these are the main ones. These these are the main ones. These these are the main ones. And so getting back to the title, we And so getting back to the title, we And so getting back to the title, we have scanned on that over 2,000 skills. have scanned on that over 2,000 skills. have scanned on that over 2,000 skills. uh now there is much more than that but uh now there is much more than that but uh now there is much more than that but this is the baseline that I brought for this is the baseline that I brought for this is the baseline that I brought for you on this presentation you on this presentation you on this presentation uh inside this we have identified uh uh inside this we have identified uh uh inside this we have identified uh more than 1,000 and half uh more than 1,000 and half uh more than 1,000 and half uh risks. So not that 1,00 skills had risk risks. So not that 1,00 skills had risk risks. So not that 1,00 skills had risk because a single skill can has many because a single skill can has many because a single skill can has many risks but these were the total risks risks but these were the total risks risks but these were the total risks that we identified over uh this amount that we identified over uh this amount that we identified over uh this amount of skills and 1,000 of them were of skills and 1,000 of them were of skills and 1,000 of them were probably probably probably remediated right after remediated right after remediated right after and there are few of them that were and there are few of them that were and there are few of them that were really hky that we were able to block really hky that we were able to block really hky that we were able to block before going to the marketplace.

  9. before going to the marketplace. before going to the marketplace. So we also had made a b uh historical So we also had made a b uh historical So we also had made a b uh historical scan looking upon the scan looking upon the scan looking upon the skills that were skills that were skills that were created before the skill v implemented. created before the skill v implemented. created before the skill v implemented. Uh over there we were able to identify Uh over there we were able to identify Uh over there we were able to identify new uh risks as well and put it them new uh risks as well and put it them new uh risks as well and put it them into the vulnerability management into the vulnerability management into the vulnerability management program so [snorts] it can be could be program so [snorts] it can be could be program so [snorts] it can be could be remediated. A few lessons that I want to bring here A few lessons that I want to bring here as well. So as well. So as well. So what things that work well is having what things that work well is having what things that work well is having both the the terminist scanners for both the the terminist scanners for both the the terminist scanners for non-risk patterns but also LLM review non-risk patterns but also LLM review non-risk patterns but also LLM review for uh behavior checking upon the for uh behavior checking upon the for uh behavior checking upon the destructive comments uh looking upon the destructive comments uh looking upon the destructive comments uh looking upon the credentials checks as well having the credentials checks as well having the credentials checks as well having the output in serif and adding comments on output in serif and adding comments on output in serif and adding comments on PRs and things that needed improvement. PRs and things that needed improvement. PRs and things that needed improvement. And we worked during the process were And we worked during the process were And we worked during the process were also there were some risks like comments also there were some risks like comments also there were some risks like comments that we were treating equally but that we were treating equally but that we were treating equally but depending on the comment it can be more depending on the comment it can be more depending on the comment it can be more or less risky. Also some signals that or less risky. Also some signals that or less risky. Also some signals that were weak and didn't have much context were weak and didn't have much context were weak and didn't have much context that were uh more troublesome than that were uh more troublesome than that were uh more troublesome than helpful.

  10. helpful. helpful. There is also the prompt level ask for There is also the prompt level ask for There is also the prompt level ask for confirmation. I there's a next slide confirmation. I there's a next slide confirmation. I there's a next slide about that that I will go deeper. That's about that that I will go deeper. That's about that that I will go deeper. That's an important one. Also, an important one. Also, an important one. Also, uh there were some warnings that seemed uh there were some warnings that seemed uh there were some warnings that seemed uh harmless, but only if it was running uh harmless, but only if it was running uh harmless, but only if it was running locally. If there were going to locally. If there were going to locally. If there were going to production, then they could be impactful production, then they could be impactful production, then they could be impactful and that we had also to look up on that. and that we had also to look up on that. and that we had also to look up on that. Uh if the finding had hadn't some clear Uh if the finding had hadn't some clear Uh if the finding had hadn't some clear guidance was troublesome as well. And guidance was troublesome as well. And guidance was troublesome as well. And last but not least, we know that other last but not least, we know that other last but not least, we know that other people could create other marketplace. people could create other marketplace. people could create other marketplace. So how can we proactively scan check So how can we proactively scan check So how can we proactively scan check there is a new marketplace and put skill there is a new marketplace and put skill there is a new marketplace and put skill vector into it as well. vector into it as well. vector into it as well. So regarding the prompt level that's So regarding the prompt level that's So regarding the prompt level that's something that's important for you to something that's important for you to something that's important for you to know people sometimes will add the know people sometimes will add the know people sometimes will add the instruction like you need to ask for instruction like you need to ask for instruction like you need to ask for confirmation but the AI may ask confirmation but the AI may ask confirmation but the AI may ask confirmation for itself. So from your confirmation for itself. So from your confirmation for itself. So from your perspective there is a human in the loop perspective there is a human in the loop perspective there is a human in the loop but for the AI perspective there is has but for the AI perspective there is has but for the AI perspective there is has been a confirmation and that's okay been a confirmation and that's okay been a confirmation and that's okay another has confirmed then let's go so another has confirmed then let's go so another has confirmed then let's go so that's something that we were scanning that's something that we were scanning that's something that we were scanning as well looking up on having proper as well looking up on having proper as well looking up on having proper human in the loop looking the tool human in the loop looking the tool human in the loop looking the tool that's executing if it's going through that's executing if it's going through that's executing if it's going through the approval gates and so on having the approval gates and so on having the approval gates and so on having hooks hooks hooks and within this as I said it's uh and within this as I said it's uh and within this as I said it's uh plug-in marketplace skill is one among

  11. plug-in marketplace skill is one among plug-in marketplace skill is one among many things that there is into that. So many things that there is into that. So many things that there is into that. So there are things that we can reuse from there are things that we can reuse from there are things that we can reuse from this lesson. So for example, this lesson. So for example, this lesson. So for example, treating these as supply chain is treating these as supply chain is treating these as supply chain is important. Reviewing what's being important. Reviewing what's being important. Reviewing what's being uploaded to the marketplace before goes uploaded to the marketplace before goes uploaded to the marketplace before goes there. letting developers to run these there. letting developers to run these there. letting developers to run these checks locally, enforcing these checks checks locally, enforcing these checks checks locally, enforcing these checks that are being run locally also in the that are being run locally also in the that are being run locally also in the CI having the termination checks CI having the termination checks CI having the termination checks together with the LLM checks and looking together with the LLM checks and looking together with the LLM checks and looking upon dangerous actions and prompting upon dangerous actions and prompting upon dangerous actions and prompting uh and having enforcement when they uh and having enforcement when they uh and having enforcement when they happen. So next steps over here is that I'm So next steps over here is that I'm talking a lot about skills here but a talking a lot about skills here but a talking a lot about skills here but a lot of these as I said could be applied lot of these as I said could be applied lot of these as I said could be applied to plugins to MCP servers rules hooks. to plugins to MCP servers rules hooks. to plugins to MCP servers rules hooks. So all of this that I'm saying here we So all of this that I'm saying here we So all of this that I'm saying here we also have the MCP vector the rules also have the MCP vector the rules also have the MCP vector the rules checks and so on that's also applicable checks and so on that's also applicable checks and so on that's also applicable here but with different risks uh having here but with different risks uh having here but with different risks uh having also different gates depending on also different gates depending on also different gates depending on policies that were implemented depending policies that were implemented depending policies that were implemented depending on the marketplace as well have some on the marketplace as well have some on the marketplace as well have some enforcements on tool level here enforcements on tool level here enforcements on tool level here enforcing that there are audit logs enforcing that there are audit logs enforcing that there are audit logs trusted gateways and so on and also last trusted gateways and so on and also last trusted gateways and so on and also last but not least Having but not least Having but not least Having the trusted trusted AI marketplace is the trusted trusted AI marketplace is the trusted trusted AI marketplace is very important. So we can have a very important. So we can have a very important. So we can have a canonical way to scan and share knowing

  12. canonical way to scan and share knowing canonical way to scan and share knowing that then are being safe and that's not that then are being safe and that's not that then are being safe and that's not about only about the skills that are about only about the skills that are about only about the skills that are being created by people but it also being created by people but it also being created by people but it also includes the third party skills or includes the third party skills or includes the third party skills or plugins and so on. So if someone plugins and so on. So if someone plugins and so on. So if someone downloads something and wants to use downloads something and wants to use downloads something and wants to use it's important to upload it on the it's important to upload it on the it's important to upload it on the marketplace. So all of this scanning can marketplace. So all of this scanning can marketplace. So all of this scanning can be done and check if it's safe or not to be done and check if it's safe or not to be done and check if it's safe or not to be used and also that way allow other be used and also that way allow other be used and also that way allow other people to use in a safe way. And that's it. Uh I'm sharing here my And that's it. Uh I'm sharing here my contact. There's my link in profile. If contact. There's my link in profile. If contact. There's my link in profile. If anybody wants to contact talk more about anybody wants to contact talk more about anybody wants to contact talk more about that the QR code will bring you to my that the QR code will bring you to my that the QR code will bring you to my profile. If you don't want to type, no profile. If you don't want to type, no profile. If you don't want to type, no problem at all. And I hope you you've problem at all. And I hope you you've problem at all. And I hope you you've enjoyed the talk.

Summary

This tech talk focuses on vetting AI skills before they're integrated into developer workflows, drawing parallels to supply chain dependencies. The speaker emphasizes the need to protect the entire workflow, not just generated code, by implementing a security review system for these AI skills. The key takeaway is to ensure the safety of the entire development pipeline by rigorously assessing all AI components.

View original episode ↗