← Back
Aaron Zisk July 10, 2026 11m

Apple’s Hidden AI Model… The Speed Test Apple Never Showed

Read full transcript 9 segments
  1. There's a new AI hidden inside Mac OS There's a new AI hidden inside Mac OS right now. Apple ships it. It runs right right now. Apple ships it. It runs right right now. Apple ships it. It runs right in your terminal and almost nobody knows in your terminal and almost nobody knows in your terminal and almost nobody knows is there. So, I had to know how fast is is there. So, I had to know how fast is is there. So, I had to know how fast is this thing really. I'm going to try it this thing really. I'm going to try it this thing really. I'm going to try it on this Mac Mini that I just upgraded to on this Mac Mini that I just upgraded to on this Mac Mini that I just upgraded to Mac OS 27, the beta version. Keep that Mac OS 27, the beta version. Keep that Mac OS 27, the beta version. Keep that in mind. It's beta. And how much faster in mind. It's beta. And how much faster in mind. It's beta. And how much faster is this tool going to be running it on is this tool going to be running it on is this tool going to be running it on an M3 Ultra Mac Studio instead of the an M3 Ultra Mac Studio instead of the an M3 Ultra Mac Studio instead of the Mac Mini? Let's find out. Quick detour Mac Mini? Let's find out. Quick detour Mac Mini? Let's find out. Quick detour before we get back to the video. I want before we get back to the video. I want before we get back to the video. I want to show you this thing called Super Nori to show you this thing called Super Nori to show you this thing called Super Nori and it's one of the first proactive and it's one of the first proactive and it's one of the first proactive family AI agents I've seen. And for a family AI agents I've seen. And for a family AI agents I've seen. And for a tech audience, I think the important tech audience, I think the important tech audience, I think the important frame here is that it's meant to run in frame here is that it's meant to run in frame here is that it's meant to run in the background around a real family the background around a real family the background around a real family context. A lot of helpful AI these days context. A lot of helpful AI these days context. A lot of helpful AI these days still depends on you noticing the issue still depends on you noticing the issue still depends on you noticing the issue first, then writing the prompt, then first, then writing the prompt, then first, then writing the prompt, then managing and following up. Even when the managing and following up. Even when the managing and following up. Even when the tools are good, they still assume the tools are good, they still assume the tools are good, they still assume the right person catches the problem in time right person catches the problem in time right person catches the problem in time and manually kicks everything off. So, and manually kicks everything off. So, and manually kicks everything off. So, now there's Super Nori from the team now there's Super Nori from the team now there's Super Nori from the team behind Nori, which is already used by behind Nori, which is already used by behind Nori, which is already used by more than 200,000 families. Instead of more than 200,000 families. Instead of more than 200,000 families. Instead of waiting for a prompt, it watches for waiting for a prompt, it watches for waiting for a prompt, it watches for situations that need attention, and it situations that need attention, and it situations that need attention, and it surfaces the next step while there's surfaces the next step while there's surfaces the next step while there's still time to do something about it. So, still time to do something about it. So, still time to do something about it. So, at 6:00 a.m., traffic is already a mess.

  2. at 6:00 a.m., traffic is already a mess. at 6:00 a.m., traffic is already a mess. It can catch that before you do and ask It can catch that before you do and ask It can catch that before you do and ask you if you want to book the ride. And you if you want to book the ride. And you if you want to book the ride. And later on in the evening, if a last later on in the evening, if a last later on in the evening, if a last minute calendar change is about to blow minute calendar change is about to blow minute calendar change is about to blow up your anniversary dinner, it doesn't up your anniversary dinner, it doesn't up your anniversary dinner, it doesn't just stop at noticing the problem. it just stop at noticing the problem. it just stop at noticing the problem. it can start finding backup options before can start finding backup options before can start finding backup options before the night falls apart, even making the night falls apart, even making the night falls apart, even making another reservation for you. What makes another reservation for you. What makes another reservation for you. What makes this interesting to me is that it's a this interesting to me is that it's a this interesting to me is that it's a cleaner example of what people mean when cleaner example of what people mean when cleaner example of what people mean when they say agent instead of just they say agent instead of just they say agent instead of just assistant. It's built to ask before assistant. It's built to ask before assistant. It's built to ask before acting, which matters when you're acting, which matters when you're acting, which matters when you're talking about realworld decisions. So talking about realworld decisions. So talking about realworld decisions. So like your coding agent takes care of like your coding agent takes care of like your coding agent takes care of your code base and proactively finds your code base and proactively finds your code base and proactively finds issues and bugs. Supernori is the agent issues and bugs. Supernori is the agent issues and bugs. Supernori is the agent that looks after you and your family. that looks after you and your family. that looks after you and your family. You can check out Supernori in the link You can check out Supernori in the link You can check out Supernori in the link in the description below and join the in the description below and join the in the description below and join the weight list from there. Okay, let me weight list from there. Okay, let me weight list from there. Okay, let me back up for a second. So, Mac OS 27 back up for a second. So, Mac OS 27 back up for a second. So, Mac OS 27 Golden Gate quietly ships this thing Golden Gate quietly ships this thing Golden Gate quietly ships this thing called FM. It's a built-in CLI that runs called FM. It's a built-in CLI that runs called FM. It's a built-in CLI that runs actual large language model models actual large language model models actual large language model models model, model, model, not sure. I think it's just one model, not sure. I think it's just one model, not sure. I think it's just one model, but there's two ways of running. There's but there's two ways of running. There's but there's two ways of running. There's a cloud version. I'm I'm going to get a cloud version. I'm I'm going to get a cloud version. I'm I'm going to get into it. Hold on a second. Basically, as into it. Hold on a second. Basically, as into it. Hold on a second. Basically, as soon as you have Mac OS 27 installed, soon as you have Mac OS 27 installed, soon as you have Mac OS 27 installed, you just have it right there in your you just have it right there in your you just have it right there in your terminal available. There's no need to terminal available. There's no need to terminal available. There's no need to install things like Olama or LM Studio install things like Olama or LM Studio install things like Olama or LM Studio or MLX. Well, hold that thought. There or MLX. Well, hold that thought. There or MLX. Well, hold that thought. There might still be a need, but for a might still be a need, but for a might still be a need, but for a developer, that's kind of wild. It's developer, that's kind of wild. It's developer, that's kind of wild. It's basically a free private AI endpoint basically a free private AI endpoint basically a free private AI endpoint just sitting right there inside the OS just sitting right there inside the OS just sitting right there inside the OS and it's all local. You can script it.

  3. and it's all local. You can script it. and it's all local. You can script it. You can build tools on top of it. You can build tools on top of it. You can build tools on top of it. Whatever you want to do. So, I had three Whatever you want to do. So, I had three Whatever you want to do. So, I had three questions. Can you actually use this questions. Can you actually use this questions. Can you actually use this thing? How fast is it? And the big one, thing? How fast is it? And the big one, thing? How fast is it? And the big one, the one I actually care about, does the one I actually care about, does the one I actually care about, does throwing more expensive hardware at it throwing more expensive hardware at it throwing more expensive hardware at it make it any faster? And to answer any of make it any faster? And to answer any of make it any faster? And to answer any of that, I needed some real numbers. So, I that, I needed some real numbers. So, I that, I needed some real numbers. So, I went to the benchmark. Easy, right? went to the benchmark. Easy, right? went to the benchmark. Easy, right? Right away, I hit a wall. I grabbed my Right away, I hit a wall. I grabbed my Right away, I hit a wall. I grabbed my favorite bencher for LLMs, Llama Beni. I favorite bencher for LLMs, Llama Beni. I favorite bencher for LLMs, Llama Beni. I pointed it at FM and it told me Apple's pointed it at FM and it told me Apple's pointed it at FM and it told me Apple's little built-in model was doing 600,000 little built-in model was doing 600,000 little built-in model was doing 600,000 tokens a second. That's for prompt tokens a second. That's for prompt tokens a second. That's for prompt processing and decode was 22,000. Come processing and decode was 22,000. Come processing and decode was 22,000. Come on, really? That's definitely not a on, really? That's definitely not a on, really? That's definitely not a benchmark. That's a typo. Here's the benchmark. That's a typo. Here's the benchmark. That's a typo. Here's the thing. It's not really the tool's fault. thing. It's not really the tool's fault. thing. It's not really the tool's fault. Apple server, Apple being Apple, they Apple server, Apple being Apple, they Apple server, Apple being Apple, they have to be a little bit different. They have to be a little bit different. They have to be a little bit different. They don't exactly comply with the Open AI don't exactly comply with the Open AI don't exactly comply with the Open AI standard all the way. Some things do standard all the way. Some things do standard all the way. Some things do like /models, but of course, they had to like /models, but of course, they had to like /models, but of course, they had to have their own thing. And Llama Beni is have their own thing. And Llama Beni is have their own thing. And Llama Beni is designed to work with OpenAI compatible designed to work with OpenAI compatible designed to work with OpenAI compatible endpoints. So, Apple's server doesn't endpoints. So, Apple's server doesn't endpoints. So, Apple's server doesn't just hand back timing numbers these just hand back timing numbers these just hand back timing numbers these benchmarks expect. Bottom line, nothing benchmarks expect. Bottom line, nothing benchmarks expect. Bottom line, nothing off the shelf can actually measure this off the shelf can actually measure this off the shelf can actually measure this thing, except maybe Apple's own thing, except maybe Apple's own thing, except maybe Apple's own proprietary internal stuff. So, fine. If proprietary internal stuff. So, fine. If proprietary internal stuff. So, fine. If nothing out there can do it, I'll build nothing out there can do it, I'll build nothing out there can do it, I'll build my own. So, I created Apple FM bench.

  4. my own. So, I created Apple FM bench. my own. So, I created Apple FM bench. But hold on, don't go uh running out But hold on, don't go uh running out But hold on, don't go uh running out there yet and grabbing it. I'll explain. there yet and grabbing it. I'll explain. there yet and grabbing it. I'll explain. I will link to it in the description, of I will link to it in the description, of I will link to it in the description, of course. Now, here are the actual numbers course. Now, here are the actual numbers course. Now, here are the actual numbers you get from this. About,50 you get from this. About,50 you get from this. About,50 for prompt processing and about 52 for for prompt processing and about 52 for for prompt processing and about 52 for decode. That's the generation stage. All decode. That's the generation stage. All decode. That's the generation stage. All on device. And I'm talking about the Mac on device. And I'm talking about the Mac on device. And I'm talking about the Mac Mini right now. So FM has two brains Mini right now. So FM has two brains Mini right now. So FM has two brains though. There's the local one running on though. There's the local one running on though. There's the local one running on your Mac and then there's the Apple your Mac and then there's the Apple your Mac and then there's the Apple cloud one uh called private cloud cloud one uh called private cloud cloud one uh called private cloud compute which is going to the cloud. So compute which is going to the cloud. So compute which is going to the cloud. So private local cloud I'm not sure okay private local cloud I'm not sure okay private local cloud I'm not sure okay how to explain that. Basically the FM how to explain that. Basically the FM how to explain that. Basically the FM tool it does both and you can specify tool it does both and you can specify tool it does both and you can specify whether you want to use the cloud one or whether you want to use the cloud one or whether you want to use the cloud one or the local one. Cool. So I tested both the local one. Cool. So I tested both the local one. Cool. So I tested both and the cloud is about three times and the cloud is about three times and the cloud is about three times faster than the local one. That's good faster than the local one. That's good faster than the local one. That's good to know. But really, the local model is to know. But really, the local model is to know. But really, the local model is the one I really care about the most the one I really care about the most the one I really care about the most because that's the one that's running on because that's the one that's running on because that's the one that's running on your hardware. And that got me thinking, your hardware. And that got me thinking, your hardware. And that got me thinking, if it runs on my hardware, wouldn't a if it runs on my hardware, wouldn't a if it runs on my hardware, wouldn't a beefier Mac just run it faster? So, on beefier Mac just run it faster? So, on beefier Mac just run it faster? So, on one side, on the left, I got the Mac one side, on the left, I got the Mac one side, on the left, I got the Mac Mini, this one. And on the right, I got Mini, this one. And on the right, I got Mini, this one. And on the right, I got the Mac Studio, that one. Now, the M4 the Mac Studio, that one. Now, the M4 the Mac Studio, that one. Now, the M4 Pro, which is what this is, it has 273 Pro, which is what this is, it has 273 Pro, which is what this is, it has 273 GB per second memory bandwidth. And GB per second memory bandwidth. And GB per second memory bandwidth. And memory bandwidth is very important when memory bandwidth is very important when memory bandwidth is very important when it comes to generating tokens. the it comes to generating tokens. the it comes to generating tokens. the decode stage. It also has a pretty nice decode stage. It also has a pretty nice decode stage. It also has a pretty nice little U GPU inside for processing.

  5. little U GPU inside for processing. little U GPU inside for processing. That's the prompt processing stage. By That's the prompt processing stage. By That's the prompt processing stage. By the way, if you don't know what I'm the way, if you don't know what I'm the way, if you don't know what I'm talking about, I explain this in a bunch talking about, I explain this in a bunch talking about, I explain this in a bunch of other my videos. Inference, when of other my videos. Inference, when of other my videos. Inference, when you're doing the LLM stuff, when you're you're doing the LLM stuff, when you're you're doing the LLM stuff, when you're using them, there's two stages. One is using them, there's two stages. One is using them, there's two stages. One is prompt processing. That's the first prompt processing. That's the first prompt processing. That's the first stage. And the second stage is the stage. And the second stage is the stage. And the second stage is the actual generation of tokens. The first actual generation of tokens. The first actual generation of tokens. The first stage is also called prefill. The second stage is also called prefill. The second stage is also called prefill. The second stage is also called decode. Now, you're stage is also called decode. Now, you're stage is also called decode. Now, you're all caught up, but you can still go all caught up, but you can still go all caught up, but you can still go watch my videos. Anyway, this one has watch my videos. Anyway, this one has watch my videos. Anyway, this one has the memory bandwidth of 819. the memory bandwidth of 819. the memory bandwidth of 819. It's the M3 Ultra. It's the biggest It's the M3 Ultra. It's the biggest It's the M3 Ultra. It's the biggest right now that Apple has, the fastest right now that Apple has, the fastest right now that Apple has, the fastest memory bandwidth. 819 versus 273. That memory bandwidth. 819 versus 273. That memory bandwidth. 819 versus 273. That should be much, much faster. So, I should be much, much faster. So, I should be much, much faster. So, I figured it was going to bury the Mac figured it was going to bury the Mac figured it was going to bury the Mac Mini. So, here's the M4 Pro Mac Mini. Mini. So, here's the M4 Pro Mac Mini. Mini. So, here's the M4 Pro Mac Mini. Prompt processing chart is on the left. Prompt processing chart is on the left. Prompt processing chart is on the left. Decode chart is on the right. Let's look Decode chart is on the right. Let's look Decode chart is on the right. Let's look at the ondevice stuff. All right. So, at the ondevice stuff. All right. So, at the ondevice stuff. All right. So, for on device 1,142 for on device 1,142 for on device 1,142 and on device for decode 56. Now, if we and on device for decode 56. Now, if we and on device for decode 56. Now, if we take a look at M3 Ultra Max Studio, what take a look at M3 Ultra Max Studio, what take a look at M3 Ultra Max Studio, what the Huh? I'm just kidding. I I already the Huh? I'm just kidding. I I already the Huh? I'm just kidding. I I already knew this cuz I already made the chart. knew this cuz I already made the chart. knew this cuz I already made the chart. So, my reaction was kind of lame, I So, my reaction was kind of lame, I So, my reaction was kind of lame, I guess. But how about your reaction? Is guess. But how about your reaction? Is guess. But how about your reaction? Is this what you expected? Cuz this is the this what you expected? Cuz this is the this what you expected? Cuz this is the first time I saw this, I was like, what?

  6. first time I saw this, I was like, what? first time I saw this, I was like, what? We got the same numbers pretty much on We got the same numbers pretty much on We got the same numbers pretty much on both machines. Those Mac Studios go for both machines. Those Mac Studios go for both machines. Those Mac Studios go for like 10 grand each. Well, these dude like 10 grand each. Well, these dude like 10 grand each. Well, these dude down here, this one I got on sale for down here, this one I got on sale for down here, this one I got on sale for 4,000. But still a huge difference in 4,000. But still a huge difference in 4,000. But still a huge difference in the hardware here that is tied to the the hardware here that is tied to the the hardware here that is tied to the Mac Mini. Are you serious? Mac Mini. Are you serious? Mac Mini. Are you serious? So now I had a kind of a mystery on my So now I had a kind of a mystery on my So now I had a kind of a mystery on my hands. Why would these two be pretty hands. Why would these two be pretty hands. Why would these two be pretty much the same? And I had a hunch that it much the same? And I had a hunch that it much the same? And I had a hunch that it was probably running on the neural was probably running on the neural was probably running on the neural engine, the A&E or Apple neural engine. engine, the A&E or Apple neural engine. engine, the A&E or Apple neural engine. That's the NPU. And the reason is it's That's the NPU. And the reason is it's That's the NPU. And the reason is it's basically the same block of silicon basically the same block of silicon basically the same block of silicon across M3 and M4. That would explain it across M3 and M4. That would explain it across M3 and M4. That would explain it perfectly. So I wanted to confirm it perfectly. So I wanted to confirm it perfectly. So I wanted to confirm it real quick. This is Mactop. And Mactop real quick. This is Mactop. And Mactop real quick. This is Mactop. And Mactop is basically one of the applications. is basically one of the applications. is basically one of the applications. There's several of them that show the There's several of them that show the There's several of them that show the current usage on the terminal of the current usage on the terminal of the current usage on the terminal of the CPU, the GPU, the A&E, and then the CPU, the GPU, the A&E, and then the CPU, the GPU, the A&E, and then the memory used. And it's loosely based on memory used. And it's loosely based on memory used. And it's loosely based on Azytop, which I have running over here. Azytop, which I have running over here. Azytop, which I have running over here. They are very similar. Actually, I don't They are very similar. Actually, I don't They are very similar. Actually, I don't have that running at the moment cuz you have that running at the moment cuz you have that running at the moment cuz you can only have one of them running at a can only have one of them running at a can only have one of them running at a time. But Azy top is an oldie. It came time. But Azy top is an oldie. It came time. But Azy top is an oldie. It came out when Apple silicon first came out out when Apple silicon first came out out when Apple silicon first came out and it showed the A&E and for a while and it showed the A&E and for a while and it showed the A&E and for a while there it showed it quite nicely until there it showed it quite nicely until there it showed it quite nicely until now. So I watched the neural engine now. So I watched the neural engine now. So I watched the neural engine while FM is running and it stayed at while FM is running and it stayed at while FM is running and it stayed at zero the whole time. 0%. I tried it with zero the whole time. 0%. I tried it with zero the whole time. 0%. I tried it with Aztop. I tried it with Mactop. So much Aztop. I tried it with Mactop. So much Aztop. I tried it with Mactop. So much for that theory that I was using A&E, for that theory that I was using A&E, for that theory that I was using A&E, right? I was now starting to think that right? I was now starting to think that right? I was now starting to think that this was the GPU that was doing it. But this was the GPU that was doing it. But this was the GPU that was doing it. But something bugged me. The numbers just something bugged me. The numbers just something bugged me. The numbers just didn't feel right. So instead of testing

  7. didn't feel right. So instead of testing didn't feel right. So instead of testing FM directly, I tested the tools. I built FM directly, I tested the tools. I built FM directly, I tested the tools. I built a little workload that has to use the a little workload that has to use the a little workload that has to use the neural engine. If you use CoreML model neural engine. If you use CoreML model neural engine. If you use CoreML model and you run it and of course I had and you run it and of course I had and you run it and of course I had Claude code help me with that cuz I Claude code help me with that cuz I Claude code help me with that cuz I never touched CoreML before. CoreML is never touched CoreML before. CoreML is never touched CoreML before. CoreML is going to be using the neural engine. I going to be using the neural engine. I going to be using the neural engine. I ran it and guess what? Azytop Mactop ran it and guess what? Azytop Mactop ran it and guess what? Azytop Mactop zero. In fact, Azy top and Mactop use zero. In fact, Azy top and Mactop use zero. In fact, Azy top and Mactop use powermetrics under the hood. That's powermetrics under the hood. That's powermetrics under the hood. That's Apple's own CLI utility. And Parametrics Apple's own CLI utility. And Parametrics Apple's own CLI utility. And Parametrics also told me the CPU was pulling zero also told me the CPU was pulling zero also told me the CPU was pulling zero watts while it was literally running the watts while it was literally running the watts while it was literally running the test, which is impossible. So I ran the test, which is impossible. So I ran the test, which is impossible. So I ran the same model two ways. One is neural same model two ways. One is neural same model two ways. One is neural engine on and the other one is neural engine on and the other one is neural engine on and the other one is neural engine off. With the neural engine on, engine off. With the neural engine on, engine off. With the neural engine on, it was over two times faster, which it was over two times faster, which it was over two times faster, which means that for FM, the neural engine is means that for FM, the neural engine is means that for FM, the neural engine is absolutely working. It's doing real absolutely working. It's doing real absolutely working. It's doing real measurable work. And every single tool measurable work. And every single tool measurable work. And every single tool that I use to measure the A&E is just that I use to measure the A&E is just that I use to measure the A&E is just not seeing it. not seeing it. not seeing it. All right, new plan. If I can't trust a All right, new plan. If I can't trust a All right, new plan. If I can't trust a single meter on these machines, I'll single meter on these machines, I'll single meter on these machines, I'll stop using the meters entirely and I'll stop using the meters entirely and I'll stop using the meters entirely and I'll make the engines fight. Here's the idea.

  8. make the engines fight. Here's the idea. make the engines fight. Here's the idea. I run FM and while it's running, I I run FM and while it's running, I I run FM and while it's running, I hammer the neural engine as hard as I hammer the neural engine as hard as I hammer the neural engine as hard as I possibly can. If FM slows down, it's possibly can. If FM slows down, it's possibly can. If FM slows down, it's fighting for the A&E. Then I'll do the fighting for the A&E. Then I'll do the fighting for the A&E. Then I'll do the whole thing again, but this time I whole thing again, but this time I whole thing again, but this time I hammer the GPU instead. Whatever drags hammer the GPU instead. Whatever drags hammer the GPU instead. Whatever drags FM down tells me which engine it's FM down tells me which engine it's FM down tells me which engine it's really using. And folks, the result was really using. And folks, the result was really using. And folks, the result was beautiful. When I hammered the neural beautiful. When I hammered the neural beautiful. When I hammered the neural engine, FM's decode slowed down. Pro engine, FM's decode slowed down. Pro engine, FM's decode slowed down. Pro processing didn't move a bit though. And processing didn't move a bit though. And processing didn't move a bit though. And when I hammered the GPU, FM's processing when I hammered the GPU, FM's processing when I hammered the GPU, FM's processing slowed down. That's this blue line right slowed down. That's this blue line right slowed down. That's this blue line right here. And decode didn't move. Each one here. And decode didn't move. Each one here. And decode didn't move. Each one hits exactly 1/2. Clean as it can be. hits exactly 1/2. Clean as it can be. hits exactly 1/2. Clean as it can be. And that right there is proof with no And that right there is proof with no And that right there is proof with no meters at all that FM's decode runs on meters at all that FM's decode runs on meters at all that FM's decode runs on the neural engine and its prompt the neural engine and its prompt the neural engine and its prompt processing runs on the GPU. So that's processing runs on the GPU. So that's processing runs on the GPU. So that's why the studio couldn't beat the Mini. why the studio couldn't beat the Mini. why the studio couldn't beat the Mini. Decode is on the neural engine. And Decode is on the neural engine. And Decode is on the neural engine. And that's basically the same exact chip that's basically the same exact chip that's basically the same exact chip across the M3 and the M4 and even the M2 across the M3 and the M4 and even the M2 across the M3 and the M4 and even the M2 and the M1 I believe. Well, the M1 I and the M1 I believe. Well, the M1 I and the M1 I believe. Well, the M1 I think I had a different one, but the M2 think I had a different one, but the M2 think I had a different one, but the M2 to M5 all have that 16 core neural to M5 all have that 16 core neural to M5 all have that 16 core neural engine. So all that extra power just engine. So all that extra power just engine. So all that extra power just sitting around doing nothing on the Mac sitting around doing nothing on the Mac sitting around doing nothing on the Mac Studio. So basically, if you're running Studio. So basically, if you're running Studio. So basically, if you're running Apple's built-in FM tool, a Mac Mini Apple's built-in FM tool, a Mac Mini Apple's built-in FM tool, a Mac Mini gives you the same exact experience as gives you the same exact experience as gives you the same exact experience as the maxed out Mac Studio. Same speed, the maxed out Mac Studio. Same speed, the maxed out Mac Studio. Same speed, fraction of the price. save that money fraction of the price. save that money fraction of the price. save that money or you can spend it on something that or you can spend it on something that or you can spend it on something that actually scales. For example, remember I actually scales. For example, remember I actually scales. For example, remember I talked about other things like Llama talked about other things like Llama talked about other things like Llama CPP, which is basically what LM Studio CPP, which is basically what LM Studio CPP, which is basically what LM Studio and Olma use under the hood, and MLX, and Olma use under the hood, and MLX, and Olma use under the hood, and MLX, especially MLX is Apple's own framework

  9. especially MLX is Apple's own framework especially MLX is Apple's own framework for doing LLMs. And MLX absolutely does for doing LLMs. And MLX absolutely does for doing LLMs. And MLX absolutely does use that big GPU. I made a bunch of use that big GPU. I made a bunch of use that big GPU. I made a bunch of other videos about it. Link down below. other videos about it. Link down below. other videos about it. Link down below. By the way, did you even know that Mac By the way, did you even know that Mac By the way, did you even know that Mac OS had this thing called FM hiding OS had this thing called FM hiding OS had this thing called FM hiding inside of it? Let me know down in the inside of it? Let me know down in the inside of it? Let me know down in the comments below. And the tool that I comments below. And the tool that I comments below. And the tool that I built to test all this, Apple FM, is built to test all this, Apple FM, is built to test all this, Apple FM, is free on GitHub, link down below, too. Go free on GitHub, link down below, too. Go free on GitHub, link down below, too. Go run it on your own Mac and tell me what run it on your own Mac and tell me what run it on your own Mac and tell me what you get. Although, I have a feeling that you get. Although, I have a feeling that you get. Although, I have a feeling that you're probably going to have the same you're probably going to have the same you're probably going to have the same results. Oh, and just a heads up, this results. Oh, and just a heads up, this results. Oh, and just a heads up, this is still in beta, of course, and it will is still in beta, of course, and it will is still in beta, of course, and it will never run on anything other than Apple never run on anything other than Apple never run on anything other than Apple Silicon. So, Intel Max, it's not going Silicon. So, Intel Max, it's not going Silicon. So, Intel Max, it's not going to work cuz Golden Gate doesn't do Intel to work cuz Golden Gate doesn't do Intel to work cuz Golden Gate doesn't do Intel Max. Also, keep in mind that since this Max. Also, keep in mind that since this Max. Also, keep in mind that since this is beta, Apple can change this anytime. is beta, Apple can change this anytime. is beta, Apple can change this anytime. They can even change it after release. They can even change it after release. They can even change it after release. they can have it use a piece of the GPU they can have it use a piece of the GPU they can have it use a piece of the GPU later on. Now, if you want to see some later on. Now, if you want to see some later on. Now, if you want to see some bigger and chunkier models running on bigger and chunkier models running on bigger and chunkier models running on Apple Silicon, watch this video here. Apple Silicon, watch this video here. Apple Silicon, watch this video here. Thanks for watching and I'll see you Thanks for watching and I'll see you Thanks for watching and I'll see you next time.

Summary

The transcript discusses a proactive AI agent called Super Nori which operates in the background of real-world family contexts, unlike traditional AI assistants that require manual prompts. Super Nori monitors situations, such as traffic or calendar conflicts, and offers solutions before issues escalate, demonstrating a more agent-like behavior. The practical takeaway is that this type of AI acts proactively and asks for confirmation before taking action, making it suitable for delicate real-world decisions.

View original episode ↗