GLM 5.2 in Claude Code is Blowing My Mind
Read full transcript 12 segments
-
So, I've been playing around with GLM So, I've been playing around with GLM 5.2 inside of Cloud Code all day, and 5.2 inside of Cloud Code all day, and 5.2 inside of Cloud Code all day, and it's incredible. It feels faster. It's it's incredible. It feels faster. It's it's incredible. It feels faster. It's significantly cheaper, and it just fits significantly cheaper, and it just fits significantly cheaper, and it just fits right into the Cloud Code harness pretty right into the Cloud Code harness pretty right into the Cloud Code harness pretty well. It's been doing so well, in fact, well. It's been doing so well, in fact, well. It's been doing so well, in fact, that it edited this entire intro that that it edited this entire intro that that it edited this entire intro that you're watching right now from raw video you're watching right now from raw video you're watching right now from raw video all the way to what you're watching. So, all the way to what you're watching. So, all the way to what you're watching. So, in today's video, I'm going to show you in today's video, I'm going to show you in today's video, I'm going to show you just how quick and easy you can get set just how quick and easy you can get set just how quick and easy you can get set up with GLM 5.2 in Cloud Code. So, let's up with GLM 5.2 in Cloud Code. So, let's up with GLM 5.2 in Cloud Code. So, let's not waste any time and just get straight not waste any time and just get straight not waste any time and just get straight into the video. So, obviously that intro into the video. So, obviously that intro into the video. So, obviously that intro wasn't perfect, but that was literally wasn't perfect, but that was literally wasn't perfect, but that was literally one prompt. It was one/goal right here. one prompt. It was one/goal right here. one prompt. It was one/goal right here. It took a little over an hour, an hour It took a little over an hour, an hour It took a little over an hour, an hour and 15 minutes for, you know, a 23 and 15 minutes for, you know, a 23 and 15 minutes for, you know, a 23 second video. So, that did take a little second video. So, that did take a little second video. So, that did take a little bit long, but this was the session that bit long, but this was the session that bit long, but this was the session that we did it in. It was GLM 5.2 1 million we did it in. It was GLM 5.2 1 million we did it in. It was GLM 5.2 1 million context. As you can see right here, it context. As you can see right here, it context. As you can see right here, it used about 357,000 tokens. But, as I've used about 357,000 tokens. But, as I've used about 357,000 tokens. But, as I've been playing around with it more and been playing around with it more and been playing around with it more and more, there are some tasks where it more, there are some tasks where it more, there are some tasks where it finishes way faster than Opus. And then finishes way faster than Opus. And then finishes way faster than Opus. And then there are some like this one where Opus there are some like this one where Opus there are some like this one where Opus would have done this much quicker. So, would have done this much quicker. So, would have done this much quicker. So, I'm going to show you guys exactly how I'm going to show you guys exactly how I'm going to show you guys exactly how to get set up. But before that, let me to get set up. But before that, let me to get set up. But before that, let me just show you a few of the things that just show you a few of the things that just show you a few of the things that I've played around with and what GLM has I've played around with and what GLM has I've played around with and what GLM has been able to do. So, it's actually like been able to do. So, it's actually like been able to do. So, it's actually like really solid at design. I want you guys really solid at design. I want you guys really solid at design. I want you guys to look right here and see which one of to look right here and see which one of to look right here and see which one of these do you think was designed by GLM these do you think was designed by GLM these do you think was designed by GLM 5.2 and which one was designed by Opus.
-
5.2 and which one was designed by Opus. 5.2 and which one was designed by Opus. So, as we sort of scroll down here, you So, as we sort of scroll down here, you So, as we sort of scroll down here, you can see that we have, you know, similar can see that we have, you know, similar can see that we have, you know, similar style branding and it's obviously the style branding and it's obviously the style branding and it's obviously the same company, but we have elements on same company, but we have elements on same company, but we have elements on both that are very similar. We have all both that are very similar. We have all both that are very similar. We have all of these things come up dynamically as of these things come up dynamically as of these things come up dynamically as well on either side. And there's even a well on either side. And there's even a well on either side. And there's even a CTA at the bottom. I think the dead CTA at the bottom. I think the dead CTA at the bottom. I think the dead giveaway here is this right side was giveaway here is this right side was giveaway here is this right side was opus because it has these weird Fs that opus because it has these weird Fs that opus because it has these weird Fs that it loves to do. It loves that font. But it loves to do. It loves that font. But it loves to do. It loves that font. But either way, these are both very solid either way, these are both very solid either way, these are both very solid for a oneshot prompt. Especially when for a oneshot prompt. Especially when for a oneshot prompt. Especially when you consider the fact that you are you consider the fact that you are you consider the fact that you are getting this output for like five times getting this output for like five times getting this output for like five times cheaper. So these are the actual cheaper. So these are the actual cheaper. So these are the actual terminal sessions where we did those terminal sessions where we did those terminal sessions where we did those website designs. On the left side with website designs. On the left side with website designs. On the left side with GLM, we got this done in 3 minutes and GLM, we got this done in 3 minutes and GLM, we got this done in 3 minutes and 59 seconds. On the right side with Opus, 59 seconds. On the right side with Opus, 59 seconds. On the right side with Opus, we got this done in 14 minutes and 59 we got this done in 14 minutes and 59 we got this done in 14 minutes and 59 seconds. And not only did the lefth hand seconds. And not only did the lefth hand seconds. And not only did the lefth hand side GLM use less tokens, but its cost side GLM use less tokens, but its cost side GLM use less tokens, but its cost per token is also five times cheaperish. per token is also five times cheaperish. per token is also five times cheaperish. So in this case, it was quicker and it So in this case, it was quicker and it So in this case, it was quicker and it was much cheaper and it was a relatively was much cheaper and it was a relatively was much cheaper and it was a relatively similar result. I also shot off this similar result. I also shot off this similar result. I also shot off this prompt on each side where I gave them a prompt on each side where I gave them a prompt on each side where I gave them a homework assignment. You can see right homework assignment. You can see right homework assignment. You can see right here we've got GLM and then right here here we've got GLM and then right here here we've got GLM and then right here we've got Opus 4.8. I had Codeex create we've got Opus 4.8. I had Codeex create we've got Opus 4.8. I had Codeex create the homework assignment just so there the homework assignment just so there the homework assignment just so there was no like crosscontamination or was no like crosscontamination or was no like crosscontamination or anything like that. And then when they anything like that. And then when they anything like that. And then when they finished, I had Codeex judge both finished, I had Codeex judge both finished, I had Codeex judge both results and tell us what it thought.
-
results and tell us what it thought. results and tell us what it thought. Now, in this case, it said that agent 2, Now, in this case, it said that agent 2, Now, in this case, it said that agent 2, which was opus, was better because it which was opus, was better because it which was opus, was better because it handled one subtle edge case that agent handled one subtle edge case that agent handled one subtle edge case that agent one missed, which were duplicate records one missed, which were duplicate records one missed, which were duplicate records with values like true versus one or one with values like true versus one or one with values like true versus one or one versus 1.0. So, the short version is versus 1.0. So, the short version is versus 1.0. So, the short version is that agent 1, GLM 5.2 was good, but that agent 1, GLM 5.2 was good, but that agent 1, GLM 5.2 was good, but agent 2 here was more precise. And agent 2 here was more precise. And agent 2 here was more precise. And generally, the way that I feel about generally, the way that I feel about generally, the way that I feel about this so far is that GLM 5.2 is really this so far is that GLM 5.2 is really this so far is that GLM 5.2 is really solid and it's pretty quick for most solid and it's pretty quick for most solid and it's pretty quick for most tasks that don't require heavy tasks that don't require heavy tasks that don't require heavy reasoning. Obviously, at the end of the reasoning. Obviously, at the end of the reasoning. Obviously, at the end of the day, Opus 4.8 is a better model. It's a day, Opus 4.8 is a better model. It's a day, Opus 4.8 is a better model. It's a closed source model. But realistically closed source model. But realistically closed source model. But realistically ask yourself how often do you actually ask yourself how often do you actually ask yourself how often do you actually need the power of Opus? Probably only need the power of Opus? Probably only need the power of Opus? Probably only maybe 10 to 20% if that of the tasks maybe 10 to 20% if that of the tasks maybe 10 to 20% if that of the tasks that you do all day. You could probably that you do all day. You could probably that you do all day. You could probably handle 80% or more of your knowledge handle 80% or more of your knowledge handle 80% or more of your knowledge work with something like GLM 5.2 or work with something like GLM 5.2 or work with something like GLM 5.2 or something more like sonnet 3.7. So something more like sonnet 3.7. So something more like sonnet 3.7. So that's really going to be a key skill as that's really going to be a key skill as that's really going to be a key skill as we move into the future of AI is we move into the future of AI is we move into the future of AI is understanding which models to use per understanding which models to use per understanding which models to use per task. But here's an example where you task. But here's an example where you task. But here's an example where you can see Opus took about 5 minutes and can see Opus took about 5 minutes and can see Opus took about 5 minutes and GLM 5.2 took about 24 minutes. And this GLM 5.2 took about 24 minutes. And this GLM 5.2 took about 24 minutes. And this was me on a $60 a month plan for um Z.AI was me on a $60 a month plan for um Z.AI was me on a $60 a month plan for um Z.AI for GLM 5.2. And I've played around with for GLM 5.2. And I've played around with for GLM 5.2. And I've played around with this for this was about four or five this for this was about four or five this for this was about four or five hours straight of just literally hours straight of just literally hours straight of just literally hammering it. Five different sessions hammering it. Five different sessions hammering it. Five different sessions open, testing GLM 5.2. And my 5-hour open, testing GLM 5.2. And my 5-hour open, testing GLM 5.2. And my 5-hour quota is a little bit over halfway used.
-
quota is a little bit over halfway used. quota is a little bit over halfway used. And my weekly quota is about 10% used. And my weekly quota is about 10% used. And my weekly quota is about 10% used. I'll talk about the billing and how to I'll talk about the billing and how to I'll talk about the billing and how to get set up in just a sec. Let me show get set up in just a sec. Let me show get set up in just a sec. Let me show you guys what else I did with it. So I you guys what else I did with it. So I you guys what else I did with it. So I did a few more/goal prompts. And I hate did a few more/goal prompts. And I hate did a few more/goal prompts. And I hate when the terminal does this, but this when the terminal does this, but this when the terminal does this, but this first one that I did was I did /goal and first one that I did was I did /goal and first one that I did was I did /goal and I literally said like, "Hey, get I literally said like, "Hey, get I literally said like, "Hey, get creative. Show me how good your design creative. Show me how good your design creative. Show me how good your design skills are and just build me whatever skills are and just build me whatever skills are and just build me whatever you want. Just make me an HTML you want. Just make me an HTML you want. Just make me an HTML document." And then this is what it gave document." And then this is what it gave document." And then this is what it gave me. It gave me the anatomy of attention. me. It gave me the anatomy of attention. me. It gave me the anatomy of attention. You can see we've got like some stars You can see we've got like some stars You can see we've got like some stars moving around in the background. This moving around in the background. This moving around in the background. This obviously looks a little bit vibe coded obviously looks a little bit vibe coded obviously looks a little bit vibe coded up here, but not too bad. And as we sort up here, but not too bad. And as we sort up here, but not too bad. And as we sort of scroll down, we can see a language of scroll down, we can see a language of scroll down, we can see a language model has no grammar book and no model has no grammar book and no model has no grammar book and no dictionary. And then on this thing, the dictionary. And then on this thing, the dictionary. And then on this thing, the animal didn't cross the street because animal didn't cross the street because animal didn't cross the street because it was too tired. This is kind of like it was too tired. This is kind of like it was too tired. This is kind of like an interactive element here. We see the an interactive element here. We see the an interactive element here. We see the query word and then we see where it query word and then we see where it query word and then we see where it points. So tired was pointing to it. It points. So tired was pointing to it. It points. So tired was pointing to it. It was pointing to the animal. Cross was was pointing to the animal. Cross was was pointing to the animal. Cross was pointing to the street. So kind of pointing to the street. So kind of pointing to the street. So kind of interesting. First the sentence is interesting. First the sentence is interesting. First the sentence is broken. Every token gets a place in broken. Every token gets a place in broken. Every token gets a place in space. So we've got a little bit of like space. So we've got a little bit of like space. So we've got a little bit of like a relationship graph here. We've got a relationship graph here. We've got a relationship graph here. We've got some charts down here and some different some charts down here and some different some charts down here and some different elements. So anyways, this is what GLM elements. So anyways, this is what GLM elements. So anyways, this is what GLM came up with when I said, "Hey, just be came up with when I said, "Hey, just be came up with when I said, "Hey, just be creative and just give me whatever you creative and just give me whatever you creative and just give me whatever you want. Show me whatever interests you, I want. Show me whatever interests you, I want. Show me whatever interests you, I And then I did the exact same prompt And then I did the exact same prompt And then I did the exact same prompt into Opus to see what kind of into Opus to see what kind of into Opus to see what kind of differences we would get. And Opus came differences we would get. And Opus came differences we would get. And Opus came up with the life of a Death Star. Here up with the life of a Death Star. Here up with the life of a Death Star. Here once again, you can see that classic F once again, you can see that classic F once again, you can see that classic F that Opus loves to do. But anyways, this that Opus loves to do. But anyways, this that Opus loves to do. But anyways, this one is more of like a timeline where we one is more of like a timeline where we one is more of like a timeline where we go through the actual kind of like life go through the actual kind of like life go through the actual kind of like life cycle of a Death Star. And we can see cycle of a Death Star. And we can see cycle of a Death Star. And we can see this one is also pretty good as well.
-
this one is also pretty good as well. this one is also pretty good as well. But once again, pay attention to the But once again, pay attention to the But once again, pay attention to the fact that from a design perspective with fact that from a design perspective with fact that from a design perspective with one shot, is Opus that much better? Is one shot, is Opus that much better? Is one shot, is Opus that much better? Is Opus five times better? Because you're Opus five times better? Because you're Opus five times better? Because you're paying five times more. However, with paying five times more. However, with paying five times more. However, with this one, the goal took about 35 this one, the goal took about 35 this one, the goal took about 35 minutes, and for Opus, this goal only minutes, and for Opus, this goal only minutes, and for Opus, this goal only took about 11. So, it's really a hit or took about 11. So, it's really a hit or took about 11. So, it's really a hit or miss as far as when is GLM 5.2 actually miss as far as when is GLM 5.2 actually miss as far as when is GLM 5.2 actually faster. Typically, the more reasoning, faster. Typically, the more reasoning, faster. Typically, the more reasoning, the slower it's going to be. And so, the slower it's going to be. And so, the slower it's going to be. And so, just remember, cloud code is a harness. just remember, cloud code is a harness. just remember, cloud code is a harness. It's a harness for AI models, and It's a harness for AI models, and It's a harness for AI models, and typically cloud models are going to use typically cloud models are going to use typically cloud models are going to use the harness the best. But, GLM 5.2 does the harness the best. But, GLM 5.2 does the harness the best. But, GLM 5.2 does pretty decent. It's able to use the SLG pretty decent. It's able to use the SLG pretty decent. It's able to use the SLG goal. It's able to read my cloudmd. It's goal. It's able to read my cloudmd. It's goal. It's able to read my cloudmd. It's able to use my skills. Right here I said able to use my skills. Right here I said able to use my skills. Right here I said /goal I need you to use the storm /goal I need you to use the storm /goal I need you to use the storm research skill which I'll have a video research skill which I'll have a video research skill which I'll have a video coming out about that very soon to coming out about that very soon to coming out about that very soon to research open source AI models versus research open source AI models versus research open source AI models versus closed source and the end deliverable is closed source and the end deliverable is closed source and the end deliverable is an HTML report. So it goes through it an HTML report. So it goes through it an HTML report. So it goes through it does a bunch of sub aents all the sub does a bunch of sub aents all the sub does a bunch of sub aents all the sub aents were using GLM 5.2 as well and it aents were using GLM 5.2 as well and it aents were using GLM 5.2 as well and it had a bunch of different personas. This had a bunch of different personas. This had a bunch of different personas. This took about 27 minutes and I got this took about 27 minutes and I got this took about 27 minutes and I got this report back. This was our storm research report back. This was our storm research report back. This was our storm research and you see it says V2. That just means and you see it says V2. That just means and you see it says V2. That just means that it had one pass and then it has that it had one pass and then it has that it had one pass and then it has another set of agents come and read that another set of agents come and read that another set of agents come and read that and then it makes some changes. So and then it makes some changes. So and then it makes some changes. So anyways, this is the HTML report that we anyways, this is the HTML report that we anyways, this is the HTML report that we got. I'm not going to read every single got. I'm not going to read every single got. I'm not going to read every single word of this, but it is very thorough.
-
word of this, but it is very thorough. word of this, but it is very thorough. It's very solid. It looks decent. You It's very solid. It looks decent. You It's very solid. It looks decent. You can see we had five different lenses can see we had five different lenses can see we had five different lenses going through this, which you'll going through this, which you'll going through this, which you'll probably see a little bit sprinkled probably see a little bit sprinkled probably see a little bit sprinkled throughout. We have a 60-second summary. throughout. We have a 60-second summary. throughout. We have a 60-second summary. We have five key findings. This one was We have five key findings. This one was We have five key findings. This one was supported by the academic and the supported by the academic and the supported by the academic and the skeptic. This one was supported by the skeptic. This one was supported by the skeptic. This one was supported by the practitioner, economist, and the practitioner, economist, and the practitioner, economist, and the academic. And it was challenged by the academic. And it was challenged by the academic. And it was challenged by the skeptic and the historian. So basically skeptic and the historian. So basically skeptic and the historian. So basically we're just leveraging like you know a we're just leveraging like you know a we're just leveraging like you know a mixture of experts here or a mixture of mixture of experts here or a mixture of mixture of experts here or a mixture of different you know styles of agents to different you know styles of agents to different you know styles of agents to help us do the research and help us help us do the research and help us help us do the research and help us debate over how this thing is actually debate over how this thing is actually debate over how this thing is actually constructed. We have the hidden constructed. We have the hidden constructed. We have the hidden connection. We have the assumption that connection. We have the assumption that connection. We have the assumption that the briefing rests on what to actually the briefing rests on what to actually the briefing rests on what to actually do different. So if you guys want to do different. So if you guys want to do different. So if you guys want to pause this and actually read through it. pause this and actually read through it. pause this and actually read through it. It has some really good information in It has some really good information in It has some really good information in here. But anyways this is what it came here. But anyways this is what it came here. But anyways this is what it came up with. This made me start to wonder up with. This made me start to wonder up with. This made me start to wonder like okay where would I actually use GLM like okay where would I actually use GLM like okay where would I actually use GLM 5.2 over something like Opus 4.8 because 5.2 over something like Opus 4.8 because 5.2 over something like Opus 4.8 because I do think that after I did a ton of I do think that after I did a ton of I do think that after I did a ton of knowledge work with this today and knowledge work with this today and knowledge work with this today and testing, it's pretty solid. And so I testing, it's pretty solid. And so I testing, it's pretty solid. And so I think for something like this, I would think for something like this, I would think for something like this, I would 100% be comfortable with GLM 5.2 doing 100% be comfortable with GLM 5.2 doing 100% be comfortable with GLM 5.2 doing this because I felt comfortable with the this because I felt comfortable with the this because I felt comfortable with the way that I orchestrated that storm way that I orchestrated that storm way that I orchestrated that storm skill. Bunch of different agents, bunch skill. Bunch of different agents, bunch skill. Bunch of different agents, bunch of different verification checks. And of different verification checks. And of different verification checks. And that is way more important ultimately that is way more important ultimately that is way more important ultimately than the model, than the underlying than the model, than the underlying than the model, than the underlying model. It's all about the way that you model. It's all about the way that you model. It's all about the way that you prompt them, the way that you use them, prompt them, the way that you use them, prompt them, the way that you use them, the way that you have your skills and the way that you have your skills and the way that you have your skills and your harness and your context layer. So your harness and your context layer. So your harness and your context layer. So I would trust GLM 5.2 here big time to I would trust GLM 5.2 here big time to I would trust GLM 5.2 here big time to do me a bunch of research. And by do me a bunch of research. And by do me a bunch of research. And by research, I mean like gathering a bunch research, I mean like gathering a bunch research, I mean like gathering a bunch of opinions and gathering data and of opinions and gathering data and of opinions and gathering data and pulling in sources. But I probably would pulling in sources. But I probably would pulling in sources. But I probably would want Opus to actually help me think want Opus to actually help me think want Opus to actually help me think through based on all this data what through based on all this data what through based on all this data what really matters and how do I apply it to really matters and how do I apply it to really matters and how do I apply it to my life. That's where I'd probably lean my life. That's where I'd probably lean my life. That's where I'd probably lean on a heavier reasoning model, a stronger
-
on a heavier reasoning model, a stronger on a heavier reasoning model, a stronger model like Opus. So I feel like everyone model like Opus. So I feel like everyone model like Opus. So I feel like everyone needs to be thinking about that kind of needs to be thinking about that kind of needs to be thinking about that kind of stuff. It's not binary. It's where in stuff. It's not binary. It's where in stuff. It's not binary. It's where in each process, what steps should I use each process, what steps should I use each process, what steps should I use what model for? Anyways, why am I what model for? Anyways, why am I what model for? Anyways, why am I talking about GLM 5.2? Because it is an talking about GLM 5.2? Because it is an talking about GLM 5.2? Because it is an open- source model, right? So like open- source model, right? So like open- source model, right? So like chatbt or claude but those are those are chatbt or claude but those are those are chatbt or claude but those are those are closed source models meaning you rent it closed source models meaning you rent it closed source models meaning you rent it you pay directly to the provider in you pay directly to the provider in you pay directly to the provider in order to access it. Now yes you guys saw order to access it. Now yes you guys saw order to access it. Now yes you guys saw earlier I am paying a subscription to earlier I am paying a subscription to earlier I am paying a subscription to access GLM 5.2 and that is just because access GLM 5.2 and that is just because access GLM 5.2 and that is just because it is so massive. It is a massive model it is so massive. It is a massive model it is so massive. It is a massive model 753 billion parameters which is very big 753 billion parameters which is very big 753 billion parameters which is very big which means that I couldn't actually run which means that I couldn't actually run which means that I couldn't actually run that on my machine. Yes it is open that on my machine. Yes it is open that on my machine. Yes it is open source. Yes you could run it locally but source. Yes you could run it locally but source. Yes you could run it locally but you would need the hardware you would you would need the hardware you would you would need the hardware you would need the infrastructure to support that. need the infrastructure to support that. need the infrastructure to support that. Most of us don't have that just lying Most of us don't have that just lying Most of us don't have that just lying around. So what we can do is we can rent around. So what we can do is we can rent around. So what we can do is we can rent it online. So kind of very similar the it online. So kind of very similar the it online. So kind of very similar the way that you pay Anthropic for Claude, way that you pay Anthropic for Claude, way that you pay Anthropic for Claude, but it's so much cheaper than Claude. So but it's so much cheaper than Claude. So but it's so much cheaper than Claude. So everyone is freaking out because it's everyone is freaking out because it's everyone is freaking out because it's basically yours. You're able to download basically yours. You're able to download basically yours. You're able to download it or get it for much cheaper. It's very it or get it for much cheaper. It's very it or get it for much cheaper. It's very smart. I'll show you guys some smart. I'll show you guys some smart. I'll show you guys some benchmarks in a sec. And it is very benchmarks in a sec. And it is very benchmarks in a sec. And it is very cheap. If you did a heavy day of coding, cheap. If you did a heavy day of coding, cheap. If you did a heavy day of coding, it'd be about five times cheaper than it'd be about five times cheaper than it'd be about five times cheaper than Opus 4.8 for the same job. You can see Opus 4.8 for the same job. You can see Opus 4.8 for the same job. You can see here, here is the input and output here, here is the input and output here, here is the input and output tokens. Opus 4.8 is $5 on the input, $25 tokens. Opus 4.8 is $5 on the input, $25 tokens. Opus 4.8 is $5 on the input, $25 on the output, whereas GM 5.2 is $1.40 on the output, whereas GM 5.2 is $1.40 on the output, whereas GM 5.2 is $1.40 40 cents on the input and $4.40 on the 40 cents on the input and $4.40 on the 40 cents on the input and $4.40 on the output. So this is where I got the output. So this is where I got the output. So this is where I got the numbers where I said it's about five numbers where I said it's about five numbers where I said it's about five times cheaper. So if you go to something times cheaper. So if you go to something times cheaper. So if you go to something like Olama, which is somewhere where you like Olama, which is somewhere where you like Olama, which is somewhere where you can actually download and pull in local can actually download and pull in local can actually download and pull in local models, you can see here that we have GM models, you can see here that we have GM models, you can see here that we have GM 5.2. It's got the 1 million context 5.2. It's got the 1 million context 5.2. It's got the 1 million context window and the size is 756 billion window and the size is 756 billion window and the size is 756 billion parameters in this case. Now they don't parameters in this case. Now they don't parameters in this case. Now they don't actually let you pull this in. They let actually let you pull this in. They let actually let you pull this in. They let you run it from their cloud, which is
-
you run it from their cloud, which is you run it from their cloud, which is nice because it is such a big model. But nice because it is such a big model. But nice because it is such a big model. But take a look at some of these benchmarks take a look at some of these benchmarks take a look at some of these benchmarks compared to Cloud Opus 4.8 8 and GBD 5.5 compared to Cloud Opus 4.8 8 and GBD 5.5 compared to Cloud Opus 4.8 8 and GBD 5.5 which are like the two best models right which are like the two best models right which are like the two best models right now. Look where GLM 5.2 is stacking up. now. Look where GLM 5.2 is stacking up. now. Look where GLM 5.2 is stacking up. It is really comparable to all of these It is really comparable to all of these It is really comparable to all of these other top tier close source models which other top tier close source models which other top tier close source models which is why obviously a lot of people are is why obviously a lot of people are is why obviously a lot of people are freaking out about it right now and freaking out about it right now and freaking out about it right now and looking at it. Think about the fact that looking at it. Think about the fact that looking at it. Think about the fact that Fable got pulled away from us, right? Fable got pulled away from us, right? Fable got pulled away from us, right? That just tells you that we are renting That just tells you that we are renting That just tells you that we are renting something that could be taken away from something that could be taken away from something that could be taken away from us for, you know, out of nowhere. And us for, you know, out of nowhere. And us for, you know, out of nowhere. And what I'm worried about is Enthropic and what I'm worried about is Enthropic and what I'm worried about is Enthropic and Openi aren't profitable companies right Openi aren't profitable companies right Openi aren't profitable companies right now. We're paying $200 a month for a now. We're paying $200 a month for a now. We're paying $200 a month for a clawed max plan, but we're getting like clawed max plan, but we're getting like clawed max plan, but we're getting like $8,000 worth of inference out of that if $8,000 worth of inference out of that if $8,000 worth of inference out of that if we actually utilize it all the way. So, we actually utilize it all the way. So, we actually utilize it all the way. So, they are not profitable. So, what they are not profitable. So, what they are not profitable. So, what happens when we finally maybe we get happens when we finally maybe we get happens when we finally maybe we get Fable back, they're obviously not Fable back, they're obviously not Fable back, they're obviously not profitable on that either. So, what they profitable on that either. So, what they profitable on that either. So, what they might do is bring it back and say, "Hey, might do is bring it back and say, "Hey, might do is bring it back and say, "Hey, but you can't use this in your but you can't use this in your but you can't use this in your subscription. You can only use this via subscription. You can only use this via subscription. You can only use this via API billing." That's more expensive than API billing." That's more expensive than API billing." That's more expensive than Opus. So, if you can start to understand Opus. So, if you can start to understand Opus. So, if you can start to understand these open source models and you can these open source models and you can these open source models and you can start to deploy them locally for start to deploy them locally for start to deploy them locally for basically like completely free, then it basically like completely free, then it basically like completely free, then it is really going to help you stay ahead is really going to help you stay ahead is really going to help you stay ahead of the game here. Take a look at this of the game here. Take a look at this of the game here. Take a look at this bench, Frontier S. SWE. It performed bench, Frontier S. SWE. It performed bench, Frontier S. SWE. It performed better than GPT 5.5 in this benchmark, better than GPT 5.5 in this benchmark, better than GPT 5.5 in this benchmark, which is just absolutely crazy. If you which is just absolutely crazy. If you which is just absolutely crazy. If you guys liked Opus 4.7, which a lot of you guys liked Opus 4.7, which a lot of you guys liked Opus 4.7, which a lot of you guys did, some of you guys didn't. GLM guys did, some of you guys didn't. GLM guys did, some of you guys didn't. GLM 5.2 was beating that model in a lot of 5.2 was beating that model in a lot of 5.2 was beating that model in a lot of these evaluations. And it beats the most these evaluations. And it beats the most these evaluations. And it beats the most recent Sonnet model in a ton of these recent Sonnet model in a ton of these recent Sonnet model in a ton of these evaluations as well. So, just think evaluations as well. So, just think evaluations as well. So, just think about it like that. Think about truly about it like that. Think about truly about it like that. Think about truly how decent this model really is. Here's how decent this model really is. Here's how decent this model really is. Here's some more benchmarks. Aenta coding. I'm some more benchmarks. Aenta coding. I'm some more benchmarks. Aenta coding. I'm not going to go through all these not going to go through all these not going to go through all these because I, you know, we all know that because I, you know, we all know that because I, you know, we all know that you should always take these with a you should always take these with a you should always take these with a grain of salt. It's more about the feel grain of salt. It's more about the feel grain of salt. It's more about the feel and how you actually use them, but they
-
and how you actually use them, but they and how you actually use them, but they are interesting to look at every once in are interesting to look at every once in are interesting to look at every once in a while. So, anyways, let's talk about a while. So, anyways, let's talk about a while. So, anyways, let's talk about how you get set up. What you do is how you get set up. What you do is how you get set up. What you do is you're going to go to z.ai and you might you're going to go to z.ai and you might you're going to go to z.ai and you might pull into something that looks like pull into something that looks like pull into something that looks like this. You can go ahead and chat. You can this. You can go ahead and chat. You can this. You can go ahead and chat. You can make landing pages right here. You can make landing pages right here. You can make landing pages right here. You can do 3D modeling. You can build your own do 3D modeling. You can build your own do 3D modeling. You can build your own mini game. And it's really good at this mini game. And it's really good at this mini game. And it's really good at this stuff. It's really impressive on the stuff. It's really impressive on the stuff. It's really impressive on the front-end design stuff. So, come in front-end design stuff. So, come in front-end design stuff. So, come in here, play around with it if you want to here, play around with it if you want to here, play around with it if you want to just test it out and see how it feels. just test it out and see how it feels. just test it out and see how it feels. But then if you want to plug it into But then if you want to plug it into But then if you want to plug it into Cloud Code or Open Code or Hermes Agent Cloud Code or Open Code or Hermes Agent Cloud Code or Open Code or Hermes Agent or wherever you want to plug it in, or wherever you want to plug it in, or wherever you want to plug it in, you're going to click on this button up you're going to click on this button up you're going to click on this button up in the top right and that's going to in the top right and that's going to in the top right and that's going to take you to the actual API console. And take you to the actual API console. And take you to the actual API console. And so you've got the option to just pay per so you've got the option to just pay per so you've got the option to just pay per token which obviously is not too token which obviously is not too token which obviously is not too expensive. If you come here and I go to expensive. If you come here and I go to expensive. If you come here and I go to my billing and I go to the model my billing and I go to the model my billing and I go to the model pricing, you can see on the input it's pricing, you can see on the input it's pricing, you can see on the input it's $1.4 and on the output it is where'd it $1.4 and on the output it is where'd it $1.4 and on the output it is where'd it go? $4.4. But you could also just get a go? $4.4. But you could also just get a go? $4.4. But you could also just get a plan. So you could go on 16 bucks a plan. So you could go on 16 bucks a plan. So you could go on 16 bucks a month, 64 bucks a month or 144 bucks a month, 64 bucks a month or 144 bucks a month, 64 bucks a month or 144 bucks a month. And if you go yearly, you can month. And if you go yearly, you can month. And if you go yearly, you can save even more money. Obviously, they're save even more money. Obviously, they're save even more money. Obviously, they're not a sponsor. I'm not working with not a sponsor. I'm not working with not a sponsor. I'm not working with them, but you can also get on a plan. them, but you can also get on a plan. them, but you can also get on a plan. So, maybe you can be on a Claude plan So, maybe you can be on a Claude plan So, maybe you can be on a Claude plan for maybe 100 bucks a month, and then for maybe 100 bucks a month, and then for maybe 100 bucks a month, and then you can be on a Z plan for 64 bucks a you can be on a Z plan for 64 bucks a you can be on a Z plan for 64 bucks a month. And then just switch between month. And then just switch between month. And then just switch between them. Whenever you need a certain type them. Whenever you need a certain type them. Whenever you need a certain type of task, then you can bounce back and of task, then you can bounce back and of task, then you can bounce back and forth. And that way, you're getting way forth. And that way, you're getting way forth. And that way, you're getting way more out of your subscriptions. Once you more out of your subscriptions. Once you more out of your subscriptions. Once you get a plan, what you would do is you get a plan, what you would do is you get a plan, what you would do is you would go to API key and you would go would go to API key and you would go would go to API key and you would go ahead and grab one. So, you would add an ahead and grab one. So, you would add an ahead and grab one. So, you would add an API key here. And then you're going to API key here. And then you're going to API key here. And then you're going to be able to start just using that inside be able to start just using that inside be able to start just using that inside of Cloud Code. and then you'll be able of Cloud Code. and then you'll be able of Cloud Code. and then you'll be able to watch your usage limit here. Now, the to watch your usage limit here. Now, the to watch your usage limit here. Now, the five- hour quota and the weekly quota five- hour quota and the weekly quota five- hour quota and the weekly quota work very similar to claude code work very similar to claude code work very similar to claude code subscription. They also have peak hours subscription. They also have peak hours subscription. They also have peak hours where it consumes, you know, a higher where it consumes, you know, a higher where it consumes, you know, a higher multiple of your quota. And then what's multiple of your quota. And then what's multiple of your quota. And then what's interesting is they have like web search
-
interesting is they have like web search interesting is they have like web search quota as well. So, if you were doing a quota as well. So, if you were doing a quota as well. So, if you were doing a lot of web searching, um it will eat lot of web searching, um it will eat lot of web searching, um it will eat this up, but you could obviously connect this up, but you could obviously connect this up, but you could obviously connect it to maybe like perplexity or a it to maybe like perplexity or a it to maybe like perplexity or a different API to do more web searching. different API to do more web searching. different API to do more web searching. So, it's not a huge deal, but I did So, it's not a huge deal, but I did So, it's not a huge deal, but I did think that that was kind of interesting. think that that was kind of interesting. think that that was kind of interesting. So anyways, all you're going to do is So anyways, all you're going to do is So anyways, all you're going to do is you are just going to edit your config you are just going to edit your config you are just going to edit your config file within cloud code, which sounds a file within cloud code, which sounds a file within cloud code, which sounds a little bit scary or technical, but it's little bit scary or technical, but it's little bit scary or technical, but it's really not. Let me just pull up my cloud really not. Let me just pull up my cloud really not. Let me just pull up my cloud real quick and show you what that looks real quick and show you what that looks real quick and show you what that looks like. So inside of yourcloud, you should like. So inside of yourcloud, you should like. So inside of yourcloud, you should have a settings.local.json. have a settings.local.json. have a settings.local.json. And this is where you can play with your And this is where you can play with your And this is where you can play with your permissions and your MCP servers and all permissions and your MCP servers and all permissions and your MCP servers and all that kind of stuff. And you can also set that kind of stuff. And you can also set that kind of stuff. And you can also set environment variables. So if you guys environment variables. So if you guys environment variables. So if you guys have activated agent teams, this might have activated agent teams, this might have activated agent teams, this might live here like mine or it might live live here like mine or it might live live here like mine or it might live globally. Wherever it lives, that's an globally. Wherever it lives, that's an globally. Wherever it lives, that's an environment variable. And all you need environment variable. And all you need environment variable. And all you need to do is you need to come in here and to do is you need to come in here and to do is you need to come in here and you need to set these. So what I'm going you need to set these. So what I'm going you need to set these. So what I'm going to do is I'm going to have this exact to do is I'm going to have this exact to do is I'm going to have this exact thing, paste it in the description of thing, paste it in the description of thing, paste it in the description of this video. So all you have to do is this video. So all you have to do is this video. So all you have to do is copy this, put it into your environment copy this, put it into your environment copy this, put it into your environment variables. You can also say, hey, cloud variables. You can also say, hey, cloud variables. You can also say, hey, cloud code, put this into my code, put this into my code, put this into my settings.local.json, and then all you settings.local.json, and then all you settings.local.json, and then all you have to do is switch out your own API have to do is switch out your own API have to do is switch out your own API key. So this is where you will put your key. So this is where you will put your key. So this is where you will put your Z API key that you just got in here when Z API key that you just got in here when Z API key that you just got in here when you go to API keys, and you added one you go to API keys, and you added one you go to API keys, and you added one right there. Because what you're doing right there. Because what you're doing right there. Because what you're doing here is you can see we have enthropic here is you can see we have enthropic here is you can see we have enthropic base URL and we are routing that to Z's base URL and we are routing that to Z's base URL and we are routing that to Z's API rather than to an anthropic API. So API rather than to an anthropic API. So API rather than to an anthropic API. So we're just switching out the engine of we're just switching out the engine of we're just switching out the engine of the car. If the harness is the car and the car. If the harness is the car and the car. If the harness is the car and the engine's AI model, we're just the engine's AI model, we're just the engine's AI model, we're just switching out the engine. And you can switching out the engine. And you can switching out the engine. And you can see right here I have left the enthropic see right here I have left the enthropic see right here I have left the enthropic API key blank. We put this in as the API key blank. We put this in as the API key blank. We put this in as the enthropic o token and that is my Z API enthropic o token and that is my Z API enthropic o token and that is my Z API key. And then we changed all of these key. And then we changed all of these key. And then we changed all of these default models to GLM 5.2. And then when default models to GLM 5.2. And then when default models to GLM 5.2. And then when you open up your new claude, if I go in
-
you open up your new claude, if I go in you open up your new claude, if I go in here, claude, we open that up and what here, claude, we open that up and what here, claude, we open that up and what do we see right here? We see GLM 5.2 do we see right here? We see GLM 5.2 do we see right here? We see GLM 5.2 with 1 million context API usage with 1 million context API usage with 1 million context API usage billing. And so that is just like on a billing. And so that is just like on a billing. And so that is just like on a per project setting. So if I come into per project setting. So if I come into per project setting. So if I come into this other one and I show you guys like this other one and I show you guys like this other one and I show you guys like you know on this lefth hand side we have you know on this lefth hand side we have you know on this lefth hand side we have GLM. On this right hand side we have um GLM. On this right hand side we have um GLM. On this right hand side we have um opus. The way I did that is I just have opus. The way I did that is I just have opus. The way I did that is I just have these in two different directories. And these in two different directories. And these in two different directories. And this is a very ugly view. I'm sorry this is a very ugly view. I'm sorry this is a very ugly view. I'm sorry about that. But what I did is in here I about that. But what I did is in here I about that. But what I did is in here I have two folders. I have GLM and I have have two folders. I have GLM and I have have two folders. I have GLM and I have opus. So in the GLM folder, my opus. So in the GLM folder, my opus. So in the GLM folder, my settings.local.json settings.local.json settings.local.json shows this stuff, right? And then in my shows this stuff, right? And then in my shows this stuff, right? And then in my opus folder, I don't even have a opus folder, I don't even have a opus folder, I don't even have a settings.local.json. And that's how I'm settings.local.json. And that's how I'm settings.local.json. And that's how I'm able to open up claude in this able to open up claude in this able to open up claude in this directory. And we just have the regular directory. And we just have the regular directory. And we just have the regular Claude. And this looks absolutely awful. Claude. And this looks absolutely awful. Claude. And this looks absolutely awful. So let me just show you. There we go. So let me just show you. There we go. So let me just show you. There we go. Because we are not in a directory that Because we are not in a directory that Because we are not in a directory that has a settings.local.json JSON with this has a settings.local.json JSON with this has a settings.local.json JSON with this routing. It just pulls up automatically routing. It just pulls up automatically routing. It just pulls up automatically with Opus on my Cloud Max plan. So with Opus on my Cloud Max plan. So with Opus on my Cloud Max plan. So that's how you can sort of, you know, that's how you can sort of, you know, that's how you can sort of, you know, tweak which projects are using which AI tweak which projects are using which AI tweak which projects are using which AI model. And also, if you guys were model. And also, if you guys were model. And also, if you guys were wondering, this entire slide deck, of wondering, this entire slide deck, of wondering, this entire slide deck, of course, was built by GLM 5.2. It's funny course, was built by GLM 5.2. It's funny course, was built by GLM 5.2. It's funny cuz it keeps referring to me as Herk, cuz it keeps referring to me as Herk, cuz it keeps referring to me as Herk, and I think by this it just means Herk and I think by this it just means Herk and I think by this it just means Herk 2, like my project called Herk 2. But 2, like my project called Herk 2. But 2, like my project called Herk 2. But this was made by GLM 5.2. And you'll this was made by GLM 5.2. And you'll this was made by GLM 5.2. And you'll notice it looks like a lot of my other notice it looks like a lot of my other notice it looks like a lot of my other slide decks because it used my skill for slide decks because it used my skill for slide decks because it used my skill for that. Anyways, I know this one was a that. Anyways, I know this one was a that. Anyways, I know this one was a quick one, but I wanted to make it quick quick one, but I wanted to make it quick quick one, but I wanted to make it quick and show you guys how to get set up so and show you guys how to get set up so and show you guys how to get set up so you can play with it. I am planning on you can play with it. I am planning on you can play with it. I am planning on bringing you guys a ton more content on bringing you guys a ton more content on bringing you guys a ton more content on local models and maybe even some other local models and maybe even some other local models and maybe even some other stuff like open code because you don't stuff like open code because you don't stuff like open code because you don't always want to be locked into maybe
-
always want to be locked into maybe always want to be locked into maybe cloud code's hardest. So, please let me cloud code's hardest. So, please let me cloud code's hardest. So, please let me know in the comments what you want to know in the comments what you want to know in the comments what you want to see around open-source AI because trust see around open-source AI because trust see around open-source AI because trust me, I can definitely see a future where me, I can definitely see a future where me, I can definitely see a future where every company is just running their own every company is just running their own every company is just running their own local models. And I think Enthropic is local models. And I think Enthropic is local models. And I think Enthropic is starting to realize that too. OpenAI is starting to realize that too. OpenAI is starting to realize that too. OpenAI is starting to realize that too. That's why starting to realize that too. That's why starting to realize that too. That's why they're investing into other things like they're investing into other things like they're investing into other things like services and other ways that they can services and other ways that they can services and other ways that they can integrate into companies like their integrate into companies like their integrate into companies like their forward deployed engineers because they forward deployed engineers because they forward deployed engineers because they know that the model might not be the know that the model might not be the know that the model might not be the moat at the end of the day. Right now moat at the end of the day. Right now moat at the end of the day. Right now there's a huge gap, but we see this gap there's a huge gap, but we see this gap there's a huge gap, but we see this gap closing super quickly and it's really closing super quickly and it's really closing super quickly and it's really fun to watch in real time. So, let's fun to watch in real time. So, let's fun to watch in real time. So, let's start playing around more with open start playing around more with open start playing around more with open source models. If you guys enjoyed the source models. If you guys enjoyed the source models. If you guys enjoyed the video or you learned something new, video or you learned something new, video or you learned something new, please give a like. Helps me out a ton. please give a like. Helps me out a ton. please give a like. Helps me out a ton. And as always, I appreciate you guys And as always, I appreciate you guys And as always, I appreciate you guys making it to the end of the video and making it to the end of the video and making it to the end of the video and I'll see you all in the next one. Thanks I'll see you all in the next one. Thanks I'll see you all in the next one. Thanks everyone.
Summary
The video explores the capabilities of GLM 5.2 within Cloud Code, highlighting its speed, cost-effectiveness, and surprising design skills, even generating a video intro. It demonstrates GLM's efficiency in tasks like website design compared to Opus, noting that GLM 5.2 is significantly cheaper and faster for many applications. The practical takeaway is that GLM 5.2 offers a powerful and budget-friendly alternative for various AI-driven tasks.