← Back
Nate B. Jones August 21, 2026 20m

Stop Paying $200 For Work An $18 Model Can Do Inside Claude Code And Codex.

Read full transcript 16 segments
  1. We are all guilty of this one. Five We are all guilty of this one. Five different AI subscriptions, two of them different AI subscriptions, two of them different AI subscriptions, two of them alone can run you 400 bucks a month. And alone can run you 400 bucks a month. And alone can run you 400 bucks a month. And now, in the middle of all that, say it's now, in the middle of all that, say it's now, in the middle of all that, say it's 2:00 in the afternoon, regular day, you 2:00 in the afternoon, regular day, you 2:00 in the afternoon, regular day, you hit your limit on one of these. If you hit your limit on one of these. If you hit your limit on one of these. If you get Codex Pro, if you get Claude's top get Codex Pro, if you get Claude's top get Codex Pro, if you get Claude's top max plan, they're going to cost you $200 max plan, they're going to cost you $200 max plan, they're going to cost you $200 a month. Z.ai's GLM coding plan, on the a month. Z.ai's GLM coding plan, on the a month. Z.ai's GLM coding plan, on the other hand, starts at $18 a month and other hand, starts at $18 a month and other hand, starts at $18 a month and officially works inside both tools, officially works inside both tools, officially works inside both tools, which most people don't realize. GLM 5.3 which most people don't realize. GLM 5.3 which most people don't realize. GLM 5.3 can take even some of your coding work can take even some of your coding work can take even some of your coding work off the shoulders of Codex or Claude. off the shoulders of Codex or Claude. off the shoulders of Codex or Claude. The savings can be huge. You've used the The savings can be huge. You've used the The savings can be huge. You've used the API, you know the API costs add up real, API, you know the API costs add up real, API, you know the API costs add up real, real fast. That $18 model can run inside real fast. That $18 model can run inside real fast. That $18 model can run inside the tool you already know. And the setup the tool you already know. And the setup the tool you already know. And the setup just takes a few lines, and I'm going to just takes a few lines, and I'm going to just takes a few lines, and I'm going to walk you through it in this video. So, walk you through it in this video. So, walk you through it in this video. So, if that's you, if you want to try GLM if that's you, if you want to try GLM if that's you, if you want to try GLM 5.3 and you don't want to switch out 5.3 and you don't want to switch out 5.3 and you don't want to switch out your harness, this is the video for you. your harness, this is the video for you. your harness, this is the video for you. I'm going to show you all about it, tell I'm going to show you all about it, tell I'm going to show you all about it, tell you how to do it in each of them, tell you how to do it in each of them, tell you how to do it in each of them, tell you the trade-offs with Claude and you the trade-offs with Claude and you the trade-offs with Claude and Codex, and I'm relying entirely on both Codex, and I'm relying entirely on both Codex, and I'm relying entirely on both Claude and Codex documentation and also Claude and Codex documentation and also Claude and Codex documentation and also Z.ai documentation, the parent for GLM Z.ai documentation, the parent for GLM Z.ai documentation, the parent for GLM 5.3. So, let's jump into it. You can 5.3. So, let's jump into it. You can 5.3. So, let's jump into it. You can keep your files and your instructions keep your files and your instructions keep your files and your instructions and tools and permissions and hooks and and tools and permissions and hooks and and tools and permissions and hooks and habits, all of it that you already habits, all of it that you already habits, all of it that you already built, intact. All you do is change built, intact. All you do is change built, intact. All you do is change which company supplies the model for the which company supplies the model for the which company supplies the model for the piece of work that you're doing. Put piece of work that you're doing. Put piece of work that you're doing. Put yourself in the obvious situation. It's yourself in the obvious situation. It's yourself in the obvious situation. It's 2:00 in the afternoon, you're halfway 2:00 in the afternoon, you're halfway 2:00 in the afternoon, you're halfway through a real project, and you hit your through a real project, and you hit your through a real project, and you hit your Claude or Codex limit. Hi, it's me, I've Claude or Codex limit. Hi, it's me, I've Claude or Codex limit. Hi, it's me, I've done that. The job still needs to ship, done that. The job still needs to ship, done that. The job still needs to ship, but you don't want to pay more, pay the but you don't want to pay more, pay the but you don't want to pay more, pay the API pricing, which is way higher than API pricing, which is way higher than API pricing, which is way higher than your monthly bill, or spend the rest of

  2. your monthly bill, or spend the rest of your monthly bill, or spend the rest of the afternoon rebuilding everything in a the afternoon rebuilding everything in a the afternoon rebuilding everything in a strange new coding tool. So, you're kind strange new coding tool. So, you're kind strange new coding tool. So, you're kind of stuck. The real question is which of stuck. The real question is which of stuck. The real question is which work you move, what context comes with work you move, what context comes with work you move, what context comes with you, and whether running a cheaper model you, and whether running a cheaper model you, and whether running a cheaper model is still going to save you money once is still going to save you money once is still going to save you money once you validate the retries, once you do you validate the retries, once you do you validate the retries, once you do review, once once you actually get the review, once once you actually get the review, once once you actually get the fully loaded cost. So, this video gives fully loaded cost. So, this video gives fully loaded cost. So, this video gives you that answer for both Claude code and you that answer for both Claude code and you that answer for both Claude code and for Codex. I'm going to show you the for Codex. I'm going to show you the for Codex. I'm going to show you the setup, what follows you when you set up setup, what follows you when you set up setup, what follows you when you set up a new model, what doesn't, and the a new model, what doesn't, and the a new model, what doesn't, and the different way I would use GLM as a different way I would use GLM as a different way I would use GLM as a worker in each of those two tools. And worker in each of those two tools. And worker in each of those two tools. And then, we're going to sort through four then, we're going to sort through four then, we're going to sort through four real kinds of coding work and figure out real kinds of coding work and figure out real kinds of coding work and figure out what goes in the cheaper model queue and what goes in the cheaper model queue and what goes in the cheaper model queue and what do you save for your stronger model what do you save for your stronger model what do you save for your stronger model queue and why? And then start to measure queue and why? And then start to measure queue and why? And then start to measure what each coding result is actually what each coding result is actually what each coding result is actually going to cost you when you start to mix going to cost you when you start to mix going to cost you when you start to mix in that cheaper model and do it smart. in that cheaper model and do it smart. in that cheaper model and do it smart. So, stick with me. The $18 plan has So, stick with me. The $18 plan has So, stick with me. The $18 plan has smaller limits than the $200 plan. So, smaller limits than the $200 plan. So, smaller limits than the $200 plan. So, this is not saying that one subscription this is not saying that one subscription this is not saying that one subscription from z.ai just magically replaces a from z.ai just magically replaces a from z.ai just magically replaces a subscription to Codex or Claude. That's subscription to Codex or Claude. That's subscription to Codex or Claude. That's not what I'm proposing. That's not the not what I'm proposing. That's not the not what I'm proposing. That's not the way most people use it. The opportunity way most people use it. The opportunity way most people use it. The opportunity here is to stop paying the most here is to stop paying the most here is to stop paying the most expensive model to do every job merely expensive model to do every job merely expensive model to do every job merely because it came bundled with your coding because it came bundled with your coding because it came bundled with your coding tool. There are four things people keep tool. There are four things people keep tool. There are four things people keep mixing together here and I want to mixing together here and I want to mixing together here and I want to separate them out because I think separate them out because I think separate them out because I think they're often misunderstood. The first they're often misunderstood. The first they're often misunderstood. The first is the model. This might be Claude. It is the model. This might be Claude. It is the model. This might be Claude. It might be an OpenAI model. Or in this might be an OpenAI model. Or in this might be an OpenAI model. Or in this case, it might be GLM 5.3. It's the part case, it might be GLM 5.3. It's the part case, it might be GLM 5.3. It's the part that does the reasoning and produces the that does the reasoning and produces the that does the reasoning and produces the response in tokens. The second is, of response in tokens. The second is, of response in tokens. The second is, of course, the coding tool. It's often

  3. course, the coding tool. It's often course, the coding tool. It's often called a harness. Claude code and Codex called a harness. Claude code and Codex called a harness. Claude code and Codex are programs you operate that are are programs you operate that are are programs you operate that are harnesses. They let the model read files harnesses. They let the model read files harnesses. They let the model read files in a particular way, run commands, use in a particular way, run commands, use in a particular way, run commands, use tools, ask for permission, make changes, tools, ask for permission, make changes, tools, ask for permission, make changes, and show you the result, what happens. I and show you the result, what happens. I and show you the result, what happens. I often call this surrounding environment often call this surrounding environment often call this surrounding environment a harness because it helps you to a harness because it helps you to a harness because it helps you to understand how a model interacts with understand how a model interacts with understand how a model interacts with the setup or the arrangement. And you the setup or the arrangement. And you the setup or the arrangement. And you can think of it as you put your horse in can think of it as you put your horse in can think of it as you put your horse in the harness and it toes the cart, right? the harness and it toes the cart, right? the harness and it toes the cart, right? It's a visual metaphor, too. The third It's a visual metaphor, too. The third It's a visual metaphor, too. The third concept is your project context. And concept is your project context. And concept is your project context. And these are things you have written down. these are things you have written down. these are things you have written down. So, it might be your Claude.markdown So, it might be your Claude.markdown So, it might be your Claude.markdown file or your agents.markdown file. It file or your agents.markdown file. It file or your agents.markdown file. It might be your tasks, your documentation, might be your tasks, your documentation, might be your tasks, your documentation, your skills, your scripts, your hooks, your skills, your scripts, your hooks, your skills, your scripts, your hooks, your project rules. your project rules. your project rules. Another session can read those things Another session can read those things Another session can read those things again just because they live in files, again just because they live in files, again just because they live in files, right? They can be read by anything right? They can be read by anything right? They can be read by anything anywhere. They're very portable. And the anywhere. They're very portable. And the anywhere. They're very portable. And the fourth concept is conversation, right? fourth concept is conversation, right? fourth concept is conversation, right? That's the easiest to understand. It's That's the easiest to understand. It's That's the easiest to understand. It's the temporary history of your particular the temporary history of your particular the temporary history of your particular session, what you asked, what the model session, what you asked, what the model session, what you asked, what the model read, the decisions you made together, read, the decisions you made together, read, the decisions you made together, and the corrections you gave it along and the corrections you gave it along and the corrections you gave it along the way. Changing the model does not the way. Changing the model does not the way. Changing the model does not mean all four of those things change in mean all four of those things change in mean all four of those things change in exactly the same way. I really want to exactly the same way. I really want to exactly the same way. I really want to underline that three times cuz it gets underline that three times cuz it gets underline that three times cuz it gets misunderstood. And this distinction is misunderstood. And this distinction is misunderstood. And this distinction is why I've warned people that sometimes why I've warned people that sometimes why I've warned people that sometimes cheaper models can be more expensive.

  4. cheaper models can be more expensive. cheaper models can be more expensive. You have to look at the fully loaded You have to look at the fully loaded You have to look at the fully loaded cost across all of these four modalities cost across all of these four modalities cost across all of these four modalities to figure out what actually works. Flo to figure out what actually works. Flo to figure out what actually works. Flo Crivello's team at Lindy had moved off Crivello's team at Lindy had moved off Crivello's team at Lindy had moved off of Claude. It ended up rebuilding the of Claude. It ended up rebuilding the of Claude. It ended up rebuilding the harness around the work they actually harness around the work they actually harness around the work they actually did so they could take advantage of an did so they could take advantage of an did so they could take advantage of an open-source model. The lesson here is open-source model. The lesson here is open-source model. The lesson here is that you want to look at the whole that you want to look at the whole that you want to look at the whole system when you're making changes. So, system when you're making changes. So, system when you're making changes. So, let's talk about Claude Code. If you use let's talk about Claude Code. If you use let's talk about Claude Code. If you use Claude Code {slash} model command to Claude Code {slash} model command to Claude Code {slash} model command to choose another model inside the current choose another model inside the current choose another model inside the current provider, Claude Code keeps the provider, Claude Code keeps the provider, Claude Code keeps the conversation. You don't have to paste conversation. You don't have to paste conversation. You don't have to paste the whole thing again. But Anthropic the whole thing again. But Anthropic the whole thing again. But Anthropic warns that the next response will reread warns that the next response will reread warns that the next response will reread the whole conversation history without the whole conversation history without the whole conversation history without the old prompt caches, essentially the old prompt caches, essentially the old prompt caches, essentially loading up fresh in the background. And loading up fresh in the background. And loading up fresh in the background. And that can make a late switch to a new that can make a late switch to a new that can make a late switch to a new model slower and much more expensive model slower and much more expensive model slower and much more expensive than people expect, even if the overall than people expect, even if the overall than people expect, even if the overall model is cheaper. model is cheaper. model is cheaper. So, moving from Anthropic to Z.ai can be So, moving from Anthropic to Z.ai can be So, moving from Anthropic to Z.ai can be a bigger change, relatively speaking, a bigger change, relatively speaking, a bigger change, relatively speaking, because you're also changing effectively because you're also changing effectively because you're also changing effectively the internet address, the key Claude the internet address, the key Claude the internet address, the key Claude Code uses for model requests. There's a Code uses for model requests. There's a Code uses for model requests. There's a lot that's going on under the surface lot that's going on under the surface lot that's going on under the surface there. So, I would say the practical way there. So, I would say the practical way there. So, I would say the practical way to switch models if you're in Claude to switch models if you're in Claude to switch models if you're in Claude Code is simply to launch a separate Code is simply to launch a separate Code is simply to launch a separate GLM-specific session. It can open the GLM-specific session. It can open the GLM-specific session. It can open the same project. It can reload instructions same project. It can reload instructions same project. It can reload instructions you saved in files. All of that works, you saved in files. All of that works, you saved in files. All of that works, right? It doesn't automatically then right? It doesn't automatically then right? It doesn't automatically then inherit the mature conversation from the inherit the mature conversation from the inherit the mature conversation from the Anthropic session that you already had Anthropic session that you already had Anthropic session that you already had open. And so, this gives us maybe our open. And so, this gives us maybe our open. And so, this gives us maybe our first useful rule of thumb. Start a first useful rule of thumb. Start a first useful rule of thumb. Start a substantial job on the model that you substantial job on the model that you substantial job on the model that you expect to finish it. Don't build 40

  5. expect to finish it. Don't build 40 expect to finish it. Don't build 40 turns of working history with a model turns of working history with a model turns of working history with a model with a single provider and then casually with a single provider and then casually with a single provider and then casually move that job to another provider in the move that job to another provider in the move that job to another provider in the final mile. This doesn't mean that you final mile. This doesn't mean that you final mile. This doesn't mean that you can never use two models on one project. can never use two models on one project. can never use two models on one project. If you're stuck, maybe it's worth If you're stuck, maybe it's worth If you're stuck, maybe it's worth jumping over because you've hit your jumping over because you've hit your jumping over because you've hit your limit. It just means that the boundary limit. It just means that the boundary limit. It just means that the boundary between them should be a really clear between them should be a really clear between them should be a really clear piece of work and not just sort of a piece of work and not just sort of a piece of work and not just sort of a hope that the second model will walk hope that the second model will walk hope that the second model will walk into the conversation magically know into the conversation magically know into the conversation magically know everything that the first one learned. everything that the first one learned. everything that the first one learned. The work that I did on token saver made The work that I did on token saver made The work that I did on token saver made this distinction very real for me, this distinction very real for me, this distinction very real for me, right? I had a tracker. I had, you know, right? I had a tracker. I had, you know, right? I had a tracker. I had, you know, days where I was passing 3 billion, 4 days where I was passing 3 billion, 4 days where I was passing 3 billion, 4 billion tokens across Codex threads and billion tokens across Codex threads and billion tokens across Codex threads and Claude threads. Um and almost 96% of Claude threads. Um and almost 96% of Claude threads. Um and almost 96% of volume, as I shared, was reused input. volume, as I shared, was reused input. volume, as I shared, was reused input. And so in that world, you have to assume And so in that world, you have to assume And so in that world, you have to assume that the calls are carrying project that the calls are carrying project that the calls are carrying project instructions and tool definitions and instructions and tool definitions and instructions and tool definitions and file context and history and other file context and history and other file context and history and other repeated material. If those important repeated material. If those important repeated material. If those important work lessons from your projects exist in work lessons from your projects exist in work lessons from your projects exist in only one long conversation thread, only one long conversation thread, only one long conversation thread, they're going to be difficult to move to they're going to be difficult to move to they're going to be difficult to move to a new model. On the other hand, if the a new model. On the other hand, if the a new model. On the other hand, if the coding standards and test commands and coding standards and test commands and coding standards and test commands and permissions and definition of done live permissions and definition of done live permissions and definition of done live in files, another model can pick them up in files, another model can pick them up in files, another model can pick them up very easily. In other words, good very easily. In other words, good very easily. In other words, good hygiene here, context hygiene, makes a hygiene here, context hygiene, makes a hygiene here, context hygiene, makes a lot of economic sense. So, staying with lot of economic sense. So, staying with lot of economic sense. So, staying with Claude code, I would keep my normal Claude code, I would keep my normal Claude code, I would keep my normal Claude setup exactly as it is. I would Claude setup exactly as it is. I would Claude setup exactly as it is. I would create a second private launch command create a second private launch command create a second private launch command called something obvious like called something obvious like called something obvious like Claude-GLN.

  6. Claude-GLN. Claude-GLN. That command supplies three things That command supplies three things That command supplies three things before Claude code opens. The z.ai API before Claude code opens. The z.ai API before Claude code opens. The z.ai API key, the z.ai address for Anthropic key, the z.ai address for Anthropic key, the z.ai address for Anthropic compatible requests, and the names that compatible requests, and the names that compatible requests, and the names that map Claude's model choices to GLN 5.3. map Claude's model choices to GLN 5.3. map Claude's model choices to GLN 5.3. Keep the API key, obviously, in your own Keep the API key, obviously, in your own Keep the API key, obviously, in your own environment or a secret manager. Don't environment or a secret manager. Don't environment or a secret manager. Don't put it in the project. I keep saying put it in the project. I keep saying put it in the project. I keep saying that. And leave the normal Claude that. And leave the normal Claude that. And leave the normal Claude settings alone. Ordinary Claude will settings alone. Ordinary Claude will settings alone. Ordinary Claude will still open an Anthropic session, while still open an Anthropic session, while still open an Anthropic session, while Claude GLM will open a z.ai session. If Claude GLM will open a z.ai session. If Claude GLM will open a z.ai session. If the GLM connection behaves badly, just the GLM connection behaves badly, just the GLM connection behaves badly, just close it down and go back to the normal close it down and go back to the normal close it down and go back to the normal command. If this all sounds like Greek command. If this all sounds like Greek command. If this all sounds like Greek to you, I've put the copy and paste to you, I've put the copy and paste to you, I've put the copy and paste command in the companion guide to this command in the companion guide to this command in the companion guide to this piece over on Substack, because we piece over on Substack, because we piece over on Substack, because we should not be trying to transcribe should not be trying to transcribe should not be trying to transcribe environment variables, environment variables, environment variables, and we should not be playing those kinds and we should not be playing those kinds and we should not be playing those kinds of games. So, if you have any doubt of games. So, if you have any doubt of games. So, if you have any doubt about whether you can do that, stick about whether you can do that, stick about whether you can do that, stick with a secure, safe command, and don't with a secure, safe command, and don't with a secure, safe command, and don't try and handle API secrets yourself on try and handle API secrets yourself on try and handle API secrets yourself on your own. Okay, what happens after the your own. Okay, what happens after the your own. Okay, what happens after the command runs? You want to open the same command runs? You want to open the same command runs? You want to open the same repository, and Claude code will see the repository, and Claude code will see the repository, and Claude code will see the same files, it will read the same same files, it will read the same same files, it will read the same Claude.markdown, and it will keep the Claude.markdown, and it will keep the Claude.markdown, and it will keep the hooks and the MCP servers, the tools, hooks and the MCP servers, the tools, hooks and the MCP servers, the tools, the permissions experience that you've the permissions experience that you've the permissions experience that you've already configured. All the stuff that already configured. All the stuff that already configured. All the stuff that goes into your Claude harness.

  7. goes into your Claude harness. goes into your Claude harness. The model answering underneath will now The model answering underneath will now The model answering underneath will now be GLM 5.3. And what comes back is the be GLM 5.3. And what comes back is the be GLM 5.3. And what comes back is the project context that you saved in files. project context that you saved in files. project context that you saved in files. Now, what won't come back is any old Now, what won't come back is any old Now, what won't come back is any old Anthropic conversation you had, any old Anthropic conversation you had, any old Anthropic conversation you had, any old prompt cache, or decision that you never prompt cache, or decision that you never prompt cache, or decision that you never wrote anywhere outside the chat. Well, wrote anywhere outside the chat. Well, wrote anywhere outside the chat. Well, that's gone, right? Suppose your that's gone, right? Suppose your that's gone, right? Suppose your Anthropic session spent an hour Anthropic session spent an hour Anthropic session spent an hour investigating a bug, right? It ruled out investigating a bug, right? It ruled out investigating a bug, right? It ruled out three causes, it learned that one log three causes, it learned that one log three causes, it learned that one log line is misleading, it agreed not to line is misleading, it agreed not to line is misleading, it agreed not to touch the authentication middleware you touch the authentication middleware you touch the authentication middleware you have. Those are all decisions, right? If have. Those are all decisions, right? If have. Those are all decisions, right? If none of that got logged into a file, a none of that got logged into a file, a none of that got logged into a file, a fresh session with a new model like GLM fresh session with a new model like GLM fresh session with a new model like GLM 5.3, it's just not going to know about 5.3, it's just not going to know about 5.3, it's just not going to know about it. And so, you're going to start from it. And so, you're going to start from it. And so, you're going to start from scratch. So, before moving a job over, scratch. So, before moving a job over, scratch. So, before moving a job over, if you have to move a job in the middle, if you have to move a job in the middle, if you have to move a job in the middle, create a handoff file. Insist that the create a handoff file. Insist that the create a handoff file. Insist that the current model document goal, current current model document goal, current current model document goal, current state, relevant files, constraints, what state, relevant files, constraints, what state, relevant files, constraints, what done means, and the checks to run. And done means, and the checks to run. And done means, and the checks to run. And yes, I have a template for that in the yes, I have a template for that in the yes, I have a template for that in the Substack guide, as well. So, for Substack guide, as well. So, for Substack guide, as well. So, for example, update these 38 API calls to example, update these 38 API calls to example, update these 38 API calls to the new field name. That might be your the new field name. That might be your the new field name. That might be your goal. The current branch is clean and goal. The current branch is clean and goal. The current branch is clean and the affected calls are in these two the affected calls are in these two the affected calls are in these two folders. That might be where the current folders. That might be where the current folders. That might be where the current state is. Don't change the public API as state is. Don't change the public API as state is. Don't change the public API as a constraint, and the job is done when a constraint, and the job is done when a constraint, and the job is done when the old field no longer appears anywhere the old field no longer appears anywhere the old field no longer appears anywhere in the existing tests all pass. And then in the existing tests all pass. And then in the existing tests all pass. And then it will say, "Run X commands before it will say, "Run X commands before it will say, "Run X commands before returning to work." Like these are returning to work." Like these are returning to work." Like these are commands that we're in the middle of.

  8. commands that we're in the middle of. commands that we're in the middle of. That is enough context to get a new That is enough context to get a new That is enough context to get a new model going. It's also much better than model going. It's also much better than model going. It's also much better than pasting an entire conversation history pasting an entire conversation history pasting an entire conversation history and asking GLM 5.3 or any new model to and asking GLM 5.3 or any new model to and asking GLM 5.3 or any new model to discover which parts matter along the discover which parts matter along the discover which parts matter along the way. And now we get to Claude Code's way. And now we get to Claude Code's way. And now we get to Claude Code's multi-agent question. Claude Code multi-agent question. Claude Code multi-agent question. Claude Code already uses and has sub-agents, right? already uses and has sub-agents, right? already uses and has sub-agents, right? But a normal sub-agent starts with fresh But a normal sub-agent starts with fresh But a normal sub-agent starts with fresh context. It receives the task that context. It receives the task that context. It receives the task that Claude delegates and applicable project Claude delegates and applicable project Claude delegates and applicable project instructions, not a complete parent instructions, not a complete parent instructions, not a complete parent conversation or every file the parent conversation or every file the parent conversation or every file the parent has read. And that's on purpose, right? has read. And that's on purpose, right? has read. And that's on purpose, right? It gives the agent bounded context. And It gives the agent bounded context. And It gives the agent bounded context. And that's one reason sub-agents are really that's one reason sub-agents are really that's one reason sub-agents are really good at specific narrow work. They can good at specific narrow work. They can good at specific narrow work. They can keep their logs and research and side keep their logs and research and side keep their logs and research and side investigations outside the main investigations outside the main investigations outside the main conversation. So, Claude Code also has conversation. So, Claude Code also has conversation. So, Claude Code also has the concept of forked sub-agents, and the concept of forked sub-agents, and the concept of forked sub-agents, and that's a different concept. A fork does that's a different concept. A fork does that's a different concept. A fork does receive the full conversation and can receive the full conversation and can receive the full conversation and can reuse the parent's prompt cache. The reuse the parent's prompt cache. The reuse the parent's prompt cache. The trade-off is that the fork must use the trade-off is that the fork must use the trade-off is that the fork must use the same model as the parent. That leaves a same model as the parent. That leaves a same model as the parent. That leaves a real limitation for this specific setup.

  9. real limitation for this specific setup. real limitation for this specific setup. Claude Code documents how a sub-agent Claude Code documents how a sub-agent Claude Code documents how a sub-agent chooses a model, but it does not chooses a model, but it does not chooses a model, but it does not document a separate provider address for document a separate provider address for document a separate provider address for one sub-agent. If the main Claude one sub-agent. If the main Claude one sub-agent. If the main Claude Code process is using Anthropic, there's Code process is using Anthropic, there's Code process is using Anthropic, there's not a simple native setting that says, not a simple native setting that says, not a simple native setting that says, "Keep Anthropic in charge, but send this "Keep Anthropic in charge, but send this "Keep Anthropic in charge, but send this one child to that with a gateway or a one child to that with a gateway or a one child to that with a gateway or a custom integration, but I would not make custom integration, but I would not make custom integration, but I would not make that a recommendation for beginners. that a recommendation for beginners. that a recommendation for beginners. Instead, I would open up two Claude Code Instead, I would open up two Claude Code Instead, I would open up two Claude Code sessions and ordinary Claude is the lead sessions and ordinary Claude is the lead sessions and ordinary Claude is the lead and Claude GLM as the worker. Give the and Claude GLM as the worker. Give the and Claude GLM as the worker. Give the GLM session that six-line handoff I GLM session that six-line handoff I GLM session that six-line handoff I talked about if both sessions will edit talked about if both sessions will edit talked about if both sessions will edit at the same time, put the worker in a at the same time, put the worker in a at the same time, put the worker in a Git work tree. All that is by the way is Git work tree. All that is by the way is Git work tree. All that is by the way is a separate copy of the repository. It a separate copy of the repository. It a separate copy of the repository. It prevents both agents from changing the prevents both agents from changing the prevents both agents from changing the same files under each other. And the GLM same files under each other. And the GLM same files under each other. And the GLM session just returns changed files and session just returns changed files and session just returns changed files and checks that it ran and anything it can't checks that it ran and anything it can't checks that it ran and anything it can't resolve, it goes and fixes. The resolve, it goes and fixes. The resolve, it goes and fixes. The Anthropic session will review the result Anthropic session will review the result Anthropic session will review the result when the job is consequential. when the job is consequential. when the job is consequential. And it's still one project and it's And it's still one project and it's And it's still one project and it's still one familiar tool. It's just two still one familiar tool. It's just two still one familiar tool. It's just two sessions with a very explicit handoff sessions with a very explicit handoff sessions with a very explicit handoff between them. Okay. Now let's tackle between them. Okay. Now let's tackle between them. Okay. Now let's tackle Codex because Codex gives us a different Codex because Codex gives us a different Codex because Codex gives us a different option. The first Codex setup looks a option. The first Codex setup looks a option. The first Codex setup looks a lot like Claude Code. You add z.ai as a lot like Claude Code. You add z.ai as a lot like Claude Code. You add z.ai as a model provider in your personal Codex model provider in your personal Codex model provider in your personal Codex configuration. In plain language, you configuration. In plain language, you configuration. In plain language, you give Codex the z.ai address and tell it give Codex the z.ai address and tell it give Codex the z.ai address and tell it which environment variable contains the which environment variable contains the which environment variable contains the key. z.ai provides a responses key. z.ai provides a responses key. z.ai provides a responses compatible address specifically for

  10. compatible address specifically for compatible address specifically for Codex and you create a GLM profile that Codex and you create a GLM profile that Codex and you create a GLM profile that says two things. Use GLM 5.3 and send says two things. Use GLM 5.3 and send says two things. Use GLM 5.3 and send the request through that provider. When the request through that provider. When the request through that provider. When you want an entire Codex job to run on you want an entire Codex job to run on you want an entire Codex job to run on GLM, you just say, "Hey, launch Codex GLM, you just say, "Hey, launch Codex GLM, you just say, "Hey, launch Codex profile GLM." And ordinary Codex will profile GLM." And ordinary Codex will profile GLM." And ordinary Codex will still use your normal OpenAI setup. So still use your normal OpenAI setup. So still use your normal OpenAI setup. So you can run both at once. This is you can run both at once. This is you can run both at once. This is similar to how we configured Anthropic, similar to how we configured Anthropic, similar to how we configured Anthropic, right? The GLM profile opens the same right? The GLM profile opens the same right? The GLM profile opens the same project, loads the same applicable project, loads the same applicable project, loads the same applicable agents.markdown files, skills, tools, agents.markdown files, skills, tools, agents.markdown files, skills, tools, project rules, all the things you're project rules, all the things you're project rules, all the things you're familiar with. And as with Claude Code, familiar with. And as with Claude Code, familiar with. And as with Claude Code, you want to treat it as a new job or a you want to treat it as a new job or a you want to treat it as a new job or a carefully handed off job. So the project carefully handed off job. So the project carefully handed off job. So the project context will reload. A separate context will reload. A separate context will reload. A separate conversation does not magically appear. conversation does not magically appear. conversation does not magically appear. Now, why does all of this matter? Let's Now, why does all of this matter? Let's Now, why does all of this matter? Let's go back to imagining a small software go back to imagining a small software go back to imagining a small software company. What are they doing? They have company. What are they doing? They have company. What are they doing? They have a technical founder, a senior engineer, a technical founder, a senior engineer, a technical founder, a senior engineer, a backlog that grows faster than the a backlog that grows faster than the a backlog that grows faster than the team because they've got customers with team because they've got customers with team because they've got customers with bugs and they've spent months getting bugs and they've spent months getting bugs and they've spent months getting Claude Coder Codex to understand their Claude Coder Codex to understand their Claude Coder Codex to understand their repository.

  11. repository. repository. And they have a bunch of jobs that they And they have a bunch of jobs that they And they have a bunch of jobs that they have to work on and they have token have to work on and they have token have to work on and they have token bills. This is the startup that I've had bills. This is the startup that I've had bills. This is the startup that I've had in mind as I've given you examples in mind as I've given you examples in mind as I've given you examples through this video. So we've talked through this video. So we've talked through this video. So we've talked about the agent having to update 38 about the agent having to update 38 about the agent having to update 38 calls to a renamed API field. Very calls to a renamed API field. Very calls to a renamed API field. Very plausible startup job. We've talked plausible startup job. We've talked plausible startup job. We've talked about the idea that there might be an about the idea that there might be an about the idea that there might be an intermittent authentication failure. intermittent authentication failure. intermittent authentication failure. That these are jobs that I've referred That these are jobs that I've referred That these are jobs that I've referred to through this video. So we've talked to through this video. So we've talked to through this video. So we've talked about the idea that the agent might have about the idea that the agent might have about the idea that the agent might have to update API calls, right? There's lots to update API calls, right? There's lots to update API calls, right? There's lots of other jobs like that that startup of other jobs like that that startup of other jobs like that that startup technical teams have to address like technical teams have to address like technical teams have to address like intermittent authentication failures. intermittent authentication failures. intermittent authentication failures. These are not glamorous things, right? These are not glamorous things, right? These are not glamorous things, right? When you are looking for how to farm out When you are looking for how to farm out When you are looking for how to farm out this work, I want you to think about it this work, I want you to think about it this work, I want you to think about it as where do you have clear definition of as where do you have clear definition of as where do you have clear definition of done, where do you have a repository done, where do you have a repository done, where do you have a repository that contains examples, and where do you that contains examples, and where do you that contains examples, and where do you have really good tests that can tell you have really good tests that can tell you have really good tests that can tell you whether the work is acceptable. In those whether the work is acceptable. In those whether the work is acceptable. In those situations, I would start to use situations, I would start to use situations, I would start to use whatever harness you're running, whether whatever harness you're running, whether whatever harness you're running, whether you're Claude or Codex, and I would you're Claude or Codex, and I would you're Claude or Codex, and I would start to assign that work out to GLM start to assign that work out to GLM start to assign that work out to GLM because it's work that I feel really because it's work that I feel really because it's work that I feel really good about GLM attacking and getting good about GLM attacking and getting good about GLM attacking and getting right. It's not too vague, it's not too right. It's not too vague, it's not too right. It's not too vague, it's not too generalizable. GLM's going to be able to generalizable. GLM's going to be able to generalizable. GLM's going to be able to go after it and you're going to be able go after it and you're going to be able go after it and you're going to be able to save a fair bit on tokens. If you're to save a fair bit on tokens. If you're to save a fair bit on tokens. If you're doing something fairly complex where doing something fairly complex where doing something fairly complex where you're root causing something like that you're root causing something like that you're root causing something like that authentication failure I talked about, authentication failure I talked about, authentication failure I talked about, that could be different because it has that could be different because it has that could be different because it has hidden state, it may have conflicting hidden state, it may have conflicting hidden state, it may have conflicting evidence as you dig again. A cheap evidence as you dig again. A cheap evidence as you dig again. A cheap worker might still help gather logs or worker might still help gather logs or worker might still help gather logs or trace code paths, but I would keep the trace code paths, but I would keep the trace code paths, but I would keep the main investigation with the strongest main investigation with the strongest main investigation with the strongest model I could trust. You see how you model I could trust. You see how you model I could trust. You see how you need your human judgment to think this need your human judgment to think this need your human judgment to think this through. In a sense, you are figuring

  12. through. In a sense, you are figuring through. In a sense, you are figuring out which model inside the harness makes out which model inside the harness makes out which model inside the harness makes sense to use in which situation. And I sense to use in which situation. And I sense to use in which situation. And I want to give you kind of some rules of want to give you kind of some rules of want to give you kind of some rules of thumb so you have that. So you'll use thumb so you have that. So you'll use thumb so you have that. So you'll use GLM when the job has a clear target, it GLM when the job has a clear target, it GLM when the job has a clear target, it has really clear permissions, and it has has really clear permissions, and it has has really clear permissions, and it has a a test objective that you can define a a test objective that you can define a a test objective that you can define really specifically. And you'll keep the really specifically. And you'll keep the really specifically. And you'll keep the smarter model in charge when the hard smarter model in charge when the hard smarter model in charge when the hard part is deciding what the job should be part is deciding what the job should be part is deciding what the job should be at all or resolving hidden state or at all or resolving hidden state or at all or resolving hidden state or investigating or weighing a risky investigating or weighing a risky investigating or weighing a risky trade-off. How difficult it is to hand trade-off. How difficult it is to hand trade-off. How difficult it is to hand off work midstream. If you have to pull off work midstream. If you have to pull off work midstream. If you have to pull an entire transcript from the parent an entire transcript from the parent an entire transcript from the parent conversation, the job is probably too conversation, the job is probably too conversation, the job is probably too unbounded to handle giving to a model unbounded to handle giving to a model unbounded to handle giving to a model like GLM. And that's also why I would like GLM. And that's also why I would like GLM. And that's also why I would not switch models repeatedly inside a not switch models repeatedly inside a not switch models repeatedly inside a conversation. Be intentional, right? A conversation. Be intentional, right? A conversation. Be intentional, right? A tool may keep visible history that you tool may keep visible history that you tool may keep visible history that you can read, but you also have to factor in can read, but you also have to factor in can read, but you also have to factor in stuff you can't read like model behavior stuff you can't read like model behavior stuff you can't read like model behavior change over the course of the change over the course of the change over the course of the conversation, prompt caching, etc. So, a conversation, prompt caching, etc. So, a conversation, prompt caching, etc. So, a new provider has to figure all of that new provider has to figure all of that new provider has to figure all of that out and that can add cost, that can add out and that can add cost, that can add out and that can add cost, that can add uncertainty, that can add more turns.

  13. uncertainty, that can add more turns. uncertainty, that can add more turns. All of that has to be factored in when All of that has to be factored in when All of that has to be factored in when you're figuring out where to apply these you're figuring out where to apply these you're figuring out where to apply these models. The other thing to figure out is models. The other thing to figure out is models. The other thing to figure out is whether the work actually is cheaper whether the work actually is cheaper whether the work actually is cheaper with an alternate model. The plans are with an alternate model. The plans are with an alternate model. The plans are not equivalent. As I called out, Z.ai not equivalent. As I called out, Z.ai not equivalent. As I called out, Z.ai has both 5-hour and weekly credit limits has both 5-hour and weekly credit limits has both 5-hour and weekly credit limits and the $18 tier is not the same token and the $18 tier is not the same token and the $18 tier is not the same token count or same problem-solving ability as count or same problem-solving ability as count or same problem-solving ability as the $200 plan. And so, you want to be the $200 plan. And so, you want to be the $200 plan. And so, you want to be thinking about in your particular thinking about in your particular thinking about in your particular instance, for your codebase, for your instance, for your codebase, for your instance, for your codebase, for your tasks, tasks, tasks, does this model make sense? And I wish I does this model make sense? And I wish I does this model make sense? And I wish I could sit there next to you and say, could sit there next to you and say, could sit there next to you and say, "This makes sense, this doesn't." But "This makes sense, this doesn't." But "This makes sense, this doesn't." But the reality is right now, you are the the reality is right now, you are the the reality is right now, you are the one that needs to figure that out for one that needs to figure that out for one that needs to figure that out for your tasks, your files, your code, and your tasks, your files, your code, and your tasks, your files, your code, and there's no substitute for actually there's no substitute for actually there's no substitute for actually testing. The best I can do is give you a testing. The best I can do is give you a testing. The best I can do is give you a rule of thumb, give you very clear rule of thumb, give you very clear rule of thumb, give you very clear instructions for how to set it up, and instructions for how to set it up, and instructions for how to set it up, and remind you always test on your own code remind you always test on your own code remind you always test on your own code because your code has its own level of because your code has its own level of because your code has its own level of complexity, and you need to make sure complexity, and you need to make sure complexity, and you need to make sure you're trueing up that level of you're trueing up that level of you're trueing up that level of complexity to the model itself. And complexity to the model itself. And complexity to the model itself. And don't be under ambitious. I don't want don't be under ambitious. I don't want don't be under ambitious. I don't want you to hear this and say, "Oh, I won't you to hear this and say, "Oh, I won't you to hear this and say, "Oh, I won't give it tough tasks." I would be really give it tough tasks." I would be really give it tough tasks." I would be really intentional about it. I would look and intentional about it. I would look and intentional about it. I would look and say, "Let's give it a couple of say, "Let's give it a couple of say, "Let's give it a couple of ambitious tasks in my code base. Let's ambitious tasks in my code base. Let's ambitious tasks in my code base. Let's see how it does. And let's back off as see how it does. And let's back off as see how it does. And let's back off as we see failure and see where the true up we see failure and see where the true up we see failure and see where the true up level for this particular model is. And level for this particular model is. And level for this particular model is. And by the way, that's also something you by the way, that's also something you by the way, that's also something you can use on less frontier versions of can use on less frontier versions of can use on less frontier versions of existing models from OpenAI and from existing models from OpenAI and from existing models from OpenAI and from Claude. So, if you want to use a Sonnet Claude. So, if you want to use a Sonnet Claude. So, if you want to use a Sonnet model, if you want to use the Terra or model, if you want to use the Terra or model, if you want to use the Terra or the Luna models, that's the same

  14. the Luna models, that's the same the Luna models, that's the same principle. You want to give them stuff principle. You want to give them stuff principle. You want to give them stuff they can tackle, see how complex they they can tackle, see how complex they they can tackle, see how complex they can go, and then back them off as you can go, and then back them off as you can go, and then back them off as you need to. This is going to ultimately need to. This is going to ultimately need to. This is going to ultimately give you your maximum token savings give you your maximum token savings give you your maximum token savings without costing you on intelligence without costing you on intelligence without costing you on intelligence capabilities. You know, when I step back capabilities. You know, when I step back capabilities. You know, when I step back here, I've spent a long, long time, here, I've spent a long, long time, here, I've spent a long, long time, years and years building products all years and years building products all years and years building products all the way from, you know, at the 100 the way from, you know, at the 100 the way from, you know, at the 100 million person and more scale down to million person and more scale down to million person and more scale down to tiny startups. tiny startups. tiny startups. And I think what's interesting about And I think what's interesting about And I think what's interesting about this moment is that for so long software this moment is that for so long software this moment is that for so long software companies have made bundles feel companies have made bundles feel companies have made bundles feel inevitable. The interface, the workflow, inevitable. The interface, the workflow, inevitable. The interface, the workflow, the underlying service all arrive the underlying service all arrive the underlying service all arrive bundled, and you can't pry them apart. bundled, and you can't pry them apart. bundled, and you can't pry them apart. And over time you stop asking whether And over time you stop asking whether And over time you stop asking whether they have to come from different they have to come from different they have to come from different companies cuz you just trained. And I companies cuz you just trained. And I companies cuz you just trained. And I think one of the things I really think one of the things I really think one of the things I really appreciate about both the Claude team appreciate about both the Claude team appreciate about both the Claude team and the Codex team is that both of them and the Codex team is that both of them and the Codex team is that both of them feel like they are committing to being feel like they are committing to being feel like they are committing to being willing to unbundle. And I think that willing to unbundle. And I think that willing to unbundle. And I think that that's a strong mark in favor of both that's a strong mark in favor of both that's a strong mark in favor of both teams. Claude code often feels like a teams. Claude code often feels like a teams. Claude code often feels like a cockpit right now where I have to stay cockpit right now where I have to stay cockpit right now where I have to stay close to the work and I steer it and I close to the work and I steer it and I close to the work and I steer it and I change direction as the problem becomes change direction as the problem becomes change direction as the problem becomes clear. Codex can feel more like kind of clear. Codex can feel more like kind of clear. Codex can feel more like kind of an ops desk, right? I can dispatch jobs, an ops desk, right? I can dispatch jobs, an ops desk, right? I can dispatch jobs, I can let them run, I can expect inspect I can let them run, I can expect inspect I can let them run, I can expect inspect what comes back, and I kind of see where what comes back, and I kind of see where what comes back, and I kind of see where I'm at. Those work styles don't I'm at. Those work styles don't I'm at. Those work styles don't disappear because another company disappear because another company disappear because another company supplies some of the intelligence, supplies some of the intelligence, supplies some of the intelligence, because the they're kind of endogenous because the they're kind of endogenous because the they're kind of endogenous to the harness. They fit with the to the harness. They fit with the to the harness. They fit with the harness. I can keep Claude code, and I harness. I can keep Claude code, and I harness. I can keep Claude code, and I can keep working with it the way I can keep working with it the way I can keep working with it the way I always do, and I'm just steering a GLM always do, and I'm just steering a GLM always do, and I'm just steering a GLM session along the way. And so,

  15. session along the way. And so, session along the way. And so, unbundling allows us to see the unbundling allows us to see the unbundling allows us to see the importance of the harness in the way we importance of the harness in the way we importance of the harness in the way we work. So, wrapping up here, I've put the work. So, wrapping up here, I've put the work. So, wrapping up here, I've put the Claude launcher, the Codex profile, the Claude launcher, the Codex profile, the Claude launcher, the Codex profile, the handoff templates, the comparison handoff templates, the comparison handoff templates, the comparison scorecards all in the companion guide on scorecards all in the companion guide on scorecards all in the companion guide on Substack. You can get them if you want Substack. You can get them if you want Substack. You can get them if you want to get started with a cheaper model to get started with a cheaper model to get started with a cheaper model inside your Codex or your Claude and inside your Codex or your Claude and inside your Codex or your Claude and save big. That's where you can go to get save big. That's where you can go to get save big. That's where you can go to get started. Those are much easier to copy started. Those are much easier to copy started. Those are much easier to copy from a page than a video, so I would go from a page than a video, so I would go from a page than a video, so I would go there if you want to do that. Listen, there if you want to do that. Listen, there if you want to do that. Listen, start with a piece of work that you care start with a piece of work that you care start with a piece of work that you care about. Figure out what you can challenge about. Figure out what you can challenge about. Figure out what you can challenge a model with that you haven't tried a model with that you haven't tried a model with that you haven't tried before with a with a subsidiary model or before with a with a subsidiary model or before with a with a subsidiary model or a new kind of open-source model and give a new kind of open-source model and give a new kind of open-source model and give it a shot. This is I I've given you how it a shot. This is I I've given you how it a shot. This is I I've given you how you can figure it. Now you have to have you can figure it. Now you have to have you can figure it. Now you have to have the boldness to actually go out there the boldness to actually go out there the boldness to actually go out there and try something interesting. And when and try something interesting. And when and try something interesting. And when you see failure, just kind of back off you see failure, just kind of back off you see failure, just kind of back off it a little bit. My rule of thumb is I it a little bit. My rule of thumb is I it a little bit. My rule of thumb is I find that bounded tasks that have clear find that bounded tasks that have clear find that bounded tasks that have clear descriptions of done tend to work descriptions of done tend to work descriptions of done tend to work better. Your mileage may vary because better. Your mileage may vary because better. Your mileage may vary because your your code complexity may vary. So, your your code complexity may vary. So, your your code complexity may vary. So, you have to look at that on your own. you have to look at that on your own. you have to look at that on your own. Best of luck to you. Go grab those Best of luck to you. Go grab those Best of luck to you. Go grab those companion guides and I will see you next companion guides and I will see you next companion guides and I will see you next time. And let me know what you are time. And let me know what you are time. And let me know what you are building with your multi-agent setups.

  16. building with your multi-agent setups. building with your multi-agent setups. What are you building with GLM 5.3? What are you building with GLM 5.3? What are you building with GLM 5.3? Which harness are you choosing for it? Which harness are you choosing for it? Which harness are you choosing for it? Are you choosing the native z.ai Are you choosing the native z.ai Are you choosing the native z.ai harness? Are you choosing Codex? Are you harness? Are you choosing Codex? Are you harness? Are you choosing Codex? Are you choosing Claude code? I'd be curious to choosing Claude code? I'd be curious to choosing Claude code? I'd be curious to hear where people come down on that and hear where people come down on that and hear where people come down on that and then how much money are you saving? You then how much money are you saving? You then how much money are you saving? You saving 200 bucks? Are you saving Cuz let saving 200 bucks? Are you saving Cuz let saving 200 bucks? Are you saving Cuz let me tell you, if you've used the API, you me tell you, if you've used the API, you me tell you, if you've used the API, you know the API costs add up real real know the API costs add up real real know the API costs add up real real fast. So, if this is going to save you fast. So, if this is going to save you fast. So, if this is going to save you API costs, it's going to save you big API costs, it's going to save you big API costs, it's going to save you big money. All right. money. All right. money. All right. Let me know what you build and let's get Let me know what you build and let's get Let me know what you build and let's get to it.

Summary

The transcript discusses the high cost and limitations of AI subscriptions like Codex Pro and Claude, referencing their pricing and API costs. The practical takeaway is to consider Z.ai's GLM 5.3, which offers a significantly cheaper alternative at $18/month and integrates with existing tools, allowing users to maintain their current workflows while reducing expenses.

View original episode ↗