← Back
Nate B. Jones July 13, 2026 13m

Your Next AI Subscription Shouldn't Be ChatGPT 5.6 Or Fable 5. It Should Be Both.

Read full transcript 11 segments
  1. Chat GPT 5.6 is a dumber model, and I Chat GPT 5.6 is a dumber model, and I love it so much. In fact, I use it all love it so much. In fact, I use it all love it so much. In fact, I use it all the time. Today, I want to tell you the time. Today, I want to tell you the time. Today, I want to tell you about which model works for you. Not about which model works for you. Not about which model works for you. Not which model works for me. I will tell which model works for me. I will tell which model works for me. I will tell you which model works for me, don't you which model works for me, don't you which model works for me, don't worry. More importantly, I'm going to worry. More importantly, I'm going to worry. More importantly, I'm going to tell you how to pick the model that tell you how to pick the model that tell you how to pick the model that works for you and why, and why your works for you and why, and why your works for you and why, and why your heuristic, why the thing you use to pick heuristic, why the thing you use to pick heuristic, why the thing you use to pick that model is not anybody's benchmark that model is not anybody's benchmark that model is not anybody's benchmark score, including mine. And I do score, including mine. And I do score, including mine. And I do benchmark these models on a private benchmark these models on a private benchmark these models on a private benchmark suite, and I'll share that, benchmark suite, and I'll share that, benchmark suite, and I'll share that, but that's not the point. The point is but that's not the point. The point is but that's not the point. The point is for you to have the tools to pick the for you to have the tools to pick the for you to have the tools to pick the model that works for you. So, let's get model that works for you. So, let's get model that works for you. So, let's get into it. Last night, I reached for Chat into it. Last night, I reached for Chat into it. Last night, I reached for Chat GPT 5.6 Soul, even though I think it's GPT 5.6 Soul, even though I think it's GPT 5.6 Soul, even though I think it's the dumber model. Now, dumber does not the dumber model. Now, dumber does not the dumber model. Now, dumber does not mean dumb, not remotely. Soul is an mean dumb, not remotely. Soul is an mean dumb, not remotely. Soul is an incredibly intelligent model. On Agent incredibly intelligent model. On Agent incredibly intelligent model. On Agent Slash exam, which measures long-running Slash exam, which measures long-running Slash exam, which measures long-running professional work across 55 different professional work across 55 different professional work across 55 different fields, Soul set a new high. And on my fields, Soul set a new high. And on my fields, Soul set a new high. And on my own benchmark, Soul scored 93 on Dingo, own benchmark, Soul scored 93 on Dingo, own benchmark, Soul scored 93 on Dingo, which is my knowledge work package. I've which is my knowledge work package. I've which is my knowledge work package. I've talked about it before. It's a really, talked about it before. It's a really, talked about it before. It's a really, really good model for doing complicated really good model for doing complicated really good model for doing complicated knowledge work. Dingo, in this case, is knowledge work. Dingo, in this case, is knowledge work. Dingo, in this case, is a test that measures whether a model can a test that measures whether a model can a test that measures whether a model can propose a business startup idea selling propose a business startup idea selling propose a business startup idea selling Dingo dogs in Alaska and get around the Dingo dogs in Alaska and get around the Dingo dogs in Alaska and get around the legal issues, get around the regulatory legal issues, get around the regulatory legal issues, get around the regulatory issues, and get into the marketing side issues, and get into the marketing side issues, and get into the marketing side of things. It's a really funny package of things. It's a really funny package of things. It's a really funny package on purpose because I like humor. It also on purpose because I like humor. It also on purpose because I like humor. It also tests whether a model is smart across a tests whether a model is smart across a tests whether a model is smart across a wide range of professional knowledge wide range of professional knowledge wide range of professional knowledge work fields. And Soul did really great work fields. And Soul did really great work fields. And Soul did really great on that. But what Soul does not have, at on that. But what Soul does not have, at on that. But what Soul does not have, at least for me, is that big model smell.

  2. least for me, is that big model smell. least for me, is that big model smell. And that matches what OpenAI has been And that matches what OpenAI has been And that matches what OpenAI has been investing in. They've been investing investing in. They've been investing investing in. They've been investing very specifically in improved very specifically in improved very specifically in improved reinforcement learning for existing reinforcement learning for existing reinforcement learning for existing model lineages, which allows them to be model lineages, which allows them to be model lineages, which allows them to be more and more useful on specific tasks. more and more useful on specific tasks. more and more useful on specific tasks. In this case, it's very strong on In this case, it's very strong on In this case, it's very strong on knowledge work, it's strong on long-run knowledge work, it's strong on long-run knowledge work, it's strong on long-run agentic coding. What it does not have, agentic coding. What it does not have, agentic coding. What it does not have, at least for me, is the same big model at least for me, is the same big model at least for me, is the same big model smell that Fable 5 has. And that makes smell that Fable 5 has. And that makes smell that Fable 5 has. And that makes sense, because Anthropic has been sense, because Anthropic has been sense, because Anthropic has been investing in pre-train for their models. investing in pre-train for their models. investing in pre-train for their models. In other words, training on larger and In other words, training on larger and In other words, training on larger and larger data sets to enable more and more larger data sets to enable more and more larger data sets to enable more and more general purpose models. We have general purpose models. We have general purpose models. We have incredibly good assistance, assistance incredibly good assistance, assistance incredibly good assistance, assistance that make our best experts better that make our best experts better that make our best experts better because they act as companions to because they act as companions to because they act as companions to research, companions to thinking, but research, companions to thinking, but research, companions to thinking, but not yet truly generalizable intelligence not yet truly generalizable intelligence not yet truly generalizable intelligence with deep recursive learning. In the with deep recursive learning. In the with deep recursive learning. In the meantime, we have to figure out what to meantime, we have to figure out what to meantime, we have to figure out what to do with the models in front of us and do with the models in front of us and do with the models in front of us and make use of them today. And that brings make use of them today. And that brings make use of them today. And that brings me to the point that I'm trying to make. me to the point that I'm trying to make. me to the point that I'm trying to make. I am picking the model I am picking I am picking the model I am picking I am picking the model I am picking because it makes it easier for me to because it makes it easier for me to because it makes it easier for me to produce my best work. For what is my produce my best work. For what is my produce my best work. For what is my model recommendation? Is I say, are you model recommendation? Is I say, are you model recommendation? Is I say, are you me? Do you have the same habits with me? Do you have the same habits with me? Do you have the same habits with your model that I have? I am someone your model that I have? I am someone your model that I have? I am someone that is extremely willing to do very that is extremely willing to do very that is extremely willing to do very lengthy, somewhat technical prompts, and lengthy, somewhat technical prompts, and lengthy, somewhat technical prompts, and I'm okay verbalizing that. And so I just I'm okay verbalizing that. And so I just I'm okay verbalizing that. And so I just talk into Whisper flow, and I just feed talk into Whisper flow, and I just feed talk into Whisper flow, and I just feed it to the model, and I'm fairly specific it to the model, and I'm fairly specific it to the model, and I'm fairly specific about it. And that fits 5.6 fairly well, about it. And that fits 5.6 fairly well, about it. And that fits 5.6 fairly well, because 5.6 will read through that whole because 5.6 will read through that whole because 5.6 will read through that whole prompt, understand all of the edges I prompt, understand all of the edges I prompt, understand all of the edges I just talked about, and come back with a just talked about, and come back with a just talked about, and come back with a full piece of work, and be really

  3. full piece of work, and be really full piece of work, and be really persistent about getting it done. But persistent about getting it done. But persistent about getting it done. But not everyone talks and works that way, not everyone talks and works that way, not everyone talks and works that way, and your best work may actually come and your best work may actually come and your best work may actually come from a different approach to prompting from a different approach to prompting from a different approach to prompting and knowledge management. have different and knowledge management. have different and knowledge management. have different tasks that you're working on. You tasks that you're working on. You tasks that you're working on. You probably do. Hint for you. My best tip probably do. Hint for you. My best tip probably do. Hint for you. My best tip when people say, "What model do I pick?" when people say, "What model do I pick?" when people say, "What model do I pick?" is to not look at the model first. is to not look at the model first. is to not look at the model first. Instead, look at your best work and look Instead, look at your best work and look Instead, look at your best work and look at how you get there. And it may be not at how you get there. And it may be not at how you get there. And it may be not with the model. Look at the process you with the model. Look at the process you with the model. Look at the process you use for thinking. And then start to ask use for thinking. And then start to ask use for thinking. And then start to ask yourself which model helps me to yourself which model helps me to yourself which model helps me to accelerate that loop that gets me to my accelerate that loop that gets me to my accelerate that loop that gets me to my best self. I find, as I've been saying, best self. I find, as I've been saying, best self. I find, as I've been saying, those lengthy prompts, the ability to those lengthy prompts, the ability to those lengthy prompts, the ability to just talk about what I want done, plus just talk about what I want done, plus just talk about what I want done, plus the harness that allows me to the harness that allows me to the harness that allows me to self-improve really easily with Codex, self-improve really easily with Codex, self-improve really easily with Codex, that gets me really far. And by that gets me really far. And by that gets me really far. And by self-improve I mean that Codex will self-improve I mean that Codex will self-improve I mean that Codex will learn from what I do and further improve learn from what I do and further improve learn from what I do and further improve skills. I mean that Codex is steerable skills. I mean that Codex is steerable skills. I mean that Codex is steerable and I have fairly high intent with my and I have fairly high intent with my and I have fairly high intent with my prompt, so I like to steer it. Fable 5 prompt, so I like to steer it. Fable 5 prompt, so I like to steer it. Fable 5 is really good out of the box at is really good out of the box at is really good out of the box at understanding intent that's a little bit understanding intent that's a little bit understanding intent that's a little bit more high-level. It has that ability to more high-level. It has that ability to more high-level. It has that ability to generalize associated with big models. I generalize associated with big models. I generalize associated with big models. I love that. It's a fantastic model. The love that. It's a fantastic model. The love that. It's a fantastic model. The Anthropic team cooked, but I'm not Anthropic team cooked, but I'm not Anthropic team cooked, but I'm not reaching for it as much because it isn't reaching for it as much because it isn't reaching for it as much because it isn't suited to my particular work patterns.

  4. suited to my particular work patterns. suited to my particular work patterns. And if your work patterns are more And if your work patterns are more And if your work patterns are more around understanding very high-level around understanding very high-level around understanding very high-level ambiguity, wrestling with concepts, ambiguity, wrestling with concepts, ambiguity, wrestling with concepts, trying to pin down ideas between ideas, trying to pin down ideas between ideas, trying to pin down ideas between ideas, then Fable may be a much better model then Fable may be a much better model then Fable may be a much better model for you. If you're more suited to for you. If you're more suited to for you. If you're more suited to understanding how to get coding done understanding how to get coding done understanding how to get coding done efficiently and quickly, honestly, you efficiently and quickly, honestly, you efficiently and quickly, honestly, you may reach for the Luna series from may reach for the Luna series from may reach for the Luna series from OpenAI, which is much cheaper to run on, OpenAI, which is much cheaper to run on, OpenAI, which is much cheaper to run on, incredibly high-powered, and they incredibly high-powered, and they incredibly high-powered, and they released it along with 5.6 Soul and 5.6 released it along with 5.6 Soul and 5.6 released it along with 5.6 Soul and 5.6 Terra. Or you may reach for Grok, which Terra. Or you may reach for Grok, which Terra. Or you may reach for Grok, which also has good frontier-ish coding also has good frontier-ish coding also has good frontier-ish coding capabilities. You may reach for GLM 5.2. capabilities. You may reach for GLM 5.2. capabilities. You may reach for GLM 5.2. You may reach for Ringer, which I built You may reach for Ringer, which I built You may reach for Ringer, which I built and talked about last week because it and talked about last week because it and talked about last week because it enables you to farm out and orchestrate enables you to farm out and orchestrate enables you to farm out and orchestrate from one central model like Fable to a from one central model like Fable to a from one central model like Fable to a bunch of cheaper models. It suits your bunch of cheaper models. It suits your bunch of cheaper models. It suits your work, right? And by the way, if you're work, right? And by the way, if you're work, right? And by the way, if you're wondering, would I still use Fable as wondering, would I still use Fable as wondering, would I still use Fable as the architect in Ringer even with 5.6 the architect in Ringer even with 5.6 the architect in Ringer even with 5.6 out? I would because Fable is good at out? I would because Fable is good at out? I would because Fable is good at understanding intent and breaking down understanding intent and breaking down understanding intent and breaking down those tasks to get that intent done. It those tasks to get that intent done. It those tasks to get that intent done. It also has a good front-end instinct.

  5. also has a good front-end instinct. also has a good front-end instinct. They're just Anthropic has been They're just Anthropic has been They're just Anthropic has been consistently good at front-end. I'm consistently good at front-end. I'm consistently good at front-end. I'm going to flash the benchmarks up here on going to flash the benchmarks up here on going to flash the benchmarks up here on the screen as I talk so you can see how the screen as I talk so you can see how the screen as I talk so you can see how I scored 5.6. I'm using the same I scored 5.6. I'm using the same I scored 5.6. I'm using the same benchmarks I've used for all of the benchmarks I've used for all of the benchmarks I've used for all of the models over the last few generations, so models over the last few generations, so models over the last few generations, so we're not changing anything. But as I do we're not changing anything. But as I do we're not changing anything. But as I do that, I want you to think about the that, I want you to think about the that, I want you to think about the larger point we've been talking about, larger point we've been talking about, larger point we've been talking about, and I want to suggest to you that given and I want to suggest to you that given and I want to suggest to you that given everything I've shared with you, we are everything I've shared with you, we are everything I've shared with you, we are missing a core insight. We are missing missing a core insight. We are missing missing a core insight. We are missing the idea that models are becoming more the idea that models are becoming more the idea that models are becoming more like like like families we need to get to know and less families we need to get to know and less families we need to get to know and less benchmarkable period. I don't believe my benchmarkable period. I don't believe my benchmarkable period. I don't believe my benchmark or any benchmark fully benchmark or any benchmark fully benchmark or any benchmark fully captures what these models do in a way captures what these models do in a way captures what these models do in a way that's useful. And that's why I make that's useful. And that's why I make that's useful. And that's why I make videos that may feel vibe-ish like this videos that may feel vibe-ish like this videos that may feel vibe-ish like this where I tell you what it actually feels where I tell you what it actually feels where I tell you what it actually feels like to use these models because I want like to use these models because I want like to use these models because I want you to get inspired to jump in and test you to get inspired to jump in and test you to get inspired to jump in and test them for yourself on your workflows and them for yourself on your workflows and them for yourself on your workflows and also to hear from someone who does that also to hear from someone who does that also to hear from someone who does that all the time. The key insight I have for all the time. The key insight I have for all the time. The key insight I have for you as someone who has touched these you as someone who has touched these you as someone who has touched these models and lived with these models is models and lived with these models is models and lived with these models is that the models really do need to be that the models really do need to be that the models really do need to be treated like family. Think of it as treated like family. Think of it as treated like family. Think of it as every new model's like a new picture in every new model's like a new picture in every new model's like a new picture in a family photo album. You're getting to a family photo album. You're getting to a family photo album. You're getting to know someone new who is a part of the know someone new who is a part of the know someone new who is a part of the family. There's family resemblance. I family. There's family resemblance. I family. There's family resemblance. I can tell you the 5.x family from OpenAI can tell you the 5.x family from OpenAI can tell you the 5.x family from OpenAI has family resemblance. They all have has family resemblance. They all have has family resemblance. They all have that preference for long-running agentic that preference for long-running agentic that preference for long-running agentic coding flows. They all have the ability

  6. coding flows. They all have the ability coding flows. They all have the ability to understand what you're saying to understand what you're saying to understand what you're saying explicitly, paint the edges very explicitly, paint the edges very explicitly, paint the edges very clearly, and just go after it. And they clearly, and just go after it. And they clearly, and just go after it. And they maybe have less of an ability to read maybe have less of an ability to read maybe have less of an ability to read between the lines. Whereas the Mythos between the lines. Whereas the Mythos between the lines. Whereas the Mythos lineage, we have Mythos and Fable there, lineage, we have Mythos and Fable there, lineage, we have Mythos and Fable there, is extremely good at ambiguous tasks, is extremely good at ambiguous tasks, is extremely good at ambiguous tasks, has extraordinary front-end taste, is has extraordinary front-end taste, is has extraordinary front-end taste, is almost philosophical in the way it almost philosophical in the way it almost philosophical in the way it approaches problems, is a deep thinker. approaches problems, is a deep thinker. approaches problems, is a deep thinker. Um and those are just fundamentally Um and those are just fundamentally Um and those are just fundamentally different approaches. It's not that one different approaches. It's not that one different approaches. It's not that one is really better or worse. And this is is really better or worse. And this is is really better or worse. And this is where I think when we say words like where I think when we say words like where I think when we say words like dumber or smarter, we say them in the dumber or smarter, we say them in the dumber or smarter, we say them in the context of benchmarks, but I don't think context of benchmarks, but I don't think context of benchmarks, but I don't think it does the models a service because the it does the models a service because the it does the models a service because the models are becoming different the way models are becoming different the way models are becoming different the way families are different. And we don't families are different. And we don't families are different. And we don't really say this family's dumb and this really say this family's dumb and this really say this family's dumb and this family's smart. We say these families family's smart. We say these families family's smart. We say these families are different. And I think that's more are different. And I think that's more are different. And I think that's more useful. And so we have the Anthropic useful. And so we have the Anthropic useful. And so we have the Anthropic family of models. It's more pre-trained. family of models. It's more pre-trained. family of models. It's more pre-trained. It's more front-end-y. It's more front-end-y. It's more front-end-y. It's more interested in character, in It's more interested in character, in It's more interested in character, in philosophy. In fact, Anthropic released philosophy. In fact, Anthropic released philosophy. In fact, Anthropic released a whole study that it did on how a whole study that it did on how a whole study that it did on how Anthropic's models think called J space, Anthropic's models think called J space, Anthropic's models think called J space, where there's this idea that these where there's this idea that these where there's this idea that these models are able to computationally models are able to computationally models are able to computationally manipulate higher-order concepts while manipulate higher-order concepts while manipulate higher-order concepts while doing autonomous processing on doing autonomous processing on doing autonomous processing on lower-order token prediction. And that lower-order token prediction. And that lower-order token prediction. And that that self-evolved in these models.

  7. that self-evolved in these models. that self-evolved in these models. Anthropic has done phenomenal work Anthropic has done phenomenal work Anthropic has done phenomenal work essentially connecting technology and essentially connecting technology and essentially connecting technology and philosophy to understand how models work philosophy to understand how models work philosophy to understand how models work at a deep level. That doesn't mean that at a deep level. That doesn't mean that at a deep level. That doesn't mean that their models always produce the smartest their models always produce the smartest their models always produce the smartest possible work. It just means that that's possible work. It just means that that's possible work. It just means that that's part of that model character and part of part of that model character and part of part of that model character and part of that model lineage. OpenAI has done that model lineage. OpenAI has done that model lineage. OpenAI has done phenomenal work on the Codex harness. phenomenal work on the Codex harness. phenomenal work on the Codex harness. And that makes it very, very easy to And that makes it very, very easy to And that makes it very, very easy to work with Codex and understand work with Codex and understand work with Codex and understand how you are going to further improve how you are going to further improve how you are going to further improve your work over time. That's where I talk your work over time. That's where I talk your work over time. That's where I talk about those self-improving loops, about those self-improving loops, about those self-improving loops, telling Codex to check what you've done telling Codex to check what you've done telling Codex to check what you've done and get better at it. And by the way, I and get better at it. And by the way, I and get better at it. And by the way, I am not leaving out ChatGPT work. I know am not leaving out ChatGPT work. I know am not leaving out ChatGPT work. I know that they launched ChatGPT work with that they launched ChatGPT work with that they launched ChatGPT work with 5.6. That's a very, very exciting 5.6. That's a very, very exciting 5.6. That's a very, very exciting development. I think that one of the development. I think that one of the development. I think that one of the things that I would be looking for with things that I would be looking for with things that I would be looking for with work work work is that we have more non-tech input into is that we have more non-tech input into is that we have more non-tech input into how these tools evolve. Right now, to be how these tools evolve. Right now, to be how these tools evolve. Right now, to be honest, we see the impact of engineering honest, we see the impact of engineering honest, we see the impact of engineering culture on how these tools are evolving.

  8. culture on how these tools are evolving. culture on how these tools are evolving. Claude code and Codex are both built by Claude code and Codex are both built by Claude code and Codex are both built by engineers for engineers, and it shows. engineers for engineers, and it shows. engineers for engineers, and it shows. They're ergonomically comfortable for They're ergonomically comfortable for They're ergonomically comfortable for engineers. But co-work and work from engineers. But co-work and work from engineers. But co-work and work from Anthropic and OpenAI respectively are Anthropic and OpenAI respectively are Anthropic and OpenAI respectively are not primarily built by non-engineers for not primarily built by non-engineers for not primarily built by non-engineers for non-engineers. And unfortunately, that non-engineers. And unfortunately, that non-engineers. And unfortunately, that means that sometimes what you get is an means that sometimes what you get is an means that sometimes what you get is an engineer's perception of what engineer's perception of what engineer's perception of what non-engineers want. And that can look non-engineers want. And that can look non-engineers want. And that can look like we need to dumb things down because like we need to dumb things down because like we need to dumb things down because these non-engineers are not as these non-engineers are not as these non-engineers are not as technical. And I get a little bit of technical. And I get a little bit of technical. And I get a little bit of that flavor with ChatGPT work, and I that flavor with ChatGPT work, and I that flavor with ChatGPT work, and I would like to see a more sophisticated would like to see a more sophisticated would like to see a more sophisticated approach because knowledge work is approach because knowledge work is approach because knowledge work is really different if you're not coding. really different if you're not coding. really different if you're not coding. Knowledge work is more about process. Knowledge work is more about process. Knowledge work is more about process. It's more about coming to a conclusion It's more about coming to a conclusion It's more about coming to a conclusion over time and thinking about something, over time and thinking about something, over time and thinking about something, and it's less about code and and it's less about code and and it's less about code and verification than engineering work is. verification than engineering work is. verification than engineering work is. And we need tools that enable AI to do And we need tools that enable AI to do And we need tools that enable AI to do that with us if we're knowledge workers. that with us if we're knowledge workers. that with us if we're knowledge workers. And we really haven't had extraordinary And we really haven't had extraordinary And we really haven't had extraordinary harnesses for that in a way that we've harnesses for that in a way that we've harnesses for that in a way that we've had for engineers with code. And so, I had for engineers with code. And so, I had for engineers with code. And so, I think ChatGPT work is a first stab at think ChatGPT work is a first stab at think ChatGPT work is a first stab at where ChatGPT and the Codex family are where ChatGPT and the Codex family are where ChatGPT and the Codex family are going. They want to get into knowledge going. They want to get into knowledge going. They want to get into knowledge work as well, definitely competing with work as well, definitely competing with work as well, definitely competing with Anthropic's co-work. I don't think it's Anthropic's co-work. I don't think it's Anthropic's co-work. I don't think it's the be-all end-all. In fact, I am still the be-all end-all. In fact, I am still the be-all end-all. In fact, I am still using Codex because I feel comfortable using Codex because I feel comfortable using Codex because I feel comfortable with it. I don't mind having the full with it. I don't mind having the full with it. I don't mind having the full range of tools. I don't want to be range of tools. I don't want to be range of tools. I don't want to be constrained. I am looking for the same constrained. I am looking for the same constrained. I am looking for the same degree of care and precision with degree of care and precision with degree of care and precision with knowledge work that I've seen with

  9. knowledge work that I've seen with knowledge work that I've seen with coding work from these model makers. And coding work from these model makers. And coding work from these model makers. And we haven't seen it yet, to be really we haven't seen it yet, to be really we haven't seen it yet, to be really honest with you. And there's an honest with you. And there's an honest with you. And there's an opportunity on the table either for a opportunity on the table either for a opportunity on the table either for a startup to go grab that or for somebody startup to go grab that or for somebody startup to go grab that or for somebody else to come in and say, "This is what else to come in and say, "This is what else to come in and say, "This is what knowledge work looks like when it's not knowledge work looks like when it's not knowledge work looks like when it's not obsessed with how code passes in a repo. obsessed with how code passes in a repo. obsessed with how code passes in a repo. I am building a tool for you that will I am building a tool for you that will I am building a tool for you that will help you to lay out, to talk, to ramble, help you to lay out, to talk, to ramble, help you to lay out, to talk, to ramble, to share what you do, what you're good to share what you do, what you're good to share what you do, what you're good at, what you're passionate about." If at, what you're passionate about." If at, what you're passionate about." If that's something you're interested in, that's something you're interested in, that's something you're interested in, absolutely come and grab it. The link is absolutely come and grab it. The link is absolutely come and grab it. The link is down below. I I want to make it easy. I down below. I I want to make it easy. I down below. I I want to make it easy. I think I've been wrestling with this idea think I've been wrestling with this idea think I've been wrestling with this idea that we traditionally have had like that we traditionally have had like that we traditionally have had like these one-to-one maps for model pickers. these one-to-one maps for model pickers. these one-to-one maps for model pickers. We've had We've had We've had choose your own adventure model pickers choose your own adventure model pickers choose your own adventure model pickers that are based on very brief quizzes. that are based on very brief quizzes. that are based on very brief quizzes. I've tried those in the past. They're I've tried those in the past. They're I've tried those in the past. They're not super useful. We have more not super useful. We have more not super useful. We have more computational power at our disposal with computational power at our disposal with computational power at our disposal with intelligence now. I wanted to use that intelligence now. I wanted to use that intelligence now. I wanted to use that to make choosing your own model mix to make choosing your own model mix to make choosing your own model mix easier over time. And so, this tool is easier over time. And so, this tool is easier over time. And so, this tool is going to keep being updated as I going to keep being updated as I going to keep being updated as I continue to benchmark more models. So, continue to benchmark more models. So, continue to benchmark more models. So, you'll have more and more models you'll have more and more models you'll have more and more models available over time including open available over time including open available over time including open source models, including the Anthropic source models, including the Anthropic source models, including the Anthropic family, the Open AI family, Meta as family, the Open AI family, Meta as family, the Open AI family, Meta as relevant, Google as relevant, Grok, etc.

  10. relevant, Google as relevant, Grok, etc. relevant, Google as relevant, Grok, etc. What I want to do is I want to have a What I want to do is I want to have a What I want to do is I want to have a much more nuanced conversational much more nuanced conversational much more nuanced conversational evolving framework for how we pick evolving framework for how we pick evolving framework for how we pick models so that we can do our best work models so that we can do our best work models so that we can do our best work with the model best suited to us. If with the model best suited to us. If with the model best suited to us. If you're still stuck or if you're like, you're still stuck or if you're like, you're still stuck or if you're like, "No, no, no, Nate, just give me the "No, no, no, Nate, just give me the "No, no, no, Nate, just give me the answer. I just want the answer." The answer. I just want the answer." The answer. I just want the answer." The answer for you is to pick the model that answer for you is to pick the model that answer for you is to pick the model that makes you feel most comfortable doing makes you feel most comfortable doing makes you feel most comfortable doing your hardest work. Because if you're if your hardest work. Because if you're if your hardest work. Because if you're if you're pushing on the model, if you're you're pushing on the model, if you're you're pushing on the model, if you're doing your hardest work and the model doing your hardest work and the model doing your hardest work and the model helps you get that done, that's the helps you get that done, that's the helps you get that done, that's the model you're going to want to lean on. model you're going to want to lean on. model you're going to want to lean on. So, when in doubt, go with the model So, when in doubt, go with the model So, when in doubt, go with the model that picks your hardest work. And by the that picks your hardest work. And by the that picks your hardest work. And by the way, for those of you that want to dig way, for those of you that want to dig way, for those of you that want to dig in and grab all of the details on 5.6, in and grab all of the details on 5.6, in and grab all of the details on 5.6, how it compares to Fable 5, grab the how it compares to Fable 5, grab the how it compares to Fable 5, grab the full test results on Grok 4.5 as well, full test results on Grok 4.5 as well, full test results on Grok 4.5 as well, and understand more deeply how all of and understand more deeply how all of and understand more deeply how all of this fits together into the new model this fits together into the new model this fits together into the new model race. I have a deeper article on race. I have a deeper article on race. I have a deeper article on Substack that really dives into those Substack that really dives into those Substack that really dives into those dynamics. And of course, you can jump in dynamics. And of course, you can jump in dynamics. And of course, you can jump in and grab the tool as well.

  11. and grab the tool as well. and grab the tool as well. All right. I will see you next time. The All right. I will see you next time. The All right. I will see you next time. The model race is going to continue to get model race is going to continue to get model race is going to continue to get more complicated, but I think that our more complicated, but I think that our more complicated, but I think that our ability to understand what we're doing ability to understand what we're doing ability to understand what we're doing can stay really consistent and that can can stay really consistent and that can can stay really consistent and that can help us stay sane. help us stay sane. help us stay sane. I'll talk to you next time.

Summary

The session discusses selecting the appropriate AI model, highlighting that less sophisticated models like Chat GPT 5.6 Soul can be highly effective for specific tasks, outperforming even larger models on benchmarks like Agent Slash exam and the author's "Dingo" knowledge work test. The key takeaway is that users should focus on understanding their own needs rather than relying solely on general benchmarks or the perceived "size" of a model, emphasizing tools for personal model selection.

View original episode ↗