What Is Fable 5?
Read full transcript 3 segments
-
We've been hearing about Mythos forever, We've been hearing about Mythos forever, and to finally have it in our hands is and to finally have it in our hands is and to finally have it in our hands is unbelievable. It's so unbelievable that unbelievable. It's so unbelievable that unbelievable. It's so unbelievable that we don't actually have it in our hands we don't actually have it in our hands we don't actually have it in our hands because Mythos 5 is not the model that because Mythos 5 is not the model that because Mythos 5 is not the model that we all have access to now. What we did we all have access to now. What we did we all have access to now. What we did end up getting is a new model called end up getting is a new model called end up getting is a new model called Fable 5, which is still Mythos, but it Fable 5, which is still Mythos, but it Fable 5, which is still Mythos, but it has a bunch of safeguards in front. I has a bunch of safeguards in front. I has a bunch of safeguards in front. I have a lot of layers to go into here have a lot of layers to go into here have a lot of layers to go into here from how it feels to build with them to from how it feels to build with them to from how it feels to build with them to all the benchmarks that we got to a all the benchmarks that we got to a all the benchmarks that we got to a handful of benchmarks that aren't even handful of benchmarks that aren't even handful of benchmarks that aren't even public yet that I was able to get early public yet that I was able to get early public yet that I was able to get early access data for so I can show you guys access data for so I can show you guys access data for so I can show you guys first. So, you should be excited. I know first. So, you should be excited. I know first. So, you should be excited. I know I am. I'll do my best to not waste time I am. I'll do my best to not waste time I am. I'll do my best to not waste time and just cover the facts. Notice that and just cover the facts. Notice that and just cover the facts. Notice that they say Mythos 5 / Fable 5 here. The they say Mythos 5 / Fable 5 here. The they say Mythos 5 / Fable 5 here. The reason for that is because a lot of reason for that is because a lot of reason for that is because a lot of these benches had questions that Fable 5 these benches had questions that Fable 5 these benches had questions that Fable 5 outright refused to answer, which would outright refused to answer, which would outright refused to answer, which would plummet the scores. Even something plummet the scores. Even something plummet the scores. Even something simple like Terminal Bench saw a simple like Terminal Bench saw a simple like Terminal Bench saw a 20-point drop compared to when the same 20-point drop compared to when the same 20-point drop compared to when the same run was done with Mythos. But when you run was done with Mythos. But when you run was done with Mythos. But when you take a look at what its capabilities are take a look at what its capabilities are take a look at what its capabilities are as in what it's able to do unrestricted, as in what it's able to do unrestricted, as in what it's able to do unrestricted, it crushes. I burned through the 5-hour it crushes. I burned through the 5-hour it crushes. I burned through the 5-hour session limits on two $200 accounts at session limits on two $200 accounts at session limits on two $200 accounts at the same time, which I never thought I the same time, which I never thought I the same time, which I never thought I would be able to do. SW Bench Pro got an would be able to do. SW Bench Pro got an would be able to do. SW Bench Pro got an 80% on, which GPT-4 6 only got a 58.6.
-
80% on, which GPT-4 6 only got a 58.6. 80% on, which GPT-4 6 only got a 58.6. Remember though, we don't like that Remember though, we don't like that Remember though, we don't like that bench. I did a detailed breakdown of why bench. I did a detailed breakdown of why bench. I did a detailed breakdown of why SW Bench Pro is kind of garbage now SW Bench Pro is kind of garbage now SW Bench Pro is kind of garbage now because most of the bench is existing because most of the bench is existing because most of the bench is existing pull requests and handing the model the pull requests and handing the model the pull requests and handing the model the description of this PR that merged and description of this PR that merged and description of this PR that merged and an old commit hash, and it just seems an old commit hash, and it just seems an old commit hash, and it just seems like Mythos 5 has so much data in its like Mythos 5 has so much data in its like Mythos 5 has so much data in its training that it has the details to training that it has the details to training that it has the details to recreate those PR's. The Frontier Code recreate those PR's. The Frontier Code recreate those PR's. The Frontier Code Bench however is much more interesting Bench however is much more interesting Bench however is much more interesting and we'll dive into that more in the and we'll dive into that more in the and we'll dive into that more in the future. For now though, it got a 30% future. For now though, it got a 30% future. For now though, it got a 30% when Opus 4 8 only got a 13 and 5 5 only when Opus 4 8 only got a 13 and 5 5 only when Opus 4 8 only got a 13 and 5 5 only got a 5.7. Its vision capabilities are a got a 5.7. Its vision capabilities are a got a 5.7. Its vision capabilities are a meaningful upgrade from the Opus line. meaningful upgrade from the Opus line. meaningful upgrade from the Opus line. It's It's really good at spatial It's It's really good at spatial It's It's really good at spatial reasoning, the first time they've taken reasoning, the first time they've taken reasoning, the first time they've taken this lead from OpenAI. And as I this lead from OpenAI. And as I this lead from OpenAI. And as I mentioned before, scored well on mentioned before, scored well on mentioned before, scored well on Terminal Bench, but the model we get Terminal Bench, but the model we get Terminal Bench, but the model we get scored much more poorly because it scored much more poorly because it scored much more poorly because it blocked a handful of the requests in the blocked a handful of the requests in the blocked a handful of the requests in the Terminal Bench tests. I'm lucky to have Terminal Bench tests. I'm lucky to have Terminal Bench tests. I'm lucky to have Data Curve sharing the the SW numbers Data Curve sharing the the SW numbers Data Curve sharing the the SW numbers with me early. If you haven't already with me early. If you haven't already with me early. If you haven't already watched the deep S T W E video, highly watched the deep S T W E video, highly watched the deep S T W E video, highly recommend it. Most of these code recommend it. Most of these code recommend it. Most of these code benchmarks suck. This is one that I benchmarks suck. This is one that I benchmarks suck. This is one that I think makes actual sense. And what it think makes actual sense. And what it think makes actual sense. And what it shows here is that Fable on X high shows here is that Fable on X high shows here is that Fable on X high performs comparably to GPT 5.5. And more performs comparably to GPT 5.5. And more performs comparably to GPT 5.5. And more importantly, they crushed all of Opus's importantly, they crushed all of Opus's importantly, they crushed all of Opus's scores with fewer dollars spent. So, scores with fewer dollars spent. So, scores with fewer dollars spent. So, even though this model is much more even though this model is much more even though this model is much more expensive, it ends up being a better expensive, it ends up being a better expensive, it ends up being a better deal overall compared to most Opus deal overall compared to most Opus deal overall compared to most Opus models because it uses way fewer tokens.
-
models because it uses way fewer tokens. models because it uses way fewer tokens. This is a great change and I'm thankful This is a great change and I'm thankful This is a great change and I'm thankful to see Anthropic token maxing a little to see Anthropic token maxing a little to see Anthropic token maxing a little less hard at the very least with the less hard at the very least with the less hard at the very least with the capabilities of this model.
Summary
The main theme is the release and initial performance analysis of Fable 5, a new iteration of the Mythos AI model. Key subjects include benchmarks like Terminal Bench and SW Bench Pro, with comparisons to GPT-4, and discussions on safeguards and unrestricted capabilities. The takeaway is that while Fable 5 has impressive raw performance and spatial reasoning, its safeguards limit it in some benchmarks, suggesting a need for practical application testing beyond standardized tests.