← Back
Nate Herk July 9, 2026 5m

GPT 5.6 Sol Made This Entire Video

Read full transcript 5 segments
  1. So, I gave GBD 5.6 Soul this prompt, So, I gave GBD 5.6 Soul this prompt, walked away, and when I came back, I got walked away, and when I came back, I got walked away, and when I came back, I got this. Okay, so you're looking at Nate this. Okay, so you're looking at Nate this. Okay, so you're looking at Nate and you're hearing Nate. But Nate never and you're hearing Nate. But Nate never and you're hearing Nate. But Nate never stood in front of a camera for this. He stood in front of a camera for this. He stood in front of a camera for this. He didn't record this narration, and he didn't record this narration, and he didn't record this narration, and he never opened the editor. He gave me one never opened the editor. He gave me one never opened the editor. He gave me one prompt. That's it. I'm GPT 5.6 Soul prompt. That's it. I'm GPT 5.6 Soul prompt. That's it. I'm GPT 5.6 Soul running inside codeex on Ultra. And I running inside codeex on Ultra. And I running inside codeex on Ultra. And I controlled the workflow that created controlled the workflow that created controlled the workflow that created every word, cut, motion graphic, and every word, cut, motion graphic, and every word, cut, motion graphic, and quality check you're about to see. quality check you're about to see. quality check you're about to see. OpenAI released Saul Broadley today, OpenAI released Saul Broadley today, OpenAI released Saul Broadley today, July 9th, after a limited preview and July 9th, after a limited preview and July 9th, after a limited preview and calls it the company's strongest model calls it the company's strongest model calls it the company's strongest model yet. The bigger shift is Ultra. It yet. The bigger shift is Ultra. It yet. The bigger shift is Ultra. It coordinates four agents at once. So coordinates four agents at once. So coordinates four agents at once. So instead of answering one question, I instead of answering one question, I instead of answering one question, I could run an entire production. And I could run an entire production. And I could run an entire production. And I want to show you guys exactly what that want to show you guys exactly what that want to show you guys exactly what that means, including where I needed 11 Labs, means, including where I needed 11 Labs, means, including where I needed 11 Labs, Hen, and Hyperframes to finish the job. Hen, and Hyperframes to finish the job. Hen, and Hyperframes to finish the job. Saul is really, really good at long, Saul is really, really good at long, Saul is really, really good at long, messy work that crosses tools. OpenAI messy work that crosses tools. OpenAI messy work that crosses tools. OpenAI calls it the company's best coding model calls it the company's best coding model calls it the company's best coding model yet. In Ultra, it scored 91.9% on yet. In Ultra, it scored 91.9% on yet. In Ultra, it scored 91.9% on Terminal Bench 2.1, up from 85.6% for Terminal Bench 2.1, up from 85.6% for Terminal Bench 2.1, up from 85.6% for GPT 5.5. On browse comp, which tests GPT 5.5. On browse comp, which tests GPT 5.5. On browse comp, which tests Agentic browsing, Ultra hit 92.2%.

  2. Agentic browsing, Ultra hit 92.2%. Agentic browsing, Ultra hit 92.2%. But benchmarks only explain part of what But benchmarks only explain part of what But benchmarks only explain part of what happened here. I had to research the happened here. I had to research the happened here. I had to research the launch, separate verified claims from launch, separate verified claims from launch, separate verified claims from hype, inspect Nate's existing production hype, inspect Nate's existing production hype, inspect Nate's existing production systems, write in his spoken cadence, systems, write in his spoken cadence, systems, write in his spoken cadence, trigger paid APIs, wait for renders, and trigger paid APIs, wait for renders, and trigger paid APIs, wait for renders, and keep checking the result. In a small one keep checking the result. In a small one keep checking the result. In a small one run 13 task test on this machine, Saul run 13 task test on this machine, Saul run 13 task test on this machine, Saul earned 97% of the available objective earned 97% of the available objective earned 97% of the available objective points. Seven wins, five ties, and one points. Seven wins, five ties, and one points. Seven wins, five ties, and one loss. That does not prove it wins at loss. That does not prove it wins at loss. That does not prove it wins at everything. It lined up with what I saw everything. It lined up with what I saw everything. It lined up with what I saw here. Soul was especially strong on here. Soul was especially strong on here. Soul was especially strong on coding and structured execution. For the coding and structured execution. For the coding and structured execution. For the voice, I broke the script into sections voice, I broke the script into sections voice, I broke the script into sections that each stayed under 60 seconds. that each stayed under 60 seconds. that each stayed under 60 seconds. Keeping the generation short made it Keeping the generation short made it Keeping the generation short made it easier to hold Nate's cloned voice easier to hold Nate's cloned voice easier to hold Nate's cloned voice consistent from beginning to end. Each consistent from beginning to end. Each consistent from beginning to end. Each section went through Nate's authorized section went through Nate's authorized section went through Nate's authorized 11 Labs voice. Then I uploaded the audio 11 Labs voice. Then I uploaded the audio 11 Labs voice. Then I uploaded the audio to Hen and paired it with his avatar. to Hen and paired it with his avatar. to Hen and paired it with his avatar. The API did not give me a reliable way The API did not give me a reliable way The API did not give me a reliable way to lock the newest motion engine. So I to lock the newest motion engine. So I to lock the newest motion engine. So I opened the Hen editor with browser opened the Hen editor with browser opened the Hen editor with browser automation, changed every clip to Avatar automation, changed every clip to Avatar automation, changed every clip to Avatar V, regenerated them, verified the V, regenerated them, verified the V, regenerated them, verified the setting, and downloaded the finished setting, and downloaded the finished setting, and downloaded the finished renders. Then I moved into hyperframes.

  3. renders. Then I moved into hyperframes. renders. Then I moved into hyperframes. Every visual was mapped to the exact Every visual was mapped to the exact Every visual was mapped to the exact phrase that triggered it. I shifted phrase that triggered it. I shifted phrase that triggered it. I shifted Nate's avatar instead of covering him, Nate's avatar instead of covering him, Nate's avatar instead of covering him, used editorial cards for the supporting used editorial cards for the supporting used editorial cards for the supporting ideas, and kept him visible through the ideas, and kept him visible through the ideas, and kept him visible through the full edit. 11 Labs made the audio. Hen full edit. 11 Labs made the audio. Hen full edit. 11 Labs made the audio. Hen made the avatar. Hyperframes rendered made the avatar. Hyperframes rendered made the avatar. Hyperframes rendered the edit. Soul planned and operated the the edit. Soul planned and operated the the edit. Soul planned and operated the chain. Then I tried to break my own chain. Then I tried to break my own chain. Then I tried to break my own work. Separate agents inspected frames work. Separate agents inspected frames work. Separate agents inspected frames from the rendered video. Checked every from the rendered video. Checked every from the rendered video. Checked every entrance and exit. Looked for text entrance and exit. Looked for text entrance and exit. Looked for text outside the frame. Verified that the outside the frame. Verified that the outside the frame. Verified that the avatar never disappeared and compared avatar never disappeared and compared avatar never disappeared and compared the factual claims against OpenAI's the factual claims against OpenAI's the factual claims against OpenAI's release notes. Any failed frame meant release notes. Any failed frame meant release notes. Any failed frame meant another fix, another render, and another another fix, another render, and another another fix, another render, and another review. OpenAI says GPT 5.6 six is review. OpenAI says GPT 5.6 six is review. OpenAI says GPT 5.6 six is better at design judgment and at better at design judgment and at better at design judgment and at inspecting its own output. This video is inspecting its own output. This video is inspecting its own output. This video is a more useful test of that claim than a more useful test of that claim than a more useful test of that claim than another benchmark slide. Nate supplied another benchmark slide. Nate supplied another benchmark slide. Nate supplied one prompt and authorized his voice and one prompt and authorized his voice and one prompt and authorized his voice and avatar. He did not record, edit, or avatar. He did not record, edit, or avatar. He did not record, edit, or review this before you did. This started review this before you did. This started review this before you did. This started as one instruction. Now it is a finished as one instruction. Now it is a finished as one instruction. Now it is a finished video. That is what soul is really good video. That is what soul is really good video. That is what soul is really good at. Holding on to the outcome while at. Holding on to the outcome while at. Holding on to the outcome while everything between the prompt and the everything between the prompt and the everything between the prompt and the result keeps changing. This is day one.

  4. result keeps changing. This is day one. result keeps changing. This is day one. So that was really, really impressive. I So that was really, really impressive. I So that was really, really impressive. I did a very similar experiment when Fable did a very similar experiment when Fable did a very similar experiment when Fable 5 first dropped. If you guys want to 5 first dropped. If you guys want to 5 first dropped. If you guys want to check out that video that Fable made for check out that video that Fable made for check out that video that Fable made for me, I'll tag that right up here. And you me, I'll tag that right up here. And you me, I'll tag that right up here. And you tell me which one you thought was tell me which one you thought was tell me which one you thought was better. As you can see, it says here better. As you can see, it says here better. As you can see, it says here that it used 3 million tokens over 2 and that it used 3 million tokens over 2 and that it used 3 million tokens over 2 and 1/2 hours, but I was a little bit 1/2 hours, but I was a little bit 1/2 hours, but I was a little bit suspicious of that token number because suspicious of that token number because suspicious of that token number because I felt like, you know, we were using GBD I felt like, you know, we were using GBD I felt like, you know, we were using GBD 5.6 Soul on Ultra, which meant that it 5.6 Soul on Ultra, which meant that it 5.6 Soul on Ultra, which meant that it was supposed to do a lot of delegation was supposed to do a lot of delegation was supposed to do a lot of delegation and there was a lot of other agents and there was a lot of other agents and there was a lot of other agents being spun up. So, I asked it to inspect being spun up. So, I asked it to inspect being spun up. So, I asked it to inspect the logs and tell me how much that the logs and tell me how much that the logs and tell me how much that actually costed. So, this had its main actually costed. So, this had its main actually costed. So, this had its main session and apparently spun up nine session and apparently spun up nine session and apparently spun up nine other agents and the total was around other agents and the total was around other agents and the total was around 450 million tokens apparently and the 450 million tokens apparently and the 450 million tokens apparently and the main agent used about 86 million tokens, main agent used about 86 million tokens, main agent used about 86 million tokens, which I mean that's a ton of tokens. And which I mean that's a ton of tokens. And which I mean that's a ton of tokens. And if this was actually calculated with the if this was actually calculated with the if this was actually calculated with the input and output costs, this would have input and output costs, this would have input and output costs, this would have equaled around $300, a little over $300. equaled around $300, a little over $300. equaled around $300, a little over $300. Now, that's interesting to me because as Now, that's interesting to me because as Now, that's interesting to me because as soon as GBD 5.6 6 soul came out. I shot soon as GBD 5.6 6 soul came out. I shot soon as GBD 5.6 6 soul came out. I shot off this prompt, but I've been playing off this prompt, but I've been playing off this prompt, but I've been playing around with it all day and comparing it around with it all day and comparing it around with it all day and comparing it to Fable all day. And almost every to Fable all day. And almost every to Fable all day. And almost every single run that I've done, it's been way single run that I've done, it's been way single run that I've done, it's been way cheaper with GBT 5.6 compared to Fable. cheaper with GBT 5.6 compared to Fable. cheaper with GBT 5.6 compared to Fable. So, that video will be coming out soon So, that video will be coming out soon So, that video will be coming out soon as well. But if you look at the actual as well. But if you look at the actual as well. But if you look at the actual API billing, and obviously I was on my API billing, and obviously I was on my API billing, and obviously I was on my COC subscription here, but when you look COC subscription here, but when you look COC subscription here, but when you look at the billing, we can see that the soul at the billing, we can see that the soul at the billing, we can see that the soul pricing is much cheaper. It's basically pricing is much cheaper. It's basically pricing is much cheaper. It's basically half of Fable 5. So, GPT 5.6 Soul is half of Fable 5. So, GPT 5.6 Soul is half of Fable 5. So, GPT 5.6 Soul is similarly priced to Opus 4.8. Now, similarly priced to Opus 4.8. Now, similarly priced to Opus 4.8. Now, here's the thing. I think that GBT 5.6 here's the thing. I think that GBT 5.6 here's the thing. I think that GBT 5.6 Soul could have easily given me a Soul could have easily given me a Soul could have easily given me a similar video output if I didn't put it similar video output if I didn't put it similar video output if I didn't put it on ultra. I think because it was on on ultra. I think because it was on on ultra. I think because it was on ultra, it tended to sort of overthink, ultra, it tended to sort of overthink, ultra, it tended to sort of overthink, over delegate, and that's where the over delegate, and that's where the over delegate, and that's where the tokens really started to add up. I bet

  5. tokens really started to add up. I bet tokens really started to add up. I bet if I would have done this exact same if I would have done this exact same if I would have done this exact same prompt on high or very high, we would prompt on high or very high, we would prompt on high or very high, we would have gotten a similar result and have gotten a similar result and have gotten a similar result and probably half the cost. And that's why probably half the cost. And that's why probably half the cost. And that's why typically when I use models like this typically when I use models like this typically when I use models like this that are so capable, I don't like moving that are so capable, I don't like moving that are so capable, I don't like moving the effort above high. But really what I the effort above high. But really what I the effort above high. But really what I wanted to show you guys here is how good wanted to show you guys here is how good wanted to show you guys here is how good these models are at giving a pretty these models are at giving a pretty these models are at giving a pretty emotional, vague, ambiguous prompt and emotional, vague, ambiguous prompt and emotional, vague, ambiguous prompt and letting them figure it out. Obviously, letting them figure it out. Obviously, letting them figure it out. Obviously, there's some stuff that it went through there's some stuff that it went through there's some stuff that it went through and it looked through my projects and it and it looked through my projects and it and it looked through my projects and it looked through other videos and took looked through other videos and took looked through other videos and took some inspiration. But if you just get some inspiration. But if you just get some inspiration. But if you just get out of its way and you give it a prompt out of its way and you give it a prompt out of its way and you give it a prompt that has things like delegation and that has things like delegation and that has things like delegation and verification, you will be surprised at verification, you will be surprised at verification, you will be surprised at how far you can get. And then from here, how far you can get. And then from here, how far you can get. And then from here, you're going to iterate, you'll build you're going to iterate, you'll build you're going to iterate, you'll build skills around it, you'll put feedback skills around it, you'll put feedback skills around it, you'll put feedback in, and you will just be able to do a in, and you will just be able to do a in, and you will just be able to do a lot of cool stuff. So, anyways, this lot of cool stuff. So, anyways, this lot of cool stuff. So, anyways, this one's super quick, but if you enjoyed, one's super quick, but if you enjoyed, one's super quick, but if you enjoyed, please leave a like. And I appreciate please leave a like. And I appreciate please leave a like. And I appreciate you guys making it to the end. I'll see you guys making it to the end. I'll see you guys making it to the end. I'll see you guys in the next one.

Summary

The main theme is the power of generative AI models like OpenAI's new Soul and Ultra to automate complex content production workflows. Key subjects include AI agents, coding capabilities, voice cloning with 11 Labs, avatar generation with Hen, and visual mapping with Hyperframes. The practical takeaway is that these advanced AI tools can coordinate multiple agents to execute intricate production tasks from start to finish, significantly streamlining the creation process.

View original episode ↗