3-Step Workflow To Make Ultra-Realistic AI Ads
By Higgsfield AI · Tech · 379.9K views · 35:37
The teardown in brief
What's working
- The hook opens on the finished commercial with zero preamble — the viewer immediately sees proof that the method works before Adele says a single word. This is textbook Result-Why-Journey execution and justifies the tutorial investment right from the start.
- The three-stage framework is stated clearly at 1:17 and used as a structural spine throughout. Viewers always know which stage they're in and why it matters, which dramatically reduces the cognitive load of following a 35-minute workflow.
- Pro tips are distributed throughout rather than front-loaded. The gray background insight, the face-locking technique, the schematic map hack — each arrives exactly at the moment the viewer needs it, making the advice feel discovered rather than delivered.
What's costing attention
- The scene generation section (Stage 3) runs on the same iterative rhythm for 20 minutes with no re-engagement mechanism between scenes. The workflow is repeated rather than escalated — scenes two through five don't feel harder or more impressive than scene one, just different.
- Stakes are entirely implicit throughout. The only stated consequence of doing things wrong is 'wasted credits' — which is mild. There's never a moment where the viewer feels 'this could go badly in a real way.' The schematic map section comes closest but doesn't lean into the tension.
- The video has almost no emotional range in the narration. The audio energy confirms this: 96% of the video sits at a single conversational level. For an enthusiast audience this is acceptable, but a few genuine moments of 'oh this actually surprised me' would create contrast that helps the video breathe.
The first 30 seconds
>> I'm on my way. Hi, I'm Adele. What you just watched is a full commercial I made completely with AI on this laptop. And here's why you'll want to stick around with. By the end of this video, you'll be able to make one of these yourself for your product, your brand, or just an idea you've been sitting on for years. >>
The strongest possible opening for a tutorial — the finished commercial plays immediately with no preamble, delivering the exact result the thumbnail promises before Adele introduces herself at 0:53. Viewer confusion is eliminated within seconds: you see the product, you understand what's being taught, and you have visual proof it works.
Where viewers drop
33:00 — Repetitive Scene Generation Loop (critical)
From the street corner through scene five, you run the same generate-identify-problem-fix-iterate cycle for the fourth time in a row. By this point the viewer has watched you debug lighting, fix faces, rewrite prompts, and patch choreography across three prior scenes. The loop is identical in structure, and there's no new mechanic introduced — just a different location.
Why it matters — After 20 minutes of the same troubleshooting rhythm, the committed viewer is still with you, but they're on autopilot. The energy cost of watching each new scene iterate from broken to fixed starts to feel like diminishing returns rather than forward progress.
12:16 — Thin Stage Two / Shot List Context Dump (moderate)
Stage two runs for about two and a half minutes explaining what a 'skill' is, how the shot list document is structured, and why prompts are connected. It's all conceptual setup — no generation happens, no problem is solved, and no result appears on screen. The viewer clicked to watch an AI ad get built, and this is the longest stretch of the video where literally nothing is generated.
Why it matters — The viewer's momentum was high after watching all the assets lock in Stage 1. Stopping for a two-minute architecture lecture before the first scene is ever generated is the closest this video comes to a classic context dump. Viewers who wanted to skip ahead to 'the good part' will do so here.
21:00 — Stadium Scene Over-Iteration (moderate)
The stadium scene runs for five full minutes covering outfit creation, a wet character variant, a lighting override, a match-cut fix, a body-rig shot, and a product close-up. Six distinct problems are introduced and solved in sequence. While each technique is genuinely useful, they arrive without any sense of which are the two or three most important — everything is weighted equally.
Why it matters — By the time the snorricam body-rig shot finally lands, the viewer has absorbed five prior problem-solution cycles for this scene alone. The payoff ('Oh yeah, look at that') is real, but it arrives after so many smaller fixes that the momentum never peaks — it just keeps going at the same altitude.
34:29 — Weak Verbal Closing After Commercial Playback (mild)
The full commercial plays back from 33:43 to about 34:09, which is the emotional peak of the video — the payoff of everything built. But the verbal wrap that follows is a sequential recap of the three stages the viewer just watched for 34 minutes. The closing line ('now go make yours') lands softly rather than with the weight the finished ad deserves.
Why it matters — After watching a genuinely impressive AI-generated commercial, the viewer is primed to feel something. A step-by-step recap immediately deflates that feeling. The outro is also the second most likely exit point after the commercial ends — viewers who got what they came for will leave the moment the ad finishes.
How the video is built
- 0:00 Hook + Promise — AI commercial plays in full, Adele introduces herself and the three-step framework
- 1:40 Stage 1 — Asset Creation — Building the character, locations, props and outfits using Soul Cinema, GPT Image 2.0, and AI Cast
- 12:17 Stage 2 — Shot List — Explaining and deploying the Claude skill to turn script and assets into a connected shot list
- 14:54 Stage 3 — Scene Generation (Scenes 1–3) — Iterative generation and debugging of kitchen, stadium, and street scenes
- 31:14 Stage 3 — Scene Generation (Scenes 4–5) — Office and boss scenes — climax of the commercial narrative
- 33:55 Final Playback + Wrap — Full commercial plays back, workflow summary, and CTA
What any creator can steal
- The scene generation loop needs one new mechanic per scene, not the same cycle
- Stage two context dump needs to be replaced with live generation
- Add one explicit stakes line before scene generation begins
- The post-commercial wrap is burying the emotional peak
- Add a progress counter between scenes
- Film a brief 'worst case' clip for every major technique — show what the output looks like when you skip the testing step, when you don't erase the face, when you just type 'he dances' without choreography. The failure examples make the solutions feel like rescues rather than procedures, and they create the contrast that holds retention through technical sections.
Want this on your own video?
Paste any YouTube URL and Retti maps every drop, spike and plateau to the moment that caused it.
Analyse a video free