This year we produced a broadcast-grade commercial almost entirely with AI video models: a scripted spot built to the standard a TV and CTV placement demands, not a sizzle reel for a conference deck. So when a marketer asks whether an AI TV commercial is genuinely viable in 2026, we can answer from the production chair rather than a vendor's showreel. The first thing we learned is that "can AI do it" is the wrong question. The right one is whether your production process can absorb the specific ways AI fails.
What we set out to make
The brief we set ourselves was deliberately unforgiving: a scripted spot at standard TV length with a person on camera, dialogue, a real product in shot, and multiple locations, at a quality that survives a 55 inch screen rather than a phone at arm's length. We make AI-led ads for paid social every week (our honest read on where AI beats real creators is here), and TV raises the bar in three ways. Shots hold longer, so artefacts that vanish in a 1.5 second social cut get time to register. Sound is on by default, so speech timing is exposed. And the ad must pass clearance, not just an ad-platform review.
We were not first through this door. The widely covered Kalshi spot, made with Google's Veo 3, aired in the YouTube TV stream of Game 3 of the 2025 NBA Finals after roughly two days of work and under $2,000 in generation costs. The detail that matched our experience was not the price but the ratio: its creator reported 300 to 400 generations to get 15 usable clips. Anyone who tells you AI video is a one-take medium has not shipped with it.
Where the models held up
Credit first. The current crop of video models (Seedance, Kling, Veo and their peers) gave us four things a traditional shoot cannot match at any sane budget.
Locations stopped being a constraint. Night exteriors, weather, a crowd, a kitchen and a street in the same spot: each one is a prompt and a reference image, not a permit and a unit move.
Concept iteration happened before commitment. We tested visual treatments of the same script against each other in days. A traditional production commits to one treatment at storyboard and finds out after the shoot whether it was the right call. Seeing several versions of the idea before locking one is, for us, the biggest genuine advantage.
Retakes cost minutes, not a reshoot. When a line reading or a composition was wrong, we regenerated it: no crew recall, no actor availability, no matching the light from three weeks ago.
Physics and motion are mostly believable now. Cloth, liquids, walking, handling objects: the uncanny failures that defined 2024-era AI video are now the exception, and the best models generate convincing ambient sound and foley with the picture.
What breaks first in an AI TV commercial?
Now the part no tool vendor will tell you: the failure modes we hit, in the order they cost us time.
1. Speech timing and lip-sync
Dialogue is where AI video is still weakest, and TV is the format least forgiving of it. Buy a clip slightly too long for the line and the model does not slow the delivery; it pads the surplus with dead air, which reads as an error on TV. Ask for a dramatic pause and you get one about twice what you asked for, with no reliable way to direct it shorter. Feed in a pre-recorded voice track and the model re-times it: gaps stretch and word endings deform. Lip-sync on the model's own native voice is good; lip-sync to audio you supplied is a gamble you check clip by clip. We lost more time to dead air than to any visual artefact.
2. On-screen text
Generated supers, price cards, captions and end-card wordmarks still come back warped often enough that we treat on-screen text as a post-production job, full stop. We generate clean framing with negative space and the edit adds the type. On broadcast this matters doubly: the legal text on a TV ad is not optional decoration.
3. Product and brand consistency
Every generation is independent: the model has no memory of the shot before, so faces drift, outfits change and rooms rearrange themselves between cuts unless you force consistency with reference assets. The product is the sharpest version of the problem. Generate a named product from a text description or a white-background packshot and it comes back at invented proportions, because a packshot carries no scale information. At TV size, a product 20% too large against the hand holding it is instantly wrong, and a truthfulness problem for clearance besides.
4. The small tells, magnified
A background described only as "a street" renders frozen, like a photo backdrop. Voices are a lottery: two clips give you two different voices, and accents wander between takes. Unnamed off-screen speakers come back male unless you say otherwise. Models invent closing ad-libs, a cheery "thank you!" mouthed into what should be a silent hold. None of this is fatal alone. On a big screen, with sound on, it compounds.
The models are good enough to make a TV ad. They are not good enough to make your TV ad without a production discipline wrapped around them.
How we worked around it
Everything that saved the project came from treating AI generation like a production department with known weaknesses, not a vending machine.
Lock the references before generating anything. We built a fixed bundle: a character sheet per person, plates of each location, and real photographs of the product, including one in a real room beside real objects so the model had true scale. Every shot was generated from that bundle, and when we could not source a real in-context photograph of the product, we did not generate the shot. That one rule killed most of our consistency problems.
Direct everything, every prompt. Leave the camera, lighting, background movement or soundscape to the model's defaults and you get its invention instead. Asking for "cinematic" footage produces drifting, orbiting camera moves no TV director would sign off; you describe the exact camera behaviour you want and what must never move.
Decide the voice route up front. Two options work, and they trade off. The model's native voice bends its performance to the picture, carries breaths and effort for free, and syncs cleanly, but you cannot hold it consistent across shots. A recorded or professionally cloned voice track gives you one controlled voice for the whole spot, but you fit picture to it in the edit rather than trusting the model's sync. For a TV spot with one narrator, we would take the recorded track and cut to it every time; for on-camera performance, native voice and careful shot selection.
QC mechanically, not by vibe. Every clip was speech-to-texted and diffed word against word with the script, because models drop and add words with total confidence. We measured dead air. We checked spoken numbers, since a figure written as digits gets read aloud however the model fancies. Before committing to any batch we generated one probe clip to check voice, accent and scale. Boring work, and the reason the finished spot held together.
Finish like a normal ad. A human edit, music, supers, grade and mix. The generation replaced the shoot. It replaced nothing else.
Will an AI-generated ad clear for broadcast?
Yes, and the UK path is better mapped than most marketers expect. Every ad on UK commercial TV is cleared against the BCAP Code; Clearcast's guidance is to leave two to three weeks for its three-stage process, and all stages must be completed even if your ad is already made. An AI ad goes through the same gates as any other ad. The rules that bite hardest concern truthfulness: demonstrations must be genuine and the ad must accurately represent the product. A generated product shot with flattering invented proportions is not a style choice to a clearance consultant; it is a misleadingness problem. Hence our refusal to generate product shots without real reference photography.
The precedent is already set at broadcaster level. Channel 4 launched a service that uses generative AI to create TV-ready ads for its streaming platform, and the first fully generative ad aired with Clearcast involved from the early stages through to delivery. Clearcast notes in the same piece that only 7,000 of the UK's 3 million advertisers currently run TV campaigns, the gap AI production is aimed at. In the US there is no single clearance body; each network and platform applies its own standards, partly why the earliest AI spots surfaced on streaming inventory first.
Clearance is not the only court, though. Coca-Cola's AI-generated Christmas ad made it to television and still took a public beating for how it looked. Audiences now recognise AI footage, and a heritage brand swapping craft for generation reads as a statement whether it means to or not. A challenger brand testing TV for the first time carries none of that baggage. Know which one you are. (How AI ads are policed on Meta and TikTok is its own question, with clearer disclosure rules.)
AI production vs a traditional shoot
The honest side-by-side for a brand weighing the two routes on a TV or CTV spot.
| Dimension | AI production | Traditional shoot |
|---|---|---|
| Concept iteration | Test several full visual treatments before committing. The decisive advantage. | One treatment locked at storyboard; you learn if it worked after the shoot. |
| Talent | Any face, any age, no availability or usage-rights renewals; consistency needs reference discipline. | Real performance and genuine emotional nuance; casting, fees and usage terms to manage. |
| Locations | Effectively unlimited; night, weather and crowds cost the same as a kitchen. | Permits, travel and unit moves; each location multiplies cost. |
| Product fidelity | The weak point. Needs real reference photography with true scale, or proportions get invented. | The camera photographs the actual product. Fidelity is free. |
| Dialogue and performance | Workable with a strict voice route and shot-by-shot checking; still the most retake-hungry area. | A director talks to an actor. Still unbeatable for performance-led spots. |
| Revisions | A regeneration. Minutes to hours, at render cost. | A reshoot. Crew, talent and location all over again. |
| Compliance and clearance | Same rules as any ad, plus extra scrutiny on whether generated footage truthfully represents the product. | Established path; clearance consultants know exactly how to read it. |
| Where it wins | Concept-led, stylised or impossible-to-shoot ideas; fast tactical spots; brands testing TV for the first time. | Close-up product craft, testimonials, emotional performance, heritage brand work. |
The verdict: should you make your TV commercial with AI?
Having shipped one, we will give the verdict as a decision rule, not a hedge.
AI is the right call when the idea carries the ad: comedy, spectacle, a stylised world, a montage driven by voiceover. It is the right call when you want to test TV or CTV as a channel without betting a production budget on one spot, and the market is moving that way: IAB projects US digital video ad spend passing $80 billion in 2026, growing 11% year on year, and describes AI for digital video as accelerating "from experimental to operational". And it is the right call when your concept needs five locations and a storm, because no other budget line gets you that.
AI is the wrong call when the sell is a close-up of the product doing its job, when the ad rests on a real person's testimony (which clearance requires to be genuine), or when the performance itself is the idea. It is also the wrong call if nobody on your side owns the unglamorous work: reference building, prompt direction, mechanical QC and the edit. The generation is maybe a third of the job. Hand the other two thirds to no one and you will ship the reason people distrust AI ads. Weighing doing that work in-house against buying it done? We wrote up the honest tools-vs-studio comparison separately.
The conclusion we did not expect: the hybrid spot beats the pure one. AI environments and b-roll around real product photography, a recorded voice, and a human edit produced something we would put on television. A single prompt-to-screen pipeline did not. Full disclosure: this is work Spark does end to end for brands (the process, the work), and everything above is why we run it as a production discipline rather than a tool subscription.
Key takeaway
AI can make a broadcast-standard TV commercial in 2026, but only inside a real production process: locked reference assets for people, places and product, a deliberate voice route, mechanical shot-by-shot QC, and a human edit. Clearance treats an AI ad like any other ad, and the truthfulness rules punish invented product footage hardest.
FAQ: AI TV commercials
Can AI actually make a TV commercial?
Yes. AI video models can now produce footage that holds up on a television screen, and AI-generated spots have already aired on national US broadcasts and UK streaming TV. What it cannot do is produce a finished commercial on its own: the script, the reference assets, the voice decision, the quality control and the edit are still human work. Treat the model as the camera department, not the agency.
Are AI-generated ads allowed on TV?
Yes, in both the UK and the US, provided they meet the same standards as any other ad. In the UK, every ad on commercial TV goes through Clearcast against the BCAP Code, and an AI-generated ad is assessed like any other: claims must be substantiated and demonstrations must be genuine, so a generated product shot cannot misrepresent the real product. In the US each network and streaming platform applies its own standards, and AI spots have already cleared them.
How long does an AI commercial take to produce?
Generation itself is fast: the widely reported Kalshi NBA Finals spot was made in about two days. But generation is only one stage. A realistic schedule adds scripting, building reference assets, several rounds of generation and rejection, a human edit with music and supers, and, for UK broadcast, two to three weeks of Clearcast clearance. AI compresses the shoot; it does not compress the thinking or the compliance.
How much does an AI TV commercial cost to make?
Far less than a traditional shoot, but more than the headline figures suggest. Kalshi put the raw cost of prompting the AI for its NBA Finals spot at under $2,000, the part that stood in for studios, directors and actors, but that figure left out the creator's fee, the scriptwriting, the hundreds of rejected generations, the edit and the media buy. The honest framing: AI shifts budget from production to iteration and judgement, rather than making the ad nearly free.
More production write-ups live in the resources hub. Rather hand the discipline to a studio that has already made the mistakes? Tell us what you are trying to make.