Skip to content
CinoidResearch
All notes
Economics6 min read

What a 30-second ad actually costs to generate

We priced three providers on the same job for the same client. The cheapest 1080p option was not the one we expected, and the most expensive mistake is a feature nobody watching will notice.

For the past few years we have run paid social for small US service businesses — pest control, HVAC, plumbing, home care, real estate. We build their sites, write their blogs and manage their campaigns.

And when a campaign needs video, we buy it. Stock licences, occasionally a freelancer. We absorb that cost rather than pass it on, because quoting the real number is how a campaign gets cancelled.

So we priced generating it instead, across four providers, on the same job: one thirty-second vertical spot for an HVAC client, rendered from their own photos.

The comparison

Video-only tiers where one exists, since these ads are watched with the sound off:

Model Provider Per second 30s spot Max res
Veo 3.1 Lite Google Vertex AI $0.03 $0.90 720p
Veo 3.1 Lite Google Vertex AI $0.05 $1.50 1080p
Wan3.0 Alibaba Model Studio $0.05 $1.50 480p
Seedance 2.5 BytePlus $0.06–0.27 $1.92–8.00 720p
Nova Reel Amazon Bedrock ~$0.08 $2.40
Veo 3.1 Fast Google Vertex AI $0.10 $3.00 1080p
Wan3.0 Alibaba Model Studio $0.20 $6.00 1080p
Veo 3.1 Google Vertex AI $0.20 $6.00 1080p

Three things surprised us.

The cheapest 1080p option is Veo 3.1 Lite, not the model we started with. At $0.05 per second it matches Wan3.0’s 480p price while delivering 1080p. We had assumed the Chinese providers would undercut on raw price and on these particular tiers they do not.

Wan3.0’s advantage is not price, it is shape. It does thirty seconds natively, which is exactly one paid-social spot with no stitching. Nova Reel 1.1 goes further — up to two minutes, multi-shot — which matters for longer formats we are not selling yet.

Seedance is the one with a range instead of a number, and that range is the interesting part. It bills per token rather than per second, and references consume tokens. Working from BytePlus’s own published plans — $32 for 5M tokens, quoted at roughly 120 to 500 seconds of 480p — the effective rate lands anywhere from $0.06 to $0.27 per second depending on how much reference material you feed it.

Which puts it in an awkward spot for us specifically. More below.

The pricing model matters more than the price

The Seedance result is the clearest lesson from this exercise, because it is the one where our use case works against us.

Seedance 2.5 accepts up to fifty image, video and audio references. That is the strongest reference control of anything we looked at, and reference fidelity is our binding constraint — the whole product depends on the client’s actual van appearing in the ad.

So the feature we want most is exactly the feature that drives the bill. A reference-light prompt sits near $0.06 per second. A reference-heavy one — which is every job we would ever run — moves toward the top of the range. The model that best serves our requirement charges us most for serving it.

Per-second billing does the opposite: on Veo or Wan, feeding reference images does not change the rate. For a workload that is reference-heavy by definition, that predictability is worth more than a lower floor.

Two further constraints on Seedance for our case: it caps at 720p, and audio is generated alongside the picture rather than sold separately.

The expensive mistake

That last point generalises. Where a provider does sell audio separately, the audio-enabled tier runs roughly double. Veo 3.1 with synchronised audio is $0.40 per second against $0.20 without — $12 instead of $6 for the same thirty seconds.

For paid social that is money set on fire. This creative is watched muted. We burn captions in by default because that is how it gets consumed, which means the generated soundtrack is a feature nobody in the target audience will ever hear.

Anyone comparing these models on a headline per-second number, without checking which tier they landed on or how the billing unit works, will overpay for nothing.

What we compare it against

A single stock clip licence — four seconds of a generic van on a generic street — runs from roughly thirty to two hundred dollars depending on the library and terms. You need several per spot.

A local production shoot starts in the high hundreds and realistically lands in four figures once you count the half day, the edit and the revisions.

Both get compared against the client’s entire media budget, which for these businesses is often three to five hundred dollars a month. Creative costing more than the media is not a rounding problem. It is why these businesses run static images against competitors running motion.

The part that changes behaviour

Cheaper creative is the obvious benefit and it is the less interesting one.

At a dollar or two per spot you stop guessing which hook works. Generate five, run them all on a small budget, let performance pick.

one $1,200 shoot      →  1 hook, no data, hope
five $1.50 renders    →  5 hooks, real data, keep the winner

Seven dollars fifty against twelve hundred, and the cheap version tells you something the expensive one cannot.

Why reference images decide this, not price

If a model cannot take the client’s own photos as input, none of the above matters, because the output is stock footage with extra steps.

Local service buyers run one check before anything else: is this a real company near me. A van in the wrong livery, a house with the wrong roofline, a crew that is too well lit — none of that registers consciously, it registers as something is off, and the thumb keeps moving.

That is why every model on our list supports an image prompt, and why we would pay a premium for reference fidelity before we would pay one for resolution. It is also why the cheapest row in the table is not automatically the answer, and why Seedance stays on the shortlist despite pricing badly against our access pattern.

Where it does not work

Generated video cannot show a specific person’s face reliably enough for a testimonial, and it cannot show a real completed job at a real address. Anything where the point is documentary proof still needs a camera, and for before-and-after restoration work that proof is the ad.

What it handles is the surrounding eighty percent: the offer spot, the seasonal promo, the recruiting ad, the explainer. Those are generic by nature and currently either do not get made, or get made out of clips that visibly belong to someone else.

Where we landed

Nowhere yet, deliberately. We are not committing before we have rendered the same brief through all four and compared output rather than spreadsheets. A dollar per spot is noise next to whether the van looks right, and no pricing page can answer that.

What the exercise settled is the shape of the decision:

  1. Video-only tiers where they exist. Muted creative does not need audio.
  2. Reference support is the gate, price is the tiebreaker — not the reverse.
  3. Check the billing unit, not just the rate. Per-token pricing behaves very differently from per-second once your workload is reference-heavy.
  4. Track cost per delivered spot, not monthly spend.

The next note will have output in it rather than arithmetic.


Veo pricing is from Google Cloud’s published Vertex AI pricing page, Wan3.0 from Alibaba Cloud Model Studio’s model listing, and the Seedance range is derived from BytePlus’s own published token plans. The Nova Reel figure is corroborated from third-party pricing trackers rather than read off AWS’s own table, which is region-gated — treat it as approximate. All of these move often; check current rates before relying on this arithmetic.