Faceless Desk
The tools behind faceless video, tested by running one.

What a minute of AI narration actually costs

Published 2026-08-28 · updated 2026-08-28 · Research-led: compiled from vendor documentation, not hands-on use.

Disclosure: some links on this page earn us a commission if you subscribe, at no extra cost to you. It never changes what we recommend — where a free or cheaper tool wins, we say so.

Comparing AI voice pricing is deliberately hard. One vendor sells credits, another characters, another audio tokens. The units don’t line up, so most comparisons give up and rank the tools by vibes.

So here’s the same question asked once, in a unit that matters: what does one minute of finished narration cost?

The conversion

Everything below assumes 900 characters per spoken minute — roughly 150 words a minute at about 6 characters per word including spaces.

That single number drives every figure on this page. If your narration is faster or denser, scale accordingly. It’s an estimate, stated up front so you can redo the arithmetic rather than trust mine.

The table

OptionPriceCost per minuteFree allowance
Google Standard / WaveNet$4 / 1M chars$0.00364M chars/mo
OpenAI tts-1$15 / 1M chars$0.0135
Google Neural2$16 / 1M chars$0.01441M chars/mo
OpenAI tts-1-hd$30 / 1M chars$0.027
Google Chirp 3: HD$30 / 1M chars$0.0271M chars/mo
Google Studio$160 / 1M chars$0.1441M chars/mo
ElevenLabs Starter$6/mo for 30,000 credits~$0.18
ElevenLabs Creator$22/mo for 121,000 credits~$0.16

The ElevenLabs figures assume its historical billing of roughly 1 credit per character, and assume you consume the entire monthly allowance — unused credits are wasted, so the real per-minute cost is higher for anyone who doesn’t hit their cap.

The spread is about 50x

A minute of Google WaveNet narration costs about a third of a cent. The same minute through ElevenLabs Starter costs about 18 cents.

Put that against a real workload. An 8-minute video is roughly 7,200 characters:

The free allowance nobody mentions

This is the finding that reframes the whole decision. Google’s free monthly usage limit for Standard and WaveNet voices is 4 million characters.

At 900 characters per minute, that’s roughly 4,400 minutes — about 74 hours of narration per month, at no cost. For comparison, ElevenLabs’ $6 Starter tier buys about 33 minutes.

Two important caveats before you get excited:

  1. Billing must be enabled. Google’s page is explicit: “You must enable billing to use Text-to-Speech, and will be automatically charged if your usage exceeds the number of free characters allowed per month.” This isn’t a no-card free plan. Overage bills silently and automatically — set a budget alert.
  2. It’s metered infrastructure, not a consumer free tier. That’s why it doesn’t carry the non-commercial restriction that ElevenLabs’ free plan does. Still confirm the Cloud terms cover your use rather than assuming it from this page.

So why would anyone pay 50x?

Because per-minute cost isn’t the only axis, and it would be dishonest to pretend the cheapest number wins.

What the money actually buys:

None of that is nothing. It’s just worth knowing you’re buying those things, rather than assuming the subscription is cheaper because its headline number is small.

The honest read

If narration is a commodity in your workflow — clear, neutral, informative — the API route is dramatically cheaper and the free allowance may cover you entirely for a long time.

If the voice is part of your channel’s identity, or you can’t or won’t write code, the premium tier is a real purchase and not a rip-off. Just size it against the $6/month floor: at ElevenLabselevenlabs.io · affiliate Starter, voice is your entire budget, and it buys about four videos.

The trap is paying premium subscription rates for narration you’re treating as a commodity anyway. That’s the case where the 50x is pure waste.

Caveats: every figure is read from vendor pricing pages, not from our own bills. The 900-characters-per-minute conversion is an estimate and drives all per-minute numbers. Google’s Gemini-TTS models price by audio tokens rather than characters and are excluded, because they aren’t directly comparable on this basis. Pricing changes often — recheck before committing.

Sources