New: conversational Urdu and Punjabi speech corpora now licensable  See the catalog →

Practical guides

How much data does a TTS voice need?

How many studio hours a text-to-speech voice actually needs, why TTS data differs completely from ASR data, and what phonetic balance means in practice.

Key takeaways

  • A production-quality single-speaker TTS voice typically needs 20–60 studio hours.
  • Few-shot voice cloning on a strong base model can work from minutes — with quality and control trade-offs.
  • TTS data requirements are the opposite of ASR: consistency and cleanliness, not variation and noise.
  • Phonetic balance matters more than raw volume. Uncovered phoneme contexts produce audible artefacts.

The headline numbers

For a high-quality single-speaker neural TTS voice trained largely from scratch, budget 20 to 60 hours of clean studio recordings from one speaker. Toward the lower end if you are fine-tuning a strong multilingual base; toward the upper end for a new language or expressive multi-style voice.

For few-shot voice cloning on a modern base model, minutes of reference audio can produce recognisable output. The trade-off is control, consistency across long passages, and prosodic quality — adequate for some products, not for a flagship brand voice.

Why TTS data is the opposite of ASR data

This trips up teams who have already built ASR. For ASR you want maximum variation: many speakers, many devices, background noise, dialect spread, spontaneous disfluency. Variation is the point.

For TTS you want the opposite. One speaker, one microphone, one room, consistent distance, consistent energy, no background noise, no disfluency. Every source of variation is a source of artefacts in the synthesised output. An ASR corpus is close to useless for TTS training, and assuming otherwise wastes a procurement cycle.

Phonetic balance in practice

The script must cover the language's phoneme inventory, and critically its phoneme contexts — the diphone and triphone combinations that actually occur. A large recording that repeatedly exercises the same contexts leaves gaps, and gaps produce audible glitches on exactly the words that fall into them.

For low-resource languages this requires building a balanced script from scratch, which means phonological analysis before recording rather than grabbing available text. It is a linguist's job, not a data-entry one.

Recording conditions that matter

Treated room, low noise floor, consistent mic and distance across every session, 48 kHz / 24-bit capture with headroom, and one speaker recorded across sessions close enough together that their voice and delivery have not drifted.

Speaker selection matters more than teams expect. You need someone who can maintain consistent energy and pace across many hours, read naturally rather than performatively, and pronounce the target variety authentically. Professional voice talent is usually worth the premium here.

Expressive and multi-style voices

If you need emotional range — neutral, warm, apologetic, urgent — each register effectively requires its own coverage. Budget additional hours per style rather than assuming a neutral corpus generalises.

The same applies to domain-specific pronunciation: product names, technical vocabulary, numbers and addresses in your specific format. These need explicit coverage in the script or the voice will mispronounce exactly the terms your product depends on.

Related pages on this site

Answers

Frequently asked questions

Typically 20–60 studio hours from a single speaker for a production-quality neural voice, toward the lower end when fine-tuning a strong multilingual base. Few-shot cloning can produce recognisable output from minutes, with trade-offs in consistency, control and prosody.

Generally no. ASR corpora maximise variation — many speakers, devices, noise conditions and disfluency — which is precisely what degrades TTS output. TTS needs single-speaker, studio-clean, consistently recorded audio. They are different products and require separate collection.

A script covering the language's phonemes and, more importantly, the phoneme contexts — diphones and triphones — that actually occur in it. Uncovered contexts produce audible artefacts when the model has to synthesise them, so balance matters more than total recording length.

Need a voice in a language nobody else offers?

We build the script, cast the speaker and run the studio sessions. Tell us the Asian language and register you need.

Average first response: under 6 business hours. NDAs signed same day.

Free samples Get a quote