What the reference clip actually decides
The embedding captures timbre, pitch range, resonance, and a good deal of accent and speaking rhythm. It does not capture intent. The model does not know that you want line three read as a question, or that the product name should land hard.
What surprises people is how much of the reference's delivery comes along. Record the reference in a bright, energetic read and every sentence tends to arrive energetic. Record it half asleep and the clone sounds tired regardless of your text. So pick the reference to match the job: a neutral, evenly-paced read for narration, and a warmer, faster one for advertising.
Reference audio checklist
- One speaker, no music, no overlapping voices.
- As little room reverb as possible. Reverb is the single most damaging artifact, since the model learns it as part of the voice and then applies it to every line.
- No aggressive noise reduction. Over-processed audio produces a clone with a metallic, gated quality that is hard to diagnose afterwards.
- Consistent level, conversational loudness, nothing clipped.
- Cut on sentence boundaries, not mid-word.
- Avoid whispers and shouting unless that is the register you want out.
Controlling delivery
Once a voice clone exists, most of your control is in the text and the segmentation, not in the model.
- Punctuation is timing. Commas, full stops, and paragraph breaks are the main pacing controls. An unpunctuated paragraph reads as a breathless run.
- Spell things out. Write numbers, units, dates, and acronyms the way they should be spoken. Respell unusual names phonetically until they land, since this is faster than any parameter.
- Segment, then stitch. Generate sentence by sentence and assemble with deliberate pauses. Long single passes drift in pitch and pace, and any flaw forces a full rerun.
- Take three. Output varies run to run. When a line matters, generate several and pick, rather than reworking the text.
Where clones still fall apart
Proper nouns and acronyms are the most common failure and the easiest to fix by respelling. Long sentences tend to lose pitch control near the end and can jump register. Emphasis is generally not where a human would place it, which is why synthetic narration can be technically clean and still sound uninvolved. Sibilance and breath noise sometimes come through exaggerated, inherited from a close-miked reference.
None of these are solved by a longer reference recording. They are solved by shorter text segments, explicit spelling, several takes, and light editing afterwards, which is exactly how a human voiceover session works too.