The first time I heard it, I almost didn’t catch it. Same character, same scene, four clips apart, and her voice had gotten a half-step brighter and about ten percent faster, like she’d swapped actresses between takes and nobody told the sound department. The fix for an AI voice that drifts between clips is to stop generating a new voice per clip and instead lock one reference source, one stability setting, and one saved configuration, then reuse that exact setup for every line the character speaks. Nothing else on my Lost Garden pipeline broke this often, and nothing else was this cheap to fix once I understood why it was happening.
This isn’t a Lost Garden problem. It’s an AI video problem. Every clip you generate is a fresh, memoryless draw from a model that has no idea what happened in the clip before it, and that’s true for the picture and just as true for the voice underneath it.
An AI voice changes between clips because most voice tools re-interpret your text and settings independently on every generation, with no memory of how the last line sounded. Nothing enforces continuity by default. Pitch drifts, pacing speeds up or slows down, and the emotional register resets to whatever a low-stability setting happens to land on for that specific sentence.
It gets worse across tools, not just across clips. If you generate dialogue in one AI voice tool for scene one and switch providers, models, or even just voice IDs for scene four, you’re not dealing with drift anymore, you’re dealing with a different voice pretending to be the same character. That’s the single most common mistake I see indie AI filmmakers make, and it’s the one I made first: treating “close enough” as good enough, then discovering in the edit that four scenes in, nobody sounds like themselves.
A face reference keeps a character’s look consistent shot to shot. Almost nobody applies the same discipline to the character’s voice, and it’s just as detectable to an audience.
Here’s the workflow I run on every Lost Garden character now, after losing a full week of dialogue to re-recording:

I do most of this dialogue work in ElevenLabs.
The stability setting controls how closely a generated voice sticks to your original reference audio, not how “calm” the character sounds. That naming trips up almost everyone the first time.
ElevenLabs documents three practical positions on this dial, and picking the wrong one is the second most common way a voice quietly stops sounding like itself:
Stability slider diagram
The tradeoff is real: a low-stability, high-expressiveness setting sounds more alive in isolation and drifts fastest across a batch of clips. A high-stability setting holds the line but can go flat on a scene that needs a genuine emotional swing. I pick per scene, not once per character, and I write the choice down next to the shot so I’m not guessing again in three weeks.
Audio tags help close the gap without abandoning stability. Bracketed cues like [whispers], [frustrated sigh], or [laughs] let you push a locked, stable voice into a specific emotional beat for one line, then return to baseline for the next, instead of loosening the stability setting for the whole scene just to get one moment right.
Lost Garden’s lead has more dialogue than any other character in the series, spread across shots generated weeks apart, on different days, sometimes on different laptops. Early on, I was regenerating her lines scene by scene, tweaking stability whenever a line felt “off” without writing down what I’d changed. By episode two, three different versions of her voice existed in the timeline, and only the strictest side-by-side listen caught it before it shipped.
The fix wasn’t a better model. It was discipline: one locked reference sample, a Natural stability setting as her default with Creative reserved for exactly two emotional peaks in the whole episode, and a single line of notes in the shot plan I keep next to her camera and lighting continuity. Batching her dialogue by character instead of by scene cut the re-generation rate on her lines by more than half, because drift got caught in the same sitting it was created, not three scenes later in the edit.
If a locked voice source is the fix, then the actual failure mode is treating each clip as its own island instead of one continuous performance spread across many separate generations. That reframing changed more of my workflow than any single setting did.
Why does my AI-generated character’s voice change between video clips?
Because most AI voice tools generate each clip independently, with no memory of previous lines, so pitch, pace, and emotional tone can shift unless you reuse the exact same reference audio and settings every time.
What does the stability setting actually control in AI voice cloning?
It controls how closely the output sticks to your reference recording. Lower stability means more expressive but less predictable output; higher stability holds closer to the original voice at the cost of emotional range.
How much reference audio do you need for a consistent AI voice?
A longer, continuous sample produces more natural pacing than a short clip; thin source audio tends to carry its own pacing problems into every line generated from it.
Can one cloned voice work for a character across a whole series?
Yes, as long as the same reference source and saved configuration are reused for every clip. The moment you swap sources or rebuild the settings from memory, consistency breaks.
None of this requires a bigger budget or a better model. It requires treating a character’s voice the way you’d treat their face: locked once, reused deliberately, and logged somewhere you’ll actually check before the next batch. I keep that log next to the rest of my shot planning in ScreenWeaver, because a voice note that lives in a separate app is a voice note nobody opens on generation day.
What’s the one continuity detail you keep losing track of across your own AI-generated shots? I’d rather compare notes than pretend I’ve solved all of them.