Descript rebuilt its dubbing pipeline around syllable counting
Descript used GPT-5 reasoning models to fix dubbing's pacing problem, lifting duration adherence by up to 43 points and boosting dubbed-video exports 15% in 30 days.

Updated
Why it matters
- Dubbed-video exports rose 15% in the first 30 days after Descript's pipeline redesign, with duration adherence improving 13 to 43 percentage points depending on language.
- Segments within the natural pacing window (playback slowed ≤10% or sped up ≤20%) rose from 40–60% to 73–83%; 85.5% of segments scored 4 or 5 out of 5 for semantic fidelity.
- Reliable syllable counting by GPT-5 series models was the key enabler; CEO Laura Burkhauser says batch dubbing and lip-sync for entire content libraries is next.
Descript increased exports of translated videos with dubbing by 15% in the first 30 days after redesigning its translation pipeline around OpenAI reasoning models — and duration adherence improved by 13 to 43 percentage points, depending on the language.
The numbers come from Descript itself, and they mark a turning point for a feature that had been stuck on a hard technical problem: translated speech that ran too long or too short for the video it was meant to fill. For a company positioning itself as an AI-native video editor — one built on the premise that if you can edit text, you should be able to edit video — dubbing at scale is a core product bet, not a side feature.
Descript has used OpenAI models for years, running Whisper for transcription and GPT series models inside its co-editor, Underlord. Translation emerged early as one of the product's most requested capabilities. Traditionally, translating video has been slow and expensive, requiring language experts to manage projects, produce translations, handle quality control, and generate corresponding audio. Large language models compress that workflow dramatically, and that compression is what makes enterprise-scale localization economically viable.
Where dubbing broke down
Descript started with captions-only translation. It worked well. But users wanted spoken audio in the target language, and once dubbing shipped, one complaint dominated.
"Probably the number one complaint we heard was that the pace of the speech was unnatural in the translated language," said Aleks Mistratov, Head of AI Product at Descript.
The root cause is structural. Different languages take different amounts of time to express the same idea. Descript observed that, on average, German is a "longer" language than English. The company's own example: the English sentence "Please review the safety guidelines before operating the machine" carries 18 syllables; its German translation, "Bitte überprüfen Sie die Sicherheitsrichtlinien, bevor Sie die Maschine bedienen," carries 24 — a 40% increase.
To fit that German audio into the fixed video segment, creators had two bad options: artificially speed up or slow down the playback, or rewrite the translation itself to fit the time budget. "You'd end up with something that sounded like chipmunks, or a sleepy giant," Mistratov explained. Both workarounds required deep timeline edits and often near-native fluency in the target language. For individual creators, that was tedious. For enterprise customers wanting to localize entire libraries, it was a blocker.
The distinction matters because captions and dubbing place different demands on the same translation system. Both require semantic fidelity — the translation must preserve the original meaning. But duration adherence is a nice-to-have for captions and a hard requirement for dubbing, because mis-timed speech sounds unnatural even when every word is correct.
A counterintuitive bottleneck: counting syllables
Descript's earlier pipeline optimized semantic fidelity first and tried to correct timing afterward. The translations were often semantically correct but routinely missed duration constraints, and the overall quality wasn't good enough.
The team's diagnosis started with a deceptively simple experiment. "We ran incremental tests, not even generating anything, just asking the model to output the number of syllables in a chunk of text," Mistratov said. "Earlier models simply weren't good at that."
Reliable syllable counting turned out to be the critical capability. If the model cannot consistently calculate syllables, it cannot reliably target a specific duration window. GPT-5 series models brought the reasoning consistency earlier models lacked, particularly on tasks like syllable counting and constraint tracking. With that improvement in hand, Descript redesigned the pipeline from scratch.
The new system works in stages. First, it breaks the transcript into chunks, guided by sentence boundaries, natural pauses, and speaking patterns in the original recording. Each chunk preserves semantic continuity but stays small enough to reason about as a single timing unit. The model then calculates the syllable count of the chunk. Using language-specific speaking-rate assumptions, the system estimates how many syllables the translated chunk should target to preserve natural pacing. The prompt instructs the model to optimize for both duration adherence and meaning preservation, with surrounding chunks passed in as context so meaning holds across segments.
The team evaluated multiple configurations to balance duration adherence, semantic fidelity, latency, and cost. The selected setup delivered strong constraint-following at production speed, enabling high-volume translation without manual retiming. Pacing is now treated as a first-class variable during generation, not something corrected after the fact.
Measuring what "natural" means
To set acceptance criteria, the team ran listening tests: they generated translated audio samples, adjusted playback speed in small increments, and asked users to identify when speech became unnatural.
"Anything that was slowed down by 10%, or sped up by 20%, generally still sounded natural," Mistratov said. Beyond that range, speech became too distorted.
Earlier systems performed poorly against that bar. Depending on the language, only 40% to 60% of segments fell within the acceptable pacing window. The redesigned pipeline raised that to between 73% and 83%.
Semantic fidelity got its own measurement, using a separate model-as-judge rating on a scale from 1 ("completely different") to 5 ("semantically equivalent"). Here Descript made a deliberate tradeoff: for dubbing, it accepted a lower semantic threshold than for caption-only translation, where duration constraints are irrelevant. Even so, 85.5% of segments were rated four or five for semantic adherence.
Because both metrics are automated, the team can continuously evaluate new model releases and prompt variations against the same benchmarks — turning what was once an art into a measurable engineering discipline.
The business stakes
The 15% jump in dubbed-video exports within 30 days signals real demand, and Descript is moving to capture it at library scale. "Dubbing is an increasingly popular use case for Descript, so we're building ways to do it in batch for companies that want to translate and lip-sync entire libraries," said Laura Burkhauser, CEO.
Translation inside Descript is one layer of a broader multimodal system. Translated text feeds into speech generation, which drives lip sync and final video rendering. The text-layer improvements make natural pacing possible, but the end experience still depends on how well the audio model preserves tone, cadence, and nonverbal characteristics of speech.
That is where the team sees the next frontier. "A lot of what's going to improve translation output is making the pipeline more multimodal: incorporating audio, video, and text together when deciding how to translate," said Mistratov. "That should better maintain the nonverbal characteristics of speech, like tone and emphasis, and preserve even more of the original delivery."
For Descript, stronger reasoning models made dubbing tractable. Once models could reliably balance the tradeoff between pacing and meaning, translation became something the team could systematically improve and deploy at scale — and as batch localization tools ship, that capability will be tested against the largest content libraries in the market.
Original: descript.com
More from Sophie Lindqvist
Show full bio
Staff writer covering marketplaces and e-commerce at AI In Context.
115 articles