What Localisation Built Into Production From Day One Actually Looks Like
The localisation guide covers what AI dubbing costs and which tools produce which quality tier across which languages. This post goes deeper into the specific production decisions that determine whether the localisation happens cleanly at the end of the pipeline or requires expensive remediation before it can begin.
Most productions treat localisation as a post-delivery activity. The finished series is delivered to the primary platform. The platform acquires it. Months later, the production company considers expanding into secondary markets and discovers that the production was not built for localisation. The audio mix is a combined stereo file with no separated stems. The scripts were written without dubbing character count constraints. The generation workflow produced character mouth movements calibrated to the English dialogue's phoneme patterns, which do not match Spanish or Hindi phoneme patterns even after AI dubbing.
Each of these discoveries requires remediation work that the localisation cost estimate does not include. AI stem separation on a combined stereo mix produces usable but not optimal stems. Script reformatting for dubbing retrospectively requires a writer's time. Re-generation of character mouth movements for dubbed versions requires returning to the generation tools. The total remediation cost frequently exceeds the original localisation estimate by 40 to 60%.
The production company that builds localisation into day one does not encounter these costs. The stems are clean because they were recorded as stems. The scripts are dubbing-ready because they were formatted that way from the first draft. The generation workflow produces mouth movements that hold across multiple phoneme systems because the character reference pack was built to accommodate variation. The localisation at the end of this pipeline is faster, cheaper, and higher quality than the localisation at the end of a pipeline that was not designed for it.
This is what building localisation into day one actually requires, decision by decision.
Pre-Production Decision 1: Script Formatting for Dubbing
The script's first localisation requirement is character count discipline. Dubbing replaces the original dialogue with a translated version that must match the timing of the original actor's lip movements. A dialogue line that takes the actor 3.2 seconds to deliver must be translated into Spanish, Hindi, and German versions that can each be delivered in approximately 3.2 seconds.
Languages expand and contract differently from English. Spanish and French typically expand by 15 to 25% relative to English word count. German expands by 20 to 30%. Japanese contracts by 10 to 20%. Hindi is variable but typically contracts relative to English for equivalent semantic content.
A script formatted without dubbing character count awareness produces English dialogue lines that are at the maximum comfortable length for the actor's delivery pace. When those lines are translated into Spanish, the translated version is 15 to 25% longer than the English version, which means the Spanish dubbing voice either speaks faster than natural pace or the translation must be truncated, losing semantic content.
The dubbing-ready script format addresses this at the writing stage through three specific constraints:
The 80% rule. Dialogue lines are written to occupy approximately 80% of the natural comfortable delivery pace for the actor. The remaining 20% is white space that the longer languages can absorb without exceeding the timing of the original delivery. A line that could naturally be delivered in 4.0 seconds is written to be deliverable in 3.2 seconds, leaving 0.8 seconds of accommodation for Spanish's 25% expansion.
The short sentence preference. Dubbing synchronisation is more forgiving on short sentences than on long ones. A 15-word sentence that expands to 18 words in translation has a relatively smaller absolute timing extension than a 25-word sentence that expands to 30 words. The dubbing-ready script prefers two 10-word sentences over one 20-word sentence when both serve the scene equally well.
The phoneme-transparent ending. Lines that end on open vowel sounds, which are visible in the actor's mouth position at the line's conclusion, are harder to synchronise in dubbing than lines that end on consonants or closed vowel sounds. The dubbing-ready script avoids line endings that produce a distinctive visible mouth position that will not match the dubbed language's translated line ending.
The writer brief for a localisation-first production includes these three constraints alongside the standard vertical drama format requirements. They do not significantly constrain the creative work. The 80% rule reduces line length by 20%, which is within normal writing variation. The short sentence preference affects style but not content. The phoneme-transparent ending affects approximately one in ten lines.
Pre-Production Decision 2: Audio Stem Discipline
The audio stem is the most critical localisation infrastructure decision in the entire production pipeline. A production that delivers a combined stereo audio mix has no localisation pathway without stem separation. A production that delivers separated stems can localise immediately without any audio remediation.
The stems that a localisation-first production records and maintains separately throughout the post-production pipeline:
The dialogue stem. The clean dialogue track containing only the performers' speech, with no music or effects layered on top. This stem is the training source for AI voice cloning in dubbing workflows that clone the original performer's voice characteristics into the target language. It is also the reference for synchronisation: the dubbing voice actor or AI dubbing system aligns the translated dialogue to the timing markers embedded in the dialogue stem.
The music stem. The music bed track containing only the series' score, with no dialogue or effects. The music stem passes unchanged into every localised version. Language localisation does not require new music; it requires replacing the dialogue over the same music.
The effects stem. The sound effects and ambient audio track containing all non-dialogue, non-music audio content. Environmental sounds, footsteps, and practical sound effects pass unchanged into every localised version.
The M&E stem. The music and effects combined stem, without dialogue. This is the standard deliverable for international distribution of live-action content and is the stem that AI dubbing systems use to produce the final localised mix. The M&E stem plus the dubbed dialogue produces the complete localised audio without requiring any additional post-production work on the music or effects.
The specific discipline required: every episode's audio post-production must be completed with these stems maintained as separate files throughout the mixing process. Stems created from a combined mix through stem separation AI are degraded versions of the original clean stems. Stems maintained separately throughout the mixing process are the original quality assets.
ElevenLabs' Professional Voice Cloning requires 30-plus minutes of clean source audio for the highest-quality voice clone. The dialogue stem, which contains 30-plus minutes of clean recorded dialogue across the full series, is precisely the source audio the professional voice cloning workflow requires. A production that maintains clean dialogue stems produces better AI dubbing quality than a production that attempts to separate dialogue from a combined mix.
Pre-Production Decision 3: Generation Workflow for Multi-Language Mouth Movements
For AI-native productions where the characters' mouth movements are generated rather than recorded, the multi-language localisation requirement creates a specific generation workflow challenge: the mouth movements generated for the English dialogue do not match the mouth movements that the same dialogue in Spanish, Hindi, or Portuguese would produce.
English dialogue produces English phoneme patterns. Spanish dialogue produces Spanish phoneme patterns. The AI-generated character whose mouth was generated to match English phoneme timing looks unconvincing when Spanish-language audio is placed over it, because the Spanish phoneme patterns require different mouth positions and timings from the English originals.
The generation workflow solution for localisation-first AI-native production:
Option 1: Phoneme-neutral generation. Generate the character's mouth movements in a partially neutral position that is compatible with multiple phoneme systems rather than specifically matched to English phonemes. This approach produces mouth movements that are slightly less precisely matched to the English dialogue but are significantly more compatible with dubbed language versions. The trade-off is acceptable when localisation into three or more language markets is planned from production outset.
Option 2: Language-specific generation. Generate separate versions of dialogue-heavy scenes with mouth movements specifically matched to each target language's phoneme patterns. This approach produces the highest quality localised output but requires generation credits for each language version's mouth movement generation. The total generation cost is proportionate to the number of language versions planned.
Option 3: AI lip-sync correction in post. Generate the primary version with English phoneme-matched mouth movements and apply AI lip-sync correction tools to each dubbed language version in post-production. Tools including HeyGen and Rask AI apply subtle mouth movement adjustments to match the dubbed audio's phoneme patterns. This is the lowest pre-production planning requirement but produces slightly lower quality than language-specific generation.
The choice between these three options depends on the number of planned language markets and the production budget's tolerance for either quality compromise or incremental generation cost. Productions planning three or more language markets typically use option 1 or option 2. Productions planning one to two language markets typically use option 3.
Pre-Production Decision 4: Character Name and Cultural Reference Neutrality
Script content that is specific to English-language cultural references creates localisation challenges that character count discipline and audio stem discipline cannot address. A character name that is naturally pronounceable in English may be unnatural in Hindi or Portuguese. A cultural reference that is immediately recognisable to an English-speaking audience may be unknown to a Hindi-speaking audience.
The localisation-first script treats character names and cultural references with neutrality as a default. Character names are selected to be pronounceable across the target language markets, avoiding English-specific phoneme combinations that the target languages cannot accommodate. Cultural references that are specific to English-language cultural contexts are replaced with references that are either universally recognisable or that localise cleanly into each target market's equivalent.
This decision is made at the writer brief stage, before any scripts are written. The writer brief specifies the target language markets for the series and instructs the writers to evaluate each character name and cultural reference against those markets' accessibility. Names and references that fail this evaluation are flagged for revision before the script is approved for production.
The Day-One Localisation Checklist
The pre-production decisions that enable day-one localisation, confirmed before any scripting, generation, or recording begins:
Script:
Writer brief includes target language markets
80% rule applied to all dialogue
Short sentence preference specified
Phoneme-transparent ending guidance included
Character name pronounceability confirmed across target languages
Cultural reference universality confirmed or market-specific variants planned
Audio:
Dialogue stem recorded separately from music and effects
Music stem recorded separately
Effects stem recorded separately
M&E stem maintained throughout post-production
Clean dialogue stem archived for AI voice clone training
Generation workflow:
Target language count determines phoneme-neutral vs language-specific generation approach
Character reference packs tested for mouth movement compatibility with target language phoneme patterns
Generation brief for dialogue scenes includes phoneme approach specification
Rights:
Platform agreement includes territory-specific exclusivity rather than worldwide
Dubbing rights confirmed in talent agreements for all live-action performers
Music licensing confirmed to cover commercial distribution in all target territory markets
The Cost Comparison: Day-One vs Retrofit
The production that builds localisation into day one and the production that retrofits localisation after delivery produce localised content at different total costs.
Day-one localisation additional pre-production cost: Writer brief adjustment, minimal additional scripting time for the 80% rule application, and stem discipline during recording and mixing. Total additional cost: $500 to $2,000 in additional production time.
AI dubbing cost per language at delivery: $130 to $450 per series per language using current AI dubbing tools at production-grade quality. Ten languages: $1,300 to $4,500.
Retrofit localisation additional remediation cost: AI stem separation on combined mix at degraded quality, script reformatting for dubbing, and lip-sync correction generation for AI-native productions. Total remediation cost per series: $2,000 to $8,000 before the AI dubbing costs are incurred.
The day-one localisation approach's total cost advantage over retrofit: $2,000 to $8,000 in avoided remediation costs per series, in addition to the higher quality output that clean stems and dubbing-formatted scripts produce.
Axis AI Studios Perspective
Localisation built into day one is the production discipline that converts a series from a single-market asset into a multi-market asset at minimal incremental cost. The decisions that enable this are pre-production decisions that cost almost nothing to make correctly and significantly more to remediate if made incorrectly.
At Axis AI Studios, the localisation checklist is reviewed before any production is commissioned. The writer brief specifies the target language markets. The stem discipline is confirmed with the audio post-production operator before mixing begins. The generation workflow for AI-native productions is selected based on the number of planned language markets. These are not production challenges that arise at delivery. They are production specifications that are set before the first script page is written.
For production companies who want to commission vertical drama content built for multi-language delivery from day one rather than localised as an afterthought, reach out at business@axisaistudios.com.
FAQ
How Many Languages Should a Production Plan For From Day One?
The practical answer depends on the production company's distribution strategy and the series' genre. A CEO romance series targeting the US primary market and Spanish-language secondary markets should plan for two languages from day one. A supernatural fantasy series with global distribution potential should plan for five to ten languages. The pre-production cost of planning for ten languages is minimal. The post-production cost of retrofitting ten-language capability on a production not built for it is significant. When in doubt, build for more languages than currently planned.
Does the 80% Rule Reduce Dialogue Quality?
Not materially. The 80% rule reduces line length by approximately 20%, which is within the natural range of dialogue variation that good writing produces. A writer who is unaware of the 80% rule and producing natural-sounding dialogue is frequently already at or near 80% of the maximum delivery pace. The explicit constraint is a reminder rather than a significant restriction for experienced writers who understand economical dialogue.
Can AI Stem Separation Produce Acceptable Results on a Combined Mix?
AI stem separation tools including Spleeter and Demucs produce usable stems from combined mixes, but at degraded quality compared to original clean stems. The degradation is most audible in the dialogue stem, where AI separation may retain residual music or effects frequencies that affect the voice clone training quality. For a production where localisation quality is commercially significant, investing in the original stem discipline during recording is preferable to relying on AI separation at the localisation stage.
Further Reading
For the AI dubbing tools and quality benchmarks that the localisation pipeline described in this post delivers into, the guide to using AI to localise your vertical drama catalog into 10 language markets covers ElevenLabs, Resemble AI, and the full AI dubbing workflow.
For the AI voice cloning workflow that the clean dialogue stems described in this post are the primary training source for, the guide to AI voice cloning for vertical drama ADR covers the voice clone training process, quality evaluation, and consent framework.
For the rights structuring that must be confirmed before localisation into new territory markets can proceed, the guide to negotiating your first platform deal covers territory-specific exclusivity and the provisions that enable multi-market distribution.

Let's set
the new standard together.
If you're working on something, we'd like to hear about it.
