How to Split Audio Stems During Production So a Vertical Drama Series Can Be Localised Without a Remix
Episode twelve of seventy. A Spanish version ordered off the back of a territory that performed. The audio exists as one stereo mix per episode, with dialogue, music, ambience and effects already bounced together. Replacing the dialogue now means re-deriving everything underneath it, which means rebuilding seventy mixes from session files that may or may not still open cleanly, and in practice means hiring a mix to be done twice. The series was always going to be localised. Nobody structured the audio for it.
This is the most reliably avoidable cost in vertical drama post production, and it is avoidable for a specific reason: the work required to prevent it happens during production, costs almost nothing at that point, and becomes impossible to retrofit once mixes are bounced and sessions are archived. What follows is the stem structure, naming scheme, mixing discipline and validation pass that turn a language swap into a replacement of one layer rather than a rebuild of all of them.
1. Why the Remix Happens
The remix happens because the deliverable the platform asks for is a finished stereo mix, and a production that builds only what is asked for builds only that. Nothing in the specification requires stems, so the session files that could produce them later get treated as working files rather than deliverables. Working files get archived carelessly, reference plugins no longer licensed, and depend on a project version nobody recorded.
The second reason is sequencing. Localisation arrives as a decision made after performance data exists, long after the audio was finished. By then the operator who built the mix has moved on, the conventions were never written down, and the only durable artefact is the bounced file. The fix is not better session archiving. It is producing the stem set as a first class deliverable alongside the mix, so the localisable version exists whether or not localisation is ever ordered.
2. The Minimum Viable Stem Split
Stem mixing and mastering in music production can run to a dozen or more buses, and importing that level of granularity into a seventy episode vertical drama pipeline is a mistake. Every additional stem multiplies by episode count and becomes a file management and quality control burden that the localisation benefit does not justify. The useful split is four stems and it should stay at four unless a specific series requirement forces a fifth.
Dialogue carries every line of spoken performance and nothing else. Music carries score, songs and any stings or transitions. Effects carries hard effects tied to on screen action: impacts, doors, vehicles, phone rings, anything that has a visible cause. Ambience carries room tone, exteriors, crowd beds and anything that establishes a space rather than an event. The reason this particular four way split works is that exactly one of them, dialogue, gets replaced in a localisation, and the other three sum to a bed that can be reused untouched. Everything beyond four stems adds management cost without adding localisation flexibility.
The fifth stem, where a series needs one, is a separated crowd layer. Background dialogue in the production language is dialogue for localisation purposes but behaves like ambience in the mix, and pulling it out is the difference between localising a crowded series and finding half the atmosphere in the wrong language.
3. Where Dialogue Has to Be Clean
A dialogue stem is only useful if it is genuinely isolated, and the most common defect is a dialogue stem with processing baked into it that belongs to the scene rather than to the voice. Reverb is the main offender. A line spoken in a stairwell needs stairwell reverb, and if that reverb is printed onto the dialogue stem then the replacement line arrives dry into a mix that has no reverb available to put it into. The discipline is to keep spatial processing on a send that renders into the effects or ambience stem, so the replacement dialogue can be routed through the same treatment.
The second defect is dialogue that has been ducked by a sidechain from the music or effects bus, with the ducking printed into the dialogue rather than applied to the thing being ducked. A localised line with different timing hits the printed ducking in the wrong places, which produces a mix that pumps against speech that is no longer there. The rule is that any processing whose shape depends on the dialogue timing belongs on the stem being affected, not on the dialogue stem itself.
Corrective processing is the exception and belongs on the dialogue stem. Noise reduction, de-essing, broadband cleanup and level matching are properties of the recording rather than the scene, so printing those is correct. Printing anything that describes the space is not.
4. Music and the Timing Problem
Music is the stem that looks easiest and causes the most trouble, because music in vertical drama is frequently cut to dialogue rather than to picture. A sting that lands on the last syllable of a line in the production language lands somewhere else when the line is a different length, and a stem that is correct as a bounce can still be wrong as a localisation input. The structural fix is to cut music to picture wherever the choice exists, and to log every place where it was deliberately cut to dialogue.
That log is the actual deliverable. A short list per episode naming which music events are dialogue locked, with timecodes, converts an invisible problem into a worklist. Without it, somebody finds the mistiming by listening to all seventy episodes. With it, the localisation pass knows which cues need nudging and can price that before starting.
Songs with lyrics in the production language are dialogue in everything but routing. A series that intends to localise should keep them instrumental, license them in a form allowing a localised vocal, or record the decision to leave them original. Any of the three is fine. Meeting the question for the first time during a localisation is not.
5. Effects, Ambience and the Layer That Gets Forgotten
Effects and ambience are the two stems that should survive localisation entirely untouched, and they do survive if the dialogue was never allowed to leak into them. The leak path is almost always a production decision rather than a mix decision: a generated clip that arrives with vocal content embedded in its audio, or a sourced ambience bed that contains intelligible speech in one language. Both are invisible in a stereo mix and both become obvious the moment the dialogue stem is replaced.
The screening discipline is a listening pass on ambience and effects sources at the point of selection, specifically for intelligible speech. Crowd beds are the main risk, because a usable crowd bed often carries a few phrases that register as language rather than texture. A bed that works in one language and nowhere else is a liability on a series with any localisation ambition, and the replacement is cheap at selection time and expensive at episode sixty.
Generated effect audio needs the same screen for a different reason. Audio arriving with a generated clip is frequently tied to that clip and not reproducible, so if it becomes the only source for an important effect there is no way to rebuild it at a different balance. Rendering it into the effects stem at the point of use, with the source retained separately, is what keeps it usable.
6. Naming and Delivery Structure
Stems are only useful if a localisation vendor who has never spoken to the production can open the delivery and understand it without asking. The naming scheme should carry the series identifier, episode number, stem type and version, in that order, with fixed field widths so that sorting works: a two digit episode on a seventy episode series sorts correctly and a one digit episode does not. Stem type should use a fixed vocabulary of four or five tokens rather than free text, because free text drifts across an episode run and drifted names defeat batch processing.
Sample rate, bit depth and channel configuration have to be identical across every stem of every episode. This sounds too obvious to state and it is the single most common defect in delivered stem sets, because stems get bounced by different operators at different points and a settings difference survives undetected into delivery. The configuration belongs in the convention document and the check belongs in the automated pass described below.
Alongside the stems, a short per episode note should record the dialogue locked music events, any intentional language retention, and the loudness target. Three facts, one file per episode, and it is what stops a localisation pass rediscovering production decisions by inference.
7. Mixing Decisions That Survive a Language Change
Some mix decisions are language independent and some are not, and the ones that are not should be made with the knowledge that they will be redone. Dialogue level relative to the bed is language dependent, because languages differ in dynamic range and intelligibility at a given level, and a bed sitting comfortably under one language can mask another. The practical consequence is to mix the bed so that it has headroom to be pulled rather than mixing it as tight as it can sit.
Pacing is the other language dependent decision. A localised line is frequently longer than the original, and a mix with no space around dialogue gives a localisation pass nowhere to put the extra syllables. Leaving a small amount of air at the head and tail of dialogue events costs nothing in the original and is the difference between a clean replacement and a reworked edit. On a format where episodes run tight by design, this is a deliberate choice rather than an accident.
Software practice in internationalization and localization calls this externalising the locale dependent parts. The audio version is the same: anything that will change goes in one replaceable layer, everything that will not sits underneath. A mix built that way needs one layer swapped rather than a redesign per territory.
8. The Phone Speaker Constraint Applies to Every Stem
Vertical drama is consumed on phone speakers, and the stems have to be built for that target rather than for a monitoring environment that no viewer has. This matters for stem splitting specifically because the frequency region where dialogue has to stay intelligible on a small speaker is the same region where ambience and music beds want to sit. A bed that is transparent on monitors can bury a replacement line on a phone, and the defect appears only after localisation if the bed was mixed against the original dialogue alone.
The discipline is to carve the bed against the region rather than against the specific dialogue. If the midrange is kept clear structurally, any dialogue stem dropped into the bed will read, regardless of language or voice. If the midrange was cleared by riding the bed against one particular performance, the clearance is specific to that performance and does not transfer. The first approach survives localisation and the second does not.
Loudness targets need stating per stem as well as per mix. A localisation pass receiving stems with no stated target will normalise to its own assumption, and a consistently targeted series can come back inconsistent across the run. One line in the per episode note prevents it.
9. Validating the Split Before the Series Scales
The validation pass belongs at episode three, not episode seventy, and it is a mechanical test rather than a review. Take one episode, mute the dialogue stem and listen to the remaining three summed. Nothing intelligible should be audible, no ducking artefacts should pump, and no reverb tail should be missing. Then drop a placeholder dialogue recording of deliberately different length into the dialogue stem and listen again. Any mistiming, masking or spatial mismatch that appears is a structural defect in the split rather than a problem with the placeholder.
The second half of the pass is automated and checks the mechanics: that every episode has the full stem set, that sample rate, bit depth and channel configuration match across all of them, that the naming scheme parses, that the per episode note exists and that the four stems sum to within a defined tolerance of the delivered stereo mix. That last check is the important one, because a stem set that does not sum to the mix means something in the mix exists outside the stems, which is the exact failure the whole structure is meant to prevent.
Running this at episode three costs an hour and catches conventions that would otherwise be wrong on every subsequent episode. Run at the end of a seventy episode series, the same pass finds those defects multiplied by seventy with no team still assembled to fix them.
Axis AI Studios Perspective
Axis AI Studios is an AI native vertical drama production studio in the Netherlands, producing series for clients including Den Tolmor and Good Fight Production LLC, and HolyWater. Stem discipline is production side practice and it is the kind of claim worth being precise about: this is work Axis controls directly, specifies in its own convention documents and runs on its own productions. The four stem split, the dialogue locked music log, the ambience screening pass for intelligible speech and the episode three validation test are standing practice rather than aspirations.
The reason to treat this as a production decision rather than a post production one is structural. Every element described here is cheap at the moment it is made and impossible to recover later. Screening a crowd bed takes a minute at selection and cannot be done after it is printed into forty episodes. Keeping reverb on a send rather than printing it is a routing choice with no cost attached. Logging the dialogue locked cues is a line per event. None of it requires a decision about whether the series will ever be localised, which is the point: the structure costs so little that building it unconditionally is cheaper than deciding.
What a platform or an IP holder should take from this is a question to ask before a series starts rather than after a territory performs. Ask whether the production is delivering stems alongside the mix, what the split is, and whether the stem set has been validated against the mix. A production that has an answer has already made the localisation cheap. A production that has to go and find out has probably not. Production enquiries, including stem specifications for a series already in progress, go to business@axisaistudios.com.
FAQ
Can stems be extracted from a finished mix instead?
Source separation tools have improved considerably and can produce a usable dialogue isolation from a stereo mix, but the result is not equivalent to a real stem set. Separation leaves artefacts in both the extracted dialogue and the residual bed, and the residual bed is the part that matters, because that is what the replacement dialogue sits in. It is a reasonable rescue for a small number of episodes where no stems exist. It is not a substitute for building the split during production on a long series, where the artefacts accumulate across the run and the quality floor is set by the worst separation rather than the best.
Does keeping four stems per episode create a storage problem at seventy episodes?
Not meaningfully. Four stems at the same configuration as the mix is roughly four times the audio volume of the mix alone, and audio is a small fraction of total deliverable size on a series carrying seventy episodes of video. The real cost of stems is management rather than storage: more files to name consistently, verify and keep in sync. That is exactly why the split stays at four rather than expanding, and why the automated completeness and configuration check is part of the structure rather than optional.
What if localisation is genuinely never going to happen?
Build the split anyway, because the localisation case is not the only thing it buys. A stem set makes a trailer cut possible without returning to the mix session, lets a platform request a dialogue level adjustment without a full remix, supports a clean version for territories with content restrictions, and survives the loss of the original session files entirely. The structure also makes a sequel cheaper, because the conventions and the bed treatment carry forward. The localisation benefit is the largest single one, and the structure pays for itself without it.
Further Reading
For the mixing decisions underneath the stem structure, particularly how the midrange has to be managed for the device the format is actually watched on, the guide to mixing audio for phone speakers covers loudness targets and intelligibility in detail.
For the wider version of the same principle applied across script, generation and delivery rather than audio alone, the piece on what localisation built into production from day one actually looks like sets out which decisions have to be made before production starts.
For where stem work sits in the full post production sequence alongside colour and visual effects, the overview of vertical drama post production covering sound design, colour and visual effects covers how the stages hand off to each other.

Let's set
the new standard together.
If you're working on something, we'd like to hear about it.
