How to Brief a Voice Actor for AI Vertical Drama ADR: The Direction Format That Produces Usable Performances in One Session

A voice director working on vertical drama needs to brief actors differently. The actor is not just matching lip movement. The actor is compensating for a visual medium that has traded emotional depth for temporal efficiency. That framing changes how an actor approaches a line: not as a line in a scene, but as a scene compressed into a line.

This distinction is the most important thing a voice actor needs to understand before an AI vertical drama ADR session begins. And it is the thing that most production companies fail to communicate before the session starts, resulting in performances that are technically clean but emotionally insufficient for the format's paywall conversion requirements.

The AI vertical drama ADR session has two variants. The first is dialogue replacement for scenes where the AI-generated audio is technically unusable — mouth movement artifacts, audio sync drift, or dialogue that was generated in a different language from the delivery target. The second is AI voice cloning and re-performance, where the production uses an AI voice model trained on a brief human performance session to produce the character's dialogue across the full episode run. Both variants require the actor to understand the format's specific performance register before they produce any audio.

This post covers the complete direction brief format for both ADR variants: what the actor needs to know before the session begins, how the scene-level direction brief is structured, and what the single-session discipline looks like that produces usable performances without multiple session callbacks.

What the Actor Needs to Know Before the Session

Brief 1: The Format's Emotional Compression

The actor who enters a vertical drama ADR session with their conventional television or film performance instincts will produce performances that are too slow, too textured, and too resolved for the format.

Vertical drama's episode runs 90 seconds. A dialogue-heavy scene within a 90-second episode has eight to fifteen lines of dialogue. Each line carries significantly more emotional information than a single line in a 60-minute episode because there are so few lines available to carry the scene's full arc. The actor cannot spend the first three lines establishing the character's emotional state. The emotional state must be present from the first word of the first line.

The specific instruction for the actor before the session begins: in this format, every line is the climactic line. There is no warm-up within a scene. The character arrives at the scene already at the emotional intensity the scene requires. The actor's job is to produce that intensity immediately, sustain it across each line, and cut it cleanly at the end of each recording rather than trailing into resolution.

Brief 2: The Character's Psychological Configuration

The actor who performs ADR for an AI-generated character needs the character's psychological configuration rather than a character description. A character description tells the actor what the character is. The psychological configuration tells the actor what the character is suppressing, how they suppress it, and what breaks through the suppression against their intention.

For the controlled alpha: the character's exterior is precise, maintained, and unrevealing. The dialogue delivery matches this exterior precisely — controlled pace, level register, no vocal variation that reveals the interior state. The emotional intensity comes from the audience's understanding that the interior state is being suppressed, not from the vocal delivery expressing it.

This is the opposite of what most theatrical training teaches: show the emotion. In vertical drama's controlled alpha character, the actor shows the suppression, not the emotion. The emotion is visible in the visual medium's close-up. The audio supports the suppression, not the revelation.

For the protagonist: the character's exterior is determined and strategically controlled, but less tightly than the alpha's. The protagonist's vocal delivery has slightly more texture and variation — not emotion, but strategic calculation visible in the delivery.

Brief 3: The Lip-Sync Standard

In conventional ADR, the actor watches the scene and synchronises their dialogue delivery to their own lip movement. In AI vertical drama ADR, the actor is synchronising to AI-generated lip movement that was produced from a generated audio track rather than from their own prior performance.

The AI-generated lip movement may not precisely match the target language's phoneme timing, particularly for sounds that have different mouth positions across languages. The actor should be briefed that approximate sync is the correct target, not perfect sync. The AI-generated visuals will be lip-sync adjusted in post-production using tools like the Jeynix lip-sync pipeline if perfect sync is required for the hero scenes. For standard dialogue scenes, the actor's best sync attempt from the performance brief is the performance standard.

The Scene-Level Direction Brief

Each scene in the ADR session requires a scene-level direction brief that specifies four elements: the emotional register, the subtext, the delivery tempo, and the end-of-line treatment.

Emotional register: Not an emotion label. A physical state description. "Jaw level, controlled, eyes forward, minimal vocal variation" is a physical state description. "Angry but hiding it" is an emotion label that the actor will interpret differently from what the production needs.

Subtext: What the character means rather than what they are saying. The controlled alpha who says "I'll have the report by tomorrow" means "I am aware of you in a way I am refusing to acknowledge and this sentence is an excuse to speak to you." The actor who delivers the line knowing the subtext produces a different performance from the actor who delivers the line as a logistics communication.

Delivery tempo: Vertical drama's dialogue tempo is faster than television drama. The beats between lines are shorter. The internal rhythm of each line is slightly clipped rather than fully breathed. The actor who has been told the delivery tempo before the session starts can calibrate from the first take rather than discovering through failed takes that their natural tempo is too slow.

End-of-line treatment: Every line should end with a clean cut rather than a trailing resolution. The actor cuts the delivery at the end of the line's final syllable. No trailing breath. No resolving inflection. The line ends. This is the single most common ADR correction in vertical drama post-production: the actor who trails their delivery into a small resolution at the end of each line produces audio that must be trimmed in post for every line rather than used as delivered.

The Single-Session Production Structure

The ADR session that produces usable performances in one session follows a specific structure that most conventional ADR sessions do not use.

Session opening — 15 minutes: The direction brief for the format, the character's psychological configuration, and the delivery tempo are communicated before any recording begins. The actor has no script in front of them during this brief. They are listening and asking questions. The questions that actors consistently ask in this session: "So I am playing the suppression rather than the emotion?" Yes. "And the line ends cleanly with no resolution?" Yes. "Even for the most emotional lines?" Especially for those.

Warm-up takes — 10 minutes: The first three lines of the session are performed twice without recording. The first performance is the actor's instinctive interpretation. The second performance is the corrected interpretation after specific direction on what changed between the first and second. This calibration session occurs before recording begins because the calibration notes from a recorded take are more useful than flagging the take as a callback — the actor hears the direction and immediately applies it rather than returning for a second session.

Recording block structure: Record in scene blocks rather than line by line. A scene block is all the lines for one character in one scene. Recording the full scene block before moving to the next scene maintains the character's emotional state across the full block rather than resetting it between individual line records.

The three-take rule: Record a maximum of three takes per line before moving on. The first take is the actor's calibrated interpretation. The second take addresses a specific note from the first. The third take addresses a specific note from the second. If the third take has not produced a usable performance, the production team needs to evaluate whether the scene-level direction brief is at fault rather than the actor's performance. A fourth take on the same note produces diminishing returns.

The AI Voice Cloning Session Variant

For productions using AI voice cloning rather than scene-by-scene ADR, the session structure changes significantly. The actor is not performing dialogue for specific scenes. They are providing a performance sample that the AI voice model is trained on.

The performance sample session requires the actor to produce the character's full emotional range across a structured script of 30 to 50 lines that cover the character's emotional register spectrum: the controlled neutral state, the involuntary recognition state, the authority confrontation state, the unearned vulnerability state. Each state is performed at the level of emotional intensity the format requires, with the clean line-end treatment that prevents the training data from including trailing resolution artifacts.

The AI voice model trained on a correctly structured performance sample produces dialogue delivery that is consistent with the character's psychological configuration across the full episode run without requiring the actor to be available for individual scene ADR. The production efficiency of this approach is significant: one four-hour session produces a voice model that covers 70 episodes rather than the multiple session days that scene-by-scene ADR requires.

The quality consideration: AI voice models trained on brief samples produce dialogue delivery that captures the character's vocal register and general performance style. The paywall episode's most emotionally precise moments, where the actor's specific involuntary performance quality is the primary commercial asset, benefit from human ADR rather than from AI voice model delivery. The production routing logic: AI voice model for standard dialogue across the episode run, human ADR for the paywall episode and two preceding episodes.

Axis AI Studios Perspective

The ADR session that produces usable performances in one session is the session that begins with the direction brief format described in this post. The actors who consistently deliver in one session are the actors who understood the format's emotional compression standard, the character's psychological configuration, and the delivery tempo before they recorded their first line.

At Axis AI Studios, the ADR direction brief is prepared alongside the direction brief for generation. The character's psychological configuration that the generation operator uses to produce the visual performance is the same configuration the ADR actor uses to produce the audio performance. Visual and audio are calibrated to the same character specification from the same brief.

For experienced voice actors who want to work on AI vertical drama productions and understand what the ADR direction format requires, the application process starts at business@axisaistudios.com. For businesses commissioning AI-native vertical drama who want to understand how ADR fits into the production workflow and timeline, the same address is the starting point.


FAQ

How Long Does a Standard ADR Session Take for a 70-Episode Series?

A 70-episode series contains approximately 800 to 1,200 individual dialogue lines across all primary characters. A single ADR session covering one primary character at the production standard described in this post covers 60 to 80 lines in a four-hour session including warm-up and direction time. A full primary character's dialogue across 70 episodes requires two to three four-hour ADR sessions. The AI voice cloning alternative covers the same scope in one four-hour session with AI delivery for standard scenes and targeted human ADR for the paywall episode.

What Qualifications Should a Voice Actor Have for AI Vertical Drama ADR?

Prior ADR experience is more valuable than prior vertical drama on-camera experience for the voice actor role. The specific skills required are: the ability to sync to pre-recorded visual content without prior rehearsal, the ability to deliver clean line-end treatments without trailing resolution, and the ability to calibrate emotional register from a brief description rather than from a full script reading. Voice actors with animation dubbing experience have the most directly applicable skill set because animation dubbing requires the same format-specific delivery calibration without the context of their own prior on-screen performance.

Can AI Voice Cloning Replace Human ADR Entirely?

For standard dialogue scenes across the middle arc of a 70-episode series, AI voice cloning at the current quality level (mid-2026) produces results that meet the platform's distribution quality standard. For the paywall episode and the two preceding episodes, and for any scene where the character's involuntary emotional precision is the primary commercial asset, human ADR produces performance quality that AI voice cloning cannot consistently replicate. The correct production approach is hybrid: AI voice model for volume, human ADR for commercial-critical scenes.


Further Reading

For the AI voice cloning capabilities and limitations that determine when human ADR is required, the guide to AI voice cloning for vertical drama ADR covers the current capability ceiling, the production routing logic, and the consent and ethics provisions that govern AI voice cloning in production.

For the audio mixing that processes the ADR output for phone speaker delivery, the guide to mixing audio for phone speakers covers the LUFS targets, frequency response limits, and dialogue priority specification that the ADR track is mixed to.

For the character's psychological configuration that the direction brief is built from, the guide to how to build a vertical drama character bible for AI generation covers the suppressed interior, the performed exterior, and the direction brief shorthand that the ADR direction brief derives from.

Stay connected

For studios moving beyond traditional production.

Let's set
the new standard together.

If you're working on something, we'd like to hear about it.