How to Write a Vertical Drama Direction Brief That a Generation Operator Can Execute Without Clarification

The direction brief is the document the generation operator reads before the first session begins. It specifies every parameter the operator needs to generate the episode: the character configuration, the scene type, the camera position, the emotional register, the dialogue sync requirement, and the visual register. A direction brief that specifies all of these correctly produces a session where the operator generates outputs that pass the quality review without needing to ask the creative director for clarification.

A direction brief that requires clarification before or during the session has failed at its primary function. Every clarification request is a production delay: the operator stops generating, the creative director is interrupted, the batch timeline extends, and the generation session's efficiency drops. At 70 episodes across multiple operators and multiple sessions, the cumulative cost of direction briefs that require clarification is significant.

This post covers the complete direction brief format. Every section. What each section specifies, why that specification matters for generation quality, and what the difference looks like between a brief that an operator can execute and one that requires follow-up.

The Direction Brief Format

A complete direction brief for an AI vertical drama episode contains seven sections: episode context, character configuration, scene type specifications, camera geometry, emotional register, dialogue sync, and visual register reference. Each section is specific enough that the operator can begin generation without asking a single question.

Section 1: Episode Context

What it specifies: The episode number, the arc position, and the narrative function of this episode within the 70-episode structure.

Why it matters: The operator who knows that episode 22 is the midpoint reversal episode where the antagonist's power is temporarily extended knows that the scene energy is different from episode 10's escalation. The operator who only knows "episode 22" has no structural context and must infer the emotional tone from the dialogue text alone. Inference produces inconsistency.

What correct specification looks like:

Episode 22. Arc position: midpoint reversal. Narrative function: the protagonist's public attempt at power inversion fails. The antagonist publicly reasserts institutional authority in front of the social witnesses who matter most to the protagonist. The protagonist leaves the scene with their position materially worse than at the episode's start. The episode ends at the protagonist's lowest point before the second act escalation begins.

What incorrect specification looks like:

Episode 22. Office confrontation scene.

The incorrect version gives the operator a setting and a scene type. It gives no structural context, no emotional trajectory, and no information about where this episode sits in the arc. The operator generates an office confrontation. Whether it captures the midpoint reversal's specific emotional register is left to chance.

Section 2: Character Configuration

What it specifies: The approved reference frames for each character at this arc position, the Soul ID model identifier, and the wardrobe state.

Why it matters: The character configuration section is the most mechanically critical section of the direction brief. An operator who loads the wrong reference frame or uses the wrong Soul ID model identifier generates outputs that will fail the character consistency check regardless of how well they execute every other brief parameter.

What correct specification looks like:

Alpha character: reference frame ALPHA_ARC2_2026-04-28.jpg. Soul ID model: [specific identifier]. Wardrobe state: dark charcoal suit, white shirt, no tie. Tie removal signals deliberate informality — a choice the character makes when they want to project casual authority rather than institutional authority.

Protagonist character: reference frame PROTAGONIST_ARC2_2026-04-28.jpg. Soul ID model: [specific identifier]. Wardrobe state: the blouse from ARC1 episodes 15 through 20. The wardrobe continuity with the pre-midpoint episodes signals that the protagonist has not yet made the wardrobe shift that marks their second-arc repositioning.

What incorrect specification looks like:

Use the standard character references. Formal wardrobe.

The incorrect version requires the operator to identify which reference frames are "standard" for arc 2 (not specified), which wardrobe from the formal wardrobe category to use (not specified), and whether the antagonist's wardrobe state carries any narrative significance (not specified). Each ambiguity is a potential generation error.

Section 3: Scene Type Specifications

What it specifies: The generation tool routing for each scene type in the episode, the batch size per scene type, and the session-open consistency check requirement.

Why it matters: Different scene types require different generation tools. Dialogue close-up scenes route to Seedance 2.5 for character reference integration. Action and physical movement scenes route to Kling 3.0 for motion quality. Hero shots at the episode's commercially critical moments route to Veo 3.1 for composition precision. An operator who generates all scene types with the same tool is not applying the tool stack correctly and will produce outputs that fail the quality criteria for one or more scene types.

What correct specification looks like:

Scene type routing for episode 22:

Dialogue close-up scenes (12 outputs): Seedance 2.5 API. Standard reference conditioning strength. Batch size: 4 outputs per session.

Button cut hero shot (1 output): Veo 3.1 frame-mode control. Maximum resolution. This is the episode's most commercially critical output — the final frame that drives the viewer's paywall unlock decision. Do not route to Seedance 2.5.

Session-open consistency check required before any production generation begins. Compare test output against ALPHA_ARC2_2026-04-28.jpg. Do not proceed if visible deviation is present.

What incorrect specification looks like:

Generate the dialogue scenes and the final shot. Use standard settings.

The incorrect version gives the operator no tool routing, no resolution guidance, and no session-open check requirement. The operator defaults to their preferred tool and standard settings. Whether those defaults match the brief's requirements is unknown until the quality review reveals a problem.

Section 4: Camera Geometry

What it specifies: The ControlNet reference for each scene's camera position and the framing parameters for each shot type.

Why it matters: Camera geometry consistency is what makes a 70-episode series look like it was produced by the same creative vision rather than by 70 different operators who each chose their own framing. The ControlNet reference encodes the exact camera geometry — the character's position in frame, the head height, the eye line, the background depth — and ensures every operator reproduces the same compositional structure.

What correct specification looks like:

Authority close-up (antagonist dialogue lines): ControlNet reference AUTHORITY_CU_ARC2.png. Character positioned in the right two-thirds of frame. Head height: upper quarter of frame. Eye line: slightly above camera axis — looking slightly down at the camera's implied viewer position. Background: dark, out-of-focus depth at 18 inches behind the character plane.

Proximity close-up (scenes where both characters are in the same frame): ControlNet reference PROXIMITY_CU_ARC2.png. Characters positioned at frame's left and right thirds. Neither character fully centred. The tension between the two characters expressed through the framing's division of the frame rather than through expression alone.

What incorrect specification looks like:

Standard close-up shots. Keep it professional.

The incorrect version gives the operator no compositional reference, no eye line guidance, and no background depth specification. "Professional" is not a ControlNet reference. Every operator will produce a different interpretation.

Section 5: Emotional Register

What it specifies: The psychological state each character is performing in each scene, described as physical state rather than emotion label.

Why it matters: Emotion labels — "angry," "sad," "jealous" — are interpreted differently by different operators. A physical state description — "jaw level, controlled, minimal facial variation, eyes forward without directional focus" — is interpreted the same way by every operator because it describes what the output should look like rather than what the character is feeling.

What correct specification looks like:

Alpha character emotional register for episode 22: controlled authority. Physical state: jaw level, controlled, minimal vocal or facial variation. The character is performing certainty about the outcome — not aggression, not satisfaction, not warmth. The performance is the absence of visible effort rather than the presence of visible emotion. The alpha does not need to perform dominance. The situation performs it for them.

Protagonist character emotional register for episode 22: suppressed resistance. Physical state: jaw set, eyes direct, no downward gaze. The character is performing dignity rather than capitulation — they are not yielding internally even though the external situation has gone against them. The performance difference between yielding and suppressed resistance is visible in the set of the jaw and the direction of the gaze.

What incorrect specification looks like:

The antagonist is powerful. The protagonist is upset but trying to hide it.

The incorrect version describes emotional states rather than physical states. Two operators reading "powerful" will produce two different physical performances.

Section 6: Dialogue Sync

What it specifies: The dialogue text for each scene, the sync priority (lip sync or audio-led), and the ADR requirement if any.

Why it matters: The generation operator's dialogue sync approach determines whether the output's lip movement is calibrated to the generated audio track or to a pre-recorded ADR track. This distinction must be specified in the brief — an operator who applies audio-led sync to a scene that will use human ADR has generated an output that the ADR session will produce incorrect sync against.

What correct specification looks like:

Scene 4 dialogue (antagonist): "The submission has been reviewed. It did not meet the threshold." Sync approach: audio-led. No ADR required for this scene.

Scene 9 dialogue (paywall episode button cut): no spoken dialogue. Visual-only button cut. No sync requirement. Hero shot generated for visual impact rather than dialogue delivery.

Button cut scene (episode 22 final frame): this output is the hero shot for the paywall position. ADR required — the vocal delivery at this moment is commercially critical and will be recorded in the ADR session. Generate the visual output for lip sync alignment with the ADR track. Do not generate a final audio track for this scene.

What incorrect specification looks like:

Use the dialogue from the script. Standard sync.

The incorrect version requires the operator to locate the correct script draft, identify which dialogue belongs to episode 22, and determine what "standard sync" means in the context of this episode. Each of these determinations is an opportunity for error.

Section 7: Visual Register Reference

What it specifies: The colour temperature, lighting direction, contrast ratio, and environment category for this episode's scenes.

Why it matters: The visual register maintains the series' visual identity across episodes generated in different sessions by different operators. An operator who generates episode 22 with a warmer colour temperature than the approved style guide specification is producing a visual register drift that the delivery quality review will flag as a fail.

What correct specification looks like:

Colour temperature: 4,200K (cool-neutral). Primary lighting direction: camera-left at 45 degrees. Contrast ratio: 4:1 key to fill. Background environment: corporate interior, dark surfaces, no natural light sources visible. Style guide reference: page 3, corporate interior specification. Do not use the residential interior specification from page 5 — the antagonist's power is expressed through the corporate environment, not through a domesticated space.

What incorrect specification looks like:

Dark and corporate. Professional office feel.

The incorrect version provides a mood description rather than a technical specification. Every operator will produce a different colour temperature, lighting direction, and contrast ratio in response to "dark and corporate."

Axis AI Studios Perspective

At Axis AI Studios, every generation session begins with a direction brief that contains all seven sections at the specification level described in this post. The creative director writes the brief. The operator executes the brief. Clarification requests before or during a session are a brief failure rather than an operator failure — they indicate a section that needed more specification.

For production coordinators, creative directors, and generation operators who want to develop direction brief writing skills, the direction brief is the production document with the most direct impact on session efficiency and quality review pass rates. A correctly written brief is the most cost-effective quality investment in AI vertical drama production.

For businesses commissioning AI-native vertical drama who want to understand how direction brief quality is verified across a managed production, reach out at business@axisaistudios.com.


FAQ

How Long Does It Take to Write a Complete Direction Brief?

A complete seven-section direction brief for one episode takes 45 to 90 minutes for an experienced creative director working from an approved arc map, character bible, and style guide. The brief writing time is the pre-production investment that prevents the session clarification delays that cost the same or more in operator time. A 75-minute brief that generates 12 approved outputs in one session is more efficient than a 30-minute brief that generates 8 outputs and requires 4 revision cycles.

Should Every Episode Have an Individual Direction Brief or Can Briefs Be Shared Across Episode Batches?

Episodes within the same arc position and the same scene type category can share a batch direction brief that specifies the parameters applicable to all episodes in the batch and notes episode-specific variations. A batch brief for episodes 15 through 22 specifies the Arc 2 character configuration, the camera geometry references, and the visual register — and notes which episodes within the batch have scene-specific variations in emotional register or dialogue sync requirements. The batch brief format reduces brief writing time without reducing specification precision.

What Is the Most Common Direction Brief Failure That Produces Quality Review Failures?

Section 5 (Emotional Register) is the most common direction brief failure. Creative directors who describe emotional states rather than physical states produce briefs that operators interpret inconsistently. "The character is suppressing their real feelings" is an emotion description. "Jaw set, eyes direct, minimal facial variation, no downward gaze" is a physical state description. The difference in generation output between the two specifications is visible in every output the operator produces.


Further Reading

For the character bible that provides the character configuration data in Section 2 of the direction brief, the guide to how to build a vertical drama character bible for AI generation covers the approved reference frame structure, the Soul ID model specification, and the wardrobe arc that the direction brief's character configuration section draws from.

For the style guide that provides the visual register specification in Section 7, the guide to how to build a vertical drama style guide covers the colour temperature, lighting direction, contrast ratio, and environment category specifications that the direction brief references.

For the ControlNet geometry references that Section 4 specifies, the guide to how to use ControlNet for consistent camera angles in AI vertical drama covers the reference library structure, the camera position categories, and how ControlNet references are built and stored for production use.

Stay connected

For studios moving beyond traditional production.

Let's set
the new standard together.

If you're working on something, we'd like to hear about it.