The Five-Layer Prompt Structure for AI Vertical Drama Generation: What Each Layer Does and Why Order Matters
Most generation operators write prompts as single-block descriptions: a sentence or paragraph that describes the scene, the characters, the setting, and the mood in natural language order. This approach produces inconsistent outputs because generation tools do not process all prompt content with equal weight. Content that appears earlier in the prompt receives more model attention than content that appears later. A character description buried after a setting description produces weaker character reference conditioning than the same description placed at the prompt's opening.
The five-layer prompt structure is the ordering framework that ensures every generation prompt is written in the sequence that AI video generation tools process most reliably. Not because the framework is what the direction brief requires — the direction brief specifies what to produce. But because the framework is how generation tools read prompts, and writing in that sequence produces outputs that match the brief more consistently than writing in natural language order.
This post covers each of the five layers: what it specifies, why it occupies its position in the sequence, what happens when it is missing or misplaced, and the specific language that makes each layer precise rather than interpretive.
Layer 1: Character Configuration
Position in sequence: First.
Why it occupies this position: Generation tools weight earlier prompt content more heavily when encoding the generation's primary subject. The character is the primary subject of every dialogue close-up and every proximity shot in AI vertical drama. Placing the character configuration first ensures the generation tool encodes the character's identity as the most prominent feature of the output rather than as a secondary detail competing with setting or action for model attention.
What it specifies: The character's visual identity as described in the approved character bible — not an external reference image (that is handled through the ControlNet or Soul ID configuration, not the text prompt) but a text description that the generation tool uses alongside the reference to anchor the character's appearance.
The correct layer 1 format:
[Character name], [specific jaw structure], [specific eye colour], [skin tone on the Fitzpatrick scale or equivalent description], [specific nose bridge proportions]. [Wardrobe state for this arc position]. [Arc position identifier so the generation tool weights the character's current arc state].
Example: The controlled alpha. Angular jaw, defined chin. Cool grey eyes. Warm medium skin tone. Straight nose bridge, medium width. Dark charcoal suit, no tie, white shirt. Arc Position 2 — institutional authority phase.
What happens when Layer 1 is missing or imprecise: The generation tool defaults to its own interpretation of the character based on contextual cues in the other layers. A character described only as "a powerful executive in a corporate setting" will produce a different face in every generation session. Character consistency failure across sessions is almost always traceable to a Layer 1 that is either absent or insufficiently specific.
Layer 2: Environment
Position in sequence: Second.
Why it occupies this position: The environment is the secondary spatial context that the generation tool uses to calibrate the character's relationship to their setting. Placing it second — after the character is anchored but before the camera position is specified — allows the generation tool to establish the spatial context before determining how to compose the frame.
What it specifies: The environment category, the colour temperature, the lighting direction, the contrast ratio, and the background depth — all at the specification level from the style guide rather than at the mood description level.
The correct layer 2 format:
[Environment category from the style guide]. Colour temperature: [Kelvin value]. Primary lighting: [direction relative to camera axis, angle in degrees]. Contrast ratio: [key to fill]. Background depth: [distance from character plane in visual estimation]. [Any specific environment details relevant to this episode's scene — a specific prop, a window, a door position].
Example: Corporate interior. Colour temperature: 4,200K. Primary lighting: camera-left at 45 degrees. Contrast ratio: 4:1. Background depth: dark, out-of-focus, approximately 18 inches behind the character plane. No natural light sources visible. Polished dark surface at the character's desk level.
What happens when Layer 2 is missing or imprecise: The generation tool selects its own environment based on the character and action descriptions. Two different operators writing "office confrontation scene" will produce two different environments — different colour temperatures, different lighting directions, different background depths — even if every other layer in their prompts is identical. Visual register drift across episodes is almost always traceable to an imprecise or absent Layer 2.
Layer 3: Camera Geometry
Position in sequence: Third.
Why it occupies this position: Camera geometry translates the established character and environment context into a specific compositional structure. Placing it third — after character and environment are anchored but before action and emotional register are specified — allows the generation tool to compose the frame before determining what is happening within it.
What it specifies: The shot type, the character's position in frame, the head height relative to the frame, the eye line relative to the camera axis, and the specific ControlNet reference if one is being used for this shot.
The correct layer 3 format:
[Shot type from the approved camera position library]. Character: [position in frame — left third, right third, centre, close centre]. Head height: [position relative to frame — upper quarter, upper third]. Eye line: [direction relative to camera axis — level, slightly above, slightly below]. ControlNet reference: [reference file name from the library].
Example: Authority close-up. Character: right two-thirds of frame. Head height: upper quarter. Eye line: slightly above camera axis — character looking marginally down toward the implied viewer position. ControlNet reference: AUTHORITY_CU_ARC2.png.
What happens when Layer 3 is missing or imprecise: The generation tool composes the frame based on the scene type implied by the action layer. An "office confrontation" action layer without a specific camera geometry layer produces unpredictable framing — sometimes a two-shot, sometimes a single close-up, sometimes a wide establishing shot. Camera composition inconsistency across episodes is almost always traceable to an absent or imprecise Layer 3.
Layer 4: Action
Position in sequence: Fourth.
Why it occupies this position: Action is what happens within the established character, environment, and compositional context. Placing it fourth — after the static visual parameters are anchored — ensures the generation tool interprets the action within the correct visual context rather than allowing the action description to determine the visual context retroactively.
What it specifies: What the character is doing in this scene — the specific physical action or the absence of action (a static hold), the duration of the action within the generation's time window, and any secondary character or prop interaction relevant to the scene.
The correct layer 4 format:
[Character name] [specific action description]. [Duration note if relevant — holds position throughout, begins from, transitions to]. [Secondary character or prop interaction if present]. [End state — how the scene resolves visually within the generation window].
Example: The controlled alpha holds position behind the desk, both hands flat on the surface. No movement throughout. Eyes forward, unblinking. The protagonist stands opposite — the second character is visible at the frame's left edge, slightly out of focus. The alpha does not move at any point in the generation window.
What happens when Layer 4 is missing or imprecise: The generation tool generates movement that seems appropriate to the scene context rather than the specific movement the direction brief required. A static authority scene generates the character shifting, gesturing, or adjusting their position — not because the brief requested it but because the generation tool defaulted to "natural" movement in the absence of specific action instruction. Unintended character movement in dialogue close-ups is almost always traceable to an imprecise Layer 4 action specification.
Layer 5: Emotional Register
Position in sequence: Fifth.
Why it occupies this position: Emotional register is the generation tool's final calibration pass — the instruction that determines the specific micro-expression and performance quality that the character's face and body produce within the already-established visual context. Placing it last ensures the emotional register is applied to the correct visual foundation rather than allowing it to drive the visual context.
What it specifies: The character's physical state — not an emotion label but a physical description that the generation tool can translate into specific facial and body expression.
The correct layer 5 format:
Emotional register: [physical state description]. [Jaw descriptor]. [Eye descriptor]. [Facial variation instruction — minimal, controlled, present]. [Body language descriptor if relevant]. [Suppression instruction if this is a controlled alpha scene].
Example: Emotional register: controlled authority. Jaw level, no tension visible. Eyes forward, direct, no softening. Minimal facial variation throughout. Shoulders level, no forward lean. The character is performing certainty — not aggression, not satisfaction. The performance is the absence of visible effort.
What happens when Layer 5 is missing or imprecise: The generation tool defaults to whatever emotional register seems contextually appropriate to the scene. A corporate confrontation scene without a specific emotional register instruction produces dramatic expressions that overplay the scene's tension — visible anger, visible dominance, visible effort — rather than the suppressed authority performance that vertical drama's controlled alpha archetype requires. Overplayed expressions in dialogue close-ups are almost always traceable to an absent or imprecise Layer 5.
The Complete Prompt in Layer Sequence
A direction brief for a single dialogue close-up in a corporate confrontation scene, written in five-layer sequence:
Layer 1 — Character: The controlled alpha. Angular jaw, defined chin. Cool grey eyes. Warm medium skin tone. Straight nose bridge, medium width. Dark charcoal suit, no tie, white shirt. Arc Position 2 — institutional authority phase.
Layer 2 — Environment: Corporate interior. Colour temperature 4,200K. Primary lighting camera-left at 45 degrees. Contrast ratio 4:1. Background dark, out-of-focus, 18 inches behind character plane. No natural light sources. Dark surface at desk level.
Layer 3 — Camera: Authority close-up. Character right two-thirds. Head height upper quarter. Eye line slightly above camera axis. ControlNet reference AUTHORITY_CU_ARC2.png.
Layer 4 — Action: The controlled alpha holds position behind the desk, both hands flat on the surface. No movement throughout. Eyes forward, unblinking. Protagonist visible at frame's left edge, slightly out of focus. No movement at any point.
Layer 5 — Emotional register: Controlled authority. Jaw level, no tension visible. Eyes direct, no softening. Minimal facial variation. Shoulders level, no forward lean. Performance is the absence of effort, not the presence of dominance.
This prompt written as a single natural language paragraph would produce a less consistent output than the five-layer structure — not because the content differs but because the sequence in which generation tools process prompt content differs from the sequence natural language uses.
Axis AI Studios Perspective
The five-layer prompt structure is the operational standard for every generation session at Axis AI Studios. Direction briefs are written in layer sequence, reviewed in layer sequence, and documented in layer sequence in the generation log. An output that fails a quality criterion is traced back to the layer that produced the failure — character inconsistency to Layer 1, visual register drift to Layer 2, framing error to Layer 3, unintended movement to Layer 4, overplayed expression to Layer 5.
The diagnostic value of the five-layer structure is as important as its generative value. A production team that knows which layer produced which type of failure can correct precisely rather than rewriting the entire prompt and hoping for a better result.
For generation operators who want to develop five-layer prompt discipline before applying for production roles, the layer structure described in this post can be practised against any publicly available vertical drama content using the scene type categories from the generation operator's role guide.
Reach out at business@axisaistudios.com for production team applications or for commissioning conversations that include prompt discipline as a verified production capability.
FAQ
Does the Five-Layer Structure Apply to All Scene Types or Only to Dialogue Close-Ups?
The structure applies to all scene types but with different Layer 4 specifications. Dialogue close-ups use static action specifications (holds position, minimal movement). Action sequences use dynamic action specifications (the character crosses the frame from left to right, beginning at the left edge and exiting at the right edge in six seconds). Environmental establishing shots minimise Layer 1 (the character may be absent or distant) and expand Layer 2 (the environment is the primary subject). The layer sequence remains the same across all scene types. The content of each layer adapts to the scene type.
Can the Five-Layer Structure Be Used for Multi-Character Scenes?
Yes, with an additional character specification added to Layer 1. Multi-character Layer 1 specifies both characters in sequence, from the primary character (the character with the most screen time in the shot) to the secondary character. Layer 3 specifies the proximity close-up camera geometry that includes both characters in frame. Layer 4 specifies both characters' actions in sequence. Layer 5 specifies both characters' emotional registers in sequence. The additional complexity makes multi-character Layer 1 and Layer 5 longer than single-character equivalents, but the layer structure remains the same.
How Does the Five-Layer Structure Interact With Soul ID or LoRA Character Reference?
The five-layer text prompt works alongside the character reference configuration rather than replacing it. The Soul ID model or LoRA reference encodes the character's visual identity at the pixel level. The Layer 1 text prompt reinforces that encoding with the generation tool's text-to-visual processing. Both operate simultaneously — the reference configuration handles the visual identity and the Layer 1 text reinforces the specific attributes the reference should anchor. Removing Layer 1 and relying only on the reference produces slightly weaker character consistency than using both, because the text reinforcement keeps the generation tool's attention on the character's specific features throughout the prompt processing.
Further Reading
For the direction brief that contains the five-layer prompt specifications for each episode's scenes, the guide to how to write a vertical drama direction brief that a generation operator can execute without clarification covers how the direction brief's seven sections produce the input for each layer of the five-layer structure.
For the ControlNet geometry references that Layer 3 specifies, the guide to how to use ControlNet for consistent camera angles in AI vertical drama covers the reference library structure, the camera position categories, and how ControlNet references interact with text prompt camera specifications.
For the character bible that supplies the Layer 1 character configuration data, the guide to how to build a vertical drama character bible for AI generation covers the psychological configuration, the visual specification, and the arc position wardrobe states that Layer 1 is built from.

Let's set
the new standard together.
If you're working on something, we'd like to hear about it.
