How to Use Wan 2.2 for Vertical Drama Generation: The Open-Weights Workflow
Wan 2.2 is an open-source AI video generation model released by Alibaba on July 28, 2025. It generates short video clips from text prompts or images at a quality level that competes with commercial tools. The model is released under the Apache 2.0 license, so you can download the weights from Hugging Face and use them commercially. What sets it apart is its Mixture-of-Experts architecture, which meaningfully changes how the model processes video. Wan 2.2 replaces the dense transformer with the MoE architecture and significantly expands the training data, adding 65.6% more images and 83.2% more videos compared to Wan 2.1. The result is better motion coherence, more consistent character appearance across frames, and stronger camera prompt responsiveness.
The proprietary tools — Seedance 2.5, Kling 3.0, Veo 3.1 — dominate the production routing discussion because they dominate the quality benchmarks. Wan 2.2 does not top any current benchmark against those tools. What it does is something none of them can: it runs locally, on your own hardware, under a license that requires no API agreement, no usage tracking, and no per-second generation cost.
For production companies and businesses evaluating vertical drama production infrastructure, Wan 2.2's open-weights status changes the build-versus-buy calculation. A production company that runs Wan 2.2 locally on a 24GB GPU has a generation infrastructure cost of hardware acquisition plus electricity rather than per-second API fees. At production scale across multiple series, the cost structure difference is commercially significant.
This is the complete guide to using Wan 2.2 in a vertical drama production workflow: the setup, the character consistency approach, the 9:16 prompt structure, the scene type routing logic, and when Wan 2.2 is the correct tool versus when the proprietary tools are the correct routing.
What Wan 2.2 Is and Why It Matters
For ComfyUI users with an RTX 4090, 7900 XTX, or Mac Studio, Wan 2.2 is the newest Wan release with open downloadable weights and the right open-source video generator in 2026. The dense TI2V-5B variant generates a 5-second 720p clip at 24fps on a single 24GB GPU in under approximately 9 minutes, while the heavier MoE A14B models push quality higher for users with more VRAM.
The Apache 2.0 license is the commercially significant fact. It means:
You can use Wan 2.2 for commercial production without a usage agreement with Alibaba. You can download the weights to your own hardware and run generation without any API call. You can fine-tune the model on your own training data. You can modify the model's architecture for your production workflow. And you are not subject to changes in API pricing, API availability, or terms of service that affect the proprietary tools.
Wan 3.0, released in early 2026, improves on previous versions in several dimensions. Wan 2.6 remains relevant for creators with limited hardware — its smaller memory footprint makes it more accessible on 8 to 12 GB GPUs. A small animation studio producing a 10-minute short film needs to generate hundreds of video clips with consistent character designs, environments, and visual styles.
Wan 2.2 specifically, rather than the newer Wan 2.6 or 2.7, is the correct starting point for production companies building local generation infrastructure for three reasons: the weights are publicly available and well-documented, the ComfyUI integration is mature, and the hardware requirement of 24GB VRAM is achievable with a single consumer RTX 4090 rather than requiring enterprise GPU access.
Hardware and Setup
The minimum viable hardware for Wan 2.2 vertical drama generation:
GPU: RTX 4090 (24GB VRAM) for the TI2V-5B variant. The MoE A14B variant requires 40GB or more VRAM and is appropriate for production companies with access to enterprise GPU hardware or cloud GPU services.
RAM: 32GB system RAM minimum. 64GB for production workflows that maintain multiple character reference packs and style LoRAs simultaneously.
Storage: 50GB for the TI2V-5B model weights. 100GB for the MoE A14B weights plus LoRA training data and reference pack storage.
To skip setup entirely, Thunder Compute offers a one-click ComfyUI template on an RTX A6000 at $0.35 per hour. For production companies that want to evaluate Wan 2.2 before hardware acquisition, cloud GPU access through Thunder Compute or equivalent services provides the correct hardware environment without capital expenditure.
Installation through ComfyUI:
The ComfyUI Manager node manager installs the Wan video generation custom nodes through the GUI without requiring command-line configuration. The model weights download from Hugging Face through the ComfyUI Manager interface. The standard installation sequence from the ComfyUI Desktop app takes approximately 30 to 60 minutes on a stable internet connection.
Character Consistency in Wan 2.2
Character consistency in Wan 2.2 requires a different approach from the Soul ID character model training that Seedance uses. The open-weights architecture means the consistency infrastructure is built through the ComfyUI node graph rather than through a proprietary reference system.
For consistent style, pair with a style LoRA. Community fine-tunes for anime, oil painting, comic book, and various film looks are on HuggingFace. Last-frame-as-first-frame chaining: generate clip A, extract its last frame, feed it as the I2V input for clip B, repeat. Maintains visual continuity but loses long-range coherence after a couple of chains.
The three-layer character consistency approach for Wan 2.2 in vertical drama production:
Layer 1: Character LoRA. Train a character-specific LoRA on 15 to 30 approved reference images of the character. The LoRA bakes the character's visual identity into the model weights at a small adapter scale. LoRA fine-tuning survives dramatic pose, lighting, and scene changes in ways that seed locking cannot. Strength: 0.65 to 0.75 in the LoRA Loader node. Alici
Layer 2: IP-Adapter reference grounding. Load the character's approved reference image in the IP-Adapter node for session-specific face anchoring alongside the character LoRA. Weight: 0.75 to 0.85. The IP-Adapter handles the session-level face identity precision that the LoRA establishes at the weight level.
Layer 3: First-last frame continuity chaining. For multi-clip episode generation, extract the last frame of each approved clip and use it as the first-frame input for the next clip generation. This maintains visual continuity between consecutive clips within the same episode. The chaining maintains short-range continuity effectively but loses consistency after four to five consecutive chains. For vertical drama production, this means the chaining approach works for a full 90-second episode across three 30-second segments but requires a full reference reload for the next episode rather than continued chaining across episodes.
The 9:16 Vertical Drama Prompt Structure for Wan 2.2
Wan 2.2 has stronger camera prompt responsiveness than Wan 2.1. The improved camera prompt responsiveness means the five-layer prompt structure used for Seedance and Kling applies to Wan 2.2 with minor modifications for the model's specific prompt interpretation patterns.
The Wan 2.2 five-layer vertical drama prompt:
Layer 1 — Frame format: "vertical 9:16 aspect ratio, mobile portrait orientation, phone screen format" — Wan 2.2's improved training data includes vertical video, but explicit format specification produces more reliable 9:16 output than relying on the aspect ratio parameter alone.
Layer 2 — Camera position: "close-up, character's face occupying upper 60% of frame, camera at eye-line height, background depth 18 to 24 inches behind character" — Wan 2.2's camera prompt responsiveness makes this layer more effective than in Wan 2.1, but combine with a ControlNet OpenPose reference for authority close-up consistency.
Layer 3 — Subject and performance: Physical behavior rather than emotional labels, consistent with the vertical drama direction brief standard. "Jaw set, controlled stillness, eyes forward, suppressed tension visible in minimal movement" rather than "looking angry but controlled."
Layer 4 — Environment: "contemporary office interior, warm key light from camera left at 45 degrees, dark background depth, single warm practical light visible at frame edge" — Wan 2.2's environment generation is strong for interior scenes. The specific lighting direction in the prompt produces consistent results.
Layer 5 — Technical: "4K quality, sharp focus on foreground character, soft bokeh background, natural skin tones, phone speaker-appropriate audio if audio generation is enabled."
The ControlNet Configuration for Wan 2.2
Wan 2.2 supports ControlNet through ComfyUI's standard ControlNet custom node. The OpenPose conditioning that enforces camera geometry in the proprietary tool stack applies to Wan 2.2 through the same ComfyUI ControlNet node configuration described in the ControlNet guide.
The specific ControlNet configuration for Wan 2.2 vertical drama close-up generation:
ControlNet model: OpenPose or DWPose model downloaded to the ComfyUI controlnet models folder.
ControlNet strength: 0.65 to 0.75 for Wan 2.2. The slightly lower strength relative to the Seedance and Kling configuration reflects Wan 2.2's different response to conditioning inputs — higher strength values in Wan 2.2 can produce rigidity in facial expression that the vertical drama close-up's emotional register requirement cannot accommodate.
Control image: The same OpenPose skeleton reference library used for Seedance and Kling generation, specifying the authority close-up, proximity close-up, and vulnerability close-up positions from the production's style guide.
The Audio Generation Capability
Wan 2.5 adds native audio-visual generation, synchronized sound and video in one pass, 1080p output at 24fps, and clips up to 10 seconds long. However, Wan 2.5 weights are not publicly available for local deployment.
Wan 2.2 does not include native audio generation. Audio post-production for Wan 2.2-generated content uses the same AI voice cloning and audio mixing workflow described in the ADR guide and the phone speaker mixing guide. The absence of native audio in Wan 2.2 is not a production limitation; it is the standard vertical drama workflow where audio is handled separately from visual generation in post-production.
For production companies that want native audio generation with open-weights access, Wan 2.5's managed API endpoints provide the capability without open-weight local deployment. The audio-visual synchronized generation in Wan 2.5 through the managed API produces a different workflow from Wan 2.2's local generation, but it is available for productions that prioritize native audio synchronization over local hardware control.
Scene Type Routing: When Wan 2.2 Is the Correct Tool
Wan 2.2's routing position in the vertical drama tool stack is determined by its specific strengths relative to the proprietary tools and by the cost structure difference that its open-weights status creates.
Wan 2.2 routes correctly for:
High-volume atmospheric and environmental scene generation where per-second API cost accumulates significantly at production scale. A 70-episode series with extensive environmental establishing shots and atmospheric transition clips generates significant API cost at Seedance or Kling pricing. Running these scene types locally on Wan 2.2 eliminates the per-second cost for the scene category with the lowest character consistency requirements.
Style-consistent B-roll generation where a production-specific style LoRA trained on the series' approved visual register produces atmospheric content in the series' specific visual language without the API cost of proprietary tools.
Pre-production concept testing where the production company wants to generate multiple premise visualizations, character configuration tests, or environment design explorations before committing to the production's primary tool stack. Wan 2.2's zero per-second cost makes extensive concept generation economically viable.
Wan 2.2 does not route correctly for:
Primary dialogue close-up generation where cross-session character consistency is the primary commercial quality requirement. Seedance 2.5's 50-reference capability and Soul ID character model architecture produce superior cross-session character consistency for the dialogue close-up scene type.
Paywall episode hero moments where the generation quality ceiling is commercially critical. Veo 3.1 remains the correct routing for the highest-commercial-consequence scene positions.
The Cost Structure Comparison
The open-weights cost structure is Wan 2.2's primary commercial advantage for production companies evaluating vertical drama generation infrastructure investment.
Proprietary tool API costs for dialogue close-up generation at production scale: approximately $2 to $5 per finished minute for Seedance 2.5, $0.07 to $0.14 per second for Kling 3.0, and higher for Veo 3.1. For a 70-episode series at 90 seconds per episode, the total visual content duration is 105 minutes. At $3 average per finished minute, the API generation cost for the full series is approximately $315 in generation credits, excluding revision attempts.
Local Wan 2.2 generation cost: electricity plus hardware amortization. At $0.10 per kWh electricity cost and an RTX 4090's approximately 350W TDP during generation, a 9-minute generation produces approximately $0.005 in electricity cost. The hardware amortization across 3,000 hours of GPU life at $2,000 acquisition cost is approximately $0.67 per GPU hour. Total per-generation cost at Wan 2.2 local rates: approximately $0.10 per 5-second clip, or approximately $1.80 per finished minute.
The local generation cost advantage over Seedance API cost: approximately $1.20 per finished minute. Across a 70-episode series, the difference is approximately $126. Across 10 series per year, the difference is approximately $1,260. The hardware acquisition cost of an RTX 4090 is recoverable across approximately 10 to 15 series of local generation for the atmospheric and B-roll scene categories alone.
Axis AI Studios Perspective
Wan 2.2 has a specific and commercially justified position in the production tool stack: local generation for high-volume atmospheric and environmental content where the open-weights cost structure reduces total production infrastructure cost without compromising the quality of the dialogue close-up and hero shot content that routes to the proprietary tools.
At Axis AI Studios, Wan 2.2 is evaluated as part of the open-source production infrastructure assessment rather than as a replacement for the primary tool routing. Its Apache 2.0 license, local deployment capability, and MoE architecture make it the correct routing for production categories where the proprietary tool cost is not justified by the scene type's commercial consequence.
For businesses who want to commission AI-native vertical drama with a production partner whose tool stack includes both open-weights local generation for cost efficiency and proprietary tools for maximum quality on commercially critical scenes, reach out at business@axisaistudios.com.
FAQ
Is Wan 2.2 Suitable for a Full 70-Episode Series as the Primary Generation Tool?
It depends on the production's quality requirements and the platform's acquisition standard. For concept test series and tier-3 platform distribution, Wan 2.2 can serve as the primary generation tool with appropriate character LoRA training and ControlNet configuration. For tier-1 and tier-2 platform acquisition at standard professional quality, Wan 2.2 is most appropriately a secondary tool for atmospheric and environmental content, with Seedance 2.5 or Kling 3.0 handling the primary dialogue close-up routing where cross-session character consistency is commercially critical.
Does the Wan 2.2 Open License Cover Commercial Production and Platform Delivery?
Yes. The Apache 2.0 license permits commercial use, modification, and distribution without royalty payments to Alibaba. Content generated using Wan 2.2 can be commercially distributed on vertical drama platforms under the same licensing terms as content from any other production tool. The delivery package's AI tool usage documentation references Wan 2.2 under the Apache 2.0 license, which provides the chain of title documentation that platform delivery requires.
What Is the Character Consistency Quality Difference Between Wan 2.2 and Seedance 2.5?
At the single-session level, Wan 2.2 with character LoRA and IP-Adapter produces character consistency comparable to Seedance 2.0 for the session's duration. The cross-session consistency difference is larger: Seedance 2.5's 50-reference capability and Soul ID character model architecture maintain character identity across sessions produced weeks apart more reliably than Wan 2.2's last-frame chaining and LoRA-based approach. For productions requiring tight cross-session character consistency across 70 episodes, Seedance 2.5 is the superior routing. For productions where session-level consistency is sufficient and cross-session variation is acceptable for the scene type, Wan 2.2 produces commercially viable results.
Further Reading
For the ComfyUI node configuration that implements the character LoRA, IP-Adapter, and ControlNet stack described in this post, the guide to ComfyUI for vertical drama production covers the complete node workflow, model loading sequence, and production workflow integration.
For the five-test battery that evaluates Wan 2.2's production readiness for a specific production's requirements, the guide to how to evaluate a new AI video tool for vertical drama covers the character consistency across sessions test and the phone display quality test that determine the routing assignment.
For the tool routing comparison that positions Wan 2.2 against Seedance 2.5 and Kling 3.0 across scene types, the guide to which tool to master first for vertical drama production covers the production routing logic and which tool competency is most commercially valuable to develop first.

Let's set
the new standard together.
If you're working on something, we'd like to hear about it.
