How to Evaluate a New AI Video Tool for Vertical Drama Before Committing to It

In January 2025 the content was just terrible, it was unwatchable. But in September, when Veo 3 was launched, there was a huge quality jump. That quality jump happened in eight months. The pace of AI video tool development in 2026 means that a tool that is not production-ready today may be production-ready in three months, and a tool that is production-ready today may be superseded by something significantly better in six months.

This creates a specific operational challenge for businesses commissioning AI-native vertical drama and for production companies building their tool stack: how do you evaluate a new AI video tool quickly enough to make a production commitment before the evaluation itself consumes the production timeline?

The useful 2026 comparison is workflow-based: video model layer for generating shots, character layer for consistency, performance layer for dialogue, editing layer for assembly, and production workflow layer for coordination. A new tool needs to be evaluated against the specific layer it would occupy in the production workflow, not against a general quality benchmark that does not reflect the production's actual requirements.

This is the complete test battery: five tests applied in sequence, with specific pass/fail criteria at each stage, that determines whether a new AI video tool is production-ready for vertical drama before any production commitment is made.

Why Generic AI Video Tool Reviews Are Not Enough

The reviews that dominate AI video tool search results evaluate tools against general creative use cases: cinematic landscape generation, character animation, social media clip production. Most tests run identical cinematic scene descriptions across tools and compare output quality, generation speed, and prompt adherence. These are the correct evaluation criteria for a social media content creator. They are insufficient for a vertical drama production.

The vertical drama production's evaluation criteria are different from general creative use cases in three specific ways:

Character consistency across sessions matters more than single-clip quality. Holding a single face, wardrobe, and proportions across a two-to-three minute sequence composed of multiple independent generations still requires manual reference locking. A tool that produces a stunning single clip with a character who looks slightly different in the next clip is not production-ready for a series requiring the same character across 70 episodes.

Native 9:16 support matters more than general video quality. Was the AI trained on tall-screen video? Or does it just make wide video and crop it? Native 9:16 keeps faces in the middle of the frame. Cropped video pushes faces to unexpected positions. A tool that crops 16:9 output to 9:16 is not the same as a tool that generates native 9:16. The face positioning and depth relationship are different in generated-native versus cropped content.

Production workflow integration matters more than standalone output quality. A tool that produces excellent standalone clips but cannot be integrated into a batch generation workflow, does not support API access for automated session management, and does not maintain reference context across a production session is a tool that reduces production efficiency rather than increasing it.

Test 1: The Native 9:16 Generation Test

The first test confirms that the tool generates native 9:16 content rather than cropping 16:9 content to the vertical format.

The test: Generate a close-up character scene in 9:16 aspect ratio with the character's face specified as the primary subject. Compare the character's face position and the background depth relationship against the same prompt in 16:9 format, then manually cropped to 9:16.

What to look for: In native 9:16 generation, the character's face occupies approximately 60% of the frame height with the eyes in the upper third and the chin at or slightly above the horizontal midpoint. The background depth is staged at 18 to 36 inches behind the character. In cropped 16:9, the face is frequently positioned slightly off-center from the crop's optimal position, and the background depth relationship is compressed or expanded by the crop rather than generated for the vertical frame's proportions.

Pass criteria: The character's face is centred in the upper frame in the native 9:16 output with proportions consistent with the vertical drama close-up standard. The background depth reads as correctly staged for phone viewing distance.

Fail criteria: The face position is off-centre or the background depth relationship looks wrong for close-up phone viewing. The tool's 9:16 output is visibly a cropped version of a wider generation rather than a natively generated vertical frame.

Test 2: The Character Consistency Across Sessions Test

The second test is the most commercially critical evaluation for vertical drama production. It tests whether the tool can maintain a character's visual identity across multiple generation sessions rather than only within a single session.

The structural test is whether episode 30 still looks like episode 1 — without burning through regeneration credits on the way there. This test approximates that structural requirement in an evaluation context.

The test: Build a character reference configuration using the tool's available character consistency mechanism, whether reference images, character model training, or persistent context. Generate five clips of the same character on session one, day one. Close the session completely. Return to the tool on day three or four, reload the character configuration from documentation alone, and generate five more clips of the same character.

What to look for: Compare the ten clips across both sessions. Specifically: does the jaw structure match between sessions, does the eye colour match, does the skin tone match? Are there subtle differences in facial proportions that indicate the tool is approximating the character from the reference rather than locking it at the model level?

Pass criteria: The character's facial structure, eye colour, and skin tone are visibly identical across all ten clips from both sessions. The difference between session one and session four clips is not detectable without side-by-side comparison.

Fail criteria: Visible character drift between sessions. The session four character looks similar to the session one character but has subtly different facial proportions, eye colour, or skin tone. Any drift that would be visible to a viewer watching episodes one and thirty in sequence is a fail.

Test 3: The Phone Display Quality Test

The third test confirms that the tool's output quality holds on the actual delivery device rather than only on a production monitor.

The test: Generate a dialogue-heavy close-up scene using the tool. Evaluate the output on the production workstation. Then evaluate the same output on a consumer phone at arm's length in ambient room light with the phone speaker audio at standard volume.

What to look for: The specific phone display quality checks applied to any vertical drama output: face legibility at arm's length, three-second silent test for emotional register communication without audio, and dialogue intelligibility on phone speaker in ambient noise.

Pass criteria: The character's face communicates emotional register clearly at arm's length on the phone display. The first three seconds communicate conflict and genre without audio. The dialogue is intelligible on the phone speaker in a room with moderate ambient noise.

Fail criteria: Any of the three checks fail on the phone display even if the output passes all three on the production workstation. A tool that produces output that looks excellent on a monitor but fails the phone display test is not suitable for vertical drama distribution.

Test 4: The ControlNet or Equivalent Geometry Enforcement Test

The fourth test confirms that the tool supports some form of camera geometry enforcement that can maintain consistent close-up framing across multiple generation sessions.

ControlNet gives you frame-by-frame control over composition, body language, and camera angles, eliminating the random composition problem that makes AI-generated scenes feel generic. Not all tools support ControlNet directly, but production-grade vertical drama generation requires some mechanism for enforcing consistent camera geometry across sessions.

The test: Identify the tool's camera geometry enforcement mechanism: ControlNet, reference frame conditioning, camera angle specification in the prompt, or equivalent. Apply the mechanism to generate ten clips of the same character from the same camera position specification. Compare the camera geometry across all ten clips: is the eye-line height consistent, is the frame distance from the character consistent, is the shoulder position relative to the frame bottom consistent?

What to look for: Geometric consistency in the character's position within the frame across all ten clips. The authority close-up standard: head at 58% frame height from top, shoulders at 52% frame height, camera at or fractionally below eye-line.

Pass criteria: The camera geometry matches the specification consistently across all ten clips with deviation below 5% of the frame dimension.

Fail criteria: Visible variation in the camera geometry across clips despite consistent specification input. If the tool requires a different prompt each time to produce the same camera position, it does not have reliable geometry enforcement for production use.

Test 5: The Production Workflow Integration Test

The fifth test evaluates whether the tool can be integrated into a production workflow rather than used only as a standalone generation tool.

The test: Attempt to: (1) access the tool via API for batch generation rather than through the interface, (2) maintain a character reference configuration that persists between API calls without manual re-specification, and (3) generate a batch of twenty clips sequentially using the API with consistent character and camera geometry across the full batch.

What to look for: API stability across a twenty-clip batch. Character consistency without manual re-loading of the reference configuration between calls. Generation credit consumption per clip at a rate that makes the production economics viable at the current pricing tier.

Pass criteria: The API produces twenty clips sequentially with consistent character identity and camera geometry, at a generation credit cost per finished minute that is within the production's per-finished-minute tool budget.

Fail criteria: API instability that requires manual intervention during the batch. Character drift across the batch despite consistent reference specification. Generation credit costs per finished minute that exceed the production's tool budget at the quality level the test produced.

What to Do With the Test Results

A tool that passes all five tests is production-ready for vertical drama and can be committed to the production workflow for the scene types its generation characteristics are strongest for.

A tool that fails Test 1 (native 9:16) is not appropriate for vertical drama production regardless of its other capabilities. The native 9:16 requirement is non-negotiable for the format.

A tool that fails Test 2 (character consistency across sessions) is appropriate only for scene types where character identity is not primary: environmental scenes, crowd scenes, abstract atmospheric sequences. It should not be used for dialogue close-ups or any scene where the same character must appear consistently across the episode run.

A tool that fails Test 3 (phone display quality) at the audio layer but passes the visual quality checks may be appropriate for generation if audio is replaced in post-production with a separately recorded or generated audio track.

A tool that fails Test 4 (geometry enforcement) is appropriate for scene types where camera geometry variation is acceptable: montage sequences, atmospheric transitions, non-dialogue visual sequences. It should not be used for primary dialogue coverage where the consistent authority close-up standard must be maintained.

A tool that fails Test 5 (production workflow integration) is appropriate for individual creative tasks but not for series-scale batch generation. It can be used for concept testing, hero shot generation, and single-scene production tasks where manual operation is acceptable.

The Evaluation Timeline

The complete five-test battery takes three to five working days to conduct properly. Test 2 requires a minimum of three days to produce the session-gap result. Tests 1, 3, 4, and 5 can be conducted in a single working day.

A tool evaluated in less than three days cannot produce a reliable character consistency result because the session gap required to test cross-session consistency cannot be compressed. Any evaluation that claims to be comprehensive without a session-gap test is not evaluating the most commercially critical production capability.

For businesses commissioning AI-native vertical drama who want to understand which tools their production partner is using and why, this test battery is the evaluation framework to apply to any production company's claimed tool stack. A production company that cannot demonstrate pass results on all five tests for its primary dialogue scene generation tool is not operating a production-grade AI workflow.

Axis AI Studios Perspective

At Axis AI Studios, every new AI video tool is evaluated against all five tests before it enters the production workflow in any role. The evaluation is documented with specific pass/fail results at each test stage. Tools that pass all five tests are assigned to the scene types where their generation characteristics are strongest: Kling 3.0 for dialogue close-ups where character identity is primary, Seedance 2.0 for scenes where native audio generation is relevant, Veo 3.1 for hero shots where generation quality ceiling matters most.

Tools that pass Tests 1, 3, and 4 but fail Test 2 are assigned to environmental and atmospheric scene types where character identity is not primary. No tool is assigned to dialogue close-up production until it has passed Test 2 with a documented session gap of at least three days.

For businesses who want to commission AI-native vertical drama with a production partner whose tool stack has been evaluated against the five-test battery, reach out at business@axisaistudios.com.


FAQ

How Often Should the Test Battery Be Rerun on Tools Already in the Production Workflow?

Every ninety days. Any claim more than 90 days old should be treated as maybe-outdated. The AI video tool market is updating fast enough that a tool's performance on any of the five tests can change significantly within a quarter. A tool that passed Test 2 in March may have introduced a model update in June that changes its character consistency behaviour. The ninety-day rerun discipline ensures that production workflow assignments remain correctly calibrated to current tool capabilities.

Does the Test Battery Need to Be Conducted on Every Tool or Only on New Entrants?

New entrant tools require the full five-test battery. Tools that have been in the production workflow for more than one model version update cycle require a targeted re-evaluation of the tests most likely to be affected by the update: typically Tests 2 and 4 for character and geometry consistency, and Test 3 for any audio capability changes. The full five-test battery is rerun quarterly as a complete workflow audit regardless of individual tool updates.

Should a Business Commissioning AI-Native Production Evaluate Tools Themselves or Rely on the Production Partner's Evaluation?

A business commissioning AI-native production does not need to conduct the five-test battery themselves. They need to confirm that their production partner has conducted it and can provide the documented results. The specific questions to ask: which tool are you using for dialogue close-up generation, when was it last evaluated against character consistency criteria, and what is the documented session gap in your character consistency test. A production partner who cannot answer these questions with specific documented results is not conducting production-grade tool evaluation.


Further Reading

For the complete tool routing logic that applies the test battery results across different scene types in production, the Seedance vs Kling vs Veo guide covers each tool's strengths, scene type routing, and production application.

For the due diligence checklist that confirms a production partner has evaluated their tools correctly, the guide to what to look for in an AI vertical drama production partner covers all six capability domains including the tool stack and workflow integration assessment.

For the character consistency infrastructure that Test 2 of this battery evaluates, the guide to using LoRA training for character consistency in vertical drama covers the complete character model build process that production-grade character consistency requires.

Stay connected

For studios moving beyond traditional production.

Let's set
the new standard together.

If you're working on something, we'd like to hear about it.