How to Test Character Consistency Across Sessions Before a Full Series Begins Production
The character consistency problem in AI vertical drama is not within a single generation session. Within a session, consistency works because the generation tool processes all outputs in the same computational context. The character's visual identity is stable across every output produced in that session.
The problem is between sessions. Each new session starts without memory of the previous one. The character reference pack — the approved reference images and the Soul ID model — is the bridge that connects sessions. If that bridge is correctly built and correctly loaded, the character looks the same in session seven as in session one. If the bridge has a gap — an incorrectly trained model, an incomplete reference image set, or a reference loading procedure that produces inconsistent results — the character drifts between sessions in ways that accumulate across 70 episodes.
The character consistency test is the pre-production protocol that identifies bridge gaps before the production begins. Not a single test output. A structured multi-session test that simulates the production's session sequence and reveals whether the character reference infrastructure holds across the kind of session transitions the full production will require.
This post covers the complete pre-production consistency testing protocol: what the test produces, how many sessions it requires, what the evaluation criteria are, and what the pass and fail thresholds mean for the production decision.
Why Pre-Production Testing Is Different From the Session-Open Consistency Check
The session-open consistency check — generating one test output at the start of each production session and comparing it against the approved reference frame — is an in-production quality control mechanism. It catches deviations as they occur and before production generation proceeds. It does not prevent deviations. It detects them.
The pre-production consistency test is different in purpose and structure. It is not a detection mechanism. It is a stress test of the character reference infrastructure that reveals whether the infrastructure is capable of maintaining consistency across the full production's session sequence before any production generation is committed.
A production that skips the pre-production consistency test and relies only on session-open checks is a production that discovers infrastructure gaps during live production — at the point where correction requires stopping, rebuilding, and regenerating rather than rebuilding before the first production session begins.
The Pre-Production Consistency Test Structure
The test simulates the production's session structure across three simulated sessions, separated by time intervals that replicate real production conditions. Each session uses the approved character reference configuration. The test evaluates whether the configuration holds across the time gaps and the re-loading procedures the real production will require.
Test Session 1: Baseline establishment.
Generate 10 test outputs using the approved character reference configuration for the primary characters at Arc Position 1. The generation tool is Seedance 2.5 at the production's standard reference conditioning strength. All 10 outputs use the approved Soul ID model for each character.
After generation, identify the single output for each character that most precisely matches the approved reference frame across all four consistency criteria: jaw structure, eye colour, skin tone, and nose bridge proportions. This output becomes the Session 1 baseline reference — a generated output from the production's own infrastructure rather than an external reference image.
Time gap: 48 to 72 hours.
Wait 48 to 72 hours before Session 2. This interval replicates the typical gap between production sessions. During this interval, close all generation tool tabs and clear any cached session state. The test must replicate cold session loading conditions — the same conditions the production's operators will face when beginning a session two days after the previous one.
Test Session 2: First consistency evaluation.
Load the approved character reference configuration from scratch — the same loading procedure the production operators will use at the start of each real session. Generate 10 test outputs per character using the identical parameters as Session 1.
Compare each Session 2 output against the Session 1 baseline reference using the four consistency criteria. The comparison is conducted on a consumer phone at arm's length — the same evaluation device the quality review uses during production.
Time gap: 48 to 72 hours.
Test Session 3: Second consistency evaluation.
Repeat the Session 2 procedure. Generate 10 outputs per character. Compare each output against the Session 1 baseline reference.
The Evaluation Criteria and Pass Thresholds
Jaw structure: The character's jaw width, jaw angle, and chin definition must match the Session 1 baseline reference within visual tolerance at arm's length. A deviation that requires side-by-side comparison to detect is within tolerance. A deviation visible on first viewing without comparison is outside tolerance.
Eye colour: Exact match required across all sessions. Eye colour is the most immediately visible character attribute and the one audiences notice most readily when it shifts. Any perceptible shift in eye colour across sessions is a fail.
Skin tone: Match within one perceptible step. A slight warmth or coolness shift between sessions that would not be noticed by a viewer watching episodes sequentially is within tolerance. A shift that produces a clearly different skin tone category across sessions is outside tolerance.
Nose bridge proportions: The width and height of the nose bridge relative to the face must match the Session 1 baseline reference within visual tolerance at arm's length. Minor variation in generation is expected. The test evaluates whether the nose bridge structure is consistent rather than whether it is pixel-identical.
Pass threshold: 8 of 10 outputs per character per session must pass all four criteria without revision. A pass rate below 80% indicates that the character reference infrastructure requires rebuilding before production begins.
Fail threshold: Any session where fewer than 6 of 10 outputs pass all four criteria indicates a systematic infrastructure failure — not random generation variance but a consistent gap in the character reference configuration that will produce visible drift across the production.
What the Test Results Tell You
All three sessions above the 80% pass threshold:
The character reference infrastructure is production-ready. The Soul ID model, the reference image set, and the loading procedure all hold across simulated production session gaps. The production can begin with confidence that the character reference infrastructure will maintain consistency across the full 70-episode session sequence.
One session between 60% and 80% pass rate:
The character reference infrastructure has a moderate vulnerability. The most likely cause is a reference image set that is not sufficiently comprehensive for the character's full range of expression or lighting context — the reference images correctly encode the character in some conditions but not in all the conditions the production will require.
The corrective action: expand the reference image set to include additional expression states and lighting contexts. Rebuild the Soul ID model from the expanded reference set. Rerun the three-session test before beginning production.
Any session below 60% pass rate:
The character reference infrastructure has a systematic failure. The most likely cause is a Soul ID model trained on insufficient data, a reference conditioning strength that is too low to reliably encode the character's identity, or a character visual design that contains features the current generation tool's character reference system cannot reliably reproduce.
The corrective action: rebuild the character reference from the beginning. Increase the reference conditioning strength. If the character's visual design contains the reproduction problem, evaluate whether the character design needs modification before the reference rebuild. This is a significant pre-production investment — but it costs considerably less than discovering the systematic failure at episode 35 and rebuilding under production pressure.
Consistent drift across all three sessions in a single direction:
A consistent drift — the character becomes progressively warmer in skin tone across sessions, or the jaw structure narrows progressively — indicates that the Soul ID model is not fully anchoring the character against the generation tool's own tendency to drift toward its base model's aesthetic preferences. The corrective action is increasing the reference conditioning strength and retraining the model with a higher-weight reference image set that more strongly anchors the character's specific features.
The Arc Position Extension
The three-session test described above covers Arc Position 1 — the character configuration at the production's opening arc. For productions where the character undergoes significant visual changes across arc positions (wardrobe changes, injury marks, physical state changes), the test should be extended to cover Arc Position 2 and Arc Position 3 reference configurations.
An Arc Position 2 character reference that passes the three-session test independently provides the same infrastructure confidence for the production's middle arc as the Arc Position 1 test provides for the opening arc. Skipping the Arc Position 2 test assumes that the Arc Position 2 reference infrastructure is as robust as the Arc Position 1 infrastructure — an assumption that the reference build quality may not support.
The total testing investment for a full arc extension — three sessions per arc position, three arc positions — is nine simulated sessions across ten to twelve days. This is a significant pre-production investment. It is also the investment that eliminates the character consistency risks that compound most damagingly across a 70-episode production.
Axis AI Studios Perspective
At Axis AI Studios, the pre-production consistency test is mandatory for every commission before the first production session begins. The three-session test is scheduled into the pre-production calendar as a fixed gate: no production generation begins until the consistency test results are documented and reviewed by the creative director.
For businesses commissioning AI-native vertical drama who want to confirm their production partner runs pre-production consistency testing before your series enters generation, this is a specific question to ask in the due diligence conversation: can the production partner describe their pre-production character consistency testing protocol, how many sessions it covers, and what the pass threshold is?
Reach out at business@axisaistudios.com for commissioning conversations or to discuss what the pre-production testing protocol looks like for your specific character design and arc structure.
FAQ
Does the Pre-Production Consistency Test Work for Supporting Characters as Well as Primary Characters?
Yes, but the investment priority should reflect the character's screen time and commercial importance. Primary characters who appear in every episode and at the paywall episode require the full three-session test across all arc positions. Supporting characters who appear in fewer than 20 episodes require a reduced version: one test session of five outputs per character at their primary arc position. Secondary characters who appear in fewer than five episodes can rely on session-open consistency checks during production rather than pre-production testing.
How Long Does the Complete Pre-Production Consistency Test Take?
Three test sessions separated by 48 to 72-hour gaps requires approximately seven to nine calendar days from the first session to the final results. This is the minimum pre-production testing timeline. Adding arc position extensions for Arc Position 2 and Arc Position 3 extends the timeline to fourteen to twenty calendar days. This timeline should be built into the pre-production schedule rather than treated as an optional addition.
What If the Character Reference Infrastructure Fails the Test Close to the Production Start Date?
A test failure close to the production start date is the correct time to discover the failure — not at episode 35 of a live production. The corrective action (reference rebuild and retest) should delay the production start rather than be bypassed because of timeline pressure. A production that begins with a known character reference infrastructure failure is a production the creative director already knows will require expensive mid-production correction. The delay cost of the retest is significantly lower than the correction cost of the mid-production failure.
Further Reading
For the character bible that provides the reference image set and the Soul ID model specification the consistency test uses, the guide to how to build a vertical drama character bible for AI generation covers the approved reference frame structure, Soul ID model training, and the session reproduction specification.
For the LoRA training that provides an alternative character reference approach alongside Soul ID, the guide to how to use LoRA training for character consistency in vertical drama covers the training process, the conditioning strength settings, and the cross-session consistency the trained model produces.
For the generation operator's session-open consistency check that the pre-production test's procedure is based on, the guide to the generation operator's role in AI-native vertical drama covers the full session discipline including how the session-open check is conducted and documented.

Let's set
the new standard together.
If you're working on something, we'd like to hear about it.
