How to Evaluate Your First Delivered Series Before Accepting It

The delivery acceptance window in most vertical drama production agreements runs 10 to 15 business days from delivery package receipt. After that window closes, the production agreement's delivery obligations are discharged and any quality issues the commissioning party identifies become revision requests rather than delivery failures — with different commercial terms.

Most first-time commissioners use the delivery acceptance window to watch some episodes and form a general impression of the quality. This is not a delivery evaluation. It is a viewing experience. A viewing experience does not systematically identify the specific quality failures that platform acquisition review will identify when the series is presented for distribution.

The structured delivery evaluation described in this post takes three to four hours for a 70-episode series. It requires a consumer phone, a quiet room, and a moderate-noise room. It applies the same five quality markers that platform acquisition teams use when they evaluate content for acquisition, so the commissioning party is applying the identical standard the platform will apply. Any quality failure the commissioning party identifies in the delivery evaluation is a quality failure the platform acquisition team would have identified — and the commissioning party has identified it while there is still a commercial mechanism to require correction.

The Evaluation Setup

What you need:

A consumer smartphone at standard display brightness. Not a calibrated monitor. Not a laptop. The phone is the delivery device. Evaluating on any other device produces an evaluation against the wrong standard.

A quiet room for the audio intelligibility baseline test.

A moderate-noise room — a kitchen with ambient appliance noise, a room with a TV audible in the background — for the ambient condition audio test.

Two to three hours of uninterrupted review time.

The delivery brief's specifications: the agreed paywall episode position, the character reference images from the pre-production approval, and the delivery technical specifications confirmed at brief stage.

What to evaluate:

Not all 70 episodes. A structured sample across the episode run. The specific episodes to review: episodes one, five, ten (the free episode window and hook quality), the paywall episode, episodes twenty-five and forty-five (mid-arc character consistency), and episodes sixty-five and seventy (resolution sequence quality). Total viewing time for the sample: approximately 12 to 16 minutes of content.

Evaluation Step 1: Character Consistency Across Sessions (15 minutes)

The most commercially significant quality check in the delivery evaluation is character consistency across the episode run. This is also the check that most commissioning parties skip because it requires side-by-side comparison rather than sequential viewing.

How to do it:

Open episode one on the phone. Pause on the controlled alpha's first close-up. Take a screenshot. Open episode twenty-five. Find the controlled alpha's first close-up. Take a screenshot. Open episode forty-five. Take a screenshot. Compare the three screenshots side by side on the phone's screen.

What to look for: the jaw structure matches across all three. The eye colour matches. The skin tone matches. The nose bridge proportions match. Any visible deviation between episode one and episode forty-five is a character consistency failure.

What it means:

Consistent across all three: the production partner has established and maintained character reference infrastructure across production sessions. This is production-grade character management.

Visible drift between episode one and episode forty-five: the character reference configuration changed between production sessions. This is a character consistency failure that the platform acquisition team will identify in their first five minutes of review. The commissioning party should flag this as a delivery issue requiring correction before acceptance.

Evaluation Step 2: Hook Effectiveness (10 minutes)

The hook effectiveness evaluation tests whether episode one's cold open works for a viewer who has no prior context for the series — which is every viewer who encounters it for the first time on a platform.

How to do it:

Open episode one on the phone. Turn off the audio completely. Watch the first 15 seconds in silence. At the 15-second mark, pause.

The questions: Is a conflict clearly visible? Can you identify the power dynamic between the characters from their physical positioning and facial expression alone? Does the visual content communicate the genre (romance, thriller, revenge) without audio support?

What it means:

Conflict visible in the first frame, legible without audio: the cold open has been designed to the hook detonation standard. The hook rate this episode produces in distribution will likely be above threshold.

Conflict not established by second 15: the cold open requires setup before conflict, which is the most common first-draft cold open failure. The commissioning party should flag this as a revision before acceptance and request a cold open redesign for episode one.

Evaluation Step 3: Audio Intelligibility (20 minutes)

The audio evaluation is the check that most production partners pass on a studio monitor and most commissioning parties fail to test on the actual delivery device.

How to do it:

Part 1 — quiet room test. Play episode one's most dialogue-dense scene on the phone speaker at standard listening volume in a quiet room. Is every line of dialogue clearly intelligible without straining?

Part 2 — ambient noise test. Take the phone into the moderate-noise room. Play the same scene at the same volume. Is the dialogue still intelligible over the background noise?

Part 3 — music-dialogue balance. In the ambient noise test, does the music bed compete with the dialogue or sit beneath it? If the music is audible at the same level as or louder than the dialogue in ambient noise, the mix has not been calibrated for the delivery environment.

What it means:

Intelligible in both tests: the audio post-production was calibrated for phone speaker delivery. Production-grade audio standard.

Intelligible in the quiet room but not in ambient noise: the audio was calibrated for studio monitors or headphones rather than phone speakers. This is the most common audio failure in AI-native vertical drama delivery. The commissioning party should flag this as a delivery issue and request a phone-calibrated audio remix before acceptance.

Evaluation Step 4: Paywall Episode Button Cut (10 minutes)

The paywall episode's button cut is the single most commercially consequential moment in the delivered series. The commissioning party should evaluate it specifically against the brief's specified paywall moment.

How to do it:

Open the paywall episode. Watch from the five-minute mark to the end. At the button cut, pause on the final frame. Compare the final frame against the paywall moment specified in the brief.

The questions: does the button cut land at the maximum unresolved tension moment? Does the final frame match the brief's specified paywall moment, or has the cut landed at a different moment from what was agreed?

What it means:

Button cut at the brief's specified moment with no partial resolution before the cut: the paywall episode has been produced correctly against the brief. The paywall conversion rate this episode produces in distribution will reflect the emotional architecture the brief was designed to create.

Button cut at a different moment from the brief's specification, or with visible partial resolution before the cut: the paywall episode's button cut does not match the agreed production standard. Flag as a delivery issue requiring correction before acceptance.

Evaluation Step 5: Technical Specification Compliance (30 minutes)

The technical specification compliance check confirms that the delivery files meet every specification confirmed in the brief's delivery requirements section.

How to do it:

Check each delivery file against the confirmed specifications: codec and container format, resolution and aspect ratio, frame rate, audio loudness (requires an audio analysis tool such as Resolve's integrated meter), subtitle format and encoding, subtitle positioning on phone display, and metadata structure.

Any file that does not meet the confirmed specification is a technical delivery failure regardless of the content quality. Flag each non-compliant specification as a delivery issue.

What it means:

All files meet confirmed specifications: the delivery package is platform-ready. The commissioning party can accept delivery and proceed to platform acquisition.

Any specification mismatch: the delivery package requires technical remediation before platform submission. Flag the specific mismatch with the specification reference from the brief and the production partner's delivery documentation.

What to Do When You Identify Issues

The delivery acceptance window is the commercial mechanism for requiring correction of delivery failures. Issues identified during the window are delivery failures. Issues identified after acceptance are revision requests.

For each identified issue during the window:

Document the specific failure with episode number, timestamp, and the brief specification it fails. Send the documented failures to the production partner with a request for correction within the remaining acceptance window or a request to extend the acceptance window by the time required for correction.

The production agreement's revision policy governs how many correction rounds are included in the commission fee and what the cost of additional rounds is. Issues identified during the acceptance window are typically addressed within the commission fee rather than as additional-cost revisions, because they are delivery failures rather than scope changes.

Axis AI Studios Perspective

The structured delivery evaluation is the commissioning party's last quality gate before the production agreement is discharged. At Axis AI Studios, the evaluation described in this post is part of the delivery process documentation we provide to every commissioning party at the beginning of the engagement — not at delivery. Understanding what the delivery evaluation will check is the information that shapes production decisions throughout the eight-week timeline.

For businesses commissioning AI-native vertical drama for the first time who want to understand the delivery evaluation before they commission, the brief development session covers the evaluation criteria alongside the brief specifications. Reach out at business@axisaistudios.com.

For generators and operators: the five evaluation steps described in this post are the same quality criteria that should be applied to every submitted output before it reaches the creative director's review queue. The delivery evaluation is the full-series version of the batch quality review.


FAQ

What Happens if the Acceptance Window Expires Before I Complete the Evaluation?

Most production agreements include a provision that the commissioning party can request a reasonable extension of the acceptance window if they identify issues requiring additional review time. Request the extension in writing before the window expires, with specific reference to the issues requiring additional review time. Do not let the window expire without either accepting or formally requesting an extension.

Should I Hire a Specialist to Conduct the Delivery Evaluation?

For a first commission, conducting the evaluation yourself using the five steps in this post is sufficient. The evaluation does not require technical expertise beyond the ability to compare screenshots side by side and use a phone in a moderate-noise environment. For subsequent commissions where the platform relationship creates higher commercial stakes for delivery failures, engaging a post-production specialist with vertical drama format experience to conduct the technical specification compliance check is a reasonable investment.

What if the Production Partner Disputes an Identified Delivery Issue?

The brief's documented specifications are the reference for dispute resolution. A character consistency failure is measurable: the screenshots either show matching jaw structure or they do not. An audio failure is measurable: the dialogue is either intelligible in ambient noise at standard phone speaker volume or it is not. If the production partner disputes an identified issue, request that they demonstrate compliance against the brief specification rather than against their own quality standard.


Further Reading

For the quality markers this post applies as a delivery evaluation, the quality assessment guide for platform buyers covers the five markers and the phone display test in complete detail.

For the delivery documentation that must accompany the content files evaluated in this post, the guide to what vertical drama data rooms look like covers the chain of title documentation, E&O certificate, and AI tool usage disclosure that complete delivery requires.

For the commission checklist that confirms every quality standard was specified before production began, the vertical drama commission checklist covers the pre-signing decisions that make the delivery evaluation productive rather than contentious.

Stay connected

For studios moving beyond traditional production.

Let's set
the new standard together.

If you're working on something, we'd like to hear about it.