What a Platform Should Ask on a Reference Call With an AI Production Partner
Thirty minutes booked. A reference contact supplied by the vendor, who has every reason to supply a friendly one. And a commissioning lead who opens with a question about whether they were happy with the work. The answer is yes. It was always going to be yes. Twenty five minutes later the call ends warmly, nothing has been learned, and the vendor selection decision rests on exactly the same evidence it rested on beforehand.
Reference calls are the only part of vendor due diligence where you speak to someone who has no commercial interest in the outcome. Everything else in the process is controlled by the partner. The deck is theirs, the showreel is curated, the case studies are written by the people who produced them, and even the delivered samples represent chosen work rather than median work. The reference call is the one input the partner cannot fully author, and most commissioning teams waste it by asking questions that can only be answered politely.
1. Why the Reference Call Carries More Weight in AI Native Production
Conventional production references are largely about reliability and temperament, because the underlying craft is legible from the finished work. AI native vertical drama is different in one specific way. A great deal of what determines whether a series is good to commission is invisible in the delivered episodes. Whether character consistency was maintained by system or by rescue work, whether the audio was built for phone playback or fixed after a complaint, whether the continuity documentation exists or was reconstructed at the end, whether the schedule held or absorbed three silent slips. All of that lands as finished video either way.
The reference is the person who watched the process rather than the output. They know how many revision rounds it took, what the partner did when a generation model updated mid production, whether problems arrived as early warnings or as delivery day surprises. Those are the variables that determine what your second and third commission will cost you in internal time, and none of them appear in a portfolio.
The market context makes this sharper. As Sensor Tower sets out in its State of Short Drama Apps 2026 analysis, the category is scaling fast across downloads, revenue and engagement, and a commissioning cycle running at that pace leaves very little room for coordination overhead. A partner who consumes twenty hours a week of your team's coordination time is expensive at any unit price, and that cost is only visible from someone who has already paid it.
2. What to Establish Before the Call Starts
Do the work on the reference sheet first, because two of your thirty minutes should not go on establishing who this person is. Find out their actual role on the engagement, when it ran, what the deliverable was, and whether they were the commissioning decision maker, the day to day contact or a reviewer. Each of those roles sees a different part of the relationship, and a glowing reference from someone who never handled a delivery is not a reference about delivery.
Establish how recent the work is. AI native production capability changes materially over a twelve month period, both in tooling and in team composition. A reference on a project from two years ago describes a company that may no longer exist in the same form. Ask the vendor directly for at least one reference from an engagement completed within the last six months, and treat reluctance on that point as information.
Establish whether the reference is a repeat client. This is the single most compressive fact available to you. A client who commissioned once and stopped is a different signal from a client who commissioned three times, regardless of how positively either speaks. Ask the vendor for the reference's commissioning history with them before the call, then verify it on the call. A discrepancy there matters more than anything else you will hear.
3. Question Set One: Scope and Actual Involvement
Open by mapping what the partner really did, because scope inflation is the most common distortion in vendor case studies. Ask what the partner was responsible for end to end and what stayed with the client team. Ask specifically who wrote the scripts, who did the sound design, who handled the platform delivery specification, and who owned the quality review. Partners frequently present a supervised engagement as a full service one, and the reference will describe the real division without being prompted to criticise anyone.
Then ask what the client team had to supply that they had not expected to supply. This question does more work than any other in the first ten minutes. Nobody experiences it as an attack, everybody has an answer, and the answer tells you exactly where the partner's operating model has gaps. If three references independently mention that they ended up writing the continuity documentation themselves, you have learned something no deck would have told you.
Follow with a question about volume context. How many episodes, over what period, and was this the client's only commission running at the time. A partner who performed well on a single seventy episode series with the client's full attention has not demonstrated the same thing as a partner who performed well on three concurrent series. Scale context is what makes the rest of the reference interpretable.
4. Question Set Two: What Went Wrong and How It Was Handled
Every production has failures. A reference who reports none is either describing a very small engagement or is being polite, and either way the call needs to move past that. The question that reliably opens this up is not what went wrong, which invites a defensive summary, but what surprised them. Surprises are neutral to describe and almost always point at a real gap in process, briefing or expectation setting.
Then ask how they found out. This is the question that separates good partners from adequate ones. The useful answer is that the partner raised it, with a proposed remedy, before it affected the schedule. The concerning answer is that the client discovered it during review, or at delivery, or after the series was published. A partner who conceals slippage until it is unrecoverable will do the same to you, and the reference will describe this pattern plainly if you give them the opening.
Ask what the partner did when a generation tool or model changed mid production. This is close to a universal event in AI native work over any multi month engagement, and how a partner handles it is diagnostic of whether they have real production infrastructure or are assembling each series ad hoc. The answer you want describes version pinning, reference libraries, a controlled migration and a client conversation. The answer that should concern you describes a period of inconsistent output and a discussion about who pays for the retakes.
Close this set by asking whether anything had to be reshot, regenerated or rebuilt at scale, and who bore that cost. Cost allocation on remediation is where commercial character shows up, and it is far easier to learn from a reference than to negotiate blind.
5. Question Set Three: Delivery, Acceptance and the Review Cycle
Move to the mechanics, because this is where your own team's time gets consumed. Ask how many rounds of revision a typical episode batch required before acceptance. There is no universally right number, but there is a very informative range. Consistent first pass acceptance suggests briefs were understood and standards were shared. Four or five rounds as the norm suggests either an unclear brief on the client side or a partner working toward a target they cannot see, and the reference will usually be able to say which.
Ask how acceptance criteria were set and whether they were written down before production started. A reference who describes a documented standard is describing a partner who can be held to something. A reference who describes review as an ongoing conversation about taste is describing a relationship that scaled badly, or would have if it had scaled.
Ask about delivery to platform specification. Whether files arrived correctly formatted, whether subtitle and metadata requirements were met without correction, whether anything was rejected downstream by the platform itself. This is unglamorous and it is exactly the kind of failure that consumes a launch window. It is also the sort of thing a reference remembers vividly, because they were the one who fixed it.
Finish with schedule. Not whether the partner delivered on time, which produces a compressed yes or no, but how far the final delivery date moved from the original agreed date, and how many times it moved. Movement is normal. Undisclosed movement is not.
6. Question Set Four: The Team Behind the Work
Ask who from the partner's side the reference actually dealt with, by role, and whether those people stayed on the engagement throughout. Continuity of team is a large part of what makes a second commission cheaper than a first, and turnover mid production is one of the most reliable predictors of quality drift. If the reference names a coordinator who left halfway through, ask what changed afterwards.
Ask how decisions were made on the partner's side when something needed a judgement call. Whether there was a named person with authority, or whether every question routed through a single founder and waited. Founder bottlenecks are common in production companies at this stage and they are not disqualifying, but they are strongly predictive of what happens when the partner takes on a second client concurrently with yours.
Ask whether the reference would work with the same team again, and then, separately, whether they would work with the same company again if the team were different. The gap between those two answers is one of the most informative things you can extract from a reference call. A client who would rehire the individuals but not the organisation has told you the capability is personal rather than systemic, which is exactly what you need to know before committing to volume.
7. The Closing Question
End every reference call with the same question. Ask what they would do differently if they were commissioning the same work again, from the same partner, starting tomorrow.
It works because it is not a question about the partner at all. It is a question about the reference's own decisions, which people answer candidly and at length. What comes back is usually a precise inventory of everything that did not go smoothly, delivered without any of the social friction of criticising a vendor. Briefs they would write differently, gates they would insert, things they would have specified in the agreement, review capacity they would have staffed. Take notes on that answer specifically, because it is a free draft of the risk register for your own engagement.
8. How to Read Hedging, Silence and Enthusiasm
Interpretation matters as much as the questions. Specific enthusiasm is credible and general enthusiasm is not. A reference who says the partner was excellent has told you nothing. A reference who says the character work held across sixty episodes without a rescue pass, and names the moment they realised it, has told you something verifiable and hard to fabricate.
Hedging clusters around real problems. Watch for the shift from concrete detail to abstraction, which usually happens exactly at the point of difficulty. When a reference has been giving precise answers and suddenly offers a general one, follow up once with a neutral request for an example. If the second answer is also abstract, note the topic and move on. You now know where the soft area is, and you can address it in the agreement rather than the call.
Pay attention to what is volunteered without prompting. References who mention responsiveness unprompted are usually signalling that responsiveness was notable, in either direction. References who describe the relationship rather than the output are often compensating for the output. And a reference who asks you what you are planning to commission, then offers advice about how to structure it, is giving you the most valuable thirty seconds of the call.
Finally, weigh a lukewarm reference heavily. Vendors supply references they expect to be positive. A merely adequate reference from a hand picked contact is a stronger negative signal than an enthusiastic one is a positive signal, because the selection bias runs entirely in one direction.
Axis AI Studios Perspective
Axis AI Studios produces AI native vertical drama for platforms, brands and IP holders, and we are on the receiving end of this process regularly. Our view is that a commissioning team should run it hard. The questions above are the ones we would want a buyer to ask, because the answers to them are where the difference between production companies in this format actually sits.
What we hold ourselves to is specific and checkable. Character consistency maintained by system rather than by rescue work. Continuity documentation that exists during production rather than being assembled at delivery. Audio built for phone playback as a standard rather than a correction. Delivery to platform specification without downstream rejection. Problems surfaced by us, early, with a proposed remedy attached. Those are the things within our control, and they are the things a reference should be able to speak to without hesitating.
We also think the closing question in section seven is the right one for buyers to keep asking, including of us. A production partner who cannot tolerate a client describing what they would do differently is a partner who has stopped improving. Recognition of that is part of how a working relationship survives a second and third commission.
If you are running a vendor selection process for AI native vertical drama and want a conversation before or after the reference stage, reach us at business@axisaistudios.com.
FAQ
How many reference calls should a platform run before selecting a partner? Three is the practical minimum, because patterns only become visible across multiple calls and a single reference is too easily an outlier in either direction. Where possible, ask for one reference from a completed engagement, one from a repeat client, and one from a client whose commission was similar in scale to yours. If a partner can only supply one, that constraint is itself a finding.
Should the vendor be told what will be asked? There is no advantage in surprise, and sharing the themes in advance tends to produce better references rather than more coached ones, because the vendor selects contacts who can actually speak to process. What you should not do is share the specific closing question, since its value comes from the reference answering it without preparation.
What if the only references available are under confidentiality restrictions? This is common with platform clients and does not have to end the conversation. Ask the reference to speak to process rather than to the title, its performance or its commercial terms. Revision cycles, escalation behaviour, team continuity and delivery reliability can all be discussed without identifying the work, and those are the areas where the reference call earns its place anyway.
Further Reading
For the structured evaluation that should sit around the reference stage, the due diligence checklist for AI vertical drama production partners covers the six capability areas worth assessing before a first commission.
For converting impressions into evidence across multiple partners, the supplier scorecard framework sets out the performance measures to track and how to weight them by commission type.
For the wider selection process a reference call sits inside, the platform buyer guide to evaluating AI native production companies covers the five domains acquisition teams should assess.

Let's set
the new standard together.
If you're working on something, we'd like to hear about it.
