The Vertical Drama Music Brief: How to Commission Score That Works on Phone Speakers

Most vertical drama scores are written by composers who have scored for cinema or television. The music sounds correct on studio monitors. It sounds wrong on the delivery device.

The problem is not the composer's craft. It is the brief. A conventional score brief tells the composer the emotional register, the tone, the genre references, and the cue-by-cue placement. It does not tell the composer that the primary playback device has an effective frequency response that begins at 200 Hz, that the listening environment contains ambient noise concentrated in the 200 to 800 Hz range that will compete with the low string register the composer considered their most effective emotional tool, or that the 90-second episode's emotional debt architecture requires the score to support the button cut's maximum unresolved tension rather than resolving it.

These are the specifications that determine whether the score works on the delivery device for the audience who watches the series. They are also the specifications that no conventional score brief contains.

This is the complete vertical drama music brief.

The Phone Speaker Constraint: What It Means for Composition

The phone speaker constraint covered in the audio mixing guide establishes the technical floor for audio delivery. The music brief extends that technical floor into compositional requirements: if the composer cannot use certain frequency ranges effectively on the delivery device, those ranges should not be the score's primary emotional carriers.

Bass amplification significantly increases emotional arousal. This research finding is commercially significant for cinema and live music. It is mostly irrelevant for vertical drama, where the delivery device cannot reproduce bass below 150 to 200 Hz at meaningful acoustic output. The emotional arousal that low-frequency content produces in cinema and live performance is not available to the vertical drama score because the delivery device eliminates it.

The practical consequence: the score cannot rely on low strings, bass instruments, or sub-bass texture as its primary emotional tools. These elements are compositionally available but commercially ineffective because the listener cannot hear them on their phone speaker in ambient light at arm's length.

The frequency ranges that the phone speaker reproduces effectively for music are: the mid-range from 200 Hz to 4,000 Hz, where the phone speaker's efficiency is highest, and the upper mid to high-mid range from 4,000 Hz to 10,000 Hz, where the phone speaker maintains reasonable output. The score that concentrates its emotional content in these ranges is a score that the delivery device can reproduce.

What this means for instrumentation:

Piano is the most phone-speaker-effective instrument for vertical drama scoring. Its frequency content is concentrated in the 200 Hz to 8,000 Hz range, its attack transients are clear at phone speaker distances, and its sustain characteristics create the harmonic tension that vertical drama's emotional debt architecture requires. A piano-led score is not a budget compromise. It is the instrumentally correct choice for the delivery device.

Solo strings, particularly violin and viola in their upper register, reproduce effectively on phone speakers because their fundamental frequencies and primary harmonics sit within the 200 Hz to 4,000 Hz effective range. Cello in its lower register loses effectiveness because the fundamental pitches drop below the phone speaker's effective output range. String arrangements for vertical drama should concentrate melodic content in the violin and upper viola range, using cello primarily as harmonic texture in the 400 Hz to 800 Hz range where phone speakers still have meaningful output.

Synthesizer pads work effectively when filtered above 200 Hz. Unfiltered synth pads that include sub-bass content are partially inaudible on phone speakers, which alters their harmonic character in ways the composer did not intend. The brief should specify that all synth pad elements be high-pass filtered at 120 Hz before mix delivery.

Percussion and rhythmic elements work effectively when the attack transients are in the 1,000 Hz to 5,000 Hz range. Kick drum fundamentals below 200 Hz are inaudible on phone speakers. A rhythmic tension element built from kick drum will lose half its character on the delivery device. A rhythmic tension element built from mid-range percussion, rim shots, high-tempo string pizzicato, or synthesiser arpeggiation in the 1,000 Hz to 3,000 Hz range, retains its character on the delivery device.

Competing With Ambient Noise: The Compositional Problem

The listening environment for most vertical drama is not a quiet room. It is a commute, a break room, a domestic evening with other sounds present. Ambient noise occupies the frequency range below 1,000 Hz primarily, with traffic and HVAC noise concentrated between 200 and 500 Hz and human speech ambient noise concentrated between 300 and 3,000 Hz.

The score element that sits in the same frequency range as the ambient noise the listener is surrounded by is the score element that the listener cannot hear clearly. A low string texture at 200 to 400 Hz is in direct competition with the traffic noise on the commuter's train. A piano melody at 1,500 Hz is above most ambient noise's primary concentration and is significantly more audible.

The ambient noise competition problem has a specific consequence for the score's relationship with the dialogue: dialogue intelligibility lives in the 1,000 to 4,000 Hz range where consonant sounds are concentrated. A score that is acoustically dense in this range competes with dialogue comprehension. The listener hears music or dialogue, not both, when both are competing for the same frequency space.

The vertical drama music brief's ambient noise specification: the score must leave the 1,000 to 3,000 Hz frequency space available for dialogue during any scene where dialogue is present. This is not a general suggestion. It is a compositional constraint that requires the score to be acoustically transparent in the dialogue frequency range rather than harmonically full.

The practical implementation: the score during dialogue scenes is mixed below the dialogue by a specific dB margin, and any score elements in the 1,000 to 3,000 Hz range are notched 2 to 3 dB during dialogue delivery through dynamic EQ triggered by the dialogue stem's presence. This is the notching technique described in the audio mixing guide, applied at the compositional stage rather than only at the mixing stage: if the composer writes score elements that are not in the dialogue frequency range, the mixing notch is less necessary and the score retains more of its musical character in the final mix.

The Emotional Debt Architecture: How Music Supports It

The vertical drama episode's emotional debt architecture has four structural positions: hook at 0 to 15 seconds, escalation at 15 to 60 seconds, spike at 60 to 80 seconds, and button cut at 80 to 90 seconds. Each position has a different musical requirement from the score.

The hook position (0 to 15 seconds).

The hook's musical requirement is orientation without resolution. The viewer encounters the episode in an established conflict or emotional state. The score's job in the hook position is to confirm the emotional register before the content has established it, not to create a new emotional experience.

The musical character that serves this function: a sustained harmonic texture, 3 to 5 seconds in duration, that establishes the emotional register without providing melodic forward motion. A long-attack piano chord with a pad layer. A sustained string tone in the series' characteristic harmonic language. The texture tells the viewer's emotional system what register they are in. It does not move within that register during the hook position because movement creates expectation that the hook's visual content needs space to create independently.

The escalation position (15 to 60 seconds).

The escalation's musical requirement is forward motion support. The episode's one forward move occurs in this range, and the score's job is to support the viewer's sense that the situation is advancing without providing the emotional release that would satisfy the tension rather than build it.

The musical character that serves this function: gradual harmonic tension increase through suspended chords, added harmonic intervals that lean toward the next resolution without arriving, or a melodic line that ascends without reaching its peak. The Picardy third as a false resolution and the suspended fourth as an unresolved harmonic moment are both compositional tools for creating forward motion that does not release. The escalation position score supports the forward move's energy without completing it.

The spike position (60 to 80 seconds).

The spike's musical requirement is maximum tension without resolution. The score must support the episode's highest emotional intensity moment, hold that intensity for 10 to 20 seconds, and then hand off to the button cut without providing any harmonic resolution that would soften the cut's unresolved tension.

This is the most compositionally demanding position in the vertical drama score. The conventional score's instinct at the dramatic peak is to provide a musical climax that resolves the tension into a satisfying emotional moment. The vertical drama score must resist this instinct completely. A musically resolved spike produces a viewer who feels satisfied at second 78. A viewer who feels satisfied at second 78 does not experience the button cut at second 84 as maximally unresolved.

The musical character that serves the spike position: sustained maximum harmonic tension using specifically unresolved intervals. The tritone, the major seventh, the suspended ninth held without resolution. Dynamic density increase through added harmonic layers rather than through melodic peak. The spike score is harmonically dense and dynamically loud, but it is harmonically arrested: it is not going anywhere because going anywhere would release the tension that the button cut needs to remain.

The button cut (80 to 90 seconds).

The button cut's musical requirement is abrupt cessation rather than fade or resolution. The score stops with the visual cut. It does not fade. It does not resolve. The specific moment of the score's ending is the same moment the picture cuts, and the silence that follows the cut is the listener's first experience of the unresolved tension's full weight.

The brief instruction for the button cut: compose to the picture cut, not past it. Any score note that extends past the button cut frame is a score note that softens the cut's emotional impact. The brief should specify the exact timecode of the button cut for every episode and require the composer to confirm that no score element extends past that timecode.

AI Music Generation for the Vertical Drama Score

AI music generation tools including Suno and Udio have reached the capability level where they can produce functional vertical drama score elements when briefed with the specific constraints described in this post.

Suno v5 introduced Musical Memory, reducing the structural amnesia of previous models where the opening motif would be forgotten by the final section. Musical memory across a short cue is the specific capability that vertical drama scoring requires: a 90-second episode cue that maintains harmonic consistency from the hook through the spike is a cue that requires musical memory across the full duration.

The AI music generation brief for vertical drama applies the same five structural specifications described above in the generation prompt:

Hook position: sustained texture, no melodic forward motion, 3 to 5 seconds, phone-speaker-optimised instrumentation.

Escalation position: gradual harmonic tension increase, forward motion without resolution, suspended intervals, 45 seconds.

Spike position: maximum harmonic tension, no resolution, dynamic density increase, 15 to 20 seconds.

Button cut: abrupt cessation at the specified timecode, no fade, no resolution in the final 5 seconds.

Frequency constraint throughout: no primary melodic or harmonic content below 200 Hz, primary emotional content in the 500 Hz to 5,000 Hz range, dialogue frequency space 1,000 to 3,000 Hz left acoustically available.

AI-generated cues that meet these specifications can function as the series' base music library alongside human-composed cues. The AI is most useful for generating the variations within established cue categories that prevent listener fatigue across 70 episodes, while human composition is most valuable for the specific cues at the highest-consequence episode positions: the paywall episode's spike and button cut, the midpoint reversal's spike, and the resolution sequence.

The Complete Music Brief Template

The vertical drama music brief delivered to any composer, human or AI, contains the following eight specifications:

1. The series sonic identity context. The character of the series' established sound, the recurring motifs, and the harmonic language the score must be consistent with. If a sonic identity document has been produced, it is attached in full.

2. The episode position mapping. The four structural positions with their timecode ranges and their specific musical requirements as described above.

3. The instrumentation brief. The approved instrumentation palette for the series, specified in terms of the frequency ranges the delivery device can reproduce. Piano, upper strings, filtered synthesiser, mid-range percussion. Approved modifiers: reverb characteristic for the intimacy scene register, resonance character for the authority scene register.

4. The ambient noise constraint. The frequency ranges the score must not dominate during dialogue: 1,000 to 3,000 Hz. The dynamic margin the score maintains below dialogue level when both are present: minimum 6 dB below dialogue peak during any dialogue scene.

5. The resolution prohibition. The button cut's mandatory abrupt cessation. The spike position's prohibition on harmonic resolution before the cut. The escalation position's prohibition on premature resolution before the spike begins.

6. The harmonic language specification. The series' specific harmonic vocabulary: the intervals that carry the emotional register, the suspended chords that carry the unresolved tension, and the harmonic territory the score does not enter because it would conflict with the established character.

7. The phone speaker delivery specification. High-pass filter at 120 Hz on all elements. No primary musical content below 200 Hz. Primary emotional content concentrated in 500 Hz to 5,000 Hz. All cues deliverable within the series' LUFS specification alongside the dialogue and effects stems.

8. The cue library structure. The number of base cues required for each episode position category, the acceptable variation range within each category, and the AI variation parameters if AI generation is used alongside human composition.

Axis AI Studios Perspective

The vertical drama music brief is the specification that converts a score from something the composer made into something the delivery device can reproduce for the audience it was made for. The composer who receives a conventional brief and produces a cinematically excellent score has not been briefed for vertical drama. The composer who receives the brief described in this post produces a score that works commercially on the device where the work is being watched.

At Axis AI Studios, the music brief is produced alongside the visual style guide and the sonic identity document in pre-production. The instrumentation palette is approved against the phone speaker delivery specification before any composition begins. The episode position mapping is provided to the composer with timecodes from the arc map. The resolution prohibition is explicit in the brief so that the composer's instinct toward harmonic satisfaction at the dramatic peak does not override the button cut's commercial requirement.

For production companies who want to commission AI-native vertical drama with a music brief that produces score elements that work on the delivery device rather than on studio monitors, reach out at business@axisaistudios.com.


FAQ

Should Every Episode Have Original Composed Score or Can a Series Use a Licensed Music Library?

A licensed music library is commercially viable for the escalation and lower-priority positions but is not correctly structured for the button cut and spike positions without significant editing. Licensed cues are composed to resolve. The button cut requires abrupt cessation at a specific timecode. The production company that uses licensed music must edit every cue at the button cut timecode to prevent the licensed track's natural resolution from softening the commercial impact of the cut. Original composition or AI generation against the brief is preferable because the composer can build the abrupt cessation into the cue from the start rather than requiring post-production editing.

Can the Same Score Cues Be Used Across Different Episodes?

Yes, within the cue library system. The authority hook cue that works in episode five's hook position also works in episode twenty-five's hook position because the hook position's musical requirements are the same regardless of the episode's narrative content. The cue library approach produces this cross-episode reuse as its intended design. The variation requirement, producing enough cue variation to prevent listener fatigue across 70 episodes, is addressed through AI variation within established cue categories rather than through entirely new cue composition per episode.

How Long Should Each Cue in the Base Library Be?

Hook cues: 5 to 12 seconds to cover the hook position's full duration with room for timing variation. Escalation cues: 45 to 55 seconds. Spike cues: 15 to 25 seconds. Button cut cues: the button cut's actual duration, typically 3 to 7 seconds, with an abrupt ending composed to the cut rather than to a natural conclusion. Each cue category should have at least three to five variations to cover the 70-episode run without direct repetition within a five-episode range.


Further Reading

For the phone speaker mixing technical foundation that this post's composition requirements are built on top of, the guide to mixing audio for phone speakers covers LUFS targets, frequency response limits, the dialogue notching technique, and the ambient noise competition problem in full technical detail.

For the sonic identity document that the music brief's harmonic language specification feeds from, the sonic identity guide covers recurring motifs, character audio tells, and the series music library architecture that the base cue library described in this post populates.

For the button cut mechanics that the resolution prohibition in this post's brief is designed to protect, the cliffhanger placement and pay conversion guide covers the structural decisions that determine whether the button cut drives coin purchase or comfortable exit.

Stay connected

For studios moving beyond traditional production.

Let's set
the new standard together.

If you're working on something, we'd like to hear about it.