The Second-Screen Phenomenon: How Viewers Watch Vertical Drama Alongside Other Activities
86% of internet users use another device alongside TV. Smartphones are the most popular choice. People use their second screens for social media browsing, looking up information, messaging friends about content, and shopping for products featured in what they are watching. Datarmatics
This research describes the second-screen phenomenon as it applies to television: the viewer is on the couch watching a primary screen with a phone as the secondary device. Vertical drama inverts this entirely. The phone is the primary screen. The television, if it is on, is the ambient background. The commute, the break room, the kitchen, or the late-night bedroom are the viewing environments. The ambient noise, the partial attention, and the activity-adjacent consumption that television describes as a problem are conditions that vertical drama was designed to operate within.
The format assumes fragmented attention, continuous scrolling, and the need to capture interest instantly — conditions that reshape how stories are written and experienced. Those conditions are not a limitation the format overcomes. They are the design brief the format was built from.
Understanding precisely what activity-adjacent viewing conditions exist for vertical drama, how those conditions affect what the hook must do in the first three seconds, and what production decisions serve a viewer whose attention is partially elsewhere is the audience research that most production conversations treat as background rather than as production specification. It is not background. It is the specification.
The Viewing Environments and What Each Requires
Vertical drama's audience does not watch in a single viewing environment. We stopped watching live TV in the same way we did before. You turn on a movie or a match, and your smartphone is already active. Psychologists explain that the multi-layered experience of content consumption originated from human desire for a quick dopamine rush multiplied by the technical ability to avoid being limited to linear, concentrated viewing.
Applied to vertical drama specifically, the research on viewing environment distribution reveals five primary contexts that each impose specific requirements on the content.
The commute context. The viewer is in transit: on a train, a bus, or walking. Ambient noise is high, between 60 and 80 decibels in transit environments. The phone is held at arm's length in portrait orientation. The viewing is interrupted by arrivals, departures, and social interactions. Session length per uninterrupted viewing window is two to eight minutes.
The production requirement: the hook must communicate conflict, genre, and emotional register without audio in the first three seconds, because audio is frequently inaudible or unavailable. The episode's emotional content must be readable from facial expression and body language in the close-up frame without dialogue support.
The break context. The viewer is on a work break: lunch, a fifteen-minute rest, a moment between tasks. The environment is quieter than transit but still ambient. The viewing is time-constrained and the viewer knows it. Session length is five to fifteen minutes.
The production requirement: the episode pacing must allow satisfying emotional engagement within the break's window without requiring the viewer to watch to a natural stopping point that does not exist. Every episode's button cut is a natural stopping point, but the button cut must create sufficient discomfort that the viewer chooses not to stop there even when the break is ending.
The domestic evening context. The viewer is at home, in a relaxed environment, often with another screen active in the background. Audio is available but the viewing is still interrupted by domestic activities. Session length is fifteen to forty minutes.
The production requirement: sustained emotional investment across multiple episodes in a single session. The series' character investment and emotional debt architecture must sustain engagement across a fifteen to forty-minute session without requiring the focused attention that the earlier viewing contexts do not provide.
The pre-sleep context. The viewer is in bed, phone at close range, often with the screen brightness reduced. This is the highest-attention vertical drama viewing context because there are fewer competing activities. Session length can extend beyond forty minutes when the content's emotional investment is strong.
The production requirement: the paywall conversion rate is highest in this context because the viewer's attention is least divided and their emotional investment is most concentrated. The paywall episode's button cut must be designed for a viewer who is in the highest-investment state they have been in across the full series.
The incidental discovery context. The viewer encounters vertical drama content through social media feeds while performing another primary activity. They are not in a viewing session. They are scrolling. The content has approximately three seconds to change their activity from scrolling to watching.
The production requirement: the hook in the first frame must stop the scroll without any genre context, platform context, or prior character investment. This is the most demanding hook requirement in any vertical drama viewing context and the context where the three-second silent test is most commercially critical.
What Activity-Adjacent Viewing Means for Hook Design
The hook's commercial function is to stop the scroll and hold the viewer through the episode. In a dedicated viewing session where the viewer has chosen to open the app, the hook's function is holding. In an activity-adjacent context where the viewer has not chosen to engage with the content, the hook's function is stopping first.
A scene becomes the hook. Not the trailer. Not the poster. Not even the premise. Television is increasingly behaving like social media, not the other way around. Short, emotionally salient bursts trigger curiosity and anticipation, pulling the user deeper into the platform ecosystem.
The activity-adjacent hook must be designed for the specific attention conditions of the viewing environment it is most likely to be encountered in. The hook that works in the pre-sleep context, where the viewer's attention is relatively concentrated, is not necessarily the hook that works in the commute context, where the viewer's attention is divided and audio may not be available.
The specific hook design decisions that serve activity-adjacent viewing across multiple contexts simultaneously:
Conflict in the first frame. Not conflict approaching. Not context before conflict. Conflict visible and legible in the first visual frame. A viewer in the commute context who glances at their phone for one second while waiting for a train stop must encounter conflict rather than setup in that one second, or the content has failed to engage them in the only moment available.
Physical behavior over stated emotion. Activity-adjacent viewing is frequently audio-limited. The hook that communicates through physical behavior in close-up, through facial expression, body language, and specific physical tells, is a hook that works without audio. The hook that communicates primarily through dialogue or sound design does not work in the 30 to 40% of viewing instances where audio is unavailable or inaudible.
Status legibility without context. The power dynamic between characters must be readable from visual cues alone in the first three seconds, without any prior narrative context. A viewer who encounters vertical drama content for the first time in a social media feed has no character context, no series context, and no platform context. The power dynamic is the only content that can communicate instantly without any of those contexts.
What Activity-Adjacent Viewing Means for Audio Requirements
The audio requirement post covers the technical floor: LUFS targets, phone speaker frequency response, dialogue intelligibility in quiet environments. Activity-adjacent viewing extends these requirements to environments that the technical floor does not fully address.
To maintain engagement, a show must keep the plot highly dynamic, vary visuals, and use different angles and scales. While this is easily achievable on TikTok, it is often not feasible in a TV show format. Consequently, the audience's attention wanders. Applied to audio: in activity-adjacent viewing contexts, the score and effects mix must not attempt to compete with ambient noise. They must be calibrated to provide emotional support to the visual content at the level of intelligibility that survives ambient noise, rather than at the level that a quiet listening environment allows.
The specific audio production decisions that serve activity-adjacent viewing:
Dialogue priority over score at all ambient noise levels. In a commute environment where ambient noise is 70 dB, a dialogue track mixed at the phone speaker's maximum output and a score mixed at half that level produces a practical intelligibility ratio. A score mixed at 80% of the dialogue level in that environment produces competition that makes the dialogue partially unintelligible.
Emotional content in the upper frequency range. Phone speakers reproduce 200 Hz to 8,000 Hz most effectively. Emotional score content that is concentrated in the 500 Hz to 4,000 Hz range survives ambient noise competition better than content with significant sub-200 Hz emotional weight that the phone speaker cannot reproduce at meaningful output.
Sound design that communicates visually concurrent events. Sound effects that reinforce the visual content happening simultaneously in the frame are more useful in activity-adjacent viewing than sound effects that communicate events happening off-frame or implicitly. A viewer whose attention is divided between the phone and their activity cannot process implicit audio storytelling the way a focused viewer can.
The Comment Section as a Second-Screen Activity
41% of viewers text or message friends and family about the content they are watching. 35% shop for products featured in the content. These are second-screen activities for television viewers. For vertical drama viewers, the equivalent activities happen within the same device: comment section participation, character discussion, and plot prediction happen on the same phone that the episode is playing on. Datarmatics
The comment section engagement that vertical drama's algorithmic distribution rewards, and that parasocial character design generates, is a second-screen activity performed on the first screen. A viewer who is watching episode twenty-five and simultaneously reading and responding to comment section discussion of the episode is a viewer whose engagement with the content has extended beyond passive viewing into active social participation.
This social participation is a production design goal, not an accidental byproduct. The character design decisions that generate comment section parasocial investment, the specific involuntary tells, the suppressed interior, the unearned vulnerability moments, are the design decisions that generate the comment section social participation that the algorithm rewards and that extends the viewing session beyond what passive viewing alone would sustain.
What Activity-Adjacent Viewing Means for Platform Strategy
For businesses and platforms commissioning AI-native vertical drama, the activity-adjacent viewing research has specific implications for distribution strategy that go beyond production design.
Stop fighting for undivided attention and start building strategies around predictable fragmentation. Viewers will use their phones during your content. Plan for it. Benefit from it. Applied to vertical drama commissioning: the series that assumes focused viewing in a quiet environment has made incorrect assumptions about its audience. The series that is designed for the specific attention and audio conditions of the commute, break, domestic evening, pre-sleep, and incidental discovery contexts is a series that has been designed for its actual audience.
The commissioning brief that specifies these viewing context assumptions is a brief that produces content calibrated to the actual distribution environment rather than to an idealised viewing condition that applies to only a minority of the audience's actual viewing sessions.
Axis AI Studios Perspective
Activity-adjacent viewing is the default condition for vertical drama. The production that does not account for it is producing for a minority of its audience's actual viewing experience. At Axis AI Studios, the hook design specification, the audio calibration standard, and the character behaviour specification all derive from the activity-adjacent viewing conditions that the audience research describes.
The hook is designed for three-second silent communication. The audio is calibrated for phone speaker output in ambient noise at 60 to 70 dB. The character tells are designed to communicate physical emotional register in close-up without audio support. These are not advanced production considerations. They are the baseline production requirements for a format whose audience is watching in the commute, the break room, and the kitchen.
For platforms and brands who want to commission AI-native vertical drama designed for the actual viewing conditions of the format's audience rather than for idealised conditions, reach out at business@axisaistudios.com.
FAQ
Does Activity-Adjacent Viewing Reduce Paywall Conversion Rates?
Not necessarily. The pre-sleep context, which is one of the five primary activity-adjacent viewing contexts, is the highest-attention context and correlates with the highest paywall conversion rates because the viewer's emotional investment is most concentrated and the competing activities are minimal. The commute and break contexts produce lower per-session paywall conversion but higher session frequency: the viewer who watches one episode on every commute for five days has generated five viewing sessions where a focused viewer watching in one evening session has generated one. The cumulative paywall conversion across multiple activity-adjacent sessions may exceed the single-session focused viewing conversion.
Should Production Companies Create Different Content Versions for Different Viewing Contexts?
No. The same content should serve all five viewing contexts through correct hook design, audio calibration, and character behavior specification. A production that requires different versions for different viewing contexts has not solved the activity-adjacent viewing design problem. It has created a production cost problem as the solution. The three-second silent hook, phone speaker audio calibration, and close-up physical behavior specification are the single-version production decisions that serve all five viewing contexts simultaneously.
How Does the Incidental Discovery Context Affect Series Structure for Social Media Distribution?
The incidental discovery context requires each episode to function as a standalone hook rather than as part of a sequence that assumes prior viewing. A viewer who encounters episode fifteen in a social media feed has no prior context for the characters or the narrative. The episode's first frame must communicate the power dynamic without that prior context. This does not require episodes to be self-contained narratively. It requires the hook to communicate the power dynamic from visual content alone, which serves the series' narrative continuity for existing viewers while also serving the incidental discovery context for new ones.
Further Reading
For the hook design specifications that the incidental discovery and commute contexts described in this post require, the guide to AI-assisted beat sheets and cold opens covers the three-second silent test, the conflict-in-first-frame requirement, and the hook testing process.
For the audio calibration that activity-adjacent ambient noise conditions require, the guide to mixing audio for phone speakers covers the LUFS targets, frequency response limits, and dialogue priority specification that all five viewing contexts require.
For the character design decisions that generate the comment section social participation that activity-adjacent viewing enables, the guide to why some characters become parasocial figures covers the specific involuntary tells and suppressed interior design that produce comment section engagement from partially-attentive viewers.

Let's set
the new standard together.
If you're working on something, we'd like to hear about it.
