Reference — glossary

Video editing glossary

42 terms that come up in every serious conversation about short-form and long-form video — defined with the numbers that matter, not dictionary padding. Written so you can hand it to a client, a new hire, or whoever is editing your video this month.

Last updated August 26, 2026

Retention & metrics

Hook rate

The percentage of viewers who are still watching a few seconds after a video starts, used as the primary signal that an opening line or frame is working.

Hook rate is usually calculated from the platform's own retention graph as the share of plays still active at the 3-second or 5-second mark, depending on which the platform surfaces natively. It is a leading indicator, not a vanity metric: a low hook rate means the algorithm sees a video as unwatchable before it has a chance to build momentum, so distribution is capped early regardless of how strong the middle or ending is.

A hook rate above roughly 60-70% at the three-second mark is commonly cited as healthy for short-form, though the number that matters is relative to a creator's own baseline, not an absolute industry figure. Editors treat anything meaningfully below that baseline as a signal to test a new opening line, visual, or on-screen text rather than to fix pacing further into the timeline.

In practice, hook rate drives what gets cut first: editors will often build 3-5 alternate openings for the same video and let the first few seconds carry a claim, a visual disruption, or a question before any logo, intro animation, or slow build-up. Anything that delays the first meaningful frame — including a title card — tends to depress this number.

Hook Structures That Actually Retain Viewers

Three-second drop-off

The rate at which viewers abandon a video within the first three seconds, the steepest and most consequential drop point on most retention curves.

Every retention curve has a characteristic cliff in the opening seconds, and the three-second mark is the point most platforms use internally to decide whether a video merits further distribution. Drop-off here is measured as the inverse of hook rate: if 35% of viewers leave by second three, that is the three-second drop-off.

This number is disproportionately sensitive to the very first frame, not the first sentence. Editors will scrub through a video's opening half-second looking for a black frame, a slow fade, or a static shot that gives the viewer nothing to process, since even a fraction of a second of dead air at the start measurably increases early exits.

Because this is the steepest part of the curve, it is also where small edits produce outsized returns: trimming a run-up, replacing a wide static shot with movement, or adding text on screen in the first frame can shift the drop-off rate more than any change made later in the timeline.

How to Read a Retention Curve

Retention curve

A graph, provided by platform analytics, showing what percentage of viewers are still watching at every point in a video's runtime.

The retention curve plots percentage of audience remaining on the y-axis against elapsed video time on the x-axis. Nearly every curve trends downward, but the shape — not just the endpoint — is what an editor reads: a smooth, gradual decline suggests even pacing, while sharp vertical drops at specific timestamps point to identifiable problems such as a slow segment, a jarring cut, or a plot reveal that lets viewers feel satisfied enough to leave.

Reading the curve well means correlating drop points with the actual timeline. A dip at 0:22 might line up with a transition into B-roll that runs too long, or a shift in speaking pace. Editors will often overlay timestamps from the curve directly onto the edit to isolate the exact cut responsible, rather than guessing from the finished video alone.

Good retention curves for short-form content tend to hold flat through the middle and sometimes tick upward near the end if a payoff or loop is well constructed. A curve that declines steadily with no recovery usually indicates the content front-loaded all its value and gave the viewer no reason to stay for the close.

How to Read a Retention Curve

Average watch time

The mean number of seconds or minutes viewers spend watching a video, calculated across all plays including partial views.

Average watch time is total watch time across all plays divided by the number of plays. Unlike watch-through rate, it is expressed in absolute time rather than a percentage, which makes it useful for comparing videos of different lengths but less useful on its own for judging whether a specific edit is holding attention proportionally.

A 90-second video with 45 seconds of average watch time and a 3-minute video with 45 seconds of average watch time both post the same raw number but represent very different outcomes — the first held half its audience to the end, the second lost most viewers a quarter of the way through. For that reason, editors typically look at average watch time alongside video length rather than in isolation.

This metric is most useful for tracking whether a specific structural change — moving a key point earlier, tightening a middle section — increases total attention across a large batch of similar videos, since it smooths out noise better than any single retention curve does.

Watch-through rate

The percentage of a video's total runtime that the average viewer watches, expressed as a proportion rather than a raw time value.

Watch-through rate is average watch time divided by total video length, giving a normalized figure that can be compared across videos of different durations. A 30-second clip with a 90% watch-through rate and a 3-minute video with a 90% watch-through rate represent comparably strong performance, even though the underlying raw watch times differ enormously.

This is the metric editors default to when deciding on ideal video length for a given format: if watch-through rate drops sharply past a certain runtime across a channel's history, that length becomes a practical ceiling for future edits, regardless of how much material is available.

It also interacts directly with pacing decisions inside the edit — a video padded with unnecessary establishing shots or repeated points will show a watch-through rate lower than a tightly cut version of the same content, even if both cover identical material.

Loop rate

The rate at which a short-form video replays automatically because viewers do not swipe away, often used as a proxy for how well an ending connects to the beginning.

Loop rate matters specifically on vertical, autoplay-looping formats such as Reels and TikTok, where a video that ends and restarts without a swipe accrues additional watch time and signals strong engagement to the distribution algorithm. It is typically inferred from average watch time exceeding the video's total length, since that is only mathematically possible if some portion of the audience watched a second pass.

A high loop rate is engineered, not accidental: editors deliberately end a video on a frame, sound, or line that flows naturally back into the opening, sometimes matching the final shot's composition or audio cue to the first shot so the seam is imperceptible on replay.

This differs from watch-through rate in that a video can have a mediocre watch-through rate but a strong loop rate if the ending is compelling enough to pull viewers back in, even after they were tempted to leave earlier in the runtime.

Editing craft

Cut density

The number of cuts per minute in an edited video, a rough proxy for pacing that varies deliberately by platform and content type.

Cut density is calculated simply as total cuts divided by runtime in minutes. Short-form vertical content aimed at younger audiences commonly runs 15-30+ cuts per minute, while long-form interview or podcast-style content might sit at 4-10 cuts per minute without feeling slow, because the spoken content itself is doing the work of holding attention.

Higher cut density is not inherently better; it is a tool matched to format and message. A tutorial with dense visual information benefits from frequent cuts to new angles or on-screen text, while a story-driven piece can lose emotional weight if it is cut so aggressively that no single shot is allowed to land.

Editors track cut density per project as a consistency check — if one video in a series runs noticeably denser or sparser than its siblings without a content-driven reason, it usually reads as inconsistent pacing to a returning audience even if they cannot articulate why.

J-cut

An edit where the audio from the next shot begins before its corresponding video appears, named for the shape it makes on a timeline.

In a J-cut, the audio track leads the video track: a viewer hears the next speaker or sound effect while still looking at the previous shot, then the picture cuts to match. On a timeline, the audio clip extends left of the video cut point, forming a shape resembling the letter J.

This technique is used constantly in interview and documentary editing to soften a hard cut between speakers or locations — hearing a new voice start before seeing the new face primes the viewer for the transition and makes it feel conversational rather than abrupt.

It is one of the most reliable tools for reducing perceived cut density without actually reducing the cut count, since the overlapping audio smooths the transition psychologically even though the visual cut still happens at the same rate.

L-cut

An edit where the audio from the current shot continues into the next shot's video, the mirror image of a J-cut.

In an L-cut, the video cuts to a new shot while the previous clip's audio keeps playing over it — commonly used to show a reaction shot or B-roll while the previous speaker's dialogue is still finishing. On the timeline, the audio clip extends right past the video cut, forming an L shape.

This is a standard tool for cutting away to B-roll without losing the thread of what is being said: the audience keeps hearing the narration or interview answer while the picture illustrates it, which lets an editor cover a static or awkward on-camera moment without any narrative gap.

L-cuts and J-cuts are frequently used together within the same sequence, and editors will often build several of each into a single interview segment to avoid the flat, mechanical feel of cutting video and audio at exactly the same frame every time.

Jump cut

A cut between two shots of the same subject from a similar angle, creating a visible, deliberate discontinuity rather than a smooth transition.

A jump cut breaks the classical continuity rule that consecutive shots of the same subject should differ enough in angle or framing to read as a scene change rather than a glitch. When that rule is broken, the subject appears to jump slightly in position or time, which used to be considered an editing error but is now a defining stylistic feature of vlog and talking-head content.

It is most commonly used to remove pauses, filler words, or dead air from a single continuous take without cutting away to B-roll: rather than covering the trim, the editor lets the visible jump stand, which has become an accepted pacing convention across YouTube and short-form content.

Overuse still reads as sloppy rather than stylistic — jump cuts work best when they are frequent and consistent throughout a piece, signaling a deliberate style, rather than appearing sporadically where they can look like an accidental skip.

B-roll

Supplementary footage cut over primary dialogue or narration to illustrate a point, cover an edit, or add visual variety.

B-roll sits underneath the audio of the main interview or narration (A-roll) rather than standing on its own, and its job is almost always functional first: covering a jump cut, illustrating a claim, or giving the eye something new after several seconds on a single talking-head shot. It ranges from cutaways filmed specifically for a piece to stock or archival footage sourced afterward.

In retention terms, well-placed B-roll is one of the most effective tools for extending watch time on longer-form content, because it resets visual interest without requiring the underlying story to change pace. Editors typically aim to introduce new visual information every few seconds on platforms with short attention spans, and B-roll is the primary supply of that variety.

Poorly matched B-roll — footage that only loosely relates to what is being said, or that repeats the same clip too often — tends to produce a measurable dip in the retention curve exactly where it appears, since it breaks the sense that every second of the video is deliberately chosen.

How to Choose B-Roll That Actually Supports the Story

A-roll

The primary footage of a video — typically the on-camera speaker or interview subject — that carries the main narrative and audio.

A-roll is the base layer of most edits: the talking head, the interview answer, the presenter delivering a line to camera. It is distinguished from B-roll not by production quality but by function — A-roll is the footage the story is built around, while B-roll supports or covers it.

In the edit, A-roll usually dictates pacing because its audio is what viewers are actually listening to; B-roll, music, and graphics are timed against it rather than the reverse. Editors will typically lock an A-roll cut for pacing and content before layering in any supporting material.

Even in heavily B-roll-driven formats, A-roll usually reappears periodically to re-anchor the viewer to a human presence, since content that cuts away from a speaker for too long without returning can lose the sense of a guiding voice.

Match cut

A transition between two shots that share a similar shape, motion, or composition, making the cut feel seamless or symbolically connected.

A match cut aligns some visual element across the cut point — a circular object in one shot lining up with a circular object in the next, or a movement in one direction continuing into the following shot — so the edit reads as continuous even though the location, subject, or time has changed.

In branded and short-form content, match cuts are frequently used for stylistic transitions between product shots or scenes, since the visual rhyme gives an otherwise arbitrary cut a sense of intentional design, which reads as higher production value without requiring additional footage.

They require more planning than most other cut types because the matching elements usually have to be considered during filming or storyboarding, not discovered after the fact — an editor can occasionally find an accidental match in existing footage, but reliable use requires shots planned with the transition in mind.

Pattern interrupt

A deliberate visual, audio, or narrative disruption inserted to reset viewer attention before it drifts, commonly used every few seconds in short-form editing.

A pattern interrupt breaks whatever rhythm the viewer has settled into — a sudden zoom, a sound effect, a change in camera angle, an unexpected on-screen graphic — specifically to prevent the passive scrolling state where attention quietly disengages without the viewer consciously deciding to leave.

On short-form platforms, editors often plan pattern interrupts on a rough cadence, commonly every 2-4 seconds in fast-paced formats, treating the retention curve as a signal for whether the interval is too wide. A flattening curve mid-video often indicates the pacing has settled into a predictable pattern the viewer has stopped actively watching.

Overusing pattern interrupts creates a different problem: if every second contains a jarring change, the technique loses its contrast and starts to read as chaotic rather than engaging, so it is typically balanced against calmer stretches where the content itself is strong enough to hold attention unaided.

Cold open

A video structure that starts directly on compelling content or a key moment before any introduction, branding, or context is given.

A cold open skips the traditional preamble — no logo animation, no 'hey guys welcome back,' no scene-setting — and instead opens on the most visually or narratively interesting moment available, often pulled from later in the footage and placed at the front as a teaser.

This structure exists specifically to counter early drop-off: since the first few seconds determine whether a video gets further distribution on most platforms, delaying any payoff behind introductory material is one of the most common and avoidable causes of a weak hook rate.

A well-built cold open typically returns to the same moment later in its proper chronological place, meaning the edit effectively shows the audience a preview and then re-earns that moment in context — a technique borrowed from broadcast television that maps directly onto short-form retention mechanics.

Hook Structures That Actually Retain Viewers

Safe zone

The area of a video frame guaranteed to remain visible after platform UI elements — captions, usernames, buttons — are overlaid on top.

Each vertical platform reserves space around the edges of the frame for its own interface: like/comment/share buttons typically sit along the right edge, captions and usernames along the bottom, and profile or sound information near the top. Anything important placed in those zones risks being obscured for a meaningful share of viewers.

Safe zone dimensions differ slightly by platform and even by app version, so editors typically work from the most conservative shared boundary across TikTok, Instagram Reels, and YouTube Shorts rather than optimizing separately for each, unless a video is genuinely platform-exclusive.

Text overlays, subtitles, and key visual details are placed within the safe zone as a matter of course, and any exported video is spot-checked against a template overlay before publishing to confirm nothing critical — a face, a number, a call to action — sits under where the interface will render.

Safe Zones by Platform

Audio

LUFS

Loudness Units Full Scale — the standard unit for measuring perceived loudness across an entire audio track, used by platforms to normalize volume.

LUFS measures integrated loudness rather than peak amplitude, meaning it reflects how loud a track sounds overall to the human ear across its full duration rather than the level of its single loudest moment. Platforms use LUFS targets to normalize audio so that content from different creators plays back at a comparable volume.

Commonly cited targets as published at the time of writing include roughly -14 LUFS integrated for platforms like YouTube, Spotify, and TikTok, and around -16 LUFS for podcast-focused platforms, though these figures are set independently by each platform and do change over time. Delivering audio significantly louder than a platform's target usually results in automatic turndown on playback, while delivering too quiet leaves a track sounding weak relative to surrounding content.

Editors check LUFS at final export using a loudness meter plugin, adjusting the overall mix level rather than individual clips, since the measurement reflects the track as a whole rather than any single section.

True peak

The highest actual amplitude an audio signal reaches after accounting for inter-sample peaks, measured in dBTP to prevent distortion during format conversion.

True peak differs from a standard peak meter reading because it estimates the waveform's real analog peak, including brief spikes that occur between digital samples and that a simple sample-peak meter can miss entirely. A track can measure safely under 0 dB on a normal peak meter and still clip once converted to a lossy format if its true peak exceeds the ceiling.

Delivery specifications commonly cap true peak at -1 dBTP or -2 dBTP, which leaves headroom for the lossy compression platforms apply on upload, preventing the distortion that shows up as harsh clipping once a track that measured clean on export is re-encoded.

In practice this means the final limiter on a mix is set with a true-peak-aware meter rather than a standard one, and the ceiling is treated as a hard constraint rather than a target — the mix is brought right up to it, not routinely left far below it, since unnecessary headroom just makes the track sound quieter than it needs to.

Audio ducking

Automatically or manually lowering the volume of background music when dialogue is present, so speech remains intelligible over the mix.

Ducking reduces the level of a secondary audio track — usually music or ambient sound — whenever a primary track, usually dialogue, is active, then restores it once the dialogue pauses. It can be done manually with keyframes on each dialogue segment or automatically using a sidechain compressor triggered by the dialogue track.

The amount of reduction matters more than the technique used to achieve it: ducking that is too subtle leaves music masking speech, particularly in the same frequency range as the human voice, while ducking that is too aggressive creates an audibly pumping effect where the music noticeably swells and drops with every sentence.

Well-executed ducking is generally inaudible as a technique — viewers should notice that they can hear the speaker clearly, not notice the music moving. Editors typically automate it across a whole video with sidechain compression for consistency, then manually adjust individual moments where the automatic response over- or under-reacts.

Audio Ducking Explained

Room tone

A short recording of the ambient sound in a location with no one speaking, used to fill gaps and smooth audio edits so silence doesn't sound artificially dead.

Every recording space has a baseline ambient noise floor — HVAC hum, distant traffic, room reflections — that is present even when nobody is talking. Room tone is a deliberate recording of that ambience, typically captured for 30-60 seconds on set immediately before or after filming, with everyone silent and still.

In the edit, room tone is used to fill gaps left by removed dialogue, coughs, or noise, because true digital silence sounds noticeably unnatural when cut against a track that otherwise carries constant ambient noise — the drop to absolute silence is audible even if the viewer cannot articulate why a cut feels wrong.

It is one of the cheapest insurance policies in production: a missed thirty seconds of room tone on set can mean an editor has no clean way to cover a bad edit later, whereas having it available makes even aggressive dialogue trimming sound seamless.

De-esser

An audio processor that reduces harsh sibilant frequencies — the 's,' 'sh,' and 't' sounds — in recorded speech without affecting the rest of the voice.

Sibilance sits in a narrow, high frequency band, typically somewhere in the 4-9 kHz range depending on the voice, and becomes exaggerated by close-proximity microphones and heavy compression, both of which are common in talking-head and podcast-style production. A de-esser is essentially a compressor tuned to react only within that band, ducking it briefly whenever it spikes rather than compressing the whole signal.

It is applied after general EQ and before final limiting in most audio chains, since applying it too early can miss sibilance introduced by later processing, and applying it too aggressively produces a lisping or muffled artifact on 's' sounds that is worse than the original harshness.

Editors typically set the de-esser by ear on the harshest line in a recording, then check the setting across the rest of the piece rather than automating it per-clip, since inconsistent de-essing between cuts is itself audible as a texture change.

Noise floor

The baseline level of unwanted background noise present in an audio recording even when no intentional sound is happening.

Every recording environment and every piece of gear introduces some amount of noise — electrical hiss from a preamp, HVAC rumble, distant street noise — and the noise floor is the level of that unwanted signal relative to the desired audio. A low noise floor means dialogue sits cleanly above the background; a high noise floor means noise reduction has to work harder and more audibly to clean the track.

Noise floor is best controlled during recording rather than fixed afterward: closer mic placement, quieter locations, and better gear all raise the ratio between wanted signal and noise floor before an editor ever touches the file, whereas noise reduction applied in post always trades some noise removal for some loss of vocal clarity.

When noise reduction is necessary, editors apply it conservatively and check the result against headphones at a realistic playback volume, since overly aggressive noise reduction introduces its own artifacts — a warbling or underwater quality — that can be more distracting than the original noise floor.

Technical delivery

Bitrate

The amount of data used to encode one second of video or audio, typically measured in megabits per second, which directly affects file size and visual quality.

Bitrate determines how much information is available to represent each second of footage — higher bitrate generally preserves more detail, especially in motion and complex textures, while lower bitrate introduces compression artifacts like blocking or banding, particularly visible in dark scenes or fast movement.

Bitrate needs vary by resolution, frame rate, and codec efficiency: a 1080p video and a 4K video require very different bitrates to look equally clean, and a more efficient codec like H.265 can achieve comparable quality to H.264 at roughly half the bitrate. Platforms also publish their own recommended upload bitrates, and exceeding them provides no visible benefit since the platform re-encodes on ingest anyway.

Editors set export bitrate based on the platform's recommended range for the resolution and frame rate in question rather than maximizing it arbitrarily, since an unnecessarily high bitrate mainly produces a larger file with a longer upload time and no perceptible quality gain after the platform's own compression is applied.

Export Settings by Platform

Codec (H.264/H.265)

The compression standard used to encode video into a manageable file size; H.264 is the long-standing default, H.265 offers similar quality at lower bitrates.

A codec defines the algorithm used to compress raw video data into a deliverable file. H.264 (AVC) has been the dominant standard for over a decade because of its broad compatibility across platforms, devices, and editing software, and remains the safest default for final delivery when compatibility is the priority.

H.265 (HEVC) achieves comparable visual quality at roughly half the bitrate of H.264, making it attractive for 4K and high-frame-rate footage where file size and bandwidth matter, but it has historically had more inconsistent support across browsers, older devices, and some platform upload pipelines, and can require more processing power to encode and decode.

The choice between them is generally use-case driven rather than fixed: H.264 for maximum compatibility on final platform uploads, H.265 or a mezzanine/intermediate codec for internal archiving and heavier source footage where storage efficiency and quality retention matter more than universal playback support.

Export Settings by Platform

Colour grading

The process of adjusting a video's color, contrast, and tone after editing to establish a consistent visual mood and correct camera inconsistencies.

Colour grading happens after colour correction, which fixes technical issues like incorrect white balance or exposure across shots so that footage from different cameras or lighting conditions matches. Grading then goes further, deliberately shaping colour and contrast to create a specific look or mood, from a warm, saturated feel to a desaturated, cinematic tone.

It is typically the last major visual step before export, applied after the cut is locked, because grading decisions are made in context of the finished sequence rather than shot by shot in isolation — a look that reads well on one clip can clash when placed next to its neighbors in the timeline.

For brand and social content, grading is also a consistency tool: applying the same grading approach or a shared LUT as a starting point across every video in a series gives a channel a recognizable visual identity that viewers register even without consciously noticing the colour work itself.

LUTs and Colour Grading for Social Video

LUT

Look-Up Table — a preset file that maps input colour values to output colour values, used to quickly apply a consistent colour style to footage.

A LUT is essentially a translation table: for every colour value that comes into it, it defines what colour should come out, applied uniformly across an image. Technical LUTs convert footage shot in a flat, low-contrast log profile into a standard viewable colour space, while creative LUTs apply a stylized look on top of that correction.

LUTs are a starting point, not a finished grade — footage shot in different lighting conditions or on different cameras will respond differently to the same LUT, so editors typically apply a LUT and then make secondary adjustments per shot to correct for exposure and white balance differences the LUT alone won't fix.

Because a LUT is just a file, it can be shared and reused across an entire series or channel, which is why many creators standardize on one or two LUTs as part of their visual brand — the tool that makes visual consistency achievable quickly across a high volume of edits without regrading every project from scratch.

LUTs and Colour Grading for Social Video

White balance

The calibration of a video's colour temperature so that white and neutral objects appear truly neutral rather than tinted orange or blue.

Different light sources emit light at different colour temperatures, measured in Kelvin — tungsten bulbs skew warm/orange, overcast daylight skews cool/blue — and a camera's white balance setting compensates for that so neutral colours render accurately rather than picking up an unwanted cast from the light source.

Incorrect white balance is usually fixable in post if the footage was shot in a flat or raw profile, but the correction becomes destructive and can introduce noise or banding if pushed too far, which is why getting white balance close to correct on set is preferable to relying entirely on post-production correction.

In multi-camera or multi-location edits, mismatched white balance between shots is one of the most common reasons footage feels inconsistent even when composition and exposure are otherwise matched, so it is typically one of the first things corrected before any creative grading begins.

Exposure / waveform

A scope that displays the brightness values of a video signal graphically, used to judge and correct exposure objectively rather than by eye alone.

A waveform monitor plots the luminance of every point across the frame's horizontal axis against a vertical brightness scale, giving an objective read on exposure that is independent of a monitor's calibration or a room's ambient lighting — both of which can make footage look brighter or darker on screen than it actually is.

Editors use it to check for clipped highlights (values pushed to the very top of the scale, where detail is permanently lost) and crushed shadows (values pushed to the very bottom), both of which can be invisible on an uncalibrated display but are immediately visible on the waveform.

This becomes especially important when correcting footage from multiple cameras shot in different lighting, since matching waveform ranges across shots is a more reliable way to achieve consistent exposure than adjusting by eye, which is easily fooled by the surrounding footage or a monitor's brightness settings.

Burned-in captions vs. closed captions

Burned-in captions are permanently rendered into the video image, while closed captions are a separate, toggleable text layer decoded at playback.

Burned-in (or open) captions are part of the video file itself — encoded directly into the pixels — meaning they display identically on every platform and device regardless of accessibility settings, but they cannot be turned off, translated automatically, or styled by the viewer. Closed captions exist as a separate data track or file that the playback device renders on top of the video and that a viewer can toggle or, on some platforms, translate.

Short-form platforms overwhelmingly favor burned-in captions because most viewing happens with sound off in public or passive-scrolling contexts, and a caption baked into the video guarantees visibility without depending on the viewer to enable anything. Long-form and platforms with strong accessibility requirements typically favor or additionally require closed captions delivered as a separate file.

Many deliverables now include both: styled, animated burned-in captions for the visual/retention benefit on short-form, alongside a standard closed-caption file for accessibility compliance and searchability on platforms like YouTube, since the two serve different functions and aren't mutually exclusive.

Subtitle Accessibility Standards · Caption Styling for Retention

SRT file

A plain-text subtitle file format that pairs blocks of text with timecodes, used to add closed captions to a video without altering the video file itself.

An SRT (SubRip Subtitle) file is a simple, numbered list of text entries, each with a start and end timecode, that a video player reads alongside the video to display captions at the correct moments. Because it's plain text, it's lightweight, easy to edit manually, and universally supported across nearly every platform and player.

SRT files are the standard deliverable for closed captioning because they keep the caption layer separate from the video: the same SRT can be swapped for a translated version, styled differently by different platforms, or updated for a correction without needing to re-export the underlying video.

Editors typically generate an initial SRT via automatic transcription and then manually correct timing and text accuracy, since auto-generated captions reliably misinterpret names, jargon, and homophones, and a caption file with visible errors undermines credibility even when the video content itself is accurate.

Subtitle Accessibility Standards

Caption safe area

The region of a vertical video frame reserved for captions so they remain readable and don't overlap platform UI elements like buttons or usernames.

On vertical formats, the bottom third of the frame is heavily contested space: platform UI typically places the caption/username bar and, on some platforms, additional interface elements in this zone, meaning captions placed too low risk being covered or cut off depending on the specific app and device.

Editors typically position burned-in captions somewhere in the middle-to-lower-middle of the frame rather than at the very bottom, leaving enough margin above the platform's own UI band, and check the placement against each target platform's interface before finalizing a caption template used repeatedly across a series.

This becomes more complex when a video is repurposed across multiple platforms simultaneously, since each has a slightly different UI footprint — a caption position validated as safe on one platform may sit uncomfortably close to an element on another, which is why templates are usually built to the most conservative shared safe area rather than optimized per platform.

Caption Styling for Retention · Safe Zones by Platform

Platform & distribution

Aspect ratio

The proportional relationship between a video frame's width and height, expressed as a ratio such as 16:9 or 9:16.

Aspect ratio determines the fundamental shape of the frame independent of resolution — a 16:9 frame is wider than it is tall, common for traditional landscape video and YouTube's default player, while 9:16 is the inverse, tall and narrow, matching how a phone is naturally held for scrolling short-form platforms.

Content shot natively in one aspect ratio and delivered on a platform built for another requires reframing, cropping, or letterboxing to fit, and each approach has tradeoffs: reframing can crop out important visual elements, and letterboxing preserves the full frame but leaves visible bars that shrink the effective image size.

Because most content is now distributed across multiple platforms with different native aspect ratios, many productions now shoot with multiple deliverable ratios in mind from the start — framing subjects and action centrally enough that both a 16:9 and a 9:16 crop remain usable without reshooting.

16:9 to 9:16 Reframing

Reframing

Adjusting the crop and position of footage to fit a different aspect ratio than it was originally shot in, most commonly converting 16:9 to 9:16.

Reframing takes source footage shot for one aspect ratio and repositions or crops it to fill a different one, typically converting wide landscape footage into the tall vertical format required by short-form platforms. This is distinct from simply adding black bars — reframing actively decides what part of the original frame remains visible.

Static reframing applies one crop position for an entire shot, which works when the subject stays roughly centered, while dynamic reframing moves the crop within the shot to follow a subject as they move, sometimes with motion tracking automating the follow. Multi-subject shots — two people in conversation, for example — are the hardest case, often requiring a split-screen or a cut between individual crops rather than a single reframed shot.

The quality of a reframe is judged by whether a viewer who never saw the original footage would assume it was shot vertically to begin with — visible zooming artifacts, awkward headroom, or cropped-out important information (a second speaker, on-screen text, a product) are the most common signs of a rushed reframe.

16:9 to 9:16 Reframing

9:16

A vertical aspect ratio, tall rather than wide, that matches how a smartphone is held in portrait orientation and is the native format for most short-form platforms.

9:16 describes a frame nine units wide for every sixteen units tall — the rotated inverse of the familiar 16:9 widescreen ratio. It became the dominant format for short-form content because it fills the entire screen of a phone held upright without any cropping or black bars, matching the default way most people hold their device while scrolling.

TikTok, Instagram Reels, and YouTube Shorts are all built around 9:16 as their primary native format, and content submitted in other ratios is typically either rejected, cropped automatically by the platform in an uncontrolled way, or displayed with black bars — none of which produce the intended framing, so 9:16 delivery is treated as a hard requirement rather than a preference for these platforms.

Shooting natively in 9:16, when a piece of content is destined only for vertical platforms, generally produces better results than shooting 16:9 and reframing afterward, since it removes any risk of losing important visual information in the crop — though many productions still shoot wide for flexibility when a piece needs to serve both horizontal and vertical deliverables.

16:9 to 9:16 Reframing

Letterboxing

Adding black bars above and below (or beside) a video's image to preserve its original aspect ratio when displayed within a differently shaped frame.

Letterboxing keeps the entirety of the original frame intact by placing it inside a differently proportioned canvas and filling the unused space with black — most commonly seen as horizontal black bars when 16:9 content is displayed within a taller 9:16 frame rather than being cropped to fill it.

The tradeoff versus reframing is straightforward: letterboxing loses no visual information from the original shot but reduces the effective size of the image on screen, which matters more on small mobile screens where a letterboxed video already occupies a shrunken portion of the display before any UI overlays are added on top.

It is generally used as a fallback when reframing isn't practical or would crop out essential content — a wide group shot, for instance, that cannot be cropped to vertical without losing people out of frame — and some editors soften the visual impact by filling the bars with a blurred, scaled-up version of the footage itself rather than solid black.

16:9 to 9:16 Reframing

Thumbnail CTR

Click-through rate on a video's thumbnail and title — the percentage of people who see it in a feed and choose to click, a major factor in early distribution.

Thumbnail CTR is calculated as clicks divided by impressions, measuring how effectively a thumbnail and title combination convinces someone browsing a feed to actually open the video. On platforms like YouTube, this is one of the first signals used to decide whether to show a video to a wider audience, independent of how good the video is once someone starts watching.

Strong thumbnails typically rely on a clear focal point, legible even at small sizes, combined with a title that creates a specific curiosity gap rather than a vague or generic promise — the two are designed together rather than separately, since a thumbnail and title that make conflicting or redundant claims tend to underperform ones that work as a single unit.

CTR and retention are related but separate problems: a thumbnail promising something the video doesn't deliver can produce a high CTR but a poor retention curve as viewers immediately recognize the mismatch, which most platforms eventually penalize even though CTR itself looked strong.

Thumbnail Design for Click-Through

Hook stack

A layered combination of verbal, visual, and text hooks used simultaneously in the opening seconds of a video to maximize the chance one of them lands.

Rather than relying on a single hook technique, a hook stack combines several at once in the first few seconds — a spoken line that creates curiosity, an on-screen text overlay reinforcing or adding to that line, and a visually disruptive shot or movement, all layered together rather than delivered sequentially.

The logic behind stacking is that different viewers process information differently and some may be watching without sound, reading quickly, or skimming visually before committing attention, so a hook relying on only one channel — audio alone, for instance — misses viewers who aren't engaging through that channel in the first place.

Building an effective hook stack requires the elements to reinforce rather than compete with each other: text that merely repeats the spoken line verbatim adds little, while text that adds a complementary detail or specific number strengthens the overall hook without requiring more time or a longer opening.

Hook Structures That Actually Retain Viewers

Production workflow

Proxy workflow

Editing with lower-resolution stand-in copies of original footage to keep playback smooth, then relinking to full-resolution files for final export.

High-resolution camera formats, particularly 4K and above shot in compressed codecs, are often too demanding for real-time playback and scrubbing on standard editing hardware. A proxy workflow transcodes the footage into a lighter-weight format at the start of a project, and the editor cuts using those lightweight files instead of the originals.

Once picture lock is reached, the editing software relinks the timeline back to the original full-resolution media for color grading and final export, so no quality is lost in the finished deliverable — the proxies exist only to make the editing process responsive, not to appear in the final output.

This workflow becomes essential rather than optional on multi-camera shoots, long-form projects, or any timeline heavy with 4K/6K footage and effects, where editing directly off originals would introduce enough lag to meaningfully slow down the cutting process and the ability to judge pacing in real time.

Shot list

A pre-production document breaking down every planned shot in a production, typically including framing, movement, and purpose, used to guide filming.

A shot list translates a script or outline into a concrete filming plan, specifying for each shot the framing (wide, medium, close-up), any camera movement, the subject, and often the purpose the shot serves in the edit — establishing context, covering a cut, delivering a key line.

Its main value shows up after filming, not during: a thorough shot list ensures that the editor has the coverage needed to build a scene without gaps, particularly enough B-roll and cutaway options to cover every planned jump cut or trim, which is far cheaper to plan for on set than to discover missing in the edit.

Shot lists are typically built collaboratively between whoever is directing the shoot and whoever will edit the footage, since the edit's pacing and structural needs — how many cutaways a scene will realistically require, what needs to be shot from multiple angles — are easier to anticipate before filming than to fix afterward with limited footage.

EDL / timeline handoff

An Edit Decision List or exported project file that records every cut, clip, and timing decision, used to move a project between editors or software.

An EDL is a structured list of every edit point in a sequence — which source clip, what in and out timecodes, and where it sits on the timeline — originally developed for handing off cuts between different editing systems without needing to rebuild the sequence from scratch. Modern handoffs more often use a full project file or an interchange format like XML or AAF, which carries additional information EDLs can't, such as effects, colour adjustments, and multiple tracks.

Timeline handoff becomes necessary whenever a project moves between people or software — a rough cut assembled by one editor handed to a colourist, or a project moving from one editing application to another for finishing — and a clean handoff depends on consistent media organization and naming, since a timeline referencing files that can't be located or matched breaks the entire handoff.

Teams working across multiple editors on the same channel typically standardize file-naming and folder-structure conventions specifically to keep handoffs reliable, since the biggest source of handoff failure is rarely the file format itself but rather relinking media that was renamed, moved, or organized inconsistently between the two systems.

File Delivery and Handoff

Turnaround time

The elapsed time between receiving raw footage or a brief and delivering a finished edit, a key operational metric for editing teams and agencies.

Turnaround time is typically measured in business days from footage/brief receipt to first-draft delivery, and is distinct from total project time when revision rounds are included, since those are usually tracked and quoted separately. Short-form content commonly turns around in 24-72 hours given its shorter runtime and simpler structure, while long-form or heavily produced content can take a week or more for a first cut.

Turnaround time is directly shaped by workflow decisions made upstream — a proxy workflow, organized shot lists, and clear briefs all shorten it, while missing footage, unclear direction, or disorganized source files extend it regardless of how efficient the actual cutting process is.

Clients and editors typically agree on turnaround expectations before a project starts, since it affects everything from how many revision rounds are realistic to whether rush delivery requires a different rate — treating it as a fixed variable in scheduling rather than an afterthought once editing begins.

Revision round

A defined cycle of feedback and changes applied to a draft edit, typically limited to a set number per project as part of a production agreement.

A revision round covers all the feedback collected and applied in a single pass — rather than treating each individual note as its own cycle, feedback is batched and delivered to the editor together, then addressed in one updated draft. This keeps the process efficient compared to responding to notes one at a time as they arrive.

Most production agreements cap the number of included revision rounds (commonly one to three) to keep scope and cost predictable, with additional rounds available at an extra cost — a structure that also encourages clients to consolidate feedback thoroughly rather than trickling in changes indefinitely.

Clear revision rounds work best when feedback is specific and timecoded rather than general, since vague notes like 'make it punchier' require the editor to guess at intent, whereas a note tied to an exact timestamp with a concrete change produces a faster and more accurate second draft.

Want this applied to your footage?

Send raw footage and a brief. We deliver a polished sample edit so you can judge hook selection, captions and audio against everything defined on this page.