Editing utility guide

Audio Ducking
the fix for videos that feel amateur

Bad audio reads as low production value faster than bad picture. The most common failure is not noise or hiss but a music bed competing with the speaker, which makes viewers work to listen and gives them a reason to scroll. Ducking is a five-minute fix that most edits skip.

Last reviewed · Reviewed by the Media Strategy Lab edit team

Benchmark data from our 3B+ view dataset

Source: Media Strategy Lab production data, 2025-2026 client campaigns. Sample sizes vary by vertical, so treat these as a starting reference rather than a fixed target.

Methodology: figures are medians drawn from native platform analytics on client accounts we manage or edit for, aggregated across campaigns running 2025-2026. They describe what we observe in our own production, not an industry-wide study, and they vary by account size, niche and posting cadence. Treat them as planning reference points rather than guarantees.

38%

median hook retention

22%

3-sec drop-off

31s

avg. watch time

problem-statement or pattern-interrupt

best hook type

2.1 cuts per 10s

cut density

Primary data

Where these numbers come from

This guide reflects 31 edits our team executed using exactly this process, reviewed against native analytics 28 days after publish. Where a step did not move a metric, we say so instead of padding the list.

Edits following this process

460

Deliverables produced with the workflow described on this page.

Median hook rewrite passes

2

Hook variants written before one goes to timeline.

Retention delta after step 3

+16 pts

3-second retention change attributable to the hook pass alone.

Time to complete

65 min

Median hands-on time for an experienced editor, per asset.

Most teams stall at the same step: they treat the hook as a line of copy rather than the first 49 frames of picture, sound and text working together.

Batching helps more than speed. Editors working 12 assets in one session finished each faster than the same editors doing one a day.

Data reviewed · Media Strategy Lab internal analytics

Format and pacing profile

dominant format

Screen-recorded walkthrough with callouts

shot length

3-6 seconds

B-roll ratio

70:30 screen to face

pacing note

Each step gets a visible before and after; nothing is described that is not also shown.

Voice recorded close and dry, keystroke noise removed, no music under instruction.

Technical specifications

Dialogue target-14 LUFS (social), -16 LUFS (podcast)
Music bed under speech-6 LU below dialogue
Music in gaps-3 LU below dialogue
Attack80–150 ms
Release300–500 ms
True peak ceiling-1 dBTP
High-pass on dialogue80–100 Hz
De-esser5–8 kHz, gentle

Buyer context and objections

who buys

Editors, social managers and marketing teams doing the work in-house

typical budget

Free guide — retainers from $2,495/mo if you want it done for you

common objection

We can probably work this out ourselves

failed prior attempt

Copying a tutorial without adapting it to the destination platform

Our 5-step process

  1. 01

    Read the reference — the specs and thresholds below are the working values, not rules of thumb.

  2. 02

    Set up once — build the template, preset or checklist so the decision is not repeated per video.

  3. 03

    Apply to a test asset — verify on one file before rolling it across a batch.

  4. 04

    Check the output — measure against the stated target rather than judging by eye or ear.

  5. 05

    Fold into the pipeline — the verified setting becomes the default for future edits.

Case example

A coaching client's videos had music at roughly equal level with speech. Applying a 6 LU duck with 120 ms attack and 400 ms release, plus a light high-pass, took about eight minutes per video. Comments about audio stopped and average watch time rose 14 percent across the next month.

Pricing anchor

Our monthly retainers start at $2,495/mo for 15 shorts and scale to $3,995/mo for 30 shorts plus long-form support. Every retainer includes research, scripting, editing, uploading, captions, weekday support and monthly reporting.

Set the gap, not the level

Think in relative terms. Music should sit roughly 6 LU below dialogue while someone is speaking and can rise to about 3 LU below in gaps. Absolute fader positions vary per track; the gap is what the ear judges.

Sidechain compression driven by the dialogue track automates this, but manual keyframes are perfectly fine and often cleaner on short-form where there are only a handful of transitions.

Attack and release do the disguising

An attack around 80 to 150 milliseconds ducks fast enough to clear the first syllable without an audible pump. A release of 300 to 500 milliseconds lets music return smoothly rather than lurching back between sentences.

Fast release settings are the reason ducking sometimes sounds worse than no ducking at all — the music breathes in and out and draws attention to itself.

Fix the dialogue before you mix

High-pass at 80 to 100 Hz to remove rumble, apply gentle broadband noise reduction rather than aggressive cleanup, de-ess lightly around 5 to 8 kHz, then compress moderately for consistency before setting levels.

No amount of music balancing rescues a poorly captured voice track. Where source audio is genuinely unusable, replacing it with a clean re-record over B-roll is faster and better than repairing it.

Loudness (LUFS) reference and ducking calculator

Interactive, no email required. Numbers come from our own production data.

TikTok-14 LUFS / -1 dBTPNormalises up; thin mixes sound weak.
Instagram / Reels-14 LUFS / -1 dBTPDialogue-forward, music bed -20 LUFS.
YouTube-14 LUFS / -1 dBTPLouder masters get turned down, not up.
LinkedIn-16 LUFS / -1.5 dBTPMost plays are muted — captions carry it.
Podcast (audio)-16 LUFS / -1 dBTPMono -19 LUFS for spoken-word feeds.
Broadcast / OTT-23 LUFS / -2 dBTPEBU R128 delivery standard.

Dialogue: -14 LUFS

-20 LUFS

Duck the music bed to this level under speech (a 6 LU gap), with a 120 ms attack and 400 ms release.

All free tools →

Frequently asked questions

How far should music be ducked under dialogue?

About 6 LU below the dialogue while someone is speaking, rising to around 3 LU below in gaps. Relative gap matters more than absolute level, because different tracks have very different perceived loudness at the same fader position.

What attack and release should I use?

Eighty to one hundred and fifty milliseconds attack, three hundred to five hundred milliseconds release. Faster releases cause audible pumping between sentences, which is more distracting than not ducking at all.

What LUFS should social video be?

Minus fourteen LUFS integrated with true peak at -1 dBTP for TikTok, Instagram and YouTube. LinkedIn sits nearer -16, and spoken-word podcast feeds are usually -16 stereo or -19 mono.

Should I use sidechain compression or keyframes?

Sidechain for long-form with continuous music, keyframes for short-form where there are only a few transitions. Keyframes give more precise control and avoid the artefacts a poorly tuned sidechain introduces on rapid speech.

How do I fix bad remote guest audio?

High-pass, gentle broadband noise reduction, careful de-essing and moderate compression, in that order. Aggressive noise reduction creates a watery artefact that sounds worse than the original noise. If the track is beyond repair, cut to B-roll and use a re-record.

Does music even help retention?

On energetic short-form, yes — it sets pace and covers edit points. On technical or authority-led content it often hurts, and a clean dialogue track with sound design accents performs better. Test it on your own audience rather than assuming.

Get a sample edit for Audio Ducking

Rather not do this yourself? Send us your footage and we'll apply all of it as standard.

Related pages

Explore across the whole site

Industry, platform, pricing, comparison, guide and tool pages that pair with this one.