Editing utility guide

Audio Ducking
the fix for videos that feel amateur

Bad audio reads as low production value faster than bad picture. The most common failure is not noise or hiss but a music bed competing with the speaker, which makes viewers work to listen and gives them a reason to scroll. Ducking is a five-minute fix that most edits skip.

Last updated · Reviewed by the Media Strategy Lab edit team

Benchmark data from our 1.2B+ view dataset

Aggregated from short-form campaigns produced by Media Strategy Lab in 2025-2026.

38%

median hook retention

22%

3-sec drop-off

31s

avg. watch time

problem-statement or pattern-interrupt

best hook type

2.1 cuts per 10s

cut density

Format and pacing profile

dominant format

Talking head + supporting B-roll

shot length

2-4 seconds

B-roll ratio

40:60 B-roll to face

pacing note

Lead with the hook, cut on breaths, use text reinforcement at 3-5s intervals.

Clean dialogue with a music bed ducking -20 LUFS under voice.

Technical specifications

Dialogue target-14 LUFS (social), -16 LUFS (podcast)
Music bed under speech-6 LU below dialogue
Music in gaps-3 LU below dialogue
Attack80–150 ms
Release300–500 ms
True peak ceiling-1 dBTP
High-pass on dialogue80–100 Hz
De-esser5–8 kHz, gentle

Buyer context and objections

who buys

Editors, social managers and marketing teams doing the work in-house

typical budget

Free guide — retainers from $2,495/mo if you want it done for you

common objection

We can probably work this out ourselves

failed prior attempt

Copying a tutorial without adapting it to the destination platform

Our 5-step process

  1. 01

    Brief and audit — we review your goals, past performance and raw material before touching a timeline.

  2. 02

    Hook extraction — every asset is scanned for the highest-retention 1-3 second opener.

  3. 03

    Native edit — pacing, captions, safe zones and sound are tuned to the destination platform.

  4. 04

    Revision rounds — two included rounds with timestamped comments, no ticket queue.

  5. 05

    Delivery pack — masters, verticals, captions, thumbnails and a posting brief in one drop.

Case example

A coaching client's videos had music at roughly equal level with speech. Applying a 6 LU duck with 120 ms attack and 400 ms release, plus a light high-pass, took about eight minutes per video. Comments about audio stopped and average watch time rose 14 percent across the next month.

Pricing anchor

Our monthly retainers start at $2,495/mo for 15 shorts and scale to $3,995/mo for 30 shorts plus long-form support. Every retainer includes research, scripting, editing, uploading, captions, weekday support and monthly reporting.

Set the gap, not the level

Think in relative terms. Music should sit roughly 6 LU below dialogue while someone is speaking and can rise to about 3 LU below in gaps. Absolute fader positions vary per track; the gap is what the ear judges.

Sidechain compression driven by the dialogue track automates this, but manual keyframes are perfectly fine and often cleaner on short-form where there are only a handful of transitions.

Attack and release do the disguising

An attack around 80 to 150 milliseconds ducks fast enough to clear the first syllable without an audible pump. A release of 300 to 500 milliseconds lets music return smoothly rather than lurching back between sentences.

Fast release settings are the reason ducking sometimes sounds worse than no ducking at all — the music breathes in and out and draws attention to itself.

Fix the dialogue before you mix

High-pass at 80 to 100 Hz to remove rumble, apply gentle broadband noise reduction rather than aggressive cleanup, de-ess lightly around 5 to 8 kHz, then compress moderately for consistency before setting levels.

No amount of music balancing rescues a poorly captured voice track. Where source audio is genuinely unusable, replacing it with a clean re-record over B-roll is faster and better than repairing it.

Loudness (LUFS) reference and ducking calculator

Interactive, no email required. Numbers come from our own production data.

TikTok-14 LUFS / -1 dBTPNormalises up; thin mixes sound weak.
Instagram / Reels-14 LUFS / -1 dBTPDialogue-forward, music bed -20 LUFS.
YouTube-14 LUFS / -1 dBTPLouder masters get turned down, not up.
LinkedIn-16 LUFS / -1.5 dBTPMost plays are muted — captions carry it.
Podcast (audio)-16 LUFS / -1 dBTPMono -19 LUFS for spoken-word feeds.
Broadcast / OTT-23 LUFS / -2 dBTPEBU R128 delivery standard.

Dialogue: -14 LUFS

-20 LUFS

Duck the music bed to this level under speech (a 6 LU gap), with a 120 ms attack and 400 ms release.

All free tools →

Frequently asked questions

How far should music be ducked under dialogue?

About 6 LU below the dialogue while someone is speaking, rising to around 3 LU below in gaps. Relative gap matters more than absolute level, because different tracks have very different perceived loudness at the same fader position.

What attack and release should I use?

Eighty to one hundred and fifty milliseconds attack, three hundred to five hundred milliseconds release. Faster releases cause audible pumping between sentences, which is more distracting than not ducking at all.

What LUFS should social video be?

Minus fourteen LUFS integrated with true peak at -1 dBTP for TikTok, Instagram and YouTube. LinkedIn sits nearer -16, and spoken-word podcast feeds are usually -16 stereo or -19 mono.

Should I use sidechain compression or keyframes?

Sidechain for long-form with continuous music, keyframes for short-form where there are only a few transitions. Keyframes give more precise control and avoid the artefacts a poorly tuned sidechain introduces on rapid speech.

How do I fix bad remote guest audio?

High-pass, gentle broadband noise reduction, careful de-essing and moderate compression, in that order. Aggressive noise reduction creates a watery artefact that sounds worse than the original noise. If the track is beyond repair, cut to B-roll and use a re-record.

Does music even help retention?

On energetic short-form, yes — it sets pace and covers edit points. On technical or authority-led content it often hurts, and a clean dialogue track with sound design accents performs better. Test it on your own audience rather than assuming.

Get a sample edit for Audio Ducking

Rather not do this yourself? Send us your footage and we'll apply all of it as standard.