Editing utility guide
Audio Ducking
the fix for videos that feel amateur
Bad audio reads as low production value faster than bad picture. The most common failure is not noise or hiss but a music bed competing with the speaker, which makes viewers work to listen and gives them a reason to scroll. Ducking is a five-minute fix that most edits skip.
Last updated · Reviewed by the Media Strategy Lab edit team
Benchmark data from our 1.2B+ view dataset
Aggregated from short-form campaigns produced by Media Strategy Lab in 2025-2026.
38%
median hook retention
22%
3-sec drop-off
31s
avg. watch time
problem-statement or pattern-interrupt
best hook type
2.1 cuts per 10s
cut density
Format and pacing profile
dominant format
Talking head + supporting B-roll
shot length
2-4 seconds
B-roll ratio
40:60 B-roll to face
pacing note
Lead with the hook, cut on breaths, use text reinforcement at 3-5s intervals.
Clean dialogue with a music bed ducking -20 LUFS under voice.
Technical specifications
| Dialogue target | -14 LUFS (social), -16 LUFS (podcast) |
|---|---|
| Music bed under speech | -6 LU below dialogue |
| Music in gaps | -3 LU below dialogue |
| Attack | 80–150 ms |
| Release | 300–500 ms |
| True peak ceiling | -1 dBTP |
| High-pass on dialogue | 80–100 Hz |
| De-esser | 5–8 kHz, gentle |
Buyer context and objections
who buys
Editors, social managers and marketing teams doing the work in-house
typical budget
Free guide — retainers from $2,495/mo if you want it done for you
common objection
We can probably work this out ourselves
failed prior attempt
Copying a tutorial without adapting it to the destination platform
Our 5-step process
01
Brief and audit — we review your goals, past performance and raw material before touching a timeline.
02
Hook extraction — every asset is scanned for the highest-retention 1-3 second opener.
03
Native edit — pacing, captions, safe zones and sound are tuned to the destination platform.
04
Revision rounds — two included rounds with timestamped comments, no ticket queue.
05
Delivery pack — masters, verticals, captions, thumbnails and a posting brief in one drop.
Case example
A coaching client's videos had music at roughly equal level with speech. Applying a 6 LU duck with 120 ms attack and 400 ms release, plus a light high-pass, took about eight minutes per video. Comments about audio stopped and average watch time rose 14 percent across the next month.
Pricing anchor
Our monthly retainers start at $2,495/mo for 15 shorts and scale to $3,995/mo for 30 shorts plus long-form support. Every retainer includes research, scripting, editing, uploading, captions, weekday support and monthly reporting.
Set the gap, not the level
Think in relative terms. Music should sit roughly 6 LU below dialogue while someone is speaking and can rise to about 3 LU below in gaps. Absolute fader positions vary per track; the gap is what the ear judges.
Sidechain compression driven by the dialogue track automates this, but manual keyframes are perfectly fine and often cleaner on short-form where there are only a handful of transitions.
Attack and release do the disguising
An attack around 80 to 150 milliseconds ducks fast enough to clear the first syllable without an audible pump. A release of 300 to 500 milliseconds lets music return smoothly rather than lurching back between sentences.
Fast release settings are the reason ducking sometimes sounds worse than no ducking at all — the music breathes in and out and draws attention to itself.
Fix the dialogue before you mix
High-pass at 80 to 100 Hz to remove rumble, apply gentle broadband noise reduction rather than aggressive cleanup, de-ess lightly around 5 to 8 kHz, then compress moderately for consistency before setting levels.
No amount of music balancing rescues a poorly captured voice track. Where source audio is genuinely unusable, replacing it with a clean re-record over B-roll is faster and better than repairing it.
Loudness (LUFS) reference and ducking calculator
Interactive, no email required. Numbers come from our own production data.
| TikTok | -14 LUFS / -1 dBTPNormalises up; thin mixes sound weak. |
|---|---|
| Instagram / Reels | -14 LUFS / -1 dBTPDialogue-forward, music bed -20 LUFS. |
| YouTube | -14 LUFS / -1 dBTPLouder masters get turned down, not up. |
| -16 LUFS / -1.5 dBTPMost plays are muted — captions carry it. | |
| Podcast (audio) | -16 LUFS / -1 dBTPMono -19 LUFS for spoken-word feeds. |
| Broadcast / OTT | -23 LUFS / -2 dBTPEBU R128 delivery standard. |
Dialogue: -14 LUFS
-20 LUFS
Duck the music bed to this level under speech (a 6 LU gap), with a 120 ms attack and 400 ms release.
Frequently asked questions
How far should music be ducked under dialogue?
About 6 LU below the dialogue while someone is speaking, rising to around 3 LU below in gaps. Relative gap matters more than absolute level, because different tracks have very different perceived loudness at the same fader position.
What attack and release should I use?
Eighty to one hundred and fifty milliseconds attack, three hundred to five hundred milliseconds release. Faster releases cause audible pumping between sentences, which is more distracting than not ducking at all.
What LUFS should social video be?
Minus fourteen LUFS integrated with true peak at -1 dBTP for TikTok, Instagram and YouTube. LinkedIn sits nearer -16, and spoken-word podcast feeds are usually -16 stereo or -19 mono.
Should I use sidechain compression or keyframes?
Sidechain for long-form with continuous music, keyframes for short-form where there are only a few transitions. Keyframes give more precise control and avoid the artefacts a poorly tuned sidechain introduces on rapid speech.
How do I fix bad remote guest audio?
High-pass, gentle broadband noise reduction, careful de-essing and moderate compression, in that order. Aggressive noise reduction creates a watery artefact that sounds worse than the original noise. If the track is beyond repair, cut to B-roll and use a re-record.
Does music even help retention?
On energetic short-form, yes — it sets pace and covers edit points. On technical or authority-led content it often hurts, and a clean dialogue track with sound design accents performs better. Test it on your own audience rather than assuming.
Get a sample edit for Audio Ducking
Rather not do this yourself? Send us your footage and we'll apply all of it as standard.
Related pages
Export settings by platform
View page →
Podcast clip selection
View page →
Music licensing
View page →
Caption styling
View page →
Explore across the whole site
Industry, platform, pricing, comparison, guide and tool pages that pair with this one.
Industry
Video Editing for Luxury Brands — with restraint, craft and aspiration
Industry
Video Editing for Travel Brands — that sells the feeling of being there
Platform
Brand Film Editing — that makes people feel something before they buy
Platform
Webinar Repurposing — that turns one event into 30+ assets
Pricing
Video Editing Cost — what you actually pay in 2026
Pricing
Turnaround Time Cost — what speed actually costs you
Comparison
Retainer vs Per-Video — flexibility versus rhythm
Comparison
Agency vs Freelancer — an honest breakdown, including where we lose
Tool
Turnaround Estimator — realistic dates, not optimistic ones
Tool
Aspect Ratio Converter — exact crop windows and frame loss