Just in

๐ŸŽš๏ธ AI Audio Enhancer Tests: Real Repair vs Just Louder

I measured what AI audio enhancers do to a signal: LUFS, DNSMOS, word error rate. Most just level and brighten. Here is how to tell repair from paint.

Gene Park

Gene Park ยท Broadcast & Audio Engineering Writer

ยท 8 min read

โœ“ Fact-checked & production-testedBased on our own paid generations and published videos. Last reviewed 2026-08-14.How we test โ†’
โšก TL;DR โ€” quick answers
Is a free AI audio enhancer good enough for podcast or video work?
For a lightly noisy room, often yes. The catch is the caps, not the quality. Adobe's free Enhance Speech tier is published at one hour of enhancement per day, 30-minute files and a 500 MB ceiling, and the quota is charged against the length you upload rather than the section you keep. Per-stem control over speech, music and ambience sits behind the paywall, which is exactly the control you want on a mixed recording.
Does running an AI audio enhancer improve transcription accuracy?
Frequently it does the opposite. A published study ran MetricGAN+ denoising ahead of four modern ASR systems across 500 medical recordings and nine noise conditions, and the untouched noisy audio scored a lower semantic word error rate than the enhanced audio in all 40 tested configurations, with degradations reported from 1.1% up to 46.6%. Feed your transcription pipeline the raw file, and keep the enhanced version for human listeners only.
Can AI actually fix clipped or distorted audio?
Declipping is reconstruction, not recovery. iZotope's own guidance is blunt about it: information destroyed at the moment of recording is gone. Mild clipping, brief peaks around +0.3 dBFS, is close to fully repairable because there is enough surrounding waveform to infer from. Sustained heavy clipping leaves no reference, so fundamental pitch survives while harmonic structure stays permanently altered, which is why heavily declipped vocals keep that plastic, over-processed edge.
A dim broadcast edit suite at night, a single engineer's hands on a fader, twin loudness meters glowing green and a spectrogram of a spoken sentence stretched across the monitor behind them.

Key takeaways

  1. Gain-match the two files before you judge anything; most of the perceived win from one-click enhancers is loudness, and platform playback normalization at roughly -14 LUFS cancels that half anyway.
  2. Mask-based denoise can only subtract what is already in the file, while generative enhancers rebuild the waveform and can quietly change the words, so always diff the transcript.
  3. Hard-clipped peaks, speech under the noise floor, late reverberation in an untreated room and overlapping speakers stay broken regardless of which tool you pay for.
  4. Score the change with integrated LUFS, DNSMOS SIG/BAK/OVRL and word error rate against a reference transcript instead of trusting headphones at two different levels.

Half the tools I ran made the file louder and called that repair. Loudness-match the original against the output and switch between them blind, and on the worst offenders the entire "studio quality in one click" transformation shrinks to a modest lift in the presence band and a slightly tamer S. I have been metering dialogue for a long time and the meters keep saying the same thing: the category is two completely different technologies sharing one marketing word, and nobody selling it will tell you which one you just ran on your file.

That matters more than the feature grids suggest, because one of those technologies can only remove things that are already in your recording, while the other rebuilds your voice from scratch and can put words in your mouth that you never said.

By the numbers

Four number cards: 40 of 40 setups where raw audio won, a 1.1 to 46.6 percent degradation range, 500 medical recordings, and 4 ASR systems tested.
Cleaning the file before transcription made the words worse in every configuration the study ran.
  • Across 500 medical recordings and nine noise conditions, the original noisy audio scored a lower semantic word error rate than the denoised audio in all 40 tested configurations, with published degradations from 1.1% to 46.6% (arXiv).
  • YouTube and Spotify normalize playback to roughly -14 LUFS integrated. YouTube applies negative gain to anything louder and adds no gain to anything quieter, so "louder" from an enhancer does not survive the trip.
  • DNSMOS P.835 scores a single file on three axes with published Pearson correlation to human ratings of 0.94 for SIG and 0.98 for BAK and OVRL, so you can score before and after without owning a clean reference.
  • Adobe's free Enhance Speech tier is capped at one hour per day, 30-minute files, 500 MB, and the quota is counted on the uploaded length, not the part you keep.
  • Adobe users report best results at 30-50% strength, not 100%, where the documented complaints are metallic timbre, cut-off word endings and sibilant mis-guesses that read as a lisp.

Two engines, one label

Two-column table setting mask-based against generative enhancers on what they do, word invention, the null test, named tools and transcript safety.
Run the file twice and null the outputs, and the tool tells you which of the two technologies you bought.

Predictive / mask-based. Spectral masking, denoise, dereverb, source separation. The model estimates which time-frequency bins are speech and attenuates the rest. It is subtractive by construction. It cannot invent a syllable because it has no generator. Krisp and NVIDIA Broadcast live here, and so does most of what runs in real time on a call.

Generative resynthesis. Diffusion models, neural vocoders, speech-synthesis engines. These do not filter your waveform; they produce a new one that resembles it. Third-party technical write-ups describe Adobe Enhance Speech as a speech-synthesis engine that regenerates a voice mimicking the original speaker's timbre, which explains why its output can differ phonetically from the input rather than merely being cleaner.

The failure modes are documented. Diffusion-based enhancement produces phonetic insertions and substitutions, spurious breathing and hissing, robotic tones and high-frequency attenuation, and these are worst at low SNR while PESQ and STOI scores stay high. The ArtiFree work reports its ensemble-inference method cutting WER by 15% in low-SNR conditions, which tells you how much damage was there to cut. Researchers have had to propose Levenshtein phoneme distance and a hallucination error rate specifically because the standard quality metrics do not see hallucination at all.

The listening test: take a 10-second clip with one clear noisy passage and run it through the tool twice, changing nothing. Null the two outputs against each other (invert one, sum). A mask-based tool nulls to near silence because it is deterministic subtraction. A generative tool leaves audible residue, because it drew a new picture both times.

So what: run that null test once per tool and write the answer on a sticky note. It decides which files you are allowed to send it.

What no enhancer can recover

Table listing five unfixable problems with a one-line reason each: clipping, speech under the noise floor, late reverb, codec damage and crosstalk.
All five are capture failures, so no paid tier or model upgrade brings the missing signal back.
  • Hard-clipped peaks. Reconstruction, not recovery. Mild clipping around +0.3 dBFS is close to fully repairable. Sustained heavy clipping leaves nothing to infer from; pitch survives, harmonic structure does not.
  • Speech below the noise floor. At negative SNR there is no signal left to unmask, so anything the tool returns is invention.
  • Late reverberation in an untreated room. Deconvolution with an unknown room impulse response is ill-posed. Early reflections do not hurt intelligibility and are usually preserved; only the chaotic late tail is targeted, and stripping all of it sounds unnaturally dry. RT60 under roughly 500 ms is the target for clarity, and late reflections are the dominant cause of severe degradation in ASR.
  • Band-limited or codec-crushed sources. Bandwidth extension generates perceptually plausible top octaves rather than the originals, which research notes brings excessive percussive transients and noticeable timbre alteration.
  • Two people talking at once. Speech-only models collapse on overlap and pick a winner.

What most "enhancement" is when you meter it

Three operations, in various proportions: loudness normalization to a LUFS target, adaptive leveling that evens out segments and speakers within the file, and a de-esser sitting under a presence-band EQ lift. Auphonic is the honest one here because it keeps global normalization and adaptive leveling as separate documented controls, with targets for EBU R128 and ATSC A/85, -16 LUFS for mobile, -14 LUFS for YouTube and Spotify, -27 LUFS for Netflix, plus short-term and momentary ceilings.

Adaptive leveling is genuinely useful. The loudness half is theatre once your file hits a platform, for the reason in the numbers above. If you want the full picture of what survives platform playback normalization, the loudness standards guide has the delivery targets laid out.

So what: if a free online enhancer's only measurable change is integrated LUFS, you paid attention for nothing.

The protocol, run it for free

Numbered five-step flow: gain-match, log LUFS and true peak, score DNSMOS, compare word error rate, then compare the two spectrograms.
Score all three axes and a tool whose only real change was loudness has nowhere left to hide.
  1. Gain-match. Normalize both files to the same integrated LUFS before you listen. Without this the A/B is invalid and the louder file wins.
  2. Log the numbers. Integrated LUFS and true peak, before and after. ffmpeg -i file -af ebur128=peak=true -f null - costs nothing.
  3. Score perception. Run DNSMOS P.835 and record SIG, BAK and OVRL separately. Aggressive denoise typically lifts BAK while dropping SIG, and that split is the whole diagnosis.
  4. Score the words. Run one fixed ASR model on both files against a reference transcript and compute WER. Microsoft Research documented the underlying trade-off years ago: the loss terms that drive noise reduction also drive speech distortion.
  5. Read the picture. Open both spectrograms. A hard horizontal cutoff means a lowpass. Room tone that vanishes between words and reappears under speech means gating.

Accuracy work on the transcript side is covered in more depth in the captions and subtitle accuracy guide.

Artifact audit before you deliver

Listen once for each item, not all at once: sibilant lisp on S and T; clipped word endings; inserted breaths or gasps that were not performed; pumping on silence; musical noise, that shimmering birdie tone around the residual; timbre drift where the speaker no longer sounds like themselves.

Then do the thing almost nobody does: run word-level transcripts of the original and the enhanced file and diff them. Phoneme substitutions slip past ears that already know what the sentence was supposed to say. Your ear is the least reliable instrument in the room precisely because it is attached to a brain that fills gaps.

When to skip enhancement entirely

Anything headed for a transcription pipeline. Anything evidentiary: interviews, legal recordings, journalism where an invented phoneme is disqualifying. Music beds, which speech-only models will happily delete. Multi-speaker crosstalk, where the model guesses. If you are building multilingual versions on top of a source, keep the raw file as the master before it reaches any dubbing and translation stage.

Dosage and order beat tool choice

Strength at 30-50% usually outperforms 100%, which matches both the Adobe user reports and my own metering: the artifact curve rises faster than the benefit curve. Never chain two enhancers, because the second one treats the first one's artifacts as signal. Correct order: repair clipping, then dereverb, then denoise, then level, and normalize last. Normalizing first means every downstream stage gets a moving target.

Tool verdicts

Actually repairs: iZotope RX (surgical, per-problem modules, the only one where you see what you changed), Auphonic (leveling and loudness done properly and documented), Adobe Enhance Speech (real repair via resynthesis, with the resynthesis risks attached), Descript Studio Sound (same family, same caution), Cleanvoice (filler and mouth-sound editing rather than spectral repair), Krisp and NVIDIA Broadcast (real-time masking, subtractive, safe for words).

Loudness plus EQ wearing an AI label: most of the browser one-click sites. Honest limitation: I could not verify what any of them actually run, because none of them publish the model. What I can report is that gain-matching removed the majority of their advantage, and the null test came back deterministic on the ones I could run twice.

Free tiers move constantly. Only Adobe's caps above are documented; treat the rest as unverified and check quotas before you build a workflow on one. Costing out a full audio stack is easier with the cost calculator.

I spent years in broadcast post where the noise reduction was a physical box and every setting was a knob you had to defend to a chief engineer. The lesson that transferred: a dialogue editor once handed back a reel I had cleaned within an inch of its life and said the room had died. He was right. The gate had removed the air between lines, and without air the cuts stopped being invisible. Modern models make the same mistake at a thousand times the speed, and the tell is identical. If the silence between your words is more silent than a real room, listeners will not name the problem, but they will feel that something is wrong with the person talking. Leave the room in.

The unrecoverable list is a capture problem, not a software problem. Get the mic closer, set gain with real headroom so nothing kisses zero, and kill the first reflection off the desk and the wall behind you. A moving blanket beats every enhancer on this page for the categories no model can fix, and it costs about the same as one month of a subscription.

Sources & further reading

Outside figures cited above. First-hand test results are our own and noted as such in the text.

  1. ArtiFree: Detecting and Reducing Generative Artifacts in Diffusion-based Speech Enhancement โ€” arXiv
  2. When De-noising Hurts: A Systematic Study of Speech Enhancement Effects on Modern Medical ASR Systems โ€” arXiv
  3. How to Fix Audio Clipping โ€” iZotope

Frequently asked questions

โ–ธIs a free AI audio enhancer good enough for podcast or video work?

For a lightly noisy room, often yes. The catch is the caps, not the quality. Adobe's free Enhance Speech tier is published at one hour of enhancement per day, 30-minute files and a 500 MB ceiling, and the quota is charged against the length you upload rather than the section you keep. Per-stem control over speech, music and ambience sits behind the paywall, which is exactly the control you want on a mixed recording.

โ–ธDoes running an AI audio enhancer improve transcription accuracy?

Frequently it does the opposite. A published study ran MetricGAN+ denoising ahead of four modern ASR systems across 500 medical recordings and nine noise conditions, and the untouched noisy audio scored a lower semantic word error rate than the enhanced audio in all 40 tested configurations, with degradations reported from 1.1% up to 46.6%. Feed your transcription pipeline the raw file, and keep the enhanced version for human listeners only.

โ–ธCan AI actually fix clipped or distorted audio?

Declipping is reconstruction, not recovery. iZotope's own guidance is blunt about it: information destroyed at the moment of recording is gone. Mild clipping, brief peaks around +0.3 dBFS, is close to fully repairable because there is enough surrounding waveform to infer from. Sustained heavy clipping leaves no reference, so fundamental pitch survives while harmonic structure stays permanently altered, which is why heavily declipped vocals keep that plastic, over-processed edge.

โ–ธHow do I prove an enhancer helped rather than just sounded different?

Gain-match both files to the same integrated LUFS first, since a louder file wins blind comparisons for reasons that have nothing to do with quality. Then log integrated LUFS and true peak, run DNSMOS P.835 for SIG, BAK and OVRL scores, and run one fixed ASR model against a reference transcript to compare word error rate. If all three axes move the right way, it helped. If only loudness moved, it did not.

The 5 best AI video finds, every week

New models, tested prompts, and what actually worked in our production โ€” one short email a week. No spam, unsubscribe anytime.

Gene Park

Written by Gene Park

Broadcast & Audio Engineering Writer

Spent a career in TV post-production before the AI wave and still trusts meters over marketing. Measures loudness, codecs, and artifacts on everything before a single word gets written.

#ai audio enhancer#ai audio enhancer free#ai audio enhancer online#enhance audio quality ai#speech enhancement measurement

Keep learning