Why ACX Rejects Files That Sound Fine
June 27, 2026 · 6 min read
You finished the book. You listened back. The voice is clear, the room sounds quiet, nothing seems obviously wrong. You upload to ACX.
The rejection email arrives.
This happens to thousands of narrators every year, and the reason almost never has anything to do with performance. It has to do with the gap between what your ear hears and what ACX's automated system measures. Those are two completely different things.
Loudness Is the Number One Failure — and the Most Misunderstood
The most common reason files get rejected is loudness. Not too loud or too quiet in an obvious way — but off in ways that only a meter can detect.
ACX specifies an RMS range of -23 to -18 dBFS. What most narrators don't realize is that RMS and integrated loudness — measured in LUFS — are not the same measurement. RMS measures raw electrical amplitude without any filtering. It treats a 20 Hz bass rumble and a 1,000 Hz vocal tone at the same dB level as equally loud. Integrated loudness applies frequency-weighting curves that account for how human hearing actually works. We're far more sensitive to midrange frequencies than to sub-bass or very high frequencies.
A deep, bass-heavy voice and a crisp, present voice can read identically on an RMS meter while sounding completely different to a listener.
Two specific scenarios trip up narrators constantly.
The long pause trap. RMS meters don't ignore silence — they factor every second of dead air into the rolling average. A narrator recording dramatic fiction with long emotional pauses between paragraphs will find that the silence drags their RMS average down. To hit the -23 dBFS floor, they push up the gain. Now every time the voice comes back in, it hits listeners like a door slamming open. LUFS meters use a gate that pauses during silence, which is why integrated loudness gives a far more accurate picture of perceived volume.
Tonal coloring. A male narrator recording close to the microphone picks up proximity effect — excess bass buildup that inflates the RMS reading without making the voice sound louder. A narrator with prominent sibilance has sharp transient spikes on every "S" and "T" that barely register on a rolling RMS average but fatigue a listener on headphones within minutes. Both files can pass the RMS check. Neither sounds right.
Understanding the ACX Noise Floor Limit (≤ -60 dBFS)
When you sit in your recording space and hear silence, your brain is doing something remarkable: filtering out everything that's always there. The refrigerator hum three rooms away. The computer fan. Traffic vibrations moving through the floorboards. Auditory adaptation is so effective that your ears register those sounds as zero.
Your microphone does not adapt. It captures all of it.
ACX requires a noise floor below -60 dBFS. A file sitting at -55 dBFS doesn't sound dramatically different when you're listening back casually in your recording space. But on noise-canceling headphones, the difference is between a distant, ignorable hiss and an active distraction — a constant fuzzy presence that surfaces every time the narration pauses. A listener's brain can't adapt to it. Every quiet moment breaks the story.
True Peak: The Clipping Your DAW Won't Show You
Standard DAW meters show sample peaks — the individual digital snapshots captured 44,100 times per second. What they don't show is what happens between those snapshots when the file is converted back to analog sound for a listener's headphones.
The reconstruction process draws a smooth curve between digital data points. That curve can overshoot the data points themselves. A file where the highest sample sits at -3.1 dBFS can produce a true peak of -2.5 dBFS — technically above the ACX limit of -3 dBFS — with zero audible distortion inside your DAW.
The problem surfaces after export. When platforms compress your WAV to MP3, the encoding process causes peaks to swell by 1–2 dB. On consumer playback hardware, that overflow clips hard. Listeners hear subtle graininess on sharp consonants. Your DAW meter said everything was fine.
Why 48 kHz Files Get Rejected — and Why Stereo Costs You Quality
A stereo file of a single voice sounds identical to a mono file — it's the same signal playing out of both speakers equally. ACX accepts stereo, but mono is strongly preferred, and the reason is mathematical: a stereo file at 192 kbps splits that bitrate across two channels, leaving 96 kbps per side, which introduces digital artifacts. ACX requires all 192 kbps dedicated to a single mono track.
Other inaudible format failures:
- Variable bit rate encoding. ACX requires constant bit rate. VBR lowers the data rate during silences — smart for file size, automatic rejection for ACX.
- 48 kHz sample rate instead of 44.1 kHz. Undetectable by ear. On legacy audio systems that can't resample properly, your audio plays back roughly 8% slower and lower in pitch.
- Corrupted ID3 metadata tags. ACX uses automated systems to inject its own chapter and tracking data. Bloated or conflicting tags from your DAW can crash the ingest process before a human ever hears the file.
- Joint Stereo encoding. Joint Stereo combines low frequencies into mono to save space. It sounds fine. The automated scanner looks for a true single-channel mono block and rejects the two-channel header.
What the Rejection Email Actually Tells You
ACX rejection emails identify which metric failed. They do not tell you where in the file the problem is.
"Chapter 4 fails peak amplitude requirements" tells you the chapter. It does not tell you whether the issue is at minute 4 or minute 44. You scroll the entire waveform hunting for a single transient spike — a sharp laugh, a dog bark in the background, one hard consonant that hit too loud.
There's a second trap in chapter-level review. Fix only the flagged chapter and leave the rest of the book untouched, and the inconsistency in processing becomes its own rejection.
Technical spec comparison
| Check | Before | After | Required |
|---|---|---|---|
| RMS Loudness | −30.4 dBFSfail | −22.8 dBFS✓ | −23 to −18 dBFS |
| True Peak | −2.1 dBFSfail | −4.2 dBFS✓ | ≤ −3 dBFS |
| Noise Floor | −55.2 dBFSfail | −71.5 dBFS✓ | ≤ −60 dBFS |
| Sample Rate | 48,000 Hzfail | 44,100 Hz✓ | 44,100 Hz |
| Channel Format | Mono✓ | Mono✓ | Mono (preferred) |
| Bit Rate | WAV/PCM✓ | 192 kbps MP3✓ | 192 kbps MP3 or WAV |
| Head Room Tone | 3.55s✓ | 1.00s✓ | 0.5 – 5.0s (1s+ recommended) |
| Tail Room Tone | 0.79sfail | 3.04s✓ | 1.0 – 5.0s |
Why Manual Fixes Almost Always Backfire
When narrators attempt to correct rejections themselves, they typically hit what I call the seesaw effect: adjusting one metric breaks another.
Turn up master gain to fix a quiet RMS reading — the peaks now clip. Apply an aggressive noise gate to silence the room between phrases — the pauses become an unnatural vacuum that a human reviewer catches immediately. The listener hears a warm, natural room sound while the narrator speaks, then a completely dead silence every time they pause. That contrast sounds wrong even to someone who can't name why. Compress an already-exported MP3 a second time — encoding artifacts compound and the noise floor rises. Fix one chapter and leave the rest alone — the book now sounds inconsistent.
Escaping the loop requires applying corrections in a specific order: high-pass filter first to remove low-end rumble, then EQ and de-essing to tame tonal problems, then gentle compression to narrow the dynamic range, then true peak limiting last. Change the sequence and you're fighting the math instead of working with it.
Getting It Right Before Submission
The goal is to make ACX compliance mechanical — a repeatable chain that produces the same result on every chapter, every time.
That means proper gain staging during recording (peaks around -12 dBFS, with headroom to spare), treating the room before turning on the mic, and running a pre-flight check instead of uploading and hoping. Your final WAV master should read -21 dB RMS and -3.5 dBFS — when converted to a 192 kbps CBR MP3, it lands comfortably inside ACX boundaries.
I built Greenlit Audio because the manual version of this process costs real money. A freelance audio engineer runs $50 to $150 per finished hour. An 8-hour book is $400 to $1,200 in engineering fees before the first sale. Turnaround adds days to a project that's already months in the making. And even experienced engineers miss a 1-millisecond inter-sample peak sitting at -2.9 dBFS. The software doesn't get tired. It applies the same math in the same order to every chapter.
Run the free diagnostic before you submit. It checks all nine spec points — chapter length, RMS, true peak, noise floor, head and tail room tone, channel format, sample rate, and bit rate — and tells you exactly what's wrong while there's still time to fix it.
See how your files measure up
Upload a chapter and get a 9-point ACX compliance report in under a minute — free, no account required.
Run a free analysis →