Can AI Narration Pass Audible's ACX Requirements?

By SpeakBreez Team Updated Jun 22, 2026 6 min read

AI narration can meet every one of ACX's audio requirements, but confirm Audible's current narration policy before you submit, because the rule on synthetic voices changes over time. We master AI audio to ACX spec in our own audiobook pipeline, so the numbers and steps below are the exact ones that pass the automated check. The audio bar is fixed and testable; the policy is the part to verify, and the cost difference versus a human narrator is large.

Key takeaways
  • ACX needs 44.1 kHz, 192 kbps CBR MP3, RMS -23 to -18 dB, peak under -3 dB, noise floor below -60 dB.
  • Raw text-to-speech usually fails on sample rate and loudness, not voice quality.
  • A five-step mastering pass fixes every technical rejection.
  • Human narration commonly runs $200-$400 per finished hour; AI narration is a fraction of that.
  • Meeting the spec does not guarantee policy approval; check current ACX terms.

What are ACX's audio requirements?

ACX accepts MP3 files at 44.1 kHz, 192 kbps CBR, with loudness, peak, and noise-floor limits, plus room tone per file. Every value is testable before upload, so rejection is preventable.

RequirementACX spec
FormatMP3, 192 kbps+, constant bitrate (CBR)
Sample rate44.1 kHz
Loudness (RMS)-23 dB to -18 dB
Peakno higher than -3 dB
Noise floorbelow -60 dB RMS
Room tone0.5-1s opening, 1-5s closing, per file
FilesOne per chapter, under 120 min, matching credits

Why does raw text-to-speech fail ACX?

Raw TTS usually fails on sample rate and loudness, not the voice. Many engines export at 24 kHz or a variable bitrate, and the loudness sits outside the -18 to -23 dB window. The clip sounds fine in a browser, then ACX's automated check rejects it. In our pipeline, more than nine of ten first-pass rejections we have seen trace to these two settings, not the narration itself.

How do you master AI audio to ACX spec?

Run the audio through five steps: resample, normalize, limit, add room tone, and export as CBR MP3.

  1. Resample to 44.1 kHz, since most TTS output is 24 kHz.
  2. Normalize loudness into the -18 to -23 dB RMS window.
  3. Apply a peak limiter set to -3 dB so no sample clips the ceiling.
  4. Add about 0.75 seconds of room tone at the start and 3 seconds at the end.
  5. Export as 192 kbps constant-bitrate MP3.

Done correctly, the file clears ACX's audio check on the first try. The self-serve route is on the text to speech for audiobooks page.

How do you handle room tone and chapters?

Give each chapter its own file with short room tone at both ends and consistent opening and closing credits. Room tone is a brief stretch of near-silence that keeps the file from starting or ending abruptly; ACX rejects files that cut in cold. Split the book one file per chapter, keep each under 120 minutes, and use the same intro and outro lines on every file so the title reads as one production.

Will Audible accept an AI voice?

Meeting the audio spec does not settle policy, so check current ACX and Audible terms before recording a full book. As of 2026, Audible has been expanding programs around virtual voices, and rules differ by program and region. Treat the spec as the controllable part and the policy as the part to confirm in writing first. This is the trade-off we are most careful to flag, because a spec-perfect file can still be the wrong path for a given title.

What does an AI audiobook cost versus a human narrator?

Human narration commonly runs $200 to $400 per finished hour, so a 6-hour book is roughly $1,200 to $2,400; AI narration is a small fraction of that. A 50,000-word book is about 300,000 characters, which fits inside one SpeakBreez Pro month at $49 with room to spare. The other saving is revision: a human re-record means re-booking and tone-matching, while AI regenerates a corrected line in seconds at no extra cost.

Should you DIY or hand it off?

DIY if you enjoy audio editing and have one book; hand it off if you want a submission-ready result without learning mastering. The mastering math is not hard, but it is fiddly and unforgiving on a 7-hour book. Book Studio writes, narrates, masters to ACX spec, and prepares your ebook and audiobook for KDP and Audible, and it can narrate in your own cloned voice.

When is AI narration the wrong call?

Skip AI narration when the book lives on vocal performance, like memoir or fiction that needs acted emotion. AI narration is strong for clear, information-led nonfiction. For deeply performed work, a human narrator still wins, and listeners notice. Match the method to the book rather than forcing every title through the same pipeline.

Frequently asked questions

Does AI narration cost less than a human narrator?

Yes, by a wide margin. Human narration runs $200-$400 per finished hour; AI narration costs a fraction and regenerates instantly when you edit.

Can I fix one word without re-recording?

Yes. Change the word in the script, regenerate that line, and splice it in. No studio session needed.

What sample rate does ACX require?

44.1 kHz. Many TTS engines default to 24 kHz, so a resample step is almost always required.

How long can each audiobook file be?

Up to 120 minutes per file, with one file per chapter and matching opening and closing credits.

Is AI narration allowed on Audible?

The audio can meet spec, but acceptance depends on Audible's current virtual-voice policy, which varies by program and region. Confirm before submitting.

What bitrate and format does ACX want?

192 kbps or higher, constant bitrate, MP3. Variable bitrate files are rejected.

How much room tone do I need?

0.5 to 1 second at the start of each file and 1 to 5 seconds at the end, with no abrupt cuts.

Next steps

Want to test narration now? Start free. Comparing voice tools first? Read our SpeakBreez vs. ElevenLabs vs. Murf comparison.

Try SpeakBreez free

800+ AI voices, 120+ languages. No credit card.

Start free →
← Back to all posts