The Audionovel: Are We Ready for a New Hybrid?
- Jul 6
- 4 min read

Steve Jobs did not invent the term podcasting when he walked on stage with the first click-wheel iPod. The word came from British journalist Ben Hammersley, writing in the Guardian, years before Apple built a business around the idea.
I've owned a Walkman, a Discman, and eventually an iPod — and I remember the specific tedium of ripping an entire music library onto it, one CD at a time, on a Sunday afternoon. Today, podcasts and audiobooks are a fully mature ecosystem. At one end of the audiobook spectrum sit fully dramatized productions — multiple voices, sound effects, music, produced like television, channeling theater-of-the-mind radio from the 1930s and '40s. At the other end sits straight narration. If you're a longtime audiobook listener, you already know how much difference a good voice makes in "selling" a story.
I spent close to a quarter century running a media production company — the kind of place with illustrators and art directors on staff, cutting narration off reel-to-reel, sitting in the booth on take sixteen just trying to get the read right. So when I looked at today's audio landscape, the question that occurred to me wasn't dramatized or narrated — it was whether there's a third lane. Something between fully produced drama and a single voice reading straight through.
I've been calling my working concept the audionovel. It's a hybrid: narration, multi-voice dialogue, music and sound design, built around the specific pleasure of good prose — the kind you hope to find in a novel, not a screenplay. Good prose can set a scene with backdrop, lighting, and even aroma. Dialogue alone can't do that. Not every manuscript lends itself to this treatment, but I think the GW Canyon books can, because the prose leans on character dialogue as much as narrative description.
Unlike adapting a novel into a film or TV series, the audionovel stays anchored in the pleasure of the prose itself. It doesn't translate the book into a different medium's grammar — it just gives the prose a voice, or several.
The practical math
Here's where theory meets budget. Depending on the voice talent, a traditionally produced audiobook runs anywhere from $2,000 to $10,000 — and the largest single line item is almost always the narrator's fee.
Like every other corner of the creative world, AI is creeping into audio production. It reminds me of when airbrush and retouching artists were replaced by pressure-sensitive tablets and Photoshop — a changeover I watched happen in real time, from the other side of the table.
A company called ElevenLabs has built a genuine beachhead in AI-generated voice and likeness. Visit their site, and you'll find narrators modeled on Michael Caine, Laurence Olivier, and John Wayne, available for licensing. We all knew this was coming. It's here.
ElevenLabs offers a large menu of voices to audition — which brought back a very specific memory of the reel-to-reel and cassette submissions we'd get from talent agencies, reviewing voice after voice for an industrial video or a slide show client. The difference now: most of these voices cost nothing to audition. No talent fee, no union rate, no residuals.
The current version also lets you add limited stage direction — type a bracketed note like [slow pace, sound fearful], and the narration adjusts. You can regenerate for a different "take," the same way you might ask a voice actor to run it again with more urgency.
What I actually found
I took the opening pages of TwinStar and fed them into the production window. It worked. It was, genuinely, a little magical. But the longer I worked the interface, the clearer it became that this is a usable 1.0 — I can only imagine what a 4.0 version will do.
And it's here that the limits of the tool showed up most clearly: years of cutting reel-to-reel takes, hours of mixing and crossfading music, the accumulated judgment of knowing which take actually serves the story — none of that transfers to a model. AI has no sound booth behind it. Human intuition is still the engine, and it will be for a long time.
It took several sessions — swapping voices, adjusting characteristics, learning the right sequence of bracketed direction — to get a performance I'd stand behind. Honestly, I could have gotten there faster with live talent and a casting call. But I did this on a Sunday afternoon, at my own desk, with no contracts and no studio time booked. For roughly the cost of lunch for two, I walked away with voices, sound effects, and music I'd designed myself.
Then came the part that hasn't changed in fifty years: the edit and the mix. I applied filters so the pilots sounded like they were on cockpit headsets. I dropped in mic clicks. I looped the ambient jet noise and balanced it against the dialogue. I checked the mix on studio speakers and again on headphones — because there's always a difference, and there always will be.
It took roughly five hours to produce two and a half minutes of finished audio. I'd chalk about 20% of that up to the learning curve — the rest is just what good audio takes, regardless of who or what generated the raw material.
So, is it viable?
The democratization of creative tools has been underway since the early 2000s. Gear that once only Hollywood and network production houses could afford is now available to anyone for a fraction of the cost — and that collapse in price is exactly what set YouTube, TikTok, and podcasting ablaze with creative output.
Paired with AI tools, that same collapse may make something like the audionovel both easier to produce and more appealing to actually buy. My answer, after this first pass: it's viable — but not for every book. It works when the prose itself is built on dialogue the way GW Canyon's is; it would fall flat on a manuscript that leans on interior narration with little spoken exchange. The test isn't the technology. It's whether the book was already halfway to a script before anyone opened a production tool.
I've posted the 2:30 pilot on the site. Have a listen and tell me — does it sound like something worth finishing?



Comments