My client loved the explainer video, but hated the AI voiceover.
Me: “No problem! I’ll do it myself.”
But my 5-year-old had other plans. She’s known for asking a question every 15 seconds. And the few times she was silent, the AC kicked on and started humming in the background.
I needed 90 seconds of clean audio. But after 30 minutes of interruptions, I rage quit and sent the AI voiceover instead.
The HyperFrames Background
I LOVE HYPERFRAMES.
Distilling a complex topic down to a minute-long explainer video has been a game-changer. They are even better with voiceovers because they can add storytelling, context, and emotion.
However, the two ways of adding this to a HyperFrames video are both… not ideal.
Stock AI. They are precise, but people hate them because they sound fake and generic.
Record your voice. Great if you have a good mic, a good voice, and a moment of silence.
Clients despise #1. I despise #2. There had to be a better way…
The Podcast Room Tax
Here is the frame I keep coming back to.
Every voiceover charges you the same three things:
A quiet room
A good mic
A clean take
Call it the podcast room tax. And it’s getting more expensive every day given that I’m generating more and more video content.
The solution? What if you could pay the tax once and benefit forever? This is what I explored this morning. I wanted to record my own voice and then feed it into AI to reproduce my tone, cadence, and voice.
The initial result? Not bad.
Yes, it needs some tweaking… but it’s getting there.
Getting Started with xAI Custom Voice
Here’s how I did it.
Go to console.x.ai
Open Custom Voices
Record your voice live in the browser by reading the passages.
Copy the ID for your voice template and hand it to your agent.
Now I know what people are thinking. But can’t everyone use this? The constraint is the consent model. IMHO it is the right one. There is no “paste in a podcast episode and become that person” path. You must show up, talk, and use your own voice.
What surprised me
My clone talks faster than the stock narrator, 3.13 words per second against 2.33, so a script written for the old voice finishes early.
TTS also reads your typos aloud (so be sure to spell-check!).
It still feels too “annoucer-y”, especially in the beginning of the recordings. But as they keep playing, I feel it sounds more and more like me.
It’s cheap! Something like $5 for an hour of audio is way cheaper than I expected.
It’s fast! From request to done takes barely longer than the audio itself.
It’s addictive! I’ve already generated 5-6 different audio takes now that I have the skill baked and ready.
My first use of HyperFrames wasn’t perfect, but I’m sharing it as is just to timestamp my first experience. I expect that this will get better with each iteration to the point where you won’t be able to tell at all.

