A voice description is a spec the AI executes, not a mood you hope it guesses. The traits that move a generated voice are specific and learnable: age band, register, accent and region, pace, texture, emotional default, and the context the voice will be used in.
This guide walks through what to include, what to leave out, and how to read the three variations Voice Design hands back on every run — because judging across three attempts, not re-rolling the first miss, is most of the skill. Start building at voice design or browse worked prompt examples as you read.
A Description Is a Spec, Not a Wish
Most weak descriptions read like a mood board: warm, friendly, professional. Those words are true of almost every voice in the library, so they narrow nothing. A spec is different. It fixes the variables that actually change what comes out, in language specific enough that two people reading it would picture the same speaker.
Think of it the way you would brief a voice actor at a casting call, not the way you would describe a vibe to a friend. A casting brief gives an age, a register, a regional accent, a pace, and a scenario. That is exactly the information Voice Design needs too.
The Traits That Actually Move the Result
Six traits do almost all the work. Get these right and the adjectives around them matter far less than most people assume.
Age band
Younger, middle-aged, or senior. Voice weight and breath control shift with age far more than most people expect.
Register
Chest-heavy and low versus light and airy. This is the single biggest lever on how deep or bright a voice reads.
Accent and region
Name a specific variety — General American, Received Pronunciation, Dublin — not a whole country. Specificity is what the model can act on.
Pace
Unhurried, brisk, or measured. Pace changes how authoritative or energetic a voice feels almost as much as tone does.
Texture
Breathy, gravelly, clear, or worn. This is the grain of the sound, independent of pitch or accent.
Emotional default and context
The resting mood the voice returns to, plus where it will be heard. Context tells the model how much energy to hold in reserve.
Judge Across Three, Not One
Voice Design returns three variations from a single description, and that is deliberate. The model reads the same brief slightly differently each time, so one run rarely lands the target exactly. Listen to all three before you touch the wording again.
If two of the three land close to what you wanted, that is a sign the description is doing its job — refine details rather than rewriting from scratch. If none of the three are close, the description itself is the problem, and changing a trait or two will do more than generating another batch from the same text.
What Doesn't Work
A few habits show up constantly in weak prompts, and all of them are avoidable.
- Naming a celebrity or character. Even asking for something "inspired by" a named figure requests a result the system won't and shouldn't reproduce. Describe the sound instead — a bouncy, energetic tone or a gravelly, weathered one — and you get a usable, original voice.
- Asking for a pitch in Hz or semitones. There is no pitch slider and no formant control. Depth comes from register, texture and pace, described in words, not from a number.
- Stacking six adjectives that fight each other. "Warm, cold, bright, gravelly, soft, commanding" cancels itself out. Pick two or three traits that could plausibly belong to the same speaker.
- Reaching for emotion tags. There is no per-line emotion control. Temperature adjusts expressiveness — how much the delivery varies — not a named emotion and not pitch.
Heads up: Naming a real or fictional character — even a public figure everyone would recognize — is off-limits and won't produce a reliable result anyway. Describe the qualities you actually want: age, register, texture, pace.
Turn a description into three voices in one pass
Write the spec, generate three variations, and pick the one that fits — no re-rolling from scratch.
Open Voice DesignWeak vs Strong: A Rewrite
Here is the same voice, described two ways. The first is a mood board. The second is a spec.
A warm, friendly, professional voice that sounds nice for videos.
✅ A woman in her late 20s, General American accent, clear and unhurried pace, warm register with a slight breathiness, explaining ideas conversationally for a YouTube tutorial.
The difference is not length. It is that the second version fixes age, accent, pace, register, texture and context — six variables instead of zero. Feed that into Voice Design and the three variations it returns will actually sound like each other, which is what makes picking the best one possible instead of arbitrary.
Writing for a Specific Job
The same six traits shift depending on what the voice is for. A narrator, a brand voice and a customer support line all want a different balance of the same ingredients.
Prompt examples worked through
See full descriptions written for narration, ads, and storytelling voices, trait by trait.
Consistent brand voice
Lock a description once so every ad, explainer and support line comes from the same spec.
Design vs clone
Understand when to write a description and when to clone an existing recording instead.
Design Now, Clone Later
A written description builds a new voice from nothing. Cloning takes an existing recording and reproduces it, accent and pitch included, which is a different tool for a different job. If you already have a voice you want to reuse rather than invent, cloning is the shorter path.
Read more in AI voice cloning for creators, or compare the two approaches directly in voice model vs. voice clone.
Audition Voices Built From Descriptions Like These
The six voices below were all designed from specs with the same structure: age, register, accent, pace, texture and context. Listen for how distinct they are from each other despite sharing a template. That is the six traits doing their job. Browse the full voice library for more.
| Voice | Accent | Tone | Rating |
|---|---|---|---|
| Intimate Male Audiobook Narrator | Neutral American | Intimate, warm | ★★★★★ |
| Curious Female Explainer Voice | Neutral American | Curious, fast-paced | ★★★★★ |
| Deep Gruff Male Narrator | Neutral American | Deep, gruff, commanding | ★★★★★ |
| Calm British Male Narrator | British | Calm, cordial | ★★★★★ |
| Bright Expressive Female Storytelling | Neutral American | Bright, expressive | ★★★★☆ |
| Gravelly Male Documentary Voice | Neutral American | Gravelly, weathered | ★★★★★ |
What It Costs to Experiment
Generating and auditioning voices costs points, priced by character count rather than a flat per-run fee — roughly one point per second of resulting audio, though the exact figure depends on how much text you generate. A free account includes a one-time trial of 60 points and one saved design, enough to test a handful of descriptions before you commit.
Lite ($9/mo) raises that to 3,000 points a month with one saved design. Pro ($24/mo) includes 12,000 points and five saved designs, which is the tier most people land on once they are iterating on more than one voice at a time. Full breakdown on the pricing page.
Frequently Asked Questions
What actually changes how a designed voice sounds?▾
Age band, register, accent and region, pace, texture, and the emotional default and context you describe. These six traits do almost all the work; extra adjectives layered on top add little.
Can I ask for a specific celebrity or character's voice?▾
No. Naming a real or fictional person is not supported, and it would not produce a reliable result anyway. Describe the qualities you want instead — age, register, texture and pace — and Voice Design builds something original from that.
Can I request a specific pitch, like a deeper or higher voice in semitones?▾
There is no pitch slider and no semitone or formant control. Depth and brightness come from how you describe register, texture and pace, not from a numeric setting.
What does the Temperature control actually do?▾
It adjusts expressiveness, meaning how much variation the delivery has. It is not emotion and not pitch, and there are no separate emotion tags or per-line emotion controls.
How many voices do I get from one description?▾
Three variations from a single run. Compare all three before rewriting the description; if two of them are close, refine details rather than starting over.
Write the spec, hear three voices back
Describe age, register, accent, pace and texture, then pick the variation that fits.









