Seed Audio 1.0
by ByteDance
ByteDance's all‑in‑one TTS model.
Voice, music, and sound effects in one generation.

Key features
Hear the range
A warm, measured documentary voice-over.
A hushed, tense line read, close and intimate.
A layered open-air market sound bed.
A rolling storm building to a distant thunderclap.
A short rising cue for strings and brass.
A relaxed beat with soft keys and vinyl crackle.
Technical specifications
ByteDance
Developed by ByteDance's Seed research team.
20
English, Chinese, Japanese, Korean, Spanish, French, German and more.
Up to 2 min
Maximum two minutes of generated audio per pass.
3,000 chars
Room for a full scene brief, dialogue included.
Up to 3
Up to three reference clips, each up to 30 seconds.
Non-streaming
Renders the complete track, not a realtime stream.
Use cases
Audiobooks
Narration, character voices, and sound design for a full book. ByteDance puts the cost near a tenth of studio recording.
Video dubbing
Describe the voice or upload a character image, then use timestamps to land each line exactly where the picture needs it.
Game audio
Character barks, scripted performances, and environmental sound effects for immersive scenes, generated from the script.
One-pass video audio
Give a video clip its narration, sound design, and score in one generation, with no separate mixing step afterward.
Ads and promos
A spoken line, sound effects, and music as one ready-to-use track, made for short-form content.
Dialogue and audio drama
Multiple characters, each with a distinct voice and delivery, in one scene with matching ambience and timing.
Prompt examples
Timestamp control
Ryan (warm, breathless): '[5.5s:8.0s] Maya! Wait, you're leaving tonight?'
Edit promptAudiobook scene
Rain on a library window. Narrator, low and unhurried: 'She read it twice.'
Edit promptSports commentary
Packed stadium, roaring crowd. Commentator, exhilarated: 'OH, HE SCORES!'
Edit promptSimple pricing
Get started for free today, with the option to upgrade or cancel anytime.
Basic
1100 monthly credits
1 user only
All models
Workflows
Standard
3625 monthly credits
1 user only
All models
Workflows
Pro
6350 shared monthly credits
1 user
All models
Workflows
Pro Max
24650 shared monthly credits
1 user
All models
Workflows
Enterprise
For higher limits
Custom
pricing and billing terms

Free
For playing around
$0
forever free
FAQs
Gemini Omni Flash 1.1
Google DeepMind
Google's video model where the edit is a sentence. Now with 4K output and 40-second scenes.
MiniMax H3 Max
fal.ai
fal's post-trained MiniMax H3. Tuned for prompt adherence, rebuilt for speed.
Runway Ruby
Runway
Runway's SDR to HDR model. Lifts finished video into 16-bit EXR frames and ProRes masters.
Hyperion 2.5
Topaz Labs
Topaz Labs' SDR to HDR model. Turns 8-bit video into 10-bit ProRes and 16-bit EXR masters.