Seed Audio logoSeed AudioSeed Audio

One prompt for voice, music, and sound effects

Seed Audio is ByteDance's non-streaming audio model that generates lifelike speech, original music, and cinematic sound effects in a single pass. Explore what the model can do, then start creating.

What Seed Audio can do

The model moves past text-to-speech to text-to-any-audio — one system for voice, music, and sound.

Speech & voice cloning

Generate natural, expressive speech directly from text, with fine control over pace, emotion, and accent. Provide a short reference clip and the zero-shot model reproduces that voice, keeping it consistent across every line.

Original music generation

Turn a written prompt into original music with the Seed-Music component — melody, arrangement, and mood composed to match your description, without samples, loops, or a session player.

Cinematic sound effects

Layer foley, ambience, and effects with SeedFoley so every scene lands with the right texture — footsteps, weather, room tone, and impacts placed naturally under the voice and music.

Cross-lingual synthesis

Speak across languages without fine-tuning. A single request can carry one voice's identity from one language into another, which makes localized narration and dubbing far quicker to produce.

Multi-character dialogue

Render a full conversation with a distinct voice for each speaker, generated together in one end-to-end pass instead of recording and stitching separate clips line by line.

Single-pass generation

Voice, music, and sound effects are produced together rather than assembled afterward, with up to two minutes of coherent audio per request and consistent character throughout.

Lifelike speech & voice cloning

Type a script and the model speaks it with natural intonation, emotion, and accent. Add a short reference clip to clone a voice, then keep that voice consistent across every line — and even across languages. It suits narration, audiobooks, explainers, and character work where a steady, believable voice matters.

Speech + voice cloning

Speech + voice cloning

Original music from a prompt

Describe a mood, genre, tempo, or scene and the Seed-Music side composes original music to match. Because the audio is generated rather than sampled, you can iterate on a direction in seconds and layer the result under narration or dialogue. Preset styles and reference audio give you a starting point when you already know the sound you want.

Text-to-music

Text-to-music

Cinematic sound effects & foley

SeedFoley adds ambience, foley, and effects so footsteps, weather, and room tone sit naturally under your voice and music. Design a full soundscape from a written description instead of hunting through effect libraries, then export the layered mix in one go.

Sound effects + foley

Sound effects + foley

Pricing

Choose the plan that works for you.

Starter

$9/mo

For individuals

  • 1 project
  • 5,000 credits
  • Email support

Pro

$29/mo

For growing teams

  • Unlimited projects
  • 50,000 credits
  • Priority support
  • API access

Enterprise

$99/mo

For large organizations

  • Everything in Pro
  • Unlimited credits
  • Dedicated support
  • Custom integrations

FAQ

Seed Audio FAQ

Common questions about the model and this platform.

What is Seed Audio?

Seed Audio is ByteDance's universal audio generation model, introduced by the ByteDance Seed team in 2026. Unlike traditional text-to-speech, which only reads text aloud, it generates human voice, music, and sound effects from a single text prompt in one end-to-end pass. The underlying research draws on the Seed team's earlier work on Seed-TTS, Seed-Music, and SeedFoley. Because everything is produced in a single pass, the voice, music, and effects stay coherent with each other instead of being mixed from separately generated tracks.

How is it different from normal text-to-speech?

Traditional TTS turns text into a single stream of speech. This is a non-streaming, text-to-any-audio model: it can produce multi-character dialogue, background music, and sound effects together in one generation, with fine-grained control over speed, volume, pitch, emotion, and accent. That means a finished scene, not just a voice track.

Can it clone a voice?

Yes. Seed Audio supports zero-shot voice cloning — provide a short reference clip and the model reproduces that voice, keeping it consistent across up to two minutes of generated audio and even across languages. Preset voices and reference-audio guidance are available when you would rather start from an existing style.

What can I build with it?

Creators use the model for narration, audiobooks, podcasts, game and video sound design, music sketches, and localized voiceovers. Developers reach it through Volcano Engine, the Doubao app, and BytePlus for international access, so it fits both quick creative experiments and production audio pipelines. Teams working on ads, e-learning, trailers, and short-form video use it to move from a rough script to a finished audio draft in minutes, then refine only the parts that need a human touch.

Is this the official ByteDance site?

No. This is an independent platform for exploring and creating with the Seed Audio model. We are not affiliated with, endorsed by, or operated by ByteDance. Seed Audio and related model names belong to their respective owners, and all trademarks remain the property of those owners.

Start creating with Seed Audio

Generate speech, music, and sound effects from a single prompt. Explore the model and bring your ideas to life.