Start For Free

Seed Audio 1.5 AI Audio Generator

Create expressive voices, conversations, music and complete sound scenes.

0/2048
60 Credits/min

Hear what a scene can become

From intimate voices to cinematic worlds. Explore the examples already available in our workspace. Seed Audio 1.0 examples

Sound Scene00:28

Futuristic Crisis Broadcast

Voiceover00:31

Li Mi's Memory

Conversation00:30

Podcast Chat

Conversation00:30

Crime-Thriller Call

6 min

Audio per generation

More room for a story to unfold.

6

Audio references

Bring more voices and sonic ideas.

30

Languages & variants

Tell the story to more audiences.

4

Input modes

Text, audio and video, together.

Beyond text to speech

A whole world of sound. One model.

Seed Audio 1.5 is ByteDance's audio generation model. It brings spoken dialogue, music, ambience and effects into one scene, with more ways to direct the performance and shape the final edit.

Start with words, guide a voice with references, or build sound around footage you already have. Longer generations and separate tracks open up a richer workflow for stories, podcasts and sound design.

Seed Audio 1.5 vs 1.0

More space. More references. More possibilities.

Seed Audio 1.5 vs 1.0
CapabilitySeed Audio 1.5 · Coming soonSeed Audio 1.0
Audio per generationUp to 6 minutesUp to 2 minutes
Audio referencesUp to 6Up to 3
Languages & regional variants30 languages & regional variants20+ languages
Input combinationsText; text + audio; text + video; text + audio + videoText; text + audio; text + image

Seed Audio 1.5 capabilities

Direct the sound, down to the detail.

Write the moment, shape the performance and leave room for the edit.

Complete scenes, together

Generate dialogue, music, ambience and sound effects together, so the scene has a shared mood and direction.

Let the story breathe

Create a scene up to six minutes long. Give a conversation, narration or soundscape time to develop.

Six references to guide you

Bring up to six reference recordings into a generation to guide voices and the sound of your scene.

Stories across languages

Support for 30 languages and regional variants expands the possibilities for dialogue, narration and localization.

A performance with character

Guide reference voices with emotion, style and pace. Direct non-speech vocal sounds such as laughter, sighs and gasps.

Your timing. Your final mix.

Use timestamp instructions for dialogue, effects and music, then work with separate dialogue, ambience, effects and music tracks.

Four ways in

Start with what you have.

A script, a voice reference, a video or all three. Seed Audio 1.5 is designed to understand the material behind your story.

Text

Describe the voices, the setting and the sound you imagine.

Audio scene

Text + audio

Add reference recordings to guide the voice and sonic direction.

Audio scene

Text + video

Build dialogue and sound design around existing footage.

Audio scene

Text + audio + video

Combine voice direction with the visual context of your scene.

Audio scene

Sound that sees the scene

Give your footage a voice. And a world.

Build dubbing and a soundtrack around existing footage. Use the visual scene as context, then direct the dialogue, effects and music with timestamp instructions.

  1. 00:02

    An alarm cuts through the engine hum.

  2. 00:06

    A calm voice: “We still have time.”

  3. 00:12

    The music builds as the ship departs.

A cinematic spacecraft scene used to illustrate video-aware sound design
Illustrative sound-design brief · Seed Audio 1.5

Built for the next edit

One generation. Room to make it yours.

Separate dialogue, ambience, effects and music tracks give you more control after generation. Bring the layers into your editor to adjust the balance, refine a cue or reshape the music around a voice.

00:0000:1500:3000:45

Dialogue

Ambience

Effects

Music

From an idea to a sound scene

Turn your first idea into a complete sound scene, then refine the final mix.

  1. 01

    Set the scene

    Write your script and describe the atmosphere, speakers and sounds.

  2. 02

    Bring your references

    Add reference recordings, video or both to guide the voice and the world around it.

  3. 03

    Direct the performance

    Specify emotion, pace and timestamp cues for dialogue, effects and music.

  4. 04

    Shape the final edit

    Work with the separate tracks in your editor to finish the mix.

Made for your next story

More than a voiceover.

Explore how SeedAudio 1.5 fits into your creative process.

Audio stories & drama

Build character conversations with expressive performances, music and a sense of place.

Podcasts & narration

Explore longer spoken scenes, reference-guided voices and natural pacing.

Dubbing & localization

Use existing footage as context and explore dialogue across languages and regional variants.

Film & game sound design

Sketch layered soundscapes, cue effects to a moment and continue the edit with separate tracks.

FAQ

Seed Audio 1.5 FAQ

Capabilities, creative workflows and access on TopMaker.

Is Seed Audio 1.5 free to try?

TopMaker offers one free trial generation. After that, further generations use credits. Check the workspace for current pricing.

When will SeedAudio 1.5 be available on TopMaker?

The current editor and examples on this page use Seed Audio 1.0. Version 1.5 is not yet integrated into TopMaker, and its launch date has not been confirmed.

How does Seed Audio 1.5 compare with 1.0?

Compared with 1.0, the maximum generation length increases from two to six minutes, and the reference limit rises from three recordings to six. Version 1.5 also supports video inputs and separate tracks, with 30 languages and regional variants across four input combinations.

Can I control a voice’s emotion and speaking style?

Yes. Describe the emotion, speaking style and pace you want in your prompt, and use voice references to guide the performance. You can also request laughter, sighs and gasps. Results can vary, so you may need to refine your instructions to get the delivery you want.

Can I generate dialogue, music and sound effects together?

Yes. You can describe the dialogue, music, ambience and sound effects in a single prompt to create a complete sound scene. Reference recordings can help guide the voices and the setting.

Can I use voice references or video clips?

Yes. You can start with text alone, add reference recordings, include a video clip, or combine all three. Recordings guide the voices and sonic direction, while video provides context for dubbing and sound design.

Can I edit individual tracks after generation?

With version 1.5, you can work with separate dialogue, ambience, effects and music tracks in your editor. These layers are not available in the current 1.0 workspace.

Your next story starts with a sound.

Explore the workspace. Try a scene, bring a voice reference and shape your next audio story.

Create with Seed Audio 1.5

TopMaker AI is an independent service and is not affiliated with or endorsed by ByteDance. TopMaker currently offers Seed Audio 1.0.

Seed Audio 1.5 | AI Audio Generator for Dialogue & Sound