HaiAI123

Curated Global AI Tools Directory

NewHot story
97
heat index
New

Suno launches Speech for end-to-end voice-and-music tracks

10/03/2026 — 10/03, 14:11·1 sources·1 reports

Story overview

On October 1, AI music platform Suno announced Speech, which it describes as the industry's first audio model capable of generating speech and music together as a single, continuous track. According to a report published on October 3, the model departs from the traditional workflow in which text-to-speech output is produced first and BGM is stitched in afterwards. Instead, Speech generates everything end to end, and Suno calls it the first closed-source end-to-end audio model able to produce a fully blended track of spoken narration and original BGM in a single pass.

Over the month before the announcement, Suno tested the feature with a small group of users. It is now open to everyone as a public beta, and it ships directly inside Suno. To use it, a person supplies a prompt — an idea, a poem, or a passage of text — and then describes the voice and the musical style they have in mind. Suno says this makes it straightforward to create spoken audio with backing music.

Suno has also acknowledged the rough edges of the current build. The company says the beta occasionally drifts in accent, offering the example of a British accent sliding into Australian partway through and then drifting back. Intonation and emotional pauses, it notes, can also come across as overly dramatic.

Those claims come from Suno itself, as relayed in the reporting, and the reports do not go into further technical detail about the model or say what comes after the beta period. The release is positioned around a single idea: rather than assembling speech and music from separate steps, Suno is offering one model that produces the combined track in one go, and letting anyone try it while the company continues to work on accuracy and delivery.

AI-generated from 1 reports · updated 17 minutes ago

Latest turnSuno has released Speech, which it calls the first audio model that generates speech and music as a single continuous track end to end, rather than running TTS and then layering background music. The feature is built into Suno and now open to all users in beta; Suno notes occasional accent drift and overly dramatic delivery in the current version.

Related tools

Reports on this story headlines open the original

Today
  1. Suno has released Speech, which it calls the first audio model that generates speech and music as a single continuous track end to end, rather than running TTS and then layering background music. The feature is built into Suno and now open to all users in beta; Suno notes occasional accent drift and overly dramatic delivery in the current version.

    IT之家AI score 85

Other stories people are talking about

How is heat calculated?About the method

Heat counts how many independent sources covered a story in the last 48 hours: one source counts once no matter how many posts it published, decaying with a 24-hour half-life. What ranks first is what many people are talking about.

This page aggregates public feeds. Headlines and summaries are machine-organized and remain the property of the original authors; verify important facts at the source.

Surge
Discussion rising fast
New
First report within 6 hours
Rising
Still gathering discussion

Back to the hot board →