Suno's Speech beta voices a script and scores it in one pass
The beta runs on web, iOS and Android, tops out near eight minutes, and lets creators switch the music off for plain narration.

Key takeaways
- Suno opened a public beta of Speech on October 1, 2026, on web, iOS and Android, after testing the feature for a month with a small group of users.
- Speech generates a spoken voice and its original background music as one audio track; Suno calls it "the first audio model that generates voice and music together as one cohesive track."
- Simple mode writes the words from a one-line description, while Advanced mode performs a custom script with controls for voice gender, speaking style and take-to-take variety.
- The Verge reported a Speech track can run to about eight minutes, and a toggle turns the background music off so the mode works as plain text-to-speech.
- Suno has not published a Speech-specific credit price: generations draw on the same credits as songs, and downloads count against the same monthly caps, 20 on Pro and 60 on Premier.
Suno opened a public beta of Speech on October 1, 2026, a mode that generates a spoken voice and its background music as one audio track rather than two files mixed later. The company calls it "the first audio model that generates voice and music together as one cohesive track," the claim chief product officer Jack Brody makes in the launch post.
Speech sits in Suno's Create screen beside Songs and Sounds, marked Beta. It is live on web, iOS and Android, according to the release note. Suno says it spent the past month testing the feature with a small group before opening it to every account.
How Speech works
There are two modes. Simple takes a one-line description and writes the words as well as performing them; the example The Verge gives is "a pirate captain rallying his crew." Advanced takes a script you supply and adds controls for the voice's gender, the speaking style and how much each generation varies, The Verge's Jess Weatherbed reported. The release note pitches three uses: bedtime stories over soft piano, hype speeches over stadium drums, and ASMR grocery lists.
Music is on by default, and a toggle turns it off, which leaves a plain voice track and makes Speech usable as an ordinary text-to-speech tool. The Verge reported a Speech track can run to about eight minutes, enough for a short story, a guided meditation or a product demo.
What it costs, and what it does not do
Suno has not published a Speech-specific price. Generations draw on the same credits as songs, and downloads count against the same monthly caps. The pricing page lists the Free plan at 50 credits a day with no commercial rights, Pro at $8 a month billed annually with 2,500 credits and 20 song downloads a month, and Premier at $24 a month billed annually with 10,000 credits and 60 downloads a month. Speech sits inside those allowances.
Three gaps matter to anyone planning a series. Suno has not listed which languages Speech supports. The beta cannot use a creator's own cloned voice; Suno's Voices feature works through the song remix menu, not inside Speech, Undetectr reported. The checked release notes describe web and mobile access, not a public Speech API.
Suno is blunt about the rough edges. "Beta really does mean beta," Brody wrote. "Occasionally, British accents can wander off to Australia and back. Dramatic pauses may be very dramatic."
How it compares
Suno enters a field of dedicated speech models. ElevenLabs launched its v4 speech models on September 28, 2026, supporting more than 90 languages and cloning a voice from a 10-second clip, TechCrunch reported. MiniMax's speech-2.8-hd adds emotion control, sound tags and zero-shot voice cloning from a short reference clip. Both are built for the voice alone. Suno's difference is the score written in the same pass, so the music can follow the delivery instead of running under it at a fixed tempo.
Suno's earlier voice work was Bark, an open-source text-to-speech model it released in 2023 before it focused on songs. Speech is the first voice model built into the app.
What this changes for someone who makes video
A narration bed and its soundtrack used to mean two tools and a mixing step: one pass for the voice, another for the score, then a session to balance them. Speech returns a single file with both layers already spaced to each other, which removes the assembly for a first cut, whether that is a channel intro, a documentary voiceover over rising strings or a trailer read over percussion. The limits are the beta's. There is no custom voice yet, no published credit price, and no API to wire into a pipeline. The practical move is to generate a short script now, listen to the full ending, and keep a separate voice-and-music workflow in reserve until the accents settle.
Speech is open to every Suno account at suno.com/create, on iOS and on Android.
Sources
- suno.com - Jack Brody's launch post: the one-track claim, the month of closed testing, the accent and pause warnings, the use-case examples
- suno.com - October 1 release note: web, iOS and Android availability, bedtime-story and hype-speech examples
- suno.com - plan credit allowances and the download caps (20 a month on Pro, 60 on Premier)
- theverge.com - the two modes, the Advanced controls, the eight-minute ceiling, the music toggle, the Create tab
- techcrunch.com - ElevenLabs v4 rival figures: 90-plus languages, 10-second voice cloning, launched September 28, 2026
- minimax.io - MiniMax's speech-2.8-hd model page: emotion control, sound tags, voice cloning
- undetectr.com - Speech credit and download accounting, the own-voice gap, the beta caveats