Products & Tools

Suno's new Speech feature generates voiceovers alongside music

Suno launches Speech in public beta on web and mobile, generating voiceovers and background music simultaneously from scripts or prompts — its first feature beyond songs.

AI music maker Suno now generates spoken words
AI music maker Suno now generates spoken wordsAI-generated
By Sophie Lindqvist2 min read

Updated

Why it matters

  • Suno launched Speech, a feature generating spoken voices from scripts or prompts, in public beta on web and mobile.
  • Speech generates voiceovers and background music simultaneously in a single output.
  • Suno CPO Jack Brody described Speech as "the first audio model that generates voice and music" and said the company's vision "has always extended to other forms of human expression."

Suno, the company best known for generating full songs from text prompts, has launched a new feature that generates spoken voices from scripts or prompted descriptions. The feature, called Speech, is available now in public beta on Suno's web and mobile platforms, the company announced on its blog.

Speech does something Suno's core product has not done before: it produces voiceovers and background music simultaneously, so a user can generate a narrated, scored audio track in one step rather than stitching together outputs from separate tools.

The move signals a strategic broadening for a company that has built its identity around AI-generated music. "Music will always be at the heart of Suno and what we build. At the same time, our vision has always extended to other forms of human expression," Suno chief product officer Jack Brody said in the announcement. "Today, we're expanding what's possible in Suno with Speech: the first audio model that generates voice and music to …" The quoted statement was truncated in the published excerpt, but the framing is unambiguous: Suno sees itself as an audio company, not strictly a music company.

That distinction matters commercially. Text-to-speech and voiceover generation is a crowded, fast-moving segment, and by combining speech with its existing music generation, Suno is betting that integrated audio output — voice plus score in a single generation — is a differentiator that standalone TTS engines and standalone music generators both lack.

The stakes are also legal. Suno's music generation has placed the company at the center of ongoing disputes over training data and copyright in generative AI, and extending its models deeper into the broader audio market — podcasts, ads, narration, whatever a script describes — expands both the opportunity and the exposure. A combined voice-and-music model aimed at general "human expression," as Brody puts it, positions Suno to compete for audio production work that currently requires multiple vendors.

For users, the practical change is workflow. Where a creator might previously have generated a voiceover in one tool, music in Suno, and then mixed the two, Speech collapses that into a single prompt or script. Suno has not yet detailed pricing tiers, model versions, or availability windows beyond the public beta designation, and the feature's quality relative to dedicated speech models remains untested in independent reviews.

The beta covers both web and mobile, meaning Suno is shipping the capability across its full user base from day one rather than gating it behind a waitlist. With Speech, Suno has made its most explicit move yet toward becoming a general-purpose audio generation platform — and the response of both incumbent speech AI providers and rights holders will shape how far that expansion goes.

Original: suno.com

Share this article:

More from Sophie Lindqvist

Sophie Lindqvist

Show full bio

Staff writer covering marketplaces and e-commerce at AI In Context.

162 articles

Related articles

  1. Google's Flash TTS models build AI voices from text descriptions
  2. Google Launches Gemini 3.1 Flash TTS With Audio Tags and 70+ Languages
  3. Sony and UMG Sue Suno Again Over v6 'Model Laundering'
  4. Google Ships Gemini 3.8 TTS Models With Voice Cloning and Direction
  5. ElevenLabs launches Eleven v4 with tighter voice control

« Previous article