Midterms 2026See who we think should earn your vote, based on our standardsThe guide →
WRITTEN IN PLAIN AMERICAN ENGLISH.
CLAY TRIBUNE.
Advertisement

Suno’s AI music tool can now speak as well as compose

Suno, the AI music maker, now speaks! A beta feature pairs voiceover with music, glitchy accents included.

By mitch·4 min read
A digital workstation glows with a ghostly voice and swirling musical notes, symbolizing AI voice synthesis.

Suno, which builds AI music, has entered voice generation. It has released a public beta that turns scripts or prompted descriptions into spoken audio, paired with AI background music. The feature is available on web and mobile.

Suno takes a prompt or a script and produces a recording with voiceover and a track underneath it. The company says it is the first audio model that builds voice and music into a single cohesive track at once. This effort pushes past its core business, though it brings along some legal baggage in the process.

Jack Brody’s announcement

Suno chief product officer Jack Brody announced the feature, framing it as a natural extension of the company’s broader ambitions. “Music will always be at the heart of Suno and what we build,” he said. “At the same time, our vision has always extended to other forms of human expression.”

Advertisement

He also added: “Today, we’re expanding what’s possible in Suno with Speech: the first audio model that generates voice and music together as one cohesive track.”

Brody also addressed the beta status directly. “Beta really does mean beta,” he said. “Occasionally, British accents can wander off to Australia and back. Dramatic pauses may be very dramatic. You will almost certainly discover uses for this that never occurred to us.”

The crowded field

Deep learning speech synthesis has been around for a while. DeepMind has been testing it for ten years now. Adobe has its own text-to-speech tool. ElevenLabs has grown into one of the most well-known platforms for it since it launched in 2023, and the field shows no sign of slowing down.

Suno is simply joining the party. The company’s music generator has attracted numerous lawsuits, and diversifying the platform makes sense as a strategic move. Pairing AI music with generated voices is Suno’s spin on text-to-speech tools — the kind of gimmick that could either win attention or simply confirm that the market is saturated.

The option to switch off the background music exists. A toggle lets you remove the soundtrack entirely, so the speech plays without accompaniment. The purpose behind it is to suit particular uses for generated spoken content — a soothing track for poems, or a livelier sound for dramatic voiceovers and inspiring addresses.

How to use it

The process is simple. Choose the “Create” tab, then go to the Speech option. Two modes are available:

  • Simple, which lets you describe what you want to create via the provided prompt box (such as “a pirate captain rallying his crew”)
  • Advanced, which lets you add a custom script if you already know exactly what you want it to say

The advanced settings let you change the voice’s gender, its speaking manner, and the amount of variation in each voice creation. A single piece of speech can last up to around eight minutes.

Suno acknowledges that the feature falls short of perfection, yet insists it will continue refining Speech based on what users tell it.

What the beta means

What matters here is the beta label. Suno is not claiming a finished product; it is describing a work-in-progress. That means occasional British accents wandering to Australia and back are part of what the company is saying comes with the package, along with dramatic pauses that may be very dramatic.

“Beta really does mean beta,” said Brody. “Occasionally, British accents can wander off to Australia and back. Dramatic pauses may be very dramatic. You will almost certainly discover uses for this that never occurred to us.”

It is refreshing to be honest about it. It also serves as a reminder that the feature remains under development. An eight-minute maximum duration gives plenty of room, though the oddities mentioned are part of what users get rather than a tally of shortcomings.

If you simply want the voice by itself, the toggle exists for that purpose. It is a modest relief. Beyond it, though, waits everything else to be found.

Source material: “AI music maker Suno now generates spoken words,” The Verge.

The Notebook

Get the Notebook.

The day's best stories and every fresh verdict, in plain English, in your inbox by seven. One email a day, no more.

We send one note to confirm. Every issue has a one-click way out.

Advertisement

Leave a Reply

Your email address will not be published. Required fields are marked *

As an Amazon Associate, Clay Tribune earns from qualifying purchases.