Controllable multilingual text to speech

Qwen Audio 3.0 TTS Voice Generator

Create expressive, multilingual speech with Qwen Audio 3.0 TTS. Direct emotion, pace, style, timbre, and accent in plain English, then fine-tune key moments with inline voice tags.

16 supported languagesPlus and Flash models86 inline voice tagsUp to 3 minutes in one pass
Qwen Audio 3.0 TTS explained

What is Qwen Audio 3.0 TTS?

Qwen Audio 3.0 TTS is an AI text-to-speech system for natural, controllable, and multilingual voice generation. Direct the delivery in plain language, refine specific moments with inline tags, and create longer speech from text or authorized reference audio. Choose Plus for quality-first production or Flash for latency-sensitive interaction.

QUALITY FIRST

Qwen Audio 3.0 TTS Plus

For professional narration, branded audio, and production workflows where naturalness and timbre fidelity matter more than speed.

High-quality generation
LATENCY FIRST

Qwen Audio 3.0 TTS Flash

For assistants, rapid previews, and interactive voice products where a faster response matters to the user experience.

Real-time interaction

Production-quality speech

Create clear, natural speech that keeps its voice, pacing, and delivery consistent across longer scripts.

Two levels of voice control

Direct the overall role, emotion, pace, timbre, style, and accent with natural-language instructions, then use 86 inline tags to shape precise moments such as pauses, laughter, or breathing.

Multilingual voice generation

Create multilingual and cross-lingual speech across 16 languages, including workflows that use noisy, reverberant, or unclear reference audio.

Qwen Audio 3.0 TTS capabilities

What are the key features of Qwen Audio 3.0 TTS?

Qwen Audio 3.0 TTS combines natural-language voice control with 86 inline tags, multilingual text to speech, long-form narration, and reference-based voice cloning. Use these features to create consistent voiceovers, dialogue, audiobooks, and localized speech without hand-tuning acoustic settings.

VOICE DIRECTIONWarm, confident, measured pace
RoleEmotionPace

Control AI voice style with natural-language prompts

Describe the speaker’s role, emotion, style, rate, timbre, and accent in plain language instead of tuning acoustic parameters. Qwen Audio 3.0 TTS turns that creative direction into repeatable delivery for voiceovers, dialogue, podcasts, and character speech.

Qwen TTS prompt control · Role · Emotion · Pace · Accent
SCRIPT / EXPRESSIVE CONTROL

[angry] I told you to lock the door.

[laughing] It was waiting for the green button.

[sarcastic, speaking slowly] Clearly a brilliant plan.

Add emotion and non-verbal details with 86 inline tags

Place fine-grained tags directly inside the script to control a phrase or individual word, including expressive transitions, laughter, breathing, coughing, and sighing. This gives game dialogue, film dubbing, and narration precise performance cues without changing the direction for the entire scene.

86 Qwen Audio inline tags · Phrase and word-level control
Qwen
Audio
English中文日本語Españolالعربية

Generate multilingual speech in 16 languages

Create multilingual and cross-lingual speech across English, Spanish, French, German, Portuguese, Japanese, Korean, Arabic, and eight additional languages. Use one controllable voice workflow for localized videos, campaigns, learning content, and global product experiences.

16 languages · Multilingual and cross-lingual TTS
DOCUMENTARY NARRATION02:36

Continue through the full passage with one consistent direction…

Create long-form narration in one pass

Generate up to three minutes of continuous speech in a single pass for audiobooks, lessons, explainers, and documentary voiceovers. Fewer stitched clips help preserve pacing and voice consistency, while vocoder super-resolution supports output up to 48 kHz.

Up to 3 minutes · Vocoder output up to 48 kHz
REFERENCE VOICE
VOICE ID READY

Rights confirmed · Speaker consent required

Clone voices from imperfect reference audio

Create a target voice from authorized reference audio, including clips with background noise, room echo, or limited bandwidth. This makes real-world voice cloning more practical when a studio recording is unavailable; always obtain the speaker’s explicit consent.

Reference-based voice cloning · Real-world audio
Official Qwen Audio 3.0 TTS samples

Qwen Audio 3.0 TTS Samples

Hear official Qwen Audio 3.0 TTS samples and compare them with source audio and CosyVoice models. Explore voice cloning, multilingual speech, expressive control, and long-form narration.

Source voice
English · Zero-shot voice

Since then, the Brazilian has featured in 53 matches for the club in all competitions and has scored 24 goals.

CosyVoice 0.5B
CosyVoice 1.5B
Qwen 3.0 TTS
Source voice
English · Hard text

Oh, no! I left my keys—again?! What am I going to do...call a locksmith?

CosyVoice 0.5B
CosyVoice 1.5B
Qwen 3.0 TTS
Choose the right Qwen TTS model

Qwen Audio 3.0 TTS Plus vs Flash

Compare Qwen Audio 3.0 TTS Plus and Flash by voice quality, first-packet latency, multilingual accuracy, voice cloning support, reference pricing, and production use case—then choose the right model for narration or real-time voice applications.

Plus

qwen-audio-3.0-tts-plus
QUALITY FIRST

Choose Plus for polished narration, audiobooks, advertising, and branded content where naturalness, timbre fidelity, and consistent speaker identity matter more than response speed.

Best fitStudio-grade narration

Choose quality, naturalness, and timbre fidelity over response speed.

First audio
Quality prioritized
Speaker similarity
82.75 average
WER / CER
3.96 average
Voice cloning
Supported
Reference API rate
$0.20 / 10K chars
Instruction control
Supported
Recommended for

Audiobooks · Documentary narration · Advertising · Branded audio

Flash

qwen-audio-3.0-tts-flash
LATENCY FIRST

Choose Flash for voice agents, customer service, interactive characters, and rapid previews where lower first-audio latency and a lower reference character rate matter most.

Best fitRealtime voice workflows

Choose faster first audio for responsive voice experiences.

First audio
300 ms-level
Speaker similarity
80.44 average
WER / CER
3.87 average
Voice cloning
Supported
Reference API rate
$0.15 / 10K chars
Instruction control
Supported
Recommended for

Voice agents · Customer service · Interactive characters · Voice cloning

Data note: Latency and model support follow the current Qwen Cloud release notes and Model Studio documentation. Speaker Similarity and WER/CER are published averages across 16 evaluated languages. Reference API rates are included only for model comparison and are not this site's pricing; availability and rates can change.

Independent comparison

Qwen Audio 3.0 TTS benchmark and ranking

See how Qwen Audio 3.0 TTS Plus compares with leading AI text-to-speech models on the Artificial Analysis Provider Voice Arena, including its current rank, Arena Elo score, voice coverage, and API price.

Global rank#1

Current confidence range: ranks 1–2

Arena Elo1,238Higher is better
95% confidence−16 / +16Reported interval
Arena samples1,479Blind comparisons
Provider voices8 voicesNative voice set
CURRENT LEAD+9 Elo

Qwen leads second-place Simba 3.2 in this snapshot. Their confidence ranges overlap, so the lead should be read as current preference—not a guaranteed permanent gap.

VOICE COVERAGE8 arena voices

The provider-voice track evaluates Qwen across eight native voices rather than presenting the result as a single hand-picked demo voice.

PRICE CONTEXT$27.6 / 1M chars

Artificial Analysis lists Qwen below Sonic 3.5 at $49, but above Simba 3.2 and Gemini 3.1 Flash TTS. It leads on Elo here, not on lowest price.

Method note: Speech Arena rankings are derived from blind user preference votes and expressed as Elo ratings. Source: Artificial Analysis Speech Arena ↗, snapshot captured July 22, 2026. Rankings update as new votes arrive, so the results reflect the tested provider voices, samples, and leaderboard methodology at the capture date.

Qwen AI voice generator use cases

Qwen Audio 3.0 TTS Use Cases

Create YouTube narration, podcast and audiobook audio, e-learning voiceovers, multilingual marketing, game dialogue, and conversational voice agents with controllable AI speech. Use Qwen Audio 3.0 TTS Plus for quality-first production and Flash for latency-sensitive voice applications.

VOICE SESSIONA voice that stays consistent
AI voiceovers for video

AI voiceovers for video

Create expressive narration for YouTube videos, product demos, explainers, training content, and short-form social media. Qwen Audio 3.0 TTS lets creators test different emotions, speaking rates, and vocal roles from the same script, making it easier to refine a voiceover before the final edit without recording every variation from scratch.

Creators · YouTube · Product video
CHAPTER 04A voice that stays consistent
Audiobooks, podcasts, and learning content

Audiobooks, podcasts, and learning content

Produce longer audiobook passages, podcast segments, guided lessons, and documentary narration with consistent vocal direction. One-pass generation of up to three minutes can reduce disruptive cuts between short clips, while natural-language instructions help maintain the intended pace, tone, and speaking style across each section of the story or lesson.

Narration · Education · Storytelling
VOICE SESSIONConfident opening.
Multilingual marketing and localization

Multilingual marketing and localization

Adapt launch videos, ads, product announcements, and branded content for international audiences. Multilingual and cross-lingual synthesis can help teams test how a voice carries across target languages, while emotion and pace controls make it easier to preserve campaign intent instead of producing the same flat delivery for every market.

Marketing · Ecommerce · Localization
LIVE RESPONSEHow should this character respond?
Voice agents and interactive characters

Voice agents and interactive characters

Prototype voices for assistants, customer service agents, games, and interactive characters with explicit control over role, emotion, pace, and speaking style. Choose Qwen Audio 3.0 TTS Flash for latency-sensitive experiences or Plus when polished voice quality matters more, then connect the selected model through a protected production API workflow.

Voice agents · Games · Conversational AI
Qwen Audio 3.0 TTS FAQ

Frequently Asked Questions

What is Qwen Audio 3.0 TTS?

Qwen Audio 3.0 TTS is a production-oriented text-to-speech system for controllable multilingual speech. It supports natural-language voice direction, 86 fine-grained inline tags, long-form generation, and voice cloning workflows for narration, localization, and interactive voice applications.

How does Qwen Audio 3.0 TTS work?

You provide text, select a voice and model, and describe the intended delivery in natural language. The model interprets instructions for role, emotion, style, pace, timbre, and accent. Inline tags can then change a specific phrase or add non-verbal events such as laughter or breathing.

What is the difference between Qwen Audio 3.0 TTS Plus and Flash?

Qwen Audio 3.0 TTS Plus prioritizes naturalness and timbre fidelity for narration, audiobooks, advertising, and other quality-focused speech production. Flash prioritizes faster first-audio response for voice agents, customer service, rapid previews, and interactive experiences. Choose based on whether polished output quality or response speed matters most for your project.

Does Qwen Audio 3.0 TTS support voice cloning?

Yes. Qwen Audio 3.0 TTS can create a cloned voice from reference audio for personalized narration and other voice-generation projects. Use only recordings you are authorized to use, and obtain the speaker’s explicit consent before creating or publishing a cloned voice.

Which languages does Qwen Audio 3.0 TTS support?

Qwen Audio 3.0 TTS supports 16 languages: Arabic, Chinese, English, French, German, Indonesian, Italian, Japanese, Korean, Malay, Portuguese, Russian, Spanish, Tagalog, Thai, and Vietnamese. This multilingual coverage supports global voiceovers, localized marketing, cross-language content, and international voice applications.

Can Qwen Audio 3.0 TTS create long-form narration?

Yes. Qwen Audio 3.0 TTS can generate up to three minutes of speech in a single pass. This is useful for audiobook passages, educational narration, explainers, podcasts, and documentary voiceovers that benefit from continuous pacing and consistent delivery.

Is Qwen Audio 3.0 TTS free to use?

You can start with complimentary credits when you create an account. Those credits let you generate and download watermark-free audio before purchasing more. Continued generation may require paid credits, and you must own or have permission to use every script, reference recording, and source voice.

Can I use Qwen Audio 3.0 TTS commercially?

Yes. Audio generated on this site can be used for advertising, marketing, client work, videos, podcasts, and other commercial projects, subject to our Terms of Service. You must own the script and have permission to use any uploaded recording or cloned voice.

Create controllable AI voices with Qwen Audio 3.0 TTS

Create multilingual speech with control over emotion, pace, style, timbre, and accent. Choose Plus for high-quality narration or Flash for low-latency voice applications.

Explore Qwen Audio 3.0 TTS features