Qwen Audio 3.0 TTS Plus
For professional narration, branded audio, and production workflows where naturalness and timbre fidelity matter more than speed.
High-quality generationCreate expressive, multilingual speech with Qwen Audio 3.0 TTS. Direct emotion, pace, style, timbre, and accent in plain English, then fine-tune key moments with inline voice tags.
Qwen Audio 3.0 TTS is an AI text-to-speech system for natural, controllable, and multilingual voice generation. Direct the delivery in plain language, refine specific moments with inline tags, and create longer speech from text or authorized reference audio. Choose Plus for quality-first production or Flash for latency-sensitive interaction.
For professional narration, branded audio, and production workflows where naturalness and timbre fidelity matter more than speed.
High-quality generationFor assistants, rapid previews, and interactive voice products where a faster response matters to the user experience.
Real-time interactionCreate clear, natural speech that keeps its voice, pacing, and delivery consistent across longer scripts.
Direct the overall role, emotion, pace, timbre, style, and accent with natural-language instructions, then use 86 inline tags to shape precise moments such as pauses, laughter, or breathing.
Create multilingual and cross-lingual speech across 16 languages, including workflows that use noisy, reverberant, or unclear reference audio.
Qwen Audio 3.0 TTS combines natural-language voice control with 86 inline tags, multilingual text to speech, long-form narration, and reference-based voice cloning. Use these features to create consistent voiceovers, dialogue, audiobooks, and localized speech without hand-tuning acoustic settings.
Describe the speaker’s role, emotion, style, rate, timbre, and accent in plain language instead of tuning acoustic parameters. Qwen Audio 3.0 TTS turns that creative direction into repeatable delivery for voiceovers, dialogue, podcasts, and character speech.
Qwen TTS prompt control · Role · Emotion · Pace · AccentPlace fine-grained tags directly inside the script to control a phrase or individual word, including expressive transitions, laughter, breathing, coughing, and sighing. This gives game dialogue, film dubbing, and narration precise performance cues without changing the direction for the entire scene.
86 Qwen Audio inline tags · Phrase and word-level controlCreate multilingual and cross-lingual speech across English, Spanish, French, German, Portuguese, Japanese, Korean, Arabic, and eight additional languages. Use one controllable voice workflow for localized videos, campaigns, learning content, and global product experiences.
16 languages · Multilingual and cross-lingual TTSContinue through the full passage with one consistent direction…
Generate up to three minutes of continuous speech in a single pass for audiobooks, lessons, explainers, and documentary voiceovers. Fewer stitched clips help preserve pacing and voice consistency, while vocoder super-resolution supports output up to 48 kHz.
Up to 3 minutes · Vocoder output up to 48 kHzRights confirmed · Speaker consent required
Create a target voice from authorized reference audio, including clips with background noise, room echo, or limited bandwidth. This makes real-world voice cloning more practical when a studio recording is unavailable; always obtain the speaker’s explicit consent.
Reference-based voice cloning · Real-world audioHear official Qwen Audio 3.0 TTS samples and compare them with source audio and CosyVoice models. Explore voice cloning, multilingual speech, expressive control, and long-form narration.
Since then, the Brazilian has featured in 53 matches for the club in all competitions and has scored 24 goals.
Oh, no! I left my keys—again?! What am I going to do...call a locksmith?
Compare Qwen Audio 3.0 TTS Plus and Flash by voice quality, first-packet latency, multilingual accuracy, voice cloning support, reference pricing, and production use case—then choose the right model for narration or real-time voice applications.
Choose Plus for polished narration, audiobooks, advertising, and branded content where naturalness, timbre fidelity, and consistent speaker identity matter more than response speed.
Choose quality, naturalness, and timbre fidelity over response speed.
Audiobooks · Documentary narration · Advertising · Branded audio
Choose Flash for voice agents, customer service, interactive characters, and rapid previews where lower first-audio latency and a lower reference character rate matter most.
Choose faster first audio for responsive voice experiences.
Voice agents · Customer service · Interactive characters · Voice cloning
Data note: Latency and model support follow the current Qwen Cloud release notes and Model Studio documentation. Speaker Similarity and WER/CER are published averages across 16 evaluated languages. Reference API rates are included only for model comparison and are not this site's pricing; availability and rates can change.
See how Qwen Audio 3.0 TTS Plus compares with leading AI text-to-speech models on the Artificial Analysis Provider Voice Arena, including its current rank, Arena Elo score, voice coverage, and API price.
Current confidence range: ranks 1–2
Current Arena Elo comparison
Qwen leads second-place Simba 3.2 in this snapshot. Their confidence ranges overlap, so the lead should be read as current preference—not a guaranteed permanent gap.
The provider-voice track evaluates Qwen across eight native voices rather than presenting the result as a single hand-picked demo voice.
Artificial Analysis lists Qwen below Sonic 3.5 at $49, but above Simba 3.2 and Gemini 3.1 Flash TTS. It leads on Elo here, not on lowest price.
Method note: Speech Arena rankings are derived from blind user preference votes and expressed as Elo ratings. Source: Artificial Analysis Speech Arena ↗, snapshot captured July 22, 2026. Rankings update as new votes arrive, so the results reflect the tested provider voices, samples, and leaderboard methodology at the capture date.
Create YouTube narration, podcast and audiobook audio, e-learning voiceovers, multilingual marketing, game dialogue, and conversational voice agents with controllable AI speech. Use Qwen Audio 3.0 TTS Plus for quality-first production and Flash for latency-sensitive voice applications.
Create expressive narration for YouTube videos, product demos, explainers, training content, and short-form social media. Qwen Audio 3.0 TTS lets creators test different emotions, speaking rates, and vocal roles from the same script, making it easier to refine a voiceover before the final edit without recording every variation from scratch.
Creators · YouTube · Product videoProduce longer audiobook passages, podcast segments, guided lessons, and documentary narration with consistent vocal direction. One-pass generation of up to three minutes can reduce disruptive cuts between short clips, while natural-language instructions help maintain the intended pace, tone, and speaking style across each section of the story or lesson.
Narration · Education · StorytellingAdapt launch videos, ads, product announcements, and branded content for international audiences. Multilingual and cross-lingual synthesis can help teams test how a voice carries across target languages, while emotion and pace controls make it easier to preserve campaign intent instead of producing the same flat delivery for every market.
Marketing · Ecommerce · LocalizationPrototype voices for assistants, customer service agents, games, and interactive characters with explicit control over role, emotion, pace, and speaking style. Choose Qwen Audio 3.0 TTS Flash for latency-sensitive experiences or Plus when polished voice quality matters more, then connect the selected model through a protected production API workflow.
Voice agents · Games · Conversational AIQwen Audio 3.0 TTS is a production-oriented text-to-speech system for controllable multilingual speech. It supports natural-language voice direction, 86 fine-grained inline tags, long-form generation, and voice cloning workflows for narration, localization, and interactive voice applications.
You provide text, select a voice and model, and describe the intended delivery in natural language. The model interprets instructions for role, emotion, style, pace, timbre, and accent. Inline tags can then change a specific phrase or add non-verbal events such as laughter or breathing.
Qwen Audio 3.0 TTS Plus prioritizes naturalness and timbre fidelity for narration, audiobooks, advertising, and other quality-focused speech production. Flash prioritizes faster first-audio response for voice agents, customer service, rapid previews, and interactive experiences. Choose based on whether polished output quality or response speed matters most for your project.
Yes. Qwen Audio 3.0 TTS can create a cloned voice from reference audio for personalized narration and other voice-generation projects. Use only recordings you are authorized to use, and obtain the speaker’s explicit consent before creating or publishing a cloned voice.
Qwen Audio 3.0 TTS supports 16 languages: Arabic, Chinese, English, French, German, Indonesian, Italian, Japanese, Korean, Malay, Portuguese, Russian, Spanish, Tagalog, Thai, and Vietnamese. This multilingual coverage supports global voiceovers, localized marketing, cross-language content, and international voice applications.
Yes. Qwen Audio 3.0 TTS can generate up to three minutes of speech in a single pass. This is useful for audiobook passages, educational narration, explainers, podcasts, and documentary voiceovers that benefit from continuous pacing and consistent delivery.
You can start with complimentary credits when you create an account. Those credits let you generate and download watermark-free audio before purchasing more. Continued generation may require paid credits, and you must own or have permission to use every script, reference recording, and source voice.
Yes. Audio generated on this site can be used for advertising, marketing, client work, videos, podcasts, and other commercial projects, subject to our Terms of Service. You must own the script and have permission to use any uploaded recording or cloned voice.
Create multilingual speech with control over emotion, pace, style, timbre, and accent. Choose Plus for high-quality narration or Flash for low-latency voice applications.
Explore Qwen Audio 3.0 TTS features