Knowledge graph · semantic discovery

Text-to-speech models and tools

Find text-to-speech models and tools for Chinese, low latency, local deployment, and open-source use cases.

Search this topic →For AI agentsOpen decision guide →

Public KG candidates · updated 2026-09-09T19:04:28.116763+00:00 · semantic retrieval remains available on the main site

Map the topic before choosing

Start by deciding whether the need is real-time interaction, long-form synthesis, or voice cloning. Then compare Chinese quality, latency, deployment, and licensing boundaries. Search can start from a model you have or from the target use case.

Discovery first
Use this page to understand the topic and discover candidates. When you are ready to compare deployment, budget, or capability constraints, open the decision guide above.

Curated KG entry points

A small set of entry points for discovering this topic, not a paid ranking or an official product review. Return to the homepage for the full semantic retrieval and second-stage summary.

Application

Text-to-speech

KG-recorded application related to Text-to-speech; detailed English notes are not yet available.

Mechanism

TTS

A sequential stage in traditional voice-plus-camera architectures.

Model

speecht5_tts

SpeechT5 model fine-tuned for speech synthesis (text-to-speech) on LibriTTS. This model was introduced in SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing by Junyi Ao, Rui Wang, Long Zhou, Chengyi Wang, Shuo Ren, Yu Wu, Shujie Liu, Tom Ko, Qing Li, Yu Zhang, Zhihua Wei, Yao Qian, Jinyu Li, Furu Wei.

Model

Qwen3-TTS-12Hz-0.6B-Base

**Qwen3-TTS Technical Report** **GitHub Repository** **Hugging Face Demo** Qwen3-TTS is a family of advanced multilingual, controllable, robust, and streaming text-to-speech models. Trained on over 5 million hours of speech data spanning 10 languages, Qwen3-TTS supports state-of-the-art 3-second voice cloning and description-based control.

Model

Qwen3-TTS-12Hz-0.6B-CustomVoice

Qwen3-TTS is a series of advanced multilingual, controllable, robust, and streaming text-to-speech models developed by the Qwen team.

Model

Step-Audio-TTS-3B

Step-Audio-TTS-3B represents the industry's first Text-to-Speech (TTS) model trained on a large-scale synthetic dataset utilizing the LLM-Chat paradigm. It has achieved SOTA Character Error Rate (CER) results on the SEED TTS Eval benchmark. The model supports multiple languages, a variety of emotional expressions, and diverse voice style controls. Notably, Step-Audio-TTS-3B is also the first TTS model in the industry capable of generating RAP and Humming, marking a significant advancement in the field of speech syn

Model

Voxtral-4B-TTS-2603

Voxtral TTS is a frontier, open-weights text-to-speech model that’s fast, instantly adaptable, and produces lifelike speech for voice agents. The model is released with BF16 weights and a set of reference voices. These voices are licensed under CC BY-NC 4, which is the license that the model inherits.

Model

Qwen-Audio-3.0-TTS

TTS model using inline control tags and natural-language style steering.

Share feedback on this topic

Anonymous feedback is attached to this topic page, stays outside the KG, and is visible to other visitors.

Explore related topics