Text-to-speech models and tools
Find text-to-speech models and tools for Chinese, low latency, local deployment, and open-source use cases.
Map the topic before choosing
Start by deciding whether the need is real-time interaction, long-form synthesis, or voice cloning. Then compare Chinese quality, latency, deployment, and licensing boundaries. Search can start from a model you have or from the target use case.
Use this page to understand the topic and discover candidates. When you are ready to compare deployment, budget, or capability constraints, open the decision guide above.
Curated KG entry points
A small set of entry points for discovering this topic, not a paid ranking or an official product review. Return to the homepage for the full semantic retrieval and second-stage summary.
Text-to-speech
KG-recorded application related to Text-to-speech; detailed English notes are not yet available.
MechanismTTS
A sequential stage in traditional voice-plus-camera architectures.
Modelspeecht5_tts
SpeechT5 model fine-tuned for speech synthesis (text-to-speech) on LibriTTS. This model was introduced in SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing by Junyi Ao, Rui Wang, Long Zhou, Chengyi Wang, Shuo Ren, Yu Wu, Shujie Liu, Tom Ko, Qing Li, Yu Zhang, Zhihua Wei, Yao Qian, Jinyu Li, Furu Wei.
ModelQwen3-TTS-12Hz-0.6B-Base
**Qwen3-TTS Technical Report** **GitHub Repository** **Hugging Face Demo** Qwen3-TTS is a family of advanced multilingual, controllable, robust, and streaming text-to-speech models. Trained on over 5 million hours of speech data spanning 10 languages, Qwen3-TTS supports state-of-the-art 3-second voice cloning and description-based control.
ModelQwen3-TTS-12Hz-0.6B-CustomVoice
Qwen3-TTS is a series of advanced multilingual, controllable, robust, and streaming text-to-speech models developed by the Qwen team.
ModelStep-Audio-TTS-3B
Step-Audio-TTS-3B represents the industry's first Text-to-Speech (TTS) model trained on a large-scale synthetic dataset utilizing the LLM-Chat paradigm. It has achieved SOTA Character Error Rate (CER) results on the SEED TTS Eval benchmark. The model supports multiple languages, a variety of emotional expressions, and diverse voice style controls. Notably, Step-Audio-TTS-3B is also the first TTS model in the industry capable of generating RAP and Humming, marking a significant advancement in the field of speech syn
ModelVoxtral-4B-TTS-2603
Voxtral TTS is a frontier, open-weights text-to-speech model that’s fast, instantly adaptable, and produces lifelike speech for voice agents. The model is released with BF16 weights and a set of reference voices. These voices are licensed under CC BY-NC 4, which is the license that the model inherits.
ModelQwen-Audio-3.0-TTS
TTS model using inline control tags and natural-language style steering.
Share feedback on this topic
Anonymous feedback is attached to this topic page, stays outside the KG, and is visible to other visitors.