Knowledge graph · semantic discovery

TTS 语音合成模型与工具

查找中文、低延迟、本地部署和开源取向的文字转语音模型与工具。

检索这个主题 →面向 AI 代理进入对应选型指南 →

公开 KG 候选 · 数据更新于 2026-09-09T19:04:28.116763+00:00 · 完整语义检索请回到首页使用

先看主题地图,再进入选型

先确定是实时交互、长文本合成还是声音克隆,再比较中文自然度、延迟、部署方式和许可边界。搜索时可以直接描述已有模型,也可以描述目标场景。

先发现,再决策
本页用于理解主题范围、发现候选和相关概念;准备按部署、预算或能力约束比较时,请进入上方的对应选型指南。

主题精选 KG 入口

这里是用于发现本主题的一组精选入口,不代表付费排名,也不等同于官方产品评测;需要完整语义召回和二层总结,请回到首页检索。

应用

Text-to-speech

公开记录已收录该实体,但还没有足够的简要说明。

机制

TTS

A sequential stage in traditional voice-plus-camera architectures.

模型

speecht5_tts

来源摘要:SpeechT5 model fine-tuned for speech synthesis (text-to-speech) on LibriTTS. This model was introduced in SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing by Junyi Ao, Rui Wang, Long Zhou, Chengyi Wang, Shuo Ren, Yu Wu, Shujie Liu, Tom Ko, Qing Li, Yu Zhang, Zhihua Wei, Yao Qian, Jinyu Li, Furu Wei.

模型

Qwen3-TTS-12Hz-0.6B-Base

来源摘要:**Qwen3-TTS Technical Report** **GitHub Repository** **Hugging Face Demo** Qwen3-TTS is a family of advanced multilingual, controllable, robust, and streaming text-to-speech models. Trained on over 5 million hours of speech data spanning 10 languages, Qwen3-TTS supports state-of-the-art 3-second voice cloning and description-based control.

模型

Qwen3-TTS-12Hz-0.6B-CustomVoice

来源摘要:Qwen3-TTS is a series of advanced multilingual, controllable, robust, and streaming text-to-speech models developed by the Qwen team.

模型

Step-Audio-TTS-3B

来源摘要:Step-Audio-TTS-3B represents the industry's first Text-to-Speech (TTS) model trained on a large-scale synthetic dataset utilizing the LLM-Chat paradigm. It has achieved SOTA Character Error Rate (CER) results on the SEED TTS Eval benchmark. The model supports multiple languages, a variety of emotional expressions, and diverse voice style controls. Notably, Step-Audio-TTS-3B is also the first TTS model in the industry capable of generating RAP and Humming, marking a significant advancement in the field of speech syn

模型

Voxtral-4B-TTS-2603

来源摘要:Voxtral TTS is a frontier, open-weights text-to-speech model that’s fast, instantly adaptable, and produces lifelike speech for voice agents. The model is released with BF16 weights and a set of reference voices. These voices are licensed under CC BY-NC 4, which is the license that the model inherits.

模型

Qwen-Audio-3.0-TTS

支持16种语言及20种方言,具备实时交互优化与高质量生成双模式。

留下这个主题的使用反馈

留言会关联到当前主题页面,保留在 KG 之外,其他访客也可以浏览和点赞。无需登录,不人工审核。

继续探索相关主题