خريطة موضوعية · اكتشاف دلالي

نماذج وأدوات تحويل النص إلى كلام

اكتشف نماذج وأدوات تحويل النص إلى كلام مع التركيز على العربية، وزمن الاستجابة المنخفض، والنشر المحلي، والبرمجيات المفتوحة.

ابحث في هذا الموضوع ←شارك ملاحظة

مرشحون من الرسم المعرفي العام · آخر تحديث: 2026-09-10T11:07:18.277315+00:00

ابدأ بالنطاق، ثم انتقل إلى القرار

حدّد أولاً ما إذا كان الاستخدام تفاعلياً وفورياً، أو لتوليد نصوص طويلة، أو لاستنساخ الصوت. ثم قارن جودة العربية، وزمن الاستجابة، وطريقة النشر، وحدود الترخيص.

اكتشاف أولاً، لا ترتيب مدفوعاً.
هذه الصفحة تساعدك على فهم الموضوع واكتشاف المرشحين. عند الحاجة إلى مقارنة القيود والميزانية والنشر، اكتب المشكلة الفعلية في البحث الدلالي.

مداخل الرسم المعرفي

الأسماء والتفاصيل المعروضة هنا تبقى قريبة من السجل العام الأصلي عندما لا تتوفر ترجمة عربية موثقة؛ لا نملأ الفجوات بتخمينات.

تطبيق

Text-to-speech

KG-recorded application related to Text-to-speech; detailed English notes are not yet available.

آلية

TTS

A sequential stage in traditional voice-plus-camera architectures.

نموذج

speecht5_tts

SpeechT5 model fine-tuned for speech synthesis (text-to-speech) on LibriTTS. This model was introduced in SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing by Junyi Ao, Rui Wang, Long Zhou, Chengyi Wang, Shuo Ren, Yu Wu, Shujie Liu, Tom Ko, Qing Li, Yu Zhang, Zhihua Wei, Yao Qian, Jinyu Li, Furu Wei.

نموذج

Qwen3-TTS-12Hz-0.6B-Base

**Qwen3-TTS Technical Report** **GitHub Repository** **Hugging Face Demo** Qwen3-TTS is a family of advanced multilingual, controllable, robust, and streaming text-to-speech models. Trained on over 5 million hours of speech data spanning 10 languages, Qwen3-TTS supports state-of-the-art 3-second voice cloning and description-based control.

نموذج

Qwen3-TTS-12Hz-0.6B-CustomVoice

Qwen3-TTS is a series of advanced multilingual, controllable, robust, and streaming text-to-speech models developed by the Qwen team.

نموذج

Step-Audio-TTS-3B

Step-Audio-TTS-3B represents the industry's first Text-to-Speech (TTS) model trained on a large-scale synthetic dataset utilizing the LLM-Chat paradigm. It has achieved SOTA Character Error Rate (CER) results on the SEED TTS Eval benchmark. The model supports multiple languages, a variety of emotional expressions, and diverse voice style controls. Notably, Step-Audio-TTS-3B is also the first TTS model in the industry capable of generating RAP and Humming, marking a significant advancement in the field of speech syn

نموذج

Voxtral-4B-TTS-2603

Voxtral TTS is a frontier, open-weights text-to-speech model that’s fast, instantly adaptable, and produces lifelike speech for voice agents. The model is released with BF16 weights and a set of reference voices. These voices are licensed under CC BY-NC 4, which is the license that the model inherits.

نموذج

Qwen-Audio-3.0-TTS

TTS model using inline control tags and natural-language style steering.

نموذج

GLM-TTS

GLM-TTS: Controllable & Emotion-Expressive Zero-shot TTS 📜 Paper       💻 GitHub Repository       🛠️ Audio.Z.AI

نموذج

magpie_tts_multilingual_357m

Model architecture Model size Language-lightgrey#model-badge) 🤗 **Hugging Face MagpieTTS Multilingual demo**: magpie_tts_multilingual_demo

نموذج

Qwen3-TTS-12Hz-1.7B-Base

Qwen3-TTS covers 10 major languages (Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian) as well as multiple dialectal voice profiles to meet global application needs. In addition, the models feature strong contextual understanding, enabling adaptive control of tone, speaking rate, and emotional expression based on instructions and text semantics, and they show markedly improved robustness to noisy input text. Key features

نموذج

Qwen3-TTS-12Hz-1.7B-CustomVoice

Qwen3-TTS covers 10 major languages (Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian) as well as multiple dialectal voice profiles to meet global application needs. In addition, the models feature strong contextual understanding, enabling adaptive control of tone, speaking rate, and emotional expression based on instructions and text semantics, and they show markedly improved robustness to noisy input text. Key features

أضف تجربة إلى هذا الموضوع

الملاحظات العامة منفصلة عن حقائق الرسم المعرفي، ويمكن للزوار مناقشتها والتصويت عليها.

افتح مساحة النقاش ←

استكشف موضوعات قريبة