نماذج الخبراء المتعددين (MoE)
تعرّف إلى معماريات الخبراء المتعددين، والتنشيط المتناثر، وآليات التوجيه، والنماذج المفتوحة المرتبطة بها.
ابدأ بالنطاق، ثم انتقل إلى القرار
عند تقييم MoE، انظر معاً إلى إجمالي المعلمات، والمعلمات النشطة، والتوجيه، والتوازي، وذاكرة الاستدلال. تجمع هذه الصفحة المعمارية والآلية والنموذج في مسار استكشاف واحد.
هذه الصفحة تساعدك على فهم الموضوع واكتشاف المرشحين. عند الحاجة إلى مقارنة القيود والميزانية والنشر، اكتب المشكلة الفعلية في البحث الدلالي.
مداخل الرسم المعرفي
الأسماء والتفاصيل المعروضة هنا تبقى قريبة من السجل العام الأصلي عندما لا تتوفر ترجمة عربية موثقة؛ لا نملأ الفجوات بتخمينات.
MoE
Model architecture with sparsely activated expert networks, feature of Seed1.6.
آليةMixture-of-Experts
Makes local frontier inference arithmetically feasible via sparsity.
آليةMoE Routing
KG-recorded mechanism related to MoE Routing; detailed English notes are not yet available.
نموذجPhi-mini-MoE-instruct
Phi-mini-MoE is a lightweight Mixture of Experts (MoE) model with 7.6B total parameters and 2.4B activated parameters. It is compressed and distilled from the base model shared by Phi-3.5-MoE and GRIN-MoE using the SlimMoE approach, then post-trained via supervised fine-tuning and direct preference optimization for instruction following and safety. The model is trained on Phi-3 synthetic data and filtered public documents, with a focus on high-quality, reasoning-dense content. It is part of the SlimMoE series, whic
نموذجPhi-tiny-MoE-instruct
Phi-tiny-MoE is a lightweight Mixture of Experts (MoE) model with 3.8B total parameters and 1.1B activated parameters. It is compressed and distilled from the base model shared by Phi-3.5-MoE and GRIN-MoE using the SlimMoE approach, then post-trained via supervised fine-tuning and direct preference optimization for instruction following and safety. The model is trained on Phi-3 synthetic data and filtered public documents, with a focus on high-quality, reasoning-dense content. It is part of the SlimMoE series, whic
آليةCursor Router
KG-recorded mechanism related to Cursor Router; detailed English notes are not yet available.
نموذجQwen1.5-MoE-A2.7B-Chat
Qwen1.5-MoE is a transformer-based MoE decoder-only language model pretrained on a large amount of data. For more details, please refer to our blog post and GitHub repo.
نموذجdeepseek-moe-16b-base
1. Introduction to DeepSeekMoE See the Introduction for more details. 2. How to Use Here give some examples of how to use our model. #### Text Completion
نموذجdeepseek-moe-16b-chat
1. Introduction to DeepSeekMoE See the Introduction for more details. 2. How to Use Here give some examples of how to use our model.
نموذجGRIN-MoE
Hugging Face     Tech Report     License     Github     Get Started   With **only 6.6B** activate parameters, GRIN MoE achieves **exceptionally good** performance across a diverse set of tasks, particularly in coding and mathematics tasks.
نموذجInstella-MoE-16B-A3B
16B total params, 2.8B active via 2 shared + 6 of 64 routed experts per token; d
تطبيقMoE模型预训练和后训练
KG-recorded application related to MoE模型预训练和后训练; detailed English notes are not yet available.
أضف تجربة إلى هذا الموضوع
الملاحظات العامة منفصلة عن حقائق الرسم المعرفي، ويمكن للزوار مناقشتها والتصويت عليها.
افتح مساحة النقاش ←