خريطة موضوعية · اكتشاف دلالي

نماذج الخبراء المتعددين (MoE)

تعرّف إلى معماريات الخبراء المتعددين، والتنشيط المتناثر، وآليات التوجيه، والنماذج المفتوحة المرتبطة بها.

ابحث في هذا الموضوع ←شارك ملاحظة

مرشحون من الرسم المعرفي العام · آخر تحديث: 2026-09-10T11:07:18.277315+00:00

ابدأ بالنطاق، ثم انتقل إلى القرار

عند تقييم MoE، انظر معاً إلى إجمالي المعلمات، والمعلمات النشطة، والتوجيه، والتوازي، وذاكرة الاستدلال. تجمع هذه الصفحة المعمارية والآلية والنموذج في مسار استكشاف واحد.

اكتشاف أولاً، لا ترتيب مدفوعاً.
هذه الصفحة تساعدك على فهم الموضوع واكتشاف المرشحين. عند الحاجة إلى مقارنة القيود والميزانية والنشر، اكتب المشكلة الفعلية في البحث الدلالي.

مداخل الرسم المعرفي

الأسماء والتفاصيل المعروضة هنا تبقى قريبة من السجل العام الأصلي عندما لا تتوفر ترجمة عربية موثقة؛ لا نملأ الفجوات بتخمينات.

آلية

MoE

Model architecture with sparsely activated expert networks, feature of Seed1.6.

آلية

Mixture-of-Experts

Makes local frontier inference arithmetically feasible via sparsity.

آلية

MoE Routing

KG-recorded mechanism related to MoE Routing; detailed English notes are not yet available.

نموذج

Phi-mini-MoE-instruct

Phi-mini-MoE is a lightweight Mixture of Experts (MoE) model with 7.6B total parameters and 2.4B activated parameters. It is compressed and distilled from the base model shared by Phi-3.5-MoE and GRIN-MoE using the SlimMoE approach, then post-trained via supervised fine-tuning and direct preference optimization for instruction following and safety. The model is trained on Phi-3 synthetic data and filtered public documents, with a focus on high-quality, reasoning-dense content. It is part of the SlimMoE series, whic

نموذج

Phi-tiny-MoE-instruct

Phi-tiny-MoE is a lightweight Mixture of Experts (MoE) model with 3.8B total parameters and 1.1B activated parameters. It is compressed and distilled from the base model shared by Phi-3.5-MoE and GRIN-MoE using the SlimMoE approach, then post-trained via supervised fine-tuning and direct preference optimization for instruction following and safety. The model is trained on Phi-3 synthetic data and filtered public documents, with a focus on high-quality, reasoning-dense content. It is part of the SlimMoE series, whic

آلية

Cursor Router

KG-recorded mechanism related to Cursor Router; detailed English notes are not yet available.

نموذج

Qwen1.5-MoE-A2.7B-Chat

Qwen1.5-MoE is a transformer-based MoE decoder-only language model pretrained on a large amount of data. For more details, please refer to our blog post and GitHub repo.

نموذج

deepseek-moe-16b-base

1. Introduction to DeepSeekMoE See the Introduction for more details. 2. How to Use Here give some examples of how to use our model. #### Text Completion

نموذج

deepseek-moe-16b-chat

1. Introduction to DeepSeekMoE See the Introduction for more details. 2. How to Use Here give some examples of how to use our model.

نموذج

GRIN-MoE

Hugging Face &nbsp &nbsp Tech Report &nbsp &nbsp License &nbsp &nbsp Github &nbsp &nbsp Get Started &nbsp With **only 6.6B** activate parameters, GRIN MoE achieves **exceptionally good** performance across a diverse set of tasks, particularly in coding and mathematics tasks.

نموذج

Instella-MoE-16B-A3B

16B total params, 2.8B active via 2 shared + 6 of 64 routed experts per token; d

تطبيق

MoE模型预训练和后训练

KG-recorded application related to MoE模型预训练和后训练; detailed English notes are not yet available.

أضف تجربة إلى هذا الموضوع

الملاحظات العامة منفصلة عن حقائق الرسم المعرفي، ويمكن للزوار مناقشتها والتصويت عليها.

افتح مساحة النقاش ←

استكشف موضوعات قريبة