MoE 混合专家模型
查找混合专家架构、稀疏激活、路由和相关开源模型。
先看主题地图,再进入选型
理解 MoE 时要同时看总参数量、激活参数量、专家路由、并行方式和推理显存;页面会把架构、机制和模型放在同一条探索路径里。
本页用于理解主题范围、发现候选和相关概念;准备按部署、预算或能力约束比较时,请进入上方的对应选型指南。
主题精选 KG 入口
这里是用于发现本主题的一组精选入口,不代表付费排名,也不等同于官方产品评测;需要完整语义召回和二层总结,请回到首页检索。
稀疏混合专家架构
Model architecture with sparsely activated expert networks, feature of Seed1.6.
机制Mixture-of-Experts
每个Token激活8个专家(6个路由专家与2个共享专家),约4.3%权重参与计算
机制MoE Routing
每层含1个共享和256个路由专家,每Token激活6个,前三层用哈希路由。
模型Phi-mini-MoE-instruct
来源摘要:Phi-mini-MoE is a lightweight Mixture of Experts (MoE) model with 7.6B total parameters and 2.4B activated parameters. It is compressed and distilled from the base model shared by Phi-3.5-MoE and GRIN-MoE using the SlimMoE approach, then post-trained via supervised fine-tuning and direct preference optimization for instruction following and safety. The model is trained on Phi-3 synthetic data and filtered public documents, with a focus on high-quality, reasoning-dense content. It is part of the SlimMoE series, whic
模型Phi-tiny-MoE-instruct
来源摘要:Phi-tiny-MoE is a lightweight Mixture of Experts (MoE) model with 3.8B total parameters and 1.1B activated parameters. It is compressed and distilled from the base model shared by Phi-3.5-MoE and GRIN-MoE using the SlimMoE approach, then post-trained via supervised fine-tuning and direct preference optimization for instruction following and safety. The model is trained on Phi-3 synthetic data and filtered public documents, with a focus on high-quality, reasoning-dense content. It is part of the SlimMoE series, whic
机制Cursor Router
请求级分类器,通过动态路由任务至不同模型实现成本优化与质量平衡。
模型Qwen1.5-MoE-A2.7B-Chat
来源摘要:Qwen1.5-MoE is a transformer-based MoE decoder-only language model pretrained on a large amount of data. For more details, please refer to our blog post and GitHub repo.
模型deepseek-moe-16b-base
来源摘要:1. Introduction to DeepSeekMoE See the Introduction for more details. 2. How to Use Here give some examples of how to use our model. #### Text Completion
留下这个主题的使用反馈
留言会关联到当前主题页面,保留在 KG 之外,其他访客也可以浏览和点赞。无需登录,不人工审核。