PreMoE: Proactive Inference for Efficient Mixture-of-Experts

ArXi:2505.17639v2 Announce Type: replace Mixture-of-Experts (MoE) models offer dynamic computation, but are typically deployed as static full-capacity models, missing opportunities for deployment-specific specialization. We