DivMoE introduces a fine-grained MoE upcycling framework that addresses routing collapse when experts come from a single source model, outperforming naive fine-grained approaches on Qwen3-1.7B.
#AIresearch #MixtureOfExperts #LLMs #ModelUpcycling
https://arxiv.org/abs/2610.11317
