Qwen3.8-Flash-Next Architecture: A Lightweight Multimodal Large Model Achieving Performance Breakthrough
Alibaba Qwen released multimodal MoE Qwen3.8-Flash and open-sourced next-gen Qwen3.8-Flash-Next (Qwen4 prototype). 125B total params, 51B N-gram embedding, 6B active per token; native 260K context, up to 1M. Training cost 1/9 of previous; API: input ¥1/M tokens, output ¥3/M.....