Alibaba's open-weight multimodal preview of the architecture being developed for Qwen4. The sparse mixture-of-experts model has a 125B-parameter backbone with 6B parameters active per token plus 51B parameters of N-gram embeddings. It combines Gated DeltaNet with Qwen Sparse Attention, supports 262,144 tokens natively and up to one million tokens with YaRN, and targets efficient coding, office, visual, and long-context workloads.
Aún no hay datos de proveedores.
Las asociaciones con proveedores aparecerán aquí cuando estén disponibles.