Sber's flagship large-scale Mixture-of-Experts model with 702B total parameters and 36B active. Designed for multilingual assistant workloads, reasoning, code generation, tool use, and large-cluster deployment. Open-weight release.
Yandex's compact 8B parameter language model trained on 15T tokens of primarily Russian and English text. Features 32K context window with strong performance on web, code, and mathematics tasks. Open-weight release.