Sber's compact Mixture-of-Experts model with 10B total parameters and 1.8B active. Designed for fast multilingual assistant workloads, reasoning, code, function calling, and product-style deployment on edge devices.
Yandex's compact 8B parameter language model trained on 15T tokens of primarily Russian and English text. Features 32K context window with strong performance on web, code, and mathematics tasks. Open-weight release.