Модели

35 канонических LLM-моделей от всех провайдеров

Показаны модели 1–24 из 35

Sarvam-M

Индия

Sarvam AI's 24B-parameter instruction-tuned model derived from Mistral-Small-3.1-24B, post-trained on English plus eleven major Indic languages (bn, hi, kn, gu, mr, ml, or, pa, ta, te). Delivers large relative gains on Indian-language, math, and programming benchmarks over its base model, with a hybrid reasoning mode for complex tasks.

Контекст
131K
Добавлена
июнь 2026 г.

DiffusionGemma

Соединенные Штаты

Google DeepMind's experimental diffusion-based member of the Gemma 4 open model family. Unlike autoregressive models that generate text one token at a time, DiffusionGemma denoises a canvas of placeholder tokens to produce up to 256 tokens in parallel, finalizing output in one block. A Mixture-of-Experts model with 26B total parameters and 3.8B active per inference, delivering roughly 4x the throughput of similarly sized autoregressive Gemma models on local hardware. Excels at non-linear tasks like in-line editing, molecular sequencing, mathematical graphing, and self-correcting puzzles.

Контекст
262K
Добавлена
июнь 2026 г.

Gemma 4 12B

Соединенные Штаты

Google's medium-size open-weight model with 12 billion parameters from the Gemma 4 family. Encoder-free unified multimodal architecture that natively processes text, image, audio, and video inputs without dedicated encoders. Features a 256K context window and supports 140+ languages. First medium-sized model capable of natively ingesting audio. Suitable for local deployment on GPUs with 16GB VRAM.

Контекст
262K
Добавлена
июнь 2026 г.

Nemotron 3 Ultra

Соединенные Штаты

NVIDIA's flagship open 550B-parameter Mixture-of-Experts model with 55B active parameters, built for frontier reasoning and orchestration in long-running agentic systems. Features hybrid Mamba-Transformer architecture, LatentMoE routing, multi-token prediction, and NVFP4 precision for 5x higher throughput. Achieves 30% lower cost-to-task-completion on agentic benchmarks. Supports 1M+ token context window with 95% accuracy on Ruler@1M.

Контекст
1.0M
Добавлена
июнь 2026 г.

MiniMax M3

Китай

MiniMax's frontier open-weight model with 1M-token context window, native multimodality (text, image, video), and strong coding capabilities. Built on MiniMax Sparse Attention (MSA) architecture, achieving 59% on SWE-Bench Pro with significantly improved efficiency at long context.

Контекст
1.0M
Добавлена
июнь 2026 г.

Ring-2.6-1T

Китай

InclusionAI's (Ant Group) trillion-parameter open-weights reasoning model with 63B active parameters per token. Built for real-world agent workflows with adaptive reasoning-effort modes. Features hybrid linear and MLA attention architecture with MIT license.

Контекст
131K
Добавлена
май 2026 г.

Falcon-H1

ОАЭ

TII's hybrid Mamba-Transformer model that outperforms comparable offerings from Meta's Llama and Alibaba's Qwen in the 30-70B parameter range. Designed for real-world AI on everyday devices and resource-limited settings with state-of-the-art efficiency.

Контекст
131K
Добавлена
май 2026 г.

Solar Pro 3

Республика Корея

Upstage's powerful Mixture-of-Experts language model with 102B total parameters and 12B active parameters per forward pass. Optimized for Korean with strong English and Japanese support. Excels at complex reasoning, structured output generation, and agentic workflows.

Контекст
128K
Добавлена
май 2026 г.

Gemini 3 Flash

Соединенные Штаты

Google's balanced model combining Gemini 3 Pro's reasoning capabilities with the Flash line's latency, efficiency, and cost. Features configurable thinking levels, multimodal function responses, and streaming function calling for complex agentic workflows.

Контекст
1.0M
Добавлена
май 2026 г.

Gemini 3.5 Flash

Соединенные Штаты

Google DeepMind's balanced Gemini 3.5 model that pairs Pro-line reasoning quality with Flash-line latency and cost. Natively multimodal across text, image, audio, and video with a 1M-token context window, configurable thinking levels, and streaming function calling, tuned for high-throughput production workloads.

Контекст
1.0M
Добавлена
май 2026 г.

Laguna M.1

Соединенные Штаты

Poolside AI's flagship agentic coding model with 225B total parameters and 23B active (MoE). Trained from scratch in-house on 30T tokens across 6,144 NVIDIA Hopper GPUs. Optimized for complex multi-step software engineering tasks including codebase exploration, file editing, test running, and iterative debugging.

Контекст
128K
Добавлена
апр. 2026 г.

DeepSeek V4 Pro

Китай

DeepSeek's flagship V4 model with 1.6T total parameters (49B activated). MoE architecture supporting 1M token context. Closes the gap with frontier proprietary models on reasoning and coding benchmarks.

Контекст
1.0M
Добавлена
апр. 2026 г.

Hy3 Preview

Китай

Tencent's flagship open-weight Mixture-of-Experts model from the Hunyuan family with 295B total parameters and 21B active. Integrates fast and slow thinking modes with configurable reasoning effort. Designed for agentic workflows, cross-file code refactoring, long-document analysis, and multi-step tool use.

Контекст
256K
Добавлена
апр. 2026 г.

Qwen 3.6 27B

Китай

Alibaba's dense 27B parameter model that outperforms its own 397B MoE predecessor on agentic coding benchmarks. Strong multilingual and reasoning capabilities released under Apache 2.0.

Контекст
131K
Добавлена
апр. 2026 г.

Qwen 3.6 35B-A3B

Китай

Alibaba's efficient Mixture-of-Experts model with 35B total parameters and 3B active per token. Frontier-level agentic coding performance with 73.4% on SWE-bench Verified and 92.7 on AIME 2026. Released under Apache 2.0.

Контекст
131K
Добавлена
апр. 2026 г.

Gemma 4 31B

Соединенные Штаты

Google's flagship open-weight dense model with 31B parameters. All parameters active per forward pass. Ranks among top open models with strong performance on AIME 2026 (89.2%) and MMLU Pro (85.2%). Supports vision and extended context.

Контекст
262K
Добавлена
апр. 2026 г.

Nemotron 3 Super 120B

Соединенные Штаты

NVIDIA's open hybrid Mamba-Transformer MoE model with 120B total parameters (12B active). Features 1M token context window and excels at agentic reasoning, coding, planning, and tool calling.

Контекст
1.0M
Добавлена
апр. 2026 г.

Kimi K2.6

Китай

Moonshot AI's latest model with ultra-long context window support, strong reasoning capabilities, and excellent performance on complex multi-step tasks. Known for reliable long-document understanding.

Контекст
1.0M
Добавлена
март 2026 г.

GLM-5.1

Китай

Zhipu AI's latest bilingual model with strong Chinese and English capabilities. Features improved reasoning, coding, and tool use with competitive performance on academic benchmarks.

Контекст
131K
Добавлена
март 2026 г.

Mistral Small 4

Франция

Mistral AI's efficient hybrid model unifying instruct, reasoning, and coding in a single model. Open-weight under Apache 2.0 with strong performance for its size class.

Контекст
128K
Добавлена
март 2026 г.

MiniMax M2.7

Китай

MiniMax's latest large language model with strong multilingual and multimodal capabilities. Competitive pricing with high-quality text generation and improved reasoning performance.

Контекст
200K
Добавлена
март 2026 г.

Qwen 3.6

Китай

Alibaba's latest Qwen model with enhanced reasoning, multilingual capabilities, and improved instruction following. Features strong performance on coding, math, and general knowledge benchmarks.

Контекст
131K
Добавлена
март 2026 г.

DeepSeek V4

Китай

DeepSeek's fourth-generation model with improved mixture-of-experts architecture, enhanced reasoning and coding capabilities, and stronger multilingual performance. Competitive with frontier proprietary models.

Контекст
256K
Добавлена
февр. 2026 г.

Mistral Medium 3.5

Франция

Mistral AI's balanced model offering strong multilingual performance with excellent price-performance ratio. Optimized for production workloads requiring reliable quality across European and global languages.

Контекст
128K
Добавлена
февр. 2026 г.