Модели
19 канонических LLM-моделей от всех провайдеров
Некоторые описания переведены автоматически в рамках пилота и пока не проверены редактором.
Inkling
Thinking Machines Lab's open-weights general-purpose multimodal Mixture-of-Experts model with 975B total parameters and 41B active parameters. Inkling accepts text, image, and audio inputs, produces text, and is designed for agentic and tool-use systems, coding assistants, chatbots, and retrieval-augmented generation.
GPT-5.6 Sol
Флагманская модель OpenAI серии GPT-5.6, развивающая программирование, научное рассуждение, долгосрочное планирование и агентные рабочие процессы и одновременно повышающая надёжность и эффективность на сложных практических задачах. Добавляет максимальный уровень глубины рассуждения и режим ultra, который запускает субагентов для сложной многоэтапной работы.
GPT-5.6 Terra
Сбалансированная модель серии GPT-5.6 от OpenAI: немного уступает по максимальному качеству, но обеспечивает заметно меньшую задержку и стоимость. Сохраняет сильные возможности рассуждения, программирования и агентного использования инструментов с настраиваемой глубиной рассуждения, поэтому подходит как основной вариант для масштабных производственных нагрузок, которым нужны передовые возможности.
Claude Sonnet 5
Самая мощная модель Anthropic класса Sonnet, переносящая передовые возможности программирования, агентной и профессиональной работы в средний сегмент и сокращающая разрыв с Opus 4.8 при меньшей цене. Поддерживает адаптивное мышление с выбором глубины рассуждения, контекстное окно 1 млн токенов и входные данные в виде текста, изображений и файлов. Кодовое имя — Fennec.
Nemotron 3 Ultra
NVIDIA's flagship open 550B-parameter Mixture-of-Experts model with 55B active parameters, built for frontier reasoning and orchestration in long-running agentic systems. Features hybrid Mamba-Transformer architecture, LatentMoE routing, multi-token prediction, and NVFP4 precision for 5x higher throughput. Achieves 30% lower cost-to-task-completion on agentic benchmarks. Supports 1M+ token context window with 95% accuracy on Ruler@1M.
Palmyra X5
Writer's most advanced adaptive reasoning model with a 1 million token context window. Processes full million-token prompts in approximately 22 seconds with multi-turn function calls in 300ms. Optimized for enterprise agentic AI workflows at 3-4x lower cost than GPT-4.1.
Gemini 3.1 Flash-Lite
Google's most cost-efficient Gemini model optimized for high-volume, low-latency use cases. Delivers 2.5x faster time to first token versus Gemini 2.5 Flash with full multimodal support. Ideal for agentic tasks, data extraction, translation, and classification.
Gemini 3.5 Flash
Google DeepMind's balanced Gemini 3.5 model that pairs Pro-line reasoning quality with Flash-line latency and cost. Natively multimodal across text, image, audio, and video with a 1M-token context window, configurable thinking levels, and streaming function calling, tuned for high-throughput production workloads.
Gemini 3 Flash
Google's balanced model combining Gemini 3 Pro's reasoning capabilities with the Flash line's latency, efficiency, and cost. Features configurable thinking levels, multimodal function responses, and streaming function calling for complex agentic workflows.
GPT-5.4 Mini
OpenAI's compact reasoning model optimized for coding, computer use, and subagent tasks. Approaches GPT-5.4 performance on several benchmarks while running more than 2x faster.
Nemotron 3 Super 120B
NVIDIA's open hybrid Mamba-Transformer MoE model with 120B total parameters (12B active). Features 1M token context window and excels at agentic reasoning, coding, planning, and tool calling.
Grok 4.3
xAI's latest and most intelligent model with strong agentic tool calling, minimal hallucinations, and configurable reasoning. Supports 1M token context window with competitive pricing.
GPT-5.4
OpenAI's frontier reasoning model combining advances in coding, reasoning, and agentic workflows. Features 1.1M token context window and strong performance on complex multi-step problems.
Grok 4.20
xAI's multi-agent capable model with 2M token context window. Available in reasoning, non-reasoning, and multi-agent variants for diverse enterprise workloads.
Grok 4.1 Fast
xAI's fast and cost-effective model with 2M token context window. Offers both reasoning and non-reasoning modes at significantly lower pricing than flagship models.
Gemini 2.5 Pro
Google's high-capability reasoning model with adaptive thinking for complex agentic and multimodal challenges. Features 1M token context window and strong performance on coding and scientific tasks.
Gemini 2.5 Flash
Google's cost-effective model optimized for high throughput tasks. Balances speed and intelligence with strong multimodal capabilities and 1M token context window.
Llama 4 Scout
Meta's efficient MoE model with 17B active parameters (109B total, 16 experts). Supports up to 10M token context — the longest of any production model. Strong performance on reasoning and multilingual tasks.