Лучшие модели для работы с изображениями

Мультимодальные модели, которые понимают изображения, документы, скриншоты и диаграммы.

54 модели

Inkling

Соединенные Штаты

Thinking Machines Lab's open-weights general-purpose multimodal Mixture-of-Experts model with 975B total parameters and 41B active parameters. Inkling accepts text, image, and audio inputs, produces text, and is designed for agentic and tool-use systems, coding assistants, chatbots, and retrieval-augmented generation.

Контекст
1.0M
Добавлена
июль 2026 г.

Kimi K3

Китай

Moonshot AI's flagship Kimi model for frontier intelligence, agentic coding, knowledge work, and deep reasoning. Kimi K3 supports a 1-million-token context window for long-running software engineering and research workflows.

Контекст
1.0M
Добавлена
июль 2026 г.

GPT-5.6 Luna

Соединенные Штаты

The fast, low-cost tier of OpenAI's GPT-5.6 series, optimized for high-volume, latency-sensitive tasks such as classification, extraction, routing, and lightweight agentic steps. Approaches the larger GPT-5.6 tiers on many benchmarks while running several times faster at a fraction of the price.

Контекст
400K
Добавлена
июль 2026 г.

Muse Spark 1.1

Соединенные Штаты

Meta Superintelligence Labs' updated flagship, building on Muse Spark with stronger agentic reasoning, more reliable multi-agent orchestration, and improved multimodal understanding across voice, text, and image. Extends the context window and reduces latency and reasoning token usage while raising coding and tool-use accuracy. Powers Meta AI across its product ecosystem.

Контекст
512K
Добавлена
июль 2026 г.

Gemini 3.5 Pro

Соединенные Штаты

Google DeepMind's flagship Gemini model, rebuilt on a new foundation with a 2M-token context window and a Deep Think reasoning mode for the hardest math, coding, and multimodal tasks. Natively multimodal across text, image, audio, and video with streaming function calling and strong long-context grounding.

Контекст
2.0M
Добавлена
июль 2026 г.

GPT-5.6 Sol

Соединенные Штаты

OpenAI's flagship model in the GPT-5.6 series, advancing coding, scientific reasoning, long-horizon planning, and agentic workflows while improving reliability and efficiency on demanding real-world tasks. Adds a max reasoning-effort setting and an ultra mode that spawns subagents for complex, multi-step work.

Контекст
1.0M
Добавлена
июль 2026 г.

GPT-5.6 Terra

Соединенные Штаты

The balanced tier of OpenAI's GPT-5.6 series, trading a small amount of peak quality for markedly lower latency and cost. Retains strong reasoning, coding, and agentic tool use with configurable reasoning effort, making it a default choice for production workloads that need frontier capability at scale.

Контекст
1.0M
Добавлена
июль 2026 г.

Grok 4.5

Соединенные Штаты

xAI's strongest model to date, built to excel at coding, agentic tasks, and knowledge work and co-developed alongside coding tools for real-world software engineering. Features real-time information access, extended reasoning, and large-context tool use with an OpenAI-compatible API.

Контекст
500K
Добавлена
июль 2026 г.

Claude Sonnet 5

Соединенные Штаты

Anthropic's most capable Sonnet-class model, bringing frontier coding, agentic, and professional-work performance to the midsize tier while closing the gap with Opus 4.8 at a lower price. Supports adaptive thinking with selectable reasoning effort levels, a 1M-token context window, and text, image, and file inputs. Codenamed Fennec.

Контекст
1.0M
Добавлена
июнь 2026 г.

Claude Fable 5

Соединенные Штаты

Anthropic's first publicly available Mythos-class model, exceeding the capabilities of any model the company has previously made generally available. State-of-the-art on nearly all tested benchmarks, with exceptional performance in software engineering, knowledge work, vision, and scientific research. Its lead grows on longer and more complex tasks. Ships with built-in safeguards that route sensitive cybersecurity, biology, chemistry, and distillation queries to Claude Opus 4.8.

Контекст
300K
Добавлена
июнь 2026 г.

Claude Mythos 5

Соединенные Штаты

Anthropic's frontier Mythos-class model — the same underlying model as Claude Fable 5 but with safeguards lifted in some areas. It has the strongest cybersecurity capabilities of any model in the world, alongside state-of-the-art performance in software engineering, knowledge work, vision, and scientific research. Access is restricted to a small group of trusted cyberdefenders and infrastructure providers through Project Glasswing.

Контекст
300K
Добавлена
июнь 2026 г.

Gemma 4 12B

Соединенные Штаты

Google's medium-size open-weight model with 12 billion parameters from the Gemma 4 family. Encoder-free unified multimodal architecture that natively processes text, image, audio, and video inputs without dedicated encoders. Features a 256K context window and supports 140+ languages. First medium-sized model capable of natively ingesting audio. Suitable for local deployment on GPUs with 16GB VRAM.

Контекст
262K
Добавлена
июнь 2026 г.

MiniMax M3

Китай

MiniMax's frontier open-weight model with 1M-token context window, native multimodality (text, image, video), and strong coding capabilities. Built on MiniMax Sparse Attention (MSA) architecture, achieving 59% on SWE-Bench Pro with significantly improved efficiency at long context.

Контекст
1.0M
Добавлена
июнь 2026 г.

Claude Opus 4.8

Соединенные Штаты

Anthropic's most advanced model, building on Opus 4.7 with improvements across benchmarks in coding, agentic skills, reasoning, and knowledge work. Features enhanced honesty, better tool use efficiency, dynamic workflows support, and improved alignment.

Контекст
300K
Добавлена
май 2026 г.

Qwen 3.7 Plus

Китай

Alibaba's multimodal variant in the Qwen 3.7 family, optimized for vision understanding and multimodal tasks. Ranked

Контекст
131K
Добавлена
май 2026 г.

Gemini 3.1 Flash-Lite

Соединенные Штаты

Google's most cost-efficient Gemini model optimized for high-volume, low-latency use cases. Delivers 2.5x faster time to first token versus Gemini 2.5 Flash with full multimodal support. Ideal for agentic tasks, data extraction, translation, and classification.

Контекст
1.0M
Добавлена
май 2026 г.

Gemini 3.5 Flash

Соединенные Штаты

Google DeepMind's balanced Gemini 3.5 model that pairs Pro-line reasoning quality with Flash-line latency and cost. Natively multimodal across text, image, audio, and video with a 1M-token context window, configurable thinking levels, and streaming function calling, tuned for high-throughput production workloads.

Контекст
1.0M
Добавлена
май 2026 г.

Gemini 3 Flash

Соединенные Штаты

Google's balanced model combining Gemini 3 Pro's reasoning capabilities with the Flash line's latency, efficiency, and cost. Features configurable thinking levels, multimodal function responses, and streaming function calling for complex agentic workflows.

Контекст
1.0M
Добавлена
май 2026 г.

MiniCPM-V 4.6

Китай

Ultra-efficient multimodal language model from OpenBMB built on SigLIP2-400M and Qwen3.5-0.8B (~1B parameters). Supports single-image, multi-image, and video understanding with mixed 4x/16x visual token compression. Designed for edge deployment on iOS, Android, and HarmonyOS.

Контекст
256K
Добавлена
май 2026 г.

GPT-5.5

Соединенные Штаты

OpenAI's most capable model designed for complex real-world work including coding, online research, information analysis, and document creation. Features advanced agentic capabilities with tool search and multi-step task execution.

Контекст
1.0M
Добавлена
апр. 2026 г.

Qwen 3.6 27B

Китай

Alibaba's dense 27B parameter model that outperforms its own 397B MoE predecessor on agentic coding benchmarks. Strong multilingual and reasoning capabilities released under Apache 2.0.

Контекст
131K
Добавлена
апр. 2026 г.

Qwen 3.6 35B-A3B

Китай

Alibaba's efficient Mixture-of-Experts model with 35B total parameters and 3B active per token. Frontier-level agentic coding performance with 73.4% on SWE-bench Verified and 92.7 on AIME 2026. Released under Apache 2.0.

Контекст
131K
Добавлена
апр. 2026 г.

GPT-5.4 Mini

Соединенные Штаты

OpenAI's compact reasoning model optimized for coding, computer use, and subagent tasks. Approaches GPT-5.4 performance on several benchmarks while running more than 2x faster.

Контекст
1.1M
Добавлена
апр. 2026 г.

Muse Spark

Соединенные Штаты

Meta Superintelligence Labs' first model, featuring advanced reasoning, multimodal understanding, and agentic capabilities. Processes voice, text, and image inputs with tool use and multi-agent orchestration. Powers Meta AI across its product ecosystem.

Контекст
256K
Добавлена
апр. 2026 г.

Qwen 3.6 Plus

Китай

Alibaba's proprietary flagship model in the Qwen 3.6 family, targeting enterprise AI workflows with stronger agentic coding capability, visual coding support, and end-to-end enterprise engineering features.

Контекст
131K
Добавлена
апр. 2026 г.

Gemma 4 31B

Соединенные Штаты

Google's flagship open-weight dense model with 31B parameters. All parameters active per forward pass. Ranks among top open models with strong performance on AIME 2026 (89.2%) and MMLU Pro (85.2%). Supports vision and extended context.

Контекст
262K
Добавлена
апр. 2026 г.

Gemma 4 26B

Соединенные Штаты

Google's high-performance open-weight dense model with 26 billion parameters from the Gemma 4 family. Supports multimodal inputs including text and images with a 256K extended context window. Strong reasoning and code generation capabilities with all parameters active per forward pass.

Контекст
262K
Добавлена
апр. 2026 г.

Gemma 4 31B

Соединенные Штаты

Google's flagship open-weight dense model with 31 billion parameters from the Gemma 4 family. All parameters active per forward pass with top-tier performance on reasoning benchmarks including AIME 2026 and MMLU Pro. Supports vision and extended 256K context window.

Контекст
262K
Добавлена
апр. 2026 г.

Claude Opus 4.7

Соединенные Штаты

Anthropic's latest and most advanced model with state-of-the-art reasoning, coding, and analysis capabilities. Features improved tool use, extended thinking, and enhanced safety alignment.

Контекст
300K
Добавлена
апр. 2026 г.

Grok 4.3

Соединенные Штаты

xAI's latest and most intelligent model with strong agentic tool calling, minimal hallucinations, and configurable reasoning. Supports 1M token context window with competitive pricing.

Контекст
1.0M
Добавлена
апр. 2026 г.

GPT-5.4

Соединенные Штаты

OpenAI's frontier reasoning model combining advances in coding, reasoning, and agentic workflows. Features 1.1M token context window and strong performance on complex multi-step problems.

Контекст
1.1M
Добавлена
март 2026 г.

GPT-5.5 Pro

Соединенные Штаты

OpenAI's premium tier model with extended reasoning capabilities, higher accuracy on complex tasks, and priority access. Optimized for professional and enterprise workloads requiring maximum quality.

Контекст
256K
Добавлена
март 2026 г.

Gemini 3.1 Pro

Соединенные Штаты

Google's latest flagship multimodal model with state-of-the-art performance on reasoning, coding, and multimodal understanding. Features native tool use, grounding, and million-token context window.

Контекст
2.0M
Добавлена
март 2026 г.

Grok 4.20

Соединенные Штаты

xAI's multi-agent capable model with 2M token context window. Available in reasoning, non-reasoning, and multi-agent variants for diverse enterprise workloads.

Контекст
2.0M
Добавлена
февр. 2026 г.

Grok 4

Соединенные Штаты

xAI's latest model with real-time information access, strong reasoning capabilities, and competitive performance on coding and analysis tasks. Features improved tool use and multimodal understanding.

Контекст
256K
Добавлена
февр. 2026 г.

Claude Opus 4.6

Соединенные Штаты

Anthropic's most capable model in the Claude 4 family, excelling at complex analysis, extended reasoning, scientific research, and advanced code generation. Features significantly improved accuracy and reduced hallucinations.

Контекст
300K
Добавлена
янв. 2026 г.

Claude Sonnet 4.6

Соединенные Штаты

Anthropic's balanced model offering strong performance at lower cost and latency than Opus. Excellent for everyday coding, analysis, and content generation tasks with good reasoning capabilities.

Контекст
200K
Добавлена
янв. 2026 г.

Mistral Large 3

Франция

Mistral AI's largest open-weight model with 41B active parameters (675B total MoE). State-of-the-art general-purpose multimodal model with 256K context window and powerful agentic capabilities. Released under Apache 2.0.

Контекст
256K
Добавлена
дек. 2025 г.

Grok 4.1 Fast

Соединенные Штаты

xAI's fast and cost-effective model with 2M token context window. Offers both reasoning and non-reasoning modes at significantly lower pricing than flagship models.

Контекст
2.0M
Добавлена
нояб. 2025 г.

Claude Haiku 4.5

Соединенные Штаты

Anthropic's fastest model with near-frontier intelligence. Optimized for high-throughput, low-latency applications requiring quick responses at minimal cost. Supports extended thinking.

Контекст
200K
Добавлена
окт. 2025 г.

Claude Sonnet 4.5

Соединенные Штаты

Anthropic's previous-generation balanced model with strong coding and analysis capabilities. Offers excellent price-performance ratio for production workloads requiring reliable quality.

Контекст
200K
Добавлена
окт. 2025 г.

Gemini 2.5 Pro

Соединенные Штаты

Google's high-capability reasoning model with adaptive thinking for complex agentic and multimodal challenges. Features 1M token context window and strong performance on coding and scientific tasks.

Контекст
1.0M
Добавлена
июнь 2025 г.

Gemini 2.5 Flash

Соединенные Штаты

Google's cost-effective model optimized for high throughput tasks. Balances speed and intelligence with strong multimodal capabilities and 1M token context window.

Контекст
1.0M
Добавлена
июнь 2025 г.

GPT-5

Соединенные Штаты

OpenAI's fifth-generation flagship model with significant improvements in reasoning, multimodal understanding, and code generation. Features enhanced instruction following and expanded context window.

Контекст
256K
Добавлена
июнь 2025 г.

Llama 4 Maverick

Соединенные Штаты

Meta's quality-focused MoE model with 17B active parameters (400B total, 128 experts). Targets quality-critical tasks with benchmark scores competitive with GPT-4o and Gemini 2.5 Pro.

Контекст
1.0M
Добавлена
апр. 2025 г.

Llama 4 Scout

Соединенные Штаты

Meta's efficient MoE model with 17B active parameters (109B total, 16 experts). Supports up to 10M token context — the longest of any production model. Strong performance on reasoning and multilingual tasks.

Контекст
10.0M
Добавлена
апр. 2025 г.

Gemma 3 12B

Соединенные Штаты

Google's mid-size open-weight model with 12 billion parameters from the Gemma 3 family. Supports multimodal inputs including text and images with a 128K context window. Strong performance on reasoning and code generation tasks at moderate compute cost.

Контекст
131K
Добавлена
март 2025 г.

Gemma 3 27B

Соединенные Штаты

Google's largest open-weight model in the Gemma 3 family with 27 billion parameters. Supports multimodal inputs including text and images with a 128K context window. Delivers strong performance across reasoning, code generation, and vision tasks, competitive with larger proprietary models.

Контекст
131K
Добавлена
март 2025 г.

Gemma 3 4B

Соединенные Штаты

Google's compact open-weight model with 4 billion parameters from the Gemma 3 family. Supports multimodal inputs including text and images with a 128K context window. Balances efficiency and capability for vision and language tasks.

Контекст
131K
Добавлена
март 2025 г.

Mistral Small 3.1

Франция

Mistral AI's Small 3.1 model with 24B parameters offering efficient multimodal capabilities including vision, function calling, and code generation with a large 128K context window.

Контекст
128K
Добавлена
март 2025 г.

Llama 3.2 11B Vision Instruct

Соединенные Штаты

Meta's multimodal open-weight model with 11 billion parameters from the Llama 3.2 family. Supports both text and image inputs, enabling visual understanding tasks alongside standard text generation. Suitable for applications requiring vision capabilities at moderate scale.

Контекст
131K
Добавлена
сент. 2024 г.

Llama 3.2 90B Vision Instruct

Соединенные Штаты

Meta's largest multimodal open-weight model with 90 billion parameters from the Llama 3.2 family. Delivers strong performance on both text and image understanding tasks with competitive results on visual reasoning benchmarks. Designed for high-quality inference requiring vision capabilities.

Контекст
131K
Добавлена
сент. 2024 г.

Claude 3 Opus

Соединенные Штаты

Anthropic's most powerful model in the Claude 3 family, excelling at complex analysis, nuanced content generation, scientific reasoning, and code generation with extended context support.

Контекст
200K
Добавлена
март 2024 г.

GPT-4

Соединенные Штаты

OpenAI's flagship large language model with advanced reasoning, instruction following, and code generation capabilities. Supports multimodal inputs including text and images.

Контекст
128K
Добавлена
янв. 2024 г.