Modelle

Durchsuchen Sie 203 kanonische LLM-Modelle aller Anbieter

Modelle 1–24 von 203

Gemini 4 Argon

Vereinigte Staaten

Google's top-tier model anchoring the Gemini 4 generation, larger than the previous Pro line and built for complex workloads, coding, and cybersecurity. Can locate, validate, and patch critical software vulnerabilities autonomously and raises the output limit to 1M tokens. At launch, access is limited to selected cybersecurity organizations through Google's Fairwind Program, with paid API customers and Google AI Ultra subscribers to follow; no public availability date has been announced.

Kontext
1.0M
Veröffentlicht
Okt. 2026

GPT-6.1 Sol

Vereinigte Staaten

OpenAI's upgrade to GPT-6 Sol, delivering near-Astra performance for complex work at a lower cost. Suited for agentic coding, computer use, and document-heavy professional work, with multimodal input, function calling, reasoning effort controls, a 1.05M-token context window, up to 128K output tokens, and cheaper cached input than GPT-6 Sol.

Kontext
1.1M
Veröffentlicht
Sept. 2026

Claude Sonnet 5.5

Vereinigte Staaten

Anthropic's Sonnet-class model offering the best combination of speed and intelligence, succeeding Claude Sonnet 5 as a direct upgrade. Especially strong at building features, fixing bugs, and well-scoped everyday coding and knowledge work, with adaptive thinking, a 1M-token context window, and up to 128K output tokens.

Kontext
1.0M
Veröffentlicht
Sept. 2026

Perceptron Mk1.5

Vereinigte Staaten

Perceptron's embodied reasoning model for physical agents. Accepts text, image, video, and audio input and answers with text plus optional structured annotations such as points, boxes, polygons, and tracks, with a 36K-token context window.

Kontext
37K
Veröffentlicht
Sept. 2026

Qwen 3.8 Max Prime

China

A higher-throughput variant of Qwen 3.8 Max from Alibaba's Qwen team, served as a separate SKU at a higher price point. Accepts text, image, and video input with reasoning, tool use, and a 1M-token context window.

Kontext
1.0M
Veröffentlicht
Sept. 2026

Ember-1

Vereinigte Staaten

A specialized reasoning model from Fireworks Research, built on Kimi K3 and designed to make every token go further by producing shorter reasoning traces. Supports image input, tool use, and a 1M-token context window.

Kontext
1.0M
Veröffentlicht
Sept. 2026

GLM-5.3-Prime

China

The high-speed variant of Z.ai's GLM-5.3, inheriting its full capabilities while delivering 1.5-2x the output throughput through inference acceleration. Text input and output with reasoning, tool use, and a 1M-token context window.

Kontext
1.0M
Veröffentlicht
Sept. 2026

Aion 3.5 Mini

The smaller, lower-cost sibling of Aion 3.5, a multi-model roleplaying and storytelling system from AionLabs built on the GLM family of models. Supports reasoning, tool use, and a 262K-token context window.

Kontext
262K
Veröffentlicht
Sept. 2026

Solar Mini 4

Südkorea

Upstage's compact, cost-efficient language model, a 35B-parameter mixture-of-experts with 3B active parameters and a 524K-token context window. Built for agentic use cases where response speed and cost matter, with reasoning and tool use.

Kontext
524K
Veröffentlicht
Sept. 2026

Aion 3.5

A multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. Uses a collaborative generation process in which multiple specialized models contribute to each response, with reasoning, tool use, and a 262K-token context window.

Kontext
262K
Veröffentlicht
Sept. 2026

GPT-6 Sol

Vereinigte Staaten

The cost-efficient high-end model in OpenAI's GPT-6 series, positioned below the flagship GPT-6 Astra and above the fast GPT-6 Luna tier. Suited for demanding professional work, agentic coding, and document-heavy tasks, with multimodal input, function calling, reasoning effort controls, a 1.05M-token context window, and up to 128K output tokens.

Kontext
1.1M
Veröffentlicht
Sept. 2026

MiMo-V2.6-Flash

China

Xiaomi's open-source mixture-of-experts foundation model with 309B total parameters and 15B activated per token, using a hybrid attention mechanism for efficient long-context inference. Multimodal across text, image, video, and audio input, with reasoning, tool use, and a 1M-token context window.

Kontext
1.1M
Veröffentlicht
Sept. 2026

GPT-6 Luna

Vereinigte Staaten

The fast, most efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol for focused, high-volume tasks. Suited for latency-sensitive workloads such as chat, classification, extraction, and lightweight agentic steps, with multimodal input, function calling, reasoning controls, a 1.05M-token context window, and up to 128K output tokens.

Kontext
1.1M
Veröffentlicht
Sept. 2026

MiMo-V2.6-Pro

China

Xiaomi's flagship foundation model, built at a scale of over 1T parameters for the most demanding workloads. Natively multimodal across text, image, video, and audio input, with reasoning, tool use, and a 1M-token context window; weights are published on Hugging Face.

Kontext
1.1M
Veröffentlicht
Sept. 2026

Claude Opus 5.5

Vereinigte Staaten

Anthropic's Opus-class model for long-running agentic coding and knowledge work, succeeding Claude Opus 5 at a lower price. Particularly strong at multi-step changes in large codebases and sustained autonomous tasks, with always-on adaptive thinking steered by an effort parameter, a 1M-token context window, and up to 128K output tokens. Anthropic's recommended starting point for most workloads.

Kontext
1.0M
Veröffentlicht
Sept. 2026

Grok 4.7

Vereinigte Staaten

xAI's flagship model for coding, agentic tasks, and knowledge work, succeeding Grok 4.6. Particularly strong at long-running software engineering tasks and verifying its own work, with image and file input, reasoning, tool use, a 500K-token context window, and an OpenAI-compatible API.

Kontext
500K
Veröffentlicht
Sept. 2026

Qwen 3.8 Omni Flash

China

Alibaba's omni-modal reasoning model and the first Qwen model built around agentic capabilities with native audio-video understanding. Suited for audio-video analysis and summarization and multimodal agents, with tool use and a 1M-token context window.

Kontext
1.0M
Veröffentlicht
Sept. 2026

Ternary Bonsai 2 27B

Vereinigte Staaten

PrismML's open-weight 27B-parameter reasoning model derived from Qwen 3.8 27B and shrunk with ternary compression. Supports coding, mathematics, tool calling, and image understanding with a 262K-token context window.

Kontext
262K
Veröffentlicht
Sept. 2026

GLM-5.3-FlashX

China

The high-speed variant of Z.ai's GLM-5.3-Flash, a natively multimodal model delivering inference speeds of up to 200 tokens per second. Built on the same hybrid sparse and linear attention architecture, with image and video input, reasoning, tool use, and a 1M-token context window.

Kontext
1.0M
Veröffentlicht
Sept. 2026

Pareto

A multimodal composite model from Unbiased built for research, coding, and agentic workflows, aiming at frontier-level performance across a broad range of general-purpose tasks. Supports image input, tool use, and a 262K-token context window.

Kontext
262K
Veröffentlicht
Sept. 2026

Schematron V2 Turbo

Vereinigte Staaten

Inference.net's open-weight 3B-parameter HTML-to-JSON extraction model, prioritizing throughput for high-volume extraction workloads. Extraction instructions are supplied through a JSON schema rather than a prompt. 128K-token context window.

Kontext
128K
Veröffentlicht
Sept. 2026

Schematron V2 Small

Vereinigte Staaten

Inference.net's open-weight 3B-parameter HTML-to-JSON extraction model, prioritizing extraction quality for complex schemas and long pages. Extraction instructions are supplied through a JSON schema rather than a prompt. 128K-token context window.

Kontext
128K
Veröffentlicht
Sept. 2026

Sakana Fugu Max

Japan

The cost-performance model in Sakana AI's Fugu family, a learned multi-agent orchestration system in which a language model routes and coordinates work across a pool of models. Supports image and file input, reasoning, tool use, and a 1M-token context window.

Kontext
1.0M
Veröffentlicht
Sept. 2026

Sakana Fugu Ultra v2

Japan

The higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system, a language model trained to route and coordinate work across a pool of models. Supports image and file input, reasoning, tool use, and a 1M-token context window.

Kontext
1.0M
Veröffentlicht
Sept. 2026