203 Modelle · 52 Anbieter · 277 Zuordnungen

Offenes Register
für KI-Infrastruktur

Entdecken und vergleichen Sie Modelle, Anbieter, MCP-Server und Agenten-Skills anhand quellentransparenter Preise, Kontextgrenzen, Funktionen, Zugangsbedingungen und Bereitstellungsdaten.

Anbieter im gesamten Ökosystem im Blick

OpenAI
Anthropic
Google
AWS Bedrock
Meta
Groq
Mistral
DeepSeek
xAI
IBM
Azure AI
Cohere
NVIDIA
Alibaba
Xiaomi
Hugging Face
Together AI
Fireworks
Replicate
SambaNova
Scaleway
Nebius
OpenAI
Anthropic
Google
AWS Bedrock
Meta
Groq
Mistral
DeepSeek
xAI
IBM
Azure AI
Cohere
NVIDIA
Alibaba
Xiaomi
Hugging Face
Together AI
Fireworks
Replicate
SambaNova
Scaleway
Nebius

203

Modelle

52

Anbieter

277

Anbieterzuordnungen

$4.36

Ø Paid I/O $/1 Mio.

Beliebte Modelle

Führende Modelle nach Relevanz und Anbieterverfügbarkeit

Alle anzeigen →

Claude Opus 5.5

Anthropic's Opus-class model for long-running agentic coding and knowledge work, succeeding Claude Opus 5 at a lower price. Particularly strong at multi-step changes in large codebases and sustained autonomous tasks, with always-on adaptive thinking steered by an effort parameter, a 1M-token context window, and up to 128K output tokens. Anthropic's recommended starting point for most workloads.

Kontext
1.0M
Veröffentlicht
Sept. 2026

GPT-6 Astra

OpenAI's most capable frontier model for the hardest end-to-end work across complex reasoning, coding, computer use, browsing, research, scientific analysis, cybersecurity defense, and professional document creation. Supports long-context multimodal input, structured outputs, function calling, reasoning effort controls, MCP, skills, computer use, hosted shell, code interpreter, web search, file search, and apply-patch workflows through the Responses API.

Kontext
1.1M
Veröffentlicht
Sept. 2026

GPT-6.1 Sol

OpenAI's upgrade to GPT-6 Sol, delivering near-Astra performance for complex work at a lower cost. Suited for agentic coding, computer use, and document-heavy professional work, with multimodal input, function calling, reasoning effort controls, a 1.05M-token context window, up to 128K output tokens, and cheaper cached input than GPT-6 Sol.

Kontext
1.1M
Veröffentlicht
Sept. 2026

Gemini 4 Argon

Google's top-tier model anchoring the Gemini 4 generation, larger than the previous Pro line and built for complex workloads, coding, and cybersecurity. Can locate, validate, and patch critical software vulnerabilities autonomously and raises the output limit to 1M tokens. At launch, access is limited to selected cybersecurity organizations through Google's Fairwind Program, with paid API customers and Google AI Ultra subscribers to follow; no public availability date has been announced.

Kontext
1.0M
Veröffentlicht
Okt. 2026

Nemotron 3 Ultra

NVIDIA's flagship open 550B-parameter Mixture-of-Experts model with 55B active parameters, built for frontier reasoning and orchestration in long-running agentic systems. Features hybrid Mamba-Transformer architecture, LatentMoE routing, multi-token prediction, and NVFP4 precision for 5x higher throughput. Achieves 30% lower cost-to-task-completion on agentic benchmarks. Supports 1M+ token context window with 95% accuracy on Ruler@1M.

Kontext
1.0M
Veröffentlicht
Juni 2026

Gemma 4 31B

Google's flagship open-weight dense model with 31B parameters. All parameters active per forward pass. Ranks among top open models with strong performance on AIME 2026 (89.2%) and MMLU Pro (85.2%). Supports vision and extended context.

Kontext
262K
Veröffentlicht
Apr. 2026

Kimi K2.7 Code

Moonshot AI's latest open-source, coding-focused model in the Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. A 1-trillion-parameter model that cuts reasoning token usage by roughly 30% versus K2.6 while improving coding and agent performance — +21.8% on Kimi Code Bench v2, +11.0% on Program Bench, and +31.5% on MLS Bench Lite for multi-language support. Released under a Modified MIT License and available via Kimi APIs and Hugging Face.

Kontext
1.0M
Veröffentlicht
Juni 2026

DeepSeek V4 Flash

DeepSeek's efficient V4 model with 284B total parameters (13B activated). Optimized for speed and cost-efficiency while maintaining strong performance. Supports 1M token context window.

Kontext
1.0M
Veröffentlicht
Apr. 2026

Gemma 4 12B

Google's medium-size open-weight model with 12 billion parameters from the Gemma 4 family. Encoder-free unified multimodal architecture that natively processes text, image, audio, and video inputs without dedicated encoders. Features a 256K context window and supports 140+ languages. First medium-sized model capable of natively ingesting audio. Suitable for local deployment on GPUs with 16GB VRAM.

Kontext
262K
Veröffentlicht
Juni 2026

MiniMax M3

MiniMax's frontier open-weight model with 1M-token context window, native multimodality (text, image, video), and strong coding capabilities. Built on MiniMax Sparse Attention (MSA) architecture, achieving 59% on SWE-Bench Pro with significantly improved efficiency at long context.

Kontext
1.0M
Veröffentlicht
Juni 2026

Qwen 3.8 27B

Alibaba's open-weight 27B dense vision-language model in the Qwen 3.8 family. It builds on the Qwen 3.6 27B line with stronger coding and office productivity capabilities across text and visual modalities, supports one million tokens of context through hosted APIs, and is suited to multimodal assistants, document work, coding, and long-running agent tasks.

Kontext
1.0M
Veröffentlicht
Aug. 2026

Granite 4.1 8B

IBM's dense decoder-only 8B parameter language model from the Granite 4.1 family. Supports 131K-token context, tool calling, RAG, code generation with fill-in-the-middle, text summarization, classification, and extraction across 12 languages. Released under Apache 2.0.

Kontext
131K
Veröffentlicht
Mai 2026

Modellvergleich

Beginnen Sie mit dem Workload. Wählen Sie dann das Modell.

Machen Sie aus einem Modellregister eine Entscheidung. Filtern Sie den gesamten Katalog, vergleichen Sie verlässliche Werte direkt und berechnen Sie Kosten anhand Ihrer eigenen Annahmen.

1

Wichtige Fakten vergleichen

Kontext, Modalitäten, Funktionen, Zugangsbedingungen und Anbieterverfügbarkeit auf einen Blick.

2

Preise dem Anbieter zuordnen

Bereitstellungsspezifische Preise für Ein- und Ausgabe werden niemals als globale Modelleigenschaften dargestellt.

3

Kosten Ihres Workloads schätzen

Nutzen Sie Ihre eigenen Tokenmengen, Cache-Annahmen und Anfragezahlen statt eines abstrakten Scores.

Live-Auszug aus dem Register

Größte dokumentierte Kontextfenster

Kapazität auf Modellebene aus dem Register. Kapazität ist kein Qualitätsmaß.

Kontext entdecken →