Offenes Register
für KI-Infrastruktur
Entdecken und vergleichen Sie Modelle, Anbieter, MCP-Server und Agenten-Skills anhand quellentransparenter Preise, Kontextgrenzen, Funktionen, Zugangsbedingungen und Bereitstellungsdaten.
Anbieter im gesamten Ökosystem im Blick
203
Modelle
52
Anbieter
277
Anbieterzuordnungen
$4.36
Ø Paid I/O $/1 Mio.
Beliebte Modelle
Führende Modelle nach Relevanz und Anbieterverfügbarkeit
Claude Opus 5.5
Anthropic's Opus-class model for long-running agentic coding and knowledge work, succeeding Claude Opus 5 at a lower price. Particularly strong at multi-step changes in large codebases and sustained autonomous tasks, with always-on adaptive thinking steered by an effort parameter, a 1M-token context window, and up to 128K output tokens. Anthropic's recommended starting point for most workloads.
GPT-6 Astra
OpenAI's most capable frontier model for the hardest end-to-end work across complex reasoning, coding, computer use, browsing, research, scientific analysis, cybersecurity defense, and professional document creation. Supports long-context multimodal input, structured outputs, function calling, reasoning effort controls, MCP, skills, computer use, hosted shell, code interpreter, web search, file search, and apply-patch workflows through the Responses API.
GPT-6.1 Sol
OpenAI's upgrade to GPT-6 Sol, delivering near-Astra performance for complex work at a lower cost. Suited for agentic coding, computer use, and document-heavy professional work, with multimodal input, function calling, reasoning effort controls, a 1.05M-token context window, up to 128K output tokens, and cheaper cached input than GPT-6 Sol.
Gemini 4 Argon
Google's top-tier model anchoring the Gemini 4 generation, larger than the previous Pro line and built for complex workloads, coding, and cybersecurity. Can locate, validate, and patch critical software vulnerabilities autonomously and raises the output limit to 1M tokens. At launch, access is limited to selected cybersecurity organizations through Google's Fairwind Program, with paid API customers and Google AI Ultra subscribers to follow; no public availability date has been announced.
Nemotron 3 Ultra
NVIDIA's flagship open 550B-parameter Mixture-of-Experts model with 55B active parameters, built for frontier reasoning and orchestration in long-running agentic systems. Features hybrid Mamba-Transformer architecture, LatentMoE routing, multi-token prediction, and NVFP4 precision for 5x higher throughput. Achieves 30% lower cost-to-task-completion on agentic benchmarks. Supports 1M+ token context window with 95% accuracy on Ruler@1M.
Gemma 4 31B
Google's flagship open-weight dense model with 31B parameters. All parameters active per forward pass. Ranks among top open models with strong performance on AIME 2026 (89.2%) and MMLU Pro (85.2%). Supports vision and extended context.
Kimi K2.7 Code
Moonshot AI's latest open-source, coding-focused model in the Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. A 1-trillion-parameter model that cuts reasoning token usage by roughly 30% versus K2.6 while improving coding and agent performance — +21.8% on Kimi Code Bench v2, +11.0% on Program Bench, and +31.5% on MLS Bench Lite for multi-language support. Released under a Modified MIT License and available via Kimi APIs and Hugging Face.
DeepSeek V4 Flash
DeepSeek's efficient V4 model with 284B total parameters (13B activated). Optimized for speed and cost-efficiency while maintaining strong performance. Supports 1M token context window.
Gemma 4 12B
Google's medium-size open-weight model with 12 billion parameters from the Gemma 4 family. Encoder-free unified multimodal architecture that natively processes text, image, audio, and video inputs without dedicated encoders. Features a 256K context window and supports 140+ languages. First medium-sized model capable of natively ingesting audio. Suitable for local deployment on GPUs with 16GB VRAM.
MiniMax M3
MiniMax's frontier open-weight model with 1M-token context window, native multimodality (text, image, video), and strong coding capabilities. Built on MiniMax Sparse Attention (MSA) architecture, achieving 59% on SWE-Bench Pro with significantly improved efficiency at long context.
Qwen 3.8 27B
Alibaba's open-weight 27B dense vision-language model in the Qwen 3.8 family. It builds on the Qwen 3.6 27B line with stronger coding and office productivity capabilities across text and visual modalities, supports one million tokens of context through hosted APIs, and is suited to multimodal assistants, document work, coding, and long-running agent tasks.
Granite 4.1 8B
IBM's dense decoder-only 8B parameter language model from the Granite 4.1 family. Supports 131K-token context, tool calling, RAG, code generation with fill-in-the-middle, text summarization, classification, and extraction across 12 languages. Released under Apache 2.0.
Modellvergleich
Beginnen Sie mit dem Workload. Wählen Sie dann das Modell.
Machen Sie aus einem Modellregister eine Entscheidung. Filtern Sie den gesamten Katalog, vergleichen Sie verlässliche Werte direkt und berechnen Sie Kosten anhand Ihrer eigenen Annahmen.
Wichtige Fakten vergleichen
Kontext, Modalitäten, Funktionen, Zugangsbedingungen und Anbieterverfügbarkeit auf einen Blick.
Preise dem Anbieter zuordnen
Bereitstellungsspezifische Preise für Ein- und Ausgabe werden niemals als globale Modelleigenschaften dargestellt.
Kosten Ihres Workloads schätzen
Nutzen Sie Ihre eigenen Tokenmengen, Cache-Annahmen und Anfragezahlen statt eines abstrakten Scores.
Live-Auszug aus dem Register
Größte dokumentierte Kontextfenster
Kapazität auf Modellebene aus dem Register. Kapazität ist kein Qualitätsmaß.
Neueste Analysen
Analysen, Benchmarks und Vergleiche im gesamten LLM-Ökosystem
Dieser Inhalt ist derzeit auf Englisch verfügbar.

GPT-5.6 and ChatGPT Work: From AI Assistant to AI Worker
OpenAI is no longer positioning ChatGPT as a conversational assistant. With GPT-5.6 and ChatGPT Work, the company is moving toward a full work execution layer across apps, files, code, and business workflows.

The AI Race Is Shifting From IQ to Agentic Economics
The AI race is shifting from benchmark scores to agentic economics. Why inference costs, latency, and open-weight models are reshaping the industry in 2026.

Stanford AI Index 2026: AI Is Scaling Faster Than Society Can Adapt
The release of the 2026 AI Index Report by Stanford HAI paints a very clear picture: artificial intelligence is no longer an emerging technology — it has become global infrastructure.