Modelos

Explora 203 modelos LLM canónicos de todos los proveedores

Mostrando 1–24 de 203 modelos

Gemini 4 Argon

Estados Unidos

Google's top-tier model anchoring the Gemini 4 generation, larger than the previous Pro line and built for complex workloads, coding, and cybersecurity. Can locate, validate, and patch critical software vulnerabilities autonomously and raises the output limit to 1M tokens. At launch, access is limited to selected cybersecurity organizations through Google's Fairwind Program, with paid API customers and Google AI Ultra subscribers to follow; no public availability date has been announced.

Contexto
1.0M
Publicado
oct 2026

GPT-6.1 Sol

Estados Unidos

OpenAI's upgrade to GPT-6 Sol, delivering near-Astra performance for complex work at a lower cost. Suited for agentic coding, computer use, and document-heavy professional work, with multimodal input, function calling, reasoning effort controls, a 1.05M-token context window, up to 128K output tokens, and cheaper cached input than GPT-6 Sol.

Contexto
1.1M
Publicado
sept 2026

Claude Sonnet 5.5

Estados Unidos

Anthropic's Sonnet-class model offering the best combination of speed and intelligence, succeeding Claude Sonnet 5 as a direct upgrade. Especially strong at building features, fixing bugs, and well-scoped everyday coding and knowledge work, with adaptive thinking, a 1M-token context window, and up to 128K output tokens.

Contexto
1.0M
Publicado
sept 2026

Perceptron Mk1.5

Estados Unidos

Perceptron's embodied reasoning model for physical agents. Accepts text, image, video, and audio input and answers with text plus optional structured annotations such as points, boxes, polygons, and tracks, with a 36K-token context window.

Contexto
37K
Publicado
sept 2026

Qwen 3.8 Max Prime

China

A higher-throughput variant of Qwen 3.8 Max from Alibaba's Qwen team, served as a separate SKU at a higher price point. Accepts text, image, and video input with reasoning, tool use, and a 1M-token context window.

Contexto
1.0M
Publicado
sept 2026

Ember-1

Estados Unidos

A specialized reasoning model from Fireworks Research, built on Kimi K3 and designed to make every token go further by producing shorter reasoning traces. Supports image input, tool use, and a 1M-token context window.

Contexto
1.0M
Publicado
sept 2026

GLM-5.3-Prime

China

The high-speed variant of Z.ai's GLM-5.3, inheriting its full capabilities while delivering 1.5-2x the output throughput through inference acceleration. Text input and output with reasoning, tool use, and a 1M-token context window.

Contexto
1.0M
Publicado
sept 2026

Aion 3.5 Mini

The smaller, lower-cost sibling of Aion 3.5, a multi-model roleplaying and storytelling system from AionLabs built on the GLM family of models. Supports reasoning, tool use, and a 262K-token context window.

Contexto
262K
Publicado
sept 2026

Solar Mini 4

Corea del Sur

Upstage's compact, cost-efficient language model, a 35B-parameter mixture-of-experts with 3B active parameters and a 524K-token context window. Built for agentic use cases where response speed and cost matter, with reasoning and tool use.

Contexto
524K
Publicado
sept 2026

Aion 3.5

A multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. Uses a collaborative generation process in which multiple specialized models contribute to each response, with reasoning, tool use, and a 262K-token context window.

Contexto
262K
Publicado
sept 2026

GPT-6 Sol

Estados Unidos

The cost-efficient high-end model in OpenAI's GPT-6 series, positioned below the flagship GPT-6 Astra and above the fast GPT-6 Luna tier. Suited for demanding professional work, agentic coding, and document-heavy tasks, with multimodal input, function calling, reasoning effort controls, a 1.05M-token context window, and up to 128K output tokens.

Contexto
1.1M
Publicado
sept 2026

MiMo-V2.6-Flash

China

Xiaomi's open-source mixture-of-experts foundation model with 309B total parameters and 15B activated per token, using a hybrid attention mechanism for efficient long-context inference. Multimodal across text, image, video, and audio input, with reasoning, tool use, and a 1M-token context window.

Contexto
1.1M
Publicado
sept 2026

GPT-6 Luna

Estados Unidos

The fast, most efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol for focused, high-volume tasks. Suited for latency-sensitive workloads such as chat, classification, extraction, and lightweight agentic steps, with multimodal input, function calling, reasoning controls, a 1.05M-token context window, and up to 128K output tokens.

Contexto
1.1M
Publicado
sept 2026

MiMo-V2.6-Pro

China

Xiaomi's flagship foundation model, built at a scale of over 1T parameters for the most demanding workloads. Natively multimodal across text, image, video, and audio input, with reasoning, tool use, and a 1M-token context window; weights are published on Hugging Face.

Contexto
1.1M
Publicado
sept 2026

Claude Opus 5.5

Estados Unidos

Anthropic's Opus-class model for long-running agentic coding and knowledge work, succeeding Claude Opus 5 at a lower price. Particularly strong at multi-step changes in large codebases and sustained autonomous tasks, with always-on adaptive thinking steered by an effort parameter, a 1M-token context window, and up to 128K output tokens. Anthropic's recommended starting point for most workloads.

Contexto
1.0M
Publicado
sept 2026

Grok 4.7

Estados Unidos

xAI's flagship model for coding, agentic tasks, and knowledge work, succeeding Grok 4.6. Particularly strong at long-running software engineering tasks and verifying its own work, with image and file input, reasoning, tool use, a 500K-token context window, and an OpenAI-compatible API.

Contexto
500K
Publicado
sept 2026

Qwen 3.8 Omni Flash

China

Alibaba's omni-modal reasoning model and the first Qwen model built around agentic capabilities with native audio-video understanding. Suited for audio-video analysis and summarization and multimodal agents, with tool use and a 1M-token context window.

Contexto
1.0M
Publicado
sept 2026

Ternary Bonsai 2 27B

Estados Unidos

PrismML's open-weight 27B-parameter reasoning model derived from Qwen 3.8 27B and shrunk with ternary compression. Supports coding, mathematics, tool calling, and image understanding with a 262K-token context window.

Contexto
262K
Publicado
sept 2026

GLM-5.3-FlashX

China

The high-speed variant of Z.ai's GLM-5.3-Flash, a natively multimodal model delivering inference speeds of up to 200 tokens per second. Built on the same hybrid sparse and linear attention architecture, with image and video input, reasoning, tool use, and a 1M-token context window.

Contexto
1.0M
Publicado
sept 2026

Pareto

A multimodal composite model from Unbiased built for research, coding, and agentic workflows, aiming at frontier-level performance across a broad range of general-purpose tasks. Supports image input, tool use, and a 262K-token context window.

Contexto
262K
Publicado
sept 2026

Schematron V2 Turbo

Estados Unidos

Inference.net's open-weight 3B-parameter HTML-to-JSON extraction model, prioritizing throughput for high-volume extraction workloads. Extraction instructions are supplied through a JSON schema rather than a prompt. 128K-token context window.

Contexto
128K
Publicado
sept 2026

Schematron V2 Small

Estados Unidos

Inference.net's open-weight 3B-parameter HTML-to-JSON extraction model, prioritizing extraction quality for complex schemas and long pages. Extraction instructions are supplied through a JSON schema rather than a prompt. 128K-token context window.

Contexto
128K
Publicado
sept 2026

Sakana Fugu Max

Japón

The cost-performance model in Sakana AI's Fugu family, a learned multi-agent orchestration system in which a language model routes and coordinates work across a pool of models. Supports image and file input, reasoning, tool use, and a 1M-token context window.

Contexto
1.0M
Publicado
sept 2026

Sakana Fugu Ultra v2

Japón

The higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system, a language model trained to route and coordinate work across a pool of models. Supports image and file input, reasoning, tool use, and a 1M-token context window.

Contexto
1.0M
Publicado
sept 2026