Modelos

Explora 129 modelos LLM canónicos de todos los proveedores

Algunas descripciones forman parte del piloto de traducción automática y aún no han sido revisadas.

Mostrando 25–48 de 129 modelos

Nemotron 3 Ultra

Estados Unidos

NVIDIA's flagship open 550B-parameter Mixture-of-Experts model with 55B active parameters, built for frontier reasoning and orchestration in long-running agentic systems. Features hybrid Mamba-Transformer architecture, LatentMoE routing, multi-token prediction, and NVFP4 precision for 5x higher throughput. Achieves 30% lower cost-to-task-completion on agentic benchmarks. Supports 1M+ token context window with 95% accuracy on Ruler@1M.

Contexto
1.0M
Publicado
jun 2026

Gemma 4 12B

Estados Unidos

Google's medium-size open-weight model with 12 billion parameters from the Gemma 4 family. Encoder-free unified multimodal architecture that natively processes text, image, audio, and video inputs without dedicated encoders. Features a 256K context window and supports 140+ languages. First medium-sized model capable of natively ingesting audio. Suitable for local deployment on GPUs with 16GB VRAM.

Contexto
262K
Publicado
jun 2026

MiniMax M3

China

MiniMax's frontier open-weight model with 1M-token context window, native multimodality (text, image, video), and strong coding capabilities. Built on MiniMax Sparse Attention (MSA) architecture, achieving 59% on SWE-Bench Pro with significantly improved efficiency at long context.

Contexto
1.0M
Publicado
jun 2026

Claude Opus 4.8

Estados Unidos

El modelo más avanzado de Anthropic, basado en Opus 4.7 y mejorado en benchmarks de programación, capacidades de agentes, razonamiento y trabajo del conocimiento. Incorpora mayor honestidad, un uso más eficiente de herramientas, compatibilidad con flujos de trabajo dinámicos y una mejor alineación.

Contexto
300K
Publicado
may 2026

Jamba Large 1.7

Israel

AI21's latest hybrid SSM-Transformer model with Mixture-of-Experts architecture. Features a 256K context window, improved grounding and instruction-following. 94B total parameters with 398B active, optimized for enterprise long-context tasks.

Contexto
262K
Publicado
may 2026

Falcon-H1

Emiratos Árabes Unidos

TII's hybrid Mamba-Transformer model that outperforms comparable offerings from Meta's Llama and Alibaba's Qwen in the 30-70B parameter range. Designed for real-world AI on everyday devices and resource-limited settings with state-of-the-art efficiency.

Contexto
131K
Publicado
may 2026

Palmyra X5

Estados Unidos

Writer's most advanced adaptive reasoning model with a 1 million token context window. Processes full million-token prompts in approximately 22 seconds with multi-turn function calls in 300ms. Optimized for enterprise agentic AI workflows at 3-4x lower cost than GPT-4.1.

Contexto
1.0M
Publicado
may 2026

Snowflake Arctic

Estados Unidos

Snowflake's enterprise-focused open LLM with 480B total parameters using a fine-grained MoE architecture with only 17B active parameters per input. Apache 2.0 licensed, excels at SQL generation, coding, and enterprise intelligence tasks with breakthrough training efficiency.

Contexto
4K
Publicado
may 2026

Falcon 3 10B

Emiratos Árabes Unidos

TII's open-source 10B parameter model from the Falcon 3 family. Achieved number one position on Hugging Face's LLM leaderboard in its size category, outperforming Meta's Llama variants and other models under 13B parameters.

Contexto
33K
Publicado
may 2026

StableLM 2 12B

Reino Unido

Stability AI's 12.1 billion parameter decoder-only language model pre-trained on 2 trillion tokens of diverse multilingual and code datasets. Supports multiple languages and offers strong performance for its compact size with instruction-tuned chat variant available.

Contexto
4K
Publicado
may 2026

DBRX

Estados Unidos

Databricks' open-source 132B parameter Mixture-of-Experts transformer model with 36B active parameters per input. Released under Databricks Open Model License, optimized for enterprise workloads including SQL generation and coding tasks.

Contexto
33K
Publicado
may 2026

Yi-Lightning

China

01.AI's flagship large language model with enhanced Mixture-of-Experts architecture. Ranked 6th on Chatbot Arena with particularly strong results in Chinese, Math, Coding, and Hard Prompts categories. Features advanced expert segmentation and optimized KV-caching.

Contexto
131K
Publicado
may 2026

Ring-2.6-1T

China

InclusionAI's (Ant Group) trillion-parameter open-weights reasoning model with 63B active parameters per token. Built for real-world agent workflows with adaptive reasoning-effort modes. Features hybrid linear and MLA attention architecture with MIT license.

Contexto
131K
Publicado
may 2026

Alloma 8B Instruct

Uzbekistán

Uzbek LLM Lab's 8B parameter instruction-tuned model optimized for the Uzbek language. Built on Llama architecture with a custom tokenizer averaging 1.7 tokens per Uzbek word versus 3.5 in original Llama, enabling 2x faster inference. Trained on 3.6B tokens with 4096 context length.

Contexto
4K
Publicado
may 2026

Solar Pro 3

Corea del Sur

Upstage's powerful Mixture-of-Experts language model with 102B total parameters and 12B active parameters per forward pass. Optimized for Korean with strong English and Japanese support. Excels at complex reasoning, structured output generation, and agentic workflows.

Contexto
128K
Publicado
may 2026

K2 Think

Emiratos Árabes Unidos

A 32 billion parameter open-weights reasoning model by LLM360/MBZUAI, built on Qwen2.5-32B. Trained with reinforcement learning and verifiable rewards for long chain-of-thought reasoning, agentic planning, and complex problem solving in math, science, and code.

Contexto
131K
Publicado
may 2026

Qwen 3.7 Max

China

Alibaba's flagship proprietary model engineered for advanced agentic coding, complex reasoning, and long-horizon task execution. Ranked

Contexto
131K
Publicado
may 2026

Qwen 3.7 Plus

China

Alibaba's multimodal variant in the Qwen 3.7 family, optimized for vision understanding and multimodal tasks. Ranked

Contexto
131K
Publicado
may 2026

Gemini 3.1 Flash-Lite

Estados Unidos

Google's most cost-efficient Gemini model optimized for high-volume, low-latency use cases. Delivers 2.5x faster time to first token versus Gemini 2.5 Flash with full multimodal support. Ideal for agentic tasks, data extraction, translation, and classification.

Contexto
1.0M
Publicado
may 2026

Gemini 3.5 Flash

Estados Unidos

Google DeepMind's balanced Gemini 3.5 model that pairs Pro-line reasoning quality with Flash-line latency and cost. Natively multimodal across text, image, audio, and video with a 1M-token context window, configurable thinking levels, and streaming function calling, tuned for high-throughput production workloads.

Contexto
1.0M
Publicado
may 2026

Gemini 3 Flash

Estados Unidos

Google's balanced model combining Gemini 3 Pro's reasoning capabilities with the Flash line's latency, efficiency, and cost. Features configurable thinking levels, multimodal function responses, and streaming function calling for complex agentic workflows.

Contexto
1.0M
Publicado
may 2026

MiniCPM-V 4.6

China

Ultra-efficient multimodal language model from OpenBMB built on SigLIP2-400M and Qwen3.5-0.8B (~1B parameters). Supports single-image, multi-image, and video understanding with mixed 4x/16x visual token compression. Designed for edge deployment on iOS, Android, and HarmonyOS.

Contexto
256K
Publicado
may 2026

Granite 4.1 8B

Estados Unidos

IBM's dense decoder-only 8B parameter language model from the Granite 4.1 family. Supports 131K-token context, tool calling, RAG, code generation with fill-in-the-middle, text summarization, classification, and extraction across 12 languages. Released under Apache 2.0.

Contexto
131K
Publicado
may 2026

Granite 4.1 30B

Estados Unidos

IBM's largest dense decoder-only 30B parameter language model from the Granite 4.1 family. Trained on approximately 15T tokens with long-context extension up to 512K tokens. Supports tool calling, RAG, code generation, multilingual tasks across 12 languages. Released under Apache 2.0.

Contexto
524K
Publicado
may 2026