Modelos

Explora 41 modelos LLM canónicos de todos los proveedores

Mostrando 1–24 de 41 modelos

Gemini 3.5 Flash-Lite1.0M ctx

Google's fastest and most cost-effective Gemini 3.5-class model, delivering around 350 output tokens per second per the Artificial Analysis Index. Designed for low-latency and high-throughput agentic workflows such as agentic search and document processing, with configurable thinking levels, built-in computer use, and full multimodal support across a 1M-token context window.

Sarvam-M131K ctx

Sarvam AI's 24B-parameter instruction-tuned model derived from Mistral-Small-3.1-24B, post-trained on English plus eleven major Indic languages (bn, hi, kn, gu, mr, ml, or, pa, ta, te). Delivers large relative gains on Indian-language, math, and programming benchmarks over its base model, with a hybrid reasoning mode for complex tasks.

Nemotron 3 Ultra1.0M ctx

NVIDIA's flagship open 550B-parameter Mixture-of-Experts model with 55B active parameters, built for frontier reasoning and orchestration in long-running agentic systems. Features hybrid Mamba-Transformer architecture, LatentMoE routing, multi-token prediction, and NVFP4 precision for 5x higher throughput. Achieves 30% lower cost-to-task-completion on agentic benchmarks. Supports 1M+ token context window with 95% accuracy on Ruler@1M.

Gemma 4 12B262K ctx

Google's medium-size open-weight model with 12 billion parameters from the Gemma 4 family. Encoder-free unified multimodal architecture that natively processes text, image, audio, and video inputs without dedicated encoders. Features a 256K context window and supports 140+ languages. First medium-sized model capable of natively ingesting audio. Suitable for local deployment on GPUs with 16GB VRAM.

MiniMax M31.0M ctx

MiniMax's frontier open-weight model with 1M-token context window, native multimodality (text, image, video), and strong coding capabilities. Built on MiniMax Sparse Attention (MSA) architecture, achieving 59% on SWE-Bench Pro with significantly improved efficiency at long context.

Falcon-H1131K ctx

TII's hybrid Mamba-Transformer model that outperforms comparable offerings from Meta's Llama and Alibaba's Qwen in the 30-70B parameter range. Designed for real-world AI on everyday devices and resource-limited settings with state-of-the-art efficiency.

Ring-2.6-1T131K ctx

InclusionAI's (Ant Group) trillion-parameter open-weights reasoning model with 63B active parameters per token. Built for real-world agent workflows with adaptive reasoning-effort modes. Features hybrid linear and MLA attention architecture with MIT license.

Solar Pro 3128K ctx

Upstage's powerful Mixture-of-Experts language model with 102B total parameters and 12B active parameters per forward pass. Optimized for Korean with strong English and Japanese support. Excels at complex reasoning, structured output generation, and agentic workflows.

Gemini 3.1 Flash-Lite1.0M ctx

Google's most cost-efficient Gemini model optimized for high-volume, low-latency use cases. Delivers 2.5x faster time to first token versus Gemini 2.5 Flash with full multimodal support. Ideal for agentic tasks, data extraction, translation, and classification.

Gemini 3.5 Flash1.0M ctx

Google DeepMind's balanced Gemini 3.5 model that pairs Pro-line reasoning quality with Flash-line latency and cost. Natively multimodal across text, image, audio, and video with a 1M-token context window, configurable thinking levels, and streaming function calling, tuned for high-throughput production workloads.

Gemini 3 Flash1.0M ctx

Google's balanced model combining Gemini 3 Pro's reasoning capabilities with the Flash line's latency, efficiency, and cost. Features configurable thinking levels, multimodal function responses, and streaming function calling for complex agentic workflows.

Granite 4.1 8B131K ctx

IBM's dense decoder-only 8B parameter language model from the Granite 4.1 family. Supports 131K-token context, tool calling, RAG, code generation with fill-in-the-middle, text summarization, classification, and extraction across 12 languages. Released under Apache 2.0.

Laguna M.1128K ctx

Poolside AI's flagship agentic coding model with 225B total parameters and 23B active (MoE). Trained from scratch in-house on 30T tokens across 6,144 NVIDIA Hopper GPUs. Optimized for complex multi-step software engineering tasks including codebase exploration, file editing, test running, and iterative debugging.

DeepSeek V4 Pro1.0M ctx

DeepSeek's flagship V4 model with 1.6T total parameters (49B activated). MoE architecture supporting 1M token context. Closes the gap with frontier proprietary models on reasoning and coding benchmarks.

DeepSeek V4 Flash1.0M ctx

DeepSeek's efficient V4 model with 284B total parameters (13B activated). Optimized for speed and cost-efficiency while maintaining strong performance. Supports 1M token context window.

Hy3 Preview256K ctx

Tencent's flagship open-weight Mixture-of-Experts model from the Hunyuan family with 295B total parameters and 21B active. Integrates fast and slow thinking modes with configurable reasoning effort. Designed for agentic workflows, cross-file code refactoring, long-document analysis, and multi-step tool use.

Qwen 3.6 27B131K ctx

Alibaba's dense 27B parameter model that outperforms its own 397B MoE predecessor on agentic coding benchmarks. Strong multilingual and reasoning capabilities released under Apache 2.0.

Qwen 3.6 35B-A3B131K ctx

Alibaba's efficient Mixture-of-Experts model with 35B total parameters and 3B active per token. Frontier-level agentic coding performance with 73.4% on SWE-bench Verified and 92.7 on AIME 2026. Released under Apache 2.0.

GPT-OSS 20B131K ctx

OpenAI's compact open-weight model with 20 billion parameters. Released under Apache 2.0 license, designed for efficient deployment on consumer hardware while maintaining strong coding and reasoning capabilities.

Nemotron 3 Super 120B1.0M ctx

NVIDIA's open hybrid Mamba-Transformer MoE model with 120B total parameters (12B active). Features 1M token context window and excels at agentic reasoning, coding, planning, and tool calling.

Mistral Small 4128K ctx

Mistral AI's efficient hybrid model unifying instruct, reasoning, and coding in a single model. Open-weight under Apache 2.0 with strong performance for its size class.

Qwen 3.6131K ctx

Alibaba's latest Qwen model with enhanced reasoning, multilingual capabilities, and improved instruction following. Features strong performance on coding, math, and general knowledge benchmarks.

Kimi K2.61.0M ctx

Moonshot AI's latest model with ultra-long context window support, strong reasoning capabilities, and excellent performance on complex multi-step tasks. Known for reliable long-document understanding.

GLM-5.1131K ctx

Zhipu AI's latest bilingual model with strong Chinese and English capabilities. Features improved reasoning, coding, and tool use with competitive performance on academic benchmarks.