模型

浏览来自所有提供商的 110 个标准化 LLM 模型

部分描述为试点机器翻译内容,尚未经过人工审核。

显示第 25–48 项,共 110 个模型

Gemma 4 12B

美国

Google's medium-size open-weight model with 12 billion parameters from the Gemma 4 family. Encoder-free unified multimodal architecture that natively processes text, image, audio, and video inputs without dedicated encoders. Features a 256K context window and supports 140+ languages. First medium-sized model capable of natively ingesting audio. Suitable for local deployment on GPUs with 16GB VRAM.

上下文
262K
发布日期
2026年6月

MiniMax M3

中国

MiniMax's frontier open-weight model with 1M-token context window, native multimodality (text, image, video), and strong coding capabilities. Built on MiniMax Sparse Attention (MSA) architecture, achieving 59% on SWE-Bench Pro with significantly improved efficiency at long context.

上下文
1.0M
发布日期
2026年6月

Claude Opus 4.8

美国

Anthropic 最先进的模型,在 Opus 4.7 基础上提升了编程、智能体技能、推理和知识工作的多项基准表现。具备更高的诚实性、更高效的工具使用、动态工作流支持和更好的对齐表现。

上下文
300K
发布日期
2026年5月

Jamba Large 1.7

以色列

AI21's latest hybrid SSM-Transformer model with Mixture-of-Experts architecture. Features a 256K context window, improved grounding and instruction-following. 94B total parameters with 398B active, optimized for enterprise long-context tasks.

上下文
262K
发布日期
2026年5月

Falcon-H1

阿拉伯联合酋长国

TII's hybrid Mamba-Transformer model that outperforms comparable offerings from Meta's Llama and Alibaba's Qwen in the 30-70B parameter range. Designed for real-world AI on everyday devices and resource-limited settings with state-of-the-art efficiency.

上下文
131K
发布日期
2026年5月

Palmyra X5

美国

Writer's most advanced adaptive reasoning model with a 1 million token context window. Processes full million-token prompts in approximately 22 seconds with multi-turn function calls in 300ms. Optimized for enterprise agentic AI workflows at 3-4x lower cost than GPT-4.1.

上下文
1.0M
发布日期
2026年5月

Ring-2.6-1T

中国

InclusionAI's (Ant Group) trillion-parameter open-weights reasoning model with 63B active parameters per token. Built for real-world agent workflows with adaptive reasoning-effort modes. Features hybrid linear and MLA attention architecture with MIT license.

上下文
131K
发布日期
2026年5月

Yi-Lightning

中国

01.AI's flagship large language model with enhanced Mixture-of-Experts architecture. Ranked 6th on Chatbot Arena with particularly strong results in Chinese, Math, Coding, and Hard Prompts categories. Features advanced expert segmentation and optimized KV-caching.

上下文
131K
发布日期
2026年5月

K2 Think

阿拉伯联合酋长国

A 32 billion parameter open-weights reasoning model by LLM360/MBZUAI, built on Qwen2.5-32B. Trained with reinforcement learning and verifiable rewards for long chain-of-thought reasoning, agentic planning, and complex problem solving in math, science, and code.

上下文
131K
发布日期
2026年5月

Solar Pro 3

韩国

Upstage's powerful Mixture-of-Experts language model with 102B total parameters and 12B active parameters per forward pass. Optimized for Korean with strong English and Japanese support. Excels at complex reasoning, structured output generation, and agentic workflows.

上下文
128K
发布日期
2026年5月

Qwen 3.7 Max

中国

Alibaba's flagship proprietary model engineered for advanced agentic coding, complex reasoning, and long-horizon task execution. Ranked

上下文
131K
发布日期
2026年5月

Qwen 3.7 Plus

中国

Alibaba's multimodal variant in the Qwen 3.7 family, optimized for vision understanding and multimodal tasks. Ranked

上下文
131K
发布日期
2026年5月

Gemini 3.5 Flash

美国

Google DeepMind's balanced Gemini 3.5 model that pairs Pro-line reasoning quality with Flash-line latency and cost. Natively multimodal across text, image, audio, and video with a 1M-token context window, configurable thinking levels, and streaming function calling, tuned for high-throughput production workloads.

上下文
1.0M
发布日期
2026年5月

Gemini 3.1 Flash-Lite

美国

Google's most cost-efficient Gemini model optimized for high-volume, low-latency use cases. Delivers 2.5x faster time to first token versus Gemini 2.5 Flash with full multimodal support. Ideal for agentic tasks, data extraction, translation, and classification.

上下文
1.0M
发布日期
2026年5月

Gemini 3 Flash

美国

Google's balanced model combining Gemini 3 Pro's reasoning capabilities with the Flash line's latency, efficiency, and cost. Features configurable thinking levels, multimodal function responses, and streaming function calling for complex agentic workflows.

上下文
1.0M
发布日期
2026年5月

MiniCPM-V 4.6

中国

Ultra-efficient multimodal language model from OpenBMB built on SigLIP2-400M and Qwen3.5-0.8B (~1B parameters). Supports single-image, multi-image, and video understanding with mixed 4x/16x visual token compression. Designed for edge deployment on iOS, Android, and HarmonyOS.

上下文
256K
发布日期
2026年5月

Granite 4.1 30B

美国

IBM's largest dense decoder-only 30B parameter language model from the Granite 4.1 family. Trained on approximately 15T tokens with long-context extension up to 512K tokens. Supports tool calling, RAG, code generation, multilingual tasks across 12 languages. Released under Apache 2.0.

上下文
524K
发布日期
2026年5月

Granite 4.1 8B

美国

IBM's dense decoder-only 8B parameter language model from the Granite 4.1 family. Supports 131K-token context, tool calling, RAG, code generation with fill-in-the-middle, text summarization, classification, and extraction across 12 languages. Released under Apache 2.0.

上下文
131K
发布日期
2026年5月

Laguna M.1

美国

Poolside AI's flagship agentic coding model with 225B total parameters and 23B active (MoE). Trained from scratch in-house on 30T tokens across 6,144 NVIDIA Hopper GPUs. Optimized for complex multi-step software engineering tasks including codebase exploration, file editing, test running, and iterative debugging.

上下文
128K
发布日期
2026年4月

DeepSeek V4 Pro

中国

DeepSeek's flagship V4 model with 1.6T total parameters (49B activated). MoE architecture supporting 1M token context. Closes the gap with frontier proprietary models on reasoning and coding benchmarks.

上下文
1.0M
发布日期
2026年4月

DeepSeek V4 Flash

中国

DeepSeek's efficient V4 model with 284B total parameters (13B activated). Optimized for speed and cost-efficiency while maintaining strong performance. Supports 1M token context window.

上下文
1.0M
发布日期
2026年4月

GPT-5.5

美国

OpenAI's most capable model designed for complex real-world work including coding, online research, information analysis, and document creation. Features advanced agentic capabilities with tool search and multi-step task execution.

上下文
1.0M
发布日期
2026年4月

Hy3 Preview

中国

Tencent's flagship open-weight Mixture-of-Experts model from the Hunyuan family with 295B total parameters and 21B active. Integrates fast and slow thinking modes with configurable reasoning effort. Designed for agentic workflows, cross-file code refactoring, long-document analysis, and multi-step tool use.

上下文
256K
发布日期
2026年4月

MiMo-V2.5-Pro

中国

Xiaomi's flagship 1.02T-parameter Mixture-of-Experts model with 42B active parameters, built on a hybrid-attention architecture with 3-layer Multi-Token Prediction. Designed for complex agentic tasks, software engineering, and long-horizon instruction following with a 1M-token context window.

上下文
1.0M
发布日期
2026年4月