模型
浏览来自所有提供商的 114 个标准化 LLM 模型
部分描述为试点机器翻译内容,尚未经过人工审核。
Google DeepMind's workhorse Flash model that builds on Gemini 3.5 Flash with better coding, knowledge work, and multimodal performance while reducing output token usage by roughly 17% per the Artificial Analysis Index. Natively multimodal across text, image, audio, and video with a 1M-token context window, configurable thinking levels, and built-in computer use, tuned for scaling agentic workflows at a lower cost per output token.
Google's fastest and most cost-effective Gemini 3.5-class model, delivering around 350 output tokens per second per the Artificial Analysis Index. Designed for low-latency and high-throughput agentic workflows such as agentic search and document processing, with configurable thinking levels, built-in computer use, and full multimodal support across a 1M-token context window.
A specialized, cyber-focused Gemini model built on top of Gemini 3.5 Flash and fine-tuned for finding and fixing cybersecurity vulnerabilities at a lower price per token than larger models. Deployed within Google's CodeMender code security agent, where multiple 3.5 Flash Cyber agents collaborate to reach competitive frontier performance on benchmarks like CyberGym. Given its dual-use nature, it is available exclusively to governments and trusted partners via CodeMender as a limited-access pilot.
Thinking Machines Lab's open-weights general-purpose multimodal Mixture-of-Experts model with 975B total parameters and 41B active parameters. Inkling accepts text, image, and audio inputs, produces text, and is designed for agentic and tool-use systems, coding assistants, chatbots, and retrieval-augmented generation.
Moonshot AI's flagship Kimi model for frontier intelligence, agentic coding, knowledge work, and deep reasoning. Kimi K3 supports a 1-million-token context window for long-running software engineering and research workflows.
Google DeepMind 的旗舰 Gemini 模型,基于全新底层架构重建,拥有 200 万 token 上下文窗口,并提供面向高难度数学、编程和多模态任务的 Deep Think 推理模式。原生支持文本、图像、音频和视频,具备流式函数调用和强大的长上下文信息关联能力。
OpenAI GPT-5.6 系列的均衡层级,以少量峰值质量换取显著更低的延迟和成本。它保留了强大的推理、编程和智能体工具使用能力,并支持可配置的推理强度,适合作为需要大规模前沿能力的生产工作负载的默认选择。
The fast, low-cost tier of OpenAI's GPT-5.6 series, optimized for high-volume, latency-sensitive tasks such as classification, extraction, routing, and lightweight agentic steps. Approaches the larger GPT-5.6 tiers on many benchmarks while running several times faster at a fraction of the price.
OpenAI GPT-5.6 系列旗舰模型,在提升可靠性和效率的同时,推进了编程、科学推理、长周期规划和智能体工作流能力。新增最高推理强度设置,以及可为复杂多步骤任务启动子智能体的 ultra 模式。
Meta Superintelligence Labs' updated flagship, building on Muse Spark with stronger agentic reasoning, more reliable multi-agent orchestration, and improved multimodal understanding across voice, text, and image. Extends the context window and reduces latency and reasoning token usage while raising coding and tool-use accuracy. Powers Meta AI across its product ecosystem.
xAI 迄今最强的模型,面向编程、智能体任务和知识工作,并与真实软件工程中的编程工具协同开发。提供实时信息访问、扩展推理和大上下文工具调用,并兼容 OpenAI API。
Anthropic 能力最强的 Sonnet 级模型,以更低价格将前沿编程、智能体和专业工作能力带到中型层级,并缩小了与 Opus 4.8 的差距。支持可选推理强度的自适应思考、100 万 token 上下文窗口,以及文本、图像和文件输入。代号 Fennec。
Sakana AI's multi-agent orchestration model from Tokyo, delivered as a single OpenAI-compatible API. Fugu is itself a language model trained to call a pool of specialist LLMs (and recursive instances of itself), handling model selection, delegation, verification, and synthesis behind one endpoint. Built on Sakana AI's TRINITY and Conductor research, its routing intelligence is learned in model weights rather than hand-configured.
The higher-quality tier of Sakana AI's Fugu multi-agent orchestration system, tuned for the hardest coding, reasoning, science, and agentic tasks. Coordinates a swappable pool of frontier LLMs through one OpenAI-compatible endpoint, delegating sub-tasks, verifying intermediate work, and synthesizing a single answer. Sakana reports strong vendor benchmarks including 93.2 on LiveCodeBench, 73.7 on SWE-Bench Pro, and 82.1 on TerminalBench.
Sarvam AI's sovereign 105B-parameter Mixture-of-Experts model activating ~9B parameters per token, with a 128K-token context window. Trained on 12 trillion tokens across 22 Indian languages using 128 sparse experts with Multi-head Latent Attention and a custom low-fertility Indic tokenizer. Wins the majority of pairwise comparisons on Indian-language and STEM benchmarks.
Z.ai's (formerly Zhipu AI) flagship open-weight coding model with a 1M-token context window. Mixture-of-Experts architecture with 753B total parameters and ~40B active per request, featuring two cost-balancing reasoning modes. Tops several coding benchmarks while remaining a fraction of the cost of comparable proprietary frontier models. MIT-licensed weights.
Sarvam AI's 24B-parameter instruction-tuned model derived from Mistral-Small-3.1-24B, post-trained on English plus eleven major Indic languages (bn, hi, kn, gu, mr, ml, or, pa, ta, te). Delivers large relative gains on Indian-language, math, and programming benchmarks over its base model, with a hybrid reasoning mode for complex tasks.
Sarvam AI's 30B-parameter Mixture-of-Experts reasoning model trained from scratch with only 2.4B active parameters per token. Optimized for real-time deployment and Indian languages, delivering strong reasoning, coding, and conversational performance while remaining efficient to serve. Open-weights.
Cohere's enterprise flagship model building on Command A with stronger reasoning, agentic tool use, and multilingual performance across 23 languages. Optimized for secure, high-throughput RAG, retrieval, and long-horizon agent workflows in regulated environments, with private and on-premise deployment options.
Moonshot AI's latest open-source, coding-focused model in the Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. A 1-trillion-parameter model that cuts reasoning token usage by roughly 30% versus K2.6 while improving coding and agent performance — +21.8% on Kimi Code Bench v2, +11.0% on Program Bench, and +31.5% on MLS Bench Lite for multi-language support. Released under a Modified MIT License and available via Kimi APIs and Hugging Face.
Google DeepMind's experimental diffusion-based member of the Gemma 4 open model family. Unlike autoregressive models that generate text one token at a time, DiffusionGemma denoises a canvas of placeholder tokens to produce up to 256 tokens in parallel, finalizing output in one block. A Mixture-of-Experts model with 26B total parameters and 3.8B active per inference, delivering roughly 4x the throughput of similarly sized autoregressive Gemma models on local hardware. Excels at non-linear tasks like in-line editing, molecular sequencing, mathematical graphing, and self-correcting puzzles.
Anthropic 首款公开提供的 Mythos 级模型,能力超过该公司此前面向公众发布的所有模型。它在几乎所有已测试的基准中达到领先水平,尤其擅长软件工程、知识工作、视觉理解和科学研究;任务越长、越复杂,优势越明显。内置安全机制会将敏感的网络安全、生物、化学和蒸馏查询路由至 Claude Opus 4.8。
Anthropic's frontier Mythos-class model — the same underlying model as Claude Fable 5 but with safeguards lifted in some areas. It has the strongest cybersecurity capabilities of any model in the world, alongside state-of-the-art performance in software engineering, knowledge work, vision, and scientific research. Access is restricted to a small group of trusted cyberdefenders and infrastructure providers through Project Glasswing.
Google's medium-size open-weight model with 12 billion parameters from the Gemma 4 family. Encoder-free unified multimodal architecture that natively processes text, image, audio, and video inputs without dedicated encoders. Features a 256K context window and supports 140+ languages. First medium-sized model capable of natively ingesting audio. Suitable for local deployment on GPUs with 16GB VRAM.