模型
浏览来自所有提供商的 12 个标准化 LLM 模型
Gemini 3.5 Flash-Lite
Google's fastest and most cost-effective Gemini 3.5-class model, delivering around 350 output tokens per second per the Artificial Analysis Index. Designed for low-latency and high-throughput agentic workflows such as agentic search and document processing, with configurable thinking levels, built-in computer use, and full multimodal support across a 1M-token context window.
GPT-5.6 Luna
The fast, low-cost tier of OpenAI's GPT-5.6 series, optimized for high-volume, latency-sensitive tasks such as classification, extraction, routing, and lightweight agentic steps. Approaches the larger GPT-5.6 tiers on many benchmarks while running several times faster at a fraction of the price.
MiniMax M3
MiniMax's frontier open-weight model with 1M-token context window, native multimodality (text, image, video), and strong coding capabilities. Built on MiniMax Sparse Attention (MSA) architecture, achieving 59% on SWE-Bench Pro with significantly improved efficiency at long context.
Qwen 3.7 Plus
Alibaba's multimodal variant in the Qwen 3.7 family, optimized for vision understanding and multimodal tasks. Ranked
Gemini 3.5 Flash
Google DeepMind's balanced Gemini 3.5 model that pairs Pro-line reasoning quality with Flash-line latency and cost. Natively multimodal across text, image, audio, and video with a 1M-token context window, configurable thinking levels, and streaming function calling, tuned for high-throughput production workloads.
Gemini 3.1 Flash-Lite
Google's most cost-efficient Gemini model optimized for high-volume, low-latency use cases. Delivers 2.5x faster time to first token versus Gemini 2.5 Flash with full multimodal support. Ideal for agentic tasks, data extraction, translation, and classification.
Gemini 3 Flash
Google's balanced model combining Gemini 3 Pro's reasoning capabilities with the Flash line's latency, efficiency, and cost. Features configurable thinking levels, multimodal function responses, and streaming function calling for complex agentic workflows.
GPT-5.4 Mini
OpenAI's compact reasoning model optimized for coding, computer use, and subagent tasks. Approaches GPT-5.4 performance on several benchmarks while running more than 2x faster.
Qwen 3.6 Plus
Alibaba's proprietary flagship model in the Qwen 3.6 family, targeting enterprise AI workflows with stronger agentic coding capability, visual coding support, and end-to-end enterprise engineering features.
Grok 4.1 Fast
xAI's fast and cost-effective model with 2M token context window. Offers both reasoning and non-reasoning modes at significantly lower pricing than flagship models.