模型

浏览来自所有提供商的 91 个标准化 LLM 模型

部分描述为试点机器翻译内容,尚未经过人工审核。

显示第 73–91 项,共 91 个模型

AlemLLM

哈萨克斯坦

Kazakhstan's flagship Mixture-of-Experts language model developed by Astana Hub with technical support from 01.AI. Features 247B total parameters with 22B active per token, achieving state-of-the-art results on Kazakh, Russian, and English benchmarks. Outperforms GPT-4o on Kazakh language tasks.

上下文
131K
发布日期
2025年8月

Trendyol LLM 8B T1

土耳其

Turkish-optimized 8B chat model developed by Trendyol, Turkey's largest e-commerce platform. Built on Qwen3-8B and fine-tuned on large-scale Turkish e-commerce datasets. Features advanced chain-of-thought reasoning in Turkish with dual operation modes (/think and /no_think), strong instruction following, summarization, coding, and attribute extraction for catalogue enrichment. English reasoning capabilities are preserved alongside Turkish.

上下文
33K
发布日期
2025年7月

Gemini 2.5 Flash

美国

Google's cost-effective model optimized for high throughput tasks. Balances speed and intelligence with strong multimodal capabilities and 1M token context window.

上下文
1.0M
发布日期
2025年6月

Gemini 2.5 Pro

美国

Google's high-capability reasoning model with adaptive thinking for complex agentic and multimodal challenges. Features 1M token context window and strong performance on coding and scientific tasks.

上下文
1.0M
发布日期
2025年6月

GPT-5

美国

OpenAI's fifth-generation flagship model with significant improvements in reasoning, multimodal understanding, and code generation. Features enhanced instruction following and expanded context window.

上下文
256K
发布日期
2025年6月

Nemotron Nano 9B v2

美国

NVIDIA's compact 9B parameter model trained from scratch for both reasoning and non-reasoning tasks. Generates reasoning traces before final responses. Efficient for edge and on-device deployment.

上下文
131K
发布日期
2025年6月

WiroAI Turkish LLM 9B

土耳其

Turkish-specialized 9B language model developed by WiroAI, built on Google's Gemma 2 architecture. Fine-tuned with Supervised Fine-Tuning (SFT) on over 500,000 carefully curated high-quality Turkish instructions, specifically adapted to Turkish culture and local context. Demonstrates superior performance on Turkish language processing tasks including conversation, reasoning, and instruction following.

上下文
8K
发布日期
2025年4月

Llama 4 Maverick

美国

Meta's quality-focused MoE model with 17B active parameters (400B total, 128 experts). Targets quality-critical tasks with benchmark scores competitive with GPT-4o and Gemini 2.5 Pro.

上下文
1.0M
发布日期
2025年4月

Qwen3 235B

中国

Alibaba's Qwen3 235B mixture-of-experts model delivering frontier-level performance with advanced reasoning, function calling, and code generation capabilities at massive scale.

上下文
131K
发布日期
2025年4月

Qwen3 32B

中国

Alibaba's Qwen3 32B dense language model with strong reasoning and multilingual capabilities, supporting function calling and code generation across diverse tasks.

上下文
131K
发布日期
2025年4月

Gemma 3 12B

美国

Google's mid-size open-weight model with 12 billion parameters from the Gemma 3 family. Supports multimodal inputs including text and images with a 128K context window. Strong performance on reasoning and code generation tasks at moderate compute cost.

上下文
131K
发布日期
2025年3月

Gemma 3 27B

美国

Google's largest open-weight model in the Gemma 3 family with 27 billion parameters. Supports multimodal inputs including text and images with a 128K context window. Delivers strong performance across reasoning, code generation, and vision tasks, competitive with larger proprietary models.

上下文
131K
发布日期
2025年3月

Command A

美国

Cohere's flagship 111B parameter model optimized for demanding enterprises requiring fast, secure, and high-quality AI. Excels at RAG, tool use, and multilingual tasks with strong reasoning capabilities.

上下文
256K
发布日期
2025年3月

QwQ 32B

中国

Alibaba's QwQ 32B reasoning-focused model designed for complex problem solving, mathematical reasoning, and step-by-step logical analysis with strong chain-of-thought capabilities.

上下文
131K
发布日期
2025年3月

DeepSeek R1

中国

DeepSeek 面向推理的模型,通过强化学习训练以处理复杂的多步骤推理任务。尤其擅长需要思维链推理的数学、科学和编程问题。

上下文
131K
发布日期
2025年1月

DeepSeek V3

中国

DeepSeek's third-generation large language model featuring mixture-of-experts architecture, strong multilingual capabilities, and competitive performance on reasoning and coding benchmarks.

上下文
128K
发布日期
2024年12月

Phi-4

美国

Microsoft's Phi-4 model with 14B parameters excelling at reasoning and code generation tasks, delivering strong performance relative to its compact size with efficient inference characteristics.

上下文
16K
发布日期
2024年12月

Aya Expanse 32B

美国

Highly performant 32B multilingual language model from Cohere For AI, designed to rival monolingual model performance across 23 languages. Built using innovations in multilingual data arbitrage, direct preference optimization, and model merging techniques. Outperforms previous multilingual models on both automatic and human evaluations.

上下文
8K
发布日期
2024年10月

Claude 3 Opus

美国

Anthropic's most powerful model in the Claude 3 family, excelling at complex analysis, nuanced content generation, scientific reasoning, and code generation with extended context support.

上下文
200K
发布日期
2024年3月