模型

浏览来自所有提供商的 9 个标准化 LLM 模型

显示第 1–9 项,共 9 个模型

Sarvam-M

印度

Sarvam AI's 24B-parameter instruction-tuned model derived from Mistral-Small-3.1-24B, post-trained on English plus eleven major Indic languages (bn, hi, kn, gu, mr, ml, or, pa, ta, te). Delivers large relative gains on Indian-language, math, and programming benchmarks over its base model, with a hybrid reasoning mode for complex tasks.

上下文
131K
发布日期
2026年6月

Gemma 4 12B

美国

Google's medium-size open-weight model with 12 billion parameters from the Gemma 4 family. Encoder-free unified multimodal architecture that natively processes text, image, audio, and video inputs without dedicated encoders. Features a 256K context window and supports 140+ languages. First medium-sized model capable of natively ingesting audio. Suitable for local deployment on GPUs with 16GB VRAM.

上下文
262K
发布日期
2026年6月

Falcon-H1

阿拉伯联合酋长国

TII's hybrid Mamba-Transformer model that outperforms comparable offerings from Meta's Llama and Alibaba's Qwen in the 30-70B parameter range. Designed for real-world AI on everyday devices and resource-limited settings with state-of-the-art efficiency.

上下文
131K
发布日期
2026年5月

Granite 4.1 8B

美国

IBM's dense decoder-only 8B parameter language model from the Granite 4.1 family. Supports 131K-token context, tool calling, RAG, code generation with fill-in-the-middle, text summarization, classification, and extraction across 12 languages. Released under Apache 2.0.

上下文
131K
发布日期
2026年5月

GPT-OSS 120B

美国

OpenAI's first open-weight large model with 120 billion parameters. Released under Apache 2.0 license, offering strong performance on reasoning and coding tasks while being fully self-hostable.

上下文
131K
发布日期
2026年4月

Mistral Small 4

法国

Mistral AI's efficient hybrid model unifying instruct, reasoning, and coding in a single model. Open-weight under Apache 2.0 with strong performance for its size class.

上下文
128K
发布日期
2026年3月

Mistral Large 3

法国

Mistral AI's largest open-weight model with 41B active parameters (675B total MoE). State-of-the-art general-purpose multimodal model with 256K context window and powerful agentic capabilities. Released under Apache 2.0.

上下文
256K
发布日期
2025年12月

Trendyol LLM 8B T1

土耳其

Turkish-optimized 8B chat model developed by Trendyol, Turkey's largest e-commerce platform. Built on Qwen3-8B and fine-tuned on large-scale Turkish e-commerce datasets. Features advanced chain-of-thought reasoning in Turkish with dual operation modes (/think and /no_think), strong instruction following, summarization, coding, and attribute extraction for catalogue enrichment. English reasoning capabilities are preserved alongside Turkish.

上下文
33K
发布日期
2025年7月

Qwen3 32B

中国

Alibaba's Qwen3 32B dense language model with strong reasoning and multilingual capabilities, supporting function calling and code generation across diverse tasks.

上下文
131K
发布日期
2025年4月