The high-speed variant of Z.ai's GLM-5.3-Flash, a natively multimodal model delivering inference speeds of up to 200 tokens per second. Built on the same hybrid sparse and linear attention architecture, with image and video input, reasoning, tool use, and a 1M-token context window.
Sortiert nach Gesamtkosten (Ein- und Ausgabe pro 1 Mio. Tokens). Wählen Sie eine Zeile, um Anbieterdetails anzuzeigen.
| Anbieter | Preise (pro 1 Mio.) | Ratenlimits | Regionen | Status | Latenz |
|---|---|---|---|---|---|
Ein: $0.37Aus: $1.25 | 60 RPM / 200K TPM | us-east-1 | Betriebsbereit | 0ms |
Verwenden Sie dieses Modell über OpenRouter mit einem OpenAI-kompatiblen SDK.
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://openrouter.ai/api/v1",
apiKey: process.env.OPENROUTER_API_KEY,
});
const response = await client.chat.completions.create({
model: "z-ai/glm-5.3-flashx",
messages: [
{ role: "user", content: "Hello!" }
],
});
console.log(response.choices[0].message.content);OpenRouter-API · OpenAI-kompatibles SDK
Alle für dieses Modell erfassten Preise je Anbieter. Preise pro 1 Mio. Tokens in USD.
Seit 1. Okt. 2026 wurden keine Preisänderungen erfasst.
| Datum | Anbieter | Eingabe | Ausgabe |
|---|---|---|---|
| 1. Okt. 2026Gelistet · Aktuell | OpenRouter | $0.37 | $1.25 |