The high-speed variant of Z.ai's GLM-5.3-Flash, a natively multimodal model delivering inference speeds of up to 200 tokens per second. Built on the same hybrid sparse and linear attention architecture, with image and video input, reasoning, tool use, and a 1M-token context window.
Sorted by total cost (input + output per 1M tokens). Select a row to view provider details.
| Provider | Pricing (per 1M) | Rate limits | Regions | Health | Latency |
|---|---|---|---|---|---|
In: $0.37Out: $1.25 | 60 RPM / 200K TPM | us-east-1 | Healthy | 0ms |
Use this model via OpenRouter with an OpenAI-compatible SDK.
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://openrouter.ai/api/v1",
apiKey: process.env.OPENROUTER_API_KEY,
});
const response = await client.chat.completions.create({
model: "z-ai/glm-5.3-flashx",
messages: [
{ role: "user", content: "Hello!" }
],
});
console.log(response.choices[0].message.content);Using OpenRouter API · OpenAI-compatible SDK
Every price recorded for this model, per provider. Prices are per 1M tokens in USD.
No price changes recorded since Oct 1, 2026.
| Date | Provider | Input | Output |
|---|---|---|---|
| Oct 1, 2026Listed · Current | OpenRouter | $0.37 | $1.25 |