GLM 5.3 Flash via the Tokenator API
What the model is for
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.
Specification
| Full name | GLM 5.3 Flash |
| API ID | glm-5.3-flash |
| Model vendor | Z-AI |
| Type | Text model (chat / completions) |
| Context window | 1M tokens |
| Max output | 131.1K tokens |
| Image input | yes |
| Tokenator usage multiplier | 1.7× |
| Input modalities | text, image, video |
| Aliases | cursor-glm-5.3-flash |
| Supported API formats | /v1/chat/completions /v1/responses /v1/messages |
| Upstream providers | 1 |
| Current status | available |
| Data updated | 2026-08-26 |
Upstream providers
| Provider | Multiplier |
|---|---|
| Unified LLM API | 1.7× |
A request goes to the first available provider by priority; if it fails, Tokenator switches to the next one.
Example request
curl https://api.tokenator.top/v1/chat/completions \ -H "Authorization: Bearer sk-your-tokenator-key" \ -H "Content-Type: application/json" \ -d '{ "model": "glm-5.3-flash", "messages": [{"role": "user", "content": "Hello"}] }'
Compatible tools
The model can be used in any tool that accepts a custom base URL: Claude Code, Codex CLI, OpenCode, Cursor, Cline, Kilo Code, Cherry Studio, OpenClaw.
See also
FAQ about GLM 5.3 Flash
What is the API ID of GLM 5.3 Flash?
glm-5.3-flash — put this into the model field of your request.
What context window does GLM 5.3 Flash have?
1M tokens; max output is 131.1K tokens.
How many tokens will a request cost?
(input + output) × 1.7×. See the documentation for details.
Which tools support this model?
Any tool that accepts a custom base URL: Claude Code, Codex CLI, OpenCode, Cursor, Cline, Kilo Code. Setup guides are in the integrations section.