Z-AI models on Tokenator
Every Z-AI model available through Tokenator: API ID, context window, usage multiplier and current availability. The key and the base URL stay the same for every model in the catalog — switching models takes one request parameter.
This model is available on a free plan with limited usage. GLM-5.3-Flash is a native multimodal model from Z.ai that demonstrates high performance in programming, agent-based tasks, and working with long-term contexts.
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.
GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering, and complex multi-step automation.
GLM 5.3 is the latest large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is built for long-horizon agent workflows, project-level software engineering, and complex multi-step automation.
4 Z-AI models in the catalog, 4 available right now, context up to 1M tokens, multiplier ×1.5–×4