Free models on Tokenator
These models are free to use — requests to them do not spend tokens from your bundle. Each one has its own daily allowance that renews every day.
This model is available on a free plan with set usage limits. Claude Opus 5.5 is Anthropic's flagship model for complex logical reasoning, programming, and long-term agent management tasks, replacing Claude Opus 5. It is particularly effective when performing multi-stage changes to large codebases, verifying and finding errors in code, performing financial and scientific analysis, and working with complex graphs, charts, and screenshots. Compared to the previous version, the model handles numerical data and sources more carefully, striving to provide only the information it can confirm. Claude Opus 5.5 completes comparable tasks in fewer steps and with lower token consumption than Opus 5. The model also provides clearer progress reporting: what has been done, what results have been obtained, and what additional information is required from the user. The thinking mode is always adaptive, so the effort parameter allows you to choose a balance between the depth of reasoning, latency, and cost. Low effort levels remain effective for tasks where high response speed is particularly important.
This model is available on a free plan with limited usage. GLM-5.3-Flash is a native multimodal model from Z.ai that demonstrates high performance in programming, agent-based tasks, and working with long-term contexts.
This model is available on a free plan with set usage limits. Gemini 3.8 Flash is the most intelligent model in Google's Flash series, significantly outperforming Gemini 3.7 Flash in programming, agent-based tasks, and complex multi-step reasoning. It combines high performance with improved capabilities for analysis, planning, and execution of complex tasks, while maintaining the efficiency characteristic of the Flash series.
This model is available on the free tier with predefined usage limits. GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon agentic tasks that involve computer and browser use.
This model is available on the free tier with predefined usage limits. MiniMax M3.1 Flash is a fast reasoning model from MiniMax designed for coding, agentic tasks, long-context workloads, and multi-step workflows. The model supports multimodal input, tool calling, and multiple reasoning effort levels, combining high generation speed with strong performance on complex tasks.
5 free models, 5 available right now, they do not spend bundle tokens, daily allowance up to 3.5M tokens
Daily allowances of the free models
When a model runs out of its daily allowance it stops answering until the allowance renews, and your paid tokens stay untouched. The other free models keep working — each one has its own allowance.
| Model | API ID | Vendor | Context | Daily allowance |
|---|---|---|---|---|
| Free GLM 5.3 Flash | free-glm-5.3-flash | Z-AI | 1M tokens | 3.5M tokens and 60 requests per day |
| Free Gemini 3.8 Flash | free-gemini-3.8-flash | Gemini | 1M tokens | 3.5M tokens and 60 requests per day |
| Free GPT-6 Astra | free-gpt-6-astra | OpenAI | 1M tokens | 3M tokens and 45 requests per day |
| Free Claude Opus 5.5 | free-claude-opus-5.5 | Anthropic | 1M tokens | 3.5M tokens and 60 requests per day |
| Free MiniMax M3.1 Flash | free-minimax-m3.1-flash | MiniMax | 1M tokens | 3M tokens and 50 requests per day |