Tokenator AI model catalog

All LLMs available through Tokenator. Pass the model ID in the model field of your request — Tokenator routes the traffic to the matching upstream provider. The chip next to each model is the billing multiplier: how much more (or less) its tokens cost relative to base.

Anthropiconline
Claude Opus 5
claude-opus-5
1.9×

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis of charts and documents, complex office deliverables, and coordinating parallel subagents. The model maintains strong instruction following and tool use across extended tasks, while remaining effective at lower effort settings for workloads that prioritize latency and token efficiency.

Context 1M tokens
Anthropiconline
Claude Opus 4.8
claude-opus-4-8
1.8×

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token context window.

Context 1M tokens
Anthropiconline
Claude Sonnet 5
claude-sonnet-5
1.7×

Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work.

Context 1M tokens
DeepSeekonline
Deepseek V4 Flash
deepseek-v4-flash
1.1 – 1.3×

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows.

Context 1.05M tokens
Anthropiconline
Claude Opus 4.7
claude-opus-4-7
1.8×

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents.

Context 1M tokens
OpenAIonline
GPT-5.6 Sol
gpt-5.6-sol
1.8×

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks and long-horizon problem solving.

Context 1M tokens
OpenAIonline
GPT 5.5
gpt-5.5
1.8×

GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks.

Context 1.05M tokens
OpenAIonline
GPT-5.6 Luna
gpt-5.6-luna
1.4 – 1.6×

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for its price tier.

Context 1M tokens
MoonshotAIonline
Kimi K2.7 Code
kimi-k2-7-code

MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts.t

Context 262.1K tokens
Minimaxonline
MiniMax M3
minimax-m3
1.5×

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding, and tool use. It is built on MiniMax Sparse Attention (MSA), which replaces full attention with KV-block selection to cut per-token compute at long context — roughly 1/20 the cost of the previous generation at 1M tokens, with substantially faster prefill and decode while retaining quality across most tasks. Trained as a native multimodal model on interleaved data and tuned for multi-turn, production-like collaboration via an interactive user-simulator framework, the model is oriented toward sustained, multi-step tasks rather than single-turn execution.

Context 1M tokens
Anthropiconline
Claude Opus 4.6
claude-opus-4-6
1.8×

Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks.

Context 1M tokens
MoonshotAIonline
Kimi K3
kimi-k3
6 – 10×

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at navigating large repositories, using tools, debugging, and iterating against images, logs, tests, and runtime feedback. Its architecture uses KDA and Attention Residuals for computational efficiency.

Context 1M tokens
Anthropiconline
Claude Sonnet 4.6
claude-sonnet-4-6
1.7×

Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work.

Context 1M tokens
Xiaomionline
MiMo-V2.5-Pro
mimo-v2.5-pro
1.7×

MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and long-horizon tasks, with top rankings on benchmarks such as ClawEval, GDPVal, and SWE-bench Pro. It can independently and autonomously complete professional tasks that would take human experts days or weeks, involving more than a thousand tool calls. Its context length of up to 1M makes it well suited for integration with a wide range of agent frameworks.

Context 1M tokens
xAIonline
Grok 4.5
grok-4.5
1.7×

Grok 4.5 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.

Context 500K tokens
Geminionline
Gemini 3.6 Flash
gemini-3.6-flash
1.7×

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and less hedging, while reducing token use and the number of model calls needed to complete a task.

Context 1.05M tokens
Geminionline
Gemini 3.1 Pro Preview
gemini-3.1-pro-preview
1.8×

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation of the Gemini 3 series, it combines high-precision reasoning across text, image, video, audio, and code with a 1M-token context window. Designed for advanced development and agentic systems, Gemini 3.1 Pro Preview improves long-horizon stability and tool orchestration while increasing token efficiency. It introduces a new medium thinking level to better balance cost, speed, and performance. The model excels in agentic coding, structured planning, multimodal analysis, and workflow automation, making it well-suited for autonomous agents, financial modeling, spreadsheet automation, and high-context enterprise tasks.

Context 1M tokens
Z-AIonline
GLM 5.2
glm-5.2
2 – 3×

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering, and complex multi-step automation.

Context 1M tokens
Geminionline
Gemini 3.5 Flash
gemini-3.5-flash
1.7 – 2.1×

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed.

Context 1M tokens
Anthropiconline
Claude Fable 5
claude-fable-5
10 – 12×

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding.

Context 1M tokens
OpenAIonline
GPT-5.6 Terra
gpt-5.6-terra
1.7×

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic tasks where capability and cost need to be balanced, offering strong performance at roughly half the cost of Sol.

Context 1M tokens
DeepSeekonline
Deepseek V4 Pro
deepseek-v4-pro
1.6×

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window.

Context 1.05M tokens
Anthropiconline
Claude Haiku 4.5
claude-haiku-4-5
1.6×

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models.

Context 200K tokens
Bytedance-seedonline
Doubao Seed 2.0 Code
doubao-seed-2.0-code
1.6×

This is a proprietary model specifically optimized for programming and agent-based tasks. It performs well in code generation, refactoring, analysis, and debugging, and supports tool calling. The model shows strong results on SWE-Bench and LiveCodeBench and is compatible with the OpenAI API through a number of services.

Context 256K tokens
Bytedance-seedonline
Doubao Seed 2.0 Lite
doubao-seed-2.0-lite
1.2×

This is a lightweight, high-performance, general-purpose multimodal model. It is optimized for low latency, high-volume requests, and enterprise workloads, while maintaining support for reasoning, agent scripting, programming, and tooling at a significantly lower cost.

Context 256K tokens
Bytedance-seedonline
Doubao Seed 2.0 Mini
doubao-seed-2.0-mini
1.1×

This is a compact, general-purpose multimodal model optimized for low latency, high throughput, and high-volume workloads. It supports four levels of reasoning intensity and is suitable for chatbots, agent-based tasks, programming, and enterprise applications where speed and cost are important.

Context 256K tokens
Bytedance-seedonline
Doubao Seed 2.0 Pro
doubao-seed-2.0-pro
1.6×

This is a flagship general-purpose multimodal model for complex reasoning, agent-based scenarios, document analysis, programming, and working with long context. It supports tool/function calling, structured inference, and context of up to 256K tokens.

Context 256K tokens
Geminionline
Gemini 3.1 Flash Lite
gemini-3.1-flash-lite
1.5×

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads.

Context 1M tokens
OpenAIonline
GPT 5.4 Mini
gpt-5.4-mini
1.6×

GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads.

Context 400K tokens
OpenAIonline
GPT Image 2
gpt-image-2
1 – 2×

GPT Image 2 combines OpenAI's GPT-5.4 model with state-of-the-art image generation capabilities from GPT Image 2.

Image genEditing
OpenAIonline
GPT-4o mini
gpt-4o-mini
1.5×

GPT-4o mini is OpenAI's newest model after GPT-4 Omni, supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable than other recent frontier models, and more than 60% cheaper than GPT-3.5 Turbo. It maintains SOTA intelligence, while being significantly more cost-effective.

Context 128K tokens
Tencentonline
Hy3
hy3
1.7×

Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort: a direct no-think mode by default, plus low and high chain-of-thought modes for complex math, coding, and multi-step problems. With a 256K context window, Hy3 targets long-horizon tasks, including improved coreference resolution, multi-turn constraint tracking, and stable tool-calling that generalizes across agent scaffoldings. Tencent positions it as a reliable, cost-effective option across coding, document processing, financial analysis, game development, and frontend design, with a strong emphasis on grounded, anti-hallucination behavior that answers when grounded and flags when evidence is missing rather than fabricating.

Context 262.1K tokens
Meituanonline
LongCat 2.0
longcat-2.0
1.4×

LongCat 2.0 is a sparse mixture-of-experts language model from Meituan, with 48B active parameters out of 1.6T total. It is suited for coding, repository-level changes, long-horizon problem solving, and agentic workflows.

Context 1M tokens
Xiaomionline
MiMo-V2.5
mimo-v2.5
1.6×

MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding tasks. Its 1M context window supports complete documents, extended conversations, and complex task contexts in a single pass, making it ideal for integration with agent frameworks where strong reasoning, rich perception, and cost efficiency all matter.

Context 1M tokens
Geminionline
Nano Banana 2
nano-banana-2

Gemini 3.1 Flash Image Preview, a.k.a. Nano Banana 2, is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed.

Image gen
Geminionline
Nano Banana Pro
nano-banana-pro

Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and high-fidelity visual synthesis. The model generates context-rich graphics, from infographics and diagrams to cinematic composites, and can incorporate real-time information via Search grounding.

Image gen
Qwenonline
Qwen 3.7 Max
qwen-3.7-max
2.5×

Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks, and long-horizon autonomous execution. The model offers notable gains in coding and agentic performance over prior Qwen generations and supports explicit prompt caching for efficient repeated context use.

Context 1M tokens
Qwenonline
Qwen 3.7 Plus
qwen-3.7-plus

Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its vision-language abilities while retaining full-stack, agent-level intelligence for coding, tool use, and productivity workflows. Its distinguishing trait is multi-modal interactive hybrid agent capability: it can perceive real-world scenes, read screens and interact with GUIs, generate code from visual references, and perform end-to-end navigation within mobile apps.

Context 1M tokens
Qwenonline
Qwen 3.8 Max
qwen-3-8-max
10×

Qwen 3.8 Max is Alibaba’s next-generation flagship model built for complex coding, analytical, and professional workflows. It excels at full-stack development, data analysis, office automation, and long-running multi-step tasks.

Context 1M tokens
Cursoroffline
Composer 2.5
composer-2.5

An enhanced auto-agent tool in the code editor designed for end-to-end project-level development.

Context 1M tokens
Bytedance-seedoffline
Seedream 4.5
seedream-4.5

Seedream 4.5 is the latest in-house image generation model developed by ByteDance.

Image gen

41 models in the catalog, 39 available right now, 14 model vendors, 4 for image generation, context up to 1.05M tokens

Models by vendor: