Tokenator AI model catalog

All LLMs available through Tokenator. Pass the model ID in the model field of your request — Tokenator routes the traffic to the matching upstream provider. The chip next to each model is the billing multiplier: how much more (or less) its tokens cost relative to base.

Anthropiconline
Claude Opus 5.5
claude-opus-5.5
×1.7–×1.8

Claude Opus 5.5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work, succeeding Claude Opus 5. It is particularly strong at multi-step changes in large codebases, code review and bug finding, financial and scientific analysis, and reading dense charts, diagrams, and screenshots, and it is more careful than its predecessor about only stating figures and citing sources it can back up. The model completes comparable tasks in fewer steps and with fewer tokens than Opus 5, and reports on its work in plainer language, with clear updates on what it did, what it found, and what it needs from the user. Thinking is always adaptive, so effort is the main lever for trading off depth, latency, and cost, and lower effort settings remain effective for latency-sensitive workloads.

Context 1M tokens
OpenAIonline
GPT-6 AstraBEST
gpt-6-astra
×0.8

GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon agentic tasks that involve computer and browser use.

Context 1M tokens
Anthropiconline
Free Claude Opus 5.5
free-claude-opus-5.5
×1.9

This model is available on a free plan with set usage limits. Claude Opus 5.5 is Anthropic's flagship model for complex logical reasoning, programming, and long-term agent management tasks, replacing Claude Opus 5. It is particularly effective when performing multi-stage changes to large codebases, verifying and finding errors in code, performing financial and scientific analysis, and working with complex graphs, charts, and screenshots. Compared to the previous version, the model handles numerical data and sources more carefully, striving to provide only the information it can confirm. Claude Opus 5.5 completes comparable tasks in fewer steps and with lower token consumption than Opus 5. The model also provides clearer progress reporting: what has been done, what results have been obtained, and what additional information is required from the user. The thinking mode is always adaptive, so the effort parameter allows you to choose a balance between the depth of reasoning, latency, and cost. Low effort levels remain effective for tasks where high response speed is particularly important.

FreeContext 1M tokens
Z-AIonline
Free GLM 5.3 Flash
free-glm-5.3-flash
×1.9

This model is available on a free plan with limited usage. GLM-5.3-Flash is a native multimodal model from Z.ai that demonstrates high performance in programming, agent-based tasks, and working with long-term contexts.

FreeContext 1M tokens
Geminionline
Free Gemini 3.8 Flash
free-gemini-3.8-flash
×1.7

This model is available on a free plan with set usage limits. Gemini 3.8 Flash is the most intelligent model in Google's Flash series, significantly outperforming Gemini 3.7 Flash in programming, agent-based tasks, and complex multi-step reasoning. It combines high performance with improved capabilities for analysis, planning, and execution of complex tasks, while maintaining the efficiency characteristic of the Flash series.

FreeContext 1M tokens
MoonshotAIonline
Kimi K3
kimi-k3
×1.8

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at navigating large repositories, using tools, debugging, and iterating against images, logs, tests, and runtime feedback. Its architecture uses KDA and Attention Residuals for computational efficiency.

Context 1M tokens
OpenAIonline
Free GPT-6 Astra
free-gpt-6-astra
×1.2

This model is available on the free tier with predefined usage limits. GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon agentic tasks that involve computer and browser use.

FreeContext 1M tokens
DeepSeekonline
Deepseek V4 Pro
deepseek-v4-pro
×1.8–×2

DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.

Context 1.05M tokens
MiniMaxonline
Free MiniMax M3.1 Flash
free-minimax-m3.1-flash
×1.5

This model is available on the free tier with predefined usage limits. MiniMax M3.1 Flash is a fast reasoning model from MiniMax designed for coding, agentic tasks, long-context workloads, and multi-step workflows. The model supports multimodal input, tool calling, and multiple reasoning effort levels, combining high generation speed with strong performance on complex tasks.

FreeContext 1M tokens
OpenAIonline
GPT-6.1 Sol
gpt-6.1-sol
×1.3–×1.6

GPT-6.1 Sol is an upgrade to GPT-6 Sol from OpenAI, positioned below the flagship GPT-6 Astra in the GPT-6 series. It is suited for agentic coding, computer use, document-heavy professional work, and multi-step business workflow automation, and approaches Astra-level results on these tasks at a much lower cost. Compared with GPT-6 Sol, it makes fewer factual errors and is more reliable at respecting explicit restrictions and user intent during agentic tasks.

Context 1.05M tokens
Anthropiconline
Claude Sonnet 5
claude-sonnet-5
×1.7

Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work.

Context 1M tokens
Anthropiconline
Claude Opus 5
claude-opus-5
×1.9

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis of charts and documents, complex office deliverables, and coordinating parallel subagents. The model maintains strong instruction following and tool use across extended tasks, while remaining effective at lower effort settings for workloads that prioritize latency and token efficiency.

Context 1M tokens
OpenAIonline
GPT-6 Sol
gpt-6-sol
×1.6–×1.7

GPT-6 Sol is the cost-efficient high-end model in OpenAI's GPT-6 series, positioned below the flagship GPT-6 Astra and above the fast GPT-6 Luna tier. It is suited for demanding professional work, agentic coding, business workflow automation, and computer use, and is particularly strong at long-horizon software engineering tasks in real codebases. It approaches Astra-level factual reliability at a much lower cost and shares Astra's clearer, more concise communication style in technical and coding conversations.

Context 1.05M tokens
OpenAIonline
GPT 5.5
gpt-5.5
×1.8

GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks.

Context 1.05M tokens
MiniMaxonline
MiniMax M3.1 Flash
minimax-m3.1-flash
×1.3

MiniMax M3.1 Flash is a fast reasoning model from MiniMax optimized for coding, agentic workflows, and everyday software development. It supports up to a 1M-token context window, text, image and video inputs, multi-step tasks, and tool calling. The model is designed for high generation speed and works well for bug fixes, feature implementation, large codebase analysis, testing, and development automation. It supports five reasoning effort levels: low, medium, high, xhigh, and max.

Context 1M tokens
Geminionline
Gemini 3.8 Flash
gemini-3.8-flash
×1.4–×1.7

Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.

Context 1M tokens
Anthropiconline
Claude Sonnet 5.5
claude-sonnet-5.5
×1.4–×1.6

Claude Sonnet 5.5 is Anthropic's Sonnet-class model for well-scoped everyday work, succeeding Claude Sonnet 5 as a direct upgrade. It is especially strong at building features, fixing bugs, and producing polished documents, slides, and spreadsheets, and it writes and communicates more clearly than its predecessor. Thinking is always on, so effort is the main lever for trading off depth, latency, and cost, and lower effort settings keep it responsive for everyday agentic loops.

Context 1M tokens
Z-AIonline
GLM 5.3
glm-5.3
×2–×4

GLM 5.3 is the latest large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is built for long-horizon agent workflows, project-level software engineering, and complex multi-step automation.

Context 1M tokens
DeepSeekonline
Deepseek V4.1 Flash
deepseek-v4.1-flash
×1.7

DeepSeek V4.1 Flash is a new high-performance model from DeepSeek that significantly outperforms V4 Pro across key metrics, including task quality, generation speed, cost efficiency, and overall execution time. The model is designed for coding, complex reasoning, agentic workflows, and multi-step tasks, delivering Pro-level capabilities with higher speed and significantly better efficiency.

Context 1.05M tokens
Anthropiconline
Claude Fable 5
claude-fable-5
×10

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding.

Context 1M tokens
Anthropiconline
Claude Fable 5.1
claude-fable-5.1
×13

Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual code generation, and finance and analysis tasks in particular. It also tends to be more concise than Fable 5 in its plans and summaries. We recommend testing it as a direct upgrade wherever you use Fable 5 today, and alongside Opus 5 on reasoning-heavy tasks.

Context 1M tokens
Anthropiconline
Claude Haiku 4.5
claude-haiku-4-5
×1.6

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models.

Context 200K tokens
Anthropiconline
Claude Haiku 5.5NEW
claude-haiku-5.5
×1.0

Claude Haiku 5.5 is Anthropic's small, fast model for high-volume, cost-sensitive work such as summarization, subagents, and browser use. It succeeds Claude Haiku 4.5 with stronger coding, computer use, and knowledge work, and accepts text and image input with a 1M-token context window. It is the first Haiku model with adjustable effort. Thinking is adaptive and on by default, with effort as the main lever for trading off depth, latency, and cost; it can be turned off at low, medium, and high effort.

Context 1M tokens
Anthropiconline
Claude Opus 4.6
claude-opus-4-6
×1.8

Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks.

Context 1M tokens
Anthropiconline
Claude Opus 4.7
claude-opus-4-7
×1.8

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents.

Context 1M tokens
Anthropiconline
Claude Opus 4.8
claude-opus-4-8
×1.8

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token context window.

Context 1M tokens
Anthropiconline
Claude Sonnet 4.6
claude-sonnet-4-6
×1.7

Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work.

Context 1M tokens
LiquidAIonline
D1NEW
d1
×1.0

D1 is Liquid AI's structured decision model, served as a System One endpoint. Send a state along with typed questions, and it returns a choice, a score, or a yes/no answer, each with a probability taken directly from the model rather than written out as text.

DecisionsContext 65.5K tokens
Geminionline
Gemini 3.1 Flash Lite
gemini-3.1-flash-lite
×1.5

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads.

UnstableContext 1M tokens
Geminionline
Gemini 3.1 Pro Preview
gemini-3.1-pro-preview
×1.7–×1.8

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation of the Gemini 3 series, it combines high-precision reasoning across text, image, video, audio, and code with a 1M-token context window. Designed for advanced development and agentic systems, Gemini 3.1 Pro Preview improves long-horizon stability and tool orchestration while increasing token efficiency. It introduces a new medium thinking level to better balance cost, speed, and performance. The model excels in agentic coding, structured planning, multimodal analysis, and workflow automation, making it well-suited for autonomous agents, financial modeling, spreadsheet automation, and high-context enterprise tasks.

Context 1M tokens
Geminionline
Gemini 3.6 Flash
gemini-3.6-flash
×1.4–×1.7

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and less hedging, while reducing token use and the number of model calls needed to complete a task.

Context 1.05M tokens
Geminionline
Gemini 3.7 Flash
gemini-3.7-flash
×1.4–×1.7

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step problem solving.

Context 1M tokens
OpenAIonline
Gemini Embedding 001
gemini-embedding-001
×1.35

gemini-embedding-001 provides a unified cutting edge experience across domains, including science, legal, finance, and coding. This embedding model has consistently held a top spot on the Massive Text Embedding Benchmark (MTEB) Multilingual leaderboard since the experimental launch in March.

EmbeddingsContext 20K tokens
Geminionline
Gemini Embedding 2
gemini-embedding-2
×1.4

Gemini Embedding 2 is Google's first multimodal embedding model. We currently support mapping text and images into a unified vector space for semantic search and retrieval-augmented generation (RAG). It supports input context up to 8,192 tokens and flexible output dimensions from 128 to 3,072 (recommended: 768, 1536, or 3,072). Designed for cross-modal similarity — you can embed a text query and retrieve the most relevant images, or vice versa — making it well-suited for multimodal search, recommendation, and document understanding pipelines.

EmbeddingsContext 8.2K tokens
Z-AIonline
GLM 5.2
glm-5.2
×2–×3

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering, and complex multi-step automation.

Context 1M tokens
Z-AIonline
GLM 5.3 Flash
glm-5.3-flash
×1.5–×1.7

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.

Context 1M tokens
Z-AIonline
GLM 5.3 FlashX
glm-5.3-flashx
×1.6

GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture (320B total parameters, 18B active), it is suited for efficient coding, visual understanding, and long-horizon agent tasks with a 1M-token context window.

Context 1.05M tokens
OpenAIonline
GPT Image 2
gpt-image-2
×1

GPT Image 2 combines OpenAI's GPT-5.4 model with state-of-the-art image generation capabilities from GPT Image 2.

Image genEditing
OpenAIonline
GPT Image 2.5
gpt-image-2.5
×1

GPT Image 2.5 is an image generation and editing model from OpenAI, positioned as the precision-oriented tier of the GPT Image 2.5 series. It is suited to detailed creative work where editing accuracy matters more than generation speed, via the dedicated Images API.

Image genEditing
OpenAIonline
GPT-4o mini
gpt-4o-mini
×1.5

GPT-4o mini is OpenAI's newest model after GPT-4 Omni, supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable than other recent frontier models, and more than 60% cheaper than GPT-3.5 Turbo. It maintains SOTA intelligence, while being significantly more cost-effective.

Context 128K tokens
OpenAIonline
GPT-5.6 Luna
gpt-5.6-luna
×1.6

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for its price tier.

Context 1M tokens
OpenAIonline
GPT-5.6 Sol
gpt-5.6-sol
×1.8–×1.9

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks and long-horizon problem solving.

Context 1M tokens
OpenAIonline
GPT-5.6 Terra
gpt-5.6-terra
×1.7–×1.8

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic tasks where capability and cost need to be balanced, offering strong performance at roughly half the cost of Sol.

Context 1M tokens
OpenAIonline
GPT-6 Luna
gpt-6-luna
×1.2–×1.55

GPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol. It is suited for high-volume and latency-sensitive workloads such as chat, classification, and lightweight agentic tasks, and at higher reasoning effort it can take on complex software engineering and computer-use tasks that previously called for a Sol-tier model. It shares the GPT-6 family's gains in factual reliability and its clearer, more concise communication style.

Context 1.05M tokens
OpenAIonline
GPT-6 Luna DecisionsNEW
gpt-6-luna-decisions
×1.1

GPT-6 Luna Decisions is GPT-6 Luna served through OpenAI's Decisions API. Instead of generating text, it reads the content passed as state (text, JSON, or images) and returns typed, probabilistic answers to named questions in a single request: the probability of yes for a yes/no question (noul), a probability for every option plus the most likely one (choice), or a probability for every level of an ordered rubric plus the expected score (score). It is built for classification, routing, moderation, and rubric grading where application code thresholds the returned numbers rather than parsing a chat reply. A request can carry up to 200 questions about the same content.

DecisionsContext 1.05M tokens
xAIonline
Grok 4.5
grok-4.5
×1.7

Grok 4.5 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.

Context 500K tokens
xAIonline
Grok 4.6
grok-4.6
×1.7

Grok 4.6 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.

Context 500K tokens
xAIonline
Grok 4.7
grok-4.7
×1.4

Grok 4.7 is SpaceXAI's flagship model for coding, complex reasoning, and agentic tasks. It features improved long-horizon performance, self-verification, and long-context handling, with support for image understanding, tool calling, and configurable reasoning effort.

Context 500K tokens
xAIonline
Grok Build 0.1
grok-build-0.1
×1.6

Grok Build 0.1 is SpaceXAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with text output, and is optimized for interactive coding agents, tool use, and multi-step development tasks. The model powers SpaceXAI’s Grok Build CLI and features a 256K context window with no text output limit, making it well suited for long-horizon coding and automation workflows. Currently in early access.

Context 256K tokens
xAIonline
Grok Imagine Image 2.0
grok-imagine-image-2.0
×3

Grok Imagine Image 2.0 is an image generation and editing model from xAI. It is suited for creating images from text prompts and editing images from references, with low and medium quality modes.

Image genEditing
xAIonline
Grok Imagine Video 1.5
grok-imagine-video-1.5
×20–×72

Grok Imagine Video 1.5 is a video generation model from SpaceXAI. It creates videos from text prompts, with an optional starting image to guide the scene. It can direct subject and camera motion, pacing, atmosphere, and physical behavior while maintaining visual continuity, and can generate synchronized sound effects, ambience, and dialogue.

Video gen1080pWith audioWith a reference
Tencentonline
Hy-MT2-1.8B
hy-mt2-1.8b
×1.4

Hy-MT2-1.8B is a compact 1.8B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glossary-based, and style-guided translation.

Context 8.2K tokens
Tencentonline
Hy-MT2-30B-A3B
hy-mt2-30b-a3b
×1.8

Hy-MT2-30B-A3B is Tencent's flagship translation model in the Hy-MT2 family. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glossary-based, and style-guided translation. It uses 3B active parameters out of 30B total.

Context 8.2K tokens
Tencentonline
Hy-MT2-7B
hy-mt2-7b
×1.8

Hy-MT2-7B is a 7B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glossary-based, and style-guided translation.

Context 8.2K tokens
Tencentonline
Hy3
hy3
×1.7–×2.1

Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort: a direct no-think mode by default, plus low and high chain-of-thought modes for complex math, coding, and multi-step problems. With a 256K context window, Hy3 targets long-horizon tasks, including improved coreference resolution, multi-turn constraint tracking, and stable tool-calling that generalizes across agent scaffoldings. Tencent positions it as a reliable, cost-effective option across coding, document processing, financial analysis, game development, and frontend design, with a strong emphasis on grounded, anti-hallucination behavior that answers when grounded and flags when evidence is missing rather than fabricating.

Context 262.1K tokens
Tencentonline
Hy4 preview
hy4-preview
×1.5

Tencent: Hy4 preview is a mixture-of-experts model from Tencent, with 49B active parameters out of 770B total. It is designed for coding agents, complex tool-use workflows, and productivity tasks that require planning, context continuity, and sustained multi-step execution.

Context 1M tokens
TypeSafeonline
Jev Latest
jev-latest
×1.1

Jev is a structured decision model from TypeSafe, and the first of its System One models. System One models make fast, structured decisions for software, returning a typed choice rather than free-form text. It is suited for routing, classification, and other decision points inside an application where a fast, predictable answer matters more than generated prose. Learn more in TypeSafe's docs: https://docs.typesafe.ai/concepts/system-one

DecisionsContext 32K tokens
MoonshotAIonline
Kimi K2.7 Code
kimi-k2-7-code
×1.4

MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts.t

Context 262.1K tokens
Klingonline
Kling 3.0 Pro
kling-3.0-pro
×45

Kling 3.0 Pro is an enhanced version of Kling 3.0 focused on higher visual quality, detail, and motion consistency. It handles complex scenes, camera movement, characters, and frame-to-frame coherence more reliably, making it suitable for demanding and cinematic video generation tasks.

Video gen1080pWith audioYour own frames
Meituanonline
LongCat 2.0
longcat-2.0
×4

LongCat 2.0 is a sparse mixture-of-experts language model from Meituan, with 48B active parameters out of 1.6T total. It is suited for coding, repository-level changes, long-horizon problem solving, and agentic workflows.

Context 1M tokens
Inceptiononline
Mercury 2.5
mercury-2.5
×1.3

Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving 1,107 tokens/sec on standard GPUs. It delivers a 10+ point jump in intelligence over Mercury 2, comparable quality to cost-optimized frontier models like GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5. Mercury 2.5 supports tunable reasoning levels, parallel tool calls, and schema-aligned JSON output. It's built for production workloads where latency compounds: search agents, voice pipelines, and coding subagents.

Context 260K tokens
Xiaomionline
MiMo-V2.6-Flash
mimo-v2.6-flash
×1.5

MiMo-V2.6-Flash is an open-source foundation model developed by Xiaomi. Built on a Mixture-of-Experts architecture with 309B total parameters and 15B activated per token, it employs a hybrid attention mechanism for greater computational efficiency. The model features a 1M-token context window and native multimodal capabilities. Optimized for agentic workflows, it delivers strong performance across coding, visual, general, and research scenarios, excelling at complex, long-horizon tasks with robust generalization across a diverse range of agent harnesses.

Context 1.05M tokens
Xiaomionline
MiMo-V2.6-Pro
mimo-v2.6-pro
×1.6–×1.7

MiMo-V2.6-Pro is the flagship foundation model developed by Xiaomi. Built at a scale of over 1T parameters, it is designed to push the ceiling of capability for the most demanding workloads. The model features a 1M-token context window and native multimodal capabilities. Optimized for agentic workflows, it delivers top-tier performance across coding, visual, general, and research scenarios, excelling at complex, long-horizon tasks with robust generalization across a diverse range of agent harnesses.

Context 1.05M tokens
MiniMaxonline
MiniMax M3
minimax-m3
×1.1–×1.5

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding, and tool use. It is built on MiniMax Sparse Attention (MSA), which replaces full attention with KV-block selection to cut per-token compute at long context — roughly 1/20 the cost of the previous generation at 1M tokens, with substantially faster prefill and decode while retaining quality across most tasks. Trained as a native multimodal model on interleaved data and tuned for multi-turn, production-like collaboration via an interactive user-simulator framework, the model is oriented toward sustained, multi-step tasks rather than single-turn execution.

Context 1M tokens
Metaonline
Muse Image
muse-image
×1

Muse Image is an image generation model from Meta that generates and edits images from text and reference images. It supports text-to-image generation, targeted image editing, multi-image composition, reference-image conditioning for style and subject consistency, and precise text rendering within generated images. Iterative editing works by passing the previous output image back with a new instruction

Image gen
Metaonline
Muse Spark 1.2
muse-spark-1.2
×9

Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, audio, and PDF documents, returns text, and offers a 1M-token context window. The model is built to support multi-agent workflows, whether as either a main agent that plans and delegates or as a subagent executing in parallel. It works across multiple coding harnesses and supports structured output, parallel function calling, and configurable reasoning effort. In Meta’s testing, it performs well on multi-file refactors, extended debugging sessions, whole-repository generation, and tasks that stretch well past a single prompt.

Context 1.05M tokens
Metaonline
Muse Spark 1.3
muse-spark-1.3
×9

Muse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows. It is designed to keep track of information across extended tasks, work through conflicting inputs, and request clarification or confirmation when needed, with an emphasis on concise execution.

Context 1M tokens
Geminionline
Nano Banana 2
nano-banana-2
×1

Gemini 3.1 Flash Image Preview, a.k.a. Nano Banana 2, is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed.

Image gen
Geminionline
Nano Banana 2.1NEW
nano-banana-2.1
×2

Nano Banana 2.1 (Gemini Nano Banana 2.1) is Google's image generation and editing model on the Flash tier, succeeding Nano Banana 2 and Nano Banana Pro. It improves product recontextualization, mask- and ink-based editing, and factual accuracy, and renders photorealistic skin tones, detailed materials, lighting, and coherent backgrounds. It accepts text and image inputs, returns images with optional text, and supports 1K and 2K.

Image gen
Geminionline
Nano Banana Pro
nano-banana-pro
×2

Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and high-fidelity visual synthesis. The model generates context-rich graphics, from infographics and diagrams to cinematic composites, and can incorporate real-time information via Search grounding.

Image gen
Geminionline
Omni
omni
×10

Omni is a multimodal video generation and editing model from Google DeepMind. It supports video creation from text, images, audio, and video references, as well as multi-turn editing through natural language. Omni is particularly strong at preserving characters, scene details, and visual style across edits, while using its understanding of physics and real-world context to produce coherent motion and scenes. It also supports first and last frame control, reference-to-video editing, and native audio generation.

Video gen720pWith audioWith a referenceYour own frames
Unbiasedonline
Pareto
pareto
×10

Pareto is a multimodal composite model built for research, coding, and agentic workflows, while delivering frontier-level performance across a broad range of general-purpose tasks.

Context 262.1K tokens
Unbiasedonline
Pareto 26.10 PreviewNEW
pareto-26.10-preview
×7

Pareto is a multimodal composite model built for research, coding, and agentic workflows, while delivering frontier-level performance across a broad range of general-purpose tasks. This is a preview of the next Pareto version and may change without notice; use pareto-26.9 for stable behaviour.

Context 1.05M tokens
Qwenonline
Qwen 3.7 Max
qwen-3.7-max
×2.5–×3

Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks, and long-horizon autonomous execution. The model offers notable gains in coding and agentic performance over prior Qwen generations and supports explicit prompt caching for efficient repeated context use.

Context 1M tokens
Qwenonline
Qwen 3.7 Plus
qwen-3.7-plus
×2–×4

Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its vision-language abilities while retaining full-stack, agent-level intelligence for coding, tool use, and productivity workflows. Its distinguishing trait is multi-modal interactive hybrid agent capability: it can perceive real-world scenes, read screens and interact with GUIs, generate code from visual references, and perform end-to-end navigation within mobile apps.

Context 1M tokens
Qwenonline
Qwen 3.8 Flash
qwen-3.8-flash
×1.4

Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.

Context 1M tokens
Qwenonline
Qwen 3.8 Max
qwen-3-8-max
×7–×10

Qwen 3.8 Max is Alibaba’s next-generation flagship model built for complex coding, analytical, and professional workflows. It excels at full-stack development, data analysis, office automation, and long-running multi-step tasks.

Context 1M tokens
ByteDanceonline
Seedance 2.0
seedance-2.0
×70

Seedance 2.0 is a video generation model from ByteDance. It supports text-to-video, image-to-video with first and last frame control, and multimodal reference-to-video. It is particularly strong at preserving character consistency, visual style, and camera movement from reference material.

Video gen720pWith audioWith a reference
ByteDanceonline
Seedance 2.0 Fast
seedance-2.0-fast
×50

Seedance 2.0 Fast is a video generation model from ByteDance. It supports text-to-video, image-to-video with first and last frame control, and multimodal reference-to-video. It prioritizes generation speed and lower cost over maximum output quality.

Video gen720pWith audioWith a reference
ByteDanceonline
Seedance 2.0 Mini
seedance-2.0-mini
×35

Seedance 2.0 Mini is a video generation model from ByteDance. It supports text-to-video, image-to-video with first and last frame control, and multimodal reference-to-video with image, video, and audio inputs. It supports 480p and 720p output for 4-15 second videos.

Video gen720pWith audioWith a reference
ByteDanceonline
Seedance 2.5
seedance-2.5
×80

Seedance 2.5 is a video generation model from ByteDance. It is suited for long-form storytelling, multimodal reference-based generation, video editing, and video extension. It supports first-frame and first-and-last-frame control.

Video gen720pWith audioWith a reference
ByteDance Seedonline
Seedream 5.0 FlashNEW
seedream-5-0-flash
×2

Seedream 5.0 Flash is an image generation and editing model from ByteDance Seed. It is the fast, cost-efficient tier of the Seedream 5.0 family, suited for high-volume production and interactive editing workflows that need precise edits at low latency.

Image genEditing
ByteDance Seedonline
Seedream 5.0 Pro
seedream-5.0-pro
×3

Seedream 5.0 Pro is an image generation and editing model from ByteDance Seed. It is suited for commercial visual-production workflows that require precise editing control, lifelike scenes, and natural rendering.

Image genEditing
Upstageonline
Solar Decide
solar-decide
×1.1

Solar Decide is Upstage's structured decision model, served as a System One endpoint on Solar Mini 4. Send a state along with typed questions, and it returns a choice, a score, or a yes/no answer, each with a calibrated probability taken directly from the model rather than written out as text. Because it generates no prose, each decision takes a single forward pass and output tokens are free. With a 512K context window, an entire document can serve as the state. Solar Decide uses the same /v1/systemone schema as Jev, bringing Solar Mini 4's strong Korean understanding to routing, classification, and policy checks.

DecisionsContext 524.3K tokens
Upstageonline
Solar Decide FlashNEW
solar-decide-flash
×1.1

Solar Decide Flash is Upstage's low-latency structured decision model, a faster variant of Solar Decide built on Solar Mini 4 and served through the System One (/v1/systemone) API. Instead of generating text, it reads a state and answers typed questions, returning a choice, a score, or a yes/no answer, each with a probability taken directly from the model. Each decision is a single forward pass, and output tokens are free. It is tuned for consistently fast response times in routing, classification, and policy checks. It keeps the 512K context window, so a full document can serve as the state, and carries Solar Mini 4's Korean-language strength.

DecisionsContext 524.3K tokens
Upstageonline
Solar Mini 4
solar-mini-4
×1.4

Solar Mini 4 is Upstage's compact, cost-efficient language model, a 35B-parameter mixture-of-experts with 3B active parameters and a 524K context window. It is built for agentic use cases where response speed and cost matter, with fluent Korean alongside strong English and Japanese, and support for long-context workloads.

Context 524.3K tokens
Upstageonline
Solar Pro 4
solar-pro-4
×1.4

Solar Pro 4 is Upstage's cost-efficient large language model, featuring a 524K context window. It is built for long-horizon tasks and agentic workflows, with strong capabilities in office productivity, document-intensive work, and coding.

Context 524.3K tokens
Respanonline
Span-01
span-01
×1.0

Span-01 is a behavior scoring model from Respan. It reads a conversation span and returns, for each plain-language behavior you define, the probability that the behavior is present. It is suited for evaluation, guardrails, and monitoring of LLM and agent outputs at scale. It is the higher-accuracy tier of the family. Span-01 Lite is the free, lighter tier.

DecisionsContext 32K tokens
PrismMLonline
Ternary Bonsai 2 27B
ternary-bonsai-2-27b
×1.5

Bonsai 2 27B is a 27B-parameter reasoning model from PrismML derived from Qwen3.8-27B. It supports coding, mathematics, tool calling, and image understanding with a 262K-token context window. Ternary compression shrinks the language-model weights to roughly 8.5 GB while retaining 98.2% of the base model's average score across PrismML's 14 thinking-mode benchmarks, enabling efficient inference on consumer hardware. The model thinks by default and defaults to xhigh reasoning effort.

Context 262.1K tokens
OpenAIonline
Text Embedding 3 Large
text-embedding-3-large
×1.2

text-embedding-3-large is OpenAI's most capable embedding model for both english and non-english tasks. Embeddings are a numerical representation of text that can be used to measure the relatedness between two pieces of text. Embeddings are useful for search, clustering, recommendations, anomaly detection, and classification tasks.

EmbeddingsContext 8.2K tokens
OpenAIonline
Text Embedding 3 Small
text-embedding-3-small
×1.0

text-embedding-3-small is OpenAI's improved, more performant version of the ada embedding model. Embeddings are a numerical representation of text that can be used to measure the relatedness between two pieces of text. Embeddings are useful for search, clustering, recommendations, anomaly detection, and classification tasks.

EmbeddingsContext 8.2K tokens
InclusionAIoffline
Ling 3.1 FlashNEW
ling-3.1-flash

Ling 3.1 Flash is a hybrid reasoning mixture-of-experts model from inclusionAI, with 25B active parameters out of 560B total.

Context 262.1K tokens
Cursoroffline
Composer 2.5
composer-2.5

An enhanced auto-agent tool in the code editor designed for end-to-end project-level development.

Context 1M tokens
MiniMaxoffline
H3
hailuo-3

MiniMax H3 — это облегченная модель генерации видео с открытыми весами от MiniMax. Она разработана для точного многомодального редактирования и контролируемой генерации контента, включая редактирование по инструкциям, рендеринг текста и брендов, а также перенос движения из видео в видео. Модель подходит для коммерческих творческих рабочих процессов в рекламе, электронной коммерции, играх и дизайне интерфейсов, с нативным аудиовизуальным выводом для генерации на основе референсов.

Video gen768pWith audioWith a referenceYour own frames
Alibabaoffline
HappyHorse 1.1
happyhorse-1.1

HappyHorse 1.1 is a video generation model from Alibaba. It generates short videos from a text prompt, a single starting image, or a set of reference images, with output up to 1080p and durations of 3 to 15 seconds. It is suited for creative content, social media clips, and image-driven animation, and improves on the prior version with stronger prompt adherence, smoother motion, and more consistent characters across frames.

Video gen1080pWith audioWith a referenceYour own frames
MoonshotAIoffline
Kimi K2.8 Preview
kimi-k2.8-preview

Kimi K2.8 Preview is a multimodal model by Moonshot AI designed for coding, complex reasoning, and autonomous agentic tasks. It supports a context window of up to 1 million tokens, image and video input, and long-running multi-step workflows with tool use. According to Moonshot AI, K2.8 Preview delivers performance close to the flagship Kimi K3 while using compute more efficiently.

Context 1.05M tokens
Klingoffline
Kling 3.0
kling-3.0

Kling 3.0 is a video generation model from Kling AI designed to create realistic and cinematic videos from text prompts or images. It provides strong motion quality, scene physics, camera movement, visual detail, and consistent results across frames.

Video gen1080pWith audioYour own frames

97 models in the catalog, 91 available right now, 26 model vendors, 9 for image generation, context up to 1.05M tokens

Models by vendor: