Gemini models on Tokenator

Every Gemini model available through Tokenator: API ID, context window, usage multiplier and current availability. The key and the base URL stay the same for every model in the catalog — switching models takes one request parameter.

Geminionline
Free Gemini 3.8 Flash
free-gemini-3.8-flash
×1.7

This model is available on a free plan with set usage limits. Gemini 3.8 Flash is the most intelligent model in Google's Flash series, significantly outperforming Gemini 3.7 Flash in programming, agent-based tasks, and complex multi-step reasoning. It combines high performance with improved capabilities for analysis, planning, and execution of complex tasks, while maintaining the efficiency characteristic of the Flash series.

FreeContext 1M tokens
Geminionline
Gemini 3.8 Flash
gemini-3.8-flash
×1.4–×1.7

Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.

Context 1M tokens
Geminionline
Gemini 3.1 Flash Lite
gemini-3.1-flash-lite
×1.5

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads.

UnstableContext 1M tokens
Geminionline
Gemini 3.1 Pro Preview
gemini-3.1-pro-preview
×1.7–×1.8

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation of the Gemini 3 series, it combines high-precision reasoning across text, image, video, audio, and code with a 1M-token context window. Designed for advanced development and agentic systems, Gemini 3.1 Pro Preview improves long-horizon stability and tool orchestration while increasing token efficiency. It introduces a new medium thinking level to better balance cost, speed, and performance. The model excels in agentic coding, structured planning, multimodal analysis, and workflow automation, making it well-suited for autonomous agents, financial modeling, spreadsheet automation, and high-context enterprise tasks.

Context 1M tokens
Geminionline
Gemini 3.6 Flash
gemini-3.6-flash
×1.4–×1.7

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and less hedging, while reducing token use and the number of model calls needed to complete a task.

Context 1.05M tokens
Geminionline
Gemini 3.7 Flash
gemini-3.7-flash
×1.4–×1.7

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step problem solving.

Context 1M tokens
Geminionline
Gemini Embedding 2
gemini-embedding-2
×1.4

Gemini Embedding 2 is Google's first multimodal embedding model. We currently support mapping text and images into a unified vector space for semantic search and retrieval-augmented generation (RAG). It supports input context up to 8,192 tokens and flexible output dimensions from 128 to 3,072 (recommended: 768, 1536, or 3,072). Designed for cross-modal similarity — you can embed a text query and retrieve the most relevant images, or vice versa — making it well-suited for multimodal search, recommendation, and document understanding pipelines.

EmbeddingsContext 8.2K tokens
Geminionline
Nano Banana 2
nano-banana-2
×1–×2

Gemini 3.1 Flash Image Preview, a.k.a. Nano Banana 2, is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed.

Image gen
Geminionline
Nano Banana Pro
nano-banana-pro
×2

Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and high-fidelity visual synthesis. The model generates context-rich graphics, from infographics and diagrams to cinematic composites, and can incorporate real-time information via Search grounding.

Image gen
Geminionline
Omni
omni
×10

Omni is a multimodal video generation and editing model from Google DeepMind. It supports video creation from text, images, audio, and video references, as well as multi-turn editing through natural language. Omni is particularly strong at preserving characters, scene details, and visual style across edits, while using its understanding of physics and real-world context to produce coherent motion and scenes. It also supports first and last frame control, reference-to-video editing, and native audio generation.

Video gen720pWith audioWith a referenceYour own frames

10 Gemini models in the catalog, 10 available right now, 2 for image generation, context up to 1.05M tokens, multiplier ×1–×2