Ling 3.1 Flash via the Tokenator API

What the model is for

Ling 3.1 Flash is a hybrid reasoning mixture-of-experts model from inclusionAI, with 25B active parameters out of 560B total.

Specification

Full nameLing 3.1 Flash
API IDling-3.1-flash
Model vendorInclusionAI
TypeText model (chat / completions)
Context window262.1K tokens
Max output32.8K tokens
Image inputno
Tokenator usage multiplier1×
BadgeNEW
Input modalitiestext
Aliasescursor-ling-3.1-flash
Supported API formats/v1/chat/completions /v1/responses /v1/messages
Current statuscurrently unavailable
Data updated2026-10-08

Benchmarks

Intelligence index41.1

Artificial Analysis data loaded into the catalog for this model.

Example request

curl
curl https://api.tokenator.top/v1/chat/completions \
  -H "Authorization: Bearer sk-your-tokenator-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "ling-3.1-flash",
    "messages": [{"role": "user", "content": "Hello"}]
  }'

Compatible tools

The model can be used in any tool that accepts a custom base URL: Claude Code, Codex CLI, OpenCode, Cursor, Cline, Kilo Code, Cherry Studio, DeepSeek Harness, Hermes, Pi, OpenClaw.

FAQ about Ling 3.1 Flash

What is the API ID of Ling 3.1 Flash?

ling-3.1-flash — put this into the model field of your request.

What context window does Ling 3.1 Flash have?

262.1K tokens; max output is 32.8K tokens.

How many tokens will a request cost?

(input + output) × 1×. See the documentation for details.

Which tools support this model?

Any tool that accepts a custom base URL: Claude Code, Codex CLI, OpenCode, Cursor, Cline, Kilo Code. Setup guides are in the integrations section.