⚡ New — Kimi K3 is live: bring your own Moonshot key →

Which model should you use?

The cheapest model, the snappiest model and the fastest-streaming model are usuallythree different models — and the winner flips with your workload. Pick your use case; we rank the catalog by what that workload actually feels:₹ per task, time to first token andsustained throughput, measured through the production gateway (sweep of 2026-07-02), not read off datasheets.

Start from your application

Agent loops make dozens of short, tool-calling turns per task — time to first token dominates how fast the agent feels, throughput matters for big diffs, and cost adds up across the loop. Reasoning support is required.

#ModelBest route₹/Mtok (blended)First tokentok/s
1gpt-oss-120bopen weights🇮🇳 India route🔧 toolsfireworks₹391.2s469
2qwen3-32bopen weights🔧 toolsgroq₹222.1s382
3gpt-oss-20bopen weights🇮🇳 India route🔧 toolskrutrim₹244.7s606
4glm-4.7open weights🔧 toolsopenrouterBYOK₹1361.9s68
5gemini-2.5-flash🔧 toolsopenrouter₹1872.2s132
6glm-4.7-flashopen weights🔧 toolsprice-rankedzhipuBYOKfree
7glm-4.5-flashopen weights🔧 toolsprice-rankedzhipuBYOKfree
8nemotron-3.5-lightning-30b-a3bopen weights🔧 toolsprice-rankednvidiaBYOKfree

Blended ₹/Mtok = cheapest route at a 1:3 input:output token mix (generation-dominant). First token and tok/s are the best measured route per model. Rankings are a weighted percentile score per lens — details in the methodology below.

The full picture — every chat model, three lenses

Click a metric column to sort by that lens. The same model often wins one and loses another.

ModelRoutes₹/Mtok ↕First token ↕tok/s ↕
glm-4.7-flashreasoning🔧 toolszhipuopenrouterfree
glm-4.5-flashreasoning🔧 toolszhipufree
nemotron-3.5-lightning-30b-a3breasoning🔧 toolsnvidiafree
qwen2.5-coder-7bbharatrouter 🇮🇳₹3.5
qwen2.5-7b-instruct🔧 toolsbharatrouter 🇮🇳₹3.5
qwen3-8breasoning🔧 toolsbharatrouter 🇮🇳₹3.5
qwen2.5-vl-7b-instructbharatrouter 🇮🇳₹5.3
gemma-4-e4b-it🔧 toolskrutrim 🇮🇳₹6.8
qwen3.5-9breasoning🔧 toolskrutrim 🇮🇳₹6.8
llama-3.1-8b-instruct🔧 toolsopenroutergroqfireworks₹7.0428ms groq548
glm-4-32b-0414-128k🔧 toolszhipu₹10
command-r7b🔧 toolscohere₹12
qwen3-32breasoning🔧 toolsgroqopenrouterfireworks₹222.1s groq382
gemma-4-26b-a4b-it🔧 toolskrutrim 🇮🇳₹23
qwen3.6-35b-a3breasoning🔧 toolskrutrim 🇮🇳₹23
gpt-oss-20breasoning🔧 toolskrutrim 🇮🇳groqfireworks₹244.7s krutrim606
deepseek-v4-flashreasoning🔧 toolsdeepseek₹24
devstral-small🔧 toolsmistral₹24
llama-3.3-70b🔧 toolsgroqopenrouterfireworks₹262.0s openrouter218
gemma-4-31b-it🔧 toolskrutrim 🇮🇳₹27
llama-4-scout🔧 toolsgroqfireworks₹28
glm-4.6v-flashxreasoning🔧 toolszhipu₹30
glm-4.7-flashxreasoning🔧 toolszhipu₹30
gemini-2.5-flash-litereasoning🔧 toolsgemini₹31
gpt-oss-120breasoning🔧 toolskrutrim 🇮🇳groqbasetenfireworks₹391.2s fireworks469
mistral-small🔧 toolsmistral₹47
command-r🔧 toolscohere₹47
gpt-4o-mini🔧 toolsopenai₹47
gpt-4o-mini-search-preview🔧 toolsopenai₹47
nemotron-superreasoning🔧 toolsbaseten₹61
glm-4.5-airreasoning🔧 toolszhipuopenrouterfireworks₹65
glm-4.6vreasoning🔧 toolszhipuopenrouter₹72
codestral🔧 toolsmistral₹72
deepseek-v4-proreasoning🔧 toolsdeepseekbasetenfireworks₹74
deepseek-v3🔧 toolsopenrouter₹8116.3s openrouter38
gpt-5.6-lunareasoning🔧 toolsopenai₹91
sonar🔧 toolsperplexity₹96
grok-code-fast-1reasoning🔧 toolsxai₹113
gemini-3.1-flash-litereasoning🔧 toolsgemini₹114
mistral-largereasoning🔧 toolsmistral₹120
glm-4.7reasoning🔧 toolsbasetenzhipuopenrouter₹1361.9s openrouter68
qwen3.6-27breasoning🔧 toolskrutrim 🇮🇳groq₹147
gpt-5-minireasoning🔧 toolsopenai₹15010.2s openai63
glm-5reasoning🔧 toolsbasetenzhipuopenrouter₹15312.2s zhipu53
glm-4.6reasoning🔧 toolszhipuopenrouter₹156
grok-build-0.1reasoning🔧 toolsxaiopenrouter₹168
glm-4.5reasoning🔧 toolszhipuopenrouter₹173
kimi-k2.5reasoning🔧 toolsbasetenopenrouter₹173
kimi-k2🔧 toolsgroqopenrouter₹180688ms groq178
nemotron-ultrareasoning🔧 toolsbaseten₹187
gemini-2.5-flashreasoning🔧 toolsgeminiopenrouter₹1872.2s openrouter132
gemini-3.5-flash-litereasoning🔧 toolsgemini₹187
grok-4.3reasoning🔧 toolsxaiopenrouter₹210
grok-4.20reasoning🔧 toolsxaiopenrouter₹210
glm-5.2reasoning🔧 toolsbasetenzhipuopenrouterfireworks₹2385.7s zhipu68
glm-5.1reasoning🔧 toolsbasetenzhipuopenrouter₹242
kimi-k2.7-codereasoning🔧 toolsmoonshotbasetenopenrouter₹27019.6s openrouter39
gemini-3.7-flashreasoning🔧 toolsgemini₹288
kimi-k2.6reasoning🔧 toolsmoonshotbasetenopenrouter₹311
glm-5-turboreasoning🔧 toolszhipu₹317
glm-5v-turboreasoning🔧 toolszhipuopenrouter₹317
glm-5.3reasoning🔧 toolszhipu₹350
glm-4.5-airxreasoning🔧 toolszhipu₹351
claude-haiku-4.5reasoning🔧 toolsanthropicopenrouter₹38416.8s openrouter102
qwen3.8-maxreasoning🔧 toolsalibabaopenrouter₹480
pixtral-large🔧 toolsmistral₹480
grok-4.6reasoning🔧 toolsxai₹480
grok-4.5reasoning🔧 toolsxai₹480
gemini-3.6-flashreasoning🔧 toolsgemini₹576
mistral-medium-3.5🔧 toolsmistral₹576
kimi-k2.7-code-highspeedreasoning🔧 toolsmoonshot₹622
sonar-reasoning-proreasoning🔧 toolsperplexity₹624
sonar-deep-researchreasoning🔧 toolsperplexity₹624
gemini-3.5-flashreasoning🔧 toolsgemini₹684
glm-4.5-xreasoning🔧 toolszhipu₹693
gemini-2.5-proreasoning🔧 toolsgeminiopenrouter₹750
gpt-5reasoning🔧 toolsopenai₹750
claude-sonnet-5reasoning🔧 toolsanthropicopenrouter₹768
command-a🔧 toolscohere₹780
gpt-4o-search-preview🔧 toolsopenai₹780
gpt-5.6-terrareasoning🔧 toolsopenai₹912
kimi-k3reasoning🔧 toolsmoonshotopenrouterbaseten₹1152
sonar-pro🔧 toolsperplexity₹1152
claude-opus-4.8reasoning🔧 toolsanthropicopenrouter₹1920
claude-opus-5reasoning🔧 toolsanthropic₹1920
gpt-5.6-solreasoning🔧 toolsopenai₹2280
gpt-5.6reasoning🔧 toolsopenai₹2280
claude-fable-5reasoning🔧 toolsanthropicopenrouter₹3840

In the open — how these numbers are made

Perf numbers are medians from a multi-run streamed sweep through the production gateway on 2026-07-02: every (model × provider) route gets the same ~300-token prompt, rounds interleaved across hosts so no provider owns a time-of-day advantage. First token counts reasoning tokens (it's what you see). tok/s is the post-first-token decode rate. Models not yet swept show “—” and rank on price with a neutral perf score. Routes we could not measure are listed openly in theAPI response(9 skipped this sweep), never silently dropped. Numbers refresh with each sweep; live per-route health is on /models.

Agents get this same chooser as JSON:GET /v1/compare/models — rankings, per-route pricing, measured perf and live failure rates, no auth required.

Get a keyBrowse the catalog