USD per 1M tokens, grouped by weight class so
open-weight heavies stand next to the big labs at their real price points. Thinking tokens bill as output.
Prices updated 2026-08-26 · capability scores 2026-08-26.
Want this in money rather than tokens? Workload costs →
Reasoning = Artificial Analysis Intelligence Index. Coding = AA Coding Index. Agentic = our mean of AA’s τ²-bench and Terminal-Bench Hard scores — AA publishes no single agentic index, so this one is ours and we say so. A dash means AA has not run that benchmark on that model yet.
Filter:
Frontier The judgment tier — best reasoning money buys
Model
In /1M
Out /1M
Cached
$/task
Context
Reasoning
Coding
Agentic
Src
MiniMax M3MiniMax
$0.30
$1.20
$0.06
$0.21
1000K
45.4
58.6
65.7
GLM-5.3 (max)Z.AI
Cheapest way to buy frontier-band intelligence right now. AA scores it 59.5 - inside the top six with Opus 5 (63.1), Fable 5 (62.1), Grok 4.6 and GPT-5.6 Sol (60.9) and Kimi K3 (59.7) - at $1.40/$4.40, against $2/$6 for the next-cheapest model in that band. Cost per task $1.32 vs $1.64 for Kimi K3, $1.95 for Grok 4.6, $4.38 for Opus 5. Same list price as GLM-5.2, which it replaces on 7 index points and 6 coding points. Z.AI's own numbers show it doing that on fewer output tokens: 34.5% agentic coding at ~75K tokens/task vs GLM-5.2's 23.4% at 96K. Slow, though - 80 tok/s and 26s to first answer token, so it is a batch model, not an interactive one. AA lists it as closed weights at launch; GLM-5.2 weights are open.
$1.40
$4.40
$0.26
$1.32
1024K
59.5
74.8
—
Grok 4.6xAI
Released 12/08/2026 as xAI's flagship for code and everything else; the docs now route every text use case to it. Same headline price as Grok 4.5 - $2/$6, 500k context - for 5.1 more points of AA intelligence (60.9 vs 55.8) and GPQA 94.9, second only to Claude Opus 5 on AA's index. Cached input is $0.50, not the $0.30 Grok 4.5 gets. Long-context surcharge applies as before: prompts at or above 200k tokens bill everything at $4/$12. Cost per task $1.95 - 2.6x Grok 4.5's $0.76, because the model reasons longer. agentic left unset: AA has run TerminalBench v2.1 (0.884) but neither tau-2 nor TerminalBench Hard, and this column is the mean of that pair. Listed regions are us-east-1 and us-west-2 only; the EU-access caveat logged on Grok 4.5 has not been cleared on the model page. Knowledge cut-off 01/02/2026.
$2.00
$6.00
$0.50
$1.95
500K
60.9
76.8
—
Grok 4.5xAI
EU access pending AI Act compliance (xAI: mid-July 2026)
$2.00
$6.00
$0.30
$0.76
500K
55.8
72.4
—
Qwen3.8 MaxAlibaba
Alibaba's flagship, launched 03/08/2026, and since 12/08/2026 the first Qwen-Max-class model with open weights. Sparse MoE: 2.446T total parameters in BF16, 512 experts with 10 active per token (~95B active), 262K native context extensible past 1M, 128K max output, multimodal. The open release is Qwen/Qwen3.8-2.4T-A95B on Hugging Face - public, not gated, custom 'qwen3.8-max' licence rather than Apache-2.0, plus an official FP8 quant. That closes the promise made at launch and reverses this entry's earlier 'weights not shipped' note. Hosted list price is unchanged at $2/$6 per 1M with $0.25 cached input; self-hosting 4.9TB of weights is a data-centre job, not a workstation one, so the API price is still the number that matters for most buyers. AA scores the open-weight build separately at intelligence index 57.7 versus 58.1 for Alibaba's hosted endpoint - same model, different host, no meaningful capability gap. Speaks both the OpenAI and Anthropic API protocols. OpenRouter's weighted-average input runs $0.634 at a 77.9% cache-hit rate, so repeat-context workloads land well under list.
$2.00
$6.00
$0.25
$2.27
1000K
58.1
71.8
—
Qwen3.8 2.4T A95BAlibaba Qwen
The open-weight twin of Qwen3.8 Max: 2.4T total parameters, 95B active, weights on Hugging Face, and the same $2/$6 hosted price. AA scores it 57.7 against Max's 58.1 and 71.9 coding against 71.8 - inside the noise. GPQA 93.5 is the second-highest tracked here after Gemini 3.7 Flash. The reason to pick it over Max is that you can run it yourself; the reason not to is that at 95B active parameters self-hosting is a cluster, not a workstation. Cost per task $2.33, roughly 1.8x GLM-5.3 for 1.8 fewer index points.
$2.00
$6.00
$0.25
$2.33
1000K
57.7
71.9
—
Qwen3.7 MaxAlibaba
Alibaba's proprietary flagship, launched 20/05/2026 — weights not released, unlike the open-weight Qwen3.x line. Highest-scoring Chinese model on the AA intelligence index at 46.0. Price shown is the International/US list rate ($2.50/$7.50); Alibaba's Global and Chinese-mainland regions list the same model at $1.65/$4.951, and a limited-time 50% promo is running on the international rate. Cached input $0.25 (90% off). Natively speaks the Anthropic API protocol, so it drops into Claude Code without an adapter.
$2.50
$7.50
$0.25
$0.93
1000K
46.7
66.0
72.7
Gemini 3.1 ProGoogle
$2.00
$12.00
$0.20
$0.53
1000K
47.7
68.8
74.7
Kimi K3Moonshot AI
Big jump from the K2.x line, released 15/07/2026. GPQA 93.5 beats Claude Opus 4.8 and nearly matches Gemini 3.1 Pro. API-only for now, weights aren't public yet, so no open-weights tag until they ship.
$3.00
$15.00
$0.30
$1.64
1000K
59.7
76.2
—
GPT-5.6 SolOpenAI
Standard tier. OpenAI's pricing page flags this as promotional pricing available at least through 21/11/2026; the previously tracked $5/$30 was gpt-5.5's list price, not Sol's. Batch and Flex tiers are $2/$10; Fast mode is $8/$40. Cost-per-task ($2.55) is AA's published figure and is still computed on the pre-promotional $5/$30, so it overstates the current run cost by roughly 1.5x; it is left as published rather than recomputed here.
$4.00
$20.00
$0.40
$2.06
400K
60.9
77.4
75.5
Claude Opus 4.8Anthropic
$5.00
$25.00
$0.50
$3.28
1000K
57.3
74.3
76.4
Claude Opus 5Anthropic
released 24/07/2026, same $5/$25 rate as Opus 4.8 which it will eventually supersede; not yet marked deprecated on vendor page
$5.00
$25.00
$0.50
$4.38
1000K
63.1
78.0
—
Claude Fable 5Anthropic
$10.00
$50.00
$1.00
$5.60
1000K
62.1
76.5
80.7
Claude Mythos 5Anthropic
Limited availability. Same list price as Fable 5; no independent benchmark yet, so the capability columns stay empty.
$10.00
$50.00
$1.00
—
1000K
—
—
—
Heavy Near-frontier capability, workhorse pricing
Model
In /1M
Out /1M
Cached
$/task
Context
Reasoning
Coding
Agentic
Src
Hunyuan Hy3Tencent
First Tencent entry tracked. Released 06/07/2026, integrated across Tencent products via Tencent Cloud TokenHub. Vendor price RMB1/RMB4 per 1M ≈ $0.15/$0.59, cached RMB0.25 ≈ $0.037; AA representative host lists $0.14/$0.58. Strong GPQA (89.7) at a low price, but AA overall intelligence index (41.2) and coding (58.8) put it mid-pack. Reasoning and non-reasoning variants exist; figures here are the standard Hy3. Weights not public — no open-weights tag until they ship. context_k unverified. AA redefined its Agentic Index between 19/08 and 21/08/2026 - the component benchmarks are now GDPval-AA v2 and tau3-Banking (previously tau-2 + TerminalBench Hard). This model dropped out of the capabilities dataset in that switch; it is back in as of 23/08/2026 at costPerTask 0.0405, which rounds to the $0.04 already carried here, so the figure is now live-sourced again rather than frozen.
$0.14
$0.56
$0.037
$0.04
—
42.2
58.8
—
MiniMax M2.7MiniMax
Corrected to vendor pricing 24/07/2026 — $0.24/$0.96 was the OpenRouter reseller rate, not MiniMax's own listed price. Vendor (minimax.io) and Artificial Analysis both confirm $0.30/$1.20, cache reads $0.06, cache writes $0.375.
$0.30
$1.20
$0.06
$0.08
205K
38.9
52.6
62.1
Inkling-Small (Preview)Thinking Machines
New sibling of the existing 'inkling' entry, announced alongside it 15/07/2026 and made available as a hosted preview 30/07/2026. 276B total / 12B active MoE (vs 975B total / 41B active for full Inkling) — vendor's own benchmark table shows it matching or slightly exceeding full Inkling on most evals (GPQA Diamond 88.3% vs 87.2%, AA intelligence index 40.2 vs 40.7) at roughly a third of the price ($0.30/$1.20 vs Inkling's corrected $1.00/$4.05; the old 1/6th figure was against the superseded $1.87/$4.68). Full open weights NOT yet released — vendor states testing is still being finished; today's pricing reflects early hosted-provider access only, not a finalized public price. context_k left null: vendor doc gives Inkling's own context options (64K/256K via Tinker, up to 1M native) but does not confirm the same for Small — not assuming equality. Re-check on full-weights GA.
$0.30
$1.20
—
$0.09
—
41.2
52.9
—
Qwen3.7 PlusAlibaba
Value tier under Qwen3.7 Max, released 01/06/2026. Natively multimodal, 1M context. Price is the 0-256K input tier; longer inputs step up to $1.2/$4.8 per Alibaba's published tier table, so treat $0.40/$1.60 as the floor, not a flat rate. Global/Chinese-mainland regions list $0.276/$1.101. GPQA 90.0 for a fifth of Max's price is the headline; AA intelligence index 39.0 puts it mid-pack overall.
$0.40
$1.60
—
$0.31
1000K
39.4
55.9
70.0
Nemotron 3 Ultra 550B A55BNVIDIA (open weights)
NVIDIA's first row in this table, and it does not win a column. AA rates it 38.3 intelligence, 49.3 coding, 59.8 on our agentic index (tau-2 83.3, TerminalBench Hard 36.4) at $0.50/$2.20 - mid-table on every one of those. The direct comparison is unkind: DeepSeek V4 Pro scores higher on all four (45.3 / 71.2 agentic / 46.2 TerminalBench Hard) and costs $0.435/$0.87, so on a hosted API there is no metric on which Nemotron is the answer. It is listed because the reasons to run it are not on the scoreboard: open weights under OpenMDW-1.1, 550B total with only 55B active on a Mamba2-Transformer LatentMoE, 1M context, and NVIDIA ships the serving recipes (vLLM, SGLang, TensorRT-LLM) for its own hardware. If you self-host on NVIDIA silicon and want a frontier-scale agent model with a permissive licence, this is a real option; if you are buying tokens, DeepSeek V4 Pro is cheaper and better. Two caveats on the numbers above: hosted routes cap context at 512K today, the 1M figure is what the weights do when you serve them yourself; and AA's host median is $0.60/$2.75, ~20-25% above the tracked OpenRouter figure - the same vendor-vs-host gap already logged on the other open-weight rows.
$0.60
$2.75
—
$0.49
1000K
38.3
49.3
59.8
LongCat 2.0Meituan
New entrant 30/07/2026 — Meituan (China's largest food-delivery/local-services platform) enters the LLM race. Sparse MoE: 48B active / 1.6T total parameters, 1M context, open weights, released 20/07/2026. OpenRouter's sole host (AtlasCloud) lists a 60%-off promo at $0.30/$1.20 — AA's representative-host price runs higher at $0.75/$2.95, so treat the low figure as promotional, not guaranteed permanent. Mid-pack capability (AA intelligence index 33.5) despite the parameter count. AA redefined its Agentic Index between 19/08 and 21/08/2026 - the component benchmarks are now GDPval-AA v2 and tau3-Banking (previously tau-2 + TerminalBench Hard). This model dropped out of the capabilities dataset in that switch, so no costPerTask is published for it today; the figure carried here is the last one AA published under the previous definition and is frozen until the model reappears.
$0.75
$2.95
—
$0.17
1000K
34.0
45.3
—
Gemini 3 FlashGoogle
$0.50
$3.00
$0.05
—
1000K
27.9
—
37.5
Qwen3.8 27BAlibaba Qwen
Dense 27B open-weights release. Repriced 23/08/2026 onto Alibaba's own endpoint - $0.50 in / $3.00 out / $0.10 cached read, 1M context - because the vendor rate is the source of truth here and it also settles the context question: every third-party host caps at 262,144, only Alibaba serves the full 1M. Third-party routes are cheaper on input and span $0.40-$0.48 in / $3.00-$3.40 out across seven hosts (CoreWeave, Chutes and AkashML at $0.40/$3.00; Reka, Venice and Parasail at $0.45/$3.20; Io Net at $0.48/$3.40 on a 65K context). Previously tracked $0.45/$3.20 was the OpenRouter default route, which has since moved. Correction: an earlier note here called this the cheapest model above 90 GPQA. That was wrong - GPT-5.6 Luna, MiniMax M3, DeepSeek V4 Flash and Qwen3.7 Plus all clear 90 GPQA at a lower blended rate. It is a cheap 90+ GPQA model, not the cheapest. Cost per task $0.380 unchanged.
$0.50
$3.00
$0.10
$0.38
1000K
52.0
68.1
—
Gemini 3.7 FlashGoogle
Released 13/08/2026, replaces Gemini 3.6 Flash. The $0.75/$3.75 is introductory pricing that Google's own page says runs through 31/12/2026 and then doubles to $1.50/$7.50 on 01/01/2027 - which is exactly what 3.6 Flash costs today. So the headline is a temporary half-price window on a model that scores higher than the one it replaces: AA intelligence 56.0 vs 51.6, coding 76.1 vs 69.2, GPQA 94.5 vs 92.8. Cost per task is flat at $0.816 vs $0.805 - the cheaper tokens are offset by more thinking tokens, so on agentic work this is not yet a cost win. Context cache $0.075 through 31/12/2026, then $0.15. Throughput 515 tok/s, TTFT 6.4s. Agentic pair stays unset: tau2 and TerminalBench Hard are null; AA does report TerminalBench v2.1 0.858 and tau-banking 0.328.
$0.75
$3.75
$0.075
$0.82
1000K
56.0
76.1
—
Gemini 3.6 FlashGoogle
Replaces Gemini 3.5 Flash, released 21/07/2026. Price halved to $0.75/$3.75 some time before 19/08/2026 - the same introductory rate Gemini 3.7 Flash launched on, and it carries the same 01/01/2027 revert to $1.50/$7.50. That removes the reason to run 3.6 over 3.7: identical price, 3.7 scores higher (AA 56.0 vs 51.6, coding 76.1 vs 69.2) and costs less per task ($0.82 vs $0.49 - 3.6 is now the cheaper of the two per task, but the gap is small enough that the index difference decides it). Google says up to 65% lower agent token cost on long-horizon engineering tasks vs the prior gen.
$0.75
$3.75
$0.075
$0.49
1000K
51.6
69.2
—
DeepSeek V4 ProDeepSeek
PRICE RISE IS LIVE. The peak/off-peak switch landed at 16:00 UTC on 16/08/2026 and the vendor page now bills only at those rates. Tracked figure flipped to PEAK today: $1.32 in / $3.96 out cache-miss, $0.044 cache-hit input. Off-peak $0.66 / $1.98, cache-hit $0.022. Peak hours are 01:00-04:00 and 06:00-10:00 UTC, 7 hours of every day. Cost per task moved with it, from the old pre-rise measurement to $0.362, and the AA figure is now adopted because AA and the vendor finally price this row on the same basis. Historical context: DeepSeek switches to peak/off-peak billing: off-peak $0.66/$1.98, peak $1.32/$3.96, cache-hit input $0.022 off-peak and $0.044 peak. Peak is 01:00-04:00 and 06:00-10:00 UTC, so 7 hours of every day. Even off-peak that is 1.5x today's input and 2.3x today's output; at peak it is 3x and 4.6x. The cheapest credible frontier-adjacent model on this table just got materially more expensive. The tracked figure stays at the live $0.435/$0.87 until the switch lands, and cost per task stays measured on that price. AA has already repriced this row at peak (cost per task $0.081 -> $0.362 on unchanged token counts). AA also dropped the agentic pair when it repointed the slug to the 0813 build, so the agentic column is blank rather than carried forward from a superseded build. Once live, peak becomes the headline number here and off-peak stays in this note. AA re-measured cost per task on 13/08/2026: $0.046 -> $0.081, on unchanged vendor pricing. AA also published a new checkpoint row, DeepSeek V4 Pro 0813 (intelligence 53.0, coding 68.8, GPQA 0.928), released 13/08/2026. It is NOT tracked as a separate row: DeepSeek sells one model id, deepseek-v4-pro, at one price, and AA carries no throughput, no tau-2, no TerminalBench Hard and no cost per task for the 0813 row yet. Repointing today would trade a fully measured row for a half-empty one. Revisit once AA characterises it. Update 2026-08-13: AA has now repointed the deepseek-v4-pro slug itself to the 0813 build and re-scored it - intelligence 45.3 -> 53.2, coding -> 68.8, GPQA -> 92.8, at the same $0.435/$0.87. Cost per task and the agentic pair (tau2 / TerminalBench Hard) are still null on the new row, so cost_per_task stays as measured on the prior build.
$1.32
$3.96
$0.044
$0.37
1000K
53.2
68.8
—
Kimi K2.6Moonshot AI
Free promo is over. Now $0.95/$4.00 per 1M.
$0.95
$4.00
$0.14
$0.46
256K
45.1
61.8
69.9
Kimi K2.7 CodeMoonshot AI
$0.95
$4.00
$0.19
$0.33
262K
43.0
60.8
67.4
Inkling (xhigh)Thinking Machines
Released 15/07/2026, first tracked entry from Thinking Machines Lab (Mira Murati's startup). Open-weight MoE, 41B active / 975B total params. Price corrected 06/08/2026 from $1.87/$4.68 to $1.00/$4.05: the higher figure was AA's representative host on 23/07 and no live provider carries it any more. Price Per Token's provider table (06/08/2026) lists all 4 hosts at $4.05 output and $0.95-$1.00 input - DeepInfra $0.95 (cheapest), Together / BaseTen / OpenRouter $1.00. Tracked figure = the $1.00 modal list price, which is also AA's current default host. Cached read $0.17 (Together/BaseTen/OpenRouter tier; DeepInfra $0.16). AA redefined its Agentic Index between 19/08 and 21/08/2026 - the component benchmarks are now GDPval-AA v2 and tau3-Banking (previously tau-2 + TerminalBench Hard). This model dropped out of the capabilities dataset in that switch, so no costPerTask is published for it today; the figure carried here is the last one AA published under the previous definition and is frozen until the model reappears.
$1.00
$4.05
$0.17
$0.44
524K
42.3
52.1
—
Muse Spark 1.1Meta
First Muse Spark entry tracked; the free-promo original never got added. This is Meta's first paid API model outside the open-weight Llama line. Zuckerberg pegged the price at roughly 25% of comparable Anthropic and OpenAI models. US-only preview for now, waitlist required, not available in the EU.
$1.25
$4.25
$0.15
$0.40
1000K
53.2
71.3
—
Muse Spark 1.2Meta
Released 05/08/2026 alongside Meta's new terminal coding agent 'Muse Code' (macOS/Linux beta). Coding-focused update to Muse Spark 1.1, co-trained with Muse Code so model behavior is calibrated to the harness (not a general-purpose model adapted for coding). Same $1.25/$4.25 pay-as-you-go pricing as 1.1; new contributor tier (opt in to share session data) priced >10x cheaper than pay-as-you-go, no exact figure published. Unlike 1.1's US-only preview, Meta says 1.2 ships with 'expanded global access' via the Meta Model API — EU availability not explicitly confirmed, flag for next check. context_k left null: not stated in the launch coverage reviewed. AA API TerminalBench v2.1 score (0.801) suggests strong agentic coding performance, but the standard agentic_from pair (tau2 + terminalbench_hard) is null for this entry so 'agentic' is left unset for cross-model comparability.
$1.25
$4.25
—
$0.61
—
56.8
72.2
—
GLM-5.2Z.AI
RXed eval winner 12/07/2026 — closest to Fable on audit work at ~9x lower cost Price realigned 19/08/2026: the $1.351/$4.29 carried since 05/08 was an AA host median that AA itself has since moved back to $1.40/$4.40, matching docs.z.ai and the OpenRouter Z.AI first-party endpoint. Third-party hosts sell it for $0.50-$0.75 input, so the spread is real, but the vendor rate is what gets tracked. Superseded by GLM-5.3 at the same list price.
$1.40
$4.40
$0.26
$0.73
1000K
52.6
68.8
74.9
Gemini 3.5 FlashGoogle
$1.50
$9.00
$0.15
$1.37
1000K
52.0
70.1
68.1
Claude Sonnet 5Anthropic
$2.00
$10.00
$0.20
$2.89
1000K
55.3
71.5
—
GPT-5.6 TerraOpenAI
$2.00
$12.00
$0.20
$0.89
400K
56.6
76.7
71.9
Claude Sonnet 4.6Anthropic
AA redefined its Agentic Index between 19/08 and 21/08/2026 - the component benchmarks are now GDPval-AA v2 and tau3-Banking (previously tau-2 + TerminalBench Hard). This model dropped out of the capabilities dataset in that switch, so no costPerTask is published for it today; the figure carried here is the last one AA published under the previous definition and is frozen until the model reappears.
$3.00
$15.00
$0.30
$1.39
1000K
36.8
—
62.9
Medium Production volume work
Model
In /1M
Out /1M
Cached
$/task
Context
Reasoning
Coding
Agentic
Src
Grok 4.1 FastxAI
DELISTING EVIDENCE, 24/08/2026 — needs an owner keep-or-remove call. Grok 4.1 Fast is absent from xAI's public model registry (globalThis.__XAI_PUBLIC_MODELS__ on docs.x.ai lists only grok-4.20 beta, 4.3, 4.5, 4.6 and grok-build-0.1), and OpenRouter now returns an EMPTY endpoints array for it — the model page still exists, but no provider serves it. That is two independent sources saying it can no longer be bought. Yesterday's run flagged this row as aggregator-only on the strength of the registry alone; the tracked figure now rests solely on pricepertoken, which lags delistings.
$0
$0
—
—
2000K
31.3
—
58.8
Ling 3.0 FlashInclusionAI (Ant Group)
Ant Group's InclusionAI lab. 124B total / ~5.1B active MoE, released 23/07/2026, added to AA 04/08/2026. Best intelligence-per-dollar entry in this table under $0.15 blended: AA intelligence 38.3 and GPQA 85.5 at $0.111 blended, against GLM-4.7 Flash's 23.3 at $0.153 and GPT-OSS 120B's 24.1 at $0.262. DeepSeek V4 Flash still scores higher (51.8) for 1.6x the blended price. Non-thinking line (Ring is the reasoning sibling); 131K context is short next to the 1M-context flash models. Single host on OpenRouter, so no cross-host price spread to reconcile yet.
$0.075
$0.22
—
$0.03
131K
37.8
50.6
—
Grok 4 FastxAI
DELISTING EVIDENCE, 24/08/2026 — needs an owner keep-or-remove call. Grok 4 Fast is absent from xAI's public model registry (globalThis.__XAI_PUBLIC_MODELS__ on docs.x.ai lists only grok-4.20 beta, 4.3, 4.5, 4.6 and grok-build-0.1), and OpenRouter now returns an EMPTY endpoints array for it — the model page still exists, but no provider serves it. That is two independent sources saying it can no longer be bought. Yesterday's run flagged this row as aggregator-only on the strength of the registry alone; the tracked figure now rests solely on pricepertoken, which lags delistings.
$0.20
$0.50
—
—
2000K
27.9
—
42.4
GPT-OSS 120BOpenAI (open)
Open-weight model sold by many hosts at different rates, so this row tracks the modal host route rather than a vendor list price. Resolved 23/08/2026: OpenRouter lists 20 endpoints and $0.15/$0.60 is the clear mode (7 of 20 hosts, incl. Amazon Bedrock, Groq, Together, Nebius, DeepInfra, Phala), matching AA's host median exactly. Output corrected $0.59 -> $0.60 to match. The cheapest route is CoreWeave at $0.03/$0.17 and the dearest is Cerebras at $0.35/$0.75 - a 12x spread, which is why no single number here is a 'price'. Cached read varies by host ($0.02 DigitalOcean to $0.35 Cerebras); the $0.04 carried here is not the modal figure and is the weakest number in this row.
$0.15
$0.60
$0.04
$0.15
131K
24.1
30.4
44.6
Llama 4 MaverickMeta (hosted)
Open-weight model with no first-party per-token price, so this row tracks the host median. Resolved 23/08/2026: OpenRouter lists 5 endpoints - DigitalOcean $0.20/$0.696, DeepInfra $0.20/$0.80, Novita $0.27/$0.85, Parasail $0.35/$1.00, Google $0.35/$1.15. The tracked $0.27/$0.85 is the exact median of both columns, so the figure is confirmed rather than merely 'not disproven'. AA's own host median reads $0.26/$0.91 - same ballpark, different weighting. Cost per task $0.073 unchanged. Cheapest route is $0.20/$0.696; context varies 128K-1M by host.
$0.27
$0.85
$0.17
$0.07
1000K
14.5
16.3
12.3
Agnes 2.5 Pro AlphaAgnes AI (Sapiens AI)
Singapore's Agnes AI (parent: Sapiens AI) enters the table on price, not on peak capability. AA scores it 39.7 intelligence and 58.8 coding at $0.45/$0.90, which AA puts at $0.18 per 1M blended 7:2:1 - second cheapest of anything scoring 38-40, behind DeepSeek V4 Flash at $0.06. The catch: DeepSeek V4 Pro hits 45.3 at that same $0.18 blended and ships open weights, so the case for Agnes rests on the 1M context and the $0.0038 cache read (third cheapest cache tier tracked here, after DeepSeek V4 Flash 0.0028 and V4 Pro 0.0036), not on the score. Proprietary weights, first-party API only (apihub.agnes-ai.com), so no second host to price-check against. The lab's reach is real - AA reports its free omni-modal API past 3M users. Still labelled alpha; the vendor shipped a GA agnes-2.5-pro on 01/08/2026 at identical published pricing that AA has not benchmarked separately, so the alpha stays the tracked entry - it is the version with a third-party score.
$0.45
$0.90
$0.004
$0.05
1000K
39.7
58.8
—
GPT-5.6 LunaOpenAI
$0.20
$1.20
$0.02
$0.08
400K
52.3
71.4
—
DeepSeek V4 FlashDeepSeek
PRICE RISE IS LIVE. The peak/off-peak switch landed at 16:00 UTC on 16/08/2026 and the vendor page now bills only at those rates. Tracked figure flipped to PEAK today: $0.44 in / $1.32 out cache-miss, $0.014 cache-hit input. Off-peak $0.22 / $0.66, cache-hit $0.007. Peak hours are 01:00-04:00 and 06:00-10:00 UTC, 7 hours of every day. Cost per task moved with it, from the old pre-rise measurement to $0.182, and the AA figure is now adopted because AA and the vendor finally price this row on the same basis. Historical context: DeepSeek switches to peak/off-peak billing: off-peak $0.22/$0.66, peak $0.44/$1.32, cache-hit input $0.007 off-peak and $0.014 peak. Peak is 01:00-04:00 and 06:00-10:00 UTC, so 7 hours of every day. Best case that is 1.6x today's input and 2.4x today's output; worst case 3.1x and 4.7x. The tracked figure stays at the live $0.14/$0.28 until the switch actually lands, and cost per task stays measured on that price so the row stays internally consistent. AA has already repriced this row at peak, which is why its cost per task jumped ~4x on unchanged token counts - that is the announcement flowing through, not a re-measurement. Once live, peak becomes the headline number here and off-peak stays in this note. vendor price confirmed $0.14/$0.28 (cache miss); PPT lists $0.077/$0.154 from 3rd-party providers
$0.44
$1.32
$0.014
$0.18
1000K
51.8
69.1
—
Gemini 3.1 Flash LiteGoogle
$0.25
$1.50
$0.025
$0.03
1000K
25.6
34.7
27.8
Muse Glimmer 30BMeta (open weights)
Released 10/08/2026 (weights on Hugging Face 09/08). Dense 30B multimodal (image+text in, text out), Apache-2.0, from Meta Superintelligence Labs, distilled from Muse Spark and aimed at autonomous agents running on consumer hardware. There is no first-party price: Meta's Model API pricing page lists only muse-spark-1.1 / 1.2 / 1.2-contributor, so the tracked $0.35/$1.50 is the hosted route (Together via OpenRouter, the only endpoint listed at time of check) - the same vendor-vs-host caveat already logged on the other open-weight rows. As of 13/08/2026 OpenRouter lists three hosts: DeepInfra $0.30/$1.20, Together $0.35/$1.50, Fireworks $0.35/$1.50. The tracked figure stays at the modal $0.35/$1.50 (two of three hosts); DeepInfra is the cheapest route. AA now carries a host median of $0.325/$1.35 - cross-check only, never the tracked figure. AA also filled in the model this run: cost per task $0.089, 99.3 tok/s. agentic stays unset: the standard pair (tau2 + TerminalBench Hard) is null; AA does report TerminalBench v2.1 0.517 and tau-banking 0.235. Weights ship in GGUF / MLX / NVFP4 / ExecuTorch on day one, which is the actual point of this row - at 30B dense it is the first Meta agent model most people can run locally.
$0.35
$1.50
$0.04
$0.09
131K
35.1
49.0
—
Grok 4.3xAI
$1.25
$2.50
$0.20
$0.12
—
37.9
42.2
67.8
Grok 4.20xAI
Off the watchlist and onto the page, because xAI finally gave it a stable name. Grok 4.20 has been benchmarked since 07/04/2026 but every alias carried -beta or -experimental until this week; the registry now resolves a plain grok-4.20 to the 0309 reasoning build, and OpenRouter lists it non-beta with four live xAI endpoints. Price is $1.25 in / $2.50 out with $0.20 cached read - identical to Grok 4.3 - and the long-context surcharge doubles all three above a 200k prompt. What you buy for that money: AA intelligence 38, GPQA 91.1, HLE 34.5, and a tau-2 of 92.98 that is the highest tool-calling score on this page, frontier rows included. The agentic column reads 65.4 because it is the mean of that tau-2 and a Terminal-Bench Hard of 37.9, and the gap between those two is the story: this model calls tools reliably and then struggles to finish a long terminal task. Read it as a cheap tool-router, not a coder. Context is 1M on xAI's own registry; OpenRouter advertises 2M and the vendor number is the one tracked here. Two caveats. AA reports no coding index, no throughput and no time-to-first-token for it, so speed is unknown. Cost per task is null because AA's capabilities page dropped to a 24-model server-rendered subset on 25/08 and this row is not in it. Also note logprobs and top_logprobs are silently ignored from grok-4.20 onward.
$1.25
$2.50
$0.20
—
1000K
38.0
—
65.4
Gemini 3.5 Flash-LiteGoogle
Released 21/07/2026, cheapest current Gemini 3.x tier. Batch mode halves price to $0.15/$1.25. AA coding index (49.3) is well below the sibling 3.1 Flash Lite's field (untracked there), so treat as a cost-optimized variant, not a capability upgrade.
$0.30
$2.50
—
$0.13
1000K
37.4
49.3
—
Qwen3 VL 235BAlibaba Qwen
Open-weight multimodal model; tracked figure now matches the FIRST-PARTY Alibaba endpoint. 20/08/2026: OpenRouter endpoints for qwen/qwen3-vl-235b-a22b-thinking list Alibaba at $0.40/$4.00 (131,072 ctx) and Novita at $0.98/$3.95; the AA API host median moved to $0.40/$4.00 on the same day. The $0.70/$8.40 carried since 26/07 is no longer supported by any live source and has been corrected down. The instruct (non-thinking) variant is cheaper again at $0.21/$1.90 on OpenRouter and $0.40/$1.60 on AA - this row tracks the thinking variant. Cached-read $0.10 is inherited from the earlier Alibaba listing and is not re-confirmed by either source today.
NVIDIA's second row here, and the first one that wins something. Released 11/08/2026 as the smallest member of the Nemotron 3 family: 30B total, 3B active, interleaved Mamba-2 and MoE layers with select attention, 1M context, weights AND training data AND recipes under OpenMDW-1.1. AA scores it 23.6 intelligence at $0.088 blended - modest on its own, but it clears 276 tokens/sec, roughly 1.7x the fastest comparable row here, and its cost per task is $0.052. That combination is the point: NVIDIA is not selling this as a thinker, it is selling it as the execution layer under a frontier planner - tool calls, result validation, subagent delegation - with NeMo Switchyard shipped the same day to route between the two. NVIDIA's own number is 86% on PinchBench, completing 10,000 tasks 30% faster than Qwen3.6 35B at comparable accuracy; treat that as vendor-measured. Caveats: the agentic column is empty because AA has no tau-2 or TerminalBench Hard score for it yet (TerminalBench v2.1 0.243 is what exists), and the tracked $0.05/$0.20 is the hosted route on OpenRouter, matching AA's host median exactly - NVIDIA sells no per-token price of its own. There is also a free OpenRouter tier and a free trial on build.nvidia.com. Runs on a single GPU: Jetson, RTX 5090, DGX Spark, via LM Studio, llama.cpp, Ollama or Unsloth, with NVFP4 and BF16 checkpoints on day one. Input price rose $0.05 -> $0.08 per 1M between 12/08 and 19/08/2026 on the OpenRouter listing cited above; output unchanged at $0.20. AA's cost per task moved 0.052 -> 0.064 in step.
$0.08
$0.20
—
$0.06
1000K
23.6
26.8
—
GLM-4.7 FlashZ.AI
Open-weight model (Z.AI), multi-host. VERIFY FLAG CLOSED 26/08/2026: the 1c input gap flagged since 04/08 is gone - tracked row was moved to $0.07/$0.40 and the AA v2 API default host reads $0.07/$0.40 today, an exact match. Multi-host spread still applies, so treat the figure as the representative Z.AI-hosted price rather than a single contractual rate. Cost per task stays null: glm-4.7-flash is not in the AA capabilities subset that still renders.
$0.07
$0.40
$0.01
—
203K
23.3
—
60.4
Llama 4 ScoutMeta (hosted)
Open-weight (Meta), multi-host. Re-verified 24/08/2026 against OpenRouter's endpoint list: DeepInfra $0.10/$0.30, Groq $0.11/$0.34, Novita $0.18/$0.59, Google $0.25/$0.70. The tracked $0.18/$0.66 sits at the upper-middle of that spread and matches the AA host median exactly; the cheapest route is roughly half of it. Unchanged, but readers self-hosting or routing to DeepInfra should expect ~$0.10/$0.30.
$0.18
$0.66
—
$0.02
328K
10.3
8.2
8.5
Qwen3 30B A3BAlibaba Qwen
Open-weight (Alibaba Qwen), multi-host. VERIFY FLAG CLOSED 24/08/2026: the gap flagged since 28/07 was a variant mismatch, not a price error. Our row tracks the THINKING-2507 variant (AA slug qwen3-30b-a3b-instruct-reasoning); OpenRouter's endpoint list for qwen/qwen3-30b-a3b-thinking-2507 shows a single first-party Alibaba endpoint at exactly $0.20/$2.40 (81,920 ctx), matching the tracked figure. The $0.048/$0.193 previously carried is the cheapest host (StreamLake) of the separate INSTRUCT-2507 row, and the $0.12-$0.13 / $0.50-$0.52 spread belongs to the base qwen3-30b-a3b row. Three different models, three price bands.
$0.20
$2.40
$0.10
$0.11
262K
9.2
—
14.1
No affiliate links. Weight classes come from published
capability indices, not marketing tiers. Capability & cost data: Artificial Analysis