Market
“Who's actually being used?” has no single honest answer — every public source measures a different slice. So this page shows four slices side by side and prints what each one can and cannot see. Where they agree, believe it.
Data fetched 17/08/2026.
Token share on OpenRouter
Deepseek V4 Flash is eating 21% of the tokens developers route through OpenRouter — 75.3T tokens in the window (10/08–16/08).
Caveat: Only traffic routed through OpenRouter — a slice of total inference, skewed toward developers building multi-model apps. Source: OpenRouter rankings, refreshed daily.
Table view
| Model | Share | Tokens |
|---|---|---|
| Deepseek V4 Flash (deepseek) | 21.3% | 16.0T |
| Hy3 (tencent) | 13.2% | 10.0T |
| Gpt 5.6 Luna (openai) | 7.1% | 5.3T |
| Glm 5.2 (z-ai) | 5.7% | 4.3T |
| Mimo V2.5 (xiaomi) | 5.0% | 3.8T |
| Deepseek V4 Pro (deepseek) | 4.5% | 3.4T |
| Claude Opus 5 (anthropic) | 3.5% | 2.7T |
| Gemini 3.6 Flash (google) | 3.0% | 2.3T |
| Nemotron 3 Ultra 550b A55b (nvidia) | 2.9% | 2.2T |
| Laguna S 2.1 (poolside) | 2.3% | 1.7T |
| Minimax M3 (minimax) | 2.0% | 1.5T |
| Kimi K3 (moonshotai) | 1.9% | 1.4T |
| Claude Sonnet 5 (anthropic) | 1.4% | 1.1T |
| Step 3.7 Flash (stepfun) | 1.3% | 965.9B |
| Gpt 5.6 Terra (openai) | 1.2% | 930.9B |
| Gemini 3 Flash Preview (google) | 1.1% | 840.3B |
| Claude 4.6 Sonnet (anthropic) | 1.0% | 735.6B |
| Gemini 2.5 Flash Lite (google) | 0.9% | 712.8B |
| Gpt 5.6 Sol (openai) | 0.9% | 706.5B |
| Nemotron 3.5 Lightning (nvidia) | 0.9% | 673.7B |
| Gpt 5.6 Luna Pro (openai) | 0.8% | 616.6B |
| Claude 4.8 Opus (anthropic) | 0.8% | 589.5B |
| Mimo V2.5 Pro (xiaomi) | 0.6% | 480.6B |
| Gemma 4 31b It (google) | 0.6% | 477.3B |
| Gemini 2.5 Flash (google) | 0.6% | 465.2B |
What the crowd prefers (LMArena)
claude-opus-5-max leads the blind-vote arena at 1508 — the axis is clamped because ten points here separate the whole frontier.
Caveat: Crowd preference votes from self-selected testers — a popularity signal, not deployment volume. Source: LMArena leaderboard dataset (CC-BY-4.0), board of 2026-08-12.
Table view
| Model | Score | Votes | Org |
|---|---|---|---|
| claude-opus-5-max | 1508 | 10k | anthropic |
| claude-opus-5-high | 1505 | 20k | anthropic |
| claude-opus-4-6-high | 1503 | 72k | anthropic |
| claude-opus-4-6 | 1497 | 76k | anthropic |
| claude-fable-5 | 1493 | 21k | anthropic |
| qwen3.8-max | 1492 | 7k | alibaba |
| claude-opus-4-7-high | 1490 | 60k | anthropic |
| muse-spark-1.2 (xHigh) | 1488 | 3k | meta |
| claude-opus-4-7 | 1483 | 61k | anthropic |
| gemini-3.5-flash-high | 1482 | 26k | |
| gemini-3.6-flash-high | 1480 | 14k | |
| gemini-3.1-pro-preview | 1480 | 95k | |
| gemini-3-pro | 1479 | 42k | |
| muse-spark-1.1 | 1477 | 17k | meta |
| kimi-k3-max | 1476 | 12k | moonshot |
| gemini-3.5-flash-medium | 1476 | 24k | |
| qwen3.7-max-preview | 1474 | 4k | alibaba |
| muse-spark | 1473 | 14k | meta |
| qwen3.5-max-preview | 1471 | 22k | alibaba |
| gpt-5.5-high | 1471 | 55k | openai |
| gpt-5.4-high | 1470 | 61k | openai |
| ernie-5.1 | 1468 | 37k | baidu |
| gpt-5.5 | 1466 | 57k | openai |
| gemini-3-flash | 1466 | 31k | |
| glm-5.2-max | 1465 | 27k | zai |
What gets pulled (Hugging Face downloads)
The most-downloaded open weights are tiny models — the workhorses of self-hosting are not the models the leaderboards talk about.
Caveat: Open-weight downloads only; says nothing about how often a model runs, and excludes GPT/Claude/Gemini entirely. Source: Hugging Face Hub API, rolling 30-day counters.
Table view
| Model | Downloads 30d |
|---|---|
| Qwen/Qwen3-0.6B | 29.2M |
| facebook/opt-125m | 17.0M |
| Qwen/Qwen3-8B | 16.0M |
| trl-internal-testing/tiny-Qwen2ForCausalLM-2.5 | 15.6M |
| openai-community/gpt2 | 13.7M |
| Qwen/Qwen2.5-7B-Instruct | 12.3M |
| nvidia/Qwen3.6-35B-A3B-NVFP4 | 12.2M |
| Qwen/Qwen2.5-1.5B-Instruct | 11.0M |
| unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF | 9.7M |
| meta-llama/Llama-3.2-1B-Instruct | 8.9M |
| Qwen/Qwen3-Embedding-0.6B | 8.1M |
| openai/gpt-oss-20b | 8.0M |
| meta-llama/Llama-3.1-8B-Instruct | 7.5M |
| deepseek-ai/DeepSeek-R1 | 7.5M |
| Qwen/Qwen3-32B | 6.9M |
| Qwen/Qwen3-1.7B | 6.9M |
| Qwen/Qwen2.5-0.5B-Instruct | 6.8M |
| Qwen/Qwen2.5-3B-Instruct | 6.8M |
| farbodtavakkoli/OTel-2.0-LLM-31B-IT | 5.3M |
| google/gemma-3-1b-it | 5.0M |
Anthropic Economic Index (watcher)
The deepest public dataset on real assistant usage — single-vendor, so it sits in its own panel.
Anthropic's Economic Index maps what people actually DO with Claude — task mix, augmentation vs automation, by occupation. Latest dataset release: 26/06/2026. We watch the dataset daily and flag the next drop here.
Caveat: Single-vendor self-reported data — depth on one model, not a neutral comparison. Source: HF dataset metadata, polled daily.
The weekly reads
What the surveys and spend trackers say — refreshed weekly, not live.
The weekly deep-research digest lands here — enterprise spend (Menlo), SMB card spend (Ramp), consumer app rankings (a16z), vendor self-reports and the annual indexes, each with its own caveat. First run pending.
Caveat: VC surveys, card-spend panels and modeled traffic measure who PAYS or VISITS — none of them measures query volume, and all of them are US-weighted. Source: weekly research task (Sundays).