RXed AI News

AI to the bone.

Market

“Who's actually being used?” has no single honest answer — every public source measures a different slice. So this page shows four slices side by side and prints what each one can and cannot see. Where they agree, believe it.

Data fetched 17/08/2026.

Token share on OpenRouter

Deepseek V4 Flash is eating 21% of the tokens developers route through OpenRouter — 75.3T tokens in the window (10/08–16/08).

Deepseek V4 Flash · deepseek21.3%Hy3 · tencent13.2%Gpt 5.6 Luna · openai7.1%Glm 5.2 · z-ai5.7%Mimo V2.5 · xiaomi5.0%Deepseek V4 Pro · deepseek4.5%Claude Opus 5 · anthropic3.5%Gemini 3.6 Flash · google3.0%Nemotron 3 Ultra 550b A55b · nvi2.9%Laguna S 2.1 · poolside2.3%Minimax M3 · minimax2.0%Kimi K3 · moonshotai1.9%Claude Sonnet 5 · anthropic1.4%Step 3.7 Flash · stepfun1.3%Gpt 5.6 Terra · openai1.2%

Caveat: Only traffic routed through OpenRouter — a slice of total inference, skewed toward developers building multi-model apps. Source: OpenRouter rankings, refreshed daily.

Table view
ModelShareTokens
Deepseek V4 Flash (deepseek)21.3%16.0T
Hy3 (tencent)13.2%10.0T
Gpt 5.6 Luna (openai)7.1%5.3T
Glm 5.2 (z-ai)5.7%4.3T
Mimo V2.5 (xiaomi)5.0%3.8T
Deepseek V4 Pro (deepseek)4.5%3.4T
Claude Opus 5 (anthropic)3.5%2.7T
Gemini 3.6 Flash (google)3.0%2.3T
Nemotron 3 Ultra 550b A55b (nvidia)2.9%2.2T
Laguna S 2.1 (poolside)2.3%1.7T
Minimax M3 (minimax)2.0%1.5T
Kimi K3 (moonshotai)1.9%1.4T
Claude Sonnet 5 (anthropic)1.4%1.1T
Step 3.7 Flash (stepfun)1.3%965.9B
Gpt 5.6 Terra (openai)1.2%930.9B
Gemini 3 Flash Preview (google)1.1%840.3B
Claude 4.6 Sonnet (anthropic)1.0%735.6B
Gemini 2.5 Flash Lite (google)0.9%712.8B
Gpt 5.6 Sol (openai)0.9%706.5B
Nemotron 3.5 Lightning (nvidia)0.9%673.7B
Gpt 5.6 Luna Pro (openai)0.8%616.6B
Claude 4.8 Opus (anthropic)0.8%589.5B
Mimo V2.5 Pro (xiaomi)0.6%480.6B
Gemma 4 31b It (google)0.6%477.3B
Gemini 2.5 Flash (google)0.6%465.2B

What the crowd prefers (LMArena)

claude-opus-5-max leads the blind-vote arena at 1508 — the axis is clamped because ten points here separate the whole frontier.

claude-opus-5-max1508claude-opus-5-high1505claude-opus-4-6-high1503claude-opus-4-61497claude-fable-51493qwen3.8-max1492claude-opus-4-7-high1490muse-spark-1.2 (xHigh)1488claude-opus-4-71483gemini-3.5-flash-high1482gemini-3.6-flash-high1480gemini-3.1-pro-preview1480gemini-3-pro1479muse-spark-1.11477kimi-k3-max1476arena score (axis clamped 1468–1516) →

Caveat: Crowd preference votes from self-selected testers — a popularity signal, not deployment volume. Source: LMArena leaderboard dataset (CC-BY-4.0), board of 2026-08-12.

Table view
ModelScoreVotesOrg
claude-opus-5-max150810kanthropic
claude-opus-5-high150520kanthropic
claude-opus-4-6-high150372kanthropic
claude-opus-4-6149776kanthropic
claude-fable-5149321kanthropic
qwen3.8-max14927kalibaba
claude-opus-4-7-high149060kanthropic
muse-spark-1.2 (xHigh)14883kmeta
claude-opus-4-7148361kanthropic
gemini-3.5-flash-high148226kgoogle
gemini-3.6-flash-high148014kgoogle
gemini-3.1-pro-preview148095kgoogle
gemini-3-pro147942kgoogle
muse-spark-1.1147717kmeta
kimi-k3-max147612kmoonshot
gemini-3.5-flash-medium147624kgoogle
qwen3.7-max-preview14744kalibaba
muse-spark147314kmeta
qwen3.5-max-preview147122kalibaba
gpt-5.5-high147155kopenai
gpt-5.4-high147061kopenai
ernie-5.1146837kbaidu
gpt-5.5146657kopenai
gemini-3-flash146631kgoogle
glm-5.2-max146527kzai

What gets pulled (Hugging Face downloads)

The most-downloaded open weights are tiny models — the workhorses of self-hosting are not the models the leaderboards talk about.

Qwen/Qwen3-0.6B29.2Mfacebook/opt-125m17.0MQwen/Qwen3-8B16.0Mtrl-internal-testing/tiny-Qwen2F15.6Mopenai-community/gpt213.7MQwen/Qwen2.5-7B-Instruct12.3Mnvidia/Qwen3.6-35B-A3B-NVFP412.2MQwen/Qwen2.5-1.5B-Instruct11.0Munsloth/Qwen3-Coder-30B-A3B-Inst9.7Mmeta-llama/Llama-3.2-1B-Instruct8.9MQwen/Qwen3-Embedding-0.6B8.1Mopenai/gpt-oss-20b8.0M

Caveat: Open-weight downloads only; says nothing about how often a model runs, and excludes GPT/Claude/Gemini entirely. Source: Hugging Face Hub API, rolling 30-day counters.

Table view
ModelDownloads 30d
Qwen/Qwen3-0.6B29.2M
facebook/opt-125m17.0M
Qwen/Qwen3-8B16.0M
trl-internal-testing/tiny-Qwen2ForCausalLM-2.515.6M
openai-community/gpt213.7M
Qwen/Qwen2.5-7B-Instruct12.3M
nvidia/Qwen3.6-35B-A3B-NVFP412.2M
Qwen/Qwen2.5-1.5B-Instruct11.0M
unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF9.7M
meta-llama/Llama-3.2-1B-Instruct8.9M
Qwen/Qwen3-Embedding-0.6B8.1M
openai/gpt-oss-20b8.0M
meta-llama/Llama-3.1-8B-Instruct7.5M
deepseek-ai/DeepSeek-R17.5M
Qwen/Qwen3-32B6.9M
Qwen/Qwen3-1.7B6.9M
Qwen/Qwen2.5-0.5B-Instruct6.8M
Qwen/Qwen2.5-3B-Instruct6.8M
farbodtavakkoli/OTel-2.0-LLM-31B-IT5.3M
google/gemma-3-1b-it5.0M

Anthropic Economic Index (watcher)

The deepest public dataset on real assistant usage — single-vendor, so it sits in its own panel.

Anthropic's Economic Index maps what people actually DO with Claude — task mix, augmentation vs automation, by occupation. Latest dataset release: 26/06/2026. We watch the dataset daily and flag the next drop here.

Dataset on Hugging Face → · June 2026 report →

Caveat: Single-vendor self-reported data — depth on one model, not a neutral comparison. Source: HF dataset metadata, polled daily.

The weekly reads

What the surveys and spend trackers say — refreshed weekly, not live.

The weekly deep-research digest lands here — enterprise spend (Menlo), SMB card spend (Ramp), consumer app rankings (a16z), vendor self-reports and the annual indexes, each with its own caveat. First run pending.

Caveat: VC surveys, card-spend panels and modeled traffic measure who PAYS or VISITS — none of them measures query volume, and all of them are US-weighted. Source: weekly research task (Sundays).