Four model numbers this week: 10 trillion, 2.4 trillion, 16 billion, 2.8 billion. Read them as a leaderboard and you get a size race with China at one end and a budget tier at the other. That reading is wrong, and it will cost you money.

The verdict: all four releases bet on the same technique, and three of the four numbers are marketing. The only figure that predicts your bill, your latency or your hardware is 2.8 billion — the parameters that actually fire per token. Total count has quietly stopped being a capability signal and become a storage requirement.

One race, dial set differently

The honest disclosure came from the small end. AMD's Instella-MoE-16B-A3B is fully open: 16 billion total, 2.8 billion active, both numbers in the model name. That is not modesty, it is a spec sheet. Same day, Thinking Machines shipped Inkling Small — under a third its flagship's size and stronger than it on coding. Meanwhile Alibaba put Qwen3.8-Max into GA at 2.4 trillion parameters and ByteDance is training at 10 trillion — both mixture-of-experts too, neither publishing the number that runs. The tell that this is settled engineering: Cursor open-sourced its MoE training megakernel. When a company gives away the tooling for the hard part, the technique stopped being its moat.

I sized heating systems for twenty years, so here is how it clicked for me: a mixture-of-experts model is a building with 128 radiators and one boiler. You pipe, valve and pay for every room — that is your memory. But only eight valves are open at once, and the boiler is sized for those eight — that is your compute. Two bills, two constraints, and the industry keeps quoting you the pipework when you asked about the boiler.

Watch the financing, not the technique

While the technique moved toward using less, the capital moved the other way. Citadel Securities forecasts over $500 billion of chip financing debt, and Google, Broadcom, Apollo, Blackstone and Morgan Stanley structured most of the Anthropic chip risk off Google's balance sheet. Read that twice: the most sophisticated compute buyer on earth kept the chips and handed someone else the risk. I am not calling a bubble. But when capability gets cheaper per unit while the financing gets more elaborate, watch the financing.

Concretely, for any business buying AI: stop asking which frontier model sits behind a subscription. Ask how many parameters activate per token and how much memory the whole model needs — those two decide whether the job runs on a machine in your own office instead of on a meter. AMD's release at 4-bit is roughly 10 GB; a 32 GB laptop holds it with room to work. Schema-shaped, high-volume work — invoices, delivery notes, contracts into structured fields, exactly what the new ExtractBench benchmark measures — was a cloud line item eighteen months ago and is now a purchase you make once. Rented frontier still wins the hardest reasoning, and it keeps getting cheaper too. The point: renting is now a choice, not a constraint.

My receipt, checked this morning: everything on this site is written and scored by GLM-4.5-Air at 4-bit on one 128 GB Mac. 128 routed experts, eight active. 53 posts live to @RXed_EU in seven days, total outside spend $1.86, all of it X's posting fee. The honest counter-number: those 53 posts earned 22 engagements. This is a cost proof, not a growth flex.

Two claims. Before end September I replace GLM-4.5-Air here with an open model under 20 billion active parameters at equal or better quality on my own scorer, and publish the numbers either way. And by end September at least one trillion-tier lab publishes its activated-parameter count, because the small models made that number a selling point and silence starts to read as an answer. Last week's two claims stay on the board.