The Hard Part Is No Longer Training the Model. It Is Serving It.
Moonshot shipped a competitive model and ran out of GPUs in 48 hours. The most interesting companies in AI right now sit between the chip and the model, and almost nobody is pricing them.
Moonshot AI launched Kimi K3 and stopped accepting new subscribers within 48 hours. Not a safety pause. Not a rollback. The GPUs were full.
That is the most clarifying data point in AI this month, because it pulls apart two things the industry keeps treating as one: having a frontier model, and being able to serve one.
## Benchmarks measure the wrong scarcity
For three years the scoreboard has been training. Who has the biggest cluster, the longest run, the best eval numbers. That framing quietly assumed the second half was easy: once you have the weights, you serve them.
Moonshot just demonstrated that the second half is where the wall is. The model was good enough to saturate its own fleet in two days. Quality was never the constraint. Serving capacity was.
This is also the sharpest evidence yet for how export controls actually bite. They do not stop a lab from building a competitive model, as the open-weight surge out of China keeps showing. They stop that lab from putting the model in front of users at scale. A capability you cannot serve is a demo.
## The money noticed before the discourse did
Look at where early-stage capital went this month and a pattern shows up that has nothing to do with model labs.
Infinity.inc raised a $15M seed at a $100M post-money valuation for an agent that automatically generates optimized inference software for any new AI chip. Earlier this month Oxmiq raised $35M for licensable GPU IP.
Neither company trains anything. Both sit in the gap between silicon and model, and that gap is the real moat in this market. CUDA is not dominant because the hardware is untouchable. It is dominant because every competing accelerator arrives without a mature software stack, and hand-writing one takes years you do not have. Compress that work into an automated pipeline and you have not built a better chip. You have made everyone else's chip usable, which is worth more.
## Capacity is not the same as throughput
Meanwhile the capital at the other end of the stack keeps getting larger and slower. Hut 8 signed a $9.8 billion, 15-year lease for 352 MW in Texas. Japan committed roughly $6.2 billion over five years to its Noetra sovereign program and a national AI factory.
Those are real commitments and they will matter. But note the units. Megawatts and years, versus a signup queue that filled in 48 hours. Racks take quarters to land and power contracts take longer. The serving layer is where you can find capacity this month rather than in 2029, which is exactly why the software between chip and model is repricing now.
The uncomfortable version: a lot of announced gigawatts will come online attached to silicon whose software stack is still immature. Whoever closes that gap captures value that the capex headlines are currently assigning to whoever poured the concrete.
## What to do with this
**Change the question you ask a model lab.** Benchmark scores are close to free now. Ask for sustained tokens per second at peak load, and what happened the last time demand spiked. Moonshot answered that question involuntarily, which makes it more honest than most disclosures. Anyone who cannot answer it is telling you they have not hit the wall yet.
**Watch the K3 reopening date.** How long it takes Moonshot to resume signups is a cleaner read on real Chinese compute availability than any policy analysis or benchmark table. Fast reopen means the controls leak more than assumed. A long silence means they bind hard. Either way you learn something that no leaderboard reports.
You can track the full flow of AI infrastructure capital in our deal tracker.