The New AI Pitch Is Not Smarter. It Is 1,664 Tokens Per Second.
Three launches in three weeks led with speed and price instead of capability. That is what happens when agents, not chatbots, become the workload. Here is the number to ask vendors for, and the tell that says the capability gap has already closed.
A San Francisco startup called Celeris announced a model on July 27 that it says generates around 1,664 tokens per second, up to 15 times faster than the alternatives, at near GPT-5 quality.
Read what the pitch does not say. It does not claim to be smarter.
That is the shift worth tracking this month. Across three weeks of launches and funding rounds in our tracker, the product claim moved from capability to throughput and price.
**Three data points, same direction**
Meta launched Muse Code on August 5, a terminal based coding agent running on its new Muse Spark 1.2 model. The company with the resources to claim the crown positioned it instead as the cheaper alternative to the agents from OpenAI and Anthropic.
Celeris sold latency, and said so explicitly: its API targets "latency sensitive workloads like agents and real time voice."
And Infinity raised a $15 million seed for Ignition, an agent that automatically generates optimized inference software stacks for new AI chips. The pitch is making any chip inference ready, which is a direct attack on the software moat that keeps inference expensive.
None of those three is a capability bet. All three are cost bets.
**Why the workload changed underneath them**
A chatbot is one call. You ask, it answers, the meter stops.
An agent is a loop. It plans, calls a tool, reads the result, revises, and goes again, often hundreds of times per task, usually with nobody watching. The same task that cost one inference call in 2024 costs a few hundred in 2026.
Look at what is being funded on that assumption. KelAI's seed round backs an engine that runs the entire hedge fund research loop, from idea generation through live signal monitoring. That is not a query. That is a process that never stops.
Then scale it to a state. The UAE set a target of moving 50 percent of federal government sectors, operations and services onto agentic AI within two years, with 80,000 employees trained alongside it.
At that volume the binding constraint stops being how clever the model is per token. It becomes how many tokens you can afford per hour.
**The hedge**
Two cautions before you trade on this.
The 15x figure is a vendor claim from a press release, not an independent benchmark, and diffusion based language models are early. Treat the direction as the signal, not the multiple.
And cheaper inference does not automatically squeeze the frontier labs. It usually expands usage instead, which is part of why Nvidia now sits on four sides of its own demand. Falling unit cost and rising total spend have coexisted in this industry for three years.
**What to do with this**
Ask vendors for cost per completed task, not cost per million tokens or a benchmark score. A cheap model that loops twenty times to finish a job is more expensive than a pricey one that finishes in three, and only the per task number exposes that.
Then watch one thing this quarter: whether OpenAI or Anthropic answers Meta on price. If they cut, it means they read the capability gap as closed too, and the agent market becomes an infrastructure margin business faster than anyone's roadmap assumes.