GPT-5.6's 10x Price Cut Drove a 13x Usage Surge

Source: 20VC with Harry Stebbings | Published: 2026-08-10T14:05:48Z

OpenRouter CEO Alex Atallah revealed that Luna's usage grew 13x after a 10x price cut on the platform — a sustained climb, not a temporary spike — landing it among the top five models by token volume for the first time.


GPT-5.6's price dropped 10x. Usage went up 13x. OpenRouter CEO Alex Atallah dropped those numbers casually in an interview, but they precisely capture why he believes the routing layer isn't going anywhere — when cost stops being a barrier, people consume more inference, not less.


A Complete Experiment in the Jevons Effect

Luna (GPT-5.6) was the subject of a rare natural experiment on OpenRouter: OpenAI cut the price 5x, then partnered with OpenRouter to cut it another 2x — a total 10x reduction within two weeks.

Usage went up 13x. Not a brief spike, but a sustained 13x baseline that then continued climbing along its original growth curve. It was a remarkably clean experiment — DeepSeek and GLM were also cutting prices at the same time, which should have pulled traffic away. Instead, Luna surpassed GLM to become one of the top three to five models on OpenRouter by token volume — the first time an OpenAI model had broken into that position on the platform.

Atallah calls this the "Jevons Paradox" — improvements in energy efficiency lead to increased energy consumption. For a business like OpenRouter that takes a cut of each transaction, this is good news: prices fall, but revenue can follow the volume.


Why the Inference Layer Wasn't Absorbed by the Hyperscalers

When OpenRouter first launched, Atallah's read on the inference provider layer was completely different — he assumed the three major hyperscalers would monopolize open-source model serving. What actually happened: companies like Fireworks and Together, focused purely on inference, ran circles around the hyperscalers. They shipped models faster and handled edge cases better.

There's a reason for this that rarely gets mentioned: Nvidia doesn't want customer concentration.

One of Nvidia's priorities is avoiding a situation where a handful of large customers absorb all the GPUs. Spreading capacity across a large number of inference providers — each taking a slice of the allocation — is Nvidia's ideal outcome, because it creates competition at the compute layer. For users, it also means that the same model running on different providers can produce measurably different results.

Atallah used Kimi K3 as an example. Moonshot published a comparison chart of service quality across providers, and the numbers diverged significantly — even on static, well-tested benchmarks, different providers reached different conclusions. OpenRouter updates its routing weights every five minutes. When a provider gets faster, cuts prices, or improves quality, traffic shifts immediately.


"Copying Isn't a Winning Strategy"

Over the past year, more and more companies have been building routing layers — well-funded startups, and established players adding routing features to existing products. Atallah's take is blunt: most of them are building routing because it's fashionable.

His concern isn't competition. It's that these entrants are optimizing to exist, not to win. "You're existing to exist, rather than existing to win." More concretely: a company that treats routing as a side feature offers users a limited set of choices — you can only access the models it supports, which means you're throwing away leverage that your users haven't even discovered yet.

OpenRouter's logic: giving users more choices means giving them more leverage, and covering the entire market is the whole point of the product. Once you narrow it down to a subset serving a particular customer type, the value proposition collapses.


The Gap With Chinese Open-Source Models May Keep Widening

Atallah's read: the U.S. is already behind, and the gap may keep growing.

His reasoning isn't technical — it's structural. When DeepSeek becomes a national-level project, the underlying logic is "put everything behind this horse." Regulatory obstacles get cleared. Policy aligns. Capital floods in. Researchers operate without constraint. Every new Chinese model lab that emerges gets this same institutional backing.

By contrast, U.S. open-source models are trying to survive within commercial logic — open-sourcing makes business models harder to design, and competing with OpenAI and Anthropic means extremely high research costs. He cited Poolside as one of the few hopeful signs he sees, but noted that the fundraising path remains far harder than what Chinese competitors face.

On GLM 5.2, he sees a genuine leap for open-source models. On Kimi, it's more of a "catching up to that level." He's always thought Kimi's writing quality was solid — the tone and voice are clean — while he's noticed that many frontier models, after improving their coding capabilities, have become increasingly painful to read: he noted that many frontier models now read like padding machines — three solid points, a pivot sentence, then filler.


It's Silicon Valley Companies That Enterprises Actually Fear

Atallah says the enterprise clients he encounters are more anxious about U.S. frontier models than Chinese models. That sounds counterintuitive, but the logic holds:

Chinese open-source models can be deployed on your own machines, you can choose the provider, and data flows are relatively controllable. But Claude, GPT, and their frontier counterparts can't run on private infrastructure. Data storage and ingestion policies aren't transparent, and the familiar enterprise question — "where is my data, who can see it" — has no definitive answer.

The emergence of Claude Design made another concern concrete: once a model lab starts serving a specific internal team directly — say, the design team — that team becomes Anthropic's user, not just an employee of their company. If Anthropic continues building dependency relationships across functional teams, the cost of that reliance accumulates fast.

On whether Figma is threatened by Claude Design, Atallah hedged: "Our designers tried it, but I haven't heard feedback about repeat usage." He did note in passing that Figma's financials are "incredibly, incredibly impressive" — then the conversation drifted to why you can't seem to have a company with both good numbers and a good stock price at the same time.


No Single Layer Can Monopolize Memory

Both the open-source community and product companies assume memory is the key retention mechanism — OpenAI knowing you live in London and do a podcast makes it more attuned to you than competitors. Atallah thinks the premise is shaky: where memory lives was never settled.

The model layer, inference layer, application layer, routing layer — every layer is trying to own memory, and each one has its own distinct contextual advantage. The application knows what you've done inside that product. The model knows how you express yourself. The routing layer knows cross-model usage patterns.

The hardest question: if you put a copy of memory in the model layer and another in the application layer, and both activate simultaneously, does it make the model more confused? He admits there's no answer yet. His view is that no single layer can monopolize all the valuable memory, because the context applications hold is something model labs can't access — unless applications voluntarily hand over the data, which requires model labs to offer sufficient incentive in return.


70 Models, One Every 10 Hours

In July alone, OpenRouter added 70 models — roughly one new model every 10 hours.

Atallah thinks this pace will accelerate. Cognition has a model. Cursor has a model. Lovable doesn't yet, but probably will soon. Every company building agents, once it reaches sufficient scale, has strong incentive to train its own model and distribute it through the agent.

Add Nvidia's active push for compute-layer diversity, add investors wanting to spread bets across different trajectories, and Atallah sees the number of new model labs increasing, not decreasing. He disagrees with the "70% of new model labs will be dead in three years" framing — if you count acquisitions, he'll accept 50%.


Distillation Isn't Cheating

One technical topic got a passing clarification. Distillation — using a large model's outputs to fine-tune another model — is sometimes characterized externally as "lazy" or "just copying." Atallah's take: it's standard practice. Sonnet itself is partly distilled from Opus. Every model lab does it.

His view: for the U.S., distilling Chinese open-source models is actually a viable catch-up path, because most Chinese open-source models permit distillation, and the distillation process lets you inspect outputs and audit alignment — an added layer of control compared to using Chinese models directly.

Closed-source model labs, of course, have the right to prohibit using their outputs to train competing models in their terms of service. That's a separate issue.


Employee Cost Is Becoming a Dynamic Number

At the end of a rapid-fire Q&A, Atallah raised something he thinks isn't getting nearly enough serious attention: companies haven't figured out how to manage employee costs in the AI era.

Historically, employee cost was static — salary is a fixed number, adjusted at quarterly reviews. Now, how much inference an employee consumes each day, which models they use, whether they're using it effectively — all of these decisions directly affect the company's operating costs. An employee's actual cost is becoming a real-time, fluctuating number.

His suggestion: put employee inference efficiency and work output on the same axis. Look at both output quality and AI usage cost. High output, low cost lands in one quadrant; mediocre output with runaway costs lands in another — and address the latter directly. The upside: every employee can control their own cost, which follows a fundamentally different logic than fixed compensation.

More articles on TLDRio