AI Costs Aren’t a Pricing Problem
The evergreen discipline separating organizations that succeed with AI from companies that just spend on it.
For a while, GitHub Copilot was easy to budget.
Copilot Enterprise costs $39 per user per month. Multiply by headcount, and you had a single line item for a budget. That predictability helped make AI feel financially safe to deploy across an organization. #tokenmaxxing
Then, on June 1, 2026, GitHub changed the economics.
Copilot plans now include AI credits, with usage tied to the tokens different models consume. For developers accustomed to a predictable subscription, the shift was jarring. Some heavy users projected dramatically higher costs, and the backlash arrived almost immediately.
But GitHub isn’t really the story.
It is an early glimpse of the economics that will govern enterprise AI.
When Cheaper Gets More Expensive
Cloud computing taught companies what happens when technology shifts from fixed pricing to consumption-based pricing. The bill starts to reflect how the system was built.
A fixed license mostly asks one question: how many people need access. AI consumption asks a dozen. Which model handled the request? How much context it received? How often it was called? Whether the result was cached? Whether the workflow retried? Whether a model was needed at all?
Each answer looks small. At scale, they add up to a significant expense.
That’s the strange economics of generative AI: tokens can get cheaper while AI gets more expensive to operate. As inference prices fall, consumption climbs to meet them. More employees get access. More products grow AI features. Models reason longer. Agent software factories run workflows that can take dozens of calls to finish a single task. Lower prices don’t guarantee lower costs when usage expands faster than efficiency does.
The market is already responding. Stripe recently acquired OpenRouter for more than $7 billion, a company valued at $1.3 billion just months earlier. OpenRouter’s job is to help developers route work across models and manage the economics around them.
The signal is hard to miss. There’s growing value in deciding what intelligence to use, when to use it, and how much a task actually needs.
That makes AI cost a design problem first, and a pricing problem second.
Can’t Penny Pinch Out Of This
When AI spending climbs, the instinct is to negotiate a lower rate, restrict the expensive models, buy committed capacity, or chase a volume discount.
Those measures can help. But they don’t address the systemic conditions that created the bill.
If a workflow makes fifteen model calls when five would do, a cheaper token doesn’t fix it. If a request sends 100,000 tokens of context when it needs 5,000, procurement can’t negotiate away the waste. If a frontier model is doing work that deterministic code could handle, the model’s price is almost beside the point.
I’ve started seeing this pattern move upstream, into the business case itself. More and more, organizations are asked to estimate and justify the token cost of an AI initiative before anyone has built it.
I understand the impulse. A cost estimate is a reasonable thing to want. But a raw token estimate is close to meaningless, because consumption depends on model choice, context size, caching, routing, retrieval, retries, agent behavior, and how people actually use the thing once it ships. None of that is known at planning time. So the number is a guess dressed as a forecast, and worse, it rewards sandbagging.
The better question was never “how much will this cost?”
It’s “what does one successful outcome cost, and which design decisions determine that number?”
That single change reframes the whole conversation. Now the questions worth asking are design constraints. Could a smaller model handle this? Could ordinary code handle part of it? Can the system retrieve only the context it needs, instead of everything it has? Can results be cached? Should the agent really try again? What gain in quality actually justifies more computation?
The useful unit of AI economics isn’t the token. It’s the outcome.
The Good News Is Also The Bad News
The real danger is dismissing the design steps as something to negotiate or wait out, and the expensive choices become defaults. Large context windows get copied into new workflows. Frontier models become the standard because nobody built routing. Agents loop freely because nobody set limits. Teams inherit the same patterns through shared platforms and templates. None of it is a decision. It’s just what happens when the design work gets skipped.
And as adoption grows, those choices spread with it. The waste isn’t added once. It’s copied into everything built next.
Here’s the prediction. Organizations that skip the design work won’t just spend more. Their spend will grow faster than the value they get for it. Every new integration inherits the same waste, so the bill climbs with adoption while the return on each dollar falls. Leadership will see rising usage and read it as success, when what’s actually rising is the cost of the same outcomes.
Falling token prices will hide it the whole way. The unit price drops, so no one looks closer, while total consumption climbs underneath. The rollout succeeds, usage expands, and the inefficiency scales right along with it. The better it looks, the worse the economics get.
Cloud computing went through this same transition. Variable infrastructure spending is what produced FinOps, born from the recognition that technology cost can’t be managed by finance or procurement alone. AI is accelerating that shift. The 2026 State of FinOps reports that 98 percent of practitioners now manage AI spend, up from 31 percent two years ago, and that AI cost management is the top skill those teams say they need to build.
What’s Next?
You won’t win this by spending the least. You’ll win by knowing why you’re spending. Cheaper Deliberate intelligence.
Nobody decides to waste money on AI. When you can prototype in minutes and have a working app by lunch, it’s no surprise design decisions get made by default instead of on purpose.
So ask the right question now, while it’s still cheap to answer: where is the intelligence earning its keep, and where is it just running?
Curious where your gaps are?
Our AI FinOps Assessment shows which design gaps may be driving your token spend, and which levers could recover it. ~4 minutes


