Every problem has a solution, and as we mature through the age of AI, it is inevitable that we start seeing clients become more savvy, more tools to help with reducing costs of usage, as well as tools to help identify where AI is being used, such as panagram here on this platform. With the Chinese rapidly producing cheaper and more efficient models, the race is on and the competition is wide open.

As users, do we really understand the cost of tokens? There is a kind of formula for it but with AI being applied to different tasks using different models, are we aware of what that prompt is costing us or how much it is costing us to deliver that task? It almost always comes as a slight surprise when we run out of tokens and have to decide whether to upgrade, top up or wait until they re-set.

Where Support Partners differs from what we are seeing in the market, comes from our longstanding mission of actively working to reduce production overheads and workflow inefficiencies. In this era of AI and Agentic technology this expands to models and tokens, specifically, the cost per deliverable for media production, which can silently creep up and throw budgets well off course.

In enterprise environments, this is not as directly visible as in the consumer space. Although AI token prices have fallen roughly 90% since 2023, the cost of enterprise AI bills have tripled or more over the same period, because agentic workflows now consume 5–30 times more tokens per task than simple chatbot use.

The solution we are seeing emerging is of course, agentic. Enterprises are adopting is agentic model routing: software that automatically picks the cheapest capable model for each step of a task instead of defaulting every request to the most expensive model available. In smaller environments, it the task falls to educating the users on what models are suitable for which tasks and attempting to involve critical thinking for human-led discernment, perhaps embedded within an AI policy. Does this necessarily produce better value?

Is our experience then akin to the IKEA effect?

You go in for one thing and end up with a cart full of random stuff you might not really need, confused because the name is in Swedish and not sure what it means, and now you get home to assemble it and it is missing a nut (plus have to fit it in your car like a game of jenga). The instructions seemed simple enough.

Token prices and total AI bills are moving in opposite directions, and both trends are real at the same time. The blended cost of a million AI tokens fell about 67% year over year, from roughly $18.40 to $6.07 between Q1 2025 and Q1 2026. Some estimates put the long-run drop even higher; tokens that cost around $60 per million in 2021 now go for pennies. By any normal reading of a market, AI should be getting cheaper to run.

This misconception is one that takes CFO’s entirely by surprise. A bit like the final price of your cart at the IKEA check out.

The average enterprise AI budget has grown from about $1.2 million a year in 2024 to roughly $7 million in 2026, and 73% of enterprises say their AI costs blew past what they’d projected. Uber’s engineers went from 32% AI-tool adoption to 84% in four months, and their monthly per-engineer AI bill went from $500 to $2,000, with, in the COO’s words, no ceiling in sight. One unnamed enterprise reportedly ran up $500 million in a single month on Claude simply because nobody had switched on spending caps.

This is Token Mania: a moment in time where the unit price of intelligence is in freefall and the total bill keeps climbing anyway. The choices are mind-boggling to muddy the waters.

We live in truly paradoxical times, research shows that introducing AI in the workplace is actually forcing employees to work harder, instead of making their jobs easier. [Futurism]. Or are they generating more data and information than actually needed? Just because we can doesn’t mean we should be generating for generating’s sake.

The latest comes from a new analysis from ActivTrak of over 164,000 workers’ digital work activity. After examining their activity 180 days before and after the employees started using AI at work, the software company found that AI “intensified” their jobs in nearly every category, the Wall Street Journal reported. The time they spent on email, messaging, and chat apps more than doubled, while their use of business software surged by 94 percent. More AI, more tokens = higher bills.

What is the Jevons paradox and why does it explain rising AI costs?

The Jevons paradox (the Coal question) describes this economic pattern whereby making something radically cheaper increases total consumption of it rather than reducing total spend, because cheaper access unlocks far more use cases than existed before. Except that it is not all cheap.

Apollo’s chief economist Torsten Slok has cited it directly to explain 2026’s AI spending: as the cost per unit of intelligence collapses, companies don’t bank the savings, they run more agents, automate more workflows, and generate more code. Cheaper tokens don’t reduce the bill. They expand what gets automated.

What are the reports on this particular cost-utilisation dilemma?

EY’s fifth-wave US AI Pulse Survey, fielded in spring 2026 across 534 senior US decision-makers, found that 82% of leaders whose organizations invest in AI are concerned about token usage and its costs, and a near-universal 98% say token costs have caused their organization to rethink its approach. Yet only 64% say their organization actually monitors token usage with clear budgetary guardrails in place, meaning most of the concern hasn’t yet turned materialised into control mechanisms.

Surprisingly this doesn’t mean that companies are encouraging less usage. EY found that 37% of leaders are actually expanding the scope of their AI rollout because of rising token costs, compared with just 15% scaling back, and more are speeding up their pace of deployment than slowing it down.

Another result from this particular survey reveals that 91% of leaders now see building AI software in-house as critical to controlling both cost and dependency on vendors, and 76% say off-the-shelf software, or “bolted on” no longer fits their needs. The understanding of how many tokens it takes to run a particular workflow is now the key way to getting a grip on what costs should be, instead of receiving a surprise invoice. Owning that internally, sometimes on the edge, is a way to rationalise the consumption by routing to the most appropriate models for the task.

Why does defaulting to “the best model” waste money?

At a point in time where there isn’t widespread understanding of the capabilities of each model, as more and more appear available, so defaulting every task to the most expensive available model is really a waste of money because most tasks: classification, simple drafting, tagging, routine lookups, don’t require frontier-level reasoning. Faced with a menu of models ranging from ultra-cheap small models (some priced under a dollar per million tokens) to frontier reasoning models costing well over a hundred dollars per million, the path of least resistance is to default to “the best model” for everything: every summarization, every classification task, every step of an agent’s reasoning loop. It feels safe. It is also enormously wasteful, because most of what an agent does in a given task doesn’t need frontier-level reasoning at all.

You are burning gas driving your Ferrari to the nearest store when you could have cycled or walked.

How are Chinese open-weight models like Kimi K3 affecting AI token prices?

Chinese open-weight models are compressing global AI token prices by offering frontier-competitive performance at a fraction of Western frontier pricing, forcing US labs to cut prices in response. Moonshot AI’s Kimi is the clearest example. Kimi K3, released in April 2026, runs at roughly $0.60 per million input tokens and $2.50 per million output tokens, priced openly under a modified MIT license, meaning any company can download the weights and run the model on its own infrastructure entirely free of per-token cost. Its successor, Kimi K, a 2.8-trillion-parameter model unveiled in July 2026 and the largest open-weight release to date, reportedly approaches the coding performance of leading US frontier systems while completing several times as much work per dollar spent. DeepSeek has moved just as aggressively: its V4-Flash model recently cut token costs by half again, and one comparison put its output pricing at roughly 1% of what a Western frontier model charges for the same task. Chinese models now occupy six of the top ten slots on OpenRouter, a widely used marketplace for model access.

The knock-on effect is a genuine price war rippling through the whole market. OpenAI cut the price of one of its efficiency-tier models by 80% just three weeks after launch. Google has released a string of new low-cost “flash” models. Whether or not an enterprise ever touches a Chinese model directly, this competition is a real reason the blended per-token price across the industry keeps falling; pricing pressure from below is doing some of the same work that internal cost discipline is meant to do.

However, all is not what it seems, and there is a lot of smoke and mirrors at play here. It’s worth noting that raw per-token price is a misleading comparison on its own, because models differ in how many tokens they burn to reach a given answer; a cheaper model that’s more verbose can end up costing about the same to actually complete a task. Some Chinese models also run slower in benchmarks even where the sticker price is lower. And using a foreign-hosted or self-hosted open-weight model for production workloads raises separate questions around data governance, compliance, and support that a token-cost comparison alone won’t answer.

If we use agentic model routing, will it necessarily reduce our costs?

Agentic model routing is the practice of using a software layer to automatically select the cheapest capable model for each step of a task, rather than sending every request to a single, expensive default model. The answer is around having choices available.

The advantage we experience in working with Microsoft as a partner, is that Azure AI Foundry’s support for a broad range of foundation models is a significant competitive advantage because it gives organizations the flexibility to choose the right model for each use case rather than being locked into a single provider. This allows for optimization of cost, performance, latency, governance, and regional compliance requirements.

In addition, enterprise routing gateways describe combining model routing with prompt caching and context optimization to bring net costs down 60–80% overall, turning, in one illustrative example, a $15-per-million-token workload into roughly $2 per million by shifting most calls to smaller models and reserving the top-tier model for the minority of steps that truly need it.

Prompt caching, storing the token cost of a stable system prompt or a set of tool definitions so it isn’t reprocessed on every call, can cut cached-input costs by as much as 90% for agents that run the same scaffolding thousands of times a day. Combined with routing, this is what actually gets an organization’s blended cost per million tokens down into the range the tiered-architecture companies are reporting, rather than the frontier-everywhere range most companies default to by accident.

A lot of the surprise bills companies get hit with come from a small slice of employees, often just the top 1–2% of the workforce, using the wrong tool for the job. A router doesn’t just save money on average; it specifically catches the expensive outliers that a flat, one-model-for-everyone policy has no way to notice, let alone stop.

What does a cost-and-governance-first agentic architecture actually look like in practice?

Most companies bolt routing and governance onto an AI system after the fact, once the bills or the compliance questions get uncomfortable. A smaller number, including Support Partners, build both in from the start, and that starting point matters, because it changes what the system does by default rather than what it does when someone remembers to configure it.

Support Partners has domain expertise and understanding of how these costs impact production, and how they are causing confusion as to how the cost of a deliverable really impacts the overall outcome.

Support Partners own content intelligence platform, AIR Fusion , is a useful real-world illustration of how we have approached cost versus value. We see platforms where the cost is high but the quality of the models and the AI does not translate to good value. This translates to user experience.

Our own experience and domain knowledge built AIR Fusion specifically for media and broadcast content operations, choosing to do so on Microsoft Azure. AIR Fusion runs semantic scene detection, object identification, OCR, transcription, and contextual tagging, with 29 core features to help you intelligently use and reuse your content. Where other platforms automatically process AI on each item ingested, AIR Fusion allows you to set the level of AI processing to reduce excessive token usage by using templates, giving the user the flexibility to decide what templates to run on their content. Every action, and transformation is tracked, auditable, and permission-controlled, with granular access, audit logs, and role-aware governance built into the platform rather than layered on top of it.

Our aim is to provide the right context for the right value, choosing the advantages of the Microsoft ecosystem for a robust secure environment, choice of where that token consumption is placed, and with the additional goal of building sustainability into the process. With our Catalyst services, we provide customisation, with a sharp focus on providing just what the client needs for that outcome, pivoting to web-grounding, region of interest selection and specialist agents to make the workflow as efficient and cost effective as possible.

This combination is exactly what enterprises worried about token costs and AI governance are asking for at once: AI processing that runs when you need it, is recorded so nobody re-triggers it by accident, and sits inside guardrails from day one rather than being retrofitted after a compliance review. It’s a small but very valid example of the same principle running through this whole piece, the token bill problem isn’t solved by cheaper tokens, it’s solved by not spending tokens on work that’s already been done or is irrelevant to the outcome you need.

What’s the takeaway on AI token costs and agentic routing?

As Token Mania subsides and the users become more discerning, the value of tokens will drive a higher understanding of value generated. Per-token prices will likely keep falling for years (Gartner has forecast up to a 90% reduction in frontier inference costs by 2030), and competition from Chinese open-weight labs is accelerating that decline. As it is commoditised more and more, the total spend will likely keep climbing right alongside it, because cheaper intelligence just means more of it gets used, it is just figuring out how we use this most efficiently. The organizations that come out ahead won’t be the ones always defaulting to the lowest or the highest price on a model. They’ll be the ones that treat model selection itself as an engineering decision, build governance , outcomes and duplicate-work detection into the architecture from the start, and hand routine model choice to an agentic routing layer rather than a human default preference or habit.

Leave a comment

Thanks for reading AIR Fusion Substack! This post is public so feel free to share it.

Share

How can I find out more on AIR Fusion or Catalyst?

Visit

https://support-partners.com



Are AI token prices actually falling? Yes. Blended AI token prices dropped roughly 67% year over year between Q1 2025 and Q1 2026, and some estimates show a near-90% decline since 2023.

Then why are enterprise AI bills going up? Because total token volume is growing faster than prices are falling, agentic workflows use 5–30 times more tokens per task than older chatbot-style usage, so total spend rises even as the per-token price drops.

What is model routing? Model routing is a software layer that automatically sends each task or step to the cheapest AI model capable of handling it, instead of defaulting every request to one expensive model.

How much can model routing save? Reported production results range from roughly 40% to 85% cost reduction, depending on the routing approach and how much of the workload actually requires a frontier-level model.

Are Chinese models like Kimi and DeepSeek cheaper than US models? Yes, generally. Models like Kimi K2.6/K3 and DeepSeek V4 are priced well below comparable Western frontier models and are putting downward pricing pressure on the whole market, though raw price comparisons should account for how many tokens each model needs to complete a task.

Can AI systems be built to avoid redundant processing? Yes. Platforms like Support Partners AIR Fusion run AI analysis (such as tagging or scene detection) once per asset at ingest, track that it’s been done, and avoid re-running the same AI work later, combining cost control with built-in governance and audit trails.



Sources: Correlation One’s 2026 Enterprise AI Enablement playbook; Forbes reporting on enterprise AI provider margins (July 2026); Ramp’s June 2026 AI Index; EY’s fifth-wave US AI Pulse Survey (July 2026); the FinOps Foundation’s 2026 State of FinOps report; analysis of enterprise API call data reported by Optimum Partners and The Source Code; pricing and market coverage of Kimi, DeepSeek, and other Chinese open-weight models from Artificial Analysis, Axios, Semafor, and Mi3; Support Partners’ public description of the AIR Fusion platform; and routing research from Zylos Research, Requesty, and the RouteLLM/arXiv literature on LLM routing, Futurism, “AI is forcing employees to work harder than ever ”By Frank Landymore Published Mar 12, 2026 8:57 AM EDT.

Support Partners
Aug 5, 2026, 12:00:00 AM

Comments