Two new executive disciplines: price your tokens right, or someone else prices your alpha for you.
By The Chiri Team
What is your token spend actually buying you?
Every other major line on the P&L, cloud compute, headcount, marketing, has decades of tooling, benchmarks, and institutional practice built around tracking what it returns, imperfect as that tracking often is. Token spend has none of that yet.
It is newer, growing faster, and split across more tools than any other line on the books. That is exactly why it is becoming one of the more expensive blind spots in the C-suite.
Two disciplines are emerging out of that gap, and neither one had a name or an owner on the org chart until very recently. The first is token optimization: getting the most value out of every dollar of inference spend by matching the right task to the right model. The second is alpha protection: making sure the work that makes your company distinct does not quietly become someone else’s training data or someone else’s product.
They are related. Understanding why is worth ten minutes of your time.
Token optimization is not a cost-cutting exercise
A useful way to think about your token bill is as a stack of casino chips. Every inference call is a spin, priced by the house, and the house sets a different price depending on which table you sit at. Not all tokens are created equal, and not all tokens are priced equally either.
Coding tokens have been heavily subsidized by the frontier labs chasing developer adoption. Back office and reasoning-heavy tokens have not. A company that treats every workload the same way, routing HR questions and legal review through the same premium model it uses for production code, is paying frontier prices for work that does not need frontier intelligence.
The market is already voting with its usage. Vercel’s AI Gateway data for June 2026 shows open-weight models running 29 percent of all gateway tokens, up from roughly 11 percent in April, on under 4 percent of total spend. DeepSeek alone accounted for 22.6 percent of token volume on that gateway, trailing only Anthropic and Google. Read that split again: nearly a third of the volume, for a twenty-fifth of the dollars. (Vercel, July 2026)
CNBC’s July 7 investigation into OpenRouter traffic found the same pattern from a different angle. Chinese-origin models have accounted for at least 30 percent of enterprise token volume on the platform every week since February 8, 2026, spiking to 46 percent in a single week. Compare that to an average of just 11 percent over the prior twelve months, and 4.5 percent in the first half of 2025. (CNBC)
That is not a fringe trend. That is a market repricing intelligence in real time, and treating it as a rounding error is how a company ends up with a token bill that scales faster than its revenue.
The discipline here is not “spend less.” It is knowing which tasks require the most capable, most expensive model available, and which tasks are being over-served by it.
- Prove the use case on the frontier model.
- Route the repeatable version of that work to a model priced for the job, once proven.
That is token optimization. Everything else is just an unmanaged AWS bill with better marketing.
The domestic labs know this is happening to them. On July 30, 2026, Sam Altman announced an 80 percent price cut on GPT-5.6 Luna, down to $0.20 per million input tokens and $1.20 per million output, plus a 20 percent cut on GPT-5.6 Terra, alongside a new Fast mode for the flagship GPT-5.6 Sol. OpenAI attributed the cuts to efficiency gains, not charity. (CNBC, VentureBeat)
That is what it looks like when a frontier lab feels the same pricing pressure this article is describing. If OpenAI is cutting prices on its own product line to compete, the market for your token spend is already moving, whether your procurement process has caught up to it or not.
Alpha protection is the newer of the two disciplines, and the harder one to see coming
Palantir CEO Alex Karp has been unusually blunt about this in public this year. He has accused frontier AI labs of “stealing weights and alpha” from enterprise customers, and said businesses are “paying for tokens that create no value” while their most sensitive workflows and data pass through someone else’s model. He is calling for what he terms AI sovereignty: companies owning their compute, their data, and their models rather than renting all three. (Alex Karp, interview on CNBC’s Squawk Box, July 1, 2026: “Palantir’s Karp bashes token-based AI model as ‘completely wrong’”)
Karp has an obvious commercial interest in that framing. He is also not wrong about the underlying mechanics.
Watch what happened when Anthropic launched a legal plugin for Claude on February 3, 2026, automating contract review, NDA triage, compliance tracking, and legal briefings:
- Thomson Reuters fell 16 percent that day.
- RELX fell 14 percent, its steepest single-day drop since 1988.
- Wolters Kluwer fell 13 percent.
A combined market reaction in the hundreds of billions of dollars. (Morningstar, “Thomson Reuters, RELX, and Wolters Stocks Crushed After Anthropic Debuts Claude Legal Plug-In,” February 2026) That is what the market thinks happens when a foundation model provider learns enough about a vertical, from aggregate usage and prompt patterns across its customer base, to compete directly in that vertical. Whether the sell-off proved out or overshot, and analysts are genuinely split on that, the reaction itself tells you the market takes the underlying risk seriously.
Even with a promise of zero data retention, there is a real, actively studied line of security research, membership inference and embedding inversion attacks, showing that a model provider can sometimes infer meaningful signal about what was sent to it even without retaining the raw prompt. This is not a settled, universal vulnerability in every system. It is a real enough risk in the underlying architecture that “we promise not to look” is a policy, not a technical guarantee, and the two should not be treated as equivalent.
The fix is not to stop using frontier models. It is to be deliberate about where your alpha actually lives, and to route the workflows that touch it through infrastructure with a real legal backstop. Zero data retention agreements that flow through a licensed subprocessor, not just a checkbox in a settings page, give you something to enforce if the promise breaks. That is the difference between a policy and a contract.
Why these two disciplines are actually one discipline
Both come back to the same root question: what happens to your business if the assumptions behind your AI spend change under you?
If your token costs keep exponentiating while your competitors find a cheaper, comparably capable model, you have a business continuity problem. If your most valuable workflows are training someone else’s next product release, you have a business continuity problem. Both are the same executive job, just pointed in different directions. One protects the bottom line. The other protects the reason customers choose you over the next company down the street.
Neither of these disciplines requires a chief AI officer or a six-month transformation program. It requires getting your CISO, your CFO, and whoever owns your AI spend in the same room, asking what your actual threat model is, and being honest about which workflows can move to a lower-cost model and which ones need a contract, not just a promise, behind them.
Each seat at the table owns a different piece of this:
- The CFO owns the number, and needs to know it is not a fixed cost, it is a usage curve that can bend the wrong way fast.
- The CTO owns the routing decision, and needs the technical case for why one workflow can move to a cheaper model and another cannot.
- The COO owns what breaks operationally if a vendor’s pricing or availability shifts under a live process.
- The CHRO owns the harder conversation about what “your alpha” actually means: the institutional knowledge sitting in your best people’s heads that a model trained on your prompts could quietly learn to replicate.
- The CEO owns the fact that this is now a competitive question, not just an IT budget line.
What does your organization actually know about where its token spend goes, and what it is protecting on the way there?
Sources cited:
- Vercel, “Open-weight models surge to 29% of volume, price per token flattens,” AI Gateway Production Index, July 2026. https://vercel.com/blog/ai-gateway-production-index-july-2026
- CNBC, “Chinese AI models are gaining ground with U.S. companies as OpenAI, Anthropic costs surge,” July 7, 2026. https://www.cnbc.com/2026/07/07/chinese-ai-models-costs-us-openai-anthropic.html
- CNBC, “Palantir’s Karp bashes token-based AI model as ‘completely wrong,’” Alex Karp on Squawk Box, July 1, 2026. https://www.cnbc.com/2026/07/01/palantir-karp-open-ai-anthropic-tokens.html
- Morningstar, “Thomson Reuters, RELX, and Wolters Stocks Crushed After Anthropic Debuts Claude Legal Plug-In,” February 2026. https://www.morningstar.com/stocks/reuters-relx-wolters-stocks-crushed-after-anthropic-debuts-claude-legal-plug-in
- CNBC, “OpenAI cuts prices for two of its GPT-5.6 AI models as companies grow sensitive to costs,” July 30, 2026. https://www.cnbc.com/2026/07/30/open-ai-price-cut-gpt.html
- VentureBeat, “AI price wars: OpenAI cuts GPT-5.6 Luna prices by 80% as model competition shifts toward cost,” July 2026. https://venturebeat.com/technology/ai-price-wars-openai-cuts-gpt-5-6-luna-prices-by-80-as-model-competition-shifts-toward-cost

