Per-Seat Pricing Is Back


By The Chiri Team

Flat, per-seat pricing built the SaaS industry. A product cost roughly the same to serve one customer as the next, so a single monthly price worked at any scale. AI broke that assumption. A capable AI product does not cost the same for every customer. It costs whatever the tokens a customer consumes actually cost, and until recently, that number could run far past what a flat seat price ever covered.

That is why so many AI products moved to metered credits and usage caps instead of a single flat price. The shift got framed as a product decision. It was a cost decision. A vendor pricing a seat has to survive its heaviest users, not just its average ones, and at frontier model prices, a heavy user could cost many times the seat price in a single month.

A new pricing promotion from a commodity model provider shows why that math is changing again.

Why flat pricing broke

The logic is simple once it is written down. A seat’s marginal cost equals the tokens a customer consumes times the rate a model charges for those tokens. Legacy SaaS software had a marginal cost close to zero. AI products do not.

Picture a heavy agentic user: someone running an AI assistant through repeated tool calls and multi-step tasks, consuming on the order of 50 million input tokens and 5 million output tokens in a month. That is a real, if demanding, usage pattern for an active daily user of an agentic product, not an edge case.

Run that workload against Anthropic’s published Claude Sonnet 5 pricing, $3 per million input tokens and $15 per million output tokens as of September 2026: the monthly token bill lands near $225. Run the same workload against a premium tier like Claude Opus or Fable 5, at $10 and $50 per million tokens, and the bill climbs past $700.

Set either number against a typical consumer software price point of $20 to $30 a month, and the problem is obvious. A single heavy user can cost a vendor many multiples of what that user pays. A flat price that has to survive its heaviest user, at frontier rates, is not a sustainable price. Credits, usage caps, and metered billing were the honest response to that math, not a product philosophy.

What changed

Z.ai launched a promotional pricing window for its GLM-5.3 Flash model this year: $0.075 per million input tokens, $0.015 per million cached input tokens, and $0.25 per million output tokens, with list pricing after the promotion set at roughly double those rates. OpenRouter lists the same model even lower, at $0.05 input and $0.167 output per million tokens.

Run the identical heavy workload, 50 million input and 5 million output tokens, against those rates. The monthly bill lands around $3.34 through OpenRouter, or about $5.00 at Z.ai’s promotional direct rate. Even at full list pricing after the promotion ends, the same workload costs roughly $10.

That is not a small improvement over frontier pricing. It is a difference of 20 to 200 times, depending on which frontier tier and which commodity rate get compared. A workload that could cost $225 to $700 a month against a frontier model costs single dollars against a commodity one.

A formula, not a guess

The shift can be written as a simple condition. A flat-priced seat is viable when a customer’s expected monthly token cost stays at or under a modest share of what that seat costs, commonly modeled around a third of the price, leaving enough margin to cover support, infrastructure, and the ordinary cost of running a business.

That condition fails hard at frontier model prices for any genuinely heavy agentic user. It holds comfortably at commodity model prices for the same workload. The model matters more than any other input to whether flat pricing works at all.

This reframes a familiar industry question. When a customer or an analyst asks whether a flat-rate AI seat will hold up if a given model gets more expensive, the honest answer is a calculation, not an argument: run the expected usage against the model’s real rate, compare it to the seat price, and the condition either holds or it does not.

Why this is not a permanent verdict

Commodity model pricing will not stay this low forever, and a promotional rate is, by definition, temporary. The point is not that any one model’s price will stay this cheap. The point is that capable, cheap models now exist at all, and that changes what a sustainable flat price can assume.

A year ago, a company building an AI product had one real option for a genuinely capable model: pay frontier prices and build a business model around the cost of doing so, usually through metering or usage caps. Today, a company has a second option: route the same class of work to a commodity-priced model built for exactly this kind of volume, and price a flat seat around that cost structure instead.

That second option is what makes flat, predictable pricing viable again for AI-native products, not as a bet that commodity prices never rise, but as evidence that the model market now supports more than one viable price tier for genuinely useful work.

What this means for how a product gets priced

Two structural choices follow from this shift, and neither depends on any one promotional window staying open.

The first is which model tier a product routes ordinary usage through. A product built entirely around frontier-model reasoning inherits frontier-model economics, and its flat pricing has to be built to survive that. A product that routes the bulk of everyday, high-volume work to commodity-priced models built for that job, reserving frontier models for the tasks that genuinely need them, inherits a materially different cost base.

The second is who bears the token cost at all. A growing pattern lets a customer connect their own model API key, sometimes called bring-your-own-key, so inference cost passes through to the customer directly rather than through the vendor’s own margin. Under that structure, a flat seat price becomes closer to pure platform rent: it covers the software, the workflow, the integrations, and the governance built around a model, while the customer’s own key absorbs whatever the underlying tokens cost. That structure insulates a flat price from model pricing swings entirely, in either direction.

Both choices point at the same underlying shift. The token, not the seat, was always the real unit of cost in an AI product. Pricing that ignores the token was never going to hold. Pricing that accounts for it, whether through routing, through commodity models, or through passing the cost through directly, is what makes a flat, predictable price sustainable again.

This lands differently depending on where you sit

A CFO or finance leader at a SaaS company selling AI features: the viability condition gives a concrete way to stress-test a flat price before committing to it. Model the heaviest realistic user’s token consumption against the actual model rate in use, not the average user’s, before setting a seat price.

A product leader deciding which model to build on: routing ordinary, high-volume work to a commodity-priced model, and reserving frontier models for tasks that specifically need that capability, is now a pricing decision as much as a technical one.

A buyer evaluating a flat-rate AI product: ask what model tier the vendor’s everyday usage actually runs on, and whether the vendor’s pricing was built to survive a heavy month at that rate, not just an average one.

An operator considering a bring-your-own-key structure: passing token cost through to a connected model key changes what a flat price is actually paying for, and it is worth understanding that distinction before assuming a flat price and a metered price mean the same thing underneath.

If a product’s flat price had to survive its heaviest user’s actual token bill this month, would it still hold?


Want more Field Notes?

Weekly dispatches on AI orchestration, ontology, and the agentic enterprise.