Direct spend with the AI model providers has crossed from experiment to enterprise line item, and it is negotiated like nothing you have bought before: priced in tokens, sold on committed spend, and moving too fast for last year's intuition. Right-size the mix, pull the levers, and benchmark the commitment.
A category that barely existed two years ago is now a standing line on the software budget: direct spend with the AI model providers. Anthropic, OpenAI, and their peers have moved from a credit card experiment in one team to an enterprise agreement negotiated by procurement, and the negotiation looks like nothing else in the portfolio. The unit is the token, a thing most buyers cannot intuitively price. The structure is committed spend, drawn down over a term. And the whole market moves fast enough that a rate that was competitive two quarters ago may be generous or stingy today, with no stable folk knowledge to anchor on.
That combination, a novel unit, a consumption commitment, and rapid change, is exactly the environment in which buyers overpay, because the usual instincts do not apply and the vendor sizes the deal from a growth story. But an AI model agreement is still a software agreement, and the same discipline that governs every other line applies here too, adapted to the units: size the commitment from your real usage, pull the levers that lower the effective price before you commit, and benchmark the rate against what comparable buyers actually pay. The units are new. The way you avoid overpaying is not.
The single largest determinant of an AI bill is usually not the rate you negotiate, it is which models you run for which work. The providers offer a range of models at very different prices, and the most capable, most expensive one is rarely required for the bulk of what an enterprise actually does. A large share of real workloads, classification, extraction, routine generation, runs perfectly well on a smaller, cheaper tier, with the top model reserved for the genuinely hard tasks. Getting that mix right, routing each workload to the cheapest model that does the job, moves the bill more than any discount you are likely to win.
This is why the optimizer starts with the mix, not the rate. Feed in your actual usage by model and the question becomes concrete: how much of your top-tier spend is doing work a mid-tier model would handle, and what would routing it down save? The vendor has no incentive to raise this, because a customer running everything on the flagship model is a more valuable customer. Right-sizing the mix is a lever entirely on your side of the table, and pulling it before you size a commitment means you commit against an optimized run rate rather than an inflated one.
Beyond the mix sit structural levers that lower the effective price of the same output without any concession from the vendor, and pulling them before you commit is often worth more than the discount. The first is caching: much of what an enterprise sends to a model is repeated context, the same system prompt, the same reference documents, and caching that repeated portion cuts its cost sharply on the providers that support it. The second is batching: work that does not need an instant answer can run through discounted asynchronous processing rather than at full real-time rates. Both are within your control, and both shrink the run rate you are about to commit against.
The discipline is to model these before sizing the commitment, because a commit built on your unoptimized usage is a commit sized to a number you could have made smaller for free. A buyer who right-sizes the mix, caches the repeated context, and batches the non-urgent work arrives at the negotiation with a materially lower and more accurate run rate, which means a smaller, safer commitment and a discount that applies to spend you will genuinely incur. The vendor prefers you commit big against your raw usage; the levers are how you commit right against your optimized one.
With the run rate optimized, the last questions are whether the rate is competitive and how large the commitment should be, and both need a benchmark because intuition fails on new units. List pricing for the frontier models is public, so unlike some vendors the starting point is visible, but the enterprise agreement, the committed-spend discount, the rate at your volume, is negotiated, and it sits in a distribution of what comparable buyers achieved. Benchmarking it tells you whether the committed-spend discount you are offered is strong or mediocre, which is impossible to judge from the quote alone when the whole category is a year old.
The commitment itself should be sized in a band, high enough to earn the discount on your optimized run rate, low enough that you will actually consume it within the term. A consumption commitment on a fast-moving AI spend is a genuine bet, and the providers size their pitch to the optimistic case, so the safe commit is one your realistic usage burns down with margin, letting growth handle the upside rather than a forfeited pre-purchase. Optimized mix, free levers pulled, rate benchmarked, and a commitment sized to a band you will burn: that is an AI agreement priced on evidence rather than accepted on a growth story.
AI pricing changes faster than any category on the estate, and a benchmark on a year-old market is thinner than on a decade-old one, so the placement is a strong guide rather than a settled verdict. Model capabilities and prices shift quarter to quarter, which means a commitment carries genuine uncertainty and the sensible response is margin, not a bet sized to today's optimism. Some of the risk here is irreducible, and a right-sized commit accepts that rather than pretending the future is knowable.
What the discipline removes is the specific way this new category separates buyers from money, which is a novel unit, a growth story, and a consumption commitment combining to produce an over-sized deal at an unbenchmarked rate. Optimize the mix, pull the free levers, benchmark the rate, and size the commit to what you will burn, and an Anthropic or OpenAI agreement becomes a line you priced deliberately. The tokens are unfamiliar; the trap and the answer are the ones procurement has always known.
Morten brings two decades of enterprise and software procurement, with stints across Oracle, IBM, SAP, and Salesforce shaping how he reads a deal. He has led sourcing through hundreds of renewals, from mid market order forms to nine figure global agreements, and learned that the buyers who win are the ones who walk in knowing the market. He built VendorBenchmark to make that pattern recognition repeatable.
What shipped on the platform, and the pricing and licensing moves worth knowing before your next renewal. One email a week, to your work address. Unsubscribe any time.