Unquantified requirement language is the cheapest thing to fix before signature and the most expensive thing to argue about afterwards. The fix is a reading discipline, not a lawyer.
Go and find the requirements annex of the last statement of work you signed. Somewhere in it there is a line that reads something like the system must respond quickly under normal operating load. It was written by an architect in week two of the process, it survived a technical review, a legal review, a security questionnaire and a procurement sign off, and nobody argued about it once. That is because nobody disagreed. Your platform team read it as a 200 millisecond ninety fifth percentile response. The implementation partner read it as sub second, which in their delivery model means anything under 800 milliseconds at the median. Both readings are defensible. Only one of them is in the price.
The gap surfaces in integration testing, roughly four months in, when someone finally puts a load generator against the environment. Your team calls it a defect. The vendor calls it an enhancement, because the contract says quickly and the build is quick. The conversation that follows is not a technical conversation. It is a commercial one, held at the worst possible moment: after the money is committed, after the deadline is public, and after your leverage has evaporated. There is a change order, a caching tier nobody budgeted, and six weeks of schedule that never comes back.
This is not a vendor problem. Every serious supplier writes requirement language that is defensible on their side of the reading, because that is what a commercial team is paid to do. It is a specification problem, and it is one of the few sourcing failures that is entirely inside your control at the moment it is created.
Ambiguity in a requirement is not an accident of drafting. It is a lubricant. It lets a deal keep moving when two parties are not yet ready to agree on a number, and both sides feel the relief of that. The architect does not want to commit to 200 milliseconds before seeing the production data model. The vendor's solution lead does not want to commit to it at all until their delivery team has sized the work. So the word quickly goes in, everyone nods, and the document advances to the next open item. The cost of that nod is deferred, not avoided.
The second reason it survives is that the people who would catch it are not reading for it. Legal reads for liability, indemnity, termination and IP. Security reads the questionnaire. Procurement reads price, term and renewal mechanics. The technical owner reads for scope, not for measurability, because they already know what they meant. Nobody in the chain owns the question: is this sentence testable? An untestable requirement passes four reviews because it is not in scope for any of them.
The third reason is structural. Requirement language migrates. A phrase written into an RFP question is copied into a vendor response, lifted into a solution summary, pasted into the SOW annex, and then referenced by the acceptance criteria. By the time it reaches the annex it carries the authority of four documents and the precision of none. Nobody rewrites it, because rewriting inherited text feels like reopening a settled point.
The honest way to think about a vague adjective is as an unpriced option that you have written and the vendor holds. They get to choose the interpretation later, when the choice is worth money. You get to argue, from a position where the alternative to agreeing is a stalled programme.
The arithmetic is easy to sketch and uncomfortable to look at. Take an illustrative mid sized implementation with a 1.8 million contract value. One performance requirement resolved in the vendor's favour, requiring additional infrastructure and rework, lands as a change order in the region of four to seven percent of contract value. Add six weeks of internal delay at, for example, eight loaded people, and the internal cost is comparable to the change order itself. That is one sentence. Most requirement annexes we decode contain somewhere between eight and thirty lines built on unquantified adjectives: fast, scalable, reliable, timely, industry standard, reasonable, appropriate, best efforts, as needed, where practicable.
Worse, these lines are invisible to the usual controls. A spend report will not show them. A savings plan will not price them. Even a thorough commercial review focused on rate cards and renewal mechanics steps straight past them, because they are not commercial terms in form, only in effect. If you already run the estate the way we describe in from spend list to savings plan, this is the category of exposure that sits underneath the ranked moves and never appears on the list.
The test for a requirement is not whether it sounds demanding. It is whether a neutral third party, handed the contract and a test environment, could determine pass or fail without asking either party what they meant. That test has three parts: a metric, a threshold, and a measurement condition. Quickly fails all three. Ninety fifth percentile API response under 300 milliseconds, measured at the load balancer, at 1,200 concurrent sessions, with the reference data volume defined in annex C, passes all three. Note that the second version is not more aggressive than the first. It is simply resolvable.
This is the same reading habit we describe in decode any contract in a minute, applied one layer deeper than the commercial clauses. And it belongs before signature for a simple reason: a threshold you propose during negotiation costs a conversation, while the same threshold proposed after signature costs a change order.
Contract and quote decoding runs the draft the way an experienced reviewer would, except it does not get tired at page forty. You load the draft SOW, the requirements annex, the vendor's response document and the quote. Six specialist agents read them together and return one list: every requirement sentence that lacks a metric, a threshold or a measurement condition, ranked by the exposure it creates if the vendor's reading prevails.
Two things make the flagged list usable rather than merely correct. The first is the substitution. A flag that only says this is vague creates work. A flag that says this is vague, here is the metric normally used for this requirement type, and here is the threshold range observed in comparable deals, closes the item. That comparison draws on 520 vendor benchmarks and 500,000+ real closed transactions, so the number you propose is a market number rather than an internal guess, which matters a great deal when the vendor pushes back. If you are unsure how to read the position you are handed, how to read a percentile covers what a market range does and does not tell you.
The second is scope. One draft is a task. The estate is the problem. Review tables let you ask the same question across every contract you hold at once: show me every performance, availability, support response and data retention requirement that has no measurable threshold attached. That turns a drafting habit into an inventory, and an inventory is something you can work down.
From there the motion is ordinary procurement work. Unquantified lines in live contracts go onto the renewal agenda, because a renewal is the one moment when reopening inherited text is cheap. Unquantified lines in drafts go into the redline. Patterns that repeat across three or more vendors become a standing position in the clause library, so the next architect inherits a template that already contains the metric, the threshold and the measurement condition. That is the same logic behind the must have coverage grid, extended from protections into precision.
Decoding tells you that a line is unmeasurable and what metric would make it measurable. It does not tell you what your threshold should be. Only your workload model, your user expectations and your tolerance for cost can decide whether 200 milliseconds is a requirement or a luxury, and a market range is an input to that judgement rather than a substitute for it. If you set the threshold too tight you will pay for headroom you never use, and no amount of flagging protects you from that.
It does not make the vendor agree. A supplier can decline a threshold, price it as an option, or accept it with carve outs that hollow it out. Precision improves the quality of the argument and moves it to a moment when you still have leverage. It does not remove the argument. Some of the time the honest outcome is a documented open item with a resolution date and a named owner, which is still materially better than an adjective.
It does not retroactively fix a signed contract. For live agreements, the output is an inventory and a renewal agenda, not a remedy. And it does not replace acceptance testing. A measurable requirement is only worth what your test regime proves, so someone still has to run the load, read the result and hold the line when the number comes in high. The platform gets the number into the document. Your team still has to check it.
None of this is glamorous work. It is one reading pass, applied at the one moment when changing a sentence is free. The reason it is worth the discipline is that the alternative is not a cleaner document. The alternative is a conversation in month four about what quickly meant, held with no leverage and a public deadline, and everyone in that room already knows how it ends.
Morten brings two decades of enterprise and software procurement, with stints across Oracle, IBM, SAP, and Salesforce shaping how he reads a deal. He has led sourcing through hundreds of renewals, from mid market order forms to nine figure global agreements, and learned that the buyers who win are the ones who walk in knowing the market. He built VendorBenchmark to make that pattern recognition repeatable.