Deferred criteria let a predetermined choice masquerade as a fair evaluation. Fixing weighted, benchmarkable criteria before anyone sees a price removes the room for that.
You ran the evaluation last month. Three proposals landed, the team booked a room, and someone opened a blank spreadsheet to build the scoring grid. That is the tell. The grid was blank because nobody had agreed the criteria yet, and by the time you filled it in, all three prices were already visible. So the weights got set with the answers in view. Integration got a heavier column the week the favoured proposal happened to integrate well. Support response time quietly dropped in weight once it became clear the preferred vendor was slower. Nobody lied. Everyone was reasonable. And the scoring landed exactly where the strongest voice in the room already wanted it to land.
This is the most common way a supposedly objective evaluation produces a foregone conclusion. Not fraud, not favouritism you could point to, just criteria that arrived after the data they were meant to judge. When that happens, the score is not measuring the vendors. It is measuring how well the committee can reverse-engineer a defensible number for the choice it has already made.
A scoring model is a set of weights. Weights are opinions about what matters. The instant those opinions are formed while the prices and features are visible, they are contaminated, because the human brain is very good at finding reasons for a preference it already holds. This is not a character flaw in your team. It is the default behaviour of any evaluation where the standard is set after the evidence.
Consider a simple example. Say a committee is choosing between two data platforms. Before quotes arrive, most people would agree that reliability and total cost matter more than, for instance, the polish of the admin console. But once the proposals are open and one vendor's console is visibly slicker while its cost is 15 percent higher, the pull is to nudge the console weight up and the cost weight down, just enough. Each nudge feels justified in isolation. Together they flip the ranking. The score now says what the room wanted it to say.
Most organisations already have a procurement policy that says criteria should be agreed up front. The problem persists anyway, for three structural reasons.
First, timing. Intake usually arrives as a request to buy a named thing, not a problem to solve, so the criteria conversation feels premature. We wrote about that pattern in the intake ticket that already names the vendor. When the request is a product name, nobody thinks to argue about weights until the RFP responses force the question. By then it is too late to be neutral.
Second, effort. Building a genuinely weighted, defensible scoring model from scratch is real work, and it is work with no visible output until much later. Under deadline, teams skip it and promise to build the grid when the proposals arrive. That promise is where objectivity dies. It is the same failure mode as setting the budget before anyone checked the market price: the decision that should anchor everything gets made last, under pressure, with the wrong information.
Third, and most quietly, deferred criteria are useful to whoever already knows what they want. A blank grid is a gift to the strongest opinion in the room. If your criteria are fixed and public before the quotes land, that person has to argue about weights honestly, in the open, before they know which vendor a given weight favours. That is exactly the argument they would rather avoid, which is why the grid stays blank.
Vera treats the scoring model as an intake artifact, not an end-of-process one. When a category is opened, she proposes a weighted criteria set drawn from the benchmark library rather than a blank sheet. Those weights are not invented for your event. They reflect what actually differentiates outcomes for that category across a wide base of transactions, so the starting point is a market standard, not one person's preference.
The committee still owns the weights. The difference is when they own them. You confirm and adjust the model before a single proposal is opened, and the model is locked once quotes are in flight. Prices are held back from the scoring view until the weighted evaluation on non-price criteria is complete. That sequence, standard first, evidence second, is the entire mechanism. It costs you almost nothing and it removes the room in which scoring bends.
Because the criteria are benchmarkable, each vendor's answer is scored not just against the others in your shortlist but against the wider distribution. A response that looks strong relative to two weak competitors can still sit at the 40th percentile of what the market delivers. That context is what stops a narrow field from flattering a mediocre winner. If your requirements themselves are the problem, that is a different failure worth reading about in the requirements document that collapses your shortlist to one.
Be honest about the limits. Fixing criteria at intake removes the mechanical way scoring bends. It does not remove human intent. A committee determined to reach a predetermined answer can still weight a criterion aggressively at intake, before the quotes, and defend it. The gain is that they now have to do so in the open, with market benchmarks visible as a counterweight, rather than quietly nudging numbers once the answer is in view. That is a large improvement in accountability, not a guarantee of neutrality.
Nor does a good scoring model tell you whether the requirement was right in the first place. If the underlying need is wrong, a rigorous, well-weighted evaluation will simply select the best answer to the wrong question. Benchmarks discipline the scoring, not the strategy behind it.
And it does not stop a vendor's terms from shifting after you score. Price lists move, and the winning proposal you scored in March may not be the offer on the table in June. That is a separate watch problem, one we cover in catching a price list change the week it happens. What fixing criteria at intake does give you is a clean, defensible, market-anchored score, produced before anyone saw a price, which is the one thing a room with a favourite can never talk its way around.
Fredrik has spent more than twenty years in enterprise software, with time at Oracle, IBM, SAP, and Salesforce before moving to the buy side. He structured and priced the kind of large agreements most buyers only see once or twice in a career, which taught him where the leverage sits and how far a vendor will actually move. He started VendorBenchmark to hand that knowledge to every sourcing team.