The agents now take instructions before they start, show their working while they run, and hand you a room the rest of the team can use. Here is what shipped, and where it still stops short.
Here is the failure mode every buyer who has used an AI agent will recognise. You point it at a vendor, it thinks for four minutes, and it returns a beautiful, well structured, entirely competent document about the wrong thing. You wanted the renewal leverage. It wrote you a company profile. You wanted the three clauses that will decide the negotiation. It gave you an even handed summary of all forty. Nothing in the output is wrong. It is simply not the document you needed, and you have no way to say so except to run it again and hope.
That is a briefing problem, not an intelligence problem. A junior analyst who returned the wrong deliverable would be told, once, what the deal actually turns on, and the second draft would land. Until today our agents had no way to receive that instruction. Four upgrades shipped in one day to close the gap: a brief you write in your own words, an agent you assemble yourself from the situation in front of you, a run you can watch page by page, and a finished room that behaves like a workroom instead of a read only artefact.
Every agent run now opens with a free text brief. Not a dropdown, not a tag picker. A box where you write what you write in a Slack message to a colleague: they are pushing a three year term and we think the price uplift is buried in the support line, focus there. That sentence is not decoration. It becomes a weighting instruction applied to every document the run touches, so the order form clauses that speak to term and support pricing are pulled forward and the sections that do not are compressed rather than expanded.
The second half of that change matters more than the first. When your stored data cannot answer something the brief asked for, the agent names it as missing. It does not reach for a plausible industry figure, and it does not quietly rephrase the question into one it can answer. If you asked about support uplift and there is no support schedule in the record, the output says the support schedule is absent and states what it would need. That is the difference between a document you can take into a room with your CFO and one you have to check line by line first. We have written before about how much of contract review is really about knowing what is not in the file, in Decode any contract in a minute, and the same principle now governs the agents.
The six specialist agents we shipped originally were built around recurring jobs: the renewal, the sourcing event, the contract decode. They still exist and they still cover most of what a desk does week to week. But procurement work is lumpy, and the situations that consume the most hours are the ones that do not fit a template. A surprise call in ninety minutes. A leadership brief requested on Thursday for Monday. A renewal where the vendor has already moved and you are answering rather than opening.
You can now assemble an agent from that situation. You describe the circumstance, you choose which specialist desks contribute, and you pick exactly which documents come out the other end. Not a bundle, a selection. If you need the price benchmark and the clause position but not the market landscape, you take two documents and the run is shorter and sharper for it. The full set of 30 background jobs is now pickable, and critically they are pickable before you hold an agreement with that vendor. That was a real constraint until today. Pre contract work is exactly when you have the least data and the most need for structure, and locking the specialist desks behind an existing contract record had it backwards.
This sits alongside rather than replaces the standing roster described in Six agents and a ghost writer. The standing agents handle the calendar driven work. The custom builds handle the ambush. If you want to see the desk map and which jobs feed which document, it lives at /workflows/agents.
The old run experience was a progress bar and a wait. Progress bars are where trust goes to die, because they are almost always a guess about elapsed time rather than a statement about delivered work. If the bar sits at 70 percent for two minutes you learn nothing about whether page four exists.
The run is now a stage. The document writes itself under the pen while you watch, section by section, so you can see the shape of the argument forming and stop early if the brief clearly landed wrong. Beside it, a checklist tracks the real state of each page: queued, writing, delivered, or blocked on missing data. The progress rail is tied to that checklist, which means it will never claim a page the agent has not actually produced. If eleven of fourteen pages are done, the rail says eleven of fourteen, and the three outstanding are named.
The practical value is timing. When you are prepping for a call in forty minutes you need to know whether to wait for the full case file or take the four pages that are already down. The stage lets you make that call on evidence. It also makes a failed run legible. A run that stalls tells you which desk it stalled on and what it was missing, which is usually a contract document nobody uploaded rather than anything mysterious.
Until today, a completed run gave you a room to read and a file to download. Reading is not where procurement work happens. Procurement work happens in the argument about whether the benchmark applies, in the legal reviewer disagreeing with the clause position, in the finance partner asking where the uplift number came from at 6pm on a Thursday.
So a finished room now carries three things. First, a team discussion sitting directly under the reader, with @mentions, so the debate about page seven lives next to page seven instead of in a thread nobody can find in March. Second, a hand off that puts the room on a teammate's desk as a real review task with an owner and a state, rather than a link in a message that may or may not be opened. Third, the whole case file emails to you as a single PDF, each page carrying its own provenance line naming the source document and the date it was read. That provenance line is what makes the PDF usable outside the platform. When someone in a board pack asks where a figure came from, the answer is on the page. This is the same discipline we apply to approval packs in the deal sign off chain, now extended to every agent output.
The brief weights, it does not conjure. If you ask the agent to focus on support uplift and your record contains no support schedule, a sharper brief produces a sharper statement of absence, not an answer. That is the correct behaviour, and it is also the most common source of disappointment in the first week. The quality of the output is still bounded by what you have uploaded. Teams that get the most from this are the ones who load the order forms and the amendments, not just the master agreement.
Second, a custom agent is a selection of existing desks, not a new capability. If the job you need does not map to one of the 30 background jobs, building an agent will not create it. You will get a competent assembly of adjacent work, which is often enough, and sometimes is not. Tell us which desk is missing rather than working around it.
Third, the live stage makes the run legible, it does not make it faster. A full case file across several desks still takes minutes, not seconds, and a run that depends on a slow external source will show as blocked for as long as that source is slow. Watching a stalled page does not unstall it. What you gain is the ability to decide early rather than find out late.
Fourth, the hand off puts a room on a colleague's desk as a task. It does not chase them. The state is visible to you, which is more than a forwarded email gives you, but ownership is still a human problem. And the emailed PDF is a snapshot at the moment you sent it. If the room is updated afterwards, the PDF in someone's inbox does not update with it. Send late, or send again.
None of this changes the underlying arithmetic of a negotiation. It changes how quickly you arrive at a defensible position, and how many people can work on it at once. For a one person desk that difference is the whole job, as we argued in the one person procurement desk. For a larger team it removes the specific friction of the analyst who has the case file and the reviewer who does not. Run one against your next renewal, write the brief as though you were messaging a colleague, and see whether the document that comes back is the one you asked for.
Morten brings two decades of enterprise and software procurement, with stints across Oracle, IBM, SAP, and Salesforce shaping how he reads a deal. He has led sourcing through hundreds of renewals, from mid market order forms to nine figure global agreements, and learned that the buyers who win are the ones who walk in knowing the market. He built VendorBenchmark to make that pattern recognition repeatable.