Agents used to forget the runs you watched them do. Now each one carries a record of what it wrote and whether it finished, and you can hand a finished piece of work to another agent with one button.
Here is a failure mode you have probably lived with. You ask an agent to work a renewal. It runs its chain, pulls the benchmark, drafts a position, writes three documents. You watch all of it happen in the transcript. Then you ask, plainly, what did you find, and it answers as though the last ten minutes never occurred. The runs were real to you. They were never real to the agent. This update closes that gap, and it changes how much you can trust an agent across a long piece of work. If you are new to how we staff a desk with agents, this is the piece that makes them dependable rather than merely fast.
An agent that forgets its own output is a liability disguised as an assistant. You cannot delegate to it, because delegation assumes the worker knows what it did. You end up re-reading the transcript yourself, copying findings out by hand, and stitching the three documents together into something coherent. The agent produced the work and then abandoned it. Worse, an agent with no record of its runs will sometimes answer a question about work it did not finish. It will state a finding as fact even though the run that would have produced that finding stopped early. That is the dangerous version of the amnesia, because a confident wrong answer costs you more than a blank one.
So we gave each agent a record of what it has done in a given conversation: what it ran, what it wrote, and whether the run actually finished. Asking about its work now returns the work. And crucially, it will not claim a finding from a run that stopped early. If the chain broke before it produced a number, it tells you it does not have the number, which is the honest answer and the one you want. This sits alongside the broader principle in what we will not let the AI do on your deals.
Memory has a ceiling. Every agent can only hold so much of a conversation at once, and a serious negotiation thread outgrows that ceiling fast. The old behaviour was to quietly lose the beginning: by month three, the thread no longer knew what was agreed in week one, which is precisely the part you most need it to remember, because week one is where the anchor and the walk-away were set.
Now, once a conversation outgrows what an agent can hold, the earlier part is folded into a running summary overnight by one of our background jobs. The thread keeps a compressed record of what was decided, so a conversation months old still knows the position it opened with. This is the same discipline we apply to your screens in every screen you work in now keeps your work. Nothing important is meant to fall off the back of the thread.
The second half of this update is handoff. Under any answer or any finished run there is now a Hand to button. Press it, pick who takes the work, and it lands in that agent's conversation as a starting point rather than as a question. The receiving agent does not have to re-derive anything. It opens with the finished work in front of it and builds from there. Both threads record the handoff, so you can always see where a piece of work came from and where it went. This is the mechanism behind the way we chain sign-off across approvers, each with the brief already in hand.
One rule matters more than any convenience here: agents never message each other on their own. You press the button. There is no autonomous chatter, no agent quietly delegating to another agent behind your back, no chain of instructions you did not author. Every handoff is a deliberate act by a person, and it is logged as one. If you have read our position on autonomy, this is that position made concrete.
The practical effect is that you can now delegate a piece of work to an agent and come back to it. The agent remembers the runs, remembers the documents, remembers whether it finished, and can pass all of that to a specialist without you acting as the courier. That is the difference between an agent that answers questions and an agent that holds a piece of work. It also raises the standard on our output, because a finished run that another agent can build on has to be a real deliverable, which is the argument in why AI reports beat AI answers.
Be clear about what this does not do. Memory is scoped to the conversation. An agent remembers the work it did in a given thread, not everything it has ever done across every thread, and that is deliberate: it keeps context clean and keeps one deal's reasoning out of another. If you want work to travel, you hand it over on purpose.
The overnight summary is a summary. Folding an old thread into a running record compresses it, and compression loses detail. The decisions and the position survive; the exact phrasing of a message from week one may not. For anything you need verbatim, keep the source document, do not rely on the summary to reproduce it word for word. The summary tells you what was decided, not necessarily every sentence that led there.
And the completion signal is only as honest as the run. The agent knows whether its own chain finished, which is why it will not claim a finding from a run that stopped early. It cannot tell you whether the finished finding is correct, only that it was produced. The check on quality is still the benchmark underneath and still your judgement. Nothing here replaces reading the deliverable before you send it, and nothing here replaces confirming that the numbers trace back to our data, which is why we keep pushing on how that data stays fresh. You can see how all of this fits together on our AI agents page.
The short version: your agents no longer forget the work you watched them do, they will not lie about a run that broke, long threads keep their memory of what was decided, and work moves between agents when you decide it should, never on its own. That is a smaller, calmer promise than autonomous AI, and a more useful one for the person who has to defend the deal.
Want to be updated when major licensing and pricing changes land? One analyst brief a week: the price rises, metric changes and audit campaigns that move software costs. Work email only.
Morten brings two decades of enterprise and software procurement, with stints across Oracle, IBM, SAP, and Salesforce shaping how he reads a deal. He has led sourcing through hundreds of renewals, from mid market order forms to nine figure global agreements, and learned that the buyers who win are the ones who walk in knowing the market. He built VendorBenchmark to make that pattern recognition repeatable.