Two friction points removed at once. The image only PDF that used to be a dead end now reads itself, and the cloud archive you kept meaning to import comes across in one pass.
Here is a problem you recognise. A vendor sends you the countersigned MSA as a scan, a photograph of paper with no text layer underneath. You drop it into your contract tooling and it comes back with a polite label that says it needs OCR. So it sits. The renewal clause inside it, the auto renew window, the cap on annual uplift, all of it is invisible to search and invisible to any review table you try to run. Multiply that by the folder of legacy agreements sitting in a shared drive nobody has migrated, and you have a real gap between what you have signed and what you can actually reason about.
This update closes that gap in two places. The scan that used to stall now queues its own transcription and finishes the job on its own. And the cloud folder you kept meaning to import walks across in full, subfolders included, instead of trickling in ten files a day. Nothing new for you to learn, the same intake you already use just reaches further.
Before, an image only PDF hit a wall. The record that referenced it stayed in an exception state, waiting for a human to do something about the missing text. Now that same document queues its own transcription automatically. The pages are read image by image in the background, the terms come out of the transcription, the record that raised the exception fills itself in, and the document becomes searchable everywhere it can appear, across the Archive, inside Migration Studio, and behind every upload door in the product.
Oversized agreements are handled too. A contract photographed at full resolution, the kind that is too large to read in one go, has its pages re rendered as compact images so the whole document goes through in a single pass rather than choking halfway. The practical effect is that the terms inside a scan are now available to the same machinery that reads a native PDF. If you want to see what that machinery does with a clean document once the text exists, decoding any contract in a minute covers the read that happens next.
The worst moment in a migration is committing a large transfer and then discovering half of it is scans that will need extra handling. This release moves that discovery to the front. During the survey, Migration Studio opens a sample of your PDFs in your own browser, with nothing uploaded to us, and tells you how many of them look like scans before you commit the transfer. You get a read on the shape of the work while your files are still sitting where they were.
That preview matters because it lets you plan. A batch that is mostly native text will land clean and fast. A batch that is mostly photographed paper will still land, but you now know it will spend time in background transcription first, so you can time the migration around a deadline rather than into one. This is the same discipline we apply to mapping your whole stack on one canvas, know the shape before you commit the effort.
The old cloud connection was polite and slow. It brought files across in a trickle, a handful per day, which is fine for keeping a live folder in sync and useless for onboarding an archive you have been sitting on for years. A connected Google Drive, SharePoint, or Box folder now has an Import everything button that walks the entire tree, subfolders included, and pulls it across in background chunks. You point it at the top of the folder and it works its way down.
Combined with self reading scans, this is what makes a real archive onboarding practical in one move. You connect the source, you press the button, the tree comes across in chunks, and any image only documents in it read themselves once they arrive. You set all of that up in one place, at settings and integrations, and then let the background work run. When it settles you have a searchable estate instead of a to do list, which is the precondition for anything useful, whether that is a review table or measuring switching costs before the vendor prices them for you.
Transcription is reading, not certainty. A scan of a clean laser printed page transcribes very well. A twenty year old fax of a fax, a coffee stained signature page, a handwritten margin note, these are harder, and the terms that come out of a poor source are only as good as the source. Treat the extracted fields on a low quality scan as a strong first draft that deserves a human glance, particularly for the numbers that will drive a negotiation. The original image stays attached to the record so you can always check it against the read.
Two more boundaries worth stating plainly. Import everything walks the tree in background chunks, which means a very large archive takes time to finish, it is a background job and not an instant load, and it runs alongside the rest of your background work rather than jumping the queue. And the local scan preview in Migration Studio is a sample, not a full audit. It tells you the likely proportion of scans so you can plan, it does not promise an exact count of every file in the tree. Both of these are deliberate. We would rather give you an honest read that runs safely than a fast number you cannot trust.
None of this replaces judgement. What it removes is the excuse for a gap between what you have signed and what you can search. Once the archive is in and the scans have read themselves, the terms are on the table, and the work you actually get paid for, comparing them against the market and holding the line at renewal, can begin.
Want to be updated when major licensing and pricing changes land? One analyst brief a week: the price rises, metric changes and audit campaigns that move software costs. Work email only.
Morten brings two decades of enterprise and software procurement, with stints across Oracle, IBM, SAP, and Salesforce shaping how he reads a deal. He has led sourcing through hundreds of renewals, from mid market order forms to nine figure global agreements, and learned that the buyers who win are the ones who walk in knowing the market. He built VendorBenchmark to make that pattern recognition repeatable.