Saffron Automations

Home · Field Notes

How to audit an automation vendor before you sign

Every demo you will ever be shown is a best-case run. These are the four questions that find out what the average run looks like — and the four links to run all of them on us first.

31 August 2026Saffron Automations6 min read

Buying automation is unusual among software purchases because you mostly cannot see what you bought. A CRM either has the field you need or it does not. An automation runs unattended, at three in the morning, against your live billing system, and the only evidence that it did the right thing is whatever record it kept — if it kept one.

That gap is why the sales process for this category is so heavily weighted toward demos. A demo is a controlled run: chosen input, clean data, someone technical in the room. It tells you the system can work. It tells you almost nothing about what happens on the four hundredth run, when the input is malformed and nobody is watching.

The four questions below are the ones that close that gap. None of them are technical, none require you to understand the stack, and all four can be asked in a first call. What matters is not the answer — a good salesperson has an answer for everything — but whether the vendor can produce the artefact behind it while you wait.

Question 01

“Show me the run you would rather I did not see.”

Ask for the worst result they have on record. Not the average, not a case study — the specific run that went badly, with its date.

This question works because of what it takes to answer it. A vendor who measures continuously has a bad run to hand, because measurement that only happens before a pitch produces exactly one kind of number. A vendor who measures for the pitch has to go and manufacture something, and you can hear it happening.

The follow-up matters as much: what did you change afterwards, and did the number move? A fix that was tried, measured, and found not to work is worth more to you than three that were shipped and never checked, because it proves the loop is real.

What a real answer sounds likeHere is the run, here is the date, here is the method we measured it with, and here is what we changed. It moved the number by this much — or it did not, and we reverted it.

What to listen for insteadDeflection to averages, to satisfaction scores, or to a case study with no date on it. “We have not had any failures” is not a good answer. It means they are not looking.

Question 02

“Who can touch my data — by name?”

Automation that is worth buying reaches into billing, CRM, scheduling and financial records. That data does not stay in one place. It passes through model providers, hosting, queueing, logging, error tracking, email delivery — each of which is a company with its own staff and its own jurisdiction.

So ask for the list. Not a compliance badge, not a framework name: the actual register of sub-processors, each with what it does and where it sits. A vendor who has one can send it before the call ends. A vendor who has to build one has never mapped where your data goes, which means they cannot tell you what happens when one of those companies has an incident.

This is also the cheapest question to verify, because a serious vendor publishes the list rather than making you ask.

What a real answer sounds likeA named list, purpose beside each entry, region beside each entry, and a stated notice period before it changes.

What to listen for insteadA certification in place of a list. Certifications describe how a company manages risk; they do not tell you who is holding your customer records this afternoon.

Question 03

“What happens when it is wrong?”

Not if. Every system that touches the real world is eventually wrong about something — a duplicate charge, a message to the wrong contact, a booking held against a slot that closed. The question that separates vendors is whether being wrong is detectable and reversible, or silent and permanent.

Three sub-questions do the work here:

  • Is every action logged, including the ones that failed? A log that only records successes is a marketing artefact.
  • Can a specific action be replayed and reversed? If the answer is “we would restore from backup”, the blast radius of a small error is your whole dataset.
  • At what point does it stop and ask a person? There should be a threshold, it should be written down, and it should be one you can set.

The last one is where most of the risk actually lives. A system with no escalation threshold is not more autonomous; it is just less supervised. The correct shape is narrow, explicit scopes, with anything outside the mandate escalating to a human before it executes rather than reported after.

What a real answer sounds likeEvery action writes to an append-only log with a timestamp and a signature. Here is how you replay one. Here is the escalation threshold, and here is where you change it.

What to listen for instead“It has a 99% accuracy rate.” That is a claim about the average run. You are asking about the other one.

Question 04

“What do I keep if we stop working together?”

Ask this in the first call, not the last. The answer tells you how the vendor thinks about the relationship, and it is the single largest hidden cost in the category.

You are looking for three things: the data itself in a format something else can read; the record of what was done, because that is your audit trail and it does not belong to the vendor; and a clear statement of what stops working on the day you leave. Some of it will stop working — that is honest and expected. What matters is whether you were told which parts, before you signed.

What a real answer sounds likeYour data and your evidence log export in a standard format on request, at any time, without a support ticket. Here is exactly which functions cease.

What to listen for insteadExport as a paid professional-services engagement, or a pause while somebody works out whether it is possible.

A vendor who cannot show you a bad run is not telling you they have never had one.

Now run all four on us

It would be a strange article to publish and then decline to answer. So here are the four artefacts, and you are welcome to start with us before you use them on anyone else.

If you want a fifth data point, the Assay runs a full diagnostic of your operation and shows you the findings before there is any engagement at all. It is free, and the output is the same output a client gets.

The same discipline applies on the other side of an engagement, with different questions. Our sister company wrote the matching set for web proposals: how to compare website quotes — published prices, a number that can fail, who owns the asset, and whether the build pays back.

Published 31 August 2026 by Saffron Automations, a division of Saffron Group. Corrections and disagreements are welcome — if any claim on this page stops being true, it gets changed here with the date on it rather than quietly removed.