}

On 30 July a customer wrote to us. Two days without shipment notices. The job had run. Status green, zero notices, 4.3 seconds. A week later, a tour export. The mail arrived at 17:00. The CSV was missing. Green again. In June an import ran for ten days with exit code 0. Not a single tour reached the target table. No step protocol in any of the three.
My position after this summer: a run with no error and no result is an outage. Treat it like a crash. "The job ran" is a sign of life. "The job delivered" is proof of health. Between the two sits a step protocol: one line per step, plus one number that measures what the run is for.
This is not carelessness. It is how the tools work. Take the UiPath Orchestrator documentation. A job is Successful there once the robot has executed it correctly and it has finished. Faulted means: did not start, or threw an unhandled error. Both are states of the program. Neither says whether a file was attached.
That is true for every vendor. Cron jobs, Windows tasks, workflow monitors: all of them ask the same thing first. When did it last run? With which exit code? And there are more of them every year. Gartner puts the RPA market for 2024 at 3.8 billion US dollars. Up 18 percent on the year before. More unattended bots than ever. Most of them are judged by a timestamp.
Google described the problem ten years ago. The SRE book separates symptoms from causes: "What's broken, and why?" Monitoring should show the symptom. The thing the user notices. For a dispatcher the symptom is not "process crashed". It is "no shipment notice in the inbox". That symptom is missing from almost every monitor I have seen at a forwarder.
The cases come from our own operations, anonymised. I list them because the patterns carry over.
First, the shipment notices at the end of July. For one due day, one of the two input files was missing. The run skipped the day in silence. No log line. The report said "no open days". The customer found it. The system did not.
Second, the tour export on 6 August. The mail reached the partner. The CSV attachment did not. A check job had noticed and written a warning. The warning mail never arrived. The mail channel was dead. Two systems, both "running".
Third, the import from 26 June. A retired module caught the call. Exit 0. The task stayed green for ten days. The only hint, in hindsight: the runtime fell from 5.5 to 2.4 seconds. Nobody looks at the runtime of a green job. A step protocol would have shown the missing step "rows written".
Fourth, a mail campaign from 11 August. A DNS hiccup the day before killed exactly the job that would have reported the gap. Follow-ups kept going. First mails stopped. The watchdog reported healthy. Weekly cadence, so a week unnoticed.
Fifth, a reporting job for a customer. The central monitor raised 20 alerts in ten days. Not one alert mail was delivered. The sender account was blocked. The alert itself was the silent failure.
(Yes, that is five. Yes, all in one summer. That is why this post exists.)
Older numbers support the picture. In 2020 Forrester Consulting surveyed firms running RPA on behalf of Tricentis. The result: 45 percent deal with bot breakage weekly or more often. The study is six years old. I have not found a newer one with the same method. The mechanics have not changed. Bots break on changed screens and missing data. And the break is often silent.
We turned this into a rule. It sits in our operations handbook. It applies to every new automation that runs without supervision. Three parts.
First, the expected output, declared when the process is set up. Exactly one number per run. Notices sent. Rows imported. Orders captured. It sits under a fixed key in the run's protocol. With it, a minimum value and an allowed idle window. In our watchdog the fields are called output_metric_key, min_output and max_idle_hours. Zero as a target is allowed. But only if it is declared.
Second, the idle reason. When a process runs without output, it writes the reason into the protocol. Machine-readable: queue empty, input missing, weekend. If the idle stretch exceeds the window, the alert fires. Same as a crash. A reason in the log without a threshold above it does not count. This is where most monitors stop. They log "nothing to do". Nobody counts for how long.
Third: a negative result is not a success. A bounce is a rejection. So is "access denied" from a portal. So is a gate that stops a send. Such runs do not end as success. At least as a warning. Or in their own metric that someone watches.
The step protocol carries all three. One line per step. Login. List loaded. File found. File attached. Mail sent. Every step with a timestamp and a result. If the step "file attached" is missing from the protocol, the run cannot be green. The shipment notice case would have surfaced on the first run.
Our watchdog knows 98 watches today. 48 of them carry an output signal: number, minimum, window. Six have declared idle reasons. Three have operating hours. That is the state on 9 September. Not the end state. The retrofit ran in two waves in August. Each watch with a simulation over 30 days. That way new thresholds do not start with permanent false alarms.
The false alarms came anyway. And they were useful. A partner pull fetches tours for today and yesterday every 15 minutes. Nobody drives on Sundays. So the watchdog reported "running but delivering nothing" every Sunday. 96 runs with zero rows. The reflex would be a 48-hour threshold. Then the alert is quieter. The risk is not. The right answer was a field for working days. An empty Sunday does not count. An empty Tuesday does.
Google has one sentence for this. "Every page should be actionable." Every alert needs an action. Otherwise it gets ignored. And an alert that goes quiet only because the time threshold went up hides the next silent failure. Unless the output signal moves with it.
These are five cases from one operation with a few dozen pipelines. Not a market study. How often silent failures happen at a carrier with 300 trucks, I do not know. What I do know: in all five cases the tool showed green. A human found the outage. Usually the customer.
And not every process has an output number. A mirror between two systems, for example. Master data from the database into a config file. There is no meaningful daily count. There, the target-actual check takes its place: number of deviations, measured at the target. Expected: zero. Every deviation alerts. Better still is to automate the manual transfer step entirely. The watchdog is the net underneath. Not the solution.
The browser agent from the last post reports two numbers per run. Orders checked. Changes booked. Zero orders checked on a working day is an alert there. Not a success. That is the whole rule in one sentence.
Which of your automations has a target number per run? And which one only reports that it ran? Name one process. I will show you on our logistics monitor what number, reason and window look like in its step protocol.