}

AI agent or RPA in logistics processes: when a script is enough and when you need an agent

By
Bodo Buschick
28/9/26
•
7 min
AI agent or RPA in logistics processes: when a script is enough and when you need an agent

"Do we need an AI agent for this, or is RPA enough?" The question comes up in almost every first call. Usually from the managing director. Usually with a quote on the table that already contains one of the two answers. My answer is always the same. AI agent or RPA is the wrong question as long as it means the whole process. A process has steps. Every step belongs to one of two classes. Only one of them needs a model.

This is not a detail. Gartner expects over 40 percent of agentic AI projects to be cancelled by the end of 2027. The reasons: rising costs, unclear business value, weak risk controls. The same June 2025 note names "agent washing". Vendors stick the word agent on RPA bots and chatbots. Of thousands of vendors, Gartner counts about 130 as real. Start with "agent or RPA" and you are buying a label. Better to take the process apart.

AI agent or RPA: a logistics process has two layers

Every process we have automated at a forwarder falls into two layers.

The first layer is deterministic. Same input, same output, every time. Log in to the portal. Find the "Confirm" button by its selector. Import the CSV. Compare the quantity in the order with the quantity in the TMS. Send the mail with the attachment. These steps have one right answer, and it lives in the code.

The second layer is about understanding. The input varies, and the step has to interpret it. An order as PDF where every customer has its own layout. A mail without a subject line. It could be a complaint, a cancellation or a question. An address that is spelled differently in the TMS and in the order. No rule covers all the cases here. This is where a model earns its cost.

Anthropic put this split into words in December 2024. Workflows are systems where models and tools run through predefined code paths. Agents direct their own steps. The advice that goes with it: find the simplest solution first, add complexity only when needed. Agents trade cost and latency for better results on open-ended tasks. A login is not an open-ended task.

The build order follows from that. Script first. Model only at the step where the script provably fails. A protocol per step in both cases.

Layer one: why the selector is cheaper than the screenshot

Our confirmation agent for a customer portal made 39 runs in eight weeks. 35 of them clean. It is pure script. It finds the button by its name in the page code, not by an image of the screen. If the customer changes the layout, it still finds it. If the customer renames the button, the step aborts. That is exactly what it should do.

A model could find the same button. It would judge a screenshot on every run. That costs time and money. And it does not always come out the same. Gartner expects 60 percent of RPA vendors to add such "computer use" capabilities by 2027. That makes sense for screens that change all the time. A customer portal changes twice a year. There, the selector is the better choice. You can test it. You can read it.

The four aborts out of 39 runs prove the point. All four had the same cause: the browser was closed during the run. A model would not have prevented that. It would only have raised the bill.

Layer two: where the script provably fails

A different project. An official document we evaluate for a customer. It comes in two layouts. The new one is a table, parsed exactly. The old one is prose with numbers in it. The parser for the table is pure script. For the prose we built sentence logic. The reported deviations were then checked by a model in context, not by the script alone.

The result against the manually captured list: 143 of 145 rows correct. And a second finding that matters more. The automatic check reported 31 deviations. Only ten were real. 21 were parser artefacts, such as numbers wrapped differently in the PDF text. Without a human at the review list, 21 wrong corrections would have reached the customer.

That is the rule for layer two. The model reads what the script cannot read. But the model does not decide alone. It puts every case into a review list with a reason. The length of that list is the metric someone watches.

Demand for exactly this step is real. In 2025 Descartes surveyed 300 executives from transport and logistics companies in Europe. 41 percent use AI for automated data entry and unstructured information. Among companies that rate themselves as high performers, it is 61 percent. That is layer two. The order from the PDF, the mail without a subject.

The cost: a licence per bot versus code plus model

The cost logic of the two layers is different. That is the second reason to split the process.

Classic RPA is licensed per bot. One public list price as an example. IBM lists an unattended bot on the UK G-Cloud marketplace at 709.20 pounds per month. Date of the price document: 1 July 2025. UiPath shows an entry price from 25 US dollars per month on its pricing page. All enterprise plans are on request. The price hangs on the bot, not on the volume. Whether the bot confirms 50 or 5,000 orders makes no difference.

A model is billed per token. Anthropic lists Claude Haiku 4.5 at 1 US dollar per million input tokens. Output: 5 US dollars per million. A two-page order is roughly 3,000 tokens in and 400 tokens out. That comes to about half a cent per order. At 300 orders a day, around 1.50 US dollars. (Yes, that little. No, that is not the whole bill.)

The whole bill includes the code around it. The operation. The protocol. The review list. For us the effort sits there, not in the model call. The point is a different one. Model costs grow with the documents. The bot licence grows with the number of bots. Run 20 fixed steps through a model and you pay tokens for work a script does for free. Force one understanding step into a bot script and you pay in errors.

The failure path is the same in both layers

This is where the two layers meet again. Script or model, every step writes a protocol line. Timestamp, input, result. And every step has a failure path that is not "exception, abort".

For the script the failure path is clear. Selector not found, import with zero rows, reconciliation with a difference. The case goes into the review list with a reason. The run ends as a warning.

For the model the failure path is uncertainty. The model returns a field and a confidence. Below the threshold, the case goes into the same review list. Dispatch sees one list in the morning, not two. That is why we do not build agents that "just do everything". An agent without a review list confirms orders nobody has checked.

An honest caveat: the split is clean on paper. In practice, steps move. A customer switches its PDF to a fixed layout, and the model step becomes a parser. A portal adds a captcha, and the script step becomes a human. The split is not a decision for good. It is the state for this year.

Where both fail

In 2025 the BVL surveyed more than 200 companies from manufacturing, logistics and retail. 54 percent name data as the biggest hurdle. More precisely: its limited availability and quality. That matches our review lists. Most entries there are neither model errors nor script errors. They are missing article numbers. Quantities without a unit. Addresses spelled differently in two systems. No agent repairs master data. It only shows where the data is missing.

That is why the review list is the most important result of both layers. It shows which customer delivers which data. That is a conversation for sales, not for IT.

The question for your process

Take one process with manual work. Order capture, portal confirmation, shipment notices. Write the steps one under the other. Mark every step where the input looks different every time. Those are the steps for a model. All the others are a script.

Is order capture from PDF the one that hurts most? Then send me one anonymised order. I will show you on our logistics example which steps the script takes. And where the model reads. And what the review list looks like afterwards.