Skip to content
Back to the blog
Agentic AI

How we watch a warehouse of 30 million operations

Twenty-four rules against live data, thresholds measured from the warehouse's own history, and the 7:06 alert that names two specific locations. Including the number we still do not publish.

An order can sit in the ERP and be absent from the warehouse without either system raising an error, because each one squares with its own data. This post covers how that gets caught by crossing the two live, where the thresholds that decide what counts as odd come from, and what it does not prove.

The order that is in the ERP and not in the warehouse

One system holds nine orders and the other holds eight. Each list squares with itself, so neither of them raises an error.

The ERP has its order, with its date and its customer, and as far as it is concerned everything is fine. The warehouse management system has its work queue, and that order is not in it, so nothing is failing there either.

Neither one is wrong, and that is exactly the kind of thing that only shows up with a model of the business sitting above the systems. We cover that in how we build a company's ontology. The order only disappears when somebody puts the two side by side, and until that happens nobody has a reason to look.

The case page names three more from the same set: open replenishments, abandoned picking and orders on hold, that last one needing a caveat the case states only for transport: telling what was held on purpose from what was held by accident. Otherwise the alert list fills with things somebody already decided and stops being read.

None of those four raises an error, and all four cost money. But they are not the same: the warehouse system can spot the first three by looking at its own data. The last one neither system can see, because it is only true once the two are crossed.

What the warehouse flagged between 6:58 and 7:06

That failure happened on one specific morning, published with the case, and in eight minutes the agent sent four lines.

The image above covers a whole morning, and these are the four lines we publish. At 6:58 the morning picking wave closes. At 7:02 1,180 packages are verified against order. At 7:04 shipments are assigned to a carrier. At 7:06 two locations do not reconcile and the shift manager gets the alert.

The first three are informational. They say what happened and they ask nothing of anybody. The fourth points at two specific locations and sends a person to look at them. It is the only one that can be wrong, which is why it is the only one worth having.

The fourth line of that morning does not say there are discrepancies. It says two, and it brings the record for each one with its export. An alert that gives the count without the cases is the same work as before, with a screen that remembers it.

A warehouse is not one place. It is five: returns, storage, picking, replenishment and dispatch. The split is ours, not the case's: we publish rules, not areas. And the failure this post is about falls into none of the five.

Mapped onto the route goods take, the four published rules fall in different places. Abandoned picking falls in picking and open replenishments in replenishment: the warehouse sees those by looking at its own data. But the order that sits in the ERP and never arrived is in no area at all, and that is the difference that matters: a failure inside one area is visible to one system; a failure between two systems is visible to neither.

Which four systems we read from before anything is switched on

This warehouse has more than 30 million operations on record, spread across four systems that do not talk to each other: the warehouse management system, the ERP, the carrier data and the commercial master data. The first step is looking at what is in each one, and which of them can be read from today.

Our agents read four sources:

  • The warehouse management system, live. That is what lets an alert at seven in the morning be about seven in the morning rather than about yesterday's close.
  • The company's ERP, which is the other side of the cross-check.
  • The carriers' data, for what happens once it leaves the warehouse.
  • The commercial master data, which is what turns an SKU into a customer with a commitment.

Those four are enough to build the cross-check. What they are not enough for is putting numbers on it, and that is the other half of the conversation.

Where each threshold comes from: that warehouse's history, not a table

Separately from the twenty-four rules, we watch transport and flag "carriers slow to collect". That "slow" is drawn from this warehouse's own history. It comes from measuring what normal looks like here, and a carrier that counts as slow here would be perfectly ordinary somewhere with a different delivery profile.

The arithmetic is all there is to it, and an example shows it. The numbers that follow are made up. Each site's real ones are not published. If 95% of collections in a warehouse's history close within forty minutes, forty is the threshold, and a collection at fifty-five falls outside it.

What makes that forty useful is that it came from there. In a warehouse serving rural routes the median may sit at fifty-five, and applying the same threshold would mean alerting on normal every morning. The number is in no vendor table: it comes from that warehouse's own percentile, which is why it cannot be shown before it has been measured.

One case makes this rule hard, and it separates a system that works from one that does not. An order created in the ERP and pushed to the warehouse minutes later is, for that gap, indistinguishable from a lost order. Both look the same: present in one place, absent in the other. A rule that does not account for that lag alerts every night about orders that are perfectly fine, and by the third night nobody reads it.

What we publish is the outcome, not the method: that the system separates what was paused on purpose from what is genuinely late. How we handle the lag we do not publish, so it does not get described here.

Put the other way round, which is how it helps when buying: thresholds shown before anybody has looked at that warehouse's data are another warehouse's thresholds.

The twenty-four rules that check the warehouse, and the four published

Twenty-four rules run as standard and only four of them are named in public.

They run against the warehouse's live data, and the four named publicly are: orders on hold, abandoned picking, open replenishments, plus orders that exist in the ERP and never reached the warehouse.

The fourth is the interesting one, and it is the one from the opening:

Published ruleWhat it catchesCan one system see it alone?
Orders on holdOrders stopped that nobody releasedYes, the warehouse system, from its own queue
Abandoned pickingPicks started and never finishedYes, the warehouse system
Open replenishmentsGaps nobody refilledYes, the warehouse system
Orders in the ERP that never arrivedAn order that exists in one system and not the otherNo. It takes crossing the two

A system looking at itself can spot the first three. Neither system can spot the fourth, because each is right about its own data.

Crossing two systems every night is not a job for a person. The agent does it without rest. Every incident it opens comes with the exact list of affected orders, with its details and an export. The detection is done by the data. The write-up is done by the model, for whoever has to read it.

There is no model detecting anything here. Twenty-four rules over live data do the detecting, and the model does the last step, turning a rule output, something like "twelve orders meet condition 4", into a sentence a shift manager understands at seven in the morning without opening a table. When this post says "the agent", it means that pair: rules that watch, a model that writes.

That split is the design decision holding up everything else. A rule can be read, argued with and changed; its threshold is a number somebody can point at, so when an alert comes back false it is clear which rule opened it, and that rule gets fixed. With a detector nobody can open, a false positive is an anecdote: there is nothing to correct, only something to try again. It has a price, and it is the limit set out below: a situation nobody wrote a rule for produces no alert at all.

How the false alert gets avoided: telling paused apart from late

The failure mode opens this post. Here is the other half: what the system does to stay out of it, which is where the real craft is. It is in not interrupting.

Four things get watched on transport, all four published. And all four are the same trap: carriers slow to collect, shipments stopped or lost, deliveries outside their usual window, returns above what is normal. The three words doing the work in there are "slow", "usual" and "normal". None of the three exists outside this warehouse.

Two destinations with different usual lead times cannot share a threshold, or one of them alerts every day. And a shipment stopped at the carrier is, in the data, identical to one somebody stopped on purpose. That is why the system separates what was paused on purpose from what is genuinely late. In the data those are identical. To whoever gets the alert they are opposites.

There is a second limit to be aware of before signing anything: the system does not know what it was not told to watch. The rules cover what they cover. A new situation produces no alert at all, and no alert looks a lot like everything being fine. The twenty-four rules cover what they cover, and only four of the twenty-four are named: nobody outside can judge the gap.

What else comes out of the same cross-check: productivity, and the size of what is watched

Watching for incidents is half of it. The other half is that the same cross-check measures the work.

The agents measure productivity per person and per machine against target, calibrated with the warehouse's real data, together with delivery times simulated day by day. It is the same cross-check used the other way round: instead of hunting the missing order, it compares what each shift did against what this warehouse's own history says is normal.

We publish it at the same level as incident watching, not as an extra.

And the same cross-check reveals the scale of what sits underneath. These are the figures we published on 20 August 2026.

30M+

warehouse operations recorded

900,000+

shipments processed

460,000

locations under control

They state the size of the problem, not the quality of the answer. The quality is at 7:06, when those two flagged locations turn out to be exactly the right ones.

What the agents do not do: they do not move an order or change a priority

"The agents detect, explain, and warn. They do not move an order or change a priority: that stays in the hands of the warehouse."

That split between what an agent may do and where it stops is the same one everywhere, and we cover it in full in read, alert, execute.

A system that reorders a picking wave on its own is a different product, with a different risk, plus a conversation still to be had about what happens when it gets it wrong at six in the morning. The limit is published precisely so it can be held against us: written before the signature, it is enforceable afterwards.

Who watches that warehouse, and what else they run

This warehouse is not the only system we have running. There are four, each with its published state, and the warehouse one is the quietest of the four.

FAQ

Does the warehouse management system or the ERP have to be replaced?
No. The agents read the systems already in place and work on top of them. What gets built first is a model of the business that crosses the two, not a system that replaces either.
How is the false-alert rate measured?
It is called false-on-closed: of the alerts a person closed last quarter, how many were closed as nothing here. The denominator is closed alerts and not opened ones, because counting opened ones lets a quarter with many left unreviewed score well without anyone doing anything. It is judged by whoever went to look, not by the system, and it is read per rule and not only in total.
Can an agent move an order or change a priority?
No, and that is published with the case. The agents detect, explain and warn; moving an order stays with the warehouse. A system that reorders a wave by itself is a different product and a different risk conversation.

What this post does not prove

All of the above comes from one installation, and that is the main limitation: one automated warehouse with its delivery profile, its carriers and its shifts. What travels to another site is the shape of the failure and the rule set, not the experience.

The weakest part of the demonstration is the four named rules. They are four out of twenty-four, and the other twenty are not published, so nobody outside can judge whether the coverage is good or whether the four published happen to be the four that look best.

The bias has a direction, so here it is: we published those figures ourselves. No third party has checked them. If they lean, they lean our way.

And there are four specific things this text does not cover. They are exactly the ones that close the post, turned into questions.

So the correct reading is this: this reads as evidence of three things: the failure exists, crossing two systems detects it, one warehouse runs it in production at that scale. It does not read as proof that it pays for itself. That would need the number from before, and we have not published that number.

Two objections press harder than anything above.

The first. This company publishes a timeline to the minute, four named rules and its size figures. So why is the one number that would judge the company itself the one missing? The sceptical reading is the obvious one: it stays quiet because the number is not good. That cannot be ruled out, and saying "nobody publishes it" does not answer it, it dodges it. What can be said is what there is: the definition is written and it is the one in this post, but the number measured with it is published on none of our case pages. Owning the yardstick and withholding the mark is harder to defend than having no yardstick. A reader who marks the post down for that is marking it down correctly.

The second is technical and it touches the foundation. The whole mechanism assumes that matching an ERP order to a warehouse order is easy, and that the hard part is tuning the threshold. In many installations it is the other way round: the two systems carry different keys, partial references and manual re-entry, and matching well is the expensive work. This post does not demonstrate that part. It shows the threshold part, which is the one the case publishes.

The false-alert rate, and why nobody publishes it

A shift manager sent to check discrepancies that are not there stops going.

One number says whether that has happened, and here it has a name: false-on-closed. Of the alerts a person closed last quarter, how many were closed as nothing here. The name carries the only part worth arguing about, which is the denominator. Over closed alerts, not opened ones: count the opened and a quarter with a pile left unreviewed scores well without anyone doing anything. And it is judged by whoever went to look, not by the system.

The schematic above shows the arithmetic with small numbers: of nine alerts opened, three nobody looked at, four with something behind them and two with nothing; the rate is those two over the six that got closed. Nobody in this sector publishes it. Before writing this, our four published case pages were checked one by one: we do not publish it either. So this post opens by failing the question it goes on to pose at the end.

What can be shown is the rest, and it starts with the failure that has to be detected.

What none of the above changes is the mechanism. An order that sits in one system and not in the other is still invisible to both, and it still only shows up when somebody puts the two side by side.

And that mechanism is not about warehouses. Two systems that each square with themselves and not with each other happens in any company running more than one system: an ERP and a warehouse management system, a sales tool and a billing one, a case file and the inbox. The table names change and so does what counts as normal there. What does not change is that nothing raises an error.

Four questions for anybody selling a system like this

Not one of them needs technical knowledge. The first three are answered by this post. The fourth is not.

Where does each threshold come from? An answer that does not mention that warehouse's own history is describing somebody else's numbers.

What happens when the alert is wrong? If there is no answer, the shift manager will provide one by ignoring it.

What gets measured before anything is switched on? The value of a system like this is a comparison, and there is no honest comparison without a baseline. If nobody measures the baseline, a year from now nobody can say whether it helped.

And the fourth is the one this post opened with, because it is the one that really decides, and it is the one we ourselves fail to answer today:

What was the false-on-closed rate last quarter? That is the number that says whether the system is still being read six months in, and none of our case pages publishes it. Anybody who puts it on the table has something we do not have yet. If nobody puts it on the table, the question for the second meeting is already obvious. All of this comes from one installation, four rules out of twenty-four, and figures we publish ourselves.

Sources

  1. 1.Published case: the warehouse watched in real time, with its figures, its twenty-four rules, the morning from 6:58 to 7:06 and its declared limit. In production, 20 August 2026. The case pages exist only in Spanish, so no English link is given here.
  2. 2.The four systems, each with its state and its declared limit, and the twenty-four rules

Share this post