All articles

You Can't Automate a Process You've Never Run

You can't automate what you don't understand: you learn a process by running it, not documenting it. Automate an unrun one and you industrialize a guess.

Paweł Bazyluk
Paweł Bazyluk Founder Athru IT & partner at Spyrosoft Innovo S.A.

A flowchart of a process you have run a hundred times and a flowchart of a process you have only imagined are the same drawing. Same boxes, same arrows, same little diamonds where a decision splits the path. Put them side by side and nobody can tell which is which, because the entire difference between them lives in the part the diagram cannot hold.

The industry has a line for this: understand your process before you automate it. Fair enough. Then it defines understanding as an artifact you can make in an afternoon: a document, a swimlane map, a task-mining recording of someone else at their desk. All three are things a confident person can produce about work they have never once done.

Automation does not create knowledge. It copies the knowledge you already have and runs it at volume, without tiring and without asking questions. Hand it a real understanding of the work and you scale something that works. Hand it your best guess and you get the guess, faster, wearing the confident face of a system that looks like it knows the work.

Before I make that case, I have to clear away the argument this one is always mistaken for.

Unrun, not broken

You have heard the warning. Don't automate a broken process, or you scale the mess and ship the dysfunction at machine speed. Good advice. Not this piece.

Can you automate a broken process? Sure, and you shouldn't. But a process can be well-designed on paper, efficient, sensibly sequenced, reviewed by people who know what good looks like, and still be a guess. Being well-designed and being run are unrelated properties. One is a fact about the diagram. The other is a fact about you.

There is a whole academic field built on that gap. Wil van der Aalst founded conformance checking, in process mining, on one stubborn observation: the process models organizations design routinely diverge from what their own event logs show actually happened. The gap is modeled versus lived.

So the danger is not a bad design. What, then, is missing when all you have done is document a process?

Understanding is not an artifact

Should you automate a process before you understand it? Everyone agrees you shouldn't. The unspoken disagreement is about what "understand" means, and the market has settled on the cheapest answer going.

Look at what each "understand first" method actually asks for:

  • A map of the workflow.
  • A plain-language description of inputs, outputs, and exceptions.
  • A screen recording of someone at their desk, reconstructed into the "real" process.

Every one is an artifact. Every one can be produced, in convincing detail, by a person who has never run the work once.

This is the reframe the whole piece turns on. A documented process is not a run process. Van der Aalst's conformance checking is the formal version of the point: the model and the behavior in the logs are separate objects. They have to be checked against each other, because they drift. The nurses are the human version. Study how care is actually delivered on a ward and you find staff routinely working around the documented protocol, not from carelessness but because the written procedure never anticipated the case in front of them.

So name what the artifact is missing, and why the only way to get it is to have done the work yourself.

The residue only running leaves behind

Here is the load-bearing claim, and I am going to argue it rather than assert it. The knowledge automation needs most is the knowledge that never made it into the document. It was never the kind of thing a document can hold.

Michael Polanyi named this sixty years ago: tacit knowledge. His one-line version is "we can know more than we can tell". His example is one everyone has lived. You can pick a familiar face out of a crowd of a million without being able to state, in advance, the rules your eye used to do it. The knowing is real. It is also not extractable as a list.

David Autor carried this straight into our subject. In his work on Polanyi's Paradox, he shows that human tacit understanding of work routinely runs ahead of anything we can write down as rules. That gap is exactly why some tasks resist automation even when they look procedurally simple. Codifying what the person doing it knows is the hard part.

Now put a real person in it. The nurse working around the documented protocol is handling the specific patient the policy never anticipated, with the supplies she actually has. That is where the tacit layer lives: in the edge cases and exceptions in workflows that only surface on the job. The case that shows up on the third Tuesday. The field marked optional that quietly breaks everything downstream when someone leaves it blank.

None of that is in the diagram. It is residue, and residue is the one thing you cannot download from someone else. You get it by running the work, or you do not get it.

Automate on the artifact anyway and you have made a mistake at industrial scale.

A guess, at volume

Let me make "industrializing" concrete, because it is doing real work in the thesis. A guess run once is a mistake. You catch it, you curse, you fix it. The same guess run ten thousand times, without tiring and without asking questions, is a pattern of failure you paid to manufacture.

This is part of why automation projects fail more often than the brochures admit. Automation does not create knowledge; it multiplies the knowledge you fed it by throughput. Feed it a wrong assumption and throughput is exactly what you get back: amplifying errors at scale, faithfully, at the speed you bought.

The failure data is not subtle. EY, in its 2016 report "Get Ready for Robots," found that as many as 30 to 50 percent of initial RPA projects fail. Not underperform. Fail. Deloitte's 2017 global survey adds the texture: among the organizations that had actually implemented RPA, 63 percent said the speed of implementation they expected had not been achieved. These are stories about people who automated a version of the work that did not match the work.

Volume is only half the trap. The other half is that the output looks authoritative enough that nobody goes back to check it.

Confident enough that no one checks

The second half of industrializing is quieter, and older than the machines we are worried about now. Automated output carries an authority a human draft never gets, and that authority suppresses the double-check. The guess is running at volume, and running unexamined.

This is one of the most durable findings in human-factors research, not a fresh observation about AI. Raja Parasuraman and Victor Riley set out the framework in 1997: operators misuse automation by over-trusting it and failing to monitor it, precisely because its outputs present as authoritative. A screen that states its answer plainly gets questioned less than a colleague who says the same thing.

Parasuraman and Dietrich Manzey sharpened it in 2010 into the exact distinction this argument needs. Complacency is merely leaning on the automation because it usually works. Automation bias is the harder failure: not independently verifying the automated output even when there are concrete reasons to check it. Nearly three decades of peer-reviewed work point the same way. The confident-looking machine is the thing least likely to be questioned, which makes it the perfect vehicle for a guess.

A rules engine at least fails loudly on the case you forgot. Generative AI does the opposite.

AI doesn't fail loudly. It fills the gap.

Here is the sharpest edge of the argument, and the reason process understanding before AI automation is not the conversation it was five years ago.

Classic automation is honest about its own ignorance. Hit the case you never encoded and it throws an error, halts the run, flags the missing field. The failure is loud, and it points more or less at the exact spot where your understanding ran out. Annoying, and a gift.

A generative AI workflow does the opposite. Hand it the gap in your guess and it does not stop. It fills the gap with fluent, plausible, confident output, in the same register as the parts you actually got right. The guess is camouflaged. The system industrializes the guess and then hides it inside prose too smooth to trigger a second look.

This is measured, not speculated. In a 2026 field study from Harvard Business School and the BCG Henderson Institute, "GenAI as a Power Persuader," researchers watched more than seventy BCG consultants try to validate model output on a real business problem. The more the professionals fact-checked and pushed back, the harder the model defended its original answer, sometimes a wrong one, escalating its persuasion instead of surfacing the error. The model argued.

If the failure hides this well, the only real defense is upstream. So how much running is actually enough before you build?

How much running is enough?

How do you know a process is ready to automate? Not by completing it once. One clean run proves only that the happy path exists. The work is everything that deviates from it.

So here is the bar, and I will not dress it up with a number no one can defend. You have run a process enough when new runs stop teaching you new exceptions. That is the whole test. Keep going until the surprises dry up: until you have hit the case that breaks the template, felt the point where the procedure hands off to your judgment, and caught yourself skipping the documented step because this time it plainly did not apply.

This is also the honest answer to when not to automate a process. Not yet, if every run still surprises you. The exceptions have not finished arriving.

There is one process this rule looks unable to cover: the one nobody has ever run.

But what about a process no one has ever run?

Here is the objection that could sink all of this. What about a genuinely novel process, one with no manual precedent, a workflow no human has run in this exact shape? You cannot gather residue from runs that never happened.

The objection does not break the rule. It names the method. The answer is: run the smallest manual version first. Do things that don't scale until you have the residue, then build.

Paul Graham built an essay around exactly this, and his line is the one to keep: do by hand the things you plan to automate later, because "when you do automate, you'll know exactly what to build".

  • Airbnb's founders went door to door in New York signing up hosts and shooting listing photos themselves.
  • Stripe ran the Collison installation, setting up each early user's payment integration by hand on the spot instead of mailing a signup link.
  • Pebble hand-assembled its first watches and learned operational detail a spec never surfaces, down to how much it mattered to source good screws.

DoorDash is the cleanest case of no precedent at all. On-demand logistics across many restaurants did not exist as a runnable process, so the founders ran deliveries themselves for six months before building the dispatch. Tony Xu's own account:

After six months of doing deliveries, you learn it's really complicated... Software can solve these problems.

The novel process is the rule at its purest. Someone still has to run it first. You have to be willing to be that someone.

Pouring concrete over a guess

So the sequence is not negotiable. Run the process by hand. Run it enough times to hit the exceptions and feel where the judgment lives. Then, and only then, do you know enough to automate it.

Skip that, and be clear about what you have actually built: you have poured concrete over a guess and called it infrastructure. In the AI era the concrete sets fluent and confident, smooth enough that you may not notice the shape of the mistake for a long time, long after it is expensive to break up.

You can't automate what you don't understand. That was always true. The trap was believing you could understand a process by describing it. The question is whether you have run it, not whether you can describe it.

Sources9