Wanna know why that cool new thing crashed and burned when you put it in the field?
The renders, promo videos, spreadsheets, and live demo all looked great. The math worked. The team you’re buying from is smart, motivated, and knowledgeable. Systems integration goes well, physical deployment goes smoothly, operators and managers receive deep and thorough training, everyone knows what the job is and how to use the new thing. And then it goes straight to hell.
The usual fingers get pointed, some folks probably get fired, maybe litigation ensues. But even after five-why’ing1 it, ultimately the root cause seems indeterminate. People just shrug their shoulders and say “it just wasn’t a good fit for us”, when the true root cause tends to be that you bought a technology based on the happy path performance.
We talk a lot about “happy paths” and “exception paths”. We don’t talk about what falls in between. What happens when a process mostly goes okay? What happens when an exception path grows big enough that it has to be operationalized? Those bits land in the “mushy middle”: a no man’s land where your technology goes to die.
Let’s visualize that with an extremely sophisticated whiteboard sketch. Intentionally bereft of numbers, and the particular positioning of the delineators changes with the specific process and point in time. But what’s important is that the area under the lines are where your operational time actually gets spent.

The head and tail of this sketch explain themselves. The highest frequency processes make up your happy path and that is what you’ve got in mind when you think about a given process path. If you’re in a warehouse and picking an item, this is where your picker gets an instruction, they go to the place the item is, they pick up the item, they scan it, then they put it in its next location. The exception paths are also what you’ve generally got in mind: maybe the item isn’t there, or the item is broken, or your picker is tired of the scanner beeping and has walked out and decided to join an ashram. In any case, you’ve got a clear judgment call that a human being needs to make2 on what happens next and how to recover from it.
Detecting the mushy middle in a manual process is devastatingly difficult precisely because human beings are so good at handling it. Anytime you encounter the Midwestern “ope!”3, that’s a human being realizing that their Plan A wasn’t working and then coming up with a way to get to the same place with a little bit of improvisation.
You can’t automate what you don’t know about. Furthermore, instrumenting everything to 100% coverage is infeasible so unless you’re out there in the field, living the process, you’ll never discover it. The liminal space of systems engineering, the mushy middle breeds informal process and empirical guesswork because the system knows less about the world around it than the human beings do, and the human beings don’t understand why the system doesn’t handle it well either4.
Even worse, the delineations do not stay static. They evolve and move over time with the business. You may get lucky and be able to improve the purity of inputs into your system and increase the amount of process falling into happy path: for example, labeling improves on your items so they scan more consistently. Or you may get unlucky and have an exception path blow up in volume and operationalize: for example, the average cube of your items has increased and you no longer have the right proportion of bin sizes to handle them. A system conceived five years ago may not be a good match as-is for today’s realities5.
This pattern is why startups, particularly robotics startups, tend to fail. A great proof-of-concept makes the sale. Seed through Series A sees engineers out in the field picking up the pieces in the exception paths. Super cool, everyone’s optimistic! Series B, when you’re looking to scale to market, is where the mushy middle problems come to bite you. You’ve got hyper focus on the happy path, you’ve started instrumenting the exception paths, but you’re soaking up the mushy middle with a bunch of guys wearing “Customer Success Team” hi-viz vests onsite6 trying to bridge the problems day in and day out. At some point, either the customer tires of seeing these vests on their site or the startup runs out of money, and you get ganked out.
Larger, more mature organizations suffer similarly, though they tend to use phase gates or some other sort of formalized product development process rather than funding rounds to describe technology maturity. No amount of “insurgent thinking” or “moving fast and breaking things” or “being agile” will overcome it. Larger organizations aren’t better at this, they’re just more able to absorb the failure and less likely to place harsh accountability on the decision-makers7.
So what do we do about it?

Design for resilience first. What happens when your technology is presented with crappy inputs or environments? Your system can have hangovers, too. Systems that can degrade performance rather than hard fail soak up lots of that mushy middle, and give you time to learn, observe, and iterate. You may not be able to pull this off for your proof of concept or your demo, but plan for a rearchitecture phase sometime in your Series B/Alpha/TRL 6 and start writing down your thoughts on how it ought to be done.
Get in the field. And not just on what I call “Ops Safari” either, where you put on a plastic hi-viz vest with the creases still in it and just take polite photos. Get into process and get dirty.
Discover the mushy middle. Spending quantity time eating shit as a picker or a tote wrangler will get you the opportunities to find all the little annoying things that human beings just handle. Map it, quantify it, count the “opes”.
Understand what pushes the delineations to the left. Now that you’ve been down in the trenches, get back to your desk and understand the business trends at play. Where are the sensitivities in the system, and what could trigger them?
The job is moving the delineators to the right because systems live and die in the mushy middle.
- Five Whys are one of the many gifts of Toyota’s manufacturing methodologies, a starting point for figuring out what went wrong retrospectively. ↩︎
- Gohan: I need an adult!
Goku: I am an adult!
Gohan: No. No, you are not. ↩︎ - Ope! Let me just scooch past you. ↩︎
- Which is how technical mythology grows up around your system. ↩︎
- If I never have to live through the Great Toilet Paper Run of COVID-19 again it will still be too soon. ↩︎
- That’s not shade you’re detecting, that’s a total eclipse of the sun. ↩︎
- Explored somewhat in Move Fast and Break Accountability. ↩︎
