The Simulation Gap Usually Means Your Data Is Lying
A Midwest lunchmeat producer wanted to know whether to spend on new equipment or cut a shift, and nobody in the room could answer without a fight.
The line the model could not match
A Midwest lunchmeat producer wanted to know whether to spend on new equipment or cut a shift, and nobody in the room could answer without a fight. So we built a digital twin off a full year of their production data and ran thousands of scenario variations against it. The model landed at 98.2% accuracy to actual plant output. Every line matched within about two points, except one. Call it Line P. It sat stuck at 88 to 89% no matter how the model ran.
The instinct in that moment is to assume the model is broken and start tuning it until the gap closes. That instinct is expensive and, here, exactly wrong. The gap on Line P was not a modeling failure. It was the model catching something the plant had lived with for weeks.
A calibrated model turns disagreement into information
Line P had seven weeks of OEE readings above 100%. That is physically impossible; a line cannot run more than 100% of the time it is available. It was a structural data error, and the plant's own team recognized it the moment we called it out. They had fought an intermittent sensor issue on that line for several weeks and moved on. The model did not move on. It refused to reconcile good scenario math with impossible inputs, and the refusal showed up as an 88% line sitting next to a wall of 98% lines.
That is the whole mechanism. A forecast is a guess about the future, and its errors are noise. An instrument is a calibrated measure, and its errors mean something. Once a model matches reality within a couple of points across the floor, you have crossed from the first thing to the second. Now every disagreement is diagnostic. The lines that match tell you the model is trustworthy. The line that does not match is not the model failing; it is the model pointing at the one place your data is lying to you.
Manufacturers throw away this signal constantly. A plant will normalize away the ugly weeks, average across the bad sensor, and present a capacity number that looks clean. The clean number is the dangerous one, because the mess it hid is still on the floor. The gap you are tempted to erase is usually the most honest reading in the dataset.
Model before you argue
The practical version of this is a sequence, and any plant can run it.
Pull a year, not a month. A single month hides seasonality and variation, and variation is where the real capacity question lives. When this producer sent thirty days first, then a full year, the year is what made the model worth trusting. Plants are not static; the value is in capturing how they actually swing.
Calibrate per line, not in aggregate. A plant-wide 98% average would have buried Line P entirely. Matching each line separately is what surfaced the one that could not be matched. If you only check the total, you never find the sensor that is feeding you fiction.
Then, and only then, run the scenarios. Here they ranged from incremental (keep the four existing lines, add auto-loaders that replace manual loading, tighten the schedule) to consolidation (shut two older lines, run two newer ones with auto-loaders across two shifts) to growth. The consolidation case modeled roughly $1.4M in labor savings and about a 14% throughput gain, sitting on top of roughly $1.4M in auto-loader equipment. Nobody argued about whether it would work, because the model had already been proven accurate against the floor it was describing. The debate shifted from "will this pay off" to "which phase do we sequence first."
The model also named the ceiling. Growth was not open-ended; the cook system becomes the binding constraint at about 18M pounds a year, roughly three times current output. That is a specific system, not a vague "the plant is getting full." Knowing the next constraint by name is what keeps a company from buying capacity in front of a wall it has not identified.
What a modeled decision looks like
The model tracks actual output within two points on every line, and any line it cannot match is flagged as a data fault inside a week, not averaged away. Every capital ask carries a scenario number rather than a vendor promise. The next capacity constraint is named as a specific system with a specific throughput, so nobody spends to unlock a line that a downstream process will immediately re-choke. And the decision to add a shift or cut a line is made against a calibrated baseline, not a hunch about how the floor "usually" runs.
The closing
The producer spent months believing Line P had a performance problem. It had a data problem, and the model was the first thing in the building to say so out loud. The gap you want to close is often the gap that is teaching you something: the one line that broke the model was the only line telling the truth about its own numbers.