A new engineer reviews the PFMEA and control plan, walks the line, and identifies a process step that makes no sense. It might be an in-process gauge adding twelve seconds to cycle time, or a redundant signature on a traveller. They ask the supervisor why it exists. Nobody knows. The quality engineer who wrote the control plan moved to another plant, and the original 8D report that triggered the control is buried in an unindexed SharePoint folder.
The team agrees the step is probably legacy overhead. They remove it. For two weeks, the productivity dashboard shows a beautiful green arrow. Cycle time drops. Then, a customer finds a defect that has not appeared in four years. The containment team scrambles, the root cause investigation traces the failure back to the exact step that was removed, and the organization realizes it just dismantled a critical safeguard.
This is Chesterton's Fence applied to manufacturing quality. The principle states that you should never remove a control until you understand why it was put there. In my experience auditing and transitioning systems at aerospace and automotive plants, this exact scenario plays out relentlessly. It is driven by a mechanical failure of documentation, and it is entirely preventable.
The mechanics of institutional knowledge erosion
Control plans list what the process does, but they rarely capture why it does it. IATF 16949 and AS9100 require documentation of current process states, not the historical context of failure modes. Work instructions describe procedures. However, the reasoning behind specific controls, such as the exact customer complaint or regulatory finding that demanded a hold point, lives in the heads of the engineers who designed it.
People are temporary. The engineers who designed the process move on to new roles or retire. The documentation they leave behind captures the mechanism but loses the intent. When institutional knowledge erodes, the visual appearance of the process becomes the only source of truth. A step that prevents a rare failure looks completely identical to a step that does nothing at all.
This erosion creates a cruel mathematical irony. When a quality control is perfectly effective, it drives the defect rate for that specific failure mode to zero. Because the problem never appears, teams begin to question whether the control was ever necessary. The control's own success becomes the primary argument for its elimination. Zero defects in the database means zero perceived risk.

The data invisibility paradox in lean transformations
This creates a genuine tension in lean manufacturing. Lean methodology actively encourages the elimination of waste. From a pure process perspective, an inspection step that never catches a defect looks exactly like non-value-added overhead. Value stream mapping and cycle-time reduction initiatives will inevitably flag these controls for elimination.
The paradox is that a critical safeguard and genuine waste look mathematically identical in your production data. If a control is perfectly effective, the defect rate for the failure mode it addresses is zero. If a control is completely unnecessary, the defect rate for that failure mode is also zero, because there was no failure mode to begin with. You cannot distinguish between the two by looking at outcomes alone.
The baseline data paradox
Cost pressure provides the motive to cut. Every process step has a quantifiable cost in cycle time, labour, and materials. When a manager proposes removing an inspection to save forty-seven seconds per unit, the savings are immediate and measurable. The risk is abstract. A potential field failure is a weak counter-argument against guaranteed throughput increases.
The disconnect between decision and consequence
The people making the decision to remove a step are rarely the people who will deal with the consequences. The engineer who deletes the inspection to streamline the process gets recognized for improving OEE. The operator who eventually catches the fallout is three organizational levels below them. The quality manager who handles the 8D containment works in a different department.
This structural disconnect makes unguided lean initiatives dangerous. The decision-maker reaps the immediate productivity gains, while the failure costs are externalized to the quality department and the customer. Without a governance mechanism that forces the decision-maker to own the downstream risk, blind optimization will always triumph over institutional memory.
I have seen a tier-one automotive supplier live through this exact scenario. A production manager noted that a visual inspection under angled light had not caught a defect in over two years. The surface treatment process had been improved, the defect rate was zero, and the fifteen-second inspection was consuming operator time. The manager removed the step to balance the line.
Within six weeks, three customer plants reported contamination failures in the fuel system. The root cause was a subtle residue that the surface treatment occasionally left on specific batches of raw material. The visual inspection was the only barrier between that residue and the customer's fuel rail. The defect had been so rare that most operators had never actually caught it, but the control was serving its purpose exactly as intended.
The real-world cost of blind optimization
The recall cost for that automotive supplier was in the seven figures. The customer-led audit that followed took three months of intensive preparation. The corrective action, which was to reinstate the exact same visual inspection, took one afternoon. The original reason for the inspection was documented in a quality alert from eleven years earlier. It was three paragraphs long, but nobody had read it.
A process step that prevents a problem and a step that does nothing can look identical in your production data.
This is the core failure mode of unmanaged process improvement. The organization pays for the safeguard twice: once when they originally install it after a painful failure, and again when they must rebuild it after a subsequent customer escape. The institutional trauma that justified the control is lost, and the organization is forced to relearn it at the customer's expense.
The solution is not to stop eliminating waste. Lean methodology is critical for remaining competitive. The solution is to establish a rigid protocol for understanding the intent of a process step before you remove it. You must investigate the function of the control before you treat it as overhead.
A practical framework for safe step removal
Organizations must streamline processes, but they must do it systematically. When a step is flagged as potential waste, the first action must be a purpose analysis, not a cycle-time calculation. This requires documenting the specific failure mode the control addresses, the customer requirement it satisfies, or the historical problem it prevents.
Protocol for evaluating legacy process controls
- 01Purpose analysisIdentify the exact failure mode, customer complaint, or audit finding that triggered the control.
- 02Knowledge auditInterview the longest-tenured operators and engineers to capture undocumented context.
- 03Pilot removalRemove the step on one line for one shift, with downstream inspection temporarily increased.
- 04Extended monitoringTrack the specific failure mode for ninety days, not two weeks, to account for rare batch variations.
- 05Risk-gated rolloutProceed with full removal only if the historical failure mode does not reappear under pilot conditions.
Check institutional knowledge before checking production data. Before you analyze defect rates and cycle times, talk to the operators and technicians who have been on the line the longest. They carry context that no database holds. They may know that an extra torque check exists because of a specific batch of bolts that arrived in 2017 with inconsistent hardness, and that knowledge is highly perishable.
If the analysis suggests a step can be safely removed, mandate a pilot removal with containment. Do not remove the step across all shifts and lines simultaneously. Remove it on one line, for one shift, with increased inspection downstream. Monitor the process for ninety days. Some failure modes tied to raw material batches are rare enough that they will not appear in a fortnight, but they will appear within a quarter.
Building a fence registry for PFMEA continuity
To permanently solve this problem, organizations must change how they document controls. Every control plan, PFMEA, and work instruction must include a brief explanation of why the control exists. This is not optional administrative overhead. Documenting the origin of a control is the difference between informed continuous improvement and reckless cost-cutting.
Make the cost of removal visible during management reviews. When someone proposes removing a step, the proposal must document the assumed risk alongside the expected savings. A proposal stating that removing an inspection saves forty-seven seconds per unit must also state that it eliminates the only detection point for residue contamination from the surface treatment process. This transparency forces engineering to make accountable decisions.
Control documentation standards
Standard documentation
- Defines the inspection method and frequency
- Lists the characteristic being measured
- Provides operator work instructions
- Disconnects the control from its history
Legacy-aware documentation
- Records the specific historical 8D or audit finding
- Identifies the exact failure mode being contained
- Links to the relevant raw material batch risks
- Mandates a purpose analysis before deletion
For critical process steps, maintain a living document, or a fence registry, that records the origin, purpose, and current relevance of each control. Review this registry annually during the management review process. Steps whose original failure mode is genuinely obsolete can be safely removed. Steps whose purpose remains relevant are preserved and reinforced.
Every process is a repository of lessons learned by your predecessors. Each control was installed because something failed badly enough to justify the cost of preventing it. The absence of an obvious purpose does not mean the absence of a purpose. The most expensive mistake a quality system can make is removing a safeguard because the organization forgot what it was protecting.
