Process Failure Mode and Effects Analysis is supposed to be the engineering backbone of preventive quality. In theory, every manufacturing process gets dissected: failure modes imagined, severity scored, occurrence ranked, detection evaluated, and a Risk Priority Number calculated that drives action on the most dangerous gaps. The reality in most plants looks nothing like the theory.

Walk into any automotive supplier and ask to see their PFMEA. You will receive a binder or a spreadsheet, often hundreds of rows long, filled with severity-occurrence-detection scores assigned during launch and untouched since. The actions column reads "operator awareness" and "improved fixture design" with no traceability to implementation. Detection scores default to 4 or 5 because nobody wants to admit their inspection system cannot catch the defect. When a new failure mode surfaces in production — one absent from the document — the PFMEA gets updated after the fact, a quiet post-mortem that prevents nothing.

The dysfunction is structural, not cultural. Engineers do not produce bad PFMEAs because they lack competence. They produce them because the system treats the document as a deliverable to pass a gate rather than a tool to manage risk. Fixing this requires rebuilding the mechanisms around the document — how it is triggered, scored, linked, and sustained. Below is that rebuild, drawn from plants I have audited and systems I have implemented under IATF 16949 and AS9100.

The Root Cause: PFMEA Positioned as a Deliverable

The dysfunction starts with how PFMEA is positioned in the APQP timeline. It sits between the process flow diagram and the control plan — a mandatory deliverable that the customer or the standard requires. Project engineers complete it because the phase gate demands it, not because they believe it will prevent defects. Once the gate is passed, the document is archived. The deadline was the trigger, and the deadline has passed.

Contrast this with how PFMEA functions in organisations that extract real value from it. There, the document is opened when a process change is proposed, when a new defect pattern appears, when a machine is relocated, when a material supplier changes, and during periodic engineering reviews. It is a working reference — a shared mental model of where the process can break and what is being done about it. The document is never finished because the process is never static.

The difference is not effort. Engineers in both organisations spend comparable hours on PFMEA. The difference is timing and trigger structure. Compliance-driven PFMEA happens once, at launch, driven by a deadline. Prevention-driven PFMEA happens continuously, driven by defined events. If your engineering change management system does not automatically flag PFMEA for review when a process or component changes, the document is already decaying.

Two PFMEA Operating Modes

Compliance-driven PFMEA

  • Triggered by APQP phase gate deadlines
  • Opened once at launch, then archived
  • Scores fixed at historical assumptions
  • Actions tracked in a separate, disconnected system

Prevention-driven PFMEA

  • Triggered by process, supplier, and defect events
  • Opened whenever the process changes
  • Scores updated as production data accumulates
  • Actions closed inside the same tracking loop
The same document, two completely different engineering tools — separated only by trigger structure.

Severity Scores Disconnected from Customer Impact

Severity is supposed to represent the impact of the failure mode on the end customer or the next operation. In practice, engineering teams assign severity based on internal inconvenience rather than external impact. A burr on a machined surface that causes assembly difficulty downstream gets a severity of 7 because it annoys the assembly team. A cosmetic defect that a customer will never notice also gets a 7 because the quality manager is aesthetically conservative. The scoring reflects organisational politics, not engineering analysis.

The fix is to build a severity evaluation table that maps specific failure consequences to specific customer effects — functional failure, safety risk, regulatory non-compliance, warranty claim, aesthetic deviation — and to score against the worst credible customer outcome, not the most likely one. AIAG-VDA aligned severity scales do this reasonably well, but only if the evaluation criteria are debated and agreed across functions, not assigned by one engineer in a spreadsheet. If your severity table has not been reviewed by a cross-functional team in the past two years, it is drifting.

Severity is the score that should change least over time, because the physics of the failure does not change. What changes is your understanding of the consequence. When a field return or a complaint reveals that a failure mode you scored at 4 actually causes a safety-critical condition, that severity must go to 9 or 10 immediately. I have seen plants where severity scores remained unchanged for years after field data proved them wrong, simply because nobody owned the update.

Occurrence Scores Based on Capability, Not Promises

Quality decisions are made at the process, not in the report that describes it afterwards.
Quality decisions are made at the process, not in the report that describes it afterwards.

Occurrence ranking should reflect the likelihood that the failure mode will happen given the current process controls. What actually happens is that engineers score occurrence based on what they believe the process is capable of, not what the data shows. A new process with no production history gets an occurrence score of 2 because the machine supplier promised capable output. Six months later, when the defect is appearing at three per cent, nobody goes back to revise the score. The gap between the assumed process and the real process widens every month.

The rebuild is straightforward: occurrence must be scored from historical data for similar processes, from capability studies on the current process, or from a structured engineering assessment when no data exists yet. When process capability data becomes available during production — through SPC, Cpk studies, or yield tracking — the occurrence scores must be updated to reflect it. A process running at Cpk 0.8 does not deserve an occurrence score of 2, regardless of what the machine supplier's brochure claimed.

If your PFMEA still shows occurrence scores from launch after a year of production data is available, the document is dead. You are managing a fiction. The production engineering team should review occurrence scores quarterly for the first year of a new process, and at least annually thereafter. This is not bureaucratic overhead — it is the mechanism that keeps the PFMEA honest.

Detection Scores and the Inspection Fallacy

Detection ranking evaluates whether current controls will catch the failure mode before the product reaches the customer. The most common corruption of this score is the assumption that because an inspection step exists, detection is good. But inspection systems have blind spots, gauges go out of calibration, operators get fatigued, and automated vision systems are tuned to specific defect signatures. The existence of a control plan line does not equal detection capability.

Every detection score should be validated against two questions: has this control actually demonstrated the ability to detect this specific failure mode — through measurement systems analysis, gauge R&R, or detection capability testing? And what happens when the control fails — is there a backup? If neither question can be answered with evidence, the detection score should be high, regardless of how many inspection steps exist on the control plan.

The detection score is where most PFMEAs are most dishonest. I have audited plants where every detection score was 4 or below, yet the plant had no MSA data for half of the gauges listed on the control plan. If you have not validated that a gauge can statistically detect the tolerance range it is guarding, your detection score is a guess. Guesses belong in brainstorming, not in a risk document that drives engineering decisions.

RPN Arithmetic and Action Prioritisation

The traditional RPN — severity multiplied by occurrence multiplied by detection — produces a number between 1 and 1000. Teams set an arbitrary threshold, often 100, and commit to actions on anything above it. This arithmetic has a well-known flaw: a high severity, low occurrence, low detection risk scoring 10-2-2 produces an RPN of 40 and gets ignored, while a medium-everything risk scoring 5-5-5 produces 125 and triggers action. The mathematics systematically masks safety-critical failure modes.

The AIAG-VDA harmonised approach replaced RPN with Action Priority, which uses a lookup matrix that weights severity more heavily and avoids the false precision of multiplication. If your organisation is still using raw RPN cut-offs, you are systematically under-prioritising safety-critical, low-frequency failure modes — exactly the ones that cause recalls, field failures, and regulatory action when they eventually occur. Transition to AP scoring is not optional for AIAG-aligned systems; it is the current standard.

If your action items are marked complete but the risk scores have not been updated, the actions either had no effect or the scoring was wrong.

The deeper problem is not the arithmetic but the false certainty it creates. A three-digit number implies a precision that the underlying scoring cannot support. Severity, occurrence, and detection are estimates based on judgement and data — sometimes good data, sometimes thin data. Treating their product as a precise ranking mechanism gives leadership a false sense that the highest risks have been addressed. In reality, the highest risks may be sitting just below the threshold, invisible to the arithmetic.

Linking PFMEA to the Control Plan and Closing the Loop

The PFMEA identifies risks. The control plan defines how those risks are managed in production. In well-functioning quality systems, every significant failure mode in the PFMEA maps to a specific control in the control plan — a reaction plan, a frequency, a method, a responsibility. In most real systems, the two documents were written by different people at different times and bear no structural relationship to each other. The risk analysis exists in one document; the process control lives in another.

This disconnect means that when a PFMEA identifies a critical failure mode with marginal detection capability, the control plan does not increase inspection frequency, does not add a contingency reaction plan, and does not escalate the characteristic to a significant or critical characteristic with additional controls. The two documents should be linked by a shared characteristic matrix, so that any change in PFMEA scoring propagates automatically to the control plan review queue.

Closed-loop action management is the mechanism that keeps both documents alive. PFMEA actions must flow into the same action tracking system used for 8D corrective actions, audit findings, and nonconformance reports. Completed actions must trigger PFMEA score revisions — and when a new failure mode appears in production that was not in the PFMEA, that gap must be investigated through root cause analysis, not simply appended as a new row.

PFMEA Living-Document Cycle

  1. 01Trigger eventProcess change, supplier change, defect pattern, complaint, or scheduled review
  2. 02Cross-functional reviewManufacturing, quality, maintenance, and operations walk the process flow together
  3. 03Data-driven scoringOccurrence from Cpk and SPC data, detection from MSA results, severity from customer impact
  4. 04Action and control plan updateActions tracked to closure, risk scores revised, control plan linked and updated
The four-phase loop that separates a prevention tool from a compliance binder. Every production event should re-enter the scoring phase.

Sustaining the Discipline Through Visibility

Documents do not maintain themselves. People maintain them, but only when the system demands it. The most effective sustaining mechanism is simple: make PFMEA currency visible to leadership. Include a PFMEA health metric in the monthly quality review — percentage of PFMEAs reviewed on schedule, percentage of actions closed on time, percentage of scores backed by current data. These are leading indicators of prevention system health.

When leadership asks about PFMEA health the same way they ask about scrap rates and on-time delivery, the document stops being a binder on a shelf. The behaviour change is immediate. Engineers start scheduling reviews before they are overdue. Scoring gets defensible because it will be questioned. Action items get owners and dates because open items appear on the dashboard. Visibility creates accountability, and accountability creates discipline.

A functional PFMEA is one of the highest-leverage tools in the prevention cost category. Every failure mode anticipated and engineered out before production avoids internal failure costs — scrap, rework, downtime — and external failure costs — warranty, recalls, reputation damage. The organisations that understand this do not view PFMEA as a compliance burden. They view it as the cheapest form of quality: thinking about problems before they happen, on paper, rather than fixing them in production or in the field.

The rebuild is not complex, but it requires consistency. Define the triggers. Score from data. Link the documents. Close the loop. Make it visible. Do these five things and the PFMEA becomes what it was always meant to be: the engineering intelligence of the process, not a compliance artifact in a binder.