The Spreadsheet That Ate
Engineering
If you’ve spent any time in manufacturing quality, you’ve seen it:
the FMEA that lives in a shared drive, opened once during APQP kickoff
and never again until the customer audit. Seventeen tabs. Four hundred
rows. Risk Priority Numbers calculated to two decimal places. Every cell
filled. Every failure mode “addressed.” And on the production floor, the
same defects the FMEA was supposed to prevent are happening every
Tuesday at a rate nobody tracks because tracking them would mean
admitting the FMEA was theater.
Nobody sets out to make FMEA theater. The intent is genuine. A
cross-functional team sits in a room, lists what could go wrong, scores
severity, occurrence, and detection on a 1-to-10 scale, multiplies them
together to get a Risk Priority Number, and prioritizes actions against
the highest RPNs. It’s logical. It’s structured. It’s exactly what the
manuals describe. And in the companies where FMEA actually works — where
it prevents problems rather than documenting them — the difference isn’t
the template. It’s what happens after the spreadsheet is filled out.
This article is about that difference. About how FMEA becomes a
compliance artifact instead of a prevention tool, and what it takes to
make it the latter.
What FMEA Actually Is
(When Done Right)
Failure Mode and Effects Analysis is a systematic, proactive method
for evaluating a process or product to identify where and how it might
fail, and to assess the relative impact of different failures. That last
part — “relative impact” — is what makes FMEA different from a simple
brainstorming list. It forces a team to prioritize.
There are three main types:
- Design FMEA (DFMEA): Focuses on product design —
what can fail in the way the product is engineered. Done during design,
before production tooling is committed. - Process FMEA (PFMEA): Focuses on the manufacturing
process — what can go wrong during fabrication, assembly, inspection,
handling, and shipping. Done during process development. - System FMEA (SFMEA): Examines the system level —
interactions between subsystems, interfaces, and external factors that
could cause system-level failures.
All three follow the same core structure: identify functions,
identify potential failure modes for each function, describe the effects
of each failure, assess severity, identify potential causes, assess
occurrence, identify current controls, assess detection, calculate RPN,
and prioritize improvement actions.
The output is a prioritized list of risks with assigned actions,
owners, and target dates. That’s the theory. The theory is sound. The
execution is where it collapses.
The Three
Numbers and the Religion Around Them
Every FMEA uses three scoring dimensions:
- Severity (S): How bad is it if this failure occurs?
A 10 means someone could die or the product violates a regulation. A 1
means the effect is so minor the customer wouldn’t notice. - Occurrence (O): How likely is the cause of this
failure to happen? A 10 means it’s inevitable without controls. A 1
means it’s remote — you’d never expect it. - Detection (D): How likely are your current controls
to catch the cause or failure mode before it reaches the customer? A 10
means you have no detection at all. A 1 means you’ll catch it every
time.
Multiply them: S × O × D = RPN (Risk Priority Number). Maximum is
1000. Minimum is 1.
And here’s where the religion starts. Some teams set a hard threshold
— “any RPN over 200 requires action” — as if 199 is fundamentally
different from 201. Others rank-order the RPNs and take action on the
top 20%. Both approaches miss the point. A failure mode with severity
10, occurrence 2, and detection 10 has an RPN of 200. A failure mode
with severity 4, occurrence 5, and detection 10 also has an RPN of 200.
The first one could kill someone. The second one is an inconvenience.
Treating them the same because the product is the same number is how
people lose faith in FMEA.
The AIAG-VDA harmonized handbook (2019) addressed some of these
issues by replacing RPN with Action Priority (AP) — High, Medium, Low —
based on the combination of S, O, and D rather than a raw product. But
in practice, many organizations still use the legacy 4th edition format,
and even those that have adopted AP often reduce it to the same
mechanical exercise: classify, fill in, move on.
The numbers are not the point. The thinking is the point. The numbers
are just a way to structure and compare the thinking. When the numbers
become the deliverable, the thinking stops.
How FMEA
Dies: Five Failure Modes of the Method Itself
1. The Lone Author
The FMEA is assigned to a quality engineer who fills it out at their
desk. They’re competent. They understand the process. But they don’t run
the machine, they don’t design the tooling, and they haven’t been on the
floor in three weeks. The failure modes they identify are the ones they
can imagine from a process flow diagram. The ones that actually happen —
the ones the operators know about, the ones the tooling designer
unconsciously designed around — never appear in the document.
Result: The FMEA is technically complete and
functionally useless. When a failure occurs that wasn’t in the document,
the quality engineer gets blamed for “missing” it. They didn’t miss it.
They were never going to find it sitting alone at a desk.
2. The Copy-Paste Cycle
New program launches. The FMEA from the previous, similar program is
opened. “Save As.” Names are changed. A few failure modes are added.
Some severity scores are adjusted. And the new FMEA is 85% identical to
the old one — including the failure modes that were irrelevant on the
old program and are equally irrelevant on the new one.
This is efficient. This is also how an FMEA accumulates scar tissue
from programs that had nothing to do with each other. The copy-paste
FMEA doesn’t reflect the new process. It reflects the history of all the
FMEAs that were copied from, stretching back to some original document
that was probably written correctly but has since been degraded by a
decade of template inheritance.
Result: The team spends time maintaining inherited
rows nobody understands while missing the new, process-specific risks
that a fresh FMEA would have caught.
3. The Post-Mortem FMEA
The FMEA is supposed to be proactive. In practice, it’s often written
after the process is already running, sometimes after the first customer
complaint. The team reverse-engineers the FMEA from what they already
know happened. The failure modes are real — they already occurred — but
the prevention logic is hollow. You can’t “prevent” something that’s
already happened. You can prevent recurrence, but that’s corrective
action, not FMEA.
Result: The FMEA documents history rather than
shaping the future. The same team that wrote it after the fact will
write the next one after the next fact.
4. The Action-Less FMEA
This is the most common failure mode. The FMEA is well-constructed.
The cross-functional team identified real failure modes with credible
causes. The RPNs are calculated. The top risks are flagged. And then…
nothing. The “Recommended Actions” column says things like “Monitor” or
“Operator training” or “TBD.” The “Action Taken” column is blank. The
“Revised RPN” column is blank.
The FMEA was the deliverable. The actions were the intention. But
nobody owned the follow-through. The document sits in the shared drive
with 47 rows of high-risk failure modes that have been known and
unaddressed since the program launched.
Result: The FMEA becomes a liability. During a
customer audit, someone opens it, sees a list of known risks with no
actions taken, and asks: “You knew about this and did nothing?” The FMEA
that was supposed to demonstrate quality discipline becomes evidence of
negligence.
5. The Frozen Document
The FMEA was written in 2023. It’s now 2026. The process has changed
three times: a new machine was installed, a material supplier changed,
and an inspection step was eliminated. The FMEA hasn’t been updated. It
still describes the 2023 process with the 2023 failure modes. When a new
failure mode emerges in 2026 — one that was introduced by the process
change — the FMEA offers no guidance because it doesn’t know the process
changed.
Result: The FMEA becomes a historical artifact. It’s
technically “on file” but functionally irrelevant to the process it’s
supposed to describe.
What Good FMEA Looks Like
The companies that get FMEA right share certain characteristics. None
of them are about the template.
It’s a conversation, not a document. The value of
FMEA is in the room when the team discusses what could go wrong and why.
The spreadsheet is the output of that conversation, not the substitute
for it. A two-hour session where a designer, a process engineer, a
quality engineer, and an operator argue about whether the occurrence
score should be a 3 or a 5 produces more insight than a perfectly
formatted spreadsheet completed by one person.
The actions are real. Not “monitor.” Not “training.”
Real actions: a design change that eliminates the failure mode entirely,
a poka-yoke device that makes the error impossible, an additional
control that improves detection from a 7 to a 2. If the recommended
action doesn’t change the S, O, or D score, it’s not an action — it’s a
hope.
The document is alive. When the process changes, the
FMEA changes. When a new failure mode is discovered (through a customer
complaint, an internal audit, or a near-miss), it’s added to the FMEA
with its causes and controls. The FMEA is reviewed at regular intervals
— not annually as a compliance ritual, but whenever the process changes
and at a minimum quarterly.
Severity is respected. A severity 10 is not a
scoring exercise. It’s a design imperative. If a failure mode has
severity 10 (safety or regulatory), the occurrence and detection scores
almost don’t matter. The action is to redesign so the severity is
reduced. FMEAs that have severity 10s with RPN-driven prioritization
that puts them below a severity 4 with high occurrence have lost the
plot.
The cross-functional team is real. Not a list of
names on a distribution list. People in the room. At minimum: the design
engineer, the manufacturing engineer, the quality engineer, and someone
who runs the process. For DFMEA, add a reliability engineer and a
service technician. For PFMEA, add an operator and a maintenance
technician. The people who know the failure modes are the people closest
to them.
The Link to Everything Else
FMEA doesn’t exist in isolation. It connects to other quality tools,
and those connections are where the value compounds:
-
FMEA feeds the Control Plan. The process
controls identified in the PFMEA — the detection controls — should
appear in the Control Plan as the inspection methods, frequencies, and
reaction plans. If your Control Plan doesn’t trace back to your PFMEA,
you have controls with no rationale and an FMEA with no
implementation. -
FMEA is informed by 8D. When a customer issue
triggers an 8D, the root cause and permanent corrective action should
flow back into the FMEA. The failure mode was either already identified
(and the controls were insufficient) or it was missed (and needs to be
added). Either way, the FMEA should be updated. -
FMEA connects to APQP. In the APQP timeline, the
DFMEA should be completed during product design, and the PFMEA during
process design — before tooling is built and before the Control Plan is
finalized. FMEA done after these milestones is documentation, not
prevention. -
FMEA supports ISO 9001 and IATF 16949. The
risk-based thinking required by ISO 9001:2015 Clause 6.1 and the
preventive action requirements throughout IATF 16949 are operationalized
through FMEA. But the auditor doesn’t want to see a spreadsheet. They
want to see that the thinking influenced decisions.
Practical Advice for
Getting Unstuck
If your FMEA program is broken (and most are), here’s how to start
fixing it:
-
Pick one FMEA — one real, current process — and redo it
properly. Don’t try to fix all of them at once. Pick a process
with known issues, assemble the right team, and build the FMEA the way
it’s supposed to be built. Use it as the reference for what “good” looks
like. -
Assign every recommended action to a named individual
with a date. Not a department. A person. Review the action list
at every program meeting. If an action is overdue, treat it with the
same urgency as a production stop. -
Establish a living document standard. The FMEA
is updated when: the process changes, a new failure mode is discovered,
a customer complaint reveals a gap, or at minimum every quarter. Version
control it. The 2023 FMEA is not the 2026 FMEA. -
Stop worshiping the RPN. Use it as a sorting
tool, not a decision rule. Discuss severity first. A severity 10 failure
mode is always high priority regardless of its RPN. Then look at the
combinations: high severity + high occurrence is critical. High
detection scores mean your controls are weak — fix them. -
Close the loop on actions. When an action is
taken, recalculate the RPN (or reclassify the Action Priority). Document
the “before” and “after” in the FMEA. This is how you demonstrate that
the FMEA drove improvement rather than just documenting risk. -
Audit your own FMEAs. Before the customer does.
Open the FMEA and ask: “Do we have failure modes with no actions? Do we
have actions that were never completed? Do we have RPNs that were never
revised?” If the answer to any of these is yes, you have a gap. Fix it
before it becomes an audit finding.
The Hard Truth
FMEA is one of the most powerful tools in quality engineering. It’s
also one of the most abused. The gap between a real FMEA and a
compliance FMEA is the gap between a company that prevents defects and a
company that documents them. The spreadsheet looks the same either way.
The factory floor tells the difference.
The next time you open an FMEA, ask one question: “Has this document
ever prevented a failure that would have otherwise occurred?” If the
answer is no — if it’s never driven a design change, never added a
control, never modified a process — then it’s not an FMEA. It’s a
spreadsheet. And the failure mode you should be analyzing is the one
where your risk analysis became a risk itself.
About the Author
Peter Stasko is a Quality Architect with over 25 years of experience
in manufacturing quality management, process improvement, and defect
prevention across automotive, electronics, and industrial sectors. He
has implemented and rescued FMEA programs in organizations ranging from
Tier 1 automotive suppliers to high-mix electronics manufacturers, and
he writes about what actually works on the shop floor — not what looks
good in a binder.