If you have spent any time in a manufacturing quality role, you have
seen the slide. It shows a beautiful bell curve sitting perfectly
between two specification limits, with a Cp of 1.67 and a Cpk of 1.33
printed underneath in bold green font. The supplier presented it during
the PPAP package. Your team nodded approvingly. The customer signed off.
Everyone felt good about the process.
Then production started, and the scrap rate hit six percent.
Process capability analysis was supposed to be the statistical proof
that your manufacturing process could consistently meet requirements.
Instead, it has become one of the most misunderstood and misused tools
in quality engineering. Not because the math is wrong — the calculations
are straightforward. But because the assumptions behind those
calculations get violated so routinely that the numbers on the report
bear almost no resemblance to what is actually happening on the shop
floor.
This article breaks down why process capability studies fail in
practice, what it takes to do them right, and how to stop using Cp and
Cpk as comfort blankets when they should be diagnostic instruments.
The Core Idea — and
Where It Goes Wrong
Process capability is simple in concept: compare the natural
variation of your process (what it actually does) to the specification
limits (what the customer requires). If your process spread is much
narrower than the specification width, you are capable. If it is not,
you are not. Cp measures potential capability — could your process fit
if it were perfectly centered? Cpk measures actual capability — does it
fit given where it is actually centered?
The problem is not the concept. It is everything that has to be true
before the numbers mean anything.
Assumption One: Statistical
Control
The fundamental requirement for capability analysis is that your
process must be in a state of statistical control. This means stable,
predictable, influenced only by common cause variation. Not drifting.
Not shifting. Not experiencing periodic special causes that come and
go.
Here is what happens in reality: an engineer collects 30 consecutive
parts from a machine, runs them through a CMM, pastes the data into a
spreadsheet template, and reports the Cpk. Nobody checked whether the
process was stable first. Nobody looked at a control chart. Nobody
verified that the 30 samples represented normal operating conditions
rather than a golden setup run specifically prepared for the capability
study.
If your process is not stable, capability indices are meaningless.
You are computing statistics on a moving target. It is like measuring
the temperature of water that is being heated and cooled at random
intervals, then declaring an average that predicts nothing about what
the next reading will be.
Assumption Two: Normal
Distribution
The standard capability formulas assume normality. Many manufacturing
processes are not normal — and not because something is wrong, but
because the physics of the process creates naturally skewed or bounded
distributions. A machining operation with a hard tooling stop produces
one-sided distributions. A filling process approaches normality only
when centered far from the fill floor. Geometric tolerances like
position, concentricity, and runout are inherently non-normal because
they are calculated as absolute values — you cannot have a negative
position deviation.
When you apply standard normal-based capability formulas to
non-normal data, you get numbers that look reassuring but are
mathematically invalid. A Cpk of 1.33 computed on heavily skewed data
might correspond to a true capability of 0.8 — meaning your actual
defect rate is ten times higher than what the index suggests.
Assumption Three:
Representative Sampling
The samples used for the capability study must represent the process
as it will actually run in production. This means normal raw material,
normal operators, normal tooling wear, normal cycle times, and normal
environmental conditions.
What typically happens: the quality team schedules the capability
study for Tuesday morning. The setup technician knows the study is
happening, so they spend extra time dialing in the machine. They use a
fresh tool. They let the process warm up longer than usual. They pick
the best operator. The 30 parts come off the line in a condition that
will never be replicated in actual production.
This is not fraud — it is human nature. But it makes the resulting
Cpk fiction. The number describes a setup condition, not a production
condition.
How
Capability Studies Actually Fail: Five Field Examples
Rather than abstract theory, let me share patterns I have seen
repeatedly across automotive, electronics, and medical device
manufacturing.
The Subgroup Size Trap
A supplier submits a Cpk of 2.0 on a critical dimension. Impressive —
until you ask how many parts were measured. The answer: 12. With only 12
samples, the confidence interval around that Cpk is enormous. The true
process capability could be anywhere from 1.2 to 3.5. The supplier is
not lying about the calculation. They simply collected so little data
that the number tells you almost nothing.
Rule of thumb: you need at least 100 individual measurements —
ideally 125 or more — before a capability index has enough statistical
power to be trusted for production decisions. With 30 samples, your Cpk
estimate has a margin of error of roughly plus or minus 0.3 at 95
percent confidence. That is the difference between “excellent” and
“marginal.”
The Cherry-Picked Timeframe
A medical device company ran capability studies quarterly. Each study
used data from the previous month. One quarter, the Cpk dropped from 1.4
to 0.9. The quality manager panicked — until an engineer discovered that
the data set included a two-day period where a worn fixture caused a
mean shift. Once those two days were removed, the Cpk jumped back to
1.4. The manager wanted to exclude the data. The engineer argued it
should stay because it reflected real process behavior.
The manager won the argument. The report showed 1.4. Six months
later, the FDA found the same fixture issue during an inspection and
issued a warning letter.
When you exclude unfavorable data from a capability study, you are
not improving your process. You are hiding a problem that will surface
later — usually at a much higher cost.
The Multi-Stream Blind Spot
An automotive supplier ran an injection molding machine with four
cavities. Their capability study combined measurements from all four
cavities into one dataset. The combined Cpk was 1.33 — acceptable. But a
cavity-to-cavity analysis revealed that cavity three was consistently
running at a Cpk of 0.7, while the other three cavities were above 1.6.
The overall number masked a serious problem because the good cavities
diluted the bad one.
This is the multi-stream problem, and it is pervasive. Any process
with multiple machines, multiple fixtures, multiple cavities, multiple
operators, or multiple shifts has multiple process streams. Combining
them into a single capability calculation averages away the variation
you most need to see.
The Short-Run Fallacy
Aerospace suppliers frequently deal with low-volume production —
sometimes fewer than 50 parts per batch. Traditional capability studies
requiring 100+ samples are impossible. So engineers use what they have,
compute a Cpk on 20 parts, and submit it.
There are legitimate statistical methods for short-run capability —
using standardized ranges, pooling data across similar part numbers, or
using Bayesian approaches that incorporate prior knowledge. But most
organizations do not use these methods. They simply run the standard
formula on too few parts and hope nobody asks questions.
The Drift Denial
A semiconductor fab monitored Cpk on wafer thickness. The number was
consistently above 1.67 — world-class. But the mean was slowly drifting
upward by 0.2 microns per month. Each monthly capability study showed
acceptable Cpk because the process was still within spec. No one noticed
the trend until the mean shifted past the upper control limit and the
yield crashed.
Cpk is a snapshot. It tells you about capability at one moment in
time. It does not tell you whether your process is stable over the long
term. A process can be capable today and incapable in three months if
tooling wears, material changes, or environmental conditions shift.
What a Real Capability
Study Looks Like
Doing process capability right is not complicated, but it requires
discipline. Here is the sequence that actually works.
Step 1: Verify Stability
First
Before computing any capability index, plot your data on a control
chart — an Individuals and Moving Range chart for one-at-a-time data, or
an X-bar and R chart for subgrouped data. Look for out-of-control
signals: points beyond three sigma, runs of seven or more on one side of
the centerline, trends, oscillations. If the chart shows instability,
stop. Fix the process first. Capability indices on an unstable process
are not just wrong — they are actively misleading.
Step 2: Check for Normality
Run a normality test — Anderson-Darling is my preference because it
is sensitive to deviations in the tails, which is where defects happen.
If the p-value is above 0.05, you can proceed with standard formulas. If
it is below 0.05, you have two options.
First, investigate whether the non-normality stems from a fixable
process issue. Sometimes non-normality is caused by measurement system
problems, operator differences, or mixed process streams. Resolving
these can restore normality.
Second, if the non-normality is inherent to the process physics, use
a non-normal capability analysis. This involves fitting an appropriate
distribution — Weibull, lognormal, largest extreme value, or a Box-Cox
transformation — and calculating capability indices based on percentiles
rather than standard deviations. Most statistical software (Minitab,
JMP, R) handles this natively.
Step 3: Separate Your Streams
If your process has multiple sources of variation — machines,
cavities, fixtures, shifts — analyze each stream independently. Report
the worst stream, not the average. If the worst stream is not capable,
the overall process is not capable. Averaging good and bad streams is
how scrap gets shipped.
Step 4: Collect Enough Data
Use at least 100 individual measurements collected over a
representative period — not a single setup run, not one shift, not the
first hour after maintenance. The data should capture normal variation
in material, operator, tooling wear, and environment. If you are
monitoring a high-volume process, collect data across at least one full
tooling life cycle.
For low-volume processes, use short-run methods or pool data from
similar part numbers after verifying that the processes are
statistically equivalent.
Step 5: Report Honestly
Include confidence intervals on every capability index. A Cpk of 1.3
plus or minus 0.4 is a very different story from a Cpk of 1.3 plus or
minus 0.05. The width of the confidence interval tells the reader how
much to trust the number.
Also report the expected defect rate — parts per million (PPM) —
alongside the index. Cpk translates directly to PPM for a stable, normal
process, and the PPM number is more intuitive for non-technical
stakeholders. “Your process will produce approximately 63 defective
parts per million” is more actionable than “your Cpk is 1.33.”
Common Misconceptions
Worth Correcting
“A high Cpk means the process is good.” Not
necessarily. It means the process is producing within specifications. If
the specifications are wrong — set too wide based on outdated
requirements, or set without regard for assembly interactions — a high
Cpk gives false confidence. I have seen processes with Cpk above 2.0
that still caused field failures because the spec limits did not reflect
the actual functional requirements.
“Cp and Cpk are the same thing.” Cp assumes the
process is centered. Cpk accounts for centering. A process with Cp of
2.0 and Cpk of 0.8 is producing within spec only because the tolerance
is wide — but the mean is so far off-center that any slight shift will
push parts out. This is a process living on borrowed time.
“Capability is a one-time calculation.” Capability
erodes. Tooling wears. Materials change. Operators rotate. Machines
drift. A capability study is valid for the period it represents — not
forever. Processes should be re-evaluated periodically, and the
frequency should depend on how much the process is capable of
changing.
“Six Sigma means Cpk of 2.0.” Six Sigma quality is
often described as a Cpk of 2.0, but this assumes a 1.5-sigma mean
shift. Without the shift, Six Sigma is actually a Cpk of 2.0 with Cp of
2.0. With the shift, it is a Cpk of 1.5. This distinction matters
because the 1.5-sigma shift assumption was a Motorola empirical
observation from the 1980s — it may or may not apply to your process.
Blindly assuming a 1.5-sigma shift can cause you to overestimate your
margin of safety.
The Connection to Real
Quality Outcomes
Process capability is not an academic exercise. It directly predicts
defect rates, customer satisfaction, warranty costs, and audit
readiness. When done correctly, it tells you whether your process can
meet customer expectations before you commit to production volumes. It
identifies which characteristics are at risk and which have margin. It
guides investment — telling you where to spend engineering effort for
maximum quality return.
When done poorly, it does something worse than tell you nothing. It
tells you everything is fine when it is not. It creates a false sense of
security that delays action until the problem becomes a crisis.
The most dangerous capability report is not the one that shows a low
Cpk. It is the one that shows a high Cpk computed on invalid data,
presented with confidence, accepted without question, and used to
justify shipping parts that will fail.
Practical Recommendations
If you are responsible for process capability in your organization,
here is what I would prioritize:
Audit your existing capability studies. Pick ten
recent reports at random. For each one, check: Was the process stable?
Was normality verified? Were enough samples used? Were process streams
separated? Were confidence intervals reported? If any answer is no, the
study needs to be redone or flagged as unreliable.
Train your team on assumptions, not just
calculations. Most engineers know how to compute Cpk. Far fewer
know the conditions under which that computation is valid. A two-hour
workshop on capability assumptions will prevent more quality problems
than a new statistical software package.
Make stability a prerequisite for reporting. No
control chart, no capability index. This single rule eliminates the
majority of misleading capability reports.
Align specifications with function. Work with design
engineering to ensure that specification limits reflect true functional
requirements, not inherited tolerances from previous drawings or
round-number defaults. Overly tight specs waste money; overly wide specs
ship defects.
Monitor capability over time. Do not treat the
capability study as a launch event. Track Cpk as a time series. Watch
for trends. Investigate drops before they become failures.
Closing Thought
Process capability analysis is one of the most powerful tools in
quality engineering — when its assumptions are respected and its results
are interpreted honestly. The math is not the hard part. The discipline
is.
Every time you sign off on a capability report, you are making a
promise to your customer: this process will produce conforming parts.
Make sure the number on the page actually supports that promise. Not the
number you wish were true. Not the number the timeline demands. The
number the data actually justifies — with its assumptions verified, its
limitations acknowledged, and its confidence intervals honest.
Anything less is not quality engineering. It is risk management
theater with statistical decorations.
Peter Stasko is a Quality Architect with over 25
years of experience in automotive, electronics, and industrial
manufacturing. He specializes in transforming statistical theory into
practical quality systems that actually work on the shop floor — not
just on slides. His work focuses on bridging the gap between what
quality tools are supposed to do and what they actually do when handed
to real production teams under real production pressure.