Process capability analysis is the moment of truth in quality engineering. You collect SPC data, plot your control charts, and calculate Cp and Cpk to determine whether your process can meet specifications. It is the definitive statistical answer to whether you can manufacture a part consistently within tolerance.
In most plants I audit, the answer is fiction. Walk the floor and ask for the capability study on a critical characteristic. The spreadsheet will show a Cpk of 1.67 or 2.0. These are beautiful numbers. They are also complete fabrications, systematically disconnected from the reality of the production line.
The deception is rarely malicious. It stems from methodological errors that transform a rigorous statistical tool into performance art. By the time the index reaches a customer's PPAP report, it has been filtered through selective sampling, ignored distribution assumptions, and irrational subgrouping. The reported figure bears no relationship to actual production variability.
The Mechanics of Cp and Cpk
Cp measures potential capability. It is the ratio of the specification width to the process spread, defined as six standard deviations. A Cp of 1.33 means the specification window is wider than your variation, giving you breathing room. A Cp of 1.0 means your process spread exactly fills the tolerance. It assumes the process mean is perfectly centred between the limits.
Cpk accounts for reality: centering. You can have a tight process spread with a high Cp, but if the mean is shifted off-target, your Cpk drops significantly. Cpk is the lesser of two values, measuring capability against the upper and lower specification limits independently.
Together, these indices should provide a clear picture of process performance. One tells you about potential. The other tells you about reality. If your process is unstable, both numbers are meaningless. You cannot statistically predict the output of an unstable process.
Selective Sampling and Unvalidated Normality

The most common capability fraud is selective sampling. An operator collects thirty consecutive parts at the start of a shift, when the machine is fresh and the tooling is new. These golden parts represent the process at its absolute peak. The engineer calculates a Cpk of 2.1, the PPAP passes, and nobody records the reality that the process drifts to a Cpk of 0.8 by hour six.
The entire mathematical framework also rests on a normal distribution. This assumption is almost never validated in practice. Engineers dump data into Excel and run the formula. If they run a normality test and the p-value is below 0.05, they often delete outliers until the test passes.
Real manufacturing data is rarely normal. It is skewed by tool wear, bounded by physical limits, or bimodal because two machines feed the same study. When your data is non-normal and you apply standard formulas, your Cpk is directionally unpredictable. You might report 1.67 when the true capability is 1.1.
Reported Cpk vs Actual Production Reality
What teams do for audits
- Sample 30 parts from a fresh tool at shift start
- Ignore normality tests or delete outliers
- Report short-term Cpk as the production standard
- Open tolerances to inflate the final index
What statistical rigor demands
- Sample across multiple lots, shifts, and tool states
- Validate distribution; use Box-Cox if non-normal
- Report long-term performance alongside short-term
- Apply engineering tolerances based on functional fit
Subgrouping and the Specification Trap
Your subgrouping strategy determines the variation you see and the variation you hide. Short-term capability studies use small subgroups over a brief period. They capture only within-subgroup variation, ignoring the between-subgroup shifts caused by material lots, operator changes, and environmental conditions.
The result is a predictable gap. Your short-term Cpk is 2.0. Your long-term Cpk, which reflects what the customer actually receives over six months, is 1.0. This gap is the well-documented 1.5 sigma shift in Six Sigma literature. The exact magnitude is debatable, but its existence is not. Almost no one reports the long-term number.
Some organizations inflate numbers by manipulating specifications. The drawing specifies a tolerance of ±0.05 mm. The initial study yields a Cpk of 1.1. Engineering reviews the tolerance, opens it to ±0.10 mm, and the Cpk jumps to 2.2. The process has not improved by a single sigma; the goalposts simply moved to pass an IATF 16949 audit.
Conditions for a Defensible Capability Study
A genuine capability study requires four conditions. Omit any one and you produce a number that actively misleads your engineering team and your customer. The first prerequisite is representative data. You need a sample size of at least 100 data points, ideally 200, spanning multiple material lots, operators, and tool states.
Second, you must verify stability. Run the control chart first. If out-of-control points indicate special causes, investigate and eliminate them. You cannot compute a meaningful capability index for an unstable process. Third, validate normality using Anderson-Darling, Shapiro-Wilk, and visual histogram inspection.
If the data is non-normal, use a Box-Cox transformation, Johnson transformation, or a distribution-specific model like Weibull. Fourth, apply rational subgrouping. Report both short-term and long-term capability. Let the gap between them quantify how much your process drifts.
A verified Cpk of 1.15 is worth more than a fabricated Cpk of 2.5. The first drives improvement; the second guarantees customer escapes.
Validating Process Capability Before PPAP Submission
- 01Verify StabilityPlot SPC control charts and eliminate special causes before any calculation.
- 02Test NormalityRun Anderson-Darling and review histograms. Transform data if non-normal.
- 03Collect Representative DataSample across shifts, lots, and tool wear cycles. Minimum 100 data points.
- 04Calculate IndicesCompute both short-term and long-term Cpk. Report the gap between them.
The Audit-Driven Incentive Structure
The technical errors in capability analysis stem from a cultural root: organizations treat indices as report cards rather than diagnostic tools. When Cpk is a deliverable required for a VDA 6.3 audit, the incentive is clear. You need the number to be high. The index becomes an administrative hurdle, not a process insight.
When the magnitude of the number matters more than its accuracy, the analysis is optimized to produce the desired result. Sampling narrows. Tolerances widen. Short-term data represents the whole. The organization lies to itself because the audit system rewards compliance over truth.
In a functional quality culture, capability indices drive conversations. A Cpk of 1.1 is information, not a failure. It identifies where the process struggles and where to invest engineering resources. A verified Cpk of 1.15 provides a foundation for improvement. A fabricated Cpk of 2.5 provides a false security that eventually collapses during a customer field failure investigation.
Rebuilding Credibility with Honest Metrics
If your organization has been producing inflated numbers, the path back to honesty is uncomfortable. Start with your most critical characteristics, tied to safety, regulatory compliance, or AS9100 core performance. Re-run those studies from scratch using proper sampling plans and verified distributions.
Accept that the numbers will likely be worse than what you reported. Share the honest data with engineering. If the true Cpk is lower than previous reports, treat it as a signal for action. Identify the long-term variation sources eroding capability, launch 8D improvement projects, and track the index over time.
Process capability was designed as a living metric to drive continuous improvement. The mathematical framework is sound. The indices work. What fails is the organizational willingness to let the tool tell the truth. Fix the methodology, and the numbers will finally mean something again.
