Someone asks whether a process can hit specification. An engineer collects thirty parts, crunches the numbers, and produces a Cpk of 1.33. The number goes into a PPAP package, a quality plan, and a management dashboard. Nobody asks whether that figure is still valid six months later, whether the data was normally distributed, or whether the gauge could even detect the variation it was measuring.

Process Capability Analysis is the statistical bridge between thinking a process works and proving it works consistently. It compares the voice of the process—its natural variation—against the voice of the customer, defined by specification limits. The mathematics are sound. The failure lies entirely in the assumptions practitioners ignore to get the study signed off and filed away.

I have audited plants where a validation Cpk of 1.67 was framed on the wall while the live process was generating scrap at 4 per cent. The snapshot taken during a controlled validation run gets treated as a permanent guarantee. Capability indices become compliance artifacts rather than indicators of process health, and the daily reality on the floor diverges further from that initial calculation every single day.

The Statistical Foundations You Must Verify

Cp measures potential capability, assuming the process is perfectly centred between specification limits. It calculates how many times the process spread fits within the tolerance width. Cpk measures actual capability, accounting for how far the process mean has drifted from the centre. A Cpk of 1.00 means the process spread exactly fills the specification with zero margin for error.

Pp and Ppk use overall variation, including between-subgroup drift, rather than just the within-subgroup variation used for Cpk. Comparing Cpk and Ppk reveals process stability. If Cpk is significantly higher than Ppk, your process is shifting between subgroups. A stable, predictable process will show minimal difference between the two indices.

A Cpk of 1.33 corresponds to roughly 63 defective parts per million. That is acceptable for many commercial applications but catastrophic for aerospace, medical devices, or automotive safety components. Conversely, demanding 1.67 on a non-critical aesthetic dimension wastes engineering resources. The threshold must be matched to the specific failure mode and customer consequence.

Standard Cpk Thresholds and Expected Yields

1.00Zero MarginProcess fills spec entirely (approx. 2,700 PPM).
1.33PPAP MinimumStandard automotive baseline (approx. 63 PPM).
1.67High RiskExpected for critical safety characteristics (approx. 0.6 PPM).
2.00Medical/AeroAggressive buffer for severe failure modes (approx. 0.002 PPM).
Capability targets must be dictated by the specific defect risk, not blindly adopted from a generic PPAP manual.

The Normality Assumption Nobody Tests

Cpk calculations are built on the assumption of normality. The entire statistical framework—the 3-sigma spread, the defect probability calculations, the indices themselves—depends on it. Most statistical software runs an Anderson-Darling or Kolmogorov-Smirnov test automatically. Engineers routinely glance at the flagged red p-value, ignore it, and proceed anyway because the submission deadline is looming.

When non-normality is ignored, the resulting indices are wrong. A right-skewed distribution, common in machining where tool wear systematically pushes dimensions in one direction, can overstate capability significantly. A bimodal distribution, caused by mixing two machines or two material lots into one study, renders the calculated Cpk entirely meaningless. The math produces a number, but that number describes no real-world process.

The correct response to non-normal data is to either transform it using a Box-Cox or Johnson transformation, fit an appropriate non-normal distribution, or fix the underlying process. Fixing the process is the most honest engineering response. But fixing the process takes time and root-cause analysis. Running the software and writing down the inflated number takes five minutes. Guess which approach dominates industry practice.

Quality decisions are made at the process, not in the report that describes it afterwards.
Quality decisions are made at the process, not in the report that describes it afterwards.

Measurement System Error and Index Engineering

A capability index reflects gauge variation as much as process variation. If the gauge has a low discrimination ratio—where the resolution is coarse relative to the process variation—you are measuring noise, not signal. The rule of thumb is that gauge resolution must be at least ten times finer than the total tolerance range. A tolerance of 0.1 mm requires a gauge that reads to 0.01 mm or better.

Despite this, capability studies are routinely launched without a prior Measurement System Analysis (MSA). The gauge is assumed adequate because a calibration certificate exists in a file. Calibration proves the gauge measures to a known standard; it does not prove the gauge can reliably detect the specific, minute variation of the manufacturing process you are trying to quantify. R&R must be verified before capability data is collected.

Beyond gauge errors, practitioners deliberately inflate Cpk values through subgroup gaming. Cpk uses within-subgroup variation. By sampling five consecutive parts every two hours from a slowly drifting process, the within-subgroup variation will be artificially minuscule. The resulting Cpk looks excellent. The Ppk, which accounts for the drift between those subgroups, tells the true, much worse story.

This is not capability analysis. It is index engineering: gaming the sampling strategy to produce a number that passes a threshold.

The Snapshot Problem and Time Decay

A capability study captures the process at one moment, under one set of conditions, with one batch of material. Tools wear, operators change, and ambient temperatures shift between seasons. A process that demonstrated a Cpk of 1.67 in a controlled validation run can easily drop below 1.00 on the production floor. Because the study was filed away in a binder, nobody knows the process is now producing defective parts.

The most honest capability statement includes a strict time component. Stating that the short-term Cpk was 1.42 based on 250 consecutive parts during a validation run on a specific date is true and useful. Stating that the process has a Cpk of 1.42 on a live dashboard is a lie of omission. It becomes more inaccurate with every passing shift, tool change, and material lot.

When a study comes back at 1.28, just below the 1.33 PPAP threshold, the engineering response is rarely to improve the process. The response is usually to collect more data, remove an unfavourable outlier, or recalculate using a different method. The threshold drives the behaviour. The behaviour is aimed at passing the gate, not at understanding the process or reducing actual variation.

The Correct Sequence for Capability Validation

  1. 011. Verify Measurement SystemConduct MSA (Gauge R&R). Confirm resolution is at least 10x finer than tolerance.
  2. 022. Establish Statistical ControlRun control charts. Identify and eliminate special causes before calculating anything.
  3. 033. Test for NormalityUse Anderson-Darling. If non-normal, transform data or fit the correct distribution.
  4. 044. Collect Sufficient DataEnsure sample size provides an acceptable confidence interval for the reported index.
  5. 055. Calculate and ContextualiseReport Cpk and Ppk together. Define the exact time window and operating conditions.
Calculating the index is the final step, not the first. Every preceding gate must be cleared or the output is invalid.

Continuous Capability Management

Real capability management is not a one-time study; it is a continuous operational practice. Before any index is computed, the process must be in statistical control. A capability index calculated on an unstable, special-cause-driven process is mathematically baseless. You are computing statistics on a moving target. Run the control chart, verify stability, and eliminate special causes first.

Distinguish explicitly between short-term and long-term capability. Report both Cpk and Ppk. If they differ significantly, investigate the drift immediately. The gap between those two numbers highlights exactly where your process is bleeding money and generating scrap. That drift is where your real continuous improvement opportunity lies, not in chasing a higher index for a PPAP submission.

Recalculate capability on a rolling basis. Set up automated systems that pull data weekly or monthly, depending on volume. Track capability on a trend chart, not a static dashboard. When capability degrades, trigger an investigation. Furthermore, report the confidence interval. A Cpk of 1.33 based on 30 parts has a massive interval; the true capability could sit anywhere from 1.05 to 1.61.

Sample Size Calculated Cpk 95% Confidence Interval Risk Interpretation
30 parts 1.33 1.05 to 1.61 Highly uncertain; actual capability may fail the 1.33 threshold.
100 parts 1.33 1.16 to 1.50 Moderate confidence; still carries marginal failure risk.
300+ parts 1.33 1.25 to 1.41 Narrow interval; reliable evidence of true process capability.
Sample size drastically alters the confidence interval. A target without statistical confidence is an assumption.

The Cultural Fix: Understanding Over Compliance

The deepest problem with process capability analysis is cultural. In many organizations, the study is treated as a compliance exercise, not an engineering effort. The customer requires a number. The PPAP package needs to be submitted. The IATF 16949 audit is approaching. When the sole purpose is to produce a number that satisfies a requirement, every incentive pushes the team toward manipulating the data to pass the gate.

Organizations that get capability analysis right ask a fundamentally different question. They do not ask what the Cpk is. They ask what they actually know about the process, how confident they are, and what it would take to be more confident. The index becomes one input among many, sitting alongside PFMEA findings, control charts, MSA results, and the operational knowledge of the operators running the line.

Quality systems must stop treating capability indices as pass/fail gates. If your only goal is hitting 1.33, you will find a way to calculate 1.33. If your goal is understanding and controlling variation, you will build a process that reliably holds 1.67 without statistical trickery. The first approach generates paperwork. The second approach generates quality, reduces scrap, and protects the customer.