An automotive components plant tracked a rising defect rate on an injection moulding line for three weeks. Overall scrap jumped from 1.8 percent to 4.2 percent. Three engineers proposed three different root causes: tooling wear, new operators, and bad raw material. They spent hours arguing in a conference room and left with an action plan that addressed all three theories simultaneously.

Three weeks later, the defect rate was 4.6 percent. The problem was not that the engineers lacked experience. The problem was that they were staring at a single aggregate number and trying to diagnose a process composed of multiple independent variables. They were arguing about the choir instead of listening to the individual voices.

Stratification is the practice of separating data into homogeneous subgroups so that patterns hidden within aggregate numbers become visible. It is one of the seven basic quality tools, yet it receives a fraction of the attention given to Pareto charts or fishbone diagrams. This is a fundamental error. Stratification is the prerequisite that makes every other analytical tool work.

Why Averages Destroy Information

An average is a summary, and summaries destroy information. If a factory reports an average equipment availability of 85 percent across three shifts, that number appears stable and acceptable. But if the morning shift runs at 95 percent, the afternoon at 85 percent, and the night shift at 75 percent, the aggregate number is actively misleading management.

In quality management, a defect rate of 4.2 percent is never a single story. It is several stories stacked on top of each other, waiting to be separated. When you treat a blended metric as a single process voice, you design countermeasures for a problem that does not actually exist in the aggregate. You correct for tooling wear, retrain operators, and audit suppliers, diluting your resources while the actual failure mode continues.

The truth lives in the layers of your operational data. Before you can apply statistical process control or design an effective 8D investigation, you must break the data down. If you do not stratify, you are flying blind while looking at a perfectly calibrated altimeter that is averaging the altitude of three different aircraft.

Where the calculation meets the floor: the gap between planned availability and the shift people actually work.
Where the calculation meets the floor: the gap between planned availability and the shift people actually work.

The Injection Moulding Case Solved

Returning to the automotive plant, a quality engineer eventually pulled the raw data from the problematic three-week period. Instead of looking at the 2,847 total defect records as a single block, she stratified the data by mould cavity number. The production tool had eight cavities operating simultaneously during each cycle.

When she charted the defect rate by individual cavity, the pattern became undeniable. Cavity 4 had a defect rate of 14.3 percent. The other seven cavities averaged 1.6 percent, which was actually better than the historical baseline. The aggregate number of 4.2 percent was a mathematical lie told by one failing cavity drowning out seven healthy ones.

The root cause was a hairline crack in the cooling channel of Cavity 4, which caused uneven solidification. This defect was invisible to the naked eye but glaringly obvious in the stratified data. The repair took two days. The overall defect rate dropped to 1.4 percent because the focused investigation also revealed minor cooling issues in Cavity 2 and Cavity 6.

The Multi-Layer Stratification Framework

Stratification requires analytical discipline. You cannot slice data randomly and hope for a revelation. You must start with the factors most likely to explain variation based on process knowledge. In most manufacturing environments, the highest-leverage starting points are machine or line, shift, and product variant. You must also ensure you collect the contextual data required to perform the analysis.

I have audited plants where the quality management system captured defect counts perfectly, but failed to log the machine ID, the operator, or the material batch. If you do not capture these attributes at the point of detection, you cannot stratify after the fact. You must design your digital and paper data collection forms with future analysis in mind.

Once you find a significant difference in one stratum, you must drill down iteratively. If the night shift shows higher defects, stratify the night shift data by machine. If Machine 3 is the problem on the night shift, stratify Machine 3 data by product variant. Keep drilling until the pattern is clear enough to act on.

Iterative Data Stratification Sequence

  1. 01Define the MetricSpecify exactly what you are measuring before touching the data.
  2. 02Capture Contextual DataEnsure records include machine, shift, operator, batch, and variant.
  3. 03Slice the Obvious SuspectsStart stratifying by time, equipment, and product family.
  4. 04Visualise Each StratumPlot box plots or control charts to look for differences in spread and mean.
  5. 05Drill Down IterativelyContinue stratifying within the anomalous subgroup until the root cause is isolated.
How to drill from a meaningless aggregate down to an actionable failure mode.

Simpson's Paradox and Supplier Evaluation

Stratification is your primary defence against a statistical trap known as Simpson's Paradox. This occurs when a trend appears in different groups of data but disappears or reverses entirely when the groups are combined. It happens frequently in supplier evaluations and process comparisons where product mixes are unevenly distributed.

Imagine evaluating two suppliers. Overall, Supplier A has a defect rate of 3.0 percent and Supplier B has 4.5 percent. The aggregate data crowns Supplier A the clear winner. But when you stratify by product complexity, a different truth emerges. Supplier B actually outperforms Supplier A on both simple and complex components.

The aggregate result favoured Supplier A only because Supplier A received mostly simple, low-tolerance products, while Supplier B received mostly complex, high-tolerance products. Without stratification, you would make the wrong sourcing decision, and it would look entirely data-driven. Stratification reveals the bias hidden in the mix.

Simpson's Paradox in Supplier Performance

3.0%Supplier A OverallAggregate rate appears superior.
4.5%Supplier B OverallAggregate rate appears worse.
1.5%Supplier B (Simple)Beats Supplier A's 2.0% on simple parts.
6.0%Supplier B (Complex)Beats Supplier A's 8.0% on complex parts.
Aggregate data crowns Supplier A, but stratified data reveals Supplier B is superior across all product types.

Integrating Stratification With Core Quality Tools

Stratification is rarely a standalone exercise. It is a meta-tool that enhances every other methodology in your quality arsenal, from IATF 16949 core tools to basic problem-solving frameworks. Applying it before you build your charts prevents misinterpretation and keeps your investigations targeted.

Consider the Pareto chart. If you build a Pareto of defect types without stratifying by machine, your top defect might actually be a localised issue. The top defect on the aggregate Pareto might not be the top defect on Machine 5, and Machine 5 might be where 60 percent of your total defects originate.

Stratification reveals correlation and pattern, but it does not confirm causation.

Control charts are equally vulnerable to unstratified data. A control chart that shows a process in a state of statistical control might be hiding two distinct subpopulations. One subpopulation runs near the upper control limit, the other near the lower limit. Averaged together, they produce a stable centre line that masks serious process instability. The same applies to histograms; a bimodal distribution is almost always a sign that two distinct processes are being blended into one dataset.

Stratification keeps these tools honest. When your scatter plot shows a weak correlation between process parameters and defects, stratify by a third variable like material supplier. The correlation might be strong within each stratum but masked when the groups are combined. Always separate the data before you trust the summary.

Operationalising the Habit on the Shop Floor

Modern manufacturing execution systems and quality management software generate enormous volumes of granular data. Every sensor reading, operator login, batch code, and timestamp is captured. This should be a golden age of stratification. Unfortunately, it rarely is. The bottleneck is not the technology; it is the analytical thinking.

Corporate dashboards are designed to show aggregates. Your daily quality dashboard shows overall scrap rate, overall OEE, and overall customer complaints. These numbers are useful for executive overviews but dangerous for problem-solving. The data is sliced and ready in your MES or Power BI, but teams fail because no one demands the breakdown.

You build a culture of stratification by changing leadership reflexes. When a production supervisor presents an aggregate scrap rate in a daily meeting, the plant manager must ask what the number looks like when broken down by machine and shift. When leaders ask the question consistently, the organisation learns to prepare the answer before being asked.

Keep a stratification log for every chronic quality issue. Document which factors have been tested, the sample sizes used, and what was found. When you try machine, shift, and product and find nothing, write it down. This prevents the next analyst from wasting time repeating dead-end analyses.

Dashboard Architecture vs Analytical Reality

What teams do

  • Review overall scrap and OEE percentages daily.
  • Brainstorm root causes based on a blended metric.
  • Launch broad countermeasures targeting all shifts.
  • Rely on single-layer data presentations.

What works

  • Demand immediate drill-down by line and operator.
  • Stratify the metric to isolate the failing subgroup first.
  • Focus 8D investigations on the specific anomalous segment.
  • Iterate through multiple layers until the pattern emerges.
Why executive summary screens fail the problem-solvers on the shop floor.

Common Execution Errors to Avoid

The most frequent error is stratifying with insufficient data. If you have 30 total defect records and you stratify by 10 machines, you have an average of three data points per machine. The pattern you see is statistical noise. As a practical rule, aim for at least 25 to 30 observations per subgroup to ensure the signal is meaningful.

Teams frequently confuse stratification with operational segmentation. Stratification is analytical; you are separating historical data to find hidden patterns. Segmentation is operational; you are dividing work into manageable chunks. They look similar but serve entirely different purposes in quality management.

Finally, never draw causal conclusions from stratification alone. If the night shift has a higher defect rate, stratification does not tell you why. The root cause could be lighting, supervision levels, or the fact that freshest material arrives in the morning. Stratification narrows the search area; root cause analysis confirms the finding.

Respect the complexity of your processes. Your production line is not a monolith. It is a collection of subprocesses, each with its own behaviour and its own story. When you insist on hearing them all at once through a single metric, you hear only noise. Stratification is the discipline of listening closely.