You know the scenario. The quality engineer walks into the Monday
review with a smile. “Gage R&R is under 10%,” she says, sliding a
report across the table. The plant manager nods. The auditors check the
box. Everyone feels good about the measurement system.
But here is what nobody asked: When was the last time anyone actually
looked at how that number was produced? Who selected the parts? Were
they hand-picked to span the range, or were they conveniently pulled
from the last production run? Were the operators chosen because they
were the best, or because they were available? Was the study done once,
years ago, and filed away — a artifact of a time when someone cared?
Measurement Systems Analysis is one of the most misunderstood tools
in quality management. Not because it is conceptually difficult — it is
not. It is misunderstood because organizations treat it as a paperwork
requirement rather than what it actually is: the foundation upon which
every other quality decision rests.
If you cannot trust your measurements, you cannot trust your data. If
you cannot trust your data, you cannot trust your control charts, your
capability studies, your inspection results, or your defect rates. And
if you cannot trust those, every decision you make — every acceptance,
every rejection, every process improvement — is built on sand.
What Measurement
Systems Analysis Actually Is
At its core, MSA is about answering one question: How much of
the variation you see in your measurements comes from the parts
themselves, and how much comes from the measurement
process?
When an inspector measures a part and gets 12.52 mm, that number is
not pure truth. It is a composite of three things:
- The true value of the part (what you actually want
to know) - The bias of the instrument (does it read
consistently high or low?) - The variation introduced by the act of measuring
(how much the number changes depending on who measures, when they
measure, and how they set up the measurement)
MSA decomposes that composite. It tells you, quantitatively, how much
noise your measurement system adds to the signal. Without that
knowledge, you are flying blind.
The most common MSA study is the Gage Repeatability and
Reproducibility (Gage R&R) study, but MSA encompasses more
than just R&R. A complete analysis examines five
characteristics:
- Bias: Does your gage read systematically different
from a known standard? - Linearity: Does the bias change across the
measurement range? - Stability: Does the bias change over time?
- Repeatability: Does the same operator get the same
result on the same part multiple times? - Reproducibility: Do different operators get the
same result on the same part?
Most organizations do R&R studies. Far fewer do bias, linearity,
and stability studies. That gap is where the trouble begins.
The Illusion of Precision
Here is a story I have seen play out in more factories than I can
count.
A precision machining operation is running tolerance bands of ±0.015
mm. The company invested in digital calipers that read to 0.001 mm.
Everyone feels confident because the instrument displays three decimal
places. The quality team runs SPC charts on the data. Capability indices
are calculated. Decisions about whether to accept or reject product are
made based on these numbers.
Then someone runs a proper Gage R&R study, and the results are
devastating.
The measurement system variation turns out to be nearly 40% of the
total tolerance. In practical terms, this means that for parts near the
specification limit, the measurement system alone could cause a good
part to be rejected or a bad part to be accepted. The digital display
with its three decimal places was an illusion — the instrument was
precise (it displayed fine increments) but the overall measurement
process was not accurate enough to distinguish good from bad product in
that tolerance range.
This is the most dangerous trap in measurement: precision of
display does not equal accuracy of measurement. A digital
readout showing 12.527 mm feels authoritative. But if the measurement
system has high R&R, that number could be 12.52 or 12.53 depending
on who held the caliper and how they positioned it. The third decimal
place is theater.
The Anatomy of a
Meaningful Gage R&R Study
A Gage R&R study that produces trustworthy results requires
careful design. Here is what a proper study looks like:
Part Selection
This is where most studies go wrong. The parts used in the study must
represent the actual production variation of the
process. If you select parts that are all nearly identical, the study
will show terrible R&R results — not because the measurement system
is bad, but because there is no part-to-part variation for the
measurement system to detect against. Conversely, if you deliberately
select parts spanning the full tolerance range (including some
deliberately out of specification), you inflate the part variation and
make the R&R look artificially good.
The correct approach is to sample parts from the actual production
process over a representative period. Take them from different shifts,
different times, different machine cycles. Let the process speak for
itself.
Operator Selection
Select at least two to three operators who routinely perform this
measurement. Do not grab your best inspector and your two most
experienced engineers. Do not run the study yourself because you are the
quality manager and you know what you are doing. Use the people who
actually do the work, because they are the ones whose variation matters
in production.
Number of Trials
Each operator should measure each part at least two to three times.
The trials should be randomized — not Operator A measures all parts,
then Operator B measures all parts. Randomize the order so operators
cannot remember their previous readings. If you hand the same part to
the same operator three times in a row, they will unconsciously try to
match their previous answer. That is not a test of the measurement
system; it is a test of memory.
Blind Measurements
The operator should not know which part they are measuring. Label the
parts on the side they cannot see during measurement. If operators know
they are measuring “Part 3, Trial 2,” they will adjust their technique —
sometimes consciously, usually not — to be consistent. Blind measurement
eliminates this bias.
Environment
Conduct the study in the actual production environment, not in the
climate-controlled quality lab. If the measurement is done on the shop
floor, the study must include shop floor conditions: temperature
variation, vibration, lighting, time pressure. A study done in ideal
conditions tells you what your measurement system could do in theory. A
study done in production conditions tells you what it actually does.
Interpreting the
Results: What the Numbers Mean
The AIAG (Automotive Industry Action Group) standard classifies Gage
R&R results into three bands:
-
Under 10% of tolerance: The measurement system
is acceptable. It contributes little variation relative to the process
tolerances, and you can trust the data it produces. -
10% to 30% of tolerance: The measurement system
is marginal. It may be acceptable depending on the application, the cost
of the measurement, and the risk of misclassification. This zone
requires judgment. If the measurement is for a critical-to-safety
characteristic, marginal is not good enough. If it is for a general
dimension with generous tolerance, marginal might be fine. -
Over 30% of tolerance: The measurement system is
unacceptable. Decisions based on this system are unreliable. You are
likely accepting bad product and rejecting good product on a regular
basis, and you do not know which decisions were wrong.
But here is what the guidelines do not tell you: the
percentage is only as honest as the study that produced it. A
6% R&R from a poorly designed study is worse than a 25% R&R from
a well-designed study. At least with the latter, you know you have a
problem.
The Type 1 Gage
Study: The Missing First Step
Before you ever run a Gage R&R study, you should run a Type 1
Gage Study. This is the simplest and most neglected form of MSA.
In a Type 1 study, one operator measures a single reference part —
ideally a calibrated standard with a known value — at least 25 times.
The results tell you:
- Bias: Is the average of the 25 measurements
different from the known reference value? - Repeatability: How much do the 25 measurements
vary? - The Cp of the measurement system: If you treated
the measurement process itself as a manufacturing process, how capable
is it?
If the Type 1 study shows significant bias or poor repeatability,
there is no point in running a full R&R study. You already know the
instrument has problems. Fix those first, then proceed.
The number of organizations that skip this step and go straight to
R&R is staggering. They run a complex crossed study with multiple
operators and parts, generate an impressive statistical report, and
never notice that the basic instrument is biased. It is like building a
second floor on a house without checking whether the foundation is
level.
Attribute
Measurement Systems: The Hidden Disaster
Most Gage R&R discussion focuses on variable data — measurements
that produce a number on a continuous scale. But many measurement
systems in manufacturing are attribute systems: pass/fail, go/no-go,
visual inspection.
Attribute measurement systems are consistently underestimated as
sources of error. The classic study involves 30 parts, two to three
inspectors, and two trials each. The parts should include borderline
cases — not obviously good and obviously bad parts, which any inspector
will classify correctly, but the difficult judgment calls that occur
near the specification limit.
The results are routinely shocking. It is not uncommon to find that
visual inspectors agree with themselves (repeatability) only 70% of the
time on borderline parts, and agree with each other (reproducibility)
only 60% of the time. This means that for parts near the specification
boundary, the accept/reject decision is essentially a coin flip.
When you consider that many visual inspection operations are the last
line of defense before product reaches the customer, this is terrifying.
And almost no one measures it.
Stability Over
Time: The Study Everyone Forgets
A Gage R&R study is a snapshot. It tells you about the
measurement system on the day of the study, with the operators who
participated, using the instrument as it was calibrated at that
time.
Measurement systems drift. Instruments wear. Operators come and go.
Environmental conditions change with the seasons. A measurement system
that was acceptable in January might be unacceptable by July, and nobody
will know because the study was filed away and forgotten.
The solution is simple and almost never implemented: run
control charts on a reference standard measured regularly by the same
operator. Once a day, once a week, once a month — whatever is
practical — measure a known reference part and plot the result on an
X-bar/R chart. When the chart shows a trend or an out-of-control point,
the measurement system has changed and needs investigation.
This is called a stability study, and it is the only
MSA component that provides ongoing assurance. Bias, linearity, and
R&R are point-in-time studies. Stability is the continuous monitor
that tells you whether your previous MSA results are still valid.
The Cost of Measurement
Error
Measurement error is not free. It imposes costs that are hidden but
real:
False rejects. When the measurement system adds
variation, parts that are actually within specification will sometimes
measure outside it. These parts are scrapped, reworked, or held for
review — all of which add cost. If your measurement system R&R is
25% of tolerance and your process is centered near the specification
limit, you could be rejecting several percent of perfectly good
product.
False accepts. The mirror image is worse. Parts that
are actually out of specification will sometimes measure within it.
These parts ship to the customer. The cost of this error shows up as
field returns, warranty claims, customer complaints, and damaged
reputation — but because the measurement system said the part was good,
nobody traces the problem back to measurement error.
Over-control. When operators see variation that
comes from the measurement system rather than the process, they adjust
the process to compensate. This adjustment adds real variation to a
process that was actually stable. The measurement system creates process
instability that would not exist if the measurements were trusted less
and the process were left alone.
Inflated capability indices. Measurement variation
adds to the total variation used in capability calculations. If your
measurement system contributes 20% of total variation, your Cpk is
understated. You may believe your process is incapable when it is
actually the measurement system that needs improvement. This leads to
unnecessary process changes, equipment purchases, and engineering effort
aimed at the wrong target.
Practical
Steps to Improve Your Measurement Systems
If this article has made you uncomfortable about your measurement
data, good. Here is what to do about it.
Start with Type 1 studies on your critical
instruments. Before investing time in full R&R studies,
check bias and repeatability on the instruments that matter most. This
is quick, cheap, and will immediately identify the worst offenders.
Prioritize by risk. You cannot fix every measurement
system at once. Focus on the measurements that drive the most important
decisions: final inspection on safety-critical characteristics,
measurements used for SPC on processes with tight tolerances, and
measurements used to calculate capability indices that drive customer or
regulatory reporting.
Involve operators in the improvement process.
Operators often know exactly what is wrong with a measurement system —
the fixture does not hold the part consistently, the instrument drifts
between calibrations, the lighting makes the visual standard hard to
compare. Ask them. They will tell you things your MSA software never
could.
Fix the system, not the operator. When R&R
studies show high reproducibility variation (operator-to-operator
differences), the instinct is often to retrain the operators. Sometimes
this helps. But more often, the root cause is a measurement procedure
that is ambiguous, a fixture that allows multiple interpretations of
correct positioning, or an instrument whose operation requires more
skill than is reasonable. Design the measurement so that any competent
operator gets the same answer. That is poka-yoke applied to
measurement.
Make stability studies part of your calibration
system. Most calibration systems verify the instrument against
a standard at defined intervals. But calibration intervals are often 12
months, and a lot can go wrong in 12 months. Stability studies using
control charts provide early warning between calibrations.
Stop treating MSA as a one-time event. The most
fundamental change is cultural. MSA is not a milestone to be completed
before a production launch and then forgotten. It is an ongoing
discipline, like calibration itself. Measurement systems are living
processes that change over time, and they need to be monitored just as
you monitor production processes.
The Measurement System
Nobody Audits
Here is a final thought that should keep quality managers awake.
ISO 9001 requires that organizations determine and provide the
resources needed to ensure valid and reliable results when monitoring
and measuring are used to demonstrate conformity to requirements. It
requires that measuring equipment be calibrated and verified. Auditors
check calibration records meticulously.
But how many auditors ask to see the MSA study? How many ask whether
the measurement system variation is understood and acceptable relative
to the tolerance? How many ask about stability, bias, or the last time
someone ran a Type 1 study on the instrument making safety-critical
measurements?
The standard says “valid and reliable results.” Calibration alone
does not guarantee valid and reliable results. Calibration tells you the
instrument was correct on the day it was checked against a standard. It
says nothing about repeatability, reproducibility, bias across the
range, or stability over time. A calibrated instrument in a poor
measurement system produces calibrated garbage.
The gap between what the standard intends and what audits verify is
where measurement systems hide their failures. The organizations that
close this gap — that treat measurement as a process to be studied and
improved, not just a tool to be calibrated — are the ones whose data can
be trusted.
Everyone else is making decisions based on numbers they have never
validated. And in quality, as in everything else, you cannot improve
what you cannot measure — and you cannot trust what you have never
analyzed.
Peter Stasko is a Quality Architect with over 25
years of experience in manufacturing quality management, from shop-floor
inspection to executive-level quality strategy. He has implemented MSA
programs across automotive, aerospace, and general manufacturing
environments, and has seen measurement system errors cost companies
millions while remaining completely invisible on their dashboards. He
writes about the realities of quality management — the gap between what
the textbooks describe and what actually happens on the production
floor.