Every quality manager eventually collects a story about two certified inspectors, same part, same drawing, same lighting booth, and opposite verdicts. I have sat through customer audits where exactly this happened: an operator rejected a surface finish sample, her supervisor passed it, and the auditor asked the only question that mattered. How do you know your inspection system works? Certificates prove someone attended training. They do not prove the judgement itself is stable, and visual judgement is where stability fails first.

An attribute agreement analysis answers a narrower question than most people expect. It does not ask whether inspectors match the truth. It asks whether they agree with themselves across repeated trials, whether they agree with each other, and whether that shared judgement aligns with a reference standard. Those are three separate failures, and conflating them is why so many studies produce confusing results. An inspector can be perfectly consistent and consistently wrong. Another can match the reference every time while a colleague wobbles around the boundary.

The method is straightforward in principle. Select thirty to fifty parts or images spanning the judgement space, including deliberately difficult boundary cases. Each appraiser evaluates every sample blind, in randomised order, at least twice — ideally three times — with no way to recall the earlier verdict. Then tally within-appraiser agreement, between-appraiser agreement, and agreement with the accepted reference. The arithmetic turns those tallies into kappa.

What Kappa Actually Tells You

Raw percent agreement flatters everybody. If ninety-five percent of parts are good and an inspector calls everything pass, he still scores ninety-five percent. Kappa corrects for that agreement you would have obtained by chance alone. It compares observed agreement against expected agreement given the marginal pass/fail rates, and expresses the excess as a proportion of the maximum possible excess. A kappa of zero means the inspectors agree no more than coin-flipping would predict; a kappa of one means perfect agreement beyond chance.

Interpreting the number requires honesty about context. Convention treats kappa above roughly 0.9 as excellent, 0.8 to 0.9 as good, and anything below 0.7 as evidence the system needs work. But kappa is sensitive to prevalence: when nearly all samples fall on one side of the boundary, even genuine agreement produces a depressed kappa, because chance agreement is already high. This is why the study sample must not mirror production proportions. You deliberately load the set with boundary cases — perhaps a third of them — because that is where the measurement system either holds or breaks.

Fleiss' kappa handles more than two appraisers; Cohen's kappa covers pairs. For pass/fail judgements you also want the confusion pattern itself: how many times a truly bad part was passed versus a good part rejected. Two systems can share a kappa while having completely different risk profiles. In my experience the false-accept cell is the one the customer's auditor will find, so report it explicitly rather than hiding it inside a summary statistic.

Reading a kappa result

0.9+ExcellentSystem suitable for critical visual characteristics
0.8GoodAcceptable with monitored boundary samples
<0.7ActionFix criteria, conditions or appraisers before relying on results
~1/3BoundaryShare of study samples placed at or near the limit
Conventional interpretation bands for kappa, plus the prevalence trap that depresses even honest agreement when the sample skews heavily to one side.

Repeatability: Can an Inspector Agree With Herself?

Within-appraiser agreement is the first gate, and I always examine it before anything else. If an inspector judges the same scratch differently on Tuesday than on Monday, the problem is not the standard, the training, or the colleague — it is the judgement itself, and no amount of cross-checking will rescue it. Intra-appraiser disagreement above a handful of trials points to three usual culprits: ambiguous acceptance criteria, inconsistent viewing conditions, or fatigue and attention drift during long inspection runs.

Viewing conditions deserve more suspicion than most plants give them. A visual defect judged under 500 lux at the bench looks different under 1,100 lux in the inspection booth, and different again at the end-of-line station where the ceiling fittings have yellowed. Illuminance, viewing angle, viewing distance, and the background colour behind the part all belong in the work instruction — specified numerically, verified periodically. When I audit a plant and find three inspection stations with three lighting conditions, I already know what the agreement study will show before running it.

Hydrogen Embrittlement Control in High-Strength Fastener Plating
Visual judgement is only as repeatable as the conditions it is made under — the same defect changes verdict with the light it is read in.

Randomisation and blinding discipline matters enormously here. Samples must be re-presented in a different order each round, shuffled by someone other than the appraisers, with any identifying marks masked. Parts get rotated or re-oriented between trials so position does not act as a memory cue. Rounds should be separated by enough time — ideally a shift boundary — that nobody can rely on recall. I have seen studies invalidated because an inspector, quite reasonably, remembered this is the one with the dent and simply repeated herself. You are measuring the judgement, not her short-term memory.

Boundary Samples: Where Agreement Goes to Die

The middle of the distribution is easy. A part with a fifteen-millimetre scratch against a two-millimetre limit passes the test of agreement trivially, and a part split open fails it equally trivially. Agreement studies pass or fail on the region just either side of the specification, and a study built from easy samples is theatre. You need parts whose defect sits at, just inside, and just outside the limit — so close that reasonable, trained, competent people must scrutinise them carefully before deciding.

Producing those samples takes deliberate effort. Some plants machine artificial defects to controlled dimensions; others quarantine borderline parts found in production and have a designated authority — usually a customer-approved reference panel or the resident material review engineer — assign the official disposition. Both approaches work. What does not work is grabbing forty consecutive good parts from a bin and adding two known rejects. Such a study tells you nothing about the judgement boundary and produces a reassuring kappa built on trivia.

Boundary samples then become physical standards, and standards need governance. Each one carries an identifier, an official classification, a date, and an authorising signature. They live in controlled storage, get inspected for degradation — corrosion, fading, handling damage — on a defined schedule, and get re-verified when the drawing, the customer requirement, or the inspection method changes. An unlabelled scratch panel in a drawer is not a standard; it is an opinion in physical form. Master samples without traceability have embarrassed more suppliers during launches than any capability study I can recall.

Running the Study Without Fooling Yourself

Practical execution starts with sample count and composition. I look for a minimum of thirty samples per characteristic studied, with a good spread across the judgement range and a deliberate cluster at the boundary. Fifty is better when the characteristic is subtle, such as weld spatter density or paint orange peel. Each of three or four appraisers rates each sample two or three times. Beyond that, the marginal information gained rarely justifies pulling inspectors off the line, though a third replicate does help separate genuine inconsistency from a single lapse.

Keep the conditions honest. Run the study at the actual station, with the actual lighting, actual fixtures, and the actual time pressure of production — not in a quiet room with the quality engineer hovering. Announced studies with management observing produce better numbers and worse truth. The appraisers must not be told which samples are which, must not discuss verdicts mid-study, and must record their call before any consultation. If your plant's practice allows asking the supervisor when unsure, document that reality in the study design, because you are then measuring the system as it operates, including its escalation path.

You are measuring the judgement system as it runs on the floor, not the one you wish you had.

Analysis software will produce the numbers, but read the cross-tabulations by hand before trusting any summary. Look for an appraiser whose verdicts flip between rounds on the same boundary sample — that flags a criterion she cannot apply consistently. Look for systematic bias, where one inspector rejects everything the others pass; that usually reveals different training lineages or a supervisor who once told her to be safe. Each pattern has a different fix: boundary samples and retraining for the first, calibration sessions against master standards for the second.

Fixing the System, Not Blaming the Inspector

When the numbers disappoint, resist the instinct to retrain everyone identically and declare victory. Diagnose first. Poor within-appraiser consistency means the criterion is not decidable from what the inspector can perceive — fix the viewing conditions, improve the limit samples, or move the characteristic to an instrument. Poor between-appraiser agreement with good individual consistency means the inspectors hold private standards; a facilitated session around physical boundary samples, ending in agreed dispositions, closes that gap faster than any slide deck.

Diagnose before you retrain

Inconsistent with herself

  • Verdicts flip across rounds on the same sample
  • Criterion not decidable from what she perceives
  • Fix: viewing conditions, limit samples, or instrumentation

Consistent but disagreeing

  • Stable individual pattern, divergent between appraisers
  • Private standards from separate training lineages
  • Fix: calibration session against agreed boundary samples
The two failure signatures point at different root causes, and applying the wrong fix wastes the study.

Sometimes the honest conclusion is that human vision is the wrong gauge. Surface texture, colour matching, and small geometric deviations migrate eventually to profilometers, spectrophotometers, or vision systems, with the human check retained only for gross anomalies. The agreement study is precisely what justifies that capital request, because it converts a vague unease into evidence that judgement at the boundary is not reliable enough for the risk. I have signed off on several such migrations, and in each case the attribute data made the argument for me.

Treat agreement as a living system, not a launch-event checkbox. Re-run the study when new inspectors qualify, when the product or its finish changes, when lighting is replaced, and on a periodic schedule regardless — annually is common practice for key visual characteristics. Keep the boundary sample library growing, because every difficult disposition production produces is a candidate standard. The plants that do this well stop arguing about whether the inspector was right and start arguing about where the boundary sits — which is the argument actually worth having, and one you can win with samples on the table.

What to Report and What to Expect

A complete attribute study report contains four layers: the within-appraiser kappa per inspector, the between-appraiser kappa, agreement with the reference standard, and the raw confusion matrix showing false accepts and false rejects. A single blended score hides the failure mode. The auditor, the customer's SQE, and your own training plan all read different rows of that matrix, which is why it should be published whole rather than summarised into one number that satisfies nobody's question.

Expect the first study to disappoint. Across automotive and aerospace plants I have seen first-pass kappa values land routinely in the 0.5 to 0.7 range for visual characteristics, even in mature operations with disciplined work instructions. That is not a failure of the method — it is the method working, exposing a boundary nobody had ever measured. The second study, run after calibration sessions and corrected viewing conditions, typically tells you whether your fixes took. The number itself matters less than the trend between studies, which is what turns agreement analysis from an audit artefact into a management tool.

The end state is worth naming. A plant with a governed boundary library, specified viewing conditions, periodic agreement studies, and boundary dispositions that people argue about openly has a visual measurement system it can defend to any auditor under IATF 16949 or AS9100. The plant next door, relying on certificates and good intentions, will keep discovering its inspection weakness during a customer claim. The difference between them is roughly a week of disciplined work per characteristic, repeated once a year. That is one of the cheapest quality investments available, and one of the few that produces evidence rather than assurance.