The andon system is a mechanism for making defects visible at the source. When an operator detects an abnormality, they pull a cord, press a button, or activate a signal. The line stops. A responder arrives. The problem is contained, investigated, and countermeasured before a single defective unit moves downstream. This is the operational core of jidoka—autonomation with a human touch.
In practice, the vast majority of andon installations fail within the first year of operation. The hardware remains functional, the cord hangs within reach, and the board displays zero activations. Management interprets the silence as process stability. The quality data, reviewed in a separate meeting by different people, tells a different story: escape defects rise steadily, internal scrap costs climb, and customer complaints accumulate.
I have audited plants where the andon board had not registered a single pull in six months, while the internal reject rate at final inspection had tripled. Nobody had connected the two datasets. The silence on the line was being rewarded as stability. The defects flowing downstream were treated as a separate problem, solved with additional inspection rather than source containment.
The Anatomy of System Collapse
Andon failure is cultural, not technical. The mechanism of collapse is consistent across industries: the response to the first few pulls determines whether the system lives or dies. When an operator activates the andon and the response is slow, dismissive, or punitive, the behaviour is extinguished. The operator learns that pulling the cord produces delay and social friction, not problem-solving.
The decline follows a predictable sequence. An operator pulls the cord for a genuine abnormality. The supervisor arrives late, asks no diagnostic questions, and orders a restart. The root cause is never addressed. On the next shift, the same defect recurs. The operator pulls again. This time, the supervisor signals frustration. By the third occurrence, the operator stops pulling. The defect continues, undetected until final inspection.
Within weeks, new operators are socialised by their peers. The formal training says pull the cord for any abnormality. The informal training, delivered in thirty seconds by the operator at the next station, says do not pull it. The real system overrides the documented one. The andon board goes quiet, and the defect stream moves downstream.

Response Time: The Standard That Determines Everything
Toyota's internal standard for andon response is ninety seconds. A team leader or supervisor reaches the operator, assesses the situation, and initiates containment within that window. This is not a suggestion. It is a measured, tracked KPI with the same weight as OEE or cycle time. The speed of response is the proof that the organisation takes the pull seriously.
Most manufacturers never set a response-time standard. The responder arrives when available—five minutes, fifteen minutes, sometimes not at all. The operator waits with the line stopped, watching colleagues stand idle. The cost of the stop is visible and immediate; the cost of the defect it prevents is invisible and deferred. Without a measured standard, the unspoken calculus favours running defective parts over stopping the line.
If your line runs at $2,000 per minute and the average andon response takes twelve minutes, each pull costs $24,000 in visible downtime. A defect caught at the source might cost $50 to contain. The same defect discovered at final inspection costs ten to twelve times more. But the $24,000 is on this shift's OEE report, and the downstream cost is on next quarter's scrap variance. The organisation optimises for the number it can see.
Andon Economics: Response Time vs. Defect Cost
The Blame Shift and the Countermeasure Vacuum
The question a supervisor asks after an andon pull reveals the culture instantly. In a functioning system, the question is: "What did you see?" In a failing system, the question is: "Why did you stop?" The first frames the operator as a quality sensor. The second frames the operator as a disruption. Operators who hear the second question stop pulling. The behaviour is extinguished within two or three incidents.
Equally lethal is the countermeasure vacuum. The supervisor arrives, confirms the defect, and applies a workaround: tag it, hold it, run the backup fixture, sort at the end of the line. The defect is deferred, not solved. The operator watches the same abnormality recur on the next shift and the shift after that. Each pull produces the same non-solution. The operator concludes that the cord is a ceremony, not a tool, and stops using it.
The production override is the final mechanism of destruction. At the end of a shift, with targets at risk, a supervisor restarts the line over the operator's objection. This interaction, observed by every operator on the line, communicates that production schedule overrides quality judgment. The authority granted in training is revoked in practice. Once this override occurs even once, the system is compromised. Operators who witness it adjust their behaviour permanently.
Reading the Pull Rate as a Health Metric
A healthy andon system has a steady, non-zero pull rate. New systems often spike as operators surface long-tolerated defects. This is the system working as designed. Management frequently misreads the spike as instability and pressures supervisors to reduce stops. Supervisors pressure operators. The pull rate drops, management is satisfied, and the hidden defect rate climbs in parallel.
The correct response to a high andon rate is accelerated problem-solving, not fewer pulls. If the same defect triggers three pulls in a week, the issue is the unresolved root cause, not the operator's willingness to report it. Track the pull rate as a process health indicator. A sudden drop below the established baseline is as diagnostic as a spike in customer complaints—it indicates that operators have stopped reporting what they see.
Segregate the data. Track first-time pulls (new issues) separately from repeat pulls (unresolved issues). A high repeat-pull rate signals that countermeasures are failing or absent. A low first-time pull rate signals that operators are not detecting or reporting abnormalities. Both conditions require intervention, but the interventions are different: one targets problem-solving speed, the other targets operator engagement and response culture.
Zero andon pulls is never a sign of a stable process. It is a sign that operators have stopped telling you what they see.
Rebuilding a Dead System
Most organisations approach andon revival backwards. They campaign for more pulls, post signs, retrain operators, and then wonder why nothing changes. The pulling rate is a dependent variable. The independent variable is the response. If the response is slow, dismissive, or absent, no amount of training will increase the pull rate. Operators respond to what happens after they pull, not to what the training manual says.
The correct sequence begins with fixing the response mechanism before asking for a single additional pull. Assign dedicated responders with no competing responsibilities during their shift. Set a response-time standard and measure it against every pull. Publish the data on the andon board itself, alongside the pull count. If the organisation cannot guarantee a response within the standard, it has no grounds to ask operators to stop the line.
Close the countermeasure loop within twenty-four hours. The operator who pulled the cord must see what happened as a result. If the root cause requires a longer investigation, communicate the timeline and the interim containment. An operator who pulls the cord and sees no visible consequence assumes nothing happened. From their perspective, nothing did. The loop closure is the evidence that the system functions.
Andon System Revival Sequence
- 01Fix the responseAssign dedicated responders, set a measured response-time standard, publish compliance
- 02Reinforce the pullThank the operator publicly for every pull, including false alarms
- 03Close the loopDeliver a visible countermeasure or interim containment within 24 hours
- 04Track the rateMonitor pull frequency as a health metric; investigate deviations from baseline
- 05Protect authorityMake supervisor override of an andon pull a documented policy violation
Structural Protection of Operator Authority
Operator authority to stop the line must be absolute and written into policy. A supervisor who overrides an andon pull for production reasons must face the same consequence as an operator who bypasses a safety interlock. This is not hyperbole. The override destroys the system's integrity more reliably than any hardware failure, because it communicates to every operator on the line that quality judgment is subordinate to the production schedule.
At WITTE Automotive, I worked with a line where the andon system had been silent for over a year. The first step was not retraining operators. It was sitting with the shift supervisors and establishing that an andon pull would be responded to within two minutes, every time, without exception. We measured compliance. We posted it. Only then did we ask operators to begin pulling again. The pull rate rose from zero to a steady baseline within three weeks, and the downstream defect rate dropped by a measurable margin.
The policy must also address the blame question. Train supervisors and team leaders to open every andon response with the same phrase: "What did you see?" Audit the interactions. If an operator reports being questioned in a way that implies blame for the stop, intervene immediately. The language used in the first thirty seconds of the response determines whether the operator will pull the cord the next time they see something wrong.
An andon system is not a hardware installation. It is a social contract backed by organisational discipline. The hardware makes the defect visible; the response makes the system credible. When the response is fast, respectful, and effective, operators use the system. When it is not, they do not. The question is not whether your factory has an andon system. The question is whether your operators believe it works.
