An Andon system is a real-time visual and auditory network that gives any operator the authority to stop production when quality is compromised. In its functional form at Toyota, a trigger activates a signal, a team leader arrives within seconds, and the problem is either resolved or the line stops. The mechanism is designed to catch defects at the source, where the cost of correction is lowest.

Most implementations fail within three months. The hardware works, but the organizational response collapses. Operators pull cords, nobody comes, and defective parts move downstream. Red lights stay on permanently. The system becomes set decoration rather than a quality tool.

I have audited plants where every tower light glowed amber and no operator could recall why. The failure is never the technology. It is the absence of response discipline, a culture that penalizes escalation, and metrics that suppress the reporting of problems rather than the problems themselves. These are predictable failure modes with specific countermeasures.

The Three Components of a Functional Andon

A functioning Andon system requires three components: a trigger, a signal, and a response. Remove any one of the three and the system collapses. Most organizations install the first two and assume the third will follow. It does not. The response protocol is the component that requires ongoing organizational discipline and the one that degrades first.

The trigger must give every operator immediate, unrestricted access to activate the system. A pull cord, a button, a foot switch — the design requirement is that activation takes one motion and requires no permission. The moment an operator identifies a problem they cannot resolve within takt time, they trigger the Andon. There is no threshold for what qualifies. If the operator cannot handle it, they pull.

The signal must direct responders to the exact station without ambiguity. A tower light, a distinctive tone, a board displaying the station number — the purpose is locatability. The signal must cut through ambient factory noise. If responders cannot identify which station triggered the alert within seconds, the signal design has failed.

The response is where implementations break down. When the Andon activates, a designated responder must arrive at the station within a predetermined time, measured in seconds. The responder assesses the situation, decides whether the line can continue, and either resolves the issue or escalates. If the problem cannot be fixed within one takt cycle, the line stops. That last sentence is where most organizations part ways with the system entirely.

Quality decisions are made at the process, not in the report that describes it afterwards.
Quality decisions are made at the process, not in the report that describes it afterwards.

Why the Line Stop Becomes Intolerable

The core design principle of an Andon system is that any operator can stop the entire production line. The logic is straightforward: a defect produced at the source costs minutes to fix. The same defect embedded in a finished assembly costs hours of investigation, sorting, rework, and potentially a customer return. Stopping the line is the cheapest correction option available.

In practice, most manufacturing organizations treat line stops as failure. Schedule attainment dominates the morning report. Operators who pull the cord receive sighs, questions about why they could not handle it, and scrutiny about whether the pull was warranted. Within weeks, an implicit hierarchy emerges: minor issues do not justify a pull, only real problems do, and the definition of real drifts upward until the cord is reserved for catastrophic failures.

The organization has redefined the Andon from a quality tool to an emergency tool. In doing so, it eliminates the system's primary value: catching small problems before they propagate. At Toyota, a new line triggers thousands of pulls in its first weeks. Each pull represents a problem found and addressed. As root causes are eliminated, frequency decreases. The system is working because problems are being solved, not because pulls are rare.

An Andon system that never triggers is not a sign of excellence. It is a sign that operators have stopped looking, stopped believing anyone will respond, or concluded that pulling the cord carries more personal cost than passing the defect downstream.

The Gradual Erosion of Response Discipline

Andon failure rarely happens in a single event. It happens through incremental compromises, each defensible in isolation. A team leader is in a meeting. A maintenance technician is at another machine. The Andon activates and response takes three minutes, then five, then ten. The operator stands at the station watching parts arrive while the line continues to move.

The operator faces a binary choice: hold the station and allow work-in-process to accumulate, or pass the defective part and maintain flow. Most operators, measured on output and trained to keep the line running, pass the part. The defect moves downstream where it is discovered — or embedded further. The Andon has failed, and the operator has learned that pulling the cord does not produce help.

Response Discipline: Designed Behaviour vs Operational Reality

What the system was designed to do

  • Responder arrives at the station within seconds, not minutes
  • Problem is assessed, resolved, or escalated within one takt cycle
  • Line stops automatically if the issue cannot be contained
  • Every pull is treated as a valid signal requiring human attention

What actually happens after 90 days

  • Response time stretches to several minutes as competing priorities win
  • Operator passes the defective part rather than holding the line
  • Line stop is treated as a management failure, not a quality safeguard
  • Pulls decrease because operators learn that nobody comes
The gap between the implemented protocol and actual shop-floor practice within 90 days of go-live.

This erosion is the most common Andon failure mode and the hardest to detect because it is gradual. The daily log records pulls and response times, but nobody reviews it because the production board is more urgent. Team leaders who initially sprinted to stations begin to walk, then finish what they are doing first, then delegate to someone else. The culture of immediate response dissolves into a culture of eventual response, which is functionally identical to no response at all.

Habituation and the Death of Signal Integrity

Habituation is a property of the human nervous system: when a stimulus repeats without consequence, the brain stops registering it. The amber light that has been on at a station for four hours has become invisible to everyone walking past it. This is not carelessness. It is neurology. Any Andon system that does not actively manage against habituation will be defeated by it.

The defence is signal discipline. An active Andon must mean something is happening right now that requires human attention. Not that something happened earlier. Not that a machine is idling. Not that someone forgot to reset a status flag. An Andon in active state represents a live problem being worked by a live person. The moment the problem is resolved or the line is stopped, the system resets.

An Andon that is permanently amber or red is a system failure, not a status update.

Walk through most factories with installed Andon systems and you will see a constellation of amber lights that have been on so long the bulbs will fail before anyone resets them. Each represents a problem detected, reported, and absorbed into the background. The signal has lost its meaning because the organization allowed it to represent everything, which means it represents nothing.

How Metrics Suppress the Reporting of Defects

Well-intentioned measurement is one of the most destructive forces acting on an immature Andon system. A manager begins tracking pull counts, response times, and downtime attributed to Andon stops. In a mature system with established psychological safety, these are useful operational metrics. In an immature system, they become weapons used against the people the system depends on.

Operators discover that high pull counts attract scrutiny. A question about why Line 3 pulls the cord three times more than Line 1, asked with implied criticism, is sufficient to suppress reporting. Supervisors coach operators to use judgment before pulling, which in practice means do not pull unless the problem is serious. Within a quarter, pull counts drop. The manager reports success. The metrics have improved.

The measurement has suppressed the reporting without suppressing the problems. Defects still occur. Operators still see them. They have stopped signalling because the cost — scrutiny, questions, criticism — exceeds the demonstrated benefit, which is zero because nobody comes when they pull anyway. The correct measure of system health is not pull frequency. It is whether problems are being solved and whether repeat issues are decreasing.

Corrective Sequence for a Degraded Andon System

  1. 01Restore response disciplineMandate immediate human response to every activation; treat missed responses as serious deviations
  2. 02Audit signal environmentReset every active light that does not represent a live issue; investigate any signal held over one shift
  3. 03Decouple operator metrics from pull frequencyMeasure response time, resolution time, and root-cause closure; never penalize an operator for pulling
  4. 04Act on what the system revealsUse pull data to identify process weaknesses, training gaps, and supplier quality issues
  5. 05Sustain through daily reviewAudit response times, signal states, and closure rates at every shift handover
The order of operations matters: restoring response discipline before addressing the hardware or the metrics prevents re-erosion.

Rebuilding Response Discipline and System Integrity

A degraded Andon system can be rebuilt, but only if leadership accepts that the failure is organizational, not technical. The system did not fail because operators stopped pulling. It failed because the organization stopped responding. Rebuilding starts with response discipline, not with new hardware or retraining. Before replacing a single bulb, establish and enforce the expectation that every activation receives an immediate human response.

This requires team leaders whose primary responsibility is Andon response. It requires coverage plans for breaks, meetings, and absences so that no activation goes unanswered. It requires measuring response time and treating a missed response with the same seriousness as a safety deviation. If the organization is not prepared to enforce this, the hardware should be removed — non-functional Andon systems are worse than none because they teach operators that escalation is futile.

Signal integrity comes next. Audit every Andon in the facility. Reset any light that does not represent a live issue. Any signal held in the same state for more than one shift without action is itself a defect and must be investigated through an 8D or equivalent root-cause process. Establish the principle that an Andon is either active or off. There is no third state, no permanent amber, no transitional condition that the organization has learned to tolerate.

Finally, decouple Andon usage from individual performance metrics. Operators should never be penalized, directly or indirectly, for pulling the cord. The metrics that matter are response time, resolution time, and root-cause closure rate aligned with Cpk targets and IATF 16949 requirements. A functioning system will show an initial spike in pulls followed by a gradual decline that corresponds to actual problem elimination — not operator disengagement. The system was never the lights. It was always the response.