Somewhere in your facility, mounted on a pillar or strung above a workstation, there is a cord. It might be yellow, red, or connected to an electronic button with a glowing indicator. Whatever its form, it represents one of the most radical ideas in lean manufacturing: that any operator, at any time, has the authority to stop production. The andon system — borrowed from the Japanese word for lantern — was one of Toyota's most consequential innovations within the Toyota Production System.
The concept is deceptively simple. When an operator detects an abnormality, they pull the cord or press the button. A signal lights up, a team leader responds immediately, and production stops if necessary. The problem is fixed, the line restarts, and the defect never reaches the customer. In theory, it is a closed loop of detection, response, and prevention. In practice, in most factories outside Toyota City, the cord hangs untouched and the defects keep coming.
The failure of an andon system is almost never technical. The cords work, the buttons work, the displays work. What fails is the cultural and organisational infrastructure that must surround them. Having audited and implemented quality systems across automotive and aerospace plants across Europe, I have seen the same pattern repeat: leadership installs the hardware, declares the culture changed, and then acts surprised when defect rates remain unchanged while the andon board stays perpetually green.
What an Andon System Actually Requires
Before examining why these implementations fail, it is worth being precise about what a functioning andon system is. It is not just a cord or a button. It is a complete communication and response mechanism consisting of three integrated components that must function together. A cord without a response protocol is theatre. A response protocol without a signal mechanism is wishful thinking. A visual board that nobody looks at is wallpaper.
The first component is the signal mechanism: the cord, button, or electronic display that an operator activates when they detect a problem. The second is the visual management board, which displays the status of each workstation and makes problems visible to everyone in the area. The third is the response protocol — the defined process by which a team leader responds, diagnoses the problem, and either resolves it or escalates it according to a standard work sequence.
Sophisticated andon systems track not just whether a cord was pulled, but why it was pulled, how quickly the response came, and whether the root cause was addressed. They generate data that reveals patterns: which stations have the most stops, which problems recur, and which team leaders resolve issues fastest. But data collection is not the purpose of andon. The purpose is to stop defects at the source, building quality into the process rather than inspecting it in at the end. When an andon system works, it transforms quality from a department into a daily operator behaviour.
All three components must function together continuously. When I assess a plant's IATF 16949 or AS9100 compliance, I do not just look for the presence of these components. I look for evidence of their integration. Are the boards updated in real time? Do team leaders know the response standard? Is the escalation path clear? Without integration, you have spent capital on visual management that signals nothing.
The Three-Month Death Cycle
Andon implementation typically follows a predictable pattern in organisations that are not Toyota. Phase one is enthusiasm: leadership reads about the system, visits a reference plant, and returns convinced this is the missing piece. Budget is allocated, cords are installed, buttons are mounted, and displays are programmed. The factory looks modern and lean. Training follows, and workers are told they have stop-the-line authority.
Phase three is the first pull. Someone activates the andon. The line stops. A team leader arrives, often visibly irritated, and asks what happened. The problem is diagnosed, but it takes longer than expected and production falls behind. The shift supervisor appears, asking why output is down. The team leader is asked — not explicitly, but clearly enough — whether it was really necessary to stop the whole line for that issue.

Phase four is the second pull, days later. The response is slower. The irritation is more visible. The operator senses that stopping the line was the wrong move socially, even if it was correct technically. By phase five, the cords hang undisturbed, displays show green across every station, and production numbers are met. The quality data tells a story nobody connects to the dark andon board: defect rates unchanged, customer complaints steady, and rework rates constant. The entire cycle takes about three months.
The Organisational Forces That Silence Workers
The forces that silence andon systems are specific and structural. Production pressure is the most powerful. In most factories, the primary metric is output. Operators know that stopping the line reduces shift units. They know their supervisor's performance is measured on throughput. They know that pulling the cord will be perceived as reducing output rather than protecting quality. When an operator must choose between meeting a target and pulling a cord for an unknown duration, the target wins every time.
Social cost reinforces this. Pulling an andon cord is a public act. It signals to everyone that you have identified a problem. If the problem turns out to be minor, or if the team leader dismisses it, or if production falls behind and coworkers must work faster to catch up, the social cost falls on the puller. Most operators will tolerate a known defect passing to the next station over the certainty of being the person who slowed everyone down.
Response quality determines whether the system survives. The first few times a cord is pulled, the response is thorough. Over time, as the novelty wears off and the same problems recur, team leaders begin to respond perfunctorily. A quick fix is applied, the line restarts, and the root cause is never addressed. Operators learn that pulling the cord produces a temporary patch, not a permanent solution. They learn the system is designed to restart the line as fast as possible, not to solve problems.
In organisations where blame is assigned rather than problems solved, pulling the cord is an act of courage. It identifies a problem and, implicitly, the process responsible for it. The operator who pulls may be asked why they did not prevent it. The operator at the previous station may be blamed for passing a defect. The team leader may be criticised for not catching the issue earlier. In a blame culture, the andon cord becomes a tool for identifying scapegoats, and operators learn quickly to keep their hands off it.
The Data That Reveals the Truth
One of the most diagnostic questions a quality professional can ask about a factory is simple: how many andon pulls happened last month? Not how many defects were found. Not how many stops occurred. How many times did an operator voluntarily stop the line? In a healthy andon culture, the number is surprisingly high — reference Toyota plants, where the system was pioneered, recorded thousands of pulls per day across their facilities.
A factory with zero andon pulls and non-zero defect rates has taught its workers that stopping the line is punished, not rewarded.
In most factories that have installed andon systems, the number is close to zero. Not because there are no abnormalities — the defect rates prove otherwise — but because operators have learned that the cord is not really for pulling. This discrepancy is extraordinarily revealing. The andon pull rate is one of the most honest metrics in manufacturing. It measures not what your quality system is designed to detect, but what your culture allows people to act on.
| Indicator | Healthy System | Decorative System |
|---|---|---|
| Andon pull frequency | Proportional to defect detection rate | Near zero despite non-zero defect rates |
| Team leader response time | Consistent, within defined standard | Increasing over time, highly variable |
| Root cause closure rate | Permanent countermeasures applied | Same problems triggering pulls weekly |
| Pull source distribution | Distributed across the team | Concentrated among one or two operators |
What Organisations Miss About Toyota
When organisations study Toyota's andon system, they focus on the visible elements: the cords, the boards, the protocols. They install replicas and expect similar results. What they miss is the invisible infrastructure that makes the visible elements work. Job security is foundational. Toyota's lifetime employment tradition creates a fundamentally different relationship between worker and company. An operator who stops the line is not risking their job — they are doing their job.
Leader standard work is the other missing element. At Toyota, team leaders are stationed on the floor, and a core component of their standardised work is responding to andon pulls. Their role depends on responding quickly and effectively. In most other organisations, responding to andon pulls is something a supervisor does in addition to their real job, which is managing production numbers. The signal this sends is unmistakable: production is the priority, and andon is an interruption.
Problem-solving capability completes the picture. Toyota invests heavily in developing root cause analysis skills at every level. When an operator pulls the cord, the responding team leader has been trained in structured problem-solving methods — 8D, 5-Why, countermeasure development, and standard work revision. They have the skills and authority to fix the problem permanently. Where team leaders lack these skills, the response is limited to restarting the line, and the same problems recur indefinitely.
Management commitment, measured in minutes rather than words, closes the loop. At Toyota, when an andon pull cannot be resolved within a defined timeframe, the escalation path leads rapidly to senior management. Plant managers are expected to leave their offices and go to the point of the problem. This is standard work, not an exceptional event. Where senior management is never seen on the shop floor during an andon event, operators correctly infer that andon is a shop-floor tool, not a company-wide commitment.
Rebuilding the System From the Floor Up
Fixing a broken andon system is not a matter of retraining operators or replacing hardware. It requires rebuilding the infrastructure that makes pulling the cord feel safe, valued, and productive. Start with leadership behaviour. Before asking operators to pull the cord, every manager and supervisor must demonstrate publicly and repeatedly that they value stops over passed defects. The plant manager should be present on the floor during andon events. Supervisors should thank operators for pulling, even when the problem turns out to be minor.
Functional Andon Response Sequence
- 01SignalOperator detects abnormality and activates cord, button, or display immediately.
- 02ResponseTeam leader arrives at station within defined time standard (typically under 60 seconds).
- 03AssessmentLeader diagnoses issue with operator, classifies severity, and decides on stop or continue.
- 04CountermeasureImmediate containment applied; escalation triggered if resolution exceeds time threshold.
- 05Root causeProblem logged with category and data; permanent countermeasure assigned with owner and date.
Redefine the response. The team leader's reaction to an andon pull should be standardised like any other production process. Define the maximum response time. Define the problem-solving steps. Define when escalation is required and what a successful resolution looks like. Make the response as structured and repeatable as the production process it interrupts. Without standard work for the response, quality degrades to whichever supervisor happens to be on shift.
Measure and publish the right numbers. Stop tracking only output and defect rates. Start measuring andon pull frequency, response times, and root cause closure rates. Publish these alongside production numbers in daily meetings. Make it visible that pulling the cord is as important as meeting the production target. When I introduced Routing Verification KPIs at a major aerospace manufacturer, the principle was the same: you get the behaviour you measure and display. If andon metrics are not on the board, they do not matter.
Protect the puller. Create explicit, enforced protections for operators who activate the andon. No disciplinary consequences, no social penalty, no impact on performance evaluations. Make it clear in policy and in practice that stopping the line to report a problem is always the correct decision. Accept the short-term cost: when an andon system begins working, production output will initially decrease. This is the system surfacing problems that were previously hidden. The output drop is the price of long-term quality improvement, and organisations that cannot accept this cost will never have a functioning andon system.
