An Andon system is a visual and auditory escalation tool that immediately informs the shop floor of a process abnormality. Derived from the Japanese word for paper lantern, it illuminates the problem. It is not simply a stack of coloured bulbs on a machine; it is a cultural mechanism that alters the power dynamic between the operator and management. The system explicitly authorises any operator to halt production when they detect an out-of-standard condition.
The core philosophy operates on a single mandate: if something is wrong, stop and report it. Crucially, the organisational structure must guarantee that no operator is punished for triggering the system. Without this psychological safety, the hardware becomes expensive decoration. I have audited plants with sophisticated IoT-connected Andon arrays that recorded zero activations per shift. The technology was flawless; the culture was paralysed by fear.
A functional Andon system enforces a structured response protocol. When an operator pulls the cord or presses the button, the signal broadcasts to a defined support group. In the traditional Toyota Production System, the mandated response time is sixty seconds. If the team leader does not arrive at the station within that window, the escalation automatically advances to the next management tier. This forced immediate response is what separates Andon from standard corrective action loops like 8D.
The Mechanics of Escalation
A standard Andon setup relies on a three-colour signal logic paired with an auditory alarm. Green indicates normal running conditions with no action required. Amber signifies a request for assistance, such as a material shortage, tooling wear, or a missing component. The line continues to move, but the situation requires immediate intervention to prevent a downstream stoppage.
Red indicates a critical quality issue or safety hazard. The line stops completely. This stoppage is deliberate. The operator has identified a deviation that will result in a nonconforming product if production continues. The red signal demands immediate presence from quality engineering, maintenance, and shift leadership to contain the defect, diagnose the root cause, and verify the correction before the line restarts.

The escalation route must be explicitly mapped to avoid confusion during an event. Vague instructions like "call maintenance" will not survive a high-pressure production environment. The response matrix must define who is summoned for an amber material call versus a red quality defect. Clarity at the point of failure ensures the problem is addressed by the correct subject matter expert instantly.
| Signal | Trigger Condition | Required Action |
|---|---|---|
| Green | Standard cycle time, in-specification output | Continue normal production |
| Amber | Operator requires assistance (material, information, tooling) | Support staff must respond within 5 minutes |
| Red | Critical defect, safety hazard, or equipment failure | Line stops immediately; shift manager required on site |
The Pathology of Silent Manufacturing
In plants with weak quality cultures, operators routinely observe defects and say nothing. They remain silent because historical precedent has taught them that reporting issues results in reprimands for slowing down the line. In this environment, defects are not prevented; they are absorbed into the finished goods inventory. The line appears highly efficient right up until the customer rejects the batch.
I took over a greenfield QA department at a plant suffering from this exact pathology. The facility had an automated Andon system integrated into the PLCs, yet activation rates were effectively zero. Meanwhile, internal scrap rates were climbing steadily. Management believed they had a defect generation problem. In reality, they had a defect reporting problem.
When escalation is punished, operators adapt by routing defective parts into rework bins without documentation. This hides the true process capability from the PFMEA and undermines the entire IATF 16949 corrective action framework. The data feeding management dashboards becomes fictional. You cannot improve a process when your baseline metrics are actively manipulated to protect the operators from blame.
Hidden defects compound geometrically. A minor misalignment at station four becomes a catastrophic assembly failure at station twenty. Implementing Andon forces the defect to be addressed at the point of creation, where the root cause is physically present and diagnosable. Stopping the line for two minutes at the injection moulding machine prevents hours of rework at the final assembly stage.
Overcoming Operator Hesitation
Implementing Andon requires dismantling years of ingrained production behaviour. You cannot install hardware and expect compliance. During implementation, I ran workshops where I asked operators a direct question: how many times did you see something wrong last shift but chose not to report it? The average answer was seven separate incidents per operator per shift.
That statistic is what quiet manufacturing actually costs an organisation. Hundreds of unrecorded abnormalities flowing through the value stream weekly. The first step is securing a guarantee from the plant manager that Andon pulls will never negatively impact operator performance reviews. This is a non-negotiable prerequisite. If management breaches this contract even once, trust collapses and the system fails permanently.
To prove the commitment, leadership must physically respond to the signals. When the amber light triggers, the shift supervisor must leave their desk and walk to the station. Their physical presence validates the operator's decision to escalate. This visible reaction trains the shop floor to understand that management prioritises process stability over blind output velocity. The system relies on this immediate feedback loop.
Implementation Approaches
Hardware-first approach
- Purchase expensive IoT-connected light stacks
- Integrate automated PLC triggers for cycle time
- Mandate usage via standard work instructions
- Track Andon pulls as a KPI on manager dashboards
Culture-first approach
- Install basic manual pull cords and analogue lights
- Train operators on out-of-standard identification
- Enforce a strict five-minute management response time
- Review pull data daily to drive Kaizen events
Interpreting the Metrics
When an Andon system begins to function correctly, the immediate reaction from production management is almost always panic. In one early implementation, a plant recorded three Andon activations in the first week. By the fourth week, the daily average had climbed to over twenty. By month two, the plant was logging more than one hundred and twenty signals weekly. Leadership immediately assumed the process was deteriorating.
The reality was the exact opposite. The process had always been generating these defects; the system was simply making them visible for the first time. A low Andon count in a high-complexity automotive plant is rarely a sign of perfection. It is a leading indicator of hidden waste, unreported rework, and suppressed process variation. The data must be framed as discovery, not deterioration.
You do not have more problems; you are simply seeing the waste that has been there all along.
Over a six-month horizon, as the major root causes are systematically eliminated through Kaizen, the activation rate will naturally decline. The initial spike represents the backlog of chronic process issues surfacing simultaneously. The subsequent decline reflects genuine process improvement. Tracking the mean time to respond alongside activation volume provides the full operational picture of system health.
Post-Implementation Stabilisation
Deployment Strategy and Failure Modes
The most common failure mode in Andon deployment is prioritising digital sophistication over operator usability. A complex touchscreen interface will see lower engagement than a simple mechanical pull cord. Systems that require operators to categorise the defect type into sub-menus before escalating create friction. The escalation must be instantaneous. Classification and root cause analysis happen after the line is secured.
Another critical failure is the erosion of the response service level agreement. If an operator pulls the cord and support staff fail to arrive within the mandated window, the system loses credibility. Within weeks, operators will revert to informal workarounds. They will bypass the Andon system and quietly stockpile defective parts to avoid waiting for unresponsive technicians. The five-minute response window must be treated as inviolable.
To scale effectively, begin with a pilot on a single line where the shift supervisor is actively engaged in lean methodology. Install manual pull cords and basic light stacks. Map the response team and measure adherence to the time targets. Refine the process over eight weeks, standardise the work instructions, and then deploy across the remaining production halls. Integrate digital MES connectivity only after the manual culture is established.
Eight-Week Pilot Sequence
- 01Week 1-2: DiagnosticsMap current reaction times and interview operators about existing escalation barriers.
- 02Week 3-4: Manual InstallInstall physical pull cords and light stacks on the pilot line; define the response team.
- 03Week 5-6: SLA EnforcementEnforce strict response windows; track adherence daily and address bottlenecks.
- 04Week 7-8: Data ReviewAnalyse pull data to identify systemic defects for immediate Kaizen action.
Integration with Lean and Quality Frameworks
Andon does not operate in isolation; it is the visual manifestation of Jidoka, the principle of autonomation with human intelligence. It provides the hard stop that prevents a machine from continuously producing scrap. Furthermore, it functions as the primary input for continuous improvement initiatives. Every red light activation should generate a corresponding entry in the site's problem-solving log.
The data captured during these events directly feeds the PDCA cycle. The problem is contained, the root cause is analysed, and the standard work is updated to prevent recurrence. Without the Andon trigger, these defects would bypass the corrective action system entirely. It bridges the gap between real-time floor execution and long-term systemic quality improvement.
Future iterations of this technology will leverage AI to predict and flag abnormalities before the operator even notices them. However, the fundamental principle will remain unchanged. Technology can detect the variance, but human intelligence is required to validate the containment and authorise the restart. The operator's authority to stop the line remains the foundation of quality assurance.
An Andon system is a structural commitment to quality over velocity. It guarantees that the organisation values defect prevention more than it fears production downtime. When management enforces that principle daily, the data takes care of itself.
