The capital request sails through because everyone agrees visual management is a hallmark of world-class operations. Six months later, those LED columns glow in permanent, flickering amber. Operators stopped pulling the cord after the second week. Supervisors glance at the colours the way commuters glance at traffic lights: reflexively, without processing.

Another andon system has died. The board that was supposed to display real-time escalation data shows yesterday's shift summary, frozen on a number nobody remembers compiling. The hardware is intact. The management system that gave it meaning never existed.

I have audited plants where the andon infrastructure cost six figures and produced worse outcomes than a 200-dollar light tower and a disciplined response protocol. The technology created the illusion of a system while masking the absence of operational foundations. The fix is never the hardware. It is the protocol, the culture, and the feedback loop.

What an Andon System Is Actually Supposed to Do

In its original Toyota Production System context, an andon is a signalling system that gives any operator the authority and the obligation to stop production when something goes wrong. It is not a suggestion or a last resort. It is an expected, encouraged action built into the standard work of the station.

The mechanics are straightforward. An operator encounters an abnormality: a defective part, a machine malfunction, a missing component. They pull a cord or press a button. A light tower activates and a specific sound plays. The team leader is supposed to arrive within a defined interval, typically 60 to 90 seconds. The problem is assessed, and production either resumes or stops for corrective action.

The philosophy underneath those mechanics is radical. The person closest to the work has the best view of the work, and stopping production to fix a problem is always cheaper than passing a defect downstream. That philosophy requires something most organisations underestimate: a management culture that genuinely prefers stopping over shipping problems.

When the andon is functioning, the escalation is structural, not discretionary. The clock is visible. If the first responder cannot resolve the issue, it escalates to the next level automatically. The operator does not have to argue for attention. The system demands it.

Where Andon Implementations Go Wrong

The most common failure mode is also the most predictable: operators stop triggering the andon because the response is worse than the problem. A team leader arrives already irritated. Questions come with a tone: is this really worth stopping for? The operator senses that pulling the cord creates friction, not support.

When the first three pulls result in a visible management reaction that ranges from annoyed to hostile, the social contract of the andon is broken. No operator will voluntarily invite that reaction twice. They adapt, handle issues quietly, work around abnormalities, and the tower becomes decoration.

Quality decisions are made at the process, not in the report that describes it afterwards. If the operator will not signal, the system is already broken.
Quality decisions are made at the process, not in the report that describes it afterwards. If the operator will not signal, the system is already broken.

Other organisations implement the hardware correctly but never define the response protocol. An operator pulls the cord, the light flashes, and nothing happens. Or someone comes eventually, but there is no time standard, no escalation path if the first responder cannot solve it, and no documentation of what occurred. An andon without a defined response cycle is a fire alarm with no fire department.

Andon systems that include digital displays require content maintenance, not just hardware maintenance. If the display shows generic messages, outdated targets, or irrelevant data, operators stop looking at it. The visual management component degrades into background noise within weeks.

The Metric That Backfires

Leadership decides to track andon pulls as a performance metric. The intention is reasonable: they want visibility into how often problems are being caught and escalated. But the moment andon pulls become a measured KPI, two distortions emerge that quietly destroy the system.

First, operators who pull frequently are perceived as troublemakers or as evidence of a poorly performing area. The data meant to encourage escalation becomes a weapon against it. Second, supervisors start managing the number of pulls rather than the problems behind them. The surest way to reduce activations is to discourage them.

Andon Metrics That Work vs Metrics That Suppress

<60sTarget response timeTeam leader arrival window, measured from activation
100%Loop closureEvery pull gets feedback to the operator within one shift
0Activation targetsNever set a quota for pulls. Measure resolution, not frequency
1.33Process Cpk triggerOptional: auto-signal when capability drops below threshold
Track response speed and resolution, not activation count. Counting pulls punishes the behaviour you are trying to encourage.

The correct metric is response time and problem resolution rate, not activation count. But that distinction is often lost by the third monthly review meeting. Once supervisors are measured on pull frequency, the data is already corrupt and the operators have already adapted to the new reality.

The Technology Trap

Modern andon systems offer sophisticated capabilities: IoT-connected towers, real-time dashboard analytics, mobile alerts to supervisors' phones, integration with MES and SCADA systems. These tools can enhance an andon system that already works. They cannot rescue one that does not.

I have watched organisations spend six figures on digital andon platforms, deploy them across multiple plants, and produce systematically worse outcomes than a simple light tower paired with a disciplined response protocol. The technology created the illusion of a system while masking the absence of the cultural and operational foundations underneath it.

A digital andon platform without a response protocol is a fire alarm with no fire department.

The sequence matters. Culture first: leadership that genuinely wants to hear about problems. Protocol second: defined response times, escalation paths, loop closure. Technology third, and only to the extent that it serves the first two. Reverse the order and you get an impressive demo followed by a silent factory floor.

Integration with MES and SCADA adds value only when the manual response cycle is already functioning. If the team leader does not arrive within 60 seconds on a manual system, the mobile alert will also be ignored. The failure mode is behavioural, and no software patch fixes it.

Rebuilding a Broken System

If your andon system has degraded, recovery is possible but requires a deliberate reset. Start with a single line or cell. Pick an area where leadership is willing to invest in the behavioural change, not just the hardware. Brief the operators personally, not via memo or a PowerPoint at the quarterly kickoff.

Andon System Reactivation Sequence

  1. 01Select pilot lineChoose a cell where leadership will commit to visible, positive response
  2. 02Brief operators in personExplain reactivation on the floor, not via memo. Make expectation explicit
  3. 03Two-week intensive responseEvery pull gets immediate, constructive engagement. No exceptions
  4. 04Share data transparentlyShow pulls, response times, and resolutions back to the operators
  5. 05Expand to next cellReplicate only after the pilot demonstrates sustained behavioural shift
A structured restart protocol for a line where the andon has gone silent. Pilot first, expand only when response discipline is proven.

Walk the floor. Explain that the andon is being reactivated and that pulling the cord is not just allowed but expected. Then do the hardest part: for the first two weeks, respond to every single pull with visible, positive engagement. Arrive quickly. Ask what is happening. Fix what you can on the spot. Document what you cannot.

Tell the operator what happened next. Every activation must generate a brief record: what was the abnormality, when did it occur, what was the immediate countermeasure, what is the follow-up action. This is a 30-second capture, not a 20-field form. It feeds the daily improvement process and proves to the operator that the signal was received.

The Cost of a Silent Andon

When an andon system fails, the cost is not just the capital investment in hardware and software. The real cost is the missed defects, the unreported near-misses, the process drifts that compound over weeks until they surface as a customer complaint or a scrap event that costs more than the entire andon implementation.

There is also an opportunity cost that is harder to quantify. Every time an operator decides not to pull the cord, a piece of information about the process is lost. That information, about recurring minor faults, about ergonomics issues, about component variability, is the raw material of continuous improvement.

A silent andon does not just fail to stop defects. It fails to generate the knowledge that prevents future ones. The 8D corrective action process depends on capturing problems at the source, when the evidence is fresh and the context is intact. By the time a defect reaches final inspection or the customer, the root cause has often been masked by rework and workaround.

You will know your andon system works when pulling the cord is unremarkable. An operator signals an abnormality, a team leader appears, the issue is addressed, and production continues. No drama, no fear, no meeting afterward to discuss why so many pulls this shift. The lantern illuminates problems so they can be solved. If it is not doing that, it is not an andon system. It is a light show.