You installed the towers. Green, yellow, red — sleek LED columns
rising from every workstation like sentinels of lean manufacturing. The
project was announced with fanfare. The supplierdemo was impressive. The
capital request sailed through because everyone agreed: visual
management is a hallmark of world-class operations.
Six months later, those towers glow in a permanent, flickering amber.
Operators stopped pulling the cord after the second week. Supervisors
glance at the colors the way commuters glance at traffic lights —
reflexively, without processing. The board that was supposed to display
real-time escalation data shows yesterday’s shift summary, frozen on a
number nobody remembers compiling.
Another andon system has died. Not with a bang, but with a
whimper.
What an Andon
System Is Actually Supposed to Do
The word comes from Japanese — literally “lantern.” In its original
Toyota Production System context, an andon is a signaling system that
gives any operator the authority — and the obligation — to stop
production when something goes wrong. Not a suggestion. Not a last
resort. An expected, encouraged, celebrated action.
The mechanics are simple. An operator encounters an abnormality: a
defective part, a machine malfunction, a missing component, a safety
concern. They pull a cord or press a button. A light tower activates. A
specific sound plays. The team leader is supposed to arrive within a
defined interval — typically 60 to 90 seconds. The problem is assessed.
Production either resumes or stops for corrective action.
The philosophy underneath those mechanics is radical: the
person closest to the work has the best view of the work, and stopping
production to fix a problem is always cheaper than passing a defect
downstream.
That philosophy requires something most organizations underestimate:
a management culture that genuinely prefers stopping over shipping
problems.
Where Andon
Implementations Go Wrong
The Cord Nobody Pulls
The most common failure mode is also the most predictable. Operators
stop triggering the andon because the response is worse than the
problem. A team leader arrives already irritated. Questions come with a
tone: “Is this really worth stopping for?” The operator senses that
pulling the cord creates friction, not support. So they adapt — they
handle issues quietly, work around abnormalities, and the andon tower
becomes decoration.
This isn’t a technology problem. It’s a leadership behavior problem.
When the first three pulls result in a visible management reaction that
ranges from annoyed to hostile, the social contract of the andon is
broken. No operator will voluntarily invite that reaction twice.
The Alarm Nobody Answers
Some organizations implement the hardware correctly — lights, sounds,
escalation tiers — but never define the response protocol. An operator
pulls the cord. The light flashes. And then… nothing happens. Or
someone comes eventually, but there’s no time standard, no escalation
path if the first responder can’t solve it, and no documentation of what
occurred.
An andon without a defined response cycle is a fire alarm with no
fire department. The hardware makes noise. The system behind it doesn’t
exist.
The Metric That Backfires
Here’s a scenario that plays out repeatedly: leadership decides to
track andon pulls as a performance metric. The intention is reasonable —
they want visibility into how often problems are being caught and
escalated. But the moment andon pulls become a measured KPI, two
distortions emerge.
First, operators who pull frequently are perceived — explicitly or
implicitly — as troublemakers or as evidence of a poorly performing
area. The data meant to encourage escalation becomes a weapon against
it. Second, supervisors start managing the number of pulls rather than
the problems behind them. “We need to reduce andon activations” becomes
the goal, and the surest way to reduce activations is to discourage
them.
The correct metric is response time and problem resolution rate, not
activation count. But that distinction is often lost by the third
monthly review meeting.
The System Nobody Updates
Andon systems that include digital displays — problem descriptions,
downtime counters, shift metrics — require maintenance. Not just
hardware maintenance, but content maintenance. If the display shows
generic messages, outdated targets, or irrelevant data, operators stop
looking at it. The visual management component degrades into background
noise.
What a
Functional Andon System Actually Looks Like
A working andon system has five characteristics that distinguish it
from a tower of colored lights:
Immediate and unambiguous signaling. When an
operator identifies an abnormality, the signal takes one action — pull,
press, toggle — and the response is instant. No forms to fill out before
pulling. No approval needed. The authority to signal is absolute and
belongs to every operator on the line.
Defined response intervals with escalation. A team
leader must arrive within a stated time — Toyota’s standard is roughly a
takt time, often 60 seconds. If the first responder cannot resolve the
issue, it escalates to the next level automatically. The clock is
visible. The escalation is structural, not discretionary.
Problem-first, not blame-first culture. When the
cord is pulled, the conversation starts with “What’s happening?” — never
“Why did you pull?” The operator is treated as the messenger, not the
cause. Leaders who model this behavior consistently are the ones whose
andon systems survive past the pilot phase.
Structured problem capture. Every activation
generates a brief record: what was the abnormality, when did it occur,
what was the immediate countermeasure, what is the follow-up action.
This isn’t a 20-field form — it’s a 30-second capture that feeds into
the daily improvement process.
Closing the loop. Operators who pull the cord hear
back. Not always immediately, but within the shift or the next day,
someone tells them: “That issue you flagged — here’s what we found,
here’s what we changed.” Without loop closure, operators conclude that
pulling the cord makes no difference. And they’re right.
The Technology Trap
Modern andon systems offer sophisticated capabilities — IoT-connected
towers, real-time dashboard analytics, mobile alerts to supervisors’
phones, integration with MES and SCADA systems. These tools can enhance
an andon system that already works. They cannot rescue one that
doesn’t.
I’ve watched organizations spend six figures on digital andon
platforms, deploy them across multiple plants, and produce worse
outcomes than a $200 light tower and a disciplined response protocol.
The technology created the illusion of a system while masking the
absence of the cultural and operational foundations underneath.
The sequence matters. Culture first — leadership that genuinely wants
to hear about problems. Protocol second — defined response times,
escalation paths, loop closure. Technology third — and only to the
extent that it serves the first two.
The Cost of a Silent Andon
When an andon system fails, the cost is not just the capital
investment in hardware and software. The real cost is the missed
defects, the unreported near-misses, the process drifts that compound
over weeks until they surface as a customer complaint or a scrap event
that costs more than the entire andon implementation.
There’s also an opportunity cost that’s harder to quantify. Every
time an operator decides not to pull the cord, a piece of information
about the process is lost. That information — about recurring minor
faults, about ergonomics issues, about component variability — is the
raw material of continuous improvement. A silent andon doesn’t just fail
to stop defects. It fails to generate the knowledge that prevents future
ones.
Rebuilding a Broken System
If your andon system has degraded — and most have — recovery is
possible, but it requires a deliberate reset.
Start with a single line or cell. Pick an area where leadership is
willing to invest in the behavioral change. Brief the operators
personally — not via memo, not via a PowerPoint at the quarterly
kickoff. Walk the floor. Explain that the andon is being reactivated and
that pulling the cord is not just allowed but expected.
Then do the hardest part: for the first two weeks, respond to every
single pull with visible, positive engagement. Arrive quickly. Ask
what’s happening. Fix what you can on the spot. Document what you can’t.
And tell the operator what happened next.
Track the data — but share it transparently. Show operators how many
pulls occurred, what the average response time was, and what problems
were solved as a result. Make the system’s value visible to the people
who power it.
Expand only when the pilot line demonstrates that the culture has
shifted. A single working andon cell is worth more than twenty
installed-but-dormant towers across a plant network.
The Real Test
You’ll know your andon system works when pulling the cord is
unremarkable. When an operator signals an abnormality, a team leader
appears, the issue is addressed, and production continues — all without
drama, without fear, and without a meeting afterward to discuss “why so
many pulls this shift.”
The lantern is supposed to illuminate problems so they can be solved,
not decorate the factory floor. If your andon system isn’t helping you
see and solve problems in real time, it’s not an andon system. It’s a
light show. And the difference between the two has nothing to do with
the lights.
Peter Stasko is a Quality Architect with over 25
years of experience in manufacturing quality, lean implementation, and
production system design. He has led andon system deployments across
automotive and industrial operations and has spent decades studying why
visual management succeeds in some plants and fails in others. He writes
about the gap between lean theory and shop-floor reality — because the
best system in the world is worthless if the operator won’t pull the
cord.