Total Productive Maintenance: When Your Equipment Reliability Program Becomes a Lubrication Schedule Nobody Owns — and the Availability You Were Supposed to Gain Became the Calendar You Followed and the Breakdowns You Never Actually Prevented

Blog

Walk into any factory that has been running for more than fifteen
years and you will find a maintenance department. It sits in a corner of
the building, staffed by people who know the equipment better than the
engineers who designed it. They carry grease guns and multimeters. They
have a whiteboard full of work orders. And every Monday morning, the
production manager walks over to ask when Machine 7 will be running
again.

This is the reality of maintenance in most manufacturing operations:
a reactive function, measured by how fast it responds to breakdowns,
rewarded for heroics rather than prevention, and treated as a cost
center whose budget gets cut the moment margins tighten.

Total Productive Maintenance was supposed to change all of that.

What TPM Actually Promised

Developed in Japan in the 1970s and formalized through the Japan
Institute of Plant Maintenance, TPM was built on a radical premise:
operators should own the daily health of their equipment. Not mechanics.
Not contractors. The people who run the machines should be the people
who maintain them — at least at the basic level.

The goal was ambitious. Ninety-eight percent equipment availability.
Zero unplanned downtime. Zero accidents. Zero defects caused by
equipment degradation. These were not aspirations scribbled on a
whiteboard. They were measurable targets tied to specific practices:
autonomous maintenance, planned maintenance, focused improvement,
quality maintenance, early equipment management, training, and
safety.

Each pillar had a defined body of knowledge. Each pillar connected to
the others. Operators cleaned, inspected, lubricated, and tightened —
the four activities that prevent the vast majority of equipment
failures. Maintenance technicians handled the complex tasks: overhauls,
precision alignments, predictive maintenance. Engineers analyzed losses
using OEE and its six big availability/performance/quality categories.
Managers tracked progress through audit systems and maturity models.

It worked. Companies that implemented TPM properly — Seiri, Seiton,
Seiso, Seiketsu, Shitsuke as the foundation, autonomous maintenance as
the structure, and focused improvement as the engine — saw dramatic
improvements. Equipment availability jumped from 70% to 90%. Unplanned
downtime dropped by 60%. Maintenance costs as a percentage of production
value fell year over year.

And then the rest of the world heard about it.

What Actually Happens

Here is what I have seen in factory after factory across three
continents.

A consultant arrives. He gives a two-day training on TPM. He shows
before-and-after photos of equipment that has been cleaned, painted, and
labeled. The plant manager is impressed. A kickoff event is scheduled.
Banners are printed. “TPM: Total Participation Makes it Happen!” — a
slogan nobody remembers the origin of.

The initial cleaning event goes well. Operators spend a Friday
afternoon wiping down machines that have accumulated years of grime.
They find leaks, loose fasteners, cracked guards, and a missing safety
cover that nobody reported. The energy is real. People feel proud of
their equipment for the first time in years.

Then Monday comes.

The production schedule does not care about TPM. There are orders to
fill. The autonomous maintenance tasks — the cleaning, inspection,
lubrication, and tightening that are supposed to happen every shift —
get skipped. Not because operators are lazy, but because nobody built
the time into the standard work. The takt time does not include ten
minutes for equipment care. The shift handover does not include a
maintenance checklist. The performance metrics reward output, not
equipment stewardship.

Within three months, the TPM boards are still on the wall, but the
columns are empty. The lubrication tags have faded. The initial cleaning
is a memory. And the maintenance department is back to fighting
fires.

The Six Losses Nobody
Actually Tracks

TPM defines six big losses that reduce equipment effectiveness. Let
me walk through how each one typically gets mistreated in practice.

Equipment failure (breakdowns). Most factories track
MTBF and MTTR, but the data is garbage. Operators log “machine stopped”
without a failure code. Mechanics log “repaired” without root cause. The
maintenance system becomes a ticket queue, not an analysis tool. When
someone finally tries to calculate reliability metrics, they discover
that 40% of the records are unusable.

Setup and adjustments. Changeover time is measured
from the last good part of the previous run to the first good part of
the new run. Except nobody actually measures it that way. Production
logs the “downtime” but excludes the ramp-up scrap, the trial parts, and
the parameter adjustments. The real changeover is twice as long as the
reported number. SMED initiatives fail because the baseline was
fictional.

Idling and minor stoppages. A machine that stops for
thirty seconds to clear a jam, then restarts, does not get logged as
downtime. But over an eight-hour shift, these micro-stops add up to
forty minutes of lost production. Nobody sees it because the andon light
resets, the counter restarts, and the shift report shows “running.” The
only way to catch it is to monitor cycle time integrity — and almost
nobody does.

Reduced speed. Every machine has a designed cycle
rate. Every machine runs slower than that rate. Sometimes intentionally,
because the equipment cannot hold tolerance at full speed. Sometimes
unintentionally, because nobody verified the actual cycle time against
the nameplate. Either way, the performance loss is invisible on most
dashboards. OEE calculations that use nameplate speed show inflated
numbers. OEE calculations that use “best demonstrated rate” show the
truth — and the truth is usually 15 to 25 percentage points worse than
what management believes.

Process defects. Equipment-related defects are
different from process defects caused by material or method. But most
quality systems do not distinguish between them. A burr on a machined
surface could be a dull tool, a worn bushing, or an incorrect feed rate.
Without equipment condition tracking, the quality team attributes all
defects to “process variation” and the maintenance team is never
engaged.

Startup losses. The parts produced during the first
thirty minutes after a cold start are more likely to be out of
specification. Everyone knows this. Nobody quantifies it. The scrap from
startup is mixed into the shift’s total scrap, masking a predictable and
preventable loss that could be eliminated with proper warm-up cycles and
startup procedures.

Autonomous
Maintenance: The Hardest Easy Thing in Manufacturing

The concept is simple: train operators to perform routine equipment
care. The execution is enormously difficult, and the reasons have
nothing to do with technical complexity.

Operators resist autonomous maintenance because they see it as extra
work without extra pay. “I run the machine, I don’t fix it” is a
sentence I have heard in every language I have worked in. And they are
not wrong — the current incentive system pays them for parts produced,
not for equipment maintained.

Maintenance technicians resist it because they see it as a threat to
their expertise. “If operators do my job, why do you need me?” is the
unspoken fear. Never mind that autonomous maintenance was never designed
to replace skilled trades — it was designed to free them up for
higher-level work.

Supervisors resist it because it does not produce parts. Ten minutes
spent on inspection and lubrication is ten minutes not spent running
production. The supervisor’s boss wants output today, not availability
tomorrow.

Breaking through this resistance requires three things, and all three
are non-negotiable.

First, leadership must change the measurement system. If you want
operators to maintain equipment, you must measure and reward equipment
health alongside production output. This means revising standard work to
include maintenance tasks, adjusting takt time calculations, and
changing the shift scorecard.

Second, maintenance must be repositioned as a partner, not a service.
The mechanic’s role shifts from “fix it when it breaks” to “prevent it
from breaking and improve it when it runs.” This is a career change, not
a task change, and it requires training, support, and a different
compensation philosophy.

Third, operators must be trained — properly, not with a one-hour
briefing and a laminated card. They need to understand how the equipment
works, what the failure modes are, what the inspection points mean, and
why each task matters. This takes weeks, not hours. It takes hands-on
practice with a maintenance mentor. And it requires ongoing coaching
until the behaviors become habit.

Predictive
Maintenance: The Technology Trap

Every TPM conference I attend features at least one presentation on
predictive maintenance. Vibration analysis. Oil analysis. Thermography.
Ultrasonic testing. Motor current signature analysis. The technology is
genuine. The physics is sound. The detection capabilities are real.

But here is what happens in practice.

A company buys a $40,000 vibration analyzer. They send a maintenance
technician to a three-day training course. He comes back, starts
collecting data on critical pumps and motors, and discovers that a
bearing on the main compressor is trending toward failure. He writes a
work order. The work order sits in the backlog for three weeks because
production will not release the equipment. The bearing fails. The
compressor goes down. The line stops for sixteen hours.

The technology worked. The organization did not.

Predictive maintenance does not fail because the sensors are
inaccurate or the algorithms are wrong. It fails because detecting an
incipient failure is useless if the organization cannot act on the
prediction. The gap between detection and action is where every PdM
program lives or dies.

Closing that gap requires a maintenance planning function with the
authority to schedule equipment outages based on condition data. It
requires a production planning function that builds flexibility into the
schedule so equipment can be taken offline without missing customer
commitments. And it requires a spare parts strategy that ensures the
right components are available when the data says they are needed.

None of that is technology. All of it is organizational.

The OEE Problem

TPM uses Overall Equipment Effectiveness as its primary metric. The
formula is simple: Availability × Performance × Quality. Multiply the
three percentages together and you get a single number that represents
how well your equipment is performing against its theoretical
potential.

The problem is that OEE has become an end rather than a means.

I have seen plants where the OEE number on the dashboard never drops
below 85%. Walk the floor, though, and the reality is different.
Equipment is down for changeovers that somehow do not get counted. Scrap
rates include only parts that are physically thrown away, not parts that
are reworked. Cycle times are based on “current standard,” which has
been adjusted downward over the years until it bears no resemblance to
the equipment’s actual capability.

The number is high because the number has to be high. The plant
manager’s bonus depends on it. The corporate scorecard ranks facilities
by it. Vendors use it to benchmark. So the definition expands, the
denominator shrinks, and the OEE becomes a performance art rather than a
performance metric.

Real OEE — OEE calculated against nameplate speed, true availability,
and first-pass yield — is almost always lower than what management
believes. Sometimes dramatically lower. A plant reporting 88% OEE might
be running at 55% when measured honestly. That gap is not a reporting
problem. It is a leadership problem.

Building TPM That Actually
Lasts

After twenty-five years of implementing, auditing, and rescuing TPM
programs, I have identified a small number of practices that separate
the programs that sustain from the ones that decay.

Start with one line, not the whole plant. Pilot
programs that focus on a single production line allow you to refine the
process, build internal expertise, and generate credible results before
scaling. Whole-plant rollouts almost always fail because the support
structure cannot keep up with the demand.

Make the first measurable goal equipment cleanliness, not
OEE.
Clean equipment reveals leaks, cracks, and wear that dirty
equipment hides. Cleaning is also the lowest-skill, highest-impact
activity in autonomous maintenance. When operators see that their
cleaning reveals problems, they begin to trust the process.

Tag every defect. When an operator finds a problem
during inspection — a leak, a vibration, a strange noise — they should
be able to tag it with a numbered red tag and know that someone will
follow up. If tags sit unresolved for weeks, the system dies. If tags
are addressed within 48 hours, trust builds quickly.

Track the six losses separately, not just OEE. An
aggregate number tells you nothing about where to act. A breakdown of
the six losses tells you exactly where to focus your improvement
resources. The breakdown should be visible to operators, not just to
engineers.

Conduct monthly equipment audits with cross-functional
teams.
The audit should include an operator, a maintenance
technician, an engineer, and a supervisor. They walk the line together,
inspect each machine against the standard, and identify gaps. The audit
is not punitive — it is diagnostic. Its purpose is to reveal where the
system is breaking down so it can be fixed.

Connect TPM to the business strategy. Equipment
reliability is not a goal in itself. It exists to serve the business:
lower costs, faster delivery, higher quality, safer operations. Every
TPM initiative should be traceable to a business outcome. If it is not,
it will be cut when the next cost reduction program comes through.

The Leadership Question

Every TPM failure I have investigated traces back to the same root
cause: leadership treated it as a program rather than a philosophy.
Programs have start dates and end dates. They have budgets and
champions. They generate reports and presentations. And when the
champion moves on, the program fades.

Philosophies do not have end dates. They become part of how the
organization thinks and acts. They are reflected in hiring criteria,
promotion decisions, capital allocation, and daily priorities. They do
not depend on a single person to sustain them.

The difference between a factory where TPM is a program and a factory
where TPM is a philosophy is visible within thirty seconds of walking
through the door. You see it in the cleanliness of the equipment, the
order of the tools, the confidence of the operators, and the quality of
the conversation between production and maintenance.

You cannot fake it. You cannot audit your way into it. You cannot buy
it from a consultant.

You build it — one machine, one operator, one standard at a time. And
you protect it every single day, because the moment you stop protecting
it, entropy takes over and the equipment goes back to the state it was
in before you started.

That is the real lesson of Total Productive Maintenance. Not the
pillars, not the metrics, not the Japanese terminology. The lesson is
that equipment reliability is a daily discipline, not a quarterly
initiative. The factories that understand this are the ones that run.
The ones that do not are the ones that call me.


Peter Stasko is a Quality Architect with over 25
years of experience in manufacturing quality, reliability engineering,
and continuous improvement. He has implemented TPM, Lean, and Six Sigma
programs across automotive, electronics, and heavy industry on three
continents. Peter writes about the gap between what quality frameworks
promise and what actually happens on the factory floor — and how to
close it.

Scroll top