Total Productive Maintenance arrived in plants with a seductive pitch: zero breakdowns, zero defects, zero accidents. The mechanism was just as appealing — blend the skills of operators and maintenance technicians so completely that the line between them disappears. Operators handle the daily care, maintenance handles the deep technical work, and everybody wins.
The goal was world-class performance. OEE above 85%, mean time between failures measured in months instead of days, and maintenance costs as a percentage of asset replacement value trending steadily downward. It sounded like the definitive answer to every breakdown that had ever cost a critical shipment.
Eighteen months in, the reality on the floor looks different. The promise of equipment excellence has been reduced to a laminated card clipped to each machine. Operators check the boxes quickly, without thinking, sometimes for equipment they have not actually looked at. The chore survived and the inspection died.
Autonomous Maintenance decayed into a checklist
The cleaning itself has become the visible signal that TPM is happening. When executives walk through, they see gleaming machines, colour-coded labels, and shadow boards with precisely placed tools. They see operators wiping down housings and think this is equipment excellence. What they do not see is that the cleaning was supposed to be the entry point to understanding — not the destination.
The act of cleaning a machine was designed to teach the operator where the heat builds up, where the chips accumulate, where the oil migrates, and where the wear shows first. It was supposed to be an inspection disguised as a chore. Instead, the checklist replaced the cognitive engagement required to actually understand the equipment.
Autonomous Maintenance was designed as a seven-step progression: initial cleaning, countermeasures to contamination sources, cleaning and lubrication standards, general inspection training, autonomous inspection, standardization, and full autonomous management. Each step was supposed to be a gate. You do not move forward until the current step is genuinely mastered. In most implementations, the gates were opened wide because the project timeline dictated it was time to move on.
Operators who never truly learned general inspection were progressed to autonomous inspection. They were handed sheets filled with technical terms they had never been taught to interpret. The sheets came back clean. Not because the machines were flawless, but because the inspectors could not tell the difference.

Focused Improvement turned into monthly theatre
Focused Improvement, or Kobetsu Kaizen, is the pillar dedicated to cross-functional teams attacking the biggest equipment losses with structured problem-solving. In theory, these are your best people spending dedicated hours analyzing why a specific machine loses hundreds of hours a quarter to minor stops, using tools like Why-Why analysis and Pareto charts to drive those losses down permanently.
In practice, Focused Improvement became a monthly meeting where people already stretched thin on their actual jobs sat in a room for forty-five minutes and brainstormed ideas on a flipchart. The action items were assigned to people who already had action items from three other pillars. The meeting adjourned with renewed commitment, but the losses continued uninterrupted.
The problem is not laziness. The problem is that Focused Improvement requires something most plants refuse to give it: protected time. Real Kobetsu Kaizen means pulling your best operator off the line for hours, or even days, to observe, measure, analyze, and experiment. It means accepting that a short-term production hit is worth the long-term reliability gain. Most plants cannot make that trade because this week's OEE target takes priority over preventing next month's breakdown.
The Seven-Step Autonomous Maintenance Progression
- 01Initial CleaningFinding latent defects through hands-on contact with the equipment.
- 02CountermeasuresEliminating sources of contamination and inaccessible areas.
- 03Provisional StandardsCreating clear, visual rules for cleaning, lubrication, and tightening.
- 04Inspection TrainingTransferring mechanical knowledge so operators understand what to look for.
- 05Autonomous InspectionOperators applying trained knowledge, not just ticking boxes.
- 06StandardizationConsolidating learnings into workable, universal standards across shifts.
- 07Full Autonomous ManagementContinuous improvement driven entirely by the operators themselves.
Planned Maintenance skipped the planning
Planned Maintenance was supposed to transition your maintenance organization from firefighting to proactive care. The system was meant to rely on condition-based monitoring, predictive maintenance, and a calendar built on actual equipment history rather than generic manufacturer estimates.
What actually happened is that maintenance teams adopted the calendar but skipped the analysis. Preventative maintenance tasks were scheduled based on convenience and available downtime, not on hard failure data. Some machines get serviced far more often than necessary, consuming parts and labour. Other machines run to failure because their PM keeps getting deferred when something more urgent arises — and in a reactive plant, something is always more urgent.
The predictive maintenance tools that were supposed to transform this cycle — vibration analysis, oil analysis, thermography — were purchased, partially deployed, and partially abandoned. The vibration analyst comes in once a month and generates reports that get filed. Critical findings get addressed when someone has time, which is often after the bearing has already failed catastrophically.
This is the reality of a system where nobody was given the dedicated time to learn the predictive tools properly, build the baseline data, and integrate the findings into the maintenance planning cycle.
Skills matrices measured intent instead of capability
The Training and Skills Development pillar was supposed to close the skills gap. The vision was a matrix of competencies for every role, with deliberate training plans to close the gaps. Most plants built a beautiful skills matrix on a whiteboard with green, yellow, and red dots showing who was proficient, developing, or untrained. The matrix was updated quarterly, then annually, then when someone remembered.
The actual transfer of knowledge from experienced to inexperienced operators was left to on-the-job training. This is manufacturing's universal euphemism for following an experienced operator around for two days and trying to pick it up. The experienced operator is also the busiest operator, which means their training consists of showing the new hire where the start button is and telling them to call if something looks weird.
This is the fault of an organization that declared a skills development pillar without allocating dedicated training time, building structured training materials, or measuring training effectiveness. A matrix measures intent. It does not measure capability.
The three structural failures that stall TPM
Strip away the individual pillar failures and you find three structural problems that explain why TPM so rarely delivers. The first is the ownership problem. TPM requires a fundamental shift in who owns equipment reliability. Operators are asked to take ownership of machines they do not fully understand, using skills they were never properly taught, with accountability for metrics they cannot meaningfully influence.
The second is the time problem. TPM is not a project you overlay on top of existing work. It requires reallocation of time. Plants that try to add TPM to an already-full workload get partial execution. The cleaning gets done because it is visible and simple. The analysis, training, and improvement work get shortchanged because there is no protected time carved out for sustained cognitive effort.
The third is the measurement problem. TPM success metrics — OEE, MTBF, MTTR, maintenance cost ratio — are lagging indicators influenced by dozens of factors. When OEE goes up, TPM gets the credit. When it goes down, the market or the aging equipment gets the blame. This measurement ambiguity makes it impossible to hold TPM accountable, which means nobody can definitively say whether their program is working or just consuming effort.
When you cannot measure whether something is working, the path of least resistance is to declare it working and move on.
The OEE trap hides underlying deterioration
A plant launches TPM, declares OEE as its north-star metric, and watches the number climb month after month. Celebrations ensue and bonuses are paid. Two years later, a major breakdown takes down a critical line for a week. The investigation reveals the machine had been deteriorating for months — visible in the vibration data, audible in the bearing noise, detectable in the rising minor-stop frequency. Nobody noticed because the OEE number looked fine.
OEE, as calculated in most plants, is a number that can be optimized in ways that mask underlying problems. Availability can be propped up by deferring preventative maintenance. Performance can be inflated by running faster than the process was designed for, trading quality for speed. Quality can be quietly redefined to exclude certain defect categories that are being addressed separately.
The OEE number that TPM was supposed to improve through genuine equipment excellence instead gets improved through accounting, and the improvement becomes self-validating evidence that TPM is working.
Vanity Metrics vs. Genuine Capability
What teams track
- OEE percentage trending upward month over month
- Checklists completed and signed off by operators daily
- Green dots dominating the skills matrix on the wall
- Number of Kobetsu Kaizen meetings held this quarter
What actually matters
- Operators catching developing faults before they become stoppages
- Maintenance backed by predictive data, not just a calendar
- Demonstrated ability to identify and troubleshoot failure modes
- Ratio of planned maintenance to reactive firefighting improving
Rebuilding equipment ownership from the ground up
If your TPM program has become a cleaning schedule with pillars, the path back starts with honesty about where you actually are. Stop measuring pillar maturity on a five-point scale and start measuring outcomes. You need to know how many of your operators can accurately describe the top three failure modes of their primary machine. You need to know how many of your PMs are based on actual failure data versus default schedules.
I have audited plants where the maintenance budget ratio to planned versus unplanned work had been flat for two years, even as the TPM dashboard glowed green. Re-invest in real training. Not a lunch-and-learn or a video module, but structured, multi-week, hands-on training with assessments that test actual capability. It is expensive, but it is the single highest-return investment in equipment reliability that exists.
Protect time for improvement work. If you cannot afford to pull people off the line for four hours a week to work on the losses costing you twenty hours of downtime, you have decided the downtime is acceptable. Say that out loud, then decide if you mean it.
Stop telling yourselves that TPM is working because the dashboard says so. The dashboard measures activity, and activity is not improvement. Improvement happens when an operator catches a developing fault that would have been a four-hour line stoppage, and the reason they caught it is that someone taught them exactly what to look for and gave them the authority to act on what they found.
