Walk into any factory that has been running for more than fifteen years and you will find a maintenance department in a corner of the building. It is staffed by people who know the equipment better than the engineers who designed it. They carry grease guns and multimeters. And every Monday morning, the production manager walks over to ask when Machine 7 will be running again.

This is the reality of maintenance in most manufacturing operations: a reactive function measured by how fast it responds to breakdowns. It is rewarded for heroics rather than prevention, and it is treated as a cost centre whose budget gets cut the moment margins tighten. Total Productive Maintenance (TPM) was supposed to change all of that.

Formalized through the Japan Institute of Plant Maintenance, TPM was built on a radical premise: operators should own the daily health of their equipment. Not mechanics, not contractors. The people who run the machines should perform the fundamental care that prevents the vast majority of equipment failures. The goal was measurable: zero unplanned downtime, zero accidents, and zero defects caused by equipment degradation.

What TPM Actually Promised Versus the Floor Reality

TPM outlined specific practices designed to transfer responsibility and predictability to the shop floor. Operators were tasked with autonomous maintenance: cleaning, inspecting, lubricating, and tightening. These four activities prevent the vast majority of mechanical failures. Maintenance technicians were then freed to handle complex tasks like precision alignments, overhauls, and predictive maintenance.

The methodology yielded dramatic results where implemented faithfully. Equipment availability jumped from 70% to 90% in documented cases. Unplanned downtime dropped significantly. The 5S foundation—Sort, Set in Order, Shine, Standardize, Sustain—provided the visual order necessary for operators to spot abnormalities instantly.

Then the rest of the world heard about it, and the dilution began. A consultant arrives, delivers a two-day training module, and shows management before-and-after photos of painted and labelled machinery. A kickoff event is scheduled, banners are printed, and operators spend a Friday afternoon wiping down machines that have accumulated years of grime.

Quality decisions are made at the process, not in the report that describes it afterwards.
Quality decisions are made at the process, not in the report that describes it afterwards.

The Six Losses Nobody Actually Tracks

TPM defines six big losses that reduce Overall Equipment Effectiveness (OEE). In my experience auditing plants across automotive and heavy industry, most factories track these losses so poorly that the resulting data is functionally useless. Operators log 'machine stopped' without a failure code. Mechanics log 'repaired' without a root cause.

Consider idling and minor stoppages. A machine that stops for thirty seconds to clear a jam, then restarts, does not get logged as downtime. Over an eight-hour shift, these micro-stops accumulate into forty minutes of lost production. Nobody sees it because the counter restarts and the shift report shows 'running.' The only way to catch this loss is to monitor cycle time integrity.

Reduced speed is equally invisible. Every machine has a designed cycle rate, and every machine runs slower than that rate. Sometimes intentionally, because the equipment cannot hold tolerance at full speed. OEE calculations that use a 'best demonstrated rate' or current standard show inflated numbers, hiding a 15 to 25 percentage point performance loss.

Loss Category Standard Reporting Failure Operational Consequence
Equipment Failure Logged without failure codes or root cause Preventative action impossible; reliability metrics unusable
Setup and Adjustments Ramp-up scrap excluded from changeover time SMED baselines are fictional; improvement targets missed
Idling and Minor Stops Micro-stops under thirty seconds unrecorded Forty minutes of hidden shift loss; cycle integrity ignored
Reduced Speed Nameplate speed ignored for current standard Performance loss invisible on management dashboards
How the six big losses are typically mistreated in standard factory reporting systems.

Autonomous Maintenance: The Hardest Easy Thing in Manufacturing

The concept of autonomous maintenance is simple: train operators to perform routine equipment care. The execution is enormously difficult, and the reasons have nothing to do with technical complexity. Operators resist it because they see it as extra work without extra pay. The current incentive system pays them for parts produced, not for equipment maintained.

Maintenance technicians resist it because they see it as a threat to their expertise. Supervisors resist it because ten minutes spent on inspection is ten minutes not spent running production. Breaking through this resistance requires changing the measurement system, repositioning maintenance as a partner, and training operators properly.

If operators do not own the daily health of the machine, your TPM programme is just a lubrication schedule waiting to fail.

If you want operators to maintain equipment, you must revise standard work to include maintenance tasks. You must adjust takt time calculations and change the shift scorecard. Operators need hands-on practice with a maintenance mentor until the behaviours become habit. A one-hour briefing and a laminated card will not sustain the discipline.

Predictive Maintenance and the Technology Trap

Every TPM conference features presentations on predictive maintenance. Vibration analysis, oil analysis, thermography, and motor current signature analysis are genuine technologies with sound physics. But detecting an incipient failure is useless if the organization cannot act on the prediction. The gap between detection and action is where every PdM program dies.

The Predictive Maintenance Execution Gap

  1. 01DetectionVibration analysis identifies a bearing trending toward failure on the main compressor.
  2. 02Work Order GenerationTechnician writes the repair order based on the condition data.
  3. 03Production BacklogOrder sits for weeks because production will not release the equipment.
  4. 04Forced FailureBearing fails, compressor goes down, line stops for sixteen hours.
Technology detects the failure, but organizational friction guarantees the breakdown occurs anyway.

Closing that execution gap requires a maintenance planning function with the authority to schedule equipment outages based on condition data. It requires a production planning function that builds flexibility into the schedule. It requires a spare parts strategy that ensures components are available when the data says they are needed. None of that is technology; all of it is organizational.

The OEE Manipulation Problem

OEE is supposed to be the primary metric of TPM: Availability multiplied by Performance, multiplied by Quality. The formula is simple, but OEE has become an end rather than a means. I have audited plants where the OEE number on the dashboard never drops below 85%, yet walking the floor reveals a completely different reality.

Equipment is down for changeovers that somehow do not get counted. Scrap rates include only parts physically thrown away, not parts reworked. Cycle times are based on a 'current standard' that has been adjusted downward over the years until it bears no resemblance to the equipment's actual capability. The number is high because the plant manager's bonus depends on it.

Real OEE—calculated against nameplate speed, true availability, and first-pass yield—is almost always lower than what management believes. A plant reporting 88% OEE might be running at 55% when measured honestly. That gap is not a reporting problem. It is a leadership problem that hides the true cost of equipment instability.

Building a TPM Programme That Actually Lasts

I have implemented and rescued enough TPM programmes across aerospace and automotive to identify the specific practices that separate sustained success from decay. Start with a single pilot line, not a whole-plant rollout. Focused pilots allow you to refine the process, build internal expertise, and generate credible results before scaling the support structure.

Make the first measurable goal equipment cleanliness, not OEE. Clean equipment reveals leaks, cracks, and wear that dirty equipment hides. Tag every defect with a numbered red tag and ensure someone follows up within 48 hours. If tags sit unresolved, the system dies immediately. If they are addressed, trust builds quickly.

Every TPM failure I have investigated traces back to the same root cause: leadership treated it as a programme rather than a philosophy. Programmes have budgets and end dates. Philosophies become part of how the organization thinks, reflected in daily priorities and capital allocation. You build reliability one machine, one operator, one standard at a time.