Every engineering student learns the bathtub curve. The elegant U-shape plots infant mortality on the left slope, a flat useful-life region across the middle, and wear-out failures climbing on the right. It is tidy, intuitive, and for a significant fraction of industrial equipment, fundamentally misleading. Maintenance programmes built on the assumption that all assets follow this pattern waste money preserving components that fail randomly.

These programmes neglect the failure modes that demand something other than time-based replacement. Reliability-Centred Maintenance (RCM) emerged in the aviation industry during the 1960s specifically because the bathtub curve kept failing to predict when complex systems actually broke. Yet in manufacturing plants today, RCM remains widely misunderstood, often reduced to a mere FMEA exercise or a static spreadsheet of failure modes.

This misunderstanding prevents organizations from translating analysis into a living maintenance strategy. The gap between theoretical reliability and actual shop-floor maintenance persists because engineers apply a simple model to complex systems. Manufacturing leaders must abandon the one-size-fits-all curve and build maintenance programmes that reflect the real failure behaviour of their specific equipment.

The Origin and Limits of Time-Based Prevention

The bathtub concept originated in reliability engineering for consumer electronics and simple mechanical components. Incandescent lamps, ball bearings, and certain electronic valves genuinely exhibit this failure pattern. A population of lightbulbs shows early failures from manufacturing defects, a long period of constant failure, and eventually a wear-out region as filaments degrade.

This model drove industrial maintenance thinking during the mid-20th century. If components wear out along a predictable curve, the logic dictated, replacing them just before the wear-out region begins should prevent failures. Time-based preventive maintenance (PM) — scheduled overhauls, calendar replacements, and running-hour interventions — became the dominant global strategy for asset management.

The fatal flaw is that this logic depends on one critical assumption: that the failure pattern actually follows the bathtub curve. In 1960, a task force at United Airlines discovered that scheduled replacement was effective for only a small fraction of components. For the majority of aircraft parts, replacing a healthy component on a fixed schedule either wasted remaining useful life or introduced installation errors that caused infant mortality.

The investigation, later formalised in the landmark Nowlan and Heap report, proved that only 11% of aircraft components exhibited a wear-out pattern consistent with the bathtub curve. Approximately 68% showed a constant or slightly increasing failure rate with no clear wear-out point. Failures were essentially random, rendering calendar-based replacements useless or actively harmful.

Six Failure Patterns That Replace One Curve

RCM classifies component failure behaviour into six distinct patterns. Understanding these patterns is the absolute foundation of any defensible maintenance strategy. Pattern A is the classical bathtub curve, featuring high infant mortality, a flat useful-life region, and a pronounced wear-out zone. This pattern is typical of simple mechanical items like seals, filters, and brake pads where scheduled replacement works perfectly.

Pattern B shows steady wear-out, where failure probability increases gradually from installation with no infant mortality spike. This is found in cutting tools and tyres. Pattern C shows a slow initial rise followed by rapid acceleration. Time-based intervention can work for both if the transition points are historically consistent and well-documented by the maintenance team.

Where the calculation meets the floor: the gap between planned availability and the shift people actually work determines true OEE.
Where the calculation meets the floor: the gap between planned availability and the shift people actually work determines true OEE.

Patterns D, E, and F dominate complex industrial equipment, and they share one critical characteristic: there is no wear-out zone where scheduled replacement reduces failure probability. Pattern D levels off at a plateau. Pattern E is a constant random failure rate throughout service life. Pattern F drops to a low constant rate after initial infant mortality. Applying calendar-based PM to these components wastes resources.

Nowlan-Heap Failure Distribution Findings

11%Wear-out (Pattern A)Components matching the bathtub curve where scheduled replacement works.
68%Random (Patterns D & E)Constant or plateau failure rates; age-based replacement has no effect.
6%Infant Mortality (Pattern F)Fail early or run forever; replacement introduces new failure risk.
15%Other CombinedVarious patterns combining gradual wear and random spikes.
Distribution of aircraft component failure patterns that shattered the time-based replacement assumption.

The Seven Core Questions of RCM

RCM is not a software package or a one-time workshop. It is a structured decision process built around seven specific questions, applied systematically to each significant system within a facility. Before discussing failures, RCM requires explicit definition of what the system is supposed to do and at what performance standard it must operate.

A pump's function is not simply 'pump fluid'. It must 'transfer 200 litres per minute of coolant at 4 bar pressure to the machining centre, with no leakage exceeding 5 drops per minute'. Performance standards are not aspirational targets; they are the operating context that determines whether the equipment is actually doing its job.

Once functions are locked down, RCM demands specificity in failure mode analysis. 'Pump failure' is entirely inadequate. 'Cavitation damage to impeller vanes due to low suction pressure during cold-start conditions' is a failure mode. The maintenance task must address the specific physical mechanism of degradation, not a generalized category that masks the true root cause of equipment downtime.

The RCM Decision Workflow

  1. 01Define FunctionsEstablish exact performance standards (flow, pressure, tolerance) required for operation.
  2. 02Identify Functional FailuresCategorise severity: total loss of function, partial loss, or degradation beyond limits.
  3. 03Determine Failure Modes & EffectsSpecify exact physical mechanisms and what the operator experiences when it happens.
  4. 04Classify ConsequencesSort by safety, environmental, operational, or non-operational impact to prioritise action.
  5. 05Assign Proactive TasksSelect scheduled restoration, condition monitoring, or run-to-failure based on data.
How maintenance strategies are logically derived from equipment functions and failure consequences.

Why Manufacturing RCM Implementations Stall

Despite its logical rigour, RCM has a mixed track record in manufacturing plants. The primary failure point is analyst-driven analysis conducted without operator input. Reliability engineers build failure mode lists at desks using OEM manuals, entirely missing the failures that actually occur on the plant floor. Operators know which sensors drift and which valves stick after weekend shutdowns.

Documentation fatigue is another serial killer of RCM. A full analysis for a mid-sized manufacturing line identifies hundreds of failure modes, each requiring consequence classification and task selection. Plants attempt to analyse everything at once, produce a thousand-row spreadsheet, and subsequently never implement it because the sheer volume becomes completely unmanageable for maintenance crews.

RCM without ground-level operator knowledge produces strategies that address theoretical failures while real ones go unmanaged.

Many plants also confuse RCM with PM optimisation. PM optimisation reviews existing scheduled tasks and adjusts intervals. It starts from existing tasks rather than from system functions. It cannot identify a failure mode that has no current task addressing it, nor can it determine that a time-based task should be entirely replaced by a condition-based approach.

The P-F Interval and Condition Monitoring

The P-F interval determines whether condition-based maintenance actually works. The potential failure condition (P) is what monitoring detects — an increase in vibration energy, a temperature spike, or a drop in oil viscosity. Functional failure (F) occurs when the bearing seizes or the pump stops producing pressure. The P-F interval is the operational window between these two points.

If a rolling element bearing on a fan motor has a P-F interval of 12 weeks, the vibration survey must be performed more frequently than every 12 weeks. A monthly survey provides three detection opportunities. A quarterly survey provides only one, and if that single measurement catches the early stages, the signal may be too weak to confidently identify, leaving no time to plan maintenance.

Many plants collect condition monitoring data on schedules that do not reflect actual P-F intervals. Monthly vibration routes on equipment with a 4-week P-F interval are useless. By the time data shows a problem, the maintenance team has zero time to execute a fix before catastrophic failure. The programme appears broken, but the real issue is an incompatible survey frequency.

Integrating Maintenance with Quality Systems

For plants operating under IATF 16949 or ISO 9001, RCM must not exist as a parallel programme. Maintenance strategy directly dictates process capability, equipment availability, and product quality. Control plans should explicitly reference maintenance triggers. When a control plan identifies equipment-related dimensions, the maintenance task must be traceable.

A dimension that drifts due to fixture wear needs a maintenance response tied to condition monitoring or scheduled restoration, not merely more frequent quality inspection. Inspecting a drifting process more often does not stop the drift. Across two decades in automotive and aerospace, I have audited plants where nonconformance investigations completely ignored maintenance history.

When a quality defect traces to equipment degradation, the corrective action must ask whether the failure mode was identified in the RCM analysis. If it was identified but the task did not prevent it, the task interval requires immediate revision. Layered process audits should expand to verify that condition monitoring tasks were actually performed and that safety-related work orders are not accumulating.

Building the Business Case from the Ground Up

Implementing a rigorous maintenance strategy requires real investment in analyst time, operator involvement, training, and condition monitoring technology. The business case rests on measurable returns. Plants that implement RCM correctly typically see significant reductions in unplanned downtime by catching failures earlier and prioritising resources on equipment that actually matters.

RCM redirects maintenance effort from blanket scheduled tasks to targeted interventions. Plants frequently find that a substantial percentage of existing PM tasks add no value. They address failure modes that do not exhibit wear-out patterns, or they duplicate coverage already provided by condition monitoring. Eliminating these useless tasks frees maintenance hours for higher-value predictive work.

Quality improvement is the final pillar. Equipment degradation causes dimensional drift, surface defects, and assembly errors. The path forward does not require a massive consulting engagement. Select one critical system, conduct a focused RCM analysis using ground-level operator knowledge, determine the P-F intervals, and measure the results over six months.