Most plants maintain equipment based on calendar intervals or run-to-failure reactions. I have audited facilities where maintenance teams spend hours replacing bearings, seals, and cables on a fixed schedule, yet still suffer unexpected downtime. The calendar approach assumes that older components are more likely to fail.
That assumption is a myth. In the 1970s, the US aviation industry faced exploding maintenance costs with no reliability improvements. Engineers questioned their own system and discovered that the vast majority of failures had no correlation to equipment age. This research created Reliability-Centered Maintenance (RCM).
RCM is a systematic maintenance planning framework that asks what a system must do to fulfil its function, and what happens when it fails. Instead of treating all equipment identically, RCM categorises each asset by its actual failure risk and operational consequences. The result is a targeted maintenance strategy for each component. The standard SAE JA1011 defines the seven core questions that govern this methodology.
The seven questions that drive RCM analysis
RCM begins with function, not failure. The first step defines exactly what the equipment must achieve: "Pump fluid X at flow Y ± 5 % at temperature Z." This specificity often reveals that a single component performs multiple functions. A cooling fan ventilaates vapours, generates negative pressure, and signals status via a vibration sensor. Each function needs distinct parameters.
Next, identify functional failures. A functional failure is not simply a broken machine; it is any state where the equipment fails to perform within specified parameters. A pump operating at 15 % reduced flow has functionally failed. A sensor measuring outside its tolerance has failed. You must identify every total and partial way the equipment can lose function.
For each functional failure, determine the failure mode. This is the specific physical or chemical cause: bearing wear, seal corrosion, fatigue fracture, or motor overload. Applying the 5 Why method here is standard practice. The goal is to reach the root physical mechanism behind the loss of function.

Categorising failure consequences
Not all failures carry the same weight. RCM categorises consequences into four distinct types. Hidden failures are not evident during normal operation, such as a backup generator that fails to start during a power cut. You only discover the failure when you need the equipment.
Safety and environmental consequences involve potential injury, loss of life, or ecological damage. These demand absolute priority. Operational consequences interrupt production, reduce quality, or delay deliveries. These are directly measurable in financial terms. Non-operational consequences only incur repair costs without affecting safety or output.
Understanding these categories forces maintenance teams to prioritise based on business impact rather than habit. A failure that halts a 1,200-vehicle-per-day automotive line requires a completely different strategy than a failure on an auxiliary pump. The consequence dictates the urgency and type of intervention required.
The age-reliability myth and failure patterns
The most significant discovery in early RCM research was that component failures follow six different patterns. Only two of these patterns relate to equipment age. The classic "bathtub curve" (Pattern A) with high initial failures, constant reliability, and end-of-life wear-out applies to roughly 4 % of components.
Pattern B, showing linearly increasing failure probability with age, applies to a mere 2 % of components. The remaining patterns show random failure distributions independent of age. Most critically, Pattern F exhibits high initial failure rates that gradually decrease over time. This pattern alone accounts for approximately 68 % of all component failures.
Component failure pattern distribution
Read those numbers carefully. If 68 % of failures are random and infant-mortality related, calendar-based replacement actively harms your operation. You remove a functional component and introduce the risk of installation error or manufacturing defect from the new part. Scheduled maintenance becomes a primary cause of failure.
Selecting the right maintenance strategy
Based on the answers to the previous questions, RCM assigns one of five specific maintenance strategies. The decision logic follows a strict hierarchy, starting with safety. If a failure threatens personnel, you must implement the most rigorous proactive strategy available before moving down to operational and economic considerations.
The objective of RCM is not to find out how often equipment fails, but to find out what must be done to prevent the consequences of failure.John Moubray
Predictive maintenance (condition-based) monitors indicators like vibration, temperature, and oil analysis, intervening only when deterioration appears. Preventive maintenance (scheduled restoration) applies only to the 6 % of components with age-related failure patterns. Failure-finding tasks test hidden functions like fire detectors and safety valves.
Run-to-failure is a deliberate choice for low-consequence components where prevention costs more than repair. Redesign is the final option when no maintenance strategy can reduce risk to an acceptable level. The final question asks whether the chosen strategy is technically and economically viable.
Targeted RCM implementation in practice
I have reviewed RCM rollouts at body-weld lines running 47 robots with 564 critical components. Before RCM, maintenance teams replaced components on a fixed schedule. Unexpected breakdowns still occurred because calendar maintenance cannot predict random failures. After implementing RCM and processing all 564 components through the seven questions, the results shifted dramatically.
| Strategy | Allocation | Rationale |
|---|---|---|
| Run-to-Failure | 73 % | Low failure consequences; fast repair times. |
| Predictive | 11 % | Vibration and thermal monitoring replaced time-based swaps. |
| Failure-Finding | 8 % | Regular testing of hidden safety and backup systems. |
| Calendar Replacement | 4 % | Only components with age-correlated failure patterns. |
| Redesign | 4 % | Material upgrades or redundant architecture required. |
By shifting 73 % of components to run-to-failure, the maintenance team stopped wasting labour on unnecessary interventions. Maintenance costs dropped by 38 %. Unexpected failures fell by 62 %. Line availability increased from 91.3 % to 96.8 %, and the maintenance team finally had time for genuine data analysis and prevention.
Avoiding common RCM implementation errors
The most frequent error is attempting RCM on every asset. Not every device warrants this level of analysis. An office printer belongs in the run-to-failure category. A critical compressor in a food-grade facility warrants full RCM. Select targets based on risk and operational impact.
The second error is conducting RCM from a desk. RCM belongs on the shop floor at the machine. An operator who runs a machine for eight hours daily understands failure modes better than any engineer reviewing a spreadsheet. You must build a team comprising maintenance technicians, operators, and process and quality engineers.
RCM implementation sequence
- 011. Select systemChoose one critical asset, not the entire plant. Assemble a cross-functional team.
- 022. Analyse componentsWork through the seven RCM questions using FMEA data and maintenance history.
- 033. Assign strategiesApply the decision logic tree, prioritising safety and operational consequences.
- 044. Implement planUpdate the maintenance schedule, configure condition monitoring, and train staff.
- 055. Review continuouslyMonitor strategy effectiveness and rework analysis when new failure modes appear.
The third error is ignoring hidden failures. Standby systems are silent until they fail under load. If you do not perform regular failure-finding tasks on backup generators and safety valves, you will discover they have been inoperative for months during an actual emergency. The final error is forgetting the people. If maintenance teams do not understand why you are changing the strategy, they will not execute it.
Integrating RCM with Industry 4.0
IoT sensors, machine learning algorithms, and digital twins give RCM new analytical power. Real-time condition monitoring feeds predictive models that identify degradation patterns long before functional failure occurs. Machine learning can spot complex failure signatures across multiple sensor inputs that human analysts will miss.
However, technology does not replace RCM methodology. IoT sensors are tools; RCM is the framework that dictates where and why you deploy them. Installing thousands of sensors without understanding the seven questions produces expensive data graveyards. The methodology must drive the technology, never the reverse.
Reliability-Centered Maintenance shifts maintenance from a reactive cost centre to a strategic operational function. It replaces calendar-driven guesswork with evidence-based intervention. Start with one critical system, ask the seven questions, and build a maintenance programme that addresses actual failure consequences.
