Mean Time Between Failures is the expected value of the time between failures of a repairable system operating in its useful-life period. It says nothing about the lifetime of an individual unit, nothing about when a unit wears out, and nothing about how failures are distributed across time. Yet it is quoted, contracted and litigated as though it were a lifespan. Across two decades in automotive and aerospace quality, I have seen this single confusion cost more money in negotiations than any other reliability error.
The classic illustration is the component with an MTBF of one million hours. Marketing departments love that figure: roughly 114 years, they imply, this thing will outlive your grandchildren. In reality, if the failure process is exponential with that rate, over a tenth of the population fails within the first eight years of service, and the probability of any single unit surviving a hundred years is effectively nil. The number is a rate, not a duration, and the difference is not academic.
The error compounds because MTBF is routinely quoted for non-repairable items — bearings, capacitors, seals — where the correct term is MTTF, mean time to failure. The distinction sounds pedantic until you realise the arithmetic can diverge sharply when renewal effects, burn-in behaviour or wear-out mechanisms are in play. A bearing replaced after seizure does not return the assembly to a statistically identical state; it returns it to a partially degraded one. This article sets out what MTBF can legitimately do, and what to ask for when it cannot.
The constant hazard assumption beneath every MTBF figure
Beneath almost every MTBF figure you will find, quietly assumed, a constant hazard rate. That is the mathematical precondition for the exponential distribution, where MTBF is simply the inverse of the failure rate and — critically — a unit that has survived ten years is statistically indistinguishable from one fresh out of the box. The mathematics is elegant. For most engineered hardware, the assumption is false, and decisions built on it inherit that falsehood.
Real hardware follows the bathtub curve: infant mortality from assembly defects, solder voids, contamination and marginal tolerances; a useful-life plateau; then wear-out as fatigue cracks propagate, insulation degrades, lubricant oxidises and elastomers harden. Naive MTBF calculations capture only the plateau. They are silent on both tails, which is precisely where most field failures actually live. The number survives because the plateau is the easiest regime to model, not the most important one.
Consider a gearbox. Its early failures come from misalignment, hard metal inclusions and poor run-in. Its late failures come from pitting, tooth bending fatigue and seal wear. Neither regime has a constant hazard rate. Pool three years of field data, blend both regimes into a single number and hand it to procurement as MTBF, and you have produced a figure that describes no physical mechanism at all — an average across two different failure populations.
What to check instead: plot the empirical hazard rate over time, or fit a Weibull distribution and read the shape parameter. A shape parameter near one means the exponential model is defensible. Well above one puts you in wear-out territory, where scheduled replacement or redesign — not MTBF arithmetic — is the correct response. Below one, you have infant mortality, and the fix belongs in process control and burn-in screening.

Mission time: the number MTBF hides
An MTBF figure on its own tells you almost nothing about survival over an actual mission. What matters is reliability at a stated time, and that requires both the distribution and the mission duration. Under the exponential assumption, reliability at time t is the exponential of minus t over MTBF. Abandon the exponential — as you must for wear-out mechanisms — and MTBF becomes nearly useless as a predictor of mission success.
Take an aerospace actuator with a Weibull wear-out distribution. Its mean life might comfortably exceed the system requirement, yet the point at which the hazard rate crosses the acceptable threshold — a B-life, the time at which a specified fraction of the population has failed — could sit squarely inside the intended service interval. Mean values smear out exactly the tail behaviour that determines whether a component stays in service or comes out prematurely.
I insist that reliability requirements be written as reliability at a mission time, not as MTBF. "Ninety-eight per cent probability of completing a 2,000-hour mission without a function-critical failure" is testable, allocatable and meaningful. "MTBF shall exceed 4,000 hours" is none of those things: it hides the distribution, the confidence level and the failure definition. When a supplier quotes MTBF without a mission time, ask what figure they expect at your actual duty cycle, and watch the conversation become uncomfortable.
Duty cycle deserves its own scrutiny. Handbook MTBF predictions are almost always quoted at reference conditions — nominal temperature, nominal load, benign environment. Real products run hot, vibrated, cycled and contaminated. Temperature acceleration alone can shift electronic failure rates dramatically under an Arrhenius model. A number derived at forty degrees ambient has no authority at ninety-five.
Where MTBF figures come from
Nearly every MTBF figure in a supplier datasheet originates from a handful of prediction handbooks — Mil-HDBK-217 and its derivatives, Telcordia, or the newer 217Plus and FIDES methodologies. Each applies component-count or stress-analysis formulas to produce a predicted failure rate. These predictions have historically disagreed with field data by wide margins, in both directions, and no competent reliability engineer treats them as absolute truth.
A prediction is a comparative tool. Used to rank design alternatives — this topology versus that one, derated versus stressed components — it has genuine value, because the biases are roughly consistent across comparisons. Used as an absolute promise of field performance, it invites disaster. I have sat through warranty negotiations where a handbook number, produced in an afternoon by a junior engineer, was treated as gospel against three years of actual return data. The field data always wins.
Then there is demonstrated MTBF, derived from life testing. Here the trap is censoring. Run twenty units for a fixed duration, note the failures and compute a mean, and you have produced a heavily censored estimate whose validity depends entirely on the assumed distribution. Fail to log the time base accurately, ignore suspended units, or mix test phases with different stresses, and the resulting figure is fiction with decimal places.
Check three things whenever a supplier presents an MTBF: the source method (handbook prediction, test demonstration or field data), the environmental and stress assumptions, and the confidence bounds. A demonstrated MTBF without a confidence interval is an estimate pretending to be a measurement. At low sample counts the interval is embarrassingly wide, and honest suppliers will say so before you ask.
An average over two different failure populations is not a conservative estimate; it is a number that describes no physical mechanism at all.
Interrogating a quoted MTBF
- 01Identify the sourceHandbook prediction, test demonstration, or field data — each carries different authority.
- 02Check the distributionWeibull shape parameter: near 1 exponential holds, above 1 wear-out, below 1 infant mortality.
- 03Confirm the environmentReference temperature and load versus your actual duty cycle and contamination.
- 04Demand the confidence boundsNo interval means an estimate posing as a measurement.
- 05Convert to mission reliabilityReliability at your stated mission time, at your confidence level.
Contract language: where the damage is done
Most MTBF disputes I have arbitrated trace back to contract wording, not engineering. Two failures of drafting appear repeatedly. The first is the undefined failure: the contract specifies an MTBF but never says what counts as a failure. Does a nuisance fault code count? A degraded mode? A cosmetic defect on a non-critical path? Without a failure definition tied to severity classification, both parties discover their disagreement twelve months into the programme, when it is expensive.
The second is the missing measurement basis. MTBF verified how — by a timed test at what confidence level, by field data over what observation window, by prediction using which handbook revision? A contract that says only "MTBF shall be no less than X hours" has specified a number with no verification path, and every exit from that ambiguity will be argued commercially rather than technically. Ambiguity in reliability clauses is always resolved by leverage, never by physics.
Liquidated damages and warranty provisions amplify the problem. When financial remedies attach to an MTBF shortfall, the incentive to massage the failure definition, the observation window or the data exclusions becomes intense. I have seen "failures" quietly reclassified as "operational events" once penalties came into view. The remedy is unglamorous: write the failure definition, severity classes, data sources, exclusion rules and confidence level into the contract itself — in an annex if necessary — before signature.
Better still, move the contractual metric. Specify reliability at mission time with a stated confidence, define B-life requirements for wear-out components, and tie preventive replacement intervals to the demonstrated hazard curve. These are harder to draft but far harder to game. A supplier who cannot sign up to a B-life has told you something important about how well they know their own product.
What to ask for instead
My working rule: MTBF may appear in analysis, but never as a standalone requirement. Internally, I want the full picture — a Weibull analysis of field returns with the shape parameter reported, hazard curves plotted against service time, and separate treatment of the three failure regimes. Each regime has a different owner, and merging them into one number guarantees that none of the three owners acts on it.
One number versus three regimes
What teams do
- Quote a single blended MTBF for all failure regimes
- Write contracts around MTBF with no failure definition
- Accept handbook predictions as field performance
- Treat a mean as a guarantee of individual unit life
What works
- Infant mortality: attack through process control and burn-in screening
- Useful life: MTBF-style rate is legitimate here
- Wear-out: B-lives drive replacement intervals and design margin
- Requirements written as reliability at mission time, with confidence
For critical components, I ask suppliers for B-lives — the operating hours at which small specified fractions of the population are expected to have failed — because those numbers drive maintenance intervals, spares provisioning and life-limited part decisions. B-ten life for a bearing, B-one for a safety-critical element, whatever the severity justifies. These figures force the supplier to know their distribution, not just their average. A supplier who cannot produce them has not characterised their product.
When testing, run time-over-stress designs that actually reveal the hazard curve, not just pass-fail at a single duration. Log suspension times scrupulously. Report confidence intervals. And when a number arrives, ask the question that matters: what does this predict about survival at my mission time, under my environment, at my confidence level? If the person across the table cannot answer it, the MTBF they quoted is decoration, and you should negotiate accordingly.
The working standard, condensed
MTBF has a legitimate narrow role: describing the failure rate of a repairable system in its useful-life plateau, at stated conditions, with stated confidence. Outside that role it misleads — as a lifespan claim, as a wear-out predictor, as a contractual metric, or as a figure transferred between environments without acceleration correction. Every one of those misuses has a documented failure trail behind it.
The practical standard is therefore short. Know the distribution, not just the mean: a Weibull shape parameter changes everything downstream. Write requirements as reliability at mission time with a confidence level. Use B-lives for wear-out components and tie maintenance intervals to the demonstrated hazard curve. Demand the source method and the bounds behind every quoted figure.
None of this requires new mathematics — Weibull analysis, censored-data methods and B-life definitions have been standard reliability engineering for decades. What it requires is the discipline to refuse the single-number shortcut in the places where it does the most damage: datasheets, requirements documents and contracts. The engineering to replace it exists; only the habit of asking for it has to be built.
