Every engineer who has moved a part from prototype shop to production line has lived through the same disappointment. Tolerances that looked comfortable in the trial run became a weekly firefight once production tooling and production rate took over. The temptation is to blame the operators, the material batch, or bad luck. In two decades of automotive and aerospace launches, I have found the root cause usually sits further upstream: the prototype process was never representative of the production process, and nobody said so loudly enough at the time.
Prototype parts are legitimate engineering artefacts. They answer questions about design, fit and function. What they do poorly, almost by construction, is answer questions about process capability. A capability study tells you what a process does when it runs the way it will actually run — and pilot parts rarely do.
Understanding precisely why pilot data flatters the process is the difference between a smooth launch and eighteen months of containment. The gap is systematic, predictable, and — with discipline — correctable before anyone commits a Cpk number to a customer.
The Fundamental Mismatch in Process Representativeness
The prototype machinist takes extra care, slows the feed, uses a fresh insert, and deburrs by hand between operations. The casting comes out of a low-pressure tool with generous draft angles and hand-fed gates. Nothing about that reflects what the machine will do at cycle time, with a tool at eighty percent of its life, tended by a night shift that has never seen the part before.
Representativeness is not a vague aspiration; it is a checklist question for each process step. Was the part cut at production feeds and speeds? Was it fixtured in the production fixture, or clamped in a vise with soft jaws? Was the material from a production supply route or a leftover bar in the stores? Was the thermal state representative — a part measured cold, straight off a slow machine, says nothing about a part coming off a hot cell at rate and measured ten minutes later.
I have reviewed launch files where every one of those answers was no, and yet the capability declaration submitted to the customer was based on those parts. The declaration was not dishonest — it was answering a different question than anyone realised. The pilot run measured the best the process could ever be, with everything tuned and everyone attentive. Production measures the process as it will be, on an ordinary Tuesday.

Soft Tooling and the Illusion of Dimensional Stability
Soft tooling — aluminium moulds, kirksite dies, 3D-printed inserts, tooling-board patterns — exists to make small quantities cheaply and quickly. Its geometry is often astonishingly good, because it is frequently machined on the same five-axis equipment that will cut the production tool. That is precisely the trap. The part looks right, so people conclude the process is right.
The failure modes of soft tooling reveal themselves only at scale. Aluminium moulds transfer heat differently than hardened steel, so the cooling regime validated on the soft tool is meaningless for production; warp and sink marks will differ. Soft tools wear fast, so dimensional drift that would take months on a P20 steel mould appears within a few hundred shots — but only if you keep running, which pilot runs rarely do. Printed inserts can have anisotropic thermal behaviour that distorts cooling-channel effectiveness in ways nobody modelled.
Fixtures deserve their own warning. A soft fixture cut from tooling board, used gently by one technician, holds the part rigidly and repeatably. Put the same locational scheme into a production fixture with hydraulic clamping at ten-second cycle times, and the clamping force itself deflects the part. The datum scheme may be identical on paper, but the force path is not, and thin-walled components will move. Compare deflection under clamp between pilot and production conditions before trusting any dimensional data from the soft-tool phase.
Capability Projection: Doing the Arithmetic Honestly
Projecting capability from pilot parts is legitimate work — done with eyes open. The correct mental model is that a pilot run gives you an estimate of the process's potential mean and spread, heavily biased toward the favourable end. Your projection must then add, explicitly and in writing, the variance sources the pilot did not contain: tool wear over its full life, material lot variation, multiple operators, ambient temperature swings across seasons, and machine-to-machine differences if the process runs in more than one cell.
Do the projection parametrically. Take the pilot data's spread, then widen it by reasoned factors for each unrepresented source — and document the reasoning so a reviewer can challenge it. A projection that says we measured this and expect production spread to be larger because of X, Y and Z is defensible. A projection that says pilot Cpk was strong, therefore we are fine, is a wish with a spreadsheet attached.
Separate the questions. Pilot parts may validate design intent, but only production-representative runs validate process. If a customer demands capability evidence before production tooling exists, the honest answer is a documented projection with its assumptions stated, not pilot data dressed up as capability. I would rather defend an explicit assumption in a review than explain an implicit one during a containment escalation.
From pilot flattery to honest capability
- 01Pilot runMeasures process potential under best-case conditions
- 02List unrepresented varianceTool wear, lots, operators, seasons, machines
- 03Widen the spread parametricallyDocumented factors a reviewer can challenge
- 04Declare projection, not capabilityAssumptions stated in the launch file
- 05Verification run at rateProduction tooling, production people, production cycle
What to Measure on the Pilot Run Itself
Even a non-representative pilot run yields useful data if you measure the right things. Measure process signatures, not just part dimensions: cutting forces, clamp pressures, injection peak pressure, cycle phase times, spindle load. These signatures travel — they tell you whether the production machine, running the same programme, is behaving like the prototype machine. A dimensional agreement between parts made under very different process signatures is a coincidence, not a transferable result.
Repeat measurements matter more than absolute measurements at this stage. Measure the same parts multiple times, on the fixture, off the fixture, after a settling interval. Thermal growth between a part fresh off the machine and one measured an hour later can consume a meaningful fraction of a tight tolerance on aluminium components; find that out on ten pilot parts, not on a Thursday production batch. Log everything with timestamps so thermal behaviour can be reconstructed.
Track the pilot run's conditions the way you would track production. Operator identity, tool serial numbers, material batch, coolant concentration, ambient temperature at the machine. When the production launch eventually diverges from the pilot results, this record is what lets you diagnose which condition changed. Without it you are arguing memory against memory, and the flattering pilot data wins by default.
Rate, Thermal State and the Night Shift
Rate is the most consistently underestimated representativeness gap. A process proven at four parts an hour behaves differently at forty. Heat soaks into fixtures and machine structure; spindles and moulds reach steady-state temperatures that pilot runs never achieve; chip evacuation, coolant delivery and part handling all change character. Dimensional drift correlated with machine warm-up is one of the most common launch surprises I have had to chase, and it is entirely predictable if anyone asks what thermal steady state means for the process.
Hand in hand with rate comes the human factor. Pilot runs are executed by the most experienced people available — often the toolmaker or the process engineer personally. Production will be run by someone trained last month, working to a standard they did not write, under time pressure the pilot never simulated. Watch a pilot run and note every judgement call, every feel-based adjustment, every instruction to load it a certain way or it marks. Each one is an uncontrolled process variable that capability statistics will eventually invoice you for.
The honest mitigation is a verification run at production rate, on production tooling, with production personnel, before committing a capability number.
Shorten that verification run if you must — but do not skip it and cite the pilot data instead. A half-day at rate will teach you more about thermal drift, fixture repeatability and handling damage than a month of pilot parts measured in a lab.
Pilot conditions versus production reality
What the pilot measured
- Fresh tool, hand-finished, slow feeds
- One expert operator, full attention
- Cold machine, cold part, no soak
- Leftover material, single lot
- Measurement after long settling
What production delivers
- Tool at mid-life wear drift
- Trained operator, time pressure
- Steady-state heat in fixtures and spindle
- Lot variation across suppliers
- Ten minutes off a hot cell
Questions to Ask Before Trusting Pilot Data
When someone presents pilot parts as evidence of production readiness, ask a short, hard sequence. Which machine made these, and is it the production machine? Which tool, at what point in its life? What were the feeds, speeds, clamping forces and cycle times, and how do they compare to the production standard? Who made the parts, and who measured them? How many settling minutes elapsed between processing and measurement? If the answer to any of these is that it was a prototype so it does not matter, you have found your false signal.
Push the projection work the same way. What variance sources did the pilot not contain, and how were they accounted for? What is the plan to verify the projection once production tooling exists, and what triggers re-evaluation — tool life milestones, a change of material supplier, a second machine? A team that has thought about the re-verification plan understands the limits of its own data. A team that has not will discover the limits on your behalf, later, at higher cost.
Finally, treat flattering pilot results with the suspicion you would treat a supplier whose parts are always exactly on nominal. Real processes have texture — drift, slight skew, occasional excursions. Data that looks too good is telling you the conditions were too good. Write that caveat into the launch file explicitly, so that when the production numbers come in lower, nobody spends a fortnight hunting a special cause that is just the ordinary process finally showing its face.
