The Experiment That Never
Ends
You’ve seen it. The team gathers in a conference room. A quality
engineer — probably the one with the Green Belt certification and the
Minitab license — presents a 32-run full factorial design with four
factors at three levels each. The slide says “Optimization Study for
Injection Molding Parameters.” Nobody in the room fully understands what
“interaction effects” means, but everyone nods. The production manager
asks when the line will be available. The engineer says “sixteen hours
of run time, plus setup.” The manager winces. The quality director says
“This is important work.” The meeting ends. The experiment begins.
Three weeks later, nobody has analyzed the data. Two of the runs were
done on a different machine because the primary was behind on
production. One operator swapped the settings on runs 14 and 15 because
he “thought there was a typo.” The engineer who designed the experiment
took a new job at a competitor. The spreadsheet sits on a shared drive
labeled “DOE_Final_v3_MACTUAL.xlsx.” Six months from now, someone will
open it during an audit, see that a DOE was “performed,” and check the
box.
This is what Design of Experiments looks like in most manufacturing
organizations. Not the textbook version. Not the ASQ certification exam
version. The real version — where statistical rigor meets operational
reality and the result is usually an expensive, incomprehensible
exercise that nobody translates into actual process improvements.
What DOE Is Actually
Supposed to Do
Let’s step back. Design of Experiments, at its core, is a systematic
method for determining cause and effect relationships between process
variables (factors) and outcomes (responses). Instead of changing one
variable at a time — which is slow, expensive, and fundamentally unable
to detect interactions — DOE lets you change multiple variables
simultaneously in a structured way, extract maximum information from
minimum runs, and build a mathematical model of how your process
actually behaves.
When done right, DOE tells you:
- Which factors actually matter (and which don’t)
- How factors interact (because real processes are not linear and
isolated) - What the optimal settings are (not just “better” — provably
optimal) - How robust those settings are (what happens when noise creeps
in)
This is powerful stuff. A well-designed experiment can replace months
of trial-and-error with a week of structured testing. It can reveal that
the variable everyone thought was critical (temperature) is actually
irrelevant, while the variable everyone ignored (humidity) is driving
60% of the variation. It can turn “we think this works” into “we know
this works, and here’s the model that proves it.”
The problem isn’t the method. The problem is everything around
it.
Where DOE Goes
Wrong: The Seven Failure Modes
1. The Kitchen Sink Design
The engineer wants to study everything. Seven factors. Three levels
each. Full factorial. That’s 2,187 runs. Nobody pushes back because
nobody understands the math well enough to say “that’s insane.” A
fractional factorial or a screening design could answer the same
questions in 16 to 32 runs, but the engineer doesn’t think about
efficiency — only completeness. The experiment collapses under its own
weight before run number 20.
The truth: Good DOE is about restraint. Start with a
screening design (Plackett-Burman or fractional factorial with 12-16
runs) to identify the vital few factors. Then run a smaller response
surface design on just those 2-3 factors. Two efficient experiments will
always beat one massive one.
2. The Wrong Response
The team spends three weeks optimizing tensile strength. They build a
beautiful model, find the optimal settings, and present the results.
Then someone asks: “Did the customer actually specify tensile strength?”
They didn’t. They specified elongation at break. The DOE optimized the
wrong output because the engineer chose the response he knew how to
measure, not the response the customer cared about.
The truth: Response selection is the most critical
decision in DOE, and it’s usually the one that gets the least thought.
Before designing a single run, ask: What does the customer actually
need? What failure mode are we trying to eliminate? What would change
our actions if we knew the answer?
3. The Uncontrolled Noise
You’re running an experiment on a production line. Shift A does runs
1-8. Shift B does runs 9-16. Shift A’s operator follows the protocol
exactly. Shift B’s operator “improves” the setup because he’s been doing
this for 25 years and knows better. The ambient temperature is 8°C
higher in the afternoon. The raw material lot changed between runs 10
and 11. Nobody recorded any of this.
When the analysis shows no significant effects, the team concludes
that “these factors don’t matter.” In reality, the noise overwhelmed the
signal. The factors might matter enormously — but the experiment was so
contaminated by uncontrolled variation that the signal was buried.
The truth: DOE requires either controlling noise
variables or accounting for them through blocking, randomization, and
replication. If you can’t control the noise, you must design the
experiment to be robust against it. Running a DOE in an uncontrolled
environment doesn’t produce insights — it produces expensive random
numbers.
4. The Analysis Paralysis
The data was collected correctly. The engineer has Minitab. The
Pareto chart of effects is right there. Three factors are significant.
Two interactions are meaningful. The path forward is clear.
And then… nothing.
The results sit in a PowerPoint for three months. The process
engineer says “we should validate these findings with a confirmation
run.” That confirmation run is never scheduled. The optimal settings are
never implemented because “we need to do more testing first.” The
organization has spent the money, done the work, generated the knowledge
— and then filed it away because acting on findings feels riskier than
acting on intuition.
The truth: A DOE that doesn’t lead to process
changes is waste. The confirmation run should be planned before the
experiment begins. The implementation timeline should be written into
the project charter. If you’re not prepared to act on the results, don’t
run the experiment.
5. The Black Box
The statistical analysis is done. The p-values are less than 0.05.
The R-squared is 0.94. The engineer presents the mathematical model to
the team. Nobody on the production floor understands it. The model says
“maximize desirability at X1=145, X2=0.8, X3=12.” The operators set
X1=140 because “close enough,” ignore X2 entirely, and set X3=10 because
“that’s what we’ve always run.”
The truth: A model that can’t be explained in plain
language won’t be followed. The best DOE practitioners translate
statistical findings into visual, intuitive, operational terms: “When
mold temperature increases by 10 degrees, shrinkage doubles. Here’s the
chart. Here’s where we need to be. Here’s why.” If the team can’t
explain the finding to a machine operator in two minutes, the finding
will never reach the floor.
6. The One-Shot Wonder
The DOE was a success. The optimal settings were found, validated,
and implemented. Yield improved by 12%. The team celebrated. The report
was filed.
And then the process drifted. The optimal settings were optimal for
the conditions that existed during the experiment — the material lot,
the machine condition, the ambient environment, the tooling age. Three
months later, those conditions no longer exist. But the settings are
locked in the work instruction because “the DOE proved this is optimal.”
Nobody runs a follow-up experiment. Nobody questions whether the optimum
has shifted.
The truth: DOE is a snapshot, not an eternal truth.
Processes evolve. Materials change. Equipment ages. The optimal settings
from last quarter’s experiment may not be optimal today. The best
organizations treat DOE as an ongoing practice — repeating key
experiments when conditions change, verifying that established optima
still hold, building a living understanding of process dynamics rather
than a static model frozen in time.
7. The Solution Looking
for a Problem
“We want to do a DOE.” That’s the opening line. Not “we have a
problem with weld penetration” or “our scrap rate on line 3 is
unacceptable” or “the customer is complaining about dimensional
variation.” Just “we want to do a DOE.” Someone read about it in a Lean
Six Sigma book. Someone attended a conference session. Someone wants to
check the box for their Black Belt project.
So they design an experiment. But since there’s no clear problem
being solved, the factors are chosen arbitrarily. The responses are
whatever’s easy to measure. The results answer a question nobody asked.
And the entire exercise becomes what every hollow quality initiative
becomes: a presentation that satisfies the audit requirement and changes
absolutely nothing.
The truth: DOE is a tool, not an initiative. It
exists to answer specific, well-defined questions. “What causes
variation in surface finish?” is a question DOE can answer. “Let’s study
the process” is not. If you can’t state the question you’re trying to
answer in one sentence, you’re not ready for DOE.
What Good DOE Looks Like
When Design of Experiments is done well — when it’s driven by a real
question, designed with discipline, executed with control, and
translated into action — it is one of the most powerful tools in
manufacturing quality. Here’s what the right approach looks like:
Start with the problem. Scrap on the CNC line has
increased from 2% to 5% in the last quarter. The team defines the
specific response (surface roughness, Ra), identifies the plausible
factors (feed rate, spindle speed, tool geometry, coolant concentration,
material hardness), and states the question: “Which of these factors
have the largest effect on surface roughness, and what are the optimal
settings?”
Screen first, optimize second. Run a 12-run
screening design. Discover that feed rate and tool geometry are the only
factors that matter. Drop the rest. Run a focused response surface
design on those two factors. Twelve more runs. Total: 24 runs, two
rounds, clear answers.
Control the environment. Run all trials on the same
machine, same shift, with the same material lot. If that’s not possible,
block by shift. If material changes are unavoidable, include material
hardness as a covariate. Randomize run order to prevent time-correlated
noise from contaminating the results.
Replicate critical points. Run the center point
three times, spread across the experiment. This gives you an estimate of
pure error and tells you whether your model is adequate or whether
you’re missing curvature.
Analyze honestly. Look at the residuals. Check for
outliers. If run 9 looks nothing like the model predicts, investigate —
don’t delete it and move on. The discrepancy might be telling you
something important about a factor you didn’t consider.
Translate and implement. Present the findings in a
way the production team can act on. “Set feed rate to 0.15 mm/rev and
use the coated insert. Here’s why. Here’s the expected improvement.
Here’s the control chart we’ll use to monitor it.” Update the work
instruction. Train the operators. Verify the improvement with actual
production data.
The
Organizational Problem Behind the Technical Problem
Every DOE failure mode described above is ultimately a people
problem, not a statistics problem. The math is sound. The method is
proven. What fails is the organization’s ability to frame the right
question, execute with discipline, and act on what the data reveals.
This is true of every quality tool, but it’s especially true for DOE
because DOE demands more than any other method. It demands statistical
literacy that most manufacturing teams don’t have. It demands production
time that operations doesn’t want to give. It demands analytical skills
that many quality engineers have only at a surface level. And it demands
the willingness to act on findings that might contradict established
practice — which requires organizational courage that most plants have
never cultivated.
The organizations that get DOE right invest in three things:
-
Deep statistical training — not just “how to use
Minitab” but “how to think experimentally.” Understanding why
randomization matters. Knowing when a fractional factorial is
appropriate and when it’s dangerous. Being able to read a residual plot
and know what it’s telling you. -
Protected experimentation time — production
schedules that include dedicated time for structured experiments. Not
“squeeze it in between orders.” Not “run it on second shift when no
one’s watching.” Real, planned, resourced experimentation
windows. -
Implementation discipline — a standard process
for converting DOE findings into updated work instructions, operator
training, and ongoing monitoring. No experiment is considered “complete”
until the optimal settings are in the standard work and the improvement
is verified on the production floor.
The Real Question
If your organization is running Design of Experiments, ask yourself
honestly: When was the last time a DOE actually changed how a process
runs? Not “produced a report.” Not “was presented in a meeting.”
Actually changed the settings on a machine, the steps in a procedure,
the understanding of how your process works.
If the answer is “I can’t remember” or “we’ve never actually
implemented the findings” — then your DOE program isn’t a quality
initiative. It’s a ritual. An expensive, time-consuming, statistically
sophisticated ritual that produces documents instead of knowledge and
presentations instead of improvements.
Design of Experiments is too powerful a tool to waste on theater.
Either do it right — with the discipline, the rigor, and the
follow-through it demands — or don’t do it at all. The worst outcome
isn’t failing to run experiments. The worst outcome is running
experiments that look like science, feel like progress, and change
absolutely nothing.
About the Author: Peter Stasko is a Quality
Architect with over 25 years of experience transforming manufacturing
operations through evidence-based quality systems. He has led Design of
Experiments initiatives across automotive, electronics, and medical
device industries, helping organizations move from trial-and-error
troubleshooting to structured, data-driven process optimization. His
work focuses on bridging the gap between statistical theory and
shop-floor reality — because the best experiment in the world is
worthless if it never reaches the production floor.