Design of Experiments: When Your DOE Becomes a Statistics Exercise Nobody Applies — and the Knowledge You Were Supposed to Gain Became the P-Values You Collected and the Process Understanding You Never Actually Achieved

Blog

Walk into most manufacturing plants and ask the quality engineer
about their last Design of Experiments. Watch the expression change.
Some will give you a confident answer about a screening design they ran
six months ago. Most will shuffle through a binder, pull out a document
dense with ANOVA tables and half-normal plots, and show you something
that looks impressive — something that earned a pat on the back from a
manager who didn’t understand it — and then quietly admit that nothing
on the production floor actually changed because of it.

That is the quiet failure of DOE in modern manufacturing. Not that
people aren’t running experiments. They are. Not that the software can’t
crunch the numbers. It can. The failure is that Design of Experiments
has been reduced to a statistical exercise — a deliverable, a milestone,
a slide in a review deck — while the actual purpose of the method, which
is to build genuine process understanding that drives better decisions,
gets lost somewhere between the regression output and the
PowerPoint.

I have spent decades watching this happen. And I can tell you exactly
why it happens, what it costs, and what to do about it.

What DOE Was Actually
Invented to Do

Ronald Fisher developed the principles of experimental design in the
1920s and 1930s at an agricultural research station in England. His
problem was simple: fertilizer trials were expensive, land was limited,
and testing one variable at a time meant that seasons passed before you
learned anything useful. Fisher needed a method to extract maximum
information from minimum resources — to test multiple factors
simultaneously, understand how they interacted, and separate real
effects from random noise.

Genichi Taguchi adapted these ideas for manufacturing in the 1950s
and built an entire engineering philosophy around robust design. George
Box refined the methods further, adding practical wisdom and response
surface methodology that made DOE accessible to industrial
practitioners. The lineage is clear: DOE has always been about efficient
learning under constraints.

Notice what none of those pioneers said. None of them said “produce
an ANOVA table.” None of them said “generate a p-value less than 0.05.”
None of them said “create a mathematical model and hand it to someone
else.”

The goal was always understanding. Decisions. Action.

How DOE Dies in Practice

Here is the typical lifecycle of a Design of Experiments in a
manufacturing organization.

Phase 1: Enthusiasm

Someone attends a training course. Maybe it is a Six Sigma Green Belt
program, maybe a vendor workshop, maybe a university short course. They
learn about full factorial designs, fractional factorials, center
points, blocking. They fire up Minitab or JMP or Design Expert. They
feel powerful. They run back to the plant determined to apply what they
learned.

So far, so good.

Phase 2: The Experiment

The engineer identifies a problem — a coating thickness that varies
too much, a welding process with inconsistent penetration, an injection
molding cycle that produces flash on some shots and short shots on
others. They design a fractional factorial. They select factors:
temperature, pressure, speed, material lot. They choose two levels for
each. They calculate the number of runs. They block by shift to account
for operator variation.

On paper, this looks exactly right.

But then reality intervenes. Production does not want to stop running
parts so the quality team can experiment. The runs get squeezed into a
four-hour window on a Saturday morning when the line would otherwise be
idle. The operators assigned to help have not been briefed on why this
matters. Material lots are swapped hastily. A temperature setpoint gets
rounded because the controller only increments in five-degree steps. One
run gets skipped entirely because the lunch break ran long and the
maintenance team needed the line for a PM.

The engineer records what they can. Some data points are incomplete.
Some factor levels deviate from the plan. The experiment that was
designed with statistical rigor is executed with industrial
compromise.

Phase 3: The Analysis

Back at the desk, the data goes into the software. A few clicks
produce a half-normal plot, a Pareto chart of effects, an ANOVA table.
Some factors are significant. Some are not. There is a possible
interaction between temperature and pressure that looks interesting on
the interaction plot.

But here is where things get murky. The engineer who designed the
experiment may not have the authority to implement the findings. The
production manager who controls the line may not understand the
analysis. The results get written up in a report — ten pages, tables,
charts, conclusions — and the report goes into a review meeting where
nobody has time to discuss the interaction plot in detail.

The recommendations from the DOE might be adopted. Or they might be
filed for “future consideration.” Or they might be overridden by someone
with more organizational authority but less process understanding.

Phase 4: Irrelevance

Six months later, the process is still running at the original
settings. The DOE report is in a binder. The engineer has moved on to
the next fire. Nobody on the production floor remembers the experiment,
let alone what it revealed.

And here is the most insidious part: the organization believes it is
doing DOE. They can point to the report. They can list it as evidence of
continuous improvement activity during customer audits. The checkbox is
marked. The capability was demonstrated. But the process did not
improve, the understanding did not transfer, and the only thing that was
actually produced was a document.

What This Costs You

The cost of performative DOE is not just the wasted time of the
experiment itself. It is the opportunity cost — the knowledge you could
have gained but did not.

A properly executed DOE on an injection molding process might reveal
that hold pressure and cooling time interact in a way that creates an
optimal window for dimensional stability. That knowledge, if acted upon,
could reduce scrap by 30 percent, eliminate a sorting operation, and
free up capacity. Instead, the process continues to drift, the sorting
station stays staffed, and the capacity remains lost.

A well-designed mixture experiment on a chemical formulation could
identify that two of the five ingredients in your adhesive recipe
contribute nothing to bond strength at the levels currently specified.
You could reformulate, reduce material cost, and simplify your supply
chain. Instead, you continue paying for ingredients that serve no
functional purpose because nobody validated their contribution.

The cost is also cultural. When engineers see that their DOE work
leads to reports but not changes, they stop designing experiments. They
start treating DOE as a resume item rather than a working tool. The
organizational muscle for structured experimentation atrophies. And when
a genuine crisis hits — a new product launch, a field failure, a
customer escalation — the team lacks both the skills and the
organizational habits to learn quickly and systematically.

The
Five Differences Between DOE That Works and DOE That Doesn’t

Having run and reviewed hundreds of experiments over twenty-five
years, I can identify the patterns that separate meaningful DOE from
statistical theater.

1. Clear Decision Before
Design

Organizations that get value from DOE start every experiment with a
decision statement: “We are trying to decide whether to change the
curing temperature from 150°C to 180°C, and if so, what the optimal
level is.” This is fundamentally different from “We want to understand
the factors affecting cure quality.”

A decision statement forces clarity. It tells you what action you
will take based on the results. If you cannot articulate the decision
you are trying to inform, you are not ready to design the experiment —
you are ready to have a conversation about what problem you are actually
trying to solve.

2.
Execution Treated as Engineering, Not Interruption

When the experiment runs, the production team treats it as real work
— because it is. The line is scheduled for it. Operators are briefed on
the purpose. Data collection sheets are prepared and verified. Time is
allocated — not stolen from between shifts or squeezed into a Saturday
morning.

I have seen organizations where the production supervisor actively
participates in the DOE planning because they own the line and they want
the answers too. That is when DOE works: when production and quality
share the same question.

3.
Analysis Focused on Practical Significance, Not Just Statistical
Significance

A factor can be statistically significant and practically irrelevant.
A p-value of 0.001 on a factor that moves your response by 0.2 percent
of tolerance is a mathematical curiosity, not an engineering
insight.

Effective DOE practitioners look at effect sizes, not just
significance flags. They ask: “How much does this factor actually move
the output, and does that movement matter relative to our specification
and our process noise?” They use the analysis to build a predictive
model that can be validated against new runs — not just a table that
gets pasted into a report.

4. Follow-Through on
Confirmation Runs

The single most neglected step in DOE is the confirmation run. After
the analysis identifies optimal factor settings, you run those settings
— actually run them, on the production line, under normal conditions —
and you verify that the predicted improvement materializes.

This step is neglected because it requires exactly the kind of
organizational commitment that performative DOE avoids. A confirmation
run means scheduling time on the line. It means collecting data over
enough parts to have statistical confidence. It means potentially
changing a process standard if the results confirm the prediction.

Organizations that get value from DOE treat the confirmation run as
non-negotiable. It is the bridge between statistical analysis and
operational change. Without it, your DOE is a hypothesis. With it, your
DOE becomes a decision.

5. Knowledge
Transfer, Not Just Knowledge Creation

When the DOE is complete and the confirmation run validates the
findings, the knowledge gets transferred to the people who run the
process. The new settings go into the control plan. The operators are
trained on why the settings changed. The reasoning — not just the
numbers — is documented in language the team can understand.

This is where DOE transitions from a project to a capability. The
operator who understands why hold pressure matters at this particular
setpoint makes better decisions when something unusual happens. The
process engineer who inherited the DOE results can build on them for the
next experiment. Knowledge compounds — but only if it is shared.

A Practical Framework
for Getting Started

If your organization has been going through the motions with DOE and
wants to extract real value, start small and start right.

Pick one process. Not the most complex process in
your plant — pick one where you have a specific decision to make and the
factors are reasonably controllable. An injection molding parameter
optimization. A machining feed and speed study. A curing temperature and
time evaluation.

Write the decision statement. Before you open any
software, write down: “This experiment will help us decide [specific
decision], and based on the results we will [specific action].”

Involve production from day one. The supervisor who
owns the line should help select the factors, review the run plan, and
commit the time. If production is not on board, do not proceed. Wait
until they are. A well-executed experiment that never runs is better
than a poorly executed one that produces misleading results.

Design for the noise you cannot control. Randomize
run order. Block by shift or day if there are known noise sources. Add
center points to check for curvature. These are not statistical niceties
— they are the difference between learning something real and fooling
yourself.

Run the confirmation. No exceptions. If the model
predicts a 20 percent reduction in variation at the new settings, run 30
parts at those settings and measure. If the prediction holds, update
your control plan and celebrate. If it does not, figure out why —
because the gap between prediction and reality is often where the most
important learning lives.

Document the story, not just the statistics. Write a
one-page summary that any engineer in your organization can read and
understand: what was the question, what did you find, what changed, what
was the result. File the full analysis separately for anyone who wants
the detail. But the story — the decision, the discovery, the action —
that is what makes the knowledge transferable.

The Real ROI of DOE

When Design of Experiments is done well, the return on investment is
not subtle. I have seen single experiments save hundreds of thousands of
dollars per year in reduced scrap. I have seen screening designs
eliminate unnecessary process steps that were adding cost without adding
value. I have seen response surface methodologies find operating windows
that simultaneously improved quality and throughput — something that
trial-and-error would never have discovered because the optimal region
was not where intuition pointed.

But the biggest return is cumulative. Organizations that learn to
experiment well develop a capability that competitors cannot easily
replicate. They make decisions based on evidence rather than opinion.
They resolve disagreements with data rather than hierarchy. They get
faster over time — not because they cut corners, but because they have
built the organizational habits and trust that make efficient learning
possible.

That capability is what Fisher was after a century ago. It is what
Taguchi brought to manufacturing. It is what every quality engineer who
has ever run a meaningful experiment was reaching for.

The statistics are a means. Understanding is the method. Better
decisions are the goal.

If your DOE program is producing reports but not decisions, it is
time to ask the hardest question in quality engineering: are we
learning, or are we performing?


About the Author: Peter Stasko is a Quality
Architect with over 25 years of experience in manufacturing quality
management, process optimization, and continuous improvement. He
specializes in helping organizations bridge the gap between statistical
theory and operational practice — turning tools like DOE, SPC, and FMEA
into decisions that actually improve processes and products.

Scroll top