12.1 Design of Experiments: 2^k Factorials, ANOVA & Taguchi Robust Design

Key Takeaways

  • Design of Experiments (DOE) studies multiple input factors simultaneously, identifying both main factor effects and non-additive interaction effects that one-factor-at-a-time (OFAT) testing cannot detect.

  • In 2k2^k factorial designs, treatments are coded as ±1\pm 1 orthogonal contrasts; main effects are computed as Contrast/(n⋅2k−1)\text{Contrast} / (n \cdot 2^{k-1}) with Sum of Squares SS=Contrast2/(n⋅2k)\text{SS} = \text{Contrast}^2 / (n \cdot 2^k).

  • Taguchi robust design crosses an inner array of control factors with an outer array of noise factors and picks the settings with the highest signal-to-noise ratio, then uses an adjustment factor to put the mean on target.

  • Analysis of Variance (ANOVA) partitions total variation into orthogonal model components (SSTotal=SSModel+SSError\text{SS}_{\text{Total}} = \text{SS}_{\text{Model}} + \text{SS}_{\text{Error}}), testing factor significance using the F0=MSFactor/MSErrorF_0 = \text{MS}_{\text{Factor}} / \text{MS}_{\text{Error}} statistic.

  • Fractional factorial designs (2k−p2^{k-p}) exploit the sparsity-of-effects principle; Resolution III aliases main effects with 2-factor interactions, Resolution IV aliases 2-factor interactions with each other, and Resolution V leaves main effects and 2-factor interactions unaliased.

Last updated: October 2026

12.1 Design of Experiments: 2^k Factorials, ANOVA & Taguchi Robust Design

Design of Experiments (DOE) is an active, structured empirical methodology for systematically manipulating process inputs (factors) and observing the resulting changes in system outputs (responses). Developed by Sir Ronald A. Fisher in agricultural research and subsequently adapted to industrial engineering by George Box, Genichi Taguchi, and Douglas Montgomery, DOE enables engineers to optimize processes, improve yield, reduce variability, and establish robust operating windows.


1. Experimental Design Fundamentals

In industrial systems, any manufacturing or service process can be modeled as a transfer function:

y=f(x1,x2,…,xk)+ϵy = f(x_1, x_2, \dots, x_k) + \epsilon

where yy is the measured performance response, xix_i represent controllable or uncontrollable input factors, and ϵ\epsilon represents random experimental error.

Core Terminology

  • Factor: An independent variable manipulated by the experimenter. Factors can be quantitative (e.g., furnace temperature in °C, feed rate in mm/min) or qualitative/categorical (e.g., machine operator, material supplier, catalyst type).
  • Level: The specific values or operational settings assigned to a factor during the experiment. In two-level designs, levels are typically designated as low (−1-1) and high (+1+1).
  • Treatment (Treatment Combination): A specific set of factor levels applied to an experimental unit. In a design with kk factors, each unique combination of settings across all kk factors constitutes one treatment.
  • Response: The dependent variable of interest measured after applying a treatment (e.g., tensile strength in MPa, surface roughness in μ\mum, assembly cycle time in seconds).
  • Experimental Error: The residual, unexplained variability among experimental units treated identically. It arises from unmeasured environmental fluctuations, material inconsistencies, and measurement system noise.

Fisher's Three Foundational Principles

To ensure valid statistical inference, any experimental plan must implement three core principles:

  1. Replication: The independent repetition of the basic experiment (treatment combinations applied to distinct experimental units). True replication must be distinguished from repeated measurements: taking five readings on a single machined part reflects measurement error, whereas machining five distinct parts under identical nominal machine settings constitutes true replication. Replication achieves two vital objectives:
    • It allows the experimenter to obtain an independent, unbiased estimate of pure experimental error variance (σ2\sigma^2), which forms the denominator of test statistics in Analysis of Variance (ANOVA).
    • It reduces the standard error of the sample mean: σyˉ=σn\sigma_{\bar{y}} = \frac{\sigma}{\sqrt{n}}, improving the precision of effect estimates.
  2. Randomization: Both the allocation of experimental material and the run sequence of treatment combinations are determined by a random mechanism (e.g., pseudo-random number generation). Randomization:
    • Neutralizes the impact of extraneous, uncontrolled, or lurking variables (e.g., ambient temperature rise throughout a workday, tool wear, raw material aging).
    • Validates the statistical assumption that the error terms ϵij\epsilon_{ij} are independently and identically distributed (i.i.d.) normal random variables: ϵ∼NID(0,σ2)\epsilon \sim \text{NID}(0, \sigma^2).
  3. Blocking: A design technique used to isolate and eliminate the variability introduced by known, controllable nuisance factors. A block is a relatively homogeneous set of experimental units (e.g., a single operator shift, a specific batch of raw chemical reagent, or a single production day). Treatments are randomized within each block, allowing the block-to-block variability to be partitioned out of the residual experimental error in the ANOVA model.

2. One-Factor-at-a-Time (OFAT) vs. Full Factorial Designs

A historically pervasive yet statistically flawed experimental approach is the One-Factor-at-a-Time (OFAT) method. Under OFAT, an engineer fixes all process variables at arbitrary baseline values, varies factor AA across its range to locate an apparent optimum, holds factor AA at that new optimum, varies factor BB, and repeats this sequentially.

The Failure Modes of OFAT

  1. Inability to Detect Interactions: OFAT implicitly assumes that factors act independently (i.e., that the system is strictly additive: f(x1,x2)=g(x1)+h(x2)f(x_1, x_2) = g(x_1) + h(x_2)). When interactions exist, the effect of factor AA changes depending on the level of factor BB. OFAT is mathematically blind to these interactions, often missing the true global optimum and converging on sub-optimal local peaks.
  2. Statistical Inefficiency: To achieve comparable precision for main effects, OFAT requires significantly more experimental runs than a factorial design. In factorial designs, every single run provides information about every factor simultaneously—a property known as hidden replication.
  3. Narrow Generalizability: Conclusions drawn from OFAT apply only under the specific, fixed levels chosen for the remaining factors. If baseline settings are changed, the observed relationships frequently collapse.
FeatureOne-Factor-at-a-Time (OFAT)Full Factorial Design (2k2^k)
Interaction DetectionImpossible; assumes zero interactionRigorously estimates all 2-way, 3-way, ..., kk-way interactions
Run EfficiencyLow; requires more runs for equal precisionMaximum; exploits hidden replication across all design points
Optimization TrajectoryEasily trapped on false ridges/local optimaExplores full multidimensional design space to find true optimum
Mathematical ModelingFails to estimate cross-product terms (xixjx_i x_j)Yields complete orthogonal first-order and interaction regression models

3. 2k2^k Factorial Designs: Matrix Architecture and Orthogonality

A 2k2^k Factorial Design investigates kk factors, each evaluated at exactly two levels: low (coded as −1-1) and high (coded as +1+1). A complete single replicate requires N=2kN = 2^k experimental runs.

Coded Variable Transformation

Quantitative factors are centered and scaled from natural engineering units (XiX_i) to dimensionless coded variables (xix_i) using the linear transformation:

xi=Xi−(Xi,high+Xi,low2)Xi,high−Xi,low2x_i = \frac{X_i - \left(\frac{X_{i,\text{high}} + X_{i,\text{low}}}{2}\right)}{\frac{X_{i,\text{high}} - X_{i,\text{low}}}{2}}

Coding normalizes all factor scales to [−1,+1][-1, +1], removing units of measurement and ensuring that the magnitudes of calculated regression coefficients directly reflect their relative physical importance.

Standard Run Order (Yates' Order) and Run Labels

Runs in a 2k2^k design are designated using standard notation where lowercase letters represent the factors present at their high (+1+1) level in that treatment combination. If all factors are at their low (−1-1) level, the treatment is denoted by the symbol (1)(1):

  • 222^2 Design (4 runs): (1),a,b,ab(1), a, b, ab
  • 232^3 Design (8 runs): (1),a,b,ab,c,ac,bc,abc(1), a, b, ab, c, ac, bc, abc

Sign Table and Design Matrix for a 232^3 Factorial

The columns for interaction effects are generated by algebraic point-by-point multiplication of the corresponding main factor columns:

Run LabelIIAABBABABCCACACBCBCABCABC
(1)(1)+1+1−1-1−1-1+1+1−1-1+1+1+1+1−1-1
aa+1+1+1+1−1-1−1-1−1-1−1-1+1+1+1+1
bb+1+1−1-1+1+1−1-1−1-1+1+1−1-1+1+1
abab+1+1+1+1+1+1+1+1−1-1−1-1−1-1−1-1
cc+1+1−1-1−1-1+1+1+1+1−1-1−1-1+1+1
acac+1+1+1+1−1-1−1-1+1+1+1+1−1-1−1-1
bcbc+1+1−1-1+1+1−1-1+1+1−1-1+1+1−1-1
abcabc+1+1+1+1+1+1+1+1+1+1+1+1+1+1+1+1

The Mathematical Property of Orthogonality

The design matrix X\mathbf{X} possesses two vital mathematical properties:

  1. Zero Sum: The sum of coefficients in any factor or interaction column (excluding identity II) is zero: ∑i=12kci=0\sum_{i=1}^{2^k} c_i = 0.
  2. Orthogonal Dot Product: The inner product of any two distinct columns cj\mathbf{c}_j and cm\mathbf{c}_m (j≠mj \neq m) is identically zero:

cjTcm=∑i=12kcijcim=0\mathbf{c}_j^T \mathbf{c}_m = \sum_{i=1}^{2^k} c_{ij} c_{im} = 0

Because the matrix XTX=NI\mathbf{X}^T \mathbf{X} = N \mathbf{I} is diagonal, all main effect and interaction estimates are statistically independent and uncorrelated. Adding or removing a factor from the analysis does not alter the numerical estimates of the remaining effects.


4. Main Effect, Contrast, and Sum of Squares Formulas

For a 2k2^k factorial experiment replicated nn times (total observations N=n⋅2kN = n \cdot 2^k):

1. Contrast

The contrast for any effect is the linear combination of treatment totals weighted by the sign vector (ci∈{−1,+1}c_i \in \{-1, +1\}) from the sign table:

ContrastA=∑i=12kci⋅yi⋅=yA(+)−yA(−)\text{Contrast}_A = \sum_{i=1}^{2^k} c_i \cdot y_{i\cdot} = y_{A(+)} - y_{A(-)}

where yi⋅y_{i\cdot} is the sum of all nn replicates for treatment combination ii, and yA(+)y_{A(+)} and yA(−)y_{A(-)} represent the aggregate responses across all runs where factor AA is at its high and low levels, respectively.

2. Main Effect

The main effect represents the average change in response produced by moving factor AA from its low level to its high level:

EffectA=ContrastAn⋅2k−1=yˉA(+)−yˉA(−)\text{Effect}_A = \frac{\text{Contrast}_A}{n \cdot 2^{k-1}} = \bar{y}_{A(+)} - \bar{y}_{A(-)}

Note that the denominator n⋅2k−1n \cdot 2^{k-1} is the total number of observations at the high level (which equals the total number at the low level).

3. Regression Coefficient

In the coded first-order regression model y=β0+∑βixi+∑∑βijxixj+…y = \beta_0 + \sum \beta_i x_i + \sum \sum \beta_{ij} x_i x_j + \dots, the regression coefficient is exactly half the effect estimate because the coded factor spans 2 units (from −1-1 to +1+1):

βA=EffectA2=ContrastAn⋅2k\beta_A = \frac{\text{Effect}_A}{2} = \frac{\text{Contrast}_A}{n \cdot 2^k}

4. Sum of Squares (SS)

The sum of squares attributable to any factor or interaction effect with 1 degree of freedom is:

SSA=(ContrastA)2n⋅2k=n⋅2k4(EffectA)2SS_A = \frac{(\text{Contrast}_A)^2}{n \cdot 2^k} = \frac{n \cdot 2^k}{4} (\text{Effect}_A)^2


5. Two-Factor and Three-Factor Interactions

An interaction occurs when the effect of one factor is contingent upon the setting of another factor.

Mathematical Formulation of a 2-Factor Interaction

The interaction effect ABAB is half the difference between the effect of AA at the high level of BB and the effect of AA at the low level of BB:

AB=12[A(B=+1)−A(B=−1)]=ContrastABn⋅2k−1AB = \frac{1}{2} \left[ A_{(B=+1)} - A_{(B=-1)} \right] = \frac{\text{Contrast}_{AB}}{n \cdot 2^{k-1}}

For a 222^2 design:

ContrastAB=[ab−b]−[a−(1)]=ab−b−a+(1)\text{Contrast}_{AB} = [ab - b] - [a - (1)] = ab - b - a + (1)

Graphical Interpretation via Interaction Plots

Interaction plots display the response yy on the vertical axis against factor AA on the horizontal axis, with separate lines plotted for each level of factor BB:

  • Parallel Lines: Zero interaction. The effect of factor AA is identical regardless of factor BB's setting. The system is purely additive: SSAB≈0SS_{AB} \approx 0.
  • Non-Parallel, Non-Intersecting Lines (Ordinal / Synergistic Interaction): The magnitude of factor AA's effect changes depending on BB, but the direction of the effect remains consistent. For instance, raising temperature always increases yield, but it increases yield far more dramatically when pressure is high.
  • Intersecting Lines (Disordinal / Antagonistic Interaction): The sign of factor AA's effect completely reverses depending on the setting of factor BB. For example, increasing cutting speed improves tool life at low feed rates, but severely degrades tool life at high feed rates. In the presence of a strong disordinal interaction, main effects have little physical meaning on their own and should never be interpreted in isolation.

6. Analysis of Variance (ANOVA) for 2k2^k Factorials

ANOVA provides the formal statistical framework for hypothesis testing in experimental design, decomposing total variability into orthogonal components.

Decomposition of Total Sum of Squares

SSTotal=SSModel+SSErrorSS_{\text{Total}} = SS_{\text{Model}} + SS_{\text{Error}}

SSTotal=∑i=12k∑j=1n(yij−yˉ⋅⋅)2=∑i=12k∑j=1nyij2−Y⋅⋅2NSS_{\text{Total}} = \sum_{i=1}^{2^k} \sum_{j=1}^n (y_{ij} - \bar{y}_{\cdot\cdot})^2 = \sum_{i=1}^{2^k} \sum_{j=1}^n y_{ij}^2 - \frac{Y_{\cdot\cdot}^2}{N}

where CF=Y⋅⋅2NCF = \frac{Y_{\cdot\cdot}^2}{N} is the correction factor for the mean, Y⋅⋅Y_{\cdot\cdot} is the grand total of all N=n⋅2kN = n \cdot 2^k observations, and yˉ⋅⋅=Y⋅⋅N\bar{y}_{\cdot\cdot} = \frac{Y_{\cdot\cdot}}{N} is the grand mean.

Because the factorial design matrix is orthogonal, SSModelSS_{\text{Model}} is the exact linear sum of the sums of squares of all individual main effects and interaction effects:

SSModel=∑j=1kSSj+∑j<mSSjm+⋯+SS12…kSS_{\text{Model}} = \sum_{j=1}^k SS_j + \sum_{j < m} SS_{jm} + \dots + SS_{12\dots k}

Error Sum of Squares (SSErrorSS_{\text{Error}})

Pure experimental error is obtained by summing the internal variation within each of the 2k2^k treatment combinations:

SSError=∑i=12k∑j=1n(yij−yˉi⋅)2=SSTotal−SSModelSS_{\text{Error}} = \sum_{i=1}^{2^k} \sum_{j=1}^n (y_{ij} - \bar{y}_{i\cdot})^2 = SS_{\text{Total}} - SS_{\text{Model}}

Degrees of Freedom (dfdf)

  • Total degrees of freedom: dfTotal=N−1=n⋅2k−1df_{\text{Total}} = N - 1 = n \cdot 2^k - 1
  • Each main effect or interaction: dfEffect=1df_{\text{Effect}} = 1
  • Model degrees of freedom: dfModel=2k−1df_{\text{Model}} = 2^k - 1
  • Error degrees of freedom: dfError=2k(n−1)=N−2kdf_{\text{Error}} = 2^k(n - 1) = N - 2^k

Mean Squares and the FF-Test Statistic

For each term, the Mean Square is its sum of squares divided by its degrees of freedom:

MSEffect=SSEffect1=SSEffectMS_{\text{Effect}} = \frac{SS_{\text{Effect}}}{1} = SS_{\text{Effect}}

MSError=SSErrordfError=SSError2k(n−1)MS_{\text{Error}} = \frac{SS_{\text{Error}}}{df_{\text{Error}}} = \frac{SS_{\text{Error}}}{2^k(n - 1)}

Under the null hypothesis H0:Effect=0H_0: \text{Effect} = 0 versus H1:Effect≠0H_1: \text{Effect} \neq 0, the test statistic:

F0=MSEffectMSErrorF_0 = \frac{MS_{\text{Effect}}}{MS_{\text{Error}}}

follows an FF-distribution with df1=1df_1 = 1 and df2=2k(n−1)df_2 = 2^k(n - 1). The null hypothesis is rejected at significance level α\alpha if:

F0>Fα,1,2k(n−1)or ifp-value<αF_0 > F_{\alpha, 1, 2^k(n-1)} \quad \text{or if} \quad p\text{-value} < \alpha

Unreplicated 2k2^k Designs and Pooling

When n=1n = 1, dfError=2k(1−1)=0df_{\text{Error}} = 2^k(1 - 1) = 0, meaning no internal estimate of pure error exists. In such cases, engineers employ two methods:

  1. Daniel's Normal Probability Plot of Effects: Unimportant effects conform to a normal distribution centered at zero and fall along a straight line on a normal probability plot. Effects that diverge significantly from the line are identified as active.
  2. Pooling (Sparsity-of-Effects Principle): High-order interactions (e.g., 3-factor and 4-factor terms) are assumed to be negligible and are pooled together to form an artificial error term with degrees of freedom equal to the sum of the pooled terms' degrees of freedom.

7. Fractional Factorial Designs (2k−p2^{k-p})

As the number of factors kk increases, the run requirements of full factorial designs grow exponentially (25=322^5 = 32, 26=642^6 = 64, 27=1282^7 = 128). According to the sparsity-of-effects principle (the Pareto rule of experimental design), systems are primarily driven by main effects and low-order (two-factor) interactions; three-factor and higher interactions are almost universally negligible.

A 2k−p2^{k-p} fractional factorial design runs a 12p\frac{1}{2^p} fraction of the full factorial, requiring 2k−p2^{k-p} runs while systematically confounding (aliasing) effects.

Design Generators and Defining Relations

To construct a fractional factorial:

  1. Choose k−pk - p basic factors to form a full 2k−p2^{k-p} design matrix.
  2. Confound the remaining pp factors with high-order interaction columns of the basic design. These assignment equations are called design generators.
  3. Multiplying both sides of a generator by the generated factor yields a word in the defining relation (since D×D=D2=ID \times D = D^2 = I, where II is the identity column of +1+1s).

Example: In a 24−12^{4-1} half-fraction (24−1=82^{4-1} = 8 runs), basic factors are A,B,CA, B, C. We set D=ABCD = ABC. Multiplying by DD gives the defining relation:

I=ABCDI = ABCD

Aliasing (Confounding) Structure

The alias of any factor is found by multiplying that factor across the defining relation using modulo 2 arithmetic (X2=IX^2 = I):

  • For AA: A⋅I=A⋅ABCD  ⟹  A=BCDA \cdot I = A \cdot ABCD \implies A = BCD
  • For BB: B⋅I=B⋅ABCD  ⟹  B=ACDB \cdot I = B \cdot ABCD \implies B = ACD
  • For ABAB: AB⋅I=AB⋅ABCD  ⟹  AB=A2B2CD  ⟹  AB=CDAB \cdot I = AB \cdot ABCD \implies AB = A^2B^2CD \implies AB = CD

Thus, the calculated contrast for column ABAB actually estimates the composite quantity βAB+βCD\beta_{AB} + \beta_{CD}. It is impossible to determine whether a significant effect is caused by ABAB or CDCD without additional runs.

Design Resolution Levels

The resolution of a design is denoted by Roman numerals (R=III, IV, VR = \text{III, IV, V}) and equals the length of the shortest word in the defining relation:

  • Resolution III Designs: No main effects are aliased with other main effects, but main effects are aliased with two-factor interactions (e.g., 23−12^{3-1} with I=ABC  ⟹  A=BCI = ABC \implies A = BC). These are screening designs used to identify whether any factors are active, assuming interactions are zero.
  • Resolution IV Designs: Main effects are unaliased with any two-factor interactions, but two-factor interactions are aliased with each other (e.g., 24−12^{4-1} with I=ABCD  ⟹  A=BCDI = ABCD \implies A = BCD, but AB=CDAB = CD). These designs cleanly isolate main effects.
  • Resolution V Designs: No main effect or two-factor interaction is aliased with any other main effect or two-factor interaction; two-factor interactions are aliased only with three-factor interactions (e.g., 25−12^{5-1} with I=ABCDE  ⟹  A=BCDE,AB=CDEI = ABCDE \implies A = BCDE, AB = CDE). These provide high-fidelity modeling capability.

8. Complete Worked Numerical 232^3 DOE Problem

Problem Formulation

An industrial manufacturing engineer wants to maximize the surface finish quality (measured as an index from 0 to 100, where higher is superior) of an aluminum aerospace component machined on a 5-axis CNC mill. The engineer investigates three controllable process parameters:

  • Factor AA (Cutting Speed): Low = 1,500 rpm (−1-1), High = 2,500 rpm (+1+1)
  • Factor BB (Feed Rate): Low = 150 mm/min (−1-1), High = 250 mm/min (+1+1)
  • Factor CC (Depth of Cut): Low = 1.0 mm (−1-1), High = 2.5 mm (+1+1)

The experiment is executed with n=2n = 2 true replicates in a completely randomized run order, yielding N=23×2=16N = 2^3 \times 2 = 16 total runs.

Raw Experimental Data

Run LabelAABBCCReplicate 1 (yi1y_{i1})Replicate 2 (yi2y_{i2})Treatment Total (yi⋅y_{i\cdot})Treatment Mean (yˉi⋅\bar{y}_{i\cdot})
(1)(1)−1-1−1-1−1-122.024.046.023.0
aa+1+1−1-1−1-132.030.062.031.0
bb−1-1+1+1−1-127.029.056.028.0
abab+1+1+1+1−1-143.045.088.044.0
cc−1-1−1-1+1+118.020.038.019.0
acac+1+1−1-1+1+126.028.054.027.0
bcbc−1-1+1+1+1+123.025.048.024.0
abcabc+1+1+1+1+1+138.040.078.039.0
TotalY⋅⋅=470.0Y_{\cdot\cdot} = 470.0yˉ⋅⋅=29.375\bar{y}_{\cdot\cdot} = 29.375

Step 1: Calculate Contrasts

Applying the sign table columns to the treatment totals yi⋅y_{i\cdot}:

  • ContrastA=−46+62−56+88−38+54−48+78=+94.0\text{Contrast}_A = -46 + 62 - 56 + 88 - 38 + 54 - 48 + 78 = +94.0
  • ContrastB=−46−62+56+88−38−54+48+78=+70.0\text{Contrast}_B = -46 - 62 + 56 + 88 - 38 - 54 + 48 + 78 = +70.0
  • ContrastC=−46−62−56−88+38+54+48+78=−34.0\text{Contrast}_C = -46 - 62 - 56 - 88 + 38 + 54 + 48 + 78 = -34.0
  • ContrastAB=+46−62−56+88+38−54−48+78=+30.0\text{Contrast}_{AB} = +46 - 62 - 56 + 88 + 38 - 54 - 48 + 78 = +30.0
  • ContrastAC=+46−62+56−88−38+54−48+78=−2.0\text{Contrast}_{AC} = +46 - 62 + 56 - 88 - 38 + 54 - 48 + 78 = -2.0
  • ContrastBC=+46+62−56−88−38−54+48+78=−2.0\text{Contrast}_{BC} = +46 + 62 - 56 - 88 - 38 - 54 + 48 + 78 = -2.0
  • ContrastABC=−46+62+56−88+38−54−48+78=−2.0\text{Contrast}_{ABC} = -46 + 62 + 56 - 88 + 38 - 54 - 48 + 78 = -2.0

Step 2: Calculate Effects and Sums of Squares

With n=2n = 2 and k=3k = 3, the effect divisor is n⋅2k−1=2⋅4=8n \cdot 2^{k-1} = 2 \cdot 4 = 8, and the sum of squares divisor is n⋅2k=2⋅8=16n \cdot 2^k = 2 \cdot 8 = 16:

  • Factor AA: EffectA=94.08=+11.75\text{Effect}_A = \frac{94.0}{8} = +11.75 SSA=94.0216=8836.016=552.25SS_A = \frac{94.0^2}{16} = \frac{8836.0}{16} = 552.25

  • Factor BB: EffectB=70.08=+8.75\text{Effect}_B = \frac{70.0}{8} = +8.75 SSB=70.0216=4900.016=306.25SS_B = \frac{70.0^2}{16} = \frac{4900.0}{16} = 306.25

  • Factor CC: EffectC=−34.08=−4.25\text{Effect}_C = \frac{-34.0}{8} = -4.25 SSC=(−34.0)216=1156.016=72.25SS_C = \frac{(-34.0)^2}{16} = \frac{1156.0}{16} = 72.25

  • Interaction ABAB: EffectAB=30.08=+3.75\text{Effect}_{AB} = \frac{30.0}{8} = +3.75 SSAB=30.0216=900.016=56.25SS_{AB} = \frac{30.0^2}{16} = \frac{900.0}{16} = 56.25

  • Interaction ACAC: EffectAC=−2.08=−0.25\text{Effect}_{AC} = \frac{-2.0}{8} = -0.25 SSAC=(−2.0)216=4.016=0.25SS_{AC} = \frac{(-2.0)^2}{16} = \frac{4.0}{16} = 0.25

  • Interaction BCBC: EffectBC=−2.08=−0.25\text{Effect}_{BC} = \frac{-2.0}{8} = -0.25 SSBC=(−2.0)216=4.016=0.25SS_{BC} = \frac{(-2.0)^2}{16} = \frac{4.0}{16} = 0.25

  • Interaction ABCABC: EffectABC=−2.08=−0.25\text{Effect}_{ABC} = \frac{-2.0}{8} = -0.25 SSABC=(−2.0)216=4.016=0.25SS_{ABC} = \frac{(-2.0)^2}{16} = \frac{4.0}{16} = 0.25

Step 3: Calculate Error Sum of Squares (SSErrorSS_{\text{Error}}) and Total Sum of Squares (SSTotalSS_{\text{Total}})

For each treatment combination, the two replicates differ by exactly 2.0 units (∣yi1−yi2∣=2.0|y_{i1} - y_{i2}| = 2.0). The sample variance within each treatment cell is:

si2=(yi1−yˉi⋅)2+(yi2−yˉi⋅)22−1=(−1.0)2+(1.0)2=2.0s_i^2 = \frac{(y_{i1} - \bar{y}_{i\cdot})^2 + (y_{i2} - \bar{y}_{i\cdot})^2}{2 - 1} = (-1.0)^2 + (1.0)^2 = 2.0

Summing across all 8 treatment cells:

SSError=∑i=18(n−1)si2=8×(1×2.0)=16.00SS_{\text{Error}} = \sum_{i=1}^8 (n - 1) s_i^2 = 8 \times (1 \times 2.0) = 16.00

dfError=23(n−1)=8×(2−1)=8df_{\text{Error}} = 2^3(n - 1) = 8 \times (2 - 1) = 8

MSError=16.008=2.00MS_{\text{Error}} = \frac{16.00}{8} = 2.00

Checking total variability:

CF=Y⋅⋅2N=470.0216=220900.016=13806.25CF = \frac{Y_{\cdot\cdot}^2}{N} = \frac{470.0^2}{16} = \frac{220900.0}{16} = 13806.25

∑i=18∑j=12yij2=222+242+322+302+⋯+382+402=14810.00\sum_{i=1}^8 \sum_{j=1}^2 y_{ij}^2 = 22^2 + 24^2 + 32^2 + 30^2 + \dots + 38^2 + 40^2 = 14810.00

SSTotal=14810.00−13806.25=1003.75SS_{\text{Total}} = 14810.00 - 13806.25 = 1003.75

Sum of all model terms:

SSModel=552.25+306.25+72.25+56.25+0.25+0.25+0.25=987.75SS_{\text{Model}} = 552.25 + 306.25 + 72.25 + 56.25 + 0.25 + 0.25 + 0.25 = 987.75

SSTotal=SSModel+SSError=987.75+16.00=1003.75(Identical Check)SS_{\text{Total}} = SS_{\text{Model}} + SS_{\text{Error}} = 987.75 + 16.00 = 1003.75 \quad \text{(Identical Check)}

Step 4: Complete ANOVA Table

Source of VariationSum of Squares (SSSS)Degrees of Freedom (dfdf)Mean Square (MSMS)F0F_0 Statistic (MS/MSEMS/MS_E)Critical F0.05,1,8F_{0.05, 1, 8}Statistical Significance (p<0.05p < 0.05)
Factor AA (Speed)552.251552.25276.135.32Significant (p<0.0001p < 0.0001)
Factor BB (Feed)306.251306.25153.135.32Significant (p<0.0001p < 0.0001)
Factor CC (Depth)72.25172.2536.135.32Significant (p=0.0003p = 0.0003)
Interaction ABAB56.25156.2528.135.32Significant (p=0.0007p = 0.0007)
Interaction ACAC0.2510.250.135.32Not Significant (p=0.732p = 0.732)
Interaction BCBC0.2510.250.135.32Not Significant (p=0.732p = 0.732)
Interaction ABCABC0.2510.250.135.32Not Significant (p=0.732p = 0.732)
Error (Pure Residual)16.0082.00———
Total1003.7515————

Engineering Conclusions & Process Optimization

  1. Active Factors: Cutting Speed (AA), Feed Rate (BB), and Depth of Cut (CC) are all highly statistically significant main effects. Furthermore, the two-factor interaction ABAB is strongly significant (F0=28.13>5.32F_0 = 28.13 > 5.32).
  2. Interaction Synergy: The positive interaction coefficient βAB=+1.875\beta_{AB} = +1.875 indicates that increasing cutting speed produces an even greater improvement in surface quality when feed rate is also set at its high level.
  3. Optimal Operating Setpoint: To maximize surface finish, the engineer sets Cutting Speed AA to high (+1=2,500+1 = 2,500 rpm), Feed Rate BB to high (+1=250+1 = 250 mm/min), and Depth of Cut CC to low (−1=1.0-1 = 1.0 mm, because EffectC=−4.25\text{Effect}_C = -4.25 is negative, meaning lower depth increases quality). This corresponds to treatment combination abab, achieving an expected mean response of yˉab=44.0\bar{y}_{ab} = 44.0 index points.

9. Taguchi Robust Parameter Design

Genichi Taguchi's approach to process improvement aims to make a product or process robust, meaning insensitive to variation it cannot control. Its main ideas are:

  • Quality loss function: loss grows with the square of the deviation from target, L(y)=k(y−T)2L(y) = k(y - T)^2, so reducing variation around the target has value even inside the specification.
  • Control factors vs. noise factors: control factors are settings the engineer chooses, such as temperature or feed rate. Noise factors vary in production or use, such as ambient humidity or material lot, and are hard or costly to control.
  • Orthogonal arrays: small, balanced fractional designs such as the L4L_4 (up to 3 two-level factors in 4 runs), L8L_8 (up to 7 two-level factors in 8 runs), L9L_9 (up to 4 three-level factors in 9 runs), and L18L_{18} (mixed two- and three-level factors in 18 runs).
  • Crossed arrays: an inner array of control-factor settings is crossed with an outer array of noise conditions, so each control setting is tested across the noise.

Signal-to-Noise (S/N) Ratios

Each inner-array row is summarized by an S/N ratio in decibels. The engineer chooses the control settings with the highest S/N ratio.

GoalS/N ratio
Smaller-the-better (wear, shrinkage)−10log⁡10(1n∑yi2)-10 \log_{10}\left(\frac{1}{n}\sum y_i^2\right)
Larger-the-better (strength, life)−10log⁡10(1n∑1yi2)-10 \log_{10}\left(\frac{1}{n}\sum \frac{1}{y_i^2}\right)
Nominal-the-best (a dimension)10log⁡10(yˉ2s2)10 \log_{10}\left(\frac{\bar{y}^2}{s^2}\right)

Example: A bond-strength test (larger-the-better) gives 42, 45, and 44 MPa across three noise conditions. The S/N ratio is −10log⁡10[13(1422+1452+1442)]≈32.8-10\log_{10}\left[\frac{1}{3}\left(\frac{1}{42^2} + \frac{1}{45^2} + \frac{1}{44^2}\right)\right] \approx 32.8 dB. A setting with higher and more consistent strength would score higher.

Two-Step Optimization (Nominal-the-Best)

  1. Choose control-factor levels that maximize the S/N ratio, which minimizes variation relative to the mean.
  2. Use an adjustment factor, one that shifts the mean but barely affects the S/N ratio, to move the mean onto target.

Criticisms to know. Statisticians note that crossed arrays can need many runs, that S/N ratios can mix up location and dispersion effects, and that saturated orthogonal arrays confound main effects with interactions. A common alternative is a single combined design that models the response directly, including control-by-noise interactions. Exam questions usually test the vocabulary: control vs. noise factors, inner and outer arrays, S/N ratio types, and the two-step method.

Test Your Knowledge

In an unreplicated 2^4 factorial design investigating four process parameters, the experimenter pools the 3-factor and 4-factor interaction sum of squares to estimate experimental error. How many degrees of freedom are available for this pooled error estimate in the ANOVA table?

A

4

B

3

C

5

D

11

Test Your Knowledge

An industrial engineer runs a 2^{5-2} fractional factorial design with design generators D = AB and E = BC, yielding the defining relation I = ABD = BCE = ACDE. What is the complete alias chain for the main effect of factor A?

A

A + BD + ABCE + CDE

B

A + B + CDE + ACDE

C

A + D + BCE + ABC

D

A + CD + BE + ADE

Sections you finish are checked off in the contents.