What is ANOVA, and why every DoE result needs it
Design of Experiments finds the best parameter combination efficiently. ANOVA tells you whether that answer is statistically real — or a well-dressed accident.
1 · What ANOVA actually answers
Analysis of Variance splits the variation in your results into two buckets: variation explained by a factor (e.g. DTE) and variation left over as noise. The F-value is the ratio of those two. A large F means the spread between factor levels is bigger than the spread inside them — the factor is doing real work, not decorating randomness.
2 · Why it prevents overfitting
A backtest will always produce a 'best' parameter. Without a significance test you cannot tell whether 14 DTE genuinely beats 30 DTE or simply won a coin flip across a handful of runs. ANOVA puts a probability on that question. Factors that fail the test are the ones you must NOT tune — tuning noise is precisely how curve-fitted strategies are born.
3 · Reading p-value, η² and R²
p-value: the probability of seeing this much separation if the factor did nothing. Below 0.05 we call the effect real at 95% confidence. η² (eta-squared): how much of the total variance the factor explains — 0.01 small, 0.06 medium, 0.14+ large. R²: variance explained by the whole model; a low R² means most of the outcome is path-dependent noise, so size smaller.
4 · Worked example — Iron Condor
Take five factors (strike distance, DTE, IV rank, short delta, gamma regime) at three levels each. Full factorial = 243 combinations. The Taguchi L18 array runs 18 of them and stays balanced. ANOVA on those 18 results might return DTE p=0.012 with η²=0.44 (real, large), IV rank p=0.031 (real, medium), and short delta p=0.48 (noise). Conclusion: fix DTE at 14 and respect IV rank; stop optimizing delta.
DoE + ANOVA vs traditional backtesting
| Traditional grid backtest | DoE + ANOVA | |
|---|---|---|
| Combinations tested | 243 full-factorial runs | 18 orthogonal runs (L18) |
| Compute time | Hours of grid backtesting | Seconds — ~93% fewer runs |
| Best parameter chosen by | Highest raw return | Highest S/N ratio, then F test |
| Noise protection | None — the peak is kept | Factors failing p < 0.05 are ignored |
| Reported confidence | Not available | 95% CI, p-value, η² and R² |
FAQ
- What is a good p-value for a trading factor?
- Below 0.05 at minimum (95% confidence). For live capital allocation many desks demand p < 0.01, especially when the number of experimental runs is small.
- Can a factor be significant but useless?
- Yes. A tiny η² with a low p-value means the effect is real but economically irrelevant. Always read significance and effect size together.
- Why one-way ANOVA per factor instead of a full model?
- Taguchi orthogonal arrays are balanced by construction, so each factor's levels are tested against the same mix of other levels. That makes a per-factor one-way F test a valid and far more interpretable screen.