3.9 Counting the Models Trained by Grid Search and Cross-Validation
Key Takeaways
The number of parameter combinations is the product of the candidate-value counts across every tuned hyperparameter.
With K-fold cross-validation, total fits are
combinations × K; scikit-learn'sGridSearchCVthen adds one final refit on the full training set whenrefit=True.Spark ML's
CrossValidatorfitsnumFolds × len(paramGrid)models, then refits the best combination once on the entire training set.TrainValidationSplitfitslen(paramGrid)models, one per combination, because there is a single validation split.This arithmetic is what makes a grid over five hyperparameters with 5-fold CV computationally impossible — the cost is multiplicative, not additive.
3.9 Counting the Models Trained by Grid Search and Cross-Validation
One exam objective is purely arithmetic: given a parameter grid and a validation scheme, state how many models are trained. It is easy marks, provided you apply the formula rather than guessing.
The Formula
Step 1 — count the combinations. For hyperparameters with candidate values each, the grid contains
combinations. A grid is a Cartesian product, so the counts multiply.
Step 2 — multiply by the validation scheme.
| Scheme | Models fitted |
|---|---|
| Single train/validation split | |
| -fold cross-validation | |
| -fold CV with a final refit on all training data | |
| Repeated -fold, repeats |
Worked Example 1 — scikit-learn GridSearchCV
An SVM is tuned with GridSearchCV, 5-fold cross-validation, and this grid:
C:[0.1, 1, 10]→ 3 valueskernel:['linear', 'rbf']→ 2 valuesgamma:[0.01, 0.1, 1]→ 3 values
90 models are trained during the search. With refit=True (the default),
GridSearchCV then trains one more model — the best combination on the entire training
set — for a total of 91 fits. When a question asks how many models the search trains,
the answer is 90; when it asks how many fits occur in total including the refit, it is
91. Read the wording carefully.
Worked Example 2 — Spark ML CrossValidator
param_grid = (ParamGridBuilder()
.addGrid(lr.regParam, [0.01, 0.1, 1.0]) # 3
.addGrid(lr.elasticNetParam, [0.0, 0.5, 1.0]) # 3
.build()) # 9 combinations
crossval = CrossValidator(estimator=pipeline,
estimatorParamMaps=param_grid,
evaluator=evaluator,
numFolds=5)
CrossValidator then refits the winning combination on the full training dataset, so
46 fits occur in total and cv_model.bestModel is that final refit.
Switching to TrainValidationSplit(trainRatio=0.8) with the same grid fits 9
models — one per combination — plus the final refit.
Why the Arithmetic Matters
Consider tuning XGBoost over five hyperparameters with four candidate values each, under 5-fold cross-validation:
At two minutes per fit that is roughly 170 hours of sequential compute. This is the
concrete reason the exam pairs this objective with the search-strategy objective: the
multiplicative blow-up is what motivates random and Bayesian search, and it is why
TrainValidationSplit exists for very large data.
A budgeting checklist
- Multiply the grid sizes to get .
- Multiply by (and by repeats, if any).
- Multiply by the average fit time.
- Divide by the tuner's
parallelismfor an approximate wall-clock estimate. - If the number is unacceptable, cut the grid, cut , or switch to Bayesian search — not the fold count alone, since makes the estimate noisy again.
Common trap:
parallelismchanges how long the search takes, never how many models are trained. A question that addsparallelism=4to a 45-model grid is testing whether you conflate the two — the answer is still 45.
Variations That Change the Count
Several details alter the arithmetic, and each one appears in question stems.
Conditional grids
scikit-learn accepts a list of grid dictionaries, and the totals add rather than multiply, because parameters that do not apply to a kernel are simply not enumerated for it:
param_grid = [
{"kernel": ["linear"], "C": [0.1, 1, 10]}, # 3 combinations
{"kernel": ["rbf"], "C": [0.1, 1, 10], "gamma": [0.01, 0.1]}, # 6 combinations
] # 9 in total
Nine combinations under 5-fold cross-validation is 45 fits. A flat product over the
union of the keys — kernel (2) × C (3) × gamma (2) — would count 12 combinations
and 60 fits, over-counting the three linear configurations that gamma does not
apply to.
Early stopping does not reduce the count
Early stopping shortens each individual fit by halting boosting rounds; it does not remove any combination from the grid. The number of models trained is unchanged.
Nested cross-validation multiplies again
An outer loop of folds wrapped around an inner tuning loop of folds trains models. With , , and , that is fits — which is why nested CV is reserved for small datasets and publication-grade estimates.
Hyperopt is budgeted, not enumerated
fmin(..., max_evals=40) trains exactly 40 models regardless of how large the search
space is, because Bayesian and random search sample rather than enumerate. If each
trial internally runs its own -fold cross-validation, the total becomes
. This contrast — a grid whose cost is set by the space, versus a
Hyperopt run whose cost is set by the budget — is the reason a large space pushes
practitioners toward fmin.
A data scientist tunes a model with GridSearchCV using 5-fold cross-validation. The grid contains C with 3 values, kernel with 2 values, and gamma with 3 values. How many models are trained during the search?
18
15
90
30
A Spark ML CrossValidator is configured with numFolds=3, a parameter grid of 12 combinations, and parallelism=4. How many models does it fit during cross-validation?
12, because parallelism divides the work across four executors
36, because each of the 12 combinations is fitted once per fold
9, because parallelism reduces the effective grid size
48, because parallelism multiplies the fold count
A team must tune 5 hyperparameters with 4 candidate values each. Under 5-fold cross-validation, roughly how many model fits does an exhaustive grid require, and what does that imply?
100 fits, so an exhaustive grid is clearly affordable
1,024 fits, matching the number of grid combinations
5,120 fits (4^5 = 1,024 combinations × 5 folds), which is usually prohibitive and motivates random or Bayesian search instead
20 fits, because the hyperparameters are evaluated independently
Sections you finish are checked off in the contents.