โ† Machine Learning and Mathematical Biology

Cross-validation

Cross-validation estimates predictive performance by repeatedly fitting a model on part of the available training data and evaluating it on held-out observations.

k-fold cross-validation

The data are divided into \(k\) folds. Each fold is used once for validation while the remaining \(k-1\) folds are used for training.

If the validation loss from fold \(j\) is \(L_j\), the average is

\[\overline L=\frac1k\sum_{j=1}^{k}L_j.\]

Biological dependence

Ordinary random folds may be inappropriate when observations are grouped by patient, site, family or time. Splits should respect the dependence structure and intended future use.

Key idea. Cross-validation is useful only when its data splits realistically represent the prediction problem.