Cross-validation
Cross-validation estimates predictive performance by repeatedly fitting a model on part of the available training data and evaluating it on held-out observations.
k-fold cross-validation
The data are divided into \(k\) folds. Each fold is used once for validation while the remaining \(k-1\) folds are used for training.
If the validation loss from fold \(j\) is \(L_j\), the average is
\[\overline L=\frac1k\sum_{j=1}^{k}L_j.\]Biological dependence
Ordinary random folds may be inappropriate when observations are grouped by patient, site, family or time. Splits should respect the dependence structure and intended future use.
Key idea. Cross-validation is useful only when its data splits realistically represent the prediction problem.