โ† Machine Learning and Mathematical Biology

Training and testing data

Training data are used to fit model parameters. Test data are kept separate and used to estimate how well the fitted model performs on unseen observations.

Why separation matters

Evaluating a model on the same data used for training can produce an overly optimistic estimate of predictive performance.

Validation data

A validation set, or cross-validation within the training data, can be used for selecting model complexity and hyperparameters. The final test set should not repeatedly guide those choices.

Data leakage

Information from test observations must not enter feature construction, scaling, model selection or training. In time-dependent biological data, future observations must not leak into predictions of the past.

Key idea. A fair test imitates the intended prediction task by keeping evaluation information genuinely unseen during model development.