Model validation
Model validation evaluates whether a calibrated model performs adequately for a stated biological purpose. It asks how well the model reproduces relevant observations or behaviours that were not simply forced by the fitting procedure.
Calibration and validation are different
Calibration chooses unknown quantities using selected data:
\[\mathbf y_{\mathrm{cal}}\longrightarrow\hat{\boldsymbol\theta}.\]Validation then evaluates the fitted model using information not used to determine those parameters whenever possible:
\[\hat{\boldsymbol\theta}\longrightarrow\widehat{\mathbf y}_{\mathrm{val}}\longleftrightarrow\mathbf y_{\mathrm{val}}.\]Good agreement with calibration data alone is not independent validation.
Why independence matters
If the same observations are repeatedly used to choose parameters, select model structure and assess performance, the evaluation becomes optimistic.
Independent or held-out information provides a more demanding test of whether the fitted model generalises beyond the data used to construct it.
Validation is purpose-specific
A model can be adequate for estimating epidemic peak timing but inadequate for predicting daily case counts.
Another model may reproduce population averages well but fail to represent extinction probabilities.
The intended scientific use should therefore be stated before deciding what successful validation means.
Validation target
Choose quantities directly connected to the intended use. Possible targets include
\[I(t),\qquad I_{\max},\qquad t_{\mathrm{peak}},\qquad R_0,\qquad P(Y>c).\]A model intended to support hospital-capacity decisions should be evaluated on quantities related to demand and threshold exceedance, not only on infection-curve fit.
Types of independent information
Validation evidence can come from a later time period, another experiment, another location, another population, another biological condition or a data stream not used for fitting.
The strongest choice depends on how the model is intended to be used.
Temporal hold-out validation
For forecasting, fit the model using observations up to a cut-off time \(t_c\):
\[\mathcal D_{\mathrm{cal}}=\{y_t:t\le t_c\}.\]Then evaluate predictions using later observations:
\[\mathcal D_{\mathrm{val}}=\{y_t:t>t_c\}.\]This prevents future observations from influencing a historical forecast.
Rolling-origin validation
A single temporal split can depend strongly on the chosen cut-off.
Rolling-origin evaluation repeats the process at several cut-offs:
\[t_c^{(1)},t_c^{(2)},\ldots,t_c^{(K)}.\]At each origin, only information available up to that time is used to fit or update the model before predicting the future.
Random train-test splits
Randomly splitting observations can be useful when observations are approximately exchangeable.
It is often inappropriate for forecasting time series because observations from the future can leak information into the fitted model.
Cross-validation
In \(K\)-fold cross-validation, observations are divided into \(K\) subsets. Each subset is held out in turn while the model is fitted to the others.
The resulting out-of-sample performance is combined across folds.
The splitting scheme must respect dependence, clustering and time order in the data.
Leave-one-out validation
Leave-one-out cross-validation is the special case in which one observational unit is held out at a time.
It can be computationally expensive for mechanistic models requiring numerical optimisation for every refit.
Grouped validation
If measurements are clustered within patients, sites or experimental replicates, individual rows should not necessarily be split independently.
Holding out complete groups can better test generalisation to new individuals or experimental units.
External validation
External validation evaluates a model using data from a genuinely different source, population, setting or study.
This can provide stronger evidence of transportability than repeatedly partitioning one dataset, although differences between settings must be interpreted biologically.
Prediction error
For continuous observations, a prediction error is
\[e_i=y_i-\hat y_i.\]Common summaries include
\[\boxed{\operatorname{MAE}=\frac1n\sum_i|e_i|}\]and
\[\boxed{\operatorname{RMSE}=\sqrt{\frac1n\sum_i e_i^2}}.\]RMSE penalises large errors more strongly than MAE.
Scale matters
MAE and RMSE retain the units of the outcome. This aids interpretation but makes direct comparisons across differently scaled outcomes difficult.
Normalisation can help, but the definition of the normalisation should always be stated.
Relative errors
Relative error measures can be useful when proportional accuracy matters, but expressions such as
\[\frac{|y_i-\hat y_i|}{|y_i|}\]become unstable or undefined when observed values are near zero.
No single error metric is appropriate for every biological dataset.
Residual structure on validation data
Validation errors should be inspected through time and against predicted values or relevant covariates.
Systematic patterns may reveal missing biological mechanisms even when an overall error score appears acceptable.
Bias
The mean prediction error
\[\overline e=\frac1n\sum_i(y_i-\hat y_i)\]can indicate systematic over- or under-prediction.
A mean near zero is not sufficient because positive and negative structured errors can cancel.
Probabilistic predictions
If a model produces a predictive distribution
\[p(y_{\mathrm{new}}\mid\mathcal D_{\mathrm{cal}}),\]validation should assess the distribution rather than only its mean.
A probabilistic forecast can be poor even when its point prediction is close to the observation.
Prediction-interval coverage
Suppose a nominal 95% prediction interval is
\[[L_i,U_i].\]Empirical coverage is
\[\boxed{\widehat C=\frac1n\sum_{i=1}^n\mathbf1\{L_i\le y_i\le U_i\}}.\]Coverage much below the nominal level suggests uncertainty is underestimated, while excessive coverage may indicate intervals that are unnecessarily wide.
Coverage is not enough
A model could achieve high coverage simply by producing extremely wide intervals.
Probabilistic validation should therefore consider both calibration of uncertainty and sharpness or informativeness of the predictive distribution.
Calibration of predicted probabilities
Suppose the model repeatedly predicts an event with probability approximately 0.7.
For a well-calibrated probability forecast, the event should occur approximately 70% of the time across comparable cases, subject to sampling variation.
This is a different meaning of “calibration” from parameter calibration.
Brier score
For binary outcomes \(y_i\in\{0,1\}\) with predicted probabilities \(p_i\), the Brier score is
\[\boxed{\operatorname{BS}=\frac1n\sum_i(p_i-y_i)^2}.\]Lower values indicate better probability predictions, but interpretation should consider the event frequency and appropriate benchmarks.
Log predictive score
A probabilistic model can also be evaluated using the log probability assigned to observed validation outcomes:
\[\boxed{\frac1n\sum_i\log p(y_i\mid\mathcal D_{\mathrm{cal}})}.\]This rewards distributions that assign high probability density or mass to what actually occurs.
Stochastic model validation
A stochastic model predicts a distribution of possible trajectories rather than one path.
It should therefore not be rejected merely because one simulated trajectory differs from the observed trajectory.
Validation should ask whether observed summaries and trajectories are plausible under the model's predictive distribution.
Distributional checks
For stochastic models, compare quantities such as means, variances, quantiles, extinction frequencies, event times, peak distributions or autocorrelation structures when scientifically relevant.
The checks should correspond to features the model is intended to reproduce.
Posterior predictive checking
In Bayesian modelling, generate replicated data
\[\mathbf y^{\mathrm{rep}}\sim p(\mathbf y^{\mathrm{rep}}\mid\mathbf y)\]and compare biologically meaningful summaries of the replicated datasets with those of the observed data.
This checks whether the fitted model can reproduce important features of the observations, although it is not the same as fully independent out-of-sample validation.
Mechanistic plausibility
Numerical predictive accuracy is not the only consideration in mathematical biology.
A fitted model should also be checked for impossible states, unrealistic rates, biologically implausible parameter values and qualitative behaviour inconsistent with established knowledge.
Conservation and constraints
If a model assumes a constant population, validate that numerical solutions satisfy
\[S(t)+E(t)+I(t)+R(t)=N\]to the expected numerical accuracy.
States representing populations should also remain non-negative when the mathematical and numerical formulation requires this.
Extreme and limiting cases
Useful validation includes checking whether the model behaves correctly in simple limiting cases.
For example, if transmission is set to zero, an epidemic model should not generate new infections through the transmission mechanism.
Such checks can reveal implementation errors even before comparison with data.
Validation across regimes
A model calibrated during rapid epidemic growth may fail during decline, after an intervention or in another season.
Testing across biologically distinct regimes can reveal where assumptions cease to be adequate.
Interpolation versus extrapolation
Predicting within conditions represented in the calibration data is interpolation.
Predicting far beyond observed times, parameter ranges, populations or environmental conditions is extrapolation and generally carries greater risk.
Validation evidence from interpolation should not automatically be interpreted as evidence for distant extrapolation.
Parameter stability
If a supposedly constant biological parameter changes dramatically whenever the calibration window changes, this may indicate weak identifiability, model misspecification or genuine time variation.
Parameter stability across relevant datasets is therefore a useful diagnostic, although exact equality should not be expected.
Robustness of conclusions
Sometimes the key question is not whether every trajectory is accurately predicted but whether a scientific conclusion remains unchanged under plausible uncertainty.
For example, does an intervention still reduce the probability of exceeding a hospital-capacity threshold across plausible parameter values and model assumptions?
Validation thresholds
There is no universal RMSE, coverage value or correlation that makes a mathematical biology model “valid”.
Acceptable performance depends on the intended decision, measurement precision, biological variability and available alternatives.
Benchmark models
A complex mechanistic model should often be compared with simpler reference predictions.
If it cannot outperform a scientifically reasonable baseline for the intended task, its additional complexity may not be justified by predictive performance alone.
Failure is informative
Validation failure can reveal which assumptions need revision.
Systematic errors may motivate a time-varying rate, an additional compartment, a different observation process or a stochastic formulation.
Validation is therefore part of model development rather than merely a final pass-or-fail step.
Avoiding repeated test-set tuning
If a held-out dataset is examined repeatedly while the model is changed to improve performance on it, that dataset gradually becomes part of model development.
A further independent evaluation may then be required for an unbiased final assessment.
Documenting validation
Report which data were used for calibration, which were used for validation, how splits were constructed, what preprocessing was performed, which metrics were chosen and why, and whether uncertainty was included in predictions.
This makes the strength and limitations of the validation evidence clear.
A practical validation workflow
State the intended use, choose validation targets and relevant independent data, generate predictions without using those outcomes for fitting, compare point and probabilistic predictions with observations, inspect systematic failures, test biological constraints and limiting cases, assess robustness across uncertainty and regimes, and compare performance with suitable benchmarks.
Transition to model comparison
Validation asks whether one model is adequate for its intended purpose. Often several plausible mathematical models are available. The next lesson develops model comparison: deciding what evidence can support choosing between competing models without relying only on goodness of fit.