← Data and Parameter Estimation

Model validation

Model validation evaluates whether a calibrated model performs adequately for a stated biological purpose. It asks how well the model reproduces relevant observations or behaviours that were not simply forced by the fitting procedure.

Core idea. Validation does not prove that a mathematical model is true. It tests whether the model is sufficiently adequate for a particular use, under specified conditions and against relevant evidence.

Calibration and validation are different

Calibration chooses unknown quantities using selected data:

\[\mathbf y_{\mathrm{cal}}\longrightarrow\hat{\boldsymbol\theta}.\]

Validation then evaluates the fitted model using information not used to determine those parameters whenever possible:

\[\hat{\boldsymbol\theta}\longrightarrow\widehat{\mathbf y}_{\mathrm{val}}\longleftrightarrow\mathbf y_{\mathrm{val}}.\]

Good agreement with calibration data alone is not independent validation.

Why independence matters

If the same observations are repeatedly used to choose parameters, select model structure and assess performance, the evaluation becomes optimistic.

Independent or held-out information provides a more demanding test of whether the fitted model generalises beyond the data used to construct it.

Validation is purpose-specific

A model can be adequate for estimating epidemic peak timing but inadequate for predicting daily case counts.

Another model may reproduce population averages well but fail to represent extinction probabilities.

The intended scientific use should therefore be stated before deciding what successful validation means.

Validation target

Choose quantities directly connected to the intended use. Possible targets include

\[I(t),\qquad I_{\max},\qquad t_{\mathrm{peak}},\qquad R_0,\qquad P(Y>c).\]

A model intended to support hospital-capacity decisions should be evaluated on quantities related to demand and threshold exceedance, not only on infection-curve fit.

Types of independent information

Validation evidence can come from a later time period, another experiment, another location, another population, another biological condition or a data stream not used for fitting.

The strongest choice depends on how the model is intended to be used.

Temporal hold-out validation

For forecasting, fit the model using observations up to a cut-off time \(t_c\):

\[\mathcal D_{\mathrm{cal}}=\{y_t:t\le t_c\}.\]

Then evaluate predictions using later observations:

\[\mathcal D_{\mathrm{val}}=\{y_t:t>t_c\}.\]

This prevents future observations from influencing a historical forecast.

Rolling-origin validation

A single temporal split can depend strongly on the chosen cut-off.

Rolling-origin evaluation repeats the process at several cut-offs:

\[t_c^{(1)},t_c^{(2)},\ldots,t_c^{(K)}.\]

At each origin, only information available up to that time is used to fit or update the model before predicting the future.

Random train-test splits

Randomly splitting observations can be useful when observations are approximately exchangeable.

It is often inappropriate for forecasting time series because observations from the future can leak information into the fitted model.

Cross-validation

In \(K\)-fold cross-validation, observations are divided into \(K\) subsets. Each subset is held out in turn while the model is fitted to the others.

The resulting out-of-sample performance is combined across folds.

The splitting scheme must respect dependence, clustering and time order in the data.

Leave-one-out validation

Leave-one-out cross-validation is the special case in which one observational unit is held out at a time.

It can be computationally expensive for mechanistic models requiring numerical optimisation for every refit.

Grouped validation

If measurements are clustered within patients, sites or experimental replicates, individual rows should not necessarily be split independently.

Holding out complete groups can better test generalisation to new individuals or experimental units.

External validation

External validation evaluates a model using data from a genuinely different source, population, setting or study.

This can provide stronger evidence of transportability than repeatedly partitioning one dataset, although differences between settings must be interpreted biologically.

Prediction error

For continuous observations, a prediction error is

\[e_i=y_i-\hat y_i.\]

Common summaries include

\[\boxed{\operatorname{MAE}=\frac1n\sum_i|e_i|}\]

and

\[\boxed{\operatorname{RMSE}=\sqrt{\frac1n\sum_i e_i^2}}.\]

RMSE penalises large errors more strongly than MAE.

Scale matters

MAE and RMSE retain the units of the outcome. This aids interpretation but makes direct comparisons across differently scaled outcomes difficult.

Normalisation can help, but the definition of the normalisation should always be stated.

Relative errors

Relative error measures can be useful when proportional accuracy matters, but expressions such as

\[\frac{|y_i-\hat y_i|}{|y_i|}\]

become unstable or undefined when observed values are near zero.

No single error metric is appropriate for every biological dataset.

Residual structure on validation data

Validation errors should be inspected through time and against predicted values or relevant covariates.

Systematic patterns may reveal missing biological mechanisms even when an overall error score appears acceptable.

Bias

The mean prediction error

\[\overline e=\frac1n\sum_i(y_i-\hat y_i)\]

can indicate systematic over- or under-prediction.

A mean near zero is not sufficient because positive and negative structured errors can cancel.

Probabilistic predictions

If a model produces a predictive distribution

\[p(y_{\mathrm{new}}\mid\mathcal D_{\mathrm{cal}}),\]

validation should assess the distribution rather than only its mean.

A probabilistic forecast can be poor even when its point prediction is close to the observation.

Prediction-interval coverage

Suppose a nominal 95% prediction interval is

\[[L_i,U_i].\]

Empirical coverage is

\[\boxed{\widehat C=\frac1n\sum_{i=1}^n\mathbf1\{L_i\le y_i\le U_i\}}.\]

Coverage much below the nominal level suggests uncertainty is underestimated, while excessive coverage may indicate intervals that are unnecessarily wide.

Coverage is not enough

A model could achieve high coverage simply by producing extremely wide intervals.

Probabilistic validation should therefore consider both calibration of uncertainty and sharpness or informativeness of the predictive distribution.

Calibration of predicted probabilities

Suppose the model repeatedly predicts an event with probability approximately 0.7.

For a well-calibrated probability forecast, the event should occur approximately 70% of the time across comparable cases, subject to sampling variation.

This is a different meaning of “calibration” from parameter calibration.

Brier score

For binary outcomes \(y_i\in\{0,1\}\) with predicted probabilities \(p_i\), the Brier score is

\[\boxed{\operatorname{BS}=\frac1n\sum_i(p_i-y_i)^2}.\]

Lower values indicate better probability predictions, but interpretation should consider the event frequency and appropriate benchmarks.

Log predictive score

A probabilistic model can also be evaluated using the log probability assigned to observed validation outcomes:

\[\boxed{\frac1n\sum_i\log p(y_i\mid\mathcal D_{\mathrm{cal}})}.\]

This rewards distributions that assign high probability density or mass to what actually occurs.

Stochastic model validation

A stochastic model predicts a distribution of possible trajectories rather than one path.

It should therefore not be rejected merely because one simulated trajectory differs from the observed trajectory.

Validation should ask whether observed summaries and trajectories are plausible under the model's predictive distribution.

Distributional checks

For stochastic models, compare quantities such as means, variances, quantiles, extinction frequencies, event times, peak distributions or autocorrelation structures when scientifically relevant.

The checks should correspond to features the model is intended to reproduce.

Posterior predictive checking

In Bayesian modelling, generate replicated data

\[\mathbf y^{\mathrm{rep}}\sim p(\mathbf y^{\mathrm{rep}}\mid\mathbf y)\]

and compare biologically meaningful summaries of the replicated datasets with those of the observed data.

This checks whether the fitted model can reproduce important features of the observations, although it is not the same as fully independent out-of-sample validation.

Mechanistic plausibility

Numerical predictive accuracy is not the only consideration in mathematical biology.

A fitted model should also be checked for impossible states, unrealistic rates, biologically implausible parameter values and qualitative behaviour inconsistent with established knowledge.

Conservation and constraints

If a model assumes a constant population, validate that numerical solutions satisfy

\[S(t)+E(t)+I(t)+R(t)=N\]

to the expected numerical accuracy.

States representing populations should also remain non-negative when the mathematical and numerical formulation requires this.

Extreme and limiting cases

Useful validation includes checking whether the model behaves correctly in simple limiting cases.

For example, if transmission is set to zero, an epidemic model should not generate new infections through the transmission mechanism.

Such checks can reveal implementation errors even before comparison with data.

Validation across regimes

A model calibrated during rapid epidemic growth may fail during decline, after an intervention or in another season.

Testing across biologically distinct regimes can reveal where assumptions cease to be adequate.

Interpolation versus extrapolation

Predicting within conditions represented in the calibration data is interpolation.

Predicting far beyond observed times, parameter ranges, populations or environmental conditions is extrapolation and generally carries greater risk.

Validation evidence from interpolation should not automatically be interpreted as evidence for distant extrapolation.

Parameter stability

If a supposedly constant biological parameter changes dramatically whenever the calibration window changes, this may indicate weak identifiability, model misspecification or genuine time variation.

Parameter stability across relevant datasets is therefore a useful diagnostic, although exact equality should not be expected.

Robustness of conclusions

Sometimes the key question is not whether every trajectory is accurately predicted but whether a scientific conclusion remains unchanged under plausible uncertainty.

For example, does an intervention still reduce the probability of exceeding a hospital-capacity threshold across plausible parameter values and model assumptions?

Validation thresholds

There is no universal RMSE, coverage value or correlation that makes a mathematical biology model “valid”.

Acceptable performance depends on the intended decision, measurement precision, biological variability and available alternatives.

Benchmark models

A complex mechanistic model should often be compared with simpler reference predictions.

If it cannot outperform a scientifically reasonable baseline for the intended task, its additional complexity may not be justified by predictive performance alone.

Failure is informative

Validation failure can reveal which assumptions need revision.

Systematic errors may motivate a time-varying rate, an additional compartment, a different observation process or a stochastic formulation.

Validation is therefore part of model development rather than merely a final pass-or-fail step.

Avoiding repeated test-set tuning

If a held-out dataset is examined repeatedly while the model is changed to improve performance on it, that dataset gradually becomes part of model development.

A further independent evaluation may then be required for an unbiased final assessment.

Validation data can be used up. Repeatedly adapting the model after seeing its validation errors weakens the independence of that evaluation.

Documenting validation

Report which data were used for calibration, which were used for validation, how splits were constructed, what preprocessing was performed, which metrics were chosen and why, and whether uncertainty was included in predictions.

This makes the strength and limitations of the validation evidence clear.

A practical validation workflow

State the intended use, choose validation targets and relevant independent data, generate predictions without using those outcomes for fitting, compare point and probabilistic predictions with observations, inspect systematic failures, test biological constraints and limiting cases, assess robustness across uncertainty and regimes, and compare performance with suitable benchmarks.

Transition to model comparison

Validation asks whether one model is adequate for its intended purpose. Often several plausible mathematical models are available. The next lesson develops model comparison: deciding what evidence can support choosing between competing models without relying only on goodness of fit.

Key idea. Model validation is evidence about adequacy for a particular purpose, not proof of truth. Strong validation combines independent predictive assessment, uncertainty evaluation, biological plausibility and transparent limits on where the model has actually been tested.