Model comparison
Model comparison evaluates alternative mathematical descriptions of the same biological system. The aim is not simply to identify which model follows one dataset most closely, but to compare fit, complexity, predictive performance, uncertainty and biological usefulness.
Competing biological models
Suppose two models are
\[M_1:\quad \frac{d\mathbf x}{dt}=\mathbf f_1(\mathbf x,\boldsymbol\theta_1),\]and
\[M_2:\quad \frac{d\mathbf x}{dt}=\mathbf f_2(\mathbf x,\boldsymbol\theta_2).\]They may differ in compartments, mechanisms, parameterisations, stochastic assumptions or observation models.
Comparison requires a common question
Models should be compared with respect to the same scientific purpose and relevant observations.
A model designed to estimate long-term equilibrium and a model designed to forecast short-term incidence need not be judged by exactly the same criterion.
Goodness of fit
For least-squares fitting, one simple quantity is the residual sum of squares
\[\mathrm{RSS}=\sum_i(y_i-\hat y_i)^2.\]A smaller RSS indicates closer in-sample fit when the same observations and scale are used.
But RSS alone always tends to favour additional flexibility.
Likelihood
For probabilistic models, compare maximised likelihoods
\[L_{\max}=L(\hat{\boldsymbol\theta};\mathbf y).\]A larger maximised likelihood means that the fitted model assigns greater likelihood to the observed data under its assumed observation distribution.
Likelihood alone does not penalise the number of fitted parameters.
Why complexity matters
Suppose model \(M_1\) has three adjustable parameters and \(M_2\) has twelve.
If \(M_2\) fits only slightly better, the improvement may reflect flexibility rather than a genuinely useful additional biological mechanism.
This motivates criteria that balance fit and complexity.
Akaike information criterion
The Akaike information criterion is
\[\boxed{\mathrm{AIC}=2k-2\ell_{\max}},\]where \(k\) is the number of estimated parameters and
\[\ell_{\max}=\log L_{\max}.\]Among models fitted to the same data under comparable likelihood definitions, smaller AIC is preferred in the AIC sense.
Interpretation of AIC
AIC can be viewed as estimating relative expected information loss for predictive approximation under its theoretical framework.
It is not a hypothesis test and does not give the probability that a model is true.
AIC differences
Let
\[\mathrm{AIC}_{\min}=\min_m\mathrm{AIC}_m.\]Define
\[\boxed{\Delta_m=\mathrm{AIC}_m-\mathrm{AIC}_{\min}}.\]The best-ranked model has \(\Delta_m=0\). Increasing values indicate progressively less support relative to the best candidate under the AIC framework.
Akaike weights
Relative AIC support can be summarised by
\[\boxed{w_m=\frac{\exp(-\Delta_m/2)}{\sum_r\exp(-\Delta_r/2)}}.\]The weights sum to one across the specified candidate set.
They are relative model weights within that set and should not automatically be interpreted as posterior probabilities of model truth.
Small-sample correction
When sample size is not large relative to the number of estimated parameters, a corrected criterion is often used:
\[\boxed{\mathrm{AICc}=\mathrm{AIC}+\frac{2k(k+1)}{n-k-1}}.\]The usual expression requires \(n>k+1\).
As \(n\) becomes large relative to \(k\), AICc approaches AIC.
What counts as sample size?
For simple independent observations, \(n\) is usually straightforward.
For correlated longitudinal, spatial or hierarchical data, the effective information content may not correspond naively to the number of rows. Care is therefore needed when applying small-sample formulas.
Bayesian information criterion
The Bayesian information criterion is
\[\boxed{\mathrm{BIC}=k\log n-2\ell_{\max}}.\]For sufficiently large \(n\), its complexity penalty per parameter is stronger than AIC's when \(\log n>2\).
BIC arises from a different asymptotic motivation and should not be treated as merely another version of AIC.
AIC and BIC can disagree
AIC is oriented toward predictive information loss, whereas BIC is associated under particular assumptions with large-sample Bayesian model evidence and consistent selection when a true finite-dimensional candidate is present.
Because their goals differ, disagreement is not necessarily an error.
Comparable likelihoods are essential
AIC and BIC comparisons require likelihoods defined for the same observed data.
Comparing a model fitted to daily incidence with another criterion computed from cumulative counts may not be meaningful even if both describe the same epidemic.
Constants in log-likelihoods
When comparing models by AIC or BIC, likelihood calculations should be mutually consistent.
Dropping constants that differ between models or observation formulations can invalidate the comparison.
Nested models
Model \(M_0\) is nested within \(M_1\) if \(M_0\) can be obtained by restricting parameters of \(M_1\).
For example, a model with separate transmission rates \(\beta_1\) and \(\beta_2\) may reduce to a constant-rate model when
\[\beta_1=\beta_2.\]Likelihood-ratio test
For suitable nested regular models, define
\[\boxed{\Lambda=2\left[\ell_1(\hat{\boldsymbol\theta}_1)-\ell_0(\hat{\boldsymbol\theta}_0)\right]}.\]Under standard regularity conditions and the null model, \(\Lambda\) is asymptotically compared with a chi-squared distribution whose degrees of freedom equal the difference in parameter dimensions.
When the standard likelihood-ratio result can fail
The usual chi-squared approximation may fail when parameters lie on boundaries, models are non-identifiable, regularity conditions fail or models are not nested.
In such cases specialised theory or simulation-based calibration of the test statistic may be required.
Non-nested models
Many biologically plausible alternatives are not nested.
Information criteria and out-of-sample predictive comparison are often more natural than a standard likelihood-ratio test in this setting.
Out-of-sample prediction
A direct way to compare models is to fit them using calibration data and evaluate predictions on held-out data.
If model \(M_1\) repeatedly predicts unseen observations better than \(M_2\), this is evidence in favour of \(M_1\) for that predictive task.
Cross-validation
Cross-validation repeats model fitting and evaluation across different held-out subsets.
For time-dependent biological data, the split should preserve temporal order when forecasting is the intended use.
Using random folds for a forecasting problem can leak future information into training.
Point-prediction comparison
Models can be compared using validation quantities such as
\[\operatorname{MAE},\qquad\operatorname{RMSE}.\]The metric should reflect the scientific cost of different errors.
A model with the smallest RMSE is not automatically best for a decision based on threshold probabilities.
Probabilistic predictive comparison
When models produce predictive distributions, compare the probability assigned to unseen observations rather than only point estimates.
One quantity is the held-out log predictive score
\[\boxed{\sum_{i\in\mathrm{val}}\log p(y_i\mid\mathcal D_{\mathrm{cal}},M)}.\]Higher predictive log score indicates that the model assigned greater probability to the observations that occurred.
Calibration and sharpness
A useful probabilistic model should produce uncertainty distributions that are well calibrated but not unnecessarily broad.
Comparing only predictive means ignores this important part of model performance.
Mechanistic comparison
Two models can have similar predictive performance while representing biology very differently.
Comparison should therefore examine assumptions such as homogeneous mixing, latent periods, density dependence, age structure, spatial movement or environmental forcing.
Parameter identifiability
A model may achieve a slightly better information criterion while introducing parameters that the data cannot separately identify.
This weakens mechanistic interpretation even if the numerical fit improves.
Identifiability should therefore be considered alongside statistical model comparison.
Parameter plausibility
A statistically competitive model whose fitted parameters imply biologically impossible or implausible rates requires investigation.
Good fit does not rescue an incoherent mechanistic interpretation.
Parsimony
Parsimony means using no more complexity than is justified by the scientific purpose and evidence.
It does not mean that the model with the fewest parameters should always be selected.
An additional mechanism is justified when it meaningfully improves explanation, prediction or decision relevance.
Overfitting
Overfitting occurs when a model adapts too closely to features of the calibration data that do not generalise.
A characteristic pattern is excellent in-sample fit but poorer performance on independent observations.
Underfitting
A model can also be too simple.
Systematic validation errors, failure to reproduce known biological behaviour or persistent residual structure can indicate that an important mechanism is missing.
Model selection uncertainty
If several models have similar support, selecting one and ignoring all others can understate uncertainty.
Scientific conclusions should reflect the fact that more than one model structure may remain plausible.
Model averaging
Under an explicit framework, predictions can be combined across candidate models:
\[\boxed{\hat Y=\sum_m w_m\hat Y_m},\qquad\sum_mw_m=1.\]The weights may come from AIC-based methods, Bayesian posterior model probabilities or another justified scheme.
The interpretation depends on how those weights were constructed.
Bayesian model comparison
Bayesian approaches can compare models through marginal likelihoods
\[p(\mathbf y\mid M)=\int p(\mathbf y\mid\boldsymbol\theta,M)p(\boldsymbol\theta\mid M)\,d\boldsymbol\theta.\]The ratio of marginal likelihoods gives a Bayes factor.
Unlike maximised likelihood, the marginal likelihood integrates over parameter uncertainty and depends on the prior distributions.
Bayes factor
For models \(M_1\) and \(M_2\),
\[\boxed{BF_{12}=\frac{p(\mathbf y\mid M_1)}{p(\mathbf y\mid M_2)}}.\]A Bayes factor measures relative evidence under the specified models and priors. It is not itself the posterior probability of either model.
Models can all be wrong
Model comparison is relative to the candidate set.
The highest-ranked model can still be scientifically inadequate if every candidate omits an important mechanism or has a poor observation model.
Comparison across stochastic and deterministic models
A deterministic and a stochastic model should be compared using observation models and predictive quantities that make their likelihoods or validation scores genuinely comparable.
Comparing one stochastic simulation with one deterministic trajectory does not constitute a meaningful model comparison.
Decision-focused comparison
The best model for a decision is the one that provides adequate information for that decision, not necessarily the one with the smallest generic error.
If the decision concerns hospital capacity, compare how accurately models estimate quantities such as
\[P(H_{\max}>H_{\mathrm{capacity}}).\]A model with slightly worse case-count RMSE may still be more useful if it represents uncertainty in capacity exceedance more realistically.
Robust conclusions across models
If several plausible models lead to the same biological conclusion, that conclusion may be more robust to structural uncertainty.
If conclusions change sharply with model structure, model uncertainty should be reported rather than hidden by choosing one preferred model.
Reporting a model comparison
State the candidate models, data used, observation models, estimated parameter counts, comparison criterion, validation design and biological assumptions.
Report actual criterion differences or predictive scores rather than only saying that one model was “better”.
A practical comparison workflow
Define the scientific purpose, specify a plausible candidate set, fit each model consistently to the same relevant data, assess identifiability and diagnostics, compare information criteria where appropriate, evaluate out-of-sample predictions, inspect biological plausibility, quantify model-selection uncertainty and determine whether the scientific conclusion is robust across plausible models.
Completing data and parameter estimation
The full workflow now runs from preparing biological data through observation models, parameter estimation, uncertainty, identifiability, sensitivity, calibration, validation and finally comparison of competing models.
The next major learning path develops the statistical foundations that support these ideas more deeply.