Model calibration
Model calibration connects a mathematical model to observed biological data by choosing unknown quantities so that model-generated observations are compatible with the real observations under a stated fitting procedure.
Mechanistic model
Suppose the biological state satisfies
\[\boxed{\frac{d\mathbf x}{dt}=\mathbf f(\mathbf x,t,\boldsymbol\theta)},\]with initial condition
\[\mathbf x(t_0)=\mathbf x_0.\]The parameter vector \(\boldsymbol\theta\) may contain rates, probabilities or other biological quantities.
Observation model
The model state is not necessarily the measured quantity. Predicted observations can be written
\[\boxed{\mu_i=h(\mathbf x(t_i;\boldsymbol\theta),\boldsymbol\theta)}.\]The function \(h\) connects latent biological states to the data.
For example, a model may contain the number infectious at each time while the dataset records new reported cases during intervals.
Calibration target
The first practical question is therefore: what exactly is being fitted?
The target might be prevalence, incidence, mortality, biomass, cell concentration, proportions or several variables simultaneously.
The model output and the data must represent the same biological quantity, time interval, population and units.
State versus incidence
Suppose \(C(t)\) is cumulative infections. If observations are new infections during intervals, the corresponding model quantity is
\[\boxed{\Delta C_i=C(t_i)-C(t_{i-1})}.\]Fitting the cumulative state directly to interval incidence would compare different quantities and can invalidate the calibration.
Choose unknown quantities
Calibration may estimate
\[\boldsymbol\theta=(\theta_1,\ldots,\theta_p)^T,\]but unknown initial conditions, reporting fractions, dispersion parameters or observation-error parameters may also need to be included.
Not every uncertain quantity should automatically be estimated from the same dataset.
Fixing known quantities
Parameters supported by strong independent evidence may be fixed rather than estimated.
This can reduce dimensionality and confounding, but fixed values should have a scientific justification. Fixing a poorly known parameter merely hides its uncertainty.
Parameter bounds
Biological restrictions can define the admissible parameter space:
\[\boldsymbol\theta\in\Theta.\]Examples include positive rates and probabilities restricted to \([0,1]\).
Bounds should reflect biological knowledge rather than being chosen solely to force an optimiser toward a preferred answer.
Define the fitting criterion
For least squares, calibration might minimise
\[\boxed{J(\boldsymbol\theta)=\sum_i[y_i-\mu_i(\boldsymbol\theta)]^2}.\]For likelihood-based calibration, it might maximise
\[\boxed{L(\boldsymbol\theta)=p(\mathbf y\mid\boldsymbol\theta)}.\]The fitting criterion should follow the observation process and data type.
Calibration is an inverse problem
The forward model maps parameters to observations:
\[\boldsymbol\theta\longrightarrow\mathbf x(t;\boldsymbol\theta)\longrightarrow\boldsymbol\mu(\boldsymbol\theta).\]Calibration works in the reverse inferential direction:
\[\mathbf y\longrightarrow\text{plausible values of }\boldsymbol\theta.\]This inverse problem may be much less stable than the forward simulation.
Numerical calibration loop
A typical deterministic calibration repeatedly performs four operations:
choose candidate parameters, solve the model, evaluate predictions at observation times, and calculate the objective or likelihood.
An optimisation algorithm then proposes new candidate values and repeats the process.
Interpolation at observation times
The numerical solver may evaluate the model at internal times different from the observation times.
Predictions should be obtained accurately at the actual measurement times rather than comparing each observation with an arbitrary nearby solver step.
Irregular observation times
Biological observations do not need to be equally spaced for calibration.
If measurements occur at \(t_1,t_2,\ldots,t_n\), the model should be evaluated at those times and the observation model should reflect the actual sampling scheme.
Missing observations
A missing measurement does not usually imply a biological value of zero.
The calibration should use the observations that genuinely exist or explicitly model the missing-data mechanism when necessary.
Multiple data streams
A model may be calibrated simultaneously to several types of data, such as cases, hospital occupancy and deaths.
A joint likelihood can be written schematically as
\[p(\mathbf y^{(1)},\mathbf y^{(2)},\mathbf y^{(3)}\mid\boldsymbol\theta).\]Each data stream may require its own observation model and scale.
Why simply adding squared errors can fail
If one dataset contains values around \(10^6\) and another around \(10^2\), raw squared errors from the larger-scale dataset may dominate the objective.
Appropriate weighting or a probabilistic observation model prevents numerical units from silently deciding which dataset matters most.
Time-varying parameters
Biological rates may change after interventions, seasonal shifts or behavioural changes.
A piecewise model might use
\[\beta(t)=\begin{cases}\beta_1,&tOverparameterisation
Adding more adjustable parameters generally makes it easier to fit calibration data.
But excessive flexibility can produce unstable estimates, strong parameter correlation and poor performance outside the calibration dataset.
Calibration quality should therefore not be judged by fit alone.
Structural identifiability before calibration
Whenever practical, determine whether the chosen model and observation scheme can theoretically distinguish the parameters.
If two parameters are structurally non-identifiable, no optimiser or larger sample of the same observations can recover unique values.
Practical identifiability after fitting
Even structurally identifiable parameters may be weakly informed by the available data.
Profile likelihoods, bootstrap distributions, confidence intervals and parameter correlations can reveal whether the fitted values are actually constrained.
Initial guesses
Nonlinear calibration can depend on the optimiser's starting point.
Using multiple biologically plausible initial guesses helps detect local minima or maxima and parameter ridges.
Convergence from one starting value is not sufficient evidence that the global optimum has been found.
Optimiser convergence
An optimiser reporting “success” means its numerical stopping conditions were satisfied.
It does not mean the biological model is correct, the parameters are identifiable, or the optimum is unique.
Solver accuracy
For ODE calibration, numerical integration error should be small compared with observational discrepancy.
Repeating the fit with tighter solver tolerances is one way to check whether numerical approximation materially affects the parameter estimates.
Residual diagnostics
For continuous-data fits, residuals
\[r_i=y_i-\mu_i(\hat{\boldsymbol\theta})\]should be inspected for systematic patterns.
Long runs of one sign, trends through time or variance increasing with fitted values can indicate an inadequate observation model or missing biological mechanism.
Diagnostics for count data
For Poisson or negative-binomial calibration, raw residuals may be difficult to compare because their variance depends on the mean.
Standardised, Pearson, deviance or simulation-based residual diagnostics can be more informative depending on the model.
Visual fit
Plotting observations together with fitted predictions is essential, but a visually close curve is only one diagnostic.
A model can look convincing while parameter estimates remain non-identifiable or the observation distribution is badly misspecified.
Calibration uncertainty
After finding \(\hat{\boldsymbol\theta}\), propagate uncertainty rather than presenting the fitted trajectory as exact.
Relevant outputs may include parameter intervals, trajectory uncertainty, predictive distributions or probabilities of biologically important events.
Calibration data versus validation data
Calibration data are used to choose parameters. Validation data are held apart from that fitting step and used to assess performance on information not used to determine the fitted parameters.
Agreement with calibration data is therefore not independent evidence of predictive ability.
Avoiding data leakage
Information from a validation dataset should not influence parameter fitting, model tuning, preprocessing decisions or repeated model selection if that dataset is intended to provide an independent evaluation.
Otherwise the apparent validation performance becomes optimistic.
Temporal calibration
For forecasting, one useful design fits the model using observations only up to time \(t_c\), then predicts later observations.
This respects the direction of time and tests the model in a way closer to its intended forecasting use.
Rolling-origin evaluation
A model can be recalibrated at several historical cut-off times and asked to predict subsequent periods.
This gives a more informative picture of forecasting performance than evaluating a single arbitrary split.
Calibration window
Parameter estimates can depend on the time period used for fitting, especially when biological processes or interventions change through time.
The calibration window should therefore be stated and justified.
Calibration under stochastic models
For a stochastic model, one parameter vector does not generate one deterministic trajectory. It generates a probability distribution over possible trajectories.
Calibration should therefore compare observed data with the stochastic model's probability structure rather than simply matching one random simulation to the observations.
Simulation-based calibration objectives
When an exact likelihood is unavailable, repeated simulations may be used to estimate summary statistics, approximate likelihoods or discrepancies.
Because such objectives contain Monte Carlo noise, optimisation and uncertainty assessment require additional care.
Calibration and sensitivity
Sensitivity analysis can show which parameters most strongly affect the calibrated outputs.
But a sensitive parameter is not necessarily identifiable, and an insensitive parameter may be impossible to estimate precisely from the chosen data.
Calibration and uncertainty quantification
Calibration produces fitted or plausible parameter values. Uncertainty quantification propagates uncertainty in those values and other inputs into biological outputs.
The two stages are connected but should not be confused.
Calibration and validation
Calibration asks whether parameter values can make the model compatible with selected observations.
Validation asks whether the calibrated model performs adequately for observations or situations not used to fit it.
A model can calibrate well and validate poorly.
Reproducible calibration
A reproducible analysis should record the data version, preprocessing, model equations, parameter definitions and units, fixed values, bounds, initial conditions, objective or likelihood, solver settings, optimisation method and random seeds when simulation is involved.
Without these details, reproducing a fitted parameter set may be impossible.
A complete calibration workflow
Define the biological question and data, verify units and observation definitions, specify the mechanistic and observation models, choose estimable parameters, check structural identifiability, define bounds and the fitting criterion, run the numerical fit from multiple starting values, inspect diagnostics, assess practical identifiability and uncertainty, and only then evaluate performance on independent information.
Transition to model validation
Calibration determines parameter values using selected observations. The next lesson asks a different question: after calibration, does the model reproduce independent biological behaviour well enough for its intended purpose?