โ† Data and Parameter Estimation

Uncertainty quantification

Uncertainty quantification studies how incomplete knowledge and random variation affect model outputs, predictions and biological conclusions. Its purpose is not simply to attach error bars to parameters, but to determine how uncertain the quantities that matter scientifically actually are.

Core idea. A mathematical model maps uncertain inputs into uncertain outputs. Uncertainty quantification identifies important sources of uncertainty, represents them appropriately and propagates them to the biological quantities used for interpretation or decisions.

A model with uncertain inputs

Write an output of interest as

\[\boxed{Y=g(\boldsymbol\theta,\mathbf x_0,\mathbf u,\boldsymbol\varepsilon)}.\]

Here \(\boldsymbol\theta\) represents parameters, \(\mathbf x_0\) initial conditions, \(\mathbf u\) external inputs or interventions, and \(\boldsymbol\varepsilon\) random components.

Uncertainty in any of these can produce uncertainty in \(Y\).

Parameter uncertainty

Parameters estimated from finite noisy data are not known exactly.

Instead of propagating only a fitted value \(\hat{\boldsymbol\theta}\), we can propagate a distribution, confidence region, bootstrap sample or other representation of plausible parameter values.

Observation uncertainty

Measurements may contain reporting error, laboratory error, sampling variation or incomplete detection.

An observation model represents this distinction:

\[Y_i\mid X_i\sim p(y_i\mid X_i,\boldsymbol\theta).\]

The observed value \(Y_i\) is therefore not necessarily identical to the underlying biological state \(X_i\).

Process uncertainty

In a stochastic biological model, the state itself can evolve randomly even when all parameter values are fixed.

For example, two epidemic simulations with identical transmission and recovery parameters can follow different paths because infection and recovery events occur randomly.

Parameter uncertainty versus process variability

These are fundamentally different.

Parameter uncertainty means we do not know the true parameter values precisely. Process variability means the biological model predicts random outcomes even if the parameters were known exactly.

Do not combine these conceptually. More data may reduce parameter uncertainty, but it does not remove genuine stochastic variability in the biological process.

Initial-condition uncertainty

The true starting state may be imperfectly known. For example, the number of infectious individuals at the beginning of an epidemic may be uncertain.

Represent this as

\[\mathbf X(0)\sim p(\mathbf x_0)\]

or through a plausible range when a full probability distribution is unavailable.

Input uncertainty

Models can depend on external quantities such as temperature, treatment dose, contact behaviour or future intervention timing.

If these inputs are uncertain, fixing them at one value can understate uncertainty in the model output.

Structural uncertainty

Several scientifically plausible model structures may describe the same system.

For example, one epidemic model may assume homogeneous mixing while another includes age groups or spatial structure.

Uncertainty caused by choosing among such formulations is structural or model-form uncertainty.

Numerical uncertainty

Numerical solvers introduce approximation error through time discretisation, tolerances, finite differences or Monte Carlo sampling.

This error should be small enough that it does not materially affect the scientific uncertainty being studied.

A useful classification

SourceMeaning
parameterunknown values of model parameters
observationmeasurement, reporting or sampling variation
processintrinsic randomness in biological dynamics
initial conditionuncertainty about the starting state
inputuncertain external drivers or interventions
structuraluncertainty about the mathematical model form
numericalapproximation introduced by computation

Aleatory and epistemic uncertainty

A broad distinction is sometimes made between aleatory uncertainty, representing intrinsic variability, and epistemic uncertainty, representing incomplete knowledge.

Process stochasticity is often treated as aleatory, while uncertain parameters or model structure are often treated as epistemic.

The distinction is useful but not always absolute; its meaning should be stated in the modelling context.

Propagating parameter uncertainty

Suppose plausible parameter vectors are

\[\boldsymbol\theta^{(1)},\ldots,\boldsymbol\theta^{(M)}.\]

For each one, calculate

\[Y^{(m)}=g(\boldsymbol\theta^{(m)}).\]

The resulting collection

\[Y^{(1)},\ldots,Y^{(M)}\]

approximates the uncertainty in the output induced by the parameter sample.

Monte Carlo propagation

If

\[\boldsymbol\Theta\sim p(\boldsymbol\theta),\]

draw independent or appropriately sampled parameter vectors and evaluate the model repeatedly.

Output summaries can then estimate quantities such as

\[E[Y],\qquad\operatorname{Var}(Y),\qquad P(Y>c).\]

Output quantiles

If simulated outputs are ordered, empirical quantiles can describe the central range of propagated outcomes.

For example, the 2.5th and 97.5th percentiles form a central 95% simulation interval under the particular uncertainty distribution being propagated.

The interpretation depends on how the input samples were generated.

Confidence uncertainty is not automatically a probability distribution

A frequentist confidence region for parameters should not automatically be treated as though it were a posterior probability distribution.

Bootstrap samples, asymptotic sampling distributions and Bayesian posterior samples have different interpretations even though all can be propagated through a model.

Bayesian propagation

In Bayesian inference, parameter uncertainty is represented by the posterior

\[p(\boldsymbol\theta\mid\mathbf y).\]

Posterior samples can be propagated through the model:

\[\boldsymbol\theta^{(m)}\sim p(\boldsymbol\theta\mid\mathbf y),\qquad Y^{(m)}=g(\boldsymbol\theta^{(m)}).\]

This produces a posterior distribution for the derived model quantity.

Predictive distributions

A future observation generally depends on both parameter uncertainty and new observational or process variation.

In Bayesian notation, a posterior predictive distribution is

\[\boxed{p(y_{\mathrm{new}}\mid\mathbf y)=\int p(y_{\mathrm{new}}\mid\boldsymbol\theta)\,p(\boldsymbol\theta\mid\mathbf y)\,d\boldsymbol\theta}.\]

This distinction explains why prediction intervals are usually wider than uncertainty bands based only on fitted parameters.

Deterministic model with uncertain parameters

A deterministic ODE model produces one trajectory for each fixed parameter vector.

If the parameters are uncertain, however, the collection of possible deterministic trajectories forms an uncertainty band.

The band comes from uncertain inputs, not intrinsic stochasticity in the ODE itself.

Stochastic model with fixed parameters

A stochastic model produces a distribution of trajectories even for one fixed parameter vector.

Repeated simulations estimate process variability conditional on those parameters.

Stochastic model with uncertain parameters

When both are present, uncertainty propagation has two levels:

\[\boldsymbol\Theta\sim p(\boldsymbol\theta),\]

followed by

\[X(t)\mid\boldsymbol\Theta\sim p(x(t)\mid\boldsymbol\Theta).\]

This combines between-parameter uncertainty with within-parameter process variability.

Law of total variance

The two contributions can be understood through

\[\boxed{\operatorname{Var}(Y)=E[\operatorname{Var}(Y\mid\boldsymbol\Theta)]+\operatorname{Var}(E[Y\mid\boldsymbol\Theta])}.\]

The first term represents average conditional variability, while the second represents variation in the conditional mean caused by uncertain parameters.

Uncertain trajectories

For a time-dependent output \(Y(t)\), uncertainty can be summarised at each time using means, medians, quantiles or other appropriate summaries.

A pointwise 95% band means that the interval is constructed separately at each time. It is not automatically a simultaneous 95% band for the entire trajectory.

Pointwise versus simultaneous bands

A simultaneous band is designed to cover an entire function or trajectory with a specified coverage property.

Because this is a stronger requirement than pointwise coverage, simultaneous bands are generally wider.

Derived biological quantities

Uncertainty should be propagated to quantities used for scientific interpretation, such as

\[R_0,\qquad I_{\max},\qquad t_{\mathrm{peak}},\qquad P(\text{extinction}),\qquad P(\text{capacity exceeded}).\]

These may be more useful for decisions than intervals for individual parameters.

Threshold probabilities

If a decision depends on whether output \(Y\) exceeds a threshold \(c\), calculate

\[\boxed{P(Y>c)}.\]

For Monte Carlo samples, estimate it by

\[\boxed{\widehat P(Y>c)=\frac1M\sum_{m=1}^{M}\mathbf1\{Y^{(m)}>c\}}.\]

This converts uncertainty into a directly interpretable risk measure.

Monte Carlo error

The estimated probability itself has simulation error because only finitely many simulations are used.

If \(\hat p\) is estimated from \(M\) independent Bernoulli indicators, its approximate Monte Carlo standard error is

\[\boxed{\operatorname{SE}_{MC}(\hat p)\approx\sqrt{\frac{\hat p(1-\hat p)}{M}}}.\]

This numerical uncertainty is separate from biological uncertainty.

Convergence checks

Monte Carlo estimates should be checked as the number of samples increases.

If estimated means, quantiles or threshold probabilities continue to change substantially, the simulation sample may be too small.

Rare events

Naive Monte Carlo can require very many simulations to estimate very small probabilities accurately.

Specialised rare-event methods may be needed when the scientific question concerns extremely unlikely but important outcomes.

Correlation among uncertain parameters

Parameters estimated jointly are often correlated.

Sampling each parameter independently from its marginal interval can create implausible combinations and distort propagated uncertainty.

The joint uncertainty structure should be retained whenever possible.

Uncertainty and identifiability

Poorly identifiable parameters often produce broad or strongly correlated uncertainty regions.

However, uncertainty in a biologically important derived quantity can still be small if that quantity depends mainly on an identifiable parameter combination.

Uncertainty and sensitivity

Sensitivity describes how strongly an output responds to an input. Uncertainty describes how uncertain that input or output is.

A useful practical question is therefore not only

\[\text{How sensitive is }Y\text{ to }\theta_j?\]

but also

\[\text{How uncertain is }\theta_j?\]

Both determine how much that parameter contributes to output uncertainty.

Structural uncertainty

Suppose two plausible models \(M_1\) and \(M_2\) produce different predictions. Reporting uncertainty conditional only on \(M_1\) ignores uncertainty about model structure.

Possible approaches include comparing models, reporting results separately, model averaging under an explicit framework, or examining whether conclusions are robust across plausible formulations.

Scenario uncertainty

Future intervention or behavioural assumptions may not have meaningful probabilities available.

In such cases it can be clearer to present separate scenarios rather than invent a probability distribution that cannot be justified.

Uncertainty versus variability

Variability describes genuine differences among individuals, environments or stochastic outcomes. Uncertainty describes incomplete knowledge about quantities or future outcomes.

They can coexist and may require different modelling approaches.

Uncertainty reduction

Different uncertainty sources require different remedies. More informative data can reduce parameter uncertainty; improved measurement can reduce observation error; better experimental design can improve identifiability; improved model formulation can address some structural uncertainty.

Intrinsic stochastic variability cannot simply be eliminated by collecting more data.

Decision-relevant uncertainty

For decision making, the most important uncertainty may concern an outcome rather than a parameter.

Examples include the probability that hospital demand exceeds capacity, the probability that a population falls below a conservation threshold, or the probability that an intervention fails.

Model outputs should therefore be connected to the decision being considered.

Communicating uncertainty

Report what sources of uncertainty were included, how they were represented, how they were propagated and what the resulting interval or probability means.

A shaded band without a clear definition can be misleading because it might represent parameter uncertainty, process variability, observation uncertainty or some combination.

Always label uncertainty bands precisely. State whether they are confidence bands, posterior credible bands, predictive intervals, simulation quantiles or another quantity.

A practical workflow

Define the biological output of interest, identify relevant uncertainty sources, estimate or specify their distributions or ranges, preserve dependencies between inputs, propagate samples through the model, quantify Monte Carlo error, summarise the resulting output distribution and interpret it in the scientific context.

Transition to model calibration

The previous lessons have developed data preparation, estimation, uncertainty, identifiability and sensitivity separately. Model calibration brings these elements together into a practical workflow for making a mathematical model consistent with observed biological data.

Key idea. Uncertainty quantification should follow uncertainty from its source through the mathematical model to the biological quantity that matters. Parameter uncertainty, observation error, process variability, structural uncertainty and numerical error should be distinguished rather than hidden inside a single unexplained band.