Parameter estimation
A mathematical model contains parameters that control biological rates, strengths, probabilities or scales. Parameter estimation uses observations to infer values of parameters that are not known directly.
Parameters in a model
Write a parameter vector as
\[\boxed{\boldsymbol\theta=(\theta_1,\theta_2,\ldots,\theta_p)^T}.\]In an epidemic model, for example, \(\boldsymbol\theta\) might contain transmission and recovery parameters. In an ecological model it might contain growth rates and interaction coefficients.
Forward model
For a given parameter vector, a deterministic model produces a prediction. We can write
\[\boxed{\mathbf x(t;\boldsymbol\theta)}.\]The notation emphasises that changing \(\boldsymbol\theta\) changes the predicted trajectory.
From model state to predicted observation
The measured quantity need not equal a model state directly. Introduce an observation function
\[\boxed{\mu_i(\boldsymbol\theta)=h\!\left(\mathbf x(t_i;\boldsymbol\theta),\boldsymbol\theta\right)}.\]Here \(h\) converts the model state into the quantity expected to be observed.
For example, an epidemic model may describe true infections while the dataset contains reported cases.
The inverse problem
Suppose observations are
\[y_1,y_2,\ldots,y_n.\]Parameter estimation asks which values of \(\boldsymbol\theta\) make the model predictions \(\mu_i(\boldsymbol\theta)\) compatible with these observations.
This is an inverse problem because the effects are observed and the unknown model quantities are inferred.
Residuals
For continuous observations, a residual is commonly defined as
\[\boxed{r_i(\boldsymbol\theta)=y_i-\mu_i(\boldsymbol\theta)}.\]A residual measures the discrepancy between an observation and the corresponding model prediction.
The sign indicates whether the observation lies above or below the prediction.
An objective function
Many estimation methods combine discrepancies into a scalar objective function
\[J(\boldsymbol\theta).\]An estimate may then be defined by
\[\boxed{\hat{\boldsymbol\theta}=\arg\min_{\boldsymbol\theta\in\Theta}J(\boldsymbol\theta)},\]where \(\Theta\) is the allowed parameter space.
Meaning of arg min
The expression
\[\arg\min_{\boldsymbol\theta}J(\boldsymbol\theta)\]means the value of \(\boldsymbol\theta\) at which the objective function is smallest. It is the parameter value, not the minimum objective value itself.
Least squares
A common objective is the residual sum of squares
\[\boxed{J(\boldsymbol\theta)=\sum_{i=1}^{n}\left[y_i-\mu_i(\boldsymbol\theta)\right]^2}.\]The least-squares estimate minimises this quantity.
The next lesson develops the assumptions and interpretation of least squares in detail.
Weighted least squares
If observations have different uncertainty, residuals can be weighted:
\[\boxed{J(\boldsymbol\theta)=\sum_{i=1}^{n}w_i\left[y_i-\mu_i(\boldsymbol\theta)\right]^2}.\]Larger weights make particular discrepancies contribute more strongly to the objective.
Weights should have a statistical or scientific justification rather than being chosen merely to improve the appearance of the fit.
Likelihood-based estimation
Instead of defining discrepancy through squared errors, we can specify a probability model for the observations:
\[Y_i\mid\boldsymbol\theta\sim p(y_i\mid\boldsymbol\theta).\]The likelihood is
\[\boxed{L(\boldsymbol\theta)=p(\mathbf y\mid\boldsymbol\theta)}.\]Maximum likelihood chooses the parameter values that make the observed data most likely under the assumed statistical model.
Different data require different observation models
Continuous measurements, counts, binary outcomes and event times need not have the same probability distribution.
For example, a count may be modelled using a Poisson or negative-binomial distribution, while a binary observation may use a Bernoulli distribution.
The estimation criterion should match the observation process.
Parameter constraints
Biological knowledge often restricts the parameter space. A rate parameter may require
\[\theta_j>0,\]while a probability must satisfy
\[0\le\theta_j\le1.\]Including such constraints prevents mathematically possible but biologically meaningless estimates.
Parameter transformations
A positive parameter can be written
\[\theta=e^{\eta},\qquad \eta\in\mathbb R.\]Optimisation can then be performed over the unconstrained variable \(\eta\), while positivity of \(\theta\) is guaranteed.
Similar transformations can constrain probabilities to \((0,1)\).
Known and unknown parameters
Not every model parameter should necessarily be estimated from the same dataset.
Some may be fixed from reliable external evidence, while others are estimated. Trying to estimate too many weakly informed parameters can make the inverse problem unstable or non-identifiable.
Initial conditions can be parameters
The initial state may also be uncertain. For example,
\[I(0)=I_0\]may be unknown and estimated jointly with transmission parameters.
This is particularly important when observations begin after the biological process has already started.
Numerical solution inside estimation
For an ODE model, each trial parameter vector may require solving
\[\frac{d\mathbf x}{dt}=\mathbf f(\mathbf x,\boldsymbol\theta)\]before the objective function can be evaluated.
Parameter estimation can therefore involve a numerical solver inside an optimisation algorithm.
Optimisation
An optimisation algorithm searches the parameter space for values that improve the objective or likelihood.
Gradient-based methods use information about local slopes, while derivative-free methods can be useful when gradients are unavailable or unreliable.
The choice of optimiser is separate from the statistical definition of the estimator.
Local minima
For nonlinear biological models, the objective surface may contain several local minima.
An optimiser starting from one initial guess may therefore converge to a different solution from an optimiser starting elsewhere.
Using multiple starting points is one useful diagnostic.
Parameter bounds
Bounds can encode biological plausibility:
\[\theta_j^{\min}\le\theta_j\le\theta_j^{\max}.\]Bounds should be chosen from biological knowledge or clearly stated modelling assumptions. Extremely restrictive bounds can make an estimate appear precise simply because the search space was artificially narrow.
Scale of parameters
Optimisation can become difficult when one parameter is around \(10^{-6}\) and another around \(10^4\).
Rescaling or transforming parameters can improve numerical behaviour without changing the underlying biological model.
Point estimate
The fitted value
\[\hat{\boldsymbol\theta}\]is a point estimate. It provides one selected parameter vector but does not describe how uncertain that estimate is.
Good fit does not imply unique parameters
Two substantially different parameter vectors can sometimes generate almost indistinguishable model predictions:
\[\mu_i(\boldsymbol\theta^{(1)})\approx\mu_i(\boldsymbol\theta^{(2)}).\]Then the data may not distinguish the parameter combinations well.
This motivates the later study of identifiability.
Good fit does not prove the model
A model can fit observations closely for the wrong mechanistic reason. Flexible models may reproduce a dataset even when important biological assumptions are incorrect.
Parameter estimation therefore does not replace model checking and validation.
Overfitting
Adding parameters can improve fit to the data used for estimation, but excessive flexibility may reduce predictive performance on new observations.
Model comparison and validation later in this section address this trade-off more directly.
Uncertainty in parameter estimates
Observed data are finite and noisy, so fitted parameters are uncertain.
Confidence intervals, profile likelihood, bootstrap methods and other approaches can quantify different aspects of uncertainty.
A narrow-looking numerical optimum should not be interpreted as certainty without such analysis.
Structural and practical identifiability
Structural identifiability asks whether ideal, noise-free observations could uniquely determine parameters under the mathematical model.
Practical identifiability asks whether the actual finite and noisy dataset contains enough information to estimate them reliably.
These are different questions and are developed later.
Parameter correlation
Parameters may compensate for one another. Increasing one parameter while decreasing another may leave model predictions almost unchanged.
Strong parameter correlation often indicates that individual parameter values are less precisely determined than the fitted trajectory suggests.
Derived biological quantities
Sometimes the main scientific quantity is a function of several fitted parameters.
For an elementary SIR model, for example,
\[R_0=\frac{\beta}{\gamma}.\]Uncertainty in \(\beta\) and \(\gamma\) therefore propagates into uncertainty in \(R_0\).
Derived quantities should not be reported as though fitted parameters were known exactly.
A simple estimation workflow
A practical workflow is to define the model and observation process, choose unknown parameters and biologically meaningful bounds, select an objective or likelihood, solve the model for candidate parameters, optimise, inspect the fit and residuals, and then assess uncertainty and identifiability.
Each stage answers a different question and should remain explicit.
Transition to least squares
The simplest widely used estimation criterion measures the squared distance between observations and model predictions. The next lesson develops least squares carefully, including when it is statistically justified and when weighting is needed.