← Data and Parameter Estimation

Maximum likelihood

Maximum likelihood estimates parameters by specifying a probability model for the observations and finding parameter values under which the observed dataset is most compatible with that model.

Core idea. Least squares measures discrepancy. Maximum likelihood begins one level deeper by specifying how observations are probabilistically generated from the model prediction.

Probability model for observations

Let the mechanistic model predict

\[\mu_i(\boldsymbol\theta).\]

An observation model specifies a distribution such as

\[Y_i\mid\boldsymbol\theta\sim p(y_i\mid\boldsymbol\theta).\]

The appropriate distribution depends on whether the data are continuous measurements, counts, proportions, binary outcomes or another data type.

What is a likelihood?

After observing data \(\mathbf y\), the likelihood is the same probability mass or density expression viewed as a function of the unknown parameters:

\[\boxed{L(\boldsymbol\theta;\mathbf y)=p(\mathbf y\mid\boldsymbol\theta)}.\]

The data are now fixed and \(\boldsymbol\theta\) varies.

A likelihood is not a probability distribution for the parameter. In frequentist maximum likelihood, \(L(\boldsymbol\theta;\mathbf y)\) measures relative support for parameter values from the observed data; it does not mean \(P(\boldsymbol\theta\mid\mathbf y)\).

Maximum-likelihood estimate

The maximum-likelihood estimate is

\[\boxed{\hat{\boldsymbol\theta}=\arg\max_{\boldsymbol\theta\in\Theta}L(\boldsymbol\theta;\mathbf y)}.\]

It is the parameter value at which the likelihood is largest within the allowed parameter space \(\Theta\).

Independent observations

If observations are conditionally independent given \(\boldsymbol\theta\), then

\[\boxed{L(\boldsymbol\theta)=\prod_{i=1}^{n}p(y_i\mid\boldsymbol\theta)}.\]

This product form is convenient but should not be used automatically for correlated time-series, spatial or repeated-measures data.

Log-likelihood

Because the logarithm is strictly increasing, maximising \(L\) is equivalent to maximising

\[\boxed{\ell(\boldsymbol\theta)=\log L(\boldsymbol\theta)}.\]

For independent observations, products become sums:

\[\boxed{\ell(\boldsymbol\theta)=\sum_{i=1}^{n}\log p(y_i\mid\boldsymbol\theta)}.\]

This is usually easier and numerically more stable.

Negative log-likelihood

Optimisation software often minimises rather than maximises. We can therefore define

\[\boxed{J(\boldsymbol\theta)=-\ell(\boldsymbol\theta)}\]

and minimise \(J\). This produces the same parameter estimate.

Bernoulli observations

For a binary outcome

\[Y_i\in\{0,1\}\]

with success probability \(p_i(\boldsymbol\theta)\),

\[P(Y_i=y_i)=p_i^{y_i}(1-p_i)^{1-y_i}.\]

The log-likelihood contribution is

\[y_i\log p_i+(1-y_i)\log(1-p_i).\]

This is appropriate when each observation represents one binary trial under the assumed model.

Binomial observations

If \(Y_i\) counts successes among \(n_i\) trials,

\[Y_i\sim\operatorname{Binomial}(n_i,p_i),\]

with probability mass function

\[P(Y_i=y_i)=\binom{n_i}{y_i}p_i^{y_i}(1-p_i)^{n_i-y_i}.\]

This retains information about the denominator that would be lost by treating only the observed proportion \(y_i/n_i\) as an ordinary continuous measurement.

Poisson observations

For count data, a simple model is

\[Y_i\sim\operatorname{Poisson}(\mu_i).\]

Then

\[P(Y_i=y_i)=\frac{e^{-\mu_i}\mu_i^{y_i}}{y_i!}.\]

The corresponding log-likelihood contribution is

\[y_i\log\mu_i-\mu_i-\log(y_i!).\]

Poisson mean and variance

The Poisson model implies

\[E[Y_i]=\mu_i,\qquad\operatorname{Var}(Y_i)=\mu_i.\]

This equality is a substantive assumption. Biological count data often have variance larger than their mean.

Overdispersion

If observed count variability substantially exceeds the Poisson assumption, the data are described as overdispersed relative to the Poisson model.

A negative-binomial observation model is one common alternative because it allows variance greater than the mean.

Negative-binomial observations

One useful parameterisation has mean \(\mu_i\) and dispersion parameter \(k>0\), with

\[\boxed{\operatorname{Var}(Y_i)=\mu_i+\frac{\mu_i^2}{k}}.\]

As \(k\) becomes large, this variance approaches the Poisson variance.

Different software packages use different negative-binomial parameterisations, so the definition should always be stated.

Gaussian observations

For continuous measurements with additive Gaussian error,

\[Y_i\sim N(\mu_i(\boldsymbol\theta),\sigma^2).\]

The log-likelihood is

\[\ell=-\frac n2\log(2\pi\sigma^2)-\frac{1}{2\sigma^2}\sum_i(y_i-\mu_i)^2.\]

If \(\sigma^2\) is constant, maximising this with respect to the mechanistic parameters is equivalent to ordinary least squares.

Unequal Gaussian variances

If

\[Y_i\sim N(\mu_i,\sigma_i^2),\]

the log-likelihood contains

\[-\frac12\sum_i\frac{(y_i-\mu_i)^2}{\sigma_i^2}.\]

This gives the inverse-variance weighting seen in weighted least squares.

Multiplicative error

Positive biological measurements sometimes vary approximately proportionally to their magnitude. A lognormal observation model may then be suitable:

\[\log Y_i\sim N(\log\mu_i,\sigma^2).\]

This is different from assuming additive Gaussian error on the original measurement scale.

Likelihood for a mechanistic ODE model

Suppose

\[\frac{d\mathbf x}{dt}=\mathbf f(\mathbf x,\boldsymbol\theta)\]

produces states \(\mathbf x(t_i;\boldsymbol\theta)\). An observation function gives

\[\mu_i=h(\mathbf x(t_i;\boldsymbol\theta),\boldsymbol\theta).\]

The likelihood is then built from the chosen distribution of \(Y_i\) around or conditional on \(\mu_i\).

Example: reported epidemic cases

If a deterministic epidemic model predicts expected new reported cases \(\mu_t(\boldsymbol\theta)\), one simple observation model is

\[Y_t\sim\operatorname{Poisson}(\mu_t(\boldsymbol\theta)).\]

A more flexible count model might use

\[Y_t\sim\operatorname{NegBin}(\mu_t(\boldsymbol\theta),k).\]

The mechanistic epidemic equations can remain the same while the observation model changes.

Latent state versus observation model

In a stochastic biological model, the latent state itself is random. Observations may then contain both process stochasticity and measurement noise.

The likelihood may require integrating over unobserved states:

\[\boxed{p(\mathbf y\mid\boldsymbol\theta)=\int p(\mathbf y,\mathbf x\mid\boldsymbol\theta)\,d\mathbf x}.\]

This can make likelihood evaluation substantially more difficult than for deterministic models.

Censored observations

Likelihood handles censored event-time data naturally. If an event is known only to occur after time \(t_i\), its likelihood contribution can involve the survival probability

\[S(t_i\mid\boldsymbol\theta)=P(T>t_i\mid\boldsymbol\theta).\]

This uses the information that is actually observed rather than pretending the event occurred at the censoring time.

Likelihood and constants

Terms that do not depend on \(\boldsymbol\theta\) can sometimes be omitted when the only aim is to find the maximiser.

However, constants may matter for comparing non-equivalent likelihoods or computing information criteria, so they should not be discarded blindly.

Numerical underflow

A product of many small probabilities can become too small for ordinary floating-point representation.

Using the log-likelihood replaces multiplication by addition and greatly reduces this problem.

Optimising the likelihood

For nonlinear biological models, the likelihood surface may have multiple local maxima, flat regions or strong parameter correlations.

Parameter bounds, transformations and multiple starting values can help diagnose numerical optimisation problems.

Boundary estimates

Sometimes the maximum occurs at a parameter boundary, such as

\[\hat\theta=0.\]

This may be biologically meaningful, may indicate weak information in the data, or may reflect an overly restrictive model.

Standard asymptotic uncertainty formulas can behave differently at boundaries.

Score function

The gradient of the log-likelihood is called the score:

\[\boxed{U(\boldsymbol\theta)=\nabla_{\boldsymbol\theta}\ell(\boldsymbol\theta)}.\]

At a smooth interior maximum,

\[U(\hat{\boldsymbol\theta})=0.\]

This generalises the derivative-equals-zero condition from elementary optimisation.

Curvature and information

The Hessian matrix of second derivatives describes local curvature:

\[H(\boldsymbol\theta)=\nabla^2\ell(\boldsymbol\theta).\]

Near a well-defined maximum, strong negative curvature indicates that the likelihood decreases rapidly away from the estimate, while shallow curvature indicates weaker local information.

Observed information

A common local information matrix is

\[\boxed{J(\hat{\boldsymbol\theta})=-\nabla^2\ell(\hat{\boldsymbol\theta})}.\]

When regularity conditions and sample size justify the approximation,

\[\operatorname{Cov}(\hat{\boldsymbol\theta})\approx J(\hat{\boldsymbol\theta})^{-1}.\]

This provides one route to approximate standard errors.

Likelihood profiles

For parameter \(\theta_j\), a profile likelihood fixes \(\theta_j\) at successive values and re-optimises all remaining parameters.

This reveals how much the likelihood deteriorates as that parameter moves away from its optimum and can expose asymmetry or weak identifiability.

Likelihood ratios

Two parameter values can be compared through the difference in log-likelihood. A common statistic is

\[\boxed{2\left[\ell(\hat{\boldsymbol\theta})-\ell(\boldsymbol\theta_0)\right]}.\]

Under suitable regularity conditions, likelihood-ratio theory gives useful asymptotic reference distributions.

The conditions matter, especially for small samples, boundary parameters or non-identifiable models.

Likelihood does not measure absolute model truth

A large likelihood value has meaning only relative to the specified observation model and dataset. Changing units, density conventions or data structure can change its numerical scale.

A maximum-likelihood fit can still be poor if the assumed model is biologically or statistically inadequate.

Compare predictions with data

After fitting, inspect residuals or suitable probability-based diagnostics and compare simulated or predicted observations with the real data.

Likelihood optimisation is not a substitute for model checking.

Likelihood and identifiability

A flat ridge in the likelihood surface means many parameter combinations are similarly supported by the data.

This often signals practical identifiability problems even when the numerical optimiser returns a single precise-looking value.

Maximum likelihood versus Bayesian inference

Maximum likelihood uses

\[p(\mathbf y\mid\boldsymbol\theta)\]

as a function of \(\boldsymbol\theta\). Bayesian inference additionally specifies a prior distribution \(p(\boldsymbol\theta)\) and forms

\[p(\boldsymbol\theta\mid\mathbf y)\propto p(\mathbf y\mid\boldsymbol\theta)p(\boldsymbol\theta).\]

The likelihood is therefore shared by both approaches, but their interpretation of parameter uncertainty differs.

Choosing an observation model

DataPossible observation model
binary outcomeBernoulli
successes from a known number of trialsbinomial
counts with variance near the meanPoisson
overdispersed countsnegative binomial
continuous additive measurement errorGaussian
positive multiplicative variationlognormal

These are starting points, not automatic rules. The distribution should reflect the actual observation mechanism as closely as practical.

Transition to confidence intervals

A maximum-likelihood estimate gives one fitted parameter value. The next question is how uncertain that value is. The next lesson introduces confidence intervals and explains what their frequentist interpretation does—and does not—mean.

Key idea. Maximum likelihood combines a mechanistic prediction with an explicit probability model for observations. The choice of likelihood should follow the data-generating process, and a likelihood maximum should always be interpreted together with uncertainty, identifiability and model checking.