Maximum likelihood
Maximum likelihood estimates parameters by specifying a probability model for the observations and finding parameter values under which the observed dataset is most compatible with that model.
Probability model for observations
Let the mechanistic model predict
\[\mu_i(\boldsymbol\theta).\]An observation model specifies a distribution such as
\[Y_i\mid\boldsymbol\theta\sim p(y_i\mid\boldsymbol\theta).\]The appropriate distribution depends on whether the data are continuous measurements, counts, proportions, binary outcomes or another data type.
What is a likelihood?
After observing data \(\mathbf y\), the likelihood is the same probability mass or density expression viewed as a function of the unknown parameters:
\[\boxed{L(\boldsymbol\theta;\mathbf y)=p(\mathbf y\mid\boldsymbol\theta)}.\]The data are now fixed and \(\boldsymbol\theta\) varies.
Maximum-likelihood estimate
The maximum-likelihood estimate is
\[\boxed{\hat{\boldsymbol\theta}=\arg\max_{\boldsymbol\theta\in\Theta}L(\boldsymbol\theta;\mathbf y)}.\]It is the parameter value at which the likelihood is largest within the allowed parameter space \(\Theta\).
Independent observations
If observations are conditionally independent given \(\boldsymbol\theta\), then
\[\boxed{L(\boldsymbol\theta)=\prod_{i=1}^{n}p(y_i\mid\boldsymbol\theta)}.\]This product form is convenient but should not be used automatically for correlated time-series, spatial or repeated-measures data.
Log-likelihood
Because the logarithm is strictly increasing, maximising \(L\) is equivalent to maximising
\[\boxed{\ell(\boldsymbol\theta)=\log L(\boldsymbol\theta)}.\]For independent observations, products become sums:
\[\boxed{\ell(\boldsymbol\theta)=\sum_{i=1}^{n}\log p(y_i\mid\boldsymbol\theta)}.\]This is usually easier and numerically more stable.
Negative log-likelihood
Optimisation software often minimises rather than maximises. We can therefore define
\[\boxed{J(\boldsymbol\theta)=-\ell(\boldsymbol\theta)}\]and minimise \(J\). This produces the same parameter estimate.
Bernoulli observations
For a binary outcome
\[Y_i\in\{0,1\}\]with success probability \(p_i(\boldsymbol\theta)\),
\[P(Y_i=y_i)=p_i^{y_i}(1-p_i)^{1-y_i}.\]The log-likelihood contribution is
\[y_i\log p_i+(1-y_i)\log(1-p_i).\]This is appropriate when each observation represents one binary trial under the assumed model.
Binomial observations
If \(Y_i\) counts successes among \(n_i\) trials,
\[Y_i\sim\operatorname{Binomial}(n_i,p_i),\]with probability mass function
\[P(Y_i=y_i)=\binom{n_i}{y_i}p_i^{y_i}(1-p_i)^{n_i-y_i}.\]This retains information about the denominator that would be lost by treating only the observed proportion \(y_i/n_i\) as an ordinary continuous measurement.
Poisson observations
For count data, a simple model is
\[Y_i\sim\operatorname{Poisson}(\mu_i).\]Then
\[P(Y_i=y_i)=\frac{e^{-\mu_i}\mu_i^{y_i}}{y_i!}.\]The corresponding log-likelihood contribution is
\[y_i\log\mu_i-\mu_i-\log(y_i!).\]Poisson mean and variance
The Poisson model implies
\[E[Y_i]=\mu_i,\qquad\operatorname{Var}(Y_i)=\mu_i.\]This equality is a substantive assumption. Biological count data often have variance larger than their mean.
Overdispersion
If observed count variability substantially exceeds the Poisson assumption, the data are described as overdispersed relative to the Poisson model.
A negative-binomial observation model is one common alternative because it allows variance greater than the mean.
Negative-binomial observations
One useful parameterisation has mean \(\mu_i\) and dispersion parameter \(k>0\), with
\[\boxed{\operatorname{Var}(Y_i)=\mu_i+\frac{\mu_i^2}{k}}.\]As \(k\) becomes large, this variance approaches the Poisson variance.
Different software packages use different negative-binomial parameterisations, so the definition should always be stated.
Gaussian observations
For continuous measurements with additive Gaussian error,
\[Y_i\sim N(\mu_i(\boldsymbol\theta),\sigma^2).\]The log-likelihood is
\[\ell=-\frac n2\log(2\pi\sigma^2)-\frac{1}{2\sigma^2}\sum_i(y_i-\mu_i)^2.\]If \(\sigma^2\) is constant, maximising this with respect to the mechanistic parameters is equivalent to ordinary least squares.
Unequal Gaussian variances
If
\[Y_i\sim N(\mu_i,\sigma_i^2),\]the log-likelihood contains
\[-\frac12\sum_i\frac{(y_i-\mu_i)^2}{\sigma_i^2}.\]This gives the inverse-variance weighting seen in weighted least squares.
Multiplicative error
Positive biological measurements sometimes vary approximately proportionally to their magnitude. A lognormal observation model may then be suitable:
\[\log Y_i\sim N(\log\mu_i,\sigma^2).\]This is different from assuming additive Gaussian error on the original measurement scale.
Likelihood for a mechanistic ODE model
Suppose
\[\frac{d\mathbf x}{dt}=\mathbf f(\mathbf x,\boldsymbol\theta)\]produces states \(\mathbf x(t_i;\boldsymbol\theta)\). An observation function gives
\[\mu_i=h(\mathbf x(t_i;\boldsymbol\theta),\boldsymbol\theta).\]The likelihood is then built from the chosen distribution of \(Y_i\) around or conditional on \(\mu_i\).
Example: reported epidemic cases
If a deterministic epidemic model predicts expected new reported cases \(\mu_t(\boldsymbol\theta)\), one simple observation model is
\[Y_t\sim\operatorname{Poisson}(\mu_t(\boldsymbol\theta)).\]A more flexible count model might use
\[Y_t\sim\operatorname{NegBin}(\mu_t(\boldsymbol\theta),k).\]The mechanistic epidemic equations can remain the same while the observation model changes.
Latent state versus observation model
In a stochastic biological model, the latent state itself is random. Observations may then contain both process stochasticity and measurement noise.
The likelihood may require integrating over unobserved states:
\[\boxed{p(\mathbf y\mid\boldsymbol\theta)=\int p(\mathbf y,\mathbf x\mid\boldsymbol\theta)\,d\mathbf x}.\]This can make likelihood evaluation substantially more difficult than for deterministic models.
Censored observations
Likelihood handles censored event-time data naturally. If an event is known only to occur after time \(t_i\), its likelihood contribution can involve the survival probability
\[S(t_i\mid\boldsymbol\theta)=P(T>t_i\mid\boldsymbol\theta).\]This uses the information that is actually observed rather than pretending the event occurred at the censoring time.
Likelihood and constants
Terms that do not depend on \(\boldsymbol\theta\) can sometimes be omitted when the only aim is to find the maximiser.
However, constants may matter for comparing non-equivalent likelihoods or computing information criteria, so they should not be discarded blindly.
Numerical underflow
A product of many small probabilities can become too small for ordinary floating-point representation.
Using the log-likelihood replaces multiplication by addition and greatly reduces this problem.
Optimising the likelihood
For nonlinear biological models, the likelihood surface may have multiple local maxima, flat regions or strong parameter correlations.
Parameter bounds, transformations and multiple starting values can help diagnose numerical optimisation problems.
Boundary estimates
Sometimes the maximum occurs at a parameter boundary, such as
\[\hat\theta=0.\]This may be biologically meaningful, may indicate weak information in the data, or may reflect an overly restrictive model.
Standard asymptotic uncertainty formulas can behave differently at boundaries.
Score function
The gradient of the log-likelihood is called the score:
\[\boxed{U(\boldsymbol\theta)=\nabla_{\boldsymbol\theta}\ell(\boldsymbol\theta)}.\]At a smooth interior maximum,
\[U(\hat{\boldsymbol\theta})=0.\]This generalises the derivative-equals-zero condition from elementary optimisation.
Curvature and information
The Hessian matrix of second derivatives describes local curvature:
\[H(\boldsymbol\theta)=\nabla^2\ell(\boldsymbol\theta).\]Near a well-defined maximum, strong negative curvature indicates that the likelihood decreases rapidly away from the estimate, while shallow curvature indicates weaker local information.
Observed information
A common local information matrix is
\[\boxed{J(\hat{\boldsymbol\theta})=-\nabla^2\ell(\hat{\boldsymbol\theta})}.\]When regularity conditions and sample size justify the approximation,
\[\operatorname{Cov}(\hat{\boldsymbol\theta})\approx J(\hat{\boldsymbol\theta})^{-1}.\]This provides one route to approximate standard errors.
Likelihood profiles
For parameter \(\theta_j\), a profile likelihood fixes \(\theta_j\) at successive values and re-optimises all remaining parameters.
This reveals how much the likelihood deteriorates as that parameter moves away from its optimum and can expose asymmetry or weak identifiability.
Likelihood ratios
Two parameter values can be compared through the difference in log-likelihood. A common statistic is
\[\boxed{2\left[\ell(\hat{\boldsymbol\theta})-\ell(\boldsymbol\theta_0)\right]}.\]Under suitable regularity conditions, likelihood-ratio theory gives useful asymptotic reference distributions.
The conditions matter, especially for small samples, boundary parameters or non-identifiable models.
Likelihood does not measure absolute model truth
A large likelihood value has meaning only relative to the specified observation model and dataset. Changing units, density conventions or data structure can change its numerical scale.
A maximum-likelihood fit can still be poor if the assumed model is biologically or statistically inadequate.
Compare predictions with data
After fitting, inspect residuals or suitable probability-based diagnostics and compare simulated or predicted observations with the real data.
Likelihood optimisation is not a substitute for model checking.
Likelihood and identifiability
A flat ridge in the likelihood surface means many parameter combinations are similarly supported by the data.
This often signals practical identifiability problems even when the numerical optimiser returns a single precise-looking value.
Maximum likelihood versus Bayesian inference
Maximum likelihood uses
\[p(\mathbf y\mid\boldsymbol\theta)\]as a function of \(\boldsymbol\theta\). Bayesian inference additionally specifies a prior distribution \(p(\boldsymbol\theta)\) and forms
\[p(\boldsymbol\theta\mid\mathbf y)\propto p(\mathbf y\mid\boldsymbol\theta)p(\boldsymbol\theta).\]The likelihood is therefore shared by both approaches, but their interpretation of parameter uncertainty differs.
Choosing an observation model
| Data | Possible observation model |
|---|---|
| binary outcome | Bernoulli |
| successes from a known number of trials | binomial |
| counts with variance near the mean | Poisson |
| overdispersed counts | negative binomial |
| continuous additive measurement error | Gaussian |
| positive multiplicative variation | lognormal |
These are starting points, not automatic rules. The distribution should reflect the actual observation mechanism as closely as practical.
Transition to confidence intervals
A maximum-likelihood estimate gives one fitted parameter value. The next question is how uncertain that value is. The next lesson introduces confidence intervals and explains what their frequentist interpretation does—and does not—mean.