← Statistics for Mathematical Biology

Probability distributions

A probability distribution describes the possible values of a random variable and how probability is assigned across those values. It provides the mathematical language for representing biological variation, sampling uncertainty and stochastic events.

Core idea. Before choosing a statistical method, identify what kind of quantity is random: a binary outcome, a count, a waiting time, a proportion or a continuous measurement. Different data-generating mechanisms lead naturally to different probability distributions.

Random variables

A random variable assigns a numerical value to the outcome of a random experiment.

Examples include

\[X=\text{number of infected individuals},\]

or

\[T=\text{time until recovery}.\]

The distribution describes uncertainty in the value before the outcome is observed.

Discrete random variables

A discrete random variable takes values in a finite or countable set, such as

\[0,1,2,\ldots.\]

Examples include numbers of mutations, infections, births or deaths.

Probability mass function

For a discrete random variable \(X\), the probability mass function is

\[\boxed{p_X(x)=P(X=x)}.\]

It satisfies

\[p_X(x)\ge0\]

and

\[\boxed{\sum_xp_X(x)=1}.\]

Continuous random variables

A continuous random variable can take values across an interval of the real line.

Examples include body mass, concentration, temperature and many measured waiting times.

Probability density function

A continuous random variable can be described by a density \(f_X(x)\) satisfying

\[f_X(x)\ge0,\qquad\boxed{\int_{-\infty}^{\infty}f_X(x)\,dx=1}.\]

Probabilities are areas under the density:

\[\boxed{P(a\le X\le b)=\int_a^bf_X(x)\,dx}.\]

Density is not probability at a point

For a continuous random variable,

\[P(X=x)=0\]

for every individual value \(x\), even though \(f_X(x)\) may be positive.

Important distinction. A probability density can be greater than 1. It is the area under the density over an interval that must be a probability between 0 and 1.

Cumulative distribution function

For either a discrete or continuous random variable, the cumulative distribution function is

\[\boxed{F_X(x)=P(X\le x)}.\]

It increases from 0 toward 1 as \(x\) moves from the lower to the upper end of the possible values.

Survival function

For waiting-time and survival variables, it is often useful to write

\[\boxed{S(t)=P(T>t)=1-F_T(t)}.\]

This is the probability that the event has not occurred by time \(t\).

Expectation

The expectation is the probability-weighted average value of a random variable.

For a discrete variable,

\[\boxed{E[X]=\sum_xxP(X=x)}.\]

For a continuous variable,

\[\boxed{E[X]=\int_{-\infty}^{\infty}xf_X(x)\,dx},\]

when the relevant expectation exists.

Expectation is not necessarily an observable value

The expected value need not be a value the random variable can actually take.

For example, the expected number of offspring could be 2.4 even though an individual organism cannot have 2.4 offspring.

Variance

Variance measures squared spread around the mean:

\[\boxed{\operatorname{Var}(X)=E[(X-E[X])^2]}.\]

An equivalent identity is

\[\boxed{\operatorname{Var}(X)=E[X^2]-E[X]^2}.\]

Standard deviation

The standard deviation is

\[\boxed{\operatorname{SD}(X)=\sqrt{\operatorname{Var}(X)}}.\]

Unlike variance, standard deviation has the same units as the random variable.

Parameters of a distribution

A distribution family is often controlled by parameters.

For example,

\[X\sim N(\mu,\sigma^2)\]

means that \(X\) has a normal distribution with mean \(\mu\) and variance \(\sigma^2\).

The distribution parameter is not the same thing as an observed data value.

Bernoulli distribution

A Bernoulli random variable represents one binary trial:

\[X\in\{0,1\}.\]

If

\[P(X=1)=p,\]

then

\[P(X=0)=1-p.\]

We write

\[X\sim\operatorname{Bernoulli}(p).\]

Its mean and variance are

\[\boxed{E[X]=p},\qquad\boxed{\operatorname{Var}(X)=p(1-p)}.\]

Biological Bernoulli examples

A Bernoulli variable can represent whether one individual is infected, whether treatment succeeds, whether a seed germinates or whether an organism survives to a specified time.

Binomial distribution

Suppose there are \(n\) independent Bernoulli trials, each with the same success probability \(p\). If \(X\) counts the successes, then

\[\boxed{X\sim\operatorname{Binomial}(n,p)}\]

with

\[\boxed{P(X=k)=\binom nkp^k(1-p)^{n-k}},\qquad k=0,\ldots,n.\]

Binomial mean and variance

For

\[X\sim\operatorname{Binomial}(n,p),\]

we have

\[\boxed{E[X]=np},\qquad\boxed{\operatorname{Var}(X)=np(1-p)}.\]

When the binomial model is appropriate

The usual binomial model assumes a fixed number of trials, binary outcomes, a common success probability and independence between trials.

If success probabilities differ substantially or outcomes are dependent, the ordinary binomial model may not be appropriate.

Poisson distribution

A Poisson variable is commonly used for counts of events occurring over a specified exposure interval when a Poisson-process approximation is reasonable.

If

\[X\sim\operatorname{Poisson}(\lambda),\]

then

\[\boxed{P(X=k)=e^{-\lambda}\frac{\lambda^k}{k!}},\qquad k=0,1,2,\ldots.\]

Poisson mean and variance

The Poisson distribution has

\[\boxed{E[X]=\lambda},\qquad\boxed{\operatorname{Var}(X)=\lambda}.\]

The equality of mean and variance is an important modelling assumption, not a universal property of biological count data.

Rates and exposure

If events occur at constant rate \(r\) per unit exposure and the exposure is \(t\), the expected count is

\[\lambda=rt.\]

Thus the Poisson parameter is an expected number of events over the specified exposure, rather than simply a rate without units.

Overdispersion

Biological counts often have

\[\operatorname{Var}(X)>E[X].\]

This is overdispersion relative to a Poisson model and can arise from heterogeneity, clustering, dependence or other mechanisms.

Negative-binomial distribution

The negative-binomial family is commonly used for overdispersed counts.

One mean-dispersion parameterisation satisfies

\[\boxed{E[X]=\mu,\qquad\operatorname{Var}(X)=\mu+\frac{\mu^2}{k}},\]

where \(k>0\) controls dispersion.

Different texts and software use different parameterisations, so the exact definition should always be stated.

Normal distribution

A normal random variable is written

\[\boxed{X\sim N(\mu,\sigma^2)}.\]

Its density is

\[\boxed{f(x)=\frac{1}{\sigma\sqrt{2\pi}}\exp\!\left[-\frac{(x-\mu)^2}{2\sigma^2}\right]}.\]

The distribution is symmetric around \(\mu\).

Normal mean and variance

For a normal variable,

\[E[X]=\mu,\qquad\operatorname{Var}(X)=\sigma^2.\]

The parameter \(\sigma\) is the standard deviation, while \(\sigma^2\) is the variance.

Standard normal distribution

If

\[Z\sim N(0,1),\]

then \(Z\) has the standard normal distribution.

A normal variable can be standardised using

\[\boxed{Z=\frac{X-\mu}{\sigma}}.\]

This expresses a value in units of standard deviations from the mean.

The 68–95 rule

For a normal distribution, approximately 68% of probability lies within one standard deviation of the mean and approximately 95% lies within about two standard deviations.

These percentages are properties of the normal distribution and should not be applied automatically to arbitrary distributions.

Normal distributions in biology

The normal distribution is often useful for approximately symmetric continuous measurements or for measurement errors produced by many small effects.

It may be unsuitable for strongly skewed data, bounded proportions or counts with many small values.

Exponential distribution

The exponential distribution models positive waiting times under a constant hazard rate.

If

\[T\sim\operatorname{Exp}(\lambda),\]

then

\[\boxed{f_T(t)=\lambda e^{-\lambda t}},\qquad t\ge0.\]

Exponential survival function

Its survival function is

\[\boxed{S(t)=P(T>t)=e^{-\lambda t}}.\]

The mean and variance are

\[\boxed{E[T]=\frac1\lambda},\qquad\boxed{\operatorname{Var}(T)=\frac1{\lambda^2}}.\]

Memorylessness

The exponential distribution satisfies

\[\boxed{P(T>s+t\mid T>s)=P(T>t)}.\]

The remaining waiting-time distribution does not depend on how long has already elapsed.

This property is central to many continuous-time Markov models.

Poisson counts and exponential waiting times

In a homogeneous Poisson process with rate \(\lambda\), event counts over an interval of length \(t\) are Poisson with mean \(\lambda t\), while waiting times between events are exponential with rate \(\lambda\).

These are two related descriptions of the same idealised event process.

Uniform distribution

If

\[X\sim U(a,b),\]

then the density is constant on \([a,b]\):

\[f(x)=\frac1{b-a},\qquad a\le x\le b.\]

Its mean is

\[E[X]=\frac{a+b}{2}.\]

Uniform random variables are also important computationally because many simulation algorithms begin with draws from \(U(0,1)\).

Distribution of a transformed variable

If

\[Y=g(X),\]

then the distribution of \(Y\) is induced by the distribution of \(X\) and the transformation \(g\).

Means and variances generally cannot be obtained simply by applying \(g\) to the mean and variance.

Joint distributions

When several random variables are considered together, a joint distribution describes their combined uncertainty:

\[p(x,y)=P(X=x,Y=y)\]

for a discrete example, or a joint density \(f(x,y)\) for continuous variables.

Marginal distributions

A marginal distribution describes one variable without conditioning on the value of another.

For a discrete joint distribution,

\[\boxed{P(X=x)=\sum_yP(X=x,Y=y)}.\]

Conditional distributions

A conditional distribution describes uncertainty in one variable after information about another is known:

\[\boxed{P(X=x\mid Y=y)=\frac{P(X=x,Y=y)}{P(Y=y)}}\]

when \(P(Y=y)>0\).

Independence

Random variables \(X\) and \(Y\) are independent if their joint distribution factorises:

\[\boxed{P(X=x,Y=y)=P(X=x)P(Y=y)}\]

in the discrete case, with the analogous density factorisation for continuous variables.

Independence is a modelling assumption and should not be inferred merely because two measurements look different.

Covariance

Covariance measures linear co-variation:

\[\boxed{\operatorname{Cov}(X,Y)=E[(X-E[X])(Y-E[Y])]}.\]

Positive covariance means larger values tend to occur together; negative covariance means larger values of one tend to accompany smaller values of the other.

Correlation

Correlation standardises covariance:

\[\boxed{\rho_{XY}=\frac{\operatorname{Cov}(X,Y)}{\sigma_X\sigma_Y}}.\]

It lies between \(-1\) and \(1\) when both standard deviations are positive.

Zero correlation does not generally imply independence.

Mixture distributions

Biological populations can contain heterogeneous subgroups. If an observation comes from subgroup \(j\) with probability \(w_j\), a mixture distribution has the form

\[\boxed{f(x)=\sum_jw_jf_j(x)},\qquad\sum_jw_j=1.\]

A mixture can be skewed or multimodal even when each component distribution is simple.

Choosing a distribution

Biological quantityCommon starting model
single yes/no outcomeBernoulli
successes among fixed trialsbinomial
event countPoisson
overdispersed event countnegative binomial
approximately symmetric continuous measurementnormal
constant-rate waiting timeexponential

These are starting points rather than automatic choices. The assumptions should be checked against how the biological data are generated.

Distribution versus observed histogram

A theoretical probability distribution is a mathematical model for a random variable. A histogram is a summary of a finite observed sample.

The histogram may resemble the underlying distribution, but sampling variability means they are not the same object.

Distribution versus sampling distribution

The distribution of individual observations is different from the sampling distribution of a statistic such as a sample mean or estimator.

This distinction becomes central when constructing standard errors, confidence intervals and hypothesis tests.

Transition to sampling

Probability distributions describe the random mechanism generating observations. The next lesson studies what happens when we observe only a finite sample from a population and use that sample to learn about the underlying biological system.

Key idea. A probability distribution connects biological assumptions to mathematical uncertainty. Its support, parameters, mean, variance and dependence structure should reflect the quantity being modelled rather than being chosen only because a distribution is familiar.