Probability distributions
A probability distribution describes the possible values of a random variable and how probability is assigned across those values. It provides the mathematical language for representing biological variation, sampling uncertainty and stochastic events.
Random variables
A random variable assigns a numerical value to the outcome of a random experiment.
Examples include
\[X=\text{number of infected individuals},\]or
\[T=\text{time until recovery}.\]The distribution describes uncertainty in the value before the outcome is observed.
Discrete random variables
A discrete random variable takes values in a finite or countable set, such as
\[0,1,2,\ldots.\]Examples include numbers of mutations, infections, births or deaths.
Probability mass function
For a discrete random variable \(X\), the probability mass function is
\[\boxed{p_X(x)=P(X=x)}.\]It satisfies
\[p_X(x)\ge0\]and
\[\boxed{\sum_xp_X(x)=1}.\]Continuous random variables
A continuous random variable can take values across an interval of the real line.
Examples include body mass, concentration, temperature and many measured waiting times.
Probability density function
A continuous random variable can be described by a density \(f_X(x)\) satisfying
\[f_X(x)\ge0,\qquad\boxed{\int_{-\infty}^{\infty}f_X(x)\,dx=1}.\]Probabilities are areas under the density:
\[\boxed{P(a\le X\le b)=\int_a^bf_X(x)\,dx}.\]Density is not probability at a point
For a continuous random variable,
\[P(X=x)=0\]for every individual value \(x\), even though \(f_X(x)\) may be positive.
Cumulative distribution function
For either a discrete or continuous random variable, the cumulative distribution function is
\[\boxed{F_X(x)=P(X\le x)}.\]It increases from 0 toward 1 as \(x\) moves from the lower to the upper end of the possible values.
Survival function
For waiting-time and survival variables, it is often useful to write
\[\boxed{S(t)=P(T>t)=1-F_T(t)}.\]This is the probability that the event has not occurred by time \(t\).
Expectation
The expectation is the probability-weighted average value of a random variable.
For a discrete variable,
\[\boxed{E[X]=\sum_xxP(X=x)}.\]For a continuous variable,
\[\boxed{E[X]=\int_{-\infty}^{\infty}xf_X(x)\,dx},\]when the relevant expectation exists.
Expectation is not necessarily an observable value
The expected value need not be a value the random variable can actually take.
For example, the expected number of offspring could be 2.4 even though an individual organism cannot have 2.4 offspring.
Variance
Variance measures squared spread around the mean:
\[\boxed{\operatorname{Var}(X)=E[(X-E[X])^2]}.\]An equivalent identity is
\[\boxed{\operatorname{Var}(X)=E[X^2]-E[X]^2}.\]Standard deviation
The standard deviation is
\[\boxed{\operatorname{SD}(X)=\sqrt{\operatorname{Var}(X)}}.\]Unlike variance, standard deviation has the same units as the random variable.
Parameters of a distribution
A distribution family is often controlled by parameters.
For example,
\[X\sim N(\mu,\sigma^2)\]means that \(X\) has a normal distribution with mean \(\mu\) and variance \(\sigma^2\).
The distribution parameter is not the same thing as an observed data value.
Bernoulli distribution
A Bernoulli random variable represents one binary trial:
\[X\in\{0,1\}.\]If
\[P(X=1)=p,\]then
\[P(X=0)=1-p.\]We write
\[X\sim\operatorname{Bernoulli}(p).\]Its mean and variance are
\[\boxed{E[X]=p},\qquad\boxed{\operatorname{Var}(X)=p(1-p)}.\]Biological Bernoulli examples
A Bernoulli variable can represent whether one individual is infected, whether treatment succeeds, whether a seed germinates or whether an organism survives to a specified time.
Binomial distribution
Suppose there are \(n\) independent Bernoulli trials, each with the same success probability \(p\). If \(X\) counts the successes, then
\[\boxed{X\sim\operatorname{Binomial}(n,p)}\]with
\[\boxed{P(X=k)=\binom nkp^k(1-p)^{n-k}},\qquad k=0,\ldots,n.\]Binomial mean and variance
For
\[X\sim\operatorname{Binomial}(n,p),\]we have
\[\boxed{E[X]=np},\qquad\boxed{\operatorname{Var}(X)=np(1-p)}.\]When the binomial model is appropriate
The usual binomial model assumes a fixed number of trials, binary outcomes, a common success probability and independence between trials.
If success probabilities differ substantially or outcomes are dependent, the ordinary binomial model may not be appropriate.
Poisson distribution
A Poisson variable is commonly used for counts of events occurring over a specified exposure interval when a Poisson-process approximation is reasonable.
If
\[X\sim\operatorname{Poisson}(\lambda),\]then
\[\boxed{P(X=k)=e^{-\lambda}\frac{\lambda^k}{k!}},\qquad k=0,1,2,\ldots.\]Poisson mean and variance
The Poisson distribution has
\[\boxed{E[X]=\lambda},\qquad\boxed{\operatorname{Var}(X)=\lambda}.\]The equality of mean and variance is an important modelling assumption, not a universal property of biological count data.
Rates and exposure
If events occur at constant rate \(r\) per unit exposure and the exposure is \(t\), the expected count is
\[\lambda=rt.\]Thus the Poisson parameter is an expected number of events over the specified exposure, rather than simply a rate without units.
Overdispersion
Biological counts often have
\[\operatorname{Var}(X)>E[X].\]This is overdispersion relative to a Poisson model and can arise from heterogeneity, clustering, dependence or other mechanisms.
Negative-binomial distribution
The negative-binomial family is commonly used for overdispersed counts.
One mean-dispersion parameterisation satisfies
\[\boxed{E[X]=\mu,\qquad\operatorname{Var}(X)=\mu+\frac{\mu^2}{k}},\]where \(k>0\) controls dispersion.
Different texts and software use different parameterisations, so the exact definition should always be stated.
Normal distribution
A normal random variable is written
\[\boxed{X\sim N(\mu,\sigma^2)}.\]Its density is
\[\boxed{f(x)=\frac{1}{\sigma\sqrt{2\pi}}\exp\!\left[-\frac{(x-\mu)^2}{2\sigma^2}\right]}.\]The distribution is symmetric around \(\mu\).
Normal mean and variance
For a normal variable,
\[E[X]=\mu,\qquad\operatorname{Var}(X)=\sigma^2.\]The parameter \(\sigma\) is the standard deviation, while \(\sigma^2\) is the variance.
Standard normal distribution
If
\[Z\sim N(0,1),\]then \(Z\) has the standard normal distribution.
A normal variable can be standardised using
\[\boxed{Z=\frac{X-\mu}{\sigma}}.\]This expresses a value in units of standard deviations from the mean.
The 68–95 rule
For a normal distribution, approximately 68% of probability lies within one standard deviation of the mean and approximately 95% lies within about two standard deviations.
These percentages are properties of the normal distribution and should not be applied automatically to arbitrary distributions.
Normal distributions in biology
The normal distribution is often useful for approximately symmetric continuous measurements or for measurement errors produced by many small effects.
It may be unsuitable for strongly skewed data, bounded proportions or counts with many small values.
Exponential distribution
The exponential distribution models positive waiting times under a constant hazard rate.
If
\[T\sim\operatorname{Exp}(\lambda),\]then
\[\boxed{f_T(t)=\lambda e^{-\lambda t}},\qquad t\ge0.\]Exponential survival function
Its survival function is
\[\boxed{S(t)=P(T>t)=e^{-\lambda t}}.\]The mean and variance are
\[\boxed{E[T]=\frac1\lambda},\qquad\boxed{\operatorname{Var}(T)=\frac1{\lambda^2}}.\]Memorylessness
The exponential distribution satisfies
\[\boxed{P(T>s+t\mid T>s)=P(T>t)}.\]The remaining waiting-time distribution does not depend on how long has already elapsed.
This property is central to many continuous-time Markov models.
Poisson counts and exponential waiting times
In a homogeneous Poisson process with rate \(\lambda\), event counts over an interval of length \(t\) are Poisson with mean \(\lambda t\), while waiting times between events are exponential with rate \(\lambda\).
These are two related descriptions of the same idealised event process.
Uniform distribution
If
\[X\sim U(a,b),\]then the density is constant on \([a,b]\):
\[f(x)=\frac1{b-a},\qquad a\le x\le b.\]Its mean is
\[E[X]=\frac{a+b}{2}.\]Uniform random variables are also important computationally because many simulation algorithms begin with draws from \(U(0,1)\).
Distribution of a transformed variable
If
\[Y=g(X),\]then the distribution of \(Y\) is induced by the distribution of \(X\) and the transformation \(g\).
Means and variances generally cannot be obtained simply by applying \(g\) to the mean and variance.
Joint distributions
When several random variables are considered together, a joint distribution describes their combined uncertainty:
\[p(x,y)=P(X=x,Y=y)\]for a discrete example, or a joint density \(f(x,y)\) for continuous variables.
Marginal distributions
A marginal distribution describes one variable without conditioning on the value of another.
For a discrete joint distribution,
\[\boxed{P(X=x)=\sum_yP(X=x,Y=y)}.\]Conditional distributions
A conditional distribution describes uncertainty in one variable after information about another is known:
\[\boxed{P(X=x\mid Y=y)=\frac{P(X=x,Y=y)}{P(Y=y)}}\]when \(P(Y=y)>0\).
Independence
Random variables \(X\) and \(Y\) are independent if their joint distribution factorises:
\[\boxed{P(X=x,Y=y)=P(X=x)P(Y=y)}\]in the discrete case, with the analogous density factorisation for continuous variables.
Independence is a modelling assumption and should not be inferred merely because two measurements look different.
Covariance
Covariance measures linear co-variation:
\[\boxed{\operatorname{Cov}(X,Y)=E[(X-E[X])(Y-E[Y])]}.\]Positive covariance means larger values tend to occur together; negative covariance means larger values of one tend to accompany smaller values of the other.
Correlation
Correlation standardises covariance:
\[\boxed{\rho_{XY}=\frac{\operatorname{Cov}(X,Y)}{\sigma_X\sigma_Y}}.\]It lies between \(-1\) and \(1\) when both standard deviations are positive.
Zero correlation does not generally imply independence.
Mixture distributions
Biological populations can contain heterogeneous subgroups. If an observation comes from subgroup \(j\) with probability \(w_j\), a mixture distribution has the form
\[\boxed{f(x)=\sum_jw_jf_j(x)},\qquad\sum_jw_j=1.\]A mixture can be skewed or multimodal even when each component distribution is simple.
Choosing a distribution
| Biological quantity | Common starting model |
|---|---|
| single yes/no outcome | Bernoulli |
| successes among fixed trials | binomial |
| event count | Poisson |
| overdispersed event count | negative binomial |
| approximately symmetric continuous measurement | normal |
| constant-rate waiting time | exponential |
These are starting points rather than automatic choices. The assumptions should be checked against how the biological data are generated.
Distribution versus observed histogram
A theoretical probability distribution is a mathematical model for a random variable. A histogram is a summary of a finite observed sample.
The histogram may resemble the underlying distribution, but sampling variability means they are not the same object.
Distribution versus sampling distribution
The distribution of individual observations is different from the sampling distribution of a statistic such as a sample mean or estimator.
This distinction becomes central when constructing standard errors, confidence intervals and hypothesis tests.
Transition to sampling
Probability distributions describe the random mechanism generating observations. The next lesson studies what happens when we observe only a finite sample from a population and use that sample to learn about the underlying biological system.