Random variables
In the previous section, we saw that a biological process can have several possible outcomes. To analyse those outcomes mathematically, we need a way to attach a number to each one. That is the purpose of a random variable.
Start with a simple epidemic example
Suppose one infectious person is observed during one day. Let
\[X=\text{number of new people infected by that person during the day}.\]At the beginning of the day, we do not know the value of \(X\). Possible values might be
\[X=0,1,2,3,\ldots\]After the day has passed, perhaps the person infected two people. For that particular outcome,
\[X=2.\]Outcome and random variable are not the same thing
Suppose we observe three susceptible contacts. The biological outcome might record exactly which people became infected:
\[\text{infected, not infected, infected}.\]The random variable \(X\) could simply count the infections. It converts that outcome into the number
\[X=2.\]Different random variables can be defined from the same experiment. We might instead define \(Y\) as the time until the first infection, or \(Z\) as whether at least one infection occurs.
Capital and lowercase notation
We usually use a capital letter such as \(X\) for the random variable and a lowercase letter such as \(x\) for one possible value.
\[X=\text{random variable},\qquad x=\text{a possible value of }X.\]Therefore
\[P(X=x)\]means “the probability that the random variable \(X\) takes the value \(x\).”
A random variable has a probability distribution
Knowing the possible values of \(X\) is not enough. We also need to know how likely those values are. The probabilities associated with all possible values form the probability distribution of \(X\).
Suppose
| Number of new infections \(x\) | \(P(X=x)\) |
|---|---|
| 0 | 0.20 |
| 1 | 0.40 |
| 2 | 0.30 |
| 3 | 0.10 |
The probabilities must add to 1:
\[0.20+0.40+0.30+0.10=1.\]This means one of the listed possibilities must occur under this simplified model.
The graph shows the whole distribution. It tells us much more than reporting only one observed value of \(X\).
Discrete random variables
A discrete random variable takes separate, countable values.
Examples in biology include:
| Random variable | Possible values |
|---|---|
| number of infectious individuals | \(0,1,2,3,\ldots\) |
| number of births in one day | \(0,1,2,3,\ldots\) |
| number of mutations in a DNA region | \(0,1,2,3,\ldots\) |
| whether infection occurs | \(0\) or \(1\) |
There cannot be 2.4 infected individuals in an individual-count model. The values occur as separate counts.
Continuous random variables
A continuous random variable can take any value within an interval.
For example, let
\[T=\text{time until the next infection event}.\]Possible values might include
\[T=0.4,\quad 1.27,\quad 3.816\text{ days},\]and all other positive real values allowed by the model.
| Discrete | Continuous |
|---|---|
| number of infection events | waiting time until an infection |
| number of cells | cell mass |
| number of deaths | time until death |
Why continuous probabilities are different
For a continuous random variable, probability is assigned to intervals rather than individual exact points.
If \(T\) is a continuous waiting time, then
\[P(1\leq T\leq2)\]is the probability that the event occurs between 1 and 2 days.
A probability density function \(f(t)\) describes how probability is distributed across possible values:
\[\boxed{P(a\leq T\leq b)=\int_a^b f(t)\,dt}.\]Expectation: the long-run average
The expected value of a random variable describes its probability-weighted average.
For a discrete random variable,
\[\boxed{E[X]=\sum_x xP(X=x)}.\]Using our infection example,
\[E[X]=0(0.20)+1(0.40)+2(0.30)+3(0.10)=1.3.\]So the expected number of new infections is
\[E[X]=1.3.\]Why expectation alone is not enough
Two random variables can have the same expected value but very different uncertainty.
For example, one epidemic process might usually produce values close to its mean, while another might frequently produce either extinction or a very large outbreak. Their averages could be similar even though their biological risks are very different.
We therefore also need a measure of how widely the values are spread.
Variance: how much outcomes vary
The variance measures the expected squared distance from the mean:
\[\boxed{\operatorname{Var}(X)=E[(X-E[X])^2]}.\]An equivalent and often convenient formula is
\[\boxed{\operatorname{Var}(X)=E[X^2]-E[X]^2}.\]For our example,
\[E[X^2]=0^2(0.20)+1^2(0.40)+2^2(0.30)+3^2(0.10)=2.5,\]so
\[\operatorname{Var}(X)=2.5-(1.3)^2=0.81.\]The standard deviation is
\[\sqrt{0.81}=0.9.\]Probability, expectation and variance answer different questions
| Quantity | Question answered |
|---|---|
| \(P(X=x)\) | How likely is this particular discrete outcome? |
| \(P(X\in A)\) | How likely is a specified range or set of outcomes? |
| \(E[X]\) | What is the probability-weighted average? |
| \(\operatorname{Var}(X)\) | How much do possible outcomes vary around the mean? |
Indicator random variables
A particularly useful random variable records whether an event occurs:
\[Y=\begin{cases}1,&\text{if infection occurs},\\0,&\text{if infection does not occur}.\end{cases}\]If the probability of infection is \(p\), then
\[P(Y=1)=p,\qquad P(Y=0)=1-p.\]Its expected value is
\[E[Y]=p.\]This simple idea is useful because counts can often be constructed by adding indicators for individual events.
Random variables in epidemic modelling
We can define many useful random variables from the same epidemic:
| Symbol | Possible meaning |
|---|---|
| \(I(10)\) | number infectious on day 10 |
| \(T\) | time until the next infection |
| \(Z\) | total number infected before the epidemic ends |
| \(M\) | maximum number infectious at any time |
| \(Y\) | indicator that a large outbreak occurs |
Each random variable answers a different biological question.
From a random variable to a stochastic process
A random variable such as \(I(10)\) describes uncertainty at one particular time. But an epidemic changes through time. We may want
\[I(0),\ I(1),\ I(2),\ldots\]or, in continuous time,
\[I(t),\qquad t\geq0.\]Now we have a collection of random variables indexed by time. This is the idea of a stochastic process, which is the subject of the next section.