← Stochastic Processes for Biology

Random variables

In the previous section, we saw that a biological process can have several possible outcomes. To analyse those outcomes mathematically, we need a way to attach a number to each one. That is the purpose of a random variable.

Core idea. A random variable is a numerical quantity whose value depends on the outcome of a random process. Before the random event occurs, we do not know which value it will take.

Start with a simple epidemic example

Suppose one infectious person is observed during one day. Let

\[X=\text{number of new people infected by that person during the day}.\]

At the beginning of the day, we do not know the value of \(X\). Possible values might be

\[X=0,1,2,3,\ldots\]

After the day has passed, perhaps the person infected two people. For that particular outcome,

\[X=2.\]
The important point. \(X\) is not “random” because the symbol itself changes unpredictably. It is random because its value is determined by an uncertain outcome.

Outcome and random variable are not the same thing

Suppose we observe three susceptible contacts. The biological outcome might record exactly which people became infected:

\[\text{infected, not infected, infected}.\]

The random variable \(X\) could simply count the infections. It converts that outcome into the number

\[X=2.\]
random biological outcome→apply a numerical rule→value of random variable

Different random variables can be defined from the same experiment. We might instead define \(Y\) as the time until the first infection, or \(Z\) as whether at least one infection occurs.

Capital and lowercase notation

We usually use a capital letter such as \(X\) for the random variable and a lowercase letter such as \(x\) for one possible value.

\[X=\text{random variable},\qquad x=\text{a possible value of }X.\]

Therefore

\[P(X=x)\]

means “the probability that the random variable \(X\) takes the value \(x\).”

Example. If \(X\) is the number of new infections, then \(P(X=2)=0.30\) means there is a 30% probability that exactly two new infections occur.

A random variable has a probability distribution

Knowing the possible values of \(X\) is not enough. We also need to know how likely those values are. The probabilities associated with all possible values form the probability distribution of \(X\).

Suppose

Number of new infections \(x\)\(P(X=x)\)
00.20
10.40
20.30
30.10

The probabilities must add to 1:

\[0.20+0.40+0.30+0.10=1.\]

This means one of the listed possibilities must occur under this simplified model.

number of new infections, xprobability P(X = x)01230.200.400.300.10

The graph shows the whole distribution. It tells us much more than reporting only one observed value of \(X\).

Discrete random variables

A discrete random variable takes separate, countable values.

Examples in biology include:

Random variablePossible values
number of infectious individuals\(0,1,2,3,\ldots\)
number of births in one day\(0,1,2,3,\ldots\)
number of mutations in a DNA region\(0,1,2,3,\ldots\)
whether infection occurs\(0\) or \(1\)

There cannot be 2.4 infected individuals in an individual-count model. The values occur as separate counts.

Continuous random variables

A continuous random variable can take any value within an interval.

For example, let

\[T=\text{time until the next infection event}.\]

Possible values might include

\[T=0.4,\quad 1.27,\quad 3.816\text{ days},\]

and all other positive real values allowed by the model.

DiscreteContinuous
number of infection eventswaiting time until an infection
number of cellscell mass
number of deathstime until death

Why continuous probabilities are different

For a continuous random variable, probability is assigned to intervals rather than individual exact points.

If \(T\) is a continuous waiting time, then

\[P(1\leq T\leq2)\]

is the probability that the event occurs between 1 and 2 days.

A probability density function \(f(t)\) describes how probability is distributed across possible values:

\[\boxed{P(a\leq T\leq b)=\int_a^b f(t)\,dt}.\]
possible value of Tdensity f(t)abarea = P(a ≤ T ≤ b)
Important. The height \(f(t)\) is a probability density, not the probability that \(T=t\). For a continuous random variable, probability comes from the area under the density curve over an interval.

Expectation: the long-run average

The expected value of a random variable describes its probability-weighted average.

For a discrete random variable,

\[\boxed{E[X]=\sum_x xP(X=x)}.\]

Using our infection example,

\[E[X]=0(0.20)+1(0.40)+2(0.30)+3(0.10)=1.3.\]

So the expected number of new infections is

\[E[X]=1.3.\]
Interpretation. This does not mean that one person produces exactly 1.3 infections. That is impossible in this count model. It means that if the same probabilistic situation were repeated many times, the average number of infections would tend towards 1.3.

Why expectation alone is not enough

Two random variables can have the same expected value but very different uncertainty.

For example, one epidemic process might usually produce values close to its mean, while another might frequently produce either extinction or a very large outbreak. Their averages could be similar even though their biological risks are very different.

We therefore also need a measure of how widely the values are spread.

Variance: how much outcomes vary

The variance measures the expected squared distance from the mean:

\[\boxed{\operatorname{Var}(X)=E[(X-E[X])^2]}.\]

An equivalent and often convenient formula is

\[\boxed{\operatorname{Var}(X)=E[X^2]-E[X]^2}.\]

For our example,

\[E[X^2]=0^2(0.20)+1^2(0.40)+2^2(0.30)+3^2(0.10)=2.5,\]

so

\[\operatorname{Var}(X)=2.5-(1.3)^2=0.81.\]

The standard deviation is

\[\sqrt{0.81}=0.9.\]
Interpretation. Expectation tells us the centre of the distribution. Variance tells us how much the possible outcomes are dispersed around that centre.

Probability, expectation and variance answer different questions

QuantityQuestion answered
\(P(X=x)\)How likely is this particular discrete outcome?
\(P(X\in A)\)How likely is a specified range or set of outcomes?
\(E[X]\)What is the probability-weighted average?
\(\operatorname{Var}(X)\)How much do possible outcomes vary around the mean?

Indicator random variables

A particularly useful random variable records whether an event occurs:

\[Y=\begin{cases}1,&\text{if infection occurs},\\0,&\text{if infection does not occur}.\end{cases}\]

If the probability of infection is \(p\), then

\[P(Y=1)=p,\qquad P(Y=0)=1-p.\]

Its expected value is

\[E[Y]=p.\]

This simple idea is useful because counts can often be constructed by adding indicators for individual events.

Random variables in epidemic modelling

We can define many useful random variables from the same epidemic:

SymbolPossible meaning
\(I(10)\)number infectious on day 10
\(T\)time until the next infection
\(Z\)total number infected before the epidemic ends
\(M\)maximum number infectious at any time
\(Y\)indicator that a large outbreak occurs

Each random variable answers a different biological question.

From a random variable to a stochastic process

A random variable such as \(I(10)\) describes uncertainty at one particular time. But an epidemic changes through time. We may want

\[I(0),\ I(1),\ I(2),\ldots\]

or, in continuous time,

\[I(t),\qquad t\geq0.\]

Now we have a collection of random variables indexed by time. This is the idea of a stochastic process, which is the subject of the next section.

random outcome→random variable→probability distribution→expectation and variance→random variables through time
Key idea. A random variable turns an uncertain biological outcome into a number. Its distribution describes the possible values and their probabilities; its expectation describes the average; and its variance describes the spread. When random variables are followed through time, they lead naturally to stochastic processes.