01
Random variables and simulated epidemic samples
Moment equations describe numerical features of random variables. We therefore begin by defining exactly what an epidemic random variable, a realisation and a simulated sample mean.
Scenario: infectious population on day 20
Consider a stochastic SIS epidemic in a closed population of 100 people. Initially, 5 are infectious. Infection and recovery occur randomly in continuous time.
Before the epidemic is simulated, the infectious population on day 20 is unknown. We denote it by
\(I(20)\) is a random variable because its numerical value depends on the random epidemic events occurring before day 20.
Random variable, outcome and realisation
| Term | Meaning in this scenario |
|---|---|
| Random experiment | Simulate one SIS epidemic from day 0 to day 20. |
| Random variable | The rule that records the infectious count \(I(20)\). |
| Outcome or realisation | The value produced in one run, such as \(I(20)=17\). |
| Sample | A collection of values from independent runs. |
| Distribution | The probabilities of the possible values before observing a run. |
The random variable is not “random” after a particular run has finished. That run has produced one observed value. Uncertainty concerns which value will occur before the random experiment is performed.
State space and possible values
Because the population is 100,
The value must be an integer in the CTMC. State 0 represents extinction and is absorbing because this model has no imported infection.
Not every mathematically possible value needs to have substantial probability under the chosen parameters and initial condition.
One variable is not one trajectory
A complete trajectory contains the epidemic state through time. The random variable \(I(20)\) extracts only one feature from that trajectory: the infectious count at the specified time.
Other random variables could be defined from the same trajectory:
- final outbreak size;
- time of extinction;
- maximum infectious count; or
- whether hospital capacity is exceeded.
Every variable needs a precise definition before samples or moments are calculated.
Interactive Python laboratory
Run 2,000 independent CTMC epidemics and record only \(I(20)\) from each run. The output shows raw sample values, a frequency table and a histogram. This page deliberately does not calculate the sample mean or variance; those belong to the next lessons.
Output
Run the code to see the result.
Understand the sampling code
| Code | Meaning |
|---|---|
return I | Extracts the value of the random variable at the observation time. |
sample[run] = ... | Stores one realisation in the sample. |
dtype=int | Records the CTMC infectious count as an integer. |
.value_counts() | Counts how often each observed value occurs. |
frequency / sample_size | Converts a count into an empirical relative frequency. |
Sample versus theoretical distribution
The theoretical distribution assigns a probability to every possible value of \(I(20)\). The simulation sample provides empirical frequencies that approximate those probabilities.
A different seed produces a different sample and slightly different frequencies. Increasing the number of independent simulations generally makes the empirical distribution more stable.
Common misunderstandings
- \(I(20)\) is a random variable; 17 is one possible realisation.
- The 2,000 stored values are a sample, not 2,000 values within one epidemic.
- A histogram describes the sample distribution, not how one trajectory changes through time.
- Relative frequency estimates probability but is not automatically the exact theoretical probability.
What this lesson adds
You can now define an epidemic random variable, distinguish it from a realised value and a full trajectory, generate an independent simulated sample, and form an empirical distribution using frequencies. The next lesson uses such samples to introduce expectation and the sample mean.