14
Histograms, sample paths, probability plots and tables
The same stochastic results can be communicated well or badly. The correct display depends on the question: individual histories, outcome variability, event probability through time or precise numerical estimates.
The purpose and boundary of this lesson
This lesson introduces no new epidemic transitions, probability formulas or simulation algorithms. It reuses the DTMC and Monte Carlo methods from Lessons 1–13 and focuses on selecting, labelling and interpreting outputs.
A display should answer one clearly stated question. Trying to make one graph communicate trajectories, distributions, thresholds, estimates and uncertainty simultaneously usually reduces clarity.
Start with the question, not the graph type
What can individual epidemics look like?
Show several sample paths against time.
Use: step-path plotHow variable is one outcome?
Collect one value per trajectory, such as final size.
Use: histogram or empirical distributionHow does an event probability change with time?
Estimate the proportion extinct by each common time.
Use: probability curveWhat exact values should be reported?
Present definitions, simulation count, estimates and uncertainty.
Use: numerical table1. Sample-path plots
A sample-path plot displays a small selection of realised integer-valued trajectories:
Use a step plot because the DTMC state changes at fixed time points. Display only enough paths to reveal variation without creating an unreadable mass of lines.
- Label the time unit and compartment count.
- State how many paths are shown and how many were simulated.
- Use transparency so overlapping paths remain visible.
- Do not describe one selected path as the predicted epidemic.
Example caption: “Fifteen of 600 simulated SIR DTMC trajectories. Each grey step line is one realisation from the same initial condition and parameter set; variation arises from random infection and recovery events.”
2. Histograms of trajectory outcomes
A histogram should use one comparable outcome per independent trajectory, such as:
- completed final epidemic size;
- peak infectious count;
- peak time; or
- extinction time.
The horizontal axis shows the outcome value or interval. The vertical axis must say whether bars represent counts, relative frequencies or probability density.
Histogram appearance depends on bin width. Very wide bins hide structure; very narrow bins create a noisy display. Check that the biological conclusion is not an artefact of one arbitrary bin choice.
If a major-outbreak threshold is relevant, mark it with a vertical line and identify it in the legend or caption.
3. Probability plots
Here “probability plot” means a probability shown against time or a threshold—not a statistical Q–Q plot.
Because extinction is absorbing in the closed SIR model:
The resulting curve must be non-decreasing. If it falls, either the event is not absorbing, the event definition changed, or the code is incorrect.
- Keep the vertical probability axis between 0 and 1.
- State the denominator \(M\).
- State the initial condition and parameter scenario.
- Do not call an estimated probability exact.
4. Numerical tables
Graphs reveal patterns; tables preserve exact reported values. A probability table should normally contain:
| Field | Why it is needed |
|---|---|
| Event | Defines precisely what was counted. |
| Threshold or time | Prevents ambiguous terms such as “major” or “early.” |
| Count | Shows the number of simulations satisfying the event. |
| Total simulations | Provides the denominator. |
| Probability estimate | Reports count divided by total simulations. |
| Monte Carlo uncertainty | Shows the numerical precision of the simulation estimate. |
Round values consistently, but retain enough digits to reflect the Monte Carlo precision. Reporting many more decimal places than the standard error supports gives false precision.
One scenario, four complementary outputs
The runnable example uses one shared dataset from 600 SIR simulations:
The major-outbreak threshold is 20 people ever infected. All outputs use the same simulations, ensuring that differences between displays arise from the questions they answer—not from different model settings.
Interactive Python laboratory
Click Run code to generate a communication-ready summary table and three deliberately separate figures.
Output
Run the code to see the result.
Understand the communication code
| Code | Communication purpose |
|---|---|
alpha=0.45 | Makes overlapping paths partly transparent so several remain visible. |
where="post" | Displays fixed-step DTMC states without suggesting continuous fractional changes. |
bins=... | Defines the outcome intervals counted by the histogram. |
plt.axvline(...) | Marks the stated major-outbreak threshold on the outcome distribution. |
(all_I_paths == 0).mean(axis=0) | Estimates extinction probability separately at every common time. |
plt.ylim(0, 1.03) | Keeps a probability axis anchored to its meaningful 0-to-1 range. |
.quantile(0.10) | Calculates an empirical lower percentile for the outcome table. |
Common display mistakes
| Mistake | Why it misleads | Correction |
|---|---|---|
| One path labelled “the prediction” | Hides stochastic variability. | Call it one realisation and show distributional summaries. |
| A smooth line through DTMC states | Suggests continuous fractional movement. | Use a step plot. |
| Histogram with no vertical-axis definition | Counts and densities have different meanings. | Label count, frequency or density explicitly. |
| Probability axis truncated near the estimate | Can exaggerate small differences. | Normally show the meaningful 0-to-1 probability range. |
| Threshold used but not displayed | The event definition is hidden. | State it in text, table and graph. |
| Too many coloured paths | Creates clutter without adding meaning. | Use a restrained selection and transparency. |
| Incomplete paths treated as final outcomes | Understates possible later infections. | Extend the horizon or report censoring. |
| Excessive decimal places | Implies unsupported precision. | Round consistently with the Monte Carlo error. |
A five-part biological interpretation
Example interpretation: “For a closed population of 100 beginning with two infectious people, the simulations produced both early fade-outs and large outbreaks. Final epidemic sizes were therefore highly variable rather than concentrated around one deterministic outcome. Using a pre-specified threshold of 20 infections, the major-outbreak probability was estimated from 600 independent trajectories and reported with its Monte Carlo interval. These results are conditional on fixed parameters, homogeneous mixing and the one-event DTMC approximation.”
Final checklist before publishing a DTMC result
- Is the biological scenario stated?
- Are parameter values and units supplied?
- Is the initial state supplied?
- Is the time step stated and valid?
- Is the number of simulations reported?
- Is the random-seed policy reproducible?
- Are paths distinguished from distributions?
- Are all graph axes labelled with units?
- Is a step plot used for state trajectories?
- Is every threshold defined and displayed?
- Are incomplete trajectories identified?
- Is Monte Carlo uncertainty reported?
- Are probability axes interpreted between 0 and 1?
- Are tables rounded consistently?
- Does the conclusion remain conditional on model assumptions?
What this lesson has added
- Choose the display from the scientific question.
- Use sample paths for individual histories.
- Use histograms for empirical outcome distributions.
- Use probability curves for event probability through time.
- Use tables for precise definitions, estimates and uncertainty.
- Keep trajectories, distribution summaries and probabilities conceptually separate.
- Interpret results through the biological scenario and stated limitations.
This completes the DTMC programming pathway. The next section introduces continuous-time Markov chains, where event times are random rather than restricted to a fixed time grid.