Markov property
A stochastic process describes a random state changing through time. To predict what may happen next, we might imagine that we need to know the system's entire history. A Markov process has an important simplifying property: once its present state is known, the model does not need the earlier history to determine the probabilities of future states.
An intuitive example
Suppose \(I_n\) is the number of infectious individuals on day \(n\). Imagine that today
\[I_n=10.\]In a simple Markov epidemic model, the probabilities for tomorrow are calculated from the current state \(I_n=10\). We do not additionally need to know whether yesterday there were 8, 12 or 20 infectious individuals.
\(I_{n-1}=8\)
\(I_n=10\)
\(I_{n+1}=?\)
If another epidemic also has \(I_n=10\) today but arrived there through a different history, the Markov model gives it the same probabilities for the next transition.
The mathematical statement
For a discrete-time Markov process,
\[P(X_{n+1}=j\mid X_n=i,X_{n-1},X_{n-2},\ldots)=P(X_{n+1}=j\mid X_n=i).\]This looks complicated, but the meaning is simple.
| Part | Meaning |
|---|---|
| \(X_n=i\) | the process is currently in state \(i\) |
| \(X_{n+1}=j\) | the process will next be in state \(j\) |
| \(X_{n-1},X_{n-2},\ldots\) | the earlier history |
| \(P(\cdot\mid\cdot)\) | conditional probability: probability given some information |
The left side says: calculate the probability of the next state when we know the present and the past. The right side says: calculate it knowing only the present. The Markov property says these are equal.
A numerical example
Suppose a simple model says that when \(I_n=10\), the next day's probabilities are
\[P(I_{n+1}=11\mid I_n=10)=0.30,\]\[P(I_{n+1}=9\mid I_n=10)=0.20,\]\[P(I_{n+1}=10\mid I_n=10)=0.50.\]Now consider two different histories:
| Epidemic A | Epidemic B |
|---|---|
| \(6\to7\to8\to10\) | \(18\to14\to12\to10\) |
| arrived at 10 while increasing | arrived at 10 while decreasing |
Both are currently at \(I_n=10\). Under this Markov model, both therefore use exactly the same next-step probabilities: 0.30 for 11, 0.20 for 9 and 0.50 for staying at 10.
It does not mean the future is independent of the present
This is an important distinction. A Markov process is not saying that the future is unrelated to the past in every sense. The present state was itself produced by the past, and the future usually depends strongly on the present.
The idea is that the effect of the past is summarised by the present state. Once that state is known, separately knowing the earlier states gives no additional information required by the model.
Why the definition of the state matters
Whether a model is Markovian can depend on what we choose to include in its state.
Suppose we describe an epidemic using only
\[X(t)=I(t).\]If recovery probability depends on how long each person has already been infected, then knowing only the total number \(I(t)\) may not be enough.
Therefore \(I(t)\) alone would not provide a Markov state for that model.
Expanding the state can recover the Markov property
The problem may be solved by including the missing information in the state. Instead of recording only the total infectious population, we could distinguish infection stages, for example
\[\mathbf X(t)=\big(S(t),I_1(t),I_2(t),I_3(t),R(t)\big),\]where \(I_1,I_2,I_3\) represent successive stages of infection.
A simple SIS Markov model
Consider a population of size \(N\), with \(I(t)=i\) infectious people and therefore \(S(t)=N-i\) susceptible people. Suppose the current infection and recovery rates are
\[b(i)=\beta\frac{(N-i)i}{N},\qquad d(i)=\gamma i.\]Both rates are determined entirely by the current value \(i\). Over a sufficiently small interval \(\Delta t\),
| Possible change | Approximate probability |
|---|---|
| \(i\to i+1\) | \(b(i)\Delta t\) |
| \(i\to i-1\) | \(d(i)\Delta t\) |
| \(i\to i\) | \(1-[b(i)+d(i)]\Delta t\) |
If we know the present state \(i\), we can calculate these probabilities without knowing the sequence of earlier values that led to \(i\). This is the Markov idea in an epidemic model.
What does “memoryless” mean?
The Markov property is often described as memorylessness. This phrase can be misleading if taken literally.
For example, immune history, age, infection duration or previous treatment may matter biologically. If they affect future behaviour, they must either be included in the state or the resulting model will generally not have the Markov property with respect to the smaller state description.
Markov property and exponential waiting times
In a continuous-time Markov chain, there is another closely related memoryless idea. If the total event rate in state \(i\) is \(a(i)\), the waiting time \(T\) until the next event is exponentially distributed:
\[T\sim\operatorname{Exp}(a(i)).\]The exponential distribution satisfies
\[P(T>s+t\mid T>s)=P(T>t).\]In words: if no event has occurred during the first \(s\) units of time, the remaining waiting-time distribution is the same as if we had just started waiting.
This exponential waiting-time property is central to continuous-time Markov chains and will be developed more carefully in the sections on CTMCs and exponential waiting times.
When might a biological model not be Markovian?
| Biological feature | Why the current simple state may be insufficient |
|---|---|
| infection age | infectiousness or recovery may depend on time since infection |
| waning immunity | future susceptibility may depend on time since recovery or vaccination |
| organism age | birth, death or disease rates may depend on age |
| previous treatment | future response may depend on treatment history |
| delayed biological response | future change may depend explicitly on an earlier state |
These features do not make stochastic modelling impossible. They mean that we must choose a richer state or use a model that explicitly incorporates history.
Markov does not mean constant probabilities
A common misunderstanding is that a Markov process must use the same transition probabilities forever. It does not.
The probabilities can change when the current state changes. For example, in the SIS rate
\[b(i)=\beta\frac{(N-i)i}{N},\]the infection rate is different for different values of \(i\). The model can still be Markovian because the rate is determined by the present state.
Some Markov models can also have transition rules that explicitly depend on time. The essential condition is still that, once the required present information is specified, additional past history is not needed.
Why is the Markov property useful?
The property greatly simplifies stochastic modelling. Instead of tracking every possible history of the biological system, we can describe transitions from the current state.
This makes it possible to construct:
| Model | What describes the transitions? |
|---|---|
| discrete-time Markov chain | transition probabilities and a transition matrix |
| continuous-time Markov chain | transition rates and a generator matrix |
From these local transition rules we can study distributions through time, extinction probabilities, outbreak probabilities, expected values and many other quantities.
Markov property versus deterministic modelling
The Markov property is not the difference between deterministic and stochastic models. It is a property of certain stochastic processes.
A deterministic model gives a fixed future trajectory once the initial state and parameters are fixed. A Markov stochastic model instead gives probabilities for possible future states, with those probabilities determined from the current state.
Where we go next
The Markov property gives the principle. The next step is to turn that principle into an actual model by specifying states and transition probabilities at discrete time steps. This leads to the discrete-time Markov chain (DTMC).