Kolmogorov equations
The Kolmogorov equations describe how the probabilities of a continuous-time Markov chain (CTMC) change through time. The generator matrix gives the instantaneous transition rates; the Kolmogorov equations turn those rates into time-dependent probabilities.
Why do we need these equations?
A CTMC can generate one random trajectory by sampling waiting times and events. But often we want a different question:
\[\text{What is the probability of being in each state at time }t?\]Instead of simulating one possible path, the Kolmogorov equations describe the whole probability distribution through time.
State probabilities
Suppose the states are \(0,1,\ldots,N\). Define
\[\pi_i(t)=P(X(t)=i).\]The row probability vector is
\[\boldsymbol\pi(t)=(\pi_0(t),\pi_1(t),\ldots,\pi_N(t)).\]Because the process must be in one of the states,
\[\sum_i\pi_i(t)=1.\]The probability-flow idea
Consider one state \(i\). Its probability can change for two reasons:
Therefore
\[\boxed{\text{rate of change} = \text{inflow rate} - \text{outflow rate}}.\]This simple balance is the intuitive basis of the forward equations.
A two-state example first
Suppose a biological system switches between states 0 and 1:
\[0\xrightarrow{\alpha}1,\qquad1\xrightarrow{\beta}0.\]Let
\[\pi_0(t)=P(X(t)=0),\qquad\pi_1(t)=P(X(t)=1).\]Deriving the equation for state 1
Probability enters state 1 from state 0. The amount of probability currently in state 0 is \(\pi_0(t)\), and its transition rate to state 1 is \(\alpha\). Therefore the inflow rate is
\[\alpha\pi_0(t).\]Probability leaves state 1 for state 0 at rate \(\beta\), so the outflow rate is
\[\beta\pi_1(t).\]Hence
\[\boxed{\frac{d\pi_1}{dt}=\alpha\pi_0-\beta\pi_1}.\]Similarly,
\[\boxed{\frac{d\pi_0}{dt}=\beta\pi_1-\alpha\pi_0}.\]Why is probability conserved?
Add the two equations:
\[\frac{d\pi_0}{dt}+\frac{d\pi_1}{dt}=(\beta\pi_1-\alpha\pi_0)+(\alpha\pi_0-\beta\pi_1)=0.\]Therefore
\[\frac{d}{dt}(\pi_0+\pi_1)=0.\]If the probabilities initially sum to 1, they continue to sum to 1. Probability moves between states but is neither created nor destroyed.
The generator for the two-state example
The generator is
\[Q=\begin{pmatrix}-\alpha&\alpha\\\beta&-\beta\end{pmatrix}.\]The probability equations can be written together as
\[\boxed{\frac{d\boldsymbol\pi(t)}{dt}=\boldsymbol\pi(t)Q}.\]This compact matrix equation is the Kolmogorov forward equation for the state distribution.
Understanding the matrix multiplication
The \(j\)-th component of \(\boldsymbol\pi Q\) is
\[\sum_i\pi_i(t)q_{ij}.\]For \(i\ne j\), each term \(\pi_iq_{ij}\) represents probability currently in state \(i\) multiplied by the rate of moving from \(i\) to \(j\). These are inflows into \(j\).
The diagonal contribution
\[\pi_jq_{jj}\]is negative because \(q_{jj}\) is minus the total leaving rate. It represents outflow from state \(j\).
The general forward equation for one state
For state \(j\),
\[\boxed{\frac{d\pi_j(t)}{dt}=\sum_{i\ne j}\pi_i(t)q_{ij}-\pi_j(t)\sum_{k\ne j}q_{jk}}.\]The first sum is all probability flowing into \(j\). The second is all probability flowing out of \(j\).
Because
\[q_{jj}=-\sum_{k\ne j}q_{jk},\]this becomes
\[\frac{d\pi_j(t)}{dt}=\sum_i\pi_i(t)q_{ij}.\]Birth–death processes
For a birth–death process, state \(i\) can move only to \(i+1\) at rate \(b(i)\) or to \(i-1\) at rate \(d(i)\).
Deriving the birth–death forward equation
Let
\[p_i(t)=P(X(t)=i).\]Probability enters state \(i\) in two ways:
\[i-1\to i\quad\text{at rate }b(i-1),\]\[i+1\to i\quad\text{at rate }d(i+1).\]Therefore the inflow rate is
\[b(i-1)p_{i-1}(t)+d(i+1)p_{i+1}(t).\]Probability leaves state \(i\) through
\[i\to i+1\quad\text{at rate }b(i),\qquad i\to i-1\quad\text{at rate }d(i).\]Therefore the outflow rate is
\[[b(i)+d(i)]p_i(t).\]Combining inflow and outflow gives
\[\boxed{\frac{dp_i}{dt}=b(i-1)p_{i-1}+d(i+1)p_{i+1}-[b(i)+d(i)]p_i}.\]SIS epidemic example
For a stochastic SIS epidemic with infectious count \(i\),
\[b(i)=\beta\frac{(N-i)i}{N},\qquad d(i)=\gamma i.\]Substituting these rates into the birth–death equation gives the evolution of the probability that exactly \(i\) people are infectious at time \(t\).
Writing the structure before expanding the rates is often clearer:
\[\frac{dp_i}{dt}=b(i-1)p_{i-1}+d(i+1)p_{i+1}-[b(i)+d(i)]p_i.\]This is fundamentally different from the deterministic SIS equation: here \(p_i(t)\) is a probability of a discrete epidemic state, not the infectious population itself.
Why this gives more than a deterministic trajectory
A deterministic model gives one value such as \(I(t)\). The Kolmogorov equations give the entire distribution
\[p_0(t),p_1(t),\ldots,p_N(t).\]From this distribution we can calculate quantities such as
\[P(I(t)=0)=p_0(t),\]which is the probability that infection is extinct by time \(t\), and
\[P(I(t)\ge m)=\sum_{i=m}^{N}p_i(t),\]which can represent the probability that an epidemic exceeds a chosen level.
Transition probabilities rather than one initial distribution
So far we have followed one probability vector \(\boldsymbol\pi(t)\). We can instead ask for transition probabilities
\[p_{ij}(t)=P(X(t)=j\mid X(0)=i).\]Collect all of them into the transition matrix
\[P(t)=(p_{ij}(t)).\]Each row corresponds to a different possible initial state.
The Kolmogorov forward equation
For a finite time-homogeneous CTMC using our row convention,
\[\boxed{\frac{dP(t)}{dt}=P(t)Q}.\]This is the matrix version of the probability-flow equations derived above.
The generator acts on the destination side: probability that has reached intermediate states then moves onward according to the rates in \(Q\).
Where does the forward equation come from?
The Markov property gives the Chapman–Kolmogorov relation
\[P(t+h)=P(t)P(h).\]For a very small \(h\),
\[P(h)=I+hQ+o(h).\]Therefore
\[P(t+h)=P(t)[I+hQ+o(h)].\]Subtract \(P(t)\):
\[P(t+h)-P(t)=hP(t)Q+o(h).\]Divide by \(h\) and let \(h\to0\):
\[\boxed{P'(t)=P(t)Q}.\]This derivation shows exactly how the local rates in \(Q\) determine probability evolution over time.
The Kolmogorov backward equation
The Chapman–Kolmogorov relation can also be written
\[P(t+h)=P(h)P(t).\]Using
\[P(h)=I+hQ+o(h),\]gives
\[P(t+h)=[I+hQ+o(h)]P(t).\]Following the same limit gives
\[\boxed{\frac{dP(t)}{dt}=QP(t)}.\]This is the Kolmogorov backward equation.
What is the intuitive difference between forward and backward?
| Forward equation | Backward equation |
|---|---|
| \(P'(t)=P(t)Q\) | \(P'(t)=QP(t)\) |
| focuses naturally on probability flow into and out of destination states | focuses naturally on how the first transition from the initial state affects later outcomes |
| especially intuitive for evolving a current probability distribution | especially useful for quantities viewed as functions of the starting state |
For a finite time-homogeneous CTMC, both equations describe the same transition matrix.
Why can both equations be true?
The solution is
\[\boxed{P(t)=e^{tQ}}.\]Because a matrix commutes with any power series in itself,
\[Qe^{tQ}=e^{tQ}Q.\]Therefore
\[P'(t)=QP(t)=P(t)Q.\]There is no contradiction between the two equations.
Initial condition
At time zero, no time has passed, so a process starting in state \(i\) is still in state \(i\). Hence
\[\boxed{P(0)=I}.\]The matrix differential equation
\[P'(t)=P(t)Q,\qquad P(0)=I\]has solution
\[P(t)=e^{tQ}.\]Probability-vector form
If the initial distribution is \(\boldsymbol\pi(0)\), then
\[\boldsymbol\pi(t)=\boldsymbol\pi(0)P(t).\]Using \(P(t)=e^{tQ}\),
\[\boxed{\boldsymbol\pi(t)=\boldsymbol\pi(0)e^{tQ}}.\]Differentiating gives
\[\boxed{\frac{d\boldsymbol\pi(t)}{dt}=\boldsymbol\pi(t)Q}.\]Boundary equations matter
The general birth–death equation applies to interior states, but boundary states must be treated according to the biology.
For example, if state 0 is absorbing, probability can enter state 0 through a transition \(1\to0\), but it cannot leave. Therefore
\[\frac{dp_0}{dt}=d(1)p_1(t).\]In an epidemic, this describes probability accumulating in the extinction state.
Kolmogorov equations versus simulation
| Kolmogorov equations | Stochastic simulation |
|---|---|
| evolve the complete probability distribution | generate one random trajectory at a time |
| deterministic equations for probabilities | random sequence of events and waiting times |
| can be exact for manageable finite state spaces | useful even when the state space is too large for all probability equations |
| state space can make the system very large | many repetitions are needed to estimate probabilities accurately |
These are not competing descriptions. They are two ways of studying the same stochastic model.
Why the equations become difficult for large biological models
For a simple SIS model with population size \(N\), there are \(N+1\) possible infectious counts. But a model with several compartments may have a very large number of possible combinations of \(S,E,I,R,\ldots\).
The Kolmogorov equations then require one probability equation for every possible state. This rapid growth of the state space is one reason simulation methods such as the Gillespie algorithm are so important.
The complete picture
What comes next?
The Kolmogorov equations describe the full probability distribution mathematically. When the state space becomes large, solving all these equations can be impractical. The next section develops the Gillespie algorithm, which generates exact CTMC sample paths directly from the event rates.