Mean and variance
At any fixed time \(t\), the solution of a stochastic differential equation is a random variable. Different simulated paths can give different values of \(X_t\), even when they start from the same initial condition and use the same model parameters.
One time, many possible values
Imagine simulating the same SDE many times and then looking only at time \(t=5\). We might obtain values such as
\[3.1,\quad 4.0,\quad 5.4,\quad 2.8,\quad 6.2,\ldots\]These are different realisations of the random variable \(X_5\).
The mean
The mean, or expectation, is
\[\boxed{m(t)=E[X_t]}.\]It is the average value we would obtain across a very large number of independent realisations of the process at time \(t\).
The mean is not one typical trajectory
This distinction matters. The curve \(E[X_t]\) is formed by averaging across many possible stochastic paths at each time. It need not itself be one path that the system actually follows.
The variance
The variance measures the average squared distance from the mean:
\[\boxed{V(t)=\operatorname{Var}(X_t)=E[(X_t-m(t))^2]}.\]If the possible values are tightly clustered around the mean, the variance is small. If they are widely dispersed, the variance is large.
Why do we square the deviations?
If we simply averaged deviations from the mean, positive and negative deviations would cancel. Squaring makes every contribution non-negative:
\[(X_t-m)^2\ge0.\]Large deviations are also given more weight than small deviations.
The equivalent variance formula
Expanding the square gives the useful identity
\[\boxed{\operatorname{Var}(X_t)=E[X_t^2]-E[X_t]^2}.\]This form is especially useful in stochastic differential equations because Itô's lemma can often be used to derive an equation for \(X_t^2\).
Standard deviation
The standard deviation is
\[\boxed{\operatorname{SD}(X_t)=\sqrt{\operatorname{Var}(X_t)}}.\]Unlike variance, standard deviation has the same units as \(X_t\), so it is often easier to interpret biologically.
Simple SDE with known mean and variance
Consider
\[dX_t=\mu\,dt+\sigma\,dW_t,\qquad X_0=x_0.\]The solution is
\[X_t=x_0+\mu t+\sigma W_t.\]Since
\[E[W_t]=0,\]the mean is
\[\boxed{E[X_t]=x_0+\mu t}.\]Since
\[\operatorname{Var}(W_t)=t,\]the variance is
\[\boxed{\operatorname{Var}(X_t)=\sigma^2t}.\]What this looks like through time
Why the band widens like \(\sqrt{t}\)
The variance grows linearly:
\[\sigma^2t,\]so the standard deviation grows as
\[\sigma\sqrt{t}.\]This is the same square-root scaling seen in Brownian motion.
Monte Carlo estimates of mean and variance
If we simulate \(M\) independent paths and record their values at time \(t\), the Monte Carlo mean is
\[\boxed{\widehat m(t)=\frac1M\sum_{r=1}^{M}X_t^{(r)}}.\]A sample variance is
\[\boxed{\widehat V(t)=\frac1{M-1}\sum_{r=1}^{M}\left(X_t^{(r)}-\widehat m(t)\right)^2}.\]As the number of independent paths increases, these estimates usually become more stable.
Mean and variance from moment equations
Simulation is not the only route. In some SDEs, we can derive differential equations for moments such as
\[E[X_t],\qquad E[X_t^2],\qquad E[X_t^3],\ldots\]Then
\[\operatorname{Var}(X_t)=E[X_t^2]-E[X_t]^2.\]Deriving the mean equation
For a general one-dimensional Itô SDE
\[dX_t=a(X_t,t)dt+b(X_t,t)dW_t,\]taking expectations gives, under suitable conditions,
\[\frac{d}{dt}E[X_t]=E[a(X_t,t)].\]The stochastic Itô integral has mean zero, so it does not appear directly in this first-moment equation.
Why nonlinear drift creates a problem
Suppose
\[a(X)=rX\left(1-\frac{X}{K}\right).\]Then
\[E[a(X)]=rE[X]-\frac{r}{K}E[X^2].\]So the equation for the mean depends on the second moment.
Second moment from Itô's lemma
Take
\[f(X)=X^2.\]For
\[dX=a(X,t)dt+b(X,t)dW_t,\]Itô's lemma gives
\[d(X^2)=\left[2Xa(X,t)+b(X,t)^2\right]dt+2Xb(X,t)dW_t.\]Taking expectations gives
\[\boxed{\frac{d}{dt}E[X_t^2]=E\left[2X_ta(X_t,t)+b(X_t,t)^2\right]}.\]This is one of the main routes to a variance equation.
Why diffusion appears explicitly in the second moment
The term
\[b^2\]comes from the Itô correction. This is how the diffusion mechanism directly contributes to the growth of the second moment and therefore to variance.
A complete check for constant coefficients
If
\[a=\mu,\qquad b=\sigma,\]then
\[\frac{d}{dt}E[X^2]=2\mu E[X]+\sigma^2.\]Since
\[E[X]=x_0+\mu t,\]we can recover
\[\operatorname{Var}(X_t)=\sigma^2t.\]Why mean alone can be misleading
Two stochastic models can have the same mean and very different uncertainty.
For example, at time \(t\), both may have
\[E[X_t]=10,\]but one may have variance 1 while another has variance 25. The second model allows a much wider range of outcomes.
Biological interpretation
| Quantity | Question |
|---|---|
| mean \(E[X_t]\) | What is the average state across possible biological histories? |
| variance \(\operatorname{Var}(X_t)\) | How uncertain is that state? |
| standard deviation | What is a typical scale of deviation from the mean? |
| higher moments | What additional features of the distribution, such as asymmetry, may matter? |
In epidemic modelling
If \(X_t=I_t\) is an infectious population, then
\[E[I_t]\]describes the average infectious level across repeated stochastic epidemics, while
\[\operatorname{Var}(I_t)\]describes how variable those epidemic outcomes are.
A mean trajectory alone cannot tell us the probability of a very large epidemic or the chance of extinction. For those questions, the distribution or additional probability calculations are needed.
Moment closure
For nonlinear models, equations for low-order moments often depend on higher-order moments. For example, the mean equation may require \(E[X^2]\), while the second-moment equation may require \(E[X^3]\).
This creates an infinite hierarchy called the moment-closure problem.
Moment-closure approximations replace higher moments with expressions involving lower moments so that a finite system can be solved.