Wright–Fisher model
The Wright–Fisher model is a classical stochastic model of genetic drift in populations with discrete, non-overlapping generations.
Basic assumptions
In the simplest neutral haploid version, the model assumes a fixed population size \(N\), discrete non-overlapping generations, random reproduction, and no selection, mutation or migration.
Each of the \(N\) allele copies in the next generation is sampled independently from the current gene pool.
State variable
Let
\[X_t=\text{number of copies of allele }A\text{ in generation }t.\]Then
\[X_t\in\{0,1,\ldots,N\}.\]The allele frequency is
\[\boxed{p_t=\frac{X_t}{N}}.\]Binomial sampling
If the current frequency is \(p_t\), each new copy is \(A\) with probability \(p_t\). Therefore
\[\boxed{X_{t+1}\mid X_t\sim\operatorname{Binomial}(N,p_t)}.\]Equivalently, if \(X_t=i\), then \(p_t=i/N\).
Transition probabilities
The probability of moving from \(i\) copies of \(A\) to \(j\) copies in one generation is
\[\boxed{P_{ij}=P(X_{t+1}=j\mid X_t=i)=\binom{N}{j}\left(\frac{i}{N}\right)^j\left(1-\frac{i}{N}\right)^{N-j}}.\]This means the Wright–Fisher model is a discrete-time Markov chain on the states \(0,1,\ldots,N\).
Expected allele frequency
For a binomial random variable,
\[E[X_{t+1}\mid p_t]=Np_t.\]Therefore
\[\boxed{E[p_{t+1}\mid p_t]=p_t}.\]So neutral drift has no systematic directional bias in expectation.
Variance of the next frequency
Because
\[\operatorname{Var}(X_{t+1}\mid p_t)=Np_t(1-p_t),\]we obtain
\[\boxed{\operatorname{Var}(p_{t+1}\mid p_t)=\frac{p_t(1-p_t)}{N}}.\]This is the haploid form. For a diploid population of \(N\) individuals there are \(2N\) gene copies, giving denominator \(2N\).
A simple example
If \(N=10\) and \(X_t=4\), then
\[p_t=0.4.\]The next generation satisfies
\[X_{t+1}\sim\operatorname{Binomial}(10,0.4).\]The probability of obtaining exactly six copies is
\[P(X_{t+1}=6)=\binom{10}{6}(0.4)^6(0.6)^4.\]The next frequency would then be \(6/10=0.6\).
Possible jumps in one generation
Unlike the Moran model, Wright–Fisher can change by many copies in a single generation. From state \(i\), the next state can be any
\[j\in\{0,1,\ldots,N\},\]although some transitions may have very small probability.
Absorbing states
If
\[X_t=0,\]then \(p_t=0\), so the next generation contains no \(A\) copies with probability one. Thus state \(0\) is absorbing.
Similarly, state \(N\) is absorbing. These correspond to allele loss and fixation.
Neutral fixation probability
Because the neutral allele frequency is a martingale and the only absorbing states are \(0\) and \(1\), the probability of eventual fixation equals the initial frequency:
\[\boxed{P(\text{fixation of }A\mid p_0)=p_0}.\]For one new neutral copy in a haploid population of size \(N\),
\[p_0=\frac1N,\]so the fixation probability is \(1/N\). In a diploid population of \(N\) individuals, a single copy begins at frequency \(1/(2N)\).
Transition matrix
For fixed \(N\), the model has a transition matrix
\[P=(P_{ij})_{i,j=0}^{N}.\]Each row sums to one:
\[\sum_{j=0}^{N}P_{ij}=1.\]Repeated multiplication of a probability distribution by \(P\) gives the distribution of allele counts in future generations.
Why the model is stochastic
Two populations can have the same \(N\), the same starting frequency and the same biological rules, yet obtain different next-generation samples.
Therefore the model predicts a distribution of possible outcomes rather than one unique trajectory.
Adding selection
Suppose allele \(A\) has fitness \(w_A\) and allele \(a\) has fitness \(w_a\). Selection first changes the effective sampling probability to
\[\boxed{p_s=\frac{p_tw_A}{p_tw_A+(1-p_t)w_a}}.\]The next generation is then sampled as
\[\boxed{X_{t+1}\sim\operatorname{Binomial}(N,p_s)}.\]Selection changes the centre of the sampling distribution, while drift remains because the final draw is still random.
Adding mutation
If \(A\to a\) occurs with probability \(\mu\) and \(a\to A\) with probability \(\nu\), a mutation-adjusted sampling probability is
\[\boxed{p_m=(1-\mu)p+\nu(1-p)}.\]The next generation can then be sampled from \(\operatorname{Binomial}(N,p_m)\), or this mutation step can be combined with selection according to the chosen life-cycle order.
Adding migration
If a fraction \(m\) of the gene pool is replaced by migrants with allele frequency \(p_M\), a simple migration step is
\[\boxed{p_{mig}=(1-m)p+mp_M}.\]This adjusted frequency can then become the sampling probability for the next generation.
Order of evolutionary processes
When selection, mutation and migration are combined, the exact recurrence depends on the order in which the model applies these processes.
A model should therefore state whether, for example, selection occurs before mutation or mutation before selection.
Large-population approximation
As \(N\) increases, the one-generation variance
\[\frac{p(1-p)}N\]decreases. This is why deterministic allele-frequency equations can become useful large-population approximations.
However, even weak drift can matter over long times or when an allele is initially rare.
Simulation
A Wright–Fisher trajectory can be simulated by repeating two steps:
calculate the current sampling probability, then draw the next allele count from a binomial distribution.
Repeated simulations reveal the range of possible trajectories, fixation probabilities and time-to-absorption behaviour.
Wright–Fisher versus Moran
| Wright–Fisher | Moran |
|---|---|
| whole generation replaced at once | one birth–death replacement event at a time |
| large jumps in allele count are possible | allele count changes by at most one per event |
| natural discrete-generation model | natural event-by-event model |
The next lesson develops the Moran process in detail.