Network centrality
Centrality measures assign numerical scores to nodes according to different ideas of structural importance. A node can be important because it has many neighbours, lies between communities, is close to the rest of the network, or is connected to other important nodes.
Degree centrality
For a simple undirected graph, the degree of node \(i\) is
\[\boxed{k_i=\sum_jA_{ij}}.\]A normalised degree centrality is
\[\boxed{C_D(i)=\frac{k_i}{n-1}}.\]This lies between zero and one for a simple graph with \(n\) nodes.
Biological interpretation of degree
In a contact network, high degree means many direct contact partners. In a gene-regulatory network, high out-degree may identify a regulator with many targets. In a food web, degree may reflect broad interaction range.
Degree is local: it does not describe where those neighbours sit in the wider network.
In-degree and out-degree
Directed networks require separate measures. With the convention \(A_{ij}=1\) for \(i\to j\),
\[k_i^{\mathrm{out}}=\sum_jA_{ij},\qquad k_i^{\mathrm{in}}=\sum_jA_{ji}.\]The two scores can have very different biological interpretations. A transcription factor may regulate many targets while itself being regulated by relatively few genes.
Strength centrality
For weighted networks, counting neighbours can discard important information. Node strength is
\[\boxed{s_i=\sum_jw_{ij}}.\]A node with a few strong interactions can therefore differ from one with many weak interactions even if ordinary degree suggests the opposite ranking.
Closeness centrality
Let \(d(i,j)\) be the shortest-path distance between nodes \(i\) and \(j\). For a connected graph, a common normalised closeness centrality is
\[\boxed{C_C(i)=\frac{n-1}{\sum_{j\ne i}d(i,j)}}.\]A node has high closeness when it can reach other nodes through relatively short paths.
Biological interpretation of closeness
Closeness may be relevant when a process can propagate through paths and short access to the whole network matters.
However, shortest paths are not always the routes followed by biological processes. Infection, signalling or ecological effects may spread stochastically along many possible paths.
Disconnected networks
Ordinary closeness becomes problematic when some distances are infinite. One alternative is harmonic centrality:
\[\boxed{C_H(i)=\sum_{j\ne i}\frac{1}{d(i,j)}},\]where disconnected pairs contribute zero.
This avoids assigning an infinite denominator when the graph contains several components.
Betweenness centrality
Let \(\sigma_{st}\) be the number of shortest paths between nodes \(s\) and \(t\), and let \(\sigma_{st}(i)\) be the number of those paths that pass through node \(i\).
Betweenness centrality is
\[\boxed{C_B(i)=\sum_{s\ne i\ne t}\frac{\sigma_{st}(i)}{\sigma_{st}}}.\]Normalisation factors can be applied when comparing networks of different sizes.
Bridge nodes
A node can have modest degree but high betweenness if it connects otherwise weakly connected regions.
In an epidemic network, such a node may provide a route between communities. In an ecological or molecular network, it may connect modules.
Removing a high-betweenness node can therefore have a larger structural effect than its degree alone suggests.
Shortest-path assumption
Betweenness assumes that shortest paths are relevant to the process. This is natural for some routing problems but less direct for biological spreading, where trajectories may not preferentially follow geodesics.
Eigenvector centrality
Eigenvector centrality gives a node a high score when it is connected to other nodes that also have high scores.
For an undirected network, the centrality vector \(x\) satisfies
\[\boxed{Ax=\lambda x}.\]The usual choice is the eigenvector associated with the largest eigenvalue of \(A\).
Recursive interpretation
The component equation is
\[\lambda x_i=\sum_jA_{ij}x_j.\]Thus the score of node \(i\) depends on the scores of its neighbours.
This differs from degree centrality, where every neighbour contributes equally.
Perron–Frobenius idea
For a connected undirected graph, the adjacency matrix is non-negative and irreducible. Perron–Frobenius theory guarantees a largest eigenvalue with an associated eigenvector that can be chosen strictly positive.
This provides the mathematical basis for using the leading eigenvector as a centrality score.
Disconnected or directed networks
Eigenvector centrality requires care when the graph is disconnected or directed. Some components or nodes may receive zero scores, and left and right eigenvectors can represent different notions of incoming and outgoing influence.
The matrix convention must therefore be stated.
Katz centrality
Katz centrality counts walks of all lengths while reducing the contribution of longer walks. One form is
\[\boxed{x=\alpha Ax+\beta\mathbf1}.\]Rearranging gives
\[\boxed{x=\beta(I-\alpha A)^{-1}\mathbf1},\]provided
\[\alpha<\frac{1}{\rho(A)},\]where \(\rho(A)\) is the spectral radius.
The baseline term \(\beta\) allows nodes to receive non-zero centrality even when ordinary eigenvector centrality is problematic.
Walk interpretation of Katz centrality
When the matrix series converges,
\[(I-\alpha A)^{-1}=I+\alpha A+\alpha^2A^2+\cdots.\]Because \((A^m)_{ij}\) counts walks of length \(m\), Katz centrality incorporates indirect connections of many lengths, discounted by powers of \(\alpha\).
PageRank-type centrality
For directed networks, PageRank-type measures combine movement along directed edges with a small probability of moving independently of the network.
In matrix form, a centrality vector may satisfy
\[\boxed{x=\alpha Px+(1-\alpha)v},\]where \(P\) is a suitable transition matrix and \(v\) is a baseline distribution.
The precise interpretation depends on how edge direction and transition probabilities are defined.
Centrality and epidemic risk
Degree can identify nodes with many immediate transmission opportunities. Eigenvector or spectral measures can reflect placement within highly connected regions, while betweenness can identify bridges between groups.
None of these scores alone gives the exact probability that a person becomes infected or causes infections. Epidemic risk also depends on timing, susceptibility, infectiousness and the states of neighbouring nodes.
Centrality and gene regulation
High out-degree may identify broad regulators, while spectral measures may identify nodes embedded in strongly interconnected regulatory structure.
But a weak regulatory edge and a strong regulatory edge should not necessarily be treated equally, so weighted and signed information may be needed.
Centrality and ecological networks
In ecological networks, degree can describe interaction breadth and betweenness can identify species linking modules.
Structural centrality should not automatically be interpreted as ecological importance. Population abundance, interaction strength and nonlinear dynamics can alter the consequences of species loss.
Centrality rankings can disagree
A hub can have high degree but low betweenness if all of its neighbours belong to one dense community. A low-degree bridge can have high betweenness. A node connected to influential hubs can have high eigenvector centrality despite moderate degree.
Different rankings are therefore expected rather than contradictory.
Centrality depends on the network definition
Changing the edge definition can change centrality dramatically. A contact network aggregated over one day may rank nodes differently from one aggregated over a month.
Likewise, using binary rather than weighted edges can alter which nodes appear important.
Static centrality on temporal networks
If contacts change through time, centrality computed from an aggregated static graph may identify paths that were never available in the correct temporal order.
Temporal centrality measures can instead account for time-respecting paths and changing node roles.
Centrality is descriptive, not automatically causal
A centrality score is calculated from network structure. It does not by itself prove that manipulating the highest-ranked node will produce the greatest biological effect.
Intervention effectiveness should be evaluated using a dynamical model or experiment whenever possible.
Choosing a centrality measure
| Question | Possible measure |
|---|---|
| Who has many direct neighbours? | degree |
| Who has strong total weighted interaction? | strength |
| Who is structurally close to many nodes? | closeness or harmonic centrality |
| Who bridges shortest paths between regions? | betweenness |
| Who connects to other well-connected nodes? | eigenvector centrality |
| Who participates in many direct and indirect walks? | Katz centrality |
Transition to network interventions
Centrality provides possible ways to rank nodes, but ranking is only the beginning of an intervention problem. The next lesson asks how vaccination, isolation, edge removal and other interventions change biological dynamics on a network.