← Data and Parameter Estimation

Biological data

Mathematical biology connects models to observations. Before estimating parameters, we must understand what was measured, how it was measured, when it was measured and how the observation relates to the variables in the model.

Core idea. A model describes a biological process, while a dataset records observations of that process. These are not automatically the same quantities. The observation mechanism forms an important link between model and data.

What counts as biological data?

Biological data can take many forms: population counts, disease cases, concentrations, cell numbers, survival times, gene-expression measurements, spatial locations, contact networks, images or repeated measurements through time.

The mathematical form of the data determines which statistical and modelling methods are appropriate.

Observational units

The observational unit is the entity on which a measurement is made. It might be an individual, cell, household, population, species, experimental well or geographical region.

Correctly identifying the observational unit is essential because repeated measurements from the same unit are usually not independent observations.

Variables and observations

Suppose a biological quantity is represented by a random variable \(Y\). The observed dataset contains realised values

\[y_1,y_2,\ldots,y_n.\]

The uppercase symbol represents the variable before observation, while lowercase values represent observed data.

Continuous data

A continuous measurement can take values over an interval. Examples include concentration, body mass, temperature and time.

A common observation model is

\[\boxed{Y_i=\mu_i+\varepsilon_i},\]

where \(\mu_i\) is the underlying model prediction and \(\varepsilon_i\) represents measurement or observation variation.

Count data

Counts take non-negative integer values:

\[Y_i\in\{0,1,2,\ldots\}.\]

Examples include numbers of infections, cells, births or captured organisms.

Because counts are discrete and their variance may depend on their mean, a Gaussian error model is not automatically appropriate.

Binary data

Some observations have two possible outcomes, such as infected/not infected or survived/died.

They can be represented by

\[Y_i\in\{0,1\}.\]

If

\[P(Y_i=1)=p_i,\]

then

\[Y_i\sim\operatorname{Bernoulli}(p_i).\]

Proportions

A proportion may arise from a numerator and denominator. If \(Y_i\) successes are observed among \(n_i\) trials, then

\[\frac{Y_i}{n_i}\]

is the observed proportion.

Keeping the denominator is important because a proportion based on 10 observations contains less information than the same proportion based on 10,000 observations.

Time-series data

Time-series observations are recorded at ordered times

\[t_1A model trajectory \(x(t;\theta)\) may then be compared with measurements

\[y_i\approx x(t_i;\theta).\]

Measurements close together in time may be correlated, so treating all residuals as independent requires justification.

Longitudinal data

Longitudinal data contain repeated observations of the same biological units through time.

If individual \(j\) is measured at several times, observations from that individual can share biological characteristics and are generally dependent.

This differs from repeatedly sampling different individuals from the population.

Cross-sectional data

Cross-sectional data describe a population or collection of units at one time or over a short observation window.

They can reveal variation between units but usually contain less direct information about individual temporal trajectories.

Spatial data

Biological observations may have locations \(s_i\), giving data such as

\[Y(s_1),Y(s_2),\ldots,Y(s_n).\]

Nearby observations may be more similar than distant observations. Spatial dependence should therefore be considered rather than automatically assuming independence.

Event-time data

Some datasets record the time until an event such as death, recovery, infection or relapse.

If \(T\) denotes event time, survival analysis studies quantities such as the survival function

\[\boxed{S(t)=P(T>t)}.\]

The observation may be incomplete if the event is not seen during the study.

Censoring

Right censoring occurs when we know only that an event time exceeds a particular value. For example, an individual may still be alive when follow-up ends.

A censored observation should not be treated as though the event occurred at the censoring time.

Exact states versus indirect measurements

A model may contain a state variable \(I(t)\) representing the true number of infectious individuals, while the available data may record reported cases rather than true prevalence.

An observation model might therefore be

\[Y_t\sim\operatorname{Binomial}(I_t,\rho),\]

where \(\rho\) is a simplified reporting probability.

This explicitly distinguishes the latent biological state from what is observed.

Incidence versus prevalence

Prevalence measures how many individuals are in a state at a particular time. Incidence measures new events occurring over an interval.

Thus an epidemic state variable \(I(t)\) is not automatically comparable with a dataset of daily new cases.

Match data to the correct model quantity. Fitting prevalence predictions directly to incidence observations is a conceptual error unless a valid observation relationship has been defined.

Process variability

Biological systems can vary even if measured perfectly. Births, infections, molecular reactions and individual responses may be intrinsically stochastic.

This variability belongs to the biological process itself.

Observation error

Measurements can also differ from the true biological state because of instrument error, reporting, detection limits, sampling or recording.

A useful conceptual decomposition is

\[\boxed{\text{observed variation}=\text{process variation}+\text{observation variation}},\]

although their precise combination depends on the statistical model.

Systematic bias

Random measurement error fluctuates around the underlying value. Systematic bias shifts observations consistently in some direction.

Examples include under-reporting, calibration error or preferential sampling.

Increasing sample size reduces random sampling uncertainty but does not automatically remove systematic bias.

Missing data

Missing observations can arise because measurements fail, individuals leave a study or records are unavailable.

The consequences depend on why values are missing. Missingness related to the unobserved biological state can create substantial bias if ignored.

Detection limits

Laboratory measurements may be recorded only above or below a detection threshold.

A value reported as “below detection limit” is not necessarily zero. Treating it as zero changes the biological meaning of the observation.

Sampling frequency

How often a system is observed affects what can be learned about its dynamics.

If measurements are too widely spaced, rapid peaks, oscillations or transitions may be missed. More frequent sampling can reveal faster dynamics but may increase cost and dependence between observations.

Sample size

Larger samples usually reduce sampling uncertainty, but sample size alone does not guarantee informative data.

Many measurements concentrated in an uninformative region may estimate a parameter less effectively than fewer strategically timed measurements.

Experimental range

Parameters are often easier to estimate when the data contain conditions under which their effects are visible.

For example, a saturation parameter may be difficult to estimate if every observation lies far below the saturation region.

Study design and identifiability are therefore closely connected.

Replication

Replicate observations help quantify biological and measurement variability.

Technical replicates repeat the measurement process, whereas biological replicates involve distinct biological units. They answer different questions and should not automatically be pooled as equivalent observations.

Independence

Many elementary statistical models assume

\[Y_1,Y_2,\ldots,Y_n\quad\text{are independent}.\]

This assumption may fail for repeated measures, spatial data, family structures, networks or time series.

The dependence structure should reflect how the data were generated.

Metadata

Useful datasets require information beyond numerical values. Units, measurement methods, dates, locations, inclusion criteria, missing-value codes and experimental conditions are part of the data-generating context.

Without these details, mathematically correct calculations can still lead to biologically incorrect conclusions.

Data quality checks

Before fitting a model, inspect units, ranges, impossible values, duplicates, missingness, time ordering, abrupt recording changes and known changes in measurement procedure.

An apparent biological signal can sometimes be a data-collection artefact.

Study design

Good study design asks what measurements are needed to distinguish biological mechanisms or estimate the parameters of interest.

Sampling times, sample size, replication, controls and experimental conditions should ideally be chosen with the intended model in mind.

Data do not identify a model automatically

Different mathematical models can sometimes produce very similar observations. A good visual fit therefore does not prove that the assumed biological mechanism is correct.

Parameter identifiability, uncertainty, validation and comparison are addressed later in this section.

Transition to data preparation

Once the biological meaning and observation process are understood, the dataset must be checked and organised without destroying that meaning. The next lesson develops data preparation for mathematical modelling.

Key idea. Biological data are observations generated by both a biological process and a measurement process. Correct modelling begins by identifying the observational unit, data type, timing, dependence, uncertainty and relationship between each measured quantity and the corresponding model variable.