โ† Data and Parameter Estimation

Data preparation

Data preparation converts raw biological observations into a form suitable for mathematical and statistical analysis while preserving their scientific meaning. It should make the data easier to analyse without silently changing what the observations represent.

Core idea. Cleaning data is not the same as making inconvenient observations disappear. Every correction, exclusion, transformation and aggregation should have a scientific or statistical reason and should be reproducible.

Keep the raw data unchanged

The original dataset should normally be preserved as a read-only source. Cleaning and transformation should produce a separate analysis dataset.

This allows every processed value to be traced back to the original observation and prevents accidental loss of information.

Understand the data dictionary

Before changing values, identify what each column means, its units, coding scheme, observational unit and allowed range.

For example, a column called cases might mean new cases per day, cumulative cases, laboratory-confirmed cases or estimated prevalence. These quantities cannot be used interchangeably.

Use a consistent table structure

A useful general structure has one observational unit per row and one variable per column, with identifiers linking repeated measurements to the same biological unit.

Longitudinal data may contain columns such as individual identifier, time and measured response.

Check data types

Numbers stored as text, dates stored in inconsistent formats and categorical labels entered with different spellings can cause analysis errors.

Each variable should be represented according to its meaning: numerical, categorical, binary, date/time or identifier.

Check units

Measurements must use compatible units before comparison with a mathematical model.

If a model parameter is expressed per day but observations are recorded weekly, the time scale must be reconciled explicitly.

Likewise, concentrations measured in different units must be converted before combining datasets.

Unit conversion

If a quantity \(x\) is converted by a known factor \(c\),

\[\boxed{x_{\mathrm{new}}=c\,x_{\mathrm{old}}}.\]

The same conversion must be reflected consistently in model parameters and labels.

Units are part of the mathematics. A numerically successful fit can still be scientifically wrong if the model and data use incompatible units.

Check time variables

Dates should be ordered and converted to a time scale appropriate for the model. If \(t_0\) is the chosen origin, elapsed time can be defined by

\[t_i=\text{date}_i-\text{date}_0.\]

This may produce time measured in days, hours or another consistent unit.

Irregular observation times

Measurements need not occur at equally spaced times:

\[t_{i+1}-t_i\ne\text{constant}.\]

Irregular sampling is not automatically a problem. Many continuous-time models can be evaluated directly at the actual observation times.

Artificially interpolating data onto a regular grid should therefore be done only when the method genuinely requires it.

Duplicate records

Two identical-looking rows may be accidental duplicates or genuine repeated measurements.

Duplicates should not be deleted merely because values match. Their identifiers, timestamps and study design must first be checked.

Impossible values

Some observations violate known constraints. Examples include negative population counts, proportions outside \([0,1]\), impossible dates or concentrations outside an instrument's physical range.

Such values should be investigated rather than automatically replaced.

Range checks

Known biological constraints can be encoded as validation rules. For a proportion \(p_i\), for example,

\[0\le p_i\le1.\]

For compartment counts in a closed population, one may check

\[S_i+E_i+I_i+R_i=N\]

when that conservation law genuinely applies to the recorded quantities.

Missing values

A missing value means that an observation is unavailable. It should be represented explicitly rather than confused with zero.

Zero can be a meaningful biological measurement; missing means that the value is unknown.

Why values are missing

The reason for missingness matters. Missing measurements caused by random equipment failure differ from measurements missing because severely ill participants left a study.

Deleting every incomplete row can introduce bias when missingness is related to the biological outcome.

Imputation

Imputation replaces missing values using an explicit statistical procedure.

Simple replacements such as the sample mean can reduce apparent variability and distort relationships between variables. More principled methods should reflect the data-generating process and uncertainty.

Do not invent biological observations merely to obtain a complete table. Sometimes the correct analysis retains missingness explicitly rather than filling every gap.

Outliers

An outlier is an observation unusually far from the rest of the data. It may represent measurement error, data-entry error or genuine biological variation.

An extreme value should not be removed solely because it makes the fitted model worse.

Investigating outliers

Useful checks include the original record, measurement units, instrument range, neighbouring time points, replicate measurements and known biological events.

If an observation is excluded, the criterion should ideally be defined independently of the desired model result.

Transformations

A transformation changes the mathematical scale of a variable. For positive measurements, a logarithmic transformation is

\[\boxed{z_i=\log y_i}.\]

This can convert multiplicative variation into approximately additive variation and reduce strong right skew.

It also changes the residual model and interpretation of fitting errors.

Log transformation and zeros

Because

\[\log 0\]

is undefined, zeros require careful treatment. Adding an arbitrary constant such as \(1\) changes the data and may have substantial effects when values are small.

The choice should be justified by the measurement process rather than used automatically.

Scaling

A variable can be rescaled by a characteristic value \(a\):

\[z_i=\frac{y_i}{a}.\]

Scaling can improve numerical optimisation when variables have very different magnitudes.

It should not be confused with changing the biological information in the observations.

Standardisation

A common statistical transformation is

\[\boxed{z_i=\frac{y_i-\bar y}{s}},\]

where \(\bar y\) is the sample mean and \(s\) is the sample standard deviation.

This gives a dimensionless variable centred near zero with sample standard deviation one.

It is useful in some statistical models but often unnecessary when mechanistic parameters have meaningful physical units.

Normalisation

The word normalisation is used for several different operations, such as dividing by a population size, total count or reference measurement.

The exact operation should therefore be stated rather than simply saying that data were normalised.

Counts and proportions

If a count \(I(t)\) is converted to a population proportion,

\[i(t)=\frac{I(t)}{N}.\]

A model fitted to \(i(t)\) must use the corresponding proportion scale.

Converting counts to proportions can be useful for comparing populations of different sizes, but it discards absolute scale unless \(N\) is retained.

Aggregation

Daily observations may be aggregated into weekly values. For incidence counts, one possible aggregation is

\[Y_w=\sum_{d\in w}Y_d.\]

For a concentration measurement, summing values would usually have no meaningful interpretation; a mean or another summary might instead be appropriate.

The aggregation rule must match the biological quantity.

Aggregation can hide dynamics

Combining observations reduces temporal or spatial resolution. Rapid peaks and short-lived changes may disappear.

Aggregation can reduce noise but also remove information needed to estimate dynamic parameters.

Cumulative data

If incidence is \(y_t\), cumulative observations are

\[C_t=\sum_{s\le t}y_s.\]

Successive cumulative values share most of the same underlying observations and are therefore strongly dependent.

Fitting cumulative data as though each point were an independent observation can give misleading uncertainty estimates.

Smoothing

Moving averages and other smoothers can reveal broad trends. A simple \(m\)-point moving average has the form

\[\bar y_t=\frac1m\sum_{j=0}^{m-1}y_{t-j}.\]

However, smoothing changes the observation process and induces dependence between neighbouring smoothed values.

Raw data should normally remain available for formal inference.

Align model output with observation times

A numerical model may be solved at many internal time points, but residuals should be constructed at the actual observation times.

If the model predicts \(x(t;\theta)\), then

\[r_i=y_i-x(t_i;\theta).\]

This is preferable to treating unobserved numerical solver points as additional data.

Match incidence intervals correctly

If the model produces a cumulative state \(C(t)\) but observations are interval incidence, the corresponding model prediction over \([t_{i-1},t_i]\) is

\[\boxed{\Delta C_i=C(t_i)-C(t_{i-1})}.\]

This should be compared with interval counts rather than directly comparing \(C(t_i)\) with incidence.

Train and validation data

When predictive assessment is required, some observations can be withheld from parameter fitting and used later for validation.

For time-series models, random shuffling may leak future information into training. A chronological split is often more meaningful.

Avoid data leakage

Any transformation estimated from the data, such as a mean, standard deviation or feature-selection rule, should be calculated using only the appropriate fitting data when independent validation is intended.

Otherwise information from the validation set can influence model construction.

Record exclusions

Every excluded observation should have a traceable reason, such as a predefined quality-control failure or documented measurement problem.

Keeping an exclusion log makes the analysis auditable.

Reproducible preparation

Preparation steps should be encoded in scripts or notebooks whenever possible rather than performed manually in a spreadsheet without a record.

A reproducible workflow can be rerun when new data arrive or when a correction is made upstream.

Separate stages of the workflow

A useful structure is

\[\boxed{\text{raw data}\longrightarrow\text{clean data}\longrightarrow\text{analysis data}\longrightarrow\text{model fitting}}.\]

Separating these stages makes it easier to identify where a change entered the analysis.

Visual inspection

Plots of observations against time, distributions, scatter plots and missingness patterns often reveal problems that summary tables do not.

Visual inspection complements formal validation checks rather than replacing them.

Document every transformation

For each transformation, record the original variable, mathematical operation, units before and after, and reason for the change.

This preserves biological interpretability when fitted parameters are reported later.

Transition to parameter estimation

Once observations have been checked, transformed appropriately and aligned with the model, the next task is to determine which parameter values make the model consistent with those observations.

Key idea. Data preparation should produce a model-ready dataset without hiding uncertainty or altering biological meaning. Preserve raw data, check definitions and units, handle missing and extreme values explicitly, align observations with model outputs, and make every processing step reproducible.