Data preparation
Data preparation converts raw biological observations into a form suitable for mathematical and statistical analysis while preserving their scientific meaning. It should make the data easier to analyse without silently changing what the observations represent.
Keep the raw data unchanged
The original dataset should normally be preserved as a read-only source. Cleaning and transformation should produce a separate analysis dataset.
This allows every processed value to be traced back to the original observation and prevents accidental loss of information.
Understand the data dictionary
Before changing values, identify what each column means, its units, coding scheme, observational unit and allowed range.
For example, a column called cases might mean new cases per day, cumulative cases, laboratory-confirmed cases or estimated prevalence. These quantities cannot be used interchangeably.
Use a consistent table structure
A useful general structure has one observational unit per row and one variable per column, with identifiers linking repeated measurements to the same biological unit.
Longitudinal data may contain columns such as individual identifier, time and measured response.
Check data types
Numbers stored as text, dates stored in inconsistent formats and categorical labels entered with different spellings can cause analysis errors.
Each variable should be represented according to its meaning: numerical, categorical, binary, date/time or identifier.
Check units
Measurements must use compatible units before comparison with a mathematical model.
If a model parameter is expressed per day but observations are recorded weekly, the time scale must be reconciled explicitly.
Likewise, concentrations measured in different units must be converted before combining datasets.
Unit conversion
If a quantity \(x\) is converted by a known factor \(c\),
\[\boxed{x_{\mathrm{new}}=c\,x_{\mathrm{old}}}.\]The same conversion must be reflected consistently in model parameters and labels.
Check time variables
Dates should be ordered and converted to a time scale appropriate for the model. If \(t_0\) is the chosen origin, elapsed time can be defined by
\[t_i=\text{date}_i-\text{date}_0.\]This may produce time measured in days, hours or another consistent unit.
Irregular observation times
Measurements need not occur at equally spaced times:
\[t_{i+1}-t_i\ne\text{constant}.\]Irregular sampling is not automatically a problem. Many continuous-time models can be evaluated directly at the actual observation times.
Artificially interpolating data onto a regular grid should therefore be done only when the method genuinely requires it.
Duplicate records
Two identical-looking rows may be accidental duplicates or genuine repeated measurements.
Duplicates should not be deleted merely because values match. Their identifiers, timestamps and study design must first be checked.
Impossible values
Some observations violate known constraints. Examples include negative population counts, proportions outside \([0,1]\), impossible dates or concentrations outside an instrument's physical range.
Such values should be investigated rather than automatically replaced.
Range checks
Known biological constraints can be encoded as validation rules. For a proportion \(p_i\), for example,
\[0\le p_i\le1.\]For compartment counts in a closed population, one may check
\[S_i+E_i+I_i+R_i=N\]when that conservation law genuinely applies to the recorded quantities.
Missing values
A missing value means that an observation is unavailable. It should be represented explicitly rather than confused with zero.
Zero can be a meaningful biological measurement; missing means that the value is unknown.
Why values are missing
The reason for missingness matters. Missing measurements caused by random equipment failure differ from measurements missing because severely ill participants left a study.
Deleting every incomplete row can introduce bias when missingness is related to the biological outcome.
Imputation
Imputation replaces missing values using an explicit statistical procedure.
Simple replacements such as the sample mean can reduce apparent variability and distort relationships between variables. More principled methods should reflect the data-generating process and uncertainty.
Outliers
An outlier is an observation unusually far from the rest of the data. It may represent measurement error, data-entry error or genuine biological variation.
An extreme value should not be removed solely because it makes the fitted model worse.
Investigating outliers
Useful checks include the original record, measurement units, instrument range, neighbouring time points, replicate measurements and known biological events.
If an observation is excluded, the criterion should ideally be defined independently of the desired model result.
Transformations
A transformation changes the mathematical scale of a variable. For positive measurements, a logarithmic transformation is
\[\boxed{z_i=\log y_i}.\]This can convert multiplicative variation into approximately additive variation and reduce strong right skew.
It also changes the residual model and interpretation of fitting errors.
Log transformation and zeros
Because
\[\log 0\]is undefined, zeros require careful treatment. Adding an arbitrary constant such as \(1\) changes the data and may have substantial effects when values are small.
The choice should be justified by the measurement process rather than used automatically.
Scaling
A variable can be rescaled by a characteristic value \(a\):
\[z_i=\frac{y_i}{a}.\]Scaling can improve numerical optimisation when variables have very different magnitudes.
It should not be confused with changing the biological information in the observations.
Standardisation
A common statistical transformation is
\[\boxed{z_i=\frac{y_i-\bar y}{s}},\]where \(\bar y\) is the sample mean and \(s\) is the sample standard deviation.
This gives a dimensionless variable centred near zero with sample standard deviation one.
It is useful in some statistical models but often unnecessary when mechanistic parameters have meaningful physical units.
Normalisation
The word normalisation is used for several different operations, such as dividing by a population size, total count or reference measurement.
The exact operation should therefore be stated rather than simply saying that data were normalised.
Counts and proportions
If a count \(I(t)\) is converted to a population proportion,
\[i(t)=\frac{I(t)}{N}.\]A model fitted to \(i(t)\) must use the corresponding proportion scale.
Converting counts to proportions can be useful for comparing populations of different sizes, but it discards absolute scale unless \(N\) is retained.
Aggregation
Daily observations may be aggregated into weekly values. For incidence counts, one possible aggregation is
\[Y_w=\sum_{d\in w}Y_d.\]For a concentration measurement, summing values would usually have no meaningful interpretation; a mean or another summary might instead be appropriate.
The aggregation rule must match the biological quantity.
Aggregation can hide dynamics
Combining observations reduces temporal or spatial resolution. Rapid peaks and short-lived changes may disappear.
Aggregation can reduce noise but also remove information needed to estimate dynamic parameters.
Cumulative data
If incidence is \(y_t\), cumulative observations are
\[C_t=\sum_{s\le t}y_s.\]Successive cumulative values share most of the same underlying observations and are therefore strongly dependent.
Fitting cumulative data as though each point were an independent observation can give misleading uncertainty estimates.
Smoothing
Moving averages and other smoothers can reveal broad trends. A simple \(m\)-point moving average has the form
\[\bar y_t=\frac1m\sum_{j=0}^{m-1}y_{t-j}.\]However, smoothing changes the observation process and induces dependence between neighbouring smoothed values.
Raw data should normally remain available for formal inference.
Align model output with observation times
A numerical model may be solved at many internal time points, but residuals should be constructed at the actual observation times.
If the model predicts \(x(t;\theta)\), then
\[r_i=y_i-x(t_i;\theta).\]This is preferable to treating unobserved numerical solver points as additional data.
Match incidence intervals correctly
If the model produces a cumulative state \(C(t)\) but observations are interval incidence, the corresponding model prediction over \([t_{i-1},t_i]\) is
\[\boxed{\Delta C_i=C(t_i)-C(t_{i-1})}.\]This should be compared with interval counts rather than directly comparing \(C(t_i)\) with incidence.
Train and validation data
When predictive assessment is required, some observations can be withheld from parameter fitting and used later for validation.
For time-series models, random shuffling may leak future information into training. A chronological split is often more meaningful.
Avoid data leakage
Any transformation estimated from the data, such as a mean, standard deviation or feature-selection rule, should be calculated using only the appropriate fitting data when independent validation is intended.
Otherwise information from the validation set can influence model construction.
Record exclusions
Every excluded observation should have a traceable reason, such as a predefined quality-control failure or documented measurement problem.
Keeping an exclusion log makes the analysis auditable.
Reproducible preparation
Preparation steps should be encoded in scripts or notebooks whenever possible rather than performed manually in a spreadsheet without a record.
A reproducible workflow can be rerun when new data arrive or when a correction is made upstream.
Separate stages of the workflow
A useful structure is
\[\boxed{\text{raw data}\longrightarrow\text{clean data}\longrightarrow\text{analysis data}\longrightarrow\text{model fitting}}.\]Separating these stages makes it easier to identify where a change entered the analysis.
Visual inspection
Plots of observations against time, distributions, scatter plots and missingness patterns often reveal problems that summary tables do not.
Visual inspection complements formal validation checks rather than replacing them.
Document every transformation
For each transformation, record the original variable, mathematical operation, units before and after, and reason for the change.
This preserves biological interpretability when fitted parameters are reported later.
Transition to parameter estimation
Once observations have been checked, transformed appropriately and aligned with the model, the next task is to determine which parameter values make the model consistent with those observations.