Overview

This report documents the cleaning and exploratory review of the DOLORisk phenotype table (GWASPheno), the clinical dataset underpinning the m.C2639T mitochondrial-variant analysis. This is first step of the pipeline: cleaning, quality control and exploratory analysis. The downstream carrier vs non-carrier comparison and the nociceptor-vs-non-nociceptor contrast build on the dataset produced here.

Data cleaning

Six data-quality issues were identified and corrected in a single pass. The raw table held 2,741 records; after de-duplication 2,740 remain.

Cleaning operations applied
Issue Detail Action
Inconsistent ID casing Sample IDs mixed DOL-/dol- prefixes Upper-cased & trimmed
Duplicate record 1 conflicting duplicate ID (PROPENG047) Kept first occurrence
Numeric columns stored as text 1992 cells with 'No bloods', '#N/A' etc. Coerced to numeric → NA
Mixed-type diabetes field Values 1 / 2 / 'LADA' / blank Factor: T1DM / T2DM / LADA
Medication flags 9 columns coded 1 / blank Recoded to 0 / 1 indicators
Untyped variables Sex, batch, grades read as numbers Cast to labelled factors

Cohort at a glance

Pain burden is summarised by the Brief Pain Inventory (BPI) average score (0 = no pain, 10 = worst), the cohort’s core pain-severity measure.

Pain phenotype: nociceptor vs non-nociceptor (QST)

QST-based nociceptor axis. The nociceptor / non-nociceptor split is taken directly from quantitative sensory testing: classifies each patient as IN :irritable nociceptor (sensitised, preserved small-fibre function) or NIN : non-irritable nociceptor (sensory loss / deafferentation).

The DN4 screening score sits slightly higher in the NIN group, consistent with NIN capturing more of the sensory-loss / deafferentation end of the spectrum.

Missing data

Missingness is substantial and highly variable — from near-complete demographics to pain-diary fields that are >90% empty. Variables are triaged into four tiers that drive how each is handled downstream.

Handling strategy by missingness
Missingness tier Variables
<5% — near-complete 24
5–40% — impute 5
40–80% — describe only 19
≥80% missing — drop 5
Free text (not modelled) 5

Strategy. Free-text fields are retained for reference but excluded from modelling. Variables ≥80% missing are dropped from the analysis set. Fields missing 5–40% (key clinical measures such as HbA1c, BMI, sensory scores) are candidates for multiple imputation; those 40–80% missing are used descriptively only. The existing carrier comparison uses pairwise deletion, so the cleaned dataset is exported with missing values preserved.

Next steps

  • Step 2: Carrier comparison: carrier vs non-carrier phenotype analysis