1PhD candidate at Program in Translation Specialistic Medicine “G.B. Morgagni”, Curriculum “Biostatistics and Clinic Epidemiology” – University of Padua. Italy;
2Dr. José Matías Delgado’s University, El Salvador;
3,4,5,6National Health Institute, El Salvador. El Salvador.
Abstract
Objective: To systematically map and synthesize the scientific evidence on the quantitative methodologies used for the identification, cleaning and redistribution of “garbage codes” in population-based mortality registries in low- and middle-income countries (LMICs), evaluating their technical applicability for the parameterization of microsimulation models in health.
Inclusion criteria: This review will consider studies utilizing population-based mortality records or vital statistics from low- and middle-income countries (Population). Eligible literature must describe or apply traditional statistical methods, multiple imputation, Bayesian models, record linkage, or machine learning algorithms to correct ill-defined causes of death and “garbage codes” (Concept). The context includes health information systems characterized by scarce or heterogeneous data, specifically where the feasibility of disaggregating results to the individual level to parameterize stochastic microsimulation models is analyzed (Context).
Methods: This scoping review will follow the JBI methodology and the PRISMA Extension for Scoping Reviews (PRISMA-ScR). Comprehensive searches will be performed across PubMed, Scopus, Web of Science, LILACS, and SciELO, supplemented by grey literature from the WHO, PAHO, and IHME. To ensure reproducibility, the workflow will be executed within the R statistical environment. We will utilize litsearchr for search syntax optimization, synthesisr for deduplication, and revtools for blinded screening supported by latent Dirichlet allocation topic clustering. Discrepancies will be resolved by discussion or a third reviewer. Following paired data extraction, findings will be synthesized narratively and visualized using ggplot2 matrix heatmaps and igraph flow networks.
Keywords: health information systems; causes of death; vital statistics; data quality; machine learning algorithms; proportional redistribution; computer simulation.
Mortality statistics derived from Civil Registration and Vital Statistics systems serve as the bedrock of global epidemiological surveillance and public health planning1,2. However, progress remains sluggish, and the utility of these databases is frequently undermined by medical certification errors3,4. The primary challenge is the assignment of deaths to “garbage codes”, defined as International Classification of Diseases codes that represent ambiguous, non-specific, or biologically implausible underlying causes5–7. Because raw mortality data frequently lack diagnostic precision, the application of indirect statistical correction remains a crucial public health priority, particularly given the escalating toll of non-communicable diseases and cancers in data-scarce regions8–11. Furthermore, because clinical interventions and medical training alone cannot fully eliminate these diagnostic errors at the source, post-hoc redistribution is absolutely necessary12,13.
To mitigate these systematic biases, a diverse array of correction methodologies has evolved over the past two decades14–18. Modern redistribution strategies are typically classified into four distinct mathematical approaches: multiple cause analysis, negative correlation, impairment reallocation, and proportional redistribution19. Beyond these foundational methods, researchers have increasingly utilized individual-level multiple cause of death data to capture hidden etiologies20–22. This led to the development of advanced computational techniques, including data linkage algorithms and coarsened exact matching23–25. Recent innovations also feature subnational regression models26–28, four-step probabilistic algorithms29, the application of pathophysiological redistribution packages30, and the integration of international rules with exact matching31. Moreover, cutting edge spatial and machine learning approaches have emerged, such as using spatial scan statistics and Local Moran’s I autocorrelation to adjust for spurious geographical clusters32,33, and deploying random forest algorithms or decision trees to map spatial inconsistencies and hierarchical clinical interactions34,35.
Before applying these algorithms, baseline data quality is now objectively assessed using standardized metrics like the Vital Statistics Performance Index and automated diagnostic tools like ANACONDA36,37. However, despite this proliferation of techniques, current literature remains fragmented38. Specifically, no evidence synthesis exists on how to evaluate the transferability of these macro-level assumptions to individual level microsimulation models39,40. Unlike aggregated demographic models that rely on macroscopic trends, microsimulations require highly disaggregated baseline mortality risks, and uncounted or misclassified deaths are rarely randomly distributed41. Consequently, this scoping review will map the current state of the art, evaluate the technical feasibility of these approaches within simulation environments, and provide a framework for future public health mathematical modeling42.
This component defines eligible data records, populations, and geographic or demographic coverages.
This component delineates the core methodologies, algorithms, and analytical tools evaluated in this review.
This component defines the data environments and final fields of application toward which the evidence synthesis is oriented.
The search strategy will locate published academic literature and online grey literature. The search will focus on evidence published from 1996 onwards, corresponding to the global implementation and widespread adoption of the International Classification of Diseases, Tenth Revision (ICD-10)—a timeframe during which the term “garbage codes” became conceptually consolidated.
The first 100 results of each structured search will be screened. Identified organizational websites will be manually evaluated for technical reports. Finally, vital registration epidemiologists, public health officials, and key collaborators with expertise in health information systems within LMICs will be consulted to identify any additional missing sources.
To ensure the methodological reproducibility and transparency
required by the PRISMA-ScR guidelines, the information retrieval process
will be systematized within the R statistical environment (v4.5 or
higher). The litsearchr package will be employed for the
algorithmic optimization of Boolean syntax via term co-occurrence
networks. Indexed databases will be interrogated programmatically
through Application APIs using the rentrez and
rscopus libraries. Data consolidation and duplicate removal
will be managed via the synthesisr package using fuzzy
string-matching models, supplemented by manual verification of
borderline cases. Finally, title and abstract screening will be
conducted through the interactive graphical interface of the
revtools package, supported by Latent Dirichlet Allocation
(LDA) algorithms for topic modeling to optimize evidence selection in
accordance with the PCC framework.
To ensure integrity, traceability, and reproducibility, we will conduct data extraction programmatically within the R statistical environment (v4.5 or higher). This automated approach mitigates the transcription errors inherent in manual processes and ensures a standardized workflow. We have designed an electronic data charting form to capture information across six dimensions:
Study metadata: Author, year of publication, country, and study design.
Geographic and health system context: Challenges encountered during implementation, such as inadequate medical training or certification capacity, and recommendations for systemic improvement12.
Mortality database characteristics: Specific baseline registry quality metrics used to evaluate the initial validity of the input data. A prime example is the five-star rating system, which objectively scores mortality data based on demographic completeness and the prevalence of major garbage codes19, as well as the algorithmic assessment of missing or unexpected values in core demographic fields, which serves as a critical proxy for systemic data quality issues in highly deprived municipalities34.
“Garbage code” typology: The specific definition of a garbage code applied by the authors10, alongside the age- and sex-specific epidemiological profiles of these codes, to understand the clinical nature of the missing data at different life stages19.
Correction methodologies: The type of redistribution model used, the algorithmic and software ecosystems utilized, and the specific performance metrics reported to demonstrate data quality improvement39.
Microsimulation applicability: The explicit capacity of the methodology to generate highly disaggregated microdata and handle parameter uncertainty for stochastic predictive modeling40.
This form will be piloted with a random sample of three included studies to refine consistency and ensure variable granularity. Two reviewers will perform the extraction independently, resolving discrepancies through consensus. The resulting data will be consolidated into a relational structure (e.g., tibble or data.frame) to facilitate transparent downstream analysis. The full extraction schema (exported as JSON/CSV) and corresponding R scripts will be made available via the Open Science Framework (OSF) upon study completion. For a detailed breakdown of the variables and the classification framework utilized, refer to the Data Charting Instrument (Appendix II).
This scoping review adheres to the Joanna Briggs Institute methodology. We will conduct all data processing and visualization programmatically within the R statistical environment.
Descriptive Quantitative Analysis: We will calculate absolute frequencies and percentages to characterize the literature by chronological distribution, geographic reach, taxonomic frameworks, cause categories, and analytical techniques. These analytical techniques will be categorized by their computational complexity, ranging from foundational deterministic methods to advanced predictive modeling, machine learning, and spatial analysis.
Thematic Narrative Synthesis: A narrative will accompany the extracted data, addressing the magnitude of garbage codes in data-scarce settings, the methodological diversity of the algorithms, and the conceptual transferability of these methods to preserve microdata validity for agent-based stochastic models. This includes addressing how uncorrected garbage codes disproportionately mask the true mortality burden in populations with higher social deprivation43 and the risk of inappropriate centralization when macro-level regression algorithms are applied to highly localized municipal datasets44.
Mapping and Visualization: To visualize the structure and knowledge gaps of the field, we will generate advanced graphics using ggplot245 and igraph46. This will include structured evidence tables cross-referencing correction methods with their input data requirements, two-dimensional heatmaps mapping analytical methods against specific ICD Chapters, and flow network diagrams illustrating the technical pathway required to connect traditional vital statistics with microsimulation parameterization.
The authors declare no competing interests. This scoping review was conducted for academic and methodological purposes only, with no commercial, financial, or institutional support that could be construed as a potential conflict of interest.
Database: MEDLINE (via Ovid)
Search conducted on: July 2026
Planned Limits: * Date restrictions:
from 1996 to present to capture historical development of algorithms
like GBD. * Language restrictions: None at search level
(any required filters based on reviewer capacities will be applied
during the programmatic screening stage). * Document
types: All source types (including technical notes, methodology
papers, and electronic articles).
| Line | Search Query (Terms, Truncations, and Syntax) | Conceptual Block / Target |
|---|---|---|
| #1 | exp Mortality/ or exp Cause of Death/ | PCC: Population / Focus (Core mortality registry filters) |
| #2 | (mortality data* or death certificate* or vital statistic* or vital registration or CRVS).tw,kf,ot. | Keywords for vital statistics and national death registry systems |
| #3 | 1 or 2 | Result: Core Mortality Data |
| #4 | (garbage code* or ill-defined or misclassified or misclassification* or undefined cause* or ill-defined cause* or unspecific cause* or vague code* or redistribution algorithm*).tw,kf,ot. | PCC: Concept (Block A) - Specific typology of Data Errors (“Garbage Codes”) |
| #5 | (data cleaning or data correction or diagnostic accuracy or underlying cause of death).tw,kf,ot. | Data quality attributes specific to mortality databases |
| #6 | 4 or 5 | Result: Garbage Code Concepts |
| #7 | exp Algorithms/ or exp Models, Statistical/ or exp Computer Simulation/ | MeSH Indexing for Analytical Methods |
| #8 | (demographic redistribution or multiple imputation or MICE or Bayesian model* or Markov chain* or machine learning or random forest* or neural network* or artificial intelligence or algorithmic correction or fractional assignment).tw,kf,ot. | PCC: Concept (Block B) - Specific statistical and computational correction frameworks |
| #9 | (microsimulation* or micro-simulation* or agent-based model* or stochastic simulation* or individual-level dynamic* or synthetic population*).tw,kf,ot. | PCC: Context / Transferability - Target micro-level application architectures |
| #10 | 7 or 8 or 9 | Result: Methodological Ecosystem |
| #11 | exp “Developing Countries”/ | MeSH Indexing for LMICs |
| #12 | (low income country or low-and-middle income countr* or LMIC or LMICs or developing nation* or resource-constrained setting* or transitional econom*).tw,kf,ot. | Geographic Context Keywords (Based on World Bank specifications) |
| #13 | (Africa* or Asia* or South America* or Central America* or Latin America*).tw,kf,ot. | Broad regional geographic filters |
| #14 | 11 or 12 or 13 | Result: Developing Context (LMIC) |
| #15 | 3 and 6 and 10 and 14 | COMBINED STRATEGY (Boolean Intersection) |
| Block | Variable | Description / Purpose | Suggested Format / Values |
|---|---|---|---|
| I. Identification & Metadata | Study ID | Unique identifier assigned by R workflow | Numeric (e.g., 001) |
| Bibliographic Data | Author, year, title, and journal/institution | Text | |
| Evidence Type | Original article, technical report, or method book | Dropdown | |
| II. Geographic & Health Context | Geographic Scope | Country, region, or sub-national area | Text |
| Economic Classification | World Bank income level | Dropdown (Low, Lower-Middle, Upper-Middle) | |
| CRVS System Status | Description of Civil Registration and Vital Statistics status | Text | |
| III. Mortality Database | Data Periodicity | Years covered by mortality records | Text / Numeric |
| Data Volume | Sample size (number of deaths analyzed) | Numeric | |
| Primary Source | Origin of data (e.g., forensic, hospital, surveys) | Text | |
| ICD Framework | Version of the ICD utilized (ICD-9, ICD-10, ICD-11) | Dropdown | |
| IV. Junk Code Typology | Definition Framework | Criteria used to identify “junk” (e.g., GBD list) | Text |
| Evaluated ICD Codes | Specific codes being targeted for correction | Text | |
| Baseline Magnitude | Percentage/volume of junk codes in original records | Percentage (%) | |
| V. Correction Methodology | Central Technique | Algorithmic classification | Dropdown (e.g., MICE, Bayesian, ML) |
| Theoretical Assumptions | Underlying mathematical principles/assumptions | Text | |
| Software Ecosystem | Tools used (R, SAS, Stata, etc.) | Text | |
| Uncertainty/Bias Report | Reported confidence intervals or limitations | Yes / No / Partial | |
| VI. Microsimulation Applicability | TRL (Technology Readiness Level) | Maturity of the development method (1-4) | Scale (1-4) |
| Output Format | Aggregate vs. individual-level data output | Dropdown | |
| Demographic Granularity | Consistency at sub-group/small-area levels | Scale (1-5) | |
| Parametrization Potential | Capability to feed agent-based/microsimulation models | Scale (1-5) | |
| Reviewer Notes | Professional notes on applicability and transferability | Free text |