1PhD candidate at Program in Translation Specialistic Medicine “G.B. Morgagni”, Curriculum “Biostatistics and Clinic Epidemiology” – University of Padua. Italy;

2Dr. José Matías Delgado’s University, El Salvador;

3,4,5,6National Health Institute, El Salvador. El Salvador.

Abstract

Objective: To systematically map and synthesize the scientific evidence on the quantitative methodologies used for the identification, cleaning and redistribution of “garbage codes” in population-based mortality registries in low- and middle-income countries (LMICs), evaluating their technical applicability for the parameterization of microsimulation models in health.

Inclusion criteria: This review will consider studies utilizing population-based mortality records or vital statistics from low- and middle-income countries (Population). Eligible literature must describe or apply traditional statistical methods, multiple imputation, Bayesian models, record linkage, or machine learning algorithms to correct ill-defined causes of death and “garbage codes” (Concept). The context includes health information systems characterized by scarce or heterogeneous data, specifically where the feasibility of disaggregating results to the individual level to parameterize stochastic microsimulation models is analyzed (Context).

Methods: This scoping review will follow the JBI methodology and the PRISMA Extension for Scoping Reviews (PRISMA-ScR). Comprehensive searches will be performed across PubMed, Scopus, Web of Science, LILACS, and SciELO, supplemented by grey literature from the WHO, PAHO, and IHME. To ensure reproducibility, the workflow will be executed within the R statistical environment. We will utilize litsearchr for search syntax optimization, synthesisr for deduplication, and revtools for blinded screening supported by latent Dirichlet allocation topic clustering. Discrepancies will be resolved by discussion or a third reviewer. Following paired data extraction, findings will be synthesized narratively and visualized using ggplot2 matrix heatmaps and igraph flow networks.

Keywords: health information systems; causes of death; vital statistics; data quality; machine learning algorithms; proportional redistribution; computer simulation.


INTRODUCTION

Mortality statistics derived from Civil Registration and Vital Statistics systems serve as the bedrock of global epidemiological surveillance and public health planning1,2. However, progress remains sluggish, and the utility of these databases is frequently undermined by medical certification errors3,4. The primary challenge is the assignment of deaths to “garbage codes”, defined as International Classification of Diseases codes that represent ambiguous, non-specific, or biologically implausible underlying causes57. Because raw mortality data frequently lack diagnostic precision, the application of indirect statistical correction remains a crucial public health priority, particularly given the escalating toll of non-communicable diseases and cancers in data-scarce regions811. Furthermore, because clinical interventions and medical training alone cannot fully eliminate these diagnostic errors at the source, post-hoc redistribution is absolutely necessary12,13.

To mitigate these systematic biases, a diverse array of correction methodologies has evolved over the past two decades1418. Modern redistribution strategies are typically classified into four distinct mathematical approaches: multiple cause analysis, negative correlation, impairment reallocation, and proportional redistribution19. Beyond these foundational methods, researchers have increasingly utilized individual-level multiple cause of death data to capture hidden etiologies2022. This led to the development of advanced computational techniques, including data linkage algorithms and coarsened exact matching2325. Recent innovations also feature subnational regression models2628, four-step probabilistic algorithms29, the application of pathophysiological redistribution packages30, and the integration of international rules with exact matching31. Moreover, cutting edge spatial and machine learning approaches have emerged, such as using spatial scan statistics and Local Moran’s I autocorrelation to adjust for spurious geographical clusters32,33, and deploying random forest algorithms or decision trees to map spatial inconsistencies and hierarchical clinical interactions34,35.

Before applying these algorithms, baseline data quality is now objectively assessed using standardized metrics like the Vital Statistics Performance Index and automated diagnostic tools like ANACONDA36,37. However, despite this proliferation of techniques, current literature remains fragmented38. Specifically, no evidence synthesis exists on how to evaluate the transferability of these macro-level assumptions to individual level microsimulation models39,40. Unlike aggregated demographic models that rely on macroscopic trends, microsimulations require highly disaggregated baseline mortality risks, and uncounted or misclassified deaths are rarely randomly distributed41. Consequently, this scoping review will map the current state of the art, evaluate the technical feasibility of these approaches within simulation environments, and provide a framework for future public health mathematical modeling42.

Research Questions

  1. What analytical methodologies (statistical, probabilistic, deterministic, linkage-based or machine learning) have been described in the literature for the correction and redistribution of “garbage codes” in mortality records?
  2. How are they classified and what types of “garbage codes” (poorly defined causes of ICD chapters, intermediate or terminal causes) are the main objective of each identified methodology?
  3. What is the documented level of applicability or compatibility of these correction methods to structure databases aimed at the development, calibration and validation of microsimulation models in health?
  4. What are the main methodological challenges, limitations and advantages reported when transferring or adapting analytical correction algorithms from the population (macro) level to the individual level (micro)?

ELIGIBILITY CRITERIA

Population

This component defines eligible data records, populations, and geographic or demographic coverages.

  • Inclusion Criteria:
    • Data Type: Studies utilizing population-based mortality databases, vital statistics systems, official national or subnational health statistics registers, or civil registration and vital statistics (CRVS) systems.
    • Geography and Income Level: Research focusing on populations within low- and middle-income countries (LMICs) as defined by historical or current World Bank classifications (e.g., regions within Latin America and the Caribbean, Sub-Saharan Africa, and South Asia).
    • Technical Exception: Global or multinational studies (e.g., Global Burden of Disease [GBD] analyses), provided they independently disaggregate methodologies and data applicable to LMIC geographies.
    • Demography: Records tracking general human mortality without restrictions on age groups (including neonatal, infant, adult, or geriatric mortality) or sex and gender.
  • Exclusion Criteria:
    • Clinical studies restricted to hyper-specific hospital specimens or controlled trials that do not utilize—or aim to correct—population-based records or broader epidemiological information systems.

Concept

This component delineates the core methodologies, algorithms, and analytical tools evaluated in this review.

  • Inclusion Criteria:
    • Studies describing, evaluating, validating, or implementing quantitative methods to identify, clean, reclassify, or redistribute “garbage codes” (levels 1 to 4) or ill-defined causes of death (as classified under ICD-10/ICD-11 Chapter XVIII, or ICD-9 equivalents).
    • Methodologies within the following analytical categories are eligible:
      • Traditional statistics and mathematical demography: Proportional redistribution (simple or stratified by age and sex), fractional algorithms, and indirect demographic adjustment methods14.
      • Advanced statistical modeling: Multiple imputation of missing or non-specific data, spatial or spatiotemporal regression modeling, and empirical or hierarchical Bayesian approaches (e.g., CODEm or CoDcorrect analytical algorithms)18,19,26.
      • Data integration (record linkage): Deterministic or probabilistic algorithms linking mortality records with complementary databases (e.g., hospital discharge records, automated verbal autopsies, or epidemiological surveillance systems)20,24.
      • Artificial intelligence and machine learning: Data-driven predictive architectures (e.g., decision trees, gradient boosting, random forests, or other data-mining frameworks) applied to classify or predict underlying causes of death masked by non-specific codes.
  • Exclusion Criteria:
    • Studies analyzing mortality data quality in a purely descriptive manner (e.g., reporting isolated completeness percentages or garbage code prevalences) without implementing or evaluating an active mathematical or algorithmic correction method.
    • Research focused solely on correcting numerical under-registration (omission of death capture), unless the methodology explicitly integrates cause-of-death redistribution.

Context

This component defines the data environments and final fields of application toward which the evidence synthesis is oriented.

  • Inclusion Criteria:
    • Data-Scarce or Heterogeneous Settings: Research situated in contexts with weak or fragmented health information systems characterized by low medical certification coverage, high regional heterogeneity, or small-area spatiotemporal misalignments.
    • Microsimulation Applicability: Studies evaluating or discussing the impact of correction techniques on individual-level microdata disaggregation, the generation of consistent mortality risk matrices, individual trajectory estimations, or parameterizing inputs for health microsimulation platforms.
    • Emergent Methodological Scope: Studies originally operating at a macro or aggregate level, provided they detail the mathematical assumptions or individual-record classification algorithms necessary to calibrate and parameterize stochastic microsimulations or agent-based models.
  • Exclusion Criteria:
    • Methodologies operating exclusively at an irreducible macro-population level whose designs preclude individual-level micro-disaggregation, thereby lacking utility for agent-based or individual-level simulations.

Sources of Evidence

  • Publication Types: Peer-reviewed original research articles published in indexed journals, methodological book chapters, and authoritative institutional gray literature (e.g., technical reports from the WHO, PAHO, or IHME).
  • Timeframe: Material published from 1996 onwards, capturing the global introduction and widespread adoption of the Tenth Revision of the International Classification of Diseases (ICD-10), which consolidated the term “garbage codes.”
  • Languages: Studies published in English, Spanish, and Portuguese, reflecting the substantial volume of methodological literature generated in Latin America (particularly Brazil).

METHODS

Search Strategy

The search strategy will locate published academic literature and online grey literature. The search will focus on evidence published from 1996 onwards, corresponding to the global implementation and widespread adoption of the International Classification of Diseases, Tenth Revision (ICD-10)—a timeframe during which the term “garbage codes” became conceptually consolidated.

  • Academic Literature: To identify academic literature, a three-step search strategy will be developed in collaboration with a medical librarian:
    1. Initial Limited Search: An initial search of MEDLINE (Ovid) will be undertaken to analyze text words in titles and abstracts, Medical Subject Headings (MeSH), and author-supplied keywords from relevant articles.
    2. Full Search Expansion: A comprehensive search strategy will be developed using these identified terms, then translated and adapted for Scopus, Web of Science, LILACS, and SciELO. The complete search syntax for MEDLINE (Ovid) is presented in Appendix I.
    3. Reference Hand-Searching: The reference lists of all selected articles and relevant systematic or scoping reviews will be hand-searched to identify primary sources missed by the electronic search.
  • Grey Literature: To locate technical reports and eligible grey literature, a targeted list of national and international health organization websites will be compiled following the two-step Google search methodology.
    1. Preliminary Term Identification: A preliminary Google search will be conducted to identify key text words, accounting for highly variable terminology such as “unspecific causes” or “ambiguous causes.”
    2. Structured Execution: Ten unique structured Google searches will be executed using distinct keyword combinations representing the population, concept, and context (PCC) of interest.

The first 100 results of each structured search will be screened. Identified organizational websites will be manually evaluated for technical reports. Finally, vital registration epidemiologists, public health officials, and key collaborators with expertise in health information systems within LMICs will be consulted to identify any additional missing sources.

Data Management and Search Automation

To ensure the methodological reproducibility and transparency required by the PRISMA-ScR guidelines, the information retrieval process will be systematized within the R statistical environment (v4.5 or higher). The litsearchr package will be employed for the algorithmic optimization of Boolean syntax via term co-occurrence networks. Indexed databases will be interrogated programmatically through Application APIs using the rentrez and rscopus libraries. Data consolidation and duplicate removal will be managed via the synthesisr package using fuzzy string-matching models, supplemented by manual verification of borderline cases. Finally, title and abstract screening will be conducted through the interactive graphical interface of the revtools package, supported by Latent Dirichlet Allocation (LDA) algorithms for topic modeling to optimize evidence selection in accordance with the PCC framework.

Data Extraction and Charting

To ensure integrity, traceability, and reproducibility, we will conduct data extraction programmatically within the R statistical environment (v4.5 or higher). This automated approach mitigates the transcription errors inherent in manual processes and ensures a standardized workflow. We have designed an electronic data charting form to capture information across six dimensions:

  • Study metadata: Author, year of publication, country, and study design.

  • Geographic and health system context: Challenges encountered during implementation, such as inadequate medical training or certification capacity, and recommendations for systemic improvement12.

  • Mortality database characteristics: Specific baseline registry quality metrics used to evaluate the initial validity of the input data. A prime example is the five-star rating system, which objectively scores mortality data based on demographic completeness and the prevalence of major garbage codes19, as well as the algorithmic assessment of missing or unexpected values in core demographic fields, which serves as a critical proxy for systemic data quality issues in highly deprived municipalities34.

  • “Garbage code” typology: The specific definition of a garbage code applied by the authors10, alongside the age- and sex-specific epidemiological profiles of these codes, to understand the clinical nature of the missing data at different life stages19.

  • Correction methodologies: The type of redistribution model used, the algorithmic and software ecosystems utilized, and the specific performance metrics reported to demonstrate data quality improvement39.

  • Microsimulation applicability: The explicit capacity of the methodology to generate highly disaggregated microdata and handle parameter uncertainty for stochastic predictive modeling40.

This form will be piloted with a random sample of three included studies to refine consistency and ensure variable granularity. Two reviewers will perform the extraction independently, resolving discrepancies through consensus. The resulting data will be consolidated into a relational structure (e.g., tibble or data.frame) to facilitate transparent downstream analysis. The full extraction schema (exported as JSON/CSV) and corresponding R scripts will be made available via the Open Science Framework (OSF) upon study completion. For a detailed breakdown of the variables and the classification framework utilized, refer to the Data Charting Instrument (Appendix II).

Data Analysis, Synthesis, and Presentation

This scoping review adheres to the Joanna Briggs Institute methodology. We will conduct all data processing and visualization programmatically within the R statistical environment.

  • Descriptive Quantitative Analysis: We will calculate absolute frequencies and percentages to characterize the literature by chronological distribution, geographic reach, taxonomic frameworks, cause categories, and analytical techniques. These analytical techniques will be categorized by their computational complexity, ranging from foundational deterministic methods to advanced predictive modeling, machine learning, and spatial analysis.

  • Thematic Narrative Synthesis: A narrative will accompany the extracted data, addressing the magnitude of garbage codes in data-scarce settings, the methodological diversity of the algorithms, and the conceptual transferability of these methods to preserve microdata validity for agent-based stochastic models. This includes addressing how uncorrected garbage codes disproportionately mask the true mortality burden in populations with higher social deprivation43 and the risk of inappropriate centralization when macro-level regression algorithms are applied to highly localized municipal datasets44.

  • Mapping and Visualization: To visualize the structure and knowledge gaps of the field, we will generate advanced graphics using ggplot245 and igraph46. This will include structured evidence tables cross-referencing correction methods with their input data requirements, two-dimensional heatmaps mapping analytical methods against specific ICD Chapters, and flow network diagrams illustrating the technical pathway required to connect traditional vital statistics with microsimulation parameterization.

Conflicts of Interest

The authors declare no competing interests. This scoping review was conducted for academic and methodological purposes only, with no commercial, financial, or institutional support that could be construed as a potential conflict of interest.


References

1.
Setel P, Macfarlane S, Szreter S, Mikkelsen L, Jha P, Stout S, et al. A scandal of invisibility: Making everyone count by counting everyone. Lancet. 2007;370:1569–77.
2.
Mathers CD, Fat DM, Inoue M, Rao C, Lopez AD. Counting the dead and what they died from: An assessment of the global status of cause of death data. Bulletin of the World Health Organization. 2005;83:171–7.
3.
AbouZahr C, Savigny D de, Mikkelsen L, Setel PW, Lozano R, Nichols E, et al. Civil registration and vital statistics: Progress in the data revolution for counting and accountability. The Lancet. 2015;386(10001):1373–85. doi:10.1016/S0140-6736(15)60173-8
4.
Iburg KM, Mikkelsen L, Adair T, Lopez AD. Are cause of death data fit for purpose? Evidence from 20 countries at different levels of socio-economic development. PLoS ONE. 2020;15(8):e0237539. doi:10.1371/journal.pone.0237539
5.
Murray C, Lopez A. The global burden of disease: A comprehensive assessment of mortality and disability from diseases, injuries and risk factors in 1990 and projected to 2020. Boston: Harvard School of Public Health; 1996.
6.
Naghavi M, Richards N, Chowdhury H, et al. Improving the quality of cause of death data for public health policy: Are all “garbage” codes equally problematic? BMC Med. 2020;18(1):55.
7.
World Health Organization. International classification of diseases for mortality and morbidity statistics eleventh revision: Reference guide [Internet]. Geneva: World Health Organization; 2022. Available from: https://icd.who.int/browse11
8.
Ahern RM, Lozano R, Naghavi M, Foreman K, Gakidou E, Murray CJ. Improving the public health utility of global cardiovascular mortality data: The rise of ischemic heart disease. Population Health Metrics. 2011 Dec;9(1):8. doi:10.1186/1478-7954-9-8
9.
Salerno P, Cotton A, Chen Z, Nascimento B, Petermann-Rocha F, Deo S. Forecasting ischemic heart disease, stroke, and peripheral artery disease mortality in brazil through 2030. Arquivos Brasileiros de Cardiologia. 2025;122(1):e2024012.
10.
Bigoni A, Cunha AR da, Antunes JLF. Redistributing deaths by ill-defined and unspecified causes on cancer mortality in brazil. Revista de Saúde Pública. 2021;55:106. doi:10.11606/s1518-8787.2021055003319
11.
Cunha AR da, Bigoni A, Antunes JLF, Hugo FN. Impact of redistributing deaths by ill-defined causes in oral and oropharyngeal cancer mortality in brazil. Brazilian Oral Research. 2022;36:e117. doi:10.1590/1807-3107bor-2022.vol36.0117
12.
Hart J et al. Improving medical certification of cause of death: Effective strategies and approaches based on experiences from the data for health initiative. BMC Public Health. 2020;20(1):1083.
13.
Miki J et al. Saving lives through certifying deaths: Assessing the impact of two interventions to improve cause of death data in peru. BMC Public Health. 2018;18(1):1329.
14.
Naghavi M, Makela S, Foreman K, et al. Algorithms for enhancing public health utility of national causes-of-death data. Popul Health Metr. 2010;8(1):9.
15.
GBD 2016 Causes of Death Collaborators. Global, regional, and national age-sex specific mortality for 264 causes of death, 1980–2016: A systematic analysis for the global burden of disease study 2016. The Lancet. 2017;390(10100):1151–210.
16.
GBD 2017 Causes of Death Collaborators. Global, regional, and national age-sex-specific mortality for 282 causes of death in 195 countries and territories, 1980–2017: A systematic analysis for the global burden of disease study 2017. The Lancet. 2018;392(10159):1736–88. doi:10.1016/S0140-6736(18)32203-7
17.
Murray CJ, Dias RH, Kulkarni SC, Lozano R, Stevens GA, Ezzati M. Improving the comparability of diabetes mortality statistics in the u.s. And mexico. Diabetes Care. 2008;31(3):451–8.
18.
Foreman K, Naghavi M, Ezzati M. Improving the usefulness of US mortality data: New methods for reclassification of underlying cause of death. Popul Health Metr. 2016;14:14.
19.
Johnson S et al. Public health utility of cause of death data: Applying empirical algorithms to improve data quality. BMC Med Inform Decis Mak. 2021;21(1):175.
20.
França E et al. Ill-defined causes of death in brazil: A redistribution method based on the investigation of such causes. Rev Saúde Pública. 2014;48(4):671–81.
21.
França EB, Ishitani LH, Teixeira RA, Cunha CC da, Marinho MF. Improving the usefulness of mortality data: Reclassification of ill-defined causes based on medical records and home interviews in brazil. Revista Brasileira de Epidemiologia. 2019;22:e190010.
22.
França EB. Garbage codes assigned as cause-of-death in health statistics. Revista Brasileira de Epidemiologia. 2019;22:e190001.
23.
Stevens GA, King G, Shibuya K. Deaths from heart failure: Using coarsened exact matching to correct cause-of-death statistics. Population Health Metrics. 2010;8(1):6. doi:10.1186/1478-7954-8-6
24.
Bierrenbach A et al. Redistribution of heart failure deaths using two methods: Linkage of hospital records with death certificates. Rev Bras Epidemiol. 2019;22(Suppl 3):e190006.
25.
Fihel A, Muszyńska-Spielauer MM. Using multiple cause of death information to eliminate garbage codes. Demographic Research. 2021;45:345–60. doi:10.4054/DemRes.2021.45.11
26.
Grigoriev P, Bonnet F, Perdrix E. Method for redistributing ill-defined causes of death. Popul Health Metr. 2024;22(1):1–14.
27.
Giraldo L, Rodrı́guez L, González-Robledo MC, Mino-León D, Cahuana-Hurtado L, Rojas-Russell ME, et al. Modelo para el análisis de la mortalidad en colombia 2000-2012. Revista de Salud Pública. 2017;19(2):166–73.
28.
Masquelier B et al. Analysis of death registers in antananarivo, madagascar. BMC. 2019.
29.
Scohy A, Lesnik T, Devleesschauwer B, Haneef R. The importance of redistribution process of ill-defined deaths in national mortality database. Archives of Public Health. 2025 Jul 11;83(1):185. doi:10.1186/s13690-025-01652-x
30.
Kyu HH et al. Estimating the burden of HIV/AIDS-related mortality: HIV/AIDS-related garbage code redistribution packages. Journal of the International AIDS Society. 2021. doi:10.1002/jia2.25791
31.
Zhang X, Yang W, Wang J, Ai L, Chen M, Wang C, et al. Reallocating diabetes-related garbage codes to improve mortality estimates: A case study in weifang, china. Population Health Metrics. 2025;23(1):38. doi:10.1186/s12963-025-00399-5
32.
Teixeira B, Toporcov T, Chiaravalloti-Neto F, Chiavegatto Filho A. Spatial clusters of cancer mortality in brazil: A machine learning modelling approach. In: Anexo do curriculo lattes. 2020.
33.
Araújo Santos Camargo JD de, Camargo SF, Souza ATB de, Oliveira Freitas AKMS de, Ferreira CA, Sarmento ACA, et al. Regional disparities in breast cancer mortality in brazil: A spatial analysis using uncorrected and adjusted data, 2000–2023. Scientific Reports. 2026;16:6770. doi:10.1038/s41598-026-37844-w
34.
Zimeo-Morais GA, Miraglia JL, Oliveira BZ de, Mistro S, Hisatugu WH, Greffin D, et al. Factors associated with the quality of death certification in brazilian municipalities: A data-driven non-linear model. PLoS One. 2023;18(8):e0290814. doi:10.1371/journal.pone.0290814
35.
Souza ARM de, Raposo LM, Lacerda GCB de, Godoy PH. Stroke deaths profile and its subtypes in brazil: Analysis using machine learning. Global Heart. 2025;20(1):85. doi:10.5334/gh.1476
36.
Mikkelsen L, Phillips D, AbouZahr C, Setel P, Savigny D de, Lozano R, et al. A global assessment of civil registration and vital statistics systems: Monitoring data quality and progress. Lancet. 2015;386:1395–406.
37.
Mikkelsen L, Moesgaard K, Hegnauer M, Lopez AD. ANACONDA: A new tool to improve mortality and cause of death data. BMC Medicine. 2020;18(1):61. doi:10.1186/s12916-020-01521-0
38.
Devleesschauwer B. Skills building seminar: Redistribution of ill-defined deaths: A methodological conundrum. In: European journal of public health. Oxford University Press; 2022. p. ckac129–618. doi:10.1093/eurpub/ckac129.618
39.
Ng TC, Lo WC, Ku CC, Lu TH, Lin HH. Improving the use of mortality data in public health: A comparison of garbage code redistribution models. American Journal of Public Health. 2020;110(2):222–9.
40.
Chrysanthopoulou SA, Rutter CM, Gatsonis CA. Bayesian versus empirical calibration of microsimulation models: A comparative analysis. Medical Decision Making. 2021 Aug;41(6):714–26. doi:10.1177/0272989X211009161
41.
Costa LFL, Montenegro M de MS, Rabello Neto D de L, Oliveira ATR de, Trindade JE de O, Adair T, et al. Estimating completeness of national and subnational death reporting in brazil: Application of record linkage methods. Population Health Metrics. 2020;18(1):22.
42.
Foreman KJ, Marquez N, Dolgert A, Fukutaki K, Fullman N, McGaughey M, et al. Forecasting life expectancy, years of life lost, and all-cause and cause-specific mortality for 250 causes of death: Reference and alternative scenarios for 201640 for 195 countries and territories. The Lancet. 2018;392(10159):2052–90. doi:10.1016/S0140-6736(18)31694-5
43.
Malta DC, Teixeira RA, Cardoso LS de M, Souza JB de, Bernal RTI, Pinheiro PC, et al. Premature mortality due to noncommunicable diseases in brazilian capitals: Redistribution of garbage causes and evolution by social deprivation strata. Revista Brasileira de Epidemiologia. 2023;26:e230002.supl.1. doi:10.1590/1980-549720230002.supl.1
44.
Liu L, Wang X, Wang C, Ma X, Meng X, Ning B, et al. A study on garbage code redistribution methods in small area: Redistributing heart failure in two chinese cities by two approaches. Research Square (Preprint). 2022. doi:10.21203/rs.3.rs-1242825/v1
45.
Wickham H. ggplot2: Elegant graphics for data analysis. New York: Springer-Verlag; 2016.
46.
Csardi G, Nepusz T. The igraph software package for complex network research. InterJournal, Complex Systems. 2006;1695:1–9.

Appendices

Appendix I: Search strategy

Database: MEDLINE (via Ovid)
Search conducted on: July 2026
Planned Limits: * Date restrictions: from 1996 to present to capture historical development of algorithms like GBD. * Language restrictions: None at search level (any required filters based on reviewer capacities will be applied during the programmatic screening stage). * Document types: All source types (including technical notes, methodology papers, and electronic articles).

Line Search Query (Terms, Truncations, and Syntax) Conceptual Block / Target
#1 exp Mortality/ or exp Cause of Death/ PCC: Population / Focus (Core mortality registry filters)
#2 (mortality data* or death certificate* or vital statistic* or vital registration or CRVS).tw,kf,ot. Keywords for vital statistics and national death registry systems
#3 1 or 2 Result: Core Mortality Data
#4 (garbage code* or ill-defined or misclassified or misclassification* or undefined cause* or ill-defined cause* or unspecific cause* or vague code* or redistribution algorithm*).tw,kf,ot. PCC: Concept (Block A) - Specific typology of Data Errors (“Garbage Codes”)
#5 (data cleaning or data correction or diagnostic accuracy or underlying cause of death).tw,kf,ot. Data quality attributes specific to mortality databases
#6 4 or 5 Result: Garbage Code Concepts
#7 exp Algorithms/ or exp Models, Statistical/ or exp Computer Simulation/ MeSH Indexing for Analytical Methods
#8 (demographic redistribution or multiple imputation or MICE or Bayesian model* or Markov chain* or machine learning or random forest* or neural network* or artificial intelligence or algorithmic correction or fractional assignment).tw,kf,ot. PCC: Concept (Block B) - Specific statistical and computational correction frameworks
#9 (microsimulation* or micro-simulation* or agent-based model* or stochastic simulation* or individual-level dynamic* or synthetic population*).tw,kf,ot. PCC: Context / Transferability - Target micro-level application architectures
#10 7 or 8 or 9 Result: Methodological Ecosystem
#11 exp “Developing Countries”/ MeSH Indexing for LMICs
#12 (low income country or low-and-middle income countr* or LMIC or LMICs or developing nation* or resource-constrained setting* or transitional econom*).tw,kf,ot. Geographic Context Keywords (Based on World Bank specifications)
#13 (Africa* or Asia* or South America* or Central America* or Latin America*).tw,kf,ot. Broad regional geographic filters
#14 11 or 12 or 13 Result: Developing Context (LMIC)
#15 3 and 6 and 10 and 14 COMBINED STRATEGY (Boolean Intersection)

Appendix II: Data extraction instrument

Block Variable Description / Purpose Suggested Format / Values
I. Identification & Metadata Study ID Unique identifier assigned by R workflow Numeric (e.g., 001)
Bibliographic Data Author, year, title, and journal/institution Text
Evidence Type Original article, technical report, or method book Dropdown
II. Geographic & Health Context Geographic Scope Country, region, or sub-national area Text
Economic Classification World Bank income level Dropdown (Low, Lower-Middle, Upper-Middle)
CRVS System Status Description of Civil Registration and Vital Statistics status Text
III. Mortality Database Data Periodicity Years covered by mortality records Text / Numeric
Data Volume Sample size (number of deaths analyzed) Numeric
Primary Source Origin of data (e.g., forensic, hospital, surveys) Text
ICD Framework Version of the ICD utilized (ICD-9, ICD-10, ICD-11) Dropdown
IV. Junk Code Typology Definition Framework Criteria used to identify “junk” (e.g., GBD list) Text
Evaluated ICD Codes Specific codes being targeted for correction Text
Baseline Magnitude Percentage/volume of junk codes in original records Percentage (%)
V. Correction Methodology Central Technique Algorithmic classification Dropdown (e.g., MICE, Bayesian, ML)
Theoretical Assumptions Underlying mathematical principles/assumptions Text
Software Ecosystem Tools used (R, SAS, Stata, etc.) Text
Uncertainty/Bias Report Reported confidence intervals or limitations Yes / No / Partial
VI. Microsimulation Applicability TRL (Technology Readiness Level) Maturity of the development method (1-4) Scale (1-4)
Output Format Aggregate vs. individual-level data output Dropdown
Demographic Granularity Consistency at sub-group/small-area levels Scale (1-5)
Parametrization Potential Capability to feed agent-based/microsimulation models Scale (1-5)
Reviewer Notes Professional notes on applicability and transferability Free text