1G.B. Morgagni PhD program in Translational Medicine, University of Padua, Padua, Italy;

2Department of Public Health, School of Medicine, Faculty of Health Sciences, Dr. José Matías Delgado University, Antiguo Cuscatlán, El Salvador;

3National Health Institute, El Salvador. El Salvador.

Abstract

Objective: To systematically map and synthesize the scientific evidence on the quantitative methodologies used for the identification, cleaning and redistribution of “garbage codes” in population-based mortality registries in low- and middle-income countries (LMICs), evaluating their technical applicability for the parameterization of microsimulation models in health.

Inclusion criteria: This review will consider studies utilizing population-based mortality records or vital statistics from low- and middle-income countries (Population). Eligible literature must describe or apply traditional statistical methods, multiple imputation, Bayesian models, record linkage, or machine learning algorithms to correct ill-defined causes of death and “garbage codes” (Concept). The context includes health information systems characterized by scarce or heterogeneous data, specifically where the feasibility of disaggregating results to the individual level to parameterize stochastic microsimulation models is analyzed (Context).

Methods: This scoping review will follow the JBI methodology and the PRISMA Extension for Scoping Reviews (PRISMA-ScR). Comprehensive searches will be performed across PubMed, Scopus, Web of Science, LILACS, and SciELO, supplemented by grey literature from the WHO, PAHO, and IHME. To ensure reproducibility, the workflow will be executed within the R statistical environment. We will utilize litsearchr for search syntax optimization, synthesisr for deduplication, and revtools for blinded screening supported by latent Dirichlet allocation topic clustering. Discrepancies will be resolved by discussion or a third reviewer. Following paired data extraction, findings will be synthesized narratively and visualized using ggplot2 matrix heatmaps and igraph flow networks.

Keywords: health information systems; causes of death; vital statistics; data quality; machine learning algorithms; proportional redistribution; computer simulation.


INTRODUCTION

Civil registration and vital statistics systems serve as the cornerstone of demographic and medium- to long-term health intelligence1,2, underpinning policy planning, health technology assessment3–5, and the quantification of the true impact of health crises6,7. However, despite decades of global advocacy, data quality progress remains sluggish1,2,8, and the analytical utility of these databases is frequently undermined by persistent medical certification errors7,9–11. The primary challenge is the assignment of deaths to “garbage codes” (GCs), defined as International Classification of Diseases codes that represent ambiguous, non-specific, or biologically implausible underlying causes12–14. Because raw mortality data frequently lack diagnostic precision, the application of indirect statistical correction remains a crucial public health priority15, particularly given the escalating toll of non-communicable diseases16 and cancers in data-scarce regions17,18. Furthermore, because clinical interventions and medical training alone cannot fully eliminate these diagnostic errors at the source7,9, robust post-hoc redistribution methods are methodologically indispensable13,19.

To mitigate these systematic biases, a diverse array of correction methodologies has evolved over the past two decades6,19–22. Modern redistribution strategies are typically classified into four distinct mathematical approaches: multiple cause analysis, negative correlation, impairment reallocation, and proportional redistribution11. Beyond these foundational methods, researchers have increasingly utilized individual-level multiple cause of death data to capture hidden etiologies23–25. This led to the development of advanced computational techniques, including clinical data linkage algorithms26 and coarsened exact matching27,28.

Recent innovations also feature subnational regression models29–31, four-step probabilistic algorithms32, the application of pathophysiological redistribution packages33, and the integration of international rules with exact matching34. Moreover, cutting edge spatial and machine learning approaches have emerged, such as using spatial scan statistics and Local Moran’s I autocorrelation to adjust for spurious geographical clusters35,36, and deploying random forest algorithms or decision trees to map spatial inconsistencies and hierarchical clinical interactions37,38.

Before applying these algorithms, baseline data quality is now objectively assessed using standardized metrics like the Vital Statistics Performance Index8,39 and automated diagnostic tools like ANACONDA10,40. However, despite this methodological proliferation, current literature exhibits a critical translational gap41. Specifically, there is a lack of evidence evaluating how macro-level redistribution assumptions can be rigorously transferred to individual-level simulation environments42,43.

Unlike aggregated demographic models that rely on macroscopic trends, stochastic and agent-based microsimulations require highly disaggregated baseline mortality risks, as uncounted or misclassified deaths are rarely distributed randomly across populations44. Naïve translation of macro-level proportions into microdata risks compounding structural errors. Consequently, this scoping review will map the current state of the art, evaluate the technical feasibility of these approaches within simulation environments, and establish a conceptual framework to guide future public health mathematical modeling5.

Despite the widespread adoption of macro-level garbage code redistribution frameworks, current literature exhibits a critical translational gap32,41. Specifically, aggregated population-level corrections assume that misclassified or ill-defined causes of death are randomly distributed across demographic strata11,31. However, stochastic microsimulations and agent-based models require highly disaggregated baseline mortality risks, as structural diagnostic errors often disproportionately mask true epidemiological burdens in vulnerable or socially deprived settings44–46. Translating uncorrected macro-level proportions directly into individual-level simulation environments risks compounding classification errors and propagating systematic bias into microdata risk matrices42,43. Therefore, this scoping review aims to evaluate the technical feasibility and methodological assumptions required to adapt macro-level correction algorithms for individual-level microsimulation parameterization5.

Research questions

  1. What analytical methodologies (statistical, probabilistic, deterministic, linkage-based or machine learning) have been described in the literature for the correction and redistribution of “garbage codes” in mortality records?
  2. How are they classified and what types of “garbage codes” (poorly defined causes of ICD chapters, intermediate or terminal causes) are the main objective of each identified methodology?
  3. What is the documented level of applicability or compatibility of these correction methods to structure databases aimed at the development, calibration and validation of microsimulation models in health?
  4. What are the main methodological challenges, limitations and advantages reported when transferring or adapting analytical correction algorithms from the population (macro) level to the individual level (micro)?

ELIGIBILITY CRITERIA

Population

This component defines eligible data records, populations, and geographic or demographic coverages.

  • Inclusion criteria:
    • Data type: Studies utilizing population-based mortality databases, vital statistics systems, official national or subnational health statistics registers, or civil registration and vital statistics (CRVS) systems.
    • Geography and income level: Research focusing on populations within low- and middle-income countries (LMICs) as defined by historical or current World Bank classifications (e.g., regions within Latin America and the Caribbean, Sub-Saharan Africa, and South Asia).
    • Technical exception: Global or multinational studies (e.g., Global Burden of Disease [GBD] analyses), provided they independently disaggregate methodologies and data applicable to LMIC geographies.
    • Demography: Records tracking general human mortality without restrictions on age groups (including neonatal, infant, adult, or geriatric mortality) or sex and gender.
  • Exclusion criteria:
    • Clinical studies restricted to hyper-specific hospital specimens or controlled trials that do not utilize—or aim to correct—population-based records or broader epidemiological information systems.

Concept

This component delineates the core methodologies, algorithms, and analytical tools evaluated in this review.

  • Inclusion criteria:
    • Studies describing, evaluating, validating, or implementing quantitative methods to identify, clean, reclassify, or redistribute “garbage codes” (levels 1 to 4) or ill-defined causes of death (as classified under ICD-10/ICD-11 Chapter XVIII, or ICD-9 equivalents).
    • Methodologies within the following analytical categories are eligible:
      • Traditional statistics and mathematical demography: Proportional redistribution (simple or stratified by age and sex), fractional algorithms, and indirect demographic adjustment methods21.
      • Advanced statistical modeling: Multiple imputation of missing or non-specific data, spatial or spatiotemporal regression modeling, and empirical or hierarchical Bayesian approaches (e.g., CODEm or CoDcorrect analytical algorithms)11,19,31.
      • Data integration (record linkage): Deterministic or probabilistic algorithms linking mortality records with complementary databases (e.g., hospital discharge records, automated verbal autopsies, or epidemiological surveillance systems)23,26.
      • Artificial intelligence and machine learning: Data-driven predictive architectures (e.g., decision trees, gradient boosting, random forests, or other data-mining frameworks) applied to classify or predict underlying causes of death masked by non-specific codes.
  • Exclusion criteria:
    • Studies analyzing mortality data quality in a purely descriptive manner (e.g., reporting isolated completeness percentages or garbage code prevalences) without implementing or evaluating an active mathematical or algorithmic correction method.
    • Research focused solely on correcting numerical under-registration (omission of death capture), unless the methodology explicitly integrates cause-of-death redistribution.

Context

This component defines the data environments and final fields of application toward which the evidence synthesis is oriented.

  • Inclusion criteria:
    • Data-scarce or heterogeneous settings: Research situated in contexts with weak or fragmented health information systems characterized by low medical certification coverage, high regional heterogeneity, or small-area spatiotemporal misalignments.
    • Microsimulation applicability: Studies evaluating or discussing the impact of correction techniques on individual-level microdata disaggregation, the generation of consistent mortality risk matrices, individual trajectory estimations, or parameterizing inputs for health microsimulation platforms.
    • Emergent methodological scope: Studies originally operating at a macro or aggregate level, provided they detail the mathematical assumptions or individual-record classification algorithms necessary to calibrate and parameterize stochastic microsimulations or agent-based models.
  • Exclusion criteria:
    • Methodologies operating exclusively at an irreducible macro-population level whose designs preclude individual-level micro-disaggregation, thereby lacking utility for agent-based or individual-level simulations.

Sources of evidence

  • Publication types: Peer-reviewed original research articles published in indexed journals, methodological book chapters, and authoritative institutional gray literature (e.g., technical reports from the WHO, PAHO, or IHME).
  • Timeframe: Material published from 1996 onwards, capturing the global introduction and widespread adoption of the Tenth Revision of the International Classification of Diseases (ICD-10), which consolidated the term “garbage codes.”
  • Languages: Studies published in English, Spanish, and Portuguese, reflecting the substantial volume of methodological literature generated in Latin America (particularly Brazil).

METHODS

Search strategy

The search strategy will locate published academic literature and online grey literature. The search will focus on evidence published from 1996 onwards, corresponding to the global implementation and widespread adoption of the International Classification of Diseases, Tenth Revision (ICD-10)—a timeframe during which the term “garbage codes” became conceptually consolidated.

  • Academic Literature: To identify academic literature, a three-step search strategy will be developed in collaboration with a medical librarian:
    1. Initial limited search: An initial search of MEDLINE (Pubmed) will be undertaken to analyze text words in titles and abstracts, Medical Subject Headings (MeSH), and author-supplied keywords from relevant articles. (Appendix I)
    2. Full search expansion: A comprehensive search strategy will be developed using these identified terms, then translated and adapted for Scopus, Web of Science, LILACS, and SciELO. The complete search syntax for MEDLINE (Pubmed) is presented in Appendix I.
    3. Reference hand-searching: The reference lists of all selected articles and relevant systematic or scoping reviews will be hand-searched to identify primary sources missed by the electronic search.
  • Grey literature: To capture unindexed institutional reports, technical documents, and policy papers while ensuring full methodological transparency, grey literature retrieval will follow the JBI guidelines for evidence syntheses.
    1. The search will encompass targeted repository exploration of key international and regional health authorities—specifically the World Health Organization (WHO), the Pan American Health Organization (PAHO), the Institute for Health Metrics and Evaluation (IHME), and the Ministry of Health of El Salvador (MINSAL).
    2. Specialized grey literature databases (e.g., CADTH GreyMatters, BASE, and OpenGrey) will be interrogated. For open web searches conducted via Google Scholar, bulk record management will be systematized using the ‘Publish or Perish’ software.
    3. In strict adherence to PRISMA-S standards, every targeted website check, exact search string, date of execution, and precise hit count will be systematically logged in an audit tracking table (Appendix II) to guarantee complete auditability and reproducibility.

Data management and search automation

To ensure the methodological reproducibility and transparency required by the PRISMA-ScR guidelines, this scoping review protocol was registered a priori on the Open Science Framework (OSF) [osf.io/374tm]. The review is reported in accordance with the Preferred Reporting Items for Systematic reviews and Meta-Analyses extension for Scoping Reviews (PRISMA-ScR) checklist (see Appendix III)

Furthermore, the information retrieval process will be systematized within the R statistical environment (v4.5 or higher). The litsearchr package will be employed for the algorithmic optimization of Boolean syntax via term co-occurrence networks. Indexed databases will be interrogated programmatically through Application APIs using the rentrez and rscopus libraries. Data consolidation and duplicate removal will be managed via the synthesisr package using fuzzy string-matching models, supplemented by manual verification of borderline cases. Finally, title and abstract screening will be conducted using the revtools package, supplemented by Latent Dirichlet Allocation (LDA) topic modeling strictly as a computational decision-support tool to assist with thematic clustering and reduce reviewer fatigue. To eliminate any risk of algorithmic bias, every single inclusion and exclusion decision is made independently by two human reviewers, with all discrepancies resolved through consensus or third-party arbitration in strict adherence to the PCC framework.

The final literature search will be re-run immediately prior to submission. Any methodological deviations from this registered protocol will be fully documented and reported in the final publication.

Data extraction and charting

To ensure integrity, traceability, and reproducibility, we will conduct data extraction programmatically within the R statistical environment (v4.5 or higher). This automated approach mitigates the transcription errors inherent in manual processes and ensures a standardized workflow. We have designed an electronic data charting form to capture information across six dimensions:

  • Study metadata: Author, year of publication, country, and study design.

  • Geographic and health system context: Challenges encountered during implementation, such as inadequate medical training or certification capacity, and recommendations for systemic improvement7.

  • Mortality database characteristics: Specific baseline registry quality metrics used to evaluate the initial validity of the input data. A prime example is the five-star rating system, which objectively scores mortality data based on demographic completeness and the prevalence of major garbage codes11, as well as the algorithmic assessment of missing or unexpected values in core demographic fields, which serves as a critical proxy for systemic data quality issues in highly deprived municipalities37.

  • “Garbage code” typology: The specific definition of a garbage code applied by the authors17, alongside the age- and sex-specific epidemiological profiles of these codes, to understand the clinical nature of the missing data at different life stages11.

  • Correction methodologies: The type of redistribution model used, the algorithmic and software ecosystems utilized, and the specific performance metrics reported to demonstrate data quality improvement42.

  • Microsimulation applicability: The explicit capacity of the methodology to generate highly disaggregated microdata and handle parameter uncertainty for stochastic predictive modeling43.

This form will be piloted with a random sample of three included studies to refine consistency and ensure variable granularity. Two reviewers will perform the extraction independently, resolving discrepancies through consensus. The resulting data will be consolidated into a relational structure (e.g., tibble or data.frame) to facilitate transparent downstream analysis. The full extraction schema (exported as JSON/CSV) and corresponding R scripts will be made available via the Open Science Framework (OSF) upon study completion. In accordance with Joanna Briggs Institute (JBI) methodology for scoping reviews, a formal critical appraisal of included sources of evidence will not be conducted, as the objective of this review is to map and synthesize the breadth of existing quantitative methodologies rather than to evaluate the internal validity of individual studies. For a detailed breakdown of the variables and the classification framework utilized, refer to the Data Charting Instrument (Appendix IV).

Data analysis, synthesis, and presentation

This scoping review adheres to the Joanna Briggs Institute methodology. We will conduct all data processing and visualization programmatically within the R statistical environment.

  • Descriptive quantitative analysis: We will calculate absolute frequencies and percentages to characterize the literature by chronological distribution, geographic reach, taxonomic frameworks, cause categories, and analytical techniques. These analytical techniques will be categorized by their computational complexity, ranging from foundational deterministic methods to advanced predictive modeling, machine learning, and spatial analysis.

  • Thematic narrative synthesis: A narrative will accompany the extracted data, addressing the magnitude of garbage codes in data-scarce settings, the methodological diversity of the algorithms, and the conceptual transferability of these methods to preserve microdata validity for agent-based stochastic models. This includes addressing how uncorrected garbage codes disproportionately mask the true mortality burden in populations with higher social deprivation46 and the risk of inappropriate centralization when macro-level regression algorithms are applied to highly localized municipal datasets47. The synthesis will specifically evaluate how included methodologies handle parameter uncertainty and individual-level risk heterogeneity when descending from population-level aggregates to stochastic microdata. Drawing upon methodological precedents from prior comparative risk assessments in Latin America (such as the 2020 baseline burden estimations of sugar-sweetened beverages in El Salvador48) the synthesis will critically appraise how uncorrected vital statistics distort baseline mortality risks for non-communicable diseases (e.g., type 2 diabetes and ischemic heart disease). Furthermore, we will synthesize evidence on whether current redistribution algorithms can be rigorously integrated into agent-based simulation platforms without violating the local validity required for policy modeling and health technology assessments.

  • Mapping and visualization: To visualize the structure and knowledge gaps of the field, we will generate advanced graphics using ggplot249 and igraph50. This will include structured evidence tables cross-referencing correction methods with their input data requirements, two-dimensional heatmaps mapping analytical methods against specific ICD Chapters, and flow network diagrams illustrating the technical pathway required to connect traditional vital statistics with microsimulation parameterization.

Ethical approval

As this scoping review relies exclusively on publicly available published scientific literature, official institutional technical reports, and aggregated population-based secondary data (Civil Registration and Vital Statistics systems), formal ethical approval by an Institutional Review Board (IRB) or ethics committee was not required.

Declaration of generative AI and AI-assisted technologies in the writing process

During the preparation of this protocol, the authors used Gemini Notebook to assist in scanning source PDF documents for the data extraction matrix—followed by 100% human verification of all data points—and for language polishing and proofreading. All evidence selection, critical appraisal, methodological synthesis, and conceptual drafting were performed independently by the authors, who assume full responsibility for the integrity and accuracy of the final content.

Availability of data and materials

All datasets generated or analyzed during this study are included in this published protocol and its supplementary information files. The complete R scripts, customized search syntaxes, and electronic data charting instruments are publicly accessible via the Open Science Framework (OSF) repository [osf.io/374tm]. Any additional customized scripts, codebooks, or intermediate extraction matrices are available from the corresponding author upon reasonable request.s of Interest.

The authors declare that they have no competing interests. This scoping review was conducted solely for academic, educational, and methodological research purposes, with no commercial, financial, corporate, or industrial support that could be construed as a potential conflict of interest.

Funding

This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors. The study was conducted entirely through the academic time, institutional affiliation resources, and professional dedication of the authors.

Authors’ Contributions


References

1.
Setel P, Macfarlane S, Szreter S, Mikkelsen L, Jha P, Stout S, et al. A scandal of invisibility: Making everyone count by counting everyone. Lancet. 2007;370:1569–77. doi:10.1016/S0140-6736(07)61307-5
2.
AbouZahr C, Savigny D de, Mikkelsen L, Setel P, Lozano R, Nichols E, et al. Civil registration and vital statistics: Progress in the data revolution for counting and accountability. Lancet. 2015;386:1373–85. doi:10.1016/S0140-6736(15)60173-8
3.
Mathers CD, Fat DM, Inoue M, Rao C, Lopez AD. Counting the dead and what they died from: An assessment of the global status of cause of death data. Bulletin of the World Health Organization. 2005;83:171–7. doi:10.1590/S0042-96862005000300009
4.
Phillips DE, AbouZahr C, Lopez AD, Mikkelsen L, Savigny D de, Lozano R, et al. Are well functioning civil registration and vital statistics systems associated with better health outcomes? The Lancet. 2015;386(10001):1386–94. doi:10.1016/S0140-6736(15)60170-2
5.
Foreman KJ, Marquez N, Dolgert A, Fukutaki K, Fullman N, McGaughey M, et al. Forecasting life expectancy, years of life lost, and all-cause and cause-specific mortality for 250 causes of death: Reference and alternative scenarios for 201640 for 195 countries and territories. The Lancet. 2018;392(10159):2052–90. doi:10.1016/S0140-6736(18)31694-5
6.
GBD 2016 Causes of Death Collaborators. Global, regional, and national age-sex specific mortality for 264 causes of death, 1980–2016: A systematic analysis for the global burden of disease study 2016. The Lancet. 2017;390(10100):1151–210. doi:10.1016/S0140-6736(17)32152-9
7.
Hart J et al. Improving medical certification of cause of death: Effective strategies and approaches based on experiences from the data for health initiative. BMC Public Health. 2020;20(1):1083. doi:10.1186/s12889-020-09122-4
8.
Mikkelsen L, Phillips D, AbouZahr C, Setel P, Savigny D de, Lozano R, et al. A global assessment of civil registration and vital statistics systems: Monitoring data quality and progress. Lancet. 2015;386:1395–406. doi:10.1016/S0140-6736(15)60171-4
9.
Miki J et al. Saving lives through certifying deaths: Assessing the impact of two interventions to improve cause of death data in peru. BMC Public Health. 2018;18(1):1329. doi:10.1186/s12889-018-6240-3
10.
Iburg KM, Mikkelsen L, Adair T, Lopez AD. Are cause of death data fit for purpose? Evidence from 20 countries at different levels of socio-economic development. PLoS ONE. 2020;15(8):e0237539. doi:10.1371/journal.pone.0237539
11.
Johnson S, Cunningham M, Dippenaar IN, Sharara F, Wool EE, Agesa KM, et al. Public health utility of cause of death data: Applying empirical algorithms to improve data quality. BMC Medical Informatics and Decision Making. 2021;21(1):175. doi:10.1186/s12911-021-01501-1
12.
Murray C, Lopez A. The global burden of disease: A comprehensive assessment of mortality and disability from diseases, injuries and risk factors in 1990 and projected to 2020. Boston: Harvard School of Public Health; 1996.
13.
Naghavi M, Richards N, Chowdhury H, et al. Improving the quality of cause of death data for public health policy: Are all “garbage” codes equally problematic? BMC Med. 2020;18(1):55. doi:10.1186/s12916-020-01509-w
14.
World Health Organization. International classification of diseases for mortality and morbidity statistics eleventh revision: Reference guide [Internet]. Geneva: World Health Organization; 2022. Available from: https://icd.who.int/browse11
15.
Ahern RM, Lozano R, Naghavi M, Foreman K, Gakidou E, Murray CJ. Improving the public health utility of global cardiovascular mortality data: The rise of ischemic heart disease. Population Health Metrics. 2011 Dec;9(1):8. doi:10.1186/1478-7954-9-8
16.
Salerno P, Cotton A, Chen Z, Nascimento B, Petermann-Rocha F, Deo S. Forecasting ischemic heart disease, stroke, and peripheral artery disease mortality in brazil through 2030. Arquivos Brasileiros de Cardiologia. 2025;122(1):e2024012. doi:10.36660/abc.20240012
17.
Bigoni A, Cunha AR da, Antunes JLF. Redistributing deaths by ill-defined and unspecified causes on cancer mortality in brazil. Rev Saude Publica. 2021 Dec;55:106–6. doi:10.11606/s1518-8787.2021055003319
18.
Cunha AR da, Bigoni A, Antunes JLF, Hugo FN. Impact of redistributing deaths by ill-defined causes in oral and oropharyngeal cancer mortality in brazil. Brazilian Oral Research. 2022;36:e117. doi:10.1590/1807-3107bor-2022.vol36.0117
19.
Foreman K, Naghavi M, Ezzati M. Improving the usefulness of US mortality data: New methods for reclassification of underlying cause of death. Popul Health Metr. 2016;14:14. doi:10.1186/s12963-016-0082-4
20.
Murray CJ, Dias RH, Kulkarni SC, Lozano R, Stevens GA, Ezzati M. Improving the comparability of diabetes mortality statistics in the u.s. And mexico. Diabetes Care. 2008;31(3):451–8. doi:10.2337/dc07-1725
21.
Naghavi M, Makela S, Foreman K, et al. Algorithms for enhancing public health utility of national causes-of-death data. Popul Health Metr. 2010;8(1):9. doi:10.1186/1478-7954-8-9
22.
GBD 2017 Causes of Death Collaborators. Global, regional, and national age-sex-specific mortality for 282 causes of death in 195 countries and territories, 1980–2017: A systematic analysis for the global burden of disease study 2017. The Lancet. 2018;392(10159):1736–88. doi:10.1016/S0140-6736(18)32203-7
23.
França E et al. Ill-defined causes of death in brazil: A redistribution method based on the investigation of such causes. Rev Saúde Pública. 2014;48(4):671–81. doi:10.1590/S0034-8910.2014048005121
24.
França EB, Ishitani LH, Teixeira RA, Cunha CC da, Marinho MF. Improving the usefulness of mortality data: Reclassification of ill-defined causes based on medical records and home interviews in brazil. Revista Brasileira de Epidemiologia. 2019;22:e190010. doi:10.1590/1980-549720190010.supl.3
25.
França EB. Garbage codes assigned as cause-of-death in health statistics. Revista Brasileira de Epidemiologia. 2019;22:e190001. doi:10.1590/1980-549720190001.supl.3
26.
Bierrenbach A et al. Redistribution of heart failure deaths using two methods: Linkage of hospital records with death certificates. Rev Bras Epidemiol. 2019;22(Suppl 3):e190006. doi:10.1590/1980-549720190006.supl.3
27.
Stevens GA, King G, Shibuya K. Deaths from heart failure: Using coarsened exact matching to correct cause-of-death statistics. Population Health Metrics. 2010;8(1):6. doi:10.1186/1478-7954-8-6
28.
Fihel A, Muszyńska-Spielauer MM. Using multiple cause of death information to eliminate garbage codes. Demographic Research. 2021;45:345–60. doi:10.4054/DemRes.2021.45.11
29.
Giraldo L, Rodrı́guez L, González-Robledo MC, Mino-León D, Cahuana-Hurtado L, Rojas-Russell ME, et al. Modelo para el análisis de la mortalidad en colombia 2000-2012. Revista de Salud Pública. 2017;19(2):166–73. doi:10.15446/rsap.v19n2.65825
30.
Masquelier B et al. Analysis of death registers in antananarivo, madagascar. BMC. 2019. doi:10.2471/BLT.18.223842
31.
Grigoriev P, Bonnet F, Perdrix E. Method for redistributing ill-defined causes of death. Popul Health Metr. 2024;22(1):1–14. doi:10.1186/s12963-024-00332-6
32.
Scohy A, Lesnik T, Devleesschauwer B, Haneef R. The importance of redistribution process of ill-defined deaths in national mortality database. Archives of Public Health. 2025 Jul 11;83(1):185. doi:10.1186/s13690-025-01652-x
33.
Kyu HH et al. Estimating the burden of HIV/AIDS-related mortality: HIV/AIDS-related garbage code redistribution packages. Journal of the International AIDS Society. 2021. doi:10.1002/jia2.25791
34.
Zhang X, Yang W, Wang J, Ai L, Chen M, Wang C, et al. Reallocating diabetes-related garbage codes to improve mortality estimates: A case study in weifang, china. Population Health Metrics. 2025;23(1):38. doi:10.1186/s12963-025-00399-5
35.
Teixeira B, Toporcov T, Chiaravalloti-Neto F, Chiavegatto Filho A. Spatial clusters of cancer mortality in brazil: A machine learning modelling approach. In: Anexo do curriculo lattes. 2020.
36.
Araújo Santos Camargo JD de, Camargo SF, Souza ATB de, Oliveira Freitas AKMS de, Ferreira CA, Sarmento ACA, et al. Regional disparities in breast cancer mortality in brazil: A spatial analysis using uncorrected and adjusted data, 2000–2023. Scientific Reports. 2026;16:6770. doi:10.1038/s41598-026-37844-w
37.
Zimeo-Morais GA, Miraglia JL, Oliveira BZ de, Mistro S, Hisatugu WH, Greffin D, et al. Factors associated with the quality of death certification in brazilian municipalities: A data-driven non-linear model. PLoS One. 2023;18(8):e0290814. doi:10.1371/journal.pone.0290814
38.
Souza ARM de, Raposo LM, Lacerda GCB de, Godoy PH. Stroke deaths profile and its subtypes in brazil: Analysis using machine learning. Global Heart. 2025;20(1):85. doi:10.5334/gh.1476
39.
Phillips DE, Lozano R, Naghavi M, Atkinson C, Gonzalez-Medina D, Mikkelsen L, et al. A composite metric for assessing data on mortality and causes of death: The vital statistics performance index. Population Health Metrics. 2014;12(1):14. doi:10.1186/1478-7954-12-14
40.
Mikkelsen L, Moesgaard K, Hegnauer M, Lopez AD. ANACONDA: A new tool to improve mortality and cause of death data. BMC Medicine. 2020;18(1):61. doi:10.1186/s12916-020-01521-0
41.
Devleesschauwer B. Skills building seminar: Redistribution of ill-defined deaths: A methodological conundrum. In: European journal of public health. Oxford University Press; 2022. p. ckac129–618. doi:10.1093/eurpub/ckac129.618
42.
Ng TC, Lo WC, Ku CC, Lu TH, Lin HH. Improving the use of mortality data in public health: A comparison of garbage code redistribution models. American Journal of Public Health. 2020;110(2):222–9. doi:10.2105/AJPH.2019.305432
43.
Chrysanthopoulou SA, Rutter CM, Gatsonis CA. Bayesian versus empirical calibration of microsimulation models: A comparative analysis. Medical Decision Making. 2021 Aug;41(6):714–26. doi:10.1177/0272989X211009161
44.
Costa LFL, Montenegro M de MS, Rabello Neto D de L, Oliveira ATR de, Trindade JE de O, Adair T, et al. Estimating completeness of national and subnational death reporting in brazil: Application of record linkage methods. Population Health Metrics. 2020;18(1):22. doi:10.1186/s12963-020-00229-a
45.
França E, Ishitani LH, Teixeira R, Duncan BB, Marinho F, Naghavi M. Changes in the quality of cause-of-death statistics in brazil: Garbage codes among registered deaths in 1996–2016. Population Health Metrics. 2020;18(1):20. doi:10.1186/s12963-020-00221-4
46.
Malta DC, Teixeira RA, Cardoso LS de M, Souza JB de, Bernal RTI, Pinheiro PC, et al. Premature mortality due to noncommunicable diseases in brazilian capitals: Redistribution of garbage causes and evolution by social deprivation strata. Revista Brasileira de Epidemiologia. 2023;26:e230002.supl.1. doi:10.1590/1980-549720230002.supl.1
47.
Liu L, Wang X, Wang C, Ma X, Meng X, Ning B, et al. A study on garbage code redistribution methods in small area: Redistributing heart failure in two chinese cities by two approaches. Research Square (Preprint). 2022. doi:10.21203/rs.3.rs-1242825/v1
48.
Rodríguez Cairoli F, Guevara Vásquez G, Bardach A, Espinola N, Perelli L, Balan D, et al. Carga de enfermedad y económica atribuible al consumo de bebidas azucaradas en el salvador. Revista Panamericana de Salud Pública. 2023;47:e80. doi:10.26633/RPSP.2023.80
49.
Wickham H. ggplot2: Elegant graphics for data analysis. New York: Springer-Verlag; 2016. doi:10.1007/978-3-319-24277-4
50.
Csardi G, Nepusz T. The igraph software package for complex network research. InterJournal, Complex Systems. 2006;1695:1–9.

Appendices

Appendix I: Search strategy

Database: MEDLINE (via Ovid)
Search conducted on: July 2026
Planned Limits:

* Date restrictions: from 1996 to present to capture historical development of algorithms like GBD.

* Language restrictions: None at search level (any required filters based on reviewer capacities will be applied during the programmatic screening stage).

* Document types: All source types (including technical notes, methodology papers, and electronic articles).

Line Search Query (Terms, Truncations, and Syntax) Conceptual Block / Target
#1 exp Mortality/ or exp Cause of Death/ PCC: Population / Focus (Core mortality registry filters)
#2 (mortality data* or death certificate* or vital statistic* or vital registration or CRVS).tw,kf,ot. Keywords for vital statistics and national death registry systems
#3 1 or 2 Result: Core Mortality Data
#4 (garbage code* or ill-defined or misclassified or misclassification* or undefined cause* or ill-defined cause* or unspecific cause* or vague code* or redistribution algorithm*).tw,kf,ot. PCC: Concept (Block A) - Specific typology of Data Errors (“Garbage Codes”)
#5 (data cleaning or data correction or diagnostic accuracy or underlying cause of death).tw,kf,ot. Data quality attributes specific to mortality databases
#6 4 or 5 Result: Garbage Code Concepts
#7 exp Algorithms/ or exp Models, Statistical/ or exp Computer Simulation/ MeSH Indexing for Analytical Methods
#8 (demographic redistribution or multiple imputation or MICE or Bayesian model* or Markov chain* or machine learning or random forest* or neural network* or artificial intelligence or algorithmic correction or fractional assignment).tw,kf,ot. PCC: Concept (Block B) - Specific statistical and computational correction frameworks
#9 (microsimulation* or micro-simulation* or agent-based model* or stochastic simulation* or individual-level dynamic* or synthetic population*).tw,kf,ot. PCC: Context / Transferability - Target micro-level application architectures
#10 7 or 8 or 9 Result: Methodological Ecosystem
#11 exp “Developing Countries”/ MeSH Indexing for LMICs
#12 (low income country or low-and-middle income countr* or LMIC or LMICs or developing nation* or resource-constrained setting* or transitional econom*).tw,kf,ot. Geographic Context Keywords (Based on World Bank specifications)
#13 (Africa* or Asia* or South America* or Central America* or Latin America*).tw,kf,ot. Broad regional geographic filters
#14 11 or 12 or 13 Result: Developing Context (LMIC)
#15 3 and 6 and 10 and 14 COMBINED STRATEGY (Boolean Intersection)

Appendix II: PRISMA-S audit and mapping table

The following table details the adherence to the 16 items of the PRISMA-S guideline for the scoping review protocol entitled “Methodologies for the correction of Garbage codes in mortality records and their applicability in microsimulation models”, ensuring complete search auditability, transparency, and computational reproducibility.

PRISMA-S Item Reporting Requirement Specific Implementation in Protocol & R Pipeline Location / Status
1. Information sources Describe all information sources consulted (databases, registries, websites, grey literature). Indexed databases: PubMed/MEDLINE, Scopus, Web of Science, LILACS, and SciELO. Institutional grey literature: WHO, PAHO, IHME, and MINSAL. Specialized grey repositories: CADTH GreyMatters, BASE, and OpenGrey. Section 3.1 & Section 3.2
2. Electronic search strategy Present the full electronic search strategy for at least one major database, including syntax and operators. Complete search query designed for MEDLINE (via Ovid) structured under the Population, Concept, and Context (PCC) framework. Section 6.1 (Appendix I)
3. Database platforms and vendors Specify the database platforms or interface vendors used for each search. Programmatic querying via Application APIs and specialized R libraries (rentrez for PubMed, rscopus for Scopus) and Ovid interface for MEDLINE. Section 3.2
4. Translation of search strategy Describe how the search strategy was adapted for different thesauri and syntaxes. Adaptation of MeSH descriptors and free-text terms optimized via term co-occurrence networks using the litsearchr package in R. Section 3.2
5. Limits and restrictions Specify any search-level restrictions applied (dates, languages, document types). Temporal restriction from 1996 onwards (marking global ICD-10 implementation and consolidation). No initial language restrictions (English, Spanish, Portuguese). Section 2.4 & Section 6.1
6. Search dates Indicate the exact date when searches were executed or scheduled for each source. Primary literature search scheduled a priori for October 2026, with a final verification re-run prior to synthesis. Section 3.1 & Section 6.1
7. Contacting authors Indicate if study authors, experts, or key researchers were contacted to identify unpublished studies. Planned consultations with vital statistics epidemiologists, public health officials, and key regional collaborators in Latin America. Section 3.1
8. Registration Cite the a priori protocol registration and registry platform used. Registered a priori on the Open Science Framework (OSF). Introduction & Section 3.3
9. Supplementary searches (Citation searching) Describe supplementary search methods (backward and forward citation tracking). Manual backward citation tracking (backwards citation searching) on all studies selected for full-text extraction. Section 3.1
10. Peer review of search strategy Indicate if the search strategy was peer-reviewed by an expert (e.g., librarian or PRESS). Search strategy developed in collaboration with a medical librarian and algorithmically optimized via term co-occurrence networks in R. Section 3.1 & Section 3.2
11. Total records identified Report the total number of records identified from each database and grey literature source. Pending empirical execution (October 2026). Will be formally documented in the final PRISMA-ScR flow diagram. Flow Diagram / Results
12. Duplicate removal Describe the process and tools used for duplicate removal. Automated consolidation using the synthesisr package in R via fuzzy string-matching, with manual verification of edge cases. Section 3.2
13. Selection process Describe the study selection and screening process (title/abstract and full-text screening). Dual independent screening by two human reviewers via revtools, assisted by Latent Dirichlet Allocation (LDA) topic modeling strictly as a human-in-the-loop decision-support tool. Section 3.2
14. Data collection process Describe the data collection process and extraction tools. Standardized programmatic data charting in R across six analytical dimensions, executed independently and in duplicate. Section 3.3
15. Data items List and define all variables planned for extraction. Electronic data charting form structured across 6 blocks: Metadata, Health System Context, Registry Characteristics, Garbage Code Typology, Correction Methodology, and Microsimulation Applicability. Section 6.2 (Appendix II)
16. Search auditability & reproducibility Ensure complete audit trail and reproducibility of the search process. Executable R scripts, syntax codes, and extraction matrices exported and permanently archived in the project’s OSF repository. Section 3.2 & Section 3.3

Appendix III: Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Reviews (PRISMA-ScR) Checklist

Section / Topic Item # PRISMA-ScR Checklist Item Reported on Page / Section Specific Implementation in Protocol
TITLE
Title 1 Identify the report as a scoping review. Title Page & Section 1 Declared as a Scoping Review Protocol in the main title.
ABSTRACT
Structured summary 2 Provide a structured summary including: background, objectives, eligibility criteria, sources of evidence, charting methods, and main synthesis plan. Section 1 (Abstract) Structured abstract following JBI and PRISMA-ScR standards.
INTRODUCTION
Rationale 3 Describe the rationale for the review in the context of what is already known. Section 1.1 (Background) Details the “garbage codes” problem, limitations of macro-level corrections, and the translational gap toward microsimulation.
Objectives 4 Provide an explicit statement of the question(s) or objective(s) being addressed. Section 1.2 (Research Questions) Formulates 4 explicit research questions using the Population, Concept, and Context (PCC) framework.
METHODS
Protocol & registration 5 Indicate if a review protocol exists; state if registered and supply registration info. Section 3.2 & OSF Link Registered a priori on the Open Science Framework (OSF ID: 10.17605/OSF.IO/ZP9AQ).
Eligibility criteria 6 Specify characteristics used to decide eligibility (e.g., PCC, study design, temporal limits). Section 2 (Eligibility Criteria) Structured according to Population (mortality records/LMICs), Concept (garbage code algorithms), and Context (microsimulation applicability).
Information sources 7 Describe all information sources (databases, grey literature, contact with authors). Section 3.1 & Appendix I/III MEDLINE/Ovid, Scopus, Web of Science, LILACS, SciELO, plus WHO, PAHO, IHME grey literature repositories.
Search strategy 8 Present full electronic search strategy for at least one database, including limits. Appendix I (Section 7.1) Complete Ovid MEDLINE syntax provided with Boolean operators, MeSH terms, and line-by-line targets.
Selection process 9 State process for selecting sources (screening, eligibility, deduplication tools). Section 3.2 Programmatic deduplication (synthesisr), dual independent screening (revtools), assisted by LDA topic modeling.
Data charting process 10 Describe methods of charting data (forms, piloting, reviewer duplication). Section 3.3 & Appendix II Dual independent programmatic charting using a 6-block relational data extraction schema in R.
Data items 11 List and define all variables for which data were sought. Section 3.3 & Appendix II Defines 20 variables across metadata, health context, registry quality, garbage code typology, methods, and TRL/microsimulation applicability.
Critical appraisal 12 State whether critical appraisal was conducted; if not, state why. Section 3.3 In accordance with JBI scoping review guidelines, formal risk-of-bias/critical appraisal is omitted as the objective is to map methodologies.
Synthesis of results 13 Describe methods of handling and summarizing data (tables, charts, narrative). Section 3.4 & 3.5 Combines descriptive quantitative metrics, thematic narrative synthesis, ggplot2 matrix heatmaps, and igraph network diagrams.
RESULTS
Selection of sources 14 Report numbers of records screened, assessed for eligibility, and included. Pending Execution To be reported in the PRISMA-ScR Flow Diagram upon study completion (September 2026).
Characteristics of sources 15 Describe characteristics of included sources of evidence. Pending Execution To be tabulated in the final review using the Appendix II extraction schema.
Critical appraisal results 16 Share results of any critical appraisal of individual sources. Not Applicable Not conducted, in strict adherence to JBI scoping review methodology.
Results of individual sources 17 For each source, present relevant data charted. Pending Execution Will be archived as open CSV/JSON matrices on the OSF project repository.
Synthesis of results 18 Summarize and/or present the charting results in relation to objectives. Pending Execution Will map analytical techniques against microsimulation parameterization potential.
DISCUSSION
Summary of evidence 19 Summarize main findings; link to objectives and broader literature. Pending Execution To be drafted upon synthesis completion.
Limitations 20 Discuss limitations of the scoping review process. Pending Execution To be discussed in the final review report.
Conclusions 21 Provide explicit conclusions linked to objectives and implications. Pending Execution To be provided in the final manuscript.
FUNDING
Funding 22 Describe sources of funding and role of funders. Section 7 Self-funded through institutional academic time; no commercial or external grants.

Appendix IV: Data extraction instrument

Block Variable Description / Purpose Suggested Format / Values
I. Identification & Metadata Study ID Unique identifier assigned by R workflow Numeric (e.g., 001)
Bibliographic Data Author, year, title, and journal/institution Text
Evidence Type Original article, technical report, or method book Dropdown
II. Geographic & Health Context Geographic Scope Country, region, or sub-national area Text
Economic Classification World Bank income level Dropdown (Low, Lower-Middle, Upper-Middle)
CRVS System Status Description of Civil Registration and Vital Statistics status Text
III. Mortality Database Data Periodicity Years covered by mortality records Text / Numeric
Data Volume Sample size (number of deaths analyzed) Numeric
Primary Source Origin of data (e.g., forensic, hospital, surveys) Text
ICD Framework Version of the ICD utilized (ICD-9, ICD-10, ICD-11) Dropdown
IV. Junk Code Typology Definition Framework Criteria used to identify “junk” (e.g., GBD list) Text
Evaluated ICD Codes Specific codes being targeted for correction Text
Baseline Magnitude Percentage/volume of junk codes in original records Percentage (%)
V. Correction Methodology Central Technique Algorithmic classification Dropdown (e.g., MICE, Bayesian, ML)
Theoretical Assumptions Underlying mathematical principles/assumptions Text
Software Ecosystem Tools used (R, SAS, Stata, etc.) Text
Uncertainty/Bias Report Reported confidence intervals or limitations Yes / No / Partial
VI. Microsimulation Applicability TRL (Technology Readiness Level) Maturity of the development method (1-4) Scale (1-4)
Output Format Aggregate vs. individual-level data output Dropdown
Demographic Granularity Consistency at sub-group/small-area levels Scale (1-5)
Parametrization Potential Capability to feed agent-based/microsimulation models Scale (1-5)
Reviewer Notes Professional notes on applicability and transferability Free text