PKP/OJS multi-source journal enrichment — simple version
This document performs Full-Cohort Journal Enrichment. It starts with every row in the PKP/OJS list. It looks for OpenAlex and Crossref records that share an exact, valid ISSN with each row. It does not use a title or publisher name to make a match.
This is Exploratory Analysis. It prepares evidence for review. It does not decide that two records describe the same journal when the evidence is uncertain.
The final Generated Artifact is a Parquet table with all input rows and all available top-level OpenAlex and Crossref fields. A second Parquet file contains the fixed ten-row Journal Enrichment Sample. Complex source values stay in JSON text so that the analysis does not discard them.
How to read the match result
An ISSN is the main journal identifier in this analysis. Identifier Availability tells us whether an input row has at least one valid ISSN. Each source can then give one of four match results:
Result
Plain-language meaning
not_attempted
The input row has no valid ISSN, so an exact-ISSN search is not possible.
unmatched
The row has a valid ISSN, but the source has no record with that ISSN.
unique
The source has one record with the ISSN. Its fields are expanded into columns.
ambiguous
The source has more than one record with the ISSN. All candidates are kept; the code does not select a winner.
The code can reuse saved data to avoid expensive API requests. A cache is only a saved copy of an earlier download or calculation. The complete download code remains in this document. Set reuse_shared_cache to FALSE and delete this document’s source-cache to request a fresh source snapshot. Missing OpenAlex batches require OPENALEX_API_KEY. A missing Crossref snapshot requires CROSSREF_MAILTO. Delete openalex-expanded.rds and crossref-expanded.rds after a change to the matching or field-expansion code.
Packages and paths
This block prepares the tools and file locations. All new files go to this analysis’s own artifact directory. reuse_shared_cache permits read-only reuse of earlier source data. The code still writes new data only to this analysis’s directory.
valid_pkp_file() compares the input file with the expected content fingerprint. This check makes sure that the analysis uses the intended PKP V7 release, not a different file with the same name.
library(dplyr)
Attaching package: 'dplyr'
The following objects are masked from 'package:stats':
filter, lag
The following objects are masked from 'package:base':
intersect, setdiff, setequal, union
This block defines how to obtain source data when a usable cache is not available. read_cache() returns a saved R object when it can read one. Otherwise, the relevant fetch_*() function calls the source API.
OpenAlex accepts the input ISSNs in groups of 100. Crossref supplies its journal directory in pages of 1,000 records. A short timeout prevents a stalled request from waiting forever. Each request can make at most four attempts. The cursor checks stop the code if an API repeats the same page position.
This block reads the PKP V7 list. Each row is one OJS record. All 19 source fields are read as text so that R does not silently change an identifier, date, or special value.
row_identity() uses the OAI address, repository name, and OAI set identifier together as a stable row identity. The two checks confirm that the fields needed for matching exist and that every input row is an OJS row. input_issns then stores the valid ISSNs for each row. tokens is the distinct ISSN list sent to the source APIs.
These functions turn exact ISSN evidence into result columns.
index_candidates() records which source objects contain each ISSN.
match_candidates() compares one PKP row with that index and assigns one of the four match results defined above.
expand_source_fields() creates one result column for every top-level source field. A simple source value keeps its number, true/false, or text type. A nested value becomes JSON text.
For a unique match, the source fields appear directly in the expanded columns. For an ambiguous match, the code leaves those single-record columns empty and stores every complete candidate in *_candidates__json. This keeps the evidence without making an unsupported choice.
This block processes OpenAlex in four steps. It reuses or retrieves each 100-ISSN batch. It removes duplicate OpenAlex source IDs. It matches each PKP row by exact ISSN. It then expands all observed OpenAlex fields.
base_result also adds identifier_status, the ISSNs used for the search, the ISSNs that matched, the match result, and the candidate count. The expanded-data cache avoids repeating the most expensive field conversion. Reuse it only when the source cache and the matching code have not changed.
used (Mb) gc trigger (Mb) limit (Mb) max used (Mb)
Ncells 62524857 3339.2 144509400 7717.7 NA 144509400 7717.7
Vcells 252579728 1927.1 369639819 2820.2 16384 294458067 2246.6
Crossref enrichment
This block applies the same exact-ISSN method to the Crossref journal directory. It first keeps only Crossref records that contain an ISSN used by the PKP list. It does not merge distinct Crossref objects merely because they have the same ISSN set. Therefore, a shared ISSN can correctly produce an ambiguous result.
The final result joins the original PKP fields, the match evidence, and every expanded OpenAlex and Crossref field. The original row order stays unchanged.
This block creates the two Generated Artifacts. sample_spec identifies ten specific rows by stable row identity. This Journal Enrichment Sample is not a random or representative sample. It is a visual review tool that shows several match situations and all expanded fields.
Before writing files, the checks confirm three points: the sensitive admin_email field is absent; the input rows and their order are unchanged; and each ambiguous JSON value contains the reported number of candidates.
Arrow writes each Parquet file to a .pending path with Zstandard compression. The code checks the full-table size and column names, and it reads the complete sample back. Only then does it give both files their final names. This short staging step protects an earlier valid result if writing stops partway through.
sample_spec <-tribble(~context_name, ~oai_url, ~repository_name, ~set_spec,"Kuwait Journal of Science","https://journalskuwait.org/kjs/index.php/index/oai","Kuwait Journal of Science", "KJS","Cakrawala","https://cakrawalajournal.org/index.php/index/oai","CAKRAWALA", "cakrawala","Sintesa: Jurnal Ilmu Pendidikan","https://sintesa.stkip-arrahmaniyah.ac.id/index.php/index/oai","Sintesa: Jurnal Ilmu Pendidikan STKIP Arrahmaniyah Depok", "sintesa","CARAKA: Jurnal Teologi Biblika dan Praktika","https://ojs.sttibc.ac.id/index.php/index/oai","CARAKA:Jurnal Teologi", "ibc","Jurnal Tata Kelola dan Akuntabilitas Keuangan Negara","https://jurnal.bpk.go.id/index.php/index/oai","Open Journal Systems", "TAKEN","JOEEL (Journal of English Education and Literature)","https://journal.stkippamanetalino.ac.id/index.php/index/oai","Rumah Jurnal STKIP Pamane Talino Ngabang", "bahasa-inggris","Malaysian Journal of Paediatrics and Child Health","https://mpaeds.my/journals/index.php/index/oai","NA", "MJPCH","Jami Scientific Research Quarterly Journal","https://journals.jami.edu.af/index.php/index/oai","Jami University Press", "jsrqj","Journal of Polymer & Composites","https://engineeringjournals.stmjournals.in/index.php/index/oai","Engineering Journals", "JoPC","JURNAL TEKNIK PERTAMBANGAN","https://e-journal.upr.ac.id/index.php/index/oai","Journal Online Universitas Palangka Raya", "JTP")sample_positions <-match(row_identity(sample_spec), row_identity(result))stopifnot(!anyNA(sample_positions))sample_result <- result[sample_positions, ]candidate_json_is_complete <-function(matches, values) {map2_lgl(matches, values, \(match, value) { candidate_count <-length(match$positions)if (candidate_count <=1L) return(is.na(value)) parsed <-tryCatch( jsonlite::fromJSON(value, simplifyVector =FALSE),error = \(error) NULL )is.list(parsed) &&length(parsed) == candidate_count })}stopifnot(!"admin_email"%in%names(result),identical(row_identity(result), row_identity(pkp_rows)),all(candidate_json_is_complete( openalex_matches, result$openalex_candidates__json )),all(candidate_json_is_complete( crossref_matches, result$crossref_candidates__json )))pending_output_path <-paste0(output_path, ".pending")pending_sample_output_path <-paste0(sample_output_path, ".pending")arrow::write_parquet(as.data.frame(result), pending_output_path,compression ="zstd")arrow::write_parquet(as.data.frame(sample_result), pending_sample_output_path,compression ="zstd")parquet_info <- nanoparquet::read_parquet_info(pending_output_path)parquet_schema <- nanoparquet::read_parquet_schema(pending_output_path)observed_sample <- nanoparquet::read_parquet(pending_sample_output_path)stopifnot( parquet_info$num_rows ==nrow(result), parquet_info$num_cols ==ncol(result),identical( parquet_schema$name[!is.na(parquet_schema$r_col)],names(result) ),isTRUE(all.equal(as.data.frame(observed_sample),as.data.frame(sample_result),check.attributes =FALSE )))stopifnot(file.rename(pending_output_path, output_path),file.rename(pending_sample_output_path, sample_output_path))
Ten-row visual review
The complete Generated Artifact contains 98,273 rows and 81 columns. The Journal Enrichment Sample contains 10 of those rows.
Journal Enrichment Sample
The interactive DT table below shows all result columns. Use its search boxes and horizontal scroll to inspect which expanded fields are useful.
This sample supports visual review. It does not represent the frequency of each match result in the complete table.
This table shows how many input rows received each match result from each source. Use the four definitions at the start of this document to interpret the counts.
This list shows every column in the Generated Artifact. It helps reviewers find the original PKP fields, the match evidence, and the expanded source fields. The display does not replace the complete Parquet file.