Abstract. Cancer-associated cachexia is a multifactorial wasting syndrome that combines skeletal-muscle loss, systemic inflammation, and altered energy and protein balance, and it remains incompletely understood at the molecular level in human muscle. This pilot report reanalyzes a public RNA-sequencing count matrix from 84 rectus abdominis muscle biopsies of patients with colorectal or pancreatic cancer, originally generated by Bhatt et al. (2025) and deposited in the NCBI Gene Expression Omnibus (GEO) under accession GSE254877/GSE254878. Because the published clinical cachexia subtype labels are not encoded in the raw count file, an independent, unsupervised workflow was used: protein-coding transcripts were filtered and converted to log2 counts per million, the 2,000 most variable genes were reduced by principal component analysis, and two exploratory clusters were assigned by k-means (Cluster_A, n = 59; Cluster_B, n = 25). A descriptive Welch t-test screen and eight pre-specified cachexia-related mechanism gene panels (covering proteolysis, inflammation, the myostatin/activin axis, muscle regeneration, autophagy, mitochondrial energy metabolism, the GDF15 axis, and the FoxP1 axis) were then applied to characterize the clusters. The inflammatory-signaling panel and several individual genes tied to cytokine and activin signaling (INHBA, TNF, IL1B, GFRAL, GDF15, FOXP1) showed the clearest, though modest, separation between clusters, while core atrophy and regeneration genes (FBXO32, TRIM63, MSTN, PAX7) did not differ meaningfully. These exploratory clusters are arbitrary and are not equivalent to the cachexia and non-cachexia subtypes reported in the source publication. The results are presented as a hypothesis-generating pilot, intended to prioritize candidates for confirmatory, covariate-adjusted differential-expression analysis and independent experimental validation rather than to support any clinical or therapeutic claim.
Keywords: cancer cachexia, skeletal muscle, RNA sequencing, transcriptomics, GDF15, FoxP1, bioinformatics pilot analysis
Scope note. This is a hypothesis-generating, unsupervised bioinformatics pilot on public, previously de-identified aggregate data. It does not reproduce the clinical cachexia subtypes reported in the original publication, does not establish causation, and does not support any clinical, diagnostic, or therapeutic claim.
Cachexia is a progressive wasting syndrome that develops in a large share of patients with advanced cancer and is defined by involuntary loss of skeletal-muscle mass, with or without loss of fat mass, that cannot be fully reversed by ordinary nutritional support. The condition is multifactorial: it involves a negative energy and protein balance, chronic low-grade inflammation, and overactivation of the two principal muscle-protein-degradation systems, the autophagy–lysosome pathway and the ubiquitin–proteasome system (Libramento et al., 2025). Beyond the physical burden of muscle loss, cachexia is associated with reduced treatment tolerance, poorer quality of life, and shorter survival, which makes understanding its molecular drivers in human muscle an important and still unresolved research problem.
A recent large-scale transcriptomic study addressed part of this gap by profiling rectus abdominis muscle biopsies from 84 patients with colorectal or pancreatic cancer using integrative non-negative matrix factorization of coding and non-coding RNA (Bhatt et al., 2025). That analysis identified two coherent molecular subtypes of human cachectic muscle, one of which was associated with high-grade weight loss, reduced muscle mass, atrophy of type IIA and IIX muscle fibers, and shorter survival. The underlying count matrix and sample metadata were deposited publicly in GEO (accessions GSE254877 and GSE254878, part of the SuperSeries GSE292053), which makes independent reanalysis possible. However, the specific per-patient subtype assignments used in the original publication are not included in the public files, so any secondary reanalysis of the raw counts must build its own exploratory grouping rather than reproduce the published subtypes directly.
This report treats that public matrix as an opportunity for a modest, transparent pilot analysis rather than an attempt to replicate the original study’s clustering or clinical findings. The guiding question is not which single drug might make a tumor less aggressive, but a narrower and more tractable one: which muscle–tumor communication pathways are associated with expression differences in this cohort, and could a candidate intervention targeting one of them plausibly preserve muscle mass and function without weakening anti-tumor immunity or interfering with standard cancer treatment. Three linked objectives follow from that question. First, an independent unsupervised clustering of the public expression matrix was performed to characterize the structure of the data without assuming the original clinical labels. Second, eight literature-motivated mechanism gene panels were scored to see whether any pathway showed a consistent difference between the resulting clusters. Third, the pilot findings are placed alongside independent lines of evidence — a completed clinical trial of an anti-GDF15 antibody (Groarke et al., 2024), a preclinical mouse model of FoxP1-driven wasting (Neyroud et al., 2021), and a recent narrative review of cachexia mechanisms (Libramento et al., 2025) — to give the exploratory results some external context.
The primary dataset was the public raw RNA-sequencing count matrix associated with GEO series GSE254877 (RNA-seq subseries) and GSE254878 (a related non-coding RNA subseries), both part of the SuperSeries GSE292053 titled Molecular Subtypes of Human Skeletal Muscle in Cancer Cachexia (National Center for Biotechnology Information [NCBI], 2025). The series includes 84 rectus abdominis muscle biopsies collected from patients with colorectal or pancreatic cancer and profiled on an Illumina NovaSeq 6000 platform. All data used here were previously de-identified and made publicly available by the original investigators; no new human participants were recruited for this pilot, and no protected health information was accessed. Cancer type (colorectal or pancreatic) was parsed directly from the public GEO sample titles associated with each accession; no other clinical variables were available in the public files used for this pilot.
The raw count matrix was restricted to protein-coding features, and duplicate gene-symbol rows were summed. Low-abundance genes were removed using a pre-specified rule requiring at least 10 counts in at least 20% of the 84 samples. This left 16,160 filtered protein-coding genes for downstream analysis. Filtered counts were converted to log2 counts per million plus one (log2-CPM+1) for all exploratory visualization, clustering, and scoring steps described below.
The 2,000 most variable genes by variance on the log2-CPM scale were selected and standardized (mean-centered, unit-scaled) across samples. Principal component analysis (PCA) was then applied to this standardized matrix, and k-means clustering (k = 2, 50 random restarts) was fitted to the first ten principal components. The resulting two clusters were labeled Cluster_A and Cluster_B, with Cluster_A assigned to the group with the lower mean score on the first principal component (PC1); the labels are arbitrary identifiers and were not chosen to represent a cachexia versus non-cachexia distinction.
For descriptive purposes only, a two-sample Welch t-test was run independently for every filtered gene, comparing log2-CPM values between Cluster_B and Cluster_A. Nominal p-values were adjusted for multiple comparisons using the Benjamini–Hochberg (BH) false discovery rate (FDR) procedure. This screen is exploratory rather than confirmatory: it does not use a count-aware model, does not adjust for cancer type or other technical covariates, and its purpose in this pilot is limited to ranking candidate genes for later, more rigorous analysis.
Eight literature-informed gene panels relevant to cancer cachexia biology were defined before the analysis was run: proteolysis/atrophy (FBXO32, TRIM63, FOXO1, FOXO3, NFE2L2), inflammatory signaling (IL6, JAK1, JAK2, STAT3, NFKB1, RELA, TNF, IL1B), the myostatin/activin axis (MSTN, ACVR2A, ACVR2B, ACVR1B, SMAD2, SMAD3, INHBA), muscle structure and regeneration (MEF2C, PAX7, MYOD1, MYOG, DES, DMD, ACTN2), autophagy and cellular stress (ULK1, BECN1, ATG5, SQSTM1, MAP1LC3B, DDIT3), mitochondrial energy metabolism (PPARGC1A, TFAM, NDUFS1, COX4I1, ATP5F, CPT1B), the GDF15 axis (GDF15, GFRAL, RET), and the FoxP1 axis (FOXP1, MEF2C, HDAC4, HDAC5). For each panel, a per-sample score was calculated as the mean log2-CPM value across the panel genes that were present in the filtered matrix; a gene’s absence from the matrix was not treated as evidence that the corresponding pathway was inactive.
Two independent, functionally equivalent implementations of this
pipeline were produced for transparency and reproducibility: a Python
implementation (pilot_analysis.py and
summarize_results.py, using numpy, pandas, matplotlib,
seaborn, scipy, and scikit-learn) and this R Markdown implementation
(cachexia_rpubs_report.Rmd, using ggplot2, dplyr, tidyr,
and scales) suitable for direct rendering and public sharing on an R
Markdown–hosting platform. Both implementations operate on the same
filtered count matrix and produce the same cluster assignments,
gene-level contrasts, and figures reported here.
After filtering, 16,160 protein-coding genes remained across the 84 samples. Unsupervised clustering on the top 2,000 variable genes produced a clearly separable two-group structure along the first principal component (Figure 1). Cluster_A contained 59 samples (34 colorectal, 25 pancreatic; mean PC1 = -15.49, mean PC2 = 0.70), and Cluster_B contained 25 samples (12 colorectal, 13 pancreatic; mean PC1 = 36.55, mean PC2 = -1.64), summarized in Table 1.
Figure 1. Exploratory principal component analysis of the 84-sample human skeletal-muscle RNA-seq cohort. Point color denotes the unsupervised k-means cluster; point shape denotes the publicly recorded cancer type (colorectal or pancreatic).
| Cluster | Samples (n) | Colorectal | Pancreatic | Mean PC1 | Mean PC2 |
|---|---|---|---|---|---|
| Cluster_A | 59 | 34 | 25 | -15.487 | 0.696 |
| Cluster_B | 25 | 12 | 13 | 36.550 | -1.642 |
Because colorectal and pancreatic samples were not evenly distributed between the two clusters, cancer type, technical batch effects, or other unmeasured factors may all contribute to this separation. For that reason, Cluster_A and Cluster_B are treated strictly as descriptive labels and are not interpreted as cachexia and non-cachexia groups in this pilot.
The Welch t-test screen ranked genes by unadjusted p-value for the Cluster_B versus Cluster_A contrast on the log2-CPM scale. Table 2 lists the top 20 ranked genes. Several of the leading genes (IL1RAP, DOCK2, HCLS1, VAV3, CR1) are associated with leukocyte and immune-cell biology, which is consistent with a difference in immune-cell content or inflammatory activity between the two clusters rather than a change confined to myofiber-intrinsic pathways alone.
| Gene | log2FC (B vs A) | p-value | BH FDR |
|---|---|---|---|
| STK19 | 0.511 | 3.37e-14 | 5.44e-10 |
| IL1RAP | 1.091 | 2.63e-13 | 2.13e-09 |
| STAB2 | 1.470 | 1.42e-12 | 7.01e-09 |
| DAPP1 | 0.612 | 1.74e-12 | 7.01e-09 |
| ALPK1 | 0.923 | 2.25e-12 | 7.20e-09 |
| ADAM28 | 1.183 | 2.79e-12 | 7.20e-09 |
| CACNB4 | 1.567 | 3.53e-12 | 7.20e-09 |
| TBXAS1 | 0.914 | 3.56e-12 | 7.20e-09 |
| CR1 | 1.157 | 6.45e-12 | 1.11e-08 |
| IQGAP2 | 0.979 | 6.89e-12 | 1.11e-08 |
| DOCK2 | 1.062 | 1.50e-11 | 1.86e-08 |
| ABCC3 | 1.222 | 1.56e-11 | 1.86e-08 |
| TMED3 | 1.290 | 1.61e-11 | 1.86e-08 |
| SLC22A15 | 0.884 | 1.66e-11 | 1.86e-08 |
| HCLS1 | 0.960 | 1.76e-11 | 1.86e-08 |
| TTPAL | 0.552 | 1.85e-11 | 1.86e-08 |
| LACC1 | 0.786 | 2.00e-11 | 1.86e-08 |
| NRCAM | 1.319 | 2.16e-11 | 1.86e-08 |
| DGKI | 1.502 | 2.22e-11 | 1.86e-08 |
| VAV3 | 0.897 | 2.31e-11 | 1.86e-08 |
Mean panel scores for the eight pre-specified mechanism panels are shown in Figure 2 and summarized by cluster in Table 3. The inflammatory-signaling panel showed the largest relative separation between clusters (mean score 3.84 in Cluster_A versus 4.37 in Cluster_B), which aligns with the immune-related genes surfaced in the differential-expression screen. The GDF15 axis and myostatin/activin axis panels showed smaller but directionally consistent increases in Cluster_B, while the mitochondrial-energy panel showed a small decrease. The remaining panels (autophagy/stress, FoxP1 axis, muscle structure/regeneration, and proteolysis/atrophy) overlapped substantially between clusters, indicating that this pilot did not detect a broad, uniform shift across every candidate pathway.
Figure 2. Exploratory mechanism-panel scores (mean log2[CPM + 1] across genes present in each panel) for the eight pre-specified cachexia-related gene panels, grouped by exploratory cluster.
| Mechanism panel | Cluster_A mean | Cluster_B mean |
|---|---|---|
| Autophagy / cellular stress | 5.350 | 5.465 |
| FoxP1 axis | 7.008 | 7.118 |
| GDF15 axis | 1.197 | 1.340 |
| Inflammatory signaling | 3.841 | 4.374 |
| Mitochondrial energy metabolism | 4.995 | 4.888 |
| Muscle structure / regeneration | 7.692 | 7.753 |
| Myostatin / activin axis | 4.452 | 4.655 |
| Proteolysis / atrophy | 7.231 | 7.343 |
Bhatt et al. (2025) applied integrative non-negative matrix factorization to the full coding and non-coding transcriptome of this same 84-patient cohort and reported two highly coherent subtypes, one of which was clinically associated with high-grade weight loss, low muscle mass, selective atrophy of type IIA and IIX fibers, and shorter survival. The exploratory clusters produced in this pilot were derived independently, using only the public raw counts and a simple PCA/k-means workflow restricted to protein-coding genes, because the original per-patient subtype labels are not included in the public files. Any correspondence between Cluster_A/Cluster_B here and the published Subtype 1/Subtype 2 groups is therefore unverified and should not be assumed; confirming or refuting that correspondence would require either the original subtype assignments or a matched non-negative matrix factorization analysis, neither of which was attempted in this pilot.
The pattern of results in the Results section is best read as a short list of candidates for further work rather than as evidence that any single pathway causes muscle wasting in this cohort. Table 5 places the four most consistently implicated mechanism areas from this pilot alongside independent supporting evidence and the caution that evidence still demands.
| Mechanism area | Supporting evidence | Caution |
|---|---|---|
| GDF15-GFRAL appetite/energy axis | A phase 2 randomized trial of the anti-GDF15 antibody ponsegromab produced dose-dependent weight gain and symptom improvement in patients with cancer cachexia and elevated GDF15 (Groarke et al., 2024). | Weight and appetite gains do not, by themselves, demonstrate reduced tumor aggressiveness or longer survival. |
| Activin/myostatin (TGF-beta family) signaling | INHBA showed the strongest cluster contrast among the pre-specified candidates in this pilot; activin-family ligands are established regulators of muscle protein turnover. | Pathway inhibition can affect tissues beyond muscle, and mass gained is not equivalent to functional strength gained. |
| Inflammatory / IL-6-JAK-STAT3 signaling | TNF, IL1B, and IL6 were among the more strongly contrasted candidate genes, consistent with reviews linking cytokine signaling to autophagy-lysosome and ubiquitin-proteasome overactivation in cachectic muscle (Libramento et al., 2025). | Broad cytokine suppression could interfere with anti-tumor immune surveillance or infection control. |
| FoxP1-linked transcriptional program | A mouse C26 tumor model showed that skeletal-muscle FoxP1 upregulation was sufficient to induce cachexia-like features and that silencing FoxP1 attenuated wasting (Neyroud et al., 2021); FOXP1 was modestly but significantly higher in Cluster_B in this human pilot. | Preclinical rodent findings do not establish human efficacy, and transcriptional targets often have wide-ranging effects. |
| Proteolysis, autophagy, and regeneration programs | Core atrophy and regeneration genes (FBXO32, TRIM63, MSTN, PAX7, ACTN2) showed little to no difference between the exploratory clusters in this dataset. | Complete inhibition of protein turnover pathways can impair normal muscle maintenance; time-resolved designs are needed to separate harmful overactivation from necessary basal turnover. |
Public RNA-seq data of this kind are well suited to generating hypotheses: they can reveal correlated expression programs, candidate biomarkers, and molecular subgroups at a scale that would be difficult to assemble locally. They are not, by themselves, able to establish that a given gene causes muscle loss, to select a safe drug dose, to show that an intervention preserves muscle during active chemotherapy, or to predict benefit for an individual patient. A muscle-preserving intervention that looked promising by transcriptomic prioritization alone would still need to be evaluated on lean mass, muscle strength or physical function, appetite and cachexia symptoms, treatment completion, quality of life, tumor response, progression-free survival, overall survival, and adverse events; an agent that increased body weight while worsening tumor control would not meet the underlying clinical goal, however favorable the expression data looked at the outset.
This pilot has several limitations that bound how its results should be used. The analysis is observational and cross-sectional, capturing a single biopsy time point per patient rather than a treatment or longitudinal comparison. The exploratory clusters are unsupervised and were not validated against the original clinical subtype labels, a phenotype table, or an outcome measure, so they may reflect cancer type, sequencing batch, or other unmeasured technical variation rather than a cachexia-specific biological axis. The differential-expression screen used simple Welch t-tests on log2-CPM values rather than a count-aware negative-binomial model such as DESeq2 or edgeR, and it did not adjust for cancer type or other covariates, which can inflate or distort apparent effect sizes. The cohort itself is limited to colorectal and pancreatic cancer, so these findings should not be extrapolated to other tumor types without direct testing. Finally, nothing in this report establishes causation, drug response, or any form of patient-specific treatment guidance; it is intended strictly as a hypothesis-generating pilot.
A reasonable next stage for this line of work would proceed in a fixed order. First, obtain the original investigators’ clinical subtype or phenotype table, where available, and reanalyze the raw counts with a count-aware differential-expression method such as DESeq2 or edgeR, adjusting explicitly for cancer type and technical covariates. Second, attempt to replicate any candidate pathway in an independent human muscle cohort or in CT-derived body-composition data, where lean-mass and sarcopenia measures can serve as an external phenotype. Third, test the most promising candidates in cancer–muscle co-culture systems and in animal models, measuring both muscle function and tumor behavior together rather than muscle mass alone. Only after those steps have produced consistent, mechanistically grounded evidence should a clinical protocol be designed, and any such protocol would require review by qualified cachexia investigators together with the appropriate ethics and regulatory bodies.
Reanalysis of a public 84-sample human skeletal-muscle RNA-seq cohort produced a clear but arbitrary two-cluster structure and pointed toward inflammatory signaling, activin/GDF15-related genes, and the FoxP1 axis as the mechanism areas most consistently associated with that structure, while the classical proteolysis and regeneration genes showed little separation. These findings are consistent with, and were partly anticipated by, independent lines of evidence: a completed phase 2 trial supporting GDF15 as a druggable driver of cachexia symptoms (Groarke et al., 2024), a preclinical mouse model implicating FoxP1 in wasting (Neyroud et al., 2021), and reviews describing inflammatory activation of proteolytic pathways in cachectic muscle (Libramento et al., 2025). None of this, however, substitutes for the confirmatory, covariate-adjusted, and experimentally validated work that a genuine therapeutic program in cancer cachexia requires. This report is offered as a transparent, reproducible pilot intended to inform the next stage of work rather than to conclude it.
The raw count matrix and sample metadata analyzed in this report are
publicly available from NCBI GEO under accessions GSE254877 and
GSE254878 (part of SuperSeries GSE292053), and the related mouse dataset
is available under accession GSE153068. The analysis pipeline is
implemented in two equivalent forms: a Python script pair
(pilot_analysis.py, summarize_results.py) and
this R Markdown report (cachexia_rpubs_report.Rmd),
suitable for rendering on an R Markdown–hosting platform for public,
reproducible sharing. Full gene-level results, including the complete
candidate-gene contrast table, per-sample mechanism-panel scores, and
the exploratory cluster summary, are provided as an accompanying
Microsoft Excel workbook,
Kenche_cachexia_supplementary_data.xlsx, alongside the two
figure files reproduced as Figures 1 and 2 in this manuscript.
Bhatt, B. J., Ghosh, S., Mazurak, V., Brun, A. Q., Bathe, O., Baracos, V. E., & Damaraju, S. (2025). Molecular subtypes of human skeletal muscle in cancer cachexia. Nature, 646(8086), 973–982. https://doi.org/10.1038/s41586-025-09502-0
Groarke, J. D., Crawford, J., Collins, S. M., Lubaczewski, S., Roeland, E. J., Naito, T., Takayama, K., Asmis, T., Hendifar, A. E., Fallon, M., Dunne, R. F., Karahanoglu, I., Northcott, C. A., Harrington, M. A., Rossulek, M., Qiu, R., & Saxena, A. R. (2024). Ponsegromab for the treatment of cancer cachexia. New England Journal of Medicine, 391(24), 2291–2303. https://doi.org/10.1056/NEJMoa2409515
Libramento, Z. P., Tichy, L., & Parry, T. L. (2025). Muscle wasting in cancer cachexia: Mechanisms and the role of exercise. Experimental Physiology. Advance online publication. https://doi.org/10.1113/EP092544
McKinney, W. (2010). Data structures for statistical computing in Python. Proceedings of the 9th Python in Science Conference, 56–61. https://doi.org/10.25080/Majora-92bf1922-00a
National Center for Biotechnology Information. (2025). Gene Expression Omnibus accession GSE254878: Molecular subtypes of human skeletal muscle in cancer cachexia [Data set]. NCBI Gene Expression Omnibus. https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE254878
National Center for Biotechnology Information. (2021). Gene Expression Omnibus accession GSE153068: Cancer-induced FoxP1 and skeletal-muscle wasting in a C26 tumor model [Data set]. NCBI Gene Expression Omnibus. https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE153068
Neyroud, D., Nosacka, R. L., Callaway, C. S., Trevino, J. G., Hu, H., Judge, S. M., & Judge, A. R. (2021). FoxP1 is a transcriptional repressor associated with cancer cachexia that induces skeletal muscle wasting and weakness. Journal of Cachexia, Sarcopenia and Muscle, 12(2), 421–442. https://doi.org/10.1002/jcsm.12666
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., & Duchesnay, É. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, 2825–2830.
Wickham, H. (2016). ggplot2: Elegant graphics for data analysis. Springer-Verlag. https://doi.org/10.1007/978-3-319-24277-4
Report generated with R Markdown · Session info available on request · September 2026