Robust prioritization of genomic features with stability selection
Clicks: 1
ID: 317654
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This
article has not been analysed, so there is no overall score —
reader engagement is measured and shown alongside.
Reader Engagement
0.0
/100
1 views
0 readers
AI Quality Assessment
Not analyzed
Readership in this journal
Ranked #811 of 829 articles by views in BMC Bioinformatics
Most read
Least read
Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 829 in total.
Mint this article as an NFT
Not yet mintedCreate a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.
5
SUSD
one-off · no wallet required
Abstract
Abstract Motivation The heterogeneity of complex diseases including cancer leads to heavy-tailed distributions in the disease traits. In such settings, non-robust variable selection methods are inherently susceptible to data contamination and can yield unstable or misleading results. This vulnerability becomes more severe for recently proposed approaches that introduce pseudo-features as negative controls, as these methods further amplify the curse of dimensionality by expanding the genotype matrix in the presence of outliers and high-dimensional genomic features. Results We develop a robust variable selection framework with stability selection to prioritize genomic features in the presence of contamination. In contrast to existing approaches that rely on pseudo-features for error control, the proposed method achieves double robustness. First, it adopts least absolute deviation (LAD) LASSO to ensure robustness against outliers and heavy-tailed errors in disease traits. Second, it avoids augmenting the genotype matrix with pseudo-features, thereby mitigating the curse of dimensionality that is particularly problematic in high-dimensional genomic data. The proposed method has been extensively evaluated in simulation studies to demonstrate its effectiveness over multiple competing methods for variable selection. In addition, we have applied the proposed method and competing approaches to two real-data case studies: the The Cancer Genome Atlas (TCGA) Skin Cutaneous Melanoma (SKCM) dataset and an eQTL dataset. The results demonstrate that the proposed method achieves superior performance by identifying genomic features with higher reproducibility. Availability and implementation The source code for implementing the proposed methods is publicly available at https://github.com/cenwu/RSS with an archival DOI https://doi.org/10.6084/m9.figshare.32306883.
| Reference Key |
openalex_W7165030377
Use this key to autocite in the manuscript while using
SciMatic Manuscript Manager or Thesis Manager
|
|---|---|
| Authors | Gongshun Yang, Xi Lu, C WU |
| Journal | BMC Bioinformatics |
| Year | 2026 |
| DOI |
10.1093/bioinformatics/btag398
|
| URL | |
| Keywords | Keywords not found |
Citations
No citations found. To add a citation, contact the admin at info@scimatic.org
Comments
No comments yet. Be the first to comment on this article.