Enhancing Feature Selection for Ordinal Outcomes Using Resampling-Based Sparse Linear Discriminant Analysis

Clicks: 1
ID: 321046
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This article has not been analysed, so there is no overall score — reader engagement is measured and shown alongside.
AI Quality Assessment
Not analyzed
Readership in this journal

Ranked #91 of 103 articles by views in Bioinformatics advances

Most read Least read

Bar heights use a square-root scale.

Mint this article as an NFT
Not yet minted

Create a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.

5 SUSD one-off · no wallet required
Abstract
Abstract Motivation High-dimensional biomedical datasets with ordinal outcomes—such as cancer stages or treatment responses—pose significant challenges for feature selection due to strong predictor correlations and limited sample sizes. Sparse Linear Discriminant Analysis (sLDA) is widely used for simultaneous classification and feature selection. However, concerns about model stability and reproducible feature selection persist, particularly in the presence of pronounced collinearity inherent in biomedical data. Consequently, direct application of sLDA often fails to capture a reproducible set of biologically coordinated markers, resulting in signatures that lack robustness and interpretability. Results We propose a resampling-based ensemble sLDA framework that integrates bootstrapping and subsampling to improve the stability of feature selection. By aggregating results across multiple resampled datasets, the method identifies features based on Variable Inclusion Probability (VIP) rather than relying on coefficients from standard sLDA. Compared with standard sLDA, this ensemble strategy reduces sensitivity to data perturbation and improves the stability and reproducibility of selected feature sets. Simulation studies demonstrate that the proposed ensemble framework achieves more accurate and consistent recovery of ground-truth predictors compared with the standard (non-resampled) sLDA. Applications to kidney renal papillary cell carcinoma (KIRP) staging and glioma grading datasets further suggest that this framework can improve predictive performance and identify biologically interpretable and reproducible feature sets, highlighting its potential utility for reliable biomarker discovery in precision medicine. Availability and Implementation The source code is available via GitHub at https://github.com/ryan-wng/RE-sLDA
Reference Key
openalex_W7168297007 Use this key to autocite in the manuscript while using SciMatic Manuscript Manager or Thesis Manager
Authors Yin Liu, Ryan Wang, Dong Si
Journal Bioinformatics advances
Year 2026
DOI
10.1093/bioadv/vbag196
URL
Keywords Keywords not found

Citations

No citations found. To add a citation, contact the admin at info@scimatic.org

No comments yet. Be the first to comment on this article.