Bayesian Hyperparameter Optimization Improves scGPT Fine-Tuning for Single-Cell Multi-Omics Integration

Clicks: 1
ID: 317263
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This article has not been analysed, so there is no overall score — reader engagement is measured and shown alongside.
AI Quality Assessment
Not analyzed
Readership in this journal

Ranked #747 of 825 articles by views in BMC Bioinformatics

Most read Least read

Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 825 in total.

Mint this article as an NFT
Not yet minted

Create a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.

5 SUSD one-off · no wallet required
Abstract
MOTIVATION: Foundation models such as scGPT have demonstrated strong potential for single-cell multi-omics integration; however, their downstream performance is highly sensitive to hyperparameter selection. Manual fine-tuning remains computationally expensive, dataset-dependent, and often irreproducible. Despite the increasing adoption of foundation models in single-cell analysis, systematic strategies for robust hyperparameter optimization remain underexplored. RESULTS: We developed a Bayesian optimization framework based on Tree-structured Parzen Estimators (TPE) for automated fine-tuning of scGPT and evaluated its performance on two benchmark bone marrow mononuclear cell (BMMC) multi-omics datasets, including CITE-seq and GSE194122 datasets. Across datasets, Bayesian optimization consistently improved biological conservation and batch integration metrics compared with default scGPT configurations. On the original BMMC benchmark, optimization improved AvgBIO from 0.59 to 0.67 and PCR from 0.33 to 0.52. On the GSE194122 dataset, the default configuration exhibited unstable convergence and weak biological preservation (AvgBIO = 0.19; ARI = 0.007), whereas Bayesian optimization substantially improved integration performance (AvgBIO = 0.60; ARI = 0.63) while reducing validation loss from 137 to 47.1. These findings demonstrate substantial dataset-specific sensitivity of scGPT fine-tuning and highlight the importance of automated optimization for stable deployment across heterogeneous multi-omics datasets. CONCLUSION: Our study demonstrates that Bayesian optimization provides an effective and reproducible strategy for stabilizing scGPT fine-tuning across diverse single-cell multi-omics datasets. Rather than introducing a new integration architecture, this work emphasizes the importance of systematic optimization for improving robustness and reproducibility of foundation-model applications in computational biology. AVAILABILITY: Our model and dataset are freely available at: https://github.com/daren642/scGPT_multiomic_tuning.
Reference Key
openalex_W7164665516 Use this key to autocite in the manuscript while using SciMatic Manuscript Manager or Thesis Manager
Authors Tay Dk, Nguyen Quoc Khanh Le, Matthew Chin Heng Chua
Journal BMC Bioinformatics
Year 2026
DOI
10.1093/bioinformatics/btag374
URL
Keywords Keywords not found

Citations

No citations found. To add a citation, contact the admin at info@scimatic.org

No comments yet. Be the first to comment on this article.