scGenoByte: a GenoByte embedding transformer with biological priors for cell type annotation

Clicks: 24
ID: 319874
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This article has not been analysed, so there is no overall score — reader engagement is measured and shown alongside.
AI Quality Assessment
Not analyzed
Readership in this journal
Emerging

Ranked #26 of 58 articles by views in Briefings in bioinformatics

Most read Least read

Bar heights use a square-root scale.

Mint this article as an NFT
Not yet minted

Create a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.

5 SUSD one-off · no wallet required
Abstract
Abstract Effective cell representation learning is crucial for accurate cell annotation and the deciphering of cellular heterogeneity in single-cell RNA sequencing (scRNA-seq) analysis. Current foundation models have achieved superior performance compared with traditional methods. However, due to data sparsity and the complexity of model, existing methods often compromise by selecting highly variable genes or filtering for nonzero expressions, which discard potentially significant genes. Thus, modeling the complete transcriptome for cell representation remains computationally challenging; we present scGenoByte, a unified framework designed to enhance cell representation learning through biologically informed full-gene modeling. To enable efficient modeling of the full transcriptome, we design GenoBytes, biologically coherent units that are constructed by leveraging biological priors in terms of protein–protein interaction network and gene paralogy network. Furthermore, considering that the information of protein and pathway is critical for analyzing cell functions and representation, scGenoByte encapsulates biological priors by harmonizing GenoByte embeddings with protein representations and leveraging an auxiliary task of pathway activity prediction to impose pathway-guided regularization. Extensive results on eight datasets have shown that scGenoByte achieves better performance than competing methods, which confirms the efficacy of combining full-gene context with biological priors.
Reference Key
openalex_W7167483721 Use this key to autocite in the manuscript while using SciMatic Manuscript Manager or Thesis Manager
Authors Jiongsen Yao, Yan Xu, Jinjin Ma, Wenjun Shen, Si Wu
Journal Briefings in bioinformatics
Year 2026
DOI
10.1093/bib/bbag369
URL
Keywords Keywords not found

Citations

No citations found. To add a citation, contact the admin at info@scimatic.org

No comments yet. Be the first to comment on this article.