Transformer-Based Operon Prediction Using Textual Representations of Gene Pairs
Clicks: 6
ID: 314591
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This
article has not been analysed, so there is no overall score —
reader engagement is measured and shown alongside.
Reader Engagement
Popular Article
1.5
/100
6 views
1 readers
AI Quality Assessment
Not analyzed
Readership in this journal
PopularRanked #31 of 106 articles by views in Bioinformatics advances
Most read
Least read
Bar heights use a square-root scale.
Mint this article as an NFT
Not yet mintedCreate a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.
5
SUSD
one-off · no wallet required
Abstract
Abstract Motivation Operons are fundamental units of gene regulation in bacteria and can provide valuable insights into genome organization, co-expression, and functional relationships between genes. Computational prediction of operons can support downstream analyses such as pathway reconstruction, comparative genomics, and gene function inference. However, many existing tools rely on rigid feature engineering or curated interaction networks, limiting scalability and applicability to poorly annotated genomes. Results We propose a transformer-based approach that reformulates operon prediction as a binary text classification task over adjacent gene pairs. By serializing genomic features, including gene orientation, intergenic distance, GC content, functional annotations, protein families, and conservation, into natural language descriptions, we enable pre-trained language models to perform operon classification using flexible, widely available inputs. A RoBERTa-based model achieves competitive predictive performance under multiple evaluation settings, including leave-one-species-out analysis over six bacterial genomes and benchmark comparisons on standard datasets. Through ablation and inference-time resilience analyses, we demonstrate that sequence-derived and annotation-based features are sufficient for competitive performance, and that performance remains stable when selected features are removed at test time.
| Reference Key |
openalex_W7161980222
Use this key to autocite in the manuscript while using
SciMatic Manuscript Manager or Thesis Manager
|
|---|---|
| Authors | Rida Assaf, Basel Fakhri |
| Journal | Bioinformatics advances |
| Year | 2026 |
| DOI |
10.1093/bioadv/vbag140
|
| URL | |
| Keywords | Keywords not found |
Citations
No citations found. To add a citation, contact the admin at info@scimatic.org
Comments
No comments yet. Be the first to comment on this article.