Transformer-Based Operon Prediction Using Textual Representations of Gene Pairs

Clicks: 6
ID: 314591
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This article has not been analysed, so there is no overall score — reader engagement is measured and shown alongside.
AI Quality Assessment
Not analyzed
Readership in this journal
Popular

Ranked #31 of 106 articles by views in Bioinformatics advances

Most read Least read

Bar heights use a square-root scale.

Mint this article as an NFT
Not yet minted

Create a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.

5 SUSD one-off · no wallet required
Abstract
Abstract Motivation Operons are fundamental units of gene regulation in bacteria and can provide valuable insights into genome organization, co-expression, and functional relationships between genes. Computational prediction of operons can support downstream analyses such as pathway reconstruction, comparative genomics, and gene function inference. However, many existing tools rely on rigid feature engineering or curated interaction networks, limiting scalability and applicability to poorly annotated genomes. Results We propose a transformer-based approach that reformulates operon prediction as a binary text classification task over adjacent gene pairs. By serializing genomic features, including gene orientation, intergenic distance, GC content, functional annotations, protein families, and conservation, into natural language descriptions, we enable pre-trained language models to perform operon classification using flexible, widely available inputs. A RoBERTa-based model achieves competitive predictive performance under multiple evaluation settings, including leave-one-species-out analysis over six bacterial genomes and benchmark comparisons on standard datasets. Through ablation and inference-time resilience analyses, we demonstrate that sequence-derived and annotation-based features are sufficient for competitive performance, and that performance remains stable when selected features are removed at test time.
Reference Key
openalex_W7161980222 Use this key to autocite in the manuscript while using SciMatic Manuscript Manager or Thesis Manager
Authors Rida Assaf, Basel Fakhri
Journal Bioinformatics advances
Year 2026
DOI
10.1093/bioadv/vbag140
URL
Keywords Keywords not found

Citations

No citations found. To add a citation, contact the admin at info@scimatic.org

No comments yet. Be the first to comment on this article.