DeepGVS: a bimodal deep learning framework integrating coding-sequence and protein-structural representations for virulence factor prediction

Clicks: 5
ID: 329449
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This article has not been analysed, so there is no overall score — reader engagement is measured and shown alongside.
AI Quality Assessment
Not analyzed
Readership in this journal
Emerging

Ranked #874 of 895 articles by views in BMC Bioinformatics

Most read Least read

Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 895 in total.

Mint this article as an NFT
Not yet minted

Create a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.

5 SUSD one-off · no wallet required
Abstract
MOTIVATION: Virulence factors (VFs) mediate host adhesion, invasion, immune evasion and toxin-mediated damage, making accurate VF prediction important for understanding bacterial pathogenesis and antimicrobial intervention. Existing predictors mainly use one-dimensional (1D) protein sequences, overlooking complementary coding DNA sequence (CDS)-level information and three-dimensional (3D) structural topology. RESULTS: We propose DeepGVS, a bimodal deep learning framework integrating CDS-derived features with protein sequence and structural representations for VF prediction. DeepGVS extracts multi-scale sequence-composition features from CDS. Concurrently, it employs a parallel graph attention network (GAT) and bidirectional Mamba (Bi-Mamba) architecture to process ESMFold-predicted structures and residue representations, capturing spatial and long-range dependencies. A neural additive model (NAM) functions as a meta-learner to integrate base-classifier predictions. DeepGVS was evaluated on the unchanged accession-level independent test set of Dataset_B and achieved an accuracy of 87.50%, corresponding to absolute improvements of 6.30, 2.60 and 1.40 percentage points over the published benchmark values of DeepVF, GTAE-VF and PLMVF, respectively. The incremental benefit of bimodal integration was dataset- and metric-dependent, and additional taxonomic analyses identified taxonomy as a potential confounding factor that does not fully reproduce the performance of the complete model. AVAILABILITY AND IMPLEMENTATION: Source code, datasets and pretrained models are available at https://github.com/guoguo26/DeepGVS. The archived code version used for the reported experiments is available at https://doi.org/10.5281/zenodo.21756890. SUPPLEMENTARY INFORMATION: Supplementary data are available online.
Reference Key
openalex_W7172200909 Use this key to autocite in the manuscript while using SciMatic Manuscript Manager or Thesis Manager
Authors guoguo26, Tingting Zou, Zhenyuan Sun, Yuming Zhao, Guohua Wang
Journal BMC Bioinformatics
Year 2026
DOI
10.1093/bioinformatics/btag698
URL
Keywords Keywords not found

Citations

No citations found. To add a citation, contact the admin at info@scimatic.org

No comments yet. Be the first to comment on this article.