pLM-SAV: A Δ-Embedding Approach for Predicting Pathogenic Single Amino Acid Variants

Clicks: 4
ID: 321530
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This article has not been analysed, so there is no overall score — reader engagement is measured and shown alongside.
AI Quality Assessment
Not analyzed
Readership in this journal

Ranked #26 of 100 articles by views in Bioinformatics advances

Most read Least read

Bar heights use a square-root scale.

Mint this article as an NFT
Not yet minted

Create a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.

5 SUSD one-off · no wallet required
Abstract
Abstract Motivation Predicting whether single amino acid variants (SAVs) in proteins lead to pathogenic outcomes is a critical challenge in molecular biology and precision medicine. Experimental determination of all possible mutation effects is infeasible, and while state-of-the-art tools such as AlphaMissense show promise, their diagnostic performance is insufficient and they are often difficult to run locally. Results We developed pLM-SAV, a simple yet effective predictor that leverages protein language models (pLMs). Δ-embeddings, computed as the difference between wild-type and mutant sequence embeddings, are used as input for a convolutional neural network. We trained our model on a well-characterized, labeled set of Eff10k and evaluated it on a non-homologous subset of ClinVar data. This approach performs exceptionally well on the Eff10k test folds and reasonably on ClinVar test sets. On ambiguity-defined subsets, pLM-SAV provides complementary predictions to AlphaMissense and REVEL, although these methods retain higher overall performance on broad ClinVar benchmarks. Our results show that a compact predictor trained on labeled variant-effect data can provide useful predictive performance across multiple benchmarks. Unlike previous methods such as VESPA, pLM-SAV uses no handcrafted features or substitution matrices, relying solely on pLM-derived representations. Δ-embedding-based features may provide useful additional signals for future mutation-effect predictors or integrative models. Availability and implementation The code is available at https://doi.org/10.5281/zenodo.15502498.
Reference Key
openalex_W4410798438 Use this key to autocite in the manuscript while using SciMatic Manuscript Manager or Thesis Manager
Authors Orsolya Gereben, Hedvig Tordai, Lana Khamisi, Erda Qorri, Tamás Hegedűs
Journal Bioinformatics advances
Year 2026
DOI
10.1093/bioadv/vbag195
URL
Keywords Keywords not found

Citations

No citations found. To add a citation, contact the admin at info@scimatic.org

No comments yet. Be the first to comment on this article.