The Physicochemical Basis of Protein Evolution: Property-Informed Evolutionary Models (PRIME)

Clicks: 6
ID: 324973
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This article has not been analysed, so there is no overall score — reader engagement is measured and shown alongside.
AI Quality Assessment
Not analyzed
Readership in this journal

Ranked #70 of 242 articles by views in molecular biology and evolution

Most read Least read

Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 242 in total.

Mint this article as an NFT
Not yet minted

Create a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.

5 SUSD one-off · no wallet required
Abstract
Abstract Standard probabilistic models of coding sequence evolution effectively identify where and when selection acts but remain agnostic to the mechanistic realization of these forces. We introduce PRIME (PRoperty Informed Models of Evolution), a framework of codon-level maximum likelihood methods—including global (G-PRIME), episodic (E-PRIME), and site-specific (S-PRIME) implementations—that explicitly model amino acid exchangeability as a function of physicochemical properties. By parameterizing attributes such as molecular volume, hydropathy, and secondary structure propensities, PRIME \revised{aims to resolve} the biophysical basis of selective constraint across both the sequence and the phylogeny. At the site level, S-PRIME leverages an explicit biophysical taxonomy to categorize residues as conserved, neutral, or changing for specific properties, resolving selective signals that \revised{are missed by} traditional rate-based metrics. Our analysis of a benchmark of 24 diverse datasets and a genome-wide screen of 18,944 mammalian genes demonstrates that \finalrevised{consideration of biophysical realism can yield} substantial improvements in model fit, acting synergistically with rate variation to explain complex evolutionary patterns. \finalrevised{We find that physicochemical constraints at individual sites can be reliably detected in datasets with sufficient information redundancy} (substitutions per unique amino acid; AUC = 0.91), with sensitivity exceeding 90 % in data-rich alignments. E-PRIME reveals a \finalrevised{distinct hierarchy in biophysical constraints}: while core packing and beta-sheet scaffolds are rigidly conserved, alpha-helix propensity and surface electrostatics serve as the primary substrates for adaptive tuning. Furthermore, PRIME importance weights align with aspects of the primary semantic axes of deep learning representations (ESM-2) and capture key features of experimental fitness landscapes. By transforming abstract evolutionary rates into interpretable biophysical rules, PRIME provides a useful framework for characterizing the mechanistic drivers of protein diversity.
Reference Key
openalex_W7202367843 Use this key to autocite in the manuscript while using SciMatic Manuscript Manager or Thesis Manager
Authors Hannah Kim, Konrad Scheffler, Anton Nekrutenko, Darren P. Martin, Steven Weaver, Ben Murrell, Sergei L. Kosakovsky Pond
Journal molecular biology and evolution
Year 2026
DOI
10.1093/molbev/msag200
URL
Keywords Keywords not found

Citations

No citations found. To add a citation, contact the admin at info@scimatic.org

No comments yet. Be the first to comment on this article.