Cross-Domain Transfer Learning from Peptides to Metabolites Using a Multi-Property Fine-Tuned LLM

Clicks: 1
ID: 319688
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This article has not been analysed, so there is no overall score — reader engagement is measured and shown alongside.
AI Quality Assessment
Not analyzed
Readership in this journal

Ranked #816 of 829 articles by views in BMC Bioinformatics

Most read Least read

Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 829 in total.

Mint this article as an NFT
Not yet minted

Create a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.

5 SUSD one-off · no wallet required
Abstract
Abstract Motivation Accurate liquid chromatography retention time (RT) prediction is a critical component of compound identification in metabolomics and lipidomics. However, existing RT prediction approaches are often limited by the scarcity of experimental RT measurements for many molecular classes, restricting model generalization and the construction of comprehensive RT libraries. Transfer learning from data-rich chemical domains offers a potential strategy to overcome these limitations, but its effectiveness for metabolite RT prediction remains insufficiently explored. Results We developed a transfer learning framework based on ChemBERTa that leverages large peptide datasets to improve metabolite RT prediction under data-sparse conditions. A peptide-pretrained model was trained using a multi-task objective that jointly predicted RT and seven RDKit-derived molecular descriptors. Compared with an RT-only model, the multi-task approach learned more robust chemical representations and demonstrated superior generalization to metabolites, achieving a median test R² of 0.842 versus 0.820. When transferred to metabolite RT prediction, the multi-task pretrained model substantially outperformed models trained from scratch at low-data regimes. Using only 3% of metabolite training data (2,129 compounds), transfer learning achieved a median test R² of 0.322 compared with 0.216 for the baseline model, while reducing MAE from 131.7 to 114.9. Significant improvements were also observed at 5% and 10% training fractions, with benefits gradually diminishing as larger metabolite datasets became available. In contrast, a peptide-pretrained single-task RT model showed performance comparable to the baseline, indicating that the observed gains arise primarily from multi-task molecular property learning rather than peptide pretraining alone. These findings demonstrate that multi-task transfer learning provides an effective and scalable strategy for improving RT prediction in metabolomics, particularly when experimental training data are limited. Availability Freely available on https://github.com/uchealex/CHEMBEDDING Supplementary information Supplementary data are available at Bioinformatics online.
Reference Key
openalex_W7167420113 Use this key to autocite in the manuscript while using SciMatic Manuscript Manager or Thesis Manager
Authors Uchenna Alex Anyaegbunam, David Teschner, Thierry Schmidlin, Andreas Hildebrandt, Johannes U Mayer, Maximilian Sprang, Miguel A. Andrade‐Navarro
Journal BMC Bioinformatics
Year 2026
DOI
10.1093/bioinformatics/btag493
URL
Keywords Keywords not found

Citations

No citations found. To add a citation, contact the admin at info@scimatic.org

No comments yet. Be the first to comment on this article.