CrossMol: Cross-Modal Mask-Predict Pre-training For 3D Molecular Data

Clicks: 3
ID: 329535
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This article has not been analysed, so there is no overall score — reader engagement is measured and shown alongside.
AI Quality Assessment
Not analyzed
Readership in this journal
Emerging

Ranked #881 of 895 articles by views in BMC Bioinformatics

Most read Least read

Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 895 in total.

Mint this article as an NFT
Not yet minted

Create a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.

5 SUSD one-off · no wallet required
Abstract
MOTIVATION: Self-supervised pre-training models for molecular data have demonstrated notable results across many downstream tasks. The inherent multimodal properties of molecules have also motivated efforts to capture information from different modalities. However, current multimodal molecular pre-training models usually treat these modalities as equal and independent, despite differences in their information content. Three-dimensional (3D) molecular structures generally contain finer-grained information than the simplified molecular-input line-entry system (SMILES), which primarily captures higher-level semantic information, such as molecular topology. RESULTS: We designed a cross-modal mask-predict pre-training model, CrossMol, to capture semantic associations between modalities with unequal information volumes. The model completes missing 3D structure using higher-level semantic information from another modality, such as SMILES. This allows it to learn cross-modal associations and better understand fine-grained 3D structural information. We also introduce a reweighted distance prediction loss to improve the modelling of short-range structural information. Experiments show that CrossMol achieves large performance gains on multiple downstream molecular tasks, attaining state-of-the-art results. AVAILABILITY AND IMPLEMENTATION: The source code and training data are available at https://github.com/zhengkangjie/crossmol. The code is archived at https://doi.org/10.5281/zenodo.19653204. SUPPLEMENTARY INFORMATION: Supplementary material contains the task-specific hyperparameters.
Reference Key
openalex_W7214113766 Use this key to autocite in the manuscript while using SciMatic Manuscript Manager or Thesis Manager
Authors Kangjie Zheng, Junwei Yang, Siyu Long, Wei Ju, Wei-Ying Ma, Hao Zhou, Ming Zhang
Journal BMC Bioinformatics
Year 2026
DOI
10.1093/bioinformatics/btag702
URL
Keywords Keywords not found

Citations

No citations found. To add a citation, contact the admin at info@scimatic.org

No comments yet. Be the first to comment on this article.