Benchmarking the Impact of Data Leakage on the Performance of Knowledge Graph Embedding Models for Biomedical Link Prediction
Clicks: 3
ID: 325411
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This
article has not been analysed, so there is no overall score —
reader engagement is measured and shown alongside.
Reader Engagement
Emerging Content
0.6
/100
3 views
2 readers
AI Quality Assessment
Not analyzed
Readership in this journal
EmergingRanked #495 of 829 articles by views in BMC Bioinformatics
Most read
Least read
Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 829 in total.
Mint this article as an NFT
Not yet mintedCreate a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.
5
SUSD
one-off · no wallet required
Abstract
Abstract Motivation Knowledge Graphs (KGs) organize complex biomedical knowledge into structured representations of entities and relations. Knowledge Graph Embedding (KGE) models learn compact representations of KGs, and are widely applied for biomedical link prediction. Despite extensive work on KGE models, current evaluations often overlook the issue of data leakage, which can artificially inflate performance and undermine benchmark validity. Data leakage can arise when (1) there is redundancy between training and test sets, (2) the model leverages illegitimate features, or (3) the test set does not accurately reflect real-world inference. Results We assess the impact of data leakage on KGE-based link prediction across three biomedical knowledge graphs, using decoder-only and GNN-based models. We first demonstrate the impact of train-test redundancies and implement a systematic procedure to detect and remove them. Using permutation experiments, we also investigate whether node degree acts as an illegitimate predictive feature, and find no evidence that predictions are driven by degree alone. Finally, we evaluate how well common test set sampling strategies reflect real-world inference in drug repurposing. We compare random and cold-start data splits with an independent test set from Orphanet, and observe a substantial performance drop on the latter, indicating that current benchmarking practices may overestimate how well KGE models generalize to practical applications. Overall, our findings highlight the importance of rigorous benchmark design and careful evaluation of the generalization ability of KGE models for biomedical link prediction. Availability and Implementation Code and results are openly available on GitHub at https://github.com/galadrielbriere/data_leakage_kge_benchmark.git and archived on Zenodo at https://doi.org/10.5281/zenodo.21885112. Supplementary Information Supplementary data are available.
| Reference Key |
openalex_W7203832691
Use this key to autocite in the manuscript while using
SciMatic Manuscript Manager or Thesis Manager
|
|---|---|
| Authors | Galadriel Bri`ere, Thomas STOSSKOPF, Benjamin Loire, Anaı̈s Baudot |
| Journal | BMC Bioinformatics |
| Year | 2026 |
| DOI |
10.1093/bioinformatics/btag608
|
| URL | |
| Keywords | Keywords not found |
Citations
No citations found. To add a citation, contact the admin at info@scimatic.org
Comments
No comments yet. Be the first to comment on this article.