EDEL: Enhancing Dense Retrievers for Curation of Biomedical Knowledge Bases
Clicks: 1
ID: 319450
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This
article has not been analysed, so there is no overall score —
reader engagement is measured and shown alongside.
Reader Engagement
0.0
/100
1 views
0 readers
AI Quality Assessment
Not analyzed
Readership in this journal
Ranked #811 of 829 articles by views in BMC Bioinformatics
Most read
Least read
Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 829 in total.
Mint this article as an NFT
Not yet mintedCreate a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.
5
SUSD
one-off · no wallet required
Abstract
Abstract Motivation Retrieval of relevant papers from the literature is the first step in curating high-quality biomedical knowledge bases. While BM25 has long been the method of choice, dense retrieval models achieved improved accuracy by embedding queries and documents into dense vector representations. Existing knowledge bases provide a natural source for deriving query-document pairs for training such models. Current training approaches, however, do not take into account that some evidence described by knowledge base entries may only be partially expressed in document abstracts, while the full evidence is often contained in inaccessible full texts, introducing noise into binary relevance labels. In addition, existing approaches only make limited use of the knowledge base structure for selecting negative samples during training. Results We propose EDEL, a novel dense bi-encoder for biomedical knowledge base curation to enable curators to find relevant papers for annotation faster. It introduces a loss function using graded relevance scores instead of binary labels to facilitate learning from partially grounded examples, together with a structured sampling strategy that exposes the model to diverse and hard negative examples during training. We evaluate EDEL’s performance in two curation settings, namely precision oncology (on CIViC and OncoKB) and post-translational modifications (on UniProt). EDEL outperforms other state-of-the-art models in NDCG@10 by 1.5 and 3.4 percentage points, respectively. Ablation studies show the effectiveness of both innovations. These results indicate that EDEL can substantially improve literature retrieval for biomedical knowledge base curation. Availability and implementation Code to reproduce our results is available at: https://github.com/WangXII/edel_repo. Supplementary information Supplementary data is attached to this manuscript.
| Reference Key |
openalex_W7167090180
Use this key to autocite in the manuscript while using
SciMatic Manuscript Manager or Thesis Manager
|
|---|---|
| Authors | Xing David Wang, Ulf Leser |
| Journal | BMC Bioinformatics |
| Year | 2026 |
| DOI |
10.1093/bioinformatics/btag490
|
| URL | |
| Keywords | Keywords not found |
Citations
No citations found. To add a citation, contact the admin at info@scimatic.org
Comments
No comments yet. Be the first to comment on this article.