GLiNER-BioMed: A Suite of Efficient Models for Open Biomedical Named Entity Recognition

Clicks: 1
ID: 314824
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This article has not been analysed, so there is no overall score — reader engagement is measured and shown alongside.
AI Quality Assessment
Not analyzed
Readership in this journal

Ranked #729 of 825 articles by views in BMC Bioinformatics

Most read Least read

Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 825 in total.

Mint this article as an NFT
Not yet minted

Create a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.

5 SUSD one-off · no wallet required
Abstract
MOTIVATION: Biomedical named entity recognition (NER) presents unique challenges due to specialized vocabularies, the sheer volume of entities, and the continuous emergence of novel entities. Traditional NER models, constrained by fixed taxonomies and human annotations, struggle to generalize beyond predefined entity types. RESULTS: To address these issues, we introduce GLiNER-BioMed, a domain-adapted suite of GLiNER models for biomedicine. Our approach first distills the annotation capabilities of large language models (LLMs) into a smaller, more efficient model, enabling the generation of high-coverage biomedical NER data. We subsequently train two GLiNER architectures, uni- and bi-encoder, at multiple scales to balance computational efficiency and performance. Experiments on eight biomedical datasets demonstrate that GLiNER-BioMed achieved state-of-the-art zero-shot performance (micro-F1 59.77%), exceeding the strongest baseline by 5.96 points (p < 0.001). In few-shot learning, the bi-encoder variant reached 70.39% (10-shot), consistently outperforming the strongest baseline across all settings (p < 0.05). Our findings show that the uni-encoder GLiNER-BioMed achieves the strongest zero-shot performance, while the bi-encoder offers superior few-shot gains and substantially higher inference throughput (+39-568%), making it well-suited to annotation-limited, latency-sensitive, or large-label-space settings. Ablation studies further indicate that combining synthetic biomedical pre-training with general-domain post-training is essential for capturing domain-specific knowledge while maintaining precision-recall balance. AVAILABILITY AND IMPLEMENTATION: The source code, datasets, and models are publicly available at https://github.com/ds4dh/GLiNER-biomed. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Reference Key
openalex_W4415320068 Use this key to autocite in the manuscript while using SciMatic Manuscript Manager or Thesis Manager
Authors Anthony Yazdani, Ihor Stepanov, Douglas Teodoro
Journal BMC Bioinformatics
Year 2026
DOI
10.1093/bioinformatics/btag322
URL
Keywords Keywords not found

Citations

No citations found. To add a citation, contact the admin at info@scimatic.org

No comments yet. Be the first to comment on this article.