A FAIR gold-standard benchmark for geographic entity recognition in cancer genomics literature for biomedical NLP
Clicks: 15
ID: 326836
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This
article has not been analysed, so there is no overall score —
reader engagement is measured and shown alongside.
Reader Engagement
Emerging Content
4.2
/100
15 views
14 readers
AI Quality Assessment
Not analyzed
Readership in this journal
EmergingRanked #25 of 115 articles by views in Bioinformatics advances
Most read
Least read
Bar heights use a square-root scale.
Mint this article as an NFT
Not yet mintedCreate a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.
5
SUSD
one-off · no wallet required
Abstract
Abstract Motivation Geographic biases in cancer genomics research limit precision oncology’s global applicability. Major consortia represent only 16 countries and sample less than 1% of 1.4 million publications worldwide. While automated literature mining could address these biases, the lack of benchmark datasets prevents rigorous development and validation of extraction models. We present a FAIR-compliant gold-standard benchmark of 129 manually curated cancer genomics studies spanning 46 countries, tripling geographic representation versus existing consortia. This dataset enables systematic evaluation of automated geographic information and molecular data extraction approaches and establishes a reference task for document-level geographic information extraction models. Results Our benchmark highlights critical challenges for AI-powered literature mining. GeoBoost2 achieves 55% macro-F1 and 46% micro-F1 for patient-origin extraction across 129 articles, exposing precision–recall tradeoffs requiring domain adaptation. Seventy percent of studies include author affiliations from countries unrelated to patient cohorts, showing that metadata-based approaches systematically misattribute geographic provenance. Of 129 studies, 96 provide molecular data for 263,472 cancer patients, revealing substantial untapped literature-derived resources absent from major consortia and supporting training and comparison of domain-specific NLP models. Availability and implementation The complete benchmark dataset, including manual annotations, automated extraction outputs, evaluation scripts, and supplementary documentation, is available at Zenodo with DOI 10.5281/zenodo.18259159.
| Reference Key |
openalex_W7204493702
Use this key to autocite in the manuscript while using
SciMatic Manuscript Manager or Thesis Manager
|
|---|---|
| Authors | Maricel G. Kann, Karen O’Connor, Petra Tembei, Graciela Gonzalez Hernandez |
| Journal | Bioinformatics advances |
| Year | 2026 |
| DOI |
10.1093/bioadv/vbag225
|
| URL | |
| Keywords | Keywords not found |
Citations
No citations found. To add a citation, contact the admin at info@scimatic.org
Comments
No comments yet. Be the first to comment on this article.