Taming the reference genome jungle: the refget sequence collection standard
Clicks: 3
ID: 323214
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This
article has not been analysed, so there is no overall score —
reader engagement is measured and shown alongside.
Reader Engagement
Emerging Content
0.6
/100
3 views
0 readers
AI Quality Assessment
Not analyzed
Readership in this journal
EmergingRanked #314 of 829 articles by views in BMC Bioinformatics
Most read
Least read
Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 829 in total.
Mint this article as an NFT
Not yet mintedCreate a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.
5
SUSD
one-off · no wallet required
Abstract
MOTIVATION: Reference genomes are foundational to genomics but suffer from widespread ambiguity and incompatibility due to inconsistent naming, undocumented differences, and lack of formal mechanisms for comparison. RESULTS: To address this, we introduce the GA4GH refget Sequence Collections (seqcol) standard. Refget seqcol is a framework for unambiguous representation, retrieval, and comparison of sequence collections such as reference genomes and transcriptomes. The seqcol standard comprises four components: a structured data schema, a canonical encoding algorithm that produces content-based, globally unique identifiers, a retrieval API, and a comparison protocol. This standard enables precise identification of sequence collections, even across decentralized or private systems, and allows compatibility assessments beyond exact identity, such as order-relaxed matches or shared coordinate systems. We applied the refget seqcol standard to 60 human and 36 mouse reference genomes sourced from major providers. Using digest-based comparisons, we quantified levels of similarity across attributes including sequence names, lengths, coordinate systems, and actual sequence content. Our analysis revealed some consistent subsets of sequences or coordinate systems, as well as substantial incompatibility among references and duplicate references under different names. This work offers a scalable, reproducible solution to the reference genome compatibility crisis, enabling improved transparency, reuse, and integration in genomic analyses. Refget seqcol enhances interoperability across tools and datasets, making genomic research more robust and reproducible. AVAILABILITY AND IMPLEMENTATION: To support adoption of refget seqcol, we provide a Python package implementing the full standard, a web API, and a comparison interface allowing users to assess local references against a curated database. The formal specification is hosted at https://ga4gh.github.io/refget/ and the reference implementation can be found at https://github.com/refgenie/refget.
| Reference Key |
openalex_W4414879798
Use this key to autocite in the manuscript while using
SciMatic Manuscript Manager or Thesis Manager
|
|---|---|
| Authors | Donald R. Campbell, Timothée Cezard, Sveinung Gundersen, Andrew Yates, Robert M. Davies, John Marshall, Sang‐Hoon Park, Alex H. Wagner, Michael I. Love, R. Thomas, Oliver Hofmann, Nathan C. Sheffield |
| Journal | BMC Bioinformatics |
| Year | 2026 |
| DOI |
10.1093/bioinformatics/btag554
|
| URL | |
| Keywords | Keywords not found |
Citations
No citations found. To add a citation, contact the admin at info@scimatic.org
Comments
No comments yet. Be the first to comment on this article.