RefSeq: expanding the Prokaryotic Genome Annotation Pipeline reach with protein family model curation

Clicks: 1
ID: 301065
2020
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This article has not been analysed, so there is no overall score — reader engagement is measured and shown alongside.
AI Quality Assessment
Not analyzed
Readership in this journal

Ranked #1,091 of 1,214 articles by views in Nucleic Acids Research

Most read Least read

Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 1,214 in total.

Mint this article as an NFT
Not yet minted

Create a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.

5 SUSD one-off · no wallet required
Abstract
Abstract The Reference Sequence (RefSeq) project at the National Center for Biotechnology Information (NCBI) contains nearly 200 000 bacterial and archaeal genomes and 150 million proteins with up-to-date annotation. Changes in the Prokaryotic Genome Annotation Pipeline (PGAP) since 2018 have resulted in a substantial reduction in spurious annotation. The hierarchical collection of protein family models (PFMs) used by PGAP as evidence for structural and functional annotation was expanded to over 35 000 protein profile hidden Markov models (HMMs), 12 300 BlastRules and 36 000 curated CDD architectures. As a result, >122 million or 79% of RefSeq proteins are now named based on a match to a curated PFM. Gene symbols, Enzyme Commission numbers or supporting publication attributes are available on over 40% of the PFMs and are inherited by the proteins and features they name, facilitating multi-genome analyses and connections to the literature. In adherence with the principles of FAIR (findable, accessible, interoperable, reusable), the PFMs are available in the Protein Family Models Entrez database to any user. Finally, the reference and representative genome set, a taxonomically diverse subset of RefSeq prokaryotic genomes, is now recalculated regularly and available for download and homology searches with BLAST. RefSeq is found at https://www.ncbi.nlm.nih.gov/refseq/.
Reference Key
openalex_W3109735703 Use this key to autocite in the manuscript while using SciMatic Manuscript Manager or Thesis Manager
Authors Wenjun Li, Kathleen O’Neill, Daniel H. Haft, Michael DiCuccio, Vyacheslav Chetvernin, Azat Badretdin, George Coulouris, Farideh Chitsaz, Myra K. Derbyshire, A. Scott Durkin, Noreen R. Gonzales, Marc Gwadz, Christopher J. Lanczycki, James S. Song, Narmada Thanki, Jiyao Wang, Roxanne A. Yamashita, Mingzhang Yang, Chanjuan Zheng, Aron Marchler‐Bauer, Françoise Thibaud‐Nissen
Journal Nucleic Acids Research
Year 2020
DOI
10.1093/nar/gkaa1105
URL
Keywords Keywords not found

Citations

No citations found. To add a citation, contact the admin at info@scimatic.org

No comments yet. Be the first to comment on this article.