CistromeMeta: A Large Language Model Powered Tool for Automated ChIP-seq Metadata Extraction
Clicks: 3
ID: 317409
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This
article has not been analysed, so there is no overall score —
reader engagement is measured and shown alongside.
Reader Engagement
Emerging Content
0.6
/100
3 views
2 readers
AI Quality Assessment
Not analyzed
Readership in this journal
EmergingRanked #811 of 835 articles by views in BMC Bioinformatics
Most read
Least read
Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 835 in total.
Mint this article as an NFT
Not yet mintedCreate a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.
5
SUSD
one-off · no wallet required
Abstract
SUMMARY: Public repositories such as NCBI's Gene Expression Omnibus (GEO) contain large numbers of ChIP-seq experiments, but their reuse is limited by heterogeneous free-text metadata describing target proteins, histone marks, cell lines, tissues, and disease states. We introduce CistromeMeta, a Python-based command-line tool that leverages large language models (LLMs) in a few-shot setting to automatically extract and standardize ChIP-seq metadata from GEO XML records without custom model training. The tool validates extracted terms against authoritative biological databases, including NCBI Gene, Harmonizome 3.0, AnimalTFDB 4.0, Cellosaurus, Experimental Factor Ontology, and Uberon, producing standardized outputs with official gene symbols and ontology identifiers for scalable metadata curation. AVAILABILITY AND IMPLEMENTATION: The Python source code is freely available at https://github.com/nickpiccaro/CistromeMetaX. An archived version of the software is available through Zenodo at DOI: 10.5281/zenodo.20244834. The tool requires Python 3.6+ and an OpenAI API key. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
| Reference Key |
openalex_W7164770084
Use this key to autocite in the manuscript while using
SciMatic Manuscript Manager or Thesis Manager
|
|---|---|
| Authors | Nicholas Piccaro, Myles Brown, Clifford Meyer |
| Journal | BMC Bioinformatics |
| Year | 2026 |
| DOI |
10.1093/bioinformatics/btag380
|
| URL | |
| Keywords | Keywords not found |
Citations
No citations found. To add a citation, contact the admin at info@scimatic.org
Comments
No comments yet. Be the first to comment on this article.