CNV-Finder: Streamlining Copy Number Variation Discovery

Clicks: 8
ID: 322593
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This article has not been analysed, so there is no overall score — reader engagement is measured and shown alongside.
AI Quality Assessment
Not analyzed
Readership in this journal

Ranked #15 of 103 articles by views in Bioinformatics advances

Most read Least read

Bar heights use a square-root scale.

Mint this article as an NFT
Not yet minted

Create a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.

5 SUSD one-off · no wallet required
Abstract
Abstract Copy Number Variations (CNVs) play pivotal roles in the etiology of complex diseases and are variable across diverse populations. Understanding the association between CNVs and disease susceptibility is significant in disease genetics research and often requires analysis of large sample sizes. One of the most cost-effective and scalable methods for detecting CNVs is based on normalized signal intensity values, such as Log R Ratio (LRR) and B Allele Frequency (BAF), from Illumina genotyping arrays. In this study, we present CNV-Finder, a novel pipeline integrating deep learning techniques on array data, specifically a Long Short-Term Memory (LSTM) network, to expedite the large-scale identification of CNVs within predefined genomic regions. This facilitates efficient prioritization of samples for time-consuming or costly subsequent analyses such as Multiplex Ligation-dependent Probe Amplification (MLPA), short-read, and long-read whole genome sequencing. We incorporate four genes to establish our methods—Parkin ( PRKN ), Leucine Rich Repeat And Ig Domain Containing 2 ( LINGO2 ), Microtubule Associated Protein Tau ( MAPT ), and alpha-Synuclein ( SNCA )—which may be relevant to neurological diseases such as Alzheimer’s disease (AD), Parkinson’s disease (PD), Progressive Supranuclear Palsy (PSP), or related disorders such as essential tremor (ET). By training our models on expert-annotated samples and validating them across diverse cohorts, including those from the Global Parkinson’s Genetics Program (GP2) and additional dementia-specific databases, we demonstrate the efficacy of CNV-Finder in accurately detecting deletions and duplications. Our pipeline outputs app-compatible files for visualization within CNV-Finder’s interactive web application. This interface enables researchers to review predictions and filter displayed samples by model prediction values, LRR range, and variant count in order to explore or confirm results. Our pipeline integrates this human feedback to enhance model performance and reduce false positive rates. Through a series of comprehensive analyses and validations using visual inspection, MLPA, short-read, and long-read sequencing data, we demonstrate the robustness and adaptability of CNV-Finder in identifying CNVs with regions of varied size, probe density, and noise. Our findings highlight the significance of contextual understanding and human expertise in enhancing the precision of CNV identification, particularly in complex genomic regions like 17q21.31. The CNV-Finder pipeline is a scalable, publicly available resource for the scientific community, available on GitHub ( https://github.com/GP2code/CNV-Finder ; DOI 10.5281/zenodo.14182563 ). CNV-Finder not only expedites accurate candidate identification but also significantly reduces the manual workload for researchers, enabling future targeted validation and downstream analyses in regions or phenotypes of interest.
Reference Key
openalex_W4404648005 Use this key to autocite in the manuscript while using SciMatic Manuscript Manager or Thesis Manager
Authors Nicole Kuznetsov, Kensuke Daida, Mary B. Makarious, Bashayer Al‐Mubarak, Kajsa Brolin, Laksh Malik, Cedric Kouam, Breeana Baker, Raquel Real, Kathryn Step, Lara M. Lange, Lesley Wu, Miriam Ostrožovičová, Katherine M. Andersh, Pin‐Jui Kung, Yasser Mecheri, Yi Wen Tay, Behloul Soundous Malek, Nada Al Tassan, María Teresa Periñán, Samantha Hong, Mathew J. Koretsky, Lana Sargeant, Kristin Levine, Cornelis Blauwendraat, Kimberley J. Billingsley, Sara Bandrés‐Ciga, Hampton L. Leonard, Soraya Bardien, Huw R. Morris, Andrew Singleton, Mike A. Nalls, Dan Vitale, The Global Parkinson’s Genetics Program (GP2)
Journal Bioinformatics advances
Year 2026
DOI
10.1093/bioadv/vbag205
URL
Keywords Keywords not found

Citations

No citations found. To add a citation, contact the admin at info@scimatic.org

No comments yet. Be the first to comment on this article.