TY - JOUR
T1 - ZIPcnv
T2 - accurate and efficient inference of copy number variations from shallow whole-genome sequencing
AU - Xue, Zhengfa
AU - Zeng, Jingyu
AU - Wang, Xuwen
AU - Yuan, Jiajing
AU - Wang, Tianci
AU - Lai, Xin
AU - Wang, Lin
AU - Wang, Yu
AU - Zhu, Huanhuan
AU - Jin, Xin
AU - Wang, Jiayin
N1 - Publisher Copyright:
© The Author(s) 2025. Published by Oxford University Press.
PY - 2025/11/1
Y1 - 2025/11/1
N2 - Motivation: Shallow whole-genome sequencing (sWGS), a rapid and cost-effective sequencing technology, has gradually been widely adopted for CNV analyses. However, with genome‑wide coverage of only 0.1–5×, sWGS data display a pronounced zero‑inflation phenomenon—a large fraction of loci has zero sequencing reads. Zero inflation causes read counts to fluctuate by several‑fold between adjacent windows. As a result, random upward blips in coverage can be misinterpreted as copy‑number gains (false positives), and true deletions often become indistinguishable from pervasive zero‑coverage noise. In addition, existing CNV detection tools developed for sWGS data often struggle to adapt across different CNV sizes. These combined effects severely constrain the accuracy of CNV inference. Results: To address above challenges, we propose ZIPcnv, a novel CNV detection tool specifically designed for sWGS data. First, we apply a segment sliding window to smooth the raw read depth signal, which transforms the original zero-inflated statistical characteristics into approximately normal distribution characteristics. We then design a statistical process model that robustly detects persistent shifts under high background noise using a cumulative sum strategy, classifying genomic regions into candidate and non-candidate CNV regions. Finally, dynamic sliding windows are used for one-pass detection of CNVs of varying lengths, with window size adapting to the CNV region size. We evaluated the performance of ZIPcnv on simulated data and 190 real whole-genome sequencing samples. Experimental results show that ZIPcnv consistently outperforms currently popular CNV detection tools.
AB - Motivation: Shallow whole-genome sequencing (sWGS), a rapid and cost-effective sequencing technology, has gradually been widely adopted for CNV analyses. However, with genome‑wide coverage of only 0.1–5×, sWGS data display a pronounced zero‑inflation phenomenon—a large fraction of loci has zero sequencing reads. Zero inflation causes read counts to fluctuate by several‑fold between adjacent windows. As a result, random upward blips in coverage can be misinterpreted as copy‑number gains (false positives), and true deletions often become indistinguishable from pervasive zero‑coverage noise. In addition, existing CNV detection tools developed for sWGS data often struggle to adapt across different CNV sizes. These combined effects severely constrain the accuracy of CNV inference. Results: To address above challenges, we propose ZIPcnv, a novel CNV detection tool specifically designed for sWGS data. First, we apply a segment sliding window to smooth the raw read depth signal, which transforms the original zero-inflated statistical characteristics into approximately normal distribution characteristics. We then design a statistical process model that robustly detects persistent shifts under high background noise using a cumulative sum strategy, classifying genomic regions into candidate and non-candidate CNV regions. Finally, dynamic sliding windows are used for one-pass detection of CNVs of varying lengths, with window size adapting to the CNV region size. We evaluated the performance of ZIPcnv on simulated data and 190 real whole-genome sequencing samples. Experimental results show that ZIPcnv consistently outperforms currently popular CNV detection tools.
UR - https://www.scopus.com/pages/publications/105022796942
U2 - 10.1093/bioinformatics/btaf592
DO - 10.1093/bioinformatics/btaf592
M3 - 文章
C2 - 41138169
AN - SCOPUS:105022796942
SN - 1367-4803
VL - 41
JO - Bioinformatics
JF - Bioinformatics
IS - 11
M1 - btaf592
ER -