TY - JOUR
T1 - An Algorithm with Base-Pair Resolution for Identifying Cancer Heterogeneity by Estimating Multiple Clonal Haplotypes
AU - Geng, Yu
AU - Zhao, Zhongmeng
AU - Liu, Jianye
AU - Xu, Jing
AU - Cui, Daibing
AU - Xiao, Xiao
AU - Wang, Jiayin
N1 - Publisher Copyright:
© 2017, Editorial Office of Journal of Xi'an Jiaotong University. All right reserved.
PY - 2017/6/10
Y1 - 2017/6/10
N2 - An algorithm for identifying haplotype heterogeneity in cancer genomes is proposed to consider somatic mutational events carried by multiple sub-clones. The algorithm is based on the genomic sequencing data with multiple libraries of tumor tissue and extracts the features from both the multi-library and the constraints of paired-end reads. A priori number of sub-clones is roughly estimated by clustering the allelic variant frequency of each somatic loci. A contig-and-extension algorithm is designed, and the haplotype sequences are assembled by traversing the reads mapping to the loci. Thus, the contigs present an identification resolution on base-pair level. The number and proportion of sub-clones and the evolution relationships among them are further estimated by maximizing the likelihood of the posterior probabilities. Simulation results show that the algorithm reaches 99% in accuracy when the sequencing based library satisfies some coverage. The proposed algorithm outperforms the existing two-stage pipeline, which is widely used in data analysis now.
AB - An algorithm for identifying haplotype heterogeneity in cancer genomes is proposed to consider somatic mutational events carried by multiple sub-clones. The algorithm is based on the genomic sequencing data with multiple libraries of tumor tissue and extracts the features from both the multi-library and the constraints of paired-end reads. A priori number of sub-clones is roughly estimated by clustering the allelic variant frequency of each somatic loci. A contig-and-extension algorithm is designed, and the haplotype sequences are assembled by traversing the reads mapping to the loci. Thus, the contigs present an identification resolution on base-pair level. The number and proportion of sub-clones and the evolution relationships among them are further estimated by maximizing the likelihood of the posterior probabilities. Simulation results show that the algorithm reaches 99% in accuracy when the sequencing based library satisfies some coverage. The proposed algorithm outperforms the existing two-stage pipeline, which is widely used in data analysis now.
KW - Cancer heterogeneity
KW - Contig-and-extension algorithm
KW - Haplotype heterogeneity
KW - Multi-library sequencing
KW - Sub-clone deconvolution
UR - https://www.scopus.com/pages/publications/85029851179
U2 - 10.7652/xjtuxb201706015
DO - 10.7652/xjtuxb201706015
M3 - 文章
AN - SCOPUS:85029851179
SN - 0253-987X
VL - 51
SP - 92
EP - 96
JO - Hsi-An Chiao Tung Ta Hsueh/Journal of Xi'an Jiaotong University
JF - Hsi-An Chiao Tung Ta Hsueh/Journal of Xi'an Jiaotong University
IS - 6
ER -