VLDB 2026 Research / reviewers in the wild / expert
Hao Chen 0064
dblp:175/3324-64
· DBLP profile ↗
6ranked-venue papers
1as first author
2since 2021 · last 2025
0000-0002-4597-2773ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorTheory of computation · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Asymptotic Distribution-Free Change-Point Detection for Modern Data Based on a New Ranking SchemeabstractChange-point detection (CPD) involves identifying distributional changes in a sequence of independent observations. Among nonparametric methods, rank-based methods are attractive due to their robustness and effectiveness, and have been extensively studied for univariate data. However, they are not well explored for high-dimensional or non-Euclidean data. This paper proposes a new method, Rank INduced by Graph Change-Point Detection (RING-CPD), which utilizes graph-induced ranks to handle high-dimensional and non-Euclidean data. The new method is asymptotically distribution-free under the null hypothesis, and an analytic p-value approximation is provided for easy type-I error control. Simulation studies show that RING-CPD effectively detects change points across a wide range of alternatives and is also robust to heavy-tailed distribution and outliers. The new method is illustrated by the detection of seizures in a functional connectivity network dataset, changes in digit images, and travel pattern changes in the New York City Taxi dataset. Doudou Zhou, Hao Chen 0064 |
IEEE Trans. Inf. Theory | 2 |
| 2023 | Likelihood Scores for Sparse Signal and Change-Point DetectionabstractWe consider here the identification of change-points on large-scale data streams. The objective is to find the most efficient way of combining information across data stream so that detection is possible under the smallest detectable change magnitude. The challenge comes from the sparsity of change-points when only a small fraction of data streams undergo change at any point in time. The most successful approach to the sparsity issue so far has been the application of hard thresholding such that only local scores from data streams exhibiting significant changes are considered and added. However the identification of an optimal threshold is a difficult one. In particular it is unlikely that the same threshold is optimal for different levels of sparsity. We propose here a sparse likelihood score for identifying a sparse signal. The score is a likelihood ratio for testing between the null hypothesis of no change against an alternative hypothesis in which the change-points or signals are barely detectable. By the Neyman-Pearson Lemma this score has maximum detection power at the given alternative. The outcome is that we have a scoring of data streams that is successful in detecting at the boundary of the detectable region of signals and change-points. The likelihood score can be seen as a soft thresholding approach to sparse signal and change-point detection in which local scores that indicate small changes are down-weighted much more than local scores indicating large changes. We are able to show sharp optimality of the sparsity likelihood score in the sense of achieving successful detection at the minimum detectable order of change magnitude as well as the best constant with respect this order of change. Shouri Hu, Jingyan Huang, Hao Chen 0064, Hock Peng Chan |
IEEE Trans. Inf. Theory | 3 |
| 2018 | DNA copy number profiling using single-cell sequencingabstractCurrently, there is a lack of software for detecting copy number variations and constructing copy number profile for the whole genome from single-cell DNA sequencing data, which are often of low coverage and high technical noises. Here we introduce a new toolkit, SCNV, which features an efficient bin-free segmentation approach and provides the highest resolution possible for breakpoint detection and the subsequent copy number calling. SCNV can auto-tune parameters based on a set of normal cells from the same batch to adjust for the technical noise level of the data, facilitating its application to data gathered from different platforms and different studies. Hao Chen 0064, Nancy Ruonan Zhang |
Briefings Bioinform. | 2 |
| 2018 | Integrative pipeline for profiling DNA copy number and inferring tumor phylogenyabstractSummary: Copy number variation is an important and abundant source of variation in the human genome, which has been associated with a number of diseases, especially cancer. Massively parallel next-generation sequencing allows copy number profiling with fine resolution. Such efforts, however, have met with mixed successes, with setbacks arising partly from the lack of reliable analytical methods to meet the diverse and unique challenges arising from the myriad experimental designs and study goals in genetic studies. In cancer genomics, detection of somatic copy number changes and profiling of allele-specific copy number (ASCN) are complicated by experimental biases and artifacts as well as normal cell contamination and cancer subclone admixture. Furthermore, careful statistical modeling is warranted to reconstruct tumor phylogeny by both somatic ASCN changes and single nucleotide variants. Here we describe a flexible computational pipeline, MARATHON, which integrates multiple related statistical software for copy number profiling and downstream analyses in disease genetic studies. Availability and implementation: MARATHON is publicly available at https://github.com/yuchaojiang/MARATHON. Supplementary information: Supplementary data are available at Bioinformatics online. Eugene Urrutia, Hao Chen 0064, Zilu Zhou, Nancy Ruonan Zhang |
Bioinform. | 2 |
| 2016 | Global copy number profiling of cancer genomesabstractUNLABELLED: In this article, we introduce a robust and efficient strategy for deriving global and allele-specific copy number alternations (CNA) from cancer whole exome sequencing data based on Log R ratios and B-allele frequencies. Applying the approach to the analysis of over 200 skin cancer samples, we demonstrate its utility for discovering distinct CNA events and for deriving ancillary information such as tumor purity. AVAILABILITY AND IMPLEMENTATION: https://github.com/xfwang/CLOSE CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mengjie Chen, Xiaoqing Yu, Natapol Pornputtapong, Hao Chen 0064, Nancy Ruonan Zhang, R. Scott Powers, Michael Krauthammer |
Bioinform. | 5 |
| 2011 | Estimation of Parent Specific DNA Copy Number in Tumors using High-Density Genotyping ArraysabstractChromosomal gains and losses comprise an important type of genetic change in tumors, and can now be assayed using microarray hybridization-based experiments. Most current statistical models for DNA copy number estimate total copy number, which do not distinguish between the underlying quantities of the two inherited chromosomes. This latter information, sometimes called parent specific copy number, is important for identifying allele-specific amplifications and deletions, for quantifying normal cell contamination, and for giving a more complete molecular portrait of the tumor. We propose a stochastic segmentation model for parent-specific DNA copy number in tumor samples, and give an estimation procedure that is computationally efficient and can be applied to data from the current high density genotyping platforms. The proposed method does not require matched normal samples, and can estimate the unknown genotypes simultaneously with the parent specific copy number. The new method is used to analyze 223 glioblastoma samples from the Cancer Genome Atlas (TCGA) project, giving a more comprehensive summary of the copy number events in these samples. Detailed case studies on these samples reveal the additional insights that can be gained from an allele-specific copy number analysis, such as the quantification of fractional gains and losses, the identification of copy neutral loss of heterozygosity, and the characterization of regions of simultaneous changes of both inherited chromosomes. Hao Chen 0064, Haipeng Xing, Nancy Ruonan Zhang |
PLoS Comput. Biol. | 1 |