VLDB 2026 Research / reviewers in the wild / expert
Ke Hao
dblp:10/2618
· DBLP profile ↗
7ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 100% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 50% Time series and sequential data · 50% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Image and video processing › super-resolution
image super-resolution |
0.9 | 1 | 2025 | StyleSRN: Scene Text Image Super-Resolution with Text Style Embedding · ICCV 2025 |
Image and video processing › super-resolution › image super-resolution
scene text image super-resolution |
0.9 | 1 | 2025 | StyleSRN: Scene Text Image Super-Resolution with Text Style Embedding · ICCV 2025 |
Machine learning › Efficient and distributed learning
dataset distillation |
0.8 | 1 | 2024 | Dataset Condensation for Time Series Classification via Dual Domain Matching · KDD 2024 |
Machine learning › Time series and sequential data › time series analysis
time series classification |
0.8 | 1 | 2024 | Dataset Condensation for Time Series Classification via Dual Domain Matching · KDD 2024 |
Bioinformatics and computational biology › functional genomics
eQTL mapping |
0.4 | 1 | 2019 | GIGSEA: genotype imputed gene set enrichment analysis using GWAS summary level data · Bioinform. 2019 |
Bioinformatics and computational biology › gene expression analysis
gene expression prediction |
0.4 | 1 | 2019 | GIGSEA: genotype imputed gene set enrichment analysis using GWAS summary level data · Bioinform. 2019 |
Bioinformatics and computational biology › functional genomics › functional enrichment analysis
gene set enrichment analysis |
0.4 | 1 | 2019 | GIGSEA: genotype imputed gene set enrichment analysis using GWAS summary level data · Bioinform. 2019 |
Bioinformatics and computational biology
genomics |
0.4 | 1 | 2019 | GIGSEA: genotype imputed gene set enrichment analysis using GWAS summary level data · Bioinform. 2019 |
Bioinformatics and computational biology › statistical genetics
post-GWAS analysis |
0.4 | 1 | 2019 | GIGSEA: genotype imputed gene set enrichment analysis using GWAS summary level data · Bioinform. 2019 |
Image and video processing
image restoration |
0.3 | 1 | 2025 | StyleSRN: Scene Text Image Super-Resolution with Text Style Embedding · ICCV 2025 |
Bioinformatics and computational biology › genomics › genome-wide association study
GWAS summary statistics |
0.1 | 1 | 2019 | GIGSEA: genotype imputed gene set enrichment analysis using GWAS summary level data · Bioinform. 2019 |
Methods — techniques the papers use, named apart from their topics
multi-view data augmentation · 0.8dual domain training · 0.8weighted linear regression · 0.4permutation test · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | StyleSRN: Scene Text Image Super-Resolution with Text Style Embedding
Shengrong Yuan, Ke Hao, Xuqi Ma, Changxin Gao, Li Liu 0002, Nong Sang |
ICCV | 3 |
| 2024 | Dataset Condensation for Time Series Classification via Dual Domain MatchingabstractTime series data has been demonstrated to be crucial in various research fields. The management of large quantities of time series data presents challenges in terms of deep learning tasks, particularly for training a deep neural network. Recently, a technique named Dataset Condensation has emerged as a solution to this problem. This technique generates a smaller synthetic dataset that has comparable performance to the full real dataset in downstream tasks such as classification. However, previous methods are primarily designed for image and graph datasets, and directly adapting them to the time series dataset leads to suboptimal performance due to their inability to effectively leverage the rich information inherent in time series data, particularly in the frequency domain. In this paper, we propose a novel framework named Dataset Condensation for Time Series Classification via Dual Domain Matching (CondTSC) which focuses on the time series classification dataset condensation task. Different from previous methods, our proposed framework aims to generate a condensed dataset that matches the surrogate objectives in both the time and frequency domains. Specifically, CondTSC incorporates multi-view data augmentation, dual domain training, and dual surrogate objectives to enhance the dataset condensation process in the time and frequency domains. Through extensive experiments, we demonstrate the effectiveness of our proposed framework, which outperforms other baselines and learns a condensed synthetic dataset that exhibits desirable characteristics such as conforming to the distribution of the original data. Zhanyu Liu, Ke Hao, Guanjie Zheng, Yanwei Yu |
KDD | 2 |
| 2019 | GIGSEA: genotype imputed gene set enrichment analysis using GWAS summary level dataabstractSummary: level data of GWAS becomes increasingly important in post-GWAS data mining. Here, we present GIGSEA (Genotype Imputed Gene Set Enrichment Analysis), a novel method that uses GWAS summary statistics and eQTL to infer differential gene expression and interrogate gene set enrichment for the trait-associated SNPs. By incorporating empirical eQTL of the disease relevant tissue, GIGSEA naturally accounts for factors such as gene size, gene boundary, SNP distal regulation and multiple-marker regulation. The weighted linear regression model was used to perform the enrichment test, properly adjusting for imputation accuracy, model incompleteness and redundancy in different gene sets. The significance level of enrichment is assessed by the permutation test, where matrix operation was employed to dramatically improve computation speed. GIGSEA has appropriate type I error rates, and discovers the plausible biological findings on the real data set. Availability and implementation: GIGSEA is implemented in R, and freely available at www.github.com/zhushijia/GIGSEA. Supplementary information: Supplementary data are available at Bioinformatics online. Shijia Zhu, Tongqi Qian, Yujin Hoshida, Ke Hao |
Bioinform. | 6 |
| 2015 | SAAS-CNV: A Joint Segmentation Approach on Aggregated and Allele Specific Signals for the Identification of Somatic Copy Number Alterations with Next-Generation Sequencing DataabstractCancer genomes exhibit profound somatic copy number alterations (SCNAs). Studying tumor SCNAs using massively parallel sequencing provides unprecedented resolution and meanwhile gives rise to new challenges in data analysis, complicated by tumor aneuploidy and heterogeneity as well as normal cell contamination. While the majority of read depth based methods utilize total sequencing depth alone for SCNA inference, the allele specific signals are undervalued. We proposed a joint segmentation and inference approach using both signals to meet some of the challenges. Our method consists of four major steps: 1) extracting read depth supporting reference and alternative alleles at each SNP/Indel locus and comparing the total read depth and alternative allele proportion between tumor and matched normal sample; 2) performing joint segmentation on the two signal dimensions; 3) correcting the copy number baseline from which the SCNA state is determined; 4) calling SCNA state for each segment based on both signal dimensions. The method is applicable to whole exome/genome sequencing (WES/WGS) as well as SNP array data in a tumor-control study. We applied the method to a dataset containing no SCNAs to test the specificity, created by pairing sequencing replicates of a single HapMap sample as normal/tumor pairs, as well as a large-scale WGS dataset consisting of 88 liver tumors along with adjacent normal tissues. Compared with representative methods, our method demonstrated improved accuracy, scalability to large cancer studies, capability in handling both sequencing and SNP array data, and the potential to improve the estimation of tumor ploidy and purity. Zhongyang Zhang, Ke Hao |
PLoS Comput. Biol. | 2 |
| 2014 | Meta-eQTL: a tool set for flexible eQTL meta-analysisabstractBACKGROUND: Increasing number of eQTL (Expression Quantitative Trait Loci) datasets facilitate genetics and systems biology research. Meta-analysis tools are in need to jointly analyze datasets of same or similar issue types to improve statistical power especially in trans-eQTL mapping. Meta-analysis framework is also necessary for ChrX eQTL discovery. RESULTS: We developed a novel tool, meta-eqtl, for fast eQTL meta-analysis of arbitrary sample size and arbitrary number of datasets. Further, this tool accommodates versatile modeling, eg. non-parametric model and mixed effect models. In addition, meta-eqtl readily handles calculation of chrX eQTLs. CONCLUSIONS: We demonstrated and validated meta-eqtl as fast and comprehensive tool to meta-analyze multiple datasets and ChrX eQTL discovery. Meta-eqtl is a set of command line utilities written in R, with some computationally intensive parts written in C. The software runs on Linux platforms and is designed to intelligently adapt to high performance computing (HPC) cluster. We applied the novel tool to liver and adipose tissue data, and revealed eSNPs underlying diabetes GWAS loci. Antonio Di Narzo, Haoxiang Cheng, Ke Hao |
BMC Bioinform. | 4 |
| 2007 | Correction: Inferring Loss-of-Heterozygosity from Unpaired Tumors Using High-Density Oligonucleotide SNP ArraysabstractIn the subsection ''Initial probabilities'' of the Results section, the first sentence was incorrect and should say: These probabilities, denoted by P 0 (RET) and P 0 (LOSS) 1 P 0 (RET), specify the probabilities of RET and LOSS for the pterminal marker on a chromosome. Rameen Beroukhim, Yuhyun Park, Ke Hao, Xiaojun Zhao, Levi A. Garraway, Edward A. Fox, Ephraim P. Hochberg, Ingo K. Mellinghoff, Matthias D. Hofer, Aurelien Descazeaud, Mark A. Rubin, Matthew Meyerson, Wing Hung Wong, William R. Sellers, Cheng Li 0016 |
PLoS Comput. Biol. | 4 |
| 2006 | Inferring Loss-of-Heterozygosity from Unpaired Tumors Using High-Density Oligonucleotide SNP ArraysabstractLoss of heterozygosity (LOH) of chromosomal regions bearing tumor suppressors is a key event in the evolution of epithelial and mesenchymal tumors. Identification of these regions usually relies on genotyping tumor and counterpart normal DNA and noting regions where heterozygous alleles in the normal DNA become homozygous in the tumor. However, paired normal samples for tumors and cell lines are often not available. With the advent of oligonucleotide arrays that simultaneously assay thousands of single-nucleotide polymorphism (SNP) markers, genotyping can now be done at high enough resolution to allow identification of LOH events by the absence of heterozygous loci, without comparison to normal controls. Here we describe a hidden Markov model-based method to identify LOH from unpaired tumor samples, taking into account SNP intermarker distances, SNP-specific heterozygosity rates, and the haplotype structure of the human genome. When we applied the method to data genotyped on 100 K arrays, we correctly identified 99% of SNP markers as either retention or loss. We also correctly identified 81% of the regions of LOH, including 98% of regions greater than 3 megabases. By integrating copy number analysis into the method, we were able to distinguish LOH from allelic imbalance. Application of this method to data from a set of prostate samples without paired normals identified known regions of prevalent LOH. We have developed a method for analyzing high-density oligonucleotide SNP array data to accurately identify of regions of LOH and retention in tumors without the need for paired normal samples. Rameen Beroukhim, Yuhyun Park, Ke Hao, Xiaojun Zhao, Levi A. Garraway, Edward A. Fox, Ephraim P. Hochberg, Ingo K. Mellinghoff, Matthias D. Hofer, Aurelien Descazeaud, Mark A. Rubin, Matthew Meyerson, Wing Hung Wong, William R. Sellers, Cheng Li 0016 |
PLoS Comput. Biol. | 4 |