Ke Hao

dblp:10/2618 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%
Computer graphics and multimedia
1 paper
Image and video processing · 100%
Artificial intelligence
1 paper
Efficient and distributed learning · 50% Time series and sequential data · 50%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Image and video processing › super-resolution
image super-resolution
0.912025
StyleSRN: Scene Text Image Super-Resolution with Text Style Embedding · ICCV 2025
Image and video processing › super-resolution › image super-resolution
scene text image super-resolution
0.912025
StyleSRN: Scene Text Image Super-Resolution with Text Style Embedding · ICCV 2025
Machine learning › Efficient and distributed learning
dataset distillation
0.812024
Dataset Condensation for Time Series Classification via Dual Domain Matching · KDD 2024
Machine learning › Time series and sequential data › time series analysis
time series classification
0.812024
Dataset Condensation for Time Series Classification via Dual Domain Matching · KDD 2024
Bioinformatics and computational biology › functional genomics
eQTL mapping
0.412019
GIGSEA: genotype imputed gene set enrichment analysis using GWAS summary level data · Bioinform. 2019
Bioinformatics and computational biology › gene expression analysis
gene expression prediction
0.412019
GIGSEA: genotype imputed gene set enrichment analysis using GWAS summary level data · Bioinform. 2019
Bioinformatics and computational biology › functional genomics › functional enrichment analysis
gene set enrichment analysis
0.412019
GIGSEA: genotype imputed gene set enrichment analysis using GWAS summary level data · Bioinform. 2019
Bioinformatics and computational biology
genomics
0.412019
GIGSEA: genotype imputed gene set enrichment analysis using GWAS summary level data · Bioinform. 2019
Bioinformatics and computational biology › statistical genetics
post-GWAS analysis
0.412019
GIGSEA: genotype imputed gene set enrichment analysis using GWAS summary level data · Bioinform. 2019
Image and video processing
image restoration
0.312025
StyleSRN: Scene Text Image Super-Resolution with Text Style Embedding · ICCV 2025
Bioinformatics and computational biology › genomics › genome-wide association study
GWAS summary statistics
0.112019
GIGSEA: genotype imputed gene set enrichment analysis using GWAS summary level data · Bioinform. 2019

Methods — techniques the papers use, named apart from their topics

multi-view data augmentation · 0.8dual domain training · 0.8weighted linear regression · 0.4permutation test · 0.4
YearPublicationVenuePosition
2025 StyleSRN: Scene Text Image Super-Resolution with Text Style Embedding
Shengrong Yuan, Ke Hao, Xuqi Ma, Changxin Gao, Li Liu 0002, Nong Sang
ICCV3
2024 Dataset Condensation for Time Series Classification via Dual Domain Matching
abstract
Time series data has been demonstrated to be crucial in various research fields. The management of large quantities of time series data presents challenges in terms of deep learning tasks, particularly for training a deep neural network. Recently, a technique named Dataset Condensation has emerged as a solution to this problem. This technique generates a smaller synthetic dataset that has comparable performance to the full real dataset in downstream tasks such as classification. However, previous methods are primarily designed for image and graph datasets, and directly adapting them to the time series dataset leads to suboptimal performance due to their inability to effectively leverage the rich information inherent in time series data, particularly in the frequency domain. In this paper, we propose a novel framework named Dataset Condensation for Time Series Classification via Dual Domain Matching (CondTSC) which focuses on the time series classification dataset condensation task. Different from previous methods, our proposed framework aims to generate a condensed dataset that matches the surrogate objectives in both the time and frequency domains. Specifically, CondTSC incorporates multi-view data augmentation, dual domain training, and dual surrogate objectives to enhance the dataset condensation process in the time and frequency domains. Through extensive experiments, we demonstrate the effectiveness of our proposed framework, which outperforms other baselines and learns a condensed synthetic dataset that exhibits desirable characteristics such as conforming to the distribution of the original data.
Zhanyu Liu, Ke Hao, Guanjie Zheng, Yanwei Yu
KDD2
2019 GIGSEA: genotype imputed gene set enrichment analysis using GWAS summary level data
abstract
Summary: level data of GWAS becomes increasingly important in post-GWAS data mining. Here, we present GIGSEA (Genotype Imputed Gene Set Enrichment Analysis), a novel method that uses GWAS summary statistics and eQTL to infer differential gene expression and interrogate gene set enrichment for the trait-associated SNPs. By incorporating empirical eQTL of the disease relevant tissue, GIGSEA naturally accounts for factors such as gene size, gene boundary, SNP distal regulation and multiple-marker regulation. The weighted linear regression model was used to perform the enrichment test, properly adjusting for imputation accuracy, model incompleteness and redundancy in different gene sets. The significance level of enrichment is assessed by the permutation test, where matrix operation was employed to dramatically improve computation speed. GIGSEA has appropriate type I error rates, and discovers the plausible biological findings on the real data set. Availability and implementation: GIGSEA is implemented in R, and freely available at www.github.com/zhushijia/GIGSEA. Supplementary information: Supplementary data are available at Bioinformatics online.
Shijia Zhu, Tongqi Qian, Yujin Hoshida, Ke Hao
Bioinform.6
2015 SAAS-CNV: A Joint Segmentation Approach on Aggregated and Allele Specific Signals for the Identification of Somatic Copy Number Alterations with Next-Generation Sequencing Data
abstract
Cancer genomes exhibit profound somatic copy number alterations (SCNAs). Studying tumor SCNAs using massively parallel sequencing provides unprecedented resolution and meanwhile gives rise to new challenges in data analysis, complicated by tumor aneuploidy and heterogeneity as well as normal cell contamination. While the majority of read depth based methods utilize total sequencing depth alone for SCNA inference, the allele specific signals are undervalued. We proposed a joint segmentation and inference approach using both signals to meet some of the challenges. Our method consists of four major steps: 1) extracting read depth supporting reference and alternative alleles at each SNP/Indel locus and comparing the total read depth and alternative allele proportion between tumor and matched normal sample; 2) performing joint segmentation on the two signal dimensions; 3) correcting the copy number baseline from which the SCNA state is determined; 4) calling SCNA state for each segment based on both signal dimensions. The method is applicable to whole exome/genome sequencing (WES/WGS) as well as SNP array data in a tumor-control study. We applied the method to a dataset containing no SCNAs to test the specificity, created by pairing sequencing replicates of a single HapMap sample as normal/tumor pairs, as well as a large-scale WGS dataset consisting of 88 liver tumors along with adjacent normal tissues. Compared with representative methods, our method demonstrated improved accuracy, scalability to large cancer studies, capability in handling both sequencing and SNP array data, and the potential to improve the estimation of tumor ploidy and purity.
Zhongyang Zhang, Ke Hao
PLoS Comput. Biol.2
2014 Meta-eQTL: a tool set for flexible eQTL meta-analysis
abstract
BACKGROUND: Increasing number of eQTL (Expression Quantitative Trait Loci) datasets facilitate genetics and systems biology research. Meta-analysis tools are in need to jointly analyze datasets of same or similar issue types to improve statistical power especially in trans-eQTL mapping. Meta-analysis framework is also necessary for ChrX eQTL discovery. RESULTS: We developed a novel tool, meta-eqtl, for fast eQTL meta-analysis of arbitrary sample size and arbitrary number of datasets. Further, this tool accommodates versatile modeling, eg. non-parametric model and mixed effect models. In addition, meta-eqtl readily handles calculation of chrX eQTLs. CONCLUSIONS: We demonstrated and validated meta-eqtl as fast and comprehensive tool to meta-analyze multiple datasets and ChrX eQTL discovery. Meta-eqtl is a set of command line utilities written in R, with some computationally intensive parts written in C. The software runs on Linux platforms and is designed to intelligently adapt to high performance computing (HPC) cluster. We applied the novel tool to liver and adipose tissue data, and revealed eSNPs underlying diabetes GWAS loci.
Antonio Di Narzo, Haoxiang Cheng, Ke Hao
BMC Bioinform.4
2007 Correction: Inferring Loss-of-Heterozygosity from Unpaired Tumors Using High-Density Oligonucleotide SNP Arrays
abstract
In the subsection ''Initial probabilities'' of the Results section, the first sentence was incorrect and should say: These probabilities, denoted by P 0 (RET) and P 0 (LOSS) 1 P 0 (RET), specify the probabilities of RET and LOSS for the pterminal marker on a chromosome.
Rameen Beroukhim, Yuhyun Park, Ke Hao, Xiaojun Zhao, Levi A. Garraway, Edward A. Fox, Ephraim P. Hochberg, Ingo K. Mellinghoff, Matthias D. Hofer, Aurelien Descazeaud, Mark A. Rubin, Matthew Meyerson, Wing Hung Wong, William R. Sellers, Cheng Li 0016
PLoS Comput. Biol.4
2006 Inferring Loss-of-Heterozygosity from Unpaired Tumors Using High-Density Oligonucleotide SNP Arrays
abstract
Loss of heterozygosity (LOH) of chromosomal regions bearing tumor suppressors is a key event in the evolution of epithelial and mesenchymal tumors. Identification of these regions usually relies on genotyping tumor and counterpart normal DNA and noting regions where heterozygous alleles in the normal DNA become homozygous in the tumor. However, paired normal samples for tumors and cell lines are often not available. With the advent of oligonucleotide arrays that simultaneously assay thousands of single-nucleotide polymorphism (SNP) markers, genotyping can now be done at high enough resolution to allow identification of LOH events by the absence of heterozygous loci, without comparison to normal controls. Here we describe a hidden Markov model-based method to identify LOH from unpaired tumor samples, taking into account SNP intermarker distances, SNP-specific heterozygosity rates, and the haplotype structure of the human genome. When we applied the method to data genotyped on 100 K arrays, we correctly identified 99% of SNP markers as either retention or loss. We also correctly identified 81% of the regions of LOH, including 98% of regions greater than 3 megabases. By integrating copy number analysis into the method, we were able to distinguish LOH from allelic imbalance. Application of this method to data from a set of prostate samples without paired normals identified known regions of prevalent LOH. We have developed a method for analyzing high-density oligonucleotide SNP array data to accurately identify of regions of LOH and retention in tumors without the need for paired normal samples.
Rameen Beroukhim, Yuhyun Park, Ke Hao, Xiaojun Zhao, Levi A. Garraway, Edward A. Fox, Ephraim P. Hochberg, Ingo K. Mellinghoff, Matthias D. Hofer, Aurelien Descazeaud, Mark A. Rubin, Matthew Meyerson, Wing Hung Wong, William R. Sellers, Cheng Li 0016
PLoS Comput. Biol.4