VLDB 2026 Research / reviewers in the wild / expert
Eun Yong Kang
dblp:38/1672
· DBLP profile ↗
8ranked-venue papers
2as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
5 papers |
Bioinformatics and computational biology · 100% | |
| Databases, data mining, and information retrieval
2 papers |
Recommender systems · 90% Data integration and cleaning · 10% | |
| Theoretical computer science
1 paper |
Algorithms and data structures · 50% Mathematical optimization · 50% | |
| Artificial intelligence
1 paper |
Probabilistic and Bayesian machine learning · 67% Knowledge representation and reasoning · 33% |
Topics — the 17 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Recommender systems › diversified recommendation
determinantal point process |
0.5 | 1 | 2021 | Diversity on the Go! Streaming Determinantal Point Processes under a Maximum Induced Cardinality Objective · WWW 2021 |
Recommender systems
diversified recommendation |
0.5 | 1 | 2021 | Diversity on the Go! Streaming Determinantal Point Processes under a Maximum Induced Cardinality Objective · WWW 2021 |
Algorithms and data structures › data streams
streaming algorithms |
0.5 | 1 | 2021 | Diversity on the Go! Streaming Determinantal Point Processes under a Maximum Induced Cardinality Objective · WWW 2021 |
Mathematical optimization › submodular optimization › submodular maximization
streaming submodular maximization |
0.5 | 1 | 2021 | Diversity on the Go! Streaming Determinantal Point Processes under a Maximum Induced Cardinality Objective · WWW 2021 |
Bioinformatics and computational biology
statistical genetics |
0.4 | 2 | 2014 | Fast pairwise IBD association testing in genome-wide association studies · Bioinform. 2014 A powerful and efficient set test for genetic markers that handles confounders · Bioinform. 2013 |
Bioinformatics and computational biology › functional genomics
eQTL mapping |
0.3 | 1 | 2017 | Applying meta-analysis to genotype-tissue expression data from multiple tissues to identify eQTLs and increase the number of eGenes · Bioinform. 2017 |
Bioinformatics and computational biology
genomics |
0.3 | 1 | 2017 | Applying meta-analysis to genotype-tissue expression data from multiple tissues to identify eQTLs and increase the number of eGenes · Bioinform. 2017 |
Bioinformatics and computational biology › population genetics
population stratification correction |
0.2 | 1 | 2015 | Efficient and Accurate Multiple-Phenotypes Regression Method for High Dimensional Data Considering Population Structure · RECOMB 2015 |
Bioinformatics and computational biology › biostatistics › statistical bioinformatics
statistical genomics |
0.2 | 1 | 2015 | Efficient and Accurate Multiple-Phenotypes Regression Method for High Dimensional Data Considering Population Structure · RECOMB 2015 |
Bioinformatics and computational biology › genomics
genome-wide association study |
0.2 | 1 | 2014 | Fast pairwise IBD association testing in genome-wide association studies · Bioinform. 2014 |
Bioinformatics and computational biology › statistical genetics
genetic association study |
0.2 | 1 | 2013 | A powerful and efficient set test for genetic markers that handles confounders · Bioinform. 2013 |
Bioinformatics and computational biology › statistical genetics
linear mixed model |
0.2 | 1 | 2013 | A powerful and efficient set test for genetic markers that handles confounders · Bioinform. 2013 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › structure learning
bayesian network structure learning |
0.1 | 1 | 2010 | Respecting Markov Equivalence in Computing Posterior Probabilities of Causal Graphical Features · AAAI 2010 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › causal reasoning
causal graphical model |
0.1 | 1 | 2010 | Respecting Markov Equivalence in Computing Posterior Probabilities of Causal Graphical Features · AAAI 2010 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal discovery
markov equivalence class |
0.1 | 1 | 2010 | Respecting Markov Equivalence in Computing Posterior Probabilities of Causal Graphical Features · AAAI 2010 |
Bioinformatics and computational biology
gene expression analysis |
0.1 | 1 | 2009 | Detecting the Presence and Absence of Causal Relationships between Expression of Yeast Genes with Very Few Samples · RECOMB 2009 |
Data integration and cleaning › data extraction
web data extraction |
0.1 | 1 | 2008 | Automatically Extracting Form Labels · ICDE 2008 |
Methods — techniques the papers use, named apart from their topics
greedy algorithm · 1.0determinantal point process · 1.0meta-analysis · 0.3mixed model · 0.2permutation testing · 0.2importance sampling · 0.2score test · 0.2linear mixed model · 0.2likelihood ratio test · 0.2dynamic programming · 0.1machine learning classification · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Diversity on the Go! Streaming Determinantal Point Processes under a Maximum Induced Cardinality ObjectiveabstractOver the past decade, Determinantal Point Processes (DPPs) have proven to be a mathematically elegant framework for modeling diversity. Given a set of items N, DPPs define a probability distribution over subsets of N, with sets of larger diversity having greater probability. Recently, DPPs have achieved success in the domain of recommendation systems, as a method to enforce diversity of recommendations in addition to relevance. In large-scale recommendation applications however, the input typically comes in the form of a stream too large to fit into main memory. However, the natural greedy algorithm for DPP-based recommendations is memory intensive, and cannot be used in a streaming setting. Paul Liu 0001, Akshay Soni, Eun Yong Kang, Yajun Wang 0001, Mehul Parsana |
WWW | 3 |
| 2017 | Applying meta-analysis to genotype-tissue expression data from multiple tissues to identify eQTLs and increase the number of eGenesabstractMOTIVATION: There is recent interest in using gene expression data to contextualize findings from traditional genome-wide association studies (GWAS). Conditioned on a tissue, expression quantitative trait loci (eQTLs) are genetic variants associated with gene expression, and eGenes are genes whose expression levels are associated with genetic variants. eQTLs and eGenes provide great supporting evidence for GWAS hits and important insights into the regulatory pathways involved in many diseases. When a significant variant or a candidate gene identified by GWAS is also an eQTL or eGene, there is strong evidence to further study this variant or gene. Multi-tissue gene expression datasets like the Gene Tissue Expression (GTEx) data are used to find eQTLs and eGenes. Unfortunately, these datasets often have small sample sizes in some tissues. For this reason, there have been many meta-analysis methods designed to combine gene expression data across many tissues to increase power for finding eQTLs and eGenes. However, these existing techniques are not scalable to datasets containing many tissues, like the GTEx data. Furthermore, these methods ignore a biological insight that the same variant may be associated with the same gene across similar tissues. RESULTS: We introduce a meta-analysis model that addresses these problems in existing methods. We focus on the problem of finding eGenes in gene expression data from many tissues, and show that our model is better than other types of meta-analyses. AVAILABILITY AND IMPLEMENTATION: Source code is at https://github.com/datduong/RECOV . CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Dat Duong, Lisa Gai, Sagi Snir, Eun Yong Kang, Buhm Han, Jae Hoon Sul, Eleazar Eskin |
Bioinform. | 4 |
| 2015 | Efficient and Accurate Multiple-Phenotypes Regression Method for High Dimensional Data Considering Population Structure
Jong Wha J. Joo, Eun Yong Kang, Elin Org, Nicholas A. Furlotte, Brian Parks, Aldons J. Lusis, Eleazar Eskin |
RECOMB | 2 |
| 2014 | Fast pairwise IBD association testing in genome-wide association studiesabstractMOTIVATION: Recently, investigators have proposed state-of-the-art Identity-by-descent (IBD) mapping methods to detect IBD segments between purportedly unrelated individuals. The IBD information can then be used for association testing in genetic association studies. One approach for this IBD association testing strategy is to test for excessive IBD between pairs of cases ('pairwise method'). However, this approach is inefficient because it requires a large number of permutations. Moreover, a limited number of permutations define a lower bound for P-values, which makes fine-mapping of associated regions difficult because, in practice, a much larger genomic region is implicated than the region that is actually associated. RESULTS: In this article, we introduce a new pairwise method 'Fast-Pairwise'. Fast-Pairwise uses importance sampling to improve efficiency and enable approximation of extremely small P-values. Fast-Pairwise method takes only days to complete a genome-wide scan. In the application to the WTCCC type 1 diabetes data, Fast-Pairwise successfully fine-maps a known human leukocyte antigen gene that is known to cause the disease. AVAILABILITY: Fast-Pairwise is publicly available at: http://genetics.cs.ucla.edu/graphibd. Buhm Han, Eun Yong Kang, Soumya Raychaudhuri, Paul I. W. de Bakker, Eleazar Eskin |
Bioinform. | 2 |
| 2013 | A powerful and efficient set test for genetic markers that handles confoundersabstractMOTIVATION: Approaches for testing sets of variants, such as a set of rare or common variants within a gene or pathway, for association with complex traits are important. In particular, set tests allow for aggregation of weak signal within a set, can capture interplay among variants and reduce the burden of multiple hypothesis testing. Until now, these approaches did not address confounding by family relatedness and population structure, a problem that is becoming more important as larger datasets are used to increase power. RESULTS: We introduce a new approach for set tests that handles confounders. Our model is based on the linear mixed model and uses two random effects-one to capture the set association signal and one to capture confounders. We also introduce a computational speedup for two random-effects models that makes this approach feasible even for extremely large cohorts. Using this model with both the likelihood ratio test and score test, we find that the former yields more power while controlling type I error. Application of our approach to richly structured Genetic Analysis Workshop 14 data demonstrates that our method successfully corrects for population structure and family relatedness, whereas application of our method to a 15 000 individual Crohn's disease case-control cohort demonstrates that it additionally recovers genes not recoverable by univariate analysis. AVAILABILITY: A Python-based library implementing our approach is available at http://mscompbio.codeplex.com. Jennifer Listgarten, Christoph Lippert, Eun Yong Kang, Jing Xiang, Carl Myers Kadie, David Heckerman |
Bioinform. | 3 |
| 2010 | Respecting Markov Equivalence in Computing Posterior Probabilities of Causal Graphical FeaturesabstractThere have been many efforts to identify causal graphical features such as directed edges between random variables from observational data. Recently, Tian et al. proposed a new dynamic programming algorithm which computes marginalized posterior probabilities of directed edge features over all the possible structures in O(n3n) time when the number of parents per node is bounded by a constant, where n is the number of variables of interest. However the main drawback of this approach is that deciding a single appropriate threshold for the existence of the directed edge feature is difficult due to the scale difference of the posterior probabilities between the directed edges forming v-structures and the directed edges not forming v-structures. We claim that computing posterior probabilities of both adjacencies and v-structures is necessary and more effective for discovering causal graphical features, since it allows us to find a single appropriate decision threshold for the existence of the feature that we are testing. For efficient computation, we provide a novel dynamic programming algorithm which computes the posterior probabilities of all of n(n – 1)/2 adjacency and n(n–1 choose 2) v-structure features in O(n3 * 3n) time. Eun Yong Kang, Ilya Shpitser, Eleazar Eskin |
AAAI | 1 |
| 2009 | Detecting the Presence and Absence of Causal Relationships between Expression of Yeast Genes with Very Few Samples
Eun Yong Kang, Ilya Shpitser, Chun Ye, Eleazar Eskin |
RECOMB | 1 |
| 2008 | Automatically Extracting Form LabelsabstractWe describe a machine-learning-based approach for extracting attribute labels from Web form interfaces. Having these labels is a requirement for several techniques that attempt to retrieve and integrate data that reside in online databases and that are hidden behind form interfaces, including schema matching and clustering, and hidden-Web crawlers. Whereas previous approaches to this problem have relied on heuristics and manually specified extraction rules, our technique makes use of learning classifiers to identify form labels. Our preliminary experiments show this approach is promising and has high accuracy. Hoa Nguyen, Eun Yong Kang, Juliana Freire |
ICDE | 2 |