Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Eun Yong Kang

dblp:38/1672 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
5 papers
Bioinformatics and computational biology · 100%
Databases, data mining, and information retrieval
2 papers
Recommender systems · 90% Data integration and cleaning · 10%
Theoretical computer science
1 paper
Algorithms and data structures · 50% Mathematical optimization · 50%
Artificial intelligence
1 paper
Probabilistic and Bayesian machine learning · 67% Knowledge representation and reasoning · 33%

Topics — the 17 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Recommender systems › diversified recommendation
determinantal point process
0.512021
Diversity on the Go! Streaming Determinantal Point Processes under a Maximum Induced Cardinality Objective · WWW 2021
Recommender systems
diversified recommendation
0.512021
Diversity on the Go! Streaming Determinantal Point Processes under a Maximum Induced Cardinality Objective · WWW 2021
Algorithms and data structures › data streams
streaming algorithms
0.512021
Diversity on the Go! Streaming Determinantal Point Processes under a Maximum Induced Cardinality Objective · WWW 2021
Mathematical optimization › submodular optimization › submodular maximization
streaming submodular maximization
0.512021
Diversity on the Go! Streaming Determinantal Point Processes under a Maximum Induced Cardinality Objective · WWW 2021
Bioinformatics and computational biology
statistical genetics
0.422014
Fast pairwise IBD association testing in genome-wide association studies · Bioinform. 2014
A powerful and efficient set test for genetic markers that handles confounders · Bioinform. 2013
Bioinformatics and computational biology › functional genomics
eQTL mapping
0.312017
Applying meta-analysis to genotype-tissue expression data from multiple tissues to identify eQTLs and increase the number of eGenes · Bioinform. 2017
Bioinformatics and computational biology
genomics
0.312017
Applying meta-analysis to genotype-tissue expression data from multiple tissues to identify eQTLs and increase the number of eGenes · Bioinform. 2017
Bioinformatics and computational biology › population genetics
population stratification correction
0.212015
Efficient and Accurate Multiple-Phenotypes Regression Method for High Dimensional Data Considering Population Structure · RECOMB 2015
Bioinformatics and computational biology › biostatistics › statistical bioinformatics
statistical genomics
0.212015
Efficient and Accurate Multiple-Phenotypes Regression Method for High Dimensional Data Considering Population Structure · RECOMB 2015
Bioinformatics and computational biology › genomics
genome-wide association study
0.212014
Fast pairwise IBD association testing in genome-wide association studies · Bioinform. 2014
Bioinformatics and computational biology › statistical genetics
genetic association study
0.212013
A powerful and efficient set test for genetic markers that handles confounders · Bioinform. 2013
Bioinformatics and computational biology › statistical genetics
linear mixed model
0.212013
A powerful and efficient set test for genetic markers that handles confounders · Bioinform. 2013
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › structure learning
bayesian network structure learning
0.112010
Respecting Markov Equivalence in Computing Posterior Probabilities of Causal Graphical Features · AAAI 2010
Knowledge, reasoning and agents › Knowledge representation and reasoning › causal reasoning
causal graphical model
0.112010
Respecting Markov Equivalence in Computing Posterior Probabilities of Causal Graphical Features · AAAI 2010
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal discovery
markov equivalence class
0.112010
Respecting Markov Equivalence in Computing Posterior Probabilities of Causal Graphical Features · AAAI 2010
Bioinformatics and computational biology
gene expression analysis
0.112009
Detecting the Presence and Absence of Causal Relationships between Expression of Yeast Genes with Very Few Samples · RECOMB 2009
Data integration and cleaning › data extraction
web data extraction
0.112008
Automatically Extracting Form Labels · ICDE 2008

Methods — techniques the papers use, named apart from their topics

greedy algorithm · 1.0determinantal point process · 1.0meta-analysis · 0.3mixed model · 0.2permutation testing · 0.2importance sampling · 0.2score test · 0.2linear mixed model · 0.2likelihood ratio test · 0.2dynamic programming · 0.1machine learning classification · 0.1
YearPublicationVenuePosition
2021 Diversity on the Go! Streaming Determinantal Point Processes under a Maximum Induced Cardinality Objective
abstract
Over the past decade, Determinantal Point Processes (DPPs) have proven to be a mathematically elegant framework for modeling diversity. Given a set of items N, DPPs define a probability distribution over subsets of N, with sets of larger diversity having greater probability. Recently, DPPs have achieved success in the domain of recommendation systems, as a method to enforce diversity of recommendations in addition to relevance. In large-scale recommendation applications however, the input typically comes in the form of a stream too large to fit into main memory. However, the natural greedy algorithm for DPP-based recommendations is memory intensive, and cannot be used in a streaming setting.
Paul Liu 0001, Akshay Soni, Eun Yong Kang, Yajun Wang 0001, Mehul Parsana
WWW3
2017 Applying meta-analysis to genotype-tissue expression data from multiple tissues to identify eQTLs and increase the number of eGenes
abstract
MOTIVATION: There is recent interest in using gene expression data to contextualize findings from traditional genome-wide association studies (GWAS). Conditioned on a tissue, expression quantitative trait loci (eQTLs) are genetic variants associated with gene expression, and eGenes are genes whose expression levels are associated with genetic variants. eQTLs and eGenes provide great supporting evidence for GWAS hits and important insights into the regulatory pathways involved in many diseases. When a significant variant or a candidate gene identified by GWAS is also an eQTL or eGene, there is strong evidence to further study this variant or gene. Multi-tissue gene expression datasets like the Gene Tissue Expression (GTEx) data are used to find eQTLs and eGenes. Unfortunately, these datasets often have small sample sizes in some tissues. For this reason, there have been many meta-analysis methods designed to combine gene expression data across many tissues to increase power for finding eQTLs and eGenes. However, these existing techniques are not scalable to datasets containing many tissues, like the GTEx data. Furthermore, these methods ignore a biological insight that the same variant may be associated with the same gene across similar tissues. RESULTS: We introduce a meta-analysis model that addresses these problems in existing methods. We focus on the problem of finding eGenes in gene expression data from many tissues, and show that our model is better than other types of meta-analyses. AVAILABILITY AND IMPLEMENTATION: Source code is at https://github.com/datduong/RECOV . CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Dat Duong, Lisa Gai, Sagi Snir, Eun Yong Kang, Buhm Han, Jae Hoon Sul, Eleazar Eskin
Bioinform.4
2015 Efficient and Accurate Multiple-Phenotypes Regression Method for High Dimensional Data Considering Population Structure
Jong Wha J. Joo, Eun Yong Kang, Elin Org, Nicholas A. Furlotte, Brian Parks, Aldons J. Lusis, Eleazar Eskin
RECOMB2
2014 Fast pairwise IBD association testing in genome-wide association studies
abstract
MOTIVATION: Recently, investigators have proposed state-of-the-art Identity-by-descent (IBD) mapping methods to detect IBD segments between purportedly unrelated individuals. The IBD information can then be used for association testing in genetic association studies. One approach for this IBD association testing strategy is to test for excessive IBD between pairs of cases ('pairwise method'). However, this approach is inefficient because it requires a large number of permutations. Moreover, a limited number of permutations define a lower bound for P-values, which makes fine-mapping of associated regions difficult because, in practice, a much larger genomic region is implicated than the region that is actually associated. RESULTS: In this article, we introduce a new pairwise method 'Fast-Pairwise'. Fast-Pairwise uses importance sampling to improve efficiency and enable approximation of extremely small P-values. Fast-Pairwise method takes only days to complete a genome-wide scan. In the application to the WTCCC type 1 diabetes data, Fast-Pairwise successfully fine-maps a known human leukocyte antigen gene that is known to cause the disease. AVAILABILITY: Fast-Pairwise is publicly available at: http://genetics.cs.ucla.edu/graphibd.
Buhm Han, Eun Yong Kang, Soumya Raychaudhuri, Paul I. W. de Bakker, Eleazar Eskin
Bioinform.2
2013 A powerful and efficient set test for genetic markers that handles confounders
abstract
MOTIVATION: Approaches for testing sets of variants, such as a set of rare or common variants within a gene or pathway, for association with complex traits are important. In particular, set tests allow for aggregation of weak signal within a set, can capture interplay among variants and reduce the burden of multiple hypothesis testing. Until now, these approaches did not address confounding by family relatedness and population structure, a problem that is becoming more important as larger datasets are used to increase power. RESULTS: We introduce a new approach for set tests that handles confounders. Our model is based on the linear mixed model and uses two random effects-one to capture the set association signal and one to capture confounders. We also introduce a computational speedup for two random-effects models that makes this approach feasible even for extremely large cohorts. Using this model with both the likelihood ratio test and score test, we find that the former yields more power while controlling type I error. Application of our approach to richly structured Genetic Analysis Workshop 14 data demonstrates that our method successfully corrects for population structure and family relatedness, whereas application of our method to a 15 000 individual Crohn's disease case-control cohort demonstrates that it additionally recovers genes not recoverable by univariate analysis. AVAILABILITY: A Python-based library implementing our approach is available at http://mscompbio.codeplex.com.
Jennifer Listgarten, Christoph Lippert, Eun Yong Kang, Jing Xiang, Carl Myers Kadie, David Heckerman
Bioinform.3
2010 Respecting Markov Equivalence in Computing Posterior Probabilities of Causal Graphical Features
abstract
There have been many efforts to identify causal graphical features such as directed edges between random variables from observational data. Recently, Tian et al. proposed a new dynamic programming algorithm which computes marginalized posterior probabilities of directed edge features over all the possible structures in O(n3n) time when the number of parents per node is bounded by a constant, where n is the number of variables of interest. However the main drawback of this approach is that deciding a single appropriate threshold for the existence of the directed edge feature is difficult due to the scale difference of the posterior probabilities between the directed edges forming v-structures and the directed edges not forming v-structures. We claim that computing posterior probabilities of both adjacencies and v-structures is necessary and more effective for discovering causal graphical features, since it allows us to find a single appropriate decision threshold for the existence of the feature that we are testing. For efficient computation, we provide a novel dynamic programming algorithm which computes the posterior probabilities of all of n(n – 1)/2 adjacency and n(n–1 choose 2) v-structure features in O(n3 * 3n) time.
Eun Yong Kang, Ilya Shpitser, Eleazar Eskin
AAAI1
2009 Detecting the Presence and Absence of Causal Relationships between Expression of Yeast Genes with Very Few Samples
Eun Yong Kang, Ilya Shpitser, Chun Ye, Eleazar Eskin
RECOMB1
2008 Automatically Extracting Form Labels
abstract
We describe a machine-learning-based approach for extracting attribute labels from Web form interfaces. Having these labels is a requirement for several techniques that attempt to retrieve and integrate data that reside in online databases and that are hidden behind form interfaces, including schema matching and clustering, and hidden-Web crawlers. Whereas previous approaches to this problem have relied on heuristics and manually specified extraction rules, our technique makes use of learning classifiers to identify form labels. Our preliminary experiments show this approach is promising and has high accuracy.
Hoa Nguyen, Eun Yong Kang, Juliana Freire
ICDE2