EDBT 2026 Demo / reviewers in the wild / expert
Taehyun Hwang
dblp:21/2077 · also Tae Hyun Hwang, Tae-Hyun Hwang, TaeHyun Hwang
· DBLP profile ↗
23ranked-venue papers
10as first author
6since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 4 first-authorArtificial intelligence and machine learning · 7 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 1 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Reinforcement learning · 97% Graph learning · 3% | |
| Theoretical computer science
1 paper |
Algorithmic game theory and mechanism design · 60% Approximation and online algorithms · 40% | |
| Interdisciplinary, comprehensive, and emerging computing
7 papers |
Bioinformatics and computational biology · 94% Medical and health informatics · 6% |
Topics — the 30 heaviest of 38, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › exploration
exploration-exploitation tradeoff |
1.3 | 2 | 2023 | Combinatorial Neural Bandits · ICML 2023 Model-Based Reinforcement Learning with Multinomial Logistic Function Approximation · AAAI 2023 |
Machine learning › Reinforcement learning
exploration |
1.0 | 2 | 2025 | Randomized Exploration for Reinforcement Learning with Multinomial Logistic Function Approximation · NeurIPS 2024 Lasso Bandit with Compatibility Condition on Optimal Arm · ICLR 2025 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
0.9 | 1 | 2025 | Biological Pathway Guided Gene Selection Through Collaborative Reinforcement Learning · KDD (2) 2025 |
Bioinformatics and computational biology › gene expression analysis
gene selection |
0.9 | 1 | 2025 | Biological Pathway Guided Gene Selection Through Collaborative Reinforcement Learning · KDD (2) 2025 |
Approximation and online algorithms › online algorithms
bandit algorithms |
0.9 | 1 | 2025 | Lasso Bandit with Compatibility Condition on Optimal Arm · ICLR 2025 |
Algorithmic game theory and mechanism design › multi-armed bandit
linear bandits |
0.9 | 1 | 2025 | Lasso Bandit with Compatibility Condition on Optimal Arm · ICLR 2025 |
Algorithmic game theory and mechanism design
multi-armed bandit |
0.9 | 1 | 2025 | Lasso Bandit with Compatibility Condition on Optimal Arm · ICLR 2025 |
Approximation and online algorithms
online algorithms |
0.9 | 1 | 2025 | Lasso Bandit with Compatibility Condition on Optimal Arm · ICLR 2025 |
Algorithmic game theory and mechanism design › multi-armed bandit › linear bandits
sparse linear bandits |
0.9 | 1 | 2025 | Lasso Bandit with Compatibility Condition on Optimal Arm · ICLR 2025 |
Machine learning › Reinforcement learning
function approximation |
0.8 | 1 | 2024 | Randomized Exploration for Reinforcement Learning with Multinomial Logistic Function Approximation · NeurIPS 2024 |
Machine learning › Reinforcement learning › exploration
randomized exploration |
0.8 | 1 | 2024 | Randomized Exploration for Reinforcement Learning with Multinomial Logistic Function Approximation · NeurIPS 2024 |
Machine learning › Reinforcement learning › multi-armed bandit
combinatorial bandits |
0.7 | 1 | 2023 | Combinatorial Neural Bandits · ICML 2023 |
Machine learning › Reinforcement learning › bandit
contextual bandit |
0.7 | 1 | 2023 | Combinatorial Neural Bandits · ICML 2023 |
Machine learning › Reinforcement learning
model-based reinforcement learning |
0.7 | 1 | 2023 | Model-Based Reinforcement Learning with Multinomial Logistic Function Approximation · AAAI 2023 |
Machine learning › Reinforcement learning › bandit › parametric bandits
neural bandit |
0.7 | 1 | 2023 | Combinatorial Neural Bandits · ICML 2023 |
Machine learning › Reinforcement learning
thompson sampling |
0.7 | 1 | 2023 | Combinatorial Neural Bandits · ICML 2023 |
Machine learning › Reinforcement learning › bandit
upper confidence bound |
0.7 | 1 | 2023 | Model-Based Reinforcement Learning with Multinomial Logistic Function Approximation · AAAI 2023 |
Machine learning › Graph learning
graph neural network |
0.3 | 1 | 2025 | Biological Pathway Guided Gene Selection Through Collaborative Reinforcement Learning · KDD (2) 2025 |
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
state representation |
0.3 | 1 | 2025 | Biological Pathway Guided Gene Selection Through Collaborative Reinforcement Learning · KDD (2) 2025 |
Bioinformatics and computational biology
cancer genomics |
0.2 | 1 | 2016 | An integrative somatic mutation analysis to identify pathways linked with survival outcomes across 19 cancer types · Bioinform. 2016 |
Bioinformatics and computational biology
biomarker discovery |
0.2 | 2 | 2008 | Robust and efficient identification of biomarkers by classifying features on graphs · Bioinform. 2008 Learning on Weighted Hypergraphs to Integrate Protein Interactions and Gene Expressions for Cancer Outcome Prediction · ICDM 2008 |
Bioinformatics and computational biology › biomarker discovery
disease-gene association |
0.1 | 1 | 2011 | Inferring disease and gene set associations with rank coherence in networks · Bioinform. 2011 |
Bioinformatics and computational biology
data integration |
0.1 | 1 | 2009 | A hypergraph-based learning algorithm for classifying gene expression and arrayCGH data with prior knowledge · Bioinform. 2009 |
Bioinformatics and computational biology › genomics › genomic data analysis
genomic classification |
0.1 | 1 | 2009 | A hypergraph-based learning algorithm for classifying gene expression and arrayCGH data with prior knowledge · Bioinform. 2009 |
Bioinformatics and computational biology › data integration
prior knowledge integration |
0.1 | 1 | 2009 | A hypergraph-based learning algorithm for classifying gene expression and arrayCGH data with prior knowledge · Bioinform. 2009 |
Bioinformatics and computational biology › knowledge representation in biology › biomedical ontology
gene ontology |
0.1 | 1 | 2017 | Transfer learning across ontologies for phenome-genome association prediction · Bioinform. 2017 |
Bioinformatics and computational biology › protein function prediction
network-based function prediction |
0.1 | 1 | 2017 | Transfer learning across ontologies for phenome-genome association prediction · Bioinform. 2017 |
Medical and health informatics
cancer prognosis |
0.1 | 1 | 2008 | Learning on Weighted Hypergraphs to Integrate Protein Interactions and Gene Expressions for Cancer Outcome Prediction · ICDM 2008 |
Bioinformatics and computational biology › genomics › computational genomics
gene prioritization |
0.1 | 1 | 2008 | Learning on Weighted Hypergraphs to Integrate Protein Interactions and Gene Expressions for Cancer Outcome Prediction · ICDM 2008 |
Bioinformatics and computational biology
multi-omics data integration |
0.1 | 1 | 2008 | Learning on Weighted Hypergraphs to Integrate Protein Interactions and Gene Expressions for Cancer Outcome Prediction · ICDM 2008 |
Methods — techniques the papers use, named apart from their topics
lasso · 2.6statistical filtering · 1.7multi-agent reinforcement learning · 1.7graph neural network · 1.7forced sampling · 1.7upper confidence bound · 1.3LASSO · 0.9optimistic sampling · 0.8local gradient information · 0.8neural tangent kernel · 0.7multinomial logistic function approximation · 0.7transfer learning · 0.3protein-protein interaction network · 0.3dual label propagation · 0.3semi-supervised learning · 0.3gene network integration · 0.2consensus clustering · 0.2simulation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Lasso Bandit with Compatibility Condition on Optimal ArmabstractWe consider a stochastic sparse linear bandit problem where only a sparse subset of context features affects the expected reward function, i.e., the unknown reward parameter has a sparse structure.
In the existing Lasso bandit literature, the compatibility conditions, together with additional diversity conditions on the context features are imposed to achieve regret bounds that only depend logarithmically on the ambient dimension $d$.
In this paper, we demonstrate that even without the additional diversity assumptions, the \textit{compatibility condition on the optimal arm} is sufficient to derive a regret bound that depends logarithmically on $d$, and our assumption is strictly weaker than those used in the lasso bandit literature under the single-parameter setting.
We propose an algorithm that adapts the forced-sampling technique and prove that the proposed algorithm achieves $\mathcal{O}(\text{poly}\log dT)$ regret under the margin condition.
To our knowledge, the proposed algorithm requires the weakest assumptions among Lasso bandit algorithms under the single-parameter setting that achieve $\mathcal{O}(\text{poly}\log dT)$ regret.
Through numerical experiments, we confirm the superior performance of our proposed algorithm. Harin Lee, Taehyun Hwang, Min-hwan Oh |
ICLR | 2 |
| 2025 | Biological Pathway Guided Gene Selection Through Collaborative Reinforcement LearningabstractGene selection in high-dimensional genomic data is essential for understanding disease mechanisms and improving therapeutic outcomes. Traditional feature selection methods effectively identify predictive genes but often ignore complex biological pathways and regulatory networks, leading to unstable and biologically irrelevant signatures. Prior approaches, such as Lasso-based methods and statistical filtering, either focus solely on individual gene-outcome associations or fail to capture pathway-level interactions, presenting a key challenge: how to integrate biological pathway knowledge while maintaining statistical rigor in gene selection? To address this gap, we propose a novel two-stage framework that integrates statistical selection with biological pathway knowledge using multi-agent reinforcement learning (MARL). First, we introduce a pathway-guided pre-filtering strategy that leverages multiple statistical methods alongside KEGG pathway information for initial dimensionality reduction. Next, for refined selection, we model genes as collaborative agents in a MARL framework, where each agent optimizes both predictive power and biological relevance. Our framework incorporates pathway knowledge through Graph Neural Network-based state representations, a reward mechanism combining prediction performance with gene centrality and pathway coverage, and collaborative learning strategies using shared memory and a centralized critic component. Extensive experiments on multiple gene expression datasets demonstrate that our approach significantly improves both prediction accuracy and biological interpretability compared to traditional methods. Ehtesamul Azim, Dongjie Wang 0001, Taehyun Hwang, Yanjie Fu, Wei Zhang 0076 |
KDD (2) | 3 |
| 2025 | Tractable Multinomial Logit Contextual Bandits with Non-Linear UtilitiesabstractWe study the _multinomial logit_ (MNL) contextual bandit problem for sequential assortment selection. Although most existing research assumes utility functions to be linear in item features, this linearity assumption restricts the modeling of intricate interactions between items and user preferences. A recent work (Zhang & Luo, 2024) has investigated general utility function classes, yet its method faces fundamental trade-offs between computational tractability and statistical efficiency. To address this limitation, we propose a computationally efficient algorithm for MNL contextual bandits leveraging the upper confidence bound principle, specifically designed for non-linear parametric utility functions, including those modeled by neural networks. Under a realizability assumption and a mild geometric condition on the utility function class, our algorithm achieves a regret bound of $\tilde{\mathcal{O}}(\sqrt{T})$, where $T$ denotes the total number of rounds. Our result establishes that sharp $\tilde{\mathcal{O}}(\sqrt{T})$-regret is attainable even with neural network-based utilities, without relying on strong assumptions such as neural tangent kernel approximations. To the best of our knowledge, our proposed method is the first computationally tractable algorithm for MNL contextual bandits with non-linear utilities that provably attains $\tilde{\mathcal{O}}(\sqrt{T})$ regret. Comprehensive numerical experiments validate the effectiveness of our approach, showing robust performance not only in realizable settings but also in scenarios with model misspecification. Taehyun Hwang, Dahngoon Kim, Min-hwan Oh |
NeurIPS | 1 |
| 2024 | Randomized Exploration for Reinforcement Learning with Multinomial Logistic Function ApproximationabstractWe study reinforcement learning with _multinomial logistic_ (MNL) function approximation where the underlying transition probability kernel of the _Markov decision processes_ (MDPs) is parametrized by an unknown transition core with features of state and action. For the finite horizon episodic setting with inhomogeneous state transitions, we propose provably efficient algorithms with randomized exploration having frequentist regret guarantees. For our first algorithm, $\texttt{RRL-MNL}$, we adapt optimistic sampling to ensure the optimism of the estimated value function with sufficient frequency and establish that $\texttt{RRL-MNL}$ is both _statistically_ and _computationally_ efficient, achieving a $\tilde{\mathcal{O}}(\kappa^{-1} d^{\frac{3}{2}} H^{\frac{3}{2}} \sqrt{T})$ frequentist regret bound with constant-time computational cost per episode. Here, $d$ is the dimension of the transition core, $H$ is the horizon length, $T$ is the total number of steps, and $\kappa$ is a problem-dependent constant. Despite the simplicity and practicality of $\texttt{RRL-MNL}$, its regret bound scales with $\kappa^{-1}$, which is potentially large in the worst case. To improve the dependence on $\kappa^{-1}$, we propose $\texttt{ORRL-MNL}$, which estimates the value function using local gradient information of the MNL transition model. We show that its frequentist regret bound is $\tilde{\mathcal{O}}(d^{\frac{3}{2}} H^{\frac{3}{2}} \sqrt{T} + \kappa^{-1} d^2 H^2)$. To the best of our knowledge, these are the first randomized RL algorithms for the MNL transition model that achieve both computational and statistical efficiency. Numerical experiments demonstrate the superior performance of the proposed algorithms. Wooseong Cho, Taehyun Hwang, Joongkyu Lee, Min-hwan Oh |
NeurIPS | 2 |
| 2023 | Model-Based Reinforcement Learning with Multinomial Logistic Function ApproximationabstractWe study model-based reinforcement learning (RL) for episodic Markov decision processes (MDP) whose transition probability is parametrized by an unknown transition core with features of state and action. Despite much recent progress in analyzing algorithms in the linear MDP setting, the understanding of more general transition models is very restrictive. In this paper, we propose a provably efficient RL algorithm for the MDP whose state transition is given by a multinomial logistic model. We show that our proposed algorithm based on the upper confidence bounds achieves O(d√(H^3 T)) regret bound where d is the dimension of the transition core, H is the horizon, and T is the total number of steps. To the best of our knowledge, this is the first model-based RL algorithm with multinomial logistic function approximation with provable guarantees. We also comprehensively evaluate our proposed algorithm numerically and show that it consistently outperforms the existing methods, hence achieving both provable efficiency and practical superior performance. Taehyun Hwang, Min-hwan Oh |
AAAI | 1 |
| 2023 | Combinatorial Neural BanditsabstractWe consider a contextual combinatorial bandit problem where in each round a learning agent selects a subset of arms and receives feedback on the selected arms according to their scores. The score of an arm is an unknown function of the arm’s feature. Approximating this unknown score function with deep neural networks, we propose algorithms: Combinatorial Neural UCB ($\texttt{CN-UCB}$) and Combinatorial Neural Thompson Sampling ($\texttt{CN-TS}$). We prove that $\texttt{CN-UCB}$ achieves $\tilde{\mathcal{O}}(\tilde{d} \sqrt{T})$ or $\tilde{\mathcal{O}}(\sqrt{\tilde{d} T K})$ regret, where $\tilde{d}$ is the effective dimension of a neural tangent kernel matrix, $K$ is the size of a subset of arms, and $T$ is the time horizon. For $\texttt{CN-TS}$, we adapt an optimistic sampling technique to ensure the optimism of the sampled combinatorial action, achieving a worst-case (frequentist) regret of $\tilde{\mathcal{O}}(\tilde{d} \sqrt{TK})$. To the best of our knowledge, these are the first combinatorial neural bandit algorithms with regret performance guarantees. In particular, $\texttt{CN-TS}$ is the first Thompson sampling algorithm with the worst-case regret guarantees for the general contextual combinatorial bandit problem. The numerical experiments demonstrate the superior performances of our proposed algorithms. Taehyun Hwang, Kyuwook Chai, Min-hwan Oh |
ICML | 1 |
| 2020 | Computerized Classification of Prostate Cancer Gleason Scores from Whole Slide ImagesabstractHistological Gleason grading of tumor patterns is one of the most powerful prognostic predictors in prostate cancer. However, manual analysis and grading performed by pathologists are typically subjective and time-consuming. In this paper, we present an automatic technique for Gleason grading of prostate cancer from H&E stained whole slide pathology images using a set of novel completed and statistical local binary pattern (CSLBP) descriptors. First, the technique divides the whole slide image (WSI) into a set of small image tiles, where salient tumor tiles with high nuclei densities are selected for analysis. The CSLBP texture features that encode pixel intensity variations from circularly surrounding neighborhoods are extracted from salient image tiles to characterize different Gleason patterns. Finally, the CSLBP texture features computed from all tiles are integrated and utilized by the multi-class support vector machine (SVM) that assigns patient slides with different Gleason scores such as 6, 7, or ≥ 8. Experiments have been performed on 312 different patient cases selected from the cancer genome atlas (TCGA) and have achieved superior performances over state-of-the-art texture descriptors and baseline methods including deep learning models for prostate cancer Gleason grading. Hongming Xu 0002, Sunho Park, Taehyun Hwang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2018 | "I'll do it!": examining the relationship between locus of control and math game retention for preschoolersabstractAcquiring simple arithmetic skills at the preschool level requires repetitive practices. One method for encouraging students to spend longer time practicing is by presenting the skills in an engaging game. As student retention on the game increases, the student will be more likely to acquire the practiced skill since she will have spent more time practicing. In this paper, we examine the relationship between internal locus of control and retention in game-based learning applications for young children using Todo Math, a mobile-based math learning application for children from Pre-K to 2nd grade. We examine 345,783 users' log data to show that when children prefer "free" mode, which has high internal locus of control, their retention on Todo Math is higher than children who prefer "daily" mode, which has high external locus of control. We present three analyses that support our findings using survival analysis, post-hoc analysis, and t-test. Bugeun Kim, Jungwook Rhim, Jihyun Rho, Taehyun Hwang, Gunho Lee, Gahgene Gweon |
LAK | 4 |
| 2017 | Transfer learning across ontologies for phenome-genome association predictionabstractMotivation: To better predict and analyze gene associations with the collection of phenotypes organized in a phenotype ontology, it is crucial to effectively model the hierarchical structure among the phenotypes in the ontology and leverage the sparse known associations with additional training information. In this paper, we first introduce Dual Label Propagation (DLP) to impose consistent associations with the entire phenotype paths in predicting phenotype-gene associations in Human Phenotype Ontology (HPO). DLP is then used as the base model in a transfer learning framework (tlDLP) to incorporate functional annotations in Gene Ontology (GO). By simultaneously reconstructing GO term-gene associations and HPO phenotype-gene associations for all the genes in a protein-protein interaction network, tlDLP benefits from the enriched training associations indirectly through relation with GO terms. Results: In the experiments to predict the associations between human genes and phenotypes in HPO based on human protein-protein interaction network, both DLP and tlDLP improved the prediction of gene associations with phenotype paths in HPO in cross-validation and the prediction of the most recent associations added after the snapshot of the training data. Moreover, the transfer learning through GO term-gene associations significantly improved association predictions for the phenotypes with no more specific known associations by a large margin. Examples are also shown to demonstrate how phenotype paths in phenotype ontology and transfer learning with gene ontology can improve the predictions. Availability and Implementation: Source code is available at http://compbio.cs.umn.edu/onto phenome . Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Raphael Petegrosso, Sunho Park, Taehyun Hwang, Rui Kuang |
Bioinform. | 3 |
| 2016 | An integrative somatic mutation analysis to identify pathways linked with survival outcomes across 19 cancer typesabstractMOTIVATION: Identification of altered pathways that are clinically relevant across human cancers is a key challenge in cancer genomics. Precise identification and understanding of these altered pathways may provide novel insights into patient stratification, therapeutic strategies and the development of new drugs. However, a challenge remains in accurately identifying pathways altered by somatic mutations across human cancers, due to the diverse mutation spectrum. We developed an innovative approach to integrate somatic mutation data with gene networks and pathways, in order to identify pathways altered by somatic mutations across cancers. RESULTS: We applied our approach to The Cancer Genome Atlas (TCGA) dataset of somatic mutations in 4790 cancer patients with 19 different types of tumors. Our analysis identified cancer-type-specific altered pathways enriched with known cancer-relevant genes and targets of currently available drugs. To investigate the clinical significance of these altered pathways, we performed consensus clustering for patient stratification using member genes in the altered pathways coupled with gene expression datasets from 4870 patients from TCGA, and multiple independent cohorts confirmed that the altered pathways could be used to stratify patients into subgroups with significantly different clinical outcomes. Of particular significance, certain patient subpopulations with poor prognosis were identified because they had specific altered pathways for which there are available targeted therapies. These findings could be used to tailor and intensify therapy in these patients, for whom current therapy is suboptimal. AVAILABILITY AND IMPLEMENTATION: The code is available at: http://www.taehyunlab.org CONTACT: [email protected] or [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Sunho Park, Seung-Jun Kim 0002, Donghyeon Yu, Samuel Peña-Llopis, Jianjiong Gao, Jinsuk Park, Jessie Norris, Xinlei Wang 0001, Min Chen 0014, Jeongsik Yong, Zabi Wardak, Kevin Choe, Michael Story, Timothy K. Starr, Jae-Ho Cheong, Taehyun Hwang |
Bioinform. | 18 |
| 2014 | Convex Optimization for Binary Classifier Aggregation in Multiclass ProblemsabstractMulticlass problems are often decomposed into multiple binary problems that are solved by individual binary classifiers whose results are integrated into a final answer. Various methods, including all-pairs (APs), one-versus-all (OVA), and error correcting output code (ECOC), have been studied, to decompose multiclass problems into binary problems. However, little study has been made to optimally aggregate binary problems to determine a final answer to the multiclass problem. In this paper we present a convex optimization method for an optimal aggregation of binary classifiers to estimate class membership probabilities in multiclass problems. We model the class membership probability as a softmax function which takes a conic combination of discrepancies induced by individual binary classifiers, as an input. With this model, we formulate the regularized maximum likelihood estimation as a convex optimization problem, which is solved by the primal-dual interior point method. Connections of our method to large margin classifiers are presented, showing that the large margin formulation can be considered as a limiting case of our convex formulation. In the experiments on human disease classification, we demonstrate that our method outperforms existing aggregation methods as well as direct methods, in terms of the classification accuracy and F-score. Sunho Park, Taehyun Hwang, Seungjin Choi 0001 |
SDM | 2 |
| 2012 | Prioritizing Disease Genes by Bi-Random Walk
Maoqiang Xie, Taehyun Hwang, Rui Kuang |
PAKDD (2) | 2 |
| 2011 | Inferring disease and gene set associations with rank coherence in networksabstractMOTIVATION: To validate the candidate disease genes identified from high-throughput genomic studies, a necessary step is to elucidate the associations between the set of candidate genes and disease phenotypes. The conventional gene set enrichment analysis often fails to reveal associations between disease phenotypes and the gene sets with a short list of poorly annotated genes, because the existing annotations of disease-causative genes are incomplete. This article introduces a network-based computational approach called rcNet to discover the associations between gene sets and disease phenotypes. A learning framework is proposed to maximize the coherence between the predicted phenotype-gene set relations and the known disease phenotype-gene associations. An efficient algorithm coupling ridge regression with label propagation and two variants are designed to find the optimal solution to the objective functions of the learning framework. RESULTS: We evaluated the rcNet algorithms with leave-one-out cross-validation on Online Mendelian Inheritance in Man (OMIM) data and an independent test set of recently discovered disease-gene associations. In the experiments, the rcNet algorithms achieved best overall rankings compared with the baselines. To further validate the reproducibility of the performance, we applied the algorithms to identify the target diseases of novel candidate disease genes obtained from recent studies of Genome-Wide Association Study (GWAS), DNA copy number variation analysis and gene expression profiling. The algorithms ranked the target disease of the candidate genes at the top of the rank list in many cases across all the three case studies. AVAILABILITY: http://compbio.cs.umn.edu/dgsa_rcNet CONTACT: [email protected]. Taehyun Hwang, Wei Zhang 0076, Maoqiang Xie, Jinfeng Liu 0003, Rui Kuang |
Bioinform. | 1 |
| 2010 | A Heterogeneous Label Propagation Algorithm for Disease Gene DiscoveryabstractLabel propagation is an effective and efficient technique to utilize local and global features in a network for semi-supervised learning. In the literature, one challenge is how to propagate information in heterogeneous networks comprising several subnetworks, each of which has its own cluster structures that need to be explored independently. In this paper, we introduce an intutitive algorithm MINProp (Mutual Interaction-based Network Propagation) and a simple regularization framework for propagating information between subnetworks in a heterogeneous network. MINProp sequentially performs label propagation on each individual subnetwork with the current label information derived from the other subnetworks and repeats this step until convergence to the global optimal solution to the convex objective function of the regularization framework. The independent label propagation on each subnetwork explores the cluster structure in the subnetwork. The label information from the other subnetworks is used to capture mutual interactions (bicluster structures) between the vertices in each pair of the subnetworks. MINProp algorithm is applied to disease gene discovery from a heterogeneus network of disease phenotypes and genes. In the experiments, MINProp significantly output-performed the original label propagation algorithm on a single network and the state-of-the-art methods for discovering disease genes. The results also suggest that MINProp is more effective in utilizing the modular structures in a heterogenous network. Finally, MINProp discovered new disease-gene associations that are only reported recently. Taehyun Hwang, Rui Kuang |
SDM | 1 |
| 2009 | A hypergraph-based learning algorithm for classifying gene expression and arrayCGH data with prior knowledgeabstractMOTIVATION: Incorporating biological prior knowledge into predictive models is a challenging data integration problem in analyzing high-dimensional genomic data. We introduce a hypergraph-based semi-supervised learning algorithm called HyperPrior to classify gene expression and array-based comparative genomic hybridization (arrayCGH) data using biological knowledge as constraints on graph-based learning. HyperPrior is a robust two-step iterative method that alternatively finds the optimal labeling of the samples and the optimal weighting of the features, guided by constraints encoding prior knowledge. The prior knowledge for analyzing gene expression data is that cancer-related genes tend to interact with each other in a protein-protein interaction network. Similarly, the prior knowledge for analyzing arrayCGH data is that probes that are spatially nearby in their layout along the chromosomes tend to be involved in the same amplification or deletion event. Based on the prior knowledge, HyperPrior imposes a consistent weighting of the correlated genomic features in graph-based learning. RESULTS: We applied HyperPrior to test two arrayCGH datasets and two gene expression datasets for both cancer classification and biomarker identification. On all the datasets, HyperPrior achieved competitive classification performance, compared with SVMs and the other baselines utilizing the same prior knowledge. HyperPrior also identified several discriminative regions on chromosomes and discriminative subnetworks in the PPI, both of which contain cancer-related genomic elements. Our results suggest that HyperPrior is promising in utilizing biological prior knowledge to achieve better classification performance and more biologically interpretable findings in gene expression and arrayCGH data. AVAILABILITY: http://compbio.cs.umn.edu/HyperPrior CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at bioinformatics online. Ze Tian, Taehyun Hwang, Rui Kuang |
Bioinform. | 2 |
| 2008 | Learning on Weighted Hypergraphs to Integrate Protein Interactions and Gene Expressions for Cancer Outcome PredictionabstractBuilding reliable predictive models from multiple complementary genomic data for cancer study is a crucial step towards successful cancer treatment and a full understanding of the underlying biological principles. To tackle this challenging data integration problem, we propose a hypergraph-based learning algorithm called HyperGene to integrate microarray gene expressions and protein-protein interactions for cancer outcome prediction and biomarker identification. HyperGene is a robust two-step iterative method that alternatively finds the optimal outcome prediction and the optimal weighting of the marker genes guided by a protein-protein interaction network. Under the hypothesis that cancer-related genes tend to interact with each other, the HyperGene algorithm uses a protein-protein interaction network as prior knowledge by imposing a consistent weighting of interacting genes. Our experimental results on two large-scale breast cancer gene expression datasets show that HyperGene utilizing a curated protein-protein interaction network achieves significantly improved cancer outcome prediction. Moreover, HyperGene can also retrieve many known cancer genes as highly weighted marker genes. Taehyun Hwang, Ze Tian, Rui Kuang, Jean-Pierre A. Kocher |
ICDM | 1 |
| 2008 | Robust and efficient identification of biomarkers by classifying features on graphsabstractMOTIVATION: A central problem in biomarker discovery from large-scale gene expression or single nucleotide polymorphism (SNP) data is the computational challenge of taking into account the dependence among all the features. Methods that ignore the dependence usually identify non-reproducible biomarkers across independent datasets. We introduce a new graph-based semi-supervised feature classification algorithm to identify discriminative disease markers by learning on bipartite graphs. Our algorithm directly classifies the feature nodes in a bipartite graph as positive, negative or neutral with network propagation to capture the dependence among both samples and features (clinical and genetic variables) by exploring bi-cluster structures in a graph. Two features of our algorithm are: (1) our algorithm can find a global optimal labeling to capture the dependence among all the features and thus, generates highly reproducible results across independent microarray or other high-thoughput datasets, (2) our algorithm is capable of handling hundreds of thousands of features and thus, is particularly useful for biomarker identification from high-throughput gene expression and SNP data. In addition, although designed for classifying features, our algorithm can also simultaneously classify test samples for disease prognosis/diagnosis. RESULTS: We applied the network propagation algorithm to study three large-scale breast cancer datasets. Our algorithm achieved competitive classification performance compared with SVMs and other baseline methods, and identified several markers with clinical or biological relevance with the disease. More importantly, our algorithm also identified highly reproducible marker genes and enriched functions from the independent datasets. AVAILABILITY: Supplementary results and source code are available at http://compbio.cs.umn.edu/Feature_Class. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Taehyun Hwang, Hugues Sicotte, Ze Tian, Baolin Wu, Jean-Pierre A. Kocher, Dennis A. Wigle, Vipin Kumar 0001, Rui Kuang |
Bioinform. | 1 |
| 2007 | MCTA: Target Tracking Algorithm Based on Minimal Contour in Wireless Sensor NetworksabstractThis paper proposes a minimal contour tracking algorithm (MCTA) that reduces energy consumption for tracking mobile targets in wireless sensor networks in terms of sensing and communication energy consumption. MCTA conserves energy by letting only a minimum number of sensor nodes participate in communication and perform sensing for target tracking. MCTA uses the minimal tracking area based on the vehicular kinematics. The modeling of target's kinematics allows for pruning out part of the tracking area that cannot be mechanically visited by the mobile target within scheduled time. So, MCTA sends the tracking area information to only the sensor nodes within minimal tracking area and wakes them up. Compared to the legacy scheme which uses circle-based tracking area, our proposed scheme uses less number of sensors for tracking in both communication and sensing without target missing. Through simulation, we show that MCTA outperforms the circle-based scheme with about 60% energy saving under certain ideal situations. Jaehoon Jeong 0001, Taehyun Hwang, Tian He 0001, David Hung-Chang Du |
INFOCOM | 2 |
| 2006 | Detection of Traffic Lights for Vision-Based Car Navigation System
Taehyun Hwang, In-Hak Joo, Seong Ik Cho |
PSIVT | 1 |
| 2005 | Updating Geospatial Database: An Automatic Approach Combining Photogrammetry and Computer Vision Techniques
In-Hak Joo, Taehyun Hwang, Kyoung-Ho Choi |
ACIVS | 2 |
| 2004 | Generation of video metadata supporting video-gis integration
In-Hak Joo, Taehyun Hwang, Kyoung-Ho Choi |
ICIP | 2 |
| 2004 | An indexing method for spatial object in video using image processing and photogrammetryabstractTo support video-map cross-reference and bi-directional search, it is necessary to index spatial object with regard to each video frame. However, it costs a great deal of effort to index spatial objects in all frames of video whose frame rate is about 30 frames/second. Therefore, we index spatial object only about 1 random frame per second, and index the spatial object in other frames by interpolation. However, when processing video frames with 1-second interval, the low spatial redundancy of spatial objects in the video makes the indexing of spatial objects difficult because object tracking is nearly impossible only with image-based method. We devise an efficient indexing method for spatial objects in video frame, by combining photogrammetric solution with image-based method. The suggested method can efficiently index spatial object in video frames and support content-based retrieval of spatial object in video. Taehyun Hwang, In-Hak Joo, Kyoung-Ho Choi |
IGARSS | 1 |
| 2003 | MPEG-7 metadata for video-based GIS applicationsabstractIn this paper, we present a MPEG-7 metadata scheme designed for a video-based Geographic Information System (VideoGIS). Recently, VideoGIS have been developed for integrated management of spatial data such as map, 3D graphic, still image, and video to provide better and more realistic information of spatial objects. Main functions of VideoGIS are to calculate 3D geographic position of a spatial object and to search image and video that correspond to the object. In this paper, a MPEG-7 metadata scheme is proposed for emerging Location-Based Services (LBS), e.g., tourist information and restaurant reservation services, which becomes more and more popular in emerging mobile environments. Taehyun Hwang, Kyoung-Ho Choi, In-Hak Joo, Jong-Hun Lee |
IGARSS | 1 |