Mingon Kang

dblp:61/9077 · DBLP profile ↗
← Back
32ranked-venue papers
10as first author
9since 2021 · last 2026
0000-0002-9565-9523ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 27 · 8 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Prediction of bacterial protein-compound interactions with only positive samples
abstract
MOTIVATION: Prediction of Compound-Protein Interactions (CPI) in bacteria is crucial to advance various pharmaceutical and chemical engineering fields, including biocatalysis, drug discovery, and industrial processing. However, current CPI models cannot be applied for bacterial CPI prediction due to the lack of curated negative interaction samples. RESULTS: We propose a novel Positive-Unlabeled (PU) learning framework, named BIN-PU, to address this limitation. BIN-PU generates pseudo positive and negative labels from known positive interaction data, enabling effective training of deep learning models for CPI prediction. We also propose a weighted positive loss function that weights to truly positive samples. We have validated BIN-PU coupled with multiple CPI backbone models, comparing the performance with the existing PU models using bacterial cytochrome P450 (CYP) data. Extensive experiments demonstrate the superiority of BIN-PU over the benchmark models in predicting CPIs with only truly positive samples. Furthermore, we have validated BIN-PU on additional bacterial proteins obtained from literature review, human CYP datasets, and uncurated data for its reproducibility. We have also validated the CPI prediction for the uncurated CYP data with biological and biophysical experiments. BIN-PU represents a significant advancement in CPI prediction for bacterial proteins, opening new possibilities for improving predictive models in related biological interaction tasks. AVAILABILITY AND IMPLEMENTATION: The source code and data are available at https://github.com/datax-lab/CYP.
Ki-Hwa Kim, Avinash Yaganapu, Sai Kosaraju, Aashish Bhatt, Yun Lyna Luo, Sai Phani Krishna Parsa, Hyun Lee 0004, Jun Hyuck Lee, Tae-Jin Oh, Mingon Kang
Bioinform.11
2025 Explainable Integrative Bipartite Graph Convolutional Neural Network for Predicting Ejection Fraction in Echocardiography
Seungeun Lee, Kyungtae Kang, Mingon Kang
MICCAI (12)4
2024 Evidential deep learning for trustworthy prediction of enzyme commission number
abstract
The rapid growth of uncharacterized enzymes and their functional diversity urge accurate and trustworthy computational functional annotation tools. However, current state-of-the-art models lack trustworthiness on the prediction of the multilabel classification problem with thousands of classes. Here, we demonstrate that a novel evidential deep learning model (named ECPICK) makes trustworthy predictions of enzyme commission (EC) numbers with data-driven domain-relevant evidence, which results in significantly enhanced predictive power and the capability to discover potential new motif sites. ECPICK learns complex sequential patterns of amino acids and their hierarchical structures from 20 million enzyme data. ECPICK identifies significant amino acids that contribute to the prediction without multiple sequence alignment. Our intensive assessment showed not only outstanding enhancement of predictive performance on the largest databases of Uniprot, Protein Data Bank (PDB) and Kyoto Encyclopedia of Genes and Genomes (KEGG), but also a capability to discover new motif sites in microorganisms. ECPICK is a reliable EC number prediction tool to identify protein functions of an increasing number of uncharacterized enzymes.
So-Ra Han, Sai Kosaraju, Jeungmin Lee, Hyun Lee 0004, Jun Hyuck Lee, Tae-Jin Oh, Mingon Kang
Briefings Bioinform.8
2024 SPIN: sex-specific and pathway-based interpretable neural network for sexual dimorphism analysis
abstract
Sexual dimorphism in prevalence, severity and genetic susceptibility exists for most common diseases. However, most genetic and clinical outcome studies are designed in sex-combined framework considering sex as a covariate. Few sex-specific studies have analyzed males and females separately, which failed to identify gene-by-sex interaction. Here, we propose a novel unified biologically interpretable deep learning-based framework (named SPIN) for sexual dimorphism analysis. We demonstrate that SPIN significantly improved the C-index up to 23.6% in TCGA cancer datasets, and it was further validated using asthma datasets. In addition, SPIN identifies sex-specific and -shared risk loci that are often missed in previous sex-combined/-separate analysis. We also show that SPIN is interpretable for explaining how biological pathways contribute to sexual dimorphism and improve risk prediction in an individual level, which can result in the development of precision medicine tailored to a specific individual's characteristics.
Euiseong Ko, Youngsoon Kim, Farhad Shokoohi, Tesfaye B. Mersha, Mingon Kang
Briefings Bioinform.5
2024 Multi-layered self-attention mechanism for weakly supervised semantic segmentation
Avinash Yaganapu, Mingon Kang
Comput. Vis. Image Underst.2
2023 Epidemic Vulnerability Index for Effective Vaccine Distribution Against Pandemic
abstract
COVID-19 vaccine distribution route directly impacts the community's mortality and infection rate. Therefore, optimal vaccination dissemination would appreciably lower the death and infection rates. This paper proposes the Epidemic Vulnerability Index (EVI) that quantitatively evaluates the subject's potential risk. Our primary aim for the suggested index is to diminish both infection rate and death rate efficiently. EVI was accordingly designed with clinical factors determining the mortality and social factors incorporating the infection rate. Through statistical COVID-19 patient dataset analysis and social network analysis with an agent-based model that is analogous to a real-world system, we define and experimentally validate the capability of EVI. Our experiments consist of nine vaccination distribution scenarios, including existing indexes which estimate the risk and stochastically proliferate the contagion and vaccine in a 300,000 agent-based graph network. We compared the outcome and variation of the three metrics in the experiments: infection case, death case, and death rate. Through this assessment, vaccination by the descending order of EVI has shown to have a significant outcome with an average of 5.0% lower infection cases, 9.4% lower death cases, and 3.5% lower death rate than other vaccine distribution routes.
Hunmin Lee, Mingon Kang, Donghyun Kim 0001, Yingshu Li 0001
IEEE ACM Trans. Comput. Biol. Bioinform.2
2022 A roadmap for multi-omics data integration using deep learning
abstract
High-throughput next-generation sequencing now makes it possible to generate a vast amount of multi-omics data for various applications. These data have revolutionized biomedical research by providing a more comprehensive understanding of the biological systems and molecular mechanisms of disease development. Recently, deep learning (DL) algorithms have become one of the most promising methods in multi-omics data analysis, due to their predictive performance and capability of capturing nonlinear and hierarchical features. While integrating and translating multi-omics data into useful functional insights remain the biggest bottleneck, there is a clear trend towards incorporating multi-omics analysis in biomedical research to help explain the complex relationships between molecular layers. Multi-omics data have a role to improve prevention, early detection and prediction; monitor progression; interpret patterns and endotyping; and design personalized treatments. In this review, we outline a roadmap of multi-omics integration using DL and offer a practical perspective into the advantages, challenges and barriers to the implementation of DL in multi-omics data.
Mingon Kang, Euiseong Ko, Tesfaye B. Mersha
Briefings Bioinform.1
2021 Epidemic Vulnerability Index for Effective Vaccine Distribution Against Pandemic
Hunmin Lee, Mingon Kang, Yingshu Li 0001, Donghyun Kim 0001
ISBRA2
2021 PathCNN: interpretable convolutional neural networks for survival prediction and pathway analysis applied to glioblastoma
abstract
MOTIVATION: Convolutional neural networks (CNNs) have achieved great success in the areas of image processing and computer vision, handling grid-structured inputs and efficiently capturing local dependencies through multiple levels of abstraction. However, a lack of interpretability remains a key barrier to the adoption of deep neural networks, particularly in predictive modeling of disease outcomes. Moreover, because biological array data are generally represented in a non-grid structured format, CNNs cannot be applied directly. RESULTS: To address these issues, we propose a novel method, called PathCNN, that constructs an interpretable CNN model on integrated multi-omics data using a newly defined pathway image. PathCNN showed promising predictive performance in differentiating between long-term survival (LTS) and non-LTS when applied to glioblastoma multiforme (GBM). The adoption of a visualization tool coupled with statistical analysis enabled the identification of plausible pathways associated with survival in GBM. In summary, PathCNN demonstrates that CNNs can be effectively applied to multi-omics data in an interpretable manner, resulting in promising predictive power while identifying key biological correlates of disease. AVAILABILITY AND IMPLEMENTATION: The source code is freely available at: https://github.com/mskspi/PathCNN.
Jung Hun Oh, Euiseong Ko, Mingon Kang, Allen R. Tannenbaum, Joseph O. Deasy
Bioinform.4
2019 Texture-based Deep Learning for Effective Histopathological Cancer Image Classification
abstract
Automatic histopathological Whole Slide Image (WSI) analysis for cancer classification has been highlighted along with the advancements in microscopic imaging techniques, since manual examination and diagnosis with WSIs are time- and cost-consuming. Recently, deep convolutional neural networks have succeeded in histopathological image analysis. However, despite the success of the development, there are still opportunities for further enhancements. In this paper, we propose a novel cancer texture-based deep neural network (CAT-Net) that learns scalable morphological features from histopathological WSIs. The innovation of CAT-Net is twofold: (1) capturing invariant spatial patterns by dilated convolutional layers and (2) improving predictive performance while reducing model complexity. Moreover, CAT-Net can provide discriminative morphological (texture) patterns formed on cancerous regions of histopathological images comparing to normal regions. We elucidated how our proposed method, CAT-Net, captures morphological patterns of interest in hierarchical levels in the model. The proposed method out-performed the current state-of-the-art benchmark methods on accuracy, precision, recall, and F1 score.
Nelson Zange Tsaku, Sai Kosaraju, Tasmia Aqila, Mohammad Masum, Dae Hyun Song, M. Mondal Ananda, Hyun Min Koh, Mingon Kang
BIBM8
2019 DoT-Net: Document Layout Classification Using Texture-Based CNN
abstract
Document Layout Analysis (DLA) is a segmentation process that decomposes a scanned document image into its blocks of interest and classifies them. DLA is essential in a large number of applications, such as Information Retrieval, Machine Translation, Optical Character Recognition (OCR) systems, and structured data extraction from documents. However, identification of document blocks in DLA is challenging due to variations of block locations, inter-and intra-class variability, and background noises. In this paper, we propose a novel texture-based convolutional neural network for document layout analysis, called DoT-Net. DoT-Net is a multiclass classifier that can effectively identify document component blocks such as text, image, table, mathematical expression, and line-diagram, whereas most related methods have focused on the text vs. non-text block classification problem. DoT-Net can capture textural variations among the multiclass regions of documents. Our proposed method DoT-Net achieved promising results outperforming state-of-the-art document layout classifiers on accuracy, F1 score, and AUC. The open-source code of DoT-Net is available at https://github.com/datax-lab/DoTNet.
Sai Kosaraju, Mohammad Masum, Nelson Zange Tsaku, Pritesh Patel, Tanju Bayramoglu, Girish Modgil, Mingon Kang
ICDAR7
2019 Semi-Supervised Discriminative Transfer Learning in Cross-Language Text Classification
abstract
Cross-Language Text Classification (CLTC) has been increasing its attention to multilingual data due to its exponentially growing. CLTC aims to classify text documents in a label-scarce language, by leveraging classification information in a label-rich language. We propose a novel semi-supervised Discriminative Transfer Learning method (DTL) for the CLTC problem of a semi-supervised setting. A small number of paired labeled data in bilingual documents constructs a discriminative transfer model that maximizes the correlations of the documents in both languages, while a large number of unlabeled data are used for accurate data reconstruction. The discriminative transfer model minimizes the discrepancy between bilingual subspaces prioritizing discriminative features to improve the text classification performance without an automatic machine translation that most state-of-the-art methods require. The performance of DTL is empirically and statistically assessed by intensive experiments with the publicly available data, Reuters RCV1/RCV2 collections. The experimental results demonstrate that DTL outperforms several representative state-of-the-art methods in CLTC in terms of accuracy and efficiency.
Mingon Kang, Ashis Kumer Biswas, Dong-Chul Kim, Jean Gao
ICMLA1
2019 Gene- and Pathway-Based Deep Neural Network for Multi-omics Data Integration to Predict Cancer Survival Outcomes
Mohammad Masum, Jung Hun Oh, Mingon Kang
ISBRA4
2019 Robust Inductive Matrix Completion Strategy to Explore Associations Between LincRNAs and Human Disease Phenotypes
abstract
Over the past few years, it has been established that a number of long intergenic non-coding RNAs (lincRNAs) are linked to a wide variety of human diseases. The relationship among many other lincRNAs still remains as puzzle. Validation of such link between the two entities through biological experiments is expensive. However, piles of information about the two are becoming available, thanks to the High Throughput Sequencing (HTS) platforms, Genome Wide Association Studies (GWAS), etc., thereby opening opportunity for cutting-edge machine learning and data mining approaches. However, there are only a few in silico lincRNA-disease association inference tools available to date, and none of these utilizes side information of both the entities. The recently developed Inductive Matrix Completion (IMC) technique provides a recommendation platform among two entities considering respective side information. But, the formulation of IMC is incapable of handling noise and outliers that may present in the dataset, while data sparsity consideration is another issue with the standard IMC method. Thus, a robust version of IMC is needed that can solve these two issues. As a remedy, in this paper, we propose Robust Inductive Matrix Completion (RIMC) using 12;1 norm loss function aswell as 12;1 norm based regularization. We applied RIMC to the available association data between human lincRNAs and OMIM disease phenotypes as well as a diverse set of side information about the lincRNAs and the diseases. Our method performs better than the state-of-the-art methods in terms of precision@k and recall@k at the top-k disease prioritization to the subject lincRNAs. We also demonstrate that RIMC is equally effective forquerying about novel lincRNAs, as well as predicting rankof a newly known disease for a set of well-characterized lincRNAs. Availability: All the supporting datasets are available at the publicly accessible URL located at http://biomecis.uta.edu/~ashis/res/RIMC/.
Ashis Kumer Biswas, Dong-Chul Kim, Mingon Kang, Jean Gao
IEEE ACM Trans. Comput. Biol. Bioinform.3
2018 Cox-PASNet: Pathway-based Sparse Deep Neural Network for Survival Analysis
Youngsoon Kim, Tejaswini Mallavarapu, Jung Hun Oh, Mingon Kang
BIBM5
2018 PASCL: Pathway-based Sparse Deep Clustering for Identifying Unknown Cancer Subtypes
Tejaswini Mallavarapu, Youngsoon Kim, Jung Hun Oh, Mingon Kang
BIBM5
2018 PASNet: pathway-associated sparse deep neural network for prognosis prediction from high-throughput data
abstract
BACKGROUND: Predicting prognosis in patients from large-scale genomic data is a fundamentally challenging problem in genomic medicine. However, the prognosis still remains poor in many diseases. The poor prognosis may be caused by high complexity of biological systems, where multiple biological components and their hierarchical relationships are involved. Moreover, it is challenging to develop robust computational solutions with high-dimension, low-sample size data. RESULTS: In this study, we propose a Pathway-Associated Sparse Deep Neural Network (PASNet) that not only predicts patients' prognoses but also describes complex biological processes regarding biological pathways for prognosis. PASNet models a multilayered, hierarchical biological system of genes and pathways to predict clinical outcomes by leveraging deep learning. The sparse solution of PASNet provides the capability of model interpretability that most conventional fully-connected neural networks lack. We applied PASNet for long-term survival prediction in Glioblastoma multiforme (GBM), which is a primary brain cancer that shows poor prognostic performance. The predictive performance of PASNet was evaluated with multiple cross-validation experiments. PASNet showed a higher Area Under the Curve (AUC) and F1-score than previous long-term survival prediction classifiers, and the significance of PASNet's performance was assessed by Wilcoxon signed-rank test. Furthermore, the biological pathways, found in PASNet, were referred to as significant pathways in GBM in previous biology and medicine research. CONCLUSIONS: PASNet can describe the different biological systems of clinical outcomes for prognostic prediction as well as predicting prognosis more accurately than the current state-of-the-art methods. PASNet is the first pathway-based deep neural network that represents hierarchical representations of genes and pathways and their nonlinear effects, to the best of our knowledge. Additionally, PASNet would be promising due to its flexible model representation and interpretability, embodying the strengths of deep learning. The open-source code of PASNet is available at https://github.com/DataX-JieHao/PASNet .
Youngsoon Kim, Mingon Kang
BMC Bioinform.4
2018 Nearest neighbor search with locally weighted linear regression for heartbeat classification
Juyoung Park, Md. Zakirul Alam Bhuiyan, Mingon Kang, Junggab Son, Kyungtae Kang
Soft Comput.3
2017 R-PathCluster: Identifying cancer subtype of glioblastoma multiforme using pathway-based restricted boltzmann machine
abstract
Glioblastoma multiforme (GBM) is the most fatal malignant type of brain tumor with a very poor prognosis with a median survival of around one year. Numerous studies have reported tumor subtypes that consider different characteristics on individual patients, which may play important roles in determining the survival rates in GBM. In this study, we present a pathway-based clustering method using Restricted Boltzmann Machine (RBM), called R-PathCluster, for identifying unknown subtypes with pathway markers of gene expressions. In order to assess the performance of R-PathCluster, we conducted experiments with several clustering methods such as k-means, hierarchical clustering, and RBM models with different input data. R-PathCluster showed the best performance in clustering long-term and short-term survivals, although its clustering score was not the highest among them in experiments. R-PathCluster provides a solution to interpret the model in biological sense, since it takes pathway markers that represent biological process of pathways. We discussed that our findings from R-PathCluster are supported by many biological literatures.
Tejaswini Mallavarapu, Youngsoon Kim, Jung Hun Oh, Mingon Kang
BIBM4
2017 Multi-Block Bipartite Graph for Integrative Genomic Analysis
abstract
Human diseases involve a sequence of complex interactions between multiple biological processes. In particular, multiple genomic data such as Single Nucleotide Polymorphism (SNP), Copy Number Variation (CNV), DNA Methylation (DM), and their interactions simultaneously play an important role in human diseases. However, despite the widely known complex multi-layer biological processes and increased availability of the heterogeneous genomic data, most research has considered only a single type of genomic data. Furthermore, recent integrative genomic studies for the multiple genomic data have also been facing difficulties due to the high-dimensionality and complexity, especially when considering their intra- and inter-block interactions. In this paper, we introduce a novel multi-block bipartite graph and its inference methods, MB2I and sMB2I, for the integrative genomic study. The proposed methods not only integrate multiple genomic data but also incorporate intra/inter-block interactions by using a multi-block bipartite graph. In addition, the methods can be used to predict quantitative traits (e.g., gene expression, survival time) from the multi-block genomic data. The performance was assessed by simulation experiments that implement practical situations. We also applied the method to the human brain data of psychiatric disorders. The experimental results were analyzed by maximum edge biclique and biclustering, and biological findings were discussed.
Mingon Kang, Juyoung Park, Dong-Chul Kim, Ashis Kumer Biswas, Chunyu Liu 0001, Jean Gao
IEEE ACM Trans. Comput. Biol. Bioinform.1
2016 Robust Inductive Matrix Completion strategy to explore associations between lincRNAs and human disease phenotypes
abstract
Long intergenic non-coding RNAs (lincRNAs) are associated with a wide variety of human diseases. Piles of data about the lincRNAs are becoming available, thanks to the High Throughput Sequencing (HTS) platforms, which open opportunity for cutting-edge machine learning and data mining approaches to analyze the disease association better. However, there are only a few in silico association inference tools available to date, and none of them utilizes the heterogeneous data about the lincRNAs and diseases. The standard Inductive Matrix Completion (IMC) technique provides with a platform among the two entities considering respective side information. But, it has two major issues pertaining to the noise and sparsity in the dataset. Thus, a robust version of IMC is needed to adequately address the issues. In this paper, we propose Robust Inductive Matrix Completion (RIMC) to address these challenges. Then, we applied RIMC to the available association dataset between the lincRNAs and OMIM disease phenotypes with a diverse set of side information of the both. The proposed method performs better than the state-of-the-art methods in terms of precision@k and recall@k at the top-k disease prioritization to the subject lincRNAs. Moreover, with an induction experiment we showed that RIMC performs superior than the standard IMC for ranking unexplored disease phenotypes to a set of known lincRNAs.
Ashis Kumer Biswas, Dong-Chul Kim, Mingon Kang, Jean Gao
BIBM3
2016 Integrative Gene Regulatory Network inference using multi-omics data
abstract
Biological network inference is of importance to understand underlying biological mechanisms. Gene regulatory networks describe molecular interactions of complex biological processes. Graph models are mainly used for gene regulatory networks, where nodes and edges represent genes and their regulations respectively. In the most research, the molecular interactions (edges) of gene regulatory networks are inferred from a single type of genomic data, e.g., gene expression data. However, gene expression is a product of sequential interactions of DNA sequence variations, single nucleotide polymorphism, copy number variation, histone modifications, transcription factor, DNA methylation, and many other factors. There are high-throughput genomic data that measure the various biological processes. We call the multiple types of genomics data as ‘multi-omics data’. In this paper, we propose an Integrative Gene Regulatory Network inference method (iGRN) that can incorporate multi-omics data and their interactions in the graph model of gene regulatory network. Copy number variation and DNA methylation were considered for multi-omics data in this paper. The proposed method, iGRN, was applied to the human brain data of psychiatric disorder. Through the experiments, iGRN showed its better performance on model representation and interpretation than other integrative methods in gene regulatory network inference.
Neda Zarayeneh, Jung Hun Oh, Donghyun Kim 0001, Chunyu Liu 0001, Jean Gao, Sang C. Suh, Mingon Kang
BIBM7
2016 Computational modeling of phagocyte transmigration for foreign body responses to subcutaneous biomaterial implants in mice
abstract
BACKGROUND: Computational modeling and simulation play an important role in analyzing the behavior of complex biological systems in response to the implantation of biomedical devices. Quantitative computational modeling discloses the nature of foreign body responses. Such understanding will shed insight on the cause of foreign body responses, which will lead to improved biomaterial design and will reduce foreign body reactions. One of the major obstacles in computational modeling is to build a mathematical model that represents the biological system and to quantitatively define the model parameters. RESULTS: In this paper, we considered quantitative inter connections and logical relationships among diverse proteins and cells, which have been reported in biological experiments and literature. Based on the established biological discovery, we have built a mathematical model while unveiling the key components that contribute to biomaterial-mediated inflammatory responses. For the parameter estimation of the mathematical model, we proposed a global optimization algorithm, called Discrete Selection Levenberg-Marquardt (DSLM). This is an extension of Levenberg-Marquardt (LM) algorithm which is a gradient-based local optimization algorithm. The proposed DSLM suggests a new approach for the selection of optimal parameters in the discrete space with fast computational convergence. CONCLUSIONS: The computational modeling not only provides critical clues to recognize current knowledge of fibrosis development but also enables the prediction of yet-to-be observed biological phenomena.
Mingon Kang, Jean Gao
BMC Bioinform.1
2015 An integrative genomic study for multimodal genomic data using multi-block bipartite graph
abstract
Human diseases involve a sequence of complex interactions in multiple biological processes. In particular, multiple genomic data such as Single Nucleotide Polymorphism (SNP), Copy Number Variation (CNV), and DNA Methylation (DM) and their interactions simultaneously play an important role in the variation of mRNA transcription in human diseases. However, despite of the widely known complex multi-layer biological processes and increased availability of the heterogeneous genomic data, most research has considered only a single type of the genomic data. Furthermore, recent integrative genomic studies for the multiple genomic data have also been facing difficulties due to the high-dimensionality and complexity, especially when considering their intra- and inter-block interactions. In this paper, we introduce a novel multi-block bipartite graph and its inference methods, MB2I and sMB2I, for the integrative genomic study. The proposed methods not only integrate the multiple genomic data but also incorporate their intra/inter-block interactions by using a multi-block bipartite graph. In addition, the methods can be used to predict quantitative traits (e.g. gene expression, survival time) from the multi-block genomic data. The outstanding performance was assessed by simulation experiments that implement practical situations.
Mingon Kang, Juyoung Park, Dong-Chul Kim, Ashis Kumer Biswas, Chunyu Liu 0001, Jean Gao
BIBM1
2015 Integrative approach for inference of gene regulatory networks using lasso-based random featuring and application to Psychiatric disorders
abstract
Inferring gene regulatory networks is one of the most interesting research areas in the systems biology. Many inference methods have been developed by using a variety of computational models and approaches. In this paper, we propose two network inference methods based on a lasso-based random feature selection algorithm (LARF). There are three main contributions. First, our z score-based method to measure gene expression variations from knockout data is more effective than similar criteria of related works. Second, we confirmed that the true regulator selection can be effectively improved by LARF. Lastly, we verified that an integrative approach can clearly outperform a single method when two different methods are effectively jointed. In the experiments, our method outperformed state of the art methods on simulated data, and LARF also was applied to the inference of gene regulatory networks associated with Psychiatric disorders.
Dong-Chul Kim, Mingon Kang, Ashis Kumer Biswas, Chunyu Liu 0001, Jean Gao
BIBM2
2015 Heartbeat classification for detecting arrhythmia using normalized beat morphology features
abstract
We propose a method of arrhythmia detection based on beat morphology, which offers a new set of features for heartbeat classification. This can be performed by nearest-neighbor search, which we applied to heartbeats from the MIT-BIH arrhythmia database. Our classifier achieved an overall accuracy of 98.18% on 103,923 heartbeats.
Juyoung Park, Mingon Kang, Kyungtae Kang
BIBM2
2015 eQTL epistasis: detecting epistatic effects and inferring hierarchical relationships of genes in biological pathways
abstract
MOTIVATION: Epistasis is the interactions among multiple genetic variants. It has emerged to explain the 'missing heritability' that a marginal genetic effect does not account for by genome-wide association studies, and also to understand the hierarchical relationships between genes in the genetic pathways. The Fisher's geometric model is common in detecting the epistatic effects. However, despite the substantial successes of many studies with the model, it often fails to discover the functional dependence between genes in an epistasis study, which is an important role in inferring hierarchical relationships of genes in the biological pathway. RESULTS: We justify the imperfectness of Fisher's model in the simulation study and its application to the biological data. Then, we propose a novel generic epistasis model that provides a flexible solution for various biological putative epistatic models in practice. The proposed method enables one to efficiently characterize the functional dependence between genes. Moreover, we suggest a statistical strategy for determining a recessive or dominant link among epistatic expression quantitative trait locus to enable the ability to infer the hierarchical relationships. The proposed method is assessed by simulation experiments of various settings and is applied to human brain data regarding schizophrenia. AVAILABILITY AND IMPLEMENTATION: The MATLAB source codes are publicly available at: http://biomecis.uta.edu/epistasis.
Mingon Kang, Chunling Zhang, Hyung-Wook Chun, Chris Ding, Chunyu Liu 0001, Jean Gao
Bioinform.1
2014 Multi-block and Multi-task Learning for Integrative Genomic Study
abstract
The importance of an integrative genomic study is steadily increasing in an emerging era of various high-throughput genomic data. Mechanisms of human diseases consist of complex interactions of multiple biological processes such as genetic, epigenetic, and transcriptional regulation. The collection of the multiple genomic data that represents the multiple processes is called 'multi-block data'. The multi-block data profiled from human disease samples provide comprehensive global snapshots of the diseases. Due to the rapid development of high-throughput technologies, the integrative genomic study using the multi-block data has been more highlighted than ever. However, in spite of its importance, there are only a few methodologies that can analyze such data. In this paper, we propose a novel Multi-Block and Multi-Task Learning (MBMTL) method for the integrative genomic study. We consider Single Nucleotide Polymorphism (SNP), Copy Number Variation (CNV), DNA methylation, and gene expression data as the multi-block data from four group samples of three major psychiatric disorders as well as data from a normal control. MBMTL identifies biomarkers that play important roles in explaining mechanisms of the human diseases from the multi-block data. We also take a multi-task problem into account so that we can identify different functions of the mechanisms. The performance of the proposed MBMTL was assessed by comparing it to a number of existing multi-block methods through simulation studies. We applied MBMTL to the multi-block data of the major psychiatric disorder samples.
Mingon Kang, Dong-Chul Kim, Chunyu Liu 0001, Baoju Zhang, Jean Gao
BIBE1
2014 Integration of DNA Methylation, Copy Number Variation, and Gene Expression for Gene Regulatory Network Inference and Application to Psychiatric Disorders
abstract
Biological network inference is a crucial problem to solve in Bioinformatics as most of biological process are based on bio molecular interactions. Many researchers have worked on especially the inference of gene regulatory networks where a node and edge represent a gene and regulation relationship respectively assuming that a gene can regulate another gene indirectly. However, a gene expression level can be influenced by not only genes and proteins but also other biological factors. Therefore, the inference could be more effective if those factors are considered in gene regulatory network inferences. In this paper, we propose an integrative approach to infer gene regulatory networks where a gene can be regulated by not only gene and but also DNA Methylation and copy number variation. It is assumed that a gene can be directly regulated by a single DNA Methylation and copy number variation at most. The simulation results show that our method outperforms popular and state-of-the-art methods of biological network inference. In addition, we applied the proposed method to psychiatric disorder data. The inferred networks provide the relationships within a set of genes that are more likely to be regulated by DNA Methylation and copy number variation of the genes.
Dong-Chul Kim, Mingon Kang, Baoju Zhang, Chunyu Liu 0001, Jean Gao
BIBE2
2013 eQTL epistasis: Detecting complex interaction effects between multiple loci from eQTL data
abstract
The identification of expression quantitative trait loci (eQTL) epistasis, which is a non-linear interaction effect between two or more genetic loci that control quantitative traits, play an essential role in understanding the mechanisms of gene interactions in the complex diseases. However, many studies have ignored the possibility of the epistasis in spite of the countless empirical evidence of epistasis in Genetics. Thus, the epistasis research is still in its infancy. Furthermore, most eQTL epistasis studies have involved only the interaction model with additive effects, which lacks the power to represent either 1) other biologically putative interaction models or 2) the genetic dominance that is a superior relationship of an allele against another one at the same locus. To tackle the problems, we propose a general eQTL epistasis model that provides the capability to incorporate diverse interaction models so that it can be widely utilized as a generic equation of epistasis in the multiple testing of eQTL studies. The extensibility of the eQTL epistasis model into multivariate methods is also proposed to reduce the computational burden of the multiple testing and to consider group effects. As the example of the extensibility, we provide the methodology that embeds the eQTL epistasis model into the sparse canonical correlation analysis (SCCA) method. A globally optimal solution is provided and the performance was assessed by realistic simulation experiments. A study of psychiatric disorder diseases with the method was conducted as a target application.
Mingon Kang, Shuo Li 0009, Chunyu Liu 0001, Jean Gao
BIBM1
2013 eQTL Mapping Study via Regularized Sparse Canonical Correlation Analysis
abstract
While genome-wide association studies (GWAS) have focused on discovering genetic loci mapped to a disease, expression quantitative trait loci (eQTL) studies combine micro array data and provide a powerful approach. Micro arrays allow one to measure thousands of gene expressions simultaneously and the advances in eQTL studies enable one to capture the insight of the genetic architecture of gene expression. A number of multivariate methods have been recently proposed to identify genetic loci which are linked to gene expression taking into account joint effects and relationships between the units rather than the single locus alone independently. However, the previous research has limitations, such as the lack of supporting the cis/tran-eQTL model into being accepted as a general genetics model. We propose a novel regularized eQTL association mapping detection (Reg-AMADE) method. We have focused on the following three problems. First, we need to take into account co-expressed genes without using clustering or partitioning techniques, as well as detecting linkage disequilibrium and the joint effect of multiple genetic markers. Secondly, we need to build a regularized model to support the cis- and trans-eQTL model observed in most association studies. Lastly, we need to discover the significant genes underlying within diseases rather than a common component. We also propose a new simulation experiment method that implements practical situations so that the results can be evaluated in the true sense instead of the assessment with random samples generated from multivariate normal distributions that most research has mainly used. The power to detect both the joint effect and grouping effect of SNPs and gene expressions is assessed in the simulation study.
Mingon Kang, Shuo Li 0009, Dong-Chul Kim, Chunyu Liu 0001, Baoju Zhang, Jean Gao
ICMLA (1)1
2010 Computational modeling of phagocyte transmigration during biomaterial-mediated foreign body responses
abstract
One of the major obstacles in computational modeling of a biological system is to determine a large number of parameters in the mathematical equations representing biological properties of the system. To tackle this problem, we have developed a global optimization method, called Discrete Selection Levenberg-Marquardt (DSLM), for parameter estimation. For fast computational convergence, DSLM suggests a new approach for the selection of optimal parameters in the discrete spaces, while other global optimization methods such as genetic algorithm and simulated annealing use heuristic approaches that do not guarantee the convergence. As a specific application example, we have targeted understanding phagocyte transmigration which is involved in the fibrosis process for biomedical device implantation. The goal of computational modeling is to construct an analyzer to understand the nature of the system. Also, the simulation by computational modeling for phagocyte transmigration provides critical clues to recognize current knowledge of the system and to predict yet-to-be observed biological phenomenon.
Mingon Kang, Jean Gao
BIBM1