VLDB 2026 Research / reviewers in the wild / expert
Hojung Nam
dblp:71/6205
· DBLP profile ↗
17ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0002-5109-9114ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AI-guided discovery and optimization of antimicrobial peptides through species-aware language modelabstractThe rise of antibiotic-resistant bacteria drives an urgent need for novel antimicrobial agents. Antimicrobial peptides (AMPs) show promise solutions due to their multiple mechanisms of action and reduced propensity for resistance development. This study introduces LLAMP (Large Language model for AMP activity prediction), a target species-aware AI model that leverages pre-trained language models to predict minimum inhibitory concentration values of AMPs. Using LLAMP, we screened approximately 5.5 million peptide sequences, identifying peptides 13 and 16 as the most selective and most potent candidates, respectively. Analysis of attention values allowed us to pinpoint critical amino acid residues (e.g., Trp, Lys, and Phe). Using the critical amino acids, the sequence of the most selective peptide 13 was engineered to increase amphipathicity through targeted modifications, yielding peptide 13-5 with an overall enhancement in antimicrobial activity but a reduction in selectively. Notably, peptides 13-5 and 16 demonstrated antimicrobial potency and selectivity comparable to the clinically investigated AMP pexiganan. Our work demonstrates the potential of AI to expedite the discovery of peptide-based antibiotics to combat antibiotic resistance. Daehun Bae, Minsang Kim, Hojung Nam |
Briefings Bioinform. | 4 |
| 2025 | DD-PRiSM: a deep learning framework for decomposition and prediction of synergistic drug combinationsabstractCombination therapies have emerged as a promising approach for treating complex diseases, particularly cancer. However, predicting the efficacy and safety profiles of these therapies remains a significant challenge, primarily because of the complex interactions among drugs and their wide-ranging effects. To address this issue, we introduce DD-PRiSM (Decomposition of Drug-Pair Response into Synergy and Monotherapy effect), a deep-learning pipeline that predicts the effects of combination therapy. DD-PRiSM consists of two predictive models. The first is the Monotherapy model, which predicts parameters of the drug response curve based on drug structure and cell line gene expression. This reconstructed curve is then used to predict cell viability at the given drug dosage. The second is the Combination therapy model, which predicts the efficacy of drug combinations by analyzing individual drug effects and their synergistic interactions with a specific dosage level of individual drugs. The efficacy of DD-PRiSM is demonstrated through its performance metrics, achieving a root mean square error of 0.0854, a Pearson correlation coefficient of 0.9063, and an R2 of 0.8209 for unseen pairs. Furthermore, DD-PRiSM distinguishes itself by its capability to decompose combination therapy efficacy, successfully identifying synergistic drug pairs. We demonstrated synergistic responses vary across cancer types and identified hub drugs that trigger synergistic effects. Finally, we suggested a promising drug pair through our case study. Iljung Jin, Songyeon Lee, Martin Schmuhalek, Hojung Nam |
Briefings Bioinform. | 4 |
| 2025 | KG-SLomics: Synthetic Lethality Prediction Using Knowledge Graph and Cancer Type-Specific Multiomics Integrated Graph Neural NetworkabstractSynthetic lethality (SL) is a phenomenon in which the simultaneous alterations of two genes evoke cell death, whereas a mutation of either gene alone does not adversely affect cell survival. After the clinical application of PARP inhibitors, SL has been a promising strategy for the undruggable cancer mutations by targeting their alternative partner genes. While various statistical and computational methods can predict SL pairs, they often overlook key challenges, including variation across cancer types and reliance on outdated networks or gene-specific data that fail to capture cancer-specific features. Recent progress has addressed these gaps, but it struggles to generalize across multiple cancer types. In this paper, we propose KG-SLomics, a relational graph attention network-based model that predicts SL using an extensively updated knowledge graph (KG) and multiple cancer cell line data. We construct a comprehensive KG incorporating newly curated biological entities, tripling its size compared to previous versions. Pre-trained KG embeddings are combined with multiomics data to capture topological and cancer-specific features. Through relational message passing, KG-SLomics calculates SL probabilities with high accuracy, allocating high attention scores to the relevant entities in KG. It outperformed advanced baselines in various evaluations and suggested novel therapeutic targets, underscoring its clinical potential. Songyeon Lee, Hojung Nam |
IEEE Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | BayeshERG: a robust, reliable and interpretable deep learning model for predicting hERG channel blockersabstractUnintended inhibition of the human ether-à-go-go-related gene (hERG) ion channel by small molecules leads to severe cardiotoxicity. Thus, hERG channel blockage is a significant concern in the development of new drugs. Several computational models have been developed to predict hERG channel blockage, including deep learning models; however, they lack robustness, reliability and interpretability. Here, we developed a graph-based Bayesian deep learning model for hERG channel blocker prediction, named BayeshERG, which has robust predictive power, high reliability and high resolution of interpretability. First, we applied transfer learning with 300 000 large data in initial pre-training to increase the predictive performance. Second, we implemented a Bayesian neural network with Monte Carlo dropout to calibrate the uncertainty of the prediction. Third, we utilized global multihead attentive pooling to augment the high resolution of structural interpretability for the hERG channel blockers and nonblockers. We conducted both internal and external validations for stringent evaluation; in particular, we benchmarked most of the publicly available hERG channel blocker prediction models. We showed that our proposed model outperformed predictive performance and uncertainty calibration performance. Furthermore, we found that our model learned to focus on the essential substructures of hERG channel blockers via an attention mechanism. Finally, we validated the prediction results of our model by conducting in vitro experiments and confirmed its high validity. In summary, BayeshERG could serve as a versatile tool for discovering hERG channel blockers and helping maximize the possibility of successful drug discovery. The data and source code are available at our GitHub repository (https://github.com/GIST-CSBL/BayeshERG). Ingoo Lee, Hojung Nam |
Briefings Bioinform. | 4 |
| 2020 | DTMBIO 2020: The Fourteenth International Workshop on Data and Text Mining in Biomedical InformaticsabstractOver a decade, as a specialized workshop in the field of text mining applied to biomedical informatics, DTMBIO (ACM international workshop on Data and Text Mining in Biomedical Informatics) has been held annually in conjunction with one of the largest data management conferences, CIKM. The purpose of DTMBIO is to foster discussions regarding the state-of-the-art applications of data and text mining on biomedical research problems. To address our purpose, we bring together researchers working on computer science and bio/medical informatics area including text mining and high throughput genomic data analysis, such as the next generation Sequencing (NGS) data. DTMBIO 2020 will help scientists navigate emerging trends and opportunities in the evolving area of informatics related techniques and problems in the context of biomedical research. Hyojung Paik, Sunyong Yoo, Hojung Nam, Mark Stevenson 0001, Albert No |
CIKM | 3 |
| 2020 | Prediction model construction of mouse stem cell pluripotency using CpG and non-CpG DNA methylation markersabstractBACKGROUND: Genome-wide studies of DNA methylation across the epigenetic landscape provide insights into the heterogeneity of pluripotent embryonic stem cells (ESCs). Differentiating into embryonic somatic and germ cells, ESCs exhibit varying degrees of pluripotency, and epigenetic changes occurring in this process have emerged as important factors explaining stem cell pluripotency. RESULTS: Here, using paired scBS-seq and scRNA-seq data of mice, we constructed a machine learning model that predicts degrees of pluripotency for mouse ESCs. Since the biological activities of non-CpG markers have yet to be clarified, we tested the predictive power of CpG and non-CpG markers, as well as a combination thereof, in the model. Through rigorous performance evaluation with both internal and external validation, we discovered that a model using both CpG and non-CpG markers predicted the pluripotency of ESCs with the highest prediction performance (0.956 AUC, external test). The prediction model consisted of 16 CpG and 33 non-CpG markers. The CpG and most of the non-CpG markers targeted depletions of methylation and were indicative of cell pluripotency, whereas only a few non-CpG markers reflected accumulations of methylation. Additionally, we confirmed that there exists the differing pluripotency between individual developmental stages, such as E3.5 and E6.5, as well as between induced mouse pluripotent stem cell (iPSC) and somatic cell. CONCLUSIONS: In this study, we investigated CpG and non-CpG methylation in relation to mouse stem cell pluripotency and developed a model thereon that successfully predicts the pluripotency of mouse ESCs. Soobok Joe, Hojung Nam |
BMC Bioinform. | 2 |
| 2019 | Drug repositioning of herbal compounds via a machine-learning approachabstractBACKGROUND: Drug repositioning, also known as drug repurposing, defines new indications for existing drugs and can be used as an alternative to drug development. In recent years, the accumulation of large volumes of information related to drugs and diseases has led to the development of various computational approaches for drug repositioning. Although herbal medicines have had a great impact on current drug discovery, there are still a large number of herbal compounds that have no definite indications. RESULTS: In the present study, we constructed a computational model to predict the unknown pharmacological effects of herbal compounds using machine learning techniques. Based on the assumption that similar diseases can be treated with similar drugs, we used four categories of drug-drug similarity (e.g., chemical structure, side-effects, gene ontology, and targets) and three categories of disease-disease similarity (e.g., phenotypes, human phenotype ontology, and gene ontology). Then, associations between drug and disease were predicted using the employed similarity features. The prediction models were constructed using classification algorithms, including logistic regression, random forest and support vector machine algorithms. Upon cross-validation, the random forest approach showed the best performance (AUC = 0.948) and also performed well in an external validation assessment using an unseen independent dataset (AUC = 0.828). Finally, the constructed model was applied to predict potential indications for existing drugs and herbal compounds. As a result, new indications for 20 existing drugs and 31 herbal compounds were predicted and validated using clinical trial data. CONCLUSIONS: The predicted results were validated manually confirming the performance and underlying mechanisms - for example, irinotecan as a treatment for neuroblastoma. From the prediction, herbal compounds were considered to be drug candidates for related diseases which is important to be further developed. The proposed prediction model can contribute to drug discovery by suggesting drug candidates from herbal compounds which have potentials but few were studied. A.-Sol Choi, Hojung Nam |
BMC Bioinform. | 3 |
| 2019 | DeepConv-DTI: Prediction of drug-target interactions via deep learning with convolution on protein sequencesabstractIdentification of drug-target interactions (DTIs) plays a key role in drug discovery. The high cost and labor-intensive nature of in vitro and in vivo experiments have highlighted the importance of in silico-based DTI prediction approaches. In several computational models, conventional protein descriptors have been shown to not be sufficiently informative to predict accurate DTIs. Thus, in this study, we propose a deep learning based DTI prediction model capturing local residue patterns of proteins participating in DTIs. When we employ a convolutional neural network (CNN) on raw protein sequences, we perform convolution on various lengths of amino acids subsequences to capture local residue patterns of generalized protein classes. We train our model with large-scale DTI information and demonstrate the performance of the proposed model using an independent dataset that is not seen during the training phase. As a result, our model performs better than previous protein descriptor-based models. Also, our model performs better than the recently developed deep learning models for massive prediction of DTIs. By examining pooled convolution results, we confirmed that our model can detect binding sites of proteins for DTIs. In conclusion, our prediction model for detecting local residue patterns of target proteins successfully enriches the protein features of a raw protein sequence, yielding better prediction results than previous approaches. Our code is available at https://github.com/GIST-CSBL/DeepConv-DTI. Ingoo Lee, Jongsoo Keum, Hojung Nam |
PLoS Comput. Biol. | 3 |
| 2018 | Identification of drug-target interaction by a random walk with restart method on an interactome networkabstractBACKGROUND: Identification of drug-target interactions acts as a key role in drug discovery. However, identifying drug-target interactions via in-vitro, in-vivo experiments are very laborious, time-consuming. Thus, predicting drug-target interactions by using computational approaches is a good alternative. In recent studies, many feature-based and similarity-based machine learning approaches have shown promising results in drug-target interaction predictions. A previous study showed that accounting connectivity information of drug-drug and protein-protein interactions increase performances of prediction by the concept of 'guilt-by-association'. However, the approach that only considers directly connected nodes often misses the information that could be derived from distance nodes. Therefore, in this study, we yield global network topology information by using a random walk with restart algorithm and apply the global topology information to the prediction model. RESULTS: As a result, our prediction model demonstrates increased prediction performance compare to the 'guilt-by-association' approach (AUC 0.89 and 0.67 in the training and independent test, respectively). In addition, we show how weighted features by a random walk with restart yields better performances than original features. Also, we confirmed that drugs and proteins that have high-degree of connectivity on the interactome network yield better performance in our model. CONCLUSIONS: The prediction models with weighted features by considering global network topology increased the prediction performances both in the training and testing compared to non-weighted models and previous a 'guilt-by-association method'. In conclusion, global network topology information on protein-protein interaction and drug-drug interaction effects to the prediction performance of drug-target interactions. Ingoo Lee, Hojung Nam |
BMC Bioinform. | 2 |
| 2018 | Predicting the Absorption Potential of Chemical Compounds Through a Deep Learning ApproachabstractThe human colorectal carcinoma cell line (Caco-2) is a commonly used in-vitro test that predicts the absorption potential of orally administered drugs. In-silico prediction methods, based on the Caco-2 assay data, may increase the effectiveness of the high-throughput screening of new drug candidates. However, previously developed in-silico models that predict the Caco-2 cellular permeability of chemical compounds use handcrafted features that may be dataset-specific and induce over-fitting problems. Deep Neural Network (DNN) generates high-level features based on non-linear transformations for raw features, which provides high discriminant power and, therefore, creates a good generalized model. We present a DNN-based binary Caco-2 permeability classifier. Our model was constructed based on 663 chemical compounds with in-vitro Caco-2 apparent permeability data. Two hundred nine molecular descriptors are used for generating the high-level features during DNN model generation. Dropout regularization is applied to solve the over-fitting problem and the non-linear activation. The Rectified Linear Unit (ReLU) is adopted to reduce the vanishing gradient problem. The results demonstrate that the high-level features generated by the DNN are more robust than handcrafted features for predicting the cellular permeability of structurally diverse chemical compounds in Caco-2 cell lines. Moonshik Shin, Donjin Jang, Hojung Nam, Kwang Hyung Lee, Doheon Lee |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2017 | In Silico Simulation of Signal Cascades in Biomedical Networks Based on the Production Rule System
Sangwoo Kim, Hojung Nam |
ISBRA | 2 |
| 2017 | Prediction models for drug-induced hepatotoxicity by using weighted molecular fingerprintsabstractBACKGROUND: Drug-induced liver injury (DILI) is a critical issue in drug development because DILI causes failures in clinical trials and the withdrawal of approved drugs from the market. There have been many attempts to predict the risk of DILI based on in vivo and in silico identification of hepatotoxic compounds. In the current study, we propose the in silico prediction model predicting DILI using weighted molecular fingerprints. RESULTS: In this study, we used 881 bits of molecular fingerprint and used as features describing presence or absence of each substructure of compounds. Then, the Bayesian probability of each substructure was calculated and labeled (positive or negative for DILI), and a weighted fingerprint was determined from the ratio of DILI-positive to DILI-negative probability values. Using weighted fingerprint features, the prediction models were trained and evaluated with the Random Forest (RF) and Support Vector Machine (SVM) algorithms. The constructed models yielded accuracies of 73.8% and 72.6%, AUCs of 0.791 and 0.768 in cross-validation. In independent tests, models achieved accuracies of 60.1% and 61.1% for RF and SVM, respectively. The results validated that weighted features helped increase overall performance of prediction models. The constructed models were further applied to the prediction of natural compounds in herbs to identify DILI potential, and 13,996 unique herbal compounds were predicted as DILI-positive with the SVM model. CONCLUSIONS: The prediction models with weighted features increased the performance compared to non-weighted models. Moreover, we predicted the DILI potential of herbs with the best performed model, and the prediction results suggest that many herbal compounds could have potential to be DILI. We can thus infer that taking natural products without detailed references about the relevant pathways may be dangerous. Considering the frequency of use of compounds in natural herbs and their increased application in drug development, DILI labeling would be very important. Hojung Nam |
BMC Bioinform. | 2 |
| 2016 | Prediction of compound-target interactions of natural products using large-scale drug and protein informationabstractBACKGROUND: Verifying the proteins that are targeted by compounds of natural herbs will be helpful to select natural herb-based drug candidates. However, this entails a great deal of effort to clarify the interaction throughout in vitro or in vivo experiments. In this light, in silico prediction of the interactions between compounds and target proteins can help ease the efforts. RESULTS: In this study, we performed in silico predictions of herbal compound target identification. First, data related to compounds, target proteins, and interactions between them are taken from the DrugBank database. Then we characterized six classes of compound-target interaction in humans including G-protein-coupled receptors (GPCRs), ion channel, enzymes, receptors, transporters, and other proteins. Also, classification-prediction models that predict the interactions between compounds and target proteins through a machine learning method were constructed using these matrices. As a result, AUC values of six classes are 0.94, 0.93, 0.90, 0.89, 0.91, and 0.76 respectively. Finally, the interactions of compounds from natural products were predicted using the constructed classification models. Furthermore, from our predicted results, we confirmed that several important disease related proteins were predicted as targets of natural herbal compounds. CONCLUSIONS: We constructed classification-prediction models that predict the interactions between compounds and target proteins. The constructed models showed good prediction performances, and numbers of potential natural compounds target proteins were predicted from our results. Jongsoo Keum, Sunyong Yoo, Doheon Lee, Hojung Nam |
BMC Bioinform. | 4 |
| 2015 | SoloDel: a probabilistic model for detecting low-frequent somatic deletions from unmatched sequencing dataabstractMOTIVATION: Finding somatic mutations from massively parallel sequencing data is becoming a standard process in genome-based biomedical studies. There are a number of robust methods developed for detecting somatic single nucleotide variations However, detection of somatic copy number alteration has been substantially less explored and remains vulnerable to frequently raised sampling issues: low frequency in cell population and absence of the matched control samples. RESULTS: We developed a novel computational method SoloDel that accurately classifies low-frequent somatic deletions from germline ones with or without matched control samples. We first constructed a probabilistic, somatic mutation progression model that describes the occurrence and propagation of the event in the cellular lineage of the sample. We then built a Gaussian mixture model to represent the mixed population of somatic and germline deletions. Parameters of the mixture model could be estimated using the expectation-maximization algorithm with the observed distribution of read-depth ratios at the points of discordant-read based initial deletion calls. Combined with conventional structural variation caller, SoloDel greatly increased the accuracy in classifying somatic mutations. Even without control, SoloDel maintained a comparable performance in a wide range of mutated subpopulation size (10-70%). SoloDel could also successfully recall experimentally validated somatic deletions from previously reported neuropsychiatric whole-genome sequencing data. AVAILABILITY AND IMPLEMENTATION: Java-based implementation of the method is available at http://sourceforge.net/projects/solodel/ CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hojung Nam, Sangwoo Kim, Doheon Lee |
Bioinform. | 3 |
| 2014 | A Systems Approach to Predict Oncometabolites via Context-Specific Genome-Scale Metabolic NetworksabstractAltered metabolism in cancer cells has been viewed as a passive response required for a malignant transformation. However, this view has changed through the recently described metabolic oncogenic factors: mutated isocitrate dehydrogenases (IDH), succinate dehydrogenase (SDH), and fumarate hydratase (FH) that produce oncometabolites that competitively inhibit epigenetic regulation. In this study, we demonstrate in silico predictions of oncometabolites that have the potential to dysregulate epigenetic controls in nine types of cancer by incorporating massive scale genetic mutation information (collected from more than 1,700 cancer genomes), expression profiling data, and deploying Recon 2 to reconstruct context-specific genome-scale metabolic models. Our analysis predicted 15 compounds and 24 substructures of potential oncometabolites that could result from the loss-of-function and gain-of-function mutations of metabolic enzymes, respectively. These results suggest a substantial potential for discovering unidentified oncometabolites in various forms of cancers. Hojung Nam, Miguel Campodonico, Aarash Bordbar, Daniel R. Hyduke, Sangwoo Kim, Daniel C. Zielinski, Bernhard O. Palsson |
PLoS Comput. Biol. | 1 |
| 2009 | Combining tissue transcriptomics and urine metabolomics for breast cancer biomarker identificationabstractMOTIVATION: For the early detection of cancer, highly sensitive and specific biomarkers are needed. Particularly, biomarkers in bio-fluids are relatively more useful because those can be used for non-biopsy tests. Although the altered metabolic activities of cancer cells have been observed in many studies, little is known about metabolic biomarkers for cancer screening. In this study, a systematic method is proposed for identifying metabolic biomarkers in urine samples by selecting candidate biomarkers from altered genome-wide gene expression signatures of cancer cells. Biomarkers identified by the present study have increased coherence and robustness because the significances of biomarkers are validated in both gene expression profiles and metabolic profiles. RESULTS: The proposed method was applied to the gene expression profiles and urine samples of 50 breast cancer patients and 50 normal persons. Nine altered metabolic pathways were identified from the breast cancer gene expression signatures. Among these altered metabolic pathways, four metabolic biomarkers (Homovanillate, 4-hydroxyphenylacetate, 5-hydroxyindoleacetate and urea) were identified to be different in cancer and normal subjects (p <0.05). In the case of the predictive performance, the identified biomarkers achieved area under the ROC curve values of 0.75, 0.79 and 0.79, according to a linear discriminate analysis, a random forest classifier and on a support vector machine, respectively. Finally, biomarkers which showed consistent significance in pathways' gene expression as well as urine samples were identified. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hojung Nam, Bong Chul Chung, Ki Young Lee, Doheon Lee |
Bioinform. | 1 |
| 2009 | Identification of temporal association rules from time-series microarray data setsabstractBACKGROUND: One of the most challenging problems in mining gene expression data is to identify how the expression of any particular gene affects the expression of other genes. To elucidate the relationships between genes, an association rule mining (ARM) method has been applied to microarray gene expression data. However, a conventional ARM method has a limit on extracting temporal dependencies between gene expressions, though the temporal information is indispensable to discover underlying regulation mechanisms in biological pathways. In this paper, we propose a novel method, referred to as temporal association rule mining (TARM), which can extract temporal dependencies among related genes. A temporal association rule has the form [gene A upward arrow, gene B downward arrow] --> (7 min) [gene C upward arrow], which represents that high expression level of gene A and significant repression of gene B followed by significant expression of gene C after 7 minutes. The proposed TARM method is tested with Saccharomyces cerevisiae cell cycle time-series microarray gene expression data set. RESULTS: In the parameter fitting phase of TARM, the fitted parameter set [threshold = +/- 0.8, support >or= 3 transactions, confidence >or= 90%] with the best precision score for KEGG cell cycle pathway has been chosen for rule mining phase. With the fitted parameter set, numbers of temporal association rules with five transcriptional time delays (0, 7, 14, 21, 28 minutes) are extracted from gene expression data of 799 genes, which are pre-identified cell cycle relevant genes. From the extracted temporal association rules, associated genes, which play same role of biological processes within short transcriptional time delay and some temporal dependencies between genes with specific biological processes are identified. CONCLUSION: In this work, we proposed TARM, which is an applied form of conventional ARM. TARM showed higher precision score than Dynamic Bayesian network and Bayesian network. Advantages of TARM are that it tells us the size of transcriptional time delay between associated genes, activation and inhibition relationship between genes, and sets of co-regulators. Hojung Nam, Ki Young Lee, Doheon Lee |
BMC Bioinform. | 1 |