VLDB 2026 Research / reviewers in the wild / expert
Inuk Jung
dblp:31/1207
· DBLP profile ↗
13ranked-venue papers
1as first author
5since 2021 · last 2025
0000-0003-0675-4244ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 2Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EpicPred: predicting phenotypes driven by epitope-binding TCRs using attention-based multiple instance learningabstractMOTIVATION: Correctly identifying epitope-binding T-cell receptors (TCRs) is important to both understand their underlying biological mechanism in association to some phenotype and accordingly develop T-cell mediated immunotherapy treatments. Although the importance of the CDR3 region in TCRs for epitope recognition is well recognized, methods for profiling their interactions in association to a certain disease or phenotype remains less studied. We developed EpicPred to identify phenotype-specific TCR-epitope interactions. EpicPred first predicts and removes unlikely TCR-epitope interactions to reduce false positives using the Open-set Recognition (OSR). Subsequently, multiple instance learning was used to identify TCR-epitope interactions specific to a cancer type or severity levels of COVID-19 infected patients. RESULTS: From six public TCR databases, 244 552 TCR sequences and 105 unique epitopes were used to predict epitope-binding TCRs and to filter out non-epitope-binding TCRs using the OSR method. The predicted interactions were used to further predict the phenotype groups in two cancer and four COVID-19 TCR-seq datasets of both bulk and single-cell resolution. EpicPred outperformed the competing methods in predicting the phenotypes, achieving an average AUROC of 0.80 ± 0.07. AVAILABILITY AND IMPLEMENTATION: The EpicPred Software is available at https://github.com/jaeminjj/EpicPred. Jaemin Jeon, Suwan Yu, Sangam Lee, Sang Cheol Kim, Hye-Yeong Jo, Inuk Jung |
Bioinform. | 6 |
| 2024 | LOCOS: A cosine based local gene expression pattern finding algorithm on time-series dataabstractIn gene expression analysis, understanding a biological event that is observed at some time instance often requires capturing genes whose expression levels modulate before and after the event. Such genes are expected to be the responders to the event and will exhibit similar expression patterns during the event. However, after the effect of the event fades, their expression levels may loose their correlation. Hence, it is a non-trivial task to identify genes that share highly similar local expression patterns nearby the event and also allowing some level of divergent expression patterns further away from the event’s time point. Here, we propose LOCOS (LOcal COSine), a novel time-course clustering algorithm tailored to cluster genes exhibiting similar expression patterns within a specified time interval. LOCOS is an extension of the traditional Non-negative Matrix Factorization (NMF), which incorporates the cosine similarity and Euclidean distances in its update procedure. LOCOS maximizes the cosine similarity within an interval of interest, while allowing a relaxed minimization of Euclidean distance outside it. Using synthetic and non-biological time-course data, we showed that LOCOS was able to correctly detect the known local patterns within a specified interval. Furthermore, we used longitudinal single-cell RNA-seq samples from four patients showing deteriorating and recovering health conditions to identify genes related to each of the phenotype. As a result, LOCOS was able to capture gene clusters with distinct expression patterns that aligned with the intervals embedding the clinical deterioration and recovery events. Youjeong Suk, Jaemin Jeon, Inuk Jung |
IEEE Big Data | 3 |
| 2024 | Denoiseit: denoising gene expression data using rank based isolation treesabstractBACKGROUND: Selecting informative genes or eliminating uninformative ones before any downstream gene expression analysis is a standard task with great impact on the results. A carefully curated gene set significantly enhances the likelihood of identifying meaningful biomarkers. METHOD: In contrast to the conventional forward gene search methods that focus on selecting highly informative genes, we propose a backward search method, DenoiseIt, that aims to remove potential outlier genes yielding a robust gene set with reduced noise. The gene set constructed by DenoiseIt is expected to capture biologically significant genes while pruning irrelevant ones to the greatest extent possible. Therefore, it also enhances the quality of downstream comparative gene expression analysis. DenoiseIt utilizes non-negative matrix factorization in conjunction with isolation forests to identify outlier rank features and remove their associated genes. RESULTS: DenoiseIt was applied to both bulk and single-cell RNA-seq data collected from TCGA and a COVID-19 cohort to show that it proficiently identified and removed genes exhibiting expression anomalies confined to specific samples rather than a known group. DenoiseIt also showed to reduce the level of technical noise while preserving a higher proportion of biologically relevant genes compared to existing methods. The DenoiseIt Software is publicly available on GitHub at https://github.com/cobi-git/DenoiseIt. Jaemin Jeon, Youjeong Suk, Sang Cheol Kim, Hye-Yeong Jo, Inuk Jung |
BMC Bioinform. | 6 |
| 2022 | Measuring single-cell level gene expression stability and variability in healthy and severe COVID-19 patients using Kullback-Leibler divergenceabstractRecent advances in single-cell RNA sequencing (scRNA-seq) technology have enabled the acquisition of RNA at the single-cell level, which showed that the expression level of genes is highly variable across and within the cell types. Even well-known housekeeping genes showed high expression variance in a single condition and within the same cell types. Previous studies made efforts to identify stably expressed genes and use them as a yardstick for robust gene expression normalization. On the other hand, drugs were shown to be less effective on genes with high expression variance. Thus, identifying both stably and variably expressed genes is an important task, especially at the single-cell level. In this study, using the Kullback-Leibler divergence method, we proposed a metric to measure the expression stability of each gene. Using private scRNA-seq data composed of 25 severe COVID-19 patients and 40 healthy individuals, we identified variably expressed genes specific to COVID-19-infected patients and healthy cohorts. Inseung Hwang, Jaeyeon Jang, Hye-Yeong Jo, Sang Cheol Kim, Inuk Jung |
BIBM | 6 |
| 2021 | IDEA: Integrating Divisive and Ensemble-Agglomerate hierarchical clustering framework for arbitrary shape dataabstractHierarchical clustering, a traditional clustering method, has been getting attention again. Among several reasons, a credit goes to a recent paper by Dasgupta in 2016 that proposed a cost function that quantitatively evaluates hierarchical clustering trees. An important question is how to combine this recent advance with existing successful clustering methods. In this paper, we propose a hierarchical clustering method to minimize the cost function of clustering tree by incorporating existing clustering techniques. First, we developed an ensemble tree-search method that finds an integrated tree with reduced cost by integrating multiple existing hierarchical clustering methods. Second, to operate on large and arbitrary shape data, we designed an efficient hierarchical clustering framework, called integrating divisive and ensemble-agglomerate (IDEA) by combining it with advanced clustering techniques such as nearest neighbor graph construction, divisive-agglomerate hybridization, and dynamic cut tree. The IDEA clustering method showed better performance in minimizing Dasgupta's cost and improving accuracy (adjusted rand index) over existing cost-minimization-based, and density-based hierarchical clustering methods in experiments using arbitrary shape datasets and complex biology-domain datasets. Hongryul Ahn, Inuk Jung, Heejoon Chae, Minsik Oh, Inyoung Kim, Sun Kim |
IEEE BigData | 2 |
| 2020 | Comprehensive and critical evaluation of individualized pathway activity measurement tools on pan-cancer dataabstractMOTIVATION: Biological pathways are extensively used for the analysis of transcriptome data to characterize biological mechanisms underlying various phenotypes. There are a number of computational tools that summarize transcriptome data at the pathway level. However, there is no comparative study on how well these tools produce useful information at the cohort level, enabling comparison of many samples or patients. RESULTS: In this study, we systematically compared and evaluated 13 different pathway activity inference tools based on 5 comparison criteria using pan-cancer data set. This study has two major contributions. First, our study provides a comprehensive survey on computational techniques used by existing pathway activity inference tools. The tools use different strategies and assume different requirements on data: input transformation, use of labels, necessity of cohort-level input data, use of gene relations and scoring metric. Second, we performed extensive evaluations on the performance of these tools. Because different tools use different methods to map samples to the pathway dimension, the tools are evaluated at the pathway level using five comparison criteria. Starting from measuring how well a tool maintains the characteristics of original gene expression values, robustness was also investigated by adding noise into gene expression data. Classification tasks on three clinical variables (tumor versus normal, survival and cancer subtypes) were performed to evaluate the utility of tools for their clinical applications. In addition, the inferred activity values were compared between the tools to see how similar they are along with the scoring schemes they use. Sangsoo Lim, Sangseon Lee, Inuk Jung, SungMin Rhee, Sun Kim |
Briefings Bioinform. | 3 |
| 2019 | HTRgene: a computational method to perform the integrated analysis of multiple heterogeneous time-series data: case analysis of cold and heat stress response signaling genes in ArabidopsisabstractBACKGROUND: Integrated analysis that uses multiple sample gene expression data measured under the same stress can detect stress response genes more accurately than analysis of individual sample data. However, the integrated analysis is challenging since experimental conditions (strength of stress and the number of time points) are heterogeneous across multiple samples. RESULTS: HTRgene is a computational method to perform the integrated analysis of multiple heterogeneous time-series data measured under the same stress condition. The goal of HTRgene is to identify "response order preserving DEGs" that are defined as genes not only which are differentially expressed but also whose response order is preserved across multiple samples. The utility of HTRgene was demonstrated using 28 and 24 time-series sample gene expression data measured under cold and heat stress in Arabidopsis. HTRgene analysis successfully reproduced known biological mechanisms of cold and heat stress in Arabidopsis. Also, HTRgene showed higher accuracy in detecting the documented stress response genes than existing tools. CONCLUSIONS: HTRgene, a method to find the ordering of response time of genes that are commonly observed among multiple time-series samples, successfully integrated multiple heterogeneous time-series gene expression datasets. It can be applied to many research problems related to the integration of time series data analysis. Hongryul Ahn, Inuk Jung, Heejoon Chae, Dongwon Kang, Woosuk Jung, Sun Kim |
BMC Bioinform. | 2 |
| 2018 | HTRgene: Integrating Multiple Heterogeneous Time-series Data to Investigate Cold and Heat Stress Response Signaling Genes in Arabidopsis
Hongryul Ahn, Inuk Jung, Heejoon Chae, Dongwon Kang, Woosuk Jung, Sun Kim |
BIBM | 2 |
| 2017 | TimesVector: a vectorized clustering approach to the analysis of time series transcriptome data from multiple phenotypesabstractMOTIVATION: Identifying biologically meaningful gene expression patterns from time series gene expression data is important to understand the underlying biological mechanisms. To identify significantly perturbed gene sets between different phenotypes, analysis of time series transcriptome data requires consideration of time and sample dimensions. Thus, the analysis of such time series data seeks to search gene sets that exhibit similar or different expression patterns between two or more sample conditions, constituting the three-dimensional data, i.e. gene-time-condition. Computational complexity for analyzing such data is very high, compared to the already difficult NP-hard two dimensional biclustering algorithms. Because of this challenge, traditional time series clustering algorithms are designed to capture co-expressed genes with similar expression pattern in two sample conditions. RESULTS: We present a triclustering algorithm, TimesVector, specifically designed for clustering three-dimensional time series data to capture distinctively similar or different gene expression patterns between two or more sample conditions. TimesVector identifies clusters with distinctive expression patterns in three steps: (i) dimension reduction and clustering of time-condition concatenated vectors, (ii) post-processing clusters for detecting similar and distinct expression patterns and (iii) rescuing genes from unclassified clusters. Using four sets of time series gene expression data, generated by both microarray and high throughput sequencing platforms, we demonstrated that TimesVector successfully detected biologically meaningful clusters of high quality. TimesVector improved the clustering quality compared to existing triclustering tools and only TimesVector detected clusters with differential expression patterns across conditions successfully. AVAILABILITY AND IMPLEMENTATION: The TimesVector software is available at http://biohealth.snu.ac.kr/software/TimesVector/. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Inuk Jung, Kyuri Jo, Hyejin Kang, Hongryul Ahn, Youngjae Yu, Sun Kim |
Bioinform. | 1 |
| 2016 | Influence maximization in time bounded network identifies transcription factors regulating perturbed pathwaysabstractMOTIVATION: To understand the dynamic nature of the biological process, it is crucial to identify perturbed pathways in an altered environment and also to infer regulators that trigger the response. Current time-series analysis methods, however, are not powerful enough to identify perturbed pathways and regulators simultaneously. Widely used methods include methods to determine gene sets such as differentially expressed genes or gene clusters and these genes sets need to be further interpreted in terms of biological pathways using other tools. Most pathway analysis methods are not designed for time series data and they do not consider gene-gene influence on the time dimension. RESULTS: In this article, we propose a novel time-series analysis method TimeTP for determining transcription factors (TFs) regulating pathway perturbation, which narrows the focus to perturbed sub-pathways and utilizes the gene regulatory network and protein-protein interaction network to locate TFs triggering the perturbation. TimeTP first identifies perturbed sub-pathways that propagate the expression changes along the time. Starting points of the perturbed sub-pathways are mapped into the network and the most influential TFs are determined by influence maximization technique. The analysis result is visually summarized in TF-PATHWAY MAP IN TIME CLOCK: TimeTP was applied to PIK3CA knock-in dataset and found significant sub-pathways and their regulators relevant to the PIP3 signaling pathway. AVAILABILITY AND IMPLEMENTATION: TimeTP is implemented in Python and available at http://biohealth.snu.ac.kr/software/TimeTP/Supplementary information: Supplementary data are available at Bioinformatics online. CONTACT: [email protected]. Kyuri Jo, Inuk Jung, Ji Hwan Moon, Sun Kim |
Bioinform. | 2 |
| 2007 | RMTool: Component-Based Network Management System for Wireless Sensor NetworksabstractThis paper introduces a component-based network management system for wireless sensor networks. The research is motivated by a variety of problems that occur due to the unpredictable behavior of wireless sensor networks. Once sensor nodes are deployed, the inner behavior of a sensor field can only be analyzed by monitoring incoming data packets at the sink node or base station, if only such a framework exists. In this paper, we present a component-based management system, RMTool, which allows developers to easily monitor and analyze the network status and interactively configure the network over unexpected problems while running applications over it. The system has been implemented on a multi-threaded sensor network operating system RETOS, which supports run-time loadable kernel modules. The preliminary evaluation shows that RMTool provides management functionalities as designed, and the implementation is efficient due to the component-based module architecture. Hojung Cha, Inuk Jung |
CCNC | 2 |
| 2007 | RETOS: resilient, expandable, and threaded operating system for wireless sensor networksabstractThis paper presents the design principles, implementation, and evaluation of the RETOS operating system which is specifically developed for micro sensor nodes. RETOS has four distinct objectives, which are to provide (1) a multithreaded programming interface, (2) system resiliency, (3) kernel extensibility with dynamic reconfiguration, and (4) WSN-oriented network abstraction. RETOS is a multithreaded operating system, hence it provides the commonly used thread model of programming interface to developers. We have used various implementation techniques to optimize the performance and resource usage of multithreading. RETOS also provides software solutions to separate kernel from user applications, and supports their robust execution on MMU-less hardware. The RETOS kernel can be dynamically reconfigured, via loadable kernel framework, so a application-optimized and resource-efficient kernel is constructed. Finally, the networking architecture in RETOS is designed with a layering concept to provide WSN-specific network abstraction. RETOS currently supports Atmel ATmega128, TI MSP430, and Chipcon CC2430 family of microcontrollers. Several real-world WSN applications are developed for RETOS and the overall evaluation of the systems is described in the paper. Hojung Cha, Sukwon Choi, Inuk Jung, Hyoseung Kim 0001, Hyojeong Shin, Jaehyun Yoo, Chanmin Yoon |
IPSN | 3 |
| 2007 | The RETOS operating system: kernel, tools and applicationsabstractThis demonstration shows the programming development suite of the RETOS operating system for sensor networks, which provides a robust and multithreaded programming interface to application programmers. We first demonstrate how to build the RETOS kernel on the TI MSP430, Atmel ATmega 128 and Chipcon CC2430 family of microcontrollers. The application or a kernel module is then compiled and disseminated, via wireless channel, to the target motes. The GUI-based RMon network management tool for RETOS is also demonstrated to monitor the networked sensors, and even to control the system's parameters or applications running on them via a remote shell. The system is demonstrated to run on a mixed set of MSP430, ATmega128 and CC2430-based motes. Overall, our demonstration will convince attendees of the programming convenience of developing sensor network applications using RETOS, which is, indeed, a mature and practical system that can be used to develop real-world applications. Hojung Cha, Sukwon Choi, Inuk Jung, Hyoseung Kim 0001, Hyojeong Shin, Jaehyun Yoo, Chanmin Yoon |
IPSN | 3 |