Minsik Oh

dblp:222/4427 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Template-assisted Contrastive Learning of Task-oriented Dialogue Sentence Embeddings
abstract
Learning high quality sentence embeddings from dialogues has drawn increasing attentions as it is essential to solve a variety of dialogue-oriented tasks with low annotation cost.Annotating and gathering utterance relationships in conversations are difficult, while token-level annotations, e.g., entities, slots and templates, are much easier to obtain.Other sentence embedding methods are usually sentence-level self-supervised frameworks and cannot utilize token-level extra knowledge.We introduce Template-aware Dialogue Sentence Embedding (TaDSE), a novel augmentation method that utilizes template information to learn utterance embeddings via self-supervised contrastive learning framework.We further enhance the effect with a synthetically augmented dataset that diversifies utterance-template association, in which slot-filling is a preliminary step.We evaluate TaDSE performance on five downstream benchmark dialogue datasets.The experiment results show that TaDSE achieves significant improvements over previous SOTA methods for dialogue.We further introduce a novel analytic instrument of semantic compression test, for which we discover a correlation with uniformity and alignment.Our code is available at https://github.com/minsik-ai/ Template-Contrastive-Embedding.
Minsik Oh, Jiwei Li 0001, Guoyin Wang 0002
ACL (1)1
2023 P5: Plug-and-Play Persona Prompting for Personalized Response Selection
abstract
The use of persona-grounded retrieval-based chatbots is crucial for personalized conversations, but there are several challenges that need to be addressed.1) In general, collecting persona-grounded corpus is very expensive.2) The chatbot system does not always respond in consideration of persona at real applications.To address these challenges, we propose a plugand-play persona prompting method.Our system can function as a standard open-domain chatbot if persona information is not available.We demonstrate that this approach performs well in the zero-shot setting, which reduces the dependence on persona-ground training data.This makes it easier to expand the system to other languages without the need to build a persona-grounded corpus.Additionally, our model can be fine-tuned for even better performance.In our experiments, the zero-shot model improved the standard model by 7.71 and 1.04 points in the original persona and revised persona, respectively.The fine-tuned model improved the previous state-of-the-art system by 1.95 and 3.39 points in the original persona and revised persona, respectively.To the best of our knowledge, this is the first attempt to solve the problem of personalized response selection using prompt sequences.Our code is available on github 1 .
Joosung Lee, Minsik Oh
EMNLP2
2023 PK-ICR: Persona-Knowledge Interactive Multi-Context Retrieval for Grounded Dialogue
abstract
Identifying relevant persona or knowledge for conversational systems is critical to grounded dialogue response generation.However, each grounding has been mostly researched in isolation with more practical multi-context dialogue tasks introduced in recent works.We define Persona and Knowledge Dual Context Identification as the task to identify persona and knowledge jointly for a given dialogue, which could be of elevated importance in complex multicontext dialogue settings.We develop a novel grounding retrieval method that utilizes all contexts of dialogue simultaneously.Our method requires less computational power via utilizing neural QA retrieval models.We further introduce our novel null-positive rank test which measures ranking performance on semantically dissimilar samples (i.e.hard negatives) in relation to data augmentation.
Minsik Oh, Joosung Lee, Jiwei Li 0001, Guoyin Wang 0002
EMNLP1
2023 GOAT: Gene-level biomarker discovery from multi-Omics data using graph ATtention neural network for eosinophilic asthma subtype
abstract
MOTIVATION: Asthma is a heterogeneous disease where various subtypes are established and molecular biomarkers of the subtypes are yet to be discovered. Recent availability of multi-omics data paved a way to discover molecular biomarkers for the subtypes. However, multi-omics biomarker discovery is challenging because of the complex interplay between different omics layers. RESULTS: We propose a deep attention model named Gene-level biomarker discovery from multi-Omics data using graph ATtention neural network (GOAT) for identifying molecular biomarkers for eosinophilic asthma subtypes with multi-omics data. GOAT identifies genes that discriminate subtypes using a graph neural network by modeling complex interactions among genes as the attention mechanism in the deep learning model. In experiments with multi-omics profiles of the COREA (Cohort for Reality and Evolution of Adult Asthma in Korea) asthma cohort of 300 patients, GOAT outperforms existing models and suggests interpretable biological mechanisms underlying asthma subtypes. Importantly, GOAT identified genes that are distinct only in terms of relationship with other genes through attention. To better understand the role of biomarkers, we further investigated two transcription factors, CTNNB1 and JUN, captured by GOAT. We were successful in showing the role of the transcription factors in eosinophilic asthma pathophysiology in a network propagation and transcriptional network analysis, which were not distinct in terms of gene expression level differences. AVAILABILITY AND IMPLEMENTATION: Source code is available https://github.com/DabinJeong/Multi-omics_biomarker. The preprocessed data underlying this article is accessible in data folder of the github repository. Raw data are available in Multi-Omics Platform at http://203.252.206.90:5566/, and it can be accessible when requested.
Dabin Jeong, Bonil Koo, Minsik Oh, Tae-Bum Kim, Sun Kim
Bioinform.3
2021 IDEA: Integrating Divisive and Ensemble-Agglomerate hierarchical clustering framework for arbitrary shape data
abstract
Hierarchical clustering, a traditional clustering method, has been getting attention again. Among several reasons, a credit goes to a recent paper by Dasgupta in 2016 that proposed a cost function that quantitatively evaluates hierarchical clustering trees. An important question is how to combine this recent advance with existing successful clustering methods. In this paper, we propose a hierarchical clustering method to minimize the cost function of clustering tree by incorporating existing clustering techniques. First, we developed an ensemble tree-search method that finds an integrated tree with reduced cost by integrating multiple existing hierarchical clustering methods. Second, to operate on large and arbitrary shape data, we designed an efficient hierarchical clustering framework, called integrating divisive and ensemble-agglomerate (IDEA) by combining it with advanced clustering techniques such as nearest neighbor graph construction, divisive-agglomerate hybridization, and dynamic cut tree. The IDEA clustering method showed better performance in minimizing Dasgupta's cost and improving accuracy (adjusted rand index) over existing cost-minimization-based, and density-based hierarchical clustering methods in experiments using arbitrary shape datasets and complex biology-domain datasets.
Hongryul Ahn, Inuk Jung, Heejoon Chae, Minsik Oh, Inyoung Kim, Sun Kim
IEEE BigData4
2021 Machine learning-based analysis of multi-omics data on the cloud for investigating gene regulations
abstract
Gene expressions are subtly regulated by quantifiable measures of genetic molecules such as interaction with other genes, methylation, mutations, transcription factor and histone modifications. Integrative analysis of multi-omics data can help scientists understand the condition or patient-specific gene regulation mechanisms. However, analysis of multi-omics data is challenging since it requires not only the analysis of multiple omics data sets but also mining complex relations among different genetic molecules by using state-of-the-art machine learning methods. In addition, analysis of multi-omics data needs quite large computing infrastructure. Moreover, interpretation of the analysis results requires collaboration among many scientists, often requiring reperforming analysis from different perspectives. Many of the aforementioned technical issues can be nicely handled when machine learning tools are deployed on the cloud. In this survey article, we first survey machine learning methods that can be used for gene regulation study, and we categorize them according to five different goals: gene regulatory subnetwork discovery, disease subtype analysis, survival analysis, clinical prediction and visualization. We also summarize the methods in terms of multi-omics input types. Then, we explain why the cloud is potentially a good solution for the analysis of multi-omics data, followed by a survey of two state-of-the-art cloud systems, Galaxy and BioVLAB. Finally, we discuss important issues when the cloud is used for the analysis of multi-omics data for the gene regulation study.
Minsik Oh, Sun Kim, Heejoon Chae
Briefings Bioinform.1
2021 Erratum to: Machine learning-based analysis of multi-omics data on the cloud for investigating gene regulations
abstract
The first version of this article neglected to identify Minsik Oh and Sungjoon Park as joint first authors. This has now been corrected.
Minsik Oh, Sun Kim, Heejoon Chae
Briefings Bioinform.1
2021 mirTime: identifying condition-specific targets of microRNA in time-series transcript data using Gaussian process model and spherical vector clustering
abstract
BACKGROUND: MicroRNAs, small noncoding RNAs, are conserved in many species, and they are key regulators that mediate post-transcriptional gene silencing. Since biologists cannot perform experiments for each of target genes of thousands of microRNAs in numerous specific conditions, prediction on microRNA target genes has been extensively investigated. A general framework is a two-step process of selecting target candidates based on sequence and binding energy features and then predicting targets based on negative correlation of microRNAs and their targets. However, there are few methods that are designed for target predictions using time-series gene expression data. RESULTS: In this article, we propose a new pipeline, mirTime, that predicts microRNA targets by integrating sequence features and time-series expression profiles in a specific experimental condition. The most important feature of mirTime is that it uses the Gaussian process regression model to measure data at unobserved or unpaired time points. In experiments with two datasets in different experimental conditions and cell types, condition-specific target modules reported in the original papers were successfully predicted with our pipeline. The context specificity of target modules was assessed with three (correlation-based, target gene-based and network-based) evaluation criteria. mirTime showed better performance than existing expression-based microRNA target prediction methods in all three criteria. AVAILABILITY AND IMPLEMENTATION: mirTime is available at https://github.com/mirTime/mirtime. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hyejin Kang, Hongryul Ahn, Kyuri Jo, Minsik Oh, Sun Kim
Bioinform.4
2020 Per-Operation Reusability Based Allocation and Migration Policy for Hybrid Cache
abstract
Recently, a hybrid cache consisting of SRAM and STT-RAM has attracted much attention as a future memory by complementing each other with different memory characteristics. Prior works focused on developing data allocation and migration techniques considering write-intensity to reduce write energy at STT-RAM. However, these works often neglect the impact of operation-specific reusability of a cache line. In this paper, we propose an energy-efficient per-operation reusability-based allocation and migration policy (ORAM) with a unified LRU replacement policy. First, to select an adequate memory type for allocation, we propose a cost function based on per-operation reusability - gain from an allocated cache line and loss from an evicted cache line for different memory types - which exploits the temporal locality. Besides, we present a migration policy, victim and target cache line selection scheme, to resolve memory type inconsistency between replacement policy and the allocation policy, with further energy reduction. Experiment results show an average energy reduction in the LLC and the main memory by 12.3 and 21.2 percent, and the improvement of latency and execution time by 21.2 and 8.8 percent, respectively, compared with a baseline hybrid cache management. In addition, the Energy-Delay Product (EDP) is improved by 36.9 percent over the baseline.
Minsik Oh, Kwangsu Kim, Duheon Choi, Hyuk-Jun Lee, Eui-Young Chung
IEEE Trans. Computers1
2018 DeepFam: deep learning based alignment-free method for protein family modeling and prediction
abstract
Motivation: A large number of newly sequenced proteins are generated by the next-generation sequencing technologies and the biochemical function assignment of the proteins is an important task. However, biological experiments are too expensive to characterize such a large number of protein sequences, thus protein function prediction is primarily done by computational modeling methods, such as profile Hidden Markov Model (pHMM) and k-mer based methods. Nevertheless, existing methods have some limitations; k-mer based methods are not accurate enough to assign protein functions and pHMM is not fast enough to handle large number of protein sequences from numerous genome projects. Therefore, a more accurate and faster protein function prediction method is needed. Results: In this paper, we introduce DeepFam, an alignment-free method that can extract functional information directly from sequences without the need of multiple sequence alignments. In extensive experiments using the Clusters of Orthologous Groups (COGs) and G protein-coupled receptor (GPCR) dataset, DeepFam achieved better performance in terms of accuracy and runtime for predicting functions of proteins compared to the state-of-the-art methods, both alignment-free and alignment-based methods. Additionally, we showed that DeepFam has a power of capturing conserved regions to model protein families. In fact, DeepFam was able to detect conserved regions documented in the Prosite database while predicting functions of proteins. Our deep learning method will be useful in characterizing functions of the ever increasing protein sequences. Availability and implementation: Codes are available at https://bhi-kimlab.github.io/DeepFam.
Seokjun Seo, Minsik Oh, Youngjune Park, Sun Kim
Bioinform.2