EDBT 2026 Demo / reviewers in the wild / expert
Eun-Sol Kim
dblp:52/10086
· DBLP profile ↗
21ranked-venue papers
5as first author
14since 2021 · last 2025
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Zero-Shot Compositional Video Learning with Coding Rate Reduction
Heeseok Jung, Jun-Hyeon Bak, Yujin Jeong, Gyugeun Lee, Jinwoo Ahn, Eun-Sol Kim |
ICCV | 6 |
| 2025 | Structure-guided sequence representation learning for generalizable protein function predictionabstractMOTIVATION: Accurately predicting protein function from sequence remains a fundamental yet challenging goal in computational biology. Although recent advances have enabled the reliable prediction of protein 3D structures from sequences, utilizing structural information alone for functional inference has shown limited success. To address this gap, previous work has explored the integration of sequence and structural data by representing proteins as graphs, where residues are modeled as nodes, and spatial proximity defines edges. However, since the number of amino acids can vary significantly between proteins, the resulting graphs, constructed based on amino acids, also differ greatly in size. This large variation poses a challenge, as it becomes extremely difficult to extract generalizable information from graphs of such differing scales accurately. In this work, we propose Structure-guided Sequence Representation Learning, a novel framework that incorporates structural knowledge to extract informative, multiscale features directly from protein sequences. By embedding structural information into a sequence-based learning paradigm, our method captures functionally meaningful representations more effectively. Furthermore, we present a generalizable model architecture designed for multitask learning and inference, offering improved performance and flexibility over traditional task-specific approaches to protein function prediction. RESULTS: In this article, we demonstrate that the proposed novel attention pooling method on protein graphs effectively integrates global structural features and local chemical properties of amino acids in various-length proteins. Through this approach, we improve performance in tasks related to predicting protein functions, functional expression sites, and their relationships with structure and sequence. By effectively extracting the information needed to predict multiple protein functions simultaneously, we improve efficiency by eliminating the need for separate learning. AVAILABILITY AND IMPLEMENTATION: The code implementation is available at https://github.com/vanha9/S2RL_protein and has also been archived on zenodo: https://doi.org/10.5281/zenodo.16441001. Seokjun On, Yujin Jeong, Eun-Sol Kim |
Bioinform. | 3 |
| 2024 | Structure-Aware Multimodal Sequential Learning for Visual DialogabstractWith the ability to collect vast amounts of image and natural language data from the web, there has been a remarkable advancement in Large-scale Language Models (LLMs). This progress has led to the emergence of chatbots and dialogue systems capable of fluent conversations with humans. As the variety of devices enabling interactions between humans and agents expands, and the performance of text-based dialogue systems improves, there has been recently proposed research on visual dialog. However, visual dialog requires understanding sequences of pairs consisting of images and sentences, making it challenging to gather sufficient data for training large-scale models from the web. In this paper, we propose a new multimodal learning method leveraging existing large-scale models designed for each modality, to enable model training for visual dialog with small visual dialog datasets. The key ideas of our approach are: 1) storing the history or context during the progression of visual dialog in the form of spatiotemporal graphs, and 2) introducing small modulation blocks between modality-specific models and the graphs to align the semantic spaces. For implementation, we introduce a novel structure-aware cross-attention method, which retrieves relevant image and text knowledge for utterance generation from the pretrained models. For experiments, we achieved a new state-of-the-art performance on three visual dialog datasets, including the most challenging one COMET. Kyunghwan An, Jinwoo Ahn, Jaeseok Kim, Yu-Jung Heo, Du-Seong Chang, Eun-Sol Kim |
AAAI | 8 |
| 2024 | Compositional Video Understanding with Spatiotemporal Structure-based TransformersabstractIn this paper, we suggest a new novel method to understand complex semantic structures through long video inputs. Conventional methods for understanding videos have been focused on short-term clips, and trained to get visual representations for the short clips using convolutional neural networks or transformer architectures. However, most real-world videos are composed of long videos ranging from minutes to hours, therefore, it essentially brings limitations to understanding the overall semantic structures of the long videos by dividing them into small clips and learning the representations of them. We suggest a new algorithm to learn the multi-granular semantic structures of videos, by defining spatiotemporal high-order relationships among object-based representations as semantic units. The proposed method includes a new transformer architecture capable of learning spatiotemporal graphs, and a compositional learning method to learn disentangled features for each semantic unit. Using the suggested method, we resolve the challenging video task, which is compositional generalization understanding of unseen videos. In experiments, we demonstrate new state-of-the-art performances for two challenging video datasets. Hoyeoung Yun, Jinwoo Ahn, Eun-Sol Kim |
CVPR | 4 |
| 2024 | Clustering-based Image-Text Graph Matching for Domain Generalization
Nokyung Park, Daewon Chae, Jeongyong Shim, Sangpil Kim, Eun-Sol Kim, Jinkyu Kim 0001 |
ICPR (10) | 5 |
| 2023 | Dense but Efficient VideoQA for Intricate Compositional ReasoningabstractIt is well known that most of the conventional video question answering (VideoQA) datasets consist of easy questions requiring simple reasoning processes. However, long videos inevitably contain complex and compositional semantic structures along with the spatio-temporal axis, which requires a model to understand the compositional structures inherent in the videos. In this paper, we suggest a new compositional VideoQA method based on transformer architecture with a deformable attention mechanism to address the complex VideoQA tasks. The deformable attentions are introduced to sample a subset of informative visual features from the dense visual feature map to cover a temporally long range of frames efficiently. Furthermore, the dependency structure within the complex question sentences is also combined with the language embeddings to readily understand the relations among question words. Extensive experiments and ablation studies show that the suggested dense but efficient model outperforms other baselines. Jihyeon Lee, Woo-Young Kang, Eun-Sol Kim |
WACV | 3 |
| 2023 | DeepFold: enhancing protein structure prediction through optimized loss functions, improved template features, and re-optimized energy functionabstractMOTIVATION: Predicting protein structures with high accuracy is a critical challenge for the broad community of life sciences and industry. Despite progress made by deep neural networks like AlphaFold2, there is a need for further improvements in the quality of detailed structures, such as side-chains, along with protein backbone structures. RESULTS: Building upon the successes of AlphaFold2, the modifications we made include changing the losses of side-chain torsion angles and frame aligned point error, adding loss functions for side chain confidence and secondary structure prediction, and replacing template feature generation with a new alignment method based on conditional random fields. We also performed re-optimization by conformational space annealing using a molecular mechanics energy function which integrates the potential energies obtained from distogram and side-chain prediction. In the CASP15 blind test for single protein and domain modeling (109 domains), DeepFold ranked fourth among 132 groups with improvements in the details of the structure in terms of backbone, side-chain, and Molprobity. In terms of protein backbone accuracy, DeepFold achieved a median GDT-TS score of 88.64 compared with 85.88 of AlphaFold2. For TBM-easy/hard targets, DeepFold ranked at the top based on Z-scores for GDT-TS. This shows its practical value to the structural biology community, which demands highly accurate structures. In addition, a thorough analysis of 55 domains from 39 targets with publicly available structures indicates that DeepFold shows superior side-chain accuracy and Molprobity scores among the top-performing groups. AVAILABILITY AND IMPLEMENTATION: DeepFold tools are open-source software available at https://github.com/newtonjoo/deepfold. Jae-Won Lee, Jong-Hyun Won, Seonggwang Jeon, Yujin Choo, Yubin Yeon, Jin-Seon Oh, Seonhwa Kim, InSuk Joung, Cheongjae Jang, Sung Jong Lee, Kyong Hwan Jin, Giltae Song, Eun-Sol Kim, Jejoong Yoo, Eunok Paek, Yung-Kyun Noh, Keehyoung Joo |
Bioinform. | 15 |
| 2022 | BaSSL: Boundary-aware Self-Supervised Learning for Video Scene Segmentation
Jonghwan Mun, Minchul Shin, Gunsoo Han, Seongsu Ha, Joonseok Lee, Eun-Sol Kim |
ACCV (4) | 7 |
| 2022 | Hypergraph Transformer: Weakly-Supervised Multi-hop Reasoning for Knowledge-based Visual Question AnsweringabstractKnowledge-based visual question answering (QA) aims to answer a question which requires visually-grounded external knowledge beyond image content itself.Answering complex questions that require multi-hop reasoning under weak supervision is considered as a challenging problem since i) no supervision is given to the reasoning process and ii) highorder semantics of multi-hop knowledge facts need to be captured.In this paper, we introduce a concept of hypergraph to encode highlevel semantics of a question and a knowledge base, and to learn high-order associations between them.The proposed model, Hypergraph Transformer, constructs a question hypergraph and a query-aware knowledge hypergraph, and infers an answer by encoding inter-associations between two hypergraphs and intra-associations in both hypergraph itself.Extensive experiments on two knowledgebased visual QA and two knowledge-based textual QA demonstrate the effectiveness of our method, especially for multi-hop reasoning problem.Our source code is available at https://github.com/yujungheo/ kbvqa-public. Yu-Jung Heo, Eun-Sol Kim, Woosuk Choi, Byoung-Tak Zhang |
ACL (1) | 2 |
| 2022 | Selective Token Generation for Few-shot Natural Language GenerationabstractNatural language modeling with limited training data is a challenging problem, and many algorithms make use of large-scale pretrained language models (PLMs) for this due to its great generalization ability. Among them, additive learning that incorporates a task-specific adapter on top of the fixed large-scale PLM has been popularly used in the few-shot setting. However, this added adapter is still easy to disregard the knowledge of the PLM especially for few-shot natural language generation (NLG) since an entire sequence is usually generated by only the newly trained adapter. Therefore, in this work, we develop a novel additive learning algorithm based on reinforcement learning (RL) that selectively outputs language tokens between the task-general PLM and the task-specific adapter during both training and inference. This output token selection over the two generators allows the adapter to take into account solely the task-relevant parts in sequence generation, and therefore makes it more robust to overfitting as well as more stable in RL training. In addition, to obtain the complementary adapter from the PLM for each few-shot task, we exploit a separate selecting module that is also simultaneously trained using RL. Experimental results on various few-shot NLG tasks including question answering, data-to-text generation and text summarization demonstrate that the proposed selective token generation significantly outperforms the previous additive learning algorithms based on the PLMs. Daejin Jo, Taehwan Kwon, Eun-Sol Kim, Sungwoong Kim |
COLING | 3 |
| 2022 | MSTR: Multi-Scale Transformer for End-to-End Human-Object Interaction DetectionabstractHuman-Object Interaction (HOI) detection is the task of identifying a set of (human, object, interaction) triplets from an image. Recent work proposed transformer encoder-decoder architectures that successfully eliminated the need for many hand-designed components in HOI detection through end-to-end training. However, they are limited to single-scale feature resolution, providing suboptimal performance in scenes containing humans, objects, and their interactions with vastly different scales and distances. To tackle this problem, we propose a Multi-Scale TRansformer (MSTR) for HOI detection powered by two novel HOI-aware deformable attention modules called Dual-Entity attention and Entity-conditioned Context attention. While existing deformable attention comes at a huge cost in HOI detection performance, our proposed attention modules of MSTR learn to effectively attend to sampling points that are essential to identify interactions. In experiments, we achieve the new state-of-the-art performance on two HOI detection benchmarks. Bumsoo Kim 0005, Jonghwan Mun, Kyoung-Woon On, Minchul Shin, Jun-Hyun Lee, Eun-Sol Kim |
CVPR | 6 |
| 2022 | Video-Text Representation Learning via Differentiable Weak Temporal AlignmentabstractLearning generic joint representations for video and text by a supervised method requires a prohibitively substantial amount of manually annotated video datasets. As a practical alternative, a large-scale but uncurated and narrated video dataset, HowTo100M, has recently been introduced. But it is still challenging to learn joint embeddings of video and text in a self-supervised manner, due to its ambiguity and non-sequential alignment. In this paper, we propose a novel multi-modal self-supervised framework Video-Text Temporally Weak Alignment-based Contrastive Learning (VT-TWINS) to capture significant information from noisy and weakly correlated data using a variant of Dynamic Time Warping (DTW). We observe that the standard DTW inherently cannot handle weakly correlated data and only considers the globally optimal alignment path. To address these problems, we develop a differentiable DTW which also reflects local information with weak temporal alignment. Moreover, our proposed model applies a contrastive learning scheme to learn feature representations on weakly correlated data. Our extensive experiments demonstrate that VT-TWINS attains significant improvements in multi-modal representation learning and outperforms various challenging downstream tasks. Code is available at https://github.com/mlvlab/VT-Twins. Dohwan Ko, Joonmyung Choi, Juyeon Ko, Shinyeong Noh, Kyoung-Woon On, Eun-Sol Kim, Hyunwoo J. Kim |
CVPR | 6 |
| 2021 | Image-to-Image Retrieval by Learning Similarity between Scene GraphsabstractAs a scene graph compactly summarizes the high-level content of an image in a structured and symbolic manner, the similarity between scene graphs of two images reflects the relevance of their contents. Based on this idea, we propose a novel approach for image-to-image retrieval using scene graph similarity measured by graph neural networks. In our approach, graph neural networks are trained to predict the proxy image relevance measure, computed from human-annotated captions using a pre-trained sentence similarity model. We collect and publish the dataset for image relevance measured by human annotators to evaluate retrieval algorithms. The collected dataset shows that our method agrees well with the human perception of image similarity than other competitive baselines. Sangwoong Yoon, Woo-Young Kang, Sungwook Jeon, SeongEun Lee, Changjin Han, Eun-Sol Kim |
AAAI | 7 |
| 2021 | HOTR: End-to-End Human-Object Interaction Detection With TransformersabstractHuman-Object Interaction (HOI) detection is a task of identifying "a set of interactions" in an image, which involves the i) localization of the subject (i.e., humans) and target (i.e., objects) of interaction, and ii) the classification of the interaction labels. Most existing methods have indirectly addressed this task by detecting human and object instances and individually inferring every pair of the detected instances. In this paper, we present a novel framework, referred by HOTR, which directly predicts a set of 〈human, object, interaction〉 triplets from an image based on a transformer encoder-decoder architecture. Through the set prediction, our method effectively exploits the inherent semantic relationships in an image and does not require time-consuming post-processing which is the main bottleneck of existing methods. Our proposed algorithm achieves the state-of-the-art performance in two HOI detection benchmarks with an inference time under 1 ms after object detection. Bumsoo Kim 0005, Jun-Hyun Lee, Jaewoo Kang, Eun-Sol Kim, Hyunwoo J. Kim |
CVPR | 4 |
| 2020 | Cut-Based Graph Learning Networks to Discover Compositional Structure of Sequential Video DataabstractConventional sequential learning methods such as Recurrent Neural Networks (RNNs) focus on interactions between consecutive inputs, i.e. first-order Markovian dependency. However, most of sequential data, as seen with videos, have complex dependency structures that imply variable-length semantic flows and their compositions, and those are hard to be captured by conventional methods. Here, we propose Cut-Based Graph Learning Networks (CB-GLNs) for learning video data by discovering these complex structures of the video. The CB-GLNs represent video data as a graph, with nodes and edges corresponding to frames of the video and their dependencies respectively. The CB-GLNs find compositional dependencies of the data in multilevel graph forms via a parameterized kernel with graph-cut and a message passing framework. We evaluate the proposed method on the two different tasks for video understanding: Video theme classification (Youtube-8M dataset (Abu-El-Haija et al. 2016)) and Video Question and Answering (TVQA dataset(Lei et al. 2018)). The experimental results show that our model efficiently learns the semantic compositional structure of video data. Furthermore, our model achieves the highest performance in comparison to other baseline methods. Kyoung-Woon On, Eun-Sol Kim, Yu-Jung Heo, Byoung-Tak Zhang |
AAAI | 2 |
| 2020 | Hypergraph Attention Networks for Multimodal LearningabstractOne of the fundamental problems that arise in multimodal learning tasks is the disparity of information levels between different modalities. To resolve this problem, we propose Hypergraph Attention Networks (HANs), which define a common semantic space among the modalities with symbolic graphs and extract a joint representation of the modalities based on a co-attention map constructed in the semantic space. HANs follow the process: constructing the common semantic space with symbolic graphs of each modality, matching the semantics between sub-structures of the symbolic graphs, constructing co-attention maps between the graphs in the semantic space, and integrating the multimodal inputs using the co-attention maps to get the final joint representation. From the qualitative analysis with two Visual Question and Answering datasets, we discover that 1) the alignment of the information levels between the modalities is important, and 2) the symbolic graphs are very powerful ways to represent the information of the low-level signals in alignment. Moreover, HANs dramatically improve the state-of-the-art accuracy on the GQA dataset from 54.6\% to 61.88\% only using the symbolic information in quantitatively. Eun-Sol Kim, Woo-Young Kang, Kyoung-Woon On, Yu-Jung Heo, Byoung-Tak Zhang |
CVPR | 1 |
| 2016 | DeepSchema: Automatic Schema Acquisition from Wearable Sensor Data in Restaurant Situations
Eun-Sol Kim, Kyoung-Woon On, Byoung-Tak Zhang |
IJCAI | 1 |
| 2015 | Analyzing Human Behavioral Data to Interact with Restaurant Server AgentsabstractIn this paper, we consider a problem of analyzing human behavioral data to predict the human cognitive states and generate corresponding actions of sever-agent. Specifically, we aim at predicting human cognitive states during meal time and generating relevant dining services for the human. For this study, we collect behavioral data using 2 kinds of wearable devices, which are an eye tracker and a watch type EDA device, during meal time. We focus on the characteristics of the behavioral data, which are heterogeneous, noisy and temporal, and suggest a novel machine learning algorithm which can analyze the data integrally. Suggested model has hierarchical structure: the bottom layer combines the multi-modal behavioral data based on causal structure of the data and extracts the feature vector. Using the extracted feature vectors, the upper layer predicts the cognitive states based on temporal correlation between feature vectors. Experimental results show that the suggested model can analyze the behavioral data efficiently and predict the human cognitive states correctly. Eun-Sol Kim, Kyoung-Woon On, Byoung-Tak Zhang |
HAI | 1 |
| 2012 | Hierarchical Slow-Feature Models of Gesture Conversation
Jiseob Kim, Sooyong Jang, Eun-Sol Kim, Byoung-Tak Zhang |
CogSci | 3 |
| 2012 | 'Is this right?' or 'Is that wrong?': Evidence from Dynamic Eye-Hand Movement in Decision Making
Eun-Sol Kim, Jiseob Kim, Thies Pfeiffer, Ipke Wachsmuth, Byoung-Tak Zhang |
CogSci | 1 |
| 2011 | Mutual information-based evolution of hypernetworks for brain data analysisabstractCortical analysis becomes increasingly important for brain research and clinical diagnosis. This problem involves a combinatorial search to find the essential modules among a large number of brain regions. Despite several statistical approaches, cortical analysis remains a formidable challenge due to high dimensionality and sparsity of data. Here we describe an evolutionary method for finding significant modules from cortical data. The method uses a hypernetwork which is encoded as a population of hyperedges, where hyperedges represent building blocks or potential modules. We develop an efficient method for evolving the hypernetwork using mutual information to generate essential hyperedges. We evaluate the method on predicting intelligence quotient (IQ) levels and finding potential significant modules on IQ from brain MRI data consisting of 62 healthy adults with over 80,000 measured points (variables). The experimental results show that our information-theoretic evolutionary hypernetworks improve the classification accuracy by 5-15%. Moreover, it extracts significant cortical modules that distinguish high IQ from low IQ groups. Eun-Sol Kim, Jung-Woo Ha 0001, Wi Hoon Jung, Joon Hwan Jang, Jun Soo Kwon, Byoung-Tak Zhang |
IEEE Congress on Evolutionary Computation | 1 |