Hyunsouk Cho

dblp:116/5184 · DBLP profile ↗
← Back
24ranked-venue papers
2as first author
12since 2021 · last 2026
0000-0002-9134-1921ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 12 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Collaborative adaptation without forgetting in source-free domain adaptation
Jisu Han, Joong-Won Hwang, Hyunsouk Cho, Wonjun Hwang
Pattern Recognit.3
2025 DCC: Differentiable Cardinality Constraints for Partial Index Tracking
abstract
Index tracking is a popular passive investment strategy aimed at optimizing portfolios, but fully replicating an index can lead to high transaction costs. To address this, partial replication have been proposed. However, the cardinality constraint renders the problem non-convex, non-differentiable, and often NP-hard, leading to the use of heuristic or neural network-based methods, which can be non-interpretable or have NP-hard complexity. To overcome these limitations, We propose a Differentiable Cardinality Constraint (DCC) for index tracking and introduce a floating-point precision-aware method to address implementation issues. We theoretically prove our methods calculate cardinality accurately and enforce actual cardinality with polynomial time complexity. We propose the range of the hyperparameter ensures that our method has no error in real implementations, based on theoretical proof and experiment. Our method applied to mathematical method outperforms baseline methods across various datasets, demonstrating the effectiveness of the identified hyperparameter.
Wooyeon Jo, Hyunsouk Cho
AAAI2
2025 Rethinking the Training Paradigm of Discrete Token-Based Multimodal LLMs: An Analysis of Text-Centric Bias
abstract
Discrete token-based multimodal large language models (MLLMs), such as AnyGPT and MIO, integrate diverse modalities into an autoregressive framework by discretizing modality inputs into tokens compatible with language models. Unlike encoder-based approaches, such as LLaVA and Flamingo, which utilize pretrained modality-specific encoders, discrete token-based MLLMs simultaneously learn modality token representations and their alignment with the language, yet are exclusively trained on modality-text paired datasets without additional unimodal training. We identify a structural limitation inherent in this training paradigm, termed text-centric bias, defined as an over-reliance on the textual context that restricts intrinsic modality understanding. To systematically analyze the existence of this bias, we propose an analytical framework involving external perplexity-based and internal neuron-level analyses. Furthermore, to verify whether the bias originates from the paired-only training paradigm, we introduce an analytical methodology named Monotune, which is a simple unimodal training stage. Our analyses demonstrate that minimal exposure to unimodal data effectively mitigates text-centric bias, providing empirical evidence that the bias is fundamentally induced by the paired-only training strategy. Through comprehensive downstream task evaluations, we further reveal that this structural bias meaningfully affects real-world multimodal task performance, particularly under limited textual contexts. Our findings highlight a fundamental limitation in current discrete token-based MLLM training paradigms and suggest directions for future multimodal training strategies. Our code and experiments are available at https://github.com/41312432/Monotune
Wansik Jo, Jooyeong Na, Soyeon Hong, Seungtaek Choi, Hyunsouk Cho
CIKM5
2025 Towards Fully-Automated Materials Discovery via Large-Scale Synthesis Dataset and Expert-Level LLM-as-a-Judge
abstract
Materials synthesis remains a critical bottleneck in developing innovations for energy storage, catalysis, electronics, and biomedical devices. Current synthesis design relies heavily on empirical trial-and-error methods guided by expert intuition, limiting the pace of materials discovery. To address this challenge, we present AlchemyBench, a comprehensive benchmark built upon a curated dataset of 17,667 expert-verified synthesis recipes from open-access literature.
Heegyu Kim, Taeyang Jeon, Seungtaek Choi, Jihoon Hong, Dongwon Jeon, Ga-Yeon Baek, Gyeong-Won Kwak, Jisu Bae, Yoon-Seo Kim, Seon-Jin Choi, Sung Beom Cho, Hyunsouk Cho
CIKM15
2025 FLEX: Expert-level False-Less EXecution Metric for Text-to-SQL Benchmark
abstract
Heegyu Kim, Jeon Taeyang, SeungHwan Choi, Seungtaek Choi, Hyunsouk Cho. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Heegyu Kim, Taeyang Jeon, Seunghwan Choi, Seungtaek Choi, Hyunsouk Cho
NAACL (Long Papers)5
2024 Can Large Language Models be Good Emotional Supporter? Mitigating Preference Bias on Emotional Support Conversation
abstract
Dongjin Kang, Sunghwan Kim, Taeyoon Kwon, Seungjun Moon, Hyunsouk Cho, Youngjae Yu, Dongha Lee, Jinyoung Yeo. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Dongjin Kang, Sunghwan Kim 0005, Taeyoon Kwon, Seungjun Moon, Hyunsouk Cho, Youngjae Yu, Dongha Lee 0003, Jinyoung Yeo
ACL (1)5
2024 OmniStitch: Depth-Aware Stitching Framework for Omnidirectional Vision with Multiple Cameras
abstract
Omnidirectional vision systems provide a 360-degree panoramic view, enabling full environmental awareness in various fields, such as Advanced Driver Assistance Systems (ADAS) and Virtual Reality (VR). Existing omnidirectional stitching methods rely on a single specialized 360-degree camera. However, due to hardware limitations such as high mounting heights and blind spots, adapting these methods to vehicles of varying sizes and geometries is challenging. These challenges include limited generalizability due to the reliance on predefined stitching regions for fixed camera arrays, performance degradation from distance parallax leading to large depth differences, and the absence of suitable datasets with ground truth for multi-camera omnidirectional systems. To overcome these challenges, we propose a novel omnidirectional stitching framework and a publicly available dataset tailored for varying distance scenarios with multiple cameras. The framework, referred to as OmniStitch, consists of a Stitching Region Maximization (SRM) module for automatic adaptation to different vehicles with multiple cameras and a Depth-Aware Stitching (DAS) module to handle depth differences caused by distance parallax between cameras. In addition, we create and release an omnidirectional stitching dataset, called GV360, which provides ground truth images that maintain the perspective of the 360-degree FOV, designed explicitly for vehicle-agnostic systems. Extensive evaluations of this dataset demonstrate that our framework outperforms state-of-the-art stitching models, especially in handling varying distance parallax. The proposed dataset and code are publicly available in https://github.com/tngh5004/Omnistitch.
Sooho Kim, Soyeon Hong, Kyungsoo Park, Hyunsouk Cho, Kyung-Ah Sohn 0001
ACM Multimedia4
2024 Multi-intent-aware Session-based Recommendation
abstract
Session-based recommendation (SBR) aims to predict the following item a user will interact with during an ongoing session. Most existing SBR models focus on designing sophisticated neural-based encoders to learn a session representation, capturing the relationship among session items. However, they tend to focus on the last item, neglecting diverse user intents that may exist within a session. This limitation leads to significant performance drops, especially for longer sessions. To address this issue, we propose a novel SBR model, called Multi-intent-aware Session-based Recommendation Model (MiaSRec). It adopts frequency embedding vectors indicating the item frequency in session to enhance the information about repeated items. MiaSRec represents various user intents by deriving multiple session representations centered on each item and dynamically selecting the important ones. Extensive experimental results show that MiaSRec outperforms existing state-of-the-art SBR models on six datasets, particularly those with longer average session length, achieving up to 6.27% and 24.56% gains for MRR@20 and Recall@20. Our code is available at https://github.com/jin530/MiaSRec.
Minjin Choi 0001, Hye-young Kim, Hyunsouk Cho, Jongwuk Lee
SIGIR3
2022 FPAdaMetric: False-Positive-Aware Adaptive Metric Learning for Session-Based Recommendation
abstract
Modern recommendation systems are mostly based on implicit feedback data which can be quite noisy due to false positives (FPs) caused by many reasons, such as misclicks or quick curiosity. Numerous recommendation algorithms based on collaborative filtering have leveraged post-click user behavior (e.g., skip) to identify false positives. They effectively involved these false positives in the model supervision as negative-like signals. Yet, false positives had not been considered in existing session-based recommendation systems (SBRs) although they provide just as deleterious effects. To resolve false positives in SBRs, we first introduce FP-Metric model which reformulates the objective of the session-based recommendation with FP constraints into metric learning regularization. In addition, we propose FP-AdaMetric that enhances the metric-learning regularization terms with an adaptive module that elaborately calculates the impact of FPs inside sequential patterns. We verify that FP-AdaMetric improves several session-based recommendation models' performances in terms of Hit Rate (HR), MRR, and NDCG on datasets from different domains including music, movie, and game. Furthermore, we show that the adaptive module plays a much more crucial role in FP-AdaMetric model than in other baselines.
Jongwon Jeong, Jeong Choi, Hyunsouk Cho, Sehee Chung
AAAI3
2022 Towards Proper Contrastive Self-Supervised Learning Strategies for Music Audio Representation
abstract
The common research goal of self-supervised learning is to extract a general representation which an arbitrary downstream task would benefit from. In this work, we investigate music audio representation learned from different contrastive self-supervised learning schemes and empirically evaluate the embedded vectors on various music information retrieval (MIR) tasks where different levels of the music perception are concerned. We analyze the results to discuss the proper direction of contrastive learning strategies for different MIR tasks. We show that these representations convey a comprehensive information about the auditory characteristics of music in general, although each of the self-supervision strategies has its own effectiveness in certain aspect of information.
Jeong Choi, Seongwon Jang, Hyunsouk Cho, Sehee Chung
ICME3
2021 Self-Supervised Multimodal Opinion Summarization
abstract
Jinbae Im, Moonki Kim, Hoyeop Lee, Hyunsouk Cho, Sehee Chung. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Jinbae Im, Moonki Kim, Hoyeop Lee, Hyunsouk Cho, Sehee Chung
ACL/IJCNLP (1)4
2021 One-Step Pixel-Level Perturbation-Based Saliency Detector
Vinnam Kim, Hyunsouk Cho, Sehee Chung
BMVC2
2020 CITIES: Contextual Inference of Tail-item Embeddings for Sequential Recommendation
abstract
Sequential recommendation techniques provide users with product recommendations fitting their current preferences by handling dynamic user preferences over time. Previous studies have focused on modeling sequential dynamics without much regard to which of the best-selling products (i.e., head items) or niche products (i.e., tail items) should be recommended. We scrutinize the structural reason for why tail items are barely served in the current sequential recommendation model, which consists of an item-embedding layer, a sequence-modeling layer, and a recommendation layer. Well-designed sequence-modeling and recommendation layers are expected to naturally learn suitable item embeddings. However, tail items are likely to fall short of this expectation because the current model structure is not suitable for learning high-quality embeddings with insufficient data. Thus, tail items are rarely recommended. To eliminate this issue, we propose a framework called CITIES, which aims to enhance the quality of the tail-item embeddings by training an embedding-inference function using multiple contextual head items so that the recommendation performance improves for not only the tail items but also for the head items. Moreover, our framework can infer new-item embeddings without an additional learning process. Extensive experiments on two real-world datasets show that applying CITIES to the state-of-the-art methods improves recommendation performance for both tail and head items. We conduct an additional experiment to verify that CITIES can infer suitable new-item embeddings as well.
Seongwon Jang, Hoyeop Lee, Hyunsouk Cho, Sehee Chung
ICDM3
2020 SQuAD2-CR: Semi-supervised Annotation for Cause and Rationales for Unanswerability in SQuAD 2.0
abstract
Existing machine reading comprehension models are reported to be brittle for adversarially perturbed questions when optimizing only for accuracy, which led to the creation of new reading comprehension benchmarks, such as SQuAD 2.0 which contains such type of questions. However, despite the super-human accuracy of existing models on such datasets, it is still unclear how the model predicts the answerability of the question, potentially due to the absence of a shared annotation for the explanation. To address such absence, we release SQuAD2-CR dataset, which contains annotations on unanswerable questions from the SQuAD 2.0 dataset, to enable an explanatory analysis of the model prediction. Specifically, we annotate (1) explanation on why the most plausible answer span cannot be the answer and (2) which part of the question causes unanswerability. We share intuitions and experimental results that how this dataset can be used to analyze and improve the interpretability of existing reading comprehension model behavior.
Gyeongbok Lee, Seung-won Hwang, Hyunsouk Cho
LREC3
2019 MeLU: Meta-Learned User Preference Estimator for Cold-Start Recommendation
abstract
This paper proposes a recommender system to alleviate the cold-start problem that can estimate user preferences based on only a small number of items. To identify a user's preference in the cold state, existing recommender systems, such as Netflix, initially provide items to a user; we call those items evidence candidates. Recommendations are then made based on the items selected by the user. Previous recommendation studies have two limitations: (1) the users who consumed a few items have poor recommendations and (2) inadequate evidence candidates are used to identify user preferences. We propose a meta-learning-based recommender system called MeLU to overcome these two limitations. From meta-learning, which can rapidly adopt new task with a few examples, MeLU can estimate new user's preferences with a few consumed items. In addition, we provide an evidence candidate selection strategy that determines distinguishing items for customized preference estimation. We validate MeLU with two benchmark datasets, and the proposed model reduces at least 5.92% mean absolute error than two comparative models on the datasets. We also conduct a user study experiment to verify the evidence selection strategy.
Hoyeop Lee, Jinbae Im, Seongwon Jang, Hyunsouk Cho, Sehee Chung
KDD4
2018 Machine-Translated Knowledge Transfer for Commonsense Causal Reasoning
abstract
This paper studies the problem of multilingual causal reasoning in resource-poor languages. Existing approaches, translating into the most probable resource-rich language such as English, suffer in the presence of translation and language gaps between different cultural area, which leads to the loss of causality. To overcome these challenges, our goal is thus to identify key techniques to construct a new causality network of cause-effect terms, targeted for the machine-translated English, but without any language-specific knowledge of resource-poor languages. In our evaluations with three languages, Korean, Chinese, and French, our proposed method consistently outperforms all baselines, achieving up-to 69.0% reasoning accuracy, which is close to the state-of-the-art accuracy 70.2% achieved on English.
Jinyoung Yeo, Hyunsouk Cho, Seungtaek Choi, Seung-won Hwang
AAAI3
2018 Visual Choice of Plausible Alternatives: An Evaluation of Image-based Commonsense Causal Reasoning
Jinyoung Yeo, Gyeongbok Lee, Seungtaek Choi, Hyunsouk Cho, Reinald Kim Amplayo, Seung-won Hwang
LREC5
2017 Gradable Adjective Embedding for Commonsense Knowledge
Kyungjae Lee 0002, Hyunsouk Cho, Seung-won Hwang
PAKDD (2)2
2017 Multimodal KB Harvesting for Emerging Spatial Entities
abstract
New entities are being created daily. Though the novelty of these entities naturally attracts mentions, due to lack of prior knowledge, it is more challenging to collect knowledge about such entities than pre-existing entities, whose KBs are comprehensively annotated through LBSNs and EBSNs. In this paper, we focus on knowledge harvesting for emerging spatial entities (ESEs), such as new businesses and venues, assuming we have only a list of ESE names. Existing techniques for knowledge base (KB) harvesting are primarily associated with information extraction from textual corpora. In contrast, we propose a multimodal method for event detection based on the complementary interaction of image, text, and user information between multi-source platforms, namely Flickr and Twitter. We empirically validate our harvesting approaches improve the quality of KB with enriched place and event knowledge.
Jinyoung Yeo, Hyunsouk Cho, Seung-won Hwang
IEEE Trans. Knowl. Data Eng.2
2016 ECO: Entity-level captioning in context
Hyunsouk Cho, Seung-won Hwang
ASONAM1
2016 Event Grounding from Multimodal Social Network Fusion
abstract
This paper studies the problem of extracting real world event information from social media streams. Although existing work focuses on event signals of bursty mentions extracted from a single-source of textual streams, these signals are likely to be noisy due to ambiguous occurrences of individual mentions. To extract accurate event signals, we propose a framework capable of "grounding" mentions to unique event using multiple social networks with complementary strength. We show that our framework jointly using multiple sources outperforms state-of-the-arts using publicly available datasets.
Hyunsouk Cho, Jinyoung Yeo, Seung-won Hwang
ICDM1
2014 Toward Scalable Indexing for Top-k Queries
abstract
A top-k query retrieves the best k tuples by assigning scores for each tuple in a target relation with respect to a user-specific scoring function. This paper studies the problem of constructing an indexing structure for supporting top-k queries over varying scoring functions and retrieval sizes. The existing research efforts can be categorized into three approaches: list-, layer-, and view-based approaches. In this paper, we mainly focus on the layer-based approach that pre-materializes tuples into consecutive multiple layers. We first propose a dual-resolution layer that consists of coarse-level and fine-level layers. Specifically, we build coarse-level layers using skylines, and divide each coarse-level layer into fine-level sublayers using convex skylines. To make our proposed dual-resolution layer scalable, we then address the following optimization directions: 1) index construction; 2) disk-based storage scheme; 3) the design of the virtual layer; and 4) index maintenance for tuple updates. Our evaluation results show that our proposed method is more scalable than the state-of-the-art methods.
Jongwuk Lee, Hyunsouk Cho, Sunyou Lee, Seung-won Hwang
IEEE Trans. Knowl. Data Eng.2
2013 Hybrid entity clustering using crowds and data
Jongwuk Lee, Hyunsouk Cho, Young-rok Cha, Seung-won Hwang, Zaiqing Nie, Ji-Rong Wen
VLDB J.2
2012 Efficient Dual-Resolution Layer Indexing for Top-k Queries
abstract
Top-k queries have gained considerable attention as an effective means for narrowing down the overwhelming amount of data. This paper studies the problem of constructing an indexing structure that efficiently supports top-k queries for varying scoring functions and retrieval sizes. The existing work can be categorized into three classes: list-, layer-, and view-based approaches. This paper focuses on the layer-based approach, pre-materializing tuples into consecutive multiple layers. The layer-based index enables us to return top-k answers efficiently by restricting access to tuples in the k layers. However, we observe that the number of tuples accessed in each layer can be reduced further. For this purpose, we propose a dual-resolution layer structure. Specifically, we iteratively build coarse-level layers using skylines, and divide each coarse-level layer into fine-level sub layers using convex skylines. The dual-resolution layer is able to leverage not only the dominance relationship between coarse-level layers, named for all-dominance, but also a relaxed dominance relationship between fine-level sub layers, named exists-dominance. Our extensive evaluation results demonstrate that our proposed method significantly reduces the number of tuples accessed than the state-of-the-art methods.
Jongwuk Lee, Hyunsouk Cho, Seung-won Hwang
ICDE2