Hyounghun Kim

dblp:228/9951 · DBLP profile ↗
← Back
17ranked-venue papers
7as first author
13since 2021 · last 2026
0009-0008-3382-7510ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 7 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Can LLMs Estimate Cognitive Complexity of Reading Comprehension Items?
abstract
Estimating the cognitive complexity of reading comprehension (RC) items is crucial for assessing item difficulty before it is administered to learners.Unlike syntactic and semantic features, such as passage length or semantic similarity between options, cognitive features that arise during answer reasoning are not readily extractable using existing NLP tools and have traditionally relied on human annotation.In this study, we examine whether large language models (LLMs) can estimate the cognitive complexity of RC items by focusing on two dimensions-Evidence Scope and Transformation Level-that indicate the degree of cognitive burden involved in reasoning about the answer.Our experimental results demonstrate that LLMs can approximate the cognitive complexity of items, indicating their potential as tools for prior difficulty analysis.Further analysis reveals a gap between LLMs' reasoning ability and their metacognitive awareness: even when they produce correct answers, they sometimes fail to correctly identify the features underlying their own reasoning process.
Seonjeong Hwang, Hyounghun Kim, Gary Geunbae Lee
ACL (1)2
2026 A Multi-Agent Framework for Feature-Constrained Difficulty Control in Reading Comprehension Item Generation
abstract
Recent studies in difficulty-controlled reading comprehension item generation have leveraged large language models (LLMs) to produce items by adjusting difficulty-related features.However, existing methods typically rely on a single-agent prompting approach, which often fails to consistently satisfy specified feature constraints, resulting in items that deviate from the target difficulty level.To address this limitation, we introduce MAFIG, a Multiagent Framework for Feature-constrained Item Generation, where multiple LLM agents and feature-specific evaluators collaborate to generate and iteratively revise items based on intended constraints.Furthermore, to verify the efficacy of MAFIG in difficulty control, we propose a method for constructing a sequence of feature constraint sets that yield items with monotonically increasing difficulty.Experimental results demonstrate that MAFIG generates items that adhere to target constraints at a significantly higher rate than baselines, achieving robust difficulty control through the difficulty-calibrated constraint sequence.
Seonjeong Hwang, Jun Seo, Hyounghun Kim, Gary Geunbae Lee
ACL (1)3
2026 Mixture-of-Experts with Intermediate CTC Supervision for Accented Speech Recognition
abstract
Accented speech remains a persistent challenge for automatic speech recognition (ASR), as most models are trained on data dominated by a few high-resource English varieties, leading to substantial performance degradation for other accents.Accent-agnostic approaches improve robustness yet struggle with heavily accented or unseen varieties, while accent-specific methods rely on limited and often noisy labels.We introduce MOE-CTC, a Mixture-of-Experts architecture with intermediate CTC supervision that jointly promotes expert specialization and generalization.During training, accent-aware routing encourages experts to capture accentspecific patterns, which gradually transitions to label-free routing for inference.Each expert is equipped with its own CTC head to align routing with transcription quality, and a routing-augmented loss further stabilizes optimization.Experiments on the MCV-ACCENT benchmark demonstrate consistent gains across both seen and unseen accents in low-and highresource conditions, achieving up to 29.3% relative WER reduction over strong FastConformer baselines.
Hyounghun Kim, Gary Geunbae Lee
ACL (1)2
2025 Enabling Chatbots with Eyes and Ears: An Immersive Multimodal Conversation System for Dynamic Interactions
abstract
As chatbots continue to evolve toward humanlike, real-world, interactions, multimodality remains an active area of research and exploration.So far, efforts to integrate multimodality into chatbots have primarily focused on image-centric tasks, such as visual dialogue and image-based instructions, placing emphasis on the "eyes" of human perception while neglecting the "ears", namely auditory aspects.Moreover, these studies often center around static interactions that focus on discussing the modality rather than naturally incorporating it into the conversation, which limits the richness of simultaneous, dynamic engagement.Furthermore, while multimodality has been explored in multi-party and multi-session conversations, task-specific constraints have hindered its seamless integration into dynamic, natural conversations.To address these challenges, this study aims to equip chatbots with "eyes and ears" capable of more immersive interactions with humans.As part of this effort, we introduce a new multimodal conversation dataset, Multimodal Multi-Session Multi-Party Conversation (M 3 C), and propose a novel multimodal conversation model featuring multimodal memory retrieval.Our model, trained on the M 3 C, demonstrates the ability to seamlessly engage in long-term conversations with multiple speakers in complex, real-world-like settings, effectively processing visual and auditory inputs to understand and respond appropriately.Human evaluations highlight the model's strong performance in maintaining coherent and dynamic interactions, demonstrating its potential for advanced multimodal conversational agents. 1
Jihyoung Jang, Minwook Bae, Dilek Hakkani-Tür, Hyounghun Kim
ACL (1)5
2025 KoBLEX: Open Legal Question Answering with Multi-hop Reasoning
abstract
Large Language Models (LLM) have achieved remarkable performances in general domains and are now extending into the expert domain of law.Several benchmarks have been proposed to evaluate LLMs' legal capabilities.However, these benchmarks fail to evaluate open-ended and provisiongrounded Question Answering (QA).To address this, we introduce a Korean Benchmark for Legal EXplainable QA (KOBLEX), designed to evaluate provision-grounded, multihop legal reasoning.KOBLEX includes 226 scenario-based QA instances and their supporting provisions, created using a hybrid LLM-human expert pipeline.We also propose a method called Parametric provisionguided Selection Retrieval (PARSER), which uses LLM-generated parametric provisions to guide legally grounded and reliable answers.PARSER facilitates multi-hop reasoning on complex legal questions by generating parametric provisions and employing a three-stage sequential retrieval process.Furthermore, to better evaluate the legal fidelity of the generated answers, we propose Legal Fidelity Evaluation (LF-EVAL).LF-EVAL is an automatic metric that jointly considers the question, answer, and supporting provisions and shows a high correlation with human judgments.Experimental results show that PARSER consistently outperforms strong baselines, achieving the best results across multiple LLMs.Notably, compared to standard retrieval with GPT-4o, PARSER achieves 37.91 higher F-1 and 30.81 higher LF-EVAL.Further analyses reveal that PARSER efficiently delivers consistent performance across reasoning depths, with ablations confirming the effectiveness of PARSER. 1 * Equal Contribution. 1 The code and dataset are available at https://github. com/daehuikim/
Jihyung Lee, Daehui Kim, Seonjeong Hwang, Hyounghun Kim, Gary Geunbae Lee
EMNLP4
2025 PanicToCalm: A Proactive Counseling Agent for Panic Attacks
abstract
Panic attacks are acute episodes of fear and distress, in which timely, appropriate intervention can significantly help individuals regain stability.However, suitable datasets for training such models remain scarce due to ethical and logistical issues.To address this, we introduce PACE, which is a dataset that includes high-distress episodes constructed from firstperson narratives, and structured around the principles of Psychological First Aid (PFA).Using this data, we train PACER, a counseling model designed to provide both empathetic and directive support, which is optimized through supervised learning and simulated preference alignment.To assess its effectiveness, we propose PANICEVAL, a multi-dimensional framework covering general counseling quality and crisis-specific strategies.Experimental results show that PACER outperforms strong baselines in both counselor-side metrics and client affect improvement.Human evaluations further confirm its practical value, with PACER consistently preferred over general, CBT-based, and GPT-4-powered models in panic scenarios 1 .
Yejin Min, Yejin Jeon, SungJun Yang, Hyounghun Kim, Gary Geunbae Lee
EMNLP6
2024 Collective Critics for Creative Story Generation
abstract
Generating a long story of several thousand words with narrative coherence using Large Language Models (LLMs) has been a challenging task.Previous research has addressed this challenge by proposing different frameworks that create a story plan and generate a long story based on that plan.However, these frameworks have been mainly focusing on maintaining narrative coherence in stories, often overlooking creativity in story planning and the expressiveness of the stories generated from those plans, which are desirable properties to captivate readers' interest.In this paper, we propose Collective Critics for Creative Story Generation framework (CRITICS), which is composed of plan refining stage (CRPLAN) and story generation stage (CRTEXT), to integrate a collective revision mechanism that promotes those properties into long-form story generation process.Specifically, in each stage, a group of LLM critics and one leader collaborate to incrementally refine drafts of plan and story throughout multiple rounds.Extensive human evaluation shows that the CRITICS can significantly enhance story creativity and reader engagement, while also maintaining narrative coherence.Furthermore, the design of the framework allows active participation from human writers in any role within the critique process, enabling interactive human-machine collaboration in story writing. 1
Minwook Bae, Hyounghun Kim
EMNLP2
2023 Conversation Chronicles: Towards Diverse Temporal and Relational Dynamics in Multi-Session Conversations
abstract
In the field of natural language processing, open-domain chatbots have emerged as an important research topic.However, a major limitation of existing open-domain chatbot research is its singular focus on short single-session dialogue, neglecting the potential need for understanding contextual information in multiple consecutive sessions that precede an ongoing dialogue.Among the elements that compose the context in multi-session conversation settings, the time intervals between sessions and the relationships between speakers would be particularly important.Despite their importance, current research efforts have not sufficiently addressed these dialogical components.In this paper, we introduce a new 1M multisession dialogue dataset, called CONVERSA-TION CHRONICLES, for implementing a longterm conversation setup in which time intervals and fine-grained speaker relationships are incorporated.Following recent works, we exploit a large language model to produce the data.The extensive human evaluation shows that dialogue episodes in CONVERSATION CHRONI-CLES reflect those properties while maintaining coherent and consistent interactions across all the sessions.We also propose a dialogue model, called REBOT, which consists of chronological summarization and dialogue generation modules using only around 630M parameters.When trained on CONVERSATION CHRONI-CLES, REBOT demonstrates long-term context understanding with a high human engagement score. 1
Jihyoung Jang, Minseong Boo, Hyounghun Kim
EMNLP3
2022 CAISE: Conversational Agent for Image Search and Editing
abstract
Demand for image editing has been increasing as users' desire for expression is also increasing. However, for most users, image editing tools are not easy to use since the tools require certain expertise in photo effects and have complex interfaces. Hence, users might need someone to help edit their images, but having a personal dedicated human assistant for every user is impossible to scale. For that reason, an automated assistant system for image editing is desirable. Additionally, users want more image sources for diverse image editing works, and integrating an image search functionality into the editing tool is a potential remedy for this demand. Thus, we propose a dataset of an automated Conversational Agent for Image Search and Editing (CAISE). To our knowledge, this is the first dataset that provides conversational image search and editing annotations, where the agent holds a grounded conversation with users and helps them to search and edit images according to their requests. To build such a system, we first collect image search and editing conversations between pairs of annotators. The assistant-annotators are equipped with a customized image search and editing tool to address the requests from the user-annotators. The functions that the assistant-annotators conduct with the tool are recorded as executable commands, allowing the trained system to be useful for real-world application execution. We also introduce a generator-extractor baseline model for this task, which can adaptively select the source of the next token (i.e., from the vocabulary or from textual/visual contexts) for the executable command. This serves as a strong starting point while still leaving a large human-machine performance gap for useful future work. Data and code are available: https://github.com/hyounghk/CAISE.
Hyounghun Kim, Doo Soon Kim, Seunghyun Yoon 0002, Franck Dernoncourt, Trung Bui, Mohit Bansal
AAAI1
2022 CoSIm: Commonsense Reasoning for Counterfactual Scene Imagination
abstract
As humans, we can modify our assumptions about a scene by imagining alternative objects or concepts in our minds. For example, we can easily anticipate the implications of the sun being overcast by rain clouds (e.g., the street will get wet) and accordingly prepare for that. In this paper, we introduce a new dataset called Commonsense Reasoning for Counterfactual Scene Imagination (COSIM) which is designed to evaluate the ability of AI systems to reason about scene change imagination. To be specific, in this multimodal task/dataset, models are given an image and an initial questionresponse pair about the image. Next, a counterfactual imagined scene change (in textual form) is applied, and the model has to predict the new response to the initial question based on this scene change. We collect 3.5K high-quality and challenging data instances, with each instance consisting of an image, a commonsense question with a response, a description of a counterfactual change, a new response to the question, and three distractor responses. Our dataset contains various complex scene change types (such as object addition/removal/state change, event description, environment change, etc.) that require models to imagine many different scenarios and reason about the changed scenes. We present a baseline model based on a vision-language Transformer (i.e., LXMERT) and ablation studies. Through human evaluation, we demonstrate a large human-model performance gap, suggesting room for promising future work on this challenging, counterfactual multimodal task.
Hyounghun Kim, Abhaysinh Zala, Mohit Bansal
NAACL-HLT1
2021 FIXMYPOSE: Pose Correctional Captioning and Retrieval
abstract
Interest in physical therapy and individual exercises such as yoga/dance has increased alongside the well-being trend, and people globally enjoy such exercises at home/office via video streaming platforms. However, such exercises are hard to follow without expert guidance. Even if experts can help, it is almost impossible to give personalized feedback to every trainee remotely. Thus, automated pose correction systems are required more than ever, and we introduce a new captioning dataset named FixMyPose to address this need. We collect natural language descriptions of correcting a “current” pose to look like a “target” pose. To support a multilingual setup, we collect descriptions in both English and Hindi. The collected descriptions have interesting linguistic properties such as egocentric relations to the environment objects, analogous references, etc., requiring an understanding of spatial relations and commonsense knowledge about postures. Further, to avoid ML biases, we maintain a balance across characters with diverse demographics, who perform a variety of movements in several interior environments (e.g., homes, offices). From our FixMyPose dataset, we introduce two tasks: the pose-correctional-captioning task and its reverse, the target-pose-retrieval task. During the correctional-captioning task, models must generate the descriptions of how to move from the current to the target pose image, whereas in the retrieval task, models should select the correct target pose given the initial pose and the correctional description. We present strong cross-attention baseline models (uni/multimodal, RL, multilingual) and also show that our baselines are competitive with other models when evaluated on other image-difference datasets. We also propose new task-specific metrics (object-match, body-part-match, direction-match) and conduct human evaluation for more reliable evaluation, and we demonstrate a large human-model performance gap suggesting room for promising future work. Finally, to verify the sim-to-real transfer of our FixMyPose dataset, we collect a set of real images and show promising performance on these images. Data and code are available: https://fixmypose-unc.github.io.
Hyounghun Kim, Abhaysinh Zala, Graham Burri, Mohit Bansal
AAAI1
2021 Continuous Language Generative Flow
abstract
Zineng Tang, Shiyue Zhang, Hyounghun Kim, Mohit Bansal. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Zineng Tang, Shiyue Zhang 0001, Hyounghun Kim, Mohit Bansal
ACL/IJCNLP (1)3
2021 NDH-Full: Learning and Evaluating Navigational Agents on Full-Length Dialogue
abstract
Communication between human and mobile agents is getting increasingly important as such agents are widely deployed in our daily lives.Vision-and-Dialogue Navigation is one of the tasks that evaluate the agent's ability to interact with humans for assistance and navigate based on natural language responses.In this paper, we explore the Navigation from Dialogue History (NDH) task, which is based on the Cooperative Vision-and-Dialogue Navigation (CVDN) dataset, and present a stateof-the-art model which is built upon Vision-Language transformers.However, despite achieving competitive performance, we find that the agent in the NDH task is not evaluated appropriately by the primary metric -Goal Progress.By analyzing the performance mismatch between Goal Progress and other metrics (e.g., normalized Dynamic Time Warping) from our state-of-the-art model, we show that NDH's sub-path based task setup (i.e., navigating partial trajectory based on its correspondent subset of the full dialogue) does not provide the agent with enough supervision signal towards the goal region.Therefore, we propose a new task setup called NDH-FULL which takes the full dialogue and the whole navigation path as one instance.We present a strong baseline model and show initial results on this new task.We further describe several approaches that we try, in order to improve the model performance (based on curriculum learning, pre-training, and data-augmentation), suggesting potential useful training methods on this new NDH-FULL task. 1
Hyounghun Kim, Jialu Li 0001, Mohit Bansal
EMNLP (1)1
2020 Modality-Balanced Models for Visual Dialogue
abstract
The Visual Dialog task requires a model to exploit both image and conversational context information to generate the next response to the dialogue. However, via manual analysis, we find that a large number of conversational questions can be answered by only looking at the image without any access to the context history, while others still need the conversation context to predict the correct answers. We demonstrate that due to this reason, previous joint-modality (history and image) models over-rely on and are more prone to memorizing the dialogue history (e.g., by extracting certain keywords or patterns in the context information), whereas image-only models are more generalizable (because they cannot memorize or extract keywords from history) and perform substantially better at the primary normalized discounted cumulative gain (NDCG) task metric which allows multiple correct answers. Hence, this observation encourages us to explicitly maintain two models, i.e., an image-only model and an image-history joint model, and combine their complementary abilities for a more balanced multimodal model. We present multiple methods for this integration of the two models, via ensemble and consensus dropout fusion with shared parameters. Empirically, our models achieve strong results on the Visual Dialog challenge 2019 (rank 3 on NDCG and high balance across metrics), and substantially outperform the winner of the Visual Dialog challenge 2018 on most metrics.
Hyounghun Kim, Hao Tan 0002, Mohit Bansal
AAAI1
2020 Dense-Caption Matching and Frame-Selection Gating for Temporal Localization in VideoQA
abstract
Videos convey rich information.Dynamic spatio-temporal relationships between people/objects, and diverse multimodal events are present in a video clip.Hence, it is important to develop automated models that can accurately extract such information from videos.Answering questions on videos is one of the tasks which can evaluate such AI abilities.In this paper, we propose a video question answering model which effectively integrates multi-modal input sources and finds the temporally relevant information to answer questions.Specifically, we first employ dense image captions to help identify objects and their detailed salient regions and actions, and hence give the model useful extra information (in explicit textual format to allow easier matching) for answering questions.Moreover, our model is also comprised of duallevel attention (word/object and frame level), multi-head self/cross-integration for different sources (video and dense captions), and gates which pass more relevant information to the classifier.Finally, we also cast the frame selection problem as a multi-label classification task and introduce two loss functions, In-and-Out Frame Score Margin (IOFSM) and Balanced Binary Cross-Entropy (BBCE), to better supervise the model with human importance annotations.We evaluate our model on the challenging TVQA dataset, where each of our model components provides significant gains, and our overall model outperforms the stateof-the-art by a large margin (74.09% versus 70.52%).We also present several word, object, and frame level visualization studies. 1Local Gate Frame Score Margin Inside Frames Outside FramesFrame-Level Att.
Hyounghun Kim, Zineng Tang, Mohit Bansal
ACL1
2019 Improving Visual Question Answering by Referring to Generated Paragraph Captions
abstract
Paragraph-style image captions describe diverse aspects of an image as opposed to the more common single-sentence captions that only provide an abstract description of the image.These paragraph captions can hence contain substantial information of the image for tasks such as visual question answering.Moreover, this textual information is complementary with visual information present in the image because it can discuss both more abstract concepts and more explicit, intermediate symbolic information about objects, events, and scenes that can directly be matched with the textual question and copied into the textual answer (i.e., via easier modality match).Hence, we propose a combined Visual and Textual Question Answering (VTQA) model which takes as input a paragraph caption as well as the corresponding image, and answers the given question based on both inputs.In our model, the inputs are fused to extract related information by cross-attention (early fusion), then fused again in the form of consensus (late fusion), and finally expected answers are given an extra score to enhance the chance of selection (later fusion).Empirical results show that paragraph captions, even when automatically generated (via an RL-based encoderdecoder model), help correctly answer more visual questions.Overall, our joint model, when trained on the Visual Genome dataset, significantly improves the VQA performance over a strong baseline model.
Hyounghun Kim, Mohit Bansal
ACL (1)1
2018 Towards Fully Mobile 3D Face, Body, and Environment Capture Using Only Head-worn Cameras
abstract
We propose a new approach for 3D reconstruction of dynamic indoor and outdoor scenes in everyday environments, leveraging only cameras worn by a user. This approach allows 3D reconstruction of experiences at any location and virtual tours from anywhere. The key innovation of the proposed ego-centric reconstruction system is to capture the wearer's body pose and facial expression from near-body views, e.g. cameras on the user's glasses, and to capture the surrounding environment using outward-facing views. The main challenge of the ego-centric reconstruction, however, is the poor coverage of the near-body views - that is, the user's body and face are observed from vantage points that are convenient for wear but inconvenient for capture. To overcome these challenges, we propose a parametric-model-based approach to user motion estimation. This approach utilizes convolutional neural networks (CNNs) for near-view body pose estimation, and we introduce a CNN-based approach for facial expression estimation that combines audio and video. For each time-point during capture, the intermediate model-based reconstructions from these systems are used to re-target a high-fidelity pre-scanned model of the user. We demonstrate that the proposed self-sufficient, head-worn capture system is capable of reconstructing the wearer's movements and their surrounding environment in both indoor and outdoor situations without any additional views. As a proof of concept, we show how the resulting 3D-plus-time reconstruction can be immersively experienced within a virtual reality system (e.g., the HTC Vive). We expect that the size of the proposed egocentric capture-and-reconstruction system will eventually be reduced to fit within future AR glasses, and will be widely useful for immersive 3D telepresence, virtual tours, and general use-anywhere 3D content creation.
Young-Woon Cha, True Price, Xinran Lu, Nicholas Rewkowski, Rohan Chabra, Zihe Qin, Hyounghun Kim, Zhaoqi Su, Yebin Liu, Adrian Ilie, Andrei State, Zhenlin Xu, Jan-Michael Frahm, Henry Fuchs
IEEE Trans. Vis. Comput. Graph.8