EDBT 2026 Demo / reviewers in the wild / expert
Jiwan Chung
dblp:277/2798
· DBLP profile ↗
19ranked-venue papers
6as first author
18since 2021 · last 2026
0009-0005-4974-0637ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 6 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Explain with Visual Keypoints Like a Real Mentor! A Benchmark for Multimodal Solution ExplanationabstractWith the rapid advancement of mathematical reasoning capabilities in Large Language Models (LLMs), AI systems are increasingly being adopted in educational settings to support students’ comprehension of problem-solving processes. However, a critical component remains underexplored in current LLM-generated explanations: multimodal explanation. In real-world instructional contexts, human tutors routinely employ visual aids, such as diagrams, markings, and highlights, to enhance conceptual clarity. To bridge this gap, we introduce the multimodal solution explanation task, designed to evaluate whether models can identify visual keypoints, such as auxiliary lines, points, angles, and generate explanations that incorporate these key elements essential for understanding. To evaluate model performance on this task, we propose ME2, a multimodal benchmark consisting of 1,000 math problems annotated with visual keypoints and corresponding explanatory text that references those elements. Our empirical results show that current models struggle to identify visual keypoints. In the task of generating keypoint-based explanations, open-source models also face notable difficulties. This highlights a significant gap in current LLMs’ ability to perform mathematical visual grounding, engage in visually grounded reasoning, and provide explanations in educational contexts. We expect that the multimodal solution explanation task and the ME2 dataset will catalyze further research on LLMs in education and promote their use as effective, explanation-oriented AI tutors. Jaewoo Park 0003, Jungyang Park, Dongju Jang, Jiwan Chung, Byungwoo Yoo, Jaewoo Shin 0003, Seonjoon Park, Taehyeong Kim 0001, Youngjae Yu |
AAAI | 4 |
| 2026 | GuideDog: A Real-World Egocentric Multimodal Dataset for Blind and Low-Vision Accessibility-Aware GuidanceabstractJunhyeok Kim, Jaewoo Park, Junhee Park, Sangeyl Lee, Jiwan Chung, Jisung Kim, Ji Hoon Joung, Youngjae Yu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Junhyeok Kim 0002, Jaewoo Park 0003, Junhee Park, Sangeyl Lee, Jiwan Chung, Jisung Kim, Ji Hoon Joung, Youngjae Yu |
ACL (1) | 5 |
| 2025 | MASS: Overcoming Language Bias in Image-Text MatchingabstractPretrained visual-language models have made significant advancements in multimodal tasks, including image-text retrieval. However, a major challenge in image-text matching lies in language bias, where models predominantly rely on language priors and neglect to adequately consider the visual content. We thus present Multimodal ASsociation Score (MASS), a framework that reduces the reliance on language priors for better visual accuracy in image-text matching problems. It can be seamlessly incorporated into existing visual-language models without necessitating additional training. Our experiments have shown that \modelname effectively lessens language bias without losing an understanding of linguistic compositionality. Overall, MASS offers a promising solution for enhancing image-text matching performance in visual-language models. Jiwan Chung, Seungwon Lim, Sangkyu Lee, Youngjae Yu |
AAAI | 1 |
| 2025 | Are Any-to-Any Models More Consistent Across Modality Transfers Than Specialists?abstractJiwan Chung, Janghan Yoon, Junhyeong Park, Sangeyl Lee, Joowon Yang, Sooyeon Park, Youngjae Yu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jiwan Chung, Janghan Yoon, Sangeyl Lee, Joowon Yang, Sooyeon Park, Youngjae Yu |
ACL (1) | 1 |
| 2025 | Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded DialoguesabstractYoungmin Kim, Jiwan Chung, Jisoo Kim, Sunghyun Lee, Sangkyu Lee, Junhyeok Kim, Cheoljong Yang, Youngjae Yu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jiwan Chung, Jisoo Kim 0006, Sunghyun Lee 0001, Sangkyu Lee, Junhyeok Kim 0002, Cheoljong Yang, Youngjae Yu |
ACL (1) | 2 |
| 2025 | VisEscape: A Benchmark for Evaluating Exploration-driven Decision-making in Virtual Escape RoomsabstractEscape rooms present a unique cognitive challenge that demands exploration-driven planning: with the sole instruction to escape the room, players must actively search their environment, collecting information, and finding solutions through repeated trial and error. Motivated by this, we introduce VisEscape, a benchmark of 20 virtual escape rooms specifically designed to evaluate AI models under these challenging conditions, where success depends not only on solving isolated puzzles but also on iteratively constructing and refining spatial-temporal knowledge of a dynamically changing environment. On VisEscape, we observe that even state-of-the-art multi-modal models generally fail to escape the rooms, showing considerable variation in their progress and problem-solving approaches. We find that integrating memory management and reasoning contributes to efficient exploration and enables successive hypothesis formulation and testing, thereby leading to significant improvements in dynamic and exploration-driven environments. Seungwon Lim, Sungwoong Kim, Jihwan Yu, Sungjae Lee 0003, Jiwan Chung, Youngjae Yu |
EMNLP | 5 |
| 2025 | VAGUE: Visual Contexts Clarify Ambiguous Expressions
Heejeong Nam, Jinwoo Ahn, Keummin Ka, Jiwan Chung, Youngjae Yu |
ICCV | 4 |
| 2025 | CANVAS: Commonsense-Aware Navigation System for Intuitive Human-Robot InteractionabstractReal-life robot navigation involves more than just reaching a destination; it requires optimizing movements while addressing scenario-specific goals. An intuitive way for humans to express these goals is through abstract cues like verbal commands or rough sketches. Such human guidance may lack details or be noisy. Nonetheless, we expect robots to navigate as intended. For robots to interpret and execute these abstract instructions in line with human expectations, they must share a common understanding of basic navigation concepts with humans. To this end, we introduce CANVAS, a novel framework that combines visual and linguistic instructions for commonsense-aware navigation. Its success is driven by imitation learning, enabling the robot to learn from human navigation behavior. We present COMMAND, a comprehensive dataset with human-annotated navigation results, spanning over 48 hours and 219 km, designed to train commonsense-aware navigation systems in simulated environments. Our experiments show that CANVAS outperforms the strong rule-based system ROS NavStack across all environments, demonstrating superior performance with noisy instructions. Notably, in the orchard environment, where ROS NavStack records a 0% total success rate, CANVAS achieves a total success rate of 67%. CANVAS also closely aligns with human demonstrations and commonsense constraints, even in unseen environments. Furthermore, real-world deployment of CANVAS showcases impressive Sim2Real transfer with a total success rate of 69%, highlighting the potential of learning from human demonstrations in simulated environments for real-world applications. Suhwan Choi, Yongjun Cho, Jaeyoon Jung, Myunchul Joe, Yubeen Park, Sungwoong Kim, Sungjae Lee 0003, Hwiseong Park, Jiwan Chung, Youngjae Yu |
ICRA | 11 |
| 2024 | Language Models as Compilers: Simulating Pseudocode Execution Improves Algorithmic Reasoning in Language ModelsabstractHyungjoo Chae, Yeonghyeon Kim, Seungone Kim, Kai Tzu-iunn Ong, Beong-woo Kwak, Moohyeon Kim, Sunghwan Kim, Taeyoon Kwon, Jiwan Chung, Youngjae Yu, Jinyoung Yeo. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Hyungjoo Chae, Yeonghyeon Kim, Seungone Kim, Kai Tzu-iunn Ong, Beong-woo Kwak, Moohyeon Kim, Sunghwan Kim 0005, Taeyoon Kwon, Jiwan Chung, Youngjae Yu, Jinyoung Yeo |
EMNLP | 9 |
| 2024 | Can visual language models resolve textual ambiguity with visual cues? Let visual puns tell you!abstractHumans possess multimodal literacy, allowing them to actively integrate information from various modalities to form reasoning.Faced with challenges like lexical ambiguity in text, we supplement this with other modalities, such as thumbnail images or textbook illustrations.Is it possible for machines to achieve a similar multimodal understanding capability?In response, we present Understanding Pun with Image Explanations ( UNPIE) 1 , a novel benchmark designed to assess the impact of multimodal inputs in resolving lexical ambiguities.Puns serve as the ideal subject for this evaluation due to their intrinsic ambiguity.Our dataset includes 1,000 puns, each accompanied by an image that explains both meanings.We pose three multimodal challenges with the annotations to assess different aspects of multimodal literacy; Pun Grounding, Disambiguation, and Reconstruction.The results 2 indicate that various Socratic Models and Visual-Language Models improve over the text-only models when given visual context, particularly as the complexity of the tasks increases. Jiwan Chung, Seungwon Lim, Jaehyun Jeon 0002, Seungbeen Lee, Youngjae Yu |
EMNLP | 1 |
| 2024 | Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument UnderstandingabstractVisual arguments, often used in advertising or social causes, rely on images to persuade viewers to do or believe something. Understanding these arguments requires selective vision: only specific visual stimuli within an image are relevant to the argument, and relevance can only be understood within the context of a broader argumentative structure. While visual arguments are readily appreciated by human audiences, we ask: are today’s AI capable of similar understanding?We present VisArgs, a dataset of 1,611 images annotated with 5,112 visual premises (with regions), 5,574 commonsense premises, and reasoning trees connecting them into structured arguments. We propose three tasks for evaluating visual argument understanding: premise localization, premise identification, and conclusion deduction.Experiments show that 1) machines struggle to capture visual cues: GPT-4-O achieved 78.5% accuracy, while humans reached 98.0%. Models also performed 19.5% worse when distinguishing between irrelevant objects within the image compared to external objects. 2) Providing relevant visual premises improved model performance significantly. Jiwan Chung, Sungjae Lee 0003, Seungju Han 0002, Ashkan Yousefpour, Jack Hessel, Youngjae Yu |
EMNLP | 1 |
| 2024 | Towards Visual Text Design Transfer Across LanguagesabstractVisual text design plays a critical role in conveying themes, emotions, and atmospheres in multimodal formats such as film posters and album covers. Translating these visual and textual elements across languages extends the concept of translation beyond mere text, requiring the adaptation of aesthetic and stylistic features. To address this, we introduce a novel task of Multimodal Style Translation (MuST-Bench), a benchmark designed to evaluate the ability of visual text generation models to perform translation across different writing systems while preserving design intent.Our initial experiments on MuST-Bench reveal that existing visual text generation models struggle with the proposed task due to the inadequacy of textual descriptions in conveying visual design.In response, we introduce SIGIL, a framework for multimodal style translation that eliminates the need for style descriptions.SIGIL enhances image generation models through three innovations: glyph latent for multilingual settings, pre-trained VAEs for stable style guidance, and an OCR model with reinforcement learning feedback for optimizing readable character generation. SIGIL outperforms existing baselines by achieving superior style consistency and legibility while maintaining visual fidelity, setting itself apart from traditional description-based approaches. We release MuST-Bench publicly for broader use and exploration https://huggingface.co/datasets/yejinc/MuST-Bench. Yejin Choi 0004, Jiwan Chung, Sumin Shim, Giyeong Oh, Youngjae Yu |
NeurIPS | 2 |
| 2023 | Long Story Short: a Summarize-then-Search Method for Prompt-Based Long Video Question Answering
Jiwan Chung, Youngjae Yu |
BMVC | 1 |
| 2023 | Fusing Pre-Trained Language Models with Multimodal Prompts through Reinforcement LearningabstractLanguage models are capable of commonsense reasoning: while domain-specific models can learn from explicit knowledge (e.g. commonsense graphs [6] ethical norms [25]), and larger models like GPT-3 [7] mani-fest broad commonsense reasoning capacity. Can their knowledge be extended to multimodal inputs such as images and audio without paired domain data? In this work, we propose‡ESPER (Extending Sensory PErception with Reinforcement learning) which enables text-only pretrained models to address multimodal tasks such as visual commonsense reasoning. Our key novelty is to use rein-forcement learning to align multimodal inputs to language model generations without direct supervision: for example, our reward optimization relies only on cosine similarity derived from CLIP [52] and requires no additional paired (image, text) data. Experiments demonstrate that ESPER outperforms baselines and prior work on a variety of multimodal text generation tasks ranging from captioning to commonsense reasoning; these include a new benchmark we collect and release, the ESP dataset, which tasks models with generating the text of several different domains for each image. Our code and data are publicly released at https://github.com/JiwanChung/esper. Youngjae Yu, Jiwan Chung, Heeseung Yun, Jack Hessel, Ximing Lu, Rowan Zellers, Prithviraj Ammanabrolu, Ronan Le Bras 0001, Gunhee Kim, Yejin Choi 0001 |
CVPR | 2 |
| 2023 | VLIS: Unimodal Language Models Guide Multimodal Language GenerationabstractMultimodal language generation, which leverages the synergy of language and vision, is a rapidly expanding field.However, existing vision-language models face challenges in tasks that require complex linguistic understanding.To address this issue, we introduce Visual-Language models as Importance Sampling weights ( VLIS), a novel framework that combines the visual conditioning capability of vision-language models with the language understanding of unimodal text-only language models without further training.It extracts pointwise mutual information of each image and text from a visual-language model and uses the value as an importance sampling weight to adjust the token likelihood from a text-only model.VLIS improves visionlanguage models on diverse tasks, including commonsense understanding (WHOOPS, OK-VQA, and ScienceQA) and complex text generation (Concadia, Image Paragraph Captioning, and ROCStories).Our results suggest that VLIS represents a promising new direction for multimodal language generation. Named EntitiesWho is this?Does he care for his family? Jiwan Chung, Youngjae Yu |
EMNLP | 1 |
| 2023 | Reading Books is Great, But Not if You Are Driving! Visually Grounded Reasoning about Defeasible Commonsense NormsabstractCommonsense norms are defeasible by context: reading books is usually great, but not when driving a car.While contexts can be explicitly described in language, in embodied scenarios, contexts are often provided visually.This type of visually grounded reasoning about defeasible commonsense norms is generally easy for humans, but (as we show) poses a challenge for machines, as it necessitates both visual understanding and reasoning about commonsense norms.We construct a new multimodal benchmark for studying visual-grounded commonsense norms: NORMLENS.NORMLENS consists of 10K human judgments accompanied by freeform explanations covering 2K multimodal situations, and serves as a probe to address two questions: (1) to what extent can models align with average human judgment?and (2) how well can models explain their predicted judgments?We find that state-of-theart model judgments and explanations are not well-aligned with human annotation.Additionally, we present a new approach to better align models with humans by distilling social commonsense knowledge from large language models.The data and code are released at https://seungjuhan.me/normlens. Seungju Han 0002, Junhyeok Kim 0002, Jack Hessel, Jiwan Chung, Yejin Son, Yejin Choi 0001, Youngjae Yu |
EMNLP | 5 |
| 2021 | Transitional Adaptation of Pretrained Models for Visual StorytellingabstractPrevious models for vision-to-language generation tasks usually pretrain a visual encoder and a language generator in the respective domains and jointly finetune them with the target task. However, this direct transfer practice may suffer from the discord between visual specificity and language fluency since they are often separately trained from large corpora of visual and text data with no common ground. In this work, we claim that a transitional adaptation task is required between pretraining and finetuning to harmonize the visual encoder and the language model for challenging downstream target tasks like visual storytelling. We propose a novel approach named Transitional Adaptation of Pre-trained Model (TAPM) that adapts the multi-modal modules to each other with a simpler alignment task between visual inputs only with no need for text labels. Through extensive experiments, we show that the adaptation step significantly improves the performance of multiple language models for sequential video and image captioning tasks. We achieve new state-of-the-art performance on both language metrics and human evaluation in the multi-sentence description task of LSMDC 2019 [50] and the image storytelling task of VIST [18]. Our experiments reveal that this improvement in caption quality does not depend on the specific choice of language models. Youngjae Yu, Jiwan Chung, Heeseung Yun, Jongseok Kim 0002, Gunhee Kim |
CVPR | 2 |
| 2021 | ACAV100M: Automatic Curation of Large-Scale Datasets for Audio-Visual Video Representation LearningabstractThe natural association between visual observations and their corresponding sound provides powerful self-supervisory signals for learning video representations, which makes the ever-growing amount of online videos an attractive source of training data. However, large portions of online videos contain irrelevant audio-visual signals because of edited/overdubbed audio, and models trained on such uncurated videos have shown to learn suboptimal representations. Therefore, existing self-supervised approaches rely on datasets with predetermined taxonomies of semantic concepts, where there is a high chance of audio-visual correspondence. Unfortunately, constructing such datasets require labor intensive manual annotation and/or verification, which severely limits the utility of online videos for large-scale learning. In this work, we present an automatic dataset curation approach based on subset optimization where the objective is to maximize the mutual information between audio and visual channels in videos. We demonstrate that our approach finds videos with high audio-visual correspondence and show that self-supervised models trained on our data achieve competitive performances compared to models trained on existing manually curated datasets. The most significant benefit of our approach is scalability: We release ACAV100M that contains 100 million videos with high audio-visual correspondence, ideal for self-supervised video representation learning. Sangho Lee 0008, Jiwan Chung, Youngjae Yu, Gunhee Kim, Thomas M. Breuel, Gal Chechik, Yale Song |
ICCV | 2 |
| 2020 | Character Grounding and Re-identification in Story of Videos and Text Descriptions
Youngjae Yu, Jongseok Kim 0002, Heeseung Yun, Jiwan Chung, Gunhee Kim |
ECCV (5) | 4 |