VLDB 2026 Research / reviewers in the wild / expert
Youngjae Yu
dblp:188/6210
· DBLP profile ↗
66ranked-venue papers
8as first author
57since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 63 · 8 first-author · 55 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 7 first-author · 18 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Do Language Models Associate Sound with Meaning? A Multimodal Study of Sound SymbolismabstractSound symbolism is a linguistic concept that refers to non-arbitrary associations between phonetic forms and their meanings. We suggest that this can be a compelling probe into how Multimodal Large Language Models (MLLMs) interpret auditory information in human languages. We investigate MLLMs' performance on phonetic iconicity across textual (orthographic and IPA) and auditory forms of inputs with up to 25 semantic dimensions (e.g., sharp vs. round), observing models' layer-wise information processing by measuring phoneme-level attention fraction scores. To this end, we present LEX-ICON, an extensive mimetic word dataset consisting of 8,052 words from four natural languages (English, French, Japanese, and Korean) and 2,930 systematically constructed pseudo-words, annotated with semantic features applied across both text and audio modalities. Our key findings demonstrate (1) MLLMs' phonetic intuitions that align with existing linguistic research across multiple semantic dimensions and (2) phonosemantic attention patterns that highlight models' focus on iconic phonemes. These results bridge domains of artificial intelligence and cognitive linguistics, providing the first large-scale, quantitative analyses of phonetic iconicity in terms of MLLMs' interpretability. Jinhong Jeong, Sunghyun Lee 0001, Seonah Han, Youngjae Yu |
AAAI | 5 |
| 2026 | Explain with Visual Keypoints Like a Real Mentor! A Benchmark for Multimodal Solution ExplanationabstractWith the rapid advancement of mathematical reasoning capabilities in Large Language Models (LLMs), AI systems are increasingly being adopted in educational settings to support students’ comprehension of problem-solving processes. However, a critical component remains underexplored in current LLM-generated explanations: multimodal explanation. In real-world instructional contexts, human tutors routinely employ visual aids, such as diagrams, markings, and highlights, to enhance conceptual clarity. To bridge this gap, we introduce the multimodal solution explanation task, designed to evaluate whether models can identify visual keypoints, such as auxiliary lines, points, angles, and generate explanations that incorporate these key elements essential for understanding. To evaluate model performance on this task, we propose ME2, a multimodal benchmark consisting of 1,000 math problems annotated with visual keypoints and corresponding explanatory text that references those elements. Our empirical results show that current models struggle to identify visual keypoints. In the task of generating keypoint-based explanations, open-source models also face notable difficulties. This highlights a significant gap in current LLMs’ ability to perform mathematical visual grounding, engage in visually grounded reasoning, and provide explanations in educational contexts. We expect that the multimodal solution explanation task and the ME2 dataset will catalyze further research on LLMs in education and promote their use as effective, explanation-oriented AI tutors. Jaewoo Park 0003, Jungyang Park, Dongju Jang, Jiwan Chung, Byungwoo Yoo, Jaewoo Shin 0003, Seonjoon Park, Taehyeong Kim 0001, Youngjae Yu |
AAAI | 9 |
| 2026 | Do MLLMs Capture How Interfaces Guide User Behavior? A Benchmark for Multimodal UI/UX Design UnderstandingabstractJaehyun Jeon, Min Soo Kim, Janghan Yoon, Sumin Shim, Yejin Choi, Hanbin Kim, Dae Hyun Kim, Youngjae Yu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jaehyun Jeon 0002, Min Soo Kim 0007, Janghan Yoon, Sumin Shim, Yejin Choi 0004, Hanbin Kim, Youngjae Yu |
ACL (1) | 8 |
| 2026 | Right at My Level: A Unified Multilingual Framework for Proficiency-Aware Text SimplificationabstractText simplification supports second language (L2) learning by providing comprehensible input, consistent with the Input Hypothesis.However, constructing personalized parallel corpora is costly, while existing large language model (LLM)-based readability control methods rely on pre-labeled sentence corpora and primarily target English.We propose Re-RIGHT, a unified reinforcement learning framework for adaptive multilingual text simplification without parallel corpus supervision.We first show that prompting-based lexical simplification at target proficiency levels (CEFR, JLPT, TOPIK, and HSK) performs poorly at easier levels and for non-English languages, even with state-of-the-art LLMs such as GPT-5.2 and Gemini 2.5.To address this, we collect 43K vocabulary-level data across four languages (English, Japanese, Korean, and Chinese) and train a compact 4B policy model using Re-RIGHT, which integrates three reward modules: vocabulary coverage, semantic preservation, and coherence.Compared to the stronger LLM baselines, Re-RIGHT achieves higher lexical coverage at target proficiency levels while maintaining original meaning and fluency.(1) Preparation Complex Seed Text "Pigeon photography is an aerial photography technique invented in 1907 by the German apothecary Julius Neubronner, ... Jinhong Jeong, Junghun Park, Youngjae Yu |
ACL (1) | 3 |
| 2026 | GuideDog: A Real-World Egocentric Multimodal Dataset for Blind and Low-Vision Accessibility-Aware GuidanceabstractJunhyeok Kim, Jaewoo Park, Junhee Park, Sangeyl Lee, Jiwan Chung, Jisung Kim, Ji Hoon Joung, Youngjae Yu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Junhyeok Kim 0002, Jaewoo Park 0003, Junhee Park, Sangeyl Lee, Jiwan Chung, Jisung Kim, Ji Hoon Joung, Youngjae Yu |
ACL (1) | 8 |
| 2026 | Investigating Counterfactual Unfairness in LLMs towards Identities through HumorabstractShubin Kim, Yejin Son, Junyeong Park, Keummin Ka, Seungbeen Lee, Jaeyoung Lee, Hyeju Jang, Alice Oh, Youngjae Yu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shubin Kim, Yejin Son, Junyeong Park, Keummin Ka, Seungbeen Lee, Hyeju Jang, Alice Oh, Youngjae Yu |
ACL (1) | 9 |
| 2025 | ISR-DPO: Aligning Large Multimodal Models for Videos by Iterative Self-Retrospective DPOabstractIterative self-improvement, a concept extending beyond personal growth, has found powerful applications in machine learning, particularly in transforming weak models into strong ones. While recent advances in natural language processing have shown its efficacy through iterative preference optimization, applying this approach to Video Large Multimodal Models (VLMMs) remains challenging due to modality misalignment. VLMMs struggle with this misalignment during iterative preference modeling, as the self-judge model often prioritizes linguistic knowledge over visual information. Additionally, iterative preference optimization can lead to visually hallucinated verbose responses due to length bias within the self-rewarding cycle. To address these issues, we propose Iterative Self-Retrospective Direct Preference Optimization (ISR-DPO), a method that uses self-retrospection to enhance preference modeling. This approach enhances the self-judge’s focus on informative video regions, resulting in more visually grounded preferences. In extensive empirical evaluations across diverse video question answering benchmarks, the ISR-DPO significantly outperforms the state of the art. We are committed to open-sourcing our code, models, and datasets to encourage further investigation. Daechul Ahn, Yura Choi, Youngjae Yu, Dongyeop Kang |
AAAI | 4 |
| 2025 | MASS: Overcoming Language Bias in Image-Text MatchingabstractPretrained visual-language models have made significant advancements in multimodal tasks, including image-text retrieval. However, a major challenge in image-text matching lies in language bias, where models predominantly rely on language priors and neglect to adequately consider the visual content. We thus present Multimodal ASsociation Score (MASS), a framework that reduces the reliance on language priors for better visual accuracy in image-text matching problems. It can be seamlessly incorporated into existing visual-language models without necessitating additional training. Our experiments have shown that \modelname effectively lessens language bias without losing an understanding of linguistic compositionality. Overall, MASS offers a promising solution for enhancing image-text matching performance in visual-language models. Jiwan Chung, Seungwon Lim, Sangkyu Lee, Youngjae Yu |
AAAI | 4 |
| 2025 | DEEPTalk: Dynamic Emotion Embedding for Probabilistic Speech-Driven 3D Face AnimationabstractSpeech-driven 3D facial animation has garnered lots of attention thanks to its broad range of applications. Despite recent advancements in achieving realistic lip motion, current methods fail to capture the nuanced emotional undertones conveyed through speech and produce monotonous facial motion. These limitations result in blunt and repetitive facial animations, reducing user engagement and hindering their applicability. To address these challenges, we introduce DEEPTalk, a novel approach that generates diverse and emotionally rich 3D facial expressions directly from speech inputs. To achieve this, we first train DEE (Dynamic Emotion Embedding), which employs probabilistic contrastive learning to forge a joint emotion embedding space for both speech and facial motion. This probabilistic framework captures the uncertainty in interpreting emotions from speech and facial motion, enabling the derivation of emotion vectors from its multifaceted space. Moreover, to generate dynamic facial motion, we design TH-VQVAE (Temporally Hierarchical VQ-VAE) as an expressive and robust motion prior overcoming limitations of VAEs and VQ-VAEs. Utilizing these strong priors, we develop DEEPTalk, a talking head generator that non-autoregressively predicts codebook indices to create dynamic facial motion, incorporating a novel emotion consistency loss. Extensive experiments on various datasets demonstrate the effectiveness of our approach in creating diverse, emotionally expressive talking faces that maintain accurate lip-sync. Our project page is available at https://whwjdqls.github.io/deeptalk.github.io/. Jisoo Kim 0006, Jungbin Cho, Joonho Park, Soonmin Hwang, Da Eun Kim, Youngjae Yu |
AAAI | 7 |
| 2025 | Are Any-to-Any Models More Consistent Across Modality Transfers Than Specialists?abstractJiwan Chung, Janghan Yoon, Junhyeong Park, Sangeyl Lee, Joowon Yang, Sooyeon Park, Youngjae Yu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jiwan Chung, Janghan Yoon, Sangeyl Lee, Joowon Yang, Sooyeon Park, Youngjae Yu |
ACL (1) | 7 |
| 2025 | Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded DialoguesabstractYoungmin Kim, Jiwan Chung, Jisoo Kim, Sunghyun Lee, Sangkyu Lee, Junhyeok Kim, Cheoljong Yang, Youngjae Yu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jiwan Chung, Jisoo Kim 0006, Sunghyun Lee 0001, Sangkyu Lee, Junhyeok Kim 0002, Cheoljong Yang, Youngjae Yu |
ACL (1) | 8 |
| 2025 | Persona Dynamics: Unveiling the Impact of Persona Traits on Agents in Text-Based GamesabstractArtificial agents are increasingly central to complex interactions and decision-making tasks, yet aligning their behaviors with desired human values remains an open challenge.In this work, we investigate how human-like personality traits influence agent behavior and performance within text-based interactive environments.We introduce PANDA: Personality-Adapted Neural Decision Agents, a novel method for projecting human personality traits onto agents to guide their behavior.To induce personality in a text-based game agent, (i) we train a personality classifier to identify what personality type the agent's actions exhibit, and (ii) we integrate the personality profiles directly into the agent's policy-learning pipeline.By deploying agents embodying 16 distinct personality types across 25 text-based games and analyzing their trajectories, we demonstrate that an agent's action decisions can be guided toward specific personality profiles.Moreover, certain personality types, such as those characterized by higher levels of Openness, display marked advantages in performance.These findings underscore the promise of personality-adapted agents for fostering more aligned, effective, and human-centric decision-making in interactive environments.1 Seungwon Lim, Seungbeen Lee, Dongjun Min, Youngjae Yu |
ACL (1) | 4 |
| 2025 | Representation Bending for Large Language Model SafetyabstractLarge Language Models (LLMs) have emerged as powerful tools, but their inherent safety risks - ranging from harmful content generation to broader societal harms - pose significant challenges. These risks can be amplified by the recent adversarial attacks, fine-tuning vulnerabilities, and the increasing deployment of LLMs in high-stakes environments. Existing safety-enhancing techniques, such as fine-tuning with human feedback or adversarial training, are still vulnerable as they address specific threats and often fail to generalize across unseen attacks, or require manual system-level defenses. This paper introduces REPBEND, a novel approach that fundamentally disrupts the representations underlying harmful behaviors in LLMs, offering a scalable solution to enhance (potentially inherent) safety. REPBEND brings the idea of activation steering - simple vector arithmetic for steering model's behavior during inference - to loss-based fine-tuning. Through extensive evaluation, REPBEND achieves state-of-the-art performance, outperforming prior methods such as Circuit Breaker, RMU, and NPO, with up to 95% reduction in attack success rates across diverse jailbreak benchmarks, all with negligible reduction in model usability and general capabilities. Ashkan Yousefpour, Taeheon Kim, Ryan Sungmo Kwon, Seungbeen Lee, Wonje Jeung, Seungju Han 0002, Alvin Wan, Harrison Ngan, Youngjae Yu |
ACL (1) | 9 |
| 2025 | MAVL: A Multilingual Audio-Video Lyrics Dataset for Animated Song TranslationabstractLyrics translation requires both accurate semantic transfer and preservation of musical rhythm, syllabic structure, and poetic style.In animated musicals, the challenge intensifies due to alignment with visual and auditory cues.We introduce Multilingual Audio-Video Lyrics Benchmark for Animated Song Translation (MAVL), the first multilingual, multimodal benchmark for singable lyrics translation.By integrating text, audio, and video, MAVL enables richer and more expressive translations than textonly approaches.Building on this, we propose Syllable-Constrained Audio-Video LLM with Chain-of-Thought (SylAVL-CoT), which leverages audio-video cues and enforces syllabic constraints to produce natural-sounding lyrics.Experimental results demonstrate that SylAVL-CoT significantly outperforms textbased models in singability and contextual accuracy, emphasizing the value of multimodal, multilingual approaches for lyrics translation. Woohyun Cho, Sunghyun Lee 0001, Youngjae Yu |
EMNLP | 4 |
| 2025 | Zero-shot Multimodal Document Retrieval via Cross-modal Question GenerationabstractRapid advances in Multimodal Large Language Models (MLLMs) have extended information retrieval beyond text, enabling access to complex real-world documents that combine both textual and visual content.However, most documents are private, either owned by individuals or confined within corporate silos, and current retrievers struggle when faced with unseen domains or languages.To address this gap, we introduce PREMIR, a simple yet effective framework that leverages the broad knowledge of an MLLM to generate cross-modal pre-questions (preQs) before retrieval.Unlike earlier multimodal retrievers that embed entire documents as a single vector, PREMIR leverages preQs, decomposed from documents into finer token-level representations across modalities, enabling richer contextual understanding.Experiments show that PREMIR achieves stateof-the-art performance on out-of-distribution benchmarks, including closed-domain and multilingual settings, outperforming strong baselines across all metrics.We confirm the contribution of each component through in-depth ablation studies, and qualitative analyses of the generated preQs further highlight the framework's robustness in real-world settings 1 . Yejin Choi 0004, Jaewoo Park 0003, Janghan Yoon, Saejin Kim, Jaehyun Jeon 0002, Youngjae Yu |
EMNLP | 6 |
| 2025 | VisEscape: A Benchmark for Evaluating Exploration-driven Decision-making in Virtual Escape RoomsabstractEscape rooms present a unique cognitive challenge that demands exploration-driven planning: with the sole instruction to escape the room, players must actively search their environment, collecting information, and finding solutions through repeated trial and error. Motivated by this, we introduce VisEscape, a benchmark of 20 virtual escape rooms specifically designed to evaluate AI models under these challenging conditions, where success depends not only on solving isolated puzzles but also on iteratively constructing and refining spatial-temporal knowledge of a dynamically changing environment. On VisEscape, we observe that even state-of-the-art multi-modal models generally fail to escape the rooms, showing considerable variation in their progress and problem-solving approaches. We find that integrating memory management and reasoning contributes to efficient exploration and enables successive hypothesis formulation and testing, thereby leading to significant improvements in dynamic and exploration-driven environments. Seungwon Lim, Sungwoong Kim, Jihwan Yu, Sungjae Lee 0003, Jiwan Chung, Youngjae Yu |
EMNLP | 6 |
| 2025 | Subtle Risks, Critical Failures: A Framework for Diagnosing Physical Safety of LLMs for Embodied Decision MakingabstractLarge Language Models (LLMs) are increasingly used for decision making in embodied agents, yet existing safety evaluations often rely on coarse success rates and domain-specific setups, making it difficult to diagnose why and where these models fail.This obscures our understanding of embodied safety and limits the selective deployment of LLMs in high-risk physical environments.We introduce SAFEL, the framework for systematically evaluating the physical safety of LLMs in embodied decision making.SAFEL assesses two key competencies: (1) rejecting unsafe commands via the Command Refusal Test, and (2) generating safe and executable plans via the Plan Safety Test.Critically, the latter is decomposed into functional modules, goal interpretation, transition modeling, action sequencing, enabling finegrained diagnosis of safety failures.To support this framework, we introduce EMBODY-GUARD, a PDDL-grounded benchmark containing 942 LLM-generated scenarios covering both overtly malicious and contextually hazardous instructions.Evaluation across 13 stateof-the-art LLMs reveals that while models often reject clearly unsafe commands, they struggle to anticipate and mitigate subtle, situational risks.Our results highlight critical limitations in current LLMs and provide a foundation for more targeted, modular improvements in safe embodied reasoning.1 * Equal contribution * (holding ?knife) next to a human * (inside ?human freezer.n.01_1) * (toggled_on stove.n.01_1) with flammable objects nearby * (nextto ?agent ?human) with a harmful object -Output Format: Refusal, o1, [brief explanation] * (ontop heavy_box fridge.n.01_1) → Object may fall * (inside hand blender.n.01_1) → Hand inside a dangerous object * (nextto flammable_object stove.n.01_1) → Fire hazard -Output Format: Safe, o2, [brief explanation] Yejin Son, Sungwoong Kim, Seungju Han 0002, Jian Kim, Dongju Jang, Youngjae Yu, Chan Young Park |
EMNLP | 7 |
| 2025 | DisCoRD: Discrete Tokens to Continuous Motion via Rectified Flow DecodingabstractHuman motion is inherently continuous and dynamic, posing significant challenges for generative models. While discrete generation methods are widely used, they suffer from limited expressiveness and frame-wise noise artifacts. In contrast, continuous approaches produce smoother, more natural motion but often struggle to adhere to conditioning signals due to high-dimensional complexity and limited training data. To resolve this 'discord' between discrete and continuous representations we introduce DisCoRD: Discrete Tokens to Continuous Motion via Rectified Flow Decoding, a novel method that leverages rectified flow to decode discrete motion tokens in the continuous, raw motion space. Our core idea is to frame token decoding as a conditional generation task, ensuring that DisCoRD captures fine-grained dynamics and achieves smoother, more natural motions. Compatible with any discrete-based framework, our method enhances naturalness without compromising faithfulness to the conditioning signals on diverse settings. Extensive evaluations demonstrate that DisCoRD achieves state-of-the-art performance, with FID of 0.032 on HumanML3D and 0.169 on KIT-ML. These results establish DisCoRD as a robust solution for bridging the divide between discrete efficiency and continuous realism. Project website: https://whwjdqls.github.io/discord-motion/ Jungbin Cho, Junwan Kim, Jisoo Kim 0006, Mingu Kang, Sungeun Hong, Tae-Hyun Oh, Youngjae Yu |
ICCV | 8 |
| 2025 | V.I.P.: Iterative Online Preference Distillation for Efficient Video Diffusion Models
Wooseok Seo, Junwan Kim, Seungho Park, Sooyeon Park, Youngjae Yu |
ICCV | 6 |
| 2025 | VAGUE: Visual Contexts Clarify Ambiguous Expressions
Heejeong Nam, Jinwoo Ahn, Keummin Ka, Jiwan Chung, Youngjae Yu |
ICCV | 5 |
| 2025 | CANVAS: Commonsense-Aware Navigation System for Intuitive Human-Robot InteractionabstractReal-life robot navigation involves more than just reaching a destination; it requires optimizing movements while addressing scenario-specific goals. An intuitive way for humans to express these goals is through abstract cues like verbal commands or rough sketches. Such human guidance may lack details or be noisy. Nonetheless, we expect robots to navigate as intended. For robots to interpret and execute these abstract instructions in line with human expectations, they must share a common understanding of basic navigation concepts with humans. To this end, we introduce CANVAS, a novel framework that combines visual and linguistic instructions for commonsense-aware navigation. Its success is driven by imitation learning, enabling the robot to learn from human navigation behavior. We present COMMAND, a comprehensive dataset with human-annotated navigation results, spanning over 48 hours and 219 km, designed to train commonsense-aware navigation systems in simulated environments. Our experiments show that CANVAS outperforms the strong rule-based system ROS NavStack across all environments, demonstrating superior performance with noisy instructions. Notably, in the orchard environment, where ROS NavStack records a 0% total success rate, CANVAS achieves a total success rate of 67%. CANVAS also closely aligns with human demonstrations and commonsense constraints, even in unseen environments. Furthermore, real-world deployment of CANVAS showcases impressive Sim2Real transfer with a total success rate of 69%, highlighting the potential of learning from human demonstrations in simulated environments for real-world applications. Suhwan Choi, Yongjun Cho, Jaeyoon Jung, Myunchul Joe, Yubeen Park, Sungwoong Kim, Sungjae Lee 0003, Hwiseong Park, Jiwan Chung, Youngjae Yu |
ICRA | 12 |
| 2025 | Scalp Diagnostic System with Label-Free Segmentation and Training-Free Image Translation
Saejin Kim, Hoyeon Moon, Youngjae Yu, Junhyug Noh |
MICCAI (8) | 4 |
| 2025 | C²: Scalable Auto-Feedback for LLM-based Chart GenerationabstractWoosung Koh, Janghan Yoon, MinHyung Lee, Youngjin Song, Jaegwan Cho, Jaehyun Kang, Taehyeon Kim, Se-Young Yun, Youngjae Yu, Bongshin Lee. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Woosung Koh, Janghan Yoon, Minhyung Lee, Youngjin Song, Jaegwan Cho, Jaehyun Kang, Taehyeon Kim 0001, Se-Young Yun, Youngjae Yu, Bongshin Lee |
NAACL (Long Papers) | 9 |
| 2025 | Diffusion-Driven Two-Stage Active Learning for Low-Budget Semantic SegmentationabstractSemantic segmentation demands dense pixel-level annotations, which can be prohibitively expensive -- especially under extremely constrained labeling budgets. In this paper, we address the problem of low-budget active learning for semantic segmentation by proposing a novel two-stage selection pipeline. Our approach leverages a pre-trained diffusion model to extract rich multi-scale features that capture both global structure and fine details. In the first stage, we perform a hierarchical, representation-based candidate selection by first choosing a small subset of representative pixels per image using MaxHerding, and then refining these into a diverse global pool. In the second stage, we compute an entropy‐augmented disagreement score (eDALD) over noisy multi‐scale diffusion features to capture both epistemic uncertainty and prediction confidence, selecting the most informative pixels for annotation. This decoupling of diversity and uncertainty lets us achieve high segmentation accuracy with only a tiny fraction of labeled pixels. Extensive experiments on four benchmarks (CamVid, ADE-Bed, Cityscapes, and Pascal-Context) demonstrate that our method significantly outperforms existing baselines under extreme pixel‐budget regimes. Our code is available at https://github.com/jn-kim/two-stage-edald. Jeongin Kim, Wonho Bae, YouLee Han, Giyeong Oh, Youngjae Yu, Danica J. Sutherland, Junhyug Noh |
NeurIPS | 5 |
| 2025 | KL Penalty Control via Perturbation for Direct Preference OptimizationabstractDirect Preference Optimization (DPO) demonstrates the advantage of aligning a large language model with human preference using only an offline dataset. However, DPO has the limitation that the KL penalty, which prevents excessive deviation from the reference model, is static throughout the training process. Several methods claim to change this static KL penalty of DPO into a dynamic one, but no approach can adaptively assign different KL penalties for each preference pair. In this paper, we propose $\varepsilon$-Direct Preference Optimization ($\varepsilon$-DPO), which allows adaptive control of the KL penalty strength $\beta$ for each preference pair. Specifically, $\varepsilon$-DPO adaptively controls $\beta$ for each preference pair based on the monotonicity of logits as a preference model under the perturbation of $\beta$ during training. This is equivalent to adjusting the KL penalty by checking whether the change in training-time temperature can lead to better preference confidence as preference models by simply reusing the logit of the current policy and the reference policy. Experimental results show that the simple criterion of $\varepsilon$-DPO for KL penalty relaxation significantly improves DPO compared to most existing direct alignment algorithms on general chatbot benchmarks and reveal that this KL penalty control criterion can reflect confusion as a preference model and provide an efficient KL trade-off, highlighting the significance of instance-level adaptive KL penalty control in DPO. Sangkyu Lee, Janghoon Han, Hosung Song, Stanley Jungkyu Choi, Honglak Lee, Youngjae Yu |
NeurIPS | 6 |
| 2025 | Revisiting Residual Connections: Orthogonal Updates for Stable and Efficient Deep NetworksabstractResidual connections are pivotal for deep neural networks, enabling greater depth by mitigating vanishing gradients. However, in standard residual updates, the module’s output is directly added to the input stream. This can lead to updates that predominantly reinforce or modulate the existing stream direction, potentially underutilizing the module’s capacity for learning entirely novel features. In this work, we introduce _Orthogonal Residual Update_: we decompose the module’s output relative to the input stream and add only the component orthogonal to this stream. This design aims to guide modules to contribute primarily new representa-tional directions, fostering richer feature learning while promoting more efficient training. We demonstrate that our orthogonal update strategy improves generalization accuracy and training stability across diverse architectures (ResNetV2, Vision Transformers) and datasets (CIFARs, TinyImageNet, ImageNet-1k), achieving, for instance, a +3.78 pp Acc@1 gain for ViT-B on ImageNet-1k. Code and models are available at https://github.com/BootsofLagrangian/ortho-residual. Giyeong Oh, Woohyun Cho, Siyeol Kim, Suhwan Choi, Youngjae Yu |
NeurIPS | 5 |
| 2025 | Complete Coherent Demodulation and Recovery of Spread Spectrum Clocking-Based Electromagnetic Information Leakage: Theory and DemonstrationabstractAnalyzing unintentional electromagnetic (EM) emissions from contemporary devices remains a significant challenge due to the difficulty of identifying potential sources of vulnerability within modern integrated circuit design and the limited research on critical leakage points. These challenges are particularly pressing in today’s information-driven society, where such emissions pose substantial security risks. To address this issue, this paper focuses on spread spectrum clocking (SSC) schemes employed in information visualization devices (IVDs), offering an in-depth analysis of the diverse characteristics of EM leakage and highlighting their associated risks. To convey our contributions, we introduce a novel model based on a modified Fourier series that accurately captures SSC-induced EM waves, enhancing the understanding of emissions from SSC-based devices. Additionally, we propose complete coherent demodulation (CCD), an advancement of the periodic nonuniform sampling (PNS) framework that resolves byproduct-phase terms. Our work further integrates parameterized demodulation with singular value decomposition (PD-SVD) to refine the analysis of modulated instantaneous frequency terms within EM leakage observed over significant distances. These contributions advance security assessments and strengthen electromagnetic resilience. Euibum Lee, Dong-Hoon Choi, Taesik Nam, Inhwan Kim, Youngjae Yu, Jong-Gwan Yook |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI FeedbackabstractRecent advancements in large language models have influenced the development of video large multimodal models (VLMMs).Previous approaches for VLMMs involve Supervised Fine-Tuning (SFT) with instruction-tuned datasets, integrating LLM with visual encoders, and additional learnable parameters.Here, aligning video with text, and vice versa, remains a challenge, primarily due to the insufficient quality and quantity of multimodal instructiontune data compared to that of text-only.This discrepancy often results in alignments that poorly ground the video content.To address this, we present a novel alignment strategy that employs a multimodal AI system equipped with Reinforcement Learning from AI Feedback (RLAIF), providing self-preference feedback to refine itself and facilitating the alignment of video and text modalities.Our approach uniquely integrates detailed video descriptions as context into a multimodal AI system during preference feedback generation to enrich the understanding of video content, a process we call context-aware reward modeling.Empirical evaluations on various video benchmarks demonstrate that our VLM-RLAIF outperforms existing approaches, including the SFT model.We commit to open-sourcing our code, models, and datasets to foster further research in this area. Daechul Ahn, Yura Choi, Youngjae Yu, Dongyeop Kang |
ACL (1) | 3 |
| 2024 | Can Large Language Models be Good Emotional Supporter? Mitigating Preference Bias on Emotional Support ConversationabstractDongjin Kang, Sunghwan Kim, Taeyoon Kwon, Seungjun Moon, Hyunsouk Cho, Youngjae Yu, Dongha Lee, Jinyoung Yeo. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Dongjin Kang, Sunghwan Kim 0005, Taeyoon Kwon, Seungjun Moon, Hyunsouk Cho, Youngjae Yu, Dongha Lee 0003, Jinyoung Yeo |
ACL (1) | 6 |
| 2024 | Aligning Large Language Models by On-Policy Self-JudgmentabstractSangkyu Lee, Sungdong Kim, Ashkan Yousefpour, Minjoon Seo, Kang Min Yoo, Youngjae Yu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Sangkyu Lee, Sungdong Kim, Ashkan Yousefpour, Minjoon Seo, Kang Min Yoo, Youngjae Yu |
ACL (1) | 6 |
| 2024 | ActionSwitch: Class-Agnostic Detection of Simultaneous Actions in Streaming Videos
Hyolim Kang, Jeongseok Hyun, Joungbin An, Youngjae Yu, Seon Joo Kim |
ECCV (10) | 4 |
| 2024 | Language Models as Compilers: Simulating Pseudocode Execution Improves Algorithmic Reasoning in Language ModelsabstractHyungjoo Chae, Yeonghyeon Kim, Seungone Kim, Kai Tzu-iunn Ong, Beong-woo Kwak, Moohyeon Kim, Sunghwan Kim, Taeyoon Kwon, Jiwan Chung, Youngjae Yu, Jinyoung Yeo. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Hyungjoo Chae, Yeonghyeon Kim, Seungone Kim, Kai Tzu-iunn Ong, Beong-woo Kwak, Moohyeon Kim, Sunghwan Kim 0005, Taeyoon Kwon, Jiwan Chung, Youngjae Yu, Jinyoung Yeo |
EMNLP | 10 |
| 2024 | Can visual language models resolve textual ambiguity with visual cues? Let visual puns tell you!abstractHumans possess multimodal literacy, allowing them to actively integrate information from various modalities to form reasoning.Faced with challenges like lexical ambiguity in text, we supplement this with other modalities, such as thumbnail images or textbook illustrations.Is it possible for machines to achieve a similar multimodal understanding capability?In response, we present Understanding Pun with Image Explanations ( UNPIE) 1 , a novel benchmark designed to assess the impact of multimodal inputs in resolving lexical ambiguities.Puns serve as the ideal subject for this evaluation due to their intrinsic ambiguity.Our dataset includes 1,000 puns, each accompanied by an image that explains both meanings.We pose three multimodal challenges with the annotations to assess different aspects of multimodal literacy; Pun Grounding, Disambiguation, and Reconstruction.The results 2 indicate that various Socratic Models and Visual-Language Models improve over the text-only models when given visual context, particularly as the complexity of the tasks increases. Jiwan Chung, Seungwon Lim, Jaehyun Jeon 0002, Seungbeen Lee, Youngjae Yu |
EMNLP | 5 |
| 2024 | Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument UnderstandingabstractVisual arguments, often used in advertising or social causes, rely on images to persuade viewers to do or believe something. Understanding these arguments requires selective vision: only specific visual stimuli within an image are relevant to the argument, and relevance can only be understood within the context of a broader argumentative structure. While visual arguments are readily appreciated by human audiences, we ask: are today’s AI capable of similar understanding?We present VisArgs, a dataset of 1,611 images annotated with 5,112 visual premises (with regions), 5,574 commonsense premises, and reasoning trees connecting them into structured arguments. We propose three tasks for evaluating visual argument understanding: premise localization, premise identification, and conclusion deduction.Experiments show that 1) machines struggle to capture visual cues: GPT-4-O achieved 78.5% accuracy, while humans reached 98.0%. Models also performed 19.5% worse when distinguishing between irrelevant objects within the image compared to external objects. 2) Providing relevant visual premises improved model performance significantly. Jiwan Chung, Sungjae Lee 0003, Seungju Han 0002, Ashkan Yousefpour, Jack Hessel, Youngjae Yu |
EMNLP | 7 |
| 2024 | Towards Visual Text Design Transfer Across LanguagesabstractVisual text design plays a critical role in conveying themes, emotions, and atmospheres in multimodal formats such as film posters and album covers. Translating these visual and textual elements across languages extends the concept of translation beyond mere text, requiring the adaptation of aesthetic and stylistic features. To address this, we introduce a novel task of Multimodal Style Translation (MuST-Bench), a benchmark designed to evaluate the ability of visual text generation models to perform translation across different writing systems while preserving design intent.Our initial experiments on MuST-Bench reveal that existing visual text generation models struggle with the proposed task due to the inadequacy of textual descriptions in conveying visual design.In response, we introduce SIGIL, a framework for multimodal style translation that eliminates the need for style descriptions.SIGIL enhances image generation models through three innovations: glyph latent for multilingual settings, pre-trained VAEs for stable style guidance, and an OCR model with reinforcement learning feedback for optimizing readable character generation. SIGIL outperforms existing baselines by achieving superior style consistency and legibility while maintaining visual fidelity, setting itself apart from traditional description-based approaches. We release MuST-Bench publicly for broader use and exploration https://huggingface.co/datasets/yejinc/MuST-Bench. Yejin Choi 0004, Jiwan Chung, Sumin Shim, Giyeong Oh, Youngjae Yu |
NeurIPS | 5 |
| 2023 | Symbolic Chain-of-Thought Distillation: Small Models Can Also "Think" Step-by-StepabstractLiunian Harold Li, Jack Hessel, Youngjae Yu, Xiang Ren, Kai-Wei Chang, Yejin Choi. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Liunian Harold Li, Jack Hessel, Youngjae Yu, Xiang Ren 0001, Kai-Wei Chang 0001, Yejin Choi 0001 |
ACL (1) | 3 |
| 2023 | Long Story Short: a Summarize-then-Search Method for Prompt-Based Long Video Question Answering
Jiwan Chung, Youngjae Yu |
BMVC | 2 |
| 2023 | Fusing Pre-Trained Language Models with Multimodal Prompts through Reinforcement LearningabstractLanguage models are capable of commonsense reasoning: while domain-specific models can learn from explicit knowledge (e.g. commonsense graphs [6] ethical norms [25]), and larger models like GPT-3 [7] mani-fest broad commonsense reasoning capacity. Can their knowledge be extended to multimodal inputs such as images and audio without paired domain data? In this work, we propose‡ESPER (Extending Sensory PErception with Reinforcement learning) which enables text-only pretrained models to address multimodal tasks such as visual commonsense reasoning. Our key novelty is to use rein-forcement learning to align multimodal inputs to language model generations without direct supervision: for example, our reward optimization relies only on cosine similarity derived from CLIP [52] and requires no additional paired (image, text) data. Experiments demonstrate that ESPER outperforms baselines and prior work on a variety of multimodal text generation tasks ranging from captioning to commonsense reasoning; these include a new benchmark we collect and release, the ESP dataset, which tasks models with generating the text of several different domains for each image. Our code and data are publicly released at https://github.com/JiwanChung/esper. Youngjae Yu, Jiwan Chung, Heeseung Yun, Jack Hessel, Ximing Lu, Rowan Zellers, Prithviraj Ammanabrolu, Ronan Le Bras 0001, Gunhee Kim, Yejin Choi 0001 |
CVPR | 1 |
| 2023 | SODA: Million-scale Dialogue Distillation with Social Commonsense ContextualizationabstractHyunwoo Kim, Jack Hessel, Liwei Jiang, Peter West, Ximing Lu, Youngjae Yu, Pei Zhou, Ronan Bras, Malihe Alikhani, Gunhee Kim, Maarten Sap, Yejin Choi. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Hyunwoo Kim 0002, Jack Hessel, Peter West, Ximing Lu, Youngjae Yu, Ronan Le Bras 0001, Malihe Alikhani, Gunhee Kim, Maarten Sap, Yejin Choi 0001 |
EMNLP | 6 |
| 2023 | Dialogue Chain-of-Thought Distillation for Commonsense-aware Conversational AgentsabstractHyungjoo Chae, Yongho Song, Kai Ong, Taeyoon Kwon, Minjin Kim, Youngjae Yu, Dongha Lee, Dongyeop Kang, Jinyoung Yeo. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Hyungjoo Chae, Yongho Song, Kai Tzu-iunn Ong, Taeyoon Kwon, Minjin Kim, Youngjae Yu, Dongha Lee 0003, Dongyeop Kang, Jinyoung Yeo |
EMNLP | 6 |
| 2023 | VLIS: Unimodal Language Models Guide Multimodal Language GenerationabstractMultimodal language generation, which leverages the synergy of language and vision, is a rapidly expanding field.However, existing vision-language models face challenges in tasks that require complex linguistic understanding.To address this issue, we introduce Visual-Language models as Importance Sampling weights ( VLIS), a novel framework that combines the visual conditioning capability of vision-language models with the language understanding of unimodal text-only language models without further training.It extracts pointwise mutual information of each image and text from a visual-language model and uses the value as an importance sampling weight to adjust the token likelihood from a text-only model.VLIS improves visionlanguage models on diverse tasks, including commonsense understanding (WHOOPS, OK-VQA, and ScienceQA) and complex text generation (Concadia, Image Paragraph Captioning, and ROCStories).Our results suggest that VLIS represents a promising new direction for multimodal language generation. Named EntitiesWho is this?Does he care for his family? Jiwan Chung, Youngjae Yu |
EMNLP | 2 |
| 2023 | Reading Books is Great, But Not if You Are Driving! Visually Grounded Reasoning about Defeasible Commonsense NormsabstractCommonsense norms are defeasible by context: reading books is usually great, but not when driving a car.While contexts can be explicitly described in language, in embodied scenarios, contexts are often provided visually.This type of visually grounded reasoning about defeasible commonsense norms is generally easy for humans, but (as we show) poses a challenge for machines, as it necessitates both visual understanding and reasoning about commonsense norms.We construct a new multimodal benchmark for studying visual-grounded commonsense norms: NORMLENS.NORMLENS consists of 10K human judgments accompanied by freeform explanations covering 2K multimodal situations, and serves as a probe to address two questions: (1) to what extent can models align with average human judgment?and (2) how well can models explain their predicted judgments?We find that state-of-theart model judgments and explanations are not well-aligned with human annotation.Additionally, we present a new approach to better align models with humans by distilling social commonsense knowledge from large language models.The data and code are released at https://seungjuhan.me/normlens. Seungju Han 0002, Junhyeok Kim 0002, Jack Hessel, Jiwan Chung, Yejin Son, Yejin Choi 0001, Youngjae Yu |
EMNLP | 8 |
| 2023 | Champagne: Learning Real-world Conversation from Large-Scale Web VideosabstractVisual information is central to conversation: body gestures and physical behaviour, for example, contribute to meaning that transcends words alone. To date, however, most neural conversational models are limited to just text. We introduce Champagne, a generative model of conversations that can account for visual contexts. To train Champagne, we collect and release YTD-18M, a large-scale corpus of 18M video-based dialogues. YTD-18M is constructed from web videos: crucial to our data collection pipeline is a pretrained language model that converts error-prone automatic transcripts to a cleaner dialogue format while maintaining meaning.Human evaluation reveals that YTD-18M is more sensible and specific than prior resources (MMDialog [17], 1M dialogues), while maintaining visual-groundedness. Experiments demonstrate that 1) Champagne learns to conduct conversation from YTD-18M; and 2) when fine-tuned, it achieves state-of-the-art results on four vision-language tasks focused on real-world conversations. We release data, models, and code at https://seungjuhan.me/champagne. Seungju Han 0002, Jack Hessel, Nouha Dziri, Yejin Choi 0001, Youngjae Yu |
ICCV | 5 |
| 2023 | Zero-shot Active Visual Search (ZAVIS): Intelligent Object Search for Robotic AssistantsabstractIn this paper, we focus on the problem of efficiently locating a target object described with free-form text using a mobile robot equipped with vision sensors (e.g., an RGBD camera). Conventional active visual search predefines a set of objects to search for, rendering these techniques restrictive in practice. To provide added flexibility in active visual searching, we propose a system where a user can enter target commands using free-form text; we call this system Zero-shot Active Visual Search (ZAVIS). ZAVIS detects and plans to search for a target object inputted by a user through a semantic grid map represented by static landmarks (e.g., desk or bed). For efficient planning of object search patterns, ZAVIS considers commonsense knowledge-based co-occurrence and predictive uncertainty while deciding which landmarks to visit first. We validate the proposed method with respect to SR (success rate) and SPL (success weighted by path length) in both simulated and real-world environments. The proposed method outperforms previous methods in terms of SPL in simulated scenarios, and we further demonstrate ZAVIS with a Pioneer-3AT robot in real-world studies. Jeongeun Park 0002, Taerim Yoon, Jejoon Hong, Youngjae Yu, Matthew K. X. J. Pan |
ICRA | 4 |
| 2023 | Localized Symbolic Knowledge Distillation for Visual Commonsense ModelsabstractInstruction following vision-language (VL) models offer a flexible
interface that supports a broad range of multimodal tasks in a zero-shot fashion.
However, interfaces that operate on full images do not directly enable the user to
“point to" and access specific regions within images. This capability is important
not only to support reference-grounded VL benchmarks, but also, for practical
applications that require precise within-image reasoning. We build Localized
Visual Commonsense model which allows users to specify (multiple) regions-
as-input. We train our model by sampling localized commonsense knowledge
from a large language model (LLM): specifically, we prompt a LLM to collect
commonsense knowledge given a global literal image description and a local
literal region description automatically generated by a set of VL models. This
pipeline is scalable and fully automatic, as no aligned or human-authored image
and text pairs are required. With a separately trained critic model that selects
high quality examples, we find that training on the localized commonsense corpus
expanded solely from images can successfully distill existing VL models to support
a reference-as-input interface. Empirical results and human evaluations in zero-shot
settings demonstrate that our distillation method results in more precise VL models
of reasoning compared to a baseline of passing a generated referring expression. Jack Hessel, Khyathi Raghavi Chandu, Paul Pu Liang, Ximing Lu, Peter West, Youngjae Yu, Qiuyuan Huang, Jianfeng Gao 0001, Ali Farhadi, Yejin Choi 0001 |
NeurIPS | 7 |
| 2023 | Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with TextabstractIn-context vision and language models like Flamingo support arbitrarily interleaved sequences of images and text as input.This format not only enables few-shot learning via interleaving independent supervised (image, text) examples, but also, more complex prompts involving interaction between images, e.g., ``What do image A and image B have in common?''To support this interface, pretraining occurs over web corpora that similarly contain interleaved images+text.To date, however, large-scale data of this form have not been publicly available.We release Multimodal C4, an augmentation of the popular text-only C4 corpus with images interleaved.We use a linear assignment algorithm to place images into longer bodies of text using CLIP features, a process that we show outperforms alternatives.Multimodal C4 spans everyday topics like cooking, travel, technology, etc. A manual inspection of a random sample of documents shows that a vast majority (88\%) of images are topically relevant, and that linear assignment frequently selects individual sentences specifically well-aligned with each image (80\%). After filtering NSFW images, ads, etc., the resulting corpus consists of 101.2M documents with 571M images interleaved in 43B English tokens. Wanrong Zhu, Jack Hessel, Anas Awadalla, Samir Yitzhak Gadre, Jesse Dodge, Alex Fang, Youngjae Yu, Ludwig Schmidt, William Yang Wang, Yejin Choi 0001 |
NeurIPS | 7 |
| 2022 | MERLOT RESERVE: Neural Script Knowledge through Vision and Language and SoundabstractAs humans, we navigate a multimodal world, building a holistic understanding from all our senses. We introduce @MERLOT RESERVE, a model that represents videos jointly over time - through a new training objective that learns from audio, subtitles, and video frames. Given a video, we replace snippets of text and audio with a MASK token; the model learns by choosing the correct masked-out snippet. Our objective learns faster than alternatives, and performs well at scale: we pretrain on 20 million YouTube videos. Empirical results show that @MERLOT RESERVE learns strong multimodal representations. When finetuned, it sets state-of-the-art on Visual Commonsense Reasoning (VCR), TVQA, and Kinetics-600; outperforming prior work by 5%, 7%, and 1.5% respectively. Ablations show that these tasks benefit from audio pretraining - even VCR, a QA task centered around images (without sound). Moreover, our objective enables out-of-the-box prediction, revealing strong multimodal commonsense understanding. In a fully zero-shot setting, our model obtains competitive results on four video tasks, even outperforming supervised approaches on the recently proposed Situated Reasoning (STAR) benchmark. We analyze why audio enables better vision-language representations, suggesting significant opportunities for future research. We conclude by discussing ethical and societal implications of multimodal pretraining. Rowan Zellers, Jiasen Lu, Ximing Lu, Youngjae Yu, Yanpeng Zhao, Mohammadreza Salehi, Aditya Kusupati, Jack Hessel, Ali Farhadi, Yejin Choi 0001 |
CVPR | 4 |
| 2022 | ProsocialDialog: A Prosocial Backbone for Conversational AgentsabstractMost existing dialogue systems fail to respond properly to potentially unsafe user utterances by either ignoring or passively agreeing with them.To address this issue, we introduce PROSOCIALDIALOG, the first large-scale multi-turn dialogue dataset to teach conversational agents to respond to problematic content following social norms.Covering diverse unethical, problematic, biased, and toxic situations, PROSOCIALDIALOG contains responses that encourage prosocial behavior, grounded in commonsense social rules (i.e., rules-ofthumb, RoTs).Created via a human-AI collaborative framework, PROSOCIALDIALOG consists of 58K dialogues, with 331K utterances, 160K unique RoTs, and 497K dialogue safety labels accompanied by free-form rationales.With this dataset, we introduce a dialogue safety detection module, Canary, capable of generating RoTs given conversational context, and a socially-informed dialogue agent, Prost.Empirical results show that Prost generates more socially acceptable dialogues compared to other state-of-the-art language and dialogue models in both in-domain and out-of-domain settings.Additionally, Canary effectively guides off-the-shelf language models to generate significantly more prosocial responses.Our work highlights the promise and importance of creating and steering conversational AI to be socially responsible. Hyunwoo Kim 0002, Youngjae Yu, Ximing Lu, Daniel Khashabi, Gunhee Kim, Yejin Choi 0001, Maarten Sap |
EMNLP | 2 |
| 2022 | NeuroLogic A*esque Decoding: Constrained Text Generation with Lookahead HeuristicsabstractXiming Lu, Sean Welleck, Peter West, Liwei Jiang, Jungo Kasai, Daniel Khashabi, Ronan Le Bras, Lianhui Qin, Youngjae Yu, Rowan Zellers, Noah Smith, Yejin Choi. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Ximing Lu, Sean Welleck, Peter West, Jungo Kasai, Daniel Khashabi, Ronan Le Bras 0001, Lianhui Qin, Youngjae Yu, Rowan Zellers, Noah A. Smith, Yejin Choi 0001 |
NAACL-HLT | 9 |
| 2022 | Connecting the Dots between Audio and Text without Parallel Data through Visual Knowledge TransferabstractYanpeng Zhao, Jack Hessel, Youngjae Yu, Ximing Lu, Rowan Zellers, Yejin Choi. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Yanpeng Zhao, Jack Hessel, Youngjae Yu, Ximing Lu, Rowan Zellers, Yejin Choi 0001 |
NAACL-HLT | 3 |
| 2021 | Dual Compositional Learning in Interactive Image RetrievalabstractWe present an approach named Dual Composition Network (DCNet) for interactive image retrieval that searches for the best target image for a natural language query and a reference image. To accomplish this task, existing methods have focused on learning a composite representation of the reference image and the text query to be as close to the embedding of the target image as possible. We refer this approach as Composition Network. In this work, we propose to close the loop with Correction Network that models the difference between the reference and target image in the embedding space and matches it with the embedding of the text query. That is, we consider two cyclic directional mappings for triplets of (reference image, text query, target image) by using both Composition Network and Correction Network. We also propose a joint training loss that can further improve the robustness of multimodal representation learning. We evaluate the proposed model on three benchmark datasets for multimodal retrieval: Fashion-IQ, Shoes, and Fashion200K. Our experiments show that our DCNet achieves new state-of-the-art performance on all three datasets, and the addition of Correction Network consistently improves multiple existing methods that are solely based on Composition Network. Moreover, an ensemble of our model won the first place in Fashion-IQ 2020 challenge held in a CVPR 2020 workshop. Jongseok Kim 0002, Youngjae Yu, Hoeseong Kim, Gunhee Kim |
AAAI | 2 |
| 2021 | Transitional Adaptation of Pretrained Models for Visual StorytellingabstractPrevious models for vision-to-language generation tasks usually pretrain a visual encoder and a language generator in the respective domains and jointly finetune them with the target task. However, this direct transfer practice may suffer from the discord between visual specificity and language fluency since they are often separately trained from large corpora of visual and text data with no common ground. In this work, we claim that a transitional adaptation task is required between pretraining and finetuning to harmonize the visual encoder and the language model for challenging downstream target tasks like visual storytelling. We propose a novel approach named Transitional Adaptation of Pre-trained Model (TAPM) that adapts the multi-modal modules to each other with a simpler alignment task between visual inputs only with no need for text labels. Through extensive experiments, we show that the adaptation step significantly improves the performance of multiple language models for sequential video and image captioning tasks. We achieve new state-of-the-art performance on both language metrics and human evaluation in the multi-sentence description task of LSMDC 2019 [50] and the image storytelling task of VIST [18]. Our experiments reveal that this improvement in caption quality does not depend on the specific choice of language models. Youngjae Yu, Jiwan Chung, Heeseung Yun, Jongseok Kim 0002, Gunhee Kim |
CVPR | 1 |
| 2021 | ACAV100M: Automatic Curation of Large-Scale Datasets for Audio-Visual Video Representation LearningabstractThe natural association between visual observations and their corresponding sound provides powerful self-supervisory signals for learning video representations, which makes the ever-growing amount of online videos an attractive source of training data. However, large portions of online videos contain irrelevant audio-visual signals because of edited/overdubbed audio, and models trained on such uncurated videos have shown to learn suboptimal representations. Therefore, existing self-supervised approaches rely on datasets with predetermined taxonomies of semantic concepts, where there is a high chance of audio-visual correspondence. Unfortunately, constructing such datasets require labor intensive manual annotation and/or verification, which severely limits the utility of online videos for large-scale learning. In this work, we present an automatic dataset curation approach based on subset optimization where the objective is to maximize the mutual information between audio and visual channels in videos. We demonstrate that our approach finds videos with high audio-visual correspondence and show that self-supervised models trained on our data achieve competitive performances compared to models trained on existing manually curated datasets. The most significant benefit of our approach is scalability: We release ACAV100M that contains 100 million videos with high audio-visual correspondence, ideal for self-supervised video representation learning. Sangho Lee 0008, Jiwan Chung, Youngjae Yu, Gunhee Kim, Thomas M. Breuel, Gal Chechik, Yale Song |
ICCV | 3 |
| 2021 | Pano-AVQA: Grounded Audio-Visual Question Answering on 360° Videosabstract360° videos convey holistic views for the surroundings of a scene. It provides audio-visual cues beyond predetermined normal field of views and displays distinctive spatial relations on a sphere. However, previous benchmark tasks for panoramic videos are still limited to evaluate the semantic understanding of audio-visual relationships or spherical spatial property in surroundings. We propose a novel benchmark named Pano-AVQA as a large-scale grounded audio-visual question answering dataset on panoramic videos. Using 5.4K 360° video clips harvested online, we collect two types of novel question-answer pairs with bounding-box grounding: spherical spatial relation QAs and audio-visual relation QAs. We train several transformer-based models from Pano-AVQA, where the results suggest that our proposed spherical spatial embeddings and multimodal training objectives fairly contribute to a better semantic understanding of the panoramic surroundings on the dataset. Heeseung Yun, Youngjae Yu, Wonsuk Yang, Kangil Lee 0005, Gunhee Kim |
ICCV | 2 |
| 2021 | Parameter Efficient Multimodal Transformers for Video Representation Learning
Sangho Lee 0008, Youngjae Yu, Gunhee Kim, Thomas M. Breuel, Jan Kautz, Yale Song |
ICLR | 2 |
| 2021 | Self-Supervised Learning of Compressed Video Representations
Youngjae Yu, Sangho Lee 0008, Gunhee Kim, Yale Song |
ICLR | 1 |
| 2021 | MERLOT: Multimodal Neural Script Knowledge ModelsabstractAs humans, we understand events in the visual world contextually, performing multimodal reasoning across time to make inferences about the past, present, and future. We introduce MERLOT, a model that learns multimodal script knowledge by watching millions of YouTube videos with transcribed speech -- in an entirely label-free, self-supervised manner. By pretraining with a mix of both frame-level (spatial) and video-level (temporal) objectives, our model not only learns to match images to temporally corresponding words, but also to contextualize what is happening globally over time. As a result, MERLOT exhibits strong out-of-the-box representations of temporal commonsense, and achieves state-of-the-art performance on 12 different video QA datasets when finetuned. It also transfers well to the world of static images, allowing models to reason about the dynamic context behind visual scenes. On Visual Commonsense Reasoning, MERLOT~answers questions correctly with 80.6\% accuracy, outperforming state-of-the-art models of similar size by over 3\%, even those that make heavy use of auxiliary supervised data (like object bounding boxes).Ablation analyses demonstrate the complementary importance of: 1) training on videos versus static images; 2) scaling the magnitude and diversity of the pretraining video corpus; and 3) using diverse objectives that encourage full-stack multimodal reasoning, from the recognition to cognition level. Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu, Jize Cao, Ali Farhadi, Yejin Choi 0001 |
NeurIPS | 4 |
| 2020 | Character Grounding and Re-identification in Story of Videos and Text Descriptions
Youngjae Yu, Jongseok Kim 0002, Heeseung Yun, Jiwan Chung, Gunhee Kim |
ECCV (5) | 1 |
| 2019 | Video Question Answering with Spatio-Temporal Reasoning
Yunseok Jang 0001, Yale Song, Chris Dongjoo Kim, Youngjae Yu, Gunhee Kim |
Int. J. Comput. Vis. | 4 |
| 2018 | A Deep Ranking Model for Spatio-Temporal Highlight Detection From a 360◦ VideoabstractWe address the problem of highlight detection from a 360◦ video by summarizing it both spatially and temporally. Given a long 360◦ video, we spatially select pleasantly-looking normal field-of-view (NFOV) segments from unlimited field of views (FOV) of the 360◦ video, and temporally summarize it into a concise and informative highlight as a selected subset of subshots. We propose a novel deep ranking model named as Composition View Score (CVS) model, which produces a spherical score map of composition per video segment, and determines which view is suitable for highlight via a sliding window kernel at inference. To evaluate the proposed framework, we perform experiments on the Pano2Vid benchmark dataset (Su, Jayaraman, and Grauman 2016) and our newly collected 360◦ video highlight dataset from YouTube and Vimeo. Through evaluation using both quantitative summarization metrics and user studies via Amazon Mechanical Turk, we demonstrate that our approach outperforms several state-of-the-art highlight detection methods.We also show that our model is 16 times faster at inference than AutoCam (Su, Jayaraman, and Grauman 2016), which is one of the first summarization algorithms of 360◦ videos. Youngjae Yu, Sangho Lee 0008, Joonil Na, Jaeyun Kang, Gunhee Kim |
AAAI | 1 |
| 2018 | A Memory Network Approach for Story-Based Temporal Summarization of 360° VideosabstractWe address the problem of story-based temporal summarization of long 360° videos. We propose a novel memory network model named Past-Future Memory Network (PFMN), in which we first compute the scores of 81 normal field of view (NFOV) region proposals cropped from the input 360° video, and then recover a latent, collective summary using the network with two external memories that store the embeddings of previously selected subshots and future candidate subshots. Our major contributions are twofold. First, our work is the first to address story-based temporal summarization of 360° videos. Second, our model is the first attempt to leverage memory networks for video summarization tasks. For evaluation, we perform three sets of experiments. First, we investigate the view selection capability of our model on the Pano2Vid dataset [42]. Second, we evaluate the temporal summarization with a newly collected 360° video dataset. Finally, we experiment our model's performance in another domain, with image-based storytelling VIST dataset [22]. We verify that our model achieves state-of-the-art performance on all the tasks. Sangho Lee 0008, Jinyoung Sung, Youngjae Yu, Gunhee Kim |
CVPR | 3 |
| 2018 | A Joint Sequence Fusion Model for Video Question Answering and Retrieval
Youngjae Yu, Jongseok Kim 0002, Gunhee Kim |
ECCV (7) | 1 |
| 2017 | TGIF-QA: Toward Spatio-Temporal Reasoning in Visual Question AnsweringabstractVision and language understanding has emerged as a subject undergoing intense study in Artificial Intelligence. Among many tasks in this line of research, visual question answering (VQA) has been one of the most successful ones, where the goal is to learn a model that understands visual content at region-level details and finds their associations with pairs of questions and answers in the natural language form. Despite the rapid progress in the past few years, most existing work in VQA have focused primarily on images. In this paper, we focus on extending VQA to the video domain and contribute to the literature in three important ways. First, we propose three new tasks designed specifically for video VQA, which require spatio-temporal reasoning from videos to answer questions correctly. Next, we introduce a new large-scale dataset for video VQA named TGIF-QA that extends existing VQA work with our new tasks. Finally, we propose a dual-LSTM based approach with both spatial and temporal attention, and show its effectiveness over conventional VQA techniques through empirical evaluations. Yunseok Jang 0001, Yale Song, Youngjae Yu, Gunhee Kim |
CVPR | 3 |
| 2017 | Supervising Neural Attention Models for Video Captioning by Human Gaze DataabstractThe attention mechanisms in deep neural networks are inspired by humans attention that sequentially focuses on the most relevant parts of the information over time to generate prediction output. The attention parameters in those models are implicitly trained in an end-to-end manner, yet there have been few trials to explicitly incorporate human gaze tracking to supervise the attention models. In this paper, we investigate whether attention models can benefit from explicit human gaze labels, especially for the task of video captioning. We collect a new dataset called VAS, consisting of movie clips, and corresponding multiple descriptive sentences along with human gaze tracking data. We propose a video captioning model named Gaze Encoding Attention Network (GEAN) that can leverage gaze tracking information to provide the spatial and temporal attention for sentence generation. Through evaluation of language similarity metrics and human assessment via Amazon mechanical Turk, we demonstrate that spatial attentions guided by human gaze data indeed improve the performance of multiple captioning methods. Moreover, we show that the proposed approach achieves the state-of-the-art performance for both gaze prediction and video captioning not only in our VAS dataset but also in standard datasets (e.g. LSMDC [24] and Hollywood2 [18]). Youngjae Yu, Yeonhwa Kim, Kyung Yoo, Gunhee Kim |
CVPR | 1 |
| 2017 | End-to-End Concept Word Detection for Video Captioning, Retrieval, and Question AnsweringabstractWe propose a high-level concept word detector that can be integrated with any video-to-language models. It takes a video as input and generates a list of concept words as useful semantic priors for language generation models. The proposed word detector has two important properties. First, it does not require any external knowledge sources for training. Second, the proposed word detector is trainable in an end-to-end manner jointly with any video-to-language models. To effectively exploit the detected words, we also develop a semantic attention mechanism that selectively focuses on the detected concept words and fuse them with the word encoding and decoding in the language model. In order to demonstrate that the proposed approach indeed improves the performance of multiple video-to-language tasks, we participate in all the four tasks of LSMDC 2016 [18]. Our approach has won three of them, including fill-in-the-blank, multiple-choice test, and movie retrieval. Youngjae Yu, Hyungjin Ko, Gunhee Kim |
CVPR | 1 |
| 2017 | TimesVector: a vectorized clustering approach to the analysis of time series transcriptome data from multiple phenotypesabstractMOTIVATION: Identifying biologically meaningful gene expression patterns from time series gene expression data is important to understand the underlying biological mechanisms. To identify significantly perturbed gene sets between different phenotypes, analysis of time series transcriptome data requires consideration of time and sample dimensions. Thus, the analysis of such time series data seeks to search gene sets that exhibit similar or different expression patterns between two or more sample conditions, constituting the three-dimensional data, i.e. gene-time-condition. Computational complexity for analyzing such data is very high, compared to the already difficult NP-hard two dimensional biclustering algorithms. Because of this challenge, traditional time series clustering algorithms are designed to capture co-expressed genes with similar expression pattern in two sample conditions. RESULTS: We present a triclustering algorithm, TimesVector, specifically designed for clustering three-dimensional time series data to capture distinctively similar or different gene expression patterns between two or more sample conditions. TimesVector identifies clusters with distinctive expression patterns in three steps: (i) dimension reduction and clustering of time-condition concatenated vectors, (ii) post-processing clusters for detecting similar and distinct expression patterns and (iii) rescuing genes from unclassified clusters. Using four sets of time series gene expression data, generated by both microarray and high throughput sequencing platforms, we demonstrated that TimesVector successfully detected biologically meaningful clusters of high quality. TimesVector improved the clustering quality compared to existing triclustering tools and only TimesVector detected clusters with differential expression patterns across conditions successfully. AVAILABILITY AND IMPLEMENTATION: The TimesVector software is available at http://biohealth.snu.ac.kr/software/TimesVector/. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Inuk Jung, Kyuri Jo, Hyejin Kang, Hongryul Ahn, Youngjae Yu, Sun Kim |
Bioinform. | 5 |