Junyeong Kim

dblp:28/9716 · DBLP profile ↗
← Back
23ranked-venue papers
7as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 10 since 2021
YearPublicationVenuePosition
2026 GranAlign: Granularity-Aware Alignment Framework for Zero-shot Video Moment Retrieval
abstract
Zero-shot video moment retrieval (ZVMR) is the task of localizing a temporal moment within an untrimmed video using a natural language query without relying on task-specific training data. The primary challenge in this setting lies in the mismatch in semantic granularity between textual queries and visual content. Previous studies in ZVMR have attempted to achieve alignment by leveraging high-quality pre-trained knowledge that represents video and language in a joint space. However, these approaches failed to balance the semantic granularity between the pre-trained knowledge provided by each modality for a given scene. As a result, despite the high quality of each modality’s representations, the mismatch in granularity led to inaccurate retrieval. In this paper, we propose a training-free framework, called Granularity-Aware Alignment (GranAlign), that bridges this gap between coarse and fine semantic representations. Our approach introduces two complementary techniques: granularity-based query rewriting to generate varied semantic granularities, and query-aware caption generation to embed query intent into video content. By pairing multi-level queries with both query-agnostic and query-aware captions, we effectively resolve semantic mismatches. As a result, our method sets a new state-of-the-art across all three major benchmarks (QVHighlights, Charades-STA, ActivityNet-Captions), with a notable 3.23% mAP@avg improvement on the QVHighlights dataset.
Mingyu Jeon, Sunjae Yoon, Jonghee Kim, Junyeong Kim
AAAI4
2026 Selective Test-Time Debiasing for CLIP via Reward Gating
abstract
Vision language models (VLMs) demonstrate strong zero-shot performance, but often perpetuate social stereotypes in person-centric queries, yielding skewed demographic distributions.Current debiasing methods apply uniform bias corrections across all input queries regardless of their bias sensitivity, creating a fundamental fairness-utility trade-off.Strong debiasing distorts semantically meaningful information in bias-insensitive queries, while weak debiasing fails to mitigate stereotypes in biassensitive ones.This one-size-fits-all approach hampers simultaneously achieving high utility on bias-insensitive queries and fairness on bias-sensitive queries.We introduce Reward-Gated Test-Time Adaptation (RG-TTA), a reinforcement learning-based test-time adaptation framework that selectively applies debiasing based on input sensitivity.RG-TTA adaptively triggers fairness regularization based on the bias sensitivity of each input during testtime policy adaptation, while focusing exclusively on optimizing cross-modal alignment for bias-insensitive inputs.Experiments on fairness benchmarks (e.g., FairFace, UTKFace) demonstrate substantial bias reduction while simultaneously improving zero-shot utility, resolving the trade-off of uniform debiasing.
Jaeho Han, Jisoo Yang, Hyeondong Woo, Mingyu Jeon, Sunjae Yoon, Junyeong Kim
ACL (1)6
2026 Video DETOX: Purifying Noisy Relevance Signals for Diverse and Long-Form Video Understanding
Sungjin Han, Minuk Ma, Trung X. Pham, Junyeong Kim
ICPR (14)4
2026 Blocking Visual Leakage: Visually-Agnostic Text Decomposition for Composed Video Retrieval
Jinkwon Hwang, Minuk Ma, Trung X. Pham, Junyeong Kim
ICPR (14)4
2026 Beyond Co-existence: Measuring Attribute Binding Hallucinations in Audio-Language Models
Minchol Kwon, Minuk Ma, Trung X. Pham, Junyeong Kim
ICPR (14)5
2026 The Pragmatic Persona: Discovering LLM Persona Through Bridging Inference
Jisoo Yang, Jongwon Ryu, Minuk Ma, Trung X. Pham, Junyeong Kim
ICPR (10)5
2025 Learning to See through Sound: From VggCaps to Multi2Cap for Richer Automated Audio Captioning
abstract
Automated Audio Captioning (AAC) aims to generate natural language descriptions of audio content, enabling machines to interpret and communicate complex acoustic scenes.However, current AAC datasets often suffer from short and simplistic captions, limiting model expressiveness and semantic depth.To address this, we introduce VggCaps, a new multimodal dataset that pairs audio with corresponding video and leverages large language models (LLMs) to generate rich, descriptive captions.VggCaps significantly outperforms existing benchmarks in caption length, lexical diversity, and human-rated quality.Furthermore, we propose Multi2Cap, a novel AAC framework that learns audio-visual representations through a AV-grounding module during pre-training and reconstructs visual semantics using audio alone at inference.This enables visually grounded captioning in audio-only scenarios.Experimental results on Clotho and AudioCaps demonstrate that Multi2Cap achieves state-of-the-art performance across multiple metrics, validating the effectiveness of cross-modal supervision and LLM-based generation in advancing AAC.
Sangyeon Cho, Jinkwon Hwang, Jaehoon Go, Minuk Ma, Sunjae Yoon, Junyeong Kim
EMNLP7
2024 ConCSE: Unified Contrastive Learning and Augmentation for Code-Switched Embeddings
Jangyeong Jeon, Sangyeon Cho, Minuk Ma, Junyeong Kim
ICPR (20)4
2024 Scalable SoftGroup for 3D Instance Segmentation on Point Clouds
abstract
This paper considers a network referred to as SoftGroup for accurate and scalable 3D instance segmentation. Existing state-of-the-art methods produce hard semantic predictions followed by grouping instance segmentation results. Unfortunately, errors stemming from hard decisions propagate into the grouping, resulting in poor overlap between predicted instances and ground truth and substantial false positives. To address the abovementioned problems, SoftGroup allows each point to be associated with multiple classes to mitigate the uncertainty stemming from semantic prediction. It also suppresses false positive instances by learning to categorize them as background. Regarding scalability, the existing fast methods require computational time on the order of tens of seconds on large-scale scenes, which is unsatisfactory and far from applicable for real-time. Our finding is that the$k$-Nearest Neighbor ($k$-NN) module, which serves as the prerequisite of grouping, introduces a computational bottleneck. SoftGroup is extended to resolve this computational bottleneck, referred to as SoftGroup++. The proposed SoftGroup++ reduces time complexity with octree$k$-NN and reduces search space with class-aware pyramid scaling and late devoxelization. Experimental results on various indoor and outdoor datasets demonstrate the efficacy and generality of the proposed SoftGroup and SoftGroup++. Their performances surpass the best-performing baseline by a large margin (6%$\sim$16%) in terms of AP$_{50}$. On datasets with large-scale scenes, SoftGroup++ achieves a 6× speed boost on average compared to SoftGroup. Furthermore, SoftGroup can be extended to perform object detection and panoptic segmentation with nontrivial improvements over existing methods.
Thang Vu, Kookhoi Kim, Tung Minh Luu, Junyeong Kim, Chang Dong Yoo
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Counterfactual Two-Stage Debiasing For Video Corpus Moment Retrieval
abstract
Video Corpus Moment Retrieval aims to select a temporal video moment pertinent to a given language query from a large video corpus. Existing systems are prone to rely on a retrieval bias as a shortcut, which hinders the systems from accurately learning vision-language association. The retrieval bias is spurious correlations between query and scene. For a given query, systems tend to retrieve incorrectly correlated scenes due to biased annotations that have predominant binding in a dataset. To this end, we present a Counterfactual Two-stage Debiasing Learning (CTDL), which incorporates a counterfactual bias network that intentionally learns the retrieval bias by providing a shortcut to learn the spurious correlation between keyword and scene, and performs two-stage debiasing learning that mitigates the bias via contrasting factual retrievals with counterfactually biased retrievals. Extensive experiments show the effectiveness of CTDL paradigm.
Sunjae Yoon, Ji Woo Hong, SooHwan Eom, Hee Suk Yoon, Eunseop Yoon, Daehyeok Kim, Junyeong Kim, Chanwoo Kim 0001, Chang Dong Yoo
ICASSP7
2022 Selective Query-Guided Debiasing for Video Corpus Moment Retrieval
Sunjae Yoon, Ji Woo Hong, Eunseop Yoon, Dahyun Kim 0002, Junyeong Kim, Hee Suk Yoon, Chang Dong Yoo
ECCV (36)5
2022 Information-Theoretic Text Hallucination Reduction for Video-grounded Dialogue
abstract
Video-grounded Dialogue (VGD) aims to decode an answer sentence to a question regarding a given video and dialogue context. Despite the recent success of multi-modal reasoning to generate answer sentences, existing dialogue systems still suffer from a text hallucination problem, which denotes indiscriminate text-copying from input texts without an understanding of the question. This is due to learning spurious correlations from the fact that answer sentences in the dataset usually include the words of input texts, thus the VGD system excessively relies on copying words from input texts by hoping those words to overlap with ground-truth texts. Hence, we design Text Hallucination Mitigating (THAM) framework, which incorporates Text Hallucination Regularization (THR) loss derived from the proposed information-theoretic text hallucination measurement approach. Applying THAM with current dialogue systems validates the effectiveness on VGD benchmarks (i.e., AVSD@DSTC7 and AVSD@DSTC8) and shows enhanced interpretability.
Sunjae Yoon, Eunseop Yoon, Hee Suk Yoon, Junyeong Kim, Chang Dong Yoo
EMNLP4
2021 Structured Co-reference Graph Attention for Video-grounded Dialogue
abstract
A video-grounded dialogue system referred to as the Structured Co-reference Graph Attention (SCGA) is presented for decoding the answer sequence to a question regarding a given video while keeping track of the dialogue context. Although recent efforts have made great strides in improving the quality of the response, performance is still far from satisfactory. The two main challenging issues are as follows: (1) how to deduce co-reference among multiple modalities and (2) how to reason on the rich underlying semantic structure of video with complex spatial and temporal dynamics. To this end, SCGA is based on (1) Structured Co-reference Resolver that performs dereferencing via building a structured graph over multiple modalities, (2) Spatio-temporal Video Reasoner that captures local-to-global dynamics of video via gradually neighboring graph attention. SCGA makes use of pointer network to dynamically replicate parts of the question for decoding the answer sequence. The validity of the proposed SCGA is demonstrated on AVSD@DSTC7 and AVSD@DSTC8 datasets, a challenging video-grounded dialogue benchmarks, and TVQA dataset, a large-scale videoQA benchmark. Our empirical results show that SCGA outperforms other state-of-the-art dialogue systems on both benchmarks, while extensive ablation study and qualitative analysis reveal performance gain and improved interpretability.
Junyeong Kim, Sunjae Yoon, Dahyun Kim 0002, Chang Dong Yoo
AAAI1
2021 Weakly-Supervised Moment Retrieval Network for Video Corpus Moment Retrieval
abstract
This paper proposes Weakly-supervised Moment Retrieval Network (WMRN) for Video Corpus Moment Retrieval (VCMR), which retrieves pertinent temporal moments related to natural language query in a large video corpus. Previous methods for VCMR require full supervision of temporal boundary information for training, which involves a labor-intensive process of annotating the boundaries in a large number of videos. To leverage this, the proposed WMRN performs VCMR in a weakly-supervised manner, where WMRN is learned without ground-truth labels but only with video and text queries. For weakly-supervised VCMR, WMRN addresses the following two limitations of prior methods: (1) Blurry attention over video features due to redundant video candidate proposals generation, (2) Insufficient learning due to weak supervision only with video-query pairs. To this end, WMRN is based on (1) Text Guided Proposal Generation (TGPG) that effectively generates text guided multi-scale video proposals in the prospective region related to query, and (2) Hard Negative Proposal Sampling (HNPS) that enhances video-language alignment via extracting negative video proposals in positive video sample for contrastive learning. Experimental results show that WMRN achieves state-of-the-art performance on TVR and DiDeMo benchmarks in the weakly-supervised setting. To validate the attainments of proposed components of WMRN, comprehensive ablation studies and qualitative analysis are conducted.
Sunjae Yoon, Dahyun Kim 0002, Ji Woo Hong, Junyeong Kim, Kookhoi Kim, Chang Dong Yoo
ICIP4
2020 Modality Shifting Attention Network for Multi-Modal Video Question Answering
abstract
This paper considers a network referred to as Modality Shifting Attention Network (MSAN) for Multimodal Video Question Answering (MVQA) task. MSAN decomposes the task into two sub-tasks: (1) localization of temporal moment relevant to the question, and (2) accurate prediction of the answer based on the localized moment. The modality required for temporal localization may be different from that for answer prediction, and this ability to shift modality is essential for performing the task. To this end, MSAN is based on (1) the moment proposal network (MPN) that attempts to locate the most appropriate temporal moment from each of the modalities, and also on (2) the heterogeneous reasoning network (HRN) that predicts the answer using an attention mechanism on both modalities. MSAN is able to place importance weight on the two modalities for each sub-task using a component referred to as Modality Importance Modulation (MIM). Experimental results show that MSAN outperforms previous state-of-the-art by achieving 71.13\% test accuracy on TVQA benchmark dataset. Extensive ablation studies and qualitative analysis are conducted to validate various components of the network.
Junyeong Kim, Minuk Ma, Trung X. Pham, Kyungsu Kim 0003, Chang Dong Yoo
CVPR1
2020 VLANet: Video-Language Alignment Network for Weakly-Supervised Video Moment Retrieval
Minuk Ma, Sunjae Yoon, Junyeong Kim, Youngjoon Lee, Sunghun Kang, Chang Dong Yoo
ECCV (28)3
2019 Progressive Attention Memory Network for Movie Story Question Answering
abstract
This paper proposes the progressive attention memory network (PAMN) for movie story question answering (QA). Movie story QA is challenging compared to VQA in two aspects: (1) pinpointing the temporal parts relevant to answer the question is difficult as the movies are typically longer than an hour, (2) it has both video and subtitle where different questions require different modality to infer the answer. To overcome these challenges, PAMN involves three main features: (1) progressive attention mechanism that utilizes cues from both question and answer to progressively prune out irrelevant temporal parts in memory, (2) dynamic modality fusion that adaptively determines the contribution of each modality for answering the current question, and (3) belief correction answering scheme that successively corrects the prediction score on each candidate answer. Experiments on publicly available benchmark datasets, MovieQA and TVQA, demonstrate that each feature contributes to our movie story QA architecture, PAMN, and improves performance to achieve the state-of-the-art result. Qualitative analysis by visualizing the inference mechanism of PAMN is also provided.
Junyeong Kim, Minuk Ma, Kyungsu Kim 0003, Chang Dong Yoo
CVPR1
2019 Gaining Extra Supervision via Multi-task learning for Multi-Modal Video Question Answering
abstract
This paper proposes a method to gain extra supervision via multi-task learning for multi-modal video question answering. Multi-modal video question answering is an important task that aims at the joint understanding of vision and language. However, establishing large scale dataset for multi-modal video question answering is expensive and the existing benchmarks are relatively small to provide sufficient supervision. To overcome this challenge, this paper proposes a multi-task learning method which is composed of three main components: (1) multi-modal video question answering network that answers the question based on the both video and subtitle feature, (2) temporal retrieval network that predicts the time in the video clip where the question was generated from and (3) modality alignment network that solves metric learning problem to find correct association of video and subtitle modalities. By simultaneously solving related auxiliary tasks with hierarchically shared intermediate layers, the extra synergistic supervisions are provided. Motivated by curriculum learning, multi-task ratio scheduling is proposed to learn easier task earlier to set inductive bias at the beginning of the training. The experiments on publicly available dataset TVQA shows state-of-the-art results, and ablation studies are conducted to prove the statistical validity.
Junyeong Kim, Minuk Ma, Kyungsu Kim 0003, Chang Dong Yoo
IJCNN1
2018 Pivot Correlational Neural Network for Multimodal Video Categorization
Sunghun Kang, Junyeong Kim, Hyunsoo Choi, Chang Dong Yoo
ECCV (14)2
2018 Action Recognition: First-and Second-Order 3D Feature in Bi-Directional Attention Network
abstract
This paper considers a 3D convolutional neural network (CNN) that learns spatial and temporal regions of higher importance through a bi-direction long short-term memory (bi-LSTM) attention for action recognition. First- and second-order differences of spatially most relevant C3D features (sp-m-C3D) are obtained, and the concatenation of the two differences with the sp-m-C3D is used to generate a temporal attention on the sp-m-C3D using a bi-LSTM. Temporally most relevant sp-m-C3D features are fed into another bi-LSTM for action recognition. Essentially, the network learns spatial and temporal regions of high importance for action recognition. We evaluate the network on two public action recognition datasets: UCF-101 (YouTube Action) and HMDB51. The proposed network performs better compared to other state-of-the-art networks.
Oh Chul Kwon, Junyeong Kim, Chang Dong Yoo
ICIP2
2017 Deep partial person re-identification via attention model
abstract
This paper considers a novel algorithm referred to as deep partial person re-identification (DPPR) for partial person re-identification where only a part of a person is observed and full body images are available for identification. The DPPR is based on an end-to-end deep model which make use of convolutional neural network (CNN), RoI Pooling layer and attention model. The RoI Pooling layer enables the extraction of feature vector corresponding to predefined part of input image. The attention model selects a subset of CNN feature vectors. For qualitative evaluation of proposed model, data from CUHK03 are randomly cropped in constructing p-CUHK03. Experimental results show that DPPR outperforms our baseline model on p-CUHK03.
Junyeong Kim, Chang Dong Yoo
ICIP1
2011 Experimental Feasibility Analysis of Primary-Shadow Replication Scheme for I/O Tansmission Fault-Tolerance in Auto-Pilot Program of Small Scale UAV
abstract
This paper treats the Primary-Shadow Replication Scheme to embody fault-tolerant capability in Operation Flight Program (OFP) of small Unmanned Aerial Vehicles (UAV). The recent increase in UAV applications to various autonomous missions demands a highly reliable and safe OFP to cope with unexpected system faults. This paper proposes a modified application of Primary-Shadow TMO's Replication (PSTR)[2]mechanism for quick detection and rectification of system failure with minimum intervention of human pilot. For this purpose the PSTR method is integrated into the Hardware-In-the-Loop Simulation (HILS) environment with a UAV model. Various failure modes in such as receiving UAV's sensor data and sending calculated data to UAV's actuator are simulated and tested to show the enhanced fault-tolerance nature of the OFP. The test results show that 96% of injected faults were successfully detected and recovered, and 94% of shadow OFP was successfully activated within given deadline.
Junyeong Kim, Doo-Hyun Kim
ISADS1
2011 Experimental Analysis of Primary-Shadow Replication Scheme for Fault-Tolerant Operational Flight Program of Small Scale UAV
abstract
This paper proposes to use a time-driven fault-tolerant mechanism motivated from Primary-Shadow TMO's Replication (PSTR)[7, 8] scheme for embodying fault-tolerant capability in Operation Flight Program (OFP) of small Unmanned Aerial Vehicles(UAV). The advantage of the time-driven fault-tolerant mechanism is considered as quick detection and rectification of system failure within minimum period. For the feasibility test, a Hardware-In-the-Loop Simulation (HILS) environment containing dynamics model of a small scaled unmanned helicopter has been developed and integrated with primary and shadow FCCs through RS-232 duplicators and switchers. Various failures and deadline violations in receiving data from sensors, calculating control logics and sending control data to actuators were simulated and tested within the HILS. This paper explains the time-driven fault-tolerant mechanism and experimental environments in details, and illustrated the results of various experiments to convince the practical applicability of the proposed mechanism.
Junyeong Kim, Doo-Hyun Kim
ISORC1