Minuk Ma

dblp:239/6086 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0003-0416-8479ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Video DETOX: Purifying Noisy Relevance Signals for Diverse and Long-Form Video Understanding
Sungjin Han, Minuk Ma, Trung X. Pham, Junyeong Kim
ICPR (14)2
2026 Blocking Visual Leakage: Visually-Agnostic Text Decomposition for Composed Video Retrieval
Jinkwon Hwang, Minuk Ma, Trung X. Pham, Junyeong Kim
ICPR (14)2
2026 Beyond Co-existence: Measuring Attribute Binding Hallucinations in Audio-Language Models
Minchol Kwon, Minuk Ma, Trung X. Pham, Junyeong Kim
ICPR (14)3
2026 The Pragmatic Persona: Discovering LLM Persona Through Bridging Inference
Jisoo Yang, Jongwon Ryu, Minuk Ma, Trung X. Pham, Junyeong Kim
ICPR (10)3
2025 Learning to See through Sound: From VggCaps to Multi2Cap for Richer Automated Audio Captioning
abstract
Automated Audio Captioning (AAC) aims to generate natural language descriptions of audio content, enabling machines to interpret and communicate complex acoustic scenes.However, current AAC datasets often suffer from short and simplistic captions, limiting model expressiveness and semantic depth.To address this, we introduce VggCaps, a new multimodal dataset that pairs audio with corresponding video and leverages large language models (LLMs) to generate rich, descriptive captions.VggCaps significantly outperforms existing benchmarks in caption length, lexical diversity, and human-rated quality.Furthermore, we propose Multi2Cap, a novel AAC framework that learns audio-visual representations through a AV-grounding module during pre-training and reconstructs visual semantics using audio alone at inference.This enables visually grounded captioning in audio-only scenarios.Experimental results on Clotho and AudioCaps demonstrate that Multi2Cap achieves state-of-the-art performance across multiple metrics, validating the effectiveness of cross-modal supervision and LLM-based generation in advancing AAC.
Sangyeon Cho, Jinkwon Hwang, Jaehoon Go, Minuk Ma, Sunjae Yoon, Junyeong Kim
EMNLP5
2024 ConCSE: Unified Contrastive Learning and Augmentation for Code-Switched Embeddings
Jangyeong Jeon, Sangyeon Cho, Minuk Ma, Junyeong Kim
ICPR (20)3
2020 Modality Shifting Attention Network for Multi-Modal Video Question Answering
abstract
This paper considers a network referred to as Modality Shifting Attention Network (MSAN) for Multimodal Video Question Answering (MVQA) task. MSAN decomposes the task into two sub-tasks: (1) localization of temporal moment relevant to the question, and (2) accurate prediction of the answer based on the localized moment. The modality required for temporal localization may be different from that for answer prediction, and this ability to shift modality is essential for performing the task. To this end, MSAN is based on (1) the moment proposal network (MPN) that attempts to locate the most appropriate temporal moment from each of the modalities, and also on (2) the heterogeneous reasoning network (HRN) that predicts the answer using an attention mechanism on both modalities. MSAN is able to place importance weight on the two modalities for each sub-task using a component referred to as Modality Importance Modulation (MIM). Experimental results show that MSAN outperforms previous state-of-the-art by achieving 71.13\% test accuracy on TVQA benchmark dataset. Extensive ablation studies and qualitative analysis are conducted to validate various components of the network.
Junyeong Kim, Minuk Ma, Trung X. Pham, Kyungsu Kim 0003, Chang Dong Yoo
CVPR2
2020 VLANet: Video-Language Alignment Network for Weakly-Supervised Video Moment Retrieval
Minuk Ma, Sunjae Yoon, Junyeong Kim, Youngjoon Lee, Sunghun Kang, Chang Dong Yoo
ECCV (28)1
2019 Progressive Attention Memory Network for Movie Story Question Answering
abstract
This paper proposes the progressive attention memory network (PAMN) for movie story question answering (QA). Movie story QA is challenging compared to VQA in two aspects: (1) pinpointing the temporal parts relevant to answer the question is difficult as the movies are typically longer than an hour, (2) it has both video and subtitle where different questions require different modality to infer the answer. To overcome these challenges, PAMN involves three main features: (1) progressive attention mechanism that utilizes cues from both question and answer to progressively prune out irrelevant temporal parts in memory, (2) dynamic modality fusion that adaptively determines the contribution of each modality for answering the current question, and (3) belief correction answering scheme that successively corrects the prediction score on each candidate answer. Experiments on publicly available benchmark datasets, MovieQA and TVQA, demonstrate that each feature contributes to our movie story QA architecture, PAMN, and improves performance to achieve the state-of-the-art result. Qualitative analysis by visualizing the inference mechanism of PAMN is also provided.
Junyeong Kim, Minuk Ma, Kyungsu Kim 0003, Chang Dong Yoo
CVPR2
2019 Gaining Extra Supervision via Multi-task learning for Multi-Modal Video Question Answering
abstract
This paper proposes a method to gain extra supervision via multi-task learning for multi-modal video question answering. Multi-modal video question answering is an important task that aims at the joint understanding of vision and language. However, establishing large scale dataset for multi-modal video question answering is expensive and the existing benchmarks are relatively small to provide sufficient supervision. To overcome this challenge, this paper proposes a multi-task learning method which is composed of three main components: (1) multi-modal video question answering network that answers the question based on the both video and subtitle feature, (2) temporal retrieval network that predicts the time in the video clip where the question was generated from and (3) modality alignment network that solves metric learning problem to find correct association of video and subtitle modalities. By simultaneously solving related auxiliary tasks with hierarchically shared intermediate layers, the extra synergistic supervisions are provided. Motivated by curriculum learning, multi-task ratio scheduling is proposed to learn easier task earlier to set inductive bias at the beginning of the training. The experiments on publicly available dataset TVQA shows state-of-the-art results, and ablation studies are conducted to prove the statistical validity.
Junyeong Kim, Minuk Ma, Kyungsu Kim 0003, Chang Dong Yoo
IJCNN2
2018 ImaGAN: Unsupervised Training of Conditional Joint CycleGAN for Transferring Style with Core Structures in Content Preserved
Kangmin Bae, Minuk Ma, Hyunjun Jang, Minjeong Ju, Hyoungwoo Park, Chang Dong Yoo
ACCV (2)2