EDBT 2026 Demo / reviewers in the wild / expert
Jae-Won Cho
dblp:15/2423
· DBLP profile ↗
19ranked-venue papers
6as first author
13since 2021 · last 2025
0000-0001-9979-8929ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 8 since 2021Security and privacy · 3 · 2 first-authorSystems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Causal Disentanglement of Treatment Effects in Single-Cell RNA Sequencing Through Counterfactual Inference
Shaokun An, Jae-Won Cho, Jiankang Xiong, Martin Hemberg |
RECOMB | 2 |
| 2024 | Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic CompositionalityabstractIn this paper, we propose a new method to enhance compositional understanding in pretrained vision and language models (VLMs) without sacrificing performance in zero-shot multi-modal tasks.Traditional fine-tuning approaches often improve compositional reasoning at the cost of degrading multi-modal capabilities, primarily due to the use of global hard negative (HN) loss, which contrasts global representations of images and texts.This global HN loss pushes HN texts that are highly similar to the original ones, damaging the model's multi-modal representations.To overcome this limitation, we propose Fine-grained Selective Calibrated CLIP (FSC-CLIP), which integrates local hard negative loss and selective calibrated regularization.These innovations provide fine-grained negative supervision while preserving the model's representational integrity.Our extensive evaluations across diverse benchmarks for both compositionality and multi-modal tasks show that FSC-CLIP not only achieves compositionality on par with state-of-the-art models but also retains strong multi-modal capabilities. Youngtaek Oh, Jae-Won Cho, Dong-Jin Kim 0003, In-So Kweon, Junmo Kim 0002 |
EMNLP | 2 |
| 2024 | Let Me Finish My Sentence: Video Temporal Grounding with Holistic Text UnderstandingabstractVideo Temporal Grounding (VTG) aims to identify visual frames in a video clip that match text queries. Recent studies in VTG employ cross-attention to correlate visual frames and text queries as individual token sequences. However, these approaches overlook a crucial aspect of the problem: a holistic understanding of the query sentence. A model may capture correlations between individual word tokens and arbitrary visual frames while possibly missing out on the global meaning. To address this, we introduce two primary contributions: (1) a visual frame-level gate mechanism that incorporates holistic textual information, (2) cross-modal alignment loss to learn the fine-grained correlation between query and relevant frames. As a result, we regularize the effect of individual word tokens and suppress irrelevant visual frames. We demonstrate that our method outperforms state-of-the-art approaches in VTG benchmarks, indicating that holistic text understanding guides the model to focus on the semantically important parts within the video. Jongbhin Woo, Hyeonggon Ryu, Youngjoon Jang 0001, Jae-Won Cho, Joon Son Chung |
ACM Multimedia | 4 |
| 2023 | Generative Bias for Robust Visual Question AnsweringabstractThe task of Visual Question Answering (VQA) is known to be plagued by the issue of VQA models exploiting biases within the dataset to make its final prediction. Various previous ensemble based debiasing methods have been proposed where an additional model is purposefully trained to be biased in order to train a robust target model. However, these methods compute the bias for a model simply from the label statistics of the training data or from single modal branches. In this work, in order to better learn the bias a target VQA model suffers from, we propose a generative method to train the bias model directly from the target model, called GenB. In particular, GenB employs a generative network to learn the bias in the target model through a combination of the adversarial objective and knowledge distillation. We then debias our target model with GenB as a bias model, and show through extensive experiments the effects of our method on various VQA bias datasets including VQA-CP2, VQA-CP1, GQA-OOD, and VQA-CE, and show state-of-the-art results with the LXMERT architecture on VQA-CP2. Jae-Won Cho, Dong-Jin Kim 0003, Hyeonggon Ryu, In-So Kweon |
CVPR | 1 |
| 2023 | Self-Sufficient Framework for Continuous Sign Language RecognitionabstractThe goal of this work is to develop self-sufficient framework for Continuous Sign Language Recognition (CSLR) that addresses key issues of sign language recognition. These include the need for complex multi-scale features such as hands, face, and mouth for understanding, and absence of frame-level annotations. To this end, we propose (1) Divide and Focus Convolution (DFConv) which extracts both manual and non-manual features without the need for additional networks or annotations, and (2) Dense Pseudo-Label Refinement (DPLR) which propagates non-spiky frame-level pseudo-labels by combining the ground truth gloss sequence labels with the predicted sequence. We demonstrate that our model achieves state-of-the-art performance among RGB-based methods on large-scale CSLR benchmarks, PHOENIX-2014 and PHOENIX-2014-T, while showing comparable results with better efficiency when compared to other approaches that use multi-modality or extra annotations. Youngjoon Jang 0001, Youngtaek Oh, Jae-Won Cho, Myungchul Kim 0002, Dong-Jin Kim 0003, In-So Kweon, Joon Son Chung |
ICASSP | 3 |
| 2023 | Empirical study on using adapters for debiased Visual Question Answering
Jae-Won Cho, Dawit Mureja Argaw, Youngtaek Oh, Dong-Jin Kim 0003, In-So Kweon |
Comput. Vis. Image Underst. | 1 |
| 2023 | MCDAL: Maximum Classifier Discrepancy for Active LearningabstractRecent state-of-the-art active learning methods have mostly leveraged generative adversarial networks (GANs) for sample acquisition; however, GAN is usually known to suffer from instability and sensitivity to hyperparameters. In contrast to these methods, in this article, we propose a novel active learning framework that we call Maximum Classifier Discrepancy for Active Learning (MCDAL) that takes the prediction discrepancies between multiple classifiers. In particular, we utilize two auxiliary classification layers that learn tighter decision boundaries by maximizing the discrepancies among them. Intuitively, the discrepancies in the auxiliary classification layers' predictions indicate the uncertainty in the prediction. In this regard, we propose a novel method to leverage the classifier discrepancies for the acquisition function for active learning. We also provide an interpretation of our idea in relation to existing GAN-based active learning methods and domain adaptation frameworks. Moreover, we empirically demonstrate the utility of our approach where the performance of our approach exceeds the state-of-the-art methods on several image classification and semantic segmentation datasets in active learning setups. Jae-Won Cho, Dong-Jin Kim 0003, Yunjae Jung, In-So Kweon |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Signing Outside the Studio: Benchmarking Background Robustness for Continuous Sign Language Recognition
Youngjoon Jang 0001, Youngtaek Oh, Jae-Won Cho, Dong-Jin Kim 0003, Joon Son Chung, In-So Kweon |
BMVC | 3 |
| 2022 | Investigating Top-k White-Box and Transferable Black-box AttackabstractExisting works have identified the limitation of top-1 attack success rate (ASR) as a metric to evaluate the attack strength but exclusively investigated it in the white-box setting, while our work extends it to a more practical black-box setting: transferable attack. It is widely reported that stronger I-FGSM transfers worse than simple FGSM, leading to a popular belief that transferability is at odds with the white-box attack strength. Our work challenges this belief with empirical finding that stronger attack actually transfers better for the general top-k ASR indicated by the interest class rank (ICR) after attack. For increasing the attack strength, with an intuitive analysis on the logit gradient from the geometric perspective, we identify that the weakness of the commonly used losses lie in prioritizing the speed to fool the network instead of maximizing its strength. To this end, we propose a new normalized CE loss that guides the logit to be updated in the direction of implicitly maximizing its rank distance from the ground-truth class. Extensive results in various settings have verified that our proposed new loss is simple yet effective for top-k attack. Code is available at: https://bit.ly/3uCiomP Chaoning Zhang, Philipp Benz, Adil Karjauv, Jae-Won Cho, Kang Zhang 0008, In-So Kweon |
CVPR | 4 |
| 2021 | Optical Flow Estimation from a Single Motion-blurred ImageabstractIn most of computer vision applications, motion blur is regarded as an undesirable artifact. However, it has been shown that motion blur in an image may have practical interests in fundamental computer vision problems. In this work, we propose a novel framework to estimate optical flow from a single motion-blurred image in an end-to-end manner. We design our network with transformer networks to learn globally and locally varying motions from encoded features of a motion-blurred input, and decode left and right frame features without explicit frame supervision. A flow estimator network is then used to estimate optical flow from the decoded features in a coarse-to-fine manner. We qualitatively and quantitatively evaluate our model through a large set of experiments on synthetic and real motion-blur datasets. We also provide in-depth analysis of our model in connection with related approaches to highlight the effectiveness and favorability of our approach. Furthermore, we showcase the applicability of the flow estimated by our method on deblurring and moving object segmentation tasks. Dawit Mureja Argaw, Junsik Kim 0001, François Rameau, Jae-Won Cho, In-So Kweon |
AAAI | 4 |
| 2021 | Single-Modal Entropy based Active Learning for Visual Question Answering
Dong-Jin Kim 0003, Jae-Won Cho, Jinsoo Choi, Yunjae Jung, In-So Kweon |
BMVC | 2 |
| 2021 | LabOR: Labeling Only if Required for Domain Adaptive Semantic SegmentationabstractUnsupervised Domain Adaptation (UDA) for semantic segmentation has been actively studied to mitigate the domain gap between label-rich source data and unlabeled target data. Despite these efforts, UDA still has a long way to go to reach the fully supervised performance. To this end, we propose a Labeling Only if Required strategy, LabOR, where we introduce a human-in-the-loop approach to adaptively give scarce labels to points that a UDA model is uncertain about. In order to find the uncertain points, we generate an inconsistency mask using the proposed adaptive pixel selector and we label these segment-based regions to achieve near supervised performance with only a small fraction (about 2.2%) ground truth points, which we call "Segment based Pixel-Labeling (SPL)." To further reduce the efforts of the human annotator, we also propose "Point based Pixel-Labeling (PPL)," which finds the most representative points for labeling within the generated inconsistency mask. This reduces efforts from 2.2% segment label → 40 points label while minimizing performance degradation. Through extensive experimentation, we show the advantages of this new framework for domain adaptive semantic segmentation while minimizing human labor costs. Inkyu Shin, Dong-Jin Kim 0003, Jae-Won Cho, Sanghyun Woo, Kwanyong Park, In-So Kweon |
ICCV | 3 |
| 2021 | Correlate-and-Excite: Real-Time Stereo Matching via Guided Cost Volume ExcitationabstractVolumetric deep learning approach towards stereo matching aggregates a cost volume computed from input left and right images using 3D convolutions. Recent works showed that utilization of extracted image features and a spatially varying cost volume aggregation complements 3D convolutions. However, existing methods with spatially varying operations are complex, cost considerable computation time, and cause memory consumption to increase. In this work, we construct Guided Cost volume Excitation (GCE) and show that simple channel excitation of cost volume guided by image can improve performance considerably. Moreover, we propose a novel method of using top-k selection prior to soft-argmin disparity regression for computing the final disparity estimate. Combining our novel contributions, we present an end-to-end network that we call Correlate-and-Excite (CoEx). Extensive experiments of our model on the SceneFlow, KITTI 2012, and KITTI 2015 datasets demonstrate the effectiveness and efficiency of our model and show that our model outperforms other speed-based algorithms while also being competitive to other state-of-the-art algorithms. Codes will be made available at https://github.com/antabangun/coex. Antyanta Bangunharcana, Jae-Won Cho, Seokju Lee, In-So Kweon, Kyung-Soo Kim 0001, Soohyun Kim 0001 |
IROS | 2 |
| 2006 | A Robust Blind Water Marking for 3D Meshes Using Distribution of Scale Coefficients in Irregular Wavelet AnalysisabstractWe propose a new method for watermarking of 3D surface mesh. After irregular wavelet analysis, the watermark is embedded in the scale coefficients. We modify the mean of the distribution of the vertex norm of the approximation mesh in order to obtain blind scheme which does not require the original mesh for detection. Experimental results show the effectiveness of the proposed algorithm both in terms of invisibility and robustness against simplification and progressive lossy compression Jae-Won Cho, Ho-Youl Jung, Rémy Prost |
ICASSP (5) | 2 |
| 2006 | 3-D Dynamic Mesh Compression using Wavelet-Based Multiresolution AnalysisabstractIn this paper, we present a wavelet-based progressive compression method for 3-D dynamic meshes. Our method exploits the spatial and temporal redundancy. We encode the geometry of base mesh, the wavelet coefficients and the connectivity of each resolution level in order to reduce the spatial redundancy of intra meshes. For inter mesh coding, we encode the differences of geometry of base meshes and of their wavelet coefficients between adjacent frames to reduce the temporal redundancy. Our proposal is based on the wavelet-based multiresolution analysis which uses a perfect reconstruction filter bank and therefore it enables not only progressive representation but also lossless compression. The simulation results demonstrate that the proposed method is applicable to lossy and lossless compression of 3-D dynamic meshes. Jae-Won Cho, Sébastien Valette, Ho-Youl Jung, Rémy Prost |
ICIP | 1 |
| 2006 | Wavelet Analysis Based Blind Watermarking for 3-D Surface Meshes
Jae-Won Cho, Rémy Prost, Ho-Youl Jung |
IWDW | 2 |
| 2004 | Audio watermarking in sub-band signals using multiple echo kernels
In-Jung Oh, Hyun-Yeol Chung, Jae-Won Cho, Ho-Youl Jung, Rémy Prost |
INTERSPEECH | 3 |
| 2004 | Robust Watermarking on Polygonal Meshes Using Distribution of Vertex Norms
Jae-Won Cho, Rémy Prost, Hyun-Yeol Chung, Ho-Youl Jung |
IWDW | 1 |
| 2003 | Echo Watermarking in Sub-band Domain
Jae-Won Cho, Ha-Joong Park, Young Huh, Hyun-Yeol Chung, Ho-Youl Jung |
IWDW | 1 |