EDBT 2026 Demo / reviewers in the wild / expert
Jongwon Choi 0002
dblp:126/0675-2
· DBLP profile ↗
29ranked-venue papers
7as first author
15since 2021 · last 2026
0000-0001-9753-8760ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 7 first-author · 8 since 2021Artificial intelligence and machine learning · 20 · 5 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Flexible Knowledge Distillation for Class-Incremental Learning via Structural Knowledge Transfer
Seungmo Seo, Jongsu Youn, Jaehyung Bae, Jongwon Choi 0002 |
ICPR (2) | 4 |
| 2026 | 3D-aware virtual try-on using only 2D inputs
Jaeyoon Lee, Hojoon Jung, Jongwon Choi 0002 |
Comput. Vis. Image Underst. | 3 |
| 2026 | Domain-generalizable face anti-spoofing with patch-based multi-tasking and artifact pattern conversion
Seungjin Jung, Yonghyun Jeong, Minha Kim, Jimin Min, Young Joon Yoo, Jongwon Choi 0002 |
Pattern Recognit. | 6 |
| 2025 | Few-shot Semantic Segmentation with Uncertainty-based Joint PrototypesabstractTo overcome the high cost of data acquisition, few-shot semantic segmentation is studied to increase the training efficiency of limited data, but it fails to detect the narrow objects well. We find that the issue is caused by two main reasons: the enlarged receptive field of the baseline models and the high-proportional noisy labels of the narrow objects. An enlarged receptive field lets the model ignore detailed information that is important for the narrow objects, which can be affected by the same amount of noisy labels more critically than the large objects. To solve the issue, we propose a novel method to improve the performance of narrow objects in few-shot semantic segmentation. First of all, we diversify the size of the receptive field by extracting multiple prototypes from multi-level pyramidal feature maps, which is helpful to consider the detailed features of narrow objects. In addition, during model training, we simultaneously update uncertainty maps that determine the pixel-wise label reliability to detect and ignore noisy labels. We validate the proposed method, which shows impressive enhancement for narrow object segmentation both quantitatively and qualitatively over the prior research. Yumin Lim, Doyoung Park, Naresh Reddy Yarram, Sunjin Kim, Seongho Joe, Youngjune Gwon, Jongwon Choi 0002 |
AVSS | 8 |
| 2025 | Style prompt tuning for bridging visual gaps in autonomous driving
Suyeon Cha, Giyun Choi, Minji Kwak, Jongwon Choi 0002 |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | Generating visual-adaptive audio representation for audio recognition
Jongsu Youn, Seungmo Seo, Sukhyun Kim, Jongwon Choi 0002 |
Pattern Recognit. Lett. | 5 |
| 2024 | Topic-VQ-VAE: Leveraging Latent Codebooks for Flexible Topic-Guided Document GenerationabstractThis paper introduces a novel approach for topic modeling utilizing latent codebooks from Vector-Quantized Variational Auto-Encoder~(VQ-VAE), discretely encapsulating the rich information of the pre-trained embeddings such as the pre-trained language model. From the novel interpretation of the latent codebooks and embeddings as conceptual bag-of-words, we propose a new generative topic model called Topic-VQ-VAE~(TVQ-VAE) which inversely generates the original documents related to the respective latent codebook. The TVQ-VAE can visualize the topics with various generative distributions including the traditional BoW distribution and the autoregressive image generation. Our experimental results on document analysis and image generation demonstrate that TVQ-VAE effectively captures the topic context which reveals the underlying structures of the dataset and supports flexible forms of document generation. Official implementation of the proposed TVQ-VAE is available at https://github.com/clovaai/TVQ-VAE. Young Joon Yoo, Jongwon Choi 0002 |
AAAI | 2 |
| 2024 | Self-supervised scheme for generalizing GAN image detection
Yonghyun Jeong, Pyounggeon Kim, Youngmin Ro, Jongwon Choi 0002 |
Pattern Recognit. Lett. | 5 |
| 2023 | Document Change Detection With Hierarchical Patch ComparisonabstractContract documents can be modified just before signing, after the consensus, with the intention of defrauding the other party, which can have serious consequences for the deal. To prevent the issue, we propose a method to detect document changes between a scanned final document and its original electronic file using image-based comparison. Our method first finds the most appropriate augmentation for various document changes, such as rotations, contrast, ratio, or brightness changes which can occur while scanning documents. Then, we employ a hierarchical search strategy from large patches to small patches in a sliding window manner, which can reduce the computational complexity to compare all the details of the documents using the deep learning model. We built a new dataset of original-scanned document pair for the validation of our method. In the experiments, we show that our method outperforms the previous approaches using segmentation and character recognition models, even when the document suffers from both non-lingual and lingual changes. Doyoung Park, Sunjin Kim, Naresh Reddy Yarram, Seongho Joe, Youngjune Gwon, Jongwon Choi 0002 |
ICIP | 7 |
| 2022 | Tracking Failure Prediction for Siamese Trackers Based on Channel Feature StatisticsabstractFailure prediction has rarely been studied for Siamese trackers due to a lack of meaningful analysis of tracking failing cases. In this paper, we provide a meaningful analysis of tracking failure in Siamese trackers. Our analysis includes the statistics of the channel-wise feature correlation between the exemplar and tracked target patches. We observe that the correlation statistics (max, mean, and std) are highly related to the overlapping ratio between tracked and ground-truth bounding boxes. Based on this observation, we devise a tracking failure prediction model that extracts more plentiful factors than simple statistics. The proposed tracking failure prediction model is validated on most-popular tracking benchmark datasets through extensive experiments. Kyuewang Lee, Hoseok Do, Taegil Ha, Jongwon Choi 0002, Jin Young Choi 0002 |
AVSS | 4 |
| 2022 | Novel-View Synthesis of Human Tourist PhotosabstractWe present a novel framework for performing novel-view synthesis on human tourist photos. Given a tourist photo from a known scene, we reconstruct the photo in 3D space through modeling the human and the background independently. We generate a deep buffer from a novel viewpoint of the reconstruction and utilize a deep network to translate the buffer into a photo-realistic rendering of the novel view. We additionally present a method to relight the renderings, allowing for relighting of both human and background to match either the provided input image or any other. The key contributions of our paper are: 1) a framework for performing novel view synthesis on human tourist photos, 2) an appearance transfer method for relighting of humans to match synthesized backgrounds, and 3) a method for estimating lighting properties from a single human photo. We demonstrate the proposed framework on photos from two different scenes of various tourists. Jonathan Freer, Kwang Moo Yi, Wei Jiang 0034, Jongwon Choi 0002, Hyung Jin Chang |
WACV | 4 |
| 2022 | BiHPF: Bilateral High-Pass Filters for Robust Deepfake DetectionabstractThe advancement in numerous generative models has a two-fold effect: a simple and easy generation of realistic synthesized images, but also an increased risk of malicious abuse of those images. Thus, it is important to develop a generalized detector for synthesized images of any GAN model or object category, including those unseen during the training phase. However, the conventional methods heavily depend on the training settings, which cause a dramatic decline in performance when tested with unknown domains. To resolve the issue and obtain a generalized detection ability, we propose Bilateral High-Pass Filters (BiHPF), which amplify the effect of the frequency-level artifacts that are generally found in the synthesized images of generative models. Also, to find the properties of the general frequency-level artifacts, we develop an additional method to adversarially extract the artifact compression map. Numerous experimental results validate that our method outperforms other state-of-the-art methods, even when tested with unseen domains. Yonghyun Jeong, Seungjai Min, Seongho Joe, Youngjune Gwon, Jongwon Choi 0002 |
WACV | 6 |
| 2022 | Rollback Ensemble With Multiple Local Minima in Fine-Tuning Deep Learning NetworksabstractImage retrieval is a challenging problem that requires learning generalized features enough to identify untrained classes, even with very few classwise training samples. In this article, to obtain generalized features further in learning retrieval data sets, we propose a novel fine-tuning method of pretrained deep networks. In the retrieval task, we discovered a phenomenon in which the loss reduction in fine-tuning deep networks is stagnated, even while weights are largely updated. To escape from the stagnated state, we propose a new fine-tuning strategy to roll back some of the weights to the pretrained values. The rollback scheme is observed to drive the learning path to a gentle basin that provides more generalized features than a sharp basin. In addition, we propose a multihead ensemble structure to create synergy among multiple local minima obtained by our rollback scheme. Experimental results show that the proposed learning method significantly improves generalization performance, achieving state-of-the-art performance on the Inshop and SOP data sets. Youngmin Ro, Jongwon Choi 0002, Byeongho Heo, Jin Young Choi 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | VaB-AL: Incorporating Class Imbalance and Difficulty With Variational Bayes for Active LearningabstractActive Learning for discriminative models has largely been studied with the focus on individual samples, with less emphasis on how classes are distributed or which classes are hard to deal with. In this work, we show that this is harmful. We propose a method based on the Bayes’ rule, that can naturally incorporate class imbalance into the Active Learning framework. We derive that three terms should be considered together when estimating the probability of a classifier making a mistake for a given sample; i) probability of mislabelling a class, ii) likelihood of the data given a predicted class, and iii) the prior probability on the abundance of a predicted class. Implementing these terms requires a generative model and an intractable likelihood estimation. Therefore, we train a Variational Auto Encoder (VAE) for this purpose. To further tie the VAE with the classifier and facilitate VAE training, we use the classifiers’ deep feature representations as input to the VAE. By considering all three probabilities, among them, especially the data imbalance, we can substantially improve the potential of existing methods under limited data budget. We show that our method can be applied to classification tasks on multiple different datasets – including one that is a real-world dataset with heavy data imbalance – significantly outperforming the state of the art. Jongwon Choi 0002, Kwang Moo Yi, Jinho Choo, Byoungjip Kim, Jin-Yeop Chang, Youngjune Gwon, Hyung Jin Chang |
CVPR | 1 |
| 2021 | Motion-aware ensemble of three-mode trackers for unmanned aerial vehicles
Kyuewang Lee, Hyung Jin Chang, Jongwon Choi 0002, Byeongho Heo, Ales Leonardis, Jin Young Choi 0002 |
Mach. Vis. Appl. | 3 |
| 2020 | Visual Domain Adaptation by Consensus-Based Transfer to Intermediate DomainabstractWe describe an unsupervised domain adaptation framework for images by a transform to an abstract intermediate domain and ensemble classifiers seeking a consensus. The intermediate domain can be thought as a latent domain where both the source and target domains can be transferred easily. The proposed framework aligns both domains to the intermediate domain, which greatly improves the adaptation performance when the source and target domains are notably dissimilar. In addition, we propose an ensemble model trained by confusing multiple classifiers and letting them make a consensus alternately to enhance the adaptation performance for ambiguous samples. To estimate the hidden intermediate domain and the unknown labels of the target domain simultaneously, we develop a training algorithm using a double-structured architecture. We validate the proposed framework in hard adaptation scenarios with real-world datasets from simple synthetic domains to complex real-world domains. The proposed algorithm outperforms the previous state-of-the-art algorithms on various environments. Jongwon Choi 0002, Youngjoon Choi, Jin-Yeop Chang, Ilhwan Kwon, Youngjune Gwon, Seungjai Min |
AAAI | 1 |
| 2020 | Associative Variational Auto-Encoder with Distributed Latent Spaces and AssociatorsabstractIn this paper, we propose a novel structure for a multi-modal data association referred to as Associative Variational Auto-Encoder (AVAE). In contrast to the existing models using a shared latent space among modalities, our structure adopts distributed latent spaces for multi-modalities which are connected through cross-modal associators. The proposed structure successfully associates even heterogeneous modality data and easily incorporates the additional modality to the entire network via the associator. Furthermore, in our structure, only a small amount of supervised (paired) data is enough to train associators after training auto-encoders in an unsupervised manner. Through experiments, the effectiveness of the proposed structure is validated on various datasets including visual and auditory data. Byeongju Lee, Jongwon Choi 0002, Haan-Ju Yoo, Jin Young Choi 0002 |
AAAI | 3 |
| 2020 | DoFNet: Depth of Field Difference Learning for Detecting Image Forgery
Yonghyun Jeong, Jongwon Choi 0002, Sehyeon Park, Minki Hong, Changhyun Park, Seungjai Min, Youngjune Gwon |
ACCV (6) | 2 |
| 2019 | Backbone Cannot Be Trained at Once: Rolling Back to Pre-Trained Network for Person Re-IdentificationabstractIn person re-identification (ReID) task, because of its shortage of trainable dataset, it is common to utilize fine-tuning method using a classification network pre-trained on a large dataset. However, it is relatively difficult to sufficiently finetune the low-level layers of the network due to the gradient vanishing problem. In this work, we propose a novel fine-tuning strategy that allows low-level layers to be sufficiently trained by rolling back the weights of high-level layers to their initial pre-trained weights. Our strategy alleviates the problem of gradient vanishing in low-level layers and robustly trains the low-level layers to fit the ReID dataset, thereby increasing the performance of ReID tasks. The improved performance of the proposed strategy is validated via several experiments. Furthermore, without any addons such as pose estimation or segmentation, our strategy exhibits state-of-the-art performance using only vanilla deep convolutional neural network architecture. Youngmin Ro, Jongwon Choi 0002, Byeongho Heo, Jongin Lim 0002, Jin Young Choi 0002 |
AAAI | 2 |
| 2018 | Context-Aware Deep Feature Compression for High-Speed Visual TrackingabstractWe propose a new context-aware correlation filter based tracking framework to achieve both high computational speed and state-of-the-art performance among real-time trackers. The major contribution to the high computational speed lies in the proposed deep feature compression that is achieved by a context-aware scheme utilizing multiple expert auto-encoders; a context in our framework refers to the coarse category of the tracking target according to appearance patterns. In the pre-training phase, one expert auto-encoder is trained per category. In the tracking phase, the best expert auto-encoder is selected for a given target, and only this auto-encoder is used. To achieve high tracking performance with the compressed feature map, we introduce extrinsic denoising processes and a new orthogonality loss term for pre-training and fine-tuning of the expert autoencoders. We validate the proposed context-aware framework through a number of experiments, where our method achieves a comparable performance to state-of-the-art trackers which cannot run in real-time, while running at a significantly fast speed of over 100 fps. Jongwon Choi 0002, Hyung Jin Chang, Tobias Fischer 0001, Sangdoo Yun, Kyuewang Lee, Jiyeoup Jeong, Yiannis Demiris, Jin Young Choi 0002 |
CVPR | 1 |
| 2018 | Selective Ensemble Network for Accurate Crowd Density EstimationabstractThis paper proposes a selective ensemble deep network architecture for crowd density estimation and people counting. In contrast to existing deep network-based methods, the proposed method incorporates two sub-networks for local density estimation: one to learn sparse density regions and one to learn dense density regions. Locally estimated density maps from the two sub-networks are selectively combined in ensemble fashion using a gating network to estimate an initial crowd density map. The initial density map is refined as a high resolution map, using another sub-network that draws on contextual information in the image. In training, a novel adaptive loss scheme is applied to resolve an ambiguity in the crowded region. the proposed scheme improves both density map accuracy and counting accuracy by adjusting the weighting value between density loss and counting loss according to the degree of crowdness and training epochs. Experiments using public datasets confirm that the proposed method outperforms state-of-the-art methods. Through self-evaluation, the effectiveness of each part in the network is also verified. Jiyeoup Jeong, Hawook Jeong, Jongin Lim 0002, Jongwon Choi 0002, Sangdoo Yun, Jin Young Choi 0002 |
ICPR | 4 |
| 2018 | Action-Driven Visual Object Tracking With Deep Reinforcement LearningabstractIn this paper, we propose an efficient visual tracker, which directly captures a bounding box containing the target object in a video by means of sequential actions learned using deep neural networks. The proposed deep neural network to control tracking actions is pretrained using various training video sequences and fine-tuned during actual tracking for online adaptation to a change of target and background. The pretraining is done by utilizing deep reinforcement learning (RL) as well as supervised learning. The use of RL enables even partially labeled data to be successfully utilized for semisupervised learning. Through the evaluation of the object tracking benchmark data set, the proposed tracker is validated to achieve a competitive performance at three times the speed of existing deep network-based trackers. The fast version of the proposed method, which operates in real time on graphics processing unit, outperforms the state-of-the-art real-time trackers with an accuracy improvement of more than 8%. Sangdoo Yun, Jongwon Choi 0002, Young Joon Yoo, Kimin Yun, Jin Young Choi 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | Attentional Correlation Filter Network for Adaptive Visual TrackingabstractWe propose a new tracking framework with an attentional mechanism that chooses a subset of the associated correlation filters for increased robustness and computational efficiency. The subset of filters is adaptively selected by a deep attentional network according to the dynamic properties of the tracking target. Our contributions are manifold, and are summarised as follows: (i) Introducing the Attentional Correlation Filter Network which allows adaptive tracking of dynamic targets. (ii) Utilising an attentional network which shifts the attention to the best candidate modules, as well as predicting the estimated accuracy of currently inactive modules. (iii) Enlarging the variety of correlation filters which cover target drift, blurriness, occlusion, scale changes, and flexible aspect ratio. (iv) Validating the robustness and efficiency of the attentional mechanism for visual tracking through a number of experiments. Our method achieves similar performance to non real-time trackers, and state-of-the-art performance amongst real-time trackers. Jongwon Choi 0002, Hyung Jin Chang, Sangdoo Yun, Tobias Fischer 0001, Yiannis Demiris, Jin Young Choi 0002 |
CVPR | 1 |
| 2017 | Action-Decision Networks for Visual Tracking with Deep Reinforcement LearningabstractThis paper proposes a novel tracker which is controlled by sequentially pursuing actions learned by deep reinforcement learning. In contrast to the existing trackers using deep networks, the proposed tracker is designed to achieve a light computation as well as satisfactory tracking accuracy in both location and scale. The deep network to control actions is pre-trained using various training sequences and fine-tuned during tracking for online adaptation to target and background changes. The pre-training is done by utilizing deep reinforcement learning as well as supervised learning. The use of reinforcement learning enables even partially labeled data to be successfully utilized for semi-supervised learning. Through evaluation of the OTB dataset, the proposed tracker is validated to achieve a competitive performance that is three times faster than state-of-the-art, deep network-based trackers. The fast version of the proposed method, which operates in real-time on GPU, outperforms the state-of-the-art real-time trackers. Sangdoo Yun, Jongwon Choi 0002, Young Joon Yoo, Kimin Yun, Jin Young Choi 0002 |
CVPR | 2 |
| 2016 | Visual Tracking Using Attention-Modulated Disintegration and IntegrationabstractIn this paper, we present a novel attention-modulated visual tracking algorithm that decomposes an object into multiple cognitive units, and trains multiple elementary trackers in order to modulate the distribution of attention according to various feature and kernel types. In the integration stage it recombines the units to memorize and recognize the target object effectively. With respect to the elementary trackers, we present a novel attentional feature-based correlation filter (AtCF) that focuses on distinctive attentional features. The effectiveness of the proposed algorithm is validated through experimental comparison with state-of-theart methods on widely-used tracking benchmark datasets. Jongwon Choi 0002, Hyung Jin Chang, Jiyeoup Jeong, Yiannis Demiris, Jin Young Choi 0002 |
CVPR | 1 |
| 2015 | Patch-based fire detection with online outlier learningabstractFire detection is one of the most interesting issues for surveillance. The existing approaches for the fire detection suffer from a high false positive ratio. To solve the problems, we present a patch-based fire detection algorithm with online outlier learning. In the proposed algorithm, the candidates of fire are obtained in the form of patch, while the classical candidates have been based on pixels or blobs. Because the patches of fire have more distinctive shape than the entire fire, the shape classifier can recognize the candidates correctly from fire-like outliers. In addition, we propose an online outlier learning scheme which handles the irregularity of fire based on the repeatability of shape in time. The proposed algorithm is experimented with new challenging dataset, consisting of 50 positive videos with fire and 44 negative ones with fire-like outliers. By evaluating on the dataset, we validate the performance of our algorithm qualitatively and quantitatively. Jongwon Choi 0002, Jin Young Choi 0002 |
AVSS | 1 |
| 2015 | Robust pan-tilt-zoom tracking via optimization combining motion features and appearance correlationsabstractThis paper proposes a new pan-tilt-zoom (PTZ) tracking method to improve the robustness against occlusions and appearance changes by using motion likelihood map and scale change estimation as well as appearance correlation filter. For this purpose, we introduce a motion likelihood map constructed from motion detection result in addition to the correlation filter. The motion likelihood map is generated by blurring the motion detection result, which shows high probability in the center of target. To combine the correlation filter and the motion likelihood map, we formulate an optimization problem. In addition, to handle the scale change of target, we repeat the combining process for various scale of bounding box. The experiments show that the proposed method outperforms the state-of-the-art methods. Byeongju Lee, Kimin Yun, Jongwon Choi 0002, Jin Young Choi 0002 |
AVSS | 3 |
| 2015 | User interactive segmentation with partially growing random forestabstractThis paper proposes a novel approach for user interactive segmentation based on graph-cut, which improves the robustness against the initial parameter setting. The existing graph-cut based segmentation uses a parametric model to estimate the color distributions of foreground/background. However, the parametric model is sensitive to the predefined number of distribution models and can be easily biased by a wrong initialization. In this paper, we develop a non-parametric approach based on random forest to handle the biased initialization problem. In addition, we design a new structure of random forest referred to as partially growing random forest to reduce the training time. We compare the proposed approach quantitatively and qualitatively to the existing graph-cut based segmentation baseline, where our method shows a remarkable performance on the new colorful dataset as well as comparable results on the classical dataset. Jongwon Choi 0002, Jin Young Choi 0002 |
ICIP | 1 |
| 2015 | Gradient preserving RGB-to-gray conversion using random forestabstractThis paper proposes a new algorithm for color-to-gray conversion preserving the gradient information in input color image. To preserve the gradient in a color image, we construct a random forest representing the relation between color intensity and gradient in an input image. The leaf nodes of random trees indicate the gray colors (single channel colors) corresponding to the input RGB colored pixels. From these initial gray colors obtained by the random forest, we determine the final gray scale by keeping the balance between intensity and luminance channels. In our experiments, we show that the proposed method outperforms the state-of-the-arts in view of color constrast preserving ratio and mean squared error versus luminance. Byeongju Lee, Jongwon Choi 0002, Kimin Yun, Jin Young Choi 0002 |
ICIP | 2 |