VLDB 2026 Research / reviewers in the wild / expert
Jae Young Lee 0002
dblp:46/985-2
· DBLP profile ↗
12ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0002-7450-5023ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DMQ: Dissecting Outliers of Diffusion Models for Post-Training QuantizationabstractDiffusion models have achieved remarkable success in image generation but come with significant computational costs, posing challenges for deployment in resource-constrained environments. Recent post-training quantization (PTQ) methods have attempted to mitigate this issue by focusing on the iterative nature of diffusion models. However, these approaches often overlook outliers, leading to degraded performance at low bit-widths. In this paper, we propose a DMQ which combines Learned Equivalent Scaling (LES) and channel-wise Power-of-Two Scaling (PTS) to effectively address these challenges. Learned Equivalent Scaling optimizes channel-wise scaling factors to redistribute quantization difficulty between weights and activations, reducing overall quantization error. Recognizing that early denoising steps, despite having small quantization errors, crucially impact the final output due to error accumulation, we incorporate an adaptive timestep weighting scheme to prioritize these critical steps during learning. Furthermore, identifying that layers such as skip connections exhibit high inter-channel variance, we introduce channel-wise Power-of-Two Scaling for activations. To ensure robust selection of PTS factors even with small calibration set, we introduce a voting algorithm that enhances reliability. Extensive experiments demonstrate that our method significantly outperforms existing works, especially at low bit-widths such as W4A6 (4-bit weight, 6-bit activation) and W4A8, maintaining high image generation quality and model stability. The code is available at https://github.com/LeeDongYeun/dmq. Dongyeun Lee, Jiwan Hur, Hyounguk Shon, Jae Young Lee 0002, Junmo Kim 0002 |
ICCV | 4 |
| 2024 | Modeling Stereo-Confidence out of the End-to-End Stereo-Matching Network via Disparity Plane SweepabstractWe propose a novel stereo-confidence that can be measured externally to various stereo-matching networks, offering an alternative input modality choice of the cost volume for learning-based approaches, especially in safety-critical systems. Grounded in the foundational concepts of disparity definition and the disparity plane sweep, the proposed stereo-confidence method is built upon the idea that any shift in a stereo-image pair should be updated in a corresponding amount shift in the disparity map. Based on this idea, the proposed stereo-confidence method can be summarized in three folds. 1) Using the disparity plane sweep, multiple disparity maps can be obtained and treated as a 3-D volume (predicted disparity volume), like the cost volume is constructed. 2) One of these disparity maps serves as an anchor, allowing us to define a desirable (or ideal) disparity profile at every spatial point. 3) By comparing the desirable and predicted disparity profiles, we can quantify the level of matching ambiguity between left and right images for confidence measurement. Extensive experimental results using various stereo-matching networks and datasets demonstrate that the proposed stereo-confidence method not only shows competitive performance on its own but also consistent performance improvements when it is used as an input modality for learning-based stereo-confidence methods. Jae Young Lee 0002, Woonghyun Ka, Jaehyun Choi, Junmo Kim 0002 |
AAAI | 1 |
| 2024 | Stereo-Matching Knowledge Distilled Monocular Depth Estimation Filtered by Multiple Disparity ConsistencyabstractIn stereo-matching knowledge distillation methods of the self-supervised monocular depth estimation, the stereo-matching network’s knowledge is distilled into a monocular depth network through pseudo-depth maps. In these methods, the learning-based stereo-confidence network is generally utilized to identify errors in the pseudo-depth maps to prevent transferring the errors. However, the learning-based stereo-confidence networks should be trained with ground truth (GT), which is not feasible in a self-supervised setting. In this paper, we propose a method to identify and filter errors in the pseudo-depth map using multiple disparity maps by checking their consistency without the need for GT and a training process. Experimental results show that the proposed method outperforms the previous methods and works well on various configurations by filtering out erroneous areas where the stereo-matching is vulnerable, especially such as textureless regions, occlusion boundaries, and reflective surfaces. Woonghyun Ka, Jae Young Lee 0002, Jaehyun Choi, Junmo Kim 0002 |
ICASSP | 2 |
| 2024 | Real-Time Polyp Detection in Colonoscopy using Lightweight TransformerabstractColorectal cancer (CRC) represents a major global health challenge, and early detection of polyps is crucial in preventing its progression. Although colonoscopy is the gold standard for polyp detection, it has limitations, such as human error and missed detection rates. In response, computer-aided detection (CADe) systems have been developed to enhance the efficiency and accuracy of polyp detection. As deep learning gained prominence, the incorporation of Convolutional Neural Networks (CNNs) into CADe systems emerged as a breakthrough approach. However, CADe systems based on CNNs often demand significant computational resources, making them unsuitable for deployment in resource-constrained environments. To mitigate this, we propose a novel and lightweight polyp detection model that integrates a Transformer layer into the You Only Look Once (YOLO) architecture, focusing on optimizing the neck part responsible for feature fusion and rescaling. Our model demonstrates a substantial reduction in computational complexity and the number of parameters, without compromising detection performances. The lightweight model makes it accessible and feasibly deployable in medically underserved regions, serving a significant public interest by potentially expanding the reach of critical diagnostic tools for CRC prevention. By optimizing the architecture to reduce resource requirements while maintaining performance, our model becomes a practical solution to assist healthcare professionals in the real-time identification of polyps, even with resource-constraint devices. Youngbeom Yoo, Jae Young Lee 0002, Jiwoon Jeon, Junmo Kim 0002 |
WACV | 2 |
| 2023 | Few-Shot Anomaly Detection with Adversarial Loss for Robust Feature Representations
Jae Young Lee 0002, Jaehyun Choi, Yongkwi Lee, Young Seog Yoon |
BMVC | 1 |
| 2023 | Fix the Noise: Disentangling Source Feature for Controllable Domain TranslationabstractRecent studies show strong generative performance in domain translation especially by using transfer learning techniques on the unconditional generator. However, the control between different domain features using a single model is still challenging. Existing methods often require additional models, which is computationally demanding and leads to unsatisfactory visual quality. In addition, they have restricted control steps, which prevents a smooth transition. In this paper, we propose a new approach for high-quality domain translation with better controllability. The key idea is to preserve source features within a disentangled subspace of a target feature space. This allows our method to smoothly control the degree to which it preserves source features while generating images from an entirely new domain using only a single model. Our extensive experiments show that the proposed method can produce more consistent and realistic images than previous works and maintain precise controllability over different levels of transformation. The code is available at LeeDongYeun/FixNoise. Dongyeun Lee, Jae Young Lee 0002, Jaehyun Choi, Jaejun Yoo 0001, Junmo Kim 0002 |
CVPR | 2 |
| 2023 | Lightweight Monocular Depth Estimation via Token-Sharing TransformerabstractDepth estimation is an important task in various robotics systems and applications. In mobile robotics systems, monocular depth estimation is desirable since a single RGB camera can be deployable at a low cost and compact size. Due to its significant and growing needs, many lightweight monocular depth estimation networks have been proposed for mobile robotics systems. While most lightweight monocular depth estimation methods have been developed using convolution neural networks, the Transformer has been gradually utilized in monocular depth estimation recently. However, massive parameters and large computational costs in the Transformer disturb the deployment to embedded devices. In this paper, we present a Token-Sharing Transformer (TST), an architecture using the Transformer for monocular depth estimation, optimized especially in embedded devices. The proposed TST utilizes global token sharing, which enables the model to obtain an accurate depth prediction with high throughput in embedded devices. Experimental results show that TST outperforms the existing lightweight monocular depth estimation methods. On the NYU Depth v2 dataset, TST can deliver depth maps up to 63.4 FPS in NVIDIA Jetson nano and 142.6 FPS in NVIDIA Jetson TX2, with lower errors than the existing methods. Furthermore, TST achieves real-time depth estimation of high-resolution images on Jetson TX2 with competitive results. Jae Young Lee 0002, Hyunguk Shon, Eojindl Yi, Yeong-Hun Park, Sung-Sik Cho, Junmo Kim 0002 |
ICRA | 2 |
| 2023 | I See-Through You: A Framework for Removing Foreground Occlusion in Both Sparse and Dense Light Field ImagesabstractLight field (LF) camera captures rich information from a scene. Using the information, the LF de-occlusion (LF-DeOcc) task aims to reconstruct the occlusion-free center view image. Existing LF-DeOcc studies mainly focus on the sparsely sampled (sparse) LF images where most of the occluded regions are visible in other views due to the large disparity. In this paper, we expand LF-DeOcc in more challenging datasets, densely sampled (dense) LF images, which are taken by a micro-lens-based portable LF camera. Due to the small disparity ranges of dense LF images, most of the background regions are invisible in any view. To apply LF-DeOcc in both LF datasets, we propose a framework, ISTY, which is defined and divided into three roles: (1) extract LF features, (2) define the occlusion, and (3) inpaint occluded regions. By dividing the framework into three specialized components according to the roles, the development and analysis can be easier. Furthermore, an explainable intermediate representation, an occlusion mask, can be obtained in the proposed framework. The occlusion mask is useful for comprehensive analysis of the model and other applications by manipulating the mask. In experiments, qualitative and quantitative results show that the proposed framework outperforms state-of-the-art LF-DeOcc methods in both sparse and dense LF datasets. Jiwan Hur, Jae Young Lee 0002, Jaehyun Choi, Junmo Kim 0002 |
WACV | 2 |
| 2023 | Multi-scale foreground-background separation for light field depth estimation with deep convolutional networks
Jae Young Lee 0002, Jiwan Hur, Jaehyun Choi, Rae-Hong Park, Junmo Kim 0002 |
Pattern Recognit. Lett. | 1 |
| 2022 | Cyclic test time augmentation with entropy weight methodabstractIn the recent studies of data augmentation of neural networks, the application of test time augmentation has been studied to extract optimal transformation policies to enhance performance with minimum cost. The policy search method with the best level of input data dependency involves training a loss predictor network to estimate suitable transformations for each of the given input image in independent manner, resulting in instance-level transformation extraction. In this work, we propose a method to utilize and modify the loss prediction pipeline to further improve the performance with the cyclic search for suitable transformations and the use of the entropy weight method. The cyclic usage of the loss predictor allows refining each input image with multiple transformations with a more flexible transformation magnitude. For cases where multiple augmentations are generated, we implement the entropy weight method to reflect the data uncertainty of each augmentation to force the final result to focus on augmentations with low uncertainty. The experimental results show convincing qualitative outcomes and robust performance for the corrupted conditions of data. Sewhan Chun, Jae Young Lee 0002, Junmo Kim 0002 |
UAI | 2 |
| 2021 | Complex-Valued Disparity: Unified Depth Model of Depth from Stereo, Depth from Focus, and Depth from Defocus Based on the Light Field GradientabstractThis paper proposes a unified depth model based on the light field gradient, in which estimated disparity is represented by the complex number. The complex-valued disparity by the proposed depth model can be represented in both the Cartesian and polar coordinates. In the Cartesian representation, the proposed depth model is represented by real and imaginary parts of the disparity. The real part can be used for disparity estimation with respect to the in-focus plane, whereas the imaginary part represents the non-Lambertian-ness. In the polar representation, the proposed depth model is expressed by the disparity magnitude and disparity angle. The disparity magnitude shows the relationship among depth from stereo, depth from focus, and depth from defocus, whereas the disparity angle shows whether or not the bundles of rays are flipped with respect to the in-focus plane. For disparity analysis, we present the real response, imaginary response, magnitude response, and angle response, which are represented by the three-dimensional volume. Experimental results on synthetic and real light field images show that the real and magnitude responses of the proposed depth model are valid for local disparity estimation. Jae Young Lee 0002, Rae-Hong Park |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Reduction of Aliasing Artifacts by Sign Function Approximation in Light Field Depth Estimation Based on Foreground-Background SeparationabstractA sign function approximation method for depth from light field (DFLF) based on the foreground–background separation (FBS) is proposed. From a signal processing viewpoint, the FBS-based method can be considered as a bridge between the cost-based DFLF methods and depth model-based ones. The proposed sign function approximation corresponds to the winner-takes-all (WTA) approaches as the cost-based methods do. Experimental results on the synthetic images show that the proposed method reasonably performs in terms of the mean squared error as the state-of-the-art methods do. Especially, by using a suitable WTA method in the framework of our previous FBS-based DFLF work, the proposed method effectively reduces the angular aliasing artifacts in the resulting disparity maps of both the synthetic and real images. Jae Young Lee 0002, Rae-Hong Park |
IEEE Signal Process. Lett. | 1 |