Janghyeon Lee 0001

dblp:203/9928-1 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0002-8599-4678ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 6 since 2021Systems, architecture and hardware · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Controllable Feature Whitening for Hyperparameter-Free Bias Mitigation
Yooshin Cho, Hanbyel Cho, Janghyeon Lee 0001, Hyeong Gwon Hong, Jaesung Ahn, Junmo Kim 0002
ICCV3
2023 Disposable Transfer Learning for Selective Source Task Unlearning
abstract
Transfer learning is widely used for training deep neural networks (DNN) for building a powerful representation. Even after the pre-trained model is adapted for the target task, the representation performance of the feature extractor is retained to some extent. As the performance of the pre-trained model can be considered the private property of the owner, it is natural to seek the exclusive right of the generalized performance of the pre-trained weight. To address this issue, we suggest a new paradigm of transfer learning called disposable transfer learning (DTL), which disposes of only the source task without degrading the performance of the target task. To achieve knowledge disposal, we propose a novel loss named Gradient Collision loss (GC loss). GC loss selectively unlearns the source knowledge by leading the gradient vectors of mini-batches in different directions. Whether the model successfully unlearns the source task is measured by piggyback learning accuracy (PL accuracy). PL accuracy estimates the vulnerability of knowledge leakage by retraining the scrubbed model on a subset of source data or new downstream data. We demonstrate that GC loss is an effective approach to the DTL problem by showing that the model trained with GC loss retains the performance on the target task with a significantly reduced PL accuracy.
Seunghee Koh, Hyounguk Shon, Janghyeon Lee 0001, Hyeong Gwon Hong, Junmo Kim 0002
ICCV3
2022 DLCFT: Deep Linear Continual Fine-Tuning for General Incremental Learning
Hyounguk Shon, Janghyeon Lee 0001, Junmo Kim 0002
ECCV (33)2
2022 On the Angular Update and Hyperparameter Tuning of a Scale-Invariant Network
Juseung Yun, Janghyeon Lee 0001, Hyounguk Shon, Eojindl Yi, Junmo Kim 0002
ECCV (12)2
2022 Multi-Scaled and Densely Connected Locally Convolutional Layers for Depth Completion
abstract
The depth completion task aims to predict a dense depth map from a sparse LiDAR point cloud and an RGB image. This task is critical because an accurate depth map can be used as prior information to solve many computer vision tasks, such as downstream tasks in autonomous vehicles and robot vision. Previous deep learning methods which focus on the local affinity have achieved impressive results. However, an architecture that is directly designed to extract local affinity has not been proposed yet. In this paper, we propose multi-scaled and densely connected locally convolutional layers to learn the affinity of the neighborhood. We set a different grid factor for each step of this module, and each step consists of several convolutional layers applied only to the local area assigned from the grid factor. In addition, each step is densely connected, sequentially, to take advantage of the multi-scale receptive fields. The proposed module effectively learns the neighbor-hood's affinity in a local area with multiple scales, while keeping the network size small. As a result, our architecture achieves state-of-the-art performance compared to published works on the KITTI depth completion benchmark. On the NYU Depth V2 completion benchmark our method achieves performance comparable to state-of-the-art approaches.
Sihaeng Lee, Eojindl Yi, Janghyeon Lee 0001, Junmo Kim 0002
IROS3
2022 Fully Convolutional Transformer with Local-Global Attention
abstract
In an attempt to imitate the success of transformers in the field of natural language processing into computer vision tasks, vision transformers (ViTs) have recently gained attention. Performance breakthroughs have been achieved in coarse-grained tasks like classification. However, dense prediction tasks, such as detection, segmentation, and depth estimation, require additional modifications and have been tackled only in an ad-hoc manner, by replacing the convolutional neural network encoder backbone of an existing architecture with a ViT. This study proposes a fully convolutional transformer that can perform both coarse and dense prediction tasks. The proposed architecture is, to the best of our knowledge, the first architecture composed of attention layers, even in the decoder part of the network. This is because our newly proposed local-global attention (LGA) can flexibly perform both downsampling and upsampling of spatial features, which are key operations required for dense prediction. Against existing ViTs on classification tasks, our architecture shows a reasonable trade-off between performance and efficiency. In the depth estimation task, our architecture achieves performance comparable to that of state-of-the-art transformer-based methods.
Sihaeng Lee, Eojindl Yi, Janghyeon Lee 0001, Jinsu Yoo, Honglak Lee
IROS3
2022 UniCLIP: Unified Framework for Contrastive Language-Image Pre-training
abstract
Pre-training vision-language models with contrastive objectives has shown promising results that are both scalable to large uncurated datasets and transferable to many downstream applications. Some following works have targeted to improve data efficiency by adding self-supervision terms, but inter-domain (image-text) contrastive loss and intra-domain (image-image) contrastive loss are defined on individual spaces in those works, so many feasible combinations of supervision are overlooked. To overcome this issue, we propose UniCLIP, a Unified framework for Contrastive Language-Image Pre-training. UniCLIP integrates the contrastive loss of both inter-domain pairs and intra-domain pairs into a single universal space. The discrepancies that occur when integrating contrastive loss between different domains are resolved by the three key components of UniCLIP: (1) augmentation-aware feature embedding, (2) MP-NCE loss, and (3) domain dependent similarity measure. UniCLIP outperforms previous vision-language pre-training methods on various single- and multi-modality downstream tasks. In our experiments, we show that each component that comprises UniCLIP contributes well to the final performance.
Janghyeon Lee 0001, Jongsuk Kim, Hyounguk Shon, Bumsoo Kim 0005, Honglak Lee, Junmo Kim 0002
NeurIPS1
2022 TricubeNet: 2D Kernel-Based Object Representation for Weakly-Occluded Oriented Object Detection
abstract
We present a novel approach for oriented object detection, named TricubeNet, which localizes oriented objects using visual cues (i.e., heatmap) instead of oriented box offsets regression. We represent each object as a 2D Tricube kernel and extract bounding boxes using simple image-processing algorithms. Our approach is able to (1) obtain well-arranged boxes from visual cues, (2) solve the angle discontinuity problem, and (3) can save computational complexity due to our anchor-free modeling. To further boost the performance, we propose some effective techniques for size-invariant loss, reducing false detections, extracting rotation-invariant features, and heatmap refinement. To demonstrate the effectiveness of our TricubeNet, we experiment on various tasks for weakly-occluded oriented object detection: detection in an aerial image, densely packed object image, and text image. The extensive experimental results show that our TricubeNet is quite effective for oriented object detection. Code is available at https://github.com/qjadud1994/TricubeNet.
Beomyoung Kim, Janghyeon Lee 0001, Sihaeng Lee, Junmo Kim 0002
WACV2
2021 Patch-Wise Attention Network for Monocular Depth Estimation
abstract
In computer vision, monocular depth estimation is the problem of obtaining a high-quality depth map from a two-dimensional image. This map provides information on three-dimensional scene geometry, which is necessary for various applications in academia and industry, such as robotics and autonomous driving. Recent studies based on convolutional neural networks achieved impressive results for this task. However, most previous studies did not consider the relationships between the neighboring pixels in a local area of the scene. To overcome the drawbacks of existing methods, we propose a patch-wise attention method for focusing on each local area. After extracting patches from an input feature map, our module generates attention maps for each local patch, using two attention modules for each patch along the channel and spatial dimensions. Subsequently, the attention maps return to their initial positions and merge into one attention feature. Our method is straightforward but effective. The experimental results on two challenging datasets, KITTI and NYU Depth V2, demonstrate that the proposed method achieves significant performance. Furthermore, our method outperforms other state-of-the-art methods on the KITTI depth estimation benchmark.
Sihaeng Lee, Janghyeon Lee 0001, Byungju Kim, Eojindl Yi, Junmo Kim 0002
AAAI2
2020 Residual Continual Learning
abstract
We propose a novel continual learning method called Residual Continual Learning (ResCL). Our method can prevent the catastrophic forgetting phenomenon in sequential learning of multiple tasks, without any source task information except the original network. ResCL reparameterizes network parameters by linearly combining each layer of the original network and a fine-tuned network; therefore, the size of the network does not increase at all. To apply the proposed method to general convolutional neural networks, the effects of batch normalization layers are also considered. By utilizing residual-learning-like reparameterization and a special weight decay loss, the trade-off between source and target performance is effectively controlled. The proposed method exhibits state-of-the-art performance in various continual learning scenarios.
Janghyeon Lee 0001, Donggyu Joo, Hyeong Gwon Hong, Junmo Kim 0002
AAAI1
2020 Continual Learning With Extended Kronecker-Factored Approximate Curvature
abstract
We propose a quadratic penalty method for continual learning of neural networks that contain batch normalization (BN) layers. The Hessian of a loss function represents the curvature of the quadratic penalty function, and a Kronecker-factored approximate curvature (K-FAC) is used widely to practically compute the Hessian of a neural network. However, the approximation is not valid if there is dependence between examples, typically caused by BN layers in deep network architectures. We extend the K-FAC method so that the inter-example relations are taken into account and the Hessian of deep neural networks can be properly approximated under practical assumptions. We also propose a method of weight merging and reparameterization to properly handle statistical parameters of BN, which plays a critical role for continual learning with BN, and a method that selects hyperparameters without source task data. Our method shows better performance than baselines in the permuted MNIST task with BN layers and in sequential learning from the ImageNet classification task to fine-grained classification tasks with ResNet-50, without any explicit or implicit use of source task data for hyperparameter selection.
Janghyeon Lee 0001, Hyeong Gwon Hong, Donggyu Joo, Junmo Kim 0002
CVPR1
2017 Refine pedestrian detections by referring to features in different ways
abstract
The performance of object detection has been improved as the success of deep architectures. The main algorithm predominantly used for general detection is Faster R-CNN because of their high accuracy and fast inference time. In pedestrian detection, Region Proposal Network (RPN) itself which is used for region proposals in Faster R-CNN can be used as a pedestrian detector. Also, RPN even shows better performance than Faster R-CNN for pedestrian detection. However, RPN generates severe false positives such as high score backgrounds and double detections because it does not have downstream classifier. From this observations, we made a network to refine results generated from the RPN. Our Refinement Network refers to the feature maps of the RPN and trains the network to rescore severe false positives. Also, we found that different type of feature referencing method is crucial for improving performance. Our network showed better accuracy than RPN with almost same speed on Caltech Pedestrian Detection benchmark.
Jaemyung Lee 0003, Sihaeng Lee, Youngdong Kim, Janghyeon Lee 0001, Junmo Kim 0002
Intelligent Vehicles Symposium4