Jong Bin Ryu

dblp:139/4068 · also Jongbin Ryu · DBLP profile ↗
← Back
21ranked-venue papers
6as first author
12since 2021 · last 2025
0000-0001-5574-5358ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 7 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Do Your Best and Get Enough Rest for Continual Learning
abstract
According to the forgetting curve theory, we can enhance memory retention by learning extensive data and taking adequate rest. This means that in order to effectively retain new knowledge, it is essential to learn it thoroughly and ensure sufficient rest so that our brain can memorize without forgetting. The main takeaway from this theory is that learning extensive data at once necessitates sufficient rest before learning the same data again. This aspect of human long-term memory retention can be effectively utilized to address the continual learning of neural networks. Retaining new knowledge for a long period of time without catastrophic forgetting is the critical problem of continual learning. Therefore, based on Ebbinghaus’ theory, we introduce the view-batch model that adjusts the learning schedules to optimize the recall interval between retraining the same samples. The proposed view-batch model allows the network to get enough rest to learn extensive knowledge from the same samples with a recall interval of sufficient length. To this end, we specifically present two approaches: 1) a replay method that guarantees the optimal recall interval, and 2) a self-supervised learning that acquires extensive knowledge from a single training sample at a time. We empirically show that these approaches of our method are aligned with the forgetting curve theory, which can enhance long-term memory. In our experiments, we also demonstrate that our method significantly improves many state-of-the-art continual learning methods in various protocols and scenarios. We open-source this project at https://github.com/hankyul2/ViewBatchModel.
Hankyul Kang, Gregor Seifer, Jong Bin Ryu
CVPR4
2025 Channel Propagation Networks for Refreshable Vision Transformer
abstract
In this paper, we introduce the Channel Propagation method, which aims to increase the channels of the Vision Transformer systematically. Skip connections are commonly acknowledged as a propagation approach that improves the stability of the performance in Vision Transformers. Nevertheless, it is important to note that these skip connections may give rise to the problem of over-smoothing, wherein similar features are represented in multiple layers. To tackle this matter, our proposed approach for Channel Propagation in Vision Transformers retains the present signal information while concurrently propagating location-specific signals in a newly introduced channel dimension. On the other hand, the proposed Channel Propagation method effectively maintains the integrity of identity representation while simultaneously including patch-wise location-specific supervision by introducing a new channel dimension. The inclusion of this approach in Vision Transformers mitigates the issue of over-smoothing while also improving the performance of visual recognition tasks. In our experiments, we confirm that the proposed method is effective for various visual recognition tasks. Specifically, our method demonstrates enhanced performance when implemented on Vision Transformer models; the classification accuracy is increased considerably for plain and hierarchical architectures on the ImageNet dataset.
Junhyeong Go, Jong Bin Ryu
WACV2
2025 Enriching Local Patterns with Multi-Token Attention for Broad-Sight Neural Networks
abstract
In neural networks, recognizing visual patterns is challenging because global average pooling disregards local patterns and solely relies on over-concentrated activation. Global average pooling enforces the network to learn objects regardless of their location, so features tend to be activated only in specific regions. To support this claim, we provide a novel analysis of the problems that over-concentration brings about in networks with extensive experiments. We analyze the over-concentration through problems arising from feature variance and dead neurons that are not activated. Based on our analysis, we introduce a multi-token attention pooling layer to alleviate the over-concentration problem. Our attention-pooling layer captures broad-sight local patterns by learning multiple tokens with the proposed distillation algorithm. It resolves the high bias and high variance errors of learned multi-tokens, which is crucial when aggregating local patterns with multi-tokens. Our method applies to various vision tasks and network architectures such as CNN, ViT, and MLP-Mixer. The proposed method improves baselines with few extra resources, and a network employing our pooling method works favorably against state-of-the-art networks. We open-source the code at https://github.com/Lab-LVM/imagenet-models.
Hankyul Kang, Jong Bin Ryu
WACV2
2024 Neural Substitution for Branch-Level Network Re-parameterization
Seungmin Oh, Jong Bin Ryu
ACCV (8)2
2024 Generative Self-supervised Learning for Medical Image Classification
Inhyuk Park, Sungeun Kim, Jong Bin Ryu
ACCV (2)3
2024 Unsupervised Hashing Network with Hyper Quantization Tree
Sungeun Kim, Jong Bin Ryu
BMVC2
2024 Spatial Bias for attention-free non-local neural networks
abstract
In this paper, we introduce the Spatial Bias to learn global knowledge without self-attention in convolutional neural networks . Owing to the limited receptive field, conventional convolutional neural networks suffer from learning long-range dependencies. Non-local neural networks have struggled to learn global knowledge, but unavoidably have too heavy a network design due to the self-attention operation. Therefore, we propose a fast and lightweight Spatial Bias that efficiently encodes global knowledge without self-attention on convolutional neural networks . Spatial Bias is stacked on the feature map and convolved together to adjust the spatial structure of the convolutional features. Because we only use the convolution operation in this process, ours is lighter and faster than traditional methods based on the heavy self-attention operation. Therefore, we learn the global knowledge on the convolution layer directly with very few additional resources. Our method is very fast and lightweight due to the attention-free non-local method while improving the performance of neural networks considerably. Compared to non-local neural networks, the Spatial Bias use about × 10 times fewer parameters while achieving comparable performance with 1.6 ∼ 3.3 times more throughput on a very little budget. Furthermore, the Spatial Bias can be used with conventional non-local neural networks to further improve the performance of the backbone model. We show that the Spatial Bias achieves competitive performance that improves the classification accuracy by + 0 . 79 % and + 1 . 5 % on ImageNet-1K and CIFAR-100 datasets. Additionally, we validate our method on the MS-COCO and ADE20K datasets for downstream tasks involving object detection and semantic segmentation.
Junhyung Go, Jong Bin Ryu
Expert Syst. Appl.2
2024 Towards long-tailed, multi-label disease classification from chest X-ray: Overview of the CXR-LT challenge
Gregory Holste, Yiliang Zhou, Song Wang 0026, Ajay Jaiswal, Mingquan Lin, Sherry Zhuge, Yuzhe Yang 0003, Dongkyun Kim, Trong-Hieu Nguyen Mau, Minh-Triet Tran, Jaehyup Jeong, Wongi Park, Jong Bin Ryu, Feng Hong 0004, Arsh Verma, Yosuke Yamagishi, Hyeryeong Seo, Myungjoo Kang, Leo A. Celi, Zhiyong Lu, Ronald M. Summers, George Shih, Zhangyang Wang, Yifan Peng 0002
Medical Image Anal.13
2023 Gramian Attention Heads are Strong yet Efficient Vision Learners
abstract
We introduce a novel architecture design that enhances expressiveness by incorporating multiple head classifiers (i.e., classification heads) instead of relying on channel expansion or additional building blocks. Our approach employs attention-based aggregation, utilizing pairwise feature similarity to enhance multiple lightweight heads with minimal resource overhead. We compute the Gramian matrices to reinforce class tokens in an attention layer for each head. This enables the heads to learn more discriminative representations, enhancing their aggregation capabilities. Furthermore, we propose a learning algorithm that encourages heads to complement each other by reducing correlation for aggregation. Our models eventually surpass state-of-the-art CNNs and ViTs regarding the accuracy-throughput trade-off on ImageNet-1K and deliver remarkable performance across various downstream tasks, such as COCO object instance segmentation, ADE20k semantic segmentation, and fine-grained visual classification datasets. The effectiveness of our framework is substantiated by practical experimental results and further underpinned by generalization error bound. We release the code publicly at: https://github.com/Lab-LVM/imagenet-models.
Jong Bin Ryu, Dongyoon Han, Jongwoo Lim
ICCV1
2021 Unsupervised feature learning for self-tuning neural networks
Jong Bin Ryu, Ming-Hsuan Yang 0001, Jongwoo Lim
Neural Networks1
2021 End-to-End Learning for Omnidirectional Stereo Matching With Uncertainty Prior
abstract
In this paper, we propose a novel end-to-end deep neural network model for omnidirectional depth estimation from a wide-baseline multi-view stereo setup. The images captured with ultra-wide field-of-view cameras on an omnidirectional rig are processed by the feature extraction module, and then the deep feature maps are warped onto the concentric spheres swept through all candidate depths using the calibrated camera parameters. The 3D encoder-decoder block takes the aligned feature volume to produce an omnidirectional depth estimate with regularization on uncertain regions utilizing the global context information. For more accurate depth estimation we also propose an uncertainty prior guidance in two ways: depth map filtering and guiding regularization. In addition, we present large-scale synthetic datasets for training and testing omnidirectional multi-view stereo algorithms. Our datasets consist of 13K ground-truth depth maps and 53K fisheye images in four orthogonal directions with various objects and environments. Experimental results show that the proposed method generates excellent results in both synthetic and real-world environments, and it outperforms the prior art and the omnidirectional versions of the state-of-the-art conventional stereo algorithms.
Changhee Won, Jong Bin Ryu, Jongwoo Lim
IEEE Trans. Pattern Anal. Mach. Intell.2
2021 Dual aggregated feature pyramid network for multi label classification
Dongjoo Yun, Jong Bin Ryu, Jongwoo Lim
Pattern Recognit. Lett.2
2020 Generalized Convolutional Forest Networks for Domain Generalization and Visual Recognition
Jong Bin Ryu, Gitaek Kwon, Ming-Hsuan Yang 0001, Jongwoo Lim
ICLR1
2020 Unsupervised Face Domain Transfer for Low-Resolution Face Recognition
abstract
Low-resolution face recognition suffers from domain shift due to the different resolution between a high-resolution gallery and a low-resolution probe set. Conventional methods use the pairwise correlation between high-resolution and low-resolution for the same subject, which requires label information for both gallery and probe sets. However, explicitly labeled low-resolution probe images are seldom available, and labeling them is labor-intensive. In this paper, we propose a novel unsupervised face domain transfer for robust low-resolution face recognition. By leveraging the attention mechanism, the proposed generative face augmentation reduces the domain shift at image-level, while spatial resolution adaptation generates domain-invariant and discriminant feature distributions. On public datasets, we demonstrate the complementarity between generative face augmentation at image-level and spatial resolution adaptation at feature-level. The proposed method outperforms the state-of-the-art supervised methods even though we do not use any label information of low-resolution probe set.
Sungeun Hong, Jong Bin Ryu
IEEE Signal Process. Lett.2
2019 OmniMVS: End-to-End Learning for Omnidirectional Stereo Matching
abstract
In this paper, we propose a novel end-to-end deep neural network model for omnidirectional depth estimation from a wide-baseline multi-view stereo setup. The images captured with ultra wide field-of-view (FOV) cameras on an omnidirectional rig are processed by the feature extraction module, and then the deep feature maps are warped onto the concentric spheres swept through all candidate depths using the calibrated camera parameters. The 3D encoder-decoder block takes the aligned feature volume to produce the omnidirectional depth estimate with regularization on uncertain regions utilizing the global context information. In addition, we present large-scale synthetic datasets for training and testing omnidirectional multi-view stereo algorithms. Our datasets consist of 11K ground-truth depth maps and 45K fisheye images in four orthogonal directions with various objects and environments. Experimental results show that the proposed method generates excellent results in both synthetic and real-world environments, and it outperforms the prior art and the omnidirectional versions of the state-of-the-art conventional stereo algorithms.
Changhee Won, Jong Bin Ryu, Jongwoo Lim
ICCV2
2019 SweepNet: Wide-baseline Omnidirectional Depth Estimation
abstract
Omnidirectional depth sensing has its advantage over the conventional stereo systems since it enables us to recognize the objects of interest in all directions without any blind regions. In this paper, we propose a novel wide-baseline omnidirectional stereo algorithm which computes the dense depth estimate from the fisheye images using a deep convolutional neural network. The capture system consists of multiple cameras mounted on a wide-baseline rig with ultra-wide field of view (FOV) lenses, and we present the calibration algorithm for the extrinsic parameters based on the bundle adjustment. Instead of estimating depth maps from multiple sets of rectified images and stitching them, our approach directly generates one dense omnidirectional depth map with full 360° coverage at the rig global coordinate system. To this end, the proposed neural network is designed to output the cost volume from the warped images in the sphere sweeping method, and the final depth map is estimated by taking the minimum cost indices of the aggregated cost volume by SGM. For training the deep neural network and testing the entire system, realistic synthetic urban datasets are rendered using Blender. The experiments using the synthetic and real-world datasets show that our algorithm outperforms the conventional depth estimation methods and generate highly accurate depth maps.
Changhee Won, Jong Bin Ryu, Jongwoo Lim
ICRA2
2018 DFT-based Transformation Invariant Pooling Layer for Visual Classification
Jong Bin Ryu, Ming-Hsuan Yang 0001, Jongwoo Lim
ECCV (14)1
2018 D3: Recognizing dynamic scenes with deep dual descriptor based on key frames and key segments
Sungeun Hong, Jong Bin Ryu, Woobin Im, Hyun Seung Yang
Neurocomputing2
2017 SSPP-DAN: Deep domain adaptation network for face recognition with single sample per person
abstract
Real-world face recognition using a single sample per person (SSPP) is a challenging task. The problem is exacerbated if the conditions under which the gallery image and the probe set are captured are completely different. To address these issues from the perspective of domain adaptation, we introduce an SSPP domain adaptation network (SSPP-DAN). In the proposed approach, domain adaptation, feature extraction, and classification are performed jointly using a deep architecture with domain-adversarial training. However, the SSPP characteristic of one training sample per class is insufficient to train the deep architecture. To overcome this shortage, we generate synthetic images with varying poses using a 3D face model. Experimental evaluations using a realistic SSPP dataset show that deep domain adaptation and image synthesis complement each other and dramatically improve accuracy. Experiments on a benchmark dataset using the proposed approach show state-of-the-art performance.
Sungeun Hong, Woobin Im, Jong Bin Ryu, Hyun Seung Yang
ICIP3
2016 Locality-preserving descriptor for robust texture feature representation
Jong Bin Ryu, Hyun Seung Yang
Neurocomputing1
2015 Sorted Consecutive Local Binary Pattern for Texture Classification
abstract
In this paper, we propose a sorted consecutive local binary pattern (scLBP) for texture classification. Conventional methods encode only patterns whose spatial transitions are not more than two, whereas scLBP encodes patterns regardless of their spatial transition. Conventional methods do not encode patterns on account of rotation-invariant encoding; on the other hand, patterns with more than two spatial transitions have discriminative power. The proposed scLBP encodes all patterns with any number of spatial transitions while maintaining their rotation-invariant nature by sorting the consecutive patterns. In addition, we introduce dictionary learning of scLBP based on kd-tree which separates data with a space partitioning strategy. Since the elements of sorted consecutive patterns lie in different space, it can be generated to a discriminative code with kd-tree. Finally, we present a framework in which scLBPs and the kd-tree can be combined and utilized. The results of experimental evaluation on five texture data sets--Outex, CUReT, UIUC, UMD, and KTH-TIPS2-a--indicate that our proposed framework achieves the best classification rate on the CUReT, UMD, and KTH-TIPS2-a data sets compared with conventional methods. The results additionally indicate that only a marginal difference exists between the best classification rate of conventional methods and that of the proposed framework on the UIUC and Outex data sets.
Jong Bin Ryu, Sungeun Hong, Hyun Seung Yang
IEEE Trans. Image Process.1