Seok Bong Yoo

dblp:29/10698 · DBLP profile ↗
← Back
27ranked-venue papers
4as first author
22since 2021 · last 2026
0000-0002-6528-701XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 9 since 2021Systems, architecture and hardware · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 MonoSTR: Structure-aware monocular 3D object detection via keypoint-based geometric representation
abstract
• Proposes a structure-aware monocular 3D detector based on semantic keypoints. • Recovers occluded object structure using a direction-guided autoencoder. • Uses dual-band depth features to mitigate depth ambiguity under occlusion. • Achieves reliable 3D perception for safety-critical transportation systems. Monocular 3D object detection provides a cost-effective alternative to multi-sensor setups, but estimating accurate depth from a single image remains inherently challenging. The difficulty becomes more pronounced when objects are partially visible. To address these challenges, we introduce skeleton keypoints. Semantic skeleton keypoints serve as interpretable structural cues that encode expert knowledge about object geometry. Our keypoint structural representation detects visible skeleton keypoints and reconstructs missing ones with an autoencoder reconstructor conditioned on coarse pose, providing a compact and shape-aware cue tolerant to occlusion and truncation. In parallel, depth features are estimated by a dual-band-guided depth predictor, which explicitly decomposes low- and high-frequency components to balance global context and fine details, thereby mitigating spectral bias and improving depth estimation for tiny and distant objects. These representations are then fused, and a scene-topology 3D detection head decodes class and box parameters while conditioning on neighborhood topology to enforce geometric consistency. Experiments show that our method provides more reliable 3D perception for decision-support systems in intelligent transportation by improving robustness under occlusion, demonstrating its effectiveness for autonomous driving perception systems.
Yeon Woo Cho, Jung Woo Cheon, Seok Bong Yoo
Expert Syst. Appl.3
2026 Privacy-preserving person re-identification through identity retrieval and hierarchical latent code protection
Seunghyeok Back, Seok Bong Yoo
Inf. Sci.2
2026 FedPure: Data poisoning attack detection and purification for federated skeleton-based action recognition
Minhyuk Kim, Eungi Lee, Seok Bong Yoo
Inf. Sci.3
2026 Unified auxiliary restoration network for robust multimodal 3D object detection in adverse conditions
Jae Hyun Yoon, Jong Won Jung, Seok Bong Yoo
Neural Networks3
2026 MonoRange: Monocular 3-D Object Detection Based on Object-Centric Range Map in Adverse Weather Conditions
abstract
Monocular 3D object detection has been studied as a promising task for diverse applications, such as autonomous driving, due to its lower cost and more straightforward configuration than multiple sensors. However, existing studies have focused on clear weather without considering diverse weather conditions with varying intensity, such as rain, snow, and fog, affecting detection performance. In this paper, we propose MonoRange, a monocular 3D object detection method that uses object-centric images and range maps in adverse weather conditions. Leveraging the 2D detection results, MonoRange generates range maps from images via an object-centric range map reconstruction. Furthermore, MonoRange flexibly removes adverse weather noise in images via weather intensity adaptive image restoration with a weight modulation transformer. Then, MonoRange fuses the range map and restored image and predicts 3D bounding boxes via the range map aligned detector. Introducing the projected box consistency loss between 2D and 3D boxes also enables to consistent and accurate 3D object detection. Experimental results on diverse weather datasets demonstrate that MonoRange surpasses existing monocular 3D object detection approaches. The source code is available athttps://github.com/jhyoon964/MonoRange
Jae Hyun Yoon, Jong Won Jung, Seok Bong Yoo
IEEE Trans. Intell. Transp. Syst.3
2025 Equirectangular Point Reconstruction for Domain Adaptive Multimodal 3D Object Detection in Adverse Weather Conditions
abstract
A multimodal fusion technique using LiDAR-camera has been developed for precise 3D object detection in autonomous driving and provides acceptable detection performance in ideal conditions with clear weather. However, the existing multimodal methods are still vulnerable to adverse weather conditions, such as snow, rain, and fog. These factors increase the point cloud sparsity due to occlusion and attenuation of the laser signal. A point cloud becomes sparser with increased distance, posing a challenge for object detection. To address these problems, we propose a point reconstruction network using equirectangular projection for multimodal 3D object detection. This network consists of distance-constrained denoising to remove adverse weather noise and an object-centric ray generator to generate distant object points flexibly. We propose a domain adaptation method that injects feature perturbations to improve detection performance by reducing the domain gap between different datasets. Furthermore, we propose a multimodal weather noise matching method for realistic data synthesis-based training to align the adverse weather noise between synthetic point clouds and images. The experimental results on adverse weather datasets confirm that the proposed approach outperforms the existing methods.
Jae Hyun Yoon, Jong Won Jung, Seok Bong Yoo
AAAI3
2025 X-FLoRA: Cross-modal Federated Learning with Modality-expert LoRA for Medical VQA
abstract
Medical visual question answering (VQA) and federated learning (FL) have emerged as vital approaches for enabling privacy-preserving, collaborative learning across clinical institutions. However, both these approaches face significant challenges in cross-modal FL scenarios, where each client possesses unpaired images from only one modality. To address this limitation, we propose X-FLoRA, a cross-modal FL framework that uses modality-expert low-rank adaptation (LoRA) for medical VQA. Specifically, X-FLoRA enables the synthesis of images from one modality to another without requiring data sharing between clients. This is achieved by training a backward translation model within a federated asymmetric translation scheme that integrates clinical semantics from textual data. Additionally, X-FLoRA introduces modality-expert LoRA, which fine-tunes separate LoRA modules to strengthen modality-specific representations in the VQA task. The server aggregates the trained backward translation models and fine-tuned LoRA modules using discriminator quality scores and expert-aware weighting, which regulate the relative contributions from different clients. Experiments were conducted on VQA datasets encompassing different medical modalities, and the results demonstrate that X-FLoRA outperforms existing FL methods in terms of VQA performance.
Minhyuk Kim, Changheon Kim, Seok Bong Yoo
EMNLP3
2025 Ego-$A^{\mathbf{3}}$: Adaptive Fusion-Based Disentangled Transformer for Egocentric Action Anticipation
abstract
Recently, egocentric action anticipation for wearable robotics cameras has gained considerable attention due to its capability to analyze nouns and verbs from a firstperson view. However, this field encounters challenges due to various uncertainties, such as action-irrelevant information and semantically fused representations of verbs and nouns. To overcome these issues, we introduce Ego-$A^{3}$, designed to improve the robustness and reliability of egocentric action anticipation systems. Ego-$A^{3}$adaptively extracts actionrelevant data to efficiently utilize additional information beyond visual data. Additionally, Ego-$A^{3}$produces effective disentangled representations for verbs and nouns by employing learnable verb and noun queries. Experiments on the EpicKitchens-100 and EGTEA Gaze+ datasets demonstrate that Ego-$A^{3}$outperforms existing methods in top-1 accuracy and mean top- 5 recall. Our code is publicly available at https://github.com/alsgur0720/egocentricanticipation.
Minhyuk Kim, Jong Won Jung, Eungi Lee, Seok Bong Yoo
ICRA4
2025 FedDet: Data Poisoning Attack Detection for Federated Skeleton-based Action Recognition
Minhyuk Kim, Eungi Lee, Seok Bong Yoo
ICRA3
2025 OPRNet: Object-Centric Point Reconstruction Network for Multimodal 3D Object Detection in Adverse Weathers
abstract
The development of a multimodal fusion technique utilizing LiDAR-camera data has enabled precise 3D object detection for self-driving vehicles, particularly in ideal conditions with clear weather. Nevertheless, adverse weathers such as fog, snow, and rain remain a challenge for existing multimodal methods. These conditions lead to a reduced density of point clouds as a result of laser signal occlusion and attenuation. Additionally, as the distance grows, the point cloud becomes sparser, further challenging object detection tasks. To address these problems, we introduce a point reconstruction network employing equirectangular projection tailored for multimodal 3D object detection. This network incorporates a range-constrained noise filter to remove noise caused by adverse weather and an object-centric point generator designed to flexibly generate points for distant objects. Moreover, we propose a dual 2D auxiliary module to enhance image features and support the point reconstruction. Experimental evaluations conducted on adverse weather datasets demonstrate that the suggested approach surpasses current techniques. The implementation can be accessed at https://github.com/jhyoon964/oprnet.
Jae Hyun Yoon, Jong Won Jung, Eungi Lee, Seok Bong Yoo
ICRA4
2025 Data Poisoning Attack Defense and Evolutionary Domain Adaptation for Federated Medical Image Segmentation
abstract
Federated learning has significant demonstrated potential in medical image segmentation to protect data privacy by retaining local data. However, its application is still hindered by two critical challenges: 1) the retained data poisoning attacks that severely compromise the accuracy of the global segmentation model and 2) domain gaps among clients, undermining its generalizability. To address these issues, we propose AdaShield-FL, a data poisoning attack defense and evolutionary domain adaptation for federated medical image segmentation. AdaShield-FL incorporates a disentangled reconstruction and segmentation module that purifies data in the k-space domain to mitigate the effects of adversarial attacks iteratively. Moreover, it introduces a data poisoning attack detection mechanism that analyzes abnormal patterns in training loss sequences to identify malicious clients. This method also aligns local and global covariance matrices via evolutionary optimization to minimize the domain gap efficiently. The experimental validation on cardiac magnetic resonance imaging datasets demonstrates the robustness and superior performance of AdaShield-FL compared with other federated learning methods.
Minhyuk Kim, Seok Bong Yoo
IJCAI2
2025 TransSino: Prior Sinogram Pattern-Based Transformer for Limited-Angle CT Image Segmentation
Jae Hyun Yoon, Yeong Jong Lee, Seok Bong Yoo
MICCAI (16)3
2025 SCOL: Style Code Orchestration in Latent Space for Proactive Face-Swapping Defense
abstract
Face-swapping deepfake poses significant risks, including privacy violations, misinformation, and defamation, amplified by the availability of pretrained models on open-source platforms. Proactive defense strategies aim to disrupt deepfake generation by modifying the original images to protect identity features. However, existing methods often introduce artifacts in facial images or rely on specific deepfake models, limiting their usability. To address these problems, we propose a style code orchestration in latent space (SCOL) method that obfuscates identity by fusing different identities in the latent space without requiring face recognition models. This study optimizes the generator to follow the original appearance while retaining the obfuscated identity via identity-preserving constraints. Further, appearance-dominant components in the latent code are aligned for visual consistency. An identity inversion attack is introduced using opposite style codes to improve the effectiveness of the defense. Experimental results demonstrate that SCOL robustly defends against various face-swapping deepfake methods, maintaining visual consistency.
Eungi Lee, Jae Hyun Yoon, Seok Bong Yoo
ACM Multimedia3
2025 Disentangled adaptive fusion transformer using adversarial perturbation for egocentric action anticipation
Minhyuk Kim, Jong Won Jung, Eungi Lee, Seok Bong Yoo
Expert Syst. Appl.4
2024 Child FER: Domain-Agnostic Facial Expression Recognition in Children Using a Secondary Image Diffusion Model
abstract
Facial expression recognition (FER) models often face challenges when generalizing across domains, such as different datasets and age groups. Despite the significance of this problem, FER in children (child FER) research remains relatively understudied, and such studies exhibit vulnerability to cross-domain evaluation. In response to the scarcity of child FER research, we propose a novel domain-agnostic approach for robust child FER. The architecture integrates a child-centric source-domain reconstructor and a child emotion feature-guided classifier. First, we use a secondary image diffusion model to reconstruct the image with source-domain childlike features while preserving emotion. Second, we recognize facial expressions based on domain-agnostic features from reconstructed images through a cross-attention mechanism. This approach counters performance degradation caused by domain discrepancies and improves the generalizability of child FER. We validate the proposed approach with diverse, publicly available datasets to highlight its effectiveness. The source code is available at https://github.com/st0421/Child-FER.
Eungi Lee, Eungjoo Lee 0001, Syed Muhammad Anwar, Seok Bong Yoo
ICASSP4
2024 Occluded Part-aware Graph Convolutional Networks for Skeleton-based Action Recognition
abstract
Recognizing human action is one of the most critical factors in the visual perception of robots. Specifically, skeletonbased action recognition has been actively researched to enhance recognition performance at a lower cost. However, action recognition in occlusion situations, where body parts are not visible, is still challenging.We propose an occluded part-aware graph convolutional network (OP-GCN) to address this challenge using the optimal occluded body parts. The proposed model uses an occluded part detector to identify occluded body parts within a human skeleton. It is based on an autoencoder trained on a nonoccluded human skeleton and exploits the symmetry and angular information of the skeleton. Then, we select an optimal group constructed considering the occluded body parts. Each group comprises five sets of joint nodes, focusing on the body parts, excluding the occluded ones. Finally, to enhance interaction within the selected groups, we apply an interpart association module, considering the fusion of global and local elements. The experimental results reveal that the proposed model outperforms others on the occluded datasets. These comparative experiments demonstrate the effectiveness of the study in addressing the challenge of action recognition in occlusion situations. Our code is publicly available at https://github.com/MJ-Kor/OP-GCN.
Minhyuk Kim, Seok Bong Yoo
ICRA3
2024 IntensPure: Attack Intensity-aware Secondary Domain Adaptive Diffusion for Adversarial Purification
Eungi Lee, Moon Seok Lee, Jae Hyun Yoon, Seok Bong Yoo
IJCAI4
2024 DenseSphere: Multimodal 3D object detection under a sparse point cloud based on spherical coordinate
Jong Won Jung, Jae Hyun Yoon, Seok Bong Yoo
Expert Syst. Appl.3
2024 Kernel adaptive memory network for blind video super-resolution
Jun-Seok Yun, Minhyuk Kim, Hyungil Kim, Seok Bong Yoo
Expert Syst. Appl.4
2023 Latent-OFER: Detect, Mask, and Reconstruct with Latent Vectors for Occluded Facial Expression Recognition
abstract
Most research on facial expression recognition (FER) is conducted in highly controlled environments, but its performance is often unacceptable when applied to real-world situations. This is because when unexpected objects occlude the face, the FER network faces difficulties extracting facial features and accurately predicting facial expressions. Therefore, occluded FER (OFER) is a challenging problem. Previous studies on occlusion-aware FER have typically required fully annotated facial images for training. However, collecting facial images with various occlusions and expression annotations is time-consuming and expensive. Latent-OFER, the proposed method, can detect occlusions, restore occluded parts of the face as if they were unoccluded, and recognize them, improving FER accuracy. This approach involves three steps: First, the vision transformer (ViT)based occlusion patch detector masks the occluded position by training only latent vectors from the unoccluded patches using the support vector data description algorithm. Second, the hybrid reconstruction network generates the masking position as a complete image using the ViT and convolutional neural network (CNN). Last, the expression-relevant latent vector extractor retrieves and uses expression-related information from all latent vectors by applying a CNN-based class activation map. This mechanism has a significant advantage in preventing performance degradation from occlusion by unseen objects. The experimental results on several databases demonstrate the superiority of the proposed method over state-of-the-art methods. This code is available at https://github.com/leeisack/Latent-OFER.
Isack Lee, Eungi Lee, Seok Bong Yoo
ICCV3
2022 LatentGaze: Cross-Domain Gaze Estimation Through Gaze-Aware Analytic Latent Code Manipulation
Isack Lee, Jun-Seok Yun, Hee Hyeon Kim, Youngju Na, Seok Bong Yoo
ACCV (5)5
2022 HAZE-Net: High-Frequency Attentive Super-Resolved Gaze Estimation in Low-Resolution Face Images
Jun-Seok Yun, Youngju Na, Hee Hyeon Kim, Hyungil Kim, Seok Bong Yoo
ACCV (5)5
2016 Texture enhancement for improving single-image super-resolution performance
Seok Bong Yoo, Kyuha Choi, Young Woo Jeon, Jong Beom Ra
Signal Process. Image Commun.1
2014 Blind Post-Processing for Ringing and Mosquito Artifact Reduction in Coded Videos
abstract
Block-based video-coding standards produce unwanted spatial and temporal artifacts in reconstructed videos. Among them, ringing and mosquito artifacts arise due to the quantization of high-frequency discrete cosine transform coefficients. Most of the existing artifact-reduction algorithms assume that coding information such as a standard quantization table and the corresponding quantization parameter for each block are available. In many multimedia applications, however, external video inputs are usually supplied without coding information. To effectively reduce the ringing and mosquito artifacts in a decoded input video sequence, it is necessary to control the filter strength block by block. In this paper, we present a blind block-based method to estimate the quantization amount and propose a novel post-processing algorithm based on the visibility of artifacts in terms of the human visual system using the estimated quantization amount. Experimental results demonstrate that the proposed algorithm better alleviates ringing and mosquito artifacts in various coded videos compared with the existing algorithms.
Seok Bong Yoo, Kyuha Choi, Jong Beom Ra
IEEE Trans. Circuits Syst. Video Technol.1
2014 Post-Processing for Blocking Artifact Reduction Based on Inter-Block Correlation
abstract
Block-based coding introduces an undesirable discontinuity between neighboring blocks in reconstructed images. This image degradation, referred to as blocking artifacts, arises mainly due to the loss of inter-block correlation in the quantization process of discrete cosine transform coefficients. In many multimedia broadcasting applications, such as a television, decoded video sequences suffer from blocking artifacts. In this paper, we present a novel post-processing algorithm based on increment of inter-block correlation aimed at reducing blocking artifacts. We first smooth the three lowest frequency discrete cosine transform (DCT) coefficients between neighboring blocks, in order to reduce blocking artifacts in the flat region, which are most sensitive to the human visual system. We then group each edge block and its matched blocks together and apply group-based filtering to increase the correlation between grouped blocks. This suppresses blocking artifacts in the edge region while preserving details. In addition, the algorithm is extended to reduce flickering artifacts as well as blocking artifacts in video sequences. Experimental results show that the proposed method successfully alleviates blocking artifacts in both images and videos coded with low bit-rates.
Seok Bong Yoo, Kyuha Choi, Jong Beom Ra
IEEE Trans. Multim.1
2012 Block Poisson Method and its application to large scale image editing
abstract
Poisson Image Editing(PIE) is one of the most well known methods for seamless image interpolation. However, direct application of PIE to large scale images(e.g. satellite image with billions of pixels) can result in hours of operation time. We propose a modification of PIE method, which enables the original method to be performed in smaller blocks, largely improving the operation time. The key idea of Block Poisson Method(BPM) is to prepare the boundary conditions for each block and perform PIE separately. Our algorithm can be effectively applied to satellite image stitching, especially for cases where the satellite images are significantly degraded at detectors' edge. Resulting images show that BPM performs fairly the same as the original PIE with considerably reduced operation time.
Dewey H. Lee, Seok Bong Yoo, Jong Beom Ra, Junmo Kim 0002
ICIP2
2011 Post processing for blocking artifact reduction
abstract
Since current coding standards rely on block based processing, reconstructed images include horizontal and vertical grid noise along the block boundaries. This image degradation, called blocking artifacts, is mainly caused due to the quantization process of discrete cosine transform (DCT) coefficients. In this paper, we present a novel post processing framework for blocking artifact reduction, which is based on the correction of a few quantized DCT coefficients. We first smooth inter-block DCT coefficients for the three lowest frequency (LF) ones, in order to reduce blocking artifacts in the flat region which are most sensitive to the human visual system (HVS). We then apply block based 3-D filtering to the edge region to reduce the remaining artifacts. Experimental results show that the proposed method successfully alleviates blocking artifacts in the images coded with low bit-rates.
Seok Bong Yoo, Kyuha Choi, Jong Beom Ra
ICIP1