EDBT 2026 Demo / reviewers in the wild / expert
Sangyoun Lee
dblp:65/1227
· DBLP profile ↗
88ranked-venue papers
2as first author
63since 2021 · last 2026
0000-0003-0394-6777ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 55 · 1 first-author · 43 since 2021Artificial intelligence and machine learning · 43 · 33 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MonoCLUE: Object-Aware Clustering Enhances Monocular 3D Object DetectionabstractMonocular 3D object detection offers a cost-effective solution for autonomous driving, but it suffers from the ill-posed depth and a limited field of view. These constraints lead to the lack of geometric cues and reduced accuracy in occluded or truncated scenes. While recent approaches incorporate additional depth information to address geometric ambiguity, they overlook the importance of visual cues essential for robust object recognition. In this paper, we propose MonoCLUE that enhances monocular 3D detection by leveraging both local clustering and generalized scene memory of visual features. First, we perform K-means clustering on visual features to capture distinct object-level appearance visual parts (e.g., bonnet, car roof), which improves the detection of partially visible objects. The clustered features are then propagated across the entire region to capture objects with similar appearances. Second, we construct a generalized scene memory by aggregating clustered features across images, providing consistent appearance representations that generalize scenes. This improves the consistency of object-level features, enabling stable detection across varying environments. Lastly, we integrate both local cluster features and generalized scene memory into object queries, guiding attention toward informative regions in the feature map. Exploiting an unified local clustering and generalized scene memory strategy, MonoCLUE enables robust monocular 3D detection under occlusion and limited visibility. Our proposed model achieves state-of-the-art performance on the KITTI benchmark. Sunghun Yang, Minhyeok Lee, Sangyoun Lee |
AAAI | 4 |
| 2026 | A foundational research framework for real-world abandoned object detection: train-free baseline and a standardized benchmark
Dong-Bum Kim, Deok-Hyun Ahn, Yong-Jin Jo, Haesol Park, Sangyoun Lee, Haksub Kim |
Expert Syst. Appl. | 5 |
| 2026 | Generalizing CLIP prompts for zero-shot anomaly detection
Donghyeong Kim, Suhwan Cho, Hyeonjeong Lim, Sangyoun Lee |
Pattern Recognit. | 7 |
| 2026 | Bidirectional token-masking autoencoder for Referring Image Segmentation
Minhyeok Lee, Dogyoon Lee, Suhwan Cho, Sangyoun Lee |
Pattern Recognit. | 5 |
| 2026 | GoP-Based Quality Enhancement on Video CompressionabstractWith recent increases in the demand for high-resolution video content, it has become increasingly challenging to transmit video data within the constraints of limited bandwidth. Due to the time-consuming nature of developing and disseminating new standard codecs, a large body of research has addressed improving low-quality videos through post-processing techniques. Previous studies have primarily concentrated on enhancing the quality of compressed video by addressing the temporal consistency of adjacent frames over short durations. However, these approaches often overlook specific characteristics of the video coding framework, such as notable variations in codec artifact patterns occurring at the Group of Pictures (GoP) level, which can result in considerable viewer discomfort. In this paper, we propose GoP-based Quality Enhancement (GQE), which aims to improve the quality of compressed videos by addressing issues at the GoP level. First, we present a GoP Guided Feature Propagation (GGFP) module, which addresses the root cause of the GoP level issue by propagating features from the I-frame of a different GoP to the frames currently undergoing enhancement. Then, we introduce a Temporal Aggregation (TA) module to efficiently and effectively aggregate features from the I-frame and the current frame. We extensively evaluate our model using diverse test sequences across a range of codecs, including HEVC, VP9, and AV1. Our approach not only achieves a significant reduction in the pattern shifts of GoP-level artifacts, but also demonstrates a substantial improvement in overall video quality. Chajin Shin, Hong-Goo Kang, Sangyoun Lee |
IEEE Trans. Image Process. | 4 |
| 2025 | Elevating Flow-Guided Video Inpainting with Reference GenerationabstractVideo inpainting (VI) is a challenging task that requires effective propagation of observable content across frames while simultaneously generating new content not present in the original video. In this study, we propose a robust and practical VI framework that leverages a large generative model for reference generation in combination with an advanced pixel propagation algorithm. Powered by a strong generative model, our method not only significantly enhances frame-level quality for object removal but also synthesizes new content in the missing areas based on user-provided text prompts. For pixel propagation, we introduce a one-shot pixel pulling method that effectively avoids error accumulation from repeated sampling while maintaining sub-pixel precision. To evaluate various VI methods in realistic scenarios, we also propose a high-quality VI benchmark, HQVI, comprising carefully generated videos using alpha matte composition. On public benchmarks and the HQVI dataset, our method demonstrates significantly higher visual quality and metric scores compared to existing solutions. Furthermore, it can process high-resolution videos exceeding 2K resolution with ease, underscoring its superiority for real-world applications. Suhwan Cho, Seoung Wug Oh, Sangyoun Lee, Joon-Young Lee |
AAAI | 3 |
| 2025 | Video Diffusion Models Are Strong Video InpainterabstractPropagation-based video inpainting using optical flow at the pixel or feature level has recently garnered significant attention. However, it has limitations such as the inaccuracy of optical flow prediction and the propagation of noise over time. These issues result in non-uniform noise and time consistency problems throughout the video, which are particularly pronounced when the removed area is large and involves substantial movement. To address these issues, we propose a novel First Frame Filling Video Diffusion Inpainting model (FFF-VDI). We design FFF-VDI inspired by the capabilities of pre-trained image-to-video diffusion models that can transform the first frame image into a highly natural video. To apply this to the video inpainting task, we propagate the noise latent information of future frames to fill the masked areas of the first frame's noise latent code. Next, we fine-tune the pre-trained image-to-video diffusion model to generate the inpainted video. The proposed model addresses the limitations of existing methods that rely on optical flow quality, producing much more natural and temporally consistent videos. This proposed approach is the first to effectively integrate image-to-video diffusion models into video inpainting tasks. Through various comparative experiments, we demonstrate that the proposed model can robustly handle diverse inpainting types with high quality. Minhyeok Lee, Suhwan Cho, Chajin Shin, Sunghun Yang, Sangyoun Lee |
AAAI | 6 |
| 2025 | CoCoGaussian: Leveraging Circle of Confusion for Gaussian Splatting from Defocused Imagesabstract3D Gaussian Splatting (3DGS) has attracted significant attention for its high-quality novel view rendering, inspiring research to address real-world challenges. While conventional methods depend on sharp images for accurate scene reconstruction, real-world scenarios are often affected by defocus blur due to finite depth of field, making it essential to account for realistic 3D scene representation. In this study, we propose CoCoGaussian, a Circle of Confusion-aware Gaussian Splatting that enables precise 3D scene representation using only defocused images. CoCoGaussian addresses the challenge of defocus blur by modeling the Circle of Confusion (CoC) through a physically grounded approach based on the principles of photographic defocus. Exploiting 3D Gaussians, we compute the CoC diameter from depth and learnable aperture information, generating multiple Gaussians to precisely capture the CoC shape. Furthermore, we introduce a learnable scaling factor to enhance robustness and provide more flexibility in handling unreliable depth in scenes with reflective or refractive surfaces. Experiments on both synthetic and real-world datasets demonstrate that CoCoGaussian achieves state-of-the-art performance across multiple benchmarks. Suhwan Cho, Taeoh Kim, Ho-Deok Jang, Minhyeok Lee, Geonho Cha, Dongyoon Wee, Dogyoon Lee, Sangyoun Lee |
CVPR | 9 |
| 2025 | Effective SAM Combination for Open-Vocabulary Semantic SegmentationabstractOpen-vocabulary semantic segmentation aims to assign pixel-level labels to images across an unlimited range of classes. Traditional methods address this by sequentially connecting a powerful mask proposal generator, such as the Segment Anything Model (SAM), with a pre-trained vision-language model like CLIP. But these two-stage approaches often suffer from high computational costs, memory inefficiencies. In this paper, we propose ESC-Net, a novel one-stage open-vocabulary segmentation model that leverages the SAM decoder blocks for class-agnostic segmentation within an efficient inference framework. By embedding pseudo prompts generated from image-text correlations into SAM’s promptable segmentation framework, ESC-Net achieves refined spatial aggregation for accurate mask predictions. Additionally, a Vision-Language Fusion (VLF) module enhances the final mask prediction through image and text guidance. ESC-Net and PASCAL-Context, outperforming prior methods in both efficiency and accuracy. Comprehensive ablation studies further demonstrate its robustness across challenging conditions. Minhyeok Lee, Suhwan Cho, Sunghun Yang, Heeseung Choi, Ig-Jae Kim, Sangyoun Lee |
CVPR | 7 |
| 2025 | CoMoGaussian: Continuous Motion-Aware Gaussian Splatting from Motion-Blurred Imagesabstract3D Gaussian Splatting (3DGS) has gained significant attention due to its high-quality novel view rendering, motivating research to address real-world challenges. A critical issue is the camera motion blur caused by movement during exposure, which hinders accurate 3D scene reconstruction. In this study, we propose CoMoGaussian, a Continuous Motion-Aware Gaussian Splatting that reconstructs precise 3D scenes from motion-blurred images while maintaining real-time rendering speed. Considering the complex motion patterns inherent in real-world camera movements, we predict continuous camera trajectories using neural ordinary differential equations (ODEs). To ensure accurate modeling, we employ rigid body transformations, preserving the shape and size of the object but rely on the discrete integration of sampled frames. To better approximate the continuous nature of motion blur, we introduce a continuous motion refinement (CMR) transformation that refines rigid transformations by incorporating additional learnable parameters. By revisiting fundamental camera theory and leveraging advanced neural ODE techniques, we achieve precise modeling of continuous camera trajectories, leading to improved reconstruction accuracy. Extensive experiments demonstrate state-of-the-art performance both quantitatively and qualitatively on benchmark datasets, which include a wide range of motion blur scenarios, from moderate to extreme blur. Donghyeong Kim, Dogyoon Lee, Suhwan Cho, Minhyeok Lee, Wonjoon Lee, Taeoh Kim, Dongyoon Wee, Sangyoun Lee |
ICCV | 9 |
| 2025 | CMTM: Cross-Modal Token Modulation for Unsupervised Video Object SegmentationabstractRecent advances in unsupervised video object segmentation have highlighted the potential of two-stream architectures that integrate appearance and motion cues. However, fully leveraging these complementary sources of information requires effectively modeling their interdependencies. In this paper, we introduce cross-modality token modulation, a novel approach designed to strengthen the interaction between appearance and motion cues. Our method establishes dense connections between tokens from each modality, enabling efficient intra-modal and inter-modal information propagation through relation transformer blocks. To improve learning efficiency, we incorporate a token masking strategy that addresses the limitations of relying solely on increased model complexity. Our approach achieves state-of-the-art performance across all public benchmarks, outperforming existing methods. The code is released on https://github.com/InSeokJeon/CMTM Inseok Jeon, Suhwan Cho, Minhyeok Lee, Donghyeong Kim, Sangyoun Lee |
ICIP | 9 |
| 2025 | Empower Words: DualGround for Structured Phrase and Sentence-Level Temporal GroundingabstractVideo Temporal Grounding (VTG) aims to localize temporal segments in long, untrimmed videos that align with a given natural language query. This task typically comprises two subtasks: \textit{Moment Retrieval (MR)} and \textit{Highlight Detection (HD)}. While recent advances have been progressed by powerful pretrained vision-language models such as CLIP and InternVideo2, existing approaches commonly treat all text tokens uniformly during cross-modal attention, disregarding their distinct semantic roles. To validate the limitations of this approach, we conduct controlled experiments demonstrating that VTG models overly rely on [EOS]-driven global semantics while failing to effectively utilize word-level signals, which limits their ability to achieve fine-grained temporal alignment. Motivated by this limitation, we propose DualGround, a dual-branch architecture that explicitly separates global and local semantics by routing the [EOS] token through a sentence-level path and clustering word tokens into phrase-level units for localized grounding. Our method introduces (1) token-role-aware cross modal interaction strategies that align video features with sentence-level and phrase-level semantics in a structurally disentangled manner, and (2) a joint modeling framework that not only improves global sentence-level alignment but also enhances fine-grained temporal grounding by leveraging structured phrase-aware context. This design allows the model to capture both coarse and localized semantics, enabling more expressive and context-aware video grounding. DualGround achieves state-of-the-art performance on both Moment Retrieval and Highlight Detection tasks across QVHighlights and Charades-STA benchmarks, demonstrating the effectiveness of disentangled semantic modeling in video-language alignment. Minhyeok Lee, Donghyeong Kim, Sangyoun Lee |
NeurIPS | 5 |
| 2025 | DualFocus: Depth from Focus with Spatio-Focal Dual Variational ConstraintsabstractDepth-from-Focus (DFF) enables precise depth estimation by analyzing focus cues across a stack of images captured at varying focal lengths. While recent learning-based approaches have advanced this field, they often struggle in complex scenes with fine textures or abrupt depth changes, where focus cues may become ambiguous or misleading. We present DualFocus, a novel DFF framework that leverages the focal stack’s unique gradient patterns induced by focus variation, jointly modeling focus changes over spatial and focal dimensions. Our approach introduces a variational formulation with dual constraints tailored to DFF: spatial constraints exploit gradient pattern changes across focus levels to distinguish true depth edges from texture artifacts, while focal constraints enforce unimodal, monotonic focus probabilities aligned with physical focus behavior. These inductive biases improve robustness and accuracy in challenging regions. Comprehensive experiments on four public datasets demonstrate that DualFocus consistently outperforms state-of-the-art methods in both depth accuracy and perceptual quality. Sungmin Woo, Sangyoun Lee |
NeurIPS | 2 |
| 2025 | Sparse-DeRF: Deblurred Neural Radiance Fields From Sparse ViewabstractRecent studies construct deblurred neural radiance fields (DeRF) using dozens of blurry images, which are not practical scenarios if only a limited number of blurry images are available. This paper focuses on constructing DeRF from sparse-view for more pragmatic real-world scenarios. As observed in our experiments, establishing DeRF from sparse views proves to be a more challenging problem due to the inherent complexity arising from the simultaneous optimization of blur kernels and NeRF from sparse view. Sparse-DeRF successfully regularizes the complicated joint optimization, presenting alleviated overfitting artifacts and enhanced quality on radiance fields. The regularization consists of three key components: Surface smoothness, helps the model accurately predict the scene structure utilizing unseen and additional hidden rays derived from the blur kernel based on statistical tendencies of real-world; Modulated gradient scaling, helps the model adjust the amount of the backpropagated gradient according to the arrangements of scene objects; Perceptual distillation improves the perceptual quality by overcoming the ill-posed multi-view inconsistency of image deblurring and distilling the pre-deblurred information, compensating for the lack of clean information in blurry images. We demonstrate the effectiveness of the Sparse-DeRF with extensive quantitative and qualitative experimental results by training DeRF from 2-view, 4-view, and 6-view blurry images. Dogyoon Lee, Donghyeong Kim, Minhyeok Lee, Seunghoon Lee 0008, Sangyoun Lee |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Spatio-temporal Feature-level Augmentation Vision Transformer for video-based person re-identification
Minjung Kim 0002, MyeongAh Cho, Heansung Lee 0001, Sangyoun Lee |
Pattern Recognit. | 4 |
| 2025 | Fast video anomaly detection via context-aware shortcut exploration and abnormal feature distance learning
Donghyeong Kim, MyeongAh Cho, Minjung Kim 0002, Minseok Lee, Seungwook Park, Sangyoun Lee |
Pattern Recognit. | 7 |
| 2025 | Treating Motion as Option With Output Selection for Unsupervised Video Object SegmentationabstractUnsupervised video object segmentation aims to detect the most salient object in a video without any external guidance regarding the object. Salient objects often exhibit distinctive movements compared to the background, and recent methods leverage this by combining motion cues from optical flow maps with appearance cues from RGB images. However, because optical flow maps are often closely correlated with segmentation masks, networks can become overly dependent on motion cues during training, leading to vulnerability when faced with confusing motion cues and resulting in unstable predictions. To address this challenge, we propose a novel motion-as-option network that treats motion cues as an optional component rather than a necessity. During training, we randomly input RGB images into the motion encoder instead of optical flow maps, which implicitly reduces the network’s reliance on motion cues. This design ensures that the motion encoder is capable of processing both RGB images and optical flow maps, leading to two distinct predictions depending on the type of input provided. To make the most of this flexibility, we introduce an adaptive output selection algorithm that determines the optimal prediction during testing. Code and models are available at https://github.com/suhwan-cho/TMO. Suhwan Cho, Minhyeok Lee, MyeongAh Cho, Seungwook Park, Jaeyeob Kim, Hyunsung Jang, Sangyoun Lee |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2024 | Dual Prototype Attention for Unsupervised Video Object SegmentationabstractUnsupervised video object segmentation (VOS) aims to detect and segment the most salient object in videos. The primary techniques used in unsupervised VOS are 1) the collaboration of appearance and motion information; and 2) temporal fusion between different frames. This paper proposes two novel prototype-based attention mechanisms, inter-modality attention (IMA) and inter-frame attention (IFA), to incorporate these techniques via dense propagation across different modalities and frames. IMA densely in-tegrates context information from different modalities based on a mutual refinement. IFA injects global context of a video to the query frame, enabling a full utilization of useful prop-erties from multiple frames. Experimental results on public benchmark datasets demonstrate that our proposed approach outperforms all existing methods by a substantial margin. The proposed two components are also thoroughly validated via ablative study. Code and models are available at https://github.com/Hydragon516/DPA. Suhwan Cho, Minhyeok Lee, Seunghoon Lee 0008, Dogyoon Lee, Heeseung Choi, Ig-Jae Kim, Sangyoun Lee |
CVPR | 7 |
| 2024 | Guided Slot Attention for Unsupervised Video Object SegmentationabstractUnsupervised video object segmentation aims to segment the most prominent object in a video sequence. However, the existence of complex backgrounds and multiple foreground objects make this task challenging. To address this issue, we propose a guided slot attention network to reinforce spatial structural information and obtain better foreground-background separation. The foreground and background slots, which are initialized with query guidance, are iteratively refined based on interactions with template information. Furthermore, to improve slot-template interaction and effectively fuse global and local features in the target and reference frames, K-nearest neighbors filtering and a feature aggregation transformer are introduced. The proposed model achieves state-of-the-art performance on two popular datasets. Additionally, we demonstrate the robustness of the proposed model in challenging scenes through various comparative experiments. Code and models are available at https://github.com/Hydragon516/GSANet. Minhyeok Lee, Suhwan Cho, Dogyoon Lee, Sangyoun Lee |
CVPR | 6 |
| 2024 | ProDepth: Boosting Self-supervised Multi-frame Monocular Depth with Probabilistic Fusion
Sungmin Woo, Wonjoon Lee, Woo Jin Kim, Dogyoon Lee, Sangyoun Lee |
ECCV (3) | 5 |
| 2024 | FIMP: Future Interaction Modeling for Multi-Agent Motion PredictionabstractMulti-agent motion prediction is a crucial concern in autonomous driving, yet it remains a challenge owing to the ambiguous intentions of dynamic agents and their intricate interactions. Existing studies have attempted to capture interactions between road entities by using the definite data in history timesteps, as future information is not available and involves high uncertainty. However, without sufficient guidance for capturing future states of interacting agents, they frequently produce unrealistic trajectory overlaps. In this work, we propose Future Interaction modeling for Motion Prediction (FIMP), which captures potential future interactions in an end-to-end manner. FIMP adopts a future decoder that implicitly extracts the potential future information in an intermediate feature-level, and identifies the interacting entity pairs through future affinity learning and top-k filtering strategy. Experiments show that our future interaction modeling improves the performance remarkably, leading to superior performance on the Argoverse motion forecasting benchmark. Sungmin Woo, Minjung Kim 0002, Donghyeong Kim, Sungjun Jang, Sangyoun Lee |
ICRA | 5 |
| 2024 | Towards Multi-Domain Learning for Generalizable Video Anomaly DetectionabstractMost of the existing Video Anomaly Detection (VAD) studies have been conducted within single-domain learning, where training and evaluation are performed on a single dataset. However, the criteria for abnormal events differ across VAD datasets, making it problematic to apply a single-domain model to other domains. In this paper, we propose a new task called Multi-Domain learning forVAD (MDVAD) to explore various real-world abnormal events using multiple datasets for a general model. MDVAD involves training on datasets from multiple domains simultaneously, and we experimentally observe that Abnormal Conflicts between domains hinder learning and generalization. The task aims to address two key objectives: (i) better distinguishing between general normal and abnormal events across multiple domains, and (ii) being aware of ambiguous abnormal conflicts. This paper is the first to tackle abnormal conflict issue and introduces a new benchmark, baselines, and evaluation protocols for MDVAD. As baselines, we propose a framework with Null(Angular)-Multiple Instance Learning and an Abnormal Conflict classifier. Through experiments on a MDVAD benchmark composed of six VAD datasets and using four different evaluation protocols, we reveal abnormal conflicts and demonstrate that the proposed baseline effectively handles these conflicts, showing robustness and adaptability across multiple domains. MyeongAh Cho, Taeoh Kim, Minho Shim, Dongyoon Wee, Sangyoun Lee |
NeurIPS | 5 |
| 2024 | Separate and Reconstruct: Asymmetric Encoder-Decoder for Speech SeparationabstractIn speech separation, time-domain approaches have successfully replaced the time-frequency domain with latent sequence feature from a learnable encoder. Conventionally, the feature is separated into speaker-specific ones at the final stage of the network. Instead, we propose a more intuitive strategy that separates features earlier by expanding the feature sequence to the number of speakers as an extra dimension. To achieve this, an asymmetric strategy is presented in which the encoder and decoder are partitioned to perform distinct processing in separation tasks. The encoder analyzes features, and the output of the encoder is split into the number of speakers to be separated. The separated sequences are then reconstructed by the weight-shared decoder, which also performs cross-speaker processing.
Without relying on speaker information, the weight-shared network in the decoder directly learns to discriminate features using a separation objective. In addition, to improve performance, traditional methods have extended the sequence length, leading to the adoption of dual-path models, which handle the much longer sequence effectively by segmenting it into chunks. To address this, we introduce global and local Transformer blocks that can directly handle long sequences more efficiently without chunking and dual-path processing. The experimental results demonstrated that this asymmetric structure is effective and that the combination of proposed global and local Transformer can sufficiently replace the role of inter- and intra-chunk processing in dual-path structure. Finally, the presented model combining both of these achieved state-of-the-art performance with much less computation in various benchmark datasets. Ui-Hyeop Shin, Sangyoun Lee, Taehan Kim, Hyung-Min Park |
NeurIPS | 2 |
| 2024 | A Nonlinear, Regularized, and Data-independent Modulation for Continuously Interactive Image Processing Network
Hyeongmin Lee, Taeoh Kim, Hanbin Son, Sangwook Baek, Minsu Cheon, Sangyoun Lee |
Int. J. Comput. Vis. | 6 |
| 2024 | Multi-Scale Structural Graph Convolutional Network for Skeleton-Based Action RecognitionabstractGraph convolutional networks (GCNs) have attracted considerable interest in skeleton-based action recognition. Existing GCN-based models have proposed methods to learn dynamic graph topologies generated from the feature information of vertices to capture inherent relationships. However, these models have two main limitations. Firstly, they struggle to effectively utilize high-dimensional or structural information, which limits their capacity for feature representation and consequently hinders performance improvement. Secondly, among these models, the multi-scale methods that aggregate information at different scales often over-capture unnecessary relationships between vertices. This leads to an over-smoothing problem where smoothed features are extracted, making it difficult to distinguish the features of each vertex. To address these limitations, we propose the multi-scale structural graph convolutional network (MSS-GCN) for skeleton-based action recognition. Within the MSS-GCN framework, the common intersection graph convolution (CI-GC) leverages the overlapped neighbor information, indicating the overlap between neighboring vertices for a given pair of root vertices. The graph topology of CI-GC is designed to compute the structural correlation between neighboring vertices corresponding to each hop, thereby enriching the context of inter-vertex relationships. Then, our proposed multi-scale spatio-temporal modeling aggregates local-global features to provide a comprehensive representation. In addition, we propose a Graph Weight Annealing (GWA) method, which is a graph scheduling method to mitigate the over-smoothing caused by multi-scale aggregation. By varying the importance between a vertex and its neighbors, we demonstrate that the over-smoothing problem can be effectively mitigated. Moreover, our proposed GWA method can easily be adapted to different GCN models to enhance performance. Combining the MSS-GCN model and the GWA method, we propose a powerful feature extractor that effectively classifies actions for skeleton-based action recognition in various datasets. We evaluate our approach on three benchmark datasets: NTU RGB+D, NTU RGB+D 120, and NW-UCLA. The proposed MSS-GCN achieves state-of-the-art performance on all three datasets, further validating the effectiveness of our approach. Sungjun Jang, Heansung Lee 0001, Woo Jin Kim, Sungmin Woo, Sangyoun Lee |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | LSHNet: Leveraging Structure-Prior With Hierarchical Features Updates for Salient Object Detection in Optical Remote Sensing ImagesabstractSalient object detection in optical remote sensing images (ORSI-SOD) is a task that detects the most prominent context in optical remote sensing images (ORSIs). ORSI-SOD is challenging due to the diverse sizes and shapes of the targets, their irregular distribution in the scene, and the occlusion caused by surrounding environments. Recently, many deep learning-based models have demonstrated promising performance in ORSI-SOD. However, there still remains considerable potential for addressing these challenges in ORSI-SOD. In this article, we propose a novel architecture, LSHNet. We propose a dual-branch architecture consisting of an edge encoder that leverages structure features using edges as the structure-prior and an image encoder that extracts context features from the image. We propose three modules. Image-structure fusion module (ISFM) integrates the two-stream features extracted from dual branch encoders through intrapatch and internal-patch attention mechanisms to utilize diverse receptive fields. Local-global feature fusion module (LGFM) transfers global features representing the target to local feature maps to discriminate the region of the targets from background clutters. The semantic cues updating module (SCUM) updates the representative features of the target from high-level to low-level. By integrating hierarchical information effectively, global features extracted from multilayers can be rectified. We experiment with the three main evaluated datasets in ORSI-SOD: ORSSD, EORSSD, and ORSI-4199. We demonstrate the promising results on the three datasets and analyze the effectiveness of the proposed modules in the ablation study. Seunghoon Lee 0008, Suhwan Cho, Seungwook Park, Jaeyeob Kim, Sangyoun Lee |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Look Around for Anomalies: Weakly-Supervised Anomaly Detection via Context-Motion Relational LearningabstractWeakly-supervised Video Anomaly Detection is the task of detecting frame-level anomalies using video-level labeled training data. It is difficult to explore class representative features using minimal supervision of weak labels with a single backbone branch. Furthermore, in real-world scenarios, the boundary between normal and abnormal is ambiguous and varies depending on the situation. For example, even for the same motion of running person, the abnormality varies depending on whether the surroundings are a playground or a roadway. Therefore, our aim is to extract discriminative features by widening the relative gap between classes' features from a single branch. In the proposed Class-Activate Feature Learning (CLAV), the features are extracted as per the weights that are implicitly activated depending on the class, and the gap is then enlarged through relative distance learning. Furthermore, as the relationship between context and motion is important in order to identify the anomalies in complex and diverse scenes, we propose a Context-Motion Interrelation Module (CoMo), which models the relationship between the appearance of the surroundings and motion, rather than utilizing only temporal dependencies or motion information. The proposed method shows SOTA performance on four benchmarks including large-scale real-world datasets, and we demonstrate the importance of relational information by analyzing the qualitative results and generalization ability. MyeongAh Cho, Minjung Kim 0002, Kyungjae Lee 0003, Sangyoun Lee |
CVPR | 6 |
| 2023 | DP-NeRF: Deblurred Neural Radiance Field with Physical Scene PriorsabstractNeural Radiance Field (NeRF) has exhibited outstanding three-dimensional (3D) reconstruction quality via the novel view synthesis from multi-view images and paired calibrated camera parameters. However, previous NeRF-based systems have been demonstrated under strictly controlled settings, with little attention paid to less ideal scenarios, including with the presence of noise such as exposure, illumination changes, and blur. In particular, though blur frequently occurs in real situations, NeRF that can handle blurred images has received little attention. The few studies that have investigated NeRF for blurred images have not considered geometric and appearance consistency in 3D space, which is one of the most important factors in 3D reconstruction. This leads to inconsistency and the degradation of the perceptual quality of the constructed scene. Hence, this paper proposes a DP-NeRF, a novel clean NeRF framework for blurred images, which is constrained with two physical priors. These priors are derived from the actual blurring process during image acquisition by the camera. DP-NeRF proposes rigid blurring kernel to impose 3D consistency utilizing the physical priors and adaptive weight proposal to refine the color composition error in consideration of the relationship between depth and blur. We present extensive experimental results for synthetic and real scenes with two types of blur: camera motion blur and defocus blur. The results demonstrate that DP-NeRF successfully improves the perceptual quality of the constructed NeRF ensuring 3D geometric and appearance consistency. We further demonstrate the effectiveness of our model with comprehensive ablation analysis.11Code: https://github.com/dogyoonlee/DP-NeRF22Project: https://dogyoonlee.github.io/dpNeRF/ Dogyoon Lee, Minhyeok Lee, Chajin Shin, Sangyoun Lee |
CVPR | 4 |
| 2023 | Exploring Discontinuity for Video Frame InterpolationabstractVideo frame interpolation (VFI) is the task that synthesizes the intermediate frame given two consecutive frames. Most of the previous studies have focused on appropriate frame warping operations and refinement modules for the warped frames. These studies have been conducted on natural videos containing only continuous motions. However, many practical videos contain various unnatural objects with discontinuous motions such as logos, user interfaces and subtitles. We propose three techniques that can make the existing deep learning-based VFI architectures robust to these elements. First is a novel data augmentation strategy called figure-text mixing (FTM) which can make the models learn discontinuous motions during training stage without any extra dataset. Second, we propose a simple but effective module that predicts a map called discontinuity map (D-map), which densely distinguishes between areas of continuous and discontinuous motions. Lastly, we propose loss functions to give supervisions of the discontinuous motion areas which can be applied along with FTM and D-map. We additionally collect a special test benchmark called Graphical Discontinuous Motion (GDM) dataset consisting of some mobile games and chatting videos. Applied to the various state-of-the-art VFI networks, our method significantly improves the interpolation qualities on the videos from not only GDM dataset, but also the existing benchmarks containing only continuous motions such as Vimeo90K, UCF101, and DAVIS. Hyeongmin Lee, Chajin Shin, Hanbin Son, Sangyoun Lee |
CVPR | 5 |
| 2023 | FAPM: Fast Adaptive Patch Memory for Real-Time Industrial Anomaly DetectionabstractFeature embedding-based methods have shown exceptional performance in detecting industrial anomalies by comparing features of target images with normal images. However, some methods do not meet the speed requirements of real-time inference, which is crucial for real-world applications. To address this issue, we propose a new method called Fast Adaptive Patch Memory (FAPM) for real-time industrial anomaly detection. FAPM utilizes patch-wise and layer-wise memory banks that store the embedding features of images at the patch and layer level, respectively, which eliminates unnecessary repetitive computations. We also propose patch-wise adaptive coreset sampling for faster and more accurate detection. FAPM performs well in both accuracy and speed compared to other state-of-the-art methods. Donghyeong Kim, Suhwan Cho, Sangyoun Lee |
ICASSP | 4 |
| 2023 | Two-Stream Decoder Feature Normality Estimating Network for Industrial Anomaly DetectionabstractImage reconstruction-based anomaly detection has recently been in the spotlight because of the difficulty of constructing anomaly datasets. These approaches work by learning to model normal features without seeing abnormal samples during training and then discriminating anomalies at test time based on the reconstructive errors. However, these models have limitations in reconstructing the abnormal samples due to their indiscriminate conveyance of features. Moreover, these approaches are not explicitly optimized for distinguishable anomalies. To address these problems, we propose a two-stream decoder network (TSDN), designed to learn both normal and abnormal features. Additionally, we propose a feature normality estimator (FNE) to eliminate abnormal features and prevent high-quality reconstruction of abnormal regions. Evaluation on a standard benchmark demonstrated performance better than state-of-the-art models. Minhyeok Lee, Suhwan Cho, Donghyeong Kim, Sangyoun Lee |
ICASSP | 5 |
| 2023 | Leveraging Spatio-Temporal Dependency for Skeleton-Based Action RecognitionabstractSkeleton-based action recognition has attracted considerable attention due to its compact representation of the human body’s skeletal sructure. Many recent methods have achieved remarkable performance using graph convolutional networks (GCNs) and convolutional neural networks (CNNs), which extract spatial and temporal features, respectively. Although spatial and temporal dependencies in the human skeleton have been explored separately, spatio-temporal dependency is rarely considered. In this paper, we propose the Spatio-Temporal Curve Network (STC-Net) to effectively leverage the spatio-temporal dependency of the human skeleton. Our proposed network consists of two novel elements: 1) The Spatio-Temporal Curve (STC) module; and 2) Dilated Kernels for Graph Convolution (DK-GC). The STC module dynamically adjusts the receptive field by identifying meaningful node connections between every adjacent frame and generating spatio-temporal curves based on the identified node connections, providing an adaptive spatio-temporal coverage. In addition, we propose DK-GC to consider long-range dependencies, which results in a large receptive field without any additional parameters by applying an extended kernel to the given adjacency matrices of the graph. Our STC-Net combines these two modules and achieves state-of-the-art performance on four skeleton-based action recognition benchmarks. Code is available at https://github.com/Jho-Yonsei/STC-Net. Minhyeok Lee, Suhwan Cho, Sungmin Woo, Sungjun Jang, Sangyoun Lee |
ICCV | 6 |
| 2023 | Hierarchically Decomposed Graph Convolutional Networks for Skeleton-Based Action RecognitionabstractGraph convolutional networks (GCNs) are the most commonly used methods for skeleton-based action recognition and have achieved remarkable performance. Generating adjacency matrices with semantically meaningful edges is particularly important for this task, but extracting such edges is challenging problem. To solve this, we propose a hierarchically decomposed graph convolutional network (HD-GCN) architecture with a novel hierarchically decomposed graph (HD-Graph). The proposed HD-GCN effectively decomposes every joint node into several sets to extract major structurally adjacent and distant edges, and uses them to construct an HD-Graph containing those edges in the same semantic spaces of a human skeleton. In addition, we introduce an attention-guided hierarchy aggregation (A-HA) module to highlight the dominant hierarchical edge sets of the HD-Graph. Furthermore, we apply a new six-way ensemble method, which uses only joint and bone stream without any motion stream. The proposed model is evaluated and achieves state-of-the-art performance on four large, popular datasets. Finally, we demonstrate the effectiveness of our model with various comparative experiments. Code is available at https://github.com/Jho-Yonsei/HD-GCN. Minhyeok Lee, Dogyoon Lee, Sangyoun Lee |
ICCV | 4 |
| 2023 | Tsanet: Temporal and Scale Alignment for Unsupervised Video Object SegmentationabstractUnsupervised Video Object Segmentation (UVOS) refers to the challenging task of segmenting the prominent object in videos without manual guidance. In recent works, two approaches for UVOS have been discussed that can be divided into: appearance and appearance-motion-based methods, which have limitations respectively. Appearance-based methods do not consider the motion of the target object due to exploiting the correlation information between randomly paired frames. Appearance-motion-based methods have the limitation that the dependency on optical flow is dominant due to fusing the appearance with motion. In this paper, we propose a novel framework for UVOS that can address the aforementioned limitations of the two approaches in terms of both time and scale. Temporal Alignment Fusion aligns the saliency information of adjacent frames with the target frame to leverage the information of adjacent frames. Scale Alignment Decoder predicts the target object mask by aggregating multi-scale feature maps via continuous mapping with implicit neural representation. We present experimental results on public benchmark datasets, DAVIS 2016 and FBMS, which demonstrate the effectiveness of our method. Furthermore, we outperform the state-of-the-art methods on DAVIS 2016. Seunghoon Lee 0008, Suhwan Cho, Dogyoon Lee, Minhyeok Lee, Sangyoun Lee |
ICIP | 5 |
| 2023 | Adaptive Graph Convolution Module for Salient Object DetectionabstractSalient object detection (SOD) is a task that involves identifying and segmenting the most visually prominent object in an image. Existing solutions can accomplish this using a multi-scale feature fusion mechanism to detect the global context of an image. However, as there is no consideration of the structures in the image nor the relations between distant pixels, conventional methods cannot deal with complex scenes effectively. In this paper, we propose an adaptive graph convolution module (AGCM) to overcome these limitations. Prototype features are initially extracted from the input image using a learnable region generation layer that spatially groups features in the image. The prototype features are then refined by propagating information between them based on a graph architecture, where each feature is regarded as a node. Experimental results show that the proposed AGCM dramatically improves the SOD performance both quantitatively and quantitatively. Minhyeok Lee, Suhwan Cho, Sangyoun Lee |
ICIP | 4 |
| 2023 | MSV-RGNN: Multiscale Voxel Graph Neural Network for 3D Object DetectionabstractThis paper proposes a two-stage 3D object detection framework, multiscale voxel graph neural network (MSV-RGNN) which aims to fully exploit multiple scale graph features by establishing global and local relationships between voxel features at different 3D convolutional neural network (CNN) layers. In contrast to conventional graph-based methods, our proposed multiscale-voxel-graph region-of-interest (RoI) pooling module constructs graphs across diverse voxel resolutions to obtain geometric structure information on voxel features. Initially, our multiscale-voxel-graph RoI pooling module sample voxel center points with voxel-wise feature vectors and 3D region proposals from backbone network. Subsequently, graphs are constructed at different scales and graph features are aggregated for second-stage refinement. The experimental results demonstrate the potential of using multiscale graphs across different voxel resolutions for 3D object detection, achieving decent experimental results with state-of-the-art methods. Wonjoon Lee, Sungmin Woo, Donghyeong Kim, Sangyoun Lee |
ICIP | 4 |
| 2023 | Exploring Temporally Dynamic Data Augmentation for Video Recognition
Taeoh Kim, Jinhyung Kim, Minho Shim, Sangdoo Yun, Myunggu Kang, Dongyoon Wee, Sangyoun Lee |
ICLR | 7 |
| 2023 | Treating Motion as Option to Reduce Motion Dependency in Unsupervised Video Object SegmentationabstractUnsupervised video object segmentation (VOS) aims to detect the most salient object in a video sequence at the pixel level. In unsupervised VOS, most state-of-the-art methods leverage motion cues obtained from optical flow maps in addition to appearance cues to exploit the property that salient objects usually have distinctive movements compared to the background. However, as they are overly dependent on motion cues, which may be unreliable in some cases, they cannot achieve stable prediction. To reduce this motion dependency of existing two-stream VOS methods, we propose a novel motion-as-option network that optionally utilizes motion cues. Additionally, to fully exploit the property of the proposed network that motion is not always required, we introduce a collaborative network learning strategy. On all the public benchmark datasets, our proposed network affords state-of-the-art performance with real-time inference speed. Code and models are available at https://github.com/suhwan-cho/TMO. Suhwan Cho, Minhyeok Lee, Seunghoon Lee 0008, Donghyeong Kim, Sangyoun Lee |
WACV | 6 |
| 2023 | Feature Disentanglement Learning with Switching and Aggregation for Video-based Person Re-IdentificationabstractIn video person re-identification (Re-ID), the network must consistently extract features of the target person from successive frames. Existing methods tend to focus only on how to use temporal information, which often leads to networks being fooled by similar appearances and same backgrounds. In this paper, we propose a Disentanglement and Switching and Aggregation Network (DSANet), which segregates the features representing identity and features based on camera characteristics, and pays more attention to ID information. We also introduce an auxiliary task that utilizes a new pair of features created through switching and aggregation to increase the network’s capability for various camera scenarios. Furthermore, we devise a Target Localization Module (TLM) that extracts robust features against a change in the position of the target according to the frame flow and a Frame Weight Generation (FWG) that reflects temporal information in the final representation. Various loss functions for disentanglement learning are designed so that each component of the network can cooperate while satisfactorily performing its own role. Quantitative and qualitative results from extensive experiments demonstrate the superiority of DSANet over state-of-the-art methods on three benchmark datasets. Minjung Kim 0002, MyeongAh Cho, Sangyoun Lee |
WACV | 3 |
| 2023 | Unsupervised Video Object Segmentation via Prototype Memory NetworkabstractUnsupervised video object segmentation aims to segment a target object in the video without a ground truth mask in the initial frame. This challenging task requires extracting features for the most salient common objects within a video sequence. This difficulty can be solved by using motion information such as optical flow, but using only the information between adjacent frames results in poor connectivity between distant frames and poor performance. To solve this problem, we propose a novel prototype memory network architecture. The proposed model effectively extracts the RGB and motion information by extracting superpixel-based component prototypes from the input RGB images and optical flow maps. In addition, the model scores the usefulness of the component prototypes in each frame based on a self-learning algorithm and adaptively stores the most useful prototypes in memory and discards obsolete proto-types. We use the prototypes in the memory bank to predict the next query frame’s mask, which enhances the association between distant frames to help with accurate mask prediction. Our method is evaluated on three datasets, achieving state-of-the-art performance. We prove the effectiveness of the proposed model with various ablation studies. Minhyeok Lee, Suhwan Cho, Seunghoon Lee 0008, Sangyoun Lee |
WACV | 5 |
| 2023 | MKConv: Multidimensional feature representation for point cloud analysisabstractDespite the remarkable success of deep learning , an optimal convolution operation on point clouds remains elusive owing to their irregular data structure . Existing methods mainly focus on designing an effective continuous kernel function that can handle an arbitrary point in continuous space. Various approaches exhibiting high performance have been proposed, but we observe that the standard pointwise feature is represented by 1D channels and can become more informative when its representation involves additional spatial feature dimensions. In this paper, we present Multidimensional Kernel Convolution (MKConv), a novel convolution operator that learns to transform the point feature representation from a vector to a multidimensional matrix. Unlike standard point convolution, MKConv proceeds via two steps. (i) It first activates the spatial dimensions of local feature representation by exploiting multidimensional kernel weights. These spatially expanded features can represent their embedded information through spatial correlation as well as channel correlation in feature space , carrying more detailed local structure information. (ii) Then, discrete convolutions are applied to the multidimensional features which can be regarded as a grid-structured matrix. In this way, we can utilize the discrete convolutions for point cloud data without voxelization that suffers from information loss. Furthermore, we propose a spatial attention module, Multidimensional Local Attention (MLA), to provide comprehensive structure awareness within the local point set by reweighting the spatial feature dimensions. We demonstrate that MKConv has excellent applicability to point cloud processing tasks including object classification, object part segmentation, and scene semantic segmentation with superior results. Sungmin Woo, Dogyoon Lee, Woo Jin Kim, Sangyoun Lee |
Pattern Recognit. | 5 |
| 2022 | Tackling Background Distraction in Video Object Segmentation
Suhwan Cho, Heansung Lee 0001, Minhyeok Lee, Sungjun Jang, Minjung Kim 0002, Sangyoun Lee |
ECCV (22) | 7 |
| 2022 | SPSN: Superpixel Prototype Sampling Network for RGB-D Salient Object Detection
Minhyeok Lee, Suhwan Cho, Sangyoun Lee |
ECCV (29) | 4 |
| 2022 | Expanded Adaptive Scaling Normalization for End to End Image Compression
Chajin Shin, Hyeongmin Lee, Hanbin Son, Dogyoon Lee, Sangyoun Lee |
ECCV (17) | 6 |
| 2022 | Occluded Person Re-Identification Via Relational Adaptive Feature Correction LearningabstractOccluded person re-identification (Re-ID) in images captured by multiple cameras is challenging because the target person is occluded by pedestrians or objects, especially in crowded scenes. In addition to the processes performed during holistic person Re-ID, occluded person Re-ID involves the removal of obstacles and the detection of partially visible body parts. Most existing methods utilize the off-the-shelf pose or parsing networks as pseudo labels, which are prone to error. To address these issues, we propose a novel Occlusion Correction Network (OCNet) that corrects features through relational-weight learning and obtains diverse and representative features without using external networks. In addition, we present a simple concept of a center feature in order to provide an intuitive solution to pedestrian occlusion scenarios. Furthermore, we suggest the idea of Separation Loss (SL) for focusing on different parts between global features and part features. We conduct extensive experiments on five challenging benchmark datasets for occluded and holistic Re-ID tasks to demonstrate that our method achieves superior performance to state-of-the-art methods especially on occluded scene. Minjung Kim 0002, MyeongAh Cho, Heansung Lee 0001, Suhwan Cho, Sangyoun Lee |
ICASSP | 5 |
| 2022 | Detection-Identification Balancing Margin Loss for One-Stage Multi-Object TrackingabstractIn recent years, one-stage multi-object tracking (MOT) methods, which jointly learn detection and identification in a single network, have attracted extensive attention, due to their efficiency. However, the negative transfer effects caused by the two conflicting objectives of detection and identification have rarely been explored. In this paper, we propose a Detection-Identification Balancing Margin (DIM) loss for minimizing the adverse effects caused by these two different objectives. The proposed DIM loss consists of Detection Margin (DM) loss and Identification Margin (IM) loss. DM loss forces features that are farther from the center of the foreground features than the defined margin due to identification learning to be converged to ensure accurate detection. IM loss enables the various feature representations that are essential for identification by intentionally spreading features that become overly clustered due to detection learning. The proposed DIM loss demonstrates competitive and balanced performance for MOT by providing a positive transfer for features that had a strong negative impact on detection and identification, respectively. (HOTA 61.5, MOTA 75.3, IDF1 75.6 on MOT16, and real-time rates of 25.9 fps were achieved) Heansung Lee 0001, Suhwan Cho, Sungjun Jang, Sungmin Woo, Sangyoun Lee |
ICIP | 6 |
| 2022 | Superpixel Group-Correlation Network for Co-Saliency DetectionabstractCo-saliency detection is a task to segment the occurring salient objects in a group of images. The biggest challenges are distracting objects in the background and ambiguity between the foreground and background. To handle these issues, we propose a novel superpixel group-correlation network (SGCN) architecture that uses a superpixel algorithm to obtain various component features from a group of images and creates a group-correlation matrix to detect the common components of those images. In this way, non-common objects can be effectively excluded from consideration, enabling a clear distinction between foreground and background. Our method outperforms current state-of-the-art methods on three popular benchmark datasets for co-saliency detection, and our extensive experiments thoroughly validate our claimed contributions. Minhyeok Lee, Suhwan Cho, Sangyoun Lee |
ICIP | 4 |
| 2022 | Saliency Detection via Global Context Enhanced Feature Fusion and Edge Weighted LossabstractUNet-based methods have shown outstanding performance in salient object detection (SOD), but are problematic in two aspects. 1) Indiscriminately integrating the encoder feature, which contains spatial information for multiple objects, and the decoder feature, which contains global information of the salient object, is likely to convey unnecessary details of non-salient objects to the decoder, hindering saliency detection. 2) To deal with ambiguous object boundaries and generate accurate saliency maps, the model needs additional branches, such as edge reconstructions, which leads to increasing computational cost. To address the problems, we propose a context fusion decoder network (CFDN) and near edge weighted loss (NEWLoss) function. The CFDN creates an accurate saliency map by integrating global context information and thus suppressing the influence of the unnecessary spatial information. NEWLoss accelerates learning of obscure boundaries without additional modules by generating weight maps on object boundaries. Our method is evaluated on four benchmarks and achieves state-of-the-art performance. We prove the effectiveness of the proposed method through comparative experiments. Minhyeok Lee, MyeongAh Cho, Sangyoun Lee |
ICIP | 4 |
| 2022 | Pixel-Level Bijective Matching for Video Object SegmentationabstractSemi-supervised video object segmentation (VOS) aims to track the designated objects present in the initial frame of a video at the pixel level. To fully exploit the appearance information of an object, pixel-level feature matching is widely used in VOS. Conventional feature matching runs in a surjective manner, i.e., only the best matches from the query frame to the reference frame are considered. Each location in the query frame refers to the optimal location in the reference frame regardless of how often each reference frame location is referenced. This works well in most cases and is robust against rapid appearance variations, but may cause critical errors when the query frame contains background distractors that look similar to the target object. To mitigate this concern, we introduce a bijective matching mechanism to find the best matches from the query frame to the reference frame and vice versa. Before finding the best matches for the query frame pixels, the optimal matches for the reference frame pixels are first considered to prevent each reference frame pixel from being overly referenced. As this mechanism operates in a strict manner, i.e., pixels are connected if and only if they are the sure matches for each other, it can effectively eliminate background distractors. In addition, we propose a mask embedding module to improve the existing mask propagation method. By embedding multiple historic masks with coordinate information, it can effectively capture the position information of a target object. Code and models are available at https://github.com/suhwan-cho/BMVOS. Suhwan Cho, Heansung Lee 0001, Minjung Kim 0002, Sungjun Jang, Sangyoun Lee |
WACV | 5 |
| 2022 | EdgeConv with Attention Module for Monocular Depth EstimationabstractMonocular depth estimation is an especially important task in robotics and autonomous driving, where 3D structural information is essential. However, extreme lighting conditions and complex surface objects make it difficult to predict depth in a single image. Therefore, to generate accurate depth maps, it is important for the model to learn structural information about the scene. We propose a novel Patch-Wise EdgeConv Module (PEM) and EdgeConv Attention Module (EAM) to solve the difficulty of monocular depth estimation. The proposed modules extract structural information by learning the relationship between image patches close to each other in space using edge convolution. Our method is evaluated on two popular datasets, the NYU Depth V2 and the KITTI Eigen split, achieving state-of-the-art performance. We prove that the proposed model predicts depth robustly in challenging scenes through various comparative experiments. Minhyeok Lee, Sangyoun Lee |
WACV | 4 |
| 2022 | Robust Lane Detection via Expanded Self AttentionabstractThe image-based lane detection algorithm is one of the key technologies in autonomous vehicles. Modern deep learning methods achieve high performance in lane detection, but it is still difficult to accurately detect lanes in challenging situations such as congested roads and extreme lighting conditions. To be robust on these challenging situations, it is important to extract global contextual information even from limited visual cues. In this paper, we propose a simple but powerful self-attention mechanism optimized for lane detection called the Expanded Self Attention (ESA) module. Inspired by the simple geometric structure of lanes, the proposed method predicts the confidence of a lane along the vertical and horizontal directions in an image. The prediction of the confidence enables estimating occluded locations by extracting global contextual information. ESA module can be easily implemented and applied to any encoder-decoder-based model without increasing the inference time. The performance of our method is evaluated on three popular lane detection benchmarks (TuSimple, CULane and BDD100K). We achieve state-of-the-art performance in CULane and BDD100K and distinct improvement on TuSimple dataset. The experimental results show that our approach is robust to occlusion and extreme lighting conditions. Minhyeok Lee, Junhyeop Lee, Dogyoon Lee, Woo Jin Kim, Sangyoun Lee |
WACV | 6 |
| 2022 | FastAno: Fast Anomaly Detection via Spatio-temporal Patch TransformationabstractVideo anomaly detection has gained significant attention due to the increasing requirements of automatic monitoring for surveillance videos. Especially, the prediction based approach is one of the most studied methods to detect anomalies by predicting frames that include abnormal events in the test set after learning with the normal frames of the training set. However, a lot of prediction networks are computationally expensive owing to the use of pre-trained optical flow networks, or fail to detect abnormal situations because of their strong generative ability to predict even the anomalies. To address these shortcomings, we propose spatial rotation transformation (SRT) and temporal mixing transformation (TMT) to generate irregular patch cuboids within normal frame cuboids in order to enhance the learning of normal features. Additionally, the proposed patch transformation is used only during the training phase, allowing our model to detect abnormal frames at fast speed during inference. Our model is evaluated on three anomaly detection benchmarks, achieving competitive accuracy and surpassing all the previous works in terms of speed. MyeongAh Cho, Minhyeok Lee, Sangyoun Lee |
WACV | 4 |
| 2022 | Unsupervised video anomaly detection via normalizing flows with implicit latent features
MyeongAh Cho, Taeoh Kim, Woo Jin Kim, Suhwan Cho, Sangyoun Lee |
Pattern Recognit. | 5 |
| 2022 | SSAT: Self-Supervised Associating Network for Multiobject TrackingabstractMulti-object tracking (MOT), which is crucial for computer vision and video processing, has immense potential for improvement. Traditional tracking-by-detection approaches include feature-based object re-identification methods that use trained features, but these methods suffer from a lack of suitable training data. In training datasets used for MOT, every object in a video sequence must have its own location and ID. However, assigning IDs to each object in every sequence is considerably labor-intensive, and hence current MOT datasets are unsuitable for training re-identification networks. To resolve this issue, this paper proposes a novel self-supervised learning method using several short videos that contain no human-added labels, based on the idea that each video is a set of temporally corresponding image frames. We then describe how to improve tracking performance using a re-identification network trained in a self-supervised manner. In addition, ablation studies were conducted in order to define the optimal parameters, such as number of clips, data augmentation, and appropriate matching algorithms. The proposed approach achieved competitive performance compared with current best-practice methods including supervised methods, achieving MOT accuracy = 62.0% and ID F1-score = 62.7% on the MOT17 benchmark. Tae-Young Chung, MyeongAh Cho, Heansung Lee 0001, Sangyoun Lee |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Geometry-Aware Deep Video Deblurring via Recurrent Feature RefinementabstractBlurring in videos is a frequent phenomenon in real-world video data owing to camera shake or object movement at different scene depths. Hence, video deblurring is an ill-posed problem that requires understanding of geometric and temporal information. Traditional model-based optimization methods first define a degradation model and then solve an optimization problem to recover the latent frames with a variational model for additional external information, such as optical flow, segmentation, depth, or camera movement. Recent deep-learning-based approaches learn from numerous training pairs of blurred and clean latent frames, with the powerful representation ability of deep convolutional neural networks. Although deep models have achieved remarkable performances without the explicit model, existing deep methods do not utilize geometrical information as strong priors. Therefore, they cannot handle extreme blurring caused by large camera shake or scene depth variations. In this paper, we propose a geometry-aware deep video deblurring method via a recurrent feature refinement module that exploits optimization-based and deep-learning-based schemes. In addition to the off-the-shelf deep geometry estimation modules, we design an effective fusion module for geometrical information with deep video features. Specifically, similar to model-based optimization, our proposed module recurrently refines video features as well as geometrical information to restore more precise latent frames. To evaluate the effectiveness and generalization of our framework, we perform tests on eight baseline networks whose structures are motivated by the previous research. The experimental results show that our framework offers greater performances than the eight baselines and produces state-of-the-art performance on four video deblurring benchmark datasets. Taeoh Kim, Sangyoun Lee |
IEEE Trans. Image Process. | 2 |
| 2022 | Enhanced Standard Compatible Image Compression Framework Based on Auxiliary Codec NetworksabstractRecent deep neural network-based research to enhance image compression performance can be divided into three categories: learnable codecs, postprocessing networks, and compact representation networks. The learnable codec has been designed for end-to-end learning beyond the conventional compression modules. The postprocessing network increases the quality of decoded images using example-based learning. The compact representation network is learned to reduce the capacity of an input image, reducing the bit rate while maintaining the quality of the decoded image. However, these approaches are not compatible with existing codecs or are not optimal for increasing coding efficiency. Specifically, it is difficult to achieve optimal learning in previous studies using a compact representation network due to the inaccurate consideration of the codecs. In this paper, we propose a novel standard compatible image compression framework based on auxiliary codec networks (ACNs). In addition, ACNs are designed to imitate image degradation operations of the existing codec, which delivers more accurate gradients to the compact representation network. Therefore, compact representation and postprocessing networks can be learned effectively and optimally. We demonstrate that the proposed framework based on the JPEG and High Efficiency Video Coding standard substantially outperforms existing image compression algorithms in a standard compatible manner. Hanbin Son, Taeoh Kim, Hyeongmin Lee, Sangyoun Lee |
IEEE Trans. Image Process. | 4 |
| 2022 | LiDAR Depth Completion Using Color-Embedded Information via Knowledge DistillationabstractDepth completion is the task of reconstructing dense depth images from sparse LiDAR data. LiDAR depth completion, for which LiDAR data is the only input, is an ill-posed and challenging problem owing to the underlying properties of LiDAR data: extremely few points, presence of discontinuities, and absence of texture information. Accordingly, most approaches are heavily dependent on guided color images, which leads to unsatisfactory results when the color images are degraded. To alleviate the dependency on color images but leverage this information during training, we present a deep convolutional neural network (CNN) consisting of depth and edge CNNs via transferring of knowledge. In order to compensate for the limitations of LiDAR data, we design the edge CNN to learn a gradient depth image from a powerful teacher network through theKnowledge-Distillationmethod. Since the teacher network is trained with color images, color-embedded information can be obtained in the test phase even if color images are not used as an input. We further propose aSelf-Distillationmethod for transferring the color-embedded features from the edge CNN to the depth CNN. Enforcing the depth features to contain edge information hardly observed in LiDAR data enables the depth CNN to generate more edge-attentive and structure-preserving results. Our novel methods show remarkable results in outdoor and indoor environments for KITTI and NYU-Depth-V2 datasets. Experiments performed with low-channel LiDAR data in KITTI and few depth points in the NYU-Depth-V2 dataset show that our method is robust to data sparsity and applicable in various scenarios. Junhyeop Lee, Woo Jin Kim, Sungmin Woo, Kyungjae Lee 0003, Sangyoun Lee |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2022 | AIBM: Accurate and Instant Background Modeling for Moving Object DetectionabstractDetecting moving objects has been widely studied since it plays vital many applications, such as video surveillance and intelligent transportation systems. It is necessary to accurately differentiate the foreground and the background in this technology to analyze object motions in the scene. Conventional detection methods use many reference frames to model the background to detect moving objects; however, the detection is inaccurate when immediate changes occur in the scene because the instant update of the background model is impossible. To be robust illumination changes and dynamic backgrounds, we propose an accurate and instant background modeling (AIBM) method that inpaints the background with superpixels and enhances it in detail with pixel-levels. Unlike the previous approaches, the proposed AIBM method utilizes spatio-temporal information of only two consecutive frames to eliminate the lengthy initialization and update period of the model. In this paper, we illustrate the importance of accurate and instant background modeling in detecting moving objects. The performance of our method is evaluated with three benchmark datasets (CDnet2014, LASIESTA, and SBI). The experimental results show that our AIBM method is robust to sudden changes in the scene and outperforms the other conventional methods with F-measure of 0.8911 and 0.9682 in detection accuracy in CDNet2014 and LASIESTA datasets, respectively. The accuracy of the generated background is also measured on the SBI dataset, which demonstrates the importance of high-quality background modeling. Woo Jin Kim, Junhyeop Lee, Sungmin Woo, Sangyoun Lee |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | SAM-Net: LiDAR Depth Inpainting for 3D Static Map GenerationabstractVarious sensors can be attached and added to autonomous vehicles, included visual cameras, radar, LiDAR (Light Detection And Ranging), and GNSS (Global Navigation Satellite System). These sensors have been studied in many research areas, in particular studies on building precise 3D maps. It is essential for autonomous driving to create an accurate 3D map of the surrounding scene. However, creating an accurate static 3D map is difficult due to changes in moving objects or dynamic environments. Spurious objects on the 3D map can be handled by removing or ignoring them for 3D mapping. Following this idea, we propose an object segmentation and inpainting network. The proposed network called SAM-Net, addresses the object duplication issue by segmenting the objects and inpainting them with the segmentation results. Conventional inpainting research has dealt with RGB images. No matter how well such approaches reconstruct holes or corrupted images, they do not establish 3D points’ relationship with the point cloud frame. Therefore, we suggest a depth inpainting method for outdoor object segmentation and inpainting tasks that utilizes a high-precision depth range sensor (Velodyne HDL-64E), which is not suggested before. Unfortunately, no dataset exists for the outdoor depth inpainting task. Thus, to train our model, we generate a new dataset by locating objects on a clean static background. Moreover, our proposed method shows outstanding depth performance compared to the previous visual inpainting method. Our dataset will be available at: “https://github.com/JunhyeopLee/lidar_inpainting”. Junhyeop Lee, Woo Jin Kim, Sangyoun Lee |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | Regularization Strategy for Point Cloud via Rigidly Mixed SampleabstractData augmentation is an effective regularization strategy to alleviate the overfitting, which is an inherent drawback of the deep neural networks. However, data augmentation is rarely considered for point cloud processing despite many studies proposing various augmentation methods for image data. Actually, regularization is essential for point clouds since lack of generality is more likely to occur in point cloud due to small datasets. This paper proposes a Rigid Subset Mix (RSMix)1, a novel data augmentation method for point clouds that generates a virtual mixed sample by replacing part of the sample with shape-preserved subsets from another sample. RSMix preserves structural information of the point cloud sample by extracting subsets from each sample without deformation using a neighboring function. The neighboring function was carefully designed considering unique properties of point cloud, unordered structure and non-grid. Experiments verified that RSMix successfully regularized the deep neural networks with remarkable improvement for shape classification. We also analyzed various combinations of data augmentations including RSMix with single and multi-view evaluations, based on abundant ablation studies. Dogyoon Lee, Jaeha Lee, Junhyeop Lee, Hyeongmin Lee, Minhyeok Lee, Sungmin Woo, Sangyoun Lee |
CVPR | 7 |
| 2021 | Test-Time Adaptation for Out-Of-Distributed Image InpaintingabstractDeep-learning-based image inpainting algorithms have shown great performance via powerful learned priors from numerous external natural images. However, they show unpleasant results for test images whose distributions are far from those of the training images because their models are biased toward the training images. In this paper, we propose a simple image inpainting algorithm with test-time adaptation named AdaFill. Given a single out-of-distributed test image, our goal is to complete hole region more naturally than the pre-trained inpainting models. To achieve this goal, we treat the remaining valid regions of the test image as an another training cue because natural images have strong internal similarities. From this test-time adaptation, our network can exploit externally learned image priors from the pre-trained features as well as the internal priors of the test image explicitly. The experimental results show that AdaFill outperforms other models on various out-of-distribution test images. Furthermore, the model named ZeroFill, which is not pre-trained also outperforms the pre-trained models sometimes. Chajin Shin, Taeoh Kim, Sangyoun Lee |
ICIP | 4 |
| 2021 | A Heterogeneous Face Recognition Via Part Adaptive And Relation Attention ModuleabstractIn the face recognition application scenario, we need to process facial images captured in various conditions, such as at night by near-infrared (NIR) surveillance cameras. The illumination difference between NIR and visible-light (VIS) images causes a domain gap, and the variations in pose and emotion also make facial matching more difficult. Since heterogeneous face recognition (HFR) has difficulties in domain discrepancy, many studies have focused on extracting domain-invariant features, such as facial part relational information. However, when pose variation occurs, the facial component position changes and a different part relation is extracted. In this paper, we propose a part relation attention module that crops facial parts obtained through a semantic mask and performs relational modeling using each of these representative features. Furthermore, we suggest component adaptive triplet loss using adaptive weights for each part to reduce the intra-class distance regardless of the domain as well as pose. Finally, our method exhibits a performance improvement in the CASIA NIR-VIS 2.0 [1] and achieves superior results in the BUAA-VisNir [2] with large pose and emotion variations. Rushuang Xu, MyeongAh Cho, Sangyoun Lee |
ICIP | 3 |
| 2021 | Relational Deep Feature Learning for Heterogeneous Face RecognitionabstractHeterogeneous Face Recognition (HFR) is a task that matches faces across two different domains such as visible light (VIS), near-infrared (NIR), or the sketch domain. Due to the lack of databases, HFR methods usually exploit the pre-trained features on a large-scale visual database that contain general facial information. However, these pre-trained features cause performance degradation due to the texture discrepancy with the visual domain. With this motivation, we propose a graph-structured module called Relational Graph Module (RGM) that extracts global relational information in addition to general facial features. Because each identity's relational information between intra-facial parts is similar in any modality, the modeling relationship between features can help cross-domain matching. Through the RGM, relation propagation diminishes texture dependency without losing its advantages from the pre-trained features. Furthermore, the RGM captures global facial geometrics from locally correlated convolutional features to identify long-range relationships. In addition, we propose a Node Attention Unit (NAU) that performs node-wise recalibration to concentrate on the more informative nodes arising from relation-based propagation. Furthermore, we suggest a novel conditional-margin loss function ($C$ -softmax) for the efficient projection learning of the embedding vector in HFR. The proposed method outperforms other state-of-the-art methods on five HFR databases. Furthermore, we demonstrate performance improvement on three backbones because our module can be plugged into any pre-trained face recognition backbone to overcome the limitations of a small HFR database. MyeongAh Cho, Taeoh Kim, Ig-Jae Kim, Kyungjae Lee 0003, Sangyoun Lee |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2020 | AdaCoF: Adaptive Collaboration of Flows for Video Frame InterpolationabstractVideo frame interpolation is one of the most challenging tasks in video processing research. Recently, many studies based on deep learning have been suggested. Most of these methods focus on finding locations with useful information to estimate each output pixel using their own frame warping operations. However, many of them have Degrees of Freedom (DoF) limitations and fail to deal with the complex motions found in real world videos. To solve this problem, we propose a new warping module named Adaptive Collaboration of Flows (AdaCoF). Our method estimates both kernel weights and offset vectors for each target pixel to synthesize the output frame. AdaCoF is one of the most generalized warping modules compared to other approaches, and covers most of them as special cases of it. Therefore, it can deal with a significantly wide domain of complex motions. To further improve our framework and synthesize more realistic outputs, we introduce dual-frame adversarial loss which is applicable only to video frame interpolation tasks. The experimental results show that our method outperforms the state-of-the-art methods for both fixed training set environments and the Middlebury benchmark. Our source code is available at https://github.com/HyeongminLEE/AdaCoF-pytorch Hyeongmin Lee, Taeoh Kim, Tae-Young Chung, Daehyun Pak, Yuseok Ban, Sangyoun Lee |
CVPR | 6 |
| 2020 | Crvos: Clue Refining Network For Video Object SegmentationabstractThe encoder-decoder based methods for semi-supervised video object segmentation (Semi-VOS) have received extensive attention due to their superior performances. However, most of them have complex intermediate networks which generate strong specifiers to be robust against challenging scenarios, and this is quite inefficient when dealing with relatively simple scenarios. To solve this problem, we propose a real-time network, Clue Refining Network for Video Object Segmentation (CRVOS), that does not have any intermediate network to efficiently deal with these scenarios. In this work, we propose a simple specifier, referred to as the Clue, which consists of the previous frame’s coarse mask and coordinates information. We also propose a novel refine module which shows the better performance compared with the general ones by using a deconvolution layer instead of a bilinear upsampling layer. Our proposed method shows the fastest speed among the existing methods with a competitive accuracy. On DAVIS 2016 validation set, our method achieves 63.5 fps and $\mathcal{J} \& \mathcal{F}$ score of 81.6%. Suhwan Cho, MyeongAh Cho, Tae-Young Chung, Heansung Lee 0001, Sangyoun Lee |
ICIP | 5 |
| 2020 | Extrapolative-Interpolative Cycle-Consistency Learning For Video Frame ExtrapolationabstractVideo frame extrapolation is a task to predict future frames when the past frames are given. Unlike previous studies that usually have been focused on the design of modules or construction of networks, we propose a novel ExtrapolativeInterpolative Cycle (EIC) loss using pre-trained frame interpolation module to improve extrapolation performance. Cycle-consistency loss has been used for stable prediction between two function spaces in many visual tasks. We formulate this cycle-consistency using two mapping functions; frame extrapolation and interpolation. Since it is easier to predict intermediate frames than to predict future frames in terms of the object occlusion and motion uncertainty, interpolation module can give guidance signal effectively for training the extrapolation function. EIC loss can be applied to any existing extrapolation algorithms and guarantee consistent prediction in the short future as well as long future frames. Experimental results show that simply adding EIC loss to the existing baseline increases extrapolation performance on both UCF101 [1] and KITTI [2] datasets. Hyeongmin Lee, Taeoh Kim, Sangyoun Lee |
ICIP | 4 |
| 2020 | False Positive Removal for 3D Vehicle Detection With Penetrated Point ClassifierabstractRecently, researchers have been leveraging LiDAR point cloud for higher accuracy in 3D vehicle detection. Most state-of-the-art methods are deep learning based, but are easily affected by the number of points generated on the object. This vulnerability leads to numerous false positive boxes at high recall positions, where objects are occasionally predicted with few points. To address the issue, we introduce Penetrated Point Classifier (PPC) based on the underlying property of LiDAR that points cannot be generated behind vehicles. It determines whether a point exists behind the vehicle of the predicted box, and if does, the box is distinguished as false positive. Our straightforward yet unprecedented approach is evaluated on KITTI dataset and achieved performance improvement of PointRCNN, one of the state-of-the-art methods. The experiment results show that precision at the highest recall position is dramatically increased by 15.46 percentage points and 14.63 percentage points on the moderate and hard difficulty of car class, respectively. Sungmin Woo, Woo Jin Kim, Junhyeop Lee, Dogyoon Lee, Sangyoun Lee |
ICIP | 6 |
| 2020 | Protuberance of depth : Detecting interest points from a depth image
Yuseok Ban, Sangyoun Lee |
Comput. Vis. Image Underst. | 2 |
| 2019 | N-RPN: Hard Example Learning For Region Proposal NetworksabstractThe region proposal task is to generate a set of candidate regions that contain an object. In this task, it is most important to propose as many candidates of ground-truth as possible in a fixed number of proposals. In a typical image, however, there are too few hard negative examples compared to the vast number of easy negatives, so region proposal networks struggle to train on hard negatives. Because of this problem, networks tend to propose hard negatives as candidates, while failing to propose ground-truth candidates, which leads to poor performance. In this paper, we propose a Negative Region Proposal Network(nRPN) to improve Region Proposal Network(RPN). The nRPN learns from the RPN's false positives and provide hard negative examples to the RPN. Our proposed nRPN leads to a reduction in false positives and better RPN performance. An RPN trained with an nRPN achieves performance improvements on the PASCAL VOC 2007 dataset. MyeongAh Cho, Tae-Young Chung, Hyeongmin Lee, Sangyoun Lee |
ICIP | 4 |
| 2019 | SF-CNN: A Fast Compression Artifacts Removal via Spatial-To-Frequency Convolutional Neural NetworksabstractIn this paper, we propose SF-CNN, a fast convolutional neural network structure for JPEG image compression artifacts removal. Recently, Convolutional Neural Network (CNN)-based image restoration has shown great performance improvement. However, its heavy computational cost makes it difficult to apply to other uses such as high-level vision tasks. Since heavy computation arises from maintaining the spatial resolution of an input image, some works make a structure that is composed of spatial downsampling and upsampling operations. SF-CNN takes Spatial input and predicts residual Frequency using downsampling operations only. Since every 8×8 pixel is grouped and spatially invariant in the JPEG DCT domain, it is possible to down sample the input by a factor of 8 to reduce the computational cost. We show this simple structure is effective for compression artifacts removal. Our scalable baseline networks achieve results comparable to to the reference networks in reduced computations. Taeoh Kim, Hyeongmin Lee, Hanbin Son, Sangyoun Lee |
ICIP | 4 |
| 2018 | Design and implementation of monitoring system for breathing and heart rate pattern using WiFi signalsabstractBreathing pattern and heart rate can be major indicators of a person's physical condition, and an easy way to measure the vital signs can be useful in health monitoring. In this paper, we propose a new method for identifying the changes in breathing and heart rate pattern of a person using commercial WiFi devices. The amplitude of signal waves can represent the periodic up-and-down chest movements caused by breathing and heartbeat, and prominent changes of the signal pattern can be detected by using the Dynamic Time Warping algorithm. We verified the feasibility of the proposed method in real testbeds and evaluated the method through various experiments with 10 participants. The proposed method achieves 94% accuracy in identifying a subjects physical status. This low-cost method will be useful for monitoring our health in everyday life. Sangyoun Lee, Young Deok Park, Young-Joo Suh, Seokseong Jeon |
CCNC | 1 |
| 2018 | Collabonet: Collaboration of Generative Models by Unsupervised ClassificationabstractDesigning models for learning dataset with complex distributions is one of the main challenges that still remains in machine learning areas. We propose CollaboNet, which can divide a large dataset into sub-datasets, train two generative models separately, and let two models work together to achieve better performance. The proposed algorithm divides a large dataset without label since the capability difference between two generative models in performing tasks on each data is the main criterion for dividing a large dataset. In other words, the classification model can be trained by unsupervised manner. Autoencoder experiments for pure MNIST and the datasets combined artificially from two image sets shows that CollaboNet successfully splits large datasets without labels, improving the performance of generative models. Hyeongmin Lee, Taeoh Kim, Eungyeol Song, Sangyoun Lee |
ICIP | 4 |
| 2017 | An adaptive local binary pattern for 3D hand tracking
Joongrock Kim, Sunjin Yu, Dongchul Kim, Kar-Ann Toh, Sangyoun Lee |
Pattern Recognit. | 5 |
| 2015 | A Fast CU Size Decision Algorithm for HEVCabstractHigh Efficiency Video Coding (HEVC) employs a coding unit (CU), prediction unit (PU), and transform unit (TU) based on the quadtree coding tree unit (CTU) structure to improve coding efficiency. However, the computational complexity increases greatly because the rate-distortion (RD) optimization process should be performed for all CUs, PUs, and TUs to obtain the optimal CTU partition. In this paper, a fast CU size decision algorithm is proposed to reduce the encoder complexity of HEVC. Based on the statistical analysis, three approaches with SKIP mode decision (SMD), CU skip estimation (CUSE), and early CU termination (ECUT) are considered. In SMD, it is determined that the remaining modes except for SKIP mode are preformed or not. CUSE and ECUT determine that larger CU sizes and smaller CU sizes are coded or not, respectively. Thresholds for SMD, CUSE, and ECUT are designed based on Bayes' rule with a complexity factor. Update process is performed to estimate the statistical parameters for SMD, CUSE, and ECUT considering the characteristic of RD cost. The experimental results demonstrate that the proposed CU size decision algorithm significantly reduces computational complexity by 69% on average with 2.99% Bjøntegaard difference bitrate (BDBR) increase for random access. The complexity reduction and BDBR increase for low delay are 68% and 2.46%, respectively. The experimental results also show that our proposed scheme performs well for various characteristic of sequences and outperforms the two previous state-of-the-art works. Seongwan Kim, Kyungmin Lim, Sangyoun Lee |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2015 | Fast PU Skip and Split Termination Algorithm for HEVC Intra PredictionabstractHigh Efficiency Video Coding (HEVC) is developed for next-generation video coding, which achieves significant improvements in coding efficiency compared with H.264/Advanced Video Coding by adopting various tools including a quadtree-based block partitioning structure. However, this causes high encoding complexity for exhaustive rate-distortion (RD) cost computation of the extended prediction unit (PU) searching. In this paper, a fast PU skip and split termination algorithm is proposed. The proposed method consists of three algorithms: 1) early skip; 2) PU skip; and 3) PU split termination. The early skip algorithm allows immediate skipping of the RD cost computation for large PUs according to the neighboring PUs. Based on Bayes's rule, the PU skip algorithm allows skipping of the full RD cost computation, and the split termination algorithm terminates further PU splitting using the RD cost of rough mode decision (RMD). The decision parameter for the PU skip and the split termination is presented as the ratio of the RMD RD costs between the current PU and the spatially adjacent or upper depth PU. The simulation results show that the proposed algorithm achieves a saving of 53.52% encoding time while maintaining almost the same RD performances as the HEVC reference software. Kyungmin Lim, Seongwan Kim, Sangyoun Lee |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2014 | Face detection based on skin color likelihood
Yuseok Ban, Sangki Kim, Kar-Ann Toh, Sangyoun Lee |
Pattern Recognit. | 5 |
| 2010 | SVM-based feature extraction for face recognition
Sangki Kim, Youn Jung Park, Kar-Ann Toh, Sangyoun Lee |
Pattern Recognit. | 4 |
| 2008 | Reversible watermarking with localization for biometric imagesabstractIn this paper, we propose a novel reversible data hiding algorithm, which can recover the original image if it is deemed authentic or detect the block-wise malicious manipulation if it is classified as manipulated. We explore the strong spatial correlation of neighboring pixels in digital images to achieve very high embedding capacity and keep the distortion low. Also, this technique provides cryptographic strength when verifying image integrity because the probability of making undetectable modifications to the image is directly related to a secure cryptographic element, such as a hash function. The algorithm has been successfully applied to a wide range of images, including commonly used images, biometric images, texture images, and aerial images. Experimental results and performance comparison with other reversible data hiding schemes are presented to demonstrate the validity of the proposed algorithm. Hyobin Lee, Seongwan Kim, Sangyoun Lee |
ICARCV | 5 |
| 2008 | Cancellable biometrics and annotations on BioHash
Andrew Beng Jin Teoh, Wai Kuan Yip, Sangyoun Lee |
Pattern Recognit. | 3 |
| 2008 | Biometric scores fusion based on total error rate minimization
Kar-Ann Toh, Jaihie Kim, Sangyoun Lee |
Pattern Recognit. | 3 |
| 2008 | Maximizing area under ROC curve for biometric scores fusion
Kar-Ann Toh, Jaihie Kim, Sangyoun Lee |
Pattern Recognit. | 3 |
| 2008 | Fusion of visual and infra-red face scores by weighted power series
Kar-Ann Toh, Youngsung Kim, Sangyoun Lee, Jaihie Kim |
Pattern Recognit. Lett. | 3 |
| 2007 | Fingerprint Image Mosaicking by Recursive Ridge MappingabstractTo obtain a large fingerprint image from several small partial images, mosaicking of fingerprint images has been recently researched. However, existing approaches cannot provide accurate transformations for mosaics when it comes to aligning images because of the plastic distortion that may occur due to the nonuniform contact between a finger and a sensor or the deficiency of the correspondences in the images. In this paper, we propose a new scheme for mosaicking fingerprint images, which iteratively matches ridges to overcome the deficiency of the correspondences and compensates for the amount of plastic distortion between two partial images by using a thin-plate spline model. The proposed method also effectively eliminates erroneous correspondences and decides how well the transformation is estimated by calculating the registration error with a normalized distance map. The proposed method consists of three phases: feature extraction, transform estimation, and mosaicking. Transform is initially estimated with matched minutia and the ridges attached to them. Unpaired ridges in the overlapping area between two images are iteratively matched by minimizing the registration error, which consists of the ridge matching error and the inverse consistency error. During the estimation, erroneous correspondences are eliminated by considering the geometric relationship between the correspondences and checking if the registration error is minimized or not. In our experiments, the proposed method was compared with three existing methods in terms of registration accuracy, image quality, minutia extraction rate, processing time, reject to fuse rate, and verification performance. The average registration error of the proposed method was less than three pixels, and the maximum error was not more than seven pixels. In a verification test, the equal error rate was reduced from 10% to 2.7% when five images were combined by our proposed method. The proposed method was superior to other compared methods in terms of registration accuracy, image quality, minutia extraction rate, and verification. Kyoungtaek Choi, Heeseung Choi, Sangyoun Lee, Jaihie Kim |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2007 | Alignment-Free Cancelable Fingerprint Templates Based on Local Minutiae InformationabstractTo replace compromised biometric templates, cancelable biometrics has recently been introduced. The concept is to transform a biometric signal or feature into a new one for enrollment and matching. For making cancelable fingerprint templates, previous approaches used either the relative position of a minutia to a core point or the absolute position of a minutia in a given fingerprint image. Thus, a query fingerprint is required to be accurately aligned to the enrolled fingerprint in order to obtain identically transformed minutiae. In this paper, we propose a new method for making cancelable fingerprint templates that do not require alignment. For each minutia, a rotation and translation invariant value is computed from the orientation information of neighboring local regions around the minutia. The invariant value is used as the input to two changing functions that output two values for the translational and rotational movements of the original minutia, respectively, in the cancelable template. When a template is compromised, it is replaced by a new one generated by different changing functions. Our approach preserves the original geometric relationships (translation and rotation) between the enrolled and query templates after they are transformed. Therefore, the transformed templates can be used to verify a person without requiring alignment of the input fingerprint images. In our experiments, we evaluated the proposed method in terms of two criteria: performance and changeability. When evaluating the performance, we examined how verification accuracy varied as the transformed templates were used for matching. When evaluating the changeability, we measured the dissimilarities between the original and transformed templates, and between two differently transformed templates, which were obtained from the same original fingerprint. The experimental results show that the two criteria mutually affect each other and can be controlled by varying the control parameters of the changing functions. Chulhan Lee, Jeung-Yoon Choi, Kar-Ann Toh, Sangyoun Lee |
IEEE Trans. Syst. Man Cybern. Part B | 4 |
| 2006 | Robust 3D Face Data Acquisition Using a Sequential Color-Coded Pattern and Stereo Camera System
Ildo Kim, Sangki Kim, Sunjin Yu, Sangyoun Lee |
ICCSA (2) | 4 |
| 2006 | Robust Design of Face Recognition Systems
Sunjin Yu, Hyobin Lee, Jaihie Kim, Sangyoun Lee |
ICCSA (2) | 4 |
| 2006 | Automatic Pose-Normalized 3D Face Modeling and Recognition Systems
Sunjin Yu, Kwontaeg Choi, Sangyoun Lee |
PSIVT | 3 |
| 1999 | Parameter optimization of robust low-bit-rate video codersabstractMost standards provide a generalized syntax and semantics framework for video coders, leaving the selection and optimization of the right parameter set (and lookup tables) to the implementation. The choice of the right parameter set that is suitable for a rich enough class of input sequences is, however, quite difficult. This difficulty is particularly amplified in the low-bit-rate video coding arena, where robust parameter sets are very important. We propose that robust parameter estimation, using the Taguchi (1979) methods, when applied to low-bit-rate video coding allows effective (near optimal) performance over a wide variety of input data streams. A number of experimental results confirm the improvement (via robustness) vis-a-vis conventional parameter estimation methods, and these methods promise a solution to the design of efficient parameter sets that support standards. Sangyoun Lee, Vijay K. Madisetti |
IEEE Trans. Circuits Syst. Video Technol. | 1 |