VLDB 2026 Research / reviewers in the wild / expert
Changick Kim
dblp:40/5999
· DBLP profile ↗
136ranked-venue papers
10as first author
66since 2021 · last 2025
0000-0001-9323-8488ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 119 · 10 first-author · 54 since 2021Artificial intelligence and machine learning · 52 · 38 since 2021Systems, architecture and hardware · 1Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enhancing Robustness in Incremental Learning with Adversarial TrainingabstractAdversarial training is one of the most effective approaches against adversarial attacks. However, adversarial training has primarily been studied in scenarios where data for all classes is provided, with limited research conducted in the context of incremental learning where knowledge is introduced sequentially. In this study, we investigate Adversarially Robust Class Incremental Learning (ARCIL), which deals with adversarial robustness in incremental learning. We first explore a series of baselines that integrate incremental learning with existing adversarial training methods, finding that they lead to conflicts between acquiring new knowledge and retaining past knowledge. Furthermore, we discover that training new knowledge causes the disappearance of a key characteristic in robust models: a flat loss landscape in input space. To address such issues, we propose a novel and robust baseline for ARCIL, named FLatness preserving Adversarial Incremental learning for Robustness (FLAIR). Experimental results demonstrate that FLAIR significantly outperforms other baselines. To the best of our knowledge, we are the first to comprehensively investigate the baselines, challenges, and solutions for ARCIL, which we believe represents a significant advance toward achieving real-world robustness. Seungju Cho, Hongsin Lee, Changick Kim |
AAAI | 3 |
| 2025 | Diffusion Model Patching via Mixture-of-PromptsabstractWe present Diffusion Model Patching (DMP), a simple method to boost the performance of pre-trained diffusion models that have already reached convergence, with a negligible increase in parameters. DMP inserts a small, learnable set of prompts into the model's input space while keeping the original model frozen. The effectiveness of DMP is not merely due to the addition of parameters but stems from its dynamic gating mechanism, which selects and combines a subset of learnable prompts at every step of the generative process (i.e., reverse denoising steps). This strategy, which we term "mixture-of-prompts'', enables the model to draw on the distinct expertise of each prompt, essentially "patching'' the model's functionality at every step with minimal yet specialized parameters. Uniquely, DMP enhances the model by further training on the original dataset already used for training, even in a scenario where significant improvements are typically not expected due to model convergence. Experiments show that DMP significantly enhances the converged FID of DiT-L/2 on FFHQ by 10.38%, achieved with only a 1.43% parameter increase and 50K additional training iterations. Seokil Ham, Sangmin Woo, Hyojun Go, Byeongjun Park, Changick Kim |
AAAI | 6 |
| 2025 | SAFIRE: Segment Any Forged Image RegionabstractMost techniques approach the problem of image forgery localization as a binary segmentation task, training neural networks to label original areas as 0 and forged areas as 1. In contrast, we tackle this issue from a more fundamental perspective by partitioning images according to their originating sources. To this end, we propose Segment Any Forged Image Region (SAFIRE), which solves forgery localization using point prompting. Each point on an image is used to segment the source region containing itself. This allows us to partition images into multiple source regions, a capability achieved for the first time. Additionally, rather than memorizing certain forgery traces, SAFIRE naturally focuses on uniform characteristics within each source region. This approach leads to more stable and effective learning, achieving superior performance in both the new task and the traditional binary forgery localization. Myung-Joon Kwon, Wonjun Lee 0006, Seung-Hun Nam, Minji Son, Changick Kim |
AAAI | 5 |
| 2025 | Difficulty-aware Balancing Margin Loss for Long-tailed RecognitionabstractWhen trained with severely imbalanced data, deep neural networks often struggle to accurately recognize classes with few samples. Previous studies in long-tailed recognition have attempted to rebalance biased learning using known sample distributions, primarily addressing different classification difficulties at the class level. However, these approaches often overlook the instance difficulty variation within each class. In this paper, we propose a difficulty-aware balancing margin (DBM) loss, which considers both class imbalance and instance difficulty. DBM loss comprises two components: a class-wise margin to mitigate learning bias caused by imbalanced class frequencies, and an instance-wise margin assigned to hard positive samples based on their individual difficulty. DBM loss improves class discriminativity by assigning larger margins to more difficult samples. Our method effortlessly combine with existing approaches and consistently improves performance across various long-tailed recognition benchmarks. Minseok Son, Inyong Koo, Jinyoung Park 0001, Changick Kim |
AAAI | 4 |
| 2025 | SplatFlow: Multi-View Rectified Flow Model for 3D Gaussian Splatting SynthesisabstractText-based generation and editing of 3D scenes hold significant potential for streamlining content creation through intuitive user interactions. While recent advances leverage 3D Gaussian Splatting (3DGS) for high-fidelity and real-time rendering, existing methods are often specialized and task-focused, lacking a unified framework for both generation and editing. In this paper, we introduce SplatFlow, a comprehensive framework that addresses this gap by enabling direct 3DGS generation and editing. SplatFlow comprises two main components: a multi-view rectified flow (RF) model and a Gaussian Splatting Decoder (GSDecoder). The multi-view RF model operates in latent space, generating multi-view images, depths, and camera poses simultaneously, conditioned on text prompts—thus addressing challenges like diverse scene scales and complex camera trajectories in real-world settings. Then, the GSDecoder efficiently translates these latent outputs into 3DGS representations through a feed-forward 3DGS method. Leveraging training-free inversion and inpainting techniques, SplatFlow enables seamless 3DGS editing and supports a broad range of 3D tasks—including object editing, novel view synthesis, and camera pose estimation—within a unified framework without requiring additional complex pipelines. We validate SplatFlow’s capabilities on the MVImgNet and DL3DV-7K datasets, demonstrating its versatility and effectiveness in various 3D generation, editing, and inpainting-based tasks. Our project page is available at https://gohyojun15.github.io/SplatFlow/. Hyojun Go, Byeongjun Park, Jiho Jang, Soonwoo Kwon, Changick Kim |
CVPR | 6 |
| 2025 | Parameter Efficient Mamba Tuning via Projector-targeted Diagonal-centric Linear TransformationabstractDespite the growing interest in Mamba architecture as a potential replacement for Transformer architecture, parameter-efficient fine-tuning (PEFT) approaches for Mamba remain largely unexplored. In our study, we introduce two key insights-driven strategies for PEFT in Mamba architecture: (1) While state-space models (SSMs) have been regarded as the cornerstone of Mamba architecture, then expected to play a primary role in transfer learning, our findings reveal that Projectors—not SSMs—are the predominant contributors to transfer learning. (2) Based on our observation, we propose a novel PEFT method specialized to Mamba architecture: Projector-targeted Diagonalcentric Linear Transformation (ProDiaL). ProDiaL focuses on optimizing only the pretrained Projectors for new tasks through diagonal-centric linear transformation matrices, without directly fine-tuning the Projector weights. This targeted approach allows efficient task adaptation, utilizing less than 1% of the total parameters, and exhibits strong performance across both vision and language Mamba models, highlighting its versatility and effectiveness. Seokil Ham, Hee-Seon Kim, Sangmin Woo, Changick Kim |
CVPR | 4 |
| 2025 | Focusing on Tracks for Online Multi-Object TrackingabstractMulti-object tracking (MOT) is a critical task in computer vision, requiring the accurate identification and continuous tracking of multiple objects across video frames. However, current state-of-the-art methods mainly rely on a global optimization technique and multi-stage cascade association strategy, and those approaches often overlook the specific characteristics of assignment task in MOT and useful detection results that may represent occluded objects. To address these challenges, we propose a novel Track-Focused Online Multi-Object Tracker (TrackTrack) with two key strategies: Track-Perspective-Based Association (TPA) and Track-Aware Initialization (TAI). The TPA strategy associates each track with the most suitable detection result by choosing the one with the minimum distance from all available detection results in a track-perspective manner. On the other hand, TAI precludes the generation of spurious tracks in the track-aware aspect by suppressing track initialization of detection results that heavily overlap with current active tracks and more confident detection results. Extensive experiments on MOT17, MOT20, and DanceTrack demonstrate that our TrackTrack outperforms current state-of-the-art trackers, offering improved robustness and accuracy across diverse and challenging tracking scenarios. Kyujin Shim, Kangwook Ko, Yujin Yang, Changick Kim |
CVPR | 4 |
| 2025 | VideoRFSplat: Direct Scene-Level Text-to-3D Gaussian Splatting Generation with Flexible Pose and Multi-View Joint ModelingabstractWe propose VideoRFSplat, a direct text-to-3D model leveraging a video generation model to generate realistic 3D Gaussian Splatting (3DGS) for unbounded real-world scenes. To generate diverse camera poses and unbounded spatial extent of real-world scenes, while ensuring generalization to arbitrary text prompts, previous methods fine-tune 2D generative models to jointly model camera poses and multi-view images. However, these methods suffer from instability when extending 2D generative models to joint modeling due to the modality gap, which necessitates additional models to stabilize training and inference. In this work, we propose an architecture and a sampling strategy to jointly model multi-view images and camera poses when fine-tuning a video generation model. Our core idea is a dual-stream architecture that attaches a dedicated pose generation model alongside a pre-trained video generation model via communication blocks, generating multi-view images and camera poses through separate streams. This design reduces interference between the pose and image modalities. Additionally, we propose an asynchronous sampling strategy that denoises camera poses faster than multi-view images, allowing rapidly denoised poses to condition multi-view generation, reducing mutual ambiguity and enhancing cross-modal consistency. Trained on multiple large-scale real-world datasets (RealEstate10K, MVImgNet, DL3DV-10K, ACID), VideoRFSplat outperforms existing text-to-3D direct generation methods that heavily depend on post-hoc refinement via score distillation sampling, achieving superior results without such refinement. Hyojun Go, Byeongjun Park, Hyelin Nam, Byung-Hoon Kim, Hyungjin Chung, Changick Kim |
ICCV | 6 |
| 2025 | SteerX: Creating Any Camera-Free 3D and 4D Scenes with Geometric Steering
Byeongjun Park, Hyojun Go, Hyelin Nam, Byung-Hoon Kim, Hyungjin Chung, Changick Kim |
ICCV | 6 |
| 2025 | Long-tailed Adversarial Training with Self-DistillationabstractAdversarial training significantly enhances adversarial robustness, yet superior performance is predominantly achieved on balanced datasets.
Addressing adversarial robustness in the context of unbalanced or long-tailed distributions is considerably more challenging, mainly due to the scarcity of tail data instances.
Previous research on adversarial robustness within long-tailed distributions has primarily focused on combining traditional long-tailed natural training with existing adversarial robustness methods.
In this study, we provide an in-depth analysis for the challenge that adversarial training struggles to achieve high performance on tail classes in long-tailed distributions.
Furthermore, we propose a simple yet effective solution to advance adversarial robustness on long-tailed distributions through a novel self-distillation technique.
Specifically, this approach leverages a balanced self-teacher model, which is trained using a balanced dataset sampled from the original long-tailed dataset.
Our extensive experiments demonstrate state-of-the-art performance in both clean and robust accuracy for long-tailed adversarial robustness, with significant improvements in tail class performance on various datasets.
We improve the accuracy against PGD attacks for tail classes by 20.3, 7.1, and 3.8 percentage points on CIFAR-10, CIFAR-100, and Tiny-ImageNet, respectively, while achieving the highest robust accuracy. Seungju Cho, Hongsin Lee, Changick Kim |
ICLR | 3 |
| 2025 | Indirect Gradient Matching for Adversarial Robust DistillationabstractAdversarial training significantly improves adversarial robustness, but superior performance is primarily attained with large models.
This substantial performance gap for smaller models has spurred active research into adversarial distillation (AD) to mitigate the difference.
Existing AD methods leverage the teacher’s logits as a guide.
In contrast to these approaches, we aim to transfer another piece of knowledge from the teacher, the input gradient.
In this paper, we propose a distillation module termed Indirect Gradient Distillation Module (IGDM) that indirectly matches the student’s input gradient with that of the teacher.
Experimental results show that IGDM seamlessly integrates with existing AD methods, significantly enhancing their performance.
Particularly, utilizing IGDM on the CIFAR-100 dataset improves the AutoAttack accuracy from 28.06\% to 30.32\% with the ResNet-18 architecture and from 26.18\% to 29.32\% with the MobileNetV2 architecture when integrated into the SOTA method without additional data augmentation. Hongsin Lee, Seungju Cho, Changick Kim |
ICLR | 3 |
| 2025 | Modality mixer exploiting complementary information for multi-modal action recognition
Sangmin Woo, Muhammad Adi Nugroho, Changick Kim |
Comput. Vis. Image Underst. | 4 |
| 2025 | FCGNet: Foreground and Class Guided Network for human parsing
Jaehyuk Jang, Yooseung Wang, Changick Kim |
Pattern Recognit. | 3 |
| 2024 | HarmonyView: Harmonizing Consistency and Diversity in One-Image-to-3DabstractRecent progress in single-image 3D generation highlights the importance of multi-view coherency, leveraging 3D priors from large-scale diffusion models pretrained on Internet-scale images. However, the aspect of novel-view diversity remains underexplored within the research landscape due to the ambiguity in converting a 2D image into 3D content, where numerous potential shapes can emerge. Here, we aim to address this research gap by simultaneously addressing both consistency and diversity. Yet, striking a balance between these two aspects poses a considerable challenge due to their inherent trade-offs. This work introduces HarmonyView, a simple yet effective diffusion sampling technique adept at decomposing two intricate aspects in single-image 3D generation: consistency and diversity. This approach paves the way for a more nuanced exploration of the two critical dimensions within the sampling process. Moreover, we propose a new evaluation metric based on CLIP image and text encoders to comprehensively assess the diversity of the generated views, which closely aligns with human evaluators' judgments. In experiments, HarmonyView achieves a harmonious balance, demonstrating a win-win scenario in both consistency and diversity. Sangmin Woo, Byeongjun Park, Hyojun Go, Changick Kim |
CVPR | 5 |
| 2024 | Spatio-Temporal Proximity-Aware Dual-Path Model for Panoramic Activity Recognition
Yooseung Wang, Sangmin Woo, Changick Kim |
ECCV (8) | 4 |
| 2024 | Flow-Assisted Motion Learning Network for Weakly-Supervised Group Activity Recognition
Muhammad Adi Nugroho, Sangmin Woo, Jinyoung Park 0001, Yooseung Wang, Changick Kim |
ECCV (48) | 7 |
| 2024 | Switch Diffusion Transformer: Synergizing Denoising Tasks with Sparse Mixture-of-Experts
Byeongjun Park, Hyojun Go, Sangmin Woo, Seokil Ham, Changick Kim |
ECCV (53) | 6 |
| 2024 | VideoMamba: Spatio-Temporal Selective State Space Model
Jinyoung Park 0001, Hee-Seon Kim, Kangwook Ko, Minbeom Kim, Changick Kim |
ECCV (25) | 5 |
| 2024 | Towards Robust Multimodal Prompting with Missing ModalitiesabstractRecently, multimodal prompting, which introduces learnable missing-aware prompts for all missing modality cases, has exhibited impressive performance. However, it encounters two critical issues: 1) The number of prompts grows exponentially as the number of modalities increases; and 2) It lacks robustness in scenarios with different missing modality settings between training and inference. In this paper, we propose a simple yet effective prompt design to address these challenges. Instead of using missing-aware prompts, we utilize prompts as modality-specific tokens, enabling them to capture the unique characteristics of each modality. Furthermore, our prompt design leverages orthogonality between prompts as a key element to learn distinct information across different modalities and promote diversity in the learned representations. Extensive experiments demonstrate that our prompt design enhances both performance and robustness while reducing the number of prompts. Jaehyuk Jang, Yooseung Wang, Changick Kim |
ICASSP | 3 |
| 2024 | A Confidence-Aware Matching Strategy For Generalized Multi-Object TrackingabstractMulti-object tracking (MOT), a crucial task in computer vision, has broad applicability, and recently, tracking-by-detection-based trackers, which separate the processes of object detection and association, are showing state-of-the-art performance. However, while techniques like feature enhancement and distance measures have been extensively explored, the matching strategy itself remains an area that requires more in-depth study. As a result, many trackers still require manual adjustment of sensitive hyper-parameters for each tracking scenario, limiting their adaptability and robustness in dynamic environments. To address these limitations, we introduce CMTrack, a new tracker featuring a novel confidence-aware matching strategy comprised of three modules: confidence-aware cascade matching (CCM), confidence-aware metric fusion (CMF), and confidence-aware feature update (CFU). Our matching strategy enables the tracker to be a generalized and practical solution for various tracking scenarios within a unified framework while obviating manual calibration of hyper-parameters. The effectiveness of CMTrack is demonstrated through comprehensive assessments of three prominent MOT datasets: MOT17, MOT20, and DanceTrack. Notably, our CMTrack consistently surpasses existing state-of-the-art trackers, showcasing its superior generalization capabilities. The source codes and models are open at https://github.com/kamkyu94/CMTrack. Kyujin Shim, Jubi Hwang, Kangwook Ko, Changick Kim |
ICIP | 4 |
| 2024 | Adaptrack: Adaptive Thresholding-Based Matching for Multi-Object TrackingabstractMulti-object tracking (MOT) plays a pivotal role in various computer vision domains with recent tracking-by detection algorithms that treat MOT as distinct detection and association tasks. However, the existing trackers often rely on sensitive thresholds to associate previous tracks and current detection results while forming complete trajectories across a video. These thresholds are crucial for tracking performance and require manual tuning for each dataset or even sequence, limiting the adaptability in real-world applications. To tackle this problem, in this paper, we introduce AdapTrack, a novel MOT algorithm designed to enable adaption on varying scenarios without handcrafted threshold configuration. With a carefully designed matching strategy, our tracker can adaptively select proper thresholds for each frame and correctly associate detected objects. Consequently, AdapTrack shows outperforming results on standard MOT benchmarks, MOT17 and MOT20, compared to existing state-of-the-art methods. Every source code is available at https://github.com/kamkyu94/AdapTrack. Kyujin Shim, Kangwook Ko, Jubi Hwang, Changick Kim |
ICIP | 4 |
| 2024 | Denoising Task Routing for Diffusion ModelsabstractDiffusion models generate highly realistic images by learning a multi-step denoising process, naturally embodying the principles of multi-task learning (MTL). Despite the inherent connection between diffusion models and MTL, there remains an unexplored area in designing neural architectures that explicitly incorporate MTL into the framework of diffusion models. In this paper, we present Denoising Task Routing (DTR), a simple add-on strategy for existing diffusion model architectures to establish distinct information pathways for individual tasks within a single architecture by selectively activating subsets of channels in the model. What makes DTR particularly compelling is its seamless integration of prior knowledge of denoising tasks into the framework: (1) Task Affinity: DTR activates similar channels for tasks at adjacent timesteps and shifts activated channels as sliding windows through timesteps, capitalizing on the inherent strong affinity between tasks at adjacent timesteps. (2) Task Weights: During the early stages (higher timesteps) of the denoising process, DTR assigns a greater number of task-specific channels, leveraging the insight that diffusion models prioritize reconstructing global structure and perceptually rich contents in earlier stages, and focus on simple noise removal in later stages. Our experiments reveal that DTR not only consistently boosts diffusion models' performance across different evaluation protocols without adding extra parameters but also accelerates training convergence. Finally, we show the complementarity between our architectural approach and existing MTL optimization techniques, providing a more complete view of MTL in the context of diffusion training. Significantly, by leveraging this complementarity, we attain matched performance of DiT-XL using the smaller DiT-L with a reduction in training iterations from 7M to 2M. Our project page is available at https://byeongjun-park.github.io/DTR/ Byeongjun Park, Sangmin Woo, Hyojun Go, Changick Kim |
ICLR | 5 |
| 2024 | FRIDAY: Mitigating Unintentional Facial Identity in Deepfake Detectors Guided by Facial RecognizersabstractPrevious Deepfake detection methods perform well within their training domains, but their effectiveness diminishes significantly with new synthesis techniques. Recent studies have revealed that detection models make decision boundaries based on facial identity instead of synthetic artifacts, leading to poor cross-domain performance. To address this issue, we propose FRIDAY, a novel training method that attenuates facial identity utilizing a face recognizer. To be specific, we first train a face recognizer using the same backbone as the Deepfake detector. We then freeze the recognizer and use it during the detector’s training to mitigate facial identity information. This is achieved by feeding input images into both the recognizer and the detector, then minimizing the similarity of their feature embeddings using our Facial Identity Attenuating loss. This process encourages the detector to produce embeddings distinct from the recognizer, effectively attenuating facial identity. Comprehensive experiments demonstrate that our approach significantly improves detection performance on both in-domain and cross-domain datasets. Younghun Kim, Myung-Joon Kwon, Wonjun Lee 0006, Changick Kim |
VCIP | 4 |
| 2024 | Anchoring Vision and Language Knowledge for Weakly Supervised Group Activity RecognitionabstractThe emergence of Foundation Vision-Language Models (VLMs) has ignited a surge of research in the computer vision field due to their robust baseline performance. Inspired by this, we propose the Anchoring Vision-Language Network (AnViL-Net), which integrates a vision language model for the challenging task of Weakly-Supervised Group Activity Recognition (WSGAR). Our network effectively incorporates VLMs into WSGAR, addressing the challenges posed by dynamic actor motions and domain-specific activity classes. AnViL-Net leverages highly generalized VLM vision features as anchors for extracting visual features. Additionally, semantically meaningful VLM language features serve as anchors for inferring the semantic relationships between actors and their activities. We demonstrate the effectiveness of AnViL-Net on multiple group activity datasets, achieving competitive state-of-the-art results. Muhammad Adi Nugroho, Jinyoung Park 0001, Changick Kim |
VCIP | 4 |
| 2024 | Point-DynRF: Point-based Dynamic Radiance Fields from a Monocular VideoabstractDynamic radiance fields have emerged as a promising approach for generating novel views from a monocular video. However, previous methods enforce the geometric consistency to dynamic radiance fields only between adjacent input frames, making it difficult to represent the global scene geometry and degenerates at the viewpoint that is spatio-temporally distant from the input camera trajectory. To solve this problem, we introduce point-based dynamic radiance fields (Point-DynRF), a novel framework where the global geometric information and the volume rendering process are trained by neural point clouds and dynamic radiance fields, respectively. Specifically, we reconstruct neural point clouds directly from geometric proxies and optimize both radiance fields and the geometric proxies using our proposed losses, allowing them to complement each other. We validate the effectiveness of our method with experiments on the NVIDIA Dynamic Scenes Dataset and several causally captured monocular video clips. Byeongjun Park, Changick Kim |
WACV | 2 |
| 2024 | Sketch-based Video Object LocalizationabstractWe introduce Sketch-based Video Object Localization (SVOL), a new task aimed at localizing spatio-temporal object boxes in video queried by the input sketch. We first outline the challenges in the SVOL task and build the Sketch-Video Attention Network (SVANet) with the following design principles: (i) to consider temporal information of video and bridge the domain gap between sketch and video; (ii) to accurately identify and localize multiple objects simultaneously; (iii) to handle various styles of sketches; (iv) to be classification-free. In particular, SVANet is equipped with a Cross-modal Transformer that models the interaction between learnable object tokens, query sketch, and video through attention operations, and learns upon a per-frame set matching strategy that enables frame-wise prediction while utilizing global video context. We evaluate SVANet on a newly curated SVOL dataset. By design, SVANet successfully learns the mapping between the query sketches and video objects, achieving state-of-the-art results on the SVOL benchmark. We further confirm the effectiveness of SVANet via extensive ablation studies and visualizations. Lastly, we demonstrate its transfer capability on unseen datasets and novel categories, suggesting its high scalability in real-world applications. Codes are available at https://github.com/sangminwoo/SVOL. Sangmin Woo, So-Yeong Jeon, Jinyoung Park 0001, Minji Son, Changick Kim |
WACV | 6 |
| 2024 | Bridging Implicit and Explicit Geometric Transformation for Single-Image View SynthesisabstractCreating novel views from a single image has achieved tremendous strides with advanced autoregressive models, as unseen regions have to be inferred from the visible scene contents. Although recent methods generate high-quality novel views, synthesizing with only one explicit or implicit 3D geometry has a trade-off between two objectives that we call the "seesaw" problem: 1) preserving reprojected contents and 2) completing realistic out-of-view regions. Also, autoregressive models require a considerable computational cost. In this paper, we propose a single-image view synthesis framework for mitigating the seesaw problem while utilizing an efficient non-autoregressive model. Motivated by the characteristics that explicit methods well preserve reprojected pixels and implicit methods complete realistic out-of-view regions, we introduce a loss function to complement two renderers. Our loss function promotes that explicit features improve the reprojected area of implicit features and implicit features improve the out-of-view area of explicit features. With the proposed architecture and loss function, we can alleviate the seesaw problem, outperforming autoregressive-based state-of-the-art methods and generating an image ≈ 100 times faster. We validate the efficiency and effectiveness of our method with experiments on RealEstate10 K and ACID datasets. Byeongjun Park, Hyojun Go, Changick Kim |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Subdivided Mask Dispersion Framework for semi-supervised semantic segmentation
Yooseung Wang, Jaehyuk Jang, Changick Kim |
Pattern Recognit. Lett. | 3 |
| 2024 | Enhancing Robustness of Multi-Object Trackers With Temporal Feature MixabstractDespite its recent advancements, multi-object tracking (MOT), one of the major research areas in video technology, still faces various challenges, including severe occlusion and diversity of tracking targets. In this paper, we introduce a novel strategy, Temporal Feature Mix (TFM), that can improve the overall robustness of multi-object trackers in diverse scenarios. More specifically, our approach simulates new and challenging scenes that can train networks to better localize the targets by blending high-level features from temporally adjacent frames with the insights that the high-level features are mainly activated on salient targets and the targets on the adjacent frames are nearly located. Therefore, our TFM can offer novel and diversified training experiences to the networks, achieved through the intensive augmentation of the high-level features of each target. As a result, our approach demonstrates notable performance improvement with three major MOT benchmarks and a newly constructed corruption dataset for MOT, underscoring its potential to enhance the robustness of MOT systems in real-world scenarios. Every related source code is released at https://github.com/kamkyu94/Temporal Feature Mix. Kyujin Shim, Junyoung Byun, Kangwook Ko, Jubi Hwang, Changick Kim |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Towards Good Practices for Missing Modality Robust Action RecognitionabstractStandard multi-modal models assume the use of the same modalities in training and inference stages. However, in practice, the environment in which multi-modal models operate may not satisfy such assumption. As such, their performances degrade drastically if any modality is missing in the inference stage. We ask: how can we train a model that is robust to missing modalities? This paper seeks a set of good practices for multi-modal action recognition, with a particular interest in circumstances where some modalities are not available at an inference time. First, we show how to effectively regularize the model during training (e.g., data augmentation). Second, we investigate on fusion methods for robustness to missing modalities: we find that transformer-based fusion shows better robustness for missing modality than summation or concatenation. Third, we propose a simple modular network, ActionMAE, which learns missing modality predictive coding by randomly dropping modality features and tries to reconstruct them with the remaining modality features. Coupling these good practices, we build a model that is not only effective in multi-modal action recognition but also robust to modality missing. Our model achieves the state-of-the-arts on multiple benchmarks and maintains competitive performances even in missing modality scenarios. Sangmin Woo, Yeonju Park, Muhammad Adi Nugroho, Changick Kim |
AAAI | 5 |
| 2023 | Introducing Competition to Boost the Transferability of Targeted Adversarial Examples Through Clean Feature MixupabstractDeep neural networks are widely known to be susceptible to adversarial examples, which can cause incorrect predictions through subtle input modifications. These adversarial examples tend to be transferable between models, but targeted attacks still have lower attack success rates due to significant variations in decision boundaries. To enhance the transferability of targeted adversarial examples, we propose introducing competition into the optimization process. Our idea is to craft adversarial perturbations in the presence of two new types of competitor noises: adversarial perturbations towards different target classes and friendly perturbations towards the correct class. With these competitors, even if an adversarial example deceives a network to extract specific features leading to the target class, this disturbance can be suppressed by other competitors. Therefore, within this competition, adversarial examples should take different attack strategies by leveraging more diverse features to overwhelm their interference, leading to improving their transferability to different models. Considering the computational complexity, we efficiently simulate various interference from these two types of competitors in feature space by randomly mixing up stored clean features in the model inference and named this method Clean Feature Mixup (CFM). Our extensive experimental results on the ImageNet-Compatible and CIFAR-10 datasets show that the proposed method outperforms the existing baselines with a clear margin. Our code is available at https://github.com/dreamflake/CFM. Junyoung Byun, Myung-Joon Kwon, Seungju Cho, Changick Kim |
CVPR | 5 |
| 2023 | Liveness Score-Based Regression Neural Networks for Face Anti-SpoofingabstractPrevious anti-spoofing methods have used either pseudo maps or user-defined labels, and the performance of each approach depends on the accuracy of the third party networks generating pseudo maps and the way in which the users define the labels. In this paper, we propose a liveness score-based regression network for overcoming the dependency on third party networks and users. First, we introduce a new labeling technique, called pseudo-discretized label encoding for generating discretized labels indicating the amount of information related to real images. Secondly, we suggest the expected liveness score based on a regression network for training the difference between the proposed supervision and the expected liveness score. Finally, extensive experiments were conducted on four face anti-spoofing benchmarks to verify our proposed method on both intra-and cross-dataset tests. The experimental results show our approach outperforms previous methods. Youngjun Kwak, Minyoung Jung, Hunjae Yoo, Jinho Shin, Changick Kim |
ICASSP | 5 |
| 2023 | ProtoFL: Unsupervised Federated Learning via Prototypical DistillationabstractFederated learning (FL) is a promising approach for enhancing data privacy preservation, particularly for authentication systems. However, limited round communications, scarce representation, and scalability pose significant challenges to its deployment, hindering its full potential. In this paper, we propose ‘ProtoFL’, Prototypical Representation Distillation based unsupervised Federated Learning to enhance the representation power of a global model and reduce round communication costs. Additionally, we introduce a local one-class classifier based on normalizing flows to improve performance with limited data. Our study represents the first investigation of using FL to improve one-class classification performance. We conduct extensive experiments on five widely used benchmarks, namely MNIST, CIFAR-10, CIFAR-100, ImageNet-30, and Keystroke-Dynamics, to demonstrate the superior performance of our proposed framework over previous methods in the literature. Youngjun Kwak, Minyoung Jung, Jinho Shin, Youngsung Kim, Changick Kim |
ICCV | 6 |
| 2023 | Breaking Temporal Consistency: Generating Video Universal Adversarial Perturbations Using Image ModelsabstractAs video analysis using deep learning models becomes more widespread, the vulnerability of such models to adversarial attacks is becoming a pressing concern. In particular, Universal Adversarial Perturbation (UAP) poses a significant threat, as a single perturbation can mislead deep learning models on entire datasets. We propose a novel video UAP using image data and image model. This enables us to take advantage of the rich image data and image model-based studies available for video applications. However, there is a challenge that image models are limited in their ability to analyze the temporal aspects of videos, which is crucial for a successful video attack. To address this challenge, we introduce the Breaking Temporal Consistancy (BTC) method, which is the first attempt to incorporate temporal information into video attacks using image models. We aim to generate adversarial videos that have opposite patterns to the original. Specifically, BTC-UAP minimizes the feature similarity between neighboring frames in videos. Our approach is simple but effective at attacking unseen video models. Additionally, it is applicable to videos of varying lengths and invariant to temporal shifts. Our approach surpasses existing methods in terms of effectiveness on various datasets, including ImageNet, UCF-101, and Kinetics-400. Hee-Seon Kim, Minji Son, Minbeom Kim, Myung-Joon Kwon, Changick Kim |
ICCV | 5 |
| 2023 | PG-RCNN: Semantic Surface Point Generation for 3D Object DetectionabstractOne of the main challenges in LiDAR-based 3D object detection is that the sensors often fail to capture the complete spatial information about the objects due to long distance and occlusion. Two-stage detectors with point cloud completion approaches tackle this problem by adding more points to the regions of interest (RoIs) with a pre-trained network. However, these methods generate dense point clouds of objects for all region proposals, assuming that objects always exist in the RoIs. This leads to the indiscriminate point generation for incorrect proposals as well. Motivated by this, we propose Point Generation R-CNN (PG-RCNN), a novel end-to-end detector that generates semantic surface points of foreground objects for accurate detection. Our method uses a jointly trained RoI point generation module to process the contextual information of RoIs and estimate the complete shape and displacement of foreground objects. For every generated point, PG-RCNN assigns a semantic feature that indicates the estimated foreground probability. Extensive experiments show that the point clouds generated by our method provide geometrically and semantically rich information for refining false positive and misaligned proposals. PG-RCNN achieves competitive performance on the KITTI benchmark, with significantly fewer parameters than state-of-the-art models. The code is available at https://github.com/quotation2520/PG-RCNN. Inyong Koo, Inyoung Lee, Se-Ho Kim, Hee-Seon Kim, Woo-Jin Jeon, Changick Kim |
ICCV | 6 |
| 2023 | Audio-Visual Glance Network for Efficient Video RecognitionabstractDeep learning has made significant strides in video understanding tasks, but the computation required to classify lengthy and massive videos using clip-level video classifiers remains impractical and prohibitively expensive. To address this issue, we propose Audio-Visual Glance Network (AVGN), which leverages the commonly available audio and visual modalities to efficiently process the spatio-temporally important parts of a video. AVGN firstly divides the video into snippets of image-audio clip pair and employs lightweight unimodal encoders to extract global visual features and audio features. To identify the important temporal segments, we use an Audio-Visual Temporal Saliency Transformer (AV-TeST) that estimates the saliency scores of each frame. To further increase efficiency in the spatial dimension, AVGN processes only the important patches instead of the whole images. We use an Audio-Enhanced Spatial Patch Attention (AESPA) module to produce a set of enhanced coarse visual features, which are fed to a policy network that produces the coordinates of the important patches. This approach enables us to focus only on the most important spatio-temporally parts of the video, leading to more efficient video recognition. Moreover, we incorporate various training techniques and multi-modal feature fusion to enhance the robustness and effectiveness of our AVGN. By combining these strategies, our AVGN sets new state-of-the-art performance in multiple video recognition benchmarks while achieving faster processing speed. Muhammad Adi Nugroho, Sangmin Woo, Changick Kim |
ICCV | 4 |
| 2023 | Improving Adversarial Transferability Via Feature TranslationabstractDeep Neural Networks (DNNs) are vulnerable to adversarial examples, which are crafted to cause the model to make wrong predictions. In real-world scenario, since adversary cannot access to target models, black-box attack has attracted great attention. Among them, many studies have been conducted on transfer-based attacks because they can effectively attack unknown target model. However, transfer-based attacks often fail to fool other models which have slightly different activation maps because adversarial examples tend to overfit to the source model. To alleviate this problem, we introduce Feature Translation Attack (FTA), which applies translation on intermediate features during optimization process. Specifically, FTA generates a new adversarial example whose feature is similar to the ensemble of translated features from the existing adversarial example. We achieved better performance than state-of-the-art methods in extensive experiments. Seungju Cho, Junyoung Byun, Myung-Joon Kwon, Changick Kim |
ICIP | 5 |
| 2023 | Multi-modal Social Group Activity Recognition in Panoramic SceneabstractGroup Activity Recognition (GAR) is a challenging problem in computer vision due to the intricate dynamics and interactions among individuals. The existing methods utilize RGB videos face challenges in panoramic environments with numerous individuals and social groups. In this paper, we propose Multimodal Group Activity Recognition network (MGAR-net), that leverages the combined power of RGB and LiDAR modalities. Our approach effectively utilizes information from both modalities thus robustly and accurately captures individual relationships and detects social groups in face of optical challenges. By harnessing the capability of LiDAR with our new fusion module, called Distance Aware Fusion Module (DAFM), MGAR-net acquires valuable 3D structure information. We conduct experiments on the JRDB-Act dataset, which contains challenging scenarios with numerous people. The results demonstrate that LiDAR data provide valuable information for social grouping and recognizing individual action and group activities, particularly in crowded group settings. For social grouping, our MGAR-net improve performance by about 12% compared to the existing state-of-the-art models in terms of the AP metric. Sangmin Woo, Jinyoung Park 0001, Muhammad Adi Nugroho, Changick Kim |
VCIP | 6 |
| 2023 | AHFu-Net: Align, Hallucinate, and Fuse Network for Missing Multimodal Action RecognitionabstractIn this work, we explore the multimodal action recognition problem, specifically in the context of RGB-Depth modalities scenario, where a subset of the learning modalities is missing at inference time. To address this issue, we construct a hallucination network to generate missing modality information from the available modality at inference time. We propose key components of an effective spatio-temporal encoder for strong unimodal performance with Local Patch Temporal Transformer (LPTT) and Spatial Encoder Transformer (SET), alignment of multi-modal features, and fusion strategy with our Multimodal Bottleneck Transformer Fusion Module (MMBTF). We incorporate these ideas into a novel framework named AHFu-Net (Align, Hallucinate, and Fuse network) for RGB-Depth action recognition. Our experiments demonstrate that AHFu Net achieves state-of-the-art performance while maintaining high accuracy in the case of missing modality on multimodal datasets of NTU-RGB+D and NWUCLA. Muhammad Adi Nugroho, Sangmin Woo, Changick Kim |
VCIP | 4 |
| 2023 | Noise-Augmented Missing Modality Aware Prompt Based Learning for Robust Visual RecognitionabstractMultimodal learning is essential for understanding interactions between different input domains. However, dealing with various modalities often leads to a high number of network parameters and extended training time. To tackle these challenges, a recent approach called "missing modality aware prompting" enhances model robustness with minimal parameters by freezing the transformer-based backbone network and introducing missing modality aware prompts. In this paper, we propose a robust missing modality aware prompting approach with the same parameter numbers as the naive prompts by adding noise. Our experiments demonstrate that robust missing modality aware prompts outperform state-of-the-art missing modality prompt-based learning in various scenarios. Additionally, our ablation study verifies the effectiveness of robust missing modality aware prompts across different signal-to-noise ratios. Yooseung Wang, Jaehyuk Jang, Changick Kim |
VCIP | 3 |
| 2023 | Modality Mixer for Multi-modal Action RecognitionabstractIn multi-modal action recognition, it is important to consider not only the complementary nature of different modalities but also global action content. In this paper, we propose a novel network, named Modality Mixer (M-Mixer) network, to leverage complementary information across modalities and temporal context of an action for multi-modal action recognition. We also introduce a simple yet effective recurrent unit, called Multi-modal Contextualization Unit (MCU), which is a core component of M-Mixer. Our MCU temporally encodes a sequence of one modality (e.g., RGB) with action content features of other modalities (e.g., depth, IR). This process encourages M-Mixer to exploit global action content and also to supplement complementary information of other modalities. As a result, our proposed method outperforms state-of-the-art methods on NTU RGB+D 60, NTU RGB+D 120, and NW-UCLA datasets. Moreover, we demonstrate the effectiveness of M-Mixer by conducting comprehensive ablation studies. Sangmin Woo, Yeonju Park, Muhammad Adi Nugroho, Changick Kim |
WACV | 5 |
| 2023 | Fast online multi-target multi-camera tracking for vehicles
Kyujin Shim, Kangwook Ko, Jubi Hwang, Hyunsung Jang, Changick Kim |
Appl. Intell. | 5 |
| 2023 | Cross-modal alignment and translation for missing modality action recognition
Yeonju Park, Sangmin Woo, Muhammad Adi Nugroho, Changick Kim |
Comput. Vis. Image Underst. | 5 |
| 2023 | Learning to Discriminate Information for Online Action Detection: Analysis and ApplicationabstractOnline action detection, which aims to identify an ongoing action from a streaming video, is an important subject in real-world applications. For this task, previous methods use recurrent neural networks for modeling temporal relations in an input sequence. However, these methods overlook the fact that the input image sequence includes not only the action of interest but background and irrelevant actions. This would induce recurrent units to accumulate unnecessary information for encoding features on the action of interest. To overcome this problem, we propose a novel recurrent unit, named Information Discrimination Unit (IDU), which explicitly discriminates the information relevancy between an ongoing action and others to decide whether to accumulate the input information. This enables learning more discriminative representations for identifying an ongoing action. In this paper, we further present a new recurrent unit, called Information Integration Unit (IIU), for action anticipation. Our IIU exploits the outputs from IDN as pseudo action labels as well as RGB frames to learn enriched features of observed actions effectively. In experiments on TVSeries and THUMOS-14, the proposed methods outperform state-of-the-art methods by a significant margin in online action detection and action anticipation. Moreover, we demonstrate the effectiveness of the proposed units by conducting comprehensive ablation studies. Hyunjun Eun, Jinyoung Moon, Seokeon Choi, Yoonhyung Kim, Chanho Jung, Changick Kim |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2022 | Improving the Transferability of Targeted Adversarial Examples through Object-Based Diverse InputabstractThe transferability of adversarial examples allows the deception on black-box models, and transfer-based targeted attacks have attracted a lot of interest due to their practical applicability. To maximize the transfer success rate, adversarial examples should avoid overfitting to the source model, and image augmentation is one of the primary approaches for this. However, prior works utilize simple image transformations such as resizing, which limits input diversity. To tackle this limitation, we propose the object-based diverse input (ODI) method that draws an adversarial image on a 3D object and induces the rendered image to be classified as the target class. Our motivation comes from the humans' superior perception of an image printed on a 3D object. If the image is clear enough, humans can recognize the image content in a variety of viewing conditions. Likewise, if an adversarial example looks like the target class to the model, the model should also classify the rendered image of the 3D object as the target class. The ODI method effectively diversifies the input by leveraging an ensemble of multiple source objects and randomizing viewing conditions. In our experimental results on the ImageNet-Compatible dataset, this method boosts the average targeted attack success rate from 28.3% to 47.0% compared to the state-of-the-art methods. We also demonstrate the applicability of the ODI method to adversarial examples on the face verification task and its superior performance improvement. Our code is available at https://github.com/dreamflake/ODI. Junyoung Byun, Seungju Cho, Myung-Joon Kwon, Hee-Seon Kim, Changick Kim |
CVPR | 5 |
| 2022 | Exploiting Doubly Adversarial Examples for Improving Adversarial RobustnessabstractDeep neural networks have shown outstanding performance in various areas, but adversarial examples can easily fool them. Although strong adversarial attacks have defeated diverse adversarial defense methods, adversarial training, which augments training data with adversarial examples, remains an effective defense strategy. To further improve adversarial robustness, this paper exploits adversarial examples of adversarial examples. We observe that these doubly adversarial examples tend to return to the original prediction on the clean images but sometimes drift toward other classes. From this finding, we propose a regularization loss that prevents these drifts, which mitigates the vulnerability against multi-targeted attacks. Experimental results on the CIFAR-10 and CIFAR-100 datasets empirically show that the proposed loss improves adversarial robustness. Junyoung Byun, Hyojun Go, Seungju Cho, Changick Kim |
ICIP | 4 |
| 2022 | Hidden Conditional Adversarial AttacksabstractDeep neural networks are vulnerable to maliciously crafted inputs called adversarial examples. Research on unprecedented adversarial attacks is significant since it can help strengthen the reliability of neural networks by alarming potential threats against them. However, since existing adversarial attacks disturb models unconditionally, the resulting adversarial examples increase their detectability through statistical observations or human inspection. To tackle this limitation, we propose hidden conditional adversarial attacks whose resultant adversarial examples disturb models only if the input images satisfy attackers’ pre-defined conditions. These hidden conditional adversarial examples have better stealthiness and controllability of their attack ability. Our experimental results on the CIFAR-10 and ImageNet datasets show their effectiveness and raise a serious concern about the vulnerability of CNNs against the novel attacks. Junyoung Byun, Kyujin Shim, Hyojun Go, Changick Kim |
ICIP | 4 |
| 2022 | Adversarial Training with Channel Attention RegularizationabstractAdversarial attack shows that deep neural networks (DNNs) are highly vulnerable to small perturbation. Currently, one of the most effective ways to defend against adversarial attacks is adversarial training, which generates adversarial examples during training and induces the models to classify them correctly. To further increase robustness, various techniques such as exploiting additional unlabeled data and novel training loss have been proposed. In this paper, we propose a novel regularization method that exploits latent features, which can be easily combined with existing approaches. We discover that particular channels are more sensitive to adversarial perturbation, motivating us to propose regularizing these channels. Specifically, we attach a channel attention module for adjusting sensitivity of each channel by reducing the difference between the latent feature of the natural image and that of the adversarial image, which we call Channel Attention Regularization (CAR). CAR can be combined with the existing adversarial training framework, showing that it improves the robustness of state-of-the-art defense models. Experiments on various existing adversarial training methods against diverse attacks show the effectiveness of our methods. Codes are available at https://github.com/sgmath12/Adversarial-Training-CAR. Seungju Cho, Junyoung Byun, Myung-Joon Kwon, Changick Kim |
ICIP | 5 |
| 2022 | Temporal Flow Mask Attention for Open-Set Long-Tailed Recognition of Wild Animals in Camera-Trap ImagesabstractCamera traps, unmanned observation devices, and deep learning-based image recognition systems have greatly reduced human effort in collecting and analyzing wildlife images. However, data collected via above apparatus exhibits 1) long-tailed and 2) open-ended distribution problems. To tackle the open-set long-tailed recognition problem, we propose the Temporal Flow Mask Attention Network that comprises three key building blocks: 1) an optical flow module, 2) an attention residual module, and 3) a meta-embedding classifier. We extract temporal features of sequential frames using the optical flow module and learn informative representation using attention residual blocks. Moreover, we show that applying the meta-embedding technique boosts the performance of the method in open-set long-tailed recognition. We apply this method on a Korean De-militarized Zone (DMZ) dataset. We conduct extensive experiments, and quantitative and qualitative analyses to prove that our method effectively tackles the open-set long-tailed recognition problem while being robust to unknown classes. Jeongsoo Kim, Sangmin Woo, Byeongjun Park, Changick Kim |
ICIP | 4 |
| 2022 | DAT: Domain Adaptive Transformer for Domain Adaptive Semantic SegmentationabstractUnsupervised domain adaptation (UDA) for semantic segmentation aims to predict class annotations on an unlabeled target dataset by training on a rich labeled source dataset. It is crucial in UDA semantic segmentation to decrease the domain gap by learning domain invariant feature representations across both domains. In this paper, we propose a novel transformer-based network, called a domain adaptive transformer (DAT), using a self-training scheme. We introduce domain invariant attention (DIA), which enables the DAT to exploit high-level domain invariant features at the patch level. Moreover, an entropy-based selective pseudo-labeling algorithm provides the DAT with reliable pseudo-labels of target samples for domain adaptive self-training, which corrects the noisy pseudo-labels online. We show that our DAT greatly improves the domain adaptability and achieves state-of-the-art results on the SYNTHIA-to-Cityscapes benchmark. Jinyoung Park 0001, Minseok Son, Changick Kim |
ICIP | 4 |
| 2022 | Adaptive Warping Network for Transferable Adversarial AttacksabstractDeep Neural Networks (DNNs) are extremely susceptible to adversarial examples, which are crafted by intentionally adding imperceptible perturbations to clean images. Due to potential threats of adversarial attacks in practice, black-box transfer-based attacks are carefully studied to identify the vulnerability of DNNs. Unfortunately, transfer-based attacks often fail to achieve high transferability because the adversarial examples tend to overfit the source model. Applying input transformation is one of the most effective methods to avoid such overfitting. However, most previous input transformation methods obtain limited transferability because these methods utilize fixed transformations for all images. To solve the problem, we propose an Adaptive Warping Network (AWN), which searches for appropriate warping to the individual data. Specifically, AWN optimizes the warping, which mitigates the effect of adversarial perturbations in each iteration. The adversarial examples are generated to become robust against such strong transformations. Extensive experimental results on the ImageNet dataset demonstrate that AWN outperforms the existing input transformation methods in terms of transferability. Minji Son, Myung-Joon Kwon, Hee-Seon Kim, Junyoung Byun, Seungju Cho, Changick Kim |
ICIP | 6 |
| 2022 | Dynamic Template Update for Visual Object TrackingabstractSiamese-based trackers have recently demonstrated impressive performance and high speed. Despite their great success, conventional siamese trackers are prone to be fooled when facing appearance variations of target objects because they refer to fixed templates captured from first frames to track target objects in the rest of videos. To address this issue, we propose a novel siamese-based tracking framework utilizing a dual template which consists of a static template and a dynamic template. The dynamic template is updated every update interval and allows the tracker to catch appearance variations of the target over time. Furthermore, we introduce a reliability score which prevents incorrect dynamic templates from degrading tracking performance to ensure reliable dynamic template updates. Experimental results show that our method possesses better discriminability and robustness than the baseline, which utilizes a single static template. Kibum Yun, Kyujin Shim, Kangwook Ko, Changick Kim |
ICIP | 4 |
| 2022 | Explanation-based Graph Neural Networks for Graph ClassificationabstractGraph Neural Network models can be used to quickly analyze interactions between multiple data expressed in a graph structure, with high accuracy. Previous studies accurately extract subgraphs which have a significant influence on the whole graph, providing accurate explanations for predictions of GNN. We noted that explanation components could help improve classification performance as unique representations of each class. Therefore, we suggest the GNN performance can be further improved by using explanation components. In this paper, we propose an Explanation-Based Graph Neural Networks (EBGNN) that utilizes contrastive learning at the instance level, by applying explanation components. In EBGNN, the explanation components ensure similarity for instances within the same class, and promote separability for instances in different classes. Finally, we conducted an evaluation on five benchmark datasets (MUTAG, IMDB-BINARY, PROTEINS, NCI1, and DD). Our experiment showed a significant increase in graph classification performance compared to state-of-the-art methods. Sangwoo Seo, Seungjun Jung, Changick Kim |
ICPR | 3 |
| 2022 | Geometrically Adaptive Dictionary Attack on Face RecognitionabstractCNN-based face recognition models have brought remarkable performance improvement, but they are vulnerable to adversarial perturbations. Recent studies have shown that adversaries can fool the models even if they can only access the models’ hard-label output. However, since many queries are needed to find imperceptible adversarial noise, reducing the number of queries is crucial for these attacks. In this paper, we point out two limitations of existing decision-based black-box attacks. We observe that they waste queries for background noise optimization, and they do not take advantage of adversarial perturbations generated for other images. We exploit 3D face alignment to overcome these limitations and propose a general strategy for query-efficient black-box attacks on face recognition named Geometrically Adaptive Dictionary Attack (GADA). Our core idea is to create an adversarial perturbation in the UV texture map and project it onto the face in the image. It greatly improves query efficiency by limiting the perturbation search space to the facial area and effectively recycling previous perturbations. We apply the GADA strategy to two existing attack methods and show overwhelming performance improvement in the experiments on the LFW and CPLFW datasets. Furthermore, we also present a novel attack strategy that can circumvent query similarity-based stateful detection that identifies the process of query-based black-box attacks. Junyoung Byun, Hyojun Go, Changick Kim |
WACV | 3 |
| 2022 | On the Effectiveness of Small Input Noise for Defending Against Query-based Black-Box AttacksabstractWhile deep neural networks show unprecedented performance in various tasks, the vulnerability to adversarial examples hinders their deployment in safety-critical systems. Many studies have shown that attacks are also possible even in a black-box setting where an adversary cannot access the target model’s internal information. Most black-box attacks are based on queries, each of which obtains the target model’s output for an input, and many recent studies focus on reducing the number of required queries. In this paper, we pay attention to an implicit assumption of query-based black-box adversarial attacks that the target model’s output exactly corresponds to the query input. If some randomness is introduced into the model, it can break the assumption, and thus, query-based attacks may have tremendous difficulty in both gradient estimation and local search, which are the core of their attack process. From this motivation, we observe even a small additive input noise can neutralize most query-based attacks and name this simple yet effective approach Small Noise Defense (SND). We analyze how SND can defend against query-based black-box attacks and demonstrate its effectiveness against eight state-of-the-art attacks with CIFAR-10 and ImageNet datasets. Even with strong defense ability, SND almost maintains the original classification accuracy and computational speed. SND is readily applicable to pre-trained models by adding only one line of code at the inference. Junyoung Byun, Hyojun Go, Changick Kim |
WACV | 3 |
| 2022 | Learning JPEG Compression Artifacts for Image Manipulation Detection and Localization
Myung-Joon Kwon, Seung-Hun Nam, In-Jae Yu, Heung-Kyu Lee, Changick Kim |
Int. J. Comput. Vis. | 5 |
| 2021 | Meta Batch-Instance Normalization for Generalizable Person Re-IdentificationabstractAlthough supervised person re-identification (Re-ID) methods have shown impressive performance, they suffer from a poor generalization capability on unseen domains. Therefore, generalizable Re-ID has recently attracted growing attention. Many existing methods have employed an instance normalization technique to reduce style variations, but the loss of discriminative information could not be avoided. In this paper, we propose a novel generalizable Re-ID framework, named Meta Batch-Instance Normalization (MetaBIN). Our main idea is to generalize normalization layers by simulating unsuccessful generalization scenarios beforehand in the meta-learning pipeline. To this end, we combine learnable batch-instance normalization layers with meta-learning and investigate the challenging cases caused by both batch and instance normalization layers. Moreover, we diversify the virtual simulations via our meta-train loss accompanied by a cyclic inner-updating manner to boost generalization capability. After all, the MetaBIN framework prevents our model from overfitting to the given source styles and improves the generalization capability to unseen domains without additional data augmentation or complicated network design. Extensive experimental results show that our model outperforms the state-of-the-art methods on the large-scale domain generalization Re-ID benchmark and the cross-domain Re-ID problem. The source code is available at: https://github.com/bismex/MetaBIN. Seokeon Choi, Taekyung Kim 0002, Minki Jeong, Hyoungseob Park, Changick Kim |
CVPR | 5 |
| 2021 | Few-Shot Open-Set Recognition by Transformation ConsistencyabstractIn this paper, we attack a few-shot open-set recognition (FSOSR) problem, which is a combination of few-shot learning (FSL) and open-set recognition (OSR). It aims to quickly adapt a model to a given small set of labeled samples while rejecting unseen class samples. Since OSR requires rich data and FSL considers closed-set classification, existing OSR and FSL methods show poor performances in solving FSOSR problems. The previous FSOSR method follows the pseudo-unseen class sample-based methods, which collect pseudo-unseen samples from the other dataset or synthesize samples to model unseen class representations. However, this approach is heavily dependent on the composition of the pseudo samples. In this paper, we propose a novel unknown class sample detector, named SnaTCHer, that does not require pseudo-unseen samples. Based on the transformation consistency, our method measures the difference between the transformed prototypes and a modified prototype set. The modified set is composed by replacing a query feature and its predicted class prototype. SnaTCHer rejects samples with large differences to the transformed prototypes. Our method alters the unseen class distribution estimation problem to a relative feature transformation problem, independent of pseudo-unseen class samples. We investigate our SnaTCHer with various prototype transformation methods and observe that our method consistently improves unseen class sample detection performance without closed-set classification reduction. Minki Jeong, Seokeon Choi, Changick Kim |
CVPR | 3 |
| 2021 | Just a Few Points are All You Need for Multi-view Stereo: A Novel Semi-supervised Learning Method for Multi-view StereoabstractWhile learning-based multi-view stereo (MVS) methods have recently shown successful performances in quality and efficiency, limited MVS data hampers generalization to unseen environments. A simple solution is to generate various large-scale MVS datasets, but generating dense ground truth for 3D structure requires a huge amount of time and resources. On the other hand, if the reliance on dense ground truth is relaxed, MVS systems will generalize more smoothly to new environments. To this end, we first introduce a novel semi-supervised multi-view stereo framework called a Sparse Ground truth-based MVS Network (SGT-MVSNet) that can reliably reconstruct the 3D structures even with a few ground truth 3D points. Our strategy is to divide the accurate and erroneous regions and individually conquer them based on our observation that a probability map can separate these regions. We propose a self-supervision loss called the 3D Point Consistency Loss to enhance the 3D reconstruction performance, which forces the 3D points back-projected from the corresponding pixels by the predicted depth values to meet at the same 3D co-ordinates. Finally, we propagate these improved depth pre-dictions toward edges and occlusions by the Coarse-to-fine Reliable Depth Propagation module. We generate the spare ground truth of the DTU dataset for evaluation and extensive experiments verify that our SGT-MVSNet outperforms the state-of-the-art MVS methods on the sparse ground truth setting. Moreover, our method shows comparable reconstruction results to the supervised MVS methods though we only used tens and hundreds of ground truth 3D points. Taekyung Kim 0002, Seokeon Choi, Dongki Jung, Changick Kim |
ICCV | 5 |
| 2021 | DnD: Dense Depth Estimation in Crowded Dynamic Indoor ScenesabstractWe present a novel approach for estimating depth from a monocular camera as it moves through complex and crowded indoor environments, e.g., a department store or a metro station. Our approach predicts absolute scale depth maps over the entire scene consisting of a static background and multiple moving people, by training on dynamic scenes. Since it is difficult to collect dense depth maps from crowded indoor environments, we design our training framework without requiring depths produced from depth sensing devices. Our network leverages RGB images and sparse depth maps generated from traditional 3D reconstruction methods to estimate dense depth maps. We use two constraints to handle depth for non-rigidly moving people without tracking their motion explicitly. We demonstrate that our approach offers consistent improvements over recent depth estimation methods on the NAVERLABS dataset, which includes complex and crowded scenes. Dongki Jung, Yonghan Lee 0001, Deokhwa Kim, Changick Kim, Dinesh Manocha |
ICCV | 5 |
| 2021 | Rethinking Training Schedules For Verifiably Robust Networks
Hyojun Go, Junyoung Byun, Changick Kim |
ICIP | 3 |
| 2021 | Fine-Grained Multi-Class Object CountingabstractMany animal species in the wild are at the risk of extinction. To deal with this situation, ecologists have monitored the population changes of endangered species. However, the current wildlife monitoring method is extremely laborious as the animals are counted manually. Automated counting of animals by species can facilitate this work and further renew the ways for ecological studies. However, to the best of our knowledge, few works and publicly available datasets have been proposed on multi-class object counting which is applicable to counting several animal species. In this paper, we propose a fine-grained multi-class object counting dataset, named KR-GRUIDAE, which contains endangered red-crowned crane and white-naped crane in the family Gruidae. We also propose a specialized network for multi-class object counting and line segment density maps, and show their effectiveness by comparing results of existing crowd counting methods on the KR-GRUIDAE dataset. Hyojun Go, Junyoung Byun, Byeongjun Park, Myung-Ae Choi, Seunghwa Yoo, Changick Kim |
ICIP | 6 |
| 2021 | A Parameter Efficient Multi-Scale Capsule NetworkabstractCapsule networks consider spatial relationships in an input image. The relationship-based feature propagation in capsule networks shows promising results. However, a large number of trainable parameters limit their widespread use. In this paper, we propose Decomposed Capsule Network (DCN) to reduce the number of training parameters in the primary capsule generation stage. Our DCN represents a capsule as a combination of basis vectors. Generating basis vectors and their coefficients notably reduce the total number of training parameters. Moreover, we introduce an extension of the DCN architecture, named Multi-scale Decomposed Capsule Network (MDCN). The MDCN architecture integrates features from multiple scales to synthesize capsules with fewer parameters. Our proposed networks show better performance on the Fashion-MNIST dataset and the CIFAR10 dataset with fewer parameters than the original network. Minki Jeong, Changick Kim |
ICIP | 2 |
| 2021 | Understanding Vqa For Negative Answers Through Visual And Linguistic InferenceabstractIn order to make Visual Question Answering (VQA) explainable, previous studies not only visualize the attended region of a VQA model, but also generate textual explanations for its answers. However, when the model’s answer is “no,” existing methods have difficulty in revealing detailed arguments that lead to that answer. In addition, previous methods are insufficient to provide logical bases when the question requires common sense to answer. In this paper, we propose a novel textual explanation method to overcome the aforementioned limitations. First, we extract keywords that are essential to infer an answer from a question. Second, we utilize a novel Variable-Constrained Beam Search (VCBS) algorithm to generate explanations that best describe the circumstances in images. Furthermore, if the answer to the question is “yes” or “no,” we apply Natural Langauge Inference (NLI) to determine if contents of the question can be inferred from the explanation using common sense. Our user study, conducted in Amazon Mechanical Turk (MTurk), shows that our proposed method generates more reliable explanations compared to the previous methods. Moreover, by modifying the VQA model’s answer through the output of the NLI model, we show that VQA performance increases by 1.1% from the original model. Seungjun Jung, Junyoung Byun, Kyujin Shim, Sanghyun Hwang, Changick Kim |
ICIP | 5 |
| 2021 | Weakly-Supervised Multiple Object Tracking Via A Masked Center Point Warping LossabstractMultiple object tracking (MOT), a popular subject in computer vision with broad application areas, aims to detect and track multiple objects across an input video. However, recent learning-based MOT methods require strong supervision on both the bounding box and the ID of each object for every frame used during training, which induces a heightened cost for obtaining labeled data. In this paper, we propose a weakly-supervised MOT framework that enables the accurate tracking of multiple objects while being trained without object ID ground truth labels. Our model is trained only with the bounding box information with a novel masked warping loss that drives the network to indirectly learn how to track objects through a video. Specifically, valid object center points in the current frame are warped with the predicted offset vector and enforced to be equal to the valid object center points in the previous frame. With this approach, we obtain an MOT accuracy on par with those of the state-of-the-art fully supervised MOT models, which use both the bounding boxes and object ID as ground truth labels, on the MOT17 dataset. Sungjoon Yoon, Kyujin Shim, Kayoung Park, Changick Kim |
ICIP | 4 |
| 2021 | Temporal filtering networks for online action detection
Hyunjun Eun, Jinyoung Moon, Jongyoul Park, Chanho Jung, Changick Kim |
Pattern Recognit. | 5 |
| 2020 | Hi-CMD: Hierarchical Cross-Modality Disentanglement for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) is an important task in night-time surveillance applications, since visible cameras are difficult to capture valid appearance information under poor illumination conditions. Compared to traditional person re-identification that handles only the intra-modality discrepancy, VI-ReID suffers from additional cross-modality discrepancy caused by different types of imaging systems. To reduce both intra- and cross-modality discrepancies, we propose a Hierarchical Cross-Modality Disentanglement (Hi-CMD) method, which automatically disentangles ID-discriminative factors and ID-excluded factors from visible-thermal images. We only use ID-discriminative factors for robust cross-modality matching without ID-excluded factors such as pose or illumination. To implement our approach, we introduce an ID-preserving person image generation network and a hierarchical feature learning module. Our generation network learns the disentangled representation by generating a new cross-modality image with different poses and illuminations while preserving a person's identity. At the same time, the feature learning module enables our model to explicitly extract the common ID-discriminative characteristic between visible-infrared images. Extensive experimental results demonstrate that our method outperforms the state-of-the-art methods on two VI-ReID datasets. The source code is available at: https://github.com/bismex/HiCMD. Seokeon Choi, Youngeun Kim, Taekyung Kim 0002, Changick Kim |
CVPR | 5 |
| 2020 | Learning to Discriminate Information for Online Action DetectionabstractFrom a streaming video, online action detection aims to identify actions in the present. For this task, previous methods use recurrent networks to model the temporal sequence of current action frames. However, these methods overlook the fact that an input image sequence includes background and irrelevant actions as well as the action of interest. For online action detection, in this paper, we propose a novel recurrent unit to explicitly discriminate the information relevant to an ongoing action from others. Our unit, named Information Discrimination Unit (IDU), decides whether to accumulate input information based on its relevance to the current action. This enables our recurrent network with IDU to learn a more discriminative representation for identifying ongoing actions. In experiments on two benchmark datasets, TVSeries and THUMOS-14, the proposed method outperforms state-of-the-art methods by a significant margin. Moreover, we demonstrate the effectiveness of our recurrent unit by conducting comprehensive ablation studies. Hyunjun Eun, Jinyoung Moon, Jongyoul Park, Chanho Jung, Changick Kim |
CVPR | 5 |
| 2020 | Attract, Perturb, and Explore: Learning a Feature Alignment Network for Semi-supervised Domain Adaptation
Taekyung Kim 0002, Changick Kim |
ECCV (14) | 2 |
| 2020 | Reinforcement Learning-Based Layer-Wise Quantization For Lightweight Deep Neural NetworksabstractNetwork quantization has been widely studied to compress the deep neural network in mobile devices. Conventional methods quantize the network parameters of all layers with the same fixed precision, regardless of the number of parameters in each layer. However, quantizing the weights of the layer with many parameters is more effective in reducing the model size. Accordingly, in this paper, we propose a novel mixed-precision quantization method based on reinforcement learning. Specifically, we utilize the number of parameters at each layer as a prior for our framework. By using the accuracy and the bit-width as a reward, the proposed framework determines the optimal quantization policy for each layer. By applying this policy sequentially, we achieve weighted-average 2.97 bits for the VGG-16 model on the CIFAR-10 dataset with no degradation of the accuracy, compared with its full-precision baseline. We also show that our framework can provide an optimal quantization policy for the VGG-Net and the ResNet to minimize the storage while preserving the accuracy. Juri Jung, Jonghee Kim, Youngeun Kim, Changick Kim |
ICIP | 4 |
| 2020 | Arbitrary Style Transfer Using Graph Instance NormalizationabstractStyle transfer is the image synthesis task, which applies a style of one image to another while preserving the content. In statistical methods, the adaptive instance normalization (AdaIN) whitens the source images and applies the style of target images through normalizing the mean and variance of features. However, computing feature statistics for each instance would neglect the inherent relationship between features, so it is hard to learn global styles while fitting to the individual training dataset. In this paper, we present a novel learnable normalization technique for style transfer using graph convolutional networks, termed Graph Instance Normalization (GrIN). This algorithm makes the style transfer approach more robust by taking into account similar information shared between instances. Besides, this simple module is also applicable to other tasks like image-to-image translation or domain adaptation. Dongki Jung, Seunghan Yang, Changick Kim |
ICIP | 4 |
| 2020 | Self-Training Of Graph Neural Networks Using Similarity Reference For Robust Training With Noisy LabelsabstractFiltering noisy labels is crucial for robust training of deep neural networks. To train networks with noisy labels, sampling methods have been introduced, which sample the reliable instances to update networks using only sampled data. Since they rarely employ the non-sampled data for training, these methods have a fundamental limitation that they reduce the amount of the training data. To alleviate this problem, our approach aims to fully utilize the whole dataset by leveraging the information of the sampled data. To this end, we propose a novel graph-based learning framework that enables networks to propagate the label information of the sampled data to adjacent data, whether they are sampled or not. Also, we propose a novel self-training strategy to utilize the non-sampled data without labels and to regularize the network update using the information of the sampled data. Our method outperforms state-of-the-art sampling methods. Hyoungseob Park, Minki Jeong, Youngeun Kim, Changick Kim |
ICIP | 4 |
| 2020 | Multi-Step Quantization Of A Multi-Scale Network For Crowd CountingabstractCrowd counting is one of the most important tasks in visual surveillance applications since it provides useful information such as the number of crowds and their distribution. However, it is very challenging due to severe occlusions, large geometrical deformations, and high visual clutter. To tackle this problem, we propose a novel CNN-based crowd density estimation network consisting of a backbone, decoder, and mapper, and also a multi-step quantization scheme to train the network more effectively. As a backbone network, ResNet is adopted, then the decoder and mapper are added to deal with multi-scale problems of crowd counting and to generate high-resolution density maps. Finally, a multi-step quantization scheme discretizes the continuous space of both predictions and ground truth density maps, and it reduces the search scope of the network and raises their matching ratio. As a result, our method outperforms recent methods in four major datasets. Kyujin Shim, Junyoung Byun, Changick Kim |
ICIP | 3 |
| 2020 | Semi-Supervised Domain Adaptation via Selective Pseudo Labeling and Progressive Self-TrainingabstractDomain adaptation (DA) is a representation learning methodology that transfers knowledge from a label-sufficient source domain to a label-scarce target domain. While most of early methods are focused on unsupervised DA (UDA), several studies on semi-supervised DA (SSDA) are recently suggested. In SSDA, a small number of labeled target images are given for training, and the effectiveness of those data is demonstrated by the previous studies. However, the previous SSDA approaches solely adopt those data for embedding ordinary supervised losses, overlooking the potential usefulness of the few yet informative clues. Based on this observation, in this paper, we propose a novel method that further exploits the labeled target images for SSDA. Specifically, we utilize labeled target images to selectively generate pseudo labels for unlabeled target images. In addition, based on the observation that pseudo labels are inevitably noisy, we apply a label noise-robust learning scheme, which progressively updates the network and the set of pseudo labels by turns. Extensive experimental results show that our proposed method outperforms other previous state-of-the-art SSDA methods. Yoonhyung Kim, Changick Kim |
ICPR | 2 |
| 2020 | The Korean Sign Language Dataset for Action Recognition
Seunghan Yang, Seungjun Jung, Heekwang Kang, Changick Kim |
MMM (1) | 4 |
| 2020 | RPM-Net: Robust Pixel-Level Matching Networks for Self-Supervised Video Object SegmentationabstractIn this paper, we introduce a self-supervised approach for video object segmentation without human labeled data. Specifically, we present Robust Pixel-level Matching Networks (RPM-Net), a novel deep architecture that matches pixels between adjacent frames, using only color information from unlabeled videos for training. Technically, RPM-Net can be separated in two main modules. The embedding module first projects input images into high dimensional embedding space. Then the matching module with deformable convolution layers matches pixels between reference and target frames based on the embedding features. Unlike previous methods using deformable convolution, our matching module adopts deformable convolution to focus on similar features in spatio-temporally neighboring pixels. Our experiments show that the selective feature sampling improves the robustness to challenging problems in video object segmentation such as camera shake, fast motion, deformation, and occlusion. Also, we carry out comprehensive experiments on three public datasets (i.e., DAVIS-2017, SegTrack-v2, and Youtube-Objects) and achieve state-of-the-art performance on self-supervised video object segmentation. Moreover, we significantly reduce the performance gap between self-supervised and fully-supervised video object segmentation (41.0% vs. 52.5% on DAVIS-2017 validation set). Youngeun Kim, Seokeon Choi, Hankyeol Lee, Taekyung Kim 0002, Changick Kim |
WACV | 5 |
| 2020 | Combinational Class Activation Maps for Weakly Supervised Object LocalizationabstractWeakly supervised object localization has recently attracted attention since it aims to identify both class labels and locations of objects by using image-level labels. Most previous methods utilize the activation map corresponding to the highest activation source. Exploiting only one activation map of the highest probability class is often biased into limited regions or sometimes even highlights background regions. To resolve these limitations, we propose to use activation maps, named combinational class activation maps (CCAM), which are linear combinations of activation maps from the highest to the lowest probability class. By using CCAM for localization, we suppress background regions to help highlighting foreground objects more accurately. In addition, we design the network architecture to consider spatial relationships for localizing relevant object regions. Specifically, we integrate non-local modules into an existing base network at both low- and high-level layers. Our final model, named non-local combinational class activation maps (NL-CCAM), obtains superior performance compared to previous methods on representative object localization benchmarks including ILSVRC 2016 and CUB- 200-2011. Furthermore, we show that the proposed method has a great capability of generalization by visualizing other datasets. Seunghan Yang, Yoonhyung Kim, Youngeun Kim, Changick Kim |
WACV | 4 |
| 2020 | Dual Back-Projection-Based Internal Learning for Blind Super-ResolutionabstractState-of-the-art super-resolution (SR) methods commonly assume that the downscaling kernel (the point spread function of the camera) is a Gaussian kernel. Therefore, the methods are particularly vulnerable to solving the problem of blind SR, which deals with a real low-resolution (LR) image that does not follow the assumption. Recently, to address this issue, several internal learning-based methods, which train an image-specific network using a single input image, have been introduced. In this approach, the blind SR is modeled as a two-stage optimization problem, which conducts downscaling kernel estimation followed by SR network training with the estimated kernel. In this letter, we assume that not only the estimated kernel can be employed for SR network training, but also the super-resolved image can contribute to downscaling kernel estimation. To that end, we propose a unified internal learning-based blind SR method that jointly trains two image-specific networks for 1) downscaling together with kernel estimation and 2) SR. More specifically, we train the two networks to minimize a dual back-projection loss; the SR network is trained to reconstruct a given input image from the LR image generated by the downscaling network, and the downscaling network is trained to downscale the high-resolution image generated by the SR network to be as close as possible to the given input image. Through the complementary training, the downscaling kernel estimation becomes more accurate, resulting in better SR performance. In the experiment, we show that the proposed method outperforms previous two-stage internal learning-based methods in terms of both SR performance and efficiency. Jonghee Kim, Chanho Jung, Changick Kim |
IEEE Signal Process. Lett. | 3 |
| 2020 | SRG: Snippet Relatedness-Based Temporal Action Proposal GeneratorabstractRecent temporal action proposal generation approaches have suggested integrating segment- and snippet score-based methodologies to produce proposals with high recall and accurate boundaries. In this paper, different from such a hybrid strategy, we focus on the potential of the snippet score-based approach. Specifically, we propose a new snippet score-based method, named Snippet Relatedness-based Generator (SRG), with a novel concept of “snippet relatedness”. Snippet relatedness represents which snippets are related to a specific action instance. To effectively learn this snippet relatedness, we present “pyramid non-local operations” for locally and globally capturing long-range dependencies among snippets. By employing these components, SRG first produces a 2D relatedness score map that enables the generation of various temporal intervals reliably covering most action instances with high overlap. Then, SRG evaluates the action confidence scores of these temporal intervals and refines their boundaries to obtain temporal action proposals. On THUMOS-14 and ActivityNet-1.3 datasets, SRG outperforms state-of-the-art methods for temporal action proposal generation. Furthermore, compared to competing proposal generators, SRG leads to significant improvements in temporal action detection. Hyunjun Eun, Jinyoung Moon, Jongyoul Park, Chanho Jung, Changick Kim |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2020 | Blind Deblurring of Text Images Using a Text-Specific Hybrid DictionaryabstractIn this paper, we propose a blind text image deblurring algorithm by using a text-specific hybrid dictionary. After careful analysis, we find that the text-specific hybrid dictionary has the great ability of providing powerful contextual information for text image deblurring. Here, it is worth noting that our proposed method is inspired by our observation that an intermediate latent image contains not only sharp regions, but also multiple types of small blurred regions. Based upon our discovery, we propose a prior for text images based on sparse representation, which models the relationship between an intermediate latent image and a desired sharp image. To this end, we carefully collect three different image patch pairs, which are 1) Gaussian blur-sharp, 2) motion blur-sharp, and 3) sharp-sharp, in order to construct the text-specific hybrid dictionary. We also propose a new optimization framework suitable for the task of text image deblurring in this paper. Extensive experiments have been conducted on a challenging dataset of synthetic and real-world text images. Our results demonstrate that the proposed method outperforms the state-of-the-art image deblurring methods both quantitatively and qualitatively. Hyukzae Lee, Chanho Jung, Changick Kim |
IEEE Trans. Image Process. | 3 |
| 2019 | Pseudo-Labeling Curriculum for Unsupervised Domain Adaptation
Minki Jeong, Taekyung Kim 0002, Changick Kim |
BMVC | 4 |
| 2019 | Bilinear Siamese Networks with Background Suppression for Visual Object Tracking
Hankyeol Lee, Seokeon Choi, Youngeun Kim, Changick Kim |
BMVC | 4 |
| 2019 | Diversify and Match: A Domain Adaptive Representation Learning Paradigm for Object DetectionabstractWe introduce a novel unsupervised domain adaptation approach for object detection. We aim to alleviate the imperfect translation problem of pixel-level adaptations, and the source-biased discriminativity problem of feature-level adaptations simultaneously. Our approach is composed of two stages, i.e., Domain Diversification (DD) and Multi-domain-invariant Representation Learning (MRL). At the DD stage, we diversify the distribution of the labeled data by generating various distinctive shifted domains from the source domain. At the MRL stage, we apply adversarial learning with a multi-domain discriminator to encourage feature to be indistinguishable among the domains. DD addresses the source-biased discriminativity, while MRL mitigates the imperfect image translation. We construct a structured domain adaptation framework for our learning paradigm and introduce a practical way of DD for implementation. Our method outperforms the state-of-the-art methods by a large margin of 3%~11% in terms of mean average precision (mAP) on various datasets. Taekyung Kim 0002, Minki Jeong, Seunghyeon Kim, Seokeon Choi, Changick Kim |
CVPR | 5 |
| 2019 | Self-Ensembling With GAN-Based Data Augmentation for Domain Adaptation in Semantic SegmentationabstractDeep learning-based semantic segmentation methods have an intrinsic limitation that training a model requires a large amount of data with pixel-level annotations. To address this challenging issue, many researchers give attention to unsupervised domain adaptation for semantic segmentation. Unsupervised domain adaptation seeks to adapt the model trained on the source domain to the target domain. In this paper, we introduce a self-ensembling technique, one of the successful methods for domain adaptation in classification. However, applying self-ensembling to semantic segmentation is very difficult because heavily-tuned manual data augmentation used in self-ensembling is not useful to reduce the large domain gap in the semantic segmentation. To overcome this limitation, we propose a novel framework consisting of two components, which are complementary to each other. First, we present a data augmentation method based on Generative Adversarial Networks (GANs), which is computationally efficient and effective to facilitate domain alignment. Given those augmented images, we apply self-ensembling to enhance the performance of the segmentation network on the target domain. The proposed method outperforms state-of-the-art semantic segmentation methods on unsupervised domain adaptation benchmarks. Taekyung Kim 0002, Changick Kim |
ICCV | 3 |
| 2019 | Self-Training and Adversarial Background Regularization for Unsupervised Domain Adaptive One-Stage Object DetectionabstractDeep learning-based object detectors have shown remarkable improvements. However, supervised learning-based methods perform poorly when the train data and the test data have different distributions. To address the issue, domain adaptation transfers knowledge from the label-sufficient domain (source domain) to the label-scarce domain (target domain). Self-training is one of the powerful ways to achieve domain adaptation since it helps class-wise domain adaptation. Unfortunately, a naive approach that utilizes pseudo-labels as ground-truth degenerates the performance due to incorrect pseudo-labels. In this paper, we introduce a weak self-training (WST) method and adversarial background score regularization (BSR) for domain adaptive one-stage object detection. WST diminishes the adverse effects of inaccurate pseudo-labels to stabilize the learning procedure. BSR helps the network extract discriminative features for target backgrounds to reduce the domain shift. Two components are complementary to each other as BSR enhances discrimination between foregrounds and backgrounds, whereas WST strengthen class-wise discrimination. Experimental results show that our approach effectively improves the performance of the one-stage object detection in unsupervised domain adaptation setting. Seunghyeon Kim, Taekyung Kim 0002, Changick Kim |
ICCV | 4 |
| 2019 | CNN-Based Semantic Segmentation Using Level Set LossabstractThesedays, Convolutional Neural Networks are widely used in semantic segmentation. However, since CNN-based segmentation networks produce low-resolution outputs with rich semantic information, it is inevitable that spatial details (e.g., small objects and fine boundary information) of segmentation results will be lost. To address this problem, motivated by a variational approach to image segmentation (i.e., level set theory), we propose a novel loss function called the level set loss which is designed to refine spatial details of segmentation results. To deal with multiple classes in an image, we first decompose the ground truth into binary images. Note that each binary image consists of background and regions belonging to a class. Then we convert level set functions into class probability maps and calculate the energy for each class. The network is trained to minimize the weighted sum of the level set loss and the cross-entropy loss. The proposed level set loss improves the spatial details of segmentation results in a time and memory efficient way. Furthermore, our experimental results show that the proposed loss function achieves better performance than previous approaches. Youngeun Kim, Seunghyeon Kim, Taekyung Kim 0002, Changick Kim |
WACV | 4 |
| 2019 | Contextual Information Based Quality Assessment for Contrast-Changed ImagesabstractIn this letter, we propose the objective metric that can precisely predict the perceptual quality of contrast-changed images using inter-pixel contextual information. The metric consists of two parts. One is a two-dimensional (2-D) histogram-based contrast quality measure that utilizes the distribution of the gray-level differences between adjacent pixels. We design the desired 2-D histogram considering the characteristic of an adequately high contrast image and predict contrast quality by comparing the desired 2-D histogram with 2-D histograms of an original image and a contrast-changed image. The other is a spatial entropy based one that uses the information of spatial location distribution of gray-levels. A comparison is carried out with many IQA metrics on five contrast related databases. Experimental results show that the proposed metric provides a more accurate prediction of human perception of contrast change than other metrics. Daeyeong Kim, Seungyoun Lee, Changick Kim |
IEEE Signal Process. Lett. | 3 |
| 2019 | Skeleton-Based Gait Recognition via Robust Frame-Level MatchingabstractGait is a useful biometric feature for human identification in video surveillance applications since it can be obtained without subject cooperation. In recent years, model-based gait recognition using a 3D skeleton has been widely studied through view-invariant modeling and kinematic gait analysis. However, existing methods integrate all frame-level feature vectors using the same criterion, even though skeleton information is highly sensitive to changes in covariate conditions such as clothing, carrying, and occlusion. The scheme inevitably reduces the frame-level discriminative power and eventually degrades performance. Instead, we propose a robust frame-level matching method for gait recognition that minimizes the influence of noisy patterns as well as secures the frame-level discriminative power. To this end, we measure the skeleton quality in terms of body symmetry for each frame. Based on the quality, we construct a quality-adjusted cost matrix between input frames and registered frames to prevent matching with noisy patterns. Our two-stage linear matching is then applied to the cost matrix to compute a frame-level discriminative score including similarity and margin. In the end, the identity of a probe is determined by a weighted majority voting scheme via frame-level scores. It enhances the robustness against inaccurate skeleton estimation results by assigning different weights for each frame based on the score. Our approach outperforms the state-of-the-art methods on three public datasets (UPCVgait, UPCVgaitK2, and SDUgait) and a new gait dataset which we create with consideration of unpredictable behaviors while walking. In addition, we demonstrate that our method is robust to skeleton estimation error, partial occlusion, and data loss. The CILgait dataset and MATLAB code are available at https://sites.google.com/site/seokeonchoi/gait-recognition. Seokeon Choi, Jonghee Kim, Wonjun Kim 0001, Changick Kim |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2019 | A Quad Edge-Based Grid Encoding Model for Content-Aware Image RetargetingabstractIn this paper, we present a novel grid encoding model for content-aware image retargeting. In contrast to previous approaches such as vertex-based and axis-aligned grid encoding models, our approach takes each horizontal/vertical distance between two adjacent vertices as an optimization variable. Upon this difference-based encoding scheme, every vertex position of a target grid is subsequently determined after optimizing the one-dimensional values. Our quad edge-based grid model has two major advantages for image retargeting. First, the model enables a grid optimization problem to be developed in a simple quadratic program while ensuring the global convexity of objective functions. Second, due to the independency of variables, spatial regularizations can be applied in a locally adaptive manner to preserve structural components. Based on this model, we propose three quadratic objective functions. Note that, in our work, their linear combination guides a grid deformation process to obtain a visually comfortable retargeting result by preserving salient regions and structural components of an input image. Comparative evaluations have been conducted with ten existing state-of-the-art image retargeting methods, and the results show that our method built upon the quad edge-based model consistently outperforms other previous methods both on qualitative and quantitative perspectives. Yoonhyung Kim, Hyunjun Eun, Chanho Jung, Changick Kim |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2018 | BitNet: Learning-Based Bit-Depth Expansion
Junyoung Byun, Kyujin Shim, Changick Kim |
ACCV (2) | 3 |
| 2018 | Water-Filling: An Efficient Algorithm for Digitized Document Shadow Removal
Seungjun Jung, Muhammad Abul Hasan, Changick Kim |
ACCV (1) | 3 |
| 2018 | Fast Korean Text Detection and Recognition in Traffic Guide SignsabstractIn this paper, we propose a fast method based on deep neural networks to detect and recognize Korean characters in traffic guide signs. To detect character candidates quickly, we first employ a region proposal network (RPN) which is in this paper ResNet-18, being relatively shallow. We also apply the Inception architecture to residual blocks for reducing parameters of the network. After character candidates are detected, we classify them into 709 Korean characters by using a classification network (CLSN). Similar to the RPN, our CLSN consists of residual blocks with the Inception architecture. In experiments, we achieved 97.69 % of accuracy at 5.9fps on both detection and recognition of Korean characters in traffic guide signs. Hyunjun Eun, Jonghee Kim, Changick Kim |
VCIP | 4 |
| 2018 | Effect of Using Object Shape Prior on Visual Object CountingabstractVisual object counting aims to count the number of objects in a given image or video. Among many object counting methods, the counting by density estimation method draws attention because of its capability of counting and its object localization ability. The method utilizes the object density map of the image with multiple objects. The density is estimated by a regression model which learns the mapping between the local features of the given image and the density map generated from its corresponding object locations. Unlike conventional methods that only rely on object locations for density map generation, in this paper, we show that the system performance can be increased by considering object shapes as well as object locations. To this end, we propose two approaches to generate the ground truth density map from the object locations. Both methods generate the density map which reflects structural features of the objects. We show that the regression models trained with density maps which reflect the object shape outperforms the models trained with density maps generated by the conventional density map generation method on several challenging benchmarks. In other words, we observe that it is essential to generate the ground truth density map according to object shape in the image. Minki Jeong, Changick Kim |
VCIP | 2 |
| 2018 | Weakly Supervised Semantic Segmentation Using Color Adjacency LossabstractLarge amount of training data is essential for deep learning-based computer vision tasks. However, in semantic segmentation, annotating pixel-wise labels for large-scale image data is laborious and time-consuming. To handle this problem, we propose a training framework for a CNN-based network using sparse labels. We propagate the sparse labels to produce the same performance as training on dense labels in the segmentation network. For effective label propagation, we take advantage of the observation that adjacent pixels sharing similar colors would be in the same class. Based on this insight, the label is propagated by our adjacency loss depending on the color similarity between the adjacent pixels. We perform on the PASCAL VOC 2012 dataset using scribbles annotations as sparse labels. The proposed algorithm achieves superior performance compared to the previous method in weakly supervised semantic segmentation task. Youngeun Kim, Taekyung Kim 0002, Seunghyeon Kim, Changick Kim |
VCIP | 4 |
| 2018 | A structure-aware axis-aligned grid deformation approach for robust image retargeting
Yoonhyung Kim, Seungjun Jung, Chanho Jung, Changick Kim |
Multim. Tools Appl. | 4 |
| 2018 | Saliency refinement: Towards a uniformly highlighted salient object
Hyunjun Eun, Yoonhyung Kim, Chanho Jung, Changick Kim |
Signal Process. Image Commun. | 4 |
| 2018 | Ramp Distribution-Based Image Enhancement Techniques for Infrared ImagesabstractA novel image enhancement method for infrared (IR) images is presented. The proposed method consists of two parts considering the characteristics of high-dynamic-range IR images. First, we attempt to enhance image contrast by introducing the ramp distribution that increases with a constant slope in an ordered histogram domain. The ramp-distributed histogram is incorporated into an optimization problem with a sorted histogram of the input image to calculate a modified histogram. Second, to deal with blurred effects on IR images, we propose a relative edge-strength index for high-boost filtering to effectively suppress noise in relatively uniform regions. Compared with various conventional and state-of-the-art algorithms, the proposed method shows highly competitive performance. Seungyoun Lee, Daeyeong Kim, Changick Kim |
IEEE Signal Process. Lett. | 3 |
| 2017 | Contrast Enhancement Using Combined 1-D and 2-D Histogram-Based TechniquesabstractThis letter presents an adaptive contrast enhancement algorithm considering both preservation of the shape of a one-dimensional (1-D) histogram and statistical information on the gray-level differences between neighboring pixels obtained by a 2-D histogram. The proposed system consists of two modules. One is to enhance the entire contrast by stretching the 1-D histogram while preserving the shape of the histogram. The other is to improve the details of nonsmooth areas occurring frequently in input images. These are formulated into a single constrained optimization problem. Compared with several state-of-the-art enhancement algorithms, the proposed algorithm shows highly competitive performance. Daeyeong Kim, Changick Kim |
IEEE Signal Process. Lett. | 2 |
| 2016 | A framework for automatic static and dynamic video thumbnail extraction
Changick Kim |
Multim. Tools Appl. | 2 |
| 2016 | Object-aware image thumbnailing using image classification and enhanced detection of ROI
Changick Kim |
Multim. Tools Appl. | 2 |
| 2016 | Superpixel-Guided Adaptive Image SmoothingabstractIn edge-preserving image smoothing, edge blurriness and structural edge attenuation have been common problems. L0smoothing successfully solves these two problems by adopting L0norm of gradients. However, a weak structural edge diminishing problem still exists because L0penalty first removes small nonzero gradients. In order to address this problem, we propose superpixel-guided adaptive image smoothing by introducing an adaptive parameter into L0smoothing framework. The adaptive smoothing parameter is efficiently computed in a cascade manner. In the first stage, we allocate smoothing parameters to the pixels consisting of details. More importantly, we then exploit similarities between a pixel and its surrounding superpixels for assigning smoothing parameters to the rest pixels. Experimental results demonstrate that our proposed method efficiently preserves structural edges regardless of their scales compared to previous methods. Hyunjun Eun, Changick Kim |
IEEE Signal Process. Lett. | 2 |
| 2016 | Face and Hair Region Labeling Using Semi-Supervised Spectral Clustering-Based Multiple SegmentationsabstractThe multiple segmentation (MS) scheme is considered to be a way to get a better spatial support for various shaped objects in image segmentation. The MS scheme assumes that the segmented regions (i.e., segments) can be treated as hypotheses for object support rather than mere partitionings of the image. As for attaining each segmentation in the MS scheme, one of the most popular methods is to employ spectral clustering (SC). When applied to image segmentation tasks, SC groups a set of pixels or small regions into unique segments. While it has been popularly used in image segmentation, it often fails to deal with images containing objects with complex boundaries. To split the image as close to the object boundaries as possible, some prior knowledge can be used to guide the clustering algorithm toward appropriate partitioning of the data. In semisupervised clustering, prior knowledge is often formulated as pairwise constraints. In this paper, we propose an MS technique combined with constrained SC to build a face and hair region labeler. To put it concretely, pairwise constraints modified to fit the problem of labeling face regions are added to SC and multiple segments are generated by the constrained SC. Then, the labeling is conducted by estimating the likelihoods for each segment to belong to the target object classes. Experiments are conducted on three datasets and the results show that the proposed scheme offers useful tools for labeling the face images. Ilkoo Ahn, Changick Kim |
IEEE Trans. Multim. | 2 |
| 2015 | A robust real time system for remote heart rate measurement via cameraabstractHeart rate (HR) is an important indicator of human health status. Traditional heart rate measurement methods rely on contact-based sensors or electrodes, which are inconvenient and troublesome for users. Remote sensing of the photoplephysmography (PPG) signal using a video camera provides a promising means to monitor vital signs of people without the need of any physical contact. However, until recently, most of the literature papers approaching this problem have only reported results from off-line recording videos taken under well controlled environments. In this paper, we propose a method to improve HR measurement accuracy under challenging environments involving factors such as subjects movement, complicated facial models (i.e., hair, glass, beards, etc.), subjects' distance to camera, and low illumination condition. We also build a framework for real-time measuring system and construct a stable model for recording and displaying results for long term heart rate monitoring. We tested our system on challenging dataset, and demonstrated that our method not only deals with real-time, on-line measurement tasks, but also outperforms others' works. Duc Nhan Tran, Hyukzae Lee, Changick Kim |
ICME | 3 |
| 2015 | Salient object detection using HOS based L0 smoothing and shape-aware region mergingabstractRecent salient object detection algorithms often involve a segmentation step to produce saliency maps preserving boundaries. However, over-segmented results that many segmentation methods produce confuse to describe object boundaries. In this paper, we present a novel salient object detection algorithm which produces reliable salient object candidates. First, the input image is processed by Higher Order Statistics (HOS) based L0smoothing to highlight strong edges and reduce texture. We then apply image segmentation to the HOS L0smoothed image to produce improved results in which the number of over-segmented regions is greatly reduced. Second, we propose shape-aware region merging with a novel region scale measure. Finally, a saliency map from the merging result is generated by taking two simple saliency cues. Extensive experiments on a challenging saliency dataset indicate that our algorithm has comparable performance against state-of-the-arts. Hyunjun Eun, Jonghee Kim, Changick Kim |
VCIP | 3 |
| 2014 | Aesthetic quality classification via subject region extractionabstractAesthetic quality classification of photos gains growing interest in recent years. In this paper, we propose an aesthetic quality classification method via subject region extraction. We extract the subject region by a combination of clear region detection and saliency detection. Once the subject regions are extracted, we extract regional features to measure contrast between the subject and background regions since people usually emphasize objects by focusing them. Global features are used to describe comprehensive properties of the image. Experimental results show that our classification performance outperforms the state-of-the-art aesthetic quality classification methods even if we do not use prior knowledge of a visual content. Jonghee Kim, Changick Kim |
ICIP | 2 |
| 2014 | Blurred image region detection and segmentationabstractEstimating both defocus blur and motion blur regions from a single monocular image is a challenging research area in modern image processing and computer vision. Most existing algorithms first divide an image into either non-blur or blur patches. Then blur-type classification is performed on the blur patches only. This means that such approaches include potential risk that incorrect blur region identification may affect the following blur-type classification. In this paper, we present a novel framework for blur region identification to overcome the deficiency of classical methods. We propose a 3-way blur identification method, which divides an image into non-blur, defocus blur, and motion blur regions at once. To this end, we employ intuitive and powerful features based on specific criteria well-suited for our 3-way classification problem. We also take a coarse-to-fine technique to produce pixelwise segmentation results. Experimental results demonstrate that our proposed method outperforms the recent algorithms. Hyukzae Lee, Changick Kim |
ICIP | 2 |
| 2014 | A novel monochromatic cue for detecting regions of visual interest
Chanho Jung, Seungwoo Yoo, Changick Kim |
Image Vis. Comput. | 4 |
| 2014 | Determining the Existence of Objects in an Image and Its Application to Image ThumbnailingabstractIn recent years, computer vision applications dealing with foreground objects are becoming more important with an increasing demand of advanced intelligent systems. Most of these applications assume that an image contains one or more objects, which often produce undesired results when noticeable objects do not appear in the image. In this letter, we address the problem of ascertaining the existence of objects in an image. In the first step, the input image is partitioned into nonoverlapping local patches, then the patches are categorized into three classes, namely natural, man-made, and object to estimate object candidates. Then a Bayesian methodology is employed to produce more reliable results by eliminating false positives. To boost the object patch detection performance, we exploit the difference between coarse and fine segmentation results. To demonstrate the effectiveness of the proposed method, extensive experiments have been conducted on several benchmark image databases. Furthermore, we have shown the usefulness of our approach by applying it to a real application (i.e., image thumbnailing). Chanho Jung, Changick Kim |
IEEE Signal Process. Lett. | 4 |
| 2014 | Spatiotemporal Saliency Detection Using Textural Contrast and Its ApplicationsabstractSaliency detection has been extensively studied due to its promising contributions for various computer vision applications. However, most existing methods are easily biased toward edges or corners, which are statistically significant, but not necessarily relevant. Moreover, they often fail to find salient regions in complex scenes due to ambiguities between salient regions and highly textured backgrounds. In this paper, we present a novel unified framework for spatiotemporal saliency detection based on textural contrast. Our method is simple and robust, yet biologically plausible; thus, it can be easily extended to various applications, such as image retargeting, object segmentation, and video surveillance. Based on various datasets, we conduct comparative evaluations of 12 representative saliency detection models presented in the literature, and the results show that the proposed scheme outperforms other previously developed methods in detecting salient regions of the static and dynamic scenes. Changick Kim |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2013 | Floor detection based depth estimation from a single indoor sceneabstractEstimating depth information from a single image has recently attracted great attention in various vision-based applications such as mobile robot navigation. Although there are numerous depth map generation methods, little effort has been done on the depth estimation from a single indoor scene. In this paper, we propose a novel method for estimating depth from a single indoor image via nonlinear diffusion and image segmentation techniques. One important advantage of our approach is that no learning scheme is required to estimate a depth map. Based on the proposed method, we obtain visually plausible depth estimation results even with the presence of occlusions or clutters in the single indoor image. From experimental results, we confirm that the proposed algorithm provides reliable depth information under various indoor environments. Changhwan Chun, Dong Jin Park, Changick Kim |
ICIP | 4 |
| 2013 | Background subtraction using hybrid feature coding in the bag-of-features framework
Seungwoo Yoo, Changick Kim |
Pattern Recognit. Lett. | 2 |
| 2013 | Active Contours Driven by the Salient Edge Energy ModelabstractIn this brief, we present a new indicator, i.e., salient edge energy, for guiding a given contour robustly and precisely toward the object boundary. Specifically, we define the salient edge energy by exploiting the higher order statistics on the diffusion space, and incorporate it into a variational level set formulation with the local region-based segmentation energy for solving the problem of curve evolution. In contrast to most previous methods, the proposed salient edge energy allows the curve to find only significant local minima relevant to the object boundary even in the noisy and cluttered background. Moreover, the segmentation performance derived from our new energy is less sensitive to the size of local windows compared with other recently developed methods, owing to the ability of our energy function to suppress diverse clutters. The proposed method has been tested on various images, and experimental results show that the salient edge energy effectively drives the active contour both qualitatively and quantitatively compared to various state-of-the-art methods. Changick Kim |
IEEE Trans. Image Process. | 2 |
| 2012 | Unsupervised detection of surface defects: A two-step approachabstractIn this paper, we focus on the problem of finding anomalies in surface images. Despite enormous research efforts and advances, it still remains a big challenge to be solved. This paper proposes a unified approach for defect detection. Our proposed method consists of two phases: (1) global estimation and (2) local refinement. First, we roughly estimate defects by applying a spectral-based approach in a global manner. We then locally refine the estimated region based on the distributions of pixel intensities derived from defect and defect-free regions. Experimental results show that the proposed method outperforms the previous defect detection methods and gives robust results even in noisy surface defect images. Changick Kim |
ICIP | 2 |
| 2012 | Depth-Based Disocclusion Filling for Virtual View SynthesisabstractFree-viewpoint rendering (FVR) has become a popular topic in 3D research. A promising technology in FVR is to generate virtual views using a single texture image and the corresponding depth image. A critical problem that occurs when generating virtual views is that the regions covered by the foreground objects in the original view may be disoccluded in the synthesized views. In this paper, a depth based disocclusion filling algorithm using patch based texture synthesis is proposed. In contrast to the existing patch based virtual view synthesis methods, the filling priority is driven by the robust structure tensor and the epipolar directional term. Moreover, the best-matched patch is searched in the background regions and finally the best-matched patch is chosen by considering the color similarity and some factors such as epipolar line and the magnitude of data term. Superiority of the proposed method over the existing methods is proved by comparing the experimental results. Ilkoo Ahn, Changick Kim |
ICME | 2 |
| 2012 | Real-time estimation of 3D scene geometry from a single image
Chanho Jung, Changick Kim |
Pattern Recognit. | 2 |
| 2012 | Background Subtraction for Dynamic Texture Scenes Using Fuzzy Color HistogramsabstractBackground subtraction is a very popular approach for detecting moving objects from a still scene. For this, most of previous methods depend on the assumption that the background is static over short time periods. However, structured motion patterns of the background (e.g., waving leaves, spouting fountain, rippling water, etc.), which are distinctive from variations due to noise, are hardly tolerated in this assumption and thus still lead to high-level false positive rates when using previous models. In this letter, we introduce a novel background subtraction algorithm for temporally dynamic texture scenes. Specifically, we propose to adopt a clustering-based feature, called fuzzy color histogram (FCH), which has an ability of greatly attenuating color variations generated by background motions while still highlighting moving objects. Experimental results demonstrate that the proposed method is effective for background subtraction in dynamic texture scenes compared to several competitive methods proposed in the literature. Changick Kim |
IEEE Signal Process. Lett. | 2 |
| 2012 | A Unified Spectral-Domain Approach for Saliency Detection and Its Application to Automatic Object SegmentationabstractIn this paper, a visual attention model is incorporated for efficient saliency detection, and the salient regions are employed as object seeds for our automatic object segmentation system. In contrast with existing interactive segmentation approaches that require considerable user interaction, the proposed method does not require it, i.e., the segmentation task is fulfilled in a fully automatic manner. First, we introduce a novel unified spectral-domain approach for saliency detection. Our visual attention model originates from a well-known property of the human visual system that the human visual perception is highly adaptive and sensitive to structural information in images rather than nonstructural information. Then, based on the saliency map, we propose an iterative self-adaptive segmentation framework for more accurate object segmentation. Extensive tests on a variety of cluttered natural images show that the proposed algorithm is an efficient indicator for characterizing the human perception and it can provide satisfying segmentation performance. Chanho Jung, Changick Kim |
IEEE Trans. Image Process. | 2 |
| 2011 | A novel image importance model for content-aware image resizingabstractThis paper presents a novel method for adaptively resizing a given image to fit the dimensions of arbitrary displays. For the success of this content-aware image resizing, the image importance model needs to be carefully defined since it guides further resizing procedures. In our work, we focus on the excellence of the local dominance for measuring the image importance in a sense of human visual perception, which tends to strongly respond to the dominant structure in a local region. In contrast to most previous approaches not allowing for underlying image structure, the proposed model effectively represents the spatial contexts, which are indeed salient regions, even under severe distortion. The proposed method has been extensively tested and the results show that the proposed scheme is more effective for the image resizing when compared to various state-of-the-art methods. Changick Kim |
ICIP | 2 |
| 2011 | A Texture-Aware Salient Edge Model for Image RetargetingabstractImage retargeting aims at adapting a given image to fit the size of arbitrary displays without severe visual distortions. To achieve this task successfully, it is essential to define a reliable image importance map (IIM) since it guides subsequent retargeting procedures. In this letter, we introduce a novel IIM for effective image retargeting. Specifically, we define our IIM by exploiting the higher order statistics (HOS) of the diffusion space for image retargeting. We call it texture-aware salient edge (TASE) map. Based on the proposed TASE map, we obtain visually acceptable retargeting results, even in the cluttered background and in the presence of noise as well. The proposed method has been extensively tested, and experimental results show that the proposed scheme is effective for image retargeting compared to other various state-of-the-art methods. Changick Kim |
IEEE Signal Process. Lett. | 2 |
| 2011 | Spatiotemporal Saliency Detection and Its Applications in Static and Dynamic ScenesabstractThis paper presents a novel method for detecting salient regions in both images and videos based on a discriminant center-surround hypothesis that the salient region stands out from its surroundings. To this end, our spatiotemporal approach combines the spatial saliency by computing distances between ordinal signatures of edge and color orientations obtained from the center and the surrounding regions and the temporal saliency by simply computing the sum of absolute difference between temporal gradients of the center and the surrounding regions. Our proposed method is computationally efficient, reliable, and simple to implement and thus it can be easily extended to various applications such as image retargeting and moving object extraction. The proposed method has been extensively tested and the results show that the proposed scheme is effective in detecting saliency compared to various state-of-the-art methods. Chanho Jung, Changick Kim |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2010 | Automatic segmentation of salient objects using iterative reversible graph cutabstractThere have been several interactive approaches to extracting objects from still images, since it is significantly difficult to automatically segment objects in complex background. In this paper, we present a novel automatic scheme for extracting salient objects from natural images. To this end, segmentation of salient objects is formulated as a global energy minimization problem in an iterative self-adaptive framework. By employing a saliency detection technique, object and background seeds are inferred automatically. The problem in this step is that the automatically generated seeds may not be reliably positioned. An iterative reversible graph cut method is introduced to overcome the problem inherent in the saliency-based seed extraction method. In the iterative self-adaptive framework, bidirectional state transitions are iteratively involved to reduce the mis-classified pixels. Experimental results show that the proposed segmentation method yields more accurate segmentation results than previous segmentation approaches. Chanho Jung, Changick Kim |
ICME | 3 |
| 2010 | Saliency detection: A self-ordinal resemblance approachabstractIn saliency detection, regions attracting visual attention need to be highlighted while effectively suppressing non-salient regions for the semantic scene understanding. However, most previous methods tend to fail in suppressing highly textured backgrounds and also high contrast edges belonging to the non-salient regions. To address this problem, we propose a method for detecting salient regions based on a self-ordinal resemblance measure (SORM). Our saliency map is defined by using the center-surround computations based on the ordinal signatures obtained from local regions centered at each pixel. It can be regarded as an energy map and thus extended to image retargeting. Our approach is fully automatic and nonparametric. To justify robustness of our approach, the proposed method is compared with the state of the art methods on various images. Chanho Jung, Changick Kim |
ICME | 3 |
| 2009 | An Efficient Correction Method of Wide-angle Lens Distortion for Surveillance SystemsabstractWide-angle or fish-eye lenses are popularly used for the surveillance system due to their large field of view (FOV). However, images obtained by wide-angle cameras tend to be nonlinearly distorted owing to lens optics. In this paper, we propose a novel framework to correct the wide-angle lens distortion for surveillance systems. Our approach is based on the FOV model, which is an efficient and simple correction method. The FOV model works well for typical wide-angle lens, whose FOV is usually smaller than 150deg. However, it begins to reveal problems as the FOV increases. First, we address two main problems of the FOV model and then improve the FOV model by refining the distortion curve. The proposed method is tested on images and videos with different FOV. Experimental results show the efficiency of our algorithm. Changick Kim |
ISCAS | 2 |
| 2009 | A New Approach for Overlay Text Detection and Extraction From Complex Video SceneabstractOverlay text brings important semantic clues in video content analysis such as video information retrieval and summarization, since the content of the scene or the editor's intention can be well represented by using inserted text. Most of the previous approaches to extracting overlay text from videos are based on low-level features, such as edge, color, and texture information. However, existing methods experience difficulties in handling texts with various contrasts or inserted in a complex background. In this paper, we propose a novel framework to detect and extract the overlay text from the video scene. Based on our observation that there exist transient colors between inserted text and its adjacent background, a transition map is first generated. Then candidate regions are extracted by a reshaping method and the overlay text regions are determined based on the occurrence of overlay text in each candidate. The detected overlay text regions are localized accurately using the projection of overlay text pixels in the transition map and the text extraction is finally conducted. The proposed method is robust to different character size, position, contrast, and color. It is also language independent. Overlay text region update between frames is also employed to reduce the processing time. Experiments are performed on diverse videos to confirm the efficiency of the proposed method. Changick Kim |
IEEE Trans. Image Process. | 2 |
| 2007 | An Intelligent Display Scheme of Soccer Video on Mobile DevicesabstractThe rapid progress of the multimedia signal processing has contributed to the extensive use of multimedia devices with small liquid crystal display (LCD) panel. With the proliferation of mobile devices with a small display, the video sequences captured for normal viewing on standard TV or HDTV may give the small-display viewers uncomfortable experience in understanding what is happening in a scene. For instance, in a soccer video sequence taken by a long-shot camera technique, the tiny objects (e.g., soccer ball and players) may not be clearly visible on the small LCD panel. Thus, an intelligent display technique is needed for small-display viewers. To this end, one of the key technologies is to determine region of interest (ROI), whenever ROI display is considered more effective. Keewon Seo, Jaeseung Ko, Ilkoo Ahn, Changick Kim |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2005 | Spatiotemporal sequence matching for efficient video copy detection
Changick Kim, Vasudev Bhaskaran |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2005 | Segmenting a Low-Depth-of-Field Image Using Morphological Filters and Region MergingabstractWe propose a novel algorithm to partition an image with low depth-of-field (DOF) into focused object-of-interest (OOI) and defocused background. The proposed algorithm unfolds into three steps. In the first step, we transform the low-DOF image into an appropriate feature space, in which the spatial distribution of the high-frequency components is represented. This is conducted by computing higher order statistics (HOS) for all pixels in the low-DOF image. Next, the obtained feature space, which is called HOS map in this paper, is simplified by removing small dark holes and bright patches using a morphological filter by reconstruction. Finally, the OOI is extracted by applying region merging to the simplified image and by thresholding. Unlike the previous methods that rely on sharp details of OOI only, the proposed algorithm complements the limitation of them by using morphological filters, which also allows perfect preservation of the contour information. Compared with the previous methods, the proposed method yields more accurate segmentation results, supporting faster processing. Changick Kim |
IEEE Trans. Image Process. | 1 |
| 2004 | Segmenting focused objects using morphological filters and region mergingabstractWe propose a novel algorithm to partition an image with low depth-of-field (DOF) into focused object-of-interest (OOI) and defocused background. The proposed algorithm consists of three steps. In the first step, we transform the low DOF image into an appropriate feature space for the partition. This is conducted by computing higher-order statistics (HOS) for all pixels in the low DOF image. Next, the obtained feature space, which is called HOS map, is simplified by removing small dark holes and bright patches using a morphological filter by reconstruction. Finally, the OOI is extracted by applying region merging and adaptive thresholding to the simplified image. Experimental results show that the proposed method yields more accurate segmentation results than the previous methods. Changick Kim |
VCIP | 1 |
| 2003 | Content-based image copy detection
Changick Kim |
Signal Process. Image Commun. | 1 |
| 2002 | Adaptive post-filtering for reducing blocking and ringing artifacts in low bit-rate video coding
Changick Kim |
Signal Process. Image Commun. | 1 |
| 2002 | Fast and automatic video object segmentation and tracking for content-based applicationsabstractThe new video-coding standard MPEG-4 enables content-based functionality, as well as high coding efficiency, by taking into account shape information of moving objects. A novel algorithm for segmentation of moving objects in video sequences and extraction of video object planes (VOPs) is proposed . For the case of multiple video objects in a scene, the extraction of a specific single video object (VO) based on connected components analysis and smoothness of VO displacement in successive frames is also discussed. Our algorithm begins with a robust double-edge map derived from the difference between two successive frames. After removing edge points which belong to the previous frame, the remaining edge map, moving edge (ME), is used to extract the VOP. The proposed algorithm is evaluated on an indoor sequence captured by a low-end camera as well as MPEG-4 test sequences and produces promising results. Changick Kim, Jenq-Neng Hwang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2002 | Object-based video abstraction for video surveillance systemsabstractKey frames are the subset of still images which best represent the content of a video sequence in an abstracted manner. In other words, video abstraction transforms an entire video clip to a small number of representative images. We present a scheme for object-based video abstraction facilitated by an efficient video-object segmentation (VOS) system. In such a framework, the concept of a "key frame" is replaced by that of a "key video-object plane (VOP)." In order to achieve an online object-based framework such as an object-based video surveillance system, it becomes essential that semantically meaningful video objects are directly accessed from video sequences. Moreover, the extraction of key VOPs needs to be automated and context dependent so that they maintain the important contents of the video while removing all redundancies. Once a VOP is extracted, the shape of the VOP needs to be well described. To this end, both region-based and contour-based shape descriptors are investigated, and the region-based descriptor is selected for the proposed system. The key VOPs are extracted in a sequential manner by successive comparison with the previously declared key VOP. Experimental results on the proposed online processing scheme combined with efficient VOS show the proposed integrated scheme generates desirable summarizations of surveillance videos. Changick Kim, Jenq-Neng Hwang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2001 | Object-based video abstraction using cluster analysisabstractAmong various semantic primitives of video, objects of interest along with their actions and generated events can play an important role in some applications such as an object-based video surveillance system and an object based video indexing/retrieval system. In this paper, we propose an object-based video abstraction algorithm by cluster analysis using the mean shift algorithm. The generated clusters, called segments in this paper, can be used as a small unit in the object-based video indexing/retrieval systems. In the proposed algorithm, Hu's (1962) seven moments are used as shape descriptors for each video object plane (VOP), and shape distance between two VOPs is measured by using weighted Euclidean distance. Promising experimental results on the proposed scheme are presented. Changick Kim, Jenq-Neng Hwang |
ICIP (2) | 1 |
| 2000 | Reliable and Fast Fingerprint Identification for Security ApplicationsabstractFingerprint identification is one of the most popular and reliable personal biometric identification methods. This paper describes an on-line fingerprint identification system consisting of fingerprint image preprocessing, feature extraction and matching. The preprocessing part includes image enhancing steps to acquire binarized and skeletonized ridges, which are needed for feature points extraction. Feature points (minutiae) such as endpoints, bifurcations, and core point are then extracted, followed by false minutiae elimination. The fast and robust matching algorithm is proposed, which is a correlation-based method that runs over the 1/8 sized feature points maps. Sanpachai Huvanandana, Changick Kim, Jenq-Neng Hwang |
ICIP | 2 |
| 2000 | An integrated scheme for object-based video abstractionabstractIn this paper, we present a novel scheme for object-based key-frame extraction facilitated by an efficient video object segmentation system. Key-frames are the subset of still images which best represent the content of a video sequence in an abstracted manner. Thus, key-frame based video abstraction transforms an entire video clip to a small number of representative images. The challenge is that the extraction of key-frames needs to be automated and context dependent so that they maintain the important contents of the video while remove all redundancy. Among various semantic primitives of video, objects of interest along with their actions and generated events can play an important role in some applications such as object-based video surveillance system. Furthermore, on-line processing combined with fast and robust video object segmentation is crucial for real-time applications to report unwanted action or event as soon as it happens. Experimental results on the proposed scheme for object-based video abstraction are presented. Changick Kim, Jenq-Neng Hwang |
ACM Multimedia | 1 |
| 1999 | A Fast and Robust Moving Object Segmentation in Video SequencesabstractThe new video coding standard MPEG-4 is enabling content-based functionalities as well as high coding efficiency considering shape information of moving objects. A novel algorithm for segmentation of moving objects in video sequences and VOP (video object planes) extraction is presented. This algorithm begins with a robust double edge map from the difference between two successive frames. After removing edges which belong to previous frame, the edge map, named ME (moving edge) is used to extract VOP. The proposed algorithm is evaluated for MPEG-4 test sequences and produces promising results. Changick Kim, Jenq-Neng Hwang |
ICIP (2) | 1 |