EDBT 2026 Demo / reviewers in the wild / expert
Junmo Kim 0002
dblp:40/240-2 · also Jun-Mo Kim 0002
· DBLP profile ↗
125ranked-venue papers
0as first author
65since 2021 · last 2026
0000-0002-7174-7932ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 93 · 44 since 2021Artificial intelligence and machine learning · 84 · 49 since 2021Systems, architecture and hardware · 6 · 5 since 2021Computer networks · 2Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Forget What Matters, Keep the Rest: Selective Unlearning of Informative TokensabstractUnlearning in large language models (LLMs) has emerged as a promising safeguard against adversarial behaviors.When the forgetting loss is applied uniformly without considering tokenlevel semantic importance, model utility can be unnecessarily degraded.Recent studies have explored token-wise loss regularizers that prioritize informative tokens, but largely rely on ground-truth confidence or external linguistic parsers, which limits their ability to capture contextual information or the model's overall predictive state.Intuitively, function words like "the" primarily serve syntactic roles and are highly predictable with little ambiguity, but informative words admit multiple plausible alternatives with greater uncertainty.Based on this intuition, we propose Entropy-guided Token Weighting (ETW), a token-level unlearning regularizer that uses entropy of the predictive distribution as a proxy for token informativeness.We demonstrate that informative tokens tend to have higher entropy, whereas structural tokens tend to have lower entropy.This behavior enables ETW to achieve more effective unlearning while better preserving model utility than existing token-level approaches.Q. Seunghee Koh, Sunghyun Baek, Youngdong Kim, Junmo Kim 0002 |
ACL (1) | 4 |
| 2025 | Efficient Dynamic Scene Editing via 4D Gaussian-based Static-Dynamic SeparationabstractRecent 4D dynamic scene editing methods require editing thousands of 2D images used for dynamic scene synthesis and updating the entire scene with additional training loops, resulting in several hours of processing to edit a single dynamic scene. Therefore, these methods are not scalable with respect to the temporal dimension of the dynamic scene (i.e., the number of timesteps). In this work, we propose Instruct-4DGS, an efficient dynamic scene editing method that is more scalable in terms of temporal dimension. To achieve computational efficiency, we leverage a 4D Gaussian representation that models a 4D dynamic scene by combining static 3D Gaussians with a Hexplane-based deformation field, which captures dynamic information. We then perform editing solely on the static 3D Gaussians, which is the minimal but sufficient component required for visual editing. To resolve the misalignment between the edited 3D Gaussians and the deformation field, which may arise from the editing process, we introduce a refinement stage using a score distillation mechanism. Extensive editing results demonstrate that Instruct-4DGS is efficient, reducing editing time by more than half compared to existing methods while achieving high-quality edits that better follow user instructions. JooHyun Kwon, Hanbyel Cho, Junmo Kim 0002 |
CVPR | 3 |
| 2025 | Controllable Feature Whitening for Hyperparameter-Free Bias Mitigation
Yooshin Cho, Hanbyel Cho, Janghyeon Lee 0001, Hyeong Gwon Hong, Jaesung Ahn, Junmo Kim 0002 |
ICCV | 6 |
| 2025 | DMQ: Dissecting Outliers of Diffusion Models for Post-Training QuantizationabstractDiffusion models have achieved remarkable success in image generation but come with significant computational costs, posing challenges for deployment in resource-constrained environments. Recent post-training quantization (PTQ) methods have attempted to mitigate this issue by focusing on the iterative nature of diffusion models. However, these approaches often overlook outliers, leading to degraded performance at low bit-widths. In this paper, we propose a DMQ which combines Learned Equivalent Scaling (LES) and channel-wise Power-of-Two Scaling (PTS) to effectively address these challenges. Learned Equivalent Scaling optimizes channel-wise scaling factors to redistribute quantization difficulty between weights and activations, reducing overall quantization error. Recognizing that early denoising steps, despite having small quantization errors, crucially impact the final output due to error accumulation, we incorporate an adaptive timestep weighting scheme to prioritize these critical steps during learning. Furthermore, identifying that layers such as skip connections exhibit high inter-channel variance, we introduce channel-wise Power-of-Two Scaling for activations. To ensure robust selection of PTS factors even with small calibration set, we introduce a voting algorithm that enhances reliability. Extensive experiments demonstrate that our method significantly outperforms existing works, especially at low bit-widths such as W4A6 (4-bit weight, 6-bit activation) and W4A8, maintaining high image generation quality and model stability. The code is available at https://github.com/LeeDongYeun/dmq. Dongyeun Lee, Jiwan Hur, Hyounguk Shon, Jae Young Lee 0002, Junmo Kim 0002 |
ICCV | 5 |
| 2025 | FairASR: Fair Audio Contrastive Learning for Automatic Speech Recognition
Jongsuk Kim, Jaemyung Yu, Minchan Kwon, Junmo Kim 0002 |
INTERSPEECH | 4 |
| 2025 | B4DL: A Benchmark for 4D LiDAR LLM in Spatio-Temporal Understanding
Changho Choi, Youngwoo Shin, Gyojin Han, Junmo Kim 0002 |
ACM Multimedia | 5 |
| 2025 | Preference Distillation via Value based Reinforcement LearningabstractDirect Preference Optimization (DPO) is a powerful paradigm to align language models with human preferences using pairwise comparisons.
However, its binary win-or-loss supervision often proves insufficient for training small models with limited capacity.
Prior works attempt to distill information from large teacher models using behavior cloning or KL divergence.
These methods often focus on mimicking current behavior and overlook distilling reward modeling.
To address this issue, we propose \textit{Teacher Value-based Knowledge Distillation} (TVKD), which introduces an auxiliary reward from the value function of the teacher model to provide a soft guide.
This auxiliary reward is formulated to satisfy potential-based reward shaping, ensuring that the global reward structure and optimal policy of DPO are preserved.
TVKD can be integrated into the standard DPO training framework and does not require additional rollouts.
Our experimental results show that TVKD consistently improves performance across various benchmarks and model sizes. Minchan Kwon, Junwon Ko, Kangil Kim, Junmo Kim 0002 |
NeurIPS | 4 |
| 2025 | Video Diffusion Models Excel at Tracking Similar-Looking Objects Without SupervisionabstractDistinguishing visually similar objects by their motion remains a critical challenge in computer vision. Although supervised trackers show promise, contemporary self-supervised trackers struggle when visual cues become ambiguous, limiting their scalability and generalization without extensive labeled data. We find that pre-trained video diffusion models inherently learn motion representations suitable for tracking without task-specific training. This ability arises because their denoising process isolates motion in early, high-noise stages, distinct from later appearance refinement. Capitalizing on this discovery, our self-supervised tracker significantly improves performance in distinguishing visually similar objects, an underexplored failure point for existing methods. Our method achieves up to a 6-point improvement over recent self-supervised approaches on established benchmarks and our newly introduced tests focused on tracking visually similar items. Visualizations confirm that these diffusion-derived motion representations enable robust tracking of even identical objects across challenging viewpoint changes and deformations. Project page: \small{\url{https://chenshuang-zhang.github.io/projects/ted}}. Chenshuang Zhang, Kang Zhang 0008, Joon Son Chung, In-So Kweon, Junmo Kim 0002, Chengzhi Mao |
NeurIPS | 5 |
| 2025 | Reducing the Content Bias for AI-generated Image Detection
Seoyeon Gye, Junwon Ko, Hyounguk Shon, Minchan Kwon, Junmo Kim 0002 |
WACV | 5 |
| 2025 | Beta Sampling is All You Need: Efficient Image Generation Strategy for Diffusion Models Using Stepwise Spectral AnalysisabstractGenerative diffusion models have emerged as a powerful tool for high-quality image synthesis, yet their iterative nature demands significant computational resources. This paper proposes an efficient time step sampling method based on an image spectral analysis of the diffusion process, aimed at optimizing the denoising process. Instead of the traditional uniform distribution-based time step sampling, we introduce a Beta distribution-like sampling technique that prioritizes critical steps in the early and late stages of the process. Our hypothesis is that certain steps exhibit significant changes in image content, while others contribute minimally. We validated our approach using Fourier transforms to measure frequency response changes at each step, revealing substantial low-frequency changes early on and high-frequency adjustments later. Experiments with ADM and Stable Diffusion demonstrated that our Beta Sampling method consistently outperforms uniform sampling, achieving better FID and IS scores, and offers competitive efficiency relative to state-of-the-art methods like AutoDiffusion. This work provides a practical framework for enhancing diffusion model efficiency by focusing computational resources on the most impactful steps, with potential for further optimization and broader application. Haeil Lee, Hansang Lee, Seoyeon Gye, Junmo Kim 0002 |
WACV | 4 |
| 2025 | Enhancing self-supervised visual representation learning through adversarially generated examplesabstractAbstract Self-supervised learning has emerged as a powerful paradigm for leveraging unlabeled data to learn rich feature representations. However, the efficacy of self-supervised models is often limited by the degree and complexity of the augmentations used during training. In this work, we propose a novel framework that enhances self-supervised learning by incorporating a generative network designed to produce adversarial examples that challenge the learning process. By integrating adversarially generated data, our method extends three well-known self-supervised architectures---SimCLR, BYOL, and SimSiam---and improves their generalization and robustness. We evaluate our approach on CIFAR-10, CIFAR-100, and Tiny ImageNet datasets, demonstrating consistent improvements in classification accuracy over baseline models. Notably, our proposed method outperforms standard self-supervised learning techniques, achieving significant gains in top-1 accuracy across all datasets and training epochs. This substantiates our hypothesis that adversarial examples can significantly contribute to the feature learning capabilities of self-supervised models. Furthermore, our findings suggest that the integration of generative networks can serve as a catalyst for the development of more advanced self-supervised learning algorithms. This study lays the groundwork for future research exploring the potential of adversarial training in self-supervised learning and its applications across diverse domains. Mintae Kang, Junmo Kim 0002 |
Neural Comput. Appl. | 2 |
| 2024 | Foreseeing Reconstruction Quality of Gradient Inversion: An Optimization PerspectiveabstractGradient inversion attacks can leak data privacy when clients share weight updates with the server in federated learning (FL). Existing studies mainly use L2 or cosine distance as the loss function for gradient matching in the attack. Our empirical investigation shows that the vulnerability ranking varies with the loss function used. Gradient norm, which is commonly used as a vulnerability proxy for gradient inversion attack, cannot explain this as it remains constant regardless of the loss function for gradient matching. In this paper, we propose a loss-aware vulnerability proxy (LAVP) for the first time. LAVP refers to either the maximum or minimum eigenvalue of the Hessian with respect to gradient matching loss at ground truth. This suggestion is based on our theoretical findings regarding the local optimization of the gradient inversion in proximity to the ground truth, which corresponds to the worst case attack scenario. We demonstrate the effectiveness of LAVP on various architectures and datasets, showing its consistent superiority over the gradient norm in capturing sample vulnerabilities. The performance of each proxy is measured in terms of Spearman's rank correlation with respect to several similarity scores. This work will contribute to enhancing FL security against any potential loss functions beyond L2 or cosine distance in the future. Hyeong Gwon Hong, Yooshin Cho, Hanbyel Cho, Jaesung Ahn, Junmo Kim 0002 |
AAAI | 5 |
| 2024 | Modeling Stereo-Confidence out of the End-to-End Stereo-Matching Network via Disparity Plane SweepabstractWe propose a novel stereo-confidence that can be measured externally to various stereo-matching networks, offering an alternative input modality choice of the cost volume for learning-based approaches, especially in safety-critical systems. Grounded in the foundational concepts of disparity definition and the disparity plane sweep, the proposed stereo-confidence method is built upon the idea that any shift in a stereo-image pair should be updated in a corresponding amount shift in the disparity map. Based on this idea, the proposed stereo-confidence method can be summarized in three folds. 1) Using the disparity plane sweep, multiple disparity maps can be obtained and treated as a 3-D volume (predicted disparity volume), like the cost volume is constructed. 2) One of these disparity maps serves as an anchor, allowing us to define a desirable (or ideal) disparity profile at every spatial point. 3) By comparing the desirable and predicted disparity profiles, we can quantify the level of matching ambiguity between left and right images for confidence measurement. Extensive experimental results using various stereo-matching networks and datasets demonstrate that the proposed stereo-confidence method not only shows competitive performance on its own but also consistent performance improvements when it is used as an input modality for learning-based stereo-confidence methods. Jae Young Lee 0002, Woonghyun Ka, Jaehyun Choi, Junmo Kim 0002 |
AAAI | 4 |
| 2024 | FRED: Towards a Full Rotation-Equivariance in Aerial Image Object DetectionabstractRotation-equivariance is an essential yet challenging property in oriented object detection. While general object detectors naturally leverage robustness to spatial shifts due to the translation-equivariance of the conventional CNNs, achieving rotation-equivariance remains an elusive goal. Current detectors deploy various alignment techniques to derive rotation-invariant features, but still rely on high capacity models and heavy data augmentation with all possible rotations. In this paper, we introduce a Fully Rotation-Equivariant Oriented Object Detector (FRED), whose entire process from the image to the bounding box prediction is strictly equivariant. Specifically, we decouple the invariant task (object classification) and the equivariant task (object localization) to achieve end-to-end equivariance. We represent the bounding box as a set of rotation-equivariant vectors to implement rotation-equivariant localization. Moreover, we utilized these rotation-equivariant vectors as offsets in the deformable convolution, thereby enhancing the existing advantages of spatial adaptation. Leveraging full rotation-equivariance, our FRED demonstrates higher robustness to image-level rotation compared to existing methods. Furthermore, we show that FRED is one step closer to non-axis aligned learning through our experiments. Compared to state-of-the-art methods, our proposed method delivers comparable performance on DOTA-v1.0 and outperforms by 1.5 mAP on DOTA-v1.5, all while significantly reducing the model parameters to 16%. Chanho Lee, Jinsu Son, Hyounguk Shon, Yunho Jeon, Junmo Kim 0002 |
AAAI | 5 |
| 2024 | ImageNet-D: Benchmarking Neural Network Robustness on Diffusion Synthetic ObjectabstractWe establish rigorous benchmarks for visual perception robustness. Synthetic images such as ImageNet-C, ImageNet-9, and Stylized ImageNet provide specific type of evaluation over synthetic corruptions, backgrounds, and textures, yet those robustness benchmarks are restricted in specified variations and have low synthetic quality. In this work, we introduce generative model as a data source for synthesizing hard images that benchmark deep models' robustness. Leveraging diffusion models, we are able to generate images with more diversified backgrounds, textures, and materials than any prior work, where we term this benchmark as ImageNet-D. Experimental results show that ImageNet-D results in a significant accuracy drop to a range of vision models, from the standard ResNet visual classifier to the latest foundation models like CLIP and MiniGPT-4, significantly reducing their accuracy by up to 60%. Our work suggests that diffusion models can be an effective source to test vision models. The code and dataset are available at https://github.com/chenshuang-zhang/imagenet_d. Chenshuang Zhang, Junmo Kim 0002, In-So Kweon, Chengzhi Mao |
CVPR | 3 |
| 2024 | Learning Neural Deformation Representation for 4D Dynamic Shape Generation
Gyojin Han, Jiwan Hur, Jaehyun Choi, Junmo Kim 0002 |
ECCV (72) | 4 |
| 2024 | Implicit Steganography Beyond the Constraints of Modality
Sojeong Song, Seoyun Yang, Chang Dong Yoo, Junmo Kim 0002 |
ECCV (86) | 4 |
| 2024 | StablePrompt : Automatic Prompt Tuning using Reinforcement Learning for Large Language ModelabstractFinding appropriate prompts for the specific task has become an important issue as the usage of Large Language Models (LLM) has expanded.Reinforcement Learning (RL) is widely used for prompt tuning, but its inherent instability and environmental dependency make it difficult to use in practice.In this paper, we propose StablePrompt, which strikes a balance between training stability and search space, mitigating the instability of RL and producing high-performance prompts.We formulate prompt tuning as an online RL problem between the agent and target LLM and introduce Adaptive Proximal Policy Optimization (APPO).APPO introduces an LLM anchor model to adaptively adjust the rate of policy updates.This allows for flexible prompt search while preserving the linguistic ability of the pre-trained LLM.StablePrompt outperforms previous methods on various tasks including text classification, question answering, and text generation.Our code can be found in github. Minchan Kwon, Gaeun Kim, Jongsuk Kim, Haeil Lee, Junmo Kim 0002 |
EMNLP | 5 |
| 2024 | Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic CompositionalityabstractIn this paper, we propose a new method to enhance compositional understanding in pretrained vision and language models (VLMs) without sacrificing performance in zero-shot multi-modal tasks.Traditional fine-tuning approaches often improve compositional reasoning at the cost of degrading multi-modal capabilities, primarily due to the use of global hard negative (HN) loss, which contrasts global representations of images and texts.This global HN loss pushes HN texts that are highly similar to the original ones, damaging the model's multi-modal representations.To overcome this limitation, we propose Fine-grained Selective Calibrated CLIP (FSC-CLIP), which integrates local hard negative loss and selective calibrated regularization.These innovations provide fine-grained negative supervision while preserving the model's representational integrity.Our extensive evaluations across diverse benchmarks for both compositionality and multi-modal tasks show that FSC-CLIP not only achieves compositionality on par with state-of-the-art models but also retains strong multi-modal capabilities. Youngtaek Oh, Jae-Won Cho, Dong-Jin Kim 0003, In-So Kweon, Junmo Kim 0002 |
EMNLP | 5 |
| 2024 | Stereo-Matching Knowledge Distilled Monocular Depth Estimation Filtered by Multiple Disparity ConsistencyabstractIn stereo-matching knowledge distillation methods of the self-supervised monocular depth estimation, the stereo-matching network’s knowledge is distilled into a monocular depth network through pseudo-depth maps. In these methods, the learning-based stereo-confidence network is generally utilized to identify errors in the pseudo-depth maps to prevent transferring the errors. However, the learning-based stereo-confidence networks should be trained with ground truth (GT), which is not feasible in a self-supervised setting. In this paper, we propose a method to identify and filter errors in the pseudo-depth map using multiple disparity maps by checking their consistency without the need for GT and a training process. Experimental results show that the proposed method outperforms the previous methods and works well on various configurations by filtering out erroneous areas where the stereo-matching is vulnerable, especially such as textureless regions, occlusion boundaries, and reflective surfaces. Woonghyun Ka, Jae Young Lee 0002, Jaehyun Choi, Junmo Kim 0002 |
ICASSP | 4 |
| 2024 | EquiAV: Leveraging Equivariance for Audio-Visual Contrastive LearningabstractRecent advancements in self-supervised audio-visual representation learning have demonstrated its potential to capture rich and comprehensive representations. However, despite the advantages of data augmentation verified in many learning methods, audio-visual learning has struggled to fully harness these benefits, as augmentations can easily disrupt the correspondence between input pairs. To address this limitation, we introduce EquiAV, a novel framework that leverages equivariance for audio-visual contrastive learning. Our approach begins with extending equivariance to audio-visual learning, facilitated by a shared attention-based transformation predictor. It enables the aggregation of features from diverse augmentations into a representative embedding, providing robust supervision. Notably, this is achieved with minimal computational overhead. Extensive ablation studies and qualitative results verify the effectiveness of our method. EquiAV outperforms previous works across various audio-visual benchmarks. The code is available on https://github.com/JongSuk1/EquiAV Jongsuk Kim, Hyeongkeun Lee, Kyeongha Rho, Junmo Kim 0002, Joon Son Chung |
ICML | 4 |
| 2024 | AVCap: Leveraging Audio-Visual Features as Text Tokens for Captioning
Jongsuk Kim, Jiwon Shin, Junmo Kim 0002 |
INTERSPEECH | 3 |
| 2024 | Unlocking the Capabilities of Masked Generative Models for Image Synthesis via Self-GuidanceabstractMasked generative models (MGMs) have shown impressive generative ability while providing an order of magnitude efficient sampling steps compared to continuous diffusion models. However, MGMs still underperform in image synthesis compared to recent well-developed continuous diffusion models with similar size in terms of quality and diversity of generated samples. A key factor in the performance of continuous diffusion models stems from the guidance methods, which enhance the sample quality at the expense of diversity. In this paper, we extend these guidance methods to generalized guidance formulation for MGMs and propose a self-guidance sampling method, which leads to better generation quality. The proposed approach leverages an auxiliary task for semantic smoothing in vector-quantized token space, analogous to the Gaussian blur in continuous pixel space. Equipped with the parameter-efficient fine-tuning method and high-temperature sampling, MGMs with the proposed self-guidance achieve a superior quality-diversity trade-off, outperforming existing sampling methods in MGMs with more efficient training and sampling costs. Extensive experiments with the various sampling hyperparameters confirm the effectiveness of the proposed self-guidance. Jiwan Hur, Gyojin Han, Jaehyun Choi, Yunho Jeon, Junmo Kim 0002 |
NeurIPS | 6 |
| 2024 | Self-supervised Transformation Learning for Equivariant RepresentationsabstractUnsupervised representation learning has significantly advanced various machine learning tasks. In the computer vision domain, state-of-the-art approaches utilize transformations like random crop and color jitter to achieve invariant representations, embedding semantically the same inputs despite transformations. However, this can degrade performance in tasks requiring precise features, such as localization or flower classification. To address this, recent research incorporates equivariant representation learning, which captures transformation-sensitive information. However, current methods depend on transformation labels and thus struggle with interdependency and complex transformations. We propose Self-supervised Transformation Learning (STL), replacing transformation labels with transformation representations derived from image pairs. The proposed method ensures transformation representation is image-invariant and learns corresponding equivariant transformations, enhancing performance without increased batch complexity. We demonstrate the approach’s effectiveness across diverse classification and detection tasks, outperforming existing methods in 7 out of 11 benchmarks and excelling in detection. By integrating complex transformations like AugMix, unusable by prior equivariant methods, this approach enhances performance across tasks, underscoring its adaptability and resilience. Additionally, its compatibility with various base models highlights its flexibility and broad applicability. The code is available at https://github.com/jaemyung-u/stl. Jaemyung Yu, Jaehyun Choi, Hyeong Gwon Hong, Junmo Kim 0002 |
NeurIPS | 5 |
| 2024 | Expanding Expressiveness of Diffusion Models with Limited Data via Self-Distillation based Fine-TuningabstractTraining diffusion models on limited datasets poses challenges in terms of limited generation capacity and expressiveness, leading to unsatisfactory results in various down-stream tasks utilizing pretrained diffusion models, such as domain translation and text-guided image manipulation. In this paper, we propose Self-Distillation for Fine-Tuning diffusion models (SDFT), a methodology to address these challenges by leveraging diverse features from diffusion models pretrained on large source datasets. SDFT distills more general features (shape, colors, etc.) and less domain-specific features (texture, fine details, etc) from the source model, allowing successful knowledge transfer without disturbing the training process on target datasets. The proposed method is not constrained by the specific architecture of the model and thus can be generally adopted to existing frameworks. Experimental results demonstrate that SDFT enhances the expressiveness of the diffusion model with limited datasets, resulting in improved generation capabilities across various downstream tasks. Jiwan Hur, Jaehyun Choi, Gyojin Han, Junmo Kim 0002 |
WACV | 5 |
| 2024 | Real-Time Polyp Detection in Colonoscopy using Lightweight TransformerabstractColorectal cancer (CRC) represents a major global health challenge, and early detection of polyps is crucial in preventing its progression. Although colonoscopy is the gold standard for polyp detection, it has limitations, such as human error and missed detection rates. In response, computer-aided detection (CADe) systems have been developed to enhance the efficiency and accuracy of polyp detection. As deep learning gained prominence, the incorporation of Convolutional Neural Networks (CNNs) into CADe systems emerged as a breakthrough approach. However, CADe systems based on CNNs often demand significant computational resources, making them unsuitable for deployment in resource-constrained environments. To mitigate this, we propose a novel and lightweight polyp detection model that integrates a Transformer layer into the You Only Look Once (YOLO) architecture, focusing on optimizing the neck part responsible for feature fusion and rescaling. Our model demonstrates a substantial reduction in computational complexity and the number of parameters, without compromising detection performances. The lightweight model makes it accessible and feasibly deployable in medically underserved regions, serving a significant public interest by potentially expanding the reach of critical diagnostic tools for CRC prevention. By optimizing the architecture to reduce resource requirements while maintaining performance, our model becomes a practical solution to assist healthcare professionals in the real-time identification of polyps, even with resource-constraint devices. Youngbeom Yoo, Jae Young Lee 0002, Jiwoon Jeon, Junmo Kim 0002 |
WACV | 5 |
| 2024 | ContextMix: A context-aware data augmentation method for industrial visual inspection systems
Hyungmin Kim 0004, Pyunghwan Ahn, Sungho Suh, Hansang Cho, Junmo Kim 0002 |
Eng. Appl. Artif. Intell. | 6 |
| 2024 | AI-KD: Adversarial learning and Implicit regularization for self-Knowledge Distillation
Hyungmin Kim 0004, Sungho Suh, Sunghyun Baek, Daun Jeong, Hansang Cho, Junmo Kim 0002 |
Knowl. Based Syst. | 7 |
| 2023 | Frequency Selective Augmentation for Video Representation LearningabstractRecent self-supervised video representation learning methods focus on maximizing the similarity between multiple augmented views from the same video and largely rely on the quality of generated views. However, most existing methods lack a mechanism to prevent representation learning from bias towards static information in the video. In this paper, we propose frequency augmentation (FreqAug), a spatio-temporal data augmentation method in the frequency domain for video representation learning. FreqAug stochastically removes specific frequency components from the video so that learned representation captures essential features more from the remaining information for various downstream tasks. Specifically, FreqAug pushes the model to focus more on dynamic features rather than static features in the video via dropping spatial or temporal low-frequency components. To verify the generality of the proposed method, we experiment with FreqAug on multiple self-supervised learning frameworks along with standard augmentations. Transferring the improved representation to five video action recognition and two temporal action localization downstream tasks shows consistent improvements over baselines. Jinhyung Kim, Taeoh Kim, Minho Shim, Dongyoon Han, Dongyoon Wee, Junmo Kim 0002 |
AAAI | 6 |
| 2023 | Implicit 3D Human Mesh Recovery using Consistency with Pose and Shape from Unseen-viewabstractFrom an image of a person, we can easily infer the natural 3D pose and shape of the person even if ambiguity exists. This is because we have a mental model that allows us to imagine a person's appearance at different viewing directions from a given image and utilize the consistency between them for inference. However, existing human mesh recovery methods only consider the direction in which the image was taken due to their structural limitations. Hence, we propose “Implicit 3D Human Mesh Recovery (ImpHMR)” that can implicitly imagine a person in 3D space at the feature-level via Neural Feature Fields. In ImpHMR, feature fields are generated by CNN-based image encoder for a given image. Then, the 2D feature map is volume-rendered from the feature field for a given viewing direction, and the pose and shape parameters are regressed from the feature. To utilize consistency with pose and shape from unseen-view, if there are 3D labels, the model predicts results including the silhouette from an arbitrary direction and makes it equal to the rotated ground-truth. In the case of only 2D labels, we perform self-supervised learning through the constraint that the pose and shape parameters inferred from different directions should be the same. Extensive evaluations show the efficacy of the proposed method. Hanbyel Cho, Yooshin Cho, Jaesung Ahn, Junmo Kim 0002 |
CVPR | 4 |
| 2023 | Reinforcement Learning-Based Black-Box Model Inversion AttacksabstractModel inversion attacks are a type of privacy attack that reconstructs private data used to train a machine learning model, solely by accessing the model. Recently, white-box model inversion attacks leveraging Generative Adversarial Networks (GANs) to distill knowledge from public datasets have been receiving great attention because of their excellent attack performance. On the other hand, current black-box model inversion attacks that utilize GANs suffer from issues such as being unable to guarantee the completion of the attack process within a predetermined number of query accesses or achieve the same level of performance as white-box attacks. To overcome these limitations, we propose a reinforcement learning-based black-box model inversion attack. We formulate the latent space search as a Markov Decision Process (MDP) problem and solve it with reinforcement learning. Our method utilizes the confidence scores of the generated images to provide rewards to an agent. Finally, the private data can be reconstructed using the latent vectors found by the agent trained in the MDP. The experiment results on various datasets and models demonstrate that our attack successfully recovers the private information of the target model by achieving state-of-the-art attack performance. We emphasize the importance of studies on privacy-preserving machine learning by proposing a more advanced black-box model inversion attack. Gyojin Han, Jaehyun Choi, Haeil Lee, Junmo Kim 0002 |
CVPR | 4 |
| 2023 | Fix the Noise: Disentangling Source Feature for Controllable Domain TranslationabstractRecent studies show strong generative performance in domain translation especially by using transfer learning techniques on the unconditional generator. However, the control between different domain features using a single model is still challenging. Existing methods often require additional models, which is computationally demanding and leads to unsatisfactory visual quality. In addition, they have restricted control steps, which prevents a smooth transition. In this paper, we propose a new approach for high-quality domain translation with better controllability. The key idea is to preserve source features within a disentangled subspace of a target feature space. This allows our method to smoothly control the degree to which it preserves source features while generating images from an entirely new domain using only a single model. Our extensive experiments show that the proposed method can produce more consistent and realistic images than previous works and maintain precise controllability over different levels of transformation. The code is available at LeeDongYeun/FixNoise. Dongyeun Lee, Jae Young Lee 0002, Jaehyun Choi, Jaejun Yoo 0001, Junmo Kim 0002 |
CVPR | 6 |
| 2023 | Video Inference for Human Mesh Recovery with Vision TransformerabstractHuman Mesh Recovery (HMR) from an image is a challenging problem because of the inherent ambiguity of the task. Existing HMR methods utilized either temporal information or kinematic relationships to achieve higher accuracy, but there is no method using both. Hence, we propose “Video Inference for Human Mesh Recovery with Vision Transformer (HMR-ViT)” that can take into account both temporal and kinematic information. In HMR-ViT, a Temporal-kinematic Feature Image is constructed using feature vectors obtained from video frames by an image encoder. When generating the feature image, we use a Channel Rearranging Matrix (CRM) so that similar kinematic features could be located spatially close together. The feature image is then further encoded using Vision Transformer, and the SMPL pose and shape parameters are finally inferred using a regression network. Extensive evaluation on the 3DPW and Human3.6M datasets indicates that our method achieves a competitive performance in HMR. Hanbyel Cho, Jaesung Ahn, Yooshin Cho, Junmo Kim 0002 |
FG | 4 |
| 2023 | Localization using Multi-Focal Spatial Attention for Masked Face RecognitionabstractSince the beginning of world-wide COVID-19 pandemic, facial masks have been recommended to limit the spread of the disease. However, these masks hide certain facial attributes. Hence, it has become difficult for existing face recognition systems to perform identity verification on masked faces. In this context, it is necessary to develop masked Face Recognition (MFR) for contactless biometric recognition systems. Thus, in this paper, we propose Complementary Attention Learning and Multi-Focal Spatial Attention that precisely removes masked region by training complementary spatial attention to focus on two distinct regions: masked regions and backgrounds. In our method, standard spatial attention and networks focus on unmasked regions, and extract mask-invariant features while minimizing the loss of the conventional Face Recognition (FR) performance. For conventional FR, we evaluate the performance on the IJB-C, Age-DB, CALFW, and CPLFW datasets. We evaluate the MFR performance on the ICCV2021-MFR/Insightface track, and demonstrate the improved performance on the both MFR and FR datasets. Additionally, we empirically verify that spatial attention of proposed method is more precisely activated in unmasked regions. Yooshin Cho, Hanbyel Cho, Hyeong Gwon Hong, Jaesung Ahn, Dongmin Cho, JungWoo Chang, Junmo Kim 0002 |
FG | 7 |
| 2023 | Proxy Anchor-based Unsupervised Learning for Continuous Generalized Category DiscoveryabstractRecent advances in deep learning have significantly improved the performance of various computer vision applications. However, discovering novel categories in an incremental learning scenario remains a challenging problem due to the lack of prior knowledge about the number and nature of new categories. Existing methods for novel category discovery are limited by their reliance on labeled datasets and prior knowledge about the number of novel categories and the proportion of novel samples in the batch. To address the limitations and more accurately reflect real-world scenarios, in this paper, we propose a novel unsupervised class incremental learning approach for discovering novel categories on unlabeled sets without prior knowledge. The proposed method fine-tunes the feature extractor and proxy anchors on labeled sets, then splits samples into old and novel categories and clusters on the unlabeled dataset. Furthermore, the proxy anchors-based exemplar generates representative category vectors to mitigate catastrophic forgetting. Experimental results demonstrate that our proposed approach outperforms the state-of-the-art methods on fine-grained datasets under real-world scenarios. Hyungmin Kim 0004, Sungho Suh, Daun Jeong, Hansang Cho, Junmo Kim 0002 |
ICCV | 6 |
| 2023 | Disposable Transfer Learning for Selective Source Task UnlearningabstractTransfer learning is widely used for training deep neural networks (DNN) for building a powerful representation. Even after the pre-trained model is adapted for the target task, the representation performance of the feature extractor is retained to some extent. As the performance of the pre-trained model can be considered the private property of the owner, it is natural to seek the exclusive right of the generalized performance of the pre-trained weight. To address this issue, we suggest a new paradigm of transfer learning called disposable transfer learning (DTL), which disposes of only the source task without degrading the performance of the target task. To achieve knowledge disposal, we propose a novel loss named Gradient Collision loss (GC loss). GC loss selectively unlearns the source knowledge by leading the gradient vectors of mini-batches in different directions. Whether the model successfully unlearns the source task is measured by piggyback learning accuracy (PL accuracy). PL accuracy estimates the vulnerability of knowledge leakage by retraining the scrubbed model on a subset of source data or new downstream data. We demonstrate that GC loss is an effective approach to the DTL problem by showing that the model trained with GC loss retains the performance on the target task with a significantly reduced PL accuracy. Seunghee Koh, Hyounguk Shon, Janghyeon Lee 0001, Hyeong Gwon Hong, Junmo Kim 0002 |
ICCV | 5 |
| 2023 | Data Poisoning Attack Aiming the Vulnerability of Continual LearningabstractGenerally, regularization-based continual learning models limit access to the previous task data to imitate the real-world constraints related to memory and privacy. However, this introduces a problem in these models by not being able to track the performance on each task. In essence, current continual learning methods are susceptible to attacks on previous tasks. We demonstrate the vulnerability of regularization-based continual learning methods by presenting a simple task-specific data poisoning attack that can be used in the learning process of a new task. Training data generated by the proposed attack causes performance degradation on a specific task targeted by the attacker. We experiment with the attack on the two representative regularization-based continual learning methods, Elastic Weight Consolidation (EWC) and Synaptic Intelligence (SI), trained with variants of MNIST dataset. The experiment results justify the vulnerability proposed in this paper and demonstrate the importance of developing continual learning models that are robust to adversarial attacks. Gyojin Han, Jaehyun Choi, Hyeong Gwon Hong, Junmo Kim 0002 |
ICIP | 4 |
| 2023 | Deep Cross-Modal Steganography Using Neural RepresentationsabstractSteganography is the process of embedding secret data into another message or data, in such a way that it is not easily noticeable. With the advancement of deep learning, Deep Neural Networks (DNNs) have recently been utilized in steganography. However, existing deep steganography techniques are limited in scope, as they focus on specific data types and are not effective for cross-modal steganography. Therefore, We propose a deep cross-modal steganography framework using Implicit Neural Representations (INRs) to hide secret data of various formats in cover images. The proposed framework employs INRs to represent the secret data, which can handle data of various modalities and resolutions. Experiments on various secret datasets of diverse types demonstrate that the proposed approach is expandable and capable of accommodating different modalities. Gyojin Han, Jiwan Hur, Jaehyun Choi, Junmo Kim 0002 |
ICIP | 5 |
| 2023 | Training Cartoonization Network without CartoonabstractPhoto cartoonization aims to translate a real-world photo into a cartoon image. Previous learning-based cartoonization studies have shown promising results, however, an alternative approach is still necessary because of the following limitations. First, the deficiency of training datasets necessitates the additional dataset acquisition process, which yields uneven network training results. Second, we observe that the created images using existing works have color transition problems. In this paper, we propose a new approach on photo cartoonization which does not use cartoon datasets and does not suffer these limitations. By focusing on the two important aspects of cartoon style, regions and edges, our work enables us to produce cartoonized images without using a manually collected dataset. We show that our framework can generate high-quality cartoonized images and achieve competitive performance with comparable cartoonization networks even without training cartoon dataset. Dongyeun Lee, Donggyu Joo, Junmo Kim 0002 |
ICIP | 4 |
| 2023 | Lightweight Monocular Depth Estimation via Token-Sharing TransformerabstractDepth estimation is an important task in various robotics systems and applications. In mobile robotics systems, monocular depth estimation is desirable since a single RGB camera can be deployable at a low cost and compact size. Due to its significant and growing needs, many lightweight monocular depth estimation networks have been proposed for mobile robotics systems. While most lightweight monocular depth estimation methods have been developed using convolution neural networks, the Transformer has been gradually utilized in monocular depth estimation recently. However, massive parameters and large computational costs in the Transformer disturb the deployment to embedded devices. In this paper, we present a Token-Sharing Transformer (TST), an architecture using the Transformer for monocular depth estimation, optimized especially in embedded devices. The proposed TST utilizes global token sharing, which enables the model to obtain an accurate depth prediction with high throughput in embedded devices. Experimental results show that TST outperforms the existing lightweight monocular depth estimation methods. On the NYU Depth v2 dataset, TST can deliver depth maps up to 63.4 FPS in NVIDIA Jetson nano and 142.6 FPS in NVIDIA Jetson TX2, with lower errors than the existing methods. Furthermore, TST achieves real-time depth estimation of high-resolution images on Jetson TX2 with competitive results. Jae Young Lee 0002, Hyunguk Shon, Eojindl Yi, Yeong-Hun Park, Sung-Sik Cho, Junmo Kim 0002 |
ICRA | 7 |
| 2023 | Test-Time Synthetic-to-Real Adaptive Depth EstimationabstractIs it possible for a synthetic to realistic domain adapted neural network in single image depth estimation to truly generalize on real world data? The resultant, adapted model will only generalize on the realistic domain dataset, which only reflects a small portion of the true, real world. As a result, the network still has to cope with the potential danger of domain shift between the realistic domain dataset and the real world data. Instead, a viable solution is to design the model to be capable of continuously adapting to the distribution of data it receives at test-time. In this paper, we propose a depth estimation method that is capable of adapting to the domain shift at test-time. Our method adapts to the unseen test-time domain, by updating the network using our proposed objective functions. Following former work, we reduce the entropy of the current prediction for refinement and adaptation. We propose a Logit Order Enforcement loss that can prevent the network from deviating into wrong solutions, which can result from the mere reduction of the aforementioned entropy. Qualitative and quantitative results show the effectiveness of our method. Our method reduces the dependency on training data by 5.8× on average, while achieving comparable performance to state-of-the-art unsupervised domain adaptation (UDA) and domain generalization methods (DG) on the KITTI dataset. Eojindl Yi, Junmo Kim 0002 |
ICRA | 2 |
| 2023 | I See-Through You: A Framework for Removing Foreground Occlusion in Both Sparse and Dense Light Field ImagesabstractLight field (LF) camera captures rich information from a scene. Using the information, the LF de-occlusion (LF-DeOcc) task aims to reconstruct the occlusion-free center view image. Existing LF-DeOcc studies mainly focus on the sparsely sampled (sparse) LF images where most of the occluded regions are visible in other views due to the large disparity. In this paper, we expand LF-DeOcc in more challenging datasets, densely sampled (dense) LF images, which are taken by a micro-lens-based portable LF camera. Due to the small disparity ranges of dense LF images, most of the background regions are invisible in any view. To apply LF-DeOcc in both LF datasets, we propose a framework, ISTY, which is defined and divided into three roles: (1) extract LF features, (2) define the occlusion, and (3) inpaint occluded regions. By dividing the framework into three specialized components according to the roles, the development and analysis can be easier. Furthermore, an explainable intermediate representation, an occlusion mask, can be obtained in the proposed framework. The occlusion mask is useful for comprehensive analysis of the model and other applications by manipulating the mask. In experiments, qualitative and quantitative results show that the proposed framework outperforms state-of-the-art LF-DeOcc methods in both sparse and dense LF datasets. Jiwan Hur, Jae Young Lee 0002, Jaehyun Choi, Junmo Kim 0002 |
WACV | 4 |
| 2023 | Multi-scale foreground-background separation for light field depth estimation with deep convolutional networks
Jae Young Lee 0002, Jiwan Hur, Jaehyun Choi, Rae-Hong Park, Junmo Kim 0002 |
Pattern Recognit. Lett. | 5 |
| 2022 | Closing the Loophole: Rethinking Reconstruction Attacks in Federated Learning from a Privacy StandpointabstractFederated Learning was deemed as a private distributed learning framework due to the separation of data from the central server. However, recent works have shown that privacy attacks can extract various forms of private information from legacy federated learning. Previous literature describe differential privacy to be effective against membership inference attacks and attribute inference attacks, but our experiments show them to be vulnerable against reconstruction attacks. To understand this outcome, we execute a systematic study of privacy attacks from the standpoint of privacy. The privacy characteristics that reconstruction attacks infringe are different from other privacy attacks, and we suggest that privacy breach occurred at different levels. From our study, reconstruction attack defense methods entail heavy computation or communication costs. To this end, we propose Fragmented Federated Learning (FFL), a lightweight solution against reconstruction attacks. This framework utilizes a simple yet novel gradient obscuring algorithm based on a newly proposed concept called the global gradient and determines which layers are safe for submission to the server. We show empirically in diverse settings that our framework improves practical data privacy of clients in federated learning with an acceptable performance trade-off without increasing communication cost. We aim to provide a new perspective to privacy in federated learning and hope this privacy differentiation can improve future privacy-preserving methods. Seung Ho Na, Hyeong Gwon Hong, Junmo Kim 0002, Seungwon Shin 0001 |
ACSAC | 3 |
| 2022 | Beyond Semantic to Instance Segmentation: Weakly-Supervised Instance Segmentation via Semantic Knowledge Transfer and Self-RefinementabstractWeakly-supervised instance segmentation (WSIS) has been considered as a more challenging task than weakly-supervised semantic segmentation (WSSS). Compared to WSSS, WSIS requires instance-wise localization, which is difficult to extract from image-level labels. To tackle the problem, most WSIS approaches use off-the-shelf proposal techniques that require pre-training with instance or object level labels, deviating the fundamental definition of the fully-image-level supervised setting. In this paper, we propose a novel approach including two innovative components. First, we propose a semantic knowledge transfer to obtain pseudo instance labels by transferring the knowledge of WSSS to WSIS while eliminating the need for the off-the-shelf proposals. Second, we propose a self-refinement method to refine the pseudo instance labels in a self-supervised scheme and to use the refined labels for training in an online manner. Here, we discover an erroneous phenomenon, semantic drift, that occurred by the missing instances in pseudo instance labels categorized as background class. This semantic drift occurs confusion between background and instance in training and consequently degrades the segmentation performance. We term this problem as semantic drift problem and show that our proposed self-refinement method eliminates the semantic drift problem. The extensive experiments on PASCAL VOC 2012 and MS COCO demonstrate the effectiveness of our approach, and we achieve a considerable performance without off-the-shelf proposal techniques. The code is available at https://github.com/clovaai/BESTIE. Beomyoung Kim, Young Joon Yoo, Chaeeun Rhee, Junmo Kim 0002 |
CVPR | 4 |
| 2022 | DLCFT: Deep Linear Continual Fine-Tuning for General Incremental Learning
Hyounguk Shon, Janghyeon Lee 0001, Junmo Kim 0002 |
ECCV (33) | 4 |
| 2022 | On the Angular Update and Hyperparameter Tuning of a Scale-Invariant Network
Juseung Yun, Janghyeon Lee 0001, Hyounguk Shon, Eojindl Yi, Junmo Kim 0002 |
ECCV (12) | 6 |
| 2022 | Rethinking Efficacy of Softmax for Lightweight Non-local Neural NetworksabstractNon-local (NL) block is a popular module that demonstrates the capability to model global contexts. However, NL block generally has heavy computation and memory costs, so it is impractical to apply the block to high-resolution feature maps. In this paper, to investigate the efficacy of NL block, we empirically analyze if the magnitude and direction of input feature vectors properly affect the attention between vectors. The results show the inefficacy of softmax operation that is generally used to normalize the attention map of the NL block. Attention maps normalized with softmax operation highly rely upon magnitude of key vectors, and performance is degenerated if the magnitude information is removed. By replacing softmax operation with the scaling factor, we demonstrate improved performance on CIFAR-10, CIFAR-100, and Tiny-ImageNet. In Addition, our method shows robustness to embedding channel reduction and embedding weight initialization. Notably, our method makes multi-head attention employable without additional computational cost. Yooshin Cho, Youngsoo Kim 0006, Hanbyel Cho, Jaesung Ahn, Hyeong Gwon Hong, Junmo Kim 0002 |
ICIP | 6 |
| 2022 | Which Metrics for Network Pruning: Final Accuracy? Or Accuracy Drop?abstractNetwork pruning enables the utilization of deep neural networks in low-resource environments by removing redundant elements in a pre-trained network. To appraise each pruning method, two evaluation metrics are generally adopted, i.e., final accuracy and accuracy drop. Final accuracy represents the ultimate performance of the pruned sub-network after the pruning completes. On the other hand, accuracy drop, a more traditional way, measures the accuracy difference between the baseline model and the final pruned model. In this work, we present several surprising observations which reveal the unfairness of both metrics when assessing the efficacy of pruning approaches. Depending on the choice of baseline network, the value of each pruning method may be completely changed. Specifically, a lower baseline tends to be advantageous for the accuracy drop, whereas a higher baseline usually yields a higher final accuracy. Moreover, to reduce the undesirable dependency on the baseline network, we propose a new reliable averaging method Average from Scratches which uses multiple distinct baselines rather than using a single baseline. Our various investigations point to the necessity for a more thorough analysis on network pruning metrics. Donggyu Joo, Sunghyun Baek, Junmo Kim 0002 |
ICIP | 3 |
| 2022 | Enhanced Prototypical Learning for Unsupervised Domain Adaptation in LiDAR Semantic SegmentationabstractDespite its importance, unsupervised domain adaptation (UDA) on LiDAR semantic segmentation is a task that has not received much attention from the research community. Only recently, a completion-based 3$D$method has been proposed to tackle the problem and formally set up the adaptive scenarios. However, the proposed pipeline is complex, voxel-based and requires multi-stage inference, which inhibits it for real-time inference. We propose a range image-based, effective and efficient method for solving UDA on LiDAR segmentation. The method exploits class prototypes from the source domain to pseudo label target domain pixels, which is a research direction showing good performance in UDA for natural image semantic segmentation. Applying such approaches to LiDAR scans has not been considered because of the severe domain shift and lack of pre-trained feature extractor that is unavailable in the LiDAR segmentation setup. However, we show that proper strategies, including reconstruction-based pre-training, enhanced prototypes, and selective pseudo labeling based on distance to prototypes, is sufficient enough to enable the use of prototypical approaches. We evaluate the performance of our method on the recently proposed LiDAR segmentation UDA scenarios. Our method achieves remarkable performance among contemporary methods. Eojindl Yi, JuYoung Yang, Junmo Kim 0002 |
ICRA | 3 |
| 2022 | Multi-Scaled and Densely Connected Locally Convolutional Layers for Depth CompletionabstractThe depth completion task aims to predict a dense depth map from a sparse LiDAR point cloud and an RGB image. This task is critical because an accurate depth map can be used as prior information to solve many computer vision tasks, such as downstream tasks in autonomous vehicles and robot vision. Previous deep learning methods which focus on the local affinity have achieved impressive results. However, an architecture that is directly designed to extract local affinity has not been proposed yet. In this paper, we propose multi-scaled and densely connected locally convolutional layers to learn the affinity of the neighborhood. We set a different grid factor for each step of this module, and each step consists of several convolutional layers applied only to the local area assigned from the grid factor. In addition, each step is densely connected, sequentially, to take advantage of the multi-scale receptive fields. The proposed module effectively learns the neighbor-hood's affinity in a local area with multiple scales, while keeping the network size small. As a result, our architecture achieves state-of-the-art performance compared to published works on the KITTI depth completion benchmark. On the NYU Depth V2 completion benchmark our method achieves performance comparable to state-of-the-art approaches. Sihaeng Lee, Eojindl Yi, Janghyeon Lee 0001, Junmo Kim 0002 |
IROS | 4 |
| 2022 | Subspace-based Feature Alignment for Unsupervised Domain AdaptationabstractAutonomous agents need to perceive the world in a robust way, such that the shift in data distribution does not lead to faulty perception results. When agents cannot be trained with abundant data, agents may need to operate on real world environments while trained on simulated data, and suffer from domain shift. This paper proposes an effective and robust unsupervised domain adaptation (UDA) method that can resolve these situations. In the UDA setup, we are given a labeled source domain and an unlabeled target domain that share the same set of classes but are sampled from different distributions. This domain shift prevents agents which employ deep neural networks from generalizing well on the target domain. Recent methods adopt the strategy of self-training the networks with pseudo labeled target samples. However, falsely labeled samples cause negative transfer and deteriorate generalization of a network. to reduce negative transfer we propose an algorithm that can filter the pseudo labels, and use the filtered labels to align the domains in the feature space. The samples whose labels have not passed the filtering process can be used as an index to tune the hyperparameters of our method. Across various benchmarks, we validate the performance of our method. Especially, our method achieves strong performance on the synthetic-to-real adaptation scenario. Eojindl Yi, Junmo Kim 0002 |
IROS | 2 |
| 2022 | UniCLIP: Unified Framework for Contrastive Language-Image Pre-trainingabstractPre-training vision-language models with contrastive objectives has shown promising results that are both scalable to large uncurated datasets and transferable to many downstream applications. Some following works have targeted to improve data efficiency by adding self-supervision terms, but inter-domain (image-text) contrastive loss and intra-domain (image-image) contrastive loss are defined on individual spaces in those works, so many feasible combinations of supervision are overlooked. To overcome this issue, we propose UniCLIP, a Unified framework for Contrastive Language-Image Pre-training. UniCLIP integrates the contrastive loss of both inter-domain pairs and intra-domain pairs into a single universal space. The discrepancies that occur when integrating contrastive loss between different domains are resolved by the three key components of UniCLIP: (1) augmentation-aware feature embedding, (2) MP-NCE loss, and (3) domain dependent similarity measure. UniCLIP outperforms previous vision-language pre-training methods on various single- and multi-modality downstream tasks. In our experiments, we show that each component that comprises UniCLIP contributes well to the final performance. Janghyeon Lee 0001, Jongsuk Kim, Hyounguk Shon, Bumsoo Kim 0005, Honglak Lee, Junmo Kim 0002 |
NeurIPS | 7 |
| 2022 | Cyclic test time augmentation with entropy weight methodabstractIn the recent studies of data augmentation of neural networks, the application of test time augmentation has been studied to extract optimal transformation policies to enhance performance with minimum cost. The policy search method with the best level of input data dependency involves training a loss predictor network to estimate suitable transformations for each of the given input image in independent manner, resulting in instance-level transformation extraction. In this work, we propose a method to utilize and modify the loss prediction pipeline to further improve the performance with the cyclic search for suitable transformations and the use of the entropy weight method. The cyclic usage of the loss predictor allows refining each input image with multiple transformations with a more flexible transformation magnitude. For cases where multiple augmentations are generated, we implement the entropy weight method to reflect the data uncertainty of each augmentation to force the final result to focus on augmentations with low uncertainty. The experimental results show convincing qualitative outcomes and robust performance for the corrupted conditions of data. Sewhan Chun, Jae Young Lee 0002, Junmo Kim 0002 |
UAI | 3 |
| 2022 | Self-Supervised Knowledge Transfer via Loosely Supervised Auxiliary TasksabstractKnowledge transfer using convolutional neural networks (CNNs) can help efficiently train a CNN with fewer parameters or maximize the generalization performance under limited supervision. To enable a more efficient transfer of pretrained knowledge under relaxed conditions, we propose a simple yet powerful knowledge transfer methodology without any restrictions regarding the network structure or dataset used, namely self-supervised knowledge transfer (SSKT), via loosely supervised auxiliary tasks. For this, we devise a training methodology that transfers previously learned knowledge to the current training process as an auxiliary task for the target task through self-supervision using a soft label. The SSKT is independent of the network structure and dataset, and is trained differently from existing knowledge transfer methods; hence, it has an advantage in that the prior knowledge acquired from various tasks can be naturally transferred during the training process to the target task. Furthermore, it can improve the generalization performance on most datasets through the proposed knowledge transfer between different problem domains from multiple source networks. SSKT outperforms the other transfer learning methods (KD, DML, and MAXL) through experiments under various knowledge transfer settings. The source code will be made available to the public1. Seungbum Hong, Jihun Yoon, Min-Kook Choi, Junmo Kim 0002 |
WACV | 4 |
| 2022 | TricubeNet: 2D Kernel-Based Object Representation for Weakly-Occluded Oriented Object DetectionabstractWe present a novel approach for oriented object detection, named TricubeNet, which localizes oriented objects using visual cues (i.e., heatmap) instead of oriented box offsets regression. We represent each object as a 2D Tricube kernel and extract bounding boxes using simple image-processing algorithms. Our approach is able to (1) obtain well-arranged boxes from visual cues, (2) solve the angle discontinuity problem, and (3) can save computational complexity due to our anchor-free modeling. To further boost the performance, we propose some effective techniques for size-invariant loss, reducing false detections, extracting rotation-invariant features, and heatmap refinement. To demonstrate the effectiveness of our TricubeNet, we experiment on various tasks for weakly-occluded oriented object detection: detection in an aerial image, densely packed object image, and text image. The extensive experimental results show that our TricubeNet is quite effective for oriented object detection. Code is available at https://github.com/qjadud1994/TricubeNet. Beomyoung Kim, Janghyeon Lee 0001, Sihaeng Lee, Junmo Kim 0002 |
WACV | 5 |
| 2021 | Linearly Replaceable Filters for Deep Network Channel PruningabstractConvolutional neural networks (CNNs) have achieved remarkable results; however, despite the development of deep learning, practical user applications are fairly limited because heavy networks can be used solely with the latest hardware and software supports. Therefore, network pruning is gaining attention for general applications in various fields. This paper proposes a novel channel pruning method, Linearly Replaceable Filter (LRF), which suggests that a filter that can be approximated by the linear combination of other filters is replaceable. Moreover, an additional method called Weights Compensation is proposed to support the LRF method. This is a technique that effectively reduces the output difference caused by removing filters via direct weight modification. Through various experiments, we have confirmed that our method achieves state-of-the-art performance in several benchmarks. In particular, on ImageNet, LRF-60 reduces approximately 56% of FLOPs on ResNet-50 without top-5 accuracy drop. Further, through extensive analyses, we proved the effectiveness of our approaches. Donggyu Joo, Eojindl Yi, Sunghyun Baek, Junmo Kim 0002 |
AAAI | 4 |
| 2021 | Discriminative Region Suppression for Weakly-Supervised Semantic SegmentationabstractWeakly-supervised semantic segmentation (WSSS) using image-level labels has recently attracted much attention for reducing annotation costs. Existing WSSS methods utilize localization maps from the classification network to generate pseudo segmentation labels. However, since localization maps obtained from the classifier focus only on sparse discriminative object regions, it is difficult to generate high-quality segmentation labels. To address this issue, we introduce discriminative region suppression (DRS) module that is a simple yet effective method to expand object activation regions. DRS suppresses the attention on discriminative regions and spreads it to adjacent non-discriminative regions, generating dense localization maps. DRS requires few or no additional parameters and can be plugged into any network. Furthermore, we introduce an additional learning strategy to give a self-enhancement of localization maps, named localization map refinement learning. Benefiting from this refinement learning, localization maps are refined and enhanced by recovering some missing parts or removing noise itself. Due to its simplicity and effectiveness, our approach achieves mIoU 71.4% on the PASCAL VOC 2012 segmentation benchmark using only image-level labels. Extensive experiments demonstrate the effectiveness of our approach. Beomyoung Kim, Sangeun Han, Junmo Kim 0002 |
AAAI | 3 |
| 2021 | Patch-Wise Attention Network for Monocular Depth EstimationabstractIn computer vision, monocular depth estimation is the problem of obtaining a high-quality depth map from a two-dimensional image. This map provides information on three-dimensional scene geometry, which is necessary for various applications in academia and industry, such as robotics and autonomous driving. Recent studies based on convolutional neural networks achieved impressive results for this task. However, most previous studies did not consider the relationships between the neighboring pixels in a local area of the scene. To overcome the drawbacks of existing methods, we propose a patch-wise attention method for focusing on each local area. After extracting patches from an input feature map, our module generates attention maps for each local patch, using two attention modules for each patch along the channel and spatial dimensions. Subsequently, the attention maps return to their initial positions and merge into one attention feature. Our method is straightforward but effective. The experimental results on two challenging datasets, KITTI and NYU Depth V2, demonstrate that the proposed method achieves significant performance. Furthermore, our method outperforms other state-of-the-art methods on the KITTI depth estimation benchmark. Sihaeng Lee, Janghyeon Lee 0001, Byungju Kim, Eojindl Yi, Junmo Kim 0002 |
AAAI | 5 |
| 2021 | Joint Negative and Positive Learning for Noisy LabelsabstractTraining of Convolutional Neural Networks (CNNs) with data with noisy labels is known to be a challenge. Based on the fact that directly providing the label to the data (Positive Learning; PL) has a risk of allowing CNNs to memorize the contaminated labels for the case of noisy data, the indirect learning approach that uses complementary labels (Negative Learning for Noisy Labels; NLNL) has proven to be highly effective in preventing overfitting to noisy data as it reduces the risk of providing faulty target. NLNL further employs a three-stage pipeline to improve convergence. As a result, filtering noisy data through the NLNL pipeline is cumbersome, increasing the training cost. In this study, we propose a novel improvement of NLNL, named Joint Negative and Positive Learning (JNPL), that unifies the filtering pipeline into a single stage. JNPL trains CNN via two losses, NL+ and PL+, which are improved upon NL and PL loss functions, respectively. We analyze the fundamental issue of NL loss function and develop new NL+ loss function producing gradient that enhances the convergence of noisy data. Furthermore, PL+ loss function is designed to enable faster convergence to expected-to-be-clean data. We show that the NL+ and PL+ train CNN simultaneously, significantly simplifying the pipeline, allowing greater ease of practical use compared to NLNL. With a simple semi-supervised training technique, our method achieves state-of-the-art accuracy for noisy data classification based on the superior filtering ability. Youngdong Kim, Juseung Yun, Hyounguk Shon, Junmo Kim 0002 |
CVPR | 4 |
| 2021 | Improving Generalization of Batch Whitening by Convolutional Unit OptimizationabstractBatch Whitening is a technique that accelerates and stabilizes training by transforming input features to have a zero mean (Centering) and a unit variance (Scaling), and by removing linear correlation between channels (Decorrelation). In commonly used structures, which are empirically optimized with Batch Normalization, the normalization layer appears between convolution and activation function. Following Batch Whitening studies have employed the same structure without further analysis; even Batch Whitening was analyzed on the premise that the input of a linear layer is whitened. To bridge the gap, we propose a new Convolutional Unit that in line with the theory, and our method generally improves the performance of Batch Whitening. Moreover, we show the inefficacy of the original Convolutional Unit by investigating rank and correlation of features. As our method is employable off-the-shelf whitening modules, we use Iterative Normalization (IterNorm), the state-of-the-art whitening module, and obtain significantly improved performance on five image classification datasets: CIFAR-10, CIFAR-100, CUB-200-2011, Stanford Dogs, and ImageNet. Notably, we verify that our method improves stability and performance of whitening when using large learning rate, group size, and iteration number. Code is available at https://github.com/YooshinCho/pytorch_ConvUnitOptimization. Yooshin Cho, Hanbyel Cho, Youngsoo Kim 0006, Junmo Kim 0002 |
ICCV | 4 |
| 2021 | Camera Distortion-aware 3D Human Pose Estimation in Video with Optimization-based Meta-LearningabstractExisting 3D human pose estimation algorithms trained on distortion-free datasets suffer performance drop when applied to new scenarios with a specific camera distortion. In this paper, we propose a simple yet effective model for 3D human pose estimation in video that can quickly adapt to any distortion environment by utilizing MAML, a representative optimization-based meta-learning algorithm. We consider a sequence of 2D keypoints in a particular distortion as a single task of MAML. However, due to the absence of a large-scale dataset in a distorted environment, we propose an efficient method to generate synthetic distorted data from undistorted 2D keypoints. For the evaluation, we assume two practical testing situations depending on whether a motion capture sensor is available or not. In particular, we propose Inference Stage Optimization using bone-length symmetry and consistency. Extensive evaluation shows that our proposed method successfully adapts to various degrees of distortion in the testing phase and outperforms the existing state-of-the-art approaches. The proposed method is useful in practice because it does not require camera calibration and additional computations in a testing set-up. Code is available at https://github.com/hanbyel0105/CamDistHumanPose3D. Hanbyel Cho, Yooshin Cho, Jaemyung Yu, Junmo Kim 0002 |
ICCV | 4 |
| 2021 | Progressive Seed Generation Auto-encoder for Unsupervised Point Cloud LearningabstractWith the development of 3D scanning technologies, 3D vision tasks have become a popular research area. Owing to the large amount of data acquired by sensors, unsupervised learning is essential for understanding and utilizing point clouds without an expensive annotation process. In this paper, we propose a novel framework and an effective auto-encoder architecture named "PSG-Net" for reconstruction-based learning of point clouds. Unlike existing studies that used fixed or random 2D points, our framework generates input-dependent point-wise features for the latent point set. PSG-Net uses the encoded input to produce point-wise features through the seed generation module and extracts richer features in multiple stages with gradually increasing resolution by applying the seed feature propagation module progressively. We prove the effectiveness of PSG-Net experimentally; PSG-Net shows state-of-the-art performances in point cloud reconstruction and unsupervised classification, and achieves comparable performance to counterpart methods in supervised completion. JuYoung Yang, Pyunghwan Ahn, Haeil Lee, Junmo Kim 0002 |
ICCV | 5 |
| 2021 | De-biasing Neural Networks with Estimated Offset for Class Imbalanced Learning
Byungju Kim, Hyeong Gwon Hong, Junmo Kim 0002 |
WACV | 3 |
| 2021 | Integrating Multiple Receptive Fields Through Grouped Active ConvolutionabstractConvolutional networks have achieved great success in various vision tasks. This is mainly due to a considerable amount of research on network structure. In this study, instead of focusing on architectures, we focused on the convolution unit itself. The existing convolution unit has a fixed shape and is limited to observing restricted receptive fields. In earlier work, we proposed the active convolution unit (ACU), which can freely define its shape and learn by itself. In this paper, we provide a detailed analysis of the previously proposed unit and show that it is an efficient representation of a sparse weight convolution. Furthermore, we extend an ACU to a grouped ACU, which can observe multiple receptive fields in one layer. We found that the performance of a naive grouped convolution is degraded by increasing the number of groups; however, the proposed unit retains the accuracy even though the number of parameters decreases. Based on this result, we suggest a depthwise ACU (DACU), and various experiments have shown that our unit is efficient and can replace the existing convolutions. Yunho Jeon, Junmo Kim 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | Residual Continual LearningabstractWe propose a novel continual learning method called Residual Continual Learning (ResCL). Our method can prevent the catastrophic forgetting phenomenon in sequential learning of multiple tasks, without any source task information except the original network. ResCL reparameterizes network parameters by linearly combining each layer of the original network and a fine-tuned network; therefore, the size of the network does not increase at all. To apply the proposed method to general convolutional neural networks, the effects of batch normalization layers are also considered. By utilizing residual-learning-like reparameterization and a special weight decay loss, the trade-off between source and target performance is effectively controlled. The proposed method exhibits state-of-the-art performance in various continual learning scenarios. Janghyeon Lee 0001, Donggyu Joo, Hyeong Gwon Hong, Junmo Kim 0002 |
AAAI | 4 |
| 2020 | Regularization on Spatio-Temporally Smoothed Feature for Action RecognitionabstractDeep neural networks for video action recognition frequently require 3D convolutional filters and often encounter overfitting due to a larger number of parameters. In this paper, we propose Random Mean Scaling (RMS), a simple and effective regularization method, to relieve the overfitting problem in 3D residual networks. The key idea of RMS is to randomly vary the magnitude of low-frequency components of the feature to regularize the model. The low-frequency component can be derived by a spatio-temporal mean on the local patch of a feature. We present that selective regularization on this locally smoothed feature makes a model handle the low-frequency and high-frequency component distinctively, resulting in performance improvement. RMS can enhance a model with little additional computation only during training, similar to other regularization methods. RMS also can be incorporated into typical training process without any bells and whistles. Experimental results show the improvement in generalization performance on a popular action recognition datasets demonstrating the effectiveness of RMS as a regularization technique, compared to other state-of-the-art regularization methods. Jinhyung Kim, Seunghwan Cha, Dongyoon Wee, Soonmin Bae, Junmo Kim 0002 |
CVPR | 5 |
| 2020 | Continual Learning With Extended Kronecker-Factored Approximate CurvatureabstractWe propose a quadratic penalty method for continual learning of neural networks that contain batch normalization (BN) layers. The Hessian of a loss function represents the curvature of the quadratic penalty function, and a Kronecker-factored approximate curvature (K-FAC) is used widely to practically compute the Hessian of a neural network. However, the approximation is not valid if there is dependence between examples, typically caused by BN layers in deep network architectures. We extend the K-FAC method so that the inter-example relations are taken into account and the Hessian of deep neural networks can be properly approximated under practical assumptions. We also propose a method of weight merging and reparameterization to properly handle statistical parameters of BN, which plays a critical role for continual learning with BN, and a method that selects hyperparameters without source task data. Our method shows better performance than baselines in the permuted MNIST task with BN layers and in sequential learning from the ImageNet classification task to fine-grained classification tasks with ResNet-50, without any explicit or implicit use of source task data for hyperparameter selection. Janghyeon Lee 0001, Hyeong Gwon Hong, Donggyu Joo, Junmo Kim 0002 |
CVPR | 4 |
| 2020 | Weight Decay Scheduling and Knowledge Distillation for Active Learning
Juseung Yun, Byungjoo Kim, Junmo Kim 0002 |
ECCV (26) | 3 |
| 2020 | Slimming ResNet by Slimming ShortcutabstractConventional network pruning methods on convolutional neural networks (CNNs) reduce the number of input or output channels of convolution layers. With these approaches, the channels in the plain network can be pruned without any restrictions. However, in the case of the ResNet based networks which have shortcuts (skip connections), the channel slimming of existing pruning methods is limited to the inside of each residual block. Since the number of Flops and parameters are also highly related to the number of channels in the shortcuts, more investigation on pruning channels in shortcuts is required. In this paper, we propose a novel pruning method, Slimming Shortcut Pruning (SSPruning), for pruning channels in shortcuts on ResNet based networks. First, we separate the long shortcut into individual regions that can be pruned independently without considering its long connections. Then, by applying our Importance Learning Gate (ILG) which learns the importance of channels globally regardless of channel type and location (i.e., in the shortcut or inside of the block), we can finally achieve an optimally pruned model. Through various experiments, we have confirmed that our method yields outstanding results when we prune the shortcuts and inside of the block together. Donggyu Joo, Junmo Kim 0002 |
ICPR | 3 |
| 2020 | Delivering Meaningful Representation for Monocular Depth Estimation
Donggyu Joo, Junmo Kim 0002 |
ICPR | 3 |
| 2020 | PBP-Net: Point Projection and Back-Projection Network for 3D Point Cloud SegmentationabstractFollowing considerable development in 3D scanning technologies, many studies have recently been proposed with various approaches for 3D vision tasks, including some methods that utilize 2D convolutional neural networks (CNNs). However, even though 2D CNNs have achieved high performance in many 2D vision tasks, existing works have not effectively applied them onto 3D vision tasks. In particular, segmentation has not been well studied because of the difficulty of dense prediction for each point, which requires rich feature representation. In this paper, we propose a simple and efficient architecture named point projection and back-projection network (PBP-Net), which leverages 2D CNNs for the 3D point cloud segmentation. 3 modules are introduced, each of which projects 3D point cloud onto 2D planes, extracts features using a 2D CNN backbone, and back-projects features onto the original 3D point cloud. To demonstrate effective 3D feature extraction using 2D CNN, we perform various experiments including comparison to recent methods. We analyze the proposed modules through ablation studies and perform experiments on object part segmentation (ShapeNet-Part dataset) and indoor scene semantic segmentation (S3DIS dataset). The experimental results show that proposed PBP-Net achieves comparable performance to existing state-of-the-art methods. JuYoung Yang, Chanho Lee, Pyunghwan Ahn, Haeil Lee, Eojindl Yi, Junmo Kim 0002 |
IROS | 6 |
| 2020 | Automatic Recognition of Children Engagement from Facial Video Using Convolutional Neural NetworksabstractAutomatic engagement recognition is a technique that is used to measure the engagement level of people in a specific task. Although previous research has utilized expensive and intrusive devices such as physiological sensors and pressure-sensing chairs, methods using RGB video cameras have become the most common because of the cost efficiency and noninvasiveness of video cameras. Automatic engagement recognition methods using video cameras are usually based on hand-crafted features and a statistical temporal dynamics modeling algorithm. This paper proposes a data-driven convolutional neural networks (CNNs)-based engagement recognition method that uses only facial images from input videos. As the amount of data in a dataset of children's engagement is insufficient for deep learning, pre-trained CNNs are utilized for low-level feature extraction from each video frame. In particular, a new layer combination for temporal dynamics modeling is employed to extract high-level features from low-level features. Experimental results on a database created using images of children from kindergarten demonstrate that the performance of the proposed method is superior to that of previous methods. The results indicate that the engagement level of children can be gauged automatically via deep learning even when the available database is deficient. Woo-han Yun, Chankyu Park, Jaehong Kim 0001, Junmo Kim 0002 |
IEEE Trans. Affect. Comput. | 5 |
| 2020 | Densely Distilled Flow-Based Knowledge Transfer in Teacher-Student Framework for Image ClassificationabstractWe propose a new teacherstudent framework (TSF)-based knowledge transfer method, in which knowledge in the form of dense flow across layers is distilled from a pre-trained "teacher" deep neural network (DNN) and transferred to another "student" DNN. In the case of distilled knowledge, multiple overlapped flow-based items of information from the pre-trained teacher DNN are densely extracted across layers. Transference of the densely extracted teacher information is then achieved in the TSF using repetitive sequential training from bottom to top between the teacher and student DNN models. In other words, to efficiently transmit extracted useful teacher information to the student DNN, we perform bottom-up step-by-step transfer of densely distilled knowledge. The performance of the proposed method in terms of image classification accuracy and fast optimization is compared with those of existing TSF-based knowledge transfer methods for application to reliable image datasets, including CIFAR-10, CIFAR-100, MNIST, and SVHN. When the dense flow-based sequential knowledge transfer scheme is employed in the TSF, the trained student ResNet more accurately reflects the rich information of the pre-trained teacher ResNet and exhibits superior accuracy to the existing TSF-based knowledge transfer methods for all benchmark datasets considered in this study. Ji-Hoon Bae 0001, Doyeob Yeo, Junho Yim, Nae-Soo Kim, Cheol-Sig Pyo, Junmo Kim 0002 |
IEEE Trans. Image Process. | 6 |
| 2019 | Learning Not to Learn: Training Deep Neural Networks With Biased DataabstractWe propose a novel regularization algorithm to train deep neural networks, in which data at training time is severely biased. Since a neural network efficiently learns data distribution, a network is likely to learn the bias information to categorize input data. It leads to poor performance at test time, if the bias is, in fact, irrelevant to the categorization. In this paper, we formulate a regularization loss based on mutual information between feature embedding and bias. Based on the idea of minimizing this mutual information, we propose an iterative algorithm to unlearn the bias information. We employ an additional network to predict the bias distribution and train the network adversarially against the feature embedding network. At the end of learning, the bias prediction network is not able to predict the bias not because it is poorly trained, but because the feature embedding network successfully unlearns the bias information. We also demonstrate quantitative and qualitative experimental results which show that our algorithm effectively removes the bias information from feature embedding. Byungju Kim, Kyungsu Kim 0003, Junmo Kim 0002 |
CVPR | 5 |
| 2019 | NLNL: Negative Learning for Noisy LabelsabstractConvolutional Neural Networks (CNNs) provide excellent performance when used for image classification. The classical method of training CNNs is by labeling images in a supervised manner as in "input image belongs to this label'' (Positive Learning; PL), which is a fast and accurate method if the labels are assigned correctly to all images. However, if inaccurate labels, or noisy labels, exist, training with PL will provide wrong information, thus severely degrading performance. To address this issue, we start with an indirect learning method called Negative Learning (NL), in which the CNNs are trained using a complementary label as in "input image does not belong to this complementary label.'' Because the chances of selecting a true label as a complementary label are low, NL decreases the risk of providing incorrect information. Furthermore, to improve convergence, we extend our method by adopting PL selectively, termed as Selective Negative Learning and Positive Learning (SelNLPL). PL is used selectively to train upon expected-to-be-clean data, whose choices become possible as NL progresses, thus resulting in superior performance of filtering out noisy data. With simple semi-supervised training technique, our method achieves state-of-the-art accuracy for noisy data classification, proving the superiority of SelNLPL's noisy data filtering ability. Youngdong Kim, Junho Yim, Juseung Yun, Junmo Kim 0002 |
ICCV | 4 |
| 2019 | Collaborative Method for Incremental Learning on Classification and GenerationabstractAlthough well-trained deep neural networks have shown remarkable performance on numerous tasks, they rapidly forget what they have learned as soon as they begin to learn with additional data with the previous data stop being provided. In this paper, we introduce a novel algorithm, Incremental Class Learning with Attribute Sharing (ICLAS), for incremental class learning with deep neural networks. As one of its component, we also introduce a generative model, incGAN, which can generate images with increased variety compared with the training data. Under challenging environment of data deficiency, ICLAS incrementally trains classification and the generation networks. Since ICLAS trains both networks, our algorithm can perform multiple times of incremental class learning. The experiments on MNIST dataset demonstrate the advantages of our algorithm. Byungju Kim, Kyungsu Kim 0003, Junmo Kim 0002 |
ICIP | 5 |
| 2019 | Capturing Long-Range Dependencies in Video CaptioningabstractMost video captioning networks rely on recurrent models, including long short-term memory (LSTM). However, these recurrent models have a long-range dependency problem; thus, they are not sufficient for video encoding. To overcome this limitation, several studies investigated the relationships between objects or entities and have shown excellent performance in video classification and video captioning. In this study, we analyze a video captioning network with a non-local block in terms of temporal capacity. We introduce a video captioning method to capture long-range temporal dependencies with a non-local block. The proposed model independently uses local and non-local features. We evaluate our approach on a Microsoft Video Description Corpus (MSVD, YouTube2Text) dataset. The experimental results show that a non-local block applied along the temporal axis can solve the long-range dependency problem of the LSTM in video captioning datasets. Yekang Lee, Sihyeon Seong, Kyungsu Kim 0003, Junmo Kim 0002 |
ICIP | 6 |
| 2019 | Learning Receptive Field Size by Learning Filter SizeabstractCovering various receptive fields within a layer is essential to effectively recognize the objects of various sizes and types at a specific layer for a convolutional neural network (CNN). In this work, we propose a novel adaptive learning method which learns the filter size (i.e. the kernel size of a convolutional filter) and distribution to learn the receptive field size. Directly optimizing with respect to the filter size is challenging because the filter size is discrete. To overcome this, we propose a masking technique, which enables the automatic allocation of resources over filters of different sizes and leads to efficient optimization. Through our proposed trainable formulation of the mask, the network self-organizes its structure through the standard backpropagation. The proposed adaptive CNN can be generalized to any single-path structures and multi-path structures as well. The effectiveness of our proposed approach is validated by several benchmark datasets compared with various previous structures on the image classification task for diverse network depths and widths. Furthermore, we demonstrate our adaptive CNN trained on a large-scale dataset can yield improved performance when applying to a transfer learning. Yegang Lee, Heechul Jung, Dongyoon Han, Kyungsu Kim 0003, Junmo Kim 0002 |
WACV | 5 |
| 2019 | Randomized Voting-Based Rigid-Body Motion SegmentationabstractIn this paper, we propose a novel rigid-body motion segmentation algorithm that uses randomized voting to assign high scores to correctly estimated models and low scores to wrongly estimated models. This algorithm is based on an epipolar geometrical representation of the camera motion, and computes scores using the distance between the feature point and the corresponding epipolar line. These scores are accumulated and utilized for motion segmentation. To evaluate the efficacy of our algorithm, we conduct a series of experiments using the Hopkins 155 data set and the UdG data set, which are representative test sets for rigid motion segmentation. Among several state-of-the-art data sets, our algorithm achieves the most accurate motion segmentation results and, in the presence of measurement noise, achieves comparable results to the other algorithms. Finally, we analyze why our motion segmentation algorithm works using probabilistic and theoretical analysis. Heechul Jung, Jeongwoo Ju, Junmo Kim 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Less-Forgetful Learning for Domain Expansion in Deep Neural NetworksabstractExpanding the domain that deep neural network has already learned without accessing old domain data is a challenging task because deep neural networks forget previously learned information when learning new data from a new domain. In this paper, we propose a less-forgetful learning method for the domain expansion scenario. While existing domain adaptation techniques solely focused on adapting to new domains, the proposed technique focuses on working well with both old and new domains without needing to know whether the input is from the old or new domain. First, we present two naive approaches which will be problematic, then we provide a new method using two proposed properties for less-forgetful learning. Finally, we prove the effectiveness of our method through experiments on image classification tasks. All datasets used in the paper, will be released on our website for someone's follow-up study. Heechul Jung, Jeongwoo Ju, Minju Jung, Junmo Kim 0002 |
AAAI | 4 |
| 2018 | Unconstrained Control of Feature Map Size Using Non-integer Strided Sampling
Donggyu Joo, Junho Yim, Junmo Kim 0002 |
BMVC | 3 |
| 2018 | Generating a Fusion Image: One's Identity and Another's ShapeabstractGenerating a novel image by manipulating two input images is an interesting research problem in the study of generative adversarial networks (GANs). We propose a new GAN-based network that generates a fusion image with the identity of input image x and the shape of input image y. Our network can simultaneously train on more than two image datasets in an unsupervised manner. We define an identity loss LI to catch the identity of image x and a shape loss LS to get the shape of y. In addition, we propose a novel training method called Min-Patch training to focus the generator on crucial parts of an image, rather than its entirety. We show qualitative results on the VGG Youtube Pose dataset, Eye dataset (MPIIGaze and UnityEyes), and the Photo-Sketch-Cartoon dataset. Donggyu Joo, Junmo Kim 0002 |
CVPR | 3 |
| 2018 | Sequential Knowledge Transfer in Teacher-Student Framework Using Densely Distilled Flow-Based InformationabstractThis paper proposes an iterative sequential knowledge transfer (KT) technique suitable for teacher-student framework (TSF)-based image classification when a state-of-the-art residual network (ResNet) is used for the TSF. To this end, we first extracted densely distilled knowledge in terms of the flow between the layers in the teacher ResNet. Subsequently, we iteratively and sequentially trained the student ResNet using the densely extracted flow-based teacher information. When using the proposed method, the trained student ResNet exhibited better performance in terms of classification accuracy and fast learning than the existing TSF-based KT methods considered in this study. Doyeob Yeo, Ji-Hoon Bae 0001, Nae-Soo Kim, Cheol-Sik Pyo, Junho Yim, Junmo Kim 0002 |
ICIP | 6 |
| 2018 | Constructing Fast Network through Deconstruction of ConvolutionabstractConvolutional neural networks have achieved great success in various vision tasks; however, they incur heavy resource costs. By using deeper and wider networks, network accuracy can be improved rapidly. However, in an environment with limited resources (e.g., mobile applications), heavy networks may not be usable. This study shows that naive convolution can be deconstructed into a shift operation and pointwise convolution. To cope with various convolutions, we propose a new shift operation called active shift layer (ASL) that formulates the amount of shift as a learnable function with shift parameters. This new layer can be optimized end-to-end through backpropagation and it can provide optimal shift values. Finally, we apply this layer to a light and fast network that surpasses existing state-of-the-art networks. Yunho Jeon, Junmo Kim 0002 |
NeurIPS | 2 |
| 2018 | Towards Flatter Loss Surface via Nonmonotonic Learning Rate Scheduling
Sihyeon Seong, Yegang Lee, Youngwook Kee, Dongyoon Han, Junmo Kim 0002 |
UAI | 5 |
| 2018 | Example image-based feature extraction for face recognition
Wonjun Hwang, Junmo Kim 0002 |
Multim. Tools Appl. | 2 |
| 2018 | ELD-Net: An Efficient Deep Learning Architecture for Accurate Saliency DetectionabstractRecent advances in saliency detection have utilized deep learning to obtain high-level features to detect salient regions in scenes. These advances have yielded results superior to those reported in past work, which involved the use of hand-crafted low-level features for saliency detection. In this paper, we propose ELD-Net, a unified deep learning framework for accurate and efficient saliency detection. We show that hand-crafted features can provide complementary information to enhance saliency detection that uses only high-level features. Our method uses both low-level and high-level features for saliency detection. High-level features are extracted using GoogLeNet, and low-level features evaluate the relative importance of a local region using its differences from other regions in an image. The two feature maps are independently encoded by the convolutional and the ReLU layers. The encoded low-level and high-level features are then combined by concatenation and convolution. Finally, a linear fully connected layer is used to evaluate the saliency of a queried region. A full resolution saliency map is obtained by querying the saliency of each local region of an image. Since the high-level features are encoded at low resolution, and the encoded high-level features can be reused for every query region, our ELD-Net is very fast. Our experiments show that our method outperforms state-of-the-art deep learning-based saliency detection methods. Gayoung Lee, Yu-Wing Tai, Junmo Kim 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Deep Facial Age Estimation Using Conditional Multitask Learning With Weak Label ExpansionabstractAccurate age estimation from a facial image is quite challenging, since physical age and apparent age can be quite different, and this difference is dependent on gender, ethnicity, and many other factors. Multitask deep learning is one of the approach to improve age estimation by employing auxiliary tasks, such as gender recognition, that are related to the primary task. However, in traditional multitask learning for age estimation, the relationship between the primary and auxiliary tasks is difficult to describe; how the auxiliary tasks enhance the model for the primary objective is ambiguous. In this letter, we propose a conditional multitask learning method that architecturally factorizes an age variable into gender-conditioned age probabilities in a deep neural network. The lack of accurate training labels with discrete age values is another critical limitation to training age estimation models. Therefore, we propose a label expansion method that increases the number of accurate labels from weakly supervised categorical labels. To verify the generality of the proposed method, we perform intensive experiments on the publicly available MORPH-II and FG-NET datasets. The proposed methods outperform state-of-the art methods in both age estimation and gender recognition accuracy. These performance gains are verified on well-known deep network architectures-VGG-16, CASIA-WebFace, and Alexnet-to confirm the proposed methods generality. ByungIn Yoo, Youngjun Kwak, Youngsung Kim, Changkyu Choi, Junmo Kim 0002 |
IEEE Signal Process. Lett. | 5 |
| 2018 | Unified Simultaneous Clustering and Feature Selection for Unlabeled and Labeled DataabstractThis paper proposes a novel feature selection method, namely, unified simultaneous clustering feature selection (USCFS). A regularized regression with a new type of target matrix is formulated to select the most discriminative features among the original features from labeled or unlabeled data. The regression with -norm regularization allows the projection matrix to represent an effective selection of discriminative features. For unsupervised feature selection, the target matrix discovers label-like information not from the original data points but rather from projected data points, which are of a reduced dimensionality. Without the aid of an affinity graph-based local structure learning method, USCFS allows the target matrix to capture latent cluster centers via orthogonal basis clustering and to simultaneously select discriminative features guided by latent cluster centers. When class labels are available, the target matrix is also able to find latent class labels by regarding the ground-truth class labels as an approximate guide. Hence, supervised feature selection is realized using these latent class labels, which may differ from the ground-truth class labels. Experimental results demonstrate the effectiveness of the proposed method. Specifically, the proposed method outperforms the state-of-the-art methods on diverse real-world data sets for both the supervised and the unsupervised feature selection. Dongyoon Han, Junmo Kim 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | Deep Pyramidal Residual NetworksabstractDeep convolutional neural networks (DCNNs) have shown remarkable performance in image classification tasks in recent years. Generally, deep neural network architectures are stacks consisting of a large number of convolutional layers, and they perform downsampling along the spatial dimension via pooling to reduce memory usage. Concurrently, the feature map dimension (i.e., the number of channels) is sharply increased at downsampling locations, which is essential to ensure effective performance because it increases the diversity of high-level attributes. This also applies to residual networks and is very closely related to their performance. In this research, instead of sharply increasing the feature map dimension at units that perform downsampling, we gradually increase the feature map dimension at all units to involve as many locations as possible. This design, which is discussed in depth together with our new insights, has proven to be an effective means of improving generalization ability. Furthermore, we propose a novel residual unit capable of further improving the classification accuracy with our new network architecture. Experiments on benchmark CIFAR-10, CIFAR-100, and ImageNet datasets have shown that our network architecture has superior generalization ability compared to the original residual networks. Dongyoon Han, Jiwhan Kim, Junmo Kim 0002 |
CVPR | 3 |
| 2017 | Active Convolution: Learning the Shape of Convolution for Image ClassificationabstractIn recent years, deep learning has achieved great success in many computer vision applications. Convolutional neural networks (CNNs) have lately emerged as a major approach to image classification. Most research on CNNs thus far has focused on developing architectures such as the Inception and residual networks. The convolution layer is the core of the CNN, but few studies have addressed the convolution unit itself. In this paper, we introduce a convolution unit called the active convolution unit (ACU). A new convolution has no fixed shape, because of which we can define any form of convolution. Its shape can be learned through backpropagation during training. Our proposed unit has a few advantages. First, the ACU is a generalization of convolution, it can define not only all conventional convolutions, but also convolutions with fractional pixel coordinates. We can freely change the shape of the convolution, which provides greater freedom to form CNN structures. Second, the shape of the convolution is learned while training and there is no need to tune it by hand. Third, the ACU can learn better than a conventional unit, where we obtained the improvement simply by changing the conventional convolution to an ACU. We tested our proposed method on plain and residual networks, and the results showed significant improvement using our method on various datasets and architectures in comparison with the baseline. Code is available at https://github.com/jyh2986/Active-Convolution. Yunho Jeon, Junmo Kim 0002 |
CVPR | 2 |
| 2017 | A Gift from Knowledge Distillation: Fast Optimization, Network Minimization and Transfer LearningabstractWe introduce a novel technique for knowledge transfer, where knowledge from a pretrained deep neural network (DNN) is distilled and transferred to another DNN. As the DNN performs a mapping from the input space to the output space through many layers sequentially, we define the distilled knowledge to be transferred in terms of flow between layers, which is calculated by computing the inner product between features from two layers. When we compare the student DNN and the original network with the same size as the student DNN but trained without a teacher network, the proposed method of transferring the distilled knowledge as the flow between two layers exhibits three important phenomena: (1) the student DNN that learns the distilled knowledge is optimized much faster than the original model, (2) the student DNN outperforms the original DNN, and (3) the student DNN can learn the distilled knowledge from a teacher DNN that is trained at a different task, and the student DNN outperforms the original DNN that is trained from scratch. Junho Yim, Donggyu Joo, Ji-Hoon Bae 0001, Junmo Kim 0002 |
CVPR | 4 |
| 2017 | Refine pedestrian detections by referring to features in different waysabstractThe performance of object detection has been improved as the success of deep architectures. The main algorithm predominantly used for general detection is Faster R-CNN because of their high accuracy and fast inference time. In pedestrian detection, Region Proposal Network (RPN) itself which is used for region proposals in Faster R-CNN can be used as a pedestrian detector. Also, RPN even shows better performance than Faster R-CNN for pedestrian detection. However, RPN generates severe false positives such as high score backgrounds and double detections because it does not have downstream classifier. From this observations, we made a network to refine results generated from the RPN. Our Refinement Network refers to the feature maps of the RPN and trains the network to rescore severe false positives. Also, we found that different type of feature referencing method is crucial for improving performance. Our network showed better accuracy than RPN with almost same speed on Caltech Pedestrian Detection benchmark. Jaemyung Lee 0003, Sihaeng Lee, Youngdong Kim, Janghyeon Lee 0001, Junmo Kim 0002 |
Intelligent Vehicles Symposium | 5 |
| 2017 | Sequential Convex Programming for Computing Information-Theoretic Minimal Partitions: Nonconvex Nonsmooth OptimizationabstractWe consider an unsupervised image segmentation problem---from figure-ground separation to multiregion partitioning---that consists of maximal distribution separation (in terms of mutual information) with spatial regularity (total variation regularization), which is what we call information-theoretic minimal partitioning. Adopting the bounded variation framework, we provide an in-depth analysis of the problem which establishes theoretical foundations and investigates the structure of the associated energy from a variational perspective. In doing so, we show that the objective exhibits a form of difference of convex functionals, which leads us to a class of large-scale nonconvex optimization problems where convex optimization techniques can be successfully applied. In this regard, we propose sequential convex programming based on the philosophy of stochastic optimization and the Chambolle--Pock primal-dual algorithm. The key idea behind it is to construct a stochastic family of convex approximations of the original nonconvex function and sequentially minimize the associated subproblems. Indeed, its stochastic nature makes it possible to often escape from bad local minima toward near-optimal solutions, where such optimality can be justified in terms of the recent findings in statistical physics regarding a striking characteristic of the high-dimensional landscapes. We experimentally demonstrate such a favorable ability of the proposed algorithm as well as show the capacity of our approach in numerous experiments. The preliminary conference paper can be found in [Y. Kee, M. Souiai, D. Cremers, and J. Kim, Proceedings of the IEEE Conference on Computer Vision and Pattern, Recognition, 2014]. Youngwook Kee, Yegang Lee, Mohamed Souiai, Daniel Cremers, Junmo Kim 0002 |
SIAM J. Imaging Sci. | 5 |
| 2017 | Tunnel Effect in CNNs: Image Reconstruction From Max Switch LocationsabstractIn this letter, we show that reconstruction of an image passed through a neural network is possible, using only the locations of the max pool activations. This was demonstrated with an architecture consisting of an encoder and a decoder. The decoder is a mirrored version of the encoder, where convolutions are replaced with deconvolutions and poolings are replaced with unpooling layers. The locations of the max pool switches are transmitted to the corresponding unpooling layer. The reconstruction is computed only from these switches without the use of feature values. Using only the max switch location information of the pool layers, a surprisingly good image reconstruction can be achieved. We examine this effect in various architectures, as well as how the quality of the reconstruction is affected by the number of features. We also compare the reconstruction with an encoder with randomly initialized weights with an encoder pretrained for classification. Finally, we give recommendations for future architecture decisions. Matthieu de La Roche Saint Andre, Laura Rieger, Morten Rieger Hannemose, Junmo Kim 0002 |
IEEE Signal Process. Lett. | 4 |
| 2017 | Uncorrelated Component Analysis-Based HashingabstractThe approximate nearest neighbor (ANN) search problem is important in applications such as information retrieval. Several hashing-based search methods that provide effective solutions to the ANN search problem have been proposed. However, most of these focus on similarity preservation and coding error minimization, and pay little attention to optimizing the precision-recall curve or receiver operating characteristic curve. In this paper, we propose a novel projection-based hashing method that attempts to maximize precision and recall. We first introduce an uncorrelated component analysis (UCA) transformation by examining precision and recall, and then propose a UCA-based hashing method. The proposed method is evaluated with a variety of data sets. The results show that UCA-based hashing outperforms state-of-the-art methods, and has computationally efficient training and encoding processes. Sungryull Sohn, Junmo Kim 0002 |
IEEE Trans. Image Process. | 3 |
| 2016 | Supervised Hashing via Uncorrelated Component AnalysisabstractThe Approximate Nearest Neighbor (ANN) search problem is important in applications such as information retrieval. Several hashing-based search methods that provide effective solutions to the ANN search problem have been proposed. However, most of these focus on similarity preservation and coding error minimization, and pay little attention to optimizing the precision-recall curve or receiver operating characteristic curve. In this paper, we propose a novel projection-based hashing method that attempts to maximize the precision and recall. We first introduce an uncorrelated component analysis (UCA) by examining the precision and recall, and then propose a UCA-based hashing method. The proposed method is evaluated with a variety of datasets. The results show that UCA-based hashing outperforms state-of-the-art methods, and has computationally efficient training and encoding processes. Sungryull Sohn, Junmo Kim 0002 |
AAAI | 3 |
| 2016 | Deep Saliency with Encoded Low Level Distance Map and High Level FeaturesabstractRecent advances in saliency detection have utilized deep learning to obtain high level features to detect salient regions in a scene. These advances have demonstrated superior results over previous works that utilize hand-crafted low level features for saliency detection. In this paper, we demonstrate that hand-crafted features can provide complementary information to enhance performance of saliency detection that utilizes only high level features. Our method utilizes both high level and low level features for saliency detection under a unified deep learning framework. The high level features are extracted using the VGG-net, and the low level features are compared with other parts of an image to form a low level distance map. The low level distance map is then encoded using a convolutional neural network(CNN) with multiple 1 1 convolutional and ReLU layers. We concatenate the encoded low level distance map and the high level features, and connect them to a fully connected neural network classifier to evaluate the saliency of a query region. Our experiments show that our method can further improve the performance of state-of-the-art deep learning-based saliency detection methods. Gayoung Lee, Yu-Wing Tai, Junmo Kim 0002 |
CVPR | 3 |
| 2016 | Plankton classification on imbalanced large scale database via convolutional neural networks with transfer learningabstractPlankton image classification plays an important role in the ocean ecosystems research. Recently, a large scale database for plankton classification with over 3 million images annotated with over 100 classes was released. However, the database suffers from imbalanced class distribution in which over 90% of images belong to only 5 classes. Due to this class-imbalance problem, the existing classification approaches are limited to label the data only to major classes, ignoring the small-sized classes. In this paper, we propose a fine-grained classification method for large scale plankton database based on convolutional neural networks (CNN). To overcome the class-imbalance problem, we incorporate transfer learning by pre-training CNN with class-normalized data and fine-tuning with original data. The class-normalized data is constructed by reducing the number of data via random sampling, for large-sized classes. In experiments, our method showed superior classification accuracy compared to both CNN without transfer learning and CNN with transfer learning via other data augmentation techniques. Hansang Lee, Minseok Park, Junmo Kim 0002 |
ICIP | 3 |
| 2016 | Seed growing for interactive image segmentation with geodesic votingabstractIn this paper, we propose a novel seed growing framework for interactive image segmentation. We first formulate the seed dependency problem in interactive segmentation and overcome it by expanding the seed automatically. To expand the user-input seed, we generate the seed distance maps based on color distribution dissimilarity, locational prior, and geodesic distance. Using these seed distance maps, we expand the seed by classifying the image into a trimap with unanimous voting. We then extract the skeleton from the foreground and background regions. Experiments show that the proposed framework provides significant support for existing interactive segmentation techniques. Sunjeong Park, Hansang Lee, Junmo Kim 0002 |
ICIP | 3 |
| 2016 | Salient Region Detection via High-Dimensional Color Transform and Local Spatial SupportabstractIn this paper, we introduce a novel approach to automatically detect salient regions in an image. Our approach consists of global and local features, which complement each other to compute a saliency map. The first key idea of our work is to create a saliency map of an image by using a linear combination of colors in a high-dimensional color space. This is based on an observation that salient regions often have distinctive colors compared with backgrounds in human perception, however, human perception is complicated and highly nonlinear. By mapping the low-dimensional red, green, and blue color to a feature vector in a high-dimensional color space, we show that we can composite an accurate saliency map by finding the optimal linear combination of color coefficients in the high-dimensional color space. To further improve the performance of our saliency estimation, our second key idea is to utilize relative location and color contrast between superpixels as features and to resolve the saliency estimation from a trimap via a learning-based algorithm. The additional local features and learning-based algorithm complement the global estimation from the high-dimensional color transform-based algorithm. The experimental results on three benchmark datasets show that our approach is effective in comparison with the previous state-of-the-art saliency estimation methods. Jiwhan Kim, Dongyoon Han, Yu-Wing Tai, Junmo Kim 0002 |
IEEE Trans. Image Process. | 4 |
| 2015 | Unsupervised Simultaneous Orthogonal basis Clustering Feature SelectionabstractIn this paper, we propose a novel unsupervised feature selection method: Simultaneous Orthogonal basis Clustering Feature Selection (SOCFS). To perform feature selection on unlabeled data effectively, a regularized regression-based formulation with a new type of target matrix is designed. The target matrix captures latent cluster centers of the projected data points by performing orthogonal basis clustering, and then guides the projection matrix to select discriminative features. Unlike the recent unsupervised feature selection methods, SOCFS does not explicitly use the pre-computed local structure information for data points represented as additional terms of their objective functions, but directly computes latent cluster information by the target matrix conducting orthogonal basis clustering in a single unified term of the proposed objective function. It turns out that the proposed objective function can be minimized by a simple optimization algorithm. Experimental results demonstrate the effectiveness of SOCFS achieving the state-of-the-art results with diverse real world datasets. Dongyoon Han, Junmo Kim 0002 |
CVPR | 2 |
| 2015 | Rotating your face using multi-task deep neural networkabstractFace recognition under viewpoint and illumination changes is a difficult problem, so many researchers have tried to solve this problem by producing the pose- and illumination- invariant feature. Zhu et al. [26] changed all arbitrary pose and illumination images to the frontal view image to use for the invariant feature. In this scheme, preserving identity while rotating pose image is a crucial issue. This paper proposes a new deep architecture based on a novel type of multitask learning, which can achieve superior performance in rotating to a target-pose face image from an arbitrary pose and illumination image while preserving identity. The target pose can be controlled by the user's intention. This novel type of multi-task model significantly improves identity preservation over the single task model. By using all the synthesized controlled pose images, called Controlled Pose Image (CPI), for the pose-illumination-invariant feature and voting among the multiple face recognition results, we clearly outperform the state-of-the-art algorithms by more than 4~6% on the MultiPIE dataset. Junho Yim, Heechul Jung, ByungIn Yoo, Changkyu Choi, Du-Sik Park, Junmo Kim 0002 |
CVPR | 6 |
| 2015 | Joint Fine-Tuning in Deep Neural Networks for Facial Expression RecognitionabstractTemporal information has useful features for recognizing facial expressions. However, to manually design useful features requires a lot of effort. In this paper, to reduce this effort, a deep learning technique, which is regarded as a tool to automatically extract useful features from raw data, is adopted. Our deep network is based on two different models. The first deep network extracts temporal appearance features from image sequences, while the other deep network extracts temporal geometry features from temporal facial landmark points. These two models are combined using a new integration method in order to boost the performance of the facial expression recognition. Through several experiments, we show that the two models cooperate with each other. As a result, we achieve superior performance to other state-of-the-art methods in the CK+ and Oulu-CASIA databases. Furthermore, we show that our new integration method gives more accurate results than traditional methods, such as a weighted summation and a feature concatenation method. Heechul Jung, Sihaeng Lee, Junho Yim, Sunjeong Park, Junmo Kim 0002 |
ICCV | 5 |
| 2015 | Entropy Minimization for Convex Relaxation ApproachesabstractDespite their enormous success in solving hard combinatorial problems, convex relaxation approaches often suffer from the fact that the computed solutions are far from binary and that subsequent heuristic binarization may substantially degrade the quality of computed solutions. In this paper, we propose a novel relaxation technique which incorporates the entropy of the objective variable as a measure of relaxation tightness. We show both theoretically and experimentally that augmenting the objective function with an entropy term gives rise to more binary solutions and consequently solutions with a substantially tighter optimality gap. We use difference of convex function (DC) programming as an efficient and provably convergent solver for the arising convex-concave minimization problem. We evaluate this approach on three prominent non-convex computer vision challenges: multi-label inpainting, image segmentation and spatio-temporal multi-view reconstruction. These experiments show that our approach consistently yields better solutions with respect to the original integral optimization problem. Mohamed Souiai, Martin R. Oswald, Youngwook Kee, Junmo Kim 0002, Marc Pollefeys, Daniel Cremers |
ICCV | 4 |
| 2015 | Facial age estimation via extended curvature Gabor filterabstractFacial age estimation is a process of identifying the age of a single face in an image or a video. Since age information can be used in many environments such as security, surveillance, and entertainment, age estimation has recently received much attention from researchers. In this paper, we propose an automatic age estimation method via extended curvature Gabor (ECG) features and a learning-based technique. Instead of conventional Gabor Filters, we use ECG filters to extract curvature information from a face image, which is useful for estimating age. We use a feature selection method to reduce the computational complexity and prove the effectiveness of ECG features at the same time. We use a regression algorithm to estimate the age of the test face image. As a result, our work achieves a competitive performance compared with other recent works in terms of age estimation. Jiwhan Kim, Dongyoon Han, Sungryull Sohn, Junmo Kim 0002 |
ICIP | 4 |
| 2015 | Face recognition using Extended Curvature Gabor classifier bunch
Wonjun Hwang, Xiangsheng Huang, Stan Z. Li, Junmo Kim 0002 |
Pattern Recognit. | 4 |
| 2015 | Entropy Minimization for Groupwise Planar Shape Co-alignment and its ApplicationsabstractWe propose an information-theoretic criterion, entropy estimate, for the joint alignment of a group of shape observations drawn from an unknown shape distribution. Employing a nonparametric density estimation technique with implicit shape representation, we minimize the entropy estimate with respect to the pose parameters of similarity transformations based on gradient descent optimization for which we provide implementation details. We demonstrate the capacity of our approach in numerous experiments with an application of building a shape prior to prostate MR image segmentation. Youngwook Kee, Hansang Lee, Junho Yim, Daniel Cremers, Junmo Kim 0002 |
IEEE Signal Process. Lett. | 5 |
| 2015 | Markov Network-Based Unified Classifier for Face RecognitionabstractIn this paper, we propose a novel unifying framework using a Markov network to learn the relationships among multiple classifiers. In face recognition, we assume that we have several complementary classifiers available, and assign observation nodes to the features of a query image and hidden nodes to those of gallery images. Under the Markov assumption, we connect each hidden node to its corresponding observation node and the hidden nodes of neighboring classifiers. For each observation-hidden node pair, we collect the set of gallery candidates most similar to the observation instance, and capture the relationship between the hidden nodes in terms of a similarity matrix among the retrieved gallery images. Posterior probabilities in the hidden nodes are computed using the belief propagation algorithm, and we use marginal probability as the new similarity value of the classifier. The novelty of our proposed framework lies in the method that considers classifier dependence using the results of each neighboring classifier. We present the extensive evaluation results for two different protocols, known and unknown image variation tests, using four publicly available databases: 1) the Face Recognition Grand Challenge ver. 2.0; 2) XM2VTS; 3) BANCA; and 4) Multi-PIE. The result shows that our framework consistently yields improved recognition rates in various situations. Wonjun Hwang, Junmo Kim 0002 |
IEEE Trans. Image Process. | 2 |
| 2014 | Rigid Motion Segmentation Using Randomized VotingabstractIn this paper, we propose a novel rigid motion segmentation algorithm called randomized voting (RV). This algorithm is based on epipolar geometry, and computes a score using the distance between the feature point and the corresponding epipolar line. This score is accumulated and utilized for final grouping. Our algorithm basically deals with two frames, so it is also applicable to the two-view motion segmentation problem. For evaluation of our algorithm, Hopkins 155 dataset, which is a representative test set for rigid motion segmentation, is adopted, it consists of two and three rigid motions. Our algorithm has provided the most accurate motion segmentation results among all of the state-of-the-art algorithms. The average error rate is 0.77%. In addition, when there is measurement noise, our algorithm is comparable with other state-of-the-art algorithms. Heechul Jung, Jeongwoo Ju, Junmo Kim 0002 |
CVPR | 3 |
| 2014 | A Convex Relaxation of the Ambrosio-Tortorelli Elliptic Functionals for the Mumford-Shah FunctionalabstractIn this paper, we revisit the phase-field approximation of Ambrosio and Tortorelli for the Mumford -- Shah functional. We then propose a convex relaxation for it to attempt to compute globally optimal solutions rather than solving the nonconvex functional directly, which is the main contribution of this paper. Inspired by McCormick's seminal work on factorable nonconvex problems, we split a nonconvex product term that appears in the Ambrosio -- Tortorelli elliptic functionals in a way that a typical alternating gradient method guarantees a globally optimal solution without completely removing coupling effects. Furthermore, not only do we provide a fruitful analysis of the proposed relaxation but also demonstrate the capacity of our relaxation in numerous experiments that show convincing results compared to a naive extension of the McCormick relaxation and its quadratic variant. Indeed, we believe the proposed relaxation and the idea behind would open up a possibility for convexifying a new class of functions in the context of energy minimization for computer vision. Youngwook Kee, Junmo Kim 0002 |
CVPR | 2 |
| 2014 | Sequential Convex Relaxation for Mutual Information-Based Unsupervised Figure-Ground SegmentationabstractWe propose an optimization algorithm for mutual information-based unsupervised figure-ground separation. The algorithm jointly estimates the color distributions of the foreground and background, and separates them based on their mutual information with geometric regularity. To this end, we revisit the notion of mutual information and reformulate it in terms of the photometric variable and the indicator function; and propose a sequential convex optimization strategy for solving the nonconvex optimization problem that arises. By minimizing a sequence of convex sub-problems for the mutual-information-based nonconvex energy, we efficiently attain high quality solutions for challenging unsupervised figure-ground segmentation problems. We demonstrate the capacity of our approach in numerous experiments that show convincing fully unsupervised figure-ground separation, in terms of both segmentation quality and robustness to initialization. Youngwook Kee, Mohamed Souiai, Daniel Cremers, Junmo Kim 0002 |
CVPR | 4 |
| 2014 | Salient Region Detection via High-Dimensional Color TransformabstractIn this paper, we introduce a novel technique to automatically detect salient regions of an image via high-dimensional color transform. Our main idea is to represent a saliency map of an image as a linear combination of high-dimensional color space where salient regions and backgrounds can be distinctively separated. This is based on an observation that salient regions often have distinctive colors compared to the background in human perception, but human perception is often complicated and highly nonlinear. By mapping a low dimensional RGB color to a feature vector in a high-dimensional color space, we show that we can linearly separate the salient regions from the background by finding an optimal linear combination of color coefficients in the high-dimensional color space. Our high dimensional color space incorporates multiple color representations including RGB, CIELab, HSV and with gamma corrections to enrich its representative power. Our experimental results on three benchmark datasets show that our technique is effective, and it is computationally efficient in comparison to previous state-of-the-art techniques. Jiwhan Kim, Dongyoon Han, Yu-Wing Tai, Junmo Kim 0002 |
CVPR | 4 |
| 2014 | Randomized decision bush: Combining global shape parameters and local scalable descriptors for human body parts recognitionabstractThis paper presents a novel method which combines global shape parameters and scalable local descriptors for accurate body parts recognition from a single depth image in real-time. Human poses are of extremely large variation in aspects of visual shapes, because human can take poses from daily activities to gymnastic actions. In order to cover wide-range of the human poses, the proposed algorithm employs a unified structure which combines pose clustering and body parts classification. We name the proposed method Randomized Decision Bush (RDB). Specifically, global shape parameters which can discriminate coarse level shapes are utilized for pose clustering while scalable local shape descriptors are employed for accurate classification. RDB splits the various human poses into multiple clusters which contain similar shapes of the poses. As a result, it provides robust clustering which enables fine level classification within the cluster. The experimental results show improvements on recognizing body parts due to the pose clustering and classification with scalable local descriptors. Additionally, we significantly reduce the complexity of training a large number of human shapes. ByungIn Yoo, Jae-Joon Han, Changkyu Choi, Du-Sik Park, Junmo Kim 0002 |
ICIP | 6 |
| 2013 | Markov Network-Based Unified Classifier for Face IdentificationabstractWe propose a novel unifying framework using a Markov network to learn the relationship between multiple classifiers in face recognition. We assume that we have several complementary classifiers and assign observation nodes to the features of a query image and hidden nodes to the features of gallery images. We connect each hidden node to its corresponding observation node and to the hidden nodes of other neighboring classifiers. For each observation-hidden node pair, we collect a set of gallery candidates that are most similar to the observation instance, and the relationship between the hidden nodes is captured in terms of the similarity matrix between the collected gallery images. Posterior probabilities in the hidden nodes are computed by the belief-propagation algorithm. The novelty of the proposed framework is the method that takes into account the classifier dependency using the results of each neighboring classifier. We present extensive results on two different evaluation protocols, known and unknown image variation tests, using three different databases, which shows that the proposed framework always leads to good accuracy in face recognition. Wonjun Hwang, Kyungshik Noh, Junmo Kim 0002 |
ICCV | 3 |
| 2013 | Markov network-based multiple classifier for face image retrievalabstractWe propose a new face-recognition framework to learn the relationship between multiple classifiers using a Markov network. For each image, we make three face models based on different distances between two eye locations. The novelty of the proposed method lies in that the method not only compares the query and target images at the three different levels, but also takes into account the statistical dependency between the three different models. This dependency is captured by a Markov network, which we describe by a graphical model, where query models are observation nodes, target models are hidden nodes, and the network line represents their relationships. For each observation-hidden node pair, we collect a set of target candidates that are most similar to the observation, and the relationship between the hidden nodes is captured in terms of the similarity between target images. Posterior probabilities at the three hidden nodes of the Markov network are computed by a belief-propagation algorithm. We evaluate the proposed method using FRGC ver 2.0, XM2VTS, BANCA, and PIE databases, which demonstrates its superiority under the untrained variations. Wonjun Hwang, Kyungshik Noh, Junmo Kim 0002 |
ICIP | 3 |
| 2013 | A novel method for salient object detection via compactness measurementabstractSalient object detection is a process of extracting an object which is visually attractive from a single image or a video. As a powerful technique for automatic image or video segmentation, saliency detection has been focused and studied recently. In this paper, we propose a novel method for salient object detection without training or learning-based techniques. The proposed framework consists of two major steps, the generation of saliency map candidates and the selection of an optimal saliency map. To generate saliency map candidates, prior maps based on combinations of RGB color components are proposed. To select the optimal saliency map among the candidates, we propose a compactness measure, which evaluates the degree to which the generated saliency maps show objects. As a result, among recent works on saliency detection, our saliency detection method achieves the highest performance in terms of saliency detection. Jiwhan Kim, Hansang Lee, Junmo Kim 0002 |
ICIP | 3 |
| 2013 | An efficient lane detection algorithm for lane departure detectionabstractIn this paper, we propose an efficient lane detection algorithm for lane departure detection; this algorithm is suitable for low computing power systems like automobile black boxes. First, we extract candidate points, which are support points, to extract a hypotheses as two lines. In this step, Haar-like features are used, and this enables us to use an integral image to remove computational redundancy. Second, our algorithm verifies the hypothesis using defined rules. These rules are based on the assumption that the camera is installed at the center of the vehicle. Finally, if a lane is detected, then a lane departure detection step is performed. As a result, our algorithm has achieved 90.16% detection rate; the processing time is approximately 0.12 milliseconds per frame without any parallel computing. Heechul Jung, Junggon Min, Junmo Kim 0002 |
Intelligent Vehicles Symposium | 3 |
| 2012 | Anterior Cruciate Ligament Segmentation from Knee MR Images Using Graph Cuts with Geometric and Probabilistic Shape Constraints
Hansang Lee, Helen Hong, Junmo Kim 0002 |
ACCV (2) | 3 |
| 2012 | Block Poisson Method and its application to large scale image editingabstractPoisson Image Editing(PIE) is one of the most well known methods for seamless image interpolation. However, direct application of PIE to large scale images(e.g. satellite image with billions of pixels) can result in hours of operation time. We propose a modification of PIE method, which enables the original method to be performed in smaller blocks, largely improving the operation time. The key idea of Block Poisson Method(BPM) is to prepare the boundary conditions for each block and perform PIE separately. Our algorithm can be effectively applied to satellite image stitching, especially for cases where the satellite images are significantly degraded at detectors' edge. Resulting images show that BPM performs fairly the same as the original PIE with considerably reduced operation time. Dewey H. Lee, Seok Bong Yoo, Jong Beom Ra, Junmo Kim 0002 |
ICIP | 5 |
| 2012 | Achievable Rates of Multi-Antenna Downlink Channels with Peak Power ConstraintsabstractThis paper considers a Gaussian multiple-input single-output (MISO) broadcast channel with a per-antenna peak power constraint (or simply peak power constraint). It is more realistic to consider the peak power constraint on each transmit antenna because in many practical implementations each antenna is equipped with its own power amplifier. Assuming the perfect knowledge of the channel state information (CSI) at the transmitter, we propose an achievable scheme using dirty-tape coding (DTC). The ideal dirty-paper coding (DPC), which is a capacity-achieving scheme for the Gaussian multiple-input multiple-output (MIMO) broadcast channel, cannot be used for our model because its optimal input distribution is Gaussian and thus the peak power constraint is violated. On the other hand, the channel input of DTC is uniformly distributed in a fixed range, which helps to control the peak power of the transmit signal easily. We also present an algorithm that finds capacity-achieving beamforming vectors and power allocation factors under a per-antenna average power constraint and use the optimized parameters in the proposed scheme. Simulation results show that the proposed scheme provides gains of 2.7dB over a non-DTC scheme based on minimum mean square error (MMSE) beamforming at high signal-to-noise ratio (SNR) when there are three receivers and the transmitter has three antennas. Ihn-Jung Baik, Sae-Young Chung, Junmo Kim 0002 |
IEEE Trans. Commun. | 3 |
| 2011 | Face Recognition System Using Multiple Face Model of Hybrid Fourier Feature Under Uncontrolled Illumination VariationabstractThe authors present a robust face recognition system for large-scale data sets taken under uncontrolled illumination variations. The proposed face recognition system consists of a novel illumination-insensitive preprocessing method, a hybrid Fourier-based facial feature extraction, and a score fusion scheme. First, in the preprocessing stage, a face image is transformed into an illumination-insensitive image, called an "integral normalized gradient image," by normalizing and integrating the smoothed gradients of a facial image. Then, for feature extraction of complementary classifiers, multiple face models based upon hybrid Fourier features are applied. The hybrid Fourier features are extracted from different Fourier domains in different frequency bandwidths, and then each feature is individually classified by linear discriminant analysis. In addition, multiple face models are generated by plural normalized face images that have different eye distances. Finally, to combine scores from multiple complementary classifiers, a log likelihood ratio-based score fusion scheme is applied. The proposed system using the face recognition grand challenge (FRGC) experimental protocols is evaluated; FRGC is a large available data set. Experimental results on the FRGC version 2.0 data sets have shown that the proposed method shows an average of 81.49% verification rate on 2-D face images under various environmental variations such as illumination changes, expression changes, and time elapses. Wonjun Hwang, Haitao Wang 0006, Seok-Cheol Kee, Junmo Kim 0002 |
IEEE Trans. Image Process. | 5 |
| 2010 | Two-User Multi-Antenna Downlink Channels with Peak Power ConstraintsabstractThis paper considers a two user Gaussian multiple-input single-output (MISO) broadcast channel with a per-antenna peak power constraint (or simply peak power constraint). It is more realistic to consider the peak power constraint on each transmit antenna because each antenna is equipped with its own power amplifier in many practical implementations. Assuming the perfect channel state information (CSI) at the transmitter, we propose an achievable scheme using a dirty-tape coding (DTC). The uniform input in a fixed range of the DTC scheme helps to control the peak power of the transmit signal easily. We also present an optimization algorithm that finds the capacity achieving beamforming vectors and power allocations under a per-antenna average power constraint used in our achievable scheme. Simulation results show that as the transmit power increases, the achievable rate region under the peak power constraint is getting close to the capacity region under the reduced per-antenna average power constraint by 1/3. Compared to a non-DTC scheme based on minimum mean square error (MMSE) beamforming, the proposed scheme performs better. Ihn-Jung Baik, Sae-Young Chung, Junmo Kim 0002 |
GLOBECOM | 3 |
| 2009 | Face recognition using gender informationabstractIn this paper, we propose a novel method using gender information for achieving better performances of face recognition systems. Gender is one of the important factors for recognizing appearance of human faces and there are many studies on gender classifications such. However, the gender information is not actively applied in vision-based face recognition tasks, because we cannot find out human identity using only gender information. Therefore, we design the face recognition system based on the gender-based facial features with global facial features, and moreover, gender-based score normalization method for verification task. For fair evaluations, we use FRGC database known as a large size face image database. Wonjun Hwang, Haibing Ren, Seok-Cheol Kee, Junmo Kim 0002 |
ICIP | 5 |