Nojun Kwak

dblp:49/2806 · DBLP profile ↗
← Back
129ranked-venue papers
19as first author
58since 2021 · last 2026
0000-0002-1792-0327ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 99 · 17 first-author · 43 since 2021Graphics, computer vision, multimedia, augmented reality and games · 72 · 1 first-author · 40 since 2021Systems, architecture and hardware · 3 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 Targeted Data Protection for Diffusion Model by Matching Training Trajectory
abstract
Recent advancements in diffusion models have made fine-tuning text-to-image models for personalization increasingly accessible, but have also raised significant concerns regarding unauthorized data usage and privacy infringement. Current protection methods are limited to passively degrading image quality, failing to achieve stable control. While Targeted Data Protection (TDP) offers a promising paradigm for active redirection toward user-specified target concepts, existing TDP attempts suffer from poor controllability due to snapshot-matching approaches that fail to account for complete learning dynamics. We introduce TAFAP (Trajectory Alignment via Fine-tuning with Adversarial Perturbations), the first method to successfully achieve effective TDP by controlling the entire training trajectory. Unlike snapshot-based methods whose protective influence is easily diluted as training progresses, TAFAP employs trajectory-matching inspired by dataset distillation to enforce persistent, verifiable transformations throughout fine-tuning. We validate our method through extensive experiments, demonstrating the first successful targeted transformation in diffusion models with simultaneous control over both identity and visual patterns. TAFAP significantly outperforms existing TDP attempts, achieving robust redirection toward target concepts while maintaining high image quality. This work enables verifiable safeguards and provides a new framework for controlling and tracing alterations in diffusion model outputs.
Hojun Lee 0002, Mijin Koo, Yeji Song, Nojun Kwak
AAAI4
2026 Knowledge Beyond Language: Bridging the Gap in Multilingual Machine Unlearning Evaluation
abstract
While LLMs are increasingly used in commercial services, they pose privacy risks such as leakage of sensitive personally identifiable information (PII).For LLMs trained on multilingual corpora, Multilingual Machine Unlearning (MMU) aims to remove information across multiple languages.However, prior MMU evaluations fail to capture such cross-linguistic distribution of information, being largely limited to direct extensions of per-language evaluation protocols.To this end, we propose two metrics to evaluate the information spread across languages: the Knowledge Separability Score (KSS) and the Knowledge Persistence Score (KPS).KSS measures the overall unlearning quality across multiple languages, while KPS more specifically aims to assess consistent removal of information among different language pairs.We evaluated various unlearning methods in the multilingual setting with these metrics and conducted comprehensive analyses.Through our investigation, we provide insights into unique phenomena exclusive to MMU and offer a new perspective on MMU evaluation.
Kyomin Hwang, Sangyeon Cho, Nojun Kwak
ACL (1)4
2026 Unlocking the Potential of Diffusion Language Models through Template Infilling
abstract
Diffusion Language Models (DLMs) have emerged as a promising alternative to Autoregressive Language Models, yet their inference strategies largely rely on prefix-based prompting inherited from the autoregressive paradigm.In this paper, we propose Template Infilling (TI), a conditioning methodology tailored for DLMs.Unlike conventional prefix prompting, TI distributes structural anchors across the target response, establishing a global template before infilling masked segments.This enables structured conditioning that leverages the bidirectional generation process of DLMs.We evaluate TI on diverse benchmarks, including mathematical reasoning, code generation, and trip planning, achieving consistent improvements of 9.40%p over baseline prompting strategies.Furthermore, TI naturally supports multi-token generation settings, providing practical speed advantages while maintaining generation quality and robustness.Overall, our results highlight a DLM-specific conditioning paradigm for structured generation, suggesting a promising direction for inference methods tailored to diffusion-based language models.
Junhoo Lee, Nojun Kwak
ACL (1)3
2026 Bi-ICE: An Inner Interpretable Framework for Image Classification via Bi-directional Interactions between Concept and Input Embeddings
abstract
Inner interpretability is a promising field aiming to uncover the internal mechanisms of AI systems through scalable, automated methods. While significant research has been conducted on large language models, limited attention has been paid to applying inner interpretability to large-scale image tasks, focusing primarily on architectural and functional levels to visualize learned concepts. In this paper, we first present a conceptual framework that supports inner interpretability and multilevel analysis for large-scale image classification tasks. Specifically, we introduce the Bi-directional Interaction between Concept and Input Embeddings (Bi-ICE) module, which facilitates interpretability across the computational, algorithmic, and implementation levels. This module enhances transparency by generating predictions based on human-understandable concepts, quantifying their contributions, and localizing them within the inputs. Finally, we showcase enhanced transparency in image classification, measuring concept contributions, and pinpointing their locations within the inputs. Our approach highlights algorithmic interpretability by demonstrating the process of concept learning and its convergence.
Jinyung Hong, Yerim Kim, Keun Hee Park, Sangyu Han, Nojun Kwak, Theodore P. Pavlic
WACV5
2026 End-to-End Multi-Entity Customization
abstract
ABSTRACT Recent advancements in text‐to‐image (T2I) models have enabled the synthesis of personalized images that align closely with user‐specified prompts, especially through the use of modifiers. However, generating multiple detailed objects with distinct modifiers in a single image remains challenging due to concept‐mixing, resulting from the difficulty of capturing interactions among text tokens. This paper proposes a modifier‐based approach to mitigate concept‐mixing by addressing the interaction among text tokens. Our method enables practical multi‐personalization while preserving the original T2I model's straightforward inference pipeline. Without structural guidance, it ensures seamless object interaction with enhanced consistency. Through a loss‐based finetuning approach, our method is adaptable to various concept‐learning algorithms, enabling plug‐and‐play functionality. Through both qualitative and quantitative evaluations, we demonstrate that our method effectively resolves concept‐mixing issues to better preserve concepts' identities and outperforms recent baselines in both quantitative and qualitative results. Our code will be publicly available.
Wonhark Park, Wonsik Shin, Junhoo Lee, Nojun Kwak
IET Image Process.5
2026 From Local to Global to Mechanistic: An iERF-Centered Unified Framework for Interpreting Vision Models
abstract
Modern vision models achieve remarkable accuracy, but explaining where evidence arises, what the model encodes, and how internal computations assemble that evidence remains fragmented. We introduce an iERF-centric framework that unifies local, global, and mechanistic interpretability around a single analysis unit: the pointwise feature vector (PFV) paired with its instance-specific Effective Receptive Field (iERF). On the local side, Sharing Ratio Decomposition (SRD) expresses each PFV as a mixture of upstream PFVs via sharing ratios and propagates iERFs to construct class-discriminative saliency maps. SRD yields high-resolution, activation-faithful explanations, is robust to targeted manipulation and noise, and remains activation-agnostic across common nonlinearities. For the global view, we introduce Concept-Anchored Feature Explanation (CAFE), which utilizes the iERF as a semantic label, grounding abstract latent vectors in verifiable pixel-level evidence. With CAFE, we address the challenge of non-localized sparse autoencoder latents-especially in Transformers, where early self-attention mixes distant context. To answer how representations are composed through depth, we propose the Interlayer Concept Graph with Interlayer Concept Attribution (ICAT), which quantifies concept-to-concept influence while isolating layer pairs; an interlayer insertion/deletion protocol identifies Integrated Gradients as the most faithful instantiation. Empirically, across ResNet50, VGG16, and ViTs, our framework outperforms baselines in both fidelity and robustness, successfully interprets dispersed SAE features, and exposes dominant concept routes in correct, misclassified, and adversarial cases. Grounded in iERFs, our approach provides a coherent, evidence-backed map from pixels to concepts to decisions.
Yerim Kim, Sangyu Han, Nojun Kwak
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Harmonizing Visual and Textual Embeddings for Zero-Shot Text-to-Image Customization
abstract
In a surge of text-to-image (T2I) models and their customization methods that generate new images of a user-provided subject, current works focus on alleviating the costs incurred by a lengthy per-subject optimization. These zero-shot customization methods encode the image of a specified subject into a visual embedding which is then utilized alongside the textual embedding for diffusion guidance. The visual embedding incorporates intrinsic information about the subject, while the textual embedding provides a new context. However, the existing methods often 1) generate images with the same pose as an input image, and 2) exhibit deterioration in the subject's identity when facing a pose variation prompt. We first pin down the problem and show that redundant pose information in the visual embedding interferes with the pose indication in the textual embedding. Conversely, the textual embedding also harms the subject's identity which is tightly entangled with the pose in the visual embedding. As a remedy, we propose text-orthogonal visual embedding which effectively harmonizes with the given textual embedding. We also adopt the visual-only embedding and inject the subject's clear features utilizing a self-attention swap. Our method is both effective and robust, offering highly flexible zero-shot generation while effectively maintaining the subject's identity.
Yeji Song, Jimyeong Kim, Wonhark Park, Wonsik Shin, Wonjong Rhee, Nojun Kwak
AAAI6
2025 Unlocking the Potential of Unlabeled Data in Semi-Supervised Domain Generalization
abstract
We address the problem of semi-supervised domain generalization (SSDG), where the distributions of train and test data differ, and only a small amount of labeled data along with a larger amount of unlabeled data are available during training. Existing SSDG methods that leverage only the unlabeled samples for which the model’s predictions are highly confident (confident-unlabeled samples), limit the full utilization of the available unlabeled data. To the best of our knowledge, we are the first to explore a method for incorporating the unconfident-unlabeled samples that were previously disregarded in SSDG setting. To this end, we propose UPCSC to utilize these unconfident-unlabeled samples in SSDG that consists of two modules: 1) Un-labeled Proxy-based Contrastive learning (UPC) module, treating unconfident-unlabeled samples as additional negative pairs and 2) Surrogate Class learning (SC) module, generating positive pairs for unconfident-unlabeled samples using their confusing class set. These modules are plug-and-play and do not require any domain labels, which can be easily integrated into existing approaches. Experiments on four widely used SSDG benchmarks demonstrate that our approach consistently improves performance when attached to baselines and outperforms competing plug-and-play methods. We also analyze the role of our method in SSDG, showing that it enhances class-level discriminability and mitigates domain gaps. The code is available at https://github.com/dongkwani/UPCSC.
Dongkwan Lee, Kyomin Hwang, Nojun Kwak
CVPR3
2025 What's Making That Sound Right Now? Video-Centric Audio-Visual Localization
abstract
Audio-Visual Localization (AVL) aims to identify sound-emitting sources within a visual scene. However, existing studies focus on image-level audio-visual associations, failing to capture temporal dynamics. Moreover, they assume simplified scenarios where sound sources are always visible and involve only a single object. To address these limitations, we propose AVATAR, a video-centric AVL benchmark that incorporates high-resolution temporal information. AVATAR introduces four distinct scenarios -- Single-sound, Mixed-sound, Multi-entity, and Off-screen -- enabling a more comprehensive evaluation of AVL models. Additionally, we present TAVLO, a novel video-centric AVL model that explicitly integrates temporal information. Experimental results show that conventional methods struggle to track temporal variations due to their reliance on global audio features and frame-level mappings. In contrast, TAVLO achieves robust and precise audio-visual alignment by leveraging high-resolution temporal modeling. Our work empirically demonstrates the importance of temporal dynamics in AVL and establishes a new standard for video-centric audio-visual localization.
Hahyeon Choi, Junhoo Lee, Nojun Kwak
ICCV3
2025 ReFlex: Text-Guided Editing of Real Images in Rectified Flow via Mid-Step Feature Extraction and Attention Adaptation
abstract
Rectified Flow text-to-image models surpass diffusion models in image quality and text alignment, but adapting ReFlow for real-image editing remains challenging. We propose a new real-image editing method for ReFlow by analyzing the intermediate representations of multimodal transformer blocks and identifying three key features. To extract these features from real images with sufficient structural preservation, we leverage mid-step latent, which is inverted only up to the mid-step. We then adapt attention during injection to improve editability and enhance alignment to the target text. Our method is training-free, requires no user-provided mask, and can be applied even without a source prompt. Extensive experiments on two benchmarks with nine baselines demonstrate its superior performance over prior methods, further validated by human evaluations confirming a strong user preference for our approach.
Jimyeong Kim, Jungwon Park, Yeji Song, Nojun Kwak, Wonjong Rhee
ICCV4
2025 Understanding Differential Transformer Unchains Pretrained Self-Attentions
abstract
Differential Transformer has recently gained significant attention for its impressive empirical performance, often attributed to its ability to perform noise canceled attention. However, precisely how differential attention achieves its empirical benefits remains poorly understood. Moreover, Differential Transformer architecture demands large-scale training from scratch, hindering utilization of open pretrained weights. In this work, we conduct an in-depth investigation of Differential Transformer, uncovering three key factors behind its success: (1) enhanced expressivity via negative attention, (2) reduced redundancy among attention heads, and (3) improved learning dynamics. Based on these findings, we propose DEX, a novel method to efficiently integrate the advantages of differential attention into pretrained language models. By reusing the softmax attention scores and adding a lightweight differential operation on the output value matrix, DEX effectively incorporates the key advantages of differential attention while remaining lightweight in both training and inference. Evaluations confirm that DEX substantially improves the pretrained LLMs across diverse benchmarks, achieving significant performance gains with minimal adaptation data (< 0.01%).
Chaerin Kong, Jiho Jang, Nojun Kwak
NeurIPS3
2025 Deep Edge Filter: Return of the Human-Crafted Layer in Deep Learning
abstract
We introduce the Deep Edge Filter, a novel approach that applies high-pass filtering to deep neural network features to improve model generalizability. Our method is motivated by our hypothesis that neural networks encode task-relevant semantic information in high-frequency components while storing domain-specific biases in low-frequency components of deep features. By subtracting low-pass filtered outputs from original features, our approach isolates generalizable representations while preserving architectural integrity. Experimental results across diverse domains such as Vision, Text, 3D, and Audio demonstrate consistent performance improvements regardless of model architecture and data modality. Analysis reveals that our method induces feature sparsification and effectively isolates high-frequency components, providing empirical validation of our core hypothesis. The code is available at \url{https://github.com/dongkwani/DeepEdgeFilter}.
Dongkwan Lee, JunHoo Lee, Nojun Kwak
NeurIPS3
2025 TPD-STR: Text Polygon Detection with Split Transformers
abstract
Regressing text in natural scenes with polygonal representations is challenging due to shape prediction difficulties. To address this, we introduce Text Polygon Detection with Split Transformers (TPD-STR), which directly regresses polygonal points. TPD-STR incorporates the Decoder Split (DS) architecture to separate polygonal point regression and textness classification, and the Positional Information Propagation (PIP) module to enhance classification. Both modules are effective and compatible with existing methods. TPD-STR achieves state-of-the-art (SOTA) performance among regression-based methods, surpassing segmentation-based methods on MSRA-TD500 without external data. Adding DS and PIP to existing models further improves performance. Experiments demonstrate the model's ability to detect text instances effectively.
Sangkuk Lee, Jeesoo Kim, Nojun Kwak
WACV4
2025 Layerwise-priority-based gradient adjustment for few-shot learning
Jangho Kim, Junhoo Lee, Donghoon Han, Nojun Kwak
Expert Syst. Appl.4
2025 Generative Head-Mounted Camera Captures for Photorealistic Avatars
abstract
Enabling photorealistic avatar animations in virtual and augmented reality (VR/AR) has been challenging because of the difficulty of obtaining ground truth state of faces. It is physically impossible to obtain synchronized images from head-mounted cameras (HMC) sensing input, which has partial observations in infrared (IR), and an array of outside-in dome cameras, which have full observations that match avatars' appearance. Prior works relying on analysis-by-synthesis methods could generate accurate ground truth, but suffer from imperfect disentanglement between expression and style in their personalized training. The reliance of extensive paired captures (HMC and dome) for the same subject makes it operationally expensive to collect large-scale datasets, which cannot be reused for different HMC viewpoints and lighting. In this work, we propose a novel generative approach, Generative HMC (GenHMC), that leverages large unpaired HMC captures , which are much easier to collect, to directly generate high-quality synthetic HMC images given any conditioning avatar state from dome captures. We show that our method is able to properly disentangle the input conditioning signal that specifies facial expression and viewpoint, from facial appearance, leading to more accurate ground truth. Furthermore, our method can generalize to unseen identities, removing the reliance on the paired captures. We demonstrate these breakthroughs by both evaluating synthetic HMC images and universal face encoders trained from these new HMC-avatar correspondences, which achieve better data efficiency and state-of-the-art accuracy.
Shaojie Bai, Seunghyeon Seo, Chenghui Li, Owen Wang, Te-Li Wang, Tianyang Ma, Jason M. Saragih, Shih-En Wei, Nojun Kwak, Hyung Jun(John) Kim
ACM Trans. Graph.10
2024 Any-Way Meta Learning
abstract
Although meta-learning seems promising performance in the realm of rapid adaptability, it is constrained by fixed cardinality. When faced with tasks of varying cardinalities that were unseen during training, the model lacks its ability. In this paper, we address and resolve this challenge by harnessing `label equivalence' emerged from stochastic numeric label assignments during episodic task sampling. Questioning what defines ``true" meta-learning, we introduce the ``any-way" learning paradigm, an innovative model training approach that liberates model from fixed cardinality constraints. Surprisingly, this model not only matches but often outperforms traditional fixed-way models in terms of performance, convergence speed, and stability. This disrupts established notions about domain generalization. Furthermore, we argue that the inherent label equivalence naturally lacks semantic information. To bridge this semantic information gap arising from label equivalence, we further propose a mechanism for infusing semantic class information into the model. This would enhance the model's comprehension and functionality. Experiments conducted on renowned architectures like MAML and ProtoNet affirm the effectiveness of our method.
Junhoo Lee, Yerim Kim, Nojun Kwak
AAAI4
2024 A Revisit to the Decoder for Camouflaged Object Detection
Seung Woo Ko 0002, Joopyo Hong, Suyoung Kim, Seungjai Bang, Sungzoon Cho, Nojun Kwak, Hyung-Sin Kim, Joonseok Lee
BMVC6
2024 What, How, and When Should Object Detectors Update in Continually Changing Test Domains?
abstract
It is a well-known fact that the performance of deep learning models deteriorates when they encounter a distribution shift at test time. Test-time adaptation (TTA) algorithms have been proposed to adapt the model online while inferring test data. However, existing research predominantly focuses on classification tasks through the optimization of batch normalization layers or classification heads, but this approach limits its applicability to various model architectures like Transformers and makes it challenging to apply to other tasks, such as object detection. In this paper, we propose a novel online adaption approach for object detection in continually changing test domains, considering which part of the model to update, how to update it, and when to perform the update. By introducing architecture-agnostic and lightweight adaptor modules and only updating these while leaving the pre-trained backbone unchanged, we can rapidly adapt to new test domains in an efficient way and prevent catastrophic forgetting. Furthermore, we present a practical and straightforward class-wise feature aligning method for object detection to resolve domain shifts. Additionally, we enhance efficiency by determining when the model is sufficiently adapted or when additional adaptation is needed due to changes in the test distribution. Our approach surpasses baselines on widely used benchmarks, achieving improvements of up to 4.9%p and 7.9%p in mAP for COCO → COCO-corrupted and SHIFT, respectively, while maintaining about 20 FPS or higher. The implementation code is available at https://github.com/natureyoo/ContinualTTA_ObjectDetection.
Jayeon Yoo, Dongkwan Lee, Inseop Chung, Nojun Kwak
CVPR5
2024 SAVE: Protagonist Diversification with Structure Agnostic Video Editing
Yeji Song, Wonsik Shin, Junsoo Lee 0002, Jeesoo Kim, Nojun Kwak
ECCV (80)5
2024 Respect the model: Fine-grained and Robust Explanation with Sharing Ratio Decomposition
abstract
The truthfulness of existing explanation methods in authentically elucidating the underlying model's decision-making process has been questioned. Existing methods have deviated from faithfully representing the model, thus susceptible to adversarial attacks. To address this, we propose a novel eXplainable AI (XAI) method called SRD (Sharing Ratio Decomposition), which sincerely reflects the model's inference process, resulting in significantly enhanced robustness in our explanations. Different from the conventional emphasis on the neuronal level, we adopt a vector perspective to consider the intricate nonlinear interactions between filters. We also introduce an interesting observation termed Activation-Pattern-Only Prediction (APOP), letting us emphasize the importance of inactive neurons and redefine relevance encapsulating all relevant information including both active and inactive neurons. Our method, SRD, allows for the recursive decomposition of a Pointwise Feature Vector (PFV), providing a high-resolution Effective Receptive Field (ERF) at any layer.
Sangyu Han, Yerim Kim, Nojun Kwak
ICLR3
2024 Advancing Beyond Identification: Multi-bit Watermark for Large Language Models
abstract
KiYoon Yoo, Wonhyuk Ahn, Nojun Kwak. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
KiYoon Yoo, Wonhyuk Ahn, Nojun Kwak
NAACL-HLT3
2024 Deep Support Vectors
abstract
Deep learning has achieved tremendous success. However, unlike SVMs, which provide direct decision criteria and can be trained with a small dataset, it still has significant weaknesses due to its requirement for massive datasets during training and the black-box characteristics on decision criteria. This paper addresses these issues by identifying support vectors in deep learning models. To this end, we propose the DeepKKT condition, an adaptation of the traditional Karush-Kuhn-Tucker (KKT) condition for deep learning models, and confirm that generated Deep Support Vectors (DSVs) using this condition exhibit properties similar to traditional support vectors. This allows us to apply our method to few-shot dataset distillation problems and alleviate the black-box characteristics of deep learning models. Additionally, we demonstrate that the DeepKKT condition can transform conventional classification models into generative models with high fidelity, particularly as latent generation models using class labels as latent variables. We validate the effectiveness of DSVs using common datasets (ImageNet, CIFAR10 and CIFAR100) on the general architectures (ResNet and ConvNet), proving their practical applicability.
JunHoo Lee, Kyomin Hwang, Nojun Kwak
NeurIPS4
2024 Fast Sun-aligned Outdoor Scene Relighting based on TensoRF
abstract
In this work, we introduce our method of outdoor scene relighting for Neural Radiance Fields (NeRF) named Sun-aligned Relighting TensoRF (SR-TensoRF). SR-TensoRF offers a lightweight and rapid pipeline aligned with the sun, thereby achieving a simplified workflow that eliminates the need for environment maps. Our sun-alignment strategy is motivated by the insight that shadows, unlike viewpoint-dependent albedo, are determined by light direction. We directly use the sun direction as an input during shadow generation, simplifying the requirements of the inference process significantly. Moreover, SR-TensoRF leverages the training efficiency of TensoRF by incorporating our proposed cubemap concept, resulting in notable acceleration in both training and rendering processes compared to existing methods.
Yeonjin Chang, Yerim Kim, Seunghyeon Seo, Jung Yi, Nojun Kwak
WACV5
2024 Discriminative subspace learning using generalized mean
abstract
Linear discriminant analysis (LDA) is one of the most popular methods to extract discriminative features because it is simple and powerful. However, LDA fails to learn a discriminative subspace in some cases. This study deals with a problem of LDA, the so-called class separation (CS) problem, which means that some classes located close to each other in the original input space tend to overlap in a learned subspace. This problem can also happen in a heteroscedastic extension of LDA, the oriented discriminant analysis (ODA). To alleviate the problem, we propose two methods to maximize the generalized mean instead of the arithmetic mean in the objective functions. Experimental results show that the proposed methods can obtain better discriminative subspaces than LDA, ODA, and other alternatives designed to solve the CS problem.
Jiyong Oh, Nojun Kwak
Signal Process.2
2023 Unifying Vision-Language Representation Space with Single-Tower Transformer
abstract
Contrastive learning is a form of distance learning that aims to learn invariant features from two related representations. In this work, we explore the hypothesis that an image and caption can be regarded as two different views of the underlying mutual information, and train a model to learn a unified vision-language representation space that encodes both modalities at once in a modality-agnostic manner. We first identify difficulties in learning a one-tower model for vision-language pretraining (VLP), and propose One Representation (OneR) as a simple yet effective framework for our goal. We discover intriguing properties that distinguish OneR from the previous works that have modality-specific representation spaces such as zero-shot localization, text-guided visual reasoning and multi-modal retrieval, and present analyses to provide insights into this new form of multi-modal representation learning. Thorough evaluations demonstrate the potential of a unified modality-agnostic VLP framework.
Jiho Jang, Chaerin Kong, Dong Hyeon Jeon, Seonhoon Kim, Nojun Kwak
AAAI5
2023 Robust Multi-bit Natural Language Watermarking through Invariant Features
abstract
Recent years have witnessed a proliferation of valuable original natural language contents found in subscription-based media outlets, web novel platforms, and outputs of large language models.However, these contents are susceptible to illegal piracy and potential misuse without proper security measures.This calls for a secure watermarking system to guarantee copyright protection through leakage tracing or ownership identification.To effectively combat piracy and protect copyrights, a multi-bit watermarking framework should be able to embed adequate bits of information and extract the watermarks in a robust manner despite possible corruption.In this work, we explore ways to advance both payload and robustness by following a well-known proposition from image watermarking and identify features in natural language that are invariant to minor corruption.Through a systematic analysis of the possible sources of errors, we further propose a corruption-resistant infill model.Our full method improves upon the previous work on robustness by +16.8% point on average on four datasets, three corruption types, and two corruption ratios. 1
KiYoon Yoo, Wonhyuk Ahn, Jiho Jang, Nojun Kwak
ACL (1)4
2023 MixNeRF: Modeling a Ray with Mixture Density for Novel View Synthesis from Sparse Inputs
abstract
Neural Radiance Field (NeRF) has broken new ground in the novel view synthesis due to its simple concept and state-of-the-art quality. However, it suffers from severe performance degradation unless trained with a dense set of images with different camera poses, which hinders its practical applications. Although previous methods addressing this problem achieved promising results, they relied heavily on the additional training resources, which goes against the philosophy of sparse-input novel-view synthesis pursuing the training efficiency. In this work, we propose MixNeRF, an effective training strategy for novel view synthesis from sparse inputs by modeling a ray with a mixture density model. Our MixNeRF estimates the joint distribution of RGB colors along the ray samples by modeling it with mixture of distributions. We also propose a new task of ray depth estimation as a useful training objective, which is highly correlated with 3D scene geometry. Moreover, we remodel the colors with regenerated blending weights based on the estimated ray depth and further improves the robustness for colors and viewpoints. Our MixNeRF outperforms other state-of-the-art methods in various standard benchmarks with superior efficiency of training and inference.
Seunghyeon Seo, Donghoon Han, Yeonjin Chang, Nojun Kwak
CVPR4
2023 Semantics-Guided Object Removal for Facial Images: with Broad Applicability and Robust Style Preservation
abstract
Object removal and image inpainting in facial images is a task in which objects that occlude a facial image are specifically targeted, removed, and replaced by a properly reconstructed facial image. Two different approaches utilize U-net-based generator and modulated approach, and they respectively have been widely endorsed but notwithstanding each method’s disadvantages of low generative capability and low reconstruction power. Here, we propose a Semantics-Guided Inpainting Network (SGIN), which is the invention of a desirable trade-off between those two methods that can be applied to any form of occluding mask while maintaining a consistent style and preserving high-fidelity details of the original image. By using the guidance of a semantic map, our model is capable of manipulating facial features and styles which grants direction to the one-to-many problem for further practicability.
Jookyung Song, Yeonjin Chang, Seonguk Park, Nojun Kwak
ICASSP4
2023 FlipNeRF: Flipped Reflection Rays for Few-shot Novel View Synthesis
abstract
Neural Radiance Field (NeRF) has been a mainstream in novel view synthesis with its remarkable quality of rendered images and simple architecture. Although NeRF has been developed in various directions improving continuously its performance, the necessity of a dense set of multi-view images still exists as a stumbling block to progress for practical application. In this work, we propose FlipNeRF, a novel regularization method for few-shot novel view synthesis by utilizing our proposed flipped reflection rays. The flipped reflection rays are explicitly derived from the input ray directions and estimated normal vectors, and play a role of effective additional training rays while enabling to estimate more accurate surface normals and learn the 3D geometry effectively. Since the surface normal and the scene depth are both derived from the estimated densities along a ray, the accurate surface normal leads to more exact depth estimation, which is a key factor for few-shot novel view synthesis. Furthermore, with our proposed Uncertainty-aware Emptiness Loss and Bottleneck Feature Consistency Loss, FlipNeRF is able to estimate more reliable outputs with reducing floating artifacts effectively across the different scene structures, and enhance the feature-level consistency between the pair of the rays cast toward the photo-consistent pixels without any additional feature extractor, respectively. Our FlipNeRF achieves the SOTA performance on the multiple benchmarks across all the scenarios.
Seunghyeon Seo, Yeonjin Chang, Nojun Kwak
ICCV3
2023 End-to-End Multi-Object Detection with a Regularized Mixture Model
abstract
Recent end-to-end multi-object detectors simplify the inference pipeline by removing hand-crafted processes such as non-maximum suppression (NMS). However, during training, they still heavily rely on heuristics and hand-crafted processes which deteriorate the reliability of the predicted confidence score. In this paper, we propose a novel framework to train an end-to-end multi-object detector consisting of only two terms: negative log-likelihood (NLL) and a regularization term. In doing so, the multi-object detection problem is treated as density estimation of the ground truth bounding boxes utilizing a regularized mixture density model. The proposed end-to-end multi-object Detection with a Regularized Mixture Model (D-RMM) is trained by minimizing the NLL with the proposed regularization term, maximum component maximization (MCM) loss, preventing duplicate predictions. Our method reduces the heuristics of the training process and improves the reliability of the predicted confidence score. Moreover, our D-RMM outperforms the previous end-to-end detectors on MS COCO dataset. Code is available at https://github.com/lhj815/D-RMM.
Jaeyoung Yoo, Hojun Lee 0002, Seunghyeon Seo, Inseop Chung, Nojun Kwak
ICML5
2023 Finding Efficient Pruned Network via Refined Gradients for Pruned Weights
abstract
With the growth of deep neural networks (DNN), the number of DNN parameters has drastically increased. This makes DNN models hard to be deployed on resource-limited embedded systems. To alleviate this problem, dynamic pruning methods have emerged, which try to find diverse sparsity patterns during training by utilizing Straight-Through-Estimator (STE) to approximate gradients of pruned weights. STE can help the pruned weights revive in the process of finding dynamic sparsity patterns. However, using these coarse gradients causes training instability and performance degradation owing to the unreliable gradient signal of the STE approximation. In this work, to tackle this issue, we introduce refined gradients to update the pruned weights by forming dual forwarding paths from two sets (pruned and unpruned) of weights. We propose a novel Dynamic Collective Intelligence Learning (DCIL) which makes use of the learning synergy between the collective intelligence of both weight sets. We verify the usefulness of the refined gradients by showing enhancements in the training stability and the model performance on the CIFAR and ImageNet datasets. DCIL outperforms various previously proposed pruning schemes including other dynamic pruning methods with enhanced stability during training. The code is provided in Github.
Jangho Kim, Jayeon Yoo, Yeji Song, KiYoon Yoo, Nojun Kwak
ACM Multimedia5
2023 SHOT: Suppressing the Hessian along the Optimization Trajectory for Gradient-Based Meta-Learning
abstract
In this paper, we hypothesize that gradient-based meta-learning (GBML) implicitly suppresses the Hessian along the optimization trajectory in the inner loop. Based on this hypothesis, we introduce an algorithm called SHOT (Suppressing the Hessian along the Optimization Trajectory) that minimizes the distance between the parameters of the target and reference models to suppress the Hessian in the inner loop. Despite dealing with high-order terms, SHOT does not increase the computational complexity of the baseline model much. It is agnostic to both the algorithm and architecture used in GBML, making it highly versatile and applicable to any GBML baseline. To validate the effectiveness of SHOT, we conduct empirical tests on standard few-shot learning tasks and qualitatively analyze its dynamics. We confirm our hypothesis empirically and demonstrate that SHOT outperforms the corresponding baseline.
Junhoo Lee, Jayeon Yoo, Nojun Kwak
NeurIPS3
2023 MDPose: real-time multi-person pose estimation via mixture density model
abstract
One of the major challenges in multi-person pose estimation is instance-aware keypoint estimation. Previous methods address this problem by leveraging an off-the-shelf detector, heuristic post-grouping process or explicit instance identification process, hindering further improvements in the inference speed which is an important factor for practical applications. From the statistical point of view, those additional processes for identifying instances are necessary to bypass learning the high-dimensional joint distribution of human keypoints, which is a critical factor for another major challenge, the occlusion scenario. In this work, we propose a novel framework of single-stage instance-aware pose estimation by modeling the joint distribution of human keypoints with a mixture density model, termed as MDPose. Our MDPose estimates the distribution of human keypoints’ coordinates using a mixture density model with an instance-aware keypoint head consisting simply of 8 convolutional layers. It is trained by minimizing the negative log-likelihood of the ground truth keypoints. Also, we propose a simple yet effective training strategy, Random Keypoint Grouping (RKG), which significantly alleviates the underflow problem leading to successful learning of relations between keypoints. On OCHuman dataset, which consists of images with highly occluded people, our MDPose achieves state-of-the-art performance by successfully learning the high-dimensional joint distribution of human keypoints. Furthermore, our MDPose shows significant improvement in inference speed with a competitive accuracy on MS COCO, a widely-used human keypoint dataset, thanks to the proposed much simpler single-stage pipeline.
Seunghyeon Seo, Jaeyoung Yoo, Jihye Hwang, Nojun Kwak
UAI4
2023 Self-Distilled Self-supervised Representation Learning
abstract
State-of-the-art frameworks in self-supervised learning have recently shown that fully utilizing transformer-based models can lead to performance boost compared to conventional CNN models. Striving to maximize the mutual information of two views of an image, existing works apply a contrastive loss to the final representations. Motivated by self-distillation in the supervised regime, we further exploit this by allowing the intermediate representations to learn from the final layer via the contrastive loss. Through self-distillation, the intermediate layers are better suited for instance discrimination, making the performance of an early-exited sub-network not much degraded from that of the full network. This renders the pretext task easier also for the final layer, leading to better representations. Our method, Self-Distilled Self-Supervised Learning (SDSSL), outperforms competitive baselines (SimCLR, BYOL and MoCo v3) using ViT on various tasks and datasets. In the linear evaluation and k-NN protocol, SDSSL not only leads to superior performance in the final layers, but also in most of the lower layers. Furthermore, qualitative and quantitative analyses show how representations are formed more effectively along the transformer layers. Code is available at https://github.com/hagiss/SDSSL.
Jiho Jang, Seonhoon Kim, KiYoon Yoo, Chaerin Kong, Jangho Kim, Nojun Kwak
WACV6
2023 Leveraging Off-the-shelf Diffusion Model for Multi-attribute Fashion Image Manipulation
abstract
Fashion attribute editing is a task that aims to convert the semantic attributes of a given fashion image while preserving the irrelevant regions. Previous works typically employ conditional GANs where the generator explicitly learns the target attributes and directly execute the conversion. These approaches, however, are neither scalable nor generic as they operate only with few limited attributes and a separate generator is required for each dataset or attribute set. Inspired by the recent advancement of diffusion models, we explore the classifier-guided diffusion that leverages the off-the-shelf diffusion model pretrained on general visual semantics such as Imagenet. In order to achieve a generic editing pipeline, we pose this as multi-attribute image manipulation task, where the attribute ranges from item category, fabric, pattern to collar and neckline. We empirically show that conventional methods fail in our challenging setting, and study efficient adaptation scheme that involves recently introduced attention-pooling technique to obtain a multi-attribute classifier guidance. Based on this, we present a mask-free fashion attribute editing framework that leverages the classifier logits and the cross-attention map for manipulation. We empirically demonstrate that our framework achieves convincing sample quality and attribute alignments.
Chaerin Kong, Dong Hyeon Jeon, Ohjoon Kwon, Nojun Kwak
WACV4
2022 Towards Efficient Neural Scene Graphs by Learning Consistency Fields
Yeji Song, Chaerin Kong, Seoyoung Lee 0001, Nojun Kwak, Joonseok Lee
BMVC4
2022 Imposing Consistency for Optical Flow Estimation
abstract
Imposing consistency through proxy tasks has been shown to enhance data-driven learning and enable self-supervision in various tasks. This paper introduces novel and effective consistency strategies for optical flow estimation, a problem where labels from real-world data are very challenging to derive. More specifically, we propose occlusion consistency and zero forcing in the forms of self-supervised learning and transformation consistency in the form of semi-supervised learning. We apply these consistency techniques in a way that the network model learns to describe pixel-level motions better while requiring no additional annotations. We demonstrate that our consistency strategies applied to a strong baseline network model using the original datasets and labels provide further improvements, attaining the state-of-the-art results on the KITTI-2015 scene flow benchmark in the non-stereo category. Our method achieves the best foreground accuracy (4.33% in Fl-all) over both the stereo and non-stereo categories, even though using only monocular image inputs.
Jisoo Jeong, Jamie Menjay Lin, Fatih Porikli, Nojun Kwak
CVPR4
2022 MUM: Mix Image Tiles and UnMix Feature Tiles for Semi-Supervised Object Detection
abstract
Many recent semi-supervised learning (SSL) studies build teacher-student architecture and train the student net-work by the generated supervisory signal from the teacher. Data augmentation strategy plays a significant role in the SSL framework since it is hard to create a weak-strong aug-mented input pair without losing label information. Espe-cially when extending SSL to semi-supervised object de-tection (SSOD), many strong augmentation methodologies related to image geometry and interpolation-regularization are hard to utilize since they possibly hurt the location information of the bounding box in the object detection task. To address this, we introduce a simple yet effective data augmentation method, Mix/UnMix (MUM), which un-mixes feature tiles for the mixed image tiles for the SSOD framework. Our proposed method makes mixed input image tiles and reconstructs them in the feature space. Thus, MUM can enjoy the interpolation-regularization effect from non-interpolated pseudo-labels and successfully generate a meaningful weak-strong pair. Furthermore, MUM can be easily equipped on top of various SSOD methods. Exten-sive experiments on MS-COCO and PASCAL VOC datasets demonstrate the superiority of MUM by consistently im-proving the mAP performance over the baseline in all the tested SSOD benchmark protocols. The code is released at https.//github.com/JongMokKim/mix-unmix.
Jongmok Kim, Jooyoung Jang, Seunghyeon Seo, Jisoo Jeong, Jongkeun Na, Nojun Kwak
CVPR6
2022 MatteFormer: Transformer-Based Image Matting via Prior-Tokens
abstract
In this paper, we propose a transformer-based image matting model called MatteFormer, which takes full advantage of trimap information in the transformer block. Our method first introduces a prior-token which is a global representation of each trimap region (e.g. foreground, background and unknown). These prior-tokens are used as global priors and participate in the self-attention mechanism of each block. Each stage of the encoder is composed of PAST (Prior-Attentive Swin Transformer) block, which is based on the Swin Transformer block, but differs in a couple of aspects: 1) It has PA-WSA (Prior-Attentive Window Self-Attention) layer, performing self-attention not only with spatial-tokens but also with prior-tokens. 2) It has prior-memory which saves prior-tokens accumulatively from the previous blocks and transfers them to the next block. We evaluate our MatteFormer on the commonly used image matting datasets: Composition-Ik and Distinctions-646. Experiment results show that our proposed method achieves state-of-the-art performance with a large margin. Our codes are available at https://github.com/webtoon/matteformer.
Gyutae Park, Sungjoon Son, Jaeyoung Yoo, Seho Kim, Nojun Kwak
CVPR5
2022 Few-Shot Image Generation with Mixup-Based Distance Learning
Chaerin Kong, Jeesoo Kim, Donghoon Han, Nojun Kwak
ECCV (15)4
2022 Unsupervised Domain Adaptation for One-Stage Object Detector Using Offsets to Bounding Box
Jayeon Yoo, Inseop Chung, Nojun Kwak
ECCV (33)3
2022 Backdoor Attacks in Federated Learning by Rare Embeddings and Gradient Ensembling
abstract
Recent advances in federated learning have demonstrated its promising capability to learn on decentralized datasets.However, a considerable amount of work has raised concerns due to the potential risks of adversaries participating in the framework to poison the global model for an adversarial purpose.This paper investigates the feasibility of model poisoning for backdoor attacks through rare word embeddings of NLP models.In text classification, less than 1% of adversary clients suffices to manipulate the model output without any drop in the performance on clean sentences.For a less complex dataset, a mere 0.1% of adversary clients is enough to poison the global model effectively.We also propose a technique specialized in the federated learning scheme called Gradient Ensemble, which enhances the backdoor performance in all our experimental settings.
KiYoon Yoo, Nojun Kwak
EMNLP2
2022 Variational On-the-Fly Personalization
abstract
With the development of deep learning (DL) technologies, the demand for DL-based services on personal devices, such as mobile phones, also increases rapidly. In this paper, we propose a novel personalization method, Variational On-the-Fly Personalization. Compared to the conventional personalization methods that require additional fine-tuning with personal data, the proposed method only requires forwarding a handful of personal data on-the-fly. Assuming even a single personal data can convey the characteristics of a target person, we develop the variational hyper-personalizer to capture the weight distribution of layers that fits the target person. In the testing phase, the hyper-personalizer estimates the model’s weights on-the-fly based on personality by forwarding only a small amount of (even a single) personal enrollment data. Hence, the proposed method can perform the personalization without any training software platform and additional cost in the edge device. In experiments, we show our approach can effectively generate reliable personalized models via forwarding (not back-propagating) a handful of samples.
Jangho Kim, Juntae Lee, Simyung Chang, Nojun Kwak
ICML4
2022 The U-Net based GLOW for Optical-Flow-Free Video Interframe Generation
abstract
Video frame interpolation is the task of creating an interframe between two adjacent frames along the time axis. So, instead of simply averaging two adjacent frames to create an intermediate image, this operation should maintain semantic continuity with the adjacent frames. Most conventional methods use optical flow, and various tools such as occlusion handling and object smoothing are indispensable. Since the use of these various tools leads to complex problems, we tried to tackle the video interframe generation problem without using problematic optical flow. To enable this, we have tried to use a deep neural network with an invertible structure, and developed an U-Net based Generative Flow which is a modified normalizing flow. In addition, we propose a learning method with a new consistency loss in the latent space to maintain semantic temporal consistency between frames. The resolution of the generated image is guaranteed to be identical to that of the original images by using an invertible network. Furthermore, as it is not a random image like the ones by generative models, our network guarantees stable outputs without flicker. Through experiments, we confirmed the feasibility of the proposed algorithm and would like to suggest the U-Net based Generative Flow as a new possibility for baseline in video frame interpolation. This paper is meaningful in that it is the new attempt to use invertible networks instead of optical flows for video interpolation.
Saem Park, Donghoon Han, Nojun Kwak
ICPRAM3
2022 Korean Language Modeling via Syntactic Guide
abstract
While pre-trained language models play a vital role in modern language processing tasks, but not every language can benefit from them. Most existing research on pre-trained language models focuses primarily on widely-used languages such as English, Chinese, and Indo-European languages. Additionally, such schemes usually require extensive computational resources alongside a large amount of data, which is infeasible for less-widely used languages. We aim to address this research niche by building a language model that understands the linguistic phenomena in the target language which can be trained with low-resources. In this paper, we discuss Korean language modeling, specifically methods for language representation and pre-training methods. With our Korean-specific language representation, we are able to build more powerful language models for Korean understanding, even with fewer resources. The paper proposes chunk-wise reconstruction of the Korean language based on a widely used transformer architecture and bidirectional language representation. We also introduce morphological features such as Part-of-Speech (PoS) into the language understanding by leveraging such information during the pre-training. Our experiment results prove that the proposed methods improve the model performance of the investigated Korean language understanding tasks.
Hyeondey Kim, Seonhoon Kim, Inho Kang, Nojun Kwak, Pascale Fung
LREC4
2022 Maximizing Cosine Similarity Between Spatial Features for Unsupervised Domain Adaptation in Semantic Segmentation
abstract
We propose a novel method that tackles the problem of unsupervised domain adaptation for semantic segmentation by maximizing the cosine similarity between the source and the target domain at the feature level. A segmentation network mainly consists of two parts, a feature extractor and a classification head. We expect that if we can make the two domains have small domain gap at the feature level, they would also have small domain discrepancy at the classification head. Our method computes a cosine similarity matrix between the source feature map and the target feature map, then we maximize the elements exceeding a threshold to guide the target features to have high similarity with the most similar source feature. Moreover, we use a class-wise source feature dictionary which stores the latest features of the source domain to prevent the unmatching problem when computing the cosine similarity matrix and be able to compare a target feature with various source features from various images. Through extensive experiments, we verify that our method gains performance on two unsupervised domain adaptation tasks (GTA5→Cityscapes and SYNTHIA→Cityscapes).
Inseop Chung, Nojun Kwak
WACV3
2022 Few-Shot Object Detection by Attending to Per-Sample-Prototype
abstract
Few-shot object detection aims to detect instances of specific categories in a query image with only a handful of support samples. Although this takes less effort than obtaining enough annotated images for supervised object detection, it results in a far inferior performance compared to the conventional object detection methods. In this paper, we propose a meta-learning-based approach that con-siders the unique characteristics of each support sample. Rather than simply averaging the information of the support samples to generate a single prototype per category, our method can better utilize the information of each support sample by treating each support sample as an individual prototype. Specifically, we introduce two types of attention mechanisms for aggregating the query and support feature maps. The first is to refine the information of few-shot samples by extracting shared information between the support samples through attention. Second, each support sample is used as a class code to leverage the information by comparing similarities between each support feature and query features. Our proposed method is complementary to the previous methods, making it easy to plug and play for further improvement. We have evaluated our method on PASCAL VOC and COCO benchmarks, and the results verify the effectiveness of our method. In particular, the advantages of our method are maximized when there is more diversity among support data.
Hojun Lee 0002, Myunggi Lee, Nojun Kwak
WACV3
2022 Dynamic Iterative Refinement for Efficient 3D Hand Pose Estimation
abstract
While hand pose estimation is a critical component of most interactive extended reality and gesture recognition systems, contemporary approaches are not optimized for computational and memory efficiency. In this paper, we propose a tiny deep neural network of which partial layers are recursively exploited for refining its previous estimations. During its iterative refinements, we employ learned gating criteria to decide whether to exit from the weight-sharing loop, allowing per-sample adaptation in our model. Our network is trained to be aware of the uncertainty in its current predictions to efficiently gate at each iteration, estimating variances after each loop for its keypoint estimates. Additionally, we investigate the effectiveness of end-to-end and progressive training protocols for our recursive structure on maximizing the model capacity. With the proposed setting, our method consistently outperforms state-of-the-art 2D/3D hand pose estimation approaches in terms of both accuracy and efficiency for widely used benchmarks.
John Yang 0001, Yash Bhalgat, Simyung Chang, Fatih Porikli, Nojun Kwak
WACV5
2021 Self-supervised Pre-training and Contrastive Representation Learning for Multiple-choice Video QA
abstract
Video Question Answering (VideoQA) requires fine-grained understanding of both video and language modalities to answer the given questions. In this paper, we propose novel training schemes for multiple-choice video question answering with a self-supervised pre-training stage and a supervised contrastive learning in the main stage as an auxiliary learning. In the self-supervised pre-training stage, we transform the original problem format of predicting the correct answer into the one that predicts the relevant question to provide a model with broader contextual inputs without any further dataset or annotation. For contrastive learning in the main stage, we add a masking noise to the input corresponding to the ground-truth answer, and consider the original input of the ground-truth answer as a positive sample, while treating the rest as negative samples. By mapping the positive sample closer to the masked input, we show that the model performance is improved. We further employ locally aligned attention to focus more effectively on the video frames that are particularly relevant to the given corresponding subtitle sentences. We evaluate our proposed model on highly competitive benchmark datasets related to multiple-choice video QA: TVQA, TVQA+, and DramaQA. Experimental results show that our model achieves state-of-the-art performance on all datasets. We also validate our approaches through further analyses.
Seonhoon Kim, Seohyeong Jeong, Eunbyul Kim, Inho Kang, Nojun Kwak
AAAI5
2021 Interpolation-Based Semi-Supervised Learning for Object Detection
abstract
Despite the data labeling cost for the object detection tasks being substantially more than that of the classification tasks, semi-supervised learning methods for object detection have not been studied much. In this paper, we propose an Interpolation-based Semi-supervised learning method for object Detection (ISD), which considers and solves the problems caused by applying conventional Interpolation Regularization (IR) directly to object detection. We divide the output of the model into two types according to the objectness scores of both original patches that are mixed in IR. Then, we apply a separate loss suitable for each type in an unsupervised manner. The proposed losses dramatically improve the performance of semi-supervised learning as well as supervised learning. In the supervised learning setting, our method improves the baseline methods by a significant margin. In the semi-supervised learning setting, our algorithm improves the performance on a benchmark dataset (PASCAL VOC and MSCOCO) in a benchmark architecture (SSD). Our code is available at https://github.com/soo89/ISD-SSD
Jisoo Jeong, Vikas Verma, Minsung Hyun, Juho Kannala, Nojun Kwak
CVPR5
2021 Learning Dynamic Network Using a Reuse Gate Function in Semi-Supervised Video Object Segmentation
abstract
Current state-of-the-art approaches for Semi-supervised Video Object Segmentation (Semi-VOS) propagates information from previous frames to generate segmentation mask for the current frame. This results in high-quality segmentation across challenging scenarios such as changes in appearance and occlusion. But it also leads to unnecessary computations for stationary or slow-moving objects where the change across frames is minimal. In this work, we exploit this observation by using temporal information to quickly identify frames with minimal change and skip the heavyweight mask generation step. To realize this efficiency, we propose a novel dynamic network that estimates change across frames and decides which path – computing a full network or reusing previous frame’s feature – to choose depending on the expected similarity. Experimental results show that our approach significantly improves inference speed without much accuracy degradation on challenging Semi-VOS datasets – DAVIS 16, DAVIS 17, and YouTube-VOS. Furthermore, our approach can be applied to multiple Semi-VOS methods demonstrating its generality. The code is available in https://github.com/HYOJINPARK/ReuseVOS
Hyojin Park 0001, Jayeon Yoo, Seohyeong Jeong, Ganesh Venkatesh, Nojun Kwak
CVPR5
2021 Prototype-Based Personalized Pruning
abstract
Nowadays, as edge devices such as smartphones become prevalent, there are increasing demands for personalized services. However, traditional personalization methods are not suitable for edge devices because retraining or finetuning is needed with limited personal data. Also, a full model might be too heavy for edge devices with limited resources. Unfortunately, model compression methods which can handle the model complexity issue also require the retraining phase. These multiple training phases generally need huge computational cost during on-device learning which can be a burden to edge devices. In this work, we propose a dynamic personalization method called prototype-based personalized pruning (PPP). PPP considers both ends of personalization and model efficiency. After training a network, PPP can easily prune the network with a prototype representing the characteristics of personal data and it performs well without retraining or finetuning. We verify the usefulness of PPP on a couple of tasks in computer vision and Keyword spotting.
Jangho Kim, Simyung Chang, Sungrack Yun, Nojun Kwak
ICASSP4
2021 Normalization Matters in Weakly Supervised Object Localization
abstract
Weakly-supervised object localization (WSOL) enables finding an object using a dataset without any localization information. By simply training a classification model using only image-level annotations, the feature map of the model can be utilized as a score map for localization. In spite of many WSOL methods proposing novel strategies, there has not been any de facto standard about how to normalize the class activation map (CAM). Consequently, many WSOL methods have failed to fully exploit their own capacity because of the misuse of a normalization method. In this paper, we review many existing normalization methods and point out that they should be used according to the property of the given dataset. Additionally, we propose a new normalization method which substantially enhances the performance of any CAM-based WSOL methods. Using the proposed normalization method, we provide a comprehensive evaluation over three datasets (CUB, ImageNet and OpenImages) on three different architectures and observe significant performance gains over the conventional min-max normalization method in all the evaluated cases (See Fig. 1).
Jeesoo Kim, Junsuk Choe, Sangdoo Yun, Nojun Kwak
ICCV4
2021 LFI-CAM: Learning Feature Importance for Better Visual Explanation
abstract
Class Activation Mapping (CAM) is a powerful technique used to understand the decision making of Convolutional Neural Network (CNN) in computer vision. Recently, there have been attempts not only to generate better visual explanations, but also to improve classification performance using visual explanations. However, previous works still have their own drawbacks. In this paper, we propose a novel architecture, LFI-CAM***(Learning Feature Importance Class Activation Mapping), which is trainable for image classification and visual explanation in an end-to-end manner. LFI-CAM generates attention map for visual explanation during forward propagation, and simultaneously uses attention map to improve classification performance through the attention mechanism. Feature Importance Network (FIN) focuses on learning the feature importance instead of directly learning the attention map to obtain a more reliable and consistent attention map. We confirmed that LFI-CAM is optimized not only by learning the feature importance but also by enhancing the backbone feature representation to focus more on important features of the input image. Experiments show that LFI-CAM outperforms baseline models’ accuracy on classification tasks as well as significantly improves on previous works in terms of attention map quality and stability over different hyper-parameters.
Kwang Hee Lee, Junghyun Oh, Nojun Kwak
ICCV4
2021 Training Multi-Object Detector by Estimating Bounding Box Distribution for Input Image
abstract
In multi-object detection using neural networks, the fundamental problem is, "How should the network learn a variable number of bounding boxes in different input images?". Previous methods train a multi-object detection network through a procedure that directly assigns the ground truth bounding boxes to the specific locations of the network’s output. However, this procedure makes the training of a multi-object detection network too heuristic and complicated. In this paper, we reformulate the multi-object detection task as a problem of density estimation of bounding boxes. Instead of assigning each ground truth to specific locations of network’s output, we train a network by estimating the probability density of bounding boxes in an input image using a mixture model. For this purpose, we propose a novel network for object detection called Mixture Density Object Detector (MDOD), and the corresponding objective function for the density-estimation-based training. We applied MDOD to MS COCO dataset. Our proposed method not only deals with multi-object detection problems in a new approach, but also improves detection performances through MDOD. The code is available: https://github.com/yoojy31/MDOD.
Jaeyoung Yoo, Hojun Lee 0002, Inseop Chung, Geonseok Seo, Nojun Kwak
ICCV5
2021 Vehicle Image Generation Going Well with the Surroundings
Jeesoo Kim, Jangho Kim, Jaeyoung Yoo, Nojun Kwak
ICONIP (4)5
2021 PQK: Model Compression via Pruning, Quantization, and Knowledge Distillation
abstract
As edge devices become prevalent, deploying Deep Neural Networks (DNN) on edge devices has become a critical issue. However, DNN requires a high computational resource which is rarely available for edge devices. To handle this, we propose a novel model compression method for the devices with limited computational resources, called PQK consisting of pruning, quantization, and knowledge distillation (KD) processes. Unlike traditional pruning and KD, PQK makes use of unimportant weights pruned in the pruning process to make a teacher network for training a better student network without pre-training the teacher model. PQK has two phases. Phase 1 exploits iterative pruning and quantization-aware training to make a lightweight and power-efficient model. In phase 2, we make a teacher network by adding unimportant weights unused in phase 1 to a pruned network. By using this teacher network, we train the pruned network as a student network. In doing so, we do not need a pre-trained teacher network for the KD framework because the teacher and the student networks coexist within the same network. We apply our method to the recognition model and verify the effectiveness of PQK on keyword spotting (KWS) and image recognition.
Jangho Kim, Simyung Chang, Nojun Kwak
Interspeech3
2021 Part-Aware Data Augmentation for 3D Object Detection in Point Cloud
abstract
Data augmentation has greatly contributed to improving the performance in image recognition tasks, and a lot of related studies have been conducted. However, data augmentation on 3D point cloud data has not been much explored. 3D label has more sophisticated and rich structural information than the 2D label, so it enables more diverse and effective data augmentation. In this paper, we propose part-aware data augmentation (PA-AUG) that can better utilize rich information of 3D label to enhance the performance of 3D object detectors. PA-AUG divides objects into partitions and stochastically applies five augmentation methods to each local region. It is compatible with existing point cloud data augmentation methods and can be used universally regardless of the detector’s architecture. PA-AUG has improved the performance of state-of-the-art 3D object detector for all classes of the KITTI dataset and has the equivalent effect of increasing the train data by about 2.5×. We also show that PA-AUG not only increases performance for a given dataset but also is robust to corrupted data. The code is available at https://github.com/sky77764/pa-aug.pytorch
Jaeseok Choi, Yeji Song, Nojun Kwak
IROS3
2020 Tell Me What They're Holding: Weakly-Supervised Object Detection with Transferable Knowledge from Human-Object Interaction
abstract
In this work, we introduce a novel weakly supervised object detection (WSOD) paradigm to detect objects belonging to rare classes that have not many examples using transferable knowledge from human-object interactions (HOI). While WSOD shows lower performance than full supervision, we mainly focus on HOI as the main context which can strongly supervise complex semantics in images. Therefore, we propose a novel module called RRPN (relational region proposal network) which outputs an object-localizing attention map only with human poses and action verbs. In the source domain, we fully train an object detector and the RRPN with full supervision of HOI. With transferred knowledge about localization map from the trained RRPN, a new object detector can learn unseen objects with weak verbal supervision of HOI without bounding box annotations in the target domain. Because the RRPN is designed as an add-on type, we can apply it not only to the object detection but also to other domains such as semantic segmentation. The experimental results on HICO-DET dataset show the possibility that the proposed method can be a cheap alternative for the current supervised object detection paradigm. Moreover, qualitative results demonstrate that our model can properly localize unseen objects on HICO-DET and V-COCO datasets.
Gyuejeong Lee, Jisoo Jeong, Nojun Kwak
AAAI4
2020 URNet: User-Resizable Residual Networks with Conditional Gating Module
abstract
Convolutional Neural Networks are widely used to process spatial scenes, but their computational cost is fixed and depends on the structure of the network used. There are methods to reduce the cost by compressing networks or varying its computational path dynamically according to the input image. However, since a user can not control the size of the learned model, it is difficult to respond dynamically if the amount of service requests suddenly increases. We propose User-Resizable Residual Networks (URNet), which allows users to adjust the computational cost of the network as needed during evaluation. URNet includes Conditional Gating Module (CGM) that determines the use of each residual block according to the input image and the desired cost. CGM is trained in a supervised manner using the newly proposed scale(cost) loss and its corresponding training methods. URNet can control the amount of computation and its inference path according to user's demand without degrading the accuracy significantly. In the experiments on ImageNet, URNet based on ResNet-101 maintains the accuracy of the baseline even when resizing it to approximately 80% of the original network, and demonstrates only about 1% accuracy degradation when using about 65% of the computation.
Simyung Chang, Nojun Kwak
AAAI3
2020 Feature-Level Ensemble Knowledge Distillation for Aggregating Knowledge from Multiple Networks
abstract
Knowledge Distillation (KD) aims to transfer knowledge in a teacher-student framework, by providing the predictions of the teacher network to the student network in the training stage to help the student network generalize better. It can use either a teacher with high capacity or an ensemble of multiple teachers. However, the latter is not convenient when one wants to use feature-map-based distillation methods. In this paper, we empirically show that using several non-linear transformation layer cope well with multiple-teacher setting compared to other kinds of feature-map-level distillation methods. Comprehensively, this paper proposes a versatile and powerful training algorithm named FEature-level Ensemble knowledge Distillation (FEED), which aims to transfer the ensemble knowledge using multiple teacher networks. In this study, we introduce a couple of training algorithms that transfer ensemble knowledge to the student at the feature-map-level. Among the feature-map-level distillation methods, using several non-linear transformations in parallel for transferring the knowledge of the multiple teachers helps the student find more generalized solutions. We name this method as parallel FEED, and experimental results on CIFAR-100 and ImageNet show that our method has clear performance enhancements, without introducing any additional parameters or computations at test time. We also show the experimental results of sequentially feeding teacher's information to the student, hence the name sequential FEED, and discuss the lessons obtained. Additionally, the empirical results on measuring the reconstruction errors at the feature map give hints for the enhancements.
Seonguk Park, Nojun Kwak
ECAI2
2020 Procrustean Regression Networks: Learning 3D Structure of Non-rigid Objects from 2D Annotations
Sungheon Park, Minsik Lee 0001, Nojun Kwak
ECCV (29)3
2020 SeqHAND: RGB-Sequence-Based 3D Hand Pose and Shape Estimation
John Yang 0001, Hyung Jin Chang, Seungeui Lee, Nojun Kwak
ECCV (12)4
2020 Kl-Divergence-Based Region Proposal Network For Object Detection
abstract
The learning of the region proposal in object detection using the deep neural networks (DNN) is divided into two tasks: binary classification and bounding box regression task. However, traditional RPN (Region Proposal Network) defines these two tasks as different problems, and they are trained independently. In this paper, we propose a new region proposal learning method that considers the bounding box offset's uncertainty in the objectness score. Our method redefines RPN to a problem of minimizing the KL-divergence, difference between the two probability distributions. We applied KLRPN, which performs region proposal using KL-Divergence, to the existing two-stage object detection framework and showed that it can improve the performance of the existing method. Experiments show that it achieves 2.6% and 2.0% AP improvements on MS COCO test-dev in Faster R-CNN with VGG-16 and R-FCN with ResNet-101 backbone, respectively.
Geonseok Seo, Jaeyoung Yoo, Jaeseok Choi, Nojun Kwak
ICIP4
2020 Unpriortized Autoencoder For Image Generation
abstract
In this paper, we treat the image generation task using an autoencoder, a representative latent model. Unlike many studies regularizing the latent variable’s distribution by assuming a manually specified prior, we approach the image generation task using an autoencoder by directly estimating the latent distribution. To this end, we introduce ‘latent density estimator’ which captures latent distribution explicitly and propose its structure. Through experiments, we show that our generative model generates images with the improved visual quality compared to previous autoencoder-based generative models.
Jaeyoung Yoo, Hojun Lee 0002, Nojun Kwak
ICIP3
2020 Feature-map-level Online Adversarial Knowledge Distillation
abstract
Feature maps contain rich information about image intensity and spatial correlation. However, previous online knowledge distillation methods only utilize the class probabilities. Thus in this paper, we propose an online knowledge distillation method that transfers not only the knowledge of the class probabilities but also that of the feature map using the adversarial training framework. We train multiple networks simultaneously by employing discriminators to distinguish the feature map distributions of different networks. Each network has its corresponding discriminator which discriminates the feature map from its own as fake while classifying that of the other network as real. By training a network to fool the corresponding discriminator, it can learn the other network’s feature map distribution. We show that our method performs better than the conventional direct alignment method such as L1 and is more suitable for online distillation. Also, we propose a novel cyclic learning scheme for training more than two networks together. We have applied our method to various network architectures on the classification task and discovered a significant improvement of performance especially in the case of training a pair of a small network and a large one.
Inseop Chung, Seonguk Park, Jangho Kim, Nojun Kwak
ICML4
2020 Feature Fusion for Online Mutual Knowledge Distillation
abstract
We propose a learning framework named Feature Fusion Learning (FFL) that efficiently trains a powerful classifier through a fusion module which combines the feature maps generated from parallel neural networks and generates meaningful feature maps. Specifically, we train a number of parallel neural networks as sub-networks, then we combine the feature maps from each sub-network using a fusion module to create a more meaningful feature map. The fused feature map is passed into the fused classifier for overall classification. Unlike existing feature fusion methods, in our framework, an ensemble of sub-network classifiers transfers its knowledge to the fused classifier and then the fused classifier delivers its knowledge back to each subnetwork, mutually teaching one another in an online-knowledge distillation manner. This mutually teaching system not only improves the performance of the fused classifier but also obtains performance gain in each sub-network. Moreover, our model is more beneficial than other alternative methods because different types of network can be used for each sub-network. We have performed a variety of experiments on multiple datasets such as CIFAR-10, CIFAR-100 and ImageNet and proved that our method is more effective than other alternative methods in terms of performances of both sub-networks and the fused classifier, and the aspect of generating meaningful feature maps. The code is available at this link1.
Jangho Kim, Minsung Hyun, Inseop Chung, Nojun Kwak
ICPR4
2020 Self-Training using Selection Network for Semi-supervised Learning
abstract
Semi-supervised learning (SSL) is a study that efficiently exploits a large amount of unlabeled data to improve performance in conditions of limited labeled data. Most of the conventional SSL methods assume that the classes of unlabeled data are included in the set of classes of labeled data. In addition, these methods do not sort out useless unlabeled samples and use all the unlabeled data for learning, which is not suitable for realistic situations. In this paper, we propose an SSL method called selective self-training (SST), which selectively decides whether to include each unlabeled sample in the training process. It is designed to be applied to a more real situation where classes of unlabeled data are different from the ones of the labeled data. For the conventional SSL problems which deal with data where both the labeled and unlabeled samples share the same class categories, the proposed method not only performs comparable to other conventional SSL algorithms but also can be combined with other SSL algorithms. While the conventional methods cannot be applied to the new SSL problems, our method does not show any performance degradation even if the classes of unlabeled data are different from those of the labeled data.
Jisoo Jeong, Seungeui Lee, Nojun Kwak
ICPRAM3
2020 Position-based Scaled Gradient for Model Quantization and Pruning
abstract
We propose the position-based scaled gradient (PSG) that scales the gradient depending on the position of a weight vector to make it more compression-friendly. First, we theoretically show that applying PSG to the standard gradient descent (GD), which is called PSGD, is equivalent to the GD in the warped weight space, a space made by warping the original weight space via an appropriately designed invertible function. Second, we empirically show that PSG acting as a regularizer to a weight vector is favorable for model compression domains such as quantization and pruning. PSG reduces the gap between the weight distributions of a full-precision model and its compressed counterpart. This enables the versatile deployment of a model either as an uncompressed mode or as a compressed mode depending on the availability of resources. The experimental results on CIFAR-10/100 and ImageNet datasets show the effectiveness of the proposed PSG in both domains of pruning and quantization even for extremely low bits. The code is released in Github.
Jangho Kim, KiYoon Yoo, Nojun Kwak
NeurIPS3
2020 SINet: Extreme Lightweight Portrait Segmentation Networks with Spatial Squeeze Modules and Information Blocking Decoder
abstract
Designing a lightweight and robust portrait segmentation algorithm is an important task for a wide range of face applications. However, the problem has been considered as a subset of the object segmentation and less handled in this field. Obviously, portrait segmentation has its unique requirements. First, because the portrait segmentation is performed in the middle of a whole process, it requires extremely lightweight models. Second, there has not been any public datasets in this domain that contain a sufficient number of images. To solve the first problem, we introduce the new extremely lightweight portrait segmentation model SINet, containing an information blocking decoder and spatial squeeze modules. The information blocking decoder uses confidence estimation to recover local spatial information without spoiling global consistency. The spatial squeeze module uses multiple receptive fields to cope with various sizes of consistency. To tackle the second problem, we propose a simple method to create additional portrait segmentation data, which can improve accuracy. In our qualitative and quantitative analysis on the EG1800 dataset, we show that our method outperforms various existing lightweight models. Our method reduces the number of parameters from 2.1M to 86.9K (around 95.9% reduction), while maintaining the accuracy under an 1% margin from the state-of-the-art method. We also show our model is successfully executed on a real mobile device with 100.6 FPS. In addition, we demonstrate that our method can be used for general semantic segmentation on the Cityscapes dataset. The code and dataset are available in https://github.com/HYOJINPARK/ExtPortraitSeg.
Hyojin Park 0001, Lars Lowe Sjösund, Young Joon Yoo, Nicolas Monet, Jihwan Bang, Nojun Kwak
WACV6
2020 Nonparametric Estimation of Probabilistic Membership for Subspace Clustering
abstract
Recent advances of subspace clustering have provided a new way of constructing affinity matrices for clustering. Unlike the kernel-based subspace clustering, which needs tedious tuning among infinitely many kernel candidates, the self-expressive models derived from linear subspace assumptions in modern subspace clustering methods are rigorously combined with sparse or low-rank optimization theory to yield an affinity matrix as a solution of an optimization problem. Despite this nice theoretical aspect, the affinity matrices of modern subspace clustering have quite different meanings from the traditional ones, and even though the affinity matrices are expected to have a rough block-diagonal structure, it is unclear whether these are good enough to apply spectral clustering. In fact, most of the subspace clustering methods perform some sort of affinity value rearrangement to apply spectral clustering, but its adequacy for the spectral clustering is not clear; even though the spectral clustering step can also have a critical impact on the overall performance. To resolve this issue, in this paper, we provide a nonparametric method to estimate the probabilistic cluster membership from these affinity matrices, which we can directly find the clusters from. The likelihood for an affinity matrix is defined nonparametrically based on histograms given the probabilistic membership, which is defined as a combination of probability simplices, and an additional prior probability is defined to regularize the membership as a Bernoulli distribution. Solving this maximum a posteriori problem replaces the spectral clustering procedure, and the final discrete cluster membership can be calculated by selecting the clusters with maximum probabilities. The proposed method provides state-of-the-art performance for well-known subspace clustering methods on popular benchmark databases.
Hyeogjin Lee, Minsik Lee 0001, Nojun Kwak
IEEE Trans. Cybern.4
2019 Semantic Sentence Matching with Densely-Connected Recurrent and Co-Attentive Information
abstract
Sentence matching is widely used in various natural language tasks such as natural language inference, paraphrase identification, and question answering. For these tasks, understanding logical and semantic relationship between two sentences is required but it is yet challenging. Although attention mechanism is useful to capture the semantic relationship and to properly align the elements of two sentences, previous methods of attention mechanism simply use a summation operation which does not retain original features enough. Inspired by DenseNet, a densely connected convolutional network, we propose a densely-connected co-attentive recurrent neural network, each layer of which uses concatenated information of attentive features as well as hidden features of all the preceding recurrent layers. It enables preserving the original and the co-attentive feature information from the bottommost word embedding layer to the uppermost recurrent layer. To alleviate the problem of an ever-increasing size of feature vectors due to dense concatenation operations, we also propose to use an autoencoder after dense concatenation. We evaluate our proposed architecture on highly competitive benchmark datasets related to sentence matching. Experimental results show that our architecture, which retains recurrent and attentive features, achieves state-of-the-art performances for most of the tasks.
Seonhoon Kim, Inho Kang, Nojun Kwak
AAAI3
2019 Textbook Question Answering with Multi-modal Context Graph Understanding and Self-supervised Open-set Comprehension
abstract
In this work, we introduce a novel algorithm for solving the textbook question answering (TQA) task which describes more realistic QA problems compared to other recent tasks.We mainly focus on two related issues with analysis of the TQA dataset.First, solving the TQA problems requires to comprehend multimodal contexts in complicated input data.To tackle this issue of extracting knowledge features from long text lessons and merging them with visual features, we establish a context graph from texts and images, and propose a new module f-GCN based on graph convolutional networks (GCN).Second, scientific terms are not spread over the chapters and subjects are split in the TQA dataset.To overcome this so called 'out-of-domain' issue, before learning QA problems, we introduce a novel self-supervised open-set learning process without any annotations.The experimental results show that our model significantly outperforms prior state-of-the-art methods.Moreover, ablation studies validate that both methods of incorporating f-GCN for extracting knowledge from multi-modal contexts and our newly proposed self-supervised learning process are effective for TQA problems. Nucleic acid classification fuction of nucleic acidDNA stores genetic information in the cells of all living things.It contains the genetic code.This is the code that instructs cells how to make proteins.
Seonhoon Kim, Nojun Kwak
ACL (1)3
2019 Towards Governing Agent's Efficacy: Action-Conditional $β$-VAE for Deep Transparent Reinforcement Learning
abstract
We tackle the blackbox issue of deep neural networks in the settings of reinforcement learning (RL) where neural agents learn towards maximizing reward gains in an uncontrollable way. Such learning approach is risky when the interacting environment includes an expanse of state space because it is then almost impossible to foresee all unwanted outcomes and penalize them with negative rewards beforehand. We propose Action-conditional $\beta$-VAE (AC-$\beta$-VAE) that allows succinct mappings of action-dependent factors in desirable dimensions of latent representations while disentangling environmental factors. Our proposed method tackles the blackbox issue by encouraging an RL policy network to learn interpretable latent features by distinguits influenshing ices from uncontrollable environmental factors, which closely resembles the way humans understand their scenes. Our experimental results show that the learned latent factors not only are interpretable, but also enable modeling the distribution of entire visited state-action space. We have experimented that this characteristic of the proposed structure can lead to ex post facto governance for desired behaviors of RL agents.
John Yang 0001, Gyuejeong Lee, Simyung Chang, Nojun Kwak
ACML4
2019 Sym-Parameterized Dynamic Inference for Mixed-Domain Image Translation
abstract
Recent advances in image-to-image translation have led to some ways to generate multiple domain images through a single network. However, there is still a limit in creating an image of a target domain without a dataset on it. We propose a method to expand the concept of `multi-domain' from data to the loss area, and to combine the characteristics of each domain to create an image. First, we introduce a sym-parameter and its learning method that can mix various losses and can synchronize them with input conditions. Then, we propose Sym-parameterized Generative Network (SGN) using it. Through experiments, we confirmed that SGN could mix the characteristics of various data and losses, and it is possible to translate images to any mixed-domain without ground truths, such as 30% Van Gogh and 20% Monet and 40% snowy.
Simyung Chang, Seonguk Park, John Yang 0001, Nojun Kwak
ICCV4
2019 A Comprehensive Overhaul of Feature Distillation
abstract
We investigate the design aspects of feature distillation methods achieving network compression and propose a novel feature distillation method in which the distillation loss is designed to make a synergy among various aspects: teacher transform, student transform, distillation feature position and distance function. Our proposed distillation loss includes a feature transform with a newly designed margin ReLU, a new distillation feature position, and a partial L2distance function to skip redundant information giving adverse effects to the compression of student. In ImageNet, our proposed method achieves 21.65% of top-1 error with ResNet50, which outperforms the performance of the teacher network, ResNet152. Our proposed method is evaluated on various tasks such as image classification, object detection and semantic segmentation and achieves a significant performance improvement in all tasks. The code is available at bhheo.github.io/overhaul.
Byeongho Heo, Jeesoo Kim, Sangdoo Yun, Hyojin Park 0001, Nojun Kwak, Jin Young Choi 0002
ICCV5
2019 BOOK: Storing Algorithm-Invariant Episodes for Deep Reinforcement Learning
abstract
We introduce a novel method to train agents of reinforcement learning (RL) by sharing knowledge in a way similar to the concept of using a book. The recorded information in the form of a book is the main means by which humans learn knowledge. Nevertheless, the conventional deep RL methods have mainly focused either on experiential learning where the agent learns through interactions with the environment from the start or on imitation learning that tries to mimic the teacher. Contrary to these, our proposed book learning shares key information among different agents in a book-like manner by delving into the following two characteristic features: (1) By defining the linguistic function, input states can be clustered semantically into a relatively small number of core clusters, which are forwarded to other RL agents in a prescribed manner. (2) By defining state priorities and the contents for recording, core experiences can be selected and stored in a small container. We call this container as 'BOOK'. Our method learns hundreds to thousand times faster than the conventional methods by learning only a handful of core cluster information, which shows that deep RL agents can effectively learn through the shared knowledge from other agents.
Simyung Chang, Young Joon Yoo, Jaeseok Choi, Nojun Kwak
ICPRAM4
2019 Two-layer Residual Feature Fusion for Object Detection
abstract
Recently, a lot of single stage detectors using multi-scale features have been actively proposed. They are much faster than two stage detectors that use region proposal networks (RPN) without much degradation in the detection performances. However, the feature maps in the lower layers close to the input which are responsible for detecting small objects in a single stage detector have a problem of insufficient representation power because they are too shallow. There is also a structural contradiction that the feature maps not only have to deliver low-level information to next layers but also have to contain high-level abstraction for prediction. In this paper, we propose a method to enrich the representation power of feature maps using a new feature fusion method which makes use of the information from the consecutive layer. It also adopts a unified prediction module which has an enhanced generalization performance. The proposed method enables more precise prediction, which achieved higher or compatible score than other competitors such as SSD and DSSD on PASCAL VOC and MS COCO. In addition, it maintains the advantage of fast computation of a single stage detector, which requires much less computation than other detectors with similar performance.
Jaeseok Choi, Jisoo Jeong, Nojun Kwak
ICPRAM4
2019 Pose estimator and tracker using temporal flow maps for limbs
abstract
For human pose estimation in videos, it is significant how to use temporal information between frames. In this paper, we propose temporal flow maps for limbs (TML) and a multi-stride method to estimate and track human poses. The proposed temporal flow maps are unit vectors describing the limbs' movements. We constructed a network to learn both spatial information and temporal information end-to-end. Spatial information such as joint heatmaps and part affinity fields is regressed in the spatial network part, and the TML is regressed in the temporal network part. We also propose a data augmentation method to learn various types of TML better. The proposed multi-stride method expands the data by randomly selecting two frames within a defined range. We demonstrate that the proposed method efficiently estimates and tracks human poses on the PoseTrack 2017 and 2018 datasets.
Jihye Hwang, Sungheon Park, Nojun Kwak
IJCNN4
2019 Consistency-based Semi-supervised Learning for Object detection
abstract
Making a precise annotation in a large dataset is crucial to the performance of object detection. While the object detection task requires a huge number of annotated samples to guarantee its performance, placing bounding boxes for every object in each sample is time-consuming and costs a lot. To alleviate this problem, we propose a Consistency-based Semi-supervised learning method for object Detection (CSD), which is a way of using consistency constraints as a tool for enhancing detection performance by making full use of available unlabeled data. Specifically, the consistency constraint is applied not only for object classification but also for the localization. We also proposed Background Elimination (BE) to avoid the negative effect of the predominant backgrounds on the detection performance. We have evaluated the proposed CSD both in single-stage and two-stage detectors and the results show the effectiveness of our method.
Jisoo Jeong, Seungeui Lee, Jeesoo Kim, Nojun Kwak
NeurIPS4
2018 3D Human Pose Estimation with Relational Networks
Sungheon Park, Nojun Kwak
BMVC2
2018 MC-GAN: Multi-conditional Generative Adversarial Network for Image Synthesis
Hyojin Park 0001, Young Joon Yoo, Nojun Kwak
BMVC3
2018 Dynamic Graph Generation Network: Generating Relational Knowledge From Diagrams
abstract
In this work, we introduce a new algorithm for analyzing a diagram, which contains visual and textual information in an Abstract and integrated way. Whereas diagrams contain richer information compared with individual image-based or language-based data, proper solutions for automatically understanding them have not been proposed due to their innate characteristics of multi-modality and arbitrariness of layouts. To tackle this problem, we propose a unified diagram-parsing network for generating knowledge from diagrams based on an object detector and a recurrent neural network designed for a graphical structure. Specifically, we propose a dynamic graph-generation network that is based on dynamic memory and graph theory. We explore the dynamics of information in a diagram with activation of gates in gated recurrent unit (GRU) cells. On publicly available diagram datasets, our model demonstrates a state-of-the-art result that outperforms other baselines. Moreover, further experiments on question answering shows potentials of the proposed method for various applications.
Young Joon Yoo, Jeesoo Kim, Sangkuk Lee, Nojun Kwak
CVPR5
2018 Image Restoration by Estimating Frequency Distribution of Local Patches
abstract
In this paper, we propose a method to solve the image restoration problem, which tries to restore the details of a corrupted image, especially due to the loss caused by JPEG compression. We have treated an image in the frequency domain to explicitly restore the frequency components lost during image compression. In doing so, the distribution in the frequency domain is learned using the cross entropy loss. Unlike recent approaches, we have reconstructed the details of an image without using the scheme of adversarial training. Rather, the image restoration problem is treated as a classification problem to determine the frequency coefficient for each frequency band in an image patch. In this paper, we show that the proposed method effectively restores a JPEG-compressed image with more detailed high frequency components, making the restored image more vivid.
Jaeyoung Yoo, Nojun Kwak
CVPR3
2018 Broadcasting Convolutional Network for Visual Relational Reasoning
Simyung Chang, John Yang 0001, Seonguk Park, Nojun Kwak
ECCV (15)4
2018 Motion Feature Network: Fixed Motion Filter for Action Recognition
Myunggi Lee, Seungeui Lee, Sung Joon Son, Gyutae Park, Nojun Kwak
ECCV (10)5
2018 Genetic-Gated Networks for Deep Reinforcement Learning
abstract
We introduce the Genetic-Gated Networks (G2Ns), simple neural networks that combine a gate vector composed of binary genetic genes in the hidden layer(s) of networks. Our method can take both advantages of gradient-free optimization and gradient-based optimization methods, of which the former is effective for problems with multiple local minima, while the latter can quickly find local minima. In addition, multiple chromosomes can define different models, making it easy to construct multiple models and can be effectively applied to problems that require multiple models. We show that this G2N can be applied to typical reinforcement learning algorithms to achieve a large improvement in sample efficiency and performance.
Simyung Chang, John Yang 0001, Jaeseok Choi, Nojun Kwak
NeurIPS4
2018 Paraphrasing Complex Network: Network Compression via Factor Transfer
abstract
Many researchers have sought ways of model compression to reduce the size of a deep neural network (DNN) with minimal performance degradation in order to use DNNs in embedded systems. Among the model compression methods, a method called knowledge transfer is to train a student network with a stronger teacher network. In this paper, we propose a novel knowledge transfer method which uses convolutional operations to paraphrase teacher's knowledge and to translate it for the student. This is done by two convolutional modules, which are called a paraphraser and a translator. The paraphraser is trained in an unsupervised manner to extract the teacher factors which are defined as paraphrased information of the teacher network. The translator located at the student network extracts the student factors and helps to translate the teacher factors by mimicking them. We observed that our student network trained with the proposed factor transfer method outperforms the ones trained with conventional knowledge transfer methods.
Jangho Kim, Seonguk Park, Nojun Kwak
NeurIPS3
2018 Independent component analysis by lp-norm optimization
Sungheon Park, Nojun Kwak
Pattern Recognit.2
2018 Procrustean Regression: A Flexible Alignment-Based Framework for Nonrigid Structure Estimation
abstract
Non-rigid structure from motion (NRSfM) is a fundamental problem of computer vision. Recently, it has been shown that incorporating shape alignment in NRSfM can improve the performance significantly compared with the other algorithms, which do not consider shape alignment. However, realizing this idea was at a cost of a heavy, complicated process, which limits its usefulness and possible extensions. In this paper, we propose a novel regression framework for NRSfM, of which the variables (3D shapes) are regularized based on their aligned shapes. We show that this can be casted into an unconstrained problem or a problem with simple bound constraints, which can be efficiently solved by existing solvers. This framework can be easily integrated with numerous existing models and assumptions, such as orthographic or perspective camera models, occlusion, low-rank assumption, smooth deformations, and so on, which makes it more practical for various real situations. The experimental results show that the proposed method gives competitive result to the state-of-the-art methods for orthographic projection with much less time complexity and memory requirement, and outperforms the existing methods for perspective projection.
Sungheon Park, Minsik Lee 0001, Nojun Kwak
IEEE Trans. Image Process.3
2017 Enhancement of SSD by concatenating feature maps for object detection
Jisoo Jeong, Hyojin Park 0001, Nojun Kwak
BMVC3
2017 Superpixel-based semantic segmentation trained by statistical process control
Hyojin Park 0001, Jisoo Jeong, Young Joon Yoo, Nojun Kwak
BMVC4
2017 Matching video net: Memory-based embedding for video action recognition
abstract
Most of recent successful researches on action recognition are based on deep learning structures. Nonetheless, training deep neural networks is notorious for requiring huge amount of data. On the other hand, not enough data can lead to an overfitted model. In this work, we propose a novel model, matching video net (MVN), which can be trained with a small amount of data. In order to avoid the problem of overfitting, we use a non-parametric setup on top of parametric networks with external memories. An input clip of video is transformed into an embedding space and matched to the memorized samples in the embedding space. Then, the similarities between the input and the memorized data are measured to determine the nearest neighbors. We perform experiments in a supervised manner on action recognition datasets, achieving state-of-the-art results. Moreover, we applied our model to one-shot learning problems with a novel training strategy. Our model achieves surprisingly good results in predicting unseen action classes from only a few examples.
Myunggi Lee, Nojun Kwak
IJCNN3
2017 Online recognition of handwritten music symbols
abstract
In this paper, we propose an effective online method to recognize handwritten music symbols. Based on the fact that most music symbols can be regarded as combinations of several basic strokes, the proposed method first classifies all the strokes comprising an input symbol and then recognizes the symbol based on the results of stroke classification. For stroke classification, we propose to use three types of features, which are the size information, the histogram of directional movement angles, and the histogram of undirected movement angles. When combining classified strokes into a music symbol, we utilize their sizes and spatial relation together with their combination. The proposed method is evaluated using two datasets including HOMUS, one of the largest music symbol datasets. As a result, it achieves a significant improvements of about 10% in recognition rates compared to the state-of-the-art method for the datasets. This shows the superiority of the proposed method in online handwritten music symbol recognition.
Jiyong Oh, Sung Joon Son, Sangkuk Lee, Ji-Won Kwon, Nojun Kwak
Int. J. Document Anal. Recognit.5
2017 Implementing Kernel Methods Incrementally by Incremental Nonlinear Projection Trick
abstract
Recently, the nonlinear projection trick (NPT) was introduced enabling direct computation of coordinates of samples in a reproducing kernel Hilbert space. With NPT, any machine learning algorithm can be extended to a kernel version without relying on the so called kernel trick. However, NPT is inherently difficult to be implemented incrementally because an ever increasing kernel matrix should be treated as additional training samples are introduced. In this paper, an incremental version of the NPT (INPT) is proposed based on the observation that the centerization step in NPT is unnecessary. Because the proposed INPT does not change the coordinates of the old data, the coordinates obtained by INPT can directly be used in any incremental methods to implement a kernel version of the incremental methods. The effectiveness of the INPT is shown by applying it to implement incremental versions of kernel methods such as, kernel singular value decomposition, kernel principal component analysis, and kernel discriminant analysis which are utilized for problems of kernel matrix reconstruction, letter classification, and face image retrieval, respectively.
Nojun Kwak
IEEE Trans. Cybern.1
2016 Analysis on the Dropout Effect in Convolutional Neural Networks
Sungheon Park, Nojun Kwak
ACCV (2)2
2016 Unregistered Bosniak Classification with Multi-phase Convolutional Neural Networks
Myunggi Lee, Hyeogjin Lee, Jiyong Oh, Hak Jong Lee, Seung Hyup Kim, Nojun Kwak
ICONIP (4)6
2016 Outdoor Context Awareness Device That Enables Mobile Phone Users to Walk Safely through Urban Intersections
abstract
Research in social science has shown that the mobile phone users pay less attention to their surroundings, which exposes them to various hazards such as collisions with vehicles than other pedestrians. In this paper, we propose a novel handheld device that assists mobile phone users to walk more safely outdoors. The proposed system is implemented on a smart phone and uses its back camera to detect the current outdoor context, e.g. traffic intersections, roadways, and sidewalks, finally alerts the user of unsafe situations using sound and vibration from the phone. The outdoor context awareness is performed by three steps: preprocessing, feature extraction, and context recognition. First, it improves the image contrast while removing image noise, and then it extracts the color and texture descriptors from each pixel. Next, each pixel is classified as an intersection, sidewalk, or roadway using a support vector machine-based classifier. Then, to support the real-time performance on the smart phone, a multi-scale classification is applied to input image, where the coarse layer first discriminates the boundary pixels from the background and the fine layer categorizes the boundary pixels as sidewalk, roadway, or intersection. In order to demonstrate the effectiveness of the proposed method, some real-world experiments were performed, then the results showed that the proposed system has the accuracy of above 98% at the various environments. © Copyright 2016 by SCITEPRESS - Science and Technology Publications, Lda. All rights reserved.
Jihye Hwang, Yeounggwang Ji, Nojun Kwak, Eun Yi Kim
ICPRAM3
2016 Generalized mean for robust principal component analysis
abstract
In this paper, we propose a robust principal component analysis (PCA) to overcome the problem that PCA is prone to outliers included in the training set. Different from the other alternatives which commonly replace L2-norm by other distance measures, the proposed method alleviates the negative effect of outliers using the characteristic of the generalized mean keeping the use of the Euclidean distance. The optimization problem based on the generalized mean is solved by a novel method. We also present a generalized sample mean, which is a generalization of the sample mean, to estimate a robust mean in the presence of outliers. The proposed method shows better or equivalent performance than the conventional PCAs in various problems such as face reconstruction, clustering, and object categorization.
Jiyong Oh, Nojun Kwak
Pattern Recognit.2
2015 Membership representation for detecting block-diagonal structure in low-rank or sparse subspace clustering
abstract
Recently, there have been many proposals with state-of-the-art results in subspace clustering that take advantages of the low-rank or sparse optimization techniques. These methods are based on self-expressive models, which have well-defined theoretical aspects. They produce matrices with (approximately) block-diagonal structure, which is then applied to spectral clustering. However, there is no definitive way to construct affinity matrices from these block-diagonal matrices and it is ambiguous how the performance will be affected by the construction method. In this paper, we propose an alternative approach to detect block-diagonal structures from these matrices. The proposed method shares the philosophy of the above subspace clustering methods, in that it is a self-expressive system based on a Hadamard product of a membership matrix. To resolve the difficulty in handling the membership matrix, we solve the convex relaxation of the problem and then transform the representation to a doubly stochastic matrix, which is closely related to spectral clustering. The result of our method has eigenvalues normalized in between zero and one, which is more reliable to estimate the number of clusters and to perform spectral clustering. The proposed method shows competitive results in our experiments, even though we simply count the number of eigenvalues larger than a certain threshold to find the number of clusters.
Minsik Lee 0001, Hyeogjin Lee, Nojun Kwak
CVPR4
2015 Character recognition for the machine reader zone of electronic identity cards
abstract
This paper proposes an overall procedure of recognizing the machine reader zone of a real world picture of a passport. To begin with, the proposed method finds the area of passport from the input image and its rotation angle is determined. With the rectified passport image by counter-rotating the area of passport, the machine reader zone is found and an inverse projective transform is performed to remove projective distortion. Then, each code is extracted and enhanced by using adaptive posterization. Template matching with improved similarity measure is applied to classify the codes. To classify the number 0 and the character O, a support vector machine is used. The experimental results show the correct character recognition rate of 99.77% and the correct recognition rate of 83.84%.
Hyeogjin Lee, Nojun Kwak
ICIP2
2015 Illumination robust optical flow estimation by illumination-chromaticity decoupling
abstract
In this paper, a novel optical flow algorithm which is robust to illumination variation is proposed. HSL color space is adopted to decouple illumination and chromaticity information. The chromaticity component is normalized by chroma and transformed to the cartesian coordinate. Then, the decoupled distance is defined using both illumination and chromaticity. Cost function for the optical flow is formulated using l1norm of the decoupled distance with Huber norm regularization term. The cost function is efficiently minimized by utilizing Legendre-Fenchel transform. Optical flow field is further refined via weighted median filter whose weight is also based on the decoupled distance. Experimental results show that the proposed method works robustly even in the presence of severe illumination variation.
Sungheon Park, Nojun Kwak
ICIP2
2015 Efficient l1-Norm-Based Low-Rank Matrix Approximations for Large-Scale Problems Using Alternating Rectified Gradient Method
abstract
Low-rank matrix approximation plays an important role in the area of computer vision and image processing. Most of the conventional low-rank matrix approximation methods are based on the l2 -norm (Frobenius norm) with principal component analysis (PCA) being the most popular among them. However, this can give a poor approximation for data contaminated by outliers (including missing data), because the l2 -norm exaggerates the negative effect of outliers. Recently, to overcome this problem, various methods based on the l1 -norm, such as robust PCA methods, have been proposed for low-rank matrix approximation. Despite the robustness of the methods, they require heavy computational effort and substantial memory for high-dimensional data, which is impractical for real-world problems. In this paper, we propose two efficient low-rank factorization methods based on the l1 -norm that find proper projection and coefficient matrices using the alternating rectified gradient method. The proposed methods are applied to a number of low-rank matrix approximation problems to demonstrate their efficiency and robustness. The experimental results show that our proposals are efficient in both execution time and reconstruction performance unlike other state-of-the-art methods.
Eunwoo Kim, Minsik Lee 0001, Chong-Ho Choi, Nojun Kwak, Songhwai Oh
IEEE Trans. Neural Networks Learn. Syst.4
2014 Optimal Conjugate Gradient Algorithm for Generalization of Linear Discriminant Analysis Based on L1 Norm
abstract
This paper analyzes a linear discriminant subspace technique from an L-1 point of view. We propose an efficient and optimal algorithm that addresses several major issues with prior work based on, not only the L-1 based LDA algorithm but also its L-2 counterpart. This includes algorithm implementation, effect of outliers and optimality of parameters used. The key idea is to use conjugate gradient to optimize the L-1 cost function and to find an optimal learning factor during the update of the weight vector in the subspace. Experimental results on UCI datasets reveal that the present method is a significant improvement over the previous work. Mathematical treatment for the proposed algorithm and calculations for learning factor are the main subject of this paper.
Kanishka Tyagi, Nojun Kwak, Michael T. Manry
ICPRAM2
2014 Principal Component Analysis by $L_{p}$ -Norm Maximization
abstract
This paper proposes several principal component analysis (PCA) methods based on Lp-norm optimization techniques. In doing so, the objective function is defined using the Lp-norm with an arbitrary p value, and the gradient of the objective function is computed on the basis of the fact that the number of training samples is finite. In the first part, an easier problem of extracting only one feature is dealt with. In this case, principal components are searched for either by a gradient ascent method or by a Lagrangian multiplier method. When more than one feature is needed, features can be extracted one by one greedily, based on the proposed method. Second, a more difficult problem is tackled that simultaneously extracts more than one feature. The proposed methods are shown to find a local optimal solution. In addition, they are easy to implement without significantly increasing computational complexity. Finally, the proposed methods are applied to several datasets with different values of p and their performances are compared with those of conventional PCA methods.
Nojun Kwak
IEEE Trans. Cybern.1
2013 Generalized mean for feature extraction in one-class classification problems
Jiyong Oh, Nojun Kwak, Minsik Lee 0001, Chong-Ho Choi
Pattern Recognit.2
2013 Generalization of linear discriminant analysis using Lp-norm
Jae Hyun Oh, Nojun Kwak
Pattern Recognit. Lett.2
2013 Nonlinear Projection Trick in Kernel Methods: An Alternative to the Kernel Trick
abstract
In kernel methods such as kernel principal component analysis (PCA) and support vector machines, the so called kernel trick is used to avoid direct calculations in a high (virtually infinite) dimensional kernel space. In this brief, based on the fact that the effective dimensionality of a kernel space is less than the number of training samples, we propose an alternative to the kernel trick that explicitly maps the input data into a reduced dimensional kernel space. This is easily obtained by the eigenvalue decomposition of the kernel matrix. The proposed method is named as the nonlinear projection trick in contrast to the kernel trick. With this technique, the applicability of the kernel methods is widened to arbitrary algorithms that do not use the dot product. The equivalence between the kernel trick and the nonlinear projection trick is shown for several conventional kernel methods. In addition, we extend PCA-L1, which uses L1-norm instead of L2-norm (or dot product), into a kernel version and show the effectiveness of the proposed approach.
Nojun Kwak
IEEE Trans. Neural Networks Learn. Syst.1
2012 Detection and Recovery of Occluded Face Images based on Correlation between Pixels
Nojun Kwak
ICPRAM (2)2
2012 Boosted-PCA for binary classification problems
abstract
In this paper, a Boosted-PCA algorithm is proposed for efficient classification of two class data. Conventionally, in classification problems, the roles of feature extraction and classification have been distinct, i.e., a feature extraction method and a classifier are applied sequentially to classify input variable into several categories. In this paper, these two steps are combined into one resulting in a good classification performance. More specifically, each principal component is treated as a weak classifier in Adaboost algorithm to constitute a strong classifier for binary classification problems. The proposed algorithm is applied to UCI data set and showed better recognition rates than sequential application of feature extraction and classification methods such as PCA+1NN and PCA+SVM.
Seaung Lok Ham, Nojun Kwak
ISCAS2
2012 Kernel discriminant analysis for regression problems
Nojun Kwak
Pattern Recognit.1
2012 Pixel selection based on discriminant features with application to face recognition
Sang-Il Choi, Chong-Ho Choi, Gu-Min Jeong, Nojun Kwak
Pattern Recognit. Lett.4
2011 Face recognition based on 2D images under illumination and pose variations
Sang-Il Choi, Chong-Ho Choi, Nojun Kwak
Pattern Recognit. Lett.3
2010 Feature extraction based on subspace methods for regression problems
Nojun Kwak, Jung-Won Lee
Neurocomputing1
2009 Feature extraction for one-class classification problems: Enhancements to biased discriminant analysis
Nojun Kwak, Jiyong Oh
Pattern Recognit.1
2008 Dimensionality reduction based on ICA for regression problems
Nojun Kwak, Chunghoon Kim, Hwangnam Kim
Neurocomputing1
2008 Principal Component Analysis Based on L1-Norm Maximization
abstract
A method of principal component analysis (PCA) based on a new L1-norm optimization technique is proposed. Unlike conventional PCA which is based on L2-norm, the proposed method is robust to outliers because it utilizes L1-norm which is less sensitive to outliers. It is invariant to rotations as well. The proposed L1-norm optimization technique is intuitive, simple, and easy to implement. It is also proven to find a locally maximal solution. The proposed method is applied to several datasets and the performances are compared with those of other conventional methods.
Nojun Kwak
IEEE Trans. Pattern Anal. Mach. Intell.1
2008 Feature extraction for classification problems and its application to face recognition
Nojun Kwak
Pattern Recognit.1
2007 Feature Extraction Based on Direct Calculation of Mutual Information
abstract
In many pattern recognition problems, it is desirable to reduce the number of input features by extracting important features related to the problems. By focusing on only the problem-relevant features, the dimension of features can be greatly reduced and thereby can result in a better generalization performance with less computational complexity. In this paper, we propose a feature extraction method for handling classification problems. The proposed algorithm is used to search for a set of linear combinations of the original features, whose mutual information with the output class can be maximized. The mutual information between the extracted features and the output class is calculated by using the probability density estimation based on the Parzen window method. A greedy algorithm using the gradient descent method is used to determine the new features. The computational load is proportional to the square of the number of samples. The proposed method was applied to several classification problems, which showed better or comparable performances than the conventional feature extraction methods.
Nojun Kwak
Int. J. Pattern Recognit. Artif. Intell.1
2006 Feature Extraction with Weighted Samples Based on Independent Component Analysis
Nojun Kwak
ICANN (2)1
2006 Dimensionality Reduction Based on ICA for Regression Problems
Nojun Kwak, Chunghoon Kim
ICANN (1)1
2003 Feature Extraction Based on ICA for Binary Classification Problems
abstract
In manipulating data such as in supervised learning, we often extract new features from the original features for the purpose of reducing the dimensions of feature space and achieving better performance. In this paper, we show how standard algorithms for independent component analysis (ICA) can be appended with binary class labels to produce a number of features that do not carry information about the class labels-these features will be discarded-and a number of features that do. We also provide a local stability analysis of the proposed algorithm. The advantage is that general ICA algorithms become available to a task of feature extraction for classification problems by maximizing the joint mutual information between class labels and new features, although only for two-class problems. Using the new features, we can greatly reduce the dimension of feature space without degrading the performance of classifying systems.
Nojun Kwak, Chong-Ho Choi
IEEE Trans. Knowl. Data Eng.1
2002 A New Method of Feature Extraction and Its Stability
Nojun Kwak, Chong-Ho Choi
ICANN1
2002 Face recognition using feature extraction based on independent component analysis
abstract
We have explored a new method of feature extraction for face recognition. It is based on independent component analysis (ICA), but unlike original ICA, one of the unsupervised learning methods, it is developed to be well suited for classification problems by utilizing class information. By using ICA in solving supervised classification problems, we can obtain new features which are made as independent from each other as possible and which convey the class information faithfully. We have applied this method on Yale face databases and AT and T face databases and compared the performance with those of conventional methods such as principal component analysis (PCA), Fisher's linear discriminant (FLD), and so on. The experimental results show that for both databases the proposed method outperforms the others.
Nojun Kwak, Chong-Ho Choi, Narendra Ahuja
ICIP (2)1
2002 Input Feature Selection by Mutual Information Based on Parzen Window
abstract
Mutual information is a good indicator of relevance between variables, and have been used as a measure in several feature selection algorithms. However, calculating the mutual information is difficult, and the performance of a feature selection algorithm depends on the accuracy of the mutual information. In this paper, we propose a new method of calculating mutual information between input and class variables based on the Parzen window, and we apply this to a feature selection algorithm for classification problems.
Nojun Kwak, Chong-Ho Choi
IEEE Trans. Pattern Anal. Mach. Intell.1
2002 Input feature selection for classification problems
abstract
Feature selection plays an important role in classifying systems such as neural networks (NNs). We use a set of attributes which are relevant, irrelevant or redundant and from the viewpoint of managing a dataset which can be huge, reducing the number of attributes by selecting only the relevant ones is desirable. In doing so, higher performances with lower computational effort is expected. In this paper, we propose two feature selection algorithms. The limitation of mutual information feature selector (MIFS) is analyzed and a method to overcome this limitation is studied. One of the proposed algorithms makes more considered use of mutual information between input attributes and output classes than the MIFS. What is demonstrated is that the proposed method can provide the performance of the ideal greedy selection algorithm when information is distributed uniformly. The computational load for this algorithm is nearly the same as that of MIFS. In addition, another feature selection algorithm using the Taguchi method is proposed. This is advanced as a solution to the question as to how to identify good features with as few experiments as possible. The proposed algorithms are applied to several classification problems and compared with MIFS. These two algorithms can be combined to complement each other's limitations. The combined algorithm performed well in several experiments and should prove to be a useful method in selecting features for classification problems.
Nojun Kwak, Chong-Ho Choi
IEEE Trans. Neural Networks1
2001 Feature Extraction Using ICA
Nojun Kwak, Chong-Ho Choi, Jin Young Choi 0002
ICANN1
2001 Disturbance Attenuation in Robot Control
abstract
We propose a model based disturbance attenuator (MBDA) with the conventional PD controller for robot manipulators. It is a generalization of the MBDA structure in Choi et al. (1999) and is applied to a robot manipulator which is nonlinear. This method does not require an accurate model of a robot manipulator and takes care of disturbances or modeling errors so that the plant output remains relatively unaffected by them. The output error due to the gravity or constant disturbance can be completely eliminated by this method in the same way as PID controllers. In addition, this can be easily implemented at a moderate computational cost. We apply this to a two-link robot manipulator and compare its performance with PD and PID controllers. Simulation results show that the proposed method is very effective in controlling robot manipulators.
Chong-Ho Choi, Nojun Kwak
ICRA2
1999 Improved mutual information feature selector for neural networks in supervised learning
abstract
In classification problems, we use a set of attributes which are relevant, irrelevant or redundant. By selecting only the relevant attributes of the data as input features of a classifying system and excluding redundant ones, higher performance is expected with smaller computational effort. We propose an algorithm of feature selection that makes more careful use of the mutual informations between input attributes and others than the mutual information feature selector (MIFS). The proposed algorithm is applied in several feature selection problems and compared with the MIFS. Experimental results show that the proposed algorithm can be well used in feature selection problems.
Nojun Kwak, Chong-Ho Choi
IJCNN1