VLDB 2026 Research / reviewers in the wild / expert
Jangho Kim
dblp:118/1777
· DBLP profile ↗
19ranked-venue papers
9as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 5 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Weight initialization based on gradient similarity for versatile machine unlearningabstractThe growing necessity for deep learning applications to adhere to rising data privacy standards has made machine unlearning crucial in removing the impact of specific examples from a given model. Although exact unlearning, which retrains the model from scratch using the remaining dataset, satisfies this objective, it is computationally expensive, leading approximate unlearning to be an active research field. However, we argue that most existing approximate algorithms fail to provide robust privacy guarantees. These methods are often "biased," designed to defend against attacks on a single model property (e.g., loss) while leaving information encoded in other properties (e.g., entropy) exposed. An adaptive attacker can simply exploit this bias, making the unlearning ineffective, as a model's privacy is only as strong as its most vulnerable point. For example, gradient descent using random labels may enhance indistinguishability with respect to loss by directly lowering the probability of the correct label, but it fails to achieve comparable results for entropy, as it does not increase the entropy to levels similar to those of unseen data. To address this problem, we propose a Weight Initialization based on Gradient similarity dubbed WIG, a novel algorithm that provides unbiased unlearning that approximates the retraining process. Instead of targeting a specific property, WIG induces catastrophic forgetting by partially initializing weights based on their gradient similarity between the train set and the data to be forgotten. WIG achieves the best indistinguishability among current state-of-the-art approximate unlearning algorithms across diverse metrics, including various membership inference attacks, Inter-class confusion test, U-LiRA, NeurIPS Machine Unlearning Challenge, and Gaussian Poison, demonstrating its versatility. Doun Lee, Jongyun Shin, Jinwoo Bae, Hyunjoon Cho, Jangho Kim |
Proc. Priv. Enhancing Technol. | 5 |
| 2025 | Exploring Diverse Sparse Network Structures via Dynamic Pruning with Weight AlignmentabstractDeep neural networks (DNNs) often require a large number of parameters, which has led to the development of model pruning techniques that remove weight connections. In this paper, we propose a new method to maximize the effect of finding sparse patterns through a gradient scaling technique that modifies the weight distribution. This approach allows for the exploration of more diverse sparse patterns compared to traditional dynamic pruning methods, leading to the discovery of stable subnetworks. Through various experiments, we demonstrate the importance of exploration in finding better sparse patterns. We achieve state-of-the-art performance across multiple network architectures on datasets such as CIFAR-10/100 and ImageNet. The code is available at https://github.com/Acasia/DWA. Jongyun Shin, Sangho An, Jangho Kim |
CIKM | 4 |
| 2025 | Sparse Structure Exploration and Re-optimization for Vision TransformerabstractVision Transformers (ViTs) achieve outstanding performance by effectively capturing long-range dependencies between image patches (tokens). However, the high computational cost and memory requirements of ViTs present challenges for model compression and deployment on edge devices. In this study, we introduce a new framework, Sparse Structure Exploration and Re-optimization (SERo), specifically designed to maximize pruning efficiency in ViTs. Our approach focuses on (1) hardware-friendly pruning that fully compresses pruned parameters instead of zeroing them out, (2) separating the exploration and re-optimization phases \red{in order to find the optimal structure among various possible sparse structures}, and (3) using a simple gradient magnitude-based criterion for pruning a pre-trained model. SERo iteratively refines pruning masks to identify optimal sparse structures and then re-optimizes the pruned structure, reducing computational costs while maintaining model performance. Experimental results indicate that SERo surpasses existing pruning methods across various ViT models in both performance and computational efficiency. For example, SERo achieves a 69% reduction in computational cost and a 2.4x increase in processing speed for DeiT-Base model, with only a 1.55% drop in accuracy. Implementation code: https://github.com/Ahnho/SERo/ Sangho An, Keonho Lee, Jingang Huh, Chanwoong Kwak, Moonsub Jin, Jangho Kim |
UAI | 8 |
| 2025 | Magnitude Attention-based Dynamic PruningabstractExisting pruning methods often rely on weight importance to identify sparse structures but typically apply this information statically, without leveraging it adaptively during training. In this work, we propose a novel approach - M agnitude A ttention-based Dynamic P runing (MAP) method, which applies the importance of weights throughout both the forward and backward paths to explore sparse model structures dynamically. Magnitude attention is defined based on the magnitude of weights as continuous real-valued numbers enabling a seamless transition from a redundant to an effective sparse network by promoting efficient exploration. Additionally, the attention mechanism ensures more effective updates for important layers within the sparse network. Later, our approach shifts from exploration to exploitation, exclusively updating the sparse model composed of crucial weights based on the explored structure, resulting in pruned models that not only achieve performance comparable to dense models but also outperform previous pruning methods on CIFAR-10/100 and ImageNet. Jihye Back, Namhyuk Ahn, Jangho Kim |
Expert Syst. Appl. | 3 |
| 2025 | Layerwise-priority-based gradient adjustment for few-shot learning
Jangho Kim, Junhoo Lee, Donghoon Han, Nojun Kwak |
Expert Syst. Appl. | 1 |
| 2025 | Entropy-Guided Meta-Initialization regularization for few-shot text classification
Jongyun Shin, Jangho Kim |
Knowl. Based Syst. | 3 |
| 2024 | Cooperative Meta-Learning with Gradient AugmentationabstractModel agnostic meta-learning (MAML) is one of the most widely used gradient-based meta-learning, consisting of two optimization loops: an inner loop and outer loop. MAML learns the new task from meta-initialization parameters with an inner update and finds the meta-initialization parameters in the outer loop. In general, the injection of noise into the gradient of the model for augmenting the gradient is one of the widely used regularization methods. In this work, we propose a novel cooperative meta-learning framework dubbed CML which leverages gradient-level regularization with gradient augmentation. We inject learnable noise into the gradient of the model for the model generalization. The key idea of CML is introducing the co-learner which has no inner update but the outer loop update to augment gradients for finding better meta-initialization parameters. Since the co-learner does not update in the inner loop, it can be easily deleted after meta-training. Therefore, CML infers with only meta-learner without additional cost and performance degradation. We demonstrate that CML is easily applicable to gradient-based meta-learning methods and CML leads to increased performance in few-shot regression, few-shot image classification and few-shot node classification tasks. Our codes are at https://github.com/JJongyn/CML. Jongyun Shin, Seungjin Han, Jangho Kim |
UAI | 3 |
| 2023 | Finding Efficient Pruned Network via Refined Gradients for Pruned WeightsabstractWith the growth of deep neural networks (DNN), the number of DNN parameters has drastically increased. This makes DNN models hard to be deployed on resource-limited embedded systems. To alleviate this problem, dynamic pruning methods have emerged, which try to find diverse sparsity patterns during training by utilizing Straight-Through-Estimator (STE) to approximate gradients of pruned weights. STE can help the pruned weights revive in the process of finding dynamic sparsity patterns. However, using these coarse gradients causes training instability and performance degradation owing to the unreliable gradient signal of the STE approximation. In this work, to tackle this issue, we introduce refined gradients to update the pruned weights by forming dual forwarding paths from two sets (pruned and unpruned) of weights. We propose a novel Dynamic Collective Intelligence Learning (DCIL) which makes use of the learning synergy between the collective intelligence of both weight sets. We verify the usefulness of the refined gradients by showing enhancements in the training stability and the model performance on the CIFAR and ImageNet datasets. DCIL outperforms various previously proposed pruning schemes including other dynamic pruning methods with enhanced stability during training. The code is provided in Github. Jangho Kim, Jayeon Yoo, Yeji Song, KiYoon Yoo, Nojun Kwak |
ACM Multimedia | 1 |
| 2023 | Self-Distilled Self-supervised Representation LearningabstractState-of-the-art frameworks in self-supervised learning have recently shown that fully utilizing transformer-based models can lead to performance boost compared to conventional CNN models. Striving to maximize the mutual information of two views of an image, existing works apply a contrastive loss to the final representations. Motivated by self-distillation in the supervised regime, we further exploit this by allowing the intermediate representations to learn from the final layer via the contrastive loss. Through self-distillation, the intermediate layers are better suited for instance discrimination, making the performance of an early-exited sub-network not much degraded from that of the full network. This renders the pretext task easier also for the final layer, leading to better representations. Our method, Self-Distilled Self-Supervised Learning (SDSSL), outperforms competitive baselines (SimCLR, BYOL and MoCo v3) using ViT on various tasks and datasets. In the linear evaluation and k-NN protocol, SDSSL not only leads to superior performance in the final layers, but also in most of the lower layers. Furthermore, qualitative and quantitative analyses show how representations are formed more effectively along the transformer layers. Code is available at https://github.com/hagiss/SDSSL. Jiho Jang, Seonhoon Kim, KiYoon Yoo, Chaerin Kong, Jangho Kim, Nojun Kwak |
WACV | 5 |
| 2022 | Variational On-the-Fly PersonalizationabstractWith the development of deep learning (DL) technologies, the demand for DL-based services on personal devices, such as mobile phones, also increases rapidly. In this paper, we propose a novel personalization method, Variational On-the-Fly Personalization. Compared to the conventional personalization methods that require additional fine-tuning with personal data, the proposed method only requires forwarding a handful of personal data on-the-fly. Assuming even a single personal data can convey the characteristics of a target person, we develop the variational hyper-personalizer to capture the weight distribution of layers that fits the target person. In the testing phase, the hyper-personalizer estimates the model’s weights on-the-fly based on personality by forwarding only a small amount of (even a single) personal enrollment data. Hence, the proposed method can perform the personalization without any training software platform and additional cost in the edge device. In experiments, we show our approach can effectively generate reliable personalized models via forwarding (not back-propagating) a handful of samples. Jangho Kim, Juntae Lee, Simyung Chang, Nojun Kwak |
ICML | 1 |
| 2022 | Domain Generalization with Relaxed Instance Frequency-wise Normalization for Multi-device Acoustic Scene ClassificationabstractWhile using two-dimensional convolutional neural networks (2D-CNNs) in image processing, it is possible to manipulate domain information using channel statistics, and instance normalization has been a promising way to get domain-invariant features. However, unlike image processing, we analyze that domain-relevant information in an audio feature is dominant in frequency statistics rather than channel statistics. Motivated by our analysis, we introduce Relaxed Instance Frequency-wise Normalization (RFN): a plug-and-play, explicit normalization module along the frequency axis which can eliminate instance-specific domain discrepancy in an audio feature while relaxing undesirable loss of useful discriminative information. Empirically, simply adding RFN to networks shows clear margins compared to previous domain generalization approaches on acoustic scene classification and yields improved robustness for multiple audio devices. Especially, the proposed RFN won the DCASE2021 challenge TASK1A, low-complexity acoustic scene classification with multiple devices, with a clear margin, and RFN is an extended work of our technical report. Byeonggeun Kim, Seunghan Yang, Jangho Kim, Hyunsin Park, Juntae Lee, Simyung Chang |
INTERSPEECH | 3 |
| 2021 | Prototype-Based Personalized PruningabstractNowadays, as edge devices such as smartphones become prevalent, there are increasing demands for personalized services. However, traditional personalization methods are not suitable for edge devices because retraining or finetuning is needed with limited personal data. Also, a full model might be too heavy for edge devices with limited resources. Unfortunately, model compression methods which can handle the model complexity issue also require the retraining phase. These multiple training phases generally need huge computational cost during on-device learning which can be a burden to edge devices. In this work, we propose a dynamic personalization method called prototype-based personalized pruning (PPP). PPP considers both ends of personalization and model efficiency. After training a network, PPP can easily prune the network with a prototype representing the characteristics of personal data and it performs well without retraining or finetuning. We verify the usefulness of PPP on a couple of tasks in computer vision and Keyword spotting. Jangho Kim, Simyung Chang, Sungrack Yun, Nojun Kwak |
ICASSP | 1 |
| 2021 | Vehicle Image Generation Going Well with the Surroundings
Jeesoo Kim, Jangho Kim, Jaeyoung Yoo, Nojun Kwak |
ICONIP (4) | 2 |
| 2021 | PQK: Model Compression via Pruning, Quantization, and Knowledge DistillationabstractAs edge devices become prevalent, deploying Deep Neural Networks (DNN) on edge devices has become a critical issue. However, DNN requires a high computational resource which is rarely available for edge devices. To handle this, we propose a novel model compression method for the devices with limited computational resources, called PQK consisting of pruning, quantization, and knowledge distillation (KD) processes. Unlike traditional pruning and KD, PQK makes use of unimportant weights pruned in the pruning process to make a teacher network for training a better student network without pre-training the teacher model. PQK has two phases. Phase 1 exploits iterative pruning and quantization-aware training to make a lightweight and power-efficient model. In phase 2, we make a teacher network by adding unimportant weights unused in phase 1 to a pruned network. By using this teacher network, we train the pruned network as a student network. In doing so, we do not need a pre-trained teacher network for the KD framework because the teacher and the student networks coexist within the same network. We apply our method to the recognition model and verify the effectiveness of PQK on keyword spotting (KWS) and image recognition. Jangho Kim, Simyung Chang, Nojun Kwak |
Interspeech | 1 |
| 2020 | Feature-map-level Online Adversarial Knowledge DistillationabstractFeature maps contain rich information about image intensity and spatial correlation. However, previous online knowledge distillation methods only utilize the class probabilities. Thus in this paper, we propose an online knowledge distillation method that transfers not only the knowledge of the class probabilities but also that of the feature map using the adversarial training framework. We train multiple networks simultaneously by employing discriminators to distinguish the feature map distributions of different networks. Each network has its corresponding discriminator which discriminates the feature map from its own as fake while classifying that of the other network as real. By training a network to fool the corresponding discriminator, it can learn the other network’s feature map distribution. We show that our method performs better than the conventional direct alignment method such as L1 and is more suitable for online distillation. Also, we propose a novel cyclic learning scheme for training more than two networks together. We have applied our method to various network architectures on the classification task and discovered a significant improvement of performance especially in the case of training a pair of a small network and a large one. Inseop Chung, Seonguk Park, Jangho Kim, Nojun Kwak |
ICML | 3 |
| 2020 | Feature Fusion for Online Mutual Knowledge DistillationabstractWe propose a learning framework named Feature Fusion Learning (FFL) that efficiently trains a powerful classifier through a fusion module which combines the feature maps generated from parallel neural networks and generates meaningful feature maps. Specifically, we train a number of parallel neural networks as sub-networks, then we combine the feature maps from each sub-network using a fusion module to create a more meaningful feature map. The fused feature map is passed into the fused classifier for overall classification. Unlike existing feature fusion methods, in our framework, an ensemble of sub-network classifiers transfers its knowledge to the fused classifier and then the fused classifier delivers its knowledge back to each subnetwork, mutually teaching one another in an online-knowledge distillation manner. This mutually teaching system not only improves the performance of the fused classifier but also obtains performance gain in each sub-network. Moreover, our model is more beneficial than other alternative methods because different types of network can be used for each sub-network. We have performed a variety of experiments on multiple datasets such as CIFAR-10, CIFAR-100 and ImageNet and proved that our method is more effective than other alternative methods in terms of performances of both sub-networks and the fused classifier, and the aspect of generating meaningful feature maps. The code is available at this link1. Jangho Kim, Minsung Hyun, Inseop Chung, Nojun Kwak |
ICPR | 1 |
| 2020 | Position-based Scaled Gradient for Model Quantization and PruningabstractWe propose the position-based scaled gradient (PSG) that scales the gradient depending on the position of a weight vector to make it more compression-friendly. First, we theoretically show that applying PSG to the standard gradient descent (GD), which is called PSGD, is equivalent to the GD in the warped weight space, a space made by warping the original weight space via an appropriately designed invertible function. Second, we empirically show that PSG acting as a regularizer to a weight vector is favorable for model compression domains such as quantization and pruning. PSG reduces the gap between the weight distributions of a full-precision model and its compressed counterpart. This enables the versatile deployment of a model either as an uncompressed mode or as a compressed mode depending on the availability of resources. The experimental results on CIFAR-10/100 and ImageNet datasets show the effectiveness of the proposed PSG in both domains of pruning and quantization even for extremely low bits. The code is released in Github. Jangho Kim, KiYoon Yoo, Nojun Kwak |
NeurIPS | 1 |
| 2018 | Paraphrasing Complex Network: Network Compression via Factor TransferabstractMany researchers have sought ways of model compression to reduce the size of a deep neural network (DNN) with minimal performance degradation in order to use DNNs in embedded systems. Among the model compression methods, a method called knowledge transfer is to train a student network with a stronger teacher network. In this paper, we propose a novel knowledge transfer method which uses convolutional operations to paraphrase teacher's knowledge and to translate it for the student. This is done by two convolutional modules, which are called a paraphraser and a translator. The paraphraser is trained in an unsupervised manner to extract the teacher factors which are defined as paraphrased information of the teacher network. The translator located at the student network extracts the student factors and helps to translate the teacher factors by mimicking them. We observed that our student network trained with the proposed factor transfer method outperforms the ones trained with conventional knowledge transfer methods. Jangho Kim, Seonguk Park, Nojun Kwak |
NeurIPS | 1 |
| 2016 | Detecting Korean characters in natural scenes by alphabet detection and agglomerative character constructionabstractThis paper considers the Korean character detection problem. Unlike English where an alphabet constitutes a character, the Korean character is composed of more than two Korean alphabets, where they could be either connected or separated, relying on the Korean character font. Also, the Korean has two character structures which constitute a nested structure. These properties make the Korean character detection problem difficult. In this paper, we divide the Korean character detection problem into two subproblems, Korean alphabet detection and Korean character construction, and redefine the Korean character structures to efficiently detect Korean characters. Based on the new structures, we train four independent Korean alphabet detectors, and perform a sequential alphabet detection process with a specific detection order, to eliminate false alarms caused during detection procedure. Finally, the detected alphabets are grouped into the Korean characters by an agglomerative character construction algorithm. To evaluate our method, we carried out some experiments on a public dataset with several alternatives, and showed that our proposed Korean character detection method has outperformed other methods. Jangho Kim, Yong-Joong Kim, Daijin Kim 0001 |
SMC | 1 |