VLDB 2026 Research / reviewers in the wild / expert
Dohyung Kim 0006
dblp:143/0175-6
· DBLP profile ↗
11ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0002-6984-6325ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Token-based dynamic bit-width assignment for ViT quantization
Dohyung Kim 0006, Jaehyeon Moon, Junghyup Lee, Jeimin Jeon, Bumsub Ham |
Pattern Recognit. | 1 |
| 2025 | Subnet-Aware Dynamic Supernet Training for Neural Architecture SearchabstractN-shot neural architecture search (NAS) exploits a supernet containing all candidate subnets for a given search space. The subnets are typically trained with a static training strategy (e.g., using the same learning rate (LR) scheduler and optimizer for all subnets). This however does not consider that individual subnets have distinct characteristics, leading to two problems: (1) The supernet training is biased towards the low-complexity subnets (unfairness); (2) the momentum update in the supernet is noisy (noisy momentum). We present a dynamic supernet training technique to address these problems by adjusting the training strategy adaptive to the subnets. Specifically, we introduce a complexity-aware LR scheduler (CaLR) that controls the decay ratio of LR adaptive to the complexities of subnets, which alleviates the unfairness problem. We also present a momentum separation technique (MS). It groups the subnets with similar structural characteristics and uses a separate momentum for each group, avoiding the noisy momentum problem. Our approach can be applicable to various N-shot NAS methods with marginal cost, while improving the search performance drastically. We validate the effectiveness of our approach on various search spaces (e.g., NAS-Bench-201, Mobilenet spaces) and datasets (e.g., CIFAR-10/100, ImageNet). Jeimin Jeon, Youngmin Oh 0001, Junghyup Lee, Donghyeon Baek, Dohyung Kim 0006, Chanho Eom, Bumsub Ham |
CVPR | 5 |
| 2025 | Scheduling Weight Transitions for Quantization-Aware Training
Junghyup Lee, Jeimin Jeon, Dohyung Kim 0006, Bumsub Ham |
ICCV | 3 |
| 2024 | Instance-Aware Group Quantization for Vision TransformersabstractPost-training quantization (PTQ) is an efficient model compression technique that quantizes a pretrained full-precision model using only a small calibration set of unla-beled samples without retraining. PTQ methods for convo-lutional neural networks (CNNs) provide quantization re-sults comparable to full-precision counterparts. Directly applying them to vision transformers (ViTs), however, in-curs severe performance degradation, mainly due to the dif-ferences in architectures between CNNs and ViTs. In par-ticular, the distribution of activations for each channel vary drastically according to input instances, making PTQ meth-ods for CNNs inappropriate for ViTs. To address this, we in-troduce instance-aware group quantization for ViTs (IGQ-ViT). To this end, we propose to split the channels of acti-vation maps into multiple groups dynamically for each in-put instance, such that activations within each group share similar statistical properties. We also extend our scheme to quantize softmax attentions across tokens. In addition, the number of groups for each layer is adjusted to minimize the discrepancies between predictions from quantized and full-precision models, under a bit-operation (BOP) constraint. We show extensive experimental results on image classification, object detection, and instance segmentation, with various transformer architectures, demonstrating the effectiveness of our approach. Jaehyeon Moon, Dohyung Kim 0006, Junyong Cheon, Bumsub Ham |
CVPR | 2 |
| 2024 | Toward INT4 Fixed-Point Training via Exploring Quantization Error for Gradients
Dohyung Kim 0006, Junghyup Lee, Jeimin Jeon, Jaehyeon Moon, Bumsub Ham |
ECCV (70) | 1 |
| 2023 | Camera-Driven Representation Learning for Unsupervised Domain Adaptive Person Re-identificationabstractWe present a novel unsupervised domain adaption method for person re-identification (reID) that generalizes a model trained on a labeled source domain to an unlabeled target domain. We introduce a camera-driven curriculum learning (CaCL) framework that leverages camera labels of person images to transfer knowledge from source to target domains progressively. To this end, we divide target domain dataset into multiple subsets based on the camera labels, and initially train our model with a single subset (i.e., images captured by a single camera). We then gradually exploit more subsets for training, according to a curriculum sequence obtained with a camera-driven scheduling rule. The scheduler considers maximum mean discrepancies (MMD) between each subset and the source domain dataset, such that the subset closer to the source domain is exploited earlier within the curriculum. For each curriculum sequence, we generate pseudo labels of person images in a target domain to train a reID model in a supervised way. We have observed that the pseudo labels are highly biased toward cameras, suggesting that person images obtained from the same camera are likely to have the same pseudo labels, even for different IDs. To address the camera bias problem, we also introduce a camera-diversity (CD) loss encouraging person images of the same pseudo label, but captured across various cameras, to involve more for discriminative feature learning, providing person representations robust to inter-camera variations. Experimental results on standard benchmarks, including real-to-real and synthetic-to-real scenarios, demonstrate the effectiveness of our framework. Dohyung Kim 0006, Younghoon Shin, Yongsang Yoon, Bumsub Ham |
ICCV | 3 |
| 2022 | Learning Semantic Correspondence Exploiting an Object-Level PriorabstractWe address the problem of semantic correspondence, that is, establishing a dense flow field between images depicting different instances of the same object or scene category. We propose to use images annotated with binary foreground masks and subjected to synthetic geometric deformations to train a convolutional neural network (CNN) for this task. Using these masks as part of the supervisory signal provides an object-level prior for the semantic correspondence task and offers a good compromise between semantic flow methods, where the amount of training data is limited by the cost of manually selecting point correspondences, and semantic alignment ones, where the regression of a single global geometric transformation between images may be sensitive to image-specific details such as background clutter. We propose a new CNN architecture, dubbed SFNet, which implements this idea. It leverages a new and differentiable version of the argmax function for end-to-end training, with a loss that combines mask and flow consistency with smoothness terms. Experimental results demonstrate the effectiveness of our approach, which significantly outperforms the state of the art on standard benchmarks. Junghyup Lee, Dohyung Kim 0006, Wonkyung Lee, Jean Ponce, Bumsub Ham |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Network Quantization With Element-Wise Gradient ScalingabstractNetwork quantization aims at reducing bit-widths of weights and/or activations, particularly important for implementing deep neural networks with limited hardware resources. Most methods use the straight-through estimator (STE) to train quantized networks, which avoids a zero-gradient problem by replacing a derivative of a discretizer (i.e., a round function) with that of an identity function. Although quantized networks exploiting the STE have shown decent performance, the STE is sub-optimal in that it simply propagates the same gradient without considering discretization errors between inputs and outputs of the discretizer. In this paper, we propose an element-wise gradient scaling (EWGS), a simple yet effective alternative to the STE, training a quantized network better than the STE in terms of stability and accuracy. Given a gradient of the discretizer output, EWGS adaptively scales up or down each gradient element, and uses the scaled gradient as the one for the discretizer input to train quantized networks via backpropagation. The scaling is performed depending on both the sign of each gradient element and an error between the continuous input and discrete output of the discretizer. We adjust a scaling factor adaptively using Hessian information of a network. We show extensive experimental results on the image classification datasets, including CIFAR-10 and ImageNet, with diverse network architectures under a wide range of bit-width settings, demonstrating the effectiveness of our method. Junghyup Lee, Dohyung Kim 0006, Bumsub Ham |
CVPR | 2 |
| 2021 | Distance-aware QuantizationabstractWe address the problem of network quantization, that is, reducing bit-widths of weights and/or activations to lighten network architectures. Quantization methods use a rounding function to map full-precision values to the nearest quantized ones, but this operation is not differentiable. There are mainly two approaches to training quantized networks with gradient-based optimizers. First, a straight-through estimator (STE) replaces the zero derivative of the rounding with that of an identity function, which causes a gradient mismatch problem. Second, soft quantizers approximate the rounding with continuous functions at training time, and exploit the rounding for quantization at test time. This alleviates the gradient mismatch, but causes a quantizer gap problem. We alleviate both problems in a unified framework. To this end, we introduce a novel quantizer, dubbed a distance-aware quantizer (DAQ), that mainly consists of a distance-aware soft rounding (DASR) and a temperature controller. To alleviate the gradient mismatch problem, DASR approximates the discrete rounding with the kernel soft argmax, which is based on our insight that the quantization can be formulated as a distance-based assignment problem between full-precision values and quantized ones. The controller adjusts the temperature parameter in DASR adaptively according to the input, addressing the quantizer gap problem. Experimental results on standard benchmarks show that DAQ outperforms the state of the art significantly for various bit-widths without bells and whistles. Dohyung Kim 0006, Junghyup Lee, Bumsub Ham |
ICCV | 1 |
| 2020 | Learning with Privileged Information for Efficient Image Super-Resolution
Wonkyung Lee, Junghyup Lee, Dohyung Kim 0006, Bumsub Ham |
ECCV (24) | 3 |
| 2019 | SFNet: Learning Object-Aware Semantic CorrespondenceabstractWe address the problem of semantic correspondence, that is, establishing a dense flow field between images depicting different instances of the same object or scene category. We propose to use images annotated with binary foreground masks and subjected to synthetic geometric deformations to train a convolutional neural network (CNN) for this task. Using these masks as part of the supervisory signal offers a good compromise between semantic flow methods, where the amount of training data is limited by the cost of manually selecting point correspondences, and semantic alignment ones, where the regression of a single global geometric transformation between images may be sensitive to image-specific details such as background clutter. We propose a new CNN architecture, dubbed SFNet, which implements this idea. It leverages a new and differentiable version of the argmax function for end-to-end training, with a loss that combines mask and flow consistency with smoothness terms. Experimental results demonstrate the effectiveness of our approach, which significantly outperforms the state of the art on standard benchmarks. Junghyup Lee, Dohyung Kim 0006, Jean Ponce, Bumsub Ham |
CVPR | 2 |