VLDB 2026 Research / reviewers in the wild / expert
Joonmyung Choi
dblp:317/4845
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0003-3673-8069ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Parameter-Efficient Fine-Tuning via Meta-RegularizerabstractAbstract Pre-trained vision-language models ( e.g ., CLIP) have shown impressive success in various computer vision tasks with their generalization capability. Recently, parameter-efficient fine-tuning (PEFT) approaches have been actively explored to effectively and efficiently adapt the pre-trained vision-language models to a variety of downstream tasks. However, most existing PEFT approaches suffer from a task overfitting issue since the general knowledge of the pre-trained models is forgotten while a small number of learnable parameters in soft prompts/adapters are fine-tuned on a small data set from a specific target task. Thus, we propose a P arameter- E fficient F ine- T uning via Meta - R egularization (PEFT-MetaR) to improve the generalizability of parameter-efficient fine-tuning methods for vision-language models. Specifically, PEFT-MetaR meta-learns both the regularizer and learnable parameters to harness the task-specific knowledge from the downstream tasks and task-agnostic general knowledge from the pretrained models. Further, PEFT-MetaR augments the task to generate multiple virtual tasks to alleviate the meta-overfitting. In addition, we provide the analysis to comprehend how PEFT-MetaR improves the generalizability from the perspective of the gradient alignment. Our experiments demonstrate that PEFT-MetaR improves the generalizability of parameter-efficient fine-tuning methods on various datasets. Jinyoung Park 0005, Juyeon Ko, Sanghyeok Lee, Joonmyung Choi, Hyunwoo J. Kim |
Int. J. Comput. Vis. | 4 |
| 2025 | EfficientViM: Efficient Vision Mamba with Hidden State Mixer based State Space DualityabstractFor the deployment of neural networks in resource-constrained environments, prior works have built lightweight architectures with convolution and attention for capturing local and global dependencies, respectively. Recently, the state space model (SSM) has emerged as an effective operation for global interaction with its favorable linear computational cost in the number of tokens. To harness the efficacy of SSM, we introduce Efficient Vision Mamba (EfficientViM), a novel architecture built on hidden state mixer-based state space duality (HSM-SSD) that efficiently captures global dependencies with further reduced computational cost. With the observation that the runtime of the SSD layer is driven by the linear projections on the input sequences, we redesign the original SSD layer to perform the channel mixing operation within compressed hidden states in the HSM-SSD layer. Additionally, we propose multi-stage hidden state fusion to reinforce the representation power of hidden states and provide the design to alleviate the bottleneck caused by the memory-bound operations. As a result, the EfficientViM family achieves a new state-of-the-art speed-accuracy trade-off on ImageNet-1k, offering up to a 0.7% performance improvement over the second-best model SHViT with faster speed. Further, we observe significant improvements in throughput and accuracy compared to prior works, when scaling images or employing distillation training. Code is available at https://github.com/mlvlab/EfficientViM. Sanghyeok Lee, Joonmyung Choi, Hyunwoo J. Kim |
CVPR | 2 |
| 2025 | Representation Shift: Unifying Token Compression with FlashattentionabstractTransformers have demonstrated remarkable success across vision, language, and video. Yet, increasing task complexity has led to larger models and more tokens, raising the quadratic cost of self-attention and the overhead of GPU memory access. To reduce the computation cost of self-attention, prior work has proposed token compression techniques that drop redundant or less informative tokens. Meanwhile, fused attention kernels such as FlashAttention have been developed to alleviate memory overhead by avoiding attention map construction and its associated I/O to HBM. This, however, makes it incompatible with most training-free token compression methods, which rely on attention maps to determine token importance. Here, we propose Representation Shift, a training-free, model-agnostic metric that measures the degree of change in each token's representation. This seamlessly integrates token compression with FlashAttention, without attention maps or retraining. Our method further generalizes beyond Transformers to CNNs and state space models. Extensive experiments show that Representation Shift enables effective token compression compatible with FlashAttention, yielding significant speedups of up to 5.5% and 4.4% in video-text retrieval and video QA, respectively. Code is available at https://github.com/mlvlab/Representation-Shift. Joonmyung Choi, Sanghyeok Lee, Byungoh Ko, Eunseo Kim, Jihyung Kil, Hyunwoo J. Kim |
ICCV | 1 |
| 2024 | vid-TLDR: Training Free Token merging for Light-Weight Video TransformerabstractVideo Transformers have become the prevalent solution for various video downstream tasks with superior expressive power and flexibility. However, these video transformers suffer from heavy computational costs induced by the massive number of tokens across the entire video frames, which has been the major barrier to train and deploy the model. Further, the patches irrelevant to the main contents, e.g., backgrounds, degrade the generalization performance of models. To tackle these issues, we propose training-free token merging for lightweight video Transformer (vid-TLDR) that aims to enhance the efficiency of video Transformers by merging the background tokens without additional training. For vid-TLDR, we introduce a novel approach to capture the salient regions in videos only with the attention map. Further, we introduce the saliency-aware token merging strategy by dropping the background tokens and sharpening the object scores. Our experiments show that vid-TLDR significantly mitigates the computational complexity of video Transformers while achieving competitive performance compared to the base model without vid-TLDR. Code is available at https://github.com/mlvlab/vid-TLDR. Joonmyung Choi, Sanghyeok Lee, Jaewon Chu, Minhyuk Choi, Hyunwoo J. Kim |
CVPR | 1 |
| 2024 | Multi-Criteria Token Fusion with One-Step-Ahead Attention for Efficient Vision TransformersabstractVision Transformer (ViT) has emerged as a prominent backbone for computer vision. For more efficient ViTs, re-cent works lessen the quadratic cost of the self-attention layer by pruning or fusing the redundant tokens. How-ever, these works faced the speed-accuracy trade-off caused by the loss of information. Here, we argue that token fusion needs to consider diverse relations between tokens to minimize information loss. In this paper, we propose a Multi-criteria Token Fusion (MCTF), that gradually fuses the tokens based on multi-criteria (i.e., similarity, informativeness, and size of fused tokens). Further, we utilize the one-step-ahead attention, which is the improved approach to capture the informativeness of the tokens. By training the model equipped with MCTF using a token reduction consistency, we achieve the best speed-accuracy trade-off in the image classification (ImageNet1K). Experimental re-sults prove that MCTF consistently surpasses the previous reduction methods with and without training. Specifically, DeiT-T and DeiT-S with MCTF reduce FLOPs by about 44% while improving the performance (+0.5%, and +0.3%) over the base model, respectively. We also demonstrate the applicability of MCTF in various Vision Transformers (e.g., T2T-ViT, LV-ViT), achieving at least 31% speedup without performance degradation. Code is available at https://github.com/mlvlab/MCTF. Sanghyeok Lee, Joonmyung Choi, Hyunwoo J. Kim |
CVPR | 2 |
| 2023 | MELTR: Meta Loss Transformer for Learning to Fine-tune Video Foundation ModelsabstractFoundation models have shown outstanding performance and generalization capabilities across domains. Since most studies on foundation models mainly focus on the pretraining phase, a naive strategy to minimize a single task-specific loss is adopted for fine-tuning. However, such fine-tuning methods do not fully leverage other losses that are potentially beneficial for the target task. Therefore, we propose MEta Loss TRansformer (MELTR), a plug-in module that automatically and non-linearly combines various loss functions to aid learning the target task via auxiliary learning. We formulate the auxiliary learning as a bi-level optimization problem and present an efficient optimization algorithm based on Approximate Implicit Differentiation (AID). For evaluation, we apply our framework to various video foundation models (UniVL, Violet and All-in-one), and show significant performance gain on all four downstream tasks: text-to-video retrieval, video question answering, video captioning, and multimodal sentiment analysis. Our qualitative analyses demonstrate that MELTR adequately 'transforms' individual loss functions and 'melts' them into an effective unified loss. Code is available at https://github.com/mlvlab/MELTR. Dohwan Ko, Joonmyung Choi, Hyeong Kyu Choi, Kyoung-Woon On, Byungseok Roh, Hyunwoo J. Kim |
CVPR | 2 |
| 2023 | Robust auxiliary learning with weighting function for biased dataabstractDeep neural networks easily suffer from weak generalization caused by overfitting on biased data. One popular remedy to alleviate this issue is sample reweighting methods that adaptively adjust the importance of biased samples. Separate from the effort to reduce bias, recent works show that the generalization power can be improved by auxiliary tasks. Inspired by the two lines of works, we extend the sample reweighting methods to auxiliary tasks. In this paper, we propose a novel auxiliary learning framework that improves the primary task by adaptively adjusting the weights of samples from multiple tasks rather than samples from a single task using a weighting function. The weighting function is optimized by meta-learning along the gradient of the loss for meta-data, which is a small unbiased validation data. We also present a task-activation score that indicates the correlation between the learning tendency of the training samples and meta-data samples. This score is utilized as a regularizer for meta-learning objective. Our framework can obtain powerful representations for the primary task on biased data by automatically identifying effective combinations of tasks. Our experiments demonstrate that our proposed method consistently outperforms all baselines and state-of-the-art methods on both corrupted labels and class imbalance settings. Dasol Hwang, Sojin Lee, Joonmyung Choi, Je-Keun Rhee, Hyunwoo J. Kim |
Inf. Sci. | 3 |
| 2022 | Video-Text Representation Learning via Differentiable Weak Temporal AlignmentabstractLearning generic joint representations for video and text by a supervised method requires a prohibitively substantial amount of manually annotated video datasets. As a practical alternative, a large-scale but uncurated and narrated video dataset, HowTo100M, has recently been introduced. But it is still challenging to learn joint embeddings of video and text in a self-supervised manner, due to its ambiguity and non-sequential alignment. In this paper, we propose a novel multi-modal self-supervised framework Video-Text Temporally Weak Alignment-based Contrastive Learning (VT-TWINS) to capture significant information from noisy and weakly correlated data using a variant of Dynamic Time Warping (DTW). We observe that the standard DTW inherently cannot handle weakly correlated data and only considers the globally optimal alignment path. To address these problems, we develop a differentiable DTW which also reflects local information with weak temporal alignment. Moreover, our proposed model applies a contrastive learning scheme to learn feature representations on weakly correlated data. Our extensive experiments demonstrate that VT-TWINS attains significant improvements in multi-modal representation learning and outperforms various challenging downstream tasks. Code is available at https://github.com/mlvlab/VT-Twins. Dohwan Ko, Joonmyung Choi, Juyeon Ko, Shinyeong Noh, Kyoung-Woon On, Eun-Sol Kim, Hyunwoo J. Kim |
CVPR | 2 |
| 2022 | TokenMixup: Efficient Attention-guided Token-level Data Augmentation for TransformersabstractMixup is a commonly adopted data augmentation technique for image classification. Recent advances in mixup methods primarily focus on mixing based on saliency. However, many saliency detectors require intense computation and are especially burdensome for parameter-heavy transformer models. To this end, we propose TokenMixup, an efficient attention-guided token-level data augmentation method that aims to maximize the saliency of a mixed set of tokens. TokenMixup provides ×15 faster saliency-aware data augmentation compared to gradient-based methods. Moreover, we introduce a variant of TokenMixup which mixes tokens within a single instance, thereby enabling multi-scale feature augmentation. Experiments show that our methods significantly improve the baseline models’ performance on CIFAR and ImageNet-1K, while being more efficient than previous methods. We also reach state-of-the-art performance on CIFAR-100 among from-scratch transformer models. Code is available at https://github.com/mlvlab/TokenMixup. Hyeong Kyu Choi, Joonmyung Choi, Hyunwoo J. Kim |
NeurIPS | 2 |