EDBT 2026 Demo / reviewers in the wild / expert
Juntae Lee
dblp:18/8673 · also Jun-Tae Lee
· DBLP profile ↗
20ranked-venue papers
11as first author
14since 2021 · last 2025
0000-0003-2953-8851ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 9 first-author · 9 since 2021Artificial intelligence and machine learning · 13 · 7 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Tripartite Weight-Space Ensemble for Few-Shot Class-Incremental LearningabstractFew-shot class incremental learning (FSCIL) enables the continual learning of new concepts with only a few training examples. In FSCIL, the model undergoes substantial updates, making it prone to forgetting previous concepts and overfitting to the limited new examples. Most recent trend is typically to disentangle the learning of the representation from the classification head of the model. A well-generalized feature extractor on the base classes (many examples and many classes) is learned, and then fixed during incremental learning. Arguing that the fixed feature extractor restricts the model’s adaptability to new classes, we introduce a novel FSCIL method to effectively address catastrophic forgetting and overfitting issues. Our method enables to seamlessly update the entire model with a few examples. We mainly propose a tripartite weight-space ensemble (Tri-WE). Tri-WE interpolates the base, immediately previous, and current models in weight-space, especially for the classification heads of the models. Then, it collaboratively maintains knowledge from the base and previous models. In addition, we recognize the challenges of distilling generalized representations from the previous model from scarce data. Hence, we suggest a regularization loss term using amplified data knowledge distillation. Simply intermixing the few-shot data, we can produce richer data enabling the distillation of critical knowledge from the previous model. Consequently, we attain state-of-the-art results on the miniImageNet, CUB200, and CIFAR100 datasets. Juntae Lee, Munawar Hayat, Sungrack Yun |
CVPR | 1 |
| 2025 | CIFLEX: Contextual Instruction Flow for Sub-task Execution in Multi-Turn Interactions with a Single On-Device LLMabstractWe present CIFLEX (Contextual Instruction FLow with EXecution), a novel execution system for efficient sub-task handling in multiturn interactions with a single on-device large language model (LLM).As LLMs become increasingly capable, a single model is expected to handle diverse sub-tasks that more effectively and comprehensively support answering user requests.Naive approach reprocesses the entire conversation context when switching between main and sub-tasks (e.g., query rewriting, summarization), incurring significant computational overhead.CIFLEX mitigates this overhead by reusing the key-value (KV) cache from the main task and injecting only task-specific instructions into isolated side paths.After sub-task execution, the model rolls back to the main path via cached context, thereby avoiding redundant prefill computation.To support sub-task selection, we also develop a hierarchical classification strategy tailored for small-scale models, decomposing multi-choice decisions into binary ones.Experiments show that CIFLEX significantly reduces computational costs without degrading task performance, enabling scalable and efficient multitask dialogue on-device. Juntae Lee, Jihwan Bang, Seunghan Yang, Simyung Chang |
EMNLP | 1 |
| 2025 | Learning Contextual Retrieval for Robust Conversational SearchabstractEffective conversational search demands a deep understanding of user intent across multiple dialogue turns.Users frequently use abbreviations and shift topics in the middle of conversations, posing challenges for conventional retrievers.While query rewriting techniques improve clarity, they often incur significant computational cost due to additional autoregressive steps.Moreover, although LLMbased retrievers demonstrate strong performance, they are not explicitly optimized to track user intent in multi-turn settings, often failing under topic drift or contextual ambiguity.To address these limitations, we propose ContextualRetriever, a novel LLM-based retriever that directly incorporates conversational context into the retrieval process.Our approach introduces: (1) a context-aware embedding mechanism that highlights the current query within the dialogue history; (2) intent-guided supervision based on high-quality rewritten queries; and (3) a training strategy that preserves the generative capabilities of the base LLM.Extensive evaluations across multiple conversational search benchmarks demonstrate that ContextualRetriever significantly outperforms existing methods while incurring no additional inference overhead. Seunghan Yang, Juntae Lee, Jihwan Bang, Kyuhong Shim, Simyung Chang |
EMNLP | 2 |
| 2024 | Crayon: Customized On-Device LLM via Instant Adapter Blending and Edge-Server Hybrid InferenceabstractThe customization of large language models (LLMs) for user-specified tasks gets important.However, maintaining all the customized LLMs on cloud servers incurs substantial memory and computational overheads, and uploading user data can also lead to privacy concerns.Ondevice LLMs can offer a promising solution by mitigating these issues.Yet, the performance of on-device LLMs is inherently constrained by the limitations of small-scaled models.To overcome these restrictions, we first propose Crayon, a novel approach for on-device LLM customization.Crayon begins by constructing a pool of diverse base adapters, and then we instantly blend them into a customized adapter without extra training.In addition, we develop a device-server hybrid inference strategy, which deftly allocates more demanding queries or non-customized tasks to a larger, more capable LLM on a server.This ensures optimal performance without sacrificing the benefits of on-device customization.We carefully craft a novel benchmark from multiple questionanswer datasets, and show the efficacy of our method in the LLM customization. *The authors contribute equally. Jihwan Bang, Juntae Lee, Kyuhong Shim, Seunghan Yang, Simyung Chang |
ACL (1) | 2 |
| 2023 | Scalable Weight Reparametrization for Efficient Transfer LearningabstractThis paper proposes a novel, efficient transfer learning method, called Scalable Weight Reparametrization (SWR) that is efficient and effective for multiple downstream tasks. Efficient transfer learning involves utilizing a pre-trained model trained on a larger dataset and repurposing it for downstream tasks with the aim of maximizing the reuse of the pre-trained model. However, previous works have led to an increase in updated parameters and task-specific modules, resulting in more computations, especially for tiny models. Additionally, there has been no practical consideration for controlling the number of updated parameters. To address these issues, we suggest learning a policy network that can decide where to reparametrize the pre-trained model, while adhering to a given constraint for the number of updated parameters. The policy network is only used during the transfer learning process and not afterward. As a result, our approach attains state-of-the-art performance in a proposed multi-lingual keyword spotting and a standard benchmark, ImageNet-to-Sketch, while requiring zero additional computations and significantly fewer additional parameters. Byeonggeun Kim, Juntae Lee, Seunghan Yang, Simyung Chang |
ICASSP | 2 |
| 2023 | Few-Shot Common Action Localization via Cross-Attentional Fusion of Context and Temporal DynamicsabstractThe goal of this paper is to localize action instances in a long untrimmed query video using just meager trimmed support videos representing a common action whose class information is not given. In this task, it is crucial to mine reliable temporal cues representing a common action from handful support videos. In our work, we develop an attention mechanism using cross-correlation. Based on this cross-attention, we first transform the support videos into query video’s context to emphasize query-relevant important frames, and suppress less relevant ones. Next, we summarize sub-sequences of support video frames to represent temporal dynamics in coarse temporal granularity, which is then propagated to the fine-grained support video features through the cross-attention. In each case, the cross-attentions are applied to each support video in the individual-to-all strategy to balance heterogeneity and compatibility of the support videos. In contrast, the candidate instances in the query video are lastly attended by the resulting support video features, at once. In addition, we also develop a relational classifier head based on the query and support video representations. We show the effectiveness of our work with the state-of-the-art (SOTA) performance in benchmark datasets (ActivityNet1.3 and THUMOS14), and analyze each component extensively. Juntae Lee, Mihir Jain, Sungrack Yun |
ICCV | 1 |
| 2023 | Task-Agnostic Open-Set Prototype for Few-Shot Open-Set RecognitionabstractIn few-shot open-set recognition (FSOSR), a network learns to recognize closed-set samples with a few support samples while rejecting open-set samples with no class cue. Unlike conventional OSR, the FSOSR considers more practical open worlds where a closed-set class can be selected as an open-set class in another testing (task) and vice versa. Existing FSOSR methods have commonly represented the open set with task-dependent extra modules. These modules decently handle the varied closed and open classes but accompany inevitable complexity increase. This paper shows that a single open-set prototype can represent open-set samples when it satisfies a specific relation in metric space: closest to open-set, and simultaneously second nearest to close-set. We propose a task-agnostic open-set prototype with distance scaling factors and design loss terms. We extensively analyze the proposed components to demonstrate their importance. Our method achieves state-of-the-art results on miniImageNet and tieredImageNet, respectively, without task-dependent extra modules. Byeonggeun Kim, Juntae Lee, Kyuhong Shim, Simyung Chang |
ICIP | 2 |
| 2023 | Multi-Scale Temporal Feature Fusion for Few-Shot Action RecognitionabstractThe aim of this paper is to recognize actions of interest that are given by a few support videos in testing (query) videos. The focus of our approach is to develop a novel temporal enrichment module where the features describing local temporal contexts in videos are enhanced by collaboratively merging important information in frame-level (no temporal context) features. We call this module a multi-scale temporal feature fusion (MSTFF) module. Utilizing multiple MSTFF modules varying the scope of local temporal context extraction, we can obtain discriminative video representation which is crucial in the few-shot tasks where support videos are not sufficient to describe an action class. For stable learning of a model with MSTFF and the performance boost, we also learn a local temporal context-level auxiliary classifier in parallel with the main classifier. We analyze the proposed components to demonstrate their importance. We achieve state-of-the-art on three few-shot action recognition benchmarks: Something-Something V2 (SSv2), HMDB51, and Kinetics. Juntae Lee, Sungrack Yun |
ICIP | 1 |
| 2022 | Multi-Head Modularization to Leverage Generalization Capability in Multi-Modal NetworksabstractIt has been crucial to leverage the rich information of multiple modalities in many tasks. Existing works have tried to design multi-modal networks with descent multi-modal fusion modules. Instead, we focus on improving generalization capability of multi-modal networks, especially the fusion module. Viewing the multi-modal data as different projections of information, we first observe that bad projection can cause poor generalization behaviors of multi-modal networks. Then, motivated by well-generalized network's low sensitivity to perturbation, we propose a novel multi-modal training method, multi-head modularization (MHM). We modularize a multi-modal network as a series of uni-modal embedding, multi-modal embedding, and task-specific head modules. Also, for training, we exploit multiple head modules learned with different datasets, swapping each other. From this, we can make the multi-modal embedding module robust to all the heads with different generalization behaviors. In testing phase, we select one of the head modules not to increase the computational cost. Owing to the perturbation of head modules, though including one selected head, the deployed network is more well-generalized compared to the simply end-to-end learned. We verify the effectiveness of MHM on various multi-modal tasks. We use the state-of-the-art methods as baselines, and show notable performance gain for all the baselines. Juntae Lee, Hyunsin Park, Sungrack Yun, Simyung Chang |
AAAI | 1 |
| 2022 | Variational On-the-Fly PersonalizationabstractWith the development of deep learning (DL) technologies, the demand for DL-based services on personal devices, such as mobile phones, also increases rapidly. In this paper, we propose a novel personalization method, Variational On-the-Fly Personalization. Compared to the conventional personalization methods that require additional fine-tuning with personal data, the proposed method only requires forwarding a handful of personal data on-the-fly. Assuming even a single personal data can convey the characteristics of a target person, we develop the variational hyper-personalizer to capture the weight distribution of layers that fits the target person. In the testing phase, the hyper-personalizer estimates the model’s weights on-the-fly based on personality by forwarding only a small amount of (even a single) personal enrollment data. Hence, the proposed method can perform the personalization without any training software platform and additional cost in the edge device. In experiments, we show our approach can effectively generate reliable personalized models via forwarding (not back-propagating) a handful of samples. Jangho Kim, Juntae Lee, Simyung Chang, Nojun Kwak |
ICML | 2 |
| 2022 | Domain Generalization with Relaxed Instance Frequency-wise Normalization for Multi-device Acoustic Scene ClassificationabstractWhile using two-dimensional convolutional neural networks (2D-CNNs) in image processing, it is possible to manipulate domain information using channel statistics, and instance normalization has been a promising way to get domain-invariant features. However, unlike image processing, we analyze that domain-relevant information in an audio feature is dominant in frequency statistics rather than channel statistics. Motivated by our analysis, we introduce Relaxed Instance Frequency-wise Normalization (RFN): a plug-and-play, explicit normalization module along the frequency axis which can eliminate instance-specific domain discrepancy in an audio feature while relaxing undesirable loss of useful discriminative information. Empirically, simply adding RFN to networks shows clear margins compared to previous domain generalization approaches on acoustic scene classification and yields improved robustness for multiple audio devices. Especially, the proposed RFN won the DCASE2021 challenge TASK1A, low-complexity acoustic scene classification with multiple devices, with a clear margin, and RFN is an extended work of our technical report. Byeonggeun Kim, Seunghan Yang, Jangho Kim, Hyunsin Park, Juntae Lee, Simyung Chang |
INTERSPEECH | 5 |
| 2022 | Leaky Gated Cross-Attention for Weakly Supervised Multi-Modal Temporal Action LocalizationabstractAs multiple modalities sometimes have a weak complementary relationship, multi-modal fusion is not always beneficial for weakly supervised action localization. Hence, to attain the adaptive multi-modal fusion, we propose a leaky gated cross-attention mechanism. In our work, we take the multi-stage cross-attention as the baseline fusion module to obtain multi-modal features. Then, for the stages of each modality, we design gates to decide the dependency on the other modality. For each input frame, if two modalities have a strong complementary relationship, the gate selects the cross-attended feature, otherwise the non-attended feature. Also, the proposed gate allows the non-selected feature to escape through it with a small intensity, we call it leaky gate. This leaky feature makes effective regularization of the selected major feature. Therefore, our leaky gating makes cross-attention more adaptable and robust even when the modalities have a weak complementary relationship. The proposed leaky gated cross-attention provides a modality fusion module that is generally compatible with various temporal action localization methods. To show its effectiveness, we do extensive experimental analysis and apply the proposed method to boost the performance of the state-of-the-art methods on two benchmark datasets (ActivityNet1.2 and THUMOS14). Juntae Lee, Sungrack Yun, Mihir Jain |
WACV | 1 |
| 2021 | Efficient Action Recognition via Dynamic Knowledge PropagationabstractEfficient action recognition has become crucial to extend the success of action recognition to many real-world applications. Contrary to most existing methods, which mainly focus on selecting salient frames to reduce the computation cost, we focus more on making the most of the selected frames. To this end, we employ two networks of different capabilities that operate in tandem to efficiently recognize actions. Given a video, the lighter network processes more frames while the heavier one only processes a few. In order to enable the effective interaction between the two, we propose dynamic knowledge propagation based on a cross-attention mechanism. This is the main component of our framework that is essentially a student-teacher architecture, but as the teacher model continues to interact with the student model during inference, we call it a dynamic student-teacher framework. Through extensive experiments, we demonstrate the effectiveness of each component of our framework. Our method outperforms competing state-of-the-art methods on two video datasets: ActivityNet-v1.3 and Mini-Kinetics. Hanul Kim 0001, Mihir Jain, Juntae Lee, Sungrack Yun, Fatih Porikli |
ICCV | 3 |
| 2021 | Cross-Attentional Audio-Visual Fusion for Weakly-Supervised Action Localization
Juntae Lee, Mihir Jain, Hyoungwoo Park, Sungrack Yun |
ICLR | 1 |
| 2020 | Semantic Line Detection Using Mirror Attention and Comparative Ranking and Matching
Dongkwon Jin, Juntae Lee, Chang-Su Kim 0001 |
ECCV (20) | 2 |
| 2019 | Image Aesthetic Assessment Based on Pairwise Comparison A Unified Approach to Score Regression, Binary Classification, and PersonalizationabstractWe propose a unified approach to three tasks of aesthetic score regression, binary aesthetic classification, and personalized aesthetics. First, we develop a comparator to estimate the ratio of aesthetic scores for two images. Then, we construct a pairwise comparison matrix for multiple reference images and an input image, and predict the aesthetic score of the input via the eigenvalue decomposition of the matrix. By varying the reference images, the proposed algorithm can be used for binary aesthetic classification and personalized aesthetics, as well as generic score regression. Experimental results demonstrate that the proposed unified algorithm provides the state-of-the-art performances in all three tasks of image aesthetics. Juntae Lee, Chang-Su Kim 0001 |
ICCV | 1 |
| 2018 | PAC-Net: Pairwise Aesthetic Comparison Network for Image Aesthetic AssessmentabstractImage aesthetic assessment is important for finding well taken and appealing photographs but is challenging due to the ambiguity and subjectivity of aesthetic criteria. We develop the pairwise aesthetic comparison network (PAC-Net), which consists of two parts: aesthetic feature extraction and pairwise feature comparison. To alleviate the ambiguity and subjectivity, we train PAC-Net to learn the relative aesthetic ranks of two images by employing a novel loss function, called aesthetic-adaptive cross entropy loss. Then, we develop simple schemes for using PAC-Net in the tasks of aesthetic ranking and aesthetic classification, respectively. Experimental results demonstrate that PAC-Net achieves the state-of-the-art performances in both the ranking and classification applications. Keunsoo Ko, Juntae Lee, Chang-Su Kim 0001 |
ICIP | 2 |
| 2018 | Photographic composition classification and dominant geometric element detection for outdoor scenes
Juntae Lee, Hanul Kim 0001, Chul Lee, Chang-Su Kim 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2017 | Semantic Line Detection and Its ApplicationsabstractSemantic lines characterize the layout of an image. Despite their importance in image analysis and scene understanding, there is no reliable research for semantic line detection. In this paper, we propose a semantic line detector using a convolutional neural network with multi-task learning, by regarding the line detection as a combination of classification and regression tasks. We use convolution and max-pooling layers to obtain multi-scale feature maps for an input image. Then, we develop the line pooling layer to extract a feature vector for each candidate line from the feature maps. Next, we feed the feature vector into the parallel classification and regression layers. The classification layer decides whether the line candidate is semant ic or not. In case of a semantic line, the regression layer determines the offset for refining the line location. Experimental results show that the proposed detector extracts semantic lines accurately and reliably. Moreover, we demonstrate that the proposed detector can be used successfully in three applications: horizon estimation, composition enhancement, and image simplification. Juntae Lee, Hanul Kim 0001, Chul Lee, Chang-Su Kim 0001 |
ICCV | 1 |
| 2014 | Depth-guided adaptive contrast enhancement using 2D histogramsabstractA novel contrast enhancement (CE) algorithm using 2-dimensional (2D) histograms, which transforms pixel values adaptively based on the depth information, is proposed in this work. In general, foreground objects convey more important visual information than background regions. Hence we assign high CE priorities to foreground pixels using the depth values and generate a depth-guided 2D histogram. Then, we stretch the gray-level differences of adjacent foreground pixels more strongly than those of adjacent background pixels. Moreover, to enhance background regions as well, we design two transformation functions for the foreground and the background separately. By combining the two functions according to pixel depths, we obtain an adaptive space-variant transformation function, which is finally used to reconstruct the output image. Experimental results show that the proposed algorithm outperforms conventional CE algorithms by enhancing salient foreground objects efficiently and preserving background details faithfully. Juntae Lee, Chulwoo Lee, Jae-Young Sim, Chang-Su Kim 0001 |
ICIP | 1 |