Jihwan Bang

dblp:221/4643 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
14since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 4 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Computer networks · 1
YearPublicationVenuePosition
2025 CIFLEX: Contextual Instruction Flow for Sub-task Execution in Multi-Turn Interactions with a Single On-Device LLM
abstract
We present CIFLEX (Contextual Instruction FLow with EXecution), a novel execution system for efficient sub-task handling in multiturn interactions with a single on-device large language model (LLM).As LLMs become increasingly capable, a single model is expected to handle diverse sub-tasks that more effectively and comprehensively support answering user requests.Naive approach reprocesses the entire conversation context when switching between main and sub-tasks (e.g., query rewriting, summarization), incurring significant computational overhead.CIFLEX mitigates this overhead by reusing the key-value (KV) cache from the main task and injecting only task-specific instructions into isolated side paths.After sub-task execution, the model rolls back to the main path via cached context, thereby avoiding redundant prefill computation.To support sub-task selection, we also develop a hierarchical classification strategy tailored for small-scale models, decomposing multi-choice decisions into binary ones.Experiments show that CIFLEX significantly reduces computational costs without degrading task performance, enabling scalable and efficient multitask dialogue on-device.
Juntae Lee, Jihwan Bang, Seunghan Yang, Simyung Chang
EMNLP2
2025 Learning Contextual Retrieval for Robust Conversational Search
abstract
Effective conversational search demands a deep understanding of user intent across multiple dialogue turns.Users frequently use abbreviations and shift topics in the middle of conversations, posing challenges for conventional retrievers.While query rewriting techniques improve clarity, they often incur significant computational cost due to additional autoregressive steps.Moreover, although LLMbased retrievers demonstrate strong performance, they are not explicitly optimized to track user intent in multi-turn settings, often failing under topic drift or contextual ambiguity.To address these limitations, we propose ContextualRetriever, a novel LLM-based retriever that directly incorporates conversational context into the retrieval process.Our approach introduces: (1) a context-aware embedding mechanism that highlights the current query within the dialogue history; (2) intent-guided supervision based on high-quality rewritten queries; and (3) a training strategy that preserves the generative capabilities of the base LLM.Extensive evaluations across multiple conversational search benchmarks demonstrate that ContextualRetriever significantly outperforms existing methods while incurring no additional inference overhead.
Seunghan Yang, Juntae Lee, Jihwan Bang, Kyuhong Shim, Simyung Chang
EMNLP3
2025 RA-TTA: Retrieval-Augmented Test-Time Adaptation for Vision-Language Models
abstract
Vision-language models (VLMs) are known to be susceptible to distribution shifts between pre-training data and test data, and test-time adaptation (TTA) methods for VLMs have been proposed to mitigate the detrimental impact of the distribution shifts. However, the existing methods solely rely on the internal knowledge encoded within the model parameters, which are constrained to pre-training data. To complement the limitation of the internal knowledge, we propose **Retrieval-Augmented-TTA (RA-TTA)** for adapting VLMs to test distribution using **external** knowledge obtained from a web-scale image database. By fully exploiting the bi-modality of VLMs, RA-TTA **adaptively** retrieves proper external images for each test image to refine VLMs' predictions using the retrieved external images, where fine-grained **text descriptions** are leveraged to extend the granularity of external knowledge. Extensive experiments on 17 datasets demonstrate that the proposed RA-TTA outperforms the state-of-the-art methods by 3.01-9.63\% on average.
Youngjun Lee, Junhyeok Kang, Jihwan Bang, Hwanjun Song, Jae-Gil Lee 0001
ICLR4
2024 Adaptive Shortcut Debiasing for Online Continual Learning
abstract
We propose a novel framework DropTop that suppresses the shortcut bias in online continual learning (OCL) while being adaptive to the varying degree of the shortcut bias incurred by continuously changing environment. By the observed high-attention property of the shortcut bias, highly-activated features are considered candidates for debiasing. More importantly, resolving the limitation of the online environment where prior knowledge and auxiliary data are not ready, two novel techniques---feature map fusion and adaptive intensity shifting---enable us to automatically determine the appropriate level and proportion of the candidate shortcut features to be dropped. Extensive experiments on five benchmark datasets demonstrate that, when combined with various OCL algorithms, DropTop increases the average accuracy by up to 10.4% and decreases the forgetting by up to 63.2%.
Dongmin Park, Yooju Shin, Jihwan Bang, Hwanjun Song, Jae-Gil Lee 0001
AAAI4
2024 Crayon: Customized On-Device LLM via Instant Adapter Blending and Edge-Server Hybrid Inference
abstract
The customization of large language models (LLMs) for user-specified tasks gets important.However, maintaining all the customized LLMs on cloud servers incurs substantial memory and computational overheads, and uploading user data can also lead to privacy concerns.Ondevice LLMs can offer a promising solution by mitigating these issues.Yet, the performance of on-device LLMs is inherently constrained by the limitations of small-scaled models.To overcome these restrictions, we first propose Crayon, a novel approach for on-device LLM customization.Crayon begins by constructing a pool of diverse base adapters, and then we instantly blend them into a customized adapter without extra training.In addition, we develop a device-server hybrid inference strategy, which deftly allocates more demanding queries or non-customized tasks to a larger, more capable LLM on a server.This ensures optimal performance without sacrificing the benefits of on-device customization.We carefully craft a novel benchmark from multiple questionanswer datasets, and show the efficacy of our method in the LLM customization. *The authors contribute equally.
Jihwan Bang, Juntae Lee, Kyuhong Shim, Seunghan Yang, Simyung Chang
ACL (1)1
2024 Active Prompt Learning in Vision Language Models
abstract
Pre-trained Vision Language Models (VLMs) have demonstrated notable progress in various zero-shot tasks, such as classification and retrieval. Despite their performance, because improving performance on new tasks requires task-specific knowledge, their adaptation is essential. While labels are needed for the adaptation, acquiring them is typically expensive. To overcome this challenge, active learning, a method of achieving a high performance by obtaining labels for a small number of samples from experts, has been studied. Active learning primarily focuses on selecting unlabeled samples for labeling and leveraging them to train models. In this study, we pose the question, “how can the pre-trained VLMs be adapted under the active learning framework?” In response to this inquiry, we observe that (1) simply applying a conventional active learning framework to pre-trained VLMs even may degrade performance compared to random selection because of the class imbalance in labeling candidates, and (2) the knowledge of VLMs can provide hints for achieving the balance before labeling. Based on these observations, we devise a novel active learning framework for VLMs, denoted as PCB. To assess the effectiveness of our approach, we conduct experiments on seven different real-world datasets, and the results demonstrate that PCB surpasses conventional active learning and random sampling methods. Code is available at https://github.com/kaist-dmlab/pcb.
Jihwan Bang, Sumyeong Ahn, Jae-Gil Lee 0001
CVPR1
2024 One Size Fits All for Semantic Shifts: Adaptive Prompt Tuning for Continual Learning
abstract
In real-world continual learning (CL) scenarios, tasks often exhibit intricate and unpredictable semantic shifts, posing challenges for fixed prompt management strategies which are tailored to only handle semantic shifts of uniform degree (i.e., uniformly mild or uniformly abrupt). To address this limitation, we propose an adaptive prompting approach that effectively accommodates semantic shifts of varying degree where mild and abrupt shifts are mixed. AdaPromptCL employs the assign-and-refine semantic grouping mechanism that dynamically manages prompt groups in accordance with the semantic similarity between tasks, enhancing the quality of grouping through continuous refinement. Our experiment results demonstrate that AdaPromptCL outperforms existing prompting methods by up to 21.3%, especially in the benchmark datasets with diverse semantic shifts between tasks.
Susik Yoon, Dongmin Park, Youngjun Lee, Hwanjun Song, Jihwan Bang, Jae-Gil Lee 0001
ICML6
2024 Prompt-guided DETR with RoI-pruned masked attention for open-vocabulary object detection
Hwanjun Song, Jihwan Bang
Pattern Recognit.2
2023 Generating Instance-level Prompts for Rehearsal-free Continual Learning
abstract
We introduce Domain-Adaptive Prompt (DAP), a novel method for continual learning using Vision Transformers (ViT). Prompt-based continual learning has recently gained attention due to its rehearsal-free nature. Currently, the prompt pool, which is suggested by prompt-based continual learning, is key to effectively exploiting the frozen pretrained ViT backbone in a sequence of tasks. However, we observe that the use of a prompt pool creates a domain scalability problem between pre-training and continual learning. This problem arises due to the inherent encoding of group-level instructions within the prompt pool. To address this problem, we propose DAP, a pool-free approach that generates a suitable prompt in an instance-level manner at inference time. We optimize an adaptive prompt generator that creates instance-specific fine-grained instructions required for each input, enabling enhanced model plasticity and reduced forgetting. Our experiments on seven datasets with varying degrees of domain similarity to ImageNet demonstrate the superiority of DAP over state-of-the-art prompt-based methods. Code is publicly available at https://github.com/naver-ai/dap-cl.
Dahuin Jung, Dongyoon Han, Jihwan Bang, Hwanjun Song
ICCV3
2023 Online Boundary-Free Continual Learning by Scheduled Data Prior
Hyunseo Koh, Minhyuk Seo, Jihwan Bang, Hwanjun Song, Deokki Hong, Seulki Park, Jung-Woo Ha 0001
ICLR3
2023 Self-Supervised Set Representation Learning for Unsupervised Meta-Learning
Dong Bok Lee, Seanie Lee, Kenji Kawaguchi, Yunji Kim, Jihwan Bang, Jung-Woo Ha 0001, Sung Ju Hwang
ICLR5
2022 Online Continual Learning on a Contaminated Data Stream with Blurry Task Boundaries
abstract
Learning under a continuously changing data distribution with incorrect labels is a desirable real-world problem yet challenging. A large body of continual learning (CL) methods, however, assumes data streams with clean labels, and online learning scenarios under noisy data streams are yet underexplored. We consider a more practical CL task setup of an online learning from blurry data stream with corrupted labels, where existing CL methods struggle. To address the task, we first argue the importance of both diversity and purity of examples in the episodic memory of continual learning models. To balance diversity and purity in the episodic memory, we propose a novel strategy to manage and use the memory by a unified approach of label noise aware diverse sampling and robust learning with semi-supervised learning. Our empirical validations on four real-world or synthetic noise datasets (CI-FAR10 and 100, mini-WebVision, and Food-101N) exhibit that our method significantly outperforms prior arts in this realistic and challenging continual learning scenario. Code and data splits are available in https://github.com/clovaai/puridiver.
Jihwan Bang, Hyunseo Koh, Seulki Park, Hwanjun Song, Jung-Woo Ha 0001
CVPR1
2022 Meta-Query-Net: Resolving Purity-Informativeness Dilemma in Open-set Active Learning
abstract
Unlabeled data examples awaiting annotations contain open-set noise inevitably. A few active learning studies have attempted to deal with this open-set noise for sample selection by filtering out the noisy examples. However, because focusing on the purity of examples in a query set leads to overlooking the informativeness of the examples, the best balancing of purity and informativeness remains an important question. In this paper, to solve this purity-informativeness dilemma in open-set active learning, we propose a novel Meta-Query-Net (MQ-Net) that adaptively finds the best balancing between the two factors. Specifically, by leveraging the multi-round property of active learning, we train MQ-Net using a query set without an additional validation set. Furthermore, a clear dominance relationship between unlabeled examples is effectively captured by MQ-Net through a novel skyline regularization. Extensive experiments on multiple open-set active learning scenarios demonstrate that the proposed MQ-Net achieves 20.14% improvement in terms of accuracy, compared with the state-of-the-art methods.
Dongmin Park, Yooju Shin, Jihwan Bang, Youngjun Lee, Hwanjun Song, Jae-Gil Lee 0001
NeurIPS3
2021 Rainbow Memory: Continual Learning With a Memory of Diverse Samples
abstract
Continual learning is a realistic learning scenario for AI models. Prevalent scenario of continual learning, however, assumes disjoint sets of classes as tasks and is less realistic rather artificial. Instead, we focus on ‘blurry’ task boundary; where tasks shares classes and is more realistic and practical. To address such task, we argue the importance of diversity of samples in an episodic memory. To enhance the sample diversity in the memory, we propose a novel memory management strategy based on per-sample classification uncertainty and data augmentation, named Rainbow Memory (RM). With extensive empirical validations on MNIST, CIFAR10, CIFAR100, and ImageNet datasets, we show that the proposed method significantly improves the accuracy in blurry continual learning setups, outperforming state of the arts by large margins despite its simplicity. Code and data splits will be available in https://github.com/clovaai/rainbow-memory.
Jihwan Bang, Heesu Kim, Young Joon Yoo, Jung-Woo Ha 0001
CVPR1
2020 SINet: Extreme Lightweight Portrait Segmentation Networks with Spatial Squeeze Modules and Information Blocking Decoder
abstract
Designing a lightweight and robust portrait segmentation algorithm is an important task for a wide range of face applications. However, the problem has been considered as a subset of the object segmentation and less handled in this field. Obviously, portrait segmentation has its unique requirements. First, because the portrait segmentation is performed in the middle of a whole process, it requires extremely lightweight models. Second, there has not been any public datasets in this domain that contain a sufficient number of images. To solve the first problem, we introduce the new extremely lightweight portrait segmentation model SINet, containing an information blocking decoder and spatial squeeze modules. The information blocking decoder uses confidence estimation to recover local spatial information without spoiling global consistency. The spatial squeeze module uses multiple receptive fields to cope with various sizes of consistency. To tackle the second problem, we propose a simple method to create additional portrait segmentation data, which can improve accuracy. In our qualitative and quantitative analysis on the EG1800 dataset, we show that our method outperforms various existing lightweight models. Our method reduces the number of parameters from 2.1M to 86.9K (around 95.9% reduction), while maintaining the accuracy under an 1% margin from the state-of-the-art method. We also show our model is successfully executed on a real mobile device with 100.6 FPS. In addition, we demonstrate that our method can be used for general semantic segmentation on the Cityscapes dataset. The code and dataset are available in https://github.com/HYOJINPARK/ExtPortraitSeg.
Hyojin Park 0001, Lars Lowe Sjösund, Young Joon Yoo, Nicolas Monet, Jihwan Bang, Nojun Kwak
WACV5
2018 Incentivizing Hosts via Multilateral Cooperation in User-Provided Networks: A Fluid Shapley Value Approach
abstract
Successful operation of User-Provided Networks (UPN) requires that both of Internet Service Provider (ISP) and self network-operating users (hosts) cooperate appropriately in terms of resource sharing and pricing strategy since ISP and hosts have a multilateral reliance on each other with respect to virtual infrastructure expansion and Internet connectivity. However, it has been underexplored whether such cooperation provides sufficient incentive to ISP and hosts under a setup where ISP and hosts are fully included, having a high dependence on how to cooperate and how to distribute the resulting cooperation worth. In this paper, we model a market of UPN, consisting of ISP, hosts, and clients via game theory, where we model various heterogeneities in terms of (i) willingness to pay and mobility pattern of clients, (ii) hosts' QoS, and (iii) type of cooperation among ISP and hosts. The key technical challenges lie in the natural mixture of cooperative and non-cooperative game theoretic angles, where the worth function---one of the crucial components in coalitional game theory---comes from the equilibrium of an embedded, non-cooperative two-stage dynamic game. We consider the Shapley value as a mechanism of revenue sharing and overcome its hardness in characterization by taking the fluid limit when the number of hosts and clients is large. Our analytical studies reveal useful implications that in UPN when and how much economic benefits can be given to the players and when they maintain their grand coalition under what conditions, referred to as stability.
Hyojung Lee, Jihwan Bang, Yung Yi
MobiHoc2