Simyung Chang

dblp:206/6540 · DBLP profile ↗
← Back
25ranked-venue papers
6as first author
18since 2021 · last 2025
0000-0001-7750-191XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 4 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 12 since 2021
YearPublicationVenuePosition
2025 CIFLEX: Contextual Instruction Flow for Sub-task Execution in Multi-Turn Interactions with a Single On-Device LLM
abstract
We present CIFLEX (Contextual Instruction FLow with EXecution), a novel execution system for efficient sub-task handling in multiturn interactions with a single on-device large language model (LLM).As LLMs become increasingly capable, a single model is expected to handle diverse sub-tasks that more effectively and comprehensively support answering user requests.Naive approach reprocesses the entire conversation context when switching between main and sub-tasks (e.g., query rewriting, summarization), incurring significant computational overhead.CIFLEX mitigates this overhead by reusing the key-value (KV) cache from the main task and injecting only task-specific instructions into isolated side paths.After sub-task execution, the model rolls back to the main path via cached context, thereby avoiding redundant prefill computation.To support sub-task selection, we also develop a hierarchical classification strategy tailored for small-scale models, decomposing multi-choice decisions into binary ones.Experiments show that CIFLEX significantly reduces computational costs without degrading task performance, enabling scalable and efficient multitask dialogue on-device.
Juntae Lee, Jihwan Bang, Seunghan Yang, Simyung Chang
EMNLP4
2025 Learning Contextual Retrieval for Robust Conversational Search
abstract
Effective conversational search demands a deep understanding of user intent across multiple dialogue turns.Users frequently use abbreviations and shift topics in the middle of conversations, posing challenges for conventional retrievers.While query rewriting techniques improve clarity, they often incur significant computational cost due to additional autoregressive steps.Moreover, although LLMbased retrievers demonstrate strong performance, they are not explicitly optimized to track user intent in multi-turn settings, often failing under topic drift or contextual ambiguity.To address these limitations, we propose ContextualRetriever, a novel LLM-based retriever that directly incorporates conversational context into the retrieval process.Our approach introduces: (1) a context-aware embedding mechanism that highlights the current query within the dialogue history; (2) intent-guided supervision based on high-quality rewritten queries; and (3) a training strategy that preserves the generative capabilities of the base LLM.Extensive evaluations across multiple conversational search benchmarks demonstrate that ContextualRetriever significantly outperforms existing methods while incurring no additional inference overhead.
Seunghan Yang, Juntae Lee, Jihwan Bang, Kyuhong Shim, Simyung Chang
EMNLP6
2025 InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding
abstract
Modern multimodal large language models (MLLMs) can reason over hour-long video, yet their key–value (KV) cache grows linearly with time—quickly exceeding the fixed memory of phones, AR glasses, and edge robots. Prior compression schemes either assume the whole video and user query are available offline or must first build the full cache, so memory still scales with stream length. InfiniPot-V is the first training-free, query-agnostic framework that enforces a hard, length-independent memory cap for \textit{streaming} video understanding. During video encoding it monitors the cache and, once a user-set threshold is reached, runs a lightweight compression pass that (i) removes temporally redundant tokens via Temporal-axis Redundancy (TaR) metric and (ii) keeps semantically significant tokens via Value-Norm (VaN) ranking. Across four open-source MLLMs and four long-video and streaming-video benchmarks, InfiniPot-V cuts peak GPU memory by up to 94\%, sustains real-time generation, and matches or surpasses full-cache accuracy—even in multi-turn dialogues. By dissolving the KV cache bottleneck without retraining or query knowledge, InfiniPot-V closes the gap for on-device streaming video assistants.
Kyuhong Shim, Jungwook Choi, Simyung Chang
NeurIPS4
2024 Crayon: Customized On-Device LLM via Instant Adapter Blending and Edge-Server Hybrid Inference
abstract
The customization of large language models (LLMs) for user-specified tasks gets important.However, maintaining all the customized LLMs on cloud servers incurs substantial memory and computational overheads, and uploading user data can also lead to privacy concerns.Ondevice LLMs can offer a promising solution by mitigating these issues.Yet, the performance of on-device LLMs is inherently constrained by the limitations of small-scaled models.To overcome these restrictions, we first propose Crayon, a novel approach for on-device LLM customization.Crayon begins by constructing a pool of diverse base adapters, and then we instantly blend them into a customized adapter without extra training.In addition, we develop a device-server hybrid inference strategy, which deftly allocates more demanding queries or non-customized tasks to a larger, more capable LLM on a server.This ensures optimal performance without sacrificing the benefits of on-device customization.We carefully craft a novel benchmark from multiple questionanswer datasets, and show the efficacy of our method in the LLM customization. *The authors contribute equally.
Jihwan Bang, Juntae Lee, Kyuhong Shim, Seunghan Yang, Simyung Chang
ACL (1)5
2024 Feature Diversification and Adaptation for Federated Domain Generalization
Seunghan Yang, Seokeon Choi, Hyunsin Park, Sungha Choi, Simyung Chang, Sungrack Yun
ECCV (72)5
2024 InfiniPot: Infinite Context Processing on Memory-Constrained LLMs
abstract
Handling long input contexts remains a significant challenge for Large Language Models (LLMs), particularly in resource-constrained environments such as mobile devices.Our work aims to address this limitation by introducing InfiniPot, a novel KV cache control framework designed to enable pre-trained LLMs to manage extensive sequences within fixed memory constraints efficiently, without requiring additional training.InfiniPot leverages Continual Context Distillation (CCD), an iterative process that compresses and retains essential information through novel importance metrics, effectively maintaining critical data even without access to future context.Our comprehensive evaluations indicate that InfiniPot significantly outperforms models trained for long contexts in various NLP tasks, establishing its efficacy and versatility.This work represents a substantial advancement toward making LLMs applicable to a broader range of real-world scenarios.
Kyuhong Shim, Jungwook Choi, Simyung Chang
EMNLP4
2023 Scalable Weight Reparametrization for Efficient Transfer Learning
abstract
This paper proposes a novel, efficient transfer learning method, called Scalable Weight Reparametrization (SWR) that is efficient and effective for multiple downstream tasks. Efficient transfer learning involves utilizing a pre-trained model trained on a larger dataset and repurposing it for downstream tasks with the aim of maximizing the reuse of the pre-trained model. However, previous works have led to an increase in updated parameters and task-specific modules, resulting in more computations, especially for tiny models. Additionally, there has been no practical consideration for controlling the number of updated parameters. To address these issues, we suggest learning a policy network that can decide where to reparametrize the pre-trained model, while adhering to a given constraint for the number of updated parameters. The policy network is only used during the transfer learning process and not afterward. As a result, our approach attains state-of-the-art performance in a proposed multi-lingual keyword spotting and a standard benchmark, ImageNet-to-Sketch, while requiring zero additional computations and significantly fewer additional parameters.
Byeonggeun Kim, Juntae Lee, Seunghan Yang, Simyung Chang
ICASSP4
2023 Task-Agnostic Open-Set Prototype for Few-Shot Open-Set Recognition
abstract
In few-shot open-set recognition (FSOSR), a network learns to recognize closed-set samples with a few support samples while rejecting open-set samples with no class cue. Unlike conventional OSR, the FSOSR considers more practical open worlds where a closed-set class can be selected as an open-set class in another testing (task) and vice versa. Existing FSOSR methods have commonly represented the open set with task-dependent extra modules. These modules decently handle the varied closed and open classes but accompany inevitable complexity increase. This paper shows that a single open-set prototype can represent open-set samples when it satisfies a specific relation in metric space: closest to open-set, and simultaneously second nearest to close-set. We propose a task-agnostic open-set prototype with distance scaling factors and design loss terms. We extensively analyze the proposed components to demonstrate their importance. Our method achieves state-of-the-art results on miniImageNet and tieredImageNet, respectively, without task-dependent extra modules.
Byeonggeun Kim, Juntae Lee, Kyuhong Shim, Simyung Chang
ICIP4
2022 Multi-Head Modularization to Leverage Generalization Capability in Multi-Modal Networks
abstract
It has been crucial to leverage the rich information of multiple modalities in many tasks. Existing works have tried to design multi-modal networks with descent multi-modal fusion modules. Instead, we focus on improving generalization capability of multi-modal networks, especially the fusion module. Viewing the multi-modal data as different projections of information, we first observe that bad projection can cause poor generalization behaviors of multi-modal networks. Then, motivated by well-generalized network's low sensitivity to perturbation, we propose a novel multi-modal training method, multi-head modularization (MHM). We modularize a multi-modal network as a series of uni-modal embedding, multi-modal embedding, and task-specific head modules. Also, for training, we exploit multiple head modules learned with different datasets, swapping each other. From this, we can make the multi-modal embedding module robust to all the heads with different generalization behaviors. In testing phase, we select one of the head modules not to increase the computational cost. Owing to the perturbation of head modules, though including one selected head, the deployed network is more well-generalized compared to the simply end-to-end learned. We verify the effectiveness of MHM on various multi-modal tasks. We use the state-of-the-art methods as baselines, and show notable performance gain for all the baselines.
Juntae Lee, Hyunsin Park, Sungrack Yun, Simyung Chang
AAAI4
2022 Variational On-the-Fly Personalization
abstract
With the development of deep learning (DL) technologies, the demand for DL-based services on personal devices, such as mobile phones, also increases rapidly. In this paper, we propose a novel personalization method, Variational On-the-Fly Personalization. Compared to the conventional personalization methods that require additional fine-tuning with personal data, the proposed method only requires forwarding a handful of personal data on-the-fly. Assuming even a single personal data can convey the characteristics of a target person, we develop the variational hyper-personalizer to capture the weight distribution of layers that fits the target person. In the testing phase, the hyper-personalizer estimates the model’s weights on-the-fly based on personality by forwarding only a small amount of (even a single) personal enrollment data. Hence, the proposed method can perform the personalization without any training software platform and additional cost in the edge device. In experiments, we show our approach can effectively generate reliable personalized models via forwarding (not back-propagating) a handful of samples.
Jangho Kim, Juntae Lee, Simyung Chang, Nojun Kwak
ICML3
2022 Dummy Prototypical Networks for Few-Shot Open-Set Keyword Spotting
abstract
Keyword spotting is the task of detecting a keyword in streaming audio. Conventional keyword spotting targets predefined keywords classification, but there is growing attention in few-shot (query-by-example) keyword spotting, e.g., N-way classification given M-shot support samples. Moreover, in real-world scenarios, there can be utterances from unexpected categories (open-set) which need to be rejected rather than classified as one of the N classes. Combining the two needs, we tackle few-shot open-set keyword spotting with a new benchmark setting, named splitGSC. We propose episode-known dummy prototypes based on metric learning to detect an open-set better and introduce a simple and powerful approach, Dummy Prototypical Networks (D-ProtoNets). Our D-ProtoNets shows clear margins compared to recent few-shot open-set recognition (FSOSR) approaches in the suggested splitGSC. We also verify our method on a standard benchmark, miniImageNet, and D-ProtoNets shows the state-of-the-art open-set detection rate in FSOSR.
Byeonggeun Kim, Seunghan Yang, Inseop Chung, Simyung Chang
INTERSPEECH4
2022 Domain Generalization with Relaxed Instance Frequency-wise Normalization for Multi-device Acoustic Scene Classification
abstract
While using two-dimensional convolutional neural networks (2D-CNNs) in image processing, it is possible to manipulate domain information using channel statistics, and instance normalization has been a promising way to get domain-invariant features. However, unlike image processing, we analyze that domain-relevant information in an audio feature is dominant in frequency statistics rather than channel statistics. Motivated by our analysis, we introduce Relaxed Instance Frequency-wise Normalization (RFN): a plug-and-play, explicit normalization module along the frequency axis which can eliminate instance-specific domain discrepancy in an audio feature while relaxing undesirable loss of useful discriminative information. Empirically, simply adding RFN to networks shows clear margins compared to previous domain generalization approaches on acoustic scene classification and yields improved robustness for multiple audio devices. Especially, the proposed RFN won the DCASE2021 challenge TASK1A, low-complexity acoustic scene classification with multiple devices, with a clear margin, and RFN is an extended work of our technical report.
Byeonggeun Kim, Seunghan Yang, Jangho Kim, Hyunsin Park, Juntae Lee, Simyung Chang
INTERSPEECH6
2022 Personalized Keyword Spotting through Multi-task Learning
abstract
Keyword spotting (KWS) plays an essential role in enabling speech-based user interaction on smart devices, and conventional KWS (C-KWS) approaches have concentrated on detecting user-agnostic pre-defined keywords.However, in practice, most user interactions come from target users enrolled in the device which motivates to construct personalized keyword spotting.We design two personalized KWS tasks; (1) Target user Biased KWS (TB-KWS) and ( 2) Target user Only KWS (TO-KWS).To solve the tasks, we propose personalized keyword spotting through multi-task learning (PK-MTL) that consists of multi-task learning and task-adaptation.First, we introduce applying multi-task learning on keyword spotting and speaker verification to leverage user information to the keyword spotting system.Next, we design task-specific scoring functions to adapt to the personalized KWS tasks thoroughly.We evaluate our framework on conventional and personalized scenarios, and the results show that PK-MTL can dramatically reduce the false alarm rate, especially in various practical scenarios.
Seunghan Yang, Byeonggeun Kim, Inseop Chung, Simyung Chang
INTERSPEECH4
2022 Dynamic Iterative Refinement for Efficient 3D Hand Pose Estimation
abstract
While hand pose estimation is a critical component of most interactive extended reality and gesture recognition systems, contemporary approaches are not optimized for computational and memory efficiency. In this paper, we propose a tiny deep neural network of which partial layers are recursively exploited for refining its previous estimations. During its iterative refinements, we employ learned gating criteria to decide whether to exit from the weight-sharing loop, allowing per-sample adaptation in our model. Our network is trained to be aware of the uncertainty in its current predictions to efficiently gate at each iteration, estimating variances after each loop for its keypoint estimates. Additionally, we investigate the effectiveness of end-to-end and progressive training protocols for our recursive structure on maximizing the model capacity. With the proposed setting, our method consistently outperforms state-of-the-art 2D/3D hand pose estimation approaches in terms of both accuracy and efficiency for widely used benchmarks.
John Yang 0001, Yash Bhalgat, Simyung Chang, Fatih Porikli, Nojun Kwak
WACV3
2021 Subspectral Normalization for Neural Audio Data Processing
abstract
Convolutional Neural Networks are widely used in various machine learning domains. In image processing, the features can be obtained by applying 2D convolution to all spatial dimensions of the input. However, in the audio case, frequency domain input like Mel-Spectrogram has different and unique characteristics in the frequency dimension. Thus, there is a need for a method that allows the 2D convolution layer to handle the frequency dimension differently. In this work, we introduce SubSpectral Normalization (SSN), which splits the input frequency dimension into several groups (sub-bands) and performs a different normalization for each group. SSN also includes an affine transformation that can be applied to each group. Our method removes the inter-frequency deflection while the network learns a frequency-aware characteristic. In the experiments with audio data, we observed that SSN can efficiently improve the network’s performance.
Simyung Chang, Hyoungwoo Park, Janghoon Cho, Hyunsin Park, Sungrack Yun, Kyuwoong Hwang
ICASSP1
2021 Prototype-Based Personalized Pruning
abstract
Nowadays, as edge devices such as smartphones become prevalent, there are increasing demands for personalized services. However, traditional personalization methods are not suitable for edge devices because retraining or finetuning is needed with limited personal data. Also, a full model might be too heavy for edge devices with limited resources. Unfortunately, model compression methods which can handle the model complexity issue also require the retraining phase. These multiple training phases generally need huge computational cost during on-device learning which can be a burden to edge devices. In this work, we propose a dynamic personalization method called prototype-based personalized pruning (PPP). PPP considers both ends of personalization and model efficiency. After training a network, PPP can easily prune the network with a prototype representing the characteristics of personal data and it performs well without retraining or finetuning. We verify the usefulness of PPP on a couple of tasks in computer vision and Keyword spotting.
Jangho Kim, Simyung Chang, Sungrack Yun, Nojun Kwak
ICASSP2
2021 Broadcasted Residual Learning for Efficient Keyword Spotting
abstract
Keyword spotting is an important research field because it plays a key role in device wake-up and user interaction on smart devices.However, it is challenging to minimize errors while operating efficiently in devices with limited resources such as mobile phones.We present a broadcasted residual learning method to achieve high accuracy with small model size and computational load.Our method configures most of the residual functions as 1D temporal convolution while still allows 2D convolution together using a broadcasted-residual connection that expands temporal output to frequency-temporal dimension.This residual mapping enables the network to effectively represent useful audio features with much less computation than conventional convolutional neural networks.We also propose a novel network architecture, Broadcasting-residual network (BC-ResNet), based on broadcasted residual learning and describe how to scale up the model according to the target device's resources.BC-ResNets achieve state-of-the-art 98.0% and 98.7% top-1 accuracy on Google speech command datasets v1 and v2, respectively, and consistently outperform previous approaches, using fewer computations and parameters.Code is available at https://github.com/Qualcomm-
Byeonggeun Kim, Simyung Chang, Jinkyu Lee 0004, Dooyong Sung
Interspeech2
2021 PQK: Model Compression via Pruning, Quantization, and Knowledge Distillation
abstract
As edge devices become prevalent, deploying Deep Neural Networks (DNN) on edge devices has become a critical issue. However, DNN requires a high computational resource which is rarely available for edge devices. To handle this, we propose a novel model compression method for the devices with limited computational resources, called PQK consisting of pruning, quantization, and knowledge distillation (KD) processes. Unlike traditional pruning and KD, PQK makes use of unimportant weights pruned in the pruning process to make a teacher network for training a better student network without pre-training the teacher model. PQK has two phases. Phase 1 exploits iterative pruning and quantization-aware training to make a lightweight and power-efficient model. In phase 2, we make a teacher network by adding unimportant weights unused in phase 1 to a pruned network. By using this teacher network, we train the pruned network as a student network. In doing so, we do not need a pre-trained teacher network for the KD framework because the teacher and the student networks coexist within the same network. We apply our method to the recognition model and verify the effectiveness of PQK on keyword spotting (KWS) and image recognition.
Jangho Kim, Simyung Chang, Nojun Kwak
Interspeech2
2020 URNet: User-Resizable Residual Networks with Conditional Gating Module
abstract
Convolutional Neural Networks are widely used to process spatial scenes, but their computational cost is fixed and depends on the structure of the network used. There are methods to reduce the cost by compressing networks or varying its computational path dynamically according to the input image. However, since a user can not control the size of the learned model, it is difficult to respond dynamically if the amount of service requests suddenly increases. We propose User-Resizable Residual Networks (URNet), which allows users to adjust the computational cost of the network as needed during evaluation. URNet includes Conditional Gating Module (CGM) that determines the use of each residual block according to the input image and the desired cost. CGM is trained in a supervised manner using the newly proposed scale(cost) loss and its corresponding training methods. URNet can control the amount of computation and its inference path according to user's demand without degrading the accuracy significantly. In the experiments on ImageNet, URNet based on ResNet-101 maintains the accuracy of the baseline even when resizing it to approximately 80% of the original network, and demonstrates only about 1% accuracy degradation when using about 65% of the computation.
Simyung Chang, Nojun Kwak
AAAI2
2020 BIBNet: An Efficient Super Resolution with Bottleneck-In-Bottleneck
abstract
Deep Neural Networks have enabled remarkable progress in the field of single image super resolution (SR). However, these models are often large and complex to be applied for real-world applications with limited resources as in mobile and embedded systems. We investigate whether the typical low latency models as MobileNet can be expected of comparable efficiency at SR tasks with recently reported SR performance in the literature. To this end, a moderate and effective architecture, Bottleneck-In-Bottleneck (BIB), is introduced in this paper. The BIB uses multiple expansion factors of the residual blocks in the form of a bottleneck, reducing computation complexity while utilizing advantageous factors of large feature dimensions. We also propose BIBNet with multiple BIB blocks, which can easily adjust its size and computational cost to create a variety of efficient and high-performance models. Extensive experiments show that, with fewer parameters and computations, BIBNet achieves highly competitive performance compared to other conventional SR methods with more complex architectures.
Simyung Chang, Keuntek Lee, Shobhit Jain, Cheul-Hee Hahm
IJCNN1
2019 Towards Governing Agent's Efficacy: Action-Conditional $β$-VAE for Deep Transparent Reinforcement Learning
abstract
We tackle the blackbox issue of deep neural networks in the settings of reinforcement learning (RL) where neural agents learn towards maximizing reward gains in an uncontrollable way. Such learning approach is risky when the interacting environment includes an expanse of state space because it is then almost impossible to foresee all unwanted outcomes and penalize them with negative rewards beforehand. We propose Action-conditional $\beta$-VAE (AC-$\beta$-VAE) that allows succinct mappings of action-dependent factors in desirable dimensions of latent representations while disentangling environmental factors. Our proposed method tackles the blackbox issue by encouraging an RL policy network to learn interpretable latent features by distinguits influenshing ices from uncontrollable environmental factors, which closely resembles the way humans understand their scenes. Our experimental results show that the learned latent factors not only are interpretable, but also enable modeling the distribution of entire visited state-action space. We have experimented that this characteristic of the proposed structure can lead to ex post facto governance for desired behaviors of RL agents.
John Yang 0001, Gyuejeong Lee, Simyung Chang, Nojun Kwak
ACML3
2019 Sym-Parameterized Dynamic Inference for Mixed-Domain Image Translation
abstract
Recent advances in image-to-image translation have led to some ways to generate multiple domain images through a single network. However, there is still a limit in creating an image of a target domain without a dataset on it. We propose a method to expand the concept of `multi-domain' from data to the loss area, and to combine the characteristics of each domain to create an image. First, we introduce a sym-parameter and its learning method that can mix various losses and can synchronize them with input conditions. Then, we propose Sym-parameterized Generative Network (SGN) using it. Through experiments, we confirmed that SGN could mix the characteristics of various data and losses, and it is possible to translate images to any mixed-domain without ground truths, such as 30% Van Gogh and 20% Monet and 40% snowy.
Simyung Chang, Seonguk Park, John Yang 0001, Nojun Kwak
ICCV1
2019 BOOK: Storing Algorithm-Invariant Episodes for Deep Reinforcement Learning
abstract
We introduce a novel method to train agents of reinforcement learning (RL) by sharing knowledge in a way similar to the concept of using a book. The recorded information in the form of a book is the main means by which humans learn knowledge. Nevertheless, the conventional deep RL methods have mainly focused either on experiential learning where the agent learns through interactions with the environment from the start or on imitation learning that tries to mimic the teacher. Contrary to these, our proposed book learning shares key information among different agents in a book-like manner by delving into the following two characteristic features: (1) By defining the linguistic function, input states can be clustered semantically into a relatively small number of core clusters, which are forwarded to other RL agents in a prescribed manner. (2) By defining state priorities and the contents for recording, core experiences can be selected and stored in a small container. We call this container as 'BOOK'. Our method learns hundreds to thousand times faster than the conventional methods by learning only a handful of core cluster information, which shows that deep RL agents can effectively learn through the shared knowledge from other agents.
Simyung Chang, Young Joon Yoo, Jaeseok Choi, Nojun Kwak
ICPRAM1
2018 Broadcasting Convolutional Network for Visual Relational Reasoning
Simyung Chang, John Yang 0001, Seonguk Park, Nojun Kwak
ECCV (15)1
2018 Genetic-Gated Networks for Deep Reinforcement Learning
abstract
We introduce the Genetic-Gated Networks (G2Ns), simple neural networks that combine a gate vector composed of binary genetic genes in the hidden layer(s) of networks. Our method can take both advantages of gradient-free optimization and gradient-based optimization methods, of which the former is effective for problems with multiple local minima, while the latter can quickly find local minima. In addition, multiple chromosomes can define different models, making it easy to construct multiple models and can be effectively applied to problems that require multiple models. We show that this G2N can be applied to typical reinforcement learning algorithms to achieve a large improvement in sample efficiency and performance.
Simyung Chang, John Yang 0001, Jaeseok Choi, Nojun Kwak
NeurIPS1