EDBT 2026 Demo / reviewers in the wild / expert
Seunghan Yang
dblp:250/9141
· DBLP profile ↗
17ranked-venue papers
7as first author
14since 2021 · last 2025
0000-0002-0411-8407ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 12 · 5 first-author · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CIFLEX: Contextual Instruction Flow for Sub-task Execution in Multi-Turn Interactions with a Single On-Device LLMabstractWe present CIFLEX (Contextual Instruction FLow with EXecution), a novel execution system for efficient sub-task handling in multiturn interactions with a single on-device large language model (LLM).As LLMs become increasingly capable, a single model is expected to handle diverse sub-tasks that more effectively and comprehensively support answering user requests.Naive approach reprocesses the entire conversation context when switching between main and sub-tasks (e.g., query rewriting, summarization), incurring significant computational overhead.CIFLEX mitigates this overhead by reusing the key-value (KV) cache from the main task and injecting only task-specific instructions into isolated side paths.After sub-task execution, the model rolls back to the main path via cached context, thereby avoiding redundant prefill computation.To support sub-task selection, we also develop a hierarchical classification strategy tailored for small-scale models, decomposing multi-choice decisions into binary ones.Experiments show that CIFLEX significantly reduces computational costs without degrading task performance, enabling scalable and efficient multitask dialogue on-device. Juntae Lee, Jihwan Bang, Seunghan Yang, Simyung Chang |
EMNLP | 3 |
| 2025 | Learning Contextual Retrieval for Robust Conversational SearchabstractEffective conversational search demands a deep understanding of user intent across multiple dialogue turns.Users frequently use abbreviations and shift topics in the middle of conversations, posing challenges for conventional retrievers.While query rewriting techniques improve clarity, they often incur significant computational cost due to additional autoregressive steps.Moreover, although LLMbased retrievers demonstrate strong performance, they are not explicitly optimized to track user intent in multi-turn settings, often failing under topic drift or contextual ambiguity.To address these limitations, we propose ContextualRetriever, a novel LLM-based retriever that directly incorporates conversational context into the retrieval process.Our approach introduces: (1) a context-aware embedding mechanism that highlights the current query within the dialogue history; (2) intent-guided supervision based on high-quality rewritten queries; and (3) a training strategy that preserves the generative capabilities of the base LLM.Extensive evaluations across multiple conversational search benchmarks demonstrate that ContextualRetriever significantly outperforms existing methods while incurring no additional inference overhead. Seunghan Yang, Juntae Lee, Jihwan Bang, Kyuhong Shim, Simyung Chang |
EMNLP | 1 |
| 2024 | Crayon: Customized On-Device LLM via Instant Adapter Blending and Edge-Server Hybrid InferenceabstractThe customization of large language models (LLMs) for user-specified tasks gets important.However, maintaining all the customized LLMs on cloud servers incurs substantial memory and computational overheads, and uploading user data can also lead to privacy concerns.Ondevice LLMs can offer a promising solution by mitigating these issues.Yet, the performance of on-device LLMs is inherently constrained by the limitations of small-scaled models.To overcome these restrictions, we first propose Crayon, a novel approach for on-device LLM customization.Crayon begins by constructing a pool of diverse base adapters, and then we instantly blend them into a customized adapter without extra training.In addition, we develop a device-server hybrid inference strategy, which deftly allocates more demanding queries or non-customized tasks to a larger, more capable LLM on a server.This ensures optimal performance without sacrificing the benefits of on-device customization.We carefully craft a novel benchmark from multiple questionanswer datasets, and show the efficacy of our method in the LLM customization. *The authors contribute equally. Jihwan Bang, Juntae Lee, Kyuhong Shim, Seunghan Yang, Simyung Chang |
ACL (1) | 4 |
| 2024 | Feature Diversification and Adaptation for Federated Domain Generalization
Seunghan Yang, Seokeon Choi, Hyunsin Park, Sungha Choi, Simyung Chang, Sungrack Yun |
ECCV (72) | 1 |
| 2024 | Balanced Learning for Multi-Domain Long-Tailed Speaker RecognitionabstractThis paper considers two types of imbalance problems commonly inherent in large-scale datasets: multiple domain and class imbalance. Class imbalance causes the algorithm to be biased toward the majority classes, and multiple-domain data results in significant performance disparities for different domains. To tackle these challenges, we propose a novel learning approach for multi-domain imbalanced datasets, featuring two techniques: (i) distribution-aware partial mask and (ii) domain-wise interprototype loss function. The distribution-aware partial mask selects negative class centers based on class-level distribution and domain labels, adjusting the ratio of positive and negative updates for prototype vectors and enhancing discriminative feature learning within each domain. Additionally, the domain-wise interprototype loss enforces orthogonality among prototype vectors within each domain, leading to increased discriminativeness. We demonstrate the superiority of our approach over baselines through experiments on publicly available speaker recognition datasets, including CN-Celeb and Mozilla Common Voice. Janghoon Cho, Hyunsin Park, Hyoungwoo Park, Seunghan Yang, Sungrack Yun |
ICASSP | 5 |
| 2023 | Progressive Random Convolutions for Single Domain GeneralizationabstractSingle domain generalization aims to train a generalizable model with only one source domain to perform well on arbitrary unseen target domains. Image augmentation based on Random Convolutions (RandConv), consisting of one convolution layer randomly initialized for each mini-batch, enables the model to learn generalizable visual representations by distorting local textures despite its simple and lightweight structure. However, RandConv has structural limitations in that the generated image easily loses semantics as the kernel size increases, and lacks the inherent diversity of a single convolution operation. To solve the problem, we propose a Progressive Random Convolution (Pro-RandConv) method that recursively stacks random convolution layers with a small kernel size instead of increasing the kernel size. This progressive approach can not only mitigate semantic distortions by reducing the influence of pixels away from the center in the theoretical receptive field, but also create more effective virtual domains by gradually increasing the style diversity. In addition, we develop a basic random convolution layer into a random convolution block including deformable offsets and affine transformation to support texture and contrast diversification, both of which are also randomly initialized. Without complex generators or adversarial learning, we demonstrate that our simple yet effective augmentation strategy outperforms state-of-the-art methods on single domain generalization benchmarks. Seokeon Choi, Debasmit Das, Sungha Choi, Seunghan Yang, Hyunsin Park, Sungrack Yun |
CVPR | 4 |
| 2023 | Scalable Weight Reparametrization for Efficient Transfer LearningabstractThis paper proposes a novel, efficient transfer learning method, called Scalable Weight Reparametrization (SWR) that is efficient and effective for multiple downstream tasks. Efficient transfer learning involves utilizing a pre-trained model trained on a larger dataset and repurposing it for downstream tasks with the aim of maximizing the reuse of the pre-trained model. However, previous works have led to an increase in updated parameters and task-specific modules, resulting in more computations, especially for tiny models. Additionally, there has been no practical consideration for controlling the number of updated parameters. To address these issues, we suggest learning a policy network that can decide where to reparametrize the pre-trained model, while adhering to a given constraint for the number of updated parameters. The policy network is only used during the transfer learning process and not afterward. As a result, our approach attains state-of-the-art performance in a proposed multi-lingual keyword spotting and a standard benchmark, ImageNet-to-Sketch, while requiring zero additional computations and significantly fewer additional parameters. Byeonggeun Kim, Juntae Lee, Seunghan Yang, Simyung Chang |
ICASSP | 3 |
| 2023 | Label Shift Adapter for Test-Time Adaptation under Covariate and Label ShiftsabstractTest-time adaptation (TTA) aims to adapt a pre-trained model to the target domain in a batch-by-batch manner during inference. While label distributions often exhibit imbalances in real-world scenarios, most previous TTA approaches typically assume that both source and target domain datasets have balanced label distribution. Due to the fact that certain classes appear more frequently in certain domains (e.g., buildings in cities, trees in forests), it is natural that the label distribution shifts as the domain changes. However, we discover that the majority of existing TTA methods fail to address the coexistence of covariate and label shifts. To tackle this challenge, we propose a novel label shift adapter that can be incorporated into existing TTA approaches to deal with label shifts during the TTA process effectively. Specifically, we estimate the label distribution of the target domain to feed it into the label shift adapter. Subsequently, the label shift adapter produces optimal parameters for target label distribution. By predicting only the parameters for a part of the pre-trained source model, our approach is computationally efficient and can be easily applied, regardless of the model architectures. Through extensive experiments, we demonstrate that integrating our strategy with TTA approaches leads to substantial performance improvements under the joint presence of label and covariate shifts. Sunghyun Park 0005, Seunghan Yang, Jaegul Choo, Sungrack Yun |
ICCV | 2 |
| 2023 | Improving Small Footprint Few-shot Keyword Spotting with Supervision on Auxiliary Data
Seunghan Yang, Byeonggeun Kim, Kyuhong Shim, Simyoung Chang |
INTERSPEECH | 1 |
| 2022 | Improving Test-Time Adaptation Via Shift-Agnostic Weight Regularization and Nearest Source Prototypes
Sungha Choi, Seunghan Yang, Seokeon Choi, Sungrack Yun |
ECCV (33) | 2 |
| 2022 | Dummy Prototypical Networks for Few-Shot Open-Set Keyword SpottingabstractKeyword spotting is the task of detecting a keyword in streaming audio. Conventional keyword spotting targets predefined keywords classification, but there is growing attention in few-shot (query-by-example) keyword spotting, e.g., N-way classification given M-shot support samples. Moreover, in real-world scenarios, there can be utterances from unexpected categories (open-set) which need to be rejected rather than classified as one of the N classes. Combining the two needs, we tackle few-shot open-set keyword spotting with a new benchmark setting, named splitGSC. We propose episode-known dummy prototypes based on metric learning to detect an open-set better and introduce a simple and powerful approach, Dummy Prototypical Networks (D-ProtoNets). Our D-ProtoNets shows clear margins compared to recent few-shot open-set recognition (FSOSR) approaches in the suggested splitGSC. We also verify our method on a standard benchmark, miniImageNet, and D-ProtoNets shows the state-of-the-art open-set detection rate in FSOSR. Byeonggeun Kim, Seunghan Yang, Inseop Chung, Simyung Chang |
INTERSPEECH | 2 |
| 2022 | Domain Generalization with Relaxed Instance Frequency-wise Normalization for Multi-device Acoustic Scene ClassificationabstractWhile using two-dimensional convolutional neural networks (2D-CNNs) in image processing, it is possible to manipulate domain information using channel statistics, and instance normalization has been a promising way to get domain-invariant features. However, unlike image processing, we analyze that domain-relevant information in an audio feature is dominant in frequency statistics rather than channel statistics. Motivated by our analysis, we introduce Relaxed Instance Frequency-wise Normalization (RFN): a plug-and-play, explicit normalization module along the frequency axis which can eliminate instance-specific domain discrepancy in an audio feature while relaxing undesirable loss of useful discriminative information. Empirically, simply adding RFN to networks shows clear margins compared to previous domain generalization approaches on acoustic scene classification and yields improved robustness for multiple audio devices. Especially, the proposed RFN won the DCASE2021 challenge TASK1A, low-complexity acoustic scene classification with multiple devices, with a clear margin, and RFN is an extended work of our technical report. Byeonggeun Kim, Seunghan Yang, Jangho Kim, Hyunsin Park, Juntae Lee, Simyung Chang |
INTERSPEECH | 2 |
| 2022 | Domain Agnostic Few-shot Learning for Speaker VerificationabstractDeep learning models for verification systems often fail to generalize to new users and new environments, even though they learn highly discriminative features.To address this problem, we propose a few-shot domain generalization framework that learns to tackle distribution shift for new users and new domains.Our framework consists of domain-specific and domainaggregation networks, which are the experts on specific and combined domains, respectively.By using these networks, we generate episodes that mimic the presence of both novel users and novel domains in the training phase to eventually produce better generalization.To save memory, we reduce the number of domain-specific networks by clustering similar domains together.Upon extensive evaluation on artificially generated noise domains, we can explicitly show generalization ability of our framework.In addition, we apply our proposed methods to the existing competitive architecture on the standard benchmark, which shows further performance improvements. Seunghan Yang, Debasmit Das, Janghoon Cho, Hyoungwoo Park, Sungrack Yun |
INTERSPEECH | 1 |
| 2022 | Personalized Keyword Spotting through Multi-task LearningabstractKeyword spotting (KWS) plays an essential role in enabling speech-based user interaction on smart devices, and conventional KWS (C-KWS) approaches have concentrated on detecting user-agnostic pre-defined keywords.However, in practice, most user interactions come from target users enrolled in the device which motivates to construct personalized keyword spotting.We design two personalized KWS tasks; (1) Target user Biased KWS (TB-KWS) and ( 2) Target user Only KWS (TO-KWS).To solve the tasks, we propose personalized keyword spotting through multi-task learning (PK-MTL) that consists of multi-task learning and task-adaptation.First, we introduce applying multi-task learning on keyword spotting and speaker verification to leverage user information to the keyword spotting system.Next, we design task-specific scoring functions to adapt to the personalized KWS tasks thoroughly.We evaluate our framework on conventional and personalized scenarios, and the results show that PK-MTL can dramatically reduce the false alarm rate, especially in various practical scenarios. Seunghan Yang, Byeonggeun Kim, Inseop Chung, Simyung Chang |
INTERSPEECH | 1 |
| 2020 | Arbitrary Style Transfer Using Graph Instance NormalizationabstractStyle transfer is the image synthesis task, which applies a style of one image to another while preserving the content. In statistical methods, the adaptive instance normalization (AdaIN) whitens the source images and applies the style of target images through normalizing the mean and variance of features. However, computing feature statistics for each instance would neglect the inherent relationship between features, so it is hard to learn global styles while fitting to the individual training dataset. In this paper, we present a novel learnable normalization technique for style transfer using graph convolutional networks, termed Graph Instance Normalization (GrIN). This algorithm makes the style transfer approach more robust by taking into account similar information shared between instances. Besides, this simple module is also applicable to other tasks like image-to-image translation or domain adaptation. Dongki Jung, Seunghan Yang, Changick Kim |
ICIP | 2 |
| 2020 | The Korean Sign Language Dataset for Action Recognition
Seunghan Yang, Seungjun Jung, Heekwang Kang, Changick Kim |
MMM (1) | 1 |
| 2020 | Combinational Class Activation Maps for Weakly Supervised Object LocalizationabstractWeakly supervised object localization has recently attracted attention since it aims to identify both class labels and locations of objects by using image-level labels. Most previous methods utilize the activation map corresponding to the highest activation source. Exploiting only one activation map of the highest probability class is often biased into limited regions or sometimes even highlights background regions. To resolve these limitations, we propose to use activation maps, named combinational class activation maps (CCAM), which are linear combinations of activation maps from the highest to the lowest probability class. By using CCAM for localization, we suppress background regions to help highlighting foreground objects more accurately. In addition, we design the network architecture to consider spatial relationships for localizing relevant object regions. Specifically, we integrate non-local modules into an existing base network at both low- and high-level layers. Our final model, named non-local combinational class activation maps (NL-CCAM), obtains superior performance compared to previous methods on representative object localization benchmarks including ILSVRC 2016 and CUB- 200-2011. Furthermore, we show that the proposed method has a great capability of generalization by visualizing other datasets. Seunghan Yang, Yoonhyung Kim, Youngeun Kim, Changick Kim |
WACV | 1 |