Kyuhong Shim

dblp:209/4981 · DBLP profile ↗
← Back
21ranked-venue papers
6as first author
20since 2021 · last 2025
0000-0002-0123-3100ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 6 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 12 since 2021
YearPublicationVenuePosition
2025 Learning Contextual Retrieval for Robust Conversational Search
abstract
Effective conversational search demands a deep understanding of user intent across multiple dialogue turns.Users frequently use abbreviations and shift topics in the middle of conversations, posing challenges for conventional retrievers.While query rewriting techniques improve clarity, they often incur significant computational cost due to additional autoregressive steps.Moreover, although LLMbased retrievers demonstrate strong performance, they are not explicitly optimized to track user intent in multi-turn settings, often failing under topic drift or contextual ambiguity.To address these limitations, we propose ContextualRetriever, a novel LLM-based retriever that directly incorporates conversational context into the retrieval process.Our approach introduces: (1) a context-aware embedding mechanism that highlights the current query within the dialogue history; (2) intent-guided supervision based on high-quality rewritten queries; and (3) a training strategy that preserves the generative capabilities of the base LLM.Extensive evaluations across multiple conversational search benchmarks demonstrate that ContextualRetriever significantly outperforms existing methods while incurring no additional inference overhead.
Seunghan Yang, Juntae Lee, Jihwan Bang, Kyuhong Shim, Simyung Chang
EMNLP4
2025 Learning Primitive Relations for Compositional Zero-Shot Learning
abstract
Compositional Zero-Shot Learning (CZSL) aims to identify unseen state-object compositions by leveraging knowledge learned from seen compositions. Existing approaches often independently predict states and objects, overlooking their relationships. In this paper, we propose a novel framework, learning primitive relations (LPR), designed to probabilistically capture the relationships between states and objects. By employing the cross-attention mechanism, LPR considers the dependencies between states and objects, enabling the model to infer the likelihood of unseen compositions. Experimental results demonstrate that LPR outperforms state-of-the-art methods on all three CZSL benchmark datasets in both closed-world and open-world settings. Through qualitative analysis, we show that LPR leverages state-object relationships for unseen composition prediction.
Insu Lee, Jiseob Kim, Kyuhong Shim, Byonghyo Shim
ICASSP3
2025 Visually Guided Decoding: Gradient-Free Hard Prompt Inversion with Language Models
abstract
Text-to-image generative models like DALL-E and Stable Diffusion have revolutionized visual content creation across various applications, including advertising, personalized media, and design prototyping. However, crafting effective textual prompts to guide these models remains challenging, often requiring extensive trial and error. Existing prompt inversion approaches, such as soft and hard prompt techniques, are not so effective due to the limited interpretability and incoherent prompt generation. To address these issues, we propose Visually Guided Decoding (VGD), a gradient-free approach that leverages large language models (LLMs) and CLIP-based guidance to generate coherent and semantically aligned prompts. In essence, VGD utilizes the robust text generation capabilities of LLMs to produce human-readable prompts. Further, by employing CLIP scores to ensure alignment with user-specified visual concepts, VGD enhances the interpretability, generalization, and flexibility of prompt generation without the need for additional training. Our experiments demonstrate that VGD outperforms existing prompt inversion techniques in generating understandable and contextually relevant prompts, facilitating more intuitive and controllable interactions with text-to-image models.
Minji Bae, Kyuhong Shim, Byonghyo Shim
ICLR3
2025 InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding
abstract
Modern multimodal large language models (MLLMs) can reason over hour-long video, yet their key–value (KV) cache grows linearly with time—quickly exceeding the fixed memory of phones, AR glasses, and edge robots. Prior compression schemes either assume the whole video and user query are available offline or must first build the full cache, so memory still scales with stream length. InfiniPot-V is the first training-free, query-agnostic framework that enforces a hard, length-independent memory cap for \textit{streaming} video understanding. During video encoding it monitors the cache and, once a user-set threshold is reached, runs a lightweight compression pass that (i) removes temporally redundant tokens via Temporal-axis Redundancy (TaR) metric and (ii) keeps semantically significant tokens via Value-Norm (VaN) ranking. Across four open-source MLLMs and four long-video and streaming-video benchmarks, InfiniPot-V cuts peak GPU memory by up to 94\%, sustains real-time generation, and matches or surpasses full-cache accuracy—even in multi-turn dialogues. By dissolving the KV cache bottleneck without retraining or query knowledge, InfiniPot-V closes the gap for on-device streaming video assistants.
Kyuhong Shim, Jungwook Choi, Simyung Chang
NeurIPS2
2025 Towards Comprehensive Scene Understanding: Integrating First and Third-Person Views for LVLMs
abstract
Large vision-language models (LVLMs) are increasingly deployed in interactive applications such as virtual and augmented reality, where a first-person (egocentric) view captured by head-mounted cameras serves as key input. While this view offers fine-grained cues about user attention and hand-object interactions, its narrow field of view and lack of global context often lead to failures on spatially or contextually demanding queries. To address this, we introduce a framework that augments egocentric inputs with third-person (exocentric) views, providing complementary information such as global scene layout and object visibility to LVLMs. We present E3VQA, the first benchmark for multi-view question answering with 4K high-quality question-answer pairs grounded in synchronized ego-exo image pairs. Additionally, we propose M3CoT, a training-free prompting technique that constructs a unified scene representation by integrating scene graphs from three complementary perspectives. M3CoT enables LVLMs to reason more effectively across views, yielding consistent performance gains (4.84\% for GPT-4o and 5.94\% for Gemini 2.0 Flash) over a recent CoT baseline. Our extensive evaluation reveals key strengths and limitations of LVLMs in multi-view reasoning and highlights the value of leveraging both egocentric and exocentric inputs. The dataset and source code are available at [https://github.com/Leeinsu1/Towards-Comprehensive-Scene-Understanding](https://github.com/Leeinsu1/Towards-Comprehensive-Scene-Understanding).
Insu Lee, Wooje Park, Jaeyun Jang, Minyoung Noh, Kyuhong Shim, Byonghyo Shim
NeurIPS5
2024 Expand-and-Quantize: Unsupervised Semantic Segmentation Using High-Dimensional Space and Product Quantization
abstract
Unsupervised semantic segmentation (USS) aims to discover and recognize meaningful categories without any labels. For a successful USS, two key abilities are required: 1) information compression and 2) clustering capability. Previous methods have relied on feature dimension reduction for information compression, however, this approach may hinder the process of clustering. In this paper, we propose a novel USS framework called Expand-and-Quantize Unsupervised Semantic Segmentation (EQUSS), which combines the benefits of high-dimensional spaces for better clustering and product quantization for effective information compression. Our extensive experiments demonstrate that EQUSS achieves state-of-the-art results on three standard benchmarks. In addition, we analyze the entropy of USS features, which is the first step towards understanding USS from the perspective of information theory.
Kyuhong Shim, Insu Lee, Byonghyo Shim
AAAI2
2024 Crayon: Customized On-Device LLM via Instant Adapter Blending and Edge-Server Hybrid Inference
abstract
The customization of large language models (LLMs) for user-specified tasks gets important.However, maintaining all the customized LLMs on cloud servers incurs substantial memory and computational overheads, and uploading user data can also lead to privacy concerns.Ondevice LLMs can offer a promising solution by mitigating these issues.Yet, the performance of on-device LLMs is inherently constrained by the limitations of small-scaled models.To overcome these restrictions, we first propose Crayon, a novel approach for on-device LLM customization.Crayon begins by constructing a pool of diverse base adapters, and then we instantly blend them into a customized adapter without extra training.In addition, we develop a device-server hybrid inference strategy, which deftly allocates more demanding queries or non-customized tasks to a larger, more capable LLM on a server.This ensures optimal performance without sacrificing the benefits of on-device customization.We carefully craft a novel benchmark from multiple questionanswer datasets, and show the efficacy of our method in the LLM customization. *The authors contribute equally.
Jihwan Bang, Juntae Lee, Kyuhong Shim, Seunghan Yang, Simyung Chang
ACL (1)3
2024 InfiniPot: Infinite Context Processing on Memory-Constrained LLMs
abstract
Handling long input contexts remains a significant challenge for Large Language Models (LLMs), particularly in resource-constrained environments such as mobile devices.Our work aims to address this limitation by introducing InfiniPot, a novel KV cache control framework designed to enable pre-trained LLMs to manage extensive sequences within fixed memory constraints efficiently, without requiring additional training.InfiniPot leverages Continual Context Distillation (CCD), an iterative process that compresses and retains essential information through novel importance metrics, effectively maintaining critical data even without access to future context.Our comprehensive evaluations indicate that InfiniPot significantly outperforms models trained for long contexts in various NLP tasks, establishing its efficacy and versatility.This work represents a substantial advancement toward making LLMs applicable to a broader range of real-world scenarios.
Kyuhong Shim, Jungwook Choi, Simyung Chang
EMNLP2
2024 Leveraging Adapter for Parameter-Efficient ASR Encoder
Kyuhong Shim, Jinkyu Lee 0004, Hyunjae Kim
INTERSPEECH1
2023 Teacher Intervention: Improving Convergence of Quantization Aware Training for Ultra-Low Precision Transformers
abstract
Pre-trained Transformer models such as BERT have shown great success in a wide range of applications, but at the cost of substantial increases in model complexity.Quantizationaware training (QAT) is a promising method to lower the implementation cost and energy consumption.However, aggressive quantization below 2-bit causes considerable accuracy degradation due to unstable convergence, especially when the downstream dataset is not abundant.This work proposes a proactive knowledge distillation method called Teacher Intervention (TI) for fast converging QAT of ultralow precision pre-trained Transformers.TI intervenes layer-wise signal propagation with the intact signal from the teacher to remove the interference of propagated quantization errors, smoothing loss surface of QAT and expediting the convergence.Furthermore, we propose a gradual intervention mechanism to stabilize the recovery of subsections of Transformer layers from quantization.The proposed schemes enable fast convergence of QAT and improve the model accuracy regardless of the diverse characteristics of downstream fine-tuning tasks.We demonstrate that TI consistently achieves superior accuracy with significantly lower finetuning iterations on well-known Transformers of natural language processing as well as computer vision compared to the state-of-the-art QAT methods.
Kyuhong Shim, Seongmin Park 0003, Wonyong Sung, Jungwook Choi
EACL2
2023 Vision Transformer-Based Feature Extraction for Generalized Zero-Shot Learning
abstract
Generalized zero-shot learning (GZSL) is a technique to train a deep learning model to identify unseen classes using the image attribute. In this paper, we put forth a new GZSL technique exploiting Vision Transformer (ViT) to maximize the attribute-related information contained in the image feature. In ViT, the entire image region is processed without the degradation of the image resolution and the local image information is preserved in patch features. To fully enjoy the benefits of ViT, we exploit patch features as well as the CLS feature in the extraction of the attribute-related image feature. In particular, we propose a novel attention-based module, called attribute attention module (AAM), to aggregate the attribute-related information in the patch features. From extensive experiments on benchmark datasets, we demonstrate that the proposed technique outperforms the state-of-the-art GZSL approaches by a large margin.
Jiseob Kim, Kyuhong Shim, Junhan Kim, Byonghyo Shim
ICASSP2
2023 Semantic-Preserving Augmentation for Robust Image-Text Retrieval
abstract
Image-text retrieval is a task to search for the proper textual descriptions of the visual world and vice versa. One challenge of this task is the vulnerability to input image/text corruptions. Such corruptions are often unobserved during the training, and degrade the retrieval model’s decision quality substantially. In this paper, we propose a novel image-text retrieval technique, referred to as robust visual semantic embedding (RVSE), which consists of novel image-based and text-based augmentation techniques called semantic-preserving augmentation for image (SPAug-I) and text (SPAug-T). Since SPAug-I and SPAug-T change the original data in a way that its semantic information is preserved, we enforce the feature extractors to generate semantic-aware embedding vectors regardless of the corruption, improving the model’s robustness significantly. From extensive experiments using benchmark datasets, we show that RVSE outperforms conventional retrieval schemes in terms of image-text retrieval performance.
Sunwoo Kim 0004, Kyuhong Shim, Luong Trung Nguyen, Byonghyo Shim
ICASSP2
2023 Spatial Cross-Attention for Transformer-Based Image Captioning
abstract
Transformer-based networks have achieved great success in image captioning because of the attention mechanism that finds relevant image locations for each word. However, the current cross-attention process, which aligns word-to-image, does not consider the spatial relationships existing in patch-to-patch. This lack of spatial information may cause incorrect descriptions that fail at generating words that correctly describe the positional relationships. In this paper, we introduce a novel cross-attention architecture that utilizes spatial information from coordinate differences between relevant image patches. In doing so, our new cross-attention process dynamically considers both the related contents and their spatial relationships in caption generation. In addition, we introduce an efficient implementation of relative spatial attention based on convolutional operations. Experimental results show that the proposed spatial cross-attention improves captions to correctly describe the spatial relationships of objects, leading to an increase of 0.7 CIDEr score on the MS-COCO dataset compared to the previous state-of-the-art.
Khoa Anh Ngo, Kyuhong Shim, Byonghyo Shim
ICASSP2
2023 Task-Agnostic Open-Set Prototype for Few-Shot Open-Set Recognition
abstract
In few-shot open-set recognition (FSOSR), a network learns to recognize closed-set samples with a few support samples while rejecting open-set samples with no class cue. Unlike conventional OSR, the FSOSR considers more practical open worlds where a closed-set class can be selected as an open-set class in another testing (task) and vice versa. Existing FSOSR methods have commonly represented the open set with task-dependent extra modules. These modules decently handle the varied closed and open classes but accompany inevitable complexity increase. This paper shows that a single open-set prototype can represent open-set samples when it satisfies a specific relation in metric space: closest to open-set, and simultaneously second nearest to close-set. We propose a task-agnostic open-set prototype with distance scaling factors and design loss terms. We extensively analyze the proposed components to demonstrate their importance. Our method achieves state-of-the-art results on miniImageNet and tieredImageNet, respectively, without task-dependent extra modules.
Byeonggeun Kim, Juntae Lee, Kyuhong Shim, Simyung Chang
ICIP3
2023 Depth-Relative Self Attention for Monocular Depth Estimation
abstract
Monocular depth estimation is very challenging because clues to the exact depth are incomplete in a single RGB image. To overcome the limitation, deep neural networks rely on various visual hints such as size, shade, and texture extracted from RGB information. However, we observe that if such hints are overly exploited, the network can be biased on RGB information without considering the comprehensive view. We propose a novel depth estimation model named RElative Depth Transformer (RED-T) that uses relative depth as guidance in self-attention. Specifically, the model assigns high attention weights to pixels of close depth and low attention weights to pixels of distant depth. As a result, the features of similar depth can become more likely to each other and thus less prone to misused visual hints. We show that the proposed model achieves competitive results in monocular depth estimation benchmarks and is less biased to RGB information. In addition, we propose a novel monocular depth estimation benchmark that limits the observable depth range during training in order to evaluate the robustness of the model for unseen depths.
Kyuhong Shim, Gusang Lee, Byonghyo Shim
IJCAI1
2023 Knowledge Distillation from Non-streaming to Streaming ASR Encoder using Auxiliary Non-streaming Layer
Kyuhong Shim, Jinkyu Lee 0004, Simyoung Chang, Kyuwoong Hwang
INTERSPEECH1
2023 Improving Small Footprint Few-shot Keyword Spotting with Supervision on Auxiliary Data
Seunghan Yang, Byeonggeun Kim, Kyuhong Shim, Simyoung Chang
INTERSPEECH3
2022 Semantic Feature Extraction for Generalized Zero-Shot Learning
abstract
Generalized zero-shot learning (GZSL) is a technique to train a deep learning model to identify unseen classes using the attribute. In this paper, we put forth a new GZSL technique that improves the GZSL classification performance greatly. Key idea of the proposed approach, henceforth referred to as semantic feature extraction-based GZSL (SE-GZSL), is to use the semantic feature containing only attribute-related information in learning the relationship between the image and the attribute. In doing so, we can remove the interference, if any, caused by the attribute-irrelevant information contained in the image feature. To train a network extracting the semantic feature, we present two novel loss functions, 1) mutual information-based loss to capture all the attribute-related information in the image feature and 2) similarity-based loss to remove unwanted attribute-irrelevant information. From extensive experiments using various datasets, we show that the proposed SE-GZSL technique outperforms conventional GZSL approaches by a large margin.
Junhan Kim, Kyuhong Shim, Byonghyo Shim
AAAI2
2022 Understanding the Role of Self Attention for Efficient Speech Recognition
Kyuhong Shim, Jungwook Choi, Wonyong Sung
ICLR1
2022 Similarity and Content-based Phonetic Self Attention for Speech Recognition
Kyuhong Shim, Wonyong Sung
INTERSPEECH1
2017 SVD-Softmax: Fast Softmax Approximation on Large Vocabulary Neural Networks
abstract
We propose a fast approximation method of a softmax function with a very large vocabulary using singular value decomposition (SVD). SVD-softmax targets fast and accurate probability estimation of the topmost probable words during inference of neural network language models. The proposed method transforms the weight matrix used in the calculation of the output vector by using SVD. The approximate probability of each word can be estimated with only a small part of the weight matrix by using a few large singular values and the corresponding elements for most of the words. We applied the technique to language modeling and neural machine translation and present a guideline for good approximation. The algorithm requires only approximately 20\% of arithmetic operations for an 800K vocabulary case and shows more than a three-fold speedup on a GPU.
Kyuhong Shim, Iksoo Choi, Yoonho Boo, Wonyong Sung
NIPS1