VLDB 2026 Research / reviewers in the wild / expert
Jiyeon Kim
dblp:08/1021
· DBLP profile ↗
15ranked-venue papers
5as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Can Large Language Models Keep Up? Benchmarking Online Adaptation to Continual Knowledge StreamsabstractJiyeon Kim, Hyunji Lee, Dylan Zhou, Sue Hyun Park, Seunghyun Yoon, Trung Bui, Franck Dernoncourt, Sungmin Cha, Minjoon Seo. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jiyeon Kim, Hyunji Lee, Dylan Zhou, Sue Hyun Park, Seunghyun Yoon 0002, Trung Bui, Franck Dernoncourt, Sungmin Cha, Minjoon Seo |
ACL (1) | 1 |
| 2025 | Knowledge Entropy Decay during Language Model Pretraining Hinders New Knowledge AcquisitionabstractIn this work, we investigate how a model's tendency to broadly integrate its parametric knowledge evolves throughout pretraining, and how this behavior affects overall performance, particularly in terms of knowledge acquisition and forgetting. We introduce the concept of knowledge entropy, which quantifies the range of memory sources the model engages with; high knowledge entropy indicates that the model utilizes a wide range of memory sources, while low knowledge entropy suggests reliance on specific sources with greater certainty. Our analysis reveals a consistent decline in knowledge entropy as pretraining advances. We also find that the decline is closely associated with a reduction in the model's ability to acquire and retain knowledge, leading us to conclude that diminishing knowledge entropy (smaller number of active memory sources) impairs the model's knowledge acquisition and retention capabilities. We find further support for this by demonstrating that increasing the activity of inactive memory sources enhances the model's capacity for knowledge acquisition and retention. Jiyeon Kim, Hyunji Lee, Hyowon Cho, Joel Jang, Hyeonbin Hwang, Seungpil Won, Youbin Ahn, Dohaeng Lee, Minjoon Seo |
ICLR | 1 |
| 2025 | How Does Vision-Language Adaptation Impact the Safety of Vision Language Models?abstractVision-Language adaptation (VL adaptation) transforms Large Language Models (LLMs) into Large Vision-Language Models (LVLMs) for multimodal tasks, but this process often compromises the inherent safety capabilities embedded in the original LLMs. Despite potential harmfulness due to weakened safety measures, in-depth analysis on the effects of VL adaptation on safety remains under-explored. This study examines how VL adaptation influences safety and evaluates the impact of safety fine-tuning methods. Our analysis reveals that safety degradation occurs during VL adaptation, even when the training data is safe. While safety tuning techniques like supervised fine-tuning with safety datasets or reinforcement learning from human feedback mitigate some risks, they still lead to safety degradation and a reduction in helpfulness due to over-rejection issues. Further analysis of internal model weights suggests that VL adaptation may impact certain safety-related layers, potentially lowering overall safety levels. Additionally, our findings demonstrate that the objectives of VL adaptation and safety tuning are divergent, which often results in their simultaneous application being suboptimal. To address this, we suggest the weight merging approach as an optimal solution effectively reducing safety degradation while maintaining helpfulness. These insights help guide the development of more reliable and secure LVLMs for real-world applications. Seongyun Lee, Geewook Kim, Jiyeon Kim, Hyunji Lee, Hoyeon Chang, Sue Hyun Park, Minjoon Seo |
ICLR | 3 |
| 2025 | Towards the Objective Characterisation of Major Depressive Disorder Using Speech Data from a 12-week Observational Study with Daily Measurements
Robert Lewis 0001, Szymon Fedor, Nelson Hidalgo Julia, Joshua Curtiss, Jiyeon Kim, Noah Jones, David Mischoulon, Thomas F. Quatieri, Nicholas Cummins, Paola Pedrelli, Rosalind W. Picard |
INTERSPEECH | 5 |
| 2024 | ListT5: Listwise Reranking with Fusion-in-Decoder Improves Zero-shot RetrievalabstractWe propose LISTT5, a novel reranking approach based on Fusion-in-Decoder (FiD) that handles multiple candidate passages at both train and inference time.We also introduce an efficient inference framework for listwise ranking based on m-ary tournament sort with output caching.We evaluate and compare our model on the BEIR benchmark for zero-shot retrieval task, demonstrating that LISTT5 (1) outperforms the state-of-the-art RankT5 baseline with a notable +1.3 gain in the average NDCG@10 score, (2) has an efficiency comparable to pointwise ranking models and surpasses the efficiency of previous listwise ranking models, and (3) overcomes the lost-in-the-middle problem of previous listwise rerankers.Our code, model checkpoints, and the evaluation framework are fully open-sourced at https: //github.com/soyoung97/ListT5. Soyoung Yoon, Eunbi Choi, Jiyeon Kim, Hyeongu Yun, Yireun Kim, Seung-won Hwang |
ACL (1) | 3 |
| 2024 | Data Driven Grapheme-to-Phoneme Representations for a Lexicon-Free Text-to-SpeechabstractGrapheme-to-Phoneme (G2P) is an essential first step in any modern, high-quality Text-to-Speech (TTS) system. Most of the current G2P systems rely on carefully hand-crafted lexicons developed by experts. This poses a two-fold problem. Firstly, the lexicons are generated using a fixed phoneme set, usually, ARPABET or IPA, which might not be the most optimal way to represent phonemes for all languages. Secondly, the man-hours required to produce such an expert lexicon are very high. In this paper, we eliminate both of these issues by using recent advances in self-supervised learning to obtain data-driven phoneme representations instead of fixed representations. We compare our lexicon-free approach against strong baselines that utilize a well-crafted lexicon. Furthermore, we show that our data-driven lexicon-free method performs as good or even marginally better than the conventional rule-based or lexicon-based neural G2Ps in terms of Mean Opinion Score (MOS) while using no prior language lexicon or phoneme set, i.e. no linguistic expertise. Abhinav Garg, Jiyeon Kim, Sushil Khyalia, Chanwoo Kim 0001, Dhananjaya Gowda |
ICASSP | 2 |
| 2023 | Self-Supervised Accent Learning for Under-Resourced Accents Using Native Language DataabstractIn this paper, we propose a novel method to improve the accuracy of an English speech recognizer for a target accent using the corresponding native language data. Collecting labeled data for all accents of English to train an end-to-end neural speech recognizer for English is a difficult and expensive task. Also, finding a pool of representative English speakers for any arbitrary accent to collect unlabeled data can be a difficult task. However, collecting unlabeled speech data for any native language is a much simpler task. It is important to note that the accents of most non-native English speakers are heavily biased by the co-articulation of sounds in their own native language. In view of this, we propose to use unlabeled native language data to learn self-supervised representations during the pre-training stage. The pre-trained model is then fine-tuned using limited labeled English data for the target accent. Experiments using native language data to pre-train an English recognizer followed by fine-tuning using target accented English show significant improvements in word error rates on four different accents (Great Britain, Korean, Chinese, Spanish). Mehul Kumar, Jiyeon Kim, Dhananjaya Gowda, Abhinav Garg, Chanwoo Kim 0001 |
ICASSP | 2 |
| 2022 | Attention-guided RGB-D Fusion Network for Category-level 6D Object Pose EstimationabstractThis work focuses on estimating 6D poses and sizes of category-level objects from a single RGB-D image. How to exploit the complementary RGB and depth features plays an important role in this task yet remains an open question. Due to the large intra-category texture and shape variations, an object instance in test may have different RGB and depth features from those of the object instances in training, which poses challenges to previous RGB-D fusion methods. To deal with such problem, an Attention-guided RGB-D Fusion Network (ARF-Net) is proposed in this work. Our key design is an ARF module that learns to adaptively fuse RGB and depth features with guidance from both structure-aware attention and relation-aware attention. Specifically, the structure-aware attention captures spatial relationship among object parts and the relation-aware attention captures the RGB-to-depth correlations between the appearance and geometric features. Our ARF -Net directly establishes canonical correspondences with a compact decoder based on the multi-modal features from our ARF module. Extensive experiments show that our method can effectively fuse RGB features to various popular point cloud encoders and provide consistent performance improvement. In particular, without reconstructing instance 3D models, our method with its relatively compact architecture outperforms all state-of-the-art models on CAMERA25 and REAL275 benchmarks by a large margin. Hao Wang 0144, Jiyeon Kim, Qiang Wang 0023 |
IROS | 3 |
| 2021 | HiTNet: Byte-to-BPE Hierarchical Transcription Network for End-to-End Speech RecognitionabstractIn this paper, we propose a new byte to byte-pair-encoding (BPE) Hierarchical Transcription Network (HiTNet) architecture for end-to-end (e2e) automatic speech recognition (ASR). The proposed HiTNet architecture simultaneously encodes as well as decodes information hierarchically at different levels of linguistic granularity such as bytes and BPE. In general this idea can be extended to any levels of granularity including phonemes or graphemes or bytes (character to sub-character in some languages), to sub-words or byte-pair encodings (BPE), to words, and so on. Existing hierarchical e2e ASR models primarily encode the acoustic information in an hierarchical manner governed by weaker linguistic constraints at each level. The language information at each level is neither embedded or used explicitly, nor is the information decoded at each level passed on to the next stage. The proposed architecture primarily decodes information in an hierarchical manner utilizing the linguistic information at each level explicitly, while at the same time utilizing the hierarchically encoded acoustic information at each level. Experiments with a two-level byte-to-BPE (b2B) hierarchical transcription show that the proposed architecture significantly reduces the word error rates of both the byte and BPE decoders compared to baseline byte and BPE based attention encoder-decoder models. Dhananjaya Gowda, Abhinav Garg, Jiyeon Kim, Mehul Kumar, Nauman Dawalatabad, Aman Maghan, Shatrughan Singh, Chanwoo Kim 0001 |
ASRU | 3 |
| 2021 | Semi-Supervised Transfer Learning for Language Expansion of End-to-End Speech Recognition Models to Low-Resource Languages
Jiyeon Kim, Mehul Kumar, Dhananjaya Gowda, Abhinav Garg, Chanwoo Kim 0001 |
ASRU | 1 |
| 2021 | A Comparison of Streaming Models and Data Augmentation Methods for Robust Speech RecognitionabstractIn this paper, we present a comparative study on the robustness of two different online streaming speech recognition models: Monotonic Chunkwise Attention (MoChA) and Recurrent Neural Network-Transducer (RNN-T). We explore three recently proposed data augmentation techniques, namely, multi-conditioned training using an acoustic simulator, Vocal Tract Length Perturbation (VTLP) for speaker variability, and SpecAugment. Experimental results show that unidirectional models are in general more sensitive to noisy examples in the training set. It is observed that the final performance of the model depends on the proportion of training examples processed by data augmentation techniques. MoChA models generally perform better than RNN-T models. However, we observe that training of MoChA models seems to be more sensitive to various factors such as the characteristics of training sets and the incorporation of additional augmentations techniques. On the other hand, RNN-T models perform better than MoChA models in terms of latency, inference time, and the stability of training. Additionally, RNN-T models are generally more robust against noise and reverberation. All these advantages make RNN-T models a better choice for streaming on-device speech recognition compared to MoChA models. Jiyeon Kim, Mehul Kumar, Dhananjaya Gowda, Abhinav Garg, Chanwoo Kim 0001 |
ASRU | 1 |
| 2020 | Streaming On-Device End-to-End ASR System for Privacy-Sensitive Voice-Typing
Abhinav Garg, Gowtham P. Vadisetti, Dhananjaya Gowda, Sichen Jin, Aditya Jayasimha, Youngho Han, Jiyeon Kim, Junmo Park, Kwangyoun Kim, Young-Yoon Lee, Kyungbo Min, Chanwoo Kim 0001 |
INTERSPEECH | 7 |
| 2020 | Utterance Invariant Training for Hybrid Two-Pass End-to-End Speech Recognition
Dhananjaya Gowda, Kwangyoun Kim, Hejung Yang, Abhinav Garg, Jiyeon Kim, Mehul Kumar, Sichen Jin, Shatrughan Singh, Chanwoo Kim 0001 |
INTERSPEECH | 7 |
| 2019 | End-to-End Training of a Large Vocabulary End-to-End Speech Recognition SystemabstractIn this paper, we present an end-to-end training framework for building state-of-the-art end-to-end speech recognition systems. Our training system utilizes a cluster of Central Processing Units (CPUs) and Graphics Processing Units (GPUs). The entire data reading, large scale data augmentation, neural network parameter updates are all performed “on-the-fly”. We use vocal tract length perturbation [1] and an acoustic simulator [2] for data augmentation. The processed features and labels are sent to the GPU cluster. The Horovod allreduce approach is employed to train neural network parameters. We evaluated the effectiveness of our system on the standard Librispeech corpus [3] and the 10,000-hr anonymized Bixby English dataset. Our end-to-end speech recognition system built using this training infrastructure showed a 2.44 % WER on test-clean of the LibriSpeech test set after applying shallow fusion with a Transformer language model (LM). For the proprietary English Bixby open domain test set, we obtained a WER of 7.92 % using a Bidirectional Full Attention (BFA) end-to-end model after applying shallow fusion with an RNN-LM. When the monotonic chunckwise attention (MoCha) based approach is employed for streaming speech recognition, we obtained a WER of 9.95 % on the same Bixby open domain test set. Chanwoo Kim 0001, Minkyoo Shin, Shatrughan Singh, Larry Heck, Dhananjaya Gowda, Kwangyoun Kim, Mehul Kumar, Jiyeon Kim, Kyungmin Lee, Abhinav Garg, Eunhyang Kim |
ASRU | 9 |
| 2015 | PBAD: Perception-Based Anomaly Detection System for Cloud DatacentersabstractDetection of anomalies in large Cloud infrastructure is challenging. Understanding operational behavior of Cloud is extremely difficult due to the heterogeneity of different technologies, virtualized platforms and complex interactions among the systems. Many of existing system models for Cloud are based on utilization metrics such as CPU, memory, network and I/O. Such system models are quite complex and their anomaly detection mechanisms are mostly based on threshold scheme. Utilization metrics exceeding a certain threshold would trigger an alarm. In fact, it is impossible to determine proper threshold for all anomalies. These system models fail to assess the state of the system accurately. We propose a novel anomaly detection system based on user perception rather than complex system models. In our Perception-Based Anomaly Detection system (PBAD), each component within multi-tier applications monitors response time and determines whether overall service response time is adequate. PBAD also locates the anomaly by analyzing component behaviors. PBAD masks the complexity of Cloud and addresses what matters, how user perceives the service provided by the Cloud applications. The key advantages of the proposed algorithm are simplicity and scalability. We implement and deploy PBAD in our production data center environment. The experimental results show that PBAD detects numerous types of anomalies as well as the combination of anomalies where existing systems fail. Jiyeon Kim, Hyong S. Kim 0001 |
CLOUD | 1 |