Jaehun Kim

dblp:208/4323 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Computer networks · 2Human-computer interaction and ubiquitous computing · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Ethane: Debloating State Data using Compact Trie for Account-based Blockchain
abstract
Account-based blockchains can suffer from huge state data as the number of accounts soars, as in the current Ethereum. This makes it hard to synchronize and operate as a full node to verify transactions, or as an archive node to maintain all archive data for provenance queries. The problem is mostly caused by the state trie, a tree-like data structure to store account states in its leaves, where an account has a path key to follow to access the account. Whenever an account is updated in the state trie, all trie nodes along the path key from the leaf to the root are newly created, which causes a data explosion. In this paper, we propose a novel state optimization technique called Ethane. Instead of assigning a fixed, hash-based path key for an account as in Ethereum, Ethane assigns a variable, counter-based path key. So, when a transaction creates or updates an account, Ethane creates a new leaf node for the account and assigns a path key of the next counter value, which has an effect of adding the leaf to the rightmost slot of the trie. This compact trie maximizes common parent nodes and minimizes the creation of new non-leaf nodes, significantly reducing the archive data size. In addition, Ethane maintains two types of compact tries: an active trie for frequently updated accounts and an inactive trie for dormant accounts not updated for a long time (e.g., three months). Dormant accounts are transferred to the inactive trie but can be reactivated at any time via a restore transaction. When synchronizing as a full node, we only need to download the active trie and the rightmost path of the inactive trie to function fully, significantly reducing both storage requirements and synchronization overhead. Our evaluation shows that Ethane can downsize the archive state data by 60–85% and the current state trie by 60–94%. It also doubles the block execution performance.
Junmo Lee 0001, Jaehun Kim, Jiyong Youn, Soo-Mook Moon
EuroSys2
2026 SCORE: Scaling Audio Generation Using Standardized COmposite REwards
abstract
The goal of this letter is to enhance Text-to-Audio (T2A) generation at inference, focusing on generating realistic audio that precisely aligns with text prompts. Despite the rapid advancements, existing approaches often fail to achieve a reliable balance between semantic alignment and perceptual quality. We address this by adopting Inference-Time Scaling for the first time in T2A generation, and propose a multi-reward guidance that places each criterion on equal footing. By standardizing each reward to zero mean and unit variance before a weighted summation, the method reduces the scale mismatch that otherwise demands weight tuning, providing stable guidance at equal weights while leaving the weight as an interpretable control over the desired aspect. Moreover, we introduce a new text-audio alignment metric using an audio language model for more robust evaluation. Empirically, our method improves both semantic alignment and perceptual quality, achieving improved performance over naive generation and existing reward guidance techniques.
Jaemin Jung, Jaehun Kim, Inkyu Shin, Joon Son Chung
IEEE Signal Process. Lett.2
2025 From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech
abstract
The objective of this study is to generate high-quality speech from silent talking face videos, a task also known as video-to-speech synthesis. A significant challenge in video-to-speech synthesis lies in the substantial modality gap between silent video and multi-faceted speech. In this paper, we propose a novel video-to-speech system that effectively bridges this modality gap, significantly enhancing the quality of synthesized speech. This is achieved by learning of hierarchical representations from video to speech. Specifically, we gradually transform silent video into acoustic feature spaces through three sequential stages – content, timbre, and prosody modeling. In each stage, we align visual factors – lip movements, face identity, and facial expressions – with corresponding acoustic counterparts to ensure the seamless transformation. Additionally, to generate realistic and coherent speech from the visual representations, we employ a flow matching model that estimates direct trajectories from a simple prior distribution to the target speech distribution. Extensive experiments demonstrate that our method achieves exceptional generation quality comparable to real utterances, outperforming existing methods by a significant margin.
Jeongsoo Choi, Jaehun Kim, Chaeyoung Jung, Joon Son Chung
CVPR3
2025 AdaptVC: High Quality Voice Conversion with Adaptive Learning
abstract
The goal of voice conversion is to transform the speech of a source speaker to sound like that of a reference speaker while preserving the original content. A key challenge is to extract disentangled linguistic content from the source and voice style from the reference. While existing approaches leverage various methods to isolate the two, a generalization still requires further attention, especially for robustness in zero-shot scenarios. In this paper, we achieve successful disentanglement of content and speaker features by tuning self-supervised speech features with adapters. The adapters are trained to dynamically encode nuanced features from rich self-supervised features, and the decoder fuses them to produce speech that accurately resembles the reference with minimal loss of content. Moreover, we leverage a conditional flow matching decoder with cross-attention speaker conditioning to further boost the synthesis quality and efficiency. Subjective and objective evaluations in a zero-shot scenario demonstrate that the proposed method outperforms existing models in speech quality and similarity to the reference speech.
Jaehun Kim, Yeunju Choi, Tan Dat Nguyen, Seongkyu Mun, Joon Son Chung
ICASSP1
2024 Let There Be Sound: Reconstructing High Quality Speech from Silent Videos
abstract
The goal of this work is to reconstruct high quality speech from lip motions alone, a task also known as lip-to-speech. A key challenge of lip-to-speech systems is the one-to-many mapping caused by (1) the existence of homophenes and (2) multiple speech variations, resulting in a mispronounced and over-smoothed speech. In this paper, we propose a novel lip-to-speech system that significantly improves the generation quality by alleviating the one-to-many mapping problem from multiple perspectives. Specifically, we incorporate (1) self-supervised speech representations to disambiguate homophenes, and (2) acoustic variance information to model diverse speech styles. Additionally, to better solve the aforementioned problem, we employ a flow based post-net which captures and refines the details of the generated speech. We perform extensive experiments on two datasets, and demonstrate that our method achieves the generation quality close to that of real human utterance, outperforming existing methods in terms of speech naturalness and intelligibility by a large margin. Synthesised samples are available at our demo page: https://mm.kaist.ac.kr/projects/LTBS.
Jaehun Kim, Joon Son Chung
AAAI2
2024 Tempo Estimation as Fully Self-Supervised Binary Classification
abstract
This paper addresses the problem of global tempo estimation in musical audio. Given that annotating tempo is time-consuming and requires certain musical expertise, few publicly available data sources exist to train machine learning models for this task. Towards alleviating this issue, we propose a fully self-supervised approach that does not rely on any human labeled data. Our method builds on the fact that generic (music) audio embeddings already encode a variety of properties, including information about tempo, making them easily adaptable for downstream tasks. While recent work in self-supervised tempo estimation aimed to learn a tempo specific representation that was subsequently used to train a supervised classifier, we reformulate the task into the binary classification problem of predicting whether a target track has the same or a different tempo compared to a reference. While the former still requires labeled training data for the final classification model, our approach uses arbitrary unlabeled music data in combination with time-stretching for model training as well as a small set of synthetically created reference samples for predicting the final tempo. Evaluation of our approach in comparison with the state-of-the-art reveals highly competitive performance when the constraint of finding the precise tempo octave is relaxed.
Florian Henkel, Jaehun Kim, Matthew C. McCallum, Samuel E. Sandberg, Matthew E. P. Davies
ICASSP2
2024 Seeing Through The Conversation: Audio-Visual Speech Separation Based on Diffusion Model
abstract
The objective of this work is to extract the target speaker’s voice from a mixture of voices using visual cues. Existing works on audio-visual speech separation have demonstrated their performance with promising intelligibility, but maintaining naturalness remains challenging. To address this issue, we propose AVDiffuSS, an audio-visual speech separation model based on a diffusion mechanism known for its capability to generate natural samples. We also propose a cross-attention-based feature fusion mechanism for an effective fusion of the two modalities for diffusion. This mechanism is specifically tailored for the speech domain to integrate the phonetic information from audio-visual correspondence in speech generation. In this way, the fusion process maintains the high temporal resolution of the features, without excessive computational requirements. We demonstrate that the proposed framework achieves state-of-the-art results on two benchmarks, including VoxCeleb2 and LRS3, producing speech with notably better naturalness. Project page with demo: https://mm.kaist.ac.kr/projects/avdiffuss/
Suyeon Lee, Chaeyoung Jung, Youngjoon Jang 0001, Jaehun Kim, Joon Son Chung
ICASSP4
2024 On The Effect Of Data-Augmentation On Local Embedding Properties In The Contrastive Learning Of Music Audio Representations
abstract
Audio embeddings are crucial tools in understanding large catalogs of music. Typically embeddings are evaluated on the basis of the performance they provide in a wide range of downstream tasks, however few studies have investigated the local properties of the embedding spaces themselves which are important in nearest neighbor algorithms, commonly used in music search and recommendation. In this work we show that when learning audio representations on music datasets via contrastive learning, musical properties that are typically homogeneous within a track (e.g., key and tempo) are reflected in the locality of neighborhoods in the resulting embedding space. By applying appropriate data augmentation strategies, localisation of such properties can not only be reduced but the localisation of other attributes is increased. For example, locality of features such as pitch and tempo that are less relevant to non-expert listeners, may be mitigated while improving the locality of more salient features such as genre and mood, achieving state-of-the-art performance in nearest neighbor retrieval accuracy. Similarly, we show that the optimal selection of data augmentation strategies for contrastive learning of music audio embeddings is dependent on the downstream task, highlighting this as an important embedding design decision.
Matthew C. McCallum, Matthew E. P. Davies, Florian Henkel, Jaehun Kim, Samuel E. Sandberg
ICASSP4
2024 Similar but Faster: Manipulation of Tempo in Music Audio Embeddings for Tempo Prediction and Search
abstract
Audio embeddings enable large scale comparisons of the similarity of audio files for applications such as search and recommendation. Due to the subjectivity of audio similarity, it can be desirable to design systems that answer not only whether audio is similar, but similar in what way (e.g., wrt. tempo, mood or genre). Previous works have proposed disentangled embedding spaces where subspaces representing specific, yet possibly correlated, attributes can be weighted to emphasize those attributes in downstream tasks. However, no research has been conducted into the independence of these subspaces, nor their manipulation, in order to retrieve tracks that are similar but different in a specific way. Here, we explore the manipulation of tempo in embedding spaces as a case-study towards this goal. We propose tempo translation functions that allow for efficient manipulation of tempo within a pre-existing embedding space whilst maintaining other properties such as genre. As this translation is specific to tempo it enables retrieval of tracks that are similar but have specifically different tempi. We show that such a function can be used as an efficient data augmentation strategy for both training of downstream tempo predictors, and improved nearest neighbor retrieval of properties largely independent of tempo.
Matthew C. McCallum, Florian Henkel, Jaehun Kim, Samuel E. Sandberg, Matthew E. P. Davies
ICASSP3
2024 Fregrad: Lightweight and Fast Frequency-Aware Diffusion Vocoder
abstract
The goal of this paper is to generate realistic audio with a lightweight and fast diffusion-based vocoder, named FreGrad. Our framework consists of the following three key components: (1) We employ discrete wavelet transform that decomposes a complicated waveform into sub-band wavelets, which helps FreGrad to operate on a simple and concise feature space, (2) We design a frequency-aware dilated convolution that elevates frequency awareness, resulting in generating speech with accurate frequency information, and (3) We introduce a bag of tricks that boosts the generation quality of the proposed model. In our experiments, FreGrad achieves 3.7 times faster training time and 2.2 times faster inference speed compared to our baseline while reducing the model size by 0.6 times (only 1.78M parameters) without sacrificing the output quality. Audio samples are available at: https://mm.kaist.ac.kr/projects/FreGrad.
Tan Dat Nguyen, Youngjoon Jang 0001, Jaehun Kim, Joon Son Chung
ICASSP4
2022 AiRS: A Large-Scale Recommender System at NAVER News
abstract
Online news providers such as Google News, Bing News, and NAVER News collect a large number of news articles from a variety of presses and distribute these articles to users via their portals. Dynamic nature of a news domain causes the problem of information overload that makes it difficult for a user to find her preferable news articles. Motivated by this situation, NAVER Corp., the largest portal company in South Korea, identified four design considerations (DCs) for news recommendation that reflect the unique characteristics of a news domain. In this paper, we introduce a large-scale news recommender system named as AiRS, present how it jointly leverages the four DCs for NAVER News service. Specifically, AiRS first generates candidate articles for recommendation to a target user based on collaborative filtering (CF), quality estimation (QE), and social impact (SI) models; then, it ranks the candidate articles based on the scores computed by considering their multi-type feature scores (e.g., user's section preference and article's recency), finally recommending the top-$k$news articles that a target user is likely to prefer. Also, we present how to build the architecture for online deployment of AiRS at NAVER News. Through extensive offline and online A/B tests using the real-world datasets, we validate that AiRS successfully reflects all of the DCs into the news recommendation process, all design choices employed in AiRS help improve the recommendation accuracy, and AiRS significantly outperforms five state-of-the-art news recommendation approaches in terms of accuracy.
Hongjun Lim, Yeon-Chang Lee, Jin-Seo Lee, Sanggyu Han, Seunghyeon Kim, Yeon Jeong Jeong, Changbong Kim, Jaehun Kim, Sunghoon Han, Solbi Choi, Hanjong Ko, Dokyeong Lee, Hong-Kyun Bae, Taeho Kim 0003, Jeewon Ahn, Hyun-Soung You, Sang-Wook Kim
ICDE8
2020 Object shape recognition using tactile sensor arrays by a spiking neural network with unsupervised learning
abstract
The tactile properties of objects are important for robotic dexterous manipulation. An increasing number of attempts have recently been made to enable tactile information processing in robotic hand via tactile sensors. However, it remains relatively unexplored how to build tactile information processing models. In this study, we aimed to develop a spiking neural network (SNN) based on neural information processing mechanisms in sensory afferents. The SNN processes electrical signals collected from tactile sensor arrays attached to the gripper of the robotic hand while grasping objects with different shapes. We converted each of 42 -channel sensor signals from 2 arrays of 21 sensors into a spike train using the Izhikevich model, which was then fed to the SNN. The synaptic weights of the SNN were learned by the Hebbian learning through pair-based spike timing- dependent plasticity (STDP) algorithm. In addition, we implemented lateral inhibition of the second-layer neurons based on unsupervised learning similar to the one used in self-organizing maps, resulting in a winner-takes-all network. By this unsupervised learning, SNN could learn to discriminate the shape of objects via tactile sensing. In particular, it demonstrated object shape recognition with 100% accuracy. The proposed model could be useful for robots manipulating objects with tactile senses.
Jaehun Kim, Sung-Phil Kim, Heeseon Hwang, Doowon Park, Unyong Jeong
SMC1
2020 One deep music representation to rule them all? A comparative analysis of different representation learning strategies
abstract
Inspired by the success of deploying deep learning in the fields of Computer Vision and Natural Language Processing, this learning paradigm has also found its way into the field of Music Information Retrieval. In order to benefit from deep learning in an effective, but also efficient manner, deep transfer learning has become a common approach. In this approach, it is possible to reuse the output of a pre-trained neural network as the basis for a new learning task. The underlying hypothesis is that if the initial and new learning tasks show commonalities and are applied to the same type of input data (e.g., music audio), the generated deep representation of the data is also informative for the new task. Since, however, most of the networks used to generate deep representations are trained using a single initial learning source, their representation is unlikely to be informative for all possible future tasks. In this paper, we present the results of our investigation of what are the most important factors to generate deep representations for the data and learning tasks in the music domain. We conducted this investigation via an extensive empirical study that involves multiple learning sources, as well as multiple deep learning architectures with varying levels of information sharing between sources, in order to learn music representations. We then validate these representations considering multiple target datasets for evaluation. The results of our experiments yield several insights into how to approach the design of methods for learning widely deployable deep data representations in the music domain.
Jaehun Kim, Julián Urbano, Cynthia C. S. Liem, Alan Hanjalic
Neural Comput. Appl.1
2019 Dissecting 802.11ac Performance - Why You Should Turn Off MU-MIMO
abstract
While the recent Wi-Fi standard 802.11ac achieves Gb/s theoretical capacity with Multi-User MIMO (MU-MIMO) technology, several studies reported that throughput of 802.11ac in practice is far from Gb/s link speed. We investigate the downlink throughput of Wi-Fi systems with commercially available 802.11ac products in multiple indoor environments to reveal the throughput of MU-MIMO system that user experiences in practice. From our experiments, Single-User MIMO (SU-MIMO) outperformed MU-MIMO at every experimental environments. We further provide analysis on our experimental results considering channel sounding overhead, user grouping, environmental impact, and transmission mode selection.
Hyunwoo Choi, Taesik Gong, Jaehun Kim, Jaemin Shin 0005, Sung-Ju Lee 0001
MobiSys3
2019 Beyond Explicit Reports: Comparing Data-Driven Approaches to Studying Underlying Dimensions of Music Preference
abstract
Prior research from the field of music psychology has suggested that there are factors common to music preference beyond individual genres. Specifically, research has shown that self-reported ratings of preference for individual musical genres can be reduced to 4 or 5 dimensions, which in turn have been shown to correlate to relevant psychological constructs, such as personality. However, the number of dimensions emerging from multiple studies has varied despite the care taken in conducting such research. Data-driven approaches offer opportunities to further this line of research with actual listening data, at a scale and scope surpassing that of traditional psychological studies. Although listening data can be considered more direct and comprehensive evidence of listening preference, transforming this data into meaningful measurements is non-trivial. In the current paper, we report on investigations seeking to find interpretable underlying dimensions of music taste, using implicit large-scale listening data. Offering a critical reflection on potential researchers' degrees of freedom, we adopt an explicit systematic approach, investigating the impact of varying different parameters, analysis, and normalization techniques. More precisely, we consider various ways to extract listening preference information from two large, openly available datasets of music listening behavior, making use of principal component analysis and variational autoencoders to extract potential underlying dimensions. Results and implications are discussed in light of prior psychological theory, and the potential of user listening data to further research on music preference.
Jaehun Kim, Andrew M. Demetriou, Sandy Manolios, Cynthia C. S. Liem
UMAP1
2019 Use MU-MIMO at your own risk - Why we don't get Gb/s Wi-Fi
Hyunwoo Choi, Taesik Gong, Jaehun Kim, Jaemin Shin 0005, Sung-Ju Lee 0001
Ad Hoc Networks3