VLDB 2026 Research / reviewers in the wild / expert
Kyogu Lee
dblp:85/6128
· DBLP profile ↗
77ranked-venue papers
2as first author
46since 2021 · last 2025
0000-0002-4210-0312ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 51 · 1 first-author · 35 since 2021Artificial intelligence and machine learning · 32 · 1 first-author · 20 since 2021Databases, data management, data science and information retrieval · 6 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Uncertainty-Aware Self-Training for CTC-Based Automatic Speech RecognitionabstractUncertainty estimation has been widely applied for trustworthy automatic speech recognition (ASR) systems across training and inference stages. In the training stage, previous studies show that uncertainty can facilitate self-training by filtering out unlabeled data samples with high uncertainty. However, the current sequence-level uncertainty estimation method for connectionist temporal classification (CTC) based ASR models drops the output probability information and depends only on the textual distance of decoded predictions. In this study, we argue that this results in limited performance improvement and propose a novel output probability-based sequence-level uncertainty estimation method. We also categorize uncertainty as pseudo-label uncertainty and in-training uncertainty for the self-training process. Finally, we present uncertainty-aware self-training for CTC-based ASR models and experimentally show the effectiveness of the proposed method compared to the baselines. Eungbeom Kim, Kyogu Lee |
AAAI | 2 |
| 2025 | Exploring the Speech-to-Song Illusion: A Comparative Study of Standard Korean and Dialects
Haesun Joung, Ahyeon Choi, Kyogu Lee |
CogSci | 3 |
| 2025 | Neural responses of Interval Judgment in the Tritone Paradox
Subeen Kim, Jusung Ham, Inyong Choi, Kyogu Lee |
CogSci | 4 |
| 2025 | Variable Bitrate Residual Vector Quantization for Audio CodingabstractRecent state-of-the-art neural audio compression models have progressively adopted residual vector quantization (RVQ). Despite this success, these models employ a fixed number of codebooks per frame, which can be suboptimal in terms of rate-distortion tradeoff, particularly in scenarios with simple input audio, such as silence. To address this limitation, we propose variable bitrate RVQ (VRVQ) for audio codecs, which allows for more efficient coding by adapting the number of codebooks used per frame. Furthermore, we propose a gradient estimation method for the non-differentiable masking operation that transforms from the importance map to the binary importance mask, improving model training via a straight-through estimator. We demonstrate that the proposed training framework achieves superior results compared to the baseline method and shows further improvement when applied to the current state-of-the-art codec. Audio samples are available at: https://yoongi43.github.io/VBRRVQ.github.io/ Yunkee Chae, Woosung Choi, Yuhta Takida, Junghyun Koo, Yukara Ikemiya, Kin Wai Cheuk, Marco A. Martínez Ramírez, Kyogu Lee, Wei-Hsiang Liao 0001, Yuki Mitsufuji |
ICASSP | 9 |
| 2025 | DOSE: Drum One-Shot Extraction from Music MixtureabstractDrum one-shot samples are crucial for music production, particularly in sound design and electronic music. This paper introduces Drum One-Shot Extraction, a task in which the goal is to extract drum one-shots that are present in the music mixture. To facilitate this, we propose the Random Mixture One-shot Dataset (RMOD), comprising large-scale, randomly arranged music mixtures paired with corresponding drum one-shot samples. Our proposed model, Drum One-Shot Extractor (DOSE), leverages neural audio codec language models for end-to-end extraction, bypassing traditional source separation steps. Additionally, we introduce a novel onset loss, designed to encourage accurate prediction of the initial transient of drum one-shots, which is essential for capturing timbral characteristics. We compare this approach against a source separation-based extraction method as a baseline. The results, evaluated using Fréchet Audio Distance (FAD) and Multi-Scale Spectral loss (MSS), demonstrate that DOSE, enhanced with onset loss, outperforms the baseline, providing more accurate and higher-quality drum one-shots from music mixtures. The code, model checkpoint, and audio examples are available at https://github.com/HSUNEH/DOSE Suntae Hwang, Seonghyeon Kang, Semin Ahn, Kyogu Lee |
ICASSP | 5 |
| 2025 | Synthetic Dataset Generation for String Ensemble SeparationabstractMost studies on music source separation have traditionally concentrated on separating popular music into vocals, bass, drums, and others, with fewer studies focusing on chamber ensemble audio. The task of separating sources from chamber ensemble audio presents greater challenges due to factors like timbral similarity, high synchronization, and spectral overlap. These complexities necessitate more refined and realistic datasets for effective chamber ensemble separation. Yet, the creation of such datasets faces numerous challenges, depending on the methods of dataset formation. Taking these challenges into account, our work seeks to enhance the performance of chamber ensemble separation tasks, with a particular focus on string quartets. Based on dataset generation using a neural synthesis model, we propose an approach that incorporates musical expressions into the dataset generation process using MusicXML files. Furthermore, our framework adapt reverberation to the audio to better match the target mixtures. Our objective evaluation shows an increase in performance of chamber ensemble separation, supported by a subjective listening test to demonstrate improvement in our dataset. We also release our dataset and source code for further usage. Joonhyeon Bae, Eunsik Shin, Kyogu Lee |
ICASSP | 4 |
| 2025 | TokenSynth: A Token-based Neural Synthesizer for Instrument Cloning and Text-to-InstrumentabstractRecent advancements in neural audio codecs have enabled the use of tokenized audio representations in various audio generation tasks, such as text-to-speech, text-to-audio, and text-to-music generation. Leveraging this approach, we propose TokenSynth, a novel neural synthesizer that utilizes a decoder-only transformer to generate desired audio tokens from MIDI tokens and CLAP (Contrastive Language-Audio Pretraining) embedding, which has timbre-related information. Our model is capable of performing instrument cloning, text-to-instrument synthesis, and text-guided timbre manipulation without any fine-tuning. This flexibility enables diverse sound design and intuitive timbre control. We evaluated the quality of the synthesized audio, the timbral similarity between synthesized and target audio/text, and synthesis accuracy (i.e., how accurately it follows the input MIDI) using objective measures.TokenSynth demonstrates the potential of leveraging advanced neural audio codecs and transformers to create powerful and versatile neural synthesizers. The source code, model weights, and audio demos are available at: https://github.com/KyungsuKim42/tokensynth Junghyun Koo, Haesun Joung, Kyogu Lee |
ICASSP | 5 |
| 2025 | Speaking Without Sound: Multi-speaker Silent Speech Voicing with Facial Inputs OnlyabstractIn this paper, we introduce a novel framework for generating multi-speaker speech without relying on any audible inputs. Our approach leverages silent electromyography (EMG) signals to capture linguistic content, while facial images are used to match with the vocal identity of the target speaker. Notably, we present a pitch-disentangled content embedding that enhances the extraction of linguistic content from EMG signals. Extensive analysis demonstrates that our method can generate multi-speaker speech without any audible inputs and confirms the effectiveness of the proposed pitch-disentanglement approach. Yoori Oh, Kyogu Lee |
ICASSP | 3 |
| 2025 | Multidimensional Adaptive Coefficient for Inference Trajectory Optimization in Flow and DiffusionabstractFlow and diffusion models have demonstrated strong performance and training stability across various tasks but lack two critical properties of simulation-based methods: freedom of dimensionality and adaptability to different inference trajectories. To address this limitation, we propose the Multidimensional Adaptive Coefficient (MAC), a plug-in module for flow and diffusion models that extends conventional unidimensional coefficients to multidimensional ones and enables inference trajectory-wise adaptation. MAC is trained via simulation-based feedback through adversarial refinement. Empirical results across diverse frameworks and datasets demonstrate that MAC enhances generative quality with high training efficiency. Consequently, our work offers a new perspective on inference trajectory optimality, encouraging future research to move beyond vector field design and to leverage training-efficient, simulation-based optimization. Dohoon Lee, Hyunwoo J. Kim, Kyogu Lee |
ICML | 4 |
| 2025 | SynthRL: Cross-domain Synthesizer Sound Matching via Reinforcement LearningabstractGeneralization of synthesizer sound matching to external instrument sounds is highly challenging due to the non-differentiability of sound synthesis process which prohibits the use of out-of-domain sounds for training with synthesis parameter loss. We propose SynthRL, a novel reinforcement learning (RL)-based approach for cross-domain synthesizer sound matching. By incorporating sound similarity into the reward function, SynthRL effectively optimizes synthesis parameters without ground-truth labels, allowing fine-tuning on out-of-domain sounds. Furthermore, we introduce a transformer-based model architecture and reward-based prioritized experience replay to enhance RL training efficiency, considering the unique characteristics of the task. Experimental results demonstrate that SynthRL outperforms state-of-the-art methods on both in-domain and out-of-domain tasks. Further experimental analysis validates the effectiveness of our reward design, showing a strong correlation with human perception of sound similarity. Wonchul Shin, Kyogu Lee |
IJCAI | 2 |
| 2025 | Towards Bitrate-Efficient and Noise-Robust Speech Coding with Variable Bitrate RVQ
Yunkee Chae, Kyogu Lee |
INTERSPEECH | 2 |
| 2025 | Song Form-aware Full-Song Text-to-Lyrics Generation with Multi-Level Granularity Syllable Count Control
Yunkee Chae, Eunsik Shin, Suntae Hwang, Seungryeol Paik, Kyogu Lee |
INTERSPEECH | 5 |
| 2025 | Few-step Adversarial Schrödinger Bridge for Generative Speech Enhancement
Seungu Han, Juheon Lee, Kyogu Lee |
INTERSPEECH | 4 |
| 2025 | Voice-Based Dysphagia Detection: Leveraging Self-Supervised Speech Representation
Injune Hwang, Jung-Min Kim, Ju Seok Ryu, Kyogu Lee |
INTERSPEECH | 4 |
| 2025 | Vo-Ve: An Explainable Voice-Vector for Speaker Identity Evaluation
Kyogu Lee |
INTERSPEECH | 2 |
| 2025 | MGE-LDM: Joint Latent Diffusion for Simultaneous Music Generation and Source ExtractionabstractWe present MGE-LDM, a unified latent diffusion framework for simultaneous music generation, source imputation, and query-driven source separation. Unlike prior approaches constrained to fixed instrument classes, MGE-LDM learns a joint distribution over full mixtures, submixtures, and individual stems within a single compact latent diffusion model. At inference, MGE-LDM enables (1) complete mixture generation, (2) partial generation (i.e., source imputation), and (3) text-conditioned extraction of arbitrary sources. By formulating both separation and imputation as conditional inpainting tasks in the latent space, our approach supports flexible, class-agnostic manipulation of arbitrary instrument sources. Notably, MGE-LDM can be trained jointly across heterogeneous multi-track datasets (e.g., Slakh2100, MUSDB18, MoisesDB) without relying on predefined instrument categories. Yunkee Chae, Kyogu Lee |
NeurIPS | 2 |
| 2025 | Understanding Audio-Text Retrieval Through Singular Value DecompositionabstractAudio-language retrieval faces a fundamental challenge due to many-to-one mappings, where multiple audio instances correspond to the same or similar textual descriptions. This ambiguity leads to overlapping representations, making it difficult for models to learn discriminative features between samples. Queue-based contrastive learning helps mitigate this ambiguity by maintaining a large pool of negatives, improving retrieval performance. However, selecting appropriate hyperparameters, such as the queue size for optimal performance and the number of epochs to prevent overfitting, remains a challenge. In this study, we propose a LAtent Space Embedding Rank (LASER) analysis. Specifically, we analyze that a low-rank property of a queue indicates the presence of informative negative samples, and demonstrate that the best performance is achieved when the queue has the lowest rank. Furthermore, the convergence of rank variation can be interpreted as the model parameters reaching an optimal state during training. Thanks to the proposed rank-based analysis, the model hyperparameters can be selected analytically rather than through trial and error. Experimental results confirm that retrieval performance depends on queue rank, with lower ranks yielding better performance. These findings suggest that our LASER analysis can provide insight into how rank-based analysis can improve retrieval performance in audio-language tasks. Yoori Oh, Yoseob Han, Joonhyeon Bae, Jaeheon Sim, Kyogu Lee |
SIGIR | 5 |
| 2024 | The Effects of Musical Factors on the Perception of Auditory Illusions
Ahyeon Choi, Younyoung Bang, Jeong Mi Park, Kyogu Lee |
CogSci | 4 |
| 2024 | Music Auto-Tagging with Robust Music Representation Learned via Domain Adversarial TrainingabstractMusic auto-tagging is crucial for enhancing music discovery and recommendation. Existing models in Music Information Retrieval (MIR) struggle with real-world noise such as environmental and speech sounds in multimedia content. This study proposes a method inspired by speech-related tasks to enhance music auto-tagging performance in noisy settings. The approach integrates Domain Adversarial Training (DAT) into the music domain, enabling robust music representations that withstand noise. Unlike previous research, this approach involves an additional pretraining phase for the domain classifier, to avoid performance degradation in the subsequent phase. Adding various synthesized noisy music data improves the model’s generalization across different noise levels. The proposed architecture demonstrates enhanced performance in music auto-tagging by effectively utilizing unlabeled noisy music data. Additional experiments with supplementary unlabeled data further improves the model’s performance, underscoring its robust generalization capabilities and broad applicability. Haesun Joung, Kyogu Lee |
ICASSP | 2 |
| 2024 | Learning Semantic Information from Raw Audio Signal Using Both Contextual and Phonetic RepresentationsabstractWe propose a framework to learn semantics from raw audio signals using two types of representations, encoding contextual and phonetic information respectively. Specifically, we introduce a speech-to-unit processing pipeline that captures two types of representations with different time resolutions. For the language model, we adopt a dual-channel architecture to incorporate both types of representation. We also present new training objectives, masked context reconstruction and masked context prediction, that push models to learn semantics effectively. Experiments on the sSIMI metric of Zero Resource Speech Benchmark 2021 and Fluent Speech Command dataset show our framework learns semantics better than models trained with only one type of representation. Jaeyeon Kim, Injune Hwang, Kyogu Lee |
ICASSP | 3 |
| 2024 | String Sound Synthesizer On Gpu-Accelerated Finite Difference SchemeabstractThis paper introduces a nonlinear string sound synthesizer, based on a finite difference simulation of the dynamic behavior of strings under various excitations. The presented synthesizer features a versatile string simulation engine capable of stochastic parameterization, encompassing fundamental frequency modulation, stiffness, tension, frequency-dependent loss, and excitation control. This open-source physical model simulator not only benefits the audio signal processing community but also contributes to the burgeoning field of neural network-based audio synthesis by serving as a novel dataset construction tool. Implemented in PyTorch, this synthesizer offers flexibility, facilitating both CPU and GPU utilization, thereby enhancing its applicability as a simulator. GPU utilization expedites computation by parallelizing operations across spatial and batch dimensions, further enhancing its utility as a data generator. Jin Woo Lee 0001, Min Jun Choi, Kyogu Lee |
ICASSP | 3 |
| 2024 | DDD: A Perceptually Superior Low-Response-Time DNN-Based DeclipperabstractClipping is a common nonlinear distortion that occurs whenever the input or output of an audio system exceeds the supported range. This phenomenon undermines not only the perception of speech quality but also downstream processes utilizing the disrupted signal. Therefore, a real-time-capable, robust, and low-response-time method for speech declipping (SD) is desired. In this work, we introduce DDD (Demucs-Discriminator-Declipper), a real-time-capable speech-declipping deep neural network (DNN) that requires less response time by design. We first observe that a previously untested real-time-capable DNN model, Demucs, exhibits a reasonable declipping performance. Then we utilize adversarial learning objectives to increase the perceptual quality of output speech without additional inference overhead. Subjective evaluations on harshly clipped speech shows that DDD outperforms the baselines by a wide margin in terms of speech quality. We perform detailed waveform and spectral analyses to gain an insight into the output behavior of DDD in comparison to the baselines. Finally, our streaming simulations also show that DDD is capable of sub-decisecond mean response times, outperforming the state-of-the-art DNN approach by a factor of six. Jayeon Yi, Junghyun Koo, Kyogu Lee |
ICASSP | 3 |
| 2024 | Guiding Frame-Level CTC Alignments Using Self-knowledge Distillation
Eungbeom Kim, Hantae Kim, Kyogu Lee |
INTERSPEECH | 3 |
| 2024 | Hear Your Face: Face-based voice conversion with F0 estimation
Yoori Oh, Injune Hwang, Kyogu Lee |
INTERSPEECH | 4 |
| 2024 | Differentiable Modal Synthesis for Physical Modeling of Planar String Sound and Motion SimulationabstractWhile significant advancements have been made in music generation and differentiable sound synthesis within machine learning and computer audition, the simulation of instrument vibration guided by physical laws has been underexplored. To address this gap, we introduce a novel model for simulating the spatio-temporal motion of nonlinear strings, integrating modal synthesis and spectral modeling within a neural network framework. Our model leverages mechanical properties and fundamental frequencies as inputs, outputting string states across time and space that solve the partial differential equation characterizing the nonlinear string. Empirical evaluations demonstrate that the proposed architecture achieves superior accuracy in string motion simulation compared to existing baseline architectures. The code and demo are available online. Jin Woo Lee 0001, Min Jun Choi, Kyogu Lee |
NeurIPS | 4 |
| 2024 | Distance Sampling-based Paraphraser Leveraging ChatGPT for Text Data ManipulationabstractThere has been growing interest in audio-language retrieval research, where the objective is to establish the correlation between audio and text modalities. However, most audio-text paired datasets often lack rich expression of the text data compared to the audio samples. One of the significant challenges facing audio-text datasets is the presence of similar or identical captions despite different audio samples. Therefore, under many-to-one mapping conditions, audio-text datasets lead to poor performance of retrieval tasks. In this paper, we propose a novel approach to tackle the data imbalance problem in audio-language retrieval task. To overcome the limitation, we introduce a method that employs a distance sampling-based paraphraser leveraging ChatGPT, utilizing distance function to generate a controllable distribution of manipulated text data. For a set of sentences with the same context, the distance is used to calculate a degree of manipulation for any two sentences, and ChatGPT's few-shot prompting is performed using a text cluster with a similar distance defined by the Jaccard similarity. Therefore, ChatGPT, when applied to few-shot prompting with text clusters, can adjust the diversity of the manipulated text based on the distance. The proposed approach is shown to significantly enhance performance in audio-text retrieval, outperforming conventional text augmentation techniques. Yoori Oh, Yoseob Han, Kyogu Lee |
SIGIR | 3 |
| 2024 | Incorporating real-world object into virtual reality: using mobile device input with augmented virtuality
Jongkyu Shin, Kyogu Lee |
Multim. Tools Appl. | 2 |
| 2023 | Pop2Piano : Pop Audio-Based Piano Cover GenerationabstractPiano covers of pop music are enjoyed by many people. However, the task of automatically generating piano covers of pop music is still understudied. This is partly due to the lack of synchronized {Pop, Piano Cover} data pairs, which made it challenging to apply the latest data-intensive deep learning-based methods. To leverage the power of the data-driven approach, we make a large amount of paired and synchronized {Pop, Piano Cover } data using an automated pipeline. In this paper, we present Pop2Piano, a Transformer network that generates piano covers given waveforms of pop music. To the best of our knowledge, this is the first model to generate a piano cover directly from pop audio without using melody and chord extraction modules. We show that Pop2Piano, trained with our dataset, is capable of producing plausible piano covers. Jongho Choi, Kyogu Lee |
ICASSP | 2 |
| 2023 | Medleyvox: An Evaluation Dataset for Multiple Singing Voices SeparationabstractSeparation of multiple singing voices into each voice is a rarely studied area in music source separation research. The absence of a benchmark dataset has hindered its progress. In this paper, we present an evaluation dataset and provide baseline studies for multiple singing voices separation. First, we introduce MedleyVox, an evaluation dataset for multiple singing voices separation. We specify the problem definition in this dataset by categorizing it into i) unison, ii) duet, iii) main vs. rest, and iv) N-singing separation. Second, to overcome the absence of existing multi-singing datasets for a training purpose, we present a strategy for construction of multiple singing mixtures using various single-singing datasets. Third, we propose the improved super-resolution network (iSRNet), which greatly enhances initial estimates of separation networks. Jointly trained with the Conv-TasNet and the multi-singing mixture construction strategy, the proposed iSRNet achieved comparable performance to ideal time-frequency masks on duet and unison subsets of MedleyVox. Audio samples, the dataset, and codes are available on our website1. Chang-Bin Jeon, Hyeongi Moon, Keunwoo Choi, Ben Sangbae Chon, Kyogu Lee |
ICASSP | 5 |
| 2023 | Show Me the Instruments: Musical Instrument Retrieval From Mixture AudioabstractAs digital music production has become mainstream, the selection of appropriate virtual instruments plays a crucial role in determining the quality of music. To search the musical instrument samples or virtual instruments that make one’s desired sound, music producers use their ears to listen and compare each instrument sample in their collection, which is time-consuming and inefficient. In this paper, we call this task as Musical Instrument Retrieval and propose a method for retrieving desired musical instruments using reference mixture audio as a query. The proposed model consists of the Single-Instrument Encoder and the Multi-Instrument Encoder, both based on convolutional neural networks. The Single-Instrument Encoder is trained to classify the instruments used in single-track audio, and we take its penultimate layer’s activation as the instrument embedding. The Multi-Instrument Encoder is trained to estimate multiple instrument embeddings using the instrument embeddings computed by the Single-Instrument Encoder as a set of target embeddings. For more generalized training and realistic evaluation, we also propose a new dataset called Nlakh. Experimental results showed that the Single-Instrument Encoder was able to learn the mapping from the audio signal of unseen instruments to the instrument embedding space and the Multi-Instrument Encoder was able to extract multiple embeddings from the mixture audio and retrieve the desired instruments successfully. The code used for the experiment and audio samples are available at: https://github.com/minju0821/musical_instrument_retrieval Minju Park, Haesun Joung, Yunkee Chae, Yeongbeom Hong, Seonghyeon Go, Kyogu Lee |
ICASSP | 7 |
| 2023 | Music Mixing Style Transfer: A Contrastive Learning Approach to Disentangle Audio EffectsabstractWe propose an end-to-end music mixing style transfer system that converts the mixing style of an input multitrack to that of a reference song. This is achieved with an encoder pre-trained with a contrastive objective to extract only audio effects related information from a reference music recording. All our models are trained in a self-supervised manner from an already-processed wet multitrack dataset with an effective data preprocessing method that alleviates the data scarcity of obtaining unprocessed dry data. We analyze the proposed encoder for the disentanglement capability of audio effects and also validate its performance for mixing style transfer through both objective and subjective evaluations. From the results, we show the proposed system not only converts the mixing style of multitrack audio close to a reference but is also robust with mixture-wise style transfer upon using a music source separation model. Junghyun Koo, Marco A. Martínez Ramírez, Wei-Hsiang Liao 0001, Stefan Uhlich, Kyogu Lee, Yuki Mitsufuji |
ICASSP | 5 |
| 2023 | Neural Fourier Shift for Binaural Speech RenderingabstractWe present a neural network for rendering binaural speech from given monaural audio, position, and orientation of the source. Most of the previous works have focused on synthesizing binaural speeches by conditioning the positions and orientations in the feature space of convolutional neural networks. These synthesis approaches are powerful in estimating the target binaural speeches even for in-the-wild data but are difficult to generalize for rendering the audio from out-of-distribution domains. To alleviate this, we propose Neural Fourier Shift (NFS), a novel network architecture that enables binaural speech rendering in the Fourier space. Specifically, utilizing a geometric time delay based on the distance between the source and the receiver, NFS is trained to predict the delays and scales of various early reflections. NFS is efficient in both memory and computational cost, is interpretable, and operates independently of the source domain by its design. Experimental results show that NFS performs comparable to the previous studies on the benchmark dataset, even with its 25 times lighter memory and 6 times fewer calculations. Jin Woo Lee 0001, Kyogu Lee |
ICASSP | 2 |
| 2023 | Global HRTF Interpolation Via Learned Affine Transformation of Hyper-Conditioned FeaturesabstractEstimating Head-Related Transfer Functions (HRTFs) of arbitrary source points is essential in immersive binaural audio rendering. Computing each individual’s HRTFs is challenging, as traditional approaches require expensive time and computational resources, while modern data-driven approaches are data-hungry. Especially for the data-driven approaches, existing HRTF datasets differ in spatial sampling distributions of source positions, posing a major problem when generalizing the method across multiple datasets. To alleviate this, we propose a deep learning method based on a novel conditioning architecture. The proposed method can predict an HRTF of any position by interpolating the HRTFs of known distributions. Experimental results show that the proposed architecture improves the model’s generalizability across datasets with various coordinate systems. Additional demonstrations show that the model robustly reconstructs the target HRTFs from the spatially downsampled HRTFs in both quantitative and perceptual measures. Jin Woo Lee 0001, Kyogu Lee |
ICASSP | 3 |
| 2023 | Blind Estimation of Audio Processing GraphabstractMusicians and audio engineers sculpt and transform their sounds by connecting multiple processors, forming an audio processing graph. However, most deep-learning methods overlook this real-world practice and assume fixed graph settings. To bridge this gap, we develop a system that reconstructs the entire graph from a given reference audio. We first generate a realistic graph-reference pair dataset and train a simple blind estimation system composed of a convolutional reference encoder and a transformer-based graph decoder. We apply our model to singing voice effects and drum mixing estimation tasks. Evaluation results show that our method can reconstruct complex signal routings, including multi-band processing and sidechaining. Seungryeol Paik, Kyogu Lee |
ICASSP | 4 |
| 2023 | Debiased Automatic Speech Recognition for Dysarthric Speech via Sample Reweighting with Sample Affinity Test
Eungbeom Kim, Yunkee Chae, Jaeheon Sim, Kyogu Lee |
INTERSPEECH | 4 |
| 2023 | Semi-supervised Learning for Continuous Emotional Intensity Controllable Speech Synthesis with Disentangled Representations
Yoori Oh, Juheon Lee, Yoseob Han, Kyogu Lee |
INTERSPEECH | 4 |
| 2023 | Exploiting Time-Frequency Conformers for Music Audio EnhancementabstractWith the proliferation of video platforms on the internet, recording musical performances by mobile devices has become commonplace. However, these recordings often suffer from degradation such as noise and reverberation, which negatively impact the listening experience. Consequently, the necessity for music audio enhancement (referred to as music enhancement from this point onward), involving the transformation of degraded audio recordings into pristine high-quality music, has surged to augment the auditory experience. To address this issue, we propose a music enhancement system based on the Conformer architecture that has demonstrated outstanding performance in speech enhancement tasks. Our approach explores the attention mechanisms of the Conformer and examines their performance to discover the best approach for the music enhancement task. Our experimental results show that our proposed model achieves state-of-the-art performance on single-stem music enhancement. Furthermore, our system can perform general music enhancement with multi-track mixtures, which has not been examined in previous work. Audio samples enhanced with our system are available at: https://tinyurl.com/smpls9999 Yunkee Chae, Junghyun Koo, Kyogu Lee |
ACM Multimedia | 4 |
| 2022 | End-To-End Music Remastering System Using Self-Supervised And Adversarial TrainingabstractMastering is an essential step in music production, but it is also a challenging task that has to go through the hands of experienced audio engineers, where they adjust tone, space, and volume of a song. Remastering follows the same technical process, in which the context lies in mastering a song for the times. As these tasks have high entry barriers, we aim to lower the barriers by proposing an end-to-end music remastering system that transforms the mastering style of input audio to that of the target. The system is trained in a self-supervised manner, in which released pop songs were used for training. We also anticipated the model to generate realistic audio reflecting the reference’s mastering style by applying a pre-trained encoder and a projection discriminator. We validate our results with quantitative metrics and a subjective listening test and show that the model generated samples of mastering style similar to the target. Junghyun Koo, Seungryeol Paik, Kyogu Lee |
ICASSP | 3 |
| 2022 | Representation Selective Self-distillation and wav2vec 2.0 Feature Exploration for Spoof-aware Speaker VerificationabstractText-to-speech and voice conversion studies are constantly improving to the extent where they can produce synthetic speech almost indistinguishable from bona fide human speech. In this regard, the importance of countermeasures (CM) against synthetic voice attacks of the automatic speaker verification (ASV) systems emerges. Nonetheless, most end-to-end spoofing detection networks are black-box systems, and the answer to what is an effective representation for finding artifacts remains veiled. In this paper, we examine which feature space can effectively represent synthetic artifacts using wav2vec 2.0, and study which architecture can effectively utilize the space. Our study allows us to analyze which attribute of speech signals is advantageous for the CM systems. The proposed CM system achieved 0.31% equal error rate (EER) on ASVspoof 2019 LA evaluation set for the spoof detection task. We further propose a simple yet effective spoofing aware speaker verification (SASV) method, which takes advantage of the disentangled representations from our countermeasure system. Evaluation performed with the SASV Challenge 2022 database show 1.08% of SASV EER. Quantitative analysis shows that using the explored feature space of wav2vec 2.0 advantages both spoofing CM and SASV. Jin Woo Lee 0001, Eungbeom Kim, Junghyun Koo, Kyogu Lee |
INTERSPEECH | 4 |
| 2022 | Exploiting Negative Preference in Content-based Music Recommendation with Contrastive LearningabstractAdvanced music recommendation systems are being introduced along with the development of machine learning. However, it is essential to design a music recommendation system that can increase user satisfaction by understanding users’ music tastes, not by the complexity of models. Although several studies related to music recommendation systems exploiting negative preferences have shown performance improvements, there was a lack of explanation on how they led to better recommendations. In this work, we analyze the role of negative preference in users’ music tastes by comparing music recommendation models with contrastive learning exploiting preference (CLEP) but with three different training strategies - exploiting preferences of both positive and negative (CLEP-PN), positive only (CLEP-P), and negative only (CLEP-N). We evaluate the effectiveness of the negative preference by validating each system with a small amount of personalized data obtained via survey and further illuminate the possibility of exploiting negative preference in music recommendations. Our experimental results show that CLEP-N outperforms the other two in accuracy and false positive rate. Furthermore, the proposed training strategies produced a consistent tendency regardless of different types of front-end musical feature extractors, proving the stability of the proposed method. Minju Park, Kyogu Lee |
RecSys | 2 |
| 2022 | Differentiable Artificial ReverberationabstractArtificial reverberation (AR) models play a central role in various audio applications. Therefore, estimating the AR model parameters (ARPs) of a reference reverberation is a crucial task. Although a few recent deep-learning-based approaches have shown promising performance, their non-end-to-end training scheme prevents them from fully exploiting the potential of deep neural networks. This motivates the introduction of differentiable artificial reverberation (DAR) models, allowing loss gradients to be back-propagated end-to-end. However, implementing the AR models with their difference equations “as is” in the deep learning framework severely bottlenecks the training speed when executed with a parallel processor like GPU due to their infinite impulse response (IIR) components. We tackle this problem by replacing the IIR filters with finite impulse response (FIR) approximations with the frequency-sampling method. Using this technique, we implement three DAR models—differentiable Filtered Velvet Noise (FVN), Advanced Filtered Velvet Noise (AFVN), and Delay Network (DN). For each AR model, we train its ARP estimation networks for analysis-synthesis (RIR-to-ARP) and blind estimation (reverberant-speech-to-ARP) task in an end-to-end manner with its DAR model counterpart. Experiment results show that the proposed method achieves consistent performance improvement over the non-end-to-end approaches in both objective metrics and subjective listening test results. Hyeong-Seok Choi, Kyogu Lee |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Neural Audio Fingerprint for High-Specific Audio Retrieval Based on Contrastive LearningabstractMost of existing audio fingerprinting systems have limitations to be used for high-specific audio retrieval at scale. In this work, we generate a low-dimensional representation from a short unit segment of audio, and couple this fingerprint with a fast maximum inner-product search. To this end, we present a contrastive learning framework that derives from the segment-level search objective. Each update in training uses a batch consisting of a set of pseudo labels, randomly selected original samples, and their augmented replicas. These replicas can simulate the degrading effects on original audio signals by applying small time offsets and various types of distortions, such as background noise and room/microphone impulse responses. In the segment-level search task, where the conventional audio fingerprinting systems used to fail, our system using 10x smaller storage has shown promising results. Our code and dataset are available at https://mimbres.github.io/neural-audio-fp/. Sungkyun Chang, Donmoon Lee, Jeongsoo Park 0001, Hyungui Lim, Kyogu Lee, Karam Ko, Yoonchang Han |
ICASSP | 5 |
| 2021 | Real-Time Denoising and Dereverberation wtih Tiny Recurrent U-NetabstractModern deep learning-based models have seen outstanding performance improvement with speech enhancement tasks. The number of parameters of state-of-the-art models, however, is often too large to be deployed on devices for real-world applications. To this end, we propose Tiny Recurrent U-Net (TRU-Net), a lightweight online inference model that matches the performance of current state-of- the-art models. The size of the quantized version of TRU-Net is 362 kilobytes, which is small enough to be deployed on edge devices. In addition, we combine the small-sized model with a new masking method called phase-aware ß-sigmoid mask, which enables simultaneous denoising and dereverberation. Results of both objective and subjective evaluations have shown that our model can achieve competitive performance with the current state-of-the-art models on benchmark datasets using fewer parameters by orders of magnitude. Hyeong-Seok Choi, Sungjin Park 0003, Jie Hwan Lee, Hoon Heo, Dongsuk Jeon, Kyogu Lee |
ICASSP | 6 |
| 2021 | Reverb Conversion Of Mixed Vocal Tracks Using An End-To-End Convolutional Deep Neural NetworkabstractReverb plays a critical role in music production, where it provides listeners with spatial realization, timbre, and texture of the music. Yet, it is challenging to reproduce the musical reverb of a reference music track even by skilled engineers. In response, we propose an end-to-end system capable of switching the musical reverb factor of two different mixed vocal tracks. This method enables us to apply the reverb of the reference track to the source track to which the effect is desired. Further, our model can perform de-reverberation when the reference track is used as a dry vocal source. The proposed model is trained in combination with an adversarial objective, which makes it possible to handle high-resolution audio samples. The perceptual evaluation confirmed that the proposed model can convert the reverb factor with the preferred rate of 64.8%. To the best of our knowledge, this is the first attempt to apply deep neural networks to converting music reverb of vocal tracks. Junghyun Koo, Seungryeol Paik, Kyogu Lee |
ICASSP | 3 |
| 2021 | Room Adaptive Conditioning Method for Sound Event Classification in Reverberant EnvironmentsabstractEnsuring performance robustness for a variety of situations that can occur in real-world environments is one of the challenging tasks in sound event classification. One of the unpredictable and detrimental factors in performance, especially in indoor environments, is reverberation. To alleviate this problem, we propose a conditioning method that provides room impulse response (RIR) information to help the network become less sensitive to environmental information and focus on classifying the desired sound. Experimental results show that the proposed method successfully reduced performance degradation caused by the reverberation of the room. In particular, our proposed method works even with similar RIR that can be inferred from the room type rather than the exact one, which has the advantage of potentially being used in real-world applications. Donmoon Lee, Hyeong-Seok Choi, Kyogu Lee |
ICASSP | 4 |
| 2021 | Neural Analysis and Synthesis: Reconstructing Speech from Self-Supervised RepresentationsabstractWe present a neural analysis and synthesis (NANSY) framework that can manipulate the voice, pitch, and speed of an arbitrary speech signal. Most of the previous works have focused on using information bottleneck to disentangle analysis features for controllable synthesis, which usually results in poor reconstruction quality. We address this issue by proposing a novel training strategy based on information perturbation. The idea is to perturb information in the original input signal (e.g., formant, pitch, and frequency response), thereby letting synthesis networks selectively take essential attributes to reconstruct the input signal. Because NANSY does not need any bottleneck structures, it enjoys both high reconstruction quality and controllability. Furthermore, NANSY does not require any labels associated with speech data such as text and speaker information, but rather uses a new set of analysis features, i.e., wav2vec feature and newly proposed pitch feature, Yingram, which allows for fully self-supervised training. Taking advantage of fully self-supervised training, NANSY can be easily extended to a multilingual setting by simply training it with a multilingual dataset. The experiments show that NANSY can achieve significant improvement in performance in several applications such as zero-shot voice conversion, pitch shift, and time-scale modification. Hyeong-Seok Choi, Juheon Lee, Jie Lee, Hoon Heo, Kyogu Lee |
NeurIPS | 6 |
| 2020 | Musical Pitch Affects Brightness Judgment of a Concurrent Visual Object
You Jeong Hong, Ahyeon Choi, Chae-Eun Lee, Kyogu Lee |
CogSci | 4 |
| 2020 | Digital Watermarking For Protecting Audio Classification DatasetsabstractIn this study, we investigate the possibility of protecting audio classification datasets used in deep learning by embedding a pattern in the magnitude of the time-frequency representation of a subset of the dataset. Previous studies on audio watermarking technologies require the actual sound of the watermarked audio to extract the information embedded in it. In our study, we propose an audio watermarking framework aimed to identify whether a deep learning based audio classification model is trained with the watermarked audio classification dataset or not by using only the classification results. The experimental results show that our proposed method can identify the usage of an audio classification dataset while having minimal effect on the overall classification performance. The results are consistent with three different audio classification datasets. The proposed method is robust to different types and parameters of time-frequency representations and classification models. Kyogu Lee |
ICASSP | 2 |
| 2020 | Disentangling Timbre and Singing Style with Multi-Singer Singing Synthesis SystemabstractIn this study, we define the identity of the singer with two independent concepts – timbre and singing style – and propose a multi-singer singing synthesis system that can model them separately. To this end, we extend our single-singer model into a multi-singer model in the following ways: first, we design a singer identity encoder that can adequately reflect the identity of a singer. Second, we use encoded singer identity to condition the two independent decoders that model timbre and singing style, respectively. Through a user study with the listening tests, we experimentally verify that the proposed framework is capable of generating a natural singing voice of high quality while independently controlling the timbre and singing style. Also, by using the method of changing singing styles while fixing the timbre, we suggest that our proposed network can produce a more expressive singing voice. Juheon Lee, Hyeong-Seok Choi, Junghyun Koo, Kyogu Lee |
ICASSP | 4 |
| 2020 | From Inference to Generation: End-to-end Fully Self-supervised Generation of Human Face from Speech
Hyeong-Seok Choi, Changdae Park, Kyogu Lee |
ICLR | 3 |
| 2020 | Exploiting Multi-Modal Features from Pre-Trained Networks for Alzheimer's Dementia RecognitionabstractCollecting and accessing a large amount of medical data is very time-consuming and laborious, not only because it is difficult to find specific patients but also because it is required to resolve the confidentiality of a patient's medical records. On the other hand, there are deep learning models, trained on easily collectible, large scale datasets such as Youtube or Wikipedia, offering useful representations. It could therefore be very advantageous to utilize the features from these pre-trained networks for handling a small amount of data at hand. In this work, we exploit various multi-modal features extracted from pre-trained networks to recognize Alzheimer's Dementia using a neural network, with a small dataset provided by the ADReSS Challenge at INTERSPEECH 2020. The challenge regards to discern patients suspicious of Alzheimer's Dementia by providing acoustic and textual data. With the multi-modal features, we modify a Convolutional Recurrent Neural Network based structure to perform classification and regression tasks simultaneously and is capable of computing conversations with variable lengths. Our test results surpass baseline's accuracy by 18.75%, and our validation result for the regression task shows the possibility of classifying 4 classes of cognitive impairment with an accuracy of 78.70%. Junghyun Koo, Jie Hwan Lee, Jaewoo Pyo, Yujin Jo, Kyogu Lee |
INTERSPEECH | 5 |
| 2020 | Do Channels Matter? Illuminating Interpersonal Influence on Music RecommendationsabstractResearchers and service providers have focused on leveraging social information acquired from interactions between users to improve the accuracy of system recommendations. However, few have explained the characteristics of music recommendations through interpersonal relationships. To investigate how interpersonal relationships affect users’ evaluation of music recommendation, we conducted a survey-based study that compared two types of recommendation channels—interpersonal (i.e., from friends) and non-interpersonal (i.e., from systems). We found that relevance was evaluated higher in music recommended from non-interpersonal channels on average, while diversity, novelty, and serendipity were higher in interpersonal channels. Non-interpersonal channels surpassed interpersonal channels in terms of convenience, frequency, and adoption rate. These results illustrate that interpersonal and non-interpersonal channels have different strengths and that digital streaming platforms, which have mainly provided system recommendations thus far, need to better support interpersonal channels for richer user experience. Hyun Jeong Kim, So Yeon Park, Minju Park, Kyogu Lee |
RecSys | 4 |
| 2019 | Enhancing Music Features by Knowledge Transfer from User-item Log DataabstractIn this paper, we propose a novel method that exploits music listening log data for general-purpose music feature extraction. Despite the wealth of information available in the log data of user-item interactions, it has been mostly used for collaborative filtering to find similar items or users and was not fully investigated for content-based music applications. We resolve this problem by extending intra-domain knowledge distillation to cross-domain: i.e., by transferring knowledge obtained from the user-item domain to the music content domain. The proposed system first trains the model that estimates log information from the audio contents; then it uses the model to improve other task-specific models. The experiments on various music classification and regression tasks show that the proposed method successfully improves the performances of the task-specific models. Donmoon Lee, Jeongsoo Park 0001, Kyogu Lee |
ICASSP | 4 |
| 2019 | Phase-Aware Speech Enhancement with Deep Complex U-Net
Hyeong-Seok Choi, Jaesung Huh, Adrian Kim, Jung-Woo Ha 0001, Kyogu Lee |
ICLR (Poster) | 6 |
| 2019 | Adversarially Trained End-to-End Korean Singing Voice Synthesis SystemabstractIn this paper, we propose an end-to-end Korean singing voice synthesis system from lyrics and a symbolic melody using the following three novel approaches: 1) phonetic enhancement masking, 2) local conditioning of text and pitch to the super-resolution network, and 3) conditional adversarial training. The proposed system consists of two main modules; a mel-synthesis network that generates a mel-spectrogram from the given input information, and a super-resolution network that upsamples the generated mel-spectrogram into a linear-spectrogram. In the mel-synthesis network, phonetic enhancement masking is applied to generate implicit formant masks solely from the input text, which enables a more accurate phonetic control of singing voice. In addition, we show that two other proposed methods -- local conditioning of text and pitch, and conditional adversarial training -- are crucial for a realistic generation of the human singing voice in the super-resolution process. Finally, both quantitative and qualitative evaluations are conducted, confirming the validity of all proposed methods. Juheon Lee, Hyeong-Seok Choi, Chang-Bin Jeon, Junghyun Koo, Kyogu Lee |
INTERSPEECH | 5 |
| 2018 | Cover Song Identification Using Song-to-Song Cross-Similarity Matrix with Convolutional Neural NetworkabstractIn this paper, we propose a cover song identification algorithm using a convolutional neural network (CNN). We first train the CNN model to classify any non-/cover relationship, by feeding a cross-similarity matrix that is generated from a pair of songs as an input. Our main idea is to use the CNN output-the cover-probabilities of one song to all other candidate songs-as a new representation vector for measuring the distance between songs. Based on this, the present algorithm searches cover songs by applying several ranking methods: 1. sorting without using the representation vectors; 2. the cosine distance between the representation vectors; and 3. the correlation between the vectors. In our experiment, the proposed algorithm significantly outperformed the algorithms used in recent studies, by achieving a mean average precision (MAP) of 93.18% in a dataset consisting of 3,300 cover-pairs and 496,200 non-cover-pairs. Juheon Lee, Sungkyun Chang, Sang Keun Choe, Kyogu Lee |
ICASSP | 4 |
| 2017 | Utilizing context-relevant keywords extracted from a large collection of user-generated documents for music discovery
Ziwon Hyung, Joon-Sang Park, Kyogu Lee |
Inf. Process. Manag. | 3 |
| 2017 | Deep Convolutional Neural Networks for Predominant Instrument Recognition in Polyphonic MusicabstractIdentifying musical instruments in polyphonic music recordings is a challenging but important problem in the field of music information retrieval. It enables music search by instrument, helps recognize musical genres, or can make music transcription easier and more accurate. In this paper, we present a convolutional neural network framework for predominant instrument recognition in real-world polyphonic music. We train our network from fixed-length music excerpts with a single-labeled predominant instrument and estimate an arbitrary number of predominant instruments from an audio signal with a variable length. To obtain the audio-excerpt-wise result, we aggregate multiple outputs from sliding windows over the test audio. In doing so, we investigated two different aggregation methods: one takes the class-wise average followed by normalization, and the other perform temporally local class-wise max-pooling on the output probability prior to averaging and normalization steps to minimize the effect of averaging process suppresses the activation of sporadically appearing instruments. In addition, we conducted extensive experiments on several important factors that affect the performance, including analysis window size, identification threshold, and activation functions for neural networks to find the optimal set of parameters. Our analysis on the instrument-wise performance found that the onset type is a critical factor for recall and precision of each instrument. Using a dataset of 10k audio excerpts from 11 instruments for evaluation, we found that convolutional neural networks are more robust than conventional methods that exploit spectral features and source separation with support vector machines. Experimental results showed that the proposed convolutional network architecture obtained an F1 measure of 0.619 for micro and 0.513 for macro, respectively, achieving 23.1% and 18.8% in performance improvement compared with the state-of-the-art algorithm. Yoonchang Han, Jae-Hun Kim, Kyogu Lee |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2017 | Exploiting Continuity/Discontinuity of Basis Vectors in Spectrogram Decomposition for Harmonic-Percussive Sound SeparationabstractIn this paper, we present a novel method for harmonic-percussive sound separation (HPSS) by exploiting continuity/discontinuity properties in a matrix decomposition framework. It is widely accepted in the HPSS research that the harmonic and percussive components have anisotropic characteristics: The spectrum of the harmonic components and the time activation of the percussive components are sparse, whereas the spectrum of the percussive components and the time activation of the harmonic components are smooth. However, conventional methods fail to fully utilize the characteristics leading to suboptimal performance. Based on the observations that not the degree of sparseness but the degree of fluctuation is an accurate measure for distinguishing the harmonic and percussive components, we propose a novel HPSS algorithm by incorporating the continuity control in the iterative update formula of the matrix decomposition algorithm. In doing so, we first utilize probabilistic latent component analysis with Dirichlet prior, and later reformulate the algorithm in the nonnegative matrix factorization framework to reduce the computational cost. The comparative evaluation results show that the proposed method outperforms conventional methods in terms of both objective and subjective evaluation. Jeongsoo Park 0001, Jaeyoung Shin, Kyogu Lee |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Application of precise indoor position tracking to immersive virtual reality with translational movement support
Jongkyu Shin, Gwangseok An, Joon-Sang Park, Seungjun Baek 0001, Kyogu Lee |
Multim. Tools Appl. | 5 |
| 2015 | Informed source separation from monaural music with limited binary time-frequency annotationabstractThis paper presents a novel informed audio source separation algorithm given a limited binary time-frequency annotation. Assuming that all the sources can be represented using a low-rank model, we derive an objective function to minimize the rank of the source spectrogram, and the error between the target and the estimated coefficients. Especially, we apply the nuclear norm and l1-norm, which allow a relaxation of the model, and represent them in the convex formulation. Experimental results show that the proposed method achieves better and more robust separation performance than the state-of-the-art under the incomplete and inexact annotation condition. Il-Young Jeong, Kyogu Lee |
ICASSP | 2 |
| 2015 | Escaping your comfort zone: A graph-based recommender system for finding novel recommendations among relevant items
Kibeom Lee, Kyogu Lee |
Expert Syst. Appl. | 2 |
| 2015 | Enhanced auditory feedback for Korean touch screen keyboards
Yongki Park, Hoon Heo, Kyogu Lee |
Int. J. Hum. Comput. Stud. | 3 |
| 2014 | A pairwise approach to simultaneous onset/offset detection for singing voice using correntropyabstractIn this paper, we propose a novel method to search for precise locations of paired note onset and offset in a singing voice signal. In comparison with the existing onset detection algorithms, our approach differs in two key respects. First, we employ Correntropy, a generalized correlation function inspired from Reyni's entropy, as a detection function to capture the instantaneous flux while preserving insensitiveness to outliers. Next, a novel peak picking algorithm is specially designed for this detection function. By calculating the fitness of a pre-defined inverse hyperbolic kernel to a detection function, it is possible to find an onset and its corresponding offset simultaneously. Experimental results show that the proposed method achieves performance significantly better than or comparable to other state-of-the-art techniques for onset detection in singing voice. Sungkyun Chang, Kyogu Lee |
ICASSP | 2 |
| 2014 | A Highly Parallelized Decoder for Random Network Coding leveraging GPGPUabstractNetwork coding has been shown to improve various performance metrics in computer networks. However, the use of network coding, especially random linear network coding, incurs serious time delay in the decoding process and thus it is imperative to use a network coding implementation that has low decoding latency characteristics, e.g. a parallelized implementation. In this paper, we investigate the problem of parallelizing Pipeline network coding, a variant of random linear coding recently developed in order to alleviate the problems of random linear coding. We propose a novel massively parallelized decoding algorithm leveraging General Purpose Graphics Processing Unit (GPGPU) and show its performance enhancement by up to 100% compared with previous GPGPU-based parallel algorithms via experiments on real systems. Joon-Sang Park, Seungjun Baek 0001, Kyogu Lee |
Comput. J. | 3 |
| 2014 | Music recommendation using text analysis on song requests to radio stations
Ziwon Hyung, Kibeom Lee, Kyogu Lee |
Expert Syst. Appl. | 3 |
| 2014 | Application of non-negative spectrogram decomposition with sparsity constraints to single-channel speech enhancement
Kyogu Lee |
Speech Commun. | 1 |
| 2014 | Vocal Separation from Monaural Music Using Temporal/Spectral Continuity and Sparsity ConstraintsabstractIn this letter, we describe a novel approach for separating a vocal signal from monaural music. We assume that the accompaniment in a music signal can be represented as the sum of the sustained harmonic and percussive sounds. Based on the observation that singing voices usually contain rapidly changing harmonic signals such as fast vibratos, slides, and/or glissandos, we propose a statistical model for the separation of harmonic/percussive and vocal sounds. To this end, we define an objective function that exploits the temporal/spectral continuities of harmonic/percussive sounds and the sparsity of vocal sounds in the spectrogram domain. Experimental results show that the proposed algorithm successfully separates the vocal from the accompaniment, resulting in a performance significantly better than that of conventional algorithms or comparable to the state-of-the-art algorithms. Il-Young Jeong, Kyogu Lee |
IEEE Signal Process. Lett. | 2 |
| 2014 | Using Dynamically Promoted Experts for Music RecommendationabstractRecommender systems have become an invaluable asset to online services with the ever-growing number of items and users. Most systems focused on recommendation accuracy, predicting likable items for each user. Such methods tend to generate popular and safe recommendations, but fail to introduce users to potentially risky, yet novel items that could help in increasing the variety of items consumed by the users. This is known as popularity bias, which is predominant in methods that adopt collaborative filtering. Recently, however, recommenders have started to improve their methods to generate lists that encompass diverse items that are both accurate and novel through specific novelty-driven algorithms or hybrid recommender systems. In this paper, we propose a recommender system that uses the concepts of Experts to find both novel and relevant recommendations. By analyzing the ratings of the users, the algorithm promotes special Experts from the user population to create novel recommendations for a target user. Thus, different users are promoted dynamically to Experts depending on who the recommendations are for. The system used data collected from Last.fm and was evaluated with several metrics. Results show that the proposed system outperforms matrix factorization methods in finding novel items and performs on par in finding simultaneously novel and relevant items. This system can also provide a means to popularity bias while preserving the advantages of collaborative filtering. Kibeom Lee, Kyogu Lee |
IEEE Trans. Multim. | 2 |
| 2013 | Using Music Notation for Teaching Computer ProgrammingabstractDespite the wealth of educational programming languages, many novice programmers face difficulties and give up in the early stages, just because they are not familiar with the programming syntax and semantics. In this paper, we propose a method for programming langauge education using music notation with an aim to entice novice programmers to write their own programs. There are two key aspects in our proposed approach: first, we use music notation as an analogy to programming, based on the observation that there are similar attributes between the two; second, we provide users with on-line auditory feedback to immediately notify potential errors in a pleasant way. We find several key concepts in programming language syntax and semantics, and translate them into music notation to help beginner programmers learn them with ease and intuition. In addition, we design examples and a learning support environment, allowing users to learn to program by themselves. Eunjeong Ko, Kyogu Lee |
ICCE | 2 |
| 2013 | Note onset detection based on harmonic cepstrum regularityabstractThis paper presents a novel onset detection algorithm based on cepstral analysis. Instead of considering unnecessary mel-scale or any interests of non-harmonic components, we selectively focus on the changes in particular cepstral coefficients that represent the harmonic structure of an input signal. In comparison with a conventional time-frequency analysis, the advantage of using cepstral coefficients is that it shows the harmonic structure more clearly, and gives a robust detection function even when the envelope of waveform fluctuates or slowly increases. As a detection function, harmonic cepstrum regularity (HCR) is derived by the summation of several harmonic cepstral coefficients, but their frequency indices are defined from the previous frame so as to reflect the temporal changes in the harmonic structure. Experiments show that the proposed algorithm achieves significant improvement in performance over other algorithms, particularly for pitched instruments with soft onsets, such as violin and singing voice. Hoon Heo, Dooyong Sung, Kyogu Lee |
ICME | 3 |
| 2013 | Music similarity-based approach to generating dance motion sequence
Kyogu Lee, Jaeheung Park |
Multim. Tools Appl. | 2 |
| 2011 | My head is your tail: applying link analysis on long-tailed music listening behavior for music recommendationabstractCollaborative filtering, being a popular method for generating recommendations, produces satisfying results for users by providing extremely relevant items. Despite being popular, however, this method is prone to many problems. One of these problems is popularity bias, in which the system becomes skewed towards items that are popular amongst the general user population. These 'obvious' items are, technically, extremely relevant items but fail to be novel. In this paper, we maintain using collaborative filtering methods while still managing to produce novel yet relevant items. This is achieved by utilizing the long-tailed distribution of listening behavior of users, in which their playlists are biased towards a few songs while the rest of the songs, those in the long tail, have relatively low play counts. In addition, we also apply a link analysis method to users and define links between them to create an increasingly fine-grained approach in calculating weights for the recommended items. The proposed recommendation method was available online as a user study in order to measure the relevancy and novelty of the recommended items. Results show that the algorithm manages to include novel recommendations that are still relevant, and shows the potential for a new way of generating novel recommendations. Kibeom Lee, Kyogu Lee |
RecSys | 2 |
| 2010 | A super-resolution spectrogram using coupled PLCAabstractThe short-time Fourier transform (STFT) based spectrogram is commonly used to analyze the time-frequency content of a signal. Depending on window size, the STFT provides a trade-off between time and frequency resolutions. This paper presents a novel method that achieves high resolution simultaneously in both time and frequency. We extend Probabilistic Latent Component Analysis (PLCA) to jointly decompose two spectrograms, one with a high time resolution and one with a high frequency resolution. Using this decomposition, a new spectrogram, maintaining high resolution in both time and frequency, is constructed. Termed the “super-resolution spectrogram”, it can be particularly useful for speech as it can simultaneously resolve both glottal pulses and individual harmonics. Juhan Nam, Gautham J. Mysore, Joachim Ganseman, Kyogu Lee, Jonathan S. Abel |
INTERSPEECH | 4 |
| 2009 | Towards a Class-Based Representation of Perceptual Tempo for Music RetrievalabstractTempo is a common criterion by which humans describe and categorize music, and this has spawned a large amount of research in the field of automatic tempo estimation. Most tempo estimation systems focus mainly on detecting the temporal repetition and periodicity present within a signal, and represent tempo as a count of beats-per-minute (BPM). However, in real-world music retrieval applications such as music navigation and playlist generation, a rough perceptual representation of tempo may be more appropriate than a BPM representation. In this paper, the problem of tempo estimation is presented as a statistical classification problem. Four perceptual tempo classes are defined which correspond to rough semantic terms that average users may use to describe tempo. Statistical models of each class are built using low-level audio features. Experimental results show that the perceptual tempo class representation outperforms several conventional BPM-based tempo estimation systems when applied to the tasks of music navigation and playlist generation. Ching-Wei Chen, Kyogu Lee, Ho-Hsiang Wu |
ICMLA | 2 |
| 2008 | Acoustic Chord Transcription and Key Extraction From Audio Using Key-Dependent HMMs Trained on Synthesized AudioabstractWe describe an acoustic chord transcription system that uses symbolic data to train hidden Markov models and gives best-of-class frame-level recognition results. We avoid the extremely laborious task of human annotation of chord names and boundaries-which must be done to provide machine learning models with ground truth-by performing automatic harmony analysis on symbolic music files. In parallel, we synthesize audio from the same symbolic files and extract acoustic feature vectors which are in perfect alignment with the labels. We, therefore, generate a large set of labeled training data with a minimal amount of human labor. This allows for richer models. Thus, we build 24 key-dependent HMMs, one for each key, using the key information derived from symbolic data. Each key model defines a unique state-transition characteristic and helps avoid confusions seen in the observation vector. Given acoustic input, we identify a musical key by choosing a key model with the maximum likelihood, and we obtain the chord sequence from the optimal state path of the corresponding key model, both of which are returned by a Viterbi decoder. This not only increases the chord recognition accuracy, but also gives key information. Experimental results show the models trained on synthesized data perform very well on real recordings, even though the labels automatically generated from symbolic data are not 100% accurate. We also demonstrate the robustness of the tonal centroid feature, which outperforms the conventional chroma feature. Kyogu Lee, Malcolm Slaney |
IEEE Trans. Speech Audio Process. | 1 |
| 2005 | Explicit onset modeling of sinusoids using time reassignmentabstractWe introduce a system that explicitly models the onsets of sinusoidal signal components. The system uses time reassignment data to detect probable onsets. When an onset is detected, the reassigned data is used to estimate the precise location in time of the original onset, allowing synthesis of the corresponding output. This is advantageous over conventional time reassignment which implicitly smears onsets. We demonstrate the efficacy of our system on synthetic and real test signals with sudden onsets. Aaron S. Master, Kyogu Lee |
ICASSP (3) | 2 |