EDBT 2026 Demo / reviewers in the wild / expert
Juhan Nam
dblp:16/8308
· DBLP profile ↗
34ranked-venue papers
2as first author
21since 2021 · last 2026
0000-0003-2664-2119ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 1 first-author · 18 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ArtistMus: A Globally Diverse, Artist-Centric Benchmark for Retrieval-Augmented Music Question AnsweringabstractRecent advances in large language models (LLMs) have transformed open-domain question answering, yet their effectiveness in music-related reasoning remains limited due to sparse music knowledge in pretraining data. While music information retrieval and computational musicology have explored structured and multimodal understanding, few resources support factual and contextual music question answering (MQA) grounded in artist metadata or historical context. We introduce MusWikiDB, a vector database of 3.2M passages from 144K music-related Wikipedia pages, and ArtistMus, a benchmark of 1,000 questions on 500 diverse artists with metadata such as genre, debut year, and topic. These resources enable systematic evaluation of retrieval-augmented generation (RAG) for MQA. Experiments show that RAG markedly improves factual accuracy; open-source models gain up to +56.8 percentage points (for example, Qwen3 8B improves from 35.0 to 91.8), approaching proprietary model performance. RAG-style fine-tuning further boosts both factual recall and contextual reasoning, improving results on both in-domain and out-of-domain benchmarks. MusWikiDB also yields approximately 6 percentage points higher accuracy and 40% faster retrieval than a general-purpose Wikipedia corpus. We release MusWikiDB and ArtistMus to advance research in music information retrieval and domain-specific question answering, establishing a foundation for retrieval-augmented reasoning in culturally rich domains such as music. Daeyong Kwon, Seungheon Doh, Juhan Nam |
LREC | 3 |
| 2025 | FlashSR: One-step Versatile Audio Super-resolution via Diffusion DistillationabstractVersatile audio super-resolution (SR) is the challenging task of restoring high-frequency components from low-resolution audio with sampling rates between 4kHz and 32kHz in various domains such as music, speech, and sound effects. Previous diffusion-based SR methods suffer from slow inference due to the need for a large number of sampling steps. In this paper, we introduce FlashSR, a single-step diffusion model for versatile audio super-resolution aimed at producing 48kHz audio. FlashSR achieves fast inference by utilizing diffusion distillation with three objectives: distillation loss, adversarial loss, and distribution-matching distillation loss. We further enhance performance by proposing the SR Vocoder, which is specifically designed for SR models operating on mel-spectrograms. FlashSR demonstrates competitive performance with the current state-of-the-art model in both objective and subjective evaluations while being approximately 22 times faster. Jaekwon Im, Juhan Nam |
ICASSP | 2 |
| 2025 | D3RM: A Discrete Denoising Diffusion Refinement Model for Piano TranscriptionabstractDiffusion models have been widely used in the generative domain due to their convincing performance in modeling complex data distributions. Moreover, they have shown competitive results on discriminative tasks, such as image segmentation. While diffusion models have also been explored for automatic music transcription, their performance has yet to reach a competitive level. In this paper, we focus on discrete diffusion model’s refinement capabilities and present a novel architecture for piano transcription. Our model utilizes Neighborhood Attention layers as the denoising module, gradually predicting the target high-resolution piano roll, conditioned on the finetuned features of a pretrained acoustic model. To further enhance refinement, we devise a novel strategy which applies distinct transition states during training and inference stage of discrete diffusion models. Experiments on the MAESTRO dataset show that our approach outperforms previous diffusion-based piano transcription models and the baseline model in terms of F1 score. Our code is available on the accompanying website1. Hounsu Kim, Taegyun Kwon, Juhan Nam |
ICASSP | 3 |
| 2025 | Twenty-Five Years of MIR Research: Achievements, Practices, Evaluations, and Future ChallengesabstractIn this paper, we trace the evolution of Music Information Retrieval (MIR) over the past 25 years. While MIR gathers all kinds of research related to music informatics, a large part of it focuses on signal processing techniques for music data, fostering a close relationship with the IEEE Audio and Acoustic Signal Processing Technical Commitee. In this paper, we reflect the main research achievements of MIR along the three EDICS related to music analysis, processing and generation. We then review a set of successful practices that fuel the rapid development of MIR research. One practice is the annual research benchmark, the Music Information Retrieval Evaluation eXchange, where participants compete on a set of research tasks. Another practice is the pursuit of reproducible and open research. The active engagement with industry research and products is another key factor for achieving large societal impacts and motivating younger generations of students to join the field. Last but not the least, the commitment to diversity, equity and inclusion ensures MIR to be a vibrant and open community where various ideas, methodologies, and career pathways collide. We finish by providing future challenges MIR will have to face. Geoffroy Peeters, Zafar Rafii, Magdalena Fuentes, Zhiyao Duan, Emmanouil Benetos, Juhan Nam, Yuki Mitsufuji |
ICASSP | 6 |
| 2025 | Designing VR Music Game for Stress ReductionabstractMany individuals experience everyday stress. Effective stress management in daily life is crucial before this stress accumulates. Music has been extensively studied as a method for reducing stress. In particular, music therapy is widely used to reduce stress and enhance the well-being of various clinical groups. However, traditional music therapy has physical constraints that require patients to visit a therapeutic location. VR music therapy has been studied to address these issues, but most research focuses on receptive music therapy, failing to utilize VR’s interactive potential fully. Additionally, the potential of applying gamification to VR active music therapy to enhance user engagement and encourage long-term use of the therapy application has not been explored. This paper proposes VR active music therapy based on conventional rhythm-following music therapy methods and VR gamified active music therapy by adding game elements based on Self-Determination Theory (SDT). Our between-subject comparative study (n=33) revealed the stress reduction effects of VR receptive music therapy, VR active music therapy, and VR gamified active music therapy. Importantly, participant interviews provided valuable insights into the user experiences with each VR content, confirming the potential for long-term use of VR gamified active music therapy. Moreover, this research delves into the effect of gamification on stress reduction. Through pilot test and experiments, we identify game elements that could potentially increase stress and provide guidelines for applying gamification to mitigate these factors, thereby enhancing the effectiveness of VR gamified active music therapy. Kirak Kim, Youjin Choi, Juhan Nam, Jeongmi Lee |
VR | 4 |
| 2024 | K-pop Lyric Translation: Dataset, Analysis, and Neural-ModellingabstractLyric translation, a field studied for over a century, is now attracting computational linguistics researchers. We identified two limitations in previous studies. Firstly, lyric translation studies have predominantly focused on Western genres and languages, with no previous study centering on K-pop despite its popularity. Second, the field of lyric translation suffers from a lack of publicly available datasets; to the best of our knowledge, no such dataset exists. To broaden the scope of genres and languages in lyric translation studies, we introduce a novel singable lyric translation dataset, approximately 89% of which consists of K-pop song lyrics. This dataset aligns Korean and English lyrics line-by-line and section-by-section. We leveraged this dataset to unveil unique characteristics of K-pop lyric translation, distinguishing it from other extensively studied genres, and to construct a neural lyric translation model, thereby underscoring the importance of a dedicated dataset for singable lyric translations. Haven Kim, Jongmin Jung, Dasaem Jeong, Juhan Nam |
LREC/COLING | 4 |
| 2024 | T-Foley: A Controllable Waveform-Domain Diffusion Model for Temporal-Event-Guided Foley Sound SynthesisabstractFoley sound, audio content inserted synchronously with videos, plays a critical role in the user experience of multimedia content. Recently, there has been active research in Foley sound synthesis, leveraging the advancements in deep generative models. However, such works mainly focus on replicating a single sound class or a textual sound description, neglecting temporal information, which is crucial in the practical applications of Foley sound. We present T-Foley, a Temporal-event-guided waveform generation model for Foley sound synthesis. T-Foley generates high-quality audio using two conditions: the sound class and temporal event feature. For temporal conditioning, we devise a temporal event feature and a novel conditioning technique named Block-FiLM. T-Foley achieves superior performance in both objective and subjective evaluation metrics and generates Foley sound well-synchronized with the temporal events. Additionally, we showcase T-Foley’s practical applications, particularly in scenarios involving vocal mimicry for temporal event control. We show the demo on our companion website.1 Yoonjin Chung, Junwon Lee, Juhan Nam |
ICASSP | 3 |
| 2024 | Enriching Music Descriptions with A Finetuned-LLM and Metadata for Text-to-Music RetrievalabstractText-to-Music Retrieval, finding music based on a given natural language query, plays a pivotal role in content discovery within extensive music databases. To address this challenge, prior research has predominantly focused on a joint embedding of music audio and text, utilizing it to retrieve music tracks that exactly match descriptive queries related to musical attributes (i.e. genre, instrument) and contextual elements (i.e. mood, theme). However, users also articulate a need to explore music that shares similarities with their favorite tracks or artists, such as I need a similar track to Superstition by Stevie Wonder. To address these concerns, this paper proposes an improved Text-to-Music Retrieval model, denoted as TTMR++, which utilizes rich text descriptions generated with a finetuned large language model and metadata. To accomplish this, we obtained various types of seed text from several existing music tag and caption datasets and a knowledge graph dataset of artists and tracks. The experimental results show the effectiveness of TTMR++ in comparison to state-of-the-art music-text joint embedding models through a comprehensive evaluation involving various musical text queries.1 Seungheon Doh, Dasaem Jeong, Juhan Nam |
ICASSP | 4 |
| 2024 | DiffRENT: A Diffusion Model for Recording Environment Transfer of SpeechabstractProperly setting up recording conditions, including microphone type and placement, room acoustics, and ambient noise, is essential to obtaining the desired acoustic characteristics of speech. In this paper, we propose Diff-R-EN-T, a Diffusion model for Recording ENvironment Transfer which transforms the input speech to have the recording conditions of a reference speech while preserving the speech content. Our model comprises the content enhancer, the recording environment encoder, and the diffusion decoder which generates the target mel-spectrogram by utilizing both enhancer and encoder as input conditions. We evaluate DiffRENT in the speech enhancement and acoustic matching scenarios. The results show that DiffRENT generalizes well to unseen environments and new speakers. Also, the proposed model achieves superior performances in objective and subjective evaluation. Sound examples of our proposed model are available online1. Jaekwon Im, Juhan Nam |
ICASSP | 2 |
| 2024 | Expressive Acoustic Guitar Sound Synthesis with an Instrument-Specific Input Representation and Diffusion OutpaintingabstractSynthesizing performing guitar sound is a highly challenging task due to the polyphony and high variability in expression. Recently, deep generative models have shown promising results in synthesizing expressive polyphonic instrument sounds from music scores, often using a generic MIDI input. In this work, we propose an expressive acoustic guitar sound synthesis model with a customized input representation to the instrument, which we call guitarroll. We implement the proposed approach using diffusion-based outpainting which can generate audio with long-term consistency. To overcome the lack of MIDI/audio-paired datasets, we used not only an existing guitar dataset but also collected data from a high quality sample-based guitar synthesizer. Through quantitative and qualitative evaluations, we show that our proposed model has higher audio quality than the baseline model and generates more realistic timbre sounds than the previous leading work. Hounsu Kim, Soonbeom Choi, Juhan Nam |
ICASSP | 3 |
| 2024 | VoiceLDM: Text-to-Speech with Environmental ContextabstractThis paper presents VoiceLDM, a model designed to produce audio that accurately follows two distinct natural language text prompts: the description prompt and the content prompt. The former provides information about the overall environmental context of the audio, while the latter conveys the linguistic content. To achieve this, we adopt a text-to-audio (TTA) model based on latent diffusion models and extend its functionality to incorporate an additional content prompt as a conditional input. By utilizing pretrained contrastive language-audio pretraining (CLAP) and Whisper, VoiceLDM is trained on large amounts of real-world audio without manual annotations or transcriptions. Additionally, we employ dual classifier-free guidance to further enhance the controllability of VoiceLDM. Experimental results demonstrate that VoiceLDM is capable of generating plausible audio that aligns well with both input conditions, even surpassing the speech intelligibility of the ground truth audio on the AudioCaps test set. Furthermore, we explore the text-to-speech (TTS) and zero-shot text-to-audio capabilities of VoiceLDM and show that it achieves competitive results. Demos and code are available at https://voiceldm.github.io. Yeonghyeon Lee, Inmo Yeon, Juhan Nam, Joon Son Chung |
ICASSP | 3 |
| 2024 | Enhancing Spatial Audio Generation with Source Separation and Channel Panning LossabstractSpatial audio is essential for many immersive content services; however, it is challenging to obtain or create it. Recently, multimodal-based ambisonic audio generation has emerged as a promising approach for addressing the limitation. It combines multiple modalities, such as audio and video, and provides more intuitive control of ambisonic audio generation. Moreover, it leverages the advantages of machine-learning methods to automatically learn the correlation between different features and generate high-quality ambisonic sounds. Herein, we propose a separation- and localization-based spatial audio generation model. First, the network extracts visual features and separates audio into sound sources. Then, it conducts localization by mapping the separated sound sources to the visual features. To overcome the performance limitation of the previous self-supervised source separation approach, we employ a pretrained source separator with superior performance. To improve the localization performance further, we propose a channel panning loss function between each channel of the ambisonic signal. We use three different types of datasets to train the model experimentally and evaluate the proposed method with four metrics. The results show that the proposed model achieves better spatialization performance than the baseline models. Wootaek Lim, Juhan Nam |
ICASSP | 2 |
| 2024 | A Real-Time Lyrics Alignment System Using Chroma and Phonetic Features for Classical Vocal PerformanceabstractThe goal of real-time lyrics alignment is to take live singing audio as input and to pinpoint the exact position within given lyrics on the fly. The task can benefit real-world applications such as the automatic subtitling of live concerts or operas. However, designing a real-time model poses a great challenge due to the constraints of only using past input and operating within a minimal latency. Furthermore, due to the lack of datasets for real-time models for lyrics alignment, previous studies have mostly evaluated with private in-house datasets, resulting in a lack of standard evaluation methods. This paper presents a real-time lyrics alignment system for classical vocal performances with two contributions. First, we improve the lyrics alignment algorithm by finding an optimal combination of chromagram and phonetic posteriorgram (PPG) that capture melodic and phonetics features of the singing voice, respectively. Second, we recast the Schubert Winterreise Dataset (SWD) which contains multiple performance renditions of the same pieces as an evaluation set for the real-time lyrics alignment. Jiyun Park, Sangeon Yong, Taegyun Kwon, Juhan Nam |
ICASSP | 4 |
| 2024 | Towards Efficient and Real-Time Piano Transcription Using Neural Autoregressive ModelsabstractIn recent years, advancements in neural network designs and the availability of large-scale labeled datasets have led to significant improvements in the accuracy of piano transcription models. However, most previous work focused on high-performance offline transcription, neglecting deliberate consideration of model size. The goal of this work is to implement real-time piano transcription with a focus on achieving both high performance and a lightweight model. To this end, we propose novel architectures for convolutional recurrent neural networks, redesigning an existing autoregressive piano transcription model. First, we extend the acoustic module by adding a frequency-conditioned FiLM layer to the CNN module to adapt the convolutional filters on the frequency axis. Second, we improve note-state sequence modeling by using a pitchwise LSTM that focuses on note-state transitions within a note. In addition, we augment the autoregressive connection with an enhanced recursive context. Using these components, we propose two types of models; one for high performance and the other for high compactness. Through extensive experiments, we demonstrate that the proposed components are necessary for achieving high performance in an autoregressive model. Additionally, we provide experiments on real-time latency. Taegyun Kwon, Dasaem Jeong, Juhan Nam |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Textless Speech-to-Music Retrieval Using Emotion SimilarityabstractWe introduce a framework that recommends music based on the emotions of speech. In content creation and daily life, speech contains information about human emotions, which can be enhanced by music. Our framework focuses on a cross-domain retrieval system to bridge the gap between speech and music via emotion labels. We explore different speech representations and report their impact on different speech types, including acting voice and wake-up words. We also propose an emotion similarity regularization term in cross-domain retrieval tasks. By incorporating the regularization term into training, similar speech-and-music pairs in the emotion space are closer in the joint embedding space. Our comprehensive experimental results show that the proposed model is effective in textless speech-to-music retrieval. Seungheon Doh, Minz Won, Keunwoo Choi, Juhan Nam |
ICASSP | 4 |
| 2023 | Toward Universal Text-To-Music RetrievalabstractThis paper introduces effective design choices for text-to-music retrieval systems. An ideal text-based retrieval system would support various input queries such as pre-defined tags, unseen tags, and sentence-level descriptions. In reality, most previous works mainly focused on a single query type (tag or sentence) which may not generalize to another input type. Hence, we review recent text-based music retrieval systems using our proposed benchmark in two main aspects: input text representation and training objectives. Our findings enable a universal text-to-music retrieval system that achieves comparable retrieval performances in both tag- and sentence-level inputs. Furthermore, the proposed multimodal representation generalizes to 9 different downstream music classification tasks. We present the code and demo online.1 Seungheon Doh, Minz Won, Keunwoo Choi, Juhan Nam |
ICASSP | 4 |
| 2023 | A Study of Audio Mixing Methods for Piano Transcription in Violin-Piano EnsemblesabstractWhile piano music transcription models have shown high performance for solo piano recordings, their performance de-grades when applied to ensemble recordings. This study aims to analyze the impact of different data augmentation methods on piano transcription performance, specifically focusing on mixing techniques applied to violin-piano ensembles. We apply mixing methods that consider both harmonic and temporal characteristics of the audio. To create datasets for this study, we generated the PFVN-synth dataset, which contains 7 hours of violin-piano ensemble audio by rendering MIDI files and corresponding labels, and also collected unaccompanied violin recordings and mixed them with the MAESTRO dataset. We evaluated the transcription results on both synthesized and real audio recordings datasets. Hyemi Kim, Jiyun Park, Taegyun Kwon, Dasaem Jeong, Juhan Nam |
ICASSP | 5 |
| 2023 | A Phoneme-Informed Neural Network Model For Note-Level Singing TranscriptionabstractNote-level automatic music transcription is one of the most representative music information retrieval (MIR) tasks and has been studied for various instruments to understand music. However, due to the lack of high-quality labeled data, transcription of many instruments is still a challenging task. In particular, in the case of singing, it is difficult to find accurate notes due to its expressiveness in pitch, timbre, and dynamics. In this paper, we propose a method of finding note onsets of singing voice more accurately by leveraging the linguistic characteristics of singing, which are not seen in other instruments. The proposed model uses mel-scaled spectrogram and phonetic posteriorgram (PPG), a frame-wise likelihood of phoneme, as an input of the onset detection network while PPG is generated by the pre-trained network with singing and speech data. To verify how linguistic features affect onset detection, we compare the evaluation results through the dataset with different languages and divide onset types for detailed analysis. Our approach substantially improves the performance of singing transcription and therefore emphasizes the importance of linguistic features in singing analysis. Sangeon Yong, Juhan Nam |
ICASSP | 3 |
| 2022 | A Melody-Unsupervision Model for Singing Voice SynthesisabstractRecent studies in singing voice synthesis have achieved high-quality results leveraging advances in text-to-speech models based on deep neural networks. One of the main issues in training singing voice synthesis models is that they require melody and lyric labels to be temporally aligned with audio data. The temporal alignment is a time-exhausting manual work in preparing for the training data. To address the issue, we propose a melody-unsupervision model that requires only audio-and-lyrics pairs without temporal alignment in training time but generates singing voice audio given a melody and lyrics input in inference time. The proposed model is composed of a phoneme classifier and a singing voice generator jointly trained in an end-to-end manner. The model can be fine-tuned by adjusting the amount of supervision with temporally aligned melody labels. Through experiments in melody-unsupervision and semi-supervision settings, we compare the audio quality of synthesized singing voice. We also show that the proposed model is capable of being trained with speech audio and text labels but can generate singing voice in inference time. Soonbeom Choi, Juhan Nam |
ICASSP | 2 |
| 2022 | Pseudo-Label Transfer from Frame-Level to Note-Level in a Teacher-Student Framework for Singing Transcription from Polyphonic MusicabstractLack of large-scale note-level labeled data is the major obstacle to singing transcription from polyphonic music. We address the issue by using pseudo labels from vocal pitch estimation models given unlabeled data. The proposed method first converts the frame-level pseudo labels to note-level through pitch and rhythm quantization steps. Then, it further improves the label quality through self-training in a teacher-student framework. To validate the method, we conduct various experiment settings by investigating two vocal pitch estimation models as pseudo-label generators, two setups of teacher-student frameworks, and the number of iterations in self-training. The results show that the proposed method can effectively leverage large-scale unlabeled audio data and self-training with the noisy student model helps to improve performance. Finally, we show that the model trained with only unlabeled data has comparable performance to previous works and the model trained with additional labeled data achieves higher accuracy than the model trained with only labeled data. Sangeun Kum, Jongpil Lee, Luke K. Kim, Juhan Nam |
ICASSP | 5 |
| 2022 | Deformable CNN and Imbalance-Aware Feature Learning for Singing Technique ClassificationabstractSinging techniques are used for expressive vocal performances by employing temporal fluctuations of the timbre, the pitch, and other components of the voice.Their classification is a challenging task, because of mainly two factors: 1) the fluctuations in singing techniques have a wide variety and are affected by many factors and 2) existing datasets are imbalanced.To deal with these problems, we developed a novel audio feature learning method based on deformable convolution with decoupled training of the feature extractor and the classifier using a classweighted loss function.The experimental results show the following: 1) the deformable convolution improves the classification results, particularly when it is applied to the last two convolutional layers, and 2) both re-training the classifier and weighting the cross-entropy loss function by a smoothed inverse frequency enhance the classification performance. Yuya Yamamoto, Juhan Nam, Hiroko Terasawa |
INTERSPEECH | 2 |
| 2020 | Korean Singing Voice Synthesis Based on Auto-Regressive Boundary Equilibrium GanabstractSinging voice synthesis is a generative task that involves not only multidimensional controls of a singer model such as phonetic modulation by lyrics and pitch control by music score but also expressive elements such as breath sounds and vibrato. Recently, end-to-end learning models based on generative adversarial network (GAN) have drawn much interest as it requires less domain-specific processing but provides high sound quality. When GAN is applied to the audio domain, it entails several issues: the choice of audio representation to generate, handling temporal continuity between two adjacent outputs, and finding an effective loss metric for the audio representation. In this paper, we propose a Korean singing voice synthesis system that addresses the issues using an auto-regressive algorithm that generates spectrogram with the boundary equilibrium GAN objective. Through the qualitative test, we show the proposed methods are superior to the original GAN objective and non-auto-regressive model. We also show that our proposed method can render natural expressions such as continuous pitch contours and breath sounds. Soonbeom Choi, Wonil Kim, Saebyul Park, Sangeon Yong, Juhan Nam |
ICASSP | 5 |
| 2020 | Disentangled Multidimensional Metric Learning for Music SimilarityabstractMusic similarity search is useful for a variety of creative tasks such as replacing one music recording with another recording with a similar "feel", a common task in video editing. For this task, it is typically necessary to define a similarity metric to compare one recording to another. Music similarity, however, is hard to define and depends on multiple simultaneous notions of similarity (i.e. genre, mood, instrument, tempo). While prior work ignore this issue, we embrace this idea and introduce the concept of multidimensional similarity and unify both global and specialized similarity metrics into a single, semantically disentangled multidimensional similarity metric. To do so, we adapt a variant of deep metric learning called conditional similarity networks to the audio domain and extend it using track-based information to control the specificity of our model. We evaluate our method and show that our single, multidimensional model outperforms both specialized similarity spaces and alternative baselines. We also run a user-study and show that our approach is favored by human annotators as well. Jongpil Lee, Nicholas J. Bryan, Justin Salamon, Zeyu Jin, Juhan Nam |
ICASSP | 5 |
| 2020 | Semantic Tagging of Singing Voices in Popular Music RecordingsabstractSinging voice is a key sound source in popular music. As recent music streaming and entertainment services call for more intelligent solutions to retrieve songs or evaluate musical characteristics, automatic analysis of popular music targeted to singing voice has been a significant research subject. The majority of studies have focused on quantitative or objective information of singing voice such as pitch, lyrics or singer identity. However, singing voice has a wide variety of dimensions that are somewhat difficult to quantify and therefore we often describe by words. In this article, we address the qualitative analysis of singing voice as a music auto-tagging task that annotates songs with a set of tag words. To this end, we build a music tag dataset dedicated to singing voice. Specifically, we define a vocabulary that describes timbre and singing styles of K-pop vocalists and collect human annotations for individual tracks. We then conduct statistical analysis to understand the global and temporal characteristics of the tag words. Using the dataset, we train a deep neural network model to automatically predict the voice-specific tags from popular music recordings and evaluate the model in different conditions. We discuss the results by comparing them to the statistical analysis of tag words. Finally, we show potential applications of the vocal tagging system in music retrieval, music thumbnailing and singing evaluation. Luke K. Kim, Jongpil Lee, Sangeun Kum, Chae Lin Park, Juhan Nam |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2019 | Graph Neural Network for Music Score Data and Modeling Expressive Piano PerformanceabstractMusic score is often handled as one-dimensional sequential data. Unlike words in a text document, notes in music score can be played simultaneously by the polyphonic nature and each of them has its own duration. In this paper, we represent the unique form of musical score using graph neural network and apply it for rendering expressive piano performance from the music score. Specifically, we design the model using note-level gated graph neural network and measure-level hierarchical attention network with bidirectional long short-term memory with an iterative feedback method. In addition, to model different styles of performance for a given input score, we employ a variational auto-encoder. The result of the listening test shows that our proposed model generated more human-like performances compared to a baseline model and a hierarchical attention network model that handles music score as a word-like sequence. Dasaem Jeong, Taegyun Kwon, Yoojin Kim, Juhan Nam |
ICML | 4 |
| 2018 | Sample-Level CNN Architectures for Music Auto-Tagging Using Raw WaveformsabstractRecent work has shown that the end-to-end approach using convolutional neural network (CNN) is effective in various types of machine learning tasks. For audio signals, the approach takes raw waveforms as input using an 1-D convolution layer. In this paper, we improve the 1-D CNN architecture for music auto-tagging by adopting building blocks from state-of-the-art image classification models, ResNets and SENets, and adding multi-level feature aggregation to it. We compare different combinations of the modules in building CNN architectures. The results show that they achieve significant improvements over previous state-of-the-art models on the MagnaTagATune dataset and comparable results on Million Song Dataset. Furthermore, we analyze and visualize our model to show how the 1-D CNN operates. Taejun Kim, Jongpil Lee, Juhan Nam |
ICASSP | 3 |
| 2018 | Singing Expression Transfer from One Voice to Another for a Given SongabstractWe present a vocal processing algorithm to automatically transfer singing expressions from one voice to another for a given song. Depending on singers' competence, a song can be rendered with great variations in terms of local tempo, pitch and dynamics. The proposed method temporally aligns a pair of singing voices using melodic and lyrical features that they have in common. Then, it conducts time-scale modification on the source voice according to the time-stretching ratio from the alignment result after smoothing. Once the two voices are aligned, the method modifies pitch and energy expressions of the source voice in a frame-by-frame manner using a pitch-synchronous overlap-add algorithm and a simple amplitude envelope matching. We designed our experiment to transfer singing expressions from a highly technical singer to a plain singer. The results show that our proposed method improves the singing quality effectively. Sangeon Yong, Juhan Nam |
ICASSP | 2 |
| 2017 | ForceClicks: Enabling Efficient Button Interaction with Single Finger TouchabstractForceClicks is a novel touch button input technique for consecutive clicking which incorporates touch force sensors. From force data of a single continuous touch over time, ForceClicks detects peaks and generates discrete clicks. Compared to typical button interaction, this is effective in a sense that consecutive clicks do not require finger positional movements. Additionally, stable force over a certain time threshold can trigger an alternate state, long press, and can be mapped to other actions. The usability of ForceClicks has been evaluated in terms of a) scattering level and b) efficiency. Results suggest higher stability than typical touch, especially when the task requires visual engagement on remote content. The relatively scatter-free characteristic of ForceClicks allows it to be applied on rapid clicking while gaming, and reduce of visual dedication allows easier control of external devices, and two applications, a shooting game and a number picker, are presented for demonstration. Sangeon Yong, Edward Jangwon Lee, Roshan Lalintha Peiris, Li-Wei Chan 0001, Juhan Nam |
TEI | 5 |
| 2017 | Multi-Level and Multi-Scale Feature Aggregation Using Pretrained Convolutional Neural Networks for Music Auto-TaggingabstractMusic auto-tagging is often handled in a similar manner to image classification by regarding the two-dimensional audio spectrogram as image data. However, music auto-tagging is distinguished from image classification in that the tags are highly diverse and have different levels of abstraction. Considering this issue, we propose a convolutional neural networks (CNN)-based architecture that embraces multi-level and multi-scaled features. The architecture is trained in three steps. First, we conduct supervised feature learning to capture local audio features using a set of CNNs with different input sizes. Second, we extract audio features from each layer of the pretrained convolutional networks separately and aggregate them altogether giving a long audio clip. Finally, we put them into fully connected networks and make final predictions of the tags. Our experiments show that using the combination of multi-level and multi-scale features is highly effective in music auto-tagging and the proposed method outperforms the previous state-of-the-art methods on the MagnaTagATune dataset and the Million Song Dataset. We further show that the proposed architecture is useful in transfer learning. Jongpil Lee, Juhan Nam |
IEEE Signal Process. Lett. | 2 |
| 2012 | Optimized Polynomial Spline Basis Function Design for Quasi-Bandlimited Classical Waveform SynthesisabstractClassical geometric waveforms used in virtual analog synthesis suffer from aliasing distortion when simple sampling is used. An efficient antialiasing technique is based on expressing the waveforms as a filtered sum of time-shifted approximately bandlimited polynomial-spline basis functions. It is shown that by optimizing the coefficients of the basis function so that the aliasing distortion is perceptually minimized, the alias-free bandwidth of classical waveforms can be expanded. With the best of the case examples given here, the generated impulse-train and sawtooth waveform are alias-free up to fundamental frequencies over 10 kHz when the sampling rate is 44.1 kHz. Jussi Pekonen, Juhan Nam, Julius O. Smith III, Vesa Välimäki |
IEEE Signal Process. Lett. | 2 |
| 2011 | Multimodal Deep Learning
Jiquan Ngiam, Aditya Khosla, Juhan Nam, Honglak Lee, Andrew Y. Ng |
ICML | 4 |
| 2010 | A super-resolution spectrogram using coupled PLCAabstractThe short-time Fourier transform (STFT) based spectrogram is commonly used to analyze the time-frequency content of a signal. Depending on window size, the STFT provides a trade-off between time and frequency resolutions. This paper presents a novel method that achieves high resolution simultaneously in both time and frequency. We extend Probabilistic Latent Component Analysis (PLCA) to jointly decompose two spectrograms, one with a high time resolution and one with a high frequency resolution. Using this decomposition, a new spectrogram, maintaining high resolution in both time and frequency, is constructed. Termed the “super-resolution spectrogram”, it can be particularly useful for speech as it can simultaneously resolve both glottal pulses and individual harmonics. Juhan Nam, Gautham J. Mysore, Joachim Ganseman, Kyogu Lee, Jonathan S. Abel |
INTERSPEECH | 1 |
| 2010 | Efficient Antialiasing Oscillator Algorithms Using Low-Order Fractional Delay FiltersabstractOne of the challenges in virtual analog synthesis is avoiding aliasing when generating classic waveforms such as sawtooth and square wave which have theoretically infinite bandwidth in their ideal forms. The human auditory system renders a certain amount of aliasing inaudible, which allows room for finding cost-effective algorithms. This paper suggests efficient algorithms to reduce the aliasing using low-order fractional delay filters in the framework of bandlimited impulse train (BLIT) synthesis. Examining Lagrange, B-spline interpolators and allpass fractional delay filters, optimized methods will be discussed for generating classic waveforms (sawtooth, square, and triangle). Techniques for generating more complicated harmonics such as pulse width modulation, hard-sync, and super-saw are also presented. The perceptual evaluation is performed by comparing the threshold of hearing and masking curve of oscillators with their aliasing levels. The result shows that the BLIT using the computationally efficient third-order B-spline generates waveforms that are perceptually free of aliasing within practically used fundamental frequencies. Juhan Nam, Vesa Välimäki, Jonathan S. Abel, Julius O. Smith III |
IEEE Trans. Speech Audio Process. | 1 |
| 2010 | Alias-Suppressed Oscillators Based on Differentiated Polynomial WaveformsabstractAn efficient approach to the generation of classical synthesizer waveforms with reduced aliasing is proposed. This paper introduces two new classes of polynomial waveforms that can be differentiated one or more times to obtain an improved version of the sampled sawtooth and triangular signals. The differentiated polynomial waveforms (DPW) extend the previous differentiated parabolic wave method to higher polynomial orders, providing improved alias-suppression. Suitable polynomials of order higher than two can be derived either by analytically integrating a previous lower order polynomial or by solving the polynomial coefficients directly from a set of equations based on constraints. We also show how rectangular waveforms can be easily produced by differentiating a triangular signal. Bandlimited impulse trains can be obtained by differentiating the sawtooth or the rectangular signal. An objective evaluation using masking and hearing threshold models shows that a fourth-order DPW method is perceptually alias-free over the whole register of the grand piano. The proposed methods are applicable in digital implementations of subtractive sound synthesis. Vesa Välimäki, Juhan Nam, Julius O. Smith III, Jonathan S. Abel |
IEEE Trans. Speech Audio Process. | 2 |