VLDB 2026 Research / reviewers in the wild / expert
Minje Kim 0001
dblp:36/3427-1
· DBLP profile ↗
26ranked-venue papers
8as first author
8since 2021 · last 2025
0000-0003-3513-8328ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 6 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Perceptual Audio Coding: A 40-Year Historical PerspectiveabstractIn the history of audio and acoustic signal processing, perceptual audio coding has certainly excelled as a bright success story by its ubiquitous deployment in virtually all digital media devices, such as computers, tablets, mobile phones, set-top-boxes, and digital radios. From a technology perspective, perceptual audio coding has undergone tremendous development from the first very basic perceptually driven coders (including the popular mp3 format) to today’s full-blown integrated coding/rendering systems. This paper provides a historical overview of this research journey by pinpointing the pivotal development steps in the evolution of perceptual audio coding. Finally, it provides thoughts about future directions in this area. Jürgen Herre, Schuyler R. Quackenbush, Minje Kim 0001, Jan Skoglund |
ICASSP | 3 |
| 2023 | The Potential of Neural Speech Synthesis-Based Data Augmentation for Personalized Speech EnhancementabstractWith the advances in deep learning, speech enhancement systems benefited from large neural network architectures and achieved state-of-the-art quality. However, speaker-agnostic methods are not always desirable, both in terms of quality and their complexity, when they are to be used in a resource-constrained environment. One promising way is personalized speech enhancement (PSE), which is a smaller and easier speech enhancement problem for small models to solve, because it focuses on a particular test-time user. To achieve the personalization goal, while dealing with the typical lack of personal data, we investigate the effect of data augmentation based on neural speech synthesis (NSS). In the proposed method, we show that the quality of the NSS system’s synthetic data matters, and if they are good enough the augmented dataset can be used to improve the PSE system that outperforms the speaker-agnostic baseline. The proposed PSE systems show significant complexity reduction while preserving the enhancement quality. Anastasia Kuznetsova, Aswin Sivaraman, Minje Kim 0001 |
ICASSP | 3 |
| 2023 | Native Multi-Band Audio Coding Within Hyper-Autoencoded Reconstruction Propagation NetworksabstractSpectral sub-bands do not portray the same perceptual relevance. In audio coding, it is therefore desirable to have independent control over each of the constituent bands so that bitrate assignment and signal reconstruction can be achieved efficiently. In this work, we present a novel neural audio coding network that natively supports a multi-band coding paradigm. Our model extends the idea of compressed skip connections in the U-Net-based codec, allowing for independent control over both core and high band-specific reconstructions and bit allocation. Our system reconstructs the full-band signal mainly from the condensed core-band code, therefore exploiting and showcasing its bandwidth extension capabilities to its fullest. Meanwhile, the low-bitrate high-band code helps the high-band reconstruction similarly to MPEG audio codecs' spectral bandwidth replication. MUSHRA tests show that the proposed model not only improves the quality of the core band by explicitly assigning more bits to it but retains a good quality in the high-band as well. Darius Petermann, Inseon Jang, Minje Kim 0001 |
ICASSP | 3 |
| 2023 | Neural Feature Predictor and Discriminative Residual Coding for Low-Bitrate Speech CodingabstractLow and ultra-low-bitrate neural speech codecs achieved unprecedented coding gain by generating speech signals from compact features. This paper introduces additional coding efficiency in speech coding by reducing the temporal redundancy existing in the frame-level feature sequence via a feature predictor. This predictor produces low-entropy residual representations, and we discriminatively code them based on their contribution to the signal reconstruction. Combining feature prediction and discriminative coding optimizes bitrate efficiency by assigning more bits to hard-to-predict events. We demonstrate the advantage of the proposed methods using the LPCNet as a neural vocoder, resulting in a scalable, lightweight, low-latency, and low-bitrate neural speech coding system. While our approach guarantees strict causality in the frame-level prediction, the subjective tests and feature space analysis show that our model achieves superior coding efficiency compared to the loosely-causal LPCNet and Lyra V2 in the very low bitrates. Haici Yang, Wootaek Lim, Minje Kim 0001 |
ICASSP | 3 |
| 2022 | Bloom-Net: Blockwise Optimization for Masking Networks Toward Scalable and Efficient Speech EnhancementabstractIn this paper, we present a blockwise optimization method for masking-based networks (BLOOM-Net) for training scalable speech enhancement networks. Here, we design our network with a residual learning scheme and train the internal separator blocks sequentially to obtain a scalable masking-based deep neural network for speech enhancement. Its scalability lets it dynamically adjust the run-time complexity depending on the test time environment. To this end, we modularize our models in that they can flexibly accommodate varying needs for enhancement performance and constraints on the resources, incurring minimal memory or training overhead due to the added scalability. Our experiments on speech enhancement demonstrate that the proposed blockwise optimization method achieves the desired scalability with only a slight performance degradation compared to corresponding models trained end-to-end. Sunwoo Kim 0003, Minje Kim 0001 |
ICASSP | 2 |
| 2022 | Boosted Locality Sensitive Hashing: Discriminative, Efficient, and Scalable Binary Codes for Source SeparationabstractWe propose a novel adaptive boosting approach to learn discriminative binary hash codes, boosted locality sensitive hashing (BLSH), that can represent audio spectra efficiently. We aim to use the learned hash codes in the single-channel speech denoising task by designing a nearest neighborhood search method that operates in the hashed feature space. To achieve the optimal denoising results given the highly compact binary feature representation, our proposed BLSH algorithm learns simple logistic regressors as the weak learners in an incremental way (i.e., one by one) so that each weak learner is trained to complement the mistake its predecessors have made. Upon testing, their binary classification results transform each spectrum of noisy speech into a bit string, where the bits are ordered based on their significance, adding scalability to the denoising system. Simple bitwise operations calculate Hamming distance to find theK-nearest matching hashed frames in the dictionary of training noisy speech spectra, whose associated ideal binary masks are averaged to estimate the denoising mask for that test mixture. In contrast to the locality sensitive hashing method's random projections, our proposed supervised learning algorithm trains the projections such that the distance between the self-similarity matrix of the hash codes and that of the original spectra is minimized. Likewise, the process conceptually aligns to the Adaboost algorithm, although ours is specialized in learning binary features for source separation rather than classification. Experimental results on speech denoising suggest that the BLSH algorithm learns more discriminative representations than Fourier or mel spectra and the nonlinear kernels derived from them. Our compact binary representation is expected to facilitate model deployment onto resource-constrained environments, where comprehensive models (e.g., deep neural networks) are unaffordable. Sunwoo Kim 0003, Minje Kim 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Scalable and Efficient Neural Speech Coding: A Hybrid DesignabstractWe present a scalable and efficient neural waveform coding system for speech compression. We formulate the speech coding problem as an autoencoding task, where a convolutional neural network (CNN) performs encoding and decoding as a neural waveform codec (NWC) during its feedforward routine. The proposed NWC also defines quantization and entropy coding as a trainable module, so the coding artifacts and bitrate control are handled during the optimization process. We achieve efficiency by introducing compact model components to NWC, such as gated residual networks and depthwise separable convolution. Furthermore, the proposed models are with a scalable architecture, cross-module residual learning (CMRL), to cover a wide range of bitrates. To this end, we employ the residual coding concept to concatenate multiple NWC autoencoding modules, where each NWC module performs residual coding to restore any reconstruction loss that its preceding modules have created. CMRL can scale down to cover lower bitrates as well, for which it employs linear predictive coding (LPC) module as its first autoencoder. The hybrid design integrates LPC and NWC by redefining LPC’s quantization as a differentiable process, making the system training an end-to-end manner. The decoder of proposed system is with either one NWC (0.12 million parameters) in low to medium bitrate ranges (12 to 20 kbps) or two NWCs in the high bitrate (32 kbps). Although the decoding complexity is not yet as low as that of conventional speech codecs, it is significantly reduced from that of other neural speech coders, such as a WaveNet-based vocoder. For wide-band speech coding quality, our system yields comparable or superior performance to AMR-WB and Opus on TIMIT test utterances at low and medium bitrates. The proposed system can scale up to higher bitrates to achieve near transparent performance. Kai Zhen, Jongmo Sung, Mi Suk Lee, Seungkwon Beack, Minje Kim 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2021 | Personalized Speech Enhancement Through Self-Supervised Data Augmentation and PurificationabstractTraining personalized speech enhancement models is innately a no-shot learning problem due to privacy constraints and limited access to noise-free speech from the target user. If there is an abundance of unlabeled noisy speech from the test-time user, a personalized speech enhancement model can be trained using self-supervised learning. One straightforward approach to model personalization is to use the target speaker's noisy recordings as pseudo-sources. Then, a pseudo denoising model learns to remove injected training noises and recover the pseudo-sources. However, this approach is volatile as it depends on the quality of the pseudo-sources, which may be too noisy. As a remedy, we propose an improvement to the self-supervised approach through data purification. We first train an SNR predictor model to estimate the frame-by-frame SNR of the pseudo-sources. Then, the predictor's estimates are converted into weights which adjust the frame-by-frame contribution of the pseudo-sources towards training the personalized model. We empirically show that the proposed data purification step improves the usability of the speaker-specific noisy data in the context of personalized speech enhancement. Without relying on any clean speech recordings or speaker embeddings, our approach may be seen as privacy-preserving. Aswin Sivaraman, Sunwoo Kim 0003, Minje Kim 0001 |
Interspeech | 3 |
| 2020 | Boosted Locality Sensitive Hashing: Discriminative Binary Codes for Source SeparationabstractSpeech enhancement tasks have seen significant improvements with the advance of deep learning technology, but with the cost of increased computational complexity. In this study, we propose an adaptive boosting approach to learning locality sensitive hash codes, which represent audio spectra efficiently. We use the learned hash codes for single-channel speech denoising tasks as an alternative to a complex machine learning model, particularly to address the resource-constrained environments. Our adaptive boosting algorithm learns simple logistic regressors as the weak learners. Once trained, their binary classification results transform each spectrum of test noisy speech into a bit string. Simple bitwise operations calculate Hamming distance to find the K-nearest matching frames in the dictionary of training noisy speech spectra, whose associated ideal binary masks are averaged to estimate the denoising mask for that test mixture. Our proposed learning algorithm differs from AdaBoost in the sense that the projections are trained to minimize the distances between the self-similarity matrix of the hash codes and that of the original spectra, rather than the misclassification rate. We evaluate our discriminative hash codes on the TIMIT corpus with various noise types, and show comparative performance to deep learning methods in terms of denoising performance and complexity. Sunwoo Kim 0003, Haici Yang, Minje Kim 0001 |
ICASSP | 3 |
| 2020 | Psychoacoustic Calibration of Loss Functions for Efficient End-to-End Neural Audio CodingabstractConventional audio coding technologies commonly leverage human perception of sound, or psychoacoustics, to reduce the bitrate while preserving the perceptual quality of the decoded audio signals. For neural audio codecs, however, the objective nature of the loss function usually leads to suboptimal sound quality as well as high run-time complexity due to the large model size. In this work, we present a psychoacoustic calibration scheme to re-define the loss functions of neural audio coding systems so that it can decode signals more perceptually similar to the reference, yet with a much lower model complexity. The proposed loss function incorporates the global masking threshold, allowing the reconstruction error that corresponds to inaudible artifacts. Experimental results show that the proposed model outperforms the baseline neural codec twice as large and consuming 23.4% more bits per second. With the proposed method, a lightweight neural codec, with only 0.9 million parameters, performs near-transparent audio coding comparable with the commercial MPEG-1 Audio Layer III codec at 112 kbps. Kai Zhen, Mi Suk Lee, Jongmo Sung, Seungkwon Beack, Minje Kim 0001 |
IEEE Signal Process. Lett. | 5 |
| 2019 | Incremental Binarization on Recurrent Neural Networks for Single-channel Source SeparationabstractThis paper proposes a Bitwise Gated Recurrent Unit (BGRU) network for the single-channel source separation task. Recurrent Neural Networks (RNN) require several sets of weights within its cells, which significantly increases the computational cost compared to the fully-connected networks. To mitigate this increased computation, we focus on the GRU cells and quantize the feedforward procedure with binarized values and bitwise operations. The BGRU network is trained in two stages. The real-valued weights are pretrained and transferred to the bitwise network, which are then incrementally binarized to minimize the potential loss that can occur from a sudden introduction of quantization. As the proposed binarization technique turns only a few randomly chosen parameters into their binary versions, it gives the network training procedure a chance to gently adapt to the partly quantized version of the network. It eventually achieves the full binarization by incrementally increasing the amount of binarization over the iterations. Our experiments show that the proposed BGRU method produces source separation results greater than that of a real-valued fully connected network, with 11-12 dB mean Signal-to-Distortion Ratio (SDR). A fully binarized BGRU still outperforms a Bitwise Neural Network (BNN) by 1-2 dB even with less number of layers. Sunwoo Kim 0003, Mrinmoy Maity, Minje Kim 0001 |
ICASSP | 3 |
| 2018 | Bitwise Neural Networks for Efficient Single-Channel Source SeparationabstractWe present Bitwise Neural Networks (BNN) as an efficient hardware-friendly solution to single-channel source separation tasks in resource-constrained environments. In the proposed BNN system, we replace all the real-valued operations during the feedforward process of a Deep Neural Network (DNN) with bitwise arithmetic (e.g. the XNOR operation between bipolar binaries in place of multiplications). Thanks to the fully bitwise run-time operations, the BNN system can serve as an alternative solution where efficient real-time processing is critical, for example real-time speech enhancement in embedded systems. Furthermore, we also propose a binarization scheme to convert the input signals into bit strings so that the BNN parameters learn the Boolean mapping between input binarized mixture signals and their target Ideal Binary Masks (IBM). Experiments on the single-channel speech denoising tasks show that the efficient BNN-based source separation system works well with an acceptable performance loss compared to a comprehensive real-valued network, while consuming a minimal amount of resources. Minje Kim 0001, Paris Smaragdis |
ICASSP | 1 |
| 2017 | On Exploiting Structured Human Interactions to Enhance Sensing Accuracy in Cyber-physical SystemsabstractIn this article, we describe a general methodology for enhancing sensing accuracy in cyber-physical systems that involve structured human interactions in noisy physical environment. We define structured human interactions as domain-specific workflow. A novel workflow-aware sensing model is proposed to jointly correct unreliable sensor data and keep track of states in a workflow. We also propose a new inference algorithm to handle cases with partially known states and objects as supervision. Our model is evaluated with extensive simulations. As a concrete application, we develop a novel log service called Emergency Transcriber , which can automatically document operational procedures followed by teams of first responders in emergency response scenarios. Evaluation shows that our system has significant improvement over commercial off-the-shelf (COTS) sensors and keeps track of workflow states with high accuracy in noisy physical environment. Shaohan Hu, Shiguang Wang, Renato Mancuso 0001, Minje Kim 0001, Po-Liang Wu, Lu Su 0001, Lui Sha, Tarek F. Abdelzaher |
ACM Trans. Cyber Phys. Syst. | 6 |
| 2016 | Efficient neighborhood-based topic modeling for collaborative audio enhancement on massive crowdsourced recordingsabstractCollaborative Audio Enhancement (CAE) aims at separating a dominant source from crowdsourced recordings of a scene. This paper proposes a CAE setup as a big ad-hoc microphone array problem, assuming hundreds of sensors scattered over a large scene, e.g. a concert hall or a street riot. An important characteristic in such cases is the fact that not all sensors capture useful information, mainly because of the existence of strong local noise interferences and recording artifacts. This renders traditional array processing techniques inadequate for tasks such as source enhancement. One way to recover the most common source while suppressing recording-specific interference, is to share latent components across simultaneous models on multiple magnitude spectrograms. The proposed method improves on the quality and the computational requirements of such a model by using a two-stage nearest-neighborhood search at every EM update. Its optional first-round search uses Hamming distance between hashed spectrograms to quickly find a redundant candidate set, and then a subsequent step narrows the set down to a subset using more appropriate cross entropy. Experimental results show that the proposed neighborhood schemes converge to the better quality solutions faster than the comprehensive model using all data. Minje Kim 0001, Paris Smaragdis |
ICASSP | 1 |
| 2015 | Efficient manifold preserving audio source separation using locality sensitive hashingabstractWe propose an efficient technique to learn probabilistic hierarchical topic models that are designed to preserve the manifold structure of audio data. The consideration of the data manifold is important, as it has been shown to provide superior performance in certain audio applications such as source separation. However, the high computational cost of a sparse encoding step due to the requirement of a large dictionary prevents it from being used in real-world applications such as real-time speech enhancement and the analysis of big audio data. In order to achieve a substantial speed-up of this step, while still respecting the data manifold, we propose to harmonize a particular type of locality sensitive hashing with the hierarchical topic model. The proposed use of hashing can reduce the computational complexity of the sparse encoding by providing candidates of non-zero activations, where the candidate set is built based on Hamming distance. The hashing step is followed by comprehensive sparse coding that considers those candidates only, rather than the entire dictionary. Experimental results show that the proposed hashing technique can provide audio source separation results comparable to the similar system without hashing, but with significantly less and cheaper computation. Minje Kim 0001, Paris Smaragdis, Gautham J. Mysore |
ICASSP | 1 |
| 2015 | Mixtures of Local Dictionaries for Unsupervised Speech EnhancementabstractWe propose a novel extension of Nonnegative Matrix Factorization (NMF) that models a signal with multiple local dictionaries activated sparsely. This set of local dictionaries for a source, e.g., speech, disjointly constitute a superset that is more discriminative than an ordinary NMF dictionary, because its local structures represent the source's manifold better. A block sparsity constraint is used to regularize the NMF solutions so that only one or a small number of blocks are active at a given time. Moreover, a concentrationz prior further regularizes each block of bases to be close to each other for better locality preservation. We test the proposed Mixture of Local Dictionaries (MLD) on single-channel speech enhancement tasks and show that it outperforms the state of the art technology by up to 2 dB in signal-to-distortion ratio, especially in the unsupervised environment where neither the speaker identity nor the type of noise is known in advance. Minje Kim 0001, Paris Smaragdis |
IEEE Signal Process. Lett. | 1 |
| 2015 | Joint Optimization of Masks and Deep Recurrent Neural Networks for Monaural Source SeparationabstractMonaural source separation is important for many real world applications. It is challenging because, with only a single channel of information available, without any constraints, an infinite number of solutions are possible. In this paper, we explore joint optimization of masking functions and deep recurrent neural networks for monaural source separation tasks, including speech separation, singing voice separation, and speech denoising. The joint optimization of the deep recurrent neural networks with an extra masking layer enforces a reconstruction constraint. Moreover, we explore a discriminative criterion for training neural networks to further enhance the separation performance. We evaluate the proposed system on the TSP, MIR-1K, and TIMIT datasets for speech separation, singing voice separation, and speech denoising tasks, respectively. Our approaches achieve 2.30-4.98 dB SDR gain compared to NMF models in the speech separation task, 2.30-2.48 dB GNSDR gain and 4.32-5.42 dB GSIR gain compared to existing models in the singing voice separation task, and outperform NMF and DNN baselines in the speech denoising task. Po-Sen Huang, Minje Kim 0001, Mark Hasegawa-Johnson, Paris Smaragdis |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Deep learning for monaural speech separationabstractMonaural source separation is useful for many real-world applications though it is a challenging problem. In this paper, we study deep learning for monaural speech separation. We propose the joint optimization of the deep learning models (deep neural networks and recurrent neural networks) with an extra masking layer, which enforces a reconstruction constraint. Moreover, we explore a discriminative training criterion for the neural networks to further enhance the separation performance. We evaluate our approaches using the TIMIT speech corpus for a monaural speech separation task. Our proposed models achieve about 3.8∼4.9 dB SIR gain compared to NMF models, while maintaining better SDRs and SARs. Po-Sen Huang, Minje Kim 0001, Mark Hasegawa-Johnson, Paris Smaragdis |
ICASSP | 2 |
| 2014 | Phase and level difference fusion for robust multichannel source separationabstractInter-channel phase (IPD) and level (ILD) differences are common features in multichannel source separation algorithms like DUET and MENUET. However, their utility depends strongly on the configuration of the array and what microphone pairs are used to calculate them. IPDs are most useful when extracted from microphones that are close together as this avoids spatial aliasing. In contrast, ILD clusters are only well separated for widely spaced microphones. We investigate this trade-off between IPD and ILD features and propose a method to best combine them for multichannel source separation. Experimental results demonstrate the utility of this approach. Johannes Traa, Minje Kim 0001, Paris Smaragdis |
ICASSP | 2 |
| 2014 | Experiments on deep learning for speech denoisingabstractIn this paper we present some experiments using a deep learn-ing model for speech denoising. We propose a very lightweight procedure that can predict clean speech spectra when presented with noisy speech inputs, and we show how various parameter choices impact the quality of the denoised signal. Through our experiments we conclude that such a structure can perform bet-ter than some comparable single-channel approaches and that it is able to generalize well across various speakers, noise types and signal-to-noise ratios. Paris Smaragdis, Minje Kim 0001 |
INTERSPEECH | 3 |
| 2013 | Collaborative audio enhancement using probabilistic latent component sharingabstractThis paper presents a collaborative audio enhancement system that aims to recover common audio sources from multiple recordings of a given audio scene. We do so in the context where each recording is uniquely corrupted. To this end, we propose a method of simultaneous probabilistic latent component analyses on synchronized inputs. In the proposed model, some of the parameters are fixed to be same during and after the learning process to capture common audio content while the rest models unwanted recording-specific interferences and artifacts. Our model also allows for prior knowledge about the parameters of the model, e.g. representative spectra of the components, to be incorporated in the factorization. A post processing scheme that consolidates the extracted sources from the set of inputs is also proposed to handle the possible loss of certain frequency regions. Experiments on commercial music signals with various artifacts show the merit of the proposed method. Minje Kim 0001, Paris Smaragdis |
ICASSP | 1 |
| 2013 | Manifold Preserving Hierarchical Topic Models for Quantization and ApproximationabstractWe present two complementary topic models to address the analysis of mixture data lying on manifolds. First, we propose a quantization method with an additional mid-layer latent variable, which selects only data points that best preserve the manifold structure of the input data. In order to address the case of modeling all the in-between parts of that manifold using this reduced representation of the input, we introduce a new model that provides a manifold-aware interpolation method. We demonstrate the advantages of these models with experiments on the hand-written digit recognition and the speech source separation tasks. Minje Kim 0001, Paris Smaragdis |
ICML (3) | 1 |
| 2013 | EMERALD: Characterization of emerging applications and algorithms for low-power devicesabstractCompute-intensive applications are emerging in intelligent home, retail store and automotive industries. These applications are becoming more sophisticated with new features rich in audio, video, image, and machine learning capabilities that demand heavy computations. We present the EMERALD (EMERging Applications and algorithms for Low power Device) workload suite. We profile the workloads to show the hotspot functions that are candidates for hardware accelerators. Chuanjun Zhang, Glenn G. Ko, Jungwook Choi, Shang-nien Tsai, Minje Kim 0001, Abner Guzmán-Rivera, Rob A. Rutenbar, Paris Smaragdis, Mi Sun Park, Narayanan Vijaykrishnan, Hongyi Xin, Onur Mutlu, Bin Li 0018, Li Zhao 0002 |
ISPASS | 5 |
| 2010 | Blind rhythmic source separation: Nonnegativity and repeatabilityabstractAn unsupervised method is proposed aiming at extracting rhythmic sources from commercial polyphonic music whose number of channels is limited to one. Commercial music signals are not usually provided with more than two channels while they often contain multiple instruments including singing voice. Therefore, instead of using conventional ways, such as modeling mixing environments or statistical characteristics, we should introduce other source-specific characteristics for separating or extracting the sources. In this paper, we concentrate on extracting rhythmic sources from the mixture with the other harmonic sources. An extension of nonnegative matrix factorization (NMF) is used to analyze multiple relationships between spectral and temporal properties in the given input matrices. Moreover, temporal repeatability of the rhythmic sound sources is implicated as common rhythmic property among segments of an input mixture signal. The proposed method shows acceptable, but not superior separation quality to the referred drum source separation systems. However, it has better applicability due to its blind manner in separation. Minje Kim 0001, Jiho Yoo, Kyeongok Kang, Seungjin Choi 0001 |
ICASSP | 1 |
| 2010 | Nonnegative matrix partial co-factorization for drum source separationabstractWe address a problem of separating drums from polyphonic music containing various pitched instruments as well as drums. Nonnegative matrix factorization (NMF) was successfully applied to spectrograms of music to learn basis vectors, followed by support vector machine (SVM) to classify basis vectors into ones associated with drums (rhythmic source) only and pitched instruments (harmonic sources). Basis vectors associated with pitched instruments are used to reconstruct drum-eliminated music. However, it is cumbersome to construct a training set for pitched instruments since various instruments are involved. In this paper, we propose a method which only incorporates prior knowledge on drums, not requiring such training sets of pitched instruments. To this end, we present nonnegative matrix partial co-factorization (NMPCF) where the target matrix (spectrograms of music) and drum-only-matrix (collected from various drums a priori) are simultaneously decomposed, sharing some factor matrix partially, to force some portion of basis vectors to be associated with drums only. We develop a simple multiplicative algorithm for NMPCF and show its usefulness empirically, with numerical experiments on real-world music signals. Jiho Yoo, Minje Kim 0001, Kyeongok Kang, Seungjin Choi 0001 |
ICASSP | 2 |
| 2005 | On Spectral Basis Selection for Single Channel Polyphonic Music Separation
Minje Kim 0001, Seungjin Choi 0001 |
ICANN (2) | 1 |