EDBT 2026 Demo / reviewers in the wild / expert
Kazuyoshi Yoshii
dblp:84/4993
· DBLP profile ↗
78ranked-venue papers
12as first author
18since 2021 · last 2025
0000-0001-8387-8609ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 43 · 7 first-author · 12 since 2021Artificial intelligence and machine learning · 37 · 5 first-author · 6 since 2021Systems, architecture and hardware · 9 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Multifaceted Multi-Agent Framework for Zero-Shot Emotion Analysis and Recognition of Symbolic MusicabstractICMI '25: International Conference on Multimodal Interaction; Canberra, Australia; October 13 - 17, 2025 Yunjia Li, Kazuyoshi Yoshii |
ICMI | 3 |
| 2024 | RIR-in-a-Box: Estimating Room Acoustics from 3D Mesh Data through Shoebox ApproximationabstractThis paper describes a method for estimating the room impulse response (RIR) for a microphone and a sound source located at arbitrary positions from the 3D mesh data of the room.Simulating realistic RIRs with pure physics-driven methods often fails the balance between physical consistency and computational efficiency, hindering application to real-time speech processing.Alternatively, one can use MESH2IR, a fast black-box estimator that consists of an encoder extracting latent code from mesh data with a graph convolutional network (GCN) and a decoder generating the RIR from the latent code.Combining these two approaches, we propose a fast yet physically coherent estimator with interpretable latent code based on differentiable digital signal processing (DDSP).Specifically, the encoder estimates a virtual shoebox room scene that acoustically approximates the real scene, accelerating physical simulation with the differentiable image-source model in the decoder.Our experiments showed that our method outperformed MESH2IR for real mesh data obtained with the depth scanner of Microsoft HoloLens 2, and can provide correct spatial consistency for binaural RIRs. Liam Kelley, Diego Di Carlo, Aditya Arie Nugraha, Mathieu Fontaine 0002, Yoshiaki Bando, Kazuyoshi Yoshii |
INTERSPEECH | 6 |
| 2023 | Neural Band-to-Piano Score Arrangement with Stepless Difficulty ControlabstractThis paper describes a music arrangement method of popular music that can convert a band score into a piano score with a steplessly-specified level of performance difficulty. The basic strategy of band-to-piano score arrangement is to select notes from an augmented band score obtained by up- and down-shifting the notes of an original band score by one octave. Given band scores and the corresponding piano scores with elementary- and advanced-levels, one can train a deep neural network (DNN) that estimates note masks conditioned by the difficulty levels. Conditioned by an intermediate level at runtime, however, the DNN tends to generate an advanced-level score. To solve this problem, assuming that an easier piano score is a subset of harder one, we estimate the basic importance of each note with a difficulty-agnostic DNN and then warp it with a power function depending on a specified difficulty level. To achieve the fine controllability of the difficulty level, we propose a training method that subjects the DNN to generating piano scores with various intermediate levels, where the note-level loss for those scores is evaluated using only the ground-truth elementary- and advanced-level scores. Considering the non-uniqueness of piano arrangement, the statistic-level loss with respect to the note density and polyphony level is also computed according to the given levels. The experimental results showed that the proposed method attained both the performance gain and the stepless difficulty control. Moyu Terao, Eita Nakamura, Kazuyoshi Yoshii |
ICASSP | 3 |
| 2022 | Flow-Based Fast Multichannel Nonnegative Matrix Factorization for Blind Source SeparationabstractThis paper describes a blind source separation method for multichannel audio signals, called NF-FastMNMF, based on the integration of the normalizing flow (NF) into the multichannel nonnegative matrix factorization with jointly-diagonalizable spatial covariance matrices, a.k.a. FastMNMF. Whereas the NF of flow-based independent vector analysis, called NF-IVA, acts as the demixing matrices to transform an M-channel mixture into M independent sources, the NF of NF- FastMNMF acts as the diagonalization matrices to transform an M- channel mixture into a spatially-independent M-channel mixture represented as a weighted sum of N source images. This diagonalization enables the NF, which has been used only for determined separation because of its bijective nature, to be applicable to non-determined separation. NF-FastMNMF has time-varying diagonalization matrices that are potentially better at handling dynamical data variation than the time-invariant ones in FastMNMF. To have an NF with richer expression capability, the dimension-wise scalings using diagonal matrices originally used in NF-IVA are replaced with linear transformations using upper triangular matrices; in both cases, the diagonal and upper triangular matrices are estimated by neural networks. The evaluation shows that NF-FastMNMF performs well for both determined and non-determined separations of multiple speech utterances by stationary or non-stationary speakers from a noisy reverberant mixture. Aditya Arie Nugraha, Kouhei Sekiguchi, Mathieu Fontaine 0002, Yoshiaki Bando, Kazuyoshi Yoshii |
ICASSP | 5 |
| 2022 | Difficulty-Aware Neural Band-to-Piano Score Arrangement based on Note- and Statistic-Level CriteriaabstractThis paper describes a neural music arrangement method that converts a given band score into a piano score with an elementary or advanced level. The major challenge of this task lies in its ill-posed nature, i.e., various piano arrangements are plausible for a band score. In this paper, we take a score reduction approach based on supervised training of a mask estimation network (U-Net) with note- and statistic-level criteria. Based on statistical analysis of existing piano arrangements, a reasonable piano score is assumed to be obtained by reducing an augmented band score obtained by up- and downshifting an original band score by one octave. This effectively narrows down a solution space. At the heart of our approach is to train a U-Net conditioned by a given difficulty level such that a piano score obtained by masking an augmented band score is close to the ground-truth piano score not only at a note level but also at a statistic level. We focus on three kinds of note statistics, i.e., a distribution of the numbers of concurrent notes, that of the intervals between the highest and lowest pitches, and that of the per-measure numbers of notes. The experimental results show the importance of both the instance- and meta-level criteria for supervised training. Moyu Terao, Yuki Hiramatsu, Ryoto Ishizuka, Yiming Wu 0003, Kazuyoshi Yoshii |
ICASSP | 5 |
| 2022 | Direction-Aware Joint Adaptation of Neural Speech Enhancement and Recognition in Real Multiparty Conversational EnvironmentsabstractThis paper describes noisy speech recognition for an augmented reality headset that helps verbal communication with in real multiparty conversational environments.A major approach that has actively been studied in simulated environments is to sequentially perform speech enhancement and automatic speech recognition (ASR) based on deep neural networks (DNNs) trained in a supervised manner.In our task, however, such a pretrained system fails to work due to the mismatch between the training and test conditions and the head movements of the user.To enhance only the utterances of a target speaker, we use beamforming based on a DNN-based speech mask estimator that can adaptively extract the speech components corresponding to a head-relative particular direction.We propose a semi-supervised adaptation method that jointly updates the mask estimator and the ASR model at run-time using clean speech signals with ground-truth transcriptions and noisy speech signals with highly-confident estimated transcriptions.Comparative experiments using the state-of-theart distant speech recognition system show that the proposed method significantly improves the ASR performance. Yicheng Du, Aditya Arie Nugraha, Kouhei Sekiguchi, Yoshiaki Bando, Mathieu Fontaine 0002, Kazuyoshi Yoshii |
INTERSPEECH | 6 |
| 2022 | Direction-Aware Adaptive Online Neural Speech Enhancement with an Augmented Reality Headset in Real Noisy Conversational EnvironmentsabstractThis paper describes the practical response- and performance-aware development of online speech enhancement for an augmented reality (AR) headset that helps a user understand conversations made in real noisy echoic environments (e.g., cocktail party). One may use a state-of-the-art blind source separation method called fast multichannel nonnegative matrix factorization (FastMNMF) that works well in various environments thanks to its unsupervised nature. Its heavy computational cost, however, prevents its application to real-time processing. In contrast, a supervised beamforming method that uses a deep neural network (DNN) for estimating spatial information of speech and noise readily fits real-time processing, but suffers from drastic performance degradation in mismatched conditions. Given such complementary characteristics, we propose a dual-process robust online speech enhancement method based on DNN-based beamforming with FastMNMF-guided adaptation. FastMNMF (back end) is performed in a mini-batch style and the noisy and enhanced speech pairs are used together with the original parallel training data for updating the direction-aware DNN (front end) with backpropagation at a computationally-allowable interval. This method is used with a blind dereverberation method called weighted prediction error (WPE) for transcribing the noisy reverberant speech of a speaker, which can be detected from video or selected by a user's hand gesture or eye gaze, in a streaming manner and spatially showing the transcriptions with an AR technique. Our experiment showed that the word error rate was improved by more than 10 points with the run-time adaptation using only twelve minutes observation. Kouhei Sekiguchi, Aditya Arie Nugraha, Yicheng Du, Yoshiaki Bando, Mathieu Fontaine 0002, Kazuyoshi Yoshii |
IROS | 6 |
| 2022 | Computationally-Efficient Overdetermined Blind Source Separation Based on Iterative Source SteeringabstractThis paper describesa computationally-efficient optimization algorithm for the blind source separation (BSS) of overdetermined mixtures. In the determined case, a matrix-inversion-free iterative source steering (ISS) algorithm has been proposed for estimating a square demixing matrix as a computationally-efficient alternative to the popular iterative projection (IP) algorithm. The IP algorithm is based on source-wise (i.e., row-wise) updates of the demixing matrix, and lends itself naturally to an extension to overdetermined independent vector analysis (IVA) called OverIVA. In contrast, the ISS algorithm changes the whole demixing matrix at every update, making its extension to the overdetermined case non-trivial. In this paper, we propose a modified ISS algorithm for OverIVA fully exploiting the computational savings of ISS. We also derive an overdetermined extension of independent low-rank matrix analysis (OverILRMA) with the modified ISS algorithm. Experimental results showed that the proposed ISS-based OverIVA and OverILRMA were comparable or superior to the conventional IP-based counterparts in speech separation performance while achieving lower computational cost. Yicheng Du, Robin Scheibler, Masahito Togami, Kazuyoshi Yoshii, Tatsuya Kawahara |
IEEE Signal Process. Lett. | 4 |
| 2022 | Generalized Fast Multichannel Nonnegative Matrix Factorization Based on Gaussian Scale Mixtures for Blind Source SeparationabstractThis paper describes heavy-tailed extensions of a state-of-the-art versatile blind source separation method called fast multichannel nonnegative matrix factorization (FastMNMF) from a unified point of view. The common way of deriving such an extension is to replace the multivariate complex Gaussian distribution in the likelihood function with its heavy-tailed generalization,e.g., the multivariate complex Student’s$t$and leptokurtic generalized Gaussian distributions, and tailor-make the corresponding parameter optimization algorithm. Using a wider class of heavy-tailed distributions called a Gaussian scale mixture (GSM),i.e., a mixture of Gaussian distributions whose variances are perturbed by positive random scalars called impulse variables, we propose GSM-FastMNMF and develop an expectation-maximization algorithm that works even when the probability density function of the impulse variables have no analytical expressions. We show that existing heavy-tailed FastMNMF extensions are instances of GSM-FastMNMF and derive a new instance based on the generalized hyperbolic distribution that include the normal-inverse Gaussian, Student’s$t$, and Gaussian distributions as the special cases. Our experiments show that the normal-inverse Gaussian FastMNMF outperforms the state-of-the-art FastMNMF extensions and ILRMA model in speech enhancement and separation in terms of the signal-to-distortion ratio. Mathieu Fontaine 0002, Kouhei Sekiguchi, Aditya Arie Nugraha, Yoshiaki Bando, Kazuyoshi Yoshii |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2022 | Autoregressive Moving Average Jointly-Diagonalizable Spatial Covariance Analysis for Joint Source Separation and DereverberationabstractThis paper describes a computationally-efficient statistical approach to joint (semi-)blind source separation and dereverberation for multichannel noisy reverberant mixture signals. A standard approach to source separation is to formulate a generative model of a multichannel mixture spectrogram that consists of source and spatial models representing the time-frequency power spectral densities (PSDs) and spatial covariance matrices (SCMs) of source images, respectively, and find the maximum-likelihood estimates of these parameters. A state-of-the-art blind source separation method in this thread of research is fast multichannel nonnegative matrix factorization (FastMNMF) based on the low-rank PSDs and jointly-diagonalizable full-rank SCMs. To perform mutually-dependent separation and dereverberation jointly, in this paper we integrate both moving average (MA) and autoregressive (AR) models that represent the early reflections and late reverberations of sources, respectively, into the FastMNMF formalism. Using a pretrained deep generative model of speech PSDs as a source model, we realize semi-blind joint speech separation and dereverberation. We derive an iterative optimization algorithm based on iterative projection or iterative source steering for jointly and efficiently updating the AR parameters and the SCMs. Our experimental results showed the superiority of the proposed ARMA extension over its AR- or MA-ablated version in a speech separation and/or dereverberation task. Kouhei Sekiguchi, Yoshiaki Bando, Aditya Arie Nugraha, Mathieu Fontaine 0002, Kazuyoshi Yoshii, Tatsuya Kawahara |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2021 | Statistical Correction of Transcribed Melody Notes Based on Probabilistic Integration of a Music Language Model and a Transcription Error ModelabstractThis paper describes a statistical post-processing method for automatic singing transcription that corrects pitch and rhythm errors included in a transcribed note sequence. Although the performance of frame-level pitch estimation has been improved drastically by deep learning techniques, note-level transcription of singing voice is still an open problem. Inspired by the standard framework of statistical machine translation, we formulate a hierarchical generative model of a transcribed note sequence that consists of a music language model describing the pitch and onset transitions of a true note sequence and a transcription error model describing the addition of deletion, insertion, and substitution errors to the true sequence. Because the length of the true sequence might be different from that of the observed transcribed sequence, the most likely sequences with possible different lengths are estimated with Viterbi decoding and the most likely length is then selected with a sophisticated language model based on a long short-term memory (LSTM) network. The experimental results show that the proposed method can correct musically unnatural transcription errors. Yuki Hiramatsu, Go Shibata, Ryo Nishikimi, Eita Nakamura, Kazuyoshi Yoshii |
ICASSP | 5 |
| 2021 | Autoregressive Fast Multichannel Nonnegative Matrix Factorization For Joint Blind Source Separation And DereverberationabstractThis paper describes a joint blind source separation and dereverberation method that works adaptively and efficiently in a reverberant noisy environment. The modern approach to blind source separation (BSS) is to formulate a probabilistic model of multichannel mixture signals that consists of a source model representing the time-frequency structures of source spectrograms and a spatial model representing the inter-channel covariance structures of source images. The cutting-edge BSS method in this thread of research is fast multi-channel nonnegative matrix factorization (FastMNMF) that consists of a low-rank source model based on nonnegative matrix factorization (NMF) and a full-rank spatial model based on jointly-diagonalizable spatial covariance matrices. Although FastMNMF is computationally efficient and can deal with both directional sources and diffuse noise simultaneously, its performance is severely degraded in a reverberant environment. To solve this problem, we propose autoregressive FastMNMF (AR-FastMNMF) based on a unified probabilistic model that combines FastMNMF with a blind dereverberation method called weighted prediction error (WPE), where all the parameters are optimized jointly such that the likelihood for observed reverberant mixture signals is maximized. Experimental results showed the superiority of AR-FastMNMF over conventional methods that perform blind dereverberation and BSS jointly or sequentially. Kouhei Sekiguchi, Yoshiaki Bando, Aditya Arie Nugraha, Mathieu Fontaine 0002, Kazuyoshi Yoshii |
ICASSP | 5 |
| 2021 | Pitch-Timbre Disentanglement Of Musical Instrument Sounds Based On Vae-Based Metric LearningabstractThis paper describes a representation learning method for disentangling an arbitrary musical instrument sound into latent pitch and timbre representations. Although such pitch-timbre disentanglement has been achieved with a variational autoencoder (VAE), especially for a predefined set of musical instruments, the latent pitch and timbre representations are outspread, making them hard to interpret. To mitigate this problem, we introduce a metric learning technique into a VAE with latent pitch and timbre spaces so that similar (different) pitches or timbres are mapped close to (far from) each other. Specifically, our VAE is trained with additional contrastive losses so that the latent distances between two arbitrary sounds of the same pitch or timbre are minimized, and those of different pitches or timbres are maximized. This training is performed under weak supervision that uses only whether the pitches and timbres of two sounds are the same or not, instead of their actual values. This improves the generalization capability for unseen musical instruments. Experimental results show that the proposed method can find better-structured disentangled representations with pitch and timbre clusters even for unseen musical instruments. Keitaro Tanaka, Ryo Nishikimi, Yoshiaki Bando, Kazuyoshi Yoshii, Shigeo Morishima |
ICASSP | 4 |
| 2021 | Alpha-Stable Autoregressive Fast Multichannel Nonnegative Matrix Factorization for Joint Speech Enhancement and Dereverberation
Mathieu Fontaine 0002, Kouhei Sekiguchi, Aditya Arie Nugraha, Yoshiaki Bando, Kazuyoshi Yoshii |
Interspeech | 5 |
| 2021 | A Real-Time Drum-Wise Volume Visualization System for Learning Volume-Balanced Drum Performance
Mitsuki Hosoya, Masanori Morise, Satoshi Nakamura 0002, Kazuyoshi Yoshii |
ICEC | 4 |
| 2021 | Musical rhythm transcription based on Bayesian piece-specific score models capturing repetitions
Eita Nakamura, Kazuyoshi Yoshii |
Inf. Sci. | 2 |
| 2021 | Non-local musical statistics as guides for audio-to-score piano transcription
Kentaro Shibata, Eita Nakamura, Kazuyoshi Yoshii |
Inf. Sci. | 3 |
| 2021 | Neural Full-Rank Spatial Covariance Analysis for Blind Source SeparationabstractThis paper describes aneural blind source separation (BSS) method based on amortized variational inference (AVI) of a non-linear generative model of mixture signals. A classical statistical approach to BSS is to fit a linear generative model that consists of spatial and source models representing the inter-channel covariances and power spectral densities of sources, respectively. Although the variational autoencoder (VAE) has successfully been used as a non-linear source model with latent features, it should be pretrained from a sufficient amount of isolated signals. Our method, in contrast, enables the VAE-based source model to be trained only from mixture signals. Specifically, we introduce a neural mixture-to-feature inference model that directly infers the latent features from the observed mixture and integrate it with a neural feature-to-mixture generative model consisting of a full-rank spatial model and a VAE-based source model. All the models are optimized jointly such that the likelihood for the training mixtures is maximized in the framework of AVI. Once the inference model is optimized, it can be used for estimating the latent features of sources included in unseen mixture signals. The experimental results show that the proposed method outperformed the state-of-the-art BSS methods based on linear generative models and was comparable to a method based on supervised learning of the VAE-based sourcemodel. Yoshiaki Bando, Kouhei Sekiguchi, Yoshiki Masuyama, Aditya Arie Nugraha, Mathieu Fontaine 0002, Kazuyoshi Yoshii |
IEEE Signal Process. Lett. | 6 |
| 2020 | Unsupervised Robust Speech Enhancement Based on Alpha-Stable Fast Multichannel Nonnegative Matrix Factorization
Mathieu Fontaine 0002, Kouhei Sekiguchi, Aditya Arie Nugraha, Kazuyoshi Yoshii |
INTERSPEECH | 4 |
| 2020 | Adaptive Neural Speech Enhancement with a Denoising Variational Autoencoder
Yoshiaki Bando, Kouhei Sekiguchi, Kazuyoshi Yoshii |
INTERSPEECH | 3 |
| 2020 | Statistical learning and estimation of piano fingering
Eita Nakamura, Yasuyuki Saito, Kazuyoshi Yoshii |
Inf. Sci. | 3 |
| 2020 | Flow-Based Independent Vector Analysis for Blind Source SeparationabstractThis letter describes a time-varying extension of independent vector analysis (IVA) based on the normalizing flow (NF), called NF-IVA, for determined blind source separation of multichannel audio signals. As in IVA, NF-IVA estimates demixing matrices that transform mixture spectra to source spectra in the complex-valued spatial domain such that the likelihood of those matrices for the mixture spectra is maximized under some non-Gaussian source model. While IVA performs a time-invariant bijective linear transformation, NF-IVA performs a series of time-varying bijective linear transformations (flow blocks) adaptively predicted by neural networks. To regularize such transformations, we introduce a soft volume-preserving (VP) constraint. Given mixture spectra, the parameters of NF-IVA are optimized by gradient descent with backpropagation in an unsupervised manner. Experimental results show that NF-IVA successfully performs speech separation in reverberant environments with different numbers of speakers and microphones and that NF-IVA with the VP constraint outperforms NF-IVA without it, standard IVA with iterative projection, and improved IVA with gradient descent. Aditya Arie Nugraha, Kouhei Sekiguchi, Mathieu Fontaine 0002, Yoshiaki Bando, Kazuyoshi Yoshii |
IEEE Signal Process. Lett. | 5 |
| 2020 | Bayesian Singing Transcription Based on a Hierarchical Generative Model of Keys, Musical Notes, and F0 TrajectoriesabstractThis article describes automatic singing transcription (AST) that estimates a human-readable musical score of a sung melody represented with quantized pitches and durations from a given music audio signal. To achieve the goal, we propose a statistical method for estimating the musical score by quantizing a trajectory of vocal fundamental frequencies (F0s) in the time and frequency directions. Since vocal F0 trajectories considerably deviate from the pitches and onset times of musical notes specified in musical scores, the local keys and rhythms of musical notes should be taken into account. In this article we propose a Bayesian hierarchical hidden semi-Markov model (HHSMM) that integrates a musical score model describing the local keys and rhythms of musical notes with an F0 trajectory model describing the temporal and frequency deviations of an F0 trajectory. Given an F0 trajectory, a sequence of musical notes, that of local keys, and the temporal and frequency deviations can be estimated jointly by using a Markov chain Monte Carlo (MCMC) method. We investigated the effect of each component of the proposed model and showed that the musical score model improves the performance of AST. Ryo Nishikimi, Eita Nakamura, Masataka Goto, Katsutoshi Itoyama, Kazuyoshi Yoshii |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2020 | A Flow-Based Deep Latent Variable Model for Speech Spectrogram Modeling and EnhancementabstractThis article describes a deep latent variable model of speech power spectrograms and its application to semi-supervised speech enhancement with a deep speech prior. By integrating two major deep generative models, a variational autoencoder (VAE) and a normalizing flow (NF), in a mutually-beneficial manner, we formulate a flexible latent variable model called the NF-VAE that can extract low-dimensional latent representations from high-dimensional observations, akin to the VAE, and does not need to explicitly represent the distribution of the observations, akin to the NF. In this article, we consider a variant of NF called the generative flow (GF a.k.a. Glow) and formulate a latent variable model called the GF-VAE. We experimentally show that the proposed GF-VAE is better than the standard VAE at capturing fine-structured harmonics of speech spectrograms, especially in the high-frequency range. A similar finding is also obtained when the GF-VAE and the VAE are used to generate speech spectrograms from latent variables randomly sampled from the standard Gaussian distribution. Lastly, when these models are used as speech priors for statistical multichannel speech enhancement, the GF-VAE outperforms the VAE and the GF. Aditya Arie Nugraha, Kouhei Sekiguchi, Kazuyoshi Yoshii |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Fast Multichannel Nonnegative Matrix Factorization With Directivity-Aware Jointly-Diagonalizable Spatial Covariance Matrices for Blind Source SeparationabstractThis article describes a computationally-efficient blind source separation (BSS) method based on the independence, low-rankness, and directivity of the sources. A typical approach to BSS is unsupervised learning of a probabilistic model that consists of a source model representing the time-frequency structure of source images and a spatial model representing their inter-channel covariance structure. Building upon the low-rank source model based on nonnegative matrix factorization (NMF), which has been considered to be effective for inter-frequency source alignment, multichannel NMF (MNMF) assumes source images to follow multivariate complex Gaussian distributions with unconstrained full-rank spatial covariance matrices (SCMs). An effective way of reducing the computational cost and initialization sensitivity of MNMF is to restrict the degree of freedom of SCMs. While a variant of MNMF called independent low-rank matrix analysis (ILRMA) severely restricts SCMs to rank-1 matrices under an idealized condition that only directional and less-echoic sources exist, we restrict SCMs to jointly-diagonalizable yet full-rank matrices in a frequency-wise manner, resulting in FastMNMF1. To help inter-frequency source alignment, we then propose FastMNMF2 that shares the directional feature of each source over all frequency bins. To explicitly consider the directivity or diffuseness of each source, we also propose rank-constrained FastMNMF that enables us to individually specify the ranks of SCMs. Our experiments showed the superiority of FastMNMF over MNMF and ILRMA in speech separation and the effectiveness of the rank constraint in speech enhancement. Kouhei Sekiguchi, Yoshiaki Bando, Aditya Arie Nugraha, Kazuyoshi Yoshii, Tatsuya Kawahara |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | Bayesian Melody Harmonization Based on a Tree-Structured Generative Model of Chord Sequences and MelodiesabstractThis article describes a melody harmonization method that generates a sequence of chords (symbols and onset positions) for a given melody (a sequence of musical notes). A typical approach to melody harmonization is to use a hidden Markov model (HMM) that represents chords and notes as latent and observed variables, respectively. This approach, however, does not consider the syntactic functions (e.g., tonic, dominant, and subdominant) and hierarchical structure of chords that play vital roles in traditional harmony theories. In this paper, we propose a unified hierarchical generative model consisting of a probabilistic context-free grammar (PCFG) model generating chord symbols associated with syntactic functions, a metrical Markov model generating chord onset positions, and a Markov model generating a melody conditioned by a chord sequence. To estimate a musically natural tree structure, the PCFG is trained in a semi-supervised manner by using chord sequences with tree structure annotations. Given a melody, a sequence of a variable number of chords can be estimated by using a Markov chain Monte Carlo method that partially and iteratively updates the symbols, onset positions, and tree structure of chords according to the posterior distribution of chord sequences. Experimental results show that the proposed method outperformed the HMM-based method and a conventional rule-based method in terms of predictive abilities. Hiroaki Tsushima, Eita Nakamura, Kazuyoshi Yoshii |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Semi-Supervised Neural Chord Estimation Based on a Variational Autoencoder With Latent Chord Labels and FeaturesabstractThis paper describes a statistically-principled semi-supervised method of automatic chord estimation (ACE) that can make effective use of music signals regardless of the availability of chord annotations. The typical approach to ACE is to train a deep classification model (neural chord estimator) in a supervised manner by using only annotated music signals. In this discriminative approach, prior knowledge about chord label sequences (model output) has scarcely been taken into account. In contrast, we propose a unified generative and discriminative approach in the framework of amortized variational inference. More specifically, we formulate a deep generative model that represents the generative process of chroma vectors (observed variables) from discrete labels and continuous features (latent variables), which are assumed to follow a Markov model favoring self-transitions and a standard Gaussian distribution, respectively. Given chroma vectors as observed data, the posterior distributions of the latent labels and features are computed approximately by using deep classification and recognition models, respectively. These three models form a variational autoencoder and can be trained jointly in a semi-supervised manner. The experimental results show that the regularization of the classification model based on the Markov prior of chord labels and the generative model of chroma vectors improved the performance of ACE even under the supervised condition. The semi-supervised learning using additional non-annotated data can further improve the performance. Yiming Wu 0003, Tristan Carsault, Eita Nakamura, Kazuyoshi Yoshii |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2019 | Improved Metrical Alignment of Midi Performance Based on a Repetition-aware Online-adapted GrammarabstractThis paper presents an improvement on an existing grammar-based method for metrical structure detection and alignment, a task which involves aligning a repeated tree structure with an input stream of musical notes. The previous method achieves state-of-the-art results, but performs poorly when it lacks training data. Data annotated as it requires is not widely available, making this drawback of the method significant. We present a novel online learning technique to improve the grammar's performance on unseen rhythmic patterns using a dynamically learned piece-specific grammar. The piece-specific grammar can measure the musical well-formedness of the underlying alignment without requiring any training data. It instead relies on musical repetition and self-similarity, enabling the model to recognize repeated rhythmic patterns, even when a similar pattern was never seen in the training data. Using it, we see improved performance on a corpus containing only Bach compositions, as well as a second corpus containing works from a variety of composers, indicating that the online-learned grammar helps the model generalize to unseen rhythms and styles. Andrew McLeod, Eita Nakamura, Kazuyoshi Yoshii |
ICASSP | 3 |
| 2019 | Unsupervised Melody Style ConversionabstractWe study a method for converting the music style of a given melody to a target style (e.g. from classical music style to pop music style) based on unsupervised statistical learning. Following the analogy with machine translation, we propose a statistical formulation of style conversion based on integration of a music language model of the target style and an edit model representing the similarity between the original and arranged melodies. In supervised-learning approaches for constructing style-specific language models, it has been crucial to use data that properly specify a music style. To reduce reliance on manual data selection and annotation, we propose a novel statistical model that can spontaneously discover styles in pitch and rhythm organization. We also point out the importance of an edit model that incorporates syntactic functions of notes such as tonic and build a model that can infer such functions unsupervisedly. We confirm that the proposed method improves the quality of arrangement by examining the results and by subjective evaluation. Eita Nakamura, Kentaro Shibata, Ryo Nishikimi, Kazuyoshi Yoshii |
ICASSP | 4 |
| 2019 | Automatic Singing Transcription Based on Encoder-decoder Recurrent Neural Networks with a Weakly-supervised Attention MechanismabstractThis paper describes neural singing transcription that estimates a sequence of musical notes directly from the audio signal of singing voice in an end-to-end manner without time-aligned training data. A conventional approach to singing transcription is to perform vocal F0 estimation followed by musical note estimation. The performance of this approach, however, is severely limited because the F0 estimation errors propagate to the note estimation step and rich acoustic information cannot be used. In addition, it is difficult and time-consuming to split continuous signals of singing voices into segments corresponding to musical notes for making precise time-aligned transcriptions. To solve these problems, we use an encoder-decoder model with an attention mechanism that can automatically learn an input-output alignment and mapping, even from non-aligned training data. The main challenge of our study is to estimate temporal categories (note values) in addition to instantaneous categories (pitches). We thus propose a novel loss function for the attention weights of time-aligned notes for semi-supervised alignment training. By gradually reducing the weight of the loss function, a better input-output alignment can be learned much more quickly. We showed that our method performed well for isolated singing voice in popular music. Ryo Nishikimi, Eita Nakamura, Satoru Fukayama, Masataka Goto, Kazuyoshi Yoshii |
ICASSP | 5 |
| 2019 | A Deep Generative Model of Speech Complex SpectrogramsabstractThis paper proposes an approach to the joint modeling of the short-time Fourier transform magnitude and phase spectrograms with a deep generative model. We assume that the magnitude follows a Gaussian distribution and the phase follows a von Mises distribution. To improve the consistency of the phase values in the time-frequency domain, we also apply the von Mises distribution to the phase derivatives, i.e., the group delay and the instantaneous frequency. Based on these assumptions, we explore and compare several combinations of loss functions for training our models. Built upon the variational autoencoder framework, our model consists of three convolutional neural networks acting as an encoder, a magnitude decoder, and a phase decoder. In addition to the latent variables, we propose to also condition the phase estimation on the estimated magnitude. Evaluated for a time-domain speech reconstruction task, our models could generate speech with a high perceptual quality and a high intelligibility. Aditya Arie Nugraha, Kouhei Sekiguchi, Kazuyoshi Yoshii |
ICASSP | 3 |
| 2019 | Joint Transcription of Lead, Bass, and Rhythm Guitars Based on a Factorial Hidden Semi-Markov ModelabstractThis paper describes a statistical method for estimating musical scores for lead, bass, and rhythm guitars from polyphonic audio signals of typical band-style music. To perform multi-instrument transcription involving multi-pitch detection and part assignment, it is crucial to formulate a musical language model that represents the characteristics of each part in order to solve the ambiguity of part assignment and estimate a musically-natural score. We propose a factorial hidden semi-Markov model that consists of three language models corresponding to the three guitar parts (three latent chains) and an acoustic model of a mixture spectrogram (emission model). The language model for rhythm guitar represents a homophonic sequence of musical notes (chord sequence) and those for lead and bass guitars represent a monophonic sequence of musical notes in a higher and lower frequency range respectively. The acoustic model represents a spectrogram as a sum of low-rank spectrograms of the three guitar parts approximated by NMF. Given a spectrogram, we estimate the note sequences using Gibbs sampling. We show that our model outperforms a state-of-the-art multi-pitch detection method in the accuracy and naturalness of the transcribed scores. Kentaro Shibata, Ryo Nishikimi, Satoru Fukayama, Masataka Goto, Eita Nakamura, Katsutoshi Itoyama, Kazuyoshi Yoshii |
ICASSP | 7 |
| 2019 | Bayesian Drum Transcription Based on Nonnegative Matrix Factor Decomposition with a Deep Score PriorabstractThis paper describes a statistical method of automatic drum transcription that estimates a musical score of bass and snare drums and hi-hats from a drum signal separated from a popular music signal. One of the most effective approaches for this problem is to apply nonnegative matrix factor deconvolution (NMFD) for estimating the temporal activations of drums and then perform thresholding for estimating a drum score. Such a pure audio-based approach, however, cannot avoid musically unnatural scores. To solve this, we propose a unified Bayesian model that integrates an NMFD-based acoustic model evaluating the likelihood of a drum score for a drum spectrogram, with a deep language model serving as a prior (constraint) of the score. The language model can be trained with existing drum scores in the framework of autoencoding variational Bayes and has more expressive power than the conventional statistical models. We derive an inference algorithm using Gibbs sampling, which is a marriage of the solid formalism of Bayesian learning with the expressive power of deep learning. It is shown that the proposed method not only slightly improved the F-measure score but also increased musical naturalness of the transcribed drum scores than NMFD. Shun Ueda, Kentaro Shibata, Yusuke Wada, Ryo Nishikimi, Eita Nakamura, Kazuyoshi Yoshii |
ICASSP | 6 |
| 2019 | Audio-Visual SLAM towards Human Tracking and Human-Robot Interaction in Indoor EnvironmentsabstractWe propose a novel audio-visual simultaneous and localization (SLAM) framework that exploits human pose and acoustic speech of human sound sources to allow a robot equipped with a microphone array and a monocular camera to track, map, and interact with human partners in an indoor environment. Since human interaction is characterized by features perceived in not only the visual modality, but the acoustic modality as well, SLAM systems must utilize information from both modalities. Using a state-of-the-art beamforming technique, we obtain sound components correspondent to speech and noise; and estimate the Direction-of-Arrival (DoA) estimates of active sound sources as useful representations of observed features in the acoustic modality. Through estimated human pose by a monocular camera, we obtain the relative positions of humans as representation of observed features in the visual modality. Using these techniques, we attempt to eliminate restrictions imposed by intermittent speech, noisy periods, reverberant periods, triangulation of sound-source range, and limited visual field-of-views; and subsequently perform early fusion on these representations. We develop a system that allows for complimentary action between audio-visual sensor modalities in the simultaneous mapping of multiple human sound sources and the localization of observer position. Aaron Chau, Kouhei Sekiguchi, Aditya Arie Nugraha, Kazuyoshi Yoshii, Kotaro Funakoshi |
RO-MAN | 4 |
| 2019 | Semi-Supervised Multichannel Speech Enhancement With a Deep Speech PriorabstractThis paper describes a semi-supervised multichannel speech enhancement method that uses clean speech data for prior training. Although multichannel nonnegative matrix factorization (MNMF) and its constrained variant called independent low-rank matrix analysis (ILRMA) have successfully been used for unsupervised speech enhancement, the low-rank assumption on the power spectral densities (PSDs) of all sources (speech and noise) does not hold in reality. To solve this problem, we replace a low-rank speech model with a deep generative speech model, i.e., formulate a probabilistic model of noisy speech by integrating a deep speech model, a low-rank noise model, and a full-rank or rank-1 model of spatial characteristics of speech and noise. The deep speech model is trained from clean speech data in an unsupervised auto-encoding variational Bayesian manner. Given multichannel noisy speech spectra, the full-rank or rank-1 spatial covariance matrices and PSDs of speech and noise are estimated in an unsupervised maximum-likelihood manner. Experimental results showed that the full-rank version of the proposed method was significantly better than MNMF, ILRMA, and the rank-1 version. We confirmed that the initialization-sensitivity and local-optimum problems of MNMF with many spatial parameters can be solved by incorporating the precise speech model. Kouhei Sekiguchi, Yoshiaki Bando, Aditya Arie Nugraha, Kazuyoshi Yoshii, Tatsuya Kawahara |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2019 | Unsupervised Speech Enhancement Based on Multichannel NMF-Informed Beamforming for Noise-Robust Automatic Speech RecognitionabstractThis paper describes multichannel speech enhancement for improving automatic speech recognition (ASR) in noisy environments. Recently, the minimum variance distortionless response (MVDR) beamforming has widely been used because it works well if the steering vector of speech and the spatial covariance matrix (SCM) of noise are given. To estimating such spatial information, conventional studies take a supervised approach that classifies each time-frequency (TF) bin into noise or speech by training a deep neural network (DNN). The performance of ASR, however, is degraded in an unknown noisy environment. To solve this problem, we take an unsupervised approach that decomposes each TF bin into the sum of speech and noise by using multichannel nonnegative matrix factorization (MNMF). This enables us to accurately estimate the SCMs of speech and noise not from observed noisy mixtures but from separated speech and noise components. In this paper, we propose online MVDR beamforming by effectively initializing and incrementally updating the parameters of MNMF. Another main contribution is to comprehensively investigate the performances of ASR obtained by various types of spatial filters, i.e., time-invariant and variant versions of MVDR beamformers and those of rank-1 and full-rank multichannel Wiener filters, in combination with MNMF. The experimental results showed that the proposed method outperformed the state-of-the-art DNN-based beamforming method in unknown environments that did not match training data. Kazuki Shimada, Yoshiaki Bando, Masato Mimura, Katsutoshi Itoyama, Kazuyoshi Yoshii, Tatsuya Kawahara |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2018 | Statistical Speech Enhancement Based on Probabilistic Integration of Variational Autoencoder and Non-Negative Matrix FactorizationabstractThis paper presents a statistical method of single-channel speech enhancement that uses a variational autoencoder (VAE) as a prior distribution on clean speech. A standard approach to speech enhancement is to train a deep neural network (DNN) to take noisy speech as input and output clean speech. Although this supervised approach requires a very large amount of pair data for training, it is not robust against unknown environments. Another approach is to use non-negative matrix factorization (NMF) based on basis spectra trained on clean speech in advance and those adapted to noise on the fly. This semi-supervised approach, however, causes considerable signal distortion in enhanced speech due to the unrealistic assumption that speech spectrograms are linear combinations of the basis spectra. Replacing the poor linear generative model of clean speech in NMF with a VAE-a powerful nonlinear deep generative model-trained on clean speech, we formulate a unified probabilistic generative model of noisy speech. Given noisy speech as observed data, we can sample clean speech from its posterior distribution. The proposed method outperformed the conventional DNN-based method in unseen noisy environments. Yoshiaki Bando, Masato Mimura, Katsutoshi Itoyama, Kazuyoshi Yoshii, Tatsuya Kawahara |
ICASSP | 4 |
| 2018 | An End-to-End Approach to Joint Social Signal Detection and Automatic Speech RecognitionabstractSocial signals such as laughter and fillers are often observed in natural conversation, and they play various roles in human-to-human communication. Detecting these events is useful for transcription systems to generate rich transcription and for dialogue systems to behave as we do such as synchronized laughing or attentive listening. We have studied an end-to-end approach to directly detect social signals from speech by using connectionist temporal classification (CTC), which is one of the end-to-end sequence labelling models. In this work, we propose a unified framework that integrates social signal detection (SSD) and automatic speech recognition (ASR). We investigate several reference labelling methods regarding social signals. Experimental evaluations demonstrate that our end-to-end framework significantly outperforms the conventional DNN-HMM system with regard to SSD performance as well as the character error rate (CER). Hirofumi Inaguma, Masato Mimura, Koji Inoue, Kazuyoshi Yoshii, Tatsuya Kawahara |
ICASSP | 4 |
| 2018 | Towards Complete Polyphonic Music Transcription: Integrating Multi-Pitch Detection and Rhythm QuantizationabstractMost work on automatic transcription produces “piano roll” data with no musical interpretation of the rhythm or pitches. We present a polyphonic transcription method that converts a music audio signal into a human-readable musical score, by integrating multi-pitch detection and rhythm quantization methods. This integration is made difficult by the fact that the multi-pitch detection produces erroneous notes such as extra notes and introduces timing errors that are added to temporal deviations due to musical expression. Thus, we propose a rhythm quantization method that can remove extra notes by extending the metrical hidden Markov model and optimize the model parameters. We also improve the note-tracking process of multi-pitch detection by refining the treatment of repeated notes and adjustment of onset times. Finally, we propose evaluation measures for transcribed scores. Systematic evaluations on commonly used classical piano data show that these treatments improve the performance of transcription, which can be used as benchmarks for further studies. Eita Nakamura, Emmanouil Benetos, Kazuyoshi Yoshii, Simon Dixon |
ICASSP | 3 |
| 2018 | Unsupervised Beamforming Based on Multichannel Nonnegative Matrix Factorization for Noisy Speech RecognitionabstractThis paper presents unsupervised multichannel speech enhancement for noisy speech recognition. Time-frequency (TF) mask estimation has actively been studied for estimating the steering vectors and spatial covariance matrices of speech and noise used for beamforming. The state-of-the-art approach to mask estimation is to use deep neural networks (DNN s) for classifying the TF bins of observed signals into speech and noise. Such a supervised approach, however, does not work well in an unknown environment. To accurately estimate the spatial covariance matrices in an unsupervised manner, we perform blind source separation (BSS) based on multichannel nonnegative matrix factorization (MNMF) for decomposing each TF bin into the components of speech and the other sources (noise). To clarify a suitable type of beamforming for MNMF, we tested both time- invariant and time-varying versions of the minimum variance distortionless response (MVDR) beamforming in addition to standard multichannel Wiener filtering (MWF). The experimental results showed that our MNMF-based beamforming approach outperformed the state-of-the-art DNN-based beamforming method in unknown environments that do not match the training data. Kazuki Shimada, Yoshiaki Bando, Masato Mimura, Katsutoshi Itoyama, Kazuyoshi Yoshii, Tatsuya Kawahara |
ICASSP | 5 |
| 2018 | Correlated Tensor Factorization for Audio Source SeparationabstractThis paper presents an ultimate extension of nonnegative matrix factorization (NMF) for audio source separation based on full covariance modeling over all the time-frequency (TF) bins of the complex spectrogram of an observed mixture signal. Although NMF has been widely used for decomposing an observed power spectrogram in a TF-wise manner, it has a critical limitation that the phase values of interdependent TF bins cannot be dealt with. This problem has been solved only partially by several phase-aware extensions of NMF that decompose an observed complex spectrogram in an time-and/or frequency-wise manner. In this paper, we propose correlated tensor factorization (CTF) that approximates the full covariance matrix over all TF bins as the sum of the Kronecker products between basis covariance matrices over frequency bands and the corresponding ones over time frames. All the TF bins of the complex spectrogram of each source signal are estimated jointly in an interdependent manner via Wiener filtering. We discuss how to reduce the computational cost of CTF and report the results of comparative evaluation of CTF with its special cases such as NMF and positive semidefinite tensor factorization (PSDTF). Kazuyoshi Yoshii |
ICASSP | 1 |
| 2018 | Speech Enhancement Based on Bayesian Low-Rank and Sparse Decomposition of Multichannel Magnitude SpectrogramsabstractThis paper presents a blind multichannel speech enhancement method that can deal with the time-varying layout of microphones and sound sources. Since nonnegative tensor factorization (NTF) separates a multichannel magnitude (or power) spectrogram into source spectrograms without phase information, it is robust against the time-varying mixing system. This method, however, requires prior information such as the spectral bases (templates) of each source spectrogram in advance. To solve this problem, we develop a Bayesian model called robust NTF (Bayesian RNTF) that decomposes a multichannel magnitude spectrogram into target speech and noise spectrograms based on their sparseness and low rankness. Bayesian RNTF is applied to the challenging task of speech enhancement for a microphone array distributed on a hose-shaped rescue robot. When the robot searches for victims under collapsed buildings, the layout of the microphones changes over time and some of them often fail to capture target speech. Our method robustly works under such situations, thanks to its characteristic of time-varying mixing system. Experiments using a 3-m hose-shaped rescue robot with eight microphones show that the proposed method outperforms conventional blind methods in enhancement performance by the signal-to-noise ratio of 1.03 dB. Yoshiaki Bando, Katsutoshi Itoyama, Masashi Konyo, Satoshi Tadokoro, Kazuhiro Nakadai, Kazuyoshi Yoshii, Tatsuya Kawahara, Hiroshi G. Okuno |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2018 | Bayesian Multichannel Audio Source Separation Based on Integrated Source and Spatial ModelsabstractThis paper presents new statistical methods of multichannel audio source separation based on unified source and spatial models that, respectively, represent the generative process of latent source spectrograms and that of observed mixture spectrograms. One possibility of the source model is a factor model based on nonnegative matrix factorization that represents each time-frequency (TF) bin as the weighted sum of basis spectra. Another possibility is a mixture model inspired by latent Dirichlet allocation that exclusively classifies each TF bin into one of basis spectra. Similarly, the spatial model can either be a factor model that represents each TF bin as the weighted sum of source spectra or a mixture model that classifies each bin into one of those spectra. To unify these models in a principled manner and incorporate prior knowledge of a microphone array, we propose hierarchical Bayesian models of all the source-spatial combinations (factor-factor, mixture-factor, factor-mixture, and mixture-mixture models) and derive efficient Gibbs sampling algorithms for posterior inference. Experimental results showed that the proposed unified models outperformed the state-of-the-art method using only the spatial mixture model. Among the four unified models, the spatial factor model tended to work better than the spatial mixture model in exchange for larger computational cost, and the choice of source models had a little impact on the performance and computational cost. Kousuke Itakura, Yoshiaki Bando, Eita Nakamura, Katsutoshi Itoyama, Kazuyoshi Yoshii, Tatsuya Kawahara |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2017 | Bayesian multichannel nonnegative matrix factorization for audio source separation and localizationabstractThis paper presents a Bayesian extension of multichannel nonnegative matrix factorization (MNMF) that decomposes the complex spectrograms of mixture signals recorded by a microphone array into basis spectra, their temporal activations, and the spatial correlation matrices of sources (directions) in the time-frequency-channel domain. Although the original MNMF can be used in a blind setting, prior knowledge of a microphone array is useful for improving source separation. The impulse response (spatial correlation matrix) of each direction can be measured in an anechoic room, however, it differs from that in a real environment where the microphone array is used. To solve this, we propose a unified Bayesian model of source separation and localization by introducing a prior distribution determined by an anechoic spatial correlation matrix on a real spatial correlation matrix with respect to each direction. This enables us to adaptively estimate a real spatial correlation matrix and the direction of each source. Experimental results showed that our method outperformed the original MNMF and the state-of-the-art methods with prior knowledge in terms of signal-to-distortion ratio (SDR) even when the method was used in an unknown environment with acoustic characteristics different from those of the anechoic room. Kousuke Itakura, Yoshiaki Bando, Eita Nakamura, Katsutoshi Itoyama, Kazuyoshi Yoshii, Tatsuya Kawahara |
ICASSP | 5 |
| 2017 | Combined Multi-Channel NMF-Based Robust Beamforming for Noisy Speech Recognition
Masato Mimura, Yoshiaki Bando, Kazuki Shimada, Shinsuke Sakai, Kazuyoshi Yoshii, Tatsuya Kawahara |
INTERSPEECH | 5 |
| 2017 | Note Value Recognition for Piano Transcription Using Markov Random FieldsabstractThis paper presents a statistical method for use in music transcription that can estimate score times of note onsets and offsets from polyphonic MIDI performance signals. Because performed note durations can deviate largely from score-indicated values, previous methods had the problem of not being able to accurately estimate offset score times (or note values) and, thus, could only output incomplete musical scores. Based on observations that the pitch context and onset score times are influential on the configuration of note values, we construct a context-tree model that provides prior distributions of note values using these features and combine it with a performance model in the framework of Markov random fields. Evaluation results show that our method reduces the average error rate by around 40 percent compared to existing/simple methods. We also confirmed that, in our model, the score model plays a more important role than the performance model, and it automatically captures the voice structure by unsupervised learning. Eita Nakamura, Kazuyoshi Yoshii, Simon Dixon |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Rhythm Transcription of Polyphonic Piano Music Based on Merged-Output HMM for Multiple VoicesabstractIn a recent conference paper, we have reported a rhythm transcription method based on a merged-output hidden Markov model (HMM) that explicitly describes the multiple-voice structure of polyphonic music. This model solves a major problem of conventional methods that could not properly describe the nature of multiple voices as in polyrhythmic scores or in the phenomenon of loose synchrony between voices. In this paper, we present a complete description of the proposed model and develop an inference technique, which is valid for any merged-output HMMs, for which output probabilities depend on past events. We also examine the influence of the architecture and parameters of the method in terms of accuracies of rhythm transcription and voice separation and perform comparative evaluations with six other algorithms. Using MIDI recordings of classical piano pieces, we found that the proposed model outperformed other methods by more than 12 points in the accuracy for polyrhythmic performances and performed almost as good as the best one for non-polyrhythmic performances. This reveals the state-of-the-art methods of rhythm transcription for the first time in the literature. Publicly available source codes are also provided for future comparisons. Eita Nakamura, Kazuyoshi Yoshii, Shigeki Sagayama |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Tree-structured probabilistic model of monophonic written music based on the generative theory of tonal musicabstractThis paper presents a probabilistic formulation of music language modelling based on the generative theory of tonal music (GTTM) named probabilistic GTTM (PGTTM). GTTM is a well-known music theory that describes the tree structure of written music in analogy with the phrase structure grammar of natural language. To develop a computational music language model incorporating GTTM and a machine-learning framework for data-driven music grammar induction, we construct a generative model of monophonic music based on probabilistic context-free grammar, in which the time-span tree proposed in GTTM corresponds to the parse tree. Applying the techniques of natural language processing, we also derive supervised and unsupervised learning algorithms based on the maximal-likelihood estimation, and a Bayesian inference algorithm based on the Gibbs sampling. Despite the conceptual simplicity of the model, we found that the model automatically acquires music grammar from data and reproduces time-span trees of written music as accurately as an analyser that required elaborate manual parameter tuning. Eita Nakamura, Masatoshi Hamanaka, Keiji Hirata 0001, Kazuyoshi Yoshii |
ICASSP | 4 |
| 2016 | Student's T nonnegative matrix factorization and positive semidefinite tensor factorization for single-channel audio source separationabstractThis paper presents a robust variant of nonnegative matrix factorization (NMF) based on complex Student's t distributions (t-NMF) for source separation of single-channel audio signals. The Itakura-Saito divergence NMF (Gaussian NMF) is justified for this purpose under an assumption that the complex spectra of source signals and those of the mixture signal are complex Gaussian distributed (the additiv-ity of power spectra holds). In fact, however, the source spectra are often heavy-tailed distributed. When the source spectra are complex Cauchy distributed, for example, the mixture spectra are also complex Cauchy distributed (the additivity of amplitude spectra holds). Using the complex t distribution that includes the complex Gaussian and Cauchy distributions as its special cases, we propose t-NMF as a unified extension of Gaussian NMF and Cauchy NMF. Furthermore, we propose the corresponding variant of positive semidefinite tensor factorization based on multivariate complex t distributions (t-PSDTF). The experimental results showed that while t-NMF and t-PSDTF were comparative to Gaussian counterparts in terms of peak performance, they worked much better on average because they are insensitive to initialization and tend to avoid local optima. Kazuyoshi Yoshii, Katsutoshi Itoyama, Masataka Goto |
ICASSP | 1 |
| 2016 | Online simultaneous localization and mapping of multiple sound sources and asynchronous microphone arraysabstractThis paper presents an online method of simultaneous localization and mapping (SLAM) for estimating the positions of multiple moving sound sources and stationary robots and synchronizing microphone arrays attached to those robots. Since each robot with a microphone array can solely estimate the directions of sound sources, the two-dimensional source positions can be estimated from the source directions estimated by multiple robots using a triangulation method. In addition, sound mixtures can be separated accurately by regarding distributed microphone arrays as one big array. To perform these tasks, some methods have been proposed for localizing and synchronizing microphone arrays. These methods, however, can be used only if a single sound source exists because the time differences of arrival (TDOAs) between microphones are assumed to be directly observed. To overcome this limitation, we propose a unified state-space model that encodes the source and robot positions and the time offsets between microphone arrays in a latent space. Given the TDOAs and directions of arrival (DOAs) estimated by separating observed mixture sounds into source sounds, the latent variables are estimated jointly in an online manner using a FastSLAM2.0 algorithm that can deal with an unknown time-varying number of moving sound sources. Kouhei Sekiguchi, Yoshiaki Bando, Keisuke Nakamura, Kazuhiro Nakadai, Katsutoshi Itoyama, Kazuyoshi Yoshii |
IROS | 6 |
| 2016 | Singing Voice Separation and Vocal F0 Estimation Based on Mutual Combination of Robust Principal Component Analysis and Subharmonic SummationabstractThis paper presents a new method of singing voice analysis that performs mutually-dependent singing voice separation and vocal fundamental frequency (F0) estimation. Vocal F0 estimation is considered to become easier if singing voices can be separated from a music audio signal, and vocal F0 contours are useful for singing voice separation. This calls for an approach that improves the performance of each of these tasks by using the results of the other. The proposed method first performs robust principal component analysis (RPCA) for roughly extracting singing voices from a target music audio signal. The F0 contour of the main melody is then estimated from the separated singing voices by finding the optimal temporal path over an F0 saliency spectrogram. Finally, the singing voices are separated again more accurately by combining a conventional time-frequency mask given by RPCA with another mask that passes only the harmonic structures of the estimated F0s. Experimental results showed that the proposed method significantly improved the performances of both singing voice separation and vocal F0 estimation. The proposed method also outperformed all the other methods of singing voice separation submitted to an international music analysis competition called MIREX 2014. Yukara Ikemiya, Katsutoshi Itoyama, Kazuyoshi Yoshii |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2015 | Challenges in deploying a microphone array to localize and separate sound sources in real auditory scenesabstractAnalyzing the auditory scene of real environments is challenging partly because an unknown number and type of sound sources are observed at the same time and partly because these sounds are observed on a significantly different sound pressure level at the microphone. These are difficult problems even with state-of-the-art sound source localization and separation methods. In this paper, we exploit two such methods using a microphone array: (1) Bayesian nonparametric microphone array processing (BNP-MAP), which is capable of separating and localizing sound sources when the number of sound sources is unspecified, and (2) robot audition software “HARK” is capable of separating and localizing in real time. Through experimentation, we found that BNP-MAP is more robust against differences in the sound pressure levels of the source signals and in the spatial closeness of source positions. Experiments analyzing real scenes of human conversations recorded in a big exhibition hall and bird calling recorded at a natural park demonstrate the efficacy and applicability of BNP-MAP. Yoshiaki Bando, Takuma Otsuka, Katsutoshi Itoyama, Kazuyoshi Yoshii, Yoko Sasaki, Satoshi Kagami, Hiroshi G. Okuno |
ICASSP | 4 |
| 2015 | Singing voice analysis and editing based on mutually dependent F0 estimation and source separationabstractThis paper presents a novel framework that improves both vocal fundamental frequency (F0) estimation and singing voice separation by making effective use of the mutual dependency of those two tasks. A typical approach to singing voice separation is to estimate the vocal F0 contour from a target music signal and then extract the singing voice by using a time-frequency mask that passes only the harmonic components of the vocal F0s and overtones. Vocal F0 estimation, on the contrary, is considered to become easier if only the singing voice can be extracted accurately from the target signal. Such mutual dependency has scarcely been focused on in most conventional studies. To overcome this limitation, our framework alternates those two tasks while using the results of each in the other. More specifically, we first extract the singing voice by using robust principal component analysis (RPCA). The F0 contour is then estimated from the separated singing voice by finding the optimal path over a F0-saliency spectrogram based on subharmonic summation (SHS). This enables us to improve singing voice separation by combining a time-frequency mask based on RPCA with a mask based on harmonic structures. Experimental results obtained when we used the proposed technique to directly edit vocal F0s in popular-music audio signals showed that it significantly improved both vocal F0 estimation and singing voice separation. Yukara Ikemiya, Kazuyoshi Yoshii, Katsutoshi Itoyama |
ICASSP | 2 |
| 2015 | A feedback framework for improved chord recognition based on NMF-based approximate note transcriptionabstractThis paper presents a feedback framework that can improve chord recognition for music audio signals by performing approximate note transcription with Bayesian non-negative matrix factorization (NMF) using prior knowledge on chords. Although the names and note compositions of chords are intrinsically linked with each other (e.g., C major chords are highly likely to include C, E, and G notes, and those notes are highly likely to be in C major chords), chord recognition and note transcription (multipitch analysis) have been studied independently. To solve this chicken-and-egg problem, our framework iterates chord recognition and approximate note transcription using each other's results. More specifically, we first perform approximate note transcription based on Bayesian NMF that forces basis spectra to respectively correspond to different semitone-level pitches covering the whole range. We then execute chord recognition based on Bayesian hidden Markov models (HMMs) that use chroma features obtained from the activation patterns of those pitches. To improve note transcription, we again perform Bayesian NMF that encourages certain kinds of pitches in each chord region to be activated. Experimental results showed that our feedback framework gradually improved the accuracy of chord recognition. Satoshi Maruo, Kazuyoshi Yoshii, Katsutoshi Itoyama, Matthias Mauch, Masataka Goto |
ICASSP | 2 |
| 2015 | Bayesian integration of sound source separation and speech recognition: a new approach to simultaneous speech recognitionabstractThis paper presents a novel Bayesian method that can directly recognize overlapping utterances without explicitly separating mixture signals into their independent components in advance of speech recognition. The conventional approach to contami-nated speech recognition in real environments uniquely extracts the clean isolated signals of individual sources (e.g., by noise reduction, dereverberation, and source separation). One of the main limitations of this cascading approach is that the accuracy of speech recognition is upper bounded by the accuracy of pre-processing. To overcome this limitation, our method marginal-izes out uncertain isolated speech signals by integrating source separation and speech recognition in a Bayesian manner. A suf-ficient number of samples are drawn from the posterior distribu-tion of isolated speech signals by using a Markov chain Monte Carlo method, and then the posterior distributions of uttered texts for those samples are integrated. Under a certain con-dition, this Monte Carlo integration is shown to reduce to the well-known method called ROVER that integrates recognized texts obtained from sampled speech signals. Results of simulta-neous speech recognition experiments showed that in terms of word accuracy the proposed method significantly outperformed conventional cascading methods. Index Terms: simultaneous speech recognition, sound source separation, Bayesian modeling, MCMC, ROVER Kousuke Itakura, Izaya Nishimuta, Yoshiaki Bando, Katsutoshi Itoyama, Kazuyoshi Yoshii |
INTERSPEECH | 5 |
| 2015 | Microphone-accelerometer based 3D posture estimation for a hose-shaped rescue robotabstract3D posture estimation for a hose-shaped robot is critical in rescue activities due to complex physical environments. Conventional sound-based posture estimation assumes rather flat physical environments and focuses only on 2D, resulting in poor performance in real world environments with rubble. This paper presents novel 3D posture estimation by exploiting microphones and accelerometers. The idea of our method is to compensate the lack of posture information obtained by sound-based time-difference-of arrival (TDOA) with the tilt information obtained from accelerometers. This compensation is formulated as a nonlinear state-space model and solved by the unscented Kalman filter. Experiments are conducted by using a 3m hose-shaped robot with eight units of a microphone and an accelerometer and seven units of a loudspeaker and a vibration motor deployed in a simple 3D structure. Experimental results demonstrate that our method reduces the errors of initial states to about 20 cm in the 3D space. If the initial errors of initial states are less than 20 %, our method can estimate the correct 3D posture in real-time. Yoshiaki Bando, Katsutoshi Itoyama, Masashi Konyo, Satoshi Tadokoro, Kazuhiro Nakadai, Kazuyoshi Yoshii, Hiroshi G. Okuno |
IROS | 6 |
| 2015 | Audio-visual beat tracking based on a state-space model for a music robot dancing with humansabstractThis paper presents an audio-visual beat-tracking method for an entertainment robot that can dance in synchronization with music and human dancers. Conventional music robots have focused on either music audio signals or dancing movements of humans for detecting and predicting beat times in real time. Since a robot needs to record music audio signals by using its own microphones, however, the signals are severely contaminated with loud environmental noise and reverberant sounds. Moreover, it is difficult to visually detect beat times from real complicated dancing movements that exhibit weaker repetitive characteristics than music audio signals do. To solve these problems, we propose a state-space model that integrates both audio and visual information in a probabilistic manner. At each frame, the method extracts acoustic features (audio tempos and onset likelihoods) from music audio signals and extracts skeleton features from movements of a human dancer. The current tempo and the next beat time are then estimated from those observed features by using a particle filter. Experimental results showed that the proposed multi-modal method using a depth sensor (Kinect) for extracting skeleton features outperformed conventional mono-modal methods by 0.20 (F measure) in terms of beat-tracking accuracy in a noisy and reverberant environment. Misato Ohkita, Yoshiaki Bando, Yukara Ikemiya, Katsutoshi Itoyama, Kazuyoshi Yoshii |
IROS | 5 |
| 2015 | Optimizing the layout of multiple mobile robots for cooperative sound source separationabstractThis paper presents a novel active audition method that enables multiple mobile robots to move to optimal positions for improving the performance of sound source separation. A main advantage of our distributed system is that each robot has its own microphone array and all mobile robots can collaborate on source separation by regarding a set of movable microphone arrays as a big reconfigurable array. To incrementally optimize the positions of the robots (the layout of the big microphone array) in an active-audition manner, it is necessary to predict the source separation performance from a possible layout of the next time step although true source signals are unknown. To solve this problem, our method simulates delay-and-sum beamforming from a possible layout for theoretically calculating the gain for each frequency component of a source signal in the corresponding separated signal. The robots are moved into a layout with the highest average gain over all sources and the whole frequency range. The experimental results showed that the harmonic mean of signal-to-distortion ratios (SDRs) was improved by 6.0 dB in simulations and by 5.7 dB in a real environment. Kouhei Sekiguchi, Yoshiaki Bando, Katsutoshi Itoyama, Kazuyoshi Yoshii |
IROS | 4 |
| 2015 | Songle Widget: Making Animation and Physical Devices Synchronized with Music Videos on the WebabstractThis paper describes a web-based multimedia development framework, Songle Widget, that makes it possible to control computer-graphic animation and physical devices such as lighting devices and robots in synchronization with music publicly available on the web. To avoid the difficulty of time-consuming manual annotation, Songle Widget makes it easy to develop web-based applications with rigid music synchronization by leveraging music-understanding technologies. Four types of musical elements (music structure, hierarchical beat structure, melody line, and chords) have been automatically annotated for more than 920,000 songs on music-or video-sharing services and can readily be used by music-synchronized applications. Since errors are inevitable when elements are annotated automatically, Songle Widget takes advantage of a user-friendly crowdsourcing interface that enables users to correct them. This is effective when applications require error-free annotation. We made Songle Widget open to the public, and its capabilities and usefulness have been demonstrated in seven music-synchronized applications. Masataka Goto, Kazuyoshi Yoshii, Tomoyasu Nakano |
ISM | 2 |
| 2015 | Musical Similarity and Commonness Estimation Based on Probabilistic Generative ModelsabstractThis paper proposes a novel concept we call musical commonness, which is the similarity of a song to a set of songs, in other words, its typicality. This commonness can be used to retrieve representative songs from a song set (e.g., songs released in the 80s or 90s). Previous research on musical similarity has compared two songs but has not evaluated the similarity of a song to a set of songs. The methods presented here for estimating the similarity and commonness of polyphonic musical audio signals are based on a unified framework of probabilistic generative modeling of four musical elements (vocal timbre, musical timbre, rhythm, and chord progression). To estimate the commonness, we use a generative model trained from a song set instead of estimating musical similarities of all possible song-pairs by using a model trained from each song. In experimental evaluation, we used 3278 popular music songs. Estimated song-pair similarities are comparable to ratings by a musician at the 0.1% significance level for vocal and musical timbre, at the 1% level for rhythm, and the 5% level for chord progression. Results of commonness evaluation show that the higher the musical commonness is, the more similar a song is to songs of a song set. Tomoyasu Nakano, Kazuyoshi Yoshii, Masataka Goto |
ISM | 2 |
| 2015 | Identification and Localization of One or Two Concurrent Speakers in a Binaural Robotic ContextabstractThis paper presents a method of identification and azimuth estimation for one or two concurrent speakers in simultaneous utterances. This method is applicable to human-machine interaction and robot audition. Identification and localization have been rarely mutually addressed and related works rely on time-frequency exploitation strategies to extract and treat each source's contribution to the received signal. The presented method relies on a training made with one speaker at a time, but it can exploit a speech segment to identify and localize two speakers. A cochlear filtering-based binaural front-end allows to extract equivalent rectangular bandwidth frequency cepstral coefficients (ERBFCC) and interaural level difference (ILD) features. Artificial neural networks (ANNs) exploit ERBFCCs to provide identity information, and a histogram-based exploitation of ILDs provides azimuth angle information. The method was evaluated in contexts including overlapping segments in the presence of noises and sound reflections and its efficiency was demonstrated. Even with fully overlapping utterances, we reached an 83% identification rate of both speakers, an 82% estimation accuracy of both azimuths and an 68% correct mutual identity and azimuth estimation rate. At least one speaker was correctly identified and localized in more than 99% of the tests for utterances lasting near 5s. Karim Youssef, Katsutoshi Itoyama, Kazuyoshi Yoshii |
SMC | 3 |
| 2014 | Timbre replacement of harmonic and drum components for music audio signalsabstractThis paper presents a system that allows users to customize an audio signal of polyphonic music (input), without using musical scores, by replacing the frequency characteristics of harmonic sounds and the timbres of drum sounds with those of another audio signal of polyphonic music (reference). To develop the system, we first use a method that can separate the amplitude spectra of the input and reference signals into harmonic and percussive spectra. We characterize frequency characteristics of the harmonic spectra by two envelopes tracing spectral dips and peaks roughly, and the input harmonic spectra are modified such that their envelopes become similar to those of the reference harmonic spectra. The input and reference percussive spectrograms are further decomposed into those of individual drum instruments, and we replace the timbres of those drum instruments in the input piece with those in the reference piece. Through the subjective experiment, we show that our system can replace drum timbres and frequency characteristics adequately. Tomohiko Nakamura, Hirokazu Kameoka, Kazuyoshi Yoshii, Masataka Goto |
ICASSP | 3 |
| 2014 | Vocal timbre analysis using latent Dirichlet allocation and cross-gender vocal timbre similarityabstractThis paper presents a vocal timbre analysis method based on topic modeling using latent Dirichlet allocation (LDA). Although many works have focused on analyzing characteristics of singing voices, none have dealt with “latent” characteristics (topics) of vocal timbre, which are shared by multiple singing voices. In the work described in this paper, we first automatically extracted vocal timbre features from polyphonic musical audio signals including vocal sounds. The extracted features were used as observed data, and mixing weights of multiple topics were estimated by LDA. Finally, the semantics of each topic were visualized by using a word-cloud-based approach. Experimental results for a singer identification task using 36 songs sung by 12 singers showed that our method achieved a mean reciprocal rank of 0.86. We also proposed a method for estimating cross-gender vocal timbre similarity by generating pitch-shifted (frequency-warped) signals of every singing voice. Experimental results for a cross-gender singer retrieval task showed that our method discovered interesting similar pitch-shifted singers. Tomoyasu Nakano, Kazuyoshi Yoshii, Masataka Goto |
ICASSP | 2 |
| 2014 | Cultivating vocal activity detection for music audio signals in a circulation-type crowdsourcing ecosystemabstractThis paper presents a crowdsourcing-based self-improvement framework of vocal activity detection (VAD) for music audio signals. A standard approach to VAD is to train a vocal-and-non-vocal classifier by using labeled audio signals (training set) and then use that classifier to label unseen signals. Using this technique, we have developed an online music-listening service called Songle that can help users better understand music by visualizing automatically estimated vocal regions and pitches of arbitrary songs existing on the Web. The accuracy of VAD is limited, however, because in general the acoustic characteristics of the training set are different from those of real songs on the Web. To overcome this limitation, we adapt a classifier by leveraging vocal regions and pitches corrected by volunteer users. UnlikeWikipedia-type crowdsourcing, our Songle-based framework can amplify user contributions: error corrections made for a limited number of songs improve VAD for all songs. This gives better music listening experiences to all users as non-monetary rewards. Kazuyoshi Yoshii, Hiromasa Fujihara, Tomoyasu Nakano, Masataka Goto |
ICASSP | 1 |
| 2014 | AutoMashUpper: automatic creation of multi-song music mashupsabstractIn this paper we present a system, AutoMashUpper, for making multi-song music mashups. Central to our system is a measure of “mashability” calculated between phrase sections of an input song and songs in a music collection. We define mashability in terms of harmonic and rhythmic similarity and a measure of spectral balance. The principal novelty in our approach centres on the determination of how elements of songs can be made fit together using key transposition and tempo modification, rather than based on their unaltered properties. In this way, the properties of two songs used to model their mashability can be altered with respect to transformations performed to maximize their perceptual compatibility. AutoMashUpper has a user interface to allow users to control the parameterization of the mashability estimation. It allows users to define ranges for key shifts and tempo as well as adding, changing or removing elements from the created mashups. We evaluate AutoMashUpper by its ability to reliably segment music signals into phrase sections, and also via a listening test to examine the relationship between estimated mashability and user enjoyment. Matthew E. P. Davies, Philippe Hamel, Kazuyoshi Yoshii, Masataka Goto |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2014 | Nonparametric Bayesian dereverberation of power spectrograms based on infinite-order autoregressive processesabstractThis paper describes a monaural audio dereverberation method that operates in the power spectrogram domain. The method is robust to different kinds of source signals such as speech or music. Moreover, it requires little manual intervention, including the complexity of room acoustics. The method is based on a non-conjugate Bayesian model of the power spectrogram. It extends the idea of multi-channel linear prediction to the power spectrogram domain, and formulates a model of reverberation as a non-negative, infinite-order autoregressive process. To this end, the power spectrogram is interpreted as a histogram count data, which allows a nonparametric Bayesian model to be used as the prior for the autoregressive process, allowing the effective number of active components to grow, without bound, with the complexity of data. In order to determine the marginal posterior distribution, a convergent algorithm, inspired by the variational Bayes method, is formulated. It employs the minorization-maximization technique to arrive at an iterative, convergent algorithm that approximates the marginal posterior distribution. Both objective and subjective evaluations show advantage over other methods based on the power spectrum. We also apply the method to a music information retrieval task and demonstrate its effectiveness. Akira Maezawa, Katsutoshi Itoyama, Kazuyoshi Yoshii, Hiroshi G. Okuno |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2013 | Infinite kernel linear prediction for joint estimation of spectral envelope and fundamental frequencyabstractThis paper presents a new probabilistic formulation of linear prediction (LP) for jointly estimating the spectral envelope and fundamental frequency (F0) of a speech signal. A main problem of classical LP is that the peaks of the estimated envelope are highly biased toward the harmonic partials of a speech spectrum. To solve this problem, we propose a nonparametric Bayesian model called infinite kernel linear prediction (IKLP) based on a Gaussian process with multiple kernel learning. Our model can represent the periodicity of a speech signal by using a weighted sum of infinitely many periodic kernels that correspond to different F0s. We put a gamma process prior on the positive weights of those kernels and perform sparse learning to determine a predominant kernel indicating the F0 at the same time of spectral envelope estimation. The experimental results showed that our model can estimate spectral envelopes and F0s of speech and singing signals while identifying pitched segments. Kazuyoshi Yoshii, Masataka Goto |
ICASSP | 1 |
| 2013 | Infinite Positive Semidefinite Tensor Factorization for Source Separation of Mixture SignalsabstractThis paper presents a new class of tensor factorization called positive semidefinite tensor factorization (PSDTF) that decomposes a set of positive semidefinite (PSD) matrices into the convex combinations of fewer PSD basis matrices. PSDTF can be viewed as a natural extension of nonnegative matrix factorization. One of the main problems of PSDTF is that an appropriate number of bases should be given in advance. To solve this problem, we propose a nonparametric Bayesian model based on a gamma process that can instantiate only a limited number of necessary bases from the infinitely many bases assumed to exist. We derive a variational Bayesian algorithm for closed-form posterior inference and a multiplicative update rule for maximum-likelihood estimation. We evaluated PSDTF on both synthetic data and real music recordings to show its superiority. Kazuyoshi Yoshii, Ryota Tomioka, Daichi Mochihashi, Masataka Goto |
ICML (3) | 1 |
| 2013 | Nested iGMM recognition and multiple hypothesis tracking of moving sound sources for mobile robot auditionabstractThe paper proposes two modules for a mobile robot audition system: 1) recognizing surrounding acoustic event, 2) tracking moving sound sources. We propose nested infinite Gaussian mixture model (iGMM) for recognizing frame based feature vectors. The main advantage is that the number of classes is allowed to increase without bound, if necessary, to represent unknown audio input. The multiple hypothesis tracking module provides time-series of separated audio stream using localized directions and recognition results at each frame. Not only for continuous sounds, the proposed tracker automatically detects appearing and disappearing point of stream from multiple hypothesis. These two modules are connected to microphone array based sound localization and separation, and the combined robot audition system achieved tracking of multiple moving sounds including intermittent sound source. Yoko Sasaki, Naotaka Hatao, Kazuyoshi Yoshii, Satoshi Kagami |
IROS | 3 |
| 2012 | Unsupervised music understanding based on nonparametric Bayesian modelsabstractThis paper presents a new research framework for unsupervised music understanding. Our goal is to recognize musical notes from polyphonic audio signals and simultaneously induce grammatical patterns from the recognized notes by integrating probabilistic acoustic and language models. Given music audio signals, both models could be jointly trained in a self-organizing manner without manually specifying the numbers of musical notes and grammatical patterns. In this paper, we introduce our nonparametric Bayesian acoustic and language models for multipitch analysis and chord progression analysis and discuss issues for integrating these models. We then provide a novel overview of various acoustic and language models whose underlying concepts are useful for implementing the framework. Kazuyoshi Yoshii, Masataka Goto |
ICASSP | 1 |
| 2012 | A Nonparametric Bayesian Multipitch Analyzer Based on Infinite Latent Harmonic AllocationabstractThe statistical multipitch analyzer described in this paper estimates multiple fundamental frequencies (F0s) in polyphonic music audio signals produced by pitched instruments. It is based on hierarchic4al nonparametric Bayesian models that can deal with uncertainty of unknown random variables such as model complexities (e.g., the number of F0s and the number of harmonic partials), model parameters (e.g., the values of F0s and the relative weights of harmonic partials), and hyperparameters (i.e., prior knowledge on complexities and parameters). Using these models, we propose a statistical method called infinite latent harmonic allocation (iLHA). To avoid model-complexity control, we allow the observed spectra to contain an unbounded number of sound sources (F0s), each of which is allowed to contain an unbounded number of harmonic partials. More specifically, to model a set of time-sliced spectra, we formulated nested infinite Gaussian mixture models based on hierarchical and generalized Dirichlet processes. To avoid manual tuning of influential hyperparameters, we put noninformative hyperprior distributions on them in a hierarchical manner. For efficient Bayesian inference, we used a modern technique called collapsed variational Bayes. In comparative experiments using audio recordings of piano and guitar solo performances, iLHA yielded promising results and we found that there would be room for improvement based on modeling of temporal continuity and spectral smoothness. Kazuyoshi Yoshii, Masataka Goto |
IEEE Trans. Speech Audio Process. | 1 |
| 2009 | MusicCommentator: Generating Comments Synchronized with Musical Audio Signals by a Joint Probabilistic Model of Acoustic and Textual Features
Kazuyoshi Yoshii, Masataka Goto |
ICEC | 1 |
| 2008 | A robot listens to music and counts its beats aloud by separating music from counting voiceabstractThis paper presents a beat-counting robot that can count musical beats aloud, i.e., speak ldquoone, two, three, four, one, two, ...rdquo along music, while listening to music by using its own ears. Music-understanding robots that interact with humans should be able not only to recognize music internally, but also to express their own internal states. To develop our beat-counting robot, we have tackled three issues: (1) recognition of hierarchical beat structures, (2) expression of these structures by counting beats, and (3) suppression of counting voice (self-generated sound) in sound mixtures recorded by ears. The main issue is (3) because the interference of counting voice in music causes the decrease of the beat recognition accuracy. So we designed the architecture for music-understanding robot that is capable of dealing with the issue of self-generated sounds. To solve these issues, we took the following approaches: (1) beat structure prediction based on musical knowledge on chords and drums, (2) speed control of counting voice according to music tempo via a vocoder called STRAIGHT, and (3) semi-blind separation of sound mixtures into music and counting voice via an adaptive filter based on ICA (independent component analysis) that uses the waveform of the counting voice as a prior knowledge. Experimental result showed that suppressing robotpsilas own voice improved music recognition capability. Takeshi Mizumoto, Ryu Takeda, Kazuyoshi Yoshii, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 3 |
| 2008 | A robot uses its own microphone to synchronize its steps to musical beats while scatting and singingabstractMusical beat tracking is one of the effective technologies for human-robot interaction such as musical sessions. Since such interaction should be performed in various environments in a natural way, musical beat tracking for a robot should cope with noise sources such as environmental noise, its own motor noises, and self voices, by using its own microphone. This paper addresses a musical beat tracking robot which can step, scat and sing according to musical beats by using its own microphone. To realize such a robot, we propose a robust beat tracking method by introducing two key techniques, that is, spectro-temporal pattern matching and echo cancellation. The former realizes robust tempo estimation with a shorter window length, thus, it can quickly adapt to tempo changes. The latter is effective to cancel self noises such as stepping, scatting, and singing. We implemented the proposed beat tracking method for Honda ASIMO. Experimental results showed ten times faster adaptation to tempo changes and high robustness in beat tracking for stepping, scatting and singing noises. We also demonstrated the robot times its steps while scatting or singing to musical beats. Kazumasa Murata, Kazuhiro Nakadai, Kazuyoshi Yoshii, Ryu Takeda, Toyotaka Torii, Hiroshi G. Okuno, Yuji Hasegawa, Hiroshi Tsujino |
IROS | 3 |
| 2008 | An Efficient Hybrid Music Recommender System Using an Incrementally Trainable Probabilistic Generative ModelabstractThis paper presents a hybrid music recommender system that ranks musical pieces while efficiently maintaining collaborative and content-based data, i.e., rating scores given by users and acoustic features of audio signals. This hybrid approach overcomes the conventional tradeoff between recommendation accuracy and variety of recommended artists. Collaborative filtering, which is used on e-commerce sites, cannot recommend nonbrated pieces and provides a narrow variety of artists. Content-based filtering does not have satisfactory accuracy because it is based on the heuristics that the user's favorite pieces will have similar musical content despite there being exceptions. To attain a higher recommendation accuracy along with a wider variety of artists, we use a probabilistic generative model that unifies the collaborative and content-based data in a principled way. This model can explain the generative mechanism of the observed data in the probability theory. The probability distribution over users, pieces, and features is decomposed into three conditionally independent ones by introducing latent variables. This decomposition enables us to efficiently and incrementally adapt the model for increasing numbers of users and rating scores. We evaluated our system by using audio signals of commercial CDs and their corresponding rating scores obtained from an e-commerce site. The results revealed that our system accurately recommended pieces including nonrated ones from a wide variety of artists and maintained a high degree of accuracy even when new users and rating scores were added. Kazuyoshi Yoshii, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IEEE Trans. Speech Audio Process. | 1 |
| 2007 | A biped robot that keeps steps in time with musical beats while listening to music with its own earsabstractWe aim at enabling a biped robot to interact with humans through real-world music in daily-life environments, e.g., to autonomously keep its steps (stamps) in time with musical beats. To achieve this, the robot should be able to robustly predict the beat times in real time while listening to musical performance with its own ears (head-embedded microphones). However, this has not previously been addressed in most studies on music-synchronized robots due to the difficulty in predicting the beat times in real-world music. To solve this problem, we implemented a beat-tracking method developed in the field of music information processing. The predicted beat times are then used by a feedback-control method that adjusts the robot's step intervals to synchronize its steps in time with the beats. The experimental results show that the robot can adjust its steps in time with the beat times as the tempo changes. The resulting robot needed about 25 [s] to recognize the tempo change after it and then synchronize its steps. Kazuyoshi Yoshii, Kazuhiro Nakadai, Toyotaka Torii, Yuji Hasegawa, Hiroshi Tsujino, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 1 |
| 2007 | Drum Sound Recognition for Polyphonic Audio Signals by Adaptation and Matching of Spectrogram Templates With Harmonic Structure SuppressionabstractThis paper describes a system that detects onsets of the bass drum, snare drum, and hi-hat cymbals in polyphonic audio signals of popular songs. Our system is based on a template-matching method that uses power spectrograms of drum sounds as templates. This method calculates the distance between a template and each spectrogram segment extracted from a song spectrogram, using Goto's distance measure originally designed to detect the onsets in drums-only signals. However, there are two main problems. The first problem is that appropriate templates are unknown for each song. The second problem is that it is more difficult to detect drum-sound onsets in sound mixtures including various sounds other than drum sounds. To solve these problems, we propose template-adaptation and harmonic-structure-suppression methods. First of all, an initial template of each drum sound, called a seed template, is prepared. The former method adapts it to actual drum-sound spectrograms appearing in the song spectrogram. To make our system robust to the overlapping of harmonic sounds with drum sounds, the latter method suppresses harmonic components in the song spectrogram before the adaptation and matching. Experimental results with 70 popular songs showed that our template-adaptation and harmonic-structure-suppression methods improved the recognition accuracy and achieved 83%, 58%, and 46% in detecting onsets of the bass drum, snare drum, and hi-hat cymbals, respectively. Kazuyoshi Yoshii, Masataka Goto, Hiroshi G. Okuno |
IEEE Trans. Speech Audio Process. | 1 |
| 2006 | An Error Correction Framework Based on Drum Pattern Periodicity for Improving Drum Sound DetectionabstractThis paper presents a framework for correcting errors of automatic drum sound detection focusing on the periodicity of drum patterns. We define drum patterns as periodic structures found in onset sequences of bass and snare drum sounds. Our framework extracts periodic drum patterns from imperfect onset sequences of detected drum sounds (bottom-up processing) and corrects errors using the periodicity of the drum patterns (top-down processing). We implemented this framework on our drum-sound detection system. We first obtained onset sequences of the drum sounds with our system and extracted drum patterns. On the basis of our observation that the same drum patterns tend to be repeated, we detected time points which deviate from the periodicity as error candidates. Finally, we verified each error candidate to judge whether it is an actual onset or not. Experiments of drum sound detection for polyphonic audio signals of popular CD recordings showed that our correction framework improved the average detection accuracy from 77.4% to 80.7% Kazuyoshi Yoshii, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
ICASSP (5) | 1 |