Katsutoshi Itoyama

dblp:77/4845 · DBLP profile ↗
← Back
56ranked-venue papers
5as first author
12since 2021 · last 2024
0000-0002-7098-3896ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 35 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 4 first-author · 5 since 2021Systems, architecture and hardware · 11 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 since 2021
YearPublicationVenuePosition
2024 UAV-Enhanced Combination to Application: Comprehensive Analysis and Benchmarking of a Human Detection Dataset for Disaster Scenarios
Ragib Amin Nihal, Benjamin Yen 0001, Katsutoshi Itoyama, Kazuhiro Nakadai
ICPR (14)3
2024 Improving Noise Robustness of Automatic Speech Recognition Based on a Parallel Adapter Model with Near-Identity Initialization
Takahiro Osaki, Yui Sudo, Katsutoshi Itoyama, Kenji Nishida, Kazuhiro Nakadai
IEA/AIE3
2024 Improving Impressions of Response Delay in AI-based Spoken Dialogue Systems
abstract
This paper aims to mitigate impression loss due to response delays in spoken dialogue systems. The system addressed in this paper consists of connected ASR, NLP, and TTS AI services, and large response delays due to processing time and network latency in each module may give a bad impression to users. Therefore, we propose three techniques to deal with the inevitable delays: filler insertion, parallel speech processing, and prompt editing. To evaluate the effectiveness of the proposed method, measurements and subject experiments were conducted. The results showed that filler insertion had statistically significant improvements of 1.1 points on a Likert scale in the impression of response delay. Parallel speech processing reduced delay by 0.6 seconds. Prompt editing was also effective in reducing delay through response length suppression.
Shuhei Asaka, Katsutoshi Itoyama, Kazuhiro Nakadai
RO-MAN2
2024 SLAM-Based Joint Calibration of Multiple Asynchronous Microphone Arrays and Sound Source Localization
abstract
Robot audition systems with multiple microphone arrays have many applications in practice. However, the accurate calibration of multiple microphone arrays remains challenging because there are many unknown parameters to be identified, including the relative transforms (i.e., orientation and translation) and asynchronous factors (i.e., initial time offset and sampling clock difference) between microphone arrays. To tackle these challenges, in this article, we adopt batch simultaneous localization and mapping (SLAM) for joint calibration of multiple asynchronous microphone arrays and sound source localization. Using the Fisher information matrix (FIM) approach, we first conduct the observability analysis (i.e., parameter identifiability) of the abovementioned calibration problem and establish necessary/sufficient conditions under which the FIM and the Jacobian matrix have full column rank, which implies the identifiability of the unknown parameters. We also discover several scenarios where the unknown parameters are not uniquely identifiable. Subsequently, we propose an effective framework to initialize the unknown parameters, which is used as the initial guess in batch SLAM for multiple microphone array calibration, aiming to further enhance optimization accuracy and convergence. Extensive numerical simulations and real experiments have been conducted to verify the performance of the proposed method. The experimental results show that the proposed pipeline achieves higher accuracy with fast convergence in comparison to methods that use the noise-corrupted ground truth of the unknown parameters as the initial guess in the optimization and other existing frameworks.
Yuanzheng He, Daobilige Su, Katsutoshi Itoyama, Kazuhiro Nakadai, Junfeng Wu 0001, Shoudong Huang, Youfu Li 0001, He Kong 0001
IEEE Trans. Robotics4
2023 miniStreamer: Enhancing Small Conformer with Chunked-Context Masking for Streaming ASR Applications on the Edge
Haris Gulzar, Monikka Roslianna Busto, Takeharu Eda, Katsutoshi Itoyama, Kazuhiro Nakadai
INTERSPEECH4
2023 Improving Sign Language Understanding Introducing Label Smoothing
abstract
Sign language is one of the most important communication methods when considering equality, diversity, and inclusion. Sign language understanding implies understanding sign language using machines, and it involves mainly two functions; sign language recognition and sign language translation. To improve sign language understanding performance, this paper proposes to use label smoothing with CTC (Connectionist Temporal Classification) loss as training criteria for the sign language understanding neural network. Experimental results showed the effectiveness of the proposed method in both sign language recognition and translation.
Sihan Tan, Khan Nabeela Khanum, Katsutoshi Itoyama, Kazuhiro Nakadai
RO-MAN3
2022 Weakly-Supervised Neural Full-Rank Spatial Covariance Analysis for a Front-End System of Distant Speech Recognition
Yoshiaki Bando, Takahiro Aizawa, Katsutoshi Itoyama, Kazuhiro Nakadai
INTERSPEECH3
2022 Spotforming by NMF Using Multiple Microphone Arrays
abstract
Sound source separation is a method to extract a target sound source from a mixture of various sound sources and noises. One of the typical sound source separation methods is beamforming, which can separate sound sources by direction based on the phase difference between channels from the recorded signal of a microphone array, a multi-channel recording system. However, beamforming is a direction-based method and cannot separate multiple sources in the same direction. In this paper, we propose a method for separating sources in the same direction using multiple microphone arrays. The proposed method performs beamforming using multiple microphone arrays and extracts only the target sound source from the separated sound by the Non-negative Matrix Factorization (NMF), thus reducing the influence of other sources in the same direction. In this paper, to investigate the effectiveness of the proposed method, experiments were conducted assuming the presence of another sound source in the same direction from an arbitrary microphone array. The results show that the proposed method outperforms the delay-sum method in a simulation environment. In addition, experiments were conducted in a real environment to verify the effect of reverberation.
Yasuhiro Kagimoto, Katsutoshi Itoyama, Kenji Nishida, Kazuhiro Nakadai
IROS2
2022 Outdoor evaluation of sound source localization for drone groups using microphone arrays
abstract
For robot and drone auditions, microphone arrays have been used for estimating sound source directions and sound source locations. By using sound source localization techniques, for example, drones can detect people calling for help even if the target person is not visible. Most sound source localization methods are based on estimated sound source directions and triangulation. However, when it comes to situations using drones, severe drone noise distorts direction estimation results which could worsen the localization results badly due to the discreteness of direction estimation. In this perspective, the authors have proposed a sound source localization method that can omit outlying triangulation points, which could improve its localization performance. In this paper, an outdoor experiment has been held, and the proposed method is evaluated whether it can localize a sound source even if real drone noise is added to the recordings. Experiment results show that the proposed method can localize with 4.15 m of estimation error for a sound source up to 50 m away, suppress the impact of outliers, and use only plausible triangulation points.
Taiki Yamada, Katsutoshi Itoyama, Kenji Nishida, Kazuhiro Nakadai
IROS2
2021 Assessment of von Mises-Bernoulli Deep Neural Network in Sound Source Localization
Katsutoshi Itoyama, Yoshiya Morimoto, Shungo Masaki, Ryosuke Kojima, Kenji Nishida, Kazuhiro Nakadai
Interspeech1
2021 Detecting earthquakes: a novel deep learning-based approach for effective disaster response
Muhammad Shakeel 0001, Katsutoshi Itoyama, Kenji Nishida, Kazuhiro Nakadai
Appl. Intell.2
2021 Multichannel environmental sound segmentation
abstract
Abstract This paper proposes a multichannel environmental sound segmentation method. Environmental sound segmentation is an integrated method to achieve sound source localization, sound source separation and classification, simultaneously. When multiple microphones are available, spatial features can be used to improve the localization and separation accuracy of sounds from different directions; however, conventional methods have three drawbacks: (a) Sound source localization and sound source separation methods using spatial features and classification using spectral features trained in the same neural network, may overfit to the relationship between the direction of arrival and the class of a sound, thereby reducing their reliability to deal with novel events. (b) Although permutation invariant training used in autonomous speech recognition could be extended, it is impractical for environmental sounds that include an unlimited number of sound sources. (c) Various features, such as complex values of short time Fourier transform and interchannel phase differences have been used as spatial features, but no study has compared them. This paper proposes a multichannel environmental sound segmentation method comprising two discrete blocks, a sound source localization and separation block and a sound source separation and classification block. By separating the blocks, overfitting to the relationship between the direction of arrival and the class is avoided. Simulation experiments using created datasets including 75-class environmental sounds showed the root mean squared error of the proposed method was lower than that of conventional methods.
Yui Sudo, Katsutoshi Itoyama, Kenji Nishida, Kazuhiro Nakadai
Appl. Intell.2
2020 Calibration of a Microphone Array Based on a Probabilistic Model of Microphone Positions
Katsuhiro Dan, Katsutoshi Itoyama, Kenji Nishida, Kazuhiro Nakadai
IEA/AIE2
2020 Synchronization of Microphones Based on Rank Minimization of Warped Spectrum for Asynchronous Distributed Recording
abstract
This paper describes a new method for synchronizing microphones based on spectral warping in an asynchronous microphone array. In an audio signal observed by an asynchronous microphone array, two factors are involved: the time lag caused by a mismatch of the sampling rate and offset between microphones, and the modulation caused by differences in spatial transfer function between the sound source and each microphone. A spectrum warping matrix representing a resampling effect in the frequency domain is formulated and an observation model of audio (spectrum) mixture in an asynchronous microphone array is constructed. The proposed synchronization method uses an iterative optimization algorithm based on gradient descent of a new objective function. The function is formulated as a logarithmic determinant of a spectrum correlation matrix that is derived from relaxation of a rank minimization problem. Experimental results showed that the proposed method effectively estimates modulated sampling rate and that the proposed method outperforms an existing synchronization method.
Katsutoshi Itoyama, Kazuhiro Nakadai
IROS1
2020 Bayesian Singing Transcription Based on a Hierarchical Generative Model of Keys, Musical Notes, and F0 Trajectories
abstract
This article describes automatic singing transcription (AST) that estimates a human-readable musical score of a sung melody represented with quantized pitches and durations from a given music audio signal. To achieve the goal, we propose a statistical method for estimating the musical score by quantizing a trajectory of vocal fundamental frequencies (F0s) in the time and frequency directions. Since vocal F0 trajectories considerably deviate from the pitches and onset times of musical notes specified in musical scores, the local keys and rhythms of musical notes should be taken into account. In this article we propose a Bayesian hierarchical hidden semi-Markov model (HHSMM) that integrates a musical score model describing the local keys and rhythms of musical notes with an F0 trajectory model describing the temporal and frequency deviations of an F0 trajectory. Given an F0 trajectory, a sequence of musical notes, that of local keys, and the temporal and frequency deviations can be estimated jointly by using a Markov chain Monte Carlo (MCMC) method. We investigated the effect of each component of the proposed model and showed that the musical score model improves the performance of AST.
Ryo Nishikimi, Eita Nakamura, Masataka Goto, Katsutoshi Itoyama, Kazuyoshi Yoshii
IEEE ACM Trans. Audio Speech Lang. Process.4
2019 Joint Transcription of Lead, Bass, and Rhythm Guitars Based on a Factorial Hidden Semi-Markov Model
abstract
This paper describes a statistical method for estimating musical scores for lead, bass, and rhythm guitars from polyphonic audio signals of typical band-style music. To perform multi-instrument transcription involving multi-pitch detection and part assignment, it is crucial to formulate a musical language model that represents the characteristics of each part in order to solve the ambiguity of part assignment and estimate a musically-natural score. We propose a factorial hidden semi-Markov model that consists of three language models corresponding to the three guitar parts (three latent chains) and an acoustic model of a mixture spectrogram (emission model). The language model for rhythm guitar represents a homophonic sequence of musical notes (chord sequence) and those for lead and bass guitars represent a monophonic sequence of musical notes in a higher and lower frequency range respectively. The acoustic model represents a spectrogram as a sum of low-rank spectrograms of the three guitar parts approximated by NMF. Given a spectrogram, we estimate the note sequences using Gibbs sampling. We show that our model outperforms a state-of-the-art multi-pitch detection method in the accuracy and naturalness of the transcribed scores.
Kentaro Shibata, Ryo Nishikimi, Satoru Fukayama, Masataka Goto, Eita Nakamura, Katsutoshi Itoyama, Kazuyoshi Yoshii
ICASSP6
2019 Environmental sound segmentation utilizing Mask U-Net
abstract
This paper proposes an environmental sound segmentation method using Mask U-Net. Recent research in robot audition has analyzed noise reduction, section detection, and sound source separation for use in a real-world environment with many noises and overlaps. However, conventional methods apply respective functions in cascades. The biggest problem of cascade systems is the accumulation of errors generated at each function block. Although many methods of human voice separation have been proposed, robots operating in a real-world environment must be able to separate not only human voices but other environmental sounds. Unlike traditional sound source separation using spatial information, environmental sound segmentation must simultaneously detect sections and separate sound sources based on pre-trained features. One such method, U-Net, which was proposed for semantic segmentation of images, has been applied to the separation of singing voices. However, this method deals only with limited classes of sounds. The current study proposes an environmental sound segmentation method using Mask U-Net, which combines segmentation using U-Net with sound event detection using CNN to 75-classes of environmental sounds. Experimental application confirmed that this method improved learning speed and sound source separation compared with the conventional method.
Yui Sudo, Katsutoshi Itoyama, Kenji Nishida, Kazuhiro Nakadai
IROS2
2019 Unsupervised Speech Enhancement Based on Multichannel NMF-Informed Beamforming for Noise-Robust Automatic Speech Recognition
abstract
This paper describes multichannel speech enhancement for improving automatic speech recognition (ASR) in noisy environments. Recently, the minimum variance distortionless response (MVDR) beamforming has widely been used because it works well if the steering vector of speech and the spatial covariance matrix (SCM) of noise are given. To estimating such spatial information, conventional studies take a supervised approach that classifies each time-frequency (TF) bin into noise or speech by training a deep neural network (DNN). The performance of ASR, however, is degraded in an unknown noisy environment. To solve this problem, we take an unsupervised approach that decomposes each TF bin into the sum of speech and noise by using multichannel nonnegative matrix factorization (MNMF). This enables us to accurately estimate the SCMs of speech and noise not from observed noisy mixtures but from separated speech and noise components. In this paper, we propose online MVDR beamforming by effectively initializing and incrementally updating the parameters of MNMF. Another main contribution is to comprehensively investigate the performances of ASR obtained by various types of spatial filters, i.e., time-invariant and variant versions of MVDR beamformers and those of rank-1 and full-rank multichannel Wiener filters, in combination with MNMF. The experimental results showed that the proposed method outperformed the state-of-the-art DNN-based beamforming method in unknown environments that did not match training data.
Kazuki Shimada, Yoshiaki Bando, Masato Mimura, Katsutoshi Itoyama, Kazuyoshi Yoshii, Tatsuya Kawahara
IEEE ACM Trans. Audio Speech Lang. Process.4
2018 Statistical Speech Enhancement Based on Probabilistic Integration of Variational Autoencoder and Non-Negative Matrix Factorization
abstract
This paper presents a statistical method of single-channel speech enhancement that uses a variational autoencoder (VAE) as a prior distribution on clean speech. A standard approach to speech enhancement is to train a deep neural network (DNN) to take noisy speech as input and output clean speech. Although this supervised approach requires a very large amount of pair data for training, it is not robust against unknown environments. Another approach is to use non-negative matrix factorization (NMF) based on basis spectra trained on clean speech in advance and those adapted to noise on the fly. This semi-supervised approach, however, causes considerable signal distortion in enhanced speech due to the unrealistic assumption that speech spectrograms are linear combinations of the basis spectra. Replacing the poor linear generative model of clean speech in NMF with a VAE-a powerful nonlinear deep generative model-trained on clean speech, we formulate a unified probabilistic generative model of noisy speech. Given noisy speech as observed data, we can sample clean speech from its posterior distribution. The proposed method outperformed the conventional DNN-based method in unseen noisy environments.
Yoshiaki Bando, Masato Mimura, Katsutoshi Itoyama, Kazuyoshi Yoshii, Tatsuya Kawahara
ICASSP3
2018 Unsupervised Beamforming Based on Multichannel Nonnegative Matrix Factorization for Noisy Speech Recognition
abstract
This paper presents unsupervised multichannel speech enhancement for noisy speech recognition. Time-frequency (TF) mask estimation has actively been studied for estimating the steering vectors and spatial covariance matrices of speech and noise used for beamforming. The state-of-the-art approach to mask estimation is to use deep neural networks (DNN s) for classifying the TF bins of observed signals into speech and noise. Such a supervised approach, however, does not work well in an unknown environment. To accurately estimate the spatial covariance matrices in an unsupervised manner, we perform blind source separation (BSS) based on multichannel nonnegative matrix factorization (MNMF) for decomposing each TF bin into the components of speech and the other sources (noise). To clarify a suitable type of beamforming for MNMF, we tested both time- invariant and time-varying versions of the minimum variance distortionless response (MVDR) beamforming in addition to standard multichannel Wiener filtering (MWF). The experimental results showed that our MNMF-based beamforming approach outperformed the state-of-the-art DNN-based beamforming method in unknown environments that do not match the training data.
Kazuki Shimada, Yoshiaki Bando, Masato Mimura, Katsutoshi Itoyama, Kazuyoshi Yoshii, Tatsuya Kawahara
ICASSP4
2018 Signal Restoration based on Bi-directional LSTM with Spectral Filtering for Robot Audition
abstract
This paper addresses restoration of acoustic signals for robot audition. A robot usually listens to target acoustic signals such as speech and music in noisy conditions. Acoustic information on such signals inevitably contaminated with noise. Even when noise reduction techniques such as sound source separation are performed, the noise-reduced acoustic signals contain distortion and/or residual noise after the noise reduction to some extent. The distortion and residual noise basically degrade the performance of recognition processes such as automatic speech recognition (ASR). We decided to use bidirectional long short-term memory (Bi-LSTM) for acoustic signal restoration since it can represent dynamic behaviors well for a temporal sequence in the forward and backward directions. When applying Bi-LSTM to recover acoustic signals, there is an issue, that is, acoustic signals tend to be sparse in high frequencies, and thus Bi-LSTM training becomes insufficient in such high frequencies due to a lack of training data. Therefore, we propose a new restoration method based on Bi-LSTM with spectral filtering. The spectral filter and the corresponding inverse filter are introduced to a Bi-LSTM framework to accelerate training in high frequencies. Preliminary results showed that the proposed Bi-LSTM with spectral filtering can perform signal restoration even when a small amount of training data is available.
Ryosuke Taniguchi, Kotaro Hoshiba, Katsutoshi Itoyama, Kenji Nishida, Kazuhiro Nakadai
RO-MAN3
2018 Speech Enhancement Based on Bayesian Low-Rank and Sparse Decomposition of Multichannel Magnitude Spectrograms
abstract
This paper presents a blind multichannel speech enhancement method that can deal with the time-varying layout of microphones and sound sources. Since nonnegative tensor factorization (NTF) separates a multichannel magnitude (or power) spectrogram into source spectrograms without phase information, it is robust against the time-varying mixing system. This method, however, requires prior information such as the spectral bases (templates) of each source spectrogram in advance. To solve this problem, we develop a Bayesian model called robust NTF (Bayesian RNTF) that decomposes a multichannel magnitude spectrogram into target speech and noise spectrograms based on their sparseness and low rankness. Bayesian RNTF is applied to the challenging task of speech enhancement for a microphone array distributed on a hose-shaped rescue robot. When the robot searches for victims under collapsed buildings, the layout of the microphones changes over time and some of them often fail to capture target speech. Our method robustly works under such situations, thanks to its characteristic of time-varying mixing system. Experiments using a 3-m hose-shaped rescue robot with eight microphones show that the proposed method outperforms conventional blind methods in enhancement performance by the signal-to-noise ratio of 1.03 dB.
Yoshiaki Bando, Katsutoshi Itoyama, Masashi Konyo, Satoshi Tadokoro, Kazuhiro Nakadai, Kazuyoshi Yoshii, Tatsuya Kawahara, Hiroshi G. Okuno
IEEE ACM Trans. Audio Speech Lang. Process.2
2018 Bayesian Multichannel Audio Source Separation Based on Integrated Source and Spatial Models
abstract
This paper presents new statistical methods of multichannel audio source separation based on unified source and spatial models that, respectively, represent the generative process of latent source spectrograms and that of observed mixture spectrograms. One possibility of the source model is a factor model based on nonnegative matrix factorization that represents each time-frequency (TF) bin as the weighted sum of basis spectra. Another possibility is a mixture model inspired by latent Dirichlet allocation that exclusively classifies each TF bin into one of basis spectra. Similarly, the spatial model can either be a factor model that represents each TF bin as the weighted sum of source spectra or a mixture model that classifies each bin into one of those spectra. To unify these models in a principled manner and incorporate prior knowledge of a microphone array, we propose hierarchical Bayesian models of all the source-spatial combinations (factor-factor, mixture-factor, factor-mixture, and mixture-mixture models) and derive efficient Gibbs sampling algorithms for posterior inference. Experimental results showed that the proposed unified models outperformed the state-of-the-art method using only the spatial mixture model. Among the four unified models, the spatial factor model tended to work better than the spatial mixture model in exchange for larger computational cost, and the choice of source models had a little impact on the performance and computational cost.
Kousuke Itakura, Yoshiaki Bando, Eita Nakamura, Katsutoshi Itoyama, Kazuyoshi Yoshii, Tatsuya Kawahara
IEEE ACM Trans. Audio Speech Lang. Process.4
2017 Bayesian multichannel nonnegative matrix factorization for audio source separation and localization
abstract
This paper presents a Bayesian extension of multichannel nonnegative matrix factorization (MNMF) that decomposes the complex spectrograms of mixture signals recorded by a microphone array into basis spectra, their temporal activations, and the spatial correlation matrices of sources (directions) in the time-frequency-channel domain. Although the original MNMF can be used in a blind setting, prior knowledge of a microphone array is useful for improving source separation. The impulse response (spatial correlation matrix) of each direction can be measured in an anechoic room, however, it differs from that in a real environment where the microphone array is used. To solve this, we propose a unified Bayesian model of source separation and localization by introducing a prior distribution determined by an anechoic spatial correlation matrix on a real spatial correlation matrix with respect to each direction. This enables us to adaptively estimate a real spatial correlation matrix and the direction of each source. Experimental results showed that our method outperformed the original MNMF and the state-of-the-art methods with prior knowledge in terms of signal-to-distortion ratio (SDR) even when the method was used in an unknown environment with acoustic characteristics different from those of the anechoic room.
Kousuke Itakura, Yoshiaki Bando, Eita Nakamura, Katsutoshi Itoyama, Kazuyoshi Yoshii, Tatsuya Kawahara
ICASSP4
2016 Student's T nonnegative matrix factorization and positive semidefinite tensor factorization for single-channel audio source separation
abstract
This paper presents a robust variant of nonnegative matrix factorization (NMF) based on complex Student's t distributions (t-NMF) for source separation of single-channel audio signals. The Itakura-Saito divergence NMF (Gaussian NMF) is justified for this purpose under an assumption that the complex spectra of source signals and those of the mixture signal are complex Gaussian distributed (the additiv-ity of power spectra holds). In fact, however, the source spectra are often heavy-tailed distributed. When the source spectra are complex Cauchy distributed, for example, the mixture spectra are also complex Cauchy distributed (the additivity of amplitude spectra holds). Using the complex t distribution that includes the complex Gaussian and Cauchy distributions as its special cases, we propose t-NMF as a unified extension of Gaussian NMF and Cauchy NMF. Furthermore, we propose the corresponding variant of positive semidefinite tensor factorization based on multivariate complex t distributions (t-PSDTF). The experimental results showed that while t-NMF and t-PSDTF were comparative to Gaussian counterparts in terms of peak performance, they worked much better on average because they are insensitive to initialization and tend to avoid local optima.
Kazuyoshi Yoshii, Katsutoshi Itoyama, Masataka Goto
ICASSP2
2016 Online simultaneous localization and mapping of multiple sound sources and asynchronous microphone arrays
abstract
This paper presents an online method of simultaneous localization and mapping (SLAM) for estimating the positions of multiple moving sound sources and stationary robots and synchronizing microphone arrays attached to those robots. Since each robot with a microphone array can solely estimate the directions of sound sources, the two-dimensional source positions can be estimated from the source directions estimated by multiple robots using a triangulation method. In addition, sound mixtures can be separated accurately by regarding distributed microphone arrays as one big array. To perform these tasks, some methods have been proposed for localizing and synchronizing microphone arrays. These methods, however, can be used only if a single sound source exists because the time differences of arrival (TDOAs) between microphones are assumed to be directly observed. To overcome this limitation, we propose a unified state-space model that encodes the source and robot positions and the time offsets between microphone arrays in a latent space. Given the TDOAs and directions of arrival (DOAs) estimated by separating observed mixture sounds into source sounds, the latent variables are estimated jointly in an online manner using a FastSLAM2.0 algorithm that can deal with an unknown time-varying number of moving sound sources.
Kouhei Sekiguchi, Yoshiaki Bando, Keisuke Nakamura, Kazuhiro Nakadai, Katsutoshi Itoyama, Kazuyoshi Yoshii
IROS5
2016 Parallel Speech Corpora of Japanese Dialects
Koichiro Yoshino, Naoki Hirayama, Shinsuke Mori, Fumihiko Takahashi, Katsutoshi Itoyama, Hiroshi G. Okuno
LREC5
2016 Singing Voice Separation and Vocal F0 Estimation Based on Mutual Combination of Robust Principal Component Analysis and Subharmonic Summation
abstract
This paper presents a new method of singing voice analysis that performs mutually-dependent singing voice separation and vocal fundamental frequency (F0) estimation. Vocal F0 estimation is considered to become easier if singing voices can be separated from a music audio signal, and vocal F0 contours are useful for singing voice separation. This calls for an approach that improves the performance of each of these tasks by using the results of the other. The proposed method first performs robust principal component analysis (RPCA) for roughly extracting singing voices from a target music audio signal. The F0 contour of the main melody is then estimated from the separated singing voices by finding the optimal temporal path over an F0 saliency spectrogram. Finally, the singing voices are separated again more accurately by combining a conventional time-frequency mask given by RPCA with another mask that passes only the harmonic structures of the estimated F0s. Experimental results showed that the proposed method significantly improved the performances of both singing voice separation and vocal F0 estimation. The proposed method also outperformed all the other methods of singing voice separation submitted to an international music analysis competition called MIREX 2014.
Yukara Ikemiya, Katsutoshi Itoyama, Kazuyoshi Yoshii
IEEE ACM Trans. Audio Speech Lang. Process.2
2015 Challenges in deploying a microphone array to localize and separate sound sources in real auditory scenes
abstract
Analyzing the auditory scene of real environments is challenging partly because an unknown number and type of sound sources are observed at the same time and partly because these sounds are observed on a significantly different sound pressure level at the microphone. These are difficult problems even with state-of-the-art sound source localization and separation methods. In this paper, we exploit two such methods using a microphone array: (1) Bayesian nonparametric microphone array processing (BNP-MAP), which is capable of separating and localizing sound sources when the number of sound sources is unspecified, and (2) robot audition software “HARK” is capable of separating and localizing in real time. Through experimentation, we found that BNP-MAP is more robust against differences in the sound pressure levels of the source signals and in the spatial closeness of source positions. Experiments analyzing real scenes of human conversations recorded in a big exhibition hall and bird calling recorded at a natural park demonstrate the efficacy and applicability of BNP-MAP.
Yoshiaki Bando, Takuma Otsuka, Katsutoshi Itoyama, Kazuyoshi Yoshii, Yoko Sasaki, Satoshi Kagami, Hiroshi G. Okuno
ICASSP3
2015 Singing voice analysis and editing based on mutually dependent F0 estimation and source separation
abstract
This paper presents a novel framework that improves both vocal fundamental frequency (F0) estimation and singing voice separation by making effective use of the mutual dependency of those two tasks. A typical approach to singing voice separation is to estimate the vocal F0 contour from a target music signal and then extract the singing voice by using a time-frequency mask that passes only the harmonic components of the vocal F0s and overtones. Vocal F0 estimation, on the contrary, is considered to become easier if only the singing voice can be extracted accurately from the target signal. Such mutual dependency has scarcely been focused on in most conventional studies. To overcome this limitation, our framework alternates those two tasks while using the results of each in the other. More specifically, we first extract the singing voice by using robust principal component analysis (RPCA). The F0 contour is then estimated from the separated singing voice by finding the optimal path over a F0-saliency spectrogram based on subharmonic summation (SHS). This enables us to improve singing voice separation by combining a time-frequency mask based on RPCA with a mask based on harmonic structures. Experimental results obtained when we used the proposed technique to directly edit vocal F0s in popular-music audio signals showed that it significantly improved both vocal F0 estimation and singing voice separation.
Yukara Ikemiya, Kazuyoshi Yoshii, Katsutoshi Itoyama
ICASSP3
2015 A feedback framework for improved chord recognition based on NMF-based approximate note transcription
abstract
This paper presents a feedback framework that can improve chord recognition for music audio signals by performing approximate note transcription with Bayesian non-negative matrix factorization (NMF) using prior knowledge on chords. Although the names and note compositions of chords are intrinsically linked with each other (e.g., C major chords are highly likely to include C, E, and G notes, and those notes are highly likely to be in C major chords), chord recognition and note transcription (multipitch analysis) have been studied independently. To solve this chicken-and-egg problem, our framework iterates chord recognition and approximate note transcription using each other's results. More specifically, we first perform approximate note transcription based on Bayesian NMF that forces basis spectra to respectively correspond to different semitone-level pitches covering the whole range. We then execute chord recognition based on Bayesian hidden Markov models (HMMs) that use chroma features obtained from the activation patterns of those pitches. To improve note transcription, we again perform Bayesian NMF that encourages certain kinds of pitches in each chord region to be activated. Experimental results showed that our feedback framework gradually improved the accuracy of chord recognition.
Satoshi Maruo, Kazuyoshi Yoshii, Katsutoshi Itoyama, Matthias Mauch, Masataka Goto
ICASSP3
2015 Bayesian integration of sound source separation and speech recognition: a new approach to simultaneous speech recognition
abstract
This paper presents a novel Bayesian method that can directly recognize overlapping utterances without explicitly separating mixture signals into their independent components in advance of speech recognition. The conventional approach to contami-nated speech recognition in real environments uniquely extracts the clean isolated signals of individual sources (e.g., by noise reduction, dereverberation, and source separation). One of the main limitations of this cascading approach is that the accuracy of speech recognition is upper bounded by the accuracy of pre-processing. To overcome this limitation, our method marginal-izes out uncertain isolated speech signals by integrating source separation and speech recognition in a Bayesian manner. A suf-ficient number of samples are drawn from the posterior distribu-tion of isolated speech signals by using a Markov chain Monte Carlo method, and then the posterior distributions of uttered texts for those samples are integrated. Under a certain con-dition, this Monte Carlo integration is shown to reduce to the well-known method called ROVER that integrates recognized texts obtained from sampled speech signals. Results of simulta-neous speech recognition experiments showed that in terms of word accuracy the proposed method significantly outperformed conventional cascading methods. Index Terms: simultaneous speech recognition, sound source separation, Bayesian modeling, MCMC, ROVER
Kousuke Itakura, Izaya Nishimuta, Yoshiaki Bando, Katsutoshi Itoyama, Kazuyoshi Yoshii
INTERSPEECH4
2015 Microphone-accelerometer based 3D posture estimation for a hose-shaped rescue robot
abstract
3D posture estimation for a hose-shaped robot is critical in rescue activities due to complex physical environments. Conventional sound-based posture estimation assumes rather flat physical environments and focuses only on 2D, resulting in poor performance in real world environments with rubble. This paper presents novel 3D posture estimation by exploiting microphones and accelerometers. The idea of our method is to compensate the lack of posture information obtained by sound-based time-difference-of arrival (TDOA) with the tilt information obtained from accelerometers. This compensation is formulated as a nonlinear state-space model and solved by the unscented Kalman filter. Experiments are conducted by using a 3m hose-shaped robot with eight units of a microphone and an accelerometer and seven units of a loudspeaker and a vibration motor deployed in a simple 3D structure. Experimental results demonstrate that our method reduces the errors of initial states to about 20 cm in the 3D space. If the initial errors of initial states are less than 20 %, our method can estimate the correct 3D posture in real-time.
Yoshiaki Bando, Katsutoshi Itoyama, Masashi Konyo, Satoshi Tadokoro, Kazuhiro Nakadai, Kazuyoshi Yoshii, Hiroshi G. Okuno
IROS2
2015 Audio-visual beat tracking based on a state-space model for a music robot dancing with humans
abstract
This paper presents an audio-visual beat-tracking method for an entertainment robot that can dance in synchronization with music and human dancers. Conventional music robots have focused on either music audio signals or dancing movements of humans for detecting and predicting beat times in real time. Since a robot needs to record music audio signals by using its own microphones, however, the signals are severely contaminated with loud environmental noise and reverberant sounds. Moreover, it is difficult to visually detect beat times from real complicated dancing movements that exhibit weaker repetitive characteristics than music audio signals do. To solve these problems, we propose a state-space model that integrates both audio and visual information in a probabilistic manner. At each frame, the method extracts acoustic features (audio tempos and onset likelihoods) from music audio signals and extracts skeleton features from movements of a human dancer. The current tempo and the next beat time are then estimated from those observed features by using a particle filter. Experimental results showed that the proposed multi-modal method using a depth sensor (Kinect) for extracting skeleton features outperformed conventional mono-modal methods by 0.20 (F measure) in terms of beat-tracking accuracy in a noisy and reverberant environment.
Misato Ohkita, Yoshiaki Bando, Yukara Ikemiya, Katsutoshi Itoyama, Kazuyoshi Yoshii
IROS4
2015 Optimizing the layout of multiple mobile robots for cooperative sound source separation
abstract
This paper presents a novel active audition method that enables multiple mobile robots to move to optimal positions for improving the performance of sound source separation. A main advantage of our distributed system is that each robot has its own microphone array and all mobile robots can collaborate on source separation by regarding a set of movable microphone arrays as a big reconfigurable array. To incrementally optimize the positions of the robots (the layout of the big microphone array) in an active-audition manner, it is necessary to predict the source separation performance from a possible layout of the next time step although true source signals are unknown. To solve this problem, our method simulates delay-and-sum beamforming from a possible layout for theoretically calculating the gain for each frequency component of a source signal in the corresponding separated signal. The robots are moved into a layout with the highest average gain over all sources and the whole frequency range. The experimental results showed that the harmonic mean of signal-to-distortion ratios (SDRs) was improved by 6.0 dB in simulations and by 5.7 dB in a real environment.
Kouhei Sekiguchi, Yoshiaki Bando, Katsutoshi Itoyama, Kazuyoshi Yoshii
IROS3
2015 Identification and Localization of One or Two Concurrent Speakers in a Binaural Robotic Context
abstract
This paper presents a method of identification and azimuth estimation for one or two concurrent speakers in simultaneous utterances. This method is applicable to human-machine interaction and robot audition. Identification and localization have been rarely mutually addressed and related works rely on time-frequency exploitation strategies to extract and treat each source's contribution to the received signal. The presented method relies on a training made with one speaker at a time, but it can exploit a speech segment to identify and localize two speakers. A cochlear filtering-based binaural front-end allows to extract equivalent rectangular bandwidth frequency cepstral coefficients (ERBFCC) and interaural level difference (ILD) features. Artificial neural networks (ANNs) exploit ERBFCCs to provide identity information, and a histogram-based exploitation of ILDs provides azimuth angle information. The method was evaluated in contexts including overlapping segments in the presence of noises and sound reflections and its efficiency was demonstrated. Even with fully overlapping utterances, we reached an 83% identification rate of both speakers, an 82% estimation accuracy of both azimuths and an 68% correct mutual identity and azimuth estimation rate. At least one speaker was correctly identified and localized in more than 99% of the tests for utterances lasting near 5s.
Karim Youssef, Katsutoshi Itoyama, Kazuyoshi Yoshii
SMC2
2015 Automatic Speech Recognition for Mixed Dialect Utterances by Mixing Dialect Language Models
abstract
This paper presents an automatic speech recognition (ASR) system that accepts a mixture of various kinds of dialects. The system recognizes dialect utterances on the basis of the statistical simulation of vocabulary transformation and combinations of several dialect models. Previous dialect ASR systems were based on handcrafted dictionaries for several dialects, which involved costly processes. The proposed system statistically trains transformation rules between a common language and dialects, and simulates a dialect corpus for ASR on the basis of a machine translation technique. The rules are trained with small sets of parallel corpora to make up for the lack of linguistic resources on dialects. The proposed system also accepts mixed dialect utterances that contain a variety of vocabularies. In fact, spoken language is not a single dialect but a mixed dialect that is affected by the circumstances of speakers’ backgrounds (e.g., native dialects of their parents or where they live). We addressed two methods to combine several dialects appropriately for each speaker. The first was recognition with language models of mixed dialects with automatically estimated weights that maximized the recognition likelihood. This method performed the best, but calculation was very expensive because it conducted grid searches of combinations of dialect mixing proportions that maximized the recognition likelihood. The second was integration of results of recognition from each single dialect language model. The improvements with this model were slightly smaller than those with the first method. Its calculation cost was, however, inexpensive and it worked in real-time on general workstations. Both methods achieved higher recognition accuracies for all speakers than those with the single dialect models and the common language model, and we could choose a suitable model for use in ASR that took into consideration the computational costs and recognition accuracies.
Naoki Hirayama, Koichiro Yoshino, Katsutoshi Itoyama, Shinsuke Mori, Hiroshi G. Okuno
IEEE ACM Trans. Audio Speech Lang. Process.3
2014 Transcribing vocal expression from polyphonic music
abstract
A method for transcribing vocal expressions such as vibrato, glissando, and kobushi separately from polyphonic music is described. The expressions appear as fluctuation in the fundamental frequency contour of the singing voice. They can be used for search and retrieval of music and for expressive singing voice synthesis based on singing style since they strongly reflect the individuality of the singer. The fundamental frequency contour of the singing voice is estimated using the Viterbi algorithm with limitation from a corresponding note sequence. Next, the notes are aligned with the fundamental frequency sequence temporally. Finally, each expression is identified and parameterized in accordance with designed rules. Experiments demonstrated that this method can transcribe expressions in the singing voice from commercial recordings.
Yukara Ikemiya, Katsutoshi Itoyama, Hiroshi G. Okuno
ICASSP2
2014 Automatic transcription of guitar tablature from audio signals in accordance with player's proficiency
abstract
We describe a method for automatically transcribing guitar tablatures from audio signals in accordance with the player's proficiency for use as support for a guitar player's practice. The system estimates the multiple pitches in each time frame and the optimal fingering considering playability and player's proficiency. It combines a conventional multipitch estimation method with a basic dynamic programming method. The difficulty of the fingerings can be changed by tuning the parameter representing the relative weights of the acoustical reproducibility and the fingering easiness. Experiments conducted using synthesized guitar audio signals to evaluate the transcribed tablatures in terms of the multipitch estimation accuracy and fingering easiness demonstrated that the system can simplify the fingering with higher precision of multipitch estimation results than the conventional method.
Kazuki Yazawa, Katsutoshi Itoyama, Hiroshi G. Okuno
ICASSP2
2014 Transferring Vocal Expression of F0 Contour Using Singing Voice Synthesizer
Yukara Ikemiya, Katsutoshi Itoyama, Hiroshi G. Okuno
IEA/AIE (2)2
2014 Visualization of auditory awareness based on sound source positions estimated by depth sensor and microphone array
abstract
We have developed a system for visualizing auditory awareness on the basis of sound source locations estimated using a depth sensor and microphone array. Previous studies on visualizing the acoustic environment viewed the level of sound pressures directly on the captured image, so the visualization was often based on a mixture of several sound sources. As a result, which targets to focus on was not intuitive. To help users selectively to find the targets and focus on the target analysis, we should extract the captured acoustic information and selectively propose it with the user demand. We have designed a three-layer visualization model for auditory awareness consisting of a sound source distribution layer, a sound location layer, and a sound saliency layer. The model extracts acoustic information by using the depth image and multi-directional sound sources captured with a depth sensor and microphone array. This model is used in the system we developed for visualizing auditory awareness.
Takahiro Iyama, Osamu Sugiyama, Takuma Otsuka, Katsutoshi Itoyama, Hiroshi G. Okuno
IROS4
2014 Sound annotation tool for multidirectional sounds based on spatial information extracted by HARK robot audition software
abstract
With the rise of inexpensive microphone array products and the robot audition software called HARK, we can record and analyze multidirectional sound sources easily. The combination of microphone array and the software enables us to separate, localize, and track multidirectional sound sources. Most of the solutions for accessing these separated sound source information provide clients for interpreting simplified information about the separated sources, but not to directly execute the semantic annotations. Since the multidirectional sound annotation requires simultaneous labeling of separated sound sources and a multidirectional overview of the sources, it is essential to have an efficient way of annotation and an intuitive view of multidirectional sounds. Our proposed sound annotation tool provides drag & drop operation of annotation with a 3D sound source view and also provides annotation autocompletion with a SVM trained with the user's annotation history. The proposed features enable users to do the annotation task intuitively and confirm its result. We also conducted an evaluation demonstrating the efficiency of annotation done using the tool.
Osamu Sugiyama, Katsutoshi Itoyama, Kazuhiro Nakadai, Hiroshi G. Okuno
SMC2
2014 Nonparametric Bayesian dereverberation of power spectrograms based on infinite-order autoregressive processes
abstract
This paper describes a monaural audio dereverberation method that operates in the power spectrogram domain. The method is robust to different kinds of source signals such as speech or music. Moreover, it requires little manual intervention, including the complexity of room acoustics. The method is based on a non-conjugate Bayesian model of the power spectrogram. It extends the idea of multi-channel linear prediction to the power spectrogram domain, and formulates a model of reverberation as a non-negative, infinite-order autoregressive process. To this end, the power spectrogram is interpreted as a histogram count data, which allows a nonparametric Bayesian model to be used as the prior for the autoregressive process, allowing the effective number of active components to grow, without bound, with the complexity of data. In order to determine the marginal posterior distribution, a convergent algorithm, inspired by the variational Bayes method, is formulated. It employs the minorization-maximization technique to arrive at an iterative, convergent algorithm that approximates the marginal posterior distribution. Both objective and subjective evaluations show advantage over other methods based on the power spectrum. We also apply the method to a music information retrieval task and demonstrate its effectiveness.
Akira Maezawa, Katsutoshi Itoyama, Kazuyoshi Yoshii, Hiroshi G. Okuno
IEEE ACM Trans. Audio Speech Lang. Process.2
2013 Multiple index combination for Japanese spoken term detection with optimum index selection based on OOV-region classifier
Naoyuki Kanda, Katsutoshi Itoyama, Hiroshi G. Okuno
ICASSP2
2013 Initialization-robust Bayesian multipitch analyzer based on psychoacoustical and musical criteria
abstract
We present a new Bayesian multipitch analyzer that dispenses with a precise optimization of parameter initialization or hyperparameters. Our method uses a new family of prior distribution, characteristic prior; it efficiently restricts the existence region of the latent variables, that is, the product of a conjugate prior and a characteristic function. The update formulas become a simple form that is actually suitable for Gibbs sampling. We construct characteristic priors of harmonic structures based on psychoacoustical and musical knowledge and apply them to nonnegative harmonic factorization. Experimental results improve 5.2 points in F-measure under a tough condition, random initialization with no hyperparameter optimization.
Daichi Sakaue, Takuma Otsuka, Katsutoshi Itoyama, Hiroshi G. Okuno
ICASSP3
2013 Audio-based guitar tablature transcription using multipitch analysis and playability constraints
abstract
This paper proposes a method of guitar tablature transcription from audio signals. Multipitch estimation and fingering configuration estimation are essential for transcribing tablatures. Conventional multipitch estimation methods, including latent harmonic allocation (LHA), often estimate combinations of pitches that people cannot play due to inherent physical constraints. Unplayable combinations of pitches are eliminated by filtering the results of LHA with three constraints. We first enumerate playable fingering configurations, and use them to suppress any undesirable combination of pitches. The optimal fingering configuration in each time frame is optimized to satisfy the need for temporal continuity by using dynamic programming. We use synthesized guitar sounds from MIDI data (ground truth) for evaluation. Experiments with them demonstrate the improvement of multipitch estimation by 5.9 points on average in F-measure and the transcribed tablatures are playable.
Kazuki Yazawa, Daichi Sakaue, Kohei Nagira, Katsutoshi Itoyama, Hiroshi G. Okuno
ICASSP4
2013 Automatic estimation of dialect mixing ratio for dialect speech recognition
abstract
This paper proposes methods for determining an appropriate mixing ratio of dialects in automatic speech recognition (ASR) for dialects. To handle ASR for various dialects, it has been reported to be effective to train a language model using a dialectmixed corpus. One reason behind this is geographical continuity of spoken dialect; we regard spoken dialect as a mixture of various dialects. This mixing ratio changes at every moment as well as depends on a speaker. We can improve recognition accuracybygivingan appropriatedialectmixingratio foraspeaker’s dialect. The mixing ratio is generally unknown and requires to be estimated and updated referring to input utterances. We handle two methods for updating it based on recognition results; one is to compute contribution of dialects for each recognized word, and the other is to predict mixture information referring to a whole recognized sentence based on topic modeling. The experimental result shows that the mixing ratio estimated by these methods realized higher recognition accuracy than a fixed mixing ratio. Index Terms: dialect, supervised latent Dirichlet allocation (sLDA), mixing ratio.
Naoki Hirayama, Koichiro Yoshino, Katsutoshi Itoyama, Shinsuke Mori, Hiroshi G. Okuno
INTERSPEECH3
2013 Posture estimation of hose-shaped robot using microphone array localization
abstract
This paper presents a posture estimation of hose-shaped robot using microphone array localization. The hose-shaped robots, one of major rescue robots, have problems with navigation because their posture is too flexible for a remote operator to control to go as far as desired. For navigational and mission usability, the posture estimation of the hose-shaped robot is essential. We developed a posture estimation method with a microphone array and small loudspeakers equipped on the hose-shaped robot. Our method consists of two steps: (1) playing a known sound from the loudspeaker one-by-one, and (2) estimating the microphone positions on the hose-shaped robot instead of estimating the posture directly. We designed a time difference of arrival (TDOA) estimation method to be robust against directional noise and implemented a prototype system using a posture model of the hose-shaped robot and an Extended Kalman Filter (EKF). The validity of our approach is evaluated by the experiments with both signals recorded in an anechoic chamber and simulated data.
Yoshiaki Bando, Takeshi Mizumoto, Katsutoshi Itoyama, Kazuhiro Nakadai, Hiroshi G. Okuno
IROS3
2013 Noise correlation matrix estimation for improving sound source localization by multirotor UAV
abstract
A method has been developed for improving sound source localization (SSL) using a microphone array from an unmanned aerial vehicle with multiple rotors, a “multirotor UAV”. One of the main problems in SSL from a multirotor UAV is that the ego noise of the rotors on the UAV interferes with the audio observation and degrades the SSL performance. We employ a generalized eigenvalue decomposition-based multiple signal classification (GEVD-MUSIC) algorithm to reduce the effect of ego noise. While GEVD-MUSIC algorithm requires a noise correlation matrix corresponding to the auto-correlation of the multichannel observation of the rotor noise, the noise correlation is nonstationary due to the aerodynamic control of the UAV. Therefore, we need an adaptive estimation method of the noise correlation matrix for a robust SSL using GEVD-MUSIC algorithm. Our method uses a Gaussian process regression to estimate the noise correlation matrix in each time period from the measurements of self-monitoring sensors attached to the UAV such as the pitch-roll-yaw tilt angles, xyz speeds, and motor control values. Experiments compare our method with existing SSL methods in terms of precision and recall rates of SSL. The results demonstrate that our method outperforms existing methods, especially under high signal-to-noise-ratio conditions.
Koutarou Furukawa, Keita Okutani, Kohei Nagira, Takuma Otsuka, Katsutoshi Itoyama, Kazuhiro Nakadai, Hiroshi G. Okuno
IROS5
2012 Initialization-robust multipitch estimation based on latent harmonic allocation using overtone corpus
abstract
We present a new method for modeling the overtone structures of musical instruments that uses an overtone corpus generated using a MIDI synthesizer. Since multipitch estimation requires a joint estimation of F0's and their overtone structures, one of the most important problems is the overtone structure modeling. Latent harmonic allocation (LHA), a promising multipitch estimation method, is difficult to use for various applications because it requires appropriate prior distributions of the overtone structures, which cannot be determined from statistical evidence. Our method uses an overtone corpus to avoid the problem of setting prior distributions and instead restricts the lower and upper bounds of each overtone weight. The bounds are determined from reference signals generated by a MIDI synthesizer. Experimental results demonstrated that the overtone structures were stably and accurately estimated for a wide variety of initial settings.
Daichi Sakaue, Katsutoshi Itoyama, Tetsuya Ogata, Hiroshi G. Okuno
ICASSP2
2012 Automatic Chord Recognition Based on Probabilistic Integration of Acoustic Features, Bass Sounds, and Chord Transition
Katsutoshi Itoyama, Tetsuya Ogata, Hiroshi G. Okuno
IEA/AIE1
2011 Simultaneous processing of sound source separation and musical instrument identification using Bayesian spectral modeling
abstract
This paper presents a method of both separating audio mixtures into sound sources and identifying the musical instruments of the sources. A statistical tone model of the power spectrogram, called an integrated model, is defined and source separation and instrument identification are carried out on the basis of Bayesian inference. Since, the parameter distributions of the integrated model depend on each instrument, the instrument name is identified by selecting the one that has the maximum relative instrument weight. Experimental results showed correct instrument identification enables precise source separation even when many overtones overlap.
Katsutoshi Itoyama, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
ICASSP1
2010 Violin Fingering Estimation Based on Violin Pedagogical Fingering Model Constrained by Bowed Sequence Estimation from Audio Input
Akira Maezawa, Katsutoshi Itoyama, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IEA/AIE (3)2
2009 Bowed String Sequence Estimation of a Violin Based on Adaptive Audio Signal Classification and Context-Dependent Error Correction
abstract
The sequence of strings played on a bowed string instrument is essential to understanding of the fingering. Thus, its estimation is required for machine understanding of violin playing. Audio-based identification is the only viable way to realize this goal for existing music recordings. A naive implementation using audio classification alone, however, is inaccurate and is not robust against variations in string or instruments. We develop a bowed string sequence estimation method by combining audio-based bowed string classification and context-dependent error correction. The robustness against different setups of instruments improves by normalizing the F0-dependent features using the average feature of a recording. The performance of error correction is evaluated using an electric violin with two different brands of strings and an acoustic violin. By incorporating mean normalization, the recognition error of recognition accuracy due to changing the string alleviates by 8 points, and that due to change of instrument by 12 points. Error correction decreases the error due to change of string by 8 points and that due to different instrument by 9 points.
Akira Maezawa, Katsutoshi Itoyama, Toru Takahashi 0001, Tetsuya Ogata, Hiroshi G. Okuno
ISM2
2009 Changing timbre and phrase in existing musical performances as you like: manipulations of single part using harmonic and inharmonic models
abstract
This paper presents a new music manipulation method that can change the timbre and phrases of an existing instrumental performance in a polyphonic sound mixture. This method consists of three primitive functions: 1) extracting and analyzing of a single instrumental part from polyphonic music signals, 2) mixing the instrument timbre with another, and 3) rendering a new phrase expression for another given score. The resulting customized part is re-mixed with the remaining parts of the original performance to generate new polyphonic music signals. A single instrumental part is extracted by using an integrated tone model that consists of harmonic and inharmonic tone models with the aid of the score of the single instrumental part. The extraction incorporates a residual model for the single instrumental part in order to avoid crosstalk between instrumental parts. The extracted model parameters are classified into their averages and deviations. The former is treated as instrument timbre and is customized by mixing, while the latter is treated as phrase expression and is customized by rendering. We evaluated our method in three experiments. The first experiment focused on introduction of the residual model, and it showed that the model parameters are estimated more accurately by 35.0 points. The second focused on timbral customization, and it showed that our method is more robust by 42.9 points in spectral distance compared with a conventional sound analysis-synthesis method, STRAIGHT. The third focused on the acoustic fidelity of customizing performance, and it showed that rendering phrase expression according to the note sequence leads to more accurate performance by 9.2 points in spectral distance in comparison with a rendering method that ignores the note sequence.
Naoki Yasuraoka, Takehiro Abe, Katsutoshi Itoyama, Toru Takahashi 0001, Tetsuya Ogata, Hiroshi G. Okuno
ACM Multimedia3
2007 Integration and Adaptation of Harmonic and Inharmonic Models for Separating Polyphonic Musical Signals
abstract
This paper describes a sound source separation method for polyphonic sound mixtures of music to build an instrument equalizer for remixing multiple tracks separated from compact-disc recordings by changing the volume level of each track. Although such mixtures usually include both harmonic and inharmonic sounds, the difficulties in dealing with both types of sounds together have not been addressed in most previous methods that have focused on either of the two types separately. We therefore developed an integrated weighted-mixture model consisting of both harmonic-structure and inharmonic-structure tone models (generative models for the power spectrogram). On the basis of the MAP estimation using the EM algorithm, we estimated all model parameters of this integrated model under several original constraints for preventing over-training and maintaining intra-instrument consistency. Using standard MIDI files as prior information of the model parameters, we applied this model to compact-disc recordings and achieved the instrument equalizer.
Katsutoshi Itoyama, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
ICASSP (1)1