VLDB 2026 Research / reviewers in the wild / expert
Nobutaka Ono
dblp:87/3755
· DBLP profile ↗
90ranked-venue papers
3as first author
19since 2021 · last 2026
0000-0003-4242-2773ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 69 · 3 first-author · 13 since 2021Artificial intelligence and machine learning · 29 · 8 since 2021Computer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An end-to-end integration of speech separation and recognition with self-supervised learning representationabstractMulti-speaker automatic speech recognition (ASR) has gained growing attention in a wide range of applications, including conversation analysis and human–computer interaction. Speech separation and enhancement (SSE) and single-speaker ASR have witnessed remarkable performance improvements with the rapid advances in deep learning. Complex spectral mapping predicts the short-time Fourier transform (STFT) coefficients of each speaker and has achieved promising results in several SSE benchmarks. Meanwhile, self-supervised learning representation (SSLR) has demonstrated its significant advantage in single-speaker ASR. In this work, we push forward the performance of multi-speaker ASR under noisy reverberant conditions by integrating powerful SSE, SSL, and ASR models in an end-to-end manner. We systematically investigate both monaural and multi-channel SSE methods and various feature representations. Our experiments demonstrate the advantages of recently proposed complex spectral mapping and SSLRs in multi-speaker ASR. The experimental results also confirm that end-to-end fine-tuning with an ASR criterion is important to achieve state-of-the-art word error rates (WERs) even with powerful pre-trained models. Moreover, we show the performance trade-off between SSE and ASE and mitigate it with a multi-task learning framework with both SSE and ASR criteria. Yoshiki Masuyama, Xuankai Chang, Wangyou Zhang, Samuele Cornell, Nobutaka Ono, Yanmin Qian, Shinji Watanabe 0001 |
Comput. Speech Lang. | 6 |
| 2026 | Guided Masked Self-Distillation Modeling for Distributed Multimedia Sensor Event AnalysisabstractThis article addresses a new task: distributed multimedia sensor event analysis (DiMSEA). DiMSEA aims to analyze a series of human and machine activities (called “events” in this article) in complex and extensive real-world environments. Since an observation from a single sensor is often missing or fragmented in such an environment, observations from multiple locations and modalities should be integrated to analyze events comprehensively. However, a learning method has yet to be established to extract joint representations that effectively combine such distributed observations. Therefore, we propose guided masked self-distillation modeling (Guided-MELD) for inter-sensor relationship modeling. The basic idea of Guided-MELD is to learn to supplement the information from the masked sensor with information from other sensors needed to detect the event. Guided-MELD is expected to effectively distill fragmented target event information from sensors without over-relying on any specific sensors. To validate the effectiveness of the proposed method in DiMSEA, we recorded two new datasets: MM-Store and MM-Office. These datasets consist of human activities in a convenience store and an office, recorded using distributed cameras and microphones. Experimental results show that the proposed Guided-MELD improves event tagging and detection performance and outperforms conventional inter-sensor relationship modeling methods. Furthermore, the proposed method performed robustly even when sensors were reduced. Masahiro Yasuda, Noboru Harada, Yasunori Ohishi, Shoichiro Saito, Akira Nakayama, Nobutaka Ono |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2025 | People-Flow Estimation from Footstep Sounds: A Feasibility Study via Simulation Using Large-Scale Pedestrian Trajectories
Yu Kitano, Nobutaka Ono |
IEEE Big Data | 2 |
| 2025 | Mel-Spectrogram Inversion via Alternating Direction Method of MultipliersabstractSignal reconstruction from its mel-spectrogram is known as mel-spectrogram inversion and has many applications, including speech and foley sound synthesis. In this paper, we propose a mel-spectrogram inversion method based on a rigorous optimization algorithm. To reconstruct a time-domain signal with inverse short-time Fourier transform (STFT), both full-band STFT magnitude and phase should be predicted from a given mel-spectrogram. Their joint estimation has outperformed the cascaded full-band magnitude prediction and phase reconstruction by preventing error accumulation. However, the existing joint estimation method requires many iterations, and there remains room for performance improvement. We present an alternating direction method of multipliers (ADMM)-based joint estimation method motivated by its success in various nonconvex optimization problems including phase reconstruction. An efficient update of each variable is derived by exploiting the conditional independence among the variables. Our experiments demonstrate the effectiveness of the proposed method on speech and foley sounds. Yoshiki Masuyama, Natsuki Ueno, Nobutaka Ono |
ICASSP | 3 |
| 2024 | Causal and Relaxed-Distortionless Response Beamforming for Online Target Source ExtractionabstractIn this paper, we propose a low-latency beamforming method for target source extraction. Beamforming has been performed in the time-frequency domain and achieved promising results in offline applications. Meanwhile, it causes a long algorithmic delay due to the frame analysis. Such a delay is unacceptable in various low-latency real-time applications, including hearing aids. To reduce this delay, we propose a causal variant of the minimum power distortionless response (MPDR) beamformer. The proposed method constraints the non-causal components of the spatial filter to be zero in the optimization of the MPDR beamformer. The algorithmic delay is reduced to zero by applying the causal spatial filter in the time domain. We further propose to relax the distortionless constraint regarding the gain, which allows us to improve the extraction performance without a phase delay. The Douglas–Rachford splitting method and its online extension are adopted to solve the optimization problems of the proposed methods. In our experiment, the relaxed method outperformed various low-latency beamforming methods in terms of extraction performance. Yoshiki Masuyama, Kouei Yamaoka, Yuma Kinoshita, Taishi Nakashima, Nobutaka Ono |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2024 | Efficient Joint Optimization of Sampling Rate Offsets Using Entire Multichannel SignalabstractIn this paper, we propose a joint estimation method for the sampling rate offsets (SROs) of multiple recording devices. In wireless acoustic sensor networks, distributed microphones are connected to different analog-to-digital converters, and thus SROs occur on non-reference channels, which degrades the performance of various array signal processing techniques. To address this problem, we propose to jointly estimate and compensate SROs of all the non-reference channels. Since the proposed method is formulated as a multivariate non-convex optimization problem, we derive an efficient optimization algorithm on the basis of the majorization-minimization and majorizationequalization algorithms. We further propose to update SROs with only low-frequency components in the initial iterations to avoid undesired local optima. Our experimental results validate the effectiveness of the joint estimation of all SROs and demonstrate its advantage in subsequent array signal processing. Yoshiki Masuyama, Kouei Yamaoka, Takao Kawamura, Nobutaka Ono |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | Multi-Channel Speaker Extraction with Adversarial Training: The Wavlab Submission to The Clarity ICASSP 2023 Grand ChallengeabstractIn this work we detail our submission to the Clarity ICASSP 2023 grand challenge, in which participants have to develop a strong target speech enhancement system for hearing-aid (HA) devices in noisy-reverberant environments. Our system builds on our previous submission at the Second Clarity Enhancement Challenge (CEC2): iNeuBe-X, which consists in an iterative neural/conventional beamforming enhancement pipeline, guided by an enrollment utterance from the target speaker. This model, which won by a large margin the CEC2, is an extension of the state-of-the-art TF-GridNet model for multi-channel, streamable target-speaker speech enhancement. Here, this approach is extended and further improved by leveraging generative adversarial training, which we show proves especially useful when the training data is limited. Using only the official 6k training scenes data, our best model achieves 0.80 hearing-aid speech perception index (HASPI) and 0.41 hearing-aid speech quality index (HASQI) scores on the synthetic evaluation set. However, our model generalized poorly on the semi-real evaluation set. This highlights the fact that our community should focus more on real-world evaluation and less on fully synthetic datasets. Samuele Cornell, Zhongqiu Wang 0001, Yoshiki Masuyama, Shinji Watanabe 0001, Manuel Pariente, Nobutaka Ono, Stefano Squartini |
ICASSP | 6 |
| 2023 | Effectiveness of Inter- and Intra-Subarray Spatial Features for Acoustic Scene ClassificationabstractIn this paper, we investigate the effectiveness of spatial features for acoustic scene classification (ASC) with distributed microphones. Assuming that multiple subarrays, each containing multiple micro-phones, are distributed and synchronized, we consider two types of generalized cross-correlation phase transform (GCC-PHAT) as spatial features: the intra- and inter-subarray GCC-PHATs. They are obtained from channels within the same subarray and between different subarrays, respectively. The log-Mel spectrogram as a spectral feature and the intra- or inter-subarray GCC-PHAT are processed in the neural network. The experimental results show that increasing the number of channels did not markedly improve the ASC performance when using the spectral features alone. However, using either of the GCC-PHATs as the spatial feature together with the spectral features successfully improved the ASC performance. Takao Kawamura, Yuma Kinoshita, Nobutaka Ono, Robin Scheibler |
ICASSP | 3 |
| 2023 | Element Selection with Wide Class of Optimization Criteria Using Non-Convex Sparse OptimizationabstractElement selection techniques for high-dimensional features have various applications in machine learning. In general, the problem of element selection is typically solved by greedy methods or convex relaxation methods. However, these algorithms are applicable to only a specific class of optimization criteria such as the minimization of the squared error loss between the original and restored data. To overcome this limitation, we propose an element selection algorithm based on non-convex sparse optimization that can be used with a wider class of optimization criteria than conventional algorithms. In the proposed method, an algorithm based on the alternating direction method of multipliers (ADMM) is de-rived by reformulating the element selection problem as a matrix optimization on a non-convex set. A numerical experiment demonstrated the effectiveness of the proposed method for element selection with a non-squared error loss compared with the conventional greedy method. Taiga Kawamura, Natsuki Ueno, Nobutaka Ono |
ICASSP | 3 |
| 2023 | Fast Online Source Steering Algorithm for Tracking Single Moving Source Using Online Independent Vector AnalysisabstractWe address the problem of separating moving sources using online independent vector analysis (IVA). To solve this problem, researchers have extended the iterative projection (IP) and iterative source steering (ISS) algorithms developed for batch auxiliary-function-based IVA (AuxIVA) to online scenarios and showed their effectiveness. However, the conventional online IP and ISS are slow because they update K × K covariance matrices for all sources, where K is the number of microphones. Here, we show that, in a target-source tracking scenario in which only one source moves, there exists an inexpensive formula for online ISS that avoids updating the full covariance matrices without changing the behavior of the algorithm. The time complexity of the proposed algorithm, which we call online source steering (OSS), is K times smaller than that of the conventional online IP and ISS for the target-source tracking task. A numerical experiment on separating a moving source demonstrates that the proposed OSS is significantly faster than the conventional online IP and ISS. Taishi Nakashima, Rintaro Ikeshita, Nobutaka Ono, Shoko Araki, Tomohiro Nakatani |
ICASSP | 3 |
| 2023 | Sound Field Interpolation for Rotation-Invariant Multichannel Array Signal ProcessingabstractIn this paper, we present a sound field interpolation for array signal processing (ASP) that is robust to rotation of a circular microphone array (CMA), and we evaluate beamforming as one of its applications. Most ASP methods assume a time-invariant acoustic transfer system (ATS) from sources to the microphone array. This assumption makes it challenging to perform ASP in real situations where sources and the microphone array can move. Therefore, considering a time-variant ATS is an essential task for the use of ASP. In this study, we focus on one such movement, the rotation of the CMA. Our method interpolates the sound field on the circumference of a circle, where microphones are equally spaced, based on the sampling theorem on the circle. The interpolation enables us to estimate the signals at the microphone positions before the rotation. Hence, conventional ASP, which assumes a time-invariant ATS, is applicable after interpolation without modification. We developed two beamforming schemes, one for batch and one for online processing, that combine the minimum power distortionless response beamformer and sound field interpolation. We evaluated the dependences of the interpolation on frequency and rotation angle using the signal-to-error ratio. Additionally, simulation results demonstrated that the two proposed schemes improve the beamformer's performance when the CMA rotates. Yukoh Wakabayashi, Kouei Yamaoka, Nobutaka Ono |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Entrainment Analysis for Assessment of Autistic Speech Prosody Using Bottleneck Features of Deep Neural NetworkabstractIn the present study, we quantify entrainment characteristics of conversation with the aim of automatic assessment of the severity of autism spectrum disorder (ASD). We focus on pairs of utterances immediately before and after turn-takings, which have prosodic/acoustic similarities.The clinical severity of ASD is estimated by the bottleneck features obtained by an hourglass-shaped deep neural network (DNN) in the neural entrainment distance (NED) method used to measure the degree of entrainment. The DNN is firstly pre-trained using a large conversation corpus in various daily situations and then fine-tuned with conversations during the Autism Diagnostic Observation Schedule (ADOS) assessment. Absolute difference vectors are calculated from the bottleneck feature vectors between a pair of utterances. Centroid and variance of the absolute difference vectors are combined with the speech features discovered in our previous study in order to estimate the scores of ASD severity.Consequently, the estimated scores significantly correlate with the actual observed ADOS ‘Reciprocity' scores with a coefficient of 0.70. This result shows the effective use of finetuning technique with data of typically developed individuals and, furthermore, reveals social communication deficits in ASD individuals represented by utterances adjacent to turn-takings. Keiko Ochi, Nobutaka Ono, Keiho Owada, Miho Kuroda, Shigeki Sagayama, Hidenori Yamasue |
ICASSP | 2 |
| 2022 | Instantaneous Linear Dimensionality Reduction of Multichannel Time-Series Signal for Array Signal ProcessingabstractLinear dimensionality reduction of signals observed by a sensor array is often useful in balancing the accuracy and speed of post-stage processing, especially in real-time systems with limited computational resources. However, for multichannel time-series signals having time-invariant intertemporal and interchannel correlations, the direct application of frequency-wise linear dimensionality reduction method requires a large number of digital filters with large filter lengths, which is still unpreferable in the viewpoint of computational cost. We propose a frequency-independent, i.e., instantaneous, linear dimensionality reduction method that achieves low computational cost and latency and high restoration accuracy. We also show several results of numerical experiments to compare the proposed method with other instantaneous linear dimensionality reduction methods, i.e., the principal component analysis and element selection method, and demonstrate the effectiveness of the proposed method. Natsuki Ueno, Nobutaka Ono |
ICASSP | 2 |
| 2022 | Joint Optimization of Sampling Rate Offsets Based on Entire Signal Relationship Among Distributed MicrophonesabstractIn this paper, we propose to simultaneously estimate all the sampling rate offsets (SROs) of multiple devices.In a distributed microphone array, the SRO is inevitable, which deteriorates the performance of array signal processing.Most of the existing SRO estimation methods focused on synchronizing two microphones.When synchronizing more than two microphones, we select one reference microphone and estimate the SRO of each non-reference microphone independently.Hence, the relationship among signals observed by non-reference microphones is not considered.To address this problem, the proposed method jointly optimizes all SROs based on a probabilistic model of a multichannel signal.The SROs and model parameters are alternately updated to increase the log-likelihood based on an auxiliary function.The effectiveness of the proposed method is validated on mixtures of various numbers of speakers. Yoshiki Masuyama, Kouei Yamaoka, Nobutaka Ono |
INTERSPEECH | 3 |
| 2022 | Use of Nods Less Synchronized with Turn-Taking and Prosody During Conversations in Adults with Autism
Keiko Ochi, Nobutaka Ono, Keiho Owada, Miho Kuroda, Shigeki Sagayama, Hidenori Yamasue |
INTERSPEECH | 2 |
| 2022 | End-to-End Integration of Speech Recognition, Dereverberation, Beamforming, and Self-Supervised Learning RepresentationabstractSelf-supervised learning representation (SSLR) has demonstrated its significant effectiveness in automatic speech recognition (ASR), mainly with clean speech. Recent work pointed out the strength of integrating SSLR with single-channel speech enhancement for ASR in noisy environments. This paper further advances this integration by dealing with multi-channel input. We propose a novel end-to-end architecture by integrating dereverberation, beamforming, SSLR, and ASR within a single neural network. Our system achieves the best performance reported in the literature on the CHiME-4 6-channel track with a word error rate (WER) of 1.77%. While the WavLM-based strong SSLR demonstrates promising results by itself, the end-to-end integration with the weighted power minimization distortionless response beamformer, which simultaneously performs dereverberation and denoising, improves WER significantly. Its effectiveness is also validated on the REVERB dataset. Yoshiki Masuyama, Xuankai Chang, Samuele Cornell, Shinji Watanabe 0001, Nobutaka Ono |
SLT | 5 |
| 2021 | Joint Dereverberation and Separation With Iterative Source SteeringabstractWe propose a new algorithm for joint dereverberation and blind source separation (DR-BSS). Our work builds upon the IRLMA-T framework that applies a unified filter combining dereverberation and separation. One drawback of this framework is that it requires several matrix inversions, an operation inherently costly and with potential stability issues. We leverage the recently introduced iterative source steering (ISS) updates to propose two algorithms mitigating this issue. Albeit derived from first principles, the first algorithm turns out to be a natural combination of weighted prediction error (WPE) dereverberation and ISS-based BSS, applied alternatingly. In this case, we manage to reduce the number of matrix inversion to only one per iteration and source. The second algorithm updates the ILRMA-T matrix using only sequential ISS updates requiring no matrix inversion at all. Its implementation is straightforward and memory efficient. Numerical experiments demonstrate that both methods achieve the same final performance as ILRMA-T in terms of several relevant objective metrics. In the important case of two sources, the number of iterations required is also similar. Taishi Nakashima, Robin Scheibler, Masahito Togami, Nobutaka Ono |
ICASSP | 4 |
| 2021 | Rotation-Robust Beamforming Based on Sound Field Interpolation with Regularly Circular Microphone ArrayabstractIn this paper, we present a novel framework of beamforming robust for a microphone array rotation. In most array signal processing methods, the time-invariant transfer system from a source to a microphone is assumed for calculating a spatial filter. This assumption makes it difficult to use the microphone array in real situations since sources and the microphone array may move. In this work, we focus on one such movement, the array’s rotation. The key in our method is to use a regularly circular microphone array where microphones are equally spaced on a circle’s circumference. Based on this, we propose a method to interpolate the sound field on the circumference. The interpolation enables us to estimate the signals at the microphone positions before the rotation even when the array rotates. Hence, the conventional array signal processing assuming the time-invariant system is applicable. Simulation results indicated that the proposed method improves the beamformer’s performance under the array rotation. Yukoh Wakabayashi, Kouei Yamaoka, Nobutaka Ono |
ICASSP | 3 |
| 2021 | Time-Frequency-Bin-Wise Linear Combination of Beamformers for Distortionless Signal EnhancementabstractIn this paper, we address signal enhancement in underdetermined situations and propose new beamforming algorithms. Beamforming in (over)determined situations can successfully reduce noise signals without distortion of a desired signal, which is known to be a desirable property, especially for automatic speech recognition systems. Even in underdetermined situations, time-frequency (TF) masking attains outstanding performance in noise reduction, although it tends to generate artifacts. Integrating these two approaches to benefit from both their advantages, we here propose time-frequency-bin-wise switching (TFS) and time-frequency-bin-wise linear combination (TFLC) beamforming. In the proposed methods, we utilize the best combination of beamformers among multiple beamformers at each TF bin, each of which suppresses a particular combination of interferers. First, we propose a general formulation of signal enhancement employing multiple spatial filters. Then a joint optimization problem of designing the spatial filters and estimating the suitable weights to combine them is considered under a unified minimum variance criterion. Finally, we present efficient algorithms to solve the problem. In experiments, we used an objective criterion that quantifies the amount of signal distortion caused by the enhancement function and confirmed that the proposed methods effectively suppress interferers without distortion of the target signal. Kouei Yamaoka, Nobutaka Ono, Shoji Makino |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Fast and Stable Blind Source Separation with Rank-1 UpdatesabstractWe propose a new algorithm for the blind source separation of acoustic sources. This algorithm is an alternative to the popular auxiliary function based independent vector analysis using iterative projection (AuxIVA-IP). It optimizes the same cost function, but instead of alternate updates of the rows of the demixing matrix, we propose a sequence of rank-1 updates. Remarkably, and unlike the previous method, the resulting updates do not require matrix inversion. Moreover, their computational complexity is quadratic in the number of microphones, rather than cubic in AuxIVA-IP. In addition, we show that the new method can be derived as alternate updates of the steering vectors of sources. Accordingly, we name the method iterative source steering (AuxIVA-ISS). Finally, we confirm in simulated experiments that the proposed algorithm separates sources just as well as AuxIVA-IP, at a lower computational cost. Robin Scheibler, Nobutaka Ono |
ICASSP | 2 |
| 2020 | Fast Independent Vector Extraction by Iterative SINR MaximizationabstractWe propose fast independent vector extraction (FIVE), a new algorithm that blindly extracts a single non-Gaussian source from a Gaussian background. The algorithm iteratively computes beam-forming weights maximizing the signal-to-interference-and-noise ratio for an approximate noise covariance matrix. We demonstrate that this procedure minimizes the negative log-likelihood of the input data according to a well-defined probabilistic model. The minimization is carried out via the auxiliary function technique whereas, unlike related methods, the auxiliary function is globally minimized at every iteration. Numerical experiments are carried out to assess the performance of FIVE. We find that it is vastly superior to competing methods in terms of convergence speed, and has high potential for real-time applications. Robin Scheibler, Nobutaka Ono |
ICASSP | 2 |
| 2020 | Independent Low-Rank Matrix Analysis Based on Time-Variant Sub-Gaussian Source Model for Determined Blind Source SeparationabstractIndependent low-rank matrix analysis (ILRMA) is a fast and stable method of blind audio source separation. Conventional ILRMAs assume time-variant (super-)Gaussian source models, which can only represent signals that follow a super-Gaussian distribution. In this article, we focus on ILRMA based on a generalized Gaussian distribution (GGD-ILRMA) and propose a new type of GGD-ILRMA that adopts a time-variant sub-Gaussian distribution for the source model. We propose a new update scheme called generalized iterative projection for homogeneous source models (GIP-HSM) and obtain a convergence-guaranteed update rule for demixing spatial parameters by combining the GIP-HSM scheme and the majorization-minimization (MM) algorithm. Furthermore, a new extension of the MM algorithm is proposed for the convergence acceleration by applying the majorization-equalization algorithm to a multivariate case. In the experimental evaluation, we show the versatility of the proposed method, i.e., the proposed time-variant sub-Gaussian source model can be applied to various types of source signal. Shinichi Mogami, Norihiro Takamune, Daichi Kitamura, Hiroshi Saruwatari, Yu Takahashi, Kazunobu Kondo, Nobutaka Ono |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2019 | Estimation of Sampling Frequency Mismatch between Distributed Asynchronous Microphones under Existence of Source Movements with Stationary Time Periods DetectionabstractIn this paper, we propose a method of estimating the sampling frequency mismatch among asynchronous recording devices, even when the sources sometimes move. For a spatially stationary source, there is a method of estimating the sampling frequency mismatch, which appears in the drift of the time difference among the observed digitized signals. When the source moves, however, the change of its location also affects the drift, and the method fails to estimate the mismatch. In the meantime, looking at the practical recording situations, sources sometimes move but sometimes do not move. That is, there should be a set of time frames in which we can assume the spatial stationarity of sources, and to which we are still able to apply the sampling frequency mismatch estimation method. Based on this idea, our proposed method first detects a set of time frames where we can assume the spatial stationary by clustering the time frames using the covariance matrix of each recording device, and then estimates the mismatch by using the detected stationary time frames. Using real recordings with several IC recorders, we show that the proposed method can estimate the sampling frequency mismatch accurately even when the sources sometimes move. Shoko Araki, Nobutaka Ono, Keisuke Kinoshita, Marc Delcroix |
ICASSP | 2 |
| 2019 | Multi-modal Blind Source Separation with Microphones and BlinkiesabstractWe propose a blind source separation algorithm that jointly exploits measurements by a conventional microphone array and an ad hoc array of low-rate sound power sensors called blinkies. While providing less information than microphones, blinkies circumvent some difficulties of microphone arrays in terms of manufacturing, synchronization, and deployment. The algorithm is derived from a joint probabilistic model of the microphone and sound power measurements. We assume the separated sources to follow a time-varying spherical Gaussian distribution, and the non-negative power measurement space-time matrix to have a low-rank structure. We show that alternating updates similar to those of independent vector analysis and Itakura-Saito non-negative matrix factorization decrease the negative log-likelihood of the joint distribution. The proposed algorithm is validated via numerical experiments. Its median separation performance is found to be up to 8 dB more than that of independent vector analysis, with significantly reduced variability. Robin Scheibler, Nobutaka Ono |
ICASSP | 2 |
| 2019 | Time-frequency-bin-wise Switching of Minimum Variance Distortionless Response Beamformer for Underdetermined SituationsabstractIn this paper, we present a speech enhancement method using two microphones in underdetermined situations. Time-frequency (TF) binary masking is a conventional method of enhancing speech in underdetermined situations by appropriately multiplying each TF component by zero or one. Extending this method, we previously proposed a new method called the time-frequency-bin-wise switching (TFS) beamformer. In this method, we switch multiple preconstructed beamformers in each TF bin, each of which suppresses a particular interferer. However, this method requires the pre-estimation of beamformer filter coefficients using the target-active period and interferer-wise-active periods as the prior information. In this paper, to overcome this limitation, we formulate the switching and construction of spatial filters as a joint optimization problem, which can be understood from two viewpoints: the clustering of the most dominant interferer signal in each TF bin and the construction of a minimum variance distortionless response beamformer using such bins. In an experiment, we confirmed that the proposed method was superior to conventional TF masking and fixed beamforming during speech enhancement regardless of the direction of interferers. Kouei Yamaoka, Nobutaka Ono, Shoji Makino, Takeshi Yamada |
ICASSP | 2 |
| 2019 | Blink-former: Light-aided beamforming for multiple targets enhancementabstractWe propose a multimodal framework to enhance multiple target sound sources using a conventional microphone array, a video camera, and sound power sensors, called Blinkies, that we have recently developed. Each Blinky consists of a microphone, LEDs, a microcontroller, and a battery. One of the LEDs intensity is varied according to sound power, that is, the Blinky works as a sound-to-light conversion sensor. They are easy to distribute over a large area, and thus, the sound power information therein can be harvested by capturing the LED signals with a video camera. Although these signals are a mixture of contributions from multiple sources, we demonstrate that they can be separated into individual source activities by non-negative matrix factorization. The obtained activities are further utilized to design maximum signal-to-interference-and-noise ratio beamformers enhancing the source signals. We conduct numerical simulations and real experiments to evaluate the performance of this method in diffuse noise environment. The experimental results show that the proposed scheme using Blinkies is superior to competing algorithms, especially at low signal-to-noise ratio. Daiki Horiike, Robin Scheibler, Yukoh Wakabayashi, Nobutaka Ono |
MMSP | 4 |
| 2019 | Bilevel Optimization Using Stationary Point of Lower-Level Objective Function for Discriminative Basis Learning in Nonnegative Matrix FactorizationabstractIn this letter, we address an audio signal separation problem and propose a new effective algorithm for solving a bilevel optimization in discriminative nonnegative matrix factorization (NMF). Recently, discriminative training of NMF bases has been developed for better signal separation in supervised NMF (SNMF), which exploits a priori training of given sample signals. The optimization in this method consists of a simultaneous minimization of two objective functions, resulting in a bilevel optimization problem with SNMF (BiSNMF), where conventional methods approximately solve this optimization. To strictly solve BiSNMF, we introduce a new algorithm with the following two features: (a) conversion of the optimization constraint into a penalty term and (b) optimization of the reformulated problem on the basis of a multiplicative steepest descent, ensuring the nonnegativity of variables. Experiments on music signal separation show the efficacy of the proposed algorithm. Hiroaki Nakajima, Daichi Kitamura, Norihiro Takamune, Hiroshi Saruwatari, Nobutaka Ono |
IEEE Signal Process. Lett. | 5 |
| 2019 | Acoustic Topic Model for Scene Analysis With Intermittently Missing ObservationsabstractWe propose a sophisticated method of acoustic scene analysis with intermittently missing observations, which analyzes acoustic scenes and restores missing observations simultaneously on the basis of the temporal correlation between acoustic words. One effective strategy for analyzing acoustic scenes is to characterize them as a combination of acoustic words. An acoustic topic model (ATM) is one of the techniques, which models the process generating multiple acoustic words. Here, an acoustic word corresponds to a sound category, while it has a homogenous time duration and is defined time frame by time frame. In the ATM, it is assumed that all acoustic words are observed, and therefore, it cannot be applied if any acoustic observations are missing. However, acoustic observations may sometimes be missing because of poor recording conditions, transmission loss, or privacy reasons. In the proposed method, focusing on the fact that acoustic words are temporally correlated, we consider the transition of acoustic words in two ways: First, by modeling the temporal transition of acoustic words directly using a Markov process and finally, by modeling the temporal transition of hidden states that generate acoustic words using a hidden Markov model. We then incorporate each transition model in a process generating acoustic words based on the ATM. The proposed method allows us to analyze acoustic scenes from acoustic words by restoring missing acoustic words. In our experiments, the proposed method exhibited a classification accuracy of acoustic scenes close to that for the case of no missing observations even when 50% of the observations were missing. Moreover, the model considering the hidden-state transition can classify acoustic scenes more accurately than the model considering the acoustic word transition directly. Keisuke Imoto, Nobutaka Ono |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Independent Deeply Learned Matrix Analysis for Determined Audio Source SeparationabstractIn this paper, we propose a new framework called independent deeply learned matrix analysis (IDLMA), which unifies a deep neural network (DNN) and independence-based multichannel audio source separation. IDLMA utilizes both pretrained DNN source models and statistical independence between sources for the separation, where the time-frequency structures of each source are iteratively optimized by a DNN while enhancing the estimation accuracy of the spatial demixing filters. As the source generative model, we introduce a complex heavy-tailed distribution to improve the separation performance. In addition, we address a semi-supervised situation; namely, a solo-recorded audio dataset can be prepared for only one source in the mixture signal. To solve the limited-data problem, we propose an appropriate data augmentation method to adapt the DNN source models to the observed signal, which enables IDLMA to work even in the semi-supervised situation. Experiments are conducted using music signals with a training dataset in both supervised and semi-supervised situations. The results show the validity of the proposed method in terms of the separation accuracy. Naoki Makishima, Shinichi Mogami, Norihiro Takamune, Daichi Kitamura, Hayato Sumino, Shinnosuke Takamichi, Hiroshi Saruwatari, Nobutaka Ono |
IEEE ACM Trans. Audio Speech Lang. Process. | 8 |
| 2018 | Deeply Learned Filter Response Functions for Hyperspectral ReconstructionabstractHyperspectral reconstruction from RGB imaging has recently achieved significant progress via sparse coding and deep learning. However, a largely ignored fact is that existing RGB cameras are tuned to mimic human trichromatic perception, thus their spectral responses are not necessarily optimal for hyperspectral reconstruction. In this paper, rather than use RGB spectral responses, we simultaneously learn optimized camera spectral response functions (to be implemented in hardware) and a mapping for spectral reconstruction by using an end-to-end network. Our core idea is that since camera spectral filters act in effect like the convolution layer, their response functions could be optimized by training standard neural networks. We propose two types of designed filters: a three-chip setup without spatial mosaicing and a single-chip setup with a Bayer-style 2x2 filter array. Numerical simulations verify the advantages of deeply learned spectral responses compared to existing RGB cameras. More interestingly, by considering physical restrictions in the design process, we are able to realize the deeply learned spectral response functions by using modern film filter production technologies, and thus construct data-inspired multispectral cameras for snapshot hyperspectral imaging. Shijie Nie, Lin Gu 0003, Yinqiang Zheng, Antony Lam, Nobutaka Ono, Imari Sato |
CVPR | 5 |
| 2018 | Meeting Recognition with Asynchronous Distributed Microphone Array Using Block-Wise Refinement of Mask-Based MVDR BeamformerabstractThis paper addresses a front-end system for speech recognition of spontaneous conversational speech signals that are recorded with asynchronous distributed microphones such as smartphones. In our previous work, we proposed combining blind synchronization and a state-of-the-art microphone array speech enhancement technique, e.g., a time-frequency mask based minimum variance distortionless response (MVDR) beamformer. This approach has provided reasonably high recognition performance even if we use asynchronous microphones. However, because the previous speech enhancement method was applied in a full-batch mode, it has been difficult to track speaker position movement in a real meeting conversation. To make it possible to handle the speaker movement, this paper describes our attempt to refine the mask-based MVDR beamformer in a blockwise manner, and reports that such a refinement reduces the word error rate from 31.4% to 28.8% for real meeting recordings. Shoko Araki, Nobutaka Ono, Keisuke Kinoshita, Marc Delcroix |
ICASSP | 2 |
| 2018 | Sonoloc: Scalable positioning of commodity mobile devicesabstractWe present Sonoloc, a mobile app and system that allows a set of co-located commodity smart devices to determine their relative positions without local infrastructure. Sonoloc enables users to address each other based on their relative positions at events like meetings, talks, or conferences. This capability can, for instance, aid spontaneous communication among users based on their relative position (e.g., in a given section of a room, at the same table, or in a given seat), facilitate interaction between speaker and audience in a lecture hall, and enable the distribution of materials, crowdsensing, and feedback collection based on users' location. Sonoloc can position any number of devices within acoustic range with a constant number of chirps emitted by a self-organized subset of devices. Our experimental evaluation shows that the system can locate up to hundreds of devices with an accuracy of tens of centimeters using up to 15 audio chirps emitted by dynamically selected devices, in actual rooms and despite substantial background noise. Viktor Erdélyi, Trung-Kien Le 0002, Bobby Bhattacharjee, Peter Druschel, Nobutaka Ono |
MobiSys | 5 |
| 2017 | Meeting recognition with asynchronous distributed microphone arrayabstractRecently, recognition of conversational speech such as meetings has widely been studied. However, most existing approaches rely on using a single close talking microphone or a distant microphone array where all the microphones are synchronous. In contrast, this paper tackles a recognition task of conversational speech recorded with asynchronous distributed microphones, to which conventional array processing is not directly applicable. We demonstrate that we can significantly improve recognition performance even when microphones are asynchronous by combining blind synchronization and state-of-the-art microphone array speech enhancement techniques such as independent vector analysis (IVA) and a time-frequency mask based minimum variance distortionless response (MVDR) beamformer. Using such a front-end, we could reduce the word error rate from 42.2 % to 29.9 % for real meeting recordings. Shoko Araki, Nobutaka Ono, Keisuke Kinoshita, Marc Delcroix |
ASRU | 2 |
| 2017 | Blind source separation based on independent low-rank matrix analysis with sparse regularization for time-series activityabstractIn this paper, we propose a new blind source separation (BSS) method based on independent low-rank matrix analysis (ILRMA) with novel sparse regularization. ILRMA is a recently proposed BSS algorithm that simultaneously estimates a demixing matrix and source spectrogram models based on nonnegative matrix factorization (NMF). To improve the separation accuracy and stability, an additional constraint such as sparseness is needed but there have been no studies on this so far. In this study, we introduce an a priori statistical model for time-series amplitudes of source spectrograms, employing a new frequency-wise sparse regularization using estimates from the Bayesian postfilter to enhance the modeling accuracy. This regularization results in a bilevel optimization problem that consists of the estimation of a sparsity-emphasized source model using NMF and the separation of sources by ILRMA. In this paper, we present two approximated optimization schemes and their combination for performing regularized ILRMA. The efficacy of the proposed method is confirmed in a BSS experiment. Yoshiki Mitsui, Daichi Kitamura, Shinnosuke Takamichi, Nobutaka Ono, Hiroshi Saruwatari |
ICASSP | 4 |
| 2017 | Low-latency real-time blind source separation for hearing aids based on time-domain implementation of online independent vector analysis with truncation of non-causal componentsabstractIn this paper, we present a low-latency scheme for real-time blind source separation (BSS) based on online auxiliary-function-based independent vector analysis (AuxIVA). In many real-time audio applications, especially hearing aids, low latency is highly desirable. Conventional frequency-domain BSS methods suffer from a delay caused by frame analysis. To reduce the delay, we implement separation filters as multiple FIR filters in the time domain, which are converted from demixing matrices estimated by online AuxIVA in the frequency domain. Also, to further reduce the latency, part of the non-causal components of the FIR filters are truncated on the basis of causality analysis for ideal separation filters using a simple model. By experimental evaluation using a head and torso simulator in a real environment, the proposed algorithm with an algorithmic delay of less than 10 ms exhibited a separation performance of 7.7 dB in terms of the signal-to-interference ratio (SIR), which was less than 1.4 dB degradation from the case of conventional frequency-domain implementation. Masahiro Sunohara, Chiho Haruta, Nobutaka Ono |
ICASSP | 3 |
| 2017 | Light transport component decomposition using multi-frequency illuminationabstractScene appearance is a mixture of light transport phenomena ranging from direct reflection to complicated effect such as inter-reflection and subsurface scattering. To decompose scene appearance into meaningful photometric components is very helpful in scene understanding and image editing. However, it has proven to be a difficult task. In this paper, we explore the difference of direct components obtained by multi-frequency illumination for light transport component decomposition. We apply independent vector analysis (IVA) to this task with no fixed constraints. Experiment results have verified the effectiveness of our method and its applicability to generic scenes. Art Subpa-Asa, Yinqiang Zheng, Nobutaka Ono, Imari Sato |
ICIP | 3 |
| 2017 | Acoustic scene classification using asynchronous multichannel observations with different lengthsabstractTo utilize asynchronous multichannel recordings with different start and end time of recordings for acoustic scene analysis, we propose a combination method for estimating unrecorded durations and extracting spatial features. Focusing on the fact that amplitude information is relatively robust to the estimation error of the unrecorded durations and the synchronization mismatch of multichannel recordings, the proposed method combines a multiple imputation and spatial cepstrum, both of which are based on the amplitude information. The proposed method allows us to estimate unrecorded durations in multichannel observations and analyze acoustic scenes using spatial information extracted from whole multichannel observations, thus achieving more accurate scene analysis. An evaluation experiment indicated that the proposed method is valid for acoustic scene classification with asynchronous multichannel recordings including different length. Keisuke Imoto, Nobutaka Ono |
MMSP | 2 |
| 2017 | Spatial Cepstrum as a Spatial Feature Using a Distributed Microphone Array for Acoustic Scene AnalysisabstractIn this paper, with the aim of using the spatial information obtained from a distributed microphone array employed for acoustic scene analysis, we propose a robust and efficient method, which is called the spatial cepstrum. In our approach, similarly to the cepstrum, which is widely used as a spectral feature, the logarithm of the amplitude in multichannel observation is converted to a feature vector by a linear orthogonal transformation. This linear orthogonal transformation is achieved by principal component analysis (PCA) in general. Moreover, we also show that for a circularly symmetric microphone arrangement with an isotropic sound field, PCA is identical to the inverse discrete Fourier transform and the spatial cepstrum exactly corresponds to the cepstrum. The proposed approach does not require the positions of the microphones and is robust against the synchronization mismatch of channels, thus ensuring its suitability for use with a distributed microphone array. Experimental results obtained using actual environmental sounds verify the validity of our approach even when a smaller feature dimension than the original one is used, which is achieved by dimensionality reduction through PCA. Additionally, experimental results also indicate that the robustness of the proposed method is satisfactory for observations that have the synchronization mismatch of channels. Keisuke Imoto, Nobutaka Ono |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Introduction to the Special Section on Sound Scene and Event AnalysisabstractThe papers in this special section are devoted to the growing field of acoustic scene classification and acoustic event recognition. Machine listening systems still have difficulties to reach the ability of human listeners in the analysis of realistic acoustic scenes. If sustained research efforts have been made for decades in speech recognition, speaker identification and to a lesser extent in music information retrieval, the analysis of other types of sounds, such as environmental sounds, is the subject of growing interest from the community and is targeting an ever increasing set of audio categories. This problem appears to be particularly challenging due to the large variety of potential sound sources in the scene, which may in addition have highly different acoustic characteristics, especially in bioacoustics. Furthermore, in realistic environments, multiple sources are often present simultaneously, and in reverberant conditions. Gaël Richard, Tuomas Virtanen, Juan Pablo Bello, Nobutaka Ono, Hervé Glotin |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2017 | Sleep Apnea Detection via Depth Video and Audio Feature LearningabstractObstructive sleep apnea, characterized by repetitive obstruction in the upper airway during sleep, is a common sleep disorder that could significantly compromise sleep quality and quality of life in general. The obstructive respiratory events can be detected by attended in-laboratory or unattended ambulatory sleep studies. Such studies require many attachments to a patient's body to track respiratory and physiological changes, which can be uncomfortable and compromise the patient's sleep quality. In this paper, we propose to record depth video and audio of a patient using a Microsoft Kinect camera during his/her sleep, and extract relevant features to correlate with obstructive respiratory events scored manually by a scientific officer based on data collected by Philips system Alice6 LDxS that is commonly used in sleep clinics. Specifically, we first propose an alternating-frame H.264 video encoding scheme and bit recovery scheme at the decoder. Next, we perform depth video temporal denoising using a motion vector graph smoothness prior. Then, we build a dual-ellipse model and track a patient's chest and abdominal movements in the denoised videos. Finally, we extract features from both depth video and audio for classifier training and respiratory event detection. Experimental results show 1) that our depth video compression scheme outperforms a competitor that records only the 8 most significant bits, 2) our graph-based temporal denoising scheme reduces the flickering effect without over-smoothing, and 3) our trained classifiers can deduce respiratory events scored manually based on data collected by system Alice6 LDxS with high accuracy. Cheng Yang 0003, Gene Cheung, Vladimir Stankovic 0001, Nobutaka Ono |
IEEE Trans. Multim. | 5 |
| 2016 | Experimental validation of TOA-based methods for microphones array positions calibrationabstractThis study is an experimental validation of a new closed-form method for automatic array position calibration, based on time of arrival (TOA) measurements between sources and sensors. An experiment with a large array composed of 121 microphones and a dozen of sources has been set up. We first show that, when considering the whole array, this calibration method gives results on par with a reference state-of-the-art acoustic method. We then show experimentally that the new method provides significantly better results when the number of sources and microphones decreases, confirming numerical simulations. We conclude the paper with a discussion on methodological issues for array position calibration. Trung-Kien Le 0002, Nobutaka Ono, Thibault Nowakowski, Laurent Daudet, Julien de Rosny |
ICASSP | 2 |
| 2016 | Automatic Discrimination of Soft Voice Onset Using Acoustic Features of Breathy Voicing
Keiko Ochi, Koichi Mori, Naomi Sakai, Nobutaka Ono |
INTERSPEECH | 4 |
| 2016 | Multi-Talker Speech Recognition Based on Blind Source Separation with ad hoc Microphone Array Using Smartphones and Cloud Storage
Keiko Ochi, Nobutaka Ono, Shigeki Miyabe, Shoji Makino |
INTERSPEECH | 2 |
| 2016 | Determined Blind Source Separation Unifying Independent Vector Analysis and Nonnegative Matrix FactorizationabstractThis paper addresses the determined blind source separation problem and proposes a new effective method unifying independent vector analysis (IVA) and nonnegative matrix factorization (NMF). IVA is a state-of-the-art technique that utilizes the statistical independence between sources in a mixture signal, and an efficient optimization scheme has been proposed for IVA. However, since the source model in IVA is based on a spherical multivariate distribution, IVA cannot utilize specific spectral structures such as the harmonic structures of pitched instrumental sounds. To solve this problem, we introduce NMF decomposition as the source model in IVA to capture the spectral structures. The formulation of the proposed method is derived from conventional multichannel NMF (MNMF), which reveals the relationship between MNMF and IVA. The proposed method can be optimized by the update rules of IVA and single-channel NMF. Experimental results show the efficacy of the proposed method compared with IVA and MNMF in terms of separation accuracy and convergence speed. Daichi Kitamura, Nobutaka Ono, Hiroshi Sawada, Hirokazu Kameoka, Hiroshi Saruwatari |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Acoustic scene analysis from acoustic event sequence with intermittent missing eventabstractWe propose a novel method for analyzing acoustic scenes that can sophisticatedly estimate acoustic scenes from an acoustic event sequence with intermittent missing events. On the basis of the idea that acoustic events are temporally correlated, we model the transition of acoustic events using a hidden Markov model (HMM) and estimate missing acoustic events. Then, we incorporate the transition of acoustic events in a generative process of acoustic event sequence associated with the acoustic scenes based on acoustic topic model (ATM). Since the proposed method allows us to analyze acoustic scenes from acoustic event sequences while estimating missing acoustic events, we can estimate acoustic scenes successfully and restore missing acoustic events. Evaluation results indicate that the proposed method achieves an estimation accuracy for acoustic scenes comparable to that obtained when there is no missing data. Additionally, the proposed model can estimate acoustic events that are strongly correlated with acoustic scenes in an acoustic event sequence. Keisuke Imoto, Nobutaka Ono |
ICASSP | 2 |
| 2015 | Efficient multichannel nonnegative matrix factorization exploiting rank-1 spatial modelabstractThis paper proposes a new efficient multichannel nonnegative matrix factorization (NMF) method. Recently, multichannel NMF (MNMF) has been proposed as a means of solving the blind source separation problem. This method estimates a mixing system of sources and attempts to separate them in a blind fashion. However, this method is strongly dependent on its initial values because there are no constraints in the spatial models. To solve this problem, we introduce a rank-1 spatial model into MNMF. The proposed method estimates a demixing matrix while representing sources using NMF bases and can be optimized by the update rules of independent vector analysis and conventional single-channel NMF. Experimental results show the efficacy of the proposed method in terms of robustness and convergence speed. Daichi Kitamura, Nobutaka Ono, Hiroshi Sawada, Hirokazu Kameoka, Hiroshi Saruwatari |
ICASSP | 2 |
| 2015 | Reference-distance estimation approach for TDOA-based source and sensor localizationabstractIn this paper, we present a new method to find solutions to the time difference of arrival (TDOA)-based source and sensor localization problem. This paper is a continuation of [1], in which sources and sensors are localized on the basis of time of arrival (TOA) measurements. Generally, the TOA is known if the TDOA and reference-distances with the sound velocity are given, where the reference-distances are defined as the distances from the first (reference) sensor to the sources. We show that when the numbers of sources and sensors are at least six and eight, respectively, the reference-distances can be computed directly from TDOA measurements. This means that in such cases, the positions of the sources and sensors can be directly estimated in closed-form solutions, except for one reference-distance, which is estimated by a grid search. The validity of our algorithm is evaluated by synthetic experiments in noise-free and noisy cases. Trung-Kien Le 0002, Nobutaka Ono |
ICASSP | 2 |
| 2015 | Designing multichannel source separation based on single-channel source separationabstractIn this paper, an extension of independent vector analysis (IVA), model-based IVA, is proposed for multichannel source separation. For obtaining better source models, we introduce a single-channel source separation method, and utilize the outputs as source variances in time-frequency-variant Gaussian source model. The demixing matrices are estimated in the same way as a state-of-the-art IVA method, auxiliary-function-based IVA (AuxIVA). Experimental evaluations show that the proposed approach is effective and improves the source separation performance of IVA. In addition, several post-filters aiming to realize multichannel Wiener filter (MWF) are investigated. This setup proves to further increase the performance of IVA. The presented method shows a potential to provide a general way to improve the separation performance from single-channel source separation to multichannel source separation. Ana Ramírez López, Nobutaka Ono, Ulpu Remes, Kalle J. Palomäki, Mikko Kurimo |
ICASSP | 2 |
| 2015 | Fast DNN training based on auxiliary function techniqueabstractDeep neural networks (DNN) are typically optimized with stochastic gradient descent (SGD) using a fixed learning rate or an adaptive learning rate approach (ADAGRAD). In this paper, we introduce a new learning rule for neural networks that is based on an auxiliary function technique without parameter tuning. Instead of minimizing the objective function, a quadratic auxiliary function is recursively introduced layer by layer which has a closed-form optimum. We prove the monotonic decrease of the new learning rule. Our experiments show that the proposed algorithm converges faster and to a better local minimum than SGD. In addition, we propose a combination of the proposed learning rule and ADAGRAD which further accelerates convergence. Experimental evaluation on the MNIST database shows the benefit of the proposed approach in terms of digit recognition accuracy. Dung T. Tran, Nobutaka Ono, Emmanuel Vincent 0001 |
ICASSP | 2 |
| 2015 | Voice liveness detection algorithms based on pop noise caused by human breath for automatic speaker verificationabstractThis paper proposes a novel countermeasure framework to detect spoofing attacks to reduce the vulnerability of automatic speaker verification (ASV) systems. Recently, ASV systems have reached equivalent performances equivalent to those of other biometric modalities. However, spoofing techniques against these systems have also progressed drastically. Experimentation using advanced speech synthesis and voice conversion techniques has showed unacceptable false acceptance rates and several new countermeasure algorithms have been explored to detect spoofing materials accurately. However, the countermeasures proposed so far are based on the acoustic differences between natural speech signals and artificial speech signals, expected to become gradually smaller in the near future. In this paper, we focus on voice liveness detection, which aims to validate whether the presented speech signals originated from a live human. We use the phenomenon of pop noise, which is a distortion that happens when human breath reaches a microphone,as liveness evidence. This paper proposes pop noise detection algorithms and shows through an experimental study that they can be used to discriminate live voice signals from artificial ones generated by means of speech synthesis techniques. Sayaka Shiota, Fernando Villavicencio, Junichi Yamagishi, Nobutaka Ono, Isao Echizen, Tomoko Matsui |
INTERSPEECH | 4 |
| 2015 | Special issue on wireless acoustic sensor networks and ad hoc microphone arrays
Alexander Bertrand, Simon Doclo, Sharon Gannot, Nobutaka Ono, Toon van Waterschoot |
Signal Process. | 4 |
| 2015 | Blind compensation of interchannel sampling frequency mismatch for ad hoc microphone array based on maximum likelihood estimationabstractIn this paper, we propose a novel method for the blind compensation of drift for the asynchronous recording of an ad hoc microphone array. Digital signals simultaneously observed by different recording devices have drift of the time differences between the observation channels because of the sampling frequency mismatch among the devices. On the basis of a model in which the time difference is constant within each short time frame but varies in proportion to the central time of the frame, the effect of the sampling frequency mismatch can be compensated in the short-time Fourier transform (STFT) domain by a linear phase shift. By assuming that the sources are motionless and have stationary amplitudes, the observation is regarded as being stationary when drift does not occur. Thus, we formulate a likelihood to evaluate the stationarity in the STFT domain to evaluate the compensation of drift. The maximum likelihood estimation is obtained effectively by a golden section search. Using the estimated parameters, we compensate the drift by STFT analysis with a noninteger frame shift. The effectiveness of the proposed blind drift compensation method is evaluated in an experiment in which artificial drift is generated. Shigeki Miyabe, Nobutaka Ono, Shoji Makino |
Signal Process. | 2 |
| 2014 | Harmonic/percussive sound separation based on anisotropic smoothness of spectrogramsabstractThis paper describes a method to separate a monaural music signal into harmonic components e.g., a guitar and percussive components, e.g., a snare drum. Separation of these two components is a useful preprocessing for many music information retrieval applications, and in addition, it can be used as a new kind of music equalizer in itself, which enables a music listener to adjust the ratio of the volume of the guitar and the drum freely by themselves. Because of these potential applications, there have been many attempts to develop such a technique, especially in the last decade. However, some of the state-of-the-art techniques have a drawback that they are based on costly operations, such as the multiplications of large-sized matrix, Monte Carlo method, etc., which may constitute barriers to the practical use on some small computers such as smart phones. In this paper, an efficient method that does not depend on these costly operations is described. In formulating the methods, the authors basically assumed only the “anisotropic smoothness” of music spectrogram, which can be one of the minimalistic model that reflects the natures of these instruments. To be specific, the authors just assumed that harmonic instruments are smooth in time, while the percussive instruments are smooth in frequency on a music spectrogram. In this paper, on the basis of the assumption, source separation methods are formulated as optimization problems that optimize the “anisotropic smoothness” under some conditions. Because of the simplicity of the model, the derived algorithms are quite simple. Experimental results show that the methods were effective compared to a state-of-the-art technique, and the computation time was much shorter than an existing method; specifically, it can process a three-minute song in around 4-20 seconds on a laptop PC. Hideyuki Tachibana, Nobutaka Ono, Hirokazu Kameoka, Shigeki Sagayama |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Singing Voice Enhancement in Monaural Music Signals Based on Two-stage Harmonic/Percussive Sound Separation on Multiple Resolution SpectrogramsabstractWe propose a novel singing voice enhancement technique for monaural music audio signals, which is a quite challenging problem. Many singing voice enhancement techniques have been proposed recently. However, our approach is based on a quite different idea from these existing methods. We focused on the fluctuation of a singing voice and considered to detect it by exploiting two differently resolved spectrograms, one has rich temporal resolution and poor frequency resolution, while the other has rich frequency resolution and poor temporal resolution. On such two spectrograms, the shapes of fluctuating components are quite different. Based on this idea, we propose a singing voice enhancement technique that we call two-stage harmonic/percussive sound separation (HPSS). In this paper, we describe the details of two-stage HPSS and evaluate the performance of the method. The experimental results show that SDR, a commonly-used criterion on the task, was improved by around 4 dB, which is a considerably higher level than existing methods. In addition, we also evaluated the performance of the method as a preprocessing for melody estimation in music. The experimental results show that our singing voice enhancement technique considerably improved the performance of a simple pitch estimation technique. These results prove the effectiveness of the proposed method. Hideyuki Tachibana, Nobutaka Ono, Shigeki Sagayama |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2013 | Blind compensation of inter-channel sampling frequency mismatch with maximum likelihood estimation in STFT domainabstractThis paper proposes a novel blind compensation of sampling frequency mismatch for asynchronous microphone array. Digital signals simultaneously observed by different recording devices have drift of the time differences between the observation channels because of the sampling frequency mismatch among the devices. Based on the model that such the time difference is constant within each time frame, but varies proportional to the time frame index, the effect of the sampling frequency mismatch can be compensated in the short-time Fourier transform domain by the linear phase shift. By assuming the sources are motionless and stationary, a likelihood of the sampling frequency mismatch is formulated. The maximum likelihood estimation is obtained effectively by a golden section search. Shigeki Miyabe, Nobutaka Ono, Shoji Makino |
ICASSP | 2 |
| 2013 | Reversible Audio Information Hiding Based on Integer DCT Coefficients with Adaptive Hiding Locations
Xuping Huang, Nobutaka Ono, Isao Echizen, Akira Nishimura |
IWDW | 2 |
| 2012 | A tandem connectionist model using combination of multi-scale spectro-temporal features for acoustic event detectionabstractAcoustic event detection systems supporting heterogeneous sets of events face the problem of having to characterize them when they have different acoustic properties (transient, stationary, both, etc.), observing this fact even within the acoustic event itself. Moreover, managing large feature vectors with features characterizing different properties of the signal is always difficult. This paper introduces the usage of spectro-temporal fluctuation features in a tandem connectionist approach, modified to generate posterior features separately for each fluctuation scale and then combine the streams to be fed to a classic GMM-HMM model. The experiments explore scale and event wise performance, as well as different stream combination methods, and show that the proposed method outperforms the GMM-HMM baseline as well as recent proposals in the CHIL 2007 evaluation campaign's related acoustic event detection tasks. Miquel Espi, Masakiyo Fujimoto, Daisuke Saito, Nobutaka Ono, Shigeki Sagayama |
ICASSP | 4 |
| 2012 | User-guided independent vector analysis with source activity tuningabstractIn this paper, user-guided source separation based on independent vector analysis is presented. In this framework, temporal power variations of sources can be tuned by a user. The information is exploited as prior distributions of source activities in independent vector analysis with time-varying Gaussian model, and source signals are separated by maximum a posteriori (MAP) estimation. Experimental evaluations show the source activity tuning is much effective to improve the separation performance in hard mixing conditions such as long reverberation or level mismatch of sources. Takuma Ono, Nobutaka Ono, Shigeki Sagayama |
ICASSP | 2 |
| 2012 | Comparative evaluations of various harmonic/percussive sound separation algorithms based on anisotropic continuity of spectrogramabstractIn this paper, we explore several algorithms to find the best performing algorithm for harmonic and percussive sound separation (HPSS) based on anisotropic continuity of spectrogram through comparative evaluation of their experimental performance. Separating harmonic and percussive sounds is useful as a preprocessor for many music analysis purposes including chord estimation, rhythm analysis, and other music information retrieval tasks. We have introduced a method called “Harmonic/Percussive Sound Separation” (HPSS), that decomposes a music signal into two components by separating the spectrogram into horizontally-continuous and vertically-continuous components, which roughly correspond to harmonic and percussive sounds, respectively. Many possible ways exist to realize the HPSS algorithm based on this concept while it has been unknown which algorithm performs best. This paper describes the details of five different HPSS algorithms and compares their performances over real music signals. Hideyuki Tachibana, Hirokazu Kameoka, Nobutaka Ono, Shigeki Sagayama |
ICASSP | 3 |
| 2011 | Multichannel harmonic and percussive component separation by joint modeling of spatial and spectral continuityabstractThis paper considers the blind separation of the harmonic and percussive components of multichannel music signals. We model the contribution of each source to all mixture channels in the time-frequency domain via a spatial covariance matrix, which encodes its spatial characteristics, and a scalar spectral variance, which represents its spectral structure. We then exploit the spatial continuity and the different spectral continuity structures of harmonic and percussive components as prior information to derive maximum a posteriori (MAP) estimates of the parameters using the expectation-maximization (EM) algorithm. Experimental results over professional musical mixtures show the effectiveness of the proposed approach. Ngoc Q. K. Duong, Hideyuki Tachibana, Emmanuel Vincent 0001, Nobutaka Ono, Rémi Gribonval, Shigeki Sagayama |
ICASSP | 4 |
| 2011 | Automatic video annotation via Hierarchical Topic Trajectory Model considering cross-modal correlationsabstractWe propose a new statistical model, named Hierarchical Topic Trajectory Model (HTTM), for acquiring a dynamically changing topic model that represents the relationship between video frames and associated text labels. Model parameter estimation, annotation and retrieval can be executed within a unified framework with a few computation. It is also easy to add new modals such as audio signal and geotags. Preliminary experiments on video annotation task with manually annotated video dataset indicate that our proposed method can improve the annotation accuracy. Takuho Nakano, Akisato Kimura, Hirokazu Kameoka, Shigeki Miyabe, Shigeki Sagayama, Nobutaka Ono, Kunio Kashino, Takuya Nishimoto |
ICASSP | 6 |
| 2011 | Infinite-state spectrum model for music signal analysisabstractThis paper presents a nonparametric Bayesian extension of non-negative matrix factorization (NMF) for music signal analysis. Instrument sounds often exhibit non-stationary spectral characteristics. We introduce infinite-state spectral bases into NMF to represent time-varying spectra in polyphonic music signals. We describe our extension of NMF with infinite-state spectral bases generated by the Dirichlet process in a statistical framework, derive an efficient optimization algorithm based on collapsed variational inference, and validate the framework on audio data. Masahiro Nakano, Jonathan Le Roux, Hirokazu Kameoka, Nobutaka Ono, Shigeki Sagayama |
ICASSP | 4 |
| 2011 | Multipitch estimation by joint modeling of harmonic and transient soundsabstractMultipitch estimation techniques are widely used for music transcription and acquisition of musical data from digital signals. In this paper, we propose a flexible harmonic temporal timbre model to decompose the spectral energy of the signal in the time-frequency domain into individual pitched notes. Each note is modeled with a 2-dimensional Gaussian mixture. Unlike previous approaches, the proposed model is able to represent not only the harmonic partials but also the inharmonic attack of each note. We derive an Expectation-Maximization (EM) algorithm to estimate the parameters of this model and illustrate the higher performance of the proposed algorithm than NMF algorithm and HTC algorithm for the task of multipitch estimation over synthetic and real-world data. Emmanuel Vincent 0001, Stanislaw Andrzej Raczynski, Takuya Nishimoto, Nobutaka Ono, Shigeki Sagayama |
ICASSP | 5 |
| 2011 | Concurrent Optimization of Context Clustering and GMM for Offline Handwritten Word Recognition Using HMMabstractContext-dependent HMMs are commonly used in speech recognition. Parameter sharing needed for this model can be realized by two methods: context clustering or tied-mixture. In speech recognition, the former is reported to be more precise. However, there is some difficulty in applying context clustering to handwritten word recognition, since the distribution of each character is typically a mixture of different distributions, such as block-printed, cursive, etc. For this reason, successful results reported so far are limited to the tied-mixture approach. To deal with this problem, we propose a novel parameter tying method ``Partial Tied-Mixture", where the Gaussian Mixture Model (GMM) consists of a portion of all Gaussians. Furthermore, we derive a method to concurrently optimize context clustering and GMM. Experiments on the CEDAR database show that the proposed method outperforms tied-mixture both in terms of precision and computational cost. Tomoyuki Hamamura, Bunpei Irie, Takuya Nishimoto, Nobutaka Ono, Shigeki Sagayama |
ICDAR | 4 |
| 2011 | Using Spectral Fluctuation of Speech in Multi-Feature HMM-Based Voice Activity DetectionabstractObservation of speech spectrum leads to the fact that speech has a specific spectral fluctuation pattern both along time and frequency. In this paper, we integrate the usage of this nature in a multi-feature approach for voice activity detection. The effect of separating such specific spectral fluctuation using multi-stage HPSS (Harmonic-Percussive Sound Separation) has been analyzed over conventional features in voice activity detection, reducing frame-wise detection error by up to 78%, depending on the SNR conditions and noise type. The multi-feature approach has been tested using Hidden Markov Models to model the features stream as a sequence, which has out-performed standard and similar VAD proposals in utterance-based tests intended for automatic speech recognition. Miquel Espi, Shigeki Miyabe, Takuya Nishimoto, Nobutaka Ono, Shigeki Sagayama |
INTERSPEECH | 4 |
| 2011 | Computational auditory induction as a missing-data model-fitting problem with Bregman divergence
Jonathan Le Roux, Hirokazu Kameoka, Nobutaka Ono, Alain de Cheveigné, Shigeki Sagayama |
Speech Commun. | 3 |
| 2011 | Diffuse Noise Suppression Using Crystal-Shaped Microphone ArraysabstractThis paper describes novel methods for diffuse noise suppression using crystal-shaped microphone arrays. The two-stage processing of the observed signals by the Minimum Variance Distortionless Response (MVDR) beamformer and the subsequent Wiener post-filter is effective for diffuse noise suppression and gives the linear minimum mean square error (LMMSE) estimator of the target signal. It is essential in this framework to accurately estimate the short-time power spectrum and the steering vectors of the target signal from the noisy observations. Our methods diagonalize the spatial noise covariance matrix and utilizes the denoised off-diagonal entries of the spatial covariance matrix to accurately estimate the short-time power spectrum and the steering vectors of the target signal. We employ crystal arrays, certain classes of crystal-shaped array geometries, which make it possible to diagonalize the unknown noise covariance matrix by a constant unitary matrix regardless of its value as long as noise meets an isotropy condition. It is shown through experiments with simulated and real environmental noise that the proposed methods outperform previous methods substantially for real world noise and in the presence of reverberation. Nobutaka Ito, Hikaru Shimizu, Nobutaka Ono, Shigeki Sagayama |
IEEE Trans. Speech Audio Process. | 3 |
| 2011 | Beyond Timbral Statistics: Improving Music Classification Using Percussive Patterns and Bass LinesabstractThis paper discusses a new approach for clustering sequences of bar-long percussive and bass-line patterns in audio music collections and its application to genre classification. Many musical genres and styles are characterized by two kinds of distinct representative patterns, i.e., percussive patterns and bass-line patterns. So far, in most automatic genre classification systems, rhythmic and bass melody information has not been effectively used. In order to extract bar-long unit rhythmic patterns for a music collection, we propose a clustering method based on one-pass dynamic programming andk-means clustering. For clustering bass-line patterns, a method based onk-means clustering capable of handling pitch-shifting is proposed. After extracting these two fundamental kinds of patterns for each style/genre, feature vectors which are suitable for representing information about the patterns are proposed for supervised learning. Experimental results show that the automatically calculated rhythmic pattern information and bass pattern information can be used to effectively classify musical genre/style and improve upon current approaches based on timbral features. Emiru Tsunoo, George Tzanetakis, Nobutaka Ono, Shigeki Sagayama |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2010 | Designing the Wiener post-filter for diffuse noise suppression using imaginary parts of inter-channel cross-spectraabstractThis paper describes a new design of the Wiener post-filter for diffuse noise suppression. The Wiener post-filter is well-known as an effective post-processing of the minimum variance distortionless response beamformer, and its output is the optimal estimate of the target signal in the sense of the minimum mean square error. It is essential to accurately estimate the target power spectrum from the observed signals contaminated by noise when designing the Wiener post-filter. In our method, it is estimated from the imaginary parts of the inter-channel observation cross-spectra, under the assumption that the inter-channel noise cross-spectra are real-valued. The post-filter is designed using the estimate and this design is shown to be effective even for a small-sized array through experiments using simulated and real environmental noise. Nobutaka Ito, Nobutaka Ono, Emmanuel Vincent 0001, Shigeki Sagayama |
ICASSP | 2 |
| 2010 | A sparse component model of source signals and its application to blind source separationabstractIn this paper, we propose a new method of blind source separation (BSS) for music signals. Our method has the following characteristics: 1) the method is a combination of the sparseness-based model of source signals and the factorized basis model in nonnegative matrix factorization (NMF), 2) it is assumed that only one basis which structure source signals is active at each time-frequency bin of the observed signals, in order to degrade the degree of freedom, 3) parameter estimation algorithm is based on the EM algorithm regarding the index of the only one active basis as the hidden variable. We develop the formulation at a different point from NMF and show source separation performance in some simulation experiments. Yu Kitano, Hirokazu Kameoka, Yosuke Izumi, Nobutaka Ono, Shigeki Sagayama |
ICASSP | 4 |
| 2010 | R-means localization: A simple iterative algorithm for range-difference-based source localizationabstractIn this paper, we present a simple iterative algorithm for range-difference (RD) based localization. The RD-based localization is a kind of nonlinear optimization problem and generally it has no closed-form solution. Through auxiliary function approach, we derive iterative update rules without any tuning parameters, which just consists of 1) averaging source-sensor distances, 2) averaging the source positions estimated by updating source-sensor distance on each sensor with the source-direction fixed. Due to the resemblance of the iterative averaging to k-means clustering, we call it r-means localization. The convergence of the algorithm is guaranteed. The acceleration of the convergence is also investigated. Nobutaka Ono, Shigeki Sagayama |
ICASSP | 1 |
| 2010 | Melody line estimation in homophonic music audio signals based on temporal-variability of melodic sourceabstractEstimation of melody line in homophonic music audio signals is a challenging subject of study. Some of the difficulties are derived from presence of accompanying components. To overcome those difficulties, we propose a method to enhance melodic components in music audio signals. The enhancement algorithm uses fluctuation and shortness of melodic components, which we call temporal-variability. We also discuss a melody tracking algorithm, which can be simple thanks to the preprocessing. In this paper, we describe the enhancement method and tracking method, and show the experimental results that supports the efficiency of our methods. Hideyuki Tachibana, Takuma Ono, Nobutaka Ono, Shigeki Sagayama |
ICASSP | 3 |
| 2010 | Music mood classification by rhythm and bass-line unit pattern analysisabstractThis paper discusses an approach for the feature extraction for audio mood classification which is an important and tough problem in the field of music information retrieval (MIR). In this task the timbral information has been widely used, however many musical moods are characterized not only by timbral information but also by musical scale and temporal features such as rhythm patterns and bass-line patterns. In particular, modern music pieces mostly have certain fixed rhythm and bass-line patterns, and these patterns can characterize the impression of songs. We have proposed the extraction of rhythm and bass-line patterns, and these unit pattern analysis are combined with statistical feature extraction for mood classification. Experimental results show that the automatically calculated unit pattern information can be used to effectively classify musical mood. Emiru Tsunoo, Taichi Akase, Nobutaka Ono, Shigeki Sagayama |
ICASSP | 3 |
| 2010 | HMM-based approach for automatic chord detection using refined acoustic featuresabstractWe discuss an HMM-based method for detecting the chord sequence from musical acoustic signals using percussion-suppressed, Fourier-transformed chroma and delta-chroma features. To reduce the interference often caused by percussive sounds in popular music, we use Harmonic/Percussive Sound Separation (HPSS) technique to suppress percussive sounds and to emphasize harmonic sound components. We also use the Fourier transform of chroma to approximately diagonalize the covariance matrix of feature parameters so as to reduce the number of model parameters without degrading performance. It is shown that HMM with the new features yields higher recognition rates (the best in MIREX 2008 audio chord detection task) than that with conventional features. Yushi Ueda, Yuuki Uchiyama, Takuya Nishimoto, Nobutaka Ono, Shigeki Sagayama |
ICASSP | 4 |
| 2010 | Flexible Harmonic Temporal Structure for Modeling Musical Instrument
Yu Kitano, Takuya Nishimoto, Nobutaka Ono, Shigeki Sagayama |
ICEC | 4 |
| 2010 | Analysis on speech characteristics for robust voice activity detectionabstractThis paper discusses about effective speech characterization for off-line voice activity detection (VAD), which is an important step prior to speech data mining. Five different natures of speech are examined; energy, spectral shape, periodicity, phonetic variation, and spectral fluctuation, the latter observed from a new point of view. Specific spectral fluctuation patterns of speech have been analyzed using multi-stage Harmonic/Percussive Sound Separation algorithm. We compared the performance of the features, and various combinations, to evaluate their robustness in multiple noise environments. The combined approach outperformed the baseline of CENSREC-1-C evaluation framework. The results suggest that the proposed feature extraction approach can improve state of the art VAD methods. Miquel Espi, Shigeki Miyabe, Takuya Nishimoto, Nobutaka Ono, Shigeki Sagayama |
SLT | 4 |
| 2010 | Speech Spectrum Modeling for Joint Estimation of Spectral Envelope and Fundamental FrequencyabstractAlthough considerable effort has been devoted to both fundamental frequency (F0) and spectral envelope estimation in the field of speech processing, the problem of determiningF0and spectral envelopes has largely been tackled independently. IfF0were known in advance, then the spectral envelope could be estimated very reliably. On the other hand, if the spectral envelope were known in advance, then we could obtain a reliableF0estimate.F0and the spectral envelope, each of which is a prerequisite of the other, should thus be estimated jointly rather than independently in succession. On this basis, we develop a parametric speech spectrum model that allows us to estimate theF0and spectral envelope simultaneously. We confirmed experimentally the significant advantage of this joint estimation approach for bothF0estimation and spectral envelope estimation. Hirokazu Kameoka, Nobutaka Ono, Shigeki Sagayama |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | Complex NMF: A new sparse representation for acoustic signalsabstractThis paper presents a new sparse representation for acoustic signals which is based on a mixing model defined in the complex-spectrum domain (where additivity holds), and allows us to extract recurrent patterns of magnitude spectra that underlie observed complex spectra and the phase estimates of constituent signals. An efficient iterative algorithm is derived, which reduces to the multiplicative update algorithm for non-negative matrix factorization developed by Lee under a particular condition. Hirokazu Kameoka, Nobutaka Ono, Kunio Kashino, Shigeki Sagayama |
ICASSP | 2 |
| 2009 | Rhythm map: Extraction of unit rhythmic patterns and analysis of rhythmic structure from music acoustic signalsabstractThis paper discusses an approach to extract constituent percussive bar-long patterns in a music piece given as acoustic signal and to analyze the music structure with a map of constituent rhythmic patterns. Possible applications include music genre classification, music information retrieval (MIR) and music modification such as replacing rhythmic patterns with others. We propose a mathematical method based on One-pass DP algorithm and k-means clustering to extract unit percussive rhythmic patterns. As the result of identifying and localization the unit patterns in the entire piece, we obtained a music structure in the form of a map of rhythmic patterns. Emiru Tsunoo, Nobutaka Ono, Shigeki Sagayama |
ICASSP | 2 |
| 2009 | Audio genre classification using percussive pattern clustering combined with timbral featuresabstractMany musical genres and styles are characterized by distinct representative rhythmic patterns. In most automatic genre classification systems global statistical features based on timbral dynamics such as mel-frequency cepstral coefficients (MFCC) are utilized but so far rhythmic information has not so effectively been used. In order to extract bar-long unit rhythmic patterns for a music collection we propose a clustering method based on one-pass dynamic programming and k-means clustering. After extracting the fundamental rhythmic patterns for each style/genre a pattern occurrence histogram is calculated and used as a feature vector for supervised learning. Experimental results show that the automatically calculated rhythmic pattern information can be used to effectively classify musical genre/style and improve upon current approaches based on timbral features. Emiru Tsunoo, George Tzanetakis, Nobutaka Ono, Shigeki Sagayama |
ICME | 3 |
| 2009 | Stereo-input speech recognition using sparseness-based time-frequency masking in a reverberant environmentabstractWe present noise robust automatic speech recognition (ASR) using sparseness-based underdetermined blind source separation (BSS) technique. As a representative underdetermined BSS method, we utilized time-frequency masking in this paper. Although time-frequency masking is able to separate target speech from interferences effectively, one should consider two problems. One is that masking does not work well in noisy or reverberant environment. Another is that masking itself might cause some distortion of the target speech. For the former, we apply our time-frequency masking method [7] which can separate the target signal robustly even in noisy and reverberant environment. Next, investigating the distortion caused by time-frequency masking, we reveal following facts through experiments: 1) soft mask is better than binary mask in terms of recognition performance and 2) cepstral mean normalization (CMN) reduces the distortion, especially for that caused by soft mask. At the end, we evaluate the recognition performance of our method in noisy and reverberant real environment. Index Terms: time-frequency mask, speech sparseness, blind source separation, stereo-input, robust ASR Yosuke Izumi, Kenta Nishiki, Shinji Watanabe 0001, Takuya Nishimoto, Nobutaka Ono, Shigeki Sagayama |
INTERSPEECH | 5 |
| 2008 | A blind noise decorrelation approach with crystal arrays on designing post-filters for diffuse noise suppressionabstractThis paper describes a new framework for extracting the target signal in diffuse noise environments. We utilize crystal arrays, a certain class of symmetrical microphone arrays with crystal-like geometries, which enable interchannel decorrelation of isotropic noise without knowing the value of its covariance matrix. We refer to this decorrelation as blind noise decorrelation. Using an improved estimation of the signal power spectrum obtained by the blind noise decorrelation, the multichannel Wiener filter is properly implemented, which is the optimal estimator of the target signal in the minimum mean square error sense. Simulated experiments have shown the effectiveness of the proposed method. Nobutaka Ito, Nobutaka Ono, Shigeki Sagayama |
ICASSP | 2 |
| 2008 | Auxiliary function approach to parameter estimation of constrained sinusoidal model for monaural speech separationabstractWe introduce in this paper an auxiliary function approach to parameter estimation of the constrained sinusoidal model, which enables us to derive a complex-spectrum-domain EM-like multiple F0estimation algorithm. Through simulations, we evaluated the performance of the presented method in the ability to avoid locally optimal solutions. We implemented a monaural speech separation system based on the presented method and confirmed its performance on compound signals of real speech. Hirokazu Kameoka, Nobutaka Ono, Shigeki Sagayama |
ICASSP | 2 |
| 2008 | Harmonic-Temporal-Timbral Clustering (HTTC) for the analysis of multi-instrument polyphonic music signalsabstractIn this paper, we discuss a new approach named Harmonic-Temporal-Timbral Clustering (HTTC) for the analysis of single-channel audio signal of multi-instrument polyphonic music to estimate the pitch, onset timing, power and duration of all the acoustic events and to classify them into timbre categories simultaneously. Each acoustic event is modeled by a harmonic structure and a smooth envelope both represented by Gaussian mixtures. Based on the similarity between these spectro-temporal structures, timbres are clustered to form timbre categories. The entire process is mathematically formulated as a minimization problem for the I-divergence between the HTTC parametric model and the observed spectrogram of the music audio signal to simultaneously update harmonic, temporal and timbral model parameters through the EM algorithm. Some experimental results are presented to discuss the performance of the algorithm. Kenichi Miyamoto, Hirokazu Kameoka, Takuya Nishimoto, Nobutaka Ono, Shigeki Sagayama |
ICASSP | 4 |
| 2008 | Modulation analysis of speech through orthogonal FIR filterbank optimizationabstractNewborns must learn to structure incoming acoustic information into segments, words, phrases, etc., before they can start to learn language. This process is thought to rely on modulation structure of the speech waveform induced by segmental or prosodic regularities within the speech heard by the infant. Here, we investigate the process by which the initial acoustic processing required by modulation analysis can itself be tuned by exposure to the regularities of speech. Starting from the classic definition of modulation, as applied within channels of the peripheral filter, we formulate a mathematical framework in which the structure of initial spectral filtering is adapted for modulation analysis. Our working hypothesis is that the human ear and brain are adapted to the analysis of modulation, via a data-driven learning process on the scale of development (or possibly evolution). Simulation results are presented and a comparison with filterbanks classically used in signal processing is done. Jonathan Le Roux, Hirokazu Kameoka, Nobutaka Ono, Shigeki Sagayama, Alain de Cheveigné |
ICASSP | 3 |
| 2007 | Harmonic-Temporal Clustering of Speech for Single and Multiple F0 Contour Estimation in Noisy EnvironmentsabstractWe present in this paper a novel F0contour estimation method based on a parametric description of the wavelet power spectrum of speech that accounts for its structure simultaneously in time and frequency directions. We model the speech spectrum as a sequence of spectral clusters governed by a smooth common F0contour expressed as a spline curve. The harmonic and temporal structure of these clusters and their common F0contour are estimated simultaneously. Through experimental comparisons with existing methods, we show that our algorithm is competitive on clean single-speaker speech, and that it outperforms existing methods both in the presence of noise and for the estimation of multiple F0contours of cochannel concurrent speech. Jonathan Le Roux, Hirokazu Kameoka, Nobutaka Ono, Alain de Cheveigné, Shigeki Sagayama |
ICASSP (4) | 3 |
| 2007 | Sound Source Localization by Asymmetrically Arrayed 2ch Microphones on a SphereabstractIn this paper, we propose a novel system to localize a sound source in any 2D directions using only two microphones. In our system, the two microphones are asymmetrically placed on a sphere, thus, (1) the diffraction by the sphere and the asymmetrical arrangement of the microphones yield the localization cue including the front-back judgment, and (2) unlike the dummy head system, no previous measurements are necessary due to the analytical representation of the sphere diffraction. To deal with reverberation or ambient noises, we consider the maximum likelihood estimation of the direction of arrival with a diffused noise model on a sphere. We present a real system that we built through the investigation of the optimal microphone arrangement for speech, and give experimental results in real environment. Nobutaka Ono, Souichiro Fukamachi, Takuya Nishimoto, Shigeki Sagayama |
MMSP | 1 |
| 2007 | Single and Multiple F0 Contour Estimation Through Parametric Spectrogram Modeling of Speech in Noisy EnvironmentsabstractThis paper proposes a novel $F_{0}$ contour estimation algorithm based on a precise parametric description of the voiced parts of speech derived from the power spectrum. The algorithm is able to perform in a wide variety of noisy environments as well as to estimate the $F_{0}$ s of cochannel concurrent speech. The speech spectrum is modeled as a sequence of spectral clusters governed by a common $F_{0}$ contour expressed as a spline curve. These clusters are obtained by an unsupervised 2-D time-frequency clustering of the power density using a new formulation of the EM algorithm, and their common $F_{0}$ contour is estimated at the same time. A smooth $F_{0}$ contour is extracted for the whole utterance, linking together its voiced parts. A noise model is used to cope with nonharmonic background noise, which would otherwise interfere with the clustering of the harmonic portions of speech. We evaluate our algorithm in comparison with existing methods on several tasks, and show 1) that it is competitive on clean single-speaker speech, 2) that it outperforms existing methods in the presence of noise, and 3) that it outperforms existing methods for the estimation of multiple $F_{0}$ contours of cochannel concurrent speech. Jonathan Le Roux, Hirokazu Kameoka, Nobutaka Ono, Alain de Cheveigné, Shigeki Sagayama |
IEEE Trans. Speech Audio Process. | 3 |
| 2006 | Speech analyzer using a joint estimation model of spectral envelope and fine structureabstractWe have been working on a new speech analyzer based on a parametric representation of speech governed by the F0 parameter, towards practical human-machine interfaces. As a precise estimation of the frequency response of the vocal tract from a real speech signal requires the power of each component of the harmonic structure to be accurately estimated, one hopes to have a high-precision estimation of F0. At the same time, under the empirical constraint that speech spectral envelopes are usually smooth in the power domain, half pitch errors can be significantly avoided. Therefore, F0 and the envelope should be estimated jointly rather than separately through an optimal estimation of the spectral envelope and the spectral fine structure. In this article, we introduce a new speech analysis method using a spectral model with a composite function of envelope and fine structure models. Index Terms: parametric speech analyzer, speech synthesis, pitch estimation, spectral envelope estimation. Hirokazu Kameoka, Jonathan Le Roux, Nobutaka Ono, Shigeki Sagayama |
INTERSPEECH | 3 |
| 1999 | AM-FM extraction based on logarithmic differential decompositionabstractThe most important features of quasi-harmonic signals such as speech or musical sounds are the instantaneously varying amplitude (AM) and pitch (FM) of the harmonics. However, when we want to recognise them after subband decomposition like a human cochlea, the problem is that AM-FM components of the original harmonics are disturbed both by the filtering and the interference between harmonics. In this paper, we propose a method of logarithmic differential decomposition (LDD) and a new AM-FM extraction algorithm based of it. It extracts separately the conventional AM-FM which are common to all the harmonics and the interferometric AM-FM which are synchronous with pitch, hence can be an efficient clue to determine the fundamental. We investigate the role of those features by experiments for speech. Nobutaka Ono, Mototsugu Abe, Shigeru Ando |
MMSP | 1 |