Yuma Koizumi

dblp:149/0012 · DBLP profile ↗
← Back
35ranked-venue papers
16as first author
6since 2021 · last 2024
0000-0003-3645-6213ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 32 · 13 first-author · 6 since 2021Artificial intelligence and machine learning · 10 · 7 first-author · 4 since 2021
YearPublicationVenuePosition
2024 FLEURS-R: A Restored Multilingual Speech Corpus for Generation Tasks
Yuma Koizumi, Shigeki Karita, Heiga Zen, Jason Riesa, Haruko Ishikawa, Michiel Bacchiani
INTERSPEECH2
2023 LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus
Yuma Koizumi, Heiga Zen, Shigeki Karita, Yifan Ding 0004, Kohei Yatabe, Nobuyuki Morioka, Michiel Bacchiani, Yu Zhang 0033, Wei Han 0002, Ankur Bapna
INTERSPEECH1
2022 SNRi Target Training for Joint Speech Enhancement and Recognition
Yuma Koizumi, Shigeki Karita, Arun Narayanan, Sankaran Panchapagesan, Michiel Bacchiani
INTERSPEECH1
2022 SpecGrad: Diffusion Probabilistic Model based Neural Vocoder with Adaptive Noise Spectral Shaping
abstract
Neural vocoder using denoising diffusion probabilistic model (DDPM) has been improved by adaptation of the diffusion noise distribution to given acoustic features.In this study, we propose SpecGrad that adapts the diffusion noise so that its timevarying spectral envelope becomes close to the conditioning log-mel spectrogram.This adaptation by time-varying filtering improves the sound quality especially in the high-frequency bands.It is processed in the time-frequency domain to keep the computational cost almost the same as the conventional DDPMbased neural vocoders.Experimental results showed that Spec-Grad generates higher-fidelity speech waveform than conventional DDPM-based neural vocoders in both analysis-synthesis and speech enhancement scenarios.Audio demos are available at wavegrad.github.io/specgrad/.
Yuma Koizumi, Heiga Zen, Kohei Yatabe, Nanxin Chen, Michiel Bacchiani
INTERSPEECH1
2022 Wavefit: an Iterative and Non-Autoregressive Neural Vocoder Based on Fixed-Point Iteration
abstract
Denoising diffusion probabilistic models (DDPMs) and generative adversarial networks (GANs) are popular generative models for neural vocoders. The DDPMs and GANs can be characterized by the iterative denoising framework and adversarial training, respectively. This study proposes a fast and high-quality neural vocoder called WaveFit, which integrates the essence of GANs into a DDPM-like iterative framework based on fixed-point iteration. WaveFit iteratively denoises an input signal, and trains a deep neural network (DNN) for minimizing an adversarial loss calculated from intermediate outputs at all iterations. Subjective (side-by-side) listening tests showed no statistically significant differences in naturalness between human natural speech and those synthesized by WaveFit with five iterations. Furthermore, the inference speed of WaveFit was more than 240 times faster than WaveRNN. Audio demos are available at google.github.io/df-conformer/wavefit/.
Yuma Koizumi, Kohei Yatabe, Heiga Zen, Michiel Bacchiani
SLT1
2022 Learning Mask Scalars for Improved Robust Automatic Speech Recognition
abstract
Improving robustness of streaming automatic speech recognition (ASR) systems using neural network based acoustic frontends is challenging because of causality constraints and the speech-distortions introduced by the frontend. Time-frequency masking based approaches are commonly used, but they need additional hyperparameters – mask scalars – to limit distortion. Mask scalars are typically hand-tuned and chosen conservatively. In this work, we present a technique to predict mask scalars using ASR loss in an end-to-end fashion, with minimal increase in model size and complexity. We evaluate the approach on two robust ASR tasks: multichannel enhancement in the presence of speech and non-speech noise, and acoustic echo cancellation (AEC). Results show that the presented algorithm consistently improves word error rate (WER) over strong baselines that use hand-tuned hyperparameters: up to 16% in noisy conditions, and up to 7% for AEC.
Arun Narayanan, James Walker, Sankaran Panchapagesan, Nathan Howard, Yuma Koizumi
SLT5
2020 Sound Event Detection by Multitask Learning of Sound Events and Scenes with Soft Scene Labels
abstract
Sound event detection (SED) and acoustic scene classification (ASC) are major tasks in environmental sound analysis. Considering that sound events and scenes are closely related to each other, some works have addressed joint analyses of sound events and acoustic scenes based on multitask learning (MTL), in which the knowledge of sound events and scenes can help in estimating them mutually. The conventional MTL-based methods utilize one-hot scene labels to train the relationship between sound events and scenes; thus, the conventional methods cannot model the extent to which sound events and scenes are related. However, in the real environment, common sound events may occur in some acoustic scenes; on the other hand, some sound events occur only in a limited acoustic scene. In this paper, we thus propose a new method for SED based on MTL of SED and ASC using the soft labels of acoustic scenes, which enable us to model the extent to which sound events and scenes are related. Experiments conducted using TUT Sound Events 2016/2017 and TUT Acoustic Scenes 2016 datasets show that the proposed method improves the SED performance by 3.80% in F-score compared with conventional MTL-based SED.
Keisuke Imoto, Noriyuki Tonami, Yuma Koizumi, Masahiro Yasuda, Ryosuke Yamanishi, Yoichi Yamashita
ICASSP3
2020 Stable Training of Dnn for Speech Enhancement Based on Perceptually-Motivated Black-Box Cost Function
abstract
Improving subjective sound quality of enhanced signals is one of the most important missions in speech enhancement. For evaluating the subjective quality, several methods related to perceptually-motivated objective sound quality assessment (OSQA) have been proposed such as PESQ (perceptual evaluation of speech quality). However, direct use of such measures for training deep neural network (DNN) is not allowed in most cases because popular OSQAs are non-differentiable with respect to DNN parameters. Therefore, the previous study has proposed to approximate the score of OS-QAs by an auxiliary DNN so that its gradient can be used for training the primary DNN. One problem with this approach is instability of the training caused by the approximation error of the score. To overcome this problem, we propose to use stabilization techniques borrowed from reinforcement learning. The experiments, aimed to increase the score of PESQ as an example, show that the proposed method (i) can stably train a DNN to increase PESQ, (ii) achieved the state-of-the-art PESQ score on a public dataset, and (iii) resulted in better sound quality than conventional methods based on subjective evaluation.
Masaki Kawanaka, Yuma Koizumi, Ryoichi Miyazaki, Kohei Yatabe
ICASSP2
2020 Speech Enhancement Using Self-Adaptation and Multi-Head Self-Attention
abstract
This paper investigates a self-adaptation method for speech enhancement using auxiliary speaker-aware features; we extract a speaker representation used for adaptation directly from the test utterance. Conventional studies of deep neural network (DNN)-based speech enhancement mainly focus on building a speaker independent model. Meanwhile, in speech applications including speech recognition and synthesis, it is known that model adaptation to the target speaker improves the accuracy. Our research question is whether a DNN for speech enhancement can be adopted to unknown speakers without any auxiliary guidance signal in test-phase. To achieve this, we adopt multi-task learning of speech enhancement and speaker identification, and use the output of the final hidden layer of speaker identification branch as an auxiliary feature. In addition, we use multi-head self-attention for capturing long-term dependencies in the speech and noise. Experimental results on a public dataset show that our strategy achieves the state-of-the-art performance and also outperform conventional methods in terms of subjective quality.
Yuma Koizumi, Kohei Yatabe, Marc Delcroix, Yoshiki Masuyama, Daiki Takeuchi
ICASSP1
2020 SPIDERnet: Attention Network For One-Shot Anomaly Detection In Sounds
abstract
We propose a similarity function for one-shot anomaly detection in sounds (ADS) called SPecific anomaly IDentifiER network (SPIDERnet). In ADS systems, since overlooking an anomaly may result in serious incidents, we need to update such systems using an (often only one) overlooked anomalous sample. A previous study proposed the use of memory-based one-shot learning. A problem with this previous method is that it can detect only short anomalous sounds such as collision sounds because its similarity function is based on a naive mean-squared-error between the input and memorized spectrogram. To detect various anomalous sounds, SPIDERnet consists of (i) a neural network-based feature extractor for measuring similarity in embedded space and (ii) attention mechanisms for absorbing time-frequency stretching. Experimental results on two public datasets indicate that SPIDERnet outperforms conventional methods and robustly detects various anomalous sounds.
Yuma Koizumi, Masahiro Yasuda, Shin Murata, Shoichiro Saito, Hisashi Uematsu, Noboru Harada
ICASSP1
2020 Phase Reconstruction Based On Recurrent Phase Unwrapping With Deep Neural Networks
abstract
Phase reconstruction, which estimates phase from a given amplitude spectrogram, is an active research field in acoustical signal processing with many applications including audio synthesis. To take advantage of rich knowledge from data, several studies presented deep neural network (DNN)–based phase reconstruction methods. However, the training of a DNN for phase reconstruction is not an easy task because phase is sensitive to the shift of a waveform. To overcome this problem, we propose a DNN-based two-stage phase reconstruction method. In the proposed method, DNNs estimate phase derivatives instead of phase itself, which allows us to avoid the sensitivity problem. Then, phase is recursively estimated based on the estimated derivatives, which is named recurrent phase unwrapping (RPU). The experimental results confirm that the proposed method outperformed the direct phase estimation by a DNN.
Yoshiki Masuyama, Kohei Yatabe, Yuma Koizumi, Yasuhiro Oikawa, Noboru Harada
ICASSP3
2020 Real-Time Speech Enhancement Using Equilibriated RNN
abstract
We propose a speech enhancement method using a causal deep neural network (DNN) for real-time applications. DNN has been widely used for estimating a time-frequency (T-F) mask which enhances a speech signal. One popular DNN structure for that is a recurrent neural network (RNN) owing to its capability of effectively modelling time-sequential data like speech. In particular, the long short-term memory (LSTM) is often used to alleviate the vanishing/exploding gradient problem which makes the training of an RNN difficult. However, the number of parameters of LSTM is increased as the price of mitigating the difficulty of training, which requires more computational resources. For real-time speech enhancement, it is preferable to use a smaller network without losing the performance. In this paper, we propose to use the equilibriated recurrent neural network (ERNN) for avoiding the vanishing/exploding gradient problem without increasing the number of parameters. The proposed structure is causal, which requires only the information from the past, in order to apply it in real-time. Compared to the uni- and bi-directional LSTM networks, the proposed method achieved the similar performance with much fewer parameters.
Daiki Takeuchi, Kohei Yatabe, Yuma Koizumi, Yasuhiro Oikawa, Noboru Harada
ICASSP3
2020 Invertible DNN-Based Nonlinear Time-Frequency Transform for Speech Enhancement
abstract
We propose an end-to-end speech enhancement method with trainable time-frequency (T-F) transform based on invertible deep neural network (DNN). The resent development of speech enhancement is brought by using DNN. The ordinary DNN-based speech enhancement employs T-F transform, typically the short-time Fourier transform (STFT), and estimates a T-F mask using DNN. On the other hand, some methods have considered end-to-end networks which directly estimate the enhanced signals without T-F transform. While end-to-end methods have shown promising results, they are black boxes and hard to understand. Therefore, some end-to-end methods used a DNN to learn the linear T-F transform which is much easier to understand. However, the learned transform may not have a property important for ordinary signal processing. In this paper, as the important property of the T-F transform, perfect reconstruction is considered. An invertible nonlinear T-F transform is constructed by DNNs and learned from data so that the obtained transform is perfectly reconstructing filterbank.
Daiki Takeuchi, Kohei Yatabe, Yuma Koizumi, Yasuhiro Oikawa, Noboru Harada
ICASSP3
2020 Sound Event Localization Based on Sound Intensity Vector Refined by Dnn-Based Denoising and Source Separation
abstract
We propose a direction-of-arrival (DOA) estimation method for Sound Event Localization and Detection (SELD). Direct estimation of DOA using a deep neural network (DNN), i.e. completely-datadriven approach, achieves high accuracy. However, there is a gap in the accuracy between DOA estimation for single and overlapping sources because they cannot incorporate physical knowledge. Meanwhile, although the accuracy of physics-based approaches is inferior to DNN-based approaches, it is robust for overlapping-source. In this study, we consider a combination of physics-based and DNN-based approaches; the sound intensity vectors (IVs) for physics-based DOA estimation is refined based on DNN-based denoising and source separation. This method enables the accurate DOA estimation for both single and overlapping sources using a spherical microphone array. Experimental results show that the proposed method achieves state-of-the-art DOA estimation accuracy on an open dataset of the SELD.
Masahiro Yasuda, Yuma Koizumi, Shoichiro Saito, Hisashi Uematsu, Keisuke Imoto
ICASSP2
2020 A Transformer-Based Audio Captioning Model with Keyword Estimation
abstract
One of the problems with automated audio captioning (AAC) is the indeterminacy in word selection corresponding to the audio event/scene.Since one acoustic event/scene can be described with several words, it results in a combinatorial explosion of possible captions and difficulty in training.To solve this problem, we propose a Transformer-based audio-captioning model with keyword estimation called TRACKE.It simultaneously solves the word-selection indeterminacy problem with the main task of AAC while executing the sub-task of acoustic event detection/acoustic scene classification (i.e., keyword estimation).TRACKE estimates keywords, which comprise a word set corresponding to audio events/scenes in the input audio, and generates the caption while referring to the estimated keywords to reduce word-selection indeterminacy.Experimental results on a public AAC dataset indicate that TRACKE achieved state-ofthe-art performance and successfully estimated both the caption and its keywords.
Yuma Koizumi, Ryo Masumura, Kyosuke Nishida, Masahiro Yasuda, Shoichiro Saito
INTERSPEECH1
2020 Listen to What You Want: Neural Network-Based Universal Sound Selector
abstract
Being able to control the acoustic events (AEs) to which we want to listen would allow the development of more controllable hearable devices.This paper addresses the AE sound selection (or removal) problems, that we define as the extraction (or suppression) of all the sounds that belong to one or multiple desired AE classes.Although this problem could be addressed with a combination of source separation followed by AE classification, this is a sub-optimal way of solving the problem.Moreover, source separation usually requires knowing the maximum number of sources, which may not be practical when dealing with AEs.In this paper, we propose instead a universal sound selection neural network that enables to directly select AE sounds from a mixture given user-specified target AE classes.The proposed framework can be explicitly optimized to simultaneously select sounds from multiple desired AE classes, independently of the number of sources in the mixture.We experimentally show that the proposed method achieves promising AE sound selection performance and could be generalized to mixtures with a number of sources that are unseen during training.
Tsubasa Ochiai, Marc Delcroix, Yuma Koizumi, Hiroaki Ito, Keisuke Kinoshita, Shoko Araki
INTERSPEECH3
2020 Crossmodal Sound Retrieval Based on Specific Target Co-Occurrence Denoted with Weak Labels
Masahiro Yasuda, Yasunori Ohishi, Yuma Koizumi, Noboru Harada
INTERSPEECH3
2019 A Two-class Hyper-spherical Autoencoder for Supervised Anomaly Detection
abstract
Supervised anomaly detection has been a tough problem due to its necessity of special handling of unseen anomalies. In this paper, we present a heuristic implementation of variational auto-encoder with von-Mises Fisher prior applied to a supervised anomaly detector. The closed latent space like sphere is suitable for detecting unseen anomalies because we have a possibility to "fill" the space with seen training samples. If it ideally works, the reconstruction error will be high for all unseen anomalies. Experiments show that our model can separate normal and anomaly samples in the spherical latent space. It is also shown that he proposed model improves the performance for seen anomalies without degrading the performance for unseen anomalies.
Yuta Kawachi, Yuma Koizumi, Shin Murata, Noboru Harada
ICASSP2
2019 Trainable Adaptive Window Switching for Speech Enhancement
abstract
This study proposes a trainable adaptive window switching (AWS) method and apply it to a deep-neural-network (DNN) for speech enhancement in the modified discrete cosine transform domain. Time-frequency (T-F) mask processing in the short-time Fourier transform (STFT)-domain is a typical speech enhancement method. To recover the target signal precisely, DNN-based short-time frequency transforms have recently been investigated and used instead of the STFT. However, since such a fixed-resolution short-time frequency transform method has a T-F resolution problem based on the uncertainty principle, not only the short-time frequency transform but also the length of the windowing function should be optimized. To overcome this problem, we incorporate AWS into the speech enhancement procedure, and the windowing function of each time-frame is manipulated using a DNN depending on the input signal. We confirmed that the proposed method achieved a higher signal-to-distortion ratio than conventional speech enhancement methods in fixed-resolution frequency domains.
Yuma Koizumi, Noboru Harada, Youichi Haneda
ICASSP1
2019 SNIPER: Few-shot Learning for Anomaly Detection to Minimize False-negative Rate with Ensured True-positive Rate
abstract
In anomaly detection systems, overlooking anomalies may result in serious incidents. Thus, when a system overlooks an anomaly, we need to update the system to never overlook the observed type of anomalies twice. There are roughly two possible approaches to solve this problem; re-training the whole system using all training data, or cascading a new specific detector for the overlooked anomaly. The first approach is the most effective solution; however, a huge computational cost and an amount of anomalous training data are required to re-train the system when it consists of a deep-learning-based anomaly detector. We focused on the latter approach and propose a training method for a cascaded specific anomaly detector using few-shot (just 1 to 3) samples. To suppress the false-negative rate of the overlooked anomaly, the proposed method works to decrease the false-positive rate under the constraint of true-positive rate equaling 1. Experimental results show that the proposed method outperformed conventional cross-entropy-based few-shot learning methods.
Yuma Koizumi, Shin Murata, Noboru Harada, Shoichiro Saito, Hisashi Uematsu
ICASSP1
2019 Deep Griffin-Lim Iteration
abstract
This paper presents a novel phase reconstruction method (only from a given amplitude spectrogram) by combining a signal-processing-based approach and a deep neural network (DNN). To retrieve a time-domain signal from its amplitude spectrogram, the corresponding phase is required. One of the popular phase reconstruction methods is the Griffin-Lim algorithm (GLA), which is based on the redundancy of the short-time Fourier transform. However, GLA often involves many iterations and produces low-quality signals owing to the lack of prior knowledge of the target signal. In order to address these issues, in this study, we propose an architecture which stacks a sub-block including two GLA-inspired fixed layers and a DNN. The number of stacked sub-blocks is adjustable, and we can trade the performance and computational load based on requirements of applications. The effectiveness of the proposed method is investigated by reconstructing phases from amplitude spectrograms of speeches.
Yoshiki Masuyama, Kohei Yatabe, Yuma Koizumi, Yasuhiro Oikawa, Noboru Harada
ICASSP3
2019 Data-driven Design of Perfect Reconstruction Filterbank for DNN-based Sound Source Enhancement
abstract
We propose a data-driven design method of perfect-reconstruction filterbank (PRFB) for sound-source enhancement (SSE) based on deep neural network (DNN). DNNs have been used to estimate a time-frequency (T-F) mask in the short-time Fourier transform (STFT) domain. Their training is more stable when a simple cost function as mean-squared error (MSE) is utilized comparing to some advanced cost such as objective sound quality assessments. However, such a simple cost function inherits strong assumptions on the statistics of the target and/or noise which is often not satisfied, and the mismatch of assumption results in degraded performance. In this paper, we propose to design the frequency scale of PRFB from training data so that the assumption on MSE is satisfied. For designing the frequency scale, the warped filterbank frame (WFBF) is considered as PRFB. The frequency characteristic of learned WFBF was in between STFT and the wavelet transform, and its effectiveness was confirmed by comparison with a standard STFT-based DNN whose input feature is compressed into the mel scale.
Daiki Takeuchi, Kohei Yatabe, Yuma Koizumi, Yasuhiro Oikawa, Noboru Harada
ICASSP3
2019 AdaFlow: Domain-adaptive Density Estimator with Application to Anomaly Detection and Unpaired Cross-domain Translation
abstract
We tackle unsupervised anomaly detection (UAD), a problem of detecting data that significantly differ from normal data. UAD is typically solved by using density estimation. Recently, deep neural network (DNN)-based density estimators, such as Normalizing Flows, have been attracting attention. However, one of their drawbacks is the difficulty in adapting them to the change in the normal data's distribution. To address this difficulty, we propose AdaFlow, a new DNN-based density estimator that can be easily adapted to the change of the distribution. AdaFlow is a unified model of a Normalizing Flow and Adaptive Batch-Normalizations, a module that enables DNNs to adapt to new distributions. AdaFlow can be adapted to a new distribution by just conducting forward propagation once per sample; hence, it can be used on devices that have limited computational resources. We have confirmed the effectiveness of the proposed model through an anomaly detection in a sound task. We also propose a method of applying AdaFlow to the unpaired cross-domain translation problem, in which one has to train a cross-domain translation model with only unpaired samples. We have confirmed that our model can be used for the cross-domain translation problem through experiments on image datasets.
Masataka Yamaguchi, Yuma Koizumi, Noboru Harada
ICASSP2
2019 Unsupervised Detection of Anomalous Sound Based on Deep Learning and the Neyman-Pearson Lemma
abstract
This paper proposes a novel optimization principle and its implementation for unsupervised anomaly detection in sound (ADS) using an autoencoder (AE). The goal of the unsupervised-ADS is to detect unknown anomalous sounds without training data of anomalous sounds. The use of an AE as a normal model is a state-of-the-art technique for the unsupervised-ADS. To decrease the false positive rate (FPR), the AE is trained to minimize the reconstruction error of normal sounds, and the anomaly score is calculated as the reconstruction error of the observed sound. Unfortunately, since this training procedure does not take into account the anomaly score for anomalous sounds, the true positive rate (TPR) does not necessarily increase. In this study, we define an objective function based on the Neyman-Pearson lemma by considering the ADS as a statistical hypothesis test. The proposed objective function trains the AE to maximize the TPR under an arbitrary low FPR condition. To calculate the TPR in the objective function, we consider that the set of anomalous sounds is the complementary set of normal sounds and simulate anomalous sounds by using a rejection sampling algorithm. Through experiments using synthetic data, we found that the proposed method improved the performance measures of the ADS under low FPR conditions. In addition, we confirmed that the proposed method could detect anomalous sounds in real environments.
Yuma Koizumi, Shoichiro Saito, Hisashi Uematsu, Yuta Kawachi, Noboru Harada
IEEE ACM Trans. Audio Speech Lang. Process.1
2018 Complementary Set Variational Autoencoder for Supervised Anomaly Detection
abstract
Anomalies have broad patterns corresponding to their causes. In industry, anomalies are typically observed as equipment failures. Anomaly detection aims to detect such failures as anomalies. Although this is usually a binary classification task, the potential existence of unseen (unknown) failures makes this task difficult. Conventional supervised approaches are suitable for detecting seen anomalies but not for unseen anomalies. Although, unsupervised neural networks for anomaly detection now detect unseen anomalies well, they cannot utilize anomalous data for detecting seen anomalies even if some data have been made available. Thus, providing an anomaly detector that finds both seen and unseen anomalies well is still a tough problem. In this paper, we introduce a novel probabilistic representation of anomalies to solve this problem. The proposed model defines the normal and anomaly distributions using the analogy between a set and the complementary set. We applied these distributions to an unsupervised variational autoencoder (VAE)-based method and turned it into a supervised VAE-based method. We tested the proposed method with well-known data and real industrial data to show that the proposed method detects seen anomalies better than the conventional unsupervised method without degrading the detection performance for unseen anomalies.
Yuta Kawachi, Yuma Koizumi, Noboru Harada
ICASSP2
2018 End-to-End Sound Source Enhancement Using Deep Neural Network in the Modified Discrete Cosine Transform Domain
abstract
This paper presents an end-to-end deep neural network (DNN)-based source enhancement on the basis of a time-frequency (T-F) mask processing in the modified discrete cosine transform (MDCT)-domain. To retrieve the target signal perfectly in the discrete Fourier transform (DFT)-domain, both amplitude and phase of the spectrum need to be manipulated. However, since it is difficult to deal with complex values by neural network straightforward way, a real-valued T-F mask is commonly estimated and only amplitude spectrum is manipulated. In this study, we use the MDCT instead of the DFT and estimate real-valued T-F masks in the MDCT-domain. The perfect retrieval can be achieved by manipulating only the real-valued MDCT-spectra. To reduce time-domain aliasing arises from manipulating the MDCT spectrum, we build an end-to-end DNN-based source enhancement using T-F mask and train the DNN to minimize an objective function defined in the time-domain. In experiments using several kinds of objective sound quality scores, we observed that the scores were significantly improved.
Yuma Koizumi, Noboru Harada, Youichi Haneda, Yusuke Hioka, Kazunori Kobayashi
ICASSP1
2018 DNN-Based Source Enhancement to Increase Objective Sound Quality Assessment Score
abstract
We propose a training method for deep neural network (DNN) based source enhancement to increase objective sound quality assessment (OSQA) scores such as the perceptual evaluation of speech quality. In many conventional studies, DNNs have been used as a mapping function to estimate time-frequency masks and trained to minimize an analytically tractable objective function such as the mean squared error (MSE). Since OSQA scores have been used widely for sound-quality evaluation, constructing DNNs to increase OSQA scores would be better than using the minimum MSE to create high-quality output signals. However, since most OSQA scores are not analytically tractable, i.e., they are black boxes, the gradient of the objective function cannot be calculated by simply applying backpropagation. To calculate the gradient of the OSQA-based objective function, we formulated a DNN optimization scheme on the basis of black-box optimization, which is used for training a computer that plays a game. For a black-box-optimization scheme, we adopt the policy gradient method for calculating the gradient on the basis of a sampling algorithm. To simulate output signals using the sampling algorithm, DNNs are used to estimate the probability density function of the output signals that maximize OSQA scores. The OSQA scores are calculated from the simulated output signals, and the DNNs are trained to increase the probability of generating the simulated output signals that achieve high OSQA scores. Through several experiments, we found that OSQA scores significantly increased by applying the proposed method, even though the MSE was not minimized.
Yuma Koizumi, Kenta Niwa, Yusuke Hioka, Kazunori Kobayashi, Youichi Haneda
IEEE ACM Trans. Audio Speech Lang. Process.1
2017 DNN-based source enhancement self-optimized by reinforcement learning using sound quality measurements
abstract
We investigated whether a deep neural network (DNN)-based source enhancement function can be self-optimized by reinforcement learning (RL). The use of a DNN is a powerful approach to describing the relationship between two sets of variables and can be useful for source enhancement function design. By training the DNN using a huge amount of training data, sound quality of output signals are improved. However, collecting a huge amount of training data is often difficult in practice. To use limited training data efficiently, we focus on the “self-optimization” of DNN-based source enhancement function in which RL is commonly utilized in the development of game playing computers. As a reward for RL, quantitative metrics that reflect a human's perceptual score (perceptual score), e.g., perceptual evaluation methods for audio source separation (PEASS), are utilized. To investigate whether the sound quality is improved by RL-based source enhancement, subjective tests were conducted. It was confirmed that the output sound quality of the RL-based source enhancement function improved as the number of iterations was increased and finally outperformed the conventional method.
Yuma Koizumi, Kenta Niwa, Yusuke Hioka, Kazunori Kobayashi, Youichi Haneda
ICASSP1
2017 Supervised source enhancement composed of nonnegative auto-encoders and complementarity subtraction
abstract
A method for constructing deep neural networks (DNNs) for accurate supervised source enhancement is proposed. Attempts were made in previous studies to estimate the power spectral densities (PSDs) of sound sources, which are used to estimate Wiener filters for source enhancement, from the output of multiple beamformings using DNNs. Although performance improved, it was not possible to guarantee accurate PSD estimation since the trained DNNs were treated as black boxes. The proposed DNN construction method uses non-negative auto-encoders and complementarity subtraction. This study also reveals that auto-encoders whose weights are non-negative correspond to non-negative matrix factorization (NMF), which decomposes source PSDs into non-negative spectral bases and their activations. It further introduces a complementarity subtraction method for estimating PSDs accurately. Through several experiments, it was confirmed that the signal-to-interference plus noise ratio improved by approximately 12 dB for datasets captured in various noisy/reverberant rooms.
Kenta Niwa, Yuma Koizumi, Tomoko Kawase, Kazunori Kobayashi, Yusuke Hioka
ICASSP2
2017 On relationships between amplitude and phase of short-time Fourier transform
abstract
The relationships between the amplitude and phase of the short-time Fourier transform (STFT) are investigated. By choosing the Gaussian window for the STFT, we reveal that the group delay and instantaneous frequency of each signal segment, both of which are derived from the phase by definition, can also be explicitly linked with the amplitude. As a result, the amplitude and phase can also be linked through the group delay or instantaneous frequency without making any assumptions for the phase property of the target signals, e.g., minimum, maximum, or linear phase. The theoretical basis is also confirmed in numerical simulations.
Suehiro Shimauchi, Shinya Kudo, Yuma Koizumi, Ken'ichi Furuya
ICASSP3
2017 Informative Acoustic Feature Selection to Maximize Mutual Information for Collecting Target Sources
abstract
An informative acoustic-feature-selection method for collecting target sources in noisy environments is proposed. Wiener filtering is a powerful framework for sound-source enhancement. For Wiener-filter estimation, statistical-mapping functions, such as deep neural network based or Gaussian mixture model based mappings, have been used. In this framework, it is essential to find informative acoustic features that provide effective cues for Wiener-filter estimation. In this study, we measured the informativeness of acoustic features using mutual information between acoustic features and supervised Wiener-filter parameters, e.g., prior signal-to-noise ratios, and developed a method for automatically selecting informative acoustic features from a large number of feature candidates. To automatically select optimum features, we derived a differentiable objective function in proportion to mutual information based on the kernel method. Since the higher order correlations between acoustic features and Wiener-filter parameters are calculated using the kernel method, the statistical dependence of these variables is accurately calculated; thus, only meaningful acoustic features are selected. Through several experiments conducted on a mock sports field, we confirmed that the signal-to-distortion ratio score improved when various types of target sources were surrounded by loud cheering noise.
Yuma Koizumi, Kenta Niwa, Yusuke Hioka, Kazunori Kobayashi, Hitoshi Ohmuro
IEEE ACM Trans. Audio Speech Lang. Process.1
2016 Integrated approach of feature extraction and sound source enhancement based on maximization of mutual information
abstract
We investigated informative acoustic feature extraction based on dimension reduction for collecting target sources on a noisy sports field. Although a Wiener filter is often used for sound source enhancement, it is difficult to accurately design the Wiener filter by simply using spatial cues because the noise on a sports field (e.g., cheering from spectators) arrives from the same direction as that of the targeted source. A statistical approach is used to estimate the Wiener filter by using pre-trained acoustic feature models. However, an informative acoustic feature, which provides a powerful clue for clear extraction of the target source, is unknown. For this study, we developed a method for optimizing a projection matrix for dimension reduction by maximizing the mutual information between acoustic features and the Wiener filter. Through experiments using two-directional microphones on a mock sports field, we confirmed that the proposed method outperformed previous methods in terms of both the noise reduction and quality of the recovered sound sources.
Yuma Koizumi, Kenta Niwa, Yusuke Hioka, Kazunori Kobayashi, Hitoshi Ohmuro
ICASSP1
2016 Pinpoint extraction of distant sound source based on DNN mapping from multiple beamforming outputs to prior SNR
abstract
We propose a method for estimating the prior signal-to-noise ratio (SNR), which is used for calculating the Wiener filter for distant sound source extraction, from output signals of beamforming using statistical mapping based on the deep neural network (DNN). Since informative features to estimate the prior SNR are included in multiple beamforming outputs, the SNR can be accurately estimated by this mapping using the DNN. The proposed method was applied to a large microphone array, the design of which was optimized to form effective directivity patterns to extract distant sound sources. Experimental results proved that the target source was clearly extracted with the proposed method.
Kenta Niwa, Yuma Koizumi, Tomoko Kawase, Kazunori Kobayashi, Yusuke Hioka
ICASSP2
2016 Binaural sound generation corresponding to omnidirectional video view using angular region-wise source enhancement
abstract
Web applications for watching omnidirectional video through head-mounted displays (HMDs) or smartphones have been widely distributed. The goal of this study was to generate binaural sounds corresponding to the user viewpoint. Assuming that a microphone array is used for sound recording, the enhanced signal for each angular region can be extracted. By convolving head-related transfer functions (HRTFs) and enhanced signals and re-synthesizing them, binaural sounds corresponding to the user viewpoint can be virtually generated. In this paper, we propose a method for achieving angular region-wise source enhancement by generating a multichannel Wiener filter based on the power spectral density (PSD)-estimation-in-beamspace method. To measure user localization when watching omnidirectional video through an HMD, we used a system that enables the generation of binaural sounds corresponding to the user viewpoint in real time. Through subjective tests, we confirmed that sound localization corresponding to the user viewpoint can be obtained when applying about a 40-degree angular region-wise source enhancement.
Kenta Niwa, Yuma Koizumi, Kazunori Kobayashi, Hisashi Uematsu
ICASSP2
2014 Intra-note segmentation via sticky HMM with DP emission
abstract
This paper presents an intra-note segmentation method for mono-phonic recordings based on acoustic feature variation; each musical note is separated into onset, steady and offset states. The task of intra-note segmentation from audio signals is detecting change points of acoustic feature. In proposed method, the Markov process is assumed on state transition, and time-varying acoustic feature is represented by three Dirichlet processes (DP) that are emitted by the each state. In order to express the generative process, the sticky hidden Markov model (HMM) with DP emission is employed. This modeling allows us to automatically estimate the state transition while avoiding the model selection problem by assuming countably infinite of possible acoustic feature in musical notes. Experimental result shows that the detection accuracy of onset-to-steady and steady-to-offset were improved 2.3 points and 20.7 points from previous method, respectively.
Yuma Koizumi, Katunobu Itou
ICASSP1