Kohei Yatabe

dblp:148/9899 · DBLP profile ↗
← Back
53ranked-venue papers
7as first author
25since 2021 · last 2026
0000-0002-1345-0663ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 47 · 6 first-author · 20 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 10 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 Stride conversion algorithms for convolutional layers and its application to sampling-frequency-independent deep neural networks
abstract
We propose interpolation-based algorithms that enable convolutional and transposed convolutional layers to operate with arbitrary (including non-integer) strides. A primary motivation for the proposed algorithms is to maintain a consistent temporal resolution when adapting deep neural networks (DNNs) to different sampling frequencies (SFs). To handle untrained SFs, we previously introduced SF-independent (SFI) convolutional layers, which adjust kernel weights in accordance with the target SF. However, achieving full consistency across SFs also requires the proportional adjustment of the stride, which results in non-integer values in many practical cases. Conventional algorithms for convolutional layers cannot handle such strides directly, and commonly used approaches (e.g., stride rounding or signal resampling) lead to performance degradation. To solve this problem, we propose a feature-domain interpolation framework that constructs continuous-time representations of intermediate features. This enables sampling at arbitrary stride intervals without modifying the network architecture. Through music source separation experiments, we show that the proposed algorithms maintain a strong performance across a range of SFs, including those where the stride becomes non-integer. Our analysis reveals that the proposed algorithms are robust to the choice of interpolation method and are especially effective for sources containing pitched sounds.
Kanami Imamura, Tomohiko Nakamura, Norihiro Takamune, Kohei Yatabe, Hiroshi Saruwatari
Signal Process.4
2024 Determined BSS by Combination of IVA and DNN via Proximal Average
abstract
This paper proposes a novel approach for determined blind source separation (BSS) assisted by deep neural network (DNN). Determined BSS algorithms, including independent vector analysis (IVA), separate source signals from multi-channel mixtures by estimating demixing filters according to their source models. Our method realizes a combined source model based on IVA and DNN in terms of proximal average. This combination allows our BSS algorithm to incorporate a single-channel denoising DNN without any specialized architecture/training, and hence the proposed method can circumvent the difficulty of training a DNN tailored for a determined BSS algorithm. Our experimental results show that the proposed method can stably and consistently improve the separation performance by combining a DNN trained with a denoising task.
Kazuki Matsumoto, Kohei Yatabe
ICASSP2
2024 Harmonic/Percussive Source Separation Based on Anisotropic Smoothness of Magnitude Spectrograms via Convex Optimization
abstract
Harmonic/percussive source separation (HPSS) is an important tool for analyzing and processing audio signals. The standard approach to HPSS takes advantage of the structural difference of sinusoidal and percussive components, calledanisotropic smoothness, in magnitude spectrograms. However, the existing methods disregard phase of the spectrograms and/or approximate the problem, which naturally limits the upper bound of the performance of HPSS. In this letter, we propose a novel approach to HPSS that regards phase without the approximation. The proposed method introduces an auxiliary variable that acts as an adaptive weight of a weighted energy minimization problem, which enables us to apply smoothing on magnitude of complex-valued spectrograms. Compared to the existing methods, the proposed method can obtain separated components having better magnitude and phase by simultaneously handling them.
Natsuki Akaishi, Koki Yamada, Kohei Yatabe
IEEE Signal Process. Lett.3
2024 PHAIN: Audio Inpainting via Phase-Aware Optimization With Instantaneous Frequency
abstract
Audio inpainting restores locally corrupted parts of digital audio signals. Sparsity-based methods achieve this by promoting sparsity in the time-frequency (T-F) domain, assuming short-time audio segments consist of a few sinusoids. However, such sparsity promotion reduces the magnitudes of the resulting waveforms; moreover, it often ignores the temporal connections of sinusoidal components. To address these problems, we propose a novel phase-aware audio inpainting method. Our method minimizes the time variations of a particular T-F representation calculated using the time derivative of the phase. This promotes sinusoidal components that coherently fit in the corrupted parts without directly suppressing the magnitudes. Both objective and subjective experiments confirmed the superiority of the proposed method compared with state-of-the-art methods.
Tomoro Tanaka, Kohei Yatabe, Yasuhiro Oikawa
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Improving Phase-Vocoder-Based Time Stretching by Time-Directional Spectrogram Squeezing
abstract
Time stretching of music signals has a crucial problem, i.e., smearing of percussive sounds. Some time stretching algorithms have addressed this problem by detecting percussive components and manipulating them differently from the other components. However, conventional methods cause artifacts. In this paper, to prevent percussion smearing, we propose a preprocessing for time stretching. The proposed algorithm aims to preserve time scale of percussive components while stretching the rest of components in the ordinary way. To do so, time-frequency bins dominated by percussive components are squeezed in time direction so as to preserve the shape of spectrogram of percussive components. Our experiment showed that our method could improve sound quality for long stretching.
Natsuki Akaishi, Kohei Yatabe, Yasuhiro Oikawa
ICASSP2
2023 Synthesizing Speech from ECoG with a Combination of Transformer-Based Encoder and Neural Vocoder
abstract
This paper reports on a novel invasive brain–computer interface (BCI) paradigm that has successfully reconstructed spoken sentences from invasive electrocorticogram (ECoG) signals using deep-neural-network-based encoders and a pre-trained neural vocoder. We recorded ECoG signals while 13 participants were speaking short sentences. Our BCI could map the ECoG recording to the log-mel spectrograms of the spoken sentences using a bidirectional long short-term memory (BLSTM) or a Transformer. The estimated log-mel spectrograms were used in Parallel WaveGAN to synthesize speech waveforms. An evaluation of the model performance revealed that the Transformer model significantly outperformed (Wilcoxon signed-rank test, p < 0.001) the BLSTM in terms of mean square error loss and Pearson correlation.
Kai Shigemi, Shuji Komeiji, Takumi Mitsuhashi, Yasushi Iimura, Hiroharu Suzuki, Hidenori Sugano, Koichi Shinoda, Kohei Yatabe, Toshihisa Tanaka 0001
ICASSP8
2023 UPGLADE: Unplugged Plug-and-Play Audio Declipper Based on Consensus Equilibrium of DNN and Sparse Optimization
abstract
In this paper, we propose a novel audio declipping method that fuses sparse-optimization-based and deep neural network (DNN)– based methods. The two methods have contrasting characteristics, depending on clipping level. Sparse-optimization-based audio de-clipping can preserve reliable samples, being suitable for precise restoration of small clipping. Besides, DNN-based methods are potent for recovering large clipping thanks to their data-driven approaches. Therefore, if these two methods are properly combined, audio declipping effective for a wide range of clipping levels can be realized. In the proposed method, we use a framework called consensus equilibrium to fuse the above two methods. Our experiments confirmed that the proposed method was superior to both conventional sparse-optimization-based and DNN-based methods.
Tomoro Tanaka, Kohei Yatabe, Yasuhiro Oikawa
ICASSP2
2023 LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus
Yuma Koizumi, Heiga Zen, Shigeki Karita, Yifan Ding 0004, Kohei Yatabe, Nobuyuki Morioka, Michiel Bacchiani, Yu Zhang 0033, Wei Han 0002, Ankur Bapna
INTERSPEECH5
2023 Versatile Time-Frequency Representations Realized by Convex Penalty on Magnitude Spectrogram
abstract
Sparse time-frequency (T-F) representations have been an important research topic for more than several decades. Among them, optimization-based methods (in particular, exten-sions of basis pursuit) allow us to design the representations through objective functions. Since acoustic signal processing uti-lizes models of spectrogram, the flexibility of optimization-based T-F representations is helpful for adjusting the representation for each application. However, acoustic applications often require models of magnitude of T-F representations obtained by discrete Gabor transform (DGT). Adjusting a T-F representation to such a magnitude model (e.g., smoothness of magnitude of DGT coefficients) results in a non-convex optimization problem that is difficult to solve. In this paper, instead of tackling difficult non-convex problems, we propose a convex optimization-based framework that realizes a T-F representation whosemagnitudehas characteristics specified by the user. We analyzed the prop-erties of the proposed method and provide numerical examples of sparse T-F representations having, e.g., low-rank or smooth magnitude, which have not been realized before.
Keidai Arai, Koki Yamada, Kohei Yatabe
IEEE Signal Process. Lett.3
2023 Online Phase Reconstruction via DNN-Based Phase Differences Estimation
abstract
This paper presents a two-stage online phase reconstruction framework using causal deep neural networks (DNNs). Phase reconstruction is a task of recovering phase of the short-time Fourier transform (STFT) coefficients only from the corresponding magnitude. However, phase is sensitive to waveform shifts and not easy to estimate from the magnitude even with a DNN. To overcome this problem, we propose to use DNNs for estimating differences of phase between adjacent time-frequency bins. We show that convolutional neural networks are suitable for phase difference estimation, according to the theoretical relation between partial derivatives of STFT phase and magnitude. The estimated phase differences are used for reconstructing phase by solving a weighted least squares problem in a frame-by-frame manner. In contrast to existing DNN-based phase reconstruction methods, the proposed framework is causal and does not require any iterative procedure. The experiments showed that the proposed method outperforms existing online methods and a DNN-based method for phase reconstruction.
Yoshiki Masuyama, Kohei Yatabe, Kento Nagatomo, Yasuhiro Oikawa
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 Harmonic and Percussive Sound Separation Based on Mixed Partial Derivative of Phase Spectrogram
abstract
Harmonic and percussive sound separation (HPSS) is a widely applied pre-processing tool that extracts distinct (harmonic and percussive) components of a signal. In the previous methods, HPSS has been performed based on the structural properties of magnitude (or power) spectrograms. However, such approach does not take advantage of phase that contains useful information of the waveform. In this paper, we propose a novel HPSS method named MipDroP that relies only on phase and does not use information of magnitude spectrograms. The proposed MipDroP algorithm effectively examines phase through its mixed partial derivative and constructs a pair of masks for the separation. Our experiments showed that MipDroP can extract percussive components better than the other methods.
Natsuki Akaishi, Kohei Yatabe, Yasuhiro Oikawa
ICASSP2
2022 Acoustic Application of Phase Reconstruction Algorithms in Optics
abstract
Phase reconstruction from amplitude spectrograms has attracted attention in recent acoustics because of its potential applications in speech synthesis and enhancement. The most well-known algorithm in acoustics is based on alternating projection and called Griffin– Lim algorithm (GLA). At the same time, GLA is known as the Gerchberg–Saxton algorithm in optics, and a lot of its variants have been proposed independently of those in acoustics. In this paper, we propose to apply phase reconstruction algorithms developed in the optics community to acoustic applications and evaluate them using acoustical metrics. Specifically, we propose to apply the averaged alternating reflections (AAR), relaxed AAR (RAAR), and hybrid input-output (HIO) algorithms to acoustic signals. Our experimental results suggested that RAAR has enough potential for acoustic applications because it clearly outperformed GLA.
Tomoki Kobayashi, Tomoro Tanaka, Kohei Yatabe, Yasuhiro Oikawa
ICASSP3
2022 Wearable Seld Dataset: Dataset For Sound Event Localization And Detection Using Wearable Devices Around Head
abstract
Sound event localization and detection (SELD) is a combined task of identifying the sound event and its direction. Deep neural networks (DNNs) are utilized to associate them with the sound signals observed by a microphone array. Although ambisonic microphones are popular in the literature of SELD, they might limits the range of applications due to their predetermined geometry. Some applications (including those for pedestrians that perform SELD while walking) require a wearable microphone array whose geometry can be designed to suit the task. In this paper, for development of such a wearable SELD, we propose a dataset named Wearable SELD dataset. It consists of data recorded by 24 microphones placed on a head and torso simulators (HATS) with some accessories mimicking wearable devices (glasses, earphones, and headphones). We also provide experimental results of SELD using the proposed dataset and SELDNet to investigate the effect of microphone configuration.
Kento Nagatomo, Masahiro Yasuda, Kohei Yatabe, Shoichiro Saito, Yasuhiro Oikawa
ICASSP3
2022 APPLADE: Adjustable Plug-and-Play Audio Declipper Combining DNN with Sparse Optimization
abstract
In this paper, we propose an audio declipping method that takes advantages of both sparse optimization and deep learning. Since sparsity-based audio declipping methods have been developed upon constrained optimization, they are adjustable and well-studied in theory. However, they always uniformly promote sparsity and ignore the individual properties of a signal. Deep neural network (DNN)– based methods can learn the properties of target signals and use them for audio declipping. Still, they cannot perform well if the training data have mismatches and/or constraints in the time domain are not imposed. In the proposed method, we use a DNN in an optimization algorithm. It is inspired by an idea called plug-and-play (PnP) and enables us to promote sparsity based on the learned information of data, considering constraints in the time domain. Our experiments confirmed that the proposed method is stable and robust to mismatches between training and test data.
Tomoro Tanaka, Kohei Yatabe, Masahiro Yasuda, Yasuhiro Oikawa
ICASSP2
2022 An objective test tool for pitch extractors' response attributes
abstract
We propose an objective measurement method for pitch extractors' responses to frequency-modulated signals.It enables us to evaluate different pitch extractors with unified criteria.The method uses extended time-stretched pulses combined by binary orthogonal sequences.It provides simultaneous measurement results consisting of the linear and the non-linear timeinvariant responses and random and time-varying responses.We tested representative pitch extractors using fundamental frequencies spanning 80 Hz to 800 Hz with 1/48 octave steps and produced more than 2000 modulation frequency response plots.We found that making scientific visualization by animating these plots enables us to understand different pitch extractors' behavior at once.Such efficient and effortless inspection is impossible by inspecting all individual plots.The proposed measurement method with visualization leads to further improvement of the performance of one of the extractors mentioned above.In other words, our procedure turns the specific pitch extractor into the best reliable measuring equipment that is crucial for scientific research.We open-sourced MATLAB codes of the proposed objective measurement method and visualization procedure.
Hideki Kawahara, Kohei Yatabe, Ken-Ichi Sakakibara, Tatsuya Kitamura, Hideki Banno, Masanori Morise
INTERSPEECH2
2022 SpecGrad: Diffusion Probabilistic Model based Neural Vocoder with Adaptive Noise Spectral Shaping
abstract
Neural vocoder using denoising diffusion probabilistic model (DDPM) has been improved by adaptation of the diffusion noise distribution to given acoustic features.In this study, we propose SpecGrad that adapts the diffusion noise so that its timevarying spectral envelope becomes close to the conditioning log-mel spectrogram.This adaptation by time-varying filtering improves the sound quality especially in the high-frequency bands.It is processed in the time-frequency domain to keep the computational cost almost the same as the conventional DDPMbased neural vocoders.Experimental results showed that Spec-Grad generates higher-fidelity speech waveform than conventional DDPM-based neural vocoders in both analysis-synthesis and speech enhancement scenarios.Audio demos are available at wavegrad.github.io/specgrad/.
Yuma Koizumi, Heiga Zen, Kohei Yatabe, Nanxin Chen, Michiel Bacchiani
INTERSPEECH3
2022 Wavefit: an Iterative and Non-Autoregressive Neural Vocoder Based on Fixed-Point Iteration
abstract
Denoising diffusion probabilistic models (DDPMs) and generative adversarial networks (GANs) are popular generative models for neural vocoders. The DDPMs and GANs can be characterized by the iterative denoising framework and adversarial training, respectively. This study proposes a fast and high-quality neural vocoder called WaveFit, which integrates the essence of GANs into a DDPM-like iterative framework based on fixed-point iteration. WaveFit iteratively denoises an input signal, and trains a deep neural network (DNN) for minimizing an adversarial loss calculated from intermediate outputs at all iterations. Subjective (side-by-side) listening tests showed no statistically significant differences in naturalness between human natural speech and those synthesized by WaveFit with five iterations. Furthermore, the inference speed of WaveFit was more than 240 times faster than WaveRNN. Audio demos are available at google.github.io/df-conformer/wavefit/.
Yuma Koizumi, Kohei Yatabe, Heiga Zen, Michiel Bacchiani
SLT2
2022 Sampling-Frequency-Independent Convolutional Layer and its Application to Audio Source Separation
abstract
Audio source separation is often used for the preprocessing of various tasks, and one of its ultimate goals is to construct a single versatile preprocessor that can handle every variety of audio signal. One of the most important varieties of the discrete-time audio signal is sampling frequency. Since it is usually task-specific, the versatile preprocessor must handle all the sampling frequencies required by the possible downstream tasks. However, conventional models based on deep neural networks (DNNs) are not designed for handling a variety of sampling frequencies. Thus, for unseen sampling frequencies, they may not work appropriately. In this paper, we propose sampling-frequency-independent (SFI) convolutional layers capable of handling various sampling frequencies. The core idea of the proposed layers comes from our finding that a convolutional layer can be viewed as a collection of digital filters and inherently depends on sampling frequency. To overcome this dependency, we propose an SFI structure that features analog filters and generates weights of a convolutional layer from the analog filters. By utilizing time- and frequency-domain analog-to-digital filter conversion techniques, we can adapt the convolutional layer for various sampling frequencies. As an example application, we construct an SFI version of a conventional source separation network. Through music source separation experiments, we show that the proposed layers enable separation networks to consistently work well for unseen sampling frequencies in objective and perceptual separation qualities. We also demonstrate that the proposed method outperforms a conventional method based on signal resampling when the sampling frequencies of input signals are significantly lower than the trained sampling frequency.
Koichi Saito, Tomohiko Nakamura, Kohei Yatabe, Hiroshi Saruwatari
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 Cascaded All-Pass Filters with Randomized Center Frequencies and Phase Polarity for Acoustic and Speech Measurement and Data Augmentation
abstract
We introduce a new member of TSP (Time Stretched Pulse) for acoustic and speech measurement infrastructure, based on a simple all-pass filter and systematic randomization. This new infrastructure fundamentally upgrades our previous measurement procedure, which enables simultaneous measurement of multiple attributes, including non-linear ones without requiring extra filtering nor post-processing. Our new proposal establishes a theoretically solid, flexible, and extensible foundation in acoustic measurement. Moreover, it is general enough to provide versatile research tools for other fields, such as biological signal analysis. We illustrate using acoustic measurements and data augmentation as representative examples among various prospective applications. We open-sourced MATLAB implementation. It consists of an interactive and real-time acoustic tool, MATLAB functions, and supporting materials.
Hideki Kawahara, Kohei Yatabe
ICASSP2
2021 Sparse Time-Frequency Representation Via Atomic Norm Minimization
abstract
Nonstationary signals are commonly analyzed and processed in the time-frequency (T-F) domain that is obtained by the discrete Gabor transform (DGT). The T-F representation obtained by DGT is spread due to windowing, which may degrade the performance of T-F domain analysis and processing. To obtain a well-localized T-F representation, sparsity-aware methods using ℓ1-norm have been studied. However, they need to discretize a continuous parameter onto a grid, which causes a model mismatch. In this paper, we propose a method of estimating a sparse T-F representation using atomic norm. The atomic norm enables sparse optimization without discretization of continuous parameters. Numerical experiments show that the T-F representation obtained by the proposed method is sparser than the conventional methods.
Tsubasa Kusano, Kohei Yatabe, Yasuhiro Oikawa
ICASSP2
2021 Linear Multichannel Blind Source Separation based on Time-Frequency Mask Obtained by Harmonic/Percussive Sound Separation
abstract
Determined blind source separation (BSS) extracts the source signals by linear multichannel filtering. Its performance depends on the accuracy of source modeling, and hence existing BSS methods have proposed several source models. Recently, a new determined BSS algorithm that incorporates a time-frequency mask has been proposed. It enables very flexible source modeling because the model is implicitly defined by a mask-generating function. Building up on this framework, in this paper, we propose a unification of determined BSS and harmonic/percussive sound separation (HPSS). HPSS is an important preprocessing for musical applications. By incorporating HPSS, both harmonic and percussive instruments can be accurately modeled for determined BSS. The resultant algorithm estimates the demixing filter using the information obtained by an HPSS method. We also propose a stabilization method that is essential for the proposed algorithm. Our experiments showed that the proposed method outperformed both HPSS and determined BSS methods including independent low-rank matrix analysis.
Soichiro Oyabu, Daichi Kitamura, Kohei Yatabe
ICASSP3
2021 Mixture of Orthogonal Sequences Made from Extended Time-Stretched Pulses Enables Measurement of Involuntary Voice Fundamental Frequency Response to Pitch Perturbation
abstract
Auditory feedback plays an essential role in the regulation of the fundamental frequency of voiced sounds. The fundamental frequency also responds to auditory stimulation other than the speaker's voice. We propose to use this response of the fundamental frequency of sustained vowels to frequency-modulated test signals for investigating involuntary control of voice pitch. This involuntary response is difficult to identify and isolate by the conventional paradigm, which uses step-shaped pitch perturbation. We recently developed a versatile measurement method using a mixture of orthogonal sequences made from a set of extended time-stretched pulses (TSP). In this article, we extended our approach and designed a set of test signals using the mixture to modulate the fundamental frequency of artificial signals. For testing the response, the experimenter presents the modulated signal aurally while the subject is voicing sustained vowels. We developed a tool for conducting this test quickly and interactively. We make the tool available as an open-source and also provide executable GUI-based applications. Preliminary tests revealed that the proposed method consistently provides compensatory responses with about 100 ms latency, representing involuntary control. Finally, we discuss future applications of the proposed method for objective and non-invasive auditory response measurements.
Hideki Kawahara, Toshie Matsui, Kohei Yatabe, Ken-Ichi Sakakibara, Minoru Tsuzaki, Masanori Morise, Toshio Irino
Interspeech3
2021 Interactive and Real-Time Acoustic Measurement Tools for Speech Data Acquisition and Presentation: Application of an Extended Member of Time Stretched Pulses
Hideki Kawahara, Kohei Yatabe, Ken-Ichi Sakakibara, Mitsunori Mizumachi, Masanori Morise, Hideki Banno, Toshio Irino
Interspeech2
2021 Gamma Boltzmann Machine for Audio Modeling
abstract
This paper presents an energy-based probabilistic model that handles nonnegative data in consideration of both linear and logarithmic scales. In audio applications, magnitude of time-frequency representation, including spectrogram, is regarded as one of the most important features. Such magnitude-based features have been extensively utilized in learning-based audio processing. Since a logarithmic scale is important in terms of auditory perception, the features are usually computed with a logarithmic function. That is, a logarithmic function is applied within the computation of features so that a learning machine does not have to explicitly model the logarithmic scale. We think in a different way and propose a restricted Boltzmann machine (RBM) that simultaneously models linear- and log-magnitude spectra. RBM is a stochastic neural network that can discover data representations without supervision. To manage both linear and logarithmic scales, we define an energy function based on both scales. This energy function results in a conditional distribution (of the observable data, given hidden units) that is written as the gamma distribution, and hence the proposed RBM is termed gamma-Bernoulli RBM. The proposed gamma-Bernoulli RBM was compared to the ordinary Gaussian-Bernoulli RBM by speech representation experiments. Both objective and subjective evaluations illustrated the advantage of the proposed model.
Toru Nakashika, Kohei Yatabe
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 Determined BSS Based on Time-Frequency Masking and Its Application to Harmonic Vector Analysis
abstract
This paper proposes harmonic vector analysis (HVA) based on a general algorithmic framework of audio blind source separation (BSS) that is also presented in this paper. BSS for a convolutive audio mixture is usually performed by multichannel linear filtering when the numbers of microphones and sources are equal (determined situation). This paper addresses such determined BSS based on batch processing. To estimate the demixing filters, effective modeling of the source signals is important. One successful example is independent vector analysis (IVA) that models the signals via co-occurrence among the frequency components in each source. To give more freedom to the source modeling, a general framework of determined BSS is presented in this paper. It is based on the plug-and-play scheme using a primal-dual splitting algorithm and enables us to model the source signals implicitly through a time-frequency mask. By using the proposed framework, determined BSS algorithms can be developed by designing masks that enhance the source signals. As an example of its application, we propose HVA by defining a time-frequency mask that enhances the harmonic structure of audio signals via sparsity of cepstrum. The experiments showed that HVA outperforms IVA and independent low-rank matrix analysis (ILRMA) for both speech and music signals. A MATLAB code is provided along with the paper for a reference.
Kohei Yatabe, Daichi Kitamura
IEEE ACM Trans. Audio Speech Lang. Process.1
2020 Stable Training of Dnn for Speech Enhancement Based on Perceptually-Motivated Black-Box Cost Function
abstract
Improving subjective sound quality of enhanced signals is one of the most important missions in speech enhancement. For evaluating the subjective quality, several methods related to perceptually-motivated objective sound quality assessment (OSQA) have been proposed such as PESQ (perceptual evaluation of speech quality). However, direct use of such measures for training deep neural network (DNN) is not allowed in most cases because popular OSQAs are non-differentiable with respect to DNN parameters. Therefore, the previous study has proposed to approximate the score of OS-QAs by an auxiliary DNN so that its gradient can be used for training the primary DNN. One problem with this approach is instability of the training caused by the approximation error of the score. To overcome this problem, we propose to use stabilization techniques borrowed from reinforcement learning. The experiments, aimed to increase the score of PESQ as an example, show that the proposed method (i) can stably train a DNN to increase PESQ, (ii) achieved the state-of-the-art PESQ score on a public dataset, and (iii) resulted in better sound quality than conventional methods based on subjective evaluation.
Masaki Kawanaka, Yuma Koizumi, Ryoichi Miyazaki, Kohei Yatabe
ICASSP4
2020 Speech Enhancement Using Self-Adaptation and Multi-Head Self-Attention
abstract
This paper investigates a self-adaptation method for speech enhancement using auxiliary speaker-aware features; we extract a speaker representation used for adaptation directly from the test utterance. Conventional studies of deep neural network (DNN)-based speech enhancement mainly focus on building a speaker independent model. Meanwhile, in speech applications including speech recognition and synthesis, it is known that model adaptation to the target speaker improves the accuracy. Our research question is whether a DNN for speech enhancement can be adopted to unknown speakers without any auxiliary guidance signal in test-phase. To achieve this, we adopt multi-task learning of speech enhancement and speaker identification, and use the output of the final hidden layer of speaker identification branch as an auxiliary feature. In addition, we use multi-head self-attention for capturing long-term dependencies in the speech and noise. Experimental results on a public dataset show that our strategy achieves the state-of-the-art performance and also outperform conventional methods in terms of subjective quality.
Yuma Koizumi, Kohei Yatabe, Marc Delcroix, Yoshiki Masuyama, Daiki Takeuchi
ICASSP2
2020 Maximally Energy-Concentrated Differential Window for Phase-Aware Signal Processing Using Instantaneous Frequency
abstract
The short-time Fourier transform (STFT) is widely employed in non-stationary signal analysis, whose property depends on window functions. Instantaneous frequency in STFT, the time-derivative of phase, is recently applied to many applications including spectrogram reassignment. The computation of instantaneous frequency requires STFT with the window and STFT with the (time-)differential window, i.e., the computation of instantaneous frequency depends on both the window function and its time derivative. To obtain the instantaneous frequency accurately, the sidelobe of frequency response of differential window should be reduced because the side-lobe causes mixing of multiple components. In this paper, we propose window functions suitable for computing the instantaneous frequency which are designed based on minimizing the sidelobe energy of the frequency response of the differential window.
Tsubasa Kusano, Kohei Yatabe, Yasuhiro Oikawa
ICASSP2
2020 Phase Reconstruction Based On Recurrent Phase Unwrapping With Deep Neural Networks
abstract
Phase reconstruction, which estimates phase from a given amplitude spectrogram, is an active research field in acoustical signal processing with many applications including audio synthesis. To take advantage of rich knowledge from data, several studies presented deep neural network (DNN)–based phase reconstruction methods. However, the training of a DNN for phase reconstruction is not an easy task because phase is sensitive to the shift of a waveform. To overcome this problem, we propose a DNN-based two-stage phase reconstruction method. In the proposed method, DNNs estimate phase derivatives instead of phase itself, which allows us to avoid the sensitivity problem. Then, phase is recursively estimated based on the estimated derivatives, which is named recurrent phase unwrapping (RPU). The experimental results confirm that the proposed method outperformed the direct phase estimation by a DNN.
Yoshiki Masuyama, Kohei Yatabe, Yuma Koizumi, Yasuhiro Oikawa, Noboru Harada
ICASSP2
2020 Real-Time Speech Enhancement Using Equilibriated RNN
abstract
We propose a speech enhancement method using a causal deep neural network (DNN) for real-time applications. DNN has been widely used for estimating a time-frequency (T-F) mask which enhances a speech signal. One popular DNN structure for that is a recurrent neural network (RNN) owing to its capability of effectively modelling time-sequential data like speech. In particular, the long short-term memory (LSTM) is often used to alleviate the vanishing/exploding gradient problem which makes the training of an RNN difficult. However, the number of parameters of LSTM is increased as the price of mitigating the difficulty of training, which requires more computational resources. For real-time speech enhancement, it is preferable to use a smaller network without losing the performance. In this paper, we propose to use the equilibriated recurrent neural network (ERNN) for avoiding the vanishing/exploding gradient problem without increasing the number of parameters. The proposed structure is causal, which requires only the information from the past, in order to apply it in real-time. Compared to the uni- and bi-directional LSTM networks, the proposed method achieved the similar performance with much fewer parameters.
Daiki Takeuchi, Kohei Yatabe, Yuma Koizumi, Yasuhiro Oikawa, Noboru Harada
ICASSP2
2020 Invertible DNN-Based Nonlinear Time-Frequency Transform for Speech Enhancement
abstract
We propose an end-to-end speech enhancement method with trainable time-frequency (T-F) transform based on invertible deep neural network (DNN). The resent development of speech enhancement is brought by using DNN. The ordinary DNN-based speech enhancement employs T-F transform, typically the short-time Fourier transform (STFT), and estimates a T-F mask using DNN. On the other hand, some methods have considered end-to-end networks which directly estimate the enhanced signals without T-F transform. While end-to-end methods have shown promising results, they are black boxes and hard to understand. Therefore, some end-to-end methods used a DNN to learn the linear T-F transform which is much easier to understand. However, the learned transform may not have a property important for ordinary signal processing. In this paper, as the important property of the T-F transform, perfect reconstruction is considered. An invertible nonlinear T-F transform is constructed by DNNs and learned from data so that the obtained transform is perfectly reconstructing filterbank.
Daiki Takeuchi, Kohei Yatabe, Yuma Koizumi, Yasuhiro Oikawa, Noboru Harada
ICASSP2
2020 Self-supervised Neural Audio-Visual Sound Source Localization via Probabilistic Spatial Modeling
abstract
Detecting sound source objects within visual observation is important for autonomous robots to comprehend surrounding environments. Since sounding objects have a large variety with different appearances in our living environments, labeling all sounding objects is impossible in practice. This calls for self-supervised learning which does not require manual labeling. Most of conventional self-supervised learning uses monaural audio signals and images and cannot distinguish sound source objects having similar appearances due to poor spatial information in audio signals. To solve this problem, this paper presents a self-supervised training method using 360° images and multichannel audio signals. By incorporating with the spatial information in multichannel audio signals, our method trains deep neural networks (DNNs) to distinguish multiple sound source objects. Our system for localizing sound source objects in the image is composed of audio and visual DNNs. The visual DNN is trained to localize sound source candidates within an input image. The audio DNN verifies whether each candidate actually produces sound or not. These DNNs are jointly trained in a self-supervised manner based on a probabilistic spatial audio model. Experimental results with simulated data showed that the DNNs trained by our method localized multiple speakers. We also demonstrate that the visual DNN detected objects including talking visitors and specific exhibits from real data recorded in a science museum.
Yoshiki Masuyama, Yoshiaki Bando, Kohei Yatabe, Yoko Sasaki, Masaki Onishi, Yasuhiro Oikawa
IROS3
2020 Joint Amplitude and Phase Refinement for Monaural Source Separation
abstract
Monaural source separation is often conducted by manipulating the amplitude spectrogram of a mixture (e.g., via time-frequency masking and spectral subtraction). The obtained amplitudes are converted back to the time domain by using the phase of the mixture or by applying phase reconstruction. Although phase reconstruction performs well for the true amplitudes, its performance is degraded when the amplitudes contain error. To deal with this problem, we propose an optimization-based method to refine both amplitudes and phases based on the given amplitudes. It aims to find time-domain signals whose amplitude spectrograms are close to the given ones in terms of the generalized alpha-beta divergences. To solve the optimization problem, the alternating direction method of multipliers (ADMM) is utilized. We confirmed the effectiveness of the proposed method through speech-nonspeech separation in various conditions.
Yoshiki Masuyama, Kohei Yatabe, Kento Nagatomo, Yasuhiro Oikawa
IEEE Signal Process. Lett.2
2020 Consistent ICA: Determined BSS Meets Spectrogram Consistency
abstract
Multichannel audio blind source separation (BSS) in the determined situation (the number of microphones is equal to that of the sources), or determined BSS, is performed by multichannel linear filtering in the time-frequency domain to handle the convolutive mixing process. Ordinarily, the filter treats each frequency independently, which causes the well-known permutation problem, i.e., the problem of how to align the frequency-wise filters so that each separated component is correctly assigned to the corresponding sources. In this paper, it is shown that the general property of the time-frequency-domain representation called spectrogram consistency can be an assistant for solving the permutation problem.
Kohei Yatabe
IEEE Signal Process. Lett.1
2019 Deep Griffin-Lim Iteration
abstract
This paper presents a novel phase reconstruction method (only from a given amplitude spectrogram) by combining a signal-processing-based approach and a deep neural network (DNN). To retrieve a time-domain signal from its amplitude spectrogram, the corresponding phase is required. One of the popular phase reconstruction methods is the Griffin-Lim algorithm (GLA), which is based on the redundancy of the short-time Fourier transform. However, GLA often involves many iterations and produces low-quality signals owing to the lack of prior knowledge of the target signal. In order to address these issues, in this study, we propose an architecture which stacks a sub-block including two GLA-inspired fixed layers and a DNN. The number of stacked sub-blocks is adjustable, and we can trade the performance and computational load based on requirements of applications. The effectiveness of the proposed method is investigated by reconstructing phases from amplitude spectrograms of speeches.
Yoshiki Masuyama, Kohei Yatabe, Yuma Koizumi, Yasuhiro Oikawa, Noboru Harada
ICASSP2
2019 Low-rankness of Complex-valued Spectrogram and Its Application to Phase-aware Audio Processing
abstract
Low-rankness of amplitude spectrograms has been effectively utilized in audio signal processing methods including non-negative matrix factorization. However, such methods have a fundamental limitation owing to their amplitude-only treatment where the phase of the observed signal is utilized for resynthesizing the estimated signal. In order to address this limitation, we directly treat a complex-valued spectrogram and show a complex-valued spectrogram of a sum of sinusoids can be approximately low-rank by modifying its phase. For evaluating the applicability of the proposed low-rank representation, we further propose a convex prior emphasizing harmonic signals, and it is applied to audio denoising.
Yoshiki Masuyama, Kohei Yatabe, Yasuhiro Oikawa
ICASSP2
2019 Phase-aware Harmonic/percussive Source Separation via Convex Optimization
abstract
Decomposition of an audio mixture into harmonic and percussive components, namely harmonic/percussive source separation (HPSS), is a useful pre-processing tool for many audio applications. Popular approaches to HPSS exploit the distinctive source-specific structures of power spectrograms. However, such approaches consider only power spectrograms, and the phase remains intact for resynthesizing the separated signals. In this paper, we propose a phase-aware HPSS method based on the structure of the phase of harmonic components. It is formulated as a convex optimization problem in the time domain, which enables the simultaneous treatment of both amplitude and phase. The numerical experiment validates the effectiveness of the proposed method.
Yoshiki Masuyama, Kohei Yatabe, Yasuhiro Oikawa
ICASSP2
2019 Data-driven Design of Perfect Reconstruction Filterbank for DNN-based Sound Source Enhancement
abstract
We propose a data-driven design method of perfect-reconstruction filterbank (PRFB) for sound-source enhancement (SSE) based on deep neural network (DNN). DNNs have been used to estimate a time-frequency (T-F) mask in the short-time Fourier transform (STFT) domain. Their training is more stable when a simple cost function as mean-squared error (MSE) is utilized comparing to some advanced cost such as objective sound quality assessments. However, such a simple cost function inherits strong assumptions on the statistics of the target and/or noise which is often not satisfied, and the mismatch of assumption results in degraded performance. In this paper, we propose to design the frequency scale of PRFB from training data so that the assumption on MSE is satisfied. For designing the frequency scale, the warped filterbank frame (WFBF) is considered as PRFB. The frequency characteristic of learned WFBF was in between STFT and the wavelet transform, and its effectiveness was confirmed by comparison with a standard STFT-based DNN whose input feature is compressed into the mel scale.
Daiki Takeuchi, Kohei Yatabe, Yuma Koizumi, Yasuhiro Oikawa, Noboru Harada
ICASSP2
2019 Guided-spatio-temporal Filtering for Extracting Sound from Optically Measured Images Containing Occluding Objects
abstract
Recent development of optical interferometry enables us to measure sound without placing any device inside the sound field. In particular, parallel phase-shifting interferometry (PPSI) has realized advanced measurement of refractive index of air. Its novel application investigated very recently is simultaneous visualization of flow and sound, which had been difficult until PPSI enabled high-speed and accurate measurement several years ago. However, for understanding aerodynamic sound, separation of air flow and sound is necessary since they are mixed up in the observed video. In this paper, guided-spatio-temporal filtering is proposed to separate sound from the optically measured images. Guided filtering is combined with a physical-model-based spatio-temporal filterbank for extracting sound-related information without the undesired effect caused by the image boundary or occluding objects. Such image boundary and occluding objects are typical difficulty arose in signal processing of an optically measured sound filed.
Risako Tanigawa, Kohei Yatabe, Yasuhiro Oikawa
ICASSP2
2019 Time-frequency-masking-based Determined BSS with Application to Sparse IVA
abstract
Most of the determined blind source separation (BSS) algorithms related to the independent component analysis (ICA) were derived from mathematical models of source signals. However, such derivation restricts the application of algorithms to explicitly definable source models, i.e., an implicit model associated with some signal-processing procedure cannot be utilized within such framework. In this paper, we propose an extension of the existing algorithm so that any time-frequency masking method (e.g., those developed in speech enhancement literature) can be incorporated into the determined BSS algorithm. As an application of the proposed algorithm, a sparse extension of the well-known independent vector analysis (IVA) is also proposed for illustrating the potentiality of the masking-based implicit source model.
Kohei Yatabe, Daichi Kitamura
ICASSP1
2019 Griffin-Lim Like Phase Recovery via Alternating Direction Method of Multipliers
abstract
Recovering a signal from its amplitude spectrogram, or phase recovery, exhibits many applications in acoustic signal processing. When only an amplitude spectrogram is available and no explicit information is given for the phases, the Griffin-Lim algorithm (GLA) is one of the most utilized methods for phase recovery. However, GLA often requires many iterations and results in low perceptual quality in some cases. In this letter, we propose two novel algorithms based on GLA and the alternating direction method of multipliers (ADMM) for better recovery with fewer iteration. Some interpretation of the existing methods and their relation to the proposed method are also provided. Evaluations are performed with both objective measure and subjective test.
Yoshiki Masuyama, Kohei Yatabe, Yasuhiro Oikawa
IEEE Signal Process. Lett.2
2018 Parametric Approximation of Piano Sound Based on Kautz Model with Sparse Linear Prediction
abstract
The piano is one of the most popular and attractive musical instruments that leads to a lot of research on it. To synthesize the piano sound in a computer, many modeling methods have been proposed from full physical models to approximated models. The focus of this paper is on the latter, approximating piano sound by an IIR filter. For stably estimating parameters, the Kautz model is chosen as the filter structure. Then, the selection of poles and excitation signal rises as the questions which are typical to the Kautz model that must be solved. In this paper, sparsity based construction of the Kautz model is proposed for approximating piano sound.
Kenji Kobayashi, Daiki Takeuchi, Mio Iwamoto, Kohei Yatabe, Yasuhiro Oikawa
ICASSP4
2018 Envelope Estimation by Tangentially Constrained Spline
abstract
Estimating envelope of a signal has various applications including empirical mode decomposition (EMD) in which the cubic C2-spline based envelope estimation is generally used. While such functional approach can easily control smoothness of an estimated envelope, the so-called undershoot problem often occurs that violates the basic requirement of envelope. In this paper, a tangentially constrained spline with tangential points optimization is proposed for avoiding the undershoot problem while maintaining smoothness. It is defined as a quartic C2-spline function constrained with first derivatives at tangential points that effectively avoids undershoot. The tangential points optimization method is proposed in combination with this spline to attain optimal smoothness of the estimated envelope.
Tsubasa Kusano, Kohei Yatabe, Yasuhiro Oikawa
ICASSP2
2018 Modal Decomposition of Musical Instrument Sound Via Alternating Direction Method of Multipliers
abstract
For a musical instrument sound containing partials, or modes, the behavior of modes around the attack time is particularly important. However, accurately decomposing it around the attack time is not an easy task, especially when the onset is sharp. This is because spectra of the modes are peaky while the sharp onsets need a broad one. In this paper, an optimization-based method of modal decomposition is proposed to achieve accurate decomposition around the attack time. The proposed method is formulated as a constrained optimization problem to enforce the perfect reconstruction property which is important for accurate decomposition. For optimization, the alternating direction method of multipliers (ADMM) is utilized, where the update of variables is calculated in closed form. The proposed method realizes accurate modal decomposition in the simulation and real piano sounds.
Yoshiki Masuyama, Tsubasa Kusano, Kohei Yatabe, Yasuhiro Oikawa
ICASSP3
2018 Realizing Directional Sound Source in FDTD Method by Estimating Initial Value
abstract
Wave-based acoustic simulation methods are studied actively for predicting acoustical phenomena. Finite-difference time-domain (FDTD) method is one of the most popular methods owing to its straightforwardness of calculating an impulse response. In an FDTD simulation, an omnidirectional sound source is usually adopted, which is not realistic because the real sound sources often have specific directivities. However, there is very little research on imposing a directional sound source into FDTD methods. In this paper, a method of realizing a directional sound source in FDTD methods is proposed. It is formulated as an estimation problem of the initial value so that the estimated result corresponds to the desired directivity. The effectiveness of the proposed method is illustrated through some numerical experiments.
Daiki Takeuchi, Kohei Yatabe, Yasuhiro Oikawa
ICASSP2
2018 Determined Blind Source Separation via Proximal Splitting Algorithm
abstract
The state-of-the-art algorithms of determined blind source separation (BSS) methods based on the independent component analysis (ICA) have gained computational efficiency by the majorization-minimization (MM) principle with a price of losing flexibility. That is, replacing and comparing different source models are not easy in such MM-based framework because it requires efforts to derive a new algorithm each time when one changes the model. In this paper, a general framework for obtaining an ICA-based BSS algorithm is proposed so that a source model can easily be replaced because only a single line of the algorithm must be modified. A sparsity-based extension of the independent vector analysis and a low-rankness-based BSS model using the nuclear norm are also proposed to demonstrate the simplicity and easiness of the proposed framework.
Kohei Yatabe, Daichi Kitamura
ICASSP1
2018 Phase Corrected Total Variation for Audio Signals
abstract
In optimization-based signal processing, the so-called prior term models the desired signal, and therefore its design is the key factor to achieve a good performance. For audio signals, the time-directional total variation applied to a spectrogram in combination with phase correction has been proposed recently to model sinusoidal components of the signal. Although it is a promising prior, its applicability might be restricted to some extent because of the mismatch of the assumption to the signal. In this paper, based upon the previously proposed one, an improved prior for audio signals named instantaneous phase corrected total variation (iPCTV) is proposed. It can handle wider range of audio signals owing to the instantaneous phase correction term calculated from the observed signal.
Kohei Yatabe, Yasuhiro Oikawa
ICASSP1
2017 Infinite-dimensional SVD for analyzing microphone array
abstract
Nowadays, various types of microphone array are used in many applications. However, it is not easy to compare arrays of different types because each array has been treated by a specific theory depending on the type of an array. Although several criteria have been proposed for microphone arrays for evaluating and/or designing an array, most of them are application-oriented criteria and the best configuration for some criterion may not be a better one in the other criterion. Therefore, an analysis and comparing method for microphone arrays which does not depend on an array configuration and application are necessary. In this paper, infinite-dimensional SVD is proposed for analyzing and comparing properties of arrays. The singular values and functions obtained by proposed method show sampling property of an array and can be unified criterion.
Yuji Koyano, Kohei Yatabe, Yasuhiro Oikawa
ICASSP2
2017 Coherence-adjusted monopole dictionary and convex clustering for 3D localization of mixed near-field and far-field sources
abstract
In this paper, 3D sound source localization method for simultaneously estimating both direction-of-arrival (DOA) and distance from the microphone array is proposed. For estimating distance, the off-grid problem must be overcome because the range of distance to be considered is quite broad and even not bounded. The proposed method estimates positions based on an extension of the convex clustering method combined with sparse coefficients estimation. A method for constructing a suitable monopole dictionary based on coherence is also proposed so that the convex clustering based method appropriately estimate distance of sound sources. Numerical experiments of distance estimation and 3D localization show possibility of the proposed method.
Tomoya Tachikawa, Kohei Yatabe, Yasuhiro Oikawa
ICASSP2
2016 Physical-model based efficient data representation for many-channel microphone array
abstract
Recent development of microphone arrays which consist of more than several tens or hundreds microphones enables acquisition of rich spatial information of sound. Although such information possibly improve performance of any array signal processing technique, the amount of data will increase as the number of microphones increases; for instance, a 1024 ch MEMS microphone array, as in Fig. 1, generates data more than 10 GB per minute. In this paper, a method constructing an orthogonal basis for efficient representation of sound data obtained by the microphone array is proposed. The proposed method can obtain a basis for arrays with any configuration including rectangle, spherical, and random microphone array. It can also be utilized for designing a microphone array because it offers a quantitative measure for comparing several array configurations.
Yuji Koyano, Kohei Yatabe, Yusuke Ikeda, Yasuhiro Oikawa
ICASSP2
2015 Visualization of sound field by means of Schlieren method with spatio-temporal filtering
abstract
Visualization of sound field using Schlieren technique provides many advantages. It enables us to investigate the change of the sound field in real-time from every point of the observing region. However, since the density gradient of air caused by the disturbance of acoustic field is very small, it is difficult to observe the audible sound field from the raw Schlieren video. In this paper, to enhance visibility of the audible sound fields from the Schlieren videos, we propose to use spatio-temporal filters for extracting sound information and for noise removal. We have utilized different filtering techniques such as the FIR bandpass filter, the Gaussian filter, the Wiener filter and the 3D Gabor filter, to do this. The results indicate that the data observed after using these signal processing methods are clearer than the raw Schlieren videos.
Nachanant Chitanont, Keita Yaginuma, Kohei Yatabe, Yasuhiro Oikawa
ICASSP3
2015 Optically visualized sound field reconstruction based on sparse selection of point sound sources
abstract
Visualization is an effective way to understand the behavior of a sound field. There are several methods for such observation including optical measurement technique which enables a non-destructive acoustical observation by detecting density variation of the medium. For audible sound propagating through the air, however, smallness of the variation requires high sensitivity of the measuring system that causes problematic noise contamination. In this paper, a method for reconstructing two-dimensional audible sound fields from noisy optical observation is proposed.
Kohei Yatabe, Yasuhiro Oikawa
ICASSP1
2014 PDE-based interpolation method for optically visualized sound field
abstract
An effective way to understand the behavior of a sound field is to visualize it. An optical measurement method is a suitable option for this as it enables contactless non-destructive measurement. After measuring a sound field, interpolation of the data is necessary for a smooth visualization. However, conventional interpolation methods cannot provide a physically meaningful result especially when the condition of the measurement causes moiré effect. In this paper, a special interpolation method for an optically visualized sound field based on the Kirchhoff-Helmholtz integral equation is proposed.
Kohei Yatabe, Yasuhiro Oikawa
ICASSP1