Yukoh Wakabayashi

dblp:158/4215 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0001-8846-3699ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 6 since 2021
YearPublicationVenuePosition
2025 Fine-tuning Parakeet-TDT for Dysarthric Speech Recognition in the Speech Accessibility Project Challenge
Kaito Takahashi, Keigo Hojo, Toshimitsu Sakai, Yukoh Wakabayashi, Norihide Kitaoka
INTERSPEECH4
2025 Domain adaptation using non-parallel target domain corpus for self-supervised learning-based automatic speech recognition
abstract
The recognition accuracy of conventional automatic speech recognition (ASR) systems depends heavily on the amount of speech and associated transcription data available in the target domain for model training. However, preparing parallel speech and text data each time a model is trained for a new domain is costly and time-consuming. To solve this problem, we propose a method of domain adaptation that does not require the use of a large amount of parallel target domain training data, as most of the data used for model training is not from the target domain. Instead, only target domain speech is used for model training, along with non-target domain speech and its parallel text data, i.e., the domains and contents of the two types of training data do not correspond to one another. Collecting this type of training data is relatively inexpensive. Domain adaptation is performed in two steps: (1) A pre-trained wav2vec 2.0 model is further pre-trained using a large amount of target domain speech data and is then fine-tuned using a large amount of non-target domain speech and its transcriptions. (2) The density ratio approach (DRA) is applied during inference to a language model (LM) trained using target domain text unrelated to, and independently from, the wav2vec 2.0 training. Experimental evaluation illustrated that the proposed domain adaptation obtained character error rate (CER) 10.4 pts lower than baseline with wav2vec 2.0 and 3.9 pts with XLS-R under the situation that the parallel target domain data is unavailable against the target domain test set, achieving 34.4% and 16.2% reductions in relative CER.
Takahiro Kinouchi, Atsunori Ogawa, Yukoh Wakabayashi, Kengo Ohta, Norihide Kitaoka
Speech Commun.3
2024 Boosting CTC-based ASR using inter-layer attention-based CTC loss
abstract
This paper addresses improving the performance of CTC-based models, which leverage the intermediate outputs of all encoder layers with an attention mechanism.Several previous studies have used the intermediate outputs of the encoder layer to modify CTC-based models.Here, we focus on the role of the Transformer encoder layer, and each encoder layer is computed for two CTC losses by weighting the intermediate outputs of its lower and upper layers using an attention mechanism.By dividing the layer into two groups, it is expected to be possible to calculate the loss, taking into account both acoustic and linguistic features.Experimental results showed that the proposed method improved the baseline recognition performance of TEDLIUM2 speech data, achieving a WER of 9.9% on the dev set and 11.8% on the test set.Our method outperformed the conventional methods for WER with only slightly increased inference speed measured by RTF.
Keigo Hojo, Yukoh Wakabayashi, Kengo Ohta, Atsunori Ogawa, Norihide Kitaoka
INTERSPEECH2
2024 Text-only Domain Adaptation for CTC-based Speech Recognition through Substitution of Implicit Linguistic Information in the Search Space
Tatsunari Takagi, Yukoh Wakabayashi, Atsunori Ogawa, Norihide Kitaoka
INTERSPEECH2
2024 Unequally Spaced Sound Field Interpolation for Rotation-Robust Beamforming
abstract
In this paper, we present an enhanced method designed to facilitate sound field interpolation (SFI) for rotation-robust beamforming using unequally spaced circular microphone arrays (unes-CMAs). Unlike the previous approach that necessitated an equally spaced circular microphone array (es-CMA), our method addresses the challenge of handling non-uniformly spaced microphones, making it suitable for real-world applications where unes-CMAs are more prevalent. Our proposed method enables the estimation of a virtual signal of an unes-CMA before rotation, derived from the observed signal after rotation. A modified SFI technique is utilized to compensate for the positional errors of microphones on an unes-CMA and to estimate a virtual signal at equally spaced positions after rotation. As an intermediate step, the previous SFI method is utilized to obtain equally spaced signals before rotation. Subsequently, the target signal of the unes-CMA before rotation is reconstructed, effectively achieving rotation-robust beamforming on the unes-CMA. Moreover, we provide an in-depth analysis of our proposed method's properties. We conducted simulated experiments, including online beamforming applications, to evaluate its performance. The experimental results demonstrated that our method effectively mitigates the adverse effects of unequal microphone placement, yielding significant improvements in estimating the signal before rotation under various conditions. Moreover, our proposed method consistently outperformed the previous approach, significantly enhancing the performance of beamforming.
Shuming Luan, Yukoh Wakabayashi, Tomoki Toda
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Weighted Von Mises Distribution-based Loss Function for Real-time STFT Phase Reconstruction Using DNN
abstract
This paper presents improvements to real-time phase reconstruction using deep neural networks (DNNs).The advantage of DNN-based approaches in phase reconstruction is that they can leverage prior knowledge from data and are adaptable to realtime applications by using causal models.However, conventional DNN-based methods do not consider the varying properties of the phase at different time-frequency bins.Our paper proposes loss functions for phase reconstruction that incorporate frequency-specific and amplitude weights to distinguish the importance of phase elements based on their properties.We also use an extension of the group delay to improve the phase connections along the frequency.To improve the generalization, we augment the data by randomly shifting the signals in the time domain for each epoch during training.Experimental results show the superior performance of the proposed methods compared to conventional DNN-based and non-DNN real-time phase reconstruction methods.
Nguyen Binh Thien, Yukoh Wakabayashi, Yuting Geng, Kenta Iwai, Takanobu Nishiura
INTERSPEECH2
2023 Inter-Frequency Phase Difference for Phase Reconstruction Using Deep Neural Networks and Maximum Likelihood
abstract
This paper presents improvements to two-stage algorithms for estimating the short-time Fourier transform (STFT) phase from only the amplitude by using deep neural networks (DNNs). The phase is difficult to reconstruct due to its sensitivity to the waveform shift and wrapping issue. To mitigate these problems, two-stage approaches indirectly estimate the phase through phase derivatives, i.e., instantaneous frequency (IF) and group delay (GD). In the first stage, the IF and GD are estimated from the amplitude using DNNs, and then in the second stage, the phase is reconstructed by maintaining the IF/GD information. Conventional methods for the second stage do not consider the importance of high-amplitude time–frequency bins, e.g., the least squares-based method, or lack a solid model, e.g., the average-based method. To address these problems, we propose improvements to the second stage of two-stage algorithms by usingvon Misesdistribution-based maximum likelihood and weighted least squares. We also provide theoretical discussions for the phase reconstruction, including the investigations of the properties of the GD and roles of the IF/GD information in the inverse STFT. On the basis of the analysis, we propose a new phase-based feature, i.e., inter-frequency phase difference (IFPD), and demonstrate its application in two-stage phase reconstruction algorithms. We conducted subjective and objective experiments to compare the performances of our proposed and conventional methods. The results confirm that the proposed method using the IFPD performs better than other methods for all metrics.
Nguyen Binh Thien, Yukoh Wakabayashi, Kenta Iwai, Takanobu Nishiura
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Sound Field Interpolation for Rotation-Invariant Multichannel Array Signal Processing
abstract
In this paper, we present a sound field interpolation for array signal processing (ASP) that is robust to rotation of a circular microphone array (CMA), and we evaluate beamforming as one of its applications. Most ASP methods assume a time-invariant acoustic transfer system (ATS) from sources to the microphone array. This assumption makes it challenging to perform ASP in real situations where sources and the microphone array can move. Therefore, considering a time-variant ATS is an essential task for the use of ASP. In this study, we focus on one such movement, the rotation of the CMA. Our method interpolates the sound field on the circumference of a circle, where microphones are equally spaced, based on the sampling theorem on the circle. The interpolation enables us to estimate the signals at the microphone positions before the rotation. Hence, conventional ASP, which assumes a time-invariant ATS, is applicable after interpolation without modification. We developed two beamforming schemes, one for batch and one for online processing, that combine the minimum power distortionless response beamformer and sound field interpolation. We evaluated the dependences of the interpolation on frequency and rotation angle using the signal-to-error ratio. Additionally, simulation results demonstrated that the two proposed schemes improve the beamformer's performance when the CMA rotates.
Yukoh Wakabayashi, Kouei Yamaoka, Nobutaka Ono
IEEE ACM Trans. Audio Speech Lang. Process.1
2021 Rotation-Robust Beamforming Based on Sound Field Interpolation with Regularly Circular Microphone Array
abstract
In this paper, we present a novel framework of beamforming robust for a microphone array rotation. In most array signal processing methods, the time-invariant transfer system from a source to a microphone is assumed for calculating a spatial filter. This assumption makes it difficult to use the microphone array in real situations since sources and the microphone array may move. In this work, we focus on one such movement, the array’s rotation. The key in our method is to use a regularly circular microphone array where microphones are equally spaced on a circle’s circumference. Based on this, we propose a method to interpolate the sound field on the circumference. The interpolation enables us to estimate the signals at the microphone positions before the rotation even when the array rotates. Hence, the conventional array signal processing assuming the time-invariant system is applicable. Simulation results indicated that the proposed method improves the beamformer’s performance under the array rotation.
Yukoh Wakabayashi, Kouei Yamaoka, Nobutaka Ono
ICASSP1
2019 Blink-former: Light-aided beamforming for multiple targets enhancement
abstract
We propose a multimodal framework to enhance multiple target sound sources using a conventional microphone array, a video camera, and sound power sensors, called Blinkies, that we have recently developed. Each Blinky consists of a microphone, LEDs, a microcontroller, and a battery. One of the LEDs intensity is varied according to sound power, that is, the Blinky works as a sound-to-light conversion sensor. They are easy to distribute over a large area, and thus, the sound power information therein can be harvested by capturing the LED signals with a video camera. Although these signals are a mixture of contributions from multiple sources, we demonstrate that they can be separated into individual source activities by non-negative matrix factorization. The obtained activities are further utilized to design maximum signal-to-interference-and-noise ratio beamformers enhancing the source signals. We conduct numerical simulations and real experiments to evaluate the performance of this method in diffuse noise environment. The experimental results show that the proposed scheme using Blinkies is superior to competing algorithms, especially at low signal-to-noise ratio.
Daiki Horiike, Robin Scheibler, Yukoh Wakabayashi, Nobutaka Ono
MMSP3
2018 Single-Channel Speech Enhancement With Phase Reconstruction Based on Phase Distortion Averaging
abstract
Speech enhancement has been widely investigated for several decades, but by modifying only the amplitude spectrum of a speech signal, ignoring the phase spectrum, which has been regarded as an unimportant feature. However, it was recently reported that the phase spectrum plays an important role in speech quality and intelligibility. In this paper, we propose a phase reconstruction method based on harmonic enhancement using the fundamental frequency and phase distortion feature. This feature is known to show fluctuations in the phase spectrum with respect to time and frequency. We estimate the speech phase spectrum by considering the relationship between harmonic phase spectra. Experimental evaluations indicate that the proposed phase reconstruction method improves speech quality in various noisy environments.
Yukoh Wakabayashi, Takahiro Fukumori, Masato Nakayama, Takanobu Nishiura, Yoichi Yamashita
IEEE ACM Trans. Audio Speech Lang. Process.1
2017 Phase reconstruction method based on time-frequency domain harmonic structure for speech enhancement
abstract
Speech enhancement in noisy environments has been widely investigated by modifying only the amplitude spectrum of the speech signal, while the phase spectrum, which is regarded as an unimportant feature, is ignored. However, it has recently been reported that the phase spectrum plays an important role in the intelligibility and quality of speech. We propose a speech-enhancement method with phase reconstruction, which estimates inartificial phase spectrum by using the time-frequency feature called phase distortion, though a conventional phase reconstruction estimates artificial one. The objective experimental results indicate improvement in speech quality with and efficiency of the proposed method.
Yukoh Wakabayashi, Takahiro Fukumori, Masato Nakayama, Takanobu Nishiura, Yoichi Yamashita
ICASSP1
2015 Enhanced speaker diarization with detection of backchannels using eye-gaze information in poster conversations
abstract
We propose multi-modal speaker diarization using acoustic and eye-gaze information in poster conversations. Eye-gaze information plays an important role in turn-taking, thus it is useful for predicting speech activity. In this paper, a variety of eyegaze features are elaborated and combined with the acoustic information by the multi-modal integration model. Moreover, we introduce another model to detect backchannels, which involve different eye-gaze behaviors. This enhances the diarization result by filtering meaningful utterances such as questions and comments. Experimental evaluations in real poster sessions demonstrate that eye-gaze information contributes to improvement of diarization accuracy under noisy environments, and its weight is automatically determined according to the Signal-toNoise Ratio (SNR). Index Terms: speaker diarization, backchannel, multi-modal, eye-gaze, poster conversation
Koji Inoue, Yukoh Wakabayashi, Hiromasa Yoshimoto, Katsuya Takanashi, Tatsuya Kawahara
INTERSPEECH2
2014 Speaker diarization using eye-gaze information in multi-party conversations
abstract
We present a novel speaker diarization method by using eye-gaze information in multi-party conversations. In real environ-ments, speaker diarization or speech activity detection of each participant of the conversation is challenging because of distant talking and ambient noise. In contrast, eye-gaze information is robust against acoustic degradation, and it is presumed that eye-gaze behavior plays an important role in turn-taking and thus in predicting utterances. The proposed method stochastically integrates eye-gaze information with acoustic information for speaker diarization. Specifically, three models are investigated for multi-modal integration in this paper. Experimental eval-uations in real poster sessions demonstrate that the proposed method improves accuracy of speaker diarization from the base-line acoustic method. Index Terms: speaker diarization, multi-modal interaction, eye-gaze
Koji Inoue, Yukoh Wakabayashi, Hiromasa Yoshimoto, Tatsuya Kawahara
INTERSPEECH2