Yusuke Kida

dblp:66/5664 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
6since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 3 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2023 Neural Diarization with Non-Autoregressive Intermediate Attractors
abstract
End-to-end neural diarization (EEND) with encoder-decoder-based attractors (EDA) is a promising method to handle the whole speaker diarization problem simultaneously with a single neural network. While the EEND model can produce all frame-level speaker labels simultaneously, it disregards output label dependency. In this work, we propose a novel EEND model that introduces the label dependency between frames. The proposed method generates non-autoregressive intermediate attractors to produce speaker labels at the lower layers and conditions the subsequent layers with these labels. While the proposed model works in a non-autoregressive manner, the speaker labels are refined by referring to the whole sequence of intermediate labels. The experiments with the two-speaker CALLHOME dataset show that the intermediate labels with the proposed non-autoregressive intermediate attractors boost the diarization performance. The proposed method with the deeper net-work benefits more from the intermediate labels, resulting in better performance and training throughput than EEND-EDA.
Yusuke Fujita, Tatsuya Komatsu, Robin Scheibler, Yusuke Kida, Tetsuji Ogawa
ICASSP4
2023 Conversation-Oriented ASR with Multi-Look-Ahead CBS Architecture
abstract
During conversations, humans are capable of inferring the intention of the speaker at any point of the speech to prepare the following action promptly. Such ability is also the key for conversational systems to achieve rhythmic and natural conversation. To perform this, the automatic speech recognition (ASR) used for transcribing the speech in real-time must achieve high accuracy without delay. In streaming ASR, high accuracy is assured by attending to look-ahead frames, which leads to delay increments. To tackle this trade-off issue, we propose a multiple latency streaming ASR to achieve high accuracy with zero look-ahead. The proposed system contains two encoders that operate in parallel, where a primary encoder generates accurate outputs utilizing look-ahead frames, and the auxiliary encoder recognizes the look-ahead portion of the primary encoder without look-ahead. The proposed system is constructed based on contextual block streaming (CBS) architecture, which leverages block processing and has a high affinity for the multiple latency architecture. Various methods are also studied for architecting the system, including shifting the network to perform as different encoders; as well as generating both encoders’ outputs in one encoding pass.
Huaibo Zhao, Shinya Fujie, Tetsuji Ogawa, Jin Sakuma, Yusuke Kida, Tetsunori Kobayashi
ICASSP5
2023 Target Vocabulary Recognition Based on Multi-Task Learning with Decomposed Teacher Sequences
Aoi Ito, Tatsuya Komatsu, Yusuke Fujita, Yusuke Kida
INTERSPEECH4
2022 Better Intermediates Improve CTC Inference
abstract
This paper proposes a method for improved CTC inference with searched intermediates and multi-pass conditioning.The paper first formulates self-conditioned CTC as a probabilistic model with an intermediate prediction as a latent representation and provides a tractable conditioning framework.We then propose two new conditioning methods based on the new formulation:(1) Searched intermediate conditioning that refines intermediate predictions with beam-search, (2) Multi-pass conditioning that uses predictions of previous inference for conditioning the next inference.These new approaches enable better conditioning than the original self-conditioned CTC during inference and improve the final performance.Experiments with the LibriSpeech dataset show relative 3%/12% performance improvement at the maximum in test clean/other sets compared to the original selfconditioned CTC.
Tatsuya Komatsu, Yusuke Fujita, Jaesong Lee, Lukas Lee, Shinji Watanabe 0001, Yusuke Kida
INTERSPEECH6
2022 InterAug: Augmenting Noisy Intermediate Predictions for CTC-based ASR
abstract
This paper proposes InterAug: a novel training method for CTC-based ASR using augmented intermediate representations for conditioning.The proposed method exploits the conditioning framework of self-conditioned CTC to train robust models by conditioning with "noisy" intermediate predictions.During the training, intermediate predictions are changed to incorrect intermediate predictions, and fed into the next layer for conditioning.The subsequent layers are trained to correct the incorrect intermediate predictions with the intermediate losses.By repeating the augmentation and the correction, iterative refinements, which generally require a special decoder, can be realized only with the audio encoder.To produce noisy intermediate predictions, we also introduce new augmentation: intermediate feature space augmentation and intermediate token space augmentation that are designed to simulate typical errors.The combination of the proposed InterAug framework with new augmentation allows explicit training of the robust audio encoders.In experiments using augmentations simulating deletion, insertion, and substitution error, we confirmed that the trained model acquires robustness to each error, boosting the speech recognition performance of the strong self-conditioned CTC baseline.
Yu Nakagome, Tatsuya Komatsu, Yusuke Fujita, Shuta Ichimura, Yusuke Kida
INTERSPEECH5
2022 Alternate Intermediate Conditioning with Syllable-Level and Character-Level Targets for Japanese ASR
Yusuke Fujita, Tatsuya Komatsu, Yusuke Kida
SLT3
2019 Simultaneous Detection and Localization of a Wake-Up Word Using Multi-Task Learning of the Duration and Endpoint
Takashi Maekaku, Yusuke Kida, Akihiko Sugiyama
INTERSPEECH2
2018 Speaker Selective Beamformer with Keyword Mask Estimation
abstract
This paper addresses the problem of automatic speech recognition (ASR) of a target speaker in background speech. The novelty of our approach is that we focus on a wakeup keyword, which is usually used for activating ASR systems like smart speakers. The proposed method firstly utilizes a DNN-based mask estimator to separate the mixture signal into the keyword signal uttered by the target speaker and the remaining background speech. Then the separated signals are used for calculating a beamforming filter to enhance the subsequent utterances from the target speaker. Experimental evaluations show that the trained DNN-based mask can selectively separate the keyword and background speech from the mixture signal. The effectiveness of the proposed method is also verified with Japanese ASR experiments, and we confirm that the character error rates are significantly improved by the proposed method for both simulated and real recorded test sets.
Yusuke Kida, Dung T. Tran, Motoi Omachi, Toru Taniguchi, Yuya Fujita
SLT1
2016 Voice Activity Detection: Merging Source and Filter-based Information
abstract
Voice Activity Detection (VAD) refers to the problem of distinguishing speech segments from background noise. Numerous approaches have been proposed for this purpose. Some are based on features derived from the power spectral density, others exploit the periodicity of the signal. The goal of this letter is to investigate the joint use of source and filter-based features. Interestingly, a mutual information-based assessment shows superior discrimination power for the source-related features, especially the proposed ones. The features are further the input of an artificial neural network-based classifier trained on a multi-condition database. Two strategies are proposed to merge source and filter information: feature and decision fusion. Our experiments indicate an absolute reduction of 3% of the equal error rate when using decision fusion. The final proposed system is compared to four state-of-the-art methods on 150 minutes of data recorded in real environments. Thanks to the robustness of its source-related features, its multi-condition training and its efficient information fusion, the proposed system yields over the best state-of-the-art VAD a substantial increase of accuracy across all conditions (24% absolute on average).
Thomas Drugman, Yannis Stylianou, Yusuke Kida, Masami Akamine
IEEE Signal Process. Lett.3
2010 Using duration and pitch for mandarin digit string recognition
abstract
Mandarin digit string recognition (MDSR) is a challenge because there exist many difficulties in acoustic discrimination for such a small vocabulary speech recognition task. In this paper, we propose to improve MDSR performance by using duration and pitch information. Speech rate dispersion is used to involve duration knowledge and is incorporated in the MDSR system by rescoring the N-best candidates in a two-pass framework. estimated with a robust pitch extraction method is also adopted to improve the acoustic discrimination among Mandarin digits. The experimental results show both duration and pitch significantly improve the performance, and the combination of them gives further improvement. Moreover, our methods are robust to background noise. In the evaluation, the sentence error rate is reduced by 50.43% on average over different SNR conditions.
Rui Zhao 0025, Yusuke Kida, Pei Ding, Lei He 0020
ICASSP2
2009 Robust F0 estimation based on log-time scale autocorrelation and its application to Mandarin tone recognition
Yusuke Kida, Masaru Sakai, Takashi Masuko, Akinori Kawamura
INTERSPEECH1
2006 Evaluation of voice activity detection by combining multiple features with weight adaptation
abstract
For noise-robust automatic speech recognition (ASR), we propose a novel voice activity detection (VAD) method based on a combination of multiple features. The scheme uses a weighted combination of four conventionalVAD features: amplitude level, zero crossing rate, spectral information, and Gaussian mixture model (GMM) likelihood. The weights for combination are adaptively updated using minimum classification error (MCE) training. In this paper, we first investigate the effect of adaptation of the combination weights and GMM parameters, and demonstrate that the weights can be effectively adapted with a single utterance. Then, we present application of the method to ASR. It is confirmed that the proposed method significantly outperforms conventional methods in various noise conditions. Index Terms: speech recognition, voice activity detection, MCE training, noise adaptation
Yusuke Kida, Tatsuya Kawahara
INTERSPEECH1
2005 Minimum Classification Error Interactive Training for Speaker Identification
abstract
This paper describes an online discriminative training algorithm aiming at achieving speaker identification on interactive robots. A robot incrementally acquires speakers' voice characteristics during the interaction with the speakers. We simulate the situation that the speakers never give their IDs and the robot can only know whether the identification decision was correct or not from the speaker's positive or negative behavioral reaction. The speaker models are adjusted based on this limited information using minimum classification error (MCE) training consisting of positive and negative adaptation. In cases of correct identification, the conventional MCE training algorithm can be used. We compare three kinds of negative adaptation algorithms for the cases of incorrect identification. Experimental results show that the combination of the positive and negative adaptation achieves faster convergence, and negative adaptation which adjusts only a misclassified speaker model reaches an identification rate of 80% four times faster than the positive adaptation alone.
Yusuke Kida, Hiroyoshi Yamamoto, Chiyomi Miyajima, Keiichi Tokuda, Tadashi Kitamura
ICASSP (1)1
2005 Voice activity detection based on optimally weighted combination of multiple features
Yusuke Kida, Tatsuya Kawahara
INTERSPEECH1
2005 Online dense local 3D world reconstruction from stereo image sequences
abstract
This paper describes an online 3D reconstruction system from stereo image sequences to obtain a dense local world model for robot navigation. The proposed method consists of three components: 1) stereo depth map calculation, 2) correspondence calculation in time sequential images by tracking raw image features, 3) 6DOF camera motion estimation by RANSAC and integrate depth map into 3D reconstructed model. We examined and evaluated our method in a motion capture environment for comparison. Finally experimental results of a humanoid robot H7 are denoted.
Satoshi Kagami, Yutaka Takaoka, Yusuke Kida, Koichi Nishiwaki, Takeo Kanade
IROS3
2005 Using visual odometry to create 3D maps for online footstep planning
abstract
This paper describes an online system for footstep planning using a 3D map reconstructed by visual odometry. This system consists of two key components: 3D reconstruction via visual odometry from a stereo image sequence to obtain a dense local world model, and a footstep planner for biped robots using the reconstructed 3D map. Visual odometry is a method to connect 3D image sequences to obtain 6DOF camera motion and dense 3D environment information. The method described in this paper consists of three components: stereo depth map calculation, 3D flow calculation from tracking raw image features, and 6DOF camera motion estimation from RANSAC. Using the resulting 3D data, an optimal sequence of footstep locations is planned. The footstep planner is provided a height map of the terrain and a discrete set of possible footstep motions. The planner then evaluates footstep locations for viability using a collection of heuristic metrics designed to encode the relative safety, effort required, and overall motion complexity. Finally, we implemented this system on the humanoid robot H7. A local 3D map is reconstructed using visual odometry at about 10 Hz and the footstep planner replans at intervals of four steps. The robot walked across a floor, avoiding obstacles and reaching the goal.
Risa Ozawa, Yutaka Takaoka, Yusuke Kida, Koichi Nishiwaki, Joel E. Chestnutt, James J. Kuffner, J. Kagami, H. Mizoguch, Hirochika Inoue
SMC3