VLDB 2026 Research / reviewers in the wild / expert
Yuki Takashima
dblp:173/6862
· DBLP profile ↗
13ranked-venue papers
7as first author
8since 2021 · last 2023
0000-0001-8489-9487ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Online Neural Diarization of Unlimited Numbers of Speakers Using Global and Local AttractorsabstractA method to perform offline and online speaker diarization for an unlimited number of speakers is described in this paper. End-to-end neural diarization (EEND) has achieved overlap-aware speaker diarization by formulating it as a multi-label classification problem. It has also been extended for a flexible number of speakers by introducing speaker-wise attractors. However, the output number of speakers of attractor-based EEND is empirically capped; it cannot deal with cases where the number of speakers appearing during inference is higher than that during training because its speaker counting is trained in a fully supervised manner. Our method, EEND-GLA, solves this problem by introducing unsupervised clustering into attractor-based EEND. In the method, the input audio is first divided into short blocks, then attractor-based diarization is performed for each block, and finally, the results of each block are clustered on the basis of the similarity between locally-calculated attractors. While the number of output speakers is limited within each block, the total number of speakers estimated for the entire input can be higher than the limitation. To use EEND-GLA in an online manner, our method also extends the speaker-tracing buffer, which was originally proposed to enable online inference of conventional EEND. We introduce a block-wise buffer update to make the speaker-tracing buffer compatible with EEND-GLA. Finally, to improve online diarization, our method improves the buffer update method and revisits the variable chunk-size training of EEND. The experimental results demonstrate that EEND-GLA can perform speaker diarization of an unseen number of speakers in both offline and online inferences. Shota Horiguchi, Shinji Watanabe 0001, L. Paola García-Perera, Yuki Takashima, Yohei Kawaguchi |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | Multi-Channel End-To-End Neural Diarization with Distributed MicrophonesabstractRecent progress on end-to-end neural diarization (EEND) has en-abled overlap-aware speaker diarization with a single neural net-work. This paper proposes to enhance EEND by using multi-channel signals from distributed microphones. We replace Transformer en-coders in EEND with two types of encoders that process a multi-channel input: spatio-temporal and co-attention encoders. Both are independent of the number and geometry of microphones and suitable for distributed microphone settings. We also propose a model adaptation method using only single-channel recordings. With simulated and real-recorded datasets, we demonstrated that the proposed method outperformed conventional EEND when a multi-channel in-put was given while maintaining comparable performance with a single-channel input. We also showed that the proposed method performed well even when spatial information is inoperative given multi-channel inputs, such as in hybrid meetings in which the utterances of multiple remote participants are played back from the same loudspeaker. Shota Horiguchi, Yuki Takashima, L. Paola García-Perera, Shinji Watanabe 0001, Yohei Kawaguchi |
ICASSP | 2 |
| 2022 | Updating Only Encoders Prevents Catastrophic Forgetting of End-to-End ASR ModelsabstractIn this paper, we present an incremental domain adaptation technique to prevent catastrophic forgetting for an end-to-end automatic speech recognition (ASR) model.Conventional approaches require extra parameters of the same size as the model for optimization, and it is difficult to apply these approaches to end-to-end ASR models because they have a huge amount of parameters.To solve this problem, we first investigate which parts of end-to-end ASR models contribute to high accuracy in the target domain while preventing catastrophic forgetting.We conduct experiments on incremental domain adaptation from the LibriSpeech dataset to the AMI meeting corpus with two popular end-to-end ASR models and found that adapting only the linear layers of their encoders can prevent catastrophic forgetting.Then, on the basis of this finding, we develop an element-wise parameter selection focused on specific layers to further reduce the number of fine-tuning parameters.Experimental results show that our approach consistently prevents catastrophic forgetting compared to parameter selection from the whole model. Yuki Takashima, Shota Horiguchi, Shinji Watanabe 0001, L. Paola García-Perera, Yohei Kawaguchi |
INTERSPEECH | 1 |
| 2022 | Mutual Learning of Single- and Multi-Channel End-to-End Neural DiarizationabstractDue to the high performance of multi-channel speech processing, we can use the outputs from a multi-channel model as teacher labels when training a single-channel model with knowledge distillation. To the contrary, it is also known that single-channel speech data can benefit multi-channel models by mixing it with multi-channel speech data during training or by using it for model pretraining. This paper focuses on speaker diarization and proposes to conduct the above bi-directional knowledge transfer alternately. We first introduce an end-to-end neural diarization model that can handle both single- and multi-channel inputs. Using this model, we alternately conduct i) knowledge distillation from a multi-channel model to a single-channel model and ii) finetuning from the distilled single-channel model to a multi-channel model. Experimental results on two-speaker data show that the proposed method mutually improved single- and multi-channel speaker diarization performances. Shota Horiguchi, Yuki Takashima, Shinji Watanabe 0001, L. Paola García-Perera |
SLT | 2 |
| 2021 | Towards Neural Diarization for Unlimited Numbers of Speakers Using Global and Local AttractorsabstractAttractor-based end-to-end diarization is achieving comparable accuracy to the carefully tuned conventional clustering-based methods on challenging datasets. However, the main drawback is that it cannot deal with the case where the number of speakers is larger than the one observed during training. This is because its speaker counting relies on supervised learning. In this work, we introduce an unsupervised clustering process embedded in the attractor-based end-to-end diarization. We first split a sequence of frame-wise embeddings into short subsequences and then perform attractor-based diarization for each subsequence. Given subsequence-wise diarization results, inter-subsequence speaker correspondence is obtained by unsupervised clustering of the vectors computed from the attractors from all the subsequences. This makes it possible to produce diarization results of a large number of speakers for the whole recording even if the number of output speakers for each subsequence is limited. Experimental results showed that our method could produce accurate diarization results of an unseen number of speakers. Our method achieved 11.84 %, 28.33 %, and 19.49 % on the CALLHOME, DI-HARD II, and DIHARD III datasets, respectively, each of which is better than the conventional end-to-end diarization methods. Shota Horiguchi, Shinji Watanabe 0001, L. Paola García-Perera, Yawen Xue, Yuki Takashima, Yohei Kawaguchi |
ASRU | 5 |
| 2021 | Semi-Supervised Training with Pseudo-Labeling for End-To-End Neural DiarizationabstractIn this paper, we present a semi-supervised training technique using pseudo-labeling for end-to-end neural diarization (EEND).The EEND system has shown promising performance compared with traditional clustering-based methods, especially in the case of overlapping speech.However, to get a welltuned model, EEND requires labeled data for all the joint speech activities of every speaker at each time frame in a recording.In this paper, we explore a pseudo-labeling approach that employs unlabeled data.First, we propose an iterative pseudolabel method for EEND, which trains the model using unlabeled data of a target condition.Then, we also propose a committeebased training method to improve the performance of EEND.To evaluate our proposed method, we conduct the experiments of model adaptation using labeled and unlabeled data.Experimental results on the CALLHOME dataset show that our proposed pseudo-label achieved a 37.4% relative diarization error rate reduction compared to a seed model.Moreover, we analyzed the results of semi-supervised adaptation with pseudo-labeling.We also show the effectiveness of our approach on the third DI-HARD dataset. Yuki Takashima, Yusuke Fujita, Shota Horiguchi, Shinji Watanabe 0001, L. Paola García-Perera, Kenji Nagamatsu |
Interspeech | 1 |
| 2021 | Online Streaming End-to-End Neural Diarization Handling Overlapping Speech and Flexible Numbers of Speakers
Yawen Xue, Shota Horiguchi, Yusuke Fujita, Yuki Takashima, Shinji Watanabe 0001, L. Paola García-Perera, Kenji Nagamatsu |
Interspeech | 4 |
| 2021 | End-to-End Speaker Diarization Conditioned on Speech Activity and Overlap DetectionabstractIn this paper, we present a conditional multitask learning method for end-to-end neural speaker diarization (EEND). The EEND system has shown promising performance compared with traditional clustering-based methods, especially in the case of overlapping speech. In this paper, to further improve the performance of the EEND system, we propose a novel multitask learning framework that solves speaker diarization and a desired subtask while explicitly considering the task dependency. We optimize speaker diarization conditioned on speech activity and overlap detection that are subtasks of speaker diarization, based on the probabilistic chain rule. Experimental results show that our proposed method can leverage a subtask to effectively model speaker diarization, and outperforms conventional EEND systems in terms of diarization error rate. Yuki Takashima, Yusuke Fujita, Shinji Watanabe 0001, Shota Horiguchi, L. Paola García-Perera, Kenji Nagamatsu |
SLT | 1 |
| 2020 | Dysarthric Speech Recognition Based on Deep Metric Learning
Yuki Takashima, Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki |
INTERSPEECH | 1 |
| 2019 | End-to-end Dysarthric Speech Recognition Using Multiple DatabasesabstractWe present in this paper an end-to-end automatic speech recognition (ASR) system for a person with an articulation disorder resulting from athetoid cerebral palsy. In the case of a person with this type of articulation disorder, the speech style is quite different from that of a physically unimpaired person, and the amount of their speech data available to train the model is limited because their burden is large due to strain on the speech muscles. Therefore, the performance of ASR systems for people with an articulation disorder degrades significantly. In this paper, we propose an end-to-end ASR framework trained by not only the speech data of a Japanese person with an articulation disorder but also the speech data of a physically unimpaired Japanese person and a non-Japanese person with an articulation disorder to relieve the lack of training data of a target speaker. An end-to-end ASR model encapsulates an acoustic and language model jointly. In our proposed model, an acoustic model portion is shared between persons with dysarthria, and a language model portion is assigned to each language regardless of dysarthria. Experimental results show the merit of our proposed approach of using multiple databases for speech recognition. Yuki Takashima, Tetsuya Takiguchi, Yasuo Ariki |
ICASSP | 1 |
| 2018 | Parallel-Data-Free Dictionary Learning for Voice Conversion Using Non-Negative Tucker DecompositionabstractVoice conversion (VC) is a technique where only speaker-specific information in source speech is converted while preserving the associated phonological information. Nonnegative Matrix Factorization (NMF)-based VC has been researched because of the natural-sounding voice it produces compared with conventional Gaussian Mixture Model-based VC. In conventional NMF- VC, parallel data are used to train the models; therefore, unnatural pre-processing of speech data to make parallel data is needed. NMF-VC also tends to be a large model because this method has many parallel exemplars for the dictionary matrix; therefore, the computational cost is high. In this paper, we propose a novel parallel dictionary learning method using non-negative Tucker decomposition (NTD) which uses tensor decomposition and decomposes an input observation into a set of mode matrices and one core tensor. Our proposed NTD-based dictionary learning method estimates the dictionary matrix for NMF- VC without using parallel data. Experimental results show that our proposed method outperforms conventional non-parallel VC methods. Yuki Takashima, Hajime Yano, Toru Nakashika, Tetsuya Takiguchi, Yasuo Ariki |
ICASSP | 1 |
| 2016 | Lip reading using a dynamic feature of lip images and convolutional neural networksabstractIn this paper, a lip-reading method using a novel dynamic feature of lip images is proposed. The dynamic feature of lip images is calculated as the first-order regression coefficients using a few neighboring frames (images). It constiutes a better representation of the time derivatives to the basic static image. The dynamic feature is processed by using convolution neural networks (CNNs), which are able to reduce the negative influence caused by shaking of the subject and face alignment blurring at the feature-extraction level. Its effectiveness has been confirmed by word-recognition experiments comparing the proposed method with the conventional static (original) image. Yuki Takashima, Tetsuya Takiguchi, Yasuo Ariki |
ICIS | 2 |
| 2016 | Audio-Visual Speech Recognition Using Bimodal-Trained Bottleneck Features for a Person with Severe Hearing Loss
Yuki Takashima, Ryo Aihara, Tetsuya Takiguchi, Yasuo Ariki, Nobuyuki Mitani, Kiyohiro Omori, Kaoru Nakazono |
INTERSPEECH | 1 |