EDBT 2026 Demo / reviewers in the wild / expert
Tianyan Zhou
dblp:198/9435
· DBLP profile ↗
18ranked-venue papers
2as first author
7since 2021 · last 2021
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | A Comparative Study of Modular and Joint Approaches for Speaker-Attributed ASR on Monaural Long-Form AudioabstractSpeaker-attributed automatic speech recognition (SA-ASR) is a task to recognize “who spoke what” from multi-talker recordings. An SA-ASR system usually consists of multiple modules such as speech separation, speaker diarization and ASR. On the other hand, considering the joint optimization, an end-to-end (E2E) SA-ASR model has recently been proposed with promising results on simulation data. In this paper, we present our recent study on the comparison of such modular and joint approaches towards SA-ASR on real monaural recordings. We develop state-of-the-art SA-ASR systems for both modular and joint approaches by leveraging large-scale training data, including 75 thousand hours of ASR training data and the VoxCeleb corpus for speaker representation learning. We also propose a new pipeline that performs the E2E SA-ASR model after speaker clustering. Our evaluation on the AMI meeting corpus reveals that after fine-tuning with a small real data, the joint system performs 8.9-29.9% better in accuracy compared to the best modular system while the modular system performs better before such fine-tuning. We also conduct various error analyses to show the remaining issues for the monaural SA-ASR. Naoyuki Kanda, Jian Wu 0027, Tianyan Zhou, Yashesh Gaur, Xiaofei Wang 0009, Zhong Meng, Zhuo Chen 0006, Takuya Yoshioka |
ASRU | 4 |
| 2021 | Dual-Path Modeling for Long Recording Speech Separation in MeetingsabstractThe continuous speech separation (CSS) is a task to separate the speech sources from a long, partially overlapped recording, which involves a varying number of speakers. A straightforward extension of conventional utterance-level speech separation to the CSS task is to segment the long recording with a size-fixed window and process each window separately. Though effective, this extension fails to model the long dependency in speech and thus leads to sub-optimum performance. The recent proposed dual-path modeling could be a remedy to this problem, thanks to its capability in jointly modeling the cross-window dependency and the local-window processing. In this work, we further extend the dual-path modeling framework for CSS task. A transformer-based dual-path system is proposed, which integrates transform layers for global modeling. The proposed models are applied to LibriCSS, a real recorded multi-talk dataset, and consistent WER reduction can be observed in the ASR evaluation for separated speech. Also, a dual-path transformer equipped with convolutional layers is proposed. It significantly reduces the computation amount by 30% with better WER evaluation. Furthermore, the online processing dual-path models are investigated, which shows 10% relative WER reduction compared to the baseline. Chenda Li, Zhuo Chen 0006, Yi Luo 0004, Cong Han 0001, Tianyan Zhou, Keisuke Kinoshita, Marc Delcroix, Shinji Watanabe 0001, Yanmin Qian |
ICASSP | 5 |
| 2021 | Rethinking The Separation Layers In Speech Separation NetworksabstractModules in all existing speech separation networks can be categorized into single-input-multi-output (SIMO) modules and single-input-single-output (SISO) modules. SIMO modules generate more outputs than input, and SISO modules keep the numbers of input and output the same. While the majority of separation models only contain SIMO architectures, it has also been shown that certain two-stage separation systems integrated with a post-enhancement SISO module can improve the separation quality. Why performance improvements can be achieved by incorporating the SISO modules? Are SIMO modules always necessary? In this paper, we empirically examine those questions by designing models with varying configurations in the SIMO and SISO modules. We show that comparing with the standard SIMO-only design, a mixed SIMO-SISO design with a same model size is able to improve the separation performance especially under low-overlap conditions. We further validate the necessity of SIMO modules and show that SISO-only models are still able to perform separation without sacrificing the performance. The observations allow us to rethink the model design paradigm and present different views on how the separation is performed. Yi Luo 0004, Zhuo Chen 0006, Cong Han 0001, Chenda Li, Tianyan Zhou, Nima Mesgarani |
ICASSP | 5 |
| 2021 | Microsoft Speaker Diarization System for the Voxceleb Speaker Recognition Challenge 2020abstractThis paper describes the Microsoft speaker diarization system for monaural multi-talker recordings in the wild, evaluated at the diarization track of the VoxCeleb Speaker Recognition Challenge (VoxSRC) 2020. We will first explain our system design to address issues in handling real multi-talker recordings. We then present the details of the components, which include Res2Net-based speaker embedding extractor, conformer-based continuous speech separation with leakage filtering, and a modified DOVER (short for Diarization Output Voting Error Reduction) method for system fusion. We evaluate the systems with the data set provided by VoxSRC challenge 2020, which contains real-life multi-talker audio collected from YouTube. Our best system achieves 3.71% and 6.23% of the diarization error rate (DER) on development set and evaluation set, respectively, being ranked the 1st at the diarization track of the challenge. Naoyuki Kanda, Zhuo Chen 0006, Tianyan Zhou, Takuya Yoshioka, Sanyuan Chen, Yong Zhao 0008, Gang Liu 0001, Yu Wu 0012, Jian Wu 0027, Shujie Liu 0001, Jinyu Li 0001, Yifan Gong 0001 |
ICASSP | 4 |
| 2021 | Continuous Speech Separation Using Speaker Inventory for Long Recording
Cong Han 0001, Yi Luo 0004, Chenda Li, Tianyan Zhou, Keisuke Kinoshita, Shinji Watanabe 0001, Marc Delcroix, Hakan Erdogan, John R. Hershey, Nima Mesgarani, Zhuo Chen 0006 |
Interspeech | 4 |
| 2021 | Dual-Path RNN for Long Recording Speech SeparationabstractContinuous speech separation (CSS) is an arising task in speech separation aiming at separating overlap-free targets from a long, partially-overlapped recording. A straightforward extension of previously proposed sentence-level separation models to this task is to segment the long recording into fixed-length blocks and perform separation on them independently. However, such simple extension does not fully address the cross-block dependencies and the separation performance may not be satisfactory. In this paper, we focus on how the block-level separation performance can be improved by exploring methods to utilize the cross-block information. Based on the recently proposed dual-path RNN (DPRNN) architecture, we investigate how DPRNN can help the block-level separation by the interleaved intra- and inter-block modules. Experiment results show that DPRNN is able to significantly outperform the baseline block-level model in both offline and block-online configurations under certain settings. Chenda Li, Yi Luo 0004, Cong Han 0001, Jinyu Li 0001, Takuya Yoshioka, Tianyan Zhou, Marc Delcroix, Keisuke Kinoshita, Christoph Böddeker, Yanmin Qian, Shinji Watanabe 0001, Zhuo Chen 0006 |
SLT | 6 |
| 2021 | ResNeXt and Res2Net Structures for Speaker VerificationabstractThe ResNet-based architecture has been widely adopted to extract speaker embeddings for text-independent speaker verification systems. By introducing the residual connections to the CNN and standardizing the residual blocks, the ResNet structure is capable of training deep networks to achieve highly competitive recognition performance. However, when the input feature space becomes more complicated, simply increasing the depth and width1of the ResNet network may not fully realize its performance potential. In this paper, we present two extensions of the ResNet architecture, ResNeXt and Res2Net, for speaker verification. Originally proposed for image recognition, the ResNeXt and Res2Net introduce two more dimensions, cardinality and scale, in addition to depth and width, to improve the model's representation capacity. By increasing the scale dimension, the Res2Net model can represent multi-scale features with various granularities, which particularly facilitates speaker verification for short utterances. We evaluate our proposed systems on three speaker verification tasks. Experiments on the VoxCeleb test set demonstrated that the ResNeXt and Res2Net can significantly outperform the conventional ResNet model. The Res2Net model achieved superior performance by reducing the EER by 18.5% relative. Experiments on the other two internal test sets of mismatched conditions further confirmed the generalization of the ResNeXt and Res2Net architectures against noisy environment and segment length variations. Tianyan Zhou, Yong Zhao 0008, Jian Wu 0027 |
SLT | 1 |
| 2020 | Continuous Speech Separation: Dataset and AnalysisabstractThis paper describes a dataset and protocols for evaluating continuous speech separation algorithms. Most prior speech separation studies use pre-segmented audio signals, which are typically generated by mixing speech utterances on computers so that they fully overlap. Also, the separation algorithms have often been evaluated based on signal-based metrics such as signal-to-distortion ratio. However, in natural conversations, speech signals are continuous and contain both overlapped and overlap-free regions. In addition, the signal-based metrics only have weak correlation with automatic speech recognition (ASR) accuracy. Not only does this make it hard to assess the practical relevance of the tested algorithms, it also hinders researchers from developing systems that can be readily applied to real scenarios. In this paper, we define continuous speech separation (CSS) as a task of generating a set of non-overlapped speech signals from a continuous audio stream that contains multiple utterances that are partially overlapped by a varying degree. A new real recording dataset, called LibriCSS, is derived from LibriSpeech by concatenating the corpus utterances to simulate conversations and capturing the audio replays with far-field microphones. A Kaldi-based ASR evaluation protocol is established by using a well-trained multi-conditional acoustic model. A recently proposed speaker-independent CSS algorithm is investigated by using LibriCSS. The dataset and evaluation scripts are made available to facilitate the research in this direction1. Zhuo Chen 0006, Takuya Yoshioka, Liang Lu 0001, Tianyan Zhou, Zhong Meng, Yi Luo 0004, Jian Wu 0027, Jinyu Li 0001 |
ICASSP | 4 |
| 2020 | Improving Deep CNN Networks with Long Temporal Context for Text-Independent Speaker VerificationabstractDeep CNN networks have shown great success in various tasks for text-independent speaker recognition. In this paper, we explore two approaches for modeling long temporal contexts to improve the performance of the ResNet networks. The first approach is simply integrating the utterance-level mean and variance normalization into the ResNet architecture. Secondly, we combine the BLSTM and ResNet into one unified architecture. The BLSTM layers model long range, supposedly phonetically aware, context information, which could facilitate the ResNet to learn the optimal attention weight and suppress the environmental variations. The BLSTM outputs are projected into multiple-channel feature maps and fed into the ResNet network. Experiments on the VoxCeleb1 and the internal MS-SV tasks show that with attentive pooling, the proposed approaches achieve up to 23-28% relative improvement in EER over a well-trained ResNet. Yong Zhao 0008, Tianyan Zhou, Zhuo Chen 0006, Jian Wu 0027 |
ICASSP | 2 |
| 2020 | Joint Speaker Counting, Speech Recognition, and Speaker Identification for Overlapped Speech of any Number of SpeakersabstractWe propose an end-to-end speaker-attributed automatic speech recognition model that unifies speaker counting, speech recognition, and speaker identification on monaural overlapped speech.Our model is built on serialized output training (SOT) with attention-based encoder-decoder, a recently proposed method for recognizing overlapped speech comprising an arbitrary number of speakers.We extend SOT by introducing a speaker inventory as an auxiliary input to produce speaker labels as well as multi-speaker transcriptions.All model parameters are optimized by speaker-attributed maximum mutual information criterion, which represents a joint probability for overlapped speech recognition and speaker identification.Experiments on LibriSpeech corpus show that our proposed method achieves significantly better speaker-attributed word error rate than the baseline that separately performs overlapped speech recognition and speaker identification. Naoyuki Kanda, Yashesh Gaur, Xiaofei Wang 0009, Zhong Meng, Zhuo Chen 0006, Tianyan Zhou, Takuya Yoshioka |
INTERSPEECH | 6 |
| 2019 | Optimizing Neural Network Embeddings Using a Pair-Wise Loss for Text-Independent Speaker VerificationabstractThis paper proposes a new loss function called the “quartet” loss for the better optimization of the neural networks for matching tasks. For such tasks, where neural network embeddings are the key component, the optimization of the network for better embeddings is critical. The embeddings are required to be class discriminative, resulting in minimal inter-class variation and maximal intra-class variation even for unseen classes for better generalization of the network. The quartet loss explicitly computes the distance metric between pairs of inputs and increases the gap between the similarity score distributions between the same class pairs and the different class pairs. We evaluate on the speaker verification task and demonstrate the performance of the loss on our proposed neural network. Hira Dhamyal, Tianyan Zhou, Bhiksha Raj, Rita Singh |
ASRU | 2 |
| 2019 | Speech Separation Using Speaker InventoryabstractOverlapped speech is one of the main challenges in conversational speech applications such as meeting transcription. Blind speech separation and speech extraction are two common approaches to this problem. Both of them, however, suffer from limitations resulting from the lack of abilities to either leverage additional information or process multiple speakers simultaneously. In this work, we propose a novel method called speech separation using speaker inventory (SSUSI), which combines the advantages of both approaches and thus solves their problems. SSUSI makes use of a speaker inventory, i.e. a pool of pre-enrolled speaker signals, and jointly separates all participating speakers. This is achieved by a specially designed attention mechanism, eliminating the need for accurate speaker identities. Experimental results show that SSUSI outperforms permutation invariant training based blind speech separation by up to 48% relatively in word error rate (WER). Compared with speech extraction, SSUSI reduces computation time by up to 70% and improves the WER by more than 13% relatively. Zhuo Chen 0006, Zhong Meng, Takuya Yoshioka, Tianyan Zhou, Liang Lu 0001, Jinyu Li 0001 |
ASRU | 6 |
| 2019 | Advances in Online Audio-Visual Meeting TranscriptionabstractThis paper describes a system that generates speaker-annotated transcripts of meetings by using a microphone array and a 360-degree camera. The hallmark of the system is its ability to handle overlapped speech, which has been an unsolved problem in realistic settings for over a decade. We show that this problem can be addressed by using a continuous speech separation approach. In addition, we describe an online audio-visual speaker diarization method that leverages face tracking and identification, sound source localization, speaker identification, and, if available, prior speaker information for robustness to various real world challenges. All components are integrated in a meeting transcription framework called SRD, which stands for “separate, recognize, and diarize”. Experimental results using recordings of natural meetings involving up to 11 attendees are reported. The continuous speech separation improves a word error rate (WER) by 16.1% compared with a highly tuned beamformer. When a complete list of meeting attendees is available, the discrepancy between WER and speaker-attributed WER is only 1.0%, indicating accurate word-to-speaker association. This increases marginally to 1.6% when 50% of the attendees are unknown to the system. Takuya Yoshioka, Yan Huang 0028, Aviv Hurvitz, Sharon Koubi, Eyal Krupka, Ido Leichter, Changliang Liu, Partha Parthasarathy, Alon Vinnikov, Lingfeng Wu, Igor Abramovski, Wayne Xiong, Huaming Wang, Jun Zhang 0066, Yong Zhao 0008, Tianyan Zhou, Cem Aksoylar, Zhuo Chen 0006, Moshe David, Dimitrios Dimitriadis, Yifan Gong 0001, Ilya Gurvich, Xuedong Huang 0001 |
ASRU | 19 |
| 2019 | CNN with Phonetic Attention for Text-Independent Speaker VerificationabstractText-independent speaker verification imposes no constraints on the spoken content and usually needs long observations to make reliable prediction. In this paper, we propose two speaker embedding approaches by integrating the phonetic information into the attention-based residual convolutional neural network (CNN). Phonetic features are extracted from the bottleneck layer of a pretrained acoustic model. In implicit phonetic attention (IPA), the phonetic features are projected by a transformation network into multi-channel feature maps, and then combined with the raw acoustic features as the input of the CNN network. In explicit phonetic attention (EPA), the phonetic features are directly connected to the attentive pooling layer through a separate 1-dim CNN to generate the attention weights. With the incorporation of spoken content and attention mechanism, the system can not only distill the speaker-discriminant frames but also actively normalize the phonetic variations. Multi-head attention and discriminative objectives are further studied to improve the system. Experiments on the VoxCeleb corpus show our proposed system could outperform the state-of-the-art by around 43% relative. Tianyan Zhou, Yong Zhao 0008, Jinyu Li 0001, Yifan Gong 0001, Jian Wu 0027 |
ASRU | 1 |
| 2019 | An automated assessment framework for atypical prosody and stereotyped idiosyncratic phrases related to autism spectrum disorder
Ming Li 0026, Dengke Tang, Junlin Zeng, Tianyan Zhou, Huilin Zhu, Biyuan Chen, Xiaobing Zou |
Comput. Speech Lang. | 4 |
| 2018 | A Corrective Learning Approach for Text-Independent Speaker VerificationabstractWe present a conceptually plausible approach for text-independent speaker verification (TISV) which treats speech recordings as a collection of segments providing incremental evidence. This approach, called corrective learning, gradually improves an initial prediction of speaker identity based on incoming speech and the latest prediction. Specifically, we propose deep corrective learning networks (CLNets) that explicitly learn a mapping from a new speech segment and the current predictions, to a correction. Intuitively, the predictions eventually converge to the ground truth after several corrections. Trained on NIST SRE datasets, CLNets outperform current CNN and the i-vector baselines. Moreover, CLNets and i-vectors are complementary, and their fusion leads to significant performance improvements compared to what can be achieved by each of them individually. Yandong Wen, Tianyan Zhou, Rita Singh, Bhiksha Raj |
ICASSP | 2 |
| 2017 | Response to name: A dataset and a multimodal machine learning framework towards autism studyabstractIn this paper, we propose a “Response to Name Dataset” for autism spectrum disorder (ASD) study as well as a multimodal ASD auxiliary screening system based on machine learning. ASD children are characterized by their impaired interpersonal communication abilities and lack of response. In the proposed dataset, the reactions of children are recorded by cameras upon calling their names. The responsiveness of each child is then evaluated by a clinician with a score among 0, 1 and 2 following the Autism Diagnostic Observation Schedule (ADOS). We then develop a rule-based multimodal framework to quantitatively evaluate each child. Our system involves speech recognition based automatic name calling detection, face detection/alignment, head pose estimation, and considers the response speed, eye contact duration and head orientation to output the final prediction. Compared to existing work, our dataset characterizes a more precise and detailed scoring system with clinical trial standards, as well as a more spontaneous setting by incorporating less lab-controlled sessions with dynamic/cluttered environments, multi-pose mobile captured videos, and flexible number of accompanying adults. Experiments show that our machine predicted scores align closely with human professional diagnosis, showing promising potential in early screening of ASD, and shedding light on future clinical applications. Wenbo Liu 0002, Tianyan Zhou, Xiaobing Zou, Ming Li 0026 |
ACII | 2 |
| 2016 | Text-independent voice conversion using deep neural network based phonetic level featuresabstractThis paper presents a phonetically-aware joint density Gaussian mixture model (JD-GMM) framework for voice conversion that no longer requires parallel data from source speaker at the training stage. Considering that the phonetic level features contain text information which should be preserved in the conversion task, we propose a method that only concatenates phonetic discriminant features and spectral features extracted from the same target speakers speech to train a JD-GMM. After the mapping relationship of these two features is trained, we can use phonetic discriminant features from source speaker to estimate target speaker's spectral features at conversion stage. The phonetic discriminant features are extracted using PCA from the output layer of a deep neural network (DNN) in an automatic speaker recognition (ASR) system. It can be seen as a low dimensional representation of the senone posteriors. We compare the proposed phonetically-aware method with conventional JD-GMM method on the Voice Conversion Challenge 2016 training database. The experimental results show that our proposed phonetically-aware feature method can obtain similar performance compared to the conventional JD-GMM in the case of using only target speech as training data. Huadi Zheng, Weicheng Cai, Tianyan Zhou, Shilei Zhang, Ming Li 0026 |
ICPR | 3 |