Huan Zhou 0008

dblp:78/6138-8 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 SCDiar: a streaming diarization system based on speaker change detection and speech recognition
abstract
In hours-long meeting scenarios, real-time speech stream often struggles with achieving accurate speaker diarization, commonly leading to speaker identification and speaker count errors. To address this challenge, we propose SCDiar, a system that operates on speech segments, split at the token level by a speaker change detection (SCD) module. Building on these segments, we introduce several enhancements to efficiently select the best available segment for each speaker. These improvements lead to significant gains across various benchmarks. Notably, on real-world meeting data involving more than ten participants, SCDiar outperforms previous systems by up to 53.6% in accuracy, substantially narrowing the performance gap between online and offline systems.
Naijun Zheng, Xucheng Wan, Kai Liu 0053, Huan Zhou 0008
ICASSP4
2024 Real-time scheme for rapid extraction of speaker embeddings in challenging recording conditions
Kai Liu 0053, Ziqing Du, Huan Zhou 0008, Xucheng Wan, Naijun Zheng
INTERSPEECH3
2024 An efficient text augmentation approach for contextualized Mandarin speech recognition
Naijun Zheng, Xucheng Wan, Kai Liu 0053, Ziqing Du, Huan Zhou 0008
INTERSPEECH5
2023 X-SEPFORMER: End-To-End Speaker Extraction Network with Explicit Optimization on Speaker Confusion
abstract
Target speech extraction (TSE) systems are designed to extract target speech from a multi-talker mixture. The popular training objective for most prior TSE networks is to enhance reconstruction performance of extracted speech waveform. However, it has been reported that a TSE system delivers high reconstruction performance may still suffer low-quality experience problems in practice. One such experience problem is wrong speaker extraction (called speaker confusion, SC), which leads to strong negative experience and hampers effective conversations. To mitigate the imperative SC issue, we reformulate the training objective and propose two novel loss schemes that explore the metric of reconstruction improvement performance defined at small chunk-level and leverage the metric associated distribution information. Both loss schemes aim to encourage a TSE network to pay attention to those SC chunks based on the said distribution information. On this basis, we present X-SepFormer, an end-to-end TSE model with proposed loss schemes and a backbone of SepFormer. Experimental results on the benchmark WSJ0-2mix dataset validate the effectiveness of our proposals, showing consistent improvements on SC errors (by 14.8% relative). Moreover, with SI-SDRi of 19.4 dB and PESQ of 3.81, our best system significantly outperforms the current SOTA systems and offers the top TSE results reported till date on the WSJ0-2mix.
Kai Liu 0053, Ziqing Du, Xucheng Wan, Huan Zhou 0008
ICASSP4
2021 Online Speaker Diarization Equipped with Discriminative Modeling and Guided Inference
Xucheng Wan, Kai Liu 0053, Huan Zhou 0008
Interspeech3
2020 Text-Independent Speaker Verification with Adversarial Learning on Short Utterances
abstract
A text-independent speaker verification system suffers severe performance degradation under short utterance condition. To address the problem, in this paper, we propose an adversarially learned embedding mapping model that directly maps a short embedding to an enhanced embedding with increased discriminability. In particular, a Wasserstein GAN with a bunch of loss criteria are investigated. These loss functions have distinct optimization objectives and some of them are less favoured for the speaker verification research area. Different from most prior studies, our main objective in this study is to investigate the effectiveness of those loss criteria by conducting numerous ablation studies. Experiments on Voxceleb dataset showed that some criteria are beneficial to the verification performance while some have trivial effects. Lastly, a Wasserstein GAN with chosen loss criteria, without finetuning, achieves meaningful advancements over the baseline, with 4% relative improvements on EER and 7% on minDCF in the challenging scenario of short 2second utterances.
Kai Liu 0053, Huan Zhou 0008
ICASSP2
2020 Speech Emotion Recognition with Discriminative Feature Learning
Huan Zhou 0008, Kai Liu 0053
INTERSPEECH1