Zhiyun Fan

dblp:251/8579 · DBLP profile ↗
← Back
7ranked-venue papers
7as first author
6since 2021 · last 2024
0000-0001-9180-7392ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 5 first-author · 4 since 2021
YearPublicationVenuePosition
2024 SA-SOT: Speaker-Aware Serialized Output Training for Multi-Talker ASR
abstract
Multi-talker automatic speech recognition plays a crucial role in scenarios involving multi-party interactions, such as meetings and conversations. Due to its inherent complexity, this task has been receiving increasing attention. Notably, the serialized output training (SOT) stands out among various approaches because of its simplistic architecture and exceptional performance. However, the frequent speaker changes in token-level SOT (t-SOT) present challenges for the autoregressive decoder in effectively utilizing context to predict output sequences. To address this issue, we introduce a masked t-SOT label, which serves as the cornerstone of an auxiliary training loss. Additionally, we utilize a speaker similarity matrix to refine the self-attention mechanism of the decoder. This strategic adjustment enhances contextual relationships within the same speaker’s tokens while minimizing interactions between different speakers’ tokens. We denote our method as speaker-aware SOT (SA-SOT). Experiments on the Librispeech datasets demonstrate that our SA-SOT obtains a relative cpWER reduction ranging from 12.75% to 22.03% on the multi-talker test sets. Furthermore, with more extensive training, our method achieves an impressive cpWER of 3.41%, establishing a new state-of-the-art result on the LibrispeechMix dataset.
Zhiyun Fan, Linhao Dong, Jun Zhang 0066, Lu Lu 0015, Zejun Ma 0001
ICASSP1
2023 Language-specific Boundary Learning for Improving Mandarin-English Code-switching Speech Recognition
Zhiyun Fan, Linhao Dong, Chen Shen 0011, Zhenlin Liang, Jun Zhang 0066, Lu Lu 0015, Zejun Ma 0001
INTERSPEECH1
2022 Token-level Speaker Change Detection Using Speaker Difference and Speech Content via Continuous Integrate-and-fire
Zhiyun Fan, Zhenlin Liang, Linhao Dong, Jun Zhang 0066, Zejun Ma 0001, Bo Xu 0002
INTERSPEECH1
2022 Sequence-Level Speaker Change Detection With Difference-Based Continuous Integrate-and-Fire
abstract
Speaker change detection is an important task in multi-party interactions such as meetings and conversations. In this paper, we address the speaker change detection task from the perspective of sequence transduction. Specifically, we propose a novel encoder-decoder framework that directly converts the input feature sequence to the speaker identity sequence. The difference-based continuous integrate-and-fire mechanism is designed to support this framework. It detects speaker changes by integrating the speaker difference between the encoder outputs frame-by-frame and transfers encoder outputs to segment-level speaker embeddings according to the detected speaker changes. The whole framework is supervised by the speaker identity sequence, a weaker label than the precise speaker change points. The experiments on the AMI and DIHARD-I corpora show that our sequence-level method consistently outperforms a strong frame-level baseline that uses the precise speaker change labels.
Zhiyun Fan, Linhao Dong, Zejun Ma 0001, Bo Xu 0002
IEEE Signal Process. Lett.1
2021 Two-Stage Pre-Training for Sequence to Sequence Speech Recognition
abstract
The attention-based encoder-decoder structure is popular in automatic speech recognition (ASR). However, it relies heavily on transcribed data. In this paper, we propose a novel pre-training strategy for the encoder-decoder sequence-to-sequence (seq2seq) model by utilizing unpaired speech and transcripts. The pre-training process consists of two stages, acoustic pre-training and linguistic pre-training. In the acoustic pre-training stage, we use a large amount of speech to pre-train the encoder by predicting masked speech feature chunks with their contexts. In the linguistic pre-training stage, we first generate synthesized speech from a large number of transcripts using a text-to-speech (TTS) system and then use the synthesized paired data to pre-train the decoder. The two-stage pre-training is conducted on the AISHELL-2 dataset, and we apply this pre-trained model to multiple subsets of AISHELL-1 and HKUST for post-training. As the size of the subset increases, we obtain relative character error rate reduction (CERR) from 38.24% to 7.88% on AISHELL-1 and from 12.00% to 1.20% on HKUST.
Zhiyun Fan, Bo Xu 0002
IJCNN1
2021 Exploring wav2vec 2.0 on Speaker Verification and Language Identification
abstract
Wav2vec 2.0 is a recently proposed self-supervised framework for speech representation learning.It follows a two-stage training process of pre-training and fine-tuning, and performs well in speech recognition tasks especially ultra-low resource cases.In this work, we attempt to extend the self-supervised framework to speaker verification and language identification.First, we use some preliminary experiments to indicate that wav2vec 2.0 can capture the information about the speaker and language.Then we demonstrate the effectiveness of wav2vec 2.0 on the two tasks respectively.For speaker verification, we obtain a new state-of-the-art result, Equal Error Rate (EER) of 3.61% on the VoxCeleb1 dataset.For language identification, we obtain an EER of 12.02% on the 1 second condition and an EER of 3.47% on the full-length condition of the AP17-OLR dataset.Finally, we utilize one model to achieve the unified modeling by the multi-task learning for the two tasks.
Zhiyun Fan, Bo Xu 0002
Interspeech1
2019 Speaker-Aware Speech-Transformer
abstract
Recently, end-to-end (E2E) models become a competitive alternative to the conventional hybrid automatic speech recognition (ASR) systems. However, they still suffer from speaker mismatch in training and testing condition. In this paper, we use Speech-Transformer (ST) as the study platform to investigate speaker aware training of E2E models. We propose a model called Speaker-Aware Speech-Transformer (SAST), which is a standard ST equipped with a speaker attention module (SAM). The SAM has a static speaker knowledge block (SKB) that is made of i-vectors. At each time step, the encoder output attends to the i-vectors in the block, and generates a weighted combined speaker embedding vector, which helps the model to normalize the speaker variations. The SAST model trained in this way becomes independent of specific training speakers and thus generalizes better to unseen testing speakers. We investigate different factors of SAM. Experimental results on the AISHELL-1 task show that SAST achieves a relative 6.5% CER reduction (CERR) over the speaker-independent (SI) baseline. Moreover, we demonstrate that SAST still works quite well even if the i-vectors in SKB all come from a different data source other than the acoustic training set.
Zhiyun Fan, Jie Li 0032, Bo Xu 0002
ASRU1