Jun Zhang 0066

dblp:29/4190-66 · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 9 since 2021
YearPublicationVenuePosition
2025 QualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and Descriptions
abstract
This paper explores a novel perspective to speech quality assessment by leveraging natural language descriptions, offering richer, more nuanced insights than traditional numerical scoring methods. Natural language feedback provides instructive recommendations and detailed evaluations, yet existing datasets lack the comprehensive annotations needed for this approach. To bridge this gap, we introduce QualiSpeech, a comprehensive low-level speech quality assessment dataset encompassing 11 key aspects and detailed natural language comments that include reasoning and contextual insights. Additionally, we propose the QualiSpeech Benchmark to evaluate the low-level speech understanding capabilities of auditory large language models (LLMs). Experimental results demonstrate that finetuned auditory LLMs can reliably generate detailed descriptions of noise and distortion, effectively identifying their types and temporal characteristics. The results further highlight the potential for incorporating reasoning to enhance the accuracy and reliability of quality assessments. The dataset can be found at https://huggingface.co/datasets/tsinghua-ee/QualiSpeech.
Siyin Wang, Wenyi Yu, Xianzhao Chen, Xiaohai Tian, Jun Zhang 0066, Lu Lu 0015, Yu Tsao 0001, Junichi Yamagishi, Yuxuan Wang 0002, Chao Zhang 0031
ACL (1)5
2024 SA-SOT: Speaker-Aware Serialized Output Training for Multi-Talker ASR
abstract
Multi-talker automatic speech recognition plays a crucial role in scenarios involving multi-party interactions, such as meetings and conversations. Due to its inherent complexity, this task has been receiving increasing attention. Notably, the serialized output training (SOT) stands out among various approaches because of its simplistic architecture and exceptional performance. However, the frequent speaker changes in token-level SOT (t-SOT) present challenges for the autoregressive decoder in effectively utilizing context to predict output sequences. To address this issue, we introduce a masked t-SOT label, which serves as the cornerstone of an auxiliary training loss. Additionally, we utilize a speaker similarity matrix to refine the self-attention mechanism of the decoder. This strategic adjustment enhances contextual relationships within the same speaker’s tokens while minimizing interactions between different speakers’ tokens. We denote our method as speaker-aware SOT (SA-SOT). Experiments on the Librispeech datasets demonstrate that our SA-SOT obtains a relative cpWER reduction ranging from 12.75% to 22.03% on the multi-talker test sets. Furthermore, with more extensive training, our method achieves an impressive cpWER of 3.41%, establishing a new state-of-the-art result on the LibrispeechMix dataset.
Zhiyun Fan, Linhao Dong, Jun Zhang 0066, Lu Lu 0015, Zejun Ma 0001
ICASSP3
2024 Can Large Language Models Understand Spatial Audio?
Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan 0019, Wei Li 0119, Jun Zhang 0066, Lu Lu 0015, Zejun Ma 0001, Yuxuan Wang 0002, Chao Zhang 0031
INTERSPEECH7
2024 SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words
abstract
Speech encompasses a wealth of information, including but not limited to content, paralinguistic, and environmental information.This comprehensive nature of speech significantly impacts communication and is crucial for human-computer interaction.Chat-Oriented Large Language Models (LLMs), known for their general-purpose assistance capabilities, have evolved to handle multi-modal inputs, including speech.Although these models can be adept at recognizing and analyzing speech, they often fall short of generating appropriate responses.We argue that this is due to the lack of principles on task definition and model development, which requires open-source datasets and metrics suitable for model evaluation.To bridge the gap, we present SD-Eval, a benchmark dataset aimed at multidimensional evaluation of spoken dialogue understanding and generation.SD-Eval focuses on paralinguistic and environmental information and includes 7,303 utterances, amounting to 8.76 hours of speech data. The data is aggregated from eight public datasets, representing four perspectives: emotion, accent, age, and background sound.To assess the SD-Eval benchmark dataset, we implement three different models and construct a training set following a process similar to that of SD-Eval. The training set contains 1,052.72 hours of speech data and 724.4k utterances. We also conduct a comprehensive evaluation using objective evaluation methods (e.g. BLEU and ROUGE), subjective evaluations and LLM-based metrics for the generated responses.Models conditioned with paralinguistic and environmental information outperform their counterparts in both objective and subjective measures.Moreover, experiments demonstrate that LLM-based metrics show a higher correlation with human evaluation compared to traditional metrics.We open-source SD-Eval at https://github.com/amphionspace/SD-Eval.
Junyi Ao, Yuancheng Wang, Xiaohai Tian, Dekun Chen, Jun Zhang 0066, Lu Lu 0015, Yuxuan Wang 0002, Haizhou Li 0001, Zhizheng Wu 0001
NeurIPS5
2023 Improving Large-Scale Deep Biasing With Phoneme Features and Text-Only Data in Streaming Transducer
abstract
Deep biasing for the Transducer can improve the recognition performance of rare words or contextual entities, which is essential in practical applications, especially for streaming Automatic Speech Recognition (ASR). However, deep biasing with large-scale rare words remains challenging, as the performance drops significantly when more distractors exist and there are words with similar grapheme sequences in the bias list. In this paper, we combine the phoneme and textual information of rare words in Transducers to distinguish words with similar pronunciation or spelling. Moreover, the introduction of training with text-only data containing more rare words benefits large-scale deep biasing. The experiments on the Librispeech corpus demonstrate that the proposed method achieves state-of-the-art performance on rare word error rate for different scales and levels of bias lists.
Jin Qiu, Jun Zhang 0066, Lu Lu 0015, Zejun Ma 0001
ASRU4
2023 Language-specific Boundary Learning for Improving Mandarin-English Code-switching Speech Recognition
Zhiyun Fan, Linhao Dong, Chen Shen 0011, Zhenlin Liang, Jun Zhang 0066, Lu Lu 0015, Zejun Ma 0001
INTERSPEECH5
2023 Text-only Domain Adaptation using Unified Speech-Text Representation in Transducer
Jun Zhang 0066, Lu Lu 0015, Zejun Ma 0001
INTERSPEECH3
2022 The Volcspeech System for the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Challenge
abstract
This paper describes our submission to ICASSP 2022 Multi-channel Multi-party Meeting Transcription (M2MeT) Challenge. For Track 1, we propose several approaches to make the clustering-based speaker diarization system enable to handle overlapped speech. Front-end dereverberation and the direction-of-arrival (DOA) estimation are used to improve the accuracy of speaker diarization. Multi-channel combination and overlap detection are applied to reduce the missed speaker error. A modified DOVER-Lap is also proposed to fuse the results from different systems. We achieve the final DER of 5.79% on the Eval set and 7.23% on the Test set, which ranks 4th in the diarization challenge. For Track 2, we develop our system using the Conformer model in a joint CTC-attention architecture. Serialized output training (SOT) is adopted to multi-speaker overlapped speech recognition. We propose a neural front-end module to model multi-channel audio and train the model end-to-end. Various data augmentation methods are utilized to mitigate over-fitting in the multi-channel multi-speaker E2E system. Transformer language model fusion is developed to achieve better performance. The final CER is 19.2% on the Eval set and 20.8% on the Test set, which ranks 2nd in the ASR challenge.
Chen Shen 0011, Wenzhi Fan, Shixue Wen, Jun Zhang 0066, Jingsheng Yang, Zejun Ma 0001
ICASSP7
2022 Token-level Speaker Change Detection Using Speaker Difference and Speech Content via Continuous Integrate-and-fire
Zhiyun Fan, Zhenlin Liang, Linhao Dong, Jun Zhang 0066, Zejun Ma 0001, Bo Xu 0002
INTERSPEECH7
2022 Bring dialogue-context into RNN-T for streaming ASR
Junfeng Hou, Jinkun Chen, Yufeng Tang, Jun Zhang 0066, Zejun Ma 0001
INTERSPEECH5
2021 HMM-Free Encoder Pre-Training for Streaming RNN Transducer
abstract
This work describes an encoder pre-training procedure using frame-wise label to improve the training of streaming recurrent neural network transducer (RNN-T) model.Streaming RNN-T trained from scratch usually performs worse than nonstreaming RNN-T.Although it is common to address this issue through pre-training components of RNN-T with other criteria or frame-wise alignment guidance, the alignment is not easily available in end-to-end manner.In this work, frame-wise alignment, used to pre-train streaming RNN-T's encoder, is generated without using a HMM-based system.Therefore an allneural framework equipping HMM-free encoder pre-training is constructed.This is achieved by expanding the spikes of CTC model to their left/right blank frames, and two expanding strategies are proposed.To our best knowledge, this is the first work to simulate HMM-based frame-wise label using CTC model for pre-training.Experiments conducted on LibriSpeech and MLS English tasks show the proposed pre-training procedure, compared with random initialization, reduces the WER by relatively 5%∼11% and the emission latency by 60 ms.Besides, the method is lexicon-free, so it is friendly to new languages without manually designed lexicon.
Jingyu Sun, Yufeng Tang, Junfeng Hou, Jinkun Chen, Jun Zhang 0066, Zejun Ma 0001
Interspeech6
2019 Advances in Online Audio-Visual Meeting Transcription
abstract
This paper describes a system that generates speaker-annotated transcripts of meetings by using a microphone array and a 360-degree camera. The hallmark of the system is its ability to handle overlapped speech, which has been an unsolved problem in realistic settings for over a decade. We show that this problem can be addressed by using a continuous speech separation approach. In addition, we describe an online audio-visual speaker diarization method that leverages face tracking and identification, sound source localization, speaker identification, and, if available, prior speaker information for robustness to various real world challenges. All components are integrated in a meeting transcription framework called SRD, which stands for “separate, recognize, and diarize”. Experimental results using recordings of natural meetings involving up to 11 attendees are reported. The continuous speech separation improves a word error rate (WER) by 16.1% compared with a highly tuned beamformer. When a complete list of meeting attendees is available, the discrepancy between WER and speaker-attributed WER is only 1.0%, indicating accurate word-to-speaker association. This increases marginally to 1.6% when 50% of the attendees are unknown to the system.
Takuya Yoshioka, Yan Huang 0028, Aviv Hurvitz, Sharon Koubi, Eyal Krupka, Ido Leichter, Changliang Liu, Partha Parthasarathy, Alon Vinnikov, Lingfeng Wu, Igor Abramovski, Wayne Xiong, Huaming Wang, Jun Zhang 0066, Yong Zhao 0008, Tianyan Zhou, Cem Aksoylar, Zhuo Chen 0006, Moshe David, Dimitrios Dimitriadis, Yifan Gong 0001, Ilya Gurvich, Xuedong Huang 0001
ASRU17