VLDB 2026 Research / reviewers in the wild / expert
Jiawen Kang 0002
dblp:145/4239-2
· DBLP profile ↗
18ranked-venue papers
4as first author
16since 2021 · last 2025
0000-0002-8218-3490ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 14 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTCabstractMulti-talker speech recognition (MTASR) faces unique challenges in disentangling and transcribing overlapping speech. To address these challenges, this paper investigates the role of Connectionist Temporal Classification (CTC) in speaker disentanglement when incorporated with Serialized Output Training (SOT) for MTASR. Our visualization reveals that CTC guides the encoder to represent different speakers in distinct temporal regions of acoustic embeddings. Leveraging this insight, we propose a novel Speaker-Aware CTC (SACTC) training objective, based on the Bayes risk CTC framework. SACTC is a tailored CTC variant for multi-talker scenarios, it explicitly models speaker disentanglement by constraining the encoder to represent different speakers’ tokens at specific time frames. When integrated with SOT, the SOT-SACTC model consistently outperforms standard SOT-CTC across various degrees of speech overlap. Specifically, we observe relative word error rate reductions of 10% overall and 15% on low-overlap speech. This work represents an initial exploration of CTC-based enhancements for MTASR tasks, offering a new perspective on speaker disentanglement in multi-talker speech recognition .1 Jiawen Kang 0002, Lingwei Meng, Yuejiao Wang, Xixin Wu, Xunying Liu, Helen M. Meng |
ICASSP | 1 |
| 2025 | Large Language Model Can Transcribe Speech in Multi-Talker Scenarios with Versatile InstructionsabstractRecent advancements in large language models (LLMs) have revolutionized various domains, bringing significant progress and new opportunities. Despite progress in speech-related tasks, LLMs have not been sufficiently explored in multi-talker scenarios. In this work, we present a pioneering effort to investigate the capability of LLMs in transcribing speech in multi-talker environments, following versatile instructions related to multi-talker automatic speech recognition (ASR), target talker ASR, and ASR based on specific talker attributes such as sex, occurrence order, language, and keyword spoken. Our approach utilizes WavLM and Whisper encoder to extract multi-faceted speech representations that are sensitive to speaker characteristics and semantic context. These representations are then fed into an LLM fine-tuned using LoRA, enabling the capabilities for speech comprehension and transcription. Comprehensive experiments reveal the promising performance of our proposed system, MT-LLM, in cocktail party scenarios, highlighting the potential of LLM to handle speech-related tasks based on user instructions in such complex settings1. Lingwei Meng, Shujie Hu, Jiawen Kang 0002, Zhaoqing Li, Yuejiao Wang, Xixin Wu, Xunying Liu, Helen M. Meng |
ICASSP | 3 |
| 2025 | On the Within-class Variation Issue in Alzheimer's Disease Detection
Jiawen Kang 0002, Dongrui Han, Lingwei Meng, Jingyan Zhou, Jinchao Li, Xixin Wu, Helen M. Meng |
INTERSPEECH | 1 |
| 2025 | Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
Yifan Yang 0005, Jiajun Deng, Jiawen Kang 0002, Shujie Hu, Tianzi Wang, Zhaoqing Li, Shiliang Zhang, Xie Chen 0001, Xunying Liu |
INTERSPEECH | 4 |
| 2024 | Cross-Speaker Encoding Network for Multi-Talker Speech RecognitionabstractEnd-to-end multi-talker speech recognition has garnered great interest as an effective approach to directly transcribe overlapped speech from multiple speakers. Current methods typically adopt either 1) single-input multiple-output (SIMO) models with a branched encoder, or 2) single-input single-output (SISO) models based on attention-based encoder-decoder architecture with serialized output training (SOT). In this work, we propose a Cross-Speaker Encoding (CSE) network to address the limitations of SIMO models by aggregating cross-speaker representations. Furthermore, the CSE model is integrated with SOT to leverage both the advantages of SIMO and SISO while mitigating their drawbacks. To the best of our knowledge, this work represents an early effort to integrate SIMO and SISO for multi-talker speech recognition. Experiments on the two-speaker LibrispeechMix dataset show that the CES model reduces word error rate (WER) by 8% over the SIMO baseline. The CSE-SOT model reduces WER by 10% overall and by 16% on high-overlap speech compared to the SOT model. Jiawen Kang 0002, Lingwei Meng, Haohan Guo, Xixin Wu, Xunying Liu, Helen M. Meng |
ICASSP | 1 |
| 2024 | Empowering Whisper as a Joint Multi-Talker and Target-Talker Speech Recognition System
Lingwei Meng, Jiawen Kang 0002, Yuejiao Wang, Zengrui Jin, Xixin Wu, Xunying Liu, Helen M. Meng |
INTERSPEECH | 2 |
| 2023 | A Sidecar Separator Can Convert A Single-Talker Speech Recognition System to A Multi-Talker OneabstractAlthough automatic speech recognition (ASR) can perform well in common non-overlapping environments, sustaining performance in multi-talker overlapping speech recognition remains challenging. Recent research revealed that ASR model’s encoder captures different levels of information with different layers – the lower layers tend to have more acoustic information, and the upper layers more linguistic. This inspires us to develop a Sidecar separator to empower a well-trained ASR model for multi-talker scenarios by separating the mixed speech embedding between two suitable layers. We experimented with a wav2vec 2.0-based ASR model with a Sidecar mounted. By freezing the parameters of the original model and training only the Sidecar (8.7 M, 8.4% of all parameters), the proposed approach outperforms the previous state-of-the-art by a large margin for the 2-speaker mixed LibriMix dataset, reaching a word error rate (WER) of 10.36%; and obtains comparable results (7.56%) for LibriSpeechMix dataset when limited training. Lingwei Meng, Jiawen Kang 0002, Yuejiao Wang, Xixin Wu, Helen M. Meng |
ICASSP | 2 |
| 2023 | Towards Effective and Compact Contextual Representation for Conformer Transducer Speech Recognition Systems
Jiawen Kang 0002, Jiajun Deng, Xi Yin 0010, Xie Chen 0001, Xunying Liu |
INTERSPEECH | 2 |
| 2023 | Unified Modeling of Multi-Talker Overlapped Speech Recognition and Diarization with a Sidecar Separator
Lingwei Meng, Jiawen Kang 0002, Xixin Wu, Helen M. Meng |
INTERSPEECH | 2 |
| 2023 | Integrated and Enhanced Pipeline System to Support Spoken Language Analytics for Screening Neurocognitive Disordersabstract24th Annual Conference of the International Speech Communication Association, INTERSPEECH 2023, Dublin, Ireland, August 20-24, 2023 Helen M. Meng, Brian Kan-Wing Mak, Man-Wai Mak, Helene H. Fung, Xianmin Gong, Timothy C. Y. Kwok, Xunying Liu, Vincent C. T. Mok, Patrick C. M. Wong, Jean Woo, Xixin Wu, Ka-Ho Wong, Sean Shensheng Xu, Naijun Zheng, Ranzo Huang, Jiawen Kang 0002, Xiaoquan Ke, Junan Li, Jinchao Li |
INTERSPEECH | 16 |
| 2022 | TalkTive: A Conversational Agent Using Backchannels to Engage Older Adults in Neurocognitive Disorders ScreeningabstractConversational agents (CAs) have the great potential in mitigating the clinicians’ burden in screening for neurocognitive disorders among older adults. It is important, therefore, to develop CAs that can be engaging, to elicit conversational speech input from older adult participants for supporting assessment of cognitive abilities. As an initial step, this paper presents research in developing the backchanneling ability in CAs in the form of a verbal response to engage the speaker. We analyzed 246 conversations of cognitive assessments between older adults and human assessors, and derived the categories of reactive backchannels (e.g. “hmm”) and proactive backchannels (e.g. “please keep going”). This is used in the development of TalkTive, a CA which can predict both timing and form of backchanneling during cognitive assessments. The study then invited 36 older adult participants to evaluate the backchanneling feature. Results show that proactive backchanneling is more appreciated by participants than reactive backchanneling. Zijian Ding, Jiawen Kang 0002, Tinky Oi Ting Ho, Ka-Ho Wong, Helene H. Fung, Helen M. Meng, Xiaojuan Ma |
CHI | 2 |
| 2022 | The CUHK-Tencent Speaker Diarization System for the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription ChallengeabstractThis paper describes our speaker diarization system submitted to the Multi-channel Multi-party Meeting Transcription (M2MeT) challenge, where Mandarin meeting data were recorded in multi-channel format for diarization and automatic speech recognition (ASR) tasks. In these meeting scenarios, the uncertainty of the speaker number and the high ratio of overlapped speech present great challenges for diarization. Based on the assumption that there is valuable complementary information between acoustic features, spatial-related and speaker-related features, we propose a multi-level feature fusion mechanism based target-speaker voice activity detection (FFM-TS-VAD) system to improve the performance of the conventional TS-VAD system. Furthermore, we propose a data augmentation method during training to improve the system robustness when the angular difference between two speakers is relatively small. We provide comparisons for different sub-systems we used in M2MeT challenge. Our submission is a fusion of several sub-systems and ranks second in the diarization task. Naijun Zheng, Na Li 0012, Xixin Wu, Lingwei Meng, Jiawen Kang 0002, Chao Weng, Dan Su 0002, Helen M. Meng |
ICASSP | 5 |
| 2022 | Spoofing-Aware Speaker Verification by Multi-Level FusionabstractRecently, many novel techniques have been introduced to deal with spoofing attacks, and achieve promising countermeasure (CM) performances.However, these works only take the standalone CM models into account.Nowadays, a spoofing aware speaker verification (SASV) challenge which aims to facilitate the research of integrated CM and ASV models, arguing that jointly optimizing CM and ASV models will lead to better performance, is taking place.In this paper, we propose a novel multi-model and multi-level fusion strategy to tackle the SASV task.Compared with purely scoring fusion and embedding fusion methods, this framework first utilizes embeddings from CM models, propagating CM embeddings into a CM block to obtain a CM score.In the second-level fusion, the CM score and ASV scores directly from ASV systems will be concatenated into a prediction block for the final decision.As a result, the best single fusion system has achieved the SASV-EER of 0.97% on the evaluation set.Then by ensembling the top-5 fusion systems, the final SASV-EER reached 0.89%. Lingwei Meng, Jiawen Kang 0002, Jinchao Li, Xu Li 0015, Xixin Wu, Hung-yi Lee, Helen M. Meng |
INTERSPEECH | 3 |
| 2022 | CN-Celeb: Multi-genre speaker recognition
Lantian Li, Jiawen Kang 0002, Yunqi Cai, Ravichander Vipperla, Thomas Fang Zheng, Dong Wang 0013 |
Speech Commun. | 3 |
| 2022 | A Principle Solution for Enroll-Test Mismatch in Speaker RecognitionabstractMismatch between enrollment and test conditions causes serious performance degradation on speaker recognition systems. This paper presents a statistics decomposition (SD) approach to solve this problem. This approach decomposes the PLDA score into three components that corresponding to enrollment, prediction and normalization respectively. Given that correct statistics are used in each component, the resultant score is theoretically optimal. A comprehensive experimental study was conducted on three datasets with different types of mismatch: (1) physical channel mismatch, (2) long-term speaker characteristics mismatch, (3) near-far recording mismatch. The results demonstrated that the proposed SD approach is highly effective, and outperforms the ad-hoc multi-condition training approach that is commonly adopted but not optimal in theory. Lantian Li, Dong Wang 0013, Jiawen Kang 0002, Renyu Wang, Zhendong Gao, Xiao Chen 0012 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Squeezing Value of Cross-Domain Labels: A Decoupled Scoring Approach for Speaker VerificationabstractDomain mismatch often occurs in real applications and causes serious performance reduction on speaker verification systems. The common wisdom is to collect cross-domain data and train a multi-domain PLDA model, with the hope to learn a domain-independent speaker subspace. In this paper, we firstly present an empirical study to show that simply adding cross-domain data does not help performance in conditions with enrollment-test mismatch. Careful analysis shows that this striking result is caused by the incoherent statistics between the enrollment and test conditions. Based on this analysis, we present a decoupled scoring approach that can maximally squeeze the value of cross-domain labels and obtain optimal verification scores in the enrollment-test mismatch condition. When the statistics are coherent, the new formulation falls back to the conventional PLDA. Experimental results on cross-channel test show that the proposed approach is highly effective and is a principal solution to domain mismatch. Lantian Li, Yang Zhang 0052, Jiawen Kang 0002, Thomas Fang Zheng, Dong Wang 0013 |
ICASSP | 3 |
| 2020 | CN-Celeb: A Challenging Chinese Speaker Recognition DatasetabstractRecently, researchers set an ambitious goal of conducting speaker recognition in unconstrained conditions where the variations on ambient, channel and emotion could be arbitrary. However, most publicly available datasets are collected under constrained environments, i.e., with little noise and limited channel variation. These datasets tend to deliver over-optimistic performance and do not meet the request of research on speaker recognition in unconstrained conditions.In this paper, we present CN-Celeb, a large-scale speaker recognition dataset collected ‘in the wild’. This dataset contains more than 130,000 utterances from 1,000 Chinese celebrities, and covers 11 different genres in real world. Experiments conducted with two state-of-the-art speaker recognition approaches (i-vector and x-vector) show that the performance on CN-Celeb is far inferior to the one obtained on Vox-Celeb, a widely used speaker recognition dataset. This result demonstrates that in real-life conditions, the performance of existing techniques might be much worse than it was thought. Our database is free for researchers and can be downloaded from http://project.cslt.org. Jiawen Kang 0002, Lantian Li, Kaicheng Li, Sitong Cheng, Pengyuan Zhang, Ziya Zhou, Yunqi Cai, Dong Wang 0013 |
ICASSP | 2 |
| 2020 | Domain-Invariant Speaker Vector Projection by Model-Agnostic Meta-LearningabstractDomain generalization remains a critical problem for speaker recognition, even with the state-of-the-art architectures based on deep neural nets. For example, a model trained on reading speech may largely fail when applied to scenarios of singing or movie. In this paper, we propose a domain-invariant projection to improve the generalizability of speaker vectors. This projection is a simple neural net and is trained following the Model-Agnostic Meta-Learning (MAML) principle, for which the objective is to classify speakers in one domain if it had been updated with speech data in another domain. We tested the proposed method on CNCeleb, a new dataset consisting of single-speaker multi-condition (SSMC) data. The results demonstrated that the MAML-based domain-invariant projection can produce more generalizable speaker vectors, and effectively improve the performance in unseen domains. Jiawen Kang 0002, Lantian Li, Yunqi Cai, Dong Wang 0013, Thomas Fang Zheng |
INTERSPEECH | 1 |