Genshun Wan

dblp:237/0015 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0002-5813-9430ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 MISP-QEKS: A Large-Scale Dataset with Multimodal Cues for Query-by-Example Keyword Spotting
Shifu Xiong, Hang Chen 0001, Shi Cheng 0001, Hengshun Zhou, Genshun Wan, Chenyue Zhang, Jun Du 0002, Li-Rong Dai 0001
ACM Multimedia6
2025 Audio-visual representation learning via knowledge distillation from speech foundation models
Jing-Xuan Zhang, Genshun Wan, Jianqing Gao, Zhen-Hua Ling
Pattern Recognit.2
2024 Implicit Enhancement of Target Speaker in Speaker-Adaptive ASR through Efficient Joint Optimization
abstract
In multi-speaker scenarios, automatic speech recognition (ASR) models rely on pre-processed audio after speaker separation. However, when the target speaker is not accurately separated, ASR models face limitations in reaching their peak performance. To address this issue, we propose a speaker-adaptive ASR framework that possesses more implicit target speaker enhancement capability by efficiently joint-optimized speaker recognition (SR) and ASR models. Our framework introduces sharing self-supervised learning representation, optimization transfer and hierarchy speaker-gated attention. In this manner, it can maximize effectiveness of embedding bias and emphasize target speaker corresponding to semantic units. In the CHiME-7 DASR sub-track, the proposed method achieves a 28.19% relative reduction in word error rate (WER) on the development sets when compared to the official baseline. Notably, this framework has also been employed in the champion system for the CHiME-7 DASR.
Haitao Tang 0001, Jiahuan Fan, Ruoyu Wang 0029, Hang Chen 0001, Yanyong Zhang, Jun Du 0002, Hengshun Zhou, Lei Sun 0010, Tian Gao 0005, Genshun Wan, Jianqing Gao
ICASSP12
2024 Lightweight Transducer Based on Frame-Level Criterion
abstract
The transducer model trained based on sequence-level criterion requires a lot of memory due to the generation of the large probability matrix. We proposed a lightweight transducer model based on frame-level criterion, which uses the results of the CTC forced alignment algorithm to determine the label for each frame. Then the encoder output can be combined with the decoder output at the corresponding time, rather than adding each element output by the encoder to each element output by the decoder as in the transducer. This significantly reduces memory and computation requirements. To address the problem of imbalanced classification caused by excessive blanks in the label, we decouple the blank and non-blank probabilities and truncate the gradient of the blank classifier to the main network. Experiments on the AISHELL-1 demonstrate that this enables the lightweight transducer to achieve similar results to transducer. Additionally, we use richer information to predict the probability of blank, achieving superior results to transducer.
Genshun Wan, Mengzhi Wang, Tingzhi Mao, Zhongfu Ye
INTERSPEECH1
2024 Collaborative Viseme Subword and End-to-End Modeling for Word-Level Lip Reading
abstract
We propose a viseme subword modeling (VSM) approach to improve the generalizability and interpretability capabilities of deep neural network based lip reading. A comprehensive analysis of preliminary experimental results reveals the complementary nature of the conventional end-to-end (E2E) and proposed VSM frameworks, especially concerning speaker head movements. To increase lip reading accuracy, we propose hybrid viseme subwords and end-to-end modeling (HVSEM), which exploits the strengths of both approaches through multitask learning. As an extension to HVSEM, we also propose collaborative viseme subword and end-to-end modeling (CVSEM), which further explores the synergy between the VSM and E2E frameworks by integrating a state-mapped temporal mask (SMTM) into joint modeling. Experimental evaluations using different model backbones on both the LRW and LRW-1000 datasets confirm the superior performance and generalizability of the proposed frameworks. Specifically, VSM outperforms the baseline E2E framework, while HVSEM outperforms VSM in a hybrid combination of VSM and E2E modeling. Building on HVSEM, CVSEM further achieves impressive accuracies on 90.75% and 58.89%, setting new benchmarks for both datasets.
Hang Chen 0001, Qing Wang 0008, Jun Du 0002, Genshun Wan, Shifu Xiong, Chin-Hui Lee 0001
IEEE Trans. Multim.4
2023 Reducing the GAP Between Streaming and Non-Streaming Transducer-Based ASR by Adaptive Two-Stage Knowledge Distillation
abstract
Transducer is one of the mainstream frameworks for streaming speech recognition. There is a performance gap between the streaming and non-streaming transducer models due to limited context. To reduce this gap, an effective way is to ensure that their hidden and output distributions are consistent, which can be achieved by hierarchical knowledge distillation. However, it is difficult to ensure the distribution consistency simultaneously because the learning of the output distribution depends on the hidden one. In this paper, we propose an adaptive two-stage knowledge distillation method consisting of hidden layer learning and output layer learning. In the former stage, we learn hidden representation with full context by applying mean square error loss function. In the latter stage, we design a power transformation based adaptive smoothness method to learn stable output distribution. It achieved 19% relative reduction in word error rate, and a faster response for the first token compared with the original streaming model in LibriSpeech corpus.
Haitao Tang 0001, Yu Fu 0008, Lei Sun 0010, Jiabin Xue, Genshun Wan, Ming'en Zhao
ICASSP10
2023 Self-Supervised Audio-Visual Speech Representations Learning by Multimodal Self-Distillation
abstract
In this work, we present a novel method, named AV2vec, for learning audio-visual speech representations by multimodal self-distillation. AV2vec has a student and a teacher module, in which the student performs a masked latent feature regression task using the multimodal target features generated online by the teacher. The parameters of the teacher model are a momentum update of the student. Since our target features are generated online, AV2vec needs no iteration step like AV-HuBERT and the total training time cost is reduced to less than one-fifth. We further propose AV2vec-MLM in this study, which augments AV2vec with a masked language model (MLM)-style loss using multitask learning. Our experimental results show that AV2vec achieved comparable performance to the AV-HuBERT baseline. When combined with an MLM-style loss, AV2vec-MLM outperformed baselines and achieved the best performance on the downstream tasks.
Jing-Xuan Zhang, Genshun Wan, Zhen-Hua Ling, Jianqing Gao, Cong Liu 0006
ICASSP2
2022 Is Lip Region-of-Interest Sufficient for Lipreading?
abstract
Lip region-of-interest (ROI) is conventionally used for visual input in the lipreading task. Few works have adopted the entire face as visual input because lip-excluded parts of the face are usually considered to be redundant and irrelevant to visual speech recognition. However, faces contain much more detailed information than lips, such as speakers’ head pose, emotion, identity etc. We argue that such information might benefit visual speech recognition if a powerful feature extractor employing the entire face is trained. In this work, we propose to adopt the entire face for lipreading with self-supervised learning. AV-HuBERT, an audio-visual multi-modal self-supervised learning framework, was adopted in our experiments. Our experimental results showed that adopting the entire face achieved 16% relative word error rate (WER) reduction on the lipreading task, compared with the baseline method using lip as visual input. Without self-supervised pretraining, the model with face input achieved a higher WER than that using lip input in the case of limited training data (30 hours), while a slightly lower WER when using large amount of training data (433 hours).
Jing-Xuan Zhang, Genshun Wan
ICMI2
2020 Speaker Adaptive Training for Speech Recognition Based on Attention-Over-Attention Mechanism
Genshun Wan, Qingran Wang, Jianqing Gao, Zhongfu Ye
INTERSPEECH1
2020 Speaker Code Based Speaker Adaptive Training Using Model Agnostic Meta-Learning
Huaxin Wu, Genshun Wan
INTERSPEECH2
2020 Online Speaker Adaptation Using Memory-Aware Networks for Speech Recognition
abstract
In our previous work, we introduced our attention-based speaker adaptation method, which has been proved to be an efficient online speaker adaptation method for real-time speech recognition. In this paper, we present a more complete framework of this method named memory-aware networks, which consists of the main network, the memory module, the attention module and the connection module. A gate mechanism and a multiple-connections strategy are presented to connect the memory with the main network in order to take full advantage of the memory. An auxiliary speaker classification task is provided to improve the accuracy of the attention module. The fixed-size ordinally forgetting encoding method is used together with average pooling to gather both short-term and long-term information. Furthermore, instead of only using traditional speaker embeddings such as i-vectors or d-vectors as the memory, we design a new form of memory called residual vectors, which can represent different pronunciation habits. Experiments on both the Switchboard and AISHELL-2 tasks show that our method can perform online speaker adaptation very well with no additional adaptation data and with only a relative 3% increase in decoding computation complexity. Under the cross-entropy criterion, our method achieves a relative word error rate reduction of 9.4% and 8.3% compared to that of the speaker-independent model on the Switchboard task and the AISHELL-2 task, respectively, and approximately 7.0% compared to that of the traditional d-vector-based speaker adaptation method.
Genshun Wan, Jun Du 0002, Zhongfu Ye
IEEE ACM Trans. Audio Speech Lang. Process.2