VLDB 2026 Research / reviewers in the wild / expert
Shifu Xiong
dblp:148/9778
· DBLP profile ↗
8ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0003-4759-147XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MISP-QEKS: A Large-Scale Dataset with Multimodal Cues for Query-by-Example Keyword Spotting
Shifu Xiong, Hang Chen 0001, Shi Cheng 0001, Hengshun Zhou, Genshun Wan, Chenyue Zhang, Jun Du 0002, Li-Rong Dai 0001 |
ACM Multimedia | 1 |
| 2025 | Lightweight Audio-Visual Wake Word Spotting With Diverse Acoustic Knowledge DistillationabstractAudio-Visual Wake Word Spotting (AVWWS) aims to accurately detect user-defined keywords by leveraging the complementary nature of different modalities in challenging acoustic environments. However, two primary challenges hinder the application of AVWWS models in real-world scenarios: increased model parameters involving the video modality and the scarcity of paired audio-visual data. To address these issues, we propose a novel diverse acoustic knowledge distillation (DAKD) framework, which utilizes easily accessible single-modality audio data to train two teacher models and employs cross-modal knowledge distillation to transfer the generalization and de-noising capabilities of the teachers to the audio-visual student model. This approach mitigates the overfitting risk associated with large parameter counts and limited data. The DAKD framework consists of an audio-visual student model based on the lightweight multi-scale temporal-spatial attention (LMTSA) architecture, a multi-conditional teacher (MCT) model, and a de-noising teacher (DNT) model. The LMTSA model integrates compact 3D and 2D blocks based on the ResNet architecture through a simple attention module and accepts multi-scale supervision from word-level and phone-level labels, achieving joint temporal-spatial modeling with minimal parameter usage. The MCT and DNT models were trained using extensive real or simulated far-field speech and paired near-field and far-field speech, respectively, to generalize unseen acoustic environments and de-noising capabilities to the audio-visual student model. The effectiveness of our proposed DAKD framework is validated through comprehensive experiments on the MISP2021 and the updated MISP2021 Eval Hard datasets, establishing new benchmarks with fewer parameters. Our code will be available athttps://github.com/wikkk-tp/AVWWS_DAKD. Hang Chen 0001, Jun Du 0002, Hengshun Zhou, Sabato Marco Siniscalchi, Shutong Niu, Shifu Xiong |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | Collaborative Viseme Subword and End-to-End Modeling for Word-Level Lip ReadingabstractWe propose a viseme subword modeling (VSM) approach to improve the generalizability and interpretability capabilities of deep neural network based lip reading. A comprehensive analysis of preliminary experimental results reveals the complementary nature of the conventional end-to-end (E2E) and proposed VSM frameworks, especially concerning speaker head movements. To increase lip reading accuracy, we propose hybrid viseme subwords and end-to-end modeling (HVSEM), which exploits the strengths of both approaches through multitask learning. As an extension to HVSEM, we also propose collaborative viseme subword and end-to-end modeling (CVSEM), which further explores the synergy between the VSM and E2E frameworks by integrating a state-mapped temporal mask (SMTM) into joint modeling. Experimental evaluations using different model backbones on both the LRW and LRW-1000 datasets confirm the superior performance and generalizability of the proposed frameworks. Specifically, VSM outperforms the baseline E2E framework, while HVSEM outperforms VSM in a hybrid combination of VSM and E2E modeling. Building on HVSEM, CVSEM further achieves impressive accuracies on 90.75% and 58.89%, setting new benchmarks for both datasets. Hang Chen 0001, Qing Wang 0008, Jun Du 0002, Genshun Wan, Shifu Xiong, Chin-Hui Lee 0001 |
IEEE Trans. Multim. | 5 |
| 2022 | A Study of Designing Compact Audio-Visual Wake Word Spotting System Based on Iterative Fine-Tuning in Neural Network PruningabstractAudio-only based wake word spotting (WWS) is challenging under noisy conditions due to the environmental interference in signal transmission. In this paper, we investigate on designing a compact audio-visual WWS system by utilizing the visual information to alleviate the degradation. Specifically, in order to use visual information, we first encode the detected lips to fixed-size vectors with MobileNet and concatenate them with acoustic features followed by the fusion network for WWS. However, the audio-visual model based on neural network requires a large footprint and a high computational complexity. To meet the application requirements, we introduce a neural network pruning strategy via the lottery ticket hypothesis in an iterative fine-tuning manner (LTH-IF), to the single-modal and multi-modal models, respectively. Tested on our in-house corpus for audio-visual WWS in a home TV scene, the proposed audiovisual system achieves significant performance improvements over the single-modality (audio-only or video-only) system under different noisy conditions. Moreover, LTH-IF pruning can largely reduce the network parameters and computations with no degradation of WWS performance, leading to a potential product solution for the TV wake-up scenario. Hengshun Zhou, Jun Du 0002, Chao-Han Huck Yang, Shifu Xiong, Chin-Hui Lee 0001 |
ICASSP | 4 |
| 2022 | Audio-Visual Wake Word Spotting in MISP2021 Challenge: Dataset Release and Deep AnalysisabstractIn this paper, we describe and release publicly the audio-visual wake word spotting (WWS) database in the MISP2021 Challenge, which covers a range of scenarios of audio and video data collected by near-, mid-, and far-field microphone arrays, and cameras, to create a shared and publicly available database for WWS. The database and the code 2 are released, which will be a valuable addition to the community for promoting WWS research using multi-modality information in realistic and complex conditions. Moreover, we investigated the different data augmentation methods for single modalities on an end-to-end WWS network. A set of audio-visual fusion experiments and analysis were conducted to observe the assistance from visual information to acoustic information based on different audio and video field configurations. The results showed that the fusion system generally improves over the single-modality (audio- or video-only) system, especially under complex noisy conditions. Hengshun Zhou, Jun Du 0002, Gongzhen Zou, Zhaoxu Nian, Chin-Hui Lee 0001, Sabato Marco Siniscalchi, Shinji Watanabe 0001, Odette Scharenborg, Jingdong Chen, Shifu Xiong, Jianqing Gao |
INTERSPEECH | 10 |
| 2021 | Audio-Visual Information Fusion Using Cross-Modal Teacher-Student Learning for Voice Activity Detection in Realistic Environments
Hengshun Zhou, Jun Du 0002, Hang Chen 0001, Zijun Jing, Shifu Xiong, Chin-Hui Lee 0001 |
Interspeech | 5 |
| 2016 | Compact Feedforward Sequential Memory Networks for Large Vocabulary Continuous Speech Recognition
Shiliang Zhang, Hui Jiang 0001, Shifu Xiong, Si Wei, Li-Rong Dai 0001 |
INTERSPEECH | 3 |
| 2014 | Lattice based optimization of bottleneck feature extractor with linear transformationabstractThis paper proposes a lattice-based sequential discriminative training method to extract more discriminative bottleneck features. In our method, the bottleneck neural network is first trained with cross entropy criteria, and then only the weights of bottleneck layer are retrained with sequential criteria. If the outputs of the layer before bottleneck are treated as the raw features, the new method is an equivalent to a linear feature transformation algorithm. This linearity makes the optimization much easier than updating the whole neural network. Just like the fMPE and RDLT, the neural network is retrained with batch mode gradient descent, making the training to be easily implemented in parallel. Meanwhile, batch mode optimization can naturally deal with the indirect gradient to make the optimization more precise. Experimental results on a Mandarin transcription task and the Switchboard task have shown the effectiveness of the proposed method with the CER decreases from 12.2% to 11.3% and the WER from 16.1% to 15.0%, respectively. Diyuan Liu, Si Wei, Wu Guo, Yebo Bao, Shifu Xiong, Li-Rong Dai 0001 |
ICASSP | 5 |