VLDB 2026 Research / reviewers in the wild / expert
Hang Chen 0001
dblp:87/3677-1
· DBLP profile ↗
28ranked-venue papers
8as first author
28since 2021 · last 2025
0000-0002-0904-8946ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 6 first-author · 26 since 2021Artificial intelligence and machine learning · 11 · 4 first-author · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MEAN-RIR: Multi-Modal Environment-Aware Network for Robust Room Impulse Response EstimationabstractThis paper presents a Multi-Modal EnvironmentAware Network (MEAN-RIR), which uses an encoder-decoder framework to predict room impulse response (RIR) based on multi-level environmental information from audio, visual, and textual sources. Specifically, reverberant speech capturing room acoustic properties serves as the primary input, which is combined with panoramic images and text descriptions as supplementary inputs. Each input is processed by its respective encoder, and the outputs are fed into cross-attention modules to enable effective interaction between different modalities. The MEAN-RIR decoder generates two distinct components: the first component captures the direct sound and early reflections, while the second produces masks that modulate learnable filtered noise to synthesize the late reverberation. These two components are mixed to reconstruct the final RIR. The results show that MEANRIR significantly improves RIR estimation, with notable gains in acoustic parameters. Jiajian Chen, Jiakang Chen, Hang Chen 0001, Qing Wang 0008, Jun Du 0002 |
ASRU | 3 |
| 2025 | Projection Valued-based Quantum Machine Learning Adapting to Differential Privacy Algorithm for Word-level LipreadingabstractDeep neural network (DNN)-based lipreading models have achieved excellent recognition accuracy but are currently facing challenges related to user privacy. To address this, we propose a novel hybrid quantum-classical neural network (HQCNN) for lipreading that balances superior performance with enhanced privacy protection. The HQCNN-based lipreading model features an innovative variational quantum circuit (VQC) back-end, which transforms the output of the DNN front-end into quantum representations and predicts the posterior probability of each word. Furthermore, we introduce projection-valued encoding (PVE) and projection-valued measurement (PVM), enabling the VQC to handle inputs and outputs of dimensions that scale exponentially with the number of qubits, thereby substantially increasing its expressive power. Additionally, we explore the privacy-preserving properties of the HQCNN-based lipreading model by integrating differentially private stochastic gradient descent (DP-SGD). Experiments conducted on the LRW dataset demonstrate the model’s exceptional recognition accuracy and privacy-preserving capabilities. Hang Chen 0001, Jun Du 0002, Chao-Han Huck Yang, Jun Qi 0002 |
ICASSP | 1 |
| 2025 | The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition
Shilong Wu, Hang Chen 0001, Jun Du 0002, Chin-Hui Lee 0001, Shinji Watanabe 0001, Jingdong Chen, Sabato Marco Siniscalchi, Odette Scharenborg |
INTERSPEECH | 3 |
| 2025 | MISP-QEKS: A Large-Scale Dataset with Multimodal Cues for Query-by-Example Keyword Spotting
Shifu Xiong, Hang Chen 0001, Shi Cheng 0001, Hengshun Zhou, Genshun Wan, Chenyue Zhang, Jun Du 0002, Li-Rong Dai 0001 |
ACM Multimedia | 2 |
| 2025 | Dual-Branch Codec With Orthogonality Constraint and Knowledge Distillation for Noisy EnvironmentabstractAudio codecs, by discretizing continuous audio signals into finite token sets, achieve high-quality reconstruction at low bitrates in clean environments. However, real-world speech often deviates from ideal conditions, particularly in noisy environments with low signal-to-noise ratios (SNRs), limiting the performance of existing codecs in restoring clean audio from noisy inputs. To address this challenge, this paper introduces a Dual-Branch Codec (DB-Codec). Leveraging the hierarchical decomposition capability of residual vector quantization (RVQ), we separate noise and speech into codebooks at different layers through dual-branch reconstruction and orthogonality constraints between noise and speech features. DB-Codec integrates enhancement and synthesis into a unified model, enabling flexible control over noise suppression or signal recovery at equivalent compression rates to conventional codecs. Experiments demonstrate that our DB-Codec achieve an average improvement of 0.83 in PESQ and 9.16 in STOI compared to traditional codecs under low SNR conditions. Hang Chen 0001, Jun Du 0002 |
IEEE Signal Process. Lett. | 2 |
| 2025 | Lightweight Audio-Visual Wake Word Spotting With Diverse Acoustic Knowledge DistillationabstractAudio-Visual Wake Word Spotting (AVWWS) aims to accurately detect user-defined keywords by leveraging the complementary nature of different modalities in challenging acoustic environments. However, two primary challenges hinder the application of AVWWS models in real-world scenarios: increased model parameters involving the video modality and the scarcity of paired audio-visual data. To address these issues, we propose a novel diverse acoustic knowledge distillation (DAKD) framework, which utilizes easily accessible single-modality audio data to train two teacher models and employs cross-modal knowledge distillation to transfer the generalization and de-noising capabilities of the teachers to the audio-visual student model. This approach mitigates the overfitting risk associated with large parameter counts and limited data. The DAKD framework consists of an audio-visual student model based on the lightweight multi-scale temporal-spatial attention (LMTSA) architecture, a multi-conditional teacher (MCT) model, and a de-noising teacher (DNT) model. The LMTSA model integrates compact 3D and 2D blocks based on the ResNet architecture through a simple attention module and accepts multi-scale supervision from word-level and phone-level labels, achieving joint temporal-spatial modeling with minimal parameter usage. The MCT and DNT models were trained using extensive real or simulated far-field speech and paired near-field and far-field speech, respectively, to generalize unseen acoustic environments and de-noising capabilities to the audio-visual student model. The effectiveness of our proposed DAKD framework is validated through comprehensive experiments on the MISP2021 and the updated MISP2021 Eval Hard datasets, establishing new benchmarks with fewer parameters. Our code will be available athttps://github.com/wikkk-tp/AVWWS_DAKD. Hang Chen 0001, Jun Du 0002, Hengshun Zhou, Sabato Marco Siniscalchi, Shutong Niu, Shifu Xiong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Video Segmentation and Tokenization for Model-Based Video Scene ClassificationabstractIn this paper, we propose a novel approach for segmenting and tokenizing a video scene recording into a sequence of cascade units, known as visual segment units and modeled with visual segment models (VSMs) for video scene classification (VSC). Specifically, the proposed VSM framework takes deep visual features extracted from pre-trained encoders as inputs and models the temporal interactions between segment units by hidden Markov models. Next, we use unit co-occurrence statistics to introduce relationships between VSM units within a video scene recording. Furthermore, the VSM approach is extended to an acoustic-visual variant, subsequently integrating itself into a deep learning-based multi-modal scene classification system. This combination serves to further exploit the complementary nature of audio and video data. By incorporating a set of visual segment units into modeling a video scene class, it captures both inter-class similarity and intra-class diversity, facilitating improved scene classification, especially within categories prone to confusion. Extensive experimental results on a benchmark published by the DCASE (Detection and Classification of Acoustic Scenes and Events) 2021 Challenge show that the proposed framework can effectively handle the confusion issue among similar video scenes. In addition, our multi-modal integration system achieves state-of-the-art performance in the audio-visual scene classification task in the DCASE 2021 Challenge, thereby demonstrating the effectiveness of our proposed approach. Qing Wang 0008, Yajian Wang, Hang Chen 0001, Jun Du 0002, Chin-Hui Lee 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | A Study of Dropout-Induced Modality Bias on Robustness to Missing Video Frames for Audio-Visual Speech RecognitionabstractAdvanced Audio- Visual Speech Recognition (AVSR) sys-tems have been observed to be sensitive to missing video frames, performing even worse than single-modality mod-els. While applying the common dropout techniques to the video modality enhances robustness to missing frames, it simultaneously results in a performance loss when dealing with complete data input. In this study, we delve into this contrasting phenomenon through the lens of modality bias and uncover that an excessive modality bias towards the audio modality induced by dropout constitutes the fun-damental cause. Next, we present the Modality Bias Hy-pothesis (MBH) to systematically describe the relationship between the modality bias and the robustness against missing modality in multimodal systems. Building on these findings, we propose a novel Multimodal Distribution Approxi-mation with Knowledge Distillation (MDA-KD)framework to reduce over-reliance on the audio modality, maintaining performance and robustness simultaneously. Finally, to address an entirely missing modality, we adopt adapters to dynamically switch decision strategies. The effective-ness of our proposed approach is evaluated through comprehensive experiments on the MISP2021 and MISP2022 datasets. Our code is available at https://github.com/dalision/ModalBiasAV5R. Yusheng Dai, Hang Chen 0001, Jun Du 0002, Ruoyu Wang 0029, Shihao Chen, Chin-Hui Lee 0001 |
CVPR | 2 |
| 2024 | Implicit Enhancement of Target Speaker in Speaker-Adaptive ASR through Efficient Joint OptimizationabstractIn multi-speaker scenarios, automatic speech recognition (ASR) models rely on pre-processed audio after speaker separation. However, when the target speaker is not accurately separated, ASR models face limitations in reaching their peak performance. To address this issue, we propose a speaker-adaptive ASR framework that possesses more implicit target speaker enhancement capability by efficiently joint-optimized speaker recognition (SR) and ASR models. Our framework introduces sharing self-supervised learning representation, optimization transfer and hierarchy speaker-gated attention. In this manner, it can maximize effectiveness of embedding bias and emphasize target speaker corresponding to semantic units. In the CHiME-7 DASR sub-track, the proposed method achieves a 28.19% relative reduction in word error rate (WER) on the development sets when compared to the official baseline. Notably, this framework has also been employed in the champion system for the CHiME-7 DASR. Haitao Tang 0001, Jiahuan Fan, Ruoyu Wang 0029, Hang Chen 0001, Yanyong Zhang, Jun Du 0002, Hengshun Zhou, Lei Sun 0010, Tian Gao 0005, Genshun Wan, Jianqing Gao |
ICASSP | 5 |
| 2024 | The Multimodal Information Based Speech Processing (MISP) 2023 Challenge: Audio-Visual Target Speaker ExtractionabstractPrevious Multimodal Information based Speech Processing (MISP) challenges mainly focused on audio-visual speech recognition (AVSR) with commendable success. However, the most advanced back-end recognition systems often hit performance limits due to the complex acoustic environments. This has prompted a shift in focus towards the Audio-Visual Target Speaker Extraction (AVTSE) task for the MISP 2023 challenge in ICASSP 2024 Signal Processing Grand Challenges. Unlike existing audio-visual speech enhancement challenges primarily focused on simulation data, the MISP 2023 challenge uniquely explores how front-end speech processing, combined with visual clues, impacts back-end tasks in real-world scenarios. This pioneering effort aims to set the first benchmark for the AVTSE task, offering fresh insights into enhancing the accuracy of back-end speech recognition systems through AVTSE in challenging and real acoustic environments. This paper delivers a thorough overview of the task setting, dataset, and baseline system of the MISP 2023 challenge. It also includes an in-depth analysis of the challenges participants may encounter. The experimental results highlight the demanding nature of this task, and we look forward to the innovative solutions participants will bring forward. Shilong Wu, Hang Chen 0001, Yusheng Dai, Chenyue Zhang, Ruoyu Wang 0029, Hongbo Lan, Jun Du 0002, Chin-Hui Lee 0001, Jingdong Chen, Sabato Marco Siniscalchi, Odette Scharenborg, Zhongqiu Wang 0001, Jianqing Gao |
ICASSP | 3 |
| 2024 | Enhancing Voice Wake-Up for Dysarthria: Mandarin Dysarthria Speech Corpus Release and Customized System Design
Hang Chen 0001, Jun Du 0002, Hongxiao Guo, Hui Bu, Jianxing Yang, Ming Li 0026, Chin-Hui Lee 0001 |
INTERSPEECH | 2 |
| 2024 | Summary of Low-Resource Dysarthria Wake-Up Word Spotting ChallengeabstractIn recent years, the rapid advancement and widespread adoption of speech technology have made smart home systems a common feature in many households. However, individuals with dysarthria face difficulties using these technologies due to inconsistent speech patterns. This paper summarizes the Low-Resource Dysarthria Wake-Up Word Spotting (LRDWWS) Challenge at SLT 2024, which aimed to develop effective voice wake-up systems for individuals with dysarthria. The challenge attracted 25 teams from 4 countries, with 7 teams submitting results and 5 providing detailed system descriptions. This paper presents an overview of the dataset, evaluation metrics, and key innovations from participating teams. Our findings highlight the potential of these systems to enhance the accessibility and usability of smart home technologies for individuals with dysarthria. The challenge results underscore the importance of developing specialized solutions to meet the unique needs of this user group. Hang Chen 0001, Jun Du 0002, Hongxiao Guo, Hui Bu, Ming Li 0026, Chin-Hui Lee 0001 |
SLT | 2 |
| 2024 | Optimizing Audio-Visual Speech Enhancement Using Multi-Level Distortion Measures for Audio-Visual Speech RecognitionabstractA multi-level distortion measure (MLDM) is proposed as an objective to optimize deep neural network-based speech enhancement (SE) in both audio-only and audio-visual scenarios. The aim is to achieve simultaneous performance improvements in speech quality, intelligibility, and recognition error reductions. Moreover, a comprehensive correlation analysis shows that these three evaluation metrics exhibit high Pearson correlation coefficient (PCC) values with three commonly used optimization objectives: the mean squared error between the ideal ratio and estimated magnitude masks, scale-invariant signal-to-noise ratio, and cross-entropy-guided measure. To further improve the performance, we leverage the complementarities of the three objectives and propose another correlated multi-level distortion measure (C-MLDM) defined as a weighted combination of MLDM and an average correlation measure based on the three PCCs. Experimental results on the TCD-TIMIT corpus corrupted by additive noise demonstrate that MLDM outperforms systems optimized with each objective in both audio-visual and audio-only scenarios, offering improved performances in all three metrics: speech quality, intelligibility, and recognition performance. C-MLDM also consistently outperforms MLDM in all test cases. Finally, the generalizability of both MLDM and C-MLDM is confirmed through extensive testing across diverse datasets, SE model architectures, and linguistic conditions. The source codes are publicly available.1 Hang Chen 0001, Qing Wang 0008, Jun Du 0002, Chin-Hui Lee 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2024 | Collaborative Viseme Subword and End-to-End Modeling for Word-Level Lip ReadingabstractWe propose a viseme subword modeling (VSM) approach to improve the generalizability and interpretability capabilities of deep neural network based lip reading. A comprehensive analysis of preliminary experimental results reveals the complementary nature of the conventional end-to-end (E2E) and proposed VSM frameworks, especially concerning speaker head movements. To increase lip reading accuracy, we propose hybrid viseme subwords and end-to-end modeling (HVSEM), which exploits the strengths of both approaches through multitask learning. As an extension to HVSEM, we also propose collaborative viseme subword and end-to-end modeling (CVSEM), which further explores the synergy between the VSM and E2E frameworks by integrating a state-mapped temporal mask (SMTM) into joint modeling. Experimental evaluations using different model backbones on both the LRW and LRW-1000 datasets confirm the superior performance and generalizability of the proposed frameworks. Specifically, VSM outperforms the baseline E2E framework, while HVSEM outperforms VSM in a hybrid combination of VSM and E2E modeling. Building on HVSEM, CVSEM further achieves impressive accuracies on 90.75% and 58.89%, setting new benchmarks for both datasets. Hang Chen 0001, Qing Wang 0008, Jun Du 0002, Genshun Wan, Shifu Xiong, Chin-Hui Lee 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Semi-Supervised Multi-Channel Speaker Diarization With Cross-Channel AttentionabstractMost neural speaker diarization systems rely on sufficient manual training data labels, which are hard to collect under real-world scenarios. This paper proposes a semi-supervised speaker diarization system to utilize large-scale multi-channel training data by generating pseudo-labels for unlabeled data. Furthermore, we introduce cross-channel attention into the Neural Speaker Diarization Using Memory-Aware Multi-Speaker Embedding (NSD-MA-MSE) to learn channel contextual information of speaker embeddings better. Experimental results on the CHiME-7 Mixer6 dataset which only contains partial speakers’ labels of the training set, show that our system achieved 57.01% relative DER reduction compared to the clustering-based model on the development set. We further conducted experiments on the CHiME- 6 dataset to simulate the scenario of missing partial training set labels. When using 80% and 50% labeled training data, our system performs comparably to the results obtained using 100% labeled data for training. Shilong Wu, Jun Du 0002, Maokui He, Shutong Niu, Hang Chen 0001, Haitao Tang 0001, Chin-Hui Lee 0001 |
ASRU | 5 |
| 2023 | Summary on the Multimodal Information Based Speech Processing (MISP) 2022 ChallengeabstractThe Multimodal Information based Speech Processing (MISP) 2022 challenge aimed to enhance speech processing performance in harsh acoustic environments by leveraging additional modalities such as video or text. The challenge included two tracks: audio-visual speaker diarization (AVSD) and audio-visual diarization and recognition (AVDR). The training material was based on previous MISP 2021 recordings, but we have accurately synchronized audio and visual data. Additionally, a new evaluation set was provided. This paper gives an overview of the challenge setup, presents the results, and summarizes the effective techniques employed by the participants. We also analyze the current technical challenges and suggest directions for future research in AVSD and AVDR. Hang Chen 0001, Shilong Wu, Yusheng Dai, Jun Du 0002, Chin-Hui Lee 0001, Jingdong Chen, Shinji Watanabe 0001, Sabato Marco Siniscalchi, Odette Scharenborg, Diyuan Liu, Jianqing Gao, Cong Liu 0006 |
ICASSP | 1 |
| 2023 | Incorporating Lip Features into Audio-Visual Multi-Speaker DOA Estimation by Gated FusionabstractThe audio-visual direction of arrival (DOA) estimation has demonstrated superior performance recently. In this paper, we present a novel audio-visual multi-speaker DOA estimation network, which for the first time incorporates multi-speaker lip features to adapt the complex overlapping and noisy scenarios. Firstly, we encode the multi-channel audio features, the reference angles and the lip Regions of Interest (RoIs) detected from the video respectively to acquire high-level representations. Then the multi-modal embeddings of audio, speaker angles and lips are fused by a tri-modal gated fusion module to balance their contributions to the output. The fused embedding is sent to the backend network to obtain the accurate DOA estimation with the combination of the predicted speaker angular vectors and the speaker activities. Experimental results show that our proposed approach can reduce the localization error by 73.48% compared to the previous work on the 2021 Multi-modal Information based Speech Processing (MISP) Challenge corpus. Meanwhile, the high accuracy and stability of localization results demonstrate the robustness of the proposed model in multi-speaker scenarios. Ya Jiang, Hang Chen 0001, Jun Du 0002, Qing Wang 0008, Chin-Hui Lee 0001 |
ICASSP | 2 |
| 2023 | The Multimodal Information Based Speech Processing (Misp) 2022 Challenge: Audio-Visual Diarization And RecognitionabstractThe Multi-modal Information based Speech Processing (MISP) challenge aims to extend the application of signal processing technology in specific scenarios by promoting the research into wake-up words, speaker diarization, speech recognition, and other technologies. The MISP2022 challenge has two tracks: 1) audio-visual speaker diarization (AVSD), aiming to solve "who spoken when" using both audio and visual data; 2) a novel audio-visual diarization and recognition (AVDR) task that focuses on addressing "who spoken what when" with audio-visual speaker diarization results. Both tracks focus on the Chinese language, and use far-field audio and video in real home-tv scenarios: 2-6 people communicating each other with TV noise in the background. This paper introduces the dataset, track settings, and baselines of the MISP2022 challenge. Our analyses of experiments and examples indicate the good performance of AVDR baseline system, and the potential difficulties in this challenge due to, e.g., the far-field video quality, the presence of TV noise in the background, and the indistinguishable speakers. Shilong Wu, Hang Chen 0001, Maokui He, Jun Du 0002, Chin-Hui Lee 0001, Jingdong Chen, Shinji Watanabe 0001, Sabato Marco Siniscalchi, Odette Scharenborg, Diyuan Liu, Jianqing Gao, Cong Liu 0006 |
ICASSP | 3 |
| 2023 | Incorporating Visual Information Reconstruction into Progressive Learning for Optimizing audio-visual Speech EnhancementabstractVideo information has been widely introduced to speech enhancement as its contribution at low signal-to-noise ratios (SNRs). Conventional audio-visual speech enhancement networks take noisy speech and video as input and learn features of clean speech directly. To reduce the large SNR gap between the learning target and input noisy speech, we propose a novel mask-based audio-visual progressive learning speech enhancement (AVPL) framework with visual information reconstruction (VIR) to increase SNRs gradually. Each stage of AVPL takes a concatenation of pre-trained visual embedding and the previous representation as input and predicts a mask with the intermediate representation of the current stage. To extract more visual information and deal with the performance distortion, the AVPL-VIR model reconstructs the visual embedding as it is fed in for each stage. Experiment on the TCD-TIMIT dataset shows that the progressive learning method significantly outperforms direct learning for both audio-only and audio-visual models. Moreover, by reconstructing video information, the VIR module provides a more accurate and comprehensive representation of the data, which in turn improves the performance of both AVDL and AVPL. Chenyue Zhang, Hang Chen 0001, Jun Du 0002, Chin-Hui Lee 0001 |
ICASSP | 2 |
| 2023 | Improving Audio-Visual Speech Recognition by Lip-Subword Correlation Based Visual Pre-training and Cross-Modal Fusion EncoderabstractIn recent research, slight performance improvement is observed from automatic speech recognition systems to audio-visual speech recognition systems in end-to-end frameworks with low-quality videos. Unmatching convergence rates and specialized input representations between audio-visual modalities are considered to cause the problem. In this paper, we propose two novel techniques to improve audio-visual speech recognition (AVSR) under a pre-training and fine-tuning training framework. First, we explore the correlation between lip shapes and syllable-level subword units in Mandarin through a frame-level subword unit classification task with visual streams as input. The fine-grained subword labels guide the network to capture temporal relationships between lip shapes and result in an accurate alignment between video and audio streams. Next, we propose an audio-guided Cross-Modal Fusion Encoder (CMFE) to utilize main training parameters for multiple cross-modal attention layers to make full use of modality complementarity. Experiments on the MISP2021-AVSR data set show the effectiveness of the two proposed techniques. Together, using only a relatively small amount of training data, the final system achieves better performances than state-of-the-art systems with more complex front-ends and back-ends. The code is released at1. Yusheng Dai, Hang Chen 0001, Jun Du 0002, Xiaofei Ding, Feijun Jiang, Chin-Hui Lee 0001 |
ICME | 2 |
| 2023 | Hierarchical Audio-Visual Information Fusion with Multi-label Joint Decoding for MER 2023abstractIn this paper, we propose a novel framework for recognizing both discrete and dimensional emotions. In our framework, deep features extracted from foundation models are used as robust acoustic and visual representations of raw video. Three different structures based on attention-guided feature gathering (AFG) are designed for deep feature fusion. Then, we introduce a joint decoding structure for emotion classification and valence regression in the decoding stage. A multi-task loss based on uncertainty is also designed to optimize the whole process. Finally, by combining three different structures on the posterior probability level, we obtain the final predictions of discrete and dimensional emotions. When tested on the dataset of multimodal emotion recognition challenge (MER 2023), the proposed framework yields consistent improvements in both emotion classification and valence regression. Our final system achieves state-of-the-art performance and ranks third on the leaderboard on MER-MULTI sub-challenge. Yuxuan Xi, Hang Chen 0001, Jun Du 0002, Yan Song 0001, Qing Wang 0008, Hengshun Zhou, Jiefeng Ma, Pengfei Hu 0006, Ya Jiang, Shi Cheng 0001, Jie Zhang 0042, Yuzhe Weng |
ACM Multimedia | 3 |
| 2023 | Space-and-speaker-aware acoustic modeling with effective data augmentation for recognition of multi-array conversational speech
Li Chai 0002, Hang Chen 0001, Jun Du 0002, Qingfeng Liu, Chin-Hui Lee 0001 |
Speech Commun. | 2 |
| 2022 | The First Multimodal Information Based Speech Processing (Misp) Challenge: Data, Tasks, Baselines And ResultsabstractIn this paper we discuss the rational of the Multi-model Information based Speech Processing (MISP) Challenge, and provide a detailed description of the data recorded, the two evaluation tasks and the corresponding baselines, followed by a summary of submitted systems and evaluation results. The MISP Challenge aims at tack-ling speech processing tasks in different scenarios by introducing information about an additional modality (e.g., video, or text), which will hopefully lead to better environmental and speaker robustness in realistic applications. In the first MISP challenge, two bench-mark datasets recorded in a real-home TV room with two reproducible open-source baseline systems have been released to promote research in audio-visual wake word spotting (AVWWS) and audio-visual speech recognition (AVSR). To our knowledge, MISP is the first open evaluation challenge to tackle real-world issues of AVWWS and AVSR in the home TV scenario. Hang Chen 0001, Hengshun Zhou, Jun Du 0002, Chin-Hui Lee 0001, Jingdong Chen, Shinji Watanabe 0001, Sabato Marco Siniscalchi, Odette Scharenborg, Diyuan Liu, Jianqing Gao, Cong Liu 0006 |
ICASSP | 1 |
| 2022 | Audio-Visual Speech Recognition in MISP2021 Challenge: Dataset Release and Deep AnalysisabstractIn this paper, we present the updated Audio-Visual Speech Recognition (AVSR) corpus of MISP2021 challenge, a large-scale audio-visual Chinese conversational corpus consisting of 141h audio and video data collected by far/middle/near microphones and far/middle cameras in 34 real-home TV rooms. To our best knowledge, our corpus is the first distant multi-microphone conversational Chinese audio-visual corpus and the first large vocabulary continuous Chinese lip-reading dataset in the adverse home-tv scenario. Moreover, we make a deep analysis of the corpus and conduct a comprehensive ablation study of all audio and video data in the audio-only/video-only/audiovisual systems. Error analysis shows video modality supplement acoustic information degraded by noise to reduce deletion errors and provide discriminative information in overlapping speech to reduce substitution errors. Finally, we also design a set of experiments such as frontend, data augmentation and end-to-end models for providing the direction of potential future work. The corpus and the code are released to promote the research not only in speech area but also for the computer vision area and cross-disciplinary research. Hang Chen 0001, Jun Du 0002, Yusheng Dai, Chin-Hui Lee 0001, Sabato Marco Siniscalchi, Shinji Watanabe 0001, Odette Scharenborg, Jingdong Chen |
INTERSPEECH | 1 |
| 2022 | Deep Segment Model for Acoustic Scene Classification
Yajian Wang, Jun Du 0002, Hang Chen 0001, Qing Wang 0008, Chin-Hui Lee 0001 |
INTERSPEECH | 3 |
| 2021 | Automatic Lip-Reading with Hierarchical Pyramidal Convolution and Self-Attention for Image Sequences with No Word Boundaries
Hang Chen 0001, Jun Du 0002, Yu Hu 0003, Li-Rong Dai 0001, Chin-Hui Lee 0001 |
Interspeech | 1 |
| 2021 | Audio-Visual Information Fusion Using Cross-Modal Teacher-Student Learning for Voice Activity Detection in Realistic Environments
Hengshun Zhou, Jun Du 0002, Hang Chen 0001, Zijun Jing, Shifu Xiong, Chin-Hui Lee 0001 |
Interspeech | 3 |
| 2021 | Correlating subword articulation with lip shapes for embedding aware audio-visual speech enhancement
Hang Chen 0001, Jun Du 0002, Yu Hu 0003, Li-Rong Dai 0001, Chin-Hui Lee 0001 |
Neural Networks | 1 |