Qing Wang 0008

dblp:97/6505-8 · DBLP profile ↗
← Back
33ranked-venue papers
7as first author
25since 2021 · last 2026
0000-0003-3843-3920ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 4 first-author · 15 since 2021Artificial intelligence and machine learning · 20 · 4 first-author · 13 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 See then tell: Enhancing key information extraction with vision grounding
Shuhang Liu, Pengfei Hu 0006, Jiefeng Ma, Jun Du 0002, Qing Wang 0008, Jianshu Zhang 0001
Neurocomputing6
2026 Reinforcement learning-powered co-optimization: Bridging critic model and multimodal LLM reasoning abilities
Qing Wang 0008, Shuhang Liu, Jun Du 0002, Jianshu Zhang 0001
Pattern Recognit.1
2025 MEAN-RIR: Multi-Modal Environment-Aware Network for Robust Room Impulse Response Estimation
abstract
This paper presents a Multi-Modal EnvironmentAware Network (MEAN-RIR), which uses an encoder-decoder framework to predict room impulse response (RIR) based on multi-level environmental information from audio, visual, and textual sources. Specifically, reverberant speech capturing room acoustic properties serves as the primary input, which is combined with panoramic images and text descriptions as supplementary inputs. Each input is processed by its respective encoder, and the outputs are fed into cross-attention modules to enable effective interaction between different modalities. The MEAN-RIR decoder generates two distinct components: the first component captures the direct sound and early reflections, while the second produces masks that modulate learnable filtered noise to synthesize the late reverberation. These two components are mixed to reconstruct the final RIR. The results show that MEANRIR significantly improves RIR estimation, with notable gains in acoustic parameters.
Jiajian Chen, Jiakang Chen, Hang Chen 0001, Qing Wang 0008, Jun Du 0002
ASRU4
2025 An Enhanced Audio Feature Tailored for Anomalous Sound Detection Based on Pre-trained Models
Guirui Zhong, Qing Wang 0008, Jun Du 0002, Mingqi Cai
ICANN (3)2
2025 Incorporating Audio-Guided Visual Attention into Sound Event Localization and Detection with Source Distance Estimation
abstract
Sound event localization and detection (SELD) is a task that involves identifying and locating sound events in a given environment, which combines sound event detection (SED) and direction-of-arrival (DOA) estimation. This study addresses the extended task of audio-visual (AV) sound event localization and detection with source distance estimation (3D SELD). To leverage effective visual information, we propose an audio-guided visual attention mechanism to extract location-based features. We use two methods to fuse audio and visual features. Additionally, we introduce a source coordinate estimation (SCE) task that integrates DOA and distance estimation. Experimental results demonstrate that our proposed model significantly outperforms the official audio-only and AV baselines of the DCASE 2024 Challenge Task 3 on the development set of the STARSS23 dataset, even surpassing the challenge’s winning method. Attention visualization further highlights the effectiveness of audio information in localizing sound sources within visual images, ultimately enhancing the 3D SELD performance. Codes are available at https://github.com/qingwang24/AGVA-3DSELD/.
Qing Wang 0008, Jun Du 0002, Hengyi Hong, Maocheng Hu, Mingqi Cai
ICME1
2025 Video Segmentation and Tokenization for Model-Based Video Scene Classification
abstract
In this paper, we propose a novel approach for segmenting and tokenizing a video scene recording into a sequence of cascade units, known as visual segment units and modeled with visual segment models (VSMs) for video scene classification (VSC). Specifically, the proposed VSM framework takes deep visual features extracted from pre-trained encoders as inputs and models the temporal interactions between segment units by hidden Markov models. Next, we use unit co-occurrence statistics to introduce relationships between VSM units within a video scene recording. Furthermore, the VSM approach is extended to an acoustic-visual variant, subsequently integrating itself into a deep learning-based multi-modal scene classification system. This combination serves to further exploit the complementary nature of audio and video data. By incorporating a set of visual segment units into modeling a video scene class, it captures both inter-class similarity and intra-class diversity, facilitating improved scene classification, especially within categories prone to confusion. Extensive experimental results on a benchmark published by the DCASE (Detection and Classification of Acoustic Scenes and Events) 2021 Challenge show that the proposed framework can effectively handle the confusion issue among similar video scenes. In addition, our multi-modal integration system achieves state-of-the-art performance in the audio-visual scene classification task in the DCASE 2021 Challenge, thereby demonstrating the effectiveness of our proposed approach.
Qing Wang 0008, Yajian Wang, Hang Chen 0001, Jun Du 0002, Chin-Hui Lee 0001
IEEE Trans. Multim.1
2024 Maths: Multimodal Transformer-Based Human-Readable Solver
abstract
Multimodal mathematical reasoning has gained increasing attention in recent times. However, previous effective methods have not tried to reason in the form of natural language. In this paper, we introduce a model named MATHS (MultimodAl Transformer-based Human-readable Solver) for visual arithmetic and geometry problems in multimodal mathematical reasoning tasks. Drawing inspiration from Multimodal Large Language Models (MLLMs), our approach involves generating problem-solving processes expressed in natural language, in order to leverage the inherent reasoning capabilities embedded within language models. To address the challenge of precise calculations for language models, our work proposes a Math-Constrained Generation (MCG) method to impose hard constraints on generated outputs. Extensive experiments demonstrate our model excels in visual arithmetic task, and achieves results that are either better or comparable to existing methods in geometry problems. Code is available at https://github.com/ycpNotFound/MATHS.
Yicheng Pan 0004, Jiefeng Ma, Pengfei Hu 0006, Jun Du 0002, Qing Wang 0008, Jianshu Zhang 0001, Dan Liu 0008, Si Wei
ICME6
2024 Exploring Audio-Visual Information Fusion for Sound Event Localization and Detection In Low-Resource Realistic Scenarios
abstract
This study presents an audio-visual information fusion approach to sound event localization and detection (SELD) in low-resource scenarios. We aim at utilizing audio and video modality information through cross-modal learning and multi-modal fusion. First, we propose a cross-modal teacher-student learning (TSL) framework to transfer information from an audio-only teacher model, trained on a rich collection of audio data with multiple data augmentation techniques, to an audiovisual student model trained with only a limited set of multimodal data. Next, we propose a two-stage audio-visual fusion strategy, consisting of an early feature fusion and a late video-guided decision fusion to exploit synergies between audio and video modalities. Finally, we introduce an innovative video pixel swapping (VPS) technique to extend an audio channel swapping (ACS) method to an audio-visual joint augmentation. Evaluation results on the Detection and Classification of Acoustic Scenes and Events (DCASE) 2023 Challenge data set demonstrate significant improvements in SELD performances. Furthermore, our submission to the SELD task of the DCASE 2023 Challenge ranks first place by effectively integrating the proposed techniques into a model ensemble.
Ya Jiang, Qing Wang 0008, Jun Du 0002, Maocheng Hu, Pengfei Hu 0006, Zeyan Liu, Shi Cheng 0001, Zhaoxu Nian, Mingqi Cai, Chin-Hui Lee 0001
ICME2
2024 Representation Learning Using Machine Attribute Information for Anomalous Sound Detection in Real Scenarios
abstract
In the previous Detection and Classification of Acoustic Scenes and Events (DCASE) Challenge Task 2: Anomalous Sound Detection (ASD) for Machine Condition Monitoring, each machine has a variety of different section IDs, which are subsets of the machine type. Therefore, section ID classification is often used to learn the representation of machine sounds for ASD. However, in real scenarios, it is both time-consuming and laborious for each machine to record data with multiple different section IDs. As such, the Task 2 of DCASE 2023 Challenge only includes one section ID for each machine, with the attribute information reflecting the machine’s working status and environment for recording. To this end, machine sound representations for ASD can be learned through the proxy task of two-stage multi-attribute classification. Specifically, the sounds of all machines are first used to pre-train a general attribute classification model. This model is then fine-tuned to obtain an attribute classification model specific to each machine, with a classification head established for each attribute that affects the acoustic characteristics of the machine in a multi-task learning framework. At the same time, data augmentation is used to improve the generalization capability caused by the limited amount of data in actual scenarios. Our approach demonstrates commendable performance on the Task 2 of DCASE 2023 Challenge. We further illustrate the effectiveness of our method through visual analysis.
Qing Wang 0008, Jun Du 0002, Fan Chu, Mingqi Cai
IJCNN2
2024 SRFUND: A Multi-Granularity Hierarchical Structure Reconstruction Benchmark in Form Understanding
abstract
Accurately identifying and organizing textual content is crucial for the automation of document processing in the field of form understanding. Existing datasets, such as FUNSD and XFUND, support entity classification and relationship prediction tasks but are typically limited to local and entity-level annotations. This limitation overlooks the hierarchically structured representation of documents, constraining comprehensive understanding of complex forms. To address this issue, we present the SRFUND, a hierarchically structured multi-task form understanding benchmark. SRFUND provides refined annotations on top of the original FUNSD and XFUND datasets, encompassing five tasks: (1) word to text-line merging, (2) text-line to entity merging, (3) entity category classification, (4) item table localization, and (5) entity-based full-document hierarchical structure recovery. We meticulously supplemented the original dataset with missing annotations at various levels of granularity and added detailed annotations for multi-item table regions within the forms. Additionally, we introduce global hierarchical structure dependencies for entity relation prediction tasks, surpassing traditional local key-value associations. The SRFUND dataset includes eight languages including English, Chinese, Japanese, German, French, Spanish, Italian, and Portuguese, making it a powerful tool for cross-lingual form understanding. Extensive experimental results demonstrate that the SRFUND dataset presents new challenges and significant opportunities in handling diverse layouts and global hierarchical structures of forms, thus providing deep insights into the field of form understanding. The original dataset and implementations of baseline methods are available at https://sprateam-ustc.github.io/SRFUND.
Jiefeng Ma, Jun Du 0002, Yu Hu 0003, Pengfei Hu 0006, Qing Wang 0008, Jianshu Zhang 0001
NeurIPS8
2024 Optimizing Audio-Visual Speech Enhancement Using Multi-Level Distortion Measures for Audio-Visual Speech Recognition
abstract
A multi-level distortion measure (MLDM) is proposed as an objective to optimize deep neural network-based speech enhancement (SE) in both audio-only and audio-visual scenarios. The aim is to achieve simultaneous performance improvements in speech quality, intelligibility, and recognition error reductions. Moreover, a comprehensive correlation analysis shows that these three evaluation metrics exhibit high Pearson correlation coefficient (PCC) values with three commonly used optimization objectives: the mean squared error between the ideal ratio and estimated magnitude masks, scale-invariant signal-to-noise ratio, and cross-entropy-guided measure. To further improve the performance, we leverage the complementarities of the three objectives and propose another correlated multi-level distortion measure (C-MLDM) defined as a weighted combination of MLDM and an average correlation measure based on the three PCCs. Experimental results on the TCD-TIMIT corpus corrupted by additive noise demonstrate that MLDM outperforms systems optimized with each objective in both audio-visual and audio-only scenarios, offering improved performances in all three metrics: speech quality, intelligibility, and recognition performance. C-MLDM also consistently outperforms MLDM in all test cases. Finally, the generalizability of both MLDM and C-MLDM is confirmed through extensive testing across diverse datasets, SE model architectures, and linguistic conditions. The source codes are publicly available.1
Hang Chen 0001, Qing Wang 0008, Jun Du 0002, Chin-Hui Lee 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 A Variance-Preserving Interpolation Approach for Diffusion Models With Applications to Single Channel Speech Enhancement and Recognition
abstract
In this paper, we propose a variance-preserving interpolation framework to improve diffusion models for single-channel speech enhancement (SE) and automatic speech recognition (ASR). This new variance-preserving interpolation diffusion model (VPIDM) approach requires only 25 iterative steps and obviates the need for a corrector, an essential element in the existing variance-exploding interpolation diffusion model (VEIDM). Two notable distinctions between VPIDM and VEIDM are the scaling function of the mean of state variables and the constraint imposed on the variance relative to the mean's scale. We conduct a systematic exploration of the theoretical mechanism underlying VPIDM, and develop insights regarding VPIDM's applications in SE and ASR using VPIDM as a frontend. Our proposed approach, evaluated on two distinct data sets, demonstrates VPIDM's superior performances over conventional discriminative SE algorithms. Furthermore, we assess the performance of the proposed model under varying signal-to-noise ratio (SNR) levels. The investigation reveals VPIDM's improved robustness in target noise elimination when compared to VEIDM. Furthermore, utilizing the mid-outputs of both VPIDM and VEIDM results in enhanced ASR accuracies, thereby highlighting the practical efficacy of our proposed approach. Code and audio examples are available onlinehttps://github.com/zelokuo/VPIDM.
Zilu Guo, Qing Wang 0008, Jun Du 0002, Qingfeng Liu, Chin-Hui Lee 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 Collaborative Viseme Subword and End-to-End Modeling for Word-Level Lip Reading
abstract
We propose a viseme subword modeling (VSM) approach to improve the generalizability and interpretability capabilities of deep neural network based lip reading. A comprehensive analysis of preliminary experimental results reveals the complementary nature of the conventional end-to-end (E2E) and proposed VSM frameworks, especially concerning speaker head movements. To increase lip reading accuracy, we propose hybrid viseme subwords and end-to-end modeling (HVSEM), which exploits the strengths of both approaches through multitask learning. As an extension to HVSEM, we also propose collaborative viseme subword and end-to-end modeling (CVSEM), which further explores the synergy between the VSM and E2E frameworks by integrating a state-mapped temporal mask (SMTM) into joint modeling. Experimental evaluations using different model backbones on both the LRW and LRW-1000 datasets confirm the superior performance and generalizability of the proposed frameworks. Specifically, VSM outperforms the baseline E2E framework, while HVSEM outperforms VSM in a hybrid combination of VSM and E2E modeling. Building on HVSEM, CVSEM further achieves impressive accuracies on 90.75% and 58.89%, setting new benchmarks for both datasets.
Hang Chen 0001, Qing Wang 0008, Jun Du 0002, Genshun Wan, Shifu Xiong, Chin-Hui Lee 0001
IEEE Trans. Multim.2
2023 Incorporating Lip Features into Audio-Visual Multi-Speaker DOA Estimation by Gated Fusion
abstract
The audio-visual direction of arrival (DOA) estimation has demonstrated superior performance recently. In this paper, we present a novel audio-visual multi-speaker DOA estimation network, which for the first time incorporates multi-speaker lip features to adapt the complex overlapping and noisy scenarios. Firstly, we encode the multi-channel audio features, the reference angles and the lip Regions of Interest (RoIs) detected from the video respectively to acquire high-level representations. Then the multi-modal embeddings of audio, speaker angles and lips are fused by a tri-modal gated fusion module to balance their contributions to the output. The fused embedding is sent to the backend network to obtain the accurate DOA estimation with the combination of the predicted speaker angular vectors and the speaker activities. Experimental results show that our proposed approach can reduce the localization error by 73.48% compared to the previous work on the 2021 Multi-modal Information based Speech Processing (MISP) Challenge corpus. Meanwhile, the high accuracy and stability of localization results demonstrate the robustness of the proposed model in multi-speaker scenarios.
Ya Jiang, Hang Chen 0001, Jun Du 0002, Qing Wang 0008, Chin-Hui Lee 0001
ICASSP4
2023 An Experimental Study on Sound Event Localization and Detection Under Realistic Testing Conditions
abstract
We study four data augmentation (DA) techniques and two model architectures on realistic data for sound event localization and detection (SELD). First, based on ResNet-Conformer (RC), we compare the four DA approaches on the realistic DCASE 2022 SELD test set which is often not easy to handle due to room reverberations and audio overlaps in spontaneous recordings. Experimental results show that, except for audio channel swapping (ACS), the other three data augmentation methods that work well on the simulated SELD data set are no longer effective due to mismatches between simulated and realistic conditions. Next, using ACS-based augmentation, the two improved ResNet-Conformer networks further enhance SELD performances in realistic conditions. By incorporating these two sets of techniques, our overall system ranked the first place in SELD task of the DCASE 2022 Challenge.
Shutong Niu, Jun Du 0002, Qing Wang 0008, Li Chai 0002, Huaxin Wu, Zhaoxu Nian, Lei Sun 0010, Chin-Hui Lee 0001
ICASSP3
2023 Loss Function Design for DNN-Based Sound Event Localization and Detection on Low-Resource Realistic Data
abstract
This study focuses on the design of a loss function for a deep neural network (DNN)-based model with two branches, which is used to solve sound event localization and detection (SELD) on low-resource realistic data. To this end, we employ a secondary network for audio classification, which provides global event information to the main network, enabling it to make robust SELD predictions. Furthermore, we suggest utilizing a momentum strategy for direction-of-arrival (DOA) estimation, taking advantage of the strong temporal consistency of sound events, thereby effectively reducing localization error. Lastly, we incorporate a regularization term into the loss function to alleviate the overfitting problem on the small dataset. We evaluate our proposed methods on the Detection and Classification of Acoustic Scenes and Events (DCASE) 2022 Task 3 dataset, and the results demonstrate consistent improvements in SELD performance. In comparison to the baseline system, the proposed loss function yields significantly improved results for both localization and detection metrics on realistic data. Moreover, the proposed loss function demonstrates its ability to generalize across different network architectures, as evidenced by the consistent improvements achieved.
Qing Wang 0008, Jun Du 0002, Zhaoxu Nian, Shutong Niu, Li Chai 0002, Huaxin Wu, Chin-Hui Lee 0001
ICASSP1
2023 The NERCSLIP-USTC System for the L3DAS23 Challenge Task2: 3D Sound Event Localization and Detection (SELD)
abstract
Sound event localization and detection (SELD) aims at identifying the temporal activities of a known set of sound event classes and estimating their locations. It remains challenging especially when there are overlapped acoustic events. In this work, a robust network architecture with data augmentation techniques is proposed to improve SELD performance, where ResNet and Conformer blocks are combined to model both local and global patterns. To address the data sparsity issue in SELD, SpecAugment, mixup and audio channel swapping (ACS) techniques are adopted. Our proposed system is evaluated in the Task2 of the L3DAS23 challenge and ranks the second place, achieving significant improvements over the baseline.
Haoyin Yan, Qing Wang 0008, Jie Zhang 0042
ICASSP3
2023 Hierarchical Audio-Visual Information Fusion with Multi-label Joint Decoding for MER 2023
abstract
In this paper, we propose a novel framework for recognizing both discrete and dimensional emotions. In our framework, deep features extracted from foundation models are used as robust acoustic and visual representations of raw video. Three different structures based on attention-guided feature gathering (AFG) are designed for deep feature fusion. Then, we introduce a joint decoding structure for emotion classification and valence regression in the decoding stage. A multi-task loss based on uncertainty is also designed to optimize the whole process. Finally, by combining three different structures on the posterior probability level, we obtain the final predictions of discrete and dimensional emotions. When tested on the dataset of multimodal emotion recognition challenge (MER 2023), the proposed framework yields consistent improvements in both emotion classification and valence regression. Our final system achieves state-of-the-art performance and ranks third on the leaderboard on MER-MULTI sub-challenge.
Yuxuan Xi, Hang Chen 0001, Jun Du 0002, Yan Song 0001, Qing Wang 0008, Hengshun Zhou, Jiefeng Ma, Pengfei Hu 0006, Ya Jiang, Shi Cheng 0001, Jie Zhang 0042, Yuzhe Weng
ACM Multimedia6
2023 Using iterative adaptation and dynamic mask for child speech extraction under real-world multilingual conditions
Shi Cheng 0001, Jun Du 0002, Shutong Niu, Alejandrina Cristià, Xin Wang 0037, Qing Wang 0008, Chin-Hui Lee 0001
Speech Commun.6
2023 A Four-Stage Data Augmentation Approach to ResNet-Conformer Based Acoustic Modeling for Sound Event Localization and Detection
abstract
In this paper, we propose a novel four-stage data augmentation approach to ResNet-Conformer based acoustic modeling for sound event localization and detection (SELD). First, we explore two spatial augmentation techniques, namely audio channel swapping (ACS) and multi-channel simulation (MCS), to deal with data sparsity in SELD. ACS and MDS focus on augmenting the limited training data with expanding direction of arrival (DOA) representations such that the acoustic models trained with the augmented data are robust to localization variations of acoustic sources. Next, time-domain mixing (TDM) and time-frequency masking (TFM) are also investigated to deal with overlapping sound events and data diversity. Finally, ACS, MCS, TDM and TFM are combined in a step-by-step manner to form an effective four-stage data augmentation scheme. Tested on the Detection and Classification of Acoustic Scenes and Events (DCASE) 2020 data set, our proposed augmentation approach greatly improves the system performance, ranking our submitted system in the first place in the SELD task of the DCASE 2020 Challenge. Furthermore, we employ a ResNet-Conformer architecture to model both global and local context dependencies of an audio sequence and win the first place in the DCASE 2022 SELD evaluations.
Qing Wang 0008, Jun Du 0002, Huaxin Wu, Chin-Hui Lee 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2022 Deep Segment Model for Acoustic Scene Classification
Yajian Wang, Jun Du 0002, Hang Chen 0001, Qing Wang 0008, Chin-Hui Lee 0001
INTERSPEECH4
2021 Speech Enhancement Autoencoder with Hierarchical Latent Structure
abstract
A new hierarchical convolutional neural network-based autoencoder architecture called SEHAE (Speech Enhancement Hierarchical AutoEncoder) is introduced, in which the latent representation is decomposed into several parts that correspond to different scales. The model consists of three functionally different components. First, a stack of encoders generates a set of latent vectors that contain information from an increasingly larger receptive field. Second, the decoders construct the clean speech in a stage-wise and additive fashion, starting from a learned initial vector. The third component, which we call funnel networks, is tasked with "knitting" together the outputs of the previous decoder and the encoder to compute latent vectors for the next decoder. Several options for initial vectors are explored. Experiments show that SEHAE achieves significant improvements for the considered speech quality and intelligibility measures, outperforming a denoising autoencoder and other step-wise models. Furthermore, its internal workings are investigated using the intermediate results from the decoders.
Koen Oostermeijer, Jun Du 0002, Qing Wang 0008, Chin-Hui Lee 0001
ICASSP3
2021 MRD: A Memory Relation Decoder for Online Handwritten Mathematical Expression Recognition
Qing Wang 0008, Jun Du 0002, Jianshu Zhang 0001, Bin Wang 0070, Bo Ren 0002
ICDAR (3)2
2021 Lightweight Causal Transformer with Local Self-Attention for Real-Time Speech Enhancement
Koen Oostermeijer, Qing Wang 0008, Jun Du 0002
Interspeech2
2021 Information Fusion in Attention Networks Using Adaptive and Multi-Level Factorized Bilinear Pooling for Audio-Visual Emotion Recognition
abstract
Multimodal emotion recognition is a challenging task in emotion computing as it is quite difficult to extract discriminative features to identify the subtle differences in human emotions with abstract concept and multiple expressions. Moreover, how to fully utilize both audio and visual information is still an open problem. In this paper, we propose a novel multimodal fusion attention network for audio-visual emotion recognition based on adaptive and multi-level factorized bilinear pooling (FBP). First, for the audio stream, a fully convolutional network (FCN) equipped with 1-D attention mechanism and local response normalization is designed for speech emotion recognition. Next, a global FBP (G-FBP) approach is presented to perform audio-visual information fusion by integrating self-attention based video stream with the proposed audio stream. To improve G-FBP, an adaptive strategy (AG-FBP) to dynamically calculate the fusion weight of two modalities is devised based on the emotion-related representation vectors from the attention mechanism of respective modalities. Finally, to fully utilize the local emotion information, adaptive and multi-level FBP (AM-FBP) is introduced by combining both global-trunk and intra-trunk data in one recording on top of AG-FBP. Tested on the IEMOCAP corpus for speech emotion recognition with only audio stream, the new FCN method outperforms the state-of-the-art results with an accuracy of 71.40%. Moreover, validated on the AFEW database of EmotiW2019 sub-challenge and the IEMOCAP corpus for audio-visual emotion recognition, the proposed AM-FBP approach achieves the best accuracy of 63.09% and 75.49% respectively on the test set.
Hengshun Zhou, Jun Du 0002, Qing Wang 0008, Qingfeng Liu, Chin-Hui Lee 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2020 Geometry Constrained Progressive Learning for Lstm-Based Speech Enhancement
abstract
In our previous work, a progressive learning framework for long short-term memory (LSTM)-based speech enhancement was proposed to improve the performance in low SNR environment, where each LSTM layer is guided to learn an intermediate target with a specific SNR gain via the MMSE criterion. However, the constraint relationship among these targets is not considered in the objective function. In this paper, we incorporate two kinds of geometric constraints among these targets into the objective function to help LSTM achieve better training. One constraint is edge constraint and the other is the centroid constraint. In addition, we propose a method for constructing the intermediate targets online. It saves device storage space and alleviates the trouble of manually constructing intermediate targets. Experiment results demonstrate these geometric constraints can bring remarkable improvements in low SNR environments.
Jun Du 0002, Li Chai 0002, Yannan Wang, Qing Wang 0008, Chin-Hui Lee 0001
ICASSP5
2020 Stroke Based Posterior Attention for Online Handwritten Mathematical Expression Recognition
abstract
Recently, many researches propose to employ attention based encoder-decoder models to convert a sequence of trajectory points into a LaTeX string for online handwritten mathematical expression recognition (OHMER), and the recognition performance of these models critically relies on the accuracy of the attention. In this paper, unlike previous methods which basically employ a soft attention model, we propose to employ a posterior attention model, which modifies the attention probabilities after observing the output probabilities generated by the soft attention model. In order to further improve the posterior attention mechanism, we propose a stroke average pooling layer to aggregate point-level features obtained from the encoder into stroke-level features. We argue that posterior attention is better to be implemented on stroke-level features than point-level features as the output probabilities generated by stroke is more convincing than generated by point, and we prove that through experimental analysis. Validated on the CROHME competition task, we demonstrate that stroke based posterior attention achieves expression recognition rates of 54.26% on CROHME 2014 and 51.75% on CROHME 2016. According to attention visualization analysis, we empirically demonstrate that the posterior attention mechanism can achieve better alignment accuracy than the soft attention mechanism.
Changjie Wu, Qing Wang 0008, Jianshu Zhang 0001, Jun Du 0002, Jiajia Wu 0003, Jin-Shui Hu
ICPR2
2020 A Transformer-based Radical Analysis Network for Chinese Character Recognition
abstract
Recently, a novel radical analysis network (RAN) has the capability of effectively recognizing unseen Chinese character classes and largely reducing the requirement of training data by treating a Chinese character as a hierarchical composition of radicals rather than a single character class. However, when dealing with more challenging issues, such as the recognition of complicated characters, low-frequency character categories, and characters in natural scenes, RAN still has a lot of room for improvement. In this paper, we explore options to further improve the structure generalization and robustness capability of RAN with the Transformer architecture, which has achieved start-of-the-art results for many sequence-to-sequence tasks. More specifically, we propose to replace the original attention module in RAN with the transformer decoder, which is named as a transformer-based radical analysis network (RTN). The experimental results show that the proposed approach can significantly outperform the RAN on both printed Chinese character database and natural scene Chinese character database. Meanwhile, further analysis proves that RTN can be better generalized to complex samples and low-frequency characters, and has better robustness in recognizing Chinese characters with different attributes.
Qing Wang 0008, Jun Du 0002, Jianshu Zhang 0001, Changjie Wu
ICPR2
2018 A Multiobjective Learning and Ensembling Approach to High-Performance Speech Enhancement With Compact Neural Network Architectures
abstract
In this study, we propose a novel deep neural network (DNN) architecture for speech enhancement (SE) via a multiobjective learning and ensembling (MOLE) framework to achieve a compact and lowlatency design, while maintaining good performance in quality evaluations. MOLE follows the boosting concept when combining weak models into a strong classifier and consists of two compact DNNs. The first, called the multiobjective learning DNN (MOL-DNN), takes multiple features, such as log-power spectra (LPS), mel-frequency cepstral coefficients (MFCCs) and Gammatone frequency cepstral coefficients (GFCCs) to predict a multiobjective set that includes clean speech feature, dynamic noise feature, and ideal ratio mask (IRM). The second, called the multiobjective ensembling DNN (MOE-DNN), takes the learned features from MOL-DNN as inputs and separately predicts clean LPS and IRM, clean MFCC and IRM, and clean GFCC and IRM using three sets of weak regression functions. Finally, a postprocessing operation can be applied to the estimated clean features by leveraging the multiple targets learned from both the MOL-DNN and the MOE-DNN. On speech corrupted by 15 noise types not seen in model training the SE results show that the MOLE approach, which features a small model size and low run-time latency, can achieve consistent improvements over both DNN- and long short-term memory (LSTM)-based techniques in terms of all the objective metrics evaluated in this study for all three cases (the input contexts contain 1-frame, 4-frame and 7-frame instances). The 1-frame MOLE-based SE system outperforms the DNN-based SE system with a 7-frame input expansion at a 3-frame delay and also achieves better performance than the LSTM-based SE system with 4-frame, no delay expansion by including only 3 previous frames, and with 170 times less processing latency.
Qing Wang 0008, Jun Du 0002, Li-Rong Dai 0001, Chin-Hui Lee 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2017 An information fusion framework with multi-channel feature concatenation and multi-perspective system combination for the deep-learning-based robust recognition of microphone array speech
Yanhui Tu, Jun Du 0002, Qing Wang 0008, Xiao Bao, Li-Rong Dai 0001, Chin-Hui Lee 0001
Comput. Speech Lang.3
2015 An information fusion approach to recognizing microphone array speech in the CHiME-3 challenge based on a deep learning framework
abstract
We present an information fusion approach to robust recognition of microphone array speech for the recently launched 3rd CHiME Challenge. It is based on a deep learning framework with a large neural network consisting of subnets with different architectures. Multiple knowledge sources are integrated via an early fusion of normalized noisy features with different beamforming techniques, speech enhanced features, speaker related features, and other auxiliary features concatenated as the input to each subnet, and a late fusion by combining the outputs of all subnets to produce one single output set. Our experiments demonstrate that all information sources are complementary in our proposed framework. Our best system achieves an average word error rate reduction of 68% from the officially released baseline results on the test set of real data.
Jun Du 0002, Qing Wang 0008, Yanhui Tu, Xiao Bao, Li-Rong Dai 0001, Chin-Hui Lee 0001
ASRU2
2015 A universal VAD based on jointly trained deep neural networks
Qing Wang 0008, Jun Du 0002, Xiao Bao, Zi-Rui Wang, Li-Rong Dai 0001, Chin-Hui Lee 0001
INTERSPEECH1
2014 Robust speech recognition with speech enhanced deep neural networks
Jun Du 0002, Qing Wang 0008, Tian Gao 0005, Yong Xu 0004, Li-Rong Dai 0001, Chin-Hui Lee 0001
INTERSPEECH2