Haoxu Wang

dblp:37/8403 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Finite-Blocklength Covert Communications for IRS-Assisted NOMA Networks With Discrete Phase Shifts and Imperfect SIC
Yuan Ren 0003, Haoxu Wang, Fan Jiang 0002, Jing Jiang 0026, Tiejun Lv
IEEE Internet Things J.2
2026 Wireless Channel as a Sensor: An Anti-Electromagnetic Interference Vehicle Detection Method Based on Wireless Sensing Technology
abstract
The management level of Smart Parking Systems (SPS) relies heavily on accurate parking occupancy information, making low-cost, high-precision wireless parking sensors (WPS), powered by batteries, widely used in urban parking lots. However, the performance of magnetometer-based WPS is often disrupted by electromagnetic interference (EMI) from underground high-voltage cables and subways, limiting their reliability in urban environments. This paper proposes an Anti-electromagnetic Interference Parking Detection (AeIPD) method to address this issue. AeIPD combines traditional Received Signal Strength (RSS) features with antenna impedance measurements, utilizing two Bluetooth Low Energy (BLE) transceivers to enhance detection robustness under EMI conditions. Compared to existing methods, AeIPD significantly improves resilience to EMI, providing a more reliable and robust solution for parking detection even in environments with severe interference. This approach offers a cost-effective, scalable solution for large-scale deployment in modern SPS, overcoming the limitations of traditional magnetometer-based systems. Experimental results demonstrate that AeIPD outperforms current parking detection methods, offering a more reliable and robust alternative for smart parking applications.
Liangliang Lou, Haoxu Wang, Miao Zhou, Hanbing Zhao, Wei He 0009
IEEE Trans. Intell. Transp. Syst.3
2026 Set-Based Asynchronous State Estimation for Networked Switched Neural Networks
abstract
This study addresses the zonotopic interval estimation for a class of discrete-time switched neural networks (NNs) under an event-triggered mechanism. Unlike the existing studies, it adopts zonotopic set-membership estimation and considers general asynchronism, thereby enhancing the accuracy and applicability of the estimation results. First, by constructing appropriate observer-mode-dependent Lyapunov functions, certain sufficient conditions are established to guarantee the stability and${\ell }_{\infty }$performance of the augmented error system. Based on these conditions, we present a co-design approach for the switching signals and event-triggered nonlinear observer. Furthermore, a time-varying state zonotope is established, and the corresponding estimated bounds are derived. Finally, the efficacy of the proposed interval estimation approach is validated through two illustrative examples.
Haoxu Wang, Xiaohua Wang 0011, Ying Zhao 0010, Xudong Wang 0008
IEEE Trans. Syst. Man Cybern. Syst.2
2025 ZipEnhancer: Dual-Path Down-Up Sampling-based Zipformer for Monaural Speech Enhancement
abstract
In contrast to other sequence tasks modeling hidden layer features with three axes, Dual-Path time and time-frequency domain speech enhancement models are effective and have low parameters but are computationally demanding due to their hidden layer features with four axes. We propose ZipEnhancer, which is Dual-Path Down-Up Sampling-based Zipformer for Monaural Speech Enhancement, incorporating time and frequency domain Down-Up sampling to reduce computational costs. We introduce the ZipformerBlock as the core block and propose the design of the Dual-Path DownSampleStacks that symmetrically scale down and scale up. Also, we introduce the ScaleAdam optimizer and Eden learning rate scheduler to improve the performance further. Our model achieves new state-of-the-art results on the DNS 2020 Challenge and Voicebank+DEMAND datasets, with a perceptual evaluation of speech quality (PESQ) of 3.69 and 3.63, using 2.04M parameters and 62.41G FLOPS, outperforming other methods with similar complexity levels.
Haoxu Wang
ICASSP1
2025 Exploring Efficient Directional and Distance Cues for Regional Speech Separation
Yiheng Jiang, Haoxu Wang, Yafeng Chen, Gang Qiao
INTERSPEECH2
2025 FLASepformer: Efficient Speech Separation with Gated Focused Linear Attention Transformer
Haoxu Wang, Yiheng Jiang, Gang Qiao, Pengteng Shi
INTERSPEECH1
2025 SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training
abstract
The efficiency of attention is important due to its quadratic time complexity. We enhance the efficiency of attention through two key contributions: First, we leverage the new $\texttt{FP4}$ Tensor Cores in Blackwell GPUs to accelerate attention computation. Our implementation achieves $\textbf{1038}$ $\texttt{TOPS}$ on $\texttt{RTX5090}$, which is a $\textbf{5}\times$ speedup over the fastest FlashAttention on $\texttt{RTX5090}$. Experiments show that our $\texttt{FP4}$ attention can accelerate inference of various models in a plug-and-play way. Second, we pioneer low-bit attention to training tasks. Existing low-bit attention works like FlashAttention3 and SageAttention focus only on inference. However, the efficiency of training large models is also important. To explore whether low-bit attention can be effectively applied to training tasks, we design an accurate and efficient $\texttt{8-bit}$ attention for both forward and backward propagation. Experiments indicate that $\texttt{8-bit}$ attention achieves lossless performance in fine-tuning tasks but exhibits slower convergence in pretraining tasks. The code is available at https://github.com/thu-ml/SageAttention.
Haoxu Wang, Pengle Zhang, Haofeng Huang, Jianfei Chen 0001, Jun Zhu 0001
NeurIPS3
2025 Store-and-forward with graph attention: Enhanced multi-agent reinforcement learning for emergency-responsive traffic signal control
Kangkang Yang, Zhiwen Wang 0003, Xinyou Meng, Yaoke Shi, Yuling Yu, Haoxu Wang, Ziheng Yao
Eng. Appl. Artif. Intell.7
2024 Robust Wake Word Spotting With Frame-Level Cross-Modal Attention Based Audio-Visual Conformer
abstract
In recent years, neural network-based Wake Word Spotting achieves good performance on clean audio samples but struggles in noisy environments. Audio-Visual Wake Word Spotting (AVWWS) receives lots of attention because visual lip movement information is not affected by complex acoustic scenes. Previous works usually use simple addition or concatenation for multi-modal fusion. The inter-modal correlation remains relatively under-explored. In this paper, we propose a novel module called Frame-Level Cross-Modal Attention (FLCMA) to improve the performance of AVWWS systems. This module can help model multi-modal information at the frame-level through synchronous lip movements and speech signals. We train the end-to-end FLCMA based Audio-Visual Conformer and further improve the performance by fine-tuning pre-trained uni-modal models for the AVWWS task. The proposed system achieves a new state-of-the-art result (4.57% WWS score) on the far-field MISP dataset.
Haoxu Wang, Ming Cheng 0005, Qiang Fu 0001, Ming Li 0026
ICASSP1
2024 SlideSpeech: A Large Scale Slide-Enriched Audio-Visual Corpus
abstract
Multi-Modal automatic speech recognition (ASR) techniques aim to leverage additional modalities to improve the performance of speech recognition systems. While existing approaches primarily focus on video or contextual information, the utilization of extra supplementary textual information has been overlooked. Recognizing the abundance of online conference videos with slides, which provide rich domain-specific information in the form of text and images, we release SlideSpeech, a large-scale audio-visual corpus enriched with slides. The corpus contains 1,705 videos, 1,000+ hours, with 473 hours of high-quality transcribed speech. Moreover, the corpus contains a significant amount of real-time synchronized slides. In this work, we present the pipeline for constructing the corpus and propose baseline methods for utilizing text information in the visual slide context. Through the application of keyword extraction and contextual ASR methods in the benchmark system, we demonstrate the potential of improving speech recognition performance by incorporating textual information from supplementary video slides.
Haoxu Wang, Fan Yu 0002, Xian Shi, Yuezhang Wang, Shiliang Zhang, Ming Li 0026
ICASSP1
2024 Hourglass-AVSR: Down-Up Sampling-Based Computational Efficiency Model for Audio-Visual Speech Recognition
abstract
Recently audio-visual speech recognition (AVSR), which better leverages video modality as additional information to extend automatic speech recognition (ASR), has shown promising results in complex acoustic environments. However, there is still substantial space to improve as complex computation of visual modules and ineffective fusion of audio-visual modalities. To eliminate these drawbacks, we propose a down-up sampling-based AVSR model (Hourglass-AVSR) to enjoy high efficiency and performance, whose time length is scaled during the intermediate processing, resembling an hourglass. Firstly, we propose a context and residual aware video upsampling approach to improve the recognition performance, which utilizes contextual information from visual representations and captures residual information between adjacent video frames. Secondly, we introduce a visual-audio alignment approach during the upsampling by explicitly incorporating boundary constraint loss. Besides, we propose a cross-layer attention fusion to capture the modality dependencies within each visual encoder layer. Experiments conducted on the MISP-AVSR dataset reveal that our proposed Hourglass-AVSR model outperforms ASR model by 12.9% and 20.8% relative concatenated minimum permutation character error rate (cpCER) reduction on far-field and middle-field test sets, respectively. Moreover, compared to other state-of-the-art AVSR models, our model exhibits the highest improvement in cpCER for the visual module. Furthermore, on the benefit of our down-up sampling approach, Hourglass-AVSR model reduces 54.2% overall computation costs with minor performance degradation.
Fan Yu 0002, Haoxu Wang, Ziyang Ma 0001, Shiliang Zhang
ICASSP2
2024 LCB-Net: Long-Context Biasing for Audio-Visual Speech Recognition
abstract
The growing prevalence of online conferences and courses presents a new challenge in improving automatic speech recognition (ASR) with enriched textual information from video slides. In contrast to rare phrase lists, the slides within videos are synchronized in real-time with the speech, enabling the extraction of long contextual bias. Therefore, we propose a novel long-context biasing network (LCB-net) for audio-visual speech recognition (AVSR) to leverage the long-context information available in videos effectively. Specifically, we adopt a bi-encoder architecture to simultaneously model audio and long-context biasing. Besides, we also propose a biasing prediction module that utilizes binary cross entropy (BCE) loss to explicitly determine biased phrases in the long-context biasing. Furthermore, we introduce a dynamic contextual phrases simulation to enhance the generalization and robustness of our LCB-net. Experiments on the SlideSpeech, a large-scale audio-visual corpus enriched with slides, reveal that our proposed LCB-net outperforms general ASR model by 9.4%/9.1%/10.9% relative WER/U-WER/B-WER reduction on test set, which enjoys high unbiased and biased performance. Moreover, we also evaluate our model on LibriSpeech corpus, leading to 23.8%/19.2%/35.4% relative WER/U-WER/B-WER reduction over the ASR model.
Fan Yu 0002, Haoxu Wang, Xian Shi, Shiliang Zhang
ICASSP2
2024 Memory-Efficient and Secure DNN Inference on TrustZone-enabled Consumer IoT Devices
abstract
Edge intelligence enables resource-demanding Deep Neural Network (DNN) inference without transferring original data, addressing concerns about data privacy in consumer Inter-net of Things (IoT) devices. For privacy-sensitive applications, deploying models in hardware-isolated trusted execution environments (TEEs) becomes essential. However, the limited secure memory in TEEs poses challenges for deploying DNN inference, and alternative techniques like model partitioning and offloading introduce performance degradation and security issues. In this paper, we present a novel approach for advanced model deployment in TrustZone that ensures comprehensive privacy preservation during model inference. We design a memory-efficient management method to support memory-demanding inference in TEEs. By adjusting the memory priority, we effectively mitigate memory leakage risks and memory overlap conflicts, resulting in 32 lines of code alterations in the trusted operating system. Additionally, we leverage two tiny libraries: S-Tinylib (2,538 LoCs), a tiny deep learning library, and Tinylibm (827 LoCs), a tiny math library, to support efficient inference in TEEs. We implemented a prototype on Raspberry Pi 3B+ and evaluated it using three well-known lightweight DNN models. The experimental results demonstrate that our design significantly improves inference speed by 3.13 times and reduces power consumption by over 66.5% compared to non-memory optimization method in TEEs.
Xueshuo Xie, Haoxu Wang, Zhaolong Jian, Tao Li 0022, Wei Wang 0012, Grace Guiling Wang
INFOCOM2
2023 The WHU-Alibaba Audio-Visual Speaker Diarization System for the MISP 2022 Challenge
abstract
This paper describes the system developed by the WHU-Alibaba team for the Multimodal Information Based Speech Processing (MISP) 2022 Challenge. We extend the Sequence-to-Sequence Target-Speaker Voice Activity Detection framework to simultaneously detect multiple speakers’ voice activities from audio-visual signals. The final system achieves a diarization error rate (DER) of 8.82% on the evaluation set of the competition database, which ranks 1st in the speaker diarization track of the MISP 2022, ICASSP Signal Processing Grand Challenge.
Ming Cheng 0005, Haoxu Wang, Qiang Fu 0001, Ming Li 0026
ICASSP2
2023 The DKU Post-Challenge Audio-Visual Wake Word Spotting System for the 2021 MISP Challenge: Deep Analysis
abstract
This paper further explores our previous wake word spotting system ranked 2-nd in Track 1 of the MISP Challenge 2021. First, we investigate a robust unimodal approach based on 3D and 2D convolution and adopt the simple attention module (SimAM) for our system to improve performance. Second, we explore different combinations of data augmentation methods for better performance. Finally, we study the fusion strategies, including score-level, cascaded and neural fusion. Our proposed multimodal system leverages multimodal features and uses the complementary visual information to mitigate the performance degradation of audio-only systems in complex acoustic scenarios. Our system obtains a false reject rate of 2.15% and a false alarm rate of 3.44% in the evaluation set of the competition database, which achieves the new state-of-the-art performance by 21% relative improvement compared to previous systems. Related resource can be found at: https://github.com/Mashiro009/DKU_WWS_MISP.
Haoxu Wang, Ming Cheng 0005, Qiang Fu 0001, Ming Li 0026
ICASSP1
2022 The DKU Audio-Visual Wake Word Spotting System for the 2021 MISP Challenge
abstract
This paper describes the system developed by the DKU team for the MISP Challenge 2021. We present a two-stage approach consisting of end-to-end neural networks for the audio-visual wake word spotting task. We first process audio and video data to give them a similar structure and then train two unimodal models with unified network architecture separately. Second, we propose a Hierarchical Modality Aggregation (HMA) module that fuses multi-scale audio-visual information from pre-trained unimodal models. Our system has a clear and concise framework consisting of end-to-end neural networks. With this framework and extensive data augmentation methods, our presented system achieves a false reject rate of 3.85% and a false alarm rate of 3.42% on far-field audio in the development set of the competition database, which ranks 2nd in the wake word spotting track of the MISP challenge.
Ming Cheng 0005, Haoxu Wang, Yechen Wang, Ming Li 0026
ICASSP2