De Hu

dblp:281/1200 · DBLP profile ↗
← Back
19ranked-venue papers
12as first author
18since 2021 · last 2026
0000-0003-0226-917XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 7 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 10 since 2021Computer networks · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Graph Neural Field with Spatial-Correlation Augmentation for HRTF Personalization
abstract
To achieve immersive spatial audio rendering on VR/AR devices, high-quality Head-Related Transfer Functions (HRTFs) are essential. In general, HRTFs are subject-dependent and position-dependent, and their measurement is time-consuming and tedious. To address this challenge, we propose the Graph Neural Field with Spatial-Correlation Augmentation (GraphNF-SCA) for HRTF personalization, which can be used to generate individual HRTFs for unseen subjects. The GraphNF-SCA consists of three key components: an HRTF personalization (HRTF-P) module, an HRTF upsampling (HRTF-U) module, and a fine-tuning stage. In the HRTF-P module, we predict HRTFs of the target subject via the Graph Neural Network (GNN) with an encoder-decoder architecture, where the encoder extracts universal features and the decoder incorporates the target-relevant features and produces individualized HRTFs. The HRTF-U module employs another GNN to model spatial correlations across HRTFs. This module is fine-tuned using the output of the HRTF-P module, thereby enhancing the spatial consistency of the predicted HRTFs. Unlike existing methods that estimate individual HRTFs position-by-position without spatial correlation modeling, the GraphNF-SCA effectively leverages inherent spatial correlations across HRTFs to enhance the performance of HRTF personalization. Experimental results demonstrate that the GraphNF-SCA achieves state-of-the-art results.
De Hu, Junsheng Hu
AAAI1
2026 Robust Self-Localization of Wireless Acoustic Sensor Networks in the Presence of TDoA Outliers
abstract
Today, we are surrounded by numerous intelligent terminals (e.g., smartphones, tablets, and laptops) that can collectively form wireless acoustic sensor networks (WASNs), which are expected to be the next-generation platform for sound processing. Since the geometric structure of WASNs is critical for tasks like source localization and acoustic beamforming, automatic self-localization of WASNs (SL-WASNs) is indispensable. Although several SL-WASN methods have been developed, most of them fail to account for the impact of time-difference-of-arrival (TDoA) outliers induced by noise and reverberation. To address this issue, we first propose a centralized-robust self-localization (CRSL) method for WASNs with a centralized network topology. Specifically, we develop a novel probabilistic model leveraging the zero-sum property of TDoAs to derive the probability of a TDoA measurement being an inlier (or outlier). Then, we construct a weighted cost function based on TDoA measurements, in which the weights are controlled by the computed inlier probabilities. Afterward, the Broyden-Fletcher-Goldfarb-Shanno (BFGS) algorithm is employed to simultaneously estimate the positions of nodes and sources, as well as the inter-node time offsets. In addition, we extend the CRSL to a distributed framework, termed distributed-robust self-localization (DRSL), which incorporates distributed computation of inlier probabilities, distributed cost function construction, and distributed optimization. Furthermore, the computational complexities of both CRSL and DRSL are analyzed in detail. The main contributions of this work can be summarized as follows: a novel TDoA outlier recognition model, the CRSL method, the DRSL method, and their complexity analysis. In real-world experiments with M = 10 nodes and N = 8 sources, the proposed methods achieved over 8% higher localization accuracy than state-of-the-art approaches.
De Hu, Xu Wang 0058
IEEE Trans. Mob. Comput.1
2025 Distributed-Robust Source Localization in Wireless Acoustic Sensor Networks
abstract
In this paper, we propose a distributed-robust sound source localization (D-R-SSL) method in wireless acoustic sensor networks (WASNs). Specifically, the space is first divided into small equal-sized grids, and the local probability of each grid containing the source is calculated by using intra-node time difference of arrivals (TDoAs). Next, we allocate weights for local probabilities from different nodes, and develop a distributed weight calculation strategy which is robust against the TDoA outliers. Then, the source position is derived from the weighted summation of local probabilities. In order to improve the SSL efficiency, we further present a fast implementation approach (named as F-D-R-SSL). Finally, we also analyse the computational complexity and communication loads of the proposed methods. Simulation and real-world experiments demonstrate that the F-D-R-SSL is as powerful or even stronger than state-of-the-art non-grid SSLs.
Xu Wang 0058, De Hu, Qintuya Si
ICASSP2
2025 Joint Rate Allocation and Sensor Selection for Speech Enhancement in Wireless Acoustic Sensor Networks
De Hu
INTERSPEECH1
2025 Joint Reference Microphone Selection and Filter Order Determination in Multi-channel Active Noise Control
De Hu, Shuyao Liu, Yanrong He
INTERSPEECH1
2025 D-GAT: Dual Graph Attention Network for Global HRTF Interpolation
Junsheng Hu, Shaojie Li 0001, Qintuya Si, De Hu
INTERSPEECH4
2025 Temporal Convolutional Network with Smoothed and Weighted Losses for Distant Voice Activity and Overlapped Speech Detection
Shaojie Li 0001, Qintuya Si, De Hu
INTERSPEECH3
2025 Robust Self-Localization of Wireless Acoustic Sensor Networks
abstract
Wireless acoustic sensor networks (WASNs), or the so-called Internet of Audio Things (IoAuT), have attracted increasing attention in the Internet of Things community. As the geometric structure of WASNs is required in audio/speech processing tasks like source localization or acoustic beamforming, automatic self-localization of sensors is necessary. However, most of the existing approaches suffer from poor stability, as their constructed cost functions involve nonconvex programming. To address this issue, we investigate the robust self-localization (or geometry calibration) of WASNs in this article. Specifically, a rough self-localization (RSL) method is first presented based on measurements including Time-Difference-of-Arrivals (TDoAs), direction-of-arrivals (DoA), and energy-rates (ERs), and its closed-form solution is further derived. As ER estimates are sensitive to acoustic environments, the performance of the RSL method is somewhat limited. Therefore, a precise self-localization (PSL) method is then developed by building a weighted (and nonconvex) TDoA-DoA cost function, after regarding the RSL approach as an initialization step. As the RSL offers better initial values compared with existing initialization strategies, the combination of RSL and PSL methods (named as RSL-PSL method) shows stronger robustness and stability. In addition, computational complexity of both RSL and PSL methods is analyzed in detail. Finally, the Cramér-Rao Bound (CRB) of the PSL method is derived to show its theoretical lower bound. The proposed RSL-PSL method outperforms the state-of-the-arts in terms of stability and accuracy, which is confirmed by numerical real-world and simulation experiments.
Xu Wang 0058, De Hu, Rui Liu 0008, Feilong Bao
IEEE Internet Things J.2
2024 Parametric Binaural Beamforming Based on Auditory Perception
abstract
Due to the compact nature, hearing aids are often equipped with only a small number of microphones. Such a restriction brings a conflict between noise reduction and spatial cue retention in binaural beamformers (BFs). To alleviate this conflict, we design a parametric binaural (PaBi) BF from the viewpoint of auditory perception. In human hearing, the binaural cues are frequency selective, i.e., the interaural phase difference (IPD) dominates at low frequencies while the interaural level difference (ILD) dominates at high frequencies. Accordingly, we construct a set of parametric IPD and ILD constraints to establish a novel cost function, which is then solved by the semi-definite relaxation strategy. By adjusting the involved parameters, a good trade-off between noise reduction and spatial cue preservation can be achieved. Moreover, the PaBi BF breaks through the degree-of-freedom limitation of existing methods. Experimental results show the superiority of the proposed method.
De Hu, Xinzhe Zhang
IEEE Signal Process. Lett.1
2024 Distributed-Robust MVDR Beamforming With Energy-Efficient Topology Control in Wireless Acoustic Sensor Networks
abstract
We are often surrounded by intelligent devices with one or more acoustic sensors, which constitute a wireless acoustic sensor network (WASN) and can be exploited for various audio/speech processing tasks. As the captured audio signals are inevitably corrupted by ambient noises, signal enhancement is vital in WASNs. To this end, this paper proposes a distributed-robust and energy-efficient MVDR beamformer (BF) for WASNs. Specifically, a distributed MVDR BF is first derived by recursively updating the inverse of the noise correlation matrix, which requires fewer data transmission without sacrificing performance. Then, its robust version is further designed to alleviate the adverse effects of parameter mismatches. Finally, an energy-efficient network topology control (EENTC) is carried out to reduce the energy consumption, by optimizing the weight matrix of the distributed averaging process. Since the proposed EENTC strategy involves non-convex programming, we transform it into a convex one and solve it via the Dinkelbach algorithm. Unlike the centralized BFs, the proposed method works without an additional central processor. Moreover, it is robust against parameter mismatch during beamforming and can reduce a large amount of data transmission. Simulation and real-world experimental results confirm the validity of the proposed method.
De Hu, Qintuya Si, Weiwei Zhang 0008
IEEE Trans. Wirel. Commun.1
2023 C2ST: Cross-modal Contextualized Sequence Transduction for Continuous Sign Language Recognition
abstract
Continuous Sign Language Recognition (CSLR) aims to transcribe the signs of an untrimmed video into written words or glosses. The mainstream framework for CSLR consists of a spatial module for visual representation learning, a temporal module aggregating the local and global temporal information of frame sequence, and the connectionist temporal classification (CTC) loss, which aligns video features with gloss sequence. Unfortunately, the language prior implicit in the gloss sequence is ignored throughout the modeling process. Furthermore, the contextualization of glosses is further ignored in alignment learning, as CTC makes an independence assumption between glosses. In this paper, we propose a Cross-modal Contextualized Sequence Transduction (C2ST) for CSLR, which effectively incorporates the knowledge of gloss sequence into the process of video representation learning and sequence transduction. Specifically, we introduce a cross-modal context learning framework for CSLR, in which the linguistic features of gloss sequences are extracted by a language model, and recurrently integrate with visual features for video modelling. Moreover, we introduce the contextualized sequence transduction loss that incorporates the contextual information of gloss sequences in label prediction, without making any independence assumptions between the glosses. Our method sets the new state of the art on three widely used large-scale sign language recognition datasets: Phoenix-2014, Phoenix-2014-T, and CSL-Daily. On CSL-Daily, our approach achieves an absolute gain of 4.9% WER compared to the best published results.
Huaiwen Zhang, Zihang Guo, Yang Yang 0121, De Hu
ICCV5
2023 Explicit Intensity Control for Accented Text-to-speech
Rui Liu 0008, Haolin Zuo, De Hu, Guanglai Gao, Haizhou Li 0001
INTERSPEECH3
2023 Distributed Self-Localization for Acoustic Transceiver Networks
abstract
Today, we are often surrounded by a lot of acoustic transceivers, such as smartphones, tablets, and smart speakers. If they constitute an acoustic transceiver network (ATN), they can be exploited for various acoustic signal processing tasks. In this letter, we propose a decentralized frame-work for geometry calibration in ATNs. Based on direction-of-arrival (DoA) and time-difference-of-arrival (TDoA) measurements, the self-localization is first implemented in local coordinate systems from different nodes. Subsequently, a distributed consensus approach is derived, which maps these local coordinates to a common one. As a result, the inconsistency among local coordinates is eliminated, and the nodes are localized in a virtual coordinate system. Finally, numerical results validate the proposed method.
Xu Wang 0058, De Hu
IEEE Signal Process. Lett.2
2023 Distributed Sensor Selection for Speech Enhancement With Acoustic Sensor Networks
abstract
In distributed acoustic sensor networks, only a few nodes make a significant contribution to speech enhancement tasks. Using these most informative nodes instead of the entire network not only avoids unnecessary energy consumption but also prolongs the lifetime of sensors. To this end, a sensor selection method for distributed speech enhancement is proposed. The best subset of microphone nodes is determined by maximizing the signal-to-noise ratio (SNR), while keeping the activated nodes connected with each other. The above criterion involves an integer and non-linear programming, which is linearized with multiple base-3 sub-optimization problems, and each of them is solved by a state-of-the-art steepest descent (SD) algorithm. In addition, a greedy searching strategy is presented to select sensors rapidly. Finally, a distributed SD algorithm is further derived, which is more suitable for distributed sensor networks. The proposed method can obtain the optimal subnetwork in noisy and reverberant environments. Unlike the existing approaches, it can select nodes from a microphone network with arbitrary communication graphs. Moreover, it requires only local communications among nodes without an external central processor. Experimental results confirm the validity of the proposed method.
De Hu, Qintuya Si, Rui Liu 0008, Feilong Bao
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 Distributed Sampling Rate Offset Estimation Over Acoustic Sensor Networks Based on Asynchronous Network Newton Optimization
abstract
Sampling rate synchronization is an inevitable issue in distributed acoustic sensor networks. In this paper, an analytical sampling rate offset (SRO) estimation approach is first proposed, and then, it is extended to a distributed method that suitable for acoustic sensor networks with arbitrary communication graphs. Specifically, a linear-phase drift model in the short-time Fourier transform domain is used to approximate the SRO between each pair of microphone nodes. Next, after unwrapping the temporally averaged phase information, SROs are recovered analytically via a new weighted-sum criterion. Based on this, a distributed cost function is established at each node to obtain the SROs of all nodes simultaneously in a distributed manner. Finally, a state-of-the-art distributed algorithm named asynchronous network Newton optimization is adopted to carry out the distributed SRO estimation. The proposed method can effectively estimate the SROs among acoustic sensor nodes in noisy and reverberant environments. Compared with the existing approaches, it does not require an external central processor, and only local communications among nodes are needed. Experimental results confirm the validity of the proposed method.
De Hu, Huaiwen Zhang, Feilong Bao, Rui Wang 0046
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 Acoustic SLAM With Moving Sound Event Based on Auxiliary Microphone Arrays
abstract
Acoustic simultaneous localization and mapping (ASLAM) aim to map the positions of sound sources while passively localizing the microphone array embedded in the robot platform. In this paper, an ASLAM method with auxiliary microphone arrays based on dual interacting multiple models and unscented Kalman filter (D-IMM-UKF) is proposed for the single moving source scenario. Firstly, a dual-unscented Kalman filter is presented, which can simultaneously track the robot and the speaker. Then, the interacting multiple models are adopted for the different motion dynamics of a robot and a speaker in space. To avoid the underdetermined condition when only the acoustic information is available, a small number of static microphone arrays are employed. Finally, the moving robot’s and speaker’s positions are estimated by the D-IMM-UKF algorithm. It can obtain the trajectories of the robot’s and speaker’s movements smoothly with good tracking accuracy. Experimental results verify the effectiveness of the proposed method.
De Hu, Zhe Chen 0005, Fuliang Yin
IEEE Trans. Intell. Transp. Syst.1
2021 Geometry Calibration for Acoustic Transceiver Networks Based on Network Newton Distributed Optimization
abstract
Geometry calibration for distributed acoustic sensor networks is becoming increasing popular in the signal processing community. In this paper, a distributed geometry calibration method based on network Newton distributed optimization is proposed for the acoustic transceiver networks where each node consists of a microphone array and a loudspeaker. After collecting the direction-of-arrival and time-difference-of-arrival measurements, a two-stage centralized cost function is formulated to estimate the geometrical configuration of networks, and the corresponding identifiability conditions are discussed. Next, to achieve the distributed calibration, a distributed cost function is established by splitting the centralized cost function into multiple local cost functions. Finally, the distributed geometry calibration is carried out by using network Newton distributed optimization. The proposed method can effectively estimate the geometry structure of acoustic transceiver networks in noisy and reverberant environments. Compared with the existing approaches, it implements the calibration process in a distributed manner, which requires only the local communication among nodes and does not need an external central processor. Experimental results show the validity of the proposed method.
De Hu, Zhe Chen 0005, Fuliang Yin
IEEE ACM Trans. Audio Speech Lang. Process.1
2021 Passive Geometry Calibration for Microphone Arrays Based on Distributed Damped Newton Optimization
abstract
Geometry calibration is an inherent challenge in distributed acoustic sensor networks. To mitigate this problem, a passive geometry calibration approach based on distributed damped Newton optimization is proposed. Specifically, a geometric cost function incorporating direction of arrivals (DoAs) and time difference of arrivals (TDoAs) is first formulated, and then its identifiability conditions are given. Next, to achieve a distributed geometry calibration, the cost function is split into multiple local cost functions that are assigned to every node. After that, a distributed damped Newton optimization is presented to retrieve the geometry of microphone nodes and synchronize the internal delay between each two neighboring nodes. Finally, computational complexity and transmission bandwidth requirements are further analyzed. Compared with the existing approaches, the proposed method estimates the geometry structure of microphone networks in a distributed manner. Moreover, it requires a small number of acoustic sources. Experimental results show the validity of the proposed method.
De Hu, Zhe Chen 0005, Fuliang Yin
IEEE ACM Trans. Audio Speech Lang. Process.1
2020 Analytical Geometry Calibration for Acoustic Transceiver Arrays
abstract
There are many intelligent devices around us that have the ability to emit and receive audio signals, such as smartphones, smartwatches, and tablets. If they constitute an acoustic transceiver network, it can boost the performance of many audio processing tasks. The speaker localization and tracking algorithms using such a network require the prior information of node positions. To acquire this knowledge, we derive an analytical solution for geometry calibration of acoustic transceiver networks where each node consists of a microphone array and a loudspeaker. Specifically, a linear cost function for node orientations is first established using the delivered direction-of-arrival (DoA) measurements, and its analytical solution is derived. Next, based on the DoA and time-of-arrival (ToA) measurements, another cost function for node positions is formulated, which can also be solved in the closed form. Experimental results show that the proposed method can successfully estimate the geometry structure of acoustic transceiver networks in noisy and reverberant environments. In contrast to most state-of-the-art geometry calibration algorithms, which iteratively solve the calibration problem by optimization methods, the proposed method can decrease the computational complexity greatly.
De Hu, Zhe Chen 0005, Fuliang Yin
IEEE Signal Process. Lett.1