Xionghu Zhong

dblp:92/7819 · DBLP profile ↗
← Back
41ranked-venue papers
10as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 20 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 12 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Computer networks · 3 · 2 since 2021
YearPublicationVenuePosition
2026 TextBFGS: A Case-Based Reasoning Approach to Code Optimization via Error-Operator Retrieval
Zizheng Zhang, Yuyang Liao, Chen Chen 0075, Dun Wu, Qianjin Yu, Yanqin Gao, Kailai Zhang, Chng Eng Siong, Xionghu Zhong
ICCBR11
2025 A Joint Time-Frequency Attention for Leakage Detection in Water Distribution Networks Using Time Series Decomposition
abstract
Detecting leakages in a water distribution network (WDN) is a challenging task due to the complexity of data patterns caused by the pipeline leakages and the volatility of the daily demands. Usually, the data under normal operations are collected and different machine learning algorithms are developed to predict anomalies due to the leaks. However, these methods are overwhelmingly rely on the time domain modeling and ignore the information in the frequency domain, and lack a comprehensive modeling of the data patterns such as shapelet, trend, seasonality and point outliers. In this paper, we propose a joint time-frequency attention (JTFA) approach to detect the WDN leakages. In essence, the received signals are decomposed into trend and residual components to represent the incipient and abrupt leaks separately. Attention models are then applied on both time and frequency domain signals to learn the corresponding patterns. In particular, the spectrum is divided into different frequency bands to better attend the detailed information in the higher frequency bands. The desired signals are subsequently reconstructed and compared to the input signal to generate an anomaly score. Experiments from simulated water supply networks are organized and the results demonstrate that the proposed approach performs better than existing leak detection methods and time-frequency analysis methods.
Juan Luo, Jielong Yang, Xionghu Zhong
ICASSP4
2025 MatCo: Computing Match Cover of Subgraph Query over Graph Data
abstract
Subgraph query can be applied in various scenarios, such as fraud detection and cyberattack pattern analysis. However, computing subgraph queries usually traverses a huge search space. Many efforts have been made to reduce this search space. The size of the answer set can be exponential, providing a substantial lower bound for the search space. Additionally, different answers may overlap, and a single vertex can occur multiple times in different matches. In this paper, we propose a new problem to compute the match cover of a subgraph query. We define the match cover as a subset of answers such that the vertices included are exactly the same as those in the entire set. There can be more than one match covers, however, we only return one, as long as we can avoid the huge overhead of searching the entire set. It is inefficient to apply traditional subgraph query methods for computing match cover. Specifically, existing methods do not prune partial matches that could grow into full matches. For match cover computation, if the vertices in those full matches are already included in previously found matches, continuing the computation over such partial matches is a waste of time. We propose a new framework, called MatCo, to compute the match cover. In MatCo, we design a new data structure, called local candidate space, to determine whether the future search scopes of partial matches have been covered. We can easily maintain local candidate space and efficiently conduct the determination. We also reduce some Cartesian products, which are inevitable in existing methods, into linear enumerations, which significantly improves performance. Extensive experiments over various datasets confirm that our method outperforms comparative ones by 1~3 orders of magnitude. Efficiently computing the minimum match cover could be an interesting future work.
Youhuan Li, Ziming Li 0004, Yuequn Dou, Xionghu Zhong, Lei Zou 0001
Proc. ACM Manag. Data5
2025 UniArray: Unified Spectral-Spatial Modeling for Array-Geometry-Agnostic Speech Separation
abstract
Array-geometry-agnostic speech separation (AGA-SS) aims to develop an effective separation method regardless of the microphone array geometry. Conventional methods rely on permutation-free operations, such as summation or attention mechanisms, to capture spatial information. However, these approaches often incur high computational costs or disrupt the effective use of spatial information during intra- and inter-channel interactions, leading to suboptimal performance. To address these issues, we propose UniArray, a novel approach that abandons the conventional interleaving manner. UniArray consists of three key components: a virtual microphone estimation (VME) module, a feature extraction and fusion module, and a hierarchical dual-path separator. The VME ensures robust performance across arrays with varying channel numbers. The feature extraction and fusion module leverages a spectral feature extraction module and a spatial dictionary learning (SDL) module to extract and fuse frequency-bin-level features, allowing the separator to focus on using the fused features. The hierarchical dual-path separator models feature dependencies along the time and frequency axes while maintaining computational efficiency. Experimental results show that UniArray outperforms state-of-the-art methods in SI-SDRi, WB-PESQ, NB-PESQ, and STOI across both seen and unseen array geometries.
Weiguang Chen, Jielong Yang, Chng Eng Siong, Xionghu Zhong
IEEE Signal Process. Lett.5
2025 Machine Unlearning for Source-Free Unsupervised Partial-Domain Adaptation in Remote Sensing
abstract
Source-Free Unsupervised Domain Adaptation (SFUDA) enables model adaptation to unlabeled target domains without accessing source data. However, when the source domain contains classes absent in the target domain, existing methods suffer from negative transfer: knowledge of irrelevant source-only classes interferes with target class recognition, significantly degrading classification accuracy. We propose Machine Unlearning-based SFUDA (MUSFUDA), which addresses this problem by selectively unlearning source-only class knowledge from the pre-trained model rather than adding compensatory mechanisms. This machine unlearning approach allows the model to focus on shared classes, fundamentally eliminating negative transfer. Remote sensing images with large intra-class variations and high inter-class similarity cause over-unlearning of target classes when forgetting source-only classes, thus we design the Model-Disruption Based Dual-Teacher Unlearning Strategy (MDUS), which uses dual teachers to manage target class preservation and source-only class erasure through knowledge distillation. MDUS is lightweight and easily integrated into existing SFUDA frameworks. Experiments on remote sensing datasets demonstrate that combining MDUS with representative baselines consistently reduces negative transfer and improves classification performance, maintaining high efficiency, validating the effectiveness and generalizability of our approach.
Jielong Yang, Xialun Yun, Xionghu Zhong, Di Wu 0050
IEEE Trans. Geosci. Remote. Sens.4
2024 Enhancing Low-Latency Speaker Diarization with Spatial Dictionary Learning
abstract
This study proposes a low-latency online speaker diarization framework. Specifically, we design a spatial dictionary learning module shared across different frequency bands, enabling spatial feature learning at each frequency bin. This contributes to reducing the latency constraints of the online diarization system. Additionally, a magnitude-weighted fusion is devised to integrate spectral features. Consequently, the system can extract discriminative speaker embeddings by simultaneously considering spectral and spatial features. Experimental results on the Alimeeting dataset demonstrate a significant improvement in diarization error rates across various latencies, with a relative improvement of 45.80% compared to single-channel online diarization. Moreover, our method surpasses offline direction-of-arrival-based diarization and achieves comparable performance to the second-ranked offline system of the Alimeeting challenge.
Weiguang Chen, Xionghu Zhong, Chng Eng Siong
ICASSP3
2024 An Attention Model Based Approach for Leakage Detection in Water Distribution Networks Using Normal Pressure Data
Juan Luo, Du Zhou, Chongxiao Wang, Jielong Yang, Xionghu Zhong
PRICAI (5)5
2024 FR-GNN: Mitigating the Impact of Distribution Shift on Graph Neural Networks via Test-Time Feature Reconstruction
abstract
Due to inappropriate sample selection and limited training data, a distribution shift often exists between the training and test sets. This shift can adversely affect the test performance of graph neural networks (GNNs). Existing approaches mitigate this issue by either enhancing the robustness of GNNs to distribution shift or reducing the shift itself. However, both approaches necessitate retraining the model, which becomes unfeasible when the model structure and parameters are inaccessible. To address this challenge, we propose FR-GNN, a general framework for GNNs to conduct feature reconstruction. FR-GNN constructs a mapping relationship between the output and input of a well-trained GNN to obtain class representative embeddings and then uses these embeddings to reconstruct the features of labeled nodes. These reconstructed features are then incorporated into the message passing mechanism of GNNs to influence the predictions of unlabeled nodes at test time. Notably, the reconstructed node features can be directly utilized for testing the well-trained model, effectively reducing the distribution shift and leading to improved test performance. This remarkable achievement is attained without any modifications to the model structure or parameters. We provide theoretical guarantees for the effectiveness of our framework. Furthermore, we conduct comprehensive experiments on various public data sets. The experimental results demonstrate the superior performance of FR-GNN in comparison to multiple categories of baseline methods.
Rui Ding 0013, Jielong Yang, Xionghu Zhong, Linbo Xie
IEEE Internet Things J.4
2024 Black-Box Attacks on Graph Neural Networks via White-Box Methods With Performance Guarantees
abstract
Graph adversarial attacks can be classified as either white-box or black-box attacks. White-box attackers typically exhibit better performance because they can exploit the known structure of victim models. However, in practical settings, most attackers generate perturbations under black-box conditions, where the victim model is unknown. A fundamental question is how to leverage a white-box attacker to attack a black-box model. Some current black-box attack approaches employ white-box techniques to attack a surrogate model, resulting in satisfactory outcomes. Nonetheless, such white-box attackers must be meticulously designed and lack theoretical assurances for attack effectiveness. In this paper, we propose a novel framework that utilizes simple white-box techniques to conduct black-box attacks and provides the lower bound for attack performance. Specifically, we first employ a more comprehensive GCN technique named BiasGCN to approximate the victim model, and subsequently, use a simple white-box approach to attack the approximate model. We provide a generalization guarantee for our BiasGCN and employ it to obtain the lower bound on attack performance. Our method is evaluated on various datasets, and the experimental results indicate that our approach surpasses recently proposed baselines.
Jielong Yang, Rui Ding 0013, Xionghu Zhong, Huarong Zhao, Linbo Xie
IEEE Internet Things J.4
2024 Uncertainty-Aware and Class-Balanced Domain Adaptation for Object Detection in Driving Scenes
abstract
This work tackles the cross-domain object detection problem which aims to generalize a pre-trained object detector to different domains (driving scenes) without labels. An uncertainty-aware and class-balanced domain adaptation method is proposed based on two motivations: 1) estimation and exploitation of model uncertainty in a new domain is critical for reliable domain adaptation; and 2) in domain adaptation the distribution alignment of two domains as well as the maintaining of category discriminability are both important. In particular, we compose a Bayesian CNN-based framework for uncertainty estimation in object detection. We propose an algorithm for generating uncertainty-aware pseudo-labels, which are then used in uncertainty-guided self-training and category-aware feature alignment. We further devise a scheme with class-balanced memory banks to address the long-tail distribution problem in category-aware feature alignment. Experiments on multiple cross-domain object detection benchmarks show that our proposed method achieves state-of-the-art performance.
Minjie Cai, Jianaresi Kezierbieke, Xionghu Zhong, Hao Chen 0002
IEEE Trans. Intell. Transp. Syst.3
2024 A Novel Composite Graph Neural Network
abstract
Graph neural networks (GNNs) have achieved great success in many fields due to their powerful capabilities of processing graph-structured data. However, most GNNs can only be applied to scenarios where graphs are known, but real-world data are often noisy or even do not have available graph structures. Recently, graph learning has attracted increasing attention in dealing with these problems. In this article, we develop a novel approach to improving the robustness of the GNNs, called composite GNN. Different from existing methods, our method uses composite graphs (C-graphs) to characterize both sample and feature relations. The C-graph is a unified graph that unifies these two kinds of relations, where edges between samples represent sample similarities, and each sample has a tree-based feature graph to model feature importance and combination preference. By jointly learning multiaspect C-graphs and neural network parameters, our method improves the performance of semisupervised node classification and ensures robustness. We conduct a series of experiments to evaluate the performance of our method and the variants of our method that only learn sample relations or feature relations. Extensive experimental results on nine benchmark datasets demonstrate that our proposed method achieves the best performance on almost all the datasets and is robust to feature noises.
Zhaogeng Liu, Jielong Yang, Xionghu Zhong, Wenwu Wang 0001, Hechang Chen, Yi Chang 0001
IEEE Trans. Neural Networks Learn. Syst.3
2023 Unifying Speech Enhancement and Separation with Gradient Modulation for End-to-End Noise-Robust Speech Separation
abstract
Recent studies in neural network-based monaural speech separation (SS) have achieved a remarkable success thanks to increasing ability of long sequence modeling. However, they would degrade significantly when put under realistic noisy conditions, as the background noise could be mistaken for speaker’s speech and thus interfere with the separated sources. To alleviate this problem, we propose a novel network to unify speech enhancement and separation with gradient modulation to improve noise-robustness. Specifically, we first build a unified network by combining speech enhancement (SE) and separation modules, with multi-task learning for optimization, where SE is supervised by parallel clean mixture to reduce noise for downstream speech separation. Furthermore, in order to avoid suppressing valid speaker information when reducing noise, we propose a gradient modulation (GM) strategy to harmonize the SE and SS tasks from optimization view. Experimental results show that our approach achieves the state-of-the-art on large-scale Libri2Mix- and Libri3Mix-noisy datasets, with SI-SNRi results of 16.0 dB and 15.8 dB respectively. Our code is available at GitHub1.
Chen Chen 0075, Heqing Zou, Xionghu Zhong, Chng Eng Siong
ICASSP4
2023 Blind Estimation of Room Impulse Response from Monaural Reverberant Speech with Segmental Generative Neural Network
Zhiheng Liao, Feifei Xiong, Juan Luo, Minjie Cai, Chng Eng Siong, Jinwei Feng, Xionghu Zhong
INTERSPEECH7
2023 Emotion-Aware Audio-Driven Face Animation via Contrastive Feature Disentanglement
Juan Luo, Xionghu Zhong, Minjie Cai
INTERSPEECH3
2023 GLAE: A graph-learnable auto-encoder for single-cell RNA-seq analysis
Yixiang Shan, Jielong Yang, Xiangtao Li, Xionghu Zhong, Yi Chang 0001
Inf. Sci.4
2023 Audio-Visual Event Localization by Learning Spatial and Semantic Co-Attention
abstract
This work aims to temporally localize events that are both audible and visible in video. Previous methods mainly focused on temporal modeling of events with simple fusion of audio and visual features. In natural scenes, a video records not only the events of interest but also ambient acoustic noise and visual background, resulting in redundant information in the raw audio and visual features. Thus, direct fusion of the two features often causes false localization of the events. In this paper, we propose a co-attention model to exploit the spatial and semantic correlations between the audio and visual features, which helps guide the extraction of discriminative features for better event localization. Our assumption is that in an audio-visual event, shared semantic information between audio and visual features exists and can be extracted by attention learning. Specifically, the proposed co-attention model is composed of a co-spatial attention module and a co-semantic attention module that are used to model the spatial and semantic correlations, respectively. The proposed co-attention model can be applied to various event localization tasks, such as cross-modality localization and multimodal event localization. Experiments on the public audio-visual event (AVE) dataset demonstrate that the proposed method achieves state-of-the-art performance by learning spatial and semantic co-attention.
Xionghu Zhong, Minjie Cai, Hao Chen 0002, Wenwu Wang 0001
IEEE Trans. Multim.2
2021 Overlapped Speech Detection Based on Spectral and Spatial Feature Fusion
Weiguang Chen, Van Tung Pham, Chng Eng Siong, Xionghu Zhong
Interspeech4
2021 Cramér-Rao Lower Bound for DOA Estimation with an Array of Directional Microphones in Reverberant Environments
Weiguang Chen, Xionghu Zhong
Interspeech3
2021 Multiple Acoustic Source Localization in Microphone Array Networks
abstract
The problem of multiple acoustic source localization using observations from a microphone array network is investigated in this article. Multiple source signals are assumed to be window-disjoint-orthogonal (WDO) on the time-frequency (TF) domain and time delay of arrival (TDOA) measurements are extracted at each TF bin. A Bayesian network model is then proposed to jointly assign the measurements to different sources and estimate the acoustic source locations. Considering that the WDO assumption is usually violated under reverberant and noisy environments, we construct a relational network by coding the distance information between the distributed microphone arrays such that adjacent arrays have higher probabilities of observing the same acoustic source, which is able to mitigate the miss detection issues in adverse environments. A Laplace approximate variational inference method is introduced to estimate the hidden variables in the proposed Bayesian network model. Both simulations and real data experiments are performed. The results show that our proposed method is able to achieve better source localization accuracy than existing methods.
Jielong Yang, Xionghu Zhong, Weiguang Chen, Wenwu Wang 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 A Comprehensive Review of Driver Behavior Analysis Utilizing Smartphones
abstract
Human factors are the primary catalyst for traffic accidents. Among different factors, fatigue, distraction, drunkenness, and/or recklessness are the most common types of abnormal driving behavior that leads to an accident. With technological advances, modern smartphones have the capabilities for driving behavior analysis. There has not yet been a comprehensive review on methodologies utilizing only a smartphone for drowsiness detection and abnormal driver behavior detection. In this paper, different methodologies proposed by different authors are discussed. It includes the sensing schemes, detection algorithms, and their corresponding accuracy and limitations. Challenges and possible solutions such as integration of the smartphone behavior classification system with the concept of context-aware, mobile crowdsensing, and active steering control are analyzed. The issue of model training and updating on the smartphone and cloud environment is also included.
Teck Kai Chan, Cheng Siong Chin, Hao Chen 0002, Xionghu Zhong
IEEE Trans. Intell. Transp. Syst.4
2019 LaIF: A Lane-Level Self-Positioning Scheme for Vehicles in GNSS-Denied Environments
abstract
Vehicle self-positioning is of significant importance for intelligent transportation applications. However, accurate positioning (e.g., with lane-level accuracy) is very difficult to obtain due to the lack of measurements with high confidence, especially in an environment without full access to a global navigation satellite system (GNSS). In this paper, a novel information fusion algorithm based on a particle filter is proposed to achieve lane-level tracking accuracy under a GNSS-denied environment. We consider the use of both coarse-scale and fine-scale signal measurements for positioning. Time-of-arrival measurements using the radio frequency signals from known transmitters or roadside units, and acceleration or gyroscope measurements from an inertial measurement unit (IMU) allow us to form a coarse estimate of the vehicle position using an extended Kalman filter. Subsequently, fine-scale measurements, including lane-change detection, radar ranging from the known obstacles (e.g., guardrails), and information from a high-resolution digital map, are incorporated to refine the position estimates. A probabilistic model is introduced to characterize the lane changing behaviors, and a multi-hypothesis model is formulated for the radar range measurements to robustly weigh the particles and refine the tracking results. Moreover, a decision fusion mechanism is proposed to achieve a higher reliability in the lane-change detection as compared to each individual detector using IMU and visual (if available) information. The posterior Cramér-Rao lower bound is also derived to provide a theoretical performance guideline. The performance of the proposed tracking framework is verified by simulations and real measured IMU data in a four-lane highway.
Ramtin Rabiee, Xionghu Zhong, Yong-Sheng Yan 0001, Wee-Peng Tay
IEEE Trans. Intell. Transp. Syst.2
2018 A Dynamic Bayesian Nonparametric Model for Blind Calibration of Sensor Networks
abstract
We consider the problem of blind calibration of a sensor network, where the sensor gains and offsets are estimated from noisy observations of unknown signals. This is in general a nonidentifiable problem, unless restrictive assumptions on the signal subspace or sensor observations are imposed. We show that if each signal observed by the sensors follows a known dynamic model with additive noise, then the sensor gains and offsets are identifiable. We propose a dynamic Bayesian nonparametric model to infer the sensors' gains and offsets. Our model allows different sensor clusters to observe different unknown signals, without knowing the sensor clusters a priori. We develop an offline algorithm using block Gibbs sampling and a linearized forward filtering backward sampling method that estimates the sensor clusters, gains, and offsets jointly. Furthermore, for practical implementation, we also propose an online inference algorithm based on particle filtering and local Markov chain Monte Carlo. Simulations using a synthetic dataset, and experiments on two real datasets suggest that our proposed methods perform better than several other blind calibration methods, including a sparse Bayesian learning approach, and methods that first cluster the sensor observations and then estimate the gains and offsets.
Jielong Yang, Xionghu Zhong, Wee-Peng Tay
IEEE Internet Things J.2
2017 A dynamic Bayesian nonparametric model for blind calibration of sensor networks
abstract
In the sensor network blind calibration problem, the gains and offsets of sensors are estimated from noisy observations of unknown underlying signals. This is in general a non-identifiable problem, unless restrictive assumptions on the signal subspace or sensor observations are imposed. To overcome these assumptions, we propose a dynamic Bayesian nonparametric model. We show that if the unknown underlying signals follow the first-order auto-regressive process, then the sensor gains and offsets are identifiable. Furthermore, our model allows sensors to form clusters, where each cluster observes the same underlying signal. The clusters are however not known a priori, and are learned through the sensor data. We present a block Gibbs sampling inference method based on the forward filtering backward sampling algorithm. Simulation results suggest that our approach can estimate the sensor gains and offsets with good accuracy, and performs better than methods that first perform clustering and then blind calibration.
Jielong Yang, Wee-Peng Tay, Xionghu Zhong
ICASSP3
2016 Virtual Multi-Antenna Array for Estimating the Angle-of-Arrival of a RF Transmitter
abstract
We consider the problem of angle-of-arrival (AoA) estimation of a RF transmitter using a mobile receiver. We develop a method very similar to synthetic aperture radar which is compatible with cellular technology. By considering the successive packets received along the receiver trajectory, we implicitly create a virtual MIMO array, which allows us to utilize conventional MIMO theory for AoA estimation. For this method to work, the first major challenge is the need to separate the phase offset due to receiver movement from the phase offset due to local oscillator (LO) offset. Two approaches are proposed to do this: i) a stop-and-start approach, where the receiver first stands still long enough to estimate the LO offset and then estimates the AoA while moving, and ii) a joint nonlinear estimator where the AoA and LO offset are estimated simultaneously. The second major difficulty is the need to estimate the receiver's relative position with sub-wavelength accuracy. We solve this by using a three-dimensional inertial measurement unit, which provides reasonably good relative position estimates if the measurement period is sufficiently short. Simulations and experimental results based on a software-defined radio platform in an anechoic chamber show the feasibility of the proposed method.
François Quitin, Vivek Govindaraj, Xionghu Zhong, Wee-Peng Tay
VTC Fall3
2016 Linear fitting Kalman filter
abstract
Dynamic estimation in signal processing and target tracking often involves non‐linear models. These non‐linear models are usually linearised through the first‐order Taylor approximation in estimation process. However, the error generated by the first‐order Taylor approximation is not negligible when the non‐linearity of a model is high or the input error is large. This study proposes a new linearisation method through minimising the error between a non‐linear function and its linear approximation. A weighted least squares (WLS) algorithm is developed to estimate a linear fitting (LF) function based on the sigma points of the random variable in non‐linear transformation. A linear fitting Kalman filter (LKF) is developed based on this principle. The accuracy of the LF transform is analysed using the Kullback–Leibler (KL) distance. The results show that the LF transform has less KL distance to the true distribution compared with the first‐order Taylor approximation. To evaluate the estimation performance, simulations are conducted and the results are compared with those of extended Kalman filter (EKF) and unscented Kalman filter (UKF). The results demonstrate that the LKF provides better accuracy than the EKF, and has similar accuracy to the UKF with lower computational cost.
Yuanbo Xiong, Xionghu Zhong
IET Signal Process.2
2015 Robust speech recognition using beamforming with adaptive microphone gains and multichannel noise reduction
abstract
This paper presents a robust speech recognition system using a microphone array for the 3rd CHiME Challenge. A minimum variance distortionless response (MVDR) beamformer with adaptive microphone gains is proposed for robust beamforming. Two microphone gain estimation methods are studied using the speech-dominant time-frequency bins. A multichannel noise reduction (MCNR) postprocessing is also proposed to further reduce the interference in the MVDR processed signal. Experimental results for the ChiME-3 challenge show that both the proposed MVDR beamformer with microphone gains and the MCNR postprocessing improve the speech recognition performance significantly. With the state-of-the-art deep neural network (DNN) based acoustic model, our system achieves a word error rate (WER) of 11.67% on the real test data of the evaluation set.
Shengkui Zhao, Thi Ngoc Tho Nguyen, Xionghu Zhong, Bo Ren 0006, Longbiao Wang, Douglas L. Jones, Chng Eng Siong, Haizhou Li 0001
ASRU5
2015 A learning-based approach to direction of arrival estimation in noisy and reverberant environments
abstract
This paper presents a learning-based approach to the task of direction of arrival estimation (DOA) from microphone array input. Traditional signal processing methods such as the classic least square (LS) method rely on strong assumptions on signal models and accurate estimations of time delay of arrival (TDOA) . They only work well in relatively clean conditions, but suffer from noise and reverberation distortions. In this paper, we propose a learning-based approach that can learn from a large amount of simulated noisy and reverberant microphone array inputs for robust DOA estimation. Specifically, we extract features from the generalised cross correlation (GCC) vectors and use a multilayer perceptron neural network to learn the nonlinear mapping from such features to the DOA. One advantage of the learning based method is that as more and more training data becomes available, the DOA estimation will become more and more accurate. Experimental results on simulated data show that the proposed learning based method produces much better results than the state-of-the-art LS method. The testing results on real data recorded in meeting rooms show improved root-mean-square error (RMSE) compared to the LS method.
Shengkui Zhao, Xionghu Zhong, Douglas L. Jones, Chng Eng Siong, Haizhou Li 0001
ICASSP3
2015 Learning to estimate reverberation time in noisy and reverberant rooms
Shengkui Zhao, Xionghu Zhong, Douglas L. Jones, Chng Eng Siong, Haizhou Li 0001
INTERSPEECH3
2015 A distributed particle filtering approach for multiple acoustic source tracking using an acoustic vector sensor network
Xionghu Zhong, Arash Mohammadi 0001, A. Benjamin Premkumar, Amir Asif
Signal Process.1
2015 Reverberant speech separation with probabilistic time-frequency masking for B-format recordings
Xiaoyi Chen 0002, Wenwu Wang 0001, Yingmin Wang, Xionghu Zhong, Atiyeh Alinaghi
Speech Commun.4
2015 A Time-Frequency Masking Based Random Finite Set Particle Filtering Method for Multiple Acoustic Source Detection and Tracking
abstract
Considering that multiple talkers may appear simultaneously, a time-frequency (TF) masking based random finite set (RFS) particle filtering (PF) method is developed for multiple acoustic source detection and tracking. The time-delay of arrival (TDOA) measurements of multiple sources are extracted by using a time-frequency masking technique, by which each source's TF bins are clustered and separated in a joint gain-ratio and time-delay histogram. Since a joint detection and tracking problem is considered, both source positions and source numbers are time-varying and need to be estimated. The tracker is built within a RFS Bayesian filtering framework. Essentially, an RFS process is used to characterize the source dynamics that include source appearance/dissappearance and motion trajectories. Latent variables are also introduced to indicate source dynamics and measurement-source associations. Subsequently, a Rao-Blackwellization PF technique is employed so that the source position state can be marginalized and only the latent variables are estimated by using the PF. The main advantage of the proposed approach is that hypothesis-pruning is formulated in a full probabilistic sense. The performance of the proposed approach is demonstrated in real speech recordings as well as in simulated room environments.
Xionghu Zhong, James R. Hopgood
IEEE ACM Trans. Audio Speech Lang. Process.1
2014 Audio-visual tracking of a variable number of speakers with a random finite set approach
Volkan Kilic, Xionghu Zhong, Mark Barnard, Wenwu Wang 0001, Josef Kittler
FUSION2
2014 A Bayesian performance bound for time-delay of arrival based acoustic source tracking in a reverberant environment
Xionghu Zhong, Wenwu Wang 0001, Syed M. Naqvi, Chng Eng Siong
FUSION1
2014 Particle filtering for TDOA based acoustic source tracking: Nonconcurrent Multiple Talkers
Xionghu Zhong, James R. Hopgood
Signal Process.1
2014 Multiple wideband source detection and tracking using a distributed acoustic vector sensor array: A random finite set approach
Xionghu Zhong, A. Benjamin Premkumar
Signal Process.1
2013 Decentralized Bayesian Estimation with Quantized Observations: Theoretical Performance Bounds
abstract
The posterior Cramέr Rao lower bound (PCRLB) has recently been proposed as an effective selection criteria for sensor resource management in large, geographically distributed sensor networks. Existing algorithms (in particular the decentralized approaches with no central fusion centre) designed for computing the PCRLB are based on raw observations resulting in significant communication overhead from the sensor nodes to the associated local processing nodes. The paper derives distributive computational techniques for determining the PCRLB for quantized sensor networks configured using decentralized architectures. We refer to the distributed computation of the PCRLB as dPCRLB. The main contribution of the paper is extending the dPCRLB algorithm [1] to quantized observations that leads to significant savings in the communication overhead over its counterparts that use raw observations. In our Monte Carlo simulations, we show that the proposed dPCRLB closely follows the centralized bound based on quantized observations. As expected, there is potential performance loss with quantization as is illustrated by the difference between the dPCRLBs computed using raw and quantized observations. The drop in the estimator's performance is, however, compensated for with an increase in the number of quantization levels associated with the observation quantizer.
Arash Mohammadi 0001, Amir Asif, Xionghu Zhong, A. Benjamin Premkumar
DCOSS3
2013 Acoustic source tracking in a reverberant environment using a pairwise synchronous microphone network
Xionghu Zhong, Arash Mohammadi 0001, Wenwu Wang 0001, A. Benjamin Premkumar, Amir Asif
FUSION1
2012 A random finite set approach for joint detection and tracking of multiple wideband sources using a distributed acoustic vector sensor array
Xionghu Zhong, A. Benjamin Premkumar
FUSION1
2012 A random finite set approach for tracking time-varying number of acoustic sources using a single acoustic vector sensor
abstract
Existing localization approaches developed using acoustic vector sensor (AVS) signals normally assume that the sources are static and the number of sources is known. In this paper, a novel approach is developed to estimate the 2-D direction of arrival (DOA) of an unknown and time-varying number of acoustic sources using a single AVS. A random finite set (RFS) is employed to characterize the randomness of the state process, i.e., the dynamics of source motion and the number of active sources. Also the measurement processes are modeled by RFS since we allow the AVS report undesired events by sending an empty set other than received signals under regular cases. We further employ particle filtering to approximate the posterior DOA distributions. The performance of the proposed approach is demonstrated by simulated experiments.
Xionghu Zhong, A. Benjamin Premkumar, A. S. Madhukumar
ICASSP1
2011 Multi-modality likelihood based particle filtering for 2-D direction of arrival tracking using a single acoustic vector sensor
abstract
The general problem addressed in this paper is tracking the 2-D direction of arrival (DOA) of an acoustic source signal by using a single acoustic vector sensor (AVS). A Bayesian framework and its particle filtering implementation are introduced to adapt to the underwater ambient noise environment, in which both the interference and background noise exist. Several innovations are explored here: 1) a particle filtering based acoustic source tracking algorithm for AVS is developed; and 2) by using a multi-modality likelihood model to model the source detection and false alarm separately, the algorithm is able to alleviate the effect due to noise and interference. Particularly, by employing additional acoustic information, the proposed approach is able to track the 2-D DOA by using a single AVS. The performance of proposed approach is fully investigated under different simulated ambient noisy environments. Experiment results show that the proposed algorithm outperforms the traditional Capon beamforming approach and is able to lock on the 2-D DOA of the source even in a very challenging environment.
Xionghu Zhong, A. Benjamin Premkumar, A. S. Madhukumar, Chiew Tong Lau
ICME1
2008 Nonconcurrent multiple speakers tracking based on extended Kalman particle filter
abstract
Acoustic reverberation introduces multipath components into an audio signal, and therefore changes the source signal statistical properties. This causes problems for source localisation and tracking since reverberation generates spurious peaks in the time delay functions, and makes the subsequent location estimator hard to track the motion trajectory. Previous time delay based tracking methods, such as the extended Kalman filter and the particle filter, are sensitive to reverberation and are unable to follow sharp changes in the source positions. In this paper, the extended Kalman filter and the particle filter are combined to solve this problem. One of the advantages of this approach is that the optimal importance function can be obtained after extended Kalman filtering. Thus, the position samples are distributed in a more accurate area than using a prior importance function. Experiment results show that the proposed algorithm outperforms the sequential importance resampling particle filter by reducing the estimation error and following the switch of speakers quickly under a moderate reverberant environment (reverberation time T60≪ 0.3s).
Xionghu Zhong, James R. Hopgood
ICASSP1