Syed M. Naqvi

dblp:57/7819 · also Syed Mohsen Naqvi · DBLP profile ↗
← Back
66ranked-venue papers
3as first author
21since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 37 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 15 · 6 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 A Frequency-aware Augmentation Network for Mental Disorders Assessment from Audio
abstract
Depression and Attention Deficit Hyperactivity Disorder (ADHD) stand out as the common mental health challenges today. In affective computing, speech signals serve as effective biomarkers for mental disorder assessment. Current research, relying on labor-intensive hand-crafted features or simplistic time-frequency representations, often overlooks critical details by not accounting for the differential impacts of various frequency bands and temporal fluctuations. Therefore, we propose a frequency-aware augmentation network with dynamic convolution for depression and ADHD assessment. In the proposed method, the spectrogram is used as the input feature and adopts a multi-scale convolution to help the network focus on discriminative frequency bands related to mental disorders. A dynamic convolution is also designed to aggregate multiple convolution kernels dynamically based upon their attentions which are input-independent to capture dynamic information. Finally, a feature augmentation block is proposed to enhance the feature representation ability and make full use of the captured information. Experimental results on AVEC 2014 and self-recorded ADHD dataset prove the robustness of our method, an RMSE of 9.23 was attained for estimating depression severity, along with an accuracy of 89.8% in detecting ADHD.
Shuanglin Li, Siyang Song, Rajesh Nair, Syed M. Naqvi
ICASSP4
2025 Efficient Long Speech Sequence Modelling for Time-Domain Depression Level Estimation
abstract
Depression significantly affects emotions, thoughts, and daily activities. Recent research indicates that speech signals contain vital cues about depression, sparking interest in audiobased deep-learning methods for estimating its severity. However, most methods rely on time-frequency representations of speech which have recently been criticized for their limitations due to the loss of information when performing time-frequency projections, e.g. Fourier transform, and Mel-scale transformation. Furthermore, segmenting real-world speech into brief intervals risks losing critical interconnections between recordings. Additionally, such an approach may not adequately reflect real-world scenarios, as individuals with depression often pause and slow down in their conversations and interactions. Building on these observations, we present an efficient method for depression level estimation using long speech signals in the time domain. The proposed method leverages a state space model coupled with the dual-path structure-based long sequence modelling module and temporal external attention module to reconstruct and enhance the detection of depression-related cues hidden in the raw audio waveforms. Experimental results on the AVEC2013 and AVEC2014 datasets show promising results in capturing consequential long-sequence depression cues and demonstrate outstanding performance over the state-of-the-art.
Shuanglin Li, Zhijie Xie, Syed M. Naqvi
ICASSP3
2025 Pose-oriented scene-adaptive matching for abnormal event detection
abstract
For intelligent surveillance systems, abnormal event detection automatically analyses surveillance video sequences and detects abnormal objects or unusual human actions at the frame level. Due to the lack of labelled data, most approaches are semi-supervised based on reconstruction or prediction methods. However, these methods may not generalize well to unseen scene contexts. To address this issue, we present a novel self and mutual scene-adaptive matching method for abnormal event detection. In the framework, we propose synergistic pose estimation and object detection, which effectively integrates human pose and object detection information to improve pose estimation accuracy. Then, the poses are resized to reduce the spatial distance between the source and target domains. The improved pose sequences are further fed into a spatio-temporal graph convolutional network to extract the geometric features. Finally, the features are embedded in a clustering layer to classify action types and compute normality scores. The training data is taken from the training part of common video anomaly detection datasets: UCSD PED1 & PED2, CHUK Avenue, and ShanghaiTech Campus. The proposed framework is evaluated on video sequences with unseen scene contexts in the UCSD PED2 and ShanghaiTech Campus datasets. The detection accuracy and efficiency are also evaluated in detail, and the proposed method for abnormal event detection achieves the highest AUC performance, 84.6%, on the ShanghaiTech Campus dataset and relatively high AUC performance, 96.9% and 74.8%, on UCSD PED2 & PED1 datasets. Compared with other state-of-the-art works, the performance analysis and results confirm the robustness and effectiveness of our proposed framework for cross-scene abnormal event detection.
Yuxing Yang, Leiyu Xie, Zeyu Fu, Jiawei Yan, Syed M. Naqvi
Neurocomputing5
2025 A Multi-Scale Feature Refinement and Dual-Attention Enhanced Dynamic Convolutional Network for Speech-Based Depression and ADHD Assessment
abstract
In the area of affective computing, speech has been identified as a promising biomarker for assessing depression and attention deficit hyperactivity disorder (ADHD). These disorders manifest as abnormalities in speech across various frequency bands and exhibit temporal variations. Most existing work on speech features relies on the magnitude spectrogram, which discards phase information and also does not consider the impact of different frequency bands on depression and ADHD detection. Inspired by these, we propose a novel multi-scale complex feature refinement and dynamic convolution attention-aware network to enhance speech-based assessment of depression and ADHD. Our approach incorporates three key components: multi-scale complex feature refinement (MSFR), dynamic convolutional neural network (Dy-CNN), and dual-attention feature enhancement (DAFE) module. The MSFR module utilizes depth-wise convolutional networks to process both magnitude and phase input, selectively emphasizing frequency bands associated with depression and ADHD. Importantly, the Dy-CNN module employs an attention mechanism to autonomously generate multiple convolution kernels that adapt to input features and capture relevant temporal dynamics linked to depression and ADHD. Additionally, the DAFE module enhances feature representation and detection performance by incorporating channel shuffle attention (CSA) and spatial axial attention (SAA) mechanisms, which leverage both inter- and intra-channel relationships and examine time-frequency characteristics of the feature map. Extensive experiments conducted on four publicly available datasets, i.e., AVEC2013, AVEC2014, E-DAIC, and a self-collected authentic ADHD dataset demonstrated that the proposed method outperforms previous approaches and exhibits superior generalization capabilities across different language settings (i.e., English, German) for speech-based depression and ADHD assessment.
Shuanglin Li, Siyang Song, Syed M. Naqvi
IEEE Trans. Affect. Comput.3
2025 Position and Orientation Aware One-Shot Learning for Medical Action Recognition From Signal Data
abstract
In this article, we propose a position and orientation-aware one-shot learning framework for medical action recognition from signal data. The proposed framework comprises two stages and each stage includes signal-level image generation (SIG), cross-attention (CsA), and dynamic time warping (DTW) modules and the information fusion between the proposed privacy-preserved position and orientation features. The proposed SIG method aims to transform the raw skeleton data into privacy-preserved features for training. The CsA module is developed to guide the network in reducing medical action recognition bias and more focusing on important human body parts for each specific action, aimed at addressing similar medical action related issues. Moreover, the DTW module is employed to minimize temporal mismatching between instances and further improve model performance. Furthermore, the proposed privacy-preserved orientation-level features are utilized to assist the position-level features in both of the two stages for enhancing medical action recognition performance. Extensive experimental results on the widely-used and well-known NTU RGB+D 60, NTU RGB+D 120, and PKU-MMD datasets all demonstrate the effectiveness of the proposed method, which outperforms the other state-of-the-art methods with general dataset partitioning by 2.7%, 6.2% and 4.1%, respectively.
Leiyu Xie, Yuxing Yang, Zeyu Fu, Syed M. Naqvi
IEEE Trans. Multim.4
2024 A Novel Audio-Visual Information Fusion System for Mental Disorders Detection
abstract
Mental disorders are among the foremost contributors to the global healthcare challenge. Research indicates that timely diagnosis and intervention are vital in treating various mental disorders. However, the early somatization symptoms of certain mental disorders may not be immediately evident, often resulting in their oversight and misdiagnosis. Additionally, the traditional diagnosis methods incur high time and cost. Deep learning methods based on fMRI and EEG have improved the efficiency of the mental disorder detection process. However, the cost of the equipment and trained staff are generally huge. Moreover, most systems are only trained for a specific mental disorder and are not general-purpose. Recently, physiological studies have shown that there are some speech and facial-related symptoms in a few mental disorders (e.g., depression and ADHD). In this paper, we focus on the emotional expression features of mental disorders and introduce a multimodal mental disorder diagnosis system based on audio-visual information input. Our proposed system is based on spatial-temporal attention networks and innovative uses a less computationally intensive pre-train audio recognition network to fine-tune the video recognition module for better results. We also apply the unified system for multiple mental disorders (ADHD and depression) for the first time. The proposed system achieves over 80% accuracy on the real multimodal ADHD dataset and achieves state-of-the-art results on the depression dataset AVEC 2014.
Yichun Li, Shuanglin Li, Syed M. Naqvi
FUSION3
2024 Cross-Modal Attention for Multimodal Information Fusion: A Novel Approach to Attention Deficit Hyperactivity Disorder Detection
abstract
This paper presents a novel method for differentiating Attention Deficit Hyperactivity Disorder subjects from control participants by multimodal data fusion, including video observations and questionnaire responses. By exploiting the well known Video Vision Transformer model, we analyse the video modality to identify the complex spatial-temporal information of ADHD symptoms. Simultaneously, a Multi-Layer Perceptron model is applied to evaluate structured questionnaire data by capturing key cognitive and emotional indicators of the ADHD symptoms. To fuse the two modalities, a cross-modal attention mechanism assigns adaptive weights to each feature based on its classification relevance. The targeted weighting significantly refines the proposed model’s decision-making capability by concentrating on the most critical elements of the aggregated information. For training and testing, our novel Multimodal ADHD dataset recorded under the Intelligent Sensing ADHD Trial in collaboration with Cumbria, Northumberland, Tyne and Wear NHS Foundation Trust UK is evaluated. The proposed model, ADViQ-AL achieves a 98.18% classification accuracy, 97.83% sensitivity, and 98.53% specificity in classifying ADHD and control groups.
Christian Nash, Rajesh Nair, Syed M. Naqvi
FUSION3
2024 NCL-DASB: GEO-Located Maritime Surveillance Labeled Dataset and Annotation API
abstract
Due to maritime transportation being the most crucial mode in international trade, maritime traffic safety significantly influences global economic development. Detecting anomalous ship behaviors (DASB) serves as a critical measure to safeguard maritime traffic safety. In recent years, data-driven deep learning technologies have witnessed remarkable advancements, and the introduction of high-quality DASB datasets facilitates the rapid and effective transformation of traditional DASB methods into intelligent ones. In this paper, we initially present a labeled DASB dataset named NCL-DASB, recorded at the Tynemouth port in Newcastle, UK. Subsequently, we propose a standard framework for processing vessel AIS data, enabling the transformation of AIS data into vessel trajectory feature information suitable for deep learning through preprocessing. Finally, we open-source an API for annotating vessel trajectory data in the NCL-DASB dataset, intended for the use of future researchers in their studies.
Leiyu Xie, Federico Angelini, Syed M. Naqvi
FUSION3
2024 Object Detection Oriented Privacy-Preserving Frame-Level Video Anomaly Detection
abstract
With the rapid development of intelligent surveillance, video anomaly detection has become a popular topic in related areas of artificial intelligence. In this work, the main focus is on those applications where the privacy of human targets is concerned, such as outdoor and indoor surveillance and smart living systems. Video frames are the most common recorded and processed information source for human anomaly detection. However, video frames also contain privacy-sensitive information such as facial information and identification of human targets. This paper provides a privacy-preserving anomaly detection framework that introduces image segmentation masks to protect the privacy of the human targets. Meanwhile, object detection is implemented to improve anomaly detection performance by incorporating contextual information. The proposed method uses the ST-AE and CONV-AE models, which were trained and tested on the popular anomaly detection datasets UCSD Ped1 and Ped2. Experiments confirm that when image segmentation masks are applied to preserve human targets' privacy information, the anomaly detection models still achieve good performances with the orientation of object detection.
Jiawei Yan, Yuxing Yang, Syed M. Naqvi
ICASSP3
2023 MEMS Gyroscope and the Ego-Motion Estimation Information Fusion for the Low-Cost Freehand Ultrasound Scanner
abstract
This paper presents the design and implementation of the fusion algorithm for a very low-cost medical ultrasound imaging system. The probe consists of only a single-element piezoelectric transducer which makes it challenging to reconstruct a geometrically correct 2-D ultrasound image. Minimal hardware has been used to reduce the manufacturing cost. Concepts of ego-motion estimation are used to estimate the lateral position of the probe rather than using a position sensor. Whereas an inexpensive microelectromechanical system (MEMS) gyroscope is used to measure the orientation of the probe. The orientation and position information are fused using the fusion algorithm proposed in this paper that gives the real-time 3-D position of the probe. This will enable a geometrically correct, 2-D, B-mode ultrasound image to be constructed from a very simple probe with a single fixed beam that is manually scanned across the skin. Previously collected in-vivo data of the fetus is analysed and 2-D image is reconstructed using the fusion algorithm that corrects the geometry of the fetus’s head.
Ayusha Abbas, Jeffrey A. Neasham, Syed M. Naqvi
BIBM3
2023 Action-Based ADHD Diagnosis in Video
abstract
Attention Deficit Hyperactivity Disorder (ADHD) causes significant impairment in various domains.Early diagnosis of ADHD and treatment could significantly improve the quality of life and functioning.Recently, machine learning methods have improved the accuracy and efficiency of the ADHD diagnosis process.However, the cost of the equipment and trained staff required by the existing methods are generally huge.Therefore, we introduce the video-based frame-level action recognition network to ADHD diagnosis for the first time.We also record a real multi-modal ADHD dataset and extract three action classes from the video modality for ADHD diagnosis.The whole process data have been reported to CNTW-NHS Foundation Trust, which would be reviewed by medical consultants/professionals and will be made public in due course.
Yichun Li, Yuxing Yang, Rajesh Nair, Syed M. Naqvi
ESANN4
2023 One-Shot Medical Action Recognition With A Cross-Attention Mechanism And Dynamic Time Warping
abstract
In this paper, we address the classification of medical actions with only one single sample by developing a novel one-shot learning framework which contains both cross-attention and dynamic time warping (DTW) modules. To be concrete, we firstly transform the raw skeleton sequence into the signal-level image representation. We exploit a metric learning approach, which is the prototypical network for the proposed one-shot learning framework and choose the residual network (ResNet18) as the backbone which is widely used in recent years. Cross-attention is applied for guiding the network to focus on the more important joints from each specific action. The cross-attention mechanism that applies between the support and query set will be adapted for mining and matching the relationships with the human body. Furthermore, a DTW module is introduced to mitigate the temporal information mismatching issue between the actions from the support and query sets. The experimental results on the NTU RGB+D 120 dataset demonstrate the effectiveness of our proposed approach and the improved performance compared to the baseline approach. The code of this work is available at1.
Leiyu Xie, Yuxing Yang, Zeyu Fu, Syed M. Naqvi
ICASSP4
2023 Enhancing ADHD Detection Using Diva Interview-Based Audio Signals and A Two-Stream Network
abstract
Attention deficit hyperactivity disorder (ADHD) is a neurodevelopmental condition that results in altered behaviour in social development and communication patterns. However, due to the dearth of medical psychiatrists globally, the diagnosis of ADHD is frequently delayed. With the burgeoning development of artificial intelligence, it is rational to introduce deep learning to facilitate the ADHD diagnosis. Previous deep learning methods mainly use functional magnetic resonance imaging (fMRI) or Electroencephalography (EEG) signals to detect ADHD, where the data are expensive to acquire, i.e. equipment cost and specialised staff for data collection. Over the past years, speech signals have gained increasing attention owing to their cost-effectiveness in data collection and non-intrusive characteristics. In this work, based on the Diagnostic Interview for ADHD in adults (DIVA), we design a questionnaire and collected the audio data of ADHD patients and normal controls in collaboration with the Cumbria, Northumberland, Tyne and Wear NHS Foundation Trust. Besides, we propose a two-stream model (TSM) to exploit local and global features to assist ADHD detection. Applying the TSM to the collected real ADHD audio data, the performance of the proposed method is promising with an average accuracy of 84.9%.
Shuanglin Li, Yang Sun 0003, Rajesh Nair, Syed M. Naqvi
IPCCC4
2023 Pose-driven human activity anomaly detection in a CCTV-like environment
abstract
Abstract Human activity anomaly detection plays a crucial role in the next generation of surveillance and assisted living systems. Most anomaly detection algorithms are generative models and learn features from raw images. This work shows that popular state‐of‐the‐art autoencoder‐based anomaly detection systems are not capable of effectively detecting human‐posture and object‐positions related anomalies. Therefore, a human pose‐driven and object‐detector‐based deep learning architecture is proposed, which simultaneously leverages human poses and raw RGB data to perform human activity anomaly detection. It is demonstrated that pose‐driven learning overcomes the raw RGB based counterpart limitations in different human activities classification. Extensive validation is provided by using popular datasets. Then, it is demonstrated that with the aid of object detection, the human activities classification can be effectively used in human activity anomaly detection. Moreover, novel challenging datasets, that is, BMbD, M‐BMbD and JBMOPbD, are proposed for single and multi‐target human posture anomaly detection and joint human posture and object position anomaly detection evaluations.
Yuxing Yang, Federico Angelini, Syed M. Naqvi
IET Image Process.3
2023 Abnormal event detection for video surveillance using an enhanced two-stream fusion method
abstract
Abnormal event detection is a critical component of intelligent surveillance systems, focusing on identifying abnormal objects or unusual human behaviours in video sequences. However, conventional methods struggle due to the scarcity of labelled data. Existing solutions typically train on normal data, establish boundaries for regular events, and identify outliers during testing. These approaches are often inadequate as they do not efficiently leverage the geometry and image texture information, and they lack a specific focus on different types of abnormal events. This paper introduces a novel two-stream fusion algorithm for abnormal event detection to address these diverse abnormal events better. We first extract the object, pose, and optical flow features. Then, the object and pose information is combined early on to eliminate occluded pose graphs. The trusted pose graphs are fed into a Spatio-Temporal Graph Convolutional Network (ST-GCN) to detect abnormal behaviours. Simultaneously, we propose a video prediction framework that identifies abnormal frames by measuring the difference between predicted and ground truth frames. Lastly, we execute a decision-level fusion between the classification and prediction streams to achieve the final results. Our results on the UCSD PED1 dataset indicate the enhanced performance of the fusion model for various abnormal events. Furthermore, experimental results on the UCSD PED2 dataset and the ShanghaiTech campus dataset underscore our approach’s effectiveness compared to other related works.
Yuxing Yang, Zeyu Fu, Syed M. Naqvi
Neurocomputing3
2023 U-Shaped Transformer With Frequency-Band Aware Attention for Speech Enhancement
Yi Li 0047, Yang Sun 0003, Wenwu Wang 0001, Syed M. Naqvi
IEEE ACM Trans. Audio Speech Lang. Process.4
2022 Machine Learning and ADHD Mental Health Detection - A Short Survey
Christian Nash, Rajesh Nair, Syed M. Naqvi
FUSION3
2022 Privacy Preserving Multi-class Fall Classification Based on Cascaded Learning And Noisy Labels Handling
Leiyu Xie, Yang Sun 0003, Jonathon A. Chambers, Syed M. Naqvi
FUSION4
2022 A Two-Stream Information Fusion Approach to Abnormal Event Detection in Video
abstract
Human abnormal activity detection for automatic surveillance systems is to detect abnormal objects and human behaviours in videos. In this paper, we propose to explicitly address different kinds of abnormal events by developing a two-stream fusion approach that integrates both geometry and image texture information. To be concrete, we firstly propose to utilize an object detector to divide the abnormal events into two catalogues: abnormal human behaviors and abnormal objects. For the detection of abnormal human behaviours, we exploit a spatial-temporal graph convolutional network (ST-GCN) which considers both spatial and temporal domains to capture the geometrical features from human pose graphs. The extracted geometric feature embeddings are further adapted with a clustering step to cluster the temporal graphs and output normality scores. For the detection of abnormal objects, the obtained from the object detector are reused to assist with generating normality scores of possible anomalies. Finally, a late fusion is performed to integrate normality scores from both screams for final decision. The experimental results on the datasets of UCSD PED2 and ShanghaiTech Campus demonstrate the effectiveness of our proposed approach and the improved performance compared to other state-of-the-art approaches.
Yuxing Yang, Zeyu Fu, Syed M. Naqvi
ICASSP3
2021 Video Anomaly Detection for Surveillance Based on Effective Frame Area
Yuxing Yang, Yang Xian, Zeyu Fu, Syed M. Naqvi
FUSION4
2021 Convolutional fusion network for monaural speech enhancement
Yang Xian, Yang Sun 0003, Wenwu Wang 0001, Syed M. Naqvi
Neural Networks4
2020 Image Segmentation Based Privacy-Preserving Human Action Recognition for Anomaly Detection
abstract
Human Action Recognition and Anomaly Detection significantly improved automatic video analysis, assisted living, and video-based surveillance. The focus of this work is on those applications where privacy protection is required, such as surveillance and assisted living. RGB video data is the most common source for human action recognition. However, RGB data also contains privacy-related data, such as the identity of the target. In this paper, we prove that human action recognition accuracy mostly depends on contextual data, rather than on privacy-related data. Therefore, human target data can be occluded by using an image segmentation mask. The proposed method achieves almost similar accuracy in comparison with the privacy case and provides the platform for privacy-preserving anomaly detection. Simulations are performed on the two popular datasets for human action recognition, i.e. UCF101 and HMDB51.
Jiawei Yan, Federico Angelini, Syed M. Naqvi
ICASSP3
2020 Single-channel dereverberation and denoising based on lower band trained SA-LSTMs
abstract
The supervised single‐channel speech enhancement presents one mixture recording at the input of the neural network and updates network parameters in order to generate an output as the reconstructed speech signal. However, current neural networks‐based single‐channel speech enhancement methods are not able to fully utilise pertinence with the specific frequency range of speech signals with limited computational complexity. In this study, the authors studied the power spectral density of mixtures with human speech and noise interferences. Based on the theory that the speech signal distributes at the lower band, they proposed a method to train signal approximation (SA) based neural networks with the lower frequency band of the speech mixture to improve the performance. To realise the lower band approach for single‐channel speech enhancement, the method uses a long short‐term memory (LSTM) block to exploit short‐time Fourier transform of the desired frequency range. Furthermore, in order to improve the speech enhancement performance within reverberant room environments, the dereverberation mask and the enhanced ratio mask are exploited as the training targets of two LSTM blocks, respectively. The detailed evaluations confirm that the proposed method outperforms the state‐of‐the‐art methods.
Yi Li 0047, Yang Sun 0003, Syed M. Naqvi
IET Signal Process.3
2020 2D Pose-Based Real-Time Human Action Recognition With Occlusion-Handling
abstract
Human Action Recognition (HAR) for CCTV-oriented applications is still a challenging problem. Real-world scenarios HAR implementations is difficult because of the gap between Deep Learning data requirements and what the CCTV-based frameworks can offer in terms of data recording equipments. We propose to reduce this gap by exploiting human poses provided by the OpenPose, which has been already proven to be an effective detector in CCTV-like recordings for tracking applications. Therefore, in this work, we first propose ActionXPose: a novel 2D pose-based approach for pose-level HAR. ActionXPose extracts low- and high-level features from body poses which are provided to a Long Short-Term Memory Neural Network and a 1D Convolutional Neural Network for the classification. We also provide a new dataset, named ISLD, for realistic pose-level HAR in a CCTV-like environment, recorded in the Intelligent Sensing Lab. ActionXPose is extensively tested on ISLD under multiple experimental settings, e.g. Dataset Augmentation and Cross-Dataset setting, as well as revising other existing datasets for HAR. ActionXPose achieves state-of-the-art performance in terms of accuracy, very high robustness to occlusions and missing data, and promising results for practical implementation in real-world applications.
Federico Angelini, Zeyu Fu, Yang Long 0001, Ling Shao 0001, Syed M. Naqvi
IEEE Trans. Multim.5
2019 Joint RGB-Pose Based Human Action Recognition for Anomaly Detection Applications
Federico Angelini, Syed M. Naqvi
FUSION2
2019 Enhanced Detection Reliability for Human Tracking Based Video Analytics
Zeyu Fu, Syed M. Naqvi
FUSION3
2019 Ego-motion Estimation for Low-cost Freehand Ultrasound Scanner
abstract
This paper describes work towards a very low-cost medical ultrasound imaging system using concepts of ego-motion estimation. This will enable a B-mode image to be constructed from a very simple probe with a single fixed beam which is either manually scanned across the skin (linear) or rotated against the skin (polar). In the case of a linear scan, an algorithm has been proposed which measures the decorrelation between successive scanlines to estimate probe velocity. With the aid of an Unscented Kalman Filter (UKF), this is used to reconstruct a geometrically correct 2D image of a resolution phantom which has well-defined image patterns for precise measurement of geometric accuracy. In the case of a polar scan, angular data obtained from a low cost MEMS gyroscope is used to reconstruct the image. Examples are also shown of data collected on human subjects, which show promising results for clinical diagnostics.
Ayusha Abbas, Jeffrey A. Neasham, Syed M. Naqvi
ICASSP3
2019 Privacy-preserving Online Human Behaviour Anomaly Detection Based on Body Movements and Objects Positions
abstract
Human behaviour anomaly detection is crucial for modern artifi-cial intelligence systems. However, privacy protection plays a great role in the realization. In this paper, an online privacy-preserving anomaly detector is presented. The proposed method is able to discriminate on human subject body movements, postures and interactions with the surrounding objects, preserving subject privacy in all the tuning, training and testing stages. ActionXPose, Single Shot MultiBox Detector and Support Vector Machine are exploited for the proposed semi-supervised anomaly detector. The method successfully detects abnormal human behaviours, including unexpected body movements and misplaced objects. A new dataset ISLD-A is also proposed1, providing suitable benchmark for performance evaluation2.
Federico Angelini, Jiawei Yan, Syed M. Naqvi
ICASSP3
2019 Enhanced pooling method for convolutional neural networks based on optimal search theory
abstract
To obtain the best pooling effect and higher accuracy in image recognition, an improved method based on optimal search theory for the pooling layer of convolutional neural networks (CNNs) is proposed. The purpose is to solve the problems of the traditional pooling method, namely that it is too simplistic and it is difficult to extract effective features. The basic principle and network structure of CNN are introduced in the study. A new optimum‐pooling method is proposed, and the authors study how to obtain the maximum probability to detect the target function under the constrained condition. Comparison experiments of different pooling methods are performed on three widely used datasets: LFW, CIFAR‐10, and ImageNet. The experimental results show that the proposed method has the characteristics of more effective feature extraction and wide adaptability, and leads to higher accuracy and lower error rate in image recognition.
Zeyu Fu, Syed M. Naqvi, Jonathon A. Chambers
IET Image Process.4
2019 Two-Stage Monaural Source Separation in Reverberant Room Environments Using Deep Neural Networks
abstract
Deep neural networks (DNNs) have been used for dereverberation and separation in the monaural source separation problem. However, the performance of current state-of-the-art methods is limited, particularly when applied in highly reverberant room environments. In this paper, we propose a two-stage approach with two DNN-based methods to address this problem. In the first stage, the dereverberation of the speech mixture is achieved with the proposed dereverberation mask (DM). In the second stage, the dereverberant speech mixture is separated with the ideal ratio mask (IRM). To realize this two-stage approach, in the first DNN-based method, the DM is integrated with the IRM to generate the enhanced time-frequency (T-F) mask, namely the ideal enhanced mask (IEM), as the training target for the single DNN. In the second DNN-based method, the DM and the IRM are predicted with two individual DNNs. The IEEE and the TIMIT corpora with real room impulse responses and noise from the NOISEX dataset are used to generate speech mixtures for evaluations. The proposed methods outperform the state-of-the-art specifically in highly reverberant room environments.
Yang Sun 0003, Wenwu Wang 0001, Jonathon A. Chambers, Syed M. Naqvi
IEEE ACM Trans. Audio Speech Lang. Process.4
2019 Multi-Level Cooperative Fusion of GM-PHD Filters for Online Multiple Human Tracking
abstract
In this paper, we propose a multi-level cooperative fusion approach to address the online multiple human tracking problem in a Gaussian mixture probability hypothesis density (GM-PHD) filter framework. The proposed fusion approach consists essentially of three steps. First, we integrate two human detectors with different characteristics (full-body and body-parts), and investigate their complementary benefits for tracking multiple targets. For each detector domain, we then propose a novel discriminative correlation matching model, and fuse it with spatio-temporal information to address ambiguous identity association in the GM-PHD filter. Finally, we develop a robust fusion center with virtual and real zones to make a global decision based on preliminary candidate targets generated by each detector. This center also mitigates the sensitivity of missed detections in the generalized covariance intersection fusion process, thereby improving the fusion performance and tracking consistency. Experiments on the MOTChallenge Benchmark demonstrate that the proposed method achieves improved performance over other state-of-the-art RFS-based tracking methods.
Zeyu Fu, Federico Angelini, Jonathon A. Chambers, Syed M. Naqvi
IEEE Trans. Multim.4
2018 Collaborative Detector Fusion of Data-Driven PHD Filter for Online Multiple Human Tracking
abstract
The use of multiple data sources (measurements) has been recently demonstrated to improve the accuracy and reliability of a tracking system as it is capable of providing redundancy in different aspects, and also eliminating interferences of individual sources. This paper focuses on addressing the multiple human tracking problem from a multi-detector approach. This approach integrates two detectors with different characteristics (full-body and body-parts) to perform robust collaborative fusion based on data-driven Gaussian Mixture Probability Hypothesis Density (GM-PHD) filters. To leverage the maximum strengths from multiple detectors, we propose a robust fusion center at the track level, which manages to perform Generalized Intersection Covariance (GCI) fusions for survival and birth tracks independently, and also eliminates false tracks caused by a cluttered environment. Moreover, an identity reassignment mechanism is also developed to address the identity mismatching problem in the target birth process, so as to enhance the fusion performance and track consistency. Experimental results on two challenging benchmark video sequences confirm the effectiveness of the proposed approach.
Zeyu Fu, Syed M. Naqvi, Jonathon A. Chambers
FUSION2
2018 3D-Hog Embedding Frameworks for Single and Multi-Viewpoints Action Recognition Based on Human Silhouettes
abstract
Given the high demand for automated systems for human action recognition, great efforts have been undertaken in recent decades to progress the field. In this paper, we present frameworks for single and multi-viewpoints action recognition based on Space-Time Volume (STV) of human silhouettes and 3D-Histogram of Oriented Gradient (3D-HOG) embedding. We exploit fast-computational approaches involving Principal Component Analysis (PCA) over the local feature spaces for compactly describing actions as combinations of local gestures and L2-Regularized Logistic Regression (L2-RLR) for learning the action model from local features. Outperforming results on Weizmann and i3DPost datasets confirm efficacy of the proposed approaches as compared to the baseline method and other works, in terms of accuracy and robustness to appearance changes.
Federico Angelini, Zeyu Fu, Sergio A. Velastin, Jonathon A. Chambers, Syed M. Naqvi
ICASSP5
2018 GM-PHD Filter Based Online Multiple Human Tracking Using Deep Discriminative Correlation Matching
abstract
In this paper, we propose deep discriminative correlation matching within the Gaussian Mixture Probability Hypothesis Density (GM-PHD) filter for online multiple human tracking. In this matching scheme, we mainly exploit the Convolutional Neural Network (CNN) based Discriminative Correlation Filter (DCF) as a target-specific classifier to discriminate the desired target from background and remaining targets. DCFs are learned through the extracted features obtained from the outputs of the last convolutional layers which are capable to encode the target appearances with better discriminativity and robustness to appearance changes. Moreover, we present a hybrid likelihood function that fuses the spatio-temporal relation and correlation matching score to collaboratively enhance the PHD association step. Experimental results on the MOT17 Challenge benchmark [1] confirm the improved performance of our proposed method as compared with other state-of-the-art techniques.
Zeyu Fu, Federico Angelini, Syed M. Naqvi, Jonathon A. Chambers
ICASSP3
2018 Geometric Information Based Monaural Speech Separation Using Deep Neural Network
abstract
The performance of deep neural network (DNN) based monaural speech separation methods is limited in reverberant and noisy room environments. In this paper, we propose a new DNN training target which incorporates geometric information describing the target speaker and microphone to improve the performance in reverberant and noisy room environments. The experiments are based on the IEEE corpus and the NOISEX database and real impulse responses (RIRs). The objective evaluations, short-time objective intelligibility (STOI) and perceptual evaluation of speech quality (PESQ) confirm the efficiency of the proposed direct path ratio mask (DRM).
Yang Xian, Yang Sun 0003, Jonathon A. Chambers, Syed M. Naqvi
ICASSP4
2018 Robust selection of the degrees of freedom in the Student's t distribution through Multiple Model Adaptive Estimation
Qian Li 0001, Yueyang Ben, Jiubin Tan, Syed M. Naqvi, Jonathon A. Chambers
Signal Process.4
2017 Particle PHD filter based multi-target tracking using discriminative group-structured dictionary learning
abstract
Structured sparse representation has been recently found to achieve better efficiency and robustness in exploiting the target appearance model in tracking systems with both holistic and local information. Therefore, to better simultaneously discriminate multi-targets from their background, we propose a novel video-based multi-target tracking system that combines the particle probability hypothesis density (PHD) filter with discriminative group-structured dictionary learning. The discriminative dictionary with group structure learned by the hierarchical K-means clustering algorithm implicitly associates the dictionary atoms with the group labels, simultaneously enforcing the target candidates from the same group (class) to share the same structured sparsity pattern. Furthermore, we propose a new joint likelihood calculation by relating the discriminative sparse codes with the maximum voting technique to enhance the particle PHD updating step. Experimental results on two publicly available benchmark video sequences confirm the improved performance of our proposed method over other state-of-the-art techniques in video-based multi-target tracking.
Zeyu Fu, Pengming Feng, Syed M. Naqvi, Jonathon A. Chambers
ICASSP3
2017 Underdetermined source separation using time-frequency masks and an adaptive combined Gaussian-Student's t probabilistic model
abstract
Time-frequency (T-F) masking algorithms are focused at separating multiple sound sources from binaural reverberant speech mixtures. The statistical modelling of binaural cues i.e. interaural phase difference (IPD) and interaural level difference (ILD) is a significant aspect of such algorithms. In this paper, a Gaussian-Student's t distribution combined mixture model is exploited for robust binaural speech separation. The weights of the distribution components are calculated adaptively with the energy of the speech mixtures. The expectation maximization (EM) algorithm is applied to calculate the parameters of the distributions. The speech signals from the TIMIT database are convolved with the real binaural room impulse responses (BRIRs) from two datasets for the evaluation of the proposed method. The objective performance measure signal to distortion ratio (SDR) confirms the improvement and robustness of the proposed method.
Yang Sun 0003, Waqas Rafique, Jonathon A. Chambers, Syed M. Naqvi
ICASSP4
2017 Social Force Model-Based MCMC-OCSVM Particle PHD Filter for Multiple Human Tracking
abstract
Video-based multiple human tracking often involves several challenges, including target number variation, object occlusions, and noise corruption in sensor measurements. In this paper, we propose a novel method to address these challenges based on probability hypothesis density (PHD) filtering with a Markov chain Monte Carlo (MCMC) implementation. More specifically, a novel social force model (SFM) for describing the interaction between the targets is used to calculate the likelihood within the MCMC resampling step in the prediction step of the PHD filter, and a one class support vector machine (OCSVM) is then used in the update step to mitigate the noise in the measurements, where the SVM is trained with features from both color and oriented gradient histograms. The proposed method is evaluated and compared with state-of-the-art techniques using sequences from the CAVIAR, TUD, and PETS2009 datasets based on the mean Euclidean tracking error on each frame, the optimal subpattern assignment metric, and the multiple object tracking precision metric. The results show improved performance of the proposed method over the baseline algorithms, including the traditional particle PHD filtering method, the traditional SFM-based particle filtering method, multi-Bernoulli filtering, and an online-learning-based tracking method.
Pengming Feng, Wenwu Wang 0001, Satnam Singh Dlay, Syed M. Naqvi, Jonathon A. Chambers
IEEE Trans. Multim.4
2016 A robust Student's t based cubature filter
Yulong Huang 0003, Yonggang Zhang 0001, Ning Li 0001, Syed M. Naqvi, Jonathon A. Chambers
FUSION4
2016 A robust and efficient system identification method for a state-space model with heavy-tailed process and measurement noises
Yulong Huang 0003, Yonggang Zhang 0001, Ning Li 0001, Syed M. Naqvi, Jonathon A. Chambers
FUSION4
2016 Social force model aided robust particle PHD filter for multiple human tracking
abstract
In this paper, we propose a novel robust multiple human tracking approach based upon processing a video signal by utilizing a social force model to enhance the particle probability hypothesis density (PHD) filter. In traditional dynamic models, the states of targets are only predicted by their own history; however, in multiple human tracking, the information from interaction between targets and the intentions of each target can be employed to obtain more robust prediction. Furthermore, such information can mitigate the problems of collision and occlusion. The cardinality of variable number of targets can also be estimated by using the PHD filter, hence improving the overall accuracy of the multiple human tracker. In this work, a background subtraction step has also been employed to identify the new born targets and provide the measurement set for the PHD filter. To evaluate tracking performance, sequences from both the CAVIAR and PETS2009 datasets are employed for evaluation, which shows clear improvement of the proposed method over the conventional particle PHD filter.
Pengming Feng, Wenwu Wang 0001, Syed M. Naqvi, Satnam Singh Dlay, Jonathon A. Chambers
ICASSP3
2016 Adaptive Retrodiction Particle PHD Filter for Multiple Human Tracking
abstract
The probability hypothesis density (PHD) filter is well known for addressing the problem of multiple human tracking for a variable number of targets, and the sequential Monte Carlo implementation of the PHD filter, known as the particle PHD filter, can give state estimates with nonlinear and non-Gaussian models. Recently, Mahler et al. have introduced a PHD smoother to gain more accurate estimates for both target states and number. However, as highlighted by Psiaki in the context of a backward-smoothing extended Kalman filter, with a nonlinear state evolution model the approximation error in the backward filtering requires careful consideration. Psiaki suggests that to minimize the aggregated least-squares error over a batch of data. We instead use the term retrodiction PHD filter to describe the backward filtering algorithm in recognition of the approximation error proposed in the original PHD smoother, and we propose an adaptive recursion step to improve the approximation accuracy. This step combines forward and backward processing through the measurement set and thereby mitigates the problems with the original PHD smoother when the target number changes significantly and the targets appear and disappear randomly. Simulation results show the improved performance of the proposed algorithm and its capability in handling a variable number of targets.
Pengming Feng, Wenwu Wang 0001, Syed M. Naqvi, Jonathon A. Chambers
IEEE Signal Process. Lett.3
2015 Real-time independent vector analysis with Student's t source prior for convolutive speech mixtures
abstract
A common approach to blind source separation is to use independent component analysis. However when dealing with realistic convolutive audio and speech mixtures, processing in the frequency domain at each frequency bin is required. As a result this introduces the permutation problem, inherent in independent component analysis, across the frequency bins. Independent vector analysis directly addresses this issue by modeling the dependencies between frequency bins, namely making use of a source prior. An alternative source prior for real-time (online) natural gradient independent vector analysis is proposed. A Student's t probability density function is known to be more suited for speech sources, due to its heavier tails, and is incorporated into a real-time version of natural gradient independent vector analysis. In addition, the importance of the degrees of freedom parameter within the Student's t distribution is highlighted. The final algorithm is realized as a real-time embedded application on a floating point Texas Instruments digital signal processor platform, where simulated recordings from a reverberant room are used for testing. Results are shown to be better than with the original (super-Gaussian) source prior.
Jack Harris, Bertrand Rivet, Syed M. Naqvi, Jonathon A. Chambers, Christian Jutten
ICASSP3
2015 IVA algorithms using a multivariate Student's t source prior for speech source separation in real room environments
abstract
The independent vector analysis (IVA) algorithm employs a multivariate source prior to retain the dependency between different frequency bins of each source and thereby avoids the permutation problem that is inherent to blind source separation (BSS). In this paper, a multivariate Student's t distribution is adopted as the source prior, which because of its heavy tail nature can better model the large amplitude information in the frequency bins. Therefore it can improve the separation performance and the convergence speed of the IVA and fast version of the IVA (FastIVA) algorithms as compared with the IVA algorithm based on another multivariate super Gaussian source prior. Separation performance with real binaural room impulse responses (BRIRs) is evaluated by detailed simulation studies when using the different source priors, and the experimental results confirm that the IVA and the FastIVA with the proposed multivariate Student's t source prior can consistently achieve improved and faster separation performance.
Waqas Rafique, Syed M. Naqvi, Philip J. B. Jackson, Jonathon A. Chambers
ICASSP2
2015 Variational EM for clustering interaural phase cues in MESSL for blind source separation of speech
abstract
The model-based expectation maximization source separation and localization (MESSL) technique is a probabilistic time-frequency masking algorithm that achieves underdetermined blind source separation of speech sources. Using only two-channel recordings, MESSL clusters spectrogram points based on their interaural spatial cues. Gaussian mixture models (GMMs) are assumed for the interaural cues and their corresponding parameters are determined by maximum likelihood estimation (MLE) via the expectation maximization (EM) framework. However, the presence of singularities and over-fitting are major drawbacks of MLE. In this paper, we investigate variational Bayesian (VB) inference for clustering spectrogram points based particularly on their interaural phase difference (IPD) cues. Variational inference overcomes the difficulties associated with the likelihood optimization and improves the separation especially when the sources are in close proximity. Simulation studies based on speech mixtures formed from the TIMIT database confirm the advantage of the proposed approach in terms of signal to distortion ratio (SDR).
Zeinab Zohny, Syed M. Naqvi, Jonathon A. Chambers
ICASSP2
2014 A Bayesian performance bound for time-delay of arrival based acoustic source tracking in a reverberant environment
Xionghu Zhong, Wenwu Wang 0001, Syed M. Naqvi, Chng Eng Siong
FUSION3
2014 Multi-target tracking by using particle filtering and a social force model
Ata ur-Rehman, Syed M. Naqvi, Lyudmila Mihaylova, Jonathon A. Chambers
FUSION2
2014 Independent vector analysis with a generalized multivariate Gaussian source prior for frequency domain blind source separation
Yanfeng Liang, Jack Harris, Syed M. Naqvi, Gaojie Chen 0001, Jonathon A. Chambers
Signal Process.3
2014 Robust Multi-Speaker Tracking via Dictionary Learning and Identity Modeling
abstract
We investigate the problem of visual tracking of multiple human speakers in an office environment. In particular, we propose novel solutions to the following challenges: (1) robust and computationally efficient modeling and classification of the changing appearance of the speakers in a variety of different lighting conditions and camera resolutions; (2) dealing with full or partial occlusions when multiple speakers cross or come into very close proximity; (3) automatic initialization of the trackers, or re-initialization when the trackers have lost lock caused by e.g. the limited camera views. First, we develop new algorithms for appearance modeling of the moving speakers based on dictionary learning (DL), using an off-line training process. In the tracking phase, the histograms (coding coefficients) of the image patches derived from the learned dictionaries are used to generate the likelihood functions based on Support Vector Machine (SVM) classification. This likelihood function is then used in the measurement step of the classical particle filtering (PF) algorithm. To improve the computational efficiency of generating the histograms, a soft voting technique based on approximate Locality-constrained Soft Assignment (LcSA) is proposed to reduce the number of dictionary atoms (codewords) used for histogram encoding. Second, an adaptive identity model is proposed to track multiple speakers whilst dealing with occlusions. This model is updated online using Maximum a Posteriori (MAP) adaptation, where we control the adaptation rate using the spatial relationship between the subjects. Third, to enable automatic initialization of the visual trackers, we exploit audio information, the Direction of Arrival (DOA) angle, derived from microphone array recordings. Such information provides, a priori, the number of speakers and constrains the search space for the speaker's faces. The proposed system is tested on a number of sequences from three publicly available and challenging data corpora (AV16.3, EPFL pedestrian data set and CLEAR) with up to five moving subjects.
Mark Barnard, Piotr Koniusz, Wenwu Wang 0001, Josef Kittler, Syed M. Naqvi, Jonathon A. Chambers
IEEE Trans. Multim.5
2013 Audio-visual face detection for tracking in a meeting room environment
Mark Barnard, Wenwu Wang 0001, Josef Kittler, Syed M. Naqvi, Jonathon A. Chambers
FUSION4
2013 Clustering and a joint probabilistic data association filter for dealing with occlusions in multi-target tracking
Ata ur-Rehman, Syed M. Naqvi, Lyudmila Mihaylova, Jonathon A. Chambers
FUSION2
2013 A new cascaded spectral subtraction approach for binaural speech dereverberation and its application in source separation
abstract
In this work we propose a new binaural spectral subtraction method for the suppression of late reverberation. The proposed approach is a cascade of three stages. The first two stages exploit distinct observations to model and suppress the late reverberation by deriving a gain function. The musical noise artifacts generated due to the processing at each stage are compensated by smoothing the spectral magnitudes of the weighting gains. The third stage linearly combines the gains obtained from the first two stages and further enhances the binaural signals. The binaural gains, obtained by independently processing the left and right channel signals are combined using a new method. Experiments on real data are performed in two contexts: dereverberation-only and joint dereverberation and source separation. Objective results verify the suitability of the proposed cascaded approach in both the contexts.
Muhammad Salman Khan 0001, Syed M. Naqvi, Jonathon A. Chambers
ICASSP2
2013 Independent vector analysis with a multivariate generalized gaussian source prior for frequency domain blind source separation
abstract
The independent vector analysis (IVA) algorithm can theoretically avoid the permutation problem in frequency domain blind source separation by using a multivariate source prior to retain the dependency between different frequency bins of each source. In this paper, a new multivariate generalized Gaussian distribution is adopted as the source prior which can exploit fourth order inter-frequency correlation, and therefore better preserve the dependency between different frequency bins to achieve an improved separation performance as compared with the original IVA algorithm. Separation performances are compared by simulation studies when using different source priors, and the experimental results confirm that IVA with the new source prior can consistently achieve improved separation performance.
Yanfeng Liang, Syed M. Naqvi, Jonathon A. Chambers
ICASSP2
2013 Variational Bayesian and belief propagation based data association for multi-target tracking
abstract
A novel two stage data association technique for multi-target tracking is proposed which assigns multiple measurements to a target to mitigate information loss. At the first stage a variational Bayesian (VB) clustering technique is used which groups the measurements automatically into a determined number of clusters. In the second stage a belief propagation (BP) based cluster to target association method is proposed to assign multiple clusters to a target. This is achieved by exploiting the inter-cluster dependency information. The proposed technique is suitable to accommodate non-rigid targets such as humans. Both location and features of clusters are used to re-identify the targets when they emerge from occlusions. The proposed technique is compared with state of the art method due to Laet et al. and evaluations are presented on a real data set.
Ata ur-Rehman, Syed M. Naqvi, Lyudmila Mihaylova, Jonathon A. Chambers
ICASSP2
2013 Video-Aided Model-Based Source Separation in Real Reverberant Rooms
abstract
Source separation algorithms that utilize only audio data can perform poorly if multiple sources or reverberation are present. In this paper we therefore propose a video-aided model-based source separation algorithm for a two-channel reverberant recording in which the sources are assumed static. By exploiting cues from video, we first localize individual speech sources in the enclosure and then estimate their directions. The interaural spatial cues, the interaural phase difference and the interaural level difference, as well as the mixing vectors are probabilistically modeled. The models make use of the source direction information and are evaluated at discrete time-frequency points. The model parameters are refined with the well-known expectation-maximization (EM) algorithm. The algorithm outputs time-frequency masks that are used to reconstruct the individual sources. Simulation results show that by utilizing the visual modality the proposed algorithm can produce better time-frequency masks thereby giving improved source estimates. We provide experimental results to test the proposed algorithm in different scenarios and provide comparisons with both other audio-only and audio-visual algorithms and achieve improved performance both on synthetic and real data. We also include dereverberation based pre-processing in our algorithm in order to suppress the late reverberant components from the observed stereo mixture and further enhance the overall output of the algorithm. This advantage makes our algorithm a suitable candidate for use in under-determined highly reverberant settings where the performance of other audio-only and audio-visual methods is limited.
Muhammad Salman Khan 0001, Syed M. Naqvi, Ata ur-Rehman, Wenwu Wang 0001, Jonathon A. Chambers
IEEE Trans. Speech Audio Process.2
2013 An Online One Class Support Vector Machine-Based Person-Specific Fall Detection System for Monitoring an Elderly Individual in a Room Environment
abstract
In this paper, we propose a novel computer vision-based fall detection system for monitoring an elderly person in a home care, assistive living application. Initially, a single camera covering the full view of the room environment is used for the video recording of an elderly person's daily activities for a certain time period. The recorded video is then manually segmented into short video clips containing normal postures, which are used to compose the normal dataset. We use the codebook background subtraction technique to extract the human body silhouettes from the video clips in the normal dataset and information from ellipse fitting and shape description, together with position information, is used to provide features to describe the extracted posture silhouettes. The features are collected and an online one class support vector machine (OCSVM) method is applied to find the region in feature space to distinguish normal daily postures and abnormal postures such as falls. The resultant OCSVM model can also be updated by using the online scheme to adapt to new emerging normal postures and certain rules are added to reduce false alarm rate and thereby improve fall detection performance. From the comprehensive experimental evaluations on datasets for 12 people, we confirm that our proposed person-specific fall detection system can achieve excellent fall detection performance with 100% fall detection rate and only 3% false detection rate with the optimally tuned parameters. This work is a semiunsupervised fall detection system from a system perspective because although an unsupervised-type algorithm (OCSVM) is applied, human intervention is needed for segmenting and selecting of video clips containing normal postures. As such, our research represents a step toward a complete unsupervised fall detection system.
Miao Yu 0001, Yuanzhang Yu, Adel Rhuma, Syed M. Naqvi, Liang Wang 0001, Jonathon A. Chambers
IEEE J. Biomed. Health Informatics4
2012 A dictionary learning approach to tracking
abstract
The problem of tracking people using multiple cameras is of much current interest as a means of providing cues for audio-visual blind source separation in dynamic environments. Here we investigate the use of one of the current state-of-the-art techniques in object recognition combined with one of the most popular methods of modelling object motion, particle filters, for tracking people. The dictionary learning or Bag-of-Words approach to object recognition has proved to be very effective in recent years, as shown in a number of large comparisons such as the PASCAL Visual Object recognition Challenge (VOC). In this paper we use this proven object recognition method within the framework of a particle filter. This provides a more accurate and robust tracking of people in a multiple camera environment. We also demonstrate that the dictionary learning approach can provide a principled method for the fusion of multiple features.
Mark Barnard, Wenwu Wang 0001, Josef Kittler, Syed M. Naqvi, Jonathon A. Chambers
ICASSP4
2012 Multimodal (audio-visual) source separation exploiting multi-speaker tracking, robust beamforming and time-frequency masking
abstract
A novel multimodal source separation approach is proposed for physically moving and stationary sources which exploits a circular microphone array, multiple video cameras, robust spatial beamforming and time-frequency masking. The challenge of separating moving sources, including higher reverberation time (RT) even for physically stationary sources, is that the mixing filters are time varying; as such the unmixing filters should also be time varying but these are difficult to determine from only audio measurements. Therefore in the proposed approach, visual modality is used to facilitate the separation for both stationary and moving sources. The movement of the sources is detected by a three-dimensional tracker based on a Markov Chain Monte Carlo particle filter. The audio separation is performed by a robust least squares frequency invariant data-independent beamformer. The uncertainties in source localisation and direction of arrival information obtained from the 3D video-based tracker are controlled by using a convex optimisation approach in the beamformer design. In the final stage, the separated audio sources are further enhanced by applying a binary time-frequency masking technique in the cepstral domain. Experimental results show that using the visual modality, the proposed algorithm cannot only achieve performance better than conventional frequency-domain source separations algorithms, but also provide acceptable separation performance for moving sources.
Syed M. Naqvi, Wenwu Wang 0001, Muhammad Salman Khan 0001, Mark Barnard, Jonathon A. Chambers
IET Signal Process.1
2012 A Posture Recognition-Based Fall Detection System for Monitoring an Elderly Person in a Smart Home Environment
abstract
We propose a novel computer vision based fall detection system for monitoring an elderly person in a home care application. Background subtraction is applied to extract the foreground human body and the result is improved by using certain post-processing. Information from ellipse fitting and a projection histogram along the axes of the ellipse are used as the features for distinguishing different postures of the human. These features are then fed into a directed acyclic graph support vector machine (DAGSVM) for posture classification, the result of which is then combined with derived floor information to detect a fall. From a dataset of 15 people, we show that our fall detection system can achieve a high fall detection rate (97.08%) and a very low false detection rate (0.8%) in a simulated home environment.
Miao Yu 0001, Adel Rhuma, Syed M. Naqvi, Liang Wang 0001, Jonathon A. Chambers
IEEE Trans. Inf. Technol. Biomed.3
2011 Multimodal blind source separation for moving sources based on robust beamforming
abstract
An improvement in the multimodal approach to the problem of blind source separation (BSS) of moving sources is proposed. The challenge of BSS for moving sources is that the mixing filters are time varying. Thus the unmixing filters should also be time varying, which are difficult to calculate from only statistical information avail able from limited number of audio samples. Therefore, in the pro posed approach a robust least square frequency invariant data independent (RLSFIDI) beamformer is implemented to perform real time speech enhancement and provide separation of the moving sources. Direction of arrival information of the sources is obtained from the visual 3-D tracker based on a Markov Chain Monte Carlo particle filter (MCMC-PF). The uncertainties in source localization and direction of arrival information are controlled by using convex optimization approach in the beamformer design. This provides robust ness with a wider main lobe for source of interest (SOI) and wider attenuation pattern to block the interference. In addition, white noise gain (WNG) constraint is used to control the beamformer sensitivity. Experimental results show that by utilizing RLSFIDI beamformer a significant improvement in BSS performance for moving sources is achieved in a low reverberant environment.
Syed M. Naqvi, Miao Yu 0001, Jonathon A. Chambers
ICASSP1
2011 Fall detection in a smart room by using a fuzzy one class support vector machine and imperfect training data
abstract
In this paper, we propose an efficient and robust fall detection system by using a fuzzy one class support vector machine based on video in formation. Two cameras are used to capture the video frames from which the features are extracted. A fuzzy one class support vector machine (FOCSVM) is used to distinguish falling from other activities, such as walking, sitting, standing, bending or lying. Compared with the traditional one class support vector machine, the FOCSVM can obtain a more accurate and tight decision boundary under a training dataset with outliers. From real video sequences, the success of the method is confirmed with less non-fall samples being misclassified as falls by the classifier under an imperfect training dataset.
Miao Yu 0001, Syed M. Naqvi, Adel Rhuma, Jonathon A. Chambers
ICASSP2
2010 A robust fall detection system for the elderly in a Smart Room
abstract
In this paper, we propose a novel and robust fall detection system by using a density method for modeling a fall event as a function of certain video feature.3-D head velocity and human shape information are extracted as feature and three types of density model, single Gaussian, mixture of Gaussians and Parzen window method, are constructed for modeling the density of fall with respect to the extracted video feature. Falls are then detected according to the corresponding obtained density model and the success of the method is confirmed on real video sequences.
Miao Yu 0001, Syed M. Naqvi, Jonathon A. Chambers
ICASSP2
2009 Multimodal blind source separation for moving sources
abstract
A novel multimodal approach is proposed to solve the problem of blind source separation (BSS) of moving sources. The challenge of BSS for moving sources is that the mixing filters are time varying, thus the unmixing filters should also be time varying, which are difficult to track in real time. In the proposed approach, the visual modality is utilized to facilitate the separation for both stationary and moving sources. The movement of the sources is detected by a 3-D tracker based on particle filtering. The full BSS solution is formed by integrating a frequency domain blind source separation algorithm and beamforming: if the sources are identified as stationary, a frequency domain BSS algorithm is implemented with an initialization derived from the visual information. Once the sources are moving, a beamforming algorithm is used to perform real time speech enhancement and provide separation of the sources. Experimental results show that by utilizing the visual modality, the proposed algorithm can not only improve the performance of the BSS algorithm and mitigate the permutation problem for stationary sources, but also provide a good BSS performance for moving sources in a low reverberant environment.
Syed M. Naqvi, Yonggang Zhang 0001, Jonathon A. Chambers
ICASSP1
2008 Blind source extraction of heart sound signals from lung sound recordings exploiting periodicity of the heart sound
abstract
A novel approach for separating heart sound signals (HSSs) from lung sound recordings is presented. The approach is based on blind source extraction (BSE) with second-order statistics (SOS), which exploits the quasi-periodicity of the HSSs. The method is evaluated on both synthetic periodic signals of known period mixed with temporally white Gaussian noise (WGN) as well as on real quasi periodic HSSs mixed with lung sound signals (LSSs). Qualitative evaluation involving comparison of the power spectral densities (PSDs) of the extracted signals, by the proposed method and by the JADE algorithm, and that of the original signal is performed for the case of real data. Separation results confirm the utility of the proposed approach, although departure from strict periodicity may impact performance.
Thato Tsalaile, Syed M. Naqvi, Kianoush Nazarpour, Saeid Sanei, Jonathon A. Chambers
ICASSP2
2007 A Geometrically Constrained Multimodal Approach for Convolutive Blind Source Separation
abstract
A novel constrained multimodal approach for convolutive blind source separation is presented which incorporates video information related to geometrical position of both the speakers and the microphones, and the directionality of the speakers into the separation algorithm. The separation is performed in the frequency domain and the constraints are incorporated through a penalty function-based formulation. The separation results show a considerable improvement over traditional frequency domain convolutive BSS systems such as that developed by Parra and Spence. Importantly, the inherent permutation problem in the frequency domain BSS is potentially solved.
Saeid Sanei, Syed M. Naqvi, Jonathon A. Chambers, Yulia Hicks
ICASSP (3)2