EDBT 2026 Demo / reviewers in the wild / expert
Simon Denman
dblp:66/2758
· DBLP profile ↗
105ranked-venue papers
11as first author
34since 2021 · last 2026
0000-0002-0983-5480ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 58 · 8 first-author · 12 since 2021Artificial intelligence and machine learning · 47 · 3 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 6 since 2021Security and privacy · 5 · 3 since 2021Databases, data management, data science and information retrieval · 3Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PDV: Prompt Directional Vectors for Zero-shot Composed Image RetrievalabstractZero-shot Composed Image Retrieval (ZS-CIR) enables image search using a reference image and a text prompt without requiring specialized text-image composition networks trained on large-scale paired data. However, current ZS-CIR approaches suffer from three critical limitations in their reliance on composed text embeddings: static query embedding representations, insufficient utilization of image embeddings, and suboptimal performance when fusing text and image embeddings. To address these challenges, we introduce the Prompt Directional Vector (PDV), a simple yet effective training-free enhancement that captures semantic modifications induced by user prompts. PDV enables three key improvements: (1) Dynamic composed text embeddings where prompt adjustments are controllable via a scaling factor, (2) composed image embeddings through semantic transfer from text prompts to image features, and (3) weighted fusion of composed text and image embeddings that enhances retrieval by balancing visual and semantic similarity. Our approach serves as a plug-and-play enhancement for existing ZS-CIR methods with minimal computational overhead. Extensive experiments across multiple benchmarks demonstrate that PDV consistently improves retrieval performance when integrated with state-of-the-art ZS-CIR approaches, particularly for methods that generate accurate compositional embeddings. The code will be released upon publication. Osman Tursun, Sinan Kalkan, Simon Denman, Clinton Fookes |
WACV | 3 |
| 2026 | Contrastive context distillation for skeleton based early action predictionabstractEarly action prediction (EAP) requires the inference of actions from partially observed sequences, which typically contain only the initial movements of an action. Compared to using complete sequences that show an action being fully executed, EAP lacks discriminative information and increased ambiguity as different actions can contain very similar initial movements. To alleviate this problem, recent methods try to distill features learned from action recognition (AR) models, which leads to sub-optimal results due to the lack of generalizability of features across AR and EAP tasks. In a different line of work, recent studies have proposed learning discriminative class-specific features, leveraging the traditional contrastive learning approach, where samples are selected and calibrated for learning discriminative features from hard-to-classify samples. This paper proposes a novel, dynamic, and context-aware framework for EAP by combining the merits of both knowledge distillation and contrastive learning. Particularly our method distills salient discriminative context from the complete action sequence to drive the EAP, with the help of a novel dynamic contrastive learning scheme. Extensive evaluations over three public datasets demonstrate state-of-the-art performance for Early Action Prediction. Chinthaka Ranasingha, Tharindu Fernando, Harshala Gammulle, Simon Denman, Sridha Sridharan, Clinton Fookes |
Expert Syst. Appl. | 4 |
| 2025 | HOTFormerLoc: Hierarchical Octree Transformer for Versatile Lidar Place Recognition Across Ground and Aerial ViewsabstractWe present HOTFormerLoc, a novel and versatile Hierarchical Octree-based TransFormer, for large-scale 3D place recognition in both ground-to-ground and ground-to-aerial scenarios across urban and forest environments. We propose an octree-based multi-scale attention mechanism that captures spatial and semantic features across granularities. To address the variable density of point distributions from spinning lidar, we present cylindrical octree attention windows to reflect the underlying distribution during attention. We introduce relay tokens to enable efficient global-local interactions and multi-scale representation learning at reduced computational cost. Our pyramid attentional pooling then synthesises a robust global descriptor for end-to-end place recognition in challenging environments. In addition, we introduce CS-WildPlaces, a novel 3D cross-source dataset featuring point cloud data from aerial and ground lidar scans captured in dense forests. Point clouds in CS-Wild-Places contain representational gaps and distinctive attributes such as varying point densities and noise patterns, making it a challenging benchmark for cross-view localisation in the wild. HOTFormerLoc achieves a top-1 average recall improvement of 5.5% – 11.5% on the CS-Wild-Places benchmark. Furthermore, it consistently outperforms SOTA 3D place recognition methods, with an average performance gain of 4.9% on well-established urban and forest datasets. The code and CS-Wild-Places benchmark is available at https://csirorobotics.github.io/HOTFormerLoc. Ethan Griffiths, Maryam Haghighat, Simon Denman, Clinton Fookes, Milad Ramezani |
CVPR | 3 |
| 2025 | RadDet: A Wideband Dataset for Real-Time Radar Spectrum DetectionabstractReal-time detection of radar signals in a wideband radio frequency spectrum is a critical situational assessment function in electronic warfare. Compute-efficient detection models have shown great promise in recent years, providing an opportunity to tackle the spectrum detection problem. However, progress in radar spectrum detection is limited by the scarcity of publicly available wideband radar signal datasets accompanied by corresponding annotations. To address this challenge, we introduce a novel and challenging dataset for radar detection (RadDet), comprising a large corpus of radar signals occupying a wideband spectrum across diverse radar density environments and signal-to-noise ratio (SNR) settings. RadDet contains 40,000 frames, each generated from 1 million in-phase and quadrature (I/Q) samples across a 500 MHz frequency band. RadDet includes 11 classes of radar signals across 6 different SNR settings, 2 radar density environments, and 3 different time-frequency resolutions, with corresponding time-frequency and class annotations. We evaluate the performance of state-of-the-art real-time detection models on RadDet and a modified radar classification dataset from NIST (NIST-CBRS) to establish a novel benchmark for wideband radar spectrum detection. Zi Huang, Simon Denman, Akila Pemasiri, Terrence Martin, Clinton Fookes |
ICASSP | 2 |
| 2025 | MTL-DFM: Multi-Task Learning and Diffusion Model for ISAC SystemsabstractDeep learning (DL) has emerged as a key enabler for unlocking the potential of integrated sensing and communication (ISAC). Despite recent progress, current DL methods primarily handle sensing and communication as independent tasks, overlooking potential performance enhancement through a joint approach. Moreover, existing methods rely on fully annotated data for training, which is often challenging to obtain, especially in multi-task scenarios where labeled data may be scarce or only exist for a subset of tasks. Motivated by these shortcomings, this paper proposes a novel scheme, MTL-DFM, to enable simultaneous sensing and communication with partially labeled training data, which leverages multi-task learning (MTL) and a diffusion model (DFM). In particular, we introduce an initial feature extraction module (IFEM) to jointly capture shared information across tasks and explore inherent cross-task connections for enhanced feature extraction. Next, we design a signal denoising with incomplete labeling (SDIL) module to effectively remove noise from extracted information and construct comprehensive feature representations for all tasks with partially labeled datasets, which is difficult for conventional DL methods. Simulation results verify the superior performance offered by MTL-DFM over prior state-of-the-art methods. Qingqing Cheng, Zhenguo Shi, Simon Denman, Clinton Fookes, Jinhong Yuan, Derrick Wing Kwan Ng |
ICC | 3 |
| 2025 | Neural Network Based Current Harmonic Predictions Using PCC Voltage Measurements in Distribution NetworksabstractConventional current harmonic prediction techniques exhibit significant limitations at a system level in distribution networks due to multivariable, complex, and time-varying models under variable operating conditions. This paper proposes a novel neural network-based technique for predicting low-order current harmonics using voltage harmonics at the point of common coupling, a practical and accurate approach. The proposed model employs voltage harmonic data at the point of common coupling, which are easily measurable in practice, to predict the corresponding amplitude of current harmonics generated by nonlinear loads such as grid-connected inverters. The results demonstrate the model's high accuracy in predicting low-order current harmonics, even under the influence of grid background voltage harmonics, variations in operational power, and grid impedance. The proposed method outperforms other models, such as support vector regression, multiple linear regression, and gaussian process regression, in terms of accuracy and consistency. This work offers a practical and reliable approach for predicting low-order current harmonics at a system level (e.g., distribution networks), contributing to the stable and efficient operation of electrical grids maintaining power quality standards. Rhns Jayathissa, Firuz Zare, Simon Denman, Amir Taghvaie, Maryam Haghighat, Dinesh Kumar 0005 |
IECON | 3 |
| 2025 | Online 6DoF Global Localisation in Forests using Semantically-Guided Re-Localisation and Cross-View Factor-Graph OptimisationabstractThis paper presents FGLoc6D, a novel approach for robust global localisation and online 6DoF pose estimation of ground robots in forest environments by leveraging deep semantically-guided re-localisation and cross-view factor graph optimisation. The proposed method addresses the challenges of aligning aerial and ground data for pose estimation, which is crucial for accurate point-to-point navigation in GPS-degraded environments. By integrating information from both perspectives into a factor graph framework, our approach effectively estimates the robot’s global position and orientation. Additionally, we enhance the repeatability of deep-learned keypoints for metric localisation in forests by incorporating a semantically-guided regression loss. This loss encourages greater attention to wooden structures, e.g., tree trunks, which serve as stable and distinguishable features, thereby improving the consistency of keypoints and increasing the success rate of global registration, a process we refer to as re-localisation. The re-localisation module along with the factor-graph structure, populated by odometry and ground-to-aerial factors over time, allows global localisation under dense canopies. We validate the performance of our method through extensive experiments in three forest scenarios, demonstrating its global localisation capability and superiority over alternative state-of-the-art in terms of accuracy and robustness in these challenging environments. Experimental results show that our proposed method can achieve drift-free localisation with bounded positioning errors, ensuring reliable and safe robot navigation through dense forests. Lucas Carvalho de Lima, Ethan Griffiths, Maryam Haghighat, Simon Denman, Clinton Fookes, Paulo Vinicius Koerich Borges, Michael Brünig, Milad Ramezani |
IROS | 4 |
| 2025 | Remembering What is Important: A Factorised Multi-Head Retrieval and Auxiliary Memory Stabilisation Scheme for Human Motion PredictionabstractHumans exhibit complex motions that vary depending on the activity they are performing, the interactions they engage in, as well as subject-specific preferences. Therefore, forecasting a human's future pose based on the history of his or her previous motion is a challenging task. This paper presents an innovative auxiliary-memory-powered deep neural network framework to improve the modelling of historical knowledge. Specifically, we disentangle subject-specific, action-specific, and other auxiliary information from the observed pose sequences and utilise these factorised features to query the memory. A novel Multi-Head knowledge retrieval scheme leverages these factorised feature embeddings to perform multiple querying operations over the historical observations captured within the auxiliary memory. Moreover, we propose a dynamic masking strategy to make this feature disentanglement process adaptive. Two novel loss functions are introduced to encourage diversity within the auxiliary memory, while ensuring the stability of the memory content such that it can locate and store salient information that aids the long-term prediction of future motion, irrespective of any data imbalances or the diversity of the input data distribution. Extensive experiments conducted on two public benchmarks, Human3.6M and CMU-Mocap, demonstrate that these design choices collectively allow the proposed approach to outperform the current state-of-the-art methods by significant margins: 17% on the Human3.6M dataset and 9% on the CMU-Mocap dataset. Tharindu Fernando, Harshala Gammulle, Sridha Sridharan, Simon Denman, Clinton Fookes |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Multi-Stage Learning for Radar Pulse Activity SegmentationabstractRadio signal recognition is a crucial function in electronic warfare. Precise identification and localisation of radar pulse activities are required by electronic warfare systems to produce effective countermeasures. Despite the importance of these tasks, deep learning-based radar pulse activity recognition methods have remained largely underexplored. While deep learning for radar modulation recognition has been explored previously, classification tasks are generally limited to short and non-interleaved IQ signals, limiting their applicability to military applications. To address this gap, we introduce an end-to-end multi-stage learning approach to detect and localise pulse activities of interleaved radar signals across an extended time horizon. We propose a simple, yet highly effective multi-stage architecture for incrementally predicting fine-grained segmentation masks that localise radar pulse activities across multiple channels. We demonstrate the performance of our approach against several reference models on a novel radar dataset, while also providing a first-of-its-kind benchmark for radar pulse activity segmentation. Zi Huang, Akila Pemasiri, Simon Denman, Clinton Fookes, Terrence Martin |
ICASSP | 3 |
| 2024 | Deep cross-domain transfer for emotion recognition via joint learningabstractAbstract Deep learning has been applied to achieve significant progress in emotion recognition from multimedia data. Despite such substantial progress, existing approaches are hindered by insufficient training data, leading to weak generalisation under mismatched conditions. To address these challenges, we propose a learning strategy which jointly transfers emotional knowledge learnt from rich datasets to source-poor datasets. Our method is also able to learn cross-domain features, leading to improved recognition performance. To demonstrate the robustness of the proposed learning strategy, we conducted extensive experiments on several benchmark datasets including eNTERFACE, SAVEE, EMODB, and RAVDESS. Experimental results show that the proposed method surpassed existing transfer learning schemes by a significant margin. Dung Nguyen 0001, Duc Thanh Nguyen, Sridha Sridharan, Mohamed Almorsy, Simon Denman, Son N. Tran, Clinton Fookes |
Multim. Tools Appl. | 5 |
| 2023 | Dual Memory Fusion for Multimodal Speech Emotion Recognition
Darshana Prisayad, Tharindu Fernando, Sridha Sridharan, Simon Denman, Clinton Fookes |
INTERSPEECH | 4 |
| 2023 | Privacy-Preserving in-bed pose monitoring: A fusion and reconstruction study
Thisun Dayarathna, Thamidu Muthukumarana, Yasiru Rathnayaka, Simon Denman, Chathura de Silva, Akila Pemasiri, David Ahmedt-Aristizabal |
Expert Syst. Appl. | 4 |
| 2023 | Meta-transfer learning for emotion recognitionabstractAbstract Deep learning has been widely adopted in automatic emotion recognition and has lead to significant progress in the field. However, due to insufficient training data, pre-trained models are limited in their generalisation ability, leading to poor performance on novel test sets. To mitigate this challenge, transfer learning performed by fine-tuning pr-etrained models on novel domains has been applied. However, the fine-tuned knowledge may overwrite and/or discard important knowledge learnt in pre-trained models. In this paper, we address this issue by proposing a PathNet-based meta-transfer learning method that is able to (i) transfer emotional knowledge learnt from one visual/audio emotion domain to another domain and (ii) transfer emotional knowledge learnt from multiple audio emotion domains to one another to improve overall emotion recognition accuracy. To show the robustness of our proposed method, extensive experiments on facial expression-based emotion recognition and speech emotion recognition are carried out on three bench-marking data sets: SAVEE, EMODB, and eNTERFACE. Experimental results show that our proposed method achieves superior performance compared with existing transfer learning methods. Dung Nguyen 0001, Duc Thanh Nguyen, Sridha Sridharan, Simon Denman, Thanh Thi Nguyen 0001, David Dean, Clinton Fookes |
Neural Comput. Appl. | 4 |
| 2023 | Multi-stage stacked temporal convolution neural networks (MS-S-TCNs) for biosignal segmentation and anomaly localizationabstractIn the computer vision domain, temporal convolution networks (TCN) have gained traction due to their lightweight, robust architectures for sequence-to-sequence prediction tasks. With that insight, in this study, we propose a novel deep learning architecture for biosignal segmentation and anomaly localization based on TCNs , named the multi-stage stacked TCN, which employs multiple TCN modules with varying dilation factors. More precisely, for each stage, our architecture uses TCN modules with multiple dilation factors, and we use convolution-based fusion to combine predictions returned from each stage. Furthermore, aiming smoothed predictions, we introduce a novel loss function based on the first-order derivative. To demonstrate the robustness of our architecture, we evaluate our model on five different tasks related to three 1D biosignal modalities (heart sounds, lung sounds and electrocardiogram). Our proposed framework achieves state-of-the-art performance for all tasks, significantly outperforming the respective state-of-the-art models having F1 score gains up to ≈ 9 %. Furthermore, the framework demonstrates competitive performance gains compared to traditional multi-stage TCN models with similar configurations yielding F1 score gains up to ≈ 5 %. Our model is also interpretable. Using neural conductance, we demonstrate the effectiveness of having TCNs with varying dilation factors. Our visualizations show that the model benefits from feature maps captured at multiple dilation factors, and the information is effectively propagated through the network such that the final stage produces the most accurate result. Theekshana Dissanayake, Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
Pattern Recognit. | 3 |
| 2023 | Pose-driven attention-guided image generation for person re-Identification
Amena Khatun, Simon Denman, Sridha Sridharan, Clinton Fookes |
Pattern Recognit. | 2 |
| 2023 | Toward On-Board Panoptic Segmentation of Multispectral Satellite ImagesabstractWith tremendous advancements in low-power embedded computing devices and remote sensing instruments, the traditional satellite image processing pipeline which includes an expensive data transfer step prior to processing data on the ground is being replaced by on- board processing of captured data. This paradigm shift enables critical and time-sensitive intelligence to be acquired in a timely manner on- board the satellite itself. However, at present, the on- board processing of multispectral satellite images is limited to classification and segmentation tasks. Extending this processing to the next logical level, we take the first step toward on- board panoptic segmentation of multispectral satellite images and evaluate the applicability of state-of-the-art panoptic segmentation models to an on- board setting. Panoptic segmentation offers major economic and environmental insights, ranging from yield estimation from agricultural lands to intelligence for complex military applications. Nevertheless, the on- board intelligence extraction poses several challenges due to the loss of temporal observations and the need to generate predictions from a single sample. To address this challenge, we propose a multimodal teacher network with a cross modality attention-based fusion strategy to improve segmentation accuracy by exploiting data from multiple modes. We also propose an online knowledge distillation framework to transfer the knowledge learned by this multimodal teacher network to a unimodal student, which receives only a single frame input, and is more appropriate for an on- board environment. We benchmark our approach against existing state-of-the-art panoptic segmentation models using the PASTIS multispectral panoptic segmentation dataset considering an on- board processing setting. Our evaluations demonstrate a substantial 10.7%, 11.9%, and 10.6% increase in segmentation quality (SQ), recognition quality (RQ), and panoptic quality (PQ) metrics compared to the existing state-of-the-art model when it is evaluated in an on- board processing setting. Tharindu Fernando, Clinton Fookes, Harshala Gammulle, Simon Denman, Sridha Sridharan |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Generalized Generative Deep Learning Models for Biosignal Synthesis and Modality TransferabstractGenerative Adversarial Networks (GANs) are a revolutionary innovation in machine learning that enables the generation of artificial data. Artificial data synthesis is valuable especially in the medical field where it is difficult to collect and annotate real data due to privacy issues, limited access to experts, and cost. While adversarial training has led to significant breakthroughs in the computer vision field, biomedical research has not yet fully exploited the capabilities of generative models for data generation, and for more complex tasks such as biosignal modality transfer. We present a broad analysis on adversarial learning on biosignal data. Our study is the first in the machine learning community to focus on synthesizing 1D biosignal data using adversarial models. We consider three types of deep generative adversarial networks: a classical GAN, an adversarial AE, and a modality transfer GAN; individually designed for biosignal synthesis and modality transfer purposes. We evaluate these methods on multiple datasets for different biosignal modalites, including phonocardiogram (PCG), electrocardiogram (ECG), vectorcardiogram and 12-lead electrocardiogram. We follow subject-independent evaluation protocols, by evaluating the proposed models' performance on completely unseen data to demonstrate generalizability. We achieve superior results in generating biosignals, specifically in conditional generation, by synthesizing realistic samples while preserving domain-relevant characteristics. We also demonstrate insightful results in biosignal modality transfer that can generate expanded representations from fewer input-leads, ultimately making the clinical monitoring setting more convenient for the patient. Furthermore our longer duration ECGs generated, maintain clear ECG rhythmic regions, which has been proven using ad-hoc segmentation models. Theekshana Dissanayake, Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
IEEE J. Biomed. Health Informatics | 3 |
| 2022 | SESS: Saliency Enhancing with Scaling and Sliding
Osman Tursun, Simon Denman, Sridha Sridharan, Clinton Fookes |
ECCV (12) | 2 |
| 2022 | Detecting Heart Failure Through Voice Analysis using Self-Supervised Mode-Based Memory FusionabstractCongestive Heart Failure (CHF) is a progressive disease that affects millions of people worldwide, severely impacting their quality of life. Missed detection of CHF and its progression affects life expectancy, thus it is critical to develop applications to continuously monitor CHF symptoms and disease progression in a patient-centric and cost-effective manner. This paper focuses on a novel non-invasive technique to identify CHF using patients' speech traits. Pulmonary congestion and breathlessness is the most common symptom of heart failure and one of the major contributors to hospitalisation. Since pulmonary congestion results in impairment of a patient's voice, we propose a novel, non invasive method for monitoring CHF through analysis of the patient's speech. We also introduce a new balanced dataset, containing voice recordings from both healthy participants and participants diagnosed with CHF, which contains voice alterations reflective of CHF status. We propose a novel deep machine learning architecture based on mode driven memory fusion for CHF recognition from audio recordings of subject's speech. We have achieved 90% accuracy under a subject-independent evaluation setting, highlighting the applicability of such methods for tele-health and home monitoring applications. Darshana Priyasad, Andi Partovi, Sridha Sridharan, Maryam Kashefpoor, Tharindu Fernando, Simon Denman, Clinton Fookes, David Kaye |
INTERSPEECH | 6 |
| 2022 | Learning test-time augmentation for content-based image retrieval
Osman Tursun, Simon Denman, Sridha Sridharan, Clinton Fookes |
Comput. Vis. Image Underst. | 2 |
| 2022 | Affect recognition from scalp-EEG using channel-wise encoder networks coupled with geometric deep learning and multi-channel feature fusion
Darshana Priyasad, Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
Knowl. Based Syst. | 3 |
| 2022 | Split 'n' merge net: A dynamic masking network for multi-task attention
Tharindu Fernando, Sridha Sridharan, Simon Denman, Clinton Fookes |
Pattern Recognit. | 3 |
| 2022 | An efficient framework for zero-shot sketch-based image retrieval
Osman Tursun, Simon Denman, Sridha Sridharan, Ethan Goan, Clinton Fookes |
Pattern Recognit. | 2 |
| 2022 | Channel Graph Regularized Correlation Filters for Visual Object TrackingabstractCorrelation Filters (CF) are a popular choice for visual object tracking due to their efficiency in the frequency domain. Convolutional and hand-crafted features are jointly used when learning a filter, however, these features are not uniformly important when tracking a target. Given this observation, spatial and temporal regularization and attention models have been investigated. However, these models do not consider the interaction between different feature channels. As a result, dissimilar weights are assigned to similar feature channels. To address this issue, we propose a channel attention model and study two different regularization methods for attention. We investigate the application of channel regularization to emphasize important feature channels; and graph regularization which increases the likelihood of similar feature channels obtaining similar weights. The proposed formulation can be efficiently solved via the alternating direction method of multipliers. We first show the advantages of using the proposed channel regularization by demonstrating its performance when applied to two existing CF trackers. This is followed by analyzing the effect of using the proposed channel-graph regularization for CF based tracking. The evaluation is performed on publicly available tracking datasets: OTB100, TC128, VOT-2017, VOT-2019, LaSOT, UAV123, and GOT-10k. Evaluation over multiple challenges and a comparative analysis with existing top-ranked trackers shows that our formulation improves the discriminative power of the learned CF, preventing tracker drift during challenging scenarios. Arjun Tyagi, A. Venkata Subramanyam, Simon Denman, Sridha Sridharan, Clinton Fookes |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Component-Based Attention for Large-Scale Trademark RetrievalabstractThe need for large-scale trademark retrieval (TR) systems has significantly increased to combat the rise in international trademark infringement. Unfortunately, the ranking accuracy of current approaches using either hand-crafted or pre-trained deep convolution neural network (DCNN) features is inadequate for large-scale deployments. We show in this paper that the ranking accuracy of TR systems can be significantly improved by incorporating hard and soft attention mechanisms, which direct attention to critical information such as figurative elements and reduce the attention given to distracting and uninformative elements such as text and background. Our proposed approach achieves state-of-the-art results on a challenging large-scale trademark dataset. Osman Tursun, Simon Denman, Sabesan Sivipalan, Sridha Sridharan, Clinton Fookes, Sandra Mau |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2022 | Geometric Deep Learning for Subject Independent Epileptic Seizure Prediction Using Scalp EEG SignalsabstractRecently, researchers in the biomedical community have introduced deep learning-based epileptic seizure prediction models using electroencephalograms (EEGs) that can anticipate an epileptic seizure by differentiating between the pre-ictal and interictal stages of the subject's brain. Despite having the appearance of a typical anomaly detection task, this problem is complicated by subject-specific characteristics in EEG data. Therefore, studies that investigate seizure prediction widely employ subject-specific models. However, this approach is not suitable in situations where a target subject has limited (or no) data for training. Subject-independent models can address this issue by learning to predict seizures from multiple subjects, and therefore are of greater value in practice. In this study, we propose a subject-independent seizure predictor using Geometric Deep Learning (GDL). In the first stage of our GDL-based method we use graphs derived from physical connections in the EEG grid. We subsequently seek to synthesize subject-specific graphs using deep learning. The models proposed in both stages achieve state-of-the-art performance using a one-hour early seizure prediction window on two benchmark datasets (CHB-MIT-EEG: 95.38% with 23 subjects and Siena-EEG: 96.05% with 15 subjects). To the best of our knowledge, this is the first study that proposes synthesizing subject-specific graphs for seizure prediction. Furthermore, through model interpretation we outline how this method can potentially contribute towards Scalp EEG-based seizure localization. Theekshana Dissanayake, Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
IEEE J. Biomed. Health Informatics | 3 |
| 2022 | Robust and Interpretable Temporal Convolution Network for Event Detection in Lung Sound RecordingsabstractOBJECTIVE: This paper proposes a novel framework for lung sound event detection, segmenting continuous lung sound recordings into discrete events and performing recognition of each event. METHODS: We propose the use of a multi-branch TCN architecture and exploit a novel fusion strategy to combine the resultant features from these branches. This not only allows the network to retain the most salient information across different temporal granularities and disregards irrelevant information, but also allows our network to process recordings of arbitrary length. RESULTS: The proposed method is evaluated on multiple public and in-house benchmarks, containing irregular and noisy recordings of the respiratory auscultation process for the identification of auscultation events including inhalation, crackles, and rhonchi. Moreover, we provide an end-to-end model interpretation pipeline. CONCLUSION: Our analysis of different feature fusion strategies shows that the proposed feature concatenation method leads to better suppression of non-informative features, which drastically reduces the classifier overhead resulting in a robust lightweight network. SIGNIFICANCE: Lung sound event detection is a primary diagnostic step for numerous respiratory diseases. The proposed method provides a cost-effective and efficient alternative to exhaustive manual segmentation, and provides more accurate segmentation than existing methods. The end-to-end model interpretability helps to build the required trust in the system for use in clinical settings. Tharindu Fernando, Sridha Sridharan, Simon Denman, Houman Ghaemmaghami, Clinton Fookes |
IEEE J. Biomed. Health Informatics | 3 |
| 2021 | Learning Regional Attention Over Multi-Resolution Deep Convolutional Features For Trademark RetrievalabstractLarge-scale trademark retrieval is an important content-based image retrieval task. A recent study shows that off-the-shelf deep features aggregated with Regional-Maximum Activation of Convolutions (R-MAC) achieve state-of-the-art results. However, R-MAC suffers in the presence of background clutter/trivial regions and scale variance, and discards important spatial information. We introduce three simple but effective modifications to R-MAC to overcome these drawbacks. First, we propose the use of both sum and max pooling to minimise the loss of spatial information. We also employ domain-specific unsupervised soft-attention to eliminate background clutter and unimportant regions. Finally, we add multi-resolution inputs to enhance the scale-invariance of R-MAC. We evaluate these three modifications on the million-scale METU dataset. Our results show that all modifications bring non-trivial improvements, and surpass previous state-of-the-art results. Osman Tursun, Simon Denman, Sridha Sridharan, Clinton Fookes |
ICIP | 2 |
| 2021 | IGSSTRCF: Importance Guided Sparse Spatio-Temporal Regularized Correlation Filters For TrackingabstractThis paper proposes a novel Importance Guided Sparse Spatio-Temporal Regularization based Correlation Filter (IGSSTRCF) tracker. Our formulation explicitly models the variations in the correlation filters and associated spatial weights in successive frames. By imposing a sparsity penalty on these variations, the formulation ensures that only relevant changes are incorporated during updates. This results in more robust filter coefficients that minimize the tracking drift. The IGSSTRCF also includes an adaptive channel importance estimation strategy that assigns an importance weight to each feature channel during training. The proposed formulation is efficiently solved via the alternating direction method of multipliers. A comparative analysis is shown on TC128, UAV123, VOT-2017, and VOT-2019 datasets; and we present an ablation study to demonstrate the contribution of each component of the IGSSTRCF. It is observed that we outperform several state-of-the-art trackers and each component of the proposed IGSSTRCF contributes positively towards tracker performance. A. Venkata Subramanyam, Simon Denman, Sridha Sridharan, Clinton Fookes |
WACV | 3 |
| 2021 | Detection of Fake and Fraudulent Faces via Neural Memory NetworksabstractAdvances in computer vision have brought us to the point where we have the ability to synthesise realistic fake content. Such approaches are seen as a source of disinformation and mistrust, and pose serious concerns to governments around the world. Convolutional Neural Networks (CNNs) demonstrate encouraging results when detecting fake images that arise from the specific type of manipulation they are trained on. However, this success has not transitioned to unseen manipulation types, resulting in a significant gap in the line-of-defense. We propose a Hierarchical Attention Memory Network (HAMN), motivated by the social cognition processes of the human brain, for the detection of fake faces. Through visual cues and by utilising knowledge stored in neural memories, we allow the network to reason about the perceived face and anticipate it's future semantic embeddings. This renders a generalisable face tampering detection framework. Experimental results demonstrate the proposed approach achieves superior performance for fake and fraudulent face detection. Tharindu Fernando, Clinton Fookes, Simon Denman, Sridha Sridharan |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2021 | End-to-End Domain Adaptive Attention Network for Cross-Domain Person Re-IdentificationabstractPerson re-identification (re-ID) remains challenging in a real-world scenario, as it requires a trained network to generalise to totally unseen target data in the presence of variations across domains. Recently, generative adversarial models have been widely adopted to enhance the diversity of training data. These approaches, however, often fail to generalise to other domains, as existing generative person re-identification models have a disconnect between the generative component and the discriminative feature learning stage. To address the on-going challenges regarding model generalisation, we propose an end-to-end domain adaptive attention network to jointly translate images between domains and learn discriminative re-id features in a single framework. To address the domain gap challenge, we introduce an attention module for image translation from source to target domains without affecting the identity of a person. More specifically, attention is directed to the background instead of the entire image of the person, ensuring identifying characteristics of the subject are preserved. The proposed joint learning network results in a significant performance improvement over state-of-the-art methods on several challenging benchmark datasets. Amena Khatun, Simon Denman, Sridha Sridharan, Clinton Fookes |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2021 | TMMF: Temporal Multi-Modal Fusion for Single-Stage Continuous Gesture RecognitionabstractGesture recognition is a much studied research area which has myriad real-world applications including robotics and human-machine interaction. Current gesture recognition methods have focused on recognising isolated gestures, and existing continuous gesture recognition methods are limited to two-stage approaches where independent models are required for detection and classification, with the performance of the latter being constrained by detection performance. In contrast, we introduce a single-stage continuous gesture recognition framework, called Temporal Multi-Modal Fusion (TMMF), that can detect and classify multiple gestures in a video via a single model. This approach learns the natural transitions between gestures and non-gestures without the need for a pre-processing segmentation step to detect individual gestures. To achieve this, we introduce a multi-modal fusion mechanism to support the integration of important information that flows from multi-modal inputs, and is scalable to any number of modes. Additionally, we propose Unimodal Feature Mapping (UFM) and Multi-modal Feature Mapping (MFM) models to map uni-modal features and the fused multi-modal features respectively. To further enhance performance, we propose a mid-point based loss function that encourages smooth alignment between the ground truth and the prediction, helping the model to learn natural gesture transitions. We demonstrate the utility of our proposed framework, which can handle variable-length input videos, and outperforms the state-of-the-art on three challenging datasets: EgoGesture, IPN hand and ChaLearn LAP Continuous Gesture Dataset (ConGD). Furthermore, ablation experiments show the importance of different components of the proposed framework. Harshala Gammulle, Simon Denman, Sridha Sridharan, Clinton Fookes |
IEEE Trans. Image Process. | 2 |
| 2021 | Identification of Children at Risk of Schizophrenia via Deep Learning and EEG ResponsesabstractThe prospective identification of children likely to develop schizophrenia is a vital tool to support early interventions that can mitigate the risk of progression to clinical psychosis. Electroencephalographic (EEG) patterns from brain activity and deep learning techniques are valuable resources in achieving this identification. We propose automated techniques that can process raw EEG waveforms to identify children who may have an increased risk of schizophrenia compared to typically developing children. We also analyse abnormal features that remain during developmental follow-up over a period of ∼ 4 years in children with a vulnerability to schizophrenia initially assessed when aged 9 to 12 years. EEG data from participants were captured during the recording of a passive auditory oddball paradigm. We undertake a holistic study to identify brain abnormalities, first by exploring traditional machine learning algorithms using classification methods applied to hand-engineered features (event-related potential components). Then, we compare the performance of these methods with end-to-end deep learning techniques applied to raw data. We demonstrate via average cross-validation performance measures that recurrent deep convolutional neural networks can outperform traditional machine learning methods for sequence modeling. We illustrate the intuitive salient information of the model with the location of the most relevant attributes of a post-stimulus window. This baseline identification system in the area of mental illness supports the evidence of developmental and disease effects in a pre-prodromal phase of psychosis. These results reinforce the benefits of deep learning to support psychiatric classification and neuroscientific research more broadly. David Ahmedt-Aristizabal, Tharindu Fernando, Simon Denman, Jonathan E. Robinson, Sridha Sridharan, Patrick J. Johnston, Kristin R. Laurens, Clinton Fookes |
IEEE J. Biomed. Health Informatics | 3 |
| 2021 | A Robust Interpretable Deep Learning Classifier for Heart Anomaly Detection Without SegmentationabstractTraditionally, abnormal heart sound classification is framed as a three-stage process. The first stage involves segmenting the phonocardiogram to detect fundamental heart sounds; after which features are extracted and classification is performed. Some researchers in the field argue the segmentation step is an unwanted computational burden, whereas others embrace it as a prior step to feature extraction. When comparing accuracies achieved by studies that have segmented heart sounds before analysis with those who have overlooked that step, the question of whether to segment heart sounds before feature extraction is still open. In this study, we explicitly examine the importance of heart sound segmentation as a prior step for heart sound classification, and then seek to apply the obtained insights to propose a robust classifier for abnormal heart sound detection. Furthermore, recognizing the pressing need for explainable Artificial Intelligence (AI) models in the medical domain, we also unveil hidden representations learned by the classifier using model interpretation techniques. Experimental results demonstrate that the segmentation which can be learned by the model plays an essential role in abnormal heart sound classification. Our new classifier is also shown to be robust, stable and most importantly, explainable, with an accuracy of almost 100% on the widely used PhysioNet dataset. Theekshana Dissanayake, Tharindu Fernando, Simon Denman, Sridha Sridharan, Houman Ghaemmaghami, Clinton Fookes |
IEEE J. Biomed. Health Informatics | 3 |
| 2020 | Geometry-Constrained Car Recognition Using a 3D Perspective NetworkabstractWe present a novel learning framework for vehicle recognition from a single RGB image. Unlike existing methods which only use attention mechanisms to locate 2D discriminative information, our work learns a novel 3D perspective feature representation of a vehicle, which is then fused with 2D appearance feature to predict the category. The framework is composed of a global network (GN), a 3D perspective network (3DPN), and a fusion network. The GN is used to locate the region of interest (RoI) and generate the 2D global feature. With the assistance of the RoI, the 3DPN estimates the 3D bounding box under the guidance of the proposed vanishing point loss, which provides a perspective geometry constraint. Then the proposed 3D representation is generated by eliminating the viewpoint variance of the 3D bounding box using perspective transformation. Finally, the 3D and 2D feature are fused to predict the category of the vehicle. We present qualitative and quantitative results on the vehicle classification and verification tasks in the BoxCars dataset. The results demonstrate that, by learning such a concise 3D representation, we can achieve superior performance to methods that only use 2D information while retain 3D meaningful information without the challenge of requiring a 3D CAD model. ZongYuan Ge, Simon Denman, Sridha Sridharan, Clinton Fookes |
AAAI | 3 |
| 2020 | Sparse Convolutions on Continuous Domains for Point Cloud and Event Stream Networks
Dominic Jack, Frédéric Maire, Simon Denman, Anders P. Eriksson |
ACCV (1) | 3 |
| 2020 | Attention Driven Fusion for Multi-Modal Emotion RecognitionabstractDeep learning has emerged as a powerful alternative to hand-crafted methods for emotion recognition on combined acoustic and text modalities. Baseline systems model emotion information in text and acoustic modes independently using Deep Convolutional Neural Networks (DCNN) and Recurrent Neural Networks (RNN), followed by applying attention, fusion, and classification. In this paper, we present a deep learning-based approach to exploit and fuse text and acoustic data for emotion classification. We utilize a SincNet layer, based on parameterized sinc functions with band-pass filters, to extract acoustic features from raw audio followed by a DCNN. This approach learns filter banks tuned for emotion recognition and provides more effective features compared to directly applying convolutions over the raw speech signal. For text processing, we use two branches (a DCNN and a Bi-direction RNN followed by a DCNN) in parallel where cross attention is introduced to infer the N-gram level correlations on hidden representations received from the Bi-RNN. Following existing state-of-the-art, we evaluate the performance of the proposed system on the IEMOCAP dataset. Experimental results indicate that the proposed system outperforms existing methods, achieving 5.2% improvement in weighted accuracy. Darshana Priyasad, Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
ICASSP | 3 |
| 2020 | Two-Stream Deep Feature Modelling for Automated Video Endoscopy Data Analysis
Harshala Gammulle, Simon Denman, Sridha Sridharan, Clinton Fookes |
MICCAI (3) | 2 |
| 2020 | Semantic Consistency and Identity Mapping Multi-Component Generative Adversarial Network for Person Re-IdentificationabstractIn a real world environment, person re-identification (Re-ID) is a challenging task due to variations in lighting conditions, viewing angles, pose and occlusions. Despite recent performance gains, current person Re-ID algorithms still suffer heavily when encountering these variations. To address this problem, we propose a semantic consistency and identity mapping multi-component generative adversarial network (SC-IMGAN) which provides style adaptation from one to many domains. To ensure that transformed images are as realistic as possible, we propose novel identity mapping and semantic consistency losses to maintain identity across the diverse domains. For the Re-ID task, we propose a joint verification-identification quartet network which is trained with generated and real images, followed by an effective quartet loss for verification. Our proposed method outperforms state-of-the-art techniques on six challenging person Re-ID datasets: CUHK01, CUHK03, VIPeR, PRID2011, iLIDS and Market-1501. Amena Khatun, Simon Denman, Sridha Sridharan, Clinton Fookes |
WACV | 2 |
| 2020 | LSTM guided ensemble correlation filter tracking with appearance model pool
A. Venkata Subramanyam, Simon Denman, Sridha Sridharan, Clinton Fookes |
Comput. Vis. Image Underst. | 3 |
| 2020 | Joint identification-verification for person re-identification: A four stream deep learning approach with improved quartet loss function
Amena Khatun, Simon Denman, Sridha Sridharan, Clinton Fookes |
Comput. Vis. Image Underst. | 2 |
| 2020 | MTRNet++: One-stage mask-based scene text eraser
Osman Tursun, Simon Denman, Sabesan Sivipalan, Sridha Sridharan, Clinton Fookes |
Comput. Vis. Image Underst. | 2 |
| 2020 | Improved reinforcement learning with curriculum
Joseph West, Frédéric Maire, Cameron Browne, Simon Denman |
Expert Syst. Appl. | 4 |
| 2020 | Neural memory plasticity for medical anomaly detection
Tharindu Fernando, Simon Denman, David Ahmedt-Aristizabal, Sridha Sridharan, Kristin R. Laurens, Patrick J. Johnston, Clinton Fookes |
Neural Networks | 2 |
| 2020 | Fine-grained action segmentation using the semi-supervised action GAN
Harshala Gammulle, Simon Denman, Sridha Sridharan, Clinton Fookes |
Pattern Recognit. | 2 |
| 2020 | Hierarchical Attention Network for Action Segmentation
Harshala Gammulle, Simon Denman, Sridha Sridharan, Clinton Fookes |
Pattern Recognit. Lett. | 2 |
| 2020 | Temporarily-Aware Context Modeling Using Generative Adversarial Networks for Speech Activity DetectionabstractThis paper presents a novel framework for Speech Activity Detection (SAD). Inspired by the recent success of multi-task learning approaches in the speech processing domain, we propose a novel joint learning framework for SAD. We utilise generative adversarial networks to automatically learn a loss function for joint prediction of the frame-wise speech/ non-speech classifications together with the next audio segment. In order to exploit the temporal relationships within the input signal, we propose a temporal discriminator which aims to ensure that the predicted signal is temporally consistent. We evaluate the proposed framework on multiple public benchmarks, including NIST OpenSAT' 17, AMI Meeting and HAVIC, where we demonstrate its capability to outperform state-of-the-art SAD approaches. Furthermore, our cross-database evaluations demonstrate the robustness of the proposed approach across different languages, accents, and acoustic environments. Tharindu Fernando, Sridha Sridharan, Mitchell McLaren, Darshana Priyasad, Simon Denman, Clinton Fookes |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2020 | Heart Sound Segmentation Using Bidirectional LSTMs With AttentionabstractOBJECTIVE: This paper proposes a novel framework for the segmentation of phonocardiogram (PCG) signals into heart states, exploiting the temporal evolution of the PCG as well as considering the salient information that it provides for the detection of the heart state. METHODS: We propose the use of recurrent neural networks and exploit recent advancements in attention based learning to segment the PCG signal. This allows the network to identify the most salient aspects of the signal and disregard uninformative information. RESULTS: The proposed method attains state-of-the-art performance on multiple benchmarks including both human and animal heart recordings. Furthermore, we empirically analyse different feature combinations including envelop features, wavelet and Mel Frequency Cepstral Coefficients (MFCC), and provide quantitative measurements that explore the importance of different features in the proposed approach. CONCLUSION: We demonstrate that a recurrent neural network coupled with attention mechanisms can effectively learn from irregular and noisy PCG recordings. Our analysis of different feature combinations shows that MFCC features and their derivatives offer the best performance compared to classical wavelet and envelop features. SIGNIFICANCE: Heart sound segmentation is a crucial pre-processing step for many diagnostic applications. The proposed method provides a cost effective alternative to labour extensive manual segmentation, and provides a more accurate segmentation than existing methods. As such, it can improve the performance of further analysis including the detection of murmurs and ejection clicks. The proposed method is also applicable for detection and segmentation of other one dimensional biomedical signals. Tharindu Fernando, Houman Ghaemmaghami, Simon Denman, Sridha Sridharan, Nayyar Hussain, Clinton Fookes |
IEEE J. Biomed. Health Informatics | 3 |
| 2020 | Memory Augmented Deep Generative Models for Forecasting the Next Shot Location in TennisabstractThis paper presents a novel framework for predicting shot location and type in tennis. Inspired by recent neuroscience discoveries, we incorporate neural memory modules to model the episodic and semantic memory components of a tennis player. We propose a Semi-Supervised Generative Adversarial Network architecture that couples these memory models with the automatic feature learning power of deep neural networks, and demonstrate methodologies for learning player level behavioral patterns with the proposed framework. We evaluate the effectiveness of the proposed model on tennis tracking data from the 2012 Australian Tennis Open and exhibit applications of the proposed method in discovering how players adapt their style depending on the match context. Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2019 | Forecasting Future Action Sequences with Neural Memory Networks
Harshala Gammulle, Simon Denman, Sridha Sridharan, Clinton Fookes |
BMVC | 2 |
| 2019 | Predicting the Future: A Jointly Learnt Model for Action AnticipationabstractInspired by human neurological structures for action anticipation, we present an action anticipation model that enables the prediction of plausible future actions by forecasting both the visual and temporal future. In contrast to current state-of-the-art methods which first learn a model to predict future video features and then perform action anticipation using these features, the proposed framework jointly learns to perform the two tasks, future visual and temporal representation synthesis, and early action anticipation. The joint learning framework ensures that the predicted future embeddings are informative to the action anticipation task. Furthermore, through extensive experimental evaluations we demonstrate the utility of using both visual and temporal semantics of the scene, and illustrate how this representation synthesis could be achieved through a recurrent Generative Adversarial Network (GAN) framework. Our model outperforms the current state-of-the-art methods on multiple datasets: UCF101, UCF101-24, UT-Interaction and TV Human Interaction. Harshala Gammulle, Simon Denman, Sridha Sridharan, Clinton Fookes |
ICCV | 2 |
| 2019 | MTRNet: A Generic Scene Text EraserabstractText removal algorithms have been proposed for uni-lingual scripts with regular shapes and layouts. However, to the best of our knowledge, a generic text removal method which is able to remove all or user-specified text regions regardless of font, script, language or shape is not available. Developing such a generic text eraser for real scenes is a challenging task, since it inherits all the challenges of multi-lingual and curved text detection and inpainting. To fill this gap, we propose a mask-based text removal network (MTRNet). MTRNet is a conditional adversarial generative network (cGAN) with an auxiliary mask. The introduced auxiliary mask not only makes the cGAN a generic text eraser, but also enables stable training and early convergence on a challenging large-scale synthetic dataset, initially proposed for text detection in real scenes. What's more, MTRNet achieves state-of-the-art results on several real-world datasets including ICDAR 2013, ICDAR 2017 MLT, and CTW1500, without being explicitly trained on this data, outperforming previous state-of-the-art methods trained directly on these datasets. Osman Tursun, Simon Denman, Sabesan Sivipalan, Sridha Sridharan, Clinton Fookes |
ICDAR | 3 |
| 2019 | Coupled Generative Adversarial Network for Continuous Fine-Grained Action SegmentationabstractWe propose a novel conditional GAN (cGAN) model for continuous fine-grained human action segmentation, that utilises multi-modal data and learned scene context information. The proposed approach utilises two GANs: termed Action GAN and Auxiliary GAN, where the Action GAN is trained to operate over the current RGB frame while the Auxiliary GAN utilises supplementary information such as depth or optical flow. The goal of both GANs is to generate similar 'action codes', a vector representation of the current action. To facilitate this process a context extractor that incorporates data and recent outputs from both modes is used to extract context information to aids recognition performance. The result is a recurrent GAN architecture which learns a task specific loss function from multiple feature modalities. Extensive evaluations on variants of the proposed model to show the importance of utilising different streams of information such as context and auxiliary information in the proposed network; and show that our model is capable of outperforming state-of-the-art methods for three widely used datasets: 50 Salads, MERL Shopping and Georgia Tech Egocentric Activities, comprising both static and dynamic camera settings. Harshala Gammulle, Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
WACV | 3 |
| 2019 | Multimodal clothing recognition for semantic search in unconstrained surveillance imagery
Michael Halstead, Simon Denman, Sridha Sridharan, Yingli Tian, Clinton Fookes |
J. Vis. Commun. Image Represent. | 2 |
| 2019 | Scene Invariant Virtual Gates Using DNNsabstractUnderstanding where people are located and how they are moving about in an environment is critical for operators of large public spaces such as shopping centers, and large public infrastructures such as airports. Automated analysis of CCTV footage is increasingly being used to address this need through techniques that can count crowd sizes, estimate their density, and estimate the through-put of people into and/or out of a choke-point. A limitation of using CCTV based approaches, however, is the need to train models specific to each view which, for large environments with 100s or 1000s of cameras, can quickly become problematic. While there is some success in developing scene-invariant crowd counting and crowd density estimation approaches, much less attention has been given to developing scene-invariant solutions for through-put estimation. In this paper, we investigate the use of convolutional neural network and long short-term memory architectures to estimate pedestrian through-put from arbitrary CCTV viewpoints. To properly develop and demonstrate our approach, we present a new 22 view database featuring 44 h of pedestrian throughput annotation, containing over 11 000 annotated people; and using this proposed approach we show that we are able to outperform a scene-dependant approach across a diverse set of challenging view-points. Simon Denman, Clinton Fookes, Prasad K. D. V. Yarlagadda, Sridha Sridharan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Understanding Patients' Behavior: Vision-Based Analysis of Seizure DisordersabstractA substantial proportion of patients with functional neurological disorders (FND) are being incorrectly diagnosed with epilepsy because their semiology resembles that of epileptic seizures (ES). Misdiagnosis may lead to unnecessary treatment and its associated complications. Diagnostic errors often result from an overreliance on specific clinical features. Furthermore, the lack of electrophysiological changes in patients with FND can also be seen in some forms of epilepsy, making diagnosis extremely challenging. Therefore, understanding semiology is an essential step for differentiating between ES and FND. Existing sensor-based and marker-based systems require physical contact with the body and are vulnerable to clinical situations such as patient positions, illumination changes, and motion discontinuities. Computer vision and deep learning are advancing to overcome these limitations encountered in the assessment of diseases and patient monitoring; however, they have not been investigated for seizure disorder scenarios. Here, we propose and compare two marker-free deep learning models, a landmark-based and a region-based model, both of which are capable of distinguishing between seizures from video recordings. We quantify semiology by using either a fusion of reference points and flow fields, or through the complete analysis of the body. Average leave-one-subject-out cross-validation accuracies for the landmark-based and region-based approaches of 68.1% and 79.6% in our dataset collected from 35 patients, reveal the benefit of video analytics to support automated identification of semiology in the challenging conditions of a hospital setting. David Ahmedt-Aristizabal, Simon Denman, Kien Nguyen Thanh, Sridha Sridharan, Sasha Dionisio, Clinton Fookes |
IEEE J. Biomed. Health Informatics | 2 |
| 2018 | GD-GAN: Generative Adversarial Networks for Trajectory Prediction and Group Detection in Crowds
Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
ACCV (1) | 2 |
| 2018 | Multi-level Sequence GAN for Group Activity Recognition
Harshala Gammulle, Simon Denman, Sridha Sridharan, Clinton Fookes |
ACCV (1) | 2 |
| 2018 | Rethinking Planar Homography Estimation Using Perspective Fields
Simon Denman, Sridha Sridharan, Clinton Fookes |
ACCV (6) | 2 |
| 2018 | Semantic Person Retrieval in Surveillance Using Soft Biometrics: AVSS 2018 Challenge IIabstractIn surveillance and security today it is a common goal to locate a subject of interest purely from a semantic description; think of an offender description form handed into a law enforcement agency. To date, these tasks are primarily undertaken by operators on the ground either by manually searching a premises or by combing through hours of video footage. Using computer vision to attempt to partially or fully automate these tasks has been gathering interest within the research community in recent years, however, to date there has been little coordinated effort to advance the field. This has motivated the challenge that is presented in this paper: the AVSS Challenge on Semantic Person Retrieval in Surveillance Using Soft Biometrics. This challenge consists of two related tasks: person re-identification from a semantic query and person search within a video from a query. In this paper, we present the publicly available data for this challenge, the evaluation framework, and the challenge results. It is our hope that the outcomes of this challenge and the availability of the data used in this challenge will expedite research and development in this societal field. Michael Halstead, Simon Denman, Clinton Fookes, Yingli Tian, Mark S. Nixon |
AVSS | 2 |
| 2018 | Calibrating Cameras in Poor-Conditioned Pitch-Based Sports GamesabstractCamera calibration is a preliminary step in sports analytics which enables us to transform player positions to standard playing area coordinates. While many camera calibration systems work well when the visual content contains sufficient clues, such as a key frame, calibrating without such information, such as may be needed when processing footage captured by a coach from the sidelines or stands, is challenging. In this paper an innovative automatic camera calibration system, which does not make use of any key frames, is presented for sports analytics. The proposed system consists of three components: a robust linear panorama module, a playing area estimation module, and a homography estimation module. It can eliminate distortion and calibrate the camera in each frame simultaneously, using correspondences between pairs of consecutive frames. Experiments on real data evaluate the performance and demonstrate the robustness of the system. Ruan Lakemond, Simon Denman, Sridha Sridharan, Clinton Fookes, Stuart Morgan |
ICASSP | 3 |
| 2018 | Pedestrian Trajectory Prediction with Structured Memory Hierarchies
Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
ECML/PKDD (1) | 2 |
| 2018 | Tracking by Prediction: A Deep Generative Model for Mutli-person Localisation and TrackingabstractCurrent multi-person localisation and tracking systems have an over reliance on the use of appearance models for target re-identification and almost no approaches employ a complete deep learning solution for both objectives. We present a novel, complete deep learning framework for multi-person localisation and tracking. In this context we first introduce a light weight sequential Generative Adversarial Network architecture for person localisation, which overcomes issues related to occlusions and noisy detections, typically found in a multi person environment. In the proposed tracking framework we build upon recent advances in pedestrian trajectory prediction approaches and propose a novel data association scheme based on predicted trajectories. This removes the need for computationally expensive person re-identification systems based on appearance features and generates human like trajectories with minimal fragmentation. The proposed method is evaluated on multiple public benchmarks including both static and dynamic cameras and is capable of generating outstanding performance, especially among other recently proposed deep neural network based approaches. Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
WACV | 2 |
| 2018 | Task Specific Visual Saliency Prediction with Memory Augmented Conditional Generative Adversarial Networks
Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
WACV | 2 |
| 2018 | A Deep Four-Stream Siamese Convolutional Neural Network with Joint Verification and Identification Loss for Person Re-DetectionabstractState-of-the-art person re-identification systems that employ a triplet based deep network suffer from a poor generalization capability. In this paper, we propose a four stream Siamese deep convolutional neural network for person redetection that jointly optimises verification and identification losses over a four image input group. Specifically, the proposed method overcomes the weakness of the typical triplet formulation by using groups of four images featuring two matched (i.e. the same identity) and two mismatched images. This allows us to jointly increase the interclass variations and reduce the intra-class variations in the learned feature space. The proposed approach also optimises over both the identification and verification losses, further minimising intra-class variation and maximising inter-class variation, improving overall performance. Extensive experiments on four challenging datasets, VIPeR, CUHK01, CUHK03 and PRID2011, demonstrates that the proposed approach achieves state-of-the-art performance. Amena Khatun, Simon Denman, Sridha Sridharan, Clinton Fookes |
WACV | 2 |
| 2018 | Tree Memory Networks for modelling long-term temporal dependencies
Tharindu Fernando, Simon Denman, Aaron McFadyen, Sridha Sridharan, Clinton Fookes |
Neurocomputing | 2 |
| 2018 | Soft + Hardwired attention: An LSTM framework for human trajectory prediction and abnormal event detection
Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
Neural Networks | 2 |
| 2017 | Deep discovery of facial motions using a shallow embedding layerabstractUnique encoding of the dynamics of facial actions has potential to provide a spontaneous facial expression recognition system. The most promising existing approaches rely on deep learning of facial actions. However, current approaches are often computationally intensive and require a great deal of memory/processing time, and typically the temporal aspect of facial actions are often ignored, despite the potential wealth of information available from the spatial dynamic movements and their temporal evolution over time from neutral state to apex state. To tackle aforementioned challenges, we propose a deep learning framework by using the 3D convolutional filters to extract spatio-temporal features, followed by the LSTM network which is able to integrate the dynamic evolution of short-duration of spatio-temporal features as an emotion progresses from the neutral state to the apex state. In order to reduce the redundancy of parameters and accelerate the learning of the recurrent neural network, we propose a shallow embedding layer to reduce the number of parameters in the LSTM by up to 98% without sacrificing recognition accuracy. As the fully connected layer approximately contains 95% of the parameters in the network, we decrease the number of parameters in this layer before passing features to the LSTM network, which significantly improves training speed and enables the possibility of deploying a state of the art deep network on real-time applications. We evaluate our proposed framework on the DISFA and UNBC-McMaster Shoulder pain datasets. Afsane Ghasemi, Mahsa Baktash, Simon Denman, Sridha Sridharan, Dung Nguyen Tien, Clinton Fookes |
ICIP | 3 |
| 2017 | Single image depth prediction using super-column super-pixel featuresabstractDepth prediction from a single monocular image is a challenging yet valuable task, as often a depth sensor is not available. The state-of-the-art approach [1] combines a deep fully convolutional network (DFCN) with a conditional random field (CRF), allowing the CRF to correct and smooth the depth values estimated by the DFCN according to efficient contextual modeling. However, using the output of the DFCN as unary input for CRF is limited by using only the last layer of the DFCN. The middle layers of the DFCN have been shown to carry useful information for other scene understanding tasks, which may help to improve the prediction quality. This paper proposes a novel super-column superpixel (SCSP) feature that is the combination of multiple layers of the DFCN after a super-pixel pooling process. The proposed approach based on the SCSP features reduces the root mean square (rms) error of the prediction by more than 16% in NYUv2 dataset. Xufeng Guo, Kien Nguyen Thanh, Simon Denman, Clinton Fookes, Sridha Sridharan |
ICIP | 3 |
| 2017 | From Affine Rank Minimization Solution to Sparse ModelingabstractCompressed sensing is a simple and efficient technique that has a number of applications in signal processing and machine learning. In machine learning it provides answers to questions such as: "under what conditions is the sparse representation of data efficient?", "when is learning a large margin classifier directly on the compressed domain possible?", and "why does a large margin classifier learn more effectively if the data is sparse?". This work tackles the problem of feature representation from the context of sparsity and affine rank minimization by leveraging compressed sensing from the learning perspective in order to provide answers to the aforementioned questions. We show, for a full-rank signal, the high dimensional sparse representation of data is efficient because from the classifiers viewpoint such a representation is in fact a low dimensional problem. We provide practical bounds on the linear classifier to investigate the relationship between the SVM classifier in the high dimensional and compressed domains and show for the high dimensional sparse signals, when the bounds are tight, directly learning in the compressed domain is possible. Iman Abbasnejad, Sridha Sridharan, Simon Denman, Clinton Fookes, Simon Lucey |
WACV | 3 |
| 2017 | Two Stream LSTM: A Deep Fusion Framework for Human Action RecognitionabstractIn this paper we address the problem of human action recognition from video sequences. Inspired by the exemplary results obtained via automatic feature learning and deep learning approaches in computer vision, we focus our attention towards learning salient spatial features via a convolutional neural network (CNN) and then map their temporal relationship with the aid of Long-Short-Term-Memory (LSTM) networks. Our contribution in this paper is a deep fusion framework that more effectively exploits spatial features from CNNs with temporal features from LSTM models. We also extensively evaluate their strengths and weaknesses. We find that by combining both the sets of features, the fully connected features effectively act as an attention mechanism to direct the LSTM to interesting parts of the convolutional feature sequence. The significance of our fusion method is its simplicity and effectiveness compared to other state-of-the-art methods. The evaluation results demonstrate that this hierarchical multi stream fusion method has higher performance compared to single stream mapping methods allowing it to achieve high accuracy outperforming current state-of-the-art methods in three widely used databases: UCF11, UCFSports, jHMDB. Harshala Gammulle, Simon Denman, Sridha Sridharan, Clinton Fookes |
WACV | 2 |
| 2016 | A robust UAV landing site detection system using mid-level discriminative patchesabstractThe forced landing problem has become one of the main impediments to UAV's entering civilian airspace. Unfortunately there is no robust forced landing site detection system that will reliably detect a safe landing site. One of the main reasons for this is the difficulty in considering the various classes of surface, to determine whether they are safe or not. We propose a robust UAV landing site detection system using midlevel discriminative patches. The training and tuning process uses a dataset containing 1600 randomly selected Google map images with weak labels.We then show how the output from multiple mid-level discriminative patch detectors can be combined to indicate the level or danger for a given region. The proposed technique reliably detects safe landing areas in UAV imagery, and achieves improved performance over the state-of-the art. The proposed system outperforms the baseline system by 29.4% for completeness and 33.9% for correctness, and is invariant to the changes of illumination, sharpness and resolution of images. Xufeng Guo, Simon Denman, Clinton Fookes, Sridha Sridharan |
ICPR | 2 |
| 2016 | Discovery of facial motions using deep machine perceptionabstractDeep, intuitive understanding of facial motions has the potential to provide an intelligent facial expression system as well as a unique encoding of the dynamics of facial actions. The most promising existing approaches rely on extracting hand crafted features; and existing approaches typically work best in constrained conditions and do not generalise well to varying environmental conditions which make them poorly suited to applications such as real-time human robot interactions. In this paper, we propose a multi-label deep learning based facial action detector, which along with a linear SVM classifier outperforms state of the art approaches such as HOG and LBP. We show that our approach can be generalized to other datasets by learning inner data structure, encoding facial actions, and providing a hierarchical representation of facial features. Our experimental results also demonstrate the efficiency of using image patches, which results in faster learning convergence while outperforms holistic approaches. We evaluate our proposed frame-work on the DISFA and CK+ datasets. Afsane Ghasemi, Simon Denman, Sridha Sridharan, Clinton Fookes |
WACV | 2 |
| 2016 | Detecting rare events using Kullback-Leibler divergence: A weakly supervised approach
Jingxin Xu, Simon Denman, Clinton Fookes, Sridha Sridharan |
Expert Syst. Appl. | 2 |
| 2015 | Large scale monitoring of crowds and building utilisation: A new database and distributed approachabstractPublic buildings and large infrastructure are typically monitored by tens or hundreds of cameras, all capturing different physical spaces and observing different types of interactions and behaviours. However to date, in large part due to limited data availability, crowd monitoring and operational surveillance research has focused on single camera scenarios which are not representative of real-world applications. In this paper we present a new, publicly available database for large scale crowd surveillance. Footage from 12 cameras for a full work day covering the main floor of a busy university campus building, including an internal and external foyer, elevator foyers, and the main external approach are provided; alongside annotation for crowd counting (single or multi-camera) and pedestrian flow analysis for 10 and 6 sites respectively. We describe how this large dataset can be used to perform distributed monitoring of building utilisation, and demonstrate the potential of this dataset to understand and learn the relationship between different areas of a building. Simon Denman, Clinton Fookes, David Ryan, Sridha Sridharan |
AVSS | 1 |
| 2015 | Searching for semantic person queries using channel representationsabstractIt is not uncommon to hear a person of interest described by their height, build, and clothing (i.e. type and colour). These semantic descriptions are commonly used by people to describe others, as they are quick to relate and easy to understand. However such queries are not easily utilised within intelligent surveillance systems as they are difficult to transform into a representation that can be searched for automatically in large camera networks. In this paper we propose a novel approach that transforms such a semantic query into an avatar that is searchable within a video stream, and demonstrate state-of-the-art performance for locating a subject in video based on a description. Simon Denman, Michael Halstead, Clinton Fookes, Sridha Sridharan |
ICASSP | 1 |
| 2015 | Detecting rare events using Kullback-Leibler divergenceabstractOne main challenge in developing a system for visual surveillance event detection is the annotation of target events in the training data. By making use of the assumption that events with security interest are often rare compared to regular behaviours, this paper presents a novel approach by using Kullback-Leibler (KL) divergence for rare event detection in a weakly supervised learning setting, where only clip-level annotation is available. It will be shown that this approach outperforms state-of-the-art methods on a popular real-world dataset, while preserving real time performance. Jingxin Xu, Simon Denman, Clinton Fookes, Sridha Sridharan |
ICASSP | 2 |
| 2015 | Class-specific sparse codes for representing activitiesabstractIn this paper we investigate the effectiveness of class specific sparse codes in the context of discriminative action classification. The bag-of-words representation is widely used in activity recognition to encode features, and although it yields state-of-the art performance with several feature descriptors it still suffers from large quantization errors and reduces the overall performance. Recently proposed sparse representation methods have been shown to effectively represent features as a linear combination of an over complete dictionary by minimizing the reconstruction error. In contrast to most of the sparse representation methods which focus on Sparse-Reconstruction based Classification (SRC), this paper focuses on a discriminative classification using a SVM by constructing class-specific sparse codes for motion and appearance separately. Experimental results demonstrates that separate motion and appearance specific sparse coefficients provide the most effective and discriminative representation for each class compared to a single class-specific sparse coefficients. Sabanadesan Umakanthan, Simon Denman, Clinton Fookes, Sridha Sridharan |
ICIP | 2 |
| 2015 | An evaluation of crowd counting methods, features and regression models
David Ryan, Simon Denman, Sridha Sridharan, Clinton Fookes |
Comput. Vis. Image Underst. | 2 |
| 2015 | Automatic surveillance in transportation hubs: No longer just about catching the bad guy
Simon Denman, Tristan Kleinschmidt, David Ryan, Paul Barnes, Sridha Sridharan, Clinton Fookes |
Expert Syst. Appl. | 1 |
| 2015 | Searching for people using semantic soft biometric descriptions
Simon Denman, Michael Halstead, Clinton Fookes, Sridha Sridharan |
Pattern Recognit. Lett. | 1 |
| 2015 | An Efficient and Robust System for Multiperson Event Detection in Real-World Indoor Surveillance ScenesabstractDue to the popularity of security cameras in public places, it is of interest to design an intelligent system that can efficiently detect events automatically. This paper proposes a novel algorithm for multiperson event detection. To ensure greater than real-time performance, features are extracted directly from compressed MPEG video. A novel histogram-based feature descriptor that captures the angles between extracted particle trajectories is proposed, which allows us to capture motion patterns for multiperson events in the video. To alleviate the need for fine-grained annotation, we propose the use of labeled latent Dirichlet allocation, a weakly supervised method that allows the use of coarse temporal annotations, which are much simpler to obtain. This novel system is able to run at ~10 times real time, while preserving state-of-the-art detection performance for multiperson events on a 100-h real-world surveillance data set (TRECVid surveillance event detection). Jingxin Xu, Simon Denman, Sridha Sridharan, Clinton Fookes |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2015 | Score-Level Multibiometric Fusion Based on Dempster-Shafer Theory Incorporating Uncertainty FactorsabstractWhile existing multibiometic Dempster-Shafer theory fusion approaches have demonstrated promising performance, they do not model the uncertainty appropriately, suggesting that further improvement can be achieved. This research seeks to develop a unified framework for multimodal biometric fusion to take advantage of the uncertainty concept of Dempster-Shafer theory, improving the performance of multibiometric authentication systems. Modeling uncertainty as a function of uncertainty factors affecting the recognition performance of the biometric systems helps to address the uncertainty of the data and the confidence of the fusion outcome. A weighted combination of quality measures and classifiers performance (equal error rate) is proposed to encode the uncertainty concept to improve the fusion. We also found that quality measures contribute unequally to the recognition performance; thus, selecting only significant factors and fusing them with a Dempster-Shafer approach to generate an overall quality score play an important role in the success of uncertainty modeling. The proposed approach achieved a competitive performance (approximate 1% EER) in comparison with other Dempster-Shafer-based approaches and other conventional fusion approaches. Kien Nguyen Thanh, Simon Denman, Sridha Sridharan, Clinton Fookes |
IEEE Trans. Hum. Mach. Syst. | 2 |
| 2014 | An MRF based abnormal event detection approach using motion and appearance featuresabstractAbnormal event detection has attracted a lot of attention in the computer vision research community during recent years due to the increased focus on automated surveillance systems to improve security in public places. Due to the scarcity of training data and the definition of an abnormality being dependent on context, abnormal event detection is generally formulated as a data-driven approach where activities are modeled in an unsupervised fashion during the training phase. In this work, we use a Gaussian mixture model (GMM) to cluster the activities during the training phase, and propose a Gaussian mixture model based Markov random field (GMM-MRF) to estimate the likelihood scores of new videos in the testing phase. Further-more, we propose two new features: optical acceleration, and the histogram of optical flow gradients; to detect the presence of any abnormal objects and speed violations in the scene. We show that our proposed method outperforms other state of the art abnormal event detection algorithms on publicly available UCSD dataset. Hajananth Nallaivarothayan, Clinton Fookes, Simon Denman, Sridha Sridharan |
AVSS | 3 |
| 2014 | Locating People in Video from Semantic Descriptions: A New Database and ApproachabstractThe location of previously unseen and unregistered individuals in complex camera networks from semantic descriptions is a time consuming and often inaccurate process carried out by human operators, or security staff on the ground. To promote the development and evaluation of automated semantic description based localisation systems, we present a new, publicly available, unconstrained 110 sequence database, collected from 6 stationary cameras. Each sequence contains detailed semantic information for a single search subject who appears in the clip (gender, age, height, build, hair and skin colour, clothing type, texture and colour), and between 21 and 290 frames for each clip are annotated with the target subject location (over 11, 000 frames are annotated in total). A novel approach for localising a person given a semantic query is also proposed and demonstrated on this database. The proposed approach incorporates clothing colour and type (for clothing worn below the waist), as well as height and build to detect people. A method to assess the quality of candidate regions, as well as a symmetry driven approach to aid in modelling clothing on the lower half of the body, is proposed within this approach. An evaluation on the proposed dataset shows that a relative improvement in localisation accuracy of up to 21% is achieved over the baseline technique. Michael Halstead, Simon Denman, Sridha Sridharan, Clinton Fookes |
ICPR | 2 |
| 2014 | Multiple Instance Dictionary Learning for Activity RepresentationabstractThis paper presents an effective feature representation method in the context of activity recognition. Efficient and effective feature representation plays a crucial role not only in activity recognition, but also in a wide range of applications such as motion analysis, tracking, 3D scene understanding etc. In the context of activity recognition, local features are increasingly popular for representing videos because of their simplicity and efficiency. While they achieve state-of-the-art performance with low computational requirements, their performance is still limited for real world applications due to a lack of contextual information and models not being tailored to specific activities. We propose a new activity representation framework to address the shortcomings of the popular, but simple bag-of-words approach. In our framework, first multiple instance SVM (mi-SVM) is used to identify positive features for each action category and the k-means algorithm is used to generate a codebook. Then locality-constrained linear coding is used to encode the features into the generated codebook, followed by spatio-temporal pyramid pooling to convey the spatio-temporal statistics. Finally, an SVM is used to classify the videos. Experiments carried out on two popular datasets with varying complexity demonstrate significant performance improvement over the base-line bag-of-feature method. Sabanadesan Umakanthan, Simon Denman, Clinton Fookes, Sridha Sridharan |
ICPR | 2 |
| 2014 | Local inter-session variability modelling for object classificationabstractObject classification is plagued by the issue of session variation. Session variation describes any variation that makes one instance of an object look different to another, for instance due to pose or illumination variation. Recent work in the challenging task of face verification has shown that session variability modelling provides a mechanism to overcome some of these limitations. However, for computer vision purposes, it has only been applied in the limited setting of face verification. In this paper we propose a local region based intersession variability (ISV) modelling approach, and apply it to challenging real-world data. We propose a region based session variability modelling approach so that local session variations can be modelled, termed Local ISV. We then demonstrate the efficacy of this technique on a challenging real-world fish image database which includes images taken underwater, providing significant real-world session variations. This Local ISV approach provides a relative performance improvement of, on average, 23% on the challenging MOBIO, Multi-PIE and SCface face databases. It also provides a relative performance improvement of 35% on our challenging fish image dataset. Kaneswaran Anantharajah, ZongYuan Ge, Chris McCool, Simon Denman, Clinton Fookes, Peter I. Corke, Dian Tjondronegoro, Sridha Sridharan |
WACV | 4 |
| 2014 | Scene invariant multi camera crowd counting
David Ryan, Simon Denman, Clinton Fookes, Sridha Sridharan |
Pattern Recognit. Lett. | 2 |
| 2014 | Real-time video event detection in crowded scenes using MPEG derived features: A multiple instance learning approach
Jingxin Xu, Simon Denman, Vikas Reddy, Clinton Fookes, Sridha Sridharan |
Pattern Recognit. Lett. | 2 |
| 2013 | Feature-domain super-resolution for iris recognition
Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Simon Denman |
Comput. Vis. Image Underst. | 4 |
| 2012 | Unusual Scene Detection Using Distributed Behaviour Model and Sparse RepresentationabstractThe ability to detect unusual events in surviellance footage as they happen is a highly desireable feature for a surveillance system. However, this problem remains challenging in crowded scenes due to occlusions and the clustering of people. In this paper, we propose using the Distributed Behavior Model (DBM), which has been widely used in computer graphics, for video event detection. Our approach does not rely on object tracking, and is robust to camera movements. We use sparse coding for classification, and test our approach on various datasets. Our proposed approach outperforms a state-of-the-art work which uses the social force model and Latent Dirichlet Allocation. Jingxin Xu, Simon Denman, Clinton Fookes, Sridha Sridharan |
AVSS | 2 |
| 2012 | Activity Analysis in Complicated Scenes Using DFT Coefficients of Particle TrajectoriesabstractModelling activities in crowded scenes is very challenging as object tracking is not robust in complicated scenes and optical flow does not capture long range motion. We propose a novel approach to analyse activities in crowded scenesusing a "bag of particle trajectories". Particle trajectoriesare extracted from foreground regions within short video clips using particle video, which estimates long rangemotion in contrast to optical flow which is only concerned with inter-frame motion. Our applications include temporal video segmentation and anomaly detection, and we perform our evaluation on several real-world datasets containing complicated scenes. We show that our approaches achieve state-of-the-art performance for both tasks. Jingxin Xu, Simon Denman, Sridha Sridharan, Clinton Fookes |
AVSS | 2 |
| 2012 | Feature-domain super-resolution framework for Gabor-based face and iris recognitionabstractThe low resolution of images has been one of the major limitations in recognising humans from a distance using their biometric traits, such as face and iris. Superresolution has been employed to improve the resolution and the recognition performance simultaneously, however the majority of techniques employed operate in the pixel domain, such that the biometric feature vectors are extracted from a super-resolved input image. Feature-domain superresolution has been proposed for face and iris, and is shown to further improve recognition performance by capitalising on direct super-resolving the features which are used for recognition. However, current feature-domain superresolution approaches are limited to simple linear features such as Principal Component Analysis (PCA) and Linear Discriminant Analysis (LDA), which are not the most discriminant features for biometrics. Gabor-based features have been shown to be one of the most discriminant features for biometrics including face and iris. This paper proposes a framework to conduct super-resolution in the non-linear Gabor feature domain to further improve the recognition performance of biometric systems. Experiments have confirmed the validity of the proposed approach, demonstrating superior performance to existing linear approaches for both face and iris biometrics. Kien Nguyen Thanh, Sridha Sridharan, Simon Denman, Clinton Fookes |
CVPR | 3 |
| 2011 | Determining operational measures from multi-camera surveillance systems using soft biometricsabstractCCTV and surveillance networks are increasingly being used for operational as well as security tasks. One emerging area of technology that lends itself to operational analytics is soft biometrics. Soft biometrics can be used to describe a person and detect them throughout a sparse multi-camera network. This enables them to be used to perform tasks such as determining the time taken to get from point to point, and the paths taken through an environment by detecting and matching people across disjoint views. However, in a busy environment where there are 100's if not 1000's of people such as an airport, attempting to monitor everyone is highly unrealistic. In this paper we propose an average soft biometric, that can be used to identity people who look distinct, and are thus suitable for monitoring through a large, sparse camera network. We demonstrate how an average soft biometric can be used to identify unique people to calculate operational measures such as the time taken to travel from point to point. Simon Denman, Alina Bialkowski, Clinton Fookes, Sridha Sridharan |
AVSS | 1 |
| 2011 | Textures of optical flow for real-time anomaly detection in crowdsabstractAutomated visual surveillance of crowds is a rapidly growing area of research. In this paper we focus on motion representation for the purpose of abnormality detection in crowded scenes. We propose a novel visual representation called textures of optical flow. The proposed representation measures the uniformity of a flow field in order to detect anomalous objects such as bicycles, vehicles and skateboarders; and can be combined with spatial information to detect other forms of abnormality. We demonstrate that the proposed approach outperforms state-of-the-art anomaly detection algorithms on a large, publicly-available dataset. David Ryan, Simon Denman, Clinton Fookes, Sridha Sridharan |
AVSS | 2 |
| 2011 | 3D ellipsoid fitting for multi-view gait recognitionabstractGait recognition approaches continue to struggle with challenges including view-invariance, low-resolution data, robustness to unconstrained environments, and fluctuating gait patterns due to subjects carrying goods or wearing different clothes. Although computationally expensive, model based techniques offer promise over appearance based techniques for these challenges as they gather gait features and interpret gait dynamics in skeleton form. In this paper, we propose a fast 3D ellipsoidal-based gait recognition algorithm using a 3D voxel model derived from multi-view silhouette images. This approach directly solves the limitations of view dependency and self-occlusion in existing ellipse fitting model-based approaches. Voxel models are segmented into four components (left and right legs, above and below the knee), and ellipsoids are fitted to each region using eigenvalue decomposition. Features derived from the ellipsoid parameters are modeled using a Fourier representation to retain the temporal dynamic pattern for classification. We demonstrate the proposed approach using the CMU MoBo database and show that an improvement of 15-20% can be achieved over a 2D ellipse fitting baseline. Sabesan Sivipalan, Daniel Chen 0002, Simon Denman, Sridha Sridharan, Clinton Fookes |
AVSS | 3 |
| 2011 | Gait energy volumes and frontal gait recognition using depth imagesabstractGait energy images (GEIs) and its variants form the basis of many recent appearance-based gait recognition systems. The GEI combines good recognition performance with a simple implementation, though it suffers problems inherent to appearance-based approaches, such as being highly view dependent. In this paper, we extend the concept of the GEI to 3D, to create what we call the gait energy volume, or GEV. A basic GEV implementation is tested on the CMU MoBo database, showing improvements over both the GEI baseline and a fused multi-view GEI approach. We also demonstrate the efficacy of this approach on partial volume reconstructions created from frontal depth images, which can be more practically acquired, for example, in biometric portals implemented with stereo cameras, or other depth acquisition systems. Experiments on frontal depth images are evaluated on an in-house developed database captured using the Microsoft Kinect, and demonstrate the validity of the proposed approach. Sabesan Sivipalan, Daniel Chen 0002, Simon Denman, Sridha Sridharan, Clinton Fookes |
IJCB | 3 |
| 2011 | Feature-domain super-resolution for iris recognitionabstractUncooperative iris identification systems at a distance suffer from poor resolution of the captured iris images, which significantly degrades iris recognition performance. Super-resolution techniques have been employed to enhance the resolution of iris images and improve the recognition performance. However, all existing super-resolution approaches proposed for the iris biometric super-resolve pixel intensity values. This paper considers transferring super-resolution of iris images from the intensity domain to the feature domain. By directly super-resolving only the features essential for recognition, and by incorporating domain specific information from iris models, improved recognition performance compared to pixel domain super-resolution can be achieved. This is the first paper to investigate the possibility of feature-domain super-resolution for iris recognition, and experiments confirm the validity of the proposed approach. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Simon Denman |
ICIP | 4 |
| 2011 | Quality-Driven Super-Resolution for Less Constrained Iris Recognition at a Distance and on the MoveabstractLess constrained iris identification systems at a distance and on the move suffer from poor resolution and poor quality of the captured iris images, which significantly degrades iris recognition performance. This paper proposes a new signal-level fusion approach which incorporates a quality score into a reconstruction-based super-resolution process to generate a high-resolution iris image from a low-resolution and quality inconsistent video sequence of an eye. A novel approach for assessing the focus level of the iris image, which is invariant to lighting and oclusion conditions, is introduced. The focus score is combined with several other quality factors to perform the quality weighted super-resolution where the highest quality frames contribute the greatest amount of information to the resulting high-resolution images without introducing spurious high-frequency components. Experiments conducted on the Multiple Biometric Grand Challenge portal dataset show that our proposed approach outperforms the traditional best quality frame selection approach and other existing state-of-the-art signal-level and score-level fusion approaches for recognition of less constrained iris at a distance and on the move. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Simon Denman |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2010 | Multi-Modal Object Tracking using Dynamic Performance MetricsabstractIntelligent surveillance systems typically use a single visual spectrum modality for their input. These systems work well in controlled conditions, but often fail when lighting is poor, or environmental effects such as shadows, dust or smoke are present. Thermal spectrum imagery is not as susceptible to environmental effects, however thermal imaging sensors are more sensitive to noise and they are only gray scale, making distinguishing between objects difficult. Several approaches to combining the visual and thermal modalities have been proposed, however they are limited by assuming that both modalities are perfuming equally well. When one modality fails, existing approaches are unable to detect the drop in performance and disregard the under performing modality. In this paper, a novel middle fusion approach for combining visual and thermal spectrum images for object tracking is proposed. Motion and object detection is performed on each modality and the object detection results for each modality are fused base on the current performance of each modality. Modality performance is determined by comparing the number of objects tracked by the system with the number detected by each mode, with a small allowance made for objects entering and exiting the scene. The tracking performance of the proposed fusion scheme is compared with performance of the visual and thermal modes individually, and a baseline middle fusion scheme. Improvement in tracking performance using the proposed fusion approach is demonstrated. The proposed approach is also shown to be able to detect the failure of an individual modality and disregard its results, ensuring performance is not degraded in such situations. Simon Denman, Clinton Fookes, Sridha Sridharan, David Ryan |
AVSS | 1 |
| 2010 | Crowd Counting Using Group Tracking and Local FeaturesabstractIn public venues, crowd size is a key indicator of crowd safety and stability. In this paper we propose a crowd counting algorithm that uses tracking and local features to count the number of people in each group as represented by a foreground blob segment, so that the total crowd estimate is the sum of the group sizes. Tracking is employed to improve the robustness of the estimate, by analysing the history of each group, including splitting and merging events. A simplified ground truth annotation strategy results in an approach with minimal setup requirements that is highly accurate. David Ryan, Simon Denman, Clinton Fookes, Sridha Sridharan |
AVSS | 2 |
| 2009 | Dynamic Performance Measures for Object Tracking SystemsabstractPerformance evaluation of object tracking systems is typically performed after the data has been processed, by comparing tracking results to ground truth. Whilst this approach is fine when performing offline testing, it does not allow for real-time analysis of the systems performance, which may be of use for live systems to either automatically tune the system or report reliability. In this paper, we propose three metrics that can be used to dynamically asses the performance of an object tracking system. Outputs and results from various stages in the tracking system are used to obtain measures that indicate the performance of motion segmentation, object detection and object matching. The proposed dynamic metrics are shown to accurately indicate tracking errors when visually comparing metric results to tracking output, and are shown to display similar trends to the ETISEO metrics when comparing different tracking configurations. Simon Denman, Clinton Fookes, Sridha Sridharan, Ruan Lakemond |
AVSS | 1 |
| 2007 | An adaptive optical flow technique for person tracking systems
Simon Denman, Vinod Chandran, Sridha Sridharan |
Pattern Recognit. Lett. | 1 |
| 2006 | A Multi-Class Tracker Using a Scalable Condensation FilterabstractTracking systems are typically targeted towards tracking a single class of object. In many real world situations, and in the ETISEO evaluation, it is advantageous to be able to track multiple classes of objects. In this paper we describe the adaptation of a single class tracking system to a multi-class tracking system, and describe a modified version of the condensation filter that can be used to track all objects, of all classes. We show that by using simple targeted detectors, we can achieve accurate tracking and can accurately distinguish between classes. Simon Denman, Vinod Chandran, Sridha Sridharan, Clinton Fookes |
AVSS | 1 |
| 2006 | Multi-view Intelligent Vehicle Surveillance SystemabstractThis paper presents a multi-view intelligent surveillance system used for the automatic tracking and monitoring of vehicles in a short-term parking lane. The system has the ability to track multiple vehicles in real-time across four cameras monitoring the area using a combination of both motion detection and optical flow modules. Automated alerts of events such as parking time violations, breaching of restricted areas or improper directional flow of traffic can be generated and communicated to attending security personnel. Results are shown using surveillance data captured from a real multi-camera network to illustrate the robust and real-time performance of the system. Simon Denman, Clinton Fookes, Jamie Cook, Chris Davoren, Anthony Mamic, Graeme Farquharson, Daniel Chen 0002, Brenden Chen, Sridha Sridharan |
AVSS | 1 |