VLDB 2026 Research / reviewers in the wild / expert
Clinton Fookes
dblp:03/5529
· DBLP profile ↗
195ranked-venue papers
3as first author
69since 2021 · last 2026
0000-0002-8515-6324ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 110 · 2 first-author · 27 since 2021Artificial intelligence and machine learning · 96 · 1 first-author · 39 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 11 since 2021Security and privacy · 11 · 7 since 2021Human-computer interaction and ubiquitous computing · 7 · 3 since 2021Systems, architecture and hardware · 6 · 5 since 2021Databases, data management, data science and information retrieval · 4Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Ilov3Splat: Instance-Level Open-Vocabulary 3D Scene Understanding in Gaussian Splatting
Binh Long Nguyen, Kien Nguyen Thanh, Sridha Sridharan, Clinton Fookes, Peyman Moghadam |
ICPR (5) | 4 |
| 2026 | PDV: Prompt Directional Vectors for Zero-shot Composed Image RetrievalabstractZero-shot Composed Image Retrieval (ZS-CIR) enables image search using a reference image and a text prompt without requiring specialized text-image composition networks trained on large-scale paired data. However, current ZS-CIR approaches suffer from three critical limitations in their reliance on composed text embeddings: static query embedding representations, insufficient utilization of image embeddings, and suboptimal performance when fusing text and image embeddings. To address these challenges, we introduce the Prompt Directional Vector (PDV), a simple yet effective training-free enhancement that captures semantic modifications induced by user prompts. PDV enables three key improvements: (1) Dynamic composed text embeddings where prompt adjustments are controllable via a scaling factor, (2) composed image embeddings through semantic transfer from text prompts to image features, and (3) weighted fusion of composed text and image embeddings that enhances retrieval by balancing visual and semantic similarity. Our approach serves as a plug-and-play enhancement for existing ZS-CIR methods with minimal computational overhead. Extensive experiments across multiple benchmarks demonstrate that PDV consistently improves retrieval performance when integrated with state-of-the-art ZS-CIR approaches, particularly for methods that generate accurate compositional embeddings. The code will be released upon publication. Osman Tursun, Sinan Kalkan, Simon Denman, Clinton Fookes |
WACV | 4 |
| 2026 | Multimodal transformer-diffusion framework for large-scale reconstruction of soccer tracking data
Harry Hughes, Patrick Lucey, Michael Horton 0001, Harshala Gammulle, Clinton Fookes, Sridha Sridharan |
Comput. Vis. Image Underst. | 5 |
| 2026 | Contrastive context distillation for skeleton based early action predictionabstractEarly action prediction (EAP) requires the inference of actions from partially observed sequences, which typically contain only the initial movements of an action. Compared to using complete sequences that show an action being fully executed, EAP lacks discriminative information and increased ambiguity as different actions can contain very similar initial movements. To alleviate this problem, recent methods try to distill features learned from action recognition (AR) models, which leads to sub-optimal results due to the lack of generalizability of features across AR and EAP tasks. In a different line of work, recent studies have proposed learning discriminative class-specific features, leveraging the traditional contrastive learning approach, where samples are selected and calibrated for learning discriminative features from hard-to-classify samples. This paper proposes a novel, dynamic, and context-aware framework for EAP by combining the merits of both knowledge distillation and contrastive learning. Particularly our method distills salient discriminative context from the complete action sequence to drive the EAP, with the help of a novel dynamic contrastive learning scheme. Extensive evaluations over three public datasets demonstrate state-of-the-art performance for Early Action Prediction. Chinthaka Ranasingha, Tharindu Fernando, Harshala Gammulle, Simon Denman, Sridha Sridharan, Clinton Fookes |
Expert Syst. Appl. | 6 |
| 2026 | Enhancing predictive performance on long-tail trajectories via clustering and specialized decodersabstractAccurate forecasting of traffic participants’ future trajectories is crucial for the advancement of autonomous driving systems. Constructing robust models for such tasks requires access to comprehensive datasets that include a variety of diverse cases. Current naturalistic trajectory prediction datasets are often imbalanced, featuring a large number of easier examples and a deficiency of more challenging instances. This long-tail distribution poses a significant challenge, resulting in inadequate model performance on the rare, yet safety-critical, parts of the data. To address this issue, we have proposed a framework that utilizes an embedding-based clustering technique and a distribution-sensitive decoder module to generate precise predictions for tail samples. In addition, the proposed framework includes a trajectory clustering module to refine predictions and improve the model’s capacity to generate multiple plausible future trajectories. Experimental results show that our framework outperforms the state-of-the-art long tail prediction method on tail samples by 19.5% on the Average Displacement error (ADE) and 25.5% on the Final Displacement error (FDE). Additionally, our approach attains state-of-the-art performance in terms of ADE metric on the ETH/UCY datasets, while only slightly trailing Y-Net in terms of the FDE metric. We further conduct ablation studies to highlight the efficacy of each of the proposed innovations. Source codes are available at his GitHub repository G. Ganeshaaraj, Tharindu Fernando, Sridha Sridharan, Clinton Fookes |
Pattern Recognit. | 4 |
| 2026 | Zoom-shot: Fast, efficient and unsupervised zero-shot knowledge transfer from CLIP to vision encodersabstractFoundation models like CLIP demonstrate exceptional capabilities over a broad domain of knowledge, such as with zero-shot classification; however, they also require significant computational resources, narrowing their real-world utility. Recent studies have shown that mapping features from pre-trained vision encoders into CLIP’s latent space can transfer some of CLIP’s abilities to smaller vision encoders, offering a promising alternative. Yet, the performance of these vision encoders still falls short of CLIP’s native capabilities, particularly in low-data regimes. In this work, we argue that enhancing training data coverage/diversity significantly improves mapping efficacy. We achieve this using tailored loss functions rather than relying on data augmentation or increasing training samples. For instance, we exploit the inherent multimodal nature of CLIP’s latent space, by incorporating cycle-consistency loss as one of our loss functions. Moreover, the mapping is learned using entirely unlabelled and unpaired data, eliminating the need for manual labelling or data pairing in novel domains. From these findings, our resulting method (Zoom-shot) offers a viable path to flexible zero-shot models for resource-limited, data-scarce settings. We test Zoom-shot’s zero-shot performance across various pre-trained vision encoders on coarse- and fine-grained datasets and achieve superior performance compared to recent works. In our ablations, we find Zoom-shot allows for a trade-off between data and compute during training; allowing for a significant reduction in required training data. All code and models are available on GitHub. Jordan Shipard, Arnold Wiliem, Kien Nguyen Thanh, Wei Xiang 0001, Clinton Fookes |
Pattern Recognit. | 5 |
| 2025 | Event2Tracking: Reconstructing Multi-Agent Soccer Trajectories Using Long-Term Multimodal ContextabstractSoccer is a rich testbed for studying multi-agent adversarial systems. In this work we focus on the task of reconstructing the noisy trajectories of soccer agents (players and the ball). Previous works that model the behaviours of agents in soccer are limited in two respects: (i) they only focus on short-term context windows (less than or equal to 10 seconds) which are not suitable for reconstructing trajectories impacted by long-term noise, and (ii) they exclusively rely on trajectory context, and do not leverage soccer's auxiliary data streams that can provide additional context. Our Event2Tracking model addresses these limitations. First, our architecture models soccer's long-term structure by processing long-term trajectories (60 seconds in duration). Secondly, our architecture is multimodal. Specifically, it fuses soccer tracking data with event data (which specifies the high-level semantic events that transpire in a game), providing rich context that cannot strictly be inferred from the raw trajectories. We evaluate our method empirically using a reconstruction loss metric. Compared to state-of-the-art approaches, our method substantially improves the accuracy of the ball's and players' reconstructed trajectories. Harry Hughes, Michael Horton 0001, Xinyu Wei 0004, Harshala Gammulle, Clinton Fookes, Sridha Sridharan, Patrick Lucey |
AAAI | 5 |
| 2025 | HOTFormerLoc: Hierarchical Octree Transformer for Versatile Lidar Place Recognition Across Ground and Aerial ViewsabstractWe present HOTFormerLoc, a novel and versatile Hierarchical Octree-based TransFormer, for large-scale 3D place recognition in both ground-to-ground and ground-to-aerial scenarios across urban and forest environments. We propose an octree-based multi-scale attention mechanism that captures spatial and semantic features across granularities. To address the variable density of point distributions from spinning lidar, we present cylindrical octree attention windows to reflect the underlying distribution during attention. We introduce relay tokens to enable efficient global-local interactions and multi-scale representation learning at reduced computational cost. Our pyramid attentional pooling then synthesises a robust global descriptor for end-to-end place recognition in challenging environments. In addition, we introduce CS-WildPlaces, a novel 3D cross-source dataset featuring point cloud data from aerial and ground lidar scans captured in dense forests. Point clouds in CS-Wild-Places contain representational gaps and distinctive attributes such as varying point densities and noise patterns, making it a challenging benchmark for cross-view localisation in the wild. HOTFormerLoc achieves a top-1 average recall improvement of 5.5% – 11.5% on the CS-Wild-Places benchmark. Furthermore, it consistently outperforms SOTA 3D place recognition methods, with an average performance gain of 4.9% on well-established urban and forest datasets. The code and CS-Wild-Places benchmark is available at https://csirorobotics.github.io/HOTFormerLoc. Ethan Griffiths, Maryam Haghighat, Simon Denman, Clinton Fookes, Milad Ramezani |
CVPR | 4 |
| 2025 | AG-VPReID: A Challenging Large-Scale Benchmark for Aerial-Ground Video-based Person Re-IdentificationabstractWe introduce AG-VPReID, a new large-scale dataset for aerial-ground video-based person re-identification (ReID) that comprises 6,632 subjects, 32,321 tracklets and over 9.6 million frames captured by drones (altitudes ranging from 15–120m), CCTV, and wearable cameras. This dataset offers a real-world benchmark for evaluating the robustness to significant viewpoint changes, scale variations, and resolution differences in cross-platform aerial-ground settings. In addition, to address these challenges, we propose AG-VPReID-Net, an end-to-end framework composed of three complementary streams: (1) an Adapted Temporal-Spatial Stream addressing motion pattern inconsistencies and facilitating temporal feature learning, (2) a Normalized Appearance Stream leveraging physics-informed techniques to tackle resolution and appearance changes, and (3) a Multi-Scale Attention Stream handling scale variations across drone altitudes. We integrate visual-semantic cues from all streams to form a robust, viewpoint-invariant whole-body representation. Extensive experiments demonstrate that AG-VPReID-Net outperforms state-of-the-art approaches on both our new dataset and existing video-based ReID benchmarks, showcasing its effectiveness and generalizability. Nevertheless, the performance gap observed on AG-VPReID across all methods underscores the dataset’s challenging nature. The dataset, code and trained models are available at AG-VPReID-Net. Kien Nguyen Thanh, Akila Pemasiri, Feng Liu 0037, Sridha Sridharan, Clinton Fookes |
CVPR | 6 |
| 2025 | RadDet: A Wideband Dataset for Real-Time Radar Spectrum DetectionabstractReal-time detection of radar signals in a wideband radio frequency spectrum is a critical situational assessment function in electronic warfare. Compute-efficient detection models have shown great promise in recent years, providing an opportunity to tackle the spectrum detection problem. However, progress in radar spectrum detection is limited by the scarcity of publicly available wideband radar signal datasets accompanied by corresponding annotations. To address this challenge, we introduce a novel and challenging dataset for radar detection (RadDet), comprising a large corpus of radar signals occupying a wideband spectrum across diverse radar density environments and signal-to-noise ratio (SNR) settings. RadDet contains 40,000 frames, each generated from 1 million in-phase and quadrature (I/Q) samples across a 500 MHz frequency band. RadDet includes 11 classes of radar signals across 6 different SNR settings, 2 radar density environments, and 3 different time-frequency resolutions, with corresponding time-frequency and class annotations. We evaluate the performance of state-of-the-art real-time detection models on RadDet and a modified radar classification dataset from NIST (NIST-CBRS) to establish a novel benchmark for wideband radar spectrum detection. Zi Huang, Simon Denman, Akila Pemasiri, Terrence Martin, Clinton Fookes |
ICASSP | 5 |
| 2025 | AG-VPReID 2025: Aerial-Ground Video-based Person Re-identification Challenge ResultsabstractPerson re-identification (ReID) across aerial and ground vantage points has become crucial for large-scale surveillance and public safety applications. Although significant progress has been made in ground-only scenarios, bridging the aerial-ground domain gap remains a formidable challenge due to extreme viewpoint differences, scale variations, and occlusions. Building upon the achievements of the AG-ReID 2023 Challenge, this paper introduces the AG-VPReID 2025 Challenge—the first large-scale video-based competition focused on high-altitude (80–120 m) aerial-ground person ReID. Constructed on the new AG-VPReID dataset with 3,027 identities, over 13,500 tracklets, and approximately 3.7 million frames captured from UAVs, CCTV, and wearable cameras, the challenge featured four international teams. These teams developed solutions ranging from multi-stream architectures to transformer-based temporal reasoning and physics-informed modeling. The leading approach, X-TFCLIP from UAM, attained 72.28% Rank-1 accuracy in the aerial-to-ground ReID setting and 70.77% in the ground-to-aerial ReID setting, surpassing existing baselines while highlighting the dataset’s complexity. For additional details, please refer to the official website at https://agvpreid25.github.io. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Feng Liu 0037, Xiaoming Liu 0002, Arun Ross, Tamás Endrei, Ivan DeAndres-Tame, Ruben Tolosana, Rubén Vera-Rodríguez, Aythami Morales, Julian Fierrez, Javier Ortega-Garcia, Zijing Gong, Xuehu Liu, Md. Rashidunnabi, Hugo Proença 0001, Kailash A. Hambarde, Saeid Rezaei |
IJCB | 2 |
| 2025 | AG-VPReID.VIR: Bridging Aerial and Ground Platforms for Video-based Visible-Infrared Person Re-IDabstractPerson re-identification (Re-ID) across visible and infrared modalities is crucial for 24-hour surveillance systems, but existing datasets primarily focus on ground-level perspectives. While ground-based IR systems offer nighttime capabilities, they suffer from occlusions, limited coverage, and vulnerability to obstructions—problems that aerial perspectives uniquely solve. To address these limitations, we introduce AG-VPReID.VIR, the first aerial-ground cross-modality video-based person Re-ID dataset. This dataset captures 1,837 identities across 4,861 tracklets (124,855 frames) using both UAV-mounted and fixed CCTV cameras in RGB and infrared modalities. AG-VPReID.VIR presents unique challenges including cross-viewpoint variations, modality discrepancies, and temporal dynamics. Additionally, we propose TCC-VPReID, a novel three-stream architecture designed to address the joint challenges of cross-platform and cross-modality person Re-ID. Our approach bridges the domain gaps between aerial-ground perspectives and RGB-IR modalities, through style-robust feature learning, memory-based cross-view adaptation, and intermediary-guided temporal modeling. Experiments show that AG-VPReID.VIR presents distinctive challenges compared to existing datasets, with our TCC-VPReID framework achieving significant performance gains across multiple evaluation protocols. Dataset and code are available at https://github.com/agvpreid25/AG-VPReID.VIR. Kien Nguyen Thanh, Akila Pemasiri, Akmal Jahan, Clinton Fookes, Sridha Sridharan |
IJCB | 5 |
| 2025 | MTL-DFM: Multi-Task Learning and Diffusion Model for ISAC SystemsabstractDeep learning (DL) has emerged as a key enabler for unlocking the potential of integrated sensing and communication (ISAC). Despite recent progress, current DL methods primarily handle sensing and communication as independent tasks, overlooking potential performance enhancement through a joint approach. Moreover, existing methods rely on fully annotated data for training, which is often challenging to obtain, especially in multi-task scenarios where labeled data may be scarce or only exist for a subset of tasks. Motivated by these shortcomings, this paper proposes a novel scheme, MTL-DFM, to enable simultaneous sensing and communication with partially labeled training data, which leverages multi-task learning (MTL) and a diffusion model (DFM). In particular, we introduce an initial feature extraction module (IFEM) to jointly capture shared information across tasks and explore inherent cross-task connections for enhanced feature extraction. Next, we design a signal denoising with incomplete labeling (SDIL) module to effectively remove noise from extracted information and construct comprehensive feature representations for all tasks with partially labeled datasets, which is difficult for conventional DL methods. Simulation results verify the superior performance offered by MTL-DFM over prior state-of-the-art methods. Qingqing Cheng, Zhenguo Shi, Simon Denman, Clinton Fookes, Jinhong Yuan, Derrick Wing Kwan Ng |
ICC | 4 |
| 2025 | Improving the Generation of VAEs with High Dimensional Latent Spaces by the use of Hyperspherical CoordinatesabstractVariational autoencoders (VAE) encode data into lower-dimensional latent vectors before decoding those vectors back to data. Once trained, decoding a random latent vector from the prior usually does not produce meaningful data, at least when the latent space has more than a dozen dimensions. In this paper, we investigate this issue by drawing insight from high dimensional statistics: in these regimes, the latent vectors of a standard VAE are by construction distributed uniformly on a hypersphere. We propose to formulate the latent variables of a VAE using hyperspherical coordinates, which allows compressing the latent vectors towards an island on the hypersphere, thereby reducing the latent sparsity and we show that this improves the generation ability of the VAE. We propose a new parameterization of the latent space with limited computational overhead. Alejandro Ascarate, Léo Lebrat, Rodrigo Santa Cruz, Clinton Fookes, Olivier Salvado |
IJCNN | 4 |
| 2025 | Online 6DoF Global Localisation in Forests using Semantically-Guided Re-Localisation and Cross-View Factor-Graph OptimisationabstractThis paper presents FGLoc6D, a novel approach for robust global localisation and online 6DoF pose estimation of ground robots in forest environments by leveraging deep semantically-guided re-localisation and cross-view factor graph optimisation. The proposed method addresses the challenges of aligning aerial and ground data for pose estimation, which is crucial for accurate point-to-point navigation in GPS-degraded environments. By integrating information from both perspectives into a factor graph framework, our approach effectively estimates the robot’s global position and orientation. Additionally, we enhance the repeatability of deep-learned keypoints for metric localisation in forests by incorporating a semantically-guided regression loss. This loss encourages greater attention to wooden structures, e.g., tree trunks, which serve as stable and distinguishable features, thereby improving the consistency of keypoints and increasing the success rate of global registration, a process we refer to as re-localisation. The re-localisation module along with the factor-graph structure, populated by odometry and ground-to-aerial factors over time, allows global localisation under dense canopies. We validate the performance of our method through extensive experiments in three forest scenarios, demonstrating its global localisation capability and superiority over alternative state-of-the-art in terms of accuracy and robustness in these challenging environments. Experimental results show that our proposed method can achieve drift-free localisation with bounded positioning errors, ensuring reliable and safe robot navigation through dense forests. Lucas Carvalho de Lima, Ethan Griffiths, Maryam Haghighat, Simon Denman, Clinton Fookes, Paulo Vinicius Koerich Borges, Michael Brünig, Milad Ramezani |
IROS | 5 |
| 2025 | SALVE: A 3D Reconstruction Benchmark of Wounds from Consumer-Grade VideosabstractManaging chronic wounds is a global challenge that can be alleviated by the adoption of automatic systems for clinical wound assessment from consumer-grade videos. While 2D image analysis approaches are insufficient for handling the 3D features of wounds, existing approaches utilizing 3D reconstruction methods have not been thor-oughly evaluated. To address this gap, this paper presents a comprehensive study on 3D wound reconstruction from consumer-grade videos. Specifically, we introduce the SALVE dataset, comprising video recordings of realistic wound phantoms captured with different cameras. Using this dataset, we assess the accuracy and precision of state-of-the-art methods for 3D reconstruction, ranging from traditional photogrammetry pipelines to advanced neural rendering approaches. In our experiments, we observe that photogrammetry approaches do not provide smooth surfaces suitable for precise clinical measurements of wounds. Neural rendering approaches show promise in addressing this issue, advancing the use of this technology in wound care practices. We encourage the readers to visit the project page: https://rcmichierchia.github.io/SALVE/. Remi Chierchia, Léo Lebrat, David Ahmedt-Aristizabal, Olivier Salvado, Clinton Fookes, Rodrigo Santa Cruz |
WACV | 5 |
| 2025 | Beyond geometry: The power of texture in interpretable 3D person ReID
Kien Nguyen Thanh, Akila Pemasiri, Sridha Sridharan, Clinton Fookes |
Comput. Vis. Image Underst. | 5 |
| 2025 | A survey on physics informed reinforcement learning: Review and open problemsabstractThe fusion of physical information in machine learning frameworks has revolutionized many application areas. This involves enhancing the learning process by incorporating physical constraints and adhering to physical laws. This work explores their utility for reinforcement learning applications. A thorough review of the literature on the fusion of physics information or physics priors in reinforcement learning approaches, commonly referred to as physics-informed reinforcement learning (PIRL), is presented. A novel taxonomy is introduced with the reinforcement learning pipeline as the backbone to classify existing works, compare and contrast them, and derive crucial insights. Existing works are analyzed with regard to the representation/form of the governing physics modeled for integration, their specific contribution to the typical reinforcement learning architecture, and their connection to the underlying reinforcement learning pipeline stages. Core learning architectures and physics incorporation biases (i.e., observational, inductive, and learning) of existing PIRL approaches are identified and used to further categorize the works for better understanding and adaptation. By providing a comprehensive perspective on the implementation of the physics-informed capability, the taxonomy presents a cohesive approach to PIRL. It identifies the areas where this approach has been applied, as well as the gaps and opportunities that exist. Additionally, the review highlights unresolved issues and challenges, while also incorporating potential and emerging solutions to guide future research. This nascent field holds great potential for enhancing reinforcement learning algorithms by increasing their physical plausibility, precision, data efficiency, and applicability in real-world scenarios. Chayan Banerjee, Kien Nguyen Thanh, Clinton Fookes, Maziar Raissi |
Expert Syst. Appl. | 3 |
| 2025 | Remembering What is Important: A Factorised Multi-Head Retrieval and Auxiliary Memory Stabilisation Scheme for Human Motion PredictionabstractHumans exhibit complex motions that vary depending on the activity they are performing, the interactions they engage in, as well as subject-specific preferences. Therefore, forecasting a human's future pose based on the history of his or her previous motion is a challenging task. This paper presents an innovative auxiliary-memory-powered deep neural network framework to improve the modelling of historical knowledge. Specifically, we disentangle subject-specific, action-specific, and other auxiliary information from the observed pose sequences and utilise these factorised features to query the memory. A novel Multi-Head knowledge retrieval scheme leverages these factorised feature embeddings to perform multiple querying operations over the historical observations captured within the auxiliary memory. Moreover, we propose a dynamic masking strategy to make this feature disentanglement process adaptive. Two novel loss functions are introduced to encourage diversity within the auxiliary memory, while ensuring the stability of the memory content such that it can locate and store salient information that aids the long-term prediction of future motion, irrespective of any data imbalances or the diversity of the input data distribution. Extensive experiments conducted on two public benchmarks, Human3.6M and CMU-Mocap, demonstrate that these design choices collectively allow the proposed approach to outperform the current state-of-the-art methods by significant margins: 17% on the Human3.6M dataset and 9% on the CMU-Mocap dataset. Tharindu Fernando, Harshala Gammulle, Sridha Sridharan, Simon Denman, Clinton Fookes |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | FICE: Text-conditioned fashion-image editing with guided GAN inversionabstractFashion-image editing is a challenging computer-vision task where the goal is to incorporate selected apparel into a given input image. Most existing techniques, known as Virtual Try-On methods, deal with this task by first selecting an example image of the desired apparel and then transferring the clothing onto the target person. Conversely, in this paper, we consider editing fashion images with text descriptions. Such an approach has several advantages over example-based virtual try-on techniques: (i) it does not require an image of the target fashion item, and (ii) it allows the expression of a wide variety of visual concepts through the use of natural language. Existing image-editing methods that work with language inputs are heavily constrained by their requirement for training sets with rich attribute annotations or they are only able to handle simple text descriptions. We address these constraints by proposing a novel text-conditioned editing model called FICE (Fashion Image CLIP Editing) that is capable of handling a wide variety of diverse text descriptions to guide the editing procedure. Specifically, with FICE, we extend the common GAN-inversion process by including semantic, pose-related, and image-level constraints when generating images. We leverage the capabilities of the CLIP model to enforce the text-provided semantics, due to its impressive image–text association capabilities. We furthermore propose a latent-code regularization technique that provides the means to better control the fidelity of the synthesized images. We validate the FICE through rigorous experiments on a combination of VITON images and Fashion-Gen text descriptions and in comparison with several state-of-the-art, text-conditioned, image-editing approaches. Experimental results demonstrate that the FICE generates very realistic fashion images and leads to better editing than existing, competing approaches. The source code is publicly available from: https://github.com/MartinPernus/FICE . • We propose a novel text-conditioned fashion image editing model called FICE. • FICE is an extended GAN inversion method tailored towards the fashion domain. • Descriptions are integrated with the GAN inversion process to manipulate appearance. • Comprehensive experiments show that FICE obtains state-of-the-art results. Martin Pernus, Clinton Fookes, Vitomir Struc, Simon Dobrisek |
Pattern Recognit. | 2 |
| 2025 | Decoupled and Explainable Associative Memory for Effective Knowledge PropagationabstractLong-term memory often plays a pivotal role in human cognition through the analysis of contextual information. Machine learning researchers have attempted to emulate this process through the development of memory-augmented neural networks (MANNs) to leverage indirectly related but resourceful historical observations during learning and inference. The area of MANN, however, is still in its infancy and significant research effort is required to enable machines to achieve performance close to the human cognition process. This article presents an innovative MANN framework for the advanced incorporation of historical knowledge into a predictive framework. Within the key-value memory structure, we propose to decouple the key representations from the learned value memory embeddings to offer improved associations between the inputs and latent memory embeddings. We argue that the keys should be static, sparse, and unique representations of a particular observation to offer robust input to memory associations, while the value embeddings could be trainable, dense latent vectors such that they can better capture historical knowledge. Moreover, we introduce a novel memory update procedure that preserves the explainability of the historical knowledge extraction process, which would enable the human end-users to interpret the deep machine learning model decisions, fostering their trust. With extensive experiments conducted on three different datasets using audio, text, and image modalities, we demonstrate that our proposed innovations collectively allow this framework to outperform the current state-of-the-art methods by significant margins, irrespective of the modalities or the downstream tasks. The code is available at https://github.com/tha725/DE-KVMN/tree/main. Tharindu Fernando, Darshana Priyasad, Sridha Sridharan, Clinton Fookes |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | NeRF Director: Revisiting View Selection in Neural Volume RenderingabstractNeural Rendering representations have significantly contributed to the field of 3D computer vision. Given their potential, considerable efforts have been invested to improve their performance. Nonetheless, the essential question of selecting training views is yet to be thoroughly investigated. This key aspect plays a vital role in achieving high-quality results and aligns with the well-known tenet of deep learning: “garbage in, garbage out”. In this paper, we first illustrate the importance of view selection by demonstrating how a simple rotation of the test views within the most pervasive NeRF dataset can lead to consequential shifts in the performance rankings of state-of-the-art techniques. To address this challenge, we introduce a unified framework for view selection methods and devise a thorough benchmark to assess its impact. Significant improvements can be achieved without leveraging error or uncertainty estimation but focusing on uniform view coverage of the reconstructed object, resulting in a training-free approach. Using this technique, we show that high-quality renderings can be achieved faster by using fewer views. We conduct extensive experiments on both synthetic datasets and realistic data to demonstrate the effectiveness of our proposed method compared with random, conventional error-based, and uncertainty-guided view selection. Wenhui Xiao, Rodrigo Santa Cruz, David Ahmedt-Aristizabal, Olivier Salvado, Clinton Fookes, Léo Lebrat |
CVPR | 5 |
| 2024 | Multi-Stage Learning for Radar Pulse Activity SegmentationabstractRadio signal recognition is a crucial function in electronic warfare. Precise identification and localisation of radar pulse activities are required by electronic warfare systems to produce effective countermeasures. Despite the importance of these tasks, deep learning-based radar pulse activity recognition methods have remained largely underexplored. While deep learning for radar modulation recognition has been explored previously, classification tasks are generally limited to short and non-interleaved IQ signals, limiting their applicability to military applications. To address this gap, we introduce an end-to-end multi-stage learning approach to detect and localise pulse activities of interleaved radar signals across an extended time horizon. We propose a simple, yet highly effective multi-stage architecture for incrementally predicting fine-grained segmentation masks that localise radar pulse activities across multiple channels. We demonstrate the performance of our approach against several reference models on a novel radar dataset, while also providing a first-of-its-kind benchmark for radar pulse activity segmentation. Zi Huang, Akila Pemasiri, Simon Denman, Clinton Fookes, Terrence Martin |
ICASSP | 4 |
| 2024 | Deep cross-domain transfer for emotion recognition via joint learningabstractAbstract Deep learning has been applied to achieve significant progress in emotion recognition from multimedia data. Despite such substantial progress, existing approaches are hindered by insufficient training data, leading to weak generalisation under mismatched conditions. To address these challenges, we propose a learning strategy which jointly transfers emotional knowledge learnt from rich datasets to source-poor datasets. Our method is also able to learn cross-domain features, leading to improved recognition performance. To demonstrate the robustness of the proposed learning strategy, we conducted extensive experiments on several benchmark datasets including eNTERFACE, SAVEE, EMODB, and RAVDESS. Experimental results show that the proposed method surpassed existing transfer learning schemes by a significant margin. Dung Nguyen 0001, Duc Thanh Nguyen, Sridha Sridharan, Mohamed Almorsy, Simon Denman, Son N. Tran, Clinton Fookes |
Multim. Tools Appl. | 8 |
| 2024 | FactoFormer: Factorized Hyperspectral Transformers With Self-Supervised PretrainingabstractHyperspectral images (HSIs) contain rich spectral and spatial information. Motivated by the success of transformers in the field of natural language processing and computer vision where they have shown the ability to learn long-range dependencies within input data, recent research has focused on using transformers for HSIs. However, current state-of-the-art hyperspectral transformers only tokenize the input HSI sample along the spectral dimension, resulting in the underutilization of spatial information. Moreover, transformers are known to be data-hungry and their performance relies heavily on large-scale pretraining, which is challenging due to limited annotated hyperspectral data. Therefore, the full potential of HSI transformers has not been fully realized. To overcome these limitations, we propose a novel factorized spectral–spatial transformer that incorporates factorized self-supervised pretraining procedures, leading to significant improvements in performance. The factorization of the inputs allows the spectral and spatial transformers to better capture the interactions within the hyperspectral data cubes. Inspired by masked image modeling (MIM) pretraining, we also devise efficient masking strategies for pretraining each of the spectral and spatial transformers. We conduct experiments on six publicly available datasets for the HSI classification task and demonstrate that our model achieves state-of-the-art performance in all the datasets. The code for our model will be made available athttps://github.com/csiro-robotics/FactoFormer. Shaheer Mohamed, Maryam Haghighat, Tharindu Fernando, Sridha Sridharan, Clinton Fookes, Peyman Moghadam |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | AG-ReID.v2: Bridging Aerial and Ground Views for Person Re-IdentificationabstractAerial-ground person re-identification (Re-ID) presents unique challenges in computer vision, stemming from the distinct differences in viewpoints, poses, and resolutions between high-altitude aerial and ground-based cameras. Existing research predominantly focuses on ground-to-ground matching, with aerial matching less explored due to a dearth of comprehensive datasets. To address this, we introduce AG-ReID.v2, a dataset specifically designed for person Re-ID in mixed aerial and ground scenarios. This dataset comprises 100,502 images of 1,615 unique individuals, each annotated with matching IDs and 15 soft attribute labels. Data were collected from diverse perspectives using a UAV, stationary CCTV, and smart glasses-integrated camera, providing a rich variety of intra-identity variations. Additionally, we have developed an explainable attention network tailored for this dataset. This network features a three-stream architecture that efficiently processes pairwise image distances, emphasizes key top-down features, and adapts to variations in appearance due to altitude differences. Comparative evaluations demonstrate the superiority of our approach over existing baselines. We plan to release the dataset and algorithm source code publicly, aiming to advance research in this specialized field of computer vision. Kien Nguyen Thanh, Sridha Sridharan, Clinton Fookes |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | Physical Adversarial Attacks for Surveillance: A SurveyabstractModern automated surveillance techniques are heavily reliant on deep learning methods. Despite the superior performance, these learning systems are inherently vulnerable to adversarial attacks-maliciously crafted inputs that are designed to mislead, or trick, models into making incorrect predictions. An adversary can physically change their appearance by wearing adversarial t-shirts, glasses, or hats or by specific behavior, to potentially avoid various forms of detection, tracking, and recognition of surveillance systems; and obtain unauthorized access to secure properties and assets. This poses a severe threat to the security and safety of modern surveillance systems. This article reviews recent attempts and findings in learning and designing physical adversarial attacks for surveillance applications. In particular, we propose a framework to analyze physical adversarial attacks and provide a comprehensive survey of physical adversarial attacks on four key surveillance tasks: detection, identification, tracking, and action recognition under this framework. Furthermore, we review and analyze strategies to defend against physical adversarial attacks and the methods for evaluating the strengths of the defense. The insights in this article present an important step in building resilience within surveillance systems to physical adversarial attacks. Kien Nguyen Thanh, Tharindu Fernando, Clinton Fookes, Sridha Sridharan |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Bias Identification with RankPix SaliencyabstractSaliency methods are critical tools that allow the estimation of the most important features of an input image that contribute to the network’s prediction. These tools are pivotal in high-stakes applications such as medical diagnosis or autonomous driving. Additionally, these tools can help identify models’ biasedness, such as a strong prior on object placement, easily distinguishable background features, or frequent object co-occurrence. We introduce RankPix, a novel saliency method for visual bias identification in image classification tasks. RankPix is a derivative-free approach that allows the identification of a minimum subset of pixels/features at a given network layer that changes the output of a classifier. Surprisingly, this approaches provides equivalent performance to gradient-based approaches on the standard pointing game benchmark. More interestingly, RankPix outperforms traditional approaches for systematic bias identification. Salamata Konate, Léo Lebrat, Rodrigo Santa Cruz, Clinton Fookes, Andrew P. Bradley, Olivier Salvado |
ICASSP | 4 |
| 2023 | AG-ReID 2023: Aerial-Ground Person Re-identification Challenge ResultsabstractPerson re-identification (Re-ID) on aerial-ground platforms has emerged as an intriguing topic within computer vision, presenting a plethora of unique challenges. Highflying altitudes of aerial cameras make persons appear differently in terms of viewpoints, poses, and resolution compared to the images of the same person viewed from ground cameras. Despite its potential, few algorithms have been developed for person re-identification on aerial-ground data, mainly due to the absence of comprehensive datasets. In response, we have collected a large-scale dataset and organized the Aerial-Ground person Re-IDentification Challenge (AG-ReID2023) to foster advancements in the field. The dataset comprises 100,502 images with 1,615 unique identities, including 51,530 training images featuring 807 identities. The test set is divided into two subsets: Aerial to Ground (808 ids, 4,348 query images, 19,259 gallery images) and Ground to Aerial (808 ids, 4,151 query images, 21,214 gallery images). In addition, we manually annotate individuals with their matching IDs across cameras and provide 15 soft attribute labels. The AG-ReID2023 Challenge in conjunction with the 7thIEEE International Joint Conference on Biometrics (IJCB) has garnered interest from numerous institutes, resulting in the submission of five distinct algorithms. We provide an in-depth examination of the evaluation outcomes and present our findings from the contest. For additional details, kindly refer to the official website1.1https://agreid23.github.io. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Feng Liu 0037, Xiaoming Liu 0002, Arun Ross, Dana Michalski, Debayan Deb, Mahak Kothari, Manisha Saini, Dawei Du, Scott McCloskey, Gabriel Bertocco, Fernanda A. Andaló, Terrance E. Boult, Anderson Rocha 0001, Haidong Zhu, Zhaoheng Zheng, Ramakant Nevatia, Zaigham A. Randhawa, Sinan Sabri, Gianfranco Doretto |
IJCB | 2 |
| 2023 | Aerial-Ground Person Re-IDabstractPerson re-ID matches persons across multiple non-overlapping cameras. Despite the increasing deployment of air-borne platforms in surveillance, current existing person re-ID benchmarks’ focus is on ground-ground matching and very limited efforts on aerial-aerial matching. We propose a new benchmark dataset - AG-ReID, which performs person re-ID matching in a new setting: across aerial and ground cameras. Our dataset contains 21,983 images of 388 identities and 15 soft attributes for each identity. The data was collected by a UAV flying at altitudes between 15 to 45 meters and a ground-based CCTV camera on a university campus. Our dataset presents a novel elevated-viewpoint challenge for person re-ID due to the significant difference in person appearance across these cameras. We propose an explainable algorithm to guide the person re-ID model’s training with soft attributes to address this challenge. Experiments demonstrate the efficacy of our method on the aerial-ground person re-ID task. The dataset will be published and the baseline codes will be open-sourced at https://github.com/huynguyen792/AG-ReID to facilitate research in this area. Kien Nguyen Thanh, Sridha Sridharan, Clinton Fookes |
ICME | 4 |
| 2023 | Wild-Places: A Large-Scale Dataset for Lidar Place Recognition in Unstructured Natural EnvironmentsabstractMany existing datasets for lidar place recognition are solely representative of structured urban environments, and have recently been saturated in performance by deep learning based approaches. Natural and unstructured environments present many additional challenges for the tasks of long-term localisation but these environments are not represented in currently available datasets. To address this we introduce Wild-Places, a challenging large-scale dataset for lidar place recognition in unstructured, natural environments. Wild-Places contains eight lidar sequences collected with a handheld sensor payload over the course of fourteen months, containing a total of 63K undistorted lidar submaps along with accurate 6DoF ground truth. This dataset contains multi-ple revisits both within and between sequences, allowing for both intra-sequence (i.e., loop closure detection) and inter-sequence (i.e., re-localisation) tasks. We also benchmark several state-of-the-art approaches to demonstrate the challenges that this dataset introduces, particularly the case of long-term place recognition due to natural environments changing over time. Our dataset and code is available at https://csiro-robotics.github.io/Wild-Places Joshua Knights, Kavisha Vidanapathirana, Milad Ramezani, Sridha Sridharan, Clinton Fookes, Peyman Moghadam |
ICRA | 5 |
| 2023 | Dual Memory Fusion for Multimodal Speech Emotion Recognition
Darshana Prisayad, Tharindu Fernando, Sridha Sridharan, Simon Denman, Clinton Fookes |
INTERSPEECH | 5 |
| 2023 | Piecewise Deterministic Markov Processes for Bayesian Neural NetworksabstractInference on modern Bayesian Neural Networks (BNNs) often relies on a variational inference treatment, imposing violated assumptions of independence and the form of the posterior. Traditional MCMC approaches avoid these assumptions at the cost of increased computation due to its incompatibility to subsampling of the likelihood. New Piecewise Deterministic Markov Process (PDMP) samplers permit subsampling, though introduce a model-specific inhomogenous Poisson Process (IPPs) which is difficult to sample from. This work introduces a new generic and adaptive thinning scheme for sampling from these IPPs, and demonstrates how this approach can accelerate the application of PDMPs for inference in BNNs. Experimentation illustrates how inference with these methods is computationally feasible, can improve predictive accuracy, MCMC mixing performance, and provide informative uncertainty measurements when compared against other approximate inference schemes. Ethan Goan, Dimitri Perrin, Kerrie L. Mengersen, Clinton Fookes |
UAI | 4 |
| 2023 | DBCE : A Saliency Method for Medical Deep Learning Through Anatomically-Consistent Free-Form DeformationsabstractDeep learning models are powerful tools for addressing challenging medical imaging problems. However, for an ever-growing range of applications, interpreting a model’s prediction remains non-trivial. Understanding decisions made by black-box algorithms is critical, and assessing their fairness and susceptibility to bias is a key step towards healthcare deployment. In this paper, we propose DBCE (Deformation Based Counterfactual Explainability). We optimise a diffeomorphic transformation that deforms a given input image to change the prediction of the model. This provides anatomically meaningful saliency maps indicating tissue atrophy and expansion, which can be easily interpreted by clinicians. In our test case, DBCE replicates the transition of a patient from healthy control (HC) to Alzheimer’s disease (AD). We benchmark DBCE against three commonly used saliency methods. We show that it provides more meaningful saliency maps when applied to one subject and disease-consistent atrophy patterns when used over a larger cohort. In addition, our method fulfils a recent sanity check and is repeatable for different model initialisations in contrast to classical sensitivity-based methods. Joshua Peters, Léo Lebrat, Rodrigo Santa Cruz, Aaron Nicolson, Gregg Belous, Salamata Konate, Parnesh Raniga, Vincent Doré, Pierrick Bourgeat, Jurgen Fripp, Clinton Fookes, Olivier Salvado |
WACV | 11 |
| 2023 | Meta-transfer learning for emotion recognitionabstractAbstract Deep learning has been widely adopted in automatic emotion recognition and has lead to significant progress in the field. However, due to insufficient training data, pre-trained models are limited in their generalisation ability, leading to poor performance on novel test sets. To mitigate this challenge, transfer learning performed by fine-tuning pr-etrained models on novel domains has been applied. However, the fine-tuned knowledge may overwrite and/or discard important knowledge learnt in pre-trained models. In this paper, we address this issue by proposing a PathNet-based meta-transfer learning method that is able to (i) transfer emotional knowledge learnt from one visual/audio emotion domain to another domain and (ii) transfer emotional knowledge learnt from multiple audio emotion domains to one another to improve overall emotion recognition accuracy. To show the robustness of our proposed method, extensive experiments on facial expression-based emotion recognition and speech emotion recognition are carried out on three bench-marking data sets: SAVEE, EMODB, and eNTERFACE. Experimental results show that our proposed method achieves superior performance compared with existing transfer learning methods. Dung Nguyen 0001, Duc Thanh Nguyen, Sridha Sridharan, Simon Denman, Thanh Thi Nguyen 0001, David Dean, Clinton Fookes |
Neural Comput. Appl. | 7 |
| 2023 | Complex-Valued Iris Recognition NetworkabstractIn this work, we design a fully complex-valued neural network for the task of iris recognition. Unlike the problem of general object recognition, where real-valued neural networks can be used to extract pertinent features, iris recognition depends on the extraction of both phase and magnitude information from the input iris texture in order to better represent its biometric content. This necessitates the extraction and processing of phase information that cannot be effectively handled by a real-valued neural network. In this regard, we design a fully complex-valued neural network that can better capture the multi-scale, multi-resolution, and multi-orientation phase and amplitude features of the iris texture. We show a strong correspondence of the proposed complex-valued iris recognition network with Gabor wavelets that are used to generate the classical IrisCode; however, the proposed method enables a new capability of automatic complex-valued feature learning that is tailored for iris recognition. We conduct experiments on three benchmark datasets - ND-CrossSensor-2013, CASIA-Iris-Thousand and UBIRIS.v2 - and show the benefit of the proposed network for the task of iris recognition. We exploit visualization schemes to convey how the complex-valued network, when compared to standard real-valued networks, extracts fundamentally different features from the iris texture. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Arun Ross |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Multi-stage stacked temporal convolution neural networks (MS-S-TCNs) for biosignal segmentation and anomaly localizationabstractIn the computer vision domain, temporal convolution networks (TCN) have gained traction due to their lightweight, robust architectures for sequence-to-sequence prediction tasks. With that insight, in this study, we propose a novel deep learning architecture for biosignal segmentation and anomaly localization based on TCNs , named the multi-stage stacked TCN, which employs multiple TCN modules with varying dilation factors. More precisely, for each stage, our architecture uses TCN modules with multiple dilation factors, and we use convolution-based fusion to combine predictions returned from each stage. Furthermore, aiming smoothed predictions, we introduce a novel loss function based on the first-order derivative. To demonstrate the robustness of our architecture, we evaluate our model on five different tasks related to three 1D biosignal modalities (heart sounds, lung sounds and electrocardiogram). Our proposed framework achieves state-of-the-art performance for all tasks, significantly outperforming the respective state-of-the-art models having F1 score gains up to ≈ 9 %. Furthermore, the framework demonstrates competitive performance gains compared to traditional multi-stage TCN models with similar configurations yielding F1 score gains up to ≈ 5 %. Our model is also interpretable. Using neural conductance, we demonstrate the effectiveness of having TCNs with varying dilation factors. Our visualizations show that the model benefits from feature maps captured at multiple dilation factors, and the information is effectively propagated through the network such that the final stage produces the most accurate result. Theekshana Dissanayake, Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
Pattern Recognit. | 5 |
| 2023 | Pose-driven attention-guided image generation for person re-Identification
Amena Khatun, Simon Denman, Sridha Sridharan, Clinton Fookes |
Pattern Recognit. | 4 |
| 2023 | Toward On-Board Panoptic Segmentation of Multispectral Satellite ImagesabstractWith tremendous advancements in low-power embedded computing devices and remote sensing instruments, the traditional satellite image processing pipeline which includes an expensive data transfer step prior to processing data on the ground is being replaced by on- board processing of captured data. This paradigm shift enables critical and time-sensitive intelligence to be acquired in a timely manner on- board the satellite itself. However, at present, the on- board processing of multispectral satellite images is limited to classification and segmentation tasks. Extending this processing to the next logical level, we take the first step toward on- board panoptic segmentation of multispectral satellite images and evaluate the applicability of state-of-the-art panoptic segmentation models to an on- board setting. Panoptic segmentation offers major economic and environmental insights, ranging from yield estimation from agricultural lands to intelligence for complex military applications. Nevertheless, the on- board intelligence extraction poses several challenges due to the loss of temporal observations and the need to generate predictions from a single sample. To address this challenge, we propose a multimodal teacher network with a cross modality attention-based fusion strategy to improve segmentation accuracy by exploiting data from multiple modes. We also propose an online knowledge distillation framework to transfer the knowledge learned by this multimodal teacher network to a unimodal student, which receives only a single frame input, and is more appropriate for an on- board environment. We benchmark our approach against existing state-of-the-art panoptic segmentation models using the PASTIS multispectral panoptic segmentation dataset considering an on- board processing setting. Our evaluations demonstrate a substantial 10.7%, 11.9%, and 10.6% increase in segmentation quality (SQ), recognition quality (RQ), and panoptic quality (PQ) metrics compared to the existing state-of-the-art model when it is evaluated in an on- board processing setting. Tharindu Fernando, Clinton Fookes, Harshala Gammulle, Simon Denman, Sridha Sridharan |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Generalized Generative Deep Learning Models for Biosignal Synthesis and Modality TransferabstractGenerative Adversarial Networks (GANs) are a revolutionary innovation in machine learning that enables the generation of artificial data. Artificial data synthesis is valuable especially in the medical field where it is difficult to collect and annotate real data due to privacy issues, limited access to experts, and cost. While adversarial training has led to significant breakthroughs in the computer vision field, biomedical research has not yet fully exploited the capabilities of generative models for data generation, and for more complex tasks such as biosignal modality transfer. We present a broad analysis on adversarial learning on biosignal data. Our study is the first in the machine learning community to focus on synthesizing 1D biosignal data using adversarial models. We consider three types of deep generative adversarial networks: a classical GAN, an adversarial AE, and a modality transfer GAN; individually designed for biosignal synthesis and modality transfer purposes. We evaluate these methods on multiple datasets for different biosignal modalites, including phonocardiogram (PCG), electrocardiogram (ECG), vectorcardiogram and 12-lead electrocardiogram. We follow subject-independent evaluation protocols, by evaluating the proposed models' performance on completely unseen data to demonstrate generalizability. We achieve superior results in generating biosignals, specifically in conditional generation, by synthesizing realistic samples while preserving domain-relevant characteristics. We also demonstrate insightful results in biosignal modality transfer that can generate expanded representations from fewer input-leads, ultimately making the clinical monitoring setting more convenient for the patient. Furthermore our longer duration ECGs generated, maintain clear ECG rhythmic regions, which has been proven using ad-hoc segmentation models. Theekshana Dissanayake, Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
IEEE J. Biomed. Health Informatics | 5 |
| 2022 | SESS: Saliency Enhancing with Scaling and Sliding
Osman Tursun, Simon Denman, Sridha Sridharan, Clinton Fookes |
ECCV (12) | 4 |
| 2022 | LoGG3D-Net: Locally Guided Global Descriptor Learning for 3D Place RecognitionabstractRetrieval-based place recognition is an efficient and effective solution for re-localization within a pre-built map, or global data association for Simultaneous Localization and Mapping (SLAM). The accuracy of such an approach is heavily dependant on the quality of the extracted scene-level representation. While end-to-end solutions - which learn a global descriptor from input point clouds - have demonstrated promising results, such approaches are limited in their ability to enforce desirable properties at the local feature level. In this paper, we introduce a local consistency loss to guide the network towards learning local features which are consistent across revisits, hence leading to more repeatable global descriptors resulting in an overall improvement in 3D place recognition performance. We formulate our approach in an end-to-end trainable architecture called LoGG3D-Net. Experiments on two large-scale public benchmarks (KITTI and MulRan) show that our method achieves mean F1maxscores of 0.939 and 0.968 on KITTI and MulRan respectively, achieving state-of-the-art performance while operating in near real-time. The open-source implementation is available at: https://github.com/csiro-robotics/LoGG3D-Net. Kavisha Vidanapathirana, Milad Ramezani, Peyman Moghadam, Sridha Sridharan, Clinton Fookes |
ICRA | 5 |
| 2022 | Detecting Heart Failure Through Voice Analysis using Self-Supervised Mode-Based Memory FusionabstractCongestive Heart Failure (CHF) is a progressive disease that affects millions of people worldwide, severely impacting their quality of life. Missed detection of CHF and its progression affects life expectancy, thus it is critical to develop applications to continuously monitor CHF symptoms and disease progression in a patient-centric and cost-effective manner. This paper focuses on a novel non-invasive technique to identify CHF using patients' speech traits. Pulmonary congestion and breathlessness is the most common symptom of heart failure and one of the major contributors to hospitalisation. Since pulmonary congestion results in impairment of a patient's voice, we propose a novel, non invasive method for monitoring CHF through analysis of the patient's speech. We also introduce a new balanced dataset, containing voice recordings from both healthy participants and participants diagnosed with CHF, which contains voice alterations reflective of CHF status. We propose a novel deep machine learning architecture based on mode driven memory fusion for CHF recognition from audio recordings of subject's speech. We have achieved 90% accuracy under a subject-independent evaluation setting, highlighting the applicability of such methods for tele-health and home monitoring applications. Darshana Priyasad, Andi Partovi, Sridha Sridharan, Maryam Kashefpoor, Tharindu Fernando, Simon Denman, Clinton Fookes, David Kaye |
INTERSPEECH | 7 |
| 2022 | InCloud: Incremental Learning for Point Cloud Place RecognitionabstractPlace recognition is a fundamental component of robotics, and has seen tremendous improvements through the use of deep learning models in recent years. Networks can experience significant drops in performance when deployed in unseen or highly dynamic environments, and require additional training on the collected data. However naively fine-tuning on new training distributions can cause severe degradation of performance on previously visited domains, a phenomenon known as catastrophic forgetting. In this paper we address the problem of incremental learning for point cloud place recognition and introduce InCloud, a structure-aware distillation-based approach which preserves the higher-order structure of the network's embedding space. We introduce several challenging new benchmarks on four popular and large-scale LiDAR datasets (Oxford, MulRan, In-house and KITTI) showing broad improvements in point cloud place recognition performance over a variety of network architectures. To the best of our knowledge, this work is the first to effectively apply incremental learning for point cloud place recognition. Data pre-processing, training and evaluation code for this paper can be found at https://github.com/csiro-robotics/InCloud. Joshua Knights, Peyman Moghadam, Milad Ramezani, Sridha Sridharan, Clinton Fookes |
IROS | 5 |
| 2022 | CorticalFlow++: Boosting Cortical Surface Reconstruction Accuracy, Regularity, and Interoperability
Rodrigo Santa Cruz, Léo Lebrat, Darren Fu, Pierrick Bourgeat, Jurgen Fripp, Clinton Fookes, Olivier Salvado |
MICCAI (5) | 6 |
| 2022 | Learning test-time augmentation for content-based image retrieval
Osman Tursun, Simon Denman, Sridha Sridharan, Clinton Fookes |
Comput. Vis. Image Underst. | 4 |
| 2022 | Affect recognition from scalp-EEG using channel-wise encoder networks coupled with geometric deep learning and multi-channel feature fusion
Darshana Priyasad, Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
Knowl. Based Syst. | 5 |
| 2022 | Quantifiable brain atrophy synthesis for benchmarking of cortical thickness estimation methodsabstractCortical thickness (CTh) is routinely used to quantify grey matter atrophy as it is a significant biomarker in studying neurodegenerative and neurological conditions. Clinical studies commonly employ one of several available CTh estimation software tools to estimate CTh from brain MRI scans. In recent years, machine learning-based methods emerged as a faster alternative to the main-stream CTh estimation methods (e.g. FreeSurfer). Evaluation and comparison of CTh estimation methods often include various metrics and downstream tasks, but none fully covers the sensitivity to sub-voxel atrophy characteristic of neurodegeneration. In addition, current evaluation methods do not provide a framework for the intra-method region-wise evaluation of CTh estimation methods. Therefore, we propose a method for brain MRI synthesis capable of generating a range of sub-voxel atrophy levels (global and local) with quantifiable changes from the baseline scan. We further create a synthetic test set and evaluate four different CTh estimation methods: FreeSurfer (cross-sectional), FreeSurfer (longitudinal), DL+DiReCT and HerstonNet. DL+DiReCT showed superior sensitivity to sub-voxel atrophy over other methods in our testing framework. The obtained results indicate that our synthetic test set is suitable for benchmarking CTh estimation methods on both global and local scales as well as regional inter-and intra-method performance comparison. Filip Rusak, Rodrigo Santa Cruz, Léo Lebrat, Ondrej Hlinka, Jurgen Fripp, Elliot Smith 0003, Clinton Fookes, Andrew P. Bradley, Pierrick Bourgeat |
Medical Image Anal. | 7 |
| 2022 | Split 'n' merge net: A dynamic masking network for multi-task attention
Tharindu Fernando, Sridha Sridharan, Simon Denman, Clinton Fookes |
Pattern Recognit. | 4 |
| 2022 | An efficient framework for zero-shot sketch-based image retrieval
Osman Tursun, Simon Denman, Sridha Sridharan, Ethan Goan, Clinton Fookes |
Pattern Recognit. | 5 |
| 2022 | Channel Graph Regularized Correlation Filters for Visual Object TrackingabstractCorrelation Filters (CF) are a popular choice for visual object tracking due to their efficiency in the frequency domain. Convolutional and hand-crafted features are jointly used when learning a filter, however, these features are not uniformly important when tracking a target. Given this observation, spatial and temporal regularization and attention models have been investigated. However, these models do not consider the interaction between different feature channels. As a result, dissimilar weights are assigned to similar feature channels. To address this issue, we propose a channel attention model and study two different regularization methods for attention. We investigate the application of channel regularization to emphasize important feature channels; and graph regularization which increases the likelihood of similar feature channels obtaining similar weights. The proposed formulation can be efficiently solved via the alternating direction method of multipliers. We first show the advantages of using the proposed channel regularization by demonstrating its performance when applied to two existing CF trackers. This is followed by analyzing the effect of using the proposed channel-graph regularization for CF based tracking. The evaluation is performed on publicly available tracking datasets: OTB100, TC128, VOT-2017, VOT-2019, LaSOT, UAV123, and GOT-10k. Evaluation over multiple challenges and a comparative analysis with existing top-ranked trackers shows that our formulation improves the discriminative power of the learned CF, preventing tracker drift during challenging scenarios. Arjun Tyagi, A. Venkata Subramanyam, Simon Denman, Sridha Sridharan, Clinton Fookes |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | 3-D Bi-directional LSTM for Satellite Soil Moisture DownscalingabstractSoil moisture is a crucial parameter of hydrological processes as it affects the exchange of water and heat at the land/atmosphere interface. Regional hydrological applications (floods and modeling of small basins) and agricultural applications (irrigation and agricultural land mapping) require daily soil moisture (SM) values having a spatial resolution of at least 1-km. This requirement is currently unmet by existing satellite missions. Notably, SM has variability over three dimensions. As such, accurate prediction of satellite SM requires multiple bidirectional spectra-spatiotemporal analyses. However, current state-of-the-art SM downscaling models can not yet fulfill this requirement. This paper proposes a new bi-directional LSTM model dubbed the three-dimensional bi-directional LSTM (3D-Bi-LSTM), which downscales the Soil Moisture Active Passive (SMAP) global daily 9-km SM to daily 1-km SM. In the proposed downscaling model, the region-specific soil moisture indices (SMIs) are first extracted using a covariance-adaptive convolutional neural network (CNN) to support the extraction of important distinctive information from multispectral data. Next, the CNN output is provided to the 3D-Bi-LSTM to perform the bi-directional analysis of spatial correlation within a feature and spectral correlation between features over multiple time instants. Experimental results demonstrate the proposed model outperforms state-of-the-art networks. An ablation study, transferability assessment, and feature importance study further demonstrate the proposed 3D-Bi-LSTM’s efficiency. Neethu Madhukumar, Eric Wang 0001, Clinton Fookes, Wei Xiang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Component-Based Attention for Large-Scale Trademark RetrievalabstractThe need for large-scale trademark retrieval (TR) systems has significantly increased to combat the rise in international trademark infringement. Unfortunately, the ranking accuracy of current approaches using either hand-crafted or pre-trained deep convolution neural network (DCNN) features is inadequate for large-scale deployments. We show in this paper that the ranking accuracy of TR systems can be significantly improved by incorporating hard and soft attention mechanisms, which direct attention to critical information such as figurative elements and reduce the attention given to distracting and uninformative elements such as text and background. Our proposed approach achieves state-of-the-art results on a challenging large-scale trademark dataset. Osman Tursun, Simon Denman, Sabesan Sivipalan, Sridha Sridharan, Clinton Fookes, Sandra Mau |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2022 | Geometric Deep Learning for Subject Independent Epileptic Seizure Prediction Using Scalp EEG SignalsabstractRecently, researchers in the biomedical community have introduced deep learning-based epileptic seizure prediction models using electroencephalograms (EEGs) that can anticipate an epileptic seizure by differentiating between the pre-ictal and interictal stages of the subject's brain. Despite having the appearance of a typical anomaly detection task, this problem is complicated by subject-specific characteristics in EEG data. Therefore, studies that investigate seizure prediction widely employ subject-specific models. However, this approach is not suitable in situations where a target subject has limited (or no) data for training. Subject-independent models can address this issue by learning to predict seizures from multiple subjects, and therefore are of greater value in practice. In this study, we propose a subject-independent seizure predictor using Geometric Deep Learning (GDL). In the first stage of our GDL-based method we use graphs derived from physical connections in the EEG grid. We subsequently seek to synthesize subject-specific graphs using deep learning. The models proposed in both stages achieve state-of-the-art performance using a one-hour early seizure prediction window on two benchmark datasets (CHB-MIT-EEG: 95.38% with 23 subjects and Siena-EEG: 96.05% with 15 subjects). To the best of our knowledge, this is the first study that proposes synthesizing subject-specific graphs for seizure prediction. Furthermore, through model interpretation we outline how this method can potentially contribute towards Scalp EEG-based seizure localization. Theekshana Dissanayake, Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
IEEE J. Biomed. Health Informatics | 5 |
| 2022 | Robust and Interpretable Temporal Convolution Network for Event Detection in Lung Sound RecordingsabstractOBJECTIVE: This paper proposes a novel framework for lung sound event detection, segmenting continuous lung sound recordings into discrete events and performing recognition of each event. METHODS: We propose the use of a multi-branch TCN architecture and exploit a novel fusion strategy to combine the resultant features from these branches. This not only allows the network to retain the most salient information across different temporal granularities and disregards irrelevant information, but also allows our network to process recordings of arbitrary length. RESULTS: The proposed method is evaluated on multiple public and in-house benchmarks, containing irregular and noisy recordings of the respiratory auscultation process for the identification of auscultation events including inhalation, crackles, and rhonchi. Moreover, we provide an end-to-end model interpretation pipeline. CONCLUSION: Our analysis of different feature fusion strategies shows that the proposed feature concatenation method leads to better suppression of non-informative features, which drastically reduces the classifier overhead resulting in a robust lightweight network. SIGNIFICANCE: Lung sound event detection is a primary diagnostic step for numerous respiratory diseases. The proposed method provides a cost-effective and efficient alternative to exhaustive manual segmentation, and provides more accurate segmentation than existing methods. The end-to-end model interpretability helps to build the required trust in the system for use in clinical settings. Tharindu Fernando, Sridha Sridharan, Simon Denman, Houman Ghaemmaghami, Clinton Fookes |
IEEE J. Biomed. Health Informatics | 5 |
| 2022 | Deep Auto-Encoders With Sequential Learning for Multimodal Dimensional Emotion RecognitionabstractMultimodal dimensional emotion recognition has drawn a great attention from the affective computing community and numerous schemes have been extensively investigated, making a significant progress in this area. However, several questions still remain unanswered for most of existing approaches including: (i) how to simultaneously learn compact yet representative features from multimodal data, (ii) how to effectively capture complementary features from multimodal streams, and (iii) how to perform all the tasks in an end-to-end manner. To address these challenges, in this paper, we propose a novel deep neural network architecture consisting of a two-stream auto-encoder and a long short term memory for effectively integrating visual and audio signal streams for emotion recognition. To validate the robustness of our proposed architecture, we carry out extensive experiments on the multimodal emotion in the wild dataset: RECOLA. Experimental results show that the proposed method achieves state-of-the-art recognition performance. Dung Nguyen 0001, Duc Thanh Nguyen, Thanh Thi Nguyen 0001, Son N. Tran, Thin Nguyen, Sridha Sridharan, Clinton Fookes |
IEEE Trans. Multim. | 8 |
| 2022 | Elasticity Meets Continuous-Time: Map-Centric Dense 3D LiDAR SLAMabstractMap-centric SLAM utilizes elasticity as a means of loop closure. This approach reduces the cost of loop closure while still providing large-scale fusion-based dense maps, when compared to trajectory-centric SLAM approaches. In this article, we present a novel framework, namedElasticLiDAR++, for multimodal map-centric SLAM. Having the advantages of a map-centric approach, our method exhibits new features to overcome the shortcomings of existing systems associated with multimodal (LiDAR-inertial-visual) sensor fusion and LiDAR motion distortion. This is accomplished through the use of a local continuous-time trajectory representation. Also, our surface resolution preserving matching algorithm and normal-inverse-Wishart-based surfel fusion model enables nonredundant yet dense mapping. Furthermore, we present a robust metric loop closure model to make the approach stable regardless of where the loop closure occurs. Finally, we demonstrate our approach through both simulation and real data experiments using multiple sensor payload configurations and environments to illustrate its utility and robustness. Chanoh Park, Peyman Moghadam, Jason Williams 0002, Soohwan Kim, Sridha Sridharan, Clinton Fookes |
IEEE Trans. Robotics | 6 |
| 2021 | MongeNet: Efficient Sampler for Geometric Deep LearningabstractRecent advances in geometric deep-learning introduce complex computational challenges for evaluating the distance between meshes. From a mesh model, point clouds are necessary along with a robust distance metric to assess surface quality or as part of the loss function for training models. Current methods often rely on a uniform random mesh discretization, which yields irregular sampling and noisy distance estimation. In this paper we introduce MongeNet, a fast and optimal transport based sampler that allows for an accurate discretization of a mesh with better approximation properties. We compare our method to the ubiquitous random uniform sampling and show that the approximation error is almost half with a very small computational overhead. Léo Lebrat, Rodrigo Santa Cruz, Clinton Fookes, Olivier Salvado |
CVPR | 3 |
| 2021 | Learning Regional Attention Over Multi-Resolution Deep Convolutional Features For Trademark RetrievalabstractLarge-scale trademark retrieval is an important content-based image retrieval task. A recent study shows that off-the-shelf deep features aggregated with Regional-Maximum Activation of Convolutions (R-MAC) achieve state-of-the-art results. However, R-MAC suffers in the presence of background clutter/trivial regions and scale variance, and discards important spatial information. We introduce three simple but effective modifications to R-MAC to overcome these drawbacks. First, we propose the use of both sum and max pooling to minimise the loss of spatial information. We also employ domain-specific unsupervised soft-attention to eliminate background clutter and unimportant regions. Finally, we add multi-resolution inputs to enhance the scale-invariance of R-MAC. We evaluate these three modifications on the million-scale METU dataset. Our results show that all modifications bring non-trivial improvements, and surpass previous state-of-the-art results. Osman Tursun, Simon Denman, Sridha Sridharan, Clinton Fookes |
ICIP | 4 |
| 2021 | Locus: LiDAR-based Place Recognition using Spatiotemporal Higher-Order PoolingabstractPlace Recognition enables the estimation of a globally consistent map and trajectory by providing non-local constraints in Simultaneous Localisation and Mapping (SLAM). This paper presents Locus, a novel place recognition method using 3D LiDAR point clouds in large-scale environments. We propose a method for extracting and encoding topological and temporal information related to components in a scene and demonstrate how the inclusion of this auxiliary information in place description leads to more robust and discriminative scene representations. Second-order pooling along with a non- linear transform is used to aggregate these multi-level features to generate a fixed-length global descriptor, which is invariant to the permutation of input features. The proposed method outperforms state-of-the-art methods on the KITTI dataset. Furthermore, Locus is demonstrated to be robust across several challenging situations such as occlusions and viewpoint changes in 3D LiDAR point clouds. The open-source implementation is available at: https://github.com/csiro-robotics/locus. Kavisha Vidanapathirana, Peyman Moghadam, Ben Harwood, Muming Zhao, Sridha Sridharan, Clinton Fookes |
ICRA | 6 |
| 2021 | CorticalFlow: A Diffeomorphic Mesh Transformer Network for Cortical Surface ReconstructionabstractIn this paper, we introduce CorticalFlow, a new geometric deep-learning model that, given a 3-dimensional image, learns to deform a reference template towards a targeted object. To conserve the template mesh’s topological properties, we train our model over a set of diffeomorphic transformations. This new implementation of a flow Ordinary Differential Equation (ODE) framework benefits from a small GPU memory footprint, allowing the generation of surfaces with several hundred thousand vertices. To reduce topological errors introduced by its discrete resolution, we derive numeric conditions which improve the manifoldness of the predicted triangle mesh. To exhibit the utility of CorticalFlow, we demonstrate its performance for the challenging task of brain cortical surface reconstruction. In contrast to the current state-of-the-art, CorticalFlow produces superior surfaces while reducing the computation time from nine and a half minutes to one second. More significantly, CorticalFlow enforces the generation of anatomically plausible surfaces; the absence of which has been a major impediment restricting the clinical relevance of such surface reconstruction methods. Léo Lebrat, Rodrigo Santa Cruz, Frédéric de Gournay, Darren Fu, Pierrick Bourgeat, Jurgen Fripp, Clinton Fookes, Olivier Salvado |
NeurIPS | 7 |
| 2021 | DeepCSR: A 3D Deep Learning Approach for Cortical Surface ReconstructionabstractThe study of neurodegenerative diseases relies on the reconstruction and analysis of the brain cortex from magnetic resonance imaging (MRI). Traditional frameworks for this task like FreeSurfer demand lengthy runtimes, while its accelerated variant FastSurfer still relies on a voxel-wise segmentation which is limited by its resolution to capture narrow continuous objects as cortical surfaces. Having these limitations in mind, we propose DeepCSR, a 3D deep learning framework for cortical surface reconstruction from MRI. Towards this end, we train a neural network model with hypercolumn features to predict implicit surface representations for points in a brain template space. After training, the cortical surface at a desired level of detail is obtained by evaluating surface representations at specific coordinates, and subsequently applying a topology correction algorithm and an isosurface extraction method. Thanks to the continuous nature of this approach and the efficacy of its hypercolumn features scheme, DeepCSR efficiently reconstructs cortical surfaces at high resolution capturing fine details in the cortical folding. Moreover, DeepCSR is as accurate, more precise, and faster than the widely used FreeSurfer toolbox and its deep learning powered variant FastSurfer on reconstructing cortical surfaces from MRI which should facilitate large-scale medical studies and new healthcare applications. Rodrigo Santa Cruz, Léo Lebrat, Pierrick Bourgeat, Clinton Fookes, Jurgen Fripp, Olivier Salvado |
WACV | 4 |
| 2021 | IGSSTRCF: Importance Guided Sparse Spatio-Temporal Regularized Correlation Filters For TrackingabstractThis paper proposes a novel Importance Guided Sparse Spatio-Temporal Regularization based Correlation Filter (IGSSTRCF) tracker. Our formulation explicitly models the variations in the correlation filters and associated spatial weights in successive frames. By imposing a sparsity penalty on these variations, the formulation ensures that only relevant changes are incorporated during updates. This results in more robust filter coefficients that minimize the tracking drift. The IGSSTRCF also includes an adaptive channel importance estimation strategy that assigns an importance weight to each feature channel during training. The proposed formulation is efficiently solved via the alternating direction method of multipliers. A comparative analysis is shown on TC128, UAV123, VOT-2017, and VOT-2019 datasets; and we present an ablation study to demonstrate the contribution of each component of the IGSSTRCF. It is observed that we outperform several state-of-the-art trackers and each component of the proposed IGSSTRCF contributes positively towards tracker performance. A. Venkata Subramanyam, Simon Denman, Sridha Sridharan, Clinton Fookes |
WACV | 5 |
| 2021 | Multi-modal semantic image segmentation
Akila Pemasiri, Kien Nguyen Thanh, Sridha Sridharan, Clinton Fookes |
Comput. Vis. Image Underst. | 4 |
| 2021 | Detection of Fake and Fraudulent Faces via Neural Memory NetworksabstractAdvances in computer vision have brought us to the point where we have the ability to synthesise realistic fake content. Such approaches are seen as a source of disinformation and mistrust, and pose serious concerns to governments around the world. Convolutional Neural Networks (CNNs) demonstrate encouraging results when detecting fake images that arise from the specific type of manipulation they are trained on. However, this success has not transitioned to unseen manipulation types, resulting in a significant gap in the line-of-defense. We propose a Hierarchical Attention Memory Network (HAMN), motivated by the social cognition processes of the human brain, for the detection of fake faces. Through visual cues and by utilising knowledge stored in neural memories, we allow the network to reason about the perceived face and anticipate it's future semantic embeddings. This renders a generalisable face tampering detection framework. Experimental results demonstrate the proposed approach achieves superior performance for fake and fraudulent face detection. Tharindu Fernando, Clinton Fookes, Simon Denman, Sridha Sridharan |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2021 | End-to-End Domain Adaptive Attention Network for Cross-Domain Person Re-IdentificationabstractPerson re-identification (re-ID) remains challenging in a real-world scenario, as it requires a trained network to generalise to totally unseen target data in the presence of variations across domains. Recently, generative adversarial models have been widely adopted to enhance the diversity of training data. These approaches, however, often fail to generalise to other domains, as existing generative person re-identification models have a disconnect between the generative component and the discriminative feature learning stage. To address the on-going challenges regarding model generalisation, we propose an end-to-end domain adaptive attention network to jointly translate images between domains and learn discriminative re-id features in a single framework. To address the domain gap challenge, we introduce an attention module for image translation from source to target domains without affecting the identity of a person. More specifically, attention is directed to the background instead of the entire image of the person, ensuring identifying characteristics of the subject are preserved. The proposed joint learning network results in a significant performance improvement over state-of-the-art methods on several challenging benchmark datasets. Amena Khatun, Simon Denman, Sridha Sridharan, Clinton Fookes |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2021 | TMMF: Temporal Multi-Modal Fusion for Single-Stage Continuous Gesture RecognitionabstractGesture recognition is a much studied research area which has myriad real-world applications including robotics and human-machine interaction. Current gesture recognition methods have focused on recognising isolated gestures, and existing continuous gesture recognition methods are limited to two-stage approaches where independent models are required for detection and classification, with the performance of the latter being constrained by detection performance. In contrast, we introduce a single-stage continuous gesture recognition framework, called Temporal Multi-Modal Fusion (TMMF), that can detect and classify multiple gestures in a video via a single model. This approach learns the natural transitions between gestures and non-gestures without the need for a pre-processing segmentation step to detect individual gestures. To achieve this, we introduce a multi-modal fusion mechanism to support the integration of important information that flows from multi-modal inputs, and is scalable to any number of modes. Additionally, we propose Unimodal Feature Mapping (UFM) and Multi-modal Feature Mapping (MFM) models to map uni-modal features and the fused multi-modal features respectively. To further enhance performance, we propose a mid-point based loss function that encourages smooth alignment between the ground truth and the prediction, helping the model to learn natural gesture transitions. We demonstrate the utility of our proposed framework, which can handle variable-length input videos, and outperforms the state-of-the-art on three challenging datasets: EgoGesture, IPN hand and ChaLearn LAP Continuous Gesture Dataset (ConGD). Furthermore, ablation experiments show the importance of different components of the proposed framework. Harshala Gammulle, Simon Denman, Sridha Sridharan, Clinton Fookes |
IEEE Trans. Image Process. | 4 |
| 2021 | Identification of Children at Risk of Schizophrenia via Deep Learning and EEG ResponsesabstractThe prospective identification of children likely to develop schizophrenia is a vital tool to support early interventions that can mitigate the risk of progression to clinical psychosis. Electroencephalographic (EEG) patterns from brain activity and deep learning techniques are valuable resources in achieving this identification. We propose automated techniques that can process raw EEG waveforms to identify children who may have an increased risk of schizophrenia compared to typically developing children. We also analyse abnormal features that remain during developmental follow-up over a period of ∼ 4 years in children with a vulnerability to schizophrenia initially assessed when aged 9 to 12 years. EEG data from participants were captured during the recording of a passive auditory oddball paradigm. We undertake a holistic study to identify brain abnormalities, first by exploring traditional machine learning algorithms using classification methods applied to hand-engineered features (event-related potential components). Then, we compare the performance of these methods with end-to-end deep learning techniques applied to raw data. We demonstrate via average cross-validation performance measures that recurrent deep convolutional neural networks can outperform traditional machine learning methods for sequence modeling. We illustrate the intuitive salient information of the model with the location of the most relevant attributes of a post-stimulus window. This baseline identification system in the area of mental illness supports the evidence of developmental and disease effects in a pre-prodromal phase of psychosis. These results reinforce the benefits of deep learning to support psychiatric classification and neuroscientific research more broadly. David Ahmedt-Aristizabal, Tharindu Fernando, Simon Denman, Jonathan E. Robinson, Sridha Sridharan, Patrick J. Johnston, Kristin R. Laurens, Clinton Fookes |
IEEE J. Biomed. Health Informatics | 8 |
| 2021 | A Robust Interpretable Deep Learning Classifier for Heart Anomaly Detection Without SegmentationabstractTraditionally, abnormal heart sound classification is framed as a three-stage process. The first stage involves segmenting the phonocardiogram to detect fundamental heart sounds; after which features are extracted and classification is performed. Some researchers in the field argue the segmentation step is an unwanted computational burden, whereas others embrace it as a prior step to feature extraction. When comparing accuracies achieved by studies that have segmented heart sounds before analysis with those who have overlooked that step, the question of whether to segment heart sounds before feature extraction is still open. In this study, we explicitly examine the importance of heart sound segmentation as a prior step for heart sound classification, and then seek to apply the obtained insights to propose a robust classifier for abnormal heart sound detection. Furthermore, recognizing the pressing need for explainable Artificial Intelligence (AI) models in the medical domain, we also unveil hidden representations learned by the classifier using model interpretation techniques. Experimental results demonstrate that the segmentation which can be learned by the model plays an essential role in abnormal heart sound classification. Our new classifier is also shown to be robust, stable and most importantly, explainable, with an accuracy of almost 100% on the widely used PhysioNet dataset. Theekshana Dissanayake, Tharindu Fernando, Simon Denman, Sridha Sridharan, Houman Ghaemmaghami, Clinton Fookes |
IEEE J. Biomed. Health Informatics | 6 |
| 2020 | Geometry-Constrained Car Recognition Using a 3D Perspective NetworkabstractWe present a novel learning framework for vehicle recognition from a single RGB image. Unlike existing methods which only use attention mechanisms to locate 2D discriminative information, our work learns a novel 3D perspective feature representation of a vehicle, which is then fused with 2D appearance feature to predict the category. The framework is composed of a global network (GN), a 3D perspective network (3DPN), and a fusion network. The GN is used to locate the region of interest (RoI) and generate the 2D global feature. With the assistance of the RoI, the 3DPN estimates the 3D bounding box under the guidance of the proposed vanishing point loss, which provides a perspective geometry constraint. Then the proposed 3D representation is generated by eliminating the viewpoint variance of the 3D bounding box using perspective transformation. Finally, the 3D and 2D feature are fused to predict the category of the vehicle. We present qualitative and quantitative results on the vehicle classification and verification tasks in the BoxCars dataset. The results demonstrate that, by learning such a concise 3D representation, we can achieve superior performance to methods that only use 2D information while retain 3D meaningful information without the challenge of requiring a 3D CAD model. ZongYuan Ge, Simon Denman, Sridha Sridharan, Clinton Fookes |
AAAI | 5 |
| 2020 | Attention Driven Fusion for Multi-Modal Emotion RecognitionabstractDeep learning has emerged as a powerful alternative to hand-crafted methods for emotion recognition on combined acoustic and text modalities. Baseline systems model emotion information in text and acoustic modes independently using Deep Convolutional Neural Networks (DCNN) and Recurrent Neural Networks (RNN), followed by applying attention, fusion, and classification. In this paper, we present a deep learning-based approach to exploit and fuse text and acoustic data for emotion classification. We utilize a SincNet layer, based on parameterized sinc functions with band-pass filters, to extract acoustic features from raw audio followed by a DCNN. This approach learns filter banks tuned for emotion recognition and provides more effective features compared to directly applying convolutions over the raw speech signal. For text processing, we use two branches (a DCNN and a Bi-direction RNN followed by a DCNN) in parallel where cross attention is introduced to infer the N-gram level correlations on hidden representations received from the Bi-RNN. Following existing state-of-the-art, we evaluate the performance of the proposed system on the IEMOCAP dataset. Experimental results indicate that the proposed system outperforms existing methods, achieving 5.2% improvement in weighted accuracy. Darshana Priyasad, Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
ICASSP | 5 |
| 2020 | Two-Stream Deep Feature Modelling for Automated Video Endoscopy Data Analysis
Harshala Gammulle, Simon Denman, Sridha Sridharan, Clinton Fookes |
MICCAI (3) | 4 |
| 2020 | Semantic Consistency and Identity Mapping Multi-Component Generative Adversarial Network for Person Re-IdentificationabstractIn a real world environment, person re-identification (Re-ID) is a challenging task due to variations in lighting conditions, viewing angles, pose and occlusions. Despite recent performance gains, current person Re-ID algorithms still suffer heavily when encountering these variations. To address this problem, we propose a semantic consistency and identity mapping multi-component generative adversarial network (SC-IMGAN) which provides style adaptation from one to many domains. To ensure that transformed images are as realistic as possible, we propose novel identity mapping and semantic consistency losses to maintain identity across the diverse domains. For the Re-ID task, we propose a joint verification-identification quartet network which is trained with generated and real images, followed by an effective quartet loss for verification. Our proposed method outperforms state-of-the-art techniques on six challenging person Re-ID datasets: CUHK01, CUHK03, VIPeR, PRID2011, iLIDS and Market-1501. Amena Khatun, Simon Denman, Sridha Sridharan, Clinton Fookes |
WACV | 4 |
| 2020 | LSTM guided ensemble correlation filter tracking with appearance model pool
A. Venkata Subramanyam, Simon Denman, Sridha Sridharan, Clinton Fookes |
Comput. Vis. Image Underst. | 5 |
| 2020 | Joint identification-verification for person re-identification: A four stream deep learning approach with improved quartet loss function
Amena Khatun, Simon Denman, Sridha Sridharan, Clinton Fookes |
Comput. Vis. Image Underst. | 4 |
| 2020 | MTRNet++: One-stage mask-based scene text eraser
Osman Tursun, Simon Denman, Sabesan Sivipalan, Sridha Sridharan, Clinton Fookes |
Comput. Vis. Image Underst. | 6 |
| 2020 | Neural memory plasticity for medical anomaly detection
Tharindu Fernando, Simon Denman, David Ahmedt-Aristizabal, Sridha Sridharan, Kristin R. Laurens, Patrick J. Johnston, Clinton Fookes |
Neural Networks | 7 |
| 2020 | Fine-grained action segmentation using the semi-supervised action GAN
Harshala Gammulle, Simon Denman, Sridha Sridharan, Clinton Fookes |
Pattern Recognit. | 4 |
| 2020 | Context from within: Hierarchical context modeling for semantic segmentation
Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan |
Pattern Recognit. | 2 |
| 2020 | Correlation-aware adversarial domain adaptation and generalization
Mohammad Mahfujur Rahman, Clinton Fookes, Mahsa Baktash, Sridha Sridharan |
Pattern Recognit. | 2 |
| 2020 | Hierarchical Attention Network for Action Segmentation
Harshala Gammulle, Simon Denman, Sridha Sridharan, Clinton Fookes |
Pattern Recognit. Lett. | 4 |
| 2020 | Temporarily-Aware Context Modeling Using Generative Adversarial Networks for Speech Activity DetectionabstractThis paper presents a novel framework for Speech Activity Detection (SAD). Inspired by the recent success of multi-task learning approaches in the speech processing domain, we propose a novel joint learning framework for SAD. We utilise generative adversarial networks to automatically learn a loss function for joint prediction of the frame-wise speech/ non-speech classifications together with the next audio segment. In order to exploit the temporal relationships within the input signal, we propose a temporal discriminator which aims to ensure that the predicted signal is temporally consistent. We evaluate the proposed framework on multiple public benchmarks, including NIST OpenSAT' 17, AMI Meeting and HAVIC, where we demonstrate its capability to outperform state-of-the-art SAD approaches. Furthermore, our cross-database evaluations demonstrate the robustness of the proposed approach across different languages, accents, and acoustic environments. Tharindu Fernando, Sridha Sridharan, Mitchell McLaren, Darshana Priyasad, Simon Denman, Clinton Fookes |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2020 | Target-Specific Siamese Attention Network for Real-Time Object TrackingabstractDeep similarity trackers are able to track above real-time speed. However, their accuracy is considerably lower than deep classification based trackers since they avoid valuable online cues. To feed the target-specific information for real-time object tracking, we propose a novel Siamese attention network. Different types of attention mechanisms are used to capture different contexts of target information and then learned knowledge is used to feed target cues at different representation levels of similarity tracking. In addition, an online learning mechanism is employed to utilise the available target-specific data. The proposed tracker reduces the impact of noise in the target template and improves the accuracy of similarity tracking by feeding target cues into the similarity search. Extensive evaluation performed on OTB-2013/50/100 and VOT2018 benchmark datasets demonstrate the proposed tracker outperforms state-of-the-art approaches while maintaining real-time tracking speed. Thanikasalam Kokul, Clinton Fookes, Sridha Sridharan, Amirthalingam Ramanan, Amalka Pinidiyaarachchi |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2020 | Constrained Design of Deep Iris NetworksabstractDespite the promise of recent deep neural networks to provide more accurate and efficient iris recognition compared to traditional techniques, there are vital properties of the classic IrisCode which are almost unable to be achieved with current deep iris networks: the compactness of model and the small number of computing operations (FLOPs). This paper casts the iris network design process as a constrained optimization problem which takes model size and computation into account as learning criteria. On one hand, this allows us to fully automate the network design process to search for the optimal iris network architecture with the highest recognition accuracy confined to the computation and model compactness constraints. On the other hand, it allows us to investigate the optimality of the classic IrisCode and recent deep iris networks. It also enables us to learn an optimal iris network and demonstrate state-of-the-art performance with less computation and memory requirements. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan |
IEEE Trans. Image Process. | 2 |
| 2020 | Heart Sound Segmentation Using Bidirectional LSTMs With AttentionabstractOBJECTIVE: This paper proposes a novel framework for the segmentation of phonocardiogram (PCG) signals into heart states, exploiting the temporal evolution of the PCG as well as considering the salient information that it provides for the detection of the heart state. METHODS: We propose the use of recurrent neural networks and exploit recent advancements in attention based learning to segment the PCG signal. This allows the network to identify the most salient aspects of the signal and disregard uninformative information. RESULTS: The proposed method attains state-of-the-art performance on multiple benchmarks including both human and animal heart recordings. Furthermore, we empirically analyse different feature combinations including envelop features, wavelet and Mel Frequency Cepstral Coefficients (MFCC), and provide quantitative measurements that explore the importance of different features in the proposed approach. CONCLUSION: We demonstrate that a recurrent neural network coupled with attention mechanisms can effectively learn from irregular and noisy PCG recordings. Our analysis of different feature combinations shows that MFCC features and their derivatives offer the best performance compared to classical wavelet and envelop features. SIGNIFICANCE: Heart sound segmentation is a crucial pre-processing step for many diagnostic applications. The proposed method provides a cost effective alternative to labour extensive manual segmentation, and provides a more accurate segmentation than existing methods. As such, it can improve the performance of further analysis including the detection of murmurs and ejection clicks. The proposed method is also applicable for detection and segmentation of other one dimensional biomedical signals. Tharindu Fernando, Houman Ghaemmaghami, Simon Denman, Sridha Sridharan, Nayyar Hussain, Clinton Fookes |
IEEE J. Biomed. Health Informatics | 6 |
| 2020 | Memory Augmented Deep Generative Models for Forecasting the Next Shot Location in TennisabstractThis paper presents a novel framework for predicting shot location and type in tennis. Inspired by recent neuroscience discoveries, we incorporate neural memory modules to model the episodic and semantic memory components of a tennis player. We propose a Semi-Supervised Generative Adversarial Network architecture that couples these memory models with the automatic feature learning power of deep neural networks, and demonstrate methodologies for learning player level behavioral patterns with the proposed framework. We evaluate the effectiveness of the proposed model on tennis tracking data from the 2012 Australian Tennis Open and exhibit applications of the proposed method in discovering how players adapt their style depending on the match context. Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2019 | Forecasting Future Action Sequences with Neural Memory Networks
Harshala Gammulle, Simon Denman, Sridha Sridharan, Clinton Fookes |
BMVC | 4 |
| 2019 | Unified 2D and 3D Hand Pose Estimation from a Single Visible or X-ray Image
Akila Pemasiri, Kien Nguyen Thanh, Sridha Sridharan, Clinton Fookes |
BMVC | 4 |
| 2019 | Investigating Domain Sensitivity of DNN Embeddings for Speaker Recognition SystemsabstractA speaker embeddings framework achieves state-of-the-art speaker recognition performance by modeling speaker discriminant information directly using deep neural networks (DNNs). After the introduction of neural network based speaker embeddings, researchers have explored the requirements for training an effective embeddings network. However, the domain of the data used for system development should match the domain of operation for optimal performance. In this paper, we investigate the sensitivity of domain mismatch in the embeddings space. Specifically, degradation in performance is observed when back-end scoring with embeddings is performed with out-domain data. To compensate for the domain mismatch, we propose two novel deep domain adaptation techniques based on autoencoder architectures trained on embeddings in an unsupervised fashion. The results show that domain mismatch can be compensated effectively using autoencoders to adapt the out-domain data to in-domain. Ivan Himawan, Sridha Sridharan, Clinton Fookes |
ICASSP | 4 |
| 2019 | Predicting the Future: A Jointly Learnt Model for Action AnticipationabstractInspired by human neurological structures for action anticipation, we present an action anticipation model that enables the prediction of plausible future actions by forecasting both the visual and temporal future. In contrast to current state-of-the-art methods which first learn a model to predict future video features and then perform action anticipation using these features, the proposed framework jointly learns to perform the two tasks, future visual and temporal representation synthesis, and early action anticipation. The joint learning framework ensures that the predicted future embeddings are informative to the action anticipation task. Furthermore, through extensive experimental evaluations we demonstrate the utility of using both visual and temporal semantics of the scene, and illustrate how this representation synthesis could be achieved through a recurrent Generative Adversarial Network (GAN) framework. Our model outperforms the current state-of-the-art methods on multiple datasets: UCF101, UCF101-24, UT-Interaction and TV Human Interaction. Harshala Gammulle, Simon Denman, Sridha Sridharan, Clinton Fookes |
ICCV | 4 |
| 2019 | MTRNet: A Generic Scene Text EraserabstractText removal algorithms have been proposed for uni-lingual scripts with regular shapes and layouts. However, to the best of our knowledge, a generic text removal method which is able to remove all or user-specified text regions regardless of font, script, language or shape is not available. Developing such a generic text eraser for real scenes is a challenging task, since it inherits all the challenges of multi-lingual and curved text detection and inpainting. To fill this gap, we propose a mask-based text removal network (MTRNet). MTRNet is a conditional adversarial generative network (cGAN) with an auxiliary mask. The introduced auxiliary mask not only makes the cGAN a generic text eraser, but also enables stable training and early convergence on a challenging large-scale synthetic dataset, initially proposed for text detection in real scenes. What's more, MTRNet achieves state-of-the-art results on several real-world datasets including ICDAR 2013, ICDAR 2017 MLT, and CTW1500, without being explicitly trained on this data, outperforming previous state-of-the-art methods trained directly on these datasets. Osman Tursun, Simon Denman, Sabesan Sivipalan, Sridha Sridharan, Clinton Fookes |
ICDAR | 6 |
| 2019 | A Study of x-Vector Based Speaker Recognition on Short UtterancesabstractThe aim of this work is to gain insights into how the deep neural network (DNN) models should be trained for short utterance evaluation conditions in an x-vector based speaker verification system. The study suggests that the speaker embedding can be extracted with reduced dimensions for short utterance evaluation conditions. When the speaker embedding is extracted from deeper layer which has lower dimension, the x-vector system achieves 14% relative improvement over baseline approach on EER on NIST2010 5sec-5sec truncated conditions. We surmise that since short utterances have less phonetic information speaker discriminative x-vectors can be extracted from a deeper layer of the DNN which captures less phonetic information. Another interesting finding is that the x-vector system achieves 5% relative improvement on NIST2010 5sec-5sec evaluation condition when the back-end PLDA is trained using short utterance development data. The results confirms the intuitive expectation that duration of development utterances and the duration of evaluation utterances should be matched. Finally, for the duration mismatch condition, we propose a variance normalization approach for PLDA training that provides a 4% relative improvement on EER over baseline approach. Ahilan Kanagasundaram, Sridha Sridharan, Sriram Ganapathy, Prachi Singh, Clinton Fookes |
INTERSPEECH | 5 |
| 2019 | Coupled Generative Adversarial Network for Continuous Fine-Grained Action SegmentationabstractWe propose a novel conditional GAN (cGAN) model for continuous fine-grained human action segmentation, that utilises multi-modal data and learned scene context information. The proposed approach utilises two GANs: termed Action GAN and Auxiliary GAN, where the Action GAN is trained to operate over the current RGB frame while the Auxiliary GAN utilises supplementary information such as depth or optical flow. The goal of both GANs is to generate similar 'action codes', a vector representation of the current action. To facilitate this process a context extractor that incorporates data and recent outputs from both modes is used to extract context information to aids recognition performance. The result is a recurrent GAN architecture which learns a task specific loss function from multiple feature modalities. Extensive evaluations on variants of the proposed model to show the importance of utilising different streams of information such as context and auxiliary information in the proposed network; and show that our model is capable of outperforming state-of-the-art methods for three widely used datasets: 50 Salads, MERL Shopping and Georgia Tech Egocentric Activities, comprising both static and dynamic camera settings. Harshala Gammulle, Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
WACV | 5 |
| 2019 | Semantic Correspondence in the WildabstractSemantic correspondence estimation where the object instances depicted are deformed extensively from one instance to the next is a challenging problem in computer vision that has received much attention. Unfortunately, all existing approaches require prior knowledge of the object classes which are present in the image environment. This is an unwanted restriction as it can prevent the establishment of semantic correspondence across object classes in wild conditions when it is uncertain which classes will be of interest. In contrast, in this paper we formulate the semantic correspondence estimation task as a key point detection process in which image-to-class classification and image-to-image correspondence are solved simultaneously. Identifying object classes within the same framework to establish correspondence, increases this approach's applicability in real world scenarios. The use of object regions in the process also enhances the accuracy while constraining the search space, thus improving overall efficiency. This new approach is compared with the state-of-the-art on publicly available datasets to validate its capability for improved semantic correspondence estimation in wild conditions. Akila Pemasiri, Kien Nguyen Thanh, Sridha Sridharan, Clinton Fookes |
WACV | 4 |
| 2019 | Multi-Component Image Translation for Deep Domain GeneralizationabstractDomain adaption (DA) and domain generalization (DG) are two closely related methods which are both concerned with the task of assigning labels to an unlabeled data set. The only dissimilarity between these approaches is that DA can access the target data during the training phase, while the target data is totally unseen during the training phase in DG. The task of DG is challenging as we have no earlier knowledge of the target samples. If DA methods are applied directly to DG by a simple exclusion of the target data from training, poor performance will result for a given task. In this paper, we tackle the domain generalization challenge in two ways. In our first approach, we propose a novel deep domain generalization architecture utilizing synthetic data generated by a Generative Adversarial Network (GAN). The discrepancy between the generated images and synthetic images is minimized using existing domain discrepancy metrics such as maximum mean discrepancy or correlation alignment. In our second approach, we introduce a protocol for applying DA methods to a DG scenario by excluding the target data from the training phase, splitting the source data to training and validation parts, and treating the validation data as target data for DA. We conduct extensive experiments on four cross-domain benchmark datasets. Experimental results signify our proposed model outperforms the current state-of-the-art methods for DG. Mohammad Mahfujur Rahman, Clinton Fookes, Mahsa Baktash, Sridha Sridharan |
WACV | 2 |
| 2019 | Deep domain adaptation for anti-spoofing in speaker verification systems
Ivan Himawan, Fernando Villavicencio, Sridha Sridharan, Clinton Fookes |
Comput. Speech Lang. | 4 |
| 2019 | Multimodal clothing recognition for semantic search in unconstrained surveillance imagery
Michael Halstead, Simon Denman, Sridha Sridharan, Yingli Tian, Clinton Fookes |
J. Vis. Commun. Image Represent. | 5 |
| 2019 | Sparse over-complete patch matching
Akila Pemasiri, Kien Nguyen Thanh, Sridha Sridharan, Clinton Fookes |
Pattern Recognit. Lett. | 4 |
| 2019 | Scene Invariant Virtual Gates Using DNNsabstractUnderstanding where people are located and how they are moving about in an environment is critical for operators of large public spaces such as shopping centers, and large public infrastructures such as airports. Automated analysis of CCTV footage is increasingly being used to address this need through techniques that can count crowd sizes, estimate their density, and estimate the through-put of people into and/or out of a choke-point. A limitation of using CCTV based approaches, however, is the need to train models specific to each view which, for large environments with 100s or 1000s of cameras, can quickly become problematic. While there is some success in developing scene-invariant crowd counting and crowd density estimation approaches, much less attention has been given to developing scene-invariant solutions for through-put estimation. In this paper, we investigate the use of convolutional neural network and long short-term memory architectures to estimate pedestrian through-put from arbitrary CCTV viewpoints. To properly develop and demonstrate our approach, we present a new 22 view database featuring 44 h of pedestrian throughput annotation, containing over 11 000 annotated people; and using this proposed approach we show that we are able to outperform a scene-dependant approach across a diverse set of challenging view-points. Simon Denman, Clinton Fookes, Prasad K. D. V. Yarlagadda, Sridha Sridharan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Understanding Patients' Behavior: Vision-Based Analysis of Seizure DisordersabstractA substantial proportion of patients with functional neurological disorders (FND) are being incorrectly diagnosed with epilepsy because their semiology resembles that of epileptic seizures (ES). Misdiagnosis may lead to unnecessary treatment and its associated complications. Diagnostic errors often result from an overreliance on specific clinical features. Furthermore, the lack of electrophysiological changes in patients with FND can also be seen in some forms of epilepsy, making diagnosis extremely challenging. Therefore, understanding semiology is an essential step for differentiating between ES and FND. Existing sensor-based and marker-based systems require physical contact with the body and are vulnerable to clinical situations such as patient positions, illumination changes, and motion discontinuities. Computer vision and deep learning are advancing to overcome these limitations encountered in the assessment of diseases and patient monitoring; however, they have not been investigated for seizure disorder scenarios. Here, we propose and compare two marker-free deep learning models, a landmark-based and a region-based model, both of which are capable of distinguishing between seizures from video recordings. We quantify semiology by using either a fusion of reference points and flow fields, or through the complete analysis of the body. Average leave-one-subject-out cross-validation accuracies for the landmark-based and region-based approaches of 68.1% and 79.6% in our dataset collected from 35 patients, reveal the benefit of video analytics to support automated identification of semiology in the challenging conditions of a hospital setting. David Ahmedt-Aristizabal, Simon Denman, Kien Nguyen Thanh, Sridha Sridharan, Sasha Dionisio, Clinton Fookes |
IEEE J. Biomed. Health Informatics | 6 |
| 2018 | GD-GAN: Generative Adversarial Networks for Trajectory Prediction and Group Detection in Crowds
Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
ACCV (1) | 4 |
| 2018 | Multi-level Sequence GAN for Group Activity Recognition
Harshala Gammulle, Simon Denman, Sridha Sridharan, Clinton Fookes |
ACCV (1) | 4 |
| 2018 | Learning Free-Form Deformations for 3D Object Reconstruction
Dominic Jack, Jhony K. Pontes, Sridha Sridharan, Clinton Fookes, Sareh Rowlands, Frédéric Maire, Anders P. Eriksson |
ACCV (2) | 4 |
| 2018 | Image2Mesh: A Learning Framework for Single Image 3D Reconstruction
Jhony K. Pontes, Chen Kong, Sridha Sridharan, Simon Lucey, Anders P. Eriksson, Clinton Fookes |
ACCV (1) | 6 |
| 2018 | Rethinking Planar Homography Estimation Using Perspective Fields
Simon Denman, Sridha Sridharan, Clinton Fookes |
ACCV (6) | 4 |
| 2018 | Semantic Person Retrieval in Surveillance Using Soft Biometrics: AVSS 2018 Challenge IIabstractIn surveillance and security today it is a common goal to locate a subject of interest purely from a semantic description; think of an offender description form handed into a law enforcement agency. To date, these tasks are primarily undertaken by operators on the ground either by manually searching a premises or by combing through hours of video footage. Using computer vision to attempt to partially or fully automate these tasks has been gathering interest within the research community in recent years, however, to date there has been little coordinated effort to advance the field. This has motivated the challenge that is presented in this paper: the AVSS Challenge on Semantic Person Retrieval in Surveillance Using Soft Biometrics. This challenge consists of two related tasks: person re-identification from a semantic query and person search within a video from a query. In this paper, we present the publicly available data for this challenge, the evaluation framework, and the challenge results. It is our hope that the outcomes of this challenge and the availability of the data used in this challenge will expedite research and development in this societal field. Michael Halstead, Simon Denman, Clinton Fookes, Yingli Tian, Mark S. Nixon |
AVSS | 3 |
| 2018 | Calibrating Cameras in Poor-Conditioned Pitch-Based Sports GamesabstractCamera calibration is a preliminary step in sports analytics which enables us to transform player positions to standard playing area coordinates. While many camera calibration systems work well when the visual content contains sufficient clues, such as a key frame, calibrating without such information, such as may be needed when processing footage captured by a coach from the sidelines or stands, is challenging. In this paper an innovative automatic camera calibration system, which does not make use of any key frames, is presented for sports analytics. The proposed system consists of three components: a robust linear panorama module, a playing area estimation module, and a homography estimation module. It can eliminate distortion and calibrate the camera in each frame simultaneously, using correspondences between pairs of consecutive frames. Experiments on real data evaluate the performance and demonstrate the robustness of the system. Ruan Lakemond, Simon Denman, Sridha Sridharan, Clinton Fookes, Stuart Morgan |
ICASSP | 5 |
| 2018 | Hierarchical Relational Attention for Video Question AnsweringabstractVideo Question Answering (VideoQA) tasks require understanding of the connection of context specific video parts which are temporally distributed. Humans are capable of focusing on temporally distributed video scenes and also to find correspondence or relationships among these segments. To achieve similar capability, a hierarchical relational attention mechanism is proposed in this paper. The proposed VideoQA model derives attention on temporal segments i.e. video features based on each of the question words. Also, contextual relevance of these temporal segments are captured to derive the final video representation which leads to a better reasoning capability. We evaluate the performance of the proposed approach on the MSRVTT-QA and the MSVD-QA datasets to establish its superior performance over the state of the art. Muhammad Iqbal Hasan Chowdhury, Kien Nguyen Thanh, Sridha Sridharan, Clinton Fookes |
ICIP | 4 |
| 2018 | Deep Match Tracker: Classifying when Dissimilar, Similarity Matching when NotabstractVisual tracking frameworks employing Convolutional Neural Networks (CNNs) have shown state-of-the-art performance due to their hierarchical feature representation. While classification and update based deep neural net tracking have shown good performance in terms of accuracy, they have poor tracking speed. On the other hand, recent matching based techniques using CNNs show higher than real-time speed in tracking but this speed is achieved at a considerably lower accuracy. To successfully manage the trade-off between accuracy and speed, we propose a novel CNN architecture for visual tracking. We achieve this trade-off balance by using an approach in which consecutive similar frames are processed with a similarity matching technique, and dissimilar frames are processed with a classification approach within the CNN architecture. The tracking speed is improved by avoiding unnecessary model updates through the measurement of similarity between adjacent frames, while the accuracy is maintained by adopting a classification approach when needed, with deeper level features. Extensive evaluation performed on a publicly available benchmark dataset demonstrates our proposed tracker shows competitive performance while maintaining near real-time speed. Thanikasalam Kokul, Clinton Fookes, Sridha Sridharan, Amirthalingam Ramanan, Amalka Pinidiyaarachchi |
ICIP | 2 |
| 2018 | Non-rigid Reconstruction with a Single Moving RGB-D CameraabstractWe present a novel non-rigid reconstruction method using a moving RGB-D camera. Current approaches use only non-rigid part of the scene and completely ignore the rigid background. Non-rigid parts often lack sufficient geometric and photometric information for tracking large frame-to-frame motion. Our approach uses camera pose estimated from the rigid background for foreground tracking. This enables robust foreground tracking in situations where large frame-to-frame motion occurs. Moreover, we are proposing a multi-scale deformation graph which improves non-rigid tracking without compromising the quality of the reconstruction. We are also contributing a synthetic dataset which is made publically available for evaluating non-rigid reconstruction methods. The dataset provides frame-by-frame ground truth geometry of the scene, the camera trajectory, and masks for background foreground. Experimental results show that our approach is more robust in handling larger frame-to-frame motions and provides better reconstruction compared to state-of-the-art approaches. Shafeeq Elanattil, Peyman Moghadam, Sridha Sridharan, Clinton Fookes, Mark Cox |
ICPR | 4 |
| 2018 | Meta Transfer Learning for Facial Emotion RecognitionabstractThe use of deep learning techniques for automatic facial expression recognition has recently attracted great interest but developed models are still unable to generalize well due to the lack of large emotion datasets for deep learning. To overcome this problem, in this paper, we propose utilizing a novel transfer learning approach relying on PathNet and investigate how knowledge can be accumulated within a given dataset and how the knowledge captured from one emotion dataset can be transferred into another in order to improve the overall performance. To evaluate the robustness of our system, we have conducted various sets of experiments on two emotion datasets: SAVEE and eNTERFACE. The experimental results demonstrate that our proposed system leads to improvement in performance of emotion recognition and performs significantly better than the recent state-of-the-art schemes adopting fine-tuning/pre-trained approaches. Dung Nguyen Tien, Kien Nguyen Thanh, Sridha Sridharan, Iman Abbasnejad, David Dean, Clinton Fookes |
ICPR | 6 |
| 2018 | Elastic LiDAR Fusion: Dense Map-Centric Continuous-Time SLAMabstractThe concept of continuous-time trajectory representation has brought increased accuracy and efficiency to multi-modal sensor fusion in modern SLAM. However, regardless of these advantages, its offline property caused by the requirement of global batch optimization is critically hindering its relevance for real-time and life-long applications. In this paper, we present a dense map-centric SLAM method based on a continuous-time trajectory to cope with this problem. The proposed system locally functions in a similar fashion to conventional Continuous-Time SLAM (CT-SLAM). However, it removes the need for global trajectory optimization by introducing map deformation. The computational complexity of the proposed approach for loop closure does not depend on the operation time, but only on the size of the space it explored before the loop closure. It is therefore more suitable for long term operation compared to the conventional CT-SLAM. Furthermore, the proposed method reduces uncertainty in the reconstructed dense map by using probabilistic surface element (surfel) fusion. We demonstrate that the proposed method produces globally consistent maps without global batch trajectory optimization, and effectively reduces LiDAR noise by surfel fusion. Chanoh Park, Peyman Moghadam, Soohwan Kim, Alberto Elfes, Clinton Fookes, Sridha Sridharan |
ICRA | 5 |
| 2018 | Employing Phonetic Information in DNN Speaker Embeddings to Improve Speaker Recognition PerformanceabstractThe recent speaker embeddings framework has been shown to provide excellent performance on the task of text-independent speaker recognition. The framework is based on a deep neural network (DNN) trained to directly discriminate between speakers from traditional acoustic features such as Mel frequency cepstral coefficients. Prior studies on speaker recognition have found that phonetic information is valuable in the task of speaker identification, with systems being based on either bottleneck features (BFs) or tied-triphone state posteriors from a DNN trained for the task of speech recognition. In this paper, we analyze the role of phonetic BFs for DNN embeddings and explore methods to enhance the BFs further. Experimental results show that exploiting phonetic information encoded in BFs is very valuable for DNN speaker embeddings. Enriching the BFs using a cascaded DNN multi-task architecture is also shown to provide further improvements to the speaker embed- ding system. Ivan Himawan, Mitchell McLaren, Clinton Fookes, Sridha Sridharan |
INTERSPEECH | 4 |
| 2018 | Pedestrian Trajectory Prediction with Structured Memory Hierarchies
Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
ECML/PKDD (1) | 4 |
| 2018 | Investigating Deep Neural Networks for Speaker Diarization in the DIHARD ChallengeabstractWe investigate the use of deep neural networks (DNNs) for the speaker diarization task to improve performance under domain mismatched conditions. Three unsupervised domain adaptation techniques, namely inter-dataset variability compensation (IDVC), domain-invariant covariance normalization (DICN), and domain mismatch modeling (DMM), are applied on DNN based speaker embeddings to compensate for the mismatch in the embedding subspace. We present results conducted on the DIHARD data, which was released for the 2018 diarization challenge. Collected from a diverse set of domains, this data provides very challenging domain mismatched conditions for the diarization task. Our results provide insights into how the performance of our proposed system could be further improved. Ivan Himawan, Sridha Sridharan, Clinton Fookes, Ahilan Kanagasundaram |
SLT | 4 |
| 2018 | Tracking by Prediction: A Deep Generative Model for Mutli-person Localisation and TrackingabstractCurrent multi-person localisation and tracking systems have an over reliance on the use of appearance models for target re-identification and almost no approaches employ a complete deep learning solution for both objectives. We present a novel, complete deep learning framework for multi-person localisation and tracking. In this context we first introduce a light weight sequential Generative Adversarial Network architecture for person localisation, which overcomes issues related to occlusions and noisy detections, typically found in a multi person environment. In the proposed tracking framework we build upon recent advances in pedestrian trajectory prediction approaches and propose a novel data association scheme based on predicted trajectories. This removes the need for computationally expensive person re-identification systems based on appearance features and generates human like trajectories with minimal fragmentation. The proposed method is evaluated on multiple public benchmarks including both static and dynamic cameras and is capable of generating outstanding performance, especially among other recently proposed deep neural network based approaches. Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
WACV | 4 |
| 2018 | Task Specific Visual Saliency Prediction with Memory Augmented Conditional Generative Adversarial Networks
Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
WACV | 4 |
| 2018 | A Deep Four-Stream Siamese Convolutional Neural Network with Joint Verification and Identification Loss for Person Re-DetectionabstractState-of-the-art person re-identification systems that employ a triplet based deep network suffer from a poor generalization capability. In this paper, we propose a four stream Siamese deep convolutional neural network for person redetection that jointly optimises verification and identification losses over a four image input group. Specifically, the proposed method overcomes the weakness of the typical triplet formulation by using groups of four images featuring two matched (i.e. the same identity) and two mismatched images. This allows us to jointly increase the interclass variations and reduce the intra-class variations in the learned feature space. The proposed approach also optimises over both the identification and verification losses, further minimising intra-class variation and maximising inter-class variation, improving overall performance. Extensive experiments on four challenging datasets, VIPeR, CUHK01, CUHK03 and PRID2011, demonstrates that the proposed approach achieves state-of-the-art performance. Amena Khatun, Simon Denman, Sridha Sridharan, Clinton Fookes |
WACV | 4 |
| 2018 | Deep spatio-temporal feature fusion with compact bilinear pooling for multimodal emotion recognition
Dung Nguyen Tien, Kien Nguyen Thanh, Sridha Sridharan, David Dean, Clinton Fookes |
Comput. Vis. Image Underst. | 5 |
| 2018 | Tree Memory Networks for modelling long-term temporal dependencies
Tharindu Fernando, Simon Denman, Aaron McFadyen, Sridha Sridharan, Clinton Fookes |
Neurocomputing | 5 |
| 2018 | Soft + Hardwired attention: An LSTM framework for human trajectory prediction and abnormal event detection
Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
Neural Networks | 4 |
| 2018 | Super-resolution for biometrics: A comprehensive survey
Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Massimo Tistarelli, Mark S. Nixon |
Pattern Recognit. | 2 |
| 2017 | Compact Model Representation for 3D Reconstructionabstract3D reconstruction from 2D images is a central problem in computer vision. Recent works have been focusing on reconstruction directly from a single image. It is well known however that only one image cannot provide enough information for such a reconstruction. A prior knowledge that has been entertained are 3D CAD models due to its online ubiquity. A fundamental question is how to compactly represent millions of CAD models while allowing generalization to new unseen objects with fine-scaled geometry. We introduce an approach to compactly represent a 3D mesh. Our method first selects a 3D model from a graph structure by using a novel free-form deformation FFD 3D-2D registration, and then the selected 3D model is refined to best fit the image silhouette. We perform a comprehensive quantitative and qualitative analysis that demonstrates impressive dense and realistic 3D reconstruction from single images. Jhony K. Pontes, Chen Kong, Anders P. Eriksson, Clinton Fookes, Sridha Sridharan, Simon Lucey |
3DV | 4 |
| 2017 | Two-stage facial age prediction using group-specific featuresabstractA novel two-stage age prediction approach with group-specific features is proposed in this paper. Aging process is captured through a highly discriminating feature representation that models shape, appearance, skin spots, and wrinkles. The two-stage method consists of a multi-class Support Vector Machine (SVM) to predict the age bracket while the final age prediction is carried out using Support Vector Regression (SVR). The novelty of our work is that the feature extraction is group-specific and can therefore be tailored to each age bracket in the specific age prediction step. The FG-NET Aging dataset was used to evaluate the proposed method and an impressive mean absolute error (MAE) of 3.98 was achieved. Our approach outperforms the current state-of-the-art while increasing the robustness to blur, expression and lighting variation with local phase features. Jhony K. Pontes, Clinton Fookes, Alceu S. Britto Jr., Alessandro L. Koerich |
ICASSP | 2 |
| 2017 | Deep features-based expression-invariant tied factor analysis for emotion recognitionabstractVideo-based facial expression recognition is an open research challenge not solved by the current state-of-the-art. On the other hand, static image based emotion recognition is highly important when videos are not available and human emotions need to be determined from a single shot only. This paper proposes sequential-based and image-based tied factor analysis frameworks with a deep network that simultaneously addresses these two problems. For video-based data, we first extract deep convolutional temporal appearance features from image sequences and then these features are fed into a generative model that constructs a low-dimensional observed space for all individuals, depending on the facial expression sequences. After learning the sequential expression components of the transition matrices among the expression manifolds, we use a Gaussian probabilistic approach to design an efficient classifier for temporal facial expression recognition. Furthermore, we analyse the utility of proposed video-based methods for image-based emotion recognition learning static tied factor analysis parameters. Meanwhile, this model can be used to predict the expressive face image sequences from given neutral faces. Recognition results achieved on three public benchmark databases: CK+, JAFFE, and FER2013, clearly indicate our approach achieves effective performance over the current techniques of handling sequential and static facial expression variations. Sarasi Munasinghe, Clinton Fookes, Sridha Sridharan |
IJCB | 2 |
| 2017 | A cascaded long short-term memory (LSTM) driven generic visual question answering (VQA)abstractA cascaded long short-term memory (LSTM) architecture with discriminant feature learning is proposed for the task of question answering on real world images. The proposed LSTM architecture jointly learns visual features and parts of speech (POS) tags of question words or tokens. Also, dimensionality of deep visual features is reduced by applying Principal Component Analysis (PCA) technique. In this manner, the proposed question answering model captures the generic pattern of question for a given context of image which is just not constricted within the training dataset. Empirical outcome shows that this kind of approach significantly improves the accuracy. It is believed that this kind of generic learning is a step towards a real-world visual question answering (VQA) system which will perform well for all possible forms of open-ended natural language queries. Iqbal Chowdhury, Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan |
ICIP | 3 |
| 2017 | Deep discovery of facial motions using a shallow embedding layerabstractUnique encoding of the dynamics of facial actions has potential to provide a spontaneous facial expression recognition system. The most promising existing approaches rely on deep learning of facial actions. However, current approaches are often computationally intensive and require a great deal of memory/processing time, and typically the temporal aspect of facial actions are often ignored, despite the potential wealth of information available from the spatial dynamic movements and their temporal evolution over time from neutral state to apex state. To tackle aforementioned challenges, we propose a deep learning framework by using the 3D convolutional filters to extract spatio-temporal features, followed by the LSTM network which is able to integrate the dynamic evolution of short-duration of spatio-temporal features as an emotion progresses from the neutral state to the apex state. In order to reduce the redundancy of parameters and accelerate the learning of the recurrent neural network, we propose a shallow embedding layer to reduce the number of parameters in the LSTM by up to 98% without sacrificing recognition accuracy. As the fully connected layer approximately contains 95% of the parameters in the network, we decrease the number of parameters in this layer before passing features to the LSTM network, which significantly improves training speed and enables the possibility of deploying a state of the art deep network on real-time applications. We evaluate our proposed framework on the DISFA and UNBC-McMaster Shoulder pain datasets. Afsane Ghasemi, Mahsa Baktash, Simon Denman, Sridha Sridharan, Dung Nguyen Tien, Clinton Fookes |
ICIP | 6 |
| 2017 | Single image depth prediction using super-column super-pixel featuresabstractDepth prediction from a single monocular image is a challenging yet valuable task, as often a depth sensor is not available. The state-of-the-art approach [1] combines a deep fully convolutional network (DFCN) with a conditional random field (CRF), allowing the CRF to correct and smooth the depth values estimated by the DFCN according to efficient contextual modeling. However, using the output of the DFCN as unary input for CRF is limited by using only the last layer of the DFCN. The middle layers of the DFCN have been shown to carry useful information for other scene understanding tasks, which may help to improve the prediction quality. This paper proposes a novel super-column superpixel (SCSP) feature that is the combination of multiple layers of the DFCN after a super-pixel pooling process. The proposed approach based on the SCSP features reduces the root mean square (rms) error of the prediction by more than 16% in NYUv2 dataset. Xufeng Guo, Kien Nguyen Thanh, Simon Denman, Clinton Fookes, Sridha Sridharan |
ICIP | 4 |
| 2017 | Facial analysis in the wild with LSTM networksabstractThe promise of computer vision systems to efficiently and accurately recognize faces and facial variations in naturally occurring circumstances still remains elusive. In this paper we present two separate systems for face analysis, both of which use Long Short Term Memory (LSTM) Networks: unconstrained video-based face verification (FaceVideoModel) and spontaneous facial expression recognition (ExpModel). Since LSTM models have influential ability to capture sequential patterns, our results prove such LSTM models have significant advantages over other proposed models in the state-of-the-art for facial analysis in the wild. On the recently introduced Youtube Faces database our FaceModel achieves an accuracy of 98.70% for face verification with a value of 99.94% for the Area Under Curve (AUC) and 1.2% Equal Error Rate (EER) which is the best performance on this database compared to other recently proposed methods. Experimental results achieved through the proposed ExpModel on the challenging FER2013 dataset, including the CK+ database, also demonstrate the effectiveness of our deep model for facial expression recognition. Sarasi Kankanamge, Clinton Fookes, Sridha Sridharan |
ICIP | 2 |
| 2017 | Gate connected convolutional neural network for object trackingabstractConvolutional neural networks (CNNs) have been employed in visual tracking due to their rich levels of feature representation. While the learning capability of a CNN increases with its depth, unfortunately spatial information is diluted in deeper layers which hinders its important ability to localise targets. To successfully manage this trade-off, we propose a novel residual network based gating CNN architecture for object tracking. Our deep model connects the front and bottom convolutional features with a gate layer. This new network learns discriminative features while reducing the spatial information lost. This architecture is pre-trained to learn generic tracking characteristics. In online tracking, an efficient domain adaptation mechanism is used to accurately learn the target appearance with limited samples. Extensive evaluation performed on a publicly available benchmark dataset demonstrates our proposed tracker outperforms state-of-the-art approaches. Thanikasalam Kokul, Clinton Fookes, Sridha Sridharan, Amirthalingam Ramanan, Amalka Pinidiyaarachchi |
ICIP | 2 |
| 2017 | From Affine Rank Minimization Solution to Sparse ModelingabstractCompressed sensing is a simple and efficient technique that has a number of applications in signal processing and machine learning. In machine learning it provides answers to questions such as: "under what conditions is the sparse representation of data efficient?", "when is learning a large margin classifier directly on the compressed domain possible?", and "why does a large margin classifier learn more effectively if the data is sparse?". This work tackles the problem of feature representation from the context of sparsity and affine rank minimization by leveraging compressed sensing from the learning perspective in order to provide answers to the aforementioned questions. We show, for a full-rank signal, the high dimensional sparse representation of data is efficient because from the classifiers viewpoint such a representation is in fact a low dimensional problem. We provide practical bounds on the linear classifier to investigate the relationship between the SVM classifier in the high dimensional and compressed domains and show for the high dimensional sparse signals, when the bounds are tight, directly learning in the compressed domain is possible. Iman Abbasnejad, Sridha Sridharan, Simon Denman, Clinton Fookes, Simon Lucey |
WACV | 4 |
| 2017 | Two Stream LSTM: A Deep Fusion Framework for Human Action RecognitionabstractIn this paper we address the problem of human action recognition from video sequences. Inspired by the exemplary results obtained via automatic feature learning and deep learning approaches in computer vision, we focus our attention towards learning salient spatial features via a convolutional neural network (CNN) and then map their temporal relationship with the aid of Long-Short-Term-Memory (LSTM) networks. Our contribution in this paper is a deep fusion framework that more effectively exploits spatial features from CNNs with temporal features from LSTM models. We also extensively evaluate their strengths and weaknesses. We find that by combining both the sets of features, the fully connected features effectively act as an attention mechanism to direct the LSTM to interesting parts of the convolutional feature sequence. The significance of our fusion method is its simplicity and effectiveness compared to other state-of-the-art methods. The evaluation results demonstrate that this hierarchical multi stream fusion method has higher performance compared to single stream mapping methods allowing it to achieve high accuracy outperforming current state-of-the-art methods in three widely used databases: UCF11, UCFSports, jHMDB. Harshala Gammulle, Simon Denman, Sridha Sridharan, Clinton Fookes |
WACV | 4 |
| 2017 | Deep Context Modeling for Semantic SegmentationabstractDeep convolutional neural networks (DCNNs) have been employed in many computer vision tasks with great success due to their robustness in feature learning. One of the advantages of DCNNs is their representation robustness to object locations, which is useful for object recognition tasks. However, this also discards spatial information, which is useful when dealing with topological information of the image (e.g. scene parsing, face recognition). Adopting graphical models (GMs) to incorporate spatial and contextual information into the DCNNs is expected to improve the performance of DCNN-based computer vision tasks. Recent research has shown that combining DCNNs and Conditional Random Fields (CRFs) can significantly improve scene parsing accuracy. This is achieved either through the combination of their independent outputs or through their application as a cascade. In this work, we propose a novel strategy to incorporate CRFs deeper inside DCNNs by modeling a CRF as a DCNN layer which is pluggable into any layer of a DCNN. This implants spatial and contextual information into the DCNN, allowing end-to-end training, better controlling the spatial constraints and improving segmentation accuracy. The new strategy for coupling graphical models with the state-of-the-art fully convolutional neural network has shown promising results on the PASCAL-Context dataset. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan |
WACV | 2 |
| 2017 | Deep Spatio-Temporal Features for Multimodal Emotion RecognitionabstractAutomatic emotion recognition has attracted great interest and numerous solutions have been proposed, most of which focus either individually on facial expression or acoustic information. While more recent research has considered multimodal approaches, individual modalities are often combined only by simple fusion at the feature and/or decision-level. In this paper, we introduce a novel approach using 3-dimensional convolutional neural networks (C3Ds) to model the spatio-temporal information, cascaded with multimodal deep-belief networks (DBNs) that can represent the audio and video streams. Experiments conducted on the eNTERFACE multimodal emotion database demonstrate that this approach leads to improved multimodal emotion recognition performance and significantly outperforms recent state-of-the-art proposals. Dung Nguyen Tien, Kien Nguyen Thanh, Sridha Sridharan, Afsane Ghasemi, David Dean, Clinton Fookes |
WACV | 6 |
| 2017 | Fine-grained action recognition of boxing punches from depth imagery
Soudeh Kasiri Bidhendi, Clinton Fookes, Sridha Sridharan, Stuart Morgan |
Comput. Vis. Image Underst. | 2 |
| 2017 | Long range iris recognition: A survey
Kien Nguyen Thanh, Clinton Fookes, Raghavender R. Jillela, Sridha Sridharan, Arun Ross |
Pattern Recognit. | 2 |
| 2016 | Deeper and wider fully convolutional network coupled with conditional random fields for scene labelingabstractDeep convolutional neural networks (DCNNs) have been employed in many computer vision tasks with great success due to their robustness in feature learning. One of the advantages of DCNNs is their representation robustness to object locations, which is useful for object recognition tasks. However, this also discards spatial information, which is useful when dealing with topological information of the image (e.g. scene labeling, face recognition). In this paper, we propose a deeper and wider network architecture to tackle the scene labeling task. The depth is achieved by incorporating predictions from multiple early layers of the DCNN. The width is achieved by combining multiple outputs of the network. We then further refine the parsing task by adopting graphical models (GMs) as a post-processing step to incorporate spatial and contextual information into the network. The new strategy for a deeper, wider convolutional network coupled with graphical models has shown promising results on the PASCAL-Context dataset. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan |
ICIP | 2 |
| 2016 | A robust UAV landing site detection system using mid-level discriminative patchesabstractThe forced landing problem has become one of the main impediments to UAV's entering civilian airspace. Unfortunately there is no robust forced landing site detection system that will reliably detect a safe landing site. One of the main reasons for this is the difficulty in considering the various classes of surface, to determine whether they are safe or not. We propose a robust UAV landing site detection system using midlevel discriminative patches. The training and tuning process uses a dataset containing 1600 randomly selected Google map images with weak labels.We then show how the output from multiple mid-level discriminative patch detectors can be combined to indicate the level or danger for a given region. The proposed technique reliably detects safe landing areas in UAV imagery, and achieves improved performance over the state-of-the art. The proposed system outperforms the baseline system by 29.4% for completeness and 33.9% for correctness, and is invariant to the changes of illumination, sharpness and resolution of images. Xufeng Guo, Simon Denman, Clinton Fookes, Sridha Sridharan |
ICPR | 3 |
| 2016 | Speakers In The Wild (SITW): The QUT Speaker Recognition SystemabstractThis paper presents the QUT speaker recognition system, as a competing system in the Speakers In The Wild (SITW) speaker recognition challenge. Our proposed system achieved an overall ranking of second place, in the main core-core condition evaluations of the SITW challenge. This system uses an ivector/ PLDA approach, with domain adaptation and a deep neural network (DNN) trained to provide feature statistics. The statistics are accumulated by using class posteriors from the DNN, in place of GMM component posteriors in a typical GMM UBM i-vector/PLDA system. Once the statistics have been collected, the i-vector computation is carried out as in a GMM-UBM based system. We apply domain adaptation to the extracted i-vectors to ensure robustness against dataset variability, PLDA modelling is used to capture speaker and session variability in the i-vector space, and the processed i-vectors are compared using the batch likelihood ratio. The final scores are calibrated to obtain the calibrated likelihood scores, which are then used to carry out speaker recognition and evaluate the performance of the system. Finally, we explore the practical application of our system to the core-multi condition recordings of the SITW data and propose a technique for speaker recognition in recordings with multiple speakers. Houman Ghaemmaghami, Ivan Himawan, David Dean, Ahilan Kanagasundaram, Sridha Sridharan, Clinton Fookes |
INTERSPEECH | 7 |
| 2016 | Short Utterance Variance Modelling and Utterance Partitioning for PLDA Speaker VerificationabstractThis paper analyses the short utterance probabilistic linear discriminant analysis (PLDA) speaker verification with utterance partitioning and short utterance variance (SUV) modelling approaches. Experimental studies have found that instead of using single long-utterance as enrolment data, if long enrolled utterance is partitioned into multiple short utterances and average of short utterance i-vectors is used as enrolled data, that improves the Gaussian PLDA (GPLDA) speaker verification. This is because short utterance i-vectors have speaker, session and utterance variations, and utterance-partitioning approach compensates the utterance variation. Subsequently, SUV-PLDA is also studied with utterance partitioning approach, and utterance partitioning-based SUV-GPLDA system shows relative improvement of 9% and 16% in EER for NIST 2008 and NIST 2010 truncated 10sec-10sec evaluation condition as utterance partitioning approach compensates the utterance variation and SUV modelling approach compensates the mismatch between full-length development data and short-length evaluation data. Ahilan Kanagasundaram, David Dean, Sridha Sridharan, Clinton Fookes, Ivan Himawan |
INTERSPEECH | 4 |
| 2016 | Discovery of facial motions using deep machine perceptionabstractDeep, intuitive understanding of facial motions has the potential to provide an intelligent facial expression system as well as a unique encoding of the dynamics of facial actions. The most promising existing approaches rely on extracting hand crafted features; and existing approaches typically work best in constrained conditions and do not generalise well to varying environmental conditions which make them poorly suited to applications such as real-time human robot interactions. In this paper, we propose a multi-label deep learning based facial action detector, which along with a linear SVM classifier outperforms state of the art approaches such as HOG and LBP. We show that our approach can be generalized to other datasets by learning inner data structure, encoding facial actions, and providing a hierarchical representation of facial features. Our experimental results also demonstrate the efficiency of using image patches, which results in faster learning convergence while outperforms holistic approaches. We evaluate our proposed frame-work on the DISFA and CK+ datasets. Afsane Ghasemi, Simon Denman, Sridha Sridharan, Clinton Fookes |
WACV | 4 |
| 2016 | Detecting rare events using Kullback-Leibler divergence: A weakly supervised approach
Jingxin Xu, Simon Denman, Clinton Fookes, Sridha Sridharan |
Expert Syst. Appl. | 3 |
| 2016 | A flexible hierarchical approach for facial age estimation based on multiple features
Jhony K. Pontes, Alceu S. Britto Jr., Clinton Fookes, Alessandro L. Koerich |
Pattern Recognit. | 3 |
| 2016 | Discovering Team Structures in Soccer from Spatiotemporal DataabstractIn team sports like soccer, utilizing tracking data for analysis is challenging due to the dynamic and multi-agent nature of the data. The biggest issue surrounds the changing of positions or “roles” between players on a frame-to-frame basis, which causes misalignment of the data and makes it difficult to perform team analysis. In this paper, we present an unsupervised method to learn a formation template which allows us to “align” the tracking data at the frame level. Not only does this approach give important contextual information to facilitate large-scale analysis (e.g., we know when a player is in the left-wing position compared to left-back), it also yields the team structure or “formation” which serves as a strong descriptor for identifying a team's style. The utility of the approach is demonstrated on a full season of player and ball tracking data from a professional soccer league consisting of over 21.5 million frames of player tracking data. Alina Bialkowski, Patrick Lucey, Peter Carr 0001, Iain A. Matthews, Sridha Sridharan, Clinton Fookes |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2015 | Large scale monitoring of crowds and building utilisation: A new database and distributed approachabstractPublic buildings and large infrastructure are typically monitored by tens or hundreds of cameras, all capturing different physical spaces and observing different types of interactions and behaviours. However to date, in large part due to limited data availability, crowd monitoring and operational surveillance research has focused on single camera scenarios which are not representative of real-world applications. In this paper we present a new, publicly available database for large scale crowd surveillance. Footage from 12 cameras for a full work day covering the main floor of a busy university campus building, including an internal and external foyer, elevator foyers, and the main external approach are provided; alongside annotation for crowd counting (single or multi-camera) and pedestrian flow analysis for 10 and 6 sites respectively. We describe how this large dataset can be used to perform distributed monitoring of building utilisation, and demonstrate the potential of this dataset to understand and learn the relationship between different areas of a building. Simon Denman, Clinton Fookes, David Ryan, Sridha Sridharan |
AVSS | 2 |
| 2015 | Searching for semantic person queries using channel representationsabstractIt is not uncommon to hear a person of interest described by their height, build, and clothing (i.e. type and colour). These semantic descriptions are commonly used by people to describe others, as they are quick to relate and easy to understand. However such queries are not easily utilised within intelligent surveillance systems as they are difficult to transform into a representation that can be searched for automatically in large camera networks. In this paper we propose a novel approach that transforms such a semantic query into an avatar that is searchable within a video stream, and demonstrate state-of-the-art performance for locating a subject in video based on a description. Simon Denman, Michael Halstead, Clinton Fookes, Sridha Sridharan |
ICASSP | 3 |
| 2015 | Detecting rare events using Kullback-Leibler divergenceabstractOne main challenge in developing a system for visual surveillance event detection is the annotation of target events in the training data. By making use of the assumption that events with security interest are often rare compared to regular behaviours, this paper presents a novel approach by using Kullback-Leibler (KL) divergence for rare event detection in a weakly supervised learning setting, where only clip-level annotation is available. It will be shown that this approach outperforms state-of-the-art methods on a popular real-world dataset, while preserving real time performance. Jingxin Xu, Simon Denman, Clinton Fookes, Sridha Sridharan |
ICASSP | 3 |
| 2015 | Combat sports analytics: Boxing punch classification using overhead depthimageryabstractIn competitive combat sporting environments like boxing, the statistics on a boxer's performance, including the amount and type of punches thrown, provide a valuable source of data and feedback which is routinely used for coaching and performance improvement purposes. This paper presents a robust framework for the automatic Classification of a boxer's punches. Overhead depth imagery is employed to alleviate challenges associated with occlusions, and robust body-part tracking is developed for the noisy time-of-flight sensors. Punch recognition is addressed through both a multi-class SVM and Random Forest classifiers. coarse-to-fine hierarchical SVM classifier is presented based on prior knowledge of boxing punches. This framework has been applied to shadow boxing image sequences taken at the Australian Institute of Sport with 8 elite boxers. Results demonstrate the effectiveness of the proposed approach, with the hierarchical SVM classifier yielding a 96% accuracy, signifying its suitability for analysing athletes punches in boxing bouts. Soudeh Kasiri Bidhendi, Clinton Fookes, Stuart Morgan, David T. Martin, Sridha Sridharan |
ICIP | 2 |
| 2015 | Improving deep convolutional neural networks with unsupervised feature learningabstractThe latest generation of Deep Convolutional Neural Networks (DCNN) have dramatically advanced challenging computer vision tasks, especially in object detection and object classification, achieving state-of-the-art performance in several computer vision tasks including text recognition, sign recognition, face recognition and scene understanding. The depth of these supervised networks has enabled learning deeper and hierarchical representation of features. In parallel, unsupervised deep learning such as Convolutional Deep Belief Network (CDBN) has also achieved state-of-the-art in many computer vision tasks. However, there is very limited research on jointly exploiting the strength of these two approaches. In this paper, we investigate the learning capability of both methods. We compare the output of individual layers and show that many learnt filters and outputs of the corresponding level layer are almost similar for both approaches. Stacking the DCNN on top of unsupervised layers or replacing layers in the DCNN with the corresponding learnt layers in the CDBN can improve the recognition/classification accuracy and training computational expense. We demonstrate the validity of the proposal on ImageNet dataset. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan |
ICIP | 2 |
| 2015 | Class-specific sparse codes for representing activitiesabstractIn this paper we investigate the effectiveness of class specific sparse codes in the context of discriminative action classification. The bag-of-words representation is widely used in activity recognition to encode features, and although it yields state-of-the art performance with several feature descriptors it still suffers from large quantization errors and reduces the overall performance. Recently proposed sparse representation methods have been shown to effectively represent features as a linear combination of an over complete dictionary by minimizing the reconstruction error. In contrast to most of the sparse representation methods which focus on Sparse-Reconstruction based Classification (SRC), this paper focuses on a discriminative classification using a SVM by constructing class-specific sparse codes for motion and appearance separately. Experimental results demonstrates that separate motion and appearance specific sparse coefficients provide the most effective and discriminative representation for each class compared to a single class-specific sparse coefficients. Sabanadesan Umakanthan, Simon Denman, Clinton Fookes, Sridha Sridharan |
ICIP | 3 |
| 2015 | Complete-linkage clustering for voice activity detection in audio and visual speechabstractWe propose a novel technique for conducting robust voice activity detection (VAD) in high-noise recordings. We use Gaussian mixture modeling (GMM) to train two generic models; speech and non-speech. We then score smaller segments of a given (unseen) recording against each of these GMMs to obtain two respective likelihood scores for each segment. These scores are used to compute a dissimilarity measure between pairs of segments and to carry out complete-linkage clustering of the segments into speech and non-speech clusters. We compare the accuracy of our method against state-of-the-art and standardised VAD techniques to demonstrate an absolute improvement of 15% in half-total error rate (HTER) over the best performing baseline system and across the QUT-NOISE-TIMIT database. We then apply our approach to the Audio-Visual Database of American English (AVDBAE) to demonstrate the performance of our algorithm in using visual, audio-visual or a proposed fusion of these features. Houman Ghaemmaghami, David Dean, Shahram Kalantari, Sridha Sridharan, Clinton Fookes |
INTERSPEECH | 5 |
| 2015 | Cross database training of audio-visual hidden Markov models for phone recognitionabstractSpeech recognition can be improved by using visual information in the form of lip movements of the speaker in addition to audio information.To date, state-of-the-art techniques for audio-visual speech recognition continue to use audio and visual data of the same database for training their models.In this paper, we present a new approach to make use of one modality of an external dataset in addition to a given audio-visual dataset.By so doing, it is possible to create more powerful models from other extensive audio-only databases and adapt them on our comparatively smaller multi-stream databases.Results show that the presented approach outperforms the widely adopted synchronous hidden Markov models (HMM) trained jointly on audio and visual data of a given audio-visual database for phone recognition by 29% relative.It also outperforms the external audio models trained on extensive external audio datasets and also internal audio models by 5.5% and 46% relative respectively.We also show that the proposed approach is beneficial in noisy environments where the audio source is affected by the environmental noise. Shahram Kalantari, David Dean, Houman Ghaemmaghami, Sridha Sridharan, Clinton Fookes |
INTERSPEECH | 5 |
| 2015 | An evaluation of crowd counting methods, features and regression models
David Ryan, Simon Denman, Sridha Sridharan, Clinton Fookes |
Comput. Vis. Image Underst. | 4 |
| 2015 | A framework for model integration and holistic modelling of socio-technical systems
Paul Pao-Yen Wu, Clinton Fookes, Jegar Pitchforth, Kerrie L. Mengersen |
Decis. Support Syst. | 2 |
| 2015 | Automatic surveillance in transportation hubs: No longer just about catching the bad guy
Simon Denman, Tristan Kleinschmidt, David Ryan, Paul Barnes, Sridha Sridharan, Clinton Fookes |
Expert Syst. Appl. | 6 |
| 2015 | Searching for people using semantic soft biometric descriptions
Simon Denman, Michael Halstead, Clinton Fookes, Sridha Sridharan |
Pattern Recognit. Lett. | 3 |
| 2015 | An Efficient and Robust System for Multiperson Event Detection in Real-World Indoor Surveillance ScenesabstractDue to the popularity of security cameras in public places, it is of interest to design an intelligent system that can efficiently detect events automatically. This paper proposes a novel algorithm for multiperson event detection. To ensure greater than real-time performance, features are extracted directly from compressed MPEG video. A novel histogram-based feature descriptor that captures the angles between extracted particle trajectories is proposed, which allows us to capture motion patterns for multiperson events in the video. To alleviate the need for fine-grained annotation, we propose the use of labeled latent Dirichlet allocation, a weakly supervised method that allows the use of coarse temporal annotations, which are much simpler to obtain. This novel system is able to run at ~10 times real time, while preserving state-of-the-art detection performance for multiperson events on a 100-h real-world surveillance data set (TRECVid surveillance event detection). Jingxin Xu, Simon Denman, Sridha Sridharan, Clinton Fookes |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2015 | Score-Level Multibiometric Fusion Based on Dempster-Shafer Theory Incorporating Uncertainty FactorsabstractWhile existing multibiometic Dempster-Shafer theory fusion approaches have demonstrated promising performance, they do not model the uncertainty appropriately, suggesting that further improvement can be achieved. This research seeks to develop a unified framework for multimodal biometric fusion to take advantage of the uncertainty concept of Dempster-Shafer theory, improving the performance of multibiometric authentication systems. Modeling uncertainty as a function of uncertainty factors affecting the recognition performance of the biometric systems helps to address the uncertainty of the data and the confidence of the fusion outcome. A weighted combination of quality measures and classifiers performance (equal error rate) is proposed to encode the uncertainty concept to improve the fusion. We also found that quality measures contribute unequally to the recognition performance; thus, selecting only significant factors and fusing them with a Dempster-Shafer approach to generate an overall quality score play an important role in the success of uncertainty modeling. The proposed approach achieved a competitive performance (approximate 1% EER) in comparison with other Dempster-Shafer-based approaches and other conventional fusion approaches. Kien Nguyen Thanh, Simon Denman, Sridha Sridharan, Clinton Fookes |
IEEE Trans. Hum. Mach. Syst. | 4 |
| 2014 | An MRF based abnormal event detection approach using motion and appearance featuresabstractAbnormal event detection has attracted a lot of attention in the computer vision research community during recent years due to the increased focus on automated surveillance systems to improve security in public places. Due to the scarcity of training data and the definition of an abnormality being dependent on context, abnormal event detection is generally formulated as a data-driven approach where activities are modeled in an unsupervised fashion during the training phase. In this work, we use a Gaussian mixture model (GMM) to cluster the activities during the training phase, and propose a Gaussian mixture model based Markov random field (GMM-MRF) to estimate the likelihood scores of new videos in the testing phase. Further-more, we propose two new features: optical acceleration, and the histogram of optical flow gradients; to detect the presence of any abnormal objects and speed violations in the scene. We show that our proposed method outperforms other state of the art abnormal event detection algorithms on publicly available UCSD dataset. Hajananth Nallaivarothayan, Clinton Fookes, Simon Denman, Sridha Sridharan |
AVSS | 2 |
| 2014 | Locating People in Video from Semantic Descriptions: A New Database and ApproachabstractThe location of previously unseen and unregistered individuals in complex camera networks from semantic descriptions is a time consuming and often inaccurate process carried out by human operators, or security staff on the ground. To promote the development and evaluation of automated semantic description based localisation systems, we present a new, publicly available, unconstrained 110 sequence database, collected from 6 stationary cameras. Each sequence contains detailed semantic information for a single search subject who appears in the clip (gender, age, height, build, hair and skin colour, clothing type, texture and colour), and between 21 and 290 frames for each clip are annotated with the target subject location (over 11, 000 frames are annotated in total). A novel approach for localising a person given a semantic query is also proposed and demonstrated on this database. The proposed approach incorporates clothing colour and type (for clothing worn below the waist), as well as height and build to detect people. A method to assess the quality of candidate regions, as well as a symmetry driven approach to aid in modelling clothing on the lower half of the body, is proposed within this approach. An evaluation on the proposed dataset shows that a relative improvement in localisation accuracy of up to 21% is achieved over the baseline technique. Michael Halstead, Simon Denman, Sridha Sridharan, Clinton Fookes |
ICPR | 4 |
| 2014 | Multiple Instance Dictionary Learning for Activity RepresentationabstractThis paper presents an effective feature representation method in the context of activity recognition. Efficient and effective feature representation plays a crucial role not only in activity recognition, but also in a wide range of applications such as motion analysis, tracking, 3D scene understanding etc. In the context of activity recognition, local features are increasingly popular for representing videos because of their simplicity and efficiency. While they achieve state-of-the-art performance with low computational requirements, their performance is still limited for real world applications due to a lack of contextual information and models not being tailored to specific activities. We propose a new activity representation framework to address the shortcomings of the popular, but simple bag-of-words approach. In our framework, first multiple instance SVM (mi-SVM) is used to identify positive features for each action category and the k-means algorithm is used to generate a codebook. Then locality-constrained linear coding is used to encode the features into the generated codebook, followed by spatio-temporal pyramid pooling to convey the spatio-temporal statistics. Finally, an SVM is used to classify the videos. Experiments carried out on two popular datasets with varying complexity demonstrate significant performance improvement over the base-line bag-of-feature method. Sabanadesan Umakanthan, Simon Denman, Clinton Fookes, Sridha Sridharan |
ICPR | 3 |
| 2014 | Analysis of passenger group behaviour and its impact on passenger flow using an agent-based model
Clinton Fookes, Vikas Reddy, Prasad K. D. V. Yarlagadda |
SIMULTECH | 2 |
| 2014 | Agent-based modelling of aircraft boarding methods
Serter Iyigunlu, Clinton Fookes, Prasad K. D. V. Yarlagadda |
SIMULTECH | 2 |
| 2014 | Local inter-session variability modelling for object classificationabstractObject classification is plagued by the issue of session variation. Session variation describes any variation that makes one instance of an object look different to another, for instance due to pose or illumination variation. Recent work in the challenging task of face verification has shown that session variability modelling provides a mechanism to overcome some of these limitations. However, for computer vision purposes, it has only been applied in the limited setting of face verification. In this paper we propose a local region based intersession variability (ISV) modelling approach, and apply it to challenging real-world data. We propose a region based session variability modelling approach so that local session variations can be modelled, termed Local ISV. We then demonstrate the efficacy of this technique on a challenging real-world fish image database which includes images taken underwater, providing significant real-world session variations. This Local ISV approach provides a relative performance improvement of, on average, 23% on the challenging MOBIO, Multi-PIE and SCface face databases. It also provides a relative performance improvement of 35% on our challenging fish image dataset. Kaneswaran Anantharajah, ZongYuan Ge, Chris McCool, Simon Denman, Clinton Fookes, Peter I. Corke, Dian Tjondronegoro, Sridha Sridharan |
WACV | 5 |
| 2014 | Scene invariant multi camera crowd counting
David Ryan, Simon Denman, Clinton Fookes, Sridha Sridharan |
Pattern Recognit. Lett. | 3 |
| 2014 | Real-time video event detection in crowded scenes using MPEG derived features: A multiple instance learning approach
Jingxin Xu, Simon Denman, Vikas Reddy, Clinton Fookes, Sridha Sridharan |
Pattern Recognit. Lett. | 4 |
| 2014 | Optimal Camera Planning Under Versatile User Constraints in Multi-Camera Image Processing SystemsabstractThe selection of optimal camera configurations (camera locations, orientations, etc.) for multi-camera networks remains an unsolved problem. Previous approaches largely focus on proposing various objective functions to achieve different tasks. Most of them, however, do not generalize well to large scale networks. To tackle this, we propose a statistical framework of the problem as well as propose a trans-dimensional simulated annealing algorithm to effectively deal with it. We compare our approach with a state-of-the-art method based on binary integer programming (BIP) and show that our approach offers similar performance on small scale problems. However, we also demonstrate the capability of our approach in dealing with large scale problems and show that our approach produces better results than two alternative heuristics designed to deal with the scalability issue of BIP. Last, we show the versatility of our approach using a number of specific scenarios. Junbin Liu, Sridha Sridharan, Clinton Fookes, Tim Wark |
IEEE Trans. Image Process. | 3 |
| 2013 | Feature-domain super-resolution for iris recognition
Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Simon Denman |
Comput. Vis. Image Underst. | 2 |
| 2013 | Evaluation of two-view geometry methods with automatic ground-truth generation
Ruan Lakemond, Clinton Fookes, Sridha Sridharan |
Image Vis. Comput. | 2 |
| 2012 | Use of brain computer interface to drive: preliminary resultsabstractThis paper reports on the implementation of a non-invasive electroencephalography-based brain-computer interface to control functions of a car in a driving simulator. The system is comprised of a Cleveland Medical Devices BioRadio 150 physiological signal recorder, a MATLAB-based BCI and an OKTAL SCANeR advanced driving experience simulator. Deanna Hood, Damian Joseph, Andry Rakotonirainy, Sridha Sridharan, Clinton Fookes |
AutomotiveUI | 5 |
| 2012 | Unusual Scene Detection Using Distributed Behaviour Model and Sparse RepresentationabstractThe ability to detect unusual events in surviellance footage as they happen is a highly desireable feature for a surveillance system. However, this problem remains challenging in crowded scenes due to occlusions and the clustering of people. In this paper, we propose using the Distributed Behavior Model (DBM), which has been widely used in computer graphics, for video event detection. Our approach does not rely on object tracking, and is robust to camera movements. We use sparse coding for classification, and test our approach on various datasets. Our proposed approach outperforms a state-of-the-art work which uses the social force model and Latent Dirichlet Allocation. Jingxin Xu, Simon Denman, Clinton Fookes, Sridha Sridharan |
AVSS | 3 |
| 2012 | Activity Analysis in Complicated Scenes Using DFT Coefficients of Particle TrajectoriesabstractModelling activities in crowded scenes is very challenging as object tracking is not robust in complicated scenes and optical flow does not capture long range motion. We propose a novel approach to analyse activities in crowded scenesusing a "bag of particle trajectories". Particle trajectoriesare extracted from foreground regions within short video clips using particle video, which estimates long rangemotion in contrast to optical flow which is only concerned with inter-frame motion. Our applications include temporal video segmentation and anomaly detection, and we perform our evaluation on several real-world datasets containing complicated scenes. We show that our approaches achieve state-of-the-art performance for both tasks. Jingxin Xu, Simon Denman, Sridha Sridharan, Clinton Fookes |
AVSS | 4 |
| 2012 | Feature-domain super-resolution framework for Gabor-based face and iris recognitionabstractThe low resolution of images has been one of the major limitations in recognising humans from a distance using their biometric traits, such as face and iris. Superresolution has been employed to improve the resolution and the recognition performance simultaneously, however the majority of techniques employed operate in the pixel domain, such that the biometric feature vectors are extracted from a super-resolved input image. Feature-domain superresolution has been proposed for face and iris, and is shown to further improve recognition performance by capitalising on direct super-resolving the features which are used for recognition. However, current feature-domain superresolution approaches are limited to simple linear features such as Principal Component Analysis (PCA) and Linear Discriminant Analysis (LDA), which are not the most discriminant features for biometrics. Gabor-based features have been shown to be one of the most discriminant features for biometrics including face and iris. This paper proposes a framework to conduct super-resolution in the non-linear Gabor feature domain to further improve the recognition performance of biometric systems. Experiments have confirmed the validity of the proposed approach, demonstrating superior performance to existing linear approaches for both face and iris biometrics. Kien Nguyen Thanh, Sridha Sridharan, Simon Denman, Clinton Fookes |
CVPR | 4 |
| 2012 | On the Statistical Determination of Optimal Camera Configurations in Large Scale Surveillance Networks
Junbin Liu, Clinton Fookes, Tim Wark, Sridha Sridharan |
ECCV (1) | 2 |
| 2012 | Modelling Passengers Flow at Airport Terminals - Individual Agent Decision Model for Stochastic Passenger Behaviour
Clinton Fookes, Tristan Kleinschmidt, Prasad K. D. V. Yarlagadda |
SIMULTECH | 2 |
| 2012 | Evaluation of image resolution and super-resolution on face recognition performance
Clinton Fookes, Frank Lin, Vinod Chandran, Sridha Sridharan |
J. Vis. Commun. Image Represent. | 1 |
| 2011 | Determining operational measures from multi-camera surveillance systems using soft biometricsabstractCCTV and surveillance networks are increasingly being used for operational as well as security tasks. One emerging area of technology that lends itself to operational analytics is soft biometrics. Soft biometrics can be used to describe a person and detect them throughout a sparse multi-camera network. This enables them to be used to perform tasks such as determining the time taken to get from point to point, and the paths taken through an environment by detecting and matching people across disjoint views. However, in a busy environment where there are 100's if not 1000's of people such as an airport, attempting to monitor everyone is highly unrealistic. In this paper we propose an average soft biometric, that can be used to identity people who look distinct, and are thus suitable for monitoring through a large, sparse camera network. We demonstrate how an average soft biometric can be used to identify unique people to calculate operational measures such as the time taken to travel from point to point. Simon Denman, Alina Bialkowski, Clinton Fookes, Sridha Sridharan |
AVSS | 3 |
| 2011 | Textures of optical flow for real-time anomaly detection in crowdsabstractAutomated visual surveillance of crowds is a rapidly growing area of research. In this paper we focus on motion representation for the purpose of abnormality detection in crowded scenes. We propose a novel visual representation called textures of optical flow. The proposed representation measures the uniformity of a flow field in order to detect anomalous objects such as bicycles, vehicles and skateboarders; and can be combined with spatial information to detect other forms of abnormality. We demonstrate that the proposed approach outperforms state-of-the-art anomaly detection algorithms on a large, publicly-available dataset. David Ryan, Simon Denman, Clinton Fookes, Sridha Sridharan |
AVSS | 3 |
| 2011 | 3D ellipsoid fitting for multi-view gait recognitionabstractGait recognition approaches continue to struggle with challenges including view-invariance, low-resolution data, robustness to unconstrained environments, and fluctuating gait patterns due to subjects carrying goods or wearing different clothes. Although computationally expensive, model based techniques offer promise over appearance based techniques for these challenges as they gather gait features and interpret gait dynamics in skeleton form. In this paper, we propose a fast 3D ellipsoidal-based gait recognition algorithm using a 3D voxel model derived from multi-view silhouette images. This approach directly solves the limitations of view dependency and self-occlusion in existing ellipse fitting model-based approaches. Voxel models are segmented into four components (left and right legs, above and below the knee), and ellipsoids are fitted to each region using eigenvalue decomposition. Features derived from the ellipsoid parameters are modeled using a Fourier representation to retain the temporal dynamic pattern for classification. We demonstrate the proposed approach using the CMU MoBo database and show that an improvement of 15-20% can be achieved over a 2D ellipse fitting baseline. Sabesan Sivipalan, Daniel Chen 0002, Simon Denman, Sridha Sridharan, Clinton Fookes |
AVSS | 5 |
| 2011 | Gait energy volumes and frontal gait recognition using depth imagesabstractGait energy images (GEIs) and its variants form the basis of many recent appearance-based gait recognition systems. The GEI combines good recognition performance with a simple implementation, though it suffers problems inherent to appearance-based approaches, such as being highly view dependent. In this paper, we extend the concept of the GEI to 3D, to create what we call the gait energy volume, or GEV. A basic GEV implementation is tested on the CMU MoBo database, showing improvements over both the GEI baseline and a fused multi-view GEI approach. We also demonstrate the efficacy of this approach on partial volume reconstructions created from frontal depth images, which can be more practically acquired, for example, in biometric portals implemented with stereo cameras, or other depth acquisition systems. Experiments on frontal depth images are evaluated on an in-house developed database captured using the Microsoft Kinect, and demonstrate the validity of the proposed approach. Sabesan Sivipalan, Daniel Chen 0002, Simon Denman, Sridha Sridharan, Clinton Fookes |
IJCB | 5 |
| 2011 | Feature-domain super-resolution for iris recognitionabstractUncooperative iris identification systems at a distance suffer from poor resolution of the captured iris images, which significantly degrades iris recognition performance. Super-resolution techniques have been employed to enhance the resolution of iris images and improve the recognition performance. However, all existing super-resolution approaches proposed for the iris biometric super-resolve pixel intensity values. This paper considers transferring super-resolution of iris images from the intensity domain to the feature domain. By directly super-resolving only the features essential for recognition, and by incorporating domain specific information from iris models, improved recognition performance compared to pixel domain super-resolution can be achieved. This is the first paper to investigate the possibility of feature-domain super-resolution for iris recognition, and experiments confirm the validity of the proposed approach. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Simon Denman |
ICIP | 2 |
| 2011 | Quality-Driven Super-Resolution for Less Constrained Iris Recognition at a Distance and on the MoveabstractLess constrained iris identification systems at a distance and on the move suffer from poor resolution and poor quality of the captured iris images, which significantly degrades iris recognition performance. This paper proposes a new signal-level fusion approach which incorporates a quality score into a reconstruction-based super-resolution process to generate a high-resolution iris image from a low-resolution and quality inconsistent video sequence of an eye. A novel approach for assessing the focus level of the iris image, which is invariant to lighting and oclusion conditions, is introduced. The focus score is combined with several other quality factors to perform the quality weighted super-resolution where the highest quality frames contribute the greatest amount of information to the resulting high-resolution images without introducing spurious high-frequency components. Experiments conducted on the Multiple Biometric Grand Challenge portal dataset show that our proposed approach outperforms the traditional best quality frame selection approach and other existing state-of-the-art signal-level and score-level fusion approaches for recognition of less constrained iris at a distance and on the move. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Simon Denman |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2010 | Multi-Modal Object Tracking using Dynamic Performance MetricsabstractIntelligent surveillance systems typically use a single visual spectrum modality for their input. These systems work well in controlled conditions, but often fail when lighting is poor, or environmental effects such as shadows, dust or smoke are present. Thermal spectrum imagery is not as susceptible to environmental effects, however thermal imaging sensors are more sensitive to noise and they are only gray scale, making distinguishing between objects difficult. Several approaches to combining the visual and thermal modalities have been proposed, however they are limited by assuming that both modalities are perfuming equally well. When one modality fails, existing approaches are unable to detect the drop in performance and disregard the under performing modality. In this paper, a novel middle fusion approach for combining visual and thermal spectrum images for object tracking is proposed. Motion and object detection is performed on each modality and the object detection results for each modality are fused base on the current performance of each modality. Modality performance is determined by comparing the number of objects tracked by the system with the number detected by each mode, with a small allowance made for objects entering and exiting the scene. The tracking performance of the proposed fusion scheme is compared with performance of the visual and thermal modes individually, and a baseline middle fusion scheme. Improvement in tracking performance using the proposed fusion approach is demonstrated. The proposed approach is also shown to be able to detect the failure of an individual modality and disregard its results, ensuring performance is not degraded in such situations. Simon Denman, Clinton Fookes, Sridha Sridharan, David Ryan |
AVSS | 2 |
| 2010 | Crowd Counting Using Group Tracking and Local FeaturesabstractIn public venues, crowd size is a key indicator of crowd safety and stability. In this paper we propose a crowd counting algorithm that uses tracking and local features to count the number of people in each group as represented by a foreground blob segment, so that the total crowd estimate is the sum of the group sizes. Tracking is employed to improve the robustness of the estimate, by analysing the history of each group, including splitting and merging events. A simplified ground truth annotation strategy results in an approach with minimal setup requirements that is highly accurate. David Ryan, Simon Denman, Clinton Fookes, Sridha Sridharan |
AVSS | 3 |
| 2009 | Dynamic Performance Measures for Object Tracking SystemsabstractPerformance evaluation of object tracking systems is typically performed after the data has been processed, by comparing tracking results to ground truth. Whilst this approach is fine when performing offline testing, it does not allow for real-time analysis of the systems performance, which may be of use for live systems to either automatically tune the system or report reliability. In this paper, we propose three metrics that can be used to dynamically asses the performance of an object tracking system. Outputs and results from various stages in the tracking system are used to obtain measures that indicate the performance of motion segmentation, object detection and object matching. The proposed dynamic metrics are shown to accurately indicate tracking errors when visually comparing metric results to tracking output, and are shown to display similar trends to the ETISEO metrics when comparing different tracking configurations. Simon Denman, Clinton Fookes, Sridha Sridharan, Ruan Lakemond |
AVSS | 2 |
| 2009 | Affine Adaptation of Local Image Features Using the Hessian MatrixabstractLocal feature detectors that make use of derivative based saliency functions to locate points of interest typically require adaptation processes after initial detection in order to achieve scale and affine covariance. Affine adaptation methods have previously been proposed that make use of the second moment matrix to iteratively estimate the affine shape of local image regions. This paper shows that it is possible to use the Hessian matrix to estimate local affine shape in a similar fashion to the second moment matrix. The Hessian matrix requires significantly less computation effort to compute than the second moment matrix, allowing more efficient affine adaptation. It may also be more convenient to use the Hessian matrix, for example, when the Determinant of Hessian detector is used. Experimental evaluation shows that the Hessian matrix is very effective in increasing the efficiency of blob detectors such as the Determinant of Hessian detector, but less effective in combination with the Harris corner detector. Ruan Lakemond, Clinton Fookes, Sridha Sridharan |
AVSS | 2 |
| 2008 | 3D face verification using a free-parts approach
Chris McCool, Vinod Chandran, Sridha Sridharan, Clinton Fookes |
Pattern Recognit. Lett. | 4 |
| 2006 | A Multi-Class Tracker Using a Scalable Condensation FilterabstractTracking systems are typically targeted towards tracking a single class of object. In many real world situations, and in the ETISEO evaluation, it is advantageous to be able to track multiple classes of objects. In this paper we describe the adaptation of a single class tracking system to a multi-class tracking system, and describe a modified version of the condensation filter that can be used to track all objects, of all classes. We show that by using simple targeted detectors, we can achieve accurate tracking and can accurately distinguish between classes. Simon Denman, Vinod Chandran, Sridha Sridharan, Clinton Fookes |
AVSS | 4 |
| 2006 | Multi-view Intelligent Vehicle Surveillance SystemabstractThis paper presents a multi-view intelligent surveillance system used for the automatic tracking and monitoring of vehicles in a short-term parking lane. The system has the ability to track multiple vehicles in real-time across four cameras monitoring the area using a combination of both motion detection and optical flow modules. Automated alerts of events such as parking time violations, breaching of restricted areas or improper directional flow of traffic can be generated and communicated to attending security personnel. Results are shown using surveillance data captured from a real multi-camera network to illustrate the robust and real-time performance of the system. Simon Denman, Clinton Fookes, Jamie Cook, Chris Davoren, Anthony Mamic, Graeme Farquharson, Daniel Chen 0002, Brenden Chen, Sridha Sridharan |
AVSS | 2 |
| 2006 | The Role of Motion Models in Super-Resolving Surveillance Video for Face RecognitionabstractAlthough the use of super-resolution techniques has demonstrated the ability to improve face recognition accuracy when compared to traditional upsampling techniques, they are difficult to implement for real-time use due to their complexity and high computational demand. As a large portion of processing time is dedicated to registering the lowresolution images, many have adopted global motion models in order to improve efficiency. The drawback of such global models is that they can not accommodate for complex local motions, such as multiple objects moving independently across and static or dynamic background as frequently occurs in a surveillance environment. Local methods like optical flow can compensate for these situations, although it is achieved at the expense of computation time. In this paper, experiments have been carried out to investigate how motion models of different super-resolution reconstruction algorithms affect reconstruction error and face recognition rates in a surveillance environment. Results show that lower reconstruction error doesn't necessarily imply better recognition rates and the use of local motion models yields better recognition rates than global motion models. Frank Lin, Clinton Fookes, Vinod Chandran, Sridha Sridharan |
AVSS | 2 |
| 2006 | Human Face Reconstruction Using Bayesian Deformable ModelsabstractThis paper presents a Bayesian framework for 3D facial reconstruction. The framework iteratively deforms a generic face mesh to fit a set of range points representing a face. The generic mesh is generated from the extensive FRGC database of face images. The deformation process is conducted within a Bayesian framework and is driven by a Markov Chain Monte Carlo (MCMC) sampler which uses information from the likelihood and prior distributions of the generic face mesh. The paper presents results on the construction of a generic face model, the deformation framework and fitting results to both synthetic and real data. The results verify the effectiveness of the proposed technique, accurately deforming a generic face mesh to captured 3D data points of human faces. George Mamic, Clinton Fookes, Sridha Sridharan |
AVSS | 2 |
| 2006 | 3D Face Recognition using Log-Gabor TemplatesabstractThe use of Three Dimensional (3D) data allows new facial recognition algorithms to overcome factors such as pose and illumination variations which have plagued traditional 2D Face Recognition. In this paper a new method for providing insensitivity to expression variation in range images based on Log-Gabor Templates is presented. By decomposing a single image of a subject into 147 observations the reliance of the algorithm upon any particular part of the face is relaxed allowing high accuracy even in the presence of occulusions, distortions and facial expressions. Using the 3D database collected by University of Notre Dame for the Face Recognition Grand Challenge (FRGC), benchmarking results are presented showing superior performance of the proposed method. Comparisons showing the relative strength of the algorithm against two commercial and two academic 3D face recognition algorithms are also presented. algoritms are also presented. 1 Introduction Jamie Cook, Vinod Chandran, Clinton Fookes |
BMVC | 3 |
| 2006 | What Is the Average Human Face?
George Mamic, Clinton Fookes, Sridha Sridharan |
PSIVT | 2 |
| 2003 | Rigid Medical Image Registration And Its Association With Mutual InformationabstractImage registration plays a crucial role in the computer vision and medical imaging field where it is used to develop a spatial mapping between different sets of data. These transformations can range from simple rigid registrations to complex nonrigid deformations. Mutual information (MI) is a popular entropy-based similarity measure which has recently experienced a prolific expansion in a number of image registration applications. Stemming from information theory, this measure generally outperforms most other intensity-based measures in multimodal applications as it only assumes a statistical dependence between images. This paper provides a thorough introduction to the MI measure and its use in rigid medical image registration. A look at the extensions proposed to the original measure will also be provided. These were developed to improve the robustness of the measure and to avoid certain cases when maximizing MI does not lead to the correct spatial alignment. Clinton Fookes, Mohammed Bennamoun |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2000 | Global 3D Rigid Registration of Medical ImagesabstractWe present in this paper an iterative algorithm for the simultaneous registration of multiple 3D medical images. The proposed algorithm is a point-based registration method and is based on global registration techniques rather than the traditional pair-wise registration methods. Corresponding feature points, known as extremal points, are first automatically extracted from the 3D images and are used as the matching features in the registration process. These extremal points are stable landmarks, as the relative positions of these points are known to be invariant according to 3D rigid transformations. The registration algorithm is based on a novel weighted least squares formulation and it also incorporates 3D noise models on the extracted feature points. Results are presented for the 3D rigid registration of three successive MR images of the same patient taken at different periods of time. Clinton Fookes, John A. Williams 0001, Mohammed Bennamoun |
ICIP | 1 |