VLDB 2026 Research / reviewers in the wild / expert
Sridha Sridharan
dblp:80/166
· DBLP profile ↗
311ranked-venue papers
3as first author
48since 2021 · last 2026
0000-0003-4316-9001ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 211 · 3 first-author · 16 since 2021Artificial intelligence and machine learning · 170 · 29 since 2021Security and privacy · 13 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 8 since 2021Human-computer interaction and ubiquitous computing · 10 · 3 since 2021Databases, data management, data science and information retrieval · 7Systems, architecture and hardware · 5 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Ilov3Splat: Instance-Level Open-Vocabulary 3D Scene Understanding in Gaussian Splatting
Binh Long Nguyen, Kien Nguyen Thanh, Sridha Sridharan, Clinton Fookes, Peyman Moghadam |
ICPR (5) | 3 |
| 2026 | Multimodal transformer-diffusion framework for large-scale reconstruction of soccer tracking data
Harry Hughes, Patrick Lucey, Michael Horton 0001, Harshala Gammulle, Clinton Fookes, Sridha Sridharan |
Comput. Vis. Image Underst. | 6 |
| 2026 | Contrastive context distillation for skeleton based early action predictionabstractEarly action prediction (EAP) requires the inference of actions from partially observed sequences, which typically contain only the initial movements of an action. Compared to using complete sequences that show an action being fully executed, EAP lacks discriminative information and increased ambiguity as different actions can contain very similar initial movements. To alleviate this problem, recent methods try to distill features learned from action recognition (AR) models, which leads to sub-optimal results due to the lack of generalizability of features across AR and EAP tasks. In a different line of work, recent studies have proposed learning discriminative class-specific features, leveraging the traditional contrastive learning approach, where samples are selected and calibrated for learning discriminative features from hard-to-classify samples. This paper proposes a novel, dynamic, and context-aware framework for EAP by combining the merits of both knowledge distillation and contrastive learning. Particularly our method distills salient discriminative context from the complete action sequence to drive the EAP, with the help of a novel dynamic contrastive learning scheme. Extensive evaluations over three public datasets demonstrate state-of-the-art performance for Early Action Prediction. Chinthaka Ranasingha, Tharindu Fernando, Harshala Gammulle, Simon Denman, Sridha Sridharan, Clinton Fookes |
Expert Syst. Appl. | 5 |
| 2026 | Enhancing predictive performance on long-tail trajectories via clustering and specialized decodersabstractAccurate forecasting of traffic participants’ future trajectories is crucial for the advancement of autonomous driving systems. Constructing robust models for such tasks requires access to comprehensive datasets that include a variety of diverse cases. Current naturalistic trajectory prediction datasets are often imbalanced, featuring a large number of easier examples and a deficiency of more challenging instances. This long-tail distribution poses a significant challenge, resulting in inadequate model performance on the rare, yet safety-critical, parts of the data. To address this issue, we have proposed a framework that utilizes an embedding-based clustering technique and a distribution-sensitive decoder module to generate precise predictions for tail samples. In addition, the proposed framework includes a trajectory clustering module to refine predictions and improve the model’s capacity to generate multiple plausible future trajectories. Experimental results show that our framework outperforms the state-of-the-art long tail prediction method on tail samples by 19.5% on the Average Displacement error (ADE) and 25.5% on the Final Displacement error (FDE). Additionally, our approach attains state-of-the-art performance in terms of ADE metric on the ETH/UCY datasets, while only slightly trailing Y-Net in terms of the FDE metric. We further conduct ablation studies to highlight the efficacy of each of the proposed innovations. Source codes are available at his GitHub repository G. Ganeshaaraj, Tharindu Fernando, Sridha Sridharan, Clinton Fookes |
Pattern Recognit. | 3 |
| 2025 | Event2Tracking: Reconstructing Multi-Agent Soccer Trajectories Using Long-Term Multimodal ContextabstractSoccer is a rich testbed for studying multi-agent adversarial systems. In this work we focus on the task of reconstructing the noisy trajectories of soccer agents (players and the ball). Previous works that model the behaviours of agents in soccer are limited in two respects: (i) they only focus on short-term context windows (less than or equal to 10 seconds) which are not suitable for reconstructing trajectories impacted by long-term noise, and (ii) they exclusively rely on trajectory context, and do not leverage soccer's auxiliary data streams that can provide additional context. Our Event2Tracking model addresses these limitations. First, our architecture models soccer's long-term structure by processing long-term trajectories (60 seconds in duration). Secondly, our architecture is multimodal. Specifically, it fuses soccer tracking data with event data (which specifies the high-level semantic events that transpire in a game), providing rich context that cannot strictly be inferred from the raw trajectories. We evaluate our method empirically using a reconstruction loss metric. Compared to state-of-the-art approaches, our method substantially improves the accuracy of the ball's and players' reconstructed trajectories. Harry Hughes, Michael Horton 0001, Xinyu Wei 0004, Harshala Gammulle, Clinton Fookes, Sridha Sridharan, Patrick Lucey |
AAAI | 6 |
| 2025 | AG-VPReID: A Challenging Large-Scale Benchmark for Aerial-Ground Video-based Person Re-IdentificationabstractWe introduce AG-VPReID, a new large-scale dataset for aerial-ground video-based person re-identification (ReID) that comprises 6,632 subjects, 32,321 tracklets and over 9.6 million frames captured by drones (altitudes ranging from 15–120m), CCTV, and wearable cameras. This dataset offers a real-world benchmark for evaluating the robustness to significant viewpoint changes, scale variations, and resolution differences in cross-platform aerial-ground settings. In addition, to address these challenges, we propose AG-VPReID-Net, an end-to-end framework composed of three complementary streams: (1) an Adapted Temporal-Spatial Stream addressing motion pattern inconsistencies and facilitating temporal feature learning, (2) a Normalized Appearance Stream leveraging physics-informed techniques to tackle resolution and appearance changes, and (3) a Multi-Scale Attention Stream handling scale variations across drone altitudes. We integrate visual-semantic cues from all streams to form a robust, viewpoint-invariant whole-body representation. Extensive experiments demonstrate that AG-VPReID-Net outperforms state-of-the-art approaches on both our new dataset and existing video-based ReID benchmarks, showcasing its effectiveness and generalizability. Nevertheless, the performance gap observed on AG-VPReID across all methods underscores the dataset’s challenging nature. The dataset, code and trained models are available at AG-VPReID-Net. Kien Nguyen Thanh, Akila Pemasiri, Feng Liu 0037, Sridha Sridharan, Clinton Fookes |
CVPR | 5 |
| 2025 | AG-VPReID 2025: Aerial-Ground Video-based Person Re-identification Challenge ResultsabstractPerson re-identification (ReID) across aerial and ground vantage points has become crucial for large-scale surveillance and public safety applications. Although significant progress has been made in ground-only scenarios, bridging the aerial-ground domain gap remains a formidable challenge due to extreme viewpoint differences, scale variations, and occlusions. Building upon the achievements of the AG-ReID 2023 Challenge, this paper introduces the AG-VPReID 2025 Challenge—the first large-scale video-based competition focused on high-altitude (80–120 m) aerial-ground person ReID. Constructed on the new AG-VPReID dataset with 3,027 identities, over 13,500 tracklets, and approximately 3.7 million frames captured from UAVs, CCTV, and wearable cameras, the challenge featured four international teams. These teams developed solutions ranging from multi-stream architectures to transformer-based temporal reasoning and physics-informed modeling. The leading approach, X-TFCLIP from UAM, attained 72.28% Rank-1 accuracy in the aerial-to-ground ReID setting and 70.77% in the ground-to-aerial ReID setting, surpassing existing baselines while highlighting the dataset’s complexity. For additional details, please refer to the official website at https://agvpreid25.github.io. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Feng Liu 0037, Xiaoming Liu 0002, Arun Ross, Tamás Endrei, Ivan DeAndres-Tame, Ruben Tolosana, Rubén Vera-Rodríguez, Aythami Morales, Julian Fierrez, Javier Ortega-Garcia, Zijing Gong, Xuehu Liu, Md. Rashidunnabi, Hugo Proença 0001, Kailash A. Hambarde, Saeid Rezaei |
IJCB | 3 |
| 2025 | AG-VPReID.VIR: Bridging Aerial and Ground Platforms for Video-based Visible-Infrared Person Re-IDabstractPerson re-identification (Re-ID) across visible and infrared modalities is crucial for 24-hour surveillance systems, but existing datasets primarily focus on ground-level perspectives. While ground-based IR systems offer nighttime capabilities, they suffer from occlusions, limited coverage, and vulnerability to obstructions—problems that aerial perspectives uniquely solve. To address these limitations, we introduce AG-VPReID.VIR, the first aerial-ground cross-modality video-based person Re-ID dataset. This dataset captures 1,837 identities across 4,861 tracklets (124,855 frames) using both UAV-mounted and fixed CCTV cameras in RGB and infrared modalities. AG-VPReID.VIR presents unique challenges including cross-viewpoint variations, modality discrepancies, and temporal dynamics. Additionally, we propose TCC-VPReID, a novel three-stream architecture designed to address the joint challenges of cross-platform and cross-modality person Re-ID. Our approach bridges the domain gaps between aerial-ground perspectives and RGB-IR modalities, through style-robust feature learning, memory-based cross-view adaptation, and intermediary-guided temporal modeling. Experiments show that AG-VPReID.VIR presents distinctive challenges compared to existing datasets, with our TCC-VPReID framework achieving significant performance gains across multiple evaluation protocols. Dataset and code are available at https://github.com/agvpreid25/AG-VPReID.VIR. Kien Nguyen Thanh, Akila Pemasiri, Akmal Jahan, Clinton Fookes, Sridha Sridharan |
IJCB | 6 |
| 2025 | Beyond geometry: The power of texture in interpretable 3D person ReID
Kien Nguyen Thanh, Akila Pemasiri, Sridha Sridharan, Clinton Fookes |
Comput. Vis. Image Underst. | 4 |
| 2025 | Remembering What is Important: A Factorised Multi-Head Retrieval and Auxiliary Memory Stabilisation Scheme for Human Motion PredictionabstractHumans exhibit complex motions that vary depending on the activity they are performing, the interactions they engage in, as well as subject-specific preferences. Therefore, forecasting a human's future pose based on the history of his or her previous motion is a challenging task. This paper presents an innovative auxiliary-memory-powered deep neural network framework to improve the modelling of historical knowledge. Specifically, we disentangle subject-specific, action-specific, and other auxiliary information from the observed pose sequences and utilise these factorised features to query the memory. A novel Multi-Head knowledge retrieval scheme leverages these factorised feature embeddings to perform multiple querying operations over the historical observations captured within the auxiliary memory. Moreover, we propose a dynamic masking strategy to make this feature disentanglement process adaptive. Two novel loss functions are introduced to encourage diversity within the auxiliary memory, while ensuring the stability of the memory content such that it can locate and store salient information that aids the long-term prediction of future motion, irrespective of any data imbalances or the diversity of the input data distribution. Extensive experiments conducted on two public benchmarks, Human3.6M and CMU-Mocap, demonstrate that these design choices collectively allow the proposed approach to outperform the current state-of-the-art methods by significant margins: 17% on the Human3.6M dataset and 9% on the CMU-Mocap dataset. Tharindu Fernando, Harshala Gammulle, Sridha Sridharan, Simon Denman, Clinton Fookes |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Decoupled and Explainable Associative Memory for Effective Knowledge PropagationabstractLong-term memory often plays a pivotal role in human cognition through the analysis of contextual information. Machine learning researchers have attempted to emulate this process through the development of memory-augmented neural networks (MANNs) to leverage indirectly related but resourceful historical observations during learning and inference. The area of MANN, however, is still in its infancy and significant research effort is required to enable machines to achieve performance close to the human cognition process. This article presents an innovative MANN framework for the advanced incorporation of historical knowledge into a predictive framework. Within the key-value memory structure, we propose to decouple the key representations from the learned value memory embeddings to offer improved associations between the inputs and latent memory embeddings. We argue that the keys should be static, sparse, and unique representations of a particular observation to offer robust input to memory associations, while the value embeddings could be trainable, dense latent vectors such that they can better capture historical knowledge. Moreover, we introduce a novel memory update procedure that preserves the explainability of the historical knowledge extraction process, which would enable the human end-users to interpret the deep machine learning model decisions, fostering their trust. With extensive experiments conducted on three different datasets using audio, text, and image modalities, we demonstrate that our proposed innovations collectively allow this framework to outperform the current state-of-the-art methods by significant margins, irrespective of the modalities or the downstream tasks. The code is available at https://github.com/tha725/DE-KVMN/tree/main. Tharindu Fernando, Darshana Priyasad, Sridha Sridharan, Clinton Fookes |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Deep cross-domain transfer for emotion recognition via joint learningabstractAbstract Deep learning has been applied to achieve significant progress in emotion recognition from multimedia data. Despite such substantial progress, existing approaches are hindered by insufficient training data, leading to weak generalisation under mismatched conditions. To address these challenges, we propose a learning strategy which jointly transfers emotional knowledge learnt from rich datasets to source-poor datasets. Our method is also able to learn cross-domain features, leading to improved recognition performance. To demonstrate the robustness of the proposed learning strategy, we conducted extensive experiments on several benchmark datasets including eNTERFACE, SAVEE, EMODB, and RAVDESS. Experimental results show that the proposed method surpassed existing transfer learning schemes by a significant margin. Dung Nguyen 0001, Duc Thanh Nguyen, Sridha Sridharan, Mohamed Almorsy, Simon Denman, Son N. Tran, Clinton Fookes |
Multim. Tools Appl. | 3 |
| 2024 | FactoFormer: Factorized Hyperspectral Transformers With Self-Supervised PretrainingabstractHyperspectral images (HSIs) contain rich spectral and spatial information. Motivated by the success of transformers in the field of natural language processing and computer vision where they have shown the ability to learn long-range dependencies within input data, recent research has focused on using transformers for HSIs. However, current state-of-the-art hyperspectral transformers only tokenize the input HSI sample along the spectral dimension, resulting in the underutilization of spatial information. Moreover, transformers are known to be data-hungry and their performance relies heavily on large-scale pretraining, which is challenging due to limited annotated hyperspectral data. Therefore, the full potential of HSI transformers has not been fully realized. To overcome these limitations, we propose a novel factorized spectral–spatial transformer that incorporates factorized self-supervised pretraining procedures, leading to significant improvements in performance. The factorization of the inputs allows the spectral and spatial transformers to better capture the interactions within the hyperspectral data cubes. Inspired by masked image modeling (MIM) pretraining, we also devise efficient masking strategies for pretraining each of the spectral and spatial transformers. We conduct experiments on six publicly available datasets for the HSI classification task and demonstrate that our model achieves state-of-the-art performance in all the datasets. The code for our model will be made available athttps://github.com/csiro-robotics/FactoFormer. Shaheer Mohamed, Maryam Haghighat, Tharindu Fernando, Sridha Sridharan, Clinton Fookes, Peyman Moghadam |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | AG-ReID.v2: Bridging Aerial and Ground Views for Person Re-IdentificationabstractAerial-ground person re-identification (Re-ID) presents unique challenges in computer vision, stemming from the distinct differences in viewpoints, poses, and resolutions between high-altitude aerial and ground-based cameras. Existing research predominantly focuses on ground-to-ground matching, with aerial matching less explored due to a dearth of comprehensive datasets. To address this, we introduce AG-ReID.v2, a dataset specifically designed for person Re-ID in mixed aerial and ground scenarios. This dataset comprises 100,502 images of 1,615 unique individuals, each annotated with matching IDs and 15 soft attribute labels. Data were collected from diverse perspectives using a UAV, stationary CCTV, and smart glasses-integrated camera, providing a rich variety of intra-identity variations. Additionally, we have developed an explainable attention network tailored for this dataset. This network features a three-stream architecture that efficiently processes pairwise image distances, emphasizes key top-down features, and adapts to variations in appearance due to altitude differences. Comparative evaluations demonstrate the superiority of our approach over existing baselines. We plan to release the dataset and algorithm source code publicly, aiming to advance research in this specialized field of computer vision. Kien Nguyen Thanh, Sridha Sridharan, Clinton Fookes |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | Physical Adversarial Attacks for Surveillance: A SurveyabstractModern automated surveillance techniques are heavily reliant on deep learning methods. Despite the superior performance, these learning systems are inherently vulnerable to adversarial attacks-maliciously crafted inputs that are designed to mislead, or trick, models into making incorrect predictions. An adversary can physically change their appearance by wearing adversarial t-shirts, glasses, or hats or by specific behavior, to potentially avoid various forms of detection, tracking, and recognition of surveillance systems; and obtain unauthorized access to secure properties and assets. This poses a severe threat to the security and safety of modern surveillance systems. This article reviews recent attempts and findings in learning and designing physical adversarial attacks for surveillance applications. In particular, we propose a framework to analyze physical adversarial attacks and provide a comprehensive survey of physical adversarial attacks on four key surveillance tasks: detection, identification, tracking, and action recognition under this framework. Furthermore, we review and analyze strategies to defend against physical adversarial attacks and the methods for evaluating the strengths of the defense. The insights in this article present an important step in building resilience within surveillance systems to physical adversarial attacks. Kien Nguyen Thanh, Tharindu Fernando, Clinton Fookes, Sridha Sridharan |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | AG-ReID 2023: Aerial-Ground Person Re-identification Challenge ResultsabstractPerson re-identification (Re-ID) on aerial-ground platforms has emerged as an intriguing topic within computer vision, presenting a plethora of unique challenges. Highflying altitudes of aerial cameras make persons appear differently in terms of viewpoints, poses, and resolution compared to the images of the same person viewed from ground cameras. Despite its potential, few algorithms have been developed for person re-identification on aerial-ground data, mainly due to the absence of comprehensive datasets. In response, we have collected a large-scale dataset and organized the Aerial-Ground person Re-IDentification Challenge (AG-ReID2023) to foster advancements in the field. The dataset comprises 100,502 images with 1,615 unique identities, including 51,530 training images featuring 807 identities. The test set is divided into two subsets: Aerial to Ground (808 ids, 4,348 query images, 19,259 gallery images) and Ground to Aerial (808 ids, 4,151 query images, 21,214 gallery images). In addition, we manually annotate individuals with their matching IDs across cameras and provide 15 soft attribute labels. The AG-ReID2023 Challenge in conjunction with the 7thIEEE International Joint Conference on Biometrics (IJCB) has garnered interest from numerous institutes, resulting in the submission of five distinct algorithms. We provide an in-depth examination of the evaluation outcomes and present our findings from the contest. For additional details, kindly refer to the official website1.1https://agreid23.github.io. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Feng Liu 0037, Xiaoming Liu 0002, Arun Ross, Dana Michalski, Debayan Deb, Mahak Kothari, Manisha Saini, Dawei Du, Scott McCloskey, Gabriel Bertocco, Fernanda A. Andaló, Terrance E. Boult, Anderson Rocha 0001, Haidong Zhu, Zhaoheng Zheng, Ramakant Nevatia, Zaigham A. Randhawa, Sinan Sabri, Gianfranco Doretto |
IJCB | 3 |
| 2023 | Aerial-Ground Person Re-IDabstractPerson re-ID matches persons across multiple non-overlapping cameras. Despite the increasing deployment of air-borne platforms in surveillance, current existing person re-ID benchmarks’ focus is on ground-ground matching and very limited efforts on aerial-aerial matching. We propose a new benchmark dataset - AG-ReID, which performs person re-ID matching in a new setting: across aerial and ground cameras. Our dataset contains 21,983 images of 388 identities and 15 soft attributes for each identity. The data was collected by a UAV flying at altitudes between 15 to 45 meters and a ground-based CCTV camera on a university campus. Our dataset presents a novel elevated-viewpoint challenge for person re-ID due to the significant difference in person appearance across these cameras. We propose an explainable algorithm to guide the person re-ID model’s training with soft attributes to address this challenge. Experiments demonstrate the efficacy of our method on the aerial-ground person re-ID task. The dataset will be published and the baseline codes will be open-sourced at https://github.com/huynguyen792/AG-ReID to facilitate research in this area. Kien Nguyen Thanh, Sridha Sridharan, Clinton Fookes |
ICME | 3 |
| 2023 | Wild-Places: A Large-Scale Dataset for Lidar Place Recognition in Unstructured Natural EnvironmentsabstractMany existing datasets for lidar place recognition are solely representative of structured urban environments, and have recently been saturated in performance by deep learning based approaches. Natural and unstructured environments present many additional challenges for the tasks of long-term localisation but these environments are not represented in currently available datasets. To address this we introduce Wild-Places, a challenging large-scale dataset for lidar place recognition in unstructured, natural environments. Wild-Places contains eight lidar sequences collected with a handheld sensor payload over the course of fourteen months, containing a total of 63K undistorted lidar submaps along with accurate 6DoF ground truth. This dataset contains multi-ple revisits both within and between sequences, allowing for both intra-sequence (i.e., loop closure detection) and inter-sequence (i.e., re-localisation) tasks. We also benchmark several state-of-the-art approaches to demonstrate the challenges that this dataset introduces, particularly the case of long-term place recognition due to natural environments changing over time. Our dataset and code is available at https://csiro-robotics.github.io/Wild-Places Joshua Knights, Kavisha Vidanapathirana, Milad Ramezani, Sridha Sridharan, Clinton Fookes, Peyman Moghadam |
ICRA | 4 |
| 2023 | Dual Memory Fusion for Multimodal Speech Emotion Recognition
Darshana Prisayad, Tharindu Fernando, Sridha Sridharan, Simon Denman, Clinton Fookes |
INTERSPEECH | 3 |
| 2023 | Meta-transfer learning for emotion recognitionabstractAbstract Deep learning has been widely adopted in automatic emotion recognition and has lead to significant progress in the field. However, due to insufficient training data, pre-trained models are limited in their generalisation ability, leading to poor performance on novel test sets. To mitigate this challenge, transfer learning performed by fine-tuning pr-etrained models on novel domains has been applied. However, the fine-tuned knowledge may overwrite and/or discard important knowledge learnt in pre-trained models. In this paper, we address this issue by proposing a PathNet-based meta-transfer learning method that is able to (i) transfer emotional knowledge learnt from one visual/audio emotion domain to another domain and (ii) transfer emotional knowledge learnt from multiple audio emotion domains to one another to improve overall emotion recognition accuracy. To show the robustness of our proposed method, extensive experiments on facial expression-based emotion recognition and speech emotion recognition are carried out on three bench-marking data sets: SAVEE, EMODB, and eNTERFACE. Experimental results show that our proposed method achieves superior performance compared with existing transfer learning methods. Dung Nguyen 0001, Duc Thanh Nguyen, Sridha Sridharan, Simon Denman, Thanh Thi Nguyen 0001, David Dean, Clinton Fookes |
Neural Comput. Appl. | 3 |
| 2023 | Complex-Valued Iris Recognition NetworkabstractIn this work, we design a fully complex-valued neural network for the task of iris recognition. Unlike the problem of general object recognition, where real-valued neural networks can be used to extract pertinent features, iris recognition depends on the extraction of both phase and magnitude information from the input iris texture in order to better represent its biometric content. This necessitates the extraction and processing of phase information that cannot be effectively handled by a real-valued neural network. In this regard, we design a fully complex-valued neural network that can better capture the multi-scale, multi-resolution, and multi-orientation phase and amplitude features of the iris texture. We show a strong correspondence of the proposed complex-valued iris recognition network with Gabor wavelets that are used to generate the classical IrisCode; however, the proposed method enables a new capability of automatic complex-valued feature learning that is tailored for iris recognition. We conduct experiments on three benchmark datasets - ND-CrossSensor-2013, CASIA-Iris-Thousand and UBIRIS.v2 - and show the benefit of the proposed network for the task of iris recognition. We exploit visualization schemes to convey how the complex-valued network, when compared to standard real-valued networks, extracts fundamentally different features from the iris texture. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Arun Ross |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Multi-stage stacked temporal convolution neural networks (MS-S-TCNs) for biosignal segmentation and anomaly localizationabstractIn the computer vision domain, temporal convolution networks (TCN) have gained traction due to their lightweight, robust architectures for sequence-to-sequence prediction tasks. With that insight, in this study, we propose a novel deep learning architecture for biosignal segmentation and anomaly localization based on TCNs , named the multi-stage stacked TCN, which employs multiple TCN modules with varying dilation factors. More precisely, for each stage, our architecture uses TCN modules with multiple dilation factors, and we use convolution-based fusion to combine predictions returned from each stage. Furthermore, aiming smoothed predictions, we introduce a novel loss function based on the first-order derivative. To demonstrate the robustness of our architecture, we evaluate our model on five different tasks related to three 1D biosignal modalities (heart sounds, lung sounds and electrocardiogram). Our proposed framework achieves state-of-the-art performance for all tasks, significantly outperforming the respective state-of-the-art models having F1 score gains up to ≈ 9 %. Furthermore, the framework demonstrates competitive performance gains compared to traditional multi-stage TCN models with similar configurations yielding F1 score gains up to ≈ 5 %. Our model is also interpretable. Using neural conductance, we demonstrate the effectiveness of having TCNs with varying dilation factors. Our visualizations show that the model benefits from feature maps captured at multiple dilation factors, and the information is effectively propagated through the network such that the final stage produces the most accurate result. Theekshana Dissanayake, Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
Pattern Recognit. | 4 |
| 2023 | Pose-driven attention-guided image generation for person re-Identification
Amena Khatun, Simon Denman, Sridha Sridharan, Clinton Fookes |
Pattern Recognit. | 3 |
| 2023 | Toward On-Board Panoptic Segmentation of Multispectral Satellite ImagesabstractWith tremendous advancements in low-power embedded computing devices and remote sensing instruments, the traditional satellite image processing pipeline which includes an expensive data transfer step prior to processing data on the ground is being replaced by on- board processing of captured data. This paradigm shift enables critical and time-sensitive intelligence to be acquired in a timely manner on- board the satellite itself. However, at present, the on- board processing of multispectral satellite images is limited to classification and segmentation tasks. Extending this processing to the next logical level, we take the first step toward on- board panoptic segmentation of multispectral satellite images and evaluate the applicability of state-of-the-art panoptic segmentation models to an on- board setting. Panoptic segmentation offers major economic and environmental insights, ranging from yield estimation from agricultural lands to intelligence for complex military applications. Nevertheless, the on- board intelligence extraction poses several challenges due to the loss of temporal observations and the need to generate predictions from a single sample. To address this challenge, we propose a multimodal teacher network with a cross modality attention-based fusion strategy to improve segmentation accuracy by exploiting data from multiple modes. We also propose an online knowledge distillation framework to transfer the knowledge learned by this multimodal teacher network to a unimodal student, which receives only a single frame input, and is more appropriate for an on- board environment. We benchmark our approach against existing state-of-the-art panoptic segmentation models using the PASTIS multispectral panoptic segmentation dataset considering an on- board processing setting. Our evaluations demonstrate a substantial 10.7%, 11.9%, and 10.6% increase in segmentation quality (SQ), recognition quality (RQ), and panoptic quality (PQ) metrics compared to the existing state-of-the-art model when it is evaluated in an on- board processing setting. Tharindu Fernando, Clinton Fookes, Harshala Gammulle, Simon Denman, Sridha Sridharan |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Generalized Generative Deep Learning Models for Biosignal Synthesis and Modality TransferabstractGenerative Adversarial Networks (GANs) are a revolutionary innovation in machine learning that enables the generation of artificial data. Artificial data synthesis is valuable especially in the medical field where it is difficult to collect and annotate real data due to privacy issues, limited access to experts, and cost. While adversarial training has led to significant breakthroughs in the computer vision field, biomedical research has not yet fully exploited the capabilities of generative models for data generation, and for more complex tasks such as biosignal modality transfer. We present a broad analysis on adversarial learning on biosignal data. Our study is the first in the machine learning community to focus on synthesizing 1D biosignal data using adversarial models. We consider three types of deep generative adversarial networks: a classical GAN, an adversarial AE, and a modality transfer GAN; individually designed for biosignal synthesis and modality transfer purposes. We evaluate these methods on multiple datasets for different biosignal modalites, including phonocardiogram (PCG), electrocardiogram (ECG), vectorcardiogram and 12-lead electrocardiogram. We follow subject-independent evaluation protocols, by evaluating the proposed models' performance on completely unseen data to demonstrate generalizability. We achieve superior results in generating biosignals, specifically in conditional generation, by synthesizing realistic samples while preserving domain-relevant characteristics. We also demonstrate insightful results in biosignal modality transfer that can generate expanded representations from fewer input-leads, ultimately making the clinical monitoring setting more convenient for the patient. Furthermore our longer duration ECGs generated, maintain clear ECG rhythmic regions, which has been proven using ad-hoc segmentation models. Theekshana Dissanayake, Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | SESS: Saliency Enhancing with Scaling and Sliding
Osman Tursun, Simon Denman, Sridha Sridharan, Clinton Fookes |
ECCV (12) | 3 |
| 2022 | LoGG3D-Net: Locally Guided Global Descriptor Learning for 3D Place RecognitionabstractRetrieval-based place recognition is an efficient and effective solution for re-localization within a pre-built map, or global data association for Simultaneous Localization and Mapping (SLAM). The accuracy of such an approach is heavily dependant on the quality of the extracted scene-level representation. While end-to-end solutions - which learn a global descriptor from input point clouds - have demonstrated promising results, such approaches are limited in their ability to enforce desirable properties at the local feature level. In this paper, we introduce a local consistency loss to guide the network towards learning local features which are consistent across revisits, hence leading to more repeatable global descriptors resulting in an overall improvement in 3D place recognition performance. We formulate our approach in an end-to-end trainable architecture called LoGG3D-Net. Experiments on two large-scale public benchmarks (KITTI and MulRan) show that our method achieves mean F1maxscores of 0.939 and 0.968 on KITTI and MulRan respectively, achieving state-of-the-art performance while operating in near real-time. The open-source implementation is available at: https://github.com/csiro-robotics/LoGG3D-Net. Kavisha Vidanapathirana, Milad Ramezani, Peyman Moghadam, Sridha Sridharan, Clinton Fookes |
ICRA | 4 |
| 2022 | Detecting Heart Failure Through Voice Analysis using Self-Supervised Mode-Based Memory FusionabstractCongestive Heart Failure (CHF) is a progressive disease that affects millions of people worldwide, severely impacting their quality of life. Missed detection of CHF and its progression affects life expectancy, thus it is critical to develop applications to continuously monitor CHF symptoms and disease progression in a patient-centric and cost-effective manner. This paper focuses on a novel non-invasive technique to identify CHF using patients' speech traits. Pulmonary congestion and breathlessness is the most common symptom of heart failure and one of the major contributors to hospitalisation. Since pulmonary congestion results in impairment of a patient's voice, we propose a novel, non invasive method for monitoring CHF through analysis of the patient's speech. We also introduce a new balanced dataset, containing voice recordings from both healthy participants and participants diagnosed with CHF, which contains voice alterations reflective of CHF status. We propose a novel deep machine learning architecture based on mode driven memory fusion for CHF recognition from audio recordings of subject's speech. We have achieved 90% accuracy under a subject-independent evaluation setting, highlighting the applicability of such methods for tele-health and home monitoring applications. Darshana Priyasad, Andi Partovi, Sridha Sridharan, Maryam Kashefpoor, Tharindu Fernando, Simon Denman, Clinton Fookes, David Kaye |
INTERSPEECH | 3 |
| 2022 | InCloud: Incremental Learning for Point Cloud Place RecognitionabstractPlace recognition is a fundamental component of robotics, and has seen tremendous improvements through the use of deep learning models in recent years. Networks can experience significant drops in performance when deployed in unseen or highly dynamic environments, and require additional training on the collected data. However naively fine-tuning on new training distributions can cause severe degradation of performance on previously visited domains, a phenomenon known as catastrophic forgetting. In this paper we address the problem of incremental learning for point cloud place recognition and introduce InCloud, a structure-aware distillation-based approach which preserves the higher-order structure of the network's embedding space. We introduce several challenging new benchmarks on four popular and large-scale LiDAR datasets (Oxford, MulRan, In-house and KITTI) showing broad improvements in point cloud place recognition performance over a variety of network architectures. To the best of our knowledge, this work is the first to effectively apply incremental learning for point cloud place recognition. Data pre-processing, training and evaluation code for this paper can be found at https://github.com/csiro-robotics/InCloud. Joshua Knights, Peyman Moghadam, Milad Ramezani, Sridha Sridharan, Clinton Fookes |
IROS | 4 |
| 2022 | Learning test-time augmentation for content-based image retrieval
Osman Tursun, Simon Denman, Sridha Sridharan, Clinton Fookes |
Comput. Vis. Image Underst. | 3 |
| 2022 | Affect recognition from scalp-EEG using channel-wise encoder networks coupled with geometric deep learning and multi-channel feature fusion
Darshana Priyasad, Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
Knowl. Based Syst. | 4 |
| 2022 | Split 'n' merge net: A dynamic masking network for multi-task attention
Tharindu Fernando, Sridha Sridharan, Simon Denman, Clinton Fookes |
Pattern Recognit. | 2 |
| 2022 | An efficient framework for zero-shot sketch-based image retrieval
Osman Tursun, Simon Denman, Sridha Sridharan, Ethan Goan, Clinton Fookes |
Pattern Recognit. | 3 |
| 2022 | Channel Graph Regularized Correlation Filters for Visual Object TrackingabstractCorrelation Filters (CF) are a popular choice for visual object tracking due to their efficiency in the frequency domain. Convolutional and hand-crafted features are jointly used when learning a filter, however, these features are not uniformly important when tracking a target. Given this observation, spatial and temporal regularization and attention models have been investigated. However, these models do not consider the interaction between different feature channels. As a result, dissimilar weights are assigned to similar feature channels. To address this issue, we propose a channel attention model and study two different regularization methods for attention. We investigate the application of channel regularization to emphasize important feature channels; and graph regularization which increases the likelihood of similar feature channels obtaining similar weights. The proposed formulation can be efficiently solved via the alternating direction method of multipliers. We first show the advantages of using the proposed channel regularization by demonstrating its performance when applied to two existing CF trackers. This is followed by analyzing the effect of using the proposed channel-graph regularization for CF based tracking. The evaluation is performed on publicly available tracking datasets: OTB100, TC128, VOT-2017, VOT-2019, LaSOT, UAV123, and GOT-10k. Evaluation over multiple challenges and a comparative analysis with existing top-ranked trackers shows that our formulation improves the discriminative power of the learned CF, preventing tracker drift during challenging scenarios. Arjun Tyagi, A. Venkata Subramanyam, Simon Denman, Sridha Sridharan, Clinton Fookes |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Component-Based Attention for Large-Scale Trademark RetrievalabstractThe need for large-scale trademark retrieval (TR) systems has significantly increased to combat the rise in international trademark infringement. Unfortunately, the ranking accuracy of current approaches using either hand-crafted or pre-trained deep convolution neural network (DCNN) features is inadequate for large-scale deployments. We show in this paper that the ranking accuracy of TR systems can be significantly improved by incorporating hard and soft attention mechanisms, which direct attention to critical information such as figurative elements and reduce the attention given to distracting and uninformative elements such as text and background. Our proposed approach achieves state-of-the-art results on a challenging large-scale trademark dataset. Osman Tursun, Simon Denman, Sabesan Sivipalan, Sridha Sridharan, Clinton Fookes, Sandra Mau |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2022 | Geometric Deep Learning for Subject Independent Epileptic Seizure Prediction Using Scalp EEG SignalsabstractRecently, researchers in the biomedical community have introduced deep learning-based epileptic seizure prediction models using electroencephalograms (EEGs) that can anticipate an epileptic seizure by differentiating between the pre-ictal and interictal stages of the subject's brain. Despite having the appearance of a typical anomaly detection task, this problem is complicated by subject-specific characteristics in EEG data. Therefore, studies that investigate seizure prediction widely employ subject-specific models. However, this approach is not suitable in situations where a target subject has limited (or no) data for training. Subject-independent models can address this issue by learning to predict seizures from multiple subjects, and therefore are of greater value in practice. In this study, we propose a subject-independent seizure predictor using Geometric Deep Learning (GDL). In the first stage of our GDL-based method we use graphs derived from physical connections in the EEG grid. We subsequently seek to synthesize subject-specific graphs using deep learning. The models proposed in both stages achieve state-of-the-art performance using a one-hour early seizure prediction window on two benchmark datasets (CHB-MIT-EEG: 95.38% with 23 subjects and Siena-EEG: 96.05% with 15 subjects). To the best of our knowledge, this is the first study that proposes synthesizing subject-specific graphs for seizure prediction. Furthermore, through model interpretation we outline how this method can potentially contribute towards Scalp EEG-based seizure localization. Theekshana Dissanayake, Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | Robust and Interpretable Temporal Convolution Network for Event Detection in Lung Sound RecordingsabstractOBJECTIVE: This paper proposes a novel framework for lung sound event detection, segmenting continuous lung sound recordings into discrete events and performing recognition of each event. METHODS: We propose the use of a multi-branch TCN architecture and exploit a novel fusion strategy to combine the resultant features from these branches. This not only allows the network to retain the most salient information across different temporal granularities and disregards irrelevant information, but also allows our network to process recordings of arbitrary length. RESULTS: The proposed method is evaluated on multiple public and in-house benchmarks, containing irregular and noisy recordings of the respiratory auscultation process for the identification of auscultation events including inhalation, crackles, and rhonchi. Moreover, we provide an end-to-end model interpretation pipeline. CONCLUSION: Our analysis of different feature fusion strategies shows that the proposed feature concatenation method leads to better suppression of non-informative features, which drastically reduces the classifier overhead resulting in a robust lightweight network. SIGNIFICANCE: Lung sound event detection is a primary diagnostic step for numerous respiratory diseases. The proposed method provides a cost-effective and efficient alternative to exhaustive manual segmentation, and provides more accurate segmentation than existing methods. The end-to-end model interpretability helps to build the required trust in the system for use in clinical settings. Tharindu Fernando, Sridha Sridharan, Simon Denman, Houman Ghaemmaghami, Clinton Fookes |
IEEE J. Biomed. Health Informatics | 2 |
| 2022 | Deep Auto-Encoders With Sequential Learning for Multimodal Dimensional Emotion RecognitionabstractMultimodal dimensional emotion recognition has drawn a great attention from the affective computing community and numerous schemes have been extensively investigated, making a significant progress in this area. However, several questions still remain unanswered for most of existing approaches including: (i) how to simultaneously learn compact yet representative features from multimodal data, (ii) how to effectively capture complementary features from multimodal streams, and (iii) how to perform all the tasks in an end-to-end manner. To address these challenges, in this paper, we propose a novel deep neural network architecture consisting of a two-stream auto-encoder and a long short term memory for effectively integrating visual and audio signal streams for emotion recognition. To validate the robustness of our proposed architecture, we carry out extensive experiments on the multimodal emotion in the wild dataset: RECOLA. Experimental results show that the proposed method achieves state-of-the-art recognition performance. Dung Nguyen 0001, Duc Thanh Nguyen, Thanh Thi Nguyen 0001, Son N. Tran, Thin Nguyen, Sridha Sridharan, Clinton Fookes |
IEEE Trans. Multim. | 7 |
| 2022 | Elasticity Meets Continuous-Time: Map-Centric Dense 3D LiDAR SLAMabstractMap-centric SLAM utilizes elasticity as a means of loop closure. This approach reduces the cost of loop closure while still providing large-scale fusion-based dense maps, when compared to trajectory-centric SLAM approaches. In this article, we present a novel framework, namedElasticLiDAR++, for multimodal map-centric SLAM. Having the advantages of a map-centric approach, our method exhibits new features to overcome the shortcomings of existing systems associated with multimodal (LiDAR-inertial-visual) sensor fusion and LiDAR motion distortion. This is accomplished through the use of a local continuous-time trajectory representation. Also, our surface resolution preserving matching algorithm and normal-inverse-Wishart-based surfel fusion model enables nonredundant yet dense mapping. Furthermore, we present a robust metric loop closure model to make the approach stable regardless of where the loop closure occurs. Finally, we demonstrate our approach through both simulation and real data experiments using multiple sensor payload configurations and environments to illustrate its utility and robustness. Chanoh Park, Peyman Moghadam, Jason Williams 0002, Soohwan Kim, Sridha Sridharan, Clinton Fookes |
IEEE Trans. Robotics | 5 |
| 2021 | Learning Regional Attention Over Multi-Resolution Deep Convolutional Features For Trademark RetrievalabstractLarge-scale trademark retrieval is an important content-based image retrieval task. A recent study shows that off-the-shelf deep features aggregated with Regional-Maximum Activation of Convolutions (R-MAC) achieve state-of-the-art results. However, R-MAC suffers in the presence of background clutter/trivial regions and scale variance, and discards important spatial information. We introduce three simple but effective modifications to R-MAC to overcome these drawbacks. First, we propose the use of both sum and max pooling to minimise the loss of spatial information. We also employ domain-specific unsupervised soft-attention to eliminate background clutter and unimportant regions. Finally, we add multi-resolution inputs to enhance the scale-invariance of R-MAC. We evaluate these three modifications on the million-scale METU dataset. Our results show that all modifications bring non-trivial improvements, and surpass previous state-of-the-art results. Osman Tursun, Simon Denman, Sridha Sridharan, Clinton Fookes |
ICIP | 3 |
| 2021 | Locus: LiDAR-based Place Recognition using Spatiotemporal Higher-Order PoolingabstractPlace Recognition enables the estimation of a globally consistent map and trajectory by providing non-local constraints in Simultaneous Localisation and Mapping (SLAM). This paper presents Locus, a novel place recognition method using 3D LiDAR point clouds in large-scale environments. We propose a method for extracting and encoding topological and temporal information related to components in a scene and demonstrate how the inclusion of this auxiliary information in place description leads to more robust and discriminative scene representations. Second-order pooling along with a non- linear transform is used to aggregate these multi-level features to generate a fixed-length global descriptor, which is invariant to the permutation of input features. The proposed method outperforms state-of-the-art methods on the KITTI dataset. Furthermore, Locus is demonstrated to be robust across several challenging situations such as occlusions and viewpoint changes in 3D LiDAR point clouds. The open-source implementation is available at: https://github.com/csiro-robotics/locus. Kavisha Vidanapathirana, Peyman Moghadam, Ben Harwood, Muming Zhao, Sridha Sridharan, Clinton Fookes |
ICRA | 5 |
| 2021 | IGSSTRCF: Importance Guided Sparse Spatio-Temporal Regularized Correlation Filters For TrackingabstractThis paper proposes a novel Importance Guided Sparse Spatio-Temporal Regularization based Correlation Filter (IGSSTRCF) tracker. Our formulation explicitly models the variations in the correlation filters and associated spatial weights in successive frames. By imposing a sparsity penalty on these variations, the formulation ensures that only relevant changes are incorporated during updates. This results in more robust filter coefficients that minimize the tracking drift. The IGSSTRCF also includes an adaptive channel importance estimation strategy that assigns an importance weight to each feature channel during training. The proposed formulation is efficiently solved via the alternating direction method of multipliers. A comparative analysis is shown on TC128, UAV123, VOT-2017, and VOT-2019 datasets; and we present an ablation study to demonstrate the contribution of each component of the IGSSTRCF. It is observed that we outperform several state-of-the-art trackers and each component of the proposed IGSSTRCF contributes positively towards tracker performance. A. Venkata Subramanyam, Simon Denman, Sridha Sridharan, Clinton Fookes |
WACV | 4 |
| 2021 | Multi-modal semantic image segmentation
Akila Pemasiri, Kien Nguyen Thanh, Sridha Sridharan, Clinton Fookes |
Comput. Vis. Image Underst. | 3 |
| 2021 | Detection of Fake and Fraudulent Faces via Neural Memory NetworksabstractAdvances in computer vision have brought us to the point where we have the ability to synthesise realistic fake content. Such approaches are seen as a source of disinformation and mistrust, and pose serious concerns to governments around the world. Convolutional Neural Networks (CNNs) demonstrate encouraging results when detecting fake images that arise from the specific type of manipulation they are trained on. However, this success has not transitioned to unseen manipulation types, resulting in a significant gap in the line-of-defense. We propose a Hierarchical Attention Memory Network (HAMN), motivated by the social cognition processes of the human brain, for the detection of fake faces. Through visual cues and by utilising knowledge stored in neural memories, we allow the network to reason about the perceived face and anticipate it's future semantic embeddings. This renders a generalisable face tampering detection framework. Experimental results demonstrate the proposed approach achieves superior performance for fake and fraudulent face detection. Tharindu Fernando, Clinton Fookes, Simon Denman, Sridha Sridharan |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2021 | End-to-End Domain Adaptive Attention Network for Cross-Domain Person Re-IdentificationabstractPerson re-identification (re-ID) remains challenging in a real-world scenario, as it requires a trained network to generalise to totally unseen target data in the presence of variations across domains. Recently, generative adversarial models have been widely adopted to enhance the diversity of training data. These approaches, however, often fail to generalise to other domains, as existing generative person re-identification models have a disconnect between the generative component and the discriminative feature learning stage. To address the on-going challenges regarding model generalisation, we propose an end-to-end domain adaptive attention network to jointly translate images between domains and learn discriminative re-id features in a single framework. To address the domain gap challenge, we introduce an attention module for image translation from source to target domains without affecting the identity of a person. More specifically, attention is directed to the background instead of the entire image of the person, ensuring identifying characteristics of the subject are preserved. The proposed joint learning network results in a significant performance improvement over state-of-the-art methods on several challenging benchmark datasets. Amena Khatun, Simon Denman, Sridha Sridharan, Clinton Fookes |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2021 | TMMF: Temporal Multi-Modal Fusion for Single-Stage Continuous Gesture RecognitionabstractGesture recognition is a much studied research area which has myriad real-world applications including robotics and human-machine interaction. Current gesture recognition methods have focused on recognising isolated gestures, and existing continuous gesture recognition methods are limited to two-stage approaches where independent models are required for detection and classification, with the performance of the latter being constrained by detection performance. In contrast, we introduce a single-stage continuous gesture recognition framework, called Temporal Multi-Modal Fusion (TMMF), that can detect and classify multiple gestures in a video via a single model. This approach learns the natural transitions between gestures and non-gestures without the need for a pre-processing segmentation step to detect individual gestures. To achieve this, we introduce a multi-modal fusion mechanism to support the integration of important information that flows from multi-modal inputs, and is scalable to any number of modes. Additionally, we propose Unimodal Feature Mapping (UFM) and Multi-modal Feature Mapping (MFM) models to map uni-modal features and the fused multi-modal features respectively. To further enhance performance, we propose a mid-point based loss function that encourages smooth alignment between the ground truth and the prediction, helping the model to learn natural gesture transitions. We demonstrate the utility of our proposed framework, which can handle variable-length input videos, and outperforms the state-of-the-art on three challenging datasets: EgoGesture, IPN hand and ChaLearn LAP Continuous Gesture Dataset (ConGD). Furthermore, ablation experiments show the importance of different components of the proposed framework. Harshala Gammulle, Simon Denman, Sridha Sridharan, Clinton Fookes |
IEEE Trans. Image Process. | 3 |
| 2021 | Identification of Children at Risk of Schizophrenia via Deep Learning and EEG ResponsesabstractThe prospective identification of children likely to develop schizophrenia is a vital tool to support early interventions that can mitigate the risk of progression to clinical psychosis. Electroencephalographic (EEG) patterns from brain activity and deep learning techniques are valuable resources in achieving this identification. We propose automated techniques that can process raw EEG waveforms to identify children who may have an increased risk of schizophrenia compared to typically developing children. We also analyse abnormal features that remain during developmental follow-up over a period of ∼ 4 years in children with a vulnerability to schizophrenia initially assessed when aged 9 to 12 years. EEG data from participants were captured during the recording of a passive auditory oddball paradigm. We undertake a holistic study to identify brain abnormalities, first by exploring traditional machine learning algorithms using classification methods applied to hand-engineered features (event-related potential components). Then, we compare the performance of these methods with end-to-end deep learning techniques applied to raw data. We demonstrate via average cross-validation performance measures that recurrent deep convolutional neural networks can outperform traditional machine learning methods for sequence modeling. We illustrate the intuitive salient information of the model with the location of the most relevant attributes of a post-stimulus window. This baseline identification system in the area of mental illness supports the evidence of developmental and disease effects in a pre-prodromal phase of psychosis. These results reinforce the benefits of deep learning to support psychiatric classification and neuroscientific research more broadly. David Ahmedt-Aristizabal, Tharindu Fernando, Simon Denman, Jonathan E. Robinson, Sridha Sridharan, Patrick J. Johnston, Kristin R. Laurens, Clinton Fookes |
IEEE J. Biomed. Health Informatics | 5 |
| 2021 | A Robust Interpretable Deep Learning Classifier for Heart Anomaly Detection Without SegmentationabstractTraditionally, abnormal heart sound classification is framed as a three-stage process. The first stage involves segmenting the phonocardiogram to detect fundamental heart sounds; after which features are extracted and classification is performed. Some researchers in the field argue the segmentation step is an unwanted computational burden, whereas others embrace it as a prior step to feature extraction. When comparing accuracies achieved by studies that have segmented heart sounds before analysis with those who have overlooked that step, the question of whether to segment heart sounds before feature extraction is still open. In this study, we explicitly examine the importance of heart sound segmentation as a prior step for heart sound classification, and then seek to apply the obtained insights to propose a robust classifier for abnormal heart sound detection. Furthermore, recognizing the pressing need for explainable Artificial Intelligence (AI) models in the medical domain, we also unveil hidden representations learned by the classifier using model interpretation techniques. Experimental results demonstrate that the segmentation which can be learned by the model plays an essential role in abnormal heart sound classification. Our new classifier is also shown to be robust, stable and most importantly, explainable, with an accuracy of almost 100% on the widely used PhysioNet dataset. Theekshana Dissanayake, Tharindu Fernando, Simon Denman, Sridha Sridharan, Houman Ghaemmaghami, Clinton Fookes |
IEEE J. Biomed. Health Informatics | 4 |
| 2020 | Geometry-Constrained Car Recognition Using a 3D Perspective NetworkabstractWe present a novel learning framework for vehicle recognition from a single RGB image. Unlike existing methods which only use attention mechanisms to locate 2D discriminative information, our work learns a novel 3D perspective feature representation of a vehicle, which is then fused with 2D appearance feature to predict the category. The framework is composed of a global network (GN), a 3D perspective network (3DPN), and a fusion network. The GN is used to locate the region of interest (RoI) and generate the 2D global feature. With the assistance of the RoI, the 3DPN estimates the 3D bounding box under the guidance of the proposed vanishing point loss, which provides a perspective geometry constraint. Then the proposed 3D representation is generated by eliminating the viewpoint variance of the 3D bounding box using perspective transformation. Finally, the 3D and 2D feature are fused to predict the category of the vehicle. We present qualitative and quantitative results on the vehicle classification and verification tasks in the BoxCars dataset. The results demonstrate that, by learning such a concise 3D representation, we can achieve superior performance to methods that only use 2D information while retain 3D meaningful information without the challenge of requiring a 3D CAD model. ZongYuan Ge, Simon Denman, Sridha Sridharan, Clinton Fookes |
AAAI | 4 |
| 2020 | Attention Driven Fusion for Multi-Modal Emotion RecognitionabstractDeep learning has emerged as a powerful alternative to hand-crafted methods for emotion recognition on combined acoustic and text modalities. Baseline systems model emotion information in text and acoustic modes independently using Deep Convolutional Neural Networks (DCNN) and Recurrent Neural Networks (RNN), followed by applying attention, fusion, and classification. In this paper, we present a deep learning-based approach to exploit and fuse text and acoustic data for emotion classification. We utilize a SincNet layer, based on parameterized sinc functions with band-pass filters, to extract acoustic features from raw audio followed by a DCNN. This approach learns filter banks tuned for emotion recognition and provides more effective features compared to directly applying convolutions over the raw speech signal. For text processing, we use two branches (a DCNN and a Bi-direction RNN followed by a DCNN) in parallel where cross attention is introduced to infer the N-gram level correlations on hidden representations received from the Bi-RNN. Following existing state-of-the-art, we evaluate the performance of the proposed system on the IEMOCAP dataset. Experimental results indicate that the proposed system outperforms existing methods, achieving 5.2% improvement in weighted accuracy. Darshana Priyasad, Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
ICASSP | 4 |
| 2020 | Two-Stream Deep Feature Modelling for Automated Video Endoscopy Data Analysis
Harshala Gammulle, Simon Denman, Sridha Sridharan, Clinton Fookes |
MICCAI (3) | 3 |
| 2020 | Semantic Consistency and Identity Mapping Multi-Component Generative Adversarial Network for Person Re-IdentificationabstractIn a real world environment, person re-identification (Re-ID) is a challenging task due to variations in lighting conditions, viewing angles, pose and occlusions. Despite recent performance gains, current person Re-ID algorithms still suffer heavily when encountering these variations. To address this problem, we propose a semantic consistency and identity mapping multi-component generative adversarial network (SC-IMGAN) which provides style adaptation from one to many domains. To ensure that transformed images are as realistic as possible, we propose novel identity mapping and semantic consistency losses to maintain identity across the diverse domains. For the Re-ID task, we propose a joint verification-identification quartet network which is trained with generated and real images, followed by an effective quartet loss for verification. Our proposed method outperforms state-of-the-art techniques on six challenging person Re-ID datasets: CUHK01, CUHK03, VIPeR, PRID2011, iLIDS and Market-1501. Amena Khatun, Simon Denman, Sridha Sridharan, Clinton Fookes |
WACV | 3 |
| 2020 | LSTM guided ensemble correlation filter tracking with appearance model pool
A. Venkata Subramanyam, Simon Denman, Sridha Sridharan, Clinton Fookes |
Comput. Vis. Image Underst. | 4 |
| 2020 | Joint identification-verification for person re-identification: A four stream deep learning approach with improved quartet loss function
Amena Khatun, Simon Denman, Sridha Sridharan, Clinton Fookes |
Comput. Vis. Image Underst. | 3 |
| 2020 | MTRNet++: One-stage mask-based scene text eraser
Osman Tursun, Simon Denman, Sabesan Sivipalan, Sridha Sridharan, Clinton Fookes |
Comput. Vis. Image Underst. | 5 |
| 2020 | Neural memory plasticity for medical anomaly detection
Tharindu Fernando, Simon Denman, David Ahmedt-Aristizabal, Sridha Sridharan, Kristin R. Laurens, Patrick J. Johnston, Clinton Fookes |
Neural Networks | 4 |
| 2020 | Fine-grained action segmentation using the semi-supervised action GAN
Harshala Gammulle, Simon Denman, Sridha Sridharan, Clinton Fookes |
Pattern Recognit. | 3 |
| 2020 | Context from within: Hierarchical context modeling for semantic segmentation
Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan |
Pattern Recognit. | 3 |
| 2020 | Correlation-aware adversarial domain adaptation and generalization
Mohammad Mahfujur Rahman, Clinton Fookes, Mahsa Baktash, Sridha Sridharan |
Pattern Recognit. | 4 |
| 2020 | Hierarchical Attention Network for Action Segmentation
Harshala Gammulle, Simon Denman, Sridha Sridharan, Clinton Fookes |
Pattern Recognit. Lett. | 3 |
| 2020 | Temporarily-Aware Context Modeling Using Generative Adversarial Networks for Speech Activity DetectionabstractThis paper presents a novel framework for Speech Activity Detection (SAD). Inspired by the recent success of multi-task learning approaches in the speech processing domain, we propose a novel joint learning framework for SAD. We utilise generative adversarial networks to automatically learn a loss function for joint prediction of the frame-wise speech/ non-speech classifications together with the next audio segment. In order to exploit the temporal relationships within the input signal, we propose a temporal discriminator which aims to ensure that the predicted signal is temporally consistent. We evaluate the proposed framework on multiple public benchmarks, including NIST OpenSAT' 17, AMI Meeting and HAVIC, where we demonstrate its capability to outperform state-of-the-art SAD approaches. Furthermore, our cross-database evaluations demonstrate the robustness of the proposed approach across different languages, accents, and acoustic environments. Tharindu Fernando, Sridha Sridharan, Mitchell McLaren, Darshana Priyasad, Simon Denman, Clinton Fookes |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Target-Specific Siamese Attention Network for Real-Time Object TrackingabstractDeep similarity trackers are able to track above real-time speed. However, their accuracy is considerably lower than deep classification based trackers since they avoid valuable online cues. To feed the target-specific information for real-time object tracking, we propose a novel Siamese attention network. Different types of attention mechanisms are used to capture different contexts of target information and then learned knowledge is used to feed target cues at different representation levels of similarity tracking. In addition, an online learning mechanism is employed to utilise the available target-specific data. The proposed tracker reduces the impact of noise in the target template and improves the accuracy of similarity tracking by feeding target cues into the similarity search. Extensive evaluation performed on OTB-2013/50/100 and VOT2018 benchmark datasets demonstrate the proposed tracker outperforms state-of-the-art approaches while maintaining real-time tracking speed. Thanikasalam Kokul, Clinton Fookes, Sridha Sridharan, Amirthalingam Ramanan, Amalka Pinidiyaarachchi |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2020 | Constrained Design of Deep Iris NetworksabstractDespite the promise of recent deep neural networks to provide more accurate and efficient iris recognition compared to traditional techniques, there are vital properties of the classic IrisCode which are almost unable to be achieved with current deep iris networks: the compactness of model and the small number of computing operations (FLOPs). This paper casts the iris network design process as a constrained optimization problem which takes model size and computation into account as learning criteria. On one hand, this allows us to fully automate the network design process to search for the optimal iris network architecture with the highest recognition accuracy confined to the computation and model compactness constraints. On the other hand, it allows us to investigate the optimality of the classic IrisCode and recent deep iris networks. It also enables us to learn an optimal iris network and demonstrate state-of-the-art performance with less computation and memory requirements. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan |
IEEE Trans. Image Process. | 3 |
| 2020 | Heart Sound Segmentation Using Bidirectional LSTMs With AttentionabstractOBJECTIVE: This paper proposes a novel framework for the segmentation of phonocardiogram (PCG) signals into heart states, exploiting the temporal evolution of the PCG as well as considering the salient information that it provides for the detection of the heart state. METHODS: We propose the use of recurrent neural networks and exploit recent advancements in attention based learning to segment the PCG signal. This allows the network to identify the most salient aspects of the signal and disregard uninformative information. RESULTS: The proposed method attains state-of-the-art performance on multiple benchmarks including both human and animal heart recordings. Furthermore, we empirically analyse different feature combinations including envelop features, wavelet and Mel Frequency Cepstral Coefficients (MFCC), and provide quantitative measurements that explore the importance of different features in the proposed approach. CONCLUSION: We demonstrate that a recurrent neural network coupled with attention mechanisms can effectively learn from irregular and noisy PCG recordings. Our analysis of different feature combinations shows that MFCC features and their derivatives offer the best performance compared to classical wavelet and envelop features. SIGNIFICANCE: Heart sound segmentation is a crucial pre-processing step for many diagnostic applications. The proposed method provides a cost effective alternative to labour extensive manual segmentation, and provides a more accurate segmentation than existing methods. As such, it can improve the performance of further analysis including the detection of murmurs and ejection clicks. The proposed method is also applicable for detection and segmentation of other one dimensional biomedical signals. Tharindu Fernando, Houman Ghaemmaghami, Simon Denman, Sridha Sridharan, Nayyar Hussain, Clinton Fookes |
IEEE J. Biomed. Health Informatics | 4 |
| 2020 | Memory Augmented Deep Generative Models for Forecasting the Next Shot Location in TennisabstractThis paper presents a novel framework for predicting shot location and type in tennis. Inspired by recent neuroscience discoveries, we incorporate neural memory modules to model the episodic and semantic memory components of a tennis player. We propose a Semi-Supervised Generative Adversarial Network architecture that couples these memory models with the automatic feature learning power of deep neural networks, and demonstrate methodologies for learning player level behavioral patterns with the proposed framework. We evaluate the effectiveness of the proposed model on tennis tracking data from the 2012 Australian Tennis Open and exhibit applications of the proposed method in discovering how players adapt their style depending on the match context. Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2019 | Forecasting Future Action Sequences with Neural Memory Networks
Harshala Gammulle, Simon Denman, Sridha Sridharan, Clinton Fookes |
BMVC | 3 |
| 2019 | Unified 2D and 3D Hand Pose Estimation from a Single Visible or X-ray Image
Akila Pemasiri, Kien Nguyen Thanh, Sridha Sridharan, Clinton Fookes |
BMVC | 3 |
| 2019 | Investigating Domain Sensitivity of DNN Embeddings for Speaker Recognition SystemsabstractA speaker embeddings framework achieves state-of-the-art speaker recognition performance by modeling speaker discriminant information directly using deep neural networks (DNNs). After the introduction of neural network based speaker embeddings, researchers have explored the requirements for training an effective embeddings network. However, the domain of the data used for system development should match the domain of operation for optimal performance. In this paper, we investigate the sensitivity of domain mismatch in the embeddings space. Specifically, degradation in performance is observed when back-end scoring with embeddings is performed with out-domain data. To compensate for the domain mismatch, we propose two novel deep domain adaptation techniques based on autoencoder architectures trained on embeddings in an unsupervised fashion. The results show that domain mismatch can be compensated effectively using autoencoders to adapt the out-domain data to in-domain. Ivan Himawan, Sridha Sridharan, Clinton Fookes |
ICASSP | 3 |
| 2019 | Predicting the Future: A Jointly Learnt Model for Action AnticipationabstractInspired by human neurological structures for action anticipation, we present an action anticipation model that enables the prediction of plausible future actions by forecasting both the visual and temporal future. In contrast to current state-of-the-art methods which first learn a model to predict future video features and then perform action anticipation using these features, the proposed framework jointly learns to perform the two tasks, future visual and temporal representation synthesis, and early action anticipation. The joint learning framework ensures that the predicted future embeddings are informative to the action anticipation task. Furthermore, through extensive experimental evaluations we demonstrate the utility of using both visual and temporal semantics of the scene, and illustrate how this representation synthesis could be achieved through a recurrent Generative Adversarial Network (GAN) framework. Our model outperforms the current state-of-the-art methods on multiple datasets: UCF101, UCF101-24, UT-Interaction and TV Human Interaction. Harshala Gammulle, Simon Denman, Sridha Sridharan, Clinton Fookes |
ICCV | 3 |
| 2019 | MTRNet: A Generic Scene Text EraserabstractText removal algorithms have been proposed for uni-lingual scripts with regular shapes and layouts. However, to the best of our knowledge, a generic text removal method which is able to remove all or user-specified text regions regardless of font, script, language or shape is not available. Developing such a generic text eraser for real scenes is a challenging task, since it inherits all the challenges of multi-lingual and curved text detection and inpainting. To fill this gap, we propose a mask-based text removal network (MTRNet). MTRNet is a conditional adversarial generative network (cGAN) with an auxiliary mask. The introduced auxiliary mask not only makes the cGAN a generic text eraser, but also enables stable training and early convergence on a challenging large-scale synthetic dataset, initially proposed for text detection in real scenes. What's more, MTRNet achieves state-of-the-art results on several real-world datasets including ICDAR 2013, ICDAR 2017 MLT, and CTW1500, without being explicitly trained on this data, outperforming previous state-of-the-art methods trained directly on these datasets. Osman Tursun, Simon Denman, Sabesan Sivipalan, Sridha Sridharan, Clinton Fookes |
ICDAR | 5 |
| 2019 | A Study of x-Vector Based Speaker Recognition on Short UtterancesabstractThe aim of this work is to gain insights into how the deep neural network (DNN) models should be trained for short utterance evaluation conditions in an x-vector based speaker verification system. The study suggests that the speaker embedding can be extracted with reduced dimensions for short utterance evaluation conditions. When the speaker embedding is extracted from deeper layer which has lower dimension, the x-vector system achieves 14% relative improvement over baseline approach on EER on NIST2010 5sec-5sec truncated conditions. We surmise that since short utterances have less phonetic information speaker discriminative x-vectors can be extracted from a deeper layer of the DNN which captures less phonetic information. Another interesting finding is that the x-vector system achieves 5% relative improvement on NIST2010 5sec-5sec evaluation condition when the back-end PLDA is trained using short utterance development data. The results confirms the intuitive expectation that duration of development utterances and the duration of evaluation utterances should be matched. Finally, for the duration mismatch condition, we propose a variance normalization approach for PLDA training that provides a 4% relative improvement on EER over baseline approach. Ahilan Kanagasundaram, Sridha Sridharan, Sriram Ganapathy, Prachi Singh, Clinton Fookes |
INTERSPEECH | 2 |
| 2019 | Coupled Generative Adversarial Network for Continuous Fine-Grained Action SegmentationabstractWe propose a novel conditional GAN (cGAN) model for continuous fine-grained human action segmentation, that utilises multi-modal data and learned scene context information. The proposed approach utilises two GANs: termed Action GAN and Auxiliary GAN, where the Action GAN is trained to operate over the current RGB frame while the Auxiliary GAN utilises supplementary information such as depth or optical flow. The goal of both GANs is to generate similar 'action codes', a vector representation of the current action. To facilitate this process a context extractor that incorporates data and recent outputs from both modes is used to extract context information to aids recognition performance. The result is a recurrent GAN architecture which learns a task specific loss function from multiple feature modalities. Extensive evaluations on variants of the proposed model to show the importance of utilising different streams of information such as context and auxiliary information in the proposed network; and show that our model is capable of outperforming state-of-the-art methods for three widely used datasets: 50 Salads, MERL Shopping and Georgia Tech Egocentric Activities, comprising both static and dynamic camera settings. Harshala Gammulle, Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
WACV | 4 |
| 2019 | Semantic Correspondence in the WildabstractSemantic correspondence estimation where the object instances depicted are deformed extensively from one instance to the next is a challenging problem in computer vision that has received much attention. Unfortunately, all existing approaches require prior knowledge of the object classes which are present in the image environment. This is an unwanted restriction as it can prevent the establishment of semantic correspondence across object classes in wild conditions when it is uncertain which classes will be of interest. In contrast, in this paper we formulate the semantic correspondence estimation task as a key point detection process in which image-to-class classification and image-to-image correspondence are solved simultaneously. Identifying object classes within the same framework to establish correspondence, increases this approach's applicability in real world scenarios. The use of object regions in the process also enhances the accuracy while constraining the search space, thus improving overall efficiency. This new approach is compared with the state-of-the-art on publicly available datasets to validate its capability for improved semantic correspondence estimation in wild conditions. Akila Pemasiri, Kien Nguyen Thanh, Sridha Sridharan, Clinton Fookes |
WACV | 3 |
| 2019 | Multi-Component Image Translation for Deep Domain GeneralizationabstractDomain adaption (DA) and domain generalization (DG) are two closely related methods which are both concerned with the task of assigning labels to an unlabeled data set. The only dissimilarity between these approaches is that DA can access the target data during the training phase, while the target data is totally unseen during the training phase in DG. The task of DG is challenging as we have no earlier knowledge of the target samples. If DA methods are applied directly to DG by a simple exclusion of the target data from training, poor performance will result for a given task. In this paper, we tackle the domain generalization challenge in two ways. In our first approach, we propose a novel deep domain generalization architecture utilizing synthetic data generated by a Generative Adversarial Network (GAN). The discrepancy between the generated images and synthetic images is minimized using existing domain discrepancy metrics such as maximum mean discrepancy or correlation alignment. In our second approach, we introduce a protocol for applying DA methods to a DG scenario by excluding the target data from the training phase, splitting the source data to training and validation parts, and treating the validation data as target data for DA. We conduct extensive experiments on four cross-domain benchmark datasets. Experimental results signify our proposed model outperforms the current state-of-the-art methods for DG. Mohammad Mahfujur Rahman, Clinton Fookes, Mahsa Baktash, Sridha Sridharan |
WACV | 4 |
| 2019 | Deep domain adaptation for anti-spoofing in speaker verification systems
Ivan Himawan, Fernando Villavicencio, Sridha Sridharan, Clinton Fookes |
Comput. Speech Lang. | 3 |
| 2019 | Multimodal clothing recognition for semantic search in unconstrained surveillance imagery
Michael Halstead, Simon Denman, Sridha Sridharan, Yingli Tian, Clinton Fookes |
J. Vis. Commun. Image Represent. | 3 |
| 2019 | Sparse over-complete patch matching
Akila Pemasiri, Kien Nguyen Thanh, Sridha Sridharan, Clinton Fookes |
Pattern Recognit. Lett. | 3 |
| 2019 | Scene Invariant Virtual Gates Using DNNsabstractUnderstanding where people are located and how they are moving about in an environment is critical for operators of large public spaces such as shopping centers, and large public infrastructures such as airports. Automated analysis of CCTV footage is increasingly being used to address this need through techniques that can count crowd sizes, estimate their density, and estimate the through-put of people into and/or out of a choke-point. A limitation of using CCTV based approaches, however, is the need to train models specific to each view which, for large environments with 100s or 1000s of cameras, can quickly become problematic. While there is some success in developing scene-invariant crowd counting and crowd density estimation approaches, much less attention has been given to developing scene-invariant solutions for through-put estimation. In this paper, we investigate the use of convolutional neural network and long short-term memory architectures to estimate pedestrian through-put from arbitrary CCTV viewpoints. To properly develop and demonstrate our approach, we present a new 22 view database featuring 44 h of pedestrian throughput annotation, containing over 11 000 annotated people; and using this proposed approach we show that we are able to outperform a scene-dependant approach across a diverse set of challenging view-points. Simon Denman, Clinton Fookes, Prasad K. D. V. Yarlagadda, Sridha Sridharan |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Understanding Patients' Behavior: Vision-Based Analysis of Seizure DisordersabstractA substantial proportion of patients with functional neurological disorders (FND) are being incorrectly diagnosed with epilepsy because their semiology resembles that of epileptic seizures (ES). Misdiagnosis may lead to unnecessary treatment and its associated complications. Diagnostic errors often result from an overreliance on specific clinical features. Furthermore, the lack of electrophysiological changes in patients with FND can also be seen in some forms of epilepsy, making diagnosis extremely challenging. Therefore, understanding semiology is an essential step for differentiating between ES and FND. Existing sensor-based and marker-based systems require physical contact with the body and are vulnerable to clinical situations such as patient positions, illumination changes, and motion discontinuities. Computer vision and deep learning are advancing to overcome these limitations encountered in the assessment of diseases and patient monitoring; however, they have not been investigated for seizure disorder scenarios. Here, we propose and compare two marker-free deep learning models, a landmark-based and a region-based model, both of which are capable of distinguishing between seizures from video recordings. We quantify semiology by using either a fusion of reference points and flow fields, or through the complete analysis of the body. Average leave-one-subject-out cross-validation accuracies for the landmark-based and region-based approaches of 68.1% and 79.6% in our dataset collected from 35 patients, reveal the benefit of video analytics to support automated identification of semiology in the challenging conditions of a hospital setting. David Ahmedt-Aristizabal, Simon Denman, Kien Nguyen Thanh, Sridha Sridharan, Sasha Dionisio, Clinton Fookes |
IEEE J. Biomed. Health Informatics | 4 |
| 2018 | GD-GAN: Generative Adversarial Networks for Trajectory Prediction and Group Detection in Crowds
Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
ACCV (1) | 3 |
| 2018 | Multi-level Sequence GAN for Group Activity Recognition
Harshala Gammulle, Simon Denman, Sridha Sridharan, Clinton Fookes |
ACCV (1) | 3 |
| 2018 | Learning Free-Form Deformations for 3D Object Reconstruction
Dominic Jack, Jhony K. Pontes, Sridha Sridharan, Clinton Fookes, Sareh Rowlands, Frédéric Maire, Anders P. Eriksson |
ACCV (2) | 3 |
| 2018 | Image2Mesh: A Learning Framework for Single Image 3D Reconstruction
Jhony K. Pontes, Chen Kong, Sridha Sridharan, Simon Lucey, Anders P. Eriksson, Clinton Fookes |
ACCV (1) | 3 |
| 2018 | Rethinking Planar Homography Estimation Using Perspective Fields
Simon Denman, Sridha Sridharan, Clinton Fookes |
ACCV (6) | 3 |
| 2018 | Calibrating Cameras in Poor-Conditioned Pitch-Based Sports GamesabstractCamera calibration is a preliminary step in sports analytics which enables us to transform player positions to standard playing area coordinates. While many camera calibration systems work well when the visual content contains sufficient clues, such as a key frame, calibrating without such information, such as may be needed when processing footage captured by a coach from the sidelines or stands, is challenging. In this paper an innovative automatic camera calibration system, which does not make use of any key frames, is presented for sports analytics. The proposed system consists of three components: a robust linear panorama module, a playing area estimation module, and a homography estimation module. It can eliminate distortion and calibrate the camera in each frame simultaneously, using correspondences between pairs of consecutive frames. Experiments on real data evaluate the performance and demonstrate the robustness of the system. Ruan Lakemond, Simon Denman, Sridha Sridharan, Clinton Fookes, Stuart Morgan |
ICASSP | 4 |
| 2018 | Hierarchical Relational Attention for Video Question AnsweringabstractVideo Question Answering (VideoQA) tasks require understanding of the connection of context specific video parts which are temporally distributed. Humans are capable of focusing on temporally distributed video scenes and also to find correspondence or relationships among these segments. To achieve similar capability, a hierarchical relational attention mechanism is proposed in this paper. The proposed VideoQA model derives attention on temporal segments i.e. video features based on each of the question words. Also, contextual relevance of these temporal segments are captured to derive the final video representation which leads to a better reasoning capability. We evaluate the performance of the proposed approach on the MSRVTT-QA and the MSVD-QA datasets to establish its superior performance over the state of the art. Muhammad Iqbal Hasan Chowdhury, Kien Nguyen Thanh, Sridha Sridharan, Clinton Fookes |
ICIP | 3 |
| 2018 | Deep Match Tracker: Classifying when Dissimilar, Similarity Matching when NotabstractVisual tracking frameworks employing Convolutional Neural Networks (CNNs) have shown state-of-the-art performance due to their hierarchical feature representation. While classification and update based deep neural net tracking have shown good performance in terms of accuracy, they have poor tracking speed. On the other hand, recent matching based techniques using CNNs show higher than real-time speed in tracking but this speed is achieved at a considerably lower accuracy. To successfully manage the trade-off between accuracy and speed, we propose a novel CNN architecture for visual tracking. We achieve this trade-off balance by using an approach in which consecutive similar frames are processed with a similarity matching technique, and dissimilar frames are processed with a classification approach within the CNN architecture. The tracking speed is improved by avoiding unnecessary model updates through the measurement of similarity between adjacent frames, while the accuracy is maintained by adopting a classification approach when needed, with deeper level features. Extensive evaluation performed on a publicly available benchmark dataset demonstrates our proposed tracker shows competitive performance while maintaining near real-time speed. Thanikasalam Kokul, Clinton Fookes, Sridha Sridharan, Amirthalingam Ramanan, Amalka Pinidiyaarachchi |
ICIP | 3 |
| 2018 | Non-rigid Reconstruction with a Single Moving RGB-D CameraabstractWe present a novel non-rigid reconstruction method using a moving RGB-D camera. Current approaches use only non-rigid part of the scene and completely ignore the rigid background. Non-rigid parts often lack sufficient geometric and photometric information for tracking large frame-to-frame motion. Our approach uses camera pose estimated from the rigid background for foreground tracking. This enables robust foreground tracking in situations where large frame-to-frame motion occurs. Moreover, we are proposing a multi-scale deformation graph which improves non-rigid tracking without compromising the quality of the reconstruction. We are also contributing a synthetic dataset which is made publically available for evaluating non-rigid reconstruction methods. The dataset provides frame-by-frame ground truth geometry of the scene, the camera trajectory, and masks for background foreground. Experimental results show that our approach is more robust in handling larger frame-to-frame motions and provides better reconstruction compared to state-of-the-art approaches. Shafeeq Elanattil, Peyman Moghadam, Sridha Sridharan, Clinton Fookes, Mark Cox |
ICPR | 3 |
| 2018 | Meta Transfer Learning for Facial Emotion RecognitionabstractThe use of deep learning techniques for automatic facial expression recognition has recently attracted great interest but developed models are still unable to generalize well due to the lack of large emotion datasets for deep learning. To overcome this problem, in this paper, we propose utilizing a novel transfer learning approach relying on PathNet and investigate how knowledge can be accumulated within a given dataset and how the knowledge captured from one emotion dataset can be transferred into another in order to improve the overall performance. To evaluate the robustness of our system, we have conducted various sets of experiments on two emotion datasets: SAVEE and eNTERFACE. The experimental results demonstrate that our proposed system leads to improvement in performance of emotion recognition and performs significantly better than the recent state-of-the-art schemes adopting fine-tuning/pre-trained approaches. Dung Nguyen Tien, Kien Nguyen Thanh, Sridha Sridharan, Iman Abbasnejad, David Dean, Clinton Fookes |
ICPR | 3 |
| 2018 | Elastic LiDAR Fusion: Dense Map-Centric Continuous-Time SLAMabstractThe concept of continuous-time trajectory representation has brought increased accuracy and efficiency to multi-modal sensor fusion in modern SLAM. However, regardless of these advantages, its offline property caused by the requirement of global batch optimization is critically hindering its relevance for real-time and life-long applications. In this paper, we present a dense map-centric SLAM method based on a continuous-time trajectory to cope with this problem. The proposed system locally functions in a similar fashion to conventional Continuous-Time SLAM (CT-SLAM). However, it removes the need for global trajectory optimization by introducing map deformation. The computational complexity of the proposed approach for loop closure does not depend on the operation time, but only on the size of the space it explored before the loop closure. It is therefore more suitable for long term operation compared to the conventional CT-SLAM. Furthermore, the proposed method reduces uncertainty in the reconstructed dense map by using probabilistic surface element (surfel) fusion. We demonstrate that the proposed method produces globally consistent maps without global batch trajectory optimization, and effectively reduces LiDAR noise by surfel fusion. Chanoh Park, Peyman Moghadam, Soohwan Kim, Alberto Elfes, Clinton Fookes, Sridha Sridharan |
ICRA | 6 |
| 2018 | Employing Phonetic Information in DNN Speaker Embeddings to Improve Speaker Recognition PerformanceabstractThe recent speaker embeddings framework has been shown to provide excellent performance on the task of text-independent speaker recognition. The framework is based on a deep neural network (DNN) trained to directly discriminate between speakers from traditional acoustic features such as Mel frequency cepstral coefficients. Prior studies on speaker recognition have found that phonetic information is valuable in the task of speaker identification, with systems being based on either bottleneck features (BFs) or tied-triphone state posteriors from a DNN trained for the task of speech recognition. In this paper, we analyze the role of phonetic BFs for DNN embeddings and explore methods to enhance the BFs further. Experimental results show that exploiting phonetic information encoded in BFs is very valuable for DNN speaker embeddings. Enriching the BFs using a cascaded DNN multi-task architecture is also shown to provide further improvements to the speaker embed- ding system. Ivan Himawan, Mitchell McLaren, Clinton Fookes, Sridha Sridharan |
INTERSPEECH | 5 |
| 2018 | Pedestrian Trajectory Prediction with Structured Memory Hierarchies
Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
ECML/PKDD (1) | 3 |
| 2018 | Investigating Deep Neural Networks for Speaker Diarization in the DIHARD ChallengeabstractWe investigate the use of deep neural networks (DNNs) for the speaker diarization task to improve performance under domain mismatched conditions. Three unsupervised domain adaptation techniques, namely inter-dataset variability compensation (IDVC), domain-invariant covariance normalization (DICN), and domain mismatch modeling (DMM), are applied on DNN based speaker embeddings to compensate for the mismatch in the embedding subspace. We present results conducted on the DIHARD data, which was released for the 2018 diarization challenge. Collected from a diverse set of domains, this data provides very challenging domain mismatched conditions for the diarization task. Our results provide insights into how the performance of our proposed system could be further improved. Ivan Himawan, Sridha Sridharan, Clinton Fookes, Ahilan Kanagasundaram |
SLT | 3 |
| 2018 | Tracking by Prediction: A Deep Generative Model for Mutli-person Localisation and TrackingabstractCurrent multi-person localisation and tracking systems have an over reliance on the use of appearance models for target re-identification and almost no approaches employ a complete deep learning solution for both objectives. We present a novel, complete deep learning framework for multi-person localisation and tracking. In this context we first introduce a light weight sequential Generative Adversarial Network architecture for person localisation, which overcomes issues related to occlusions and noisy detections, typically found in a multi person environment. In the proposed tracking framework we build upon recent advances in pedestrian trajectory prediction approaches and propose a novel data association scheme based on predicted trajectories. This removes the need for computationally expensive person re-identification systems based on appearance features and generates human like trajectories with minimal fragmentation. The proposed method is evaluated on multiple public benchmarks including both static and dynamic cameras and is capable of generating outstanding performance, especially among other recently proposed deep neural network based approaches. Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
WACV | 3 |
| 2018 | Task Specific Visual Saliency Prediction with Memory Augmented Conditional Generative Adversarial Networks
Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
WACV | 3 |
| 2018 | A Deep Four-Stream Siamese Convolutional Neural Network with Joint Verification and Identification Loss for Person Re-DetectionabstractState-of-the-art person re-identification systems that employ a triplet based deep network suffer from a poor generalization capability. In this paper, we propose a four stream Siamese deep convolutional neural network for person redetection that jointly optimises verification and identification losses over a four image input group. Specifically, the proposed method overcomes the weakness of the typical triplet formulation by using groups of four images featuring two matched (i.e. the same identity) and two mismatched images. This allows us to jointly increase the interclass variations and reduce the intra-class variations in the learned feature space. The proposed approach also optimises over both the identification and verification losses, further minimising intra-class variation and maximising inter-class variation, improving overall performance. Extensive experiments on four challenging datasets, VIPeR, CUHK01, CUHK03 and PRID2011, demonstrates that the proposed approach achieves state-of-the-art performance. Amena Khatun, Simon Denman, Sridha Sridharan, Clinton Fookes |
WACV | 3 |
| 2018 | Improving PLDA speaker verification performance using domain mismatch compensation techniques
Ahilan Kanagasundaram, Ivan Himawan, David Dean, Sridha Sridharan |
Comput. Speech Lang. | 5 |
| 2018 | Deep spatio-temporal feature fusion with compact bilinear pooling for multimodal emotion recognition
Dung Nguyen Tien, Kien Nguyen Thanh, Sridha Sridharan, David Dean, Clinton Fookes |
Comput. Vis. Image Underst. | 3 |
| 2018 | Tree Memory Networks for modelling long-term temporal dependencies
Tharindu Fernando, Simon Denman, Aaron McFadyen, Sridha Sridharan, Clinton Fookes |
Neurocomputing | 4 |
| 2018 | Soft + Hardwired attention: An LSTM framework for human trajectory prediction and abnormal event detection
Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes |
Neural Networks | 3 |
| 2018 | Super-resolution for biometrics: A comprehensive survey
Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Massimo Tistarelli, Mark S. Nixon |
Pattern Recognit. | 3 |
| 2018 | Interactive Sports Analytics: An Intelligent Interface for Utilizing Trajectories for Interactive Sports Play Retrieval and AnalyticsabstractAnalytics in professional sports has experienced a dramatic growth in the last decade due to the wide deployment of player and ball tracking systems in team sports, such as basketball and soccer. With the massive amount of fine-grained data being generated, new data-points are being generated, which can shed light on player and team performance. However, due to the complexity of plays in continuous sports, these data-points often lack the specificity and context to enable meaningful retrieval and analytics. In this article, we present an intelligent human--computer interface that utilizes trajectories instead of words, which enables specific play retrieval in sports. Various techniques of alignment, templating, and hashing were utilized by our system and they are tailored to multi-agent scenario so that interactive speeds can be achieved. We conduct a user study to compare our method to the conventional keywords-based system and the results show that our method significantly improves the retrieval quality. We also show how our interface can be utilized for broadcast purposes, where a user can draw and interact with trajectories on a broadcast view using computer vision techniques. Additionally, we show that our method can also be used for interactive analytics of player performance, which enables the users to move players around and see how performance changes as a function of position and proximity to other players. Long Sha, Patrick Lucey, Yisong Yue, Xinyu Wei 0004, Jennifer A. Hobbs, Charlie Rohlf, Sridha Sridharan |
ACM Trans. Comput. Hum. Interact. | 7 |
| 2017 | Compact Model Representation for 3D Reconstructionabstract3D reconstruction from 2D images is a central problem in computer vision. Recent works have been focusing on reconstruction directly from a single image. It is well known however that only one image cannot provide enough information for such a reconstruction. A prior knowledge that has been entertained are 3D CAD models due to its online ubiquity. A fundamental question is how to compactly represent millions of CAD models while allowing generalization to new unseen objects with fine-scaled geometry. We introduce an approach to compactly represent a 3D mesh. Our method first selects a 3D model from a graph structure by using a novel free-form deformation FFD 3D-2D registration, and then the selected 3D model is refined to best fit the image silhouette. We perform a comprehensive quantitative and qualitative analysis that demonstrates impressive dense and realistic 3D reconstruction from single images. Jhony K. Pontes, Chen Kong, Anders P. Eriksson, Clinton Fookes, Sridha Sridharan, Simon Lucey |
3DV | 5 |
| 2017 | Fast, Dense Feature SDM on an iPhoneabstractIn this paper, we present our method for enabling dense SDM to run at over 90 FPS on a mobile device. Our contributions are two-fold. Drawing inspiration from the FFT, we propose a Sparse Compositional Regression (SCR) framework, which enables a significant speed up over classical dense regressors. Second, we propose a binary approximation to SIFT features. Binary Approximated SIFT (BASIFT) features, which are a computationally efficient approximation to SIFT, a commonly used feature with SDM. We demonstrate the performance of our algorithm on an iPhone 7, and show that we achieve similar accuracy to SDM. Ashton Fagg, Simon Lucey, Sridha Sridharan |
FG | 3 |
| 2017 | Deep features-based expression-invariant tied factor analysis for emotion recognitionabstractVideo-based facial expression recognition is an open research challenge not solved by the current state-of-the-art. On the other hand, static image based emotion recognition is highly important when videos are not available and human emotions need to be determined from a single shot only. This paper proposes sequential-based and image-based tied factor analysis frameworks with a deep network that simultaneously addresses these two problems. For video-based data, we first extract deep convolutional temporal appearance features from image sequences and then these features are fed into a generative model that constructs a low-dimensional observed space for all individuals, depending on the facial expression sequences. After learning the sequential expression components of the transition matrices among the expression manifolds, we use a Gaussian probabilistic approach to design an efficient classifier for temporal facial expression recognition. Furthermore, we analyse the utility of proposed video-based methods for image-based emotion recognition learning static tied factor analysis parameters. Meanwhile, this model can be used to predict the expressive face image sequences from given neutral faces. Recognition results achieved on three public benchmark databases: CK+, JAFFE, and FER2013, clearly indicate our approach achieves effective performance over the current techniques of handling sequential and static facial expression variations. Sarasi Munasinghe, Clinton Fookes, Sridha Sridharan |
IJCB | 3 |
| 2017 | A cascaded long short-term memory (LSTM) driven generic visual question answering (VQA)abstractA cascaded long short-term memory (LSTM) architecture with discriminant feature learning is proposed for the task of question answering on real world images. The proposed LSTM architecture jointly learns visual features and parts of speech (POS) tags of question words or tokens. Also, dimensionality of deep visual features is reduced by applying Principal Component Analysis (PCA) technique. In this manner, the proposed question answering model captures the generic pattern of question for a given context of image which is just not constricted within the training dataset. Empirical outcome shows that this kind of approach significantly improves the accuracy. It is believed that this kind of generic learning is a step towards a real-world visual question answering (VQA) system which will perform well for all possible forms of open-ended natural language queries. Iqbal Chowdhury, Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan |
ICIP | 4 |
| 2017 | Deep discovery of facial motions using a shallow embedding layerabstractUnique encoding of the dynamics of facial actions has potential to provide a spontaneous facial expression recognition system. The most promising existing approaches rely on deep learning of facial actions. However, current approaches are often computationally intensive and require a great deal of memory/processing time, and typically the temporal aspect of facial actions are often ignored, despite the potential wealth of information available from the spatial dynamic movements and their temporal evolution over time from neutral state to apex state. To tackle aforementioned challenges, we propose a deep learning framework by using the 3D convolutional filters to extract spatio-temporal features, followed by the LSTM network which is able to integrate the dynamic evolution of short-duration of spatio-temporal features as an emotion progresses from the neutral state to the apex state. In order to reduce the redundancy of parameters and accelerate the learning of the recurrent neural network, we propose a shallow embedding layer to reduce the number of parameters in the LSTM by up to 98% without sacrificing recognition accuracy. As the fully connected layer approximately contains 95% of the parameters in the network, we decrease the number of parameters in this layer before passing features to the LSTM network, which significantly improves training speed and enables the possibility of deploying a state of the art deep network on real-time applications. We evaluate our proposed framework on the DISFA and UNBC-McMaster Shoulder pain datasets. Afsane Ghasemi, Mahsa Baktash, Simon Denman, Sridha Sridharan, Dung Nguyen Tien, Clinton Fookes |
ICIP | 4 |
| 2017 | Single image depth prediction using super-column super-pixel featuresabstractDepth prediction from a single monocular image is a challenging yet valuable task, as often a depth sensor is not available. The state-of-the-art approach [1] combines a deep fully convolutional network (DFCN) with a conditional random field (CRF), allowing the CRF to correct and smooth the depth values estimated by the DFCN according to efficient contextual modeling. However, using the output of the DFCN as unary input for CRF is limited by using only the last layer of the DFCN. The middle layers of the DFCN have been shown to carry useful information for other scene understanding tasks, which may help to improve the prediction quality. This paper proposes a novel super-column superpixel (SCSP) feature that is the combination of multiple layers of the DFCN after a super-pixel pooling process. The proposed approach based on the SCSP features reduces the root mean square (rms) error of the prediction by more than 16% in NYUv2 dataset. Xufeng Guo, Kien Nguyen Thanh, Simon Denman, Clinton Fookes, Sridha Sridharan |
ICIP | 5 |
| 2017 | Facial analysis in the wild with LSTM networksabstractThe promise of computer vision systems to efficiently and accurately recognize faces and facial variations in naturally occurring circumstances still remains elusive. In this paper we present two separate systems for face analysis, both of which use Long Short Term Memory (LSTM) Networks: unconstrained video-based face verification (FaceVideoModel) and spontaneous facial expression recognition (ExpModel). Since LSTM models have influential ability to capture sequential patterns, our results prove such LSTM models have significant advantages over other proposed models in the state-of-the-art for facial analysis in the wild. On the recently introduced Youtube Faces database our FaceModel achieves an accuracy of 98.70% for face verification with a value of 99.94% for the Area Under Curve (AUC) and 1.2% Equal Error Rate (EER) which is the best performance on this database compared to other recently proposed methods. Experimental results achieved through the proposed ExpModel on the challenging FER2013 dataset, including the CK+ database, also demonstrate the effectiveness of our deep model for facial expression recognition. Sarasi Kankanamge, Clinton Fookes, Sridha Sridharan |
ICIP | 3 |
| 2017 | Gate connected convolutional neural network for object trackingabstractConvolutional neural networks (CNNs) have been employed in visual tracking due to their rich levels of feature representation. While the learning capability of a CNN increases with its depth, unfortunately spatial information is diluted in deeper layers which hinders its important ability to localise targets. To successfully manage this trade-off, we propose a novel residual network based gating CNN architecture for object tracking. Our deep model connects the front and bottom convolutional features with a gate layer. This new network learns discriminative features while reducing the spatial information lost. This architecture is pre-trained to learn generic tracking characteristics. In online tracking, an efficient domain adaptation mechanism is used to accurately learn the target appearance with limited samples. Extensive evaluation performed on a publicly available benchmark dataset demonstrates our proposed tracker outperforms state-of-the-art approaches. Thanikasalam Kokul, Clinton Fookes, Sridha Sridharan, Amirthalingam Ramanan, Amalka Pinidiyaarachchi |
ICIP | 3 |
| 2017 | Domain Mismatch Modeling of Out-Domain i-Vectors for PLDA Speaker VerificationabstractThe state-of-the-art i-vector based probabilistic linear discriminant analysis (PLDA) trained on non-target (or out- domain) data significantly affects the speaker verification performance due to the domain mismatch between training and evaluation data. To improve the speaker verification performance, sufficient amount of domain mismatch compensated out-domain data must be used to train the PLDA models successfully. In this paper, we propose a domain mismatch model- ing (DMM) technique using maximum-a-posteriori (MAP) estimation to model and compensate the domain variability from the out-domain training i-vectors. From our experimental results, we found that the DMM technique can achieve at least a 24% improvement in EER over an out-domain only base- line when speaker labels are available. Further improvement of 3% is obtained when combining DMM with domain-invariant covariance normalization (DICN) approach. The DMM/DICN combined technique is shown to perform better than in-domain PLDA system with only 200 labeled speakers or 2,000 unlabeled i-vectors. Ivan Himawan, David Dean, Sridha Sridharan |
INTERSPEECH | 4 |
| 2017 | From Affine Rank Minimization Solution to Sparse ModelingabstractCompressed sensing is a simple and efficient technique that has a number of applications in signal processing and machine learning. In machine learning it provides answers to questions such as: "under what conditions is the sparse representation of data efficient?", "when is learning a large margin classifier directly on the compressed domain possible?", and "why does a large margin classifier learn more effectively if the data is sparse?". This work tackles the problem of feature representation from the context of sparsity and affine rank minimization by leveraging compressed sensing from the learning perspective in order to provide answers to the aforementioned questions. We show, for a full-rank signal, the high dimensional sparse representation of data is efficient because from the classifiers viewpoint such a representation is in fact a low dimensional problem. We provide practical bounds on the linear classifier to investigate the relationship between the SVM classifier in the high dimensional and compressed domains and show for the high dimensional sparse signals, when the bounds are tight, directly learning in the compressed domain is possible. Iman Abbasnejad, Sridha Sridharan, Simon Denman, Clinton Fookes, Simon Lucey |
WACV | 2 |
| 2017 | Two Stream LSTM: A Deep Fusion Framework for Human Action RecognitionabstractIn this paper we address the problem of human action recognition from video sequences. Inspired by the exemplary results obtained via automatic feature learning and deep learning approaches in computer vision, we focus our attention towards learning salient spatial features via a convolutional neural network (CNN) and then map their temporal relationship with the aid of Long-Short-Term-Memory (LSTM) networks. Our contribution in this paper is a deep fusion framework that more effectively exploits spatial features from CNNs with temporal features from LSTM models. We also extensively evaluate their strengths and weaknesses. We find that by combining both the sets of features, the fully connected features effectively act as an attention mechanism to direct the LSTM to interesting parts of the convolutional feature sequence. The significance of our fusion method is its simplicity and effectiveness compared to other state-of-the-art methods. The evaluation results demonstrate that this hierarchical multi stream fusion method has higher performance compared to single stream mapping methods allowing it to achieve high accuracy outperforming current state-of-the-art methods in three widely used databases: UCF11, UCFSports, jHMDB. Harshala Gammulle, Simon Denman, Sridha Sridharan, Clinton Fookes |
WACV | 3 |
| 2017 | Deep Context Modeling for Semantic SegmentationabstractDeep convolutional neural networks (DCNNs) have been employed in many computer vision tasks with great success due to their robustness in feature learning. One of the advantages of DCNNs is their representation robustness to object locations, which is useful for object recognition tasks. However, this also discards spatial information, which is useful when dealing with topological information of the image (e.g. scene parsing, face recognition). Adopting graphical models (GMs) to incorporate spatial and contextual information into the DCNNs is expected to improve the performance of DCNN-based computer vision tasks. Recent research has shown that combining DCNNs and Conditional Random Fields (CRFs) can significantly improve scene parsing accuracy. This is achieved either through the combination of their independent outputs or through their application as a cascade. In this work, we propose a novel strategy to incorporate CRFs deeper inside DCNNs by modeling a CRF as a DCNN layer which is pluggable into any layer of a DCNN. This implants spatial and contextual information into the DCNN, allowing end-to-end training, better controlling the spatial constraints and improving segmentation accuracy. The new strategy for coupling graphical models with the state-of-the-art fully convolutional neural network has shown promising results on the PASCAL-Context dataset. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan |
WACV | 3 |
| 2017 | Deep Spatio-Temporal Features for Multimodal Emotion RecognitionabstractAutomatic emotion recognition has attracted great interest and numerous solutions have been proposed, most of which focus either individually on facial expression or acoustic information. While more recent research has considered multimodal approaches, individual modalities are often combined only by simple fusion at the feature and/or decision-level. In this paper, we introduce a novel approach using 3-dimensional convolutional neural networks (C3Ds) to model the spatio-temporal information, cascaded with multimodal deep-belief networks (DBNs) that can represent the audio and video streams. Experiments conducted on the eNTERFACE multimodal emotion database demonstrate that this approach leads to improved multimodal emotion recognition performance and significantly outperforms recent state-of-the-art proposals. Dung Nguyen Tien, Kien Nguyen Thanh, Sridha Sridharan, Afsane Ghasemi, David Dean, Clinton Fookes |
WACV | 3 |
| 2017 | Cross database audio visual speech adaptation for phonetic spoken term detection
Shahram Kalantari, David Dean, Sridha Sridharan |
Comput. Speech Lang. | 3 |
| 2017 | Fine-grained action recognition of boxing punches from depth imagery
Soudeh Kasiri Bidhendi, Clinton Fookes, Sridha Sridharan, Stuart Morgan |
Comput. Vis. Image Underst. | 3 |
| 2017 | Long range iris recognition: A survey
Kien Nguyen Thanh, Clinton Fookes, Raghavender R. Jillela, Sridha Sridharan, Arun Ross |
Pattern Recognit. | 4 |
| 2016 | Deeper and wider fully convolutional network coupled with conditional random fields for scene labelingabstractDeep convolutional neural networks (DCNNs) have been employed in many computer vision tasks with great success due to their robustness in feature learning. One of the advantages of DCNNs is their representation robustness to object locations, which is useful for object recognition tasks. However, this also discards spatial information, which is useful when dealing with topological information of the image (e.g. scene labeling, face recognition). In this paper, we propose a deeper and wider network architecture to tackle the scene labeling task. The depth is achieved by incorporating predictions from multiple early layers of the DCNN. The width is achieved by combining multiple outputs of the network. We then further refine the parsing task by adopting graphical models (GMs) as a post-processing step to incorporate spatial and contextual information into the network. The new strategy for a deeper, wider convolutional network coupled with graphical models has shown promising results on the PASCAL-Context dataset. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan |
ICIP | 3 |
| 2016 | A robust UAV landing site detection system using mid-level discriminative patchesabstractThe forced landing problem has become one of the main impediments to UAV's entering civilian airspace. Unfortunately there is no robust forced landing site detection system that will reliably detect a safe landing site. One of the main reasons for this is the difficulty in considering the various classes of surface, to determine whether they are safe or not. We propose a robust UAV landing site detection system using midlevel discriminative patches. The training and tuning process uses a dataset containing 1600 randomly selected Google map images with weak labels.We then show how the output from multiple mid-level discriminative patch detectors can be combined to indicate the level or danger for a given region. The proposed technique reliably detects safe landing areas in UAV imagery, and achieves improved performance over the state-of-the art. The proposed system outperforms the baseline system by 29.4% for completeness and 33.9% for correctness, and is invariant to the changes of illumination, sharpness and resolution of images. Xufeng Guo, Simon Denman, Clinton Fookes, Sridha Sridharan |
ICPR | 4 |
| 2016 | Speakers In The Wild (SITW): The QUT Speaker Recognition SystemabstractThis paper presents the QUT speaker recognition system, as a competing system in the Speakers In The Wild (SITW) speaker recognition challenge. Our proposed system achieved an overall ranking of second place, in the main core-core condition evaluations of the SITW challenge. This system uses an ivector/ PLDA approach, with domain adaptation and a deep neural network (DNN) trained to provide feature statistics. The statistics are accumulated by using class posteriors from the DNN, in place of GMM component posteriors in a typical GMM UBM i-vector/PLDA system. Once the statistics have been collected, the i-vector computation is carried out as in a GMM-UBM based system. We apply domain adaptation to the extracted i-vectors to ensure robustness against dataset variability, PLDA modelling is used to capture speaker and session variability in the i-vector space, and the processed i-vectors are compared using the batch likelihood ratio. The final scores are calibrated to obtain the calibrated likelihood scores, which are then used to carry out speaker recognition and evaluate the performance of the system. Finally, we explore the practical application of our system to the core-multi condition recordings of the SITW data and propose a technique for speaker recognition in recordings with multiple speakers. Houman Ghaemmaghami, Ivan Himawan, David Dean, Ahilan Kanagasundaram, Sridha Sridharan, Clinton Fookes |
INTERSPEECH | 6 |
| 2016 | Short Utterance Variance Modelling and Utterance Partitioning for PLDA Speaker VerificationabstractThis paper analyses the short utterance probabilistic linear discriminant analysis (PLDA) speaker verification with utterance partitioning and short utterance variance (SUV) modelling approaches. Experimental studies have found that instead of using single long-utterance as enrolment data, if long enrolled utterance is partitioned into multiple short utterances and average of short utterance i-vectors is used as enrolled data, that improves the Gaussian PLDA (GPLDA) speaker verification. This is because short utterance i-vectors have speaker, session and utterance variations, and utterance-partitioning approach compensates the utterance variation. Subsequently, SUV-PLDA is also studied with utterance partitioning approach, and utterance partitioning-based SUV-GPLDA system shows relative improvement of 9% and 16% in EER for NIST 2008 and NIST 2010 truncated 10sec-10sec evaluation condition as utterance partitioning approach compensates the utterance variation and SUV modelling approach compensates the mismatch between full-length development data and short-length evaluation data. Ahilan Kanagasundaram, David Dean, Sridha Sridharan, Clinton Fookes, Ivan Himawan |
INTERSPEECH | 3 |
| 2016 | Discovery of facial motions using deep machine perceptionabstractDeep, intuitive understanding of facial motions has the potential to provide an intelligent facial expression system as well as a unique encoding of the dynamics of facial actions. The most promising existing approaches rely on extracting hand crafted features; and existing approaches typically work best in constrained conditions and do not generalise well to varying environmental conditions which make them poorly suited to applications such as real-time human robot interactions. In this paper, we propose a multi-label deep learning based facial action detector, which along with a linear SVM classifier outperforms state of the art approaches such as HOG and LBP. We show that our approach can be generalized to other datasets by learning inner data structure, encoding facial actions, and providing a hierarchical representation of facial features. Our experimental results also demonstrate the efficiency of using image patches, which results in faster learning convergence while outperforms holistic approaches. We evaluate our proposed frame-work on the DISFA and CK+ datasets. Afsane Ghasemi, Simon Denman, Sridha Sridharan, Clinton Fookes |
WACV | 3 |
| 2016 | A study of speaker clustering for speaker attribution in large telephone conversation datasets
Houman Ghaemmaghami, David Dean, Sridha Sridharan, David A. van Leeuwen |
Comput. Speech Lang. | 3 |
| 2016 | Detecting rare events using Kullback-Leibler divergence: A weakly supervised approach
Jingxin Xu, Simon Denman, Clinton Fookes, Sridha Sridharan |
Expert Syst. Appl. | 4 |
| 2016 | Feature mapping using far-field microphones for distant speech recognition
Ivan Himawan, Petr Motlícek, David Imseng, Sridha Sridharan |
Speech Commun. | 4 |
| 2016 | Discovering Team Structures in Soccer from Spatiotemporal DataabstractIn team sports like soccer, utilizing tracking data for analysis is challenging due to the dynamic and multi-agent nature of the data. The biggest issue surrounds the changing of positions or “roles” between players on a frame-to-frame basis, which causes misalignment of the data and makes it difficult to perform team analysis. In this paper, we present an unsupervised method to learn a formation template which allows us to “align” the tracking data at the frame level. Not only does this approach give important contextual information to facilitate large-scale analysis (e.g., we know when a player is in the left-wing position compared to left-back), it also yields the team structure or “formation” which serves as a strong descriptor for identifying a team's style. The utility of the approach is demonstrated on a full season of player and ball tracking data from a professional soccer league consisting of over 21.5 million frames of player tracking data. Alina Bialkowski, Patrick Lucey, Peter Carr 0001, Iain A. Matthews, Sridha Sridharan, Clinton Fookes |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2016 | Forecasting the Next Shot Location in Tennis Using Fine-Grained Spatiotemporal Tracking DataabstractIn professional sport, an enormous amount of fine-grain performance data can be generated at near millisecond intervals in the form of vision-based tracking data. One of the first sports to embrace this technology has been tennis, where Hawk-Eye technology has been used to both aid umpiring decisions, and to visualize shot trajectories for broadcast purposes. Despite the high-level of accuracy of the tracking systems and the sheer volume of spatiotemporal data they generate, the use of this data for player performance analysis and prediction has been lacking. In this research, we use ball and player tracking data from “Hawk-Eye” to discover unique player styles and predict within-point events. We move beyond current analysis that only incorporates coarse match statistics (i.e., serves, winners, number of shots, and volleys) and use spatial and temporal information which better characterizes the tactics and tendencies of each player. Using a probabilistic graphical model, we are able to model player behaviors which enables us to: 1) find the factors such as location and speed of the incoming shot which are most conducive to a player hitting a winner (i.e., “sweet-spot”) or cause an error, and 2) do “live in-point” prediction - based on the shots being played during a rally we estimate the probability of the outcome (e.g., winner, continuation, or error) and the location of the next shot. As player behavior depends on the opponent, we use model adaptation to enhance our prediction. We show the utility of our approach by analyzing the play of Djokovic, Nadal, and Federer at the 2012 Australian Tennis Open. Xinyu Wei 0004, Patrick Lucey, Stuart Morgan, Sridha Sridharan |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2015 | Large scale monitoring of crowds and building utilisation: A new database and distributed approachabstractPublic buildings and large infrastructure are typically monitored by tens or hundreds of cameras, all capturing different physical spaces and observing different types of interactions and behaviours. However to date, in large part due to limited data availability, crowd monitoring and operational surveillance research has focused on single camera scenarios which are not representative of real-world applications. In this paper we present a new, publicly available database for large scale crowd surveillance. Footage from 12 cameras for a full work day covering the main floor of a busy university campus building, including an internal and external foyer, elevator foyers, and the main external approach are provided; alongside annotation for crowd counting (single or multi-camera) and pedestrian flow analysis for 10 and 6 sites respectively. We describe how this large dataset can be used to perform distributed monitoring of building utilisation, and demonstrate the potential of this dataset to understand and learn the relationship between different areas of a building. Simon Denman, Clinton Fookes, David Ryan, Sridha Sridharan |
AVSS | 4 |
| 2015 | Searching for semantic person queries using channel representationsabstractIt is not uncommon to hear a person of interest described by their height, build, and clothing (i.e. type and colour). These semantic descriptions are commonly used by people to describe others, as they are quick to relate and easy to understand. However such queries are not easily utilised within intelligent surveillance systems as they are difficult to transform into a representation that can be searched for automatically in large camera networks. In this paper we propose a novel approach that transforms such a semantic query into an avatar that is searchable within a video stream, and demonstrate state-of-the-art performance for locating a subject in video based on a description. Simon Denman, Michael Halstead, Clinton Fookes, Sridha Sridharan |
ICASSP | 4 |
| 2015 | A cluster-voting approach for speaker diarization and linking of Australian broadcast news recordingsabstractWe present a clustering-only approach to the problem of speaker diarization to eliminate the need for the commonly employed and computationally expensive Viterbi segmentation and realignment stage. We use multiple linear segmentations of a recording and carry out complete-linkage clustering within each segmentation scenario to obtain a set of clustering decisions for each case. We then collect all clustering decisions, across all cases, to compute a pairwise vote between the segments and conduct complete-linkage clustering to cluster them at a resolution equal to the minimum segment length used in the linear segmentations. We use our proposed cluster-voting approach to carry out speaker diarization and linking across the SAIVT-BNEWS corpus of Australian broadcast news data. We compare our technique to an equivalent baseline system with Viterbi realignment and show that our approach can outperform the baseline technique with respect to the diarization error rate (DER) and attribution error rate (AER). Houman Ghaemmaghami, David Dean, Sridha Sridharan |
ICASSP | 3 |
| 2015 | Improving out-domain PLDA speaker verification using unsupervised inter-dataset variability compensation approachabstractExperimental studies have found that when the state-of-the-art probabilistic linear discriminant analysis (PLDA) speaker verification systems are trained using out-domain data, it significantly affects speaker verification performance due to the mismatch between development data and evaluation data. To overcome this problem we propose a novel unsupervised inter dataset variability (IDV) compensation approach to compensate the dataset mismatch. IDV-compensated PLDA system achieves over 10% relative improvement in EER values over out-domain PLDA system by effectively compensating the mismatch between in-domain and out-domain data. Ahilan Kanagasundaram, David Dean, Sridha Sridharan |
ICASSP | 3 |
| 2015 | Detecting rare events using Kullback-Leibler divergenceabstractOne main challenge in developing a system for visual surveillance event detection is the annotation of target events in the training data. By making use of the assumption that events with security interest are often rare compared to regular behaviours, this paper presents a novel approach by using Kullback-Leibler (KL) divergence for rare event detection in a weakly supervised learning setting, where only clip-level annotation is available. It will be shown that this approach outperforms state-of-the-art methods on a popular real-world dataset, while preserving real time performance. Jingxin Xu, Simon Denman, Clinton Fookes, Sridha Sridharan |
ICASSP | 4 |
| 2015 | Combat sports analytics: Boxing punch classification using overhead depthimageryabstractIn competitive combat sporting environments like boxing, the statistics on a boxer's performance, including the amount and type of punches thrown, provide a valuable source of data and feedback which is routinely used for coaching and performance improvement purposes. This paper presents a robust framework for the automatic Classification of a boxer's punches. Overhead depth imagery is employed to alleviate challenges associated with occlusions, and robust body-part tracking is developed for the noisy time-of-flight sensors. Punch recognition is addressed through both a multi-class SVM and Random Forest classifiers. coarse-to-fine hierarchical SVM classifier is presented based on prior knowledge of boxing punches. This framework has been applied to shadow boxing image sequences taken at the Australian Institute of Sport with 8 elite boxers. Results demonstrate the effectiveness of the proposed approach, with the hierarchical SVM classifier yielding a 96% accuracy, signifying its suitability for analysing athletes punches in boxing bouts. Soudeh Kasiri Bidhendi, Clinton Fookes, Stuart Morgan, David T. Martin, Sridha Sridharan |
ICIP | 5 |
| 2015 | Improving deep convolutional neural networks with unsupervised feature learningabstractThe latest generation of Deep Convolutional Neural Networks (DCNN) have dramatically advanced challenging computer vision tasks, especially in object detection and object classification, achieving state-of-the-art performance in several computer vision tasks including text recognition, sign recognition, face recognition and scene understanding. The depth of these supervised networks has enabled learning deeper and hierarchical representation of features. In parallel, unsupervised deep learning such as Convolutional Deep Belief Network (CDBN) has also achieved state-of-the-art in many computer vision tasks. However, there is very limited research on jointly exploiting the strength of these two approaches. In this paper, we investigate the learning capability of both methods. We compare the output of individual layers and show that many learnt filters and outputs of the corresponding level layer are almost similar for both approaches. Stacking the DCNN on top of unsupervised layers or replacing layers in the DCNN with the corresponding learnt layers in the CDBN can improve the recognition/classification accuracy and training computational expense. We demonstrate the validity of the proposal on ImageNet dataset. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan |
ICIP | 3 |
| 2015 | Class-specific sparse codes for representing activitiesabstractIn this paper we investigate the effectiveness of class specific sparse codes in the context of discriminative action classification. The bag-of-words representation is widely used in activity recognition to encode features, and although it yields state-of-the art performance with several feature descriptors it still suffers from large quantization errors and reduces the overall performance. Recently proposed sparse representation methods have been shown to effectively represent features as a linear combination of an over complete dictionary by minimizing the reconstruction error. In contrast to most of the sparse representation methods which focus on Sparse-Reconstruction based Classification (SRC), this paper focuses on a discriminative classification using a SVM by constructing class-specific sparse codes for motion and appearance separately. Experimental results demonstrates that separate motion and appearance specific sparse coefficients provide the most effective and discriminative representation for each class compared to a single class-specific sparse coefficients. Sabanadesan Umakanthan, Simon Denman, Clinton Fookes, Sridha Sridharan |
ICIP | 4 |
| 2015 | The QUT-NOISE-SRE protocol for the evaluation of noisy speaker recognitionabstractThe QUT-NOISE-SRE protocol is designed to mix the large QUT-NOISE database, consisting of over 10 hours of back- ground noise, collected across 10 unique locations covering 5 common noise scenarios, with commonly used speaker recognition datasets such as Switchboard, Mixer and the speaker recognition evaluation (SRE) datasets provided by NIST. By allowing common, clean, speech corpora to be mixed with a wide variety of noise conditions, environmental reverberant responses, and signal-to-noise ratios, this protocol provides a solid basis for the development, evaluation and benchmarking of robust speaker recognition algorithms, and is freely available to download alongside the QUT-NOISE database. In this work, we use the QUT-NOISE-SRE protocol to evaluate a state-of-the-art PLDA i-vector speaker recognition system, demonstrating the importance of designing voice-activity-detection front-ends specifically for speaker recognition, rather than aiming for perfect coherence with the true speech/non-speech boundaries. David Dean, Ahilan Kanagasundaram, Houman Ghaemmaghami, Sridha Sridharan |
INTERSPEECH | 5 |
| 2015 | Complete-linkage clustering for voice activity detection in audio and visual speechabstractWe propose a novel technique for conducting robust voice activity detection (VAD) in high-noise recordings. We use Gaussian mixture modeling (GMM) to train two generic models; speech and non-speech. We then score smaller segments of a given (unseen) recording against each of these GMMs to obtain two respective likelihood scores for each segment. These scores are used to compute a dissimilarity measure between pairs of segments and to carry out complete-linkage clustering of the segments into speech and non-speech clusters. We compare the accuracy of our method against state-of-the-art and standardised VAD techniques to demonstrate an absolute improvement of 15% in half-total error rate (HTER) over the best performing baseline system and across the QUT-NOISE-TIMIT database. We then apply our approach to the Audio-Visual Database of American English (AVDBAE) to demonstrate the performance of our algorithm in using visual, audio-visual or a proposed fusion of these features. Houman Ghaemmaghami, David Dean, Shahram Kalantari, Sridha Sridharan, Clinton Fookes |
INTERSPEECH | 4 |
| 2015 | Channel selection in the short-time modulation domain for distant speech recognitionabstractAutomatic speech recognition from multiple distant microphones poses significant challenges because of noise and reverberations.The quality of speech acquisition may vary between microphones because of movements of speakers and channel distortions.This paper proposes a channel selection approach for selecting reliable channels based on selection criterion operating in the short-term modulation spectrum domain.The proposed approach quantifies the relative strength of speech from each microphone and speech obtained from beamforming modulations.The new technique is compared experimentally in the real reverb conditions in terms of perceptual evaluation of speech quality (PESQ) measures and word error rate (WER).Overall improvement in recognition rate is observed using delay-sum and superdirective beamformers compared to the case when the channel is selected randomly using circular microphone arrays. Ivan Himawan, Petr Motlícek, Sridha Sridharan, David Dean, Dian Tjondronegoro |
INTERSPEECH | 3 |
| 2015 | Cross database training of audio-visual hidden Markov models for phone recognitionabstractSpeech recognition can be improved by using visual information in the form of lip movements of the speaker in addition to audio information.To date, state-of-the-art techniques for audio-visual speech recognition continue to use audio and visual data of the same database for training their models.In this paper, we present a new approach to make use of one modality of an external dataset in addition to a given audio-visual dataset.By so doing, it is possible to create more powerful models from other extensive audio-only databases and adapt them on our comparatively smaller multi-stream databases.Results show that the presented approach outperforms the widely adopted synchronous hidden Markov models (HMM) trained jointly on audio and visual data of a given audio-visual database for phone recognition by 29% relative.It also outperforms the external audio models trained on extensive external audio datasets and also internal audio models by 5.5% and 46% relative respectively.We also show that the proposed approach is beneficial in noisy environments where the audio source is affected by the environmental noise. Shahram Kalantari, David Dean, Houman Ghaemmaghami, Sridha Sridharan, Clinton Fookes |
INTERSPEECH | 4 |
| 2015 | Incorporating visual information for spoken term detectionabstractSpoken term detection (STD) is the task of looking up a spoken term in a large volume of speech segments. In order to provide fast search, speech segments are first indexed into an intermediate representation using speech recognition engines which provide multiple hypotheses for each speech segment. Approximate matching techniques are usually applied at the search stage to compensate the poor performance of automatic speech recognition engines during indexing. Recently, using visual information in addition to audio information has been shown to improve phone recognition performance, particularly in noisy environments. In this paper, we will make use of visual information in the form of lip movements of the speaker in indexing stage and will investigate its effect on STD performance. Particularly, we will investigate if gains in phone recognition accuracy will carry through the approximate matching stage to provide similar gains in the final audio-visual STD system over a traditional audio only approach. We will also investigate the effect of using visual information on STD performance in different noise environments. Shahram Kalantari, David Dean, Sridha Sridharan |
INTERSPEECH | 3 |
| 2015 | Improving PLDA speaker verification using WMFD and linear-weighted approaches in limited microphone data conditionsabstractThis paper proposes the addition of a weighted median Fisher discriminator (WMFD) projection prior to length-normalised Gaussian probabilistic linear discriminant analysis (GPLDA) modelling in order to compensate the additional session variation. In limited microphone data conditions, a linear-weighted approach is introduced to increase the influence of microphone speech dataset. The linear-weighted WMFD-projected GPLDA system shows improvements in EER and DCF values over the pooled LDA- and WMFD-projected GPLDA systems in inter-view-interview condition as WMFD projection extracts more speaker discriminant information with limited number of sessions/ speaker data, and linear-weighted GPLDA approach estimates reliable model parameters with limited microphone data. Ahilan Kanagasundaram, David Dean, Sridha Sridharan |
INTERSPEECH | 3 |
| 2015 | Investigating in-domain data requirements for PLDA trainingabstractThis paper analyzes the limitations upon the amount of indomain (NIST SREs) data required for training a probabilistic linear discriminant analysis (PLDA) speaker verification system based on out-domain (Switchboard) total variability subspaces.By limiting the number of speakers, the number of sessions per speaker and the length of active speech per session available in the target domain for PLDA training, we investigated the relative effect of these three parameters on PLDA speaker verification performance in the NIST 2008 and NIST 2010 speaker recognition evaluation datasets.Experimental results indicate that while these parameters depend highly on each other, to beat out-domain PLDA training, more than 10 seconds of active speech should be available for at least 4 sessions/speaker for a minimum of 800 speakers.If further data is available, considerable improvement can be made over solely out-domain PLDA training. David Dean, Ahilan Kanagasundaram, Sridha Sridharan |
INTERSPEECH | 4 |
| 2015 | Dataset-invariant covariance normalization for out-domain PLDA speaker verificationabstractIn this paper we introduce a novel domain-invariant covariance normalization (DICN) technique to relocate both in-domain and out-domain i-vectors into a third dataset-invariant space, providing an improvement for out-domain PLDA speaker verification with a very small number of unlabelled in-domain adaptation i-vectors.By capturing the dataset variance from a global mean using both development out-domain i-vectors and limited unlabelled in-domain i-vectors, we could obtain domaininvariant representations of PLDA training data.The DICNcompensated out-domain PLDA system is shown to perform as well as in-domain PLDA training with as few as 500 unlabelled in-domain i-vectors for NIST-2010 SRE and 2000 unlabelled in-domain i-vectors for NIST-2008 SRE, and considerable relative improvement over both out-domain and in-domain PLDA development if more are available. Ahilan Kanagasundaram, David Dean, Sridha Sridharan |
INTERSPEECH | 4 |
| 2015 | Predicting Serves in Tennis using Style PriorsabstractIn professional sport, an enormous amount of fine-grain performance data can be generated at near millisecond intervals in the form of vision-based tracking data. One of the first sports to embrace this technology has been tennis, where Hawk-Eye technology has been used to both aid umpiring decisions, and to visualize shot trajectories for broadcast purposes. These data have tremendous untapped applications in terms of "opponent planning'', where a large amount of recent data is used to learn contextual behavior patterns of individual players, and ultimately predict the likelihood of a particular type of serve. Since the type of serve selected by a player may be contingent on the match context (i.e., is the player down break-point, or is serving for the match etc.), the characteristics of the player (i.e., the player may have a very fast serve, hit heavy with topspin or kick, or slice serves into the body) as well as the characteristics of the opponent (e.g., the opponent may prefer to play from the baseline or "chip-and-charge'' into the net). In this paper we present a method which recommends the most likely serves of a player in a given context. We show by utilizing a "style prior", we can improve the prediction/recommendation. Such an approach also allows us to quantify the similarity between players, which is useful in enriching the dataset for future prediction. We conduct our analysis on Hawk-Eye data collected from three recent Australian Open Grand-Slam Tournaments and show how our approach can be used in practice. Xinyu Wei 0004, Patrick Lucey, Stuart Morgan, Peter Carr 0001, Machar Reid, Sridha Sridharan |
KDD | 6 |
| 2015 | An evaluation of crowd counting methods, features and regression models
David Ryan, Simon Denman, Sridha Sridharan, Clinton Fookes |
Comput. Vis. Image Underst. | 3 |
| 2015 | Automatic surveillance in transportation hubs: No longer just about catching the bad guy
Simon Denman, Tristan Kleinschmidt, David Ryan, Paul Barnes, Sridha Sridharan, Clinton Fookes |
Expert Syst. Appl. | 5 |
| 2015 | Searching for people using semantic soft biometric descriptions
Simon Denman, Michael Halstead, Clinton Fookes, Sridha Sridharan |
Pattern Recognit. Lett. | 4 |
| 2015 | An Efficient and Robust System for Multiperson Event Detection in Real-World Indoor Surveillance ScenesabstractDue to the popularity of security cameras in public places, it is of interest to design an intelligent system that can efficiently detect events automatically. This paper proposes a novel algorithm for multiperson event detection. To ensure greater than real-time performance, features are extracted directly from compressed MPEG video. A novel histogram-based feature descriptor that captures the angles between extracted particle trajectories is proposed, which allows us to capture motion patterns for multiperson events in the video. To alleviate the need for fine-grained annotation, we propose the use of labeled latent Dirichlet allocation, a weakly supervised method that allows the use of coarse temporal annotations, which are much simpler to obtain. This novel system is able to run at ~10 times real time, while preserving state-of-the-art detection performance for multiperson events on a 100-h real-world surveillance data set (TRECVid surveillance event detection). Jingxin Xu, Simon Denman, Sridha Sridharan, Clinton Fookes |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2015 | Score-Level Multibiometric Fusion Based on Dempster-Shafer Theory Incorporating Uncertainty FactorsabstractWhile existing multibiometic Dempster-Shafer theory fusion approaches have demonstrated promising performance, they do not model the uncertainty appropriately, suggesting that further improvement can be achieved. This research seeks to develop a unified framework for multimodal biometric fusion to take advantage of the uncertainty concept of Dempster-Shafer theory, improving the performance of multibiometric authentication systems. Modeling uncertainty as a function of uncertainty factors affecting the recognition performance of the biometric systems helps to address the uncertainty of the data and the confidence of the fusion outcome. A weighted combination of quality measures and classifiers performance (equal error rate) is proposed to encode the uncertainty concept to improve the fusion. We also found that quality measures contribute unequally to the recognition performance; thus, selecting only significant factors and fusing them with a Dempster-Shafer approach to generate an overall quality score play an important role in the success of uncertainty modeling. The proposed approach achieved a competitive performance (approximate 1% EER) in comparison with other Dempster-Shafer-based approaches and other conventional fusion approaches. Kien Nguyen Thanh, Simon Denman, Sridha Sridharan, Clinton Fookes |
IEEE Trans. Hum. Mach. Syst. | 3 |
| 2014 | Learning Detectors Quickly with Stationary Statistics
Jack Valmadre, Sridha Sridharan, Simon Lucey |
ACCV (1) | 2 |
| 2014 | Forecasting Events Using an Augmented Hidden Conditional Random Field
Xinyu Wei 0004, Patrick Lucey, Stephen Vidas, Stuart Morgan, Sridha Sridharan |
ACCV (4) | 5 |
| 2014 | An MRF based abnormal event detection approach using motion and appearance featuresabstractAbnormal event detection has attracted a lot of attention in the computer vision research community during recent years due to the increased focus on automated surveillance systems to improve security in public places. Due to the scarcity of training data and the definition of an abnormality being dependent on context, abnormal event detection is generally formulated as a data-driven approach where activities are modeled in an unsupervised fashion during the training phase. In this work, we use a Gaussian mixture model (GMM) to cluster the activities during the training phase, and propose a Gaussian mixture model based Markov random field (GMM-MRF) to estimate the likelihood scores of new videos in the testing phase. Further-more, we propose two new features: optical acceleration, and the histogram of optical flow gradients; to detect the presence of any abnormal objects and speed violations in the scene. We show that our proposed method outperforms other state of the art abnormal event detection algorithms on publicly available UCSD dataset. Hajananth Nallaivarothayan, Clinton Fookes, Simon Denman, Sridha Sridharan |
AVSS | 4 |
| 2014 | Improving PLDA speaker verification with limited development dataabstractThis paper analyses the probabilistic linear discriminant analysis (PLDA) speaker verification approach with limited development data. This paper investigates the use of the median as the central tendency of a speaker's i-vector representation, and the effectiveness of weighted discriminative techniques on the performance of state-of-the-art length-normalised Gaussian PLDA (GPLDA) speaker verification systems. The analysis within shows that the median (using a median fisher discriminator (MFD)) provides a better representation of a speaker when the number of representative i-vectors available during development is reduced, and that further, usage of the pair-wise weighting approach in weighted LDA and weighted MFD provides further improvement in limited development conditions. Best performance is obtained using a weighted MFD approach, which shows over 10% relative improvement in EER over the baseline GPLDA system on mismatched and interview-interview conditions. Ahilan Kanagasundaram, David Dean, Sridha Sridharan |
ICASSP | 3 |
| 2014 | Large-Scale Analysis of Soccer Matches Using Spatiotemporal Tracking DataabstractAlthough the collection of player and ball tracking data is fast becoming the norm in professional sports, large-scale mining of such spatiotemporal data has yet to surface. In this paper, given an entire season's worth of player and ball tracking data from a professional soccer league (≈400,000,000 data points), we present a method which can conduct both individual player and team analysis. Due to the dynamic, continuous and multi-player nature of team sports like soccer, a major issue is aligning player positions over time. We present a "role-based" representation that dynamically updates each player's relative role at each frame and demonstrate how this captures the short-term context to enable both individual player and team analysis. We discover role directly from data by utilizing a minimum entropy data partitioning method and show how this can be used to accurately detect and visualize formations, as well as analyze individual player behavior. Alina Bialkowski, Patrick Lucey, Peter Carr 0001, Yisong Yue, Sridha Sridharan, Iain A. Matthews |
ICDM | 5 |
| 2014 | Locating People in Video from Semantic Descriptions: A New Database and ApproachabstractThe location of previously unseen and unregistered individuals in complex camera networks from semantic descriptions is a time consuming and often inaccurate process carried out by human operators, or security staff on the ground. To promote the development and evaluation of automated semantic description based localisation systems, we present a new, publicly available, unconstrained 110 sequence database, collected from 6 stationary cameras. Each sequence contains detailed semantic information for a single search subject who appears in the clip (gender, age, height, build, hair and skin colour, clothing type, texture and colour), and between 21 and 290 frames for each clip are annotated with the target subject location (over 11, 000 frames are annotated in total). A novel approach for localising a person given a semantic query is also proposed and demonstrated on this database. The proposed approach incorporates clothing colour and type (for clothing worn below the waist), as well as height and build to detect people. A method to assess the quality of candidate regions, as well as a symmetry driven approach to aid in modelling clothing on the lower half of the body, is proposed within this approach. An evaluation on the proposed dataset shows that a relative improvement in localisation accuracy of up to 21% is achieved over the baseline technique. Michael Halstead, Simon Denman, Sridha Sridharan, Clinton Fookes |
ICPR | 3 |
| 2014 | Multiple Instance Dictionary Learning for Activity RepresentationabstractThis paper presents an effective feature representation method in the context of activity recognition. Efficient and effective feature representation plays a crucial role not only in activity recognition, but also in a wide range of applications such as motion analysis, tracking, 3D scene understanding etc. In the context of activity recognition, local features are increasingly popular for representing videos because of their simplicity and efficiency. While they achieve state-of-the-art performance with low computational requirements, their performance is still limited for real world applications due to a lack of contextual information and models not being tailored to specific activities. We propose a new activity representation framework to address the shortcomings of the popular, but simple bag-of-words approach. In our framework, first multiple instance SVM (mi-SVM) is used to identify positive features for each action category and the k-means algorithm is used to generate a codebook. Then locality-constrained linear coding is used to encode the features into the generated codebook, followed by spatio-temporal pyramid pooling to convey the spatio-temporal statistics. Finally, an SVM is used to classify the videos. Experiments carried out on two popular datasets with varying complexity demonstrate significant performance improvement over the base-line bag-of-feature method. Sabanadesan Umakanthan, Simon Denman, Clinton Fookes, Sridha Sridharan |
ICPR | 4 |
| 2014 | An iterative speaker re-diarization scheme for improving speaker-based entity extraction in multimedia archivesabstractIn this paper we present a novel scheme for improving speaker diarization by making use of repeating speakers across multiple recordings within a large corpus.We call this technique speaker re-diarization and demonstrate that it is possible to reuse the initial speaker-linked diarization outputs to boost diarization accuracy within individual recordings.We first propose and evaluate two novel re-diarization techniques.We demonstrate their complementary characteristics and fuse the two techniques to successfully conduct speaker re-diarization across the SAIVT-BNEWS corpus of Australian broadcast data.This corpus contains recurring speakers in various independent recordings that need to be linked across the dataset.We show that our speaker re-diarization approach can provide a relative improvement of 23% in diarization error rate (DER), over the original diarization results, as well as improve the estimated number of speakers and the cluster purity and coverage metrics. Houman Ghaemmaghami, David Dean, Sridha Sridharan |
INTERSPEECH | 3 |
| 2014 | Local inter-session variability modelling for object classificationabstractObject classification is plagued by the issue of session variation. Session variation describes any variation that makes one instance of an object look different to another, for instance due to pose or illumination variation. Recent work in the challenging task of face verification has shown that session variability modelling provides a mechanism to overcome some of these limitations. However, for computer vision purposes, it has only been applied in the limited setting of face verification. In this paper we propose a local region based intersession variability (ISV) modelling approach, and apply it to challenging real-world data. We propose a region based session variability modelling approach so that local session variations can be modelled, termed Local ISV. We then demonstrate the efficacy of this technique on a challenging real-world fish image database which includes images taken underwater, providing significant real-world session variations. This Local ISV approach provides a relative performance improvement of, on average, 23% on the challenging MOBIO, Multi-PIE and SCface face databases. It also provides a relative performance improvement of 35% on our challenging fish image dataset. Kaneswaran Anantharajah, ZongYuan Ge, Chris McCool, Simon Denman, Clinton Fookes, Peter I. Corke, Dian Tjondronegoro, Sridha Sridharan |
WACV | 8 |
| 2014 | Predicting movie ratings from audience behaviorsabstractWe propose a method of representing audience behavior through facial and body motions from a single video stream, and use these features to predict the rating for feature-length movies. This is a very challenging problem as: i) the movie viewing environment is dark and contains views of people at different scales and viewpoints; ii) the duration of feature-length movies is long (80-120 mins) so tracking people uninterrupted for this length of time is still an unsolved problem; and iii) expressions and motions of audience members are subtle, short and sparse making labeling of activities unreliable. To circumvent these issues, we use an infrared illuminated test-bed to obtain a visually uniform input. We then utilize motion-history features which capture the subtle movements of a person within a pre-defined volume, and then form a group representation of the audience by a histogram of pair-wise correlations over a small-window of time. Using this group representation, we learn our movie rating classifier from crowd-sourced ratings collected by rottentomatoes.com and show our prediction capability on audiences from 30 movies across 250 subjects (> 50 hrs). Rajitha Navarathna, Patrick Lucey, Peter Carr 0001, Elizabeth J. Carter, Sridha Sridharan, Iain A. Matthews |
WACV | 5 |
| 2014 | Understanding and analyzing a large collection of archived swimming videosabstractIn elite sports, nearly all performances are captured on video. Despite the massive amounts of video that has been captured in this domain over the last 10-15 years, most of it remains in an “unstructured” or “raw” form, meaning it can only be viewed or manually annotated/tagged with higher-level event labels which is time consuming and subjective. As such, depending on the detail or depth of annotation, the value of the collected repositories of archived data is minimal as it does not lend itself to large-scale analysis and retrieval. One such example is swimming, where each race of a swimmer is captured on a camcorder and in-addition to the split-times (i.e., the time it takes for each lap), stroke rate and stroke-lengths are manually annotated. In this paper, we propose a vision-based system which effectively “digitizes” a large collection of archived swimming races by estimating the location of the swimmer in each frame, as well as detecting the stroke rate. As the videos are captured from moving hand-held cameras which are located at different positions and angles, we show our hierarchical-based approach to tracking the swimmer and their different parts is robust to these issues and allows us to accurately estimate the swimmer location and stroke rates. Long Sha, Patrick Lucey, Sridha Sridharan, Stuart Morgan, Dave Pease |
WACV | 3 |
| 2014 | I-vector based speaker recognition using advanced channel compensation techniques
Ahilan Kanagasundaram, David Dean, Sridha Sridharan, Mitchell McLaren, Robbie Vogt |
Comput. Speech Lang. | 3 |
| 2014 | Scene invariant multi camera crowd counting
David Ryan, Simon Denman, Clinton Fookes, Sridha Sridharan |
Pattern Recognit. Lett. | 4 |
| 2014 | Real-time video event detection in crowded scenes using MPEG derived features: A multiple instance learning approach
Jingxin Xu, Simon Denman, Vikas Reddy, Clinton Fookes, Sridha Sridharan |
Pattern Recognit. Lett. | 5 |
| 2014 | Improving short utterance i-vector speaker verification using utterance variance modelling and compensation techniques
Ahilan Kanagasundaram, David Dean, Sridha Sridharan, Javier Gonzalez-Dominguez, Joaquín González-Rodríguez, Daniel Ramos-Castro |
Speech Commun. | 3 |
| 2014 | Optimal Camera Planning Under Versatile User Constraints in Multi-Camera Image Processing SystemsabstractThe selection of optimal camera configurations (camera locations, orientations, etc.) for multi-camera networks remains an unsolved problem. Previous approaches largely focus on proposing various objective functions to achieve different tasks. Most of them, however, do not generalize well to large scale networks. To tackle this, we propose a statistical framework of the problem as well as propose a trans-dimensional simulated annealing algorithm to effectively deal with it. We compare our approach with a state-of-the-art method based on binary integer programming (BIP) and show that our approach offers similar performance on small scale problems. However, we also demonstrate the capability of our approach in dealing with large scale problems and show that our approach produces better results than two alternative heuristics designed to deal with the scalability issue of BIP. Last, we show the versatility of our approach using a number of specific scenarios. Junbin Liu, Sridha Sridharan, Clinton Fookes, Tim Wark |
IEEE Trans. Image Process. | 2 |
| 2013 | Rank Minimization across Appearance and Shape for AAM Ensemble FittingabstractActive Appearance Models (AAMs) employ a paradigm of inverting a synthesis model of how an object can vary in terms of shape and appearance. As a result, the ability of AAMs to register an unseen object image is intrinsically linked to two factors. First, how well the synthesis model can reconstruct the object image. Second, the degrees of freedom in the model. Fewer degrees of freedom yield a higher likelihood of good fitting performance. In this paper we look at how these seemingly contrasting factors can complement one another for the problem of AAM fitting of an ensemble of images stemming from a constrained set (e.g. an ensemble of face images of the same person). Sridha Sridharan, Jason Saraghi, Simon Lucey |
ICCV | 2 |
| 2013 | Improving short utterance based i-vector speaker recognition using source and utterance-duration normalization techniquesabstractA significant amount of speech is typically required for speaker verification system development and evaluation, especially in the presence of large intersession variability. This paper in-troduces a source and utterance-duration normalized linear dis-criminant analysis (SUN-LDA) approaches to compensate ses-sion variability in short-utterance i-vector speaker verification systems. Two variations of SUN-LDA are proposed where normalization techniques are used to capture source variation from both short and full-length development i-vectors, one based upon pooling (SUN-LDA-pooled) and the other on con-catenation (SUN-LDA-concat) across the duration and source-dependent session variation. Both the SUN-LDA-pooled and SUN-LDA-concat techniques are shown to provide improve-ment over traditional LDA on NIST 08 truncated 10sec-10sec evaluation conditions, with the highest improvement obtained with the SUN-LDA-concat technique achieving a relative im-provement of 8 % in EER for mis-matched conditions and over 3 % for matched conditions over traditional LDA approaches. Index Terms: speaker verification, i-vector, total-variability, LDA, WCCN Ahilan Kanagasundaram, David Dean, Javier Gonzalez-Dominguez, Sridha Sridharan, Daniel Ramos-Castro, Joaquín González-Rodríguez |
INTERSPEECH | 4 |
| 2013 | Improving the PLDA based speaker verification in limited microphone data conditionsabstractA significant amount of speech data is required to develop a robust speaker verification system, but it is difficult to find enough development speech to match all expected conditions.In this paper we introduce a new approach to Gaussian probabilistic linear discriminant analysis (GPLDA) to estimate reliable model parameters as a linearly weighted model taking more input from the large volume of available telephone data and smaller proportional input from limited microphone data.In comparison to a traditional pooled training approach, where the GPLDA model is trained over both telephone and microphone speech, this linear-weighted GPLDA approach is shown to provide better EER and DCF performance in microphone and mixed conditions in both the NIST 2008 and NIST 2010 evaluation corpora.Based upon these results, we believe that linear-weighted GPLDA will provide a better approach than pooled GPLDA, allowing for the further improvement of GPLDA speaker verification in conditions with limited development data. Ahilan Kanagasundaram, David Dean, Javier Gonzalez-Dominguez, Sridha Sridharan, Daniel Ramos-Castro, Joaquín González-Rodríguez |
INTERSPEECH | 4 |
| 2013 | Multiple cameras for audio-visual speech recognition in an automotive environment
Rajitha Navarathna, David Dean, Sridha Sridharan, Patrick Lucey |
Comput. Speech Lang. | 3 |
| 2013 | Eigenvoice modelling for cross likelihood ratio based speaker clustering: A Bayesian approach
David Wang 0002, Robbie Vogt, Sridha Sridharan |
Comput. Speech Lang. | 3 |
| 2013 | Feature-domain super-resolution for iris recognition
Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Simon Denman |
Comput. Vis. Image Underst. | 3 |
| 2013 | Evaluation of two-view geometry methods with automatic ground-truth generation
Ruan Lakemond, Clinton Fookes, Sridha Sridharan |
Image Vis. Comput. | 3 |
| 2013 | Fourier Lucas-Kanade AlgorithmabstractIn this paper, we propose a framework for both gradient descent image and object alignment in the Fourier domain. Our method centers upon the classical Lucas & Kanade (LK) algorithm where we represent the source and template/model in the complex 2D Fourier domain rather than in the spatial 2D domain. We refer to our approach as the Fourier LK (FLK) algorithm. The FLK formulation is advantageous when one preprocesses the source image and template/model with a bank of filters (e.g., oriented edges, Gabor, etc.) as 1) it can handle substantial illumination variations, 2) the inefficient preprocessing filter bank step can be subsumed within the FLK algorithm as a sparse diagonal weighting matrix, 3) unlike traditional LK, the computational cost is invariant to the number of filters and as a result is far more efficient, and 4) this approach can be extended to the Inverse Compositional (IC) form of the LK algorithm where nearly all steps (including Fourier transform and filter bank preprocessing) can be precomputed, leading to an extremely efficient and robust approach to gradient descent image matching. Further, these computational savings translate to nonrigid object alignment tasks that are considered extensions of the LK algorithm, such as those found in Active Appearance Models (AAMs). Simon Lucey, Rajitha Navarathna, Ahmed Ashraf 0001, Sridha Sridharan |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2012 | Use of brain computer interface to drive: preliminary resultsabstractThis paper reports on the implementation of a non-invasive electroencephalography-based brain-computer interface to control functions of a car in a driving simulator. The system is comprised of a Cleveland Medical Devices BioRadio 150 physiological signal recorder, a MATLAB-based BCI and an OKTAL SCANeR advanced driving experience simulator. Deanna Hood, Damian Joseph, Andry Rakotonirainy, Sridha Sridharan, Clinton Fookes |
AutomotiveUI | 4 |
| 2012 | Unusual Scene Detection Using Distributed Behaviour Model and Sparse RepresentationabstractThe ability to detect unusual events in surviellance footage as they happen is a highly desireable feature for a surveillance system. However, this problem remains challenging in crowded scenes due to occlusions and the clustering of people. In this paper, we propose using the Distributed Behavior Model (DBM), which has been widely used in computer graphics, for video event detection. Our approach does not rely on object tracking, and is robust to camera movements. We use sparse coding for classification, and test our approach on various datasets. Our proposed approach outperforms a state-of-the-art work which uses the social force model and Latent Dirichlet Allocation. Jingxin Xu, Simon Denman, Clinton Fookes, Sridha Sridharan |
AVSS | 4 |
| 2012 | Activity Analysis in Complicated Scenes Using DFT Coefficients of Particle TrajectoriesabstractModelling activities in crowded scenes is very challenging as object tracking is not robust in complicated scenes and optical flow does not capture long range motion. We propose a novel approach to analyse activities in crowded scenesusing a "bag of particle trajectories". Particle trajectoriesare extracted from foreground regions within short video clips using particle video, which estimates long rangemotion in contrast to optical flow which is only concerned with inter-frame motion. Our applications include temporal video segmentation and anomaly detection, and we perform our evaluation on several real-world datasets containing complicated scenes. We show that our approaches achieve state-of-the-art performance for both tasks. Jingxin Xu, Simon Denman, Sridha Sridharan, Clinton Fookes |
AVSS | 3 |
| 2012 | Improved facial expression recognition via uni-hyperplane classificationabstractLarge margin learning approaches, such as support vector machines (SVM), have been successfully applied to numerous classification tasks, especially for automatic facial expression recognition. The risk of such approaches however, is their sensitivity to large margin losses due to the influence from noisy training examples and outliers which is a common problem in the area of affective computing (i.e., manual coding at the frame level is tedious so coarse labels are normally assigned). In this paper, we leverage the relaxation of the parallel-hyperplanes constraint and propose the use of modified correlation filters (MCF). The MCF is similar in spirit to SVMs and correlation filters, but with the key difference of optimizing only a single hyperplane. We demonstrate the superiority of MCF over current techniques on a battery of experiments. Sien W. Chew, Simon Lucey, Patrick Lucey, Sridha Sridharan, Jeff F. Conn |
CVPR | 4 |
| 2012 | Feature-domain super-resolution framework for Gabor-based face and iris recognitionabstractThe low resolution of images has been one of the major limitations in recognising humans from a distance using their biometric traits, such as face and iris. Superresolution has been employed to improve the resolution and the recognition performance simultaneously, however the majority of techniques employed operate in the pixel domain, such that the biometric feature vectors are extracted from a super-resolved input image. Feature-domain superresolution has been proposed for face and iris, and is shown to further improve recognition performance by capitalising on direct super-resolving the features which are used for recognition. However, current feature-domain superresolution approaches are limited to simple linear features such as Principal Component Analysis (PCA) and Linear Discriminant Analysis (LDA), which are not the most discriminant features for biometrics. Gabor-based features have been shown to be one of the most discriminant features for biometrics including face and iris. This paper proposes a framework to conduct super-resolution in the non-linear Gabor feature domain to further improve the recognition performance of biometric systems. Experiments have confirmed the validity of the proposed approach, demonstrating superior performance to existing linear approaches for both face and iris biometrics. Kien Nguyen Thanh, Sridha Sridharan, Simon Denman, Clinton Fookes |
CVPR | 2 |
| 2012 | On the Statistical Determination of Optimal Camera Configurations in Large Scale Surveillance Networks
Junbin Liu, Clinton Fookes, Tim Wark, Sridha Sridharan |
ECCV (1) | 4 |
| 2012 | Efficient Articulated Trajectory Reconstruction Using Dynamic Programming and Filters
Jack Valmadre, Yingying Zhu 0004, Sridha Sridharan, Simon Lucey |
ECCV (1) | 3 |
| 2012 | Hand-held monocular SLAM in thermal-infraredabstractThermal-infrared imagery is relatively robust to many of the failure conditions of visual and laser-based SLAM systems, such as fog, dust and smoke. The ability to use thermal-infrared video for localization is therefore highly appealing for many applications. However, operating in thermal-infrared is beyond the capacity of existing SLAM implementations. This paper presents the first known monocular SLAM system designed and tested for hand-held use in the thermal-infrared modality. The implementation includes a flexible feature detection layer able to achieve robust feature tracking in high-noise, low-texture thermal images. A novel approach for structure initialization is also presented. The system is robust to irregular motion and capable of handling the unique mechanical shutter interruptions common to thermal-infrared cameras. The evaluation demonstrates promising performance of the algorithm in several environments. Stephen Vidas, Sridha Sridharan |
ICARCV | 2 |
| 2012 | Speaker attribution of multiple telephone conversations using a complete-linkage clustering approachabstractIn this paper we propose and evaluate a speaker attribution system using a complete-linkage clustering method. Speaker attribution refers to the annotation of a collection of spoken audio based on speaker identities. This can be achieved using diarization and speaker linking. The main challenge associated with attribution is achieving computational efficiency when dealing with large audio archives. Traditional agglomerative clustering methods with model merging and retraining are not feasible for this purpose. This has motivated the use of linkage clustering methods without retraining. We first propose a diarization system using complete-linkage clustering and show that it outperforms traditional agglomerative and single-linkage clustering based diarization systems with a relative improvement of 40% and 68%, respectively. We then propose a complete-linkage speaker linking system to achieve attribution and demonstrate a 26% relative improvement in attribution error rate (AER) over the single-linkage speaker linking approach. Houman Ghaemmaghami, David Dean, Robbie Vogt, Sridha Sridharan |
ICASSP | 4 |
| 2012 | Self-calibration of wireless cameras with restricted degrees of freedom
Junbin Liu, Tim Wark, Ruan Lakemond, Sridha Sridharan |
Comput. Vis. Image Underst. | 4 |
| 2012 | Evaluation of image resolution and super-resolution on face recognition performance
Clinton Fookes, Frank Lin, Vinod Chandran, Sridha Sridharan |
J. Vis. Commun. Image Represent. | 4 |
| 2012 | In the Pursuit of Effective Affective Computing: The Relationship Between Features and RegistrationabstractFor facial expression recognition systems to be applicable in the real world, they need to be able to detect and track a previously unseen person's face and its facial movements accurately in realistic environments. A highly plausible solution involves performing a "dense" form of alignment, where 60-70 fiducial facial points are tracked with high accuracy. The problem is that, in practice, this type of dense alignment had so far been impossible to achieve in a generic sense, mainly due to poor reliability and robustness. Instead, many expression detection methods have opted for a "coarse" form of face alignment, followed by an application of a biologically inspired appearance descriptor such as the histogram of oriented gradients or Gabor magnitudes. Encouragingly, recent advances to a number of dense alignment algorithms have demonstrated both high reliability and accuracy for unseen subjects [e.g., constrained local models (CLMs)]. This begs the question: Aside from countering against illumination variation, what do these appearance descriptors do that standard pixel representations do not? In this paper, we show that, when close to perfect alignment is obtained, there is no real benefit in employing these different appearance-based representations (under consistent illumination conditions). In fact, when misalignment does occur, we show that these appearance descriptors do work well by encoding robustness to alignment error. For this work, we compared two popular methods for dense alignment-subject-dependent active appearance models versus subject-independent CLMs-on the task of action-unit detection. These comparisons were conducted through a battery of experiments across various publicly available data sets (i.e., CK+, Pain, M3, and GEMEP-FERA). We also report our performance in the recent 2011 Facial Expression Recognition and Analysis Challenge for the subject-independent task. Sien W. Chew, Patrick Lucey, Simon Lucey, Jason M. Saragih, Jeffrey F. Cohn, Iain A. Matthews, Sridha Sridharan |
IEEE Trans. Syst. Man Cybern. Part B | 7 |
| 2011 | Determining operational measures from multi-camera surveillance systems using soft biometricsabstractCCTV and surveillance networks are increasingly being used for operational as well as security tasks. One emerging area of technology that lends itself to operational analytics is soft biometrics. Soft biometrics can be used to describe a person and detect them throughout a sparse multi-camera network. This enables them to be used to perform tasks such as determining the time taken to get from point to point, and the paths taken through an environment by detecting and matching people across disjoint views. However, in a busy environment where there are 100's if not 1000's of people such as an airport, attempting to monitor everyone is highly unrealistic. In this paper we propose an average soft biometric, that can be used to identity people who look distinct, and are thus suitable for monitoring through a large, sparse camera network. We demonstrate how an average soft biometric can be used to identify unique people to calculate operational measures such as the time taken to travel from point to point. Simon Denman, Alina Bialkowski, Clinton Fookes, Sridha Sridharan |
AVSS | 4 |
| 2011 | Textures of optical flow for real-time anomaly detection in crowdsabstractAutomated visual surveillance of crowds is a rapidly growing area of research. In this paper we focus on motion representation for the purpose of abnormality detection in crowded scenes. We propose a novel visual representation called textures of optical flow. The proposed representation measures the uniformity of a flow field in order to detect anomalous objects such as bicycles, vehicles and skateboarders; and can be combined with spatial information to detect other forms of abnormality. We demonstrate that the proposed approach outperforms state-of-the-art anomaly detection algorithms on a large, publicly-available dataset. David Ryan, Simon Denman, Clinton Fookes, Sridha Sridharan |
AVSS | 4 |
| 2011 | 3D ellipsoid fitting for multi-view gait recognitionabstractGait recognition approaches continue to struggle with challenges including view-invariance, low-resolution data, robustness to unconstrained environments, and fluctuating gait patterns due to subjects carrying goods or wearing different clothes. Although computationally expensive, model based techniques offer promise over appearance based techniques for these challenges as they gather gait features and interpret gait dynamics in skeleton form. In this paper, we propose a fast 3D ellipsoidal-based gait recognition algorithm using a 3D voxel model derived from multi-view silhouette images. This approach directly solves the limitations of view dependency and self-occlusion in existing ellipse fitting model-based approaches. Voxel models are segmented into four components (left and right legs, above and below the knee), and ellipsoids are fitted to each region using eigenvalue decomposition. Features derived from the ellipsoid parameters are modeled using a Fourier representation to retain the temporal dynamic pattern for classification. We demonstrate the proposed approach using the CMU MoBo database and show that an improvement of 15-20% can be achieved over a 2D ellipse fitting baseline. Sabesan Sivipalan, Daniel Chen 0002, Simon Denman, Sridha Sridharan, Clinton Fookes |
AVSS | 4 |
| 2011 | Person-independent facial expression detection using Constrained Local ModelsabstractIn automatic facial expression detection, very accurate registration is desired which can be achieved via a deformable model approach where a dense mesh of 60-70 points on the face is used, such as an active appearance model (AAM). However, for applications where manually labeling frames is prohibitive, AAMs do not work well as they do not generalize well to unseen subjects. As such, a more coarse approach is taken for person-independent facial expression detection, where just a couple of key features (such as face and eyes) are tracked using a Viola-Jones type approach. The tracked image is normally post-processed to encode for shift and illumination invariance using a linear bank of filters. Recently, it was shown that this preprocessing step is of no benefit when close to ideal registration has been obtained. In this paper, we present a system based on the Constrained Local Model (CLM) method which is a generic or person-independent face alignment algorithm which gains high accuracy. We show these results against the LBP feature extraction on the CK+ and GEMEP-FERA datasets. Sien W. Chew, Patrick Lucey, Simon Lucey, Jason M. Saragih, Jeffrey F. Cohn, Sridha Sridharan |
FG | 6 |
| 2011 | Gait energy volumes and frontal gait recognition using depth imagesabstractGait energy images (GEIs) and its variants form the basis of many recent appearance-based gait recognition systems. The GEI combines good recognition performance with a simple implementation, though it suffers problems inherent to appearance-based approaches, such as being highly view dependent. In this paper, we extend the concept of the GEI to 3D, to create what we call the gait energy volume, or GEV. A basic GEV implementation is tested on the CMU MoBo database, showing improvements over both the GEI baseline and a fused multi-view GEI approach. We also demonstrate the efficacy of this approach on partial volume reconstructions created from frontal depth images, which can be more practically acquired, for example, in biometric portals implemented with stereo cameras, or other depth acquisition systems. Experiments on frontal depth images are evaluated on an in-house developed database captured using the Microsoft Kinect, and demonstrate the validity of the proposed approach. Sabesan Sivipalan, Daniel Chen 0002, Simon Denman, Sridha Sridharan, Clinton Fookes |
IJCB | 4 |
| 2011 | Fourier Active Appearance ModelsabstractGaining invariance to camera and illumination variations has been a well investigated topic in Active Appearance Model (AAM) fitting literature. The major problem lies in the inability of the appearance parameters of the AAM to generalize to unseen conditions. An attractive approach for gaining invariance is to fit an AAM to a multiple filter response (e.g. Gabor) representation of the input image. Naively applying this concept with a traditional AAM is computationally prohibitive, especially as the number of filter responses increase. In this paper, we present a computationally efficient AAM fitting algorithm based on the Lucas-Kanade (LK) algorithm posed in the Fourier domain that affords invariance to both expression and illumination. We refer to this as a Fourier AAM (FAAM), and show that this method gives substantial improvement in person specific AAM fitting performance over traditional AAM fitting methods. Rajitha Navarathna, Sridha Sridharan, Simon Lucey |
ICCV | 2 |
| 2011 | Feature-domain super-resolution for iris recognitionabstractUncooperative iris identification systems at a distance suffer from poor resolution of the captured iris images, which significantly degrades iris recognition performance. Super-resolution techniques have been employed to enhance the resolution of iris images and improve the recognition performance. However, all existing super-resolution approaches proposed for the iris biometric super-resolve pixel intensity values. This paper considers transferring super-resolution of iris images from the intensity domain to the feature domain. By directly super-resolving only the features essential for recognition, and by incorporating domain specific information from iris models, improved recognition performance compared to pixel domain super-resolution can be achieved. This is the first paper to investigate the possibility of feature-domain super-resolution for iris recognition, and experiments confirm the validity of the proposed approach. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Simon Denman |
ICIP | 3 |
| 2011 | Extending the Task of Diarization to Speaker AttributionabstractIn this paper we extend the concept of speaker annotation within a single-recording, or speaker diarization, to a collection wide approach we call speaker attribution.Accordingly, speaker attribution is the task of clustering expectantly homogenous intersession clusters obtained using diarization according to common cross-recording identities.The result of attribution is a collection of spoken audio across multiple recordings attributed to speaker identities.In this paper, an attribution system is proposed using mean-only MAP adaptation of a combined-gender UBM to model clusters from a perfect diarization system, as well as a JFA-based system with session variability compensation.The normalized cross-likelihood ratio is calculated for each pair of clusters to construct an attribution matrix and the complete linkage algorithm is employed to conduct clustering of the inter-session clusters.A matched cluster purity and coverage of 87.1% was obtained on the NIST 2008 SRE corpus. Houman Ghaemmaghami, David Dean, Robbie Vogt, Sridha Sridharan |
INTERSPEECH | 4 |
| 2011 | i-vector Based Speaker Recognition on Short UtterancesabstractRobust speaker verification on short utterances remains a key consideration when deploying automatic speaker recognition, as many real world applications often have access to only limited duration speech data. This paper explores how the recent technologies focused around total variability modeling behave when training and testing utterance lengths are reduced. Results are presented which provide a comparison of Joint Factor Analysis (JFA) and i-vector based systems including various compensation techniques; Within-Class Covariance Normalization (WCCN), LDA, Scatter Difference Nuisance Attribute Projection (SDNAP) and Gaussian Probabilistic Linear Discriminant Analysis (GPLDA). Speaker verification performance for utterances with as little as 2 sec of data taken from the NIST Speaker Recognition Evaluations are presented to provide a clearer picture of the current performance characteristics of these techniques in short utterance conditions. Ahilan Kanagasundaram, Robbie Vogt, David Dean, Sridha Sridharan, Michael Mason |
INTERSPEECH | 4 |
| 2011 | Can Audio-Visual Speech Recognition Outperform Acoustically Enhanced Speech Recognition in Automotive Environment?abstractThe use of visual features in the form of lip movements to improve the performance of acoustic speech recognition has been shown to work well, particularly in noisy acoustic conditions.However, whether this technique can outperform speech recognition incorporating well-known acoustic enhancement techniques, such as spectral subtraction, or multi-channel beamforming is not known.This is an important question to be answered especially in an automotive environment, for the design of an efficient human-vehicle computer interface.We perform a variety of speech recognition experiments on a challenging automotive speech dataset and results show that synchronous HMM-based audio-visual fusion can outperform traditional single as well as multi-channel acoustic speech enhancement techniques.We also show that further improvement in recognition performance can be obtained by fusing speech-enhanced audio with the visual modality, demonstrating the complementary nature of the two robust speech recognition approaches. Rajitha Navarathna, Tristan Kleinschmidt, David Dean, Sridha Sridharan, Patrick Lucey |
INTERSPEECH | 4 |
| 2011 | Cross Likelihood Ratio Based Speaker Clustering Using Eigenvoice ModelsabstractThis paper proposes the use of eigenvoice modeling techniques with the Cross Likelihood Ratio (CLR) as a criterion for speaker clustering within a speaker diarization system.The CLR has previously been shown to be a robust decision criterion for speaker clustering using Gaussian Mixture Models.Recently, eigenvoice modeling techniques have become increasingly popular, due to its ability to adequately represent a speaker based on sparse training data, as well as an improved capture of differences in speaker characteristics.This paper hence proposes that it would be beneficial to capitalize on the advantages of eigenvoice modeling in a CLR framework.Results obtained on the 2002 Rich Transcription (RT-02) Evaluation dataset show an improved clustering performance, resulting in a 35.1% relative improvement in the overall Diarization Error Rate (DER) compared to the baseline system. David Wang 0002, Robbie Vogt, Sridha Sridharan, David Dean |
INTERSPEECH | 3 |
| 2011 | Sparse Temporal Representations for Facial Expression Recognition
Sien W. Chew, Rajib Rana, Patrick Lucey, Simon Lucey, Sridha Sridharan |
PSIVT (2) | 5 |
| 2011 | The use of phase in complex spectrum subtraction for robust speech recognition
Tristan Kleinschmidt, Sridha Sridharan, Michael Mason |
Comput. Speech Lang. | 2 |
| 2011 | Clustered Blind Beamforming From Ad-Hoc Microphone ArraysabstractMicrophone arrays have been used in various applications to capture conversations, such as in meetings and teleconferences. In many cases, the microphone and likely source locations are known a priori, and calculating beamforming filters is therefore straightforward. In ad-hoc situations, however, when the microphones have not been systematically positioned, this information is not available and beamforming must be achieved blindly. In achieving this, a commonly neglected issue is whether it is optimal to use all of the available microphones, or only an advantageous subset of these. This paper commences by reviewing different approaches to blind beamforming, characterizing them by the way they estimate the signal propagation vector and the spatial coherence of noise in the absence of prior knowledge of microphone and speaker locations. Following this, a novel clustered approach to blind beamforming is motivated and developed. Without using any prior geometrical information, microphones are first grouped into localized clusters, which are then ranked according to their relative distance from a speaker. Beamforming is then performed using either the closest microphone cluster, or a weighted combination of clusters. The clustered algorithms are compared to the full set of microphones in experiments on a database recorded on different ad-hoc array geometries. These experiments evaluate the methods in terms of signal enhancement as well as performance on a large vocabulary speech recognition task. Ivan Himawan, Iain McCowan, Sridha Sridharan |
IEEE Trans. Speech Audio Process. | 3 |
| 2011 | The Delta-Phase Spectrum With Application to Voice Activity Detection and Speaker RecognitionabstractFor several reasons, the Fourier phase domain is less favored than the magnitude domain in signal processing and modeling of speech. To correctly analyze the phase, several factors must be considered and compensated, including the effect of the step size, windowing function and other processing parameters. Building on a review of these factors, this paper investigates a spectral representation based on the Instantaneous Frequency Deviation, but in which the step size between processing frames is used in calculating phase changes, rather than the traditional single sample interval. Reflecting these longer intervals, the term delta-phase spectrum is used to distinguish this from instantaneous derivatives. Experiments show that mel-frequency cepstral coefficients features derived from the delta-phase spectrum (termed Mel-Frequency delta-phase features) can produce broadly similar performance to equivalent magnitude domain features for both voice activity detection and speaker recognition tasks. Further, it is shown that the fusion of the magnitude and phase representations yields performance benefits over either in isolation. Iain McCowan, David Dean, Mitchell McLaren, Robbie Vogt, Sridha Sridharan |
IEEE Trans. Speech Audio Process. | 5 |
| 2011 | Discriminative Optimization of the Figure of Merit for Phonetic Spoken Term DetectionabstractThis paper proposes to improve spoken term detection (STD) accuracy by optimizing the figure of merit (FOM). In this paper, the index takes the form of a phonetic posterior-feature matrix. Accuracy is improved by formulating STD as a discriminative training problem and directly optimizing the FOM, through its use as an objective function to train a transformation of the index. The outcome of indexing is then a matrix of enhanced posterior-features that are directly tailored for the STD task. The technique is shown to improve the FOM by up to 13% on held-out data. Additional analysis explores the effect of the technique on phone recognition accuracy, examines the actual values of the learned transform, and demonstrates that using an extended training data set results in further improvement in the FOM. Roy Wallace, Brendan Baker, Robbie Vogt, Sridha Sridharan |
IEEE Trans. Speech Audio Process. | 4 |
| 2011 | Quality-Driven Super-Resolution for Less Constrained Iris Recognition at a Distance and on the MoveabstractLess constrained iris identification systems at a distance and on the move suffer from poor resolution and poor quality of the captured iris images, which significantly degrades iris recognition performance. This paper proposes a new signal-level fusion approach which incorporates a quality score into a reconstruction-based super-resolution process to generate a high-resolution iris image from a low-resolution and quality inconsistent video sequence of an eye. A novel approach for assessing the focus level of the iris image, which is invariant to lighting and oclusion conditions, is introduced. The focus score is combined with several other quality factors to perform the quality weighted super-resolution where the highest quality frames contribute the greatest amount of information to the resulting high-resolution images without introducing spurious high-frequency components. Experiments conducted on the Multiple Biometric Grand Challenge portal dataset show that our proposed approach outperforms the traditional best quality frame selection approach and other existing state-of-the-art signal-level and score-level fusion approaches for recognition of less constrained iris at a distance and on the move. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Simon Denman |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2011 | Automatically Detecting Pain in Video Through Facial Action UnitsabstractIn a clinical setting, pain is reported either through patient self-report or via an observer. Such measures are problematic as they are: 1) subjective, and 2) give no specific timing information. Coding pain as a series of facial action units (AUs) can avoid these issues as it can be used to gain an objective measure of pain on a frame-by-frame basis. Using video data from patients with shoulder injuries, in this paper, we describe an active appearance model (AAM)-based system that can automatically detect the frames in video in which a patient is in pain. This pain data set highlights the many challenges associated with spontaneous emotion detection, particularly that of expression and head movement due to the patient's reaction to pain. In this paper, we show that the AAM can deal with these movements and can achieve significant improvements in both the AU and pain detection performance compared to the current-state-of-the-art approaches which utilize similarity-normalized appearance features only. Patrick Lucey, Jeffrey F. Cohn, Iain A. Matthews, Simon Lucey, Sridha Sridharan, Jessica Howlett, Kenneth M. Prkachin |
IEEE Trans. Syst. Man Cybern. Part B | 5 |
| 2010 | Multi-Modal Object Tracking using Dynamic Performance MetricsabstractIntelligent surveillance systems typically use a single visual spectrum modality for their input. These systems work well in controlled conditions, but often fail when lighting is poor, or environmental effects such as shadows, dust or smoke are present. Thermal spectrum imagery is not as susceptible to environmental effects, however thermal imaging sensors are more sensitive to noise and they are only gray scale, making distinguishing between objects difficult. Several approaches to combining the visual and thermal modalities have been proposed, however they are limited by assuming that both modalities are perfuming equally well. When one modality fails, existing approaches are unable to detect the drop in performance and disregard the under performing modality. In this paper, a novel middle fusion approach for combining visual and thermal spectrum images for object tracking is proposed. Motion and object detection is performed on each modality and the object detection results for each modality are fused base on the current performance of each modality. Modality performance is determined by comparing the number of objects tracked by the system with the number detected by each mode, with a small allowance made for objects entering and exiting the scene. The tracking performance of the proposed fusion scheme is compared with performance of the visual and thermal modes individually, and a baseline middle fusion scheme. Improvement in tracking performance using the proposed fusion approach is demonstrated. The proposed approach is also shown to be able to detect the failure of an individual modality and disregard its results, ensuring performance is not degraded in such situations. Simon Denman, Clinton Fookes, Sridha Sridharan, David Ryan |
AVSS | 3 |
| 2010 | Crowd Counting Using Group Tracking and Local FeaturesabstractIn public venues, crowd size is a key indicator of crowd safety and stability. In this paper we propose a crowd counting algorithm that uses tracking and local features to count the number of people in each group as represented by a foreground blob segment, so that the total crowd estimate is the sum of the group sizes. Tracking is employed to improve the robustness of the estimate, by analysing the history of each group, including splitting and merging events. A simplified ground truth annotation strategy results in an approach with minimal setup requirements that is highly accurate. David Ryan, Simon Denman, Clinton Fookes, Sridha Sridharan |
AVSS | 4 |
| 2010 | Noise robust voice activity detection using normal probability testing and time-domain histogram analysisabstractThis paper presents a method of voice activity detection (VAD) suitable for high noise scenarios, based on the fusion of two complementary systems. The first system uses a proposed non-Gaussianity score (NGS) feature based on normal probability testing. The second system employs a histogram distance score (HDS) feature that detects changes in the signal through conducting a template-based similarity measure between adjacent frames. The decision outputs by the two systems are then merged using an open-by-reconstruction fusion stage. Accuracy of the proposed method was compared to several baseline VAD methods on a database created using real recordings of a variety of high-noise environments. Houman Ghaemmaghami, David Dean, Sridha Sridharan, Iain McCowan |
ICASSP | 3 |
| 2010 | Clustering of ad-hoc microphone arrays for robust blind beamformingabstractThis paper proposes a clustered approach for blind beamfoming from ad-hoc microphone arrays. In such arrangements, microphone placement is arbitrary and the speaker may be close to one, all or a subset of microphones at a given time. Practical issues with such a configuration mean that some microphones might be better discarded due to poor input signal to noise ratio (SNR) or undesirable spatial aliasing effects from large inter-element spacings when beamforming. Large inter-microphone spacings may also lead to inaccuracies in delay estimation during blind beamforming. In such situations, using a cluster of microphones (ie, a sub-array), closely located both to each other and to the desired speech source, may provide more robust enhancement than the full array. This paper proposes a method for blind clustering of microphones based on the magnitude square coherence function, and evaluates the method on a database recorded using various ad-hoc microphone arrangements. Ivan Himawan, Iain McCowan, Sridha Sridharan |
ICASSP | 3 |
| 2010 | Exploiting multiple feature sets in data-driven impostor dataset selection for speaker verificationabstractThis study assesses the recently proposed data-driven background dataset refinement technique for speaker verification using alternate SVM feature sets to the GMM supervector features for which it was originally designed. The performance improvements brought about in each trialled SVM configuration demonstrate the versatility of background dataset refinement. This work also extends on the originally proposed technique to exploit support vector coefficients as an impostor suitability metric in the data-driven selection process. Using support vector coefficients improved the performance of the refined datasets in the evaluation of unseen data. Further, attempts are made to exploit the differences in impostor example suitability measures from varying features spaces to provide added robustness. Mitchell McLaren, Brendan Baker, Robbie Vogt, Sridha Sridharan |
ICASSP | 4 |
| 2010 | Optimising Figure of Merit for phonetic spoken term detectionabstractThis paper introduces a novel technique to directly optimise the Figure of Merit (FOM) for phonetic spoken term detection. The FOM is a popular measure of STD accuracy, making it an ideal candidate for use as an objective function. A simple linear model is introduced to transform the phone log-posterior probabilities output by a phone classifier to produce enhanced log-posterior features that are more suitable for the STD task. Direct optimisation of the FOM is then performed by training the parameters of this model using a nonlinear gradient descent algorithm. Substantial FOM improvements of 11% relative are achieved on held-out evaluation data, demonstrating the generalisability of the approach. Roy Wallace, Robbie Vogt, Brendan Baker, Sridha Sridharan |
ICASSP | 4 |
| 2010 | The QUT-NOISE-TIMIT corpus for the evaluation of voice activity detection algorithmsabstractThe QUT-NOISE-TIMIT corpus consists of 600 hours of noisy speech sequences designed to enable a thorough evaluation of voice activity detection (VAD) algorithms across a wide variety of common background noise scenarios. In order to construct the final mixed-speech database, a collection of over 10 hours of background noise was conducted across 10 unique locations covering 5 common noise scenarios, to create the QUT-NOISE corpus. This background noise corpus was then mixed with speech events chosen from the TIMIT clean speech corpus over a wide variety of noise lengths, signal-to-noise ratios (SNRs) and active speech proportions to form the mixed-speech QUT-NOISE-TIMIT corpus. The evaluation of five baseline VAD systems on the QUT-NOISE-TIMIT corpus is conducted to validate the data and show that the variety of noise available will allow for better evaluation of VAD systems than existing approaches in the literature. David Dean, Sridha Sridharan, Robbie Vogt, Michael Mason |
INTERSPEECH | 2 |
| 2010 | Noise robust voice activity detection using features extracted from the time-domain autocorrelation functionabstractThis paper presents a method of voice activity detection (VAD) for high noise scenarios, using a noise robust voiced speech detection feature. The developed method is based on the fusion of two systems. The first system utilises the maximum peak of the normalised time-domain autocorrelation function (MaxPeak). The second zone system uses a novel combination of cross-correlation and zero-crossing rate of the normalised autocorrelation to approximate a measure of signal pitch and periodicity (CrossCorr) that is hypothesised to be noise robust. The score outputs by the two systems are then merged using weighted sum fusion to create the proposed autocorrelation zero-crossing rate (AZR) VAD. Accuracy of AZR was compared to state of the art and standardised VAD methods and was shown to outperform the best performing system with an average relative improvement of 24.8% in half-total error rate (HTER) on the QUT-NOISE-TIMIT database created using real recordings from high-noise environments. Houman Ghaemmaghami, Brendan Baker, Robbie Vogt, Sridha Sridharan |
INTERSPEECH | 4 |
| 2010 | Bayes factor based speaker segmentation for speaker diarizationabstractThis paper proposes the use of the Bayes Factor as a distance metric for speaker segmentation within a speaker diarization system. The proposed approach uses a pair of constant sized, sliding windows to compute the value of the Bayes Factor between the adjacent windows over the entire audio. Results obtained on the 2002 Rich Transcription Evaluation dataset show an improved segmentation performance compared to previous approaches reported in literature using the Generalized Likelihood Ratio. When applied in a speaker diarization system, this approach results in a 5.1% relative improvement in the overall Diarization Error Rate compared to the baseline. David Wang 0002, Robbie Vogt, Sridha Sridharan |
INTERSPEECH | 3 |
| 2010 | Dynamic visual features for audio-visual speaker verification
David Dean, Sridha Sridharan |
Comput. Speech Lang. | 2 |
| 2010 | Data-Driven Background Dataset Selection for SVM-Based Speaker VerificationabstractThe recently proposed data-driven background dataset refinement technique provides a means of selecting an informative background for support vector machine (SVM)-based speaker verification systems. This paper investigates the characteristics of the impostor examples in such highly informative background datasets. Data-driven dataset refinement individually evaluates the suitability of candidate impostor examples for the SVM background prior to selecting the highest-ranking examples as arefinedbackground dataset. Further, the characteristics of the refined dataset were analyzed to investigate the desired traits of an informative SVM background. The most informative examples of the refined dataset were found to consist of large amounts of active speech and distinctive language characteristics. The data-driven refinement technique was shown to filter the set of candidate impostor examples to produce a more disperse representation of the impostor population in the SVM kernel space, thereby reducing the number of redundant and less-informative examples in the background dataset. Furthermore, data-driven refinement was shown to provide performance gains when applied to the difficult task of refining a small candidate dataset that was mismatched to the evaluation conditions. Mitchell McLaren, Robbie Vogt, Brendan Baker, Sridha Sridharan |
IEEE Trans. Speech Audio Process. | 4 |
| 2010 | Making Confident Speaker Verification Decisions With Minimal SpeechabstractProposed is an approach to estimating confidence measures on the verification score produced by a Gaussian mixture model (GMM)-based automatic speaker verification system with applications to drastically reducing the typical data requirements for producing a confident verification decision. The confidence measures are based on estimating the distribution of the observed frame scores. The confidence estimation procedure is also extended to produce robust results with very limited and highly correlated frame scores as well as in the presence of score normalization. The proposed Early Verification Decision method utilizes the developed confidence measures in a sequential hypothesis testing framework, demonstrating that as little as 2–10 s of speech on average was able to produce verification results approaching that of using an average of over 100 s of speech on the 2005 NIST SRE protocol. Robbie Vogt, Sridha Sridharan, Michael Mason |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | A Comparison of Session Variability Compensation Approaches for Speaker VerificationabstractThis paper compares two of the leading techniques for session variability compensation in the context of support vector machine (SVM) speaker verification using Gaussian mixture model (GMM) mean supervectors: joint factor analysis (JFA) modeling and nuisance attribute projection (NAP). Motivation for this comparison comes from the distinctly different domains in which these techniques are employed-the probabilistic GMM domain versus the discriminative SVM kernel. A theoretical analysis is given comparing the JFA and NAP approaches to variability compensation. The role of speaker factors in the factor analysis model is also contrasted against the scatter difference NAP objective of retaining speaker information in the SVM kernel space. These methods for retaining speaker variation are found to provide improved verification performance over the removal of channel effects alone. Overall, experimental results on the NIST 2006 and 2008 SRE corpora demonstrate the effectiveness of both JFA and NAP techniques for reducing the effects of variability. However, the overheads associated with the implementation of JFA may make NAP a more attractive technique due to its simple yet effective approach to variability compensation. Mitchell McLaren, Robbie Vogt, Brendan Baker, Sridha Sridharan |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2009 | Dynamic Performance Measures for Object Tracking SystemsabstractPerformance evaluation of object tracking systems is typically performed after the data has been processed, by comparing tracking results to ground truth. Whilst this approach is fine when performing offline testing, it does not allow for real-time analysis of the systems performance, which may be of use for live systems to either automatically tune the system or report reliability. In this paper, we propose three metrics that can be used to dynamically asses the performance of an object tracking system. Outputs and results from various stages in the tracking system are used to obtain measures that indicate the performance of motion segmentation, object detection and object matching. The proposed dynamic metrics are shown to accurately indicate tracking errors when visually comparing metric results to tracking output, and are shown to display similar trends to the ETISEO metrics when comparing different tracking configurations. Simon Denman, Clinton Fookes, Sridha Sridharan, Ruan Lakemond |
AVSS | 3 |
| 2009 | Affine Adaptation of Local Image Features Using the Hessian MatrixabstractLocal feature detectors that make use of derivative based saliency functions to locate points of interest typically require adaptation processes after initial detection in order to achieve scale and affine covariance. Affine adaptation methods have previously been proposed that make use of the second moment matrix to iteratively estimate the affine shape of local image regions. This paper shows that it is possible to use the Hessian matrix to estimate local affine shape in a similar fashion to the second moment matrix. The Hessian matrix requires significantly less computation effort to compute than the second moment matrix, allowing more efficient affine adaptation. It may also be more convenient to use the Hessian matrix, for example, when the Determinant of Hessian detector is used. Experimental evaluation shows that the Hessian matrix is very effective in increasing the efficiency of blob detectors such as the Determinant of Hessian detector, but less effective in combination with the Harris corner detector. Ruan Lakemond, Clinton Fookes, Sridha Sridharan |
AVSS | 3 |
| 2009 | The Australian English Speech Corpus for In-Car Speech processingabstractThe Australian in-car speech corpus is a multi-channel recording of a series of prompts from an in-car navigation task collected over a range of speakers in a variety of driving conditions. Its purpose is to provide a significant resource of speech data appropriate for investigating speech processing needs in the adverse environment of a car. Utterances spoken by 50 speakers were collected in seven different driving conditions, providing the foundation for investigation into noisy, speaker-independent speech processing. Speech recognition experiments are performed to validate the data, to provide baseline results for in-car speech recognition research, and to show that this data can improve speech recognition performance under adverse in-car conditions for Australian English when adapting from American English acoustic models. Tristan Kleinschmidt, Michael Mason, Eddie Wong, Sridha Sridharan |
ICASSP | 4 |
| 2009 | Improved SVM speaker verification through data-driven background dataset collectionabstractThe problem of background dataset selection in SVM-based speaker verification is addressed through the proposal of a new data-driven selection technique. Based on support vector selection, the proposed approach introduces a method to individually assess the suitability of each candidate impostor example for use in the background dataset. The technique can then produce a refined background dataset by selecting only the most informative impostor examples. Improvements of 13% in min. DCF and 10% in EER were found on the SRE 2006 development corpus when using the proposed method over the best heuristically chosen set. The technique was also shown to generalise to the unseen NIST 2008 SRE corpus. Mitchell McLaren, Brendan Baker, Robbie Vogt, Sridha Sridharan |
ICASSP | 4 |
| 2009 | Spoken term detection using fast phonetic decodingabstractWhile spoken term detection (STD) systems based on word indices provide good accuracy, there are several practical applications where it is infeasible or too costly to employ an LVCSR engine. An STD system is presented, which is designed to incorporate a fast phonetic decoding front-end and be robust to decoding errors whilst still allowing for rapid search speeds. This goal is achieved through monophone open-loop decoding coupled with fast hierarchical phone lattice search. Results demonstrate that an STD system that is designed with the constraint of a fast and simple phonetic decoding front-end requires a compromise to be made between search speed and search accuracy. Roy Wallace, Robbie Vogt, Sridha Sridharan |
ICASSP | 3 |
| 2009 | Least-squares congealing for large numbers of imagesabstractIn this paper we pursue the task of aligning an ensemble of images in an unsupervised manner. This task has been commonly referred to as “congealing” in literature. A form of congealing, using a least-squares criteria, has been recently demonstrated to have desirable properties over conventional congealing. Least-squares congealing can be viewed as an extension of the Lucas & Kanade (LK) image alignment algorithm. It is well understood that the alignment performance for the LK algorithm, when aligning a single image with another, is theoretically and empirically equivalent for additive and compositional warps. In this paper we: (i) demonstrate that this equivalence does not hold for the extended case of congealing, (ii) characterize the inherent drawbacks associated with least-squares congealing when dealing with large numbers of images, and (iii) propose a novel method for circumventing these limitations through the application of an inverse-compositional strategy that maintains the attractive properties of the original method while being able to handle very large numbers of images. Mark Cox, Sridha Sridharan, Simon Lucey, Jeffrey F. Cohn |
ICCV | 2 |
| 2009 | Improved GMM-based speaker verification using SVM-driven impostor dataset selectionabstractThe problem of impostor dataset selection for GMM-based speaker verification is addressed through the recently proposed data-driven background dataset refinement technique. The SVM-based refinement technique selects from a candidate impostor dataset those examples that are most frequently selected as support vectors when training a set of SVMs on a development corpus. This study demonstrates the versatility of dataset refinement in the task of selecting suitable impostor datasets for use in GMM-based speaker verification. The use of refined Z- and T-norm datasets provided performance gains of 15% in EER in the NIST 2006 SRE over the use of heuristically selected datasets. The refined datasets were shown to generalise well to the unseen data of the NIST 2008 SRE. Mitchell McLaren, Robbie Vogt, Brendan Baker, Sridha Sridharan |
INTERSPEECH | 4 |
| 2009 | Within-session variability modelling for factor analysis speaker verificationabstractThis work presents an extended Joint Factor Analysis model including explicit modelling of unwanted within-session variability. The goals of the proposed extended JFA model are to improve verification performance with short utterances by compensating for the effects of limited or imbalanced phonetic coverage, and to produce a flexible JFA model that is effective over a wide range of utterance lengths without adjusting model parameters such as retraining session subspaces. Experimental results on the 2006 NIST SRE corpus demonstrate the flexibility of the proposed model by providing competitive results over a wide range of utterance lengths without retraining and also yielding modest improvements in a number of conditions over current state-of-the-art. Robbie Vogt, Jason W. Pelecanos, Nicolas Scheffer, Sachin S. Kajarekar, Sridha Sridharan |
INTERSPEECH | 5 |
| 2009 | Efficient constrained local model fitting for non-rigid face alignment
Simon Lucey, Yang Wang 0001, Mark Cox, Sridha Sridharan, Jeffrey F. Cohn |
Image Vis. Comput. | 4 |
| 2008 | Least squares congealing for unsupervised alignment of imagesabstractIn this paper, we present an approach we refer to as "least squares congealing" which provides a solution to the problem of aligning an ensemble of images in an unsupervised manner. Our approach circumvents many of the limitations existing in the canonical "congealing" algorithm. Specifically, we present an algorithm that:- (i) is able to simultaneously, rather than sequentially, estimate warp parameter updates, (ii) exhibits fast convergence and (iii) requires no pre-defined step size. We present alignment results which show an improvement in performance for the removal of unwanted spatial variation when compared with the related work of Learned-Miller on two datasets, the MNIST hand written digit database and the MultiPIE face database. Mark Cox, Sridha Sridharan, Simon Lucey, Jeffrey F. Cohn |
CVPR | 2 |
| 2008 | Dealing with uncertainty in microphone placement in a microphone array speech recognition systemabstractThis paper investigates robustness to uncertain microphone placements in an array beamformer front-end to a speech recognition system. There are two general approaches to handling the placement uncertainty: using the approximately known geometry in a robust beamforming technique, or using techniques that require no prior knowledge of geometry. Experiments in this paper compare the robustness of different techniques for both of these approaches in terms of speech recognition accuracy. To benefit from existing microphone array speech recognition data corpora for experimentation, microphone placement uncertainty is simulated by introducing random perturbations in the assumed geometry. Experimental results show that robust beamforming yields stable performance to a certain degree of placement error, but thereafter techniques such as automatic calibration are beneficial. Ivan Himawan, Sridha Sridharan, Iain McCowan |
ICASSP | 2 |
| 2008 | Cascading appearance-based features for visual speaker verificationabstractThe cascading appearance-based (CAB) feature extraction technique has established itself as the state of the art in extracting dynamic visual speech features for speech recognition. In this paper, we will focus on investigating the effectiveness of this technique for the related speaker verification application. By investigating the speaker verification ability of each stage of the cascade we will demonstrate that the same steps taken to reduce static speaker and environmental information for the speech recognition application also provide similar improvements for speaker recognition. These results suggest that visual speaker recognition can improve considerable when conducted solely through a consideration of the dynamic speech information rather than the static appearance of the speaker's mouth region. David Dean, Sridha Sridharan, Patrick Lucey |
INTERSPEECH | 2 |
| 2008 | Continuous pose-invariant lipreadingabstractIn audio-visual automatic speech recognition (AVASR), no research to date has been conducted into the problem of recognising visual speech whilst the speaker is moving their head. In this paper, we extend our current system to deal with this task, which we entitle continuous pose-invariant lipreading. By developing an AVASR system which can deal with such a scenario, we believe we are making the system effectively "real-world" as it requires little cooperation from the user and as such can be used in a host of realistic applications (e.g. mobile phones, in-vehicles etc.). In this proof of concept paper, we show via our experiments on the CUAVE database, that recognising visual speech whilst a speaker is moving their head during the utterance is feasible. Patrick Lucey, Sridha Sridharan, David Dean |
INTERSPEECH | 2 |
| 2008 | Factor analysis subspace estimation for speaker verification with short utterancesabstractTraining the speaker and session subspaces is an integral problem in developing a joint factor analysis GMM speaker verification system. This work investigates and compares several alternative procedures for this task with a particular focus on training and testing with short utterances. Experiments show that better performance can be obtained when an independent rather than simultaneous optimisation of the two core variability subspaces is used. It is additionally shown that for verification trials on short utterances it is important for the session subspace to be trained with matched length utterances. Conversely, the speaker transform should always be trained with as much data as possible. Index Terms: speaker verification, factor analysis, probabilistic PCA. Robbie Vogt, Brendan Baker, Sridha Sridharan |
INTERSPEECH | 3 |
| 2008 | Making confident speaker verification decisions with minimal speechabstractDrastic reductions in the typical data requirements for produc-ing confident decisions in an automatic speaker verification sys-tem are demonstrated through the application of a novel ap-proach of score confidence interval estimation. The confidence estimation procedure is also extended to produce robust re-sults with very limited and highly correlated frame scores. The early verification decision method evaluated on the 2005 NIST SRE protocol demonstrates that an average of 2–10 seconds of speech is sufficient to produce verification results approaching those achieved previously using an average of over 100 seconds of speech. Index Terms: automatic speaker verification, confidence mea-sures, verification decision confidence. Robbie Vogt, Sridha Sridharan, Michael Mason |
INTERSPEECH | 2 |
| 2008 | Comparing object alignment algorithms with appearance variation: Forward-additive vs inverse-compositionabstractA common problem that affects object alignment algorithms is when they have to deal with objects with unseen intra-class appearance variation. Several variants based on gradient-decent algorithms, such as the Lucas-Kanade (or forward-additive) and inverse-compositional algorithms, have been proposed to deal with this issue by solving for both alignment and appearance simultaneously. In [1], Baker and Matthews showed that without appearance variation, the inverse-compositional (IC) algorithm was theoretically and empirically equivalent to the forward-additive (FA) algorithm, whilst achieving significant improvement in computational efficiency. With appearance variation, it would be intuitive that a similar benefit of the IC algorithm would be experienced over the FA counterpart. However, to date no such comparison has been performed. In this paper we remedy this situation by performing such a comparison. In this comparison we show that the two algorithms are not equivalent due to the inclusion of the appearance variation parameters. Through a number of experiments on the MultiPIE face database, we show that we can gain greater refinement using the FA algorithm due to it being a truer solution than the IC approach. Patrick Lucey, Simon Lucey, Mark Cox, Sridha Sridharan, Jeffrey F. Cohn |
MMSP | 4 |
| 2008 | Explicit modelling of session variability for speaker verification
Robbie Vogt, Sridha Sridharan |
Comput. Speech Lang. | 2 |
| 2008 | 3D face verification using a free-parts approach
Chris McCool, Vinod Chandran, Sridha Sridharan, Clinton Fookes |
Pattern Recognit. Lett. | 3 |
| 2007 | Fused HMM-adaptation of multi-stream HMMs for audio-visual speech recognitionabstractA technique known as fused hidden Markov models (FHMMs) was recently proposed as an alternative multi-stream modelling technique for audio-visual speaker recognition. In this paper we show that for audio-visual speech recognition (AVSR), FHMMs can be adopted as a novel method of training synchronous MSHMMs. MSHMMs, as proposed by several authors for use in AVSR, are jointly trained on both the audio and visual modalities. In contrast our proposed FHMM adaptation method can be used to adapt the multi-stream models from single-stream audio HMMs, and in the process, better model the video speech in the final model when compared to jointly-trained MSHMMs. By experiments conducted on the XM2VTS database we show that the improved video performance of the FHMM-adapted MSHMMs results in an improvement in AVSR performance over jointly-trained MSHMMs at all levels of audio noise, and provide significant advantage in high noise environments. David Dean, Patrick Lucey, Sridha Sridharan, Tim Wark |
INTERSPEECH | 3 |
| 2007 | A unified approach to multi-pose audio-visual ASRabstractThe vast majority of studies in the field of audio-visual automatic speech recognition (AVASR) assumes frontal images of a speaker's face, but this cannot always be guaranteed in practice. Hence our recent research efforts have concentrated on extracting visual speech information from non-frontal faces, in particular the profile view. The introduction of additional views to an AVASR system increases the complexity of the system, as it has to deal with the different visual features associated with the various views. In this paper, we propose the use of linear regression to find a transformation matrix based on synchronous frontal and profile visual speech data, which is used to normalize the visual speech in each viewpoint into a single uniform view. In our experiments for the task of multi-speaker lipreading, we show that this "pose-invariant" technique reduces train/test mismatch between visual speech features of different views, and is of particular benefit when there is more training data for one viewpoint over another (e.g. frontal over profile). Patrick Lucey, Gerasimos Potamianos, Sridha Sridharan |
INTERSPEECH | 3 |
| 2007 | A comparison of session variability compensation techniques for SVM-based speaker recognitionabstractThis paper compares two of the leading techniques for session variability compensation in the context of GMM mean supervector SVM classifiers for speaker recognition: inter-session variability modelling and nuisance attribute projection.The former is incorporated in the GMM model training while the latter is employed as a modified SVM kernel.Results on both the NIST 2005 and 2006 corpora demonstrate the effectiveness of both techniques for reducing the effects of session variation.Further, system-and score-level fusion experiments show that the combination of the two methods provides improved performance. Mitchell McLaren, Robbie Vogt, Brendan Baker, Sridha Sridharan |
INTERSPEECH | 4 |
| 2007 | A phonetic search approach to the 2006 NIST spoken term detection evaluationabstractThis paper details the submission from the Speech and Audio Research Lab of Queensland University of Technology (QUT) to the inaugural 2006 NIST Spoken Term Detection Evaluation. The task involved accurately locating the occurrences of a specified list of English terms in a given corpus of broadcast news and conversational telephone speech. The QUT system uses phonetic decoding and Dynamic Match Lattice Spotting to rapidly locate search terms, combined with a neural network-based verification stage. The use of phonetic search means the system is open vocabulary and performs usefully (Actual Term-Weighted Value of 0.23) whilst avoiding the cost of a large vocabulary speech recognition engine. Roy Wallace, Robbie Vogt, Sridha Sridharan |
INTERSPEECH | 3 |
| 2007 | An adaptive optical flow technique for person tracking systems
Simon Denman, Vinod Chandran, Sridha Sridharan |
Pattern Recognit. Lett. | 3 |
| 2007 | Rapid Yet Accurate Speech Indexing Using Dynamic Match Lattice SpottingabstractThe support for typically out-of-vocabulary query terms such as names, acronyms, and foreign words is an important requirement of many speech indexing applications. However, to date many unrestricted vocabulary indexing systems have struggled to provide a balance between good detection rate and fast query speeds. This paper presents a fast and accurate unrestricted vocabulary speech indexing technique named Dynamic Match Lattice Spotting (DMLS). The proposed method augments the conventional lattice spotting technique with dynamic sequence matching, together with a number of other novel algorithmic enhancements, to obtain a system that is capable of searching hours of speech in seconds while maintaining excellent detection performance Kishan Thambiratnam, Sridha Sridharan |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | Multiscale Representation for 3-D Face RecognitionabstractThe eigenfaces algorithm has long been a mainstay in the field of face recognition due to the high dimensionality of face images. While providing minimal reconstruction error, the eigenface-based transform space de-emphasizes high-frequency information, effectively reducing the information available for classification. Methods such as linear discriminant analysis (also known as fisherfaces) allow the construction of subspaces which preserve the discriminatory information. In this article, multiscale techniques are used to partition the information contained in the frequency domain prior to dimensionality reduction. In this manner, it is possible to increase the information available for classification and, hence, increase the discriminative performance of both eigenfaces and fisherfaces techniques. Motivated by biological systems, Gabor filters are a natural choice for such a partitioning scheme. However, a comprehensive filter bank will dramatically increase the already high dimensionality of extracted features. In this article, a new method for intelligently reducing the dimensionality of Gabor features is presented. The face recognition grand challenge dataset of 3-D face images is used to examine the performance of Gabor filter banks for face recognition and to compare them against other multiscale partitioning methods such as the discrete wavelet transform and the discrete cosine transform. Jamie Cook, Vinod Chandran, Sridha Sridharan |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2006 | On the Performance and Use of Speaker Recognition Systems for SurveillanceabstractWe model the performance of a speaker recognition system used for surveillance to prioritize a large number of candidate speakers in search of a single target speaker. It is assumed that the system operates by ordering all speakers in order from best match to worst match, with the goal of having the true speaker sample positioned as high as possible on the list. Some performance measures for prioritization systems are given and are applied to a real speaker recognition system. An analytic expression for the probability density function of the true speaker's position on the list is found, subject to basic assumptions concerning the distribution of true speaker and false speaker scores. A comparison is made to the performance of a system which is operating only by making verification decisions, and it is shown that making soft decisions results in significantly better surveillance performance. Peter J. Barger, Sridha Sridharan |
AVSS | 2 |
| 2006 | Combined 2D/3D Face Recognition Using Log-Gabor TemplatesabstractThe addition of Three Dimensional (3D) data has the potential to greatly improve the accuracy of Face Recognition Technologies by providing complementary information. In this paper a new method combining intensity and range images and providing insensitivity to expression variation based on Log-Gabor Templates is presented. By breaking a single image into 75 semi-independent observations the reliance of the algorithm upon any particular part of the face is relaxed allowing robustness in the presence of occulusions, distortions and facial expressions. Also presented is a new distance measure based on the Mahalanobis Cosine metric which has desirable discriminatory characteristics in both the 2D and 3D domains. Using the 3D database collected by University of Notre Dame for the Face Recognition Grand Challenge (FRGC), benchmarking results are presented demonstrating the performance of the proposed methods. Jamie Cook, Chris McCool, Vinod Chandran, Sridha Sridharan |
AVSS | 4 |
| 2006 | A Multi-Class Tracker Using a Scalable Condensation FilterabstractTracking systems are typically targeted towards tracking a single class of object. In many real world situations, and in the ETISEO evaluation, it is advantageous to be able to track multiple classes of objects. In this paper we describe the adaptation of a single class tracking system to a multi-class tracking system, and describe a modified version of the condensation filter that can be used to track all objects, of all classes. We show that by using simple targeted detectors, we can achieve accurate tracking and can accurately distinguish between classes. Simon Denman, Vinod Chandran, Sridha Sridharan, Clinton Fookes |
AVSS | 3 |
| 2006 | Multi-view Intelligent Vehicle Surveillance SystemabstractThis paper presents a multi-view intelligent surveillance system used for the automatic tracking and monitoring of vehicles in a short-term parking lane. The system has the ability to track multiple vehicles in real-time across four cameras monitoring the area using a combination of both motion detection and optical flow modules. Automated alerts of events such as parking time violations, breaching of restricted areas or improper directional flow of traffic can be generated and communicated to attending security personnel. Results are shown using surveillance data captured from a real multi-camera network to illustrate the robust and real-time performance of the system. Simon Denman, Clinton Fookes, Jamie Cook, Chris Davoren, Anthony Mamic, Graeme Farquharson, Daniel Chen 0002, Brenden Chen, Sridha Sridharan |
AVSS | 9 |
| 2006 | The Role of Motion Models in Super-Resolving Surveillance Video for Face RecognitionabstractAlthough the use of super-resolution techniques has demonstrated the ability to improve face recognition accuracy when compared to traditional upsampling techniques, they are difficult to implement for real-time use due to their complexity and high computational demand. As a large portion of processing time is dedicated to registering the lowresolution images, many have adopted global motion models in order to improve efficiency. The drawback of such global models is that they can not accommodate for complex local motions, such as multiple objects moving independently across and static or dynamic background as frequently occurs in a surveillance environment. Local methods like optical flow can compensate for these situations, although it is achieved at the expense of computation time. In this paper, experiments have been carried out to investigate how motion models of different super-resolution reconstruction algorithms affect reconstruction error and face recognition rates in a surveillance environment. Results show that lower reconstruction error doesn't necessarily imply better recognition rates and the use of local motion models yields better recognition rates than global motion models. Frank Lin, Clinton Fookes, Vinod Chandran, Sridha Sridharan |
AVSS | 4 |
| 2006 | Human Face Reconstruction Using Bayesian Deformable ModelsabstractThis paper presents a Bayesian framework for 3D facial reconstruction. The framework iteratively deforms a generic face mesh to fit a set of range points representing a face. The generic mesh is generated from the extensive FRGC database of face images. The deformation process is conducted within a Bayesian framework and is driven by a Markov Chain Monte Carlo (MCMC) sampler which uses information from the likelihood and prior distributions of the generic face mesh. The paper presents results on the construction of a generic face model, the deformation framework and fitting results to both synthetic and real data. The results verify the effectiveness of the proposed technique, accurately deforming a generic face mesh to captured 3D data points of human faces. George Mamic, Clinton Fookes, Sridha Sridharan |
AVSS | 3 |
| 2006 | Feature Modelling of PCA Difference Vectors for 2D and 3D Face RecognitionabstractThis paper examines the the effectiveness of feature modelling to conduct 2D and 3D face recognition. In particular, PCA difference vectors are modelled using Gaussian Mixture Models (GMMs) which describe Intra-Personal (IP) and Extra-Personal (EP) variations. Two classifiers, an IP and IPEP classifier, are formed using these GMMs and their performance is compared to that of the Mahalanobis cosine metric (MahCosine). The best results for the 2D and 3D face modalities are obtained with the IP and IPEP classifiers respectively. The multi-modal fusion of these two systems provided consistent performance improvement across the FRGC database v2.0. Chris McCool, Jamie Cook, Vinod Chandran, Sridha Sridharan |
AVSS | 4 |
| 2006 | Experiments in Session Variability Modelling for Speaker VerificationabstractPresented is an approach to modelling session variability for GMM-based text-independent speaker verification incorporating a constrained session variability component in both the training and testing procedures. The proposed technique reduces the data labelling requirements and removes discrete categorisation needed by previous techniques and provides superior performance. Experiments on Mixer conversational telephony data show improvements of as much as 46% in equal error rate over a baseline system. In this paper the algorithm used for the enrollment procedure is described in detail. Results are also presented investigating the response of the technique to short test utterances and varying session subspace dimension Robbie Vogt, Sridha Sridharan |
ICASSP (1) | 2 |
| 2006 | What Is the Average Human Face?
George Mamic, Clinton Fookes, Sridha Sridharan |
PSIVT | 3 |
| 2006 | A syllable-scale framework for language identification
Terrence Martin, Brendan Baker, Eddie Wong, Sridha Sridharan |
Comput. Speech Lang. | 4 |
| 2006 | Gaze tracking for region of interest coding in JPEG 2000
Anthony N. Nguyen, Vinod Chandran, Sridha Sridharan |
Signal Process. Image Commun. | 3 |
| 2005 | Cross-language Acoustic Model Refinement for the Indonesian LanguageabstractPorting ASR capabilities to many languages is hindered by a lack of transcribed acoustic data. Cross-language adaptation techniques seek to address this problem by substituting models trained in resource-rich source languages to recognise speech in resource-poor target languages. The differences in coarticulatory effects between the source and target languages, together with unwanted pronunciation and channel variation, result in recognition rates that are typically much worse then those achieved by well trained monolingual systems. We present a technique which makes more effective use of limited adaptation data by structuring the state distributions to suit the coarticulatory occurrences in the target language. Additionally, the proposed technique provides a more suitable method for synthesising unseen contexts. Evaluation of this technique is presented for a word recognition task using English and Spanish source language acoustic models trained using Switchboard and CallHome databases, respectively. Using 25 minutes of Indonesian speech for target language adaptation data, this technique achieved absolute improvements of 3.69% and 6.31% for English and Spanish sources, respectively, when compared to traditional adaptation techniques. Using 90 minutes of adaptation data, absolute improvements of 3.22% and 3.07% were achieved. Terrence Martin, Sridha Sridharan |
ICASSP (1) | 2 |
| 2005 | Dynamic Match Phone-Lattice Searches For Very Fast And Accurate Unrestricted Vocabulary Keyword SpottingabstractThe ability to search for typically out-of-vocabulary terms such as names, acronyms and foreign words is a requirement of many audio indexing applications. To date, such applications have employed unrestricted vocabulary keyword spotting approaches that unfortunately suffer from poor miss rates or slow query speeds. This paper proposes a very fast and accurate keyword spotting approach named dynamic match phone-lattice keyword spotting. Reported experiments on conversational telephone speech and microphone speech demonstrate that the proposed method dramatically outperforms conventional methods and is capable of searching at speeds in excess of 300 times real-time while maintaining low miss rate performance. Kishan Thambiratnam, Sridha Sridharan |
ICASSP (1) | 2 |
| 2005 | Gaussian mixture modelling of broad phonetic and syllabic events for text-independent speaker verificationabstractThis paper examines the usefulness of a multilingual broad syllable-based framework for text-independent speaker verification.Syllabic segmentation is used in order to obtain a convenient unit for constrained and more detailed model generation.Gaussian mixture models are chosen as a suitable modelling paradigm for initial testing of the framework.Promising results are presented for the NIST 2003 speaker recognition evaluation corpus.The syllable-based modelling technique is shown to outperform a state-of-the-art baseline GMM system.A simple selective reduction of the syllable set is also shown to give further improvement in performance.Overall, the syllable based framework presents itself as valid alternative to text-constrained speaker verification systems, with the advantage of being multilingual.The framework allows for future testing of alternative modelling paradigms, feature sets and qualitative analysis. Brendan Baker, Robbie Vogt, Sridha Sridharan |
INTERSPEECH | 3 |
| 2005 | Data-driven clustering for blind feature mapping in speaker verificationabstractHandset and channel mismatch degrades the performance of automatic speaker recognition systems significantly. This paper enhances the feature mapping technique by proposing an iterative clustering approach to context model generation which offers an improvement in the performance of feature mapping trained on labelled data and offers the potential to train feature mapping in the absence of correctly labelled background data. The performance of the clustered feature mapping models is demonstrated on an expanded version of the NIST 2003 Extended Data Task (EDT) protocol. Michael Mason, Robbie Vogt, Brendan Baker, Sridha Sridharan |
INTERSPEECH | 4 |
| 2005 | Modelling session variability in text-independent speaker verificationabstractPresented is an approach to modelling session variability for GMM-based text-independent speaker verification incorporating a constrained session variability component in both the training and testing procedures. The proposed technique reduces the data labelling requirements and removes discrete categorisation needed by techniques such as feature mapping and H-Norm, while providing superior performance. Experiments on Switchboard-II conversational telephony data show improvements of as much as 48% in detection cost with a single training utterance and 68% with multiple training utterances over a baseline system. Robbie Vogt, Brendan Baker, Sridha Sridharan |
INTERSPEECH | 3 |
| 2005 | Texture for Script IdentificationabstractThe problem of determining the script and language of a document image has a number of important applications in the field of document analysis, such as indexing and sorting of large collections of such images, or as a precursor to optical character recognition (OCR). In this paper, we investigate the use of texture as a tool for determining the script of a document image, based on the observation that text has a distinct visual texture. An experimental evaluation of a number of commonly used texture features is conducted on a newly created script database, providing a qualitative measure of which features are most appropriate for this task. Strategies for improving classification results in situations with limited training data and multiple font types are also proposed. Andrew Busch, Wageeh W. Boles, Sridha Sridharan |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2005 | Integration strategies for audio-visual speech processing: applied to text-dependent speaker recognitionabstractIn this paper, an in-depth analysis is undertaken into effective strategies for integrating the audio-visual speech modalities with respect to two major questions. Firstly, at what level should integration occur? Secondly, given a level of integration how should this integration be implemented? Our work is based around the well-known hidden Markov model (HMM) classifier framework for modeling speech. A novel framework for modeling the mismatch between train and test observation sets is proposed, so as to provide effective classifier combination performance between the acoustic and visual HMM classifiers. From this framework, it can be shown that strategies for combining independent classifiers, such as the weighted product or sum rules, naturally emerge depending on the influence of the mismatch. Based on the assumption that poor performance in most audio-visual speech processing applications can be attributed to train/test mismatches we propose that the main impetus of practical audio-visual integration is to dampen the independent errors, resulting from the mismatch, rather than trying to model any bimodal speech dependencies. To this end a strategy is recommended, based on theory and empirical evidence, using a hybrid between the weighted product and weighted sum rules in the presence of varying acoustic noise for the task of text-dependent speaker recognition. Simon Lucey, Tsuhan Chen, Sridha Sridharan, Vinod Chandran |
IEEE Trans. Multim. | 3 |
| 2004 | Logarithmic quantisation of wavelet coefficients for improved texture classification performanceabstractThe coefficients of the wavelet transform have been widely used for texture analysis tasks, including segmentation, classification and synthesis. Second order statistics of such values have been shown to give excellent performance in these applications, and are typically calculated using co-occurrence matrices, which require quantisation of the coefficients. In this paper, we propose a non-linear quantisation function which is experimentally shown to better characterise textured images, and use this to formulate a new set of texture features, the wavelet log co-occurrence signatures. Andrew Busch, Wageeh W. Boles, Sridha Sridharan |
ICASSP (3) | 3 |
| 2004 | Techniques for improving stereo depth maps of facesabstractThis paper presents an improved technique for determining the 3D structure of the human face form stereo images. The approach targets specific regions of the face individually. An application of the approach to estimate the 3D structure of the nose is presented. Stereo matching is extended and applied to colour in addition to grayscale, which improves matching in a low texture image like parts of the skin of a human face. Left-right consistency check is used to determine erroneous disparity estimates. Prior knowledge of the nose structure is used in conjunction to improve the quality of the 3D nose estimates. Jason Baker, Vinod Chandran, Sridha Sridharan |
ICIP | 3 |
| 2004 | Visual attention based roi maps from gaze tracking dataabstractThe use of visual attention (VA) spatial and temporal characteristics, monitored by a gaze-tracking device, to generate a region of interest (ROI) 'importance' map is proposed. A K-means clustering approach is adopted to group gaze location points into a number of clusters to represent the loci of regions of VA (or ROIs). Several metrics are then derived from the gaze positions and sequences to quantify the relative importance of the K-means clusters. An entropy-weighting strategy is adopted for the combination of these metrics to generate the ROI map. Results show that the ROI map is robust to the number of clusters and different gaze patterns, and can be used in progressive image coding/decoding to enhance the image quality in regions of interest. Anthony N. Nguyen, Vinod Chandran, Sridha Sridharan |
ICIP | 3 |
| 2004 | Importance prioritisation in JPEG 2000 for improved interpretability
Anthony N. Nguyen, Vinod Chandran, Sridha Sridharan |
Signal Process. Image Commun. | 3 |
| 2003 | Real-time adaptive background segmentationabstractAutomatic analysis of digital video scenes often requires the segmentation of moving objects from the background. Historically, algorithms developed for this purpose have been restricted to small frame sizes, low frame rates or offline processing. The simplest approach involves subtracting the current frame from the known background. However, as the background is unknown, the key is how to learn and model it. The paper proposes a new algorithm that represents each pixel in the frame by a group of clusters. The clusters are ordered according the likelihood that they model the background and are adapted to deal with background and lighting variations. Incoming pixels are matched against the corresponding cluster group and are classified according to whether the matching cluster is considered part of the background. The algorithm has been subjectively evaluated against three other techniques. It demonstrates equal or better segmentation than the other techniques and proves capable of processing 320/spl times/240 video at 28 fps, excluding post-processing. Darren E. Butler, Sridha Sridharan, V. Michael Bove Jr. |
ICASSP (3) | 2 |
| 2003 | Three approaches to multilingual phone recognitionabstractThis paper investigates and compares three different approaches of multilingual phone recognition (MPR). Two types of MPR approach are defined according to the language identification (LID) process of the system: explicit-LID where language identification is mandatory, and implicit-LID where LID is an integrated part of the MPR process. The OGI-TS database is employed to perform the isolated and continuous MPR experiments. Three of the world's most spoken languages, English, Mandarin and Spanish, are selected as the target languages for the system. Experimental results indicate that different MPR approaches should be employed for different applications according to the degree of LID accuracy that can be achieved from the input test utterance. If high LID accuracy is achievable, the MPR approach that depends on LID can obtain better performance. Conversely, the implicit-LID MPR approach is more appropriate. Eddie Wong, Sridha Sridharan |
ICASSP (1) | 2 |
| 2003 | Real-time adaptive background segmentationabstractAutomatic analysis of digital video scenes often requires the segmentation of moving objects from the background. Historically, algorithms developed for this purpose have been restricted to small frame sizes, low frame rates or offline processing. The simplest approach involves subtracting the current frame from the known background. However, as the background is unknown, the key is how to learn and model it. This paper proposes a new algorithm that represents each pixel in the frame by a group of clusters. The clusters are ordered according the likelihood they model the background and are adapted to deal with background and lighting variations. Incoming pixels are matched against the corresponding cluster group and are classified according to whether the matching cluster is considered part of the background. The algorithm has been subjectively evaluated against three other techniques. It demonstrated equal or better segmentation than the other techniques and proved capable of processing 320 /spl times/ 240 video at 28 fps, excluding post-processing. Darren E. Butler, Sridha Sridharan, V. Michael Bove Jr. |
ICME | 2 |
| 2003 | Cross-lingual pronunciation modelling for indonesian speech recognitionabstractThe resources necessary to produce Automatic Speech Recognition systems for a new language are considerable, and for many languages these resources are not available. This emphasizes the need for the development of generic techniques which overcome this data shortage. Indonesian is one language which suffers from this problem and whose population and importance suggest it could benefit from speech enabled technology. Accordingly, we investigate using English acoustic models to recognize Indonesian speech. The mapping process, where the symbolic representation of the Source language acoustic models is equated to the Target language phonetic units, has typically been achieved using one to one mapping techniques. This mapping method does not allow for the incorporation of predictable allophonic variation in the lexicon. Accordingly, in this paper we present the use of cross-lingual pronunciation modelling to extract context dependant mapping rules, which are subsequently used to produce a more accurate cross lingual lexicon. Terrence Martin, Torbjørn Svendsen, Sridha Sridharan |
INTERSPEECH | 3 |
| 2003 | Isolated word verification using cohort word-level verification
Kishan Thambiratnam, Sridha Sridharan |
INTERSPEECH | 2 |
| 2003 | Dependence of GMM adaptation on feature post-processing for speaker recognition
Robbie Vogt, Jason W. Pelecanos, Sridha Sridharan |
INTERSPEECH | 3 |
| 2003 | Multilingual phone clustering for recognition of spontaneous indonesian speech utilising pronunciation modelling techniquesabstractIn this paper, a multilingual acoustic model set derived from English, Hindi, and Spanish is utilised to recognise speech in Indonesian. In order to achieve this task we incorporate a two tiered approach to perform the cross-lingual porting of the multilingual models to a new language. In the first stage, we use an entropy based decision tree to merge similar phones from different languages intoclustersto forma newmultilingual model set. In the second stage, we propose the use of a cross-lingual pronunciation modelling technique to perform the mapping from the multilingual models to the Indonesian phone set. A set of mapping rules are derived from this process and are employed to convert the original Indonesian lexicon into a pronunciation lexicon in terms of the multilingual model set. Preliminary experimental results show that, compared to the common knowledge based approach, both of these techniques reduce the word error rate in a spontaneous speech recognition task. Eddie Wong, Terrence Martin, Torbjørn Svendsen, Sridha Sridharan |
INTERSPEECH | 4 |
| 2003 | Importance prioritization coding in JPEG2000 for interpretability with application to surveillance imagery
Anthony N. Nguyen, Vinod Chandran, Sridha Sridharan, Robert Prandolini |
VCIP | 3 |
| 2002 | Chromatic colour spaces for skin detection using GMMSabstractSkin detection often forms part of both face detectors and pornography detectors for Internet filters. This paper introduces a trivial skin detection algorithm that uses Gaussian Mixture Models to classify 8 × 8 blocks according to their skin-likeness. The absence of structural features in skin regions means luminance information alone is unsuitable for accurate skin detection. Therefore, the affect of the chosen colour components on the classification accuracy is investigated. Fused features vectors extracted from the Cb and Cr chrominance components were found to give the best results, achieving an equal error rate of 5.5%. By utilising these components and an 8 × 8 block size, skin regions in JPEG and MPEG files can be located without full decompression. Darren E. Butler, Sridha Sridharan, Vinod Chandran |
ICASSP | 2 |
| 2002 | A link between cepstral shrinking and the weighted product rule in audio-visual speech recognitionabstractThe weighted product rule has been shown empirically to be of great benefit in audio-visual speech recognition (AVSR), for isolated word recognition tasks. A firm theoretical ba-sis for the selection of effective weights is of considerable interest to the audio-visual speech processing community. In this paper a clear link is established between the selec-tion of effective weightings and the approximately isotropic shrinkage that the distribution of acoustic cepstral features undergo in the presence of additive noise. An elucidation of the theoretical relationship between the cepstral shrinkage and the variance of the HMM audio log-likelihoods is then explored. 1. Simon Lucey, Sridha Sridharan, Vinod Chandran |
INTERSPEECH | 2 |
| 2002 | Methods to improve Gaussian mixture model based language identification system
Eddie Wong, Sridha Sridharan |
INTERSPEECH | 2 |
| 2002 | Adaptive mouth segmentation using chromatic features
Simon Lucey, Sridha Sridharan, Vinod Chandran |
Pattern Recognit. Lett. | 2 |
| 2001 | Trainable speech synthesis with trended hidden Markov modelsabstractWe present a trainable speech synthesis system that uses the trended hidden Markov model to generate the trajectories of spectral features of synthesis units. The synthesis units are trained from a transcribed continuous speech corpus, making the speech more natural than that produced by conventional diphone synthesisers which are generally, trained from a highly articulated speech database and require a large investment of time and effort in order to train a new voice. The,overall system has been incorporated into a PSOLA synthesiser to produce speech that is natural sounding and preserves the identity of the source speaker. John Dines, Sridha Sridharan |
ICASSP | 2 |
| 2001 | Microphone array sub-band speech recognitionabstractProposes the integration of sub-band speech recognition with a microphone array. A broadband beamforming microphone array allows for natural integration with sub-band speech recognition as the beamformer is typically implemented as a combination of band-limited sub-arrays. In the paper, rather than recombining the sub-array outputs to give a single enhanced output, we propose the fusion of separate hidden Markov models trained on each subarray frequency band. In addition, a dynamic sub-band weighting scheme is proposed in which the cross- and auto-spectral densities of the microphone array inputs are used to estimate the reliability of each frequency band. The microphone array sub-band system is evaluated on an isolated digit recognition task and compared to the standard full-band approach. The results of the proposed dynamic weighting scheme are compared to those obtained using both fixed equal sub-band weights, as well as optimal sub-band weights calculated from a priori knowledge of the correct results. Iain McCowan, Sridha Sridharan |
ICASSP | 2 |
| 2001 | Face recognition using fractal codesabstractIn this paper we propose a new method for face recognition using fractal codes. Fractal codes represent local contractive, affine transformations which when iteratively applied to range-domain pairs in an arbitrary initial image result in a fixed point close to a given image. The transformation parameters such as brightness offset, contrast factor, orientation and the address of the corresponding domain for each range are used directly as features in our method. Features of an unknown face image are compared with those pre-computed for images in a database. There is no need to iterate, use fractal neighbor distances or fractal dimensions for comparison in the proposed method. This method is robust to scale change, frame size change and rotations as well as to some noise, facial expressions and blur distortion in the image. Hossein Ebrahimpour-Komleh, Vinod Chandran, Sridha Sridharan |
ICIP (3) | 3 |
| 2001 | A suitability metric for mouth tracking through chromatic segmentationabstractRecently, the use of chromatic segmentation has come very much into vogue for mouth tracking. Our recent work has endeavored to show under what conditions and representations chromatic segmentation works. Results are presented showing that, for some members of the population, chromatic segmentation does not work satisfactorily, irrespective of recording conditions. A suitability metric is proposed that can give a quantitative measure on how well chromatic mouth tracking will work for a given subject. Simon Lucey, Sridha Sridharan, Vinod Chandran |
ICIP (3) | 2 |
| 2001 | Importance coding of still imagery based on importance maps of visually interpretable regionsabstractThe paper proposes a general framework for the importance coding of still images to maximise the interpretability versus bitrate performance. The interpretability of an image to achieve maximum content recognition is important in a diverse range of applications such as surveillance and medical. Importance coding aims to address this problem by prioritisation of the encoded image bit-stream based on the importance of regions in an image. Consequently, the most important features required for interpretability are encoded and transmitted earlier in the encoded image bit-stream. The notion of importance maps, which provide a systematic approach for the assignment of relative importance of regions in an image, are presented and its use in importance coding is developed. One highly desirable advantage of the proposed importance coding framework is that it can be implemented by the EBCOT (embedded block coding with optimised truncation) coder, which has been selected as the core-coding algorithm in JPEG 2000. Anthony N. Nguyen, Vinod Chandran, Sridha Sridharan, Robert Prandolini |
ICIP (3) | 3 |
| 2001 | Application of the trended hidden Markov model to speech synthesis
John Dines, Sridha Sridharan, Miles Moody |
INTERSPEECH | 2 |
| 2001 | An investigation of HMM classifier combination strategies for improved audio-visual speech recognition
Simon Lucey, Sridha Sridharan, Vinod Chandran |
INTERSPEECH | 2 |
| 2000 | Improving the performance of a small microphone array at low frequencies using critical band and LPC codebooksabstractThis paper presents a novel method for improving the low frequency performance of a microphone array by use of a twin codebook consisting of critical band energy and LPC distribution of speech. A modified Wiener filter created using these codebooks eliminates residual noise at low frequencies, which may result from the limitation of the array size or the beamforming algorithm. The preliminary simulation tests have shown that the proposed system is effective, and may provide good performance for those microphone array systems which are installed on a computer in an office, or in a moving vehicle. Yuchang Cao, Sridha Sridharan |
ICASSP | 2 |
| 2000 | Hybrid coding of mixed signals for digital covert audio surveillanceabstractAn overview of the requirements of digital covert audio acquisition (DCAA) is provided and a scalable hybrid coder tuned for the coding of intelligible speech in a covert environment is proposed. The codec incorporates linear predictive coding (LPC) and an M-band discrete wavelet transform (DWT) to offer effective intelligible speech coding in the presence of multiple signal sources at bit rates between 8 kbps and 32 kbps. The results of informal intelligibility testing and an analysis of the algorithm's complexity are presented to demonstrate the performance of the proposed coder. Michael Mason, Sridha Sridharan, Vinod Chandran |
ICASSP | 2 |
| 2000 | The use of temporal speech and lip information for multi-modal speaker identification via multi-stream HMMsabstractInvestigates the use of temporal lip information, in conjunction with speech information, for robust, text-dependent speaker identification. We propose that significant speaker-dependent information can be obtained from moving lips, enabling speaker recognition systems to be highly robust in the presence of noise. The fusion structure for the audio and visual information is based around the use of multi-stream hidden Markov models (MSHMM), with audio and visual features forming two independent data streams. Recent work with multi-modal MSHMMs has been performed successfully for the task of speech recognition. The use of temporal lip information for speaker identification has been performed previously (T.J. Wark et al., 1998), however this has been restricted to output fusion via single-stream HMMs. We present an extension to this previous work, and show that a MSHMM is a valid structure for multi-modal speaker identification. Tim Wark, Sridha Sridharan, Vinod Chandran |
ICASSP | 2 |
| 2000 | Initialized Eigenlip Estimator for Fast Lip Tracking Using Linear RegressionabstractMultimodal speech processing in which visual facial features are jointly processed with audio features is a rapidly advancing field. Lip movements and configurations provide useful information to improve speech and speaker recognition. However, the use of this visual information requires accurate and fast lip tracking algorithms. A new technique is outlined that is able to estimate the outer lip contour directly from a given lip intensity image via linear regression. This estimate can be improved by a active shape model that is able to track a speakers lips without requiring time consuming iterative energy minimization techniques. Results of performance are presented against known tracking algorithms using the M2VTS database. Simon Lucey, Sridha Sridharan, Vinod Chandran |
ICPR | 2 |
| 2000 | Vector Quantization Based Gaussian Modeling for Speaker VerificationabstractGaussian mixture models (GMMs) have become an established means of modeling feature distributions in speaker recognition systems. It is useful for experimentation and practical implementation purposes to develop and test these models in an efficient manner particularly when computational resources are limited. A method of combining vector quantization (VQ) with single multi-dimensional Gaussians is proposed to rapidly generate a robust model approximation to the Gaussian mixture model. A fast method of testing these systems is also proposed and implemented. Results on the NIST 1996 Speaker Recognition Database suggest comparable and in some cases an improved verification performance to the traditional GMM based analysis scheme. In addition, previous research for the task of speaker identification indicated a similar system perfomance between the VQ Gaussian based technique and GMMs. Jason W. Pelecanos, S. Myers, Sridha Sridharan, Vinod Chandran |
ICPR | 3 |
| 1999 | Robust speaker verification via fusion of speech and lip modalitiesabstractThis paper investigates the use of lip information, in conjunction with speech information, for robust speaker verification in the presence of background noise. It has been previously shown in our own work, and in the work of others, that features extracted from a speaker's moving lips hold speaker dependencies which are complementary with speech features. We demonstrate that the fusion of lip and speech information allows for a highly robust speaker verification system which outperforms the performance of either sub-system. We present a new technique for determining the weighting to be applied to each modality so as to optimize the performance of the fused system. Given a correct weighting, lip information is shown to be highly effective for reducing the false acceptance and false rejection error rates in the presence of background noise. Tim Wark, Sridha Sridharan, Vinod Chandran |
ICASSP | 2 |
| 1999 | Modelling output probability distributions for enhancing speaker recognition
Jason W. Pelecanos, Sridha Sridharan |
EUROSPEECH | 2 |
| 1998 | Two novel lossless algorithms to exploit index redundancy in VQ speech compressionabstractWe address the problem of speech compression at very low rates, with the short-term spectrum compressed to less than 20 bits per frame. Current techniques apply structured vector quantization (VQ) to the short-term synthesis filter coefficients to achieve rates of the order of 24 to 26 bits per frame. We show that temporal correlations in the VQ index stream can be introduced by dynamic codebook ordering, and that these correlations can be exploited by lossless coding approaches to reduce the number of bits per frame of the VQ scheme. The use of lossless coding ensures that no additional distortion is introduced, unlike other interframe techniques. We then detail two constructive algorithms which are able to exploit this redundancy. The first method is a delayed-decision approach, which dynamically adapts the VQ codebook to allow for efficient entropy coding of the index stream. The second is based on a vector sub-codebook approach, and does not incur any additional delay. Experimental results are presented for both methods to validate the approach. Sridha Sridharan, John Leis |
ICASSP | 1 |
| 1998 | A syntactic approach to automatic lip feature extraction for speaker identificationabstractThis paper presents a novel technique for the tracking and extraction of features from lips for the purpose of speaker identification. In noisy or other adverse conditions, identification performance via the speech signal can significantly reduce, hence additional information which can complement the speech signal is of particular interest. In our system, syntactic information is derived from chromatic information in the lip region. A model of the lip contour is formed directly from the syntactic information, with no minimization procedure required to refine estimates. Colour features are then extracted from the lips via profiles taken around the lip contour. Further improvement in lip features is obtained via linear discriminant analysis (LDA). Speaker models are built from the lip features based on the Gaussian mixture model (GMM). Identification experiments are performed on the M2VTS database, with encouraging results. Tim Wark, Sridha Sridharan |
ICASSP | 2 |
| 1998 | An approach to statistical lip modelling for speaker identification via chromatic feature extractionabstractThis paper presents a novel technique for the tracking of moving lips for the purpose of speaker identification. In our system, a model of the lip contour is formed directly from chromatic information in the lip region. Iterative refinement of contour point estimates is not required. Colour features are extracted from the lips via concatenated profiles taken around the lip contour. Reduction of order in lip features is obtained via principal component analysis (PCA) followed by linear discriminant analysis (LDA). Statistical speaker models are built from the lip features based on the Gaussian mixture model (GMM). Identification experiments performed on the M2VTS/sup 1/ database, show encouraging results. Tim Wark, Sridha Sridharan, Vinod Chandran |
ICPR | 2 |
| 1998 | Hierarchical temporal decomposition: a novel approach to efficient compression of spectral characteristics of speech
Shahrokh Ghaemmaghami, Mohamed Deriche 0001, Sridha Sridharan |
ICSLP | 3 |
| 1998 | On the convergence of Gaussian mixture models: improvements through vector quantization
James Moody, Stefan Slomka, Jason W. Pelecanos, Sridha Sridharan |
ICSLP | 4 |
| 1998 | Speech enhancement using critical band spectral subtractionabstractThis paper proposes a new enhancement technique for the enhancement of broadband noise corrupted speech. The technique exploits the human auditory systems inability to distinguish between individual frequency components within critical frequency bands. Spectral subtraction is used and the spectrum is considered as critical frequency bands rather than individual frequency components. The proposed technique is compared with the existing spectral subtraction technique, using both subjective and objective speech assessment measures. Results are quoted and indicate that there is a significant increase in intelligibility and quality. Latchman Singh, Sridha Sridharan |
ICSLP | 2 |
| 1998 | A comparison of fusion techniques in mel-cepstral based speaker identification
Stefan Slomka, Sridha Sridharan, Vinod Chandran |
ICSLP | 2 |
| 1998 | Modeling of output probability distribution to improve small vocabulary speech recognition in adverse environments
David P. Thambiratnam, Sridha Sridharan |
ICSLP | 2 |
| 1998 | Improving speaker identification performance in reverberant conditions using lip information
Tim Wark, Sridha Sridharan |
ICSLP | 2 |
| 1997 | Speech separation by simulating the cocktail party effect with a neural network controlled Wiener filterabstractA novel speech separation structure which simulates the cocktail party effect using a modified iterative Wiener filter and a multi-layer perceptron neural network is presented. The neural network is used as a speaker recognition system to control the iterative Wiener filter. The neural network is a modified perceptron with a hidden layer using feature data extracted from LPC cepstral analysis. The proposed technique has been successfully used for speech separation when the interference is competing speech or broad band noise. Yuchang Cao, Sridha Sridharan, Miles Moody |
ICASSP | 2 |
| 1997 | Telephone based speaker recognition using multiple binary classifier and Gaussian mixture modelsabstractThe present study evaluates multiple binary classifier model (MBCM) and Gaussian mixture model (GMM) solutions for both automatic speaker verification (ASV) and automatic speaker identification (ASI) problems involving text-independent telephone speech from the King speech database. The MBCM's accuracy is enhanced by selectively removing those classifiers within the model which perform worst (pruning). An unpruned MBCM outperforms a GMM for ASV and speakers taken from within the same dialectic region (San Diego, CA). Once pruned, the MBCM is found to be 2.6 times more accurate than the GMM. For closed set ASI, based on the same data, the MBCM is roughly twice as accurate as the GMM but only after pruning. Pierre Castellano, Stefan Slomka, Sridha Sridharan |
ICASSP | 3 |
| 1997 | Speech compression with preservation of speaker identityabstractAlthough much effort has been directed recently towards speech compression at rates below 4 kb/s, the primary metric for comparison has, understandably, been the amount of spectral distortion in the decompressed speech. However, an aspect which is becoming important in some applications is the ability to identify the original speaker from the coded speech algorithmically. We investigate here the effect of speech compression using multistage vector quantization of the short-term (formant) filter parameters on text-independent speaker identification. It is demonstrated that in cases where the speech is stored in a compressed database for retrieval, the speaker model should be constructed from the raw speech before spectral compression. Additionally, Gaussian models of sufficiently high order are able to reduce the negative effects of spectral vector quantization upon speaker identification accuracy. John Leis, Mark Phythian, Sridha Sridharan |
ICASSP | 3 |
| 1997 | Robust enhancement of reverberant speech using iterative noise removal
David R. Cole, Miles Moody, Sridha Sridharan |
EUROSPEECH | 3 |
| 1997 | Automatic gender identification under adverse conditions
Stefan Slomka, Sridha Sridharan |
EUROSPEECH | 2 |
| 1997 | Multichannel speech separation by eigendecomposition and its application to co-talker interference removalabstractThis paper describes the concept of eigendecomposition for multichannel signal separation, an alternative method of enhancing the desired signal corrupted by interference. The method uses two observations that come from a pair of sensors, both of which contain the desired signal and the undesired signal(s). The method assumes that the desired signal and the undesired signal(s) are uncorrelated, and the signal-to-noise ratios (SNR's) of each observation are different and not necessarily greater than unity. The technique has been successfully used to separate speech signals corrupted heavily by ambient noise, co-talker interference, and other sources such as background music. The method can also be applied to instances where more than two sensors are used by first forming two beams and then processing the output of the two beamformers. Yuchang Cao, Sridha Sridharan, Miles Moody |
IEEE Trans. Speech Audio Process. | 2 |
| 1996 | Speaker recognition in reverberant enclosuresabstractThis paper evaluates the effects of room reverberation on two automatic speaker verification (ASV) applications. Reverberation is simulated using an image method. AV is conducted with a multiple binary classifier model. Results are strongly dependent on speaker location, room size and reverberation time. ASV is poor given anechoic training but reverberant test speech. In voice activated security access, speaker locations can be identical for ASV training and testing, eliminating differences between respective impulse responses. Reverberation does not affect ASV even when it renders speech unintelligible. In covert speaker verification, speakers are uncooperative and mobile, thus test room responses unknown. By subjecting training speech to an impulse response corresponding to the centre of the room, ASV is degraded by 5.45 instead of 13.7 percent should anechoic training speech be used. Subjecting training speech to responses corresponding to two room locations does not further improve ASV performance. Pierre Castellano, Sridha Sridharan, David R. Cole |
ICASSP | 2 |
| 1996 | A two stage fuzzy decision classifier for speaker identification
Pierre Castellano, Sridha Sridharan |
Speech Commun. | 2 |
| 1995 | Speech-seeking microphone array with multi-stage processing
Yuchang Cao, Sridha Sridharan, Miles Moody |
EUROSPEECH | 2 |
| 1995 | Speech enhancement by eigen decomposition with two-channel observations
Yuchang Cao, Sridha Sridharan, Miles Moody |
EUROSPEECH | 2 |
| 1994 | The design and development of an undergraduate signal processing laboratoryabstractThe School of Electrical and Electronic Systems Engineering of Queensland University of Technology (like many other universities around the world) has recognised the importance of complementing the teaching of signal processing with computer based experiments. A laboratory has been developed to provide a "hands-on" approach to the teaching of signal processing techniques. The motivation for the development of this laboratory was the cliche "What I hear I remember but what I do I understand." The laboratory has been named as the "Signal Computing and Real-time DSP Laboratory" and provides practical training to approximately 150 final year undergraduate students each year. The paper describes the novel features of the laboratory, techniques used in the laboratory based teaching, interesting aspects of the experiments that have been developed and student evaluation of the teaching techniques.> Sridha Sridharan, Vinod Chandran, M. Dawson |
ICASSP (6) | 1 |
| 1987 | Implementation of state-space digital filter structures using block floating-point arithmeticabstractBlock floating-point arithmetic is considered as an alternative to fixed-point and floating-point arithmetic in the implementation of recursive digital filters. Block floating-point implementation of state-space digital structures is shown to have improved signal-to-noise ratio compared to fixed-point implementation and can be designed to be free of overflow. It is shown that the filter cannot support zero input limit cycle oscillations of period higher than one. An architecture suitable for the VLSI implementation of a block floating point co-processor is described. Sridha Sridharan |
ICASSP | 1 |