EDBT 2026 Demo / reviewers in the wild / expert
Kien Nguyen Thanh
dblp:237/1074 · also Kien Nguyen 0001
· DBLP profile ↗
42ranked-venue papers
16as first author
18since 2021 · last 2026
0000-0002-3466-9218ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 9 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 8 first-author · 6 since 2021Security and privacy · 6 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Ilov3Splat: Instance-Level Open-Vocabulary 3D Scene Understanding in Gaussian Splatting
Binh Long Nguyen, Kien Nguyen Thanh, Sridha Sridharan, Clinton Fookes, Peyman Moghadam |
ICPR (5) | 2 |
| 2026 | Multi-agent reinforcement curriculum learning for real unmanned ground vehiclesabstractThis paper investigates the use of deep reinforcement learning (DRL) for the control of mobile robot teams within the context of navigation and task-based collaborative scenarios. We apply a DRL policy with a tailored neural network architecture as a solution to control, path planning, and higher-level guidance tasks. Our network architecture was trained using a unique multi-stage curriculum that progresses from single-agent navigation, to multi-agent pathfinding with obstacles, and finally to a complex collaborative firefighting scenario. This structured approach accelerates training convergence by systematically building sophisticated collaborative behaviours upon foundational skills, which enhances training stability and guides the agents towards learning effective and coordinated strategies The policy evaluation was conducted in both simulation and hybrid simulation-physical demonstrations utilising a real unmanned ground vehicle (UGV). The policy presented is capable of achieving multi-agent navigation tasks with a 95.83% accuracy in our testing environments, and has demonstrated emergent multi-agent behaviours. In more complex collaborative firefighting scenarios, the policy also demonstrated superior performance than baselines in reaching goals, e.g., navigating and extinguishing two fires with a 99.67% success rate, suggesting its strong potential for real-world deployment. Timothy Mead, Zhe Wang 0001, Ernest Foo, Jin Song Dong 0001, Naipeng Dong, Ryan Kok Leong Ko, Abigail M. Y. Koay, Kien Nguyen Thanh, Yue Xu 0001, Junae Kim, Stephen Bornstein |
Eng. Appl. Artif. Intell. | 8 |
| 2026 | Zoom-shot: Fast, efficient and unsupervised zero-shot knowledge transfer from CLIP to vision encodersabstractFoundation models like CLIP demonstrate exceptional capabilities over a broad domain of knowledge, such as with zero-shot classification; however, they also require significant computational resources, narrowing their real-world utility. Recent studies have shown that mapping features from pre-trained vision encoders into CLIP’s latent space can transfer some of CLIP’s abilities to smaller vision encoders, offering a promising alternative. Yet, the performance of these vision encoders still falls short of CLIP’s native capabilities, particularly in low-data regimes. In this work, we argue that enhancing training data coverage/diversity significantly improves mapping efficacy. We achieve this using tailored loss functions rather than relying on data augmentation or increasing training samples. For instance, we exploit the inherent multimodal nature of CLIP’s latent space, by incorporating cycle-consistency loss as one of our loss functions. Moreover, the mapping is learned using entirely unlabelled and unpaired data, eliminating the need for manual labelling or data pairing in novel domains. From these findings, our resulting method (Zoom-shot) offers a viable path to flexible zero-shot models for resource-limited, data-scarce settings. We test Zoom-shot’s zero-shot performance across various pre-trained vision encoders on coarse- and fine-grained datasets and achieve superior performance compared to recent works. In our ablations, we find Zoom-shot allows for a trade-off between data and compute during training; allowing for a significant reduction in required training data. All code and models are available on GitHub. Jordan Shipard, Arnold Wiliem, Kien Nguyen Thanh, Wei Xiang 0001, Clinton Fookes |
Pattern Recognit. | 3 |
| 2025 | AG-VPReID: A Challenging Large-Scale Benchmark for Aerial-Ground Video-based Person Re-IdentificationabstractWe introduce AG-VPReID, a new large-scale dataset for aerial-ground video-based person re-identification (ReID) that comprises 6,632 subjects, 32,321 tracklets and over 9.6 million frames captured by drones (altitudes ranging from 15–120m), CCTV, and wearable cameras. This dataset offers a real-world benchmark for evaluating the robustness to significant viewpoint changes, scale variations, and resolution differences in cross-platform aerial-ground settings. In addition, to address these challenges, we propose AG-VPReID-Net, an end-to-end framework composed of three complementary streams: (1) an Adapted Temporal-Spatial Stream addressing motion pattern inconsistencies and facilitating temporal feature learning, (2) a Normalized Appearance Stream leveraging physics-informed techniques to tackle resolution and appearance changes, and (3) a Multi-Scale Attention Stream handling scale variations across drone altitudes. We integrate visual-semantic cues from all streams to form a robust, viewpoint-invariant whole-body representation. Extensive experiments demonstrate that AG-VPReID-Net outperforms state-of-the-art approaches on both our new dataset and existing video-based ReID benchmarks, showcasing its effectiveness and generalizability. Nevertheless, the performance gap observed on AG-VPReID across all methods underscores the dataset’s challenging nature. The dataset, code and trained models are available at AG-VPReID-Net. Kien Nguyen Thanh, Akila Pemasiri, Feng Liu 0037, Sridha Sridharan, Clinton Fookes |
CVPR | 2 |
| 2025 | AG-VPReID 2025: Aerial-Ground Video-based Person Re-identification Challenge ResultsabstractPerson re-identification (ReID) across aerial and ground vantage points has become crucial for large-scale surveillance and public safety applications. Although significant progress has been made in ground-only scenarios, bridging the aerial-ground domain gap remains a formidable challenge due to extreme viewpoint differences, scale variations, and occlusions. Building upon the achievements of the AG-ReID 2023 Challenge, this paper introduces the AG-VPReID 2025 Challenge—the first large-scale video-based competition focused on high-altitude (80–120 m) aerial-ground person ReID. Constructed on the new AG-VPReID dataset with 3,027 identities, over 13,500 tracklets, and approximately 3.7 million frames captured from UAVs, CCTV, and wearable cameras, the challenge featured four international teams. These teams developed solutions ranging from multi-stream architectures to transformer-based temporal reasoning and physics-informed modeling. The leading approach, X-TFCLIP from UAM, attained 72.28% Rank-1 accuracy in the aerial-to-ground ReID setting and 70.77% in the ground-to-aerial ReID setting, surpassing existing baselines while highlighting the dataset’s complexity. For additional details, please refer to the official website at https://agvpreid25.github.io. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Feng Liu 0037, Xiaoming Liu 0002, Arun Ross, Tamás Endrei, Ivan DeAndres-Tame, Ruben Tolosana, Rubén Vera-Rodríguez, Aythami Morales, Julian Fierrez, Javier Ortega-Garcia, Zijing Gong, Xuehu Liu, Md. Rashidunnabi, Hugo Proença 0001, Kailash A. Hambarde, Saeid Rezaei |
IJCB | 1 |
| 2025 | AG-VPReID.VIR: Bridging Aerial and Ground Platforms for Video-based Visible-Infrared Person Re-IDabstractPerson re-identification (Re-ID) across visible and infrared modalities is crucial for 24-hour surveillance systems, but existing datasets primarily focus on ground-level perspectives. While ground-based IR systems offer nighttime capabilities, they suffer from occlusions, limited coverage, and vulnerability to obstructions—problems that aerial perspectives uniquely solve. To address these limitations, we introduce AG-VPReID.VIR, the first aerial-ground cross-modality video-based person Re-ID dataset. This dataset captures 1,837 identities across 4,861 tracklets (124,855 frames) using both UAV-mounted and fixed CCTV cameras in RGB and infrared modalities. AG-VPReID.VIR presents unique challenges including cross-viewpoint variations, modality discrepancies, and temporal dynamics. Additionally, we propose TCC-VPReID, a novel three-stream architecture designed to address the joint challenges of cross-platform and cross-modality person Re-ID. Our approach bridges the domain gaps between aerial-ground perspectives and RGB-IR modalities, through style-robust feature learning, memory-based cross-view adaptation, and intermediary-guided temporal modeling. Experiments show that AG-VPReID.VIR presents distinctive challenges compared to existing datasets, with our TCC-VPReID framework achieving significant performance gains across multiple evaluation protocols. Dataset and code are available at https://github.com/agvpreid25/AG-VPReID.VIR. Kien Nguyen Thanh, Akila Pemasiri, Akmal Jahan, Clinton Fookes, Sridha Sridharan |
IJCB | 2 |
| 2025 | Beyond geometry: The power of texture in interpretable 3D person ReID
Kien Nguyen Thanh, Akila Pemasiri, Sridha Sridharan, Clinton Fookes |
Comput. Vis. Image Underst. | 2 |
| 2025 | A survey on physics informed reinforcement learning: Review and open problemsabstractThe fusion of physical information in machine learning frameworks has revolutionized many application areas. This involves enhancing the learning process by incorporating physical constraints and adhering to physical laws. This work explores their utility for reinforcement learning applications. A thorough review of the literature on the fusion of physics information or physics priors in reinforcement learning approaches, commonly referred to as physics-informed reinforcement learning (PIRL), is presented. A novel taxonomy is introduced with the reinforcement learning pipeline as the backbone to classify existing works, compare and contrast them, and derive crucial insights. Existing works are analyzed with regard to the representation/form of the governing physics modeled for integration, their specific contribution to the typical reinforcement learning architecture, and their connection to the underlying reinforcement learning pipeline stages. Core learning architectures and physics incorporation biases (i.e., observational, inductive, and learning) of existing PIRL approaches are identified and used to further categorize the works for better understanding and adaptation. By providing a comprehensive perspective on the implementation of the physics-informed capability, the taxonomy presents a cohesive approach to PIRL. It identifies the areas where this approach has been applied, as well as the gaps and opportunities that exist. Additionally, the review highlights unresolved issues and challenges, while also incorporating potential and emerging solutions to guide future research. This nascent field holds great potential for enhancing reinforcement learning algorithms by increasing their physical plausibility, precision, data efficiency, and applicability in real-world scenarios. Chayan Banerjee, Kien Nguyen Thanh, Clinton Fookes, Maziar Raissi |
Expert Syst. Appl. | 2 |
| 2024 | Improved Packet-Level Synthetic Network Traffic GenerationabstractWhile using generative models to create synthetic network traffic is faster and cheaper than traditional testbeds, synthetic traffic suffers from problems with realism and structural completeness. State of the art traffic generation frameworks usually omit payloads because of the difficulties in representing their high-dimensional data, which makes the synthetic traffic unrealistic and limits its usefulness. This work proposes a two-stage process that takes advantage of the high repetition of some protocols, particularly those used by Industrial Control Systems, to selectively simplify payloads, greatly reducing the number of classes and reducing model loss and consequently the ability of the model to handle sequences of payloads. Model training loss was reduced by 47.796%, and payload class selection was improved up to 69% over state of the art approaches, allowing for more realistic synthetic network traffic with reduced memory and computation overheads. Jacob Soper, Yue Xu 0001, Ernest Foo, Zahra Jadidi, Kien Nguyen Thanh |
TrustCom | 5 |
| 2024 | Unlocking visual data to enhance the accuracy of AI-enabled mass valuation of urban houses: An Australian city case studyabstractDespite earlier praises of hedonic price models (HPMs) in predicting the prices of residential properties, stakeholders in the property market have raised concerns regarding inaccuracy in automated valuation models (AVMs) using conventional HPMs. Furthermore, despite significant advancements in integrating artificial intelligence and image data into AVMs, research in the Australian context is absent. This paper investigates how image data capturing visual features of properties could enhance valuation accuracy using the data of actual sales of 34,399 properties from 2018 to 2020 across 128 urban suburbs in Brisbane, the third largest city in Australia. We develop a convolutional neural network model to extract visual features from large-scale data of 320,000 street-view and aerial-view images. We develop fusion models to integrate these additional visual features in HPMs. Using several experimental designs, our preferred fusion models generate a reduction of 28.35% in root-mean-square errors. Less predictive errors of AI-enabled AVMs through the use of visual data would enhance business confidence for AVMs’ end-users for investment, reporting, risk management, tax estimation, and urban planning purposes. Viet-Ngu Hoang, Kien Nguyen Thanh, Manh Thang Nguyen, Andrea Blake |
Expert Syst. Appl. | 2 |
| 2024 | AG-ReID.v2: Bridging Aerial and Ground Views for Person Re-IdentificationabstractAerial-ground person re-identification (Re-ID) presents unique challenges in computer vision, stemming from the distinct differences in viewpoints, poses, and resolutions between high-altitude aerial and ground-based cameras. Existing research predominantly focuses on ground-to-ground matching, with aerial matching less explored due to a dearth of comprehensive datasets. To address this, we introduce AG-ReID.v2, a dataset specifically designed for person Re-ID in mixed aerial and ground scenarios. This dataset comprises 100,502 images of 1,615 unique individuals, each annotated with matching IDs and 15 soft attribute labels. Data were collected from diverse perspectives using a UAV, stationary CCTV, and smart glasses-integrated camera, providing a rich variety of intra-identity variations. Additionally, we have developed an explainable attention network tailored for this dataset. This network features a three-stream architecture that efficiently processes pairwise image distances, emphasizes key top-down features, and adapts to variations in appearance due to altitude differences. Comparative evaluations demonstrate the superiority of our approach over existing baselines. We plan to release the dataset and algorithm source code publicly, aiming to advance research in this specialized field of computer vision. Kien Nguyen Thanh, Sridha Sridharan, Clinton Fookes |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Physical Adversarial Attacks for Surveillance: A SurveyabstractModern automated surveillance techniques are heavily reliant on deep learning methods. Despite the superior performance, these learning systems are inherently vulnerable to adversarial attacks-maliciously crafted inputs that are designed to mislead, or trick, models into making incorrect predictions. An adversary can physically change their appearance by wearing adversarial t-shirts, glasses, or hats or by specific behavior, to potentially avoid various forms of detection, tracking, and recognition of surveillance systems; and obtain unauthorized access to secure properties and assets. This poses a severe threat to the security and safety of modern surveillance systems. This article reviews recent attempts and findings in learning and designing physical adversarial attacks for surveillance applications. In particular, we propose a framework to analyze physical adversarial attacks and provide a comprehensive survey of physical adversarial attacks on four key surveillance tasks: detection, identification, tracking, and action recognition under this framework. Furthermore, we review and analyze strategies to defend against physical adversarial attacks and the methods for evaluating the strengths of the defense. The insights in this article present an important step in building resilience within surveillance systems to physical adversarial attacks. Kien Nguyen Thanh, Tharindu Fernando, Clinton Fookes, Sridha Sridharan |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | AG-ReID 2023: Aerial-Ground Person Re-identification Challenge ResultsabstractPerson re-identification (Re-ID) on aerial-ground platforms has emerged as an intriguing topic within computer vision, presenting a plethora of unique challenges. Highflying altitudes of aerial cameras make persons appear differently in terms of viewpoints, poses, and resolution compared to the images of the same person viewed from ground cameras. Despite its potential, few algorithms have been developed for person re-identification on aerial-ground data, mainly due to the absence of comprehensive datasets. In response, we have collected a large-scale dataset and organized the Aerial-Ground person Re-IDentification Challenge (AG-ReID2023) to foster advancements in the field. The dataset comprises 100,502 images with 1,615 unique identities, including 51,530 training images featuring 807 identities. The test set is divided into two subsets: Aerial to Ground (808 ids, 4,348 query images, 19,259 gallery images) and Ground to Aerial (808 ids, 4,151 query images, 21,214 gallery images). In addition, we manually annotate individuals with their matching IDs across cameras and provide 15 soft attribute labels. The AG-ReID2023 Challenge in conjunction with the 7thIEEE International Joint Conference on Biometrics (IJCB) has garnered interest from numerous institutes, resulting in the submission of five distinct algorithms. We provide an in-depth examination of the evaluation outcomes and present our findings from the contest. For additional details, kindly refer to the official website1.1https://agreid23.github.io. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Feng Liu 0037, Xiaoming Liu 0002, Arun Ross, Dana Michalski, Debayan Deb, Mahak Kothari, Manisha Saini, Dawei Du, Scott McCloskey, Gabriel Bertocco, Fernanda A. Andaló, Terrance E. Boult, Anderson Rocha 0001, Haidong Zhu, Zhaoheng Zheng, Ramakant Nevatia, Zaigham A. Randhawa, Sinan Sabri, Gianfranco Doretto |
IJCB | 1 |
| 2023 | Aerial-Ground Person Re-IDabstractPerson re-ID matches persons across multiple non-overlapping cameras. Despite the increasing deployment of air-borne platforms in surveillance, current existing person re-ID benchmarks’ focus is on ground-ground matching and very limited efforts on aerial-aerial matching. We propose a new benchmark dataset - AG-ReID, which performs person re-ID matching in a new setting: across aerial and ground cameras. Our dataset contains 21,983 images of 388 identities and 15 soft attributes for each identity. The data was collected by a UAV flying at altitudes between 15 to 45 meters and a ground-based CCTV camera on a university campus. Our dataset presents a novel elevated-viewpoint challenge for person re-ID due to the significant difference in person appearance across these cameras. We propose an explainable algorithm to guide the person re-ID model’s training with soft attributes to address this challenge. Experiments demonstrate the efficacy of our method on the aerial-ground person re-ID task. The dataset will be published and the baseline codes will be open-sourced at https://github.com/huynguyen792/AG-ReID to facilitate research in this area. Kien Nguyen Thanh, Sridha Sridharan, Clinton Fookes |
ICME | 2 |
| 2023 | Complex-Valued Iris Recognition NetworkabstractIn this work, we design a fully complex-valued neural network for the task of iris recognition. Unlike the problem of general object recognition, where real-valued neural networks can be used to extract pertinent features, iris recognition depends on the extraction of both phase and magnitude information from the input iris texture in order to better represent its biometric content. This necessitates the extraction and processing of phase information that cannot be effectively handled by a real-valued neural network. In this regard, we design a fully complex-valued neural network that can better capture the multi-scale, multi-resolution, and multi-orientation phase and amplitude features of the iris texture. We show a strong correspondence of the proposed complex-valued iris recognition network with Gabor wavelets that are used to generate the classical IrisCode; however, the proposed method enables a new capability of automatic complex-valued feature learning that is tailored for iris recognition. We conduct experiments on three benchmark datasets - ND-CrossSensor-2013, CASIA-Iris-Thousand and UBIRIS.v2 - and show the benefit of the proposed network for the task of iris recognition. We exploit visualization schemes to convey how the complex-valued network, when compared to standard real-valued networks, extracts fundamentally different features from the iris texture. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Arun Ross |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | HOS-FingerCode: Bispectral invariants based contactless multi-finger recognition system using ridge orientation and feature fusion
M. A. C. Akmal Jahan, Kien Nguyen Thanh, Jasmine Banks, Vinod Chandran |
Expert Syst. Appl. | 2 |
| 2021 | Multi-modal semantic image segmentation
Akila Pemasiri, Kien Nguyen Thanh, Sridha Sridharan, Clinton Fookes |
Comput. Vis. Image Underst. | 2 |
| 2021 | Hierarchical fusion network for periocular and iris by neural network approximation and sparse autoencoder
Faisal AlGashaam, Kien Nguyen Thanh, Jasmine Banks, Vinod Chandran, Mohamed I. Alkanhal |
Mach. Vis. Appl. | 2 |
| 2020 | Context from within: Hierarchical context modeling for semantic segmentation
Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan |
Pattern Recognit. | 1 |
| 2020 | Constrained Design of Deep Iris NetworksabstractDespite the promise of recent deep neural networks to provide more accurate and efficient iris recognition compared to traditional techniques, there are vital properties of the classic IrisCode which are almost unable to be achieved with current deep iris networks: the compactness of model and the small number of computing operations (FLOPs). This paper casts the iris network design process as a constrained optimization problem which takes model size and computation into account as learning criteria. On one hand, this allows us to fully automate the network design process to search for the optimal iris network architecture with the highest recognition accuracy confined to the computation and model compactness constraints. On the other hand, it allows us to investigate the optimality of the classic IrisCode and recent deep iris networks. It also enables us to learn an optimal iris network and demonstrate state-of-the-art performance with less computation and memory requirements. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan |
IEEE Trans. Image Process. | 1 |
| 2019 | Unified 2D and 3D Hand Pose Estimation from a Single Visible or X-ray Image
Akila Pemasiri, Kien Nguyen Thanh, Sridha Sridharan, Clinton Fookes |
BMVC | 2 |
| 2019 | Semantic Correspondence in the WildabstractSemantic correspondence estimation where the object instances depicted are deformed extensively from one instance to the next is a challenging problem in computer vision that has received much attention. Unfortunately, all existing approaches require prior knowledge of the object classes which are present in the image environment. This is an unwanted restriction as it can prevent the establishment of semantic correspondence across object classes in wild conditions when it is uncertain which classes will be of interest. In contrast, in this paper we formulate the semantic correspondence estimation task as a key point detection process in which image-to-class classification and image-to-image correspondence are solved simultaneously. Identifying object classes within the same framework to establish correspondence, increases this approach's applicability in real world scenarios. The use of object regions in the process also enhances the accuracy while constraining the search space, thus improving overall efficiency. This new approach is compared with the state-of-the-art on publicly available datasets to validate its capability for improved semantic correspondence estimation in wild conditions. Akila Pemasiri, Kien Nguyen Thanh, Sridha Sridharan, Clinton Fookes |
WACV | 2 |
| 2019 | Benchmarking HEp-2 specimen cells classification using linear discriminant analysis on higher order spectra features of cell shape
Khamael Al-Dulaimi, Vinod Chandran, Kien Nguyen Thanh, Jasmine Banks, Inmaculada Tomeo-Reyes |
Pattern Recognit. Lett. | 3 |
| 2019 | Sparse over-complete patch matching
Akila Pemasiri, Kien Nguyen Thanh, Sridha Sridharan, Clinton Fookes |
Pattern Recognit. Lett. | 2 |
| 2019 | Understanding Patients' Behavior: Vision-Based Analysis of Seizure DisordersabstractA substantial proportion of patients with functional neurological disorders (FND) are being incorrectly diagnosed with epilepsy because their semiology resembles that of epileptic seizures (ES). Misdiagnosis may lead to unnecessary treatment and its associated complications. Diagnostic errors often result from an overreliance on specific clinical features. Furthermore, the lack of electrophysiological changes in patients with FND can also be seen in some forms of epilepsy, making diagnosis extremely challenging. Therefore, understanding semiology is an essential step for differentiating between ES and FND. Existing sensor-based and marker-based systems require physical contact with the body and are vulnerable to clinical situations such as patient positions, illumination changes, and motion discontinuities. Computer vision and deep learning are advancing to overcome these limitations encountered in the assessment of diseases and patient monitoring; however, they have not been investigated for seizure disorder scenarios. Here, we propose and compare two marker-free deep learning models, a landmark-based and a region-based model, both of which are capable of distinguishing between seizures from video recordings. We quantify semiology by using either a fusion of reference points and flow fields, or through the complete analysis of the body. Average leave-one-subject-out cross-validation accuracies for the landmark-based and region-based approaches of 68.1% and 79.6% in our dataset collected from 35 patients, reveal the benefit of video analytics to support automated identification of semiology in the challenging conditions of a hospital setting. David Ahmedt-Aristizabal, Simon Denman, Kien Nguyen Thanh, Sridha Sridharan, Sasha Dionisio, Clinton Fookes |
IEEE J. Biomed. Health Informatics | 3 |
| 2018 | Contactless Multiple Finger Segments based Identity Verification using Information Fusion from Higher Order Spectral InvariantsabstractA methodology for identity verification from contactless finger images using ridge orientation profiles and bispectral invariant features is extended to use fusion at data, feature levels and combination of both levels using multiple finger segments. Performance is evaluated on 1341 images (selected 24 Megapixel video frames obtained with the finger to camera distance between 12 and 20cms, reduced to finger regions of about 1250 x 3000 pixels) from 41 individuals of different ethnicities. Features are extracted from profiles along key lines between landmarks that facilitate segmentation. The methodology is designed by means of the segmentation procedure, the invariant features and fusion techniques to be robust to geometric and photometric transformations and partial occlusion. Feature fusion, data fusion and a combination of the two are tested. The methodology yields around 5% EER with a combination of data and feature fusion providing the best performance. It can be applied to soft, on-the-move biometric systems. M. A. C. Akmal Jahan, Kien Nguyen Thanh, Jasmine Banks, Vinod Chandran |
AVSS | 2 |
| 2018 | Hierarchical Relational Attention for Video Question AnsweringabstractVideo Question Answering (VideoQA) tasks require understanding of the connection of context specific video parts which are temporally distributed. Humans are capable of focusing on temporally distributed video scenes and also to find correspondence or relationships among these segments. To achieve similar capability, a hierarchical relational attention mechanism is proposed in this paper. The proposed VideoQA model derives attention on temporal segments i.e. video features based on each of the question words. Also, contextual relevance of these temporal segments are captured to derive the final video representation which leads to a better reasoning capability. We evaluate the performance of the proposed approach on the MSRVTT-QA and the MSVD-QA datasets to establish its superior performance over the state of the art. Muhammad Iqbal Hasan Chowdhury, Kien Nguyen Thanh, Sridha Sridharan, Clinton Fookes |
ICIP | 2 |
| 2018 | Meta Transfer Learning for Facial Emotion RecognitionabstractThe use of deep learning techniques for automatic facial expression recognition has recently attracted great interest but developed models are still unable to generalize well due to the lack of large emotion datasets for deep learning. To overcome this problem, in this paper, we propose utilizing a novel transfer learning approach relying on PathNet and investigate how knowledge can be accumulated within a given dataset and how the knowledge captured from one emotion dataset can be transferred into another in order to improve the overall performance. To evaluate the robustness of our system, we have conducted various sets of experiments on two emotion datasets: SAVEE and eNTERFACE. The experimental results demonstrate that our proposed system leads to improvement in performance of emotion recognition and performs significantly better than the recent state-of-the-art schemes adopting fine-tuning/pre-trained approaches. Dung Nguyen Tien, Kien Nguyen Thanh, Sridha Sridharan, Iman Abbasnejad, David Dean, Clinton Fookes |
ICPR | 2 |
| 2018 | Deep spatio-temporal feature fusion with compact bilinear pooling for multimodal emotion recognition
Dung Nguyen Tien, Kien Nguyen Thanh, Sridha Sridharan, David Dean, Clinton Fookes |
Comput. Vis. Image Underst. | 2 |
| 2018 | Super-resolution for biometrics: A comprehensive survey
Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Massimo Tistarelli, Mark S. Nixon |
Pattern Recognit. | 1 |
| 2017 | A cascaded long short-term memory (LSTM) driven generic visual question answering (VQA)abstractA cascaded long short-term memory (LSTM) architecture with discriminant feature learning is proposed for the task of question answering on real world images. The proposed LSTM architecture jointly learns visual features and parts of speech (POS) tags of question words or tokens. Also, dimensionality of deep visual features is reduced by applying Principal Component Analysis (PCA) technique. In this manner, the proposed question answering model captures the generic pattern of question for a given context of image which is just not constricted within the training dataset. Empirical outcome shows that this kind of approach significantly improves the accuracy. It is believed that this kind of generic learning is a step towards a real-world visual question answering (VQA) system which will perform well for all possible forms of open-ended natural language queries. Iqbal Chowdhury, Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan |
ICIP | 2 |
| 2017 | Single image depth prediction using super-column super-pixel featuresabstractDepth prediction from a single monocular image is a challenging yet valuable task, as often a depth sensor is not available. The state-of-the-art approach [1] combines a deep fully convolutional network (DFCN) with a conditional random field (CRF), allowing the CRF to correct and smooth the depth values estimated by the DFCN according to efficient contextual modeling. However, using the output of the DFCN as unary input for CRF is limited by using only the last layer of the DFCN. The middle layers of the DFCN have been shown to carry useful information for other scene understanding tasks, which may help to improve the prediction quality. This paper proposes a novel super-column superpixel (SCSP) feature that is the combination of multiple layers of the DFCN after a super-pixel pooling process. The proposed approach based on the SCSP features reduces the root mean square (rms) error of the prediction by more than 16% in NYUv2 dataset. Xufeng Guo, Kien Nguyen Thanh, Simon Denman, Clinton Fookes, Sridha Sridharan |
ICIP | 2 |
| 2017 | Deep Context Modeling for Semantic SegmentationabstractDeep convolutional neural networks (DCNNs) have been employed in many computer vision tasks with great success due to their robustness in feature learning. One of the advantages of DCNNs is their representation robustness to object locations, which is useful for object recognition tasks. However, this also discards spatial information, which is useful when dealing with topological information of the image (e.g. scene parsing, face recognition). Adopting graphical models (GMs) to incorporate spatial and contextual information into the DCNNs is expected to improve the performance of DCNN-based computer vision tasks. Recent research has shown that combining DCNNs and Conditional Random Fields (CRFs) can significantly improve scene parsing accuracy. This is achieved either through the combination of their independent outputs or through their application as a cascade. In this work, we propose a novel strategy to incorporate CRFs deeper inside DCNNs by modeling a CRF as a DCNN layer which is pluggable into any layer of a DCNN. This implants spatial and contextual information into the DCNN, allowing end-to-end training, better controlling the spatial constraints and improving segmentation accuracy. The new strategy for coupling graphical models with the state-of-the-art fully convolutional neural network has shown promising results on the PASCAL-Context dataset. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan |
WACV | 1 |
| 2017 | Deep Spatio-Temporal Features for Multimodal Emotion RecognitionabstractAutomatic emotion recognition has attracted great interest and numerous solutions have been proposed, most of which focus either individually on facial expression or acoustic information. While more recent research has considered multimodal approaches, individual modalities are often combined only by simple fusion at the feature and/or decision-level. In this paper, we introduce a novel approach using 3-dimensional convolutional neural networks (C3Ds) to model the spatio-temporal information, cascaded with multimodal deep-belief networks (DBNs) that can represent the audio and video streams. Experiments conducted on the eNTERFACE multimodal emotion database demonstrate that this approach leads to improved multimodal emotion recognition performance and significantly outperforms recent state-of-the-art proposals. Dung Nguyen Tien, Kien Nguyen Thanh, Sridha Sridharan, Afsane Ghasemi, David Dean, Clinton Fookes |
WACV | 2 |
| 2017 | Long range iris recognition: A survey
Kien Nguyen Thanh, Clinton Fookes, Raghavender R. Jillela, Sridha Sridharan, Arun Ross |
Pattern Recognit. | 1 |
| 2016 | Deeper and wider fully convolutional network coupled with conditional random fields for scene labelingabstractDeep convolutional neural networks (DCNNs) have been employed in many computer vision tasks with great success due to their robustness in feature learning. One of the advantages of DCNNs is their representation robustness to object locations, which is useful for object recognition tasks. However, this also discards spatial information, which is useful when dealing with topological information of the image (e.g. scene labeling, face recognition). In this paper, we propose a deeper and wider network architecture to tackle the scene labeling task. The depth is achieved by incorporating predictions from multiple early layers of the DCNN. The width is achieved by combining multiple outputs of the network. We then further refine the parsing task by adopting graphical models (GMs) as a post-processing step to incorporate spatial and contextual information into the network. The new strategy for a deeper, wider convolutional network coupled with graphical models has shown promising results on the PASCAL-Context dataset. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan |
ICIP | 1 |
| 2015 | Improving deep convolutional neural networks with unsupervised feature learningabstractThe latest generation of Deep Convolutional Neural Networks (DCNN) have dramatically advanced challenging computer vision tasks, especially in object detection and object classification, achieving state-of-the-art performance in several computer vision tasks including text recognition, sign recognition, face recognition and scene understanding. The depth of these supervised networks has enabled learning deeper and hierarchical representation of features. In parallel, unsupervised deep learning such as Convolutional Deep Belief Network (CDBN) has also achieved state-of-the-art in many computer vision tasks. However, there is very limited research on jointly exploiting the strength of these two approaches. In this paper, we investigate the learning capability of both methods. We compare the output of individual layers and show that many learnt filters and outputs of the corresponding level layer are almost similar for both approaches. Stacking the DCNN on top of unsupervised layers or replacing layers in the DCNN with the corresponding learnt layers in the CDBN can improve the recognition/classification accuracy and training computational expense. We demonstrate the validity of the proposal on ImageNet dataset. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan |
ICIP | 1 |
| 2015 | Score-Level Multibiometric Fusion Based on Dempster-Shafer Theory Incorporating Uncertainty FactorsabstractWhile existing multibiometic Dempster-Shafer theory fusion approaches have demonstrated promising performance, they do not model the uncertainty appropriately, suggesting that further improvement can be achieved. This research seeks to develop a unified framework for multimodal biometric fusion to take advantage of the uncertainty concept of Dempster-Shafer theory, improving the performance of multibiometric authentication systems. Modeling uncertainty as a function of uncertainty factors affecting the recognition performance of the biometric systems helps to address the uncertainty of the data and the confidence of the fusion outcome. A weighted combination of quality measures and classifiers performance (equal error rate) is proposed to encode the uncertainty concept to improve the fusion. We also found that quality measures contribute unequally to the recognition performance; thus, selecting only significant factors and fusing them with a Dempster-Shafer approach to generate an overall quality score play an important role in the success of uncertainty modeling. The proposed approach achieved a competitive performance (approximate 1% EER) in comparison with other Dempster-Shafer-based approaches and other conventional fusion approaches. Kien Nguyen Thanh, Simon Denman, Sridha Sridharan, Clinton Fookes |
IEEE Trans. Hum. Mach. Syst. | 1 |
| 2013 | Feature-domain super-resolution for iris recognition
Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Simon Denman |
Comput. Vis. Image Underst. | 1 |
| 2012 | Feature-domain super-resolution framework for Gabor-based face and iris recognitionabstractThe low resolution of images has been one of the major limitations in recognising humans from a distance using their biometric traits, such as face and iris. Superresolution has been employed to improve the resolution and the recognition performance simultaneously, however the majority of techniques employed operate in the pixel domain, such that the biometric feature vectors are extracted from a super-resolved input image. Feature-domain superresolution has been proposed for face and iris, and is shown to further improve recognition performance by capitalising on direct super-resolving the features which are used for recognition. However, current feature-domain superresolution approaches are limited to simple linear features such as Principal Component Analysis (PCA) and Linear Discriminant Analysis (LDA), which are not the most discriminant features for biometrics. Gabor-based features have been shown to be one of the most discriminant features for biometrics including face and iris. This paper proposes a framework to conduct super-resolution in the non-linear Gabor feature domain to further improve the recognition performance of biometric systems. Experiments have confirmed the validity of the proposed approach, demonstrating superior performance to existing linear approaches for both face and iris biometrics. Kien Nguyen Thanh, Sridha Sridharan, Simon Denman, Clinton Fookes |
CVPR | 1 |
| 2011 | Feature-domain super-resolution for iris recognitionabstractUncooperative iris identification systems at a distance suffer from poor resolution of the captured iris images, which significantly degrades iris recognition performance. Super-resolution techniques have been employed to enhance the resolution of iris images and improve the recognition performance. However, all existing super-resolution approaches proposed for the iris biometric super-resolve pixel intensity values. This paper considers transferring super-resolution of iris images from the intensity domain to the feature domain. By directly super-resolving only the features essential for recognition, and by incorporating domain specific information from iris models, improved recognition performance compared to pixel domain super-resolution can be achieved. This is the first paper to investigate the possibility of feature-domain super-resolution for iris recognition, and experiments confirm the validity of the proposed approach. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Simon Denman |
ICIP | 1 |
| 2011 | Quality-Driven Super-Resolution for Less Constrained Iris Recognition at a Distance and on the MoveabstractLess constrained iris identification systems at a distance and on the move suffer from poor resolution and poor quality of the captured iris images, which significantly degrades iris recognition performance. This paper proposes a new signal-level fusion approach which incorporates a quality score into a reconstruction-based super-resolution process to generate a high-resolution iris image from a low-resolution and quality inconsistent video sequence of an eye. A novel approach for assessing the focus level of the iris image, which is invariant to lighting and oclusion conditions, is introduced. The focus score is combined with several other quality factors to perform the quality weighted super-resolution where the highest quality frames contribute the greatest amount of information to the resulting high-resolution images without introducing spurious high-frequency components. Experiments conducted on the Multiple Biometric Grand Challenge portal dataset show that our proposed approach outperforms the traditional best quality frame selection approach and other existing state-of-the-art signal-level and score-level fusion approaches for recognition of less constrained iris at a distance and on the move. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Simon Denman |
IEEE Trans. Inf. Forensics Secur. | 1 |