VLDB 2026 Research / reviewers in the wild / expert
Akshay Asthana
dblp:47/1891
· DBLP profile ↗
31ranked-venue papers
10as first author
9since 2021 · last 2025
0000-0001-6871-346XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 8 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 8 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Context-Dependent Anomaly Action RecognitionabstractWe explore the problem of unsupervised anomaly action recognition, focusing on identifying abnormal human behavior that takes into account the surrounding context. Unlike conventional approaches that rely solely on observing the human action, we recognize that anomalies can be context-dependent—texting while driving, for example, may be considered an anomaly, but not if the car is parked safely and the driver is waiting to pick up a passenger. To this end, we propose a simple framework that processes concurrent data streams of human action and contextual information. To learn the concept of normal behavior, we employ normalizing flows to model the joint probability distribution over pre-defined normal actioncontext sample pairs. Our method is flexible in admitting the use of different action and context streams, including the ability to leverage pretrained models for extracting features describing the action and context. We conduct a series of experiments on publicly available data, demonstrating the effectiveness of our approach in identifying context-dependent anomaly actions, including application to driving scenarios. Our code is available at https://github.com/knbandit/CDAAR. Nutthadech Banditakkarakul, Ming Xu 0015, Akshay Asthana, Liang Zheng 0001, Stephen Gould |
FG | 4 |
| 2025 | Flexible Geometric Guidance for Probabilistic Human Pose Estimation with Diffusion Modelsabstract3D human pose estimation from 2D images is a challenging problem due to depth ambiguity and occlusion. Because of these challenges the task is underdetermined, where there exists multiple—possibly infinite—poses that are plausible given the image. Despite this, many prior works assume the existence of a deterministic mapping and estimate a single pose given an image. Furthermore, methods based on machine learning require a large amount of paired 2D-3D data to train and suffer from generalization issues to unseen scenarios. To address both of these issues, we propose a framework for pose estimation using diffusion models, which enables sampling from a probability distribution over plausible poses which are consistent with a 2D image. Our approach falls under the guidance framework for conditional generation, and guides samples from an unconditional diffusion model, trained only on 3D data, using the gradients of the heatmaps from a 2D keypoint detector. We evaluate our method on the Human 3.6M dataset under best-of- m multiple hypothesis evaluation, showing state-of-the-art performance among methods which do not require paired 2D-3D data for training. We additionally evaluate the generalization ability using the MPI-INF-3DHP and 3DPW datasets and demonstrate competitive performance. Finally, we demonstrate the flexibility of our framework by using it for novel tasks including pose generation and pose completion, without the need to train bespoke conditional models. We make code available at https://github.com/fsnelgar/diffusion_pose. Francis Snelgar, Ming Xu 0015, Stephen Gould, Liang Zheng 0001, Akshay Asthana |
FG | 5 |
| 2025 | Can We Predict Performance of Large Models across Vision-Language Tasks?abstractEvaluating large vision-language models (LVLMs) is very expensive, due to high computational cost and the wide variety of tasks. The good news is that if we already have some observed performance scores, we may be able to infer unknown ones. In this study, we propose a new framework for predicting unknown performance scores based on observed ones from other LVLMs or tasks. We first formulate the performance prediction as a matrix completion task. Specifically, we construct a sparse performance matrix $\boldsymbol{R}$, where each entry $R_{mn}$ represents the performance score of the $m$-th model on the $n$-th dataset. By applying probabilistic matrix factorization (PMF) with Markov chain Monte Carlo (MCMC), we can complete the performance matrix, i.e., predict unknown scores. Additionally, we estimate the uncertainty of performance prediction based on MCMC. Practitioners can evaluate their models on untested tasks with higher uncertainty first, which quickly reduces the prediction errors. We further introduce several improvements to enhance PMF for scenarios with sparse observed performance scores. Our experiments demonstrate the accuracy of PMF in predicting unknown scores, the reliability of uncertainty estimates in ordering evaluations, and the effectiveness of our enhancements for handling sparse data. Our code is available at https://github.com/Qinyu-Allen-Zhao/CrossPred-LVLM. Qinyu Zhao, Ming Xu 0015, Kartik Gupta, Akshay Asthana, Liang Zheng 0001, Stephen Gould |
ICML | 4 |
| 2024 | The First to Know: How Token Distributions Reveal Hidden Knowledge in Large Vision-Language Models?
Qinyu Zhao, Ming Xu 0015, Kartik Gupta, Akshay Asthana, Liang Zheng 0001, Stephen Gould |
ECCV (48) | 4 |
| 2024 | Towards Optimal Feature-Shaping Methods for Out-of-Distribution DetectionabstractFeature shaping refers to a family of methods that exhibit state-of-the-art performance for out-of-distribution (OOD) detection. These approaches manipulate the feature representation, typically from the penultimate layer of a pre-trained deep learning model, so as to better differentiate between in-distribution (ID) and OOD samples. However, existing feature-shaping methods usually employ rules manually designed for specific model architectures and OOD datasets, which consequently limit their generalization ability. To address this gap, we first formulate an abstract optimization framework for studying feature-shaping methods. We then propose a concrete reduction of the framework with a simple piecewise constant shaping function and show that existing feature-shaping methods approximate the optimal solution to the concrete optimization problem. Further, assuming that OOD data is inaccessible, we propose a formulation that yields a closed-form solution for the piecewise constant shaping function, utilizing solely the ID data. Through extensive experiments, we show that the feature-shaping function optimized by our method improves the generalization ability of OOD detection across a large variety of datasets and model architectures. Our code is available at https://github.com/Qinyu-Allen-Zhao/OptFSOOD. Qinyu Zhao, Ming Xu 0015, Kartik Gupta, Akshay Asthana, Liang Zheng 0001, Stephen Gould |
ICLR | 4 |
| 2024 | Reducing the Side-Effects of Oscillations in Training of Quantized YOLO NetworksabstractQuantized networks use less computational and memory resources and are suitable for deployment on edge devices. While quantization-aware training (QAT) is a well-studied approach to quantize the networks at low precision, most research focuses on over-parameterized networks for classification with limited studies on popular and edge device friendly single-shot object detection and semantic segmentation methods like YOLO. Moreover, majority of QAT methods rely on Straight Through Estimator (STE) approximation which suffers from an oscillation phenomenon resulting in sub-optimal network quantization. In this paper, we show that it is difficult to achieve extremely low precision (4-bit and lower) for efficient YOLO models even with SOTA QAT methods due to oscillation issue and existing methods to overcome this problem are not effective on these models. To mitigate the effect of oscillation, we first propose Exponentially Moving Average (EMA) based update to the QAT model. Further, we propose a simple QAT correction method, namely QC, that takes only a single epoch of training after standard Quantization-Aware Training (QAT) procedure to correct the error induced by oscillating weights and activations resulting in a more accurate quantized model. With extensive evaluation on COCO dataset using various YOLO5 and YOLO7 variants, we show that our correction method improves quantized YOLO networks consistently on both object detection and segmentation tasks at low-precision (4-bit and 3-bit). Kartik Gupta, Akshay Asthana |
WACV | 2 |
| 2023 | A Weakly Supervised Approach to Emotion-change Prediction and Improved Mood InferenceabstractWhilst a majority of affective computing research focuses on inferring emotions, examining mood or understanding the mood-emotion interplay has received significantly less attention. Building on prior work, we (a) deduce and incorporate emotion-change ($\Delta$) information for inferring mood, without resorting to annotated labels, and (b) attempt mood prediction for long duration video clips, in alignment with the characterisation of mood. We generate the emotion-change ($\Delta$) labels via metric learning from a pre-trained Siamese Network, and use these in addition to mood labels for mood classification. Experiments evaluating unimodal (training only using mood labels) vs muttimodat (training using mood plus $\Delta$ labels) models show that mood prediction benefits from the incorporation of emotion-change information, emphasising the importance of modelling the moodemotion interplay for effective mood inference. Soujanya Narayana, Ibrahim Radwan, Ravikiran Parameshwara, Iman Abbasnejad, Akshay Asthana, Subramanian Ramanathan, Roland Göcke |
ACII | 5 |
| 2023 | Efficient Labelling of Affective Video Datasets via Few-Shot & Multi-Task Contrastive LearningabstractWhilst deep learning techniques have achieved excellent emotion prediction, they still require large amounts of labelled training data, which are (a) onerous and tedious to compile, and (b) prone to errors and biases. We propose Multi-Task Contrastive Learning for Affect Representation (MT-CLAR) for few-shot affect inference. MT-CLAR combines multi-task learning with a Siamese network trained via contrastive learning to infer from a pair of expressive facial images (a) the (dis)similarity between the facial expressions, and (b) the difference in valence and arousal levels of the two faces. We further extend the image-based MT-CLAR framework for automated video labelling where, given one or a few labelled video frames (termed support-set), MT-CLAR labels the remainder of the video for valence and arousal. Experiments are performed on the AFEW-VA dataset with multiple support-set configurations; moreover, supervised learning on representations learnt via MT-CLAR are used for valence, arousal and categorical emotion prediction on the AffectNet and AFEW-VA datasets. The results show that valence and arousal predictions via MT-CLAR are very comparable to the state-of-the-art (SOTA), and we significantly outperform SOTA with a support-set ≈6% the size of the video dataset. Ravikiran Parameshwara, Ibrahim Radwan, Akshay Asthana, Iman Abbasnejad, Subramanian Ramanathan, Roland Göcke |
ACM Multimedia | 3 |
| 2021 | Does Keypoint Estimation Benefit Object Detection? An Empirical Study of One-stage and Two-stage DetectorsabstractThis paper studies the benefit of keypoint estimation on object detection. In particular, we focus on the paradigmatic one-stage and two-stage methods, two main categories in the object detection community. We note that while there has been remarkable progress on object detection and keypoint detection, insights on how the latter would benefit the former are somehow lacking. In this paper, we make two contributions. As a major contribution, we point out that one-stage and two-stage detectors have different abilities in accommodating keypoint description. The difference is clearly shown in our experiment where multiple detectors are compared in various detection tasks. Our essential observation is that one-stage detectors benefit consistently from the inclusion of a keypoint detection branch, while for two-stage detectors such benefit is obscure. As a minor contribution, we make several variant designs to improve the trade-off between efficiency and accuracy of the one-stage CenterNet [39] on multiple detection tasks. Yang Yang 0223, Akshay Asthana, Liang Zheng 0001 |
FG | 2 |
| 2019 | Estimation of Missing Human Body Parts Via Bidirectional LSTMabstractIn this paper, a bi-directional long-short term memory (LSTM) based approach is proposed for the estimation of missing body parts in a human pose estimation context. Accurate human pose estimation is often a key component for accurate human action and activity recognition. The key idea of our algorithm is to learn the temporal consistencies of the human body poses between previous and subsequent frames. This helps in estimating missing body parts and improves the general smoothness of the pose detection results. The approach acts as a post-processing step after the application of any off-the-shelf body part detector and has been evaluated on the PoseTrack dataset for both validation and testing sequences. The results show consistent improvement in the detection across all body parts. Ibrahim Radwan, Akshay Asthana, Hafsa Ismail, Byron W. Keating, Roland Göcke |
FG | 2 |
| 2018 | A Comprehensive Performance Evaluation of Deformable Face Tracking "In-the-Wild"abstractRecently, technologies such as face detection, facial landmark localisation and face recognition and verification have matured enough to provide effective and efficient solutions for imagery captured under arbitrary conditions (referred to as "in-the-wild"). This is partially attributed to the fact that comprehensive "in-the-wild" benchmarks have been developed for face detection, landmark localisation and recognition/verification. A very important technology that has not been thoroughly evaluated yet is deformable face tracking "in-the-wild". Until now, the performance has mainly been assessed qualitatively by visually assessing the result of a deformable face tracking technology on short videos. In this paper, we perform the first, to the best of our knowledge, thorough evaluation of state-of-the-art deformable face tracking pipelines using the recently introduced 300 VW benchmark. We evaluate many different architectures focusing mainly on the task of on-line deformable face tracking. In particular, we compare the following general strategies: (a) generic face detection plus generic facial landmark localisation, (b) generic model free tracking plus generic facial landmark localisation, as well as (c) hybrid approaches using state-of-the-art face detection, model free tracking and facial landmark localisation technologies. Our evaluation reveals future avenues for further research on the topic. Grigorios Chrysos 0002, Epameinondas Antonakos, Patrick Snape, Akshay Asthana, Stefanos Zafeiriou |
Int. J. Comput. Vis. | 4 |
| 2016 | Face fiducial detection by consensus of exemplarsabstractFacial fiducial detection is a challenging problem for several reasons like varying pose, appearance, expression, partial occlusion and others. In the past, several approaches like mixture of trees [32], regression based methods [8], exemplar based methods [7] have been proposed to tackle this challenge. In this paper, we propose an exemplar based approach to select the best solution from among outputs of regression and mixture of trees based algorithms (which we call candidate algorithms). We show that by using a very simple SIFT and HOG based descriptor, it is possible to identify the most accurate fiducial outputs from a set of results produced by candidate algorithms on any given test image. Our approach manifests as two algorithms, one based on optimizing an objective function with quadratic terms and the other based on simple kNN. Both algorithms take as input fiducial locations produced by running state-of-the-art candidate algorithms on an input image, and output accurate fiducials using a set of automatically selected exemplar images with annotations. Our surprising result is that in this case, a simple algorithm like kNN is able to take advantage of the seemingly huge complementarity of these candidate algorithms, better than optimization based algorithms. We do extensive experiments on several datasets, and show that our approach outperforms state-of-the-art consistently. In some cases, we report as much as a 10% improvement in accuracy. We also extensively analyze each component of our approach, to illustrate its efficacy. An implementation and extended technical report of our approach is available www.sites.google.com/site/wacv2016facefiducialexemplars. B. R. Mallikarjun, Visesh Chari, C. V. Jawahar, Akshay Asthana |
WACV | 4 |
| 2015 | From Pixels to Response Maps: Discriminative Image Filtering for Face Alignment in the WildabstractWe propose a face alignment framework that relies on the texture model generated by the responses of discriminatively trained part-based filters. Unlike standard texture models built from pixel intensities or responses generated by generic filters (e.g. Gabor), our framework has two important advantages. First, by virtue of discriminative training, invariance to external variations (like identity, pose, illumination and expression) is achieved. Second, we show that the responses generated by discriminatively trained filters (or patch-experts) are sparse and can be modeled using a very small number of parameters. As a result, the optimization methods based on the proposed texture model can better cope with unseen variations. We illustrate this point by formulating both part-based and holistic approaches for generic face alignment and show that our framework outperforms the state-of-the-art on multiple "wild" databases. The code and dataset annotations are available for research purposes from http://ibug.doc.ic.ac.uk/resources. Akshay Asthana, Stefanos Zafeiriou, Georgios Tzimiropoulos, Shiyang Cheng 0001, Maja Pantic |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2014 | Incremental Face Alignment in the WildabstractThe development of facial databases with an abundance of annotated facial data captured under unconstrained 'in-the-wild' conditions have made discriminative facial deformable models the de facto choice for generic facial landmark localization. Even though very good performance for the facial landmark localization has been shown by many recently proposed discriminative techniques, when it comes to the applications that require excellent accuracy, such as facial behaviour analysis and facial motion capture, the semi-automatic person-specific or even tedious manual tracking is still the preferred choice. One way to construct a person-specific model automatically is through incremental updating of the generic model. This paper deals with the problem of updating a discriminative facial deformable model, a problem that has not been thoroughly studied in the literature. In particular, we study for the first time, to the best of our knowledge, the strategies to update a discriminative model that is trained by a cascade of regressors. We propose very efficient strategies to update the model and we show that is possible to automatically construct robust discriminative person and imaging condition specific models 'in-the-wild' that outperform state-of-the-art generic face alignment strategies. Akshay Asthana, Stefanos Zafeiriou, Shiyang Cheng 0001, Maja Pantic |
CVPR | 1 |
| 2014 | 3D facial geometric features for constrained local modelabstractWe propose a 3D Constrained Local Model framework for deformable face alignment in depth image. Our framework exploits the intrinsic 3D geometric information in depth data by utilizing robust histogram-based 3D geometric features that are based on normal vectors. In addition, we demonstrate the fusion of intensity data and 3D features that further improves the facial landmark localization accuracy. The experiments are conducted on publicly available FRGC database. The results show that our 3D features based CLM completely outperforms the raw depth features based CLM in term of fitting accuracy and robustness, and the fusion of intensity and 3D depth feature further improves the performance. Another benefit is that the proposed 3D features in our framework do not require any pre-processing procedure on the data. Shiyang Cheng 0001, Stefanos Zafeiriou, Akshay Asthana, Maja Pantic |
ICIP | 3 |
| 2014 | Hybrid Decision Forests for Prostate Segmentation in Multi-channel MR ImagesabstractWe propose a fully automatic learning-based multi-atlas approach to segment the prostate using multi-channel (T1 and T2) MR images. After affine transformation to the template space, multi-scale features are extracted and separate random forest classifiers are learnt for the prostate region from the most similar T1 and T2 atlases. The probabilities from these two classifiers (T1 and T2) are then fused to obtain a robust probabilistic atlas. Finally, using the probabilistic representation for each voxel, the multi-image graph cuts algorithm is applied on these multi-channel images simultaneously to get the final segmentation. The novelty of the proposed method lies in the use of multi-channel MR images, a decision forest learnt from only the most similar MR images, and the fusion of global and local template-based classifiers for prostate segmentation. We apply this method to a set of 107 prostate images, with 77 randomly selected images used for training and the remaining 30 images for testing. The results are compared to the radiologist's labeled ground truth using cross-validation. The best result is obtained via hybrid approach in which the global classifier trained on T1 images and local template-based classifiers trained on T2 images are fused to obtain the final probability for each voxel. Our results indicate that the proposed method is robust, capable of producing accurate segmentation automatically and most importantly, not patient-specific. Qinquan Gao, Akshay Asthana, Tong Tong 0001, Yipeng Hu, Daniel Rueckert, Philip J. Edwards |
ICPR | 2 |
| 2014 | Real-time generic face tracking in the wild with CUDAabstractWe present a robust real-time face tracking system based on the Constrained Local Models framework by adopting the novel regression-based Discriminative Response Map Fitting (DRMF) method. By exploiting the algorithm's potential parallelism, we present a hybrid CPU-GPU implementation capable of achieving real-time performance at 30 to 45 FPS, on ordinary consumer-grade computers. We have made the software publicly available for research purposes Shiyang Cheng 0001, Akshay Asthana, Stefanos Zafeiriou, Jie Shen 0008, Maja Pantic |
MMSys | 2 |
| 2013 | Robust Discriminative Response Map Fitting with Constrained Local ModelsabstractWe present a novel discriminative regression based approach for the Constrained Local Models (CLMs) framework, referred to as the Discriminative Response Map Fitting (DRMF) method, which shows impressive performance in the generic face fitting scenario. The motivation behind this approach is that, unlike the holistic texture based features used in the discriminative AAM approaches, the response map can be represented by a small set of parameters and these parameters can be very efficiently used for reconstructing unseen response maps. Furthermore, we show that by adopting very simple off-the-shelf regression techniques, it is possible to learn robust functions from response maps to the shape parameters updates. The experiments, conducted on Multi-PIE, XM2VTS and LFPW database, show that the proposed DRMF method outperforms state-of-the-art algorithms for the task of generic face fitting. Moreover, the DRMF method is computationally very efficient and is real-time capable. The current MATLAB implementation takes 1 second per image. To facilitate future comparisons, we release the MATLAB code and the pre-trained models for research purposes. Akshay Asthana, Stefanos Zafeiriou, Shiyang Cheng 0001, Maja Pantic |
CVPR | 1 |
| 2012 | Facial Performance Transfer via Deformable Models and Parametric CorrespondenceabstractThe issue of transferring facial performance from one person's face to another's has been an area of interest for the movie industry and the computer graphics community for quite some time. In recent years, deformable face models, such as the Active Appearance Model (AAM), have made it possible to track and synthesize faces in real time. Not surprisingly, deformable face model-based approaches for facial performance transfer have gained tremendous interest in the computer vision and graphics community. In this paper, we focus on the problem of real-time facial performance transfer using the AAM framework. We propose a novel approach of learning the mapping between the parameters of two completely independent AAMs, using them to facilitate the facial performance transfer in a more realistic manner than previous approaches. The main advantage of modeling this parametric correspondence is that it allows a "meaningful" transfer of both the nonrigid shape and texture across faces irrespective of the speakers' gender, shape, and size of the faces, and illumination conditions. We explore linear and nonlinear methods for modeling the parametric correspondence between the AAMs and show that the sparse linear regression method performs the best. Moreover, we show the utility of the proposed framework for a cross-language facial performance transfer that is an area of interest for the movie dubbing industry. Akshay Asthana, Miles de la Hunty, Abhinav Dhall, Roland Göcke |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2011 | Pose Normalization via Learned 2D Warping for Fully Automatic Face RecognitionabstractWe present a novel approach to pose-invariant face recognition that handles continuous pose variations, is not database-specific, and achieves high accuracy without any manual intervention. Our method uses multidimensional Gaussian process regression to learn a nonlinear mapping function from the 2D shapes of faces at any non-frontal pose to the corresponding 2D frontal face shapes. We use this mapping to take an input image of a new face at an arbitrary pose and pose-normalize it, generating a synthetic frontal image of the face that is then used for recognition. Our fully automatic system for face recognition includes automatic methods for extracting 2D facial feature points and accurately estimating 3D head pose, and this information is used as input to the 2D pose-normalization algorithm. The current system can handle pose variation up to 45 degrees to the left or right (yaw angle) and up to 30 degrees up or down (pitch angle). The system demonstrates high accuracy in recognition experiments on the CMU-PIE, USF 3D, and Multi-PIE databases, showing excellent generalization across databases and convincingly outperforming other automatic methods. Akshay Asthana, Michael J. Jones 0001, Tim K. Marks, Kinh H. Tieu, Roland Göcke |
BMVC | 1 |
| 2011 | A SSIM-based approach for finding similar facial expressionsabstractThere are various scenarios where finding the most similar expression is the requirement rather than classifying one into discrete, pre-defined classes, for example, for facial expression transfer and facial expression based automatic album generation. This paper proposes a novel method for finding the most similar facial expression. Instead of the regular L2 norm distance, we investigate the use of the Structural SIMilarity (SSIM) metric for similarity comparison as a distance metric in a nearest neighbour unsupervised algorithm. The feature vectors are generated using Active Appearance Models (AAM). We also demonstrate how this technique can be extended and used for finding corresponding facial expression images across two or more subjects, which is useful in applications such as facial animation and automatic expression transfer. Person-independent facial expression performance results are shown on the Multi-PIE, FEEDTUM and AVOZES databases. We also compare the performance of the SSIM metric versus other distance metrics in a nearest neighbour search for finding the most similar facial expression to a given image. Abhinav Dhall, Akshay Asthana, Roland Göcke |
FG | 2 |
| 2011 | Emotion recognition using PHOG and LPQ featuresabstractWe propose a method for automatic emotion recognition as part of the FERA 2011 competition. The system extracts pyramid of histogram of gradients (PHOG) and local phase quantisation (LPQ) features for encoding the shape and appearance information. For selecting the key frames, K-means clustering is applied to the normalised shape vectors derived from constraint local model (CLM) based face tracking on the image sequences. Shape vectors closest to the cluster centers are then used to extract the shape and appearance features. We demonstrate the results on the SSPNET GEMEP-FERA dataset. It comprises of both person specific and person independent partitions. For emotion classification we use support vector machine (SVM) and largest margin nearest neighbour (LMNN) and compare our results to the pre-computed FERA 2011 emotion challenge baseline. Abhinav Dhall, Akshay Asthana, Roland Göcke, Tom Gedeon |
FG | 2 |
| 2011 | Fully automatic pose-invariant face recognition via 3D pose normalizationabstractAn ideal approach to the problem of pose-invariant face recognition would handle continuous pose variations, would not be database specific, and would achieve high accuracy without any manual intervention. Most of the existing approaches fail to match one or more of these goals. In this paper, we present a fully automatic system for pose-invariant face recognition that not only meets these requirements but also outperforms other comparable methods. We propose a 3D pose normalization method that is completely automatic and leverages the accurate 2D facial feature points found by the system. The current system can handle 3D pose variation up to ±45° in yaw and ±30° in pitch angles. Recognition experiments were conducted on the USF 3D, Multi-PIE, CMU-PIE, FERET, and FacePix databases. Our system not only shows excellent generalization by achieving high accuracy on all 5 databases but also outperforms other methods convincingly. Akshay Asthana, Tim K. Marks, Michael J. Jones 0001, Kinh H. Tieu, M. V. Rohith |
ICCV | 1 |
| 2011 | Regression based automatic face annotation for deformable model building
Akshay Asthana, Simon Lucey, Roland Göcke |
Pattern Recognit. | 1 |
| 2010 | Facial Expression Based Automatic Album Creation
Abhinav Dhall, Akshay Asthana, Roland Göcke |
ICONIP (2) | 2 |
| 2010 | Linear Facial Expression Transfer with Active Appearance ModelsabstractThe issue of transferring facial expressions from one person's face to another's has been an area of interest for the movie industry and the computer graphics community for quite some time. In recent years, with the proliferation of online image and video collections and web applications, such as Google Street View, the question of preserving privacy through face de-identification has gained interest in the computer vision community. In this paper, we focus on the problem of real-time dynamic facial expression transfer using an Active Appearance Model framework. We provide a theoretical foundation for a generalisation of two well-known expression transfer methods and demonstrate the improved visual quality of the proposed linear extrapolation transfer method on examples of face swapping and expression transfer using the AVOZES data corpus. Realistic talking faces can be generated in real-time at low computational cost. Miles de la Hunty, Akshay Asthana, Roland Göcke |
ICPR | 2 |
| 2010 | Illumination and Expression Invariant Recognition Using SSIM Based Sparse RepresentationabstractThe sparse representation technique has provided a new way of looking at object recognition. As we demonstrate in this paper, however, the mean-squared error (MSE) measure, which is at the heart of this technique, is not a very robust measure when it comes to comparing facial images, which differ significantly in luminance values, as it only performs pixel-by-pixel comparisons. This requires a significantly large training set with enough variations in it to offset the drawback of the MSE measure. A large training set, however, is often not available. We propose the replacement of the MSE measure by the structural similarity (SSIM) measure in the sparse representation algorithm, which performs a more robust comparison using only one training sample per subject. In addition, since the off-the-shelf sparsifiers are also written using the MSE measure, we developed our own sparsifier using genetic algorithms that use the SSIM measure. We applied the modified algorithm to the Extended Yale Face B database as well as to the Multi-PIE database with expression and illumination variations. The improved performance demonstrates the effectiveness of the proposed modifications. Asim A. Khwaja, Akshay Asthana, Roland Göcke |
ICPR | 2 |
| 2009 | Learning-based Face Synthesis for Pose-Robust Recognition from Single ImageabstractFace recognition in real-world conditions requires the ability to deal with a number of conditions, such as variations in pose, illumination and expression. In this paper, we focus on variations in head pose and use a computationally efficient regression-based approach for synthesising face images in different poses, which are used to extend the face recognition training set. In this data-driven approach, the correspondences between facial landmark points in frontal and non-frontal views are learnt offline from manually annotated training data via Gaussian Process Regression. We then use this learner to synthesise non-frontal face images from any unseen frontal image. To demonstrate the utility of this approach, two frontal face recognition systems (the commonly used PCA and the recent Multi-Region Histograms) are augmented with synthesised non-frontal views for each person. This synthesis and augmentation approach is experimentally validated on the FERET dataset, showing a considerable improvement in recognition rates for ±40° and ±60° views, while maintaining high recognition rates for ±15° and ±25° views. Akshay Asthana, Conrad Sanderson, Tom Gedeon, Roland Göcke |
BMVC | 1 |
| 2009 | Learning based automatic face annotation for arbitrary poses and expressions from frontal images onlyabstractStatistical approaches for building non-rigid deformable models, such as the active appearance model (AAM), have enjoyed great popularity in recent years, but typically require tedious manual annotation of training images. In this paper, a learning based approach for the automatic annotation of visually deformable objects from a single annotated frontal image is presented and demonstrated on the example of automatically annotating face images that can be used for building AAMs for fitting and tracking. This approach employs the idea of initially learning the correspondences between landmarks in a frontal image and a set of training images with a face in arbitrary poses. Using this learner, virtual images of unseen faces at any arbitrary pose for which the learner was trained can be reconstructed by predicting the new landmark locations and warping the texture from the frontal image. View-based AAMs are then built from the virtual images and used for automatically annotating unseen images, including images of different facial expressions, at any random pose within the maximum range spanned by the virtually reconstructed images. The approach is experimentally validated by automatically annotating face images from three different databases. Akshay Asthana, Roland Göcke, Novi Quadrianto, Tom Gedeon |
CVPR | 1 |
| 2009 | Automatic frontal face annotation and AAM building for arbitrary expressions from a single frontal image onlyabstractStatistically motivated approaches for the registration and tracking of non-rigid objects, such as the active appearance model (AAM), have become very popular. A major drawback of these approaches is that they require manual annotation of all training images which can be tedious and error prone. In this paper, a MPEG-4 based approach for the automatic annotation of frontal face images, having any arbitrary facial expression, from a single annotated frontal image is presented. This approach utilises the MPEG-4 based facial animation system to generate virtual images having different expressions and uses the existing AAM framework to automatically annotate unseen images. The approach demonstrates an excellent generalisability by automatically annotating face images from two different databases. Akshay Asthana, Asim A. Khwaja, Roland Göcke |
ICIP | 1 |
| 2008 | A Hybrid Fuzzy Approach for Human Eye Gaze Pattern Recognition
Dingyun Zhu, B. Sumudu U. Mendis, Tom Gedeon, Akshay Asthana, Roland Göcke |
ICONIP (2) | 4 |