Bappaditya Mandal

dblp:86/777 · DBLP profile ↗
← Back
32ranked-venue papers
10as first author
9since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 7 first-author · 5 since 2021Artificial intelligence and machine learning · 11 · 5 first-author · 1 since 2021Computer networks · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2024 Interpretative Attention Networks for Structural Component Recognition
Abhishek Uniyal, Bappaditya Mandal, Niladri B. Puhan, Padmalochan Bera
ICPR (16)2
2024 Security and Privacy for Fat Intra-Body Communication: Mechanisms and Protocol Stack
abstract
Innovative medical applications based on networked implants foster the development of in-body communication technologies. Among the in-body communication technologies that are being considered, fat intra-body communication (Fat-IBC) is a very recent approach. Its main advantage lies in its higher data rate compared to earlier approaches based on capacitive and galvanic coupling. However, Fat-IBC faces privacy-, security-, as well as safety-related attacks. In this paper, we discuss security and privacy concerns about Fat-IBC, as well as corresponding countermeasures. Furthermore, we present our secure protocol stack for Fat-IBC and suggest directions for future research.
Johan Engstrand, Konrad-Felix Krentz, Noor Badariah Asan, Madhushanka Padmal, Wenqing Yan, Laya Joseph, Pramod K. B. Rangaiah, Bappaditya Mandal, Christian Rohner, Maria Mani, Robin Augustine, Thiemo Voigt
LCN8
2024 Grid LSTM based Attention Modelling for Traffic Flow Prediction
abstract
Traffic flow prediction is an important task that can directly impact the control of traffic flow positively and improve the overall traffic throughput. Although a large number of studies have been performed to improve traffic flow prediction, there are very few works on purely temporal prediction models, which is important for execution on an edge device that does not have access to spatial flow information. In order to explore the temporal prediction models further, we propose an innovative hybrid long short-term memory (LSTM) model, which we call Grid LSTM based Attention Modelling for Traffic Flow Prediction (GLSTM-A), that helps to encode temporal information better at various levels/scenarios. The proposed architecture incorporates a Grid LSTM to capture historical dependencies and a simple LSTM layer dedicated to the short-term analysis of recent data. Moreover, an innovative attention mechanism is designed to focus on the importance of data features automatically for further enhancing the model's predictive capabilities. Our proposed GLSTM-A outperforms other popular temporal prediction models such as temporal convolutional network (TCN), Bi-LSTM and LSTM, in terms of prediction accuracy and memory efficiency as mentioned in the experimental results. Experimental results and ablation studies on benchmark datasets demonstrate the superior performance of the proposed model over existing state-of-the-art models in various time series prediction tasks.
Rahul Biju, Sai Usha Nagasri Goparaju, Deepak Gangadharan, Bappaditya Mandal
VTC Spring4
2023 Visual Attention Assisted Games
abstract
In this work, we propose a committee of attention models developed for improving the deep reinforcement learning frequently used for games. The game environment is manifested with spatial and temporal attention mechanisms so as to focus on important regions while playing the games. We propose that when an agent’s visual attention space is streamlined and strategic temporally coherent representations are used, generalisation will be faster than other traditional architectures. The proposed spatial attention mechanism’s output enables direct analysis of the information recognised by the agent to choose its actions, allowing for a more straightforward interpretation of the provided state. We evaluate several techniques across a variety of games to reinforce our argument. Extensive experimental results on five Atari 2600 games demonstrate that an agent that makes use of this framework is capable of outperforming state-of-the-art models on ATARI tasks while being interpretable.
Bappaditya Mandal, Niladri B. Puhan, Varma Homi Anil
CoG1
2023 Optimization and Performance Evaluation of Hybrid Deep Learning Models for Traffic Flow Prediction
abstract
Traffic flow prediction has been regarded as a critical problem in intelligent transportation systems. An accurate prediction can help mitigate congestion and other societal problems while facilitating safer, cost and time-efficient travel. However, this requires the prediction algorithm to consider several complex characteristics of traffic flow data. These complex characteristics are an amalgamation of the spatial, temporal and periodic features exhibited by traffic flow data. To extract and leverage these features for traffic flow prediction, several hybrid deep learning models have been developed recently; however, there are still some challenges to determine the optimal architecture considering both spatial and temporal features. In this work, we perform an extensive comparison of hybrid deep learning models with and without periodicity to understand the prediction accuracy of these popular models. We propose an optimization framework that unifies genetic algorithm (GA) embedded optimization with prediction models in order to derive optimized deep learning architectures for hybrid traffic flow prediction exploring 2D spatial and temporal information. The framework enables the improvement of prediction performance and eliminates the hand-tuning process. An improved temporal convolutional network (TCN) architecture is derived using the GA driven optimization, which achieves superior traffic flow prediction accuracy compared to all other existing hybrid deep learning models on the freeway and urban traffic data from the PeMS traffic data set. We also evaluate the performance of the derived hybrid deep learning algorithms on the Raspberry PI embedded platform.
Sai Usha Nagasri Goparaju, Rahul Biju, Pravalika M, Bhavana MC, Deepak Gangadharan, Bappaditya Mandal, Pradeep C
VTC2023-Spring6
2022 Kernelized dynamic convolution routing in spatial and channel interaction for attentive concrete defect recognition
Gaurab Bhattacharya, Niladri B. Puhan, Bappaditya Mandal
Signal Process. Image Commun.3
2021 Enabling Offline Tuning of Fat Channel Communication
abstract
Though fat channel communication has advantages over earlier intra-body communication (IBC) technologies based on galvanic or capacitive coupling, the development of a protocol stack on top of fat channel communication is still at its infancy. In this paper, we consider Krentz's denial-of-sleep-resilient multi-channel medium access control (MAC) layer for IEEE 802.15.4 networks as a starting point for such a protocol stack. In brief, we conducted the following experiment with a phantom that mimics human tissues. Two devices exchanged IEEE 802.15.4 radio frames in a ping-pong manner on the phantom's fat tissue using Krentz's MAC layer. The data collected from this experiment lends itself to two purposes. First, it can serve to benchmark and tune algorithms for selecting radio channels. Second, it can also serve to benchmark and tune schemes for deriving cryptographic keys from received signal strength indicator (RSSI) readings. We made the data available at https://uppsala.box.com/s/z2a6jpigswpoifd5l73yophokcfwd88b.
Konrad-Felix Krentz, Madhushanka Padmal, Bappaditya Mandal, Robin Augustine, Thiemo Voigt
SenSys3
2021 Multi-Deformation Aware Attention Learning for Concrete Structural Defect Classification
abstract
In this work, we propose a deep multi-deformation aware attention learning (MDAL) architecture comprising of multi-scale committee of attention (MSCA) and fine-grained feature induced attention (FGIA) modules to classify multi-target multi-class defects in concrete structures found in civil infrastructures. The MDAL network is composed of interleaved MSCA and FGIA modules to encode crucial fine-grained deformation-aware information from concrete images. The novel attention mechanism is able to localize specific defect regions within an image and extracts crucial discriminative information in multi-scale fashion ranging from coarser to finer features without using any preprocessing step, such as region-of-interest selection or denoising. Our proposed attention mechanism enables the MDAL architecture to automatically classify multiple overlapping defect classes present in the concrete images and leads to an end-to-end trainable deep network. Experimental results on three large concrete defect datasets and ablation studies show that our MDAL network outperforms the current state-of-the-art methodologies significantly.
Gaurab Bhattacharya, Bappaditya Mandal, Niladri B. Puhan
IEEE Trans. Circuits Syst. Video Technol.2
2021 Interleaved Deep Artifacts-Aware Attention Mechanism for Concrete Structural Defect Classification
abstract
Automatic machine classification of concrete structural defects in images poses significant challenges because of multitude of problems arising from the surface texture, such as presence of stains, holes, colors, poster remains, graffiti, marking and painting, along with uncontrolled weather conditions and illuminations. In this paper, we propose an interleaved deep artifacts-aware attention mechanism (iDAAM) to classify multi-target multi-class and single-class defects from structural defect images. Our novel architecture is composed of interleaved fine-grained dense modules (FGDM) and concurrent dual attention modules (CDAM) to extract local discriminative features from concrete defect images. FGDM helps to aggregate multi-layer robust information with wide range of scales to describe visually-similar overlapping defects. On the other hand, CDAM selects multiple representations of highly localized overlapping defect features and encodes the crucial spatial regions from discriminative channels to address variations in texture, viewing angle, shape and size of overlapping defect classes. Within iDAAM, FGDM and CDAM are interleaved to extract salient discriminative features from multiple scales by constructing an end-to-end trainable network without any preprocessing steps, making the process fully automatic. Experimental results and extensive ablation studies on three publicly available large concrete defect datasets show that our proposed approach outperforms the current state-of-the-art methodologies.
Gaurab Bhattacharya, Bappaditya Mandal, Niladri B. Puhan
IEEE Trans. Image Process.2
2020 Towards secure backscatter-based in-body sensor networks: poster abstract
abstract
In the near future more and more people will have multiple implants to handle their diseases. The implants benefit from being connected using in-body sensor networks. We have previously shown that RF communication through human adipose (fat) tissue is feasible. In this poster, we argue why we believe that backscatter communication within this fat channel is possible. As security is of utmost importance for in-body communication, we also discuss how backscatter-based in-body networks can be secured.
Thiemo Voigt, Christian Rohner, Wenqing Yan, Laya Joseph, Sam Hylamia, Noor Badariah Asan, Bappaditya Mandal, Mauricio David Pérez, Robin Augustine
SenSys7
2020 Variance-guided attention-based twin deep network for cross-spectral periocular recognition
Sushree Sangeeta Behera, Sapna S. Mishra, Bappaditya Mandal, Niladri B. Puhan
Image Vis. Comput.3
2019 Towards Automatic Screening of Typical and Atypical Behaviors in Children With Autism
abstract
Autism spectrum disorders (ASD) impact the cognitive, social, communicative and behavioral abilities of an individual. The development of new clinical decision support systems is of importance in reducing the delay between presentation of symptoms and an accurate diagnosis. In this work, we contribute a new database consisting of video clips of typical (normal) and atypical (such as hand flapping, spinning or rocking) behaviors, displayed in natural settings, which have been collected from the YouTube video website. We propose a preliminary non-intrusive approach based on skeleton keypoint identification using pretrained deep neural networks on human body video clips to extract features and perform body movement analysis that differentiates typical and atypical behaviors of children. Experimental results on the newly contributed database show that our platform performs best with decision tree as the classifier when compared to other popular methodologies and offers a baseline against which alternate approaches may developed and tested.
Andrew Cook, Bappaditya Mandal, Donna Berry
DSAA2
2019 Multi-Level Dual-Attention Based CNN for Macular Optical Coherence Tomography Classification
abstract
In this letter, we propose a multi-level dual-attention model to classify two common macular diseases, age-related macular degeneration (AMD) and diabetic macular edema (DME) from normal macular eye conditions using optical coherence tomography (OCT) imaging technique. Our approach unifies the dual-attention mechanism at multi-levels of the pre-trained deep convolutional neural network (CNN). It provides a focused learning mechanism by taking into account both multi-level features based attention focusing on the salient coarser features and self-attention mechanism attending higher entropy regions of the finer features. Our proposed method enables the network to automatically focus on the relevant parts of the input images at different levels of feature subspaces. This leads to a more locally deformation-aware feature generation and classification. The proposed approach does not require pre-processing steps such as extraction of region of interest, denoising, and retinal flattening, making the network more robust and fully automatic. Experimental results on two macular OCT databases show the superior performance of our proposed approach as compared to the current state-of-the-art methodologies.
Sapna S. Mishra, Bappaditya Mandal, Niladri B. Puhan
IEEE Signal Process. Lett.2
2018 Deep Residual Network with Subclass Discriminant Analysis for Crowd Behavior Recognition
abstract
In this work, we extract rich representations of crowd behavior from video using a fine-tuned deep convolutional neural residual network. Using spatial partitioning trees we create subclasses within the feature maps from each of the crowd behavior attributes (classes). Features from these subclasses are then regularized using an eigen modeling scheme. This enables to model the variance appearing from the intra-subclass information. Low dimensional discriminative features are extracted after using the total subclass scatter information. Dynamic time warping is used on the cosine distance measure to find the similarity measure between videos. A 1-nearest neighbor (NN) classifier is used to find the respective crowd behavior attribute classes from the normal videos. Experimental results on large crowd behavior video database show the superior performance of our proposed framework as compared to the baseline and current state-of-the-art methodologies for the crowd behavior recognition task.
Bappaditya Mandal, Jiri Fajtl, Vasileios Argyriou, Dorothy Ndedi Monekosso, Paolo Remagnino
ICIP1
2018 Deep Adaptive Temporal Pooling for Activity Recognition
abstract
Deep neural networks have recently achieved competitive accuracy for human activity recognition. However, there is room for improvement, especially in modeling of long-term temporal importance and determining the activity relevance of different temporal segments in a video. To address this problem, we propose a learnable and differentiable module: Deep Adaptive Temporal Pooling (DATP). DATP applies a self-attention mechanism to adaptively pool the classification scores of different video segments. Specifically, using frame-level features, DATP regresses importance of different temporal segments, and generates weights for them. Remarkably, DATP is trained using only the video-level label. There is no need of additional supervision except video-level activity class label. We conduct extensive experiments to investigate various input features and different weight models. Experimental results show that DATP can learn to assign large weights to key video segments. More importantly, DATP can improve training of frame-level feature extractor. This is because relevant temporal segments are assigned large weights during back-propagation. Overall, we achieve state-of-the-art performance on UCF101, HMDB51 and Kinetics datasets.
Sibo Song, Ngai-Man Cheung, Vijay Chandrasekhar 0001, Bappaditya Mandal
ACM Multimedia4
2018 Deep residual network with regularised fisher framework for detection of melanoma
abstract
Of all the skin cancer that is prevalent, melanoma has the highest mortality rates. Melanoma becomes life threatening when it penetrates deep into the dermis layer unless detected at an early stage, it becomes fatal since it has a tendency to migrate to other parts of our body. This study presents an automated non‐invasive methodology to assist the clinicians and dermatologists for detection of melanoma. Unlike conventional computational methods which require (expensive) domain expertise for segmentation and hand crafted feature computation and/or selection, a deep convolutional neural network‐based regularised discriminant learning framework which extracts low‐dimensional discriminative features for melanoma detection is proposed. Their approach minimises the whole of within‐class variance information and maximises the total class variance information. The importance of various subspaces arising in the within‐class scatter matrix followed by dimensionality reduction using total class variance information is analysed for melanoma detection. Experimental results on ISBI 2016, MED‐NODE, PH2 and the recent ISBI 2017 databases show the efficacy of their proposed approach as compared to other state‐of‐the‐art methodologies.
Nazneen N. Sultana, Bappaditya Mandal, Niladri B. Puhan
IET Comput. Vis.2
2017 FoodNet: Recognizing Foods Using Ensemble of Deep Networks
abstract
In this letter, we propose a protocol for an automatic food recognition system that identifies the contents of the meal from the images of the food. We developed a multilayered convolutional neural network (CNN) pipeline that takes advantages of the features from other deep networks and improves the efficiency. Numerous traditional handcrafted features and methods are explored, among which CNNs are chosen as the best performing features. Networks are trained and fine-tuned using preprocessed images and the filter outputs are fused to achieve higher accuracy. Experimental results on the largest real-world food recognition database ETH Food-101 and newly contributed Indian food image database demonstrate the effectiveness of the proposed methodology as compared to many other benchmark deep learned CNN frameworks.
Paritosh Pandey, Akella Deepthi, Bappaditya Mandal, Niladri B. Puhan
IEEE Signal Process. Lett.3
2017 Person Reidentification Using Multiple Egocentric Views
abstract
Development of a robust and scalable multicamera surveillance system is the need of the hour to ensure public safety and security. Being able to reidentify and track one or more targets over multiple nonoverlapping camera field of views in a crowded environment remains an important and challenging problem because of occlusions, large change in the viewpoints, and illumination across cameras. However, the rise of wearable imaging devices has led to new avenues in solving the reidentification (re-id) problem. Unlike static cameras, where the views are often restricted or low resolution and occlusions are common scenarios, egocentric/first person views (FPVs) mostly get zoomed in, unoccluded face images. In this paper, we present a person re-id framework designed for a network of multiple wearable devices. The proposed framework builds on commonly used facial feature extraction and similarity computation methods between camera pairs and utilizes a data association method to yield globally optimal and consistent re-id results with much improved accuracy. Moreover, to ensure its utility in practical applications where a large amount of observations are available every instant, an online scheme is proposed as a direct extension of the batch method. This can dynamically associate new observations to already observed and labeled targets in an iterative fashion. We tested both the offline and online methods on realistic FPV video databases, collected using multiple wearable cameras in a complex office environment and observed large improvements in performance when compared with the state of the arts.
Anirban Chakraborty 0001, Bappaditya Mandal, Junsong Yuan 0001
IEEE Trans. Circuits Syst. Video Technol.2
2017 Towards Detection of Bus Driver Fatigue Based on Robust Visual Analysis of Eye State
abstract
Driver's fatigue is one of the major causes of traffic accidents, particularly for drivers of large vehicles (such as buses and heavy trucks) due to prolonged driving periods and boredom in working conditions. In this paper, we propose a vision-based fatigue detection system for bus driver monitoring, which is easy and flexible for deployment in buses and large vehicles. The system consists of modules of head-shoulder detection, face detection, eye detection, eye openness estimation, fusion, drowsiness measure percentage of eyelid closure (PERCLOS) estimation, and fatigue level classification. The core innovative techniques are as follows: 1) an approach to estimate the continuous level of eye openness based on spectral regression; and 2) a fusion algorithm to estimate the eye state based on adaptive integration on the multimodel detections of both eyes. A robust measure of PERCLOS on the continuous level of eye openness is defined, and the driver states are classified on it. In experiments, systematic evaluations and analysis of proposed algorithms, as well as comparison with ground truth on PERCLOS measurements, are performed. The experimental results show the advantages of the system on accuracy and robustness for the challenging situations when a camera of an oblique viewing angle to the driver's face is used for driving state monitoring.
Bappaditya Mandal, Liyuan Li, Gang S. Wang, Jie Lin 0001
IEEE Trans. Intell. Transp. Syst.1
2016 Face recognition: Perspectives from the real world
abstract
In this paper, we analyze some of our real-world deployment of face recognition (FR) systems for various applications and discuss the gaps between expectations of the user and what the system can deliver. We evaluate some of the existing algorithms with modifications for applications including FR on wearable devices (like Google Glass) for improving social interactions, monitoring of elderly people in senior citizens centers, FR of children in child care centers and face matching between a scanned IC/passport face image and few live webcam images for automatic hotel/resort checkout or clearance. Each of these applications poses unique challenges and demands specific research components so as to adapt in the actual sites.
Bappaditya Mandal
ICARCV1
2016 Appearance based robot activity recognition system
abstract
In this work, we present an appearance based human activity recognition system. It uses background modeling to segment the foreground object and extracts useful discriminative features for representing activities performed by humans and robots. Subspace based method like principal component analysis is used to extract low dimensional features from large voluminous activity images. These low dimensional features are then used to classify an activity. An apparatus is designed using a webcam, which watches a robot replicating a human fall under indoor environment. In this apparatus, a robot performs various activities (like walking, bending, moving arms) replicating humans, which also includes a sudden fall. Experimental results on robot performing various activities and standard human activity recognition databases show the efficacy of our proposed method.
Bappaditya Mandal
ICARCV1
2016 Egocentric activity recognition with multimodal fisher vector
abstract
With the increasing availability of wearable devices, research on egocentric activity recognition has received much attention recently. In this paper, we build a Multimodal Egocentric Activity dataset which includes egocentric videos and sensor data of 20 fine-grained and diverse activity categories. We present a novel strategy to extract temporal trajectory-like features from sensor data. We propose to apply the Fisher Kernel framework to fuse video and temporal enhanced sensor features. Experiment results show that with careful design of feature extraction and fusion algorithm, sensor data can enhance information-rich video data. We make publicly available the Multimodal Egocentric Activity dataset to facilitate future research.
Sibo Song, Ngai-Man Cheung, Vijay Chandrasekhar 0001, Bappaditya Mandal, Jie Lin 0001
ICASSP4
2016 Person re-identification using multiple first-person-views on wearable devices
abstract
The rise of wearable devices has led to many new ways of re-identifying an individual. Unlike static cameras, where the views are often restricted or zoomed out and occlusions are common scenarios, first-person-views (FPVs) or ego-centric views see people closely and mostly get un-occluded face images. In this paper, we propose a face re-identification framework designed for a network of multiple wearable devices. This framework utilizes a global data association method termed as Network Consistent Reidentification (NCR) that not only helps in maintaining consistency in association results across the network, but also improves the pair-wise face re-identification accuracy. To test the proposed pipeline, we collected a database of FPV videos of 72 persons using multiple wearable devices (such as Google Glasses) in a multi-storied office environment. Experimental results indicate that NCR is able to consistently achieve large performance gains when compared to the state-of-the-art methodologies.
Anirban Chakraborty 0001, Bappaditya Mandal, Hamed Kiani Galoogahi
WACV2
2016 Performance evaluation of local descriptors and distance measures on benchmarks and first-person-view videos for face identification
Bappaditya Mandal, Liyuan Li, Ashraf A. Kassim
Neurocomputing1
2015 Whole space subclass discriminant analysis for face recognition
abstract
In this work, we propose to divide each class (a person) into subclasses using spatial partition trees which helps in better capturing the intra-personal variances arising from the appearances of the same individual. We perform a comprehensive analysis on within-class and within-subclass eigen-spectrums of face images and propose a novel method of eigen-spectrum modeling which extracts discriminative features of faces from both within-subclass and total or between-subclass scatter matrices. Effective low-dimensional face discriminative features are extracted for face recognition (FR) after performing discriminant evaluation in the entire eigenspace. Experimental results on popular face databases (AR, FERET) and the challenging unconstrained YouTube Face database show the superiority of our proposed approach on all three databases.
Bappaditya Mandal, Liyuan Li, Vijay Chandrasekhar 0001, Joo-Hwee Lim
ICIP1
2015 Multi-sensor Self-Quantification of Presentations
abstract
Presentations have been an effective means of delivering information to groups for ages. Over the past few decades, technological advancements have revolutionized the way humans deliver presentations. Despite that, the quality of presentations can be varied and affected by a variety of reasons. Conventional presentation evaluation usually requires painstaking manual analysis by experts. Although the expert feedback can definitely assist users in improving their presentation skills, manual evaluation suffers from high cost and is often not accessible to most people. In this work, we propose a novel multi-sensor self-quantification framework for presentations. Utilizing conventional ambient sensors (i.e., static cameras, Kinect sensor) and the emerging wearable egocentric sensors (i.e., Google Glass), we first analyze the efficacy of each type of sensor with various nonverbal assessment rubrics, which is followed by our proposed multi-sensor presentation analytics framework. The proposed framework is evaluated on a new presentation dataset, namely NUS Multi-Sensor Presentation (NUSMSP) dataset, which consists of 51 presentations covering a diverse set of topics. The dataset was recorded with ambient static cameras, Kinect sensor, and Google Glass. In addition to multi-sensor analytics, we have conducted a user study with the speakers to verify the effectiveness of our system generated analytics, which has received positive and promising feedback.
Tian Gan 0002, Yongkang Wong, Bappaditya Mandal, Vijay Chandrasekhar 0001, Mohan Kankanhalli
ACM Multimedia3
2010 3-parameter based eigenfeature regularization for human activity recognition
abstract
We propose an appearance based eigenfeature regularization methodology for recognizing human activities. This regularization utilizes a 3-parameter based eigenmodel derived from the variances of within-class (activity) scatter matrix. Original eigenvalues are replaced by the model eigenvalues which facilitates in regularizing eigenfeatures corresponding to very small and zero eigenvalues and perform discriminant evaluation in the whole eigenspace. This is done directly from the intensity information appearing in activity images. After this regularization, low dimensional discriminative features are extracted and used for recognizing various activities. Experimental results on two benchmark databases, Weizmann and INRIA-IXMAS show the superiority of our proposed approach over other popular methods.
Bappaditya Mandal, How-Lung Eng
ICASSP1
2010 Prediction of eigenvalues and regularization of eigenfeatures for human face verification
Bappaditya Mandal, Xudong Jiang 0001, How-Lung Eng, Alex Chichung Kot
Pattern Recognit. Lett.1
2009 Complete discriminant evaluation and feature extraction in kernel space for face recognition
Xudong Jiang 0001, Bappaditya Mandal, Alex Chichung Kot
Mach. Vis. Appl.2
2008 Verification of human faces using predicted eigenvalues
abstract
To alleviate the conventional problems of LDA and its variants, we propose a procedure of predicting eigenvalues using few reliable eigenvalues from the range space. Partitioning of entire eigenspace is performed using two control points, however, the effective low dimensional discriminative vectors are extracted from the whole eigenspace. This prediction strategy enables to perform discriminant evaluation in the full eigenspace. The proposed method is evaluated and compared with 8 popular subspace based methods for face verification task. Experimental results on popular face databases show that our method consistently outperforms others.
Bappaditya Mandal, Xudong Jiang 0001, Alex Chichung Kot
ICPR1
2008 Eigenfeature Regularization and Extraction in Face Recognition
abstract
This work proposes a subspace approach that regularizes and extracts eigenfeatures from the face image. Eigenspace of the within-class scatter matrix is decomposed into three subspaces: a reliable subspace spanned mainly by the facial variation, an unstable subspace due to noise and finite number of training samples and a null subspace. Eigenfeatures are regularized differently in these three subspaces based on an eigenspectrum model to alleviate problems of instability, over-fitting or poor generalization. This also enables the discriminant evaluation performed in the whole space. Feature extraction or dimensionality reduction occurs only at the final stage after the discriminant assessment. These efforts facilitate a discriminative and stable low-dimensional feature representation of the face image. Experiments comparing the proposed approach with some other popular subspace methods on the FERET, ORL, AR and GT databases show that our method consistently outperforms others.
Xudong Jiang 0001, Bappaditya Mandal, Alex Chichung Kot
IEEE Trans. Pattern Anal. Mach. Intell.2
2007 Face Recognition Based on Discriminant Evaluation in the Whole Space
abstract
This paper proposes a face recognition approach that performs linear discriminant analysis in the whole eigenspace. It decomposes the eigenspace into two subspaces: a reliable subspace spanned mainly by the facial variation and an unstable subspace due to finite number of training samples. Eigenvalues in the unstable subspace are replaced by a constant. This alleviates the over-fitting problem and enables the discriminant evaluation in the whole space. Feature extraction or dimensionality reduction occurs only at the final stage after the discriminant assessment. These efforts facilitate a discriminative and stable low-dimensional feature representation of the face image. Experimental results comparing some popular subspace methods on FERET and ORL databases show that our approach consistently outperforms others.
Xudong Jiang 0001, Bappaditya Mandal, Alex Chichung Kot
ICASSP (2)2