Alireza Sepas-Moghaddam

dblp:84/10882 · DBLP profile ↗
← Back
18ranked-venue papers
11as first author
7since 2021 · last 2023
0000-0002-4881-5386ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 4 since 2021Security and privacy · 3 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2023 Deep Gait Recognition: A Survey
abstract
Gait recognition is an appealing biometric modality which aims to identify individuals based on the way they walk. Deep learning has reshaped the research landscape in this area since 2015 through the ability to automatically learn discriminative representations. Gait recognition methods based on deep learning now dominate the state-of-the-art in the field and have fostered real-world applications. In this paper, we present a comprehensive overview of breakthroughs and recent developments in gait recognition with deep learning, and cover broad topics including datasets, test protocols, state-of-the-art solutions, challenges, and future research directions. We first review the commonly used gait datasets along with the principles designed for evaluating them. We then propose a novel taxonomy made up of four separate dimensions namely body representation, temporal representation, feature representation, and neural architecture, to help characterize and organize the research landscape and literature in this area. Following our proposed taxonomy, a comprehensive survey of gait recognition methods using deep learning is presented with discussions on their performances, characteristics, advantages, and limitations. We conclude this survey with a discussion on current challenges and mention a number of promising directions for future research in gait recognition.
Alireza Sepas-Moghaddam, Ali Etemad
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 Multi-Perspective LSTM for Joint Visual Representation Learning
abstract
We present a novel LSTM cell architecture capable of learning both intra- and inter-perspective relationships available in visual sequences captured from multiple perspectives. Our architecture adopts a novel recurrent joint learning strategy that uses additional gates and memories at the cell level. We demonstrate that by using the pro-posed cell to create a network, more effective and richer visual representations are learned for recognition tasks. We validate the performance of our proposed architecture in the context of two multi-perspective visual recognition tasks namely lip reading and face recognition. Three relevant datasets are considered and the results are compared against fusion strategies, other existing multi-input LSTM architectures, and alternative recognition solutions. The experiments show the superior performance of our solution over the considered benchmarks, both in terms of recognition accuracy and complexity. We make our code publicly available at https://github.com/arsm/MPLSTM.
Alireza Sepas-Moghaddam, Fernando Pereira 0001, Paulo Lobato Correia, Ali Etemad
CVPR1
2021 Face Trees for Expression Recognition
abstract
We propose an end-to-end architecture for facial expression recognition. Our model learns an optimal tree topology for facial landmarks, whose traversal generates a sequence from which we obtain an embedding to feed a sequential learner. The proposed architecture incorporates two main streams, one focusing on landmark positions to learn the structure of the face, while the other focuses on patches around the landmarks to learn texture information. Each stream is followed by an attention mechanism and the outputs are fed to a two-stream fusion component to perform the final classification. We conduct extensive experiments on two large-scale publicly available facial expression datasets, AffectNet and FER2013, to evaluate the efficacy of our approach. Our method outperforms other solutions in the area and sets new state-of-the-art expression recognition rates on these datasets.
Mojtaba Kolahdouzi, Alireza Sepas-Moghaddam, Ali Etemad
FG2
2021 Teacher-Student Adversarial Depth Hallucination to Improve Face Recognition
abstract
We present the Teacher-Student Generative Adversarial Network (TS-GAN) to generate depth images from single RGB images in order to boost the performance of face recognition systems. For our method to generalize well across unseen datasets, we design two components in the architecture, a teacher and a student. The teacher, which itself consists of a generator and a discriminator, learns a latent mapping between input RGB and paired depth images in a supervised fashion. The student, which consists of two generators (one shared with the teacher) and a discriminator, learns from new RGB data with no available paired depth information, for improved generalization. The fully trained shared generator can then be used in runtime to hallucinate depth from RGB for downstream applications such as face recognition. We perform rigorous experiments to show the superiority of TS-GAN over other methods in generating synthetic depth images. Moreover, face recognition experiments demonstrate that our hallucinated depth along with the input RGB images boost performance across various architectures when compared to a single RGB modality by average values of +1.2%, +2.6%, and +2.6% for IIIT-D, EURECOM, and LFW datasets respectively. We make our implementation public at: https://github.com/hardik-uppal/teacher-student-gan.git.
Hardik Uppal, Alireza Sepas-Moghaddam, Michael A. Greenspan, Ali Etemad
ICCV2
2021 Long Short-Term Memory With Gate and State Level Fusion for Light Field-Based Face Recognition
abstract
Long Short-Term Memory (LSTM) is a prominent recurrent neural network for extracting dependencies from sequential data such as time-series and multi-view data, having achieved impressive results for different visual recognition tasks. A conventional LSTM network, hereafter referred only as LSTM network, can learn a model to posteriorly extract information from one input sequence. However, if two or more dependent sequences of data are simultaneously acquired, the LSTM networks may only process those sequences consecutively, not taking benefit of the information carried out by their mutual dependencies. In this context, this paper proposes two novel LSTM cell architectures that are able to jointly learn from multiple sequences simultaneously acquired, targeting to create richer and more effective models for recognition tasks. The efficacy of the novel LSTM cell architectures is assessed by integrating them into deep learning-based methods for face recognition with multi-view, light field images. The new cell architectures jointly learn the scene horizontal and vertical parallaxes available in a light field image, to capture richer spatio-angular information from both directions. A comprehensive evaluation, with the IST-EURECOM LFFD dataset using three challenging evaluation protocols, shows the advantage of using the novel LSTM cell architectures for face recognition over the state-of-the-art light field-based methods. These results highlight the added value of the novel cell architectures when learning from correlated input sequences.
Alireza Sepas-Moghaddam, Ali Etemad, Fernando Pereira 0001, Paulo Lobato Correia
IEEE Trans. Inf. Forensics Secur.1
2021 Depth as Attention for Face Representation Learning
abstract
Face representation learning solutions have recently achieved great success for various applications such as verification and identification. However, face recognition approaches that are based purely on RGB images rely solely on intensity information, and therefore are more sensitive to facial variations, notably pose, occlusions, and environmental changes such as illumination and background. A novel depth-guided attention mechanism is proposed for deep multi-modal face recognition using low-cost RGB-D sensors. Our novel attention mechanism directs the deep network “where to look” for visual features in the RGB image by focusing the attention of the network using depth features extracted by a Convolution Neural Network (CNN). The depth features help the network focus on regions of the face in the RGB image that contain more prominent person-specific information. Our attention mechanism then uses this correlation to generate an attention map for RGB images from the depth features extracted by the CNN. We test our network on four public datasets, showing that the features obtained by our proposed solution yield better results on the Lock3DFace, CurtinFaces, IIIT-D RGB-D, and KaspAROV datasets which include challenging variations in pose, occlusion, illumination, expression, and time lapse. Our solution achieves average (increased) accuracies of 87.3% (+5.0%), 99.1% (+0.9%), 99.7% (+0.6%) and 95.3%(+0.5%) for the four datasets respectively, thereby improving the state-of-the-art. We also perform additional experiments with thermal images, instead of depth images, showing the high generalization ability of our solution when adopting other modalities for guiding the attention mechanism instead of depth information.
Hardik Uppal, Alireza Sepas-Moghaddam, Michael A. Greenspan, Ali Etemad
IEEE Trans. Inf. Forensics Secur.2
2021 CapsField: Light Field-Based Face and Expression Recognition in the Wild Using Capsule Routing
abstract
Light field (LF) cameras provide rich spatio-angular visual representations by sensing the visual scene from multiple perspectives and have recently emerged as a promising technology to boost the performance of human-machine systems such as biometrics and affective computing. Despite the significant success of LF representation for constrained facial image analysis, this technology has never been used for face and expression recognition in the wild. In this context, this paper proposes a new deep face and expression recognition solution, called CapsField, based on a convolutional neural network and an additional capsule network that utilizes dynamic routing to learn hierarchical relations between capsules. CapsField extracts the spatial features from facial images and learns the angular part-whole relations for a selected set of 2D sub-aperture images rendered from each LF image. To analyze the performance of the proposed solution in the wild, the first in the wild LF face dataset, along with a new complementary constrained face dataset captured from the same subjects recorded earlier have been captured and are made available. A subset of the in the wild dataset contains facial images with different expressions, annotated for usage in the context of face expression recognition tests. An extensive performance assessment study using the new datasets has been conducted for the proposed and relevant prior solutions, showing that the CapsField proposed solution achieves superior performance for both face and expression recognition tasks when compared to the state-of-the-art.
Alireza Sepas-Moghaddam, Ali Etemad, Fernando Pereira 0001, Paulo Lobato Correia
IEEE Trans. Image Process.1
2020 Facial Emotion Recognition Using Light Field Images with Deep Attention-Based Bidirectional LSTM
abstract
Light field cameras are able to capture the intensity of light rays coming from multiple directions, thus representing the visual scene from multiple viewpoints. This paper exploits the rich spatio-angular information available in light field images for facial emotion recognition. In this context, a new deep network is proposed that first extracts spatial features using a VGG16 convolutional neural network. Then, a Bidirectional Long Short-Term Memory (Bi-LSTM) recurrent neural network is used to learn spatio-angular features from viewpoint feature sequences, exploring both forward and backward angular relationships. Additionally, an attention mechanism allows our model to selectively focus on the most important spatio-angular features, thus enabling a more effective learning outcome. Finally, a fusion scheme is adopted to obtain the emotion recognition classification results. Comprehensive experiments have been conducted on the IST-EURECOM Light Field Face database using two challenging evaluation protocols, showing the superiority of our method over the state-of-the-art.
Alireza Sepas-Moghaddam, Ali Etemad, Fernando Pereira 0001, Paulo Lobato Correia
ICASSP1
2020 Gait Recognition using Multi-Scale Partial Representation Transformation with Capsules
abstract
Gait recognition, referring to the identification of individuals based on the manner in which they walk, can be very challenging due to the variations in the viewpoint of the camera and the appearance of individuals. Current methods for gait recognition have been dominated by deep learning models, notably those based on partial feature representations. In this context, we propose a novel deep network, learning to transfer multi-scale partial gait representations using capsules to obtain more discriminative gait features. Our network first obtains multi-scale partial representations using a state-of-the-art deep partial feature extractor. It then recurrently learns the correlations and co-occurrences of the patterns among the partial features in forward and backward directions using Bidirectional Gated Recurrent Units (BGRU). Finally, a capsule network is adopted to learn deeper part-whole relationships and assigns more weights to the more relevant features while ignoring the spurious dimensions. That way, we obtain final features that are more robust to both viewing and appearance changes. The performance of our method has been extensively tested on two gait recognition datasets, CASIA-B and OU-MVLP, using four challenging test protocols. The results of our method have been compared to the state-of-the-art gait recognition solutions, showing the superiority of our model, notably when facing challenging viewing and carrying conditions.
Alireza Sepas-Moghaddam, Saeed Ghorbani, Nikolaus F. Troje, Ali Etemad
ICPR1
2020 Two-Level Attention-based Fusion Learning for RGB-D Face Recognition
abstract
With recent advances in RGB-D sensing technologies as well as improvements in machine learning and fusion techniques, RGB-D facial recognition has become an active area of research. A novel attention aware method is proposed to fuse two image modalities, RGB and depth, for enhanced RGB-D facial recognition. The proposed method first extracts features from both modalities using a convolutional feature extractor. These features are then fused using a two layer attention mechanism. The first layer focuses on the fused feature maps generated by the feature extractor, exploiting the relationship between feature maps using LSTM recurrent learning. The second layer focuses on the spatial features of those maps using convolution. The training database is preprocessed and augmented through a set of geometric transformations, and the learning process is further aided using transfer learning from a pure 2D RGB image training process. Comparative evaluations demonstrate that the proposed method outperforms other state-of-the-art approaches, including both traditional and deep neural network-based methods, on the challenging CurtinFaces and IIIT-D RGB-D benchmark databases, achieving classification accuracies over 98.2% and 99.3% respectively. The proposed attention mechanism is also compared with other attention mechanisms, demonstrating more accurate results.
Hardik Uppal, Alireza Sepas-Moghaddam, Michael A. Greenspan, Ali Etemad
ICPR2
2020 A Double-Deep Spatio-Angular Learning Framework for Light Field-Based Face Recognition
abstract
Face recognition has attracted increasing attention due to its wide range of applications, but it is still challenging when facing large variations in the biometric data characteristics. Lenslet light field cameras have recently come into prominence to capture rich spatio-angular information, thus offering new possibilities for advanced biometric recognition systems. This paper proposes a double-deep spatio-angular learning framework for light field-based face recognition, which is able to model both the intra-view/spatial and inter-view/angular information using two deep networks in sequence. This is a novel recognition framework that has never been proposed in the literature for face recognition or any other visual recognition task. The proposed double-deep learning framework includes a long short-term memory (LSTM) recurrent network, whose inputs are VGG-Face descriptions, computed using a VGG-16 convolutional neural network (CNN). The VGG-Face spatial descriptions are extracted from a selected set of 2D sub-aperture (SA) images rendered from the light field image, corresponding to different observation angles. A sequence of the VGG-Face spatial descriptions is then analyzed by the LSTM network. A comprehensive set of experiments has been conducted using the IST-EURECOM light field face database, addressing varied and challenging recognition tasks. The results show that the proposed framework achieves superior face recognition performance when compared to the state of the art.
Alireza Sepas-Moghaddam, Mohammad A. Haque, Paulo Lobato Correia, Kamal Nasrollahi, Thomas B. Moeslund, Fernando Pereira 0001
IEEE Trans. Circuits Syst. Video Technol.1
2019 A Deep Framework for Facial Emotion Recognition using Light Field Images
abstract
Light field cameras capture the intensity of light rays coming from multiple directions, thus allowing a set of 2D images, named sub-aperture (SA) images, to be rendered. These images correspond to observations of the scene from slightly different angles. The rich spatio-angular information obtained using these cameras is exploited in this paper, for the first time, in the context of facial emotion recognition. A deep learning spatio-angular fusion framework is adopted which is able to model both the intra-view/spatial and inter-view/angular information, using a VGG-16 convolutional neural network and a long short-term memory (LSTM) recurrent network. The proposed solution, based on the adopted deep spatio-angular fusion framework, creates two view sequences, horizontal and vertical, with selected SA images, for which VGG-Face descriptions are extracted. The resulting descriptions are fed to two LSTM networks, with the aim of independently learning horizontal and vertical classification models. The softmax classifier scores obtained for the horizontal and vertical descriptors are then fused to obtain the final emotion recognition labels. A comprehensive set of experiments has been conducted on the IST-EURECOM light field face database using two assessment protocols. The adopted framework achieves superior emotion recognition performance when compared with state-of-the-art benchmarking methods.
Alireza Sepas-Moghaddam, Ali Etemad, Paulo Lobato Correia, Fernando Pereira 0001
ACII1
2018 Light Field-Based Face Presentation Attack Detection: Reviewing, Benchmarking and One Step Further
abstract
Vulnerability of face recognition systems to presentation attacks has attracted increasing attention from the biometrics and forensics communities. Moreover, the recent availability of light field cameras is opening new possibilities for designing improved face presentation attack detection solutions. In this context, this paper provides the first review and benchmarking study in the literature on light field-based face presentation attack detection solutions. State-of-the-art solutions are assessed in terms of accuracy, generalization and complexity, using a common, representative evaluation framework. This paper also proposes a novel face presentation attack detection solution, based on a histogram of oriented gradients descriptor, which exploits the disparity information available in light field imaging. The evaluation of the proposed face presentation attack detection solution for different presentation attack types shows a very effective and stable performance, notably performing better than the state-of-the-art alternatives.
Alireza Sepas-Moghaddam, Fernando Pereira 0001, Paulo Lobato Correia
IEEE Trans. Inf. Forensics Secur.1
2017 Light field local binary patterns description for face recognition
abstract
Light field cameras are emerging as powerful sensor devices to capture the full spatio-angular visual information in a viewing range. As more information should allow better analysis performance, this paper proposes a simple, yet effective descriptor, named Light Field Local Binary Patterns (LFLBP), able to exploit the richer information available in light field images for face recognition. The LFLBP descriptor combines two main components, the spatial, local LBP and the angular LBP, to capture not only the usual spatial information but also the light field angular information associated to the set of sub-aperture images, corresponding to different viewpoints. Experiments were conducted with the novel IST-EURECOM light field face database. When compared with competing methods, the proposed descriptor has shown superior face recognition performance under varied and challenging acquisition conditions. Moreover, the proposed light field angular LBP descriptor can be flexibly combined with any available spatial descriptor to derive combined descriptors for enhanced light field based face recognition performance.
Alireza Sepas-Moghaddam, Paulo Lobato Correia, Fernando Pereira 0001
ICIP1
2016 A Novel Approach for Optimization in Dynamic Environments Based on Modified Artificial Fish Swarm Algorithm
abstract
Swarm intelligence algorithms are amongst the most efficient approaches toward solving optimization problems. Up to now, most of swarm intelligence approaches have been proposed for optimization in static environments. However, numerous real-world problems are dynamic which could not be solved using static approaches. In this paper, a novel approach based on artificial fish swarm algorithm (AFSA) has been proposed for optimization in dynamic environments in which changes in the problem space occur in discrete intervals. The proposed algorithm can quickly find the peaks in the problem space and track them after an environment change. In this algorithm, artificial fish swarms are responsible for finding and tracking peaks and several behaviors and mechanisms are employed to cope with the dynamic environment. Extensive experiments show that the proposed algorithm significantly outperforms previous algorithms in most of tested dynamic environments modeled by moving peaks benchmark.
Danial Yazdani, Alireza Sepas-Moghaddam, Atabak Dehban, Nuno Horta
Int. J. Comput. Intell. Appl.2
2012 Counting the number of cells in immunocytochemical images using genetic algorithm
abstract
Immunocytochemistry (ICC) is a microscopic imaging technique that is used to assess the presence of a specific antigen in cells utilizing a specific antibody for allowing visualization and examination processes. Number of cells in an ICC image is considered as one of the most important indicators in the examination process. In this paper, an image analysis approach is proposed in order to count the number of cells in an ICC images. For this purpose, morphological filtering is done to clean up noise. Then, nucleuses and antibodies are separated by classifying relevant colors using a Nearest Neighbor classifier. Finally, the adherent cells are segmented by learning a Genetic model and the number of cells has been counted. The experiments have been conducted on a dataset of ICC images which are collected for this research. The results show the high efficiency of the proposed method.
Marjan Ramin, Payam Ahmadvand, Alireza Sepas-Moghaddam, Mohammad Mahdi Dehshibi
HIS3
2012 A novel hybrid algorithm for optimization in multimodal Dynamic environments
abstract
Objective function or the constraints and consequently the optimal value of the problem can be changed during time in Dynamic optimization problems. There are several challenges in dynamic environments, so that algorithms designed for optimization in these environments would utilize several mechanisms in order to conquer the challenges. In this paper, a novel hybrid algorithm for optimization in dynamic environments, called HPSOLS, is proposed based on particle swarm optimization and local search approaches. In this approach, it aims to increase the ability of local search around optimum with focusing on best found peak in each environment. The results of the proposed approach are evaluated using moving peak benchmark, which is currently the most well-known benchmark for evaluating dynamic environments, and are compared with results of several state-of-the-art algorithms in this domain. Experimental results show that the efficiency of the proposed method outperforms that of other algorithms in this domain.
Alireza Sepas-Moghaddam, Alireza Arabshahi, Danial Yazdani, Mohammad Mahdi Dehshibi
HIS1
2012 A multilevel thresholding method for image segmentation using a novel hybrid intelligent approach
abstract
Swarm intelligence algorithms have been extensively used in clustering based applications e.g. image segmentation which is one of the fundamental components in image analysis and pattern recognition domains. Particle swarm optimization is amongst swarm intelligence algorithms that performs based on population and random search. In this paper, a hybrid algorithm based on PSO, k-means and learning automata is proposed for image segmentation. In the proposed algorithm, learning automata is responsible for activating and deactivating PSO and k-means methods based on current conditions of the segmentation problem. The proposed approach along with other comparative studies has been applied for segmenting benchmark images. Efficiency of the proposed method has been compared with that of other methods and experimental results show the superiority proposed algorithm.
Danial Yazdani, Alireza Arabshahi, Alireza Sepas-Moghaddam, Mohammad Mahdi Dehshibi
HIS3