Mohamed Abdel-Mottaleb

dblp:20/559 · DBLP profile ↗
← Back
82ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0002-4163-7230ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 46 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 28 · 3 first-author · 3 since 2021Security and privacy · 8Computer networks · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 3Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-Level Volumetric Transformer for Early Prediction of Response to Neoadjuvant Chemotherapy in Locally Advanced Breast Cancer
abstract
Neoadjuvant Chemotherapy (NAC) is a standard treatment for locally advanced breast cancer, where achieving a Pathological Complete Response (pCR) is the primary goal. Early identification of patient response is vital for personalizing treatment and avoiding unnecessary toxicity. Recent advances in deep learning have produced models predicting NST effectiveness using breast MRI scans. However, many of these models rely on time-consuming tumor segmentation, requiring expert involvement. To address these challenges and enhance clinical applicability, we propose a Multi-Level Volumetric Transformer (MVT-Former), a novel model that predicts NAC response directly from non-segmented, full-field breast MRI data combined with relevant clinical information. The primary novelty of this work lies in its specialized dual-transformer design: (1) the Multi-Level Convolutional Spatial Transformer (MLCS-Former), which utilizes multi-scale convolutions and a Global Convolutional Attention (GCA) mechanism to extract fine-grained textural and morphological features from 2D MRI slices without manual annotations; and (2) the Volume Feature Learning Transformer (VFL-Former), which captures 3D structural changes and long-range dependencies across the entire MRI volume. We evaluated the MVT-Former on the I-SPY-1 TRIAL dataset and results demonstrate that the proposed model outperforms state-of-the-art methods, achieving superior performance across key metrics, including area under the curve, accuracy, sensitivity, and specificity.
Monu Verma, Fernando Collado-Mesa, Mohamed Abdel-Mottaleb
ACM Trans. Comput. Heal.3
2026 A motion flow guided MicroNet framework for micro expression recognition
Monu Verma, Santosh Kumar Vipparthi, Mohamed Abdel-Mottaleb
J. Vis. Commun. Image Represent.3
2026 FedHC: Enhanced federated learning with Hessian and cosine correlation for proximal correlation
Kushall Singh, Monu Verma, Dinesh Kumar Tyagi, Santosh Kumar Vipparthi, G. Shankara Raju Kosuru, M. Subrahmanyam 0001, Mohamed Abdel-Mottaleb
Knowl. Based Syst.7
2026 FedHMed: Adaptive progressive loss and KL-divergence regularization for federated heterogeneous medical image classification tasks
Kushall Singh, Monu Verma, Dinesh Kumar Tyagi, Santosh Kumar Vipparthi, M. Subrahmanyam 0001, Mohamed Abdel-Mottaleb
Knowl. Based Syst.6
2026 ME-NAS: A Micro Expression Feature Adaptive Neural Architecture Search
abstract
Convolution neural networks (CNN) have emerged as a prevailing paradigm for micro-expression recognition (MER) yet, it is inefficient and time-intensive to design optimal CNN-based MER models manually. In recent times, the neural architecture search (NAS) has garnered attention due to its automatic CNN architecture searching ability. However, the performance of NAS in MER is limited by challenges such as rapid duration, subtle intensity, and a mismatch between architecture and cell-level search. The existing search space, which stacks 12 cells with 3 transition paths (downsample, upsample, and same resolution), creates deep networks that may diminish minute spatiotemporal features due to progressive convolution and pooling. Therefore, motivated by these factors, in this article, we introduce a novel approach, the Micro-Expression Feature Adaptive NAS (ME-NAS), to analyze true human emotions through MER. While NAS has gained attention for its automatic CNN architecture search ability, its application in MER faces challenges due to ingrained challenges (rapid duration, subtle and low intensity) and the discrepancy between architecture and cell-level search. The existing NAS architecture search space is designed by stacking 12 cells with 3 transition paths (downsample, upsample, and same resolution), resulting in a deep network. Such deep networks may diminish minute spatiotemporal features due to the progressive convolution and pooling operations. Motivated by these factors, we designed a new NAS algorithm: ME-NAS. The ME-NAS comprises f (EXPERT) in architecture search, along with refined and complementary feature derivative (ReCODE) operations in cell-level search. The EXPERT aims to trace the optimal paths instead of covering all possible paths between cells. The ReCODE operations capture micro-level variations from spatial and temporal domains by introducing 24 3D convolution operations. The proposed ReCODE and EXPERT search space jointly lead to the search for a robust and shallow CNN architecture for micro-expressions (MEs). The proposed ME-NAS is evaluated on six datasets: CASME-I, CASME-II, CAS(ME) \({}^{2}\) , SAMM, SMIC, and MEGC-19 composite, with two evaluation strategies: LOSO and cross-domain, respectively. The experimental results manifest that the proposed ME-NAS outperformed the state-of-the-art approaches on both evaluation strategies.
Monu Verma, Santosh Kumar Vipparthi, M. Subrahmanyam 0001, Mohamed Abdel-Mottaleb
ACM Trans. Intell. Syst. Technol.4
2023 RNAS-MER: A Refined Neural Architecture Search with Hybrid Spatiotemporal Operations for Micro-Expression Recognition
abstract
Existing neural architecture search (NAS) methods comprise linear connected convolution operations and use ample search space to search task-driven convolution neural networks (CNN). These CNN models are computationally expensive and diminish the quality of receptive fields for tasks like micro-expression recognition (MER) with limited training samples. Therefore, we propose a refined neural architecture search strategy to search for a tiny CNN architecture for MER. In addition, we introduced a refined hybrid module (RHM) for inner-level search space and an optimal path explore network (OPEN) for outer-level search space. The RHM focuses on discovering optimal cell structures by incorporating a multilateral hybrid spatiotemporal operation space. Also, spatiotemporal attention blocks are embedded to refine the aggregated cell features. The OPEN search space aims to trace an optimal path between the cells to generate a tiny spatiotemporal CNN architecture instead of covering all possible tracks. The aggregate mix of RHM and OPEN search space availed the NAS method to robustly search and design an effective and efficient framework for MER. Compared with contemporary works, experiments reveal that the RNAS-MER is capable of bridging the gap between NAS algorithms and MER tasks. Furthermore, RNAS-MER achieves new state-of-the-art performances on challenging MER benchmarks, including 0.8511%, 0.7620%, 0.9078% and 0.8235% UAR on COMPOSITE, SMIC, CASME-II and SAMM datasets respectively.
Monu Verma, Priyanka Lubal, Santosh Kumar Vipparthi, Mohamed Abdel-Mottaleb
WACV4
2023 Difference-of-Gaussian generative adversarial network for segmenting breast arterial calcifications in mammograms
Manal AlAmir, Manal Al Ghamdi, Fernando Collado-Mesa, Mohamed Abdel-Mottaleb
Expert Syst. Appl.4
2022 Automated Ensemble for Deep Learning Inference on Edge Computing Platforms
abstract
Advances in deep learning (DL) have triggered an explosion of mobile intelligence, posing a soaring demand for computing resources that cannot be satisfied by mobile devices. In this article, we employedge computingto deliver better DL inference services to end users. The key is to leverage deep neural network (DNN) ensemble techniques that provide state-of-the-art performance for many machine learning applications in terms of inference accuracy and robustness. Compared to end devices, the edge computing platform is endowed with more powerful computing resources, making it feasible to implement DNN ensembles for DL inferences. However, due to the constrained computing capacity of edge servers and the possible service response deadline, an edge server can only use a limited number of DNNs to construct DNN ensembles. This poses a unique problem, namely, DNN ensemble selection, for identifying the best-fit DNN ensembles. We propose a novel algorithm called automated DNN ensemble selection (AES) algorithm to solve this problem. Because DNNs exhibit performance variations over different distributions of input data, AES adaptively determines a DNN ensemble according to the features of admitted inference tasks. AES is an online learning algorithm that learns DNNs’ in-use performance over time. An ensemble selection rule is further designed as a subroutine of AES to recruit members into the DNN ensemble based on the accuracy and diversity of DNNs. In particular, we theoretically prove that AES can achieve asymptotic optimality. We carry out experiments on real-world data sets. The results show that using the DNN ensemble technique on edge computing platforms dramatically improves the DL inference quality, and AES outperforms other benchmark schemes.
Yang Bai 0010, Lixing Chen, Mohamed Abdel-Mottaleb, Jie Xu 0001
IEEE Internet Things J.3
2022 Towards a better understanding of annotation tools for medical imaging: a survey
Manar Aljabri, Manal AlAmir, Manal Al Ghamdi, Mohamed Abdel-Mottaleb, Fernando Collado-Mesa
Multim. Tools Appl.4
2021 Adaptive Deep Neural Network Ensemble for Inference-as-a-Service on Edge Computing Platforms
abstract
The momentous enabling of deep learning (DL)-powered mobile application is posing a soaring demand for computing resources that can hardly be satisfied by mobile devices. In this paper, we employ Edge Computing to deliver DL inference services to mobile users, where Deep Neural Networks (DNNs) are configured on edge servers, processing inference tasks received from mobile devices. A novel method called Adaptive DNN Ensemble (ADE) is proposed to enhance the performance of DL inference services. The core of ADE is the DNN ensemble technique which improves the stability and accuracy of DL inference. Due to the limited computing resources and service response deadline, ADE needs to judiciously determine DNNs to be included in the DNN ensemble, which poses a unique DNN ensemble selection problem. In addition, because DNNs exhibit performance variations for tasks with different features, DNN ensemble selection also aims to reconFigure DNN ensembles according to the feature of admitted tasks. We design an online learning algorithm, Contextual Combinatorial Multi-Armed Bandit (CC-MAB), to learn the DNN performance for tasks with different features. We rigorously prove that the proposed online learning algorithm is able to achieve asymptotic optimality. Experiments are carried out on an edge computing testbed to evaluate our method. Various implementation concerns, including memory usage, time complexity, and DNN switching cost, are considered. The results show that ADE outperforms other benchmarks in terms of inference accuracy and can provide real-time responses.
Yang Bai 0010, Lixing Chen, Mohamed Abdel-Mottaleb, Jie Xu 0001
MASS4
2021 Correcting Higher Order Aberrations Using Image Processing
abstract
Higher Order Aberrations (HOAs) are complex refractive errors in the human eye that cannot be corrected by regular lens systems. Researchers have developed numerous approaches to analyze the effect of these refractive errors; the most popular among these approaches use Zernike polynomial approximation to describe the shape of the wavefront of light exiting the pupil after it has been altered by the refractive errors. We use this wavefront shape to create a linear imaging system that simulates how the eye perceives source images at the retina. With phase information from this system, we create a second linear imaging system to modify source images so that they would be perceived by the retina without distortion. By modifying source images, the visual process cascades two optical systems before the light reaches the retina, a technique that counteracts the effect of the refractive errors. While our method effectively compensates for distortions induced by HOAs, it also introduces blurring and loss of contrast; a problem that we address with Total Variation Regularization. With this technique, we optimize source images so that they are perceived at the retina as close as possible to the original source image. To measure the effectiveness of our methods, we compute the Euclidean error between the source images and the images perceived at the retina. When comparing our results with existing corrective methods that use deconvolution and total variation regularization, we achieve an average of 50% reduction in error with lower computational costs.
Olga E. Jumbo, Shihab S. Asfour, Ahmed M. Sayed, Mohamed Abdel-Mottaleb
IEEE Trans. Image Process.4
2021 3DCD: Scene Independent End-to-End Spatiotemporal Feature Learning Framework for Change Detection in Unseen Videos
abstract
Change detection is an elementary task in computer vision and video processing applications. Recently, a number of supervised methods based on convolutional neural networks have reported high performance over the benchmark dataset. However, their success depends upon the availability of certain proportions of annotated frames from test video during training. Thus, their performance on completely unseen videos or scene independent setup is undocumented in the literature. In this work, we present a scene independent evaluation (SIE) framework to test the supervised methods in completely unseen videos to obtain generalized models for change detection. In addition, a scene dependent evaluation (SDE) is also performed to document the comparative analysis with the existing approaches. We propose a fast (speed-25 fps) and lightweight (0.13 million parameters, model size-1.16 MB) end-to-end 3D-CNN based change detection network (3DCD) with multiple spatiotemporal learning blocks. The proposed 3DCD consists of a gradual reductionist block for background estimation from past temporal history. It also enables motion saliency estimation, multi-schematic feature encoding-decoding, and finally foreground segmentation through several modular blocks. The proposed 3DCD outperforms the existing state-of-the-art approaches evaluated in both SIE and SDE setup over the benchmark CDnet 2014, LASIESTA and SBMI2015 datasets. To the best of our knowledge, this is a first attempt to present results in clearly defined SDE and SIE setups in three change detection datasets.
Murari Mandal, Vansh Dhar, Santosh Kumar Vipparthi, Mohamed Abdel-Mottaleb
IEEE Trans. Image Process.5
2020 Multi-Resolution Overlapping Stripes Network for Person Re-Identification
abstract
This paper addresses the person re-identification (PReID) problem by combining global and local information at multiple feature resolutions with different loss functions. Many previous studies address this problem using either part-based features or global features. In case of part-based representation, the spatial correlation between these parts is not considered, while global-based representation are not sensitive to spatial variations. This paper presents a part-based model with a multi-resolution network that uses different level of features. The output of the last two conv blocks is then partitioned horizontally and processed in pairs with overlapping stripes to cover the important information that might lie between parts. We use different loss functions to combine local and global information for classification. Experimental results on a benchmark dataset demonstrate that the presented method outperforms the state-of-the-art methods.
Arda Efe Okay, Manal Al Ghamdi, Robert Westendorp, Mohamed Abdel-Mottaleb
ICASSP4
2020 DU-Net: Convolutional Network for the Detection of Arterial Calcifications in Mammograms
abstract
Breast arterial calcifications (BACs) are part of several benign findings present on some mammograms. Previous studies have indicated that BAC may provide evidence of general atherosclerotic vascular disease, and potentially be a useful marker of cardiovascular disease (CVD). Currently, there is no technique in use for the automatic detection of BAC in mammograms. Since a majority of women over the age of 40 already undergo breast cancer screening with mammography, detecting BAC may offer a method to screen women for CVD in a way that is effective, efficient, and broad reaching, at no additional cost or radiation. In this paper, we present a deep learning approach for detecting BACs in mammograms. Inspired by the promising results achieved using the U-Net model in many biomedical segmentation problems and the DenseNet in semantic segmentation, we extend the U-Net model with dense connectivity to automatically detect BACs in mammograms. The presented model helps to facilitate the reuse of computation and improve the flow of gradients, leading to better accuracy and easier training of the model. We evaluate the performance using a set of full-field digital mammograms collected and prepared for this task from a publicly available dataset. Experimental results demonstrate that the presented model outperforms human experts as well as the other related deep learning models. This confirms the effectiveness of our model in the BACs detection task, which is a promising step in providing a cost-effective risk assessment tool for CVD.
Manal Al Ghamdi, Mohamed Abdel-Mottaleb, Fernando Collado-Mesa
IEEE Trans. Medical Imaging2
2019 Semi-supervised Transfer Learning for Convolutional Neural Networks for Glaucoma Detection
abstract
Convolutional neural network (CNN) can be applied in glaucoma detection for achieving good performance. However, its performance depends on the availability of a large number of the labelled samples for its training phase. To solve this problem, this paper present a semi-supervised transfer learning CNN model for automatic glaucoma detection based on both labeled and unlabeled data. First, a pre-trained CNN from non-medical data is fine-tuned and trained in a supervised fashion using the labeled data. The self-learning approach is then used to predict the labels for the unlabeled data and utilize it for training. The experimental results on the RIM-ONE database demonstrate the effectiveness of the proposed algorithm despite the lack of initial labeled samples.
Manal Al Ghamdi, Mohamed Abdel-Mottaleb, Mohamed Abou Shousha
ICASSP3
2018 Kernel Discriminant Correlation Analysis: Feature Level Fusion for Nonlinear Biometric Recognition
abstract
In biometric recognition, feature fusion is an important area of research due to the fact that multiple types of features contain richer and complementary information. Discriminative Correlation Analysis (DCA) is a recently proposed feature fusion method, which incorporates the class association into correlation analysis so that the features not only have the maximum intrinsic correlation between feature sets but also have class structure information. However, DCA is a linear technique, that finds a linear transformation of the original space. For highly nonlinearly distributed data, classification with nonlinear techniques works better than the linear ones. In this paper, we propose Kernel-DCA which generalizes DCA in order to handle nonlinear problems. Similar to Kernel-SVM, Kernel-DCA utilizes a kernel method to map feature sets to a high-dimensional space in which features are linearly separable. Experimental results, for the fusion of ear and face feature, using the WVU database with large variations in pose, show that Kernel-DCA achieves better results on nonlinearly distributed data than DCA and other feature fusion methods.
Mohammad Haghighat, Mohamed Abdel-Mottaleb
ICPR3
2017 Lower Resolution Face Recognition in Surveillance Systems Using Discriminant Correlation Analysis
abstract
Due to large distances between surveillance cameras and subjects, the captured images usually have low resolution in addition to uncontrolled poses and illumination conditions that adversely affect the performance of face recognition algorithms. In this paper, we present a low-resolution face recognition technique based on Discriminant Correlation Analysis (DCA). DCA analyzes the correlation of the features in high-resolution and low-resolution images and aims to find projections that maximize the pair-wise correlations between the two feature sets and at the same time, separate the classes within each set. This makes it possible to project the features extracted from high-resolution and low-resolution images into a common space, in which we can apply matching. The proposed method is computationally efficient and can be applied to challenging real-time applications such as recognition of several faces appearing in a crowded frame of a surveillance video. Extensive experiments performed on low-resolution surveillance images from the SCface database as well as FRGC database demonstrated the efficacy of our proposed approach in the recognition of low-resolution face images, which outperformed other state-of-the-art techniques.
Mohammad Haghighat, Mohamed Abdel-Mottaleb
FG2
2016 Discriminant correlation analysis for feature level fusion with application to multimodal biometrics
abstract
In this paper, we present Discriminant Correlation Analysis (DCA), a feature level fusion technique that incorporates the class associations in correlation analysis of the feature sets. DCA performs an effective feature fusion by maximizing the pair-wise correlations across the two feature sets, and at the same time, eliminating the between-class correlations and restricting the correlations to be within classes. Our proposed method can be used in pattern recognition applications for fusing features extracted from multiple modalities or combining different feature vectors extracted from a single modality. It is noteworthy that DCA is the first technique that considers class structure in feature fusion. Moreover, it has a very low computational complexity and it can be employed in realtime applications. Multiple sets of experiments performed on various biometric databases show the effectiveness of our proposed method, which outperforms other state-of-the-art approaches.
Mohammad Haghighat, Mohamed Abdel-Mottaleb, Wadee Alhalabi
ICASSP2
2016 Fully automatic face normalization and single sample face recognition in unconstrained environments
Mohammad Haghighat, Mohamed Abdel-Mottaleb, Wadee Alhalabi
Expert Syst. Appl.2
2016 Discriminant Correlation Analysis: Real-Time Feature Level Fusion for Multimodal Biometric Recognition
abstract
Information fusion is a key step in multimodal biometric systems. The fusion of information can occur at different levels of a recognition system, i.e., at the feature level, matching-score level, or decision level. However, feature level fusion is believed to be more effective owing to the fact that a feature set contains richer information about the input biometric data than the matching score or the output decision of a classifier. The goal of feature fusion for recognition is to combine relevant information from two or more feature vectors into a single one with more discriminative power than any of the input feature vectors. In pattern recognition problems, we are also interested in separating the classes. In this paper, we present discriminant correlation analysis (DCA), a feature level fusion technique that incorporates the class associations into the correlation analysis of the feature sets. DCA performs an effective feature fusion by maximizing the pairwise correlations across the two feature sets and, at the same time, eliminating the between-class correlations and restricting the correlations to be within the classes. Our proposed method can be used in pattern recognition applications for fusing the features extracted from multiple modalities or combining different feature vectors extracted from a single modality. It is noteworthy that DCA is the first technique that considers class structure in feature fusion. Moreover, it has a very low computational complexity and it can be employed in real-time applications. Multiple sets of experiments performed on various biometric databases and using different feature extraction techniques, show the effectiveness of our proposed method, which outperforms other state-of-the-art approaches.
Mohammad Haghighat, Mohamed Abdel-Mottaleb, Wadee Alhalabi
IEEE Trans. Inf. Forensics Secur.2
2016 Automatic Ear Landmark Localization, Segmentation, and Pose Classification in Range Images
abstract
Multibiometric systems using face and ear features are increasingly adopted for forensic and civilian applications to address the challenges of facial expressions and occlusions. Although numerous ear and face recognition techniques have been proposed, not much work has been conducted in the field of 3-D fiducial points localization and 3-D ear detection. This paper presents an effective and efficient system of ear landmark localization, ear detection, and pose classification based on 3-D ears captured under large yaw variations. By utilizing the symmetrical property of human heads and classifying the ear with respect to its pose, all three tasks can be fulfilled given either left or right ears, without any prior pose information. A novel ear tree-structured graph (ETG) is proposed to represent the 3-D ear, after which a 3-D flexible mixture model is trained to locate the landmarks automatically. Then, the ear region is segmented based on them and the pose of the ear, i.e., whether it is a left or right ear, is classified based on the detected ETG. To the best of our knowledge, this paper is the first to present automatic landmark localization of 3-D ears extracted from facial scans with significant pose variations. Experiments were conducted at the University of Notre Dame collection F, G and J2, which contain large occlusion and pose variations, validating the effectiveness of the proposed methods.
Jiajia Lei, Xinge You, Mohamed Abdel-Mottaleb
IEEE Trans. Syst. Man Cybern. Syst.3
2015 Classification based on weighted sparse representation using smoothed L0 norm with non-negative coefficients
abstract
We present a novel classification technique based on sparse representation. The main idea of sparse representation for classification is the assumption that the training samples, or atoms, for a particular class form a linear basis for any new test sample that belongs to that class. Currently, most of the methods for sparse representation classification do not apply constraints to the coefficients that form the linear combination of the atoms, which leads to coefficients that can be positive or negative. In addition, all the training samples in the dictionary are treated equally. In this paper, we impose non-negative constraint on the components of the coefficient vector to ensure that the coefficient vector represents the contributions of the training samples towards the query, which is more natural for classification purposes. We also use the mutual information between the query sample and each of the training samples to obtain a weight for each of the atoms in the dictionary. These weights have the effect of reducing the search space and speeding the convergence of the algorithm in finding the coefficient vector. Experiments conducted on the Extended Yale B database for face recognition and on the University of Notre Dame (UND) database for ear recognition show that the proposed non-negative weighted sparse representation obtained by smoothed l0norm outperforms other state-of-the-art classifiers.
Rahman Khorsandi, Mohamed Abdel-Mottaleb
ICIP2
2015 CloudID: Trustworthy cloud-based and cross-enterprise biometric identification
Mohammad Haghighat, Saman A. Zonouz, Mohamed Abdel-Mottaleb
Expert Syst. Appl.3
2015 Multimedia event detection with ℓ2-regularized logistic Gaussian mixture regression
Changyu Liu, Shoubin Dong, Mohamed Abdel-Mottaleb
Neural Comput. Appl.4
2015 Unsupervised segmentation of highly dynamic scenes through global optimization of multiscale cues
Yinhui Zhang, Mohamed Abdel-Mottaleb, Zifen He
Pattern Recognit.2
2015 3D Ear Segmentation and Classification Through Indexing
abstract
Current growth trends in different biometrics applications present challenges to researchers. To address these challenges, we need new data storage and retrieval techniques to make the recognition process time efficient. This paper presents a system for time efficient 3D ear biometrics from a large biometrics database. The proposed system has two components that are primarily responsible for: 1) automatic 3D ear segmentation and 2) hierarchical categorization of the 3D ear database using the shape information and surface depth information, respectively. We use an active contour algorithm along with a tree-structured graph to segment the ear region from the 3D profile images. The segmented 3D ear database is then categorized based on the geometrical feature values, computed from the ear shape, into oval, round, rectangular, and triangular categories. For the categorization based on the depth information, the feature space is partitioned using tree-based indexing techniques. We used indexing techniques with balanced split (k-dimensional (KD) tree) and unbalanced split (pyramid tree) data structures to categorize the database separately and then compared their retrieval efficiency. Experiments are conducted to compare the average computation time per query when performing recognition through hierarchical categorization with the average computation time when recognition is based on sequential search. Experimental results without indexing conducted on the University of Notre Dame Collection J2 data set yielded a rank-one recognition rate of 98.5%. We applied the indexing technique to compute the rank-one recognition accuracy with a 10%-50% search space reduction using a 10% step size. With 10%, 20%, 30%, 40%, and 50% search space reduction the rank-one recognition accuracy gracefully degrades to 96.87%, 96.14%, 95.18%, 94.21%, and 93.49%, respectively, while performing nearly 3, 3.3, 4, 4.2, and 5 times faster than the state-of-the-art technique that uses sequential search.
Sayan Maity, Mohamed Abdel-Mottaleb
IEEE Trans. Inf. Forensics Secur.2
2014 Ear Biometrics and Sparse Representation Based on Smoothed L0 norm
abstract
Ear biometrics attracted the attention of researchers in computer vision and machine learning for its use in many applications. In this paper, we present a fully automated system for recognition from ear images based upon sparse representation. In sparse representation, extracted features from the training data is used to develop a dictionary. Classification is achieved by representing the extracted features of the test data as a linear combination of entries in the dictionary. In fact, there are many solutions for this problem and the goal is to find the sparsest solution. We use a relatively new algorithm named smoothed l0 norm to find the sparsest solution and Gabor wavelet features are used for building the dictionary. Furthermore, we expand the proposed approach for gender classification from ear images. Several researches have addressed this issue based on facial images. We introduce a novel approach based on majority voting for gender classification. Experimental results conducted on the University of Notre Dame (UND) collection J data set, containing large appearance, pose, and lighting variations, resulted in a gender classification rate of 89.49%. Furthermore, the proposed method is evaluated on the WVU data set and classification rates for different view angles are presented. Results show improvement and great robustness in gender classification over existing methods.
Rahman Khorsandi, Mohamed Abdel-Mottaleb
Int. J. Pattern Recognit. Artif. Intell.2
2013 Identification Using Encrypted Biometrics
Mohammad Haghighat, Saman A. Zonouz, Mohamed Abdel-Mottaleb
CAIP (2)3
2013 Gender Classification Using Facial Images and Basis Pursuit
Rahman Khorsandi, Mohamed Abdel-Mottaleb
CAIP (1)2
2013 A novel shape-based interest point descriptor (SIP) for 3D ear recognition
abstract
In this paper, we introduce a novel shape-based interest point (SIP) descriptor to encode local surface shapes for three-dimensional (3D) ear recognition; the descriptor provides an advantage over previous descriptors by capturing greater details of the macro-shape patterns surrounding an interest point. Using the SIP descriptor, a function is developed to measure the shape dissimilarity between any two interest points. Finally, in the recognition stage, a probe and a gallery pair are compared by applying the matching algorithm on the interest points, with the similarity score set as the number of matched interest points. The proposed method has been tested on the University of Notre Dame(UND) collection J2 dataset, containing range images of 415 subjects. The experimental results demonstrate that our method achieves a 97.4% rank-one recognition rate and a 2.0% Equal Error Rate (EER), which outperforms the state-of-the-art methods.
Jiajia Lei, Jindan Zhou, Mohamed Abdel-Mottaleb
ICIP3
2013 Detection, localization and pose classification of ear in 3D face profile images
abstract
We present an efficient and robust system for landmark localization, segmentation and pose classification of ears from 3D profile facial range data. After defining 18 landmarks on the ear, including Triangular Fossa and Incisure Intertragica, a novel Ear Tree-structured Graph (ETG) is proposed to represent the 3D ear. We trained a flexible mixture model to locate these landmarks automatically. Afterwards, the ear region is outlined as the minimum rectangle including all landmarks. Finally, by calculating the turning angle between landmarks on the helix, the ear is classified as either a left or a right ear. To the best of our knowledge, there is no previous work on automatic landmark localization for 3D ear on 3D facial profile depth images. Experiments are conducted on University of Notre Dame Collection F and Collection J2 datasets, containing large occlusion, scale and pose variations. Results demonstrate the effectiveness of the proposed techniques.
Jiajia Lei, Jindan Zhou, Mohamed Abdel-Mottaleb, Xinge You
ICIP3
2013 Gender classification using 2-D ear images and sparse representation
abstract
Gender classification attracted the attention of researchers in computer vision for its use in many applications. Researches have addressed this issue based on facial images. In this paper, we present the first approach for gender classification using 2-D ear images based upon sparse representation. In sparse representation, the training data is used to develop a dictionary based on extracted features. In this work, Gabor filters are used for feature extraction. Classification is achieved by representing the test data using the dictionary based upon the extracted features. Experimental results conducted on the University of Notre Dame (UND) collection J dataset, containing large appearance, pose, and lighting variability, yielded gender classification rate of 89.49%.
Rahman Khorsandi, Mohamed Abdel-Mottaleb
WACV2
2012 Exploiting visual quasi-periodicity for real-time chewing event detection using active appearance models and support vector machines
Steven Cadavid, Mohamed Abdel-Mottaleb, Abdelsalam Helal
Pers. Ubiquitous Comput.2
2012 An Efficient 3-D Ear Recognition System Employing Local and Holistic Features
abstract
We present a complete three-dimensional (3-D) ear recognition system combining local and holistic features in a computationally efficient manner. The system is comprised of four primary components, namely: 1) ear image segmentation; 2) local feature extraction and matching; 3) holistic feature extraction and matching; and 4) a fusion framework combining local and holistic features at the match score level. For the segmentation component, we introduce a novel shape-based feature set, termed the Histograms of Indexed Shapes (HIS), to localize a rectangular region containing the ear. For the local feature extraction and representation component, we extend the HIS feature descriptor to an object-centered 3-D shape descriptor, the Surface Patch Histogram of Indexed Shapes (SPHIS), for local ear surface representation and matching. For the holistic matching component, we introduce a voxelization scheme for holistic ear representation from which an efficient, voxel-wise comparison of gallery-probe model pairs can be made. The match scores obtained from both the local and holistic matching components are fused to generate the final match scores. Experimental results conducted on the University of Notre Dame (UND) collection G dataset, containing range images of 415 subjects yielded a rank-one recognition rate of 98.3% and an equal error rate of 1.7%. These results demonstrate that the proposed approach outperforms state-of-the-art 3-D ear biometric systems. Additionally, the method is considerably more efficient compared to the state-of-the-art because it employs a sparse set of features rather than using the dense model.
Jindan Zhou, Steven Cadavid, Mohamed Abdel-Mottaleb
IEEE Trans. Inf. Forensics Secur.3
2011 An adaptive resolution voxelization framework for 3D ear recognition
abstract
We present a novel voxelization framework for holistic Three-Dimensional (3D) object representation that accounts for distinct surface features. A voxelization of an object is performed by encoding an attribute or set of attributes of the surface region contained within each voxel occupying the space that the object resides in. To our knowledge, the voxel structures employed in previous methods consist of uniformly-sized voxels. The proposed framework, in contrast, generates structures consisting of variable-sized voxels that are adoptively distributed in higher concentration near distinct surface features. The primary advantage of the proposed method over its fixed resolution counterparts is that it yields a significantly more concise feature representation that is demonstrated to achieve a superior recognition performance. An evaluation of the method is conducted on a 3D ear recognition task. The ear provides a challenging case study be- cause of its high degree of inter-subject similarity.
Steven Cadavid, Sherin Fathy, Jindan Zhou, Mohamed Abdel-Mottaleb
IJCB4
2011 Exploiting color SIFT features for 2D ear recognition
abstract
In this paper, we present a robust method for 2D ear recognition using color SIFT features. Firstly, we extend the Scale Invariant Feature Transform (SIFT) algorithm originally performed on the intensity channel [1] to the RGB color channels to maximize the robustness of the SIFT feature descriptor. Secondly, a feature matching algorithm for ear recognition is proposed by fusion of the features extracted from the different color channels. Experiments conducted on the University of Notre Dame (UND) and the West Virginia University (WVU) ear biometric datasets indicate that our method can achieve better recognition rates than the state-of-the-art methods applied on the same datasets.
Jindan Zhou, Steven Cadavid, Mohamed Abdel-Mottaleb
ICIP3
2010 Exploiting Visual Quasi-periodicity for Automated Chewing Event Detection Using Active Appearance Models and Support Vector Machines
abstract
We present a method that automatically detects chewing events in surveillance video of a subject. Firstly, an Active Appearance Model (AAM) is used to track a subject's face across the video sequence. It is observed that the variations in the AAM parameters across chewing events demonstrate a distinct periodicity. We utilize this property to discriminate between chewing and non-chewing facial actions such as talking. A feature representation is constructed by applying spectral analysis to a temporal window of model parameter values. The estimated power spectra subsequently undergo non-linear dimensionality reduction via spectral regression. The low-dimensional representations of the power spectra are employed to train a Support Vector Machine (SVM) binary classifier to detect chewing events. Experimental results yielded a cross validated percentage agreement of 93.4%, indicating that the proposed system provides an efficient approach to automated chewing detection.
Steven Cadavid, Mohamed Abdel-Mottaleb
ICPR2
2009 Detecting Local Audio-visual Synchrony in Monologues Utilizing Vocal Pitch and Facial Landmark Trajectories
abstract
We describe a novel approach for determining the audio-visual synchrony of a monologue video sequence utilizing vocal pitch and facial landmark trajectories as descriptors of the audio and visual modalities, respectively. The visual component is represented by the horizontal and vertical displacement of corresponding facial landmarks between subsequent frames. These facial landmarks are acquired using the statistical modeling technique, known as the Active Shape Model (ASM). The audio component is represented by the fundamental frequency, or pitch, obtained using the subharmonic-to-harmonic ratio (SHR). The synchrony between the audio and visual feature vectors is computed using Gaussian mutual information. The raw synchrony estimates obtained using this method may contain spurious synchrony values due to over-sensitivity. A filtering method is employed for discarding synchrony values that occur during non-associated audio/visual events. The human visual system is capable of distinguishing rigid and non-rigid motion of an articulator during speech. In an attempt to emulate this process, we separate rigid and non-rigid motion and compute the synchrony attributed to each. Experiments are conducted on a dataset of monologue video clip pairs. Each pair is composed of an asynchronous and synchronous version of the video clip. For the asynchronous video clips, the audio signal is displaced with respect to the visual signal. Experimental results indicate that the proposed approach is successful in detecting facial regions that demonstrate synchrony, and in distinguishing between synchronous and asynchronous sequences. © 2009. The copyright of this document resides with its authors.
Steven Cadavid, Mohamed Abdel-Mottaleb, Daniel S. Messinger, Mohammad H. Mahoor, Lorraine E. Bahrick
BMVC2
2009 Determining discriminative anatomical point pairings using adaboost for 3D face recognition
abstract
In this paper, we present a novel method for 3D face recognition using adaboosted geodesic distance features. Firstly, a generic model is finely conformed to each face model contained within a 3D face dataset. Secondly, the geodesic distance between anatomical point pairs are computed across each conformed generic model. Adaboost then generates a strong-classifier based on a collection of geodesic distances that are most discriminative for face recognition. Experiments conducted on the face recognition grand challenge (FRGC) database D collection indicate that the system can achieve over a 95% rank-one recognition rate.
Steven Cadavid, Jindan Zhou, Mohamed Abdel-Mottaleb
ICIP3
2009 Multi-modal ear and face modeling and recognition
abstract
In this paper we describe a multi-modal ear and face biometric system. The system is comprised of two components: a 3D ear recognition component and a 2D face recognition component. For the 3D ear recognition, a series of frames is extracted from a video clip and the region of interest (i.e., ear) in each frame is independently reconstructed in 3D using Shape From Shading. The resulting 3D models are then registered using the iterative closest point algorithm. We iteratively consider each model in the series as a reference model and calculate the similarity between the reference model and every model in the series using a similarity cost function. Cross validation is performed to assess the relative fidelity of each 3D model. The model that demonstrates the greatest overall similarity is determined to be the most stable 3D model and is subsequently enrolled in the database. For the 2D face recognition, a set of facial landmarks is extracted from frontal facial images using the Active Shape Model. Then, the response of facial images to a series of Gabor filters at the locations of facial landmarks are calculated. The Gabor features (attributes) are stored in the database as the face model for recognition. The similarity between the Gabor features of a probe facial image and the reference models are utilized to determine the best match. The match scores of the ear recognition and face recognition modalities are fused to boost the overall recognition rate of the system. Experiments are conducted using a gallery set of 402 video clips and a probe of 60 video clips (images). As a result, a rank-one identification rate of 100% was achieved using the weighted sum technique for fusion.
Mohammad H. Mahoor, Steven Cadavid, Mohamed Abdel-Mottaleb
ICIP3
2009 A multimodal approach for 3D face modeling and recognition using 3D deformable facial mask
A-Nasser Ansari, Mohamed Abdel-Mottaleb, Mohammad H. Mahoor
Mach. Vis. Appl.2
2009 Face recognition based on 3D ridge images obtained from range data
Mohammad H. Mahoor, Mohamed Abdel-Mottaleb
Pattern Recognit.2
2008 Multi-modal (2-D and 3-D) face modeling and recognition using Attributed Relational Graph
abstract
In this paper we present a unified graph model, called Attributed Relational Graph (ARG), for multi-modal face modeling and recognition. Based on the ARG model, the 2-D and 3-D data are included in a single model. The developed ARG model consists of nodes, edges, and mutual relations. The nodes of the graph correspond to the landmark points that are extracted by an improved Active Shape Model (ASM) technique. Then, at each node of the graph, the responses of a set of log-Gabor filters to the facial image texture and shape information (depth values) are calculated; the filter responses are used to model the local structure of the face at each node of the graph. The edges of the graph are defined based on Delaunay triangulation and a set of mutual relations between the sides of the triangles are defined. The mutual relations boost the final performance of the system. The results of face matching using the 2-D and 3-D attributes and the mutual relations are fused at the score level. A rank-one identification rate of 99% is achieved by experimenting on the University of Miami face database.
Mohammad H. Mahoor, A-Nasser Ansari, Mohamed Abdel-Mottaleb
ICIP3
2008 3D ear modeling and recognition from video sequences using shape from shading
abstract
We describe a novel approach for 3D ear biometrics using video. A series of frames are extracted from a video clip and the region-of-interest (ROI) in each frame is independently reconstructed in 3D using Shape from Shading (SFS). The resulting 3D models are then registered using the Iterative Closest Point (ICP) algorithm. We iteratively consider each model in the series as a reference and calculate the similarity between the reference model and every model in the series using a similarity cost function. Cross validation is performed to assess the relative fidelity of each 3D model. The model that demonstrates the greatest overall similarity is determined to be the most stable 3D model and is subsequently enrolled in the database. Experiments are conducted using a gallery set of 402 video clips and a probe of 60 video clips. The results (95.0% rank-1 recognition rate) indicate that the proposed approach can produce recognition rates comparable to systems that use 3D range data. To the best of our knowledge, we are the first to develop a 3D ear biometric system that obtains 3D ear structure from a video sequence.
Steven Cadavid, Mohamed Abdel-Mottaleb
ICPR2
2008 Hierarchical contour matching for dental X-ray radiographs
Omaima Nomir, Mohamed Abdel-Mottaleb
Pattern Recognit.2
2008 3-D Ear Modeling and Recognition From Video Sequences Using Shape From Shading
abstract
We describe a novel approach for 3-D ear biometrics using video. A series of frames is extracted from a video clip and the region of interest in each frame is independently reconstructed in 3-D using shape from shading. The resulting 3-D models are then registered using the iterative closest point algorithm. We iteratively consider each model in the series as a reference model and calculate the similarity between the reference model and every model in the series using a similarity cost function. Cross validation is performed to assess the relative fidelity of each 3-D model. The model that demonstrates the greatest overall similarity is determined to be the most stable 3-D model and is subsequently enrolled in the database. Experiments are conducted using a gallery set of 402 video clips and a probe of 60 video clips. The results (95.0% rank-1 recognition rate and 3.3% equal error rate) indicate that the proposed approach can produce recognition rates comparable to systems that use 3-D range data. To the best of our knowledge, we are the first to develop a 3-D ear biometric system that obtains a 3-D ear structure from a video sequence.
Steven Cadavid, Mohamed Abdel-Mottaleb
IEEE Trans. Inf. Forensics Secur.2
2008 A Multimodal Approach for Face Modeling and Recognition
abstract
In this paper, we present a fully automated multi- modal (3-D and 2-D) face recognition system. For the 3-D modality, we model the facial image as a 3-D binary ridge image that contains the ridge lines on the face. We use the principal curvature to extract the locations of the ridge lines around the important facial regions on the range image (i.e., the eyes, the nose, and the mouth.) For matching, we utilize a fast variant of the iterative closest point to match the ridge image of a given probe image to the archived ridge images in the database. The main advantage of this approach is reducing the computational complexity by two orders of magnitude by relying on the ridge lines. For the 2-D modality, we model the face by an attributed relational graph (ARG), where each node of the graph corresponds to a facial feature point. At each facial feature point, a set of attributes is extracted by applying Gabor wavelets to the 2-D image and assigned to the node of the graph. The edges of the graph are defined based on Delaunay triangulation and a set of geometrical features that defines the mutual relations between the edges is extracted from the Delaunay triangles and stored in the ARG model. The similarity measure between the ARG models that represent the probe and gallery images is used for 2-D face recognition. Finally, we fuse the matching results of the 3-D and the 2-D modalities at the score level to improve the overall performance of the system. Different techniques for fusion, such as the Dempster-Shafer theory of evidence and weighted sum of scores are employed and tested using the facial images in the third experiment dataset of the Face Recognition Grand Challenge version 2.0.
Mohammad H. Mahoor, Mohamed Abdel-Mottaleb
IEEE Trans. Inf. Forensics Secur.2
2008 Fusion of Matching Algorithms for Human Identification Using Dental X-Ray Radiographs
abstract
The goal of forensic dentistry is to identify individuals based on their dental characteristics. In this paper, we introduce a system that uses some scenarios to fuse three matching techniques for identifying individuals based on their dental X-ray images. The system integrates a method for teeth segmentation, and three different methods for representing and matching teeth. The first method for matching antemortem (AM) and postmortem (PM) images represents each tooth contour by a set of signature vectors obtained at salient points on the contour of the tooth. The second method uses hierarchical chamfer distance for matching AM and PM teeth to reduce the search space and accordingly reduce the retrieval time. The third matching method represents each tooth by a small set of features extracted using the forcefield energy function and Fourier descriptors. For each matcher, given a query PM image, AM radiographs that are mostly similar to the PM image, are found and presented to the user. To improve the performance of the system, we present different scenarios to fuse the three matchers. We fuse the matchers using three different approaches at the matching level, the decision level, and using the Bayesian framework. Preliminarily results demonstrate that fusing the matching techniques improves the overall performance of the dental identification system.
Omaima Nomir, Mohamed Abdel-Mottaleb
IEEE Trans. Inf. Forensics Secur.2
2007 3D Face Mesh Modeling from Range Images for 3D Face Recognition
abstract
We present an algorithm for 3D face deformation and modeling using range data captured by a 3D scanner. Using only three facial feature points extracted from the range images and a 3D generic face model, the algorithm first aligns the 3D model to the entire range data of a given subject's face. Then each aligned triangle of the mesh model, with three vertices, is treated as a surface plane which is then fitted to the corresponding interior 3D range data, using least squares plane fitting. Via triangular vertices subdivisions, a higher resolution model is generated from the coordinates of the aligned and fitted model. Finally the model and its triangular surfaces are fitted once again resulting in a smoother mesh model that resembles and captures the surface characteristic of the face. Application of the final deformed model in 3D face recognition, using a publicly available database, shows promising results.
A-Nasser Ansari, Mohamed Abdel-Mottaleb, Mohammad H. Mahoor
ICIP (4)2
2007 3D Face Recognition Based on 3D Ridge Lines in Range Data
abstract
In this paper we present an approach for 3D face recognition from range data based on the principal curvature, kmax, and Hausdorff distance. We use the principal curvature, kmax, to represent the face image as a 3D binary image called ridge image. The ridge image shows the locations of the ridge lines around the important facial regions on the face (i.e. the eyes, the nose, and the mouth). We utilize Hausdorff distance to match the ridge image of a given probe to the created ridge images of the subjects in the gallery. For pose alignment, we extract the locations of three feature points, the inner corners of the two eyes and the tip of the nose using Gaussian curvature. These three feature points plus an auxiliary point in the center of the triangle, made by averaging the coordinates of the three feature points, are used for initial 3D face alignment. In the face recognition stage, we find the optimum pose alignment between the probe image and the gallery, which gives the minimum Hasusdorff distance between the two sets of features. This approach is used for identification of both neutral faces and faces with smile expression. Experiments on a public face database of 61 subjects resulted in 93.5% ranked one recognition rate for neutral expression and 82.0% for the faces with smile expression.
Mohammad H. Mahoor, Mohamed Abdel-Mottaleb
ICIP (1)2
2007 Combining Matching Algorithms for Human Identification using Dental X-Ray Radiographs
abstract
The goal of forensic dentistry is to identify individuals based on their dental characteristics. In this paper we present a system for identifying individuals from their dental X-ray records. Given a dental record, usually a postmortem (PM) radiograph, the system searches a database of ante mortem (AM) radiographs and retrieves the best matches from the database. The system automatically segments dental X-ray images into individual teeth and extracts representative feature vectors for each tooth, which are later used for retrieval. The system integrates one method for teeth segmentation, and two different methods for representing and matching teeth. The first matching method represents each tooth contour by signature vectors obtained at salient points on the contour of the tooth. The second method uses hierarchical Chamfer distance for matching AM and PM teeth to reduce the search space and accordingly reduce the retrieval time. Given a query PM image, and according to a matching distance, AM radiographs that are most similar to the PM image, are found and presented to the user using the two matching methods. The experimental results show that the system is robust. We studied the performance of the different modules of the system as well as the results effusing the matching techniques.
Omaima Nomir, Mohamed Abdel-Mottaleb
ICIP (2)2
2007 Human Identification From Dental X-Ray Images Based on the Shape and Appearance of the Teeth
abstract
Dental biometrics deal with human identification from dental characteristics. In this paper, we present a new technique for identifying people based upon shapes and appearances of their teeth from dental X-ray radiographs. The new technique represents each tooth by a feature vector obtained from the forcefield energy function of the grayscale image of the tooth and Fourier descriptors of the contour of the tooth. The feature vector is composed of the distances between a small number of potential energy wells as well as a small number of Fourier descriptors. Given a query image (i.e., postmortem radiograph), each tooth is matched with the archived teeth in the database (antemortem radiographs) that have the same tooth number. Then, voting is used to obtain a list of best matches for the query image based upon the matching results of the individual teeth. Our goal of using appearance and shape-based features together is to overcome the drawback of using only the contour of the tooth, which can be strongly affected by the quality of the images. The experimental results on a database of 162 antemortem images show that our method is effective in identifying individuals based on their dental radiographs
Omaima Nomir, Mohamed Abdel-Mottaleb
IEEE Trans. Inf. Forensics Secur.2
2006 Disparity-Based 3D Face Modeling for 3D Face Recognition
abstract
We present an automatic disparity-based approach for 3D face modeling, from two frontal and one profile view stereo images, for 3D face recognition applications. Once the images are captured, the algorithm starts by extracting selected 2D facial features from one of the frontal views and computes a dense disparity map from the two frontal images. We then align a low resolution 2D mesh model to the selected features, adjust some of its vertices along the profile line using the profile view, increase its triangular vertices to a higher resolution, and re-project them back on the frontal image. Using the coordinates of the re-projected vertices and their corresponding disparities, we capture and compute the 3D facial shape variations using stereo vision. The final result is a deformed 3D model specific to a given subject's face. Application of the model in 3D face recognition validates the algorithm and shows a promising 98 % recognition rate.
A-Nasser Ansari, Mohamed Abdel-Mottaleb, Mohammad H. Mahoor
ICIP2
2006 Hierarchical Dental X-Ray Radiographs Matching
abstract
The goal of forensic dentistry is to identify individuals based on their dental characteristics. In this paper we present a new matching technique for identifying missing, and wanted individuals from their dental X-ray records. Given a dental record, usually a postmortem (PM) radiograph, the proposed technique searches a database of ante mortem (AM) radiographs and retrieves the best matches from the database. The technique is based on matching teeth contours using hierarchical Chamfer distance. The proposed technique has two main stages: feature extraction, and teeth matching. During retrieval, according to a matching distance between the AM and PM teeth, AM radiographs that are most similar to a given PM image, are found and presented to the user. The experimental results on a database of 162 AM images show that the technique is robust for identifying individuals based on their dental records.
Omaima Nomir, Mohamed Abdel-Mottaleb
ICIP2
2006 Disparity-Based 3D Face Modeling using 3D Deformable Facial Mask for 3D Face Recognition
abstract
We present an automatic disparity-based approach for 3D face modeling, from two frontal and one profile view stereo images, for 3D face recognition applications. Once the images are captured, the algorithm starts by extracting selected 2D facial features from one of the frontal views and computes a dense disparity map from the two frontal images. Using the extracted 2D features plus their corresponding disparities in the disparity map, we compute their 3D coordinates. We next align a low resolution 3D mesh model to the 3D features, re-project its vertices on the frontal 2D image and adjust its profile line vertices using the profile view. We increase the resolutions of the resulting 2D model only at its center region to obtain a facial mask model covering distinctive features of the face. The computation of the 2D vertices coordinates with their disparities results in a deformed 3D model mask specific to a give subject face. Application of the model in 3D face recognition validates the algorithm and shows a high recognition rate
A-Nasser Ansari, Mohamed Abdel-Mottaleb, Mohammad H. Mahoor
ICME2
2005 HMM-Based Segmentation and Recognition of Human Activities from Video Sequences
abstract
Recognizing human activities from image sequences is an active area of research in computer vision. Most of the previous work on activity recognition focuses on recognition from video clips that show only single activities. There are few published algorithms for segmenting and recognizing complex activities that are composed of more than one single activity. In this paper, we present a novel HMM-based approach that uses threshold and voting to automatically and effectively segment and recognize complex activities. Experiments on a database of video clips of different activities show that our method is effective
Feng Niu, Mohamed Abdel-Mottaleb
ICME2
2005 Automatic facial feature extraction and 3D face modeling using two orthogonal views with application to 3D face recognition
A-Nasser Ansari, Mohamed Abdel-Mottaleb
Pattern Recognit.2
2005 Tracking multiple people with recovery from partial and total occlusion
Charay Lerdsudwichai, Mohamed Abdel-Mottaleb, A-Nasser Ansari
Pattern Recognit.2
2005 Classification and numbering of teeth in dental bitewing images
Mohammad H. Mahoor, Mohamed Abdel-Mottaleb
Pattern Recognit.2
2005 A system for human identification from X-ray dental radiographs
Omaima Nomir, Mohamed Abdel-Mottaleb
Pattern Recognit.2
2005 A content-based system for human identification based on bitewing dental X-ray images
Jindan Zhou, Mohamed Abdel-Mottaleb
Pattern Recognit.2
2004 Automatic classification of teeth in bitewing dental images
abstract
We present an automated algorithm to classify teeth in bitewing dental images, using Bayesian classification, and assign an absolute number to each tooth based on common numbering system used in dentistry. Fourier descriptors of the contours of the molar and the premolar teeth in bitewing images are used in the Bayesian classification of these two types of the teeth. Then, the spatial relation between the two types of the teeth is considered to number each tooth and correct the misclassification of some teeth in order to obtain high precision results. Experiments with 50 bitewing images containing more than 400 teeth show that our method is capable of classifying and assigning absolute index number to the teeth with high accuracy.
Mohammad H. Mahoor, Mohamed Abdel-Mottaleb
ICIP2
2004 Content-based photo album management using faces' arrangement
abstract
Photo album management, an application of content-based image browsing/retrieval, has attracted much attention in recent years. Identities of individuals appearing in photos are the most important aspect of photo browsing. However, face recognition generally does not work effectively in such situations due to the large variations in pose, illumination and sometimes poor quality of images. We present a system for browsing photo albums. The system automatically detects and stores information about the locations of faces in the photos. A similarity function based on face arrangement is defined. Photos are then clustered based on the similarity function using a proposed clustering algorithm. The system also represents photos of an event using a composite image. The composite image is built from representative faces and an image that represents the event. Experiments indicate that the face arrangement features are effective in representing the semantic content of the photos and are appropriate for photo albums.
Mohamed Abdel-Mottaleb, Longbin Chen
ICME1
2004 Multimodal detection of highlights for multimedia content
Serhan Dagtas, Mohamed Abdel-Mottaleb
Multim. Syst.2
2004 Multimedia descriptions based on MPEG-7: extraction and applications
abstract
The amount of digital multimedia content available to consumers is growing because of the existence of digital capturing devices such as digital cameras, camcorders and the advent of digital video broadcast. With this increase in content, it becomes important for users to be able to browse and search for content in a timely manner. Descriptions and annotations of the content are needed to enable searching and browsing of content. MPEG-7 is a recent ISO standard for multimedia content description. In this paper, we present three descriptors, which we had proposed to MPEG-7, and have been accepted to the standard. In addition, we describe algorithms for automatically extracting these descriptors. We also present the indexing and retrieval algorithms that we developed for these descriptors; these algorithms are fast and scalable for large databases. The results presented in the paper for image and video segment matching using these descriptors show their usefulness in real applications.
Mohamed Abdel-Mottaleb, Santhana Krishnamachari
IEEE Trans. Multim.1
2003 3-D Face Modeling Using Two Views and a Generic Face Model with Application to 3-D Face Recognition
abstract
We present an algorithm for 3D face modeling from a frontal image and a profile image of a person's face. The algorithm starts by computing the 3D coordinates of automatically extracted facial feature points. The coordinates of the selected feature points are then used to deform a 3D generic face model to obtain a 3D face model for that person. Procrustes analysis is used to minimize globally the distances between the facial feature vertices in the model and the corresponding 3D points obtained from the images. Then, local deformation is performed on the facial feature vertices to obtain a more realistic 3D model for the person. Preliminary experiments to asses the applicability of the models for face recognition show encouraging results.
A-Nasser Ansari, Mohamed Abdel-Mottaleb
AVSS2
2003 3D face modeling using two orthogonal views and a generic face model
abstract
We present an algorithm for 3-D face modeling from a frontal and a profile view images of a person's face. The algorithm starts by computing the 3D coordinates of automatically extracted facial feature points. The coordinates of the selected feature points are then used to deform a 3D generic face model to obtain a 3D face model for that person. Procrustes analysis is used to globally minimize the distance between facial feature vertices in the model and the corresponding 3D points obtained from the images. Then, local deformation is performed on the facial feature vertices to obtain a more realistic 3D model for the person. Preliminary experiments to asses the applicability of the models for face recognition show encouraging results.
A-Nasser Ansari, Mohamed Abdel-Mottaleb
ICME2
2003 Algorithm for multiple faces tracking
abstract
Robust real-time face tracking is a challenging task. This paper presents an algorithm for tracking faces of multiple people even in case of total occlusion. The method uses the color distribution of each face with the mean shift tracking method. Mean shift tracking is fast and robust to partial occlusion, and it is rotation invariant and computationally efficient. Since the mean shift algorithm does not deal with the problem of total occlusion, we overcome this problem by using an occlusion grid to detect occlusion. We then use the color distribution of the occluded person's clothes to distinguish that person after the occlusion ends. We also use the speed and the trajectory of the occluded person to predict the locations that should be searched after occlusion ends. The proposed face tracking method integrates multiple features to handle tracking of multiple people even in case of occlusion. The experiments show the robustness of the algorithm.
Charay Lerdsudwichai, Mohamed Abdel-Mottaleb
ICME2
2002 Face Detection in Color Images
abstract
Human face detection plays an important role in applications such as video surveillance, human computer interface, face recognition, and face image database management. We propose a face detection algorithm for color images in the presence of varying lighting conditions as well as complex backgrounds. Based on a novel lighting compensation technique and a nonlinear color transformation, our method detects skin regions over the entire image and then generates face candidates based on the spatial arrangement of these skin patches. The algorithm constructs eye, mouth, and boundary maps for verifying each face candidate. Experimental results demonstrate successful face detection over a wide range of facial variations in color, position, scale, orientation, 3D pose, and expression in images from several photo collections (both indoors and outdoors).
Rein-Lien Hsu, Mohamed Abdel-Mottaleb, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2001 Face detection in color images
abstract
Human face detection is often the first step in applications such as video surveillance, human computer interface, face recognition, and image database management. We propose a face detection algorithm for color images in the presence of varying lighting conditions as well as complex backgrounds. Our method detects skin regions over the entire image, and then generates face candidates based on the spatial arrangement of these skin patches. The algorithm constructs eye, mouth, and boundary maps for verifying each face candidate. Experimental results demonstrate successful detection over a wide variety of facial variations in color, position, scale, rotation, pose, and expression from several photo collections.
Rein-Lien Hsu, Mohamed Abdel-Mottaleb, Anil K. Jain 0001
ICIP (1)2
2001 Extraction of TV highlights using multimedia features
abstract
The amount of digital multimedia content available to consumers is growing due to the digital capturing devices such as digital cameras and camcorders. With this increase in content, it becomes important for people to be able to browse and search for content in a timely manner. To enable efficient content-based video browsing, extracted descriptions and annotations are needed to represent the content. We present algorithms for the automatic extraction of highlights in video using audio, text and visual features. These extracted descriptions can be used for selective browsing of sports videos. We also present the experimental results for the proposed algorithms on several hours of sports programs.
Serhan Dagtas, Mohamed Abdel-Mottaleb
MMSP2
2000 Image Retrieval Based on Edge Representation
abstract
Content-based image retrieval from large collections of images is a challenging task; especially retrieval based on edges. We present a fast algorithm for image retrieval by sketch or image example. The images in the database are archived in a hash table, where the addresses of the entries in the table are derived from the edge orientations in predefined regions of the images. In the retrieval stage, images that have regions similar to the regions of the query image are retrieved from the hash table and a voting is performed to rank the images according to their similarity to the query. Experimental results show that the algorithm is robust.
Mohamed Abdel-Mottaleb
ICIP1
2000 Group-of-Frames/Pictures Color Histogram Descriptors for Multimedia Applications
abstract
Joint representation of color-based features for multiple images or a video segment is an important task for visual information management systems. Generally, a key-frame or key-image is selected from such a group, and the color-related features of the entire collection are represented with those of the chosen sample. Such methods are highly dependent on the quality of the representative sample, and may lead to unreliable results. We present a set of histogram-based descriptors that reliably capture the color content of multiple images or video frames. These descriptors are defined for a group-of-frames (GoF) or a group-of-pictures (GoP). A single representation for the entire collection is obtained by combining individual frame or image histograms in various ways. We demonstrate the efficacy of GoF-histograms for video segment retrieval and the GoP-histograms for fast image search. This descriptor has been accepted to the Working Draft of MPEG-7, the evolving ISO standard for multimedia content description.
A. Müfit Ferman, Santhana Krishnamachari, A. Murat Tekalp, Mohamed Abdel-Mottaleb, Rajiv Mehrotra
ICIP4
2000 SmartWatch: an automated video event finder
Serhan Dagtas, Thomas McGee, Mohamed Abdel-Mottaleb
ACM Multimedia3
1999 Face Detection in Complex Environments from Color Images
abstract
The detection of faces in color images is important for many multimedia applications. It is the first step for face recognition and it can be used for classifying specific shots such as anchorperson and talk show shots. In this paper, we present an algorithm for the detection of faces in color images. The algorithm works by first detecting areas of skin color, then it applies a top-down and a bottom-up analysis to the skin colored areas. The algorithm has been tested on images from news clips and other television programs. The results show that the algorithm is robust and works even for cases where there are objects in the background that have colors similar to the skin.
Mohamed Abdel-Mottaleb, Ahmed M. Elgammal
ICIP (3)1
1999 Image Browsing using Hierarchical Clustering
abstract
Digital images and video clips are becoming popular due to the increase in the availability of consumer devices that capture digital images and video clips. Digital content is also growing over the Internet. The increase of the digital content creates a need for user-friendly tools to browse through large volumes of digital material. We present a clustering-based browsing algorithm. Images are automatically clustered using a hierarchical clustering algorithm and users can then browse through the images by navigating the tree structure that results from the clustering. We have tested the algorithm on a large number of images.
Santhana Krishnamachari, Mohamed Abdel-Mottaleb
ISCC2
1998 A Scalable Algorithm for Image Retrieval by Color
Santhana Krishnamachari, Mohamed Abdel-Mottaleb
ICIP (3)2
1996 Exploiting the JPEG Compression Scheme for Image Retrieval
abstract
We address the problem of retrieving images from a large database using an image as a query. The method is specifically aimed at databases that store images in JPEG format, and works in the compressed domain to create index keys. A key is generated for each image in the database and is matched with the key generated for the query image. The keys are independent of the size of the image. Images that have similar keys are assumed to be similar, but there is no semantic meaning to the similarity.
Michael Shneier, Mohamed Abdel-Mottaleb
IEEE Trans. Pattern Anal. Mach. Intell.2
1993 Binocular motion stereo using MAP estimation
abstract
An algorithm for fusing monocular and stereo cues to get robust estimates of both motion and structure is presented. The authors' algorithm assumes the motion to be along a smooth trajectory and the sequence of images to be dense. The algorithm starts by calculating the instantaneous FOE (focus of expansion). Knowing the FOE, a MAP estimate of the displacement at each pixel and an associated confidence measure are calculated. Using the displacement estimates, a relative depth map is calculated from one of the two frame sequences. By calculating the disparities at some feature points and using information about their relative depths, the instantaneous component of velocity in the direction perpendicular to the image plane (the Z direction) is computed. Using this information, a depth map is calculated; this depth map is then used to derive a prior probability distribution for disparity that is used in matching the two frames of the stereo pairs. This method is used to estimate the disparity of each pixel independently. Experimental results on a real image sequence are given.>
Mohamed Abdel-Mottaleb, Rama Chellappa, Azriel Rosenfeld
CVPR1
1992 Inexact Bayesian estimation
Mohamed Abdel-Mottaleb, Azriel Rosenfeld
Pattern Recognit.1
1992 "Qualitative" Bayesian estimation of digital signals and images
Mohamed Abdel-Mottaleb, Azriel Rosenfeld
Pattern Recognit.1
1989 A multistage algorithm for fast classification of patterns
Hisham El-Shishiny, Mohamed Abdel-Mottaleb, Mohamed El-Rayes, Amin A. Shoukry
Pattern Recognit. Lett.2