VLDB 2026 Research / reviewers in the wild / expert
Mohamed E. Hussein 0001
dblp:70/7061 · also Mohamed Elsayed Ahmed Hussein, Mohamed Elsayed Hussein
· DBLP profile ↗
25ranked-venue papers
8as first author
7since 2021 · last 2024
0000-0002-4707-9313ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 16 · 4 first-author · 6 since 2021Systems, architecture and hardware · 3 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Multi-Scope Representation Learning for Causal Relation Discovery with new Challenging Datasets
Jiageng Zhu, Hanchen Xie, Mohamed E. Hussein 0001, Mahyar Khayatkhoei, Jiazhi Li 0001, Wael Abd-Almageed |
BMVC | 4 |
| 2024 | Large Multimodal Models Thrive with Little Data for Image Emotion Prediction
Mohamed E. Hussein 0001, Wael Abd-Almageed |
ICPR (1) | 2 |
| 2024 | TRIGS: Trojan Identification from Gradient-Based Signatures
Mohamed E. Hussein 0001, Sudharshan Subramaniam Janakiraman, Wael Abd-Almageed |
ICPR (3) | 1 |
| 2023 | A Critical View of Vision-Based Long-Term Dynamics Prediction Under Environment MisalignmentabstractDynamics prediction, which is the problem of predicting future states of scene objects based on current and prior states, is drawing increasing attention as an instance of learning physics. To solve this problem, Region Proposal Convolutional Interaction Network (RPCIN), a vision-based model, was proposed and achieved state-of-the-art performance in long-term prediction. RPCIN only takes raw images and simple object descriptions, such as the bounding box and segmentation mask of each object, as input. However, despite its success, the model’s capability can be compromised under conditions of environment misalignment. In this paper, we investigate two challenging conditions for environment misalignment: Cross-Domain and Cross-Context by proposing four datasets that are designed for these challenges: SimB-Border, SimB-Split, BlenB-Border, and BlenB-Split. The datasets cover two domains and two contexts. Using RPCIN as a probe, experiments conducted on the combinations of the proposed datasets reveal potential weaknesses of the vision-based long-term dynamics prediction model. Furthermore, we propose a promising direction to mitigate the Cross-Domain challenge and provide concrete evidence supporting such a direction, which provides dramatic alleviation of the challenge on the proposed datasets. Hanchen Xie, Jiageng Zhu, Mahyar Khayatkhoei, Jiazhi Li 0001, Mohamed E. Hussein 0001, Wael Abd-Almageed |
ICML | 5 |
| 2021 | Explaining Face Presentation Attack Detection Using Natural LanguageabstractA large number of deep neural network based techniques have been developed to address the challenging problem of face presentation attack detection (PAD). Whereas such techniques' focus has been on improving PAD performance in terms of classification accuracy and robustness against unseen attacks and environmental conditions, there exists little attention on the explainability of PAD predictions. In this paper, we tackle the problem of explaining PAD predictions through natural language. Our approach passes feature representations of a deep layer of the PAD model to a language model to generate text describing the reasoning behind the PAD prediction. Due to the limited amount of annotated data in our study, we apply a light-weight LSTM network as our natural language generation model. We investigate how the quality of the generated explanations is affected by different loss functions, including the commonly used word-wise cross entropy loss, a sentence discriminative loss, and a sentence semantic loss. We perform our experiments using face images from a dataset consisting of 1,105 bona-fide and 924 presentation attack samples. Our quantitative and qualitative results show the effectiveness of our model for generating proper PAD explanations through text as well as the power of the sentence-wise losses. To the best of our knowledge, this is the first introduction of a joint biometrics-NLP task. Our dataset can be obtained through our GitHub page11https://github.com/ISICV/PADISI_USC_Dataset . Hengameh Mirzaalian, Mohamed E. Hussein 0001, Leonidas Spinoulas, Jonathan May, Wael Abd-Almageed |
FG | 2 |
| 2021 | Detection and Continual Learning of Novel Face Presentation AttacksabstractAdvances in deep learning, combined with availability of large datasets, have led to impressive improvements in face presentation attack detection research. However, state-of-the-art face antispoofing systems are still vulnerable to novel types of attacks that are never seen during training. Moreover, even if such attacks are correctly detected, these systems lack the ability to adapt to newly encountered attacks. The post-training ability of continually detecting new types of attacks and self-adaptation to identify these attack types, after the initial detection phase, is highly appealing. In this paper, we enable a deep neural network to detect anomalies in the observed input data points as potential new types of attacks by suppressing the confidence-level of the network outside the training samples’ distribution. We then use experience replay to update the model to incorporate knowledge about new types of attacks without forgetting the past learned attack types. Experimental results are provided to demonstrate the effectiveness of the proposed method on two benchmark datasets as well as a newly introduced dataset which exhibits a large variety of attack types.1 Leonidas Spinoulas, Mohamed E. Hussein 0001, Joe Mathai, Wael Abd-Almageed |
ICCV | 3 |
| 2021 | MUSCLE: Strengthening Semi-Supervised Learning Via Concurrent Unsupervised Learning Using Mutual Information MaximizationabstractDeep neural networks are powerful, massively parameterized machine learning models that have been shown to perform well in supervised learning tasks. However, very large amounts of labeled data are usually needed to train deep neural networks. Several semi-supervised learning approaches have been proposed to train neural networks using smaller amounts of labeled data with a large amount of unlabeled data. The performance of these semisupervised methods significantly degrades as the size of labeled data decreases. We introduce Mutual-information-based Unsupervised & Semi-supervised Concurrent LEarning (MUSCLE), a hybrid learning approach that uses mutual information to combine both unsupervised and semisupervised learning. MUSCLE can be used as a standalone training scheme for neural networks, and can also be incorporated into other learning approaches. We show that the proposed hybrid model outperforms state of the art on several standard benchmarks, including CIFAR-10, CIFAR-100, and Mini-Imagenet. Furthermore, the performance gain consistently increases with the reduction in the amount of labeled data, as well as in the presence of bias. We also show that MUSCLE has the potential to boost the classification performance when used in the fine-tuning phase for a model pre-trained only on unlabeled data. Hanchen Xie, Mohamed E. Hussein 0001, Aram Galstyan, Wael Abd-Almageed |
WACV | 2 |
| 2020 | Revealing True Identity: Detecting Makeup Attacks in Face-based Biometric SystemsabstractFace-based authentication systems are among the most commonly used biometric systems, because of the ease of capturing face images at a distance and in non-intrusive way. These systems are, however, susceptible to various presentation attacks, including printed faces, artificial masks, and makeup attacks. In this paper, we propose a novel solution to address makeup attacks, which are the hardest to detect in such systems because makeup can substantially alter the facial features of a person, including making them appear older/younger by adding/hiding wrinkles, modifying the shape of eyebrows, beard, and moustache, and changing the color of lips and cheeks. In our solution, we design a generative adversarial network for removing the makeup from face images while retaining their essential facial features and then compare the face images before and after removing makeup. We collect a large dataset of various types of makeup, especially malicious makeup that can be used to break into remote unattended security systems. This dataset is quite different from existing makeup datasets that mostly focus on cosmetic aspects. We conduct an extensive experimental study to evaluate our method and compare it against the state-of-the art using standard objective metrics commonly used in biometric systems as well as subjective metrics collected through a user study. Our results show that the proposed solution produces high accuracy and substantially outperforms the closest works in the literature. Mohammad Amin Arab, Puria Azadi Moghadam, Mohamed E. Hussein 0001, Wael Abd-Almageed, Mohamed Hefeeda |
ACM Multimedia | 3 |
| 2020 | Identifying motion pathways in highly crowded scenes: A non-parametric tracklet clustering approach
Allam S. Hassanein, Mohamed E. Hussein 0001, Walid Gomaa 0001, Yasushi Makihara, Yasushi Yagi |
Comput. Vis. Image Underst. | 2 |
| 2020 | Efficient and Robust Skeleton-Based Quality Assessment and Abnormality Detection in Human Action PerformanceabstractElderly people can be provided with safer and more independent living by the early detection of abnormalities in their performing actions and the frequent assessment of the quality of their motion. Low-cost depth sensing is one of the emerging technologies that can be used for unobtrusive and inexpensive motion abnormality detection and quality assessment. In this study, we develop and evaluate vision-based methods to detect and assess neuromusculoskeletal disorders manifested in common daily activities using three-dimensional skeletal data provided by the SDK of a depth camera (e.g., MS Kinect and Asus Xtion PRO). The proposed methods are based on extracting medically -justified features to compose a simple descriptor. Thereafter, a probabilistic normalcy model is trained on normal motion patterns. For abnormality detection, a test sequence is classified as either normal or abnormal based on its likelihood, which is calculated from the trained normalcy model. For motion quality assessment, a linear regression model is built using the proposed descriptor in order to quantitatively assess the motion quality. The proposed methods were evaluated on four common daily actions-sit to stand, stand to sit, flat walk, and gait on stairs-from two datasets, a publicly released dataset and our dataset that was collected in a clinic from 32 patients suffering from different neuromusculoskeletal disorders and 11 healthy individuals. Experimental results demonstrate promising results, which is a step toward having convenient in-home automatic health care services. Amr Elkholy, Mohamed E. Hussein 0001, Walid Gomaa 0001, Dima Damen, Emmanuel Saba |
IEEE J. Biomed. Health Informatics | 2 |
| 2018 | Multi-modality-based Arabic sign language recognitionabstractWith the increase in the number of deaf‐mute people in the Arab world and the lack of Arabic sign language (ArSL) recognition benchmark data sets, there is a pressing need for publishing a large‐volume and realistic ArSL data set. This study presents such a data set, which consists of 150 isolated ArSL signs. The data set is challenging due to the great similarity among hand shapes and motions in the collected signs. Along with the data set, a sign language recognition algorithm is presented. The authors’ proposed method consists of three major stages: hand segmentation, hand shape sequence and body motion description, and sign classification. The hand shape segmentation is based on the depth and position of the hand joints. Histograms of oriented gradients and principal component analysis are applied on the segmented hand shapes to obtain the hand shape sequence descriptor. The covariance of the three‐dimensional joints of the upper half of the skeleton in addition to the hand states and face properties are adopted for motion sequence description. The canonical correlation analysis and random forest classifiers are used for classification. The achieved accuracy is 55.57% over 150 ArSL signs, which is considered promising. Marwa S. Elpeltagy, Moataz M. Abdelwahab, Mohamed E. Hussein 0001, Amin A. Shoukry, Asmaa Shoala, Moustafa Galal |
IET Comput. Vis. | 3 |
| 2017 | Human Action Recognition Using A Multi-Modal Hybrid Deep Learning Model
Hany A. El-Ghaish, Mohamed E. Hussein 0001, Amin A. Shoukry |
BMVC | 2 |
| 2016 | Semantic Analysis for Crowded Scenes Based on Non-Parametric Tracklet Clustering
Allam S. Hassanein, Mohamed E. Hussein 0001, Walid Gomaa 0001 |
IJCAI | 2 |
| 2016 | Linear-time online action detection from 3D skeletal data using bags of gestureletsabstractSliding window is one direct way to extend a successful recognition system to handle the more challenging detection problem. While action recognition decides only whether or not an action is present in a pre-segmented video sequence, action detection identifies the time interval where the action occurred in an unsegmented video stream. Sliding window approaches can however be slow as they maximize a classifier score over all possible sub-intervals. Even though new schemes utilize dynamic programming to speed up the search for the optimal sub-interval, they require offline processing on the whole video sequence. In this paper, we propose a novel approach for online action detection based on 3D skeleton sequences extracted from depth data. It identifies the sub-interval with the maximum classifier score in linear time. Furthermore, it is suitable for real-time applications with low latency. Moustafa Meshry, Mohamed E. Hussein 0001, Marwan Torki |
WACV | 2 |
| 2015 | Visual Comparison of Images Using Multiple Kernel Learning for RankingabstractRanking is the central problem for many applications such as web search, recommendation systems, and visual comparison of images. In this paper, the multiple kernel learning framework is generalized for the learning to rank problem. This approach extends the existing learning to rank algorithms by considering multiple kernel learning and consequently improves their effectiveness. The proposed approach provides the convenience of fusing different features for describing the underlying data. As an application to our approach, the problem of visual image comparison is studied. Several visual features are used for describing the images and multiple kernel learning is adopted to find an optimal feature fusion. Experimental results on three challenging datasets show that our approach outperforms the state-of-the art and is significantly more efficient in runtime. Amr Sharaf, Mohamed E. Hussein 0001, Mohamed A. Ismail |
BMVC | 2 |
| 2015 | Real-Time Multi-scale Action Detection from 3D Skeleton DataabstractIn this paper we introduce a real-time system for action detection. The system uses a small set of robust features extracted from 3D skeleton data. Features are effectively described based on the probability distribution of skeleton data. The descriptor computes a pyramid of sample covariance matrices and mean vectors to encode the relationship between the features. For handling the intra-class variations of actions, such as action temporal scale variations, the descriptor is computed using different window scales for each action. Discriminative elements of the descriptor are mined using feature selection. The system achieves accurate detection results on difficult unsegmented sequences. Experiments on MSRC-12 and G3D datasets show that the proposed system outperforms the state-of-the-art in detection accuracy with very low latency. To the best of our knowledge, we are the first to propose using multi-scale description in action detection from 3D skeleton data. Amr Sharaf, Marwan Torki, Mohamed E. Hussein 0001, Motaz El-Saban |
WACV | 3 |
| 2013 | Histogram of Oriented Displacements (HOD): Describing Trajectories of Human Joints for Action Recognition
Mohammad Abdelaziz Gowayyed, Marwan Torki, Mohamed E. Hussein 0001, Motaz El-Saban |
IJCAI | 3 |
| 2013 | Human Action Recognition Using a Temporal Hierarchy of Covariance Descriptors on 3D Joint Locations
Mohamed E. Hussein 0001, Marwan Torki, Mohammad Abdelaziz Gowayyed, Motaz El-Saban |
IJCAI | 1 |
| 2011 | CrossTrack: Robust 3D tracking from two cross-sectional viewsabstractOne of the challenges in radiotherapy of moving tumors is to determine the location of the tumor accurately. Existing solutions to the problem are either invasive or inaccurate. We introduce a non-invasive solution to the problem by tracking the tumor in 3D using bi-plane ultrasound image sequences. We present CrossTrack, a novel tracking algorithm in this framework. We pose the problem as recursive inference of 3D location and tumor boundary segmentation in the two ultrasound views using the tumor 3D model as a prior. For the segmentation task, a robust graph-based approach is deployed as follows: First, robust segmentation priors are obtained through the tumor 3D model. Second, a unified graph combining information across time and multiple views is constructed with a robust weighting function. For the tracking task, an effective mechanism for recovery from respiration-induced occlusion is introduced. Our experiments show the robustness of CrossTrack in handling challenging tumor shapes and disappearance scenarios, with sub-voxel accuracy, and almost 100% precision and recall, significantly outperforming baseline solutions. Mohamed E. Hussein 0001, Fatih Porikli, Rui Li 0053, Suayb S. Arslan |
CVPR | 1 |
| 2009 | Approximate kernel matrix computation on GPUs forlarge scale learning applicationsabstractKernel-based learning methods require quadratic space and time complexities to compute the kernel matrix. These complexities limit the applicability of kernel methods to large scale problems with millions of data points. In this paper, we introduce a novel representation of kernel matrices on Graphics Processing Units (GPU). The novel representation exploits the sparseness of the kernel matrix to address the space complexity problem. It also respects the guidelines for memory access on GPUs, which are critical for good performance, to address the time complexity problem. Our representation utilizes the locality preserving properties of space filling curves to obtain a band approximation of the kernel matrix. To prove the validity of the representation, we use Affinity Propagation, an unsupervised clustering algorithm, as an example of kernel methods. Experimental results show a 40x speedup of AP using our representation without degradation in clustering performance. Mohamed E. Hussein 0001, Wael Abd-Almageed |
ICS | 1 |
| 2009 | Efficient band approximation of Gram matrices for large scale kernel methods on GPUsabstractKernel-based methods require O(N2) time and space complexities to compute and store non-sparse Gram matrices, which is prohibitively expensive for large scale problems. We introduce a novel method to approximate a Gram matrix with a band matrix. Our method relies on the locality preserving properties of space filling curves, and the special structure of Gram matrices. Our approach has several important merits. First, it computes only those elements of the Gram matrix that lie within the projected band. Second, it is simple to parallelize. Third, using the special band matrix structure makes it space efficient and GPU-friendly. We developed GPU implementations for the Affinity Propagation (AP) clustering algorithm using both our method and the COO sparse representation. Our band approximation is about 5 times more space efficient and faster to construct than COO. AP gains up to 6x speedup using our method without any degradation in its clustering performance. Mohamed E. Hussein 0001, Wael Abd-Almageed |
SC | 1 |
| 2009 | A Comprehensive Evaluation Framework and a Comparative Study for Human DetectorsabstractWe introduce a framework for evaluating human detectors that considers the practical application of a detector on a full image using multisize sliding-window scanning. We produce detection error tradeoff (DET) curves relating the miss detection rate and the false-alarm rate computed by deploying the detector on cropped windows and whole images, using, in the latter, either image resize or feature resize. Plots for cascade classifiers are generated based on confidence scores instead of on variation of the number of layers. To assess a method's overall performance on a given test, we use the average log miss rate (ALMR) as an aggregate performance score. To analyze the significance of the obtained results, we conduct 10-fold cross-validation experiments. We applied our evaluation framework to two state-of-the-art cascade-based detectors on the standard INRIA person dataset and a local dataset of near-infrared images. We used our evaluation framework to study the differences between the two detectors on the two datasets with different evaluation methods. Our results show the utility of our framework. They also suggest that the descriptors used to represent features and the training window size are more important in predicting the detection performance than the nature of the imaging process, and that the choice between resizing images or features can have serious consequences. Mohamed E. Hussein 0001, Fatih Porikli, Larry Davis 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2008 | Kernel integral images: A framework for fast non-uniform filteringabstractIntegral images are commonly used in computer vision and computer graphics applications. Evaluation of box filters via integral images can be performed in constant time, regardless of the filter size. Although Heckbert (1986) extended the integral image approach for more complex filters, its usage has been very limited, in practice. In this paper, we present an extension to integral images that allows for application of a wide class of non-uniform filters. Our approach is superior to Heckbertpsilas in terms of precision requirements and suitability for parallelization. We explain the theoretical basis of the approach and instantiate two concrete examples: filtering with bilinear interpolation, and filtering with approximated Gaussian weighting. Our experiments show the significant speedups we achieve, and the higher accuracy of our approach compared to Heckbertpsilas. Mohamed E. Hussein 0001, Fatih Porikli, Larry Davis 0001 |
CVPR | 1 |
| 2006 | Real-Time Human Detection, Tracking, and Verification in Uncontrolled Camera Motion EnvironmentsabstractIn environments where a camera is installed on a freely moving platform, e.g. a vehicle or a robot, object detection and tracking becomes much more difficult. In this paper, we presents a real time system for human detection, tracking, and verification in such challenging environments. To deliver a robust performance, the system integrates several computer vision algorithms to perform its function: a human detection algorithm, an object tracking algorithm, and a motion analysis algorithm. To utilize the available computing resources to the maximum possible extent, each of the system components is designed to work in a separate thread that communicates with the other threads through shared data structures. The focus of this paper is more on the implementation issues than on the algorithmic issues of the system. Object oriented design was adopted to abstract algorithmic details away from the system structure. Mohamed E. Hussein 0001, Wael Abd-Almageed, Yang Ran, Larry Davis 0001 |
ICVS | 1 |
| 2006 | Tracking Articulating Objects from Ground Vehicles using Mixtures of MixturesabstractAn algorithm for tracking articulating objects from moving camera platforms is presented. Mixtures of mixtures are used to model the appearance of the object and the background. The state of the object is tracked using a particle filter. Egomotion information are estimated and used to set the state variance of the particle filter. Results of tracking human objects from an unmanned ground vehicle are used to evaluate the tracking algorithm Wael Abd-Almageed, Mohamed E. Hussein 0001, Larry Davis 0001 |
IROS | 2 |