VLDB 2026 Research / reviewers in the wild / expert
Wael Abd-Almageed
dblp:99/222 · also Wael AbdAlmageed
· DBLP profile ↗
66ranked-venue papers
11as first author
19since 2021 · last 2025
0000-0002-8320-8530ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 44 · 7 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 42 · 4 first-author · 14 since 2021Systems, architecture and hardware · 7 · 3 first-author · 1 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Attention-Driven Causal Discovery: From Transformer Matrices to Granger Causal Graphs for Non-Stationary Time-series DataabstractCausal discovery in non-stationary time series data is crucial for understanding complex systems but remains challenging due to evolving relationships over time. This paper presents a novel two-stage approach for causal discovery in non-stationary multivariate time series data. The first stage employs a Temporal Attention Forecasting Network (TAFNet), a modified Transformer architecture, to capture complex temporal dependencies and generate informative attention matrices. The second stage utilizes these matrices in an iterative process for Granger causality discovery, refining the predicted causal graph while improving forecasting accuracy. The proposed method addresses the limitations of existing approaches and provides a more complete understanding of causal relationships in non-stationary systems. Extensive experiments demonstrate the method’s superior performance compared to state-of-the-art approaches, particularly in handling non-linear relationships and scaling to high-dimensional data. Jiageng Zhu, Kehao Li, Zheda Mai, Hanchen Xie, Wael Abd-Almageed, Zubin Abraham |
ICASSP | 5 |
| 2024 | Multi-Scope Representation Learning for Causal Relation Discovery with new Challenging Datasets
Jiageng Zhu, Hanchen Xie, Mohamed E. Hussein 0001, Mahyar Khayatkhoei, Jiazhi Li 0001, Wael Abd-Almageed |
BMVC | 7 |
| 2024 | ManiFPT: Defining and Analyzing Fingerprints of Generative ModelsabstractRecent works have shown that generative models leave traces of their underlying generative process on the generated samples, broadly referred to as fingerprints of a generative model, and have studied their utility in detecting synthetic images from real ones. However, the extend to which these fingerprints can distinguish between various types of synthetic image and help identify the underlying generative process remain under-explored. In particular, the very definition of a fingerprint remains unclear, to our knowledge. To that end, in this work, we formalize the definition of artifact and fingerprint in generative models, propose an algorithm for computing them in practice, and finally study its effectiveness in distinguishing a large array of different generative models. We find that using our proposed definition can significantly improve the performance on the task of identifying the underlying generative process from samples (model attribution) compared to existing methods. Additionally, we study the structure of the fingerprints, and observe that it is very predictive of the effect of different design choices on the generative process. Hae Jin Song, Mahyar Khayatkhoei, Wael Abd-Almageed |
CVPR | 3 |
| 2024 | Large Multimodal Models Thrive with Little Data for Image Emotion Prediction
Mohamed E. Hussein 0001, Wael Abd-Almageed |
ICPR (1) | 3 |
| 2024 | TRIGS: Trojan Identification from Gradient-Based Signatures
Mohamed E. Hussein 0001, Sudharshan Subramaniam Janakiraman, Wael Abd-Almageed |
ICPR (3) | 3 |
| 2023 | Emergent Asymmetry of Precision and Recall for Measuring Fidelity and Diversity of Generative Models in High DimensionsabstractPrecision and Recall are two prominent metrics of generative performance, which were proposed to separately measure the fidelity and diversity of generative models. Given their central role in comparing and improving generative models, understanding their limitations are crucially important. To that end, in this work, we identify a critical flaw in the common approximation of these metrics using k-nearest-neighbors, namely, that the very interpretations of fidelity and diversity that are assigned to Precision and Recall can fail in high dimensions, resulting in very misleading conclusions. Specifically, we empirically and theoretically show that as the number of dimensions grows, two model distributions with supports at equal point-wise distance from the support of the real distribution, can have vastly different Precision and Recall regardless of their respective distributions, hence an emergent asymmetry in high dimensions. Based on our theoretical insights, we then provide simple yet effective modifications to these metrics to construct symmetric metrics regardless of the number of dimensions. Finally, we provide experiments on real-world datasets to illustrate that the identified flaw is not merely a pathological case, and that our proposed metrics are effective in alleviating its impact. Mahyar Khayatkhoei, Wael Abd-Almageed |
ICML | 2 |
| 2023 | A Critical View of Vision-Based Long-Term Dynamics Prediction Under Environment MisalignmentabstractDynamics prediction, which is the problem of predicting future states of scene objects based on current and prior states, is drawing increasing attention as an instance of learning physics. To solve this problem, Region Proposal Convolutional Interaction Network (RPCIN), a vision-based model, was proposed and achieved state-of-the-art performance in long-term prediction. RPCIN only takes raw images and simple object descriptions, such as the bounding box and segmentation mask of each object, as input. However, despite its success, the model’s capability can be compromised under conditions of environment misalignment. In this paper, we investigate two challenging conditions for environment misalignment: Cross-Domain and Cross-Context by proposing four datasets that are designed for these challenges: SimB-Border, SimB-Split, BlenB-Border, and BlenB-Split. The datasets cover two domains and two contexts. Using RPCIN as a probe, experiments conducted on the combinations of the proposed datasets reveal potential weaknesses of the vision-based long-term dynamics prediction model. Furthermore, we propose a promising direction to mitigate the Cross-Domain challenge and provide concrete evidence supporting such a direction, which provides dramatic alleviation of the challenge on the proposed datasets. Hanchen Xie, Jiageng Zhu, Mahyar Khayatkhoei, Jiazhi Li 0001, Mohamed E. Hussein 0001, Wael Abd-Almageed |
ICML | 6 |
| 2022 | MONet: Multi-Scale Overlap Network for Duplication Detection in Biomedical ImagesabstractManipulation of biomedical images to misrepresent experimental results has plagued the biomedical community for a while. Recent interest in the problem led to the curation of a dataset and associated tasks to promote the development of biomedical forensic methods. Of these, the largest manipulation detection task focuses on the detection of duplicated regions between images. Traditional computer-vision based forensic models trained on natural images are not designed to overcome the challenges presented by biomedical images. We propose a multi-scale overlap detection model to detect duplicated image regions. Our model is structured to find duplication hierarchically, so as to reduce the number of patch operations. It achieves state-of-the-art performance overall and on multiple biomedical image categories. Ekraam Sabir, Soumyaroop Nandi, Wael Abd-Almageed, Premkumar Natarajan |
ICIP | 3 |
| 2022 | P2M-DeTrack: Processing-in-Pixel-in-Memory for Energy-efficient and Real-Time Multi-Object Detection and TrackingabstractToday’s high resolution, high frame rate cameras in autonomous vehicles generate a large volume of data that needs to be transferred and processed by a downstream processor or machine learning (ML) accelerator to enable intelligent computing tasks, such as multi-object detection and tracking. The massive amount of data transfer incurs significant energy, latency, and bandwidth bottlenecks, which hinders real-time processing. To mitigate this problem, we propose an algorithm-hardware co-design framework called Processing-in-Pixel-in-Memory-based object Detection and Tracking (P2M-DeTrack). P2M-DeTrack is based on a custom faster R-CNN-based model that is distributed partly inside the pixel array (front-end) and partly in a separate FPGA/ASIC (back-end). The proposed front-end in-pixel processing down-samples the input feature maps significantly with judiciously optimized strided convolution and pooling. Compared to a conventional baseline design that transfers frames of RGB pixels to the back-end, the resulting P2M-DeTrack designs reduce the data bandwidth between sensor and back-end by up to 24×. The designs also reduce the sensor and total energy (obtained from in-house circuit simulations at Globalfoundries 22nm technology node) per frame by 5.7× and 1.14×, respectively. Lastly, they reduce the sensing and total frame latency by an estimated 1.7× and 3×, respectively. We evaluate our approach on the multi-object object detection (tracking) task of the large-scale BDD100K dataset and observe only a 0.5% reduction in the mean average precision (0.8% reduction in the identification F1 score) compared to the state-of-the-art. Gourav Datta, Souvik Kundu 0002, Zihan Yin, Joe Mathai, Zeyu Liu 0003, Mulin Tian, Shunlin Lu, Ravi Teja Lakkireddy, Andrew G. Schmidt, Wael Abd-Almageed, Ajey P. Jacob, Akhilesh Jaiswal 0001, Peter A. Beerel |
VLSI-SoC | 11 |
| 2021 | Information-Theoretic Bias Assessment Of Learned Representations Of Pretrained Face RecognitionabstractAs equality issues in the use of face recognition have garnered a lot of attention lately, greater efforts have been made to debiased deep learning models to improve fairness to minorities. However, there is still no clear definition nor sufficient analysis for bias assessment metrics. We propose an information-theoretic, independent bias assessment metric to identify degree of bias against protected demographic attributes from learned representations of pretrained facial recognition systems. Our metric differs from other methods that rely on classification accuracy or examine the differences between ground truth and predicted labels of protected attributes predicted using a shallow network. Also, we argue, theoretically and experimentally, that logits-level loss is not adequate to explain bias since predictors based on neural networks will always find correlations. Further, we present a synthetic dataset that mitigates the issue of insufficient samples in certain cohorts. Lastly, we establish a benchmark metric by presenting advantages in clear discrimination and small variation comparing with other metrics, and evaluate the performance of different debiased models with the proposed metric. Jiazhi Li 0001, Wael Abd-Almageed |
FG | 2 |
| 2021 | Explaining Face Presentation Attack Detection Using Natural LanguageabstractA large number of deep neural network based techniques have been developed to address the challenging problem of face presentation attack detection (PAD). Whereas such techniques' focus has been on improving PAD performance in terms of classification accuracy and robustness against unseen attacks and environmental conditions, there exists little attention on the explainability of PAD predictions. In this paper, we tackle the problem of explaining PAD predictions through natural language. Our approach passes feature representations of a deep layer of the PAD model to a language model to generate text describing the reasoning behind the PAD prediction. Due to the limited amount of annotated data in our study, we apply a light-weight LSTM network as our natural language generation model. We investigate how the quality of the generated explanations is affected by different loss functions, including the commonly used word-wise cross entropy loss, a sentence discriminative loss, and a sentence semantic loss. We perform our experiments using face images from a dataset consisting of 1,105 bona-fide and 924 presentation attack samples. Our quantitative and qualitative results show the effectiveness of our model for generating proper PAD explanations through text as well as the power of the sentence-wise losses. To the best of our knowledge, this is the first introduction of a joint biometrics-NLP task. Our dataset can be obtained through our GitHub page11https://github.com/ISICV/PADISI_USC_Dataset . Hengameh Mirzaalian, Mohamed E. Hussein 0001, Leonidas Spinoulas, Jonathan May, Wael Abd-Almageed |
FG | 5 |
| 2021 | Adversarial Defense for Deep Speaker Recognition Using Hybrid Adversarial TrainingabstractDeep neural network based speaker recognition systems can easily be deceived by an adversary using minuscule imperceptible perturbations to the input speech samples. These adversarial attacks pose serious security threats to the speaker recognition systems that use speech biometric. To address this concern, in this work, we propose a new defense mechanism based on a hybrid adversarial training (HAT) setup. In contrast to existing works on countermeasures against adversarial attacks in deep speaker recognition that only use class-boundary information by supervised cross-entropy (CE) loss, we propose to exploit additional information from supervised and unsupervised cues to craft diverse and stronger perturbations for adversarial training. Specifically, we employ multi-task objectives using CE, feature-scattering (FS), and margin losses to create adversarial perturbations and include them for adversarial training to enhance the robustness of the model. We conduct speaker recognition experiments on the Librispeech dataset, and compare the performance with state-of-the-art projected gradient descent (PGD)-based adversarial training which employs only CE objective. The proposed HAT improves adversarial accuracy by absolute 3.29% and 3.18% for PGD and Carlini-Wagner (CW) attacks respectively, while retaining high accuracy on benign examples. Monisankha Pal, Arindam Jati, Raghuveer Peri, Chin-Cheng Hsu, Wael Abd-Almageed, Shri Narayanan |
ICASSP | 5 |
| 2021 | SIGN: Spatial-information Incorporated Generative Network for Generalized Zero-shot Semantic SegmentationabstractUnlike conventional zero-shot classification, zero-shot semantic segmentation predicts a class label at the pixel level instead of the image level. When solving zero-shot semantic segmentation problems, the need for pixel-level prediction with surrounding context motivates us to incorporate spatial information using positional encoding. We improve standard positional encoding by introducing the concept of Relative Positional Encoding, which integrates spatial information at the feature level and can handle arbitrary image sizes. Furthermore, while self-training is widely used in zero-shot semantic segmentation to generate pseudo-labels, we propose a new knowledge-distillation-inspired self-training strategy, namely Annealed Self-Training, which can automatically assign different importance to pseudo-labels to improve performance. We systematically study the proposed Relative Positional Encoding and Annealed Self-Training in a comprehensive experimental evaluation, and our empirical results confirm the effectiveness of our method on three benchmark datasets. Jiaxin Cheng, Soumyaroop Nandi, Premkumar Natarajan, Wael Abd-Almageed |
ICCV | 4 |
| 2021 | Partner-Assisted Learning for Few-Shot Image ClassificationabstractFew-shot Learning has been studied to mimic human visual capabilities and learn effective models without the need of exhaustive human annotation. Even though the idea of meta-learning for adaptation has dominated the few-shot learning methods, how to train a feature extractor is still a challenge. In this paper, we focus on the design of training strategy to obtain an elemental representation such that the prototype of each novel class can be estimated from a few labeled samples. We propose a two-stage training scheme, Partner-Assisted Learning (PAL), which first trains a Partner Encoder to model pair-wise similarities and extract features serving as soft-anchors, and then trains a Main Encoder by aligning its outputs with soft-anchors while attempting to maximize classification performance. Two alignment constraints from logit-level and feature-level are designed individually. For each few-shot task, we perform prototype classification. Our method consistently outperforms the state-of-the-art methods on four benchmarks. Detailed ablation studies of PAL are provided to justify the selection of each component involved in training. Jiawei Ma, Hanchen Xie, Guangxing Han, Shih-Fu Chang, Aram Galstyan, Wael Abd-Almageed |
ICCV | 6 |
| 2021 | Detection and Continual Learning of Novel Face Presentation AttacksabstractAdvances in deep learning, combined with availability of large datasets, have led to impressive improvements in face presentation attack detection research. However, state-of-the-art face antispoofing systems are still vulnerable to novel types of attacks that are never seen during training. Moreover, even if such attacks are correctly detected, these systems lack the ability to adapt to newly encountered attacks. The post-training ability of continually detecting new types of attacks and self-adaptation to identify these attack types, after the initial detection phase, is highly appealing. In this paper, we enable a deep neural network to detect anomalies in the observed input data points as potential new types of attacks by suppressing the confidence-level of the network outside the training samples’ distribution. We then use experience replay to update the model to incorporate knowledge about new types of attacks without forgetting the past learned attack types. Experimental results are provided to demonstrate the effectiveness of the proposed method on two benchmark datasets as well as a newly introduced dataset which exhibits a large variety of attack types.1 Leonidas Spinoulas, Mohamed E. Hussein 0001, Joe Mathai, Wael Abd-Almageed |
ICCV | 5 |
| 2021 | BioFors: A Large Biomedical Image Forensics DatasetabstractResearch in media forensics has gained traction to combat the spread of misinformation. However, most of this research has been directed towards content generated on social media. Biomedical image forensics is a related problem, where manipulation or misuse of images reported in biomedical research documents is of serious concern. The problem has failed to gain momentum beyond an academic discussion due to an absence of benchmark datasets and standardized tasks. In this paper we present BioFors1– the first dataset for benchmarking common biomedical image manipulations. BioFors comprises 47,805 images extracted from 1,031 open-source research papers. Images in BioFors are divided into four categories – Microscopy, Blot/Gel, FACS and Macroscopy. We also propose three tasks for forensic analysis – external duplication detection, internal duplication detection and cut/sharp-transition detection. We benchmark BioFors on all tasks with suitable state-of-the-art algorithms. Our results and analysis show that existing algorithms developed on common computer vision datasets are not robust when applied to biomedical images, validating that more research is required to address the unique challenges of biomedical image forensics. Ekraam Sabir, Soumyaroop Nandi, Wael Abd-Almageed, Premkumar Natarajan |
ICCV | 3 |
| 2021 | Customizable Camera Verification for Media Forensic
Huaigu Cao, Wael Abd-Almageed |
ICDAR (3) | 2 |
| 2021 | MUSCLE: Strengthening Semi-Supervised Learning Via Concurrent Unsupervised Learning Using Mutual Information MaximizationabstractDeep neural networks are powerful, massively parameterized machine learning models that have been shown to perform well in supervised learning tasks. However, very large amounts of labeled data are usually needed to train deep neural networks. Several semi-supervised learning approaches have been proposed to train neural networks using smaller amounts of labeled data with a large amount of unlabeled data. The performance of these semisupervised methods significantly degrades as the size of labeled data decreases. We introduce Mutual-information-based Unsupervised & Semi-supervised Concurrent LEarning (MUSCLE), a hybrid learning approach that uses mutual information to combine both unsupervised and semisupervised learning. MUSCLE can be used as a standalone training scheme for neural networks, and can also be incorporated into other learning approaches. We show that the proposed hybrid model outperforms state of the art on several standard benchmarks, including CIFAR-10, CIFAR-100, and Mini-Imagenet. Furthermore, the performance gain consistently increases with the reduction in the amount of labeled data, as well as in the presence of bias. We also show that MUSCLE has the potential to boost the classification performance when used in the fine-tuning phase for a model pre-trained only on unlabeled data. Hanchen Xie, Mohamed E. Hussein 0001, Aram Galstyan, Wael Abd-Almageed |
WACV | 4 |
| 2021 | Adversarial attack and defense strategies for deep speaker recognition systems
Arindam Jati, Chin-Cheng Hsu, Monisankha Pal, Raghuveer Peri, Wael Abd-Almageed, Shri Narayanan |
Comput. Speech Lang. | 5 |
| 2020 | Invariant Representations through Adversarial ForgettingabstractWe propose a novel approach to achieving invariance for deep neural networks in the form of inducing amnesia to unwanted factors of data through a new adversarial forgetting mechanism. We show that the forgetting mechanism serves as an information-bottleneck, which is manipulated by the adversarial training to learn invariance to unwanted factors. Empirical results show that the proposed framework achieves state-of-the-art performance at learning invariance in both nuisance and bias settings on a diverse collection of datasets and tasks. Ayush Jaiswal, Daniel Moyer, Greg Ver Steeg, Wael Abd-Almageed, Premkumar Natarajan |
AAAI | 4 |
| 2020 | Towards Learning Structure via Consensus for Face Segmentation and ParsingabstractFace segmentation is the task of densely labeling pixels on the face according to their semantics. While current methods place an emphasis on developing sophisticated architectures, use conditional random fields for smoothness, or rather employ adversarial training, we follow an alternative path towards robust face segmentation and parsing. Occlusions, along with other parts of the face, have a proper structure that needs to be propagated in the model during training. Unlike state-of-the-art methods that treat face segmentation as an independent pixel prediction problem, we argue instead that it should hold highly correlated outputs within the same object pixels. We thereby offer a novel learning mechanism to enforce structure in the prediction via consensus, guided by a robust loss function that forces pixel objects to be consistent with each other. Our face parser is trained by transferring knowledge from another model, yet it encourages spatial consistency while fitting the labels. Different than current practice, our method enjoys pixel-wise predictions, yet paves the way for fewer artifacts, less sparse masks, and spatially coherent outputs. Iacopo Masi, Joe Mathai, Wael Abd-Almageed |
CVPR | 3 |
| 2020 | Two-Branch Recurrent Network for Isolating Deepfakes in Videos
Iacopo Masi, Aditya Killekar, Royston Marian Mascarenhas, Shenoy Pratik Gurudatt, Wael Abd-Almageed |
ECCV (7) | 5 |
| 2020 | MEG: Multi-Evidence GNN for Multimodal Semantic ForensicsabstractFake news often involves semantic manipulations across modalities such as image, text, location etc and requires the development of multimodal semantic forensics for its detection. Recent research has centered the problem around images, calling it image repurposing - where a digitally unmanipulated image is semantically misrepresented by means of its accompanying multimodal metadata such as captions, location, etc. The image and metadata together comprise a multimedia package. The problem setup requires algorithms to perform multimodal semantic forensics to authenticate a query multimedia package using a reference dataset of potentially related packages as evidences. Existing methods are limited to using a single evidence (retrieved package), which ignores potential performance improvement from the use of multiple evidences. In this work, we introduce a novel graph neural network based model for multimodal semantic forensics, which effectively utilizes multiple retrieved packages as evidences and is scalable with the number of evidences. We compare the scalability and performance of our model against existing methods. Experimental results show that the proposed model outperforms existing state-of-the-art algorithms with an error reduction of up to 25 %. Ekraam Sabir, Ayush Jaiswal, Wael Abd-Almageed, Premkumar Natarajan |
ICPR | 3 |
| 2020 | Revealing True Identity: Detecting Makeup Attacks in Face-based Biometric SystemsabstractFace-based authentication systems are among the most commonly used biometric systems, because of the ease of capturing face images at a distance and in non-intrusive way. These systems are, however, susceptible to various presentation attacks, including printed faces, artificial masks, and makeup attacks. In this paper, we propose a novel solution to address makeup attacks, which are the hardest to detect in such systems because makeup can substantially alter the facial features of a person, including making them appear older/younger by adding/hiding wrinkles, modifying the shape of eyebrows, beard, and moustache, and changing the color of lips and cheeks. In our solution, we design a generative adversarial network for removing the makeup from face images while retaining their essential facial features and then compare the face images before and after removing makeup. We collect a large dataset of various types of makeup, especially malicious makeup that can be used to break into remote unattended security systems. This dataset is quite different from existing makeup datasets that mostly focus on cosmetic aspects. We conduct an extensive experimental study to evaluate our method and compare it against the state-of-the art using standard objective metrics commonly used in biometric systems as well as subjective metrics collected through a user study. Our results show that the proposed solution produces high accuracy and substantially outperforms the closest works in the literature. Mohammad Amin Arab, Puria Azadi Moghadam, Mohamed E. Hussein 0001, Wael Abd-Almageed, Mohamed Hefeeda |
ACM Multimedia | 4 |
| 2019 | ManTra-Net: Manipulation Tracing Network for Detection and Localization of Image Forgeries With Anomalous FeaturesabstractTo fight against real-life image forgery, which commonly involves different types and combined manipulations, we propose a unified deep neural architecture called ManTraNet. Unlike many existing solutions, ManTra-Net is an end-to-end network that performs both detection and localization without extra preprocessing and postprocessing. ManTra-Net is a fully convolutional network and handles images of arbitrary sizes and many known forgery types such splicing, copy-move, removal, enhancement, and even unknown types. This paper has three salient contributions. We design a simple yet effective self-supervised learning task to learn robust image manipulation traces from classifying 385 image manipulation types. Further, we formulate the forgery localization problem as a local anomaly detection problem, design a Z-score feature to capture local anomaly, and propose a novel long short-term memory solution to assess local anomalies. Finally, we carefully conduct ablation experiments to systematically optimize the proposed network design. Our extensive experimental results demonstrate the generalizability, robustness and superiority of ManTra-Net, not only in single types of manipulations/forgeries, but also in their complicated combinations. Yue Wu 0001, Wael Abd-Almageed, Premkumar Natarajan |
CVPR | 2 |
| 2019 | QATM: Quality-Aware Template Matching for Deep LearningabstractFinding a template in a search image is one of the core problems in many computer vision applications, such as template matching, image semantic alignment, image-to-GPS verification \etc. In this paper, we propose a novel quality-aware template matching method, which is not only used as a standalone template matching algorithm, but also a trainable layer that can be easily plugged in any deep neural network. Specifically, we assess the quality of a matching pair as its soft-ranking among all matching pairs, and thus different matching scenarios like 1-to-1, 1-to-many, and many-to-many will be all reflected to different values. Our extensive studies in the classic template matching problem and deep learning tasks demonstrate the effectiveness of QATM: it not only outperforms SOTA template matching methods when used alone, but also largely improves existing DNN solutions when used in DNN. Jiaxin Cheng, Yue Wu 0001, Wael Abd-Almageed, Premkumar Natarajan |
CVPR | 3 |
| 2019 | AIRD: Adversarial Learning Framework for Image Repurposing DetectionabstractImage repurposing is a commonly used method for spreading misinformation on social media and online forums, which involves publishing untampered images with modified metadata to create rumors and further propaganda. While manual verification is possible, given vast amounts of verified knowledge available on the internet, the increasing prevalence and ease of this form of semantic manipulation call for the development of robust automatic ways of assessing the semantic integrity of multimedia data. In this paper, we present a novel method for image repurposing detection that is based on the real-world adversarial interplay between a bad actor who repurposes images with counterfeit metadata and a watchdog who verifies the semantic consistency between images and their accompanying metadata, where both players have access to a reference dataset of verified content, which they can use to achieve their goals. The proposed method exhibits state-of-the-art performance on location-identity, subject-identity and painting-artist verification, showing its efficacy across a diverse set of scenarios. Ayush Jaiswal, Yue Wu 0001, Wael Abd-Almageed, Iacopo Masi, Premkumar Natarajan |
CVPR | 3 |
| 2019 | Learning Pose-Aware Models for Pose-Invariant Face Recognition in the WildabstractWe propose a method designed to push the frontiers of unconstrained face recognition in the wild with an emphasis on extreme out-of-plane pose variations. Existing methods either expect a single model to learn pose invariance by training on massive amounts of data or else normalize images by aligning faces to a single frontal pose. Contrary to these, our method is designed to explicitly tackle pose variations. Our proposed Pose-Aware Models (PAM) process a face image using several pose-specific, deep convolutional neural networks (CNN). 3D rendering is used to synthesize multiple face poses from input images to both train these models and to provide additional robustness to pose variations at test time. Our paper presents an extensive analysis of the IARPA Janus Benchmark A (IJB-A), evaluating the effects that landmark detection accuracy, CNN layer selection, and pose model selection all have on the performance of the recognition pipeline. It further provides comparative evaluations on IJB-A and the PIPA dataset. These tests show that our approach outperforms existing methods, even surprisingly matching the accuracy of methods that were specifically fine-tuned to the target dataset. Parts of this work previously appeared in [1] and [2]. Iacopo Masi, Feng-Ju Chang, Jongmoo Choi, Shai Harel, Jungyeon Kim, KangGeon Kim, Jatuporn Toy Leksut, Stephen Rawls, Yue Wu 0001, Tal Hassner, Wael Abd-Almageed, Gérard G. Medioni, Louis-Philippe Morency, Premkumar Natarajan, Ramakant Nevatia |
IEEE Trans. Pattern Anal. Mach. Intell. | 11 |
| 2018 | Image-to-GPS Verification Through a Bottom-Up Pattern Matching Network
Jiaxin Cheng, Yue Wu 0001, Wael Abd-Almageed, Premkumar Natarajan |
ACCV (5) | 3 |
| 2018 | Bidirectional Conditional Generative Adversarial Networks
Ayush Jaiswal, Wael Abd-Almageed, Yue Wu 0001, Premkumar Natarajan |
ACCV (3) | 2 |
| 2018 | Weighted Feature Pooling Network in Template-Based Recognition
Zekun Li 0007, Yue Wu 0001, Wael Abd-Almageed, Premkumar Natarajan |
ACCV (5) | 3 |
| 2018 | BusterNet: Detecting Copy-Move Image Forgery with Source/Target Localization
Yue Wu 0001, Wael Abd-Almageed, Premkumar Natarajan |
ECCV (6) | 2 |
| 2018 | Deep Multimodal Image-Repurposing DetectionabstractNefarious actors on social media and other platforms often spread rumors and falsehoods through images whose metadata (e.g., captions) have been modified to provide visual substantiation of the rumor/falsehood. This type of modification is referred to as image repurposing, in which often an unmanipulated image is published along with incorrect or manipulated metadata to serve the actor's ulterior motives. We present the Multimodal Entity Image Repurposing (MEIR) dataset, a substantially challenging dataset over that which has been previously available to support research into image repurposing detection. The new dataset includes location, person, and organization manipulations on real-world data sourced from Flickr. We also present a novel, end-to-end, deep multimodal learning model for assessing the integrity of an image by combining information extracted from the image with related information from a knowledge base. The proposed method is compared against state-of-the-art techniques on existing datasets as well as MEIR, where it outperforms existing methods across the board, with AUC improvement up to 0.23. Ekraam Sabir, Wael Abd-Almageed, Yue Wu 0001, Premkumar Natarajan |
ACM Multimedia | 2 |
| 2018 | Unsupervised Adversarial InvarianceabstractData representations that contain all the information about target variables but are invariant to nuisance factors benefit supervised learning algorithms by preventing them from learning associations between these factors and the targets, thus reducing overfitting. We present a novel unsupervised invariance induction framework for neural networks that learns a split representation of data through competitive training between the prediction task and a reconstruction task coupled with disentanglement, without needing any labeled information about nuisance factors or domain knowledge. We describe an adversarial instantiation of this framework and provide analysis of its working. Our unsupervised model outperforms state-of-the-art methods, which are supervised, at inducing invariance to inherent nuisance factors, effectively using synthetic data augmentation to learn invariance, and domain adaptation. Our method can be applied to any prediction task, eg., binary/multi-class classification or regression, without loss of generality. Ayush Jaiswal, Yue Wu 0001, Wael Abd-Almageed, Premkumar Natarajan |
NeurIPS | 3 |
| 2018 | Image Copy-Move Forgery Detection via an End-to-End Deep Neural NetworkabstractIn this paper, for the first time, we introduce a new end-to-end deep neural network predicting forgery masks to the image copy-move forgery detection problem. Specifically, we use a convolutional neural network to extract block-like features from an image, compute self-correlations between different blocks, use a pointwise feature extractor to locate matching points, and reconstruct a forgery mask through a deconvolutional network. Unlike classic solutions requiring multiple stages of training and parameter tuning, ranging from feature extraction to postprocessing, the proposed solution is fully trainable and can be jointly optimized for the forgery mask reconstruction loss. Our experimental results demonstrate that the proposed method achieves better forgery detection performance than classic approaches relying on different features and matching schemes, and it is more robust against various known attacks like affine transformation, JPEG compression, blurring, etc. Yue Wu 0001, Wael Abd-Almageed, Premkumar Natarajan |
WACV | 2 |
| 2017 | EPAT: Euclidean Perturbation Analysis and Transform - An Agnostic Data Adaptation Framework for Improving Facial Landmark DetectorsabstractWe propose EPAT, (Euclidean Perturbation Analysis and Transform) a novel unsupervised adaptation approach for improving the accuracy of any facial landmark detector by characterizing the stability of landmark prediction on test images. In EPAT, a test image is transformed several times using a set of Euclidean transforms, producing several perturbed images. The black box landmark detector is used to find facial landmarks on each perturbed version of the test image. Subsequently, inverse transforms are applied to the corresponding landmarks in order to map them back to the original image. Mean and variance are calculated for all inversely transformed detection. Mean and variance represent the new ensemble prediction and the sensitivity of the underlying landmark detector, respectively. We also introduce affine variance (AV) of facial landmarks. AV is used as a measure of the stability of the predicted landmarks and a criterion for selecting a good data adaptation model which effectively addresses potential mismatches between test and training data of the underlying landmark detector. EPAT is evaluated using four state-of-the-art landmark detectors on the standard 300W dataset and also incorporated into a face recognition pipeline to show improved recognition accuracy on the challenging IJB-A dataset. Yue Wu 0001, Wael Abd-Almageed, Stephen Rawls, Premkumar Natarajan |
FG | 2 |
| 2017 | Adversarial Auto-Encoders for Speech Based Emotion RecognitionabstractRecently, generative adversarial networks and adversarial autoencoders have gained a lot of attention in machine learning community due to their exceptional performance in tasks such as digit classification and face recognition. They map the autoencoder's bottleneck layer output (termed as code vectors) to different noise Probability Distribution Functions (PDFs), that can be further regularized to cluster based on class information. In addition, they also allow a generation of synthetic samples by sampling the code vectors from the mapped PDFs. Inspired by these properties, we investigate the application of adversarial autoencoders to the domain of emotion recognition. Specifically, we conduct experiments on the following two aspects: (i) their ability to encode high dimensional feature vector representations for emotional utterances into a compressed space (with a minimal loss of emotion class discriminability in the compressed space), and (ii) their ability to regenerate synthetic samples in the original feature space, to be later used for purposes such as training emotion recognition classifiers. We demonstrate the promise of adversarial autoencoders with regards to these aspects on the Interactive Emotional Dyadic Motion Capture (IEMOCAP) corpus and present our analysis. Saurabh Sahu, Rahul Gupta 0001, Ganesh Sivaraman, Wael Abd-Almageed, Carol Y. Espy-Wilson |
INTERSPEECH | 4 |
| 2017 | Multimedia Semantic Integrity Assessment Using Joint Embedding Of Images And TextabstractReal-world multimedia data is often composed of multiple modalities such as an image or a video with associated text (e.g., captions, user comments, etc.) and metadata. Such multimodal data packages are prone to manipulations, where a subset of these modalities can be altered to misrepresent or repurpose data packages, with possible malicious intent. It is therefore important to develop methods to assess or verify the integrity of these multimedia packages. Using computer vision and natural language processing methods to directly compare the image (or video) and the associated caption to verify the integrity of a media package is only possible for a limited set of objects and scenes. In this paper we present a novel deep-learning-based approach that uses a reference set of multimedia packages to assess the semantic integrity of multimedia packages containing images and captions. We construct a joint embedding of images and captions with deep multimodal representation learning on the reference dataset in a framework that also provides image-caption consistency scores (ICCSs). The integrity of query media packages is assessed as the inlierness of the query ICCSs with respect to the reference dataset. We present the MultimodAl Information Manipulation dataset (MAIM), a new dataset of media packages from Flickr, which we are making available to the research community. We use both the newly created dataset as well as Flickr30K and MS COCO datasets to quantitatively evaluate our proposed approach. The reference dataset does not contain unmanipulated versions of tampered query packages. Our method is able to achieve F-1 scores of 0.75, 0.89 and 0.94 on MAIM, Flickr30K and MS COCO, respectively, for detecting semantically incoherent media packages. Ayush Jaiswal, Ekraam Sabir, Wael Abd-Almageed, Premkumar Natarajan |
ACM Multimedia | 3 |
| 2017 | Deep Matching and Validation Network: An End-to-End Solution to Constrained Image Splicing Localization and DetectionabstractImage splicing is a very common image manipulation technique that is sometimes used for malicious purposes. A splicing detection and localization algorithm usually takes an input image and produces a binary decision indicating whether the input image has been manipulated, and also a segmentation mask that corresponds to the spliced region. Most existing splicing detection and localization pipelines suffer from two main shortcomings: 1) they use handcrafted features that are not robust against subsequent processing (e.g., compression), and 2) each stage of the pipeline is usually optimized independently. In this paper we extend the formulation of the underlying splicing problem to consider two input images, a query image and a potential donor image. Here the task is to estimate the probability that the donor image has been used to splice the query image, and obtain the splicing masks for both the query and donor images. We introduce a novel deep convolutional neural network architecture, called Deep Matching and Validation Network (DMVN), which simultaneously localizes and detects image splicing. The proposed approach does not depend on handcrafted features and uses raw input images to create deep learned representations. Furthermore, the DMVN is end-to-end optimized to produce the probability estimates and the segmentation masks. Our extensive experiments demonstrate that this approach outperforms state-of-the-art splicing detection methods by a large margin in terms of both AUC score and speed. Yue Wu 0001, Wael Abd-Almageed, Premkumar Natarajan |
ACM Multimedia | 2 |
| 2016 | Learning document image binarization from dataabstractWe present a fully trainable solution for binarization of degraded document images using extremely randomized trees. Unlike previous attempts that often use simple features, our method encodes all heuristics about whether or not a pixel is foreground text into a high-dimensional feature vector and learns a more complicated decision function. We introduce two novel features, the Logarithm Intensity Percentile (LIP) and the Relative Darkness Index (RDI), and combine them with low level features, and reformulated features from existing binarization methods. Experimental results show that using small sample size (about 1.5% of all available training data), we can achieve a binarization performance comparable to manually-tuned, state-of-the-art methods. Additionally, the trained document binarization classifier shows good generalization capabilities on out-of-domain data. Yue Wu 0001, Premkumar Natarajan, Stephen Rawls, Wael Abd-Almageed |
ICIP | 4 |
| 2016 | Computationally efficient template-based face recognitionabstractClassically, face recognition depends on computing the similarity (or distance) between a pair of face images and/or their respective representations, where each subject is represented by one image. Template-based face recognition was introduced by the release of IARPA's Janus Benchmark-A (IJB-A) dataset, in which each enrolled subject is represented by a group of one or more images, called a template. The group of images comprising a template might have been acquired using different head poses, illuminations, ages and facial expressions. Template images could come from still images or video frames. Therefore, measuring the similarity between templates representing two subjects significantly increases the number of pairwise image comparisons (i.e., O(NM), where N and M are the number of image templates being compared). As the number of enrolled subjects, K, increases, both computational and space requirements become computationally prohibitive. To address this challenge, we present a novel approximate nearest-neighbor (ANN) search-based solution. Given a query template, ANN methods are used to find similar face images. Retrieved images are used to construct a template pool that is used to find the correct identity of the query subject. The proposed approach largely reduces the number of imposter template-pair comparisons. Experimental results on the IJB-A dataset show that the proposed approach achieves significant speed-up and storage savings, without sacrificing accuracy. Yue Wu 0001, Wael Abd-Almageed, Stephen Rawls, Premkumar Natarajan |
ICPR | 2 |
| 2016 | Face recognition using deep multi-pose representationsabstractWe introduce our method and system for face recognition using multiple pose-aware deep learning models. In our representation, a face image is processed by several pose-specific deep convolutional neural network (CNN) models to generate multiple pose-specific features. 3D rendering is used to generate multiple face poses from the input image. Sensitivity of the recognition system to pose variations is reduced since we use an ensemble of pose-specific CNN features. The paper presents extensive experimental results on the effect of landmark detection, CNN layer selection and pose model selection on the performance of the recognition pipeline. Our novel representation achieves better results than the state-of-the-art on IARPA's CS2 and NIST's IJB-A in both verification and identification (i.e. search) tasks. Wael Abd-Almageed, Yue Wu 0001, Stephen Rawls, Shai Harel, Tal Hassner, Iacopo Masi, Jongmoo Choi, Jatuporn Toy Leksut, Jungyeon Kim, Premkumar Natarajan, Ramakant Nevatia, Gérard G. Medioni |
WACV | 1 |
| 2015 | A graphical model approach for matching partial signaturesabstractIn this paper, we present a novel partial signature matching method using graphical models. Shape context features are extracted from the contour of signatures to capture local variations, and K-means clustering is used to build a visual vocabulary from a set of reference signatures. To describe the signatures, supervised latent Dirichlet allocation is used to learn the latent distributions of the salient regions over the visual vocabulary and hierarchical Dirichlet processes are implemented to infer the number of salient regions needed. Our work is evaluated on three datasets derived from the DS-I Tobacco signature dataset with clean signatures and the DS-II UMD dataset with signatures with different degradations. The results show the effectiveness of the approach for both the partial and full signature matching. Xianzhi Du, David S. Doermann, Wael Abd-Almageed |
CVPR | 3 |
| 2015 | Feature Selection using Partial Least Squares regression and optimal experiment designabstractWe propose a supervised feature selection technique called the Optimal Loadings, that is based on applying the theory of Optimal Experiment Design (OED) to Partial Least Squares (PLS) regression. We apply the OED criterions to PLS with the goal of selecting an optimal feature subset that minimizes the variance of the regression model and hence minimize its prediction error. We show that the variance of the PLS model can be minimized by employing the OED criterions on the loadings covariance matrix obtained from PLS. We also provide an intuitive viewpoint to the technique by deriving the Aoptimality version of the Optimal Loadings criterion using the properties of maximum relevance and minimum redundancy for PLS models. In our experiments we use the D-optimality version of the criterion which maximizes the determinant of the loadings covariance matrix. To overcome the computational challenges in this criterion, we provide an approximate D-optimality criterion along with the theoretical justification. Varun K. Nagaraja, Wael Abd-Almageed |
IJCNN | 2 |
| 2014 | Signature Matching Using Supervised Topic ModelsabstractIn this paper, we present a novel signature matching method based on supervised topic models. Shape Context features are extracted from signature shape contours which capture the local variations in signature properties. We then use the concept of topic models to learn the shape context features which correspond to individual authors. The approach consists of three primary steps. First, K-means is used to cluster shape context features to form term frequency histograms which correspond to a vocabulary for the set of signatures in the gallery. Second, a supervised topic model is used to construct an observation/author correspondence. Finally, the correspondence is used to classify query signatures and return the corresponding author. Two datasets are used to test our algorithm: DS-I Tobacco signature dataset with clean signatures and DS-II UMD dataset with noisy signatures. We demonstrate considerable improvement over state of the art methods. Xianzhi Du, David S. Doermann, Wael Abd-Almageed |
ICPR | 3 |
| 2013 | Large-Scale Signature Matching Using Multi-stage HashingabstractIn this paper, we propose a fast large-scale signature matching method based on locality sensitive hashing (LSH). Shape Context features are used to describe the structure of signatures. Two stages of hashing are performed to find the nearest neighbours for query signatures. In the first stage, we use M randomly generated hyper planes to separate shape context feature points into different bins, and compute a term-frequency histogram to represent the feature point distribution as a feature vector. In the second stage we again use LSH to categorize the high-level features into different classes. The experiments are carried out on two datasets - DS-I, a small dataset contains 189 signatures, and DS-II, a large dataset created by our group which contains 26,000 signatures. We show that our algorithm can achieve a high accuracy even when few signatures are collected from one same person and perform fast matching when dealing with a large dataset. Xianzhi Du, Wael Abd-Almageed, David S. Doermann |
ICDAR | 2 |
| 2013 | Distributed Kernel Matrix Approximation and Implementation Using Message Passing InterfaceabstractWe propose a distributed method to compute similarity (also known as kernel and Gram) matrices used in various kernel-based machine learning algorithms. Current methods for computing similarity matrices have quadratic time and space complexities, which make them not scalable to large-scale data sets. To reduce these quadratic complexities, the proposed method first partitions the data into smaller subsets using various families of locality sensitive hashing, including random project and spectral hashing. Then, the method computes the similarity values among points in the smaller subsets to result in approximated similarity matrices. We analytically show that the time and space complexities of the proposed method are sub quadratic. We implemented the proposed method using the Message Passing Interface (MPI) framework and ran it on a cluster. Our results with real large-scale data sets show that the proposed method does not significantly impact the accuracy of the computed similarity matrices and it achieves substantial savings in running time and memory requirements. Taher A. Dameh, Wael Abd-Almageed, Mohamed Hefeeda |
ICMLA (1) | 2 |
| 2012 | Distributed approximate spectral clustering for large-scale datasetsabstractData-intensive applications are becoming important in many science and engineering fields, because of the high rates in which data are being generated and the numerous opportunities offered by the sheer amount of these data. Large-scale datasets, however, are challenging to process using many of the current machine learning algorithms due to their high time and space complexities. In this paper, we propose a novel approximation algorithm that enables kernel-based machine learning algorithms to efficiently process very large-scale datasets. While important in many applications, current kernel-based algorithms suffer from a scalability problem as they require computing a kernel matrix which takes O(N2) in time and space to compute and store. The proposed algorithm yields substantial reduction in computation and memory overhead required to compute the kernel matrix, and it does not significantly impact the accuracy of the results. In addition, the level of approximation can be controlled to tradeoff some accuracy of the results with the required computing resources. The algorithm is designed such that it is independent of the subsequently used kernel-based machine learning algorithm, and thus can be used with many of them. To illustrate the effect of the approximation algorithm, we developed a variant of the spectral clustering algorithm on top of it. Furthermore, we present the design of a MapReduce-based implementation of the proposed algorithm. We have implemented this design and run it on our own Hadoop cluster as well as on the Amazon Elastic MapReduce service. Experimental results on synthetic and real datasets demonstrate that significant time and memory savings can be achieved using our algorithm. Mohamed Hefeeda, Wael Abd-Almageed |
HPDC | 3 |
| 2011 | Segmentation of Handwritten Textlines in Presence of Touching ComponentsabstractThis paper presents an approach to text line extraction in handwritten document images which combines local and global techniques. We propose a graph-based technique to detect touching and proximity errors that are common with handwritten text lines. In a refinement step, we use Expectation-Maximization (EM) to iteratively split the error segments to obtain correct text-lines. We show improvement in accuracies using our correction method on datasets of Arabic document images. Results on a set of artificially generated proximity images show that the method is effective for handling touching errors in handwritten document images. Jayant Kumar, David S. Doermann, Wael Abd-Almageed |
ICDAR | 4 |
| 2011 | Silhouette-based gesture and action recognition via modeling trajectories on Riemannian shape manifolds
Mohamed F. Abdelkader, Wael Abd-Almageed, Anuj Srivastava, Rama Chellappa |
Comput. Vis. Image Underst. | 2 |
| 2010 | Handwritten Arabic text line segmentation using affinity propagationabstractIn this paper, we present a novel graph-based method for extracting handwritten text lines in monochromatic Arabic document images. Our approach consists of two steps - Coarse text line estimation using primary components which define the line and assignment of diacritic components which are more difficult to associate with a given line. We first estimate local orientation at each primary component to build a sparse similarity graph. We then, use a shortest path algorithm to compute similarities between non-neighboring components. From this graph, we obtain coarse text lines using two estimates obtained from Affinity propagation and Breadth-first search. In the second step, we assign secondary components to each text line. The proposed method is very fast and robust to non-uniform skew and character size variations, normally present in handwritten text lines. We evaluate our method using a pixel-matching criteria, and report 96% accuracy on a dataset of 125 Arabic document images. We also present a proximity analysis on datasets generated by artificially decreasing the spacings between text lines to demonstrate the robustness of our approach. Jayant Kumar, Wael Abd-Almageed, David S. Doermann |
Document Analysis Systems | 2 |
| 2009 | Page Rule-Line Removal Using Linear Subspaces in Monochromatic Handwritten Arabic DocumentsabstractIn this paper we present a novel method for removing page rule lines in monochromatic handwritten Arabic documents using subspace methods with minimal effect on the quality of the foreground text. We use moment and histogram properties to extract features that represent the characteristics of the underlying rule lines. A linear subspace is incrementally built to obtain a line model that can be used to identify rule line pixels. We also introduce a novel scheme for evaluating noise removal algorithms in general and we use it to assess the quality of our rule line removal algorithm. Experimental results presented on a data set of 50 Arabic documents, handwritten by different writers, demonstrate the effectiveness of the proposed method. Wael Abd-Almageed, Jayant Kumar, David S. Doermann |
ICDAR | 1 |
| 2009 | Approximate kernel matrix computation on GPUs forlarge scale learning applicationsabstractKernel-based learning methods require quadratic space and time complexities to compute the kernel matrix. These complexities limit the applicability of kernel methods to large scale problems with millions of data points. In this paper, we introduce a novel representation of kernel matrices on Graphics Processing Units (GPU). The novel representation exploits the sparseness of the kernel matrix to address the space complexity problem. It also respects the guidelines for memory access on GPUs, which are critical for good performance, to address the time complexity problem. Our representation utilizes the locality preserving properties of space filling curves to obtain a band approximation of the kernel matrix. To prove the validity of the representation, we use Affinity Propagation, an unsupervised clustering algorithm, as an example of kernel methods. Experimental results show a 40x speedup of AP using our representation without degradation in clustering performance. Mohamed E. Hussein 0001, Wael Abd-Almageed |
ICS | 2 |
| 2009 | Efficient band approximation of Gram matrices for large scale kernel methods on GPUsabstractKernel-based methods require O(N2) time and space complexities to compute and store non-sparse Gram matrices, which is prohibitively expensive for large scale problems. We introduce a novel method to approximate a Gram matrix with a band matrix. Our method relies on the locality preserving properties of space filling curves, and the special structure of Gram matrices. Our approach has several important merits. First, it computes only those elements of the Gram matrix that lie within the projected band. Second, it is simple to parallelize. Third, using the special band matrix structure makes it space efficient and GPU-friendly. We developed GPU implementations for the Affinity Propagation (AP) clustering algorithm using both our method and the COO sparse representation. Our band approximation is about 5 times more space efficient and faster to construct than COO. AP gains up to 6x speedup using our method without any degradation in its clustering performance. Mohamed E. Hussein 0001, Wael Abd-Almageed |
SC | 2 |
| 2008 | Online, simultaneous shot boundary detection and key frame extraction for sports videos using rank tracingabstractIn this paper, we present a novel algorithm for simultaneously detecting shot boundaries and extracting key frames from video sequences or streams in real-time. Multivariate feature vectors are extracted from the video frames and arranged in a feature matrix. Singular value decomposition is then used to factorize the feature matrix and compute the significant singular vectors. The rank of the singular vectors is traced using a sliding window approach. By tracing the computed rank, we are able to determine shot boundaries and extract key frames of the video. Results of experiments conducted on soccer videos show that the algorithm is robust to a wide range of digital effects used during shot transition. Moreover, the algorithm is shown to run in real time making it suitable for online video summarization and multimedia networking applications. Wael Abd-Almageed |
ICIP | 1 |
| 2008 | Document-zone classification using partial least squares and hybrid classifiersabstractThis paper introduces a novel document-zone classification algorithm. Low level image features are first extracted from document zones and partial least squares is used on pairs of classes to compute discriminating pairwise features. Rather than using the popular one-against-all and one-against-one voting schemes, we introduce a novel hybrid method which combines the benefits of the two schemes. The algorithm is applied on the University of Washington dataset and 97.3% classification accuracy is obtained. Wael Abd-Almageed, Mudit Agrawal, Wontaek Seo, David S. Doermann |
ICPR | 1 |
| 2008 | Human detection using iterative feature selection and logistic principal component analysisabstractWe present a fast feature selection algorithm suitable for object detection applications where the image being tested must be scanned repeatedly to detected the object of interest at different locations and scales. The algorithm iteratively estimates the belongness probability of image pixels to foreground of the image. To prove the validity of the algorithm, we apply it to a human detection problem. The edge map is filtered using a feature selection algorithm. The filtered edge map is then projected onto an eigen space of human shapes to determine if the image contains a human. Since the edge maps are binary in nature, Logistic Principal Component Analysis is used to obtain the eigen human shape space. Experimental results illustrate the accuracy of the human detector. Wael Abd-Almageed, Larry Davis 0001 |
ICRA | 1 |
| 2006 | Density Estimation Using Mixtures of Mixtures of Gaussians
Wael Abd-Almageed, Larry Davis 0001 |
ECCV (4) | 1 |
| 2006 | Real-Time Human Detection, Tracking, and Verification in Uncontrolled Camera Motion EnvironmentsabstractIn environments where a camera is installed on a freely moving platform, e.g. a vehicle or a robot, object detection and tracking becomes much more difficult. In this paper, we presents a real time system for human detection, tracking, and verification in such challenging environments. To deliver a robust performance, the system integrates several computer vision algorithms to perform its function: a human detection algorithm, an object tracking algorithm, and a motion analysis algorithm. To utilize the available computing resources to the maximum possible extent, each of the system components is designed to work in a separate thread that communicates with the other threads through shared data structures. The focus of this paper is more on the implementation issues than on the algorithmic issues of the system. Object oriented design was adopted to abstract algorithmic details away from the system structure. Mohamed E. Hussein 0001, Wael Abd-Almageed, Yang Ran, Larry Davis 0001 |
ICVS | 2 |
| 2006 | Tracking Articulating Objects from Ground Vehicles using Mixtures of MixturesabstractAn algorithm for tracking articulating objects from moving camera platforms is presented. Mixtures of mixtures are used to model the appearance of the object and the background. The state of the object is tracked using a particle filter. Egomotion information are estimated and used to set the state variance of the particle filter. Results of tracking human objects from an unmanned ground vehicle are used to evaluate the tracking algorithm Wael Abd-Almageed, Mohamed E. Hussein 0001, Larry Davis 0001 |
IROS | 1 |
| 2006 | Estimating time-varying densities using a stochastic learning automaton
Wael Abd-Almageed, Aly I. El-Osery, Christopher E. Smith |
Soft Comput. | 1 |
| 2005 | Pedestrian classification from moving platforms using cyclic motion patternabstractThis paper describes an efficient pedestrian detection system for videos acquired from moving platforms. Given a detected and tracked object as a sequence of images within a bounding box, we describe the periodic signature of its motion pattern using a twin-pendulum model. Then a principle gait angle is extracted in every frame providing gait phase information. By estimating the periodicity from the phase data using a digital phase locked loop (dPLL), we quantify the cyclic pattern of the object, which helps us to continuously classify it as a pedestrian. Past approaches have used shape detectors applied to a single image or classifiers based on human body pixel oscillations, but ours is the first to integrate a global cyclic motion model and periodicity analysis. Novel contributions of this paper include: i) development of a compact shape representation of cyclic motion as a signature for a pedestrian, ii) estimation of gait period via a feedback loop module, and iii) implementation of a fast online pedestrian classification system which operates on videos acquired from moving platforms. Yang Ran, Qinfen Zheng, Isaac Weiss, Larry Davis 0001, Wael Abd-Almageed |
ICIP (2) | 5 |
| 2005 | A learning automata based power management for ad-hoc networksabstractPower management is a very important aspect of ad-hoc networks. It directly impacts the network throughput among other network metrics. On the other hand, transmission power management may result in disconnected networks and increased level of collisions. In this paper, we introduce a transmission power control based on stochastic learning automata (SLA) to modify the transmission power. Based on the level of successful transmissions and the level of packet retransmissions, the SLA will modify the transmission power level either by increasing it or decreasing it. The probabilistic nature of SLA makes it a useful choice for ad-hoc networks. Using the network simulator NS, we show that using SLA for transmission power will result in an increased system bandwidth and a decrease in the collision levels. Aly I. El-Osery, David Baird, Wael Abd-Almageed |
SMC | 3 |
| 2003 | Non-parametric expectation maximization: a learning automata approachabstractThe famous expectation maximization technique suffers two major drawbacks. First, the number of components has to be specified apriori. Also, the expectation maximization is sensitive to initialization. In this paper, we present a new stochastic technique for estimating the mixture parameters. Parzen Window is used to estimate a discrete estimate of the PDF of the given data. Stochastic learning automata is then used to select the mixture parameters that minimize the distance between the discrete estimate of the PDF and the estimate of the expectation maximization. The validity of the proposed approach is verified using bivariate simulation data. Wael Abd-Almageed, Aly I. El-Osery, Christopher E. Smith |
SMC | 1 |
| 2003 | Kernel snakes: non-parametric active contour modelsabstractIn this paper, a new non-parametric generalized formulation to statistical pressure snakes is presented. We discuss the shortcomings of the traditional pressure snakes. We then introduce a new generic pressure model that alleviates these shortcomings, based on the Bayesian decision theory. Non-parametric techniques are used to obtain the statistical models that drive the snake. We discuss the advantages of using the proposed non-parametric model compared to other parametric techniques. Multi-colored-target tracking is used to demonstrate the performance of the proposed approach. Experimental results show enhanced, real-time performance. Wael Abd-Almageed, Christopher E. Smith, Samah Ramadan |
SMC | 1 |
| 2002 | Contour migration: solving object ambiguity with shape-space visual guidanceabstractA fundamental problem in computer vision is the issue of shape ambiguity. Simply stated, a silhouette cannot uniquely identify an object or an object's classification since many unique objects can present identical occluding contours. This problem has no solution in the general case for a monocular vision system. This paper presents a method for disambiguating objects during silhouette matching using a visual servoing system. This method identifies the camera motion(s) that gives disambiguating views of the objects. These motions are identified through a new technique called contour migration. The occluding contour's shape is used to identify objects or object classes that are potential matches for that shape. A contour migration is then determined that disambiguates the possible matches by purposive viewpoint adjustment. The technique is demonstrated using an example set of objects. Wael Abd-Almageed, Christopher E. Smith |
IROS | 1 |