VLDB 2026 Research / reviewers in the wild / expert
Ahmed M. Elgammal
dblp:e/AhmedMElgammal · also Ahmed Elgammal
· DBLP profile ↗
111ranked-venue papers
20as first author
7since 2021 · last 2024
0000-0003-2761-4822ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 91 · 17 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 80 · 11 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-authorDatabases, data management, data science and information retrieval · 2 · 2 first-authorSystems, architecture and hardware · 1Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | MoMA: Multimodal LLM Adapter for Fast Personalized Image Generation
Kunpeng Song, Yizhe Zhu, Ahmed M. Elgammal |
ECCV (40) | 5 |
| 2024 | StyleGAN-Fusion: Diffusion Guided Domain Adaptation of Image GeneratorsabstractCan a text-to-image diffusion model be used as a training objective for adapting a GAN generator to another domain? In this paper, we show that the classifier-free guidance can be leveraged as a critic and enable generators to distill knowledge from large-scale text-to-image diffusion models. Generators can be efficiently shifted into new domains indicated by text prompts without access to groundtruth samples from target domains. We demonstrate the effectiveness and controllability of our method through extensive experiments. Although not trained to minimize CLIP loss, our model achieves equally high CLIP scores and significantly lower FID than prior work on short prompts, and outperforms the baseline qualitatively and quantitatively on long and complicated prompts. To our best knowledge, the proposed method is the first attempt at incorporating large-scale pre-trained diffusion models and distillation sampling for text-driven image generator domain adaptation and gives a quality previously beyond possible. Moreover, we extend our work to 3D-aware style-based generators and DreamBooth guidance. For code and more visual samples, please visit our Project Webpage. Kunpeng Song, Ligong Han, Dimitris N. Metaxas, Ahmed M. Elgammal |
WACV | 5 |
| 2022 | Spatial Frequency Bias in Convolutional Generative Adversarial NetworksabstractUnderstanding the capability of Generative Adversarial Networks (GANs) in learning the full spectrum of spatial frequencies, that is, beyond the low-frequency dominant spectrum of natural images, is critical for assessing the reliability of GAN-generated data in any detail-sensitive application. In this work, we show that the ability of convolutional GANs to learn an image distribution depends on the spatial frequency of the underlying carrier signal, that is, they have a bias against learning high spatial frequencies. Our findings are consistent with the recent observations of high-frequency artifacts in GAN-generated images, but further suggest that such artifacts are the consequence of an underlying bias. We also provide a theoretical explanation for this bias as the manifestation of linear dependencies present in the spectrum of filters of a typical generative Convolutional Neural Network (CNN). Finally, by proposing a proof-of-concept method that can effectively manipulate this bias towards other spatial frequencies, we show that the bias is not fixed and can be exploited to explicitly direct computational resources towards any specific spatial frequency of interest in a dataset, with minimal computational overhead. Mahyar Khayatkhoei, Ahmed M. Elgammal |
AAAI | 2 |
| 2022 | Proxy Learning of Visual Concepts of Fine Art Paintings from Styles through Language ModelsabstractWe present a machine learning system that can quantify fine art paintings with a set of visual elements and principles of art. The formal analysis is fundamental for understanding art, but developing such a system is challenging. Paintings have high visual complexities, but it is also difficult to collect enough training data with direct labels. To resolve these practical limitations, we introduce a novel mechanism, called proxy learning, which learns visual concepts in paintings through their general relation to styles. This framework does not require any visual annotation, but only uses style labels and a general relationship between visual concepts and style. In this paper, we propose a novel proxy model and reformulate four pre-existing methods in the context of proxy learning. Through quantitative and qualitative comparison, we evaluate these methods and compare their effectiveness in quantifying the artistic visual concepts, where the general relationship is estimated by language models; GloVe or BERT. The language modeling is a practical and scalable solution requiring no labeling, but it is inevitably imperfect. We demonstrate how the new proxy model is robust to the imperfection, while the other methods are sensitively affected by it. Diana Kim, Ahmed M. Elgammal, Marian Mazzone |
AAAI | 2 |
| 2021 | TIME: Text and Image Mutual-Translation Adversarial NetworksabstractFocusing on text-to-image (T2I) generation, we propose Text and Image Mutual-Translation Adversarial Networks (TIME), a lightweight but effective model that jointly learns a T2I generator G and an image captioning discriminator D under the Generative Adversarial Network framework. While previous methods tackle the T2I problem as a uni-directional task and use pre-trained language models to enforce the image--text consistency, TIME requires neither extra modules nor pre-training. We show that the performance of G can be boosted substantially by training it jointly with D as a language model. Specifically, we adopt Transformers to model the cross-modal connections between the image features and word embeddings, and design an annealing conditional hinge loss that dynamically balances the adversarial learning. In our experiments, TIME achieves state-of-the-art (SOTA) performance on the CUB dataset (Inception Score of 4.91 and Fréchet Inception Distance of 14.3 on CUB), and shows promising performance on MS-COCO dataset on image captioning and downstream vision-language tasks. Kunpeng Song, Yizhe Zhu, Gerard de Melo, Ahmed M. Elgammal |
AAAI | 5 |
| 2021 | Self-Supervised Sketch-to-Image SynthesisabstractImagining a colored realistic image from an arbitrary-drawn sketch is one of human capabilities that we eager machines to mimic. Unlike previous methods that either require the sketch-image pairs or utilize low-quantity detected edges as sketches, we study the exemplar-based sketch-to-image (s2i) synthesis task in a self-supervised learning manner, eliminating the necessity of the paired sketch data. To this end, we first propose an unsupervised method to efficiently synthesize line-sketches for general RGB-only datasets. With the synthetic paired-data, we then present a self-supervised Auto-Encoder (AE) to decouple the content/style features from sketches and RGB-images, and synthesize images both content-faithful to the sketches and style-consistent to the RGB-images. While prior works employ either the cycle-consistence loss or dedicated attentional modules to enforce the content/style fidelity, we show AE's superior performance with pure self-supervisions. To further improve the synthesis quality in high resolution, we also leverage an adversarial network to refine the details of synthetic images. Extensive experiments on $1024^2$ resolution demonstrate a new state-of-art-art performance of the proposed model on CelebA-HQ and Wiki-Art datasets. Moreover, with the proposed sketch generator, the model shows a promising performance on style mixing and style transfer, which the synthesized images are not only style-consistent but also semantically meaningful. Yizhe Zhu, Kunpeng Song, Ahmed M. Elgammal |
AAAI | 4 |
| 2021 | Towards Faster and Stabilized GAN Training for High-fidelity Few-shot Image Synthesis
Yizhe Zhu, Kunpeng Song, Ahmed M. Elgammal |
ICLR | 4 |
| 2020 | OOGAN: Disentangling GAN with One-Hot Sampling and Orthogonal RegularizationabstractExploring the potential of GANs for unsupervised disentanglement learning, this paper proposes a novel GAN-based disentanglement framework with One-Hot Sampling and Orthogonal Regularization (OOGAN). While previous works mostly attempt to tackle disentanglement learning through VAE and seek to implicitly minimize the Total Correlation (TC) objective with various sorts of approximation methods, we show that GANs have a natural advantage in disentangling with an alternating latent variable (noise) sampling method that is straightforward and robust. Furthermore, we provide a brand-new perspective on designing the structure of the generator and discriminator, demonstrating that a minor structural change and an orthogonal regularization on model weights entails an improved disentanglement. Instead of experimenting on simple toy datasets, we conduct experiments on higher-resolution images and show that OOGAN greatly pushes the boundary of unsupervised disentanglement. Yizhe Zhu, Zuohui Fu, Gerard de Melo, Ahmed M. Elgammal |
AAAI | 5 |
| 2020 | Sketch-to-Art: Synthesizing Stylized Art Images from Sketches
Kunpeng Song, Yizhe Zhu, Ahmed M. Elgammal |
ACCV (6) | 4 |
| 2020 | ULMFiT replicationabstractAuthors: Mohamed Abdellatif and Ahmed Elgammal Gitlab URL: https://gitlab.com/abdollatif/lrec_app Commit hash: 3f20b2ddb96d8c865e5f56f5566edf371214785f Tag name: Splits2 Dataset file md5: 5aee3dac5e48d1ac3d279083212734c9 Dataset URL: https://drive.google.com/file/d/1cv5HuQhgFVizupFI40dzreemS2gMM498/view?usp=sharing Mohamed Abdellatif, Ahmed M. Elgammal |
LREC | 2 |
| 2019 | Large-Scale Visual Relationship UnderstandingabstractLarge scale visual understanding is challenging, as it requires a model to handle the widely-spread and imbalanced distribution of 〈subject, relation, object〉 triples. In real-world scenarios with large numbers of objects and relations, some are seen very commonly while others are barely seen. We develop a new relationship detection model that embeds objects and relations into two vector spaces where both discriminative capability and semantic affinity are preserved. We learn a visual and a semantic module that map features from the two modalities into a shared space, where matched pairs of features have to discriminate against those unmatched, but also maintain close distances to semantically similar ones. Benefiting from that, our model can achieve superior performance even when the visual entity categories scale up to more than 80,000, with extremely skewed class distribution. We demonstrate the efficacy of our model on a large and imbalanced benchmark based of Visual Genome that comprises 53,000+ objects and 29,000+ relations, a scale at which no previous work has been evaluated at. We show superiority of our model over competitive baselines on the original Visual Genome dataset with 80,000+ categories. We also show state-of-the-art performance on the VRD dataset and the scene graph dataset which is a subset of Visual Genome with 200 categories. Yannis Kalantidis, Marcus Rohrbach, Manohar Paluri, Ahmed M. Elgammal, Mohamed Elhoseiny 0001 |
AAAI | 5 |
| 2019 | Graphical Contrastive Losses for Scene Graph ParsingabstractMost scene graph parsers use a two-stage pipeline to detect visual relationships: the first stage detects entities, and the second predicts the predicate for each entity pair using a softmax distribution. We find that such pipelines, trained with only a cross entropy loss over predicate classes, suffer from two common errors. The first, Entity Instance Confusion, occurs when the model confuses multiple instances of the same type of entity (e.g. multiple cups). The second, Proximal Relationship Ambiguity, arises when multiple subject-predicate-object triplets appear in close proximity with the same predicate, and the model struggles to infer the correct subject-object pairings (e.g. mis-pairing musicians and their instruments). We propose a set of contrastive loss formulations that specifically target these types of errors within the scene graph parsing problem, collectively termed the Graphical Contrastive Losses. These losses explicitly force the model to disambiguate related and unrelated instances through margin constraints specific to each type of confusion. We further construct a relationship detector, called RelDN, using the aforementioned pipeline to demonstrate the efficacy of our proposed losses. Our model outperforms the winning method of the OpenImages Relationship Detection Challenge by 4.7\% (16.5\% relatively) on the test set. We also show improved results over the best previous methods on the Visual Genome and Visual Relationship Detection datasets. Kevin J. Shih, Ahmed M. Elgammal, Andrew Tao, Bryan Catanzaro |
CVPR | 3 |
| 2019 | Computational Analysis of Content in Fine Art Paintings
Diana Kim, Ahmed M. Elgammal, Marian Mazzone |
ICCC | 3 |
| 2019 | Learning Feature-to-Feature Translator by Alternating Back-Propagation for Generative Zero-Shot LearningabstractWe investigate learning feature-to-feature translator networks by alternating back-propagation as a general-purpose solution to zero-shot learning (ZSL) problems. It is a generative model-based ZSL framework. In contrast to models based on generative adversarial networks (GAN) or variational autoencoders (VAE) that require auxiliary networks to assist the training, our model consists of a single conditional generator that maps class-level semantic features and Gaussian white noise vectors accounting for instance-level latent factors to visual features, and is trained by maximum likelihood estimation. The training process is a simple yet effective alternating back-propagation process that iterates the following two steps: (i) the inferential back-propagation to infer the latent noise vector of each observed example, and (ii) the learning back-propagation to update the model parameters. We show that, with slight modifications, our model is capable of learning from incomplete visual features for ZSL. We conduct extensive comparisons with existing generative ZSL methods on five benchmarks, demonstrating the superiority of our method in not only ZSL performance but also convergence speed and computational cost. Specifically, our model outperforms the existing state-of-the-art methods by a remarkable margin up to 3.1% and 4.0% in ZSL and generalized ZSL settings, respectively. Yizhe Zhu, Jianwen Xie, Ahmed M. Elgammal |
ICCV | 4 |
| 2019 | Semantic-Guided Multi-Attention Localization for Zero-Shot LearningabstractZero-shot learning extends the conventional object classification to the unseen class recognition by introducing semantic representations of classes. Existing approaches predominantly focus on learning the proper mapping function for visual-semantic embedding, while neglecting the effect of learning discriminative visual features. In this paper, we study the significance of the discriminative region localization. We propose a semantic-guided multi-attention localization model, which automatically discovers the most discriminative parts of objects for zero-shot learning without any human annotations. Our model jointly learns cooperative global and local features from the whole object as well as the detected parts to categorize objects based on semantic descriptions. Moreover, with the joint supervision of embedding softmax loss and class-center triplet loss, the model is encouraged to learn features with high inter-class dispersion and intra-class compactness. Through comprehensive experiments on three widely used zero-shot learning benchmarks, we show the efficacy of the multi-attention localization and our proposed approach improves the state-of-the-art results by a considerable margin. Yizhe Zhu, Jianwen Xie, Zhiqiang Tang 0001, Xi Peng 0005, Ahmed M. Elgammal |
NeurIPS | 5 |
| 2018 | Picasso, Matisse, or a Fake? Automated Analysis of Drawings at the Stroke Level for Attribution and AuthenticationabstractThis paper proposes a computational approach for analysis of strokes in line drawings by artists. We aim at developing an AI methodology that facilitates attribution of drawings of unknown authors in a way that is not easy to be deceived by forged art. The methodology used is based on quantifying the characteristics of individual strokes in drawings. We propose a novel algorithm for segmenting individual strokes. We propose an approach that combines different hand-crafted and learned features for the task of quantifying stroke characteristics. We experimented with a dataset of 300 digitized drawings with over 80 thousands strokes. The collection mainly consisted of drawings of Pablo Picasso, Henry Matisse, and Egon Schiele, besides a small number of representative works of other artists. The experiments shows that the proposed methodology can classify individual strokes with accuracy 70%-90%, and aggregate over drawings with accuracy above 80%, while being robust to be deceived by fakes. Ahmed M. Elgammal, Yan Kang 0004, Milko Den Leeuw |
AAAI | 1 |
| 2018 | The Shape of Art History in the Eyes of the MachineabstractHow does the machine classify styles in art? And how does it relate to art historians' methods for analyzing style? Several studies showed the ability of the machine to learn and predict styles, such as Renaissance, Baroque, Impressionism, etc., from images of paintings. This implies that the machine can learn an internal representation encoding discriminative features through its visual analysis. However, such a representation is not necessarily interpretable. We conducted a comprehensive study of several of the state-of-the-art convolutional neural networks applied to the task of style classification on 67K images of paintings, and analyzed the learned representation through correlation analysis with concepts derived from art history. Surprisingly, the networks could place the works of art in a smooth temporal arrangement mainly based on learning style labels, without any a priori knowledge of time of creation, the historical time and context of styles, or relations between styles. The learned representations showed that there are a few underlying factors that explain the visual variations of style in art. Some of these factors were found to correlate with style patterns suggested by Heinrich Wölfflin (1846-1945). The learned representations also consistently highlighted certain artists as the extreme distinctive representative of their styles, which quantitatively confirms art historian observations. Ahmed M. Elgammal, Diana Kim, Marian Mazzone |
AAAI | 1 |
| 2018 | A Generative Adversarial Approach for Zero-Shot Learning From Noisy TextsabstractMost existing zero-shot learning methods consider the problem as a visual semantic embedding one. Given the demonstrated capability of Generative Adversarial Networks(GANs) to generate images, we instead leverage GANs to imagine unseen categories from text descriptions and hence recognize novel classes with no examples being seen. Specifically, we propose a simple yet effective generative model that takes as input noisy text descriptions about an unseen class (e.g. Wikipedia articles) and generates synthesized visual features for this class. With added pseudo data, zero-shot learning is naturally converted to a traditional classification problem. Additionally, to preserve the inter-class discrimination of the generated features, a visual pivot regularization is proposed as an explicit supervision. Unlike previous methods using complex engineered regularizers, our approach can suppress the noise well without additional regularization. Empirically, we show that our method consistently outperforms the state of the art on the largest available benchmarks on Text-based Zero-shot Learning. Yizhe Zhu, Xi Peng 0005, Ahmed M. Elgammal |
CVPR | 5 |
| 2018 | Disconnected Manifold Learning for Generative Adversarial NetworksabstractNatural images may lie on a union of disjoint manifolds rather than one globally connected manifold, and this can cause several difficulties for the training of common Generative Adversarial Networks (GANs). In this work, we first show that single generator GANs are unable to correctly model a distribution supported on a disconnected manifold, and investigate how sample quality, mode dropping and local convergence are affected by this. Next, we show how using a collection of generators can address this problem, providing new insights into the success of such multi-generator GANs. Finally, we explain the serious issues caused by considering a fixed prior over the collection of generators and propose a novel approach for learning the prior and inferring the necessary number of generators without any supervision. Our proposed modifications can be applied on top of any other GAN model to enable learning of distributions supported on disconnected manifolds. We conduct several experiments to illustrate the aforementioned shortcoming of GANs, its consequences in practice, and the effectiveness of our proposed modifications in alleviating these issues. Mahyar Khayatkhoei, Maneesh Kumar Singh 0001, Ahmed M. Elgammal |
NeurIPS | 3 |
| 2017 | Sherlock: Scalable Fact Learning in ImagesabstractWe study scalable and uniform understanding of facts in images. Existing visual recognition systems are typically modeled differently for each fact type such as objects, actions, and interactions. We propose a setting where all these facts can be modeled simultaneously with a capacity to understand an unbounded number of facts in a structured way. The training data comes as structured facts in images, including (1) objects (e.g., ), (2) attributes (e.g., ), (3) actions (e.g., ), and (4) interactions (e.g., ). Each fact has a semantic language view (e.g., < boy, playing>) and a visual view (an image with this fact). We show that learning visual facts in a structured way enables not only a uniform but also generalizable visual understanding. We propose and investigate recent and strong approaches from the multiview learning literature and also introduce two learning representation models as potential baselines. We applied the investigated methods on several datasets that we augmented with structured facts and a large scale dataset of more than 202,000 facts and 814,000 images. Our experiments show the advantage of relating facts by the structure by the proposed models compared to the designed baselines on bidirectional fact retrieval. Scott Cohen, Walter Chang, Brian L. Price, Ahmed M. Elgammal |
AAAI | 5 |
| 2017 | Link the Head to the "Beak": Zero Shot Learning from Noisy Text Description at Part PrecisionabstractIn this paper, we study learning visual classifiers from unstructured text descriptions at part precision with no training images. We propose a learning framework that is able to connect text terms to its relevant parts and suppress connections to non-visual text terms without any part-text annotations. For instance, this learning process enables terms like beak to be sparsely linked to the visual representation of parts like head, while reduces the effect of non-visual terms like migrate on classifier prediction. Images are encoded by a part-based CNN that detect bird parts and learn part-specific representation. Part-based visual classifiers are predicted from text descriptions of unseen visual classifiers to facilitate classification without training images (also known as zero-shot recognition). We performed our experiments on CUBirds 2011 dataset and improves the state-of-the-art textbased zero-shot recognition results from 34.7% to 43.6%. We also created large scale benchmarks on North American Bird Images augmented with text descriptions, where we also show that our approach outperforms existing methods. Our code, data, and models are publically available link [1]. Yizhe Zhu, Han Zhang 0010, Ahmed M. Elgammal |
CVPR | 4 |
| 2017 | Relationship Proposal NetworksabstractImage scene understanding requires learning the relationships between objects in the scene. A scene with many objects may have only a few individual interacting objects (e.g., in a party image with many people, only a handful of people might be speaking with each other). To detect all relationships, it would be inefficient to first detect all individual objects and then classify all pairs, not only is the number of all pairs quadratic, but classification requires limited object categories, which is not scalable for real-world images. In this paper we address these challenges by using pairs of related regions in images to train a relationship proposer that at test time produces a manageable number of related regions. We name our model the Relationship Proposal Network (Rel-PN). Like object proposals, our Rel-PN is class-agnostic and thus scalable to an open vocabulary of objects. We demonstrate the ability of our Rel-PN to localize relationships with only a few thousand proposals. We demonstrate its performance on the Visual Genome dataset and compare to other baselines that we designed. We also conduct experiments on a smaller subset of 5,000 images with over 37,000 related regions and show promising results. Scott Cohen, Walter Chang, Ahmed M. Elgammal |
CVPR | 5 |
| 2017 | CAN: Creative Adversarial Networks, Generating "Art" by Learning About Styles and Deviating from Style Norms
Ahmed M. Elgammal, Marian Mazzone |
ICCC | 1 |
| 2017 | A Multilayer-Based Framework for Online Background Subtraction with Freely Moving CamerasabstractThe exponentially increasing use of moving platforms for video capture introduces the urgent need to develop the general background subtraction algorithms with the capability to deal with the moving background. In this paper, we propose a multilayer-based framework for online background subtraction for videos captured by moving cameras. Unlike the previous treatments of the problem, the proposed method is not restricted to binary segmentation of background and foreground, but formulates it as a multi-label segmentation problem by modeling multiple foreground objects in different layers when they appear simultaneously in the scene. We assign an independent processing layer to each foreground object, as well as the background, where both motion and appearance models are estimated, and a probability map is inferred using a Bayesian filtering framework. Finally, Multi-label Graph-cut on Markov Random Field is employed to perform pixel-wise labeling. Extensive evaluation results show that the proposed method outperforms state-of-the-art methods on challenging video sequences. Yizhe Zhu, Ahmed M. Elgammal |
ICCV | 2 |
| 2017 | On the effect of hyperedge weights on hypergraph learning
Sheng Huang 0001, Ahmed M. Elgammal, Dan Yang 0001 |
Image Vis. Comput. | 2 |
| 2017 | Write a Classifier: Predicting Visual Classifiers from Unstructured TextabstractPeople typically learn through exposure to visual concepts associated with linguistic descriptions. For instance, teaching visual object categories to children is often accompanied by descriptions in text or speech. In a machine learning context, these observations motivates us to ask whether this learning process could be computationally modeled to learn visual classifiers. More specifically, the main question of this work is how to utilize purely textual description of visual classes with no training images, to learn explicit visual classifiers for them. We propose and investigate two baseline formulations, based on regression and domain transfer, that predict a linear classifier. Then, we propose a new constrained optimization formulation that combines a regression function and a knowledge transfer function with additional constraints to predict the parameters of a linear classifier. We also propose a generic kernelized models where a kernel classifier is predicted in the form defined by the representer theorem. The kernelized models allow defining and utilizing any two Reproducing Kernel Hilbert Space (RKHS) kernel functions in the visual space and text space, respectively. We finally propose a kernel function between unstructured text descriptions that builds on distributional semantics, which shows an advantage in our setting and could be useful for other applications. We applied all the studied models to predict visual classifiers on two fine-grained and challenging categorization datasets (CU Birds and Flower Datasets), and the results indicate successful predictions of our final model over several baselines that we designed. Mohamed Elhoseiny 0002, Ahmed M. Elgammal, Babak Saleh |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2017 | Modeling depth for nonparametric foreground segmentation using RGBD devices
Biel Moyà Alcover, Ahmed M. Elgammal, Antoni Jaume-i-Capó, Javier Varona |
Pattern Recognit. Lett. | 2 |
| 2016 | Zero-Shot Event Detection by Multimodal Distributional Semantic Embedding of VideosabstractWe propose a new zero-shot Event-Detection method by Multi-modal Distributional Semantic embedding of videos. Our model embeds object and action concepts as well as other available modalities from videos into a distributional semantic space. To our knowledge, this is the first Zero-Shot event detection model that is built on top of distributional semantics and extends it in the following directions: (a) semantic embedding of multimodal information in videos (with focus on the visual modalities), (b) semantic embedding of concepts definitions, and (c) retrieve videos by free text event query (e.g., "changing a vehicle tire") based on their content. We first embed the video into the multi-modal semantic space and then measure the similarity between videos with the event query in free text form. We validated our method on the large TRECVID MED (Multimedia Event Detection) challenge. Using only the event title as a query, our method outperformed the state-the-art that uses big descriptions from 12.6\% to 13.5\% with MAP metric and from 0.73 to 0.83 with ROC-AUC metric. It is also an order of magnitude faster. Jingen Liu, Harpreet Sawhney, Ahmed M. Elgammal |
AAAI | 5 |
| 2016 | Toward a Taxonomy and Computational Models of Abnormalities in ImagesabstractThe human visual system can spot an abnormal image, and reason about what makes it strange. This task has not received enough attention in computer vision. In this paper we study various types of atypicalities in images in a more comprehensive way than has been done before. We propose a new dataset of abnormal images showing a wide range of atypicalities. We design human subject experiments to discover a coarse taxonomy of the reasons for abnormality. Our experiments reveal three major categories of abnormality: object-centric, scene-centric, and contextual. Based on this taxonomy, we propose a comprehensive computational model that can predict all different types of abnormality in images and outperform prior arts in abnormality recognition. Babak Saleh, Ahmed M. Elgammal, Jacob Feldman, Ali Farhadi |
AAAI | 2 |
| 2016 | SPDA-CNN: Unifying Semantic Part Detection and Abstraction for Fine-Grained RecognitionabstractMost convolutional neural networks (CNNs) lack midlevel layers that model semantic parts of objects. This limits CNN-based methods from reaching their full potential in detecting and utilizing small semantic parts in recognition. Introducing such mid-level layers can facilitate the extraction of part-specific features which can be utilized for better recognition performance. This is particularly important in the domain of fine-grained recognition. In this paper, we propose a new CNN architecture that integrates semantic part detection and abstraction (SPDACNN) for fine-grained classification. The proposed network has two sub-networks: one for detection and one for recognition. The detection sub-network has a novel top-down proposal method to generate small semantic part candidates for detection. The classification sub-network introduces novel part layers that extract features from parts detected by the detection sub-network, and combine them for recognition. As a result, the proposed architecture provides an end-to-end network that performs detection, localization of multiple semantic parts, and whole object recognition within one framework that shares the computation of convolutional filters. Our method outperforms state-of-theart methods with a large margin for small parts detection (e.g. our precision of 93.40% vs the best previous precision of 74.00% for detecting the head on CUB-2011). It also compares favorably to the existing state-of-the-art on finegrained classification, e.g. it achieves 85.14% accuracy on CUB-2011. Han Zhang 0010, Tao Xu 0029, Sharon X. Huang, Shaoting Zhang 0001, Ahmed M. Elgammal, Dimitris N. Metaxas |
CVPR | 6 |
| 2016 | A Comparative Analysis and Study of Multiview CNN Models for Joint Object Categorization and Pose EstimationabstractIn the Object Recognition task, there exists a dichotomy between the categorization of objects and estimating object pose, where the former necessitates a view-invariant representation, while the latter requires a representation capable of capturing pose information over different categories of objects. With the rise of deep architectures, the prime focus has been on object category recognition. Deep learning methods have achieved wide success in this task. In contrast, object pose estimation using these approaches has received relatively less attention. In this work, we study how Convolutional Neural Networks (CNN) architectures can be adapted to the task of simultaneous object recognition and pose estimation. We investigate and analyze the layers of various CNN models and extensively compare between them with the goal of discovering how the layers of distributed representations within CNNs represent object pose information and how this contradicts with object category representations. We extensively experiment on two recent large and challenging multi-view datasets and we achieve better than the state-of-the-art. Tarek El-Gaaly, Amr Bakry, Ahmed M. Elgammal |
ICML | 4 |
| 2016 | Incorporating Prototype Theory in Convolutional Neural Networks
Babak Saleh, Ahmed M. Elgammal, Jacob Feldman |
IJCAI | 2 |
| 2016 | Joint object recognition and pose estimation using a nonlinear view-invariant latent generative modelabstractObject recognition and pose estimation are two fundamental problems in the field of computer vision. Recognizing objects and their poses/viewpoints are critical components of ample vision and robotic systems. Multiple viewpoints of an object lie on an intrinsic low-dimensional manifold in the input space (i.e. descriptor space). Different objects captured from the same set of viewpoints have manifolds with a common topology. In this paper we utilize this common topology between object manifolds by learning a low-dimensional latent space which non-linearly maps between a common unified manifold and the object manifold in the input space. Using a supervised embedding approach, the latent space is computed and used to jointly infer the category and pose of objects. We empirically validate our model by using multiple inference approaches and testing on multiple challenging datasets. We compare our results with the state-of-the-art and present our increased category recognition and pose estimation accuracy. Amr Bakry, Tarek El-Gaaly, Ahmed M. Elgammal |
WACV | 4 |
| 2016 | Text to multi-level MindMaps - A novel method for hierarchical visual abstraction of natural language text
Ahmed M. Elgammal |
Multim. Tools Appl. | 2 |
| 2016 | Toward automated discovery of artistic influence
Babak Saleh, Kanako Abe, Ravneet Singh Arora, Ahmed M. Elgammal |
Multim. Tools Appl. | 4 |
| 2016 | Learning representations from multiple manifolds
Chan-Su Lee, Ahmed M. Elgammal, Marwan Torki |
Pattern Recognit. | 2 |
| 2016 | Collaborative Graph Embedding: A Simple Way to Generally Enhance Subspace Learning AlgorithmsabstractCollaborative representation (CR), known as an effective way to address the signal representation (regression) problem, has achieved remarkable success in visual classification. According to our theoretical analysis, the subspace learning issue can also be deemed as a signal representation problem. Therefore, we extend the graph embedding (GE) framework as a CR model to improve the discriminating power of the subspace learning algorithm. The new GE framework, which is named collaborative GE (CGE) framework, enjoys many desirable properties of CR. From theoretical analysis, CGE is robust to the noise and has the same computational complexity as GE. From experimental analysis, CGE can generally enhance the subspace learning algorithms and a reasonable regularization parameter can be inferred from its intrinsic graph. Several state-of-the-art subspace learning algorithms are plugged into our framework to produce their collaborative versions. Meanwhile, by exploring the intrinsic relation among GE methods, we present a new collaborative method named collaborative class-scattering locality preserving projections (CCSLPPs). The results of extensive experiments on ORL, AR, Scene15, Caltech256, LFW-A, and OU-ISIR-A databases demonstrate that the collaborative versions consistently outperform their original algorithms with a remarkable improvement and CCSLPP gets the best performance compared with all used methods. Sheng Huang 0001, Yu Yang 0010, Dan Yang 0001, Ahmed M. Elgammal |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2015 | A Bayesian Approach to Perceptual 3D Object-Part Decomposition Using Skeleton-Based RepresentationsabstractWe present a probabilistic approach to shape decomposition that creates a skeleton-based shape representation of a 3D object while simultaneously decomposing it into constituent parts. Our approach probabilistically combines two prominent threads from the shape literature: skeleton-based (medial axis) representations of shape, and part-based representations of shape, in which shapes are combinations of primitive parts. Our approach recasts skeleton-based shape representation as a mixture estimation problem, allowing us to apply probabilistic estimation techniques to the problem of 3D shape decomposition, extending earlier work on the 2D case. The estimated 3D shape decompositions approximate human shape decomposition judgments. We present a tractable implementation of the framework, which begins by over-segmenting objects at concavities, and then probabilistically merges them to create a distribution over possible decompositions. This results in a hierarchy of decompositions at different structural scales, again closely matching known properties of human shape representation. The probabilistic estimation procedures that arise naturally in the model allow effective prediction of missing parts. We present results on shapes from a standard database illustrating the effectiveness of the approach. Tarek El-Gaaly, Vicky Froyen, Ahmed M. Elgammal, Jacob Feldman, Manish Singh 0001 |
AAAI | 3 |
| 2015 | Overlapping Domain Cover for Scalable and Accurate Regression Kernel MachinesabstractIn this paper, we present the Overlapping Domain Cover (ODC) notion for kernel machines, as a set of overlapping subsets of the data that covers the entire training set and optimized to be spatially cohesive as possible. We propose an efficient ODC framework, which is applicable to various regression models and in particular reduces the complexity of Twin Gaussian Processes (TGP) regression from cubic to quadratic. We also theoretically justified the idea behind our method. We validated and analyzed our method on three human pose estimation datasets and interesting findings are discussed. Ahmed M. Elgammal |
BMVC | 2 |
| 2015 | Learning Hypergraph-regularized Attribute PredictorsabstractWe present a novel attribute learning framework named Hypergraph-based Attribute Predictor (HAP). In HAP, a hypergraph is leveraged to depict the attribute relations in the data. Then the attribute prediction problem is casted as a regularized hypergraph cut problem, in which a collection of attribute projections is jointly learnt from the feature space to a hypergraph embedding space aligned with the attributes. The learned projections directly act as attribute classifiers (linear and kernelized). This formulation leads to a very efficient approach. By considering our model as a multi-graph cut task, our framework can flexibly incorporate other available information, in particular class label. We apply our approach to attribute prediction, Zero-shot and N-shot learning tasks. The results on AWA, USAA and CUB databases demonstrate the value of our methods in comparison with the state-of-the-art approaches. Sheng Huang 0001, Ahmed M. Elgammal, Dan Yang 0001 |
CVPR | 3 |
| 2015 | Quantifying Creativity in Art Networks
Ahmed M. Elgammal, Babak Saleh |
ICCC | 1 |
| 2015 | Weather classification with deep convolutional neural networksabstractIn this paper, we study weather classification from images using Convolutional Neural Networks (CNNs). Our approach outperforms the state of the art by a huge margin in the weather classification task. Our approach achieves 82.2% normalized classification accuracy instead of 53.1% for the state of the art (i.e., 54.8% relative improvement). We also studied the behavior of all the layers of the Convolutional Neural Networks, we adopted, and interesting findings are discussed. Sheng Huang 0001, Ahmed M. Elgammal |
ICIP | 3 |
| 2015 | From circle to 3-sphere: Head pose estimation by instance parameterization
Xi Peng 0005, Junzhou Huang, Qiong Hu 0001, Shaoting Zhang 0001, Ahmed M. Elgammal, Dimitris N. Metaxas |
Comput. Vis. Image Underst. | 5 |
| 2015 | Factorization of view-object manifolds for joint object recognition and pose estimation
Haopeng Zhang 0001, Tarek El-Gaaly, Ahmed M. Elgammal, Zhiguo Jiang 0001 |
Comput. Vis. Image Underst. | 3 |
| 2015 | Generalized Twin Gaussian processes using Sharma-Mittal divergence
Ahmed M. Elgammal |
Mach. Learn. | 2 |
| 2015 | Cross-Speed Gait Recognition Using Speed-Invariant Gait Templates and Globality-Locality Preserving ProjectionsabstractWe present a novel manifold-based approach for cross-speed gait recognition. In our approach, the walking action is considered as residing on a manifold, in the feature space, that is homomorphic to a unit circle. We employ thin plate spline (TPS) kernel-based radial basis function (RBF) interpolation to fit such manifold. TPS kernel-based RBF interpolation separates the learned coefficients into an affine component and a nonaffine component, which, respectively, encodes the dynamic and static characteristics of the gait manifold. We introduce the use of the nonaffine component as a cross-speed gait representation, and denote it speed invariant gait template (SIGT). We also propose an enhanced locality preserving projections (LPP) algorithm named globality LPP (GLPP) for reducing the dimension of SIGT. In GLPP, the graph Laplacians of intrasubject part and intersubjects part are separately constructed, and then to combine as a new graph Laplacian. Finally, a manifold learning-based classifier named normalized hypergraph classifier is employed for classification. Experimental results on two gait databases demonstrate the effectiveness of our proposed approach in comparison with the state-of-the-art gait recognition methods. Sheng Huang 0001, Ahmed M. Elgammal, Jiwen Lu, Dan Yang 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2014 | Simultaneous Twin Kernel Learning Using Polynomial Transformations for Structured PredictionabstractMany learning problems in computer vision can be posed as structured prediction problems, where the input and output instances are structured objects such as trees, graphs or strings rather than, single labels {+1, -1} or scalars. Kernel methods such as Structured Support Vector Machines, Twin Gaussian Processes (TGP), Structured Gaussian Processes, and vector-valued Reproducing Kernel Hilbert Spaces (RKHS), offer powerful ways to perform learning and inference over these domains. Positive definite kernel functions allow us to quantitatively capture similarity between a pair of instances over these arbitrary domains. A poor choice of the kernel function, which decides the RKHS feature space, often results in poor performance. Automatic kernel selection methods have been developed, but have focused only on kernels on the input domain (i.e.'one-way'). In this work, we propose a novel and efficient algorithm for learning kernel functions simultaneously, on both input and output domains. We introduce the idea of learning polynomial kernel transformations, and call this method Simultaneous Twin Kernel Learning (STKL). STKL can learn arbitrary, but continuous kernel functions, including 'one-way' kernel learning as a special case. We formulate this problem for learning covariances kernels of Twin Gaussian Processes. Our experimental evaluation using learned kernels on synthetic and several real-world datasets demonstrate consistent improvement in performance of TGP's. Chetan Tonde, Ahmed M. Elgammal |
CVPR | 2 |
| 2014 | Untangling Object-View Manifold for Multiview Recognition and Pose Estimation
Amr Bakry, Ahmed M. Elgammal |
ECCV (4) | 2 |
| 2014 | Knowledge Discovery of Artistic Influences: A Metric Learning Approach
Babak Saleh, Kanako Abe, Ahmed M. Elgammal |
ICCC | 3 |
| 2014 | Improving non-negative matrix factorization via ranking its basesabstractAs a considerable technique in image processing and computer vision, Nonnegative Matrix Factorization (NMF) generates its bases by iteratively multiplicative update with two initial random nonnegative matrices W and H, that leads to the randomness of the bases selection. For this reason, the potentials of NMF algorithms are not completely exploited. To address this issue, we present a novel framework which uses the feature selection techniques to evaluate and rank the bases of the NMF algorithms to enhance the NMF algorithms. We adopted the well known Fisher criterion and Least Reconstruction Error criterion, which is proposed by us, as two instances to show how that works successfully under our framework. Moreover, in order to avoid the hard combinatorial optimization issue in ranking procedure, a de-correlation constraint can be optionally imposed to the NMF algorithms for giving a better approximation to the global optimum of the NMF projections. We evaluate our works in face recognition, object recognition and image reconstruction on ORL and ETH-80 databases and the results demonstrate the enhancement of the state-of-the-art NMF under our framework. Sheng Huang 0001, Ahmed M. Elgammal, Dan Yang 0001 |
ICIP | 3 |
| 2014 | Spatial-Visual Label Propagation for Local Feature ClassificationabstractIn this paper we present a novel approach to integrate feature similarity and spatial consistency of local features to achieve the goal of localizing an object of interest in an image. The goal is to achieve coherent and accurate labeling of feature points in a simple and effective way. We introduced our Spatial-Visual Label Propagation algorithm to infer the labels of local features in a test image from known labels. This is done in a transductive manner to provide spatial and feature smoothing over the learned labels. We show the value of our novel approach by a diverse set of experiments with successful improvements over previous methods and baseline classifiers. Tarek El-Gaaly, Marwan Torki, Ahmed M. Elgammal |
ICPR | 3 |
| 2014 | Hierarchical Semantic Hashing: Visual Localization from Buildings on MapsabstractIn this paper we present a vision-based method for instant global localization from a given aerial image. The approach mimics how humans localize themselves on maps using spatial layout of semantic elements on the map. Unlike other matching and localization methods that use visual appearance or feature matching, our method relies on robust and consistently detectable semantic elements that are invariant to illumination, temporal variations and occlusions. We use the buildings on the map and on the given aerial query image as our semantic elements. Spatial relations between these elements are efficiently stored and queried under a hierarchical semantic version of the Geometric Hashing algorithm that is inherently rotation and scale invariant. We also present a method to obtain building locations from a given query image using image classification and processing techniques. Overall this approach provides fast and robust localization over large areas. We show our experimental results for localizing satellite image tiles from a 16.5 km sq dense city map with over 7,000 buildings. Turgay Senlet, Tarek El-Gaaly, Ahmed M. Elgammal |
ICPR | 3 |
| 2013 | Joint Object and Pose Recognition Using Homeomorphic Manifold AnalysisabstractObject recognition is a key precursory challenge in the fields of object manipulation and robotic/AI visual reasoning in general. Recognizing object categories, particular instances of objects and viewpoints/poses of objects are three critical subproblems robots must solve in order to accurately grasp/manipulate objects and reason about their environ- ments. Multi-view images of the same object lie on intrinsic low-dimensional manifolds in descriptor spaces (e.g. visual/depth descriptor spaces). These object manifolds share the same topology despite being geometrically different. Each object manifold can be represented as a deformed version of a unified manifold. The object manifolds can thus be parametrized by its homeomorphic mapping/reconstruction from the unified manifold. In this work, we construct a manifold descriptor from this mapping between homeomorphic manifolds and use it to jointly solve the three challenging recognition sub-problems. We extensively experiment on a challenging multi-modal (i.e. RGBD) dataset and other object pose datasets and achieve state-of-the-art results. Haopeng Zhang 0001, Tarek El-Gaaly, Ahmed M. Elgammal, Zhiguo Jiang 0001 |
AAAI | 3 |
| 2013 | Learning Speed Invariant Gait Template via Thin Plate Spline Kernel Manifold FittingabstractWe present a novel approach for cross-speed gait recognition. In our approach, the cyclic walking action is considered as residing on a manifold which is homeomorphic to a unit circle in the gait space. Thin Plate Spline (TPS) kernel-based Radial Basis Function (RBF) interpolation is used to fit the walking manifold for each gait sequence. The sub-ject related kernel mapping coefficients are learned for representing the gait. According to the property of TPS, the coefficients can be naturally separated as an affine component and a non-affine component. The affine component is the style factor corresponding to the deformation of the homeomorphic manifold caused by the walking action, while the non-affine component is the shape factor, invariant to the walking speed. We denote this non-affine component as Speed Invariant Gait Template (SIGT) and use it as cross-speed gait feature. To address the curse of dimensionality issue and speed up the recognition, we use Globality Locality Preserving Projections (GLPP) to reduce the dimensions of SIGTs. Two walking speeds related gait databases are employed for evaluating our pro-posed method. The experimental results demonstrate the superiority of our method over the state-of-the-art. 1 Sheng Huang 0001, Ahmed M. Elgammal, Dan Yang 0001 |
BMVC | 2 |
| 2013 | MKPLS: Manifold Kernel Partial Least Squares for Lipreading and Speaker IdentificationabstractVisual speech recognition is a challenging problem, due to confusion between visual speech features. The speaker identification problem is usually coupled with speech recognition. Moreover, speaker identification is important to several applications, such as automatic access control, biometrics, authentication, and personal privacy issues. In this paper, we propose a novel approach for lip reading and speaker identification. We propose a new approach for manifold parameterization in a low-dimensional latent space, where each manifold is represented as a point in that space. We initially parameterize each instance manifold using a nonlinear mapping from a unified manifold representation. We then factorize the parameter space using Kernel Partial Least Squares (KPLS) to achieve a low-dimension manifold latent space. We use two-way projections to achieve two manifold latent spaces, one for the speech content and one for the speaker. We apply our approach on two public databases: AVLetters and OuluVS. We show the results for three different settings of lip reading: speaker independent, speaker dependent, and speaker semi-dependent. Our approach outperforms for the speaker semi-dependent setting by at least 15% of the baseline, and competes in the other two settings. Amr Bakry, Ahmed M. Elgammal |
CVPR | 2 |
| 2013 | Object-Centric Anomaly Detection by Attribute-Based ReasoningabstractWhen describing images, humans tend not to talk about the obvious, but rather mention what they find interesting. We argue that abnormalities and deviations from typicalities are among the most important components that form what is worth mentioning. In this paper we introduce the abnormality detection as a recognition problem and show how to model typicalities and, consequently, meaningful deviations from prototypical properties of categories. Our model can recognize abnormalities and report the main reasons of any recognized abnormality. We also show that abnormality predictions can help image categorization. We introduce the abnormality detection dataset and show interesting results on how to reason about abnormalities. Babak Saleh, Ali Farhadi, Ahmed M. Elgammal |
CVPR | 3 |
| 2013 | Write a Classifier: Zero-Shot Learning Using Purely Textual DescriptionsabstractThe main question we address in this paper is how to use purely textual description of categories with no training images to learn visual classifiers for these categories. We propose an approach for zero-shot learning of object categories where the description of unseen categories comes in the form of typical text such as an encyclopedia entry, without the need to explicitly defined attributes. We propose and investigate two baseline formulations, based on regression and domain adaptation. Then, we propose a new constrained optimization formulation that combines a regression function and a knowledge transfer function with additional constraints to predict the classifier parameters for new classes. We applied the proposed approach on two fine-grained categorization datasets, and the results indicate successful classifier prediction. Babak Saleh, Ahmed M. Elgammal |
ICCV | 3 |
| 2013 | Online Motion Segmentation Using Dynamic Label PropagationabstractThe vast majority of work on motion segmentation adopts the affine camera model due to its simplicity. Under the affine model, the motion segmentation problem becomes that of subspace separation. Due to this assumption, such methods are mainly offline and exhibit poor performance when the assumption is not satisfied. This is made evident in state-of-the-art methods that relax this assumption by using piecewise affine spaces and spectral clustering techniques to achieve better results. In this paper, we formulate the problem of motion segmentation as that of manifold separation. We then show how label propagation can be used in an online framework to achieve manifold separation. The performance of our framework is evaluated on a benchmark dataset and achieves competitive performance while being online. Ali Elqursh, Ahmed M. Elgammal |
ICCV | 2 |
| 2013 | Homeomorphic Manifold Analysis (HMA): Generalized separation of style and content on manifolds
Ahmed M. Elgammal, Chan-Su Lee |
Image Vis. Comput. | 1 |
| 2012 | Online Moving Camera Background Subtraction
Ali Elqursh, Ahmed M. Elgammal |
ECCV (6) | 2 |
| 2012 | Towards automated classification of fine-art painting style: A comparative study
Ravneet Singh Arora, Ahmed M. Elgammal |
ICPR | 2 |
| 2012 | Video figure ground labeling
Ali Elqursh, Ahmed M. Elgammal |
ICPR | 2 |
| 2012 | Single axis relative rotation from orthogonal lines
Ali Elqursh, Ahmed M. Elgammal |
ICPR | 2 |
| 2012 | Segmentation of occluded sidewalks in satellite images
Turgay Senlet, Ahmed M. Elgammal |
ICPR | 2 |
| 2012 | Satellite image based precise robot localization on sidewalksabstractIn this paper, we present a novel computer vision framework for precise localization of a mobile robot on sidewalks. In our framework, we combine stereo camera images, visual odometry, satellite map matching, and a sidewalk probability transfer function obtained from street maps in order to attain globally corrected localization results. The framework is capable of precisely localizing a mobile robot platform that navigates on sidewalks, without the use of traditional wheel odometry, GPS or INS inputs. On a complex 570-meter sidewalk route, we show that we obtain superior localization results compared to visual odometry and GPS. Turgay Senlet, Ahmed M. Elgammal |
ICRA | 2 |
| 2012 | English2MindMap: An Automated System for MindMap Generation from English TextabstractMind Mapping is a well-known technique used in note taking and is known to encourage learning and studying. Besides, Mind Mapping can be a very good way to present knowledge and concepts in a visual form. Unfortunately there is no reliable automated tool that can generate Mind Maps from Natural Language text. This paper fills in this gap by developing the first evaluated automated system that takes a text input and generates a Mind Map visualization out of it. The system also could visualize large text documents in multilevel Mind Maps in which a high level Mind Map node could be expanded into child Mind Maps. The proposed approach involves understanding of the input text converting it into intermediate Detailed Meaning Representation (DMR). The DMR is then visualized with two proposed approaches, Single level or Multiple levels which is convenient for larger text. The generated Mind Maps from both approaches were evaluated based on Human Subject experiments performed on Amazon Mechanical Turk with various parameter settings. Ahmed M. Elgammal |
ISM | 2 |
| 2012 | Style adaptive contour tracking of human gait using explicit manifold models
Chan-Su Lee, Ahmed M. Elgammal |
Mach. Vis. Appl. | 2 |
| 2011 | Line-based relative pose estimationabstractWe present an algorithm for calibrated camera relative pose estimation from lines. Given three lines with two of the lines parallel and orthogonal to the third we can compute the relative rotation between two images. We can also compute the relative translation from two intersection points. We also present a framework in which such lines can be detected. We evaluate the performance of the algorithm using synthetic and real data. The intended use of the algorithm is with robust hypothesize-and-test frameworks such as RANSAC. Our approach is suitable for urban and indoor environments where most lines are either parallel or orthogonal to each other. Ali Elqursh, Ahmed M. Elgammal |
CVPR | 2 |
| 2011 | Supervised hypergraph labelingabstractWe address the problem of labeling individual datapoints given some knowledge about (small) subsets or groups of them. The knowledge we have for a group is the likelihood value for each group member to satisfy a certain model. This problem is equivalent to hypergraph labeling problem where each datapoint corresponds to a node and the each subset correspond to a hyperedge with likelihood value as its weight. We propose a novel method to model the label dependence using an Undirected Graphical Model and reduce the problem of hypergraph labeling into an inference problem. This paper describes the structure and necessary components of such model and proposes useful cost functions. We discuss the behavior of proposed algorithm with different forms of the cost functions, identify suitable algorithms for inference and analyze required properties when it is theoretically guaranteed to have exact solution. Examples of several real world problems are shown as applications of the proposed method. Toufiq Parag, Ahmed M. Elgammal |
CVPR | 2 |
| 2011 | Regression from local features for viewpoint and pose estimationabstractIn this paper we propose a framework for learning a regression function form a set of local features in an image. The regression is learned from an embedded representation that reflects the local features and their spatial arrangement as well as enforces supervised manifold constraints on the data. We applied the approach for viewpoint estimation on a Multiview car dataset, a head pose dataset and arm posture dataset. The experimental results show that this approach has superior results (up to 67% improvement) to the state-of-the-art approaches in very challenging datasets. Marwan Torki, Ahmed M. Elgammal |
ICCV | 2 |
| 2010 | Putting local features on a manifoldabstractLocal features have proven very useful for recognition. Manifold learning has proven to be a very powerful tool in data analysis. However, manifold learning application for images are mainly based on holistic vectorized representations of images. The challenging question that we address in this paper is how can we learn image manifolds from a punch of local features in a smooth way that captures the feature similarity and spatial arrangement variability between images. We introduce a novel framework for learning a manifold representation from collections of local features in images. We first show how we can learn a feature embedding representation that preserves both the local appearance similarity as well as the spatial structure of the features. We also show how we can embed features from a new image by introducing a solution for the out-of-sample that is suitable for this context. By solving these two problems and defining a proper distance measure in the feature embedding space, we can reach an image manifold embedding space. Marwan Torki, Ahmed M. Elgammal |
CVPR | 2 |
| 2010 | One-shot multi-set non-rigid feature-spatial matchingabstractWe introduce a novel framework for nonrigid feature matching among multiple sets in a way that takes into consideration both the feature descriptor and the features spatial arrangement. We learn an embedded representation that combines both the descriptor similarity and the spatial arrangement in a unified Euclidean embedding space. This unified embedding is reached by minimizing an objective function that has two sources of weights; the feature spatial arrangement and the feature descriptor similarity scores across the different sets. The solution can be obtained directly by solving one Eigen-value problem that is linear in the number of features. Therefore, the framework is very efficient and can scale up to handle a large number of features. Experimental evaluation is done using different sets showing outstanding results compared to the state of the art; up to 100% accuracy is achieved in the case of the well known `Hotel' sequence. Marwan Torki, Ahmed M. Elgammal |
CVPR | 2 |
| 2010 | Object Localization by Propagating Connectivity via SuperfeaturesabstractIn this paper, we propose a part-based approach to localize objects in cluttered images. We represent object parts as boundary segments and image patches. A semi-local grouping of parts named superfeatures encodes appearance and connectivity within a neighborhood. To match parts, we integrate inter-feature similarities and intra-feature connectivity via a relaxation labeling framework. Additionally, we use a global elliptical shape prior to match the shape of the solution space to that of the object. To this end, we demonstrate the efficacy of the method for detecting various objects in cluttered images by comparing them to simple object models. Ishani Chakraborty, Ahmed M. Elgammal |
ICPR | 2 |
| 2010 | Learning a Joint Manifold Representation from Multiple Data SetsabstractThe problem we address in the paper is how to learn a joint representation from data lying on multiple manifolds. We are given multiple data sets and there is an underlying common manifold among the different data set. We propose a framework to learn an embedding of all the points on all the manifolds in a way that preserves the local structure on each manifold and, in the same time, collapses all the different manifolds into one manifold in the embedding space, while preserving the implicit correspondences between the points across different data sets. The proposed solution works as extensions to current state of the art spectral-embedding approaches to handle multiple manifolds. Marwan Torki, Ahmed M. Elgammal, Chan-Su Lee |
ICPR | 2 |
| 2010 | Coupled Visual and Kinematic Manifold Models for Tracking
Chan-Su Lee, Ahmed M. Elgammal |
Int. J. Comput. Vis. | 2 |
| 2010 | Dynamic Shape Style Analysis: Bilinear and Multilinear Human Identification with Temporal NormalizationabstractModeling and analyzing the dynamic shape of human motion is a challenging task owing to temporal variations in the shape and multiple sources of observed shape variations such as viewpoint, motion speed, clothing, etc. We present a new framework for dynamic shape analysis based on temporal normalization and factorized shape style analysis. Using a nonlinear generative model with motion manifold embedding in a low-dimensional space, we detect cycles of periodic motion like gait in different views and synthesize temporally-aligned shape sequences from the same type of motion at different speeds. The bilinear analysis of temporally-aligned shape sequences decomposes dynamic motion into time-invariant shape style factors and time-dependent motion factors. We extend the bilinear model into a tensor shape model, a multilinear decomposition of dynamic shape sequences for view-invariant shape style representations. The shape style is a view-invariant, time-invariant, and speed-invariant shape signature and is used as a feature vector for human identification. The shape style can be adapted to new environmental conditions by iterative estimation of style and content factors to reflect new observation conditions. We present the experimental results of gait recognition using the CMU Mobo gait database and the USF gait challenging database. Chan-Su Lee, Ahmed M. Elgammal |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2009 | Contour Segment Matching by Integrating Intra and Inter Shape Cues of ObjectsabstractIn this paper we propose an algorithm for contour-based object detection in cluttered images. Contour of an object shape is approximated as a set of line segments and object detection is framed as matching contour segments of an image (i.e.,an edge image) to a boundary model of an object (i.e., a line drawing). Local shape is abstracted as a group of k-adjacent segments. We use a multi-level shape description (with different k’s) to capture complexity variations in local shape. Between images, shape descriptors are matched to give inter-shape correspondences and within images the underlying segment grouping enforces intra-shape contextual constraints. We use an efficient relaxation labeling approach that integrates these shape cues to qualify a contour match. To this end, we propose a novel framework that solves the problem of object detection as a contour segments correspondence problem. We then demonstrate the efficacy of the method for detecting various objects in cluttered images by comparing them to simple line drawings. Ishani Chakraborty, Ahmed M. Elgammal |
BMVC | 2 |
| 2009 | Dynamic shape outlier detection for human locomotion
Chan-Su Lee, Ahmed M. Elgammal |
Comput. Vis. Image Underst. | 2 |
| 2009 | Introduction to computer vision and image understanding the special issue on video analysis
Qingshan Liu 0001, Xuelong Li 0001, Ahmed M. Elgammal, Xian-Sheng Hua 0001, Dong Xu 0001, Dacheng Tao |
Comput. Vis. Image Underst. | 3 |
| 2009 | Tracking People on a TorusabstractWe present a framework for monocular 3D kinematic pose tracking and viewpoint estimation of periodic and quasi-periodic human motions from an uncalibrated camera. The approach we introduce here is based on learning both the visual observation manifold and the kinematic manifold of the motion using a joint representation. We show that the visual manifold of the observed shape of a human performing a periodic motion, observed from different viewpoints, is topologically equivalent to a torus manifold. The approach we introduce here is based on the supervised learning of both the visual and kinematic manifolds. Instead of learning an embedding of the manifold, we learn the geometric deformation between an ideal manifold (conceptual equivalent topological structure) and a twisted version of the manifold (the data). Experimental results show accurate estimation of the 3D body posture and the viewpoint from a single uncalibrated camera. Ahmed M. Elgammal, Chan-Su Lee |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2008 | Information Theoretic Key Frame Selection for Action RecognitionabstractThis paper presents an approach for human action recognition by finding the discriminative key frames from a video sequence and representing them with the distribution of local motion features and their spatiotemporal arrangements. In this approach, the key frames of the video sequence are selected by their discriminative power and represented by the local motion features detected in them and integrated from their temporal neighbors. In the key frame’s representation, the spatial arrangements of the motion features are captured in a hierarchical spatial pyramid structure. By using frame by frame voting for the recognition, experiments have demonstrated improved performances over most of the other known methods on the popular benchmark data sets. 1 Ahmed M. Elgammal |
BMVC | 2 |
| 2008 | Boosting adaptive linear weak classifiers for online learning and trackingabstractOnline boosting methods have recently been used successfully for tracking, background subtraction etc. Conventional online boosting algorithms emphasize on interchanging new weak classifiers/features to adapt with the change over time. We are proposing a new online boosting algorithm where the form of the weak classifiers themselves are modified to cope with scene changes. Instead of replacement, the parameters of the weak classifiers are altered in accordance with the new data subset presented to the online boosting process at each time step. Thus we may avoid altogether the issue of how many weak classifiers to be replaced to capture the change in the data or which efficient search algorithm to use for a fast retrieval of weak classifiers. A computationally efficient method has been used in this paper for the adaptation of linear weak classifiers. The proposed algorithm has been implemented to be used both as an online learning and a tracking method. We show quantitative and qualitative results on both UCI datasets and several video sequences to demonstrate improved performance of our algorithm. Toufiq Parag, Fatih Porikli, Ahmed M. Elgammal |
CVPR | 3 |
| 2008 | Spatiotemporal pyramid representation for recognition of facial expressions and hand gesturesabstractThis paper presents a spatiotemporal pyramid representation for recognizing facial expressions and hand gestures. This approach works by partitioning video sequence into increasingly fine subdivisions in the space and time domains and modeling the distribution of the local motion features inside each subdivision such that the set of motion features are mapped into spatial and temporal multi-resolution histograms. This spatiotemporal pyramid is built by weighting the histograms from the different layers of the subdivisions. The proposed approach is an extension of the orderless “bag-of-words” model by approximately capturing geometric and temporal arrangements of the local motion features. The experiments on facial expression and hand gesture data sets have demonstrated the significantly improved performance over state of art results on human activity recognition tasks by using our representation. Ahmed M. Elgammal |
FG | 2 |
| 2008 | Object localization using affine invariant substructure constraintsabstractIn this paper we propose a novel method for generic object localization. The method is based on modeling the object as a graph at two levels: a local substructural representation and a global object graph. In the first level, an object substructure is a quasi affine-invariant canonical encoding of a set of four straight contour lines of the object. The second level is a connectivity graph of these substructures that defines the object. The candidate substructures in an observed image are selected probabilistically using the model distribution. To extract the object graph from these candidates, we exploit the strong inter-structural affinities within the object. We consider the connected graph of all candidates and find a bi-partition of this graph. Finally, the partition with higher density (and hence with higher affinity) is selected and labeled as the object structure. This method is independent of affine transformations of objects and robust to intra-class variability and partial occlusion. Ishani Chakraborty, Ahmed M. Elgammal |
ICPR | 2 |
| 2008 | Human activity recognition from frame's spatiotemporal representationabstractThis paper presents an approach for human activity recognition by representing the frames of the video sequence with the distribution of local motion features and their spatiotemporal arrangements. In this approach, the local motion features used for the representation of a frame are integrated from the ones detected in this frame and its temporal neighbors. The features’ spatial arrangements are captured in a hierarchical spatial pyramid structure. By using frame by frame voting for the recognition, experiments have demonstrated improved performances over most of the other known methods on the popular benchmark data sets while approaching the best known results. Ahmed M. Elgammal |
ICPR | 2 |
| 2007 | Modeling View and Posture Manifolds for TrackingabstractIn this paper we consider modeling data lying on multiple continuous manifolds. In particular, we model the shape manifold of a person performing a motion observed from different view points along a view circle at fixed camera height. We introduce a model that ties together the body configuration (kinematics) manifold and the visual manifold (observations) in a way that facilitates tracking the 3D configuration with continuous relative view variability. The model exploits the low dimensionality nature of both the body configuration manifold and the view manifold where each of them are represented separately. Chan-Su Lee, Ahmed M. Elgammal |
ICCV | 2 |
| 2007 | Nonlinear Dynamic Shape and Appearance Models for Facial Motion Tracking
Chan-Su Lee, Ahmed M. Elgammal, Dimitris N. Metaxas |
PSIVT | 2 |
| 2007 | Nonlinear manifold learning for dynamic shape and dynamic appearance
Ahmed M. Elgammal, Chan-Su Lee |
Comput. Vis. Image Underst. | 1 |
| 2006 | Unsupervised Learning of Boosted Tree Classifier Using Graph Cuts for Hand Pose RecognitionabstractThis study proposes an unsupervised learning approach for the task of hand pose recognition. Considering the large variation in hand poses, classification using a decision tree seems highly suitable for this purpose. Various research works have used boosted decision trees and have shown encouraging results for pose recognition. This work also employs a boosted classifier tree learned in an unsupervised manner for hand pose recognition. We use a recursive two way spectral clustering method, namely the Normalized Cut method (NCut), to generate the decision tree. A binary boosting classifier is then learned at each node of the tree generated by the clustering algorithm. Since the output of the clustering algorithm may contain outliers in practice, the variant of boosting algorithm applied at each node is the Soft Margin version of AdaBoost, which was developed to maximize the classifier margin in a noisy environment. We propose a novel approach to learn the weak classifiers of the boosting process using the partitioning vector given by the NCut algorithm. The algorithm applies a linear regression of feature responses with the partitioning vector and utilizes the sample weights used in boosting to learn the weak hypotheses. Initial result shows satisfactory performances in recognizing complex hand poses with large variations in background and illumination. This framework of tree classifier can also be applied to general multi-class object recognition. 1 Toufiq Parag, Ahmed M. Elgammal |
BMVC | 2 |
| 2006 | A Framework for Feature Selection for Background SubtractionabstractBackground subtraction is a widely used paradigm to detect moving objects in video taken from a static camera and is used for various important applications such as video surveillance, human motion analysis, etc. Various statistical approaches have been proposed for modeling a given scene background. However, there is no theoretical framework for choosing which features to use to model different regions of the scene background. In this paper we introduce a novel framework for feature selection for background modeling and subtraction. A boosting algorithm, namely RealBoost, is used to choose the best combination of features at each pixel. Given the probability estimates from a pool of features calculated by Kernel Density Estimate (KDE) over a certain time period, the algorithm selects the most useful ones to discriminate foreground objects from the scene background. The results show that the proposed framework successfully selects appropriate features for different parts of the image. Toufiq Parag, Ahmed M. Elgammal, Anurag Mittal |
CVPR (2) | 2 |
| 2006 | Learning to Detect Objects of Many Classes Using Binary Classifiers
Ramana Isukapalli, Ahmed M. Elgammal, Russell Greiner |
ECCV (1) | 2 |
| 2006 | Synthesis and Control of High Resolution Facial Expressions for Visual InteractionsabstractThe synthesis of facial expression with control of intensity and personal styles is important in intelligent and affective human-computer interaction, especially in face-to-face interaction between human and intelligent agent. We present a facial expression animation system that facilitates control of expressiveness and style. We learn a decomposable generative model for the nonlinear deformation of facial expressions by analyzing the mapping space between low dimensional embedded representation and high resolution tracking data. Bilinear analysis of the mapping space provides a compact representation of the nonlinear generative model for facial expressions. The decomposition allows synthesis of new facial expressions by control of geometry and expression style. The generative model provides control of expressiveness preserving nonlinear deformation in the expressions with simple parameters and allows synthesis of stylized facial geometry. In addition, we can directly extract the MPEG-4 facial animation parameters (FAPs) from the synthesized data, which allows using any animation engine that supports FAPs to animate new synthesized expressions Chan-Su Lee, Ahmed M. Elgammal, Dimitris N. Metaxas |
ICME | 2 |
| 2006 | Edge affinity for pose-contour matching
V. Shiv Naga Prasad, Larry Davis 0001, Son Dinh Tran, Ahmed M. Elgammal |
Comput. Vis. Image Underst. | 4 |
| 2005 | Style Adaptive Bayesian Tracking Using Explicit Manifold LearningabstractCharacteristics of the 2D contour shape deformation in human motion con-tain rich information and can be useful for human identification, gender classi-fication, 3D pose reconstruction and so on. In this paper we introduce a new approach for contour tracking for human motion using an explicit modeling of the motion manifold and learning a decomposable generative model. We use nonlinear dimensionality reduction to embed the motion manifold in a low di-mensional configuration space utilizing the constraints imposed by the human motion. Given such embedding, we learn an explicit representation of the mani-fold, which reduces the problem to a one-dimensional tracking problem and also facilitates linear dynamics on the manifold. We also utilize a generative model through learning a nonlinear mapping between the embedding space and the vi-sual input space, which facilitates capturing global deformation characteristics. The contour tracking problem is formulated as states estimation in the decom-posed generative model parameter within a Bayesian tracking framework. The result is a robust, adaptive gait tracking with shape style estimation. 1 Chan-Su Lee, Ahmed M. Elgammal |
BMVC | 2 |
| 2005 | Learning to Track: Conceptual Manifold Map for Closed-Form TrackingabstractOur objective is to model the visual manifold of object appearance corresponding to geometric transformation. We learn a generative model for object appearance where the appearance of the object at each new frame is a function that maps from a conceptual representation of the geometric transformation space into the visual manifold. By learning such generative model we can infer the geometric transformation (track) directly from the tracked object appearance. As a result tracking can be achieved in a closed form and therefore can be done very efficiently. Ahmed M. Elgammal |
CVPR (1) | 1 |
| 2004 | Separating Style and Content on a Nonlinear Manifold
Ahmed M. Elgammal, Chan-Su Lee |
CVPR (1) | 1 |
| 2004 | Inferring 3D Body Pose from Silhouettes Using Activity Manifold Learning
Ahmed M. Elgammal, Chan-Su Lee |
CVPR (2) | 1 |
| 2004 | High Resolution Acquisition, Learning and Transfer of Dynamic 3D Facial ExpressionsabstractAbstract Synthesis and re‐targeting of facial expressions is central to facial animation and often involves significant manual work in order to achieve realistic expressions, due to the difficulty of capturing high quality dynamic expression data. In this paper we address fundamental issues regarding the use of high quality dense 3‐D data samples undergoing motions at video speeds, e.g. human facial expressions. In order to utilize such data for motion analysis and re‐targeting, correspondences must be established between data in different frames of the same faces as well as between different faces. We present a data driven approach that consists of four parts: 1) High speed, high accuracy capture of moving faces without the use of markers, 2) Very precise tracking of facial motion using a multi‐resolution deformable mesh, 3) A unified low dimensional mapping of dynamic facial motion that can separate expression style, and 4) Synthesis of novel expressions as a combination of expression styles. The accuracy and resolution of our method allows us to capture and track subtle expression details. The low dimensional representation of motion data in a unified embedding for all the subjects in the database allows for learning the most discriminating characteristics of each individual's expressions as that person's “expression style”. Thus new expressions can be synthesized, either as dynamic morphing between individuals, or as expression transfer from a source face to a target face, as demonstrated in a series of experiments. Categories and Subject Descriptors (according to ACM CCS): I.3.7 [Computer Graphics]: Animation; I.3.5 [Computer Graphics]: Curve, surface, solid, and object representations; I.3.3 [Computer Graphics]: Digitizing and scanning; I.2.10 [Artificial intelligence]: Motion ; I.2.10 [Artificial intelligence]: Representations, data structures, and transforms; I.2.10 [Artificial intelligence]: Shape; I.2.6 [Artificial intelligence]: Concept learning Yang Wang 0001, Sharon X. Huang, Chan-Su Lee, Song Zhang 0002, Dimitris Samaras, Dimitris N. Metaxas, Ahmed M. Elgammal, Peisen Huang |
Comput. Graph. Forum | 8 |
| 2003 | A Scalable Image-Based Multi-Camera Visual Surveillance SystemabstractWe describe the design of a scalable and wide coverage visual surveillance system. Scalability (the ability to add and remove cameras easily during system operation with minimal overhead and system degradation) is achieved by utilizing only image-based information for camera control. We show that when a pan-tilt-zoom camera pans and tilts, a given image point moves in a circular and a linear trajectory, respectively. We create a scene model using a plan view of the scene. The scene model makes it easy for us to handle occlusion prediction and schedule video acquisition tasks subject to visibility constraints. We describe a maximum weight matching algorithm to assign cameras to tasks that meet the visibility constraints. The system is illustrated both through simulations and real video from a 6-camera configuration. Ser-Nam Lim, Larry Davis 0001, Ahmed M. Elgammal |
AVSS | 3 |
| 2003 | Probabilistic Tracking in Joint Feature-Spatial SpacesabstractIn this paper, we present a probabilistic framework for tracking regions based on their appearance. We exploit the feature-spatial distribution of a region representing an object as a probabilistic constraint to track that region over time. The tracking is achieved by maximizing a similarity-based objective function over transformation space given a nonparametric representation of the joint feature-spatial distribution. Such a representation imposes a probabilistic constraint on the region feature distribution coupled with the region structure, which yields an appearance tracker that is robust to small local deformations and partial occlusion. We present the approach for the general form of joint feature-spatial distributions and apply it to tracking with different types of image features including row intensity, color and image gradient. Ahmed M. Elgammal, Ramani Duraiswami, Larry Davis 0001 |
CVPR (1) | 1 |
| 2003 | Learning Dynamics for Exemplar-based Gesture Recognition abstractThis paper addresses the problem of capturing the dynamics for exemplar-based recognition systems. Traditional HMM provides a probabilistic tool to capture system dynamics and in exemplar paradigm, HMM states are typically coupled with the exemplars. Alternatively, we propose a non-parametric HMM approach that uses a discrete HMM with arbitrary states (decoupled from exemplars) to capture the dynamics over a large exemplar space where a nonparametric estimation approach is used to model the exemplar distribution. This reduces the need for lengthy and non-optimal training of the HMM observation model. We used the proposed approach for view-based recognition of gestures. The approach is based on representing each gesture as a sequence of learned body poses (exemplars). The gestures are recognized through a probabilistic framework for matching these body poses and for imposing temporal constraints between different poses using the proposed non-parametric HMM. Ahmed M. Elgammal, Vinay D. Shet, Yaser Yacoob, Larry Davis 0001 |
CVPR (1) | 1 |
| 2003 | Image-based pan-tilt camera control in a multi-camera surveillance environmentabstractIn automated surveillance systems with multiple cameras, the system must be able to position the cameras accurately. Each camera must be able to pan-tilt such that an object detected in the scene is in a vantage position in the camera's image plane and subsequently capture images of that object. Typically, camera calibration is required. We propose an approach that uses only image-based information. Each camera is assigned a pan-tilt zero-position. Position of an object detected in one camera is related to the other cameras by homographies between the zero-positions while different pan-tilt positions of the same camera are related in the form of projective rotations. We then derive that the trajectories in the image plane corresponding to these projective rotations are approximately circular for pan and linear for tilt. The camera control technique is subsequently tested in a working prototype. Ser-Nam Lim, Ahmed M. Elgammal, Larry Davis 0001 |
ICME | 2 |
| 2003 | Efficient Kernel Density Estimation Using the Fast Gauss Transform with Applications to Color Modeling and TrackingabstractMany vision algorithms depend on the estimation of a probability density function from observations. Kernel density estimation techniques are quite general and powerful methods for this problem, but have a significant disadvantage in that they are computationally intensive. In this paper, we explore the use of kernel density estimation with the fast Gauss transform (FGT) for problems in vision. The FGT allows the summation of a mixture of ill Gaussians at N evaluation points in O(M+N) time, as opposed to O(MN) time for a naive evaluation and can be used to considerably speed up kernel density estimation. We present applications of the technique to problems from image segmentation and tracking and show that the algorithm allows application of advanced statistical techniques to solve practical vision problems in real-time with today's computers. Ahmed M. Elgammal, Ramani Duraiswami, Larry Davis 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2002 | Gesture recognition using a probabilistic framework for pose matchingabstractThis paper presents an approach for view-based recognition of gestures. The approach is based on representing each gesture as a sequence of learned body poses. The gestures are recognized through a probabilistic framework for matching these body poses and for imposing temporal constrains between different poses. Matching individual poses to image data is performed using a probabilistic formulation for edge matching to obtain a likelihood measurement for each individual pose. The paper introduces a weighted matching scheme for edge templates that emphasize discriminating features in the matching. The weighting does not require establishing correspondences between the different pose models. The probabilistic framework also imposes temporal constrains between different pose through a learned Hidden Markov Model (HMM)for each gesture. Ahmed M. Elgammal, Vhay Shet, Yaser Yacoob, Larry Davis 0001 |
ICARCV | 1 |
| 2002 | Background and foreground modeling using nonparametric kernel density estimation for visual surveillanceabstractAutomatic understanding of events happening at a site is the ultimate goal for many visual surveillance systems. Higher level understanding of events requires that certain lower level computer vision tasks be performed. These may include detection of unusual motion, tracking targets, labeling body parts, and understanding the interactions between people. To achieve many of these tasks, it is necessary to build representations of the appearance of objects in the scene. This paper focuses on two issues related to this problem. First, we construct a statistical representation of the scene background that supports sensitive detection of moving objects in the scene, but is robust to clutter arising out of natural scene variations. Second, we build statistical representations of the foreground regions (moving objects) that support their tracking and support occlusion reasoning. The probability density functions (pdfs) associated with the background and foreground are likely to vary from image to image and will not in general have a known parametric form. We accordingly utilize general nonparametric kernel density estimation techniques for building these statistical representations of the background and the foreground. These techniques estimate the pdf directly from the data without any assumptions about the underlying distributions. Example results from applications are presented. Ahmed M. Elgammal, Ramani Duraiswami, David Harwood, Larry Davis 0001 |
Proc. IEEE | 1 |
| 2001 | Efficient Non-Parametric Adaptive Color Modeling Using Fast Gauss TransformabstractModeling the color distribution of a homogeneous region is used extensively for object tracking and recognition applications. The color distribution of an object represents a feature that is robust to partial occlusion, scaling and object deformation. A variety of parametric and non-parametric statistical techniques have been used to model color distributions. In this paper we present a non-parametric color modeling approach based on kernel density estimation as well as a computational framework for efficient density estimation. Theoretically, our approach is general since kernel density estimators can converge to any density shape with sufficient samples. Therefore, this approach is suitable to model the color distribution of regions with patterns and mixture of colors. Since kernel density estimation techniques are computationally expensive, the paper introduces the use of the fast Gauss transform for efficient computation of the color densities. We show that this approach can be used successfully for color-based segmentation of body parts as well as segmentation of many people under occlusion. Ahmed M. Elgammal, Ramani Duraiswami, Larry Davis 0001 |
CVPR (2) | 1 |
| 2001 | Probabilistic Framework for Segmenting People Under OcclusionabstractIn this paper we address the problem of segmenting foreground regions corresponding to a group of people given models of their appearance that were initialized before occlusion. We present a general framework that uses maximum likelihood estimation to estimate the best arrangement for people in terms of 2D translation that yields a segmentation for the foreground region. Given the segmentation result we conduct occlusion reasoning to recover relative depth information and we show how to utilize this depth information in the same segmentation framework. We also present a more practical solution for the segmentation problem that is online to avoid searching an exponential space of hypothesis. The person model is based on segmenting the body into regions in order to spatially localize the color-features corresponding to the way people are dressed. Modeling these regions involves modeling their appearance (color distributions) as well us their spatial distribution with respect to the body. We use a non-parametric approach bused on kernel density estimation to represent the color distribution of each region and therefore we do not restrict the clothing to be of uniform color instead it can be any mixture of colors and/or patterns. We also present a method to automatically initialize these models and learn them before the occlusion. Ahmed M. Elgammal, Larry Davis 0001 |
ICCV | 1 |
| 2001 | A Graph-Based Segmentation and Feature-Extraction Framework for Arabic Text RecognitionabstractThis paper presents a graph-based framework for the segmentation of Arabic text. The same framework is used to extract font independent structural features from the text that are used in the recognition. The major contribution of this paper is a new graph-based structural segmentation approach based on the topological relation between the baseline and the line adjacency graph representation of the text. The text is segmented to sub-character units that we call "scripts". A structure analysis approach is used for recognition of these units. A different classifier is used to recognize dots and diacritic signs. The final character recognition is achieved by using a regular grammar that describes how characters are composed from scripts. Ahmed M. Elgammal, Mohamed A. Ismail |
ICDAR | 1 |
| 2001 | Techniques for Language Identification for Hybrid Arabic-English Document ImagesabstractBecause of the different characteristics of Arabic language and Romance and Anglo Saxon languages, recognition of documents written in hybrids of these languages requires that the language of the text is to be identified prior to the recognition phase. In this paper, three efficient techniques that can be used to discriminate between text written in Arabic script and text written in English script are presented and evaluated. These techniques address the language identification problem on the word level and on text level. The characteristics of horizontal projection profiles as well as runlength histograms for text written in both languages are the basic features underlying these techniques. Solving this problem is very important in building bilingual document image analysis systems which are capable of processing documents containing hybrid Arabic/Romance and Anglo Saxon languages. Ahmed M. Elgammal, Mohamed A. Ismail |
ICDAR | 1 |
| 2000 | Non-parametric Model for Background Subtraction
Ahmed M. Elgammal, David Harwood, Larry Davis 0001 |
ECCV (2) | 1 |
| 1999 | Face Detection in Complex Environments from Color ImagesabstractThe detection of faces in color images is important for many multimedia applications. It is the first step for face recognition and it can be used for classifying specific shots such as anchorperson and talk show shots. In this paper, we present an algorithm for the detection of faces in color images. The algorithm works by first detecting areas of skin color, then it applies a top-down and a bottom-up analysis to the skin colored areas. The algorithm has been tested on images from news clips and other television programs. The results show that the algorithm is robust and works even for cases where there are objects in the background that have colors similar to the skin. Mohamed Abdel-Mottaleb, Ahmed M. Elgammal |
ICIP (3) | 2 |