Mohamed R. Amer

dblp:87/10771 · also Mohamed Rabie Amer · DBLP profile ↗
← Back
21ranked-venue papers
12as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 8 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 9 first-authorSecurity and privacy · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
11 papers
Video understanding and tracking · 38% Graph learning · 17% Trustworthy machine learning · 15%
Human-computer interaction and pervasive computing
2 papers
Human-AI interaction · 62% Haptics and multimodal interaction · 38%

Topics — the 21 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
activity recognition
0.742016
Sum Product Networks for Activity Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Monte Carlo Tree Search for Scheduling Activity Recognition · ICCV 2013
Cost-Sensitive Top-Down/Bottom-Up Inference for Multiscale Activity Recognition · ECCV (4) 2012
Machine learning › Trustworthy machine learning › interpretability
attention analysis
0.412019
Understanding Attention and Generalization in Graph Neural Networks · NeurIPS 2019
Machine learning › Graph learning › graph neural network
graph attention
0.412019
Understanding Attention and Generalization in Graph Neural Networks · NeurIPS 2019
Machine learning › Graph learning
graph neural network
0.412019
Understanding Attention and Generalization in Graph Neural Networks · NeurIPS 2019
Machine learning › Trustworthy machine learning
interpretability
0.412019
Understanding Attention and Generalization in Graph Neural Networks · NeurIPS 2019
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models
0.322014
HiRF: Hierarchical Random Field for Collective Activity Recognition in Videos · ECCV (6) 2014
Cost-Sensitive Top-Down/Bottom-Up Inference for Multiscale Activity Recognition · ECCV (4) 2012
Computer vision › Vision and language
multimodal fusion
0.312018
Deep Multimodal Fusion: A Hybrid Approach · Int. J. Comput. Vis. 2018
Human-AI interaction
conversational agents
0.312018
Aesop: A Visual Storytelling Platform for Conversational AI · IJCAI 2018
Haptics and multimodal interaction
multimodal interaction
0.312018
Aesop: A Visual Storytelling Platform for Conversational AI · IJCAI 2018
Computer vision › Video understanding and tracking › activity recognition
group activity recognition
0.322014
HiRF: Hierarchical Random Field for Collective Activity Recognition in Videos · ECCV (6) 2014
A chains model for localizing participants of group activities in videos · ICCV 2011
Machine learning › Probabilistic and Bayesian machine learning › tractable probabilistic model
sum-product networks
0.322016
Sum Product Networks for Activity Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Sum-product networks for modeling activities with stochastic structure · CVPR 2012
Computer vision › Video understanding and tracking › action detection
temporal action localization
0.322013
Monte Carlo Tree Search for Scheduling Activity Recognition · ICCV 2013
A chains model for localizing participants of group activities in videos · ICCV 2011
Computer vision › Video understanding and tracking › activity recognition
human activity localization
0.212016
Sum Product Networks for Activity Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Machine learning › Efficient and distributed learning › inference efficiency
cost-aware inference
0.112012
Cost-Sensitive Top-Down/Bottom-Up Inference for Multiscale Activity Recognition · ECCV (4) 2012
Computer vision › Video understanding and tracking › multi-object tracking
data association
0.112011
Multiobject tracking as maximum weight independent set · CVPR 2011
Computer vision › Video understanding and tracking
multi-object tracking
0.112011
Multiobject tracking as maximum weight independent set · CVPR 2011
Computer vision › Video understanding and tracking › object tracking
occlusion handling
0.112011
Multiobject tracking as maximum weight independent set · CVPR 2011
Machine learning › Graph learning
graph classification
0.112019
Understanding Attention and Generalization in Graph Neural Networks · NeurIPS 2019
Machine learning › Deep learning architectures and training
hybrid deep learning
0.112018
Deep Multimodal Fusion: A Hybrid Approach · Int. J. Comput. Vis. 2018
Knowledge, reasoning and agents › Knowledge representation and reasoning › temporal reasoning
spatio-temporal reasoning
0.112018
Aesop: A Visual Storytelling Platform for Conversational AI · IJCAI 2018
Computer vision › 3D vision › depth estimation
monocular depth estimation
0.112015
Monocular Extraction of 2.1D Sketch Using Constrained Convex Optimization · Int. J. Comput. Vis. 2015

Methods — techniques the papers use, named apart from their topics

language parsing · 1.0gesture recognition · 1.0dialogue state management · 1.0weakly-supervised training · 0.4graph neural network · 0.4attention pooling · 0.4multimodal fusion · 0.3deep learning · 0.3sum-product networks · 0.3counting grid model · 0.2
YearPublicationVenuePosition
2021 Learn, Generate, Rank, Explain: A Case Study of Visual Explanation by Generative Machine Learning
abstract
While the computer vision problem of searching for activities in videos is usually addressed by using discriminative models, their decisions tend to be opaque and difficult for people to understand. We propose a case study of a novel machine learning approach for generative searching and ranking of motion capture activities with visual explanation. Instead of directly ranking videos in the database given a text query, our approach uses a variant of Generative Adversarial Networks (GANs) to generate exemplars based on the query and uses them to search for the activity of interest in a large database. Our model is able to achieve comparable results to its discriminative counterpart, while being able to dynamically generate visual explanations. In addition to our searching and ranking method, we present an explanation interface that enables the user to successfully explore the model’s explanations and its confidence by revealing query-based, model-generated motion capture clips that contributed to the model’s decision. Finally, we conducted a user study with 44 participants to show that by using our model and interface, participants benefit from a deeper understanding of the model’s conceptualization of the search query. We discovered that the XAI system yielded a comparable level of efficiency, accuracy, and user-machine synchronization as its black-box counterpart, if the user exhibited a high level of trust for AI explanation.
Chris Kim, Graham W. Taylor, Mohamed R. Amer
ACM Trans. Interact. Intell. Syst.5
2020 Chameleons' Oblivion: Complex-Valued Deep Neural Networks for Protocol-Agnostic RF Device Fingerprinting
abstract
Prior work has demonstrated techniques for fingerprinting devices based on their network traffic or transmitted signals, which use software artifacts or characteristics of the underlying protocol. However these approaches are not robust or applicable in many real-world scenarios. In this paper we explore the feasibility of device fingerprinting under challenging realistic settings, by identifying artifacts in the transmitted signals caused by devices' unique hardware “imperfections”. We develop RF-DCN, a novel Deep Complex-valued Neural Network (DCN) that operates on raw RF signals and is completely agnostic of the underlying applications and protocols. We introduce two DCN variations: a retrofitted Convolutional DCN (CDCN) originally created for acoustic signals, and a novel Recurrent DCN (RDCN) for modeling time series. Our work demonstrates the feasibility of operating on raw I/Q data collected within a narrowband spectrum from open air captures across vastly different modulation schemes. In contrast to prior work, we do not utilize knowledge of the modulation scheme or protocol intricacies such as carrier frequencies. We conduct an extensive experimental evaluation on large and diverse datasets as part of a DARPA red team evaluation, and investigate the effects of different environmental factors as well as neural network architectures and hyperparameters on our system's performance. Our novel RDCN consistently outperforms all baseline neural network architectures, is robust to noise, and can identify a target device even when numerous devices are concurrently transmitting within the band of interest under the same or different protocols. While our experiments demonstrate the applicability of our techniques under challenging conditions where other neural network architectures break down, we identify additional challenges in signal-based fingerprinting and provide guidelines for future explorations.
Ioannis Agadakos, Nikolaos Agadakos, Iasonas Polakis, Mohamed R. Amer
EuroS&P4
2019 Image Classification with Hierarchical Multigraph Networks
Boris Knyazev 0001, Mohamed R. Amer, Graham W. Taylor
BMVC3
2019 Understanding Attention and Generalization in Graph Neural Networks
abstract
We aim to better understand attention over nodes in graph neural networks (GNNs) and identify factors influencing its effectiveness. We particularly focus on the ability of attention GNNs to generalize to larger, more complex or noisy graphs. Motivated by insights from the work on Graph Isomorphism Networks, we design simple graph reasoning tasks that allow us to study attention in a controlled environment. We find that under typical conditions the effect of attention is negligible or even harmful, but under certain conditions it provides an exceptional gain in performance of more than 60% in some of our classification tasks. Satisfying these conditions in practice is challenging and often requires optimal initialization or supervised training of attention. We propose an alternative recipe and train attention in a weakly-supervised fashion that approaches the performance of supervised models, and, compared to unsupervised models, improves results on several synthetic as well as real datasets. Source code and datasets are available at https://github.com/bknyaz/graphattentionpool.
Boris Knyazev 0001, Graham W. Taylor, Mohamed R. Amer
NeurIPS3
2018 Aesop: A Visual Storytelling Platform for Conversational AI
abstract
We present a new collaborative visual storytelling platform, Aesop, for direction and animation. Aesop consists of a language parser, human gesture monitoring, composition graphs, dialogue state manager, and an interactive 3D animation software. Aesop thus enables 3D spatial and temporal reasoning which are both essential for storytelling. Our key innovation is to enable conversational AI using both verbal and non-verbal communication, which enables research in language, vision, and planning.
Timothy J. Meo, Aswin Raghavan, David A. Salter, Alex Tozzo, Amir Tamrakar, Mohamed R. Amer
IJCAI6
2018 Deep Multimodal Fusion: A Hybrid Approach
Mohamed R. Amer, Timothy J. Shields, Behjat Siddiquie, Amir Tamrakar, Ajay Divakaran, Sek M. Chai
Int. J. Comput. Vis.1
2018 Bayesian optimization on graph-structured search spaces: Optimizing deep multimodal fusion architectures
Dhanesh Ramachandram, Michal Lisicki, Timothy J. Shields, Mohamed R. Amer, Graham W. Taylor
Neurocomputing4
2017 Structure optimization for deep multimodal fusion networks using graph-induced kernels
Dhanesh Ramachandram, Michal Lisicki, Timothy J. Shields, Mohamed R. Amer, Graham W. Taylor
ESANN4
2016 Sum Product Networks for Activity Recognition
abstract
This paper addresses detection and localization of human activities in videos. We focus on activities that may have variable spatiotemporal arrangements of parts, and numbers of actors. Such activities are represented by a sum-product network (SPN). A product node in SPN represents a particular arrangement of parts, and a sum node represents alternative arrangements. The sums and products are hierarchically organized, and grounded onto space-time windows covering the video. The windows provide evidence about the activity classes based on the Counting Grid (CG) model of visual words. This evidence is propagated bottom-up and top-down to parse the SPN graph for the explanation of the video. The node connectivity and model parameters of SPN and CG are jointly learned under two settings, weakly supervised, and supervised. For evaluation, we use our new Volleyball dataset, along with the benchmark datasets VIRAT, UT-Interactions, KTH, and TRECVID MED 2011. Our video classification and activity localization are superior to those of the state of the art on these datasets.
Mohamed R. Amer, Sinisa Todorovic
IEEE Trans. Pattern Anal. Mach. Intell.1
2015 The Tower Game Dataset: A multimodal dataset for analyzing social interaction predicates
abstract
We introduce the Tower Game Dataset for computational modeling of social interaction predicates. Existing research in affective computing has focused primarily on recognizing the emotional and mental state of a human based on external behaviors. Recent research in the social science community argues that engaged and sustained social interactions require the participants to jointly coordinate their verbal and non-verbal behaviors. With this as our guiding principle, we collected the Tower Game Dataset consisting of multimodal recordings of two players participating in a tower building game, in the process communicating and collaborating with each other. The format of the game was specifically chosen as it elicits spontaneous communication from the participants through social interaction predicates such as joint attention and entrainment. The dataset will be made public and we believe that it will foster new research in the area of computational social interaction modeling.
David A. Salter, Amir Tamrakar, Behjat Siddiquie, Mohamed R. Amer, Ajay Divakaran, Brian Lande, Darius Mehri
ACII4
2015 2nd Workshop on Computational Models of Social Interactions: Human-Computer-Media Communication (HCMC2015)
abstract
Communicating ideas and information from and to humans is a very important subject. In our daily life, human interact with variety of entities, such as, other humans, machines, media. Constructive interactions are needed for good communication, which would result in successful outcomes, such as answering a query, learning a new skill, getting a service done, and communicating emotions. Each of these entities invokes a set of signals. Current research has focused on analyzing one entity's signals with no respect to the other entities in a unidirectional manner. The computer vision community focused on detection, classification and recognition of humans and their poses and gestures progressing onto actions, activities, and events but it does not go beyond that. The signal processing community focused on emotion recognition from facial expressions or audio or both combined. The HCI community focused on making easier interfaces for machines to ease their usage. The goal of this workshop is to bring multiple disciplines together, to process human directed signals holistically, in a bidirectional manner, rather than isolation. This workshop is positioned to display this rich domain of applications, which will provide the necessary next boost for these technologies. At the same time, it seeks to ground computational models on theory that would help achieve the technology goals. This would allow us to leverage decades of research in different fields and to spur interdisciplinary research thereby opening up new problem domains for the multimedia community.
Mohamed R. Amer, Ajay Divakaran, Shih-Fu Chang, Nicu Sebe
ACM Multimedia1
2015 Monocular Extraction of 2.1D Sketch Using Constrained Convex Optimization
Mohamed R. Amer, Siavash Yousefi, Raviv Raich, Sinisa Todorovic
Int. J. Comput. Vis.1
2014 HiRF: Hierarchical Random Field for Collective Activity Recognition in Videos
Mohamed R. Amer, Sinisa Todorovic
ECCV (6)1
2014 Emotion detection in speech using deep networks
abstract
We propose a novel staged hybrid model for emotion detection in speech. Hybrid models exploit the strength of discriminative classifiers along with the representational power of generative models. Discriminative classifiers have been shown to achieve higher performances than the corresponding generative likelihood-based classifiers. On the other hand, generative models learn a rich informative representations. Our proposed hybrid model consists of a generative model, which is used for unsupervised representation learning of short term temporal phenomena and a discriminative model, which is used for event detection and classification of long range temporal dynamics. We evaluate our approach on multiple audio-visual datasets (AVEC, VAM, and SPD) and demonstrate its superiority compared to the state-of-the-art.
Mohamed R. Amer, Behjat Siddiquie, Colleen Richey, Ajay Divakaran
ICASSP1
2014 Multimodal fusion using dynamic hybrid models
abstract
We propose a novel hybrid model that exploits the strength of discriminative classifiers along with the representational power of generative models. Our focus is on detecting multimodal events in time varying sequences. Discriminative classifiers have been shown to achieve higher performances than the corresponding generative likelihood-based classifiers. On the other hand, generative models learn a rich informative space which allows for data generation and joint feature representation that discriminative models lack. We employ a deep temporal generative model for unsupervised learning of a shared representation across multiple modalities with time varying data. The temporal generative model takes into account short term temporal phenomena and allows for filling in missing data by generating data within or across modalities. The hybrid model involves augmenting the temporal generative model with a temporal discriminative model for event detection, and classification, which enables modeling long range temporal dynamics. We evaluate our approach on audio-visual datasets (AVEC, AVLetters, and CUAVE) and demonstrate its superiority compared to the state-of-the-art.
Mohamed R. Amer, Behjat Siddiquie, Saad M. Khan, Ajay Divakaran, Harpreet Sawhney
WACV1
2013 Monte Carlo Tree Search for Scheduling Activity Recognition
abstract
This paper addresses recognition of human activities with stochastic structure, characterized by variable space-time arrangements of primitive actions, and conducted by a variable number of actors. Our approach classifies the activity of interest as well as identifies the relevant foreground in the video. Each activity representation is considered as a mixture distribution of BoWs captured by a Sum-Product Network (SPN). In our approach, SPN represents a linear mixture of many bags-of-words (BoWs) where each BoW represents an important foreground part of the activity. This mixture distribution is efficiently computed by organizing the BoWs in a hierarchy, where children BoWs are nested within parent BoWs. SPN allows us to model this mixture since it consists of terminal nodes representing BoWs, product nodes, and sum nodes organized in a number of layers. The products are aimed at encoding particular configurations of primitive actions, and the sums serve to capture their alternative configurations. SPN inference amounts to parsing the SPN graph, which yields the most probable explanation (MPE) of the video foreground. SPN inference has linear complexity in the number of nodes, under fairly general conditions, enabling fast and scalable recognition. The connectivity of SPN and the parameters of BoW distributions are learned under weak supervision using a variational EM algorithm. For our evaluation, we have compiled and annotated a new Volleyball dataset. Our classification accuracy and localization results are superior to those of the state of the art on current benchmarks as well as our Volleyball datasets.
Mohamed R. Amer, Sinisa Todorovic, Alan Fern, Song-Chun Zhu
ICCV1
2012 Sum-product networks for modeling activities with stochastic structure
abstract
This paper addresses recognition of human activities with stochastic structure, characterized by variable spacetime arrangements of primitive actions, and conducted by a variable number of actors. We demonstrate that modeling aggregate counts of visual words is surprisingly expressive enough for such a challenging recognition task. An activity is represented by a sum-product network (SPN). SPN is a mixture of bags-of-words (BoWs) with exponentially many mixture components, where subcomponents are reused by larger ones. SPN consists of terminal nodes representing BoWs, and product and sum nodes organized in a number of layers. The products are aimed at encoding particular configurations of primitive actions, and the sums serve to capture their alternative configurations. The connectivity of SPN and parameters of BoW distributions are learned under weak supervision using the EM algorithm. SPN inference amounts to parsing the SPN graph, which yields the most probable explanation (MPE) of the video in terms of activity detection and localization. SPN inference has linear complexity in the number of nodes, under fairly general conditions, enabling fast and scalable recognition. A new Volleyball dataset is compiled and annotated for evaluation. Our classification accuracy and localization precision and recall are superior to those of the state-of-the-art on the benchmark and our Volleyball datasets.
Mohamed R. Amer, Sinisa Todorovic
CVPR1
2012 Cost-Sensitive Top-Down/Bottom-Up Inference for Multiscale Activity Recognition
Mohamed R. Amer, Dan Xie 0005, Mingtian Zhao, Sinisa Todorovic, Song-Chun Zhu
ECCV (4)1
2011 Multiobject tracking as maximum weight independent set
abstract
This paper addresses the problem of simultaneous tracking of multiple targets in a video. We first apply object detectors to every video frame. Pairs of detection responses from every two consecutive frames are then used to build a graph of tracklets. The graph helps transitively link the best matching tracklets that do not violate hard and soft contextual constraints between the resulting tracks. We prove that this data association problem can be formulated as finding the maximum-weight independent set (MWIS) of the graph. We present a new, polynomial-time MWIS algorithm, and prove that it converges to an optimum. Similarity and contextual constraints between object detections, used for data association, are learned online from object appearance and motion properties. Long-term occlusions are addressed by iteratively repeating MWIS to hierarchically merge smaller tracks into longer ones. Our results demonstrate advantages of simultaneously accounting for soft and hard contextual constraints in multitarget tracking. We outperform the state of the art on the benchmark datasets.
William Brendel, Mohamed R. Amer, Sinisa Todorovic
CVPR2
2011 A chains model for localizing participants of group activities in videos
abstract
Given a video, we would like to recognize group activities, localize video parts where these activities occur, and detect actors involved in them. This advances prior work that typically focuses only on video classification. We make a number of contributions. First, we specify a new, mid-level, video feature aimed at summarizing local visual cues into bags of the right detections (BORDs). BORDs seek to identify the right people who participate in a target group activity among many noisy people detections. Second, we formulate a new, generative, chains model of group activities. Inference of the chains model identifies a subset of BORDs in the video that belong to occurrences of the activity, and organizes them in an ensemble of temporal chains. The chains extend over, and thus localize, the time intervals occupied by the activity. We formulate a new MAP inference algorithm that iterates two steps: i) Warps the chains of BORDs in space and time to their expected locations, so the transformed BORDs can better summarize local visual cues; and ii) Maximizes the posterior probability of the chains. We outperform the state of the art on benchmark UT-Human Interaction and Collective Activities datasets, under reasonable running times.
Mohamed R. Amer, Sinisa Todorovic
ICCV1
2010 Monocular Extraction of 2.1D Sketch
Mohamed R. Amer, Raviv Raich, Sinisa Todorovic
ICIP1