Gamaleldin F. Elsayed

dblp:215/4903 · also Gamaleldin Fathy Elsayed · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
6since 2021 · last 2023
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 6 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
11 papers
Trustworthy machine learning · 31% Representation and self-supervised learning · 16% Deep learning architectures and training · 16%
Network and information security
1 paper
Security and privacy of machine learning · 100%

Topics — the 29 heaviest of 33, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
robustness
1.332023
Teacher-generated spatial-attention labels boost robustness and accuracy of contrastive models · CVPR 2023
Adversarial Examples that Fool both Computer Vision and Time-Limited Humans · NeurIPS 2018
Large Margin Deep Networks for Classification · NeurIPS 2018
Computer vision › Video understanding and tracking › video representation learning
object-centric video learning
1.122022
SAVi++: Towards End-to-End Object-Centric Learning from Real-World Videos · NeurIPS 2022
Conditional Object-Centric Learning from Video · ICLR 2022
Machine learning › Representation and self-supervised learning
contrastive learning
0.712023
Teacher-generated spatial-attention labels boost robustness and accuracy of contrastive models · CVPR 2023
Machine learning › Representation and self-supervised learning
equivariance
0.712023
Invariant Slot Attention: Object Discovery with Slot-Centric Reference Frames · ICML 2023
Computer vision › Image recognition and object detection
image classification
0.712023
Teacher-generated spatial-attention labels boost robustness and accuracy of contrastive models · CVPR 2023
Machine learning › Efficient and distributed learning › large-scale learning
large-scale model training
0.712023
Scaling Vision Transformers to 22 Billion Parameters · ICML 2023
Computer vision › Image recognition and object detection
object discovery
0.712023
Invariant Slot Attention: Object Discovery with Slot-Centric Reference Frames · ICML 2023
Machine learning › Trustworthy machine learning › robustness › robust learning
robust classification
0.712023
Teacher-generated spatial-attention labels boost robustness and accuracy of contrastive models · CVPR 2023
Machine learning › Representation and self-supervised learning › representation learning › object-centric representation learning
slot attention
0.712023
Invariant Slot Attention: Object Discovery with Slot-Centric Reference Frames · ICML 2023
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.712023
Scaling Vision Transformers to 22 Billion Parameters · ICML 2023
Machine learning › Deep learning architectures and training › transformer › vision transformer
vision transformer scaling
0.712023
Scaling Vision Transformers to 22 Billion Parameters · ICML 2023
Computer vision › 3D vision
depth estimation
0.612022
SAVi++: Towards End-to-End Object-Centric Learning from Real-World Videos · NeurIPS 2022
Machine learning › Representation and self-supervised learning › representation learning
object-centric representation learning
0.612022
Conditional Object-Centric Learning from Video · ICLR 2022
Computer vision › Segmentation and scene understanding › object segmentation
unsupervised object segmentation
0.612022
SAVi++: Towards End-to-End Object-Centric Learning from Real-World Videos · NeurIPS 2022
Machine learning › Deep learning architectures and training
convolutional neural network
0.412020
Revisiting Spatial Invariance with Low-Rank Local Connectivity · ICML 2020
Machine learning › Learning theory
inductive bias
0.412020
Revisiting Spatial Invariance with Low-Rank Local Connectivity · ICML 2020
Machine learning › Deep learning architectures and training › symmetry-aware learning
spatial invariance
0.412020
Revisiting Spatial Invariance with Low-Rank Local Connectivity · ICML 2020
Machine learning › Trustworthy machine learning › robustness
adversarial attack
0.412019
Adversarial Reprogramming of Neural Networks · ICLR (Poster) 2019
Machine learning › Trustworthy machine learning › robustness › adversarial attack
adversarial reprogramming
0.412019
Adversarial Reprogramming of Neural Networks · ICLR (Poster) 2019
Machine learning › Trustworthy machine learning › interpretability
attention analysis
0.412019
Saccader: Improving Accuracy of Hard Attention Models for Vision · NeurIPS 2019
Machine learning › Trustworthy machine learning
interpretability
0.412019
Saccader: Improving Accuracy of Hard Attention Models for Vision · NeurIPS 2019
Security and privacy of machine learning
adversarial attack
0.412019
Adversarial Reprogramming of Neural Networks · ICLR (Poster) 2019
Machine learning › Trustworthy machine learning › robustness
adversarial examples
0.312018
Adversarial Examples that Fool both Computer Vision and Time-Limited Humans · NeurIPS 2018
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.312018
Large Margin Deep Networks for Classification · NeurIPS 2018
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial transferability
0.312018
Adversarial Examples that Fool both Computer Vision and Time-Limited Humans · NeurIPS 2018
Machine learning › Deep learning architectures and training › loss function design
margin-based loss
0.312018
Large Margin Deep Networks for Classification · NeurIPS 2018
Machine learning › Trustworthy machine learning › fairness › fairness trade-off
fairness-accuracy trade-off
0.212023
Scaling Vision Transformers to 22 Billion Parameters · ICML 2023
Machine learning › Trustworthy machine learning › fairness
fairness and robustness
0.212023
Scaling Vision Transformers to 22 Billion Parameters · ICML 2023
Machine learning › Learning theory
generalization
0.112018
Large Margin Deep Networks for Classification · NeurIPS 2018

Methods — techniques the papers use, named apart from their topics

slot attention · 1.2contrastive learning · 1.2pseudo-labeling · 0.7linear probing on frozen features · 0.7large-scale pretraining · 0.7knowledge distillation · 0.7equivariant attention · 0.7attention prediction · 0.7self-supervised learning · 0.6LiDAR depth signals · 0.6adversarial reprogramming · 0.4
YearPublicationVenuePosition
2023 Teacher-generated spatial-attention labels boost robustness and accuracy of contrastive models
abstract
Human spatial attention conveys information about the regions of visual scenes that are important for performing visual tasks. Prior work has shown that the information about human attention can be leveraged to benefit various supervised vision tasks. Might providing this weak form of supervision be useful for self-supervised representation learning? Addressing this question requires collecting large datasets with human attention labels. Yet, collecting such large scale data is very expensive. To address this challenge, we construct an auxiliary teacher model to predict human attention, trained on a relatively small labeled dataset. This teacher model allows us to generate image (pseudo) attention labels for ImageNet. We then train a model with a primary contrastive objective; to this standard configuration, we add a simple output head trained to predict the attention map for each image, guided by the pseudo labels from teacher model. We measure the quality of learned representations by evaluating classification performance from the frozen learned embeddings as well as performance on image retrieval tasks (see supplementary material). We find that the spatial-attention maps predicted from the contrastive model trained with teacher guidance aligns better with human attention compared to vanilla contrastive models. Moreover, we find that our approach improves classification accuracy and robustness of the contrastive models on ImageNet and ImageNet-C. Further, we find that model representations become more useful for image retrieval task as measured by precision-recall performance on ImageNet, ImageNet-C, CIFAR10, and CIFAR10-C datasets.
Yushi Yao, Chang Ye, Junfeng He, Gamaleldin F. Elsayed
CVPR4
2023 Learning in temporally structured environments
Matt Jones 0001, Tyler R. Scott, Mengye Ren, Gamaleldin F. Elsayed, Katherine L. Hermann, David Mayo, Michael C. Mozer
ICLR4
2023 Scaling Vision Transformers to 22 Billion Parameters
abstract
The scaling of Transformers has driven breakthrough capabilities for language models. At present, the largest large language models (LLMs) contain upwards of 100B parameters. Vision Transformers (ViT) have introduced the same architecture to image and video modelling, but these have not yet been successfully scaled to nearly the same degree; the largest dense ViT contains 4B parameters (Chen et al., 2022). We present a recipe for highly efficient and stable training of a 22B-parameter ViT (ViT-22B) and perform a wide variety of experiments on the resulting model. When evaluated on downstream tasks (often with a lightweight linear model on frozen features), ViT-22B demonstrates increasing performance with scale. We further observe other interesting benefits of scale, including an improved tradeoff between fairness and performance, state-of-the-art alignment to human visual perception in terms of shape/texture bias, and improved robustness. ViT-22B demonstrates the potential for "LLM-like" scaling in vision, and provides key steps towards getting there.
Mostafa Dehghani 0001, Josip Djolonga, Basil Mustafa, Piotr Padlewski, Jonathan Heek, Justin Gilmer, Andreas Steiner 0001, Mathilde Caron, Robert Geirhos, Ibrahim Alabdulmohsin, Rodolphe Jenatton, Lucas Beyer, Michael Tschannen, Anurag Arnab, Xiao Wang 0038, Carlos Riquelme, Matthias Minderer, Joan Puigcerver, Utku Evci, Sjoerd van Steenkiste, Gamaleldin F. Elsayed, Aravindh Mahendran, Fisher Yu 0001, Avital Oliver, Fantine Huot, Jasmijn Bastings, Mark Collier, Alexey A. Gritsenko, Vighnesh Birodkar, Cristina Nader Vasconcelos, Yi Tay, Thomas Mensink, Alexander Kolesnikov 0003, Filip Pavetic, Dustin Tran, Thomas Kipf, Mario Lucic, Xiaohua Zhai, Daniel Keysers, Jeremiah J. Harmsen, Neil Houlsby
ICML22
2023 Invariant Slot Attention: Object Discovery with Slot-Centric Reference Frames
abstract
Automatically discovering composable abstractions from raw perceptual data is a long-standing challenge in machine learning. Recent slot-based neural networks that learn about objects in a self-supervised manner have made exciting progress in this direction. However, they typically fall short at adequately capturing spatial symmetries present in the visual world, which leads to sample inefficiency, such as when entangling object appearance and pose. In this paper, we present a simple yet highly effective method for incorporating spatial symmetries via slot-centric reference frames. We incorporate equivariance to per-object pose transformations into the attention and generation mechanism of Slot Attention by translating, scaling, and rotating position encodings. These changes result in little computational overhead, are easy to implement, and can result in large gains in terms of data efficiency and overall improvements to object discovery. We evaluate our method on a wide range of synthetic object discovery benchmarks namely CLEVR, Tetrominoes, CLEVRTex, Objects Room and MultiShapeNet, and show promising improvements on the challenging real-world Waymo Open dataset.
Ondrej Biza, Sjoerd van Steenkiste, Mehdi S. M. Sajjadi, Gamaleldin F. Elsayed, Aravindh Mahendran, Thomas Kipf
ICML4
2022 Conditional Object-Centric Learning from Video
Thomas Kipf, Gamaleldin F. Elsayed, Aravindh Mahendran, Austin Stone, Sara Sabour, Georg Heigold, Rico Jonschkowski, Alexey Dosovitskiy, Klaus Greff
ICLR2
2022 SAVi++: Towards End-to-End Object-Centric Learning from Real-World Videos
abstract
The visual world can be parsimoniously characterized in terms of distinct entities with sparse interactions. Discovering this compositional structure in dynamic visual scenes has proven challenging for end-to-end computer vision approaches unless explicit instance-level supervision is provided. Slot-based models leveraging motion cues have recently shown great promise in learning to represent, segment, and track objects without direct supervision, but they still fail to scale to complex real-world multi-object videos. In an effort to bridge this gap, we take inspiration from human development and hypothesize that information about scene geometry in the form of depth signals can facilitate object-centric learning. We introduce SAVi++, an object-centric video model which is trained to predict depth signals from a slot-based video representation. By further leveraging best practices for model scaling, we are able to train SAVi++ to segment complex dynamic scenes recorded with moving cameras, containing both static and moving objects of diverse appearance on naturalistic backgrounds, without the need for segmentation supervision. Finally, we demonstrate that by using sparse depth signals obtained from LiDAR, SAVi++ is able to learn emergent object segmentation and tracking from videos in the real-world Waymo Open dataset.
Gamaleldin F. Elsayed, Aravindh Mahendran, Sjoerd van Steenkiste, Klaus Greff, Michael C. Mozer, Thomas Kipf
NeurIPS1
2020 Revisiting Spatial Invariance with Low-Rank Local Connectivity
abstract
Convolutional neural networks are among the most successful architectures in deep learning with this success at least partially attributable to the efficacy of spatial invariance as an inductive bias. Locally connected layers, which differ from convolutional layers only in their lack of spatial invariance, usually perform poorly in practice. However, these observations still leave open the possibility that some degree of relaxation of spatial invariance may yield a better inductive bias than either convolution or local connectivity. To test this hypothesis, we design a method to relax the spatial invariance of a network layer in a controlled manner; we create a \emph{low-rank} locally connected layer, where the filter bank applied at each position is constructed as a linear combination of basis set of filter banks with spatially varying combining weights. By varying the number of basis filter banks, we can control the degree of relaxation of spatial invariance. In experiments with small convolutional networks, we find that relaxing spatial invariance improves classification accuracy over both convolution and locally connected layers across MNIST, CIFAR-10, and CelebA datasets, thus suggesting that spatial invariance may be an overly restrictive prior.
Gamaleldin F. Elsayed, Prajit Ramachandran, Jonathon Shlens, Simon Kornblith
ICML1
2019 Adversarial Reprogramming of Neural Networks
Gamaleldin F. Elsayed, Ian J. Goodfellow, Jascha Sohl-Dickstein
ICLR (Poster)1
2019 Saccader: Improving Accuracy of Hard Attention Models for Vision
abstract
Although deep convolutional neural networks achieve state-of-the-art performance across nearly all image classification tasks, their decisions are difficult to interpret. One approach that offers some level of interpretability by design is \textit{hard attention}, which uses only relevant portions of the image. However, training hard attention models with only class label supervision is challenging, and hard attention has proved difficult to scale to complex datasets. Here, we propose a novel hard attention model, which we term Saccader. Key to Saccader is a pretraining step that requires only class labels and provides initial attention locations for policy gradient optimization. Our best models narrow the gap to common ImageNet baselines, achieving $75\%$ top-1 and $91\%$ top-5 while attending to less than one-third of the image.
Gamaleldin F. Elsayed, Simon Kornblith, Quoc V. Le
NeurIPS1
2018 Large Margin Deep Networks for Classification
abstract
We present a formulation of deep learning that aims at producing a large margin classifier. The notion of \emc{margin}, minimum distance to a decision boundary, has served as the foundation of several theoretically profound and empirically successful results for both classification and regression tasks. However, most large margin algorithms are applicable only to shallow models with a preset feature representation; and conventional margin methods for neural networks only enforce margin at the output layer. Such methods are therefore not well suited for deep networks. In this work, we propose a novel loss function to impose a margin on any chosen set of layers of a deep network (including input and hidden layers). Our formulation allows choosing any $l_p$ norm ($p \geq 1$) on the metric measuring the margin. We demonstrate that the decision boundary obtained by our loss has nice properties compared to standard classification loss functions. Specifically, we show improved empirical results on the MNIST, CIFAR-10 and ImageNet datasets on multiple tasks: generalization from small training sets, corrupted labels, and robustness against adversarial perturbations. The resulting loss is general and complementary to existing data augmentation (such as random/adversarial input transform) and regularization techniques such as weight decay, dropout, and batch norm. \footnote{Code for the large margin loss function is released at \url{https://github.com/google-research/google-research/tree/master/large_margin}}
Gamaleldin F. Elsayed, Dilip Krishnan, Hossein Mobahi, Kevin Regan 0001, Samy Bengio
NeurIPS1
2018 Adversarial Examples that Fool both Computer Vision and Time-Limited Humans
abstract
Machine learning models are vulnerable to adversarial examples: small changes to images can cause computer vision models to make mistakes such as identifying a school bus as an ostrich. However, it is still an open question whether humans are prone to similar mistakes. Here, we address this question by leveraging recent techniques that transfer adversarial examples from computer vision models with known parameters and architecture to other models with unknown parameters and architecture, and by matching the initial processing of the human visual system. We find that adversarial examples that strongly transfer across computer vision models influence the classifications made by time-limited human observers.
Gamaleldin F. Elsayed, Shreya Shankar, Brian Cheung, Nicolas Papernot, Alexey Kurakin, Ian J. Goodfellow, Jascha Sohl-Dickstein
NeurIPS1