Allan Jabri

dblp:172/0858 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
7since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
15 papers
3D vision · 18% Generative modeling · 18% Reinforcement learning · 16%
Computer graphics and multimedia
2 papers
Visual content generation and editing · 74% Image and video processing · 26%

Topics — the 30 heaviest of 47, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
2.132024
DORSal: Diffusion for Object-centric Representations of Scenes et al · ICLR 2024
Diffusion Self-Guidance for Controllable Image Generation · NeurIPS 2023
Scalable Adaptive Computation for Iterative Generation · ICML 2023
Computer vision › Video understanding and tracking
video object segmentation
1.022022
Learning Pixel Trajectories with Multiscale Contrastive Random Walks · CVPR 2022
Learning Correspondence From the Cycle-Consistency of Time · CVPR 2019
Machine learning › Generative modeling › scene generation
object-centric scene generation
0.812024
DORSal: Diffusion for Object-centric Representations of Scenes et al · ICLR 2024
Computer vision › 3D vision
object representation
0.812024
DORSal: Diffusion for Object-centric Representations of Scenes et al · ICLR 2024
Computer vision › 3D vision › 3d scene modeling
scene representation
0.812024
DORSal: Diffusion for Object-centric Representations of Scenes et al · ICLR 2024
Computer vision › 3D vision › motion estimation
optical flow
0.722022
Learning Pixel Trajectories with Multiscale Contrastive Random Walks · CVPR 2022
Learning Correspondence From the Cycle-Consistency of Time · CVPR 2019
Machine learning › Efficient and distributed learning
adaptive computation
0.712023
Scalable Adaptive Computation for Iterative Generation · ICML 2023
Machine learning › Generative modeling › diffusion model › controllable generation
controllable image generation
0.712023
Diffusion Self-Guidance for Controllable Image Generation · NeurIPS 2023
Machine learning › Reinforcement learning
exploration
0.712023
MIMEx: Intrinsic Rewards from Masked Input Modeling · NeurIPS 2023
Machine learning › Reinforcement learning › exploration
intrinsic motivation
0.712023
MIMEx: Intrinsic Rewards from Masked Input Modeling · NeurIPS 2023
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
masked modeling
0.712023
MIMEx: Intrinsic Rewards from Masked Input Modeling · NeurIPS 2023
Visual content generation and editing
image editing
0.712023
Diffusion Self-Guidance for Controllable Image Generation · NeurIPS 2023
Computer vision › Video understanding and tracking
motion segmentation
0.612022
Discovering Objects that Can Move · CVPR 2022
Computer vision › Image recognition and object detection
object discovery
0.612022
Discovering Objects that Can Move · CVPR 2022
Computer vision › Video understanding and tracking › video analytics › video object analysis › object-centric video understanding
object permanence
0.612022
Object Permanence Emerges in a Random Walk along Memory · ICML 2022
Computer vision › Video understanding and tracking
spatio-temporal correspondence
0.612022
Learning Pixel Trajectories with Multiscale Contrastive Random Walks · CVPR 2022
Computer vision › Image recognition and object detection › object discovery
unsupervised object discovery
0.612022
Discovering Objects that Can Move · CVPR 2022
Computer vision › 3D vision › motion estimation › optical flow
unsupervised optical flow
0.612022
Learning Pixel Trajectories with Multiscale Contrastive Random Walks · CVPR 2022
Robotics › Robot manipulation › assembly › object assembly
block stacking
0.412020
Towards Practical Multi-Object Manipulation using Relational Reinforcement Learning · ICRA 2020
Machine learning › Representation and self-supervised learning
cycle consistency
0.412020
Space-Time Correspondence as a Contrastive Random Walk · NeurIPS 2020
Robotics › Robot manipulation › object manipulation
multi-object manipulation
0.412020
Towards Practical Multi-Object Manipulation using Relational Reinforcement Learning · ICRA 2020
Machine learning › Reinforcement learning
relational reinforcement learning
0.412020
Towards Practical Multi-Object Manipulation using Relational Reinforcement Learning · ICRA 2020
Machine learning › Representation and self-supervised learning › representation learning › visual representation learning
visual correspondence learning
0.412020
Space-Time Correspondence as a Contrastive Random Walk · NeurIPS 2020
Machine learning › Reinforcement learning › reinforcement learning environment › environment design
automatic curriculum learning
0.412019
Unsupervised Curricula for Visual Meta-Reinforcement Learning · NeurIPS 2019
Computer vision › 3D vision › correspondence estimation
image correspondence
0.412019
Learning Correspondence From the Cycle-Consistency of Time · CVPR 2019
Computer vision › Video understanding and tracking › object tracking
keypoint tracking
0.412019
Learning Correspondence From the Cycle-Consistency of Time · CVPR 2019
Machine learning › Reinforcement learning
meta-reinforcement learning
0.412019
Unsupervised Curricula for Visual Meta-Reinforcement Learning · NeurIPS 2019
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
self-supervised visual representation learning
0.412019
Learning Correspondence From the Cycle-Consistency of Time · CVPR 2019
Machine learning › Reinforcement learning › goal-conditioned reinforcement learning
goal-conditioned policy learning
0.312018
Universal Planning Networks: Learning Generalizable Representations for Visuomotor Control · ICML 2018
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
plan representation
0.312018
Universal Planning Networks: Learning Generalizable Representations for Visuomotor Control · ICML 2018

Methods — techniques the papers use, named apart from their topics

object-centric encoding · 1.5diffusion · 1.5self-guidance · 1.3internal representation steering · 1.3weak supervision · 0.8self-attention · 0.7pseudo-likelihood estimation · 0.7masked autoencoding · 0.7latent self-conditioning · 0.7cross-attention · 0.7
YearPublicationVenuePosition
2024 DORSal: Diffusion for Object-centric Representations of Scenes et al
Allan Jabri, Sjoerd van Steenkiste, Emiel Hoogeboom, Mehdi S. M. Sajjadi, Thomas Kipf
ICLR1
2023 Scalable Adaptive Computation for Iterative Generation
abstract
Natural data is redundant yet predominant architectures tile computation uniformly across their input and output space. We propose the Recurrent Interface Network (RIN), an attention-based architecture that decouples its core computation from the dimensionality of the data, enabling adaptive computation for more scalable generation of high-dimensional data. RINs focus the bulk of computation (i.e. global self-attention) on a set of latent tokens, using cross-attention to read and write (i.e. route) information between latent and data tokens. Stacking RIN blocks allows bottom-up (data to latent) and top-down (latent to data) feedback, leading to deeper and more expressive routing. While this routing introduces challenges, this is less problematic in recurrent computation settings where the task (and routing problem) changes gradually, such as iterative generation with diffusion models. We show how to leverage recurrence by conditioning the latent tokens at each forward pass of the reverse diffusion process with those from prior computation, i.e. latent self-conditioning. RINs yield state-of-the-art pixel diffusion models for image and video generation, scaling to1024×1024 images without cascades or guidance, while being domain-agnostic and up to 10× more efficient than 2D and 3D U-Nets.
Allan Jabri, David J. Fleet, Ting Chen 0007
ICML1
2023 Diffusion Self-Guidance for Controllable Image Generation
abstract
Large-scale generative models are capable of producing high-quality images from detailed prompts. However, many aspects of an image are difficult or impossible to convey through text. We introduce self-guidance, a method that provides precise control over properties of the generated image by guiding the internal representations of diffusion models. We demonstrate that the size, location, and appearance of objects can be extracted from these representations, and show how to use them to steer the sampling process. Self-guidance operates similarly to standard classifier guidance, but uses signals present in the pretrained model itself, requiring no additional models or training. We demonstrate the flexibility and effectiveness of self-guided generation through a wide range of challenging image manipulations, such as modifying the position or size of a single object (keeping the rest of the image unchanged), merging the appearance of objects in one image with the layout of another, composing objects from multiple images into one, and more. We also propose a new method for reconstruction using self-guidance, which allows extending our approach to editing real images.
Dave Epstein, Allan Jabri, Ben Poole, Alexei A. Efros, Aleksander Holynski
NeurIPS2
2023 MIMEx: Intrinsic Rewards from Masked Input Modeling
abstract
Exploring in environments with high-dimensional observations is hard. One promising approach for exploration is to use intrinsic rewards, which often boils down to estimating "novelty" of states, transitions, or trajectories with deep networks. Prior works have shown that conditional prediction objectives such as masked autoencoding can be seen as stochastic estimation of pseudo-likelihood. We show how this perspective naturally leads to a unified view on existing intrinsic reward approaches: they are special cases of conditional prediction, where the estimation of novelty can be seen as pseudo-likelihood estimation with different mask distributions. From this view, we propose a general framework for deriving intrinsic rewards -- Masked Input Modeling for Exploration (MIMEx) -- where the mask distribution can be flexibly tuned to control the difficulty of the underlying conditional prediction task. We demonstrate that MIMEx can achieve superior results when compared against competitive baselines on a suite of challenging sparse-reward visuomotor tasks.
Toru Lin, Allan Jabri
NeurIPS2
2022 Discovering Objects that Can Move
abstract
This paper studies the problem of object discovery - separating objects from the background without manual labels. Existing approaches utilize appearance cues, such as color, texture, and location, to group pixels into object-like regions. However, by relying on appearance alone, these methods fail to separate objects from the background in cluttered scenes. This is a fundamental limitation since the definition of an object is inherently ambiguous and context-dependent. To resolve this ambiguity, we choose to focus on dynamic objects - entities that can move independently in the world. We then scale the recent auto-encoder based frameworks for unsuper-vised object discovery from toy synthetic images to complex real-world scenes. To this end, we simplify their architecture, and augment the resulting model with a weak learning signal from general motion segmentation algorithms. Our experiments demonstrate that, despite only capturing a small subset of the objects that move, this signal is enough to generalize to segment both moving and static instances of dynamic objects. We show that our model scales to a newly collected, photo- realistic synthetic dataset with street driving scenarios. Additionally, we leverage ground truth segmentation and flow annotations in this dataset for thorough ablation and evaluation. Finally, our experiments on the real-world KITTI benchmark demonstrate that the proposed approach outperforms both heuristic- and learning-based methods by capitalizing on motion cues.
Zhipeng Bao, Pavel Tokmakov, Allan Jabri, Yu-Xiong Wang, Adrien Gaidon, Martial Hebert
CVPR3
2022 Learning Pixel Trajectories with Multiscale Contrastive Random Walks
abstract
A range of video modeling tasks, from optical flow to multiple object tracking, share the same fundamental challenge: establishing space-time correspondence. Yet, approaches that dominate each space differ. We take a step to-wards bridging this gap by extending the recent contrastive random walk formulation to much denser, pixel-level spacetime graphs. The main contribution is introducing hierarchy into the search problem by computing the transition matrix between two frames in a coarse-to-fine manner, forming a multiscale contrastive random walk when ex-tended in time. This establishes a unified technique for self-supervised learning of optical flow, keypoint tracking, and video object segmentation. Experiments demonstrate that, for each of these tasks, the unified model achieves performance competitive with strong self-supervised approaches specific to that task.11Project page at https://jasonbian97.github.io/flowwalk
Zhangxing Bian, Allan Jabri, Alexei A. Efros, Andrew Owens
CVPR2
2022 Object Permanence Emerges in a Random Walk along Memory
abstract
This paper proposes a self-supervised objective for learning representations that localize objects under occlusion - a property known as object permanence. A central question is the choice of learning signal in cases of total occlusion. Rather than directly supervising the locations of invisible objects, we propose a self-supervised objective that requires neither human annotation, nor assumptions about object dynamics. We show that object permanence can emerge by optimizing for temporal coherence of memory: we fit a Markov walk along a space-time graph of memories, where the states in each time step are non-Markovian features from a sequence encoder. This leads to a memory representation that stores occluded objects and predicts their motion, to better localize them. The resulting model outperforms existing approaches on several datasets of increasing complexity and realism, despite requiring minimal supervision, and hence being broadly applicable.
Pavel Tokmakov, Allan Jabri, Jie Li 0031, Adrien Gaidon
ICML2
2020 Towards Practical Multi-Object Manipulation using Relational Reinforcement Learning
abstract
Learning robotic manipulation tasks using reinforcement learning with sparse rewards is currently impractical due to the outrageous data requirements. Many practical tasks require manipulation of multiple objects, and the complexity of such tasks increases with the number of objects. Learning from a curriculum of increasingly complex tasks appears to be a natural solution, but unfortunately, does not work for many scenarios. We hypothesize that the inability of the state- of-the-art algorithms to effectively utilize a task curriculum stems from the absence of inductive biases for transferring knowledge from simpler to complex tasks. We show that graph-based relational architectures overcome this limitation and enable learning of complex tasks when provided with a simple curriculum of tasks with increasing numbers of objects. We demonstrate the utility of our framework on a simulated block stacking task. Starting from scratch, our agent learns to stack six blocks into a tower. Despite using step-wise sparse rewards, our method is orders of magnitude more data- efficient and outperforms the existing state-of-the-art method that utilizes human demonstrations. Furthermore, the learned policy exhibits zero-shot generalization, successfully stacking blocks into taller towers and previously unseen configurations such as pyramids, without any further training.
Allan Jabri, Trevor Darrell, Pulkit Agrawal 0001
ICRA2
2020 Space-Time Correspondence as a Contrastive Random Walk
abstract
This paper proposes a simple self-supervised approach for learning a representation for visual correspondence from raw video. We cast correspondence as prediction of links in a space-time graph constructed from video. In this graph, the nodes are patches sampled from each frame, and nodes adjacent in time can share a directed edge. We learn a representation in which pairwise similarity defines transition probability of a random walk, such that prediction of long-range correspondence is computed as a walk along the graph. We optimize the representation to place high probability along paths of similarity. Targets for learning are formed without supervision, by cycle-consistency: the objective is to maximize the likelihood of returning to the initial node when walking along a graph constructed from a palindrome of frames. Thus, a single path-level constraint implicitly supervises chains of intermediate comparisons. When used as a similarity metric without adaptation, the learned representation outperforms the self-supervised state-of-the-art on label propagation tasks involving objects, semantic parts, and pose. Moreover, we demonstrate that a technique we call edge dropout, as well as self-supervised adaptation at test-time, further improve transfer for object-centric correspondence.
Allan Jabri, Andrew Owens, Alexei A. Efros
NeurIPS1
2019 Learning Correspondence From the Cycle-Consistency of Time
abstract
We introduce a self-supervised method for learning visual correspondence from unlabeled video. The main idea is to use cycle-consistency in time as free supervisory signal for learning visual representations from scratch. At training time, our model learns a feature map representation to be useful for performing cycle-consistent tracking. At test time, we use the acquired representation to find nearest neighbors across space and time. We demonstrate the generalizability of the representation -- without finetuning -- across a range of visual correspondence tasks, including video object segmentation, keypoint tracking, and optical flow. Our approach outperforms previous self-supervised methods and performs competitively with strongly supervised methods.
Xiaolong Wang 0004, Allan Jabri, Alexei A. Efros
CVPR2
2019 Unsupervised Curricula for Visual Meta-Reinforcement Learning
abstract
In principle, meta-reinforcement learning algorithms leverage experience across many tasks to learn fast and effective reinforcement learning (RL) strategies. However, current meta-RL approaches rely on manually-defined distributions of training tasks, and hand-crafting these task distributions can be challenging and time-consuming. Can ``useful'' pre-training tasks be discovered in an unsupervised manner? We develop an unsupervised algorithm for inducing an adaptive meta-training task distribution, i.e. an automatic curriculum, by modeling unsupervised interaction in a visual environment. The task distribution is scaffolded by a parametric density model of the meta-learner's trajectory distribution. We formulate unsupervised meta-RL as information maximization between a latent task variable and the meta-learner’s data distribution, and describe a practical instantiation which alternates between integration of recent experience into the task distribution and meta-learning of the updated tasks. Repeating this procedure leads to iterative reorganization such that the curriculum adapts as the meta-learner's data distribution shifts. Moreover, we show how discriminative clustering frameworks for visual representations can support trajectory-level task acquisition and exploration in domains with pixel observations, avoiding the pitfalls of alternatives. In experiments on vision-based navigation and manipulation domains, we show that the algorithm allows for unsupervised meta-learning that both transfers to downstream tasks specified by hand-crafted reward functions and serves as pre-training for more efficient meta-learning of test task distributions.
Allan Jabri, Kyle Hsu, Abhishek Gupta 0004, Benjamin Eysenbach, Sergey Levine, Chelsea Finn
NeurIPS1
2018 Universal Planning Networks: Learning Generalizable Representations for Visuomotor Control
abstract
A key challenge in complex visuomotor control is learning abstract representations that are effective for specifying goals, planning, and generalization. To this end, we introduce universal planning networks (UPN). UPNs embed differentiable planning within a goal-directed policy. This planning computation unrolls a forward model in a latent space and infers an optimal action plan through gradient descent trajectory optimization. The plan-by-gradient-descent process and its underlying representations are learned end-to-end to directly optimize a supervised imitation learning objective. We find that the representations learned are not only effective for goal-directed visual imitation via gradient-based trajectory optimization, but can also provide a metric for specifying goals using images. The learned representations can be leveraged to specify distance-based rewards to reach new target states for model-free reinforcement learning, resulting in substantially more effective learning when solving new tasks described via image based goals. We were able to achieve successful transfer of visuomotor planning strategies across robots with significantly different morphologies and actuation capabilities. Visit https://sites.google. com/view/upn-public/home for video highlights.
Aravind Srinivas, Allan Jabri, Pieter Abbeel, Sergey Levine, Chelsea Finn
ICML2
2018 Learning Visually Grounded Sentence Representations
abstract
Douwe Kiela, Alexis Conneau, Allan Jabri, Maximilian Nickel. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Douwe Kiela, Alexis Conneau, Allan Jabri, Maximilian Nickel
NAACL-HLT3
2017 Learning Visual N-Grams from Web Data
Allan Jabri, Armand Joulin, Laurens van der Maaten
ICCV2
2016 Revisiting Visual Question Answering Baselines
Allan Jabri, Armand Joulin, Laurens van der Maaten
ECCV (8)1
2016 Learning Visual Features from Large Weakly Supervised Data
Armand Joulin, Laurens van der Maaten, Allan Jabri, Nicolas Vasilache
ECCV (7)3