VLDB 2026 Research / reviewers in the wild / expert
Allan Jabri
dblp:172/0858
· DBLP profile ↗
16ranked-venue papers
5as first author
7since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
15 papers |
3D vision · 18% Generative modeling · 18% Reinforcement learning · 16% | |
| Computer graphics and multimedia
2 papers |
Visual content generation and editing · 74% Image and video processing · 26% |
Topics — the 30 heaviest of 47, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
2.1 | 3 | 2024 | DORSal: Diffusion for Object-centric Representations of Scenes et al · ICLR 2024 Diffusion Self-Guidance for Controllable Image Generation · NeurIPS 2023 Scalable Adaptive Computation for Iterative Generation · ICML 2023 |
Computer vision › Video understanding and tracking
video object segmentation |
1.0 | 2 | 2022 | Learning Pixel Trajectories with Multiscale Contrastive Random Walks · CVPR 2022 Learning Correspondence From the Cycle-Consistency of Time · CVPR 2019 |
Machine learning › Generative modeling › scene generation
object-centric scene generation |
0.8 | 1 | 2024 | DORSal: Diffusion for Object-centric Representations of Scenes et al · ICLR 2024 |
Computer vision › 3D vision
object representation |
0.8 | 1 | 2024 | DORSal: Diffusion for Object-centric Representations of Scenes et al · ICLR 2024 |
Computer vision › 3D vision › 3d scene modeling
scene representation |
0.8 | 1 | 2024 | DORSal: Diffusion for Object-centric Representations of Scenes et al · ICLR 2024 |
Computer vision › 3D vision › motion estimation
optical flow |
0.7 | 2 | 2022 | Learning Pixel Trajectories with Multiscale Contrastive Random Walks · CVPR 2022 Learning Correspondence From the Cycle-Consistency of Time · CVPR 2019 |
Machine learning › Efficient and distributed learning
adaptive computation |
0.7 | 1 | 2023 | Scalable Adaptive Computation for Iterative Generation · ICML 2023 |
Machine learning › Generative modeling › diffusion model › controllable generation
controllable image generation |
0.7 | 1 | 2023 | Diffusion Self-Guidance for Controllable Image Generation · NeurIPS 2023 |
Machine learning › Reinforcement learning
exploration |
0.7 | 1 | 2023 | MIMEx: Intrinsic Rewards from Masked Input Modeling · NeurIPS 2023 |
Machine learning › Reinforcement learning › exploration
intrinsic motivation |
0.7 | 1 | 2023 | MIMEx: Intrinsic Rewards from Masked Input Modeling · NeurIPS 2023 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
masked modeling |
0.7 | 1 | 2023 | MIMEx: Intrinsic Rewards from Masked Input Modeling · NeurIPS 2023 |
Visual content generation and editing
image editing |
0.7 | 1 | 2023 | Diffusion Self-Guidance for Controllable Image Generation · NeurIPS 2023 |
Computer vision › Video understanding and tracking
motion segmentation |
0.6 | 1 | 2022 | Discovering Objects that Can Move · CVPR 2022 |
Computer vision › Image recognition and object detection
object discovery |
0.6 | 1 | 2022 | Discovering Objects that Can Move · CVPR 2022 |
Computer vision › Video understanding and tracking › video analytics › video object analysis › object-centric video understanding
object permanence |
0.6 | 1 | 2022 | Object Permanence Emerges in a Random Walk along Memory · ICML 2022 |
Computer vision › Video understanding and tracking
spatio-temporal correspondence |
0.6 | 1 | 2022 | Learning Pixel Trajectories with Multiscale Contrastive Random Walks · CVPR 2022 |
Computer vision › Image recognition and object detection › object discovery
unsupervised object discovery |
0.6 | 1 | 2022 | Discovering Objects that Can Move · CVPR 2022 |
Computer vision › 3D vision › motion estimation › optical flow
unsupervised optical flow |
0.6 | 1 | 2022 | Learning Pixel Trajectories with Multiscale Contrastive Random Walks · CVPR 2022 |
Robotics › Robot manipulation › assembly › object assembly
block stacking |
0.4 | 1 | 2020 | Towards Practical Multi-Object Manipulation using Relational Reinforcement Learning · ICRA 2020 |
Machine learning › Representation and self-supervised learning
cycle consistency |
0.4 | 1 | 2020 | Space-Time Correspondence as a Contrastive Random Walk · NeurIPS 2020 |
Robotics › Robot manipulation › object manipulation
multi-object manipulation |
0.4 | 1 | 2020 | Towards Practical Multi-Object Manipulation using Relational Reinforcement Learning · ICRA 2020 |
Machine learning › Reinforcement learning
relational reinforcement learning |
0.4 | 1 | 2020 | Towards Practical Multi-Object Manipulation using Relational Reinforcement Learning · ICRA 2020 |
Machine learning › Representation and self-supervised learning › representation learning › visual representation learning
visual correspondence learning |
0.4 | 1 | 2020 | Space-Time Correspondence as a Contrastive Random Walk · NeurIPS 2020 |
Machine learning › Reinforcement learning › reinforcement learning environment › environment design
automatic curriculum learning |
0.4 | 1 | 2019 | Unsupervised Curricula for Visual Meta-Reinforcement Learning · NeurIPS 2019 |
Computer vision › 3D vision › correspondence estimation
image correspondence |
0.4 | 1 | 2019 | Learning Correspondence From the Cycle-Consistency of Time · CVPR 2019 |
Computer vision › Video understanding and tracking › object tracking
keypoint tracking |
0.4 | 1 | 2019 | Learning Correspondence From the Cycle-Consistency of Time · CVPR 2019 |
Machine learning › Reinforcement learning
meta-reinforcement learning |
0.4 | 1 | 2019 | Unsupervised Curricula for Visual Meta-Reinforcement Learning · NeurIPS 2019 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
self-supervised visual representation learning |
0.4 | 1 | 2019 | Learning Correspondence From the Cycle-Consistency of Time · CVPR 2019 |
Machine learning › Reinforcement learning › goal-conditioned reinforcement learning
goal-conditioned policy learning |
0.3 | 1 | 2018 | Universal Planning Networks: Learning Generalizable Representations for Visuomotor Control · ICML 2018 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
plan representation |
0.3 | 1 | 2018 | Universal Planning Networks: Learning Generalizable Representations for Visuomotor Control · ICML 2018 |
Methods — techniques the papers use, named apart from their topics
object-centric encoding · 1.5diffusion · 1.5self-guidance · 1.3internal representation steering · 1.3weak supervision · 0.8self-attention · 0.7pseudo-likelihood estimation · 0.7masked autoencoding · 0.7latent self-conditioning · 0.7cross-attention · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | DORSal: Diffusion for Object-centric Representations of Scenes et al
Allan Jabri, Sjoerd van Steenkiste, Emiel Hoogeboom, Mehdi S. M. Sajjadi, Thomas Kipf |
ICLR | 1 |
| 2023 | Scalable Adaptive Computation for Iterative GenerationabstractNatural data is redundant yet predominant architectures tile computation uniformly across their input and output space. We propose the Recurrent Interface Network (RIN), an attention-based architecture that decouples its core computation from the dimensionality of the data, enabling adaptive computation for more scalable generation of high-dimensional data. RINs focus the bulk of computation (i.e. global self-attention) on a set of latent tokens, using cross-attention to read and write (i.e. route) information between latent and data tokens. Stacking RIN blocks allows bottom-up (data to latent) and top-down (latent to data) feedback, leading to deeper and more expressive routing. While this routing introduces challenges, this is less problematic in recurrent computation settings where the task (and routing problem) changes gradually, such as iterative generation with diffusion models. We show how to leverage recurrence by conditioning the latent tokens at each forward pass of the reverse diffusion process with those from prior computation, i.e. latent self-conditioning. RINs yield state-of-the-art pixel diffusion models for image and video generation, scaling to1024×1024 images without cascades or guidance, while being domain-agnostic and up to 10× more efficient than 2D and 3D U-Nets. Allan Jabri, David J. Fleet, Ting Chen 0007 |
ICML | 1 |
| 2023 | Diffusion Self-Guidance for Controllable Image GenerationabstractLarge-scale generative models are capable of producing high-quality images from detailed prompts. However, many aspects of an image are difficult or impossible to convey through text. We introduce self-guidance, a method that provides precise control over properties of the generated image by guiding the internal representations of diffusion models. We demonstrate that the size, location, and appearance of objects can be extracted from these representations, and show how to use them to steer the sampling process. Self-guidance operates similarly to standard classifier guidance, but uses signals present in the pretrained model itself, requiring no additional models or training. We demonstrate the flexibility and effectiveness of self-guided generation through a wide range of challenging image manipulations, such as modifying the position or size of a single object (keeping the rest of the image unchanged), merging the appearance of objects in one image with the layout of another, composing objects from multiple images into one, and more. We also propose a new method for reconstruction using self-guidance, which allows extending our approach to editing real images. Dave Epstein, Allan Jabri, Ben Poole, Alexei A. Efros, Aleksander Holynski |
NeurIPS | 2 |
| 2023 | MIMEx: Intrinsic Rewards from Masked Input ModelingabstractExploring in environments with high-dimensional observations is hard. One promising approach for exploration is to use intrinsic rewards, which often boils down to estimating "novelty" of states, transitions, or trajectories with deep networks. Prior works have shown that conditional prediction objectives such as masked autoencoding can be seen as stochastic estimation of pseudo-likelihood. We show how this perspective naturally leads to a unified view on existing intrinsic reward approaches: they are special cases of conditional prediction, where the estimation of novelty can be seen as pseudo-likelihood estimation with different mask distributions. From this view, we propose a general framework for deriving intrinsic rewards -- Masked Input Modeling for Exploration (MIMEx) -- where the mask distribution can be flexibly tuned to control the difficulty of the underlying conditional prediction task. We demonstrate that MIMEx can achieve superior results when compared against competitive baselines on a suite of challenging sparse-reward visuomotor tasks. Toru Lin, Allan Jabri |
NeurIPS | 2 |
| 2022 | Discovering Objects that Can MoveabstractThis paper studies the problem of object discovery - separating objects from the background without manual labels. Existing approaches utilize appearance cues, such as color, texture, and location, to group pixels into object-like regions. However, by relying on appearance alone, these methods fail to separate objects from the background in cluttered scenes. This is a fundamental limitation since the definition of an object is inherently ambiguous and context-dependent. To resolve this ambiguity, we choose to focus on dynamic objects - entities that can move independently in the world. We then scale the recent auto-encoder based frameworks for unsuper-vised object discovery from toy synthetic images to complex real-world scenes. To this end, we simplify their architecture, and augment the resulting model with a weak learning signal from general motion segmentation algorithms. Our experiments demonstrate that, despite only capturing a small subset of the objects that move, this signal is enough to generalize to segment both moving and static instances of dynamic objects. We show that our model scales to a newly collected, photo- realistic synthetic dataset with street driving scenarios. Additionally, we leverage ground truth segmentation and flow annotations in this dataset for thorough ablation and evaluation. Finally, our experiments on the real-world KITTI benchmark demonstrate that the proposed approach outperforms both heuristic- and learning-based methods by capitalizing on motion cues. Zhipeng Bao, Pavel Tokmakov, Allan Jabri, Yu-Xiong Wang, Adrien Gaidon, Martial Hebert |
CVPR | 3 |
| 2022 | Learning Pixel Trajectories with Multiscale Contrastive Random WalksabstractA range of video modeling tasks, from optical flow to multiple object tracking, share the same fundamental challenge: establishing space-time correspondence. Yet, approaches that dominate each space differ. We take a step to-wards bridging this gap by extending the recent contrastive random walk formulation to much denser, pixel-level spacetime graphs. The main contribution is introducing hierarchy into the search problem by computing the transition matrix between two frames in a coarse-to-fine manner, forming a multiscale contrastive random walk when ex-tended in time. This establishes a unified technique for self-supervised learning of optical flow, keypoint tracking, and video object segmentation. Experiments demonstrate that, for each of these tasks, the unified model achieves performance competitive with strong self-supervised approaches specific to that task.11Project page at https://jasonbian97.github.io/flowwalk Zhangxing Bian, Allan Jabri, Alexei A. Efros, Andrew Owens |
CVPR | 2 |
| 2022 | Object Permanence Emerges in a Random Walk along MemoryabstractThis paper proposes a self-supervised objective for learning representations that localize objects under occlusion - a property known as object permanence. A central question is the choice of learning signal in cases of total occlusion. Rather than directly supervising the locations of invisible objects, we propose a self-supervised objective that requires neither human annotation, nor assumptions about object dynamics. We show that object permanence can emerge by optimizing for temporal coherence of memory: we fit a Markov walk along a space-time graph of memories, where the states in each time step are non-Markovian features from a sequence encoder. This leads to a memory representation that stores occluded objects and predicts their motion, to better localize them. The resulting model outperforms existing approaches on several datasets of increasing complexity and realism, despite requiring minimal supervision, and hence being broadly applicable. Pavel Tokmakov, Allan Jabri, Jie Li 0031, Adrien Gaidon |
ICML | 2 |
| 2020 | Towards Practical Multi-Object Manipulation using Relational Reinforcement LearningabstractLearning robotic manipulation tasks using reinforcement learning with sparse rewards is currently impractical due to the outrageous data requirements. Many practical tasks require manipulation of multiple objects, and the complexity of such tasks increases with the number of objects. Learning from a curriculum of increasingly complex tasks appears to be a natural solution, but unfortunately, does not work for many scenarios. We hypothesize that the inability of the state- of-the-art algorithms to effectively utilize a task curriculum stems from the absence of inductive biases for transferring knowledge from simpler to complex tasks. We show that graph-based relational architectures overcome this limitation and enable learning of complex tasks when provided with a simple curriculum of tasks with increasing numbers of objects. We demonstrate the utility of our framework on a simulated block stacking task. Starting from scratch, our agent learns to stack six blocks into a tower. Despite using step-wise sparse rewards, our method is orders of magnitude more data- efficient and outperforms the existing state-of-the-art method that utilizes human demonstrations. Furthermore, the learned policy exhibits zero-shot generalization, successfully stacking blocks into taller towers and previously unseen configurations such as pyramids, without any further training. Allan Jabri, Trevor Darrell, Pulkit Agrawal 0001 |
ICRA | 2 |
| 2020 | Space-Time Correspondence as a Contrastive Random WalkabstractThis paper proposes a simple self-supervised approach for learning a representation for visual correspondence from raw video. We cast correspondence as prediction of links in a space-time graph constructed from video. In this graph, the nodes are patches sampled from each frame, and nodes adjacent in time can share a directed edge. We learn a representation in which pairwise similarity defines transition probability of a random walk, such that prediction of long-range correspondence is computed as a walk along the graph. We optimize the representation to place high probability along paths of similarity. Targets for learning are formed without supervision, by cycle-consistency: the objective is to maximize the likelihood of returning to the initial node when walking along a graph constructed from a palindrome of frames. Thus, a single path-level constraint implicitly supervises chains of intermediate comparisons. When used as a similarity metric without adaptation, the learned representation outperforms the self-supervised state-of-the-art on label propagation tasks involving objects, semantic parts, and pose. Moreover, we demonstrate that a technique we call edge dropout, as well as self-supervised adaptation at test-time, further improve transfer for object-centric correspondence. Allan Jabri, Andrew Owens, Alexei A. Efros |
NeurIPS | 1 |
| 2019 | Learning Correspondence From the Cycle-Consistency of TimeabstractWe introduce a self-supervised method for learning visual correspondence from unlabeled video. The main idea is to use cycle-consistency in time as free supervisory signal for learning visual representations from scratch. At training time, our model learns a feature map representation to be useful for performing cycle-consistent tracking. At test time, we use the acquired representation to find nearest neighbors across space and time. We demonstrate the generalizability of the representation -- without finetuning -- across a range of visual correspondence tasks, including video object segmentation, keypoint tracking, and optical flow. Our approach outperforms previous self-supervised methods and performs competitively with strongly supervised methods. Xiaolong Wang 0004, Allan Jabri, Alexei A. Efros |
CVPR | 2 |
| 2019 | Unsupervised Curricula for Visual Meta-Reinforcement LearningabstractIn principle, meta-reinforcement learning algorithms leverage experience across many tasks to learn fast and effective reinforcement learning (RL) strategies. However, current meta-RL approaches rely on manually-defined distributions of training tasks, and hand-crafting these task distributions can be challenging and time-consuming. Can ``useful'' pre-training tasks be discovered in an unsupervised manner? We develop an unsupervised algorithm for inducing an adaptive meta-training task distribution, i.e. an automatic curriculum, by modeling unsupervised interaction in a visual environment. The task distribution is scaffolded by a parametric density model of the meta-learner's trajectory distribution. We formulate unsupervised meta-RL as information maximization between a latent task variable and the meta-learner’s data distribution, and describe a practical instantiation which alternates between integration of recent experience into the task distribution and meta-learning of the updated tasks. Repeating this procedure leads to iterative reorganization such that the curriculum adapts as the meta-learner's data distribution shifts. Moreover, we show how discriminative clustering frameworks for visual representations can support trajectory-level task acquisition and exploration in domains with pixel observations, avoiding the pitfalls of alternatives. In experiments on vision-based navigation and manipulation domains, we show that the algorithm allows for unsupervised meta-learning that both transfers to downstream tasks specified by hand-crafted reward functions and serves as pre-training for more efficient meta-learning of test task distributions. Allan Jabri, Kyle Hsu, Abhishek Gupta 0004, Benjamin Eysenbach, Sergey Levine, Chelsea Finn |
NeurIPS | 1 |
| 2018 | Universal Planning Networks: Learning Generalizable Representations for Visuomotor ControlabstractA key challenge in complex visuomotor control is learning abstract representations that are effective for specifying goals, planning, and generalization. To this end, we introduce universal planning networks (UPN). UPNs embed differentiable planning within a goal-directed policy. This planning computation unrolls a forward model in a latent space and infers an optimal action plan through gradient descent trajectory optimization. The plan-by-gradient-descent process and its underlying representations are learned end-to-end to directly optimize a supervised imitation learning objective. We find that the representations learned are not only effective for goal-directed visual imitation via gradient-based trajectory optimization, but can also provide a metric for specifying goals using images. The learned representations can be leveraged to specify distance-based rewards to reach new target states for model-free reinforcement learning, resulting in substantially more effective learning when solving new tasks described via image based goals. We were able to achieve successful transfer of visuomotor planning strategies across robots with significantly different morphologies and actuation capabilities. Visit https://sites.google. com/view/upn-public/home for video highlights. Aravind Srinivas, Allan Jabri, Pieter Abbeel, Sergey Levine, Chelsea Finn |
ICML | 2 |
| 2018 | Learning Visually Grounded Sentence RepresentationsabstractDouwe Kiela, Alexis Conneau, Allan Jabri, Maximilian Nickel. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Douwe Kiela, Alexis Conneau, Allan Jabri, Maximilian Nickel |
NAACL-HLT | 3 |
| 2017 | Learning Visual N-Grams from Web Data
Allan Jabri, Armand Joulin, Laurens van der Maaten |
ICCV | 2 |
| 2016 | Revisiting Visual Question Answering Baselines
Allan Jabri, Armand Joulin, Laurens van der Maaten |
ECCV (8) | 1 |
| 2016 | Learning Visual Features from Large Weakly Supervised Data
Armand Joulin, Laurens van der Maaten, Allan Jabri, Nicolas Vasilache |
ECCV (7) | 3 |