VLDB 2026 Research / reviewers in the wild / expert
Rishabh Kabra
dblp:234/8010
· DBLP profile ↗
10ranked-venue papers
2as first author
8since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Generative modeling · 28% Video understanding and tracking · 20% 3D vision · 19% | |
| Computer graphics and multimedia
1 paper |
Visual content generation and editing · 100% |
Topics — the 19 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Video understanding and tracking
video representation learning |
1.8 | 3 | 2025 | From Image to Video: An Empirical Study of Diffusion Representations · ICCV 2025 Moving Off-the-Grid: Scene-Grounded Video Representations · NeurIPS 2024 SIMONe: View-Invariant, Temporally-Abstracted Object Representations via Unsupervised Video Decomposition · NeurIPS 2021 |
Machine learning › Generative modeling
diffusion model |
1.0 | 2 | 2025 | Neural Assets: 3D-Aware Multi-Object Scene Synthesis with Image Diffusion Models · NeurIPS 2024 From Image to Video: An Empirical Study of Diffusion Representations · ICCV 2025 |
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning |
1.0 | 2 | 2021 | SIMONe: View-Invariant, Temporally-Abstracted Object Representations via Unsupervised Video Decomposition · NeurIPS 2021 PARTS: Unsupervised segmentation with slots, attention and independence maximization · ICCV 2021 |
Computer vision › Segmentation and scene understanding › object segmentation
unsupervised object segmentation |
0.9 | 2 | 2021 | PARTS: Unsupervised segmentation with slots, attention and independence maximization · ICCV 2021 Multi-Object Representation Learning with Iterative Variational Inference · ICML 2019 |
Machine learning › Generative modeling › diffusion model › diffusion-based representation learning
diffusion model features |
0.9 | 1 | 2025 | From Image to Video: An Empirical Study of Diffusion Representations · ICCV 2025 |
Machine learning › Generative modeling › generative adversarial network
3d-aware image synthesis |
0.8 | 1 | 2024 | Neural Assets: 3D-Aware Multi-Object Scene Synthesis with Image Diffusion Models · NeurIPS 2024 |
Computer vision › 3D vision › 3d shape analysis
3d shape understanding |
0.8 | 1 | 2024 | Leveraging VLM-Based Pipelines to Annotate 3D Objects · ICML 2024 |
Machine learning › Generative modeling › diffusion model › text-to-image generation
text-to-image diffusion model |
0.8 | 1 | 2024 | Neural Assets: 3D-Aware Multi-Object Scene Synthesis with Image Diffusion Models · NeurIPS 2024 |
Computer vision › Vision and language
vision-language model |
0.8 | 1 | 2024 | Leveraging VLM-Based Pipelines to Annotate 3D Objects · ICML 2024 |
Visual content generation and editing
image editing |
0.8 | 1 | 2024 | Neural Assets: 3D-Aware Multi-Object Scene Synthesis with Image Diffusion Models · NeurIPS 2024 |
Computer vision › 3D vision
object representation |
0.7 | 2 | 2021 | PARTS: Unsupervised segmentation with slots, attention and independence maximization · ICCV 2021 Unsupervised Object-Based Transition Models For 3D Partially Observable Environments · NeurIPS 2021 |
Machine learning › Reinforcement learning
model-based reinforcement learning |
0.5 | 1 | 2021 | Unsupervised Object-Based Transition Models For 3D Partially Observable Environments · NeurIPS 2021 |
Computer vision › 3D vision › 3d scene understanding
scene decomposition |
0.5 | 1 | 2021 | SIMONe: View-Invariant, Temporally-Abstracted Object Representations via Unsupervised Video Decomposition · NeurIPS 2021 |
Machine learning › Reinforcement learning
model-free reinforcement learning |
0.4 | 1 | 2019 | An Investigation of Model-Free Planning · ICML 2019 |
Machine learning › Representation and self-supervised learning › representation learning
object-centric representation learning |
0.4 | 1 | 2019 | Multi-Object Representation Learning with Iterative Variational Inference · ICML 2019 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
search-based planning |
0.4 | 1 | 2019 | An Investigation of Model-Free Planning · ICML 2019 |
Machine learning › Generative modeling
variational autoencoder |
0.4 | 1 | 2019 | Multi-Object Representation Learning with Iterative Variational Inference · ICML 2019 |
Computer vision › 3D vision › 3d scene understanding
scene structure |
0.2 | 1 | 2024 | Moving Off-the-Grid: Scene-Grounded Video Representations · NeurIPS 2024 |
Computer vision › Video understanding and tracking
video prediction |
0.1 | 1 | 2021 | PARTS: Unsupervised segmentation with slots, attention and independence maximization · ICCV 2021 |
Methods — techniques the papers use, named apart from their topics
diffusion model · 2.4fine-tuning · 1.5slot attention · 1.0representation probing · 0.9probabilistic aggregation · 0.8positional embedding · 0.8next-frame prediction · 0.8joint image-text likelihood · 0.8cross-attention · 0.8generative model · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | From Image to Video: An Empirical Study of Diffusion RepresentationsabstractDiffusion models have revolutionized generative modeling, enabling unprecedented realism in image and video synthesis. This success has sparked interest in leveraging their representations for visual understanding tasks. While recent works have explored this potential for image generation, the visual understanding capabilities of video diffusion models remain largely uncharted. To address this gap, we systematically compare the same model architecture trained for video versus image generation, analyzing the performance of their latent representations on various downstream tasks including image classification, action recognition, depth estimation, and tracking. Results show that video diffusion models consistently outperform their image counterparts, though we find a striking range in the extent of this superiority. We further analyze features extracted from different layers and with varying noise levels, as well as the effect of model size and training budget on representation and generation quality. This work marks the first direct comparison of video and image diffusion objectives for visual understanding, offering insights into the role of temporal information in representation learning. Pedro Vélez, Luisa F. Polanía, Yi Yang 0007, Rishabh Kabra, Anurag Arnab, Mehdi S. M. Sajjadi |
ICCV | 5 |
| 2025 | OpenWorldSAM: Extending SAM2 for Universal Image Segmentation with Language PromptsabstractThe ability to segment objects based on open-ended language prompts remains a critical challenge, requiring models to ground textual semantics into precise spatial masks while handling diverse and unseen categories. We present OpenWorldSAM, a framework that extends the prompt-driven Segment Anything Model v2 (SAM2) to open-vocabulary scenarios by integrating multi-modal embeddings extracted from a lightweight vision-language model (VLM). Our approach is guided by four key principles: i) Unified prompting: OpenWorldSAM supports a diverse range of prompts, including category-level and sentence-level language descriptions, providing a flexible interface for various segmentation tasks. ii) Efficiency: By freezing the pre-trained components of SAM2 and the VLM, we train only 4.5 million parameters on the COCO-stuff dataset, achieving remarkable resource efficiency. iii) Instance Awareness: We enhance the model's spatial understanding through novel positional tie-breaker embeddings and cross-attention layers, enabling effective segmentation of multiple instances. iv) Generalization: OpenWorldSAM exhibits strong zero-shot capabilities, generalizing well on unseen categories and an open vocabulary of concepts without additional training. Extensive experiments demonstrate that OpenWorldSAM achieves state-of-the-art performance in open-vocabulary semantic, instance, and panoptic segmentation across multiple benchmarks. Code is available at https://github.com/GinnyXiao/OpenWorldSAM. Shiting Xiao, Rishabh Kabra, Yuhang Li 0001, Donghyun Lee 0002, João Carreira 0001, Priyadarshini Panda |
NeurIPS | 2 |
| 2024 | Leveraging VLM-Based Pipelines to Annotate 3D ObjectsabstractPretrained vision language models (VLMs) present an opportunity to caption unlabeled 3D objects at scale. The leading approach to summarize VLM descriptions from different views of an object (Luo et al., 2023) relies on a language model (GPT4) to produce the final output. This text-based aggregation is susceptible to hallucinations as it merges potentially contradictory descriptions. We propose an alternative algorithm to marginalize over factors such as the viewpoint that affect the VLM’s response. Instead of merging text-only responses, we utilize the VLM’s joint image-text likelihoods. We show our probabilistic aggregation is not only more reliable and efficient, but sets the SoTA on inferring object types with respect to human-verified labels. The aggregated annotations are also useful for conditional inference; they improve downstream predictions (e.g., of object material) when the object’s type is specified as an auxiliary text-based input. Such auxiliary inputs allow ablating the contribution of visual reasoning over visionless reasoning in an unsupervised setting. With these supervised and unsupervised evaluations, we show how a VLM-based pipeline can be leveraged to produce reliable annotations for 764K objects from the Objaverse dataset. Rishabh Kabra, Loïc Matthey, Alexander Lerchner, Niloy J. Mitra |
ICML | 1 |
| 2024 | Moving Off-the-Grid: Scene-Grounded Video RepresentationsabstractCurrent vision models typically maintain a fixed correspondence between their representation structure and image space.
Each layer comprises a set of tokens arranged “on-the-grid,” which biases patches or tokens to encode information at a specific spatio(-temporal) location. In this work we present *Moving Off-the-Grid* (MooG), a self-supervised video representation model that offers an alternative approach, allowing tokens to move “off-the-grid” to better enable them to represent scene elements consistently, even as they move across the image plane through time. By using a combination of cross-attention and positional embeddings we disentangle the representation structure and image structure. We find that a simple self-supervised objective—next frame prediction—trained on video data, results in a set of latent tokens which bind to specific scene structures and track them as they move. We demonstrate the usefulness of MooG’s learned representation both qualitatively and quantitatively by training readouts on top of the learned representation on a variety of downstream tasks. We show that MooG can provide a strong foundation for different vision tasks when compared to “on-the-grid” baselines. Sjoerd van Steenkiste, Daniel Zoran, Yi Yang 0007, Yulia Rubanova, Rishabh Kabra, Carl Doersch, Dilara Gokay, Joseph Heyward, Etienne Pot, Klaus Greff, Drew A. Hudson, Thomas Keck, João Carreira 0001, Alexey Dosovitskiy, Mehdi S. M. Sajjadi, Thomas Kipf |
NeurIPS | 5 |
| 2024 | Neural Assets: 3D-Aware Multi-Object Scene Synthesis with Image Diffusion ModelsabstractWe address the problem of multi-object 3D pose control in image diffusion models. Instead of conditioning on a sequence of text tokens, we propose to use a set of per-object representations, *Neural Assets*, to control the 3D pose of individual objects in a scene. Neural Assets are obtained by pooling visual representations of objects from a reference image, such as a frame in a video, and are trained to reconstruct the respective objects in a different image, e.g., a later frame in the video. Importantly, we encode object visuals from the reference image while conditioning on object poses from the target frame, which enables learning disentangled appearance and position features. Combining visual and 3D pose representations in a sequence-of-tokens format allows us to keep the text-to-image interface of existing models, with Neural Assets in place of text tokens. By fine-tuning a pre-trained text-to-image diffusion model with this information, our approach enables fine-grained 3D pose and placement control of individual objects in a scene. We further demonstrate that Neural Assets can be transferred and recomposed across different scenes. Our model achieves state-of-the-art multi-object editing results on both synthetic 3D scene datasets, as well as two real-world video datasets (Objectron, Waymo Open). Ziyi Wu 0002, Yulia Rubanova, Rishabh Kabra, Drew A. Hudson, Igor Gilitschenski, Yusuf Aytar, Sjoerd van Steenkiste, Kelsey R. Allen, Thomas Kipf |
NeurIPS | 3 |
| 2021 | PARTS: Unsupervised segmentation with slots, attention and independence maximizationabstractFrom an early age, humans perceive the visual world as composed of coherent objects with distinctive properties such as shape, size, and color. There is great interest in building models that are able to learn similar structure, ideally in an unsupervised manner. Learning such structure from complex 3D scenes that include clutter, occlusions, interactions, and camera motion is still an open challenge. We present a model that is able to segment visual scenes from complex 3D environments into distinct objects, learn disentangled representations of individual objects, and form consistent and coherent predictions of future frames, in a fully unsupervised manner. Our model (named PARTS) builds on recent approaches that utilize iterative amortized inference and transition dynamics for deep generative models. We achieve dramatic improvements in performance by introducing several novel contributions. We introduce a recurrent slot-attention like encoder which allows for top-down influence during inference. We argue that when inferring scene structure from image sequences it is better to use a fixed prior which is shared across the sequence rather than an auto-regressive prior as often used in prior work. We demonstrate our model’s success on three different video datasets (the popular benchmark CLEVRER; a simulated 3D Playroom environment; and a real-world Robotics Arm dataset). Finally, we analyze the contributions of the various model components and the representations learned by the model. Daniel Zoran, Rishabh Kabra, Alexander Lerchner, Danilo Jimenez Rezende |
ICCV | 2 |
| 2021 | Unsupervised Object-Based Transition Models For 3D Partially Observable EnvironmentsabstractWe present a slot-wise, object-based transition model that decomposes a scene into objects, aligns them (with respect to a slot-wise object memory) to maintain a consistent order across time, and predicts how those objects evolve over successive frames. The model is trained end-to-end without supervision using transition losses at the level of the object-structured representation rather than pixels. Thanks to the introduction of our novel alignment module, the model deals properly with two issues that are not handled satisfactorily by other transition models, namely object persistence and object identity. We show that the combination of an object-level loss and correct object alignment over time enables the model to outperform a state-of-the-art baseline, and allows it to deal well with object occlusion and re-appearance in partially observable environments. Antonia Creswell, Rishabh Kabra, Chris Burgess 0001, Murray Shanahan |
NeurIPS | 2 |
| 2021 | SIMONe: View-Invariant, Temporally-Abstracted Object Representations via Unsupervised Video DecompositionabstractTo help agents reason about scenes in terms of their building blocks, we wish to extract the compositional structure of any given scene (in particular, the configuration and characteristics of objects comprising the scene). This problem is especially difficult when scene structure needs to be inferred while also estimating the agent’s location/viewpoint, as the two variables jointly give rise to the agent’s observations. We present an unsupervised variational approach to this problem. Leveraging the shared structure that exists across different scenes, our model learns to infer two sets of latent representations from RGB video input alone: a set of "object" latents, corresponding to the time-invariant, object-level contents of the scene, as well as a set of "frame" latents, corresponding to global time-varying elements such as viewpoint. This factorization of latents allows our model, SIMONe, to represent object attributes in an allocentric manner which does not depend on viewpoint. Moreover, it allows us to disentangle object dynamics and summarize their trajectories as time-abstracted, view-invariant, per-object properties. We demonstrate these capabilities, as well as the model's performance in terms of view synthesis and instance segmentation, across three procedurally generated video datasets. Rishabh Kabra, Daniel Zoran, Goker Erdogan, Loïc Matthey, Antonia Creswell, Matt M. Botvinick, Alexander Lerchner, Chris Burgess 0001 |
NeurIPS | 1 |
| 2019 | Multi-Object Representation Learning with Iterative Variational InferenceabstractHuman perception is structured around objects which form the basis for our higher-level cognition and impressive systematic generalization abilities. Yet most work on representation learning focuses on feature learning without even considering multiple objects, or treats segmentation as an (often supervised) preprocessing step. Instead, we argue for the importance of learning to segment and represent objects jointly. We demonstrate that, starting from the simple assumption that a scene is composed of multiple entities, it is possible to learn to segment images into interpretable objects with disentangled representations. Our method learns – without supervision – to inpaint occluded parts, and extrapolates to scenes with more objects and to unseen objects with novel feature combinations. We also show that, due to the use of iterative variational inference, our system is able to learn multi-modal posteriors for ambiguous inputs and extends naturally to sequences. Klaus Greff, Raphael Lopez Kaufman, Rishabh Kabra, Nicholas Watters, Chris Burgess 0001, Daniel Zoran, Loïc Matthey, Matt M. Botvinick, Alexander Lerchner |
ICML | 3 |
| 2019 | An Investigation of Model-Free PlanningabstractThe field of reinforcement learning (RL) is facing increasingly challenging domains with combinatorial complexity. For an RL agent to address these challenges, it is essential that it can plan effectively. Prior work has typically utilized an explicit model of the environment, combined with a specific planning algorithm (such as tree search). More recently, a new family of methods have been proposed that learn how to plan, by providing the structure for planning via an inductive bias in the function approximator (such as a tree structured neural network), trained end-to-end by a model-free RL algorithm. In this paper, we go even further, and demonstrate empirically that an entirely model-free approach, without special structure beyond standard neural network components such as convolutional networks and LSTMs, can learn to exhibit many of the characteristics typically associated with a model-based planner. We measure our agent’s effectiveness at planning in terms of its ability to generalize across a combinatorial and irreversible state space, its data efficiency, and its ability to utilize additional thinking time. We find that our agent has many of the characteristics that one might expect to find in a planning algorithm. Furthermore, it exceeds the state-of-the-art in challenging combinatorial domains such as Sokoban and outperforms other model-free approaches that utilize strong inductive biases toward planning. Arthur Guez, Mehdi Mirza, Karol Gregor, Rishabh Kabra, Sébastien Racanière, Theophane Weber, David Raposo, Adam Santoro, Laurent Orseau, Tom Eccles, Greg Wayne, David Silver 0001, Timothy P. Lillicrap |
ICML | 4 |