VLDB 2026 Research / reviewers in the wild / expert
Andrii Zadaianchuk
dblp:274/9441
· DBLP profile ↗
10ranked-venue papers
3as first author
9since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Systems, architecture and hardware · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Reinforcement learning · 27% Representation and self-supervised learning · 24% Video understanding and tracking · 12% |
Topics — the 22 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning › representation learning
object-centric representation learning |
3.9 | 5 | 2025 | On the Transfer of Object-Centric Representation Learning · ICLR 2025 Temporally Consistent Object-Centric Learning by Contrasting Slots · CVPR 2025 CTRL-O: Language-Controllable Object-Centric Visual Representation Learning · CVPR 2025 |
Computer vision › 3D vision
object representation |
1.2 | 2 | 2023 | Unsupervised Semantic Segmentation with Self-supervised Object-centric Representations · ICLR 2023 Self-supervised Visual Reinforcement Learning with Object-centric Representations · ICLR 2021 |
Computer vision › 3D vision › neural rendering
3d gaussian splatting |
0.9 | 1 | 2025 | Dream to Manipulate: Compositional World Models Empowering Robot Imitation Learning with Imagination · ICLR 2025 |
Machine learning › Reinforcement learning
exploration |
0.9 | 1 | 2025 | SENSEI: Semantic Exploration Guided by Foundation Models to Learn Versatile World Models · ICML 2025 |
Machine learning › Reinforcement learning
imitation learning |
0.9 | 1 | 2025 | Dream to Manipulate: Compositional World Models Empowering Robot Imitation Learning with Imagination · ICLR 2025 |
Machine learning › Reinforcement learning › exploration
intrinsic motivation |
0.9 | 1 | 2025 | SENSEI: Semantic Exploration Guided by Foundation Models to Learn Versatile World Models · ICML 2025 |
Computer vision › Image recognition and object detection
object discovery |
0.9 | 1 | 2025 | On the Transfer of Object-Centric Representation Learning · ICLR 2025 |
Machine learning › Reinforcement learning › model-based reinforcement learning
world model |
0.9 | 1 | 2025 | Dream to Manipulate: Compositional World Models Empowering Robot Imitation Learning with Imagination · ICLR 2025 |
Machine learning › Transfer learning and domain adaptation
zero-shot transfer |
0.9 | 1 | 2025 | On the Transfer of Object-Centric Representation Learning · ICLR 2025 |
Machine learning › Reinforcement learning
model-based reinforcement learning |
0.7 | 2 | 2025 | A Real-Robot Dataset for Assessing Transferability of Learned Dynamics Models · ICRA 2020 SENSEI: Semantic Exploration Guided by Foundation Models to Learn Versatile World Models · ICML 2025 |
Computer vision › Segmentation and scene understanding
scene understanding |
0.7 | 1 | 2023 | Bridging the Gap to Real-World Object-Centric Learning · ICLR 2023 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.7 | 1 | 2023 | Unsupervised Semantic Segmentation with Self-supervised Object-centric Representations · ICLR 2023 |
Computer vision › Segmentation and scene understanding › annotation-efficient segmentation
unsupervised semantic segmentation |
0.7 | 1 | 2023 | Unsupervised Semantic Segmentation with Self-supervised Object-centric Representations · ICLR 2023 |
Computer vision › Video understanding and tracking › video object segmentation
unsupervised video object segmentation |
0.7 | 1 | 2023 | Object-Centric Learning for Real-World Videos by Predicting Temporal Feature Similarities · NeurIPS 2023 |
Computer vision › Video understanding and tracking › video analytics › video object analysis › object-centric video understanding
video object discovery |
0.7 | 1 | 2023 | Object-Centric Learning for Real-World Videos by Predicting Temporal Feature Similarities · NeurIPS 2023 |
Machine learning › Reinforcement learning › deep reinforcement learning
visual reinforcement learning |
0.5 | 1 | 2021 | Self-supervised Visual Reinforcement Learning with Object-centric Representations · ICLR 2021 |
Machine learning › Reinforcement learning › model-based reinforcement learning › world model
learned dynamics models |
0.4 | 1 | 2020 | A Real-Robot Dataset for Assessing Transferability of Learned Dynamics Models · ICRA 2020 |
Robotics › Motion planning and robot control
robot control |
0.4 | 1 | 2020 | A Real-Robot Dataset for Assessing Transferability of Learned Dynamics Models · ICRA 2020 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.3 | 1 | 2025 | Temporally Consistent Object-Centric Learning by Contrasting Slots · CVPR 2025 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.3 | 1 | 2025 | CTRL-O: Language-Controllable Object-Centric Visual Representation Learning · CVPR 2025 |
Computer vision › Vision and language
vision-language model |
0.3 | 1 | 2025 | SENSEI: Semantic Exploration Guided by Foundation Models to Learn Versatile World Models · ICML 2025 |
Computer vision › Vision and language
visual question answering |
0.3 | 1 | 2025 | CTRL-O: Language-Controllable Object-Centric Visual Representation Learning · CVPR 2025 |
Methods — techniques the papers use, named apart from their topics
slot attention · 2.4self-supervised learning · 1.8model-based reinforcement learning · 1.3object-centric representation learning · 1.2temporal contrastive loss · 0.9physics simulation · 0.9language conditioning · 0.9gaussian splatting · 0.9fine-tuning · 0.9equivariant transformation · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CTRL-O: Language-Controllable Object-Centric Visual Representation LearningabstractObject-centric representation learning aims to decompose visual scenes into fixed-size vectors called "slots" or "object files", where each slot captures a distinct object. Current state-of-the-art object-centric models have shown remarkable success in object discovery in diverse domains, including complex real-world scenes. However, these models suffer from a key limitation: they lack controllability. Specifically, current object-centric models learn representations based on their preconceived understanding of objects, without allowing user input to guide which objects are represented. Introducing controllability into object-centric models could unlock a range of useful capabilities, such as the ability to extract instance-specific representations from a scene. In this work, we propose a novel approach for user-directed control over slot representations by conditioning slots on language descriptions. The proposed CONTROLLABLE OBJECT-CENTRIC REPRESENTATION LEARNING approach, which we term CTRL- O, achieves targeted object-language binding in complex real-world scenes without requiring mask supervision. Next, we apply these controllable slot representations on two downstream vision language tasks: text-to-image generation and visual question answering. The proposed approach enables instance-specific text-to-image generation and also achieves strong performance on visual question answering. Aniket Didolkar, Andrii Zadaianchuk, Rabiul Awal, Maximilian Seitzer, Efstratios Gavves, Aishwarya Agrawal |
CVPR | 2 |
| 2025 | Temporally Consistent Object-Centric Learning by Contrasting SlotsabstractUnsupervised object-centric learning from videos is a promising approach to extract structured representations from large, unlabeled collections of videos. To support downstream tasks like autonomous control, these representations must be both compositional and temporally consistent. Existing approaches based on recurrent processing often lack long-term stability across frames because their training objective does not enforce temporal consistency. In this work, we introduce a novel object-level temporal contrastive loss for video object-centric models that explicitly promotes temporal consistency. Our method significantly improves the temporal consistency of the learned object-centric representations, yielding more reliable video decompositions that facilitate challenging downstream tasks such as unsupervised object dynamics prediction. Furthermore, the inductive bias added by our loss strongly improves object discovery, leading to state-of-the-art results on both synthetic and real-world datasets, outperforming even weakly-supervised methods that leverage motion masks as additional cues. Visit slotcontrast.github.io for videos and further details. Anna Manasyan, Maximilian Seitzer, Filip Radovic, Georg Martius, Andrii Zadaianchuk |
CVPR | 5 |
| 2025 | Dream to Manipulate: Compositional World Models Empowering Robot Imitation Learning with ImaginationabstractA world model provides an agent with a representation of its environment, enabling it to predict the causal consequences of its actions. Current world models typically cannot directly and explicitly imitate the actual environment in front of a robot, often resulting in unrealistic behaviors and hallucinations that make them unsuitable for real-world robotics applications.
To overcome those challenges, we propose to rethink robot world models as learnable digital twins. We introduce DreMa, a new approach for constructing digital twins automatically using learned explicit representations of the real world and its dynamics, bridging the gap between traditional digital twins and world models.
DreMa replicates the observed world and its structure by integrating Gaussian Splatting and physics simulators, allowing robots to imagine novel configurations of objects and to predict the future consequences of robot actions thanks to its compositionality.
We leverage this capability to generate new data for imitation learning by applying equivariant transformations to a small set of demonstrations. Our evaluations across various settings demonstrate significant improvements in accuracy and robustness by incrementing actions and object distributions, reducing the data needed to learn a policy and improving the generalization of the agents.
As a highlight, we show that a real Franka Emika Panda robot, powered by DreMa’s imagination, can successfully learn novel physical tasks from just a single example per task variation (one-shot policy learning).
Our project page can be found in: https://dreamtomanipulate.github.io/. Leonardo Barcellona, Andrii Zadaianchuk, Davide Allegro, Samuele Papa, Stefano Ghidoni, Efstratios Gavves |
ICLR | 2 |
| 2025 | On the Transfer of Object-Centric Representation LearningabstractThe goal of object-centric representation learning is to decompose visual scenes into a structured representation that isolates the entities into individual vectors. Recent successes have shown that object-centric representation learning can be scaled to real-world scenes by utilizing features from pre-trained foundation models like DINO. However, so far, these object-centric methods have mostly been applied in-distribution, with models trained and evaluated on the same dataset. This is in contrast to the underlying foundation models, which have been shown to be applicable to a wide range of data and tasks. Thus, in this work, we answer the question of whether current real-world capable object-centric methods exhibit similar levels of transferability by introducing a benchmark comprising seven different synthetic and real-world datasets. We analyze the factors influencing performance under transfer and find that training on diverse real-world images improves generalization to unseen scenarios. Furthermore, inspired by the success of task-specific fine-tuning in foundation models, we introduce a novel fine-tuning strategy to adapt pre-trained vision encoders for the task of object discovery. We find that the proposed approach results in state-of-the-art performance for unsupervised object discovery, exhibiting strong zero-shot transfer to unseen datasets. Aniket Didolkar, Andrii Zadaianchuk, Anirudh Goyal, Michael C. Mozer, Yoshua Bengio, Georg Martius, Maximilian Seitzer |
ICLR | 2 |
| 2025 | SENSEI: Semantic Exploration Guided by Foundation Models to Learn Versatile World ModelsabstractExploration is a cornerstone of reinforcement learning (RL). Intrinsic motivation attempts to decouple exploration from external, task-based rewards. However, established approaches to intrinsic motivation that follow general principles such as information gain, often only uncover low-level interactions. In contrast, children’s play suggests that they engage in meaningful high-level behavior by imitating or interacting with their caregivers. Recent work has focused on using foundation models to inject these semantic biases into exploration. However, these methods often rely on unrealistic assumptions, such as language-embedded environments or access to high-level actions. We propose SEmaNtically Sensible ExploratIon (SENSEI), a framework to equip model-based RL agents with an intrinsic motivation for semantically meaningful behavior. SENSEI distills a reward signal of interestingness from Vision Language Model (VLM) annotations, enabling an agent to predict these rewards through a world model. Using model-based RL, SENSEI trains an exploration policy that jointly maximizes semantic rewards and uncertainty. We show that in both robotic and video game-like simulations SENSEI discovers a variety of meaningful behaviors from image observations and low-level actions. SENSEI provides a general tool for learning from foundation model feedback, a crucial research direction, as VLMs become more powerful. Cansu Sancaktar, Christian Gumbsch, Andrii Zadaianchuk, Pavel Kolev, Georg Martius |
ICML | 3 |
| 2023 | Bridging the Gap to Real-World Object-Centric Learning
Maximilian Seitzer, Max Horn, Andrii Zadaianchuk, Dominik Zietlow, Tianjun Xiao, Carl-Johann Simon-Gabriel, Tong He 0002, Zheng Zhang 0001, Bernhard Schölkopf, Thomas Brox, Francesco Locatello |
ICLR | 3 |
| 2023 | Unsupervised Semantic Segmentation with Self-supervised Object-centric Representations
Andrii Zadaianchuk, Matthäus Kleindessner, Francesco Locatello, Thomas Brox |
ICLR | 1 |
| 2023 | Object-Centric Learning for Real-World Videos by Predicting Temporal Feature SimilaritiesabstractUnsupervised video-based object-centric learning is a promising avenue to learn structured representations from large, unlabeled video collections, but previous approaches have only managed to scale to real-world datasets in restricted domains.
Recently, it was shown that the reconstruction of pre-trained self-supervised features leads to object-centric representations on unconstrained real-world image datasets.
Building on this approach, we propose a novel way to use such pre-trained features in the form of a temporal feature similarity loss.
This loss encodes semantic and temporal correlations between image patches and is a natural way to introduce a motion bias for object discovery.
We demonstrate that this loss leads to state-of-the-art performance on the challenging synthetic MOVi datasets.
When used in combination with the feature reconstruction loss, our model is the first object-centric video model that scales to unconstrained video datasets such as YouTube-VIS.
https://martius-lab.github.io/videosaur/ Andrii Zadaianchuk, Maximilian Seitzer, Georg Martius |
NeurIPS | 1 |
| 2021 | Self-supervised Visual Reinforcement Learning with Object-centric Representations
Andrii Zadaianchuk, Maximilian Seitzer, Georg Martius |
ICLR | 1 |
| 2020 | A Real-Robot Dataset for Assessing Transferability of Learned Dynamics ModelsabstractIn the context of model-based reinforcement learning and control, a large number of methods for learning system dynamics have been proposed in recent years. The purpose of these learned models is to synthesize new control policies. An important open question is how robust current dynamics-learning methods are to shifts in the data distribution due to changes in the control policy. We present a real-robot dataset which allows to systematically investigate this question. This dataset contains trajectories of a 3 degrees-of-freedom (DOF) robot being controlled by a diverse set of policies. For comparison, we also provide a simulated version of the dataset. Finally, we benchmark a few widely-used dynamics-learning methods using the proposed dataset. Our results show that the iid test error of a learned model is not necessarily a good indicator of its accuracy under control policies different from the one which generated the training data. This suggests that it may be important to evaluate dynamics-learning methods in terms of their transfer performance, rather than only their iid error. Diego Agudelo-España, Andrii Zadaianchuk, Philippe Wenk, Aditya Garg, Joel Akpo, Felix Grimminger, Julian Viereck, Maximilien Naveau, Ludovic Righetti, Georg Martius, Andreas Krause 0001, Bernhard Schölkopf, Stefan Bauer, Manuel Wüthrich |
ICRA | 2 |