Sara Sabour

dblp:172/1315 · also Sara Sabour Rouh Aghdam · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0001-7965-2113ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
12 papers
3D vision · 39% Video understanding and tracking · 17% Deep learning architectures and training · 17%
Computer graphics and multimedia
2 papers
Rendering · 100%

Topics — the 29 heaviest of 34, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
capsule network
1.542021
Unsupervised Part Representation by Flow Capsules · ICML 2021
Stacked Capsule Autoencoders · NeurIPS 2019
Matrix capsules with EM routing · ICLR (Poster) 2018
Computer vision › 3D vision › neural rendering
3d gaussian splatting
0.912025
SpotLessSplats: Ignoring Distractors in 3D Gaussian Splatting · ACM Trans. Graph. 2025
Computer vision › 3D vision
3d reconstruction
0.912025
SpotLessSplats: Ignoring Distractors in 3D Gaussian Splatting · ACM Trans. Graph. 2025
Computer vision › Video understanding and tracking
motion segmentation
0.912025
RoMo: Robust Motion Segmentation Improves Structure from Motion · ICCV 2025
Computer vision › Video understanding and tracking › motion segmentation
multi-frame motion segmentation
0.912025
RoMo: Robust Motion Segmentation Improves Structure from Motion · ICCV 2025
Computer vision › 3D vision
structure from motion
0.912025
RoMo: Robust Motion Segmentation Improves Structure from Motion · ICCV 2025
Rendering
neural rendering
0.912025
SpotLessSplats: Ignoring Distractors in 3D Gaussian Splatting · ACM Trans. Graph. 2025
Computer vision › 3D vision › neural radiance field
neural radiance field registration
0.712023
nerf2nerf: Pairwise Registration of Neural Radiance Fields · ICRA 2023
Rendering
neural radiance fields
0.712023
RobustNeRF: Ignoring Distractors with Robust Losses · CVPR 2023
Machine learning › Representation and self-supervised learning › representation learning
object-centric representation learning
0.612022
Conditional Object-Centric Learning from Video · ICLR 2022
Computer vision › Video understanding and tracking › video representation learning
object-centric video learning
0.612022
Conditional Object-Centric Learning from Video · ICLR 2022
Computer vision › 3D vision › motion estimation
optical flow
0.612022
Kubric: A scalable dataset generator · CVPR 2022
Machine learning › Generative modeling
synthetic data generation
0.612022
Kubric: A scalable dataset generator · CVPR 2022
Machine learning › Representation and self-supervised learning › equivariance
equivariant representation learning
0.512021
Canonical Capsules: Self-Supervised Capsules in Canonical Pose · NeurIPS 2021
Computer vision › Segmentation and scene understanding
part segmentation
0.512021
Unsupervised Part Representation by Flow Capsules · ICML 2021
Computer vision › 3D vision
point cloud processing
0.512021
Canonical Capsules: Self-Supervised Capsules in Canonical Pose · NeurIPS 2021
Computer vision › 3D vision › point cloud analysis › point cloud learning
point cloud representation learning
0.512021
Canonical Capsules: Self-Supervised Capsules in Canonical Pose · NeurIPS 2021
Machine learning › Representation and self-supervised learning › representation learning › part-based representation learning
unsupervised part discovery
0.512021
Unsupervised Part Representation by Flow Capsules · ICML 2021
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.412020
Detecting and Diagnosing Adversarial Images with Class-Conditional Capsule Reconstructions · ICLR 2020
Machine learning › Trustworthy machine learning
interpretability
0.412020
Detecting and Diagnosing Adversarial Images with Class-Conditional Capsule Reconstructions · ICLR 2020
Machine learning › Deep learning architectures and training › sequence modeling
sequence generation
0.412019
Optimal Completion Distillation for Sequence Learning · ICLR (Poster) 2019
Machine learning › Deep learning architectures and training
sequence modeling
0.412019
Optimal Completion Distillation for Sequence Learning · ICLR (Poster) 2019
Machine learning › Optimization for machine learning
robust optimization
0.312025
SpotLessSplats: Ignoring Distractors in 3D Gaussian Splatting · ACM Trans. Graph. 2025
Machine learning › Deep learning architectures and training › data-centric deep learning
training data generation
0.212022
Kubric: A scalable dataset generator · CVPR 2022
Computer vision › 3D vision
3d shape analysis
0.112021
Canonical Capsules: Self-Supervised Capsules in Canonical Pose · NeurIPS 2021
Computer vision › 3D vision
canonicalization
0.112021
Canonical Capsules: Self-Supervised Capsules in Canonical Pose · NeurIPS 2021
Computer vision › Video understanding and tracking › motion analysis
motion cues
0.112021
Unsupervised Part Representation by Flow Capsules · ICML 2021
Computer vision › Image recognition and object detection › image classification › label-efficient image classification
unsupervised image classification
0.112019
Stacked Capsule Autoencoders · NeurIPS 2019
Computer vision › Image recognition and object detection › character recognition
handwritten digit recognition
0.112017
Dynamic Routing Between Capsules · NIPS 2017

Methods — techniques the papers use, named apart from their topics

robust optimization · 2.4pre-trained general-purpose features · 1.7capsule network · 1.0video segmentation model · 0.9optical flow · 0.9epipolar geometry · 0.9robust loss · 0.7outlier modeling · 0.7iterative closest point · 0.7slot attention · 0.6contrastive learning · 0.6motion-based learning · 0.5
YearPublicationVenuePosition
2025 RoMo: Robust Motion Segmentation Improves Structure from Motion
abstract
There has been extensive progress in the reconstruction and generation of 4D scenes from monocular casually-captured video. While these tasks rely heavily on known camera poses, the problem of finding such poses using structure-from-motion (SfM) often depends on robustly separating static from dynamic parts of a video. The lack of a robust solution to this problem limits the performance of SfM camera-calibration pipelines. We propose a novel approach to video-based motion segmentation to identify the components of a scene that are moving w.r.t. a fixed world frame. Our simple but effective iterative method, RoMo, combines optical flow and epipolar cues with a pre-trained video segmentation model. It outperforms unsupervised baselines for motion segmentation as well as supervised baselines trained from synthetic data. More importantly, the combination of an off-the-shelf SfM pipeline with our segmentation masks establishes a new state-of-the-art on camera calibration for scenes with dynamic content, outperforming existing methods by a substantial margin.
Lily Goli, Sara Sabour, Mark J. Matthews, Marcus A. Brubaker, Dmitry Lagun, Alec Jacobson, David J. Fleet, Saurabh Saxena, Andrea Tagliasacchi
ICCV2
2025 SpotLessSplats: Ignoring Distractors in 3D Gaussian Splatting
abstract
Three-dimensional Gaussian Splatting (3DGS) is a promising technique for 3D reconstruction, offering efficient training and rendering speeds, making it suitable for real-time applications. However, current methods require highly controlled environments–no moving people or wind-blown elements, and consistent lighting–to meet the interview consistency assumption of 3DGS. This makes reconstruction of real-world captures problematic. We present SpotLessSplats, an approach that leverages pre-trained and general-purpose features coupled with robust optimization to effectively ignore transient distractors. Our method achieves state-of-the-art reconstruction quality both visually and quantitatively, on casual captures.
Sara Sabour, Lily Goli, Georgios Kopanas, Mark J. Matthews, Dmitry Lagun, Leonidas J. Guibas, Alec Jacobson, David J. Fleet, Andrea Tagliasacchi
ACM Trans. Graph.1
2023 RobustNeRF: Ignoring Distractors with Robust Losses
abstract
Neural radiance fields (NeRF) excel at synthesizing new views given multi-view, calibrated images of a static scene. When scenes include distractors, which are not persistent during image capture (moving objects, lighting variations, shadows), artifacts appear as view-dependent effects or ‘floaters’. To cope with distractors, we advocate a form of robust estimation for NeRF training, modeling distractors in training data as outliers of an optimization problem. Our method successfully removes outliers from a scene and improves upon our baselines, on synthetic and real-world scenes. Our technique is simple to incorporate in modern NeRF frameworks, with few hyper-parameters. It does not assume a priori knowledge of the types of distractors, and is instead focused on the optimization problem rather than pre-processing or modeling transient objects. More results at https://robustnerf.github.io/public.
Sara Sabour, Suhani Vora, Daniel Duckworth, Ivan Krasin, David J. Fleet, Andrea Tagliasacchi
CVPR1
2023 nerf2nerf: Pairwise Registration of Neural Radiance Fields
abstract
We introduce a technique for pairwise registration of neural fields that extends classical optimization-based local registration (i.e. ICP) to operate on Neural Radiance Fields (NeRF)-neural 3D scene representations trained from collections of calibrated images. NeRF does not decompose illumination and color, so to make registration invariant to illumination, we introduce the concept of a “surface field” - a field distilled from a pre-trained NeRF model that measures the likelihood of a point being on the surface of an object. We then cast nerf2nerf registration as a robust optimization that iteratively seeks a rigid transformation that aligns the surface fields of the two scenes. We evaluate the effectiveness of our technique by introducing a dataset of pre-trained NeRF scenes - our synthetic scenes enable quantitative evaluations and comparisons to classical registration techniques, while our real scenes demonstrate the validity of our technique in real-world scenarios. Additional results available at: https://nerf2nerf.github.io
Lily Goli, Daniel Rebain, Sara Sabour, Animesh Garg, Andrea Tagliasacchi
ICRA3
2022 Kubric: A scalable dataset generator
abstract
Data is the driving force of machine learning, with the amount and quality of training data often being more important for the performance of a system than architecture and training details. But collecting, processing and annotating real data at scale is difficult, expensive, and frequently raises additional privacy, fairness and legal concerns. Synthetic data is a powerful tool with the potential to address these shortcomings: 1) it is cheap 2) supports rich ground-truth annotations 3) offers full control over data and 4) can circumvent or mitigate problems regarding bias, privacy and licensing. Unfortunately, software tools for effective data generation are less mature than those for architecture design and training, which leads to fragmented generation efforts. To address these problems we introduce Kubric, an open-source Python framework that interfaces with PyBullet and Blender to generate photo-realistic scenes, with rich annotations, and seamlessly scales to large jobs distributed over thousands of machines, and generating TBs of data. We demonstrate the effectiveness of Kubric by presenting a series of 13 different generated datasets for tasks ranging from studying 3D NeRF models to optical flow estimation. We release Kubric, the used assets, all of the generation code, as well as the rendered datasets for reuse and modification.
Klaus Greff, Francois Belletti, Lucas Beyer, Carl Doersch, Yilun Du, Daniel Duckworth, David J. Fleet, Dan Gnanapragasam, Florian Golemo, Charles Herrmann, Thomas Kipf, Abhijit Kundu, Dmitry Lagun, Issam H. Laradji, Hsueh-Ti Derek Liu, Henning Meyer, Yishu Miao, Derek Nowrouzezahrai, A. Cengiz Öztireli, Etienne Pot, Noha Radwan, Daniel Rebain, Sara Sabour, Mehdi S. M. Sajjadi, Matan Sela, Vincent Sitzmann, Austin Stone, Deqing Sun, Suhani Vora, Tianhao Wu 0003, Kwang Moo Yi, Fangcheng Zhong, Andrea Tagliasacchi
CVPR23
2022 Conditional Object-Centric Learning from Video
Thomas Kipf, Gamaleldin F. Elsayed, Aravindh Mahendran, Austin Stone, Sara Sabour, Georg Heigold, Rico Jonschkowski, Alexey Dosovitskiy, Klaus Greff
ICLR5
2021 Unsupervised Part Representation by Flow Capsules
abstract
Capsule networks aim to parse images into a hierarchy of objects, parts and relations. While promising, they remain limited by an inability to learn effective low level part descriptions. To address this issue we propose a way to learn primary capsule encoders that detect atomic parts from a single image. During training we exploit motion as a powerful perceptual cue for part definition, with an expressive decoder for part generation within a layered image model with occlusion. Experiments demonstrate robust part discovery in the presence of multiple objects, cluttered backgrounds, and occlusion. The learned part decoder is shown to infer the underlying shape masks, effectively filling in occluded regions of the detected shapes. We evaluate FlowCapsules on unsupervised part segmentation and unsupervised image classification.
Sara Sabour, Andrea Tagliasacchi, Soroosh Yazdani, Geoffrey E. Hinton, David J. Fleet
ICML1
2021 Canonical Capsules: Self-Supervised Capsules in Canonical Pose
abstract
We propose a self-supervised capsule architecture for 3D point clouds. We compute capsule decompositions of objects through permutation-equivariant attention, and self-supervise the process by training with pairs of randomly rotated objects. Our key idea is to aggregate the attention masks into semantic keypoints, and use these to supervise a decomposition that satisfies the capsule invariance/equivariance properties. This not only enables the training of a semantically consistent decomposition, but also allows us to learn a canonicalization operation that enables object-centric reasoning. To train our neural network we require neither classification labels nor manually-aligned training datasets. Yet, by learning an object-centric representation in a self-supervised manner, our method outperforms the state-of-the-art on 3D point cloud reconstruction, canonicalization, and unsupervised classification.
Andrea Tagliasacchi, Boyang Deng, Sara Sabour, Soroosh Yazdani, Geoffrey E. Hinton, Kwang Moo Yi
NeurIPS4
2020 Detecting and Diagnosing Adversarial Images with Class-Conditional Capsule Reconstructions
Yao Qin 0001, Nicholas Frosst, Sara Sabour, Colin Raffel, Garrison W. Cottrell, Geoffrey E. Hinton
ICLR3
2019 Optimal Completion Distillation for Sequence Learning
Sara Sabour, Mohammad Norouzi 0002
ICLR (Poster)1
2019 Stacked Capsule Autoencoders
abstract
Objects are composed of a set of geometrically organized parts. We introduce an unsupervised capsule autoencoder (SCAE), which explicitly uses geometric relationships between parts to reason about objects. Since these relationships do not depend on the viewpoint, our model is robust to viewpoint changes. SCAE consists of two stages. In the first stage, the model predicts presences and poses of part templates directly from the image and tries to reconstruct the image by appropriately arranging the templates. In the second stage, the SCAE predicts parameters of a few object capsules, which are then used to reconstruct part poses. Inference in this model is amortized and performed by off-the-shelf neural encoders, unlike in previous capsule networks. We find that object capsule presences are highly informative of the object class, which leads to state-of-the-art results for unsupervised classification on SVHN (55%) and MNIST (98.7%).
Adam R. Kosiorek, Sara Sabour, Yee Whye Teh, Geoffrey E. Hinton
NeurIPS2
2018 Matrix capsules with EM routing
Geoffrey E. Hinton, Sara Sabour, Nicholas Frosst
ICLR (Poster)2
2017 Dynamic Routing Between Capsules
abstract
A capsule is a group of neurons whose activity vector represents the instantiation parameters of a specific type of entity such as an object or object part. We use the length of the activity vector to represent the probability that the entity exists and its orientation to represent the instantiation parameters. Active capsules at one level make predictions, via transformation matrices, for the instantiation parameters of higher-level capsules. When multiple predictions agree, a higher level capsule becomes active. We show that a discrimininatively trained, multi-layer capsule system achieves state-of-the-art performance on MNIST and is considerably better than a convolutional net at recognizing highly overlapping digits. To achieve these results we use an iterative routing-by-agreement mechanism: A lower-level capsule prefers to send its output to higher level capsules whose activity vectors have a big scalar product with the prediction coming from the lower-level capsule.
Sara Sabour, Nicholas Frosst, Geoffrey E. Hinton
NIPS1