EDBT 2026 Demo / reviewers in the wild / expert
S. M. Ali Eslami
dblp:117/4847
· DBLP profile ↗
21ranked-venue papers
5as first author
4since 2021 · last 2023
0000-0003-1838-5589ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 5 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
18 papers |
Probabilistic and Bayesian machine learning · 20% 3D vision · 16% Transfer learning and domain adaptation · 13% | |
| Computer graphics and multimedia
2 papers |
Geometric modeling and processing · 72% Visual content generation and editing · 28% |
Topics — the 30 heaviest of 48, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
neural processes |
0.7 | 2 | 2019 | Attentive Neural Processes · ICLR (Poster) 2019 Conditional Neural Processes · ICML 2018 |
Machine learning › Transfer learning and domain adaptation
few-shot learning |
0.7 | 2 | 2018 | Machine Theory of Mind · ICML 2018 Few-shot Autoregressive Density Estimation: Towards Learning to Learn Distributions · ICLR (Poster) 2018 |
Machine learning › Transfer learning and domain adaptation
meta-learning |
0.7 | 2 | 2018 | Machine Theory of Mind · ICML 2018 Few-shot Autoregressive Density Estimation: Towards Learning to Learn Distributions · ICLR (Poster) 2018 |
Machine learning › Representation and self-supervised learning › pre-training › spatio-temporal pre-training
video pretraining |
0.7 | 1 | 2023 | Self-supervised video pretraining yields robust and more human-aligned visual representations · NeurIPS 2023 |
Computer vision › 3D vision
implicit neural representation |
0.6 | 1 | 2022 | From data to functa: Your data point is a function and you can treat it like one · ICML 2022 |
Computer vision › 3D vision › implicit neural representation
neural field generation |
0.6 | 1 | 2022 | From data to functa: Your data point is a function and you can treat it like one · ICML 2022 |
Machine learning › Transfer learning and domain adaptation › few-shot learning
cross-modal few-shot learning |
0.5 | 1 | 2021 | Multimodal Few-Shot Learning with Frozen Language Models · NeurIPS 2021 |
Natural language and speech › Language models and text generation › pre-trained language model › efficient pre-trained language model
frozen language model |
0.5 | 1 | 2021 | Multimodal Few-Shot Learning with Frozen Language Models · NeurIPS 2021 |
Machine learning › Probabilistic and Bayesian machine learning
boltzmann machine |
0.5 | 3 | 2014 | The Shape Boltzmann Machine: A Strong Model of Object Shape · Int. J. Comput. Vis. 2014 A Generative Model for Parts-based Object Segmentation · NIPS 2012 The Shape Boltzmann Machine: A strong model of object shape · CVPR 2012 |
Machine learning › Generative modeling
autoregressive model |
0.4 | 1 | 2020 | PolyGen: An Autoregressive Generative Model of 3D Meshes · ICML 2020 |
Machine learning › Generative modeling › 3d generative model
mesh generative model |
0.4 | 1 | 2020 | PolyGen: An Autoregressive Generative Model of 3D Meshes · ICML 2020 |
Geometric modeling and processing › mesh generation
autoregressive mesh generation |
0.4 | 1 | 2020 | PolyGen: An Autoregressive Generative Model of 3D Meshes · ICML 2020 |
Geometric modeling and processing
mesh generation |
0.4 | 1 | 2020 | PolyGen: An Autoregressive Generative Model of 3D Meshes · ICML 2020 |
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training |
0.3 | 1 | 2018 | Synthesizing Programs for Images using Reinforced Adversarial Learning · ICML 2018 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process › neural processes
conditional neural process |
0.3 | 1 | 2018 | Conditional Neural Processes · ICML 2018 |
Machine learning › Generative modeling › variational autoencoder
conditional variational autoencoder |
0.3 | 1 | 2018 | A Probabilistic U-Net for Segmentation of Ambiguous Images · NeurIPS 2018 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
density estimation |
0.3 | 1 | 2018 | Few-shot Autoregressive Density Estimation: Towards Learning to Learn Distributions · ICLR (Poster) 2018 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process |
0.3 | 1 | 2018 | Conditional Neural Processes · ICML 2018 |
Computer vision › 3D vision
inverse rendering |
0.3 | 1 | 2018 | Synthesizing Programs for Images using Reinforced Adversarial Learning · ICML 2018 |
Machine learning › Representation and self-supervised learning › representation learning › latent representation learning › state representation learning
latent dynamics model |
0.3 | 1 | 2018 | Generative Temporal Models with Spatial Memory for Partially Observed Environments · ICML 2018 |
Machine learning › Reinforcement learning
model-based reinforcement learning |
0.3 | 1 | 2018 | Generative Temporal Models with Spatial Memory for Partially Observed Environments · ICML 2018 |
Computer vision › Segmentation and scene understanding › image segmentation
probabilistic segmentation |
0.3 | 1 | 2018 | A Probabilistic U-Net for Segmentation of Ambiguous Images · NeurIPS 2018 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.3 | 1 | 2018 | A Probabilistic U-Net for Segmentation of Ambiguous Images · NeurIPS 2018 |
Robotics › Robot navigation and mapping › spatial representation
spatial memory |
0.3 | 1 | 2018 | Generative Temporal Models with Spatial Memory for Partially Observed Environments · ICML 2018 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
theory of mind |
0.3 | 1 | 2018 | Machine Theory of Mind · ICML 2018 |
Machine learning › Generative modeling
variational autoencoder |
0.3 | 1 | 2018 | Generative Temporal Models with Spatial Memory for Partially Observed Environments · ICML 2018 |
Computer vision › Video understanding and tracking
video prediction |
0.3 | 1 | 2018 | Generative Temporal Models with Spatial Memory for Partially Observed Environments · ICML 2018 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › deep latent variable model
object-centric generative model |
0.2 | 1 | 2016 | Attend, Infer, Repeat: Fast Scene Understanding with Generative Models · NIPS 2016 |
Machine learning › Probabilistic and Bayesian machine learning
probabilistic inference |
0.2 | 1 | 2016 | Unsupervised Learning of 3D Structure from Images · NIPS 2016 |
Computer vision › Segmentation and scene understanding
scene understanding |
0.2 | 1 | 2016 | Attend, Infer, Repeat: Fast Scene Understanding with Generative Models · NIPS 2016 |
Methods — techniques the papers use, named apart from their topics
transformer · 0.9autoregressive modeling · 0.9contrastive learning · 0.7probabilistic inference · 0.6neural field · 0.6implicit neural representation · 0.6autoencoder · 0.6vision encoder · 0.5few-shot prompting · 0.5attention · 0.4u-net · 0.3reinforcement learning · 0.3generative segmentation · 0.3discriminator reward · 0.3conditional variational autoencoder · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Self-supervised video pretraining yields robust and more human-aligned visual representationsabstractHumans learn powerful representations of objects and scenes by observing how they evolve over time. Yet, outside of specific tasks that require explicit temporal understanding, static image pretraining remains the dominant paradigm for learning visual foundation models. We question this mismatch, and ask whether video pretraining can yield visual representations that bear the hallmarks of human perception: generalisation across tasks, robustness to perturbations, and consistency with human judgements. To that end we propose a novel procedure for curating videos, and develop a contrastive framework which learns from the complex transformations therein. This simple paradigm for distilling knowledge from videos, called VITO, yields general representations that far outperform prior video pretraining methods on image understanding tasks, and image pretraining methods on video understanding tasks. Moreover, VITO representations are significantly more robust to natural and synthetic deformations than image-, video-, and adversarially-trained
ones. Finally, VITO’s predictions are strongly aligned with human judgements, surpassing models that were specifically trained for that purpose. Together, these results suggest that video pretraining could be a simple way of learning unified, robust, and human-aligned representations of the visual world. Nikhil Parthasarathy, S. M. Ali Eslami, João Carreira 0001, Olivier J. Hénaff |
NeurIPS | 2 |
| 2022 | From data to functa: Your data point is a function and you can treat it like oneabstractIt is common practice in deep learning to represent a measurement of the world on a discrete grid, e.g. a 2D grid of pixels. However, the underlying signal represented by these measurements is often continuous, e.g. the scene depicted in an image. A powerful continuous alternative is then to represent these measurements using an implicit neural representation, a neural function trained to output the appropriate measurement value for any input spatial location. In this paper, we take this idea to its next level: what would it take to perform deep learning on these functions instead, treating them as data? In this context we refer to the data as functa, and propose a framework for deep learning on functa. This view presents a number of challenges around efficient conversion from data to functa, compact representation of functa, and effectively solving downstream tasks on functa. We outline a recipe to overcome these challenges and apply it to a wide range of data modalities including images, 3D shapes, neural radiance fields (NeRF) and data on manifolds. We demonstrate that this approach has various compelling properties across data modalities, in particular on the canonical tasks of generative modeling, data imputation, novel view synthesis and classification. Emilien Dupont, Hyunjik Kim, S. M. Ali Eslami, Danilo Jimenez Rezende, Dan Rosenbaum |
ICML | 3 |
| 2021 | Multimodal Few-Shot Learning with Frozen Language ModelsabstractWhen trained at sufficient scale, auto-regressive language models exhibit the notable ability to learn a new language task after being prompted with just a few examples. Here, we present a simple, yet effective, approach for transferring this few-shot learning ability to a multimodal setting (vision and language). Using aligned image and caption data, we train a vision encoder to represent each image as a sequence of continuous embeddings, such that a pre-trained, frozen language model presented with this prefix generates the appropriate caption. The resulting system is a multimodal few-shot learner, with the surprising ability to learn a variety of new tasks when conditioned on examples, represented as a sequence of any number of interleaved image and text embeddings. We demonstrate that it can rapidly learn words for new objects and novel visual categories, do visual question-answering with only a handful of examples, and make use of outside knowledge, by measuring a single model on a variety of established and new benchmarks. Maria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami, Oriol Vinyals, Felix Hill |
NeurIPS | 4 |
| 2021 | Game Plan: What AI can do for Football, and What Football can do for AIabstractThe rapid progress in artificial intelligence (AI) and machine learning has opened unprecedented analytics possibilities in various team and individual sports, including baseball, basketball, and tennis. More recently, AI techniques have been applied to football, due to a huge increase in data collection by professional teams, increased computational power, and advances in machine learning, with the goal of better addressing new scientific challenges involved in the analysis of both individual players’ and coordinated teams’ behaviors. The research challenges associated with predictive and prescriptive football analytics require new developments and progress at the intersection of statistical learning, game theory, and computer vision. In this paper, we provide an overarching perspective highlighting how the combination of these fields, in particular, forms a unique microcosm for AI research, while offering mutual benefits for professional teams, spectators, and broadcasters in the years to come. We illustrate that this duality makes football analytics a game changer of tremendous value, in terms of not only changing the game of football itself, but also in terms of what this domain can mean for the field of AI. We review the state-of-the-art and exemplify the types of analysis enabled by combining the aforementioned fields, including illustrative examples of counterfactual analysis using predictive models, and the combination of game-theoretic analysis of penalty kicks with statistical learning of player attributes. We conclude by highlighting envisioned downstream impacts, including possibilities for extensions to other sports (real and virtual). Karl Tuyls, Shayegan Omidshafiei, Paul Muller, Zhe Wang 0055, Jerome T. Connor, Daniel Hennes, Ian Graham, William Spearman, Tim Waskett, Dafydd Steele, Pauline Luc, Adrià Recasens, Alexandre Galashov, Gregory Thornton, Romuald Elie, Pablo Sprechmann, Pol Moreno, Kris Cao, Marta Garnelo, Praneet Dutta, Michal Valko, Nicolas Heess, Alex Bridgland, Julien Pérolat, Bart De Vylder, S. M. Ali Eslami, Mark Rowland 0001, Andrew Jaegle, Rémi Munos, Trevor Back, Razia Ahamed, Simon Bouton, Nathalie Beauguerlange, Jackson Broshear, Thore Graepel, Demis Hassabis |
J. Artif. Intell. Res. | 26 |
| 2020 | PolyGen: An Autoregressive Generative Model of 3D MeshesabstractPolygon meshes are an efficient representation of 3D geometry, and are of central importance in computer graphics, robotics and games development. Existing learning-based approaches for object synthesis have avoided the challenges of working with 3D meshes, instead using alternative object representations that are more compatible with neural architectures and training approaches. We present PolyGen, a generative model of 3D objects which models the mesh directly, predicting vertices and faces sequentially using a Transformer-based architecture. Our model can condition on a range of inputs, including object classes, voxels, and images, and because the model is probabilistic it can produce samples that capture uncertainty in ambiguous scenarios. We show that the model is capable of producing high-quality, usable meshes, and establish log-likelihood benchmarks for the mesh-modelling task. We also evaluate the conditional models on surface reconstruction metrics against alternative methods, and demonstrate competitive performance despite not training directly on this task. Charlie Nash, Yaroslav Ganin, S. M. Ali Eslami, Peter W. Battaglia |
ICML | 3 |
| 2019 | Attentive Neural Processes
Hyunjik Kim, Andriy Mnih, Jonathan Schwarz, Marta Garnelo, S. M. Ali Eslami, Dan Rosenbaum, Oriol Vinyals, Yee Whye Teh |
ICLR (Poster) | 5 |
| 2018 | Few-shot Autoregressive Density Estimation: Towards Learning to Learn Distributions
Scott E. Reed, Yutian Chen 0001, Thomas Paine, Aäron van den Oord, S. M. Ali Eslami, Danilo Jimenez Rezende, Oriol Vinyals, Nando de Freitas |
ICLR (Poster) | 5 |
| 2018 | Generative Temporal Models with Spatial Memory for Partially Observed EnvironmentsabstractIn model-based reinforcement learning, generative and temporal models of environments can be leveraged to boost agent performance, either by tuning the agent’s representations during training or via use as part of an explicit planning mechanism. However, their application in practice has been limited to simplistic environments, due to the difficulty of training such models in larger, potentially partially-observed and 3D environments. In this work we introduce a novel action-conditioned generative model of such challenging environments. The model features a non-parametric spatial memory system in which we store learned, disentangled representations of the environment. Low-dimensional spatial updates are computed using a state-space model that makes use of knowledge on the prior dynamics of the moving agent, and high-dimensional visual observations are modelled with a Variational Auto-Encoder. The result is a scalable architecture capable of performing coherent predictions over hundreds of time steps across a range of partially observed 2D and 3D environments. Marco Fraccaro, Danilo Jimenez Rezende, Yori Zwols, Alexander Pritzel, S. M. Ali Eslami, Fabio Viola |
ICML | 5 |
| 2018 | Synthesizing Programs for Images using Reinforced Adversarial LearningabstractAdvances in deep generative networks have led to impressive results in recent years. Nevertheless, such models can often waste their capacity on the minutiae of datasets, presumably due to weak inductive biases in their decoders. This is where graphics engines may come in handy since they abstract away low-level details and represent images as high-level programs. Current methods that combine deep learning and renderers are limited by hand-crafted likelihood or distance functions, a need for large amounts of supervision, or difficulties in scaling their inference algorithms to richer datasets. To mitigate these issues, we present SPIRAL, an adversarially trained agent that generates a program which is executed by a graphics engine to interpret and sample images. The goal of this agent is to fool a discriminator network that distinguishes between real and rendered data, trained with a distributed reinforcement learning setup without any supervision. A surprising finding is that using the discriminator’s output as a reward signal is the key to allow the agent to make meaningful progress at matching the desired output rendering. To the best of our knowledge, this is the first demonstration of an end-to-end, unsupervised and adversarial inverse graphics agent on challenging real world (MNIST, Omniglot, CelebA) and synthetic 3D datasets. A video of the agent can be found at https://youtu.be/iSyvwAwa7vk. Yaroslav Ganin, Tejas Kulkarni, Igor Babuschkin, S. M. Ali Eslami, Oriol Vinyals |
ICML | 4 |
| 2018 | Conditional Neural ProcessesabstractDeep neural networks excel at function approximation, yet they are typically trained from scratch for each new function. On the other hand, Bayesian methods, such as Gaussian Processes (GPs), exploit prior knowledge to quickly infer the shape of a new function at test time. Yet, GPs are computationally expensive, and it can be hard to design appropriate priors. In this paper we propose a family of neural models, Conditional Neural Processes (CNPs), that combine the benefits of both. CNPs are inspired by the flexibility of stochastic processes such as GPs, but are structured as neural networks and trained via gradient descent. CNPs make accurate predictions after observing only a handful of training data points, yet scale to complex functions and large datasets. We demonstrate the performance and versatility of the approach on a range of canonical machine learning tasks, including regression, classification and image completion. Marta Garnelo, Dan Rosenbaum, Chris J. Maddison, Tiago Ramalho, David Saxton, Murray Shanahan, Yee Whye Teh, Danilo Jimenez Rezende, S. M. Ali Eslami |
ICML | 9 |
| 2018 | Machine Theory of MindabstractTheory of mind (ToM) broadly refers to humans’ ability to represent the mental states of others, including their desires, beliefs, and intentions. We design a Theory of Mind neural network {–} a ToMnet {–} which uses meta-learning to build such models of the agents it encounters. The ToMnet learns a strong prior model for agents’ future behaviour, and, using only a small number of behavioural observations, can bootstrap to richer predictions about agents’ characteristics and mental states. We apply the ToMnet to agents behaving in simple gridworld environments, showing that it learns to model random, algorithmic, and deep RL agents from varied populations, and that it passes classic ToM tasks such as the "Sally-Anne" test of recognising that others can hold false beliefs about the world. Neil C. Rabinowitz, Frank Perbet, H. Francis Song, Chiyuan Zhang, S. M. Ali Eslami, Matt M. Botvinick |
ICML | 5 |
| 2018 | A Probabilistic U-Net for Segmentation of Ambiguous ImagesabstractMany real-world vision problems suffer from inherent ambiguities. In clinical applications for example, it might not be clear from a CT scan alone which particular region is cancer tissue. Therefore a group of graders typically produces a set of diverse but plausible segmentations. We consider the task of learning a distribution over segmentations given an input. To this end we propose a generative segmentation model based on a combination of a U-Net with a conditional variational autoencoder that is capable of efficiently producing an unlimited number of plausible hypotheses. We show on a lung abnormalities segmentation task and on a Cityscapes segmentation task that our model reproduces the possible segmentation variants as well as the frequencies with which they occur, doing so significantly better than published approaches. These models could have a high impact in real-world applications, such as being used as clinical decision-making algorithms accounting for multiple plausible semantic segmentation hypotheses to provide possible diagnoses and recommend further actions to resolve the present ambiguities. Simon Kohl, Bernardino Romera-Paredes, Clemens Meyer, Jeffrey De Fauw, Joseph R. Ledsam, Klaus H. Maier-Hein, S. M. Ali Eslami, Danilo Jimenez Rezende, Olaf Ronneberger |
NeurIPS | 7 |
| 2016 | Attend, Infer, Repeat: Fast Scene Understanding with Generative ModelsabstractWe present a framework for efficient inference in structured image models that explicitly reason about objects. We achieve this by performing probabilistic inference using a recurrent neural network that attends to scene elements and processes them one at a time. Crucially, the model itself learns to choose the appropriate number of inference steps. We use this scheme to learn to perform inference in partially specified 2D models (variable-sized variational auto-encoders) and fully specified 3D models (probabilistic renderers). We show that such models learn to identify multiple objects - counting, locating and classifying the elements of a scene - without any supervision, e.g., decomposing 3D images with various numbers of objects in a single forward pass of a neural network at unprecedented speed. We further show that the networks produce accurate inferences when compared to supervised counterparts, and that their structure leads to improved generalization. S. M. Ali Eslami, Nicolas Heess, Theophane Weber, Yuval Tassa, David Szepesvari, Koray Kavukcuoglu, Geoffrey E. Hinton |
NIPS | 1 |
| 2016 | Unsupervised Learning of 3D Structure from ImagesabstractA key goal of computer vision is to recover the underlying 3D structure that gives rise to 2D observations of the world. If endowed with 3D understanding, agents can abstract away from the complexity of the rendering process to form stable, disentangled representations of scene elements. In this paper we learn strong deep generative models of 3D structures, and recover these structures from 2D images via probabilistic inference. We demonstrate high-quality samples and report log-likelihoods on several datasets, including ShapeNet, and establish the first benchmarks in the literature. We also show how these models and their inference networks can be trained jointly, end-to-end, and directly from 2D images without any use of ground-truth 3D labels. This demonstrates for the first time the feasibility of learning to infer 3D representations of the world in a purely unsupervised manner. Danilo Jimenez Rezende, S. M. Ali Eslami, Shakir Mohamed, Peter W. Battaglia, Max Jaderberg, Nicolas Heess |
NIPS | 2 |
| 2015 | Consensus Message Passing for Layered Graphical ModelsabstractGenerative models provide a powerful framework for probabilistic reasoning. However, in many domains their use has been hampered by the practical difficulties of inference. This is particularly the case in computer vision, where models of the imaging process tend to be large, loopy and layered. For this reason bottom-up conditional models have traditionally dominated in such domains. We find that widely-used, general-purpose message passing inference algorithms such as Expectation Propagation (EP) and Variational Message Passing (VMP) fail on the simplest of vision models. With these models in mind, we introduce a modification to message passing that learns to exploit their layered structure by passing ’consensus’ messages that guide inference towards good solutions. Experiments on a variety of problems show that the proposed technique leads to significantly more accurate inference results, not only when compared to standard EP and VMP, but also when compared to competitive bottom-up conditional models. Varun Jampani, S. M. Ali Eslami, Daniel Tarlow, Pushmeet Kohli, John M. Winn |
AISTATS | 2 |
| 2015 | Kernel-Based Just-In-Time Learning for Passing Expectation Propagation Messages
Wittawat Jitkrittum, Arthur Gretton, Nicolas Heess, S. M. Ali Eslami, Balaji Lakshminarayanan, Dino Sejdinovic, Zoltán Szabó 0001 |
UAI | 4 |
| 2015 | The Pascal Visual Object Classes Challenge: A Retrospective
Mark Everingham, S. M. Ali Eslami, Luc Van Gool, Christopher K. I. Williams, John M. Winn, Andrew Zisserman |
Int. J. Comput. Vis. | 2 |
| 2014 | Just-In-Time Learning for Fast and Flexible Inference
S. M. Ali Eslami, Daniel Tarlow, Pushmeet Kohli, John M. Winn |
NIPS | 1 |
| 2014 | The Shape Boltzmann Machine: A Strong Model of Object Shape
S. M. Ali Eslami, Nicolas Heess, Christopher K. I. Williams, John M. Winn |
Int. J. Comput. Vis. | 1 |
| 2012 | The Shape Boltzmann Machine: A strong model of object shapeabstractA good model of object shape is essential in applications such as segmentation, object detection, inpainting and graphics. For example, when performing segmentation, local constraints on the shape can help where the object boundary is noisy or unclear, and global constraints can resolve ambiguities where background clutter looks similar to part of the object. In general, the stronger the model of shape, the more performance is improved. In this paper, we use a type of Deep Boltzmann Machine [22] that we call a Shape Boltzmann Machine (ShapeBM) for the task of modeling binary shape images. We show that the ShapeBM characterizes a strong model of shape, in that samples from the model look realistic and it can generalize to generate samples that differ from training examples. We find that the ShapeBM learns distributions that are qualitatively and quantitatively better than existing models for this task. S. M. Ali Eslami, Nicolas Heess, John M. Winn |
CVPR | 1 |
| 2012 | A Generative Model for Parts-based Object SegmentationabstractThe Shape Boltzmann Machine (SBM) has recently been introduced as a state-of-the-art model of foreground/background object shape. We extend the SBM to account for the foreground object's parts. Our model, the Multinomial SBM (MSBM), can capture both local and global statistics of part shapes accurately. We combine the MSBM with an appearance model to form a fully generative model of images of objects. Parts-based image segmentations are obtained simply by performing probabilistic inference in the model. We apply the model to two challenging datasets which exhibit significant shape and appearance variability, and find that it obtains results that are comparable to the state-of-the-art. S. M. Ali Eslami, Christopher K. I. Williams |
NIPS | 1 |