EDBT 2026 Demo / reviewers in the wild / expert
Silvia Zuffi
dblp:14/1794
· DBLP profile ↗
24ranked-venue papers
9as first author
10since 2021 · last 2026
0000-0003-1358-0828ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 8 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 8 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Model-based Metric 3D Shape and Motion Reconstruction of Wild Bottlenose Dolphins in Drone-Shot Videos
Daniele Baieri, Riccardo Cicciarella, Michael Krützen, Emanuele Rodolà, Silvia Zuffi |
Int. J. Comput. Vis. | 5 |
| 2025 | Reconstructing Animals and the WildabstractThe notion of 3D reconstruction as scene understanding is foundational in computer vision. Reconstructing 3D scenes from 2D visual observations necessitates strong priors to disambiguate structure. Much work has been focused on the anthropocentric, which, characterized by smooth surfaces, coherent normals, and regular edges, allows for the integration of strong geometric inductive biases. Here, we consider a more challenging problem where such assumptions do not hold: the reconstruction of natural scenes containing trees, bushes, boulders, and animals. While numerous works have attempted to tackle the problem of reconstructing animals in the wild, they have focused solely on the animal, neglecting environmental context. This limits their usefulness for analysis tasks, as animals exist inherently within the 3D world, and information is lost when environmental factors are disregarded. We propose a method to reconstruct natural scenes from single images. We base our approach on recent advances leveraging the strong world priors ingrained in Large Language Models and train an autoregressive model to decode a CLIP embedding into a structured, compositional scene representation, encompassing both animals and the wild. To enable this, we propose a synthetic dataset comprising one million images and thousands of assets. Through our introduction of a CLIP-projection head, we demonstrate that our approach generalizes to the task of reconstructing animals and their environments in real-world images, despite having been trained solely on synthetic data. We release our dataset and code to encourage future research at https://raw.is.tue.mpg.de/. Peter Kulits, Michael J. Black, Silvia Zuffi |
CVPR | 3 |
| 2025 | Generative Zoo
Tomasz Niewiadomski, Anastasios Yiannakidis, Hanz Cuevas-Velasquez, Soubhik Sanyal, Michael J. Black, Silvia Zuffi, Peter Kulits |
ICCV | 6 |
| 2024 | Dessie: Disentanglement for Articulated 3D Horse Shape and Pose Estimation from Images
Ci Li, Yi Yang 0095, Zehang Weng, Elin Hernlund, Silvia Zuffi, Hedvig Kjellström |
ACCV (10) | 5 |
| 2024 | VAREN: Very Accurate and Realistic Equine NetworkabstractData-driven three-dimensional parametric shape mod-els of the human body have gained enormous popularity both for the analysis of visual data and for the generation of synthetic humans. Following a similar approach for animals does not scale to the multitude of existing ani-mal species, not to mention the difficulty of accessing sub-jects to scan in 3D. However, we argue that for domestic species of great importance, like the horse, it is a highly valuable investment to put effort into gathering a large dataset of real 3D scans, and learn a realistic 3D articu-lated shape model. We introduce VAREN, a novel 3D ar-ticulated parametric shape model learned from 3D scans of many real horses. VAREN bridges synthesis and analysis tasks, as the generated model instances have unprecedented realism, while being able to represent horses of different sizes and shapes. Differently from previous body models, VAREN has two resolutions, an anatomical skeleton, and interpretable, learned pose-dependent deformations, which are related to the body muscles. We show with experiments that this formulation has superior performance with respect to previous strategies for modeling pose-dependent deformations in the human body case, while also being more compact and allowing an analysis of the relationship be-tween articulation and muscle deformation during articu-lated motion. The VAREN model and data are available at https://varen.is.tue.mpg.de. Silvia Zuffi, Ylva Mellbin, Ci Li, Markus Höschle, Hedvig Kjellström, Senya Polikovsky, Elin Hernlund, Michael J. Black |
CVPR | 1 |
| 2024 | AWOL: Analysis WithOut Synthesis Using Language
Silvia Zuffi, Michael J. Black |
ECCV (58) | 1 |
| 2023 | BITE: Beyond Priors for Improved Three-D Dog Pose EstimationabstractWe address the problem of inferring the 3D shape and pose of dogs from images. Given the lack of 3D training data, this problem is challenging, and the best methods lag behind those designed to estimate human shape and pose. To make progress, we attack the problem from multiple sides at once. First, we need a good 3D shape prior, like those available for humans. To that end, we learn a dog-specific 3D parametric model, called D-SMAL. Second, existing methods focus on dogs in standing poses because when they sit or lie down, their legs are self occluded and their bodies deform. Without access to a good pose prior or 3D data, we need an alternative approach. To that end, we exploit contact with the ground as a form of side information. We consider an existing large dataset of dog images and label any 3D contact of the dog with the ground. We exploit body-ground contact in estimating dog pose and find that it significantly improves results. Third, we develop a novel neural network architecture to infer and exploit this contact information. Fourth, to make progress, we have to be able to measure it. Current evaluation metrics are based on 2D features like keypoints and silhouettes, which do not directly correlate with 3D errors. To address this, we create a synthetic dataset containing rendered images of scanned 3D dogs. With these advances, our method recovers significantly better dog shape and pose than the state of the art, and we evaluate this improvement in 3D. Our code, model and test dataset are publicly available for research purposes at https://bite.is.tue.mpg.de/. Nadine Bertsch, Shashank Tripathi, Konrad Schindler, Michael J. Black, Silvia Zuffi |
CVPR | 5 |
| 2023 | BARC: Breed-Augmented Regression Using Classification for 3D Dog Reconstruction from ImagesabstractAbstract The goal of this work is to reconstruct 3D dogs from monocular images. We take a model-based approach, where we estimate the shape and pose parameters of a 3D articulated shape model for dogs. We consider dogs as they constitute a challenging problem, given they are highly articulated and come in a variety of shapes and appearances. Recent work has considered a similar task using the multi-animal SMAL model, with additional limb scale parameters, obtaining reconstructions that are limited in terms of realism. Like previous work, we observe that the original SMAL model is not expressive enough to represent dogs of many different breeds. Moreover, we make the hypothesis that the supervision signal used to train the network, that is 2D keypoints and silhouettes, is not sufficient to learn a regressor that can distinguish between the large variety of dog breeds. We therefore go beyond previous work in two important ways. First, we modify the SMAL shape space to be more appropriate for representing dog shape. Second, we formulate novel losses that exploit information about dog breeds. In particular, we exploit the fact that dogs of the same breed have similar body shapes. We formulate a novel breed similarity loss, consisting of two parts: One term is a triplet loss, that encourages the shape of dogs from the same breed to be more similar than dogs of different breeds. The second one is a breed classification loss. With our approach we obtain 3D dogs that, compared to previous work, are quantitatively better in terms of 2D reconstruction, and significantly better according to subjective and quantitative 3D evaluations. Our work shows that a-priori side information about similarity of shape and appearance, as provided by breed labels, can help to compensate for the lack of 3D training data. This concept may be applicable to other animal species or groups of species. We call our method BARC (Breed-Augmented Regression using Classification). Our code is publicly available for research purposes at https://barc.is.tue.mpg.de/ . Nadine Bertsch, Silvia Zuffi, Konrad Schindler, Michael J. Black |
Int. J. Comput. Vis. | 2 |
| 2022 | OSSO: Obtaining Skeletal Shape from OutsideabstractWe address the problem of inferring the anatomic skeleton of a person, in an arbitrary pose, from the 3D surface of the body; i.e. we predict the inside (bones) from the outside (skin). This has many applications in medicine and biomechanics. Existing state-of-the-art biomechanical skeletons are detailed but do not easily generalize to new subjects. Additionally, computer vision and graphics methods that predict skeletons are typically heuristic, not learned from data, do not leverage the full 3D body surface, and are not validated against ground truth. To our knowledge, our system, called OSSO (Obtaining Skeletal Shape from Outside), is the first to learn the mapping from the 3D body surface to the internal skeleton from real data. We do so using 1000 male and 1000 female dual-energy X-ray absorptiometry (DXA) scans. To these, we fit a parametric 3D body shape model (STAR) to capture the body surface and a novel part-based 3D skeleton model to capture the bones. This provides inside/outside training pairs. We model the statistical variation of full skeletons using PCA in a pose-normalized space and train a regressor from body shape parameters to skeleton shape parameters. Given an arbitrary 3D body shape and pose, OSSO predicts a realistic skeleton inside. In contrast to previous work, we evaluate the accuracy of the skeleton shape quantitatively on held out DXA scans, outperforming the state-of-the art. We also show 3D skeleton prediction from varied and challenging 3D bodies. The code to infer a skeleton from a body shape is available at https://osso.is.tue.mpg.de, and the dataset of paired outer surface (skin) and skeleton (bone) meshes is available as a Biobank Returned Dataset. This research has been conducted using the UK Biobank Resource. Marilyn Keller, Silvia Zuffi, Michael J. Black, Sergi Pujades |
CVPR | 2 |
| 2022 | BARC: Learning to Regress 3D Dog Shape from Images by Exploiting Breed InformationabstractOur goal is to recover the 3D shape and pose of dogs from a single image. This is a challenging task because dogs exhibit a wide range of shapes and appearances, and are highly articulated. Recent work has proposed to directly regress the SMAL animal model, with additional limb scale parameters, from images. Our method, called BARC (Breed-Augmented Regression using Classification), goes beyond prior work in several important ways. First, we modify the SMAL shape space to be more appropriate for representing dog shape. But, even with a better shape model, the problem of regressing dog shape from an image is still challenging because we lack paired images with 3D ground truth. To compensate for the lack of paired data, we formulate novel losses that exploit information about dog breeds. In particular, we exploit the fact that dogs of the same breed have similar body shapes. We formulate a novel breed similarity loss consisting of two parts: One term encourages the shape of dogs from the same breed to be more similar than dogs of different breeds. The second one, a breed classification loss, helps to produce recognizable breed-specific shapes. Through ablation studies, we find that our breed losses significantly improve shape accuracy over a baseline without them. We also compare BARC qualitatively to WLDO with a perceptual study and find that our approach produces dogs that are significantly more realistic. This work shows that a-priori information about genetic similarity can help to compensate for the lack of 3D training data. This concept may be applicable to other animal species or groups of species. Our code is publicly available for research purposes at https://barc.is.tue.mpg.de/. Nadine Bertsch, Silvia Zuffi, Konrad Schindler, Michael J. Black |
CVPR | 2 |
| 2019 | Three-D Safari: Learning to Estimate Zebra Pose, Shape, and Texture From Images "In the Wild"abstractWe present the first method to perform automatic 3D pose, shape and texture capture of animals from images acquired in-the-wild. In particular, we focus on the problem of capturing 3D information about Grevy's zebras from a collection of images. The Grevy's zebra is one of the most endangered species in Africa, with only a few thousand individuals left. Capturing the shape and pose of these animals can provide biologists and conservationists with information about animal health and behavior. In contrast to research on human pose, shape and texture estimation, training data for endangered species is limited, the animals are in complex natural scenes with occlusion, they are naturally camouflaged, travel in herds, and look similar to each other. To overcome these challenges, we integrate the recent SMAL animal model into a network-based regression pipeline, which we train end-to-end on synthetically generated images with pose, shape, and background variation. Going beyond state-of-the-art methods for human shape and pose estimation, our method learns a shape space for zebras during training. Learning such a shape space from images using only a photometric loss is novel, and the approach can be used to learn shape in other settings with limited 3D supervision. Moreover, we couple 3D pose and shape prediction with the task of texture synthesis, obtaining a full texture map of the animal from a single image. We show that the predicted texture map allows a novel per-instance unsupervised optimization over the network features. This method, SMALST (SMAL with learned Shape and Texture) goes beyond previous work, which assumed manual keypoints and/or segmentation, to regress directly from pixels to 3D animal shape, pose and texture. Silvia Zuffi, Angjoo Kanazawa, Tanya Y. Berger-Wolf, Michael J. Black |
ICCV | 1 |
| 2018 | Lions and Tigers and Bears: Capturing Non-Rigid, 3D, Articulated Shape From ImagesabstractAnimals are widespread in nature and the analysis of their shape and motion is important in many fields and industries. Modeling 3D animal shape, however, is difficult because the 3D scanning methods used to capture human shape are not applicable to wild animals or natural settings. Consequently, we propose a method to capture the detailed 3D shape of animals from images alone. The articulated and deformable nature of animals makes this problem extremely challenging, particularly in unconstrained environments with moving and uncalibrated cameras. To make this possible, we use a strong prior model of articulated animal shape that we fit to the image data. We then deform the animal shape in a canonical reference pose such that it matches image evidence when articulated and projected into multiple images. Our method extracts significantly more 3D shape detail than previous methods and is able to model new species, including the shape of an extinct animal, using only a few video frames. Additionally, the projected 3D shapes are accurate enough to facilitate the extraction of a realistic texture map from multiple frames. Silvia Zuffi, Angjoo Kanazawa, Michael J. Black |
CVPR | 1 |
| 2017 | 3D Menagerie: Modeling the 3D Shape and Pose of AnimalsabstractThere has been significant work on learning realistic, articulated, 3D models of the human body. In contrast, there are few such models of animals, despite many applications. The main challenge is that animals are much less cooperative than humans. The best human body models are learned from thousands of 3D scans of people in specific poses, which is infeasible with live animals. Consequently, we learn our model from a small set of 3D scans of toy figurines in arbitrary poses. We employ a novel part-based shape model to compute an initial registration to the scans. We then normalize their pose, learn a statistical shape model, and refine the registrations and the model together. In this way, we accurately align animal scans from different quadruped families with very different shapes and poses. With the registration to a common template we learn a shape space representing animals including lions, cats, dogs, horses, cows and hippos. Animal shapes can be sampled from the model, posed, animated, and fit to data. We demonstrate generalization by fitting it to images of real animals including species not seen in training. Silvia Zuffi, Angjoo Kanazawa, David Jacobs 0001, Michael J. Black |
CVPR | 1 |
| 2016 | Body talk: crowdshaping realistic 3D avatars with wordsabstractRealistic, metrically accurate, 3D human avatars are useful for games, shopping, virtual reality, and health applications. Such avatars are not in wide use because solutions for creating them from high-end scanners, low-cost range cameras, and tailoring measurements all have limitations. Here we propose a simple solution and show that it is surprisingly accurate. We use crowdsourcing to generate attribute ratings of 3D body shapes corresponding to standard linguistic descriptions of 3D shape. We then learn a linear function relating these ratings to 3D human shape parameters. Given an image of a new body, we again turn to the crowd for ratings of the body shape. The collection of linguistic ratings of a photograph provides remarkably strong constraints on the metric 3D shape. We call the process crowdshaping and show that our Body Talk system produces shapes that are perceptually indistinguishable from bodies created from high-resolution scans and that the metric accuracy is sufficient for many tasks. This makes body "scanning" practical without a scanner, opening up new applications including database search, visualization, and extracting avatars from books. Stephan Streuber, Maria Alejandra Quiros-Ramirez, Matthew Q. Hill, Carina A. Hahn, Silvia Zuffi, Alice J. O'Toole, Michael J. Black |
ACM Trans. Graph. | 5 |
| 2015 | The stitched puppet: A graphical model of 3D human shape and poseabstractWe propose a new 3D model of the human body that is both realistic and part-based. The body is represented by a graphical model in which nodes of the graph correspond to body parts that can independently translate and rotate in 3D and deform to represent different body shapes and to capture pose-dependent shape variations. Pairwise potentials define a “stitching cost” for pulling the limbs apart, giving rise to the stitched puppet (SP) model. Unlike existing realistic 3D body models, the distributed representation facilitates inference by allowing the model to more effectively explore the space of poses, much like existing 2D pictorial structures models. We infer pose and body shape using a form of particle-based max-product belief propagation. This gives SP the realism of recent 3D body models with the computational advantages of part-based models. We apply SP to two challenging problems involving estimating human shape and pose from 3D data. The first is the FAUST mesh alignment challenge, where ours is the first method to successfully align all 3D meshes with no pose prior. The second involves estimating pose and shape from crude visual hull representations of complex body movements. Silvia Zuffi, Michael J. Black |
CVPR | 1 |
| 2014 | Preserving Modes and Messages via Diverse Particle SelectionabstractIn applications of graphical models arising in domains such as computer vision and signal processing, we often seek the most likely configurations of high-dimensional, continuous variables. We develop a particle-based max-product algorithm which maintains a diverse set of posterior mode hypotheses, and is robust to initialization. At each iteration, the set of hypotheses at each node is augmented via stochastic proposals, and then reduced via an efficient selection algorithm. The integer program underlying our optimization-based particle selection minimizes errors in subsequent max-product message updates. This objective automatically encourages diversity in the maintained hypotheses, without requiring tuning of application-specific distances among hypotheses. By avoiding the stochastic resampling steps underlying particle sum-product algorithms, we also avoid common degeneracies where particles collapse onto a single hypothesis. Our approach significantly outperforms previous particle-based algorithms in experiments focusing on the estimation of human pose from single images. Jason L. Pacheco, Silvia Zuffi, Michael J. Black, Erik B. Sudderth |
ICML | 2 |
| 2013 | Towards Understanding Action RecognitionabstractAlthough action recognition in videos is widely studied, current methods often fail on real-world datasets. Many recent approaches improve accuracy and robustness to cope with challenging video sequences, but it is often unclear what affects the results most. This paper attempts to provide insights based on a systematic performance evaluation using thoroughly-annotated data of human actions. We annotate human Joints for the HMDB dataset (J-HMDB). This annotation can be used to derive ground truth optical flow and segmentation. We evaluate current methods using this dataset and systematically replace the output of various algorithms with ground truth. This enables us to discover what is important - for example, should we work on improving flow algorithms, estimating human bounding boxes, or enabling pose estimation? In summary, we find that high-level pose features greatly outperform low/mid level features, in particular, pose over time is critical, but current pose estimation algorithms are not yet reliable enough to provide this information. We also find that the accuracy of a top-performing action recognition framework can be greatly increased by refining the underlying low/mid level features, this suggests it is important to improve optical flow and human detection algorithms. Our analysis and J-HMDB dataset should facilitate a deeper understanding of action recognition algorithms. Hueihan Jhuang, Juergen Gall, Silvia Zuffi, Cordelia Schmid, Michael J. Black |
ICCV | 3 |
| 2013 | Estimating Human Pose with Flowing PuppetsabstractWe address the problem of upper-body human pose estimation in uncontrolled monocular video sequences, without manual initialization. Most current methods focus on isolated video frames and often fail to correctly localize arms and hands. Inferring pose over a video sequence is advantageous because poses of people in adjacent frames exhibit properties of smooth variation due to the nature of human and camera motion. To exploit this, previous methods have used prior knowledge about distinctive actions or generic temporal priors combined with static image likelihoods to track people in motion. Here we take a different approach based on a simple observation: Information about how a person moves from frame to frame is present in the optical flow field. We develop an approach for tracking articulated motions that "links" articulated shape models of people in adjacent frames through the dense optical flow. Key to this approach is a 2D shape model of the body that we use to compute how the body moves over time. The resulting "flowing puppets" provide a way of integrating image evidence across frames to improve pose inference. We apply our method on a challenging dataset of TV video sequences and show state-of-the-art performance. Silvia Zuffi, Javier Romero 0002, Cordelia Schmid, Michael J. Black |
ICCV | 1 |
| 2012 | From Pictorial Structures to deformable structuresabstractPictorial Structures (PS) define a probabilistic model of 2D articulated objects in images. Typical PS models assume an object can be represented by a set of rigid parts connected with pairwise constraints that define the prior probability of part configurations. These models are widely used to represent non-rigid articulated objects such as humans and animals despite the fact that such objects have parts that deform non-rigidly. Here we define a new Deformable Structures (DS) model that is a natural extension of previous PS models and that captures the non-rigid shape deformation of the parts. Each part in a DS model is represented by a low-dimensional shape deformation space and pairwise potentials between parts capture how the shape varies with pose and the shape of neighboring parts. A key advantage of such a model is that it more accurately models object boundaries. This enables image likelihood models that are more discriminative than previous PS likelihoods. This likelihood is learned using training imagery annotated using a DS “puppet.” We focus on a human DS model learned from 2D projections of a realistic 3D human body model and use it to infer human poses in images using a form of non-parametric belief propagation. Silvia Zuffi, Oren Freifeld, Michael J. Black |
CVPR | 1 |
| 2010 | Contour people: A parameterized model of 2D articulated human shapeabstractWe define a new “contour person” model of the human body that has the expressive power of a detailed 3D model and the computational benefits of a simple 2D part-based model. The contour person (CP) model is learned from a 3D SCAPE model of the human body that captures natural shape and pose variations; the projected contours of this model, along with their segmentation into parts forms the training set. The CP model factors deformations of the body into three components: shape variation, viewpoint change and part rotation. This latter model also incorporates a learned non-rigid deformation model. The result is a 2D articulated model that is compact to represent, simple to compute with and more expressive than previous models. We demonstrate the value of such a model in 2D pose estimation and segmentation. Given an initial pose from a standard pictorial-structures method, we refine the pose and shape using an objective function that segments the scene into foreground and background regions. The result is a parametric, human-specific, image segmentation. Oren Freifeld, Alexander Weiss, Silvia Zuffi, Michael J. Black |
CVPR | 3 |
| 2007 | A computational strategy exploiting genetic algorithms to recover color surface reflectance functions
Raimondo Schettini, Silvia Zuffi |
Neural Comput. Appl. | 2 |
| 2006 | Promoting Cultural Tourism across Mediterranean Countries through ICT technologies: The Daedalus Project
Alfredo De Massis, Anna Della Ventura, Triantafillos Karathanasis, Giammarco Tosi, Silvia Zuffi |
ENTER | 5 |
| 2001 | A system for the automatic selection of conspicuous color sets for qualitative data displayabstractThe authors describe the main features of a system supporting the selection of color palettes for qualitative data representation, such as in supervised or unsupervised image classification. Based on visual interaction, the system provides effective tools for browsing the Munsell color space and setting perceptual constraints on the colors, which it then selects automatically. The system is now available for academic and nonprofit purposes. Paola Campadelli, Raimondo Schettini, Silvia Zuffi |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 1999 | A Model-Based Method for the Reconstruction of Total Knee Replacement KinematicsabstractA better knowledge of the kinematics behavior of total knee replacement (TKR) during activity still remains a crucial issue to validate innovative prosthesis designs and different surgical strategies. Tools for more accurate measurement of in vivo kinematics of knee prosthesis components are therefore fundamental to improve the clinical outcome of knee replacement. In the present study, a novel model-based method for the estimation of the three-dimensional (3-D) position and orientation (pose) of both the femoral and tibial knee prosthesis components during activity is presented. The knowledge of the 3-D geometry of the components and a single plane projection view in a fluoroscopic image are sufficient to reconstruct the absolute and relative pose of the components in space. The technique is based on the best alignment of the component designs with the corresponding projection on the image plane. The image generation process is modeled and an iterative procedure localizes the spatial pose of the object by minimizing the Euclidean distance of the projection rays from the object surface. Computer simulation and static/dynamic in vitro tests using real knee prosthesis show that the accuracy with which relative orientation and position of the components can be estimated is better than 1.5 degrees and 1.5 mm, respectively. In vivo tests demonstrate that the method is well suited for kinematics analysis on TKR patients and that good quality images can be obtained with a carefully positioning of the fluoroscope and an appropriate dosage. With respect to previously adopted template matching techniques, the present method overcomes the complete segmentation of the components on the projected image and also features the simultaneous evaluation of all the six degrees of freedom (DOF) of the object. The expected small difference between successive poses in in vivo sequences strongly reduces the frequency of false poses and both the operator and computation time. Silvia Zuffi, Alberto Leardini, Fabio Catani, Silvia Fantozzi, Angelo Cappello |
IEEE Trans. Medical Imaging | 1 |