EDBT 2026 Demo / reviewers in the wild / expert
Federica Bogo
dblp:132/2146
· DBLP profile ↗
23ranked-venue papers
6as first author
12since 2021 · last 2026
0009-0003-4991-185XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 18 · 4 first-author · 9 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Masked Modeling for Human Motion Recovery Under OcclusionsabstractHuman motion reconstruction from monocular videos is a fundamental challenge in computer vision, with broad applications in AR/VR, robotics, and digital content creation, but remains challenging under frequent occlusions in real-world settings. Existing regression-based methods are efficient but fragile to missing observations, while optimization- and diffusion-based approaches improve robustness at the cost of slow inference speed and heavy preprocessing steps. To address these limitations, we leverage recent advances in generative masked modeling and present MoRo: Masked mOdeling for human motion Recovery under Occlusions. MoRo is an occlusion-robust, end-to-end generative framework that formulates motion reconstruction as a video-conditioned task, and efficiently recover human motion in a consistent global coordinate system from RGB videos. By masked modeling, MoRo naturally handles occlusions while enabling efficient, end-to-end inference. To overcome the scarcity of paired video-motion data, we design a cross-modality learning scheme that learns multimodal priors from a set of heterogeneous datasets: (i) a trajectory-aware motion prior trained on MoCap datasets, (ii) an image-conditioned pose prior trained on image-pose datasets, capturing diverse frame-level poses, and (iii) a video-conditioned masked transformer that fuses motion and pose priors, finetuned on video-motion datasets to integrate visual cues with motion dynamics for robust inference. Extensive experiments on EgoBody and RICH demonstrate that MoRo substantially outperforms state-of-the-art methods in accuracy and motion realism under occlusions, while performing on-par in non-occluded scenarios. MoRo achieves real-time inference at 70 FPS on a single H200 GPU. Project page: https://mikeqzy.github.io/MoRo. Zhiyin Qian, Bharat Lal Bhatnagar, Federica Bogo, Siyu Tang 0001 |
3DV | 4 |
| 2025 | ATLAS: Decoupling Skeletal and Shape Parameters for Expressive Parametric Human ModelingabstractParametric body models offer expressive 3D representation of humans across a wide range of poses, shapes, and facial expressions, typically derived by learning a basis over registered 3D meshes. However, existing human mesh modeling approaches struggle to capture detailed variations across diverse body poses and shapes, largely due to limited training data diversity and restrictive modeling assumptions. Moreover, the common paradigm first optimizes the external body surface using a linear basis, then regresses internal skeletal joints from surface vertices. This approach introduces problematic dependencies between internal skeleton and outer soft tissue, limiting direct control over body height and bone lengths. To address these issues, we present ATLAS, a high-fidelity body model learned from 600k high-resolution scans captured using 240 synchronized cameras. Unlike previous methods, we explicitly decouple the shape and skeleton bases by grounding our mesh representation in the human skeleton. This decoupling enables enhanced shape expressivity, fine-grained customization of body attributes, and keypoint fitting independent of external soft-tissue characteristics. ATLAS outperforms existing methods by fitting unseen subjects in diverse poses more accurately, and quantitative evaluations show that our non-linear pose correctives more effectively capture complex poses compared to linear models. Jinhyung Park, Javier Romero 0002, Shunsuke Saito, Fabian Prada, Takaaki Shiratori, Federica Bogo, Shoou-I Yu, Kris Makoto Kitani, Rawal Khirodkar |
ICCV | 7 |
| 2025 | Generating High-Fidelity Clothed Human Dynamics with Temporal DiffusionabstractClothed human modeling plays a crucial role in multimedia research, with applications spanning virtual reality, gaming, and fashion design. The goal is to learn clothed human dynamics from observations and then generate humans with high-fidelity clothing details for motion animation. Despite tremendous advancements in clothing shape analysis by existing approaches, the community still faces challenges in generating convincing visual effects of cloth dynamics, maintaining temporally smooth clothing details, and handling diverse clothing patterns. To address these challenges, we introduce ClothDiffuse, a temporal diffusion model that seamlessly integrates three key components into this task—temporal dynamics modeling, iterative refinement, and diversified generation. Our approach begins by using an encoder to extract high-level temporal features from input human body motions. These features are combined with a learnable pixel-aligned garment feature, serving as prior conditions for the shape decoder. The decoder then iteratively denoise Gaussian noise to produce clothing deformations over time on the input unclothed human bodies. To ensure that the results align with observations and adhere to physical plausibility for clothing shape inference, we propose two physics-inspired loss functions that preserve the intra-frame distances and inter-frame forces of clothing points. Additionally, the stochastic nature of the denoising process allows for the generation of diverse and plausible clothing shapes. Experiments show that our approach outperforms state-of-the-art methods in chamfer distance and visual effects, particularly for loose clothing such as dresses and skirts. Furthermore, our approach effectively adapts to out-of-domain clothing types and generates realistic clothes dynamics. Shihao Zou, Yuanlu Xu, Nikolaos Sarafianos, Federica Bogo, Tony Tung, Weixin Si, Li Cheng 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | RoHM: Robust Human Motion Reconstruction via DiffusionabstractWe propose RoHM, an approach for robust 3D human motion reconstruction from monocular RGB(-D) videos in the presence of noise and occlusions. Most previous approaches either train neural networks to directly regress motion in 3D or learn data-driven motion priors and com-bine them with optimization at test time. The former do not recover globally coherent motion and fail under occlusions; the latter are time-consuming, prone to local minima, and require manual tuning. To overcome these shortcomings, we exploit the iterative, denoising nature of diffusion models. RoHM is a novel diffusion-based motion model that, conditioned on noisy and occluded input data, reconstructs complete, plausible motions in consistent global co-ordinates. Given the complexity of the problem - requiring one to address different tasks (denoising and infilling) in different solution spaces (local and global motion) - we de-compose it into two sub-tasks and learn two models, one for global trajectory and one for local motion. To capture the correlations between the two, we then introduce a novel conditioning module, combining it with an iterative inference scheme. We apply RoHM to a variety of tasks from motion reconstruction and denoising to spatial and temporal infilling. Extensive experiments on three popular datasets show that our method outperforms state-of-the-art approaches qualitatively and quantitatively, while being faster at test time. The code is available at https://sanweiliti.github.io/ROHM/ROHM.html. Bharat Lal Bhatnagar, Yuanlu Xu, Alexander Winkler, Petr Kadlecek, Siyu Tang 0001, Federica Bogo |
CVPR | 7 |
| 2024 | SplatFields: Neural Gaussian Splats for Sparse 3D and 4D Reconstruction
Marko Mihajlovic, Sergey Prokudin, Siyu Tang 0001, Robert Maier 0001, Federica Bogo, Tony Tung, Edmond Boyer |
ECCV (2) | 5 |
| 2023 | Domain-Specific Priors and Meta Learning for Few-Shot First-Person Action RecognitionabstractThe lack of large-scale real datasets with annotations makes transfer learning a necessity for video activity understanding. We aim to develop an effective method for few-shot transfer learning for first-person action classification. We leverage independently trained local visual cues to learn representations that can be transferred from a source domain, which provides primitive action labels, to a different target domain using only a handful of examples. Visual cues we employ include object-object interactions, hand grasps and motion within regions that are a function of hand locations. We employ a framework based on meta-learning to extract the distinctive and domain invariant components of the deployed visual cues. This enables transfer of action classification models across public datasets captured with diverse scene and action configurations. We present comparative results of our transfer learning methodology and report superior results over state-of-the-art action classification approaches for both inter-class and inter-dataset transfer. Huseyin Coskun, M. Zeeshan Zia, Bugra Tekin, Federica Bogo, Nassir Navab, Federico Tombari, Harpreet Sawhney |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | FLAG: Flow-based 3D Avatar Generation from Sparse ObservationsabstractTo represent people in mixed reality applications for collaboration and communication, we need to generate realistic and faithful avatar poses. However, the signal streams that can be applied for this task from head-mounted devices (HMDs) are typically limited to head pose and hand pose estimates. While these signals are valuable, they are an incomplete representation of the human body, making it challenging to generate a faithful full-body avatar. We address this challenge by developing a flow-based generative model of the 3D human body from sparse observations, wherein we learn not only a conditional distribution of 3D human pose, but also a probabilistic mapping from observations to the latent space from which we can generate a plausible pose along with uncertainty estimates for the joints. We show that our approach is not only a strong predictive model, but can also act as an efficient pose prior in different optimization settings where a good initial latent code plays a major role. Mohammad Sadegh Ali Akbarian, Pashmina Cameron, Federica Bogo, Andrew W. Fitzgibbon, Thomas J. Cashman 0001 |
CVPR | 3 |
| 2022 | Learning to Fit Morphable Models
Vasileios Choutas, Federica Bogo, Jingjing Shen, Julien P. C. Valentin |
ECCV (6) | 2 |
| 2022 | EgoBody: Human Body Shape and Motion of Interacting People from Head-Mounted Devices
Qianli Ma 0007, Yan Zhang 0054, Zhiyin Qian, Taein Kwon, Marc Pollefeys, Federica Bogo, Siyu Tang 0001 |
ECCV (6) | 7 |
| 2021 | Spatio-Temporal Human Shape Completion With Implicit Function NetworksabstractWe address the problem of inferring a human shape from partial observations, such as depth images, in temporal sequences. Deep Neural Networks (DNN) have been shown successful to estimate detailed shapes on a frame-by-frame basis but consider yet little or no temporal information over frame sequences for detailed shape estimation. Recently, networks that implicitly encode shape occupancy using MLP layers have shown very promising results for such single-frame shape inference, with the advantage of reducing the dimensionality of the problem and providing continuously encoded results. In this work we propose to generalize implicit encoding to spatio-temporal shape inference with spatio-temporal implicit function networks or STIF-Nets, where temporal redundancy and continuity is expected to improve the shape and motion quality. To validate these added benefits, we collect and train with motion data from CAPE for dressed humans, and DFAUST for body shapes with no clothing. We show our model’s ability to estimate shapes for a set of input frames, and interpolate between them. Our results show that our method outpetforms existing state of the art methods, in particular the single-frame methods for detailed shape estimation. Boyao Zhou, Jean-Sébastien Franco, Federica Bogo, Edmond Boyer |
3DV | 3 |
| 2021 | H2O: Two Hands Manipulating Objects for First Person Interaction RecognitionabstractWe present a comprehensive framework for egocentric interaction recognition using markerless 3D annotations of two hands manipulating objects. To this end, we propose a method to create a unified dataset for egocentric 3D interaction recognition. Our method produces annotations of the 3D pose of two hands and the 6D pose of the manipulated objects, along with their interaction labels for each frame. Our dataset, called H2O (2 Hands and Objects), provides synchronized multi-view RGB-D images, interaction labels, object classes, ground-truth 3D poses for left & right hands, 6D object poses, ground-truth camera poses, object meshes and scene point clouds. To the best of our knowledge, this is the first benchmark that enables the study of first-person actions with the use of the pose of both left and right hands manipulating objects and presents an unprecedented level of detail for egocentric 3D interaction recognition. We further propose the method to predict interaction classes by estimating the 3D pose of two hands and the 6D pose of the manipulated objects, jointly from RGB images. Our method models both inter- and intra-dependencies between both hands and objects by learning the topology of a graph convolutional network that predicts interactions. We show that our method facilitated by this dataset establishes a strong baseline for joint hand-object pose estimation and achieves state-of-the-art accuracy for first person interaction recognition. Taein Kwon, Bugra Tekin, Jan Stühmer, Federica Bogo, Marc Pollefeys |
ICCV | 4 |
| 2021 | Learning Motion Priors for 4D Human Body Capture in 3D ScenesabstractRecovering high-quality 3D human motion in complex scenes from monocular videos is important for many applications, ranging from AR/VR to robotics. However, capturing realistic human-scene interactions, while dealing with occlusions and partial views, is challenging; current approaches are still far from achieving compelling results. We address this problem by proposing LEMO: LEarning human MOtion priors for 4D human body capture. By leveraging the large-scale motion capture dataset AMASS [38], we introduce a novel motion smoothness prior, which strongly reduces the jitters exhibited by poses recovered over a sequence. Furthermore, to handle contacts and occlusions occurring frequently in body-scene interactions, we design a contact friction term and a contact-aware motion infiller obtained via per-instance self-supervised training. To prove the effectiveness of the proposed motion priors, we combine them into a novel pipeline for 4D human body capture in 3D scenes. With our pipeline, we demonstrate high-quality 4D human body capture, reconstructing smooth motions and physically plausible body-scene interactions. The code and data are available at https://sanweiliti.github.io/LEMO/LEMO.html. Yan Zhang 0054, Federica Bogo, Marc Pollefeys, Siyu Tang 0001 |
ICCV | 3 |
| 2020 | Reconstructing Human Body Mesh from Point Clouds by Adversarial GP Network
Boyao Zhou, Jean-Sébastien Franco, Federica Bogo, Bugra Tekin, Edmond Boyer |
ACCV (1) | 3 |
| 2020 | Leveraging Photometric Consistency Over Time for Sparsely Supervised Hand-Object ReconstructionabstractModeling hand-object manipulations is essential for understanding how humans interact with their environment. While of practical importance, estimating the pose of hands and objects during interactions is challenging due to the large mutual occlusions that occur during manipulation. Recent efforts have been directed towards fully-supervised methods that require large amounts of labeled training samples. Collecting 3D ground-truth data for hand-object interactions, however, is costly, tedious, and error-prone. To overcome this challenge we present a method to leverage photometric consistency across time when annotations are only available for a sparse subset of frames in a video. Our model is trained end-to-end on color images to jointly reconstruct hands and objects in 3D by inferring their poses. Given our estimated reconstructions, we differentiably render the optical flow between pairs of adjacent images and use it within the network to warp one frame to another. We then apply a self-supervised photometric loss that relies on the visual consistency between nearby images. We achieve state-of-the-art results on 3D hand-object reconstruction benchmarks and demonstrate that our approach allows us to improve the pose estimation accuracy by leveraging information from neighboring frames in low-data regimes. Yana Hasson, Bugra Tekin, Federica Bogo, Ivan Laptev, Marc Pollefeys, Cordelia Schmid |
CVPR | 3 |
| 2020 | The Phong Surface: Efficient 3D Model Fitting Using Lifted Optimization
Jingjing Shen, Thomas J. Cashman 0001, Qi Ye 0001, Tim Hutton, Toby Sharp, Federica Bogo, Andrew W. Fitzgibbon, Jamie Shotton |
ECCV (1) | 6 |
| 2019 | H+O: Unified Egocentric Recognition of 3D Hand-Object Poses and InteractionsabstractWe present a unified framework for understanding 3D hand and object interactions in raw image sequences from egocentric RGB cameras. Given a single RGB image, our model jointly estimates the 3D hand and object poses, models their interactions, and recognizes the object and action classes with a single feed-forward pass through a neural network. We propose a single architecture that does not rely on external detection algorithms but rather is trained end-to-end on single images. We further merge and propagate information in the temporal domain to infer interactions between hand and object trajectories and recognize actions. The complete model takes as input a sequence of frames and outputs per-frame 3D hand and object pose predictions along with the estimates of object and action categories for the entire sequence. We demonstrate state-of-the-art performance of our algorithm even in comparison to the approaches that work on depth data and ground-truth annotations. Bugra Tekin, Federica Bogo, Marc Pollefeys |
CVPR | 2 |
| 2017 | Dynamic FAUST: Registering Human Bodies in MotionabstractWhile the ready availability of 3D scan data has influenced research throughout computer vision, less attention has focused on 4D data, that is 3D scans of moving non-rigid objects, captured over time. To be useful for vision research, such 4D scans need to be registered, or aligned, to a common topology. Consequently, extending mesh registration methods to 4D is important. Unfortunately, no ground-truth datasets are available for quantitative evaluation and comparison of 4D registration methods. To address this we create a novel dataset of high-resolution 4D scans of human subjects in motion, captured at 60 fps. We propose a new mesh registration method that uses both 3D geometry and texture information to register all scans in a sequence to a common reference topology. The approach exploits consistency in texture over both short and long time intervals and deals with temporal offsets between shape and texture capture. We show how using geometry alone results in significant errors in alignment when the motions are fast and non-rigid. We evaluate the accuracy of our registration and provide a dataset of 40,000 raw and aligned meshes. Dynamic FAUST extends the popular FAUST dataset to dynamic 4D data, and is available for research purposes at http://dfaust.is.tue.mpg.de. Federica Bogo, Javier Romero 0002, Gerard Pons-Moll, Michael J. Black |
CVPR | 1 |
| 2017 | Unite the People: Closing the Loop Between 3D and 2D Human Representationsabstract3D models provide a common ground for different representations of human bodies. In turn, robust 2D estimation has proven to be a powerful tool to obtain 3D fits in-the-wild. However, depending on the level of detail, it can be hard to impossible to acquire labeled data for training 2D estimators on large scale. We propose a hybrid approach to this problem: with an extended version of the recently introduced SMPLify method, we obtain high quality 3D body model fits for multiple human pose datasets. Human annotators solely sort good and bad fits. This procedure leads to an initial dataset, UP-3D, with rich annotations. With a comprehensive set of experiments, we show how this data can be used to train discriminative models that produce results with an unprecedented level of detail: our models predict 31 segments and 91 landmark locations on the body. Using the 91 landmark pose estimator, we present state-of-the art results for 3D human pose and shape estimation using an order of magnitude less training data and without assumptions about gender or pose in the fitting procedure. We show that UP-3D can be enhanced with these improved fits to grow in quantity and quality, which makes the system deployable on large scale. The data, code and models are available for research purposes. Christoph Lassner, Javier Romero 0002, Martin Kiefel, Federica Bogo, Michael J. Black, Peter V. Gehler |
CVPR | 4 |
| 2016 | Keep It SMPL: Automatic Estimation of 3D Human Pose and Shape from a Single Image
Federica Bogo, Angjoo Kanazawa, Christoph Lassner, Peter V. Gehler, Javier Romero 0002, Michael J. Black |
ECCV (5) | 1 |
| 2015 | Detailed Full-Body Reconstructions of Moving People from Monocular RGB-D SequencesabstractWe accurately estimate the 3D geometry and appearance of the human body from a monocular RGB-D sequence of a user moving freely in front of the sensor. Range data in each frame is first brought into alignment with a multi-resolution 3D body model in a coarse-to-fine process. The method then uses geometry and image texture over time to obtain accurate shape, pose, and appearance information despite unconstrained motion, partial views, varying resolution, occlusion, and soft tissue deformation. Our novel body model has variable shape detail, allowing it to capture faces with a high-resolution deformable head model and body shape with lower-resolution. Finally we combine range data from an entire sequence to estimate a high-resolution displacement map that captures fine shape details. We compare our recovered models with high-resolution scans from a professional system and with avatars created by a commercial product. We extract accurate 3D avatars from challenging motion sequences and even capture soft tissue dynamics. Federica Bogo, Michael J. Black, Matthew Loper, Javier Romero 0002 |
ICCV | 1 |
| 2014 | FAUST: Dataset and Evaluation for 3D Mesh RegistrationabstractNew scanning technologies are increasing the importance of 3D mesh data and the need for algorithms that can reliably align it. Surface registration is important for building full 3D models from partial scans, creating statistical shape models, shape retrieval, and tracking. The problem is particularly challenging for non-rigid and articulated objects like human bodies. While the challenges of real-world data registration are not present in existing synthetic datasets, establishing ground-truth correspondences for real 3D scans is difficult. We address this with a novel mesh registration technique that combines 3D shape and appearance information to produce high-quality alignments. We define a new dataset called FAUST that contains 300 scans of 10 people in a wide range of poses together with an evaluation methodology. To achieve accurate registration, we paint the subjects with high-frequency textures and use an extensive validation process to ensure accurate ground truth. We find that current shape registration methods have trouble with this real-world data. The dataset and evaluation website are available for research purposes at http://faust.is.tue.mpg.de. Federica Bogo, Javier Romero 0002, Matthew Loper, Michael J. Black |
CVPR | 1 |
| 2014 | Automated Detection of New or Evolving Melanocytic Lesions Using a 3D Body Model
Federica Bogo, Javier Romero 0002, Enoch Peserico, Michael J. Black |
MICCAI (1) | 1 |
| 2013 | Optimal throughput and delay in delay-tolerant networks with ballistic mobilityabstractThis work studies delay and throughput achievable in delay-tolerant networks with ballistic mobility -- informally, when the average distance a node travels before changing direction does not become vanishingly small as the number of nodes in the deployment area grows. Ballistic mobility is a simple condition satisfied by a large number of well-studied mobility models, including the i.i.d. model, the random waypoint model, the uniform mobility model and Levy walks with exponent less than 1. Our contribution is twofold. First, we show that, under some very mild and natural hypotheses satisfied by all models in the literature, ballistic mobility is strictly necessary to achieve simultaneously, as the number of nodes grows, a) per-node throughput that does not become vanishingly small and b) communication delay that does not become infinitely large. Any network whose nodes exhibit a more "local" mobility pattern (e.g. Levy walks with exponent greater than 1, or Brownian motion) must sacrifice either a) or b), regardless of the communication scheme adopted -- even with network coding. Federica Bogo, Enoch Peserico |
MobiCom | 1 |