Diogo C. Luvizon

dblp:148/9918 · also Diogo Carbonera Luvizon · DBLP profile ↗
← Back
25ranked-venue papers
11as first author
19since 2021 · last 2025
0000-0002-5055-500XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 6 first-author · 14 since 2021Artificial intelligence and machine learning · 14 · 6 first-author · 12 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Ego4o: Egocentric Human Motion Capture and Understanding from Multi-Modal Input
abstract
This work focuses on tracking and understanding human motion using consumer wearable devices, such as VR/AR headsets, smart glasses, cellphones, and smartwatches. These devices provide diverse, multi-modal sensor inputs, including egocentric images, and 1-3 sparse IMU sensors in varied combinations. Motion descriptions can also accompany these signals. The diverse input modalities and their intermittent availability pose challenges for consistent motion capture and understanding. In this work, we present Ego4o (o for omni), a new framework for simultaneous human motion capture and understanding from multi-modal egocentric inputs. This method maintains performance with partial inputs while achieving better results when multiple modalities are combined. First, the IMU sensor inputs, the optional egocentric image, and text description of human motion are encoded into the latent space of a motion VQ-VAE. Next, the latent vectors are sent to the VQ-VAE decoder and optimized to track human motion. When motion descriptions are unavailable, the latent vectors can be input into a multi-modal LLM to generate human motion descriptions, which can further enhance motion capture accuracy. Quantitative and qualitative evaluations demonstrate the effectiveness of our method in predicting accurate human motion and high-quality motion descriptions. Project page: https://jianwang-mpi.github.io/ego4o.
Jian Wang 0100, Rishabh Dabral, Diogo C. Luvizon, Zhe Cao 0003, Lingjie Liu, Thabo Beeler, Christian Theobalt
CVPR3
2025 PocoLoco: A Point Cloud Diffusion Model of Human Shape in Loose Clothing
abstract
Modeling a human avatar that can plausibly deform to articulations is an active area of research. We present PocoLoco - the first template-free, point-based, pose-conditioned generative model for 3D humans in loose clothing. We motivate our work by noting that most methods require a parametric model of the human body to ground pose-dependent deformations. Consequently, they are restricted to modeling clothing that is topologically similar to the naked body and do not extend well to loose clothing. The few methods that attempt to model loose clothing typically require either canonicalization or a UV-parameterization and need to address the challenging problem of explicitly estimating correspondences for the deforming clothes. In this work, we formulate avatar clothing deformation as a conditional point-cloud generation task within the denoising diffusion framework. Crucially, our framework operates directly on unordered point clouds, eliminating the need for a parametric model or a clothing template. This also enables a variety of practical applications, such as point-cloud completion and pose-based editing - important features for virtual human animation. As current datasets for human avatars in loose clothing are far too small for training diffusion models, we release a dataset of two subjects performing various poses in loose clothing with a total of 75K point clouds. By contributing towards tackling the challenging task of effectively modeling loose clothing and expanding the available data for training these models, we aim to set the stage for further innovation in digital humans. The source code is available at https://github.com/sidsunny/pocoloco.
Siddharth Seth, Rishabh Dabral, Diogo C. Luvizon, Marc Habermann, Ming-Hsuan Yang 0001, Christian Theobalt, Adam Kortylewski
WACV3
2025 EventEgo3D++: 3D Human Motion Capture from A Head-Mounted Event Camera
abstract
Monocular egocentric 3D human motion capture remains a significant challenge, particularly under conditions of low lighting and fast movements, which are common in head-mounted device applications. Existing methods that rely on RGB cameras often fail under these conditions. To address these limitations, we introduce EventEgo3D++, the first approach that leverages a monocular event camera with a fisheye lens for 3D human motion capture. Event cameras excel in high-speed scenarios and varying illumination due to their high temporal resolution, providing reliable cues for accurate 3D human motion capture. EventEgo3D++ leverages the LNES representation of event streams to enable precise 3D reconstructions. We have also developed a mobile head-mounted device (HMD) prototype equipped with an event camera, capturing a comprehensive dataset that includes real event observations from both controlled studio environments and in-the-wild settings, in addition to a synthetic dataset. Additionally, to provide a more holistic dataset, we include allocentric RGB streams that offer different perspectives of the HMD wearer, along with their corresponding SMPL body model. Our experiments demonstrate that EventEgo3D++ achieves superior 3D accuracy and robustness compared to existing solutions, even in challenging conditions. Moreover, our method supports real-time 3D pose updates at a rate of 140Hz. This work is an extension of the EventEgo3D approach (CVPR 2024) and further advances the state of the art in egocentric 3D human motion capture. For more details, visit the project page at https://eventego3d.mpi-inf.mpg.de.
Christen Millerdurai, Hiroyasu Akada, Jian Wang 0042, Diogo C. Luvizon, Alain Pagani, Didier Stricker, Christian Theobalt, Vladislav Golyanik
Int. J. Comput. Vis.4
2024 3D Pose Estimation of Two Interacting Hands from a Monocular Event Camera
abstract
3D hand tracking from a monocular video is a very challenging problem due to hand interactions, occlusions, left-right hand ambiguity, and fast motion. Most existing methods rely on RGB inputs, which have severe limitations under low-light conditions and suffer from motion blur. In contrast, event cameras capture local brightness changes instead of full image frames and do not suffer from the described effects. Unfortunately, existing image-based techniques cannot be directly applied to events due to significant differences in the data modalities. In response to these challenges, this paper introduces the first framework for 3D tracking of two fast-moving and interacting hands from a single monocular event camera. Our approach tackles the left-right hand ambiguity with a novel semi-supervised feature-wise attention mechanism and integrates an intersection loss to fix hand collisions. To facilitate advances in this research domain, we release a new synthetic large-scale dataset of two interacting hands, Ev2Hands-S, and a new real benchmark with real event streams and ground-truth 3D annotations, Ev2Hands-R. Our approach outperforms existing methods in terms of the 3D reconstruction accuracy and generalises to real data under severe light conditions1.1https://4dqv.mpi-inf.mpg.de/Ev2Hands/
Christen Millerdurai, Diogo C. Luvizon, Viktor Rudnev, André Jonas, Jiayi Wang 0001, Christian Theobalt, Vladislav Golyanik
3DV2
2024 EventEgo3D: 3D Human Motion Capture from Egocentric Event Streams
abstract
Monocular egocentric 3D human motion capture is a challenging and actively researched problem. Existing methods use synchronously operating visual sensors (e.g. RGB cameras) and often fail under low lighting and fast motions, which can be restricting in many applications involving head-mounted devices. In response to the existing limitations, this paper 1) introduces a new problem, i.e. 3D human motion capture from an egocentric monoc-ular event camera with a fish eye lens, and 2) proposes the first approach to it called EventEgo 3D (EE3D). Event streams have high temporal resolution and provide reliable cues for 3D human motion capture under high-speed hu-man motions and rapidly changing illumination. The pro-posed EE3D framework is specifically tailored for learning with event streams in the LNES representation, enabling high 3D reconstruction accuracy. We also design a proto-type of a mobile head-mounted device with an event cam-era and record a real dataset with event observations and the ground-truth 3D human poses (in addition to the syn-thetic dataset). Our EE3D demonstrates robustness and su-perior 3D accuracy compared to existing solutions across various challenging experiments while supporting real-time 3D pose update rates of 140Hz.11https://4dqv.mpi-inf.mpg.de/EventEgo3D/
Christen Millerdurai, Hiroyasu Akada, Jian Wang 0042, Diogo C. Luvizon, Christian Theobalt, Vladislav Golyanik
CVPR4
2024 Holoported Characters: Real-Time Free-Viewpoint Rendering of Humans from Sparse RGB Cameras
abstract
We present the first approach to render highly realistic free-viewpoint videos of a human actor in general apparel, from sparse multi-view recording to display, in real-time at an unprecedented 4K resolution. At inference, our method only requires four camera views of the moving actor and the respective 3D skeletal pose. It handles actors in wide clothing, and reproduces even fine-scale dynamic detail, e.g. clothing wrinkles, face expressions, and hand gestures. At training time, our learning-based approach expects dense multi- view video and a rigged static surface scan of the actor. Our method comprises three main stages. Stage 1 is a skeleton-driven neural approach for high-quality capture of the detailed dynamic mesh geometry. Stage 2 is a novel solution to create a view-dependent texture using four test-time camera views as input. Finally, stage 3 comprises a new image-based refinement network rendering the final 4K image given the output from the previous stages. Our approach establishes a new benchmark for real-time rendering resolution and quality using sparse input camera views, unlocking possibilities for immersive telepresence. Code and data is available on our project page.
Ashwath Shetty, Marc Habermann, Guoxing Sun 0001, Diogo C. Luvizon, Vladislav Golyanik, Christian Theobalt
CVPR4
2024 Egocentric Whole-Body Motion Capture with FisheyeViT and Diffusion-Based Motion Refinement
abstract
In this work, we explore egocentric whole-body motion capture using a single fisheye camera, which simultane-ously estimates human body and hand motion. This task presents significant challenges due to three factors: the lack of high-quality datasets, fisheye camera distortion, and hu-man body self-occlusion. To address these challenges, we propose a novel approach that leverages Fisheye ViT to ex-tract fisheye image features, which are subsequently con-verted into pixel-aligned 3D heatmap representations for 3D human body pose prediction. For hand tracking, we incorporate dedicated hand detection and hand pose esti-mation networks for regressing 3D hand poses. Finally, we develop a diffusion-based whole-body motion prior model to refine the estimated whole-body motion while accounting for joint uncertainties. To train these networks, we col-lect a large synthetic dataset, Ego WholeBody, comprising 840,000 high-quality egocentric images captured across a diverse range of whole-body motion sequences. Quantitative and qualitative evaluations demonstrate the effective-ness of our method in producing high-quality whole-body motion estimates from a single egocentric camera.
Jian Wang 0042, Zhe Cao 0003, Diogo C. Luvizon, Lingjie Liu, Kripasindhu Sarkar, Danhang Tang, Thabo Beeler, Christian Theobalt
CVPR3
2024 Relightable Neural Actor with Intrinsic Decomposition and Pose Control
Diogo C. Luvizon, Vladislav Golyanik, Adam Kortylewski, Marc Habermann, Christian Theobalt
ECCV (59)1
2023 Scene-Aware Egocentric 3D Human Pose Estimation
abstract
Egocentric 3D human pose estimation with a single head-mounted fisheye camera has recently attracted attention due to its numerous applications in virtual and augmented reality. Existing methods still struggle in challenging poses where the human body is highly occluded or is closely interacting with the scene. To address this issue, we propose a scene-aware egocentric pose estimation method that guides the prediction of the egocentric pose with scene constraints. To this end, we propose an egocentric depth estimation network to predict the scene depth map from a wide-view egocentric fisheye camera while mitigating the occlusion of the human body with a depth-inpainting network. Next, we propose a scene-aware pose estimation network that projects the 2D image features and estimated depth map of the scene into a voxel space and regresses the 3D pose with a V2V network. The voxel-based feature representation provides the direct geometric connection between 2D image features and scene geometry, and further facilitates the V2V network to constrain the predicted pose based on the estimated scene geometry. To enable the training of the aforementioned networks, we also generated a synthetic dataset, called EgoGTA, and an in-the-wild dataset based on EgoPW, called EgoPW-Scene. The experimental results of our new evaluation sequences show that the predicted 3D egocentric poses are accurate and physically plausible in terms of human-scene interaction, demonstrating that our method outperforms the state-of-the-art methods both quantitatively and qualitatively.
Jian Wang 0042, Diogo C. Luvizon, Weipeng Xu, Lingjie Liu, Kripasindhu Sarkar, Christian Theobalt
CVPR2
2023 Computational Design of Personalized Wearable Robotic Limbs
abstract
Wearable robotic limbs (WRLs) augment human capabilities through robotic structures that attach to the user’s body. While WRLs are intensely researched and various device designs have been presented, it remains difficult for non-roboticists to engage with this exciting field. We aim to empower interaction designers and application domain experts to explore novel designs and applications by rapidly prototyping personalized WRLs that are customized for different tasks, different body locations, or different users. In this paper, we present WRLKit, an interactive computational design approach that enables designers to rapidly prototype a personalized WRL without requiring extensive robotics and ergonomics expertise. The body-aware optimization approach starts by capturing the user’s body dimensions and dynamic body poses. Then, an optimized fabricable structure of the WRL is generated for a desired mounting location and workspace of the WRL, to fit the user’s body and intended task. The results of a user study and several implemented prototypes demonstrate the practical feasibility and versatility of WRLKit.
Artin Saberpour, Ata Otaran, Martin Schmitz 0001, Marie Muehlhaus, Rishabh Dabral, Diogo C. Luvizon, Azumi Maekawa, Masahiko Inami, Christian Theobalt, Jürgen Steimle
UIST6
2023 Scene-Aware 3D Multi-Human Motion Capture from a Single Camera
abstract
Abstract In this work, we consider the problem of estimating the 3D position of multiple humans in a scene as well as their body shape and articulation from a single RGB video recorded with a static camera. In contrast to expensive marker‐based or multi‐view systems, our lightweight setup is ideal for private users as it enables an affordable 3D motion capture that is easy to install and does not require expert knowledge. To deal with this challenging setting, we leverage recent advances in computer vision using large‐scale pre‐trained models for a variety of modalities, including 2D body joints, joint angles, normalized disparity maps, and human segmentation masks. Thus, we introduce the first non‐linear optimization‐based approach that jointly solves for the 3D position of each human, their articulated pose, their individual shapes as well as the scale of the scene. In particular, we estimate the scene depth and person scale from normalized disparity predictions using the 2D body joints and joint angles. Given the per‐frame scene depth, we reconstruct a point‐cloud of the static scene in 3D space. Finally, given the per‐frame 3D estimates of the humans and scene point‐cloud, we perform a space‐time coherent optimization over the video to ensure temporal, spatial and physical plausibility. We evaluate our method on established multi‐person 3D human pose benchmarks where we consistently outperform previous methods and we qualitatively demonstrate that our method is robust to in‐the‐wild conditions including challenging scenes with people of different sizes. Code: https://github.com/dluvizon/scene‐aware‐3d‐multi‐human
Diogo C. Luvizon, Marc Habermann, Vladislav Golyanik, Adam Kortylewski, Christian Theobalt
Comput. Graph. Forum1
2023 SSP-Net: Scalable sequential pyramid networks for real-Time 3D human pose regression
Diogo C. Luvizon, Hedi Tabia, David Picard
Pattern Recognit.1
2022 Estimating Egocentric 3D Human Pose in the Wild with External Weak Supervision
abstract
Egocentric 3D human pose estimation with a single fisheye camera has drawn a significant amount of attention recently. However, existing methods struggle with pose estimation from in-the-wild images, because they can only be trained on synthetic data due to the unavailability of large-scale in-the-wild egocentric datasets. Furthermore, these methods easily fail when the body parts are occluded by or interacting with the surrounding scene. To address the shortage of in-the-wild data, we collect a large-scale in-the-wild egocentric dataset called Egocentric Poses in the Wild (EgoPW). This dataset is captured by a head-mounted fisheye camera and an auxiliary external camera, which provides an additional observation of the human body from a third-person perspective during training. We present a new egocentric pose estimation method, which can be trained on the new dataset with weak external supervision. Specifically, we first generate pseudo labels for the EgoPW dataset with a spatio-temporal optimization method by incorporating the external-view supervision. The pseudo labels are then used to train an egocentric pose estimation network. To facilitate the network training, we propose a novel learning strategy to supervise the egocentric features with the high-quality features extracted by a pretrained external-view pose estimation model. The experiments show that our method predicts accurate 3D poses from a single in-the-wild egocentric image and outperforms the state-of-the-art methods both quantitatively and qualitatively.
Jian Wang 0042, Lingjie Liu, Weipeng Xu, Kripasindhu Sarkar, Diogo C. Luvizon, Christian Theobalt
CVPR5
2022 Consensus-Based Optimization for 3D Human Pose Estimation in Camera Coordinates
Diogo C. Luvizon, David Picard, Hedi Tabia
Int. J. Comput. Vis.1
2021 Learning multiplane images from single views with self-supervision
Gustavo Sutter 0002, Diogo C. Luvizon, Antonio Joia, André G. C. Pacheco, Otávio A. B. Penatti
BMVC2
2021 Pyramidal Layered Scene Inference with Image Outpainting for Monocular View Synthesis
Marcos Roberto e Souza, Jhonatas Santos de Jesus Conceição, Jose L. Flores-Campana, Luis G. L. Decker, Diogo C. Luvizon, Gustavo Sutter 0002, Helena Almeida Maia, Hélio Pedrini
CAIP (1)5
2021 Improving Deep Learning Sound Events Classifiers Using Gram Matrix Feature-Wise Correlations
abstract
In this paper, we propose a new Sound Event Classification (SEC) method which is inspired in recent works for out-of-distribution detection. In our method, we analyse all the activations of a generic CNN in order to produce feature representations using Gram Matrices. The similarity metrics are evaluated considering all possible classes, and the final prediction is defined as the class that minimizes the deviation with respect to the features seeing during training. The proposed approach can be applied to any CNN and our experimental evaluation of four different architectures on two datasets demonstrated that our method consistently improves the baseline models.
Antonio Joia, André G. C. Pacheco, Diogo C. Luvizon
ICASSP3
2021 Adaptive Multiplane Image Generation from a Single Internet Picture
abstract
In the last few years, several works have tackled the problem of novel view synthesis from stereo images or even from a single picture. However, previous methods are computationally expensive, specially for high-resolution images. In this paper, we address the problem of generating a multiplane image (MPI) from a single high-resolution picture. We present the adaptive-MPI representation, which allows rendering novel views with low computational requirements. To this end, we propose an adaptive slicing algorithm that produces an MPI with a variable number of image planes. We present a new lightweight CNN for depth estimation, which is learned by knowledge distillation from a larger network. Occluded regions in the adaptive-MPI are inpainted also by a lightweight CNN. We show that our method is capable of producing high-quality predictions with one order of magnitude less parameters compared to previous approaches. The robustness of our method is evidenced on challenging pictures from the Internet.
Diogo C. Luvizon, Gustavo Sutter 0002, Andreza A. dos Santos, Jhonatas Santos de Jesus Conceição, Jose L. Flores-Campana, Luis G. L. Decker, Marcos Roberto e Souza, Hélio Pedrini, Antonio Joia, Otávio A. B. Penatti
WACV1
2021 Multi-Task Deep Learning for Real-Time 3D Human Pose Estimation and Action Recognition
abstract
Human pose estimation and action recognition are related tasks since both problems are strongly dependent on the human body representation and analysis. Nonetheless, most recent methods in the literature handle the two problems separately. In this article, we propose a multi-task framework for jointly estimating 2D or 3D human poses from monocular color images and classifying human actions from video sequences. We show that a single architecture can be used to solve both problems in an efficient way and still achieves state-of-the-art or comparable results at each task while running with a throughput of more than 100 frames per second. The proposed method benefits from high parameters sharing between the two tasks by unifying still images and video clips processing in a single pipeline, allowing the model to be trained with data from different categories simultaneously and in a seamlessly way. Additionally, we provide important insights for end-to-end training the proposed multi-task model by decoupling key prediction parts, which consistently leads to better accuracy on both tasks. The reported results on four datasets (MPII, Human3.6M, Penn Action and NTU RGB+D) demonstrate the effectiveness of our method on the targeted tasks. Our source code and trained weights are publicly available at https://github.com/dluvizon/deephar.
Diogo C. Luvizon, David Picard, Hedi Tabia
IEEE Trans. Pattern Anal. Mach. Intell.1
2020 Parallax Motion Effect Generation Through Instance Segmentation And Depth Estimation
abstract
Stereo vision is a growing topic in computer vision due to the innumerable opportunities and applications this technology offers for the development of modern solutions, such as virtual and augmented reality applications. To enhance the user's experience in three-dimensional virtual environments, the motion parallax estimation is a promising technique to achieve this objective. In this paper, we propose an algorithm for generating parallax motion effects from a single image, taking advantage of state-of-the-art instance segmentation and depth estimation approaches. This work also presents a comparison against such algorithms to investigate the trade-off between efficiency and quality of the parallax motion effects, taking into consideration a multi-task learning network capable of estimating instance segmentation and depth estimation at once. Experimental results and visual quality assessment indicate that the PyD-Net network (depth estimation) combined with Mask R-CNN or FBNet networks (instance segmentation) can produce parallax motion effects with good visual quality.
Allan Pinto, Manuel Alberto Cordova Neira, Luis G. L. Decker, Jose L. Flores-Campana, Marcos Roberto e Souza, Andreza A. dos Santos, Jhonatas Santos de Jesus Conceição, Henrique F. Gagliardi, Diogo C. Luvizon, Ricardo da Silva Torres, Hélio Pedrini
ICIP9
2019 Human pose regression by combining indirect part detection and contextual information
Diogo C. Luvizon, Hedi Tabia, David Picard
Comput. Graph.1
2018 2D/3D Pose Estimation and Action Recognition Using Multitask Deep Learning
abstract
Action recognition and human pose estimation are closely related but both problems are generally handled as distinct tasks in the literature. In this work, we propose a multitask framework for jointly 2D and 3D pose estimation from still images and human action recognition from video sequences. We show that a single architecture can be used to solve the two problems in an efficient way and still achieves state-of-the-art results. Additionally, we demonstrate that optimization from end-to-end leads to significantly higher accuracy than separated learning. The proposed architecture can be trained with data from different categories simultaneously in a seamlessly way. The reported results on four datasets (MPII, Human3.6M, Penn Action and NTU) demonstrate the effectiveness of our method on the targeted tasks.
Diogo C. Luvizon, David Picard, Hedi Tabia
CVPR1
2017 Learning features combination for human action recognition from skeleton sequences
Diogo C. Luvizon, Hedi Tabia, David Picard
Pattern Recognit. Lett.1
2017 A Video-Based System for Vehicle Speed Measurement in Urban Roadways
abstract
In this paper, we propose a nonintrusive video-based system for vehicle speed measurement in urban roadways. Our system uses an optimized motion detector and a novel text detector to efficiently locate vehicle license plates in image regions containing motion. Distinctive features are then selected on the license plate regions, tracked across multiple frames, and rectified for perspective distortion. Vehicle speed is measured by comparing the trajectories of the tracked features to known real-world measures. The proposed system was tested on a data set containing approximately 5 h of videos recorded in different weather conditions by a single low-cost camera, with associated ground truth speeds obtained by an inductive loop detector. Our data set is freely available for research purposes. The measured speeds have an average error of -0.5 km/h, staying inside the [-3, +2] km/h limit determined by regulatory authorities in several countries in over 96.0% of the cases. To the authors' knowledge, there are no other video-based systems able to achieve results comparable to those produced by an inductive loop detector. We also show that our license plate detector outperforms two other published state-of-the-art text detectors, as well as a well-known license plate detector, achieving a precision of 0.93 and a recall of 0.87.
Diogo C. Luvizon, Bogdan Tomoyuki Nassu, Rodrigo Minetto
IEEE Trans. Intell. Transp. Syst.1
2014 Vehicle speed estimation by license plate detection and tracking
abstract
We describe a novel system for vehicle speed estimation from videos captured in urban roadways. Our system uses text detection to locate the license plates of passing vehicles, which are then used to select stable features for tracking. The tracked features are then filtered and rectified for perspective distortion. Vehicle speed is estimated by comparing the trajectory of the tracked features to known real world measures. In experiments performed on videos captured under real operation conditions, our system attained a precision of 0.87 and a recall of 0.92 for license plate detection. Vehicle speeds were estimated with an average error of 0.59 km/h, staying inside the +2/-3 km/h limit, determined by regulatory authorities in several countries, in over 75% of the cases.
Diogo C. Luvizon, Bogdan Tomoyuki Nassu, Rodrigo Minetto
ICASSP1