EDBT 2026 Demo / reviewers in the wild / expert
Denis Tomè
dblp:169/9835
· DBLP profile ↗
7ranked-venue papers
5as first author
3since 2021 · last 2024
0000-0001-5733-1079ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 2 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
3D vision · 41% Face, body and person analysis · 36% Video understanding and tracking · 23% | |
| Computer graphics and multimedia
3 papers |
Computer animation and physical simulation · 65% Virtual and augmented reality · 35% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
3d human pose estimation |
1.3 | 3 | 2023 | SelfPose: 3D Egocentric Pose Estimation From a Headset Mounted Camera · IEEE Trans. Pattern Anal. Mach. Intell. 2023 xR-EgoPose: Egocentric 3D Human Pose From an HMD Camera · ICCV 2019 Lifting from the Deep: Convolutional 3D Pose Estimation from a Single Image · CVPR 2017 |
Computer vision › Face, body and person analysis › human pose estimation › 3d pose estimation
egocentric pose estimation |
1.0 | 2 | 2023 | SelfPose: 3D Egocentric Pose Estimation From a Headset Mounted Camera · IEEE Trans. Pattern Anal. Mach. Intell. 2023 xR-EgoPose: Egocentric 3D Human Pose From an HMD Camera · ICCV 2019 |
Computer vision › Face, body and person analysis
human pose estimation |
1.0 | 2 | 2023 | SelfPose: 3D Egocentric Pose Estimation From a Headset Mounted Camera · IEEE Trans. Pattern Anal. Mach. Intell. 2023 xR-EgoPose: Egocentric 3D Human Pose From an HMD Camera · ICCV 2019 |
Computer vision › Video understanding and tracking
action recognition |
0.8 | 1 | 2024 | HumMUSS: Human Motion Understanding Using State Space Models · CVPR 2024 |
Computer vision › Video understanding and tracking › motion analysis
human motion understanding |
0.8 | 1 | 2024 | HumMUSS: Human Motion Understanding Using State Space Models · CVPR 2024 |
Computer vision › 3D vision
pose estimation |
0.8 | 1 | 2024 | HumMUSS: Human Motion Understanding Using State Space Models · CVPR 2024 |
Computer animation and physical simulation › audio-driven animation
speech-driven animation |
0.6 | 1 | 2022 | Speech Driven Tongue Animation · CVPR 2022 |
Virtual and augmented reality › immersive display
head-mounted display |
0.3 | 2 | 2023 | SelfPose: 3D Egocentric Pose Estimation From a Headset Mounted Camera · IEEE Trans. Pattern Anal. Mach. Intell. 2023 xR-EgoPose: Egocentric 3D Human Pose From an HMD Camera · ICCV 2019 |
Computer vision › 3D vision › pose estimation › rotation and translation estimation
2d-3d pose estimation |
0.3 | 1 | 2017 | Lifting from the Deep: Convolutional 3D Pose Estimation from a Single Image · CVPR 2017 |
Computer vision › Face, body and person analysis › human pose estimation
2d human pose estimation |
0.3 | 1 | 2017 | Lifting from the Deep: Convolutional 3D Pose Estimation from a Single Image · CVPR 2017 |
Computer vision › 3D vision › 3d human pose estimation
single-image 3d pose estimation |
0.3 | 1 | 2017 | Lifting from the Deep: Convolutional 3D Pose Estimation from a Single Image · CVPR 2017 |
Methods — techniques the papers use, named apart from their topics
synthetic data generation · 2.1encoder-decoder architecture · 2.1multi-branch decoder · 1.3state space model · 0.8self-supervised audio feature encoder · 0.6encoder-decoder network · 0.6probabilistic pose prior · 0.3convolutional neural network · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | HumMUSS: Human Motion Understanding Using State Space ModelsabstractUnderstanding human motion from video is essential for a range of applications, including pose estimation, mesh recovery and action recognition. While state-of-the-art methods predominantly rely on transformer-based architectures, these approaches have limitations in practical scenarios. Transformers are slower when sequentially predicting on a continuous stream of frames in real-time, and do not generalize to new frame rates. In light of these constraints, we propose a novel attention-free spatiotemporal model for human motion understanding building upon recent advancements in state space models. Our model not only matches the performance of transformer-based models in various motion understanding tasks but also brings added benefits like adaptability to different video frame rates and enhanced training speed when working with longer sequences of keypoints. Moreover, the proposed model supports both offline and real-time applications. For real-time sequential prediction, our model is both memory efficient and several times faster than transformer-based approaches while maintaining their high accuracy. Arnab Kumar Mondal, Stefano Alletto, Denis Tomè |
CVPR | 3 |
| 2023 | SelfPose: 3D Egocentric Pose Estimation From a Headset Mounted CameraabstractWe present a new solution to egocentric 3D body pose estimation from monocular images captured from a downward looking fish-eye camera installed on the rim of a head mounted virtual reality device. This unusual viewpoint leads to images with unique visual appearance, characterized by severe self-occlusions and strong perspective distortions that result in a drastic difference in resolution between lower and upper body. We propose a new encoder-decoder architecture with a novel multi-branch decoder designed specifically to account for the varying uncertainty in 2D joint locations. Our quantitative evaluation, both on synthetic and real-world datasets, shows that our strategy leads to substantial improvements in accuracy over state of the art egocentric pose estimation approaches. To tackle the severe lack of labelled training data for egocentric 3D pose estimation we also introduced a large-scale photo-realistic synthetic dataset. xR-EgoPose offers 383K frames of high quality renderings of people with diverse skin tones, body shapes and clothing, in a variety of backgrounds and lighting conditions, performing a range of actions. Our experiments show that the high variability in our new synthetic training corpus leads to good generalization to real world footage and to state of the art results on real world datasets with ground truth. Moreover, an evaluation on the Human3.6M benchmark shows that the performance of our method is on par with top performing approaches on the more classic problem of 3D human pose from a third person viewpoint. Denis Tomè, Thiemo Alldieck, Patrick Peluse, Gerard Pons-Moll, Lourdes Agapito, Hernán Badino, Fernando De la Torre |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Speech Driven Tongue AnimationabstractAdvances in speech driven animation techniques allow the creation of convincing animations for virtual characters solely from audio data. Many existing approaches focus on facial and lip motion and they often do not provide realistic animation of the inner mouth. This paper addresses the problem of speech-driven inner mouth animation. Obtaining performance capture data of the tongue and jaw from video alone is difficult because the inner mouth is only partially observable during speech. In this work, we introduce a large-scale speech and mocap dataset that focuses on capturing tongue, jaw, and lip motion. This dataset enables research using data-driven techniques to generate realistic inner mouth animation from speech. We then propose a deep-learning based method for accurate and generalizable speech to tongue and jaw animation, and evaluate several encoder-decoder network architectures and audio feature encoders. We find that recent self-supervised deep learning based audio feature encoders are robust, generalize well to unseen speakers and content, and work best for our task. To demonstrate the practical application of our approach, we show animations on high-quality parametric 3D face models driven by the landmarks generated from our speech-to-tongue animation method. Denis Tomè, Carsten Stoll, Mark K. Tiede, Kevin Munhall, Alex Hauptmann 0001, Iain A. Matthews |
CVPR | 2 |
| 2019 | xR-EgoPose: Egocentric 3D Human Pose From an HMD CameraabstractWe present a new solution to egocentric 3D body pose estimation from monocular images captured from a downward looking fish-eye camera installed on the rim of a head mounted virtual reality device. This unusual viewpoint, just 2 cm away from the user's face, leads to images with unique visual appearance, characterized by severe self-occlusions and strong perspective distortions that result in a drastic difference in resolution between lower and upper body. Our contribution is two-fold. Firstly, we propose a new encoder-decoder architecture with a novel dual branch decoder designed specifically to account for the varying uncertainty in the 2D joint locations. Our quantitative evaluation, both on synthetic and real-world datasets, shows that our strategy leads to substantial improvements in accuracy over state of the art egocentric pose estimation approaches. Our second contribution is a new large-scale photorealistic synthetic dataset - xR-EgoPose - offering 383K frames of high quality renderings ofpeople with a diversity of skin tones, body shapes, clothing, in a variety of backgrounds and lighting conditions, performing a range of actions. Our experiments show that the high variability in our new synthetic training corpus leads to good generalization to real world footage and to state of the art results on real world datasets with ground truth. Moreover, an evaluation on the Human3.6M benchmark shows that the performance of our method is on par with top performing approaches on the more classic problem of 3D human pose from a third person viewpoint. Denis Tomè, Patrick Peluse, Lourdes Agapito, Hernán Badino |
ICCV | 1 |
| 2018 | Rethinking Pose in 3D: Multi-stage Refinement and Recovery for Markerless Motion CaptureabstractWe propose a CNN-based approach for multi-camera markerless motion capture of the human body. Unlike existing methods that first perform pose estimation on individual cameras and generate 3D models as post-processing, our approach makes use of 3D reasoning throughout a multi-stage approach. This novelty allows us to use provisional 3D models of human pose to rethink where the joints should be located in the image and to recover from past mistakes. Our principled refinement of 3D human poses lets us make use of image cues, even from images where we previously misdetected joints, to refine our estimates as part of an end-to-end approach. Finally, we demonstrate how the high-quality output of our multi-camera setup can be used as an additional training source to improve the accuracy of existing single camera models. Denis Tomè, Matteo Toso, Lourdes Agapito, Chris Russell 0001 |
3DV | 1 |
| 2017 | Lifting from the Deep: Convolutional 3D Pose Estimation from a Single ImageabstractWe propose a unified formulation for the problem of 3D human pose estimation from a single raw RGB image that reasons jointly about 2D joint estimation and 3D pose reconstruction to improve both tasks. We take an integrated approach that fuses probabilistic knowledge of 3D human pose with a multi-stage CNN architecture and uses the knowledge of plausible 3D landmark locations to refine the search for better 2D locations. The entire process is trained end-to-end, is extremely efficient and obtains state-of-the-art results on Human3.6M outperforming previous approaches both on 2D and 3D errors. Denis Tomè, Chris Russell 0001, Lourdes Agapito |
CVPR | 1 |
| 2016 | Deep Convolutional Neural Networks for pedestrian detection
Denis Tomè, Federico Monti, Luca Baroffio, Luca Bondi, Marco Tagliasacchi, Stefano Tubaro |
Signal Process. Image Commun. | 1 |