EDBT 2026 Demo / reviewers in the wild / expert
Giorgio Becherini
dblp:363/7591
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2026
0009-0005-8770-8144ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Generative modeling · 49% 3D vision · 45% Video understanding and tracking · 6% | |
| Computer graphics and multimedia
4 papers |
Computer animation and physical simulation · 40% Image and video processing · 14% Rendering · 14% |
Topics — the 18 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer animation and physical simulation › gesture generation
co-speech gesture generation |
1.5 | 2 | 2024 | EMAGE: Towards Unified Holistic Co-Speech Gesture Generation via Expressive Masked Audio Gesture Modeling · CVPR 2024 Emotional Speech-Driven 3D Body Animation via Disentangled Latent Diffusion · CVPR 2024 |
Computer vision › 3D vision
3d human reconstruction |
0.9 | 1 | 2025 | Im2Haircut: Single-View Strand-Based Hair Reconstruction for Human Avatars · ICCV 2025 |
Machine learning › Generative modeling › video generation
diffusion-based video generation |
0.9 | 1 | 2025 | GenLit: Reformulating Single-Image Relighting as Video Generation · SIGGRAPH Asia 2025 |
Computer vision › 3D vision › 3d human reconstruction
human avatar reconstruction |
0.9 | 1 | 2025 | Im2Haircut: Single-View Strand-Based Hair Reconstruction for Human Avatars · ICCV 2025 |
Computer vision › 3D vision › pose estimation
human pose and motion estimation |
0.9 | 1 | 2025 | BEDLAM2.0: Synthetic humans and cameras in motion · NeurIPS 2025 |
Computer vision › 3D vision › 3d reconstruction
single-view 3d reconstruction |
0.9 | 1 | 2025 | Im2Haircut: Single-View Strand-Based Hair Reconstruction for Human Avatars · ICCV 2025 |
Machine learning › Generative modeling
video generation |
0.9 | 1 | 2025 | GenLit: Reformulating Single-Image Relighting as Video Generation · SIGGRAPH Asia 2025 |
Computer animation and physical simulation › motion capture
human motion capture |
0.9 | 1 | 2025 | BEDLAM2.0: Synthetic humans and cameras in motion · NeurIPS 2025 |
Image and video processing › motion analysis › human motion analysis
human motion estimation |
0.9 | 1 | 2025 | BEDLAM2.0: Synthetic humans and cameras in motion · NeurIPS 2025 |
Rendering
relighting |
0.9 | 1 | 2025 | GenLit: Reformulating Single-Image Relighting as Video Generation · SIGGRAPH Asia 2025 |
Computational photography and imaging › image relighting
single-image relighting |
0.9 | 1 | 2025 | GenLit: Reformulating Single-Image Relighting as Video Generation · SIGGRAPH Asia 2025 |
Machine learning › Generative modeling
generative model |
0.8 | 1 | 2024 | EMAGE: Towards Unified Holistic Co-Speech Gesture Generation via Expressive Masked Audio Gesture Modeling · CVPR 2024 |
Machine learning › Generative modeling › diffusion model
latent diffusion model |
0.8 | 1 | 2024 | Emotional Speech-Driven 3D Body Animation via Disentangled Latent Diffusion · CVPR 2024 |
Audio and music processing
speech processing |
0.8 | 1 | 2024 | EMAGE: Towards Unified Holistic Co-Speech Gesture Generation via Expressive Masked Audio Gesture Modeling · CVPR 2024 |
Computer vision › Video understanding and tracking › motion analysis
human motion analysis |
0.5 | 2 | 2024 | EMAGE: Towards Unified Holistic Co-Speech Gesture Generation via Expressive Masked Audio Gesture Modeling · CVPR 2024 Emotional Speech-Driven 3D Body Animation via Disentangled Latent Diffusion · CVPR 2024 |
Machine learning › Generative modeling
diffusion model |
0.3 | 1 | 2025 | GenLit: Reformulating Single-Image Relighting as Video Generation · SIGGRAPH Asia 2025 |
Machine learning › Generative modeling › diffusion model
video diffusion model |
0.3 | 1 | 2025 | GenLit: Reformulating Single-Image Relighting as Video Generation · SIGGRAPH Asia 2025 |
Visual content generation and editing
synthetic data generation |
0.3 | 1 | 2025 | BEDLAM2.0: Synthetic humans and cameras in motion · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
synthetic data rendering · 1.7synthetic data · 1.7score distillation · 1.7motion capture · 1.7fine-tuning · 1.7masked transformer · 1.5masked gesture reconstruction · 1.5latent diffusion · 1.5disentangled representation learning · 1.5VQ-VAE · 1.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NeuralFur: Animal Fur Reconstruction from Multi-View ImagesabstractReconstructing realistic animal fur geometry from images is a challenging task due to the fine-scale details, selfocclusion, and view-dependent appearance of fur. In contrast to human hairstyle reconstruction, there are also no datasets that can be leveraged to learn a fur prior for different animals. In this work, we present a first multi-viewbased method for high-fidelity 3D fur modeling of animals using a strand-based representation, leveraging the general knowledge of a vision language model. Given multiview RGB images, we first reconstruct a coarse surface geometry using traditional multi-view stereo techniques. We then use a vision language model (VLM) system to retrieve information about the realistic length structure of the fur for each part of the body. We use this knowledge to construct the animal's furless geometry and grow strands atop it. The fur reconstruction is supervised with both geometric and photometric losses computed from multi-view images. To mitigate orientation ambiguities stemming from the Gabor filters that are applied to the input images, we additionally utilize the VLM to guide the strands' growth direction and their relation to the gravity vector that we incorporate as a loss. With this new schema of using a VLM to guide 3D reconstruction from multi-view inputs, we show generalization across a variety of animals with different fur types. For additional results and code, please refer to https://neuralfur.is.tue.mpg.de. Vanessa Sklyarova, Berna Kabadayi, Anastasios Yiannakidis, Giorgio Becherini, Michael J. Black, Justus Thies |
3DV | 4 |
| 2025 | Im2Haircut: Single-View Strand-Based Hair Reconstruction for Human Avatars
Vanessa Sklyarova, Egor Zakharov, Malte Prinzler, Giorgio Becherini, Michael J. Black, Justus Thies |
ICCV | 4 |
| 2025 | BEDLAM2.0: Synthetic humans and cameras in motionabstractInferring 3D human motion from video remains a challenging problem with many applications. While traditional methods estimate the human in image coordinates, many applications require human motion to be estimated in world coordinates. This is particularly challenging when there is both human and camera motion. Progress on this topic has been limited by the lack of rich video data with ground truth human and camera movement. We address this with BEDLAM2.0, a new dataset that goes beyond the popular BEDLAM dataset in important ways. In addition to introducing more diverse and realistic cameras and camera motions, BEDLAM2.0 increases diversity and realism of body shape, motions, clothing, hair, and 3D environments. Additionally, it adds shoes, which were missing in BEDLAM. BEDLAM has become a key resource for training 3D human pose and motion regressors today and we show that BEDLAM2.0 is significantly better, particularly for training methods that estimate humans in world coordinates. We compare state-of-the art methods trained on BEDLAM and BEDLAM2.0, and find that BEDLAM2.0 significantly improves accuracy over BEDLAM. For research purposes, we provide the rendered videos, ground truth body parameters, and camera motions. We also provide the 3D assets to which we have rights and links to those from third parties. Joachim Tesch, Giorgio Becherini, Prerana Achar, Anastasios Yiannakidis, Muhammed Kocabas, Priyanka Patel, Michael J. Black |
NeurIPS | 2 |
| 2025 | GenLit: Reformulating Single-Image Relighting as Video GenerationabstractManipulating the illumination of a 3D scene within a single image represents a fundamental challenge in computer vision and graphics. This problem has traditionally been addressed using inverse rendering techniques, which involve explicit 3D asset reconstruction and costly ray-tracing simulations. Meanwhile, recent advancements in visual foundation models suggest that a new paradigm could soon be possible – one that replaces explicit physical models with networks that are trained on large amounts of image and video data. In this paper, we exploit the implicit scene understanding of a video diffusion model, particularly Stable Video Diffusion, to relight a single image. We introduce GenLit, a framework that distills the ability of a graphics engine to perform light manipulation into a video-generation model, enabling users to directly insert and manipulate a point light in the 3D world within a given image and generate results directly as a video sequence. We find that a model fine-tuned on only a small synthetic dataset generalizes to real-world scenes, enabling single-image relighting with plausible and convincing shadows and inter-reflections. Our results highlight the ability of video foundation models to capture rich information about lighting, material, and shape, and our findings indicate that such models, with minimal training, can be used to perform relighting without explicit asset reconstruction or ray-tracing. Shrisha Bharadwaj, Haiwen Feng, Giorgio Becherini, Victoria Fernández Abrevaya, Michael J. Black |
SIGGRAPH Asia | 3 |
| 2024 | Emotional Speech-Driven 3D Body Animation via Disentangled Latent DiffusionabstractExisting methods for synthesizing 3D human gestures from speech have shown promising results, but they do not explicitly model the impact of emotions on the generated gestures. Instead, these methods directly output animations from speech without control over the expressed emotion. To address this limitation, we present AMUSE, an emotional speech-driven body animation model based on latent diffusion. Our observation is that content (i.e., gestures related to speech rhythm and word utterances), emotion, and personal style are separable. To account for this, AMUSE maps the driving audio to three disentangled latent vectors: one for content, one for emotion, and one for personal style. A latent diffusion model, trained to generate gesture motion sequences, is then conditioned on these latent vectors. Once trained, AMUSE synthesizes 3D human gestures directly from speech with control over the expressed emotions and style by combining the content from the driving speech with the emotion and style of another speech sequence. Randomly sampling the noise of the diffusion model further generates variations of the gesture with the same emotional expressivity. Qualitative, quantitative, and perceptual evaluations demonstrate that AMUSE outputs realistic gesture sequences. Compared to the state of the art, the generated gestures are better synchronized with the speech content, and better represent the emotion expressed by the input speech. Our code is available at amuse.is.tue.mpg.de. Kiran Chhatre, Radek Danecek, Nikos Athanasiou, Giorgio Becherini, Christopher Peters 0001, Michael J. Black, Timo Bolkart |
CVPR | 4 |
| 2024 | EMAGE: Towards Unified Holistic Co-Speech Gesture Generation via Expressive Masked Audio Gesture ModelingabstractWe propose EMAGE, a framework to generate full-body human gestures from audio and masked gestures, encompassing facial, local body, hands, and global movements. To achieve this, we first introduce BEAT2 (BEAT-SMPLX-FLAME), a new mesh-level holistic co-speech dataset. BEAT2 combines a MoShed SMPL-X body with FLAME head parameters and further refines the modeling of head, neck, and finger movements, offering a community-standardized, high-quality 3D motion captured dataset. EMAGE leverages masked body gesture priors during training to boost inference performance. It involves a Masked Audio Gesture Transformer, facilitating joint training on audio-to-gesture generation and masked gesture reconstruction to effectively encode audio and body gesture hints. Encoded body hints from masked gestures are then separately employed to generate facial and body movements. Moreover, EMAGE adaptively merges speech features from the audio's rhythm and content and utilizes four compositional VQ-VAEs to enhance the results' fidelity and diversity. Experiments demonstrate that EMAGE generates holistic gestures with state-of-the-art performance and is flexible in accepting predefined spatial-temporal gesture inputs, generating complete, audio-synchronized results. Our code and dataset are available.1 Giorgio Becherini, Yichen Peng, Mingyang Su, Xuefei Zhe, Naoya Iwamoto, Michael J. Black |
CVPR | 3 |