Christian Diller

dblp:255/5074 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
3since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
3D vision · 36% Video understanding and tracking · 31% Face, body and person analysis · 24%
Computer graphics and multimedia
2 papers
Computer animation and physical simulation · 100%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
action anticipation
0.812024
FutureHuman3D: Forecasting Complex Long-Term 3D Human Behavior from Video Observations · CVPR 2024
Computer vision › Video understanding and tracking
human motion prediction
0.812024
FutureHuman3D: Forecasting Complex Long-Term 3D Human Behavior from Video Observations · CVPR 2024
Computer animation and physical simulation › motion synthesis
human motion synthesis
0.812024
CG-HOI: Contact-Guided 3D Human-Object Interaction Generation · CVPR 2024
Computer animation and physical simulation › human-object interaction
human-object interaction generation
0.812024
CG-HOI: Contact-Guided 3D Human-Object Interaction Generation · CVPR 2024
Computer vision › Face, body and person analysis › human pose estimation
3d pose forecasting
0.612022
Forecasting Characteristic 3D Poses of Human Actions · CVPR 2022
Computer vision › Face, body and person analysis
human pose estimation
0.612022
Forecasting Characteristic 3D Poses of Human Actions · CVPR 2022
Computer animation and physical simulation
human motion prediction
0.612022
Forecasting Characteristic 3D Poses of Human Actions · CVPR 2022
Computer vision › 3D vision
3d scene reconstruction
0.412020
SG-NN: Sparse Generative Neural Networks for Self-Supervised Scene Completion of RGB-D Scans · CVPR 2020
Computer vision › 3D vision › 3d scene understanding
scene completion
0.412020
SG-NN: Sparse Generative Neural Networks for Self-Supervised Scene Completion of RGB-D Scans · CVPR 2020
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training
0.212024
FutureHuman3D: Forecasting Complex Long-Term 3D Human Behavior from Video Observations · CVPR 2024
Machine learning › Generative modeling
diffusion model
0.212024
CG-HOI: Contact-Guided 3D Human-Object Interaction Generation · CVPR 2024
Computer vision › 3D vision › 3d reconstruction
surface reconstruction
0.112020
SG-NN: Sparse Generative Neural Networks for Self-Supervised Scene Completion of RGB-D Scans · CVPR 2020

Methods — techniques the papers use, named apart from their topics

diffusion model · 1.5cross-attention · 1.5probabilistic modeling · 1.1autoregressive sampling · 1.1differentiable 2d projection · 0.8autoregressive model · 0.8adversarial loss · 0.8sparse generative convolutional neural network · 0.4self-supervision · 0.4
YearPublicationVenuePosition
2024 CG-HOI: Contact-Guided 3D Human-Object Interaction Generation
abstract
We propose CG-HOI, the first method to address the task of generating dynamic 3D human-object interactions (HOIs) from text. We model the motion of both human and object in an interdependent fashion, as semantically rich human motion rarely happens in isolation without any interactions. Our key insight is that explicitly modeling contact between the human body surface and object geometry can be used as strong proxy guidance, both during training and inference. Using this guidance to bridge human and object motion enables generating more realistic and physically plausible interaction sequences, where the human body and corresponding object move in a coherent manner. Our method first learns to model human motion, object motion, and contact in a joint diffusion process, inter-correlated through cross-attention. We then leverage this learned contact for guidance during inference to synthesize realistic and coherent HOIs. Extensive evaluation shows that our joint contact-based human-object interaction approach generates realistic and physically plausible sequences, and we show two applications highlighting the capabilities of our method. Conditioned on a given object trajectory, we can generate the corresponding human motion without re-training, demonstrating strong human-object interdependency learning. Our approach is also flexible, and can be applied to static realworld 3D scene scans.
Christian Diller, Angela Dai
CVPR1
2024 FutureHuman3D: Forecasting Complex Long-Term 3D Human Behavior from Video Observations
abstract
We present a generative approach to forecast long-term future human behavior in 3D, requiring only weak supervision from readily available 2D human action data. This is a fundamental task enabling many downstream applications. The required ground-truth data is hard to capture in 3D (mocap suits, expensive setups) but easy to acquire in 2D (simple RGB cameras). Thus, we design our method to only require 2D RGB data at inference time while being able to generate 3D human motion sequences. We use a differentiable 2D projection scheme in an autoregressive manner for weak supervision, and an adversarial loss for 3D regularization. Our method predicts long and complex human behavior sequences (e.g., cooking, assembly) consisting of multiple sub-actions. We tackle this in a semantically hierarchical manner, jointly predicting high-level coarse action labels together with their low-level fine-grained real-izations as characteristic 3D human poses. We observe that these two action representations are coupled in nature, and joint prediction benefits both action and pose forecasting. Our experiments demonstrate the complementary nature of joint action and 3D pose prediction: our joint approach outperforms each task treated individually, enables robust longer-term sequence prediction, and improves over alter-native approaches to forecast actions and characteristic 3D poses.
Christian Diller, Thomas A. Funkhouser, Angela Dai
CVPR1
2022 Forecasting Characteristic 3D Poses of Human Actions
abstract
We propose the task of forecasting characteristic 3d poses: from a short sequence observation of a person, predict a future 3d pose of that person in a likely action-defining, characteristic pose - for instance, from observing a person picking up an apple, predict the pose of the person eating the apple. Prior work on human motion prediction estimates future poses at fixed time intervals. Although easy to define, this frame-by-frame formulation confounds temporal and intentional aspects of human action. Instead, we define a semantically meaningful pose prediction task that decouples the predicted pose from time, taking inspiration from goal-directed behavior. To predict characteristic poses, we propose a probabilistic approach that models the possible multimodality in the distribution of likely characteristic poses. We then sample future pose hypotheses from the predicted distribution in an autoregressive fashion to model dependencies between joints. To evaluate our method, we construct a dataset of manually annotated characteristic 3d poses. Our experiments with this dataset suggest that our proposed probabilistic approach outperforms state-of-the-art methods by 26% on average.
Christian Diller, Thomas A. Funkhouser, Angela Dai
CVPR1
2020 SG-NN: Sparse Generative Neural Networks for Self-Supervised Scene Completion of RGB-D Scans
abstract
We present a novel approach that converts partial and noisy RGB-D scans into high-quality 3D scene reconstructions by inferring unobserved scene geometry. Our approach is fully self-supervised and can hence be trained solely on incomplete, real-world scans. To achieve, self-supervision, we remove frames from a given (incomplete) 3D scan in order to make it even more incomplete; self-supervision is then formulated by correlating the two levels of partialness of the same scan while masking out regions that have never been observed. Through generalization across a large training set, we can then predict 3D scene completions even without seeing any 3D scan of entirely complete geometry. Combined with a new 3D sparse generative convolutional neural network architecture, our method is able to predict highly detailed surfaces in a coarse-to-fine hierarchical fashion that outperform existing state-of-the-art methods by a significant margin in terms of reconstruction quality.
Angela Dai, Christian Diller, Matthias Nießner
CVPR2