David Forsyth

dblp:405/4763 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Video understanding and tracking · 44% Face, body and person analysis · 44% 3D vision · 13%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking › human motion prediction
hand motion prediction
0.912025
How Do I Do That? Synthesizing 3D Hand Motion and Contacts for Everyday Interactions · CVPR 2025
Computer vision › Face, body and person analysis › human pose estimation › articulated pose estimation
hand pose estimation
0.912025
How Do I Do That? Synthesizing 3D Hand Motion and Contacts for Everyday Interactions · CVPR 2025
Computer vision › 3D vision › human body modeling
contact map prediction
0.312025
How Do I Do That? Synthesizing 3D Hand Motion and Contacts for Everyday Interactions · CVPR 2025

Methods — techniques the papers use, named apart from their topics

transformer decoder · 0.9codebook learning · 0.9VQ-VAE · 0.9
YearPublicationVenuePosition
2025 How Do I Do That? Synthesizing 3D Hand Motion and Contacts for Everyday Interactions
abstract
We tackle the novel problem of predicting 3D hand motion and contact maps (or Interaction Trajectories) given a single RGB view, action text, and a 3D contact point on the object as input. Our approach consists of (1) Interaction Codebook: a VQVAE model to learn a latent codebook of hand poses and contact points, effectively tokenizing interaction trajectories, (2) Interaction Predictor: a transformer-decoder module to predict the interaction trajectory from test time inputs by using an indexer module to retrieve a latent affordance from the learned codebook. To train our model, we develop a data engine that extracts 3D hand poses and contact trajectories from the diverse HoloAssist dataset. We evaluate our model on a benchmark that is 2.5-10× larger than existing works, in terms of diversity of objects and interactions observed, and test for generalization of the model across object categories, action categories, tasks, and scenes. Experimental results show the effectiveness of our approach over transformer & diffusion baselines across all settings.
Ben Lundell, Dmitry Andreychuk, David Forsyth, Saurabh Gupta 0001, Harpreet Sawhney
CVPR4