Yin Wang 0005

dblp:70/4828-5 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
5since 2021 · last 2026
0009-0005-6088-3794ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
3 papers
Computer animation and physical simulation · 95% Multimedia analysis and retrieval · 5%
Artificial intelligence
2 papers
3D vision · 39% Generative modeling · 34% Deep learning architectures and training · 13%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer animation and physical simulation › motion synthesis
human motion synthesis
1.522025
MOST: Motion Diffusion Model for Rare Text via Temporal Clip Banzhaf Interaction · IEEE Trans. Vis. Comput. Graph. 2025
Fg-T2M: Fine-Grained Text-Driven Human Motion Generation via Diffusion Model · ICCV 2023
Computer animation and physical simulation › motion synthesis › human motion synthesis
text-to-motion generation
1.522025
MOST: Motion Diffusion Model for Rare Text via Temporal Clip Banzhaf Interaction · IEEE Trans. Vis. Comput. Graph. 2025
Fg-T2M: Fine-Grained Text-Driven Human Motion Generation via Diffusion Model · ICCV 2023
Computer animation and physical simulation
human-scene interaction
1.012026
Dynamic Worlds, Dynamic Humans: Generating Virtual Human-Scene Interaction Motion in Dynamic Scenes · IEEE Trans. Vis. Comput. Graph. 2026
Computer animation and physical simulation
motion synthesis
1.012026
Dynamic Worlds, Dynamic Humans: Generating Virtual Human-Scene Interaction Motion in Dynamic Scenes · IEEE Trans. Vis. Comput. Graph. 2026
Machine learning › Generative modeling › diffusion model
human motion generation
0.912025
Fg-T2M++: LLMs-Augmented Fine-Grained Text Driven Human Motion Generation · Int. J. Comput. Vis. 2025
Machine learning › Generative modeling › motion generation
text-driven motion generation
0.912025
Fg-T2M++: LLMs-Augmented Fine-Grained Text Driven Human Motion Generation · Int. J. Comput. Vis. 2025
Computer vision › 3D vision
3d reconstruction
0.712023
Dynamic Hyperbolic Attention Network for Fine Hand-object Reconstruction · ICCV 2023
Machine learning › Graph learning › graph neural network
graph convolution
0.712023
Dynamic Hyperbolic Attention Network for Fine Hand-object Reconstruction · ICCV 2023
Computer vision › 3D vision › 3d reconstruction › object reconstruction
hand-object reconstruction
0.712023
Dynamic Hyperbolic Attention Network for Fine Hand-object Reconstruction · ICCV 2023
Machine learning › Deep learning architectures and training › attention mechanism
hyperbolic attention
0.712023
Dynamic Hyperbolic Attention Network for Fine Hand-object Reconstruction · ICCV 2023
Computer vision › 3D vision › 3d reconstruction
surface reconstruction
0.712023
Dynamic Hyperbolic Attention Network for Fine Hand-object Reconstruction · ICCV 2023
Multimedia analysis and retrieval › multimedia retrieval › content-based retrieval
motion retrieval
0.312025
MOST: Motion Diffusion Model for Rare Text via Temporal Clip Banzhaf Interaction · IEEE Trans. Vis. Comput. Graph. 2025

Methods — techniques the papers use, named apart from their topics

diffusion model · 2.5world model · 1.0hierarchical experience memory · 1.0temporal clip banzhaf interaction · 0.9motion prompt module · 0.9large language model · 0.9linguistics-structure assisted module · 0.7hyperbolic space · 0.7graph neural network · 0.7graph convolution · 0.7context-aware progressive reasoning · 0.7attention · 0.7
YearPublicationVenuePosition
2026 Dynamic Worlds, Dynamic Humans: Generating Virtual Human-Scene Interaction Motion in Dynamic Scenes
abstract
Scenes are continuously undergoing dynamic changes in the real world. However, existing human-scene interaction generation methods typically treat the scene as static, which deviates from reality. Inspired by world models, we introduce Dyn-HSI, the first cognitive architecture for dynamic human-scene interaction, which endows virtual humans with three humanoid components. (1) Vision (human eyes): we equip the virtual human with a Dynamic Scene-Aware Navigation, which continuously perceives changes in the surrounding environment and adaptively predicts the next waypoint. (2) Memory (human brain): we equip the virtual human with a Hierarchical Experience Memory, which stores and updates experiential data accumulated during training. This allows the model to leverage prior knowledge during inference for context-aware motion priming, thereby enhancing both motion quality and generalization. (3) Control (human body): we equip the virtual human with Human-Scene Interaction Diffusion Model, which generates high-fidelity interaction motions conditioned on multimodal inputs. To evaluate performance in dynamic scenes, we extend the existing static human-scene interaction datasets to construct a dynamic benchmark, Dyn-Scenes. We conduct extensive qualitative and quantitative experiments to validate Dyn-HSI, showing that our method consistently outperforms existing approaches and generates high-quality human-scene interaction motions in both static and dynamic settings.
Yin Wang 0005, Zhiying Leng, Frederick W. B. Li, Xiaohui Liang 0001
IEEE Trans. Vis. Comput. Graph.1
2025 Fg-T2M++: LLMs-Augmented Fine-Grained Text Driven Human Motion Generation
Yin Wang 0005, Zhiying Leng, Frederick W. B. Li, Xiaohui Liang 0001
Int. J. Comput. Vis.1
2025 MOST: Motion Diffusion Model for Rare Text via Temporal Clip Banzhaf Interaction
abstract
We introduce MOST, a novel MOtion diffuSion model via Temporal clip Banzhaf interaction, aimed at addressing the persistent challenge of generating human motion from rare language prompts. While previous approaches struggle with coarse-grained matching and overlook important semantic cues due to motion redundancy, our key insight lies in leveraging fine-grained clip relationships to mitigate these issues. MOST's retrieval stage presents the first formulation of its kind - temporal clip Banzhaf interaction - which precisely quantifies textual-motion coherence at the clip level. This facilitates direct, fine-grained text-to-motion clip matching and eliminates prevalent redundancy. In the generation stage, a motion prompt module effectively utilizes retrieved motion clips to produce semantically consistent movements. Extensive evaluations confirm that MOST achieves state-of-the-art text-to-motion retrieval and generation performance by comprehensively addressing previous challenges, as demonstrated through quantitative and qualitative results highlighting its effectiveness, especially for rare prompts.
Yin Wang 0005, Zhiying Leng, Frederick W. B. Li, Xiaohui Liang 0001
IEEE Trans. Vis. Comput. Graph.1
2023 Dynamic Hyperbolic Attention Network for Fine Hand-object Reconstruction
abstract
Reconstructing both objects and hands in 3D from a single RGB image is complex. Existing methods rely on manually defined hand-object constraints in Euclidean space, leading to suboptimal feature learning. Compared with Euclidean space, hyperbolic space better preserves the geometric properties of meshes thanks to its exponentially-growing space distance, which amplifies the differences between the features based on similarity. In this work, we propose the first precise hand-object reconstruction method in hyperbolic space, namely Dynamic Hyperbolic Attention Network (DHANet), which leverages intrinsic properties of hyperbolic space to learn representative features. Our method that projects mesh and image features into a unified hyperbolic space includes two modules, i.e. dynamic hyperbolic graph convolution and image-attention hyperbolic graph convolution. With these two modules, our method learns mesh features with rich geometry-image multi-modal information and models better hand-object interaction. Our method provides a promising alternative for fine hand-object reconstruction in hyperbolic space. Extensive experiments on three public datasets demonstrate that our method outperforms most state-of-the-art methods.
Zhiying Leng, Mahdi Saleh, Antonio Montanaro, Hao Yu 0010, Yin Wang 0005, Nassir Navab, Xiaohui Liang 0001, Federico Tombari
ICCV6
2023 Fg-T2M: Fine-Grained Text-Driven Human Motion Generation via Diffusion Model
abstract
Text-driven human motion generation in computer vision is both significant and challenging. However, current methods are limited to producing either deterministic or imprecise motion sequences, failing to effectively control the temporal and spatial relationships required to conform to a given text description. In this work, we propose a fine-grained method for generating high-quality, conditional human motion sequences supporting precise text description. Our approach consists of two key components: 1) a linguistics-structure assisted module that constructs accurate and complete language feature to fully utilize text information; and 2) a context-aware progressive reasoning module that learns neighborhood and overall semantic linguistics features from shallow and deep graph neural networks to achieve a multi-step inference. Experiments show that our approach outperforms text-driven motion generation methods on HumanML3D and KIT test sets and generates better visually confirmed motion to the text conditions.
Yin Wang 0005, Zhiying Leng, Frederick W. B. Li, Xiaohui Liang 0001
ICCV1