Jiepeng Wang 0005

dblp:405/2803 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2026
0000-0002-6049-4458ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
2 papers
Computer animation and physical simulation · 57% Visual content generation and editing · 33% Multimedia analysis and retrieval · 10%
Artificial intelligence
2 papers
Generative modeling · 70% Video understanding and tracking · 30%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visual content generation and editing › video generation
controllable video generation
1.012026
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding · AAAI 2026
Computer animation and physical simulation › motion synthesis
human motion synthesis
0.912025
Uni-Inter: Unifying 3D Human Motion Synthesis Across Diverse Interaction Contexts · SIGGRAPH Asia 2025
Computer animation and physical simulation
motion synthesis
0.912025
Uni-Inter: Unifying 3D Human Motion Synthesis Across Diverse Interaction Contexts · SIGGRAPH Asia 2025
Machine learning › Generative modeling
diffusion model
0.312026
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding · AAAI 2026
Machine learning › Generative modeling › diffusion model
video diffusion model
0.312026
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding · AAAI 2026
Multimedia analysis and retrieval
video understanding
0.312026
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding · AAAI 2026
Computer vision › Video understanding and tracking › dynamic scene analysis › video scene understanding › human-centric scene understanding
human-scene interaction
0.312025
Uni-Inter: Unifying 3D Human Motion Synthesis Across Diverse Interaction Contexts · SIGGRAPH Asia 2025

Methods — techniques the papers use, named apart from their topics

diffusion model · 2.0adaptive modality control · 2.0volumetric representation · 1.7probabilistic prediction · 1.7
YearPublicationVenuePosition
2026 OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding
abstract
In this paper, we propose a novel framework for controllable video diffusion, OmniVDiff , aiming to synthesize and comprehend multiple video visual content in a single diffusion model. To achieve this, OmniVDiff treats all video visual modalities in the color space to learn a joint distribution, while employing an adaptive control strategy that dynamically adjusts the role of each visual modality during the diffusion process, either as a generation modality or a conditioning modality. Our framework supports three key capabilities: (1) Text-conditioned video generation, where all modalities are jointly synthesized from a textual prompt; (2) Video understanding, where structural modalities are predicted from rgb inputs in a coherent manner; and (3) X-conditioned video generation, where video synthesis is guided by finegrained inputs such as depth, canny and segmentation. Extensive experiments demonstrate that OmniVDiff achieves state-of-the-art performance in video generation tasks and competitive results in video understanding. Its flexibility and scalability make it well-suited for downstream applications such as video-to-video translation, modality adaptation for visual tasks, and scene reconstruction.
Dianbing Xi, Jiepeng Wang 0005, Yuanzhi Liang, Xi Qiu, Yuchi Huo, Rui Wang 0004, Chi Zhang 0012, Xuelong Li 0001
AAAI2
2025 Uni-Inter: Unifying 3D Human Motion Synthesis Across Diverse Interaction Contexts
abstract
We present Uni-Inter, a unified framework for human motion generation that supports a wide range of interaction scenarios: including human-human, human-object, and human-scene—within a single, task-agnostic architecture. In contrast to existing methods that rely on task-specific designs and exhibit limited generalization, Uni-Inter introduces the Unified Interactive Volume (UIV), a volumetric representation that encodes heterogeneous interactive entities into a shared spatial field. This enables consistent relational reasoning and compound interaction modeling. Motion generation is formulated as joint-wise probabilistic prediction over the UIV, allowing the model to capture fine-grained spatial dependencies and produce coherent, context-aware behaviors. Experiments across three representative interaction tasks demonstrate that Uni-Inter achieves competitive performance and generalizes well to novel combinations of entities. These results suggest that unified modeling of compound interactions offers a promising direction for scalable motion synthesis in complex environments.
Sheng Liu 0013, Yuanzhi Liang, Jiepeng Wang 0005, Sidan Du, Chi Zhang 0067, Xuelong Li 0001
SIGGRAPH Asia3