Sheng Liu 0013

dblp:03/5747-13 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0001-6612-5023ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Computer animation and physical simulation · 100%
Artificial intelligence
1 paper
Video understanding and tracking · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer animation and physical simulation › motion synthesis
human motion synthesis
0.912025
Uni-Inter: Unifying 3D Human Motion Synthesis Across Diverse Interaction Contexts · SIGGRAPH Asia 2025
Computer animation and physical simulation
motion synthesis
0.912025
Uni-Inter: Unifying 3D Human Motion Synthesis Across Diverse Interaction Contexts · SIGGRAPH Asia 2025
Computer vision › Video understanding and tracking › dynamic scene analysis › video scene understanding › human-centric scene understanding
human-scene interaction
0.312025
Uni-Inter: Unifying 3D Human Motion Synthesis Across Diverse Interaction Contexts · SIGGRAPH Asia 2025

Methods — techniques the papers use, named apart from their topics

volumetric representation · 1.7probabilistic prediction · 1.7
YearPublicationVenuePosition
2026 URPose: The model with unbiased rectified projection and reconstruction error for monocular unsupervised 3D human pose estimation
Sheng Liu 0013, Yang Li 0063, Sidan Du
Neurocomputing1
2025 Uni-Inter: Unifying 3D Human Motion Synthesis Across Diverse Interaction Contexts
abstract
We present Uni-Inter, a unified framework for human motion generation that supports a wide range of interaction scenarios: including human-human, human-object, and human-scene—within a single, task-agnostic architecture. In contrast to existing methods that rely on task-specific designs and exhibit limited generalization, Uni-Inter introduces the Unified Interactive Volume (UIV), a volumetric representation that encodes heterogeneous interactive entities into a shared spatial field. This enables consistent relational reasoning and compound interaction modeling. Motion generation is formulated as joint-wise probabilistic prediction over the UIV, allowing the model to capture fine-grained spatial dependencies and produce coherent, context-aware behaviors. Experiments across three representative interaction tasks demonstrate that Uni-Inter achieves competitive performance and generalizes well to novel combinations of entities. These results suggest that unified modeling of compound interactions offers a promising direction for scalable motion synthesis in complex environments.
Sheng Liu 0013, Yuanzhi Liang, Jiepeng Wang 0005, Sidan Du, Chi Zhang 0067, Xuelong Li 0001
SIGGRAPH Asia1
2025 ATHENA - Autonomous Vehicle Trajectory Planning Considered Human Action Awareness
abstract
Large language models have brought revolutionary changes to autonomous driving algorithms, ushering them into the era of multimodality. However, existing vehicle trajectory planning methods primarily focus on obstacle avoidance in autonomous driving scenarios, overlooking interactions with entities within the scene, such as humans. In this letter, we propose a new research direction: vehicle trajectory planning that takes into account human actions. We establish ATHENA, the first autonomous driving dataset that integrates multimodal human actions, comprising 33,855 scenarios. Each scenario contains status information about ego vehicle, as well as pedestrian actions that actively or passively interact with the vehicle, such as signaling the vehicle to proceed by waving and the unexpected falls by pedestrians that force the vehicle to stop. Based on each type of interaction, ATHENA also provides the corresponding driving suggestions and the reasons behind them. Moreover, we present an LLM-based baseline that consists of two agents: the Action Understanding Agent and the Vehicle Control Agent. Our baseline implements the generation of driving recommendations and vehicle control functions, which are guided by pedestrian actions. Experiments demonstrate the effectiveness and strong performance of our method. Our dataset and code will be publicly available at https://github.com/dogooooo/ATHENA.
Jinghao Cao, Sheng Liu 0013, Chaofan Wu, Yang Li 0063, Sidan Du
IEEE Signal Process. Lett.2
2024 CASSC: Context-aware method for depth guided semantic scene completion
abstract
Abstract Semantic scene completion is a crucial end‐to‐end 3D perception task, and the 3D information perception subjects is vital for autonomous driving. This paper presents CASSC, a novel adaptive context‐aware method based on Transformer networks, aimed at realizing camera‐based semantic scene completion algorithms. The key idea is to leverage rich context information from images to obtain pixel‐level label proposals, followed by designing a multiscale fusion mechanism to merge this information and match it with voxel space. A weakly supervised training strategy is proposed to obtain semantic label distribution features from images and introduce an adaptive multiscale fusion module to fuse and adaptively match these features with voxel space. Here, CASSC achieves state‐of‐the‐art performance on the SemanticKITTI dataset and demonstrates excellent performance on the SSC‐Bench dataset. Ablation experiments validate the rationality and effectiveness of our design, and the model and code of CASSC will be open‐sourced on https://github.com/dogooooo/CASSC .
Jinghao Cao, Ming Li 0069, Sheng Liu 0013, Yang Li 0063, Sidan Du
IET Image Process.3
2024 ARES: Text-Driven Automatic Realistic Simulator for Autonomous Traffic
abstract
The large-scale generation of real-world scenario datasets is a pivotal task in the field of autonomous driving. Existing methods emphasize solely on single-frame rendering, which need complex inputs for continuous scenario rendering. In this letter, ARES: a text-driven automatic realistic simulator is proposed, which can generate extensive realistic datasets with just a single text input. Its core idea is to generate vehicle trajectories based on the textual description, and then render the scenario by vehicle attributes associated with these trajectories. For learning trajectories generating, supervisory signal temporal logic is proposed to assist conditional diffusion model, which incorporates prior physical information. We annotate textual descriptions for KITTI-MOT dataset and establish an objective quantitative evaluation system. The superiority of our method is demonstrated by its high performance, which is reflected in a matching score of 3.54 and an FID of 8.93in the trajectory reconstruction task, along with a speed accuracy of 0.99 and a direction accuracy of 0.93in the trajectory editing task. The scenarios rendered by the proposed method exhibit high quality and realism, which indicates its great potential in testing of autonomous driving algorithms with vehicle-in-the-loop simulations.
Jinghao Cao, Sheng Liu 0013, Yang Li 0063, Sidan Du
IEEE Signal Process. Lett.2
2023 MMDA: Multi-person marginal distribution awareness for monocular 3D pose estimation
abstract
Abstract Most existing 3D pose representations cannot completely decouple the overlapping two or more human joints of the same type. In this paper, the authors propose a novel 2.5 D representation of the human pose by projecting human joints in 3D space onto the three orthogonal planes. The authors apply for the first time the permutation module to a multi‐person 3D human pose estimation task and use Geometric Constraints Loss (GCL) to guide the learning of the model. The authors overcome the negative effects of the inductive bias of convolutional neural networks (CNNs) by aligning the intermediate feature space with the output feature space. The effectiveness of the authors’ approach is validated on the carnegie mellon university (CMU) panoptic dataset and MuPoTS‐3D dataset. The authors’ proposed representations can effectively decouple the human joints in their selected data from overlapping human joints.
Sheng Liu 0013, Jianghai Shuai, Yang Li 0063, Sidan Du
IET Image Process.1