Pengxiang Zhu

dblp:186/1770 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 4 since 2021Systems, architecture and hardware · 3 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
3D vision · 64% Robot navigation and mapping · 36%
Computer graphics and multimedia
1 paper
Computer animation and physical simulation · 67% Virtual and augmented reality · 33%
Human-computer interaction and pervasive computing
1 paper
Human-robot interaction · 100%

Topics — the 6 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › 3d human reconstruction
hand mesh reconstruction
0.912025
Multi-View Hand Reconstruction With a Point-Embedded Transformer · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Computer vision › 3D vision › 3d reconstruction
multi-view stereo
0.912025
Multi-View Hand Reconstruction With a Point-Embedded Transformer · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Computer animation and physical simulation › motion synthesis
human motion synthesis
0.812024
FAVOR: Full-Body AR-Driven Virtual Object Rearrangement Guided by Instruction Text · AAAI 2024
Computer animation and physical simulation
motion synthesis
0.812024
FAVOR: Full-Body AR-Driven Virtual Object Rearrangement Guided by Instruction Text · AAAI 2024
Robotics › Robot navigation and mapping › localization › multi-robot localization
cooperative localization
0.512021
Cooperative Visual-Inertial Odometry · ICRA 2021
Robotics › Robot navigation and mapping › visual odometry
visual-inertial odometry
0.512021
Cooperative Visual-Inertial Odometry · ICRA 2021

Methods — techniques the papers use, named apart from their topics

motion capture · 1.5instruction-guided motion generation · 1.5point-embedded transformer · 0.9multi-view fusion · 0.9multi-state constraint kalman filter · 0.5covariance intersection · 0.5
YearPublicationVenuePosition
2025 Multi-View Hand Reconstruction With a Point-Embedded Transformer
abstract
This work introduces a novel and generalizable multi-view Hand Mesh Reconstruction (HMR) model, named POEM, designed for practical use in real-world hand motion capture scenarios. The advances of the POEM model consist of two main aspects. First, concerning the modeling of the problem, we propose embedding a static basis point within the multi-view stereo space. A point represents a natural form of 3D information and serves as an ideal medium for fusing features across different views, given its varied projections across these views. Consequently, our method harnesses a simple yet effective idea: a complex 3D hand mesh can be represented by a set of 3D basis points that 1) are embedded in the multi-view stereo, 2) carry features from the multi-view images, and 3) encompass the hand in it. The second advance lies in the training strategy. We utilize a combination of five large-scale multi-view datasets and employ randomization in the number, order, and poses of the cameras. By processing such a vast amount of data and a diverse array of camera configurations, our model demonstrates notable generalizability in the real-world applications. As a result, POEM presents a highly practical, plug-and-play solution that enables user-friendly, cost-effective multi-view motion capture for both left and right hands.
Lixin Yang 0001, Licheng Zhong, Pengxiang Zhu, Xinyu Zhan 0001, Junxiao Kong, Jian Xu 0027, Cewu Lu
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 FAVOR: Full-Body AR-Driven Virtual Object Rearrangement Guided by Instruction Text
abstract
Rearrangement operations form the crux of interactions between humans and their environment. The ability to generate natural, fluid sequences of this operation is of essential value in AR/VR and CG. Bridging a gap in the field, our study introduces FAVOR: a novel dataset for Full-body AR-driven Virtual Object Rearrangement that uniquely employs motion capture systems and AR eyeglasses. Comprising 3k diverse motion rearrangement sequences and 7.17 million interaction data frames, this dataset breaks new ground in research data. We also present a pipeline FAVORITE for producing digital human rearrangement motion sequences guided by instructions. Experimental results, both qualitative and quantitative, suggest that this dataset and pipeline deliver high-quality motion sequences. Our dataset, code, and appendix are available at https://kailinli.github.io/FAVOR.
Kailin Li 0001, Lixin Yang 0001, Zenan Lin, Jian Xu 0027, Xinyu Zhan 0001, Yifei Zhao 0003, Pengxiang Zhu, Wenxiong Kang, Kejian Wu, Cewu Lu
AAAI7
2021 Cooperative Visual-Inertial Odometry
abstract
This paper studies the problem of multi-robot cooperative visual-inertial localization where each robot is equipped with only a single camera and IMU. We develop two cooperative visual-inertial odometry (C-VIO) algorithms within the multi-state constraint Kalman filter (MSCKF) framework, in which each robot utilizes not only its own measurements but constraints of common features co-observed with its neighbors within the current sliding window in order to improve the localization accuracy. The first centralized-equivalent algorithm tracks the robot-to-robot cross correlations and prioritizes the pose accuracy while requiring full capacity communication among all the robots during update. The second distributed algorithm ignores the robot-to-robot cross correlations to obtain a scalable, robust and efficient fully distributed structure where each robot only keeps its own states and communicates with its neighbors, while a covariance intersection (CI)-based update strategy is leveraged to guarantee consistency. The proposed algorithms are validated extensively in both Monte-Carlo simulations and real-world datasets, and shown to be able to achieve better accuracy with competitive efficiency.
Pengxiang Zhu, Wei Ren 0001, Guoquan Huang 0001
ICRA1
2021 Distributed Visual-Inertial Cooperative Localization
abstract
In this paper we present a consistent and distributed state estimator for multi-robot cooperative localization (CL) which efficiently fuses environmental features and loop-closure constraints across time and robots. In particular, we leverage covariance intersection (CI) to allow each robot to only estimate its own state and autocovariance and compensate for the unknown correlations between robots. Two novel multi-robot methods for utilizing common environmental SLAM features are introduced and evaluated in terms of accuracy and efficiency. Moreover, we adapt CI to enable drift-free estimation through the use of loop-closure measurement constraints to other robots’ historical poses without a significant increase in computational cost. The proposed distributed CL estimator is validated against its non-realtime centralized counterpart extensively in both simulations and real-world experiments.
Pengxiang Zhu, Patrick Geneva, Wei Ren 0001, Guoquan Huang 0001
IROS1
2020 Multi-Robot Joint Visual-Inertial Localization and 3-D Moving Object Tracking
abstract
In this paper, we present a novel distributed algorithm to track a moving object's state by utilizing a heterogenous mobile robot network in a three-dimensional (3-D) environment, wherein the robots' poses (positions and orientations) are unknown. Each robot is equipped with a monocular camera and an inertial measurement unit (IMU), and has the ability to communicate with its neighbors. Rather than assuming a known common global frame for all the robots (which is often the case in the literature regarding multi-robot systems), we allow each robot to perform motion estimation locally. For localization, we propose a multi-robot visual-inertial navigation systems (VINS) where one robot builds a prior map and then the map is used to bound the long-term drifts of the visual-inertial odometry (VIO) running on the other robots. Moreover, a novel distributed Kalman filter is introduced and employed to cooperatively track the six degree-of-freedom (6-DoF) motion of the object which is represented as a point cloud. Further, the object can be totally invisible to some robots during the tracking period. The proposed algorithm is extensively validated in Monte-Carlo simulations.
Pengxiang Zhu, Wei Ren 0001
IROS1