Pengwei Xie

dblp:229/0254 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0003-1005-9252ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Rethinking 6-DoF grasp detection: A flexible framework for high-quality grasping
Pengwei Xie, Siang Chen, Kaiqin Yang, Guijin Wang
Pattern Recognit.1
2025 Region-Centric 6-Dof Grasp Detection: A Data-Efficient Solution for Cluttered Scenes
abstract
Robotic grasping, serving as the cornerstone of robot manipulation, is fundamental for embodied intelligence. Manipulation in challenging scenarios demands grasp detection algorithms with higher efficiency and generalizability. However, for general 6-Dof grasp detection, most data-driven methods directly extract scene-level features to generate grasp prediction, relying on a relatively heavy scene-level feature encoder and a significant amount of data with dense grasp labels for model training. In this letter, we propose a novel data-efficient 6-Dof grasp detection framework in cluttered scenes, named Region-Centric Grasp Detection (RCGD), consisting of an Iterative Search Module (ISM) and a Region Grasp Model (RGM). Concretely, ISM aims to retrieve potential region centers and aggregate multiple regions in a coarse-to-fine way. Then, RGM extracts aligned grasp-related embeddings and predicts grasps within these local regions. Benefiting from the region-centric paradigm and the training-free location strategy, RCGD significantly outperforms previous methods and shows minimal performance loss with even a very small portion of training data or labels. Furthermore, real-world robotic experiments in two distinct settings highlight the effectiveness of our method with a 95% success rate.
Siang Chen, Pengwei Xie, Dingchang Hu, Wenming Yang, Guijin Wang
IROS3
2025 FEG-VON: Frontier Embedding Graph for Efficient Visual Object Navigation
abstract
Visual object navigation, requiring agents to locate target objects in novel environments through egocentric visual observation, remains a critical challenge in Embodied AI. We propose FEG-VON, a training-free framework that constructs and maintains a Frontier Embedding Graph for efficient Visual Object Navigation. The graph initializes frontier embeddings using Vision Language Models (VLMs), where visual observations are encoded into spatially anchored semantic embeddings through cross-modal alignment with target text descriptors. We then update the graph by aggregating spatio-temporal semantic relations across frontiers, enabling online adaptation to new targets via similarity scoring without remapping. The evaluation results in public benchmarks demonstrate the superior performance of FEG-VON in both single- and multi-object navigation tasks compared with state-of-the-art methods. Crucially, FEG-VON eliminates dependency on task-specific training for exploration and advances the feasibility of zero-shot navigation in open-world environments.
Yingru Dai, Pengwei Xie, Yikai Liu, Siang Chen, Wenming Yang, Guijin Wang
IROS2
2025 A General Zero-Shot Joint Training Framework for Pansharpening
Pengwei Xie, Kangqing Shen, Xingce Wang, Zhongke Wu
PRCV (15)1
2024 GenH2R: Learning Generalizable Human-to-Robot Handover via Scalable Simulation, Demonstration, and Imitation
abstract
This paper presents GenH2R, a framework for learning generalizable vision-based human-to-robot (H2R) handover skills. The goal is to equip robots with the ability to reliably receive objects with unseen geometry handed over by humans in various complex trajectories. We acquire such generalizability by learning H2R handover at scale with a comprehensive solution including procedural simulation assets creation, automated demonstration generation, and effective imitation learning. We leverage large-scale 3D model repositories, dexterous grasp generation methods, and curve-based 3D animation to create an H2R handover simulation environment named GenH2R-Sim, surpassing the number of scenes in existing simulators by three orders of magnitude. We further introduce a distillation-friendly demonstration generation method that automati-cally generates a million high-quality demonstrations suitable for learning. Finally, we present a 4D imitation learning method augmented by a future forecasting objective to distill demonstrations into a visuo-motor handover policy. Experimental evaluations in both simulators and the real world demonstrate significant improvements (at least +10% success rate) over baselines in all cases.
Junyu Chen 0003, Ziqing Chen, Pengwei Xie, Rui Chen 0019, Li Yi 0001
CVPR4
2024 Category-Agnostic Pose Estimation for Point Clouds
abstract
The goal of object pose estimation is to visually determine the pose of a specific object in the RGB-D input. Unfortunately, when faced with new categories, both instance-based and category-based methods are unable to deal with unseen objects of unseen categories, which is a challenge for pose estimation. To address this issue, this paper proposes a method to introduce geometric features for pose estimation of point clouds without requiring category information. The method is based only on the patch feature of the point cloud, a geometric feature with rotation invariance. After training without category information, our method achieves as good results as other category-based methods. Our method successfully achieved pose annotation of no category information instances on the CAMERA25 dataset and ModelNet40 dataset.
Siang Chen, Pengwei Xie, Guijin Wang
ICIP4
2023 ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills
Jiayuan Gu, Fanbo Xiang, Zhan Ling, Xiqiang Liu, Tongzhou Mu, Yihe Tang, Stone Tao, Xinyue Wei, Yunchao Yao, Xiaodi Yuan, Pengwei Xie, Zhiao Huang, Rui Chen 0019, Hao Su 0001
ICLR12
2022 Multi-scale Cross-Modal Transformer Network for RGB-D Object Detection
Pengwei Xie, Guijin Wang
MMM (1)2
2021 FETNet: Feature Exchange Transformer Network for RGB-D Object Detection
Jing-Hao Xue, Pengwei Xie, Guijin Wang
BMVC3
2021 Non-Local Aggregation for RGB-D Semantic Segmentation
abstract
Exploiting both RGB (2D appearance) and Depth (3D geometry) information can improve the performance of semantic segmentation. However, due to the inherent difference between the RGB and Depth information, it remains a challenging problem in how to integrate RGB-D features effectively. In this letter, to address this issue, we propose a Non-local Aggregation Network (NANet), with a well-designed Multi-modality Non-local Aggregation Module (MNAM), to better exploit the non-local context of RGB-D features at multi-stage. Compared with most existing RGB-D semantic segmentation schemes, which only exploit local RGB-D features, the MNAM enables the aggregation of non-local RGB-D information along both spatial and channel dimensions. The proposed NANet achieves comparable performances with state-of-the-art methods on popular RGB-D benchmarks, NYUDv2 and SUN-RGBD.
Guodong Zhang 0004, Jing-Hao Xue, Pengwei Xie, Sifan Yang, Guijin Wang
IEEE Signal Process. Lett.3
2020 Weakly Supervised Segmentation Guided Hand Pose Estimation During Interaction with Unknown Objects
abstract
Hand pose estimation is important for human computer interaction, but the performance is not satisfying when the hand is interacting with objects. To alleviate the influence of unknown objects, we propose a novel weakly supervised segmentation guided scheme to estimate hand poses. Approximate hand masks generated from annotations of sparse hand joints are used to supervise the segmentation task. Better features can be extracted since they are shared between the two tasks of hand segmentation and hand pose estimation. With the guidance of weakly supervised segmentation, the network can learn intermediate features balanced between focusing on the foreground and preserving contextual information. Finally the xy and z coordinates are estimated in different branches but utilizing shared feature maps. Experimental results of three different tasks on the publicly available FHAD dataset demonstrate the effectiveness of the proposed architecture.
Cairong Zhang, Guijin Wang, Xinghao Chen 0001, Pengwei Xie, Toshihiko Yamasaki
ICASSP4