EDBT 2026 Demo / reviewers in the wild / expert
Peng Wang 0024
dblp:95/4442-24
· DBLP profile ↗
31ranked-venue papers
3as first author
13since 2021 · last 2025
0000-0002-8265-9866ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 2 since 2021Systems, architecture and hardware · 7 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Refer and Grasp: Vision-Language Guided Continuous Dexterous GraspingabstractRobotic grasping guided by natural language instructions faces challenges due to ambiguities in object descriptions and the need to interpret complex spatial context. Existing visual grounding methods often rely on datasets that fail to capture these complexities, particularly when object categories are vague or undefined. To address these challenges, we make three key contributions. First, we present an automated dataset generation engine for visual grounding in tabletop grasping, combining procedural scene synthesis with template-based referring expression generation, requiring no manual labeling. Second, we introduce the RefGrasp dataset, featuring diverse indoor environments and linguistically challenging expressions for robotic grasping tasks. Third, we propose a visually grounded dexterous grasping framework with continuous grasp generation, validated through extensive real-world robotic experiments. Our work offers a novel approach for language-guided robotic manipulation, providing both a challenging dataset and an effective grasping framework for real-world applications. Project website: https://refer-and-grasp.github.io. Yayu Huang, Dongxuan Fan, Wen Qi 0002, Daheng Li, Yongkang Luo 0001, Jia Sun 0008, Peng Wang 0024 |
IROS | 9 |
| 2025 | Active Semi-supervised Continual Learning for Robotic Object Recognition
Sangzhou Xia, Xiangli Nie, Peng Wang 0024 |
PRCV (1) | 3 |
| 2025 | Reactive Human-to-Robot Dexterous Handovers for Anthropomorphic HandabstractHuman-robot object handovers are essential for robots to effectively serve human needs in various domains of human–robot interaction and collaboration, yet remain a significant challenge. Remarkable progress has been made by parallel-jaw gripper robots in grasp generation and motion planning for handovers, while few studies address this issue using anthropomorphic hands, necessitating the ability to handle higher collision probabilities and lower approaching space under occluded situations. In this article, we present a reactive human-to-robot dexterous handover framework for anthropomorphic hands. The closed-loop framework employs an effective collision detection and grasp selection approach to ensure safe and smooth motion in unstructured environments. We implement a handover system using a UR5 robot arm and a Schunk SVH Hand based on the presented framework, which can react to human motion during the handover process, generalize to diverse objects with different 6-DoF poses, and execute suitable grasp configurations. The generalizability, reliability, and efficiency of our method are demonstrated through the handover of 30 novel objects, a system ablation study for submodule evaluation, and a user study assessment involving eight participants. Haonan Duan 0001, Peng Wang 0024, Daheng Li, Wei Wei 0062, Yongkang Luo 0001, Guoqiang Deng |
IEEE Trans. Robotics | 2 |
| 2024 | Learning Realistic and Reasonable Grasps for Anthropomorphic Hand in Cluttered ScenesabstractGrasping is one of the most fundamental skills for humans to interact with objects. However, it remains a challenging problem for anthropomorphic hands, due to the lack of object affordance understanding and high-dimensional grasp planning. In this work, we propose an anthropomorphic hand grasping framework to learn realistic and reasonable grasps in cluttered scenes, which tackles the problem in three items: 1) graspable point segmentation; 2) hand grasp generation and 3) grasp optimization. Specifically, our method generates high-quality hand grasps efficiently without complete object models by learning graspable points, associated grasp configurations from observed point cloud in a parallel manner and optimizing predicted grasps based on hand-object contacts. Simulation experiments show that our model generates physical plausible grasps for the anthropomorphic hand effectively with over 70% success rate. Real-world experiments demonstrate that the model trained in simulation performs satisfactorily in real-world scenarios for unseen objects. Haonan Duan 0001, Daheng Li, Wei Wei 0062, Yayu Huang, Peng Wang 0024 |
ICRA | 6 |
| 2024 | Learning Human-Like Functional Grasping for Multifinger Hands From Few DemonstrationsabstractThis article investigates the challenge of enabling multifinger hands to perform human-like functional grasping for various intentions. However, accomplishing functional grasping in real robot hands present many challenges, including handling generalization ability for kinematically diverse robot hands, generating intention-conditioned grasps for a large variety of objects, and incomplete perception from a single-view camera. In this work, we first propose a six-step functional grasp synthesis algorithm based on fine-grained contact modeling. With the fine-grained contact-based optimization and learned dense shape correspondence, the algorithm is adaptable to various objects of the same category and a wide range of multifinger hands using few demonstrations. Second, over 10 k functional grasps are synthesized to train our neural network, named DexFG-Net, which generates intention-conditioned grasps based on reconstructed object. Extensive experiments in the simulation and physical grasps indicate that the grasp synthesis algorithm can produce human-like functional grasp with robust stability and functionality, and the DexFG-Net can generate plausible and human-like intention-conditioned grasping postures for anthropomorphic hands. Wei Wei 0062, Peng Wang 0024, Yongkang Luo 0001, Wanyi Li 0002, Daheng Li, Yayu Huang, Haonan Duan 0001 |
IEEE Trans. Robotics | 2 |
| 2023 | IPC-Net: Incomplete point cloud classification network based on data augmentation and similarity measurement
Yunqian He, Zhi Zhang 0008, Yongkang Luo 0001, Wanyi Li 0002, Peng Wang 0024 |
J. Vis. Commun. Image Represent. | 7 |
| 2023 | OS-DS tracker: Orientation-variant Siamese 3D tracking with Detection based Sampling
Lipeng Wang 0002, Wanyi Li 0002, Zhi Zhang 0008, Peng Wang 0024 |
Pattern Recognit. Lett. | 7 |
| 2022 | HGC-Net: Deep Anthropomorphic Hand Grasping in ClutterabstractGrasping in cluttered environments is one of the most fundamental skills in robotic manipulation. Most of the current works focus on estimating grasp poses for parallel-jaw or suction-cup end effectors. However, the study for dexterous anthropomorphic hand grasping in clutter remains a great challenge. In this paper, we propose HGC-Net, a single-shot network that learns to predict dense hand grasp configurations in clutter from single-view point cloud input. Our end-to-end neural network can predict hand grasp proposals efficiently and effectively. To enhance generalization, we built a large-scale synthetic grasping dataset with 179 household objects, 5K cluttered scenes and over 10M hand annotations. Experiments in simulation show that our model can predict dense and robust hand grasps and clear over 78% of unseen objects in clutter without any post-processing and outperform baseline methods by a large margin. Experiments on the real robot platform also demonstrate that the model trained on synthetic data performs well in natural environments. Code is available at https://github.com/yimingli1998/hgc_net. Wei Wei 0062, Daheng Li, Peng Wang 0024, Wanyi Li 0002 |
ICRA | 4 |
| 2022 | Global Mask R-CNN for marine ship instance segmentation
Yongkang Luo 0001, Wanyi Li 0002, Zhi Zhang 0008, Peng Wang 0024 |
Neurocomputing | 7 |
| 2021 | GPR: Grasp Pose Refinement Network for Cluttered ScenesabstractObject grasping in cluttered scenes is a widely investigated field of robot manipulation. Most of the current works focus on estimating grasp pose from point clouds based on an efficient single-shot grasp detection network. However, due to the lack of geometry awareness of the local grasping area, it may cause severe collisions and unstable grasp configurations. In this paper, we propose a two-stage grasp pose refinement network which detects grasps globally while fine-tuning low-quality grasps and filtering noisy grasps locally. Furthermore, we extend the 6-DoF grasp with an extra dimension as grasp width which is critical for collisionless grasping in cluttered scenes. It takes a single-view point cloud as input and predicts dense and precise grasp configurations. To enhance the generalization ability, we build a synthetic single-object grasp dataset including 150 commodities of various shapes, and a complex multi-object cluttered scene dataset including 100k point clouds with robust, dense grasp poses and mask annotations. Experiments conducted on Yumi IRB-1400 Robot demonstrate that the model trained on our dataset performs well in real environments and outperforms previous methods by a large margin. Wei Wei 0062, Yongkang Luo 0001, Fuyu Li, Guangyun Xu, Wanyi Li 0002, Peng Wang 0024 |
ICRA | 7 |
| 2021 | POIS: Policy-Oriented Instance Segmentation for Ambidextrous Robot PickingabstractRobots with a parallel-jaw gripper and suction cup is an adaptive and efficient robotic picking system. This paper proposed Policy-Oriented Instance Segmentation (POIS) for ambidextrous robots. POIS can generate a pair of target masks that allows ambidextrous robots to pick in parallel. It takes a depth image and predicts initial mask, center offset, and policy confidence map through three paralleled branches. We incorporate the initial mask with center offset to obtain candidate instances, from which we select masks of target objects for policy execution (decided with policy confidence map). We also provide a dataset that contains 6k synthetic scenes and 100 real scenes for ambidextrous picking. Trained on synthetic scenes, POIS generalizes well in real scene and is capable of handling novel objects in cluttered scenes. Our dataset and video are available at https://bit.ly/3oJj8Tu. Guangyun Xu, Peng Wang 0024, Yongkang Luo 0001 |
ICRA | 4 |
| 2021 | Simultaneous Semantic and Collision Learning for 6-DoF Grasp Pose EstimationabstractGrasping in cluttered scenes has always been a great challenge for robots, due to the requirement of the ability to well understand the scene and object information. Previous works usually assume that the geometry information of the objects is available, or utilize a step-wise, multi-stage strategy to predict the feasible 6-DoF grasp poses. In this work, we propose to formalize the 6-DoF grasp pose estimation as a simultaneous multi-task learning problem. In a unified framework, we jointly predict the feasible 6-DoF grasp poses, instance semantic segmentation, and collision information. The whole framework is jointly optimized and end-to-end differentiable. Our model is evaluated on large-scale benchmarks as well as the real robot system. On the public dataset, our method outperforms prior state-of-the-art methods by a large margin (+4.08 AP). We also demonstrate the implementation of our model on a real robotic platform and show that the robot can accurately grasp target objects in cluttered scenarios with a high success rate. Project link: https://openbyterobotics.github.io/sscl. Tao Kong, Ruihang Chu, Peng Wang 0024, Lei Li 0005 |
IROS | 5 |
| 2021 | DVFENet: Dual-branch voxel feature extraction network for 3D object detection
Yunqian He, Guihua Xia, Yongkang Luo 0001, Zhi Zhang 0008, Wanyi Li 0002, Peng Wang 0024 |
Neurocomputing | 7 |
| 2020 | An integrated ship segmentation method based on discriminator and extractor
Xujie He, Wanyi Li 0002, Zhi Zhang 0008, Yongkang Luo 0001, Peng Wang 0024 |
Image Vis. Comput. | 7 |
| 2019 | S3OD: Single Stage Small Object Detector from Scratch for Remote Sensing Images
Feng Yang 0001, Wentong Li 0001, Wanyi Li 0002, Peng Wang 0024 |
ICIG (3) | 4 |
| 2019 | Multi-Scale Object Detection in Satellite Imagery Based On YOLTabstractMulti-scale object detection (MOD) is one of the remaining challenges for satellite imagery. To improve the performance of MOD task, YOLT (You Only Look Twice) has achieved a good accuracy in high resolution remote sensing images. Motivated by the state-of-art object detection method for satellite imagery, we explored and achieved the state-of-the-art accuracy based on the standard YOLT for MOD task by providing a novel method with enough experimental results and model comparison on the typical multi-scale satellite imagery dataset. First, we divide objects into three categories according to the scale of objects. Then, different training strategies are used to train the classifier and detector for different scale objects. Finally, multi-scale detection chips are stitched and fused to get more accurate localization and classification as the final predicted results for MOD in satellite imagery. Experiments have been conducted over dataset from the second stage of AIIA1Cup Competition of Typical Object Recognition for Satellite Imagery in Small Samples compared with the standard YOLT and Faster R-CNN, which demonstrates the effectiveness and the comparable detection performance of our proposed pipeline. Wentong Li 0001, Wanyi Li 0002, Feng Yang 0001, Peng Wang 0024 |
IGARSS | 4 |
| 2019 | Salient object detection based on an efficient End-to-End Saliency Regression Network
Xuanyang Xi, Yongkang Luo 0001, Peng Wang 0024, Hong Qiao |
Neurocomputing | 3 |
| 2018 | Efficient beyond-birthday-bound secure authenticated encryption modes
Ping Zhang 0020, Honggang Hu, Peng Wang 0024 |
Sci. China Inf. Sci. | 3 |
| 2017 | Partially Decoupled Image-Based Visual Servoing Using Different Sensitive FeaturesabstractA new image-based visual servoing method based on sensitive features is presented to separately realize the position control and orientation control. Line features are used for the orientation control because of their sensitivities to rotational motions. Point features and area size features are employed to realize the position control since area size features are very sensitive to the objects' depths. The translations resulting from rotational motions are introduced into the position control as the compensation in order to eliminate the influence of the camera's motions on the point features. The depths for all active features are estimated via interaction matrices, features variations, and the executed camera motions. The proposed method can keep the tracked objects in the camera's field of view in the visual servoing process. In addition, the determination methods of the interaction matrices for point, line, and area size features are proposed. Comparing to the traditional method, the proposed determination method of the interaction matrix for line is independent from the parameters of the plane containing the line. Experimental results verify the effectiveness of the proposed methods. De Xu, Jinyan Lu, Peng Wang 0024, Zhengtao Zhang, Zi-ze Liang |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2016 | Adaptive probabilistic tracking with discriminative feature selection for mobile robotabstractObject tracking is one of the important tasks for mobile robot, and developing a robust and real-time visual tracking algorithm which can adaptively capture the varying appearance of target under challenging conditions for mobile robot is still an open problem. The main challenges of visual tracking for mobile robot come from variation of target's appearance and disturbance of environment. To cope with these problems, one of the most important topics is how to select the best tracking features. In this paper, we propose a novel adaptive probabilistic tracking method with discriminative feature selection for mobile robot Different from the existing adaptive tracking algorithms which select the discriminative features in a finite feature set, the proposed method treats feature selection as an estimation problem of the best feature tunable parameters in a continuous space. The estimation of the best tunable parameters and object tracking are implemented via different particle filters with novel observation models. A novel target model updating strategy is also proposed to adapt to the varying appearance of target and resist gradual drift. Experiments show the robustness of the proposed method under challenging conditions. Peng Wang 0024, Yongkang Luo 0001, Wanyi Li 0002, Hong Qiao |
SMC | 1 |
| 2016 | Top-down visual attention integrated particle filter for robust object tracking
Wanyi Li 0002, Peng Wang 0024, Hong Qiao |
Signal Process. Image Commun. | 2 |
| 2015 | Robust object tracking guided by top-down spectral analysis visual attention
Wanyi Li 0002, Peng Wang 0024, Hong Qiao |
Neurocomputing | 2 |
| 2015 | Sparse-Distinctive Saliency DetectionabstractIn this letter, we propose a novel saliency model for saliency detection, named sparse-distinctive (SD) saliency model. Different from the existing models that only consider sparsity or distinctness of image, the proposed model computes saliency based on sparsity and distinctness. The basic idea is that sparsity and distinctness contribute to saliency simultaneously and play different roles under different scenes. This sparse-distinctive saliency model is based on some key ideas introduced in this letter and supported by psychological evidence. Experimental results on public benchmark eye-tracking datasets show that considering the sparsity and distinctness for saliency can improve the accuracy of predicting human fixations, and the proposed model outperforms the mainstream models on predicting human fixations. Yongkang Luo 0001, Peng Wang 0024, Hong Qiao |
IEEE Signal Process. Lett. | 2 |
| 2014 | Visual Tracking via Saliency Weighted Sparse Coding Appearance ModelabstractSparse coding has been used for target appearance modeling and applied successfully in visual tracking. However, noise may be inevitably introduced into the representation due to background clutter. To cope with this problem, we propose a saliency weighted sparse coding appearance model for visual tracking. Firstly, a spectral filtering based visual attention computational model, which combines both bottom-up and top-down visual attention, is proposed to calculate saliency map. Secondly, pooling operation in sparse coding is weighted by calculated saliency map to help target representation focus on distinctive features and suppress background clutter. Extensive experiments on a recently proposed tracking benchmark demonstrate that the proposed algorithm outperforms state-of-the-art methods in tracking objects under background clutter. Wanyi Li 0002, Peng Wang 0024, Hong Qiao |
ICPR | 2 |
| 2014 | Salient region detection based on local and global saliencyabstractA new and effective salient region detection method based on local and global saliency information is proposed. To keep the completeness of salient regions, the input image is segmented into several regions firstly. Then for each region, local saliency and global saliency are generated respectively. The local saliency is computed by multi-scale neighborhood contrast, and the global saliency is measured according to global spatial distribution and inter-region isolation of features. Based on the local saliency and global saliency, the final saliency can be obtained by the weighted combination of them. The comparison experiment results demonstrate the effective performance of the proposed algorithm on salient region detection. Peng Wang 0024, Hong Qiao |
ICRA | 1 |
| 2014 | Introducing Memory and Association Mechanism Into a Biologically Inspired Visual ModelabstractA famous biologically inspired hierarchical model (HMAX model), which was proposed recently and corresponds to V1 to V4 of the ventral pathway in primate visual cortex, has been successfully applied to multiple visual recognition tasks. The model is able to achieve a set of position- and scale-tolerant recognition, which is a central problem in pattern recognition. In this paper, based on some other biological experimental evidence, we introduce the memory and association mechanism into the HMAX model. The main contributions of the work are: 1) mimicking the active memory and association mechanism and adding the top down adjustment to the HMAX model, which is the first try to add the active adjustment to this famous model and 2) from the perspective of information, algorithms based on the new model can reduce the computation storage and have a good recognition performance. The new model is also applied to object recognition processes. The primary experimental results show that our method is efficient with a much lower memory requirement. Hong Qiao, Yinlin Li, Peng Wang 0024 |
IEEE Trans. Cybern. | 4 |
| 2013 | A biologically inspired model of emotion eliciting from visual stimuli
Dongchun Ren, Peng Wang 0024, Hong Qiao, Suiwu Zheng |
Neurocomputing | 2 |
| 2012 | Sub-pattern bilinear model and its application in pose estimation of work-pieces
Zhicai Ou, Peng Wang 0024, Jianhua Su, Hong Qiao |
Neurocomputing | 2 |
| 2012 | Part-based adaptive detection of workpieces using differential evolution
Peng Wang 0024, Hong Qiao |
Signal Process. | 2 |
| 2012 | Corrigendum to "Part-based adaptive detection of workpieces using differential evolution" [Signal Processing 92 (2012) 301-307]
Peng Wang 0024, Hong Qiao |
Signal Process. | 2 |
| 2011 | Online Appearance Model Learning and Generation for Adaptive Visual TrackingabstractSeveral adaptive visual tracking algorithms have been recently proposed to capture the varying appearance of target. However, adaptability may also result in the problem of gradual drift, especially when the target appearance changes drastically. This paper gives some theoretical principles for online learning of target model, and then presents a novel adaptive tracking algorithm which is able to effectively cope with drastic variations in target appearance and resist gradual drift. Once target is localized in each frame, the patches sampled from target observation are first classified into foreground and background using an effective classifier. Then the adaptive, pure and time-continuous target model is extracted online through two processes: absorption process and rejection process, through which only the reliable features with high separability are absorbed in the new target model, while the “dangerous” features which may cause interfusion of background patterns are rejected. To minimize the influence of background and keep the temporal continuity of target model, two collaborative models dominant model and continuous model are designed. The proposed learning and generation mechanisms of target model are finally embedded in an adaptive tracking system. Experimental results demonstrate the robust performance of the proposed algorithm under challenging conditions. Peng Wang 0024, Hong Qiao |
IEEE Trans. Circuits Syst. Video Technol. | 1 |