Pengwen Xiong

dblp:133/0644 · DBLP profile ↗
← Back
20ranked-venue papers
4as first author
20since 2021 · last 2025
0000-0002-0623-8592ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Task-Specific Embodied Tactile Sensing for Dexterous Hand
abstract
In order to obtain a good tactile sensing, traditional dexterous hands always enable all the sensing units installed on them all the time, even if just a few sensor units are actually used, which make the tactile sensing system resource-wasting and energy consuming. In order to reduce their complexities by placing the tactile sensing units only at critical locations, this work proposes an embodied tactile dexterous hand (ET-Hand) and a novel multimodal sensor placement framework that learns multiple tasks to generate optimal placement proposal. Furthermore, our ET-Hand can dynamically adjust the perceived tactile sensor positions, types and numbers during robotic manipulation, providing novel tools and methods for investigating the tactile channels and placement scale required for robot exploration. In the object recognition and slip detection tasks, the results show that our proposed method performs close to or even better than traditional sensing way with large-scale placement.
Pengwen Xiong, Aiguo Song
ICRA2
2025 A Natural Human-Robot Interaction System for Teleoperation Based on Noncontact Haptic Feedback
abstract
In order to provide natural and immersive interactive experience for teleoperation in the context of human-robot collaboration and interaction, this work introduces a natural human-robot interaction system for teleoperation based on ultrasonic haptic feedback. Specifically, our system can accurately capture an operator's hand movements and replicate these actions on the remote robot with low latency and high fidelity. It utilizes an ultrasonic phased array to achieve non-contact haptic feedback. We propose a dynamic ultrasonic array acoustic field customization method based on interactive feature information image. This method can dynamically adjust the acoustic field according to the operator's hand characteristics, focus on multiple target points in real time, and project them onto an operator's fingertips, thereby providing force-controllable non-contact haptic feedback to the operator. The operator is integrated into the feedback loop of our system, controlling the system through multimodal feedback to form a high-quality human-in-the-loop closed control system. The system's performance is validated in two classic robotic tasks: block pick-and-place and nut-tightening. The experimental results show that the system exhibits excellent accuracy and dexterity, and can efficiently complete tasks with high accuracy while providing great interactive experience for operators.
Letian Wei, Pengwen Xiong, Aiguo Song, MengChu Zhou
IROS2
2025 Dual-Modal Magnetic Skin for Robust Tactile Sensing
abstract
Traditional magnetic tactile sensors are highly susceptible to external magnetic field interference, limiting their reliability in practical applications. To address this challenge, we propose a dual-modal soft magnetic skin capable of simultaneously acquiring magnetic and force tactile information across spatiotemporal domains, inspired by the sensory mechanisms of human skin. The system integrates a Convolutional Neural Network-Convolutional Neural Network-Multilayer Perceptron (CNN-CNN-MLP) architecture to fuse these dual-modal signals effectively. Furthermore, we introduce a novel Dynamic Weighting Coefficient Layer (DWCL) to dynamically optimize fusion weights for each modality based on real-time input characteristics, thereby enhancing robustness against magnetic interference. The DWCL leverages temporal discrepancies between modalities during pre-contact sensing and quantifies the magnetic field strength of target objects to autonomously adjust fusion ratios, prioritizing the more reliable modality under varying interference conditions. Extensive experimental evaluations demonstrate that the proposed DWCL significantly improves interference resistance compared to conventional fusion methods, advancing the feasibility of magnetic tactile sensing in real-world environments.
Pengwen Xiong, Huan Peng, Aiguo Song, Peter Xiaoping Liu
IROS1
2025 Diffusion-based vision-language model for zero-shot anomaly detection in medical images
abstract
With the rapid advancement of diagnostic technology, the ability to detect pathological areas such as tumors and polyps has significantly improved. This progress provides medical imaging specialists with more precise visual information to support anomaly identification, diagnosis, treatment planning, and patient monitoring. However, existing unsupervised and semi-supervised anomaly detection methods struggle with data privacy constraints, limited annotated medical datasets, and challenges in generalization. Zero-Shot Anomaly Detection (ZSAD), which enables the detection of unseen categories without requiring class-specific training, has emerged as a promising solution by leveraging the vision-language alignment capabilities of Vision-Language Models (VLMs), such as Contrastive Language-Image Pretraining (CLIP). Despite recent progress, ZSAD remains hindered by high noise levels, sparse targets, and poor adaptability in complex medical imaging scenarios. To address these issues, we propose a novel framework: DiffusionCLIP, a diffusion-based VLM for zero-shot anomaly detection in two-dimensional medical images. Specifically, DiffusionCLIP integrates diffusion models into the VLM to progressively denoise multi-level features extracted from the CLIP visual encoder, enhancing feature robustness and discriminability. A multi-level feature fusion strategy is designed to aggregate multi-scale representations from different depths of the visual encoder, ensuring complementary semantic alignment across layers. In addition, a dynamically modulated weight loss function is introduced to adaptively balance the learning of hard and easy samples, further improving model generalization. Extensive experiments on multiple benchmark medical imaging datasets, demonstrate that the proposed method significantly outperforms existing zero-shot anomaly detection approaches in terms of accuracy, robustness, and generalization.
Yanhui Chen, Hongkang Tao, Zan Yang, Yunkang Cao, Longhua Hu, Pengwen Xiong, Haobo Qiu
Eng. Appl. Artif. Intell.7
2025 Distributed GNE seeking for aggregative games under event-triggered communication: Predefined-time convergent algorithm design
Lingwei Zeng, Jinlei Cheng, Pengwen Xiong, Qian Li 0039
Neurocomputing4
2025 Mobile-DeepRFB: A Lightweight Terrain Classifier for Automatic Mars Rover Navigation
abstract
It requires terrain classification for unmanned Mars Rover to identify the safe areas. The current deep learning-based semantic segmentation and object recognition suffer from a large number of parameters and long training time. In this paper, a lightweight segmentation framework called Mobile-DeepRFB is proposed for the Martian terrain classification. It improves from the DeepLabV3$+$by taking the MobileNetV3 as the backbone module to decrease the parameters and the Receptive Field Block (RFB) module to strengthen the feature extraction capability as well as to enlarge the receptive field. Experimental results on the NASA Mars terrain dataset AI4MARS show that the presented method reduces 94% about the parameter number and improves the mean pixel accuracy by 2% compared to the existing ResNet101 and Xception backbone networks. The deployment of this framework on a low-computing power embedded platform (NVIDIA Jetson Xavier) demonstrates its great potential to apply to Mars rovers.Note to Practitioners—This paper was motivated by the problem of terrain classification of planetary rovers. Existing methods are typically based on semantic segmentation technology to recognize various terrains while suffering from the drawback of a large number of parameters. We propose a lightweight segmentation framework to address this issue. In particular, the lightweight backbone network is applied to significantly reduce the number of parameters. The receptive field module is substantially improved to enhance the feature extraction capability. Eventually, we deploy the framework on a low-computing platform. Experimental tests show that the framework can significantly reduce the number of network parameters and it can be used for planetary rovers with limited computational resources.
Lihang Feng, Sui Wang, Dong Wang 0036, Pengwen Xiong, Jinjin Xie, Miaomiao Zhang 0001, Qi Wu 0003, Aiguo Song
IEEE Trans Autom. Sci. Eng.4
2025 Adversarial Subgraph Contrastive Learning for Predicting Grasp Stability of Robotic Hands With Multimodal Signals
abstract
Accurate prediction of grasp stability is crucial for reliable and precise operations with multi-fingered robotic hands. Traditional methods tend to oversimplify tactile information and pay equal attention to all regions of the data. This can obscure subtle yet critical variations and introduce noise, increasing the risk of stability assessment errors. To address these challenges, a novel self-supervised method, Adversarial Subgraph Contrastive Learning (ASCL), is proposed. It constructs an instance graph from the spatial distribution and features of perceptual nodes. It employs a bi-level adversarial strategy to enhance latent data representations by maximizing the mutual information between the instance graph and its semantic subgraphs, while minimizing it with its noisy subgraphs. To prevent trivial solutions and continuous relaxation of semantic subgraphs, node confidence and edge connection terms are incorporated to ensure stabilization. From an information-theoretic perspective, ASCL exhibits notable advantages on unlabeled or sparsely labeled data, well outperforming existing methods in empirical tests with robotic hands.
Pengwen Xiong, MengChu Zhou, Peter Xiaoping Liu, Aiguo Song
IEEE Trans Autom. Sci. Eng.2
2025 Combining YOLO and background subtraction for small dynamic target detection
Pengwen Xiong, Yushui Huang
Vis. Comput.4
2024 Vision-Locomotion Coordination Control for a Powered Lower-Limb Prosthesis Using Fuzzy-Based Dynamic Movement Primitives
abstract
An amputee cannot directly use the perceptual visual information to control the movements and gait patterns of his worn prosthesis. In order to help an amputee walk and cross over obstacles smoothly, the paper proposes a vision-locomotion coordination control method for a powered lower-limb prosthesis (PLLP), in which a vision system is proposed to detect obstacles, and a complete vision-locomotion loop is then constructed. With deep learning techniques, the vision system can recognize common obstacles (e.g., garbage cans, bricks and boxes) and obtain the features of obstacles (e.g., distance from obstacles to the depth camera and height of obstacles). Through integrating the vision system into the locomotion control system, the PLLP can make obstacle avoidance decisions and use dynamic movement primitives with type-2 fuzzy models (T2FDMPs) to help amputees cross over obstacles simultaneously. Utilizing the type-2 fuzzy models, smooth trajectories for crossing over obstacles can be obtained. The experimental results show that the PLLP with visual information can switch an amputee’s gaits between level walking and obstacle avoidance adaptively, which demonstrates the effectiveness of the vision-locomotion coordination control system.Note to Practitioners—The paper is motivated by the challenge of obstacle avoidance of prostheses. Traditional prostheses cannot achieve autonomous obstacle avoidance because they lack of environmental perception capability. In addition, most prostheses always utilize use the human-robot interaction between amputees and prostheses to recognize the environments. However, due to noise and individual differences, recognition results are not accurate enough. We found that the integration of vision can improve the accuracy and efficiency of environmental recognition. Therefore, it is necessary to construct a complete vision and locomotion closed-loop. In the paper, a vision-locomotion coordination control is proposed, and to make the PLLP cross over obstacles smoothly, a novel trajectory shaping with fuzzy-based dynamic movement primitives is developed. The coordination control is partitioned into the obstacle detection, trajectory shaping and joint control, which can help the PLLP fulfill several obstacle avoidance tasks.
Zhouyang Hong, Shiyuan Bian, Pengwen Xiong, Zhijun Li 0001
IEEE Trans Autom. Sci. Eng.3
2024 An Interpretable Nonlinear Decoupling and Calibration Approach to Wheel Force Transducers
abstract
The multi-dimensional force/torque decoupling and calibration is extremely crucial to increase the accuracy of the Wheel Force Transducer/Sensor (WFT). A novel interpretable nonlinear decoupling and calibration approach to WFT is presented. A physical interpretable prime-error framework is developed such that the linear prime part accounts for most force-voltage responses while the nonlinear error part accounts for the gross error deviation. The conventional least-square decoupling is improved with the delicate nonlinear error modeling using a polynomial base module and a hyperbolic activation function. The developed framework is proved to be mathematically solvable and physically feasible by a two-step calibration scheme. A two-axis WFT is tested and compared with the proposed interpretable nonlinear decoupling model (IND), the least-square-based method (LSM), and the error-based neural network model (eNN). Results demonstrate that the proposed IND provides an accurate, practical, and effective scheme for modeling and calibrating WFTs and maintains a good balance among accuracy, generalization ability, and computational efficiency for real applications.
Lihang Feng, Sui Wang, Pengwen Xiong, Aiguo Song, Peter Xiaoping Liu
IEEE Trans. Intell. Transp. Syst.4
2024 An Improved Level Set Method for Reachability Problems in Differential Games
abstract
This study focuses on reachability problems in differential games. An improved level set (LS) method for computing reachable tubes (RTs) is proposed in this article. The RT is described as a sub-LS of a value function, which is the viscosity solution of a Hamilton–Jacobi (HJ) equation with running cost. We generalize the concept of RTs and propose a new class of RTs, which are referred to as cost-limited one. In particular, a performance index can be specified for the system, and A set of initial states of the system’s evolutions that can reach the target set before the performance index grows to a given allowable cost is referred to as a cost-limited RT (CRT). Such an RT can be obtained by specifying the corresponding running cost function for the HJ equation. Different nonzero sub-LSs of the viscosity solution of the HJ equation at a certain time point can be used to characterize the CRTs with different allowable costs (or the RTs with different time horizons), thus reducing the storage space consumption. The validity and accuracy of the suggested technique are demonstrated via some examples.
Taotao Liang, Pengwen Xiong, Chen Wang 0115, Aiguo Song, Peter Xiaoping Liu
IEEE Trans. Syst. Man Cybern. Syst.3
2023 Robotic haptic adjective perception based on coupled sparse coding
Pengwen Xiong, Kongfei He, Aiguo Song, Peter Xiaoping Liu
Sci. China Inf. Sci.1
2023 Few-Sample Generation of Amount in Figures for Financial Multi-Bill Scene Based on GAN
abstract
Recognition of amount in figures in the financial multi-bill scenes is crucial for the automatic banking business. However, the diversity of banking business and the limitation of customer data privacy determine that it is difficult to collect a large number of sample datasets. Aiming at the problem of insufficient training data in multi-bill scenes and the low accuracy of the detection model, this article proposes a new generative adversarial network (GAN) to generate new samples and to expand the bill dataset, which is then adopted to train a framework for recognition of the bill amount. In the proposed WGAN-SA, a residual block is adopted as the basic structure of the generator and the discriminator, and the self-attention mechanism is also utilized to improve the generation performance. In addition, Wasserstein distance is utilized to measure the distance between real and synthetic samples. Experimental results on the benchmark dataset and comparisons with state-of-the-art works show that our proposed WGAN-SA can effectively improve the few-sample learning performance. Besides, experiments on the bill dataset verify that our method can solve the problem of model collapse and has the ability to generate images of the amount in figures with better fidelity and variety, which is also helpful to achieve better bill amount recognition performance compared with other latest works.
Qi-Qi Chen, Zhao-Hui Sun, Pengwen Xiong, Qi Wu 0003
IEEE Trans. Comput. Soc. Syst.4
2023 Deeply Supervised Subspace Learning for Cross-Modal Material Perception of Known and Unknown Objects
abstract
In order to help robots understand and perceive an object's properties during noncontact robot-object interaction, this article proposes a deeply supervised subspace learning method. In contrast to previous work, it takes the advantages of low noise and fast response of noncontact sensors and extracts novel contactless feature information to retrieve cross-modal information, so as to estimate and infer material properties of known as well as unknown objects. Specifically, a depth-supervised subspace cross-modal material retrieval model is trained to learn a common low-dimensional feature representation to capture the clustering structure among different modal features of the same class of objects. Meanwhile, all of unknown objects are accurately perceived by an energy-based model, which forces an unlabeled novel object's features to be mapped beyond the common low-dimensional features. The experimental results show that our approach is effective in comparison with other advanced methods.
Pengwen Xiong, MengChu Zhou, Aiguo Song, Peter Xiaoping Liu
IEEE Trans. Ind. Informatics1
2023 AGV-Based Vehicle Transportation in Automated Container Terminals: A Survey
abstract
To respond to the rapid growth of shipping container throughput, terminals urgently need to improve the efficiency of thier operations and reduce operational costs through automation and intellectualization upgrades, thereby improving service levels and enhancing market competitiveness. Due to the advantages of reliable transportation, efficient operation, and environmental friendliness, AGV-based automated container terminal (ACT) has become the development trend of container terminals. To help ACT improve its operational management capabilities, plenty of scholars have explored the transportation system of ACT. Through the analysis of operational management issues, the paper defines the four main research topics in vehicle transportation of the ACT including equipment scheduling, path planning, exception handling, and vehicle management. Then, in each topic, the works in the recent 25 years are summarized and several research opportunities for possible follow-up research directions in different fields are proposed. We expect our survey could not only provide references for more scholars on the research of operation and management of terminals, but also provide guidance for system evaluation and improvement for terminal system engineers and operation managers.
Zhao-Hui Sun, Jiapeng You, Siqi Qiu, Qi Wu 0003, Pengwen Xiong, Aiguo Song, Hanzhong Zhang
IEEE Trans. Intell. Transp. Syst.5
2023 Robotic Object Perception Based on Multispectral Few-Shot Coupled Learning
abstract
In order to enable intelligent robots to recognize unknown objects as accurately as human beings, object perception research is of great significance in service and industrial robot application scenarios. However, object perception using spectral measurements under few-shot learning usually leads to a poor result because of inadequate training samples. To overcome this problem, this work proposes a novel few-shot learning with coupled dictionary learning (FSL-CDL) framework. First, a hybrid feature fusion method is developed to extract the multiple dimension-reduced features of original spectral measurements to build the hybrid features. Then, based on the hybrid features, a multitask coupled learning method is developed to effectively recognize unknown objects under few-shot learning. In this method, two coupling patterns, i.e., interspectroscopy coupling and intraspectroscopy coupling, effectively bridge the gap between two spectral measurements. Finally, the proposed FSL-CDL is compared with other advanced algorithms on the SMM50 dataset, and reaches 97.5% and 98.4% recognition accuracy under one-shot and five-shot learning, respectively, which are better than other algorithms. Besides, FSL-CDL can be extended to other perception tasks which contains multiple heterogeneous measurements.
Pengwen Xiong, Xiaobao Tong, Peter Xiaoping Liu, Aiguo Song, Zhijun Li 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2022 Scalable Gamma-Driven Multilayer Network for Brain Workload Detection Through Functional Near-Infrared Spectroscopy
abstract
This work proposes a scalable gamma non-negative matrix network (SGNMN), which uses a Poisson randomized Gamma factor analysis to obtain the neurons of the first layer of a network. These neurons obey Gamma distribution whose shape parameter infers the neurons of the next layer of the network and their related weights. Upsampling the connection weights follows a Dirichlet distribution. Downsampling hidden units obey Gamma distribution. This work performs up-down sampling on each layer to learn the parameters of SGNMN. Experimental results indicate that the width and depth of SGNMN are closely related, and a reasonable network structure for accurately detecting brain fatigue through functional near-infrared spectroscopy can be obtained by considering network width, depth, and parameters.
Qi Wu 0003, Xu-Yi Qiu, Ping-Yu Deng, Pengwen Xiong, Aiguo Song, Limin Zhu 0001, MengChu Zhou
IEEE Trans. Cybern.6
2022 Sliding Mode Impedance Control for Dual Hand Master Single Slave Teleoperation Systems
abstract
For the purpose of avoiding injury and realizing precise operations, the multilateral teleoperation system is the most efficient way to transport trace toxic or radioactive substances, to perform minimally invasive surgery, etc. It is essential to enhance the transparency of a multilateral teleoperation system including multiple masters and the single slave manipulator. However, there are few researchers focus on the allocation of the contact force of the single slave manipulator to different master manipulators. In this paper, we firstly introduce the concept of force translation for teleoperation systems consisting of dual hand master (left and right hands) manipulators and a single slave manipulator. Force translation reflects how the impedance on the single slave side is translated or allocated to contact forces on different master sides. To maintain the stability of the system and to improve transparency, we elucidate the mechanism of the force translation and analyze the relation among masters and the slave. Furthermore, the force translation mechanism is analyzed through numerical simulations and CHAI 3D virtual physical simulations. It is used to propose a multilateral impedance control, and the Lyapunov function is used to analyze the system stability. The results of numerical simulations and real robot experiments verify the effectiveness of the proposed control methods based on the proposed force translation mechanism.
Ting Wang 0013, Zhenxing Sun, Aiguo Song, Pengwen Xiong, Peter Xiaoping Liu
IEEE Trans. Intell. Transp. Syst.4
2022 Inferring Flight Performance Under Different Maneuvers With Pilot's Multi-Physiological Parameters
abstract
The relationship between flight performance and multi-physiological parameters under different flight operating patterns is unknown. This work proposes a Stacked Gaussian Process Network (SGPN) to reveal it. SGPN is a multi-layer network model formed by recursion from a regular Gaussian process and random disturbance. This work constructs an auxiliary variable strategy with the induced points to improve its learning efficiency, thus leading to a sparse SGPN model. In it, a Gaussian process acts as an activation function of each node, but the entire model is no longer a Gaussian process and thus very challenging to solve it. This work presents its solution via variational approximate inference. Experimental results of pilot flight performance evaluation show that the proposed model has stronger learning and generalization ability than its seven competitive peers. It is able to approximate non-linear coupling relationship between multi-physiological parameters and flight height differences.
Qi Wu 0003, MengChu Zhou, Pengwen Xiong, Ruihan Hu, Yu-Wen Jie
IEEE Trans. Intell. Transp. Syst.3
2022 Detecting Dynamic Behavior of Brain Fatigue Through 3-D-CNN-LSTM
abstract
This article proposes a four-dimensional brain mapping method, which can represent the continuous process of a person’s fatigue state in the form of image frames in a space-time range. This work couples 3-D-convolutional neural networks and long-short-term memory networks to form a cognitive detection model of brain fatigue dynamics, which can simulate the continuous process of a person’s brain fatigue dynamics and accurately identify different cognitive fatigue states. Our approach can be applied to any type of brain fatigue detection.
Qi Wu 0003, Pengwen Xiong, Gui-Jiang Li, Aiguo Song, Limin Zhu 0001
IEEE Trans. Syst. Man Cybern. Syst.2