Pengxiang Su

dblp:271/7988 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
12since 2021 · last 2026
0009-0008-0257-0900ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Residual Mamba-Driven Multiscale Attentive Network With Boundary Enhancement for IoT-Enabled Medical Image Segmentation
abstract
In IoT-enabled intelligent healthcare systems, medical images are frequently acquired in real-time from heterogeneous imaging sensors such as dermoscopic devices and MRI scanners. In resource-limited or edge-deployed settings, achieving precise and rapid image segmentation plays a crucial role in facilitating early diagnosis and supporting clinical decision-making. This paper proposes a medical image segmentation method based on a residual Mamba backbone network, combining a multi-scale gated attention (MGA) module with a boundary enhancement (BE) module to effectively enhance the model’s feature representation and boundary localization capabilities. Specifically, the R-Mamba backbone network combines the advantages of statespace modeling and convolutional feature extraction, achieving efficient fusion of global context and local details. The MGA module dynamically captures multi-scale semantic information through dilated convolutions and gating mechanisms, enhancing the model’s adaptability to targets of different scales and shapes. The BE module significantly strengthens boundary representation and fine-grained structural segmentation through multi-scale convolutions and channel-spatial dual attention mechanisms. Additionally, this paper designs a multi-loss function joint optimization strategy to comprehensively constrain region overlap, pixel classification, and structural consistency. Experimental validation on ISIC skin lesion and LGG brain tumor datasets shows competitive performance compared to several mainstream models under the tested conditions.
Guoqiang Ren, Qi Wang 0061, Jieying Tu, Pengxiang Su, Hengrui Liu, Di Gai, Peng Luo 0005, Shuxiao Li
IEEE Internet Things J.4
2026 Decoupled dual-branch prototype network with fine-grained context mining for self-supervised few-shot abdominal image segmentation
Di Gai, Yuxuan Zou, Jieying Tu, Yuhan Geng, Dengyun Xu, Pengxiang Su
Image Vis. Comput.7
2025 Learning Two-factor Representation for Magnetic Resonance Image Super-resolution
abstract
Magnetic Resonance Imaging (MRI) requires a trade-off between resolution, signal-to-noise ratio, and scan time, making high-resolution (HR) acquisition challenging. Therefore, super-resolution for MR image is a feasible solution. However, most existing methods face challenges in accurately learning a continuous volumetric representation from low-resolution image or require HR image for supervision. To solve these challenges, we propose a novel method for MR image super-resolution based on two-factor representation. Specifically, we factorize intensity signals into a linear combination of learnable basis and coefficient factors, enabling efficient continuous volumetric representation from low-resolution MR image. Besides, we introduce a coordinate-based encoding to capture structural relationships between sparse voxels, facilitating smooth completion in unobserved regions. Experiments on BraTS 2019 and MSSEG 2016 datasets demonstrate that our method achieves state-of-the-art performance, providing superior visual fidelity and robustness, particularly in large up-sampling scale MR image super-resolution.
Weifeng Wei, Pengxiang Su
ICASSP3
2025 Adaptive blank compensation for few-shot image classification
Baozhe Wang, Huan Wan, Pengxiang Su, Xin Wei 0002
Neurocomputing4
2023 M-AResNet: a novel multi-scale attention residual network for melting curve image classification
Pengxiang Su, Xuanjing Shen, Haipeng Chen 0002, Di Gai, Yu Liu 0004
Multim. Tools Appl.1
2023 Spatiotemporal Consistency Learning From Momentum Cues for Human Motion Prediction
abstract
Extrapolating future human motion based on the historical human pose sequence is the foundation of various intelligent applications. Numerous deep learning-based algorithms have been designed to address this task, achieving state-of-the-art performance on different human motion benchmark datasets. However, most existing methods employ three-dimensional coordinates of joints to demonstrate dynamic motion contexts implicitly. Unfortunately, it remains challenging in capturing motion information from the pose sequence. In this paper, we advocate explicitly describing dynamic contexts via the momentum of human motion mechanic space, as the momentum of a joint is explicit, temporal consistent, and can provide abundant information to the model. In addition, the single-stream methods play a dominant role in the field of human motion prediction. They usually capture motion information via the strategy of continuous or sparse sampling, which might obviate global or detailed local information. Therefore, we present a simple yet effective dual-stream method that can consider both the detailed and global temporal information through a combination of continuous and sparse sampling. The proposed dual-stream paradigm enables the improvement of computational efficiency and the short-term prediction accuracy concurrently. Furthermore, we present a novel temporal attention-based graph convolutional network (TA-GCN) to derive a spatiotemporally consistent motion representation, which can adequately consider the rationality of human body topology. Extensive experiments on two large motion prediction benchmark datasets (i.e., Human 3.6M and CMU Mocap) show that our algorithm achieves state-of-the-art performance both qualitatively and quantitatively.
Haipeng Chen 0002, Wenyin Zhang, Pengxiang Su
IEEE Trans. Circuits Syst. Video Technol.4
2023 Spatiotemporal Learning Transformer for Video-Based Human Pose Estimation
abstract
Multi-frame human pose estimation has long been an appealing and fundamental issue in visual perception. Owing to the frequent rapid motion and pose occlusion in videos, this task is extremely challenging. Current state-of-the-art methods seek to model spatiotemporal features by equally fusing each frame in the local sequence, which weakens the target frame information. In addition, existing approaches usually emphasize more on deep features while ignoring the detailed information implied in the shallow feature maps, resulting in the dropping of crucial features. To address the above problems, we propose an effective framework, namely spatiotemporal learning transformer for video-based human pose estimation (SLT-Pose), which consists of a Personalized Feature Extraction Module (PFEM), Self-feature Refinement Module (SRM), Cross-frame Temporal Learning Module (CTLM) and Disentangled Keypoint Detector (DKD). To be specific, we propose PFEM which extracts and modulates the individual frame features to adapt to the varying human shape, and integrates single-frame features to obtain the spatiotemporal features. We further present SRM to establish global correlation spatial cues on the target frame to attain the refinement feature. Then, a CTLM is designed to search for the information most closely related to the target frame from the spatiotemporal features to intensify the interaction between the target frame and the local sequence, using both the shallow detailed and the deep semantic representations. Finally, we employ DKD to extract the disentangled characteristics of each joint and encode the articulated joint pairs in the human body, promoting the model to reasonably and accurately predict the keypoint heatmaps. Extensive experiments on three huamn motion benchmarks, including PoseTrack2017, PoseTrack2018, and Sub-JHMDB dataset, demonstrate that SLT-Pose plays favorably against state-of-the-art approaches in terms of both objective evaluation and subjective visual performance.
Di Gai, Runyang Feng, Weidong Min, Xiaosong Yang, Pengxiang Su, Qi Wang 0061
IEEE Trans. Circuits Syst. Video Technol.5
2022 Multiscale Spatial and Temporal Learning for Human Motion Prediction
Pengxiang Su, Xuanjing Shen, Haipeng Chen 0002
ICANN (2)1
2022 Spatial-Temporal Correlation Modeling for Motion Prediction
abstract
Human motion prediction is fundamental for many applications in computer vision. Current methods typically handle motion prediction with seqential models, which ignore the fact that joint movement is driven by forces. In this paper, we provide a novel mechanical view to decompose force into magnitude and direction, which contributes to modeling the temporal evolution of joints. Moreover, existing graph convolution-based methods merely utilize the deep-level features, which is difficult to capture the complex spatial dependencies contexts. We introduce a novel spatial connections encoding model to capture the multi-level spatial dependencies between joints. Finally, to encode abundant temporal dependencies, we present a multi-head temporal encoding module. Comprehensive experiments show that our model sets the state-of-the-art performance on the largest human motion benchmark datasets.
Yingying Jiao, Haipeng Chen 0002, Chang Yao 0001, Pengxiang Su, Chong Fu 0002, Xiang Wang 0010
ICME4
2022 Adaptive Multi-Order Graph Neural Networks for Human Motion Prediction
abstract
Human motion prediction aims at capturing the hidden temporal correlations between historical motion and future poses. Various graph convolution networks have been presented for encoding the spatial dependencies between joints. Empirically, the crucial shortcoming of these methods is that they fail to extract enough spatially relevant information. In this paper, we propose an adaptive multi-order context fusion architecture that consists of two components. A novel message propagation module encodes the interaction between joints, while highlighting contexts from closely related joints. An adaptive aggregation module fuses various information from different-order joint features. Our model is evaluated on Human 3.6 Million dataset. Extensive experiments show that our method achieves state-of-the-art performance on short-term and long-term predictions.
Pengxiang Su, Xuanjing Shen, Zenan Shi
ICME1
2021 Motion Prediction using Trajectory Cues
abstract
Predicting human motion from a historical pose sequence is at the core of many applications in computer vision. Current state-of-the-art methods concentrate on learning motion contexts in the pose space, however, the high dimensionality and complex nature of human pose invoke inherent difficulties in extracting such contexts. In this paper, we instead advocate to model motion contexts in the joint trajectory space, as the trajectory of a joint is smooth, vectorial, and gives sufficient information to the model. Moreover, most existing methods consider only the dependencies between skeletal connected joints, disregarding prior knowledge and the hidden connections between geometrically separated joints. Motivated by this, we present a semi-constrained graph to explicitly encode skeletal connections and prior knowledge, while adaptively learn implicit dependencies between joints.We also explore the applications of our approach to a range of objects including human, fish, and mouse. Surprisingly, our method sets the new state-of-the-art performance on 4 different benchmark datasets, a remarkable highlight is that it achieves a 19.1% accuracy improvement over current state-of-the-art in average. To facilitate future research, we have released our code at https://github.com/Pose-Group/MPT.
Zhenguang Liu, Pengxiang Su, Shuang Wu 0002, Xuanjing Shen, Haipeng Chen 0002, Yanbin Hao, Meng Wang 0001
ICCV2
2021 Motion Prediction via Joint Dependency Modeling in Phase Space
abstract
Motion prediction is a classic problem in computer vision, which aims at forecasting future motion given the observed pose sequence. Various deep learning models have been proposed, achieving state-of-the-art performance on motion prediction. However, existing methods typically focus on modeling temporal dynamics in the pose space. Unfortunately, the complicated and high dimensionality nature of human motion brings inherent challenges for dynamic context capturing. Therefore, we move away from the conventional pose based representation and present a novel approach employing a phase space trajectory representation of individual joints. Moreover, current methods tend to only consider the dependencies between physically connected joints. In this paper, we introduce a novel convolutional neural model to effectively leverage explicit prior knowledge of motion anatomy, and simultaneously capture both spatial and temporal information of joint trajectory dynamics. We then propose a global optimization module that learns the implicit relationships between individual joint features. Empirically, our method is evaluated on large-scale 3D human motion benchmark datasets (i.e., Human3.6M, CMU MoCap). These results demonstrate that our method sets the new state-of-the-art on the benchmark datasets. Our code is released at https://github.com/Pose-Group/TEID.
Pengxiang Su, Zhenguang Liu, Shuang Wu 0002, Lei Zhu 0002, Yifang Yin, Xuanjing Shen
ACM Multimedia1
2020 Medical image fusion using the PCNN based on IQPSO in NSST domain
abstract
In this study, an improved quantum‐behaved particle swarm optimisation based pulse‐coupled neural network (IQPSO‐PCNN) is proposed in the non‐subsampled shearlet transform (NSST) domain for medical image fusion. First, NSST tool is used to decompose the source image into low‐frequency and high‐frequency subbands. Then, for low‐frequency subbands, the fusion rules of two different functions are presented, which simultaneously addresses two key issues of energy preservation and detail extraction. For high‐frequency subbands, unlike conventional PCNN‐based methods, parameters are manually set based on experience, and the decomposed high‐frequency subbands share a set of parameters. The IQPSO‐PCNN model can obtain the optimal parameters for each high‐frequency subband adaptively according to its own information. Finally, the fused low‐frequency subband and high‐frequency subbands are inversely transformed by NSST to acquire the final fused image. The proposed algorithm uses >90 pairs of images with four different modalities. In addition, fusion experiments are performed on different sequences of the three modes. The experimental results demonstrate that the proposed method is superior to existing state‐of‐art methods in subjective visual performance and objective evaluation.
Di Gai, Xuanjing Shen, Haipeng Chen 0002, Zeyu Xie, Pengxiang Su
IET Image Process.5
2020 Multi-focus image fusion method based on two stage of convolutional neural network
Di Gai, Xuanjing Shen, Haipeng Chen 0002, Pengxiang Su
Signal Process.4