Zhongcai Pei

dblp:69/8488 · DBLP profile ↗
← Back
18ranked-venue papers
1as first author
17since 2021 · last 2025
0000-0001-7748-8591ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Quadruped robot locomotion via soft actor-critic with muti-head critic and dynamic policy gradient
Yanan Fan, Zhongcai Pei, Hongbing Shi, Tianyuan Guo, Zhiyong Tang
Appl. Intell.2
2025 Domain generalization for zero-calibration brain-computer interfaces with knowledge distillation-based phase invariant feature extraction
Zilin Liang, Zheng Zheng 0001, Weihai Chen, Xinzhi Ma, Zhongcai Pei, Xiantao Sun
Eng. Appl. Artif. Intell.5
2025 Learnable patchmatch and self-teaching for multi-frame depth estimation in monocular endoscopy
Shuwei Shao, Zhongcai Pei, Weihai Chen, Xingming Wu, Zhong Liu 0005
Eng. Appl. Artif. Intell.2
2025 IEBins: Iterative Elastic Bins for Monocular Depth Estimation and Completion
Shuwei Shao, Zhongcai Pei, Weihai Chen, Peter C. Y. Chen, Zhengguo Li
Int. J. Comput. Vis.2
2025 Cross-Modality Self-Attention and Fusion-Based Neural Network for Lower Limb Locomotion Mode Recognition
abstract
Although there are many wearable sensors that make the acquisition of multi-modality data easier, effective feature extraction and fusion of the data is still challenging for lower limb locomotion mode recognition. In this article, a novel neural network is proposed for accurate prediction of five common lower limb locomotion modes including level walking, ramp ascent, ramp descent, stair ascent, and stair descent. First, the encoder-decoder structure is employed to enrich the channel diversity for the separation of the useful patterns from combined patterns. Second, a self-attention based cross-modality interaction module is proposed, which enables bilateral information flow between two encoding paths to fully exploit the interdependencies and to find complementary information between modalities. Third, a multi-modality fusion module is designed where the complementary features are fused by a channel-wise weighted summation whose coefficients are learned end-to-end. A benchmark dataset is collected from 10 health subjects containing EMG and IMU signals and five locomotion modes. Extensive experiments are conducted on one publicly available dataset ENABL3S and one self-collected dataset. The results show that the proposed method outperforms the compared methods with higher classification accuracy. The proposed method achieves a classification accuracy of 98.25% on ENABL3S dataset and 95.51% on the self-collected dataset. Note to Practitioners—This article aims to solve the real challenges encountered when intelligent recognition algorithms are applied in wearable robots: how to effectively and efficiently fuse the multi-modality data for better decision-making. First, most existing methods directly concatenate the multi-modality data, which increases the data dimensionality and brings computational burden. Second, existing recognition neural networks continuously compress the feature size such that the discriminative patterns are submerged in the noise and thus difficult to be identified. This research decomposes the mixed input signals on the channel dimension such that the useful patterns can be separated. Moreover, this research employs self-attention mechanism to associate correlations between two modalities and use this correlation as a new feature for subsequent representation learning, generating new, compact, and complementary features for classification. We demonstrate that the proposed network achieves 98.25% accuracy and 3.5 ms prediction time. We anticipate that the proposed network could be a general scientific and practical methodology of multi-modality signal fusion and feature learning for intelligent systems.
Changchen Zhao, Wenbo Song, Zhongcai Pei, Weihai Chen
IEEE Trans Autom. Sci. Eng.5
2025 MonoDiffusion: Self-Supervised Monocular Depth Estimation Using Diffusion Model
abstract
Over the past few years, self-supervised monocular depth estimation has received widespread attention. Most efforts focus on designing different types of network architectures and loss functions or handling edge cases, for example, occlusion and dynamic objects. In this work, we take another path and propose a novel conditional diffusion-based generative framework for self-supervised monocular depth estimation, dubbed MonoDiffusion. Because the depth ground-truth is unavailable in a self-supervised setting, we develop a new pseudo ground-truth diffusion process to assist the diffusion for training. Instead of diffusing at a fixed high resolution, we perform diffusion in a coarse-to-fine manner that allows for faster inference time without sacrificing accuracy or even better accuracy. Furthermore, we develop a simple yet effective contrastive depth reconstruction mechanism to enhance the denoising ability of model. It is worth noting that the proposed MonoDiffusion has the property of naturally acquiring the depth uncertainty that is essential to be implemented in safety-critical cases. Extensive experiments on the KITTI, Make3D and DIML datasets indicate that our MonoDiffusion outperforms prior state-of-the-art self-supervised competitors. The source code will be publicly available upon the acceptance.
Shuwei Shao, Zhongcai Pei, Weihai Chen, Dingchi Sun, Peter C. Y. Chen, Zhengguo Li
IEEE Trans. Circuits Syst. Video Technol.2
2024 A wearable knee rehabilitation system based on graphene textile composite sensor: Implementation and validation
Zhongcai Pei, Weihai Chen, Xingming Wu, Jianer Chen
Eng. Appl. Artif. Intell.2
2024 NDDepth: Normal-Distance Assisted Monocular Depth Estimation and Completion
abstract
Over the past few years, monocular depth estimation and completion have been paid more and more attention from the computer vision community because of their widespread applications. In this paper, we introduce novel physics (geometry)-driven deep learning frameworks for these two tasks by assuming that 3D scenes are constituted with piece-wise planes. Instead of directly estimating the depth map or completing the sparse depth map, we propose to estimate the surface normal and plane-to-origin distance maps or complete the sparse surface normal and distance maps as intermediate outputs. To this end, we develop a normal-distance head that outputs pixel-level surface normal and distance. Afterthat, the surface normal and distance maps are regularized by a developed plane-aware consistency constraint, which are then transformed into depth maps. Furthermore, we integrate an additional depth head to strengthen the robustness of the proposed frameworks. Extensive experiments on the NYU-Depth-v2, KITTI and SUN RGB-D datasets demonstrate that our method exceeds in performance prior state-of-the-art monocular depth estimation and completion competitors.
Shuwei Shao, Zhongcai Pei, Weihai Chen, Peter C. Y. Chen, Zhengguo Li
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 URCDC-Depth: Uncertainty Rectified Cross-Distillation With CutFlip for Monocular Depth Estimation
abstract
This work aims to estimate a high-quality depth map from a single RGB image. Due to the lack of depth clues, making full use of the long-range correlation and local information is critical for accurate depth estimation. To this end, we introduce an uncertainty rectified cross-distillation between the Transformer and convolutional neural network (CNN) to achieve a comprehensive depth estimator. Specifically, we utilize the depth estimates from the Transformer branch and CNN branch as pseudo labels to teach each other. At the same time, the pixel-wise depth uncertainty is modeled to mitigate the negative impact of noisy pseudo labels. To avoid the large capacity gap induced by the strong Transformer branch deteriorating the cross-distillation, we transfer the feature maps from the Transformer to the CNN and develop coupling units to assist the weak CNN branch in leveraging the transferred features. Furthermore, we introduce CutFlip, a surprisingly simple yet highly effective data augmentation technique, which forces the model to focus on more valuable depth reasoning clues apart from the vertical image position. Extensive experiments demonstrate that our model, termedURCDC-Depth, exceeds in performance previous state-of-the-art approaches on the KITTI, NYU-Depth-v2 and SUN RGB-D datasets, with no additional computational burden in the evaluation phase. The source code will be publicly available upon acceptance. The source code is available athttps://github.com/ShuweiShao/URCDC-Depth.
Shuwei Shao, Zhongcai Pei, Weihai Chen, Zhong Liu 0005, Zhengguo Li
IEEE Trans. Multim.2
2024 StairNetV3: depth-aware stair modeling using deep learning
Chen Wang 0056, Zhongcai Pei, Yachun Wang, Zhiyong Tang
Vis. Comput.2
2023 NDDepth: Normal-Distance Assisted Monocular Depth Estimation
abstract
Monocular depth estimation has drawn widespread attention from the vision community due to its broad applications. In this paper, we propose a novel physics (geometry)-driven deep learning framework for monocular depth estimation by assuming that 3D scenes are constituted by piece-wise planes. Particularly, we introduce a new normal-distance head that outputs pixel-level surface normal and plane-to-origin distance for deriving depth at each position. Meanwhile, the normal and distance are regularized by a developed plane-aware consistency constraint. We further integrate an additional depth head to improve the robustness of the proposed framework. To fully exploit the strengths of these two heads, we develop an effective contrastive iterative refinement module that refines depth in a complementary manner according to the depth uncertainty. Extensive experiments indicate that the proposed method exceeds previous state-of-the-art competitors on the NYU-Depth-v2, KITTI and SUN RGB-D datasets. Notably, it ranks 1st among all submissions on the KITTI depth prediction online benchmark at the submission time. The source code is available at https://github.com/ShuweiShao/NDDepth.
Shuwei Shao, Zhongcai Pei, Weihai Chen, Xingming Wu, Zhengguo Li
ICCV2
2023 Simultaneous Gait Event Intention Detection Using Single sEMG Sensor for Lower Limb Exoskeleton
abstract
Accurate detection of gait event intention and sending it to lower limb exoskeleton (LLE) is the key to achieve active rehabilitation. Most existing surface electromyography (sEMG)-based gait event intention detection methods suffer from insufficient generalization and complex detection. In this paper, we propose a novel approach for detecting gait event intention using a single sEMG sensor. The gait event intention is obtained by detecting the peak activity of the rectus femoris during the stance period. First, the root mean square (RMS) features are extracted from the sEMG data of the rectus femoris. Then, the data groups composed of the RMS features are smoothed and all extreme points are calculated. Finally, the midstance (MSt) events are discovered when the latest maximum point satisfies the preset condition. The experimental results of three different gait speeds showed that the proposed approach could adapt to different walking speeds and maintain a high detection accuracy of gait event intention detection. This study provides a convenient and reliable detection approach for gait research of LLE.
Zhongcai Pei, Weihai Chen, Wen Duan, Jianer Chen
IECON2
2023 IEBins: Iterative Elastic Bins for Monocular Depth Estimation
abstract
Monocular depth estimation (MDE) is a fundamental topic of geometric computer vision and a core technique for many downstream applications. Recently, several methods reframe the MDE as a classification-regression problem where a linear combination of probabilistic distribution and bin centers is used to predict depth. In this paper, we propose a novel concept of iterative elastic bins (IEBins) for the classification-regression-based MDE. The proposed IEBins aims to search for high-quality depth by progressively optimizing the search range, which involves multiple stages and each stage performs a finer-grained depth search in the target bin on top of its previous stage. To alleviate the possible error accumulation during the iterative process, we utilize a novel elastic target bin to replace the original target bin, the width of which is adjusted elastically based on the depth uncertainty. Furthermore, we develop a dedicated framework composed of a feature extractor and an iterative optimizer that has powerful temporal context modeling capabilities benefiting from the GRU-based architecture. Extensive experiments on the KITTI, NYU-Depth-v2 and SUN RGB-D datasets demonstrate that the proposed method surpasses prior state-of-the-art competitors. The source code is publicly available at https://github.com/ShuweiShao/IEBins.
Shuwei Shao, Zhongcai Pei, Xingming Wu, Zhong Liu 0005, Weihai Chen, Zhengguo Li
NeurIPS2
2023 Towards Comprehensive Monocular Depth Estimation: Multiple Heads are Better Than One
abstract
Depth estimation attracts widespread attention in the computer vision community. However, it is still quite difficult to recover an accurate depth map using only one RGB image. We observe a phenomenon that existing methods tend to fail in different cases, caused by differences in network architecture, loss function and so on. In this work, we investigate into the phenomenon and propose to integrate the strengths of multiple weak depth predictor to build a comprehensive and accurate depth predictor, which is critical for many real-world applications, e.g., 3D reconstruction. Specifically, we construct multiple base (weak) depth predictors by utilizing different Transformer-based and convolutional neural network (CNN)-based architectures. Transformer establishes long-range correlation while CNN preserves local information ignored by Transformer due to the spatial inductive bias. Therefore, the coupling of Transformer and CNN contributes to the generation of complementary depth estimates, which are essential to achieve a comprehensive depth predictor. Then, we design mixers to learn from multiple weak predictions and adaptively fuse them into a strong depth estimate. The resultant model, which we refer to as Transformer-assisted depth ensembles (TEDepth). On the standard NYU-Depth-v2 and KITTI datasets, we thoroughly explore how the neural ensembles affect the depth estimation and demonstrate that our TEDepth achieves better results than previous state-of-the-art approaches. To validate the generalizability across cameras, we directly apply the models trained on NYU-Depth-v2 to the SUN RGB-D dataset without any fine-tuning, and the superior results emphasize its strong generalizability.
Shuwei Shao, Zhongcai Pei, Zhong Liu 0005, Weihai Chen, Wentao Zhu 0001, Xingming Wu, Baochang Zhang 0001
IEEE Trans. Multim.3
2022 Self-Supervised monocular depth and ego-Motion estimation in endoscopy: Appearance flow to the rescue
Shuwei Shao, Zhongcai Pei, Weihai Chen, Wentao Zhu 0001, Xingming Wu, Dianmin Sun, Baochang Zhang 0001
Medical Image Anal.2
2021 Self-Supervised Learning for Monocular Depth Estimation on Minimally Invasive Surgery Scenes
abstract
Self-supervised learning algorithms that compute depth map from monocular videos have achieved remarkable performance on urban scenes and have been applied extensively. These techniques still face significant challenges, however, when applied directly to endoscopic videos because of the brightness variations from frame to frame and inadequate representation learning during the training phase. Inspired by the optical flow for motion alignment between adjacent frames, we design a AFNet with structural stability loss and residual-based smoothness loss to learn the appearance flow across adjacent frames, which handles the brightness inconsistency issue efficaciously. In addition, we propose a novel self-attention mechanism named feature scaling module to alleviate the inadequate representation learning problem. In a comparison study to the current state-of-the-art self-supervised methods explored for urban videos on the SCARED dataset, the developed model surpasses existing methods by a large margin.
Shuwei Shao, Zhongcai Pei, Weihai Chen, Baochang Zhang 0001, Xingming Wu, Dianmin Sun, David S. Doermann
ICRA2
2021 Coordinate-based anchor-free module for object detection
Zhiyong Tang, Jianbing Yang, Zhongcai Pei, Xiao Song 0001
Appl. Intell.3
2012 Adaptive control of a quadruped robot based on Central Pattern Generators
abstract
Biocybernetics method is developed to solve robot motion control problem in recent years. Biocybernetics method can realize the robot's rhythmic movement by simulating, simplifying and improving rhythmic movement control area of animal. The rhythmic movement has many advantages, such as regular expression form, high stability and adaptability, so it is always used to control the motion of legged robots. Animal's rhythmic movement is controlled by Central Pattern Generators (CPG). This paper designed a general hydraulic quadruped robot platform adopting Matsuoka's CPG model to complete the robot motion intelligent control. Four gaits of quadruped robot were simulated based on this CPG control model and the foot endpoints trajectory planning problems of quadruped robot was analyzed.
Zhongcai Pei
INDIN1