VLDB 2026 Research / reviewers in the wild / expert
Shenghao Zhang 0001
dblp:191/1814-1
· DBLP profile ↗
14ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0002-5804-5012ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TFGait - Stable and Efficient Adaptive Gait Planning With Terrain Recognition and Froude Number for Quadruped RobotabstractGait planning is one of the most critical technologies for quadruped robots. However, far too little attention has been paid to the tight coupling mechanism of gait planning with terrain understanding and energy efficiency. To date, it is still challenging to plan optimal gait strategies that are highly adapted to terrain features with stable and efficient transitions. Accordingly, this paper proposes an adaptive gait control framework for quadruped robots that combines terrain recognition, Cost of Transport (CoT), and the Froude (Fr) number. More specifically, an optimal gait selection strategy for quadruped robots is designed based on different terrain texture features and the CoT characteristics of different gaits. To address the gait transition process induced thereby, an adaptive method for gait parameters based on the Fr number is further proposed, which can make the process more stable. Besides, model predictive control (MPC) and whole-body control (WBC) are employed as the motion controllers for the quadruped robot. Furthermore, simulation and experimental results indicate that the proposed method possesses superior terrain adaptability, energy efficiency, and motion stability during gait transitions, which is beneficial for the quadruped robots to maintain stable motion and reduce energy consumption when performing tasks in changeable terrains.Note to Practitioners—This paper is motivated by the problem of adaptive gait planning for quadruped robots that walks through different terrains. We propose a method that ensures optimal gait selection by robots facing diverse terrains and maintains the stability of gait transition. The proposed control framework, upon testing in a simulated environment, can be directly deployed on real-world robot without further adjustments and allows the robot to traverse various terrains with minimal sim-to-real issues. Hopefully, our proposed method can provide valuable guidance and support for facilitating the enhancement of capabilities in performing prolonged endurance tasks in unstructured environments for quadruped robots. Aocheng Luo, Qifeng Wan, Shihan Kong, Wanchao Chi, Shenghao Zhang 0001, Qiuguo Zhu, Junzhi Yu 0001 |
IEEE Trans Autom. Sci. Eng. | 7 |
| 2025 | A Novel ViDAR Device With Visual Inertial Encoder Odometry and Reinforcement Learning-Based Active SLAM MethodabstractIn the field of multisensor fusion for simultaneous localization and mapping (SLAM), monocular cameras and IMUs are widely used to build simple and effective visual-inertial systems. However, limited research has explored the integration of motor-encoder devices to enhance SLAM performance. By incorporating such devices, it is possible to significantly improve active capability and field of view (FOV) with minimal additional cost and structural complexity. This article proposes a novel visual-inertial-encoder tightly coupled odometry (VIEO) based on a video detection and ranging (ViDAR) device. A ViDAR calibration method is introduced to ensure accurate initialization for VIEO. In addition, a platform motion decoupled active SLAM method based on deep reinforcement learning (DRL) is proposed. Experimental data demonstrate that the proposed ViDAR and the VIEO algorithm significantly increase cross-frame co-visibility relationships compared to its corresponding visual-inertial odometry (VIO) algorithm, improving state estimation accuracy. Additionally, the DRL-based active SLAM algorithm, with the ability to decouple from platform motion, can increase the diversity weight of the feature points and further enhance the VIEO algorithm's performance. The proposed methodology sheds fresh insights into both the updated platform design and decoupled approach of active SLAM systems in complex environments. Zhanhua Xin, Shenghao Zhang 0001, Wanchao Chi, Shihan Kong, Junzhi Yu 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2024 | Deep Fusion for Multi-Modal 6D Pose Estimationabstract6D pose estimation with individual modality encounters difficulties due to the limitations of modalities, such as RGB information on textureless objects and depth on reflective objects. This can be improved by exploiting the complementarity between modalities. Most of the previous methods only consider the correspondence between point clouds and RGB images and directly extract the features of the corresponding two modalities for fusion, which ignore the information of the modality itself and are negatively affected by erroneous background information when introducing more features for fusion. To enhance the complementarities between multiple modalities, we propose a neighbor-based cross-modalities attention mechanism for multi-modal 6D pose estimation. Neighbors represent that the RGB features of multiple neighbor are applied for fusion, which expands the receptive field. The cross-modalities attention mechanism leverages the similarities between the different modal features to help modal feature fusion, which reduces the negative impact of incorrect background information. Moreover, we design some features between the rendered image and the original image to obtain the confidence of pose estimation results. Experimental results on LM, LM-O and YCB-V datasets demonstrate the effectiveness of our methods. Video is available at https://www.youtube.com/watch?v=ApNBcX6NEGs.Note to Practitioners—Introducing the information of surrounding points during multi-modal fusion improves the performance of 6D pose estimation. For example, the RGB image corresponding to some point clouds on the object may lack rich texture features while the neighbors exist. However, most methods of modal fusion based on RGBD for 6D pose estimation only simply consider the corresponding between RGB images and point clouds for feature fusion, which may bring redundant information or the wrong background information when introducing neighbor information. In this paper, we propose a cross-modal attention mechanism based on neighbor information. By introducing the information of the modality itself to obtain the weight of the neighbor information of another modality in the encoding and decoding stages, the receptive field is expanded and the complementarities between different modalities are enhanced. The experiment shows our effectiveness. In addition, we provide a pose confidence estimator for predicted pose results. Specifically, the rendered image with the predicted pose and the real image are applied to extract features for the decision tree. The experimental results show that the result of the wrong estimation can be eliminated with high accuracy and recall. The 6D pose confidence can provide a reference for real-world grasping. However, the current method can only estimate objects with known models. In the future, we will consider applying the method to unseen objects. Shifeng Lin, Zunran Wang, Shenghao Zhang 0001, Yonggen Ling, Chenguang Yang 0001 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2024 | Max: A Wheeled-Legged Quadruped Robot for Multimodal Agile LocomotionabstractTo enrich legged robots with fast energy-efficient mobility on even terrain, wheeled-legged robots have emerged as a valued robot form in robotics research. This paper describes the complete development of a new wheeled-legged quadruped robot named Max, ranging from its mechanical design over system architecture to core algorithms implemented for it to realize various motion behaviors. Instead of attaching wheels to the distal ends of legs as in the existing wheeled-legged robot designs, this robot has wheels installed on the knees with a special switching mechanism to convert a leg between the legged and wheeled locomotion modes. This design keeps the wheeled leg lightweight, enabling the robot to preserve the motion agility as a quadruped robot while gaining the energy-efficiency as a four-wheel or even two-wheel mobile robot. An online locomotion generation method is proposed to compute the 6-D body trajectory of the robot in walking on the perceived terrain, while dynamic movements such as leaps and flips are generated by a unified trajectory optimizer, which is also used to generate the transition motions of the robot to transform into the wheeled mode. The diverse mobility of the proposed robot Max is verified with extensive experiments.Note to Practitioners—Empowering robots with all-terrain mobility is a fundamental open problem in developing a new generation of robots. To this end, combinations of wheels and legs have been explored for robots to possess both traversability on uneven terrains and efficiency on even terrains. This paper proposes a new wheeled-legged quadruped robot with focuses on the integrated design of wheeled legs, system architecture, and core algorithms implemented for various legged and wheeled locomotion behaviors. To embed wheels without adding additional motors and keep the light weight of original legs, a special switching mechanism is designed and integrated at the knee joints where wheels are installed. Algorithms for generating quadrupedal walk according to online perceived terrain information as well as other dynamic legged and wheeled motions are discussed and demonstrated. The system architecture for allocating all vision and motion algorithms is also presented. This work is intended to provide a whole picture of developing this new robot including both hardware and software aspects. Qinqin Zhou 0002, Xinyang Jiang, Wanchao Chi, Shenghao Zhang 0001, Jingfan Zhang, Rui Wang 0193, Jingchen Li 0001, Shuai Wang 0007, Lingzhu Xiang, Yu Zheng 0001, Zhengyou Zhang |
IEEE Trans Autom. Sci. Eng. | 7 |
| 2023 | Mx2M: Masked Cross-Modality Modeling in Domain Adaptation for 3D Semantic SegmentationabstractExisting methods of cross-modal domain adaptation for 3D semantic segmentation predict results only via 2D-3D complementarity that is obtained by cross-modal feature matching. However, as lacking supervision in the target domain, the complementarity is not always reliable. The results are not ideal when the domain gap is large. To solve the problem of lacking supervision, we introduce masked modeling into this task and propose a method Mx2M, which utilizes masked cross-modality modeling to reduce the large domain gap. Our Mx2M contains two components. One is the core solution, cross-modal removal and prediction (xMRP), which makes the Mx2M adapt to various scenarios and provides cross-modal self-supervision. The other is a new way of cross-modal feature matching, the dynamic cross-modal filter (DxMF) that ensures the whole method dynamically uses more suitable 2D-3D complementarity. Evaluation of the Mx2M on three DA scenarios, including Day/Night, USA/Singapore, and A2D2/SemanticKITTI, brings large improvements over previous methods on many metrics. Boxiang Zhang, Zunran Wang, Yonggen Ling, Yuanyuan Guan, Shenghao Zhang 0001, Wenhui Li 0002 |
AAAI | 5 |
| 2023 | ShuffleTrans: Patch-wise weight shuffle for transparent object segmentation
Boxiang Zhang, Zunran Wang, Yonggen Ling, Yuanyuan Guan, Shenghao Zhang 0001, Wenhui Li 0002, Lei Wei 0002, Chunxu Zhang |
Neural Networks | 5 |
| 2022 | RECCraft System: Towards Reliable and Efficient Collective Robotic ConstructionabstractThis research presents a novel Collective Robotic Construction (CRC) system named RECCraft. The RECCraft hardware system is composed of the mobile manipulation vehicles, the cubic blocks, and the folding ramp blocks. Solid connection and easy removal of the blocks are achieved by an electropermanent magnet and silicon steel sheets. With one degree of freedom (DOF) lifting manipulator, the robot can carry a block 3.7 times its volume. An active folding ramp block can provide a robust passage to the upper level for the robot. Our study focuses on systemic improvement of the construction speed and reliability of the robotic construction system. Visual perception system realized by Apritag is adopted, featured by convenient deployment and high precision, to provide a reliable guarantee for robotic construction. RL-based planner provides end-to-end solution for planning tasks of building multi-layer constructions, which is validated by simulation platform and real prototype. Compared with construction speed of existing robotic construction systems, our proposed RECCraft system achieves state-of-the-art level. The robot builds a 2-layer construction by RL-based planner in 4 minutes and 16 seconds, which achieves construction volumetric throughput of 6.7×105mm3/s. Qiwei Xu, Yizheng Zhang, Shenghao Zhang 0001, Zhuoxing Wu, Xiong Li 0001, Jiahong Chen, Zengjun Zhao, Luyang Tang, Zhengyou Zhang, Lei Han 0001 |
IROS | 3 |
| 2022 | Unsupervised Occlusion-Aware Stereo Matching With Directed Disparity SmoothingabstractWhen handling occlusion in unsupervised stereo matching, existing methods tend to neglect the supportive role of occlusion and to perform inappropriate disparity smoothing around the occlusion. To address these problems, we propose an occlusion-aware stereo network that contains a specific module to first estimate occlusion as an additional depth cue. In the occlusion inference module, a pixel is classified with a three-category label based on whether an area is occluded by an object on the left, occluded by an object on the right, or unoccluded. After the occluders are detected, we introduce a directed disparity smoothing loss that allows valid disparity estimates to be propagated to fill the occluded region, while ambiguous matches in the occluded region do not affect other regions. Disparity and occlusion are trained alternately in an unsupervised manner with detached backpropagation to enable the directed smoothness. Experiments show that our method achieves 3-pixel threshold error rates of 6.51% and 5.69% on the KITTI 2015 and KITTI 2012 validation sets, state-of-the-art results among unsupervised learning networks at the time of submission. Ang Li 0025, Zejian Yuan, Yonggen Ling, Wanchao Chi, Shenghao Zhang 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2020 | Domain Adaptation Gaze Estimation by Embedding with Prediction Consistency
Zidong Guo, Zejian Yuan, Wanchao Chi, Yonggen Ling, Shenghao Zhang 0001 |
ACCV (5) | 6 |
| 2020 | Learning End-to-End Action Interaction by Paired-Embedding Data Augmentation
Zejian Yuan, Wanchao Chi, Yonggen Ling, Shenghao Zhang 0001 |
ACCV (6) | 6 |
| 2020 | FastCompletion: A Cascade Network with Multiscale Group-Fused Inputs for Real-Time Depth CompletionabstractCompleting sparse data captured with commercial depth sensors is a vital and fundamental procedure for many computer vision applications. For execution in real-world scenarios, a good trade-off between accuracy and speed is increasingly in demand for depth completion methods. Most previous methods achieve satisfactory accuracy on standard benchmarks. However, they extensively rely on heavy models to handle diverse structures and require additional run time on multimodal data. In this paper, we present an efficient method of depth completion. We propose a grouped fusion strategy for efficiently extracting depth and guidance features in parallel and fusing them naturally in the feature spaces to achieve high performance. Instead of a monolithic architecture, we employ cascaded hourglass networks, each of which is specialized for certain structures and has a lightweight architecture. Given the sparsity of the depth maps, we downsample the inputs to multiple scales to further accelerate the computation. Our model runs at over 39 FPS on an embedded GPU with high-resolution inputs. Evaluations on the KITTI benchmark demonstrate that the proposed model is an ideal approach for real-world applications. Ang Li 0025, Zejian Yuan, Yonggen Ling, Wanchao Chi, Shenghao Zhang 0001 |
ICPR | 5 |
| 2020 | Attention-Oriented Action Recognition for Real- Time Human-Robot InteractionabstractDespite the notable progress made in action recognition tasks, not much work has been done in action recognition specifically for human-robot interaction. In this paper, we deeply explore the characteristics of the action recognition task in interaction scenes and propose an attention-oriented multi-level network framework to meet the need for real-time interaction. Specifically, a Pre-Attention network is employed to roughly focus on the interactor in the scene at low resolution firstly and then perform fine-grained pose estimation at high resolution. The other compact CNN receives the extracted skeleton sequence as input for action recognition, utilizing attention-like mechanisms to capture local spatial-temporal patterns and global semantic information effectively. To evaluate our approach, we construct a new action dataset specially for the recognition task in interaction scenes. Experimental results on our dataset and high efficiency (112 fps at 640 × 480 RGBD) on the mobile computing platform (Nvidia Jetson AGX Xavier) demonstrate excellent applicability of our method on action recognition in real-time human-robot interaction. Ziyi Yin 0001, Zejian Yuan, Wanchao Chi, Yonggen Ling, Shenghao Zhang 0001 |
ICPR | 7 |
| 2020 | A Multi-Scale Guided Cascade Hourglass Network for Depth CompletionabstractDepth completion, a task to estimate the dense depth map from sparse measurement under the guidance from the high-resolution image, is essential to many computer vision applications. Most previous methods building on fully convolutional networks can not handle diverse patterns in the depth map efficiently and effectively. We propose a multi-scale guided cascade hourglass network to tackle this problem. Structures at different levels are captured by specialized hourglasses in the cascade network with sparse inputs in various sizes. An encoder extracts multi-scale features from color image to provide deep guidance for all the hourglasses. A multi-scale training strategy further activates the effect of cascade stages. With the role of each sub-module divided explicitly, we can implement components with simple architectures. Extensive experiments show that our lightweight model achieves competitive results compared with state-of-the-art in KITTI depth completion benchmark, with low complexity in run-time. Ang Li 0025, Zejian Yuan, Yonggen Ling, Wanchao Chi, Shenghao Zhang 0001 |
WACV | 5 |
| 2020 | Crowded Human Detection via an Anchor-pair NetworkabstractThis paper presents an anchor-pair network for crowded human detection, which can overcome and solve the difficulties caused by occlusion in crowded scenes. Specifically, we use a function-aware network structure to extract more distinctive and discriminative features for head and full-body respectively, and then a CNN module is also exploited to fuse the features by learning the correlations between head and full-body to reduce crowd errors. Meanwhile, a novel paired form for anchors, denoted as anchor-pair, is proposed to estimate the head regions and full-body regions simultaneously. Furthermore, a new ingenious Joint-NMS is introduced to perform on the detected head and full-body box pairs, which produces significant performance improvement in heavily occluded scenarios at tiny computational cost. Our anchor-pair network achieves a state-of-the-art result on the CrowdHuman dataset which reduces the MR−2to 55.43%, achieving 11.59% relative improvement over our dataset baseline. Jinguo Zhu, Zejian Yuan, Wanchao Chi, Yonggen Ling, Shenghao Zhang 0001 |
WACV | 6 |