Cheng Zhao 0002

dblp:93/3598-2 · also Henry Cheng Zhao · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
13since 2021 · last 2025
0000-0001-8502-3233ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 8 since 2021Systems, architecture and hardware · 5 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2025 SplatFlow: Self-Supervised Dynamic Gaussian Splatting in Neural Motion Flow Field for Autonomous Driving
abstract
Most existing Dynamic Gaussian Splatting methods for complex dynamic urban scenarios rely on accurate object-level supervision from expensive manual labeling, limiting their scalability in real-world applications. In this paper, we introduce SplatFlow, a Self-Supervised Dynamic Gaussian Splatting within Neural Motion Flow Fields (NMFF) to learn 4D space-time representations without requiring tracked 3D bounding boxes, enabling accurate dynamic scene reconstruction and novel view RGB/depth/flow synthesis. SplatFlow designs a unified framework to seamlessly integrate time-dependent 4D Gaussian representation within NMFF, where NMFF is a set of implicit functions to model temporal motions of both LiDAR points and Gaussians as continuous motion flow fields. Leveraging NMFF, SplatFlow effectively decomposes static background and dynamic objects, representing them with 3D and 4D Gaussian primitives, respectively. NMFF also models the correspondences of each 4D Gaussian across time, which aggregates temporal features to enhance cross-view consistency of dynamic components. SplatFlow further improves dynamic object identification by distilling features from 2D foundation models into 4D space-time representation. Comprehensive evaluations conducted on the Waymo and KITTI Datasets validate SplatFlow’s state-of-the-art (SOTA) performance for both image reconstruction and novel view synthesis in dynamic urban scenarios.
Su Sun, Cheng Zhao 0002, Zhuoyang Sun, Victor Y. Chen
CVPR2
2024 Behind the Veil: Enhanced Indoor 3D Scene Reconstruction with Occluded Surfaces Completion
abstract
In this paper, we present a novel indoor 3D reconstruction method with occluded surface completion, given a sequence of depth readings. Prior state-of-the-art (SOTA) methods only focus on the reconstruction of the visible areas in a scene, neglecting the invisible areas due to the occlusions, e.g., the contact surface between furniture, occluded wall and floor. Our method tackles the task of completing the occluded scene surfaces, resulting in a complete 3D scene mesh. The core idea of our method is learning 3D geometry prior from various complete scenes to infer the occluded geometry of an unseen scene from solely depth measurements. We design a coarse-fine hierarchical octree representation coupled with a dual-decoder architecture, i.e., Geo-decoder and 3D Inpainter, which jointly reconstructs the complete 3D scene geometry. The Geo-decoder with detailed representation at fine levels is optimized online for each scene to reconstruct visible surfaces. The 3D Inpainter with abstract representation at coarse levels is trained offline using various scenes to complete occluded surfaces. As a result, while the Geo-decoder is specialized for an individual scene, the 3D Inpainter can be generally applied across different scenes. We evaluate the proposed method on the 3D Completed Room Scene (3D-CRS) and iTHOR datasets, significantly outperforming the SOTA methods by a gain of 16.8% and 24.2% in terms of the completeness of 3D reconstruction. 3D-CRS dataset including a complete 3D mesh of each scene is provided on project webpage11https://github.com/BoschRHI3NA/3D-CRS-dataset.
Su Sun, Cheng Zhao 0002, Yuliang Guo, Ruoyu Wang 0012, Xinyu Huang 0001, Victor Y. Chen, Liu Ren 0001
CVPR2
2024 SUP-NeRF: A Streamlined Unification of Pose Estimation and NeRF for Monocular 3D Object Reconstruction
Yuliang Guo, Abhinav Kumar 0004, Cheng Zhao 0002, Ruoyu Wang 0012, Xinyu Huang 0001, Liu Ren 0001
ECCV (69)3
2024 TCLC-GS: Tightly Coupled LiDAR-Camera Gaussian Splatting for Autonomous Driving: Supplementary Materials
Cheng Zhao 0002, Su Sun, Ruoyu Wang 0012, Yuliang Guo, Jun-Jun Wan, Xinyu Huang 0001, Victor Y. Chen, Liu Ren 0001
ECCV (63)1
2023 IF-Based Trajectory Planning and Cooperative Control for Transportation System of Cable Suspended Payload With Multi UAVs
abstract
In this paper, we tackle the control and trajectory planning problems for the cooperative transportation system of cable-suspended payload with multi Unmanned Aerial Vehicles (UAVs). Firstly, a payload controller is presented considering the dynamic coupling between the UAV and the payload to accomplish the active suppression of payload swing and the complex payload trajectory tracking. Secondly, different from the simplification of obstacles in most approaches, we propose three Insetting Formation (IF) algorithms for the complete obstacle shape to generate collision-free waypoints for the cooperative transportation system. An IF strategy is proposed by integrating three IF algorithms to improve the success rate of obstacle avoidance and reduce the algorithm complexity for performing the aggressive flight. Finally, we verify the robustness and high performance of the proposed algorithm through benchmark comparison and real-world experiments. Moreover, our source code is released as an open-source ros package.
Yu Zhang 0091, Cheng Zhao 0002, Jiuxiang Dong
IROS3
2023 3D Copy-Paste: Physically Plausible Object Insertion for Monocular 3D Detection
abstract
A major challenge in monocular 3D object detection is the limited diversity and quantity of objects in real datasets. While augmenting real scenes with virtual objects holds promise to improve both the diversity and quantity of the objects, it remains elusive due to the lack of an effective 3D object insertion method in complex real captured scenes. In this work, we study augmenting complex real indoor scenes with virtual objects for monocular 3D object detection. The main challenge is to automatically identify plausible physical properties for virtual assets (e.g., locations, appearances, sizes, etc.) in cluttered real scenes. To address this challenge, we propose a physically plausible indoor 3D object insertion approach to automatically copy virtual objects and paste them into real scenes. The resulting objects in scenes have 3D bounding boxes with plausible physical locations and appearances. In particular, our method first identifies physically feasible locations and poses for the inserted objects to prevent collisions with the existing room layout. Subsequently, it estimates spatially-varying illumination for the insertion location, enabling the immersive blending of the virtual objects into the original scene with plausible appearances and cast shadows. We show that our augmentation method significantly improves existing monocular 3D object models and achieves state-of-the-art performance. For the first time, we demonstrate that a physically plausible 3D object insertion, serving as a generative data augmentation technique, can lead to significant improvements for discriminative downstream tasks such as monocular 3D object detection. Project website: https://gyhandy.github.io/3D-Copy-Paste/.
Yunhao Ge, Hong-Xing Yu, Cheng Zhao 0002, Yuliang Guo, Xinyu Huang 0001, Liu Ren 0001, Laurent Itti, Jiajun Wu 0001
NeurIPS3
2023 EventPoint: Self-Supervised Interest Point Detection and Description for Event-based Camera
abstract
This paper proposes a self-supervised learned local detector and descriptor, called EventPoint, for event stream/camera tracking and registration. Event-based cameras have grown in popularity because of their biological inspiration and low power consumption. Despite this, applying local features directly to the event stream is difficult due to its peculiar data structure. We propose a new time-surface-like event stream representation method called Ten-code. The event stream data processed by Tencode can obtain the pixel-level positioning of interest points while also simultaneously extracting descriptors through a neural network. Instead of using costly and unreliable manual annotation, our network leverages the prior knowledge of local feature extraction on color images and conducts self-supervised learning via homographic and spatio-temporal adaptation. To the best of our knowledge, our proposed method is the first research on event-based local features learning using a deep neural network. We provide comprehensive experiments of feature point detection and matching, and three public datasets are used for evaluation (i.e. DSEC, N-Caltech101, and HVGA ATIS Corner Dataset). The experimental findings demonstrate that our method outperforms SOTA in terms of feature point detection and description.
Ze Huang, Li Sun 0005, Cheng Zhao 0002, Songzhi Su
WACV3
2022 VEFNet: an Event-RGB Cross Modality Fusion Network for Visual Place Recognition
abstract
Visual Place Recognition (VPR) on natural image is challenging due to the illumination variance and seasonal changes. In terms of long-term localization, the emerging event stream cameras are naturally resilient to appearance changes. In this paper, we propose a novel multi-modal network, e.g. VEFNet for VPR by learning location-specific cross RGB-event modality feature representations. Specifically, we firstly extract dense visual features via shared Convolutional Neural Network (CNN) backbone from RGB and event frames separately. Then, two branch features are fed to the cross-modality attention module to establish correspondences between the dual-modality. We also employ a self-attention module to enhance the contextual integration within densely encoded features. Finally, the learned global descriptor is used as the place representation of the dual-modality inputs for VPR. Experimental results demonstrate the state-of-the-art (SOTA) performance on the public datasets
Ze Huang, Li Sun 0005, Cheng Zhao 0002, Min Huang 0004, Songzhi Su
ICIP4
2022 USPoint: Self-Supervised Interest Point Detection and Description for Ultrasound-Probe Motion Estimation During Fine-Adjustment Standard Fetal Plane Finding
Cheng Zhao 0002, Richard Droste, Lior Drukker, Aris T. Papageorghiou, J. Alison Noble
MICCAI (8)1
2021 NDT-Transformer: Large-Scale 3D Point Cloud Localisation using the Normal Distribution Transform Representation
abstract
3D point cloud-based place recognition is highly demanded by autonomous driving in GPS-challenged environments and serves as an essential component (i.e. loop-closure detection) in lidar-based SLAM systems. This paper proposes a novel approach, named NDT-Transformer, for real-time and large-scale place recognition using 3D point clouds. Specifically, a 3D Normal Distribution Transform (NDT) representation is employed to condense the raw, dense 3D point cloud as probabilistic distributions (NDT cells) to provide the geometrical shape description. Then a novel NDT-Transformer network learns a global descriptor from a set of 3D NDT cell representations. Benefiting from the NDT representation and NDT-Transformer network, the learned global descriptors are enriched with both geometrical and contextual information. Finally, descriptor retrieval is achieved using a query-database for place recognition. Compared to the state-of-the-art methods, the proposed approach achieves an improvement of 7.52% on average top 1 recall and 2.73% on average top 1% recall on the Oxford Robotcar benchmark.
Cheng Zhao 0002, Daniel Adolfsson, Songzhi Su, Yang Gao 0002, Tom Duckett, Li Sun 0005
ICRA2
2021 Robust and Long-term Monocular Teach and Repeat Navigation using a Single-experience Map
abstract
This paper presents a robust monocular visual teach-and-repeat (VT&R) navigation system for long-term operation in outdoor environments. The approach leverages deep-learned descriptors to deal with the high illumination variance of the real world. In particular, a tailored self-supervised descriptor, DarkPoint, is proposed for autonomous navigation in outdoor environments. We seamlessly integrate the localisation with control, in which proportional–integral control is used to eliminate the visual error with the pitfall of the unknown depth. Consequently, our approach achieves day-to-night navigation using a single-experience map and is able to repeat complex and fast manoeuvres. To verify our approach, we performed a vast array of navigation experiments in various outdoor environments, where both navigation accuracy and robustness of the proposed system are investigated. The experimental results show that our approach is superior to the baseline method with regards to accuracy and robustness.
Li Sun 0005, Marwan Taher, Christopher Wild, Cheng Zhao 0002, Yu Zhang 0091, Filip Majer, Zhi Yan 0001, Tomás Krajník, Tony J. Prescott, Tom Duckett
IROS4
2021 Monocular Teach-and-Repeat Navigation using a Deep Steering Network with Scale Estimation
abstract
This paper proposes a novel monocular teach-and-repeat navigation system with the capability of scale awareness, i.e. the absolute distance between observation and goal images. It decomposes the navigation task into a sequence of visual servoing sub-tasks to approach consecutive goal/node images in a topological map. To be specific, a novel hybrid model, named deep steering network is proposed to infer the navigation primitives according to the learned local feature and scale for each visual servoing sub-task. A novel architecture, Scale-Transformer, is developed to estimate the absolute scale between the observation and goal image pair from a set of matched deep representations to assist repeating navigation. The experiments demonstrate that our scale-aware teach-and-repeat method achieves satisfying navigation accuracy, and converges faster than the monocular methods without scale correction given an inaccurate initial pose. The proposed network is integrated into an onboard system deployed on a real robot to achieve real-time navigation in a real environment. A demonstration video can be found online: https://youtu.be/ctlwDaMKnHw
Cheng Zhao 0002, Li Sun 0005, Tomás Krajník, Tom Duckett, Zhi Yan 0001
IROS1
2021 Visual-Assisted Probe Movement Guidance for Obstetric Ultrasound Scanning Using Landmark Retrieval
Cheng Zhao 0002, Richard Droste, Lior Drukker, Aris T. Papageorghiou, J. Alison Noble
MICCAI (8)1
2019 Recurrent Kalman Networks: Factorized Inference in High-Dimensional Deep Feature Spaces
abstract
In order to integrate uncertainty estimates into deep time-series modelling, Kalman Filters (KFs) (Kalman et al., 1960) have been integrated with deep learning models, however, such approaches typically rely on approximate inference tech- niques such as variational inference which makes learning more complex and often less scalable due to approximation errors. We propose a new deep approach to Kalman filtering which can be learned directly in an end-to-end manner using backpropagation without additional approximations. Our approach uses a high-dimensional factorized latent state representation for which the Kalman updates simplify to scalar operations and thus avoids hard to backpropagate, computationally heavy and potentially unstable matrix inversions. Moreover, we use locally linear dynamic models to efficiently propagate the latent state to the next time step. The resulting network architecture, which we call Recurrent Kalman Network (RKN), can be used for any time-series data, similar to a LSTM (Hochreiter & Schmidhuber, 1997) but uses an explicit representation of uncertainty. As shown by our experiments, the RKN obtains much more accurate uncertainty estimates than an LSTM or Gated Recurrent Units (GRUs) (Cho et al., 2014) while also showing a slightly improved prediction performance and outperforms various recent generative models on an image imputation task.
Philipp Becker, Harit Pandya, Gregor H. W. Gebhardt, Cheng Zhao 0002, C. James Taylor, Gerhard Neumann
ICML4
2018 Synthetic View Generation for Absolute Pose Regression and Image Synthesis
Pulak Purkait, Cheng Zhao 0002, Christopher Zach
BMVC2
2018 Learning Monocular Visual Odometry with Dense 3D Mapping from Dense 3D Flow
abstract
This paper introduces a fully deep learning approach to monocular SLAM, which can perform simultaneous localization using a neural network for learning visual odometry (L-VO) and dense 3D mapping. Dense 2D flow and a depth image are generated from monocular images by sub-networks, which are then used by a 3D flow associated layer in the L-VO network to generate dense 3D flow. Given this 3D flow, the dual-stream L-VO network can then predict the 6DOF relative pose and furthermore reconstruct the vehicle trajectory. In order to learn the correlation between motion directions, the Bivariate Gaussian modeling is employed in the loss function. The L-VO network achieves an overall performance of 2.68 % for average translational error and 0.0143°/m for average rotational error on the KITTI odometry benchmark. Moreover, the learned depth is leveraged to generate a dense 3D map. As a result, an entire visual SLAM system, that is, learning monocular odometry combined with dense 3D mapping, is achieved.
Cheng Zhao 0002, Li Sun 0005, Pulak Purkait, Tom Duckett, Rustam Stolkin
IROS1