Tongtong Cao

dblp:229/5950 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
9since 2021 · last 2026
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 7 since 2021Systems, architecture and hardware · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2026 Unigaussian: Driving Scene Reconstruction From Multiple Camera Models Via Unified Gaussian Representations
abstract
Urban scene reconstruction is crucial for real-world autonomous driving simulators. Although existing methods have achieved photorealistic reconstruction, they mostly focus on pinhole cameras and neglect fisheye cameras. In fact, how to effectively simulate fisheye cameras in driving scene remains an unsolved problem. In this work, we propose UniGaussian, a novel approach that learns a unified 3D Gaussian representation from multiple camera models for urban scene reconstruction in autonomous driving. Our contributions are two-fold. First, we propose a new differentiable rendering method that distorts 3D Gaussians using a series of affine transformations tailored to fisheye camera models. This addresses the compatibility issue of 3D Gaussian splatting with fisheye cameras, which is hindered by light ray distortion caused by lenses or mirrors. Besides, our method maintains real-time rendering while ensuring differentiability. Second, built on the differentiable rendering method, we design a new framework that learns a unified Gaussian representation from multiple camera models. By applying affine transformations to adapt different camera models and regularizing the shared Gaussians with supervision from different modalities, our framework learns a unified 3D Gaussian representation with input data from multiple sources and achieves holistic driving scene understanding. As a result, our approach models multiple sensors (pinhole and fisheye cameras) and modalities (depth, semantic, normal and LiDAR point clouds). Our experiments show that our method achieves superior rendering quality and fast rendering speed for driving scene simulation.
Guile Wu, Runhao Li, Zheyuan Yang, Tongtong Cao, Xingxin Chen
3DV6
2026 WALDO: Where Unseen Model-based 6D Pose Estimation Meets Occlusion
abstract
Accurate 6D object pose estimation is vital for robotics, augmented reality, and scene understanding. For seen objects, high accuracy is often attainable via per-object fine-tuning but generalizing to unseen objects remains a challenge. To address this problem, past arts assume access to CAD models at test time and typically follow a multistage pipeline to estimate poses: detect and segment the object, propose an initial pose, and then refine it. Under occlusion, however, the early-stage of such pipelines are prone to errors, which can propagate through the sequential processing, and consequently degrade the performance. To remedy this shortcoming, we propose four novel extensions to model-based 6D pose estimation methods: (i) a dynamic non-uniform dense sampling strategy that focuses computation on visible regions, reducing occlusion-induced errors; (ii) a multi-hypothesis inference mechanism that retains several confidence-ranked pose candidates, mitigating brittle single-path failures; (iii) iterative refinement to progressively improve pose accuracy; and (iv) series of occlusion-focused training augmentations that strengthen robustness and generalization. Furthermore, we propose a new weighted by visibility metric for evaluation under occlusion to minimize the bias in the existing protocols. Via extensive empirical evaluations, we show that our proposed approach achieves more than 5% improvement in accuracy on ICBIN and more than 2% on BOP dataset benchmarks, while achieving ≈ 3× faster inference.
Sajjad Pakdamansavoji, Yintao Ma, Amir Rasouli, Tongtong Cao
WACV4
2025 Validity Learning on Failures: Mitigating the Distribution Shift in Autonomous Vehicle Planning
abstract
The planning problem constitutes a fundamental aspect of the autonomous driving framework. Recent strides in representation learning have empowered vehicles to comprehend their surrounding environments, thereby facilitating the integration of learning-based planning strategies. Among these approaches, Imitation Learning stands out due to its notable training efficiency. However, traditional Imitation Learning methodologies encounter challenges associated with the co-variate shift phenomenon. We propose Validity Learning on Failures, VL(on failure), as a remedy to address this issue. The essence of our method lies in deploying a pre-trained planner across diverse scenarios. Instances where the planner deviates from its immediate objectives, such as maintaining a safe distance from obstacles or adhering to traffic rules, are flagged as failures. The states corresponding to these failures are compiled into a new dataset, termed the failure dataset. Notably, the absence of expert annotations for this data precludes the applicability of standard imitation learning approaches. To facilitate learning from the closed-loop mistakes, we introduce the VL objective which aims to discern valid trajectories within the current environmental context. Experimental evaluations conducted on both reactive CARLA simulation and non-reactive log-replay simulations reveal substantial enhancements in closed-loop metrics such as Score, Progress, and Success Rate, underscoring the effectiveness of the proposed methodology. Further evaluations against the Bench2Drive benchmark demonstrate that VL(on failure) outperforms the state-of-the-art methods by a large margin.
Fazel Arasteh, Mohammed Elmahgiubi, Behzad Khamidehi, Hamidreza Mirkhani, Weize Zhang, Tongtong Cao, Kasra Rezaee
ICRA6
2025 AutoSplat: Constrained Gaussian Splatting for Autonomous Driving Scene Reconstruction
abstract
Realistic scene reconstruction and view synthesis are essential for advancing autonomous driving systems by simulating safety-critical scenarios. 3D Gaussian Splatting (3DGS) excels in real-time rendering and static scene reconstructions but struggles with modeling driving scenarios due to complex backgrounds, dynamic objects, and sparse camera views. We propose AutoSplat, a framework employing Gaussian splatting to realistically reconstruct autonomous driving scenes. By imposing geometric constraints on Gaussians representing the road and sky regions, our method enables multi-view consistent simulation of challenging scenarios, including lane changes. Leveraging 3D templates, we introduce a reflected Gaussian consistency constraint to supervise both the visible and unseen side of foreground objects. Moreover, to model the dynamic appearance of foreground objects, we estimate temporally-dependent residual spherical harmonics for each foreground Gaussian. Extensive experiments on Pandaset [1] and KITTI [2] demonstrate that AutoSplat outperforms state-of-the-art methods in scene reconstruction and novel view synthesis across diverse driving scenarios. Our project page can be found here: https://autosplat.github.io/
Mustafa Khan, Hamid R. Fazlali, Dhruv Sharma, Tongtong Cao, Dongfeng Bai
ICRA4
2025 ET-Plan-Bench: Embodied Task-level Planning Benchmark Towards Spatial-Temporal Cognition with Foundation Models
abstract
Recent advancements in Large Language Models (LLMs) have catalyzed numerous efforts to apply these technologies to embodied tasks, with a particular focus on high-level task planning and task decomposition. LLMs face challenges in understanding the physical world, especially regarding spatial, temporal, and causal relationships among objects and actions. Moreover, the current benchmarks for evaluating these relationships are limited. To further investigate this domain, we introduce a novel embodied task planning benchmark, ET-Plan-Bench. This benchmark features a controllable and diverse array of embodied tasks, varying in levels of difficulty and complexity. It is designed to evaluate two critical dimensions of LLMs’ application in embodied task understanding: spatial understanding (including relation constraints and occlusion of target objects) and temporal and causal comprehension of sequences of actions within an environment. Utilizing multi-source simulators as the backend simulator, ET-Plan-Bench provides immediate environmental feedback to LLMs, enabling dynamic interaction with the environment and the capacity for re-planning as necessary. We evaluated state-of-the-art open-source and closed-source foundational models, including GPT-4, Llama, and Mistral, using our proposed benchmark. While these models perform adequately on simple navigation tasks, their performance significantly deteriorates when con-fronted with tasks that demand a deeper understanding of spatial, temporal, and causal relationships. Consequently, our benchmark distinguishes itself as a large-scale, quantifiable, highly automated, and fine-grained diagnostic framework that presents a substantial challenge to the latest foundational models. We hope it will inspire and propel further research in embodied task planning utilizing foundational models. Code available at: https://github.com/ET-Plan-Bench/ET-Plan-Bench
Yuening Wang, Hongjian Gu, Atia Hamidizadeh, Zhanguang Zhang, Yuecheng Liu, David Gamaliel Arcos Bravo, Junyi Dong, Shunbo Zhou, Tongtong Cao, Xingyue Quan, Yuzheng Zhuang, Yingxue Zhang 0001, Jianye Hao
IROS11
2025 3DArticCyclists: Generating Synthetic Articulated 8D Pose-Controllable Cyclist Data for Computer Vision Applications
abstract
In Autonomous Driving (AD) Perception, cyclists are considered safety-critical scene objects. Commonly used publicly-available AD datasets typically contain large amounts of car and vehicle object instances but a low number of cyclist instances, usually with limited appearance and pose diversity. This cyclist training data scarcity problem not only limits the generalization of deep-learning perception models for cyclist semantic segmentation, pose estimation, and cyclist crossing intention prediction, but also limits research on new cyclist-related tasks such as fine-grained cyclist pose estimation and spatio-temporal analysis under complex interactions between humans and articulated objects. To address this data scarcity problem, in this paper we propose a framework to generate synthetic dynamic 3D cyclist data assets that can be used to generate training data for different tasks. In our framework, we designed a methodology for creating a new part-based multi-view articulated synthetic 3D bicycle dataset that we call 3DArticBikes that we use to train a 3D Gaussian Splatting (3DGS)-based reconstruction and image rendering method. We then propose a parametric bicycle 3DGS composition model to assemble 8-DoF pose-controllable 3D bicycles. Finally, using dynamic information from cyclist videos, we build a complete synthetic dynamic 3D cyclist (rider pedaling a bicycle) by reposing a selectable synthetic 3D person, while automatically placing the rider onto one of our new articulated 3D bicycles using a proposed 3D Keypoint optimization-based Inverse Kinematics pose refinement. We present both, qualitative and quantitative results where we compare our generated cyclists against those from a recent stable diffusion-based method.
Eduardo R. Corral-Soto, Dongfeng Bai, Tongtong Cao
IV5
2023 GPA-3D: Geometry-aware Prototype Alignment for Unsupervised Domain Adaptive 3D Object Detection from Point Clouds
abstract
LiDAR-based 3D detection has made great progress in recent years. However, the performance of 3D detectors is considerably limited when deployed in unseen environments, owing to the severe domain gap problem. Existing domain adaptive 3D detection methods do not adequately consider the problem of the distributional discrepancy in feature space, thereby hindering generalization of detectors across domains. In this work, we propose a novel unsupervised domain adaptive 3D detection framework, namely Geometry-aware Prototype Alignment (GPA-3D), which explicitly leverages the intrinsic geometric relationship from point cloud objects to reduce the feature discrepancy, thus facilitating cross-domain transferring. Specifically, GPA-3D assigns a series of tailored and learnable prototypes to point cloud objects with distinct geometric structures. Each prototype aligns BEV (bird’s-eye-view) features derived from corresponding point cloud objects on source and target domains, reducing the distributional discrepancy and achieving better adaptation. The evaluation results obtained on various benchmarks, including Waymo, nuScenes and KITTI, demonstrate the superiority of our GPA-3D over the state-of-the-art approaches for different adaptation scenarios. The MindSpore version code will be publicly available at https://github.com/Liz66666/GPA3D.
Tongtong Cao, Wankou Yang
ICCV3
2023 Towards Universal LiDAR-Based 3D Object Detection by Multi-Domain Knowledge Transfer
abstract
Contemporary LiDAR-based 3D object detection methods mostly focus on single-domain learning or cross-domain adaptive learning. However, for autonomous driving systems, optimizing a specific LiDAR-based 3D object detector for each domain is costly and lacks of scalability in real-world deployment. It is desirable to train a universal LiDAR-based 3D object detector from multiple domains. In this work, we propose the first attempt to explore multi-domain learning and generalization for LiDAR-based 3D object detection. We show that jointly optimizing a 3D object detector from multiple domains achieves better generalization capability compared to the conventional single-domain learning model. To explore informative knowledge across domains towards a universal 3D object detector, we propose a multi-domain knowledge transfer framework with universal feature transformation. This approach leverages spatial-wise and channel-wise knowledge across domains to learn universal feature representations, so it facilitates to optimize a universal 3D object detector for deployment at different domains. Extensive experiments on four benchmark datasets (Waymo, KITTI, NuScenes and ONCE) show the superiority of our approach over the state-of-the-art approaches for multi-domain learning and generalization in LiDAR-based 3D object detection.
Guile Wu, Tongtong Cao, Xingxin Chen
ICCV2
2022 How to Build a Curb Dataset with LiDAR Data for Autonomous Driving
abstract
Curbs are one of the essential elements of urban and highway traffic environments. Robust curb detection provides road structure information for motion planning in an autonomous driving system. Commonly, video cameras and 3D LiDARs are mounted on autonomous vehicles for curb detection. However, camera-based methods suffer from challenging illumination conditions. During the long period of time before wide application of Deep Neural Network (DNN) with point clouds, LiDAR-based curb detection methods are based on hand-crafted features, which suffer from poor detection in some complex scenes. Recently, DNN-based dynamic object detection using LiDAR data has become prevalent, while few works pay attention to curb detection with a DNN approach due to lack of labeled data. A dataset with curb annotations or an efficient curb labeling approach, hence, is of high demand. In this paper, we present how to build a curb dataset with LiDAR data for autonomous driving highly automatically. Firstly, a Semantic High Definition map (SHD map) in a global coordinate frame is generated by applying both SLAM and semantic segmentation on consecutive LiDAR frames. Next, a Road HD map (RHD map) is generated from the SHD map by removing its dynamic noise caused by road users e.g. cars. After that, a Curb Instance map (CI map) can be obtained from the filtered RHD map by a series of curb point extraction and growing. Finally, the CI map can be projected back to single frames for direct, highly automatic curb labeling. In order to validate our proposed labeling method, on top of an open public LiDAR semantic dataset SemanticKITTI [1], an additional curb dataset is built. We run both semantic segmentation and instance segmentation methods on this built dataset. Experimental results show that the curb annotations have good consistency and accuracy. We released this dataset and it is publicly available at https://download.mindspore.cn.
Dongfeng Bai, Tongtong Cao
ICRA2
2020 Improvement of Decorrelation-Based OCT Angiography by an Adaptive Spatial-Temporal Kernel in Monitoring Stimulus-Evoked Hemodynamic Responses
abstract
Complex decorrelation-based OCT angiography (OCTA) has the potential for monitoring hemodynamic activities in a label-free, high-resolution, and quantitative manner. To improve the measurement dynamic range and uncertainty of blood flow, an adaptive spatial-temporal (ST) kernel was proposed for decorrelation estimation and it was validated through a theoretical simulation and experimental measurements. The ensemble size in the decorrelation computation was effectively enlarged by collecting samples of the phasor pair in both the spatial and temporal dimensions. The spatial sub-kernel size was adaptively changed to suppress the influence of bulk motion in the temporal dimension by solving a maximum entropy model. Using the flow phantom, it was observed that the decorrelation dynamic range presented an increase of ~49% and the uncertainty exhibited a decrease of ~40% and ~38% in the saturation and background limits, respectively. In monitoring the stimulus-evoked hemodynamic response, the extended dynamic range enabled an improvement of ~180% in the separability between different stimulation modes. Furthermore, the suppressed uncertainty and motion artifacts allowed a reliable temporal analysis of the hemodynamic response. The proposed adaptive ST-kernel will greatly promote the development of decorrelation-based quantitative OCTA in hemodynamic studies.
Ruixiang Chen, Tongtong Cao, Huakun Li, Peng Li 0034
IEEE Trans. Medical Imaging4
2018 A Grid Projection Method Based on Ultrasonic Sensor for Parking Space Detection
abstract
In this paper a parking space detection method is proposed. It is based on the grid map projection and utilizes the ultrasonic sensors. Firstly, it applies a virtual grid map to quantify the observing target space and build a coordinate system. Then the probable outline of the targets can be deduced from the ultrasonic wave echo signal and projected into the grid map. The boundary of the object can be obtained by comparing the overlap number of the outline in the grid with a threshold. In this way, the size of the target space can be calculated and determined whether it is proper for parking a vehicle. The method is simple while remaining effective. It can be well validated by the experimental results and the accuracy is better than 0.2m.
Yunfeng Shao 0001, Pengzhen Chen, Tongtong Cao
IGARSS3