EDBT 2026 Demo / reviewers in the wild / expert
Chao Sun 0006
dblp:54/3957-6
· DBLP profile ↗
17ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0002-9324-0892ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 7 since 2021Computer networks · 7 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robo-DETR: Robustness-Aware Depth-Guided Transformer for Monocular 3-D Object Detection Under Adverse Visual ConditionsabstractMonocular 3D object detection is attractive for autonomous driving due to its low cost, but its performance degrades severely under adverse visual conditions that corrupts appearance cues and destabilizes depth-related representations. We propose Robo-DETR, a robustness-aware depth-guided transformer framework that combines minimal visual-conditioned input preconditioning with detector-internal depth reliability correction. Specifically, a lightweight Condition Recognition and Enhancement Block stabilizes low-level visual cues, while a Condition-aware Depth-guided Block refines depth logits via foreground-aware supervision and structured dynamic/region-level adjustment to suppress visual-induced depth noise. A dual-branch transformer decoder further promotes depth-appearance consistency through modality-specific cross-attention and fusion. Experiments on KITTI-C and the real-world TJ4DRadSet demonstrate consistent improvements over competitive monocular detectors under diverse degradations, and ablations validate the contribution of each component. Jiaru Zhong, Zitong Chen, Yueran Zhao, Bo Wang 0143, Jianghao Leng, Chao Sun 0006 |
IEEE Internet Things J. | 8 |
| 2026 | UT-Planner: Energy-Efficient Trajectory Planner for Autonomous Ground Vehicles on Uneven TerrainabstractInternet-of-Things (IoT) sensing and V2X connectivity are expanding the deployment of electrified unmanned ground vehicles (UGVs) in rugged off-road environments. However, safe and energy-efficient navigation on uneven terrain remains challenging because terrain understanding, vehicle attitude prediction, and trajectory refinement are often handled in isolation. To address this limitation, this paper presents UT-Planner, an integrated energy- and stability-aware planning framework for IoT-enabled UGVs that exploits high-fidelity point-cloud maps. First, a wheel-contact-based attitude prediction module simulates four-wheel interactions on local point-cloud patches to estimate future body attitude and terrain roughness before traversal, thereby forming a predictive traversability layer. Second, an Eco-A∗ search augments Hybrid-A∗ with a Dubins-based energy heuristic that combines path length with terrain-aware energy surrogates to generate short and energy-favorable coarse paths. Third, a multi-objective trajectory optimizer refines the coarse path by jointly minimizing smoothness, gravity-aligned energy expenditure, and attitude deviation. Experiments on four representative uneven-terrain maps show that, relative to strong baselines, UT-Planner shortens the final 3D path length by 9.4%, reduces real energy consumption by 12.4%, improves trajectory smoothness by 13.3%, lowers the maximum tracking error by 19.8%, and also reduces peak body tilt. These results demonstrate a practical route to safe and energy-efficient off-road autonomous navigation. Changjiu Ning, Chao Sun 0006, Xiongji Yang, Zhishuai Huang, Da Wen, Zitong Chen, Jianghao Leng |
IEEE Internet Things J. | 2 |
| 2026 | Delving Into the Secrets of BEV 3D Object Detection in Autonomous Driving: A Comprehensive Surveyabstract3D object detection plays a crucial role in autonomous driving, with Bird’s Eye View (BEV) becoming increasingly popular for its rich contextual information, ease of multi-modal fusion, and scalability. Despite its advantages, current BEV-based 3D detection methods still face significant challenges, including multi-modal fusion, communication bottlenecks, robustness under varying conditions, and safety concerns. This survey provides a systematic review of recent advancements in BEV perception, and meanwhile organizes a research based on these focal areas. It spans a broad range of perspectives, offering valuable insights for future perception research. Additionally, this survey explores the influence of emerging technologies, such as large language models and end-to-end frameworks on enhancing BEV perception capabilities, focusing on improving performance and robustness. Key future directions would include: 1) advancement from isolated vehicle perception to vehicle-to-everything (V2X) cooperative perception; 2) evolution from single-modal to integrated multi-modal fusion; 3) shift from simulated environments to real-world applications; and 4) transition from hierarchical perception frameworks to interpretable, end-to-end large-scale models. Yueran Zhao, Jiaru Zhong, Bo Wang 0143, Chao Sun 0006, Fengchun Sun |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2026 | PillarID: Rethinking Backbone Network Designs for Pillar-Based 3D Object Detection in Infrastructure Point CloudabstractIn recent years, vehicle-centric point cloud 3D object detection has been widely explored and effectively developed. However, due to differences in the placement of sensors, infrastructure-centric point cloud 3D object detection, which is an important component of the Intelligent Transportation System (ITS), has not received sufficient attention as well as effective network architecture design. Based on the difference in perspective of the infrastructure point cloud, We discover that the roadside point cloud is denser and with a higher coverage compared to the vehicle-side in the pillar representation, resulting in a narrowing of the performance difference between dense pillar and sparse pillar backbone networks in roadside scenes. Inspired by this insight, a network based on the dense backbone is proposed, dubbed PillarID. It utilizes Single-stride Cross-stage Dense-backbone (SCD) to obtains efficient computation through channel degradation, split, and cross-stage connection, and benefits from the rich context of the roadside point cloud based on single-stride. Further, Hierarchical Receptive-field Expansion (HRE) are used to address the receptive field constraints of single-stride backbone. Extensive experiments reveal that our PillarID achieves effective designs in terms of architecture and renders the state-of-the-art performance on the popular large-scale roadside benchmark: DAIR-V2X-I and RCooper. The code is available athttps://github.com/zhangzhang2024/PillarID Zhang Zhang 0006, Chao Sun 0006, Bo Wang 0143, Da Wen |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2026 | HeightFormer: Learning Height Prediction in Voxel Features for Roadside Vision Centric 3D Object Detection via TransformerabstractRoadside vision centric 3D object detection has received increasing attention in recent years. It expands the perception range of autonomous vehicles, enhances the road safety. Previous methods focused on predicting per-pixel height rather than depth, making significant gains in roadside visual perception. While it is limited by the perspective property of near-large and far-small on image features, making it difficult for network to understand real dimension of objects in the 3D world. Bird’s Eye View (BEV) features and voxel features present the real distribution of objects in 3D world compared to the image features. However, BEV features tend to lose details due to the lack of explicit height information, and voxel features are computationally expensive. Inspired by this insight, an efficient framework learning height prediction in voxel features via transformer is proposed, dubbed HeightFormer. It groups the voxel features into local height sequences, and utilize attention mechanism to obtain height distribution prediction. Subsequently, the local height sequences are reassembled to generate accurate voxel features. The proposed method is applied to two large-scale roadside benchmarks, DAIR-V2X-I and Rope3D. Extensive experiments are performed and the HeightFormer outperforms the state-of-the-art methods in roadside vision centric 3D object detection task. Code will be athttps://github.com/zhangzhang2024/HeightFormer Zhang Zhang 0006, Chao Sun 0006, Da Wen, Tianze Wang, Jianghao Leng |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | UT-MPC: Manifold-Based Model Predictive Control With Dynamic Weighting and Feedback for Vehicle Trajectory Tracking on Uneven TerrainabstractAs autonomous vehicle (AV) technology advance, their ability to navigate uneven terrains, such as mountainous areas, becomes increasingly important. However, current trajectory tracking methods struggle with tracking accuracy and stability due to insufficient consideration of terrain slopes and vehicle kinematics. In this article, we propose a manifold-based model predictive control framework designed for Ackermann steering electrified vehicles on uneven terrains. This method models the trajectory tracking control on manifolds, utilizing acceleration and front-wheel steering angle as control inputs to enhance control stability. To mitigate model inaccuracies and enhance the controller’s adaptability, we dynamically adjust the controller’s objective function weights based on trajectory curvature and integrate a PID feedback mechanism to provide real-time compensation for vehicle speed and steering angle. When fitting the surface terrain equation, we select sparse key points along the reference trajectory, achieving lightweight computation while maintaining high-fitting accuracy. The experimental and simulation results demonstrate that the proposed UT-MPC controller improves tracking performance by 53.74% and 42.68% compared to four baseline methods when tracking a trajectory on different uneven terrain maps. Real-world experiments have also demonstrated the effectiveness of UT-MPC. Changjiu Ning, Bo Wang 0143, Jianghao Leng, Zitong Chen, Da Wen, Chao Sun 0006 |
IEEE Internet Things J. | 6 |
| 2025 | HSIGCN: Hierarchical Spatial Interaction Graph Convolutional Network Considering Group Behavior for Pedestrian Trajectory PredictionabstractPedestrian trajectory prediction is crucial in various fields, but remains challenging due to complex spatial interactions. While existing pedestrian trajectory prediction methods show promise, they often fail to capture these dynamics effectively. To address this limitation, a Hierarchical Spatial Interaction Graph Convolutional Network (HSIGCN) is proposed to handle both group interactions and spatial interactions. Although previous methods have attempted to model group behaviors, they lack a comprehensive consideration of group interactions and often oversimplify the complex social dynamics in groups. HSIGCN introduces a novel group interaction mechanism that encompasses four types of interactions: all-pedestrian, intra-group, out-group, and inter-group interactions, enhancing the expressiveness in group behavior prediction. Furthermore, current approaches to spatial interaction sparsification either rely solely on prior-based or on learning-based methods. HSIGCN innovatively combines both approaches to form a mixed sparsification mechanism, effectively filtering all-pedestrian and out-group interactions. Additionally, existing prior-based methods fail to consider social factors comprehensively. HSIGCN takes into account the field of view (FOV), collision awareness, and distance factors to establish a more robust prior-based sparse function. Experimental results on ETH and UCY datasets demonstrate that the proposed method significantly outperforms baseline models, showcasing its potential to accurately predict pedestrian trajectories by effectively handling complex spatial interactions. Bo Wang 0143, Chao Sun 0006, Jianghao Leng, Zhishuai Huang, Zitong Chen |
IEEE Internet Things J. | 2 |
| 2025 | BEVMamba: Time Sequence Dense Bird's-Eye-View Perception Modeling With State Space ModelabstractBEV-based 3D perception with multi-frame images input is crucial for autonomous driving. However, current methods for temporal BEV perception fail to fully utilize long sequence features because of local fusion or high complexity. Recently, Mamba, a powerful temporal modeling network with linear complexity, has shown exceptional performance in various 2D vision tasks, but its application to 3D perception tasks remains unexplored. Therefore, this paper proposes a general BEV perception backbone named BEVMamba, which is the first work to leverage State Space Model for 3D perception. Built upon the BEVFormer, to adapt Mamba for 3D perception we first add Hybrid Positional Encoding to the BEV features, enabling the networks to be aware of their spatial-temporal position. In the Temporal SSM block, the proposed 3D Factorized Scan ensures that historical BEV features are enriched with global temporal-spatial information. Subsequently, the Spatial-Temporal Corridor Fusion aggregates all BEV features in a physically meaningful manner, achieving precise feature fusion. The reliable BEV features obtained by BEVMamba are used for various perception tasks, including 3D object detection and 3D occupancy prediction. Results on the nuScenes and Occ-3D nuScenes datasets show that BEVMamba outperforms its baseline BEVFormer in both dense and sparse perception tasks and demonstrates competitive performance compared to other methods, highlighting the potential of Mamba in 3D perception tasks. The code will be available at https://github.com/Liuxiaoaaa/bevmamba Jiaru Zhong, Chao Sun 0006 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | Height3D: A Roadside Visual Framework Based on Height Prediction in Real 3-D SpaceabstractIn recent years, vision-based roadside 3D object detection has received a great deal of attention, which is an important part of the Intelligent Transportation System (ITS). It extends the perception range beyond the limitations of Autonomous Vehicle (AV) and enhances road safety. While previous work mainly focuses on height prediction in image 2D space, which is limited by the perspective property of near-large and far-small on images, making it difficult for network to understand real dimension of targets in the 3D world. Inspired by this insight, a roadside visual framework Height3D based on height prediction in real 3D space, is proposed. Height Prediction Block (HPB) with explicit height supervision is proposed in real 3D space instead of in image 2D space to predict the height distribution of targets for roadside view transform. Also, Spatial Aware Block (SAB) is used to further extracts spatial context information in BEV space and enhances fine-grained BEV features. The proposed method is applied to two large-scale roadside benchmarks, DAIR-V2X-I and Rope3D. Extensive experiments are performed to verify its effectiveness. The proposed Height3D outperforms the state-of-the-art methods of (1.15, 7.37, 4.03) Average Precision (AP) for Vehicle, Pedestrian and Cyclist categories in 3D object detection task, respectively. Meanwhile, the proposed method achieves 31.55 FPS without using any CUDA or TensorRT acceleration. The code is available at https://github.com/zhangzhang2024/Height3D Zhang Zhang 0006, Chao Sun 0006, Bo Wang 0143, Da Wen, Qili Ning |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | LaneMapNet: Lane Network Recognization and HD Map Construction Using Curve Region Aware Temporal Bird's-Eye-View PerceptionabstractThe construction of local HD (High Definition) Map and Lane Network with onboard sensors is critical for autonomous vechicles and facilitates downstream tasks. In contrast to previous studies that treated building HD Map and Lane Network as two individual tasks, in this paper a unified BEV (Bird’s-Eye-View) perception framework is proposed with seperate decoders to realize two tasks simultaneously. In this paper, the gap between object detection and curve regression when using a DETR-like decoder is discussed and a curve region aware method is proposed to make up for the above gap. Specifically, a mechanism called Curve Region Aware Deformable Attention is designed with a bezier grid sampling module to guide the attention learning in bev features and structual loss regarding shapes of lanelines is also included. Moreover, a BEV spatial-temporal fusion method is introduced to better utilize historical features with minimal loss of spatial information. The results on NuScenes dataset show that our work has been close to or exceeded SOTAs (state-of-the-art) on both two tasks simultaneously. Jianghao Leng, Jiaru Zhong, Zhang Zhang 0006, Chao Sun 0006 |
IV | 5 |
| 2024 | SDAGCN: Sparse Directed Attention Graph Convolutional Network for Spatial Interaction in Pedestrian Trajectory PredictionabstractPedestrian trajectory prediction is crucial across various domains, but remains challenging due to complex spatial interactions. Existing Graph Convolutional Network (GCN) methods show promise but often fail to capture these dynamics effectively. To address this limitation, a Sparse Directed Attention Graph Convolutional Network (SDAGCN) is proposed to handle both social interactions among pedestrians and self-interactions within individuals. Traditional GCN-based methods often model social interactions as undirected or dense graphs. However, due to the field of view and awareness of collision avoidance of pedestrians, they tend to focus unilaterally on specific neighbors. To reflect this, SDAGCN constructs a sparse and directed spatial graph that considers these attributes innovatively. Furthermore, the attention weights of pedestrians towards their neighbors are closely tied to spatial conflicts. The conflicts are deeply influenced by relative velocity and distance. Therefore, these attributes are leveraged to calculate the attention weights. These two components form the Sparse Directed Attention (SDA) mechanism, which effectively discerns the influence of neighbors on a target pedestrian in various situations. Additionally, the self-interaction of each pedestrian is significantly influenced by their speed. To capture variations in self-interaction across different states, SDAGCN employs a single-layer perceptron with the square of pedestrian speed as input. Experiments conducted on the ETH and UCY datasets demonstrate that our method outperforms other GCN-based spatial interaction methods, showcasing its potential in accurately predicting pedestrian trajectories by effectively handling complex social and self-interactions. Chao Sun 0006, Bo Wang 0143, Jianghao Leng, Xiangchao Zhang |
IEEE Internet Things J. | 1 |
| 2024 | Interactive Left-Turning of Autonomous Vehicles at Uncontrolled IntersectionsabstractThis paper presents a novel interactive motion planning approach for left-turning of autonomous vehicles at uncontrolled intersections. A typical left-turning scenario that consists of agent vehicles and an autonomous ego vehicle is established. The autonomous ego vehicle aims to cross the uncontrolled intersection quickly and safely when meeting with the agent vehicles. A modified obstacle reciprocal collision avoidance (MORCA) prediction model is proposed to predict the trajectory of the agent vehicles with the considerations of longitudinal and lateral interaction-aware behaviors. The parameters of the MORCA prediction model are optimized using particle swarm optimization (PSO) with the inD datasets. Based on MORCA, a partially observable Markov decision process (POMDP) framework is established. The action and state spaces of the POMDP are extended reasonably to satisfy the requirement of real-time application. The simulation results show that the prediction precision is improved by 15% using the MORCA compared with the obstacle reciprocal collision avoidance (ORCA) prediction model. Meanwhile, the results manifest that the proposed MORCA-based POMDP planning method improves 43.8% in commuting efficiency compared with a traditional rule-based method, nearly as good as the vehicle-to-vehicle (V2V) communication is equipped. Note to Practitioners—The motivation of this article is to establish a safe and efficient planning framework for the left-turning of autonomous vehicles at uncontrolled intersections. Rule-based approaches were widely used in the previous studies but they are not optimal. In this paper, the lateral movements of vehicles at the intersection are considered for the first time. A MORCA prediction model is established and optimized to formulate a POMDP planning framework. The CARLA simulation demonstrates the effectiveness of the proposed MORCA-based POMDP planner. In future work, we will migrate the algorithm to a real-world test platform for experimentation. Chao Sun 0006, Jianghao Leng, Bing Lu 0005 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2023 | Dust concentration prediction model in thermal power plant using improved genetic algorithm
Bo Wang 0143, Xuliang Yao, Yongqing Jiang, Chao Sun 0006 |
Soft Comput. | 4 |
| 2022 | A Fast Optimal Speed Planning System in Arterial Roads for Intelligent and Connected VehiclesabstractSpeed planning system is generally equipped for intelligent and connected vehicles (ICVs). Under the circumstances of autonomous driving, an energy-optimal speed trajectory is usually desired, particularly on urban arterial roads with complex traffic conditions involved. However, the existing speed planning solutions in the literature have not dealt with the problem of time consuming. Also, the ego vehicle could not perfectly track a precalculated speed reference because of the dynamically varying traffic. Thus, optimal speed planning cannot always be guaranteed. In this article, a fast optimal speed planning system for complex urban driving situations is established through an adaptive hierarchical control framework. In the planning layer, dynamic programming (DP) and the interior-point optimizer are jointly used to compute the global speed trajectory with access to signal phase and timing (SPaT) information. The computational burden is greatly alleviated based on a weighted orientation graph assumption and problem decomposition. The following layer utilizes the Informer, which is a transformer-based model, to predict preceding vehicle speed. Then, a target-switching model-predictive controller (MPC) is adopted for global speed trajectory following and adaption. The proposed approach significantly reduces speed planning computation time compared to previous solutions. Simulation results based on real road traffic scenes manifest that 22.0% of energy is saved compared with human driving. Chao Sun 0006, Jianghao Leng, Fengchun Sun |
IEEE Internet Things J. | 1 |
| 2022 | An Eco-Driving Approach With Flow Uncertainty Tolerance for Connected Vehicles Against Waiting Queue Dynamics on Arterial RoadsabstractEco-driving incorporating multiple signalized intersections simultaneously has been proven to substantially benefit connected vehicles (CVs) in energy performance. However, ignoring the dynamic variation of waiting queues before downstream intersections may prevent CVs from following the obtained speed profile on security grounds. In this article, the dynamic variation of the waiting queue is modeled and predicted based on shockwave theory and data-driven-based traffic flow prediction. To formulate the waiting queues as additional time-varying constraints for optimization problems, an extended traffic signal model is constructed based on the prediction. Furthermore, a hierarchical optimization framework is proposed, under which the hybrid optimization problem is decomposed into a discrete problem and a continuous one. Monte Carlo simulation demonstrates that if the proposed eco-driving approach is implemented, failure to follow the reference speed profile decreases by 79.4%. Also, the fuel consumption can be saved by over 4% compared with approaches ignoring the waiting queue. Chao Sun 0006, Chuntao Zhang, Haiyang Yu 0002, Weiqiang Liang, Jianwei Li 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2020 | Optimal Eco-Driving Control of Connected and Autonomous Vehicles Through Signalized IntersectionsabstractThis article focuses on the speed planning problem for connected and automated vehicles (CAVs) communicating to traffic lights. The uncertainty of traffic signal timing for signalized intersections on the road is considered. The eco-driving problem is formulated as a data-driven chance-constrained robust optimization problem. Effective red-light duration (ERD) is defined as a random variable, and describes the feasible passing time through the signalized intersections. Usually, the true probability distribution for ERD is unknown. Consequently, a data-driven approach is adopted to formulate chance constraints based on empirical sample data. This incorporates robustness into the eco-driving control problem with respect to uncertain signal timing. Dynamic programming (DP) is employed to solve the optimization problem. The simulation results demonstrate that the proposed method can generate optimal speed reference trajectories with 40% less vehicle fuel consumption, while maintaining the arrival time at a similar level compared to a modified intelligent driver model (IDM). The proposed control approach significantly improves the controller's robustness in the face of uncertain signal timing, without requiring to know the distribution of the random variable a priori. Chao Sun 0006, Jacopo Guanetti, Francesco Borrelli, Scott J. Moura |
IEEE Internet Things J. | 1 |
| 2018 | Stochastic Model Predictive Control of Air Conditioning System for Electric Vehicles: Sensitivity Study, Comparison, and ImprovementabstractA stochastic model predictive controller (SMPC) of air conditioning (AC) system is proposed to improve the energy efficiency of electric vehicles (EVs). A Markov-chain based velocity predictor is adopted to provide a sense of the future disturbances over the SMPC control horizon. The sensitivity of electrified AC plant to solar radiation, ambient temperature, and relative air flow speed is quantificationally analyzed from an energy efficiency perspective. Three control approaches are compared in terms of the electricity consumption, cabin temperature, and comfort fluctuation, which include the proposed SMPC method, a generally used bang-bang controller, and dynamic programming as the benchmark. Real solar radiation and ambient temperature data are measured to validate the effectiveness of the SMPC. Comparison results illustrate that SMPC is able to improve the AC energy economy by 12% compared to the rule-based controller. The cabin temperature variation is reduced by more than 50.4%, resulting with a much better cabin comfort. Hongwen He, Hui Jia, Chao Sun 0006, Fengchun Sun |
IEEE Trans. Ind. Informatics | 3 |