Wenjie Song 0001

dblp:174/4345-1 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
13since 2021 · last 2025
0000-0002-4121-908XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 5 since 2021Systems, architecture and hardware · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 TASeg: Text-aware RGB-T Semantic Segmentation based on Fine-tuning Vision Foundation Models
abstract
Reliable semantic segmentation of open environments is essential for intelligent systems, yet significant problems remain: 1) Existing RGB-T semantic segmentation models mainly rely on low-level visual features and lack high-level textual information, which struggle with accurate segmentation when categories share similar visual characteristics. 2) While SAM excels in instance-level segmentation, integrating it with thermal images and text is hindered by modality heterogeneity and computational inefficiency. To address these, we propose TASeg, a text-aware RGB-T segmentation framework by using Low-Rank Adaptation (LoRA) fine-tuning technology to adapt vision foundation models. Specifically, we propose a Dynamic Feature Fusion Module (DFFM) in the image encoder, which effectively merges features from multiple visual modalities while freezing SAM’s original transformer blocks. Additionally, we incorporate CLIP-generated text embeddings in the mask decoder to enable semantic alignment, which further rectifies the classification error and improves the semantic understanding accuracy. Experimental results across diverse datasets demonstrate that our method achieves superior performance in challenging scenarios with fewer trainable parameters.
Te Cui, Qitong Chu, Wenjie Song 0001, Yi Yang 0009, Yufeng Yue
IROS4
2025 Motion Control of a Hybrid Self-Reconfigurable Wheel-Legged Dual-Arm Robot
abstract
Current wheeled bipedal robots face significant mobility challenges when traversing discontinuous terrain such as gaps and step-like obstacles, and suffer from substantial dynamic inefficiencies. This paper presents a hybrid self-reconfigurable wheel-legged dual-arm robot equipped with an active docking mechanism, enabling transitions between wheeled bipedal and multi-wheel-legged configurations. Based on a self-developed robotic platform, this work addresses key control challenges in articulated multi-wheel-legged mode and proposes a novel distributed operation paradigm for wheeled bipedal robots. Each module utilizes its manipulators for stable grasping of elevated objects and collaborative tasks, while the multi-unit system achieves efficient, high-load, and stable locomotion. To manage the control complexities in multimodal operation, we develop a unified modular control architecture integrating Virtual Model Control (VMC) and Linear Quadratic Regulator (LQR). For the articulated multi-wheel-legged mode, a body-posture controller regulates global body configuration, and a turning controller adjusts the wheelbase and roll angle via distributed actuation to manage the passive degrees of freedom (DoF) at the articulation points. Experimental validation using a physical prototype confirms the effectiveness and practicality of the proposed approach.
Hong Du, Peng Qiu, Yi Yang 0009, Wenjie Song 0001
IROS5
2024 Risk-Inspired Aerial Active Exploration for Enhancing Autonomous Driving of UGV in Unknown Off-Road Environments
abstract
Unknown area exploration is a crucial but challenging task for autonomous driving of unmanned ground vehicles (UGV) in unknown off-road environments. However, the exploration efficiency of a single UGV is low due to its limited sensing range. To solve this problem, this paper proposes a risk-inspired aerial active exploration system, which utilizes the flexibility and field of view advantages of Unmanned Aerial Vehicles (UAV) to guide the UGV in unknown off-road environments. Firstly, a fast terrain risk mapping method that can be used for both UAV and UGV is developed. This method efficiently combines quadtree and hash table data structure to enable UAV to analyze large scale terrain point cloud in real time. Based on the risk mapping result, a risk-inspired active exploration method is proposed to actively search a safe reference path for the UGV, which introduces terrain risk information into the process of travel point selection. Finally, the reference path is gradually generated and optimized, so that the UGV can safely and smoothly follow the path to the target location. Compared with single UGV exploration system, our approach reduces the overall path risk by 26.8% in simulated experiments, showing that the proposed system can enhance autonomous driving of the UGV and help it effectively avoid high-risk areas in unknown off-road environments.
Rongchuan Wang, Mengyin Fu, Yi Yang 0009, Wenjie Song 0001
ICRA5
2024 Self-supervised Monocular Depth Estimation in Challenging Environments Based on Illumination Compensation PoseNet
abstract
Self-supervised depth estimation has attracted much attention due to its ability to improve the 3D perception capabilities of unmanned systems. However, existing unsupervised frameworks rely on the assumption of photometric consistency, which may not hold in challenging environments such as night-time, rainy nights, or snowy winters due to complex lighting and reflections, resulting in inconsistent photometry across different frames for the same pixel. To address this problem, we propose a self-supervised monocular depth estimation unified framework that can handle these complex scenarios, which has the following characteristics: (1) an Illumination Compensation PoseNet (ICP) is designed, which is based on the classic Phong illumination theory and compensates for lighting changes in adjacent frames by estimating per-pixel transformations; (2) a Dual-Axis Transformer (DAT) block is proposed as the backbone network of the depth encoder, which infers the depth of local repeat-texture areas through spatial-channel dual-dimensional global context information of images. Experimental results demonstrate that our approach achieves state-of-the-art depth estimation results in complex environments on the challenging Oxford RobotCar dataset.
Shengyu Hou, Wenjie Song 0001, Rongchuan Wang, Meiling Wang 0002, Yi Yang 0009, Mengyin Fu
IROS2
2024 Self-Supervised Monocular Depth Estimation for All-Day Images Based on Dual-Axis Transformer
abstract
All-day self-supervised monocular depth estimation has strong practical significance for autonomous systems to continuously perceive the 3D information of the world. However, night-time scenes pose challenges of weak texture and violating the brightness consistency assumption due to low illumination and varying lighting, respectively, which easily leads to most existing self-supervised models only being able to handle day-time scenes. To address this problem, we propose a self-supervised monocular depth estimation unified framework that can handle all-day scenarios, which has three features: (1) an Illumination Compensation PoseNet (ICP) is designed, which is based on the classic Phong illumination theory and compensates for lighting changes in adjacent frames by estimating per-pixel transformations; (2) a Dual-Axis Transformer (DAT) block is proposed as the backbone network of the depth encoder, which infers the depth of local low-illumination areas through spatial-channel dual-dimensional global context information of night-time images; (3) a cross-layer Adaptive Fusion Module (AFM) is introduced between multiple DAT blocks, which learns attention weights between different layer features and adaptively fuses cross-layer features using the learned weights, enhancing the complementarity of different layer features. This work was evaluated on multiple datasets, including: RobotCar, Waymo and KITTI datasets, achieving state-of-the-art results in both day-time and night-time scenarios.
Shengyu Hou, Mengyin Fu, Rongchuan Wang, Yi Yang 0009, Wenjie Song 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 Dynamic Voxels Based on Ego-Conditioned Prediction: An Integrated Spatio-Temporal Framework for Motion Planning
abstract
Prediction is a vital component of motion planning for autonomous vehicles (AVs). By reasoning about the possible behavior of other target agents, the ego vehicle (EV) can navigate safely, efficiently, and politely. However, most of the existing work overlooks the interdependencies of the prediction and planning module, only connecting them in a sequential pipeline or underexploring the prediction results in the planning module. In this work, we propose a framework that integrates the prediction and planning module with three highlights. First, we propose an ego-conditioned model for causal prediction, with the introduced edge-featured graph transformer model, the impact the ego future maneuver poses to the target vehicles is demonstrated. Second, we develop a motion planner based on ‘dynamic voxels’ in the spatio-temporal domain, enabling the time-to-collision criterion evaluation and the optimal trajectory generation in continuous space. Third, the prediction and planning modules are coupled in a closed-loop and efficient form. Specifically, taking each maneuver as a cluster, representative trajectory primitives are generated for conditional prediction, and conversely, prediction results are used to score the primitives as guidance, which alleviates the duplicated callback of the prediction module. The simulations are conducted in overtaking, merging, unprotected left turns, and also scenarios with imperfect social behaviors. The comparison studies demonstrate the better safety assurance and efficiency of the proposed model, and the ablation experiments further reveal the effectiveness of the new ideas.
Ting Zhang 0014, Mengyin Fu, Wenjie Song 0001, Yi Yang 0009, Alexandre Alahi
IEEE Trans. Intell. Transp. Syst.3
2023 Conflict-constrained Multi-agent Reinforcement Learning Method for Parking Trajectory Planning
abstract
Automated Valet Parking (AVP) has been exten-sively researched as an important application of autonomous driving. Considering the high dynamics and density of real parking lots, a system that considers multiple vehicles simultaneously is more robust and efficient than a single vehicle setting as in most studies. In this paper, we propose a dis-tributed Multi-agent Reinforcement Learning(MARL) method for coordinating multiple vehicles in the framework of an AVP system. This method utilizes traditional trajectory planning to accelerate the learning process and introduces collision conflict constraints for policy optimization to mitigate the path conflict problem. In contrast to other centralized multi-agent path finding methods, the proposed approach is scalable, distributed, and adapts to dynamic stochastic scenarios. We train the models in random scenarios and validate in several artificially designed complex parking scenarios where vehicles are always disturbed by dynamic and static obstacles. Experimental results show that our approach mitigates path conflicts and excels in terms of success rate and efficiency.
Meiling Wang 0002, Yi Yang 0009, Wenjie Song 0001
ICRA4
2023 Joint Learning of Image Deblurring and Depth Estimation Through Adversarial Multi-Task Network
abstract
Self-supervised monocular depth estimation methods have achieved remarkable results on natural clear images. However, it is still a serious challenge to directly recover depth information from blurred images caused by long-time exposure while camera fast moving. To address this issue, we propose a unified framework for simultaneous deblurring and depth estimation (SDDE), which has higher coupling performance and flexibility compared with the simple concatenation strategy of deblurring model and depth estimation model. This framework mainly benefits from three features: 1) a novel Task-aware Fusion Module (TFM) to adaptively select the most relevant intermediate shared features for the dual decoder network by aggregating multi-scale features, 2) a unique Spatial Interaction Module (SIM) to learn higher-order representation in the encoder stage to better describe complex boundaries of different classes in high-dimensional space, and focuses on the task-related region by modeling the pairwise spatial correlation of the holistic tensor, 3) a Priors-Based Composite Regularization term to jointly optimize the shared encoder-dual decoder network. This work was evaluated on multiple datasets, including: Stereo blur, KITTI,NYUv2, REDS and our own large-scale stereo blur dataset, resulting in state-of-the-art results for depth estimation and image deblurring, respectively.
Shengyu Hou, Mengyin Fu, Wenjie Song 0001
IEEE Trans. Circuits Syst. Video Technol.3
2023 Risk-Aware Decision-Making and Planning Using Prediction-Guided Strategy Tree for the Uncontrolled Intersections
abstract
Uncontrolled intersections with interaction and uncertainties are challenging for autonomous vehicles (AV) to manage. In this work, we propose a decision-making model specific to intersections with emphasis on three aspects. First, behavior estimation of the social vehicles’ (SVs) is essential for risk avoidance. We try to improve prediction accuracy by predicting the intentions and driving styles of SVs in advance and doing adaptive goal sampling. Second, the uncertainty from the prediction results should be considered in the decision-making process. For this, a risk-aware framework is developed, composed of a Subordinate Driver (SD) and a Primary Driver (PD) for decision-making and planning. Particularly, in SD, the prediction-guided strategy tree is built to search for an optimal strategy with observation and action branch trimming, which employs the prediction results for risk assessment. In PD, to mimic the both-way negotiation among vehicles, the level-k game model is deployed to determine the action in the players’ best interest and update the estimation of driving styles. Third, the generated maneuver is required to be evaluated in a closed-loop simulation. A ‘semi-autonomous’ control model is designed, which is a combination of the dataset and the stochastic sampling model. The results of ablation experiments verify the function of each module. The case studies and comparison experiments demonstrate the effectiveness of the framework in highly interactive intersections.
Ting Zhang 0014, Mengyin Fu, Wenjie Song 0001
IEEE Trans. Intell. Transp. Syst.3
2022 Trajectory Prediction-Based Local Spatio-Temporal Navigation Map for Autonomous Driving in Dynamic Highway Environments
abstract
Autonomous driving, including intelligent decision-making and path planning, in dynamic environments (like highway) is significantly more difficult than the navigation in static scenarios because of the additional time dimension. Therefore, correlating the time dimension and the space dimension through prediction to create a spatio-temporal navigation map can make decision-making and path planning in such kinds of environment much easier. In this article, NGSIM data is analysed and processed from the perspective of the ego-vehicle (using the data as an ego-vehicle’s perception results). Based on the data, we develop an LSTM (Long-Short Term Memory)-based framework to predict possible trajectories of multiple surrounding vehicles within a certain range of the ego-vehicle. Then, the multiple predicted trajectories in a series of continuous dynamic highway scenes are projected into a spatio-temporal domain to create an octree map. Thus, dynamic targets and static obstacles can be unified into the same domain or map so that the dynamic disturbance problem for autonomous driving in highway environments can be resolved. Experimental results show that the proposed model is capable of predicting all the future trajectories around the ego-vehicle efficiently and the corresponding spatio-temporal map can be generated accurately in different dynamic scenarios.
Mengyin Fu, Ting Zhang 0014, Wenjie Song 0001, Yi Yang 0009, Meiling Wang 0002
IEEE Trans. Intell. Transp. Syst.3
2022 Action-State Joint Learning-Based Vehicle Taillight Recognition in Diverse Actual Traffic Scenes
abstract
As the vital factor of vehicle behavior understanding and prediction, vehicle taillight recognition is an important technology for autonomous driving, especially in diverse actual traffic scenes full of dynamic interactive traffic participants. However, in practical application, it always faces many challenges, such as ‘variable lighting conditions’, ‘non-uniform taillight standards’ and ‘random relative observation pose’, which lead to few mature solutions in current common autopilot systems. This work proposes an action-state joint learning-based vehicle taillight recognition method on the basis of vehicles detection and tracking, which takes both taillight state features and time series features into account, consequently getting practicable results even in complex actual scenes. In detail, vehicle tracking sequence is used as input and split into pieces through a sliding window. Then, a CNN-LSTM model is applied to simultaneously identify the action features of brake lights and turn signals, dividing taillight actions into five categories: None, Brake_on, Brake_off, Left_turn, Right_turn. Next, the brightness of high-position brake light is extracted through semantic segmentation and combined with taillight actions to form higher-level features for taillight state sequence analysis. Finally, an undirected graph model is used to establish the long-term dependence between successive pieces by analysing the higher-level features, thus inferring the continuous taillight state into:$off$,$brake$,$left$,$right$. Datasets including daytime, nighttime, congested road, highway, etc. were collected, tested and published in our work to demonstrate its effectiveness and practicability.
Wenjie Song 0001, Shixian Liu, Ting Zhang 0014, Yi Yang 0009, Mengyin Fu
IEEE Trans. Intell. Transp. Syst.1
2022 Trajectory Planning Based on Spatio-Temporal Map With Collision Avoidance Guaranteed by Safety Strip
abstract
Trajectory planning for the unmanned vehicle in the complex environment has always been a challenging task. Planned trajectory with the corresponding target velocity or acceleration sequence must be collision-free guaranteed and as comfortable as possible on the premise of obeying the traffic rules and interaction with other dynamic social vehicles. To meet this requirement, this paper proposes a framework for trajectory planning based on spatio-temporal map. Due to the time layer architecture in the map, the trajectory can be generated with velocity and acceleration simultaneously, and the whole trajectory is constrained within a ‘safety strip’, resulting in an efficient and safety guaranteed trajectory. The framework is composed of three sections: rough search, fine optimization and safety strip-based collision avoidance. For rough search, we propose an improved A* algorithm implemented in the discrete time layer to find out the suboptimal states efficiently. In fine optimization, the B-spline curve is exploited to connect the searched states into a continuous trajectory. And the optimal control points of B-spline are further grouped into several segments, forming the safety strip which is actually the distribution space of the planned trajectory. If necessary, an adjustment will be applied to keep the strip away from the collision zone, making the entire trajectory completely collision-free. Experiments on both public dataset and self-driving simulator show that the proposed framework can adapt to different kinds of complex traffic scenes well.
Ting Zhang 0014, Mengyin Fu, Wenjie Song 0001, Yi Yang 0009, Meiling Wang 0002
IEEE Trans. Intell. Transp. Syst.3
2022 A Unified Framework Integrating Decision Making and Trajectory Planning Based on Spatio-Temporal Voxels for Highway Autonomous Driving
abstract
Intelligent decision making and efficient trajectory planning are closely related in autonomous driving technology, especially in highway environment full of dynamic interactive traffic participants. This work integrates them into a unified hierarchical framework with long-term behavior planning (LTBP) and short-term dynamic planning (STDP) running in two parallel threads with different horizon, consequently forming a closed-loop maneuver and trajectory planning system that can react to the dynamic environment effectively and efficiently. In LTBP, a novel voxel structure and the ‘voxel expansion’ algorithm are proposed for the generation of driving corridors in 3D configuration, which involves the prediction states of surrounding vehicles. By using Dijkstra search, the maneuver with minimal cost is determined in form of voxel sequences, then a quadratic programming (QP) problem is constructed for solving the optimal trajectory. And in STDP, another small-scaled QP problem is performed to track or adjust the reference trajectory from LTBP in response to the dynamic obstacles. Meanwhile, a Responsibility-Sensitive Safety (RSS) Checker keeps running at high frequency for real-time feedback to ensure security. Experiments on real data collected in different highway scenarios demonstrate the effectiveness and efficiency of our work.
Ting Zhang 0014, Wenjie Song 0001, Mengyin Fu, Yi Yang 0009, Xiaohui Tian, Meiling Wang 0002
IEEE Trans. Intell. Transp. Syst.2
2020 Trajectory Prediction based on Constraints of Vehicle Kinematics and Social Interaction†
abstract
Trajectory prediction for vehicles is a popular subject since it is beneficial for efficient and secure trajectory planning. In structured traffic scenarios, the behaviour and motion of vehicles are heavily dependent on the social interaction constraints, such as road geometry and surrounding vehicles, and the kinematics model constraints, such as continuous heading and maximum acceleration. To take these factors into account, we analyse the particular characteristics of driving vehicles and propose a model that predicts the possible and feasible trajectory for host vehicle in 3 seconds. In this model, the trajectory of host vehicle takes the center-line as reference, imitates the leader vehicle and focuses on the social vehicles through attention concentration mechanism (ACM) with spatial and temporal information encoded in a fusion hidden state. Furthermore, in order to make the trajectory feasible for vehicle dynamics and kinematics, we introduce a prediction diagnosis method to check the continuous heading and maximum acceleration condition, pruning and adjusting the prediction candidates. Experiments on released public datasets show that this framework can well evaluate the traffic interactions and forecast the trajectory more accurately than common networks.
Ting Zhang 0014, Mengyin Fu, Wenjie Song 0001, Yi Yang 0009, Meiling Wang 0002
SMC3
2018 Real-Time Obstacles Detection and Status Classification for Collision Warning in a Vehicle Active Safety System
abstract
This paper presents real-time obstacles detection and their status classification method for collision warning in the vehicle active safety system. Specifically, stereo cameras and millimeter wave (mmw)-radar are fused to help the driving ego-vehicle to find “Danger” or “Potential Danger” in a timely way through combining with the vehicle kinematic model. The proposed method makes full use of the unique advantages of stereo cameras and mmw-radar to sense the environment through several modules. Cameras are mainly used to detect the near or lateral dynamic objects and to obtain the obstacles region of interest (ROI) considering its rich information and high sensitivity to the lateral displacement, while far or longitudinal relative dynamic objects are detected by mmw-radar according to its observational ability to make up for the disadvantage of cameras. In detail, a cameras detector utilizes ”error vectors” rather than the optical flow to obtain dynamic classes through two times clustering. Mmw-radar mainly detects relative dynamic objects, whose absolute speed can be computed according to the ego-vehicle's state. Then, the detected objects of these two detectors are integrated in an obstacles ROI map, which is obtained through an UV-disparity obstacles detection algorithm to get the final dynamic and relative dynamic objects. Finally, they are classified by comparing them with a dangerous area that is acquired according to the vehicle kinematic model in a special vehicle coordinate system, which is fixed to the ground temporarily. This method is tested on our mobile platforms and the results prove that it can work effectively even though the ego-vehicle drives quickly.
Wenjie Song 0001, Yi Yang 0009, Mengyin Fu, Fan Qiu, Meiling Wang 0002
IEEE Trans. Intell. Transp. Syst.1
2017 Real-time lane detection and forward collision warning system based on stereo vision
abstract
This paper presents a real-time and robust lane detection and forward collision warning technique based on stereo cameras. First, obstacles image is obtained through stereo matching and UV-disparity segmentation algorithm. Then, Inverse Perspective Mapping(IPM) and Sobel filtering are conducted to generate a low-noise top view of the road by fusing the obstacles image and the original image. Next, Hough Transformation for the top view map is completed and the extreme points(poles) are calculated as the detected lanes according to the traffic lanes model. Besides, the host lane is selected or supplemented among all the detected lanes and the nearest obstacle in this host lane is detected for the forward collision warning. Experimental results on the public data set indicate that our method can work effectively and real-timely in the normal structured environment.
Wenjie Song 0001, Mengyin Fu, Yi Yang 0009, Meiling Wang 0002, Xinyu Wang 0018, Alain L. Kornhauser
Intelligent Vehicles Symposium1
2017 Intersection scan model and probability inference for vision based small-scale urban intersection detection
abstract
Large-scale intersections stamped on maps have diverse visual features for detection, while small-scale urban intersections are hard to be identified especially when GPS signals are missing. In this paper, we propose a Hidden Markov Model (HMM) based small-scale intersection detection method utilizing monocular vision. We extract visual cues of road transformations and dynamic vehicles' tracks, and then design an Intersection Scan Model to obtain the potential traversable direction of the current road, which is the primary criterion of the intersection estimation. For better performances, we take the detections of consecutive frames into consideration and finally integrate them into HMM to estimate the probabilities of intersections. Results from KITTI datasets and real-world experiments have shown the functionality of the presented approach.
Yi Yang 0009, Hao Li 0075, Hao Zhu 0002, Songtian Shang, Ningyi Lyu, Wenjie Song 0001
Intelligent Vehicles Symposium6