EDBT 2026 Demo / reviewers in the wild / expert
Jing Lian 0002
dblp:135/1468-2
· DBLP profile ↗
16ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0002-2095-5563ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hierarchical Recursive Interaction and Multi-Stage Goal-Guided Mechanism for Multimodal Trajectory PredictionabstractIn highly dynamic and complex autonomous driving environments, accurately predicting agents’ future multimodal trajectories still faces challenges such as modeling diverse social interactions, capturing dynamic intents, and ensuring prediction consistency. To address these issues, this paper proposes a novel trajectory prediction model that integrates a Hierarchical Recursive Interaction Network (HRINet) and a multi-stage goal-guided mechanism (GoalNet), aiming to improve prediction accuracy, stability, and plausibility. Specifically, we design a HRINet with local and global attention mechanisms to recursively model various social interactions, while progressively integrating map semantic information to enhance the model’s understanding of traffic scenes. Meanwhile, inspired by the divide-and-conquer approach, the proposed GoalNet first estimates fine-grained multi-stage goal lane segments along the path. These goals are then used to continuously guide and constrain the trajectory generation process, effectively reducing error accumulation and improving stability. In addition, we construct a dynamic goal candidate area that combines domain knowledge and traffic rules to filter out unreasonable goals, thereby enhancing the plausibility and consistency of the predictions. Experimental results on nuScenes, INTERACTION, and Waymo Open Motion Dataset (WOMD) show that our model achieves state-of-the-art performance in multiple key metrics, maintains a trade-off between prediction accuracy, model complexity, and inference latency, and shows high stability and consistency in predictions. Jing Lian 0002, Zhenfeng Wang, Jian Zhao 0029, Jun Hu 0020 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | Pedestrian Facial Attention Detection Using Deep Fusion and Multi-Modal Fusion ClassifierabstractPedestrian facial attention plays an essential role in autonomous driving scenarios where a vehicle has to handle complex interactions with pedestrians. By inferring whether pedestrians are making eye contact with the ego-vehicle, the intention of pedestrians can be deduced. However, traditional gaze estimation and eye detection algorithms have limitations in complex traffic scenes due to the lower resolution caused by spatial distance and the lack of visual features caused by occlusion. To address these limitations, this study proposes an innovative pedestrian facial attention detection framework. The proposed framework adopts a deep feature fusion strategy to achieve a deep-level fusion of visual features and semantic pose features. Moreover, a multi-modal fusion classifier that helps discover the cross-model spatial interactive representation from the feature maps, thus enhancing the robustness of model generalization, is proposed. The proposed framework is verified by experiments on public JAAD and LOOK datasets. The experimental results demonstrate the effectiveness of the proposed framework, indicating that it can achieve better performance compared to the existing methods. Jing Lian 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Dual-STGAT: Dual Spatio-Temporal Graph Attention Networks With Feature Fusion for Pedestrian Crossing Intention PredictionabstractPedestrian intent prediction is critical for autonomous driving, as accurately predicting crossing intentions helps prevent collisions and ensures the safety of both pedestrians and passengers. Recent research has focused on vision-based deep neural networks for this task, but challenges remain. First, current methods suffer from low efficiency in multi-feature fusion and unreliable predictions under challenging conditions. Additionally, real-time performance is essential in practical applications, so the efficiency of the algorithm is crucial. To address these issues, we propose a novel architecture, Dual-STGAT, which uses a dual-level spatio-temporal graph network to extract pedestrian pose and scene interaction features, reducing information loss and improving feature fusion efficiency. The model captures key features of pedestrian behavior and the surrounding environment through two modules: the Pedestrian Module and the Scene Module. The Pedestrian Module extracts pedestrian motion features using a spatio-temporal graph attention network, while the Scene Module models interactions between pedestrians and surrounding objects by integrating visual, semantic, and motion information through a graph network. Extensive experiments conducted on the PIE and JAAD datasets show that Dual-STGAT achieves over 90% accuracy in pedestrian crossing intention prediction, with inference latency close to 5ms, making it well-suited for large-scale production autonomous driving systems that demand both performance and computational efficiency. Jing Lian 0002, Yiyang Luo, Ge Guo 0001, Weiwei Ren |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | DFT-Net: A Bimodal Object Detection Algorithm for Complex Traffic EnvironmentsabstractIn response to the challenge of single-modal sensors struggling to rapidly and accurately detect in complex traffic environments, a DFT-Net detection algorithm based on the fusion of visible light and infrared bimodal features is proposed. Firstly, the algorithm constructs a dual modal feature extraction network and designs an efficient aggregation module with stronger feature extraction capability. Secondly, the algorithm utilizes a framework of self-attention and cross-attention mechanisms to fuse the features of visible light and infrared modalities. This framework facilitates information interaction between different modalities through query-guided mechanisms. To reduce the computational load of the dual modal feature transformer module, a spatial feature compression module is introduced. Finally, in order to mitigate visible noise and low infrared resolution, this paper combines Transformer and CBAM attention mechanism. The designed algorithm is compared with various detection algorithms on multiple datasets, and the results demonstrate its effectiveness. On the FLIR-gan dataset, compared to the YOLOv7 detection algorithms for visible light and infrared, our DFT-Net achieved an accuracy improvement of 26.8% and 26.7% respectively, and 2.8% improvement compared to the bimodal detection algorithm CFT. This indicates that the algorithm exhibits good detection performance in complex traffic environments. Jing Lian 0002, Jun Hu 0020 |
INDIN | 1 |
| 2024 | Cascade fusion of multi-modal and multi-source feature fusion by the attention for three-dimensional object detection
Fengning Yu, Jing Lian 0002, Jian Zhao 0029 |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | A Novel Framework for Scene Graph Generation via Prior KnowledgeabstractThe scene graph generation aims to recognize objects and infer the relationships between them, which can provide a comprehensive understanding of image visual perception. However, the long-tailed issue of relations remains challenging for scene graph generation. This paper proposes a novel framework based on knowledge-driven data-driven joining to address the long-tail issues in scene graph generation. The proposed framework consists of two modules: the relation inference module and the prior knowledge learning module. The relation inference module aims to learn the relational features of entity pairs in images and the structural features of scene graphs. The prior knowledge learning module aims to learn the triplet representation from the knowledge graph and use it as prior knowledge to provide logical guidance and constraints for relation inference. This provides prior bias for relation inference to transfer the bias towards head categories to reasonable categories, thereby mitigating the long-tail problem. Experiment results indicate that the proposed framework outperforms on Visual Genome datasets and that the generated scene graph relation is logically reasonable. Jing Lian 0002, Jian Zhao 0029 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Efficient Vehicle Trajectory Prediction With Goal Lane Segments and Dual-Stream Cross AttentionabstractReal-time and efficient prediction of plausible trajectories of surrounding traffic agents is essential for autonomous driving. In reality, the motion of agents depends not only on their goal intents, but is also constrained by road topology. Especially for vehicles, lane geometry can significantly influence their future trajectories. In this paper, the constraining and guiding roles of lane networks are investigated, and a novel vehicle trajectory prediction model based on the goal lane segment is proposed. Specifically, a Dual-Stream Cross Attention Module (DSCAM) is developed that incorporates goal lane segment prediction into the process of collecting various interaction information. This achieves scene-consistent predictions while reducing resource consumption and inference latency. Then, learnable refinement tokens are used to adaptively refine coarse-grained trajectories during the feature decoding process, with the goal of improving prediction accuracy and ensuring temporal consistency. Extensive experiments with the Argoverse-1 and nuScenes datasets demonstrate that our model performs competitively and well-balanced. Notably, on the nuScenes, our model with only 0.44 million (M) parameters has an inference latency of less than six milliseconds (ms), making it ideal for large-scale production autonomous driving systems that require high computational resources, operational efficiency, and prediction performance. Jing Lian 0002, Jian Zhao 0029, Jun Hu 0020 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | PTP-STGCN: Pedestrian Trajectory Prediction Based on a Spatio-temporal Graph Convolutional Neural Network
Jing Lian 0002, Weiwei Ren, Yafu Zhou |
Appl. Intell. | 1 |
| 2023 | Early intention prediction of pedestrians using contextual attention-based LSTM
Jing Lian 0002, Fengning Yu, Yafu Zhou |
Multim. Tools Appl. | 1 |
| 2023 | Robustly Integrated Design of Plug-in Fuel Cell Electric Buses Considering the Noise DisturbanceabstractPlug-in fuel cell electric bus (PFCEB) is taken as one of the most promising technical ways for the sustainable and healthy development of the vehicles. However, the energy saving potential has not been exhaustively tapped, because of the limited integrated design (component matching design together with the energy management) level. This paper proposes a robustly integrated design method to address the issues. Firstly, a Taguchi robust design (TRD) method considering the noise of components, stochastic vehicle mass and historical driving cycles is proposed to find the robust design parameters that are suitable for different component matching and can resist the noise disturbance. Then, a component matching design zone reduced work is finished through analyzing the mechanism of action between components and the evaluation index, based on the TRD theory and the Sensitivity analysis method (SAM). Finally, a robust component matching design method is investigated by considering the noises, based on the Design for Six Sigma (DFSS) theory. The simulation results demonstrate that the TRD and DFSS methods are reasonable and applicable in the integrated design of PFCEB, where the robust component matching can be found; the energy economy of the proposed adaptive Pontryagin’s Minimum Principle (A-PMP) based energy management strategy is better than the rule-based control strategy. Daizheng Hou, Yafu Zhou, Jing Lian 0002 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | The Understanding of Traffic Police Intention Based on Visual Awareness
Jing Lian 0002, Yafu Zhou |
Neural Process. Lett. | 1 |
| 2022 | Multi-PPTP: Multiple Probabilistic Pedestrian Trajectory Prediction in the Complex Junction SceneabstractPedestrian trajectory prediction in dynamic, multi-agent traffic junction scene is an important problem in the context of self-driving cars. Accurately predicting the trajectory of the agent, especially the pedestrian with high randomness, is of great significance to autonomous driving technology. In this paper, we propose Multi-PPTP, a novel trajectory prediction model that first utilizes a composite rasterized map to model the complex representations and interactions of road components, including dynamic obstacles (e.g., pedestrians, vehicles, and cyclists) and static road information (e.g., lanes, traffic lights, and sidewalks). Then it uses MapNet and AgentNet to extract spatio-temporal features by deep convolutional networks and LSTM to automatically derive relevant features, and next exploits Interaction-AttNet aggregate features to learn the high interactions among all components by affine transformation and multi-head attention mechanism. Additionally, we also propose a series of unique loss functions to predict multiple possible trajectories of each pedestrian while estimating their probabilities. Following extensive offline evaluation and comparison to the state-of-the-art models, our approach outperforms significantly the state-of-the-art models on our large scale in-house and public nuScenes dataset. Jing Lian 0002, Yafu Zhou |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Causal Temporal-Spatial Pedestrian Trajectory Prediction With Goal Point Estimation and Contextual InteractionabstractForecasting pedestrian trajectories in complex dynamic environments is highly critical for the application of autonomous vehicles and robots. Accordingly, this paper proposes a novel pedestrian trajectory prediction model called CTSGI, which, utilizes self-attention mechanism to construct an interactive graph between pedestrians and their neighbors based on their spatial relationship, to model the crowd’s interaction. At the same time, it uses self-attention to extract the temporal dependence for a single pedestrian. In order to effectively model the interaction between the pedestrian and the context, such as where a pedestrian can walk or approach, the semantic segmentation of background image is utilized. CTSGI estimates the goal points of pedestrian and their neighbors to assist in predicting the future trajectory. In addition, the causal structure model is used to analyze the confounding factors existing in the encoding stage, and Do-calculus is introduced for eliminating the confounding impact to improve the prediction performance. Moreover, extensive experiments are conducted for the proposed model on ETH and UCY datasets, and clearly, the experimental results reveal that the model reported herein outperforms the comparative state-of-the-art methods by 27.59% in Average Displacement Error (ADE) and 8.33% in Final Displacement Error (FDE). Furthermore, visualization of attention indicates that our model can capture the interaction between some specific pedestrian and their neighbors better. Jing Lian 0002, Fengning Yu, Yafu Zhou |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2018 | Dense 3D Semantic SLAM of traffic environment based on stereo visionabstractTo solve the intelligent vehicles’ problems of ‘where am I?’ and ‘what is around me?’, a dense 3D sematic Simultaneous Localization and Mapping (SLAM) system is proposed to evaluate the pose of the intelligent vehicles and build the dense 3D semantic map. We address these challenges by combining a state of art Stereo-ORB-SLAM system and Convolutional Neural Networks. Firstly, we build a dense 3D point cloud map by using a four thread Stereo-ORB-SLAM system. Subsequently, a fully convolutional neural network architecture which uses RGB-D image as input is used to obtain pixel-wise segmentation. Finally, we fuse the geometric information and semantic information to get the semantic map. We test our method on the KITTI dataset and our dataset made with the Fpgalena stereo camera. Results indicate the system was effective in the real-time building of a semantic map, the speed of the entire system is about 10Hz, and the loop closing function can eliminate most of the drifting errors. Ümit Özgüner, Jing Lian 0002, Yafu Zhou, Yibing Zhao |
Intelligent Vehicles Symposium | 4 |
| 2018 | Real-time Traffic Scene Segmentation Based on Multi-Feature Map and Deep LearningabstractVisual-based semantic segmentation for traffic scene plays an important role in intelligent vehicles. In this paper, we present a new real-time deep fully convolution neural network (FCNN) for pixel-wise segmentation with six channel inputs. The six channel inputs include the RGB three channel color image, the Disparity (D) image generated by stereo vision sensor, the image to describe the Height (H) of each pixel above road ground, and the image to describe the Angle (A) between each pixel normal direction and the predicted direction of gravity, which are defined as a RGB-DHA multi-feature map. The FCNN is simplified and modified based on AlexNet to meet the real-time requirements of intelligent vehicle for environmental perception. The proposed algorithm is tested and compared in Cityscapes dataset, yields global accuracies 73.4% and 22ms for $400 \times 200$ resolution image with one Titan X GPU. Wei-Na Zheng, Lingchao Kong, Ümit Özgüner, Wenbin Hou, Jing Lian 0002 |
Intelligent Vehicles Symposium | 6 |
| 2018 | Traffic Scene Segmentation Based on RGB-D Image and Deep LearningabstractSemantic segmentation of traffic scenes has potential applications in intelligent transportation systems. Deep learning techniques can improve segmentation accuracy, especially when the information from depth maps is introduced. However, little research has been done on the application of depth maps to the segmentation of traffic scene. In this paper, we propose a method for semantic segmentation of traffic scenes based on RGB-D images and deep learning. The semi-global stereo matching algorithm and the fast global image smoothing method are employed to obtain a smooth disparity map. We present a new deep fully convolutional neural network architecture for semantic pixel-wise segmentation. We test the performance of the proposed network architecture using RGB-D images as input and compare the results with the method that only takes RGB images as input. The experimental results show that the introduction of the disparity map can help to improve the semantic segmentation accuracy and that our proposed network architecture achieves good real-time performance and competitive segmentation accuracy. Jing Lian 0002, Wei-Na Zheng, Yafu Zhou |
IEEE Trans. Intell. Transp. Syst. | 3 |