VLDB 2026 Research / reviewers in the wild / expert
Jian Zhao 0029
dblp:70/2932-29
· DBLP profile ↗
17ranked-venue papers
0as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hierarchical Recursive Interaction and Multi-Stage Goal-Guided Mechanism for Multimodal Trajectory PredictionabstractIn highly dynamic and complex autonomous driving environments, accurately predicting agents’ future multimodal trajectories still faces challenges such as modeling diverse social interactions, capturing dynamic intents, and ensuring prediction consistency. To address these issues, this paper proposes a novel trajectory prediction model that integrates a Hierarchical Recursive Interaction Network (HRINet) and a multi-stage goal-guided mechanism (GoalNet), aiming to improve prediction accuracy, stability, and plausibility. Specifically, we design a HRINet with local and global attention mechanisms to recursively model various social interactions, while progressively integrating map semantic information to enhance the model’s understanding of traffic scenes. Meanwhile, inspired by the divide-and-conquer approach, the proposed GoalNet first estimates fine-grained multi-stage goal lane segments along the path. These goals are then used to continuously guide and constrain the trajectory generation process, effectively reducing error accumulation and improving stability. In addition, we construct a dynamic goal candidate area that combines domain knowledge and traffic rules to filter out unreasonable goals, thereby enhancing the plausibility and consistency of the predictions. Experimental results on nuScenes, INTERACTION, and Waymo Open Motion Dataset (WOMD) show that our model achieves state-of-the-art performance in multiple key metrics, maintains a trade-off between prediction accuracy, model complexity, and inference latency, and shows high stability and consistency in predictions. Jing Lian 0002, Zhenfeng Wang, Jian Zhao 0029, Jun Hu 0020 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | Entity and relationship extraction based on span contribution evaluation and focusing framework
Qibin Li, Nianmin Yao, Nai Zhou, Jian Zhao 0029 |
Comput. Speech Lang. | 4 |
| 2025 | Coarse-to-fine medical image registration with landmarks and deformable networks
Nianmin Yao, Linqi Meng, Jingyi Fang, Jian Zhao 0029 |
J. Supercomput. | 5 |
| 2024 | Cascade fusion of multi-modal and multi-source feature fusion by the attention for three-dimensional object detection
Fengning Yu, Jing Lian 0002, Jian Zhao 0029 |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | Alignment and fusion for adaptive domain nighttime semantic segmentation
Nianmin Yao, Jian Zhao 0029 |
Image Vis. Comput. | 3 |
| 2024 | CDGAN-BERT: Adversarial constraint and diversity discriminator for semi-supervised text classification
Nai Zhou, Nianmin Yao, Nannan Hu, Jian Zhao 0029 |
Knowl. Based Syst. | 4 |
| 2024 | A Novel Framework for Scene Graph Generation via Prior KnowledgeabstractThe scene graph generation aims to recognize objects and infer the relationships between them, which can provide a comprehensive understanding of image visual perception. However, the long-tailed issue of relations remains challenging for scene graph generation. This paper proposes a novel framework based on knowledge-driven data-driven joining to address the long-tail issues in scene graph generation. The proposed framework consists of two modules: the relation inference module and the prior knowledge learning module. The relation inference module aims to learn the relational features of entity pairs in images and the structural features of scene graphs. The prior knowledge learning module aims to learn the triplet representation from the knowledge graph and use it as prior knowledge to provide logical guidance and constraints for relation inference. This provides prior bias for relation inference to transfer the bias towards head categories to reasonable categories, thereby mitigating the long-tail problem. Experiment results indicate that the proposed framework outperforms on Visual Genome datasets and that the generated scene graph relation is logically reasonable. Jing Lian 0002, Jian Zhao 0029 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Efficient Vehicle Trajectory Prediction With Goal Lane Segments and Dual-Stream Cross AttentionabstractReal-time and efficient prediction of plausible trajectories of surrounding traffic agents is essential for autonomous driving. In reality, the motion of agents depends not only on their goal intents, but is also constrained by road topology. Especially for vehicles, lane geometry can significantly influence their future trajectories. In this paper, the constraining and guiding roles of lane networks are investigated, and a novel vehicle trajectory prediction model based on the goal lane segment is proposed. Specifically, a Dual-Stream Cross Attention Module (DSCAM) is developed that incorporates goal lane segment prediction into the process of collecting various interaction information. This achieves scene-consistent predictions while reducing resource consumption and inference latency. Then, learnable refinement tokens are used to adaptively refine coarse-grained trajectories during the feature decoding process, with the goal of improving prediction accuracy and ensuring temporal consistency. Extensive experiments with the Argoverse-1 and nuScenes datasets demonstrate that our model performs competitively and well-balanced. Notably, on the nuScenes, our model with only 0.44 million (M) parameters has an inference latency of less than six milliseconds (ms), making it ideal for large-scale production autonomous driving systems that require high computational resources, operational efficiency, and prediction performance. Jing Lian 0002, Jian Zhao 0029, Jun Hu 0020 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Mobile Phone Use Driver Distraction Detection Based on MSaE of Multi-Modality Physiological SignalsabstractDriver distraction, a major cause of traffic crashes, is reported to reduce driving performance and be detected with vehicle behavioral features. It also induces physiological responses. Time and frequency-domain features of physiological signals have been used to study distraction, but they are susceptible to residual noise and tend to overlook complexity. Moreover, the resampling problem arises while analyzing physiological signals at multiple time scales. This paper proposes a novel framework based on multiscale entropy on absolute time scales (MSaE) and bidirectional long short-term memory (BiLSTM) network to mine the distraction information in multi-modality physiological signals and detect distraction automatically. Firstly, an entropy-based resampling method is adopted to find the suitable downsampling rates of electroencephalography (EEG), electrocardiogram (ECG), and electromyography (EMG). Then, calculating entropy with absolute time scales instead of relative time scales in a sliding window is utilized to explore the fluctuations of each signal while distraction. Afterward, ReliefF is selected from conventional feature selectors to identify the optimal feature set for each signal. Finally, BiLSTM with time dependency is designed to detect driver distraction with the selected feature set. The results illustrate significant distinctions in the MSaE of multiple physiological signals between normal and distracted driving. Additionally, MSaE, superior to traditional features, is selected as the most discriminative feature for each signal in distraction mining. Furthermore, the accuracy is further improved by about 8%, incorporating multi-modality features rather than vehicle behavioral features. This study indicates the potential of employing various signals to understand and detect driver distraction effectively. Chi Zhang 0002, Fengyu Cong, Jian Zhao 0029, Timo Hämäläinen 0002 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | Multi-MCCR: Multiple models regularization for semi-supervised text classification with few labels
Nai Zhou, Nianmin Yao, Qibin Li, Jian Zhao 0029 |
Knowl. Based Syst. | 4 |
| 2023 | Optimal Trajectory Planning Method for the Navigation of WIP Vehicles in Unknown Environments: Theory and ExperimentabstractNavigation of underactuated wheeled inverted pendulum (WIP) vehicles in unknown environments is still facing great difficulties, especially when the optimal motion is required. This article proposes an optimal trajectory planning method for the navigation of WIP vehicles in unknown environments, where various performance demands, such as security, smoothness, efficiency, etc., are all considered. First, a map-building algorithm based on the improved Rao–Blackwellized particle filter is applied for the WIP vehicle to construct the environmental map. Then, a multiobjective optimization using the genetic algorithm is performed to find an optimized path between the given start and target point with path length, path curvature, and safe distance being taken into consideration simultaneously. Moreover, on the basis of kinematical and dynamical analysis, velocity, and acceleration constraints are parameterized with a path parameter, and the minimum-time trajectory along the optimized path is further planned with a sequence of maximum acceleration and deceleration trajectories. Finally, a WIP vehicle platform based on the robot operating system is designed, and related experiments in a real obstacle environment are conducted to validate the feasibility of the proposed method. Yigao Ning, Ming Yue 0001, Jinyong Shangguan, Jian Zhao 0029 |
IEEE Trans. Cybern. | 4 |
| 2023 | A Joint Entity and Relation Extraction Model based on Efficient Sampling and Explicit InteractionabstractJoint entity and relation extraction (RE) construct a framework for unifying entity recognition and relationship extraction, and the approach can exploit the dependencies between the two tasks to improve the performance of the task. However, the existing tasks still have the following two problems. First, when the model extracts entity information, the boundary is blurred. Secondly, there are mostly implicit interactions between modules, that is, the interactive information is hidden inside the model, and the implicit interactions are often insufficient in the degree of interaction and lack of interpretability. To this end, this study proposes a joint entity and relation extraction model (ESEI) based on E fficient S ampling and E xplicit I nteraction. We innovatively divide negative samples into sentences based on whether they overlap with positive samples, which improves the model’s ability to extract entity word boundary information by controlling the sampling ratio. In order to increase the explicit interaction ability between the models, we introduce a heterogeneous graph neural network (GNN) into the model, which will serve as a bridge linking the entity recognition module and the relation extraction module, and enhance the interaction between the modules through information transfer. Our method substantially improves the model’s discriminative power on entity extraction tasks and enhances the interaction between relation extraction tasks and entity extraction tasks. Experiments show that the method is effective, we validate our method on four datasets, and for joint entity and relation extraction, our model improves the F1 score on multiple datasets. Qibin Li, Nianmin Yao, Nai Zhou, Jian Zhao 0029 |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2023 | Scene-Adaptive Real-Time Fast Dehazing and Detection in Driving EnvironmentabstractReal-time and effective dehazing is crucial to ensure safe and smooth operations of driving in foggy conditions. In this paper, a novel vision-based scene-adaptive real-time dehazing method for continuous video frames is developed, and the enhanced frames are fed into a detector in order to assess the value of image enhancement for detection tasks. The defogging stage is established by importing scene adaptation in order to improve the effectiveness of continuous frames defogging. Considering the distribution of atmospheric light intensity within the frame, the positions of the frame’s key points are determined in order to distinguish the scene differences between frames. Then, an algorithm is proposed for the inter-frame perception of atmospheric light intensity based on key points (“inverted triangle”), which employs a dynamic updating to adjust the estimated update frequency of atmospheric light intensity in time based on the scene change. Further, frames are inverted to estimate the transmission map based on the scattering model, which can quickly recover free of fog. Finally, the adaptive region of interest is set based on the obtained key point of frames, and the enhanced image is fed to the detector. The experimental results indicate that the proposed method can effectively handle sudden scene changes, as the defogging accuracy can reach 98.58%, the average defogging time of each frame is only a few milliseconds, and the detection accuracy can be improved by 4.11%. Nana Lyu, Jian Zhao 0029, Pengbo Liu 0007, Tianfa Su, Jianxi Wen |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Robust Fault-Tolerant Estimation of Sideslip and Roll Angles for Distributed Drive Electric Buses With Stochastic Passenger MassabstractFaults or data loss of onboard sensors are a critical concern for active safety control of distributed drive electric buses (DDEBs). In this article, a robust fault-tolerant estimation (RFTE) method is proposed to estimate sideslip and roll angles of DDEB in spite of stochastic disturbances and sensor faults. Considering the effects of sensor faults, stochastic passenger mass and modeling errors on DDEB active safety control, an augmented system consisting of DDEB states and associated sensor faults is firstly constructed, where the disturbances caused by passenger mass and modeling errors are merged as unknown input vector. Then, by decoupling partial disturbances, an unknown input observer is developed to estimate the augmented system states, while a linear matrix inequality and an estimated states correction method are introduced to further attenuate the effects of residual disturbances and thus guarantee the accuracy of the estimation. Finally, the validity of the proposed RFTE method is evaluated by a Trucksim-Simulink co-simulation platform and a hardware-in-the-loop platform. The experimental results demonstrated that the proposed RFTE method can achieve effective estimation in the presence of stochastic disturbances and sensor faults, exhibiting better robustness and reliability than existing baseline estimation method. Jinyong Shangguan, Ming Yue 0001, Jian Zhao 0029 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Self attention mechanism of bidirectional information enhancement
Qibin Li, Nianmin Yao, Jian Zhao 0029 |
Appl. Intell. | 3 |
| 2022 | Rule-based adversarial sample generation for text classification
Nai Zhou, Nianmin Yao, Jian Zhao 0029 |
Neural Comput. Appl. | 3 |
| 2022 | Driver Distraction Detection Using Bidirectional Long Short-Term Network Based on Multiscale Entropy of EEGabstractDriver distraction diverting drivers’ attention to unrelated tasks and decreasing the ability to control vehicles, has aroused widespread concern about driving safety. Previous studies have found that driving performance decreases after distraction and have used vehicle behavioral features to detect distraction. But how brain activity changes while distraction remains unknown. Electroencephalography (EEG), a reliable indicator of brain activities has been widely employed in many fields. However, challenges still exist in mining the distraction information of EEG in realistic driving scenarios with uncertain information. In this paper, we propose a novel framework based on Multi-scale entropy (MSE) in a sliding window and Bidirectional Long Short-term Memory Network (BiLSTM) to explore the distraction information of EEG to detect driver distraction based on multi-modality signals in real traffic. Firstly, MSE with sliding window is implemented to extract the EEG features to determine the distraction position. Statistical analysis of vehicle behavioral data is then performed to validate driving performance indeed changes around distraction position. Finally, we use BiLSTM to detect driver distraction with MSE and other traditional features. Our results show that MSE notably decreases after distraction. Consistent with the result of MSE, driving performance significantly deviates from the normal state after distraction. Besides, BiLSTM performance of MSE outperforms other entropy-based methods and is better than behavioral features. Additionally, the accuracy is improved again after adding MSE feature to behavioral features with a 3% increasement. The proposed framework is useful for mining brain activity information and driver distraction detection applications in realistic driving scenarios. Chi Zhang 0002, Fengyu Cong, Jian Zhao 0029, Timo Hämäläinen 0002 |
IEEE Trans. Intell. Transp. Syst. | 4 |