VLDB 2026 Research / reviewers in the wild / expert
Jun Hu 0020
dblp:28/441-20
· DBLP profile ↗
19ranked-venue papers
1as first author
18since 2021 · last 2026
0000-0002-7094-1901ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Distributed NN adaptive optimal resilient control for nonlinear vehicle platoon system under DoS attacks
Guowei Dong, Shuai Cheng 0001, Jun Hu 0020, Kewen Li 0001, Yongming Li 0002 |
Neurocomputing | 4 |
| 2026 | TP-LReID: Lifelong person re-identification using text prompts
Zhaoshuo Liu, Chaolu Feng, Wei Li 0117, Kun Yu 0002, Jun Hu 0020, Jinzhu Yang |
Pattern Recognit. | 6 |
| 2026 | Hierarchical Recursive Interaction and Multi-Stage Goal-Guided Mechanism for Multimodal Trajectory PredictionabstractIn highly dynamic and complex autonomous driving environments, accurately predicting agents’ future multimodal trajectories still faces challenges such as modeling diverse social interactions, capturing dynamic intents, and ensuring prediction consistency. To address these issues, this paper proposes a novel trajectory prediction model that integrates a Hierarchical Recursive Interaction Network (HRINet) and a multi-stage goal-guided mechanism (GoalNet), aiming to improve prediction accuracy, stability, and plausibility. Specifically, we design a HRINet with local and global attention mechanisms to recursively model various social interactions, while progressively integrating map semantic information to enhance the model’s understanding of traffic scenes. Meanwhile, inspired by the divide-and-conquer approach, the proposed GoalNet first estimates fine-grained multi-stage goal lane segments along the path. These goals are then used to continuously guide and constrain the trajectory generation process, effectively reducing error accumulation and improving stability. In addition, we construct a dynamic goal candidate area that combines domain knowledge and traffic rules to filter out unreasonable goals, thereby enhancing the plausibility and consistency of the predictions. Experimental results on nuScenes, INTERACTION, and Waymo Open Motion Dataset (WOMD) show that our model achieves state-of-the-art performance in multiple key metrics, maintains a trade-off between prediction accuracy, model complexity, and inference latency, and shows high stability and consistency in predictions. Jing Lian 0002, Zhenfeng Wang, Jian Zhao 0029, Jun Hu 0020 |
IEEE Trans. Intell. Transp. Syst. | 8 |
| 2026 | Dynamic Query Management and Internal Consistency Representation Based Transformer for Online Vectorized HD Map ConstructionabstractThe online vectorized map construction technique employs a neural network model to forecast the vectorized representation of a specific region around automobiles, using data obtained from sensors mounted on automobiles. Due to advances in end-to-end object detection with transformers framework, the research on query-based online mapping has attracted substantial attention. However, the fixed number of queries and the random initialization of query embeddings constrain the model's performance. Moreover, the transformer architecture for object detection is based on the assumption that queries are identically distributed and independent, a premise that is not entirely applicable to map point queries which possess established subordinate relationships with map element instances. To address these issues, we initially incorporate a simplified transformer layer that utilizes semantic priors in bird's-eye view features for query initialization. The queries are then sent to a transformer-based map decoder for optimization and combined with a dynamic query management mechanism to eliminate low-confidence queries, hence maintaining computational efficiency. Furthermore, to guarantee that point queries within each instance preserve a consistent representation and avoid feature confusion among map element instances, we proposed an instance internal consistency map decoder. We conduct extensive experiments on commonly used map construction datasets to evaluate the proposed method. The experimental results demonstrate that our proposed method achieves state-of-the-art performance on the nuScenes and Argoverse 2 datasets. Wenjing Bai, Yunzhou Zhang, Wei Liu 0022, Shangwei Du, Jun Hu 0020, Shuai Cheng 0001, Zuotao Ning |
IEEE Trans. Multim. | 6 |
| 2025 | PriorsFusionMap: A Unified Framework for Robust Online Vectorized Map Construction with Temporal Aggregation and Historical Global Map Interaction for Autonomous Driving
Wenjing Bai, Zuotao Ning, Jun Hu 0020, Shuai Cheng 0001, Wei Liu 0022 |
PRCV (11) | 4 |
| 2025 | LLFormer4D: LiDAR-based lane detection method by temporal feature fusion and sparse transformerabstractAbstract Lane detection is a fundamental problem in autonomous driving, which provides vehicles with essential road information. Despite the attention from scholars and engineers, lane detection based on LiDAR meets challenges such as unsatisfactory detection accuracy and significant computation overhead. In this paper, the authors propose LLFormer4D to overcome these technical challenges by leveraging the strengths of both Convolutional Neural Network and Transformer networks. Specifically, the Temporal Feature Fusion module is introduced to enhance accuracy and robustness by integrating features from multi‐frame point clouds. In addition, a sparse Transformer decoder based on Lane Key‐point Query is designed, which introduces key‐point supervision for each lane line to streamline the post‐processing. The authors conduct experiments and evaluate the proposed method on the K‐Lane and nuScenes map datasets respectively. The results demonstrate the effectiveness of the presented method, achieving second place with an F1 score of 82.39 and a processing speed of 16.03 Frames Per Seconds on the K‐Lane dataset. Furthermore, this algorithm attains the best mAP of 70.66 for lane detection on the nuScenes map dataset. Jun Hu 0020, Chaolu Feng, Haoxiang Jie, Zuotao Ning, Xinyi Zuo |
IET Comput. Vis. | 1 |
| 2024 | DFT-Net: A Bimodal Object Detection Algorithm for Complex Traffic EnvironmentsabstractIn response to the challenge of single-modal sensors struggling to rapidly and accurately detect in complex traffic environments, a DFT-Net detection algorithm based on the fusion of visible light and infrared bimodal features is proposed. Firstly, the algorithm constructs a dual modal feature extraction network and designs an efficient aggregation module with stronger feature extraction capability. Secondly, the algorithm utilizes a framework of self-attention and cross-attention mechanisms to fuse the features of visible light and infrared modalities. This framework facilitates information interaction between different modalities through query-guided mechanisms. To reduce the computational load of the dual modal feature transformer module, a spatial feature compression module is introduced. Finally, in order to mitigate visible noise and low infrared resolution, this paper combines Transformer and CBAM attention mechanism. The designed algorithm is compared with various detection algorithms on multiple datasets, and the results demonstrate its effectiveness. On the FLIR-gan dataset, compared to the YOLOv7 detection algorithms for visible light and infrared, our DFT-Net achieved an accuracy improvement of 26.8% and 26.7% respectively, and 2.8% improvement compared to the bimodal detection algorithm CFT. This indicates that the algorithm exhibits good detection performance in complex traffic environments. Jing Lian 0002, Jun Hu 0020 |
INDIN | 4 |
| 2024 | Attention-based adaptive structured continuous sparse network pruning
Wei Liu 0022, Yongming Li 0002, Jun Hu 0020, Shuai Cheng 0001, Wenxing Yang |
Neurocomputing | 4 |
| 2024 | SADGFeat: Learning local features with layer spatial attention and domain generalization
Wenjing Bai, Yunzhou Zhang, Li Wang 0160, Wei Liu 0022, Jun Hu 0020 |
Image Vis. Comput. | 5 |
| 2024 | GCReID: Generalized continual person re-identification via meta learning and knowledge accumulation
Zhaoshuo Liu, Chaolu Feng, Kun Yu 0002, Jun Hu 0020, Jinzhu Yang |
Neural Networks | 4 |
| 2024 | Efficient Vehicle Trajectory Prediction With Goal Lane Segments and Dual-Stream Cross AttentionabstractReal-time and efficient prediction of plausible trajectories of surrounding traffic agents is essential for autonomous driving. In reality, the motion of agents depends not only on their goal intents, but is also constrained by road topology. Especially for vehicles, lane geometry can significantly influence their future trajectories. In this paper, the constraining and guiding roles of lane networks are investigated, and a novel vehicle trajectory prediction model based on the goal lane segment is proposed. Specifically, a Dual-Stream Cross Attention Module (DSCAM) is developed that incorporates goal lane segment prediction into the process of collecting various interaction information. This achieves scene-consistent predictions while reducing resource consumption and inference latency. Then, learnable refinement tokens are used to adaptively refine coarse-grained trajectories during the feature decoding process, with the goal of improving prediction accuracy and ensuring temporal consistency. Extensive experiments with the Argoverse-1 and nuScenes datasets demonstrate that our model performs competitively and well-balanced. Notably, on the nuScenes, our model with only 0.44 million (M) parameters has an inference latency of less than six milliseconds (ms), making it ideal for large-scale production autonomous driving systems that require high computational resources, operational efficiency, and prediction performance. Jing Lian 0002, Jian Zhao 0029, Jun Hu 0020 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | A Training-Free, Lightweight Global Image Descriptor for Long-Term Visual Place Recognition Toward Autonomous VehiclesabstractLong-term visual place recognition (VPR) has recently become a popular research topic in the field of autonomous driving. In urban scenarios, variations in scene appearance due to the change in seasons and illumination bring great challenges for scene description. Several learning-based VPR techniques can learn latent invariant descriptors for appearance variations and show excellent performance in long-term VPR tasks. However, these methods require huge datasets and computational resources (e.g., GPUs) for training and inference. Mobile platforms such as autonomous vehicles often cannot provide sufficient computing power. To address this issue, in this paper, a training-free lightweight global image descriptor named SSR-VLAD is proposed for VPR. This descriptor is able to work accurately in real-time without GPUs, even on embedded platforms. The contribution of this work has two aspects. (1) A novel semantic skeleton representation (SSR) is proposed to describe the semantic spatial distribution of scenes by using the semantic spatial context; (2) Inspired by the Vector of Locally Aggregated Descriptors (VLAD), a spatial-temporal aggregation framework for SSR features is constructed to aggregate all SSR features into one SSR-VLAD descriptor, which encodes the spatial and temporal information into a fixed-size global descriptor. SSR-VLAD shows robust performance towards the appearance variations of scenes. Specifically, on three public datasets with challenging urban scenes, experimental results show that SSR-VLAD has competitive VPR performance compared to several state-of-the-art (SoTA) VPR methods. Additionally, SSR-VLAD achieves SoTA real-time computational performance with lower RAM consumption in computationally constrained scenarios. Jiwei Nie, Joe-Mei Feng, Dingyu Xue, Wei Liu 0022, Jun Hu 0020, Shuai Cheng 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2023 | LCDeT: LiDAR Curb Detection Network with TransformerabstractCurb detection can be used to determine road boundary information, which plays a crucial role in intelligent driving. In this paper, we propose an efficient 3D curb detection network combined with the Transformer (LCDeT), which realizes efficient and stable curb extraction from the mobile laser scanning data end-to-end. Different from the most existing algorithms that project the 3D point cloud to the 2D image, like height map or density map images before processing, we directly extract the point cloud features from the 3D point cloud to avoid the loss of spatial information of the 3D point cloud. Furthermore, we introduce SpatioTemporal Window(STWin) attention operations in the Transformer module to extract continuous, smooth curb features. In the temporal dimension, we perform a cross-attention operation on the point cloud of the historical frame and the point cloud of the current frame to improve the stability and continuity of the road edge detection results between the multi-frame point clouds. In the spatial dimension, we introduce the hybrid-attention operation on the point cloud of the current frame to extract spatially related features from the axial and local positions, respectively, to improve the detection accuracy. At last, to verify the performance, we firstly test it based on the only public curb dataset of 32-line LiDAR. The proposed LCDet achieves the state-of-the-art performance, an F1 score of 97.59%, with LCDet. At the same time, it is considered that the industry lacks roadside datasets for high-resolution LiDARs. We have collected, organized and published the industry's first 128-line laser roadside dataset, NRS-Dataset, which contains 6200 frames of point clouds in Urban during daytime and nighttime. And the NRS-Dataset will be available at [github11https://github.com/STWin1/curb]. Furthermore, we test the algorithm based on this dataset. The experimental results prove that our approach achieves high accuracy and recall in complex scenarios, which further shows the higher effective and robust performance of the LCDet algorithm than previous studies. Jian Gao 0016, Haoxiang Jie, Bingqing Xu, Lifeng Liu, Jun Hu 0020, Wei Liu 0022 |
IJCNN | 5 |
| 2023 | Adaptive Channel Pruning for Trainability Protection
Dazong Zhang, Wei Liu 0022, Yongming Li 0002, Jun Hu 0020, Shuai Cheng 0001, Wenxing Yang |
PRCV (10) | 5 |
| 2023 | ITCNN: Incremental Learning Network Based on ITDA and Tree Hierarchical CNN
Pengyu Wang 0008, Tao Ren 0002, Wei Liu 0022, Jun Hu 0020, Shuai Cheng 0001, Dazong Zhang |
PRCV (8) | 5 |
| 2023 | Knowledge-Preserving continual person re-identification using Graph Attention Network
Zhaoshuo Liu, Chaolu Feng, Shuaizheng Chen, Jun Hu 0020 |
Neural Networks | 4 |
| 2023 | Adaptive Fuzzy Predefined-Time Control for Third-Order Heterogeneous Vehicular Platoon Systems With Dead ZoneabstractThis article investigates the problem of fuzzy adaptive predefined-time terminal sliding mode (TSM) control for a third-order heterogeneous vehicular platoon system with an unknown dead zone. For the purpose of approximating unknown nonlinear functions, fuzzy logic systems (FLSs) are utilized. In addition, the impact of the dead zone on the performance of the control may be lessened by building the dead-zone compensation. A tracking error based on the modified constant time headway policy is built to get rid of the assumption of zero initial spacing and decrease the distance between vehicles at the same time. A unique nonsingular TSM control system is then built using the predefined-time stability criterion; with the help of Lyapunov functions, it is possible to demonstrate both the individual and the string stability of the whole heterogeneous vehicle platoon in a predefined time. Finally, a series of simulations are shown to demonstrate the validity of the proposed results. Yongming Li 0002, Yongyan Zhao, Wei Liu 0022, Jun Hu 0020 |
IEEE Trans. Ind. Informatics | 4 |
| 2022 | A Novel Image Descriptor with Aggregated Semantic Skeleton Representation for Long-term Visual Place RecognitionabstractIn a Simultaneous Localization and Mapping (SLAM) system, loop-closure can eliminate accumulated errors, which is accomplished by Visual Place Recognition (VPR), a task that retrieves current scene from a set of pre-stored sequential images through matching specific scene-descriptors. In urban scenes, the appearance variation caused by seasons and illumination have brought great challenges to the robustness of scene descriptors. Semantic segmentation images can not only deliver the shape information of objects, but also their categories and spatial relations that will not be affected by the appearance-variation of the scene. Innovated by the Vector of Locally Aggregated Descriptor (VLAD), in this paper, we propose a novel image descriptor with aggregated semantic skeleton representation (SSR), dubbed SSR-VLAD, for the VPR under drastic appearance-variation of environments. The SSR-VLAD of one image aggregates the semantic skeleton features of each category, and encodes the spatial-temporal distribution information of the image semantic information. We conduct a series of experiments on three public datasets of challenging urban scenes. Compared with three state-of-the-art VPR methods- CoHog, NetVLAD, and Region-VLAD, VPR by matching SSR-VLAD outperforms those methods and maintains competitive real-time performance at the same time. Jiwei Nie, Joe-Mei Feng, Dingyu Xue, Wei Liu 0022, Jun Hu 0020, Shuai Cheng 0001 |
ICPR | 6 |
| 2020 | BCEFCM_S: Bias correction embedded fuzzy c-means with spatial constraint to segment multiple spectral images with intensity inhomogeneities and noises
Chaolu Feng, Wei Li 0117, Jun Hu 0020, Kun Yu 0002, Dazhe Zhao |
Signal Process. | 3 |