VLDB 2026 Research / reviewers in the wild / expert
Bangquan Xie
dblp:326/3897
· DBLP profile ↗
13ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0002-4500-8941ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Task-Aware Retrieval Augmentation for Dynamic RecommendationabstractDynamic recommendation systems aim to provide personalized suggestions by modeling temporal user-item interactions across time-series behavioral data. Recent studies have leveraged pre-trained dynamic graph neural networks (GNNs) to learn user-item representations over temporal snapshot graphs. However, fine-tuning GNNs on these graphs often results in generalization issues due to temporal discrepancies between pre-training and fine-tuning stages, limiting the model’s ability to capture evolving user preferences. To address this, we propose TarDGR, a task-aware retrieval-augmented framework designed to enhance generalization capability by incorporating task-aware model and retrieval-augmentation. Specifically, TarDGR introduces a Task-Aware Evaluation Mechanism to identify semantically relevant historical subgraphs, enabling the construction of task-specific datasets without manual labeling. It also presents a Graph Transformer-based Task-Aware Model that integrates semantic and structural encodings to assess subgraph relevance. During inference, TarDGR retrieves and fuses task-aware subgraphs with the query subgraph, enriching its representation and mitigating temporal generalization issues. Experiments on multiple large-scale dynamic graph datasets demonstrate that TarDGR consistently outperforms state-of-the-art methods, with extensive empirical evidence underscoring its superior accuracy and generalization capabilities. Xinke Jiang, Qingshuai Feng, Lun Du, Yuchen Fang 0001, Hao Miao 0001, Bangquan Xie, Qingqiang Sun |
AAAI | 8 |
| 2026 | Reinforcement Learning Neural Network Observer-Based Adaptive Optimal Time-Varying Formation Control for Uncertain Multi-Agent Systems With External DisturbanceabstractThe traditional formation control methods of multi-agent systems (MASs) often rely on restrictive linear matrix inequalities and strong assumptions. To overcome this limitation, this work proposes an adaptive optimal time-varying formation control protocol for uncertain MASs, integrating reinforcement learning(RL)-based neural network observer. First, a radial basis function neural network(RBFNN)with adaptive law is designed to approximate the unknown nonlinear dynamics. An RBFNN-based state estimator and a novel disturbance estimator are developed to reconstruct the unmeasurable system state and the disturbance generated by an exosystem, respectively. Second, using the state estimate, a distributed formation control error system is established, and a performance index based on a Hamilton–Jacobi–Bellman(HJB)equation is introduced to achieve the objective of optimal formation control. Third, leveraging the state and disturbance estimates, an actor-critic RL-based optimal time-varying formation control protocol with adaptive update laws is proposed. Actor and critic networks are constructed to approximate unknown nonlinear terms in HJB equation. It is proved that estimator error systems and RBFNN approximation errors are semi-globally uniformly ultimately bounded(SGUUB). Simulation examples are provided to verify the effectiveness of the proposed estimators and optimal formation control protocol. Moreover, the method is successfully applied to unmanned aerial vehicles (UAVs) formation control in the Isaac Sim simulation environment, demonstrating the method’s practical feasibility. Yanzhou Li, Shenghuang He, Bangquan Xie, Wenjian Zhong, Yongkang Lu |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2026 | Enhanced Robustness in Simultaneous Fault Estimation and Distributed Fault-Tolerant Consensus Tracking Control for Multiagent Systems Utilizing Extended Observer ApproachabstractIn the context of consensus tracking control for multiagent systems (MASs), the system performance may be significantly degraded by multiple potential factors, including actuator/sensor faults and external disturbances. However, the development of a distributed fault estimation (FE) mechanism integrated with a fault-tolerant consensus control protocol remains a critical research challenge. This work proposes a novel distributed extended simultaneous observer (DESO) and observer-based fault-tolerant consensus tracking control (OB-FTCT) protocol for nonlinear MASs. First, system state as well as actuator and sensor faults for each agent are integrated into a new augmented vector. The original system is transformed into a descriptor system, and a DESO, tailored to the augmented vector, is designed to simultaneously estimate the system state and faults. Second, a novel OB-FTCT protocol is designed to achieve theH∞consensus tracking and compensate the impact of the two faults. Through the establishment of Lyapunov functions and an algorithm, the stability of consensus tracking error system is proved and ensured. Third, the effectiveness of the proposed method is verified by simulation examples. Also, we apply the proposed method to single-link flexible joint robot, and compare it with existing method to prove the effectiveness. Yanzhou Li, Shenghuang He, Bangquan Xie, Wenjian Zhong |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2025 | CalibMutiL: Online Calibration Of LiDAR-Camera Based On Multi-level Visual Feature FusionabstractMulti-sensor fusion is a key technology in the field of autonomous driving and robotics. Traditional offline multi-sensor fusion calibration methods rely on manual operations and fail to meet real-time requirements, while recent online calibration technologies have limited generalization capabilities. This paper proposes CalibMutiL, an end-to-end calibration network that departs from conventional deep feature fusion by leveraging multi-level RGB image features to guide point cloud alignment. CalibMutiL introduces a Multi-level Fusion module (MLF) that effectively utilizes the rich visual features of the image. In addition, we regard the alignment process as a sequence prediction problem and further improve the performance through an Iterative Refinement module (IRM). Evaluation of the KITTI odometry and raw dataset demonstrates the average calibration error reaches 0.81cm and 0.09°. The generalization tests resulted in errors of 4.24cm and 0.13°, outperforming existing methods. Our implementation will be publicly available at https://github.com/VIP-G/CalibMutiL. Eksan Firkat, Eliyas Suleyman, Bangquan Xie, Fengze Li, Askar Hamdulla |
IROS | 4 |
| 2024 | SA-BiGCN: Bi-Stream Graph Convolution Networks With Spatial Attentions for the Eye Contact Detection in the WildabstractEye contact is essential in transmitting information and intention in the wild environment (e.g., urban streets or parking lots) with mixed vehicles and pedestrians. Compared with the vision image data, the human skeleton data are deemed to be robust to unconstrained surroundings and illumination. However, the skeleton graph-based approaches are mainly used for the action recognition. It is challenging to directly apply them to the eye detection task, which is momentary and dynamic given the complex wild environment. This paper proposes a Bi-stream Spatial Attention Graph Convolution Network (SA-BiGCN) for eye contact detection in the wild. We design a directed, nose-centric skeleton graph to capture relevant and hierarchical information and their interactions. We also propose a Bi-stream graph convolution network model with spatial attention to dynamically extract and fuse skeleton joints and bones information. The model was validated by comparing with state-of-art models on three large-scale public datasets, including JAAD, PIE, and LOOK. The results highlight the accuracy and generalization performance of the proposed SA-BiGCN model in detecting the eye contact in the wild environment. The ablation analysis validates the importance of the skeleton graph design, the spatial attention mechanism in the feature fusion process, as well as the model robustness against noisy skeleton data in terms of part occlusions, block occlusions, random occlusions, and random deviations. Yancheng Ling, Zhenliang Ma, Bangquan Xie, Qi Zhang 0086, Xiaoxiong Weng |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | PedAST-GCN: Fast Pedestrian Crossing Intention Prediction Using Spatial-Temporal Attention Graph Convolution NetworksabstractAccurately and timely predicting pedestrian crossing intentions in real-time is critical for operating intelligent vehicles on roads. Although existing models achieve promising accuracy using complex models and video image data, they are constrained for real-time practical use given the high model complexity, time-consuming data preprocessing, and low-quality image data in the wild. To address these, the paper proposes a Spatial-Temporal Attention Graph Convolution Network model for fast pedestrian crossing intention prediction (PedAST-GCN). It uses a lightweight GCN model as the backbone network with simple but robust graph representations of pedestrian crossing intention modality features, including pedestrian pose, bounding box, and vehicle speeds. The model is validated by comparing it with state-of-the-art models on two large-scale public datasets (JAAD and PIE). The results highlight the better performance of the PedAST-GCN model for pedestrian crossing intention prediction in terms of accuracy and computation times. The ablation analysis confirms the value of the backbone layer and graph design, the designed modality features, the effectiveness of attention mechanisms in capturing long-term dependencies (spatial-temporal attention) and fusing heterogeneous features (modality attention), and the robust performance across various observation lengths and in the presence of noisy data. Yancheng Ling, Zhenliang Ma, Qi Zhang 0086, Bangquan Xie, Xiaoxiong Weng |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | SPD: Semi-Supervised Learning and Progressive Distillation for 3-D DetectionabstractCurrent learning-based 3-D object detection accuracy is heavily impacted by the annotation quality. It is still a challenge to expect an overall high detection accuracy for all classes under different scenarios given the dataset sparsity. To mitigate this challenge, this article proposes a novel method called semi-supervised learning and progressive distillation (SPD), which uses semi-supervised learning (SSL) and knowledge distillation to improve label efficiency. The SPD uses two big backbones to hand the unlabeled/labeled input data augmented by the periodic IO augmentation (PA). Then the backbones are compressed using progressive distillation (PD). Precisely, PA periodically shifts the data augmentation operations between the input and output of the big backbone, aiming to improve the network's generalization of the unseen and unlabeled data. Using the big backbone can benefit from large-scale augmented data better than the small one. And two backbones are trained by the data scale and ratio-sensitive loss (data-loss). It solves the over-flat caused by the large-scale unlabeled data from PA and helps the big backbone prevent overfitting on the limited-scale labeled data. Hence, using the PA and data loss during SSL training dramatically improves the label efficiency. Next, the trained big backbone set as the teacher CNN is progressively distilled to obtain a small student model, referenced as PD. PD mitigates the problem that student CNN performance degrades when the gap between the student and the teacher is oversized. Extensive experiments are conducted on the indoor datasets SUN RGB-D and ScanNetV2 and outdoor dataset KITTI. Using only 50% labeled data and a 27% smaller model size, SPD performs 0.32 higher than the fully supervised VoteNet [1] which is adopted as our backbone. Besides, using only 2% labeled data, compared to the other fully supervised backbone PV-RCNN [2], SPD accomplishes a similar accuracy (84.1 and 84.83) and 30% less inference time. Bangquan Xie, Zongming Yang, Ruifa Luo, Ailin Wei, Xiaoxiong Weng, Bing Li 0008 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | FourStr: When Multi-sensor Fusion Meets Semi-supervised LearningabstractThis research proposes a novel semi-supervised learning framework FourStr (Four-Stream formed by two two-stream models) that focuses on the improvement of fusion and labeling efficiency for 3D multi-sensor detector. FourStr adopts a multi-sensor single-stage detector named adaptive fusion network (AFNet) as the backbone and trains it through the semi-supervision learning (SSL) strategy Stereo Fusion. Note that multi-sensor AFNet and SSL Stereo Fusion can benefit each other. On the one hand, the Four-stream composed of two AFNets naturally provides rich inputs and large models for SSL Stereo Fusion. While other SSL works have to use massive augmentation to obtain rich inputs, and deepen and widen the network for large models. On the other hand, by the novel three fusion stages and Loss Pruning, Stereo Fusion improves the fusion and labeling efficiency for AFNet. Finally, extensive experiments demonstrate that FourStr performs excellently on outdoor dataset (KITTI and Waymo Open Dataset) and indoor dataset (SUN RGB-D), especially for the small contour objects. And compared to the fully-supervised methods, FourStr achieves similar accuracy with only 2% labeled data on KITTI (or with 50% labeled data on SUN RGB-D). Bangquan Xie, Zongming Yang, Ailin Wei, Xiaoxiong Weng, Bing Li 0008 |
ICRA | 1 |
| 2023 | ANAS: Asymptotic NAS for large-scale proxyless search and multi-task transfer learning
Bangquan Xie, Zongming Yang, Ruifa Luo, Ailin Wei, Xiaoxiong Weng, Bing Li 0008 |
Pattern Recognit. | 1 |
| 2023 | MuTrans: Multiple Transformers for Fusing Feature Pyramid on 2D and 3D Object DetectionabstractOne of the major components of the neural network, the feature pyramid plays a vital part in perception tasks, like object detection in autonomous driving. But it is a challenge to fuse multi-level and multi-sensor feature pyramids for object detection. This paper proposes a simple yet effective framework namedMuTrans(MultipleTransformers) to fuse feature pyramid in single-stream 2D detector or two-stream 3D detector. The MuTrans based on encoder-decoder focuses on the significant features via multiple Transformers. MuTrans encoder uses three innovative self-attention mechanisms:Spatial-wiseBoxAlign attention (SB) for low-level spatial locations,Context-wiseAffinity attention (CA) for high-level context information, and high-level attention for multi-level features. Then MuTrans decoder processes these significant proposals including the RoI and context affinity. Besides, theLow andHigh-levelFusion (LHF) in the encoder reduces the number of computational parameters. And the Pre-LN [1] is utilized to accelerate the training convergence. LHF and Pre-LN are proven to reduce self-attention’s computational complexity and slow training convergence. Our result demonstrates the higher detection accuracy of MuTrans than that of the baseline method, particularly in small object detection. MuTrans demonstrates a 2.1 higher detection accuracy onAPSindex in small object detection on MS-COCO 2017 with ResNeXt-101 backbone, a 2.18 higher 3D detection accuracy (moderate difficulty) for small object-pedestrian on KITTI, and 6.85 higher RC index (Town05 Long) on CARLA urban driving simulator platform. Bangquan Xie, Ailin Wei, Xiaoxiong Weng, Bing Li 0008 |
IEEE Trans. Image Process. | 1 |
| 2023 | AMMF: Attention-Based Multi-Phase Multi-Task Fusion for Small Contour Object 3D DetectionabstractRecently significant progress has been made in 3D detection. However, it is still challenging to detect small contour objects under complex scenes. This paper proposes a novel Attention-based Multi-phase Multi-task Fusion (AMMF) that uses point-level, RoI-level, and multi-task fusions to complement the disadvantages of LiDAR and camera, to solve this challenge. First, at the feature extraction phase, AMMF uses the Low and High-level Fusion with Matching Attention (LHF-MA) and efficient FPN (eFPN) to perform point-level fusion for cross sensors and single sensor, respectively. Instead of merging each level and using expensive 3D CNN like other methods, LHF-MA fuses low-level spatial location and high-level contextual feature of 2D CNN customized feature extractors and ignores the fusion of middle levels, reducing the computational cost. Then, at the proposal generation phase, Progressive Proposal Fusion (PPF) with learned attention map is used to perform coarse-to-fine RoI-level fusion, instead of only combining coarse-grained features at high-level of network. PPF using progressively increasing IoU thresholds could avoid overfitting and improve the performance. Note that the matching attentions and learned attention maps are utilized to weigh the priority of different sensors. Moreover, to solve the sparseness of point-wise fusion between LiDAR BEV and RGB image, AMMF uses multi-task fusion that generates pseudo-LiDAR from camera by depth estimation task, to guide this point-wise fusion. Finally, AMMF performs excellently for detecting small contour objects like pedestrians, cyclists, and distant cars. On the KITTI, AMMF finishes 3.62% improvements in the moderate instance for pedestrians. It achieves a 2.21% improvement in the$>$50 instance of LEVEL$_{-}$2 level for vehicle on the Waymo Open Dataset. And AMMF is further verified on our customized dataset consisting of challenging scenarios like strong illumination and heavy shadow cases. Bangquan Xie, Zongming Yang, Ailin Wei, Xiaoxiong Weng, Bing Li 0008 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | FocusTR: Focusing on Valuable Feature by Multiple Transformers for Fusing Feature Pyramid on Object DetectionabstractThe feature pyramid, which is a vital component of the convolutional neural networks, plays a significant role in several perception tasks, including object detection for autonomous driving. However, how to better fuse multi-level and multi-sensor feature pyramids is still a significant challenge, especially for object detection. This paper presents a FocusTR (Focusing on the valuable features by multiple Transformers), which is a simple yet effective architecture, to fuse feature pyramid for the single-stream 2D detector and two-stream 3D detector. Specifically, FocusTR encompasses several novel self-attention mechanisms, including the spatial-wise boxAlign attention (SB) for low-level spatial locations, context-wise affinity attention (CA) for high-level context information, and level-wise attention for the multi-level feature. To alleviate self-attention's computational complexity and slow training convergence, Fo-cusTR introduces a low and high-level fusion (LHF) to reduce the computational parameters, and the Pre- Ln [1]to accelerate the training convergence. Bangquan Xie, Zongming Yang, Ailin Wei, Xiaoxiong Weng, Bing Li 0008 |
IROS | 1 |
| 2022 | Multi-Scale Fusion With Matching Attention Model: A Novel Decoding Network Cooperated With NAS for Real-Time Semantic SegmentationabstractThis paper proposes a real-time multi-scale semantic segmentation network (MsNet). MsNet is a combination of our novel multi-scale fusion with matching attention model (MFMA) as the decoding network and the network searched by asymptotic neural architecture search (ANAS) or MobileNetV3 as the encoding network. The MFMA not only extracts low level spatial features from multi-scale inputs but also decodes the contextual features extracted by ANAS. Specifically, considering the advantages and disadvantages of the addition fusion and concatenation fusion, we design multi-scale fusion (MF) that balances speed and accuracy. Then we creatively design two matching attention mechanisms (MA), including matching attention with low calculation (MALC) mechanism and matching attention with strong global context modeling (MASG) mechanism, to match varying resolutions and information of features at different levels of a network. Besides, the ANAS performs the deep neural network search by employing an asymptotic method and provide an efficient encoding network for MsNet, releasing researchers from those tedious mechanical trials. Through extensive experiments, we prove that MFMA, which can be applied to numerous recognition tasks, possesses excellent decoding ability. And we demonstrate the effectiveness and necessity of implementing the “matching” attention mechanism. Finally, the proposed two versions, MsNet_ANAS and MsNet_M achieve a new state-of-the-art trade-off between accuracy and speed on the CamVid and Cityscapes datasets. More remarkably, on the Nvidia Tesla V100 GPU, our MsNet_ANAS achieves 74.1% mIoU with the speed of 184.2 FPS on the CamVid while 72.9% mIoU with the speed of 119.9 FPS on the Cityscapes. Bangquan Xie, Zongming Yang, Ruifa Luo, Ailin Wei, Xiaoxiong Weng, Bing Li 0008 |
IEEE Trans. Intell. Transp. Syst. | 1 |