VLDB 2026 Research / reviewers in the wild / expert
Ming Gao 0012
dblp:71/4173-12
· DBLP profile ↗
9ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0003-4199-4470ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 5 since 2021Computer networks · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LVMSOD: Lightweight Visual Mamba Small Object Detection for Autonomous VehiclesabstractThe application of object detection in industrial transportation has witnessed substantial advancements, yielding significant enhancements in both safety and efficiency. While Transformer-based detectors have demonstrated remarkable success in object detection for autonomous driving, their quadratic computational complexity and limited small-object perception capabilities remain significant challenges. To address these limitations, we propose lightweight visual Mamba small object detection (LVMSOD) method, a novel selective state space model designed for small object detection task. The proposed framework employs a multi-level cascade of dual-layer nested Mamba modules to comprehensively capture contextual information across different scales, coupled with a hierarchical feature fusion strategy to enhance multi-scale feature integration for improved small object detection. The LVMSOD framework incorporates two key components to enhance feature representation: Visual Mamba Detection (VDM) block that preserves fine-grained image details, and Lightweight Gated Multi-Layer Perceptron (LGMLP) designed to model local feature dependencies efficiently. Furthermore, we propose an optimized feature extraction mechanism employing depthwise and pointwise convolutions with distribution factors, significantly reducing computational overhead while maintaining detection accuracy. Comprehensive evaluations on benchmark datasets demonstrate the effectiveness of our approach, achieving [email protected] scores of 93.3% on KITTI and 45.3% on VisDrone while maintaining superior computational efficiency. These results highlight the potential of LVMSOD for efficient and accurate small object detection. Ming Gao 0012, Yinlin Wang, Yang Li 0093, Manjiang Hu, Yougang Bian, Rongjun Ding |
IEEE Internet Things J. | 2 |
| 2026 | MOAT: Multi-Scale Group Interaction Transformer for Trajectory Prediction in Crowded ScenesabstractPredicting pedestrian trajectories in complex environments presents a significant challenge for autonomous driving systems, primarily due to the intricate social interactions among pedestrians and their surrounding groups. Existing methods often struggle to fully capture the impact of group behavior on individual movement. To address these limitations, we propose the Multi-scale Group Interaction Transformer (MOAT) for pedestrian trajectory prediction in high-density, complex scenarios. Our approach introduces a dynamic group interaction module (DGIM) that clusters pedestrians based on proximity, determined by a distance matrix of neighboring pedestrians. In constructing interaction representations, we go beyond traditional features like speed, distance, and direction by incorporating crowd density, thus providing a more comprehensive understanding of group dynamics. To effectively process these features, we employ a multi-branch attention fusion (MBAF) module, which independently analyzes each feature set to capture the unique dynamics and density characteristics of each group and their varying effects on the target pedestrian. These spatial features are then combined with temporal information, allowing our model to account for both spatial and temporal dependencies. Additionally, we leverage a multi-scale Transformer to adaptively partition input trajectories, enhancing the model’s ability to capture dynamic patterns across various scales. Extensive evaluations on benchmark datasets show that our approach consistently outperforms state-of-the-art methods in terms of both prediction accuracy and robustness. Ming Gao 0012, Yinlin Wang, Guotao Xie, Chi Ding, Yougang Bian |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2026 | Criticality Assessment Model for Intelligent Vehicle Test Scenario Based on Interactive Field Feature and Hypergraph LearningabstractScenario-based testing is an important part in intelligent vehicle (IV) development. The data volume of collected test scenarios is extremely large, and directly using all collected scenarios to test IVs will lead to extremely low testing efficiency. To solve this problem, a criticality assessment model (CAM) for IV test scenario based on interactive field feature (IFF) and hypergraph learning is proposed to quantify the test scenario criticality to improve the test efficiency. The IFF is constructed based on the potential field-based method to integrally consider the multidimensional coupling of scenario elements. In addition, the interaction between the fields generated by the ego vehicle and the driving environment is modeled based on Delaunay triangulation discretization method to accurately quantify the driving environment risk to the ego vehicle. The node and hyperedge of the hypergraph are used to model the individual dynamic evolution and group interaction characteristics of vehicles, respectively. Subsequently, a hypergraph learning network is constructed to extract features from the built hypergraph, IFF and traffic elements. Finally, the effectiveness, reasonableness and accuracy validation experiments are designed to validate the proposed CAM. Ablation experiment results show that the proposed IFF and hypergraph learning network enhance the CAM accuracy. The reasonableness validation results show that the proposed CAM can better find critical test scenarios than time-to-collision and time-head-way methods. The accuracy of the proposed CAM is compared through four real validation scenarios in the proving ground. The comparison results show that the proposed CAM can accurately output the quantified test scenario criticality. Yinzi Huang, Bing Zhu 0006, Jian Zhao 0007, Jiayi Han, Dongjian Song, Peixing Zhang, Shizheng Jia, Ming Gao 0012 |
IEEE Trans. Intell. Transp. Syst. | 8 |
| 2025 | DiffWT: Diffusion-Based Pedestrian Trajectory Prediction With Time-Frequency Wavelet TransformabstractAccurate pedestrian trajectory prediction is a crucial task for ensuring the safety of autonomous driving. However, most of the existing methods only model pedestrian trajectories in the spatial-temporal domain, which results in a lack of analysis of motion at different scales. In this work, we propose a framework based on wavelet transform and diffusion model, which is called DiffWT. Different from previous approaches, our method employs a discrete wavelet transform (DWT) to perform time-frequency analysis of trajectories. The high-frequency component of the DWT indicates local motion details of the trajectory, while the low-frequency one represents overall motion trends. Second, we propose a trajectory decoder comprising a conditional diffusion model and a cross-constrained bidirectional trajectory generator (C2Bid). The diffusion model generates the distribution of implicit pedestrian behaviors by taking the multiscale motion behaviors as conditions. Furthermore, the C2Bid module is designed as a cross-constrained bidirectional structure to decode behavioral distribution into multimodal trajectories. This trajectory decoder can generate precise distribution of trajectories and reduce accumulation of prediction errors. Extensive experimental results on the ETH/UCY and the Stanford drone datasets (SDDs) demonstrate that our method achieves better performance as well as higher efficiency compared to other state-of-the-art approaches. Xin Chen 0125, Ming Gao 0012, Chi Ding, Yougang Bian |
IEEE Internet Things J. | 3 |
| 2025 | Hybrid Matching Teacher Framework for Cross-Domain Visual Detection TransformerabstractObject detection is a critical component of autonomous vehicle perception systems. However, domain shifts between training environments and real-world scenarios often degrade detector performance. Cross-domain object detection aims to adapt detectors to unlabeled target domains utilizing only labeled source data. Recent popular cross-domain object detection methods employ the mean teacher framework, which uses pseudo-labels generated by the teacher model to guide training on unlabeled real-world data. Despite its effectiveness, continuous training with noisy pseudo-labels leads to abnormal performance degradation in the later stages of training. To address this issue, we propose a novel Hybrid Matching Teacher (HMT) framework for cross-domain visual detection transformers, which enhances cross-domain knowledge transfer across pseudo-label generation, filtering, and training processes. Specifically, we design a Feature Sparse Alignment (FSA) module to adapt DETR tokens and queries, generate domain-adaptive weights to initialize the teacher-student models, and mitigate the inherent initial source bias in the teacher model. Next, a Localization-aware Pseudo-label Filtering (LPF) module ensures high-quality pseudo-labels by considering the consistency between localization and classification tasks. Furthermore, to improve the efficiency of pseudo-label training, the Cross-view Hybrid Matching (CHM) module introduces an auxiliary matching branch to increase the number of positive queries that match with pseudo-labels. Extensive experiments demonstrate that our approach achieves state-of-the-art performance, outperforming previous benchmarks by 3.1%, 8.5%, and 4.4% in adverse weather, diverse scenes, and synthetic-to-real, respectively. Xiaowei Wang 0001, Jinhui Suo, Yang Li 0093, Ming Gao 0012, Peiwen Jiang, Pengwen Dai |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Dynamic multi-scale spatial-temporal graph convolutional network for traffic flow predictionabstractThis paper proposes a dynamic multi-scale spatial-temporal graph convolutional network (DS-STGCN) for traffic flow prediction . The network aims to comprehensively extract global and local dependencies in dynamic spatial-temporal data by inputting traffic network flow data to construct node feature graphs, topology graphs , and time slot feature graphs, capturing the complexity and dynamics of traffic flow. DS-STGCN interprets feature information of the traffic network from both spatial and temporal dimensions through dynamic multi-scale graph convolutional blocks. In the spatial dimension, these blocks use constraints at different levels to balance fine-grained local features and extensive global features, revealing the intrinsic structure of traffic flow data. In the temporal dimension, these blocks jointly learn with temporal convolutional blocks to capture multi-frequency time patterns and handle long sequence data, effectively extracting potential dependencies of time series . Furthermore, DS-STGCN effectively models the changing spatial-temporal relationships in road network flow by constructing dynamically adaptive updated adjacency tensors, generating dynamic graph structures to address the challenge of changing spatial-temporal relationships in the transportation system. Experimental results show that our method significantly outperforms other competing methods on five real traffic datasets (PEMS03, PEMS04, PEMS07, PEMS08 and METR-LA). Ming Gao 0012, Zhuoran Du, Hongmao Qin, Guangyin Jin, Guotao Xie |
Knowl. Based Syst. | 1 |
| 2024 | Progressive Critical Region Transfer for Cross-Domain Visual Object DetectionabstractWell-trained visual object detectors are generally confronted with a severe performance decline when deployed in a novel driving scenario due to the impact of domain shift. Despite excellent improvements in unsupervised domain adaptive object detection achieved by adversarial training, those approaches fail to capture the transfer core underlying the holistic scenes. To solve this problem, we propose a progressive critical region transfer framework for cross-domain visual object detection. Specifically, we exploit a potential foreground mining (PFM) module and a semantic-specific RoI aggregation (SRA) module to improve the robustness of the cross-domain detection framework. Upon the critical regions in the broad sense, the PFM module first highlights the foreground regions by reweighting the hierarchical feature maps in sequence, and then modifies location biases at the downstream position of the backbone network for more accurate upstream predictions. Deep into the critical regions in the narrow sense, the SRA module concentrates on establishing an appropriate matching between batch-wise RoIs and all semantic centers, and further strengthens the aggregation of cross-domain identical semantic with the complement of context references. Together these modules are obligated to transform the adaptation importance from the whole scope to the latent foreground areas, and afterward to the informative regions of interest along the detection pipeline. Experiments show that our progressive critical region transfer framework achieves a state-of-the-art performance in adverse weather, camera configuration, and complicated scene adaptation, which outperforms the baselines by 19.4%, 5.0%, and 6.1%, respectively. Xiaowei Wang 0001, Peiwen Jiang, Yang Li 0093, Manjiang Hu, Ming Gao 0012, Dongpu Cao, Rongjun Ding |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | PPF-Det: Point-Pixel Fusion for Multi-Modal 3D Object DetectionabstractMulti-modal fusion can take advantage of the LiDAR and camera to boost the robustness and performance of 3D object detection. However, there are still of great challenges to comprehensively exploit image information and perform accurate diverse feature interaction fusion. In this paper, we proposed a novel multi-modal framework, namely Point-Pixel Fusion for Multi-Modal 3D Object Detection (PPF-Det). The PPF-Det consists of three submodules, Multi Pixel Perception (MPP), Shared Combined Point Feature Encoder (SCPFE), and Point-Voxel-Wise Triple Attention Fusion (PVW-TAF) to address the above problems. Firstly, MPP can make full use of image semantic information to mitigate the problem of resolution mismatch between point cloud and image. In addition, we proposed SCPFE to preliminary extract point cloud features and point-pixel features simultaneously reducing time-consuming on 3D space. Lastly, we proposed a fine alignment fusion strategy PVW-TAF to generate multi-level voxel-fused features based on attention mechanism. Extensive experiments on KITTI benchmarks, conducted on September 24, 2023, demonstrate that our method shows excellent performance. Guotao Xie, Ming Gao 0012, Manjiang Hu, Xiaohui Qin 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2020 | Manifold Siamese Network: A Novel Visual Tracking ConvNet for Autonomous VehiclesabstractVisual tracking is a vital component of autonomous driving perception system. Siamese networks have achieved great success in both accuracy and speed for visual tracking tasks. These Siamese trackers share a similar framework in which each tracker consists of two network branches for exploring semantic information. However, the performance of Siamese trackers is limited by an insufficient semantic template and an unsatisfactory updating strategy. To tackle these problems, we propose a manifold Siamese network for visual tracking that can simultaneously utilize semantic and geometric information. A manifold sample pool is constructed to exploit the manifold structure of image object sequences. This sample pool is dynamically learned via a fast Gaussian mixture model (GMM). After obtaining a manifold sample template, we design a deep architecture based on a correlation filter (CF) network and append a novel manifold feature branch. The network remains fully convolutional and can train a template to discriminate exemplar image and arbitrarily size search image. Then, a triplet occlusion score function cooperates with an effective update method that is established to prevent model drift. Extensive experiments show that the proposed tracking algorithm performs favorably compared with the state-of-the-art methods on three standard benchmark datasets at a high framerate, which is very suitable for autonomous driving. Ming Gao 0012, Lisheng Jin, Baicang Guo |
IEEE Trans. Intell. Transp. Syst. | 1 |