VLDB 2026 Research / reviewers in the wild / expert
Libo Weng
dblp:184/7384
· DBLP profile ↗
24ranked-venue papers
6as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Spatial-temporal domain generalization for cross-city traffic prediction
Shengzhe You, Libo Weng, Yanjing Lei, Fei Gao 0014 |
Expert Syst. Appl. | 2 |
| 2026 | Injecting image text structure and edge priors into segment anything for scene text segmentation
Qian Shao, Libo Weng, Yanjing Lei, Xianxun Zhu, Hui Chen 0026 |
Image Vis. Comput. | 2 |
| 2025 | An Automatic Extrinsic Calibration Method for LiDAR-Camera Fusion via Combining Semantic and Geometric FeaturesabstractPrecise extrinsic calibration is one of the key techniques for LiDAR-camera fusion system. In current methods, the extrinsic calibration is usually not automatic. To address this, an automatic calibration method via combining semantic and geometric features is proposed, which is not dependent on any specific calibration object. First, extrinsics are automatically initialized; semantic objects are utilized to formulate the edge constraints and projection boundary constraints. Then, an efficient global optimization algorithm that synergizes the Jacobian matrix and the stochastic strategy of simulated annealing is put forward to calculate precise extrinsics. A feedback mechanism is designed to evaluate the reliability of the proposed method. Experiments on the KITTI dataset show that the proposed method achieves a rotation error of 0.14°and a translation error of 4.5cm, outperforming most current methods. Besides, the experiment on the proposed optimization algorithm is also conducted to verify its effectiveness and efficiency. Minqian Wang, Libo Weng, Fei Gao 0014 |
ICASSP | 2 |
| 2025 | SGAD: An Unsupervised Secondary-Guided Diffusion Model for Industrial Anomaly DetectionabstractReconstruction-based anomaly detection methods often struggle with invariant reconstruction of abnormal regions and the unintended reconstruction of novel anomalies. To address these limitations, this study proposes a novel guided training and reconstruction framework (SGAD) to enhance anomaly reconstruction quality. The proposed approach integrates a training paradigm built on target images and fusion loss, along with target-guided and secondary reconstruction strategies utilizing a diffusion model, achieving superior anomaly detection performance. Additionally, a new DTY anomaly detection dataset is introduced to benchmark the approach. Extensive experiments were conducted on the DTY and MVTec datasets, demonstrating that SGAD achieves state-of-the-art performance, with mean scores of 93.7% I-AUROC and 86.2% P-AUROC. These results highlight the effectiveness and robustness of SGAD in addressing complex anomaly detection challenges, underscoring its potential for deployment in practical production environments. Wenze Kang, Libo Weng, Zhenbo Cheng, Fei Gao 0014 |
ICME | 3 |
| 2025 | ViTraj: Learning Dual-Side Representations for Vehicle-Infrastructure Cooperative Trajectory PredictionabstractWhile autonomous driving has made substantial progress, accurately predicting the trajectories of surrounding traffic agents remains a fundamental challenge for ensuring safety. Integrating both infrastructure-side and vehicle-side information has the potential to enhance perception and prediction capabilities. However, existing methods overlook the challenges in Vehicle-Infrastructure Cooperative Trajectory Prediction. To bridge this gap, we propose ViTraj, a model-agnostic framework for VIC-TP that leverages infrastructure-side trajectories to mitigate the inherent limitations of vehicle-side forecasting. ViTraj introduces a Feature-Side Selection and a Cooperative Interaction to aggregate complementary features from both sides, effectively expanding the perceptual horizon of prediction models. In addition, we present a Vehicle-Infrastructure Knowledge Distillation strategy to enforce consistency between multi-side predictions, which efficient global-local feature alignment through a single backward pass. Extensive experiments on large-scale public datasets demonstrate that ViTraj consistently improves advanced trajectory prediction models, achieving the state-of-the-art performance compared to existing vehicle-infrastructure cooperative methods. We believe this work provides a promising step toward the practical deployment of V2X-based autonomous driving systems. Shengzhe You, Libo Weng, Fei Gao 0014 |
ACM Multimedia | 2 |
| 2025 | Incremental few-shot instance segmentation without fine-tuning on novel classes
Luofeng Zhang, Libo Weng, Fei Gao 0014 |
Comput. Vis. Image Underst. | 2 |
| 2025 | Lightweight binary convolutional-transformers fusion network for facial expression recognition
Xiyin Wu, Libo Weng, Qiaolin Ye |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | Weakly Supervised Few-Shot Segmentation Through Textual PromptabstractRecently, significant progress has been made in few-shot segmentation (FSS), which aims to segment unknown objects with only a few support images. However, during both training and testing, FSS still requires pixel-level annotations. When only image-level labels are available, FSS will become a more challenging task, namely weakly supervised few-shot segmentation (WS-FSS). To address this problem, this paper proposes a novel text-driven approach, which replaces pixel-level labels with textual prompts. To guide the model in selecting the target features and capturing the inter-class correlations, a Text-Image Matching Module (TIMM) and a Text Supervision Scheme (TSS) are designed for the feature matching and decoding stages, respectively. Extensive experiments are conducted on two public datasets, PASCAL-5iand COCO-20i. The experimental results demonstrate that our method not only outperforms existing state-of-the-art WS-FSS methods but also achieves comparable or even superior performance to advanced FSS models. The code can be available at https: //github.com/Joseph-Lee-V/Text-WS-FSS. Shengzhe You, Libo Weng, Fei Gao 0014 |
ICASSP | 2 |
| 2024 | BFIDet: A YOLOv7-improved Vehicle and Pedestrian Detector via Balancing Feature IntegrationabstractAccurate vehicle and pedestrian detection are fundamental for safe driving and maintenance of traffic order. In this paper, a YOLOv7-improved vehicle and pedestrian detector via balancing feature integration (BFIDet) is proposed. First, EFFM module is designed to facilitate feature map fusion across layers. Second, GSRFConv is utilized to expand the receptive field of the intrinsic feature map as a way to improve the feature discriminability and robustness. VFBM module is then introduced to guide the propagation of the information flow as a way to solve the problem of dilution of features in non-adjacent layers and semantic differences between cross-scale features. In the experiments, the proposed method achieves 93.9% and 69.4% [email protected] and [email protected]:0.95 metric on the KITTI dataset, which are 2.1% and 1.5% better than YOLOv7, respectively, and the [email protected] metric on the SODA10M dataset reaches 63.1% with an improvement of 0.9% and 1.9% over YOLOv7 and YOLOv8m, respectively. The experimental results demonstrate that the proposed BFIDet is more accuracy than that of other mainstream models with controllable computational consumption. Anrui Wang, Libo Weng, Fei Gao 0014 |
ICMR | 2 |
| 2024 | FP3Seg: Point Cloud Panoptic Segmentation via LiDAR-Camera Fusion and Progressive DecoderabstractPoint cloud panoptic segmentation is a 3D scene perception task that provides a holistic solution for both semantic and instance segmentation. The sparsity and lack of texture features in LiDAR point cloud, coupled with the relatively narrow field of view of camera, make multi-modal fusion challenging. In this paper, we propose a novel multi-modal fusion based point cloud panoptic segmentation method, named FP3Seg, with main contributions including Hybrid Domain Adaptive Fusion (HDAF) module and Progressive Decoder. HDAF employs learnable weights to adaptively fuse multi-modal features in both the channel and spatial domains. Through knowledge distillation, FP3Seg extends the benefits of multi-modal fusion beyond the camera field of view. Pro-gressive Decoder embeds semantic and instance information into the input of panoptic decoder, assisting the decoder in understanding the distinctions between stuff and thing classes. The proposed method is benchmarked on the SemanticKITTI test set, achieving 57.7% PQ, showing a 1.7% improvement over baseline. Experimental results demonstrate that FP3Seg possesses advantages over single-modal approaches in multiple aspects, especially for the segmentation of thing classes. Xianyou Dai, Libo Weng, Fei Gao 0014 |
SMC | 2 |
| 2024 | EFFDet: A Crack Detector via Boundary Preservation and Cross-Attention IntegrationabstractThe complexity of scenes and the topology of cracks make road crack detection a challenging task. Compared to other semantic segmentation tasks, this mission places a greater demand on the network's ability to preserve detailed boundary information. To address this, a novel road crack detection network architecture EFFDet is proposed in this paper. Firstly, we redesign the encoding-decoding module based on large-scale convolutional kernels and attention mechanisms to reduce the loss of detailed information caused by downsampling. Secondly, the Cross Attention module is proposed to integrate more precise details into the output of the decoding layer. In comparative experiments on four datasets, CRACK500, Volker, CrackLS315 and DeepCrack, EFFDet achieves ODS values of 0.7434, 0.6758, 0.6449 and 0.8708, respectively. The experimental results show that EFFDet demonstrates stronger detection capabilities in road crack detection. Linhua Gao, Libo Weng, Fei Gao 0014 |
SMC | 2 |
| 2024 | SIF-TF: A Scene-Interaction fusion Transformer for trajectory prediction
Fei Gao 0014, Wanjun Huang, Libo Weng |
Knowl. Based Syst. | 3 |
| 2023 | ECDet: A Real-Time Vehicle Detection Network for CPU-Only Devices
Fei Gao 0014, Jianwen Shao, Xinyang Dong, Libo Weng |
ICANN (7) | 5 |
| 2023 | Language Guided Graph Transformer for Skeleton Action Recognition
Libo Weng, Weidong Lou, Fei Gao 0014 |
ICONIP (10) | 1 |
| 2023 | Traffic Sign Recognition Model Based on Small Object Detection
Fei Gao 0014, Wanjun Huang, Xiuqi Chen, Libo Weng |
PRICAI (3) | 4 |
| 2023 | Whether and how is a surveillance camera jittering? A ROR perception based framework and method
Fei Gao 0014, Kaitao Mei, Libo Weng, Yaozhong Zhuang |
Appl. Intell. | 3 |
| 2023 | A 3D graph convolutional networks model for 2D skeleton-based human action recognitionabstractAbstract With the popularity of cameras, the application of action recognition is more and more extensive. After the emergence of RGB‐D cameras and human pose estimation algorithms, human actions can be represented by a sequence of skeleton joints. Therefore, skeleton‐based action recognition has been a research hotspot. In this paper, a novel 3D Graph Convolutional Network model (3D‐GCN) with space‐time attention mechanism for 2D skeleton data is proposed. Three‐dimensional graph convolution is employed to extract spatiotemporal features of skeleton descriptor that is composed of joint coordinates, frame differences and angles. Meanwhile, different joints and different frames are given different attention to achieve action classification. A zebra crossing pedestrian dataset named ZCP is also provided, which simulates possible pedestrian actions on the zebra crossing in real scenes. Experimental evaluation is carried out on ZCP dataset and NTU RGB+D dataset. Experimental results show that our method is better than current 2D‐based methods and is comparable with 3D methods. Libo Weng, Weidong Lou, Fei Gao 0014 |
IET Image Process. | 1 |
| 2023 | A semantic-aware monocular projection model for accurate pose measurement
Libo Weng, Xiuqi Chen, Qi Qiu, Yaozhong Zhuang, Fei Gao 0014 |
Pattern Anal. Appl. | 1 |
| 2022 | Traffic Scene Perception Based on Joint Object Detection and Semantic Segmentation
Libo Weng, Fei Gao 0014 |
Neural Process. Lett. | 1 |
| 2022 | A Trajectory Evaluator by Sub-tracks for Detecting VOT-based Anomalous TrajectoryabstractWith the popularization of visual object tracking (VOT), more and more trajectory data are obtained and have begun to gain widespread attention in the fields of mobile robots, intelligent video surveillance, and the like. How to clean the anomalous trajectories hidden in the massive data has become one of the research hotspots. Anomalous trajectories should be detected and cleaned before the trajectory data can be effectively used. In this article, a Trajectory Evaluator by Sub-tracks (TES) for detecting VOT-based anomalous trajectory is proposed. Feature of Anomalousness is defined and described as the Eigenvector of classifier to filter Track Lets anomalous trajectory and IDentity Switch anomalous trajectory, which includes Feature of Anomalous Pose and Feature of Anomalous Sub-tracks (FAS). In the comparative experiments, TES achieves better results on different scenes than state-of-the-art methods. Moreover, FAS makes better performance than point flow, least square method fitting and Chebyshev Polynomial Fitting. It is verified that TES is more accurate and effective and is conducive to the sub-tracks trajectory data analysis. Fei Gao 0014, Jiada Li, Yisu Ge, Jianwen Shao, Shufang Lu, Libo Weng |
ACM Trans. Knowl. Discov. Data | 6 |
| 2019 | Sparse graphs with smoothness constraints: Application to dimensionality reduction and semi-supervised classification
Fadi Dornaika, Libo Weng |
Pattern Recognit. | 2 |
| 2018 | Structured sparse graphs using manifold constraints for visual data analysis
Fadi Dornaika, Libo Weng, Zhong Jin |
Neurocomputing | 2 |
| 2016 | Graph construction based on data self-representativeness and Laplacian smoothness
Libo Weng, Fadi Dornaika, Zhong Jin |
Neurocomputing | 1 |
| 2016 | Flexible constrained sparsity preserving embedding
Libo Weng, Fadi Dornaika, Zhong Jin |
Pattern Recognit. | 1 |