Libo Weng

dblp:184/7384 · DBLP profile ↗
← Back
24ranked-venue papers
6as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Spatial-temporal domain generalization for cross-city traffic prediction
Shengzhe You, Libo Weng, Yanjing Lei, Fei Gao 0014
Expert Syst. Appl.2
2026 Injecting image text structure and edge priors into segment anything for scene text segmentation
Qian Shao, Libo Weng, Yanjing Lei, Xianxun Zhu, Hui Chen 0026
Image Vis. Comput.2
2025 An Automatic Extrinsic Calibration Method for LiDAR-Camera Fusion via Combining Semantic and Geometric Features
abstract
Precise extrinsic calibration is one of the key techniques for LiDAR-camera fusion system. In current methods, the extrinsic calibration is usually not automatic. To address this, an automatic calibration method via combining semantic and geometric features is proposed, which is not dependent on any specific calibration object. First, extrinsics are automatically initialized; semantic objects are utilized to formulate the edge constraints and projection boundary constraints. Then, an efficient global optimization algorithm that synergizes the Jacobian matrix and the stochastic strategy of simulated annealing is put forward to calculate precise extrinsics. A feedback mechanism is designed to evaluate the reliability of the proposed method. Experiments on the KITTI dataset show that the proposed method achieves a rotation error of 0.14°and a translation error of 4.5cm, outperforming most current methods. Besides, the experiment on the proposed optimization algorithm is also conducted to verify its effectiveness and efficiency.
Minqian Wang, Libo Weng, Fei Gao 0014
ICASSP2
2025 SGAD: An Unsupervised Secondary-Guided Diffusion Model for Industrial Anomaly Detection
abstract
Reconstruction-based anomaly detection methods often struggle with invariant reconstruction of abnormal regions and the unintended reconstruction of novel anomalies. To address these limitations, this study proposes a novel guided training and reconstruction framework (SGAD) to enhance anomaly reconstruction quality. The proposed approach integrates a training paradigm built on target images and fusion loss, along with target-guided and secondary reconstruction strategies utilizing a diffusion model, achieving superior anomaly detection performance. Additionally, a new DTY anomaly detection dataset is introduced to benchmark the approach. Extensive experiments were conducted on the DTY and MVTec datasets, demonstrating that SGAD achieves state-of-the-art performance, with mean scores of 93.7% I-AUROC and 86.2% P-AUROC. These results highlight the effectiveness and robustness of SGAD in addressing complex anomaly detection challenges, underscoring its potential for deployment in practical production environments.
Wenze Kang, Libo Weng, Zhenbo Cheng, Fei Gao 0014
ICME3
2025 ViTraj: Learning Dual-Side Representations for Vehicle-Infrastructure Cooperative Trajectory Prediction
abstract
While autonomous driving has made substantial progress, accurately predicting the trajectories of surrounding traffic agents remains a fundamental challenge for ensuring safety. Integrating both infrastructure-side and vehicle-side information has the potential to enhance perception and prediction capabilities. However, existing methods overlook the challenges in Vehicle-Infrastructure Cooperative Trajectory Prediction. To bridge this gap, we propose ViTraj, a model-agnostic framework for VIC-TP that leverages infrastructure-side trajectories to mitigate the inherent limitations of vehicle-side forecasting. ViTraj introduces a Feature-Side Selection and a Cooperative Interaction to aggregate complementary features from both sides, effectively expanding the perceptual horizon of prediction models. In addition, we present a Vehicle-Infrastructure Knowledge Distillation strategy to enforce consistency between multi-side predictions, which efficient global-local feature alignment through a single backward pass. Extensive experiments on large-scale public datasets demonstrate that ViTraj consistently improves advanced trajectory prediction models, achieving the state-of-the-art performance compared to existing vehicle-infrastructure cooperative methods. We believe this work provides a promising step toward the practical deployment of V2X-based autonomous driving systems.
Shengzhe You, Libo Weng, Fei Gao 0014
ACM Multimedia2
2025 Incremental few-shot instance segmentation without fine-tuning on novel classes
Luofeng Zhang, Libo Weng, Fei Gao 0014
Comput. Vis. Image Underst.2
2025 Lightweight binary convolutional-transformers fusion network for facial expression recognition
Xiyin Wu, Libo Weng, Qiaolin Ye
Eng. Appl. Artif. Intell.3
2024 Weakly Supervised Few-Shot Segmentation Through Textual Prompt
abstract
Recently, significant progress has been made in few-shot segmentation (FSS), which aims to segment unknown objects with only a few support images. However, during both training and testing, FSS still requires pixel-level annotations. When only image-level labels are available, FSS will become a more challenging task, namely weakly supervised few-shot segmentation (WS-FSS). To address this problem, this paper proposes a novel text-driven approach, which replaces pixel-level labels with textual prompts. To guide the model in selecting the target features and capturing the inter-class correlations, a Text-Image Matching Module (TIMM) and a Text Supervision Scheme (TSS) are designed for the feature matching and decoding stages, respectively. Extensive experiments are conducted on two public datasets, PASCAL-5iand COCO-20i. The experimental results demonstrate that our method not only outperforms existing state-of-the-art WS-FSS methods but also achieves comparable or even superior performance to advanced FSS models. The code can be available at https: //github.com/Joseph-Lee-V/Text-WS-FSS.
Shengzhe You, Libo Weng, Fei Gao 0014
ICASSP2
2024 BFIDet: A YOLOv7-improved Vehicle and Pedestrian Detector via Balancing Feature Integration
abstract
Accurate vehicle and pedestrian detection are fundamental for safe driving and maintenance of traffic order. In this paper, a YOLOv7-improved vehicle and pedestrian detector via balancing feature integration (BFIDet) is proposed. First, EFFM module is designed to facilitate feature map fusion across layers. Second, GSRFConv is utilized to expand the receptive field of the intrinsic feature map as a way to improve the feature discriminability and robustness. VFBM module is then introduced to guide the propagation of the information flow as a way to solve the problem of dilution of features in non-adjacent layers and semantic differences between cross-scale features. In the experiments, the proposed method achieves 93.9% and 69.4% [email protected] and [email protected]:0.95 metric on the KITTI dataset, which are 2.1% and 1.5% better than YOLOv7, respectively, and the [email protected] metric on the SODA10M dataset reaches 63.1% with an improvement of 0.9% and 1.9% over YOLOv7 and YOLOv8m, respectively. The experimental results demonstrate that the proposed BFIDet is more accuracy than that of other mainstream models with controllable computational consumption.
Anrui Wang, Libo Weng, Fei Gao 0014
ICMR2
2024 FP3Seg: Point Cloud Panoptic Segmentation via LiDAR-Camera Fusion and Progressive Decoder
abstract
Point cloud panoptic segmentation is a 3D scene perception task that provides a holistic solution for both semantic and instance segmentation. The sparsity and lack of texture features in LiDAR point cloud, coupled with the relatively narrow field of view of camera, make multi-modal fusion challenging. In this paper, we propose a novel multi-modal fusion based point cloud panoptic segmentation method, named FP3Seg, with main contributions including Hybrid Domain Adaptive Fusion (HDAF) module and Progressive Decoder. HDAF employs learnable weights to adaptively fuse multi-modal features in both the channel and spatial domains. Through knowledge distillation, FP3Seg extends the benefits of multi-modal fusion beyond the camera field of view. Pro-gressive Decoder embeds semantic and instance information into the input of panoptic decoder, assisting the decoder in understanding the distinctions between stuff and thing classes. The proposed method is benchmarked on the SemanticKITTI test set, achieving 57.7% PQ, showing a 1.7% improvement over baseline. Experimental results demonstrate that FP3Seg possesses advantages over single-modal approaches in multiple aspects, especially for the segmentation of thing classes.
Xianyou Dai, Libo Weng, Fei Gao 0014
SMC2
2024 EFFDet: A Crack Detector via Boundary Preservation and Cross-Attention Integration
abstract
The complexity of scenes and the topology of cracks make road crack detection a challenging task. Compared to other semantic segmentation tasks, this mission places a greater demand on the network's ability to preserve detailed boundary information. To address this, a novel road crack detection network architecture EFFDet is proposed in this paper. Firstly, we redesign the encoding-decoding module based on large-scale convolutional kernels and attention mechanisms to reduce the loss of detailed information caused by downsampling. Secondly, the Cross Attention module is proposed to integrate more precise details into the output of the decoding layer. In comparative experiments on four datasets, CRACK500, Volker, CrackLS315 and DeepCrack, EFFDet achieves ODS values of 0.7434, 0.6758, 0.6449 and 0.8708, respectively. The experimental results show that EFFDet demonstrates stronger detection capabilities in road crack detection.
Linhua Gao, Libo Weng, Fei Gao 0014
SMC2
2024 SIF-TF: A Scene-Interaction fusion Transformer for trajectory prediction
Fei Gao 0014, Wanjun Huang, Libo Weng
Knowl. Based Syst.3
2023 ECDet: A Real-Time Vehicle Detection Network for CPU-Only Devices
Fei Gao 0014, Jianwen Shao, Xinyang Dong, Libo Weng
ICANN (7)5
2023 Language Guided Graph Transformer for Skeleton Action Recognition
Libo Weng, Weidong Lou, Fei Gao 0014
ICONIP (10)1
2023 Traffic Sign Recognition Model Based on Small Object Detection
Fei Gao 0014, Wanjun Huang, Xiuqi Chen, Libo Weng
PRICAI (3)4
2023 Whether and how is a surveillance camera jittering? A ROR perception based framework and method
Fei Gao 0014, Kaitao Mei, Libo Weng, Yaozhong Zhuang
Appl. Intell.3
2023 A 3D graph convolutional networks model for 2D skeleton-based human action recognition
abstract
Abstract With the popularity of cameras, the application of action recognition is more and more extensive. After the emergence of RGB‐D cameras and human pose estimation algorithms, human actions can be represented by a sequence of skeleton joints. Therefore, skeleton‐based action recognition has been a research hotspot. In this paper, a novel 3D Graph Convolutional Network model (3D‐GCN) with space‐time attention mechanism for 2D skeleton data is proposed. Three‐dimensional graph convolution is employed to extract spatiotemporal features of skeleton descriptor that is composed of joint coordinates, frame differences and angles. Meanwhile, different joints and different frames are given different attention to achieve action classification. A zebra crossing pedestrian dataset named ZCP is also provided, which simulates possible pedestrian actions on the zebra crossing in real scenes. Experimental evaluation is carried out on ZCP dataset and NTU RGB+D dataset. Experimental results show that our method is better than current 2D‐based methods and is comparable with 3D methods.
Libo Weng, Weidong Lou, Fei Gao 0014
IET Image Process.1
2023 A semantic-aware monocular projection model for accurate pose measurement
Libo Weng, Xiuqi Chen, Qi Qiu, Yaozhong Zhuang, Fei Gao 0014
Pattern Anal. Appl.1
2022 Traffic Scene Perception Based on Joint Object Detection and Semantic Segmentation
Libo Weng, Fei Gao 0014
Neural Process. Lett.1
2022 A Trajectory Evaluator by Sub-tracks for Detecting VOT-based Anomalous Trajectory
abstract
With the popularization of visual object tracking (VOT), more and more trajectory data are obtained and have begun to gain widespread attention in the fields of mobile robots, intelligent video surveillance, and the like. How to clean the anomalous trajectories hidden in the massive data has become one of the research hotspots. Anomalous trajectories should be detected and cleaned before the trajectory data can be effectively used. In this article, a Trajectory Evaluator by Sub-tracks (TES) for detecting VOT-based anomalous trajectory is proposed. Feature of Anomalousness is defined and described as the Eigenvector of classifier to filter Track Lets anomalous trajectory and IDentity Switch anomalous trajectory, which includes Feature of Anomalous Pose and Feature of Anomalous Sub-tracks (FAS). In the comparative experiments, TES achieves better results on different scenes than state-of-the-art methods. Moreover, FAS makes better performance than point flow, least square method fitting and Chebyshev Polynomial Fitting. It is verified that TES is more accurate and effective and is conducive to the sub-tracks trajectory data analysis.
Fei Gao 0014, Jiada Li, Yisu Ge, Jianwen Shao, Shufang Lu, Libo Weng
ACM Trans. Knowl. Discov. Data6
2019 Sparse graphs with smoothness constraints: Application to dimensionality reduction and semi-supervised classification
Fadi Dornaika, Libo Weng
Pattern Recognit.2
2018 Structured sparse graphs using manifold constraints for visual data analysis
Fadi Dornaika, Libo Weng, Zhong Jin
Neurocomputing2
2016 Graph construction based on data self-representativeness and Laplacian smoothness
Libo Weng, Fadi Dornaika, Zhong Jin
Neurocomputing1
2016 Flexible constrained sparsity preserving embedding
Libo Weng, Fadi Dornaika, Zhong Jin
Pattern Recognit.1