EDBT 2026 Demo / reviewers in the wild / expert
Fei Gao 0014
dblp:16/722-14
· DBLP profile ↗
47ranked-venue papers
17as first author
38since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 7 first-author · 17 since 2021Artificial intelligence and machine learning · 17 · 9 first-author · 14 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Text-Guided Prototype Replay and Classifier Guidance for Incremental Few-Shot Semantic SegmentationabstractIncremental few-shot semantic segmentation (iFSS) aims to incrementally acquire new knowledge from limited labeled samples while retaining previously learned concepts, without relying on large-scale manual annotations. Despite its practical importance, effective iFSS methods remain limited. In this paper, we propose a multimodal framework to address this challenge. First, textual features are leveraged to supervise the classifier weights, mitigating forgetting of base classes and overfitting to novel classes. Second, class prototype images are generated from textual features to support low-cost replay of previous knowledge. Finally, a momentum-based updating strategy is introduced to decouple the background and novel class weights within the classifier of the old model used for knowledge distillation. Extensive experiments and ablation studies validate the effectiveness of the approach. The method achieves strong performance in continual learning of novel classes while preserving knowledge of old ones, closely mirroring human-like few-shot learning over time. Luofeng Zhang, Shengzhe You, Qian Shao, Yanjing Lei, Fei Gao 0014 |
ICMR | 6 |
| 2026 | Vector sketch animation generation with differentialable motion trajectoriesabstractAbstract Sketching is a direct and inexpensive means of visual expression. Though image‐based sketching has been well studied, video‐based sketch animation generation is still very challenging due to the temporal coherence requirement. In this paper, we propose a novel end‐to‐end automatic generation approach for vector sketch animation. To solve the flickering issue, we introduce a Differentiable Motion Trajectory (DMT) representation that describes the frame‐wise movement of stroke control points using differentiable polynomial‐based trajectories. DMT enables global semantic gradient propagation across multiple frames, significantly improving the semantic consistency and temporal coherence, and producing high‐framerate output. DMT employs a Bernstein basis to balance the sensitivity of polynomial parameters, thus achieving more stable optimization. Instead of implicit fields, we introduce sparse track points for explicit spatial modeling, which improves efficiency and supports long‐duration video processing. Evaluations on DAVIS and LVOS datasets demonstrate the superiority of our approach over SOTA methods. Cross‐domain validation on 3D models and text‐to‐video data confirms the robustness and compatibility of our approach. Xinding Zhu, Shuyang Zheng, Zhexin Zhang, Fei Gao 0014, Jiazhou Chen 0002 |
Comput. Graph. Forum | 5 |
| 2026 | Spatial-temporal domain generalization for cross-city traffic prediction
Shengzhe You, Libo Weng, Yanjing Lei, Fei Gao 0014 |
Expert Syst. Appl. | 4 |
| 2025 | An Automatic Extrinsic Calibration Method for LiDAR-Camera Fusion via Combining Semantic and Geometric FeaturesabstractPrecise extrinsic calibration is one of the key techniques for LiDAR-camera fusion system. In current methods, the extrinsic calibration is usually not automatic. To address this, an automatic calibration method via combining semantic and geometric features is proposed, which is not dependent on any specific calibration object. First, extrinsics are automatically initialized; semantic objects are utilized to formulate the edge constraints and projection boundary constraints. Then, an efficient global optimization algorithm that synergizes the Jacobian matrix and the stochastic strategy of simulated annealing is put forward to calculate precise extrinsics. A feedback mechanism is designed to evaluate the reliability of the proposed method. Experiments on the KITTI dataset show that the proposed method achieves a rotation error of 0.14°and a translation error of 4.5cm, outperforming most current methods. Besides, the experiment on the proposed optimization algorithm is also conducted to verify its effectiveness and efficiency. Minqian Wang, Libo Weng, Fei Gao 0014 |
ICASSP | 3 |
| 2025 | SGAD: An Unsupervised Secondary-Guided Diffusion Model for Industrial Anomaly DetectionabstractReconstruction-based anomaly detection methods often struggle with invariant reconstruction of abnormal regions and the unintended reconstruction of novel anomalies. To address these limitations, this study proposes a novel guided training and reconstruction framework (SGAD) to enhance anomaly reconstruction quality. The proposed approach integrates a training paradigm built on target images and fusion loss, along with target-guided and secondary reconstruction strategies utilizing a diffusion model, achieving superior anomaly detection performance. Additionally, a new DTY anomaly detection dataset is introduced to benchmark the approach. Extensive experiments were conducted on the DTY and MVTec datasets, demonstrating that SGAD achieves state-of-the-art performance, with mean scores of 93.7% I-AUROC and 86.2% P-AUROC. These results highlight the effectiveness and robustness of SGAD in addressing complex anomaly detection challenges, underscoring its potential for deployment in practical production environments. Wenze Kang, Libo Weng, Zhenbo Cheng, Fei Gao 0014 |
ICME | 5 |
| 2025 | EPNet: Efficient Part Segmentation for Dense Point CloudsabstractThe segmentation of dense point clouds from industrial LiDAR scans presents challenges in computational overhead and VRAM usage, hindering the development of automated fast measurement systems. To address this, we propose EPNet, an efficient model for part segmentation of dense point clouds. EPNet employs a U-Net-like architecture with skip connections to merge original and recovered features, enhancing local feature extraction via KNN and cosine similarity. Factorization-dimensionality-reduction module based on self-attention overcomes the limitations of trilinear interpolation in feature recovery, improving both local and global feature fusion. In experiments on the LVPC dataset of dense vehicle point clouds, EPNet outperforms models from the past three years, achieving a 1.7% accuracy improvement and a 9.7% increase in average Instance IoU compared to PointNet++. EPNet also achieves a single-file inference time of under 1 second while requiring minimal GPU VRAM resources, demonstrating its potential for real-world industrial high-precision fast automated measurements. The code is available at https://github.com/duskNNNN/EPNet. Wulong Hu, Minqian Wang, Zhenbo Cheng, Fei Gao 0014 |
ICMR | 6 |
| 2025 | ViTraj: Learning Dual-Side Representations for Vehicle-Infrastructure Cooperative Trajectory PredictionabstractWhile autonomous driving has made substantial progress, accurately predicting the trajectories of surrounding traffic agents remains a fundamental challenge for ensuring safety. Integrating both infrastructure-side and vehicle-side information has the potential to enhance perception and prediction capabilities. However, existing methods overlook the challenges in Vehicle-Infrastructure Cooperative Trajectory Prediction. To bridge this gap, we propose ViTraj, a model-agnostic framework for VIC-TP that leverages infrastructure-side trajectories to mitigate the inherent limitations of vehicle-side forecasting. ViTraj introduces a Feature-Side Selection and a Cooperative Interaction to aggregate complementary features from both sides, effectively expanding the perceptual horizon of prediction models. In addition, we present a Vehicle-Infrastructure Knowledge Distillation strategy to enforce consistency between multi-side predictions, which efficient global-local feature alignment through a single backward pass. Extensive experiments on large-scale public datasets demonstrate that ViTraj consistently improves advanced trajectory prediction models, achieving the state-of-the-art performance compared to existing vehicle-infrastructure cooperative methods. We believe this work provides a promising step toward the practical deployment of V2X-based autonomous driving systems. Shengzhe You, Libo Weng, Fei Gao 0014 |
ACM Multimedia | 3 |
| 2025 | Incremental few-shot instance segmentation without fine-tuning on novel classes
Luofeng Zhang, Libo Weng, Fei Gao 0014 |
Comput. Vis. Image Underst. | 4 |
| 2025 | Learning Combinatorial Prompts for Universal Controllable Image CaptioningabstractAbstract Controllable Image Captioning (CIC)—generating natural language descriptions about images under the guidance of given control signals—is one of the most promising directions toward next-generation captioning systems. Till now, various kinds of control signals for CIC have been proposed, ranging from content-related control to structure-related control. However, due to the format and target gaps of different control signals, all existing CIC works (or architectures) only focus on one certain control signal, and overlook the human-like combinatorial ability. By “combinatorial", we mean that our humans can easily meet multiple needs (or constraints) simultaneously when generating descriptions. To this end, we propose a novel prompt-based framework for CIC by learning Com binatorial Pro mpts, dubbed as ComPro . Specifically, we directly utilize a pretrained language model GPT-2 Radford et al. (OpenAI blog 1:9, 2019) as our language model, which can help to bridge the gap between different signal-specific CIC architectures. Then, we reformulate the CIC as a prompt-guide sentence generation problem, and propose a new lightweight prompt generation network to generate the combinatorial prompts for different kinds of control signals. For different control signals, we further design a new mask attention mechanism to realize the prompt-based CIC. Due to its simplicity, our ComPro can be further extended to more kinds of combined control signals by concatenating these prompts. Extensive experiments on two prevalent CIC benchmarks have verified the effectiveness and efficiency of our ComPro on both single and combined control signals. Zhen Wang 0004, Jun Xiao 0001, Yueting Zhuang, Fei Gao 0014, Jian Shao 0001, Long Chen 0016 |
Int. J. Comput. Vis. | 4 |
| 2025 | Highly Condensed All-MLP Architecture for Long-Term Human Motion PredictionabstractIn artificial intelligence (AI) scenarios where computational resources are constrained, such as in autonomous driving systems, it is challenging to construct a lightweight model that can accurately predict human motion overextended duration. To tackle this challenge, we introduce a highly condensed all-multilayer perceptron (HCMLP) architecture that is engineered for supreme lightweight efficiency. This design facilitates extended-range motion predictions while maintaining uncompromised performance. First, the spatiotemporal dynamic perception (STDP) block enhances operational efficiency while maintaining a simple structure. In STDP, the distinct but parallel spatial multilayer perceptron (SMLP) and temporal multilayer perceptron (TMLP) simultaneously capture the spatial correlations between pose joints and the temporal dynamics of each joint. The subsequent dynamic aggregation (DA), coupled with the channel multilayer perceptron (CMLP), dynamically consolidates and refines spatial and temporal features, leading to improved predictive accuracy. Second, the multiterm union prediction (MTUP) block directly delivers precise predictions for periods ranging from 0 to 4000 ms, eliminating the need for repetitive short-term (ST) prediction iterations. Our experimental results on the Human3.6M, AMASS, 3DPW, and CMU-Mocap datasets demonstrate that HCMLP outperforms existing state-of-the-art (SOTA) methods in ST prediction, long-term (LT) prediction, and especially in extended and extra extended LT (ELT) predictions, all while utilizing the fewest parameters. Sheng Liu 0002, Shaobo Zhang 0005, Fei Gao 0014, Yuan Feng 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | FDNet: A Novel Multivariate Time Series Classification Model Through Fusing Feature and DifferenceabstractThe classification problem of multivariate time series has been widely used in many fields, but current methods still cannot achieve high accuracy. In this paper, a novel multivariate time series classification model named FDNet (Feature and Difference encoding fused Network) is proposed. A structure that comprehensively considers global feature extraction and local multi-scale feature extraction is used in the feature encoding, which can better represent the closeness between the sequence and true label. In the difference encoding, distance difference and shape difference are combined to further enhance the FDNet’s discriminative power. Also, a representative sequence selection algorithm is investigated to speed up the calculation of difference encoding, as well as a k-fold filtering algorithm based on density masking to enhance the accuracy of shape difference. In the comparative experiments on 23 public datasets, FDNet outperforms other baseline methods and state-of-the-art methods in terms of average ranking, average accuracy, and the number of wins/ties, which verifies FDNet has high accuracy and good generalization. Fei Gao 0014, Luofeng Zhang |
ICASSP | 1 |
| 2024 | Weakly Supervised Few-Shot Segmentation Through Textual PromptabstractRecently, significant progress has been made in few-shot segmentation (FSS), which aims to segment unknown objects with only a few support images. However, during both training and testing, FSS still requires pixel-level annotations. When only image-level labels are available, FSS will become a more challenging task, namely weakly supervised few-shot segmentation (WS-FSS). To address this problem, this paper proposes a novel text-driven approach, which replaces pixel-level labels with textual prompts. To guide the model in selecting the target features and capturing the inter-class correlations, a Text-Image Matching Module (TIMM) and a Text Supervision Scheme (TSS) are designed for the feature matching and decoding stages, respectively. Extensive experiments are conducted on two public datasets, PASCAL-5iand COCO-20i. The experimental results demonstrate that our method not only outperforms existing state-of-the-art WS-FSS methods but also achieves comparable or even superior performance to advanced FSS models. The code can be available at https: //github.com/Joseph-Lee-V/Text-WS-FSS. Shengzhe You, Libo Weng, Fei Gao 0014 |
ICASSP | 3 |
| 2024 | Dynamic Mutual-Activated Transformer for Human Motion PredictionabstractAccurate human motion prediction is vital for diverse artificial intelligence applications, and recent research has yielded substantial advancements. Despite this, the prediction process often encounters abrupt discontinuities and accumulates errors over the long term due to insufficient modeling of spatial and temporal correlations, which significantly impacts predictive accuracy. To tackle these challenges, we introduce the Dynamic Mutual-Activated Transformer (DyMAT). This innovative approach learns spatial correlation among joints in pose and temporal correlation of each joint. It is achieved through separate yet concurrent Pose-wise Spatial Attention (PSA) and Joint-specific Temporal Attention (JTA). The dynamic mutual-activation block (DMA) adeptly combines spatio-temporal features, significantly enhancing DyMAT’s representational capacity. Moreover, we integrate a Temporal Self-Enhancement (TSE) block with JTA, serving as a supplement for refining temporal correlation learning. Our experiments conducted on Human3.6M and CMU Mocap datasets underscore that DyMAT consistently outperforms state-of-the-art methods in terms of prediction accuracy. Code is available at https://github.com/alanzhangv123/DyMAT. Shaobo Zhang 0005, Sheng Liu 0002, Fei Gao 0014, Yuan Feng 0002 |
ICASSP | 3 |
| 2024 | BFIDet: A YOLOv7-improved Vehicle and Pedestrian Detector via Balancing Feature IntegrationabstractAccurate vehicle and pedestrian detection are fundamental for safe driving and maintenance of traffic order. In this paper, a YOLOv7-improved vehicle and pedestrian detector via balancing feature integration (BFIDet) is proposed. First, EFFM module is designed to facilitate feature map fusion across layers. Second, GSRFConv is utilized to expand the receptive field of the intrinsic feature map as a way to improve the feature discriminability and robustness. VFBM module is then introduced to guide the propagation of the information flow as a way to solve the problem of dilution of features in non-adjacent layers and semantic differences between cross-scale features. In the experiments, the proposed method achieves 93.9% and 69.4% [email protected] and [email protected]:0.95 metric on the KITTI dataset, which are 2.1% and 1.5% better than YOLOv7, respectively, and the [email protected] metric on the SODA10M dataset reaches 63.1% with an improvement of 0.9% and 1.9% over YOLOv7 and YOLOv8m, respectively. The experimental results demonstrate that the proposed BFIDet is more accuracy than that of other mainstream models with controllable computational consumption. Anrui Wang, Libo Weng, Fei Gao 0014 |
ICMR | 3 |
| 2024 | FP3Seg: Point Cloud Panoptic Segmentation via LiDAR-Camera Fusion and Progressive DecoderabstractPoint cloud panoptic segmentation is a 3D scene perception task that provides a holistic solution for both semantic and instance segmentation. The sparsity and lack of texture features in LiDAR point cloud, coupled with the relatively narrow field of view of camera, make multi-modal fusion challenging. In this paper, we propose a novel multi-modal fusion based point cloud panoptic segmentation method, named FP3Seg, with main contributions including Hybrid Domain Adaptive Fusion (HDAF) module and Progressive Decoder. HDAF employs learnable weights to adaptively fuse multi-modal features in both the channel and spatial domains. Through knowledge distillation, FP3Seg extends the benefits of multi-modal fusion beyond the camera field of view. Pro-gressive Decoder embeds semantic and instance information into the input of panoptic decoder, assisting the decoder in understanding the distinctions between stuff and thing classes. The proposed method is benchmarked on the SemanticKITTI test set, achieving 57.7% PQ, showing a 1.7% improvement over baseline. Experimental results demonstrate that FP3Seg possesses advantages over single-modal approaches in multiple aspects, especially for the segmentation of thing classes. Xianyou Dai, Libo Weng, Fei Gao 0014 |
SMC | 3 |
| 2024 | EFFDet: A Crack Detector via Boundary Preservation and Cross-Attention IntegrationabstractThe complexity of scenes and the topology of cracks make road crack detection a challenging task. Compared to other semantic segmentation tasks, this mission places a greater demand on the network's ability to preserve detailed boundary information. To address this, a novel road crack detection network architecture EFFDet is proposed in this paper. Firstly, we redesign the encoding-decoding module based on large-scale convolutional kernels and attention mechanisms to reduce the loss of detailed information caused by downsampling. Secondly, the Cross Attention module is proposed to integrate more precise details into the output of the decoding layer. In comparative experiments on four datasets, CRACK500, Volker, CrackLS315 and DeepCrack, EFFDet achieves ODS values of 0.7434, 0.6758, 0.6449 and 0.8708, respectively. The experimental results show that EFFDet demonstrates stronger detection capabilities in road crack detection. Linhua Gao, Libo Weng, Fei Gao 0014 |
SMC | 3 |
| 2024 | HCMLP: A Highly Condensed All-MLP Architecture for Extended Long-term Human Motion PredictionabstractAccurate human motion prediction has significant potential in various artificial intelligence applications. To accommodate the demands of applications such as autonomous driving on mobile devices, it is essential to utilize models that are both lightweight and capable of performing extended-duration predictions to ensure the system remains swift and reliable. To address these challenges, we present the HCMLP, a highly condensed all-MLP architecture designed for optimal lightweight efficiency, enabling extended long-term predictions without compromising performance. This pioneering method simultaneously captures the spatial correlations between pose joints and the temporal dynamics of each joint by employing distinct but parallel spatial and temporal MLPs. Then, Dynamic Aggregation component dynamically assimilates the spatial and temporal correlations. Finally, channel MLP synergizes and refines these spatio-temporal features for enhanced prediction accuracy. Our experiments on the Human3.6M, AMASS, and 3DPW datasets reveal that HCMLP surpasses the performance of current state-of-the-art methods in short-term, long-term, and particularly extended long-term predictions, while maintaining the least parameters. Code will be available at https://github.com/alanzhangv123/HCMLP. Shaobo Zhang 0005, Sheng Liu 0002, Fei Gao 0014, Yuan Feng 0002 |
SMC | 3 |
| 2024 | SIF-TF: A Scene-Interaction fusion Transformer for trajectory prediction
Fei Gao 0014, Wanjun Huang, Libo Weng |
Knowl. Based Syst. | 1 |
| 2024 | Occluded person re-identification based on feature fusion and sparse reconstruction
Fei Gao 0014, Yisu Ge, Shufang Lu |
Multim. Tools Appl. | 1 |
| 2024 | Anomaly Detection of Tire Tiny Text: Mechanism and MethodabstractDue to the variety of tire specifications and poor contrast, it is difficult for ordinary 2D vision-based methods to accurately and autonomously detect anomalies in the tire text embossed on the tire sidewalls. In this paper, a 3D vision-based autonomous scanning mechanism and method is proposed, including three core components: a three-degree-of-freedom (3-DoF) robotic arm, a high-precision turntable and a composite vision probe. Based on the designed vision probe, the mechanism can autonomously scan the tire sidewall through a series of optimal viewpoints to obtain high-quality tire images even in the absence of tire CAD models. To obtain the optimal viewpoints, an information-driven viewpoint planning method is proposed. Combined with the real-time viewpoint adjustment strategy, the viewpoint of the vision probe is dynamically fine-tuned, which can effectively balance the scanning precision and motion cost. In the text anomaly detection phase, a three-stage tire text anomaly detection method based on deep learning and statistics strategy is proposed, which can accurately detect and recognize tire text, identify sub-mold categories, and judge text anomalies. Experimental results show that the proposed mechanism and method perform well in terms of accuracy ($>$96%) and efficiency (time per lap for a tire is less than 9 s); it shows that the proposed mechanism and method have certain industrial application value and can be applied to tire quality inspection of on-line production. Note to Practitioners—Tire text may be incorrectly embossed during tire manufacturing. Due to the wide variety of tire specifications, poor contrast, and thin text strokes, it is difficult to detect abnormal tire text. In this paper, a 3D vision-based mechanism and method is developed to address this issue. Using a specially designed composite vision probe and scanning method, the mechanism can observe tiny text information on tires of different specifications from multiple angles and directions. By using our tire text anomaly detection method based deep learning and statistics strategy, the mechanism can accurately and robustly detect the region and content of abnormal text. The mechanism is efficient (with a detection time of less than 9 s per lap for one tire) and accurate ($>$96%), indicating that it has great potential in the first inspection or quality inspection step of tire manufacturing in industrial scenarios. Ailing Cheng, Shufang Lu, Fei Gao 0014 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2024 | Knowledge-Guided Causal Intervention for Weakly-Supervised Object LocalizationabstractPrevious weakly-supervised object localization (WSOL) methods aim to expand activation map discriminative areas to cover the whole objects, yet neglect two inherent challenges when relying solely on image-level labels. First, the “entangled context” issue arises from object-context co-occurrence (e.g., fish and water), making the model inspection hard to distinguish object boundaries clearly. Second, the “C-L dilemma” issue results from the information decay caused by the pooling layers, which struggle to retain both the semantic information for precise classification and those essential details for accurate localization, leading to a trade-off in performance. In this paper, we propose a knowledge-guided causal intervention method, dubbed KG-CI-CAM, to address these two under-explored issues in one go. More specifically, we tackle the co-occurrence context confounder problem via causal intervention, which explores the causalities among image features, contexts, and categories to eliminate the biased object-context entanglement in the class activation maps. Based on the disentangled object feature, we introduce a multi-source knowledge guidance framework to strike a balance between absorbing classification knowledge and localization knowledge during model training. Extensive experiments conducted on several benchmark datasets demonstrate the effectiveness of KG-CI-CAM in learning distinct object boundaries amidst confounding contexts and mitigating the dilemma between classification and localization performance. Feifei Shao, Yawei Luo, Fei Gao 0014, Yi Yang 0001, Jun Xiao 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2024 | Gaussian-based adaptive frame skipping for visual object tracking
Fei Gao 0014, Shengzhe You, Yisu Ge |
Vis. Comput. | 1 |
| 2023 | ECDet: A Real-Time Vehicle Detection Network for CPU-Only Devices
Fei Gao 0014, Jianwen Shao, Xinyang Dong, Libo Weng |
ICANN (7) | 1 |
| 2023 | Language Guided Graph Transformer for Skeleton Action Recognition
Libo Weng, Weidong Lou, Fei Gao 0014 |
ICONIP (10) | 3 |
| 2023 | Triple Correlations-Guided Label Supplementation for Unbiased Video Scene Graph GenerationabstractVideo-based scene graph generation (VidSGG) is an approach that aims to represent video content in a dynamic graph by identifying visual entities and their relationships. Due to the inherently biased distribution and missing annotations in the training data, current VidSGG methods have been found to perform poorly on less-represented predicates. In this paper, we propose an explicit solution to address this under-explored issue by supplementing missing predicates that should be included in the ground-truth annotations. Dubbed Trico, our method seeks to supplement the missing predicates that are supposed to appear in the ground-truth annotations, by exploring three complementary spatio-temporal correlations. Guided by these correlations, the missing labels can be effectively supplemented thus achieving an unbiased predicate predictions. We validate the effectiveness of Trico on the most widely used VidSGG datasets, i.e., VidVRD and VidOR. Extensive experiments demonstrate the state-of-the-art performance achieved by Trico, particularly on those tail predicates. The code is available in the supplementary material. Kaifeng Gao, Yawei Luo, Tao Jiang 0042, Fei Gao 0014, Jian Shao 0001, Jun Xiao 0001 |
ACM Multimedia | 5 |
| 2023 | FAFVTC: A Real-Time Network for Vehicle Tracking and Counting
Fei Gao 0014 |
PRCV (12) | 3 |
| 2023 | An Automatic Fabric Defect Detector Using an Efficient Multi-scale Network
Fei Gao 0014, Xiaolu Cao, Yaozhong Zhuang |
PRICAI (3) | 1 |
| 2023 | Traffic Sign Recognition Model Based on Small Object Detection
Fei Gao 0014, Wanjun Huang, Xiuqi Chen, Libo Weng |
PRICAI (3) | 1 |
| 2023 | SEAT: A Spatiotemporal Encode-Again Transformer for Traffic PredictionabstractCurrently, many networks like recurrent neural networks and graph convolution networks are paying more attention to traffic prediction. However, there are still some limitations like lack of consideration of the dynamics between spatial and temporal features, loss of short-term to long-term prediction correlation, and dimensional information destroyed by self-attention. To address these issues, a novel transformer, i.e., Spatiotemporal Encode-Again Transformer (SEAT), is proposed for traffic prediction. In the SEAT, two components, spatial-temporal cross attention, and encode-again strategy are designed to learn spatiotemporal features and capture the relationship among forecasting series. We conducted experiments on several public datasets, METR-LA, PeMS-Bay, and PeMS-S. In particular, SEAT outperforms existing models by up to 6% improvement in RMSE measurement. The experimental results verify that SEAT can better learn the spatiotemporal features and can help lead to more efficient traffic control and management, Shengzhe You, Jianwen Shao, Fei Gao 0014 |
SMC | 4 |
| 2023 | Whether and how is a surveillance camera jittering? A ROR perception based framework and method
Fei Gao 0014, Kaitao Mei, Libo Weng, Yaozhong Zhuang |
Appl. Intell. | 1 |
| 2023 | A 3D graph convolutional networks model for 2D skeleton-based human action recognitionabstractAbstract With the popularity of cameras, the application of action recognition is more and more extensive. After the emergence of RGB‐D cameras and human pose estimation algorithms, human actions can be represented by a sequence of skeleton joints. Therefore, skeleton‐based action recognition has been a research hotspot. In this paper, a novel 3D Graph Convolutional Network model (3D‐GCN) with space‐time attention mechanism for 2D skeleton data is proposed. Three‐dimensional graph convolution is employed to extract spatiotemporal features of skeleton descriptor that is composed of joint coordinates, frame differences and angles. Meanwhile, different joints and different frames are given different attention to achieve action classification. A zebra crossing pedestrian dataset named ZCP is also provided, which simulates possible pedestrian actions on the zebra crossing in real scenes. Experimental evaluation is carried out on ZCP dataset and NTU RGB+D dataset. Experimental results show that our method is better than current 2D‐based methods and is comparable with 3D methods. Libo Weng, Weidong Lou, Fei Gao 0014 |
IET Image Process. | 4 |
| 2023 | A semantic-aware monocular projection model for accurate pose measurement
Libo Weng, Xiuqi Chen, Qi Qiu, Yaozhong Zhuang, Fei Gao 0014 |
Pattern Anal. Appl. | 5 |
| 2022 | Mirror invariant convolutional neural networks for image classificationabstractAbstract Deep convolutional neural networks (DCNNs) have been developed rapidly and they perform well on both classification and object detection tasks. However, its strong performance makes people ignore to study the invariance of the DCNNs, such as mirror invariance. In fact, the ability of DCNNs in handling mirror‐symmetrical images remains limited. In this paper, a mirror transformation convolutional layer is proposed, which transforms several feature maps to produce mirror‐symmetrical feature maps based on the traditional convolutional layer. By combining with the mirror transformation convolutional layer, the DCNNs will have mirror invariance and the performance of neural networks can be improved on classification tasks. A dataset for driver's and passenger's seatbelt detection has been collected, which is used to verify the effectiveness of the proposed convolutional layer. In the experiments, one of the state‐of‐the‐art DCNNs, GoogLeNet, is collaborated with the mirror transformation convolutional layer to form a mirror invariant networks (MINets). The experimental results show that the MINets can achieve better classification performance than the original GoogLeNet. MINets can also reduce the risk of over‐fitting caused by applying data augmentations to the dataset. Shufang Lu, Minqian Wang, Fei Gao 0014 |
IET Image Process. | 4 |
| 2022 | Traffic Scene Perception Based on Joint Object Detection and Semantic Segmentation
Libo Weng, Fei Gao 0014 |
Neural Process. Lett. | 3 |
| 2022 | A Trajectory Evaluator by Sub-tracks for Detecting VOT-based Anomalous TrajectoryabstractWith the popularization of visual object tracking (VOT), more and more trajectory data are obtained and have begun to gain widespread attention in the fields of mobile robots, intelligent video surveillance, and the like. How to clean the anomalous trajectories hidden in the massive data has become one of the research hotspots. Anomalous trajectories should be detected and cleaned before the trajectory data can be effectively used. In this article, a Trajectory Evaluator by Sub-tracks (TES) for detecting VOT-based anomalous trajectory is proposed. Feature of Anomalousness is defined and described as the Eigenvector of classifier to filter Track Lets anomalous trajectory and IDentity Switch anomalous trajectory, which includes Feature of Anomalous Pose and Feature of Anomalous Sub-tracks (FAS). In the comparative experiments, TES achieves better results on different scenes than state-of-the-art methods. Moreover, FAS makes better performance than point flow, least square method fitting and Chebyshev Polynomial Fitting. It is verified that TES is more accurate and effective and is conducive to the sub-tracks trajectory data analysis. Fei Gao 0014, Jiada Li, Yisu Ge, Jianwen Shao, Shufang Lu, Libo Weng |
ACM Trans. Knowl. Discov. Data | 1 |
| 2021 | How frontal is a face? Quantitative estimation of face pose based on CNN and geometric projection
Fei Gao 0014, Shufang Lu |
Neural Comput. Appl. | 1 |
| 2021 | Outfit compatibility prediction with multi-layered feature fusion network
Shufang Lu, Xianmei Wan, Fei Gao 0014 |
Pattern Recognit. Lett. | 5 |
| 2021 | Extracting Moving Objects More Accurately: A CDA Contour OptimizerabstractIn the area of change detection, there were a rare number of optimization methods. Most of the optimization methods that are used by change detection are morphological transformation or median filtering, which cannot best optimize change detection algorithm. In this paper, a general post-processing algorithm for change detection is proposed. We believe that some problems cannot be avoided in the area of change detection such as 1) region of moving object generated by change detection is slightly larger than the ground-truth and 2) there are always some disjoint and small regions that are independent from the moving objects. To address the problem, our method can optimize the change detection algorithm bases on the idea of edge detection, which can remove the wrong edge or pixel. In the experiments, more than 20 change detection algorithms that include the best algorithm inChangeDetection.netare selected. Most of these change detection algorithms are optimized by the proposed method on PWC, Precision, and FMeasure, where, our optimized algorithm named FgSegNet_v2 is better than all other algorithms in the CDnet. The best-optimized margin of PWC is 0.64, and the fast speed is 548FPS on CPU. Our approach can better resolve the afore-mentioned problems that cannot be avoided and is general and fast. The experiments can be reproduced with C++ on Githubhttps://github.com/walty19950301/CDA-contour-optimizer. Fei Gao 0014, Yunyang Li, Shufang Lu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Property-based shadow detection and removal method for licence plate imageabstractShadow detection and removal is a classic research in the field of image processing. This work aims to address the problem of shadow detection and removal for licence plate. The shadow detection and removal method based on inherent properties of licence plate is presented in this study. First, an S–V‐channel‐property‐based shadow detection algorithm is investigated to extract the background‐shaded area of licence plate. Afterwards, a shadow edge location algorithm is proposed to detect the complete shadow edge. Then, the shadow is removed on H, S, and V channels with different strategies, respectively. Also, a filter‐based median filtering algorithm is employed to remove the pseudo‐edge. Finally, the proposed method is verified through comparing with other generic shadow detection and removal methods quantitatively and qualitatively. Fei Gao 0014, Yunjing Xu, Yisu Ge, Shufang Lu |
IET Image Process. | 1 |
| 2019 | Dense Receptive Field Network: A Backbone Network for Object Detection
Fei Gao 0014, Chengguang Yang, Yisu Ge, Shufang Lu, Qike Shao |
ICANN (3) | 1 |
| 2019 | Extracting closed object contour in the image: remove, connect and fit
Fei Gao 0014, Shufang Lu |
Pattern Anal. Appl. | 1 |
| 2019 | Depth-aware image vectorization and editing
Shufang Lu, Wei Jiang 0034, Craig S. Kaplan, Xiaogang Jin 0001, Fei Gao 0014, Jiazhou Chen 0002 |
Vis. Comput. | 6 |
| 2018 | A Two-stage Vehicle Type Recognition MethodabstractVehicle type recognition is a common question in modern intelligent transportation system. Although many methods have been proposed in literatures, how to recognize the vehicle type in the case of small samples is still troubling. To solve the problem mentioned above, this paper presents a method for vehicle type recognition based on a two-stage strategy that includes data preprocessing phrase and training phrase. The main purpose of the first stage is to remove the background of the vehicle image to reduce the interference of redundant information on subsequent stage. In second stage, a concept of multi-scale fusion feature which integrates handcrafted features with learning-based feature is to put forward to describe the characteristics of vehicles. The combined features are input to Support Vector Machine with Racial Basis Function (RBF-SVM) to train the recognition model. The proposed method is tested on MVVTR and the recognition accuracy is up to 90.96%, which is better than other methods. It is verified that strong generalization ability can be achieved in the case of small samples by using the two- stage strategy. Fei Gao 0014, Zhijing He, Yisu Ge, Shufang Lu |
IJCNN | 1 |
| 2016 | A Pose-Driven Physically-Based Interactive System Using KinectabstractIn this paper we propose a novel interactive framework based on Microsoft Kinect. The edges silhouette of the user's rigid body is extracted from the depth image captured by the depth sensor. At the same time, the skeleton and the positions of 15 joints of the body are also tracked which are used for defining interactive rules for the system. Then, the edges silhouette, interactive rules, and particle objects are involved to produce vivid results by a physical simulation engine. This system can achieve real-time feedback and believable visual results. The efficiency and effectiveness of our system are demonstrated via various results including supporting multi-player. Shufang Lu, Taoran Xu, Fei Gao 0014 |
CW | 3 |
| 2010 | Principal axis and crease detection for slap fingerprint segmentationabstractIn slap fingerprint segmentation, crease is the most difficult edge to correctly detect. In this paper, we present a novel yet simple and accurate algorithm for the principal axis and crease detection. Firstly, the principal axis of each foreground region is detected using the minimal rotational inertia; Secondly, the crease detection is done based on cost function minimization. This algorithm has been incorporated in a slap fingerprint segmentation scheme, previously developed by the authors, producing successful results. Yong-Liang Zhang, Yan-Miao Li, Gang Xiao 0001, Fei Gao 0014 |
ICIP | 6 |
| 2009 | Semi-similarity design of motorcycle-hydraulic-disk brake: strategy and applicationabstractSemi-similarity design is a reasonable and effective approach for motorcycle-hydraulic-disk brake design, which is an indeterminate problem because of incomplete required parameters. First, equation of semi-similar level scale is set up through analyzing the design model of motorcycle-hydraulic-disk brake. Second, algorithm of semi-similarity design and a fuzzy evaluation model for design scheme are proposed. Third, work flowchart of semi-similarity design based on the above strategy is given and a software system of motorcycle-hydraulic-disk brake semi-similarity design is developed. Finally, the proposed strategy and method are effectively demonstrated in an instance of motorcycle-hydraulic-disk brake semi-similarity design. Fei Gao 0014, Gang Xiao 0001 |
CAD/Graphics | 1 |
| 2008 | Product interface reengineering using fuzzy clustering
Fei Gao 0014, Gang Xiao 0001, Jiu-jun Chen |
Comput. Aided Des. | 1 |