Gang Li 0005

dblp:62/2655-5 · DBLP profile ↗
← Back
42ranked-venue papers
9as first author
40since 2021 · last 2026
0000-0002-7896-4833ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 8 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Normality-enhanced knowledge distillation network for unsupervised industrial anomaly detection
Gang Li 0005, Tianjiao Chen, Jin Wan, Mingle Zhou, Delong Han, Min Li 0033
Eng. Appl. Artif. Intell.1
2026 Enhancing mixture-of-experts model with prior knowledge for infrared and visible image fusion in complex degraded environments
Gang Li 0005, Chengrun Jiang, Jin Wan, Mingle Zhou, Delong Han
Expert Syst. Appl.1
2026 A Novel Dataset and Lightweight Distillation Baseline for Highlight Transparent Object Detection
Gang Li 0005, Qinghui Chen, Qunshu Zhang, Jin Wan, Maomao Xiong, Cong Bai, Dagang Li 0001, Wenyin Zhang, Jinglin Zhang 0004, Shengyong Chen
Int. J. Comput. Vis.2
2026 Unification of Closed-Open Industrial Detection Scenarios: New Large-Scale Benchmarks, Challenges and Baselines
abstract
Large-scale Visual-Language Models (LVLMs) have achieved remarkable success in natural visual tasks, yet their application to industrial defect detection remains challenging due to two fundamental limitations: (i) the scarcity of large-scale industrial datasets that cover diverse defect categories across multiple domains, and (ii) the reliance on manual prompts (points, boxes, masks) that introduce subjective noise and lack text-visual interaction for fine-grained understanding. To address these challenges, we introduce a Large-Scale Multi-Modal Industrial Open-Closed benchmark (MMIOC-1 M) containing over one million samples across 14 super-categories, 29 industrial scenes, and 351 defect subcategories. To our knowledge, MMIOC-1 M is the first unified largest benchmark supporting both open-vocabulary and closed-set industrial detection, providing valuable pre-training data for LVLMs in industrial scenarios. Furthermore, we propose a Refined Text-Visual Prompt Network (RTVPNet) that incorporates three key innovations: (1) an expert-assisted domain projection mechanism that enables rapid adaptation of general vision models to industrial domains, (2) an energy-based sparse sampling strategy that automatically generates refined visual prompts without manual intervention, and (3) a bidirectional text-visual interaction module that enhances cross-modal semantic alignment and understanding. Extensive experiments demonstrate that RTVPNet achieves state-of-the-art performance on MMIOC-1 M, LVIS, and COCO benchmarks while maintaining computational efficiency.
Jinglin Zhang 0001, Qinghui Chen, Gang Li 0005, Da Chen 0002, Shuainan Jing, Dagang Li 0001, Cong Liu 0012, Cong Bai, Shengyong Chen
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 Boosting Small Object Detection via High-Frequency Feature Oriented Network
abstract
Small Object Detection (SOD) aims to accurately identify and locate small objects in images. However, existing methods usually focus on exploring spatial domain features, neglecting high-frequency features that preserve fine-grained details such as texture and edge information. To overcome this limitation, we propose a High-Frequency Feature-Oriented Network (HFFO-Net). First, we introduce the Channel- wise Frequency Modulation Module (CFMM), which leverages the 2D Discrete Cosine Transform (DCT) to accentuate salient frequency components while mitigating noise interference. Second, we design a High-Frequency Oriented Module (HFOM), which utilizes the Channel Selection Branch (CSB) and Spatial Selection Branch (SSB) to highlight small objects in the channel and spatial region. Third, we introduce a Dual-Query Attention Fusion Mechanism (DQAFM), which reduces the semantic gap between spatial and frequency features and achieves better feature fusion through bidirectional cross-attention. Extensive experiments are implemented, and the corresponding results demonstrate that HFFO-Net excels at detecting small objects.
Min Li 0033, Zhaofei Hao, Gang Li 0005, Jin Wan, Delong Han, Mingle Zhou
IEEE Signal Process. Lett.3
2026 Multimodal Industrial Anomaly Detection via Geometric Prior
abstract
The purpose of multimodal industrial anomaly detection is to detect complex geometric shape defects such as subtle surface deformations and irregular contours that are difficult to detect in 2D-based methods. However, current multimodal industrial anomaly detection lacks the effective use of crucial geometric information like surface normal vectors and 3D shape topology, resulting in low detection accuracy. In this paper, we propose a novel Geometric Prior-based Anomaly Detection network (GPAD). Firstly, we propose a point cloud expert model to perform fine-grained geometric feature extraction, employing differential normal vector computation to enhance the geometric details of the extracted features and generate geometric prior. Secondly, we propose a two-stage fusion strategy to efficiently leverage the complementarity of multimodal data as well as the geometric prior inherent in 3D points. We further propose attention fusion and anomaly regions segmentation based on geometric prior, which enhance the model’s ability to perceive geometric defects. Extensive experiments show that our multimodal industrial anomaly detection model outperforms the State-of-the-art (SOTA) methods in detection accuracy on both MVTec-3D AD and Eyecandies datasets.
Min Li 0033, Gang Li 0005, Jin Wan, Delong Han
IEEE Trans. Circuits Syst. Video Technol.3
2025 Few-Shot Relation Extraction via Semantically Related Negative Samples
abstract
Few-shot relation extraction aims to identify and classify specific semantic relations between entities from text with a small number of annotated examples. Recent studies have shown that innovative model designs and learning strategies can significantly enhance the model’s generalization ability and performance, even with limited training samples. However, the samples used for training face the dual challenges of scarce annotated data and semantic ambiguity. Given the limitations of the existing SaCon framework in negative sample generation and discrimination efficiency, we propose a dynamic adversarial negative sample enhancement strategy. This strategy introduces adversarial perturbations before encoding the pre-trained language model by constructing a multi-granularity semantic space. Specifically, we first design an entity permutation mechanism to randomly exchange the subject/object entities of sentences in a small batch to generate negative sample clusters with similar semantics but misplaced relations, then, we integrate a multi-view contrastive learning framework to embed adversarial samples into the feature space topology optimization process.To strengthen boundary-sensitive features, an adaptive margin ranking loss function is proposed to dynamically adjust the representation distance constraints of positive and negative samples, forcing the model to capture the deep semantic invariance of relational predicates under limited samples. This method aims to optimize the traditional negative sample random sampling paradigm, actively explore the semantic space through an adversarial generation mechanism, and construct “difficult samples” with minimal semantic deviation through gradient back propagation, thereby improving the model’s ability to parse implicit relational patterns. The experiment result shows that this framework effectively alleviates the risk of overfitting in small sample scenarios by decoupling relational semantics and surface syntactic features, and its dynamic loss design provides mathematical guarantees for orthogonal separation of feature space. Our code is available online at https://github.com/hhy-test/ESCR.
Delong Han, Hongyu Hao, Jin Wan, Gang Li 0005, Min Li 0033, Mingle Zhou
ECAI4
2025 HGCF: Hierarchical Geometry-Color Fusion for Multimodal Industrial Anomaly Detection
abstract
While current multimodal anomaly detection methods predominantly employ intermediate fusion strategies, they often suffer from inadequate cross-modal interaction and irreversible information loss during feature alignment processes. To overcome these limitations, we propose Hierarchical Geometry-Color Fusion (HGCF), a novel framework that establishes deep synergistic relationships between RGB texture features and point cloud geometric representations. Firstly, we propose a bidirectional cross-modal early fusion mechanism that enables complementary information exchange between point cloud and RGB modalities at the input level. Secondly, we introduce a local self-supervised geometric color reconstruction network with group-wise feature alignment, enhancing fine-grained feature extraction through joint color-geometry reconstruction tasks. Finally, we propose a local window spatial-consistent attention fusion, which achieves semantic consistency and spatial consistency by emphasizing local mutation features to improve the detection of subtle anomalies. Extensive experiments show our model achieves 99.1% I-AUROC on MVTec 3D-AD and 91.7% on Eyecandies, both surpassing state-of-the-art methods.
Min Li 0033, Delong Han, Jin Wan, Gang Li 0005
ACM Multimedia6
2025 Exploring Multimodal Prompts For Unsupervised Continuous Anomaly Detection
abstract
Unsupervised Continuous Anomaly Detection (UCAD) is gaining attention for effectively addressing the catastrophic forgetting and heavy computational burden issues in traditional Unsupervised Anomaly Detection (UAD). However, existing UCAD approaches that rely solely on visual information are insufficient to capture the manifold of normality in complex scenes, thereby impeding further gains in anomaly detection accuracy. To overcome this limitation, we propose an unsupervised continual anomaly detection framework grounded in multimodal prompting. Specifically, we introduce a Continual Multimodal Prompt Memory Bank (CMPMB) that progressively distills and retains prototypical normal patterns from both visual and textual domains across consecutive tasks, yielding a richer representation of normality. Furthermore, we devise a Defect-Semantic-Guided Adaptive Fusion Mechanism (DSG-AFM) that integrates an Adaptive Normalization Module (ANM) with a Dynamic Fusion Strategy (DFS) to jointly enhance detection accuracy and adversarial robustness. Benchmark experiments on MVTec AD and VisA datasets show that our approach achieves state-of-the-art (SOTA) performance on image-level AUROC and pixel-level AUPR metrics.
Mingle Zhou, Jin Wan, Gang Li 0005, Min Li 0033
ACM Multimedia4
2025 Unsupervised Dual-Domain Memory Model for Time Series Anomaly Detection
abstract
The continuous advancement of multimedia technology has led to the exponential accumulation of massive time-stamped data. However, accurately identifying anomalies in such data remains a major challenge. Current anomaly detection methods still face serious limitations, including the difficulty in handling complex time series data and the anomaly masking phenomenon caused by overlapping temporal patterns. Existing methods cannot effectively address these challenges. To overcome these limitations, we propose an unsupervised time series anomaly detection algorithm DMemAD based on a dual-domain memory module. Specifically, we design an STD Mamba structure that can effectively extract trend and seasonal components in the series and enhance the connection between elements in each component through bidirectional learning. Second, we design a dual-domain memory module to avoid anomaly masking by independently storing trend and seasonal patterns. Additionally, we propose a residual-based memory update mechanism to enhance the accuracy of memory updates, ensuring that prototype patterns are stored precisely. Extensive experiments on four datasets from different domains show that DMemAD achieves an average F1 score of 96.81%, outperforming 17 baseline methods and establishing state-of-the-art performance.
Mingle Zhou, Xingli Wang, Delong Han, Gang Li 0005
ACM Multimedia5
2025 Multimodal feature cooperative refinement for few-shot anomaly detection
Delong Han, Gang Li 0005, Mingle Zhou, Jin Wan, Min Li 0033
Adv. Eng. Informatics3
2025 Industrial-application-oriented 2D image and 3D object anomaly detection technology: a comprehensive review
Gang Li 0005, Chengrun Jiang, Min Li 0033, Delong Han, Mingle Zhou
Appl. Intell.1
2025 Three-dimensional reconstruction and fracture segmentation based on X-ray and computed tomography paired dataset
abstract
In some orthopedic surgeries, the use of three-dimensional (3D) computed tomography (CT) scanning technology is not feasible due to scene limitations, leaving doctors to rely on two-dimensional (2D) X-ray images for real-time diagnosis. However, X-ray images lack 3D information, making accurate diagnosis challenging. Developing an algorithm to convert 2D X-ray images into 3D CT images, while simultaneously combining high-quality 3D reconstruction with precise fracture segmentation, offers a promising solution to the problem. In this study, we propose a novel artificial intelligence (AI)-driven framework named 3D reconstruction and segment anything model (3DRecSAM). The reconstruction image enhancer (RIE) is designed to achieve high-precision 3D reconstruction and provide high-quality feature initialization for fracture segmentation. Meanwhile, the mamba segment anything model (MSAM), based on the segment anything model (SAM) architecture, is developed for accurate fracture segmentation. We introduce a Kolmogorov–Arnold network (KAN)-based attention fusion module (KAF), which facilitates the joint optimization of the RIE reconstruction network and the MSAM segmentation network. Furthermore, the selective scanning mamba with KAN (SKM) is incorporated to enhance feature extraction for both RIE and MSAM. Mamba efficiently captures long-range dependencies and sequential patterns, while KAN’s learnable activation functions facilitate adaptive feature fusion and non-linear representation. To train and evaluate 3DRecSAM, we introduce the real X-ray and CT paired dataset (XCPData), which is publicly available on GitHub: https://github.com/YuanGao1201/XCPData .
Yuan Gao 0033, Da Chen 0002, Mingle Zhou, Gang Li 0005, Yunbo Gu, Jean-Louis Coatrieux, Yang Chen 0008
Eng. Appl. Artif. Intell.6
2025 MemMambaAD: Memory-augmented state space model for multivariate time series anomaly detection
Gang Li 0005, Mingchao Ge, Jin Wan, Delong Han, Min Li 0033, Mingle Zhou
Eng. Appl. Artif. Intell.1
2025 Scene text image super-resolution with semantic-aware interaction
Mingle Zhou, Jin Wan, Delong Han, Min Li 0033, Gang Li 0005
Eng. Appl. Artif. Intell.6
2025 PFRNet: Progressive multi-scale feature fusion and refinement for RGB-D salient object detection
Zhengqian Feng, Wei Wang 0441, Mingle Zhou, Yuan Gao 0033, Gang Li 0005
Neurocomputing7
2025 Reconsidering learnable fine-grained text prompts for few-shot anomaly detection in visual-language models
Delong Han, Mingle Zhou, Jin Wan, Min Li 0033, Gang Li 0005
Neural Networks6
2025 Multiagent Confrontation Method Based on Three-Party Dynamic Multistrategy Evolutionary Game
abstract
Unmanned agents represent a significant advancement in unmanned control and constitute an important element in the future agent warfare. Their autonomous decision-making capabilities are integral to accomplishing tasks independently. To address challenges inherent in multiparty game scenarios that traditional method struggle with and enhance the applicability and accuracy of game decision-making, this article proposes a novel multiagent confrontation method for unmanned vessels, tailored to a three-party dynamic multistrategy evolutionary game in incomplete information scenarios. The approach introduces a new incentive mechanism designed to enhance both individual and collective profits of agents. Using evolutionary game theory, a three-party model is developed, incorporating interactions among player, enemy, and neutral agents. The model tracks the evolution of strategies to identify stable equilibria across various perceptual conditions. Simulations validate the effectiveness of the proposed method in selecting optimal strategies for unmanned vessels in complex battlefield scenarios, demonstrating its potential for improving autonomous decision-making in multiparty confrontations.
Shilong Jin, Yingjie Wang 0002, Peiyong Duan, Haijing Zhang, Gang Li 0005, Zhipeng Cai 0001
IEEE Trans. Comput. Soc. Syst.5
2025 Determining Task Assignments for Candidate Workers Based on Trajectory Prediction
abstract
With the rise of sensor-equipped mobile devices, Mobile Crowd Sensing (MCS) has emerged as an efficient method for information gathering. In smart city environmental sensing, workers can acquire data by merely being within the sensing area. Currently, most studies select opportunistic workers based on the workers’ prior preferences and ignore the effect of movement trajectories on potential opportunistic workers. This may result in the selected opportunistic workers being less-than-ideal, or even ignoring the failure of some tasks to be accomplished, thus resulting in a waste of resources. Therefore, this paper proposes a Recruitment Framework for judging Opportunistic Workers based on Movement Trajectories (RFOW-MT), a two-phase framework for worker recruitment. In the offline phase, combining the neural network model Long Short-Term Memory (LSTM) and Geohash algorithm, an algorithm to detect the set of candidate opportunistic workers is proposed, solving the problems of location privacy and search efficiency. In the online phase, in order to maximize the task spatial coverage under the task budget constraint, a task allocation algorithm based on geographic location packed grouping is proposed. Finally, RFOW-MT outperforms other methods in terms of task spatial coverage and runtime as verified by experiments on real datasets.
Yahong Li, Yingjie Wang 0002, Gang Li 0005, Xiangrong Tong, Zhipeng Cai 0001
IEEE Trans. Mob. Comput.3
2025 Representation Learning Based on Co-Evolutionary Combined With Probability Distribution Optimization for Precise Defect Location
abstract
Visual defect detection methods based on representation learning play an important role in industrial scenarios. Defect detection technology based on representation learning has made significant progress. However, existing defect detection methods still face three challenges: first, the extreme scarcity of industrial defect samples makes training difficult. Second, due to the characteristics of industrial defects, such as blur and background interference, it is challenging to obtain fuzzy defect separation edges and context information. Third, industrial defects cannot obtain accurate positioning information. This article proposes feature co-evolution interaction architecture (CIA) and glass container defect dataset to address the above challenges. Specifically, the contributions of this article are as follows: first, this article designs a glass container image acquisition system that combines RGB and polarization information to create a glass container defect dataset containing more than 60000 samples to alleviate the sample scarcity problem in industrial scenarios. Subsequently, this article designs the CIA. CIA optimizes the probability distribution of features through the co-evolution of edge and context features, thereby improving detection accuracy in blurred defects and noisy environments. Finally, this article proposes a novel inforced IoU loss (IIoU loss), which can obtain more accurate position information by being aware of the scale changes of the predicted box. Defect detection experiments in three mainstream industrial manufacturing categories (Northeastern University (NEU)-Det, glass containers, wood) show that CIA only uses 22.5 GFLOPs, and mean average precision (mAP) (NEU-Det: 88.74%, glass containers: 95.38%, wood: 68.42%) outperforms state-of-the-art methods.
Jinglin Zhang 0001, Qinghui Chen, Gang Li 0005, Shijiao Ding, Maomao Xiong, Shengyong Chen
IEEE Trans. Neural Networks Learn. Syst.4
2024 PDT: Uav Target Detection Dataset for Pests and Diseases Tree
Mingle Zhou, Delong Han, Zhiyong Qi, Gang Li 0005
ECCV (76)5
2024 A Multi-Granularity Semantic Extraction Method for Text Classification
Min Li 0033, Gang Li 0005, Delong Han
ICIC (13)3
2024 Time Series Anomaly Detection via Temporal Dependencies and Multivariate Correlations Integrating
Gang Li 0005, Mingchao Ge, Mingle Zhou, Jin Wan, Delong Han
ICONIP (3)1
2024 Gradient-YOLO: Exploring the integration of gradient architecture into the YOLO network
abstract
Object detection is an essential task in the field of computer vision. The one-stage object detection model directly completes object detection through a single forward propagation and has fast real-time object detection capabilities, making it widely used—especially the model of the You Only Look Once (YOLO) series. Improving the detection accuracy of the YOLO model has always been a research topic. The gradient architecture cascades feature information extraction and aggregates all features at the end, which can improve the feature fusion ability of the network. Combining the idea of gradient architecture, this paper proposes a YOLO based on gradient architecture: Gradient-YOLO. Specifically, this paper integrates gradient architecture into the backbone, neck, and bottleneck structures in the YOLO network, obtaining gradient backbone, gradient neck, and gradient bottleneck, respectively. By combining gradient architecture, various network parts in Gradient-YOLO can fully aggregate multi-layer features, reduce feature loss, and thus improve detection accuracy. This paper takes the mature YOLOv5 model and the latest YOLOv8 model as the baseline and combines gradient thinking to obtain Gradient-YOLOv5 and GradientYOLOv8. Moreover, conduct experimental testing on the MS COCO dataset. Regarding the [email protected] detection indicator, compared to YOLOv5s, Gradient-YOLOv5s increased by 5.07%. Compared to YOLOv8n, Gradient-YOLOv8n has increased by 4.63%. Therefore, the combination of gradient architecture can improve the detection accuracy of YOLO networks.
Gang Li 0005, Delong Han, Letian Gao, Mingle Zhou
IJCNN1
2024 Color and Feature Space Classification Methods Solve the Problem of Covariate Semantic Information Deviation in Complex Scene Detection Tasks
abstract
The field weed detection task under the perspective of Plant Protection UAVs (PPU) is characterized by multi-scale targets and rich background semantic information, which belongs to the target detection task in complex scenarios. Blindly boosting the network size and overusing the attention mechanism in such tasks are not effective in improving the model accuracy. The core of the problem lies in the fact that the covariate features of the target are shifted during the transfer of the three layers of image semantic information (visual, object and conceptual layers). This phenomenon is particularly evident at the conceptual level, resulting in models that do not accurately understand the target information. Redundant semantic information and irrational attention mechanisms for complex scenarios are the cause of the problem. In this paper, we propose the Color and Feature Space Classification (CFSC) method, which aims to construct a covariate semantic information matrix that spans the visual, object and conceptual layers. Mitigating the problem of offsetting semantic information of covariate features in complex scenarios. The color and feature space classification method is applied to a one-stage detection model, YOLOv8, and experiments are conducted on three complex scene PPU weed datasets. The experimental results show that the CFSC method can improve the model accuracy, loss computation capability and accelerate the gradient descent.
Mingle Zhou, Min Li 0033, Delong Han, Gang Li 0005
IJCNN5
2024 Knowledge Distillation Using Global Fusion and Feature Restoration for Industrial Defect Detectors
abstract
Deep learning technology has been widely applied in industrial quality inspection tasks to improve the accuracy of defect detection, object recognition, and classification. However, in the task of object detection, both feature-based and regression-based traditional knowledge distillation methods impose overly strict constraints on the student model. To address the above issues, this article proposes Global Fusion and Feature Restoration Knowledge Distillation(FRD). FRD integrates contextual information into the channel through a Global Fusion Module (GFM), and further utilizes attention mechanisms to adaptively focus on the distillation region after separating the foreground background. FRD also uses a Mask Feature Restoration (MFR) to mask and restore a portion of student features, improving the learnability of the model. At the same time, FRD adopts a comprehensive approach of multiple loss superposition constraints, rather than simply using MSE losses to imitate the features of teachers. Experiments have shown that FRD can effectively improve model performance in industrial defect detection tasks. On the aluminum surface defect dataset, FRD increased the mAP index of RetinaNet-Res50 from 57.7% to 61.7%. We also confirmed the effectiveness of FRD for general object detection on the Coco dataset. On a randomly selected COCO dataset containing 4000 images, FRD increased the mAP metric of RetinaNet-Res50 from 21.5% to 43.6%.
Zhengqian Feng, Xiyao Yue, Mingle Zhou, Delong Han, Gang Li 0005
SMC6
2024 FGSNet: A Finer-Grained Siamese Network for Industrial Few-Shot Anomaly Detection
abstract
Image anomaly detection usually relies on a large set of training samples and abnormal samples are relatively rare and challenging to obtain in daily industrial scenarios. In the scenario of Few-Shot Anomaly Detection (FSAD), how to better utilize the fine-grained features of the few images is a key issue. To solve these problems, a Finer-Grained Siamese Network (FGSNet) is proposed in this paper. FGSNet considers the setting of Few-Shot Anomaly Detection (FSAD) and consists of two stages. The first stage is responsible for extracting finer-grained features and roughly aligning the image features, while incorporating an efficient Fine-grained Feature Fusion Module$(\mathbf{F}^{3}\mathbf{M})$to enhance feature representation. Meanwhile, we design a Deep-supervised Loss to improve the level of fine-grained information extraction. The second is primarily used to refine and denoise the fused features, followed by detailed alignment operations for subsequent feature distribution modeling. Compared to most existing FASD models that use single-class training and single-model evaluation, FGSNet is more suitable for existing industrial scenarios. Extensive experiments are conducted on the publicly available MVTec AD and MPDD datasets in this paper. The results show that, for 2, 4 and 8-shot cases, FGSNet can increase the average Image-AUROC on MVTec AD by 0.87%, 1.27% and 1.33% respectively, compared to RegAD.
Delong Han, Mingle Zhou, Gang Li 0005, Min Li 0033
SMC4
2024 Anomaly-Free Prior Guided Knowledge Distillation for Industrial Anomaly Detection
abstract
In industrial manufacturing, visual anomaly detection is critical for maintaining product quality by detecting and preventing production anomalies. Anomaly detection methods based on knowledge distillation demonstrate promising performance in addressing the unpredictability and diversity of anomalies. However, they suffer from a lack of effective guidance from anomaly-free priors when handling anomalous features and underutilize multi-scale features during the segmentation scoring stage, yielding suboptimal detection results. To alleviate these issues, we propose an Anomaly-free Prior Guided knowledge distillation (APG) for industrial anomaly detection. Firstly, it filters the abnormal features by training the de-noising target network with knowledge distillation structure. Concurrently, we propose the Prior Perception Propagation Module (P3M), which extracts more efficient anomaly-free features by imposing constraints on anomalous features. Secondly, we propose the Multi-scale Prior Guided Fusion Module (MPGFM) to improve anomaly detection accuracy by utilizing anomaly-free features from the target network as priors to guide the generation and fusion of cross-scale differential features. Finally, the Global Perception Enhancement Module (GPEM) is proposed to construct an anomaly scoring network, leveraging comprehensive scene features to enhance the detection and localization performance of numerous small-target anomalies in industrial manufacturing. Extensive experiments on the MVTecAD and BTAD datasets show that the proposed method demonstrates a consistent and significant outperformance against competing methods.
Gang Li 0005, Tianjiao Chen, Min Li 0033, Delong Han, Mingle Zhou
SMC1
2024 MemADet: A Representative Memory Bank Approach for Industrial Image Anomaly Detection
abstract
In the field of industrial production, anomaly detection is crucial for ensuring product quality and maintaining production efficiency. With the continuous advancement of computer vision technology, it has shown tremendous potential in industrial applications. However, the scarcity of labeled anomaly samples in real-world operating environments poses significant challenges for traditional anomaly detection techniques. To address this, we propose MemADet, a novel anomaly detection model that employs an unsupervised approach and leverages a representative memory bank. It employs a dynamic decision mechanism to control the representativeness of the features stored in the memory bank, and employs a weighted anomaly score calculation mechanism to further enhance the performance of image anomaly detection. Our evaluations indicate that MemADet performs robustly in industrial image anomaly detection across three datasets, with particularly no-table detection accuracy on the MVTec AD dataset. Its efficacy is further validated by competitive results on two additional datasets, highlighting its consistent and effective performance in various settings.
Min Li 0033, Zuobin Ying, Gang Li 0005, Mingle Zhou
SMC4
2024 Revisiting the application of twin connected parallel networks and regression loss functions in industrial defect detection
Zhanzhi Su, Mingle Zhou, Min Li 0033, Delong Han, Gang Li 0005
Adv. Eng. Informatics6
2024 IDP-Net: Industrial defect perception network based on cross-layer semantic information guidance and context concentration enhancement
abstract
Applications in Engineering: In industry, surface defect detection is crucial for improving product quality . However, there are many challenges in industrial inspection scenarios, such as interference from background noise, complex small-target problems, significant variations in target objects, and the problem of finding a balance between inspection speed and accuracy. To address the above problems, this paper proposes an industrial defect-aware network based on cross-layer semantic information guidance and contextual attention enhancement (IDP-Net). Specifically, IDP-Net has four different new features. The contribution of artificial intelligence : Firstly, to solve the industrial surface context and defect similarity problem, this paper proposes a Lightweight Local Global Feature Extraction Network (LLG-Net), unlike other methods, the effective combination of self-attention blocks and convolution blocks ensures gradual integration of global and local features across multiple layers, to improve the detection ability of targets with significant changes in scale, this paper designs a Multiscale Perceptual Feature Aggregation Network (MPA-Net), adequately fuses the shallow fine-grained information and the deep semantic information. Then, to enhance the connection between multi-scale semantic information, an adaptive cross-layer feature fusion module (ACFF) is proposed, which is novel in integrating the characteristics of multiple adjacent levels to help the model better capture the different scale characterisation of the target. Finally, a Region Attention Module (RAM) is proposed and introduced in the detector to enhance the attention to the critical regions around the target object. In particular, this paper proposes a new localisation loss function (MEIoU) that enhances the network’s attention to objects at different scales. The experimental results show that 94.3%, 98.7% and 99.5% of [email protected] are obtained on steel, PCB and aluminium surface defect datasets, respectively, and 50 FPS is achieved, which is better than the current mainstream detectors and meets the demand of practical industrial production.
Gang Li 0005, ShiLong Zhao, Min Li 0033, Mingle Zhou, Zuobin Ying
Eng. Appl. Artif. Intell.1
2024 MFUR-Net: Multimodal feature fusion and unimodal feature refinement for RGB-D salient object detection
Zhengqian Feng, Wei Wang 0441, Gang Li 0005, Min Li 0033, Mingle Zhou
Knowl. Based Syst.4
2024 CSDD-Net: A cross semi-supervised dual-feature distillation network for industrial defect detection
Mingle Zhou, Zhanzhi Su, Min Li 0033, Yingjie Wang 0002, Gang Li 0005
Knowl. Based Syst.5
2023 Transforming Limitations into Advantages: Improving Small Object Detection Accuracy with SC-AttentionIoU Loss Function
Mingle Zhou, Changle Yi, Min Li 0033, Honglin Wan, Gang Li 0005, Delong Han
ICANN (7)5
2023 Exploring Adaptive Regression Loss and Feature Focusing in Industrial Scenarios
Mingle Zhou, Zhanzhi Su, Min Li 0033, Delong Han, Gang Li 0005
ICONIP (6)5
2023 DCP-Net: The Defect Detection Method of Industrial Product based on Dual Collaborative Paths
abstract
With the rapid development of industrial automation, large-scale production of automated pipelines has become a trend. In the production process of industrial products, the quality detection of industrial products is an essential means to avoid economic losses and ensure personal safety. Improving industrial product quality detection accuracy and speed has been an important research topic in this field. In recent years, one-stage and two-stage deep learning object detection methods have been applied in defect detection. However, problems such as significant changes in object scale, high similarity, and the balance of speed and accuracy in industrial detection scenarios result in poor performance of common object detection algorithms. To address these issues, this paper proposes a Glass Bottle Bottom Mold Sequence Recognition Data Set and designs a novel dual collaborative paths feature network (DCP- Net) for industrial defect detection. The DCP-Net is designed to extract object features with varied scales from two branches, which can effectively focus on multi-scale industrial product defect features. Moreover, this paper presents the feature filter to obtain fine-grained information about objects, significantly improving similar objects' classification. In the experiments of the open data set of steel defects and the point sequence data set of glass bottle bottom mold collected in this paper, this method achieves higher accuracy than other methods.
Mingle Zhou, Honglin Wan, Min Li 0033, Gang Li 0005
IJCNN5
2023 A Relational Classification Network Integrating Multi-scale Semantic Features
Gang Li 0005, Jiakai Tian, Mingle Zhou, Min Li 0033, Delong Han
NLPCC (2)1
2023 PCB Defect Detection Model with Convolutional Modules Instead of Self-Attention Mechanism
abstract
The detection of defects in printed circuit boards requires high accuracy and realtime performance. Existing industrial detection models generally adopt a pure convolutional structure for ease of deployment. However, the detection accuracy of these models is often insufficient to meet the requirements of the scene. To improve the detection model accuracy and ease of deployment, this paper proposes a convolutional merging Transformer network(CMTRNet). The CMTR-Net model proposes a backbone network (CNN-Former) that uses convolutional modules to replace self-attention, combining the Transformer architecture with a convolutional structure. This approach not only avoids the drawback of self-attention high computation complexity that is detrimental to deployment but also improves the model detection accuracy. Based on CNN-Former, this paper also proposes a feature fusion module that can better fuse the features extracted by CNN-Former. Further-more, based on the CMTRNet model and the characteristics of circuit board defects, this paper proposes a loss function called Melt-IoU, which can make the initial training phase smoother and further improve detection accuracy. Experiments have shown that CMTRNet outperforms existing advanced models on both datasets.
Mingle Zhou, Changle Yi, Honglin Wan, Min Li 0033, Gang Li 0005, Delong Han
SMC5
2023 IDD-Net: Industrial defect detection method based on Deep-Learning
Mingle Zhou, Honglin Wan, Min Li 0033, Gang Li 0005, Delong Han
Eng. Appl. Artif. Intell.5
2023 ICA-Net: Industrial defect detection network based on convolutional attention guidance and aggregation of multiscale features
ShiLong Zhao, Gang Li 0005, Mingle Zhou, Min Li 0033
Eng. Appl. Artif. Intell.2
2008 Adaptive Image Segmentation Using Modified Pulse Coupled Neural Network
Gang Li 0005, Min Li 0033
ISNN (2)2
2008 A Novel Pixel-Level and Feature-Level Combined Multisensor Image Fusion Scheme
Min Li 0033, Gang Li 0005
ISNN (2)2