Delong Han

dblp:330/4601 · DBLP profile ↗
← Back
31ranked-venue papers
4as first author
31since 2021 · last 2026
0000-0001-7195-3413ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 2 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 UniGRe-3D: Unified Geometric Reconstruction for Multi-category 3D Anomaly Detection
Delong Han, Yuan Gao 0033, Min Li 0033, Mingle Zhou
Comput. Graph. Forum1
2026 Normality-enhanced knowledge distillation network for unsupervised industrial anomaly detection
Gang Li 0005, Tianjiao Chen, Jin Wan, Mingle Zhou, Delong Han, Min Li 0033
Eng. Appl. Artif. Intell.5
2026 Enhancing mixture-of-experts model with prior knowledge for infrared and visible image fusion in complex degraded environments
Gang Li 0005, Chengrun Jiang, Jin Wan, Mingle Zhou, Delong Han
Expert Syst. Appl.6
2026 Boosting Small Object Detection via High-Frequency Feature Oriented Network
abstract
Small Object Detection (SOD) aims to accurately identify and locate small objects in images. However, existing methods usually focus on exploring spatial domain features, neglecting high-frequency features that preserve fine-grained details such as texture and edge information. To overcome this limitation, we propose a High-Frequency Feature-Oriented Network (HFFO-Net). First, we introduce the Channel- wise Frequency Modulation Module (CFMM), which leverages the 2D Discrete Cosine Transform (DCT) to accentuate salient frequency components while mitigating noise interference. Second, we design a High-Frequency Oriented Module (HFOM), which utilizes the Channel Selection Branch (CSB) and Spatial Selection Branch (SSB) to highlight small objects in the channel and spatial region. Third, we introduce a Dual-Query Attention Fusion Mechanism (DQAFM), which reduces the semantic gap between spatial and frequency features and achieves better feature fusion through bidirectional cross-attention. Extensive experiments are implemented, and the corresponding results demonstrate that HFFO-Net excels at detecting small objects.
Min Li 0033, Zhaofei Hao, Gang Li 0005, Jin Wan, Delong Han, Mingle Zhou
IEEE Signal Process. Lett.5
2026 Multimodal Industrial Anomaly Detection via Geometric Prior
abstract
The purpose of multimodal industrial anomaly detection is to detect complex geometric shape defects such as subtle surface deformations and irregular contours that are difficult to detect in 2D-based methods. However, current multimodal industrial anomaly detection lacks the effective use of crucial geometric information like surface normal vectors and 3D shape topology, resulting in low detection accuracy. In this paper, we propose a novel Geometric Prior-based Anomaly Detection network (GPAD). Firstly, we propose a point cloud expert model to perform fine-grained geometric feature extraction, employing differential normal vector computation to enhance the geometric details of the extracted features and generate geometric prior. Secondly, we propose a two-stage fusion strategy to efficiently leverage the complementarity of multimodal data as well as the geometric prior inherent in 3D points. We further propose attention fusion and anomaly regions segmentation based on geometric prior, which enhance the model’s ability to perceive geometric defects. Extensive experiments show that our multimodal industrial anomaly detection model outperforms the State-of-the-art (SOTA) methods in detection accuracy on both MVTec-3D AD and Eyecandies datasets.
Min Li 0033, Gang Li 0005, Jin Wan, Delong Han
IEEE Trans. Circuits Syst. Video Technol.6
2025 Few-Shot Relation Extraction via Semantically Related Negative Samples
abstract
Few-shot relation extraction aims to identify and classify specific semantic relations between entities from text with a small number of annotated examples. Recent studies have shown that innovative model designs and learning strategies can significantly enhance the model’s generalization ability and performance, even with limited training samples. However, the samples used for training face the dual challenges of scarce annotated data and semantic ambiguity. Given the limitations of the existing SaCon framework in negative sample generation and discrimination efficiency, we propose a dynamic adversarial negative sample enhancement strategy. This strategy introduces adversarial perturbations before encoding the pre-trained language model by constructing a multi-granularity semantic space. Specifically, we first design an entity permutation mechanism to randomly exchange the subject/object entities of sentences in a small batch to generate negative sample clusters with similar semantics but misplaced relations, then, we integrate a multi-view contrastive learning framework to embed adversarial samples into the feature space topology optimization process.To strengthen boundary-sensitive features, an adaptive margin ranking loss function is proposed to dynamically adjust the representation distance constraints of positive and negative samples, forcing the model to capture the deep semantic invariance of relational predicates under limited samples. This method aims to optimize the traditional negative sample random sampling paradigm, actively explore the semantic space through an adversarial generation mechanism, and construct “difficult samples” with minimal semantic deviation through gradient back propagation, thereby improving the model’s ability to parse implicit relational patterns. The experiment result shows that this framework effectively alleviates the risk of overfitting in small sample scenarios by decoupling relational semantics and surface syntactic features, and its dynamic loss design provides mathematical guarantees for orthogonal separation of feature space. Our code is available online at https://github.com/hhy-test/ESCR.
Delong Han, Hongyu Hao, Jin Wan, Gang Li 0005, Min Li 0033, Mingle Zhou
ECAI1
2025 STAD: Joint Spatial-Temporal Dimension and Channel Correlation for Time Series Anomaly Detection
abstract
Accurately identifying real anomalies and pseudo-anomalies in complex multi-dimensional time series data has been a difficult problem in time series anomaly detection. To solve this problem, this paper proposes a new framework, STAD, that joint temporal and spatial dimensions. This framework guides the model to capture the correlation information between channels It also aims to learn the deep feature representation of sequences by mining potential information in the spatialtemporal dimension. It can effectively distinguish between true and false anomalies by comparing information from spatial and temporal dimensions. STAD identifies and integrates correlated channels by using a correlation aggregation mechanism to join multiple channels and detect anomalies.In addition, the KAN mixer designed in this paper can effectively extract features from different spatial locations in the spatial-temporal dimension. Through extensive experiments on several public datasets, STAD demonstrates its superiority in terms of accuracy and robustness.
Mingle Zhou, Xingli Wang, Delong Han, Jin Wan
ICASSP3
2025 CSF-DSRE:A Distantly Supervised Relation Extraction Model Based on Contextual Semantic Fusion
abstract
In the Distantly Supervised Relation Extraction (DSRE) task, it is typically assumed that sentences containing the same entity pair reflect the same relation. While this assumption simplifies the learning process to some extent, it inevitably introduces substantial noise. Previous extensive research has utilized a series of selective attention mechanisms on sentences in the bag to extract relation features, but has seldom considered leveraging the information surrounding the entities and different features. We argue that tokens surrounding the entities play a significant role in relation classification, though their importance varies. In this paper, we propose the Dynamic Semantic Interaction (DSI) module, which uses token-level features to calculate the similarity between the entity and other tokens in the sentence as weights, assigning different importance to each token. Additionally, the Responsiveness Feature Fusion Strategy (RFFS) uses Euclidean distance to calculate reliability scores, adaptively adjusts the information fusion process, and weights different features to reduce the impact of noisy information. Finally, the Relation-Aware Contextual Modeling (RACM) module initializes and trains a set of seed vectors to guide the model in focusing on relation-relevant contextual features, thereby improving relation recognition. Combining these methods, our model is able to better predict the relations between entity pairs. Experimental results show that our model demonstrates significant advantages across various mainstream DSRE datasets.
Min Li 0033, Mingle Zhou, Delong Han
IJCNN5
2025 HGCF: Hierarchical Geometry-Color Fusion for Multimodal Industrial Anomaly Detection
abstract
While current multimodal anomaly detection methods predominantly employ intermediate fusion strategies, they often suffer from inadequate cross-modal interaction and irreversible information loss during feature alignment processes. To overcome these limitations, we propose Hierarchical Geometry-Color Fusion (HGCF), a novel framework that establishes deep synergistic relationships between RGB texture features and point cloud geometric representations. Firstly, we propose a bidirectional cross-modal early fusion mechanism that enables complementary information exchange between point cloud and RGB modalities at the input level. Secondly, we introduce a local self-supervised geometric color reconstruction network with group-wise feature alignment, enhancing fine-grained feature extraction through joint color-geometry reconstruction tasks. Finally, we propose a local window spatial-consistent attention fusion, which achieves semantic consistency and spatial consistency by emphasizing local mutation features to improve the detection of subtle anomalies. Extensive experiments show our model achieves 99.1% I-AUROC on MVTec 3D-AD and 91.7% on Eyecandies, both surpassing state-of-the-art methods.
Min Li 0033, Delong Han, Jin Wan, Gang Li 0005
ACM Multimedia4
2025 Unsupervised Dual-Domain Memory Model for Time Series Anomaly Detection
abstract
The continuous advancement of multimedia technology has led to the exponential accumulation of massive time-stamped data. However, accurately identifying anomalies in such data remains a major challenge. Current anomaly detection methods still face serious limitations, including the difficulty in handling complex time series data and the anomaly masking phenomenon caused by overlapping temporal patterns. Existing methods cannot effectively address these challenges. To overcome these limitations, we propose an unsupervised time series anomaly detection algorithm DMemAD based on a dual-domain memory module. Specifically, we design an STD Mamba structure that can effectively extract trend and seasonal components in the series and enhance the connection between elements in each component through bidirectional learning. Second, we design a dual-domain memory module to avoid anomaly masking by independently storing trend and seasonal patterns. Additionally, we propose a residual-based memory update mechanism to enhance the accuracy of memory updates, ensuring that prototype patterns are stored precisely. Extensive experiments on four datasets from different domains show that DMemAD achieves an average F1 score of 96.81%, outperforming 17 baseline methods and establishing state-of-the-art performance.
Mingle Zhou, Xingli Wang, Delong Han, Gang Li 0005
ACM Multimedia4
2025 Multimodal feature cooperative refinement for few-shot anomaly detection
Delong Han, Gang Li 0005, Mingle Zhou, Jin Wan, Min Li 0033
Adv. Eng. Informatics2
2025 Industrial-application-oriented 2D image and 3D object anomaly detection technology: a comprehensive review
Gang Li 0005, Chengrun Jiang, Min Li 0033, Delong Han, Mingle Zhou
Appl. Intell.5
2025 DualPhys-GS: Dual physically-guided 3D Gaussian splatting for underwater scene reconstruction
Guangzhi Han, Jin Wan, Yuan Gao 0033, Delong Han
Comput. Graph.5
2025 MemMambaAD: Memory-augmented state space model for multivariate time series anomaly detection
Gang Li 0005, Mingchao Ge, Jin Wan, Delong Han, Min Li 0033, Mingle Zhou
Eng. Appl. Artif. Intell.4
2025 Scene text image super-resolution with semantic-aware interaction
Mingle Zhou, Jin Wan, Delong Han, Min Li 0033, Gang Li 0005
Eng. Appl. Artif. Intell.4
2025 Reconsidering learnable fine-grained text prompts for few-shot anomaly detection in visual-language models
Delong Han, Mingle Zhou, Jin Wan, Min Li 0033, Gang Li 0005
Neural Networks1
2024 PDT: Uav Target Detection Dataset for Pests and Diseases Tree
Mingle Zhou, Delong Han, Zhiyong Qi, Gang Li 0005
ECCV (76)3
2024 A Multi-Granularity Semantic Extraction Method for Text Classification
Min Li 0033, Gang Li 0005, Delong Han
ICIC (13)4
2024 Structural Optimization and Sequence Interaction Enhancement for Hyper-Relational Knowledge Graphs
Delong Han, Zhengqian Feng, Mingle Zhou
ICIC (13)3
2024 Time Series Anomaly Detection via Temporal Dependencies and Multivariate Correlations Integrating
Gang Li 0005, Mingchao Ge, Mingle Zhou, Jin Wan, Delong Han
ICONIP (3)5
2024 Gradient-YOLO: Exploring the integration of gradient architecture into the YOLO network
abstract
Object detection is an essential task in the field of computer vision. The one-stage object detection model directly completes object detection through a single forward propagation and has fast real-time object detection capabilities, making it widely used—especially the model of the You Only Look Once (YOLO) series. Improving the detection accuracy of the YOLO model has always been a research topic. The gradient architecture cascades feature information extraction and aggregates all features at the end, which can improve the feature fusion ability of the network. Combining the idea of gradient architecture, this paper proposes a YOLO based on gradient architecture: Gradient-YOLO. Specifically, this paper integrates gradient architecture into the backbone, neck, and bottleneck structures in the YOLO network, obtaining gradient backbone, gradient neck, and gradient bottleneck, respectively. By combining gradient architecture, various network parts in Gradient-YOLO can fully aggregate multi-layer features, reduce feature loss, and thus improve detection accuracy. This paper takes the mature YOLOv5 model and the latest YOLOv8 model as the baseline and combines gradient thinking to obtain Gradient-YOLOv5 and GradientYOLOv8. Moreover, conduct experimental testing on the MS COCO dataset. Regarding the [email protected] detection indicator, compared to YOLOv5s, Gradient-YOLOv5s increased by 5.07%. Compared to YOLOv8n, Gradient-YOLOv8n has increased by 4.63%. Therefore, the combination of gradient architecture can improve the detection accuracy of YOLO networks.
Gang Li 0005, Delong Han, Letian Gao, Mingle Zhou
IJCNN3
2024 Color and Feature Space Classification Methods Solve the Problem of Covariate Semantic Information Deviation in Complex Scene Detection Tasks
abstract
The field weed detection task under the perspective of Plant Protection UAVs (PPU) is characterized by multi-scale targets and rich background semantic information, which belongs to the target detection task in complex scenarios. Blindly boosting the network size and overusing the attention mechanism in such tasks are not effective in improving the model accuracy. The core of the problem lies in the fact that the covariate features of the target are shifted during the transfer of the three layers of image semantic information (visual, object and conceptual layers). This phenomenon is particularly evident at the conceptual level, resulting in models that do not accurately understand the target information. Redundant semantic information and irrational attention mechanisms for complex scenarios are the cause of the problem. In this paper, we propose the Color and Feature Space Classification (CFSC) method, which aims to construct a covariate semantic information matrix that spans the visual, object and conceptual layers. Mitigating the problem of offsetting semantic information of covariate features in complex scenarios. The color and feature space classification method is applied to a one-stage detection model, YOLOv8, and experiments are conducted on three complex scene PPU weed datasets. The experimental results show that the CFSC method can improve the model accuracy, loss computation capability and accelerate the gradient descent.
Mingle Zhou, Min Li 0033, Delong Han, Gang Li 0005
IJCNN4
2024 Knowledge Distillation Using Global Fusion and Feature Restoration for Industrial Defect Detectors
abstract
Deep learning technology has been widely applied in industrial quality inspection tasks to improve the accuracy of defect detection, object recognition, and classification. However, in the task of object detection, both feature-based and regression-based traditional knowledge distillation methods impose overly strict constraints on the student model. To address the above issues, this article proposes Global Fusion and Feature Restoration Knowledge Distillation(FRD). FRD integrates contextual information into the channel through a Global Fusion Module (GFM), and further utilizes attention mechanisms to adaptively focus on the distillation region after separating the foreground background. FRD also uses a Mask Feature Restoration (MFR) to mask and restore a portion of student features, improving the learnability of the model. At the same time, FRD adopts a comprehensive approach of multiple loss superposition constraints, rather than simply using MSE losses to imitate the features of teachers. Experiments have shown that FRD can effectively improve model performance in industrial defect detection tasks. On the aluminum surface defect dataset, FRD increased the mAP index of RetinaNet-Res50 from 57.7% to 61.7%. We also confirmed the effectiveness of FRD for general object detection on the Coco dataset. On a randomly selected COCO dataset containing 4000 images, FRD increased the mAP metric of RetinaNet-Res50 from 21.5% to 43.6%.
Zhengqian Feng, Xiyao Yue, Mingle Zhou, Delong Han, Gang Li 0005
SMC5
2024 FGSNet: A Finer-Grained Siamese Network for Industrial Few-Shot Anomaly Detection
abstract
Image anomaly detection usually relies on a large set of training samples and abnormal samples are relatively rare and challenging to obtain in daily industrial scenarios. In the scenario of Few-Shot Anomaly Detection (FSAD), how to better utilize the fine-grained features of the few images is a key issue. To solve these problems, a Finer-Grained Siamese Network (FGSNet) is proposed in this paper. FGSNet considers the setting of Few-Shot Anomaly Detection (FSAD) and consists of two stages. The first stage is responsible for extracting finer-grained features and roughly aligning the image features, while incorporating an efficient Fine-grained Feature Fusion Module$(\mathbf{F}^{3}\mathbf{M})$to enhance feature representation. Meanwhile, we design a Deep-supervised Loss to improve the level of fine-grained information extraction. The second is primarily used to refine and denoise the fused features, followed by detailed alignment operations for subsequent feature distribution modeling. Compared to most existing FASD models that use single-class training and single-model evaluation, FGSNet is more suitable for existing industrial scenarios. Extensive experiments are conducted on the publicly available MVTec AD and MPDD datasets in this paper. The results show that, for 2, 4 and 8-shot cases, FGSNet can increase the average Image-AUROC on MVTec AD by 0.87%, 1.27% and 1.33% respectively, compared to RegAD.
Delong Han, Mingle Zhou, Gang Li 0005, Min Li 0033
SMC1
2024 Anomaly-Free Prior Guided Knowledge Distillation for Industrial Anomaly Detection
abstract
In industrial manufacturing, visual anomaly detection is critical for maintaining product quality by detecting and preventing production anomalies. Anomaly detection methods based on knowledge distillation demonstrate promising performance in addressing the unpredictability and diversity of anomalies. However, they suffer from a lack of effective guidance from anomaly-free priors when handling anomalous features and underutilize multi-scale features during the segmentation scoring stage, yielding suboptimal detection results. To alleviate these issues, we propose an Anomaly-free Prior Guided knowledge distillation (APG) for industrial anomaly detection. Firstly, it filters the abnormal features by training the de-noising target network with knowledge distillation structure. Concurrently, we propose the Prior Perception Propagation Module (P3M), which extracts more efficient anomaly-free features by imposing constraints on anomalous features. Secondly, we propose the Multi-scale Prior Guided Fusion Module (MPGFM) to improve anomaly detection accuracy by utilizing anomaly-free features from the target network as priors to guide the generation and fusion of cross-scale differential features. Finally, the Global Perception Enhancement Module (GPEM) is proposed to construct an anomaly scoring network, leveraging comprehensive scene features to enhance the detection and localization performance of numerous small-target anomalies in industrial manufacturing. Extensive experiments on the MVTecAD and BTAD datasets show that the proposed method demonstrates a consistent and significant outperformance against competing methods.
Gang Li 0005, Tianjiao Chen, Min Li 0033, Delong Han, Mingle Zhou
SMC4
2024 Revisiting the application of twin connected parallel networks and regression loss functions in industrial defect detection
Zhanzhi Su, Mingle Zhou, Min Li 0033, Delong Han, Gang Li 0005
Adv. Eng. Informatics5
2023 Transforming Limitations into Advantages: Improving Small Object Detection Accuracy with SC-AttentionIoU Loss Function
Mingle Zhou, Changle Yi, Min Li 0033, Honglin Wan, Gang Li 0005, Delong Han
ICANN (7)6
2023 Exploring Adaptive Regression Loss and Feature Focusing in Industrial Scenarios
Mingle Zhou, Zhanzhi Su, Min Li 0033, Delong Han, Gang Li 0005
ICONIP (6)4
2023 A Relational Classification Network Integrating Multi-scale Semantic Features
Gang Li 0005, Jiakai Tian, Mingle Zhou, Min Li 0033, Delong Han
NLPCC (2)5
2023 PCB Defect Detection Model with Convolutional Modules Instead of Self-Attention Mechanism
abstract
The detection of defects in printed circuit boards requires high accuracy and realtime performance. Existing industrial detection models generally adopt a pure convolutional structure for ease of deployment. However, the detection accuracy of these models is often insufficient to meet the requirements of the scene. To improve the detection model accuracy and ease of deployment, this paper proposes a convolutional merging Transformer network(CMTRNet). The CMTR-Net model proposes a backbone network (CNN-Former) that uses convolutional modules to replace self-attention, combining the Transformer architecture with a convolutional structure. This approach not only avoids the drawback of self-attention high computation complexity that is detrimental to deployment but also improves the model detection accuracy. Based on CNN-Former, this paper also proposes a feature fusion module that can better fuse the features extracted by CNN-Former. Further-more, based on the CMTRNet model and the characteristics of circuit board defects, this paper proposes a loss function called Melt-IoU, which can make the initial training phase smoother and further improve detection accuracy. Experiments have shown that CMTRNet outperforms existing advanced models on both datasets.
Mingle Zhou, Changle Yi, Honglin Wan, Min Li 0033, Gang Li 0005, Delong Han
SMC6
2023 IDD-Net: Industrial defect detection method based on Deep-Learning
Mingle Zhou, Honglin Wan, Min Li 0033, Gang Li 0005, Delong Han
Eng. Appl. Artif. Intell.6