Wenyong Yu

dblp:246/1223 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
14since 2021 · last 2026
0000-0003-4012-1264ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-dimensional logic anomaly inspection method for assembly components based on virtual domain contrastive pre-training
Yangfeng Wang, Changyang Yu, Wenyong Yu
Eng. Appl. Artif. Intell.5
2026 Corrigendum to "Multi-dimensional logic anomaly inspection method for assembly components based on virtual domain contrastive pre-training" [Eng. Appl. Artif. Intell., 173 (2026) 14431]
Yangfeng Wang, Changyang Yu, Wenyong Yu
Eng. Appl. Artif. Intell.5
2026 FDFR-Net: A fruit ripeness classification method using multimodal learning technique
Yangfeng Wang, Xinyi Jin, Wenyong Yu
Expert Syst. Appl.5
2025 Exploring Intrinsic Normal Prototypes within a Single Image for Universal Anomaly Detection
abstract
Anomaly detection (AD) is essential for industrial inspection, yet existing methods typically rely on “comparing” test images to normal references from a training set. However, variations in appearance and positioning often complicate the alignment of these references with the test image, limiting detection accuracy. We observe that most anomalies manifest as local variations, meaning that even within anomalous images, valuable normal information remains. We argue that this information is useful and may be more aligned with the anomalies since both the anomalies and the normal information originate from the same image. Therefore, rather than relying on external normality from the training set, we propose INP-Former, a novel method that extracts Intrinsic Normal Prototypes (INPs) directly from the test image. Specifically, we introduce the INP Extractor, which linearly combines normal tokens to represent INPs. We further propose an INP Coherence Loss to ensure INPs can faithfully represent normality for the testing image. These INPs then guide the INP-Guided Decoder to reconstruct only normal tokens, with reconstruction errors serving as anomaly scores. Additionally, we propose a Soft Mining Loss to prioritize hard-to-optimize samples during training. INP-Former achieves state-of-the-art performance in single-class, multi-class, and few-shot AD tasks across MVTec-AD, VisA, and Real-IAD, positioning it as a versatile and universal solution for AD. Remarkably, INP-Former also demonstrates some zero-shot AD capability. Code is available at: https://github.com/luow23/INPFormer.
Yunkang Cao, Haiming Yao, Jianan Lou, Weiming Shen 0001, Wenyong Yu
CVPR8
2025 AMI-Net: Adaptive Mask Inpainting Network for Industrial Anomaly Detection and Localization
abstract
Unsupervised visual anomaly detection is crucial for enhancing industrial production quality and efficiency. Among unsupervised methods, reconstruction approaches are popular due to their simplicity and effectiveness. The key aspect of reconstruction methods lies in the restoration of anomalous regions, which current methods have not satisfactorily achieved. To tackle this issue, we introduce a novel Adaptive Mask Inpainting Network (AMI-Net) from the perspective of adaptive mask-inpainting. In contrast to traditional reconstruction methods that treat non-semantic image pixels as targets, our method uses a pre-trained network to extract multi-scale semantic features as reconstruction targets. Given the multiscale nature of industrial defects, we incorporate a training strategy involving random positional and quantitative masking. Moreover, we propose an innovative adaptive mask generator capable of generating adaptive masks that effectively mask anomalous regions while preserving normal regions. In this manner, the model can leverage the visible normal global contextual information to restore the masked anomalous regions, thereby effectively suppressing the reconstruction of defects. Extensive experimental results on the MVTec AD and BTAD industrial datasets validate the effectiveness of the proposed method. Additionally, AMI-Net exhibits exceptional real-time performance, striking a favorable balance between detection accuracy and speed, rendering it highly suitable for industrial applications.Note to Practitioners—AMI-Net restores defective images to normal ones and subsequently detects defects by leveraging the differences between them. This method only needs to collect about a few hundred defect-free samples for training, without the need for additional defect samples. It is noteworthy that AMI-Net is applicable not only to the detection of simple texture surface defects, such as carpet, leather, and tile, but also to the detection of surface defects in objects with posture diversity, such as cable, transistor, and screw. The trained model not only exhibits high detection accuracy but also demonstrates superior real-time performance, showcasing significant potential in practical industrial settings.
Haiming Yao, Wenyong Yu, Zhengyong Li
IEEE Trans Autom. Sci. Eng.3
2025 Global-Regularized Neighborhood Regression for Efficient Zero-Shot Texture Anomaly Detection
abstract
Texture surface anomaly detection finds widespread applications in industrial settings. However, existing methods often necessitate gathering numerous samples for model training. Moreover, they predominantly operate within a closed-set detection framework, limiting their ability to identify anomalies beyond the training dataset. To tackle these challenges, this article introduces a novel zero-shot texture anomaly detection method named global-regularized neighborhood regression (GRNR). Unlike conventional approaches, GRNR can detect anomalies on arbitrary textured surfaces without any training data or cost. Drawing from human visual cognition, GRNR derives two intrinsic prior supports directly from the test texture image: local neighborhood priors characterized by coherent similarities and global normality priors featuring typical normal patterns. The fundamental principle of GRNR involves utilizing the two extracted intrinsic support priors for self-reconstructive regression of the query sample. This process employs the transformation facilitated by local neighbor support while being regularized by global normality support, aiming to not only achieve visually consistent reconstruction results but also preserve normality properties. We validate the effectiveness of GRNR across various industrial scenarios using eight benchmark datasets, demonstrating its superior detection performance without the need for training data. Remarkably, our method is applicable for open-set texture defect detection and can even surpass existing vanilla approaches that require extensive training.
Haiming Yao, Yunkang Cao, Wenyong Yu, Weiming Shen 0001
IEEE Trans. Syst. Man Cybern. Syst.5
2024 Few-shot unseen defect segmentation for polycrystalline silicon panels with an interpretable dual subspace attention variational learning framework
Haiming Yao, Wenyong Yu, Zhenfeng Qiang, Donghao Luo 0002
Adv. Eng. Informatics3
2024 Template-based Feature Aggregation Network for industrial anomaly detection
Haiming Yao, Wenyong Yu
Eng. Appl. Artif. Intell.3
2024 Dual-Attention Transformer and Discriminative Flow for Industrial Visual Anomaly Detection
abstract
In this paper, we introduce the novel state-of-the-art Dual-attention Transformer and Discriminative Flow (DADF) framework for visual anomaly detection. Based on only normal knowledge, visual anomaly detection has wide applications in industrial scenarios and has attracted significant attention. However, most existing methods fail to meet the requirements of logic defect detection under complex semantic conditions. In contrast, the proposed DADF presents a new paradigm: it firstly leverages a pre-trained network to acquire multi-scale prior embeddings, followed by the development of a vision Transformer with dual attention mechanisms, namely self-attention and memorial-attention, to achieve global-local two-level reconstruction for prior embeddings with the sequential and normality association. Additionally, we propose using normalizing flow to establish discriminative likelihood for the joint distribution of prior and reconstructions at each scale. The experimental results validate the effectiveness of the proposed DADF approach, as evidenced by the impressive performance metrics obtained across various benchmarks, especially for logic defects with complex semantics. Specifically, DADF achieves image-level and pixel-level AUROC scores of 98.3 and 98.4, respectively, on the Mvtec AD benchmark, and an image-level AUROC score of 83.7 and a pixel sPRO score of 67.4 on the Mvtec LOCO AD benchmark. Additionally, we applied DADF to a real-world Printed Circuit Board (PCB) industrial defect inspection task, further demonstrating its efficacy in practical scenarios. The source code of DADF is available at https://github.com/hmyao22/DADF.Note to Practitioners—Most of the current industrial visual inspection techniques can only detect structural defects under uncomplicated semantic settings. Detecting anomalies in products featuring intricate components and logical defects with high-level semantics remains a considerable challenge. The presented DADF is a robust model that can effectively identify defects in products with complex components, such as Printed Circuit Boards (PCBs). Furthermore, it can also accurately detect both structural and logical defects, which is of significant importance for practical industrial applications.
Haiming Yao, Wenyong Yu, Zhenfeng Qiang, Donghao Luo 0002
IEEE Trans Autom. Sci. Eng.3
2024 Learning Global-Local Correspondence With Semantic Bottleneck for Logical Anomaly Detection
abstract
This paper presents a novel framework, named Global-Local Correspondence Framework (GLCF), for visual anomaly detection with logical constraints. Visual anomaly detection has become an active research area in various real-world applications, such as industrial anomaly detection and medical disease diagnosis. However, most existing methods focus on identifying local structural degeneration anomalies and often fail to detect high-level functional anomalies that involve logical constraints. To address this issue, we propose a two-branch approach that consists of a local branch for detecting structural anomalies and a global branch for detecting logical anomalies. To facilitate local-global feature correspondence, we introduce a novel semantic bottleneck enabled by the visual Transformer. Moreover, we develop feature estimation networks for each branch separately to detect anomalies. Our proposed framework is validated using various benchmarks, including industrial datasets, Mvtec AD, Mvtec Loco AD, the logical dataset DigitAnatomy, and the newly proposed Mvtec AAD dataset. Experimental results show that our method outperforms existing methods, particularly in detecting logical anomalies.
Haiming Yao, Wenyong Yu, Zhenfeng Qiang, Donghao Luo 0002
IEEE Trans. Circuits Syst. Video Technol.2
2024 Prior Normality Prompt Transformer for Multiclass Industrial Image Anomaly Detection
abstract
Image anomaly detection plays a pivotal role in industrial inspection. Traditional approaches often demand distinct models for specific categories, resulting in substantial deployment costs. This raises concerns about multiclass anomaly detection, where a unified model is developed for multiple classes. However, applying conventional methods, particularly reconstruction-based models, directly to multiclass scenarios encounters challenges, such as identical shortcut learning, hindering effective discrimination between normal and abnormal instances. To tackle this issue, our study introduces the prior normality prompt transformer (PNPT) method for multiclass image anomaly detection. PNPT strategically incorporates normal semantics prompting to mitigate the “identical mapping” problem. This entails integrating a prior normality prompt into the reconstruction process, yielding a dual-stream model. This innovative architecture combines normal prior semantics with abnormal samples, enabling dual-stream reconstruction grounded in both prior knowledge and intrinsic sample characteristics. PNPT comprises four essential modules: 1) class-specific normality prompting pool, 2) hierarchical patch embedding, 3) semantic alignment coupling encoding, and 4) contextual semantic conditional decoding. Experimental validation on diverse benchmark datasets and real-world industrial applications highlights PNPT's superior performance in multiclass industrial anomaly detection.
Haiming Yao, Yunkang Cao, Wenyong Yu, Weiming Shen 0001
IEEE Trans. Ind. Informatics5
2023 A Feature Memory Rearrangement Network for Visual Inspection of Textured Surface Defects Toward Edge Intelligent Manufacturing
abstract
Recent advances in the industrial inspection of textured surfaces—in the form of visual inspection—have made such inspections possible for efficient, flexible manufacturing systems. However, establishing a unified manual-feature-based inspection model for homogeneous and nonregularly textured surfaces presents an enormous challenge. Furthermore, in real industrial scenarios, collecting and labeling sufficient defective samples is impracticable due to the scarcity of defects and the endless variety of defect types, thus limiting the performance of supervised deep learning methods. To address these challenges, we propose an unsupervised feature memory rearrangement network (FMR-Net) to accurately detect various textural defects simultaneously. Consistent with mainstream methods, we adopt the idea of background reconstruction; however, we innovatively utilize artificial synthetic defects to enable the model to recognize anomalies, while traditional wisdom relies only on defect-free samples. First, we employ an encoding module to obtain multiscale features of the textured surface. Subsequently, a contrastive-learning-based memory feature module (CMFM) is proposed to obtain discriminative representations and construct a normal feature memory bank in the latent space, which can be employed as a substitute for defects and fast anomaly scores at the patch level. Next, a novel global feature rearrangement module (GFRM) is proposed to further suppress the reconstruction of residual defects. Finally, a decoding module utilizes the restored features to reconstruct the normal texture background. In addition, to improve inspection performance, a two-phase training strategy is utilized for accurate defect restoration refinement, and we exploit a multimodal inspection method to achieve noise-robust defect localization. We verify our method through extensive experiments and test its practical deployment in collaborative edge–cloud intelligent manufacturing scenarios by means of a multilevel detection method, demonstrating that FMR-Net exhibits state-of-the-art inspection accuracy and shows great potential for use in edge-computing-enabled smart industries. Note to Practitioners—Most conventional visual inspection methods rely on supervised training and consequently require a large amount of labeled data and can detect only specific types of texture defects. In contrast, the proposed FMR-Net is a robust model for the simultaneous and accurate inspection of textured surfaces for various defects that does not require any real labeled defect samples. Furthermore, this model can also support a different fine-grained detection method that is very suitable in the edge computing paradigm. These two characteristics are both extremely important for practical industrial applications. To the best of our knowledge, this is the first unsupervised edge intelligent vision inspection framework. As such, it can provide inspiration and serve as a reference for intelligent industry.
Haiming Yao, Wenyong Yu, Xue Wang 0001
IEEE Trans Autom. Sci. Eng.2
2022 Joint weakly and fully supervised learning for surface defect segmentation from images
Bin Hu 0020, Xinggang Wang, Wenyong Yu
Signal Process. Image Commun.3
2022 Separable Coupled Dictionary Learning for Large-Scene Precise Classification of Multispectral Images
abstract
Large-scene precise classification of multispectral images (MSIs) has become one of the hot topics in remote sensing field. MSIs usually have wide swath and a meter or even submeter level of spatial resolution, which make large-scene observation possible. However, the limited number of spectral bands leads to the confusion of land covers in classification, especially for the large-scene conditions with abundant land cover types. Therefore, overlapped hyperspectral images (HSIs) can be used to improve the precision degree of classification. To achieve this purpose, coupled dictionary learning has been proposed as a major means. Aiming at separating the class-specific characteristics and mutual patterns among different land covers, this paper proposed a separable coupled dictionary learning (SCDL) method, which converts the separation of mutual features into the construction of separable coupled dictionaries and learns both class-specific coupled dictionaries and mutual coupled dictionaries simultaneously with the aid of label information. More specifically, the proposed method uses the labels of training samples to construct class-specific reconstruction error constraint, class-specificity constraint and separable dictionary incoherence constraint as regularization terms, to make sure that the learned coupled dictionaries to be both compact and discriminative. The learned separable coupled dictionaries facilitate pixels belong to the same category to be represented by the mutual dictionary and the class-specific sub-dictionary of corresponding class. The experiments compared with several state-of-the-art methods on three pairs of HSI and MSI have shown better classification performance.
Tianzhu Liu, Yanfeng Gu, Wenyong Yu, Xiuping Jia, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.3