VLDB 2026 Research / reviewers in the wild / expert
Yunkang Cao
dblp:321/4825
· DBLP profile ↗
36ranked-venue papers
10as first author
36since 2021 · last 2026
0000-0001-7619-6618ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 6 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 8 · 2 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards High-Resolution 3D Anomaly Detection: A Scalable Dataset and Real-Time Framework for Subtle Industrial DefectsabstractIn industrial point cloud analysis, detecting subtle anomalies demands high-resolution spatial data, yet prevailing benchmarks emphasize low-resolution inputs. To address this disparity, we propose a scalable pipeline for generating realistic and subtle 3D anomalies. Employing this pipeline, we developed MiniShift, the inaugural high-resolution 3D anomaly detection dataset, encompassing 2,577 point clouds, each with 500,000 points and anomalies occupying less than 1% of the total. We further introduce Simple3D, an efficient framework integrating Multi-scale Neighborhood Descriptors (MSND) and Local Feature Spatial Aggregation (LFSA) to capture intricate geometric details with minimal computational overhead, achieving real-time inference exceeding 20 fps. Extensive evaluations on MiniShift and established benchmarks demonstrate that Simple3D surpasses state-of-the-art methods in both accuracy and speed, highlighting the pivotal role of high-resolution data and effective feature aggregation in advancing practical 3D anomaly detection. Yihan Sun 0007, Hui Zhang 0023, Weiming Shen 0001, Yunkang Cao |
AAAI | 5 |
| 2026 | Anomagic: Crossmodal Prompt-driven Zero-shot Anomaly GenerationabstractWe propose Anomagic, a zero-shot anomaly generation method that produces semantically coherent anomalies without requiring any exemplar anomalies. By unifying both visual and textual cues through a crossmodal prompt encoding scheme, Anomagic leverages rich contextual information to steer an inpainting‐based generation pipeline. A subsequent contrastive refinement strategy enforces precise alignment between synthesized anomalies and their masks, thereby bolstering downstream anomaly detection accuracy. To facilitate training, we introduce AnomVerse, a collection of 12,987 anomaly–mask–caption triplets assembled from 13 publicly available datasets, where captions are automatically generated by multimodal large language models using structured visual prompts and template‐based textual hints. Extensive experiments demonstrate that Anomagic trained on AnomVerse can synthesize more realistic and varied anomalies than prior methods, yielding superior improvements in downstream anomaly detection. Furthermore, Anomagic can generate anomalies for any normal‐category image using user‐defined prompts, establishing a versatile foundation model for anomaly generation. Hui Zhang 0023, Qiyu Chen 0002, Haiming Yao, Weiming Shen 0001, Yunkang Cao |
AAAI | 7 |
| 2026 | IAD-R1: Reinforcing Consistent Reasoning in Industrial Anomaly DetectionabstractIndustrial anomaly detection is a critical component of modern manufacturing, yet the scarcity of defective samples restricts traditional detection methods to scenario-specific applications. Although Vision-Language Models (VLMs) demonstrate significant advantages in generalization capabilities, their performance in industrial anomaly detection remains limited. To address this challenge, we propose IAD-R1, a universal post-training framework applicable to VLMs of different architectures and parameter scales, which substantially enhances their anomaly detection capabilities. IAD-R1 employs a two-stage training strategy: the Perception Activation Supervised Fine-Tuning (PA-SFT) stage utilizes a meticulously constructed high-quality Chain-of-Thought dataset (Expert-AD) for training, enhancing anomaly perception capabilities and establishing reasoning-to-answer correlations; the Structured Control Group Relative Policy Optimization (SC-GRPO) stage employs carefully designed reward functions to achieve a capability leap from "Anomaly Perception" to "Anomaly Interpretation". Experimental results demonstrate that IAD-R1 achieves significant improvements across 7 VLMs, the largest improvement was on the DAGM dataset, with average accuracy 43.3% higher than the 0.5B baseline. Notably, the 0.5B parameter model trained with IAD-R1 surpasses commercial models including GPT-4.1 and Claude-Sonnet-4 in zero-shot settings, demonstrating the effectiveness and superiority of IAD-R1. Yunkang Cao, Chengliang Liu 0003, Yuan Xiong, Xinghui Dong, Chao Huang 0008 |
AAAI | 2 |
| 2026 | Bidirectional adaptive transformers for multimodal anomaly detection
Yunkang Cao, Weiming Shen 0001 |
Expert Syst. Appl. | 2 |
| 2026 | Visual anomaly detection under complex view-illumination interplay: A large-scale benchmark
Yunkang Cao, Xiaohao Xu, Yihan Sun 0007, Yuxiang Tan, Xiaonan Huang, Chao Huang 0008, Weiming Shen 0001 |
Pattern Recognit. | 1 |
| 2026 | Cross-source medical anomaly detection via prompt-guided diffusion representations
Yunkang Cao, Haiming Yao, Hui Zhang 0023, Weiming Shen 0001 |
Pattern Recognit. | 1 |
| 2026 | URA-Net: Uncertainty-Integrated Anomaly Perception and Restoration Attention Network for Unsupervised Anomaly DetectionabstractUnsupervised anomaly detection plays a pivotal role in industrial defect inspection and medical image analysis, with most methods relying on the reconstruction framework. However, these methods may suffer from over-generalization, enabling them to reconstruct anomalies well, which leads to poor detection performance. To address this issue, instead of focusing solely on normality reconstruction, we propose an innovative Uncertainty-Integrated Anomaly Perception and Restoration Attention Network (URA-Net), which explicitly restores abnormal patterns to their corresponding normality. First, unlike traditional image reconstruction methods, we utilize a pre-trained convolutional neural network to extract multi-level semantic features as the reconstruction target. To assist the URA-Net learning to restore anomalies, we introduce a novel feature-level artificial anomaly synthesis module to generate anomalous samples for training. Subsequently, a novel uncertainty-integrated anomaly perception module based on Bayesian neural networks is introduced to learn the distributions of anomalous and normal features. This facilitates the estimation of anomalous regions and ambiguous boundaries, laying the foundation for subsequent anomaly restoration. Then, we propose a novel restoration attention mechanism that leverages global normal semantic information to restore detected anomalous regions, thereby obtaining defect-free restored features. Finally, we employ residual maps between input features and restored features for anomaly detection and localization. The comprehensive experimental results on two industrial datasets, MVTec AD and BTAD, along with a medical image dataset, OCT-2017, unequivocally demonstrate the effectiveness and superiority of the proposed method. Peng Xing, Yunkang Cao, Haiming Yao, Weiming Shen 0001, Zechao Li |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | VTFusion: A Vision-Text Multimodal Fusion Network for Few-Shot Anomaly DetectionabstractFew-shot anomaly detection (FSAD) has emerged as a critical paradigm for identifying irregularities using scarce normal references. While recent methods have integrated textual semantics to complement visual data, they predominantly rely on features pretrained on natural scenes, thereby neglecting the granular, domain-specific semantics essential for industrial inspection. Furthermore, prevalent fusion strategies often resort to superficial concatenation, failing to address the inherent semantic misalignment between visual and textual modalities, which compromises robustness against cross-modal interference. To bridge these gaps, this study proposes VTFusion, a vision-text multimodal fusion framework tailored for FSAD. The framework rests on two core designs. First, adaptive feature extractors for both image and text modalities are introduced to learn task-specific representations, bridging the domain gap between pretrained models and industrial data; this is further augmented by generating diverse synthetic anomalies to enhance feature discriminability. Second, a dedicated multimodal prediction fusion module is developed, comprising a fusion block that facilitates rich cross-modal information exchange and a segmentation network that generates refined pixel-level anomaly maps under multimodal guidance. VTFusion significantly advances FSAD performance, achieving image-level area under the receiver operating characteristics (AUROCs) of 96.8% and 86.2% in the 2-shot scenario on the MVTec AD and VisA datasets, respectively. Furthermore, VTFusion achieves an AUPRO of 93.5% on a real-world dataset of industrial automotive plastic parts introduced in this article, further demonstrating its practical applicability in demanding industrial scenarios. Yunkang Cao, Weiming Shen 0001 |
IEEE Trans. Cybern. | 2 |
| 2026 | FPF: A Focused Perception Framework for Small Defect Identification in Complex Power Scenarios
Hui Zhang 0023, Baheti Biekezat, Yunkang Cao, Kaining Zhang, Tongzhi Niu, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2026 | SAF: A Structure-Aware Framework for Radial Ice Thickness Detection on Overhead Transmission LinesabstractIce thickness estimation on overhead transmission lines (OHTL) is essential for mitigating icing-induced mechanical failures and ensuring safe grid operation. To address the challenges of detecting radial ice thickness in complex power line corridors, particularly geometric fragmentation of slender conductors and semantic ambiguity near occluded boundaries, this work proposes a structure-aware framework (SAF) based on 3-D point cloud segmentation and geometry-guided modeling. SAF introduces a structure-aware segmentation network, which integrates a cross-level spatial encoding module to preserve geometric continuity and a partition-aware loss to improve boundary localization under vegetation or tower occlusion. Building on accurate segmentation, a geometry-guided module performs centerline fitting and cross-sectional reconstruction to infer slice-level ice thickness. To support evaluation, a large-scale uncrewed aerial vehicle (UAV)-based point cloud dataset covering 32 OHTL is constructed, including six lines with ground-truth ice labels. Experimental results demonstrate that SAF achieves robust and accurate ice estimation across varied voltage levels and terrains, supporting its practical application in intelligent transmission line inspection and icing risk prevention. Hui Zhang 0023, Youyuan Tang, Yihong Cao, Kaining Zhang, Yunkang Cao, Tongzhi Niu, Jianxu Mao, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 6 |
| 2026 | Toward Zero-Shot Point Cloud Anomaly Detection: A Multiview Projection FrameworkabstractDetecting anomalies within point clouds is crucial for various industrial applications, but traditional unsupervised methods face challenges due to data acquisition costs, early stage production constraints, and limited generalization across product categories. To overcome these challenges, we introduce the multiview projection (MVP) framework, leveraging pretrained vision-language models (VLMs) to detect anomalies. Specifically, MVP projects point cloud data into multiview depth images, thereby translating point cloud anomaly detection into image anomaly detection. Following zero-shot image anomaly detection methods, pretrained VLMs are utilized to detect anomalies on these depth images. Given that pretrained VLMs are not inherently tailored for zero-shot point cloud anomaly detection and may lack specificity, we propose the integration of learnable visual and adaptive text prompting techniques to fine-tune these VLMs, thereby enhancing their detection performance. Extensive experiments on the MVTec 3-D-AD and Real3D-AD demonstrate our proposed MVP framework’s superior zero-shot anomaly detection performance and the prompting techniques’ effectiveness. Real-world evaluations on automotive plastic part inspection further showcase that the proposed method can also be generalized to practical, unseen scenarios. Yunkang Cao, Guoyang Xie, Zhichao Lu, Weiming Shen 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2025 | Customizing Visual-Language Foundation Models for Multi-Modal Anomaly Detection and ReasoningabstractAnomaly detection is vital in various industrial scenarios, including the identification of unusual patterns in production lines and the detection of manufacturing defects for quality control. Existing techniques tend to be specialized in individual scenarios and lack generalization capacities. In this study, our objective is to develop a generic anomaly detection model that can be applied in multiple scenarios. To achieve this, we custom-build generic visual language foundation models that possess extensive knowledge and robust reasoning abilities as anomaly detectors and reasoners. Specifically, we introduce a multi-modal prompting strategy that incorporates domain knowledge from experts as conditions to guide the models. Our approach considers diverse prompt types, including task descriptions, class context, normality rules, and reference images. In addition, we unify the input representation of multi-modality into a 2D image format, enabling multi-modal anomaly detection and reasoning. Our preliminary studies demonstrate that combining visual and language prompts as conditions for customizing the models enhances anomaly detection performance. The customized models showcase the ability to detect anomalies across different data modalities such as images, point clouds, and videos. Qualitative case studies further highlight the anomaly detection and reasoning capabilities, particularly for multi-object scenes and temporal data. Our code is publicly available at https://github.com/Xiaohac-Xu/Customizable-VLM.11More insights of customized foundation models for broader anomaly detection settings are available at Github repo: https://github.com/caoyunkang/GPT4V-for-Generic-Anomaly-Detection. Xiaohao Xu, Yunkang Cao, Huaxin Zhang, Nong Sang, Xiaonan Huang |
CSCWD | 2 |
| 2025 | Exploring Intrinsic Normal Prototypes within a Single Image for Universal Anomaly DetectionabstractAnomaly detection (AD) is essential for industrial inspection, yet existing methods typically rely on “comparing” test images to normal references from a training set. However, variations in appearance and positioning often complicate the alignment of these references with the test image, limiting detection accuracy. We observe that most anomalies manifest as local variations, meaning that even within anomalous images, valuable normal information remains. We argue that this information is useful and may be more aligned with the anomalies since both the anomalies and the normal information originate from the same image. Therefore, rather than relying on external normality from the training set, we propose INP-Former, a novel method that extracts Intrinsic Normal Prototypes (INPs) directly from the test image. Specifically, we introduce the INP Extractor, which linearly combines normal tokens to represent INPs. We further propose an INP Coherence Loss to ensure INPs can faithfully represent normality for the testing image. These INPs then guide the INP-Guided Decoder to reconstruct only normal tokens, with reconstruction errors serving as anomaly scores. Additionally, we propose a Soft Mining Loss to prioritize hard-to-optimize samples during training. INP-Former achieves state-of-the-art performance in single-class, multi-class, and few-shot AD tasks across MVTec-AD, VisA, and Real-IAD, positioning it as a versatile and universal solution for AD. Remarkably, INP-Former also demonstrates some zero-shot AD capability. Code is available at: https://github.com/luow23/INPFormer. Yunkang Cao, Haiming Yao, Jianan Lou, Weiming Shen 0001, Wenyong Yu |
CVPR | 2 |
| 2025 | Unseen Visual Anomaly GenerationabstractVisual anomaly detection (AD) presents significant challenges due to the scarcity of anomalous data samples. While numerous works have been proposed to synthesize anomalous samples, these synthetic anomalies often lack authenticity or require extensive training data, limiting their applicability in real-world scenarios. In this work, we propose Anomaly Anything (AnomalyAny), a novel framework that leverages Stable Diffusion (SD)’s image generation capabilities to generate diverse and realistic unseen anomalies. By conditioning on a single normal sample during test time, AnomalyAny is able to generate unseen anomalies for arbitrary object types with text descriptions. Within AnomalyAny, we propose attention-guided anomaly optimization to direct SD’s attention on generating hard anomaly concepts. Additionally, we introduce prompt-guided anomaly refinement, incorporating detailed descriptions to further improve the generation quality. Extensive experiments on MVTec AD and VisA datasets demonstrate AnomalyAny’s ability in generating high-quality unseen anomalies and its effectiveness in enhancing downstream AD performance. Our demo and code are available at https://hansunhayden.github.io/CUT.github.io/. Yunkang Cao, Hao Dong 0011, Olga Fink |
CVPR | 2 |
| 2025 | Towards VLM-based Hybrid Explainable Prompt Enhancement for Zero-Shot Industrial Anomaly DetectionabstractZero-Shot Industrial Anomaly Detection (ZSIAD) aims to identify and localize anomalies in industrial images from unseen categories. Owing to the powerful generalization capabilities, Vision-Language Models (VLMs) have achieved growing interest in ZSIAD. To guide the model toward understanding and localizing the semantically complex industrial anomalies, existing VLM-based methods have attempted to provide additional prompts to the model through learnable text prompt templates. However, these zero-shot methods lack detailed descriptions of specific anomalies, making it difficult to classify and segment the diverse range of industrial anomalies accurately. To address the aforementioned issue, we firstly propose the multi-stage prompt generation agent for ZSIAD. Specifically, we leverage the Multi-modal Language Large Model (MLLM) to articulate the detailed differential information between normal and test samples, which can provide detailed text prompts to the model through further refinement and anti-false alarm constraint. Moreover, we introduce the Visual Fundamental Model (VFM) to generate anomaly-related attention prompts for more accurate localization of anomalies with varying sizes and shapes. Extensive experiments on seven real-world industrial anomaly detection datasets have shown that the proposed method not only outperforms recent SOTA methods, but also its explainable prompts provide the model with a more intuitive basis for anomaly identification. Weichao Cai, Weiliang Huang, Yunkang Cao, Chao Huang 0008, Bob Zhang 0001, Jie Wen 0001 |
IJCAI | 3 |
| 2025 | Multi-View Reconstruction with Global Context for 3D Anomaly Detection*abstract3D anomaly detection is critical in industrial quality inspection. While existing methods achieve notable progress, their performance degrades in high-precision 3D anomaly detection due to insufficient global information. To address this, we propose Multi-View Reconstruction (MVR), a method that losslessly converts high-resolution point clouds into multi-view images and employs a reconstruction-based anomaly detection framework to enhance global information learning. Extensive experiments demonstrate the effectiveness of MVR, achieving 89.6% object-wise AU-ROC and 95.7% point-wise AU-ROC on the Real3D-AD benchmark. Yihan Sun 0007, Yunkang Cao, Weiming Shen 0001 |
SMC | 3 |
| 2025 | Leveraging Learning Bias for Noisy Anomaly DetectionabstractThis paper addresses the challenge of fully unsupervised image anomaly detection (FUIAD), where training data may contain unlabeled anomalies. Conventional methods assume anomaly-free training data, but real-world contamination leads models to absorb anomalies as normal, degrading detection performance. To mitigate this, we propose a two-stage framework that systematically exploits inherent learning bias in models. The learning bias stems from: (1) the statistical dominance of normal samples, driving models to prioritize learning stable normal patterns over sparse anomalies, and (2) feature-space divergence, where normal data exhibit high intra-class consistency while anomalies display high diversity, leading to unstable model responses. Leveraging the learning bias, stage 1 partitions the training set into subsets, trains sub-models, and aggregates cross-model anomaly scores to filter a purified dataset. Stage 2 trains the final detector on this dataset. Experiments on the Real-IAD benchmark demonstrate superior anomaly detection and localization performance under different noise conditions. Ablation studies further validate the framework’s contamination resilience, emphasizing the critical role of learning bias exploitation. The model-agnostic design ensures compatibility with diverse unsupervised backbones, offering a practical solution for real-world scenarios with imperfect training data. Code is available at https://github.com/hustzhangyuxin/LLBNAD. Yunkang Cao, Yihan Sun 0007, Weiming Shen 0001 |
SMC | 2 |
| 2025 | Diffusion-based vision-language model for zero-shot anomaly detection in medical imagesabstractWith the rapid advancement of diagnostic technology, the ability to detect pathological areas such as tumors and polyps has significantly improved. This progress provides medical imaging specialists with more precise visual information to support anomaly identification, diagnosis, treatment planning, and patient monitoring. However, existing unsupervised and semi-supervised anomaly detection methods struggle with data privacy constraints, limited annotated medical datasets, and challenges in generalization. Zero-Shot Anomaly Detection (ZSAD), which enables the detection of unseen categories without requiring class-specific training, has emerged as a promising solution by leveraging the vision-language alignment capabilities of Vision-Language Models (VLMs), such as Contrastive Language-Image Pretraining (CLIP). Despite recent progress, ZSAD remains hindered by high noise levels, sparse targets, and poor adaptability in complex medical imaging scenarios. To address these issues, we propose a novel framework: DiffusionCLIP, a diffusion-based VLM for zero-shot anomaly detection in two-dimensional medical images. Specifically, DiffusionCLIP integrates diffusion models into the VLM to progressively denoise multi-level features extracted from the CLIP visual encoder, enhancing feature robustness and discriminability. A multi-level feature fusion strategy is designed to aggregate multi-scale representations from different depths of the visual encoder, ensuring complementary semantic alignment across layers. In addition, a dynamically modulated weight loss function is introduced to adaptively balance the learning of hard and easy samples, further improving model generalization. Extensive experiments on multiple benchmark medical imaging datasets, demonstrate that the proposed method significantly outperforms existing zero-shot anomaly detection approaches in terms of accuracy, robustness, and generalization. Yanhui Chen, Hongkang Tao, Zan Yang, Yunkang Cao, Longhua Hu, Pengwen Xiong, Haobo Qiu |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | Boosting Global-Local Feature Matching via Anomaly Synthesis for Multi-Class Point Cloud Anomaly DetectionabstractPoint cloud anomaly detection is essential for various industrial applications. The huge computation and storage costs caused by the increasing product classes limit the application of single-class unsupervised methods, necessitating the development of multi-class unsupervised methods. However, the feature similarity between normal and anomalous points from different class data leads to the feature confusion problem, which greatly hinders the performance of multi-class methods. Therefore, we introduce a multi-class point cloud anomaly detection method, named GLFM, leveraging global-local feature matching to progressively separate data that are prone to confusion across multiple classes. Specifically, GLFM is structured into three stages: Stage-I proposes an anomaly synthesis pipeline that stretches point clouds to create abundant anomaly data that are utilized to adapt the point cloud feature extractor for better feature representation. Stage-II establishes the global and local memory banks according to the global and local feature distributions of all the training data, weakening the impact of feature confusion on the establishment of the memory bank. Stage-III implements anomaly detection of test data leveraging its feature distance from global and local memory banks. Extensive experiments on the MVTec 3D-AD, Real3D-AD and actual industry parts dataset showcase our proposed GLFM’s superior point cloud anomaly detection performance.Note to Practitioners—The proposed GLFM is employed for point cloud anomaly detection in industrial inspection, capable of simultaneously processing data across multiple classes. GLFM requires the collection of a set of normal product samples for model training, where the features of these samples are stored. If the feature distribution of a test sample deviates substantially from that of the normal samples, it is flagged as anomalous. GLFM not only exhibits outstanding performance on public datasets but has also been validated on a real-world industrial parts point cloud dataset. Yunkang Cao, Weiming Shen 0001, Wenlong Li 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | LogiCode: An LLM-Driven Framework for Logical Anomaly DetectionabstractThis paper presents LogiCode, a novel framework that leverages Large Language Models (LLMs) for identifying logical anomalies in industrial settings, moving beyond the traditional focus on structural inconsistencies. By harnessing LLMs for logical reasoning, LogiCode autonomously generates Python codes to pinpoint anomalies such as incorrect component quantities or missing elements, marking a significant leap forward in anomaly detection technologies. A custom dataset “LOCO-Annotations” and a benchmark “LogiBench” are introduced to evaluate the LogiCode’s performance across various metrics including binary classification accuracy, code generation success rate, and precision in reasoning. Findings demonstrate LogiCode’s enhanced interpretability, significantly improving the accuracy of logical anomaly detection and offering detailed explanations for identified anomalies. This represents a notable shift towards more intelligent, LLM-driven approaches in industrial anomaly detection, promising substantial impacts on industry-specific applications. Our code are available athttps://github.com/22strongestme/LOCO-Annotations. Note to Practitioners—This work introduces LogiCode, an innovative system leveraging Large Language Models (LLMs) for logical anomaly detection in industrial settings, shifting the paradigm from traditional visual inspection methods. LogiCode autonomously generates Python codes for logical anomaly detection, enhancing interpretability and accuracy. Our novel approach, validated through the “LOCO-Annotations” dataset and LogiBench benchmark, demonstrates superior performance in identifying logical anomalies, a challenge often encountered in complex industrial components like assembly and packaging. LogiCode provides a significant advancement in addressing the nuanced requirements of detecting logical anomalies, offering a robust and interpretable solution to practitioners seeking to enhance quality control and reduce manual inspection efforts. Yunkang Cao, Xiaohao Xu, Weiming Shen 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Personalizing Vision-Language Models With Hybrid Prompts for Zero-Shot Anomaly DetectionabstractZero-shot anomaly detection (ZSAD) aims to develop a foundational model capable of detecting anomalies across arbitrary categories without relying on reference images. However, since "abnormality" is inherently defined in relation to "normality" within specific categories, detecting anomalies without reference images describing the corresponding normal context remains a significant challenge. As an alternative to reference images, this study explores the use of widely available product standards to characterize normal contexts and potential abnormal states. Specifically, this study introduces AnomalyVLM, which leverages generalized pretrained vision-language models (VLMs) to interpret these standards and detect anomalies. Given the current limitations of VLMs in comprehending complex textual information, AnomalyVLM generates hybrid prompts-comprising prompts for abnormal regions, symbolic rules, and region numbers-from the standards to facilitate more effective understanding. These hybrid prompts are incorporated into various stages of the anomaly detection process within the selected VLMs, including an anomaly region generator and an anomaly region refiner. By utilizing hybrid prompts, VLMs are personalized as anomaly detectors for specific categories, offering users flexibility and control in detecting anomalies across novel categories without the need for training data. Experimental results on four public industrial anomaly detection datasets, as well as a practical automotive part inspection task, highlight the superior performance and enhanced generalization capability of AnomalyVLM, especially in texture categories. An online demo of AnomalyVLM is available at https://github.com/caoyunkang/Segment-Any-Anomaly. Yunkang Cao, Xiaohao Xu, Chen Sun 0015, Zongwei Du, Liang Gao 0001, Weiming Shen 0001 |
IEEE Trans. Cybern. | 1 |
| 2025 | VarAD: Lightweight High-Resolution Image Anomaly Detection via Visual Autoregressive ModelingabstractThis article addresses a practical task: high-resolution image anomaly detection (HRIAD). In comparison to conventional image anomaly detection for low-resolution images, HRIAD imposes a heavier computational burden and necessitates superior global information capture capacity. To tackle HRIAD, this article translates image anomaly detection into visual token prediction and proposes visual autoregressive modeling-based anomaly detection (VarAD) based on visual autoregressive modeling for token prediction. Specifically, VarAD first extracts multihierarchy and multidirectional visual token sequences, and then employs an advanced model, Mamba, for visual autoregressive modeling and token prediction. During the prediction process, VarAD effectively exploits information from all preceding tokens to predict the target token. Finally, the discrepancies between predicted tokens and original tokens are utilized to score anomalies. Comprehensive experiments on four publicly available datasets and a real-world button inspection dataset demonstrate that the proposed VarAD achieves superior HRIAD performance while maintaining lightweight, rendering VarAD a viable solution for HRIAD. Yunkang Cao, Haiming Yao, Weiming Shen 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2025 | Prototypical Learning Guided Context-Aware Segmentation Network for Few-Shot Anomaly DetectionabstractFew-shot anomaly detection (FSAD) denotes the identification of anomalies within a target category with a limited number of normal samples. Existing FSAD methods largely rely on pretrained feature representations to detect anomalies, but the inherent domain gap between pretrained representations and target FSAD scenarios is often overlooked. This study proposes a prototypical learning-guided context-aware segmentation network (PCSNet) to address the domain gap, thereby improving feature descriptiveness in target scenarios and enhancing FSAD performance. In particular, PCSNet comprises a prototypical feature adaption (PFA) subnetwork and a context-aware segmentation (CAS) subnetwork. PFA extracts prototypical features as guidance to ensure better feature compactness for normal data while distinct separation from anomalies. A pixel-level disparity classification (PDC) loss is also designed to make subtle anomalies more distinguishable. Then a CAS subnetwork is introduced for pixel-level anomaly localization, where pseudo anomalies are exploited to facilitate the training process. Experimental results on MVTec AD and metal part defect detection (MPDD) demonstrate the superior FSAD performance of PCSNet, with 94.9% and 80.2% image-level area under the receiver operating characteristics (AUROCs) in an eight-shot scenario, respectively. Real-world applications on automotive plastic part inspection further demonstrate that PCSNet can achieve promising results with limited training samples. The code is available at https://github.com/yuxin-jiang/PCSNet. Yunkang Cao, Weiming Shen 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Global-Regularized Neighborhood Regression for Efficient Zero-Shot Texture Anomaly DetectionabstractTexture surface anomaly detection finds widespread applications in industrial settings. However, existing methods often necessitate gathering numerous samples for model training. Moreover, they predominantly operate within a closed-set detection framework, limiting their ability to identify anomalies beyond the training dataset. To tackle these challenges, this article introduces a novel zero-shot texture anomaly detection method named global-regularized neighborhood regression (GRNR). Unlike conventional approaches, GRNR can detect anomalies on arbitrary textured surfaces without any training data or cost. Drawing from human visual cognition, GRNR derives two intrinsic prior supports directly from the test texture image: local neighborhood priors characterized by coherent similarities and global normality priors featuring typical normal patterns. The fundamental principle of GRNR involves utilizing the two extracted intrinsic support priors for self-reconstructive regression of the query sample. This process employs the transformation facilitated by local neighbor support while being regularized by global normality support, aiming to not only achieve visually consistent reconstruction results but also preserve normality properties. We validate the effectiveness of GRNR across various industrial scenarios using eight benchmark datasets, demonstrating its superior detection performance without the need for training data. Remarkably, our method is applicable for open-set texture defect detection and can even surpass existing vanilla approaches that require extensive training. Haiming Yao, Yunkang Cao, Wenyong Yu, Weiming Shen 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2024 | AdaCLIP: Adapting CLIP with Hybrid Learnable Prompts for Zero-Shot Anomaly Detection
Yunkang Cao, Jiangning Zhang, Luca Frittoli, Weiming Shen 0001, Giacomo Boracchi |
ECCV (35) | 1 |
| 2024 | Dual-path Frequency Discriminators for few-shot anomaly detection
Yuhu Bai, Jiangning Zhang, Zhaofeng Chen, Yunkang Cao, Guanzhong Tian |
Knowl. Based Syst. | 5 |
| 2024 | Generative Denoise Distillation: Simple stochastic noises induce efficient knowledge transfer for dense prediction
Zhaoge Liu, Xiaohao Xu, Yunkang Cao, Weiming Shen 0001 |
Knowl. Based Syst. | 3 |
| 2024 | Complementary pseudo multimodal feature for point cloud anomaly detection
Yunkang Cao, Xiaohao Xu, Weiming Shen 0001 |
Pattern Recognit. | 1 |
| 2024 | Prior Normality Prompt Transformer for Multiclass Industrial Image Anomaly DetectionabstractImage anomaly detection plays a pivotal role in industrial inspection. Traditional approaches often demand distinct models for specific categories, resulting in substantial deployment costs. This raises concerns about multiclass anomaly detection, where a unified model is developed for multiple classes. However, applying conventional methods, particularly reconstruction-based models, directly to multiclass scenarios encounters challenges, such as identical shortcut learning, hindering effective discrimination between normal and abnormal instances. To tackle this issue, our study introduces the prior normality prompt transformer (PNPT) method for multiclass image anomaly detection. PNPT strategically incorporates normal semantics prompting to mitigate the “identical mapping” problem. This entails integrating a prior normality prompt into the reconstruction process, yielding a dual-stream model. This innovative architecture combines normal prior semantics with abnormal samples, enabling dual-stream reconstruction grounded in both prior knowledge and intrinsic sample characteristics. PNPT comprises four essential modules: 1) class-specific normality prompting pool, 2) hierarchical patch embedding, 3) semantic alignment coupling encoding, and 4) contextual semantic conditional decoding. Experimental validation on diverse benchmark datasets and real-world industrial applications highlights PNPT's superior performance in multiclass industrial anomaly detection. Haiming Yao, Yunkang Cao, Wenyong Yu, Weiming Shen 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | BiaS: Incorporating Biased Knowledge to Boost Unsupervised Image Anomaly LocalizationabstractImage anomaly localization is a pivotal technique in industrial inspection, often manifesting as a supervised task where abundant normal samples coexist with rare abnormal samples. Existing supervised methods in this context are prone to overfitting, as they primarily encounter anomalies that represent only a fraction of the open-world anomalies. Conversely, unsupervised methods excel in performance, yet they disregard the essential biased knowledge pertaining to both seen and unseen anomalies within the open world. To bridge this gap and refine unsupervised methods for supervised applications, this study introduces a comprehensive framework called biased students (BiaS), mainly comprising a three-step strategy. This strategy encompasses biased knowledge generation, transfer, and fusion. BiaS effectively segregates the vast anomaly space into two subsets: 1) unseen anomalies and 2) seen anomalies. Subsequently, it generates specialized biased knowledge for these subsets and transfers this knowledge to two distinct subnetworks. As a result, one subnetwork becomes adept at detecting unseen anomalies, while the other excels in localizing seen anomalies. To optimize their capabilities, BiaS synergistically fuses these subnetworks based on their expertise. Rigorous experimentation has empirically validated the effectiveness, generality, and scalability of BiaS, underscoring its potential to enhance unsupervised methods and effectively address the challenges of supervised anomaly localization. Yunkang Cao, Xiaohao Xu, Chen Sun 0015, Liang Gao 0001, Weiming Shen 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2023 | A masked reverse knowledge distillation method incorporating global and local information for image anomaly detection
Yunkang Cao, Weiming Shen 0001 |
Knowl. Based Syst. | 2 |
| 2023 | Collaborative Discrepancy Optimization for Reliable Image Anomaly LocalizationabstractMost unsupervised image anomaly localization methods suffer from overgeneralization because of the high generalization abilities of convolutional neural networks, leading to unreliable predictions. To mitigate the overgeneralization, this article proposes to collaboratively optimize normal and abnormal feature distributions with the assistance of synthetic anomalies, namely collaborative discrepancy optimization (CDO). CDO introduces a margin optimization module and an overlap optimization module to optimize the two key factors determining the localization performance, i.e., the margin and the overlap between the discrepancy distributions (DDs) of normal and abnormal samples. With CDO, a large margin and a small overlap between normal and abnormal DDs are obtained, and the prediction reliability is boosted. Experiments on MVTec2D and MVTec3D show that CDO effectively mitigates the overgeneralization and achieves great anomaly localization performance with real-time computation efficiency. A real-world automotive plastic parts inspection application further demonstrates the capability of the proposed CDO. Yunkang Cao, Xiaohao Xu, Zhaoge Liu, Weiming Shen 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2022 | Semi-supervised Knowledge Distillation for Tiny Defect DetectionabstractImage anomaly detection can automatically detect defects using images of products, which is crucial for product quality controls. Because of insufficient abnormal data, unsupervised image anomaly detection based on knowledge distillation has attracted broad attention recently. However, fully unsupervised methods suffer from detecting tiny anomalies that widely exist in industrial products because the features of tiny anomalies and normal features extracted by the teacher network are similar. This paper extends current unsupervised anomaly detection methods into a semi-supervised manner, simultaneously leveraging normal data and a limited amount of abnormal data. An automobile plastic parts dataset is established to prove the effectiveness of the proposed method. Experiments show that the proposed method can accurately detect small anomalies and largely surpass a powerful baseline (6% in AU-ROC, 10% in F1-score, 11% in Accuracy). Yunkang Cao, Yanan Song, Xiaohao Xu, Shuya Li, Yuhao Yu, Yifeng Zhang 0007, Weiming Shen 0001 |
CSCWD | 1 |
| 2022 | An Outlier-Aware Method for UWB Indoor Positioning in NLoS SituationsabstractUltra-wideband (UWB) technology has been widely applied in the high-precision indoor positioning system. However, the complicated indoor environment makes signals propagate in non-line-of-sight (NLoS) situations, which seriously deteriorates the positioning accuracy. This work proposes an outlier-aware method to improve the positioning accuracy under NLoS scenarios. End-to-end optimization and positioning are achieved by combining the measurement error mitigation process with the positioning process. Experiments on public benchmarks illustrate that the proposed method enhances the performance of indoor positioning in NLoS situations. Chuan Liu 0001, Yunkang Cao, Chen Sun 0015, Weiming Shen 0001, Xinyu Li 0001, Liang Gao 0001 |
CSCWD | 2 |
| 2022 | Informative knowledge distillation for image anomaly segmentation
Yunkang Cao, Weiming Shen 0001, Liang Gao 0001 |
Knowl. Based Syst. | 1 |
| 2022 | GON: End-to-end optimization framework for constraint graph optimization problems
Chuan Liu 0001, Jingwei Wang 0001, Yunkang Cao, Min Liu 0002, Weiming Shen 0001 |
Knowl. Based Syst. | 3 |