EDBT 2026 Demo / reviewers in the wild / expert
Haiming Yao
dblp:286/0652
· DBLP profile ↗
20ranked-venue papers
10as first author
20since 2021 · last 2026
0000-0003-1419-5489ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Anomagic: Crossmodal Prompt-driven Zero-shot Anomaly GenerationabstractWe propose Anomagic, a zero-shot anomaly generation method that produces semantically coherent anomalies without requiring any exemplar anomalies. By unifying both visual and textual cues through a crossmodal prompt encoding scheme, Anomagic leverages rich contextual information to steer an inpainting‐based generation pipeline. A subsequent contrastive refinement strategy enforces precise alignment between synthesized anomalies and their masks, thereby bolstering downstream anomaly detection accuracy. To facilitate training, we introduce AnomVerse, a collection of 12,987 anomaly–mask–caption triplets assembled from 13 publicly available datasets, where captions are automatically generated by multimodal large language models using structured visual prompts and template‐based textual hints. Extensive experiments demonstrate that Anomagic trained on AnomVerse can synthesize more realistic and varied anomalies than prior methods, yielding superior improvements in downstream anomaly detection. Furthermore, Anomagic can generate anomalies for any normal‐category image using user‐defined prompts, establishing a versatile foundation model for anomaly generation. Hui Zhang 0023, Qiyu Chen 0002, Haiming Yao, Weiming Shen 0001, Yunkang Cao |
AAAI | 5 |
| 2026 | TDSS: Task Dynamic-Synergistic Skill Adaptation for Boosting Efficient and Scalable Multi-Task Learning in Dense Visual PredictionabstractThe transfer of knowledge from large-scale pre-trained models to diverse downstream tasks has achieved remarkable success. Beyond the traditional full fine-tuning paradigm, Parameter-Efficient Fine-Tuning (PEFT) has emerged as a more efficient model adaptation approach. However, applying existing PEFT methods to adapt dense vision models, particularly in multi-task settings, remains inadequately explored due to their low efficiency, limited task scalability, and neglect of cross-task fine-tuning interactions. To address these challenges, we propose the Task Dynamic-Synergistic Skill Adaptation, termed TDSS, an efficient and scalable multi-task model adaptation framework for dense visual predictions. TDSS comprises two key components: Task-Dynamic Skill Adapters (TDSA) and Task-Synergistic Adaptation Interaction (TSAI). Specifically, TDSA are inserted in parallel into pre-trained vision models to extract task-specific adapted features through the construction of skill representation experts and task dynamic gating. TSAI is developed to enhance cross-task adaptation interaction by bridging global generic and task-specific adapted features. Extensive experiments on multi-task dense visual predictions demonstrate that TDSS surpasses existing state-of-the-art parameter-efficient fine-tuning methods, while exhibiting remarkable efficiency and scalability in parameters and computational complexity. Haiming Yao, Qiyu Chen 0002, Jianxing Liao |
AAAI | 1 |
| 2026 | Parameter-, Memory-, Time-Efficient Multi-Task Dense Vision AdaptationabstractWhile adapting pretrained vision models to downstream dense prediction tasks is widely used, current methods often overlook adaptation efficiency, especially in the context of multi-task learning (MTL). Although parameter-efficient fine-tuning (PEFT) methods can enhance parameter efficiency, broader aspects such as GPU memory and training time efficiency remain underexplored. In this paper, we propose a new paradigm that simultaneously achieves efficiency in Parameters, GPU Memory, and Training Time for Multi-Task Dense Vision Adaptation. Specifically, we propose a dual-branch framework, in which a frozen pretrained backbone serves as the generic main branch, and the proposed Bi-Directional Task Adaptation (BDTA) modules are integrated in parallel to form a task bypass branch that extracts adaptation features required by multiple specific tasks. This adaptation module is lightweight, efficient, and does not require backpropagation through the large pre-trained backbone, thus avoiding resource-intensive gradient computations. Moreover, a Mixture of Task Experts mechanism (MoTE) is further proposed to integrate adaptation features across tasks and scales, thereby obtaining more robust representations tailored for dense prediction tasks. On the PASCAL-Context benchmark, our method achieves over 2× relative performance improvement compared to the best prior multi-task PEFT method, while using only ~30% of the parameters, ~50% of the memory, and ~60% of the training time, demonstrating superior overall adaptation efficiency. Haiming Yao, Qiyu Chen 0002, Jianxing Liao |
AAAI | 1 |
| 2026 | Cross-source medical anomaly detection via prompt-guided diffusion representations
Yunkang Cao, Haiming Yao, Hui Zhang 0023, Weiming Shen 0001 |
Pattern Recognit. | 2 |
| 2026 | URA-Net: Uncertainty-Integrated Anomaly Perception and Restoration Attention Network for Unsupervised Anomaly DetectionabstractUnsupervised anomaly detection plays a pivotal role in industrial defect inspection and medical image analysis, with most methods relying on the reconstruction framework. However, these methods may suffer from over-generalization, enabling them to reconstruct anomalies well, which leads to poor detection performance. To address this issue, instead of focusing solely on normality reconstruction, we propose an innovative Uncertainty-Integrated Anomaly Perception and Restoration Attention Network (URA-Net), which explicitly restores abnormal patterns to their corresponding normality. First, unlike traditional image reconstruction methods, we utilize a pre-trained convolutional neural network to extract multi-level semantic features as the reconstruction target. To assist the URA-Net learning to restore anomalies, we introduce a novel feature-level artificial anomaly synthesis module to generate anomalous samples for training. Subsequently, a novel uncertainty-integrated anomaly perception module based on Bayesian neural networks is introduced to learn the distributions of anomalous and normal features. This facilitates the estimation of anomalous regions and ambiguous boundaries, laying the foundation for subsequent anomaly restoration. Then, we propose a novel restoration attention mechanism that leverages global normal semantic information to restore detected anomalous regions, thereby obtaining defect-free restored features. Finally, we employ residual maps between input features and restored features for anomaly detection and localization. The comprehensive experimental results on two industrial datasets, MVTec AD and BTAD, along with a medical image dataset, OCT-2017, unequivocally demonstrate the effectiveness and superiority of the proposed method. Peng Xing, Yunkang Cao, Haiming Yao, Weiming Shen 0001, Zechao Li |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Mining Global and Local Semantics From Unlabeled Spectra for Spectral ClassificationabstractNon-destructive detection methods based on molecular vibrational spectroscopy are pivotal in fields such as analytical chemistry and medical diagnostics. Recent advances have integrated deep learning with vibrational spectroscopy, significantly enhancing spectral recognition accuracy. However, these methods often rely on large annotated spectral datasets, limiting their general applicability. To address this limitation, we propose a novel approach, Global and Local Semantics Mining (GLSM), which leverages self-supervised learning to capture the global and local semantic information of unlabeled spectra, obviating the need for extensive annotated data. We devise two proxy tasks: global semantic mining and local semantic mining. The global semantic mining task is based on the premise that different views of the same spectrum can be mutually transformed, enabling the model to capture domain-invariant features across various perspectives and thereby develop a global understanding of the spectral data. This, in turn, enhances the model's robustness to variations in peak positions. Meanwhile, the local semantic mining task posits that noisy spectra can be reconstructed into noise-free spectra, thereby facilitating the extraction of local patterns and fine-grained details, such as subtle variations in peak intensities. By combining both self-supervised tasks, our model effectively captures the global and local semantic information of the spectrum. The pre-trained model can be fine-tuned with a limited amount of labeled homologous or heterologous spectral data for semi-supervised or transfer learning-based spectral classification. Extensive experiments on three datasets in semi-supervised and transfer learning-based spectral recognition tasks comprehensively validate the effectiveness of our GLSM method, demonstrating its significant potential for real-world spectral analysis applications. Haiming Yao, Xue Wang 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | Exploring Intrinsic Normal Prototypes within a Single Image for Universal Anomaly DetectionabstractAnomaly detection (AD) is essential for industrial inspection, yet existing methods typically rely on “comparing” test images to normal references from a training set. However, variations in appearance and positioning often complicate the alignment of these references with the test image, limiting detection accuracy. We observe that most anomalies manifest as local variations, meaning that even within anomalous images, valuable normal information remains. We argue that this information is useful and may be more aligned with the anomalies since both the anomalies and the normal information originate from the same image. Therefore, rather than relying on external normality from the training set, we propose INP-Former, a novel method that extracts Intrinsic Normal Prototypes (INPs) directly from the test image. Specifically, we introduce the INP Extractor, which linearly combines normal tokens to represent INPs. We further propose an INP Coherence Loss to ensure INPs can faithfully represent normality for the testing image. These INPs then guide the INP-Guided Decoder to reconstruct only normal tokens, with reconstruction errors serving as anomaly scores. Additionally, we propose a Soft Mining Loss to prioritize hard-to-optimize samples during training. INP-Former achieves state-of-the-art performance in single-class, multi-class, and few-shot AD tasks across MVTec-AD, VisA, and Real-IAD, positioning it as a versatile and universal solution for AD. Remarkably, INP-Former also demonstrates some zero-shot AD capability. Code is available at: https://github.com/luow23/INPFormer. Yunkang Cao, Haiming Yao, Jianan Lou, Weiming Shen 0001, Wenyong Yu |
CVPR | 3 |
| 2025 | Adversarial contrastive domain-generative learning for bacteria Raman spectrum joint denoising and cross-domain identification
Haiming Yao, Xue Wang 0001 |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | A feature shuffling and restoration strategy for universal unsupervised anomaly detection
Haiming Yao, Zhenfeng Qiang |
Knowl. Based Syst. | 2 |
| 2025 | AMI-Net: Adaptive Mask Inpainting Network for Industrial Anomaly Detection and LocalizationabstractUnsupervised visual anomaly detection is crucial for enhancing industrial production quality and efficiency. Among unsupervised methods, reconstruction approaches are popular due to their simplicity and effectiveness. The key aspect of reconstruction methods lies in the restoration of anomalous regions, which current methods have not satisfactorily achieved. To tackle this issue, we introduce a novel Adaptive Mask Inpainting Network (AMI-Net) from the perspective of adaptive mask-inpainting. In contrast to traditional reconstruction methods that treat non-semantic image pixels as targets, our method uses a pre-trained network to extract multi-scale semantic features as reconstruction targets. Given the multiscale nature of industrial defects, we incorporate a training strategy involving random positional and quantitative masking. Moreover, we propose an innovative adaptive mask generator capable of generating adaptive masks that effectively mask anomalous regions while preserving normal regions. In this manner, the model can leverage the visible normal global contextual information to restore the masked anomalous regions, thereby effectively suppressing the reconstruction of defects. Extensive experimental results on the MVTec AD and BTAD industrial datasets validate the effectiveness of the proposed method. Additionally, AMI-Net exhibits exceptional real-time performance, striking a favorable balance between detection accuracy and speed, rendering it highly suitable for industrial applications.Note to Practitioners—AMI-Net restores defective images to normal ones and subsequently detects defects by leveraging the differences between them. This method only needs to collect about a few hundred defect-free samples for training, without the need for additional defect samples. It is noteworthy that AMI-Net is applicable not only to the detection of simple texture surface defects, such as carpet, leather, and tile, but also to the detection of surface defects in objects with posture diversity, such as cable, transistor, and screw. The trained model not only exhibits high detection accuracy but also demonstrates superior real-time performance, showcasing significant potential in practical industrial settings. Haiming Yao, Wenyong Yu, Zhengyong Li |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | VarAD: Lightweight High-Resolution Image Anomaly Detection via Visual Autoregressive ModelingabstractThis article addresses a practical task: high-resolution image anomaly detection (HRIAD). In comparison to conventional image anomaly detection for low-resolution images, HRIAD imposes a heavier computational burden and necessitates superior global information capture capacity. To tackle HRIAD, this article translates image anomaly detection into visual token prediction and proposes visual autoregressive modeling-based anomaly detection (VarAD) based on visual autoregressive modeling for token prediction. Specifically, VarAD first extracts multihierarchy and multidirectional visual token sequences, and then employs an advanced model, Mamba, for visual autoregressive modeling and token prediction. During the prediction process, VarAD effectively exploits information from all preceding tokens to predict the target token. Finally, the discrepancies between predicted tokens and original tokens are utilized to score anomalies. Comprehensive experiments on four publicly available datasets and a real-world button inspection dataset demonstrate that the proposed VarAD achieves superior HRIAD performance while maintaining lightweight, rendering VarAD a viable solution for HRIAD. Yunkang Cao, Haiming Yao, Weiming Shen 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2025 | Center-Aware Residual Anomaly Synthesis for Multiclass Industrial Anomaly DetectionabstractAnomaly detection plays a vital role in the inspection of industrial images. Most existing methods require separate models for each category, resulting in multiplied deployment costs. This highlights the challenge of developing a unified model for multiclass anomaly detection. However, the significant increase in interclass interference leads to severe missed detections. Furthermore, the intraclass overlap between normal and abnormal samples, particularly in synthesis-based methods, cannot be ignored and may lead to over-detection. To tackle these issues, we propose a novel center-aware residual anomaly synthesis (CRAS) method for multiclass anomaly detection. CRAS leverages center-aware residual learning to couple samples from different categories into a unified center, mitigating the effects of interclass interference. To further reduce intraclass overlap, CRAS introduces distance-guided anomaly synthesis that adaptively adjusts noise variance based on normal data distribution. Experimental results on diverse datasets and real-world industrial applications demonstrate the superior detection accuracy and competitive inference speed of CRAS. Qiyu Chen 0002, Huiyuan Luo, Haiming Yao, Zhen Qu, Chengkan Lv, Zhengtao Zhang |
IEEE Trans. Ind. Informatics | 3 |
| 2025 | Global-Regularized Neighborhood Regression for Efficient Zero-Shot Texture Anomaly DetectionabstractTexture surface anomaly detection finds widespread applications in industrial settings. However, existing methods often necessitate gathering numerous samples for model training. Moreover, they predominantly operate within a closed-set detection framework, limiting their ability to identify anomalies beyond the training dataset. To tackle these challenges, this article introduces a novel zero-shot texture anomaly detection method named global-regularized neighborhood regression (GRNR). Unlike conventional approaches, GRNR can detect anomalies on arbitrary textured surfaces without any training data or cost. Drawing from human visual cognition, GRNR derives two intrinsic prior supports directly from the test texture image: local neighborhood priors characterized by coherent similarities and global normality priors featuring typical normal patterns. The fundamental principle of GRNR involves utilizing the two extracted intrinsic support priors for self-reconstructive regression of the query sample. This process employs the transformation facilitated by local neighbor support while being regularized by global normality support, aiming to not only achieve visually consistent reconstruction results but also preserve normality properties. We validate the effectiveness of GRNR across various industrial scenarios using eight benchmark datasets, demonstrating its superior detection performance without the need for training data. Remarkably, our method is applicable for open-set texture defect detection and can even surpass existing vanilla approaches that require extensive training. Haiming Yao, Yunkang Cao, Wenyong Yu, Weiming Shen 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2024 | Few-shot unseen defect segmentation for polycrystalline silicon panels with an interpretable dual subspace attention variational learning framework
Haiming Yao, Wenyong Yu, Zhenfeng Qiang, Donghao Luo 0002 |
Adv. Eng. Informatics | 1 |
| 2024 | Template-based Feature Aggregation Network for industrial anomaly detection
Haiming Yao, Wenyong Yu |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Local-global normality learning and discrepancy normalizing flow for unsupervised image anomaly detection
Haiming Yao, Zhenfeng Qiang, Donghao Luo 0002 |
Eng. Appl. Artif. Intell. | 1 |
| 2024 | Dual-Attention Transformer and Discriminative Flow for Industrial Visual Anomaly DetectionabstractIn this paper, we introduce the novel state-of-the-art Dual-attention Transformer and Discriminative Flow (DADF) framework for visual anomaly detection. Based on only normal knowledge, visual anomaly detection has wide applications in industrial scenarios and has attracted significant attention. However, most existing methods fail to meet the requirements of logic defect detection under complex semantic conditions. In contrast, the proposed DADF presents a new paradigm: it firstly leverages a pre-trained network to acquire multi-scale prior embeddings, followed by the development of a vision Transformer with dual attention mechanisms, namely self-attention and memorial-attention, to achieve global-local two-level reconstruction for prior embeddings with the sequential and normality association. Additionally, we propose using normalizing flow to establish discriminative likelihood for the joint distribution of prior and reconstructions at each scale. The experimental results validate the effectiveness of the proposed DADF approach, as evidenced by the impressive performance metrics obtained across various benchmarks, especially for logic defects with complex semantics. Specifically, DADF achieves image-level and pixel-level AUROC scores of 98.3 and 98.4, respectively, on the Mvtec AD benchmark, and an image-level AUROC score of 83.7 and a pixel sPRO score of 67.4 on the Mvtec LOCO AD benchmark. Additionally, we applied DADF to a real-world Printed Circuit Board (PCB) industrial defect inspection task, further demonstrating its efficacy in practical scenarios. The source code of DADF is available at https://github.com/hmyao22/DADF.Note to Practitioners—Most of the current industrial visual inspection techniques can only detect structural defects under uncomplicated semantic settings. Detecting anomalies in products featuring intricate components and logical defects with high-level semantics remains a considerable challenge. The presented DADF is a robust model that can effectively identify defects in products with complex components, such as Printed Circuit Boards (PCBs). Furthermore, it can also accurately detect both structural and logical defects, which is of significant importance for practical industrial applications. Haiming Yao, Wenyong Yu, Zhenfeng Qiang, Donghao Luo 0002 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2024 | Learning Global-Local Correspondence With Semantic Bottleneck for Logical Anomaly DetectionabstractThis paper presents a novel framework, named Global-Local Correspondence Framework (GLCF), for visual anomaly detection with logical constraints. Visual anomaly detection has become an active research area in various real-world applications, such as industrial anomaly detection and medical disease diagnosis. However, most existing methods focus on identifying local structural degeneration anomalies and often fail to detect high-level functional anomalies that involve logical constraints. To address this issue, we propose a two-branch approach that consists of a local branch for detecting structural anomalies and a global branch for detecting logical anomalies. To facilitate local-global feature correspondence, we introduce a novel semantic bottleneck enabled by the visual Transformer. Moreover, we develop feature estimation networks for each branch separately to detect anomalies. Our proposed framework is validated using various benchmarks, including industrial datasets, Mvtec AD, Mvtec Loco AD, the logical dataset DigitAnatomy, and the newly proposed Mvtec AAD dataset. Experimental results show that our method outperforms existing methods, particularly in detecting logical anomalies. Haiming Yao, Wenyong Yu, Zhenfeng Qiang, Donghao Luo 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Prior Normality Prompt Transformer for Multiclass Industrial Image Anomaly DetectionabstractImage anomaly detection plays a pivotal role in industrial inspection. Traditional approaches often demand distinct models for specific categories, resulting in substantial deployment costs. This raises concerns about multiclass anomaly detection, where a unified model is developed for multiple classes. However, applying conventional methods, particularly reconstruction-based models, directly to multiclass scenarios encounters challenges, such as identical shortcut learning, hindering effective discrimination between normal and abnormal instances. To tackle this issue, our study introduces the prior normality prompt transformer (PNPT) method for multiclass image anomaly detection. PNPT strategically incorporates normal semantics prompting to mitigate the “identical mapping” problem. This entails integrating a prior normality prompt into the reconstruction process, yielding a dual-stream model. This innovative architecture combines normal prior semantics with abnormal samples, enabling dual-stream reconstruction grounded in both prior knowledge and intrinsic sample characteristics. PNPT comprises four essential modules: 1) class-specific normality prompting pool, 2) hierarchical patch embedding, 3) semantic alignment coupling encoding, and 4) contextual semantic conditional decoding. Experimental validation on diverse benchmark datasets and real-world industrial applications highlights PNPT's superior performance in multiclass industrial anomaly detection. Haiming Yao, Yunkang Cao, Wenyong Yu, Weiming Shen 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2023 | A Feature Memory Rearrangement Network for Visual Inspection of Textured Surface Defects Toward Edge Intelligent ManufacturingabstractRecent advances in the industrial inspection of textured surfaces—in the form of visual inspection—have made such inspections possible for efficient, flexible manufacturing systems. However, establishing a unified manual-feature-based inspection model for homogeneous and nonregularly textured surfaces presents an enormous challenge. Furthermore, in real industrial scenarios, collecting and labeling sufficient defective samples is impracticable due to the scarcity of defects and the endless variety of defect types, thus limiting the performance of supervised deep learning methods. To address these challenges, we propose an unsupervised feature memory rearrangement network (FMR-Net) to accurately detect various textural defects simultaneously. Consistent with mainstream methods, we adopt the idea of background reconstruction; however, we innovatively utilize artificial synthetic defects to enable the model to recognize anomalies, while traditional wisdom relies only on defect-free samples. First, we employ an encoding module to obtain multiscale features of the textured surface. Subsequently, a contrastive-learning-based memory feature module (CMFM) is proposed to obtain discriminative representations and construct a normal feature memory bank in the latent space, which can be employed as a substitute for defects and fast anomaly scores at the patch level. Next, a novel global feature rearrangement module (GFRM) is proposed to further suppress the reconstruction of residual defects. Finally, a decoding module utilizes the restored features to reconstruct the normal texture background. In addition, to improve inspection performance, a two-phase training strategy is utilized for accurate defect restoration refinement, and we exploit a multimodal inspection method to achieve noise-robust defect localization. We verify our method through extensive experiments and test its practical deployment in collaborative edge–cloud intelligent manufacturing scenarios by means of a multilevel detection method, demonstrating that FMR-Net exhibits state-of-the-art inspection accuracy and shows great potential for use in edge-computing-enabled smart industries. Note to Practitioners—Most conventional visual inspection methods rely on supervised training and consequently require a large amount of labeled data and can detect only specific types of texture defects. In contrast, the proposed FMR-Net is a robust model for the simultaneous and accurate inspection of textured surfaces for various defects that does not require any real labeled defect samples. Furthermore, this model can also support a different fine-grained detection method that is very suitable in the edge computing paradigm. These two characteristics are both extremely important for practical industrial applications. To the best of our knowledge, this is the first unsupervised edge intelligent vision inspection framework. As such, it can provide inspiration and serve as a reference for intelligent industry. Haiming Yao, Wenyong Yu, Xue Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 1 |