Yunfeng Ma

dblp:51/3979 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 1 since 2021Systems, architecture and hardware · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ZUMA: Training-Free Zero-Shot Unified Multimodal Anomaly Detection
abstract
Multimodal anomaly detection (MAD) aims to exploit both texture and spatial attributes to identify deviations from normal patterns in complex scenarios. However, zero-shot (ZS) settings arising from privacy concerns or confidentiality constraints present significant challenges to existing MAD methods. To address this issue, we introduce ZUMA, a training-free, Zero-shot Unified Multimodal Anomaly detection framework that unleashes CLIP's cross-modal potential to perform ZS MAD. To mitigate the domain gap between CLIP's pretraining space and point clouds, we propose cross-domain calibration (CDC), which efficiently bridges the manifold misalignment through source-domain semantic transfer and establishes a hybrid semantic space, enabling a joint embedding of 2D and 3D representations. Subsequently, ZUMA performs dynamic semantic interaction (DSI) to enable structural decoupling of anomaly regions in the high-dimensional embedding space constructed by CDC, where natural languages serve as semantic anchors to help DSI establish discriminative hyperplanes within hybrid modality representations. Within this framework, ZUMA enables plug-and-play detection of 2D, 3D or multimodal anomalies, without training or fine-tuning even for cross-dataset or incomplete-modality scenarios. Additionally, to further investigate the potential of the training-free ZUMA within the training-based paradigm, we develop ZUMA-FT, a fine-tuned variant that achieves notable improvements with minimal parameter trade-off. Extensive experiments are conducted on two MAD benchmarks, MVTec 3D-AD and Eyecandies. Notably, the training-free ZUMA achieves state-of-the-art (SOTA) performance on both datasets, outperforming existing ZS MAD methods, including training-based approaches. Moreover, ZUMA-FT further extends the performance boundary of ZUMA with only 6.75 M learnable parameters.
Yunfeng Ma, Min Liu 0008, Jingyu Zhou, Yuan Bian 0002, Yaonan Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 Graph Fuzzy Competitive Equilibrium Framework for Interpretable and Robust State Estimation Under FDI Attacks in Smart Grids
abstract
The integration of graph-based fuzzy systems into power system state estimation (SE) remains underexplored, yet it offers strong potential for interpretable and resilient grid monitoring under cyber threats. As modern power grids become increasingly complex, ensuring accurate SE while defending against false data injection (FDI) attacks poses a critical challenge. Traditional SE techniques and bad-data detection methods, relying on linearized models and single-estimator residual checks, often fail to capture nonlinear topology-coupled dependencies and noise in practice, thereby limiting their ability to detect stealthy and coordinated attacks. To bridge this gap, we propose a Graph Fuzzy Competitive Equilibrium State Estimation (GF-CESE) framework that combines graph fuzzy reasoning with power grid topology to enhance interpretability and anomaly detection. By constructing fuzzy rules via virtual center nodes (VCNs) and embedding network connectivity into rule consequents through message propagation, GF-CESE models nonlinear state evolution in a topology-consistent manner. By further introducing a multi-estimator competitive-equilibrium mechanism and cross-estimator residual fusion with a sliding-window test, GF-CESE improves estimation stability and robustness against stealthy FDI attacks. Extensive experiments on IEEE 14-, 118-, and 300-bus systems validate the effectiveness, scalability, and robustness of the proposed approach under high measurement noise and large-scale coordinated multi-node attacks.
Yunfeng Ma, Fuping Hu, Qi Wang 0035
IEEE Trans Autom. Sci. Eng.2
2026 SCAP: Semantic Prototype Alignment for Robust Point Cloud Registration
abstract
Point cloud rigid registration is a fundamental problem in robotics, 3D reconstruction, and augmented reality. However, existing methods predominantly rely on local geometric neighborhoods, which fail to capture higher-order semantic structures and thus degrade performance under noisy or complex geometry conditions. To address these limitations, we propose SCAP, a new point cloud registration paradigm that transforms feature interaction from geometry-driven to semantics–geometric co-driven. Specifically, a semantic prototype extractor is devised to abstract high-level semantic prototypes through graph embedding and clustering, thereby mitigating sensitivity to local feature noise. Since semantic abstraction alone cannot guarantee consistent correspondences across point clouds, SCAP performs a prototype alignment path learning to infer reliable semantic mappings through optimal transport. To enhance cross-layer feature integration and prevent redundant attention, an alignment-driven cross-layer transformer is proposed to incorporate the learned priors into the attention mechanism, thereby enabling feature aggregation with improved semantic coherence and local precision. Extensive experiments on ModelNet, ModelLoNet, 3DMatch, and 3DLoMatch demonstrate that our SCAP consistently surpasses state-of-the-art approaches, showing superior robustness and generalization in challenging scenarios with noise and partial overlap. The code will be available at https://github.com/Zhou-111jy/SCAP.git.
Jingyu Zhou, Yunfeng Ma, Yaonan Wang 0001, Min Liu 0008
IEEE Trans. Circuits Syst. Video Technol.2
2026 Cross-View Dynamic Learning-Based Multi-Class Industrial Anomaly Detection
abstract
Industrial anomaly detection plays a crucial role in smart manufacturing. Traditional methods typically train separate models for each category, leading to substantial memory demands and computational cost. Moreover, relying solely on single-view images is prone to detection blind spots and poor sensitivity to subtle defects. To address these problems, this study proposes CVDL, a cross-view dynamic learning-based multi-class industrial anomaly detection method. Specifically, the CVDL leverages a proposed cross-view dynamic attention in conjunction with intra-view self-attention to dynamically modulate the model’s attention on multi-view information, thereby enhancing the detection performance of subtle defects. Furthermore, a category-guided prompt is developed to utilize object category information, which improves the model’s class-aware detection accuracy. To enhance the model’s robustness, we introduce a structured noise injection strategy and a region-wise mask into the CVDL, mitigating the “identity shortcut” that preserves anomalies during reconstruction. Extensive experiments on the authentic multi-view industrial datasets (Real-IAD) and well-known datasets (MVTec-AD and VisA) confirm the superior detection capability and robustness of the proposed CVDL, and the overall performance of CVDL is superior to all advanced approaches on Real-IAD, achieving SoTA performance of 90.1% image-level and 99.0% pixel-level AUROC. The code will be available at https://github.com/zfinn1/CVDL.git.
Jingyu Zhou, Yunfeng Ma, Yaonan Wang 0001, Min Liu 0008
IEEE Trans. Circuits Syst. Video Technol.2
2026 PANDA: Progressive Adaptive Network for Defect-Aware Few-Shot Segmentation
abstract
Few-shot semantic segmentation aims to reduce reliance on dense annotations, while enhancing generalization to unseen categories. However, most methods are constrained by static global prototypes, which fail to represent subtle defect details, resulting in pronounced support–query misalignment. To address this issue, we propose progressive adaptive network for defect-aware few-shot segmentation (PANDA), a few-shot segmentation framework that integrates representational modulation with semantic consistency constraints. Specifically, to capture subtle and scale-sensitive variations in defect patterns, we design anchored representational modulation (ARM), which overcomes the rigidity of static prototypes by dynamically adjusting representations. In addition, we develop hierarchical semantic coherence (HSC), which enforces consistency across representation hierarchies to suppress the accumulation of semantic drift as depth increases. Collectively, ARM and HSC mitigate support–query misalignment and stabilize representations in few-shot defect segmentation. PANDA achieves state-of-the-art performance on MetFS-18, with 55.1% and 56.3% mean intersection over union under the one-shot and five-shot settings. Moreover, PANDA has been integrated into a real-time industrial inspection platform, where it delivers accurate segmentation across diverse defect types, highlighting its robustness in practical application.
Yunfeng Ma, Min Liu 0008, Xiangfei Meng, Yaonan Wang 0001
IEEE Trans. Ind. Informatics2
2026 Unified Multimodal Industrial Anomaly Detection via Few Normal Samples
abstract
Multimodal industrial anomaly detection (MIAD) is the process of integrating multiple sensor data and utilizing visual intelligence to identify abnormal states in industrial production. In this article, we focus on two main practical but challenging issues in MIAD, i.e., a unified model for multiclass anomaly detection, and model training with only few normal samples. The current mainstream “one-for-one” paradigm requires training time that grows exponentially, and it relies on a sufficient number of samples (even just normal samples), which cannot adapt to practical industrial scenarios with rich abnormal classes. To this end, we offer aUnifiedMIAD model that trained using onlyFew (e.g., 1, 2, and 4) normal samples, termed UniMF. Specifically, we propose a fusion-guided prompt engineering process that generates paired antithetical instance-specific prompts with the assistance of multimodal fusion at both query and token levels. To enable cross-modal prompt learning under multimodal conditions, UniMF performs multi-proxy pairwise matching that involves alignment among multimodal feature patches, embeddings, and tokens of antithetical prompts. Experimental results show that UniMF stands state-of-the-art performance while remaining “one-for-all” paradigm, and even outperforms “one-for-one” methods under certain settings. Cross-dataset evaluation between MVTec 3D-AD and Eyecandies datasets also shows the transferability of UniMF.
Yunfeng Ma, Jingyu Zhou, Yaonan Wang 0001, Min Liu 0008
IEEE Trans. Ind. Informatics2
2026 CiSeg: Unsupervised Cross-Modality Adaptation for 3D Medical Image Segmentation via Causal Intervention
abstract
Unsupervised domain adaptation (UDA) addresses the domain shift problem by transferring knowledge from labeled source domain data (e.g. CT) to unlabeled target domain data (e.g. MRI). While state-of-the-art methods reduce domain gaps via image- or feature-level alignment, their reliance on spurious correlations in the training data often limits generalization across domains. To overcome this limitation, we propose the Causal Intervention Segmentation Network (CiSeg), a novel framework that first integrates causal inference into UDA. A Structural Causal Model (SCM) is first constructed for the source domain to disentangle causal variables from bias variables, alleviating the impact of spurious correlations. Based on this SCM, we introduce a Counterfactual Disentanglement (CD) module to decompose the source domain's latent features into distinct causal and bias components, effectively eliminating their mutual dependencies. To enhance cross-domain consistency, two auxiliary components are introduced: Prototype-guided Contrastive Learning (PCL) and Causal-bias Residual Alignment (CBRA). PCL aligns pixel-level representations with their corresponding semantic prototypes, promoting stronger intra-class consistency and clearer inter-class separability. CBRA employs adversarial learning to align causal and bias residual features across domains, further enhancing feature-level invariance. Extensive experiments on cardiac, abdominal multi-organ, and BraTS18 segmentation tasks demonstrate that CiSeg outperforms state-of-the-art methods, achieving superior segmentation performance and robust cross-domain generalization. Code and models are available at https://github.com/lvpeiqing/CiSeg.
Peiqing Lv, Yaonan Wang 0001, Min Liu 0008, Zhe Zhang 0022, Yunfeng Ma, Licheng Liu, Erik Meijering
IEEE Trans. Medical Imaging5
2025 Multi-Context Aggregation Network With Foreground Correction for Automated Few-Shot Defect Segmentation
abstract
State-of-the-art defect segmentation methods rely on sufficient training data and struggle to generalize to unseen categories. Few-Shot Semantic Segmentation (FSS) is introduced to specifically address these issues. However, existing FSS models still face two challenges in the industry. 1) Defects usually present as weak features, resulting in incomplete segmentation; 2) Severe background interference often leads to incorrect segmentation. To tackle these problems, we propose the Multi-Context Aggregation Network (MCANet). Specifically, we design a Cross-Layer Multi-Level Feature Aggregation Module (CMAM). CMAM effectively aggregates discretely distributed multi-level defect features across different layers and guides the query image to perceive defects from the pixel level, which avoids incomplete segmentation caused by weak features. Additionally, a Foreground Correction Module (FCM) is developed, which is equipped with a dedicated background predictor (BP) and a foreground corrector (FC). BP places more emphasis on learning features from backgrounds rather than defects. FC achieves efficient feature ensemble and further suppresses the backgrounds misidentified as defects in CMAM. They collaborate to prevent incorrect segmentation caused by background interference. Extensive experiments demonstrate the effectiveness of our method. We achieve state-of-the-art results on both FSSD-12, a public benchmark FSS dataset for strip steel, and FSS-AEB, an FSS dataset for aero-engine blades. Specifically, with 1/5 support images, we achieve 64.6%/65.6% mIoU on FSSD-12 and 55.0%/57.8% mIoU on FSS-AEB. Note to Practitioners—Surface defect segmentation has always been a hot topic in the industry. However, existing methods rely on sufficient training data and struggle to generalize to unseen categories, which significantly hinders the automation of defect segmentation. To address this problem, we propose MCANet for automated few-shot defect segmentation. It achieves effective segmentation for surface defects with limited data, even for unseen categories. Furthermore, MCANet achieves state-of-the-art results on two datasets from real-world industrial scenarios and also delivers significant improvements over the widely concerned large vision models. Finally, we integrate MCANet into an automated surface defect inspection platform consisting of an imaging system and a high-performance computing server for real-world performance validation.
Yunfeng Ma, Min Liu 0008, Yuan Bian 0002, Yaonan Wang 0001
IEEE Trans Autom. Sci. Eng.1
2025 SPDP-Net: A Semantic Prior Guided Defect Perception Network for Automated Aero-Engine Blades Surface Visual Inspection
abstract
Automated surface defect detection is essential to manufacturing automation. However, automated inspection of aero-engine blades remains challenging due to tiny defects and weak features. To address this issue, we propose a semantic prior guided defect perception network, named SPDP-Net, which is ultimately integrated into an automated system to achieve efficient detection of defects. Firstly, a semantic prior mining module (SPM) is developed to capture finer-grained pixel-level location priors of defects by leveraging image feature mapping relations which facilitates the precise perception of tiny defects. Subsequently, we propose a defect enhancement perception module (DEP) to separate weak defects from complex backgrounds by utilizing defect location priors provided by SPM to enhance the features of defects while suppressing the values of non-defect regions, which makes the weak defects present as more obvious outliers. Finally, the global information extraction module (GIE) extracts the global features of defects, which helps to further improve the predicted results. When equipped with SPM, DEP and GIE, SPDP-Net can accurately identify and locate defects, exhibiting more competitive recognition and feature extraction capabilities for tiny defects and weak defects. To evaluate the effectiveness of our method, we construct an aero-engine blade surface defect detection dataset from real industrial scenarios called ABSDD with the collaboration of senior engineers. We achieve 95.9% precision, 94.0% recall and 94.9% F1 score on ABSDD. In addition, we also achieve state-of-the-art results on two public benchmark datasets, KSDD2 and DAGM. Finally, we have applied the developed SPDP-Net to an automated system and have conducted actual tests in collaboration with a well-known aero-engine production company.Note to Practitioners—At present, the automated detection system for surface defects in industrial manufacturing has not been well developed, which is particularly trailing behind in the field of aero-engine manufacturing. To the best of our knowledge, the surface defect detection of aero-engine blades is still carried out manually. To address this problem, we propose SPDP-Net for the automated detection of surface defects in aero-engine blades. It can accurately perceive and capture tiny and weak defects with excellent performance. SPDP-Net achieves state-of-the-art results on three tasks from different industrial manufacturing fields, which demonstrates its good transfer application capabilities. In addition, we integrate SPDP-Net into an automated detection system consisting of an autonomous imaging system and a high-performance computing server and conduct actual tests in an aero-engine production company. The test achieves remarkable results, indicating the good application prospects of the automated detection system.
Yunfeng Ma, Min Liu 0008, Yiqiong Zhang, Yaonan Wang 0001
IEEE Trans Autom. Sci. Eng.1
2025 Multi-Estimator Framework for Robust FDI Attack Detection in Smart Grids
abstract
With the increasing incorporation of information and communication technologies, power grids have transitioned into complex cyber-physical systems. While this transformation has brought numerous advantages, it has also exposed the grid to heightened cyber risks, particularly false data injection (FDI) attacks. Traditional state estimation (SE) methods, often employed for FDI detection, face challenges due to their dependence on a single estimator, making them susceptible to sophisticated attacks. To address this, we propose a novel FDI detection method that integrates multiple state estimators within a game-theoretic framework. This approach improves estimator coordination by solving for global optimal and competitive equilibrium, and enhances detection accuracy through an anomaly detector. Validation using the IEEE-14 and IEEE-118 bus systems demonstrates the method’s increased reliability and robustness in real-time grid monitoring.
Yunfeng Ma
IEEE Trans. Circuits Syst. I Regul. Pap.2
2025 Modality Unified Attack for Omni-Modality Person Re-Identification
abstract
Deep learning based person re-identification (re-id) models have been widely employed in surveillance systems. Recent studies have demonstrated that black-box single-modality and cross-modality re-id models are vulnerable to adversarial examples (AEs), leaving the robustness of multi-modality re-id models unexplored. Due to the lack of knowledge about the specific type of model deployed in the target black-box surveillance system, we aim to generate modality unified AEs for omni-modality (single-, cross- and multi-modality) re-id models. Specifically, we propose a novel Modality Unified Attack method to train modality-specific adversarial generators to generate AEs that effectively attack different omni-modality models. A multi-modality model is adopted as the surrogate model, wherein the features of each modality are perturbed by metric disruption loss before fusion. To collapse the common features of omnimodality models, Cross Modality Simulated Disruption approach is introduced to mimic the cross-modality feature embeddings by intentionally feeding images to non-corresponding modality-specific subnetworks of the surrogate model. Moreover, Multi Modality Collaborative Disruption strategy is devised to facilitate the attacker to comprehensively corrupt the informative content of person images by leveraging a multi modality feature collaborative metric disruption loss. Extensive experiments show that our MUA method can effectively attack the omni-modality re-id models, achieving 55.9%, 24.4%, 49.0% and 62.7% mean mAP Drop Rate, respectively.
Yuan Bian 0002, Min Liu 0008, Yunqi Yi, Yunfeng Ma, Yaonan Wang 0001
IEEE Trans. Inf. Forensics Secur.5
2025 Learning to Learn Transferable Generative Attack for Person Re-Identification
abstract
Deep learning-based person re-identification (re-id) models are widely employed in surveillance systems and inevitably inherit the vulnerability of deep networks to adversarial attacks. Existing attacks merely consider cross-dataset and cross-model transferability, ignoring the cross-test capability to perturb models trained in different domains. To powerfully examine the robustness of real-world re-id models, the Meta Transferable Generative Attack (MTGA) method is proposed, which adopts meta-learning optimization to promote the generative attacker producing highly transferable adversarial examples by learning comprehensively simulated transfer-based cross-model&dataset&test black-box meta attack tasks. Specifically, cross-model&dataset black-box attack tasks are first mimicked by selecting different re-id models and datasets for meta-train and meta-test attack processes. As different models may focus on different feature regions, the Perturbation Random Erasing module is further devised to prevent the attacker from learning to only corrupt model-specific features. To boost the attacker learning to possess cross-test transferability, the Normalization Mix strategy is introduced to imitate diverse feature embedding spaces by mixing multi-domain statistics of target models. Extensive experiments show the superiority of MTGA, especially in cross-model&dataset and cross-model&dataset&test attacks, our MTGA outperforms the SOTA methods by 20.0% and 11.3% on mean mAP drop rate, respectively. The source codes are available at https://github.com/yuanbianGit/MTGA.
Yuan Bian 0002, Min Liu 0008, Yunfeng Ma, Yaonan Wang 0001
IEEE Trans. Image Process.4
2023 Towards Hard Real-Time and Energy-Efficient Virtualization for Many-Core Embedded Systems
abstract
In safety-critical computing systems, the I/O virtualization must simultaneously satisfy different requirements, including time-predictability, performance, and energy-efficiency. However, these requirements are challenging to achieve due to complex I/O access path and resource management at the system level, lack of support from preemptive scheduling at I/O hardware level, and missing an effective energy management method. In this paper, we propose a new framework, I/O-GUARD, which reconstructs the system architecture of I/O virtualization, bringing a dedicated hardware hypervisor to handle resource management throughout the system. The hypervisor improves system real-time performance by enabling preemptive scheduling in I/O virtualization with both analytical and experimental real-time guarantees. Furthermore, we also present a dedicated energy management unit to adjustI/O-GUARD's dynamic energy using frequency scaling. Associated with that, a frequency identification algorithm is proposed to find the appropriate executing frequency at run-time. As shown in experiments,I/O-GUARDsimultaneously improves the predictability, performance and energy-efficiency compared to the state-of-the-art I/O virtualization.
Zhe Jiang 0004, Kecheng Yang 0001, Yunfeng Ma, Nathan Fisher, Neil C. Audsley, Zheng Dong 0002
IEEE Trans. Computers3
2021 I/O-GUARD: Hardware/Software Co-Design for I/O Virtualization with Guaranteed Real-time Performance
abstract
For safety-critical| computer systems, time-predictability and performance are usually required simultaneously in I/O virtualization. However, both requirements are challenging to achieve due to complex I/O access path and resource management at system level and lack of support from preemptive scheduling at I/O hardware level. In this paper, we propose a new framework, I/O-GUARD, which reconstructs the system architecture of I/O virtualization, bringing a dedicated hardware hypervisor to handle resource management throughout the system. The hypervisor improves system real-time performance by enabling preemptive scheduling in I/O virtualization with both analytical and experimental real-time guarantees. Specifically, I/O-GUARD is a First-of-Its-Kind framework for multi-/many-core I/O virtualization.
Zhe Jiang 0004, Kecheng Yang 0001, Yunfeng Ma, Nathan Fisher, Neil C. Audsley, Zheng Dong 0002
DAC3
2016 Hardware-Accelerated Parallel Genetic Algorithm for Fitness Functions with Variable Execution Times
abstract
Genetic Algorithms (GAs) following a parallel master-slave architecture can be effectively used to reduce searching time when fitness functions have fixed execution time. This paper presents a parallel GA architecture along with two accelerated GA operators to enhance the performance of masterslave GAs, specially when considering fitness functions with variable execution times. We explore the performance of the proposed approach, and analyse its effectiveness against the state-of-the-art. The results show a significant improvement in search times and fitness function utilisation, thus potentially enabling the use of this approach as a faster searching tool for timing-sensitive optimisation processes such as those found in dynamic real-time systems.
Yunfeng Ma, Leandro Soares Indrusiak
GECCO1
2006 Two Artificial Intelligence Heuristics in Solving Multiple Allocation Hub Maximal Covering Problem
Ke-rui Weng, Yunfeng Ma
ICIC (1)3
2006 A Study on Decision Model of Bottleneck Capacity Expansion with Fuzzy Demand
Mingming Ren, Yunfeng Ma
ICONIP (3)4