EDBT 2026 Demo / reviewers in the wild / expert
Yurong Chen 0003
dblp:02/41-3
· DBLP profile ↗
19ranked-venue papers
8as first author
19since 2021 · last 2026
0000-0002-6171-4555ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DDIP: Mutual-Regularized Dual Deep Image Prior for Self-Supervised Compressive Spectral ImagingabstractIn recent years, numerous hyperspectral image (HSI) reconstruction methods have been proposed to enhance the imaging quality of coded aperture snapshot spectral compressive imaging (CASSI) systems. Among these methods, self-supervised Deep Image Prior (DIP)-based approaches have gained attention for their ability to reconstruct three-dimensional (3D) HSIs without the need for external training data. However, DIP methods often suffer from overfitting to high-frequency noise during the optimization process, leading to artifacts and loss of fine details. To address these challenges, we propose a Mutual-Regularized Dual Deep Image Prior (DDIP) framework that employs implicit mutual regularization between two DIP networks. By encouraging mutual constraints, DDIP effectively mitigates high-frequency learning bias and suppresses noise amplification. Additionally, we employ a Half Quadratic Splitting (HQS) optimization strategy to ensure stable and efficient convergence, progressively integrating complementary information from the dual networks. We provide a comprehensive convergence analysis of the DDIP framework and establish theoretical conditions to guide the progressive fusion of the dual networks, ensuring robust and reliable reconstruction. Based on the insights from the convergence analysis, we introduce an Adaptive Deep Image Prior inner-loop strategy that dynamically adjusts the inner-loop updates, ensuring balanced learning of low- and high-frequency components. Moreover, a Residual Spectral-Spatial Feature Attention Network (SSFAN) is designed to enhance spectral-spatial feature extraction, further improving reconstruction accuracy. Extensive experiments on benchmark datasets demonstrate that DDIP achieves competitive HSI reconstruction quality compared to state-of-the-art unsupervised and self-supervised methods. Lizhu Liu, Yaonan Wang 0001, Yurong Chen 0003, Hui Zhang 0023 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Balancing Accuracy and Efficiency With a Multiscale Uncertainty-Aware Knowledge-Based Network for Transmission Line InspectionabstractReal-world transmission line inspections (RTLIs) ensure power stability and safety. Deep learning (DL) models have become prevalent approaches for performing RTLI tasks. However, the high computational demands and substantial parameter requirements of DL models limit their real-world applicability. This article introduces a novel approach, a multiscale uncertainty-aware knowledge-based network, which is designed to balance the accuracy and efficiency in RTLI tasks. Specifically, we propose an uncertainty-aware knowledge distillation method that incorporates pixel-level uncertainty into the knowledge transfer process, mitigating the impact of noisy knowledge derived from extra background information contained in ground truths. In addition, our method integrates a multiscale relationship distillation technique, thus enhancing the transfer of multiscale information between the teacher and student models. Consequently, RTLI tasks can be efficiently accomplished using the well-learned lightweight student model. Comprehensive experiments conducted on a real-world dataset collected via uncrewedaerial vehicles demonstrate the efficacy of our proposed approach in terms of achieving high detection accuracy with reduced computational costs. Junfei Yi, Jianxu Mao, Hui Zhang 0023, Yurong Chen 0003, Tengfei Liu 0005, Kai Zeng 0010, He Xie, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2025 | LCTC: Lightweight Convolutional Thresholding Sparse Coding Network Prior for Compressive Hyperspectral ImagingabstractCompressive spectral imaging has garnered significant attention for its ability to effectively enhance the captured spatial and spectral information. Predominant methods, based on compressive sensing, typically formulate the imaging task as a constrained optimization problem and rely on hand-crafted priors to model the sparsity of spectral images. However, these approaches often suffer from suboptimal performance due to the inherent difficulty of identifying an appropriate transform space where spectral images exhibit sparsity. To overcome this limitation, we propose a novel convolutional sparse coding-inspired untrained network prior for fast and adaptive identification of the sparse transform domain and compressible signal. Specifically, a Lightweight Convolutional Thresholding sparse Coding (LCTC) network is designed as the sparse transform domain, with its inputs interpreted as sparse coefficients. Crucially, both the transform domain and its coefficients are solved in a self-supervised learning manner. Furthermore, we demonstrate that LCTC prior can be seamlessly incorporated into the iterative optimization algorithm as a Plug-and-Play (PnP) regularization. Both the LCTC and PnP-LCTC exhibit superior performance compared to previous methods. Experiments under various scenarios validate the effectiveness and efficiency of our approach. Yurong Chen 0003, Yaonan Wang 0001, Xiaodong Wang 0026, Xin Yuan 0002, Hui Zhang 0023 |
IEEE Trans. Image Process. | 1 |
| 2025 | Unsupervised Range-Nullspace Learning Prior for Multispectral Images ReconstructionabstractSnapshot Spectral Imaging (SSI) techniques, with the ability to capture both spectral and spatial information in a single exposure, have been found useful in a wide range of applications. SSI systems generally operate within the 'encoding-decoding' framework, leveraging the synergism of optical hardware and reconstruction algorithms. Typically, reconstructing desired spectral images from SSI measurements is an ill-posed and challenging problem. Existing studies utilize either model-based or deep learning-based methods, but both have their drawbacks. Model-based algorithms suffer from high computational costs, while supervised learning-based methods rely on large paired training data. In this paper, we propose a novel Unsupervised range-Nullspace learning (UnNull) prior for spectral image reconstruction. UnNull explicitly models the data via subspace decomposition, offering enhanced interpretability and generalization ability. Specifically, UnNull considers that the spectral images can be decomposed into the range and null subspaces. The features projected onto the range subspace are mainly low-frequency information, while features in the nullspace represent high-frequency information. Comprehensive multispectral demosaicing and reconstruction experiments demonstrate the superior performance of our proposed algorithm. Yurong Chen 0003, Yaonan Wang 0001, Hui Zhang 0023 |
IEEE Trans. Image Process. | 1 |
| 2025 | SSPD: Spatial-Spectral Prior Decoupling Model for Spectral Snapshot Compressive ImagingabstractCoded aperture snapshot spectral imaging (CASSI) captures 3D hyperspectral images (HSIs) in a single shot by encoding incident light into 2D measurements. However, recovering the original hyperspectral data from these measurements is a severely ill-posed inverse problem due to significant information loss during compression. Recent deep learning methods, especially deep unfolding networks, have demonstrated promising reconstruction results by embedding learnable priors into iterative optimization frameworks. However, most existing approaches use a single network to jointly estimate spatial and spectral priors, limiting their ability to handle the distinct properties of HSIs. To overcome this limitation, we propose the Spatial-Spectral Prior Decoupling Model (SSPD), which reformulates HSI reconstruction as a prior absorption problem, enabling independent modeling of spatial and spectral priors with specialized network architectures. To achieve this, we design two attention mechanisms tailored for hyperspectral data: one for capturing spatial correlations and another for preserving spectral signatures. Additionally, we develop a hybrid loss function that combines convergence constraints and cross-prior interactions, ensuring accurate prior fusion and stable reconstruction. Experiments on synthetic and real-world datasets confirm that SSPD outperforms existing methods in spectral snapshot compressive imaging. Lizhu Liu, Yaonan Wang 0001, Yurong Chen 0003, Jiwen Lu, Hui Zhang 0023 |
IEEE Trans. Multim. | 3 |
| 2025 | Reliable Wind Turbine Blade Performance Monitoring System Using Aerodynamic Audio Signals and Deep Learning ApproachesabstractWind turbines have emerged as a prominent and environmentally friendly energy generation solution. However, with the widespread use of new materials, ensuring the reliability of these devices has become as a critical issue. Developing efficient and cost-effective monitoring methods for the wind turbine's blades (WTBs), the most expensive components of wind turbine, has become a focal point of research. In this article, we present a novel monitoring system for WTBs that employs a deep convolutional neural network approach based on the medical auscultatory method. The system is designed to balance economic efficiency and engineering reliability. First, we proposed a lightweight WTBs monitoring framework based on edge computing that leverages the signals from the programmable logic controller output of wind turbine to enable efficient collection of relevant aerodynamic audio signals while filtering out irrelevant data. Second, we present a set of audio enhancement algorithms that employ multiscale feature extraction, self-adaptive mask targeting, and deep neural networks to reduce noise in the audio signals generated by WTBs. Third, we introduce a new approach for compressing deep convolution neural networks that makes them suitable for resource-constrained edge computing devices and efficiently utilizes audio-generated spectrograms to diagnose faults in WTBs. Baheti Biekezat, Hui Zhang 0023, Yihong Cao, Yurong Chen 0003, Yaonan Wang 0001 |
IEEE Trans. Reliab. | 4 |
| 2024 | Prior Images Guided Generative Autoencoder Model for Dual-Camera Compressive Spectral ImagingabstractCompressive Spectral Imaging (CSI) techniques have attracted considerable attention among researchers for their ability to simultaneously capture spatial and spectral information using low-cost, compact optical components. A prominent example of CSI techniques is the Dual-Camera Coded Aperture Snapshot Spectral Imaging (DC-CASSI), which involves reconstructing hyperspectral images from CASSI measurements and uncoded panchromatic or RGB images. Despite its significance, the reconstruction process in DC-CASSI is challenging. Conventional DC-CASSI techniques rely on different models to explore the similarity between uncoded images and hyperspectral images. Nevertheless, two main issues persist: i) the effective utilization ofspatial informationfrom RGB images to guide the reconstruction process, and ii) the enhancement ofspectral consistencyof recovered images when using panchromatic/RGB images, which inherently lack precise spectral information. To address these challenges, we propose a novel Prior images guided generative autoEncoder (PiE) model. The PiE model leverages RGB images as prior information to enhance spatial details and designs a generative model to improve spectral quality. Notably, the generative model is optimized in a self-supervised manner. Comprehensive experimental results demonstrate that the proposed PiE method outperforms existing techniques, achieving state-of-the-art performance. Yurong Chen 0003, Yaonan Wang 0001, Hui Zhang 0023 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | ADMM-DSP: A Deep Spectral Image Prior for Snapshot Spectral Image DemosaicingabstractSpectral imaging, with the ability to simultaneously capture the spectral and spatial information of scenes, has obtained researchers' significant interest. Traditional spectral imaging techniques typically suffer from high costs and slow imaging speed. Multispectral filter array-based snapshot imaging is a cutting-edge technology for mitigating these problems. The spectral image demosaicing algorithm plays a key role in reconstructing high-resolution spectral images from the raw measurement. In this article, we first formulate the multispectral demosaicing problem as a compressive spectral imaging problem. Then, a novel deep spectral image prior is introduced as the regularization, which assumes that neural networks can generate the desired spectral image from the raw image in a self-supervised learning manner. Finally, the constructed constrained problem is solved by the alternating direction method of multipliers optimization algorithm. Compared with existing methods, experiments on various scenes demonstrate that the proposed method obtains superior performance. Yurong Chen 0003, Hui Zhang 0023, Yaonan Wang 0001, Ating Yin |
IEEE Trans. Ind. Informatics | 1 |
| 2024 | Flex-DLD: Deep Low-Rank Decomposition Model With Flexible Priors for Hyperspectral Image Denoising and RestorationabstractHyperspectral images (HSIs) are composed of hundreds of contiguous waveband images, offering a wealth of spatial and spectral information. However, the practical use of HSIs is often hindered by the presence of complicated noise caused by various factors such as non-uniform sensor response and dark current. Traditional methods for denoising HSIs rely on constrained optimization approaches, where selecting appropriate prior knowledge is critical for achieving satisfactory results. Nevertheless, these traditional algorithms are limited by hand-crafted priors, leaving room for improvement in their denoising performance. Recently, the supervised deep learning technique has emerged as a promising approach for HSI denoising. However, their requirement for paired training data and poor generalization ability on untrained noise distributions pose challenges in practical applications. In this paper, we design a novel algorithm by the synergism of optimization-based methods and deep learning techniques. Specifically, we introduce a plug-and-play Deep Low-rank Decomposition (DLD) model into the optimization framework. Furthermore, we propose an effective mechanism to incorporate traditional prior knowledge into the DLD model. Finally, we provide a detailed analysis of the optimization process and convergence of the proposed method. Empirical evaluations on various tasks, including hyperspectral image denoising and spectral compressive imaging, demonstrate the superiority of our approach over state-of-the-art methods. Yurong Chen 0003, Hui Zhang 0023, Yaonan Wang 0001, Yimin Yang 0001, Q. M. Jonathan Wu |
IEEE Trans. Image Process. | 1 |
| 2024 | R2D2-GAN: Robust Dual Discriminator Generative Adversarial Network for Microscopy Hyperspectral Image Super-ResolutionabstractHigh-resolution microscopy hyperspectral (HS) images can provide highly detailed spatial and spectral information, enabling the identification and analysis of biological tissues at a microscale level. Recently, significant efforts have been devoted to enhancing the resolution of HS images by leveraging high spatial resolution multispectral (MS) images. However, the inherent hardware constraints lead to a significant distribution gap between HS and MS images, posing challenges for image super-resolution within biomedical domains. This discrepancy may arise from various factors, including variations in camera imaging principles (e.g., snapshot and push-broom imaging), shooting positions, and the presence of noise interference. To address these challenges, we introduced a unique unsupervised super-resolution framework named R2D2-GAN. This framework utilizes a generative adversarial network (GAN) to efficiently merge the two data modalities and improve the resolution of microscopy HS images. Traditionally, supervised approaches have relied on intuitive and sensitive loss functions, such as mean squared error (MSE). Our method, trained in a real-world unsupervised setting, benefits from exploiting consistent information across the two modalities. It employs a game-theoretic strategy and dynamic adversarial loss, rather than relying solely on fixed training strategies for reconstruction loss. Furthermore, we have augmented our proposed model with a central consistency regularization (CCR) module, aiming to further enhance the robustness of the R2D2-GAN. Our experimental results show that the proposed method is accurate and robust for super-resolution images. We specifically tested our proposed method on both a real and a synthetic dataset, obtaining promising results in comparison to other state-of-the-art methods. Hui Zhang 0023, Jiang-Huai Tian, Yingjian Su, Yurong Chen 0003, Yaonan Wang 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2023 | Prior Image Guided Snapshot Compressive Spectral ImagingabstractSpectral images with rich spatial and spectral information have wide usage, however, traditional spectral imaging techniques undeniably take a long time to capture scenes. We consider the computational imaging problem of the snapshot spectral spectrometer, i.e., the Coded Aperture Snapshot Spectral Imaging (CASSI) system. For the sake of a fast and generalized reconstruction algorithm, we propose a prior image guidance-based snapshot compressive imaging method. Typically, the prior image denotes the RGB measurement captured by the additional uncoded panchromatic camera of the dual-camera CASSI system. We argue that the RGB image as a prior image can provide valuable semantic information. More importantly, we design the Prior Image Semantic Similarity (PIDS) regularization term to enhance the reconstructed spectral image fidelity. In particular, the PIDS is formulated as the difference between the total variation of the prior image and the recovered spectral image. Then, we solve the PIDS regularized reconstruction problem by the Alternating Direction Method of Multipliers (ADMM) optimization algorithm. Comprehensive experiments on various datasets demonstrate the superior performance of our method. Yurong Chen 0003, Yaonan Wang 0001, Hui Zhang 0023 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Adaptive Refining-Aggregation-Separation Framework for Unsupervised Domain Adaptation Semantic SegmentationabstractUnsupervised domain adaptation has attracted widespread attention as a promising method to solve the labeling difficulties of semantic segmentation tasks. It trains a segmentation network for unlabeled real target images using easily available labeled virtual source images. To improve performance, clustering is used to obtain domain-invariant feature representations. However, most clustering-based methods indiscriminately cluster all features mapped by category from both domains, causing the centroid shift and affecting the generation of discriminative features. We propose a novel clustering-based method that uses an adaptive refining-aggregation-separation framework, which learns the discriminative features by designing different adaptive schemes for different domains and features. The clustering does not require any tunable thresholds. To estimate more accurate domain-invariant centroids, we design different ways to guide the adaptive refinement of different domain features. A critic is proposed to directly evaluate the confidence of target features to solve the absence of target labels. We introduce a domain-balanced aggregation loss and two adaptive separation losses for distance and similarity respectively, which can discriminate clustering features by combining the refinement strategy to improve segmentation performance. Experimental results on GTA$5\rightarrow $Cityscapes and SYNTHIA$\rightarrow $Cityscapes benchmarks show that our method outperforms existing state-of-the-art methods. Yihong Cao, Hui Zhang 0023, Xiao Lu 0002, Yurong Chen 0003, Yaonan Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | D-BIN: A Generalized Disentangling Batch Instance Normalization for Domain AdaptationabstractPattern recognition is significantly challenging in real-world scenarios by the variability of visual statistics. Therefore, most existing algorithms relying on the independent identically distributed assumption of training and test data suffer from the poor generalization capability of inference on unseen testing datasets. Although numerous studies, including domain discriminator or domain-invariant feature learning, are proposed to alleviate this problem, the data-driven property and lack of interpretation of their principle throw researchers and developers off. Consequently, this dilemma incurs us to rethink the essence of networks' generalization. An observation that visual patterns cannot be discriminative after style transfer inspires us to take careful consideration of the importance of style features and content features. Does the style information related to the domain bias? How to effectively disentangle content and style features across domains? In this article, we first investigate the effect of feature normalization on domain adaptation. Based on it, we propose a novel normalization module to adaptively leverage the propagated information through each channel and batch of features called disentangling batch instance normalization (D-BIN). In this module, we explicitly explore domain-specific and domaininvariant feature disentanglement. We maneuver contrastive learning to encourage images with the same semantics from different domains to have similar content representations while having dissimilar style representations. Furthermore, we construct both self-form and dual-form regularizers for preserving the mutual information (MI) between feature representations of the normalization layer in order to compensate for the loss of discriminative information and effectively match the distributions across domains. D-BIN and the constrained term can be simply plugged into state-of-the-art (SOTA) networks to improve their performance. In the end, experiments, including domain adaptation and generalization, conducted on different datasets have proven their effectiveness. Yurong Chen 0003, Hui Zhang 0023, Yaonan Wang 0001, Weixing Peng, Wangdong Zhang, Q. M. Jonathan Wu, Yimin Yang 0001 |
IEEE Trans. Cybern. | 1 |
| 2023 | SSAPN: Spectral-Spatial Anomaly Perception Network for Unsupervised Vaccine DetectionabstractVaccines are the most significant and effective way to prevent disease and safeguard human health. However, it is easy to produce or mix foreign matters during the manufacturing process. Moreover, foreign matters are extremely faint that it is difficult to obtain images and detect them accurately. To tackle imaging challenges, in this article, we built a hyperspectral imaging system to construct a first-of-its-kind HSI dataset with pixel-level annotation for vaccine anomaly detection, where the vaccine comes from the actual pharmaceutical company. To address the problem of low detection accuracy, we propose a spectral–spatial anomaly perception module joint with an unsupervised autoencoder network (SSAPN), in which nonlinear features learned from the encoder are divided into nonoverlapping patches and mapped to efficiently encode spectral and spatial feature information. The spectral–spatial multilayer perceptrons (MLP) module consists of continuous and alternating spectral MLP with spatial MLP, which achieves spectral with spatial perception in the global receptive field, captures long-range dependencies, and extracts the most discriminative spectral–spatial features. Experimental results show that our SSAPN model outperforms other state-of-the-art anomaly detection methods in terms of both detection and generalization performance. This work will help speed up the production process in the vaccine pharmaceutical industry and ensure vaccine quality. Ating Yin, Yaonan Wang 0001, Yurong Chen 0003, Kai Zeng 0010, Hui Zhang 0023, Jianxu Mao |
IEEE Trans. Ind. Informatics | 3 |
| 2023 | Cycle Consistency Based Pseudo Label and Fine Alignment for Unsupervised Domain AdaptationabstractUnsupervised Domain Adaptation (UDA) aims to transfer knowledge from a well-labeled source domain to an unlabeled target domain with a correlative distribution. Numerous existing approaches process this hard nut by directly matching the marginal distribution between two domains, which confront the obstacle of rough alignment and blurred decision boundary. Recent advances in UDA introduce target pseudo-label and subdomain adaptation to reduce misalignment and distribution discrepancy. Whereas, they frequently ignore that the production of target pseudo-label is so dependent on the source-trained classifier, which without reasonable restriction to discriminate generated pseudo-label is whether confident. Meanwhile, many methods in the subdomain alignment metric ignore exploring the potential distribution discrepancy between same-class samples of the intra-domain. To address these two issues simultaneously, this paper proposes a Cycle Consistency based Pseudo Label and Fine Alignment (CCPLFA) approach for UDA. In particular, firstly, a novel cycle-consistency based pseudo label module is designed, which is a simple yet effective way to alleviate the noise of pseudo labels and improve their semantic correctness. Secondly, we develop a Fine-Alignment distribution matching metric. Which can maximize the feature distribution density of intra-class cross-domains and not overlook the distribution structure of the global aspect. Comprehensive experiment results on four benchmarks demonstrate the capability of plug and play and the well generalization performance of our proposed method. Hui Zhang 0023, Junkun Tang, Yihong Cao, Yurong Chen 0003, Yaonan Wang 0001, Q. M. Jonathan Wu |
IEEE Trans. Multim. | 4 |
| 2022 | Review on the COVID-19 pandemic prevention and control system based on AI
Junfei Yi, Hui Zhang 0023, Jianxu Mao, Yurong Chen 0003, Hang Zhong, Yaonan Wang 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2022 | MRSDI-CNN: Multi-Model Rail Surface Defect Inspection System Based on Convolutional Neural NetworksabstractDefects on rail surfaces, which have become critical problems, need to be detected and removed as quickly as possible to ensure the fast, safe, and stable operation of trains. At present, although many solutions have been proposed to address these problems, the comprehensiveness, rapidity, and accuracy of defect detection remain unsatisfactory. This study aims to resolve these existing problems and accordingly proposes a multi-model rail surface defect detection system based on convolutional neural networks (MRSDI-CNN) from the standpoint of studying the squat on the rail surface. The convolutional neural networks utilized include the improved Single Shot MultiBox Detector (SSD) and You Only Look Once version 3(YOLOv3)—two types of one-stage networks. We expounded and analyzed the performance of the convolutional neural networks as well as their applicability to rail surface defect detection. We used a diverse range of rail defect sizes to improve the detection performance of the two deep learning networks, following which they could identify three types of squats in parallel with improved accuracy and without reduction of the detection speed. The experimental results confirm the effectiveness and superiority of the proposed method over those of previous studies. Hui Zhang 0023, Yanan Song, Yurong Chen 0003, Hang Zhong, Li Liu 0060, Yaonan Wang 0001, Akilan Thangarajah, Q. M. Jonathan Wu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Semi-supervised Cloud Edge Collaborative Power Transmission Line Insulator Anomaly Detection Framework
Yanqing Yang, Jianxu Mao, Hui Zhang 0023, Yurong Chen 0003, Hang Zhong, Yaonan Wang 0001 |
ICIG (1) | 4 |
| 2021 | MAMA Net: Multi-Scale Attention Memory Autoencoder Network for Anomaly DetectionabstractAnomaly detection refers to the identification of cases that do not conform to the expected pattern, which takes a key role in diverse research areas and application domains. Most of existing methods can be summarized as anomaly object detection-based and reconstruction error-based techniques. However, due to the bottleneck of defining encompasses of real-world high-diversity outliers and inaccessible inference process, individually, most of them have not derived groundbreaking progress. To deal with those imperfectness, and motivated by memory-based decision-making and visual attention mechanism as a filter to select environmental information in human vision perceptual system, in this paper, we propose a Multi-scale Attention Memory with hash addressing Autoencoder network (MAMA Net) for anomaly detection. First, to overcome a battery of problems result from the restricted stationary receptive field of convolution operator, we coin the multi-scale global spatial attention block which can be straightforwardly plugged into any networks as sampling, upsampling and downsampling function. On account of its efficient features representation ability, networks can achieve competitive results with only several level blocks. Second, it's observed that traditional autoencoder can only learn an ambiguous model that also reconstructs anomalies "well" due to lack of constraints in training and inference process. To mitigate this challenge, we design a hash addressing memory module that proves abnormalities to produce higher reconstruction error for classification. In addition, we couple the mean square error (MSE) with Wasserstein loss to improve the encoding data distribution. Experiments on various datasets, including two different COVID-19 datasets and one brain MRI (RIDER) dataset prove the robustness and excellent generalization of the proposed MAMA Net. Yurong Chen 0003, Hui Zhang 0023, Yaonan Wang 0001, Yimin Yang 0001, Xianen Zhou, Q. M. Jonathan Wu |
IEEE Trans. Medical Imaging | 1 |