Yi Zhang 0018

dblp:64/6544-18 · DBLP profile ↗
← Back
87ranked-venue papers
3as first author
68since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 45 · 34 since 2021Artificial intelligence and machine learning · 25 · 3 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 20 since 2021Security and privacy · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Beyond Class Boundaries: Federated Visual Primitive Sharing with Text-Guided Adaptation
abstract
Personalized Federated Learning (pFL) effectively addresses the challenge of statistical heterogeneity in traditional Federated Learning (FL), with feature alignment methods (e.g., FedProto) standing out due to their communication efficiency and model-agnostic design, making them practically viable in real-world non-IID scenarios. These methods directly align class-level features across clients without requiring model parameter transmission. However, they represent each class as a holistic prototype, which limits the diversity and expressiveness of shared features. This restriction hampers the model's ability to generalize across clients and impedes personalized adaptation, as clients lack sufficient semantic components to reconstruct discriminative features tailored to their local data distributions. To overcome these limitations, we propose Federated Visual Primitive Learning (FedVPL), a novel framework comprising two key components: (1) Visual Primitive Space Sharing, which decomposes class-level features into semantically meaningful and reusable visual primitives, enabling cross-client and cross-class sharing to enrich feature diversity and decouple communication cost from the number of classes, significantly improving efficiency; and (2) Text-Guided Semantic Alignment, a parameter-free personalization mechanism that leverages external language priors to align shared primitives with client-specific semantics, without requiring additional communication overhead. Extensive experiments across diverse non-IID benchmarks demonstrate that FedVPL substantially outperforms state-of-the-art baselines, achieving up to a 5.62% improvement in accuracy, reducing communication overhead by at least 12.5x, and effectively addressing generalization and personalization challenges in heterogeneous federated environments.
Yongqiang Huang 0003, Tao Wang 0167, Zerui Shao, Beibei Li 0002, Yi Zhang 0018
WWW7
2026 Universal pre-training for generalizable incomplete-view CT reconstruction
Chenglong Ma 0002, Zilong Li 0001, Junjun He, Junping Zhang, Yi Zhang 0018, Hongming Shan
Pattern Recognit.5
2026 CPL-IQA: Blind image quality assessment via convolutional prototype learning
Hui Wang 0135, Guangcheng Wang, Ziyuan Yang 0001, Yi Zhang 0018
Pattern Recognit.5
2026 Enhancing federated learning through exploring filter-aware relationships and personalizing local structures
Ziyuan Yang 0001, Zerui Shao, Huijie Huangfu, Andrew Beng Jin Teoh, Hongming Shan, Yi Zhang 0018
Pattern Recognit.8
2026 Tuning Less, Learning More: Toward a More Generalizable Image Manipulation Detector
Ziyuan Yang 0001, Yi Zhang 0018
IEEE Signal Process. Lett.3
2026 Multi-Task Learning Network for Medical Image Analysis Guided by Lesion Regions and Spatial Relationships of Tissues
abstract
Medical image analysis plays key role in computer-aided diagnosis, where segmentation and classification are essential and interconnected tasks. While multi-task learning (MTL) has been widely explored to leverage inter-task synergies, effectively guiding knowledge transfer to prevent task conflict and negative transfer remains a key challenge, particularly in anatomically complex diagnostic scenarios. This paper presents LTRMTL-Net, a novel multi-task learning framework for medical image analysis that simultaneously addresses segmentation and classification tasks guided by lesion regions and spatial relationships of tissues. The proposed architecture integrates an Enhanced Lesion Region Fusion (ELRF) module that leverages GradCAM-guided attention mechanisms to precisely locate and enhance lesion regions, providing critical prior knowledge for both tasks. Tissue Space Structure Prediction (TSSP) component captures local-global spatial dependencies through contrastive learning, establishing effective anatomical context modeling. The core encoder employs Hybrid Wavelet-State Attention blocks that combine modulated wavelet transform convolutions with structured state space models to extract multi-scale features while maintaining computational efficiency. Dual-stream inputs with symmetric architecture accommodate single-source scenarios across diverse medical imaging applications. Experimental results on mammography and breast ultrasound datasets demonstrate that the proposed method captures fine-grained lesion boundary details while providing accurate malignancy classification. Harnessing cooperative knowledge transfer between segmentation and classification, guided by anatomical priors, boosts diagnostic performance and provides comprehensive, interpretable clinical insights.
Guowei Dai 0001, Duwei Dai, Chaoyu Wang 0001, Qingfeng Tang 0001, Hu Chen 0002, Yi Zhang 0018
IEEE Trans. Circuits Syst. Video Technol.7
2026 Dynamically Perceived Forgery Conditional Diffusion Model for Scientific Image Tampering Localization
abstract
Recently, image tampering localization techniques for scientific publications have attracted increasing attention due to the prevalence of data manipulation and the integrity issue of image content. However, existing methods are still inefficient to expose tampering traces in scientific images due to their unique properties, such as acquisition noise and ambiguous edges. To address these limitations, we propose a Dynamically Perceived Forgery Conditional Diffusion Model, which formulates the prediction of the localization mask as a noise-state aware denoising process. This process progressively localizes the tampered regions by involving time-step guidance to dynamically perceive tampering traces under the variation of diffusion noise, which is jointly controlled by two conditions, including a forgery condition with hierarchically aggregated forensic clues and an enhanced edge condition with multilevel spatial attention. To conduct dynamic controls efficiently, two conditions are fused and then applied to the denoising process via a channel-cross attention module. Furthermore, in the inference stage, a salient element ensemble-based sampling strategy is developed to further improve the reliability against undesired factors of scientific images. Extensive experiments have been conducted on several scientific image tampering datasets, compared with state-of-the-art methods, which demonstrates our superiority in aspects of intra-/cross-dataset evaluations and robustness against post-processing operations.
Jialing Xu, Peisong He, Haoliang Li, Shiqi Wang 0001, Yi Zhang 0018, Xinghao Jiang
IEEE Trans. Circuits Syst. Video Technol.5
2026 MedSAM-U: Uncertainty-Guided Auto Multi-Prompt Adaptation for Reliable MedSAM
abstract
The Medical Segment Anything Model (MedSAM) has demonstrated strong performance in medical image segmentation, attracting increasing attention in the medical imaging domain. However, as with many prompt-based segmentation models, its performance is highly sensitive to the type and location of input prompts. This sensitivity often leads to suboptimal segmentation outcomes and necessitates labor-intensive manual prompt tuning, which hampers both efficiency and robustness. To address this challenge, this paper proposes MedSAM-U, an uncertainty-guided framework designed to automatically refine prompt inputs and enhance segmentation reliability. Specifically, a Multi-Prompt Adapter is integrated into MedSAM, resulting in MPA-MedSAM, which enables the model to effectively accommodate diverse multi-prompt inputs. An uncertainty estimation module is then introduced to evaluate the reliability of the prompts and their initial segmentation results. Based on this, a novel uncertainty-guided prompt adaptation strategy is applied to automatically generate refined prompts and more accurate segmentation outputs. The proposed MedSAM-U framework is evaluated across multiple medical imaging modalities. Experimental results on five diverse datasets demonstrate that MedSAM-U achieves consistent performance improvements ranging from 1.7% to 20.5% over the baseline MedSAM, confirming its effectiveness and practicality for robust and efficient medical image segmentation.
Ke Zou, Mengting Luo, Linchao He, Meng Wang 0038, Yi Zhang 0018, Hu Chen 0002, Huazhu Fu
IEEE Trans. Circuits Syst. Video Technol.8
2026 FedPalm: A General Federated Learning Framework for Closed- and Open-Set Palmprint Verification
abstract
Current deep learning (DL)-based palmprint verification models rely on centralized training with large datasets, which raises significant privacy concerns due to the sensitive and immutable nature of biometric data. Federated learning (FL), a privacy-preserving distributed learning paradigm, offers a compelling alternative by enabling collaborative model training without the need for data sharing. However, FL-based palmprint verification faces critical challenges, including data heterogeneity from diverse identities and the absence of standardized evaluation benchmarks. This paper addresses these gaps by establishing a comprehensive benchmark for FL-based palmprint verification, which explicitly defines and evaluates two practical scenarios: closed-set and open-set verification. We propose FedPalm, a unified FL framework that balances local adaptability with global generalization. Each client trains a personalized textural expert tailored to local data and collaboratively contributes to a shared global textural expert for extracting generalized features. To further enhance verification performance, we introduce a Textural Expert Interaction Module that dynamically routes textural features among experts to generate refined side textural features. Learnable parameters are employed to model relationships between original and side features, fostering cross-texture-expert interaction and improving feature discrimination. Extensive experiments validate the effectiveness of FedPalm, demonstrating robust performance across both scenarios and providing a promising foundation for advancing FL-based palm-print verification research. The related code has been publicly available at https://github.com/Zi-YuanYang/FedPalm.
Ziyuan Yang 0001, Chengrui Gao, Andrew Beng Jin Teoh, Bob Zhang 0001, Yi Zhang 0018
IEEE Trans. Inf. Forensics Secur.6
2026 FoundDiff: Foundational Diffusion Model for Generalizable Low-Dose CT Denoising
abstract
Low-dose computed tomography (CT) denoising is crucial for reduced radiation exposure while ensuring diagnostically acceptable image quality. Despite significant advancements driven by deep learning (DL) in recent years, existing DL-based methods, typically trained on a specific dose level and anatomical region, struggle to handle diverse noise characteristics and anatomical heterogeneity during varied scanning conditions, limiting their generalizability and robustness in clinical scenarios. In this paper, we propose FoundDiff, a foundational diffusion model for unified and generalizable LDCT denoising across various dose levels and anatomical regions. FoundDiff employs a two-stage strategy: (i) dose-anatomy perception and (ii) adaptive denoising. First, we develop a dose- and anatomy-aware contrastive language-image pre-training model (DA-CLIP) to achieve robust dose and anatomy perception by leveraging specialized contrastive learning strategies to learn continuous representations that quantify ordinal dose variations and identify salient anatomical regions. Second, we design a dose- and anatomy-aware diffusion model (DA-Diff) to perform adaptive and generalizable denoising by synergistically integrating the learned dose and anatomy embeddings from DA-CLIP into diffusion process via a novel dose and anatomy conditional block (DACB) based on Mamba. Extensive experiments on a large simulated multi-dose CT dataset spanning three anatomical regions, together with cross-dataset evaluations on Mayo-2016, CQ500, and piglet datasets, demonstrate superior denoising performance and strong generalization to unseen dose levels and anatomical regions. The codes and models are available at https://github.com/hao1635/FoundDiff.
Zilong Li 0001, Junping Zhang, Yi Zhang 0018, Jun Zhao 0010, Hongming Shan
IEEE Trans. Medical Imaging5
2026 DeepSparse: A Foundation Model for Sparse-View CBCT Reconstruction
abstract
Cone-beam computed tomography (CBCT) is a critical 3D imaging technology in the medical field, while the high radiation exposure required for high-quality imaging raises significant concerns, particularly for vulnerable populations. Sparse-view reconstruction reduces radiation by using fewer X-ray projections while maintaining image quality, yet existing methods face challenges such as high computational demands and poor generalizability to different datasets. To overcome these limitations, we propose DeepSparse, the first foundation model for sparse-view CBCT reconstruction, featuring DiCE (Dual-Dimensional Cross-Scale Embedding), a novel network that integrates multi-view 2D features and multi-scale 3D features. Additionally, we introduce the HyViP (Hybrid View Sampling Pretraining) framework, which pretrains the model on large datasets with both sparse-view and dense-view projections, and a two-step finetuning strategy to adapt and refine the model for new datasets. Extensive experiments and ablation studies demonstrate that our proposed DeepSparse achieves superior reconstruction quality compared to state-of-the-art methods, paving the way for safer and more efficient CBCT imaging. The code will be publicly available at https://github.com/xmed-lab/DeepSparse.
Yiqun Lin, Jixiang Chen 0001, Hualiang Wang, Jiewen Yang, Jiarong Guo, Yi Zhang 0018, Xiaomeng Li 0001
IEEE Trans. Medical Imaging6
2026 SMART: Self-Supervised Learning for Metal Artifact Reduction in Computed Tomography Using Range Null Space Decomposition
abstract
Metal artifacts in computed tomography (CT) imaging significantly hinder diagnostic accuracy and clinical decision-making. While deep learning-based metal artifact reduction (MAR) methods have demonstrated promising progress, their clinical application is still constrained by three major challenges: 1) balancing metal artifact reduction with the preservation of critical anatomical structures, 2) effectively capturing the clinical priors of metal artifacts, and 3) dynamically adapting to polychromatic spectral variations. To address these limitations, in this paper, we propose a Self-supervised MAR method for computed Tomography (SMART) that leverages range-null space decomposition (RND) to model metal and tissue LACs separately, and employs implicit neural representation (INR) to learn their respective clinical characteristics without explicit supervision. Specifically, RND decouples metal and tissue LACs into a residual range component for metal LAC modeling, which captures metal artifacts, thus facilitating metal artifact reduction, and a null component for tissue LAC modeling, which focuses on preserving tissue details. To deal with the lack of paired data in clinical settings, we utilize INR to learn the clinical characteristics of these components in a self-supervised manner. Furthermore, SMART incorporates polychromatic spectra into the implicit representation, allowing dynamic adaptation to spectral variations across different imaging conditions. Extensive experiments on one synthetic and two clinical datasets demonstrate the strong potential of SMART in real-world scenarios. By flexibly adapting to spectral variations, it achieves superior generalizability to out-of-distribution clinical data.
Yanxin Cao, Yongqiang Huang 0003, Jingfeng Lu, Fenglei Fan, Hongming Shan, Yi Zhang 0018
IEEE Trans. Medical Imaging8
2026 UPGRADE-Net: Unsupervised Sinogram-Domain Data-Consistent Network for Metal Artifact Reduction
abstract
Computed tomography (CT) scanners are widely used to obtain detailed internal images in clinical diagnosis. Highly attenuated metallic implants resulting from strong and energy-dependent attenuation cause metal artifacts in CT scanning. However, current supervised deep network-based metal artifact reduction (MAR) methods hardly generalize in clinical diagnosis and treatment because of difficult acquisition for the paired artifact-affected and artifact-free data. In addition, these deep model-based methods cannot ensure the sinogram-domain data consistency for the exact metal trace inpainting. To address the above problems, we propose an UnsuPervised sinoGRam-domAin Data-consistEnt network for MAR, i.e., UPGRADE-Net. First, UPGRADE-Net fully leverages the prior knowledge to guide the generative conditional diffusion model for fine-grained metal trace inpainting. Second, without the artifact-free ground truth, a deep unsupervised MAR framework in the reverse process is constructed to contextually learn the known background data distribution for the unknown metal trace restoration in sinogram-domain. Third, to further maintain the sinogram-domain data consistency, two physics-based consistency constraint loss functions, including conjugate-ray and accumulation-ray consistency loss, are designed for the conjugate point constraint and the accumulation constraint. The proposed UPGRADE-Net is trained and evaluated on a publicly available dataset and a clinical dataset. Extensive experimental results validate that the proposed method outperforms the state-of-the-art competing methods for MAR.
Zhan Wu, Yikun Zhang 0001, Yongjie Guo, Huazhong Shu, Yan Xi, Yi Zhang 0018, Gouenou Coatrieux, Yang Chen 0008
IEEE Trans. Medical Imaging8
2025 Trustworthy Disentangled Framework for Multi-Label Medical Image Classification with Multimodal Refinement
abstract
Clinical practice reveals that patients frequently suffer from multiple co-occurring diseases, making multi-label classification (MLC) essential for accurate diagnosis. However, current MLC methods face two major challenges: (1) Disease-specific feature entanglement arising from the complex interdisease correlations among comorbidities; and (2) Untrustworthy results due to single-point estimates that lack confidence measurement. In this paper, we attempt to address these challenges at both the model and optimization levels. Specifically, at the model level, we introduce an improved transformer architecture with multi-CLS tokens for feature disentanglement. This architecture effectively captures the relationships among different diseases, while each CLS token integrates class-wise features, further refined by a multimodal method using a vision language model (VLM). At the optimization level, we propose a novel trustworthy MLC loss that aggregates positive/negative evidence for each class, modeling a multi-Beta distribution based on the Theory of Evidence, to generate reliable predictions with uncertainty estimations. Extensive experiments are conducted on publicly available clinical datasets, and the results demonstrate the effectiveness of our proposed method11The code is available at: https://github.com/CYYukio/Trustworthy-Disentangled-Framework..
Ziyuan Yang 0001, Yongqiang Huang 0003, Xulei Yang, Siyong Yeo, Yi Zhang 0018
BIBM6
2025 Patient-Level Anatomy Meets Scanning-Level Physics: Personalized Federated Low-Dose CT Denoising Empowered by Large Language Model
abstract
Reducing radiation doses benefits patients, but the resultant low-dose computed tomography (LDCT) images often suffer from clinically unacceptable noise and artifacts. While deep learning (DL) has shown promise in LDCT reconstruction, it requires large-scale data collection from multiple clients, raising privacy concerns. Federated learning (FL) has been introduced to mitigate these privacy concerns; however, current methods are typically tailored to specific scanning protocols, which limits their generalizability and makes them less effective for unseen protocols. To address these issues, we propose SCANPhysFed, a novel SCanning- and ANatomy-level personalized Physics-Driven Federated learning paradigm for LDCT reconstruction. Since the noise distribution in LDCT data is closely tied to scanning protocols and anatomical structures, we propose a dual-level physics-informed way to address these challenges. Specifically, we incorporate physical and anatomical prompts into our physics-informed hypernetworks to capture scanning- and anatomy-specific information, enabling dual-level physics-driven personalization of imaging features. These prompts are derived from the scanning protocol and the radiology report generated by a medical large language model (MLLM). Subsequently, client-specific decoders project these dual-level personalized imaging features back into the image domain. Besides, to tackle the challenge of unseen data, we introduce a novel protocol vector-quantization strategy (PVQS), which ensures consistent performance across new clients by quantifying unseen scanning codes to the closest match in the scanning codebook. Extensive experimental results demonstrate the superior performance of SCAN-PhysFed on public datasets1.
Ziyuan Yang 0001, Zhiwen Wang 0002, Hongming Shan, Yang Chen 0008, Yi Zhang 0018
CVPR6
2025 Modality Modulation and Dual Consistency for Multi-Modality Semi-Supervised Medical Image Segmentation
abstract
Multi-modality (MM) semi-supervised learning (SSL) based medical image segmentation has recently gained increasing attention due to its ability to utilize MM data and low dependency on labeled images. However, current MM-SSL methods face two major challenges: (1) Complex network designs make it difficult to apply these methods to scenarios involving more than two modalities. (2) The use of generative methods to leverage unlabeled data may not be reliable for SSL learning. To address these challenges, we propose Modality Modulation Dual Consistency, dubbed MM-DC. Specifically, we design a modality all-in-one network to process data from all modalities, with learnable plug-in Modality Modulation Layers (MML) to gradually modulate features from different modalities into a modality-invariant feature space, enabling unified segmentation. Additionally, we propose a dual-consistency strategy that enforces consistency at both the image and feature levels, which eliminates the requirements for generative methods. Extensive experiments demonstrate that MM-DC outperforms other state-of-the-art methods on open-source datasets with 2- and 4-modalities. The code is available1.
Deng Xiong, Yi Zhang 0018
ICASSP4
2025 A Unified Multimodal Multi-Granularity Pre-Training Framework for Fine-Grained Medical Image Analysis
abstract
Medical image diagnosis relies heavily on subtle, fine-grained details, making it crucial to integrate global imaging representations with localized pathological information. Existing vision-language pretraining models have shown promise in utilizing free-text radiology reports for deeper semantic insights; however, they face significant challenges in precisely capturing and aligning region-specific, fine-grained information, which often results in suboptimal feature extraction for localized pathologies and reduced model interpretability. In this paper, we propose a unified multimodal multi-granularity pre-training framework (UMMPF) that explicitly models the local anatomical regions in chest X-ray images through a Region Querying Module (RQM), thereby establishing stronger correspondences between image sub-regions and textual descriptions. We further incorporate cross-modal bidirectional attention and unified multiple training objectives, facilitating the interaction of both global and local features across modalities. Our method enhances interpretability by explicitly mapping pathological findings to anatomical regions. Experimental results on benchmark datasets (RSNA, SIIM, and ChestX-ray14) demonstrate a significant improvement in diagnostic accuracy.
Zenan Gong, Linchao He, Peixi Liao, Hongjie Yang, Hu Chen 0002, Yi Zhang 0018
IJCNN8
2025 FedRIR: Rethinking Information Representation in Federated Learning
abstract
Mobile and Web-of-Things (WoT) devices at the network edge generate vast amounts of data for machine learning applications, yet privacy concerns hinder centralized model training. Federated Learning (FL) allows clients (devices) to collaboratively train a shared model coordinated by a central server without transferring private data. However, inherent statistical heterogeneity among clients presents challenges, often leading to a dilemma between clients' need for personalized local models and the server's goal of building a generalized global model. Existing FL methods typically prioritize either global generalization or local personalization, resulting in a trade-off between these objectives and limiting the full potential of diverse client data. To address this challenge, we propose a novel framework that enhances both global generalization and local personalization by Rethinking Information Representation in the Federated learning process (FedRIR). Specifically, we introduce Masked Client-Specific Learning (MCSL), which isolates and extracts fine-grained client-specific features tailored to each client's unique data characteristics, thereby enhancing personalization. Meanwhile, the Information Distillation Module (IDM) refines global shared features by filtering out redundant client-specific information, resulting in a purer and more robust global representation that enhances generalization. By integrating refined global features with isolated client-specific features, we construct enriched representations that effectively capture both global patterns and local nuances, thereby improving the performance of downstream tasks on the client. Extensive experiments on diverse datasets demonstrate that FedRIR significantly outperforms state-of-the-art FL methods, achieving up to a 3.93% improvement in accuracy while ensuring robustness and stability in heterogeneous environments. The code is publicly available at https://github.com/Deep-Imaging-Group/FedRIR.
Yongqiang Huang 0003, Zerui Shao, Ziyuan Yang 0001, Yi Zhang 0018
WWW5
2025 Lightweight Vision Transformer With Lite-AVPSO Hyperparameter Optimization for Agricultural Disease Recognition
abstract
Plant disease identification and management are vital for crop protection, productivity, and sustainable agriculture. This paper presents RepAgrViT, a lightweight vision transformer architecture for agricultural disease recognition in Internet of Things edge environments. The proposed model integrates CNN with vision transformer principles through a novel dual-stream design comprising token mixer and channel mixer arranged in series. A Bilinear Attention Transformation module performs extensive local-global attention via direct feature position connections. This enables comprehensive capture of long-range dependencies between healthy and diseased leaf regions. We introduce an efficient parameter fusion methodology that optimizes depthwise separable convolutions while integrating normalized weights with biases, reducing model complexity. Furthermore, we propose Lite-AVPSO, a hyperparameter optimization algorithm for precise model configuration. It incorporates adaptive weighted delayed velocity optimization and neighborhood-based local search strategies. Visualization techniques using multi-channel color heatmaps and three-dimensional feature spaces provide interpretable diagnostic results. Experiments across diverse plant disease datasets demonstrate RepAgrViT efficacy in capturing discriminative disease features while maintaining computational efficiency suitable for resource-constrained agricultural monitoring systems, establishing a foundation for efficient vision-based monitoring in real-time agricultural applications.
Guowei Dai 0001, Zhimin Tian, Chaoyu Wang 0001, Qingfeng Tang 0001, Hu Chen 0002, Yi Zhang 0018
IEEE Internet Things J.6
2025 ROPREM: a prototype-free retrival method for automatic check-out
Huijie Huangfu, Maosong Ran, Jingfeng Lu, Yi Zhang 0018
Neural Comput. Appl.6
2025 Beyond Static Features: A Novel Dynamic Palmprint Verification Framework Empowered by Generative Models
abstract
Palmprint recognition has received considerable attention due to its inherent discriminative characteristics. However, conventional methods largely rely on static features extracted from individual images, which limits their representational richness. To address this, we propose a dynamic palmprint verification framework that harnesses generative models to enhance feature representations through dynamic construction and matching strategies. During training, a classifier-guided generative model synthesizes class-aware pairs, and a regularization term is introduced to expand the feature space, while mitigating overfitting. For matching, we reformulate the process as a subspace projection within a locally adaptive feature space, where the original and class-conditioned generated features form the basis of the subspace. This enables the model to capture latent inter-individual relationships and achieve stronger discriminability. Extensive experiments across multiple backbones and public benchmarks validate the effectiveness and robustness of the proposed framework.
Ziyuan Yang 0001, Lu Leng, Andrew Beng Jin Teoh, Bob Zhang 0001, Yi Zhang 0018
IEEE Signal Process. Lett.5
2025 Boosting Geometric Invariants for Discriminative Forensics of Large-Scale Generated Visual Content
abstract
Generative artificial intelligence has shown great success in visual content synthesis such that humans struggle to distinguish between real and synthesized images. Forensic research seeks to reveal artifacts in such generated images, ensuring information security or improving generation capability. In this regard, the robustness and interpretability are important for the trustworthy purpose of forensic tasks. However, typical forensic models and their underlying data representations rely on empirical learning algorithms, which cannot effectively handle the high robustness and interpretability requirements beyond experience. As an effective solution, we extend the classical geometric invariants to the forensic research of large-scale generated images. Invariants are handcrafted representations with robust and interpretable geometric principles. However, their discriminability is far from the large scale of today's forensic tasks. We boost the discriminability by extending the classical invariants to the hierarchical architecture of convolutional neural networks. The resulting overcompleteness allows for an automatic selection of task-discriminative features, while retaining the previous advantages of robustness and interpretability. From generative adversarial networks to diffusion models, the forensic with our boosted invariants demonstrates state-of-the-art discriminability against large-scale content diversity. It also exhibits high efficiency on training examples, intrinsic invariance to geometric variations, and better interpretability of the forensic process.
Chao Wang 0028, Yushu Zhang 0001, Xiangyu Chen 0006, Yi Zhang 0018, Tieyong Zeng, Fenglei Fan
IEEE Trans. Image Process.6
2025 Solving Zero-Shot Sparse-View CT Reconstruction With Variational Score Solver
abstract
Computed tomography (CT) stands as a ubiquitous medical diagnostic tool. Nonetheless, the radiation-related concerns associated with CT scans have raised public apprehensions. Mitigating radiation dosage in CT imaging poses an inherent challenge as it inevitably compromises the fidelity of CT reconstructions, impacting diagnostic accuracy. While previous deep learning techniques have exhibited promise in enhancing CT reconstruction quality, they remain hindered by the reliance on paired data, which is arduous to procure. In this study, we present a novel approach named Variational Score Solver (VSS) for sparse-view reconstruction without paired data. Our approach entails the acquisition of a probability distribution from densely sampled CT reconstructions, employing a latent diffusion model. High-quality reconstruction outcomes are achieved through an iterative process, wherein the diffusion model serves as the prior term, subsequently integrated with the data consistency term. Notably, rather than directly employing the prior diffusion model, we distill prior knowledge by finding the fixed point of the diffusion model. This framework empowers us to exercise precise control over the process. Moreover, we depart from modeling the reconstruction outcomes as deterministic values, opting instead for a distribution-based approach. This enables us to achieve more accurate reconstructions utilizing a trainable model. Our approach introduces a fresh perspective to the realm of zero-shot CT reconstruction, circumventing the constraints of supervised learning. Extensive qualitative and quantitative experiments unequivocally demonstrate that VSS surpasses other contemporary unsupervised and achieves comparable results compared to the most advanced supervised methods in sparse-view reconstruction tasks. Codes are available in https://github.com/fpsandnoob/vss.
Linchao He, Wenchao Du, Peixi Liao, Fenglei Fan, Hu Chen 0002, Hongyu Yang 0002, Yi Zhang 0018
IEEE Trans. Medical Imaging7
2025 Bi-Constraints Diffusion: A Conditional Diffusion Model With Degradation Guidance for Metal Artifact Reduction
abstract
In recent years, score-based diffusion models have emerged as effective tools for estimating score functions from empirical data distributions, particularly in integrating implicit priors with inverse problems like CT reconstruction. However, score-based diffusion models are rarely explored in challenging tasks such as metal artifact reduction (MAR). In this paper, we introduce a Bi-Constraints Diffusion Model for Metal Artifact Reduction (BCDMAR), an innovative approach that enhances iterative reconstruction with a conditional diffusion model for MAR. This method employs a metal artifact degradation operator in place of the traditional metal-excluded projection operator in the data-fidelity term, thereby preserving structure details around metal regions. However, score-based diffusion models tend to be susceptible to grayscale shifts and unreliable structures, making it challenging to reach an optimal solution. To address this, we utilize a pre-corrected image as a prior constraint, guiding the generation of the score-based diffusion model. By iteratively applying the score-based diffusion model and the data-fidelity step in each sampling iteration, BCDMAR effectively maintains reliable tissue representation around metal regions and produces highly consistent structures in non-metal regions. Through extensive experiments focused on metal artifact reduction tasks, BCDMAR demonstrates superior performance over other state-of-the-art unsupervised and supervised methods, both quantitatively and qualitatively.
Mengting Luo, Tao Wang 0167, Linchao He, Wang Wang, Hu Chen 0002, Peixi Liao, Yi Zhang 0018
IEEE Trans. Medical Imaging8
2025 Radiologist-in-the-Loop Self-Training for Generalizable CT Metal Artifact Reduction
abstract
Metal artifacts in computed tomography (CT) images can significantly degrade image quality and impede accurate diagnosis. Supervised metal artifact reduction (MAR) methods, trained using simulated datasets, often struggle to perform well on real clinical CT images due to a substantial domain gap. Although state-of-the-art semi-supervised methods use pseudo ground-truths generated by a prior network to mitigate this issue, their reliance on a fixed prior limits both the quality and quantity of these pseudo ground-truths, introducing confirmation bias and reducing clinical applicability. To address these limitations, we propose a novel radiologist-in-the-loop self-training framework for MAR, termed RISE-MAR, which can integrate radiologists' feedback into the semi-supervised learning process, progressively improving the quality and quantity of pseudo ground-truths for enhanced generalization on real clinical CT images. For quality assurance, we introduce a clinical quality assessor model that emulates radiologist evaluations, effectively selecting high-quality pseudo ground-truths for semi-supervised training. For quantity assurance, our self-training framework iteratively generates additional high-quality pseudo ground-truths, expanding the clinical dataset and further improving model generalization. Extensive experimental results on multiple clinical datasets demonstrate the superior generalization performance of our RISE-MAR over state-of-the-art methods, advancing the development of MAR models for practical application. The source code is available at https://github.com/Masaaki-75/rise-mar.
Chenglong Ma 0002, Zilong Li 0001, Junping Zhang, Yi Zhang 0018, Jiannan Liu, Hongming Shan
IEEE Trans. Medical Imaging6
2025 UniAda: Domain Unifying and Adapting Network for Generalizable Medical Image Segmentation
abstract
Learning a generalizable medical image segmentation model is an important but challenging task since the unseen (testing) domains may have significant discrepancies from seen (training) domains due to different vendors and scanning protocols. Existing segmentation methods, typically built upon domain generalization (DG), aim to learn multi-source domain-invariant features through data or feature augmentation techniques, but the resulting models either fail to characterize global domains during training or cannot sense unseen domain information during testing. To tackle these challenges, we propose a domain Unifying and Adapting network (UniAda) for generalizable medical image segmentation, a novel "unifying while training, adapting while testing" paradigm that can learn a domain-aware base model during training and dynamically adapt it to unseen target domains during testing. First, we propose to unify the multi-source domains into a global inter-source domain via a novel feature statistics update mechanism, which can sample new features for the unseen domains, facilitating the training of a domain base model. Second, we leverage the uncertainty map to guide the adaptation of the trained model for each testing sample, considering the specific target domain may be outside the global inter-source domain. Extensive experimental results on two public cross-domain medical datasets and one in-house cross-domain dataset demonstrate the strong generalization capacity of the proposed UniAda over state-of-the-art DG methods. The source code of our UniAda is available at https://github.com/ZhouZhang233/UniAda.
Zhongzhou Zhang, Zhiwen Wang 0002, Shanshan Wang 0008, Fenglei Fan, Hongming Shan, Yi Zhang 0018
IEEE Trans. Medical Imaging8
2025 Hypernetwork-Based Physics-Driven Personalized Federated Learning for CT Imaging
abstract
In clinical practice, computed tomography (CT) is an important noninvasive inspection technology to provide patients' anatomical information. However, its potential radiation risk is an unavoidable problem that raises people's concerns. Recently, deep learning (DL)-based methods have achieved promising results in CT reconstruction, but these methods usually require the centralized collection of large amounts of data for training from specific scanning protocols, which leads to serious domain shift and privacy concerns. To relieve these problems, in this article, we propose a hypernetwork-based physics-driven personalized federated learning method (HyperFed) for CT imaging. The basic assumption of the proposed HyperFed is that the optimization problem for each domain can be divided into two subproblems: local data adaption and global CT imaging problems, which are implemented by an institution-specific physics-driven hypernetwork and a global-sharing imaging network, respectively. Learning stable and effective invariant features from different data distributions is the main purpose of global-sharing imaging network. Inspired by the physical process of CT imaging, we carefully design physics-driven hypernetwork for each domain to obtain hyperparameters from specific physical scanning protocol to condition the global-sharing imaging network, so that we can achieve personalized local CT reconstruction. Experiments show that HyperFed achieves competitive performance in comparison with several other state-of-the-art methods. It is believed as a promising direction to improve CT imaging quality and personalize the needs of different institutions or scanners without data sharing. Related codes have been released at https://github.com/Zi-YuanYang/HyperFed.
Ziyuan Yang 0001, Wenjun Xia, Xiaoxiao Li 0001, Yi Zhang 0018
IEEE Trans. Neural Networks Learn. Syst.6
2025 FedLoRE: Communication-Efficient and Personalized Edge Intelligence Framework via Federated Low-Rank Estimation
abstract
Federated learning (FL) has recently garnered significant attention in edge intelligence. However, FL faces two major challenges: First, statistical heterogeneity can adversely impact the performance of the global model on each client. Second, the model transmission between server and clients leads to substantial communication overhead. Previous works often suffer from the trade-off issue between these seemingly competing goals, yet we show that it is possible to address both challenges simultaneously. We propose a novel communication-efficient personalized FL framework for edge intelligence that estimates the low-rank component of the training model gradient and stores the residual component at each client. The low-rank components obtained across communication rounds have high similarity, and sharing these components with the server can significantly reduce communication overhead. Specifically, we highlight the importance of previously neglected residual components in tackling statistical heterogeneity, and retaining them locally for training model updates can effectively improve the personalization performance. Moreover, we provide a theoretical analysis of the convergence guarantee of our framework. Extensive experimental results demonstrate that our framework outperforms state-of-the-art approaches, achieving up to 89.18% reduction in communication overhead and 91.00% reduction in computation overhead while maintaining comparable personalization accuracy compared to previous works.
Zerui Shao, Beibei Li 0002, Peiran Wang, Yi Zhang 0018, Kim-Kwang Raymond Choo
IEEE Trans. Parallel Distributed Syst.4
2024 Point, Segment and Count: A Generalized Framework for Object Counting
abstract
Class-agnostic object counting aims to count all objects in an image with respect to example boxes or class names, a.k.a few-shot and zero-shot counting. In this paper, we propose a generalized framework for both few-shot and zero-shot object counting based on detection. Our framework combines the superior advantages of two foundation models without compromising their zero-shot capability: (i) SAM to segment all possible objects as mask proposals, and (ii) CLIP to classify proposals to obtain accurate object counts. However, this strategy meets the obstacles of efficiency over-head and the small crowded objects that cannot be localized and distinguished. To address these issues, our framework, termed PseCo, follows three steps: point, segment, and count. Specifically, we first propose a class-agnostic object localization to provide accurate but least point prompts for SAM, which consequently not only reduces computation costs but also avoids missing small objects. Furthermore, we propose a generalized object classification that leverages CLIP image/text embeddings as the classifier, following a hierarhical knowledge distillation to obtain discriminative classifications among hierarchical mask proposals. Extensive experimental results on FSC-147, COCO, and LVIS demonstrate that PseCo achieves state-of-the-art performance in both few-shot/zero-shot object counting/detection.
Zhizhong Huang, Mingliang Dai, Yi Zhang 0018, Junping Zhang, Hongming Shan
CVPR3
2024 Data Currency Quality Assessment Based on Multi-sensor
Zhaoxin Zhu, Xuanzhi Feng, Dongxu Fan, Yi Zhang 0018, Dasha Hu, Xue-Feng Ding 0003, Yuming Jiang 0004
ICIC (13)4
2024 Physics-Driven Spectrum-Consistent Federated Learning for Palmprint Verification
Ziyuan Yang 0001, Andrew Beng Jin Teoh, Bob Zhang 0001, Lu Leng, Yi Zhang 0018
Int. J. Comput. Vis.5
2024 Progressive dual-domain-transfer cycleGAN for unsupervised MRI reconstruction
Zhiwen Wang 0002, Ziyuan Yang 0001, Wenjun Xia, Yi Zhang 0018
Neurocomputing5
2024 Hierarchical disentangled representation for image denoising and beyond
Wenchao Du, Hu Chen 0002, Yi Zhang 0018, Hongyu Yang 0002
Image Vis. Comput.3
2024 Generalizable MRI Motion Correction via Compressed Sensing Equivariant Imaging Prior
abstract
Existing deep learning (DL)-based magnetic resonance imaging (MRI) retrospective motion correction (MoCo) models are typically task-specific, which makes them challenging to generalize to different scenarios w.r.t motions, modalities, planes, and scanner centers. This limitation occurs since the motions of each patient vary, and collecting diverse paired/unpaired motion data is generally costly and infeasible. To deal with this problem, we propose the Equivariant Imaging Prior (EIP) framework to generalize the MoCo tasks toward various scenarios.In this paper, the traditional MRI MoCo tasks, specifically for the multi-scenarios, can be treated as a mask-varying compressed sensing self-supervised problem for MRI reconstruction with corrupted k-space data.To the best of our knowledge, this framework is the first attempt to handle multiple MRI MoCo scenarios with one single DL model. Specifically, stochastic subsampling and modality augmentation are employed for data preparation. Then, a domain generalization-friendly net is carefully designed and an equivariant imaging task is leveraged to learn the mapping from corrupted data to clean images. The experimental results show that the proposed EIP framework achieves impressive adaptability across generalizable MoCo tasks, including but not limited to multi-motion, multi-modality, multi-center, and multi-plane. Furthermore, our EIP demonstrates similar or superior performance to several state-of-the-art models trained in a supervised manner, extending to even motion estimation on the multi-coil raw data. The code is available:https://github.com/wangzhiwen-scu/EIP4MoCo.
Zhiwen Wang 0002, Maosong Ran, Ziyuan Yang 0001, Jie Jing 0001, Tao Wang 0167, Jingfeng Lu, Yi Zhang 0018
IEEE Trans. Circuits Syst. Video Technol.8
2024 A Dual-Level Cancelable Framework for Palmprint Verification and Hack-Proof Data Storage
abstract
In recent years, palmprints have been extensively utilized for individual verification. The abundance of sensitive information in palmprint data necessitates robust protection to ensure security and privacy without compromising system performance. Existing systems frequently use cancelable transformations to protect palmprint templates. However, if an adversary gains access to the stored database, they could initiate a replay attack before the system detects the breach and can revoke and replace the reference template. To address replay attacks while meeting template protection criteria, we propose a dual-level cancelable palmprint verification framework. In this framework, the reference template is initially transformed using a cancelable competition hashing network with a first-level token, enabling the end-to-end generation of cancelable templates. During enrollment, the system creates a negative database (NDB) using a second-level token for further protection. Due to the unique NDB-to-vector matching characteristic, a replay attack involving the matching between the reference template and a compromised instance in NDB form is infeasible. This approach effectively addresses the replay attack problem at its root. Furthermore, the dual-level protected reference template enjoys heightened security, as reversing the NDB is NP-hard. We also propose a novel NDB-to-vector matching algorithm based on matrix operations to expedite the matching process, addressing the inefficiencies of previous NDB methods reliant on dictionary-based matching rules. Extensive experiments conducted on public palmprint datasets confirm the effectiveness and generality of the proposed framework. Upon acceptance of the paper, the code will be accessible athttps://github.com/Zi-YuanYang/DCPV.
Ziyuan Yang 0001, Ming Kang 0007, Andrew Beng Jin Teoh, Chengrui Gao, Bob Zhang 0001, Yi Zhang 0018
IEEE Trans. Inf. Forensics Secur.7
2024 Gradient-Guided Network With Fourier Enhancement for Glioma Segmentation in Multimodal 3D MRI
abstract
Glioma segmentation is a crucial task in computer-aided diagnosis, requiring precise discrimination between lesions and normal tissue at the pixel level. Popular methods neglect crucial edge information, leading to inaccurate contour delineation. Moreover, global information has been proven beneficial for segmentation. The feature representations extracted by convolution neural networks often struggle with local-related information owing to the limited receptive fields. To address these issues, we propose a novel edge-aware segmentation network that incorporates a dual-path gradient-guided training strategy with Fourier edge-enhancement for precise glioma segmentation, a.k.a. GFNet. First, we introduce a Dual-path Gradient-guided Training strategy (DGT) based on a Siamese network guiding the optimizing direction of one path by the gradient from the other path. DGT pays attention to the indistinguishable pixels with large weight-updating gradient, such as the pixels near the boundary, to guide the network training, addressing hard samples. Second, to further perceive the edge information, we derive a Fourier Edge-enhancement Module (FEM) to augment feature edges with high-frequency representations from the spectral domain, providing global information and edge details. Extensive experiments on public glioma segmentation datasets, BraTS2020 and Medical Segmentation Decathlon (MSD) glioma and prostate segmentation, demonstrate that GFNet achieves competitive performance compared to other state-of-the-art methods, both qualitatively and quantitatively.
Zhongzhou Zhang, Zhongxian Wang, Zhiwen Wang 0002, Jingfeng Lu, Yan Liu 0052, Yi Zhang 0018
IEEE J. Biomed. Health Informatics7
2024 CoreDiff: Contextual Error-Modulated Generalized Diffusion Model for Low-Dose CT Denoising and Generalization
abstract
Low-dose computed tomography (CT) images suffer from noise and artifacts due to photon starvation and electronic noise. Recently, some works have attempted to use diffusion models to address the over-smoothness and training instability encountered by previous deep-learning-based denoising models. However, diffusion models suffer from long inference time due to a large number of sampling steps involved. Very recently, cold diffusion model generalizes classical diffusion models and has greater flexibility. Inspired by cold diffusion, this paper presents a novel COntextual eRror-modulated gEneralized Diffusion model for low-dose CT (LDCT) denoising, termed CoreDiff. First, CoreDiff utilizes LDCT images to displace the random Gaussian noise and employs a novel mean-preserving degradation operator to mimic the physical process of CT degradation, significantly reducing sampling steps thanks to the informative LDCT images as the starting point of the sampling process. Second, to alleviate the error accumulation problem caused by the imperfect restoration operator in the sampling process, we propose a novel ContextuaL Error-modulAted Restoration Network (CLEAR-Net), which can leverage contextual information to constrain the sampling process from structural distortion and modulate time step embedding features for better alignment with the input at the next time step. Third, to rapidly generalize the trained model to a new, unseen dose level with as few resources as possible, we devise a one-shot learning framework to make CoreDiff generalize faster and better using only one single LDCT image (un)paired with normal-dose CT (NDCT). Extensive experimental results on four datasets demonstrate that our CoreDiff outperforms competing methods in denoising and generalization performance, with clinically acceptable inference time. Source code is made available at https://github.com/qgao21/CoreDiff.
Zilong Li 0001, Junping Zhang, Yi Zhang 0018, Hongming Shan
IEEE Trans. Medical Imaging4
2024 SOUL-Net: A Sparse and Low-Rank Unrolling Network for Spectral CT Image Reconstruction
abstract
Spectral computed tomography (CT) is an emerging technology, that generates a multienergy attenuation map for the interior of an object and extends the traditional image volume into a 4-D form. Compared with traditional CT based on energy-integrating detectors, spectral CT can make full use of spectral information, resulting in high resolution and providing accurate material quantification. Numerous model-based iterative reconstruction methods have been proposed for spectral CT reconstruction. However, these methods usually suffer from difficulties such as laborious parameter selection and expensive computational costs. In addition, due to the image similarity of different energy bins, spectral CT usually implies a strong low-rank prior, which has been widely adopted in current iterative reconstruction models. Singular value thresholding (SVT) is an effective algorithm to solve the low-rank constrained model. However, the SVT method requires a manual selection of thresholds, which may lead to suboptimal results. To relieve these problems, in this article, we propose a sparse and low-rank unrolling network (SOUL-Net) for spectral CT image reconstruction, that learns the parameters and thresholds in a data-driven manner. Furthermore, a Taylor expansion-based neural network backpropagation method is introduced to improve the numerical stability. The qualitative and quantitative results demonstrate that the proposed method outperforms several representative state-of-the-art algorithms in terms of detail preservation and artifact reduction.
Xiang Chen 0015, Wenjun Xia, Ziyuan Yang 0001, Hu Chen 0002, Yan Liu 0052, Jiliu Zhou, Yang Chen 0008, Bihan Wen, Yi Zhang 0018
IEEE Trans. Neural Networks Learn. Syst.10
2023 Stay In The Middle: A Semi-Supervised Model for CT Metal Artifact Reduction
abstract
Metal artifacts degrade CT image’s quality. Recently, some deep learning-based metal artifact reduction (MAR) methods have been developed. Supervised MAR methods don’t perform well in clinical due to the domain gap between simulated and clinical data. Although this problem can be avoided in an unsupervised way, severe artifacts cannot be well suppressed. Semi-supervised MAR methods can alleviate the domain gap problem. However, the existing ones are usually accompanied by boosted model scale, which is challenging for optimization. In this paper, we propose a novel semi-supervised framework for MAR, termed SemiMAR. First, we only use one generator to learn the clean part, instead of multiple encoders and decoders disentangling artifacts. Thus, the model naturally becomes much smaller. To recover more tissue details, the advanced dual-domain MAR network knowledge is distilled into our model in both the image domain and latent feature space. Extensive experiments demonstrate the efficiency and robustness of our model.
Tao Wang 0167, Zhongzhou Zhang, Jiliu Zhou, Yi Zhang 0018
ICASSP6
2023 SynFacePAD 2023: Competition on Face Presentation Attack Detection Based on Privacy-aware Synthetic Training Data
abstract
This paper presents a summary of the Competition on Face Presentation Attack Detection Based on Privacy-aware Synthetic Training Data (SynFacePAD 2023) held at the 2023 International Joint Conference on Biometrics (IJCB 2023). The competition attracted a total of 8 participating teams with valid submissions from academia and industry. The competition aimed to motivate and attract solutions that target detecting face presentation attacks while considering synthetic-based training data motivated by privacy, legal and ethical concerns associated with personal data. To achieve that, the training data used by the participants was limited to synthetic data provided by the organizers. The submitted solutions presented innovations and novel approaches that led to outperforming the considered baseline in the investigated benchmarks.
Meiling Fang, Marco Huber, Julian Fierrez, Ramachandra Raghavendra, Naser Damer, Alhasan Alkhaddour, Maksim Kasantcev, Vasiliy Pryadchenko, Ziyuan Yang 0001, Huijie Huangfu, Yi Zhang 0018, Junjun Jiang, Xianming Liu 0005, Xianyun Sun, Caiyong Wang, Zhaohua Chang, Guangzhe Zhao, Juan E. Tapia, Lázaro J. González Soler, Carlos M. Aravena, Daniel Schulz
IJCB12
2023 A Robust Prototype-Free Retrieval Method for Automatic Check-Out
abstract
In recent years, automatic check-out (ACO) gains increasing interest and has been widely used in daily life. However, current works mainly rely on both counter and product prototype images in the training phase, and it is hard to maintain the performance in an incremental setting. To deal with this problem, we propose a robust prototype-free retrieval method (ROPREM) for ACO, which is a cascaded framework composed of a product detector module and a product retrieval module. We use the product detector module without product class information to locate products. Additionally, we first attempt the check-out process as a retrieval process rather than a classification process. The retrieval result is considered as the product class by comparing the feature similarity between a query image and gallery templates. As a result, our method require much fewer training samples and achieves state-of-the-art performance on the public Retail Product Checkout (RPC) dataset.
Huijie Huangfu, Ziyuan Yang 0001, Maosong Ran, Jingfeng Lu, Yi Zhang 0018
ISCC6
2023 ASCON: Anatomy-Aware Supervised Contrastive Learning Framework for Low-Dose CT Denoising
Yi Zhang 0018, Hongming Shan
MICCAI (10)3
2023 Inter-slice Consistency for Unpaired Low-Dose CT Denoising Using Boosted Contrastive Learning
Jie Jing 0001, Tao Wang 0167, Yi Zhang 0018
MICCAI (1)5
2023 FreeSeed: Frequency-Band-Aware and Self-guided Network for Sparse-View CT Reconstruction
Chenglong Ma 0002, Zilong Li 0001, Junping Zhang, Yi Zhang 0018, Hongming Shan
MICCAI (10)4
2023 PET Image Denoising with Score-Based Diffusion Probabilistic Models
Chenyu Shen, Ziyuan Yang 0001, Yi Zhang 0018
MICCAI (1)3
2023 Cross-database attack of different coding-based palmprint templates
Ziyuan Yang 0001, Lu Leng, Andrew Beng Jin Teoh, Bob Zhang 0001, Yi Zhang 0018
Knowl. Based Syst.5
2023 TIME-Net: Transformer-Integrated Multi-Encoder Network for limited-angle artifact removal in dual-energy CBCT
Yikun Zhang 0001, Dianlin Hu, Zhihong Yan, Qingxian Zhao, Guotao Quan, Shouhua Luo, Yi Zhang 0018, Yang Chen 0008
Medical Image Anal.7
2023 SemiMAR: Semi-Supervised Learning for CT Metal Artifact Reduction
abstract
Metal artifacts lead to CT imaging quality degradation. With the success of deep learning (DL) in medical imaging, a number of DL-based supervised methods have been developed for metal artifact reduction (MAR). Nonetheless, fully-supervised MAR methods based on simulated data do not perform well on clinical data due to the domain gap. Although this problem can be avoided in an unsupervised way to a certain degree, severe artifacts cannot be well suppressed in clinical practice. Recently, semi-supervised metal artifact reduction (MAR) methods have gained wide attention due to their ability in narrowing the domain gap and improving MAR performance in clinical data. However, these methods typically require large model sizes, posing challenges for optimization. To address this issue, we propose a novel semi-supervised MAR framework. In our framework, only the artifact-free parts are learned, and the artifacts are inferred by subtracting these clean parts from the metal-corrupted CT images. Our approach leverages a single generator to execute all complex transformations, thereby reducing the model's scale and preventing overlap between clean part and artifacts. To recover more tissue details, we distill the knowledge from the advanced dual-domain MAR network into our model in both image domain and latent feature space. The latent space constraint is achieved via contrastive learning. We also evaluate the impact of different generator architectures by investigating several mainstream deep learning-based MAR backbones. Our experiments demonstrate that the proposed method competes favorably with several state-of-the-art semi-supervised MAR techniques in both qualitative and quantitative aspects.
Tao Wang 0167, Zhiwen Wang 0002, Hu Chen 0002, Yan Liu 0052, Jingfeng Lu, Yi Zhang 0018
IEEE J. Biomed. Health Informatics7
2023 Dynamic Corrected Split Federated Learning With Homomorphic Encryption for U-Shaped Medical Image Networks
abstract
U-shaped networks have become prevalent in various medical image tasks such as segmentation, and restoration. However, most existing U-shaped networks rely on centralized learning which raises privacy concerns. To address these issues, federated learning (FL) and split learning (SL) have been proposed. However, achieving a balance between the local computational cost, model privacy, and parallel training remains a challenge. In this articler, we propose a novel hybrid learning paradigm called Dynamic Corrected Split Federated Learning (DC-SFL) for U-shaped medical image networks. To preserve data privacy, including the input, model parameters, label and output simultaneously, we propose to split the network into three parts hosted by different parties. We propose a Dynamic Weight Correction Strategy (DWCS) to stabilize the training process and avoid the model drift problem due to data heterogeneity. To further enhance privacy protection and establish a trustworthy distributed learning paradigm, we propose to introduce additively homomorphic encryption into the aggregation process of client-side model, which helps prevent potential collusion between parties and provides a better privacy guarantee for our proposed method. The proposed DC-SFL is evaluated on various medical image tasks, and the experimental results demonstrate its effectiveness. In comparison with state-of-the-art distributed learning methods, our method achieves competitive performance.
Ziyuan Yang 0001, Huijie Huangfu, Maosong Ran, Hui Wang 0135, Xiaoxiao Li 0001, Yi Zhang 0018
IEEE J. Biomed. Health Informatics7
2023 DREAM-Net: Deep Residual Error Iterative Minimization Network for Sparse-View CT Reconstruction
abstract
Sparse-view Computed Tomography (CT) has the ability to reduce radiation dose and shorten the scan time, while the severe streak artifacts will compromise anatomical information. How to reconstruct high-quality images from sparsely sampled projections is a challenging ill-posed problem. In this context, we propose the unrolled Deep Residual Error iterAtive Minimization Network (DREAM-Net) based on a novel iterative reconstruction framework to synergize the merits of deep learning and iterative reconstruction. DREAM-Net performs constraints using deep neural networks in the projection domain, residual space, and image domain simultaneously, which is different from the routine practice in deep iterative reconstruction frameworks. First, a projection inpainting module completes the missing views to fully explore the latent relationship between projection data and reconstructed images. Then, the residual awareness module attempts to estimate the accurate residual image after transforming the projection error into the image space. Finally, the image refinement module learns a non-standard regularizer to further fine-tune the intermediate image. There is no need to empirically adjust the weights of different terms in DREAM-Net because the hyper-parameters are embedded implicitly in network modules. Qualitative and quantitative results have demonstrated the promising performance of DREAM-Net in artifact removal and structural fidelity.
Yikun Zhang 0001, Dianlin Hu, Shilei Hao, Jin Liu 0019, Guotao Quan, Yi Zhang 0018, Yang Chen 0008
IEEE J. Biomed. Health Informatics6
2023 M3NAS: Multi-Scale and Multi-Level Memory-Efficient Neural Architecture Search for Low-Dose CT Denoising
abstract
Lowering the radiation dose in computed tomography (CT) can greatly reduce the potential risk to public health. However, the reconstructed images from dose-reduced CT or low-dose CT (LDCT) suffer from severe noise which compromises the subsequent diagnosis and analysis. Recently, convolutional neural networks have achieved promising results in removing noise from LDCT images. The network architectures that are used are either handcrafted or built on top of conventional networks such as ResNet and U-Net. Recent advances in neural network architecture search (NAS) have shown that the network architecture has a dramatic effect on the model performance. This indicates that current network architectures for LDCT may be suboptimal. Therefore, in this paper, we make the first attempt to apply NAS to LDCT and propose a multi-scale and multi-level memory-efficient NAS for LDCT denoising, termed M3NAS. On the one hand, the proposed M3NAS fuses features extracted by different scale cells to capture multi-scale image structural details. On the other hand, the proposed M3NAS can search a hybrid cell- and network-level structure for better performance. In addition, M3NAS can effectively reduce the number of model parameters and increase the speed of inference. Extensive experimental results on two different datasets demonstrate that the proposed M3NAS can achieve better performance and fewer parameters than several state-of-the-art methods. In addition, we also validate the effectiveness of the multi-scale and multi-level architecture for LDCT denoising, and present further analysis for different configurations of super-net.
Wenjun Xia, Yongqiang Huang 0003, Mingzheng Hou, Hu Chen 0002, Jiliu Zhou, Hongming Shan, Yi Zhang 0018
IEEE Trans. Medical Imaging8
2023 MLF-IOSC: Multi-Level Fusion Network With Independent Operation Search Cell for Low-Dose CT Denoising
abstract
Computed tomography (CT) is widely used in clinical medicine, and low-dose CT (LDCT) has become popular to reduce potential patient harm during CT acquisition. However, LDCT aggravates the problem of noise and artifacts in CT images, increasing diagnosis difficulty. Through deep learning, denoising CT images by artificial neural network has aroused great interest for medical imaging and has been hugely successful. We propose a framework to achieve excellent LDCT noise reduction using independent operation search cells, inspired by neural architecture search, and introduce the Laplacian to further improve image quality. Employing patch-based training, the proposed method can effectively eliminate CT image noise while retaining the original structures and details, hence significantly improving diagnosis efficiency and promoting LDCT clinical applications.
Jinbo Shen, Mengting Luo, Peixi Liao, Hu Chen 0002, Yi Zhang 0018
IEEE Trans. Medical Imaging6
2023 Hyper RPCA: Joint Maximum Correntropy Criterion and Laplacian Scale Mixture Modeling on-the-Fly for Moving Object Detection
abstract
Moving object detection is critical for automated video analysis in many vision-related tasks, such as surveillance tracking, video compression coding, etc. Robust Principal Component Analysis (RPCA), as one of the most popular moving object modelling methods, aims to separate the temporally-varying (i.e., moving) foreground objects from the static background in video, assuming the background frames to be low-rank while the foreground to be spatially sparse. Classic RPCA imposes sparsity of the foreground component using$\ell _1$-norm, and minimizes the modeling error via$\ell _2$-norm. We show that such assumptions can be too restrictive in practice, which limits the effectiveness of the classic RPCA, especially when processing videos with dynamic background, camera jitter, camouflaged moving object, etc. In this paper, we propose a novel RPCA-based model, called Hyper RPCA, to detect moving objects on the fly. Different from classic RPCA, the proposed Hyper RPCA jointly applies the maximum correntropy criterion (MCC) for the modeling error, and Laplacian scale mixture (LSM) model for foreground objects. Extensive experiments have been conducted, and the results demonstrate that the proposed Hyper RPCA has competitive performance for foreground detection to the state-of-the-art algorithms on several well-known benchmark datasets.
Zerui Shao, Yi-Fei Pu, Jiliu Zhou, Bihan Wen, Yi Zhang 0018
IEEE Trans. Multim.5
2023 DHI-GAN: Improving Dental-Based Human Identification Using Generative Adversarial Networks
abstract
In this work, a novel semisupervised framework is proposed to tackle the small-sample problem of dental-based human identification (DHI), achieving enhanced performance via a "classifying while generating" paradigm. A generative adversarial network (GAN), called the DHI-GAN, is presented to implement this idea, in which an extra classifier is also dedicatedly proposed to achieve an efficient training procedure. Considering the complex specificities of this problem, except for the noise input of the generator, an identity embedding-guided architecture is proposed to retain informative features for each individual. A parallel spatial and channel fusion attention block is innovatively designed to encourage the model to learn discriminative and informative features by focusing on different regional details and abstract concepts. The attention block is also widely applied to the overall classifier to learn identity-dependent information. A loss combination of the ArcFace and focal loss is utilized to address the small-sample problem. Two parameters are proposed to control the generated samples that are fed into the classifier during the optimization procedure. The proposed DHI-GAN framework is finally validated on a real-world dataset, and the experimental results demonstrate that it outperforms other baselines, achieving a 92.5% top-one accuracy rate. Most importantly, the proposed GAN-based semisupervised training strategy is able to reduce the required number of training samples (individuals) and can also be incorporated into other classification models. Our code will be available at https://github.com/sculyi/MedicalImages/.
Yi Lin 0006, Jianwei Zhang 0013, Jizhe Zhou 0001, Peixi Liao, Hu Chen 0002, Zhenhua Deng, Yi Zhang 0018
IEEE Trans. Neural Networks Learn. Syst.8
2022 Depth Completion Using Geometry-Aware Embedding
abstract
Exploiting internal spatial geometric constraints of sparse LiDARs is beneficial to depth completion, however, has been not explored well. This paper proposes an efficient method to learn geometry-aware embedding, which encodes the local and global geometric structure information from 3D points, e.g., scene layout, object's sizes and shapes, to guide dense depth estimation. Specifically, we utilize the dynamic graph representation to model generalized geometric relationship from irregular point clouds in a flexible and efficient manner. Further, we joint this embedding and corresponded RGB appearance information to infer missing depths of the scene with well structure-preserved details. The key to our method is to integrate implicit 3D geometric representation into a 2D learning architecture, which leads to a better trade-off between the performance and efficiency. Extensive experiments demonstrate that the proposed method outperforms previous works and could reconstruct fine depths with crisp boundaries in regions that are over-smoothed by them. The ablation study gives more insights into our method that could achieve significant gains with a simple design, while having better generalization capability and stability. The code is available at https://github.com/Wenchao-Du/GAENet.
Wenchao Du, Hu Chen 0002, Hongyu Yang 0002, Yi Zhang 0018
ICRA4
2022 A Transformer-Based Iterative Reconstruction Model for Sparse-View CT Reconstruction
Wenjun Xia, Ziyuan Yang 0001, Qizheng Zhou, Zhongxian Wang, Yi Zhang 0018
MICCAI (6)6
2022 NFANet: A Novel Method for Weakly Supervised Water Extraction From High-Resolution Remote-Sensing Imagery
abstract
The use of deep learning for water extraction requires precise pixel-level labels. However, it is very difficult to label high-resolution remote-sensing images at the pixel level. Therefore, we study how to utilize point labels to extract water bodies and propose a novel method called the neighbor feature aggregation network (NFANet). Compared with pixel-level labels, point labels are much easier to obtain, but they will lose much information. In this article, we take advantage of the similarity between the adjacent pixels of a local water body, and propose a neighbor sampler to resample remote-sensing images. Then, the sampled images are sent to the network for feature aggregation. In addition, we use an improved recursive training algorithm to further improve the extraction accuracy, making the water boundary more natural. Furthermore, our method utilizes neighboring features instead of global or local features to learn more representative features. The experimental results show that the proposed NFANet method not only outperforms other studied weakly supervised approaches, but also obtains similar results as the state-of-the-art ones.
Leyuan Fang, Muxing Li, Bob Zhang 0001, Yi Zhang 0018, Pedram Ghamisi
IEEE Trans. Geosci. Remote. Sens.5
2022 PRIOR: Prior-Regularized Iterative Optimization Reconstruction For 4D CBCT
abstract
4D cone-beam computed tomography (CBCT) is an important imaging modality in image-guided radiation therapy to address the motion-induced artifacts caused by organ movements during the respiratory process. However, due to the extremely sparse projection data for each temporal phase, 4D CBCT reconstructions will suffer from severe streaking artifacts. Therefore, to tackle the streak artifacts and provide high-quality images, we proposed a framework termed Prior-Regularized Iterative Optimization Reconstruction (PRIOR) for 4D CBCT. The PRIOR framework combines the physics-based model and data-driven method simultaneously, with powerful feature extracting capacity, significantly promoting the image quality compared to single model-based or deep learning-based methods. Besides, we designed a specialized deep learning model named PRIOR-Net, which can effectively excavate the static information in the prior image reconstructed from the fully-sampled projections at the encoding stage to improve the reconstruction performance for individual phase-resolved images. Both the simulated and clinical 4D CBCT datasets were performed to evaluate the performance of the PRIOR-Net and the PRIOR framework. Compared with the advanced 4D CBCT reconstruction methods, the proposed methods achieve promising results quantitatively and qualitatively in streak artifact suppression, soft tissue restoration, and tiny detail preservation.
Dianlin Hu, Yikun Zhang 0001, Jin Liu 0019, Yi Zhang 0018, Jean-Louis Coatrieux, Yang Chen 0008
IEEE J. Biomed. Health Informatics4
2022 FONT-SIR: Fourth-Order Nonlocal Tensor Decomposition Model for Spectral CT Image Reconstruction
abstract
Spectral computed tomography (CT) reconstructs images from different spectral data through photon counting detectors (PCDs). However, due to the limited number of photons and the counting rate in the corresponding spectral segment, the reconstructed spectral images are usually affected by severe noise. In this paper, we propose a fourth-order nonlocal tensor decomposition model for spectral CT image reconstruction (FONT-SIR). To maintain the original spatial relationships among similar patches and improve the imaging quality, similar patches without vectorization are grouped in both spectral and spatial domains simultaneously to form the fourth-order processing tensor unit. The similarity of different patches is measured with the cosine similarity of latent features extracted using principal component analysis (PCA). By imposing the constraints of the weighted nuclear and total variation (TV) norms, each fourth-order tensor unit is decomposed into a low-rank component and a sparse component, which can efficiently remove noise and artifacts while preserving the structural details. Moreover, the alternating direction method of multipliers (ADMM) is employed to solve the decomposition model. Extensive experimental results on both simulated and real data sets demonstrate that the proposed FONT-SIR achieves superior qualitative and quantitative performance compared with several state-of-the-art methods.
Xiang Chen 0015, Wenjun Xia, Yan Liu 0052, Hu Chen 0002, Jiliu Zhou, Zhiyuan Zha, Bihan Wen, Yi Zhang 0018
IEEE Trans. Medical Imaging8
2021 Dual-Domain Adaptive-Scaling Non-local Network for CT Metal Artifact Reduction
Tao Wang 0167, Wenjun Xia, Yongqiang Huang 0003, Huaiqiang Sun, Yan Liu 0052, Hu Chen 0002, Jiliu Zhou, Yi Zhang 0018
MICCAI (6)8
2021 Heterogeneous Face Recognition with Attention-guided Feature Disentangling
abstract
This paper proposes an attention-guided feature disentangling framework (AgFD) to eliminate the large cross-modality discrepancy for Heterogeneous Face Recognition (HFR). Existing HFR methods either focus only on extracting identity features or impose linear/no independence constraints on the decomposed components. Instead, our AgFD disentangles the facial representation and forces intrinsic independence between identity features and identity-irrelevant variations. To this end, an Attention-based Residual Decomposition Module (AbRDM) and an Adversarial Decorrelation Module (ADM) are presented. AbRDM provides hierarchical complementary feature disentanglement, while ADM is introduced for decorrelation learning. Extensive experiments on the challenging CASIA NIR-VIS 2.0 Database, Oulu-CASIA NIR&VIS Database, BUAA-VisNir Database, and IIIT-D Viewed Sketch Database demonstrate the generalization ability and competitive performance of the proposed method.
Shanmin Yang, Xiao Yang 0029, Yi Lin 0006, Peng Cheng 0006, Yi Zhang 0018, Jianwei Zhang 0013
ACM Multimedia5
2021 Identity-and-pose-guided generative adversarial network for face rotation
Yi Zhang 0018, Keren Fu, Peng Cheng 0006
Neurocomputing1
2021 PGM-face: Pose-guided margin loss for cross-pose face recognition
Yi Zhang 0018, Keren Fu, Peng Cheng 0006, Shanmin Yang, Xiao Yang 0029
Neurocomputing1
2021 Noise-Powered Disentangled Representation for Unsupervised Speckle Reduction of Optical Coherence Tomography Images
abstract
Due to its noninvasive character, optical coherence tomography (OCT) has become a popular diagnostic method in clinical settings. However, the low-coherence interferometric imaging procedure is inevitably contaminated by heavy speckle noise, which impairs both visual quality and diagnosis of various ocular diseases. Although deep learning has been applied for image denoising and achieved promising results, the lack of well-registered clean and noisy image pairs makes it impractical for supervised learning-based approaches to achieve satisfactory OCT image denoising results. In this paper, we propose an unsupervised OCT image speckle reduction algorithm that does not rely on well-registered image pairs. Specifically, by employing the ideas of disentangled representation and generative adversarial network, the proposed method first disentangles the noisy image into content and noise spaces by corresponding encoders. Then, the generator is used to predict the denoised OCT image with the extracted content features. In addition, the noise patches cropped from the noisy image are utilized to facilitate more accurate disentanglement. Extensive experiments have been conducted, and the results suggest that our proposed method is superior to the classic methods and demonstrates competitive performance to several recently proposed learning-based approaches in both quantitative and qualitative aspects. Code is available at: https://github.com/tsmotlp/DRGAN-OCT.
Yongqiang Huang 0003, Wenjun Xia, Yan Liu 0052, Hu Chen 0002, Jiliu Zhou, Leyuan Fang, Yi Zhang 0018
IEEE Trans. Medical Imaging8
2021 LCANet: Learnable Connected Attention Network for Human Identification Using Dental Images
abstract
Forensic odontology is regarded as an important branch of forensics dealing with human identification based on dental identification. This paper proposes a novel method that uses deep convolution neural networks to assist in human identification by automatically and accurately matching 2-D panoramic dental X-ray images. Designed as a top-down architecture, the network incorporates an improved channel attention module and a learnable connected module to better extract features for matching. By integrating associated features among all channel maps, the channel attention module can selectively emphasize interdependent channel information, which contributes to more precise recognition results. The learnable connected module not only connects different layers in a feed-forward fashion but also searches the optimal connections for each connected layer, resulting in automatically and adaptively learning the connections among layers. Extensive experiments demonstrate that our method can achieve new state-of-the-art performance in human identification using dental images. Specifically, the method is tested on a dataset including 1,168 dental panoramic images of 503 different subjects, and its dental image recognition accuracy for human identification reaches 87.21% rank-1 accuracy and 95.34% rank-5 accuracy. Code has been released on Github. (https://github.com/cclaiyc/TIdentify).
Yancun Lai, Qingsong Wu, Wenchi Ke, Peixi Liao, Zhenhua Deng, Hu Chen 0002, Yi Zhang 0018
IEEE Trans. Medical Imaging8
2021 CT Reconstruction With PDF: Parameter-Dependent Framework for Data From Multiple Geometries and Dose Levels
abstract
The current mainstream computed tomography (CT) reconstruction methods based on deep learning usually need to fix the scanning geometry and dose level, which significantly aggravates the training costs and requires more training data for real clinical applications. In this paper, we propose a parameter-dependent framework (PDF) that trains a reconstruction network with data originating from multiple alternative geometries and dose levels simultaneously. In the proposed PDF, the geometry and dose level are parameterized and fed into two multilayer perceptrons (MLPs). The outputs of the MLPs are used to modulate the feature maps of the CT reconstruction network, which condition the network outputs on different geometries and dose levels. The experiments show that our proposed method can obtain competitive performance compared to the original network trained with either specific or mixed geometry and dose level, which can efficiently save extra training costs for multiple geometries and dose levels.
Wenjun Xia, Yongqiang Huang 0003, Yan Liu 0052, Hu Chen 0002, Jiliu Zhou, Yi Zhang 0018
IEEE Trans. Medical Imaging7
2021 MAGIC: Manifold and Graph Integrative Convolutional Network for Low-Dose CT Reconstruction
abstract
Low-dose computed tomography (LDCT) scans, which can effectively alleviate the radiation problem, will degrade the imaging quality. In this paper, we propose a novel LDCT reconstruction network that unrolls the iterative scheme and performs in both image and manifold spaces. Because patch manifolds of medical images have low-dimensional structures, we can build graphs from the manifolds. Then, we simultaneously leverage the spatial convolution to extract the local pixel-level features from the images and incorporate the graph convolution to analyze the nonlocal topological features in manifold space. The experiments show that our proposed method outperforms both the quantitative and qualitative aspects of state-of-the-art methods. In addition, aided by a projection loss component, our proposed method also demonstrates superior performance for semi-supervised learning. The network can remove most noise while maintaining the details of only 10% (40 slices) of the training data labeled.
Wenjun Xia, Yongqiang Huang 0003, Zuoqiang Shi, Yan Liu 0052, Hu Chen 0002, Yang Chen 0008, Jiliu Zhou, Yi Zhang 0018
IEEE Trans. Medical Imaging9
2021 CLEAR: Comprehensive Learning Enabled Adversarial Reconstruction for Subtle Structure Enhanced Low-Dose CT Imaging
abstract
X-ray computed tomography (CT) is of great clinical significance in medical practice because it can provide anatomical information about the human body without invasion, while its radiation risk has continued to attract public concerns. Reducing the radiation dose may induce noise and artifacts to the reconstructed images, which will interfere with the judgments of radiologists. Previous studies have confirmed that deep learning (DL) is promising for improving low-dose CT imaging. However, almost all the DL-based methods suffer from subtle structure degeneration and blurring effect after aggressive denoising, which has become the general challenging issue. This paper develops the Comprehensive Learning Enabled Adversarial Reconstruction (CLEAR) method to tackle the above problems. CLEAR achieves subtle structure enhanced low-dose CT imaging through a progressive improvement strategy. First, the generator established on the comprehensive domain can extract more features than the one built on degraded CT images and directly map raw projections to high-quality CT images, which is significantly different from the routine GAN practice. Second, a multi-level loss is assigned to the generator to push all the network components to be updated towards high-quality reconstruction, preserving the consistency between generated images and gold-standard images. Finally, following the WGAN-GP modality, CLEAR can migrate the real statistical properties to the generated images to alleviate over-smoothing. Qualitative and quantitative analyses have demonstrated the competitive performance of CLEAR in terms of noise suppression, structural fidelity and visual perception improvement.
Yikun Zhang 0001, Dianlin Hu, Qianlong Zhao, Guotao Quan, Jin Liu 0019, Qiegen Liu, Yi Zhang 0018, Gouenou Coatrieux, Yang Chen 0008, Hengyong Yu
IEEE Trans. Medical Imaging7
2020 Disentanglement Network for Unsupervised Speckle Reduction of Optical Coherence Tomography Images
Yongqiang Huang 0003, Wenjun Xia, Yan Liu 0052, Jiliu Zhou, Leyuan Fang, Yi Zhang 0018
MICCAI (5)7
2020 Learning from discrete Gaussian label distribution and spatial channel-aware residual attention for head pose estimation
Yi Zhang 0018, Keren Fu, Jiang Wang 0005, Peng Cheng 0006
Neurocomputing1
2020 Residual Encoder-Decoder Conditional Generative Adversarial Network for Pansharpening
abstract
Due to the limitation of the satellite sensor, it is difficult to acquire a high-resolution (HR) multispectral (HRMS) image directly. The aim of pansharpening (PNN) is to fuse the spatial in panchromatic (PAN) with the spectral information in multispectral (MS). Recently, deep learning has drawn much attention, and in the field of remote sensing, several pioneering attempts have been made related to PNN. However, the big size of remote sensing data will produce more training samples, which require a deeper neural network. Most current networks are relatively shallow and raise the possibility of detail loss. In this letter, we propose a residual encoder-decoder conditional generative adversarial network (RED-cGAN) for PNN to produce more details with sharpened images. The proposed method combines the idea of an autoencoder with generative adversarial network (GAN), which can effectively preserve the spatial and spectral information of the PAN and MS images simultaneously. First, the residual encoder-decoder module is adopted to extract the multiscale features from the last step to yield pansharpened images and relieve the training difficulty caused by deepening the network layers. Second, to further enhance the performance of the generator to preserve more spatial information, a conditional discriminator network with the input of PAN and MS images is proposed to encourage that the estimated MS images share the same distribution as that of the referenced HRMS images. The experiments conducted on the Worldview2 (WV2) and Worldview3 (WV3) images demonstrate that our proposed method provides better results than several state-of-the-art PNN methods.
Zhimin Shao, Maosong Ran, Leyuan Fang, Jiliu Zhou, Yi Zhang 0018
IEEE Geosci. Remote. Sens. Lett.6
2020 CT Super-Resolution GAN Constrained by the Identical, Residual, and Cycle Learning Ensemble (GAN-CIRCLE)
abstract
In this paper, we present a semi-supervised deep learning approach to accurately recover high-resolution (HR) CT images from low-resolution (LR) counterparts. Specifically, with the generative adversarial network (GAN) as the building block, we enforce the cycle-consistency in terms of the Wasserstein distance to establish a nonlinear end-to-end mapping from noisy LR input images to denoised and deblurred HR outputs. We also include the joint constraints in the loss function to facilitate structural preservation. In this process, we incorporate deep convolutional neural network (CNN), residual learning, and network in network techniques for feature extraction and restoration. In contrast to the current trend of increasing network depth and complexity to boost the imaging performance, we apply a parallel 1×1 CNN to compress the output of the hidden layer and optimize the number of layers and the number of filters for each convolutional layer. The quantitative and qualitative evaluative results demonstrate that our proposed model is accurate, efficient and robust for super-resolution (SR) image restoration from noisy LR input images. In particular, we validate our composite SR networks on three large-scale CT datasets, and obtain promising results as compared to the other state-of-the-art methods.
Chenyu You, Wenxiang Cong, Michael W. Vannier, Punam K. Saha, Eric A. Hoffman, Ge Wang 0001, Guang Li 0011, Yi Zhang 0018, Xiaoliu Zhang, Hongming Shan, Mengzhou Li, Shenghong Ju, Zhen Zhao 0003, Zhuiyang Zhang
IEEE Trans. Medical Imaging8
2019 Denoising of 3D magnetic resonance images using a residual encoder-decoder Wasserstein generative adversarial network
Maosong Ran, Jinrong Hu, Yang Chen 0008, Hu Chen 0002, Huaiqiang Sun, Jiliu Zhou, Yi Zhang 0018
Medical Image Anal.7
2019 Visual Attention Network for Low-Dose CT
abstract
Noise and artifacts are intrinsic to low-dose computed tomography (LDCT) data acquisition, and will significantly affect the imaging performance. Perfect noise removal and image restoration is intractable in the context of LDCT due to the statistical and the technical uncertainties. In this letter, we apply the generative adversarial network (GAN) framework with a visual attention mechanism to deal with this problem in a data-driven/machine learning fashion. Our main idea is to inject visual attention knowledge into the learning process of GAN to provide a powerful prior of the noise distribution. By doing this, both the generator and discriminator networks are empowered with visual attention information so that they will not only pay special attention to noisy regions and surrounding structures but also explicitly assess the local consistency of the recovered regions. Our experiments qualitatively and quantitatively demonstrate the effectiveness of the proposed method with clinic CT images.
Wenchao Du, Hu Chen 0002, Peixi Liao, Hongyu Yang 0002, Ge Wang 0001, Yi Zhang 0018
IEEE Signal Process. Lett.6
2019 Convolutional Sparse Coding for Compressed Sensing CT Reconstruction
abstract
Over the past few years, dictionary learning (DL)-based methods have been successfully used in various image reconstruction problems. However, the traditional DL-based computed tomography (CT) reconstruction methods are patch-based and ignore the consistency of pixels in overlapped patches. In addition, the features learned by these methods always contain shifted versions of the same features. In recent years, convolutional sparse coding (CSC) has been developed to address these problems. In this paper, inspired by several successful applications of CSC in the field of signal processing, we explore the potential of CSC in sparse-view CT reconstruction. By directly working on the whole image, without the necessity of dividing the image into overlapped patches in DL-based methods, the proposed methods can maintain more details and avoid artifacts caused by patch aggregation. With predetermined filters, an alternating scheme is developed to optimize the objective function. Extensive experiments with simulated and real CT data were performed to validate the effectiveness of the proposed methods. The qualitative and quantitative results demonstrate that the proposed methods achieve better performance than the several existing state-of-the-art methods.
Peng Bao 0001, Huaiqiang Sun, Zhangyang Wang, Yi Zhang 0018, Wenjun Xia, Mianyi Chen, Yan Xi, Shanzhou Niu, Jiliu Zhou, He Zhang 0004
IEEE Trans. Medical Imaging4
2018 Low-dose CT restoration via stacked sparse denoising autoencoders
Yan Liu 0052, Yi Zhang 0018
Neurocomputing2
2018 LEARN: Learned Experts' Assessment-Based Reconstruction Network for Sparse-Data CT
abstract
Compressive sensing (CS) has proved effective for tomographic reconstruction from sparsely collected data or under-sampled measurements, which are practically important for few-view computed tomography (CT), tomosynthesis, interior tomography, and so on. To perform sparse-data CT, the iterative reconstruction commonly uses regularizers in the CS framework. Currently, how to choose the parameters adaptively for regularization is a major open problem. In this paper, inspired by the idea of machine learning especially deep learning, we unfold the state-of-the-art "fields of experts"-based iterative reconstruction scheme up to a number of iterations for data-driven training, construct a learned experts' assessment-based reconstruction network (LEARN) for sparse-data CT, and demonstrate the feasibility and merits of our LEARN network. The experimental results with our proposed LEARN network produces a superior performance with the well-known Mayo Clinic low-dose challenge data set relative to the several state-of-the-art methods, in terms of artifact reduction, feature preservation, and computational speed. This is consistent to our insight that because all the regularization terms and parameters used in the iterative reconstruction are now learned from the training data, our LEARN network utilizes application-oriented knowledge more effectively and recovers underlying images more favorably than competing algorithms. Also, the number of layers in the LEARN network is only 50, reducing the computational complexity of typical iterative algorithms by orders of magnitude.
Hu Chen 0002, Yi Zhang 0018, Yunjin Chen, Huaiqiang Sun, Yang Lu 0011, Peixi Liao, Jiliu Zhou, Ge Wang 0001
IEEE Trans. Medical Imaging2
2018 3-D Convolutional Encoder-Decoder Network for Low-Dose CT via Transfer Learning From a 2-D Trained Network
abstract
Low-dose computed tomography (LDCT) has attracted major attention in the medical imaging field, since CT-associated X-ray radiation carries health risks for patients. The reduction of the CT radiation dose, however, compromises the signal-to-noise ratio, which affects image quality and diagnostic performance. Recently, deep-learning-based algorithms have achieved promising results in LDCT denoising, especially convolutional neural network (CNN) and generative adversarial network (GAN) architectures. This paper introduces a conveying path-based convolutional encoder-decoder (CPCE) network in 2-D and 3-D configurations within the GAN framework for LDCT denoising. A novel feature of this approach is that an initial 3-D CPCE denoising model can be directly obtained by extending a trained 2-D CNN, which is then fine-tuned to incorporate 3-D spatial information from adjacent slices. Based on the transfer learning from 2-D to 3-D, the 3-D network converges faster and achieves a better denoising performance when compared with a training from scratch. By comparing the CPCE network with recently published work based on the simulated Mayo data set and the real MGH data set, we demonstrate that the 3-D CPCE denoising model has a better performance in that it suppresses image noise and preserves subtle structures.
Hongming Shan, Yi Zhang 0018, Qingsong Yang, Uwe Krüger 0001, Mannudeep K. Kalra, Ling Sun 0006, Wenxiang Cong, Ge Wang 0001
IEEE Trans. Medical Imaging2
2018 Correction for "3D Convolutional Encoder-Decoder Network for Low-Dose CT via Transfer Learning From a 2D Trained Network"
abstract
In[1], please note the updated figure captions for Figures 5, 6, 7, and 8 as follows:
Hongming Shan, Yi Zhang 0018, Qingsong Yang, Uwe Krüger 0001, Mannudeep K. Kalra, Ling Sun 0006, Wenxiang Cong, Ge Wang 0001
IEEE Trans. Medical Imaging2
2018 Low-Dose CT Image Denoising Using a Generative Adversarial Network With Wasserstein Distance and Perceptual Loss
abstract
The continuous development and extensive use of computed tomography (CT) in medical practice has raised a public concern over the associated radiation dose to the patient. Reducing the radiation dose may lead to increased noise and artifacts, which can adversely affect the radiologists' judgment and confidence. Hence, advanced image reconstruction from low-dose CT data is needed to improve the diagnostic performance, which is a challenging problem due to its ill-posed nature. Over the past years, various low-dose CT methods have produced impressive results. However, most of the algorithms developed for this application, including the recently popularized deep learning techniques, aim for minimizing the mean-squared error (MSE) between a denoised CT image and the ground truth under generic penalties. Although the peak signal-to-noise ratio is improved, MSE- or weighted-MSE-based methods can compromise the visibility of important structural details after aggressive denoising. This paper introduces a new CT image denoising method based on the generative adversarial network (GAN) with Wasserstein distance and perceptual similarity. The Wasserstein distance is a key concept of the optimal transport theory and promises to improve the performance of GAN. The perceptual loss suppresses noise by comparing the perceptual features of a denoised output against those of the ground truth in an established feature space, while the GAN focuses more on migrating the data noise distribution from strong to weak statistically. Therefore, our proposed method transfers our knowledge of visual perception to the image denoising task and is capable of not only reducing the image noise level but also trying to keep the critical information at the same time. Promising results have been obtained in our experiments with clinical CT images.
Qingsong Yang, Pingkun Yan, Hengyong Yu, Yongyi Shi, Xuanqin Mou, Mannudeep K. Kalra, Yi Zhang 0018, Ling Sun 0006, Ge Wang 0001
IEEE Trans. Medical Imaging8
2017 Defense Against Chip Cloning Attacks Based on Fractional Hopfield Neural Networks
abstract
This paper presents a state-of-the-art application of fractional hopfield neural networks (FHNNs) to defend against chip cloning attacks, and provides insight into the reason that the proposed method is superior to physically unclonable functions (PUFs). In the past decade, PUFs have been evolving as one of the best types of hardware security. However, the development of the PUFs has been somewhat limited by its implementation cost, its temperature variation effect, its electromagnetic interference effect, the amount of entropy in it, etc. Therefore, it is imperative to discover, through promising mathematical methods and physical modules, some novel mechanisms to overcome the aforementioned weaknesses of the PUFs. Motivated by this need, in this paper, we propose applying the FHNNs to defend against chip cloning attacks. At first, we implement the arbitrary-order fractor of a FHNN. Secondly, we describe the implementation cost of the FHNNs. Thirdly, we propose the achievement of the constant-order performance of a FHNN when ambient temperature varies. Fourthly, we analyze the electrical performance stability of the FHNNs under electromagnetic disturbance conditions. Fifthly, we study the amount of entropy of the FHNNs. Lastly, we perform experiments to analyze the pass-band width of the fractor of an arbitrary-order FHNN and the defense against chip cloning attacks capability of the FHNNs. In particular, the capabilities of defense against chip cloning attacks, anti-electromagnetic interference, and anti-temperature variation of a FHNN are illustrated experimentally in detail. Some significant advantages of the FHNNs are that their implementation cost is considerably lower than that of the PUFs, their electrical performance is much more stable than that of the PUFs under different temperature conditions, their electrical performance stability of the FHNNs under electromagnetic disturbance conditions is much more robust than that of the PUFs, and their amount of entropy is significantly higher than that of the PUFs with the same rank circuit scale.
Yi-Fei Pu, Yi Zhang 0018, Jiliu Zhou
Int. J. Neural Syst.2
2017 Low-Dose CT With a Residual Encoder-Decoder Convolutional Neural Network
abstract
Given the potential risk of X-ray radiation to the patient, low-dose CT has attracted a considerable interest in the medical imaging field. Currently, the main stream low-dose CT methods include vendor-specific sinogram domain filtration and iterative reconstruction algorithms, but they need to access raw data, whose formats are not transparent to most users. Due to the difficulty of modeling the statistical characteristics in the image domain, the existing methods for directly processing reconstructed images cannot eliminate image noise very well while keeping structural details. Inspired by the idea of deep learning, here we combine the autoencoder, deconvolution network, and shortcut connections into the residual encoder-decoder convolutional neural network (RED-CNN) for low-dose CT imaging. After patch-based training, the proposed RED-CNN achieves a competitive performance relative to the-state-of-art methods in both simulated and clinical cases. Especially, our method has been favorably evaluated in terms of noise suppression, structural preservation, and lesion detection.
Hu Chen 0002, Yi Zhang 0018, Mannudeep K. Kalra, Feng Lin 0010, Yang Chen 0008, Peixi Liao, Jiliu Zhou, Ge Wang 0001
IEEE Trans. Medical Imaging2
2017 Discriminative Feature Representation to Improve Projection Data Inconsistency for Low Dose CT Imaging
abstract
In low dose computed tomography (LDCT) imaging, the data inconsistency of measured noisy projections can significantly deteriorate reconstruction images. To deal with this problem, we propose here a new sinogram restoration approach, the sinogram- discriminative feature representation (S-DFR) method. Different from other sinogram restoration methods, the proposed method works through a 3-D representation-based feature decomposition of the projected attenuation component and the noise component using a well-designed composite dictionary containing atoms with discriminative features. This method can be easily implemented with good robustness in parameter setting. Its comparison to other competing methods through experiments on simulated and real data demonstrated that the S-DFR method offers a sound alternative in LDCT.
Jin Liu 0019, Jianhua Ma 0001, Yi Zhang 0018, Yang Chen 0008, Jian Yang 0009, Huazhong Shu, Limin Luo 0001, Gouenou Coatrieux, Wei Yang 0006, Qianjin Feng 0004, Wufan Chen
IEEE Trans. Medical Imaging3
2016 Analysis of micro-Doppler signatures of vibration targets using EMD and SPWVD
Yan Wang 0015, Xi Wu 0004, Wenzao Li, Yi Zhang 0018, Jiliu Zhou
Neurocomputing5
2016 A texture image denoising approach based on fractional developmental mathematics
Yi-Fei Pu, Yi Zhang 0018, Jiliu Zhou
Pattern Anal. Appl.3
2015 Fractional Extreme Value Adaptive Training Method: Fractional Steepest Descent Approach
abstract
The application of fractional calculus to signal processing and adaptive learning is an emerging area of research. A novel fractional adaptive learning approach that utilizes fractional calculus is presented in this paper. In particular, a fractional steepest descent approach is proposed. A fractional quadratic energy norm is studied, and the stability and convergence of our proposed method are analyzed in detail. The fractional steepest descent approach is implemented numerically and its stability is analyzed experimentally.
Yi-Fei Pu, Jiliu Zhou, Yi Zhang 0018, Guo Huang, Patrick Siarry
IEEE Trans. Neural Networks Learn. Syst.3
2014 A homography transform based higher-order MRF model for stereo matching
Menglong Yang, Yiguang Liu, Zhisheng You, Yi Zhang 0018
Pattern Recognit. Lett.5