Wenjun Xia

dblp:24/11285 · DBLP profile ↗
← Back
22ranked-venue papers
6as first author
18since 2021 · last 2026
0000-0002-0428-1490ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 15 · 4 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author
YearPublicationVenuePosition
2026 Information-Maximized Soft Variable Discretization for Self-Supervised Image Representation Learning
abstract
Self-supervised learning (SSL) has emerged as a crucial technique in image processing, encoding, and understanding, especially for developing today's vision foundation models that utilize large-scale datasets without annotations to enhance various downstream tasks. This study introduces a novel SSL approach, Information-Maximized Soft Variable Discretization (IMSVD), for image representation learning. Specifically, IMSVD softly discretizes each variable in the latent space, enabling the estimation of their probability distributions over training batches and allowing the learning process to be directly guided by information measures. Motivated by the MultiView assumption, we propose an information-theoretic objective function to learn transform-invariant, non-trivial, and redundancy-minimized representation features. We then derive a cross-joint entropy loss function for self-supervised image representation learning, which theoretically enjoys superiority over the existing methods in reducing feature redundancy. Notably, our non-contrastive IMSVD method statistically performs contrastive learning. Extensive experimental results demonstrate the effectiveness of IMSVD on various downstream tasks in terms of both accuracy and efficiency. Thanks to our variable discretization, the embedding features optimized by IMSVD offer unique explainability at the variable level. IMSVD has the potential to be adapted to other learning paradigms. Our code is publicly available at https://github.com/niuchuangnn/IMSVD.
Chuang Niu, Wenjun Xia, Hongming Shan, Ge Wang 0001
IEEE Trans. Image Process.2
2026 Privacy-Preserving Latent Diffusion-Based Synthetic Medical Image Generation
abstract
Deep learning methods have impacted almost every research field, demonstrating notable successes in medical imaging tasks such as denoising and super-resolution. However, the prerequisite for deep learning is data at scale, but data sharing is expensive yet at risk of privacy leakage. As cutting-edge AI generative models, diffusion models have now become dominant because of their rigorous foundation and unprecedented outcomes. Here we propose a latent diffusion approach for data synthesis without compromising patient privacy. In our exemplary case studies, we develop a latent diffusion model to generate medical CT, MRI, and PET images using publicly available datasets. We demonstrate that state-of-the-art deep learning-based denoising/super-resolution networks can be trained on our synthetic data to achieve image quality with no significant difference from what the same network can achieve after being trained on the original data. In our advanced diffusion model, we specifically embed a safeguard mechanism to protect patient privacy effectively and efficiently. Our approach enables privacy-preserving public sharing of diverse big datasets for development of deep models, potentially enabling federated learning at the level of input data instead of local network weights.
Yongyi Shi, Wenjun Xia, Chuang Niu, Christopher Wiedeman, Ge Wang 0001
IEEE Trans. Medical Imaging2
2026 Tomographic Foundation Model - FORCE: Flow-Oriented Reconstruction Conditioning Engine
abstract
Computed tomography (CT) is a major medical imaging modality. Clinical CT scenarios, such as low-dose screening, sparse-view scanning, and metal implants, often lead to severe noise and artifacts in reconstructed images, requiring improved reconstruction techniques. The introduction of deep learning has significantly advanced CT image reconstruction. However, obtaining paired training data remains rather challenging due to patient motion and other constraints. Although deep learning methods can still perform well with approximately paired data, they inherently carry the risk of hallucination due to data inconsistencies and model instability. In this paper, we integrate the data fidelity with the state-of-the-art generative AI model, referred to as the Poisson flow generative model (PFGM) with a generalized version PFGM++, and propose a novel CT framework: Flow-Oriented Reconstruction Conditioning Engine (FORCE). In our experiments, the proposed method shows superior performance in various CT imaging tasks, outperforming existing unsupervised reconstruction approaches.
Wenjun Xia, Chuang Niu, Ge Wang 0001
IEEE Trans. Medical Imaging1
2025 Hybrid Mamba-Transformer with Frequency Enhancement for Single Image Deraining
Yue Que 0001, Wenjun Xia, Xue Xia 0005
PRCV (9)2
2025 Low-dose computed tomography perceptual image quality assessment
abstract
In computed tomography (CT) imaging, optimizing the balance between radiation dose and image quality is crucial due to the potentially harmful effects of radiation on patients. Although subjective assessments by radiologists are considered the gold standard in medical imaging, these evaluations can be time-consuming and costly. Thus, objective methods, such as the peak signal-to-noise ratio and structural similarity index measure, are often employed as alternatives. However, these metrics, initially developed for natural images, may not fully encapsulate the radiologists' assessment process. Consequently, interest in developing deep learning-based image quality assessment (IQA) methods that more closely align with radiologists' perceptions is growing. A significant barrier to this development has been the absence of open-source datasets and benchmark models specific to CT IQA. Addressing these challenges, we organized the Low-dose Computed Tomography Perceptual Image Quality Assessment Challenge in conjunction with the Medical Image Computing and Computer Assisted Intervention 2023. This event introduced the first open-source CT IQA dataset, consisting of 1,000 CT images of various quality, annotated with radiologists' assessment scores. As a benchmark, this challenge offers a comprehensive analysis of six submitted methods, providing valuable insight into their performance. This paper presents a summary of these methods and insights. This challenge underscores the potential for developing no-reference IQA methods that could exceed the capabilities of full-reference IQA methods, making a significant contribution to the research community with this novel dataset. The dataset is accessible at https://zenodo.org/records/7833096.
Wonkyeong Lee, Fabian Wagner, Adrian Galdran, Yongyi Shi, Wenjun Xia, Ge Wang 0001, Xuanqin Mou, Md. Atik Ahamed, Abdullah-Al-Zubaer Imran, Jieun Oh, Kyung Sang Kim, Jong Tak Baek, Dongheon Lee 0002, Boohwi Hong, Philip Tempelman, Donghang Lyu, Adrian Kuiper, Lars van Blokland, Maria Baldeon Calisto, Scott S. Hsieh, Minah Han, Jongduk Baek, Andreas K. Maier, Adam S. Wang, Garry Gold, Jang Hwan Choi 0001
Medical Image Anal.5
2025 Hypernetwork-Based Physics-Driven Personalized Federated Learning for CT Imaging
abstract
In clinical practice, computed tomography (CT) is an important noninvasive inspection technology to provide patients' anatomical information. However, its potential radiation risk is an unavoidable problem that raises people's concerns. Recently, deep learning (DL)-based methods have achieved promising results in CT reconstruction, but these methods usually require the centralized collection of large amounts of data for training from specific scanning protocols, which leads to serious domain shift and privacy concerns. To relieve these problems, in this article, we propose a hypernetwork-based physics-driven personalized federated learning method (HyperFed) for CT imaging. The basic assumption of the proposed HyperFed is that the optimization problem for each domain can be divided into two subproblems: local data adaption and global CT imaging problems, which are implemented by an institution-specific physics-driven hypernetwork and a global-sharing imaging network, respectively. Learning stable and effective invariant features from different data distributions is the main purpose of global-sharing imaging network. Inspired by the physical process of CT imaging, we carefully design physics-driven hypernetwork for each domain to obtain hyperparameters from specific physical scanning protocol to condition the global-sharing imaging network, so that we can achieve personalized local CT reconstruction. Experiments show that HyperFed achieves competitive performance in comparison with several other state-of-the-art methods. It is believed as a promising direction to improve CT imaging quality and personalize the needs of different institutions or scanners without data sharing. Related codes have been released at https://github.com/Zi-YuanYang/HyperFed.
Ziyuan Yang 0001, Wenjun Xia, Xiaoxiao Li 0001, Yi Zhang 0018
IEEE Trans. Neural Networks Learn. Syst.2
2024 Progressive dual-domain-transfer cycleGAN for unsupervised MRI reconstruction
Zhiwen Wang 0002, Ziyuan Yang 0001, Wenjun Xia, Yi Zhang 0018
Neurocomputing4
2024 A Denoising Diffusion Probabilistic Model for Metal Artifact Reduction in CT
abstract
The presence of metal objects leads to corrupted CT projection measurements, resulting in metal artifacts in the reconstructed CT images. AI promises to offer improved solutions to estimate missing sinogram data for metal artifact reduction (MAR), as previously shown with convolutional neural networks (CNNs) and generative adversarial networks (GANs). Recently, denoising diffusion probabilistic models (DDPM) have shown great promise in image generation tasks, potentially outperforming GANs. In this study, a DDPM-based approach is proposed for inpainting of missing sinogram data for improved MAR. The proposed model is unconditionally trained, free from information on metal objects, which can potentially enhance its generalization capabilities across different types of metal implants compared to conditionally trained approaches. The performance of the proposed technique was evaluated and compared to the state-of-the-art normalized MAR (NMAR) approach as well as to CNN-based and GAN-based MAR approaches. The DDPM-based approach provided significantly higher SSIM and PSNR, as compared to NMAR (SSIM: p [Formula: see text]; PSNR: p [Formula: see text]), the CNN (SSIM: p [Formula: see text]; PSNR: p [Formula: see text]) and the GAN (SSIM: p [Formula: see text]; PSNR: p <0.05) methods. The DDPM-MAR technique was further evaluated based on clinically relevant image quality metrics on clinical CT images with virtually introduced metal objects and metal artifacts, demonstrating superior quality relative to the other three models. In general, the AI-based techniques showed improved MAR performance compared to the non-AI-based NMAR approach. The proposed methodology shows promise in enhancing the effectiveness of MAR, and therefore improving the diagnostic accuracy of CT.
Grigorios M. Karageorgos, Jiayong Zhang, Nils Peters, Wenjun Xia, Chuang Niu, Harald Paganetti, Ge Wang 0001, Bruno De Man
IEEE Trans. Medical Imaging4
2024 Blind CT Image Quality Assessment Using DDPM-Derived Content and Transformer-Based Evaluator
abstract
Lowering radiation dose per view and utilizing sparse views per scan are two common CT scan modes, albeit often leading to distorted images characterized by noise and streak artifacts. Blind image quality assessment (BIQA) strives to evaluate perceptual quality in alignment with what radiologists perceive, which plays an important role in advancing low-dose CT reconstruction techniques. An intriguing direction involves developing BIQA methods that mimic the operational characteristic of the human visual system (HVS). The internal generative mechanism (IGM) theory reveals that the HVS actively deduces primary content to enhance comprehension. In this study, we introduce an innovative BIQA metric that emulates the active inference process of IGM. Initially, an active inference module, implemented as a denoising diffusion probabilistic model (DDPM), is constructed to anticipate the primary content. Then, the dissimilarity map is derived by assessing the interrelation between the distorted image and its primary content. Subsequently, the distorted image and dissimilarity map are combined into a multi-channel image, which is inputted into a transformer-based image quality evaluator. By leveraging the DDPM-derived primary content, our approach achieves competitive performance on a low-dose CT dataset.
Yongyi Shi, Wenjun Xia, Ge Wang 0001, Xuanqin Mou
IEEE Trans. Medical Imaging2
2024 SOUL-Net: A Sparse and Low-Rank Unrolling Network for Spectral CT Image Reconstruction
abstract
Spectral computed tomography (CT) is an emerging technology, that generates a multienergy attenuation map for the interior of an object and extends the traditional image volume into a 4-D form. Compared with traditional CT based on energy-integrating detectors, spectral CT can make full use of spectral information, resulting in high resolution and providing accurate material quantification. Numerous model-based iterative reconstruction methods have been proposed for spectral CT reconstruction. However, these methods usually suffer from difficulties such as laborious parameter selection and expensive computational costs. In addition, due to the image similarity of different energy bins, spectral CT usually implies a strong low-rank prior, which has been widely adopted in current iterative reconstruction models. Singular value thresholding (SVT) is an effective algorithm to solve the low-rank constrained model. However, the SVT method requires a manual selection of thresholds, which may lead to suboptimal results. To relieve these problems, in this article, we propose a sparse and low-rank unrolling network (SOUL-Net) for spectral CT image reconstruction, that learns the parameters and thresholds in a data-driven manner. Furthermore, a Taylor expansion-based neural network backpropagation method is introduced to improve the numerical stability. The qualitative and quantitative results demonstrate that the proposed method outperforms several representative state-of-the-art algorithms in terms of detail preservation and artifact reduction.
Xiang Chen 0015, Wenjun Xia, Ziyuan Yang 0001, Hu Chen 0002, Yan Liu 0052, Jiliu Zhou, Yang Chen 0008, Bihan Wen, Yi Zhang 0018
IEEE Trans. Neural Networks Learn. Syst.2
2023 Root canal treatment planning by automatic tooth and root canal segmentation in dental CBCT with deep multi-task feature learning
Wenjun Xia, Zhennan Yan, Liang Zhao 0018, Xiaohe Bian, Zhengnan Qi, Shaoting Zhang 0001, Zisheng Tang
Medical Image Anal.2
2023 M3NAS: Multi-Scale and Multi-Level Memory-Efficient Neural Architecture Search for Low-Dose CT Denoising
abstract
Lowering the radiation dose in computed tomography (CT) can greatly reduce the potential risk to public health. However, the reconstructed images from dose-reduced CT or low-dose CT (LDCT) suffer from severe noise which compromises the subsequent diagnosis and analysis. Recently, convolutional neural networks have achieved promising results in removing noise from LDCT images. The network architectures that are used are either handcrafted or built on top of conventional networks such as ResNet and U-Net. Recent advances in neural network architecture search (NAS) have shown that the network architecture has a dramatic effect on the model performance. This indicates that current network architectures for LDCT may be suboptimal. Therefore, in this paper, we make the first attempt to apply NAS to LDCT and propose a multi-scale and multi-level memory-efficient NAS for LDCT denoising, termed M3NAS. On the one hand, the proposed M3NAS fuses features extracted by different scale cells to capture multi-scale image structural details. On the other hand, the proposed M3NAS can search a hybrid cell- and network-level structure for better performance. In addition, M3NAS can effectively reduce the number of model parameters and increase the speed of inference. Extensive experimental results on two different datasets demonstrate that the proposed M3NAS can achieve better performance and fewer parameters than several state-of-the-art methods. In addition, we also validate the effectiveness of the multi-scale and multi-level architecture for LDCT denoising, and present further analysis for different configurations of super-net.
Wenjun Xia, Yongqiang Huang 0003, Mingzheng Hou, Hu Chen 0002, Jiliu Zhou, Hongming Shan, Yi Zhang 0018
IEEE Trans. Medical Imaging2
2022 A Transformer-Based Iterative Reconstruction Model for Sparse-View CT Reconstruction
Wenjun Xia, Ziyuan Yang 0001, Qizheng Zhou, Zhongxian Wang, Yi Zhang 0018
MICCAI (6)1
2022 FONT-SIR: Fourth-Order Nonlocal Tensor Decomposition Model for Spectral CT Image Reconstruction
abstract
Spectral computed tomography (CT) reconstructs images from different spectral data through photon counting detectors (PCDs). However, due to the limited number of photons and the counting rate in the corresponding spectral segment, the reconstructed spectral images are usually affected by severe noise. In this paper, we propose a fourth-order nonlocal tensor decomposition model for spectral CT image reconstruction (FONT-SIR). To maintain the original spatial relationships among similar patches and improve the imaging quality, similar patches without vectorization are grouped in both spectral and spatial domains simultaneously to form the fourth-order processing tensor unit. The similarity of different patches is measured with the cosine similarity of latent features extracted using principal component analysis (PCA). By imposing the constraints of the weighted nuclear and total variation (TV) norms, each fourth-order tensor unit is decomposed into a low-rank component and a sparse component, which can efficiently remove noise and artifacts while preserving the structural details. Moreover, the alternating direction method of multipliers (ADMM) is employed to solve the decomposition model. Extensive experimental results on both simulated and real data sets demonstrate that the proposed FONT-SIR achieves superior qualitative and quantitative performance compared with several state-of-the-art methods.
Xiang Chen 0015, Wenjun Xia, Yan Liu 0052, Hu Chen 0002, Jiliu Zhou, Zhiyuan Zha, Bihan Wen, Yi Zhang 0018
IEEE Trans. Medical Imaging2
2021 Dual-Domain Adaptive-Scaling Non-local Network for CT Metal Artifact Reduction
Tao Wang 0167, Wenjun Xia, Yongqiang Huang 0003, Huaiqiang Sun, Yan Liu 0052, Hu Chen 0002, Jiliu Zhou, Yi Zhang 0018
MICCAI (6)2
2021 Noise-Powered Disentangled Representation for Unsupervised Speckle Reduction of Optical Coherence Tomography Images
abstract
Due to its noninvasive character, optical coherence tomography (OCT) has become a popular diagnostic method in clinical settings. However, the low-coherence interferometric imaging procedure is inevitably contaminated by heavy speckle noise, which impairs both visual quality and diagnosis of various ocular diseases. Although deep learning has been applied for image denoising and achieved promising results, the lack of well-registered clean and noisy image pairs makes it impractical for supervised learning-based approaches to achieve satisfactory OCT image denoising results. In this paper, we propose an unsupervised OCT image speckle reduction algorithm that does not rely on well-registered image pairs. Specifically, by employing the ideas of disentangled representation and generative adversarial network, the proposed method first disentangles the noisy image into content and noise spaces by corresponding encoders. Then, the generator is used to predict the denoised OCT image with the extracted content features. In addition, the noise patches cropped from the noisy image are utilized to facilitate more accurate disentanglement. Extensive experiments have been conducted, and the results suggest that our proposed method is superior to the classic methods and demonstrates competitive performance to several recently proposed learning-based approaches in both quantitative and qualitative aspects. Code is available at: https://github.com/tsmotlp/DRGAN-OCT.
Yongqiang Huang 0003, Wenjun Xia, Yan Liu 0052, Hu Chen 0002, Jiliu Zhou, Leyuan Fang, Yi Zhang 0018
IEEE Trans. Medical Imaging2
2021 CT Reconstruction With PDF: Parameter-Dependent Framework for Data From Multiple Geometries and Dose Levels
abstract
The current mainstream computed tomography (CT) reconstruction methods based on deep learning usually need to fix the scanning geometry and dose level, which significantly aggravates the training costs and requires more training data for real clinical applications. In this paper, we propose a parameter-dependent framework (PDF) that trains a reconstruction network with data originating from multiple alternative geometries and dose levels simultaneously. In the proposed PDF, the geometry and dose level are parameterized and fed into two multilayer perceptrons (MLPs). The outputs of the MLPs are used to modulate the feature maps of the CT reconstruction network, which condition the network outputs on different geometries and dose levels. The experiments show that our proposed method can obtain competitive performance compared to the original network trained with either specific or mixed geometry and dose level, which can efficiently save extra training costs for multiple geometries and dose levels.
Wenjun Xia, Yongqiang Huang 0003, Yan Liu 0052, Hu Chen 0002, Jiliu Zhou, Yi Zhang 0018
IEEE Trans. Medical Imaging1
2021 MAGIC: Manifold and Graph Integrative Convolutional Network for Low-Dose CT Reconstruction
abstract
Low-dose computed tomography (LDCT) scans, which can effectively alleviate the radiation problem, will degrade the imaging quality. In this paper, we propose a novel LDCT reconstruction network that unrolls the iterative scheme and performs in both image and manifold spaces. Because patch manifolds of medical images have low-dimensional structures, we can build graphs from the manifolds. Then, we simultaneously leverage the spatial convolution to extract the local pixel-level features from the images and incorporate the graph convolution to analyze the nonlocal topological features in manifold space. The experiments show that our proposed method outperforms both the quantitative and qualitative aspects of state-of-the-art methods. In addition, aided by a projection loss component, our proposed method also demonstrates superior performance for semi-supervised learning. The network can remove most noise while maintaining the details of only 10% (40 slices) of the training data labeled.
Wenjun Xia, Yongqiang Huang 0003, Zuoqiang Shi, Yan Liu 0052, Hu Chen 0002, Yang Chen 0008, Jiliu Zhou, Yi Zhang 0018
IEEE Trans. Medical Imaging1
2020 Disentanglement Network for Unsupervised Speckle Reduction of Optical Coherence Tomography Images
Yongqiang Huang 0003, Wenjun Xia, Yan Liu 0052, Jiliu Zhou, Leyuan Fang, Yi Zhang 0018
MICCAI (5)2
2019 Convolutional Sparse Coding for Compressed Sensing CT Reconstruction
abstract
Over the past few years, dictionary learning (DL)-based methods have been successfully used in various image reconstruction problems. However, the traditional DL-based computed tomography (CT) reconstruction methods are patch-based and ignore the consistency of pixels in overlapped patches. In addition, the features learned by these methods always contain shifted versions of the same features. In recent years, convolutional sparse coding (CSC) has been developed to address these problems. In this paper, inspired by several successful applications of CSC in the field of signal processing, we explore the potential of CSC in sparse-view CT reconstruction. By directly working on the whole image, without the necessity of dividing the image into overlapped patches in DL-based methods, the proposed methods can maintain more details and avoid artifacts caused by patch aggregation. With predetermined filters, an alternating scheme is developed to optimize the objective function. Extensive experiments with simulated and real CT data were performed to validate the effectiveness of the proposed methods. The qualitative and quantitative results demonstrate that the proposed methods achieve better performance than the several existing state-of-the-art methods.
Peng Bao 0001, Huaiqiang Sun, Zhangyang Wang, Yi Zhang 0018, Wenjun Xia, Mianyi Chen, Yan Xi, Shanzhou Niu, Jiliu Zhou, He Zhang 0004
IEEE Trans. Medical Imaging5
2016 A Nearest Neighbor Classifier Employing Critical Boundary Vectors for Efficient On-Chip Template Reduction
abstract
Aiming at efficient data condensation and improving accuracy, this paper presents a hardware-friendly template reduction (TR) method for the nearest neighbor (NN) classifiers by introducing the concept of critical boundary vectors. A hardware system is also implemented to demonstrate the feasibility of using an field-programmable gate array (FPGA) to accelerate the proposed method. Initially, k -means centers are used as substitutes for the entire template set. Then, to enhance the classification performance, critical boundary vectors are selected by a novel learning algorithm, which is completed within a single iteration. Moreover, to remove noisy boundary vectors that can mislead the classification in a generalized manner, a global categorization scheme has been explored and applied to the algorithm. The global characterization automatically categorizes each classification problem and rapidly selects the boundary vectors according to the nature of the problem. Finally, only critical boundary vectors and k -means centers are used as the new template set for classification. Experimental results for 24 data sets show that the proposed algorithm can effectively reduce the number of template vectors for classification with a high learning speed. At the same time, it improves the accuracy by an average of 2.17% compared with the traditional NN classifiers and also shows greater accuracy than seven other TR methods. We have shown the feasibility of using a proof-of-concept FPGA system of 256 64-D vectors to accelerate the proposed method on hardware. At a 50-MHz clock frequency, the proposed system achieves a 3.86 times higher learning speed than on a 3.4-GHz PC, while consuming only 1% of the power of that used by the PC.
Wenjun Xia, Yoshio Mita, Tadashi Shibata
IEEE Trans. Neural Networks Learn. Syst.1
2012 Self-adaptive quasi-Gaussian circuits for analog on-chip-trainable multi-class classifiers
abstract
Self-adaptive quasi-Gaussian circuits have been developed and introduced to an analog multi-class classifier in order to enhance its classification performance. By applying a floating threshold scheme to the quasi-Gaussian kernel, the kernel can extend its tail region adaptively according to the characteristics of input data. As a result, the misclassification problem due to the zero tail region in the quasi-Gaussian kernel has been completely eliminated, and the classification accuracy is significantly improved. Software simulation showed the performance is comparable to complex Gaussian-kernel Support Vector Machines. A proof-of-concept chip implementing an analog on-chip-trainable multi-class classifier which employs 64-dimensional self-adaptive quasi-Gaussian circuits was designed in a 0.18-μm CMOS technology and is now under fabrication. Its successful operation was confirmed by Nanosim simulation.
Wenjun Xia, Tadashi Shibata
ISCAS1