Jian Zheng 0001

dblp:49/3161-1 · DBLP profile ↗
← Back
20ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0003-4257-331XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021
YearPublicationVenuePosition
2026 Sparse-View CT Reconstruction via Implicit Neural Representation Learning Powered by Dual-Domain Vision Foundation Models
abstract
Sparse-view computed tomography (SVCT) offers the advantages of accelerated scanning and reduced X-ray radiation dose in different clinical applications. However, it faces a challenge due to incomplete data acquisition, resulting in streak artifacts in the analytically reconstructed CT images. Utilizing self-supervised learning, implicit neural representation (INR) recently has shown great promise in addressing inverse problems such as SVCT reconstruction. Nonetheless, given that the input of original INR only contains coordinate information, it is limited to represent one SVCT instance at a time, and its performance significantly declines when performing cross-instance reconstruction. In this study, we propose a novel self-supervised framework named VFMINR, which leverages generalizable representations extracted from the visual foundation models (VFMs) to tackle the cross-instance reconstruction issue of INR. Specifically, VFMINR first utilizes VFMs to effectively capture the spatial and frequency domain representations of sinograms, and then a fusion module is applied to fuse two domain features into complementary representations. This combination maximizes the utilization of local detail information from the spatial domain and the global structural information from the frequency domain. Subsequently, an adaptive cell decoding strategy is designed to map representations into variable resolution hybrid feature grids, which are integrated into the learning of the INR to enhance its generalizability for different SVCT instances. The VFMs and VFMINR are trained by using only SV sinogram data, and extensive results confirm that the proposed method can effectively handle the generalization problem of INR, while achieving superior performance in image fidelity and artifact suppression. The code is available at: https://github.com/nightastars/VFMINR-main.
Yang Chen 0008, Yangchuan Liu, Zhongyi Wu, Hengyong Yu, Jian Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.10
2026 MaKAN-Mixer: Channel Interaction-Based Mamba Method for rPPG Extraction
abstract
Remote photoplethysmography (rPPG) achieves non-contact heart rate monitoring by detecting subtle skin color variations in facial videos, offering significant potential in healthcare, fitness, and security applications.However, accurately extracting rPPG signals in complex environments-especially under variable lighting and motion artifacts-remains challenging. The main difficulties are capturing spatio-temporal dynamics and modeling long-term dependencies across channels. To address these limitations, we propose MaKAN-Mixer, a novel end-to-end network designed to enhance the robustness and accuracy of rPPG signal extraction. First, MaKAN-Mixer integrates a Hybrid of Eulerian Video Magnification and Temporal Shift Module Amplification (HETA) to amplify subtle physiological signals and enhance temporal information without relying on explicit region-of-interest (ROI) selection. Additionally, we propose the Mamba-KAN Fusion Module (MKFM), which leverages Mamba's ability to efficiently model long-term dependencies in temporal sequences. By incorporating the Kolmogorov-Arnold Network (KAN) for effective channel mixing, MKFM ensures the comprehensive fusion of relevant spatio-temporal features across different channels. Finally, we employ a KAN Feedforward Neural Network (KFN) to capture complex, nonlinear, and periodic physiological patterns, improving heart rate estimation. Extensive experiments conducted on four benchmark datasets demonstrate that MaKAN-Mixer achieves superior performance in both intra- and cross-dataset testing, exhibiting exceptional robustness in challenging scenarios, particularly with compressed video data and complex environments. In comparison to the best-performing existing method, which reported RMSE values of 0.78/0.47/4.57/6.81 on the four datasets, MaKAN-Mixer significantly improves the RMSE to 0.66/0.40/0.32/6.25, highlighting its effectiveness across diverse conditions. Furthermore, novel visualization techniques were employed for qualitative validation of the results, underscoring its potential for accurate, real-world rPPG monitoring.
Feiyang Liao, Haoyang Jin, Biao Xie, Mingcui Fu, Jian Zheng 0001
IEEE J. Biomed. Health Informatics8
2025 FDF-VQVAE: A Frequency Disentanglement and Fusion Learning Framework for Multi-sequence MRI Enhancement
Xinghe Xie, Luyi Han, Yue Sun 0001, Chi Kin Lam, Jian Zheng 0001, Tong Tong 0001, Wei Ke 0001, Chan-Tong Lam, Tao Tan 0002
MICCAI (3)5
2025 Paired Image Generation with Diffusion-Guided Diffusion Models
Haoxuan Zhang, Wenju Cui, Yuzhu Cao, Tao Tan 0002, Yunsong Peng, Jian Zheng 0001
MICCAI (4)7
2025 Adaptive critical subgraph mining for cognitive impairment conversion prediction with T1-MRI-based brain network
Yilin Leng, Wenju Cui, Xi Jiang 0001, Yunsong Peng, Jian Zheng 0001
Expert Syst. Appl.6
2025 Domain Progressive Low-Dose CT Imaging Using Iterative Partial Diffusion Model
abstract
Traditional deep learning reconstruction (DLR) methods have been sparsely applied in practical low-dose computed tomography (LDCT) imaging, as they heavily rely on the similarity between the latent distributions of data features. However, in real LDCT imaging scenarios, the distribution of data features is highly diverse and complex, which limits the generalizability of existing DLR methods. Recently, diffusion models have shown great potential in the field of LDCT imaging, and some early studies have used them to address the domain generalization problem. However, they still face challenges such as high time consumption, difficulties in training with high resolution, and performance degradation in denoising scenario. In this paper, we propose a novel domain progressive LDCT imaging framework with an iterative partial diffusion model (IPDM) as the core. Firstly, the derived IPDM theoretical framework supports completing the denoising task by iterating a small part of the complete diffusion model, utilizing the strong generation ability of the diffusion model while alleviating time consumption and convergence difficulties. Secondly, a derived condition guided sampling method alleviates sampling bias caused by deviations of the predictive data gradient and Langevin dynamics. Finally, an adaptive weight strategy based on pixel-wise noise estimation can gradually adjust guided intensity. Extensive testing on diverse datasets reveals that our method outperforms traditional iterative reconstructions, unsupervised, and some supervised DLR methods in visual and quantitative evaluations, closely matching the performance of state-of-the-art supervised DLR techniques. Additionally, our IPDM was trained using practical normal-dose CT data, rather than the tested LDCT data. This enables our method to have better generalization ability compared to traditional DLR methods in practical imaging scenarios. Source code is available at https://github.com/LFY1998/IPDM-PyTorch.
Feiyang Liao, Yufei Tang, Jian Zheng 0001
IEEE Trans. Medical Imaging6
2025 A Novel Dynamic Neural Network for Heterogeneity-Aware Structural Brain Network Exploration and Alzheimer's Disease Diagnosis
abstract
Heterogeneity is a fundamental characteristic of brain diseases, distinguished by variability not only in brain atrophy but also in the complexity of neural connectivity and brain networks. However, existing data-driven methods fail to provide a comprehensive analysis of brain heterogeneity. Recently, dynamic neural networks (DNNs) have shown significant advantages in capturing sample-wise heterogeneity. Therefore, in this article, we first propose a novel dynamic heterogeneity-aware network (DHANet) to identify critical heterogeneous brain regions, explore heterogeneous connectivity between them, and construct a heterogeneous-aware structural brain network (HGA-SBN) using structural magnetic resonance imaging (sMRI). Specifically, we develop a 3-D dynamic convmixer to extract abundant heterogeneous features from sMRI first. Subsequently, the critical brain atrophy regions are identified by dynamic prototype learning with embedding the hierarchical brain semantic structure. Finally, we employ a joint dynamic edge-correlation (JDE) modeling approach to construct the heterogeneous connectivity between these regions and analyze the HGA-SBN. To evaluate the effectiveness of the DHANet, we conduct elaborate experiments on three public datasets and the method achieves state-of-the-art (SOTA) performance on two classification tasks.
Wenju Cui, Yilin Leng, Yunsong Peng, Lei Li 0058, Xi Jiang 0001, Jian Zheng 0001
IEEE Trans. Neural Networks Learn. Syst.8
2024 Domain adaptive noise reduction with iterative knowledge transfer and style generalization learning
Yufei Tang, Tianling Lyu, Haoyang Jin, Yang Chen 0008, Jian Zheng 0001
Medical Image Anal.9
2024 SAH-NET: Structure-Aware Hierarchical Network for Clustered Microcalcification Classification in Digital Breast Tomosynthesis
abstract
Benign and malignant classification of clustered microcalcifications (MCs) in digital breast tomosynthesis (DBT) is an essential task in computer-aided diagnosis. However, due to the anisotropic resolution of DBT, three-dimensional (3-D) convolutional neural network (CNN)-based methods cannot extract hierarchical features efficiently. Moreover, the sparse distribution of MC points in the cluster makes it difficult for the CNN to extract discriminative structural information for classification. To comprehensively address these challenges, we propose a novel structure-aware hierarchical network (SAH-Net) for benign and malignant classification of clustered MC in a DBT volume. Specifically, the two-dimensional (2-D) group convolution is used to extract intraslice features. The one-to-one correspondence between group convolutions and slices ensures the independence of hierarchical feature extraction. Then, a partial deformable Transformer-based 3-D structural feature learning module is proposed to capture the long-range dependency between MC points in the cluster. We evaluate the proposed method on an in-house dataset with 495 clustered MCs collected from 462 DBT images. Experimental results confirm the validity of our proposed modules. The results also show that the proposed SAH-Net outperforms several other representative methods on this topic, and achieves the best classification result, with an area under the receiver operation curve (AUC) of 86.87%. The implementation of the proposed model is available at https://github.com/sunhaotian130911/SAHNet.
Shandong Wu, Xinjian Chen 0001, Lingji Kong, Xiaodong Yang 0005, You Meng, Shuangqing Chen, Jian Zheng 0001
IEEE Trans. Cybern.9
2024 NeighborNet: Learning Intra- and Inter-Image Pixel Neighbor Representation for Breast Lesion Segmentation
abstract
Breast lesion segmentation from ultrasound images is essential in computer-aided breast cancer diagnosis. To alleviate the problems of blurry lesion boundaries and irregular morphologies, common practices combine CNN and attention to integrate global and local information. However, previous methods use two independent modules to extract global and local features separately, such feature-wise inflexible integration ignores the semantic gap between them, resulting in representation redundancy/insufficiency and undesirable restrictions in clinic practices. Moreover, medical images are highly similar to each other due to the imaging methods and human tissues, but the captured global information by transformer-based methods in the medical domain is limited within images, the semantic relations and common knowledge across images are largely ignored. To alleviate the above problems, in the neighbor view, this paper develops a pixel neighbor representation learning method (NeighborNet) to flexibly integrate global and local context within and across images for lesion morphology and boundary modeling. Concretely, we design two neighbor layers to investigate two properties (i.e., number and distribution) of neighbors. The neighbor number for each pixel is not fixed but determined by itself. The neighbor distribution is extended from one image to all images in the datasets. With the two properties, for each pixel at each feature level, the proposed NeighborNet can evolve into the transformer or degenerate into the CNN for adaptive context representation learning to cope with the irregular lesion morphologies and blurry boundaries. The state-of-the-art performances on three ultrasound datasets prove the effectiveness of the proposed NeighborNet.
Xiaohui You, Lei Li 0058, Wenju Cui, Yuzhu Cao, Xinjian Chen 0001, Jian Zheng 0001
IEEE J. Biomed. Health Informatics9
2024 Adversarial Learning Based Node-Edge Graph Attention Networks for Autism Spectrum Disorder Identification
abstract
Graph neural networks (GNNs) have received increasing interest in the medical imaging field given their powerful graph embedding ability to characterize the non-Euclidean structure of brain networks based on magnetic resonance imaging (MRI) data. However, previous studies are largely node-centralized and ignore edge features for graph classification tasks, resulting in moderate performance of graph classification accuracy. Moreover, the generalizability of GNN model is still far from satisfactory in brain disorder [e.g., autism spectrum disorder (ASD)] identification due to considerable individual differences in symptoms among patients as well as data heterogeneity among different sites. In order to address the above limitations, this study proposes a novel adversarial learning-based node-edge graph attention network (AL-NEGAT) for ASD identification based on multimodal MRI data. First, both node and edge features are modeled based on structural and functional MRI data to leverage complementary brain information and preserved in the constructed weighted adjacent matrix for individuals through the attention mechanism in the proposed NEGAT. Second, two AL methods are employed to improve the generalizability of NEGAT. Finally, a gradient-based saliency map strategy is utilized for model interpretation to identify important brain regions and connections contributing to the classification. Experimental results based on the public Autism Brain Imaging Data Exchange I (ABIDE I) data demonstrate that the proposed framework achieves a classification accuracy of 74.7% between ASD and typical developing (TD) groups based on 1007 subjects across 17 different sites and outperforms the state-of-the-art methods, indicating satisfying classification ability and generalizability of the proposed AL-NEGAT model. Our work provides a powerful tool for brain disorder identification.
Yuzhong Chen 0002, Jiadong Yan, Mingxin Jiang, Zhongbo Zhao, Weihua Zhao, Jian Zheng 0001, Dezhong Yao 0001, Keith M. Kendrick, Xi Jiang 0001
IEEE Trans. Neural Networks Learn. Syst.7
2023 Dynamic Structural Brain Network Construction by Hierarchical Prototype Embedding GCN Using T1-MRI
Yilin Leng, Wenju Cui, Jian Zheng 0001
MICCAI (8)6
2023 ICL-Net: Global and Local Inter-Pixel Correlations Learning Network for Skin Lesion Segmentation
abstract
Skin lesion segmentation is a fundamental procedure in computer-aided melanoma diagnosis. However, due to the diverse shape, variable size, blurry boundary, and noise interference of lesion regions, existing methods may struggle with the challenge of inconsistency within classes and indiscrimination between classes. In view of this, we propose a novel method to learn and model inter-pixel correlations from both global and local aspects, which can increase inter-class variances and intra-class similarities. Specifically, under the encoder-decoder architecture, we first design a pyramid transformer inter-pixel correlations (PTIC) module, aiming at capturing the non-local context information of different levels and further exploring the global pixel-level relationship to deal with the large variance of shape and size. Further, we devise a local neighborhood metric learning (LNML) module to strengthen the local semantic correlations learning capability and increase the separability between classes in the feature space. These two modules can complementarily strengthen the feature representation capability via exploiting the inter-pixel semantic correlations, thus further improving intra-class consistency and inter-class variance. Comprehensive experiments are performed on public skin lesion segmentation datasets: ISIC 2018, ISIC2016, and PH2, and experimental results demonstrate that the proposed method achieves better segmentation performance than other state-of-the-art methods.
Qi Liu 0003, Chengtao Peng, Xiaodong Yang 0005, Xinye Ni, Jian Zheng 0001
IEEE J. Biomed. Health Informatics8
2023 Low-Dose CT Image Synthesis for Domain Adaptation Imaging Using a Generative Adversarial Network With Noise Encoding Transfer Learning
abstract
Deep learning (DL) based image processing methods have been successfully applied to low-dose x-ray images based on the assumption that the feature distribution of the training data is consistent with that of the test data. However, low-dose computed tomography (LDCT) images from different commercial scanners may contain different amounts and types of image noise, violating this assumption. Moreover, in the application of DL based image processing methods to LDCT, the feature distributions of LDCT images from simulation and clinical CT examination can be quite different. Therefore, the network models trained with simulated image data or LDCT images from one specific scanner may not work well for another CT scanner and image processing task. To solve such domain adaptation problem, in this study, a novel generative adversarial network (GAN) with noise encoding transfer learning (NETL), or GAN-NETL, is proposed to generate a paired dataset with a different noise style. Specifically, we proposed a method to perform noise encoding operator and incorporate it into the generator to extract a noise style. Meanwhile, with a transfer learning (TL) approach, the image noise encoding operator transformed the noise type of the source domain to that of the target domain for realistic noise generation. One public and two private datasets are used to evaluate the proposed method. Experiment results demonstrated the feasibility and effectiveness of our proposed GAN-NETL model in LDCT image synthesis. In addition, we conduct additional image denoising study using the synthesized clinical LDCT data, which verified the merit of the proposed synthesis in improving the performance of the DL based LDCT processing method.
Yang Chen 0008, Yufei Tang, Zhongyi Wu, Yujin Qi, Haochuan Jiang, Jian Zheng 0001, Benjamin M. W. Tsui
IEEE Trans. Medical Imaging8
2021 3D Context-Aware Convolutional Neural Network for False Positive Reduction in Clustered Microcalcifications Detection
abstract
False positives (FPs) reduction is indispensable for clustered microcalcifications (MCs) detection in digital breast tomosynthesis (DBT), since there might be excessive false candidates in the detection stage. Considering that DBT volume has an anisotropic resolution, we proposed a novel 3D context-aware convolutional neural network (CNN) to reduce FPs, which consists of a 2D intra-slices feature extraction branch and a 3D inter-slice features fusion branch. In particular, 3D anisotropic convolutions were designed to learn representations from DBT volumes and inter-slice information fusion is only performed on the feature map level, which could avoid the influence of anisotropic resolution of DBT volume. The proposed method was evaluated on a large-scale Chinese women population of 877 cases with 1754 DBT volumes and compared with 8 related methods. Experimental results show that the proposed network achieved the best performance with an accuracy of 92.68% for FPs reduction with an AUC of 97.65%, and the FPs are 0.0512 per DBT volume at a sensitivity of 90%. This also proved that making full use of 3D contextual information of DBT volume can improve the performance of the classification algorithm.
Jian Zheng 0001, Shandong Wu, Yunsong Peng, Xiaodong Yang 0005
IEEE J. Biomed. Health Informatics1
2020 OCTRexpert: A Feature-Based 3D Registration Method for Retinal OCT Images
abstract
Medical image registration can be used for studying longitudinal and cross-sectional data, quantitatively monitoring disease progression and guiding computer assisted diagnosis and treatments. However, deformable registration which enables more precise and quantitative comparison has not been well developed for retinal optical coherence tomography (OCT) images. This paper proposes a new 3D registration approach for retinal OCT data called OCTRexpert. To the best of our knowledge, the proposed algorithm is the first full 3D registration approach for retinal OCT images which can be applied to longitudinal OCT images for both normal and serious pathological subjects. In this approach, a pre-processing method is first performed to remove eye motion artifact and then a novel design-detection-deformation strategy is applied for the registration. In the design step, a couple of features are designed for each voxel in the image. In the detection step, active voxels are selected and the point-to-point correspondences between the subject and template images are established. In the deformation step, the image is hierarchically deformed according to the detected correspondences in multi-resolution. The proposed method is evaluated on a dataset with longitudinal OCT images from 20 healthy subjects and 4 subjects diagnosed with serious Choroidal Neovascularization (CNV). Experimental results show that the proposed registration algorithm consistently yields statistically significant improvements in both Dice similarity coefficient and the average unsigned surface error compared with the other registration methods.
Lingjiao Pan, Dehui Xiang, Kai Yu 0009, Luwen Duan, Jian Zheng 0001, Xinjian Chen 0001
IEEE Trans. Image Process.6
2020 A Cross-Domain Metal Trace Restoring Network for Reducing X-Ray CT Metal Artifacts
abstract
Metal artifacts commonly appear in computed tomography (CT) images of the patient body with metal implants and can affect disease diagnosis. Known deep learning and traditional metal trace restoring methods did not effectively restore details and sinogram consistency information in X-ray CT sinograms, hence often causing considerable secondary artifacts in CT images. In this paper, we propose a new cross-domain metal trace restoring network which promotes sinogram consistency while reducing metal artifacts and recovering tissue details in CT images. Our new approach includes a cross-domain procedure that ensures information exchange between the image domain and the sinogram domain in order to help them promote and complement each other. Under this cross-domain structure, we develop a hierarchical analytic network (HAN) to recover fine details of metal trace, and utilize the perceptual loss to guide HAN to concentrate on the absorption of sinogram consistency information of metal trace. To allow our entire cross-domain network to be trained end-to-end efficiently and reduce the graphic memory usage and time cost, we propose effective and differentiable forward projection (FP) and filtered back-projection (FBP) layers based on FP and FBP algorithms. We use both simulated and clinical datasets in three different clinical scenarios to evaluate our proposed network's practicality and universality. Both quantitative and qualitative evaluation results show that our new network outperforms state-of-the-art metal artifact reduction methods. In addition, the elapsed time analysis shows that our proposed method meets the clinical time requirement.
Chengtao Peng, Bin Li 0025, Peixian Liang, Jian Zheng 0001, Yizhe Zhang 0001, Bensheng Qiu, Danny Ziyi Chen
IEEE Trans. Medical Imaging4
2019 Nonrigid Image Registration Using Spatially Region-Weighted Correlation Ratio and GPU-Acceleration
abstract
OBJECTIVE: Nonrigid image registration with high accuracy and efficiency remains a challenging task for medical image analysis. In this paper, we present the spatially region-weighted correlation ratio (SRWCR) as a novel similarity measure to improve the registration performance. METHODS: SRWCR is rigorously deduced from a three-dimension joint probability density function combining the intensity channels with an extra spatial information channel. SRWCR estimates the optimal functional dependence between the intensities for each spatial bin, in which the spatial distribution modeled by a cubic B-spline function is used to differentiate the contribution of voxels. We also analytically derive the gradient of SRWCR with respect to the transformation parameters and optimize it using a quasi-Newton approach. Furthermore, we propose a GPU-based parallel mechanism to accelerate the computation of SRWCR and its derivatives. RESULTS: The experiments on synthetic images, public four-dimensional thoracic computed tomography (CT) dataset, retinal optical coherence tomography data, and clinical CT and positron emission tomography images confirm that SRWCR significantly outperforms some state-of-the-art techniques such as spatially encoded mutual information and Robust PaTch-based cOrrelation Ration. CONCLUSION: This study demonstrates the advantages of SRWCR in tackling the practical difficulties due to distinct intensity changes, serious speckle noise, or different imaging modalities. SIGNIFICANCE: The proposed registration framework might be more reliable to correct the nonrigid deformations and more potential for clinical applications.
Lun Gong, Luwen Duan, Xueying Du, Hanqiu Liu, Xinjian Chen 0001, Jian Zheng 0001
IEEE J. Biomed. Health Informatics7
2011 Salient Feature Region: A New Method for Retinal Image Registration
abstract
Retinal image registration is crucial for the diagnoses and treatments of various eye diseases. A great number of methods have been developed to solve this problem; however, fast and accurate registration of low-quality retinal images is still a challenging problem since the low content contrast, large intensity variance as well as deterioration of unhealthy retina caused by various pathologies. This paper provides a new retinal image registration method based on salient feature region (SFR). We first propose a well-defined region saliency measure that consists of both local adaptive variance and gradient field entropy to extract the SFRs in each image. Next, an innovative local feature descriptor that combines gradient field distribution with corresponding geometric information is then computed to match the SFRs accurately. After that, normalized cross-correlation-based local rigid registration is performed on those matched SFRs to refine the accuracy of local alignment. Finally, the two images are registered by adopting high-order global transformation model with locally well-aligned region centers as control points. Experimental results show that our method is quite effective for retinal image registration.
Jian Zheng 0001, Jie Tian 0001, Kexin Deng, Xiaoqian Dai
IEEE Trans. Inf. Technol. Biomed.1
2008 A Novel Software Platform for Medical Image Processing and Analyzing
abstract
The design of software platform for medical imaging application has been increasingly prioritized as the sophisticated application of medical imaging. With this demand, we have designed and implemented a novel software platform in traditional object-oriented fashion with some common design patterns. This platform integrates the mainstream algorithms for medical image processing and analyzing within a consistent framework, including reconstruction, segmentation, registration, visualization, etc., and provides a powerful tool for both scientists and engineers. The overall framework and certain key technologies are introduced in detail. Presented experiment examples, numerous downloads, extensive uses, and practical applications commendably demonstrate the validity and flexibility of the platform.
Jie Tian 0001, Jian Xue 0002, Yakang Dai, Jian Chen 0014, Jian Zheng 0001
IEEE Trans. Inf. Technol. Biomed.5