Zixuan Pan

dblp:319/2483 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0003-0177-1520ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 NeFT: Negative Feedback Training to Improve Robustness of Compute-in-Memory DNN Accelerators
abstract
Compute-in-memory accelerators built upon non-volatile memory devices excel in energy efficiency and latency when performing deep neural network (DNN) inference, thanks to their in-situ data processing capability. However, the stochastic nature and intrinsic variations of non-volatile memory devices often result in performance degradation during DNN inference. Introducing these non-ideal device behaviors in DNN training enhances robustness, but drawbacks include limited accuracy improvement, reduced prediction confidence, and convergence issues. This arises from a mismatch between the deterministic training and non-deterministic device variations, as such training, though considering variations, relies solely on the model’s final output. In this work, inspired by control theory, we propose Negative Feedback Training (NeFT)—a novel concept supported by theoretical analysis—to more effectively capture the multi-scale noisy information throughout the network. We instantiate this concept with two specific instances, oriented variational forward (OVF) and intermediate representation snapshot (IRS). Based on device variation models extracted from measured data, extensive experiments show that our NeFT outperforms existing state-of-the-art methods with up to a 45.08% improvement in inference accuracy while reducing epistemic uncertainty, boosting output confidence, and improving convergence probability. These results underline the generality and practicality of our NeFT framework for increasing the robustness of DNNs against device variations. The source code for these two instances is available at https://github.com/YifanQin-ND/NeFT_CIM.
Zheyu Yan, Dailin Gan, Jun Xia 0003, Zixuan Pan, Wujie Wen, Xiaobo Sharon Hu, Yiyu Shi 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2025 H-CNN-ViT: A Hierarchical Gated Attention Multi-Branch Model for Bladder Cancer Recurrence Prediction
abstract
Bladder cancer is one of the most prevalent malignancies worldwide, with a recurrence rate of up to 78 %, necessitating accurate post-operative monitoring for effective patient management. Multi-sequence contrast-enhanced MRI is commonly used for recurrence detection; however, interpreting these scans remains challenging, even for experienced radiologists, due to post-surgical alterations such as scarring, swelling, and tissue remodeling. AI-assisted diagnostic tools have shown promise in improving bladder cancer recurrence prediction, yet progress in this field is hindered by the lack of dedicated multi-sequence MRI datasets for recurrence assessment study. In this work, we first introduce a curated multi-sequence, multimodal MRI dataset specifically designed for bladder cancer recurrence prediction, establishing a valuable benchmark for future research. We then propose H-CNN-ViT, a new Hierarchical Gated Attention Multi-Branch model that enables selective weighting of features from the global (ViT) and local (CNN) paths based on contextual demands, achieving a balanced and targeted feature fusion. Our multi-branch architecture processes each modality independently, ensuring that the unique properties of each imaging channel are optimally captured and integrated. Evaluated on our dataset, H-CNN-ViT achieves an AUC of 78.6 %, surpassing state-of-the-art models. Our model is publicly available at https://github.com/XLIAaron/H-CNN-ViT.
Zongren Wang, Zixuan Pan, Nishchal Sapkota, Gelei Xu, Danny Ziyi Chen, Yiyu Shi 0001
BIBM4
2025 Multi-Exit Class Activation Map Guided Feature Masking for Unsupervised Out-of-Distribution Detection in Medical Imaging
abstract
Out-of-distribution (OOD) detection is crucial for ensuring the safety and reliability of deep learning models in high-stakes domains such as medical imaging. However, existing methods often struggle to detect subtle or localized anomalies, which are common in clinical settings. We hypothesize that such challenges stem in part from a limited understanding of how models focus on different image regions under ID and OOD inputs. To investigate this, we analyze the behavior of deep models under different inputs, and observe that class activation maps (CAMs) for in-distribution (ID) data typically emphasize regions that are highly relevant to the prediction of a model, whereas OOD data often lacks such focused activations. Building on this, we find that masking input images with inverted CAMs induces larger shifts in feature representations for ID than OOD data, a signal that can be leveraged for robust detection. Based on this insight, we propose Multi-Exit Class Activation Map (MECAM), a novel unsupervised OOD detection framework that integrates aggregated multi-exit CAMs and CAM-guided feature masking. By combining CAMs from multiple network depths, our method captures both global and local feature representations, thereby enhancing the robustness of OOD detection. We evaluate MECAM on two ID datasets, including ISIC19 and PathMNIST, and test its performance against three medical OOD datasets, RSNA Pneumonia, COVID-19, and HeadCT, and one natural image OOD dataset, iSUN. Comprehensive experiments demonstrate that MECAM consistently outperforms state-of-theart OOD detection methods, validating its effectiveness. These findings highlight the potential of multi-exit architectures and CAM-guided feature masking in advancing unsupervised OOD detection for medical imaging, paving the way for more reliable and interpretable models in clinical practice. The source code is available at https://github.com/zx-pan/MECAM-OOD.
Zixuan Pan, Jun Xia 0003, Max Ficco, Jianxu Chen 0001, Tsung-Yi Ho, Yiyu Shi 0001
BIBM1
2025 Rethinking Medical Anomaly Detection in Brain MRI: An Image Quality Assessment Perspective
abstract
Reconstruction-based methods, particularly those leveraging autoencoders, have been widely adopted for anomaly detection task in brain MRI. Unlike most existing works try to improve the task accuracy through architectural or algorithmic innovations, we tackle this task from image quality assessment (IQA) perspective, an under-explored direction in the field. Due to the limitations of conventional metrics such as £1 in capturing the nuanced differences in reconstructed images for medical anomaly detection, we propose fusion quality, a novel metric that wisely integrates the structure-level sensitivity of Structural Similarity Index Measure (SSIM) with the pixel-level precision of £1. The metric offers a more comprehensive assessment of reconstruction quality, considering intensity (subtractive property of l1and divisive property of SSIM), contrast, and structural similarity. Furthermore, the proposed metric makes subtle regional variations more impactful in the final assessment. Thus, considering the inherent divisive properties of SSIM, we design an average intensity ratio (AIR)-based data transformation that amplifies the divisive discrepancies between normal and abnormal regions, thereby enhancing anomaly detection. By fusing the aforementioned two components, we devise the IQA approach. Experimental results on two distinct brain MRI datasets show that our IQA approach significantly enhances medical anomaly detection performance when integrated with state-of-the-art baselines. Code is provided here.
Zixuan Pan, Jun Xia 0003, Zheyu Yan, Guoyue Xu, Yawen Wu, Zhenge Jia, Jianxu Chen 0001, Yiyu Shi 0001
BIBM1
2024 Efficient Vision-Language Pre-Training by Cluster Masking
abstract
We propose a simple strategy for masking image patches during visual-language contrastive learning that improves the quality of the learned representations and the training speed. During each iteration of training, we randomly mask clusters of visually similar image patches, as measured by their raw pixel intensities. This provides an extra learning signal, beyond the contrastive training itself, since it forces a model to predict words for masked visual structures solely from context. It also speeds up training by reducing the amount of data used in each image. We evaluate the effectiveness of our model by pre-training on a number of bench-marks, finding that it outperforms other masking strategies, such as FLIP, on the quality of the learned representation.
Zihao Wei, Zixuan Pan, Andrew Owens
CVPR2
2024 TSB: Tiny Shared Block for Efficient DNN Deployment on NVCIM Accelerators
abstract
Compute-in-memory (CIM) accelerators using non-volatile memory (NVM) devices offer promising solutions for energy-efficient and low-latency Deep Neural Network (DNN) inference execution. However, practical deployment is often hindered by the challenge of dealing with the massive amount of model weight parameters impacted by the inherent device variations within non-volatile computing-in-memory (NVCIM) accelerators. This issue significantly offsets their advantages by increasing training overhead, the time and energy needed for mapping weights to device states, and diminishing inference accuracy. To mitigate these challenges, we propose the "Tiny Shared Block (TSB)" method, which integrates a small shared 1 × 1 convolution block into the DNN architecture. This block is designed to stabilize feature processing across the network, effectively reducing the impact of device variation. Extensive experimental results show that TSB achieves over 20× inference accuracy gap improvement, over 5× training speedup, and weights-to-device mapping cost reduction while requiring less than 0.4% of the original weights to be write-verified during programming, when compared with state-of-the-art baseline solutions. Our approach provides a practical and efficient solution for deploying robust DNN models on NVCIM accelerators, making it a valuable contribution to the field of energy-efficient AI hardware.
Zheyu Yan, Zixuan Pan, Wujie Wen, Xiaobo Sharon Hu, Yiyu Shi 0001
ICCAD3
2023 Partial Unbalanced Feature Transport for Cross-Modality Cardiac Image Segmentation
abstract
Deep learning based approaches have achieved great success on the automatic cardiac image segmentation task. However, the achieved segmentation performance remains limited due to the significant difference across image domains, which is referred to as domain shift. Unsupervised domain adaptation (UDA), as a promising method to mitigate this effect, trains a model to reduce the domain discrepancy between the source (with labels) and the target (without labels) domains in a common latent feature space. In this work, we propose a novel framework, named Partial Unbalanced Feature Transport (PUFT), for cross-modality cardiac image segmentation. Our model facilities UDA leveraging two Continuous Normalizing Flow-based Variational Auto-Encoders (CNF-VAE) and a Partial Unbalanced Optimal Transport (PUOT) strategy. Instead of directly using VAE for UDA in previous works where the latent features from both domains are approximated by a parameterized variational form, we introduce continuous normalizing flows (CNF) into the extended VAE to estimate the probabilistic posterior and alleviate the inference bias. To remove the remaining domain shift, PUOT exploits the label information in the source domain to constrain the OT plan and extracts structural information of both domains, which are often neglected in classical OT for UDA. We evaluate our proposed model on two cardiac datasets and an abdominal dataset. The experimental results demonstrate that PUFT achieves superior performance compared with state-of-the-art segmentation methods for most structural segmentation.
Shunjie Dong, Zixuan Pan, Yu Fu 0008, Dongwei Xu, Kuangyu Shi, Qianqian Yang 0002, Yiyu Shi 0001, Cheng Zhuo
IEEE Trans. Medical Imaging2
2022 DeU-Net 2.0: Enhanced deformable U-Net for 3D cardiac cine MRI segmentation
Shunjie Dong, Zixuan Pan, Yu Fu 0008, Qianqian Yang 0002, Yuanxue Gao, Tianbai Yu, Yiyu Shi 0001, Cheng Zhuo
Medical Image Anal.2