Jiajin Zhang

dblp:124/0291 · DBLP profile ↗
← Back
20ranked-venue papers
8as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 5 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Computer networks · 2Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Clinical Knowledge-Guided PET/CT Lesion Segmentation With Interpretable Fusion of Metabolic and Structural Cues
abstract
F-FDG PET/CT images marks a pivotal breakthrough in oncological diagnostics, substantially improving the accuracy and efficiency of tumor burden assessment. Manual segmentation is often plagued by significant inter-observer variability, underscoring the necessity for automated solutions. The synergistic combination of PET's exceptional sensitivity for detecting metabolic activity with CT's anatomical precision renders accurate segmentation crucial for achieving quantitative and reproducible clinical workflows. However, current methodologies frequently grapple with challenges such as over-segmentation or under-segmentation, inadvertently delineating normal tissues with elevated uptake or neglecting lesions characterized by subtle intensity variations, primarily due to a lack of integrated metabolic and anatomical insights. To address these limitations, we present a novel framework that adeptly integrates clinical expertise regarding anatomical and metabolic cues to refine PET/CT lesion segmentation. Our innovative mixture-of-experts (MoE) based interpretable fusion module skillfully merges complementary modality information while explicitly elucidating the pixel-level contributions of each modality to the final segmentation outcome. Rigorous evaluations across three in-domain benchmarks and two external datasets demonstrate our model's superior segmentation performance and generalizability. Furthermore, our visualizations provide compelling insights into the pivotal role each modality plays in the decision-making process, highlighting our approach's transformative potential in enhancing PET/CT lesion segmentation. Building on this foundation, we further validated the prognostic significance of the features extracted from our proposed framework in the context of PET/CT-based prognosis predictions.
Jiajin Zhang, Liheng Qiu, Wei Liu 0127, Dakai Jin, Wenpei Jiao, Le Lu 0001, Tzu-Chen Yen, Shenmiao Yang, Ke Yan 0006
IEEE Trans. Medical Imaging2
2025 Anatomy-Aware Low-Dose CT Denoising via Pretrained Vision Models and Semantic-Guided Contrastive Learning
Zeli Chen, Zhiyun Song, Wei Fang 0005, Jiajin Zhang, Danyang Tu, Yuxing Tang, Minfeng Xu, Xianghua Ye, Le Lu 0001, Dakai Jin
MICCAI (2)5
2025 Lymphoma Prognosis with Lesion-Anatomy Context Fusion and Attention-Based Multi-lesion Aggregation
Jiajin Zhang, Liheng Qiu, Wei Liu 0127, Dakai Jin, Le Lu 0001, Shenmiao Yang, Ke Yan 0006
MICCAI (1)2
2025 DINO-Reg: Efficient Multimodal Image Registration With Distilled Features
abstract
Medical image registration is a crucial process for aligning anatomical structures, enabling applications such as atlas mapping, longitudinal analysis, and multimodal data fusion. This paper introduces DINO-Reg, an adaptation-free registration method leveraging the vision foundation model, DINOv2, to extract features for deformable 3D medical image alignment. Although DINOv2 was originally trained on natural images, our study links the vision foundation model with medical image registration and demonstrates that the generic image encoder could readily generalize to medical images with state-of-the-art performance. We further propose DINO-Reg-Eco, a knowledge-distilled version using a UNet-structured 3D convolutional neural network (CNN) for feature extraction. The Eco model reduces encoding time by 99% while maintaining state-of-the-art performance, which is essential for resource-limited settings and significantly lowers the carbon footprint associated with intensive computational demands. Benchmarking across diverse datasets shows that both methods outperform existing supervised and unsupervised approaches without fine-tuning, demonstrating the transformative potential of foundation models in medical image registration. Our code is open-sourced at https://github.com/RPIDIAL/DINO-Reg.
Xinrui Song, Xuanang Xu, Jiajin Zhang, Diego Machado Reyes, Pingkun Yan
IEEE Trans. Medical Imaging3
2025 Chest X-Ray Foundation Model With Global and Local Representations Integration
abstract
Chest X-ray (CXR) is the most frequently ordered imaging test, supporting diverse clinical tasks from thoracic disease detection to postoperative monitoring. However, task-specific classification models are limited in scope, require costly labeled data, and lack generalizability to out-of-distribution datasets. To address these challenges, we introduce CheXFound, a self-supervised vision foundation model that learns robust CXR representations and generalizes effectively across a wide range of downstream tasks. We pretrained CheXFound on a curated CXR-987K dataset, comprising over approximately 987K unique CXRs from 12 publicly available sources. We propose a Global and Local Representations Integration (GLoRI) head for downstream adaptations, by incorporating fine- and coarse-grained disease-specific local features with global image features for enhanced performance in multilabel classification. Our experimental results showed that CheXFound outperformed state-of-the-art models in classifying 40 disease findings across different prevalence levels on the CXR-LT 24 dataset and exhibited superior label efficiency on downstream tasks with limited training data. Additionally, CheXFound achieved significant improvements on downstream tasks with out-of-distribution datasets, including opportunistic cardiovascular disease risk estimation, mortality prediction, malpositioned tube detection, and anatomical structure segmentation. The above results demonstrate CheXFound's strong generalization capabilities, which will enable diverse downstream adaptations with improved label efficiency in future applications. The project source code is publicly available at https://github.com/RPIDIAL/CheXFound.
Zefan Yang, Xuanang Xu, Jiajin Zhang, Ge Wang 0001, Mannudeep K. Kalra, Pingkun Yan
IEEE Trans. Medical Imaging3
2025 Disease-Informed Adaptation of Vision-Language Models
abstract
Expertise scarcity and high cost of data annotation hinder the development of artificial intelligence (AI) foundation models for medical image analysis. Transfer learning provides a way to utilize the off-the-shelf foundation models to address the clinical challenges. However, such models encounter difficulties when adapting to new diseases not presented in their original pre-training datasets. Compounding this challenge is the limited availability of example cases for a new disease, which further leads to the poor performance of the existing transfer learning techniques. This paper proposes a novel method for transfer learning of foundation Vision-Language Models (VLMs) to efficiently adapt them to a new disease with only a few examples. Such an effective adaptation of VLMs hinges on learning the nuanced representation of new disease concepts. By capitalizing on the joint visual-linguistic capabilities of VLMs, we introduce disease-informed contextual prompting in a novel disease prototype learning framework, which enables VLMs to quickly grasp the concept of the new disease, even with limited data. Extensive experiments across multiple pre-trained medical VLMs and multiple tasks showcase the notable enhancements in performance compared to other existing adaptation techniques. The code will be made publicly available at https://github.com/RPIDIAL/Disease-informed-VLM-Adaptation.
Jiajin Zhang, Ge Wang 0001, Mannudeep K. Kalra, Pingkun Yan
IEEE Trans. Medical Imaging1
2024 Cardiovascular Disease Detection from Multi-view Chest X-Rays with BI-Mamba
Zefan Yang, Jiajin Zhang, Ge Wang 0001, Mannudeep K. Kalra, Pingkun Yan
MICCAI (5)2
2024 Disease-Informed Adaptation of Vision-Language Models
Jiajin Zhang, Ge Wang 0001, Mannudeep K. Kalra, Pingkun Yan
MICCAI (11)1
2023 When Neural Networks Fail to Generalize? A Model Sensitivity Perspective
abstract
Domain generalization (DG) aims to train a model to perform well in unseen domains under different distributions. This paper considers a more realistic yet more challenging scenario, namely Single Domain Generalization (Single-DG), where only a single source domain is available for training. To tackle this challenge, we first try to understand when neural networks fail to generalize? We empirically ascertain a property of a model that correlates strongly with its generalization that we coin as "model sensitivity". Based on our analysis, we propose a novel strategy of Spectral Adversarial Data Augmentation (SADA) to generate augmented images targeted at the highly sensitive frequencies. Models trained with these hard-to-learn samples can effectively suppress the sensitivity in the frequency space, which leads to improved generalization performance. Extensive experiments on multiple public datasets demonstrate the superiority of our approach, which surpasses the state-of-the-art single-DG methods by up to 2.55%. The source code is available at https://github.com/DIAL-RPI/Spectral-Adversarial-Data-Augmentation.
Jiajin Zhang, Hanqing Chao, Amit Dhurandhar, Ali Tajer, Pingkun Yan
AAAI1
2023 Dynamic Ensemble of Low-Fidelity Experts: Mitigating NAS "Cold-Start"
abstract
Predictor-based Neural Architecture Search (NAS) employs an architecture performance predictor to improve the sample efficiency. However, predictor-based NAS suffers from the severe ``cold-start'' problem, since a large amount of architecture-performance data is required to get a working predictor. In this paper, we focus on exploiting information in cheaper-to-obtain performance estimations (i.e., low-fidelity information) to mitigate the large data requirements of predictor training. Despite the intuitiveness of this idea, we observe that using inappropriate low-fidelity information even damages the prediction ability and different search spaces have different preferences for low-fidelity information types. To solve the problem and better fuse beneficial information provided by different types of low-fidelity information, we propose a novel dynamic ensemble predictor framework that comprises two steps. In the first step, we train different sub-predictors on different types of available low-fidelity information to extract beneficial knowledge as low-fidelity experts. In the second step, we learn a gating network to dynamically output a set of weighting coefficients conditioned on each input neural architecture, which will be used to combine the predictions of different low-fidelity experts in a weighted sum. The overall predictor is optimized on a small set of actual architecture-performance data to fuse the knowledge from different low-fidelity experts to make the final prediction. We conduct extensive experiments across five search spaces with different architecture encoders under various experimental settings. For example, our methods can improve the Kendall's Tau correlation coefficient between actual performance and predicted scores from 0.2549 to 0.7064 with only 25 actual architecture-performance data on NDS-ResNet. Our method can easily be incorporated into existing predictor-based NAS frameworks to discover better architectures. Our method will be implemented in Mindspore (Huawei 2020), and the example code is published at https://github.com/A-LinCui/DELE.
Junbo Zhao 0007, Xuefei Ning, Enshu Liu, Binxin Ru, Tianchen Zhao, Chen Chen 0077, Jiajin Zhang, Qingmin Liao, Yu Wang 0002
AAAI8
2023 Spectral Adversarial MixUp for Few-Shot Unsupervised Domain Adaptation
Jiajin Zhang, Hanqing Chao, Amit Dhurandhar, Ali Tajer, Pingkun Yan
MICCAI (1)1
2023 Toward Adversarial Robustness in Unlabeled Target Domains
abstract
In the past several years, various adversarial training (AT) approaches have been invented to robustify deep learning model against adversarial attacks. However, mainstream AT methods assume the training and testing data are drawn from the same distribution and the training data are annotated. When the two assumptions are violated, existing AT methods fail because either they cannot pass knowledge learnt from a source domain to an unlabeled target domain or they are confused by the adversarial samples in that unlabeled space. In this paper, we first point out this new and challenging problem- adversarial training in unlabeled target domain. We then propose a novel framework named Unsupervised Cross-domain Adversarial Training (UCAT) to address this problem. UCAT effectively leverages the knowledge of the labeled source domain to prevent the adversarial samples from misleading the training process, under the guidance of automatically selected high quality pseudo labels of the unannotated target domain data together with the discriminative and robust anchor representations of the source domain data. The experiments on four public benchmarks show that models trained with UCAT can achieve both high accuracy and strong robustness. The effectiveness of the proposed components is demonstrated through a large set of ablation studies. The source code is publicly available at https://github.com/DIAL-RPI/UCAT.
Jiajin Zhang, Hanqing Chao, Pingkun Yan
IEEE Trans. Image Process.1
2022 Regression Metric Loss: Learning a Semantic Representation Space for Medical Images
Hanqing Chao, Jiajin Zhang, Pingkun Yan
MICCAI (8)2
2022 Overlooked Trustworthiness of Saliency Maps
Jiajin Zhang, Hanqing Chao, Giridhar Dasegowda, Ge Wang 0001, Mannudeep K. Kalra, Pingkun Yan
MICCAI (3)1
2021 Task-Oriented Low-Dose CT Image Denoising
Jiajin Zhang, Hanqing Chao, Xuanang Xu, Chuang Niu, Ge Wang 0001, Pingkun Yan
MICCAI (6)1
2021 Integrative analysis for COVID-19 patient outcome prediction
Hanqing Chao, Xi Fang 0002, Jiajin Zhang, Fatemeh Homayounieh, Chiara Daniela Arru, Subba R. Digumarthy, Rosa Babaei, Hadi Karimi Mobin, Iman Mohseni, Luca Saba, Alessandro Carriero, Zeno Falaschi, Alessio Pasche, Ge Wang 0001, Mannudeep K. Kalra, Pingkun Yan
Medical Image Anal.3
2020 A sensor attack detection method based on fusion interval and historical measurement in CPS
abstract
Cyber-physical systems (CPS) is a next-generation intelligent system that realizes close integration of computing, communication and physical elements based on environmental perception. The interaction between information technology and the physical world makes CPS vulnerable to various malicious attacks and damage Its security. The paper designs a sensor attack detection method based on fusion interval and historical measurement. The method first builds different fault models for different sensors, and uses system dynamics equations to integrate historical measurements into the attack detection method. Different aspects of sensor measurement are analyzed. In addition, the use of historical measurement and fusion interval solves the problem of whether there is a failure when the measurements of two sensors intersect. The core idea of this method is to use the paired inconsistent relationship between sensors to detect and identify attacks .
Xiaobo Cai, Jiajin Zhang
IPCCC5
2020 Research on Security Estimation and Control of Cyber-Physical System
abstract
Cyber-Physical Systems (CPS) is a multidimensional complex system that integrates computing, network, and physical environments. Due to the existence of network communications and embedded computers, CPS is vulnerable to attacks, so its security has become an important issue. The paper studies the main types of network attacks and the methods of detection, including the security estimation and control of CPS. Finally, the paper puts forward the key issues and challenges faced by the CPS network attack research.
Xiaobo Cai, Huihui Wang 0001, Jiajin Zhang
IPCCC5
2017 Extremely Fast Decision Tree Mining for Evolving Data Streams
abstract
Nowadays real-time industrial applications are generating a huge amount of data continuously every day. To process these large data streams, we need fast and efficient methodologies and systems. A useful feature desired for data scientists and analysts is to have easy to visualize and understand machine learning models. Decision trees are preferred in many real-time applications for this reason, and also, because combined in an ensemble, they are one of the most powerful methods in machine learning.
Albert Bifet, Jiajin Zhang, Wei Fan 0001, Jianfeng Qian, Geoff Holmes 0001, Bernhard Pfahringer
KDD2
2014 Design and implementation of gaze tracking system with iPad
abstract
Gaze tracking is the process of measuring gaze point or the motion of an eye relative to the head. Gaze tracking technique provides us a brand new way of human computer interaction. In addition, eye gaze tracking can be also applied to support seriously disabled people in using computer. In this paper, we explored the use of eye gaze tracking technology on a tablet device, designed and implemented an eye tracking system on an iPad device. In our system, gaze estimation is based on analyzing the appearances of eyes which are retrieved by the built-in camera on the iPad. Artificial neural networks are employed to estimate the location of the user gaze from the eye region image. The results indicate that it is possible to obtain an accuracy of 82.5% with the proposed system.
Jiajin Zhang, Liu Di, Lichang Chen
RTCSA1