VLDB 2026 Research / reviewers in the wild / expert
Harrison X. Bai
dblp:165/6544
· DBLP profile ↗
23ranked-venue papers
0as first author
23since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond the LUMIR challenge: The pathway to foundational registration models
Junyu Chen 0002, Shuwen Wei, Joel Honkamaa, Pekka Marttinen, Hang Zhang 0010, Min Liu 0008, Yichao Zhou 0002, Zuopeng Tan, Yi Wang 0028, Hongchao Zhou, Shunbo Hu, Yi Zhang 0120, Lukas Förner, Thomas Wendler 0001, Bailiang Jian, Benedikt Wiestler, Tim Hable, Dan Ruan, Frederic Madesta, Thilo Sentker, Wiebke Heyer, Lianrui Zuo, Yuwei Dai, Jerry L. Prince, Harrison X. Bai, Yong Du 0002, Yihao Liu 0003, Alessa Hering, Reuben Dorent, Lasse Hansen, Mattias P. Heinrich, Aaron Carass |
Medical Image Anal. | 29 |
| 2026 | Unsupervised learning of spatially varying regularization for diffeomorphic image registration
Junyu Chen 0002, Shuwen Wei, Yihao Liu 0003, Zhangxing Bian, Yufan He, Aaron Carass, Harrison X. Bai, Yong Du 0002 |
Medical Image Anal. | 7 |
| 2026 | Abn-BLIP: Abnormality-aligned Bootstrapping Language-Image Pre-training for pulmonary embolism diagnosis and report generation from CTPAabstractMedical imaging plays a pivotal role in modern healthcare, with computed tomography pulmonary angiography (CTPA) being a critical tool for diagnosing pulmonary embolism and other thoracic conditions. However, the complexity of interpreting CTPA scans and generating accurate radiology reports remains a significant challenge. This paper introduces Abn-BLIP (Abnormality-aligned Bootstrapping Language-Image Pretraining), an advanced diagnosis model designed to align abnormal findings to generate the accuracy and comprehensiveness of radiology reports. By leveraging learnable queries and cross-modal attention mechanisms, our model demonstrates superior performance in detecting abnormalities, reducing missed findings, and generating structured reports compared to existing methods. Our experiments show that Abn-BLIP outperforms state-of-the-art medical vision-language models and 3D report generation methods in both accuracy and clinical relevance. These results highlight the potential of integrating multimodal learning strategies for improving radiology reporting. The source code is available at https://github.com/zzs95/abn-blip. Zhusi Zhong, Yuli Wang, Lulu Bi, Zhuoqi Ma, Sun Ho Ahn, Christopher J. Mullin, Colin Greineder, Michael Atalay, Scott Collins, Grayson Baird, Cheng Ting Lin, J. Webster Stayman, Todd M. Kolb, Ihab Kamel, Harrison X. Bai, Zhicheng Jiao |
Medical Image Anal. | 15 |
| 2026 | PSC-UDA: Point-cloud Structure Constrained Unsupervised Domain Adaptation for contour-based kidney segmentationabstractCross-domain medical image segmentation has gained increasing interest for its potential to reduce annotation efforts and improve clinical generalization capabilities. Domain adaptation aims to tackle the domain shift that appears in different image modalities. In cross-domain segmentation, generative models often suffer from limited accuracy due to their lack of domain-specific representations. Besides, many transfer learning approaches rely on additional manual annotations for supervision, emerging paradigms such as Unsupervised Domain Adaptation (UDA) facilitate effective knowledge transfer even when labels in the target domain are entirely absent. In this study, we propose a novel Point-cloud Structure Constrained Unsupervised Domain Adaptation (PSC-UDA) framework based on a Contour-Aware Segmentation (CAS) model with a 3D contour point cloud to bridge the domain gaps appearing in cross-site and cross-domain medical images. The CAS model distills the domain-invariant kidney structure from image texture to distinguish the point cloud and characterize the kidney contour in a coarse-to-fine way. With point-to-voxel self-learning on 3D structure constraints, the proposed PSC-UDA framework addresses visual domain shift, adapting discriminative information of the kidney from the labeled source domain (CT) to the unlabeled target domain (CT/MRI), so that it realizes precise cross-domain kidney segmentation with limited labels. Experimental results prove that the proposed method outperforms the generative UDA methods and the source-free methods on three cross-domain kidney segmentation datasets, outperforming even without a target domain adaptation strategy. The source code is available at https://github.com/zzs95/PSC-UDA . Yang Li 0111, Zhusi Zhong, Jie Li 0001, Helen Zhang, Mihir Khunte, Lulu Bi, Scott Collins, Harrison X. Bai, Michael Atalay, Ihab Kamel, Xinbo Gao 0001, Zhicheng Jiao |
Pattern Recognit. | 8 |
| 2025 | RFID-Based Vital Sign Monitoring Under Motion Using Physics-Informed Generative ModelsabstractWireless signals are widely used for human sensing, but they require devices and targets to remain stationary, especially for fine-grained motions like respiration. To enable vital sign monitoring under motion using RFID, we employ dual tags to create a relative coordinate system that reduces motion interference. We also propose physics-informed generative models with frequency domain constraints to improve noise reduction, capturing both time and frequency features. Our method, tested across dynamic scenarios including walking, treadmill exercises, and driving and validated using real patient data, demonstrates superior performance compared to traditional approaches in accurately matching real respiratory signals and exhibits robustness against time shifts. Tianya Zhao, Yuwei Dai, Harrison X. Bai, Karthik Suresh 0006, Zhicheng Jiao, Shiwen Mao, Xuyu Wang |
MASS | 5 |
| 2025 | Inspired by pathogenic mechanisms: A novel gradual multi-modal fusion framework for mild cognitive impairment diagnosis
Hong-Dong Li, Hanhe Lin, Chao Li 0031, Harrison X. Bai, Wei Lan 0001, Jin Liu 0012 |
Neural Networks | 6 |
| 2025 | Multi-Modality Regional Alignment Network for Covid X-Ray Survival Prediction and Report GenerationabstractIn response to the worldwide COVID-19 pandemic, advanced automated technologies have emerged as valuable tools to aid healthcare professionals in managing an increased workload by improving radiology report generation and prognostic analysis. This study proposes a Multi-modality Regional Alignment Network (MRANet), an explainable model for radiology report generation and survival prediction that focuses on high-risk regions. By learning spatial correlation in the detector, MRANet visually grounds region-specific descriptions, providing robust anatomical regions with a completion strategy. The visual features of each region are embedded using a novel survival attention mechanism, offering spatially and risk-aware features for sentence encoding while maintaining global coherence across tasks. A cross-domain LLMs-Alignment is employed to enhance the image-to-text transfer process, resulting in sentences rich with clinical detail and improved explainability for radiologists. Multi-center experiments validate the overall performance and each module's composition within the model, encouraging further advancements in radiology report generation research emphasizing clinical interpretation and trustworthiness in AI models applied to medical studies. Zhusi Zhong, Jie Li 0001, John Sollee, Scott Collins, Harrison X. Bai, Terrance Healey, Michael Atalay, Xinbo Gao 0001, Zhicheng Jiao |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | ECG-grained Cardiac Monitoring Using RFIDabstractHeartbeat signals are useful to disease prediction, sub-health diagnosis, fatigue warning, and even emotion estimation. There is a compelling need for contactless, easy-to-deploy, and long-term heartbeat monitoring. This paper presents a contactless Radio Frequency Identification (RFID) based system for heartbeat monitoring that leverages the insight that RFID signal fluctuations induced by chest motion are synchronous with both respiration and heartbeat. The proposed system collects the temporal phase information from the tag pair on the body to extract heartbeat signals using a sequence of signal processing techniques. We propose a signal separation method based on empirical mode decomposition (EMD) to obtain heart rate after preprocessing. Furthermore, the estimated signal is input to an enhanced variational autoencoder (VAE) model to recover the heartbeat waveform. Implemented with commercial off-the-shelf (COTS) RFID devices, the system achieves accurate heart rate monitoring with less than 3% relative errors. The detected waveform exhibits a median cosine similarity of 0.83 as compared with the ground truth, which validate the system’s wide applicability and high reliability for fine-grained, contactless heartbeat monitoring. Tianya Zhao, Shiwen Mao, Harrison X. Bai, Zhicheng Jiao, Xuyu Wang |
ICCCN | 4 |
| 2024 | Structural Entities Extraction and Patient Indications Incorporation for Chest X-Ray Report Generation
Kang Liu 0025, Zhuoqi Ma, Xiaolu Kang, Zhusi Zhong, Zhicheng Jiao, Grayson Baird, Harrison X. Bai, Qiguang Miao |
MICCAI (3) | 7 |
| 2024 | Car-Dcros: A Dataset and Benchmark for Enhancing Cardiovascular Artery Segmentation Through Disconnected Components Repair and Open Curve Snake
Yuli Wang, Wen-Chi Hsu, Victoria Shi, Gigin Lin, Cheng Ting Lin, Harrison X. Bai |
MICCAI (1) | 7 |
| 2024 | Enhancing vision-language models for medical imaging: bridging the 3D gap with innovative slice selectionabstractRecent approaches to vision-language tasks are built on the remarkable capabilities of large vision-language models (VLMs). These models excel in zero-shot and few-shot learning, enabling them to learn new tasks without parameter updates. However, their primary challenge lies in their design, which primarily accommodates 2D input, thus limiting their effectiveness for medical images, particularly radiological images like MRI and CT, which are typically 3D. To bridge the gap between state-of-the-art 2D VLMs and 3D medical image data, we developed an innovative, one-pass, unsupervised representative slice selection method called Vote-MI, which selects representative 2D slices from 3D medical imaging. To evaluate the effectiveness of vote-MI when implemented with VLMs, we introduce BrainMD, a robust, multimodal dataset comprising 2,453 annotated 3D MRI brain scans with corresponding textual radiology reports and electronic health records. Based on BrainMD, we further develop two benchmarks, BrainMD-select (including the most representative 2D slice of 3D image) and BrainBench (including various vision-language downstream tasks). Extensive experiments on the BrainMD dataset and its two corresponding benchmarks demonstrate that our representative selection method significantly improves performance in zero-shot and few-shot learning tasks. On average, Vote-MI achieves a 14.6\% and 16.6\% absolute gain for zero-shot and few-shot learning, respectively, compared to randomly selecting examples. Our studies represent a significant step toward integrating AI in medical imaging to enhance patient care and facilitate medical research. We hope this work will serve as a foundation for data selection as vision-language models are increasingly applied to new tasks. Yuli Wang, Peng jian, Yuwei Dai, Craig K. Jones, Haris I. Sair, Jinglai Shen, Nicolas Loizou, Wen-Chi Hsu, Maliha R. Imami, Zhicheng Jiao, Harrison X. Bai |
NeurIPS | 13 |
| 2024 | Evidential Uncertainty Quantification: A Variance-Based PerspectiveabstractUncertainty quantification of deep neural networks has become an active field of research and plays a crucial role in various downstream tasks such as active learning. Recent advances in evidential deep learning shed light on the direct quantification of aleatoric and epistemic uncertainties with a single forward pass of the model. Most traditional approaches adopt an entropy-based method to derive evidential uncertainty in classification, quantifying uncertainty at the sample level. However, the variance-based method that has been widely applied in regression problems is seldom used in the classification setting. In this work, we adapt the variance-based approach from regression to classification, quantifying classification uncertainty at the class level. The variance decomposition technique in regression is extended to class covariance decomposition in classification based on the law of total covariance, and the class correlation is also derived from the covariance. Experiments on cross-domain datasets are conducted to illustrate that the variance-based approach not only results in similar accuracy as the entropy-based one in active domain adaptation but also brings information about class-wise uncertainties as well as between-class correlations. The code is available at https://github.com/KerryDRX/EvidentialADA. This alternative means of evidential uncertainty quantification will give researchers more options when class uncertainties and correlations are important in their applications. Ruxiao Duan, Brian Caffo, Harrison X. Bai, Haris I. Sair, Craig K. Jones |
WACV | 3 |
| 2024 | MMGK: Multimodality Multiview Graph Representations and Knowledge Embedding for Mild Cognitive Impairment DiagnosisabstractThe diagnosis of mild cognitive impairment (MCI), which is an early stage of Alzheimer’s disease (AD), has great clinical significance. Medical imaging and gene sequencing technologies have provided sufficient multimodality data for MCI diagnostic studies. However, how to effectively extract the rich representations from multimodality data remains a challenging task. To address this challenging task, we propose a new multimodality multiview graph representations and knowledge embedding (MMGK) framework to diagnose MCI. First, to obtain rich information from multimodality data, we extract multiview feature representations from magnetic resonance imaging (MRI) and genetic data. Afterward, considering the correlations between subjects, all subjects are constructed into a graph based on the different single-view feature representations, respectively. To further enrich the correlations between subjects, demographic data are utilized through knowledge embedding. Finally, to perform MCI diagnosis on multiview graphs, graph convolutional networks (GCNs) are utilized. In addition, to further improve the performance of MCI diagnosis, a two-step ensemble learning method is proposed. The proposed framework is evaluated on 188 subjects from the AD Neuroimaging Initiative (ADNI). Experimental results show that our proposed framework achieves good performance with accuracy reaching 0.888, and outperforms some state-of-the-art (SOTA) methods. In addition, the proposed framework is applied to Parkinson’s disease (PD) diagnosis and achieves 0.856 accuracy. Overall, our proposed method has potential for clinical application in MCI diagnosis and other diseases via integrating MRI, genetic data, and demographic data. Our code is available at:https://github.com/miacsu/MMGK. Jin Liu 0012, Rui Guo 0009, Harrison X. Bai, Hulin Kuang, Jianxin Wang 0001 |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2024 | De-Biased Disentanglement Learning for Pulmonary Embolism Survival Prediction on Multimodal DataabstractHealth disparities among marginalized populations with lower socioeconomic status significantly impact the fairness and effectiveness of healthcare delivery. The increasing integration of artificial intelligence (AI) into healthcare presents an opportunity to address these inequalities, provided that AI models are free from bias. This paper aims to address the bias challenges by population disparities within healthcare systems, existing in the presentation of and development of algorithms, leading to inequitable medical implementation for conditions such as pulmonary embolism (PE) prognosis. In this study, we explore the diverse bias in healthcare systems, which highlights the demand for a holistic framework to reducing bias by complementary aggregation. By leveraging de-biasing deep survival prediction models, we propose a framework that disentangles identifiable information from images, text reports, and clinical variables to mitigate potential biases within multimodal datasets. Our study offers several advantages over traditional clinical-based survival prediction methods, including richer survival-related characteristics and bias-complementary predicted results. By improving the robustness of survival analysis through this framework, we aim to benefit patients, clinicians, and researchers by enhancing fairness and accuracy in healthcare AI systems. Zhusi Zhong, Jie Li 0001, Helen Zhang, Fayez H. Fayad, Yang Li 0111, Scott Collins, Harrison X. Bai, Sun Ho Ahn, Michael Atalay, Xinbo Gao 0001, Zhicheng Jiao |
IEEE J. Biomed. Health Informatics | 8 |
| 2023 | Improving Outcome Prediction of Pulmonary Embolism by De-biased Multi-modality Model
Zhusi Zhong, Jie Li 0001, Yang Li 0111, Fayez H. Fayad, Helen Zhang, Sun Ho Ahn, Harrison X. Bai, Xinbo Gao 0001, Michael Atalay, Zhicheng Jiao |
MICCAI (5) | 8 |
| 2023 | AC-E Network: Attentive Context-Enhanced Network for Liver SegmentationabstractSegmentation of liver from CT scans is essential in computer-aided liver disease diagnosis and treatment. However, the 2DCNN ignores the 3D context, and the 3DCNN suffers from numerous learnable parameters and high computational cost. In order to overcome this limitation, we propose an Attentive Context-Enhanced Network (AC-E Network) consisting of 1) an attentive context encoding module (ACEM) that can be integrated into the 2D backbone to extract 3D context without a sharp increase in the number of learnable parameters; 2) a dual segmentation branch including complemental loss making the network attend to both the liver region and boundary so that getting the segmented liver surface with high accuracy. Extensive experiments on the LiTS and the 3D-IRCADb datasets demonstrate that our method outperforms existing approaches and is competitive to the state-of-the-art 2D-3D hybrid method on the equilibrium of the segmentation precision and the number of model parameters. Yang Li 0111, Beiji Zou 0001, Peishan Dai, Miao Liao, Harrison X. Bai, Zhicheng Jiao |
IEEE J. Biomed. Health Informatics | 5 |
| 2022 | A dynamic multi-modal fusion network for ovarian tumor differentiationabstractAccurate ovarian tumor differentiation is a challenging task where the benign and malignant tumors share similar T1C and T2WI MRI appearances. Therefore, it is necessary to leverage additional multi-modal data, e.g., the age, CA125level, and other clinical information, which are helpful but rarely exploited. In this paper, we propose a dynamic fusion network that can adaptively make full use of multi-modal data, including MRI and clinical information, to realize precise ovarian tumor differentiation. Specifically, we design a dynamic nonlinear module (D-Non-L module) on the top of the image representation. The D-Non-L module is formulated as an iterative nonlinear projection parameterized by the learned features of the patient-wise clinical information. With the help of this module, the interaction between clinical features and image features could be achieved to adaptively improve the discrimination of visual representations. Moreover, we construct a dual-path-based architecture to fully exploit the complementary information from T1C and T2WI MRIs. Extensive experimental results on the locally organized ovarian tumor dataset demonstrate that our methods are superior to the single-modal and single-path-based methods. And the proposed dynamic non-linear module obtains the best performance compared with other multi-modal fusion strategies. Yang Li 0111, Beiji Zou 0001, Yulan Dai, Harrison X. Bai, Zhicheng Jiao |
BIBM | 5 |
| 2022 | Parameter-Free Latent Space Transformer for Zero-Shot Bidirectional Cross-modality Liver Segmentation
Yang Li 0111, Beiji Zou 0001, Yulan Dai, Chengzhang Zhu, Fan Yang 0054, Xin Li 0079, Harrison X. Bai, Zhicheng Jiao |
MICCAI (4) | 7 |
| 2022 | Dynamic prototypical feature representation learning framework for semi-supervised skin lesion segmentation
Chunna Tian, Xinbo Gao 0001, Xue Feng 0001, Harrison X. Bai, Zhicheng Jiao |
Neurocomputing | 6 |
| 2022 | DARC: Deep adaptive regularized clustering for histopathological image classification
Junjian Li, Jin Liu 0012, Hailin Yue, Jianhong Cheng, Hulin Kuang, Harrison X. Bai, Yu-Ping Wang 0002, Jianxin Wang 0001 |
Medical Image Anal. | 6 |
| 2022 | MLDRL: Multi-loss disentangled representation learning for predicting esophageal cancer response to neoadjuvant chemoradiotherapy using longitudinal CT images
Hailin Yue, Jin Liu 0012, Junjian Li, Hulin Kuang, Jinyi Lang, Jianhong Cheng, Yongtao Han, Harrison X. Bai, Yu-Ping Wang 0002, Jianxin Wang 0001 |
Medical Image Anal. | 9 |
| 2022 | Discriminative error prediction network for semi-supervised colon gland segmentation
Chunna Tian, Harrison X. Bai, Zhicheng Jiao, Xilan Tian |
Medical Image Anal. | 3 |
| 2022 | Prediction of Glioma Grade Using Intratumoral and Peritumoral Radiomic Features From Multiparametric MRI ImagesabstractThe accurate prediction of glioma grade before surgery is essential for treatment planning and prognosis. Since the gold standard (i.e., biopsy)for grading gliomas is both highly invasive and expensive, and there is a need for a noninvasive and accurate method. In this study, we proposed a novel radiomics-based pipeline by incorporating the intratumoral and peritumoral features extracted from preoperative mpMRI scans to accurately and noninvasively predict glioma grade. To address the unclear peritumoral boundary, we designed an algorithm to capture the peritumoral region with a specified radius. The mpMRI scans of 285 patients derived from a multi-institutional study were adopted. A total of 2153 radiomic features were calculated separately from intratumoral volumes (ITVs)and peritumoral volumes (PTVs)on mpMRI scans, and then refined using LASSO and mRMR feature ranking methods. The top-ranking radiomic features were entered into the classifiers to build radiomic signatures for predicting glioma grade. The prediction performance was evaluated with five-fold cross-validation on a patient-level split. The radiomic signatures utilizing the features of ITV and PTV both show a high accuracy in predicting glioma grade, with AUCs reaching 0.968. By incorporating the features of ITV and PTV, the AUC of IPTV radiomic signature can be increased to 0.975, which outperforms the state-of-the-art methods. Additionally, our proposed method was further demonstrated to have strong generalization performance in an external validation dataset with 65 patients. The source code of our implementation is made publicly available at https://github.com/chengjianhong/glioma_grading.git. Jianhong Cheng, Jin Liu 0012, Hailin Yue, Harrison X. Bai, Yi Pan 0001, Jianxin Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |