VLDB 2026 Research / reviewers in the wild / expert
Jing Cai 0001
dblp:53/252-1
· DBLP profile ↗
42ranked-venue papers
0as first author
41since 2021 · last 2026
0000-0001-6934-0108ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 19 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adapting vision-Language foundation model for next generation medical ultrasound image analysisabstractVision-Language Foundation Models (VLFMs) exhibit remarkable generalization, yet their direct application to medical ultrasound is severely hindered by a profound modality gap. The unique acoustic physics of ultrasound, characterized by speckle noise, shadowing, and heterogeneous textures, often degrades the performance of off-the-shelf VLFMs. To bridge this gap, we propose a novel Hybrid Tuning (HT) strategy for the parameter-efficient adaptation of CLIP-based models to ultrasound analysis. Instead of updating the pre-trained weights, HT freezes the visual backbone and integrates a specialized lightweight adapter. This adapter features a Frequency Filtering module to suppress domain-specific periodic artifacts and a Noise Estimation module to dynamically calibrate feature representations. Extensive evaluations across six multi-center datasets demonstrate that our HT-enhanced models significantly outperform existing state-of-the-art adapters and medical VLFMs in both segmentation and classification tasks. Notably, HT exhibits exceptional data efficiency in few-shot scenarios and robust cross-dataset generalization. Our findings prove that preserving pre-trained semantic priors while explicitly modeling ultrasound-specific noise is key to unlocking foundational intelligence in automated ultrasound diagnosis. The source code will be made publicly available. Jingguo Qu, Jia Ai, Tonghuan Xiao, Sheng Ning, Harry Qin, Ann Dorothy King, Winnie Chiu-Wing Chu, Jing Cai 0001, Michael T. C. Ying |
Expert Syst. Appl. | 12 |
| 2026 | Fully & partially-transmitted-rule fusion: A novel hierarchical fuzzy classification with application to nasopharyngeal cancer's metastasis prediction
Ta Zhou, Yuanqing Yang, Wei Yan 0030, Weiqin Liu, Xibei Yang, Weiping Ding 0001, Jing Cai 0001, Shitong Wang 0001 |
Fuzzy Sets Syst. | 7 |
| 2026 | Linguistically interpretable hierarchical fuzzy classifier with dynamic adjustable generalizability gradient
Ta Zhou, Jinghao Chen, Shuihua Wang, Weiqin Liu, Xibei Yang, Jing Cai 0001, Shitong Wang 0001 |
Fuzzy Sets Syst. | 9 |
| 2026 | Designing a hybrid optimization methodology for delineating boundary of ultrasound prostate cancer with an explainable mathematical model
Tao Peng 0013, Dehui Xiang, Binbin Jiang, Baoqing Nie, Derun Li, Caishan Wang, Weifang Zhu, Jing Cai 0001, Enting Gao, Xinjian Chen 0001 |
Neurocomputing | 10 |
| 2026 | Self-supervised reconstruction framework via motion- and physics-informed learning for four-dimensional magnetic resonance fingerprinting
Weihang Liao, Xinzhi Teng, Jiarui Zhu, Junyi Yan, Yat-Lam Wong, Victor Ho-fun Lee, Harry Qin, Tian Li 0012, Jing Cai 0001 |
Medical Image Anal. | 16 |
| 2026 | ZA-Net: A universal zero-annotation nuclei segmentation network for pathology images via vision-language pre-trained model
Fuqiang Chen, Kun Ru, Miaoxia He, Qizhai Li, Yao Pu, Jing Cai 0001, Wenjian Qin |
Pattern Recognit. Lett. | 8 |
| 2026 | DA-TSK-PLR-FS: Domain Adaptive Takagi-Sugeno-Kang Fuzzy System via Pseudolabel Refinement for CCTA-Based Vulnerable Coronary Plaques RecognitionabstractArtificial intelligence has shown great promise in noninvasive recognition of vulnerable coronary plaques. However, practical data issues in multicenter studies, such as inconsistent data distribution and insufficient or missing data labels, could significantly affect the recognition accuracy. Unsupervised domain adaptation (UDA) can be introduced to address this challenge, but several limits still remain. First, many existing UDA models are black boxes, hindering healthcare professionals' ability to interpret and trust the model's decision. Second, some methods use pseudolabel to enhance performance, but often overlook the quality assessment of these pseudolabels, potentially leading to negative knowledge transfer. To this end, based on the interpretable Takagi–Sugeno–Kang fuzzy system (TSK-FS), a novel domain adaptive method is proposed to improve model generalizability for vulnerable coronary plaques recognition in multicenter data. First of all, TSK-FS is employed to construct a shared fuzzy feature space for the source domain and the target domain, aiming to better align data distribution. To make full use of the information of unlabeled target domain data and further reduce the negative knowledge transfer, the enhanced pseudolabel learning mechanism is further introduced by combining the graph-based random walking and label filtering. Moreover, Multicenter data of 910 patients with suspected or diagnosed coronary artery disease were collected from three hospitals for experiments. Experimental results demonstrate that the proposed DA-TSK-PLR-FS achieves the promising generalizability across multicenter datasets Yuanpeng Zhang 0001, Wei Zhang 0221, Zhaoheng Huang, Saikit Lam, Shitong Wang 0001, Jing Cai 0001 |
IEEE Trans. Fuzzy Syst. | 9 |
| 2026 | Stabilization of Hierarchical Supervised Adversarial Mechanism With Dynamic Compensation for Two-View Soft StimulationabstractIn this study, we propose a hierarchical two-view Takagi–Sugeno–Kang (TSK) fuzzy classifier based on a supervised adversarial mechanism (ADVML-FC) designed to address the limitations of traditional single- and two-view fuzzy classifiers, such as data underutilization, challenges in feature fusion, and poor generalization. The proposed model leverages the output of one view as perturbation information to challenge the other view, thereby promoting information exchange and fusion between the two views. Unlike conventional approaches that rely on random noise perturbations, the model uses signals derived from the opposing view outputs and dynamically adjusts the attack success rate (ρ) to prevent overfitting and enhance generalization. The model simplifies the objective function design and integrates the least-squares method to efficiently estimate the consequent parameters of fuzzy rules, thereby substantially reducing computational complexity. Moreover, the model retains a zero-order TSK fuzzy rule structure and integrates fuzzy C-means clustering with Gaussian membership functions to ensure semantic interpretability. Experimental results demonstrate that ADVML-FC achieves superior performance in terms of average training accuracy, average testing accuracy, F1-score, Kappa, and Matthews correlation coefficient across nine two-view datasets. The proposed model not only delivers exceptional classification performance in terms of interpretability, highlighting its strong potential for broad applicability. Ta Zhou, Weiqin Liu, Xibei Yang, Jing Cai 0001 |
IEEE Trans. Ind. Informatics | 6 |
| 2026 | USIGAN: Unbalanced Self-Information Feature Transport for Weakly Paired Image IHC Virtual StainingabstractImmunohistochemical (IHC) virtual staining is a task that generates virtual IHC images from H&E images while maintaining pathological semantic consistency with adjacent slices. This task aims to achieve cross-domain mapping between morphological structures and staining patterns through generative models, providing an efficient and cost-effective solution for pathological analysis. However, under weakly paired conditions, spatial heterogeneity between adjacent slices presents significant challenges. This can lead to inaccurate one-to-many mappings and generate results that are inconsistent with the pathological semantics of adjacent slices. To address this issue, we propose a novel unbalanced self-information feature transport for IHC virtual staining, named USIGAN, which extracts global morphological semantics without relying on positional correspondence. By removing weakly paired terms in the joint marginal distribution, we effectively mitigate the impact of weak pairing on joint distributions, thereby significantly improving the content consistency and pathological semantic consistency of the generated results. Moreover, we design the Unbalanced Optimal Transport Consistency Mining (UOT-CTM) mechanism and the Pathology Self-Correspondence Mining (PC-SCM) mechanism to construct correlation matrices between H&E and generated IHC in image-level and real IHC and generated IHC image sets in intra-group level. Experiments conducted on two publicly available datasets demonstrate that our method achieves superior performance across multiple clinically significant metrics, such as IoD and Pearson-R correlation, demonstrating better clinical relevance. The code is available at: https://github.com/MIXAILAB/USIGAN. Bing Xiong 0004, Fuqiang Chen, Deboch Eyob Abera, Wanming Hu, Jing Cai 0001, Wenjian Qin |
IEEE Trans. Image Process. | 7 |
| 2026 | UTADC-Net: Unsupervised Topological-Aware Diffusion Condensation Network for Medical Image SegmentationabstractMedical image segmentation plays a crucial role in computer-aided diagnosis and treatment planning. Unsupervised segmentation methods that can effectively leverage unlabeled data bring significant promise in clinical application. However, they remain a challenging task in maintaining anatomical structure topological consistency that often produces anatomical structure breaks, connectivity errors, or boundary discontinuities. To address these issues, we propose a novel Unsupervised Topological-Aware Diffusion Condensation Network (UTADC-Net) for medical image segmentation. Specifically, we design a diffusion condensation-based framework that achieves structural consistency in segmentation results by effectively modeling long-range dependencies between pixels and incorporating topological constraints. First, to effectively fuse local details and global semantic information, we employ a pixel-centric patch embedding module by simultaneously modeling local structural features and inter-region interactions. Second, to enhance the topological consistency of segmentation results, we introduce an adaptive topological constraint mechanism that guides the network to learn anatomically aligned structural representations through pixel-level topological relationships and corresponding loss functions. Extensive experiments conducted on three public medical image datasets demonstrate that our proposed UTADC-Net significantly outperforms existing unsupervised methods in terms of segmentation accuracy and topological structure preservation. Notably, our method demonstrates segmentation results with excellent anatomical structural consistency. These results indicate that our framework provides a novel and practical solution for unsupervised medical image segmentation. Ruodai Wu, Bing Xiong 0004, Fuqiang Chen, Yaoqin Xie, Jing Cai 0001, Wenjian Qin |
IEEE J. Biomed. Health Informatics | 7 |
| 2026 | PGVMS: A Prompt-Guided Unified Framework for Virtual Multiplex IHC Staining With Pathological Semantic LearningabstractImmunohistochemical (IHC) staining enables precise molecular profiling of protein expression, with over 200 clinically available antibody-based tests in modern pathology. However, comprehensive IHC analysis is frequently limited by insufficient tissue quantities in small biopsies. Therefore, virtual multiplex staining emerges as an innovative solution to digitally transform H&E images into multiple IHC representations, yet current methods still face three critical challenges: 1) inadequate semantic guidance for multi-staining, 2) inconsistent distribution of immunochemistry staining, and 3) spatial misalignment across different stain modalities. To overcome these limitations, we present a prompt-guided framework for virtual multiplex IHC staining using only uniplex training data (PGVMS). Our framework introduces three key innovations corresponding to each challenge: First, an adaptive prompt guidance mechanism employing a pathological visual language model dynamically adjusts staining prompts to resolve semantic guidance limitations (Challenge 1). Second, our protein-aware learning strategy (PALS) maintains precise protein expression patterns by direct quantification and constraint of protein distributions (Challenge 2). Third, the prototype-consistent learning strategy (PCLS) establishes cross-image semantic interaction to correct spatial misalignments (Challenge 3). Evaluated on two benchmark datasets, PGVMS demonstrates superior performance in pathological consistency. In general, PGVMS represents a paradigm shift from dedicated single-task models toward unified virtual staining systems. Fuqiang Chen, Wanming Hu, Deboch Eyob Abera, Boyun Zheng, Jing Cai 0001, Wenjian Qin |
IEEE Trans. Medical Imaging | 8 |
| 2025 | REACT-KD: Region-Aware Cross-Modal Topological Knowledge Distillation for Interpretable Medical Image ClassificationabstractReliable and interpretable tumor classification from clinical imaging remains a core challenge. The main difficulties arise from heterogeneous modality quality, limited annotations, and the absence of structured anatomical guidance. We present REACT-KD, a Region-Aware Cross-modal Topological Knowledge Distillation framework that transfers supervision from highfidelity multi-modal sources into a lightweight CT-based student model. The framework employs a dual teacher design. One branch captures structure-function relationships through dualtracer PET/CT, while the other models dose-aware features using synthetically degraded low-dose CT. These branches jointly guide the student model through two complementary objectives. The first achieves semantic alignment through logits distillation, and the second models anatomical topology through region graph distillation. A shared CBAM3D module ensures consistent attention across modalities. To improve reliability in deployment, REACTKD introduces modality dropout during training, which enables robust inference under partial or noisy inputs. As a case study, we applied REACT-KD to hepatocellular carcinoma staging. The framework achieved an average AUC of 93.5 % on an internal PET/CT cohort and maintained$\mathbf{7 6. 6 \%}$to$\mathbf{8 1. 5 \%}$AUC across varying levels of dose degradation in external CT testing. Decision curve analysis further shows that REACT-KD consistently provides the highest net clinical benefit across all thresholds, confirming its value in real-world diagnostic practice. Code is available at: https://github.com/Kinetics-JOJO/REACT-KD. Hongzhao Chen, Hexiao Ding, Jing Lan, Ka Chun Li, Gerald W. Y. Cheng, Nga-Chun Ng, Yao Pu, Jing Cai 0001, Liang-Ting Lin, Jung Sun Yoo |
BIBM | 9 |
| 2025 | FuzzyMIL: Decoupling Pathological Phenotypes through Deep Fuzzy Clustering for Efficient Whole Slide Image AnalysisabstractIn Multiple Instance Learning (MIL) for Whole Slide Image (WSI) analysis, attention mechanisms are often employed to weigh the importance of different instances. However, global attention may lead to feature homogenization and overlook tissue differences. Adding local attention can capture these variations but is more parameter-intensive. To tackle the above challenges, in this work, we propose a deep fuzzy clustering framework (FuzzyMIL) based on a learnable Fuzzy C-means variant named FCM to analyze WSIs in a compact and efficient manner. By iteratively updating the fuzzy clustering centers, FCM decouples the morphological features of WSIs, resulting in more distinct and less correlated phenotypes. In this learning process, FCM compresses the feature representation space of WSI, guiding the features to gradually converge toward the representation prototypes. These prototypes, influenced by the soft assignment mechanism, take into account all updated features, enabling the model to retain both global information and local awareness. We evaluated our approach using three public datasets for diagnosis and sub-typing. Experimental results show that our approach achieves competitive performance while significantly reducing the downstream task framework parameters, striking a good balance between accuracy and model complexity. We release our code at https://github.com/Liuanana/FuzzyMIL. Anran Liu 0001, Jing Cai 0001, Srinivasa Sampath Veer Vajrala |
ICASSP | 3 |
| 2025 | DB-GNN: Dual-Branch Graph Neural Network with Multi-Level Contrastive Learning for Jointly Identifying Within- and Cross-Frequency Coupled Brain NetworksabstractWithin-frequency coupling (WFC) and cross-frequency coupling (CFC) in brain networks reflect neural synchronization within the same frequency band and cross-band oscillatory interactions, respectively. Their synergy provides a comprehensive understanding of neural mechanisms underlying cognitive states such as emotion. However, existing multi-channel EEG studies often analyze WFC or CFC separately, failing to fully leverage their complementary properties. This study proposes a dual-branch graph neural network (DB-GNN) to jointly identify within- and cross-frequency coupled brain networks. Firstly, DB-GNN leverages its unique dual-branch learning architecture to efficiently mine global collaborative information and local cross-frequency and within-frequency coupling information. Secondly, to more fully perceive the global information of cross-frequency and within-frequency coupling, the global perception branch of DB-GNN adopts a Transformer architecture. To prevent overfitting of the Transformer architecture, this study integrates prior within- and cross-frequency coupling information into the Transformer inference process, thereby enhancing the generalization capability of DB-GNN. Finally, a multi-scale graph contrastive learning regularization term is introduced to constrain the global and local perception branches of DB-GNN at both graph-level and node-level, enhancing its joint perception ability and further improving its generalization performance. Experimental validation on the emotion recognition dataset shows that DB-GNN achieves a testing accuracy of 97.88% and an F1-score of 97.87%, reaching the state-of-the-art performance. Jing Cai 0001, Ta Zhou, Xibei Yang |
IJCNN | 3 |
| 2025 | Boosting Generalizability in NPC ART Prediction via Multi-omics Feature Mapping
Jiabao Sheng, Zhe Li 0030, Saikit Lam, Jing Cai 0001 |
MICCAI (15) | 7 |
| 2025 | SynMSE: A multimodal similarity evaluator for complex distribution discrepancy in unsupervised deformable multimodal medical image registration
Jingke Zhu, Boyun Zheng, Bing Xiong 0004, Ming Cui, Deyu Sun, Jing Cai 0001, Yaoqin Xie, Wenjian Qin |
Medical Image Anal. | 7 |
| 2025 | Improved two-view interactional fuzzy learning based on mutual-rectification and knowledge-mergenceabstractNasopharyngeal carcinoma (NPC) is a malignant tumor that originates from the back of the nasal canal from above the soft palate to the upper larynx. Because the nasopharyngeal location is deeply hidden, it is often difficult for a single imaging means to clarify its complex adjacency. In addition, there exist some differences and uncertainties in its clinical manifestations. Although two-view fuzzy classifiers can effectively tap into the nasopharyngeal location for hidden information and exhibit good classification performance, existing fuzzy reasoning for predicting whether or not a nasopharyngeal cancer often stems from the inability to reuse the one-sided rules. Therefore, a novel two-view mutual rectification and knowledge mergence Takagi-Sugeno-Kang fuzzy classifier (TVRM-TFC) is proposed here to address the challenge of using imaging means to fine-tune the organ tissues. Firstly, Kullback-Leibler divergence (KLIC) is used to select important features from various imaging sections (i.e., pieces of knowledge). Secondly, the interpretable zero-order Takagi-Sugeno-Kang (TSK) fuzzy classifier is used as the basic training unit to simultaneously obtain satisfactory accuracies and concise linguistic interpretability. Thirdly, from the perspective of both imaging means and the organ, this study fine-tunes the information required for decision-making between different imaging means, so that the complementary advantages of the different views may improve the decision-making information and thus increase decision accuracies. Finally, the perspective of imaging technology and the organ are merged to capture decision-making knowledge. These decision-making advantages from different views are organically integrated to compensate information and further optimize the decision-making information. The merits of the proposed classifier are demonstrated through comparative experimental analysis on CT and MRI data. Ta Zhou, Wei Yan 0030, Zhengxin Xia, Shuihua Wang, Bing Li 0001, Weiping Ding 0001, Jing Cai 0001 |
Neural Networks | 8 |
| 2025 | A Lightweight TSK Fuzzy Classifier With Quantitative Equivalent Fuzzy Rules via Adaptive WeightingabstractThe first-order Takagi-Sugeno-Kang (TSK) fuzzy classifier with a fully combined fuzzy rule base (FuCo-FRB) is a potent and interpretable classifier for multiple input and multiple output (MIMO) tasks. However, FuCo-FRB possess an exponential increase in the number of fuzzy rules, and it poses challenges for efficient identification of the parameter matrix in MIMO tasks. A lightweight TSK fuzzy classifier (LW-TSK-FC) is proposed to achieve the balance of training efficiency and predictive performance especially on MIMO tasks. It has the following advantages: (1) An adaptive weighting method based on directly connected FRB enables the efficient generation of fuzzy rules, meanwhile retaining the distribution characteristic of FuCo-FRB to reduce information loss. (2) The consequent network of first-order TSK is optimized to a novel series structure by using matrix factorization. This new series structure increases the depth of consequent network and enable the implementation of kernel function and least learning machine (LLM), enhancing the calculation efficiency and predictive performance. (3) There are only weight parameters in consequent network of LW-TSK-FC and these parameters can be identified quantitatively by LLM, leading to a great training efficiency. The comparison experiments with other 14 power classifiers on 13 public UCI datasets shown comparable predictive performance and outstanding training and testing efficiency in both MISO and MIMO tasks. The experiments on a real-world clinical task demonstrated the significant capability of LW-TSK-FC on handling imbalanced small data, and can provide the semantic interpretability for reasoning process. Ta Zhou, Saikit Lam, Yuanpeng Zhang 0001, Defeng Sun, Jing Cai 0001 |
IEEE Trans. Fuzzy Syst. | 8 |
| 2025 | Bridging MRI Cross-Modality Synthesis and Multi-Contrast Super-Resolution by Fine-Grained Difference LearningabstractIn multi-modal magnetic resonance imaging (MRI), the tasks of imputing or reconstructing the target modality share a common obstacle: the accurate modeling of fine-grained inter-modal differences, which has been sparingly addressed in current literature. These differences stem from two sources: 1) spatial misalignment remaining after coarse registration and 2) structural distinction arising from modality-specific signal manifestations. This paper integrates the previously separate research trajectories of cross-modality synthesis (CMS) and multi-contrast super-resolution (MCSR) to address this pervasive challenge within a unified framework. Connected through generalized down-sampling ratios, this unification not only emphasizes their common goal in reducing structural differences, but also identifies the key task distinguishing MCSR from CMS: modeling the structural distinctions using the limited information from the misaligned target input. Specifically, we propose a composite network architecture with several key components: a label correction module to align the coordinates of multi-modal training pairs, a CMS module serving as the base model, an SR branch to handle target inputs, and a difference projection discriminator for structural distinction-centered adversarial training. When training the SR branch as the generator, the adversarial learning is enhanced with distinction-aware incremental modulation to ensure better-controlled generation. Moreover, the SR branch integrates deformable convolutions to address cross-modal spatial misalignment at the feature level. Experiments conducted on three public datasets demonstrate that our approach effectively balances structural accuracy and realism, exhibiting overall superiority in comprehensive evaluations for both tasks over current state-of-the-art approaches. The code is available at https://github.com/papshare/FGDL. Yidan Feng, Jing Cai 0001, Mingqiang Wei, Harry Qin |
IEEE Trans. Medical Imaging | 4 |
| 2024 | Organ boundary delineation for automated diagnosis from multi-center using ultrasound images
Tao Peng 0013, Yiyun Wu, Caishan Wang, Qingrong Jackie Wu, Jing Cai 0001 |
Expert Syst. Appl. | 6 |
| 2024 | Fuzzy inference system with interpretable fuzzy rules: Advancing explainable artificial intelligence for disease diagnosis - A comprehensive reviewabstractInterpretable artificial intelligence (AI), also known as explainable AI, is indispensable in establishing trustable AI for bench-to-bedside translation, with substantial implications for human well-being. However, the majority of existing research in this area has centered on designing complex and sophisticated methods, regardless of their interpretability. Consequently, the main prerequisite for implementing trustworthy AI in medical domains has not been met. Scientists have developed various explanation methods for interpretable AI. Among these methods, fuzzy rules embedded in a fuzzy inference system (FIS) have emerged as a novel and powerful tool to bridge the communication gap between humans and advanced AI machines. However, there have been few reviews of the use of FISs in medical diagnosis. In addition, the application of fuzzy rules to different kinds of multimodal medical data has received insufficient attention, despite the potential use of fuzzy rules in designing appropriate methodologies for available datasets. This review provides a fundamental understanding of interpretability and fuzzy rules, conducts comparative analyses of the use of fuzzy rules and other explanation methods in handling three major types of multimodal data (i.e., sequence signals, medical images, and tabular data), and offers insights into appropriate fuzzy rule application scenarios and recommendations for future research. Ta Zhou, Shaohua Zhi, Saikit Lam, Yuanpeng Zhang 0001, Yanjing Dong, Jing Cai 0001 |
Inf. Sci. | 9 |
| 2024 | Dual domain distribution disruption with semantics preservation: Unsupervised domain adaptation for medical image segmentation
Boyun Zheng, Songhui Diao, Jingke Zhu, Yixuan Yuan, Jing Cai 0001, Shuo Li 0001, Wenjian Qin |
Medical Image Anal. | 6 |
| 2024 | A multi-center study of ultrasound images using a fully automated segmentation architecture
Tao Peng 0013, Caishan Wang, Caiyin Tang, Yidong Gu, Jing Cai 0001 |
Pattern Recognit. | 7 |
| 2024 | Deep Reconciled and Self-Paced TSK Fuzzy System Ensemble for Imbalanced Data Classification: Architecture, Interpretability, and TheoryabstractStacking-based takagi-sugeno-kang (TSK) fuzzy system ensemble has been successfully applied to imbalanced data classification. However, there still exist many challenges that need to be further addressed. For example, during stacking, augmenting output variables into the input feature space reduces the interpretability of antecedents of fuzzy rules. During sampling for balancing, discovering informative samples usually only relies on training samples, which may reduce generalizability. More importantly, there is no theory to support the reliability of stacking. To address the aforementioned challenges, in this study, we propose a deep reconciled and self-paced TSK fuzzy system ensemble framework termed D-RSP-TSKE for imbalanced data classification. Compared with the existing ensemble frameworks, its superiorities can be exhibited from the following three aspects. First, in the first layer, we use random undersampling to generate a class-balanced training set to train an initial zero-order TSK fuzzy classifier. Based on the TSK fuzzy classifier, then we define classifier-specific and testing-compatible sample sensitivity to discover informative (high-sensitive) samples and design a reconciled and self-paced sampling approach to balance the minority class for the training of the following layers. Second, to improve the interpretability of antecedents of fuzzy rules, we propose to transfer the output variables from antecedents to consequents through equivalent mathematical transformations while keeping the final output unchanged. These transferred output variables are interpreted as the dynamic fuzzy rule confidence. Third, furthermore, we engage in a comprehensive theoretical examination of our stacking-based ensemble to elucidate the underlying mechanisms that enable the stacking strategy to consistently deliver superior performance. We conduct tests and comparisons on 7 artificial datasets and 30 real-world datasets to evaluate D-RSP-TSKE. The experimental results demonstrate the effectiveness and interpretability of D-RSP-TSKE for imbalanced data classification. Yuanpeng Zhang 0001, Guanjin Wang, Ta Zhou, Saikit Lam, Weiping Ding 0001, Jing Cai 0001 |
IEEE Trans. Fuzzy Syst. | 7 |
| 2024 | Model Generalizability Investigation for GFCE-MRI Synthesis in NPC Radiotherapy Using Multi-Institutional Patient-Based Data NormalizationabstractRecently, deep learning has been demonstrated to be feasible in eliminating the use of gadoliniumbased contrast agents (GBCAs) through synthesizing gadolinium-free contrast-enhanced MRI (GFCE-MRI) from contrast-free MRI sequences, providing the community with an alternative to get rid of GBCAs-associated safety issues in patients. Nevertheless, generalizability assessment of the GFCE-MRI model has been largely challenged by the high inter-institutional heterogeneity of MRI data, on top of the scarcity of multi-institutional data itself. Although various data normalization methods have been adopted to address the heterogeneity issue, it has been limited to single-institutional investigation and there is no standard normalization approach presently. In this study, we aimed at investigating generalizability of GFCE-MRI model using data from seven institutions by manipulating heterogeneity of MRI data under five popular normalization approaches. Three state-of-the-art neural networks were applied to map from T1-weighted and T2-weighted MRI to contrast-enhanced MRI (CE-MRI) for GFCE-MRI synthesis in patients with nasopharyngeal carcinoma. MRI data from three institutions were used separately to generate three uni-institution models and jointly for a tri-institution model. The five normalization methods were applied to normalize the data of each model. MRI data from the remaining four institutions served as external cohorts for model generalizability assessment. Quality of GFCE-MRI was quantitatively evaluated against ground-truth CE-MRI using mean absolute error (MAE) and peak signal-to-noise ratio(PSNR). Results showed that performance of all uni-institution models remarkably dropped on the external cohorts. By contrast, model trained using multi-institutional data with Z-Score normalization yielded the best model generalizability improvement. Wen Li 0010, Saikit Lam, Yinghui Wang 0003, Tian Li 0012, Jens Kleesiek, Andy Lai-Yin Cheung, Ying Sun 0015, Francis Kar-ho Lee, Kwok-hung Au, Victor Ho-fun Lee, Jing Cai 0001 |
IEEE J. Biomed. Health Informatics | 12 |
| 2024 | Coarse-Super-Resolution-Fine Network (CoSF-Net): A Unified End-to-End Neural Network for 4D-MRI With Simultaneous Motion Estimation and Super-ResolutionabstractFour-dimensional magnetic resonance imaging (4D-MRI) is an emerging technique for tumor motion management in image-guided radiation therapy (IGRT). However, current 4D-MRI suffers from low spatial resolution and strong motion artifacts owing to the long acquisition time and patients' respiratory variations. If not managed properly, these limitations can adversely affect treatment planning and delivery in IGRT. In this study, we developed a novel deep learning framework called the coarse-super-resolution-fine network (CoSF-Net) to achieve simultaneous motion estimation and super-resolution within a unified model. We designed CoSF-Net by fully excavating the inherent properties of 4D-MRI, with consideration of limited and imperfectly matched training datasets. We conducted extensive experiments on multiple real patient datasets to assess the feasibility and robustness of the developed network. Compared with existing networks and three state-of-the-art conventional algorithms, CoSF-Net not only accurately estimated the deformable vector fields between the respiratory phases of 4D-MRI but also simultaneously improved the spatial resolution of 4D-MRI, enhancing anatomical features and producing 4D-MR images with high spatiotemporal resolution. Shaohua Zhi, Yinghui Wang 0003, Haonan Xiao, Ti Bai, Bing Li 0001, Yunsong Tang, Wen Li 0010, Tian Li 0012, Jing Cai 0001 |
IEEE Trans. Medical Imaging | 11 |
| 2023 | Radiomics-Dosiomics-Contouromics Collaborative Learning for Adaptive Radiotherapy Eligibility Prediction in Nasopharyngeal CarcinomaabstractRadiomics, dosiomics and contouromics have been combined to predict adaptive radiotherapy eligibility in nasopharyngeal carcinoma. However, the commonly-used feature concatenation ignores the complementary or consistency relationship across different omics feature spaces. Also, the number of features increases with the concatenation of omics, leading to the curse of dimensionality and potential overfitting. To address the issues, in this study, multi-omics collaborative learning MOCL is developed. In MOCL, a priori knowledge-driven consistency regularization associated with Shannon entropy is designed to automatically explore the weighting consistency across different omics feature spaces. In addition, a label soften strategy is adopted to enlarge the margins between different classes, rendering more freedom for the models to fit the soft label matrix. To avoid overfitting, we design a regularized term deduced from manifold learning to keep samples in the label space as close as possible if they are in the same manifold in the feature space. Experimental results on 311 nasopharyngeal carcinoma patients collected from the Hong Kong Queen Elizabeth Hospital demonstrate the promising performance of MOCL. Yuanpeng Zhang 0001, Saikit Lam, Xinzhi Teng, Chengyu Qiu, Jing Cai 0001 |
BIBM | 6 |
| 2023 | Delineation of Prostate Boundary from Medical Images via a Mathematical Formula-Based Hybrid Algorithm
Tao Peng 0013, Daqiang Xu, Yiyun Wu, Jing Cai 0001 |
ICANN (8) | 6 |
| 2023 | Clinical Evaluation of AI-Assisted Virtual Contrast Enhanced MRI in Primary Gross Tumor Volume Delineation for Radiotherapy of Nasopharyngeal Carcinoma
Wen Li 0010, Saikit Lam, Yaoqin Xie, Wenjian Qin, Andy Lai-Yin Cheung, Haonan Xiao, Francis Kar-ho Lee, Kwok-hung Au, Victor Ho-fun Lee, Jing Cai 0001, Tian Li 0012 |
MICCAI (7) | 14 |
| 2023 | Multi-view Contrastive Learning with Additive Margin for Adaptive Nasopharyngeal Carcinoma Radiotherapy PredictionabstractThe accurate prediction of adaptive radiation therapy (ART) for nasopharyngeal carcinoma (NPC) patients before radiation therapy (RT) is crucial for minimizing toxicity and enhancing patient survival rates. Owing to the complexity of the tumor micro-environment, a single high-resolution image offers only limited insight. Furthermore, the traditional softmax-based loss falls short in quantifying a model’s discriminative power. To address these challenges, we introduce a supervised multi-view contrastive learning approach with an additive margin (MMCon). For each patient, we consider four medical images to form multi-view positive pairs, which supply supplementary information and bolster the representation of medical images. We employ supervised contrastive learning to determine the embedding space, ensuring that NPC samples from the same patient or with the same labels stay in close proximity while NPC samples with different labels are distant. To enhance the discriminative ability of the loss function, we incorporate a margin into the contrastive learning process. Experimental results show that this novel learning objective effectively identifies an embedding space with superior discriminative abilities for NPC images. Jiabao Sheng, Saikit Lam, Zhe Li 0030, Xinzhi Teng, Yuanpeng Zhang 0001, Jing Cai 0001 |
ICMR | 7 |
| 2023 | Contour Detection from Ultrasound Kidney Images with A Coarse-to-Fine ApproachabstractUltrasound kidney image segmentation presents significant challenges due to missing or ambiguous boundaries. In this study, we introduce a coarse-to-refinement approach incorporating four novel aspects. Firstly, we leverage the properties of a principal curve (PC) to automatically fine-tune the curve shape and employ a neural network's learning ability to reduce model error. Secondly, a deep fusion learning network is utilized for the coarse segmentation step, incorporating a parallel architecture to enhance deep-learning performance. Thirdly, addressing the limitation of standard PC-based methods in determining the number of vertices automatically, we propose an automatic searching polygon tracking method using a mean shift clustering-based approach to replace the projection and vertex extension step in standard PC-based methods. Lastly, we develop an explainable mathematical map function for the kidney contour, as denoted by the neural network output (i.e., optimized vertices), which aligns well with the ground truth contour. We conducted various experiments to evaluate our method's performance, demonstrating its effectiveness in ultrasound kidney image segmentation. Tao Peng 0013, Yidong Gu, Caishan Wang, Jing Cai 0001 |
SMC | 6 |
| 2023 | Hybrid Intelligent-Annotation Organ Segmentation on Medical DatasetsabstractUltrasound image segmentation is crucial for early disease detection and treatment planning but remains a challenging task due to the low contrast of organ boundaries and varying image quality. Current methods often require manual intervention or have limited accuracy. In this paper, we propose a novel hybrid framework that combines an automatic option polygon segment (AOPS) algorithm and a distributed- and memory-based evolution (DME) algorithm for precise ultrasound organ segmentation. Our pipeline consists of two cascaded stages: (1) a coarse segmentation step using the AOPS algorithm, which determines the number of vertices/clusters without human intervention, and (2) a refinement step using the DME algorithm to hunt for the optimal neural network, which is then used to represent a smooth, explainable mathematical expression of the organ boundary. We employ the fractional backpropagation learning network with L2 regularization (FBLN) for training and use the scaled exponential linear unit (SELU) activation function to address the vanishing gradient problem. This is a new attempt such a hybrid framework is applied to ultrasound organ segmentation tasks, and it demonstrates significant contributions in terms of accuracy, smoothness, and computational efficiency. Tao Peng 0013, Yidong Gu, Gongye Di, Jing Cai 0001 |
SMC | 6 |
| 2023 | Coarse-to-fine tuning knowledgeable system for boundary delineation in medical images
Tao Peng 0013, Yiyun Wu, Caishan Wang, Yuntian Shen, Jing Cai 0001 |
Appl. Intell. | 7 |
| 2023 | Automatic coarse-to-refinement-based ultrasound prostate segmentation using optimal polyline segment tracking method and deep learning
Tao Peng 0013, Daqiang Xu, Caiyin Tang, Yuntian Shen, Jing Cai 0001 |
Appl. Intell. | 7 |
| 2022 | Explainability-guided Mathematical Model-Based Segmentation of Transrectal Ultrasound Images for Prostate BrachytherapyabstractAccurate segmentation of the prostate is important to image-guided prostate biopsy and brachytherapy treatment planning. However, the incompleteness of prostate boundary increases the challenges in the automatic ultrasound prostate segmentation task. In this work, an automatic coarse-to-fine framework for prostate segmentation was developed and tested. Our framework has four metrics: first, it combines the ability of deep learning model to automatically locate the prostate and integrates the characteristics of principal curve that can automatically fit the data center for refinement. Second, to well balance the accuracy and efficiency of our method, we proposed an intelligent determination of the data radius algorithm-based modified polygon tracking method. Third, we modified the traditional quantum evolution network by adding the numerous-operator scheme and global optimum search scheme for ensuring population diversity and achieving the optimal model parameters. Fourth, we found a suitable mathematical function expressed by the parameters of the machine learning model to smooth the contour of the prostate. Results on the multiple datasets demonstrate that our method has good segmentation performance. Tao Peng 0013, Yiyun Wu, Jin Wang 0009, Jing Cai 0001 |
BIBM | 6 |
| 2022 | Multi-institutional Investigation of Model Generalizability for Virtual Contrast-Enhanced MRI Synthesis
Wen Li 0010, Saikit Lam, Tian Li 0012, Andy Lai-Yin Cheung, Haonan Xiao, Xinzhi Teng, Shaohua Zhi, Francis Kar-ho Lee, Kwok-hung Au, Victor Ho-fun Lee, Amy Tien Yee Chang, Jing Cai 0001 |
MICCAI (8) | 15 |
| 2022 | H-SegMed: A Hybrid Method for Prostate Segmentation in TRUS Images via Improved Closed Principal Curve and Improved Enhanced Machine Learning
Tao Peng 0013, Caiyin Tang, Yiyun Wu, Jing Cai 0001 |
Int. J. Comput. Vis. | 4 |
| 2022 | Integration of an imbalance framework with novel high-generalizable classifiers for radiomics-based distant metastases prediction of advanced nasopharyngeal carcinoma
Yuanpeng Zhang 0001, Saikit Lam, Xinzhi Teng, Francis Kar-ho Lee, Kwok-hung Au, Celia Wai-yi Yip, Shitong Wang 0001, Jing Cai 0001 |
Knowl. Based Syst. | 10 |
| 2022 | Glioma segmentation of optimized 3D U-net and prediction of multi-modal survival time
Qihong Liu, Kai Liu 0039, Antonio Bolufé Röhler, Jing Cai 0001 |
Neural Comput. Appl. | 4 |
| 2022 | H-ProMed: Ultrasound image segmentation based on the evolutionary neural network and an improved principal curve
Tao Peng 0013, Yidong Gu, Caishan Wang, Yiyun Wu, Xiuxiu Cheng, Jing Cai 0001 |
Pattern Recognit. | 7 |
| 2021 | Fuzzy Clustering Based on Automated Feature Pattern-Driven Similarity Matrix ReductionabstractMost of the medoid-based fuzzy clustering algorithms only use one similarity matrix to organize objects into groups. The similarity matrix is often constructed by equally employing all features which may ignore the contribution differences existing among the features. In this study, we also propose a medoid-based fuzzy clustering algorithm feature pattern-driven similarity matrices-reduction based fuzzy clustering (FP-SMR-FC) which is different from the existing ones in the following two aspects. First, multiple similarity matrices are constructed to represent the similarity between objects. Additionally, a feature pattern-driven Shannon entropy which combines nondeterminacy information contained in the similarity matrices and statistical information contained in the features together is used to learn the weight of each similarity matrix. Second, during the clustering processes of FP-SMR-FC, a new schema for eliminating some of the similarity matrices with very few contributions is developed for similarity matrices reduction. The comparison studies in terms of time complexity and clustering accuracy for FP-SMR-FC with various medoid-based clustering algorithms on real-life data sets are done. In addition, FP-SMR-FC is applied to head pose estimation of human behavior analysis. Comparisons, indeed, demonstrate the promising performance of FP-SMR-FC in practice. Yuanpeng Zhang 0001, Jing Cai 0001 |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2018 | Application of the 4-D XCAT Phantoms in Biomedical Imaging and BeyondabstractThe four-dimensional (4-D) eXtended CArdiac-Torso (XCAT) series of phantoms was developed to provide accurate computerized models of the human anatomy and physiology. The XCAT series encompasses a vast population of phantoms of varying ages from newborn to adult, each including parameterized models for the cardiac and respiratory motions. With great flexibility in the XCAT's design, any number of body sizes, different anatomies, cardiac or respiratory motions or patterns, patient positions and orientations, and spatial resolutions can be simulated. As such, the XCAT phantoms are gaining a wide use in biomedical imaging research. There they can provide a virtual patient base from which to quantitatively evaluate and improve imaging instrumentation, data acquisition, techniques, and image reconstruction and processing methods which can lead to improved image quality and more accurate clinical diagnoses. The phantoms have also found great use in radiation dosimetry, radiation therapy, medical device design, and even the security and defense industry. This review paper highlights some specific areas in which the XCAT phantoms have found use within biomedical imaging and other fields. From these examples, we illustrate the increasingly important role that computerized phantoms and computer simulation are playing in the research community. William Paul Segars, Benjamin M. W. Tsui, Jing Cai 0001, Fang-Fang Yin, George S. K. Fung, Ehsan Samei |
IEEE Trans. Medical Imaging | 3 |