Zhiyuan Cai

dblp:52/5232 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Knowledge-Enhanced Explainable Prompting for Vision-Language Models
abstract
Large-scale vision-language models (VLMs) embedded with expansive representations and visual concepts have showcased significant potential in image and text understanding. Efficiently adapting VLMs such as CLIP to downstream tasks like few-shot image classification has garnered growing attention, with prompt learning emerging as a representative approach. However, most existing prompt-based adaptation methods, which rely solely on coarse-grained textual prompts, suffer from limited performance and interpretability when handling domain tasks that require specific knowledge. This results in a failure to satisfy the stringent trustworthiness requirements of Explainable Artificial Intelligence (XAI) in high-risk scenarios like healthcare. To address this issue, we propose a Knowledge-Enhanced Explainable Prompting (KEEP) framework that leverages fine-grained domain-specific knowledge to enhance the adaptation process of VLMs across various domains and image modalities. By incorporating retrieval augmented generation and domain foundation models, our framework can provide more reliable image-wise knowledge for prompt learning in various domains, alleviating the lack of fine-grained annotations, while offering both visual and textual explanations. Extensive experiments and explainability analyses conducted on eight datasets of different domains and image modalities demonstrate that our method simultaneously achieves superior performance and interpretability, highlighting the effectiveness of the collaboration between foundation models and XAI.
Yequan Bie, Andong Tan, Zhixuan Chen, Zhiyuan Cai, Luyang Luo, Hao Chen 0011
AAAI4
2026 An efficient constrained multi-objective evolutionary algorithm with a spatial discretization evaluation mechanism for unmanned aerial vehicle path planning
Zhiyuan Cai, Chaoda Peng, Junyan Lin, Yueting Xu, Haoyu Luo
Eng. Appl. Artif. Intell.1
2026 SurgPETL: Parameter-Efficient Image-to-Surgical-Video Transfer Learning for Surgical Phase Recognition
abstract
Capitalizing on image-level pre-trained models for various downstream tasks has recently emerged with promising performance. However, the paradigm of "image pre-training followed by video fine-tuning" for high-dimensional video data inevitably introduces significant performance bottlenecks. Furthermore, in the medical domain, many surgical video tasks encounter additional challenges posed by the limited availability of video data and the necessity for comprehensive spatiotemporal modeling. Recently, Parameter-Efficient Image-to-Video Transfer Learning (PEIVTL) has emerged as an efficient and effective paradigm for video action recognition tasks, which employs image-level pre-trained models with promising feature transferability and involves cross-modality temporal modeling with minimal fine-tuning. Nevertheless, the effectiveness and generalizability of this paradigm within intricate surgical domain remain unexplored. In this paper, we delve into a novel problem of efficiently adapting image-level pre-trained models to specialize in fine-grained surgical phase recognition, termed Parameter-Efficient Image-to-Surgical-Video Transfer Learning. First, we develop SurgPETL, a parameter-efficient transfer learning framework for surgical phase recognition, and conduct extensive experiments with three advanced methods based on ViTs of two distinct scales pre-trained on five large-scale natural and medical datasets. Then, we introduce the Adaptive Spatiotemporal Representation Modulation (ASRM) module, integrating a standard spatial adapter with a novel temporal adapter to capture detailed spatial features and establish connections across temporal sequences for robust spatiotemporal modeling. Extensive experiments on three challenging datasets spanning various surgical procedures demonstrate the effectiveness of SurgPETL with ASRM. SurgPETL-ASRM outperforms both parameter-efficient alternatives and state-of-the-art surgical phase recognition methods while maintaining parameter efficiency and minimizing overhead.
Shu Yang 0004, Zhiyuan Cai, Luyang Luo, Shuchang Xu, Hao Chen 0011
IEEE Trans. Medical Imaging2
2026 QoS-Aware Deep Reinforcement Learning for Dynamic CPU Pinning of Co-Located Cloud Workloads
abstract
In cloud computing, static resource configurations create a trade-off: tenants overprovision to avoid resource starvation, causing inefficiency and cost, while providers suffer low utilization despite high allocations. To improve efficiency, providers often use overcommitted environments where multiple workloads share hosts, but this leads to interference and potential Quality-of-Service (QoS) violations. This paper introduces a realtime dynamic control framework that mitigates interference by adaptively pinning workloads to CPU groups. Using deep reinforcement learning (DRL) with the Proximal Policy Optimization (PPO) algorithm, an intelligent agent continuously adjusts CPU pinning based on real-time feedback to maintain Service- Level-Agreement (SLA) compliance. Experiments under two optimization objectives—overall-performance-first and priorityperformance- first—show that the proposed approach improves overall QoS by$\approx$25% compared with static pinning. When prioritization is enabled, high-priority workloads gain significant performance improvements while lower-priority ones remain within SLA limits. These results demonstrate that a DRL-based CPU-pinning strategy effectively manages resource contention in overcommitted clouds, enhancing utilization while upholding tenant SLAs.
Dongji Lu, Weipeng Cao, Jiongjiong Gu, Zhiyuan Cai, Chuanfei Xu, Liang-Jie Zhang, Zhong Ming 0001
IEEE Trans. Serv. Comput.5
2025 FedLPPA: Learning Personalized Prompt and Aggregation for Federated Weakly-Supervised Medical Image Segmentation
abstract
Federated learning (FL) effectively mitigates the data silo challenge brought about by policies and privacy concerns, implicitly harnessing more data for deep model training. However, traditional centralized FL models grapple with diverse multi-center data, especially in the face of significant data heterogeneity, notably in medical contexts. In the realm of medical image segmentation, the growing imperative to curtail annotation costs has amplified the importance of weakly-supervised techniques which utilize sparse annotations such as points, scribbles, etc. A pragmatic FL paradigm shall accommodate diverse annotation formats across different sites, which research topic remains under-investigated. In such context, we propose a novel personalized FL framework with learnable prompt and aggregation (FedLPPA) to uniformly leverage heterogeneous weak supervision for medical image segmentation. In FedLPPA, a learnable universal knowledge prompt is maintained, complemented by multiple learnable personalized data distribution prompts and prompts representing the supervision sparsity. Integrated with sample features through a dual-attention mechanism, those prompts empower each local task decoder to adeptly adjust to both the local distribution and the supervision form. Concurrently, a dual-decoder strategy, predicated on prompt similarity, is introduced for enhancing the generation of pseudo-labels in weakly-supervised learning, alleviating overfitting and noise accumulation inherent to local data, while an adaptable aggregation method is employed to customize the task decoder on a parameter-wise basis. Extensive experiments on four distinct medical image segmentation tasks involving different modalities underscore the superiority of FedLPPA, with its efficacy closely parallels that of fully supervised centralized training. Our code and data will be available at https://github.com/llmir/FedLPPA.
Li Lin 0006, Yixiang Liu, Jiewei Wu, Pujin Cheng, Zhiyuan Cai, Kenneth K. Y. Wong, Xiaoying Tang 0001
IEEE Trans. Medical Imaging5
2025 Flexible Computing: A New Framework for Improving Resource Allocation and Scheduling in Elastic Computing
abstract
Since the advent of cloud computing, Elastic Computing (EC) has become the standard architecture for resource allocation and scheduling. EC typically allocates computing resources based on predefined specifications, such as virtual machine or container flavors. However, these flavors are often constrained by fixed CPU-to-memory ratios, which frequently fail to match the actual resource needs of applications. As a result, cloud providers experience high resource allocation rates nearing saturation ($> $80%) but with low utilization ($< $25%). This study introduces Flexible Computing (FC), a novel approach to resource allocation and scheduling. Unlike EC, FC allocates resources based on an application resource usage profile, derived from the historical resource consumption of workloads, rather than relying on fixed specifications. Additionally, FC incorporates a real-time performance degradation detection mechanism to address performance issues caused by the noisy-neighbor effect when colocated workloads interfere with each other. FC dynamically adjusts resource allocation according to actual usage, ensuring that application performance meets Service Level Agreements (SLAs), while preventing resource waste and performance degradation from improper resource over-commitment. Large-scale experimental validations conducted on the FC architecture within Huawei Cloud data centers demonstrate that, compared to EC, FC can reduce computing resource consumption by over 33% while managing the same workloads. Furthermore, FC's real-time performance degradation detection model achieves a prediction error of less than 5% across various testing environments, highlighting its commercial viability.
Weipeng Cao, Jiongjiong Gu, Zhong Ming 0001, Zhiyuan Cai, Yuzhao Wang, Changping Ji, Zhijiao Xiao, Yuhong Feng, Liang-Jie Zhang
IEEE Trans. Serv. Comput.4
2024 BPaCo: Balanced Parametric Contrastive Learning for Long-Tailed Medical Image Classification
Zhiyuan Cai, Tianyunxi Wei, Li Lin 0006, Hao Chen 0011, Xiaoying Tang 0001
MICCAI (1)1
2024 Safe Reinforcement Learning-Based Motion Planning for Functional Mobile Robots Suffering Uncontrollable Mobile Robots
abstract
An increasing number of Autonomous Mobile Robots (AMRs) are used in warehouses and factories in recent years. The risk of some of the AMRs being out of control is surging. Although Reinforcement Learning (RL)-based approaches have achieved dramatic success in the motion planning of a large number of AMRs, the available RL-based motion planning approaches cannot provide a safety guarantee for the remaining functional AMRs if some of the AMRs are out of control. To this end, this paper develops a scalable Multi-agent RL (MARL) with Control Barrier Function (CBF)-based shields algorithm. The MARL with CBF-based shields algorithm can address complex high-level tasks by MARL and deal with the safety issue of every single functional AMR by a low-level CBF-based shield. A CBF-based shield is designed for every single functional AMR to ensure that the action of the functional AMR is safe, even if an uncontrollable AMR is pursuing the functional AMR. Experiments are conducted based on simulated warehouse environments to evaluate the effectiveness and scalability of a safe RL-based motion planning approach (The safe RL-based motion planning approach developed in this study is demonstrated in a video: https://youtu.be/I7ja5nFVpY4). developed according to the MARL with CBF-based shields algorithm.
Huanhui Cao, Hao Xiong 0004, Weifeng Zeng, Hantao Jiang, Zhiyuan Cai, Liang Hu 0002, Lin Zhang 0060, Wenjie Lu 0004
IEEE Trans. Intell. Transp. Syst.5
2024 Uni4Eye++: A General Masked Image Modeling Multi-Modal Pre-Training Framework for Ophthalmic Image Classification and Segmentation
abstract
A large-scale labeled dataset is a key factor for the success of supervised deep learning in most ophthalmic image analysis scenarios. However, limited annotated data is very common in ophthalmic image analysis, since manual annotation is time-consuming and labor-intensive. Self-supervised learning (SSL) methods bring huge opportunities for better utilizing unlabeled data, as they do not require massive annotations. To utilize as many unlabeled ophthalmic images as possible, it is necessary to break the dimension barrier, simultaneously making use of both 2D and 3D images as well as alleviating the issue of catastrophic forgetting. In this paper, we propose a universal self-supervised Transformer framework named Uni4Eye++ to discover the intrinsic image characteristic and capture domain-specific feature embedding in ophthalmic images. Uni4Eye++ can serve as a global feature extractor, which builds its basis on a Masked Image Modeling task with a Vision Transformer architecture. On the basis of our previous work Uni4Eye, we further employ an image entropy guided masking strategy to reconstruct more-informative patches and a dynamic head generator module to alleviate modality confusion. We evaluate the performance of our pre-trained Uni4Eye++ encoder by fine-tuning it on multiple downstream ophthalmic image classification and segmentation tasks. The superiority of Uni4Eye++ is successfully established through comparisons to other state-of-the-art SSL pre-training methods. Our code is available at https://github.com/Davidczy/Uni4Eye++.
Zhiyuan Cai, Li Lin 0006, Huaqing He, Pujin Cheng, Xiaoying Tang 0001
IEEE Trans. Medical Imaging1
2022 Uni4Eye: Unified 2D and 3D Self-supervised Pre-training via Masked Image Modeling Transformer for Ophthalmic Image Classification
Zhiyuan Cai, Li Lin 0006, Huaqing He, Xiaoying Tang 0001
MICCAI (8)1
2014 Clustering Image Search Results by Entity Disambiguation
Kaiqi Zhao 0001, Zhiyuan Cai, Qingyu Sui, Enxun Wei, Kenny Q. Zhu
ECML/PKDD (3)2
2013 Wikification via link co-occurrence
abstract
Wikification, which stands for the process of linking terms in a plain text document to Wikipedia articles which represent the correct meanings of the terms, can be thought of as a generalized Word Sense Disambiguation problem. It disambiguates multi-word expressions (MWEs) in addition to single words. Existing Wikification techniques either models the context of a given term as well as the Wikipedia article as bags of words, or compute global constraints among Wikipedia concepts by the link graph or link distributions. The first method doesn't achieve good results because the MWEs can have very different meanings than its constituent words which themselves are ambiguous. The second method doesn't produce high accuracy because the link structure or link distribution is often biased or incomplete by themselves due to the fact that Wikipedia pages are often sparsely linked. In this paper, we present a simple but powerful framework of sense disambiguation using co-occurrences of Wikipedia links in the Wikipedia corpus. We propose an iterative method to enrich the sparsely-linked articles by adding more links and then use the resulting link co-occurrence matrix to disambiguate an input document by a sliding window algorithm. Our prototype system achieves 89.97% precision and 76.43% recall on average for three benchmark data and compares favorably against four state-of-the-art wikification techniques.
Zhiyuan Cai, Kaiqi Zhao 0001, Kenny Q. Zhu, Haixun Wang
CIKM1
2007 Time Series Prediction of Short Circuit Current for Synchronous Control of Synthetic Test Based on Delay Coordinate Embedding
Zhiyuan Cai, Yangyang Ge, Erzhi Wang
ICIC (3)1
2007 An Approach of Combustion Diagnosis in Boiler Furnace Based on Phase Space Reconstruction
Zhiyuan Cai, Ying Hua, Yangyang Ge
ICIC (3)2
2003 M-Kernel Merging: Towards Density Estimation over Data Streams
abstract
Density estimation is a costly operation for computing distribution information of data sets underlying many important data mining applications, such as clustering and biased sampling. However, traditional density estimation methods are inapplicable for streaming data, which are continuously arriving large volume of data, because of their request for linear storage and square size calculation. The shortcoming limits the application of many existing effective algorithms on data streams, for which the mining problem is an emergency for applications and a challenge for research. In this paper, the problem of computing density functions over data streams is examined. A novel method attacking this shortcoming of existing methods is developed to enable density estimation for large volume of data in linear time, fixed size memory, and without lose of accuracy. The method is based on M-Kernel merging, so that limited kernel functions to be maintained are determined intelligently, The application of the new method on different streaming data models is discussed, and the result of intensive experiments is presented. The analytical and empirical result show that this new density estimation algorithm for data streams can calculate density functions on demand at any time with high accuracy for different streaming data models.
Aoying Zhou, Zhiyuan Cai, Weining Qian
DASFAA2