VLDB 2026 Research / reviewers in the wild / expert
Zhiyi Shi
dblp:318/2817
· DBLP profile ↗
13ranked-venue papers
2as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Virtual Multiplex Staining for Histological Images Using a Marker-Wise Conditioned Diffusion ModelabstractMultiplex imaging is revolutionizing pathology by enabling the simultaneous visualization of multiple biomarkers within tissue samples, providing molecular-level insights that traditional hematoxylin and eosin (H&E) staining cannot provide. However, the complexity and cost of multiplex data acquisition have hindered its widespread adoption. Additionally, most existing large repositories of H&E images lack corresponding multiplex images, limiting opportunities for multi-modal analysis. To address these challenges, we leverage recent advances in latent diffusion models (LDMs), which excel at modeling complex data distributions by utilizing their powerful priors for fine-tuning to a target domain. In this paper, we introduce a novel framework for virtual multiplex staining that utilizes pretrained LDM parameters to generate multiplex images from H&E images using a conditional diffusion model. Our approach enables marker-by-marker generation by conditioning the diffusion model on each marker, while sharing the same architecture across all markers. To tackle the challenge of varying pixel value distributions across different marker stains and to improve inference speed, we fine-tune the model for single-step sampling, enhancing both color contrast fidelity and inference efficiency through pixel-level loss functions. We validate our framework on two publicly available datasets, notably demonstrating its effectiveness in generating up to 18 different marker types with improved accuracy, a substantial increase over the 2-3 marker types achieved in previous approaches. This validation highlights the potential of our framework, pioneering virtual multiplex staining. Finally, this paper bridges the gap between H&E and multiplex imaging, potentially enabling retrospective studies and large-scale analyses of existing H&E image repositories. Hyun-Jic Oh, Junsik Kim 0001, Zhiyi Shi, Yu-An Chen, Peter K. Sorger, Hanspeter Pfister, Won-Ki Jeong |
AAAI | 3 |
| 2026 | Task-Specific Directions: Definition, Exploration, and Utilization in Parameter Efficient Fine-TuningabstractLarge language models demonstrate impressive performance on downstream tasks, yet requiring extensive resource consumption when fully fine-tuning all parameters. To mitigate this, Parameter Efficient Fine-Tuning (PEFT) strategies, such as LoRA, have been developed. In this paper, we delve into the concept of task-specific directions (TSDs)-critical for transitioning large models from pretrained states to task-specific enhancements in PEFT. We propose a framework to clearly define these directions and explore their properties, and practical utilization challenges. We then introduce a novel approach, LoRA-Dash, which aims to maximize the impact of TSDs during the fine-tuning process, thereby enhancing model performance on targeted tasks. Additionally, based on our exploration of TSD, we focus on an important issue in PEFT: the initialization of LoRA. While some works have pointed out the significance of initialization for LoRA's performance and proposed various strategies, these methods are often empirical and not task-specific. To address this issue, we propose LoRA-Init. Starting from TSD, we identify the directions that require the most adjustment during fine-tuning for downstream tasks. By initializing the matrices in LoRA with these directions, LoRA-Init significantly enhances LoRA's performance. Moreover, we can combine LoRA-Dash and LoRA-Init to create the final version of LoRA based on TSDs, which we refer to as LoRA-TSD. Extensive experiments have conclusively demonstrated the effectiveness of these methods, and in-depth analyses further reveal the underlying mechanisms of these methods. Chongjie Si, Zhiyi Shi, Shifan Zhang, Xiaokang Yang 0001, Hanspeter Pfister, Wei Shen 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Evolutionary Contrastive Ensemble With Conditional Redundancy Fitness Evaluation for Goal-Conditioned Humanoid LocomotionabstractGoal-conditioned humanoid locomotion in reinforcement learning (RL) remains challenging due to sparse reward signals and the single-goal overfitting problem. Although contrastive reinforcement learning (CRL) has achieved considerable success in this setting, it can suffer from pronounced estimation variance, since epistemic uncertainty is difficult to reduce given limited task-specific information and model capacity. Ensemble-based critics can partially alleviate this issue. However, sufficient ensemble diversity and accurate individual estimates are not necessarily guaranteed during training, resulting in unstructured exploration. To address these challenges, we propose Conditional Redundancy-Guided Evolutionary Contrastive Ensemble with Direct Preference Optimization weighting (CRECE-DPO), which augments CRL with a vectorized critic ensemble and refines the ensemble via an evolutionary algorithm guided by a tailored fitness metric. Specifically, we design a DPO-weighted conditional redundancy fitness score, to prune redundant representations while promoting effective exploration of the parameter space. Simulation results on challenging goal-conditioned benchmarks, including humanoid locomotion, demonstrate consistent improvements over CRL and other baselines. Zhiyi Shi, Haoyu Pan, Ruihao Zhu, Changyu Li, Shuai Wu 0004, Qi Wu 0003 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2026 | Channel-Robust Radio Frequency Fingerprint Extraction Based on Channel ReciprocityabstractRecently, Radio Frequency Fingerprint (RFF) technology has emerged as a promising technique for physical layer authentication. However, overcoming the interference from wireless multipath channels remains a key challenge for robust RFF extraction. Existing channel-robust studies often suppress channel effects at the expense of intrinsic RFF information, and the retained RFF features may lack sufficient discrimination. This paper proposes a channel-robust RFF extraction scheme based on channel reciprocity. Firstly, a Channel State Information (CSI) feedback mechanism is proposed to collect reciprocal uplink and downlink CSI. Then, CSI preprocessing methods and the channel reciprocity judging method based on Mean Absolute Distance (MAD), are exploited to further enhance CSI reciprocity. Finally, a novel RFF extraction algorithm eliminates reciprocal channel components by calculating the quotient of uplink and downlink CSI, thereby retaining the differentiated device-specific RFF features. A unified identification framework supports both in-library classification and unknown device detection. Extensive experiments using 41 ESP32 development kits under various Wi-Fi channel environments show that the extracted RFF features are channel-robust, highly distinctive, and stable over time. Specifically, in a static scenario with strong multipath effects, the in-library classification accuracy reaches 96.86%. And the same metric in a dynamic scenario with significant interference remains 93.48%. Furthermore, new-device recognition accuracy remains above 94% across all scenarios, with a False Negative Rate (FNR) below 1.3%. Bingshu Dong, Aiqun Hu, Jiabao Yu, Zhiyi Shi |
IEEE Trans. Commun. | 5 |
| 2025 | Generalized Tensor-Based Parameter-Efficient Fine-Tuning via Lie Group TransformationsabstractAdapting pre-trained foundation models for diverse downstream tasks is a core practice in artificial intelligence. However, the wide range of tasks and high computational costs make full fine-tuning impractical. To overcome this, parameter-efficient fine-tuning (PEFT) methods like LoRA have emerged and are becoming a growing research focus. Despite the success of these methods, they are primarily designed for linear layers, focusing on two-dimensional matrices while largely ignoring higher-dimensional parameter spaces like convolutional kernels. Moreover, directly applying these methods to higher-dimensional parameter spaces often disrupts their structural relationships. Given the rapid advancements in matrix-based PEFT methods, rather than designing a specialized strategy, we propose a generalization that extends matrix-based PEFT methods to higher-dimensional parameter spaces without compromising their structural properties. Specifically, we treat parameters as elements of a Lie group, with updates modeled as perturbations in the corresponding Lie algebra. These perturbations are mapped back to the Lie group through the exponential map, ensuring smooth, consistent updates that preserve the inherent structure of the parameter space. Extensive experiments on computer vision and natural language processing validate the effectiveness and versatility of our approach, demonstrating clear improvements over existing methods. Chongjie Si, Zhiyi Shi, Yichen Xiao, Xiaokang Yang 0001, Wei Shen 0002 |
ICCV | 2 |
| 2025 | Unleashing the Power of Task-Specific Directions in Parameter Efficient Fine-tuningabstractLarge language models demonstrate impressive performance on downstream tasks, yet requiring extensive resource consumption when fully fine-tuning all parameters. To mitigate this, Parameter Efficient Fine-Tuning (PEFT) strategies, such as LoRA, have been developed.
In this paper, we delve into the concept of task-specific directions (TSDs)—critical for transitioning large models from pretrained states to task-specific enhancements in PEFT. We propose a framework to clearly define these directions and explore their properties, and practical utilization challenges. We then introduce a novel approach, LoRA-Dash, which aims to maximize the impact of TSDs during the fine-tuning process, thereby enhancing model performance on targeted tasks. Extensive experiments have conclusively demonstrated the effectiveness of LoRA-Dash, and in-depth analyses further reveal the underlying mechanisms of LoRA-Dash. Chongjie Si, Zhiyi Shi, Shifan Zhang, Xiaokang Yang 0001, Hanspeter Pfister, Wei Shen 0002 |
ICLR | 2 |
| 2025 | "Jumpingly" Perceive Time Series: Image Generation Approach to Modeling Functional Brain ActivationabstractThis paper presents a novel Linear Mapping Field (LMF) to map time series into two-dimensional images. The LMF extracts deeper features of fNIRS signals, which makes fNIRS less reliant on some prior. The developed convolution neural networks detect more prominent features than the state-of-the-art methods. The experimental results indicate that different from RNNs which can only perceive the time series in a “sequential” manner, LMF’s characteristic of “jumpingly” perception is the key to achieving excellent results.Note to Practitioners—As an optical and non-invasive technique to obtain the changes of oxyhemoglobin (O2Hb) and deoxyhemoglobin (HHb), fNIRS can be used to measure the changes of cerebral hemodynamics related to brain activities. This work proposes a linear mapping field and other mapping fields to map fNIRS signals to two-dimensional images, and the deep features of these generated images can be further extracted by convolutional neural networks, establishing an end-to-end bridge to detect the activation of brain functions under different tasks. Compared with mainstream methods, this work can be used as an effective mapping for fNIRS with low computational complexity and good performance. Kevin W. Tong, Miaomiao Zhang 0001, Zhiyi Shi, Yuhong Hou |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2024 | Multimodal Learning for Embryo Viability Prediction in Clinical IVF
Junsik Kim 0001, Zhiyi Shi, Davin Jeong, Johannes Knittel, Helen Y. Yang, Yonghyun Song, Wanhua Li 0001, Yicong Li 0002, Dalit Ben-Yosef, Daniel Needleman, Hanspeter Pfister |
MICCAI (5) | 2 |
| 2024 | MoRA: LoRA Guided Multi-modal Disease Diagnosis with Missing Modality
Zhiyi Shi, Junsik Kim 0001, Wanhua Li 0001, Yicong Li 0002, Hanspeter Pfister |
MICCAI (3) | 1 |
| 2024 | A Robust Radio Frequency Fingerprint Extraction Method Based on Channel ReciprocityabstractRadio Frequency Fingerprint (RFF) identification is a promising technique for physical layer identification that can enhance wireless security. However, interference of wireless channel characteristics is a key challenge hindering its robustness. To solve this problem, we propose a channel-robust RFF extraction method. First, we design a challenge-response mechanism-based framework to satisfy the uplink and downlink channel reciprocity. Then, we propose a novel RFF extraction method named Quotient of the Estimated Channel State Information (QoECSI) that exploits channel reciprocity to eliminate channel effects. We implemented the QoECSI with the ESP32 development kits that support 2.4GHz Wi-Fi, Experimental results show that the extracted RFF features have high discrimination and long-term stability, and are robust to channel variations and noise. The accuracy rate is higher than 98% when the Signal-to-Noise Ratio (SNR) exceeds 25 dB. Specifically, in a Non-Line-of-Sight (NLOS) scenario with SNR = 35 dB, the average recognition accuracy of cross-validation is 98.56%. In a dynamic scenario where the terminal moves slowly indoors along a fixed route, the highest cross-time-validation accuracy is 98.57%. Bingshu Dong, Aiqun Hu, Jiabao Yu, Hongxia Chen, Zhiyi Shi |
WCNC | 5 |
| 2024 | Efficient Reinforcement Learning With the Novel N-Step Method and V-NetworkabstractThe application of reinforcement learning (RL) in artificial intelligence has become increasingly widespread. However, its drawbacks are also apparent, as it requires a large number of samples for support, making the enhancement of sample efficiency a research focus. To address this issue, we propose a novel N-step method. This method extends the horizon of the agent, enabling it to acquire more long-term effective information, thus resolving the issue of data inefficiency in RL. Additionally, this N-step method can reduce the estimation variance of Q-function, which is one of the factors contributing to estimation errors in Q-function estimation. Apart from high variance, estimation bias in Q-function estimation is another factor leading to estimation errors. To mitigate the estimation bias of Q-function, we design a regularization method based on the V-function, which has been underexplored. The combination of these two methods perfectly addresses the problems of low sample efficiency and inaccurate Q-function estimation in RL. Finally, extensive experiments conducted in discrete and continuous action spaces demonstrate that the proposed novel N-step method, when combined with classical deep Q-network, deep deterministic policy gradient, and TD3 algorithms, is effective, consistently outperforming the classical algorithms. Miaomiao Zhang 0001, Shuo Zhang 0023, Zhiyi Shi, Xiangyang Deng, Qi Wu 0003, Xin Xu 0001 |
IEEE Trans. Cybern. | 4 |
| 2023 | Brainnetformer: Decoding Brain Cognitive States with Spatial-Temporal Cross AttentionabstractLearning about the cognitive state of the brain has always been a popular topic. Based on the fact that fluctuations of brain signals and functional connectome (FC) relate to specific human behaviors, deep learning based methods have shown promising results on the prediction of such behaviors by analyzing biological signals. Existing methods either model from static perspectives or apply spatial-temporal graph convolution to extract dynamic properties. However, the static information and dynamic information can reflect global brain activities and local brain activities respectively. Thus, we propose BrainNetFormer to incorporate both static and dynamic properties for human behavior prediction. To be specific, a spatial cross attention module and a temporal cross attention module are introduced for information fusion. In addition, since a specific behavior of subjects can be decomposed into a series of subtasks, we introduce a sub-task regularization loss to assist in training and empower the model to recognize subtasks at each moment. Experiments on the HCP-Task dataset demonstrate the superior performance of the proposed model. Leheng Sheng, Wenhan Wang, Zhiyi Shi, Jichao Zhan, Youyong Kong |
ICASSP | 3 |
| 2023 | Strolling in Room-Scale VR: Hex-Core-MK1 Omnidirectional TreadmillabstractThe natural locomotion interface is critical to the development of many VR applications. For household VR applications, there are two basic requirements: natural immersive experience and minimized space occupation. The existing locomotion strategies generally do not simultaneously satisfy these two requirements well. This article presents a novel omnidirectional treadmill (ODT) system named Hex-Core-MK1 (HCMK1). By implementing two kinds of mirror-symmetrical spiral rollers to generate the omnidirectional velocity field, this proposed system is capable of providing real walking experiences with a full-degree of freedom in an area as small as 1.76 m$^{2}$, while delivering great advantages over several existing ODT systems in terms of weight, volume, latency and dynamic performance. Compared with the sizes of Infinadeck and HCP, the two best motor-driven ODTs so far, the 8 cm height of HCMK1 is only 20% of Infinadeck and 50% of HCP. In addition, HCMK1 is a lightweight device weighing only 110 kg, which provides possibilities for further expanding VR scenarios, such as terrain simulation. The system latency of HCMK1 is only 9ms. The experiments show that HCMK1 can deliver a starting acceleration of 16.00 m/s$^{2}$and a braking acceleration of 30.00 m/s$^{2}$. Chiyi Liu, Dazheng Fang, Zhiyi Shi, Yiye Wang, Kan-Jian Zhang, Haikun Wei |
IEEE Trans. Vis. Comput. Graph. | 6 |