Yifan Tan

dblp:196/9159 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
3since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 PipeLLM: Fast and Confidential Large Language Model Services with Speculative Pipelined Encryption
abstract
Confidential computing on GPUs, like NVIDIA H100, mitigates the security risks of outsourced Large Language Models (LLMs) by implementing strong isolation and data encryption. Nonetheless, this encryption incurs a significant performance overhead, reaching up to 52.8% and 88.2% throughput drop when serving OPT-30B and OPT-66B, respectively. To address this challenge, we introduce PipeLLM, a user-transparent runtime system. PipeLLM removes the overhead by overlapping the encryption and GPU computation through pipelining-an idea inspired by the CPU instruction pipelining-thereby effectively concealing the latency increase caused by encryption. The primary technical challenge is that, unlike CPUs, the encryption module lacks prior knowledge of the specific data needing encryption until it is requested by the GPUs. To this end, we propose speculative pipelined encryption to predict the data requiring encryption by analyzing the serving patterns of LLMs. Further, we have developed an efficient, low-cost pipeline relinquishing approach for instances of incorrect predictions. Our experiments show that compared with vanilla systems without confidential computing (e.g., vLLM, PEFT, and FlexGen), PipeLLM incurs modest overhead ( < 19.6% in throughput) across various LLM sizes, from 13B to 175B. PipeLLM's source code is available at https://github.com/SJTU-IPADS/PipeLLM.
Yifan Tan, Cheng Tan 0005, Zeyu Mi, Haibo Chen 0001
ASPLOS (1)1
2024 Performance Analysis and Optimization of Nvidia H100 Confidential Computing for AI Workloads
abstract
NVIDIA’s H100 Confidential Computing (CC) counters the security hazards inherent in cloud AI workloads. It enforces data encryption to achieve data confidentiality, which leads to substantial throughput reductions as high as 93% in various AI workloads (such as TensorRT, PEFT and vLLM). Confronting this substantial overhead issue, we first delve into the underlying causes through meticulous analysis. This groundwork enables us to devise an innovative runtime system that operates seamlessly in the background, completely transparent to end-users. The cornerstone of our system lies in leveraging multiple encryption workers. Experiments demonstrate that our solution effectively reduces throughput drop to less than 28.1%.
Yifan Tan, Zeyu Mi
ISPA1
2023 Bifrost: Analysis and Optimization of Network I/O Tax in Confidential Virtual Machines
Dingji Li, Zeyu Mi, Chenhui Ji, Yifan Tan, Binyu Zang, Haibing Guan, Haibo Chen 0001
USENIX ATC4
2020 Learning Latent Semantic Attributes for Zero-Shot Object Detection
abstract
Zero-shot Object Detection (ZSD) aims to locate and classify instances of unseen categories. Existing methods focus on learning the mapping from visual space to semantic space, while the learning of discriminative representations for ZSD has not gained enough attention. In this paper, we demonstrate the necessity to learn discriminative semantic representations for ZSD, and propose a new end-to-end framework for this task. Our framework is able to learn discriminative semantic representations in an augmented space introduced for both user-defined and latent attributes, and refine the user-defined attributes with the help of unseen and external classes. The proposed method is extensively evaluated on two challenging ZSD datasets, and the experimental results show that our method significantly outperforms several existing methods.
Lu Zhang 0060, Yifan Tan, Shuigeng Zhou
ICTAI3
2017 On the Practical Design of a High Power Density SiC Single-Phase Uninterrupted Power Supply System
abstract
This paper proposes a high power density SiC single-phase system potential for uninterrupted power supply applications. To get the high power density, the semiconductors, packaging, circuit topology, and thermal design are synthetically considered. To increase the switching frequency and reduce the size of the passive components, the SiC MOSFETs and diodes are chosen; to minimize the parasitic inductances and eliminate the snubbers, the SiC bare dies are packaged as the half-bridge (HB) modules; to remove the bulky dc-link capacitors, the full-bridge inverter and the active power filter are designed, and they are structured by using the fabricated SiC HB modules; and finally to dissipate the heat from such a compact enclosure in the cost-efficient way, the heat sink of the modules and the forced air cool system are well designed, and the thermal 3-D finite-element analysis model is built to survey the best cooling configuration. A 2-kVA prototype is built and tested, and the power density of the system is up to 58 W/in3and the maximal efficiency is up to 98.3%.
Yu Chen 0025, Yifan Tan, Jianming Fang, Yong Kang
IEEE Trans. Ind. Informatics3