Cheng Tan 0005

dblp:70/1533-5 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0002-1420-5125ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 7 · 2 first-author · 5 since 2021Systems, architecture and hardware · 6 · 5 since 2021Computer networks · 1 · 1 first-authorTheory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2025 PipeLLM: Fast and Confidential Large Language Model Services with Speculative Pipelined Encryption
abstract
Confidential computing on GPUs, like NVIDIA H100, mitigates the security risks of outsourced Large Language Models (LLMs) by implementing strong isolation and data encryption. Nonetheless, this encryption incurs a significant performance overhead, reaching up to 52.8% and 88.2% throughput drop when serving OPT-30B and OPT-66B, respectively. To address this challenge, we introduce PipeLLM, a user-transparent runtime system. PipeLLM removes the overhead by overlapping the encryption and GPU computation through pipelining-an idea inspired by the CPU instruction pipelining-thereby effectively concealing the latency increase caused by encryption. The primary technical challenge is that, unlike CPUs, the encryption module lacks prior knowledge of the specific data needing encryption until it is requested by the GPUs. To this end, we propose speculative pipelined encryption to predict the data requiring encryption by analyzing the serving patterns of LLMs. Further, we have developed an efficient, low-cost pipeline relinquishing approach for instances of incorrect predictions. Our experiments show that compared with vanilla systems without confidential computing (e.g., vLLM, PEFT, and FlexGen), PipeLLM incurs modest overhead ( < 19.6% in throughput) across various LLM sizes, from 13B to 175B. PipeLLM's source code is available at https://github.com/SJTU-IPADS/PipeLLM.
Yifan Tan, Cheng Tan 0005, Zeyu Mi, Haibo Chen 0001
ASPLOS (1)2
2025 TrainVerify: Equivalence-Based Verification for Distributed LLM Training
abstract
Training large language models (LLMs) at scale requires parallel execution across thousands of devices, incurring enormous computational costs. Yet, these costly distributed trainings are prone to correctness bugs, causing silent errors and potentially wasting millions of GPU hours. These bugs are challenging to expose through testing.
Yunchi Lu, Youshan Miao, Cheng Tan 0005, Peng Huang 0005, Xian Zhang 0001, Fan Yang 0024
SOSP3
2024 Scheduling Splittable Jobs on Configurable Machines
abstract
Motivated by modern architectures allowing for the partitioning of a GPU into hardware separated instances, we initiate the study of scheduling splittable jobs on configurable machines. We consider machines that can be configured into smaller instances, which we call blocks, in multiple ways, each of which is referred to as a configuration. We introduce the Configurable Machine Scheduling (cms) problem, where we are given n jobs and a set C of configurations. A schedule consists of a set of machines, each assigned some configuration in C with each block in the configuration assigned to process one job. The amount of a job’s demand that is satisfied by a block is given by an arbitrary function of the job and block. The objective is to construct a schedule using as few machines as possible. We provide a tight logarithmic factor approximation algorithm for this problem in the general setting, a factor (3 + ε) approximation algorithm for arbitrary ε > 0 when there are O(1) input configurations, and a polynomial time approximation scheme when both the number and size of configurations are O(1). Finally, we utilize a technique for finding conic integer combinations in fixed dimension to develop an optimal polynomial time algorithm in the case with O(1) jobs, O(1) blocks, and every configuration up to a given size.
Matthew M. Casey, Rajmohan Rajaraman, David Stalfa, Cheng Tan 0005
APPROX/RANDOM4
2024 Optimizing GPU Sharing for Container-Based DNN Serving with Multi-Instance GPUs
abstract
The trade-off of throughput versus latency is important in serving deep neural networks (DNNs) on GPUs. A hardware feature---Multi-Instance GPU (MIG)---provides an additional dimension to this trade-off. This paper studies the GPU sharing problem for serving DNNs with MIG. We propose a system Jormungandr that allocates MIG-enabled GPUs to optimize a user-define utility function, and places DNN serving containers on heterogeneous MIG instances. Jormungandr introduces a utility-first search that yields efficient solutions in practice. Additionally, it presents a robust transition protocol that guarantees a seamless switch between two GPU-cluster configurations, minimizing throughput fluctuations. We implement Jormungandr on Kubernetes. Our experiments show that Jormungandr provides near-optimal solutions, using < 6% more GPUs than the optimal solutions.
Xinpeng Wei, Cheng Tan 0005
SYSTOR3
2023 NNSmith: Generating Diverse and Valid Test Cases for Deep Learning Compilers
abstract
Deep-learning (DL) compilers such as TVM and TensorRT are increasingly being used to optimize deep neural network (DNN) models to meet performance, resource utilization and other requirements. Bugs in these compilers can result in models whose semantics differ from the original ones, producing incorrect results that corrupt the correctness of downstream applications. However, finding bugs in these compilers is challenging due to their complexity. In this work, we propose a new fuzz testing approach for finding bugs in deep-learning compilers. Our core approach consists of (i) generating diverse yet valid DNN test models that can exercise a large part of the compiler's transformation logic using light-weight operator specifications; (ii) performing gradient-based search to find model inputs that avoid any floating-point exceptional values during model execution, reducing the chance of missed bugs or false alarms; and (iii) using differential testing to identify bugs. We implemented this approach in NNSmith which has found 72 new bugs for TVM, TensorRT, ONNXRuntime, and PyTorch to date. Of these 58 have been confirmed and 51 have been fixed by their respective project maintainers.
Jiawei Liu 0004, Jinkun Lin, Fabian Ruffy, Cheng Tan 0005, Jinyang Li 0001, Aurojit Panda, Lingming Zhang 0001
ASPLOS (2)4
2023 Viper: A Fast Snapshot Isolation Checker
abstract
Snapshot isolation (SI) is supported by most commercial databases and is widely used by applications. However, checking SI today---given a set of transactions, checking if they obey SI---is either slow or gives up soundness.
Jian Zhang 0102, Ye Ji 0003, Shuai Mu 0001, Cheng Tan 0005
EuroSys4
2023 Encrypted Databases Made Secure Yet Maintainable
Cheng Tan 0005, Huorong Li, Sheng Wang 0011, Zeyu Mi, Yubin Xia, Feifei Li 0001, Haibo Chen 0001
OSDI4
2023 Predicting GPU Failures With High Precision Under Deep Learning Workloads
abstract
Graphics processing units (GPUs) are the de facto standard for processing deep learning (DL) tasks. In large-scale GPU clusters, GPU failures are inevitable and may cause severe consequences. For example, GPU failures disrupt distributed training, crash inference services, and result in service level agreement violations. In this paper, we study the problem of predicting GPU failures using machine learning (ML) models to mitigate their damages.
Heting Liu, Cheng Tan 0005, Rongqiu Yang, Guohong Cao, Zherui Liu, Chuanxiong Guo
SYSTOR3
2021 Bringing Decentralized Search to Decentralized Services
Jinhao Zhu, Tianxu Zhang, Cheng Tan 0005, Yubin Xia, Sebastian Angel, Haibo Chen 0001
OSDI4
2020 Cobra: Making Transactional Key-Value Stores Verifiably Serializable
Cheng Tan 0005, Changgeng Zhao, Shuai Mu 0001, Michael Walfish
OSDI1
2019 NetBouncer: Active Device and Link Failure Localization in Data Center Networks
Cheng Tan 0005, Ze Jin, Chuanxiong Guo, Tianrong Zhang, Karl Deng, Dongming Bi
NSDI1
2017 The Efficient Server Audit Problem, Deduplicated Re-execution, and the Web
abstract
You put a program on a concurrent server, but you don't trust the server; later, you get a trace of the actual requests that the server received from its clients and the responses that it delivered. You separately get logs from the server; these are untrusted. How can you use the logs to efficiently verify that the responses were derived from running the program on the requests? This is the Efficient Server Audit Problem, which abstracts real-world scenarios, including running a web application on an untrusted provider. We give a solution based on several new techniques, including simultaneous replay and efficient verification of concurrent executions. We implement the solution for PHP web applications. For several applications, our verifier achieves 5.6-10.9x speedup versus simply re-executing, with <10% overhead for the server.
Cheng Tan 0005, Lingfan Yu, Joshua B. Leners, Michael Walfish
SOSP1
2015 TinMan: eliminating confidential mobile data exposure with security oriented offloading
abstract
The wide adoption of smart devices has stimulated a fast shift of security-critical data from desktop to mobile devices. However, recurrent device theft and loss expose mobile devices to various security threats and even physical attacks. This paper presents TinMan, a system that protects confidential data such as web site password and credit card number (we use the term cor to represent these data, which is short for Confidential Record) from being leaked or abused even under device theft. TinMan separates accesses of cor from the rest of the functionalities of an app, by introducing a trusted node to store cor and offloading any code from a mobile device to the trusted node to access cor. This completely eliminates the exposure of cor on the mobile devices. The key challenges to TinMan include deciding when and how to efficiently and transparently offload execution; TinMan addresses these challenges with security-oriented offloading with a low-overhead tainting scheme called asymmetric tainting to track accesses to cor to trigger offloading, as well as transparent SSL session injection and TCP pay-load replacement to offload accesses to cor. We have implemented a prototype of TinMan based on Android and demonstrated how TinMan protects the information of user's bank account and credit card number without modifying the apps. Evaluation results also show that TinMan incurs only a small amount of performance and power overhead.
Yubin Xia, Cheng Tan 0005, Haibing Guan, Binyu Zang, Haibo Chen 0001
EuroSys3