Suyi Li 0002

dblp:98/7732-2 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0003-3921-9037ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 3
YearPublicationVenuePosition
2026 FlashPS: Efficient Generative Image Editing with Mask-aware Caching and Scheduling
abstract
Generative image editing using diffusion models has become a prevalent application in today's AI cloud services. In production environments, image editing typically involves a mask that specifies the regions of an image template to be edited. The use of mask provides direct control over the editing process and introduces sparsity in the model inference. In this paper, we present FlashPS, a system that efficiently serves image editing requests. The key insight behind FlashPS is that image editing only modifies the masked regions of image templates, while preserving the original content in the unmasked areas. Driven by this insight, FlashPS judiciously skips redundant computations associated with the unmask areas by reusing cached intermediate activations from previous inferences. To mitigate the high cache loading overhead, FlashPS employs a bubble-free pipeline scheme that overlaps computation with cache loading. Additionally, to reduce queuing latency in online serving while improving the GPU utilization, FlashPS proposes a novel continuous batching strategy for diffusion model serving, allowing newly arrived requests to join the running batch in just one step of denoising computation, without waiting for the entire batch to complete. As heterogenous masks induce imbalanced load, FlashPS also develops a load balancing strategy that takes into account the loads of both computation and cache loading. Collectively, FlashPS outperforms state-of-the-art diffusion serving systems for image editing, achieving up to 3× higher throughput and reducing average request latency by up to 14.7× while ensuring image quality.
Xiaoxiao Jiang, Suyi Li 0002, Lingyun Yang, Tianyu Feng, Zhipeng Di, Weiyi Lu, Guoxuan Zhu, Xiu Lin, Yinghao Yu, Tao Lan, Lin Qu, Liping Zhang 0013, Wei Wang 0030
EuroSys2
2025 Toppings: CPU-Assisted, Rank-Aware Adapter Serving for LLM Inference
Suyi Li 0002, Hanfeng Lu, Tianyuan Wu, Minchen Yu, Qizhen Weng 0001, Xusheng Chen, Yizhou Shan, Binhang Yuan, Wei Wang 0030
USENIX ATC1
2025 Katz: Efficient Workflow Serving for Diffusion Models with Many Adapters
Suyi Li 0002, Lingyun Yang, Xiaoxiao Jiang, Hanfeng Lu, Dakai An, Zhipeng Di, Weiyi Lu, Yinghao Yu, Tao Lan, Lin Qu, Liping Zhang 0013, Wei Wang 0030
USENIX ATC1
2023 Golgi: Performance-Aware, Resource-Efficient Function Scheduling for Serverless Computing
abstract
This paper introduces Golgi, a novel scheduling system designed for serverless functions, with the goal of minimizing resource provisioning costs while meeting the function latency requirements. To achieve this, Golgi judiciously over-commits functions based on their past resource usage. To ensure overcommitment does not cause significant performance degradation, Golgi identifies nine low-level metrics to capture the runtime performance of functions, encompassing factors like request load, resource allocation, and contention on shared resources. These metrics enable accurate prediction of function performance using the Mondrian Forest, a classification model that is continuously updated in real-time for optimal accuracy without extensive offline training. Golgi employs a conservative exploration-exploitation strategy for request routing. By default, it routes requests to non-overcommitted instances to ensure satisfactory performance. However, it actively explores opportunities for using more resource-efficient overcommitted instances, while maintaining the specified latency SLOs. Golgi also performs vertical scaling to dynamically adjust the concurrency of overcommitted instances, maximizing request throughput and enhancing system robustness to prediction errors. We have prototyped Golgi and evaluated it in both EC2 cluster and a small production cluster. The results show that Golgi can meet the SLOs while reducing the resource provisioning cost by 42% (30%) in EC2 cluster (our production cluster).
Suyi Li 0002, Wei Wang 0030, Guangzhen Chen, Daohe Lu
SoCC1
2022 Owl: performance-aware scheduling for resource-efficient function-as-a-service cloud
abstract
This work documents our experience of improving the scheduler in Alibaba Function Compute, a public FaaS platform. It commences with our observation that memory and CPU are under-utilized in most FaaS sandboxes. A natural solution is to overcommit VM resources when allocating sandboxes, whereas the ensuing contention may cause performance degradation and compromise user experience. To complicate matters, the degradation in FaaS can arise from external factors, such as failed dependencies of user functions.
Huangshi Tian, Suyi Li 0002, Wei Wang 0030, Tianlong Wu
SoCC2
2021 George: Learning to Place Long-Lived Containers in Large Clusters with Operation Constraints
abstract
Online cloud services are widely deployed as Long-Running Applications (LRAs) hosted in containers. Placing LRA containers turns out to be particularly challenging due to the complex interference between co-located containers and the operation constraints in production clusters such as fault tolerance, disaster avoidance and incremental deployment. Existing schedulers typically provide APIs for operators to manually specify the container scheduling requirements and offer only qualitative scheduling guidelines for container placement. Such schedulers, do not perform well in terms of both performance and scale, while also requiring manual intervention.
Suyi Li 0002, Wei Wang 0030, Yinghao Yu, Bo Li 0001
SoCC1
2020 BatchCrypt: Efficient Homomorphic Encryption for Cross-Silo Federated Learning
Chengliang Zhang, Suyi Li 0002, Junzhe Xia, Wei Wang 0030, Feng Yan 0001, Yang Liu 0165
USENIX ATC2
2019 Multi-News: A Large-Scale Multi-Document Summarization Dataset and Abstractive Hierarchical Model
abstract
Automatic generation of summaries from multiple news articles is a valuable tool as the number of online publications grows rapidly.Single document summarization (SDS) systems have benefited from advances in neural encoder-decoder model thanks to the availability of large datasets.However, multidocument summarization (MDS) of news articles has been limited to datasets of a couple of hundred examples.In this paper, we introduce Multi-News, the first large-scale MDS news dataset.Additionally, we propose an end-to-end model which incorporates a traditional extractive summarization model with a standard SDS model and achieves competitive results on MDS datasets.We benchmark several methods on Multi-News and release our data and code in hope that this work will promote advances in summarization in the multidocument setting 1 .
Alexander R. Fabbri, Irene Li, Tianwei She, Suyi Li 0002, Dragomir R. Radev
ACL (1)4
2019 SParC: Cross-Domain Semantic Parsing in Context
abstract
Tao Yu, Rui Zhang, Michihiro Yasunaga, Yi Chern Tan, Xi Victoria Lin, Suyi Li, Heyang Er, Irene Li, Bo Pang, Tao Chen, Emily Ji, Shreya Dixit, David Proctor, Sungrok Shim, Jonathan Kraft, Vincent Zhang, Caiming Xiong, Richard Socher, Dragomir Radev. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Tao Yu 0009, Rui Zhang 0037, Michihiro Yasunaga, Yi Chern Tan, Xi Victoria Lin, Suyi Li 0002, Heyang Er, Irene Li, Bo Pang 0004, Emily Ji, Shreya Dixit, David Proctor, Sungrok Shim, Jonathan Kraft, Caiming Xiong, Richard Socher, Dragomir R. Radev
ACL (1)6
2019 CoSQL: A Conversational Text-to-SQL Challenge Towards Cross-Domain Natural Language Interfaces to Databases
abstract
Tao Yu, Rui Zhang, Heyang Er, Suyi Li, Eric Xue, Bo Pang, Xi Victoria Lin, Yi Chern Tan, Tianze Shi, Zihan Li, Youxuan Jiang, Michihiro Yasunaga, Sungrok Shim, Tao Chen, Alexander Fabbri, Zifan Li, Luyao Chen, Yuwen Zhang, Shreya Dixit, Vincent Zhang, Caiming Xiong, Richard Socher, Walter Lasecki, Dragomir Radev. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Tao Yu 0009, Rui Zhang 0037, Heyang Er, Suyi Li 0002, Eric Xue 0001, Bo Pang 0004, Xi Victoria Lin, Yi Chern Tan, Tianze Shi, Youxuan Jiang, Michihiro Yasunaga, Sungrok Shim, Alexander R. Fabbri, Zifan Li, Shreya Dixit, Caiming Xiong, Richard Socher, Walter S. Lasecki, Dragomir R. Radev
EMNLP/IJCNLP (1)4