EDBT 2026 Demo / reviewers in the wild / expert
Suyi Li 0002
dblp:98/7732-2
· DBLP profile ↗
10ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0003-3921-9037ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FlashPS: Efficient Generative Image Editing with Mask-aware Caching and SchedulingabstractGenerative image editing using diffusion models has become a prevalent application in today's AI cloud services. In production environments, image editing typically involves a mask that specifies the regions of an image template to be edited. The use of mask provides direct control over the editing process and introduces sparsity in the model inference. In this paper, we present FlashPS, a system that efficiently serves image editing requests. The key insight behind FlashPS is that image editing only modifies the masked regions of image templates, while preserving the original content in the unmasked areas. Driven by this insight, FlashPS judiciously skips redundant computations associated with the unmask areas by reusing cached intermediate activations from previous inferences. To mitigate the high cache loading overhead, FlashPS employs a bubble-free pipeline scheme that overlaps computation with cache loading. Additionally, to reduce queuing latency in online serving while improving the GPU utilization, FlashPS proposes a novel continuous batching strategy for diffusion model serving, allowing newly arrived requests to join the running batch in just one step of denoising computation, without waiting for the entire batch to complete. As heterogenous masks induce imbalanced load, FlashPS also develops a load balancing strategy that takes into account the loads of both computation and cache loading. Collectively, FlashPS outperforms state-of-the-art diffusion serving systems for image editing, achieving up to 3× higher throughput and reducing average request latency by up to 14.7× while ensuring image quality. Xiaoxiao Jiang, Suyi Li 0002, Lingyun Yang, Tianyu Feng, Zhipeng Di, Weiyi Lu, Guoxuan Zhu, Xiu Lin, Yinghao Yu, Tao Lan, Lin Qu, Liping Zhang 0013, Wei Wang 0030 |
EuroSys | 2 |
| 2025 | Toppings: CPU-Assisted, Rank-Aware Adapter Serving for LLM Inference
Suyi Li 0002, Hanfeng Lu, Tianyuan Wu, Minchen Yu, Qizhen Weng 0001, Xusheng Chen, Yizhou Shan, Binhang Yuan, Wei Wang 0030 |
USENIX ATC | 1 |
| 2025 | Katz: Efficient Workflow Serving for Diffusion Models with Many Adapters
Suyi Li 0002, Lingyun Yang, Xiaoxiao Jiang, Hanfeng Lu, Dakai An, Zhipeng Di, Weiyi Lu, Yinghao Yu, Tao Lan, Lin Qu, Liping Zhang 0013, Wei Wang 0030 |
USENIX ATC | 1 |
| 2023 | Golgi: Performance-Aware, Resource-Efficient Function Scheduling for Serverless ComputingabstractThis paper introduces Golgi, a novel scheduling system designed for serverless functions, with the goal of minimizing resource provisioning costs while meeting the function latency requirements. To achieve this, Golgi judiciously over-commits functions based on their past resource usage. To ensure overcommitment does not cause significant performance degradation, Golgi identifies nine low-level metrics to capture the runtime performance of functions, encompassing factors like request load, resource allocation, and contention on shared resources. These metrics enable accurate prediction of function performance using the Mondrian Forest, a classification model that is continuously updated in real-time for optimal accuracy without extensive offline training. Golgi employs a conservative exploration-exploitation strategy for request routing. By default, it routes requests to non-overcommitted instances to ensure satisfactory performance. However, it actively explores opportunities for using more resource-efficient overcommitted instances, while maintaining the specified latency SLOs. Golgi also performs vertical scaling to dynamically adjust the concurrency of overcommitted instances, maximizing request throughput and enhancing system robustness to prediction errors. We have prototyped Golgi and evaluated it in both EC2 cluster and a small production cluster. The results show that Golgi can meet the SLOs while reducing the resource provisioning cost by 42% (30%) in EC2 cluster (our production cluster). Suyi Li 0002, Wei Wang 0030, Guangzhen Chen, Daohe Lu |
SoCC | 1 |
| 2022 | Owl: performance-aware scheduling for resource-efficient function-as-a-service cloudabstractThis work documents our experience of improving the scheduler in Alibaba Function Compute, a public FaaS platform. It commences with our observation that memory and CPU are under-utilized in most FaaS sandboxes. A natural solution is to overcommit VM resources when allocating sandboxes, whereas the ensuing contention may cause performance degradation and compromise user experience. To complicate matters, the degradation in FaaS can arise from external factors, such as failed dependencies of user functions. Huangshi Tian, Suyi Li 0002, Wei Wang 0030, Tianlong Wu |
SoCC | 2 |
| 2021 | George: Learning to Place Long-Lived Containers in Large Clusters with Operation ConstraintsabstractOnline cloud services are widely deployed as Long-Running Applications (LRAs) hosted in containers. Placing LRA containers turns out to be particularly challenging due to the complex interference between co-located containers and the operation constraints in production clusters such as fault tolerance, disaster avoidance and incremental deployment. Existing schedulers typically provide APIs for operators to manually specify the container scheduling requirements and offer only qualitative scheduling guidelines for container placement. Such schedulers, do not perform well in terms of both performance and scale, while also requiring manual intervention. Suyi Li 0002, Wei Wang 0030, Yinghao Yu, Bo Li 0001 |
SoCC | 1 |
| 2020 | BatchCrypt: Efficient Homomorphic Encryption for Cross-Silo Federated Learning
Chengliang Zhang, Suyi Li 0002, Junzhe Xia, Wei Wang 0030, Feng Yan 0001, Yang Liu 0165 |
USENIX ATC | 2 |
| 2019 | Multi-News: A Large-Scale Multi-Document Summarization Dataset and Abstractive Hierarchical ModelabstractAutomatic generation of summaries from multiple news articles is a valuable tool as the number of online publications grows rapidly.Single document summarization (SDS) systems have benefited from advances in neural encoder-decoder model thanks to the availability of large datasets.However, multidocument summarization (MDS) of news articles has been limited to datasets of a couple of hundred examples.In this paper, we introduce Multi-News, the first large-scale MDS news dataset.Additionally, we propose an end-to-end model which incorporates a traditional extractive summarization model with a standard SDS model and achieves competitive results on MDS datasets.We benchmark several methods on Multi-News and release our data and code in hope that this work will promote advances in summarization in the multidocument setting 1 . Alexander R. Fabbri, Irene Li, Tianwei She, Suyi Li 0002, Dragomir R. Radev |
ACL (1) | 4 |
| 2019 | SParC: Cross-Domain Semantic Parsing in ContextabstractTao Yu, Rui Zhang, Michihiro Yasunaga, Yi Chern Tan, Xi Victoria Lin, Suyi Li, Heyang Er, Irene Li, Bo Pang, Tao Chen, Emily Ji, Shreya Dixit, David Proctor, Sungrok Shim, Jonathan Kraft, Vincent Zhang, Caiming Xiong, Richard Socher, Dragomir Radev. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. Tao Yu 0009, Rui Zhang 0037, Michihiro Yasunaga, Yi Chern Tan, Xi Victoria Lin, Suyi Li 0002, Heyang Er, Irene Li, Bo Pang 0004, Emily Ji, Shreya Dixit, David Proctor, Sungrok Shim, Jonathan Kraft, Caiming Xiong, Richard Socher, Dragomir R. Radev |
ACL (1) | 6 |
| 2019 | CoSQL: A Conversational Text-to-SQL Challenge Towards Cross-Domain Natural Language Interfaces to DatabasesabstractTao Yu, Rui Zhang, Heyang Er, Suyi Li, Eric Xue, Bo Pang, Xi Victoria Lin, Yi Chern Tan, Tianze Shi, Zihan Li, Youxuan Jiang, Michihiro Yasunaga, Sungrok Shim, Tao Chen, Alexander Fabbri, Zifan Li, Luyao Chen, Yuwen Zhang, Shreya Dixit, Vincent Zhang, Caiming Xiong, Richard Socher, Walter Lasecki, Dragomir Radev. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Tao Yu 0009, Rui Zhang 0037, Heyang Er, Suyi Li 0002, Eric Xue 0001, Bo Pang 0004, Xi Victoria Lin, Yi Chern Tan, Tianze Shi, Youxuan Jiang, Michihiro Yasunaga, Sungrok Shim, Alexander R. Fabbri, Zifan Li, Shreya Dixit, Caiming Xiong, Richard Socher, Walter S. Lasecki, Dragomir R. Radev |
EMNLP/IJCNLP (1) | 4 |