EDBT 2026 Demo / reviewers in the wild / expert
Chao Jin 0007
dblp:19/4764-7
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2026
0009-0006-1355-4995ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 7 · 1 first-author · 7 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in ProductionabstractWe present MegaScale-MoE, a production system tailored for the efficient training of large-scale mixture-of-experts (MoE) models. MoE emerges as a promising architecture to scale large language models (LLMs) to unprecedented sizes, thereby enhancing model performance. However, existing MoE training systems experience a degradation in training efficiency, exacerbated by the escalating scale of MoE models and the continuous evolution of hardware. Chao Jin 0007, Ziheng Jiang, Zhihao Bai, Juncai Liu, Xiang Li 0067, Ningxin Zheng, Qi Huang 0001, Wen Heng, Yiyuan Ma, Wenlei Bao, Size Zheng 0001, Xuegui Zheng, Yanghua Peng, Haibin Lin, Xuanzhe Liu, Xin Jin 0008, Xin Liu 0086 |
EuroSys | 1 |
| 2026 | HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds
Chiheng Lou, Chao Jin 0007, Dapeng Nie, Xuanzhe Liu, Xin Jin 0008 |
NSDI | 3 |
| 2026 | RAGCache: Efficient Knowledge Caching for Retrieval-Augmented GenerationabstractRetrieval-Augmented Generation (RAG) has demonstrated substantial advancements in various natural language processing tasks by integrating the strengths of large language models (LLMs) and external knowledge databases. However, the retrieval step introduces long sequence generation and extra data dependency, resulting in long end-to-end latency. Our analysis benchmarks current RAG systems and reveals that, while the retrieval step poses performance challenges, it also offers optimization opportunities through its retrieval pattern and streaming search behavior. We propose RAGCache, a latency-optimized serving system tailored for RAG. RAGCache leverages the retrieval pattern to organize and cache the intermediate states of retrieved knowledge in a knowledge tree across the GPU and host memory hierarchy, reducing LLM generation time. RAGCache employs dynamic speculative pipelining to exploit the streaming search behavior, overlapping retrieval with LLM generation to minimize end-to-end latency. We implement RAGCache based on vLLM and Faiss, and evaluate it on both open-source and production datasets. Experimental results demonstrate that RAGCache reduces the time to first token (TTFT) by up to 4× and improves the throughput by up to 2.1× compared to vLLM integrated with Faiss. Chao Jin 0007, Xuanlin Jiang, Fangyue Liu, Shufan Liu, Xuanzhe Liu, Xin Jin 0008 |
ACM Trans. Comput. Syst. | 1 |
| 2025 | MegaScale-Infer: Efficient Mixture-of-Experts Model Serving with Disaggregated Expert ParallelismabstractMixture-of-Experts (MoE) showcases tremendous potential to scale large language models (LLMs) with enhanced performance and reduced computational complexity. However, its sparsely activated architecture shifts feed-forward networks (FFNs) from being compute-intensive to memory-intensive during inference, leading to substantially lower GPU utilization and increased operational costs. Ruidong Zhu, Ziheng Jiang, Chao Jin 0007, Cesar A. Stuardo, Huaping Zhou, Jianzhe Xiao, Lingjun Liu, Haibin Lin, Li-Wen Chang, Jianxi Ye, Xuanzhe Liu, Xin Jin 0008, Xin Liu 0086 |
SIGCOMM | 3 |
| 2025 | FaaSPR: Latency-Oriented Placement and Routing Optimization for Serverless Workflow ProcessingabstractWorkflow processing enhances the applicability of serverless computing while retaining the characteristics of fine-grained resource management and elastic scalability. However, current serverless platforms lack targeted optimization of placement and routing strategies for workflow processing, leading to high overheads due to inter-server data transmission, instance cold starts, and function request queuing. We propose FaaSPR, a serverless scheduling system that exploits placement and routing optimizations to minimize workflow processing latency. FaaSPR groups instances with potential data transmission and proportionally distributes groups with heterogeneous instances across multiple servers, taking into account resource constraints and historical placement traces. This method addresses the issues of poor scalability and frequent instance migrations in existing solutions. Utilizing a routing algorithm based on multi-stage linear programming, FaaSPR minimizes cross-server data transmission within and between instance groups while ensuring load balancing among instances. Experiments show that, compared to the state-of-the-art solution FaaSFlow, FaaSPR decreases the average and 99th percentile tail latency by up to 68.03% and 93.46%, respectively. Additionally, reducing workflow processing latency leads to up to 46.18% decrease in resource consumption for FaaS users. Yunshan Jia, Chao Jin 0007, Qing Li 0028, Xuanzhe Liu, Xin Jin 0008 |
IEEE Trans. Netw. | 2 |
| 2024 | Jolteon: Unleashing the Promise of Serverless for Serverless Workflows
Chao Jin 0007, Xin Jin 0008 |
NSDI | 2 |
| 2024 | Pyxis: Scheduling Mixed Tasks in Disaggregated DatacentersabstractDisaggregating compute from storage is an emerging trend in cloud computing. Effectively utilizing resources in both compute and storage pool is the key to high performance. The state-of-the-art scheduler provides optimal scheduling decisions for workloads with homogeneous tasks. However, cloud applications often generate a mix of tasks with diverse compute and IO characteristics, resulting in sub-optimal performance for existing solutions. We present Pyxis, a system that provides optimal scheduling decisions for mixed workloads in disaggregated datacenters with theoretical guarantees. Pyxis is capable of maximizing overall throughput while meeting latency SLOs. Pyxis decouples the scheduling of different tasks. Our insight is that the optimal solution has an “all-or-nothing” structure that can be captured by a singleturning pointin the spectrum of tasks. Based on task characteristics, the turning point partitions the tasks either all to storage nodes or all to compute nodes (none to storage nodes). We theoretically prove that the optimal solution has such a structure, and design an online algorithm with sub-second convergence. We implement a prototype of Pyxis. Experiments on CloudLab with various synthetic and application workloads show that Pyxis improves the throughput by 3–21× over the state-of-the-art solution. Chao Jin 0007, Mosharaf Chowdhury, Zhenming Liu, Xuanzhe Liu, Xin Jin 0008 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2023 | Fast, Approximate Vector Queries on Very Large Unstructured Datasets
Chao Jin 0007, Linpeng Tang, Xuanzhe Liu, Xin Jin 0008 |
NSDI | 2 |
| 2023 | Ditto: Efficient Serverless Analytics with Elastic ParallelismabstractServerless computing provides fine-grained resource elasticity for data analytics---a job can flexibly scale its resources for each stage, instead of sticking to a fixed pool of resources throughout its lifetime. Due to different data dependencies and different shuffling overheads caused by intra- and inter-server communication, the best degree of parallelism (DoP) for each stage varies based on runtime conditions. Chao Jin 0007, Xingyu Xiang, Songyun Zou, Gang Huang 0001, Xuanzhe Liu, Xin Jin 0008 |
SIGCOMM | 1 |
| 2022 | Melon: breaking the memory wall for resource-efficient on-device machine learningabstractOn-device learning is a promising technique for emerging privacy-preserving machine learning paradigms. However, through quantitative experiments, we find that commodity mobile devices cannot well support state-of-the-art DNN training with a large enough batch size, due to the limited local memory capacity. To fill the gap, we propose Melon, a memory-friendly on-device learning framework that enables the training tasks with large batch size beyond the physical memory capacity. Melon judiciously retrofits existing memory saving techniques to fit into resource-constrained mobile devices, i.e., recomputation and micro-batch. Melon further incorporates novel techniques to deal with the high memory fragmentation and memory adaptation. We implement and evaluate Melon with various typical DNN models on commodity mobile devices. The results show that Melon can achieve up to 4.33× larger batch size under the same memory budget. Given the same batch size, Melon achieves 1.89× on average (up to 4.01×) higher training throughput, and saves up to 49.43% energy compared to competitive alternatives. Furthermore, Melon reduces 78.59% computation on average in terms of memory budget adaptation. Qipeng Wang 0001, Mengwei Xu 0001, Chao Jin 0007, Xinran Dong, Jinliang Yuan, Xin Jin 0008, Gang Huang 0001, Yunxin Liu 0001, Xuanzhe Liu |
MobiSys | 3 |