VLDB 2026 Research / reviewers in the wild / expert
Xin Jin 0008
dblp:68/3340-8
· DBLP profile ↗
108ranked-venue papers
11as first author
77since 2021 · last 2026
0000-0001-8741-5847ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 54 · 6 first-author · 38 since 2021Software engineering, systems software and programming languages · 24 · 1 first-author · 19 since 2021Systems, architecture and hardware · 17 · 13 since 2021Databases, data management, data science and information retrieval · 9 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Security and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in ProductionabstractWe present MegaScale-MoE, a production system tailored for the efficient training of large-scale mixture-of-experts (MoE) models. MoE emerges as a promising architecture to scale large language models (LLMs) to unprecedented sizes, thereby enhancing model performance. However, existing MoE training systems experience a degradation in training efficiency, exacerbated by the escalating scale of MoE models and the continuous evolution of hardware. Chao Jin 0007, Ziheng Jiang, Zhihao Bai, Juncai Liu, Xiang Li 0067, Ningxin Zheng, Qi Huang 0001, Wen Heng, Yiyuan Ma, Wenlei Bao, Size Zheng 0001, Xuegui Zheng, Yanghua Peng, Haibin Lin, Xuanzhe Liu, Xin Jin 0008, Xin Liu 0086 |
EuroSys | 19 |
| 2026 | Serverless Replication of Object Storage across Multi-Vendor Clouds and Regions
Junyi Shu, Gang Huang 0001, Hong Mei 0001, Xuanzhe Liu, Xin Jin 0008 |
EuroSys | 6 |
| 2026 | HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds
Chiheng Lou, Chao Jin 0007, Dapeng Nie, Xuanzhe Liu, Xin Jin 0008 |
NSDI | 8 |
| 2026 | A Composable Emulation Framework for Whitebox Switches
Congcong Miao, Xianneng Zou, Chuwen Zhang, Qihang Liu, Zhijie Yan, Yanke Zhang, Yong Jiang 0001, Qiao Xiang, Xin Jin 0008, Zili Meng, Ang Chen 0001 |
NSDI | 10 |
| 2026 | KUBEDIRECT: Unleashing the Full Power of the Cluster Manager for Serverless Computing
Zhiquan Zhang, Xuanzhe Liu, Xin Jin 0008 |
NSDI | 4 |
| 2026 | FastServe: Iteration-Level Preemptive Scheduling for Large Language Model Inference
Bingyang Wu, Yinmin Zhong, Fangyue Liu, Yuanhang Sun, Gang Huang 0001, Xuanzhe Liu, Xin Jin 0008 |
NSDI | 9 |
| 2026 | Iceberg: Automated Verification of DNS Authoritative Engines via Just-in-Time Summarization
Yuxing Xiang, Rilin Huang, Naiqian Zheng, Xin Jin 0008 |
NSDI | 4 |
| 2026 | ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production
Yuxing Xiang, Xue Li 0024, Kun Qian 0004, Yan Zhang 0117, Wenyuan Yu, Ennan Zhai, Xin Jin 0008, Jingren Zhou 0001 |
NSDI | 7 |
| 2026 | DistRS: Disaggregated Reward Service for RLVR with Batch-Level Constraint
Ruidong Zhu, Mingcong Han, Yinmin Zhong, Wencong Xiao, Xuanzhe Liu, Xin Jin 0008 |
NSDI | 6 |
| 2026 | DualPath: Accelerating Agentic LLM Inference by Harvesting Disaggregated KV-Cache Storage I/OabstractThe performance of multi-turn, agentic LLM inference is increasingly dominated by KV-Cache storage I/O rather than computation. In prevalent disaggregated architectures, loading the massive KV-Cache from external storage creates a fundamental imbalance: storage NICs on prefill engines become bandwidth-saturated, while those on decoding engines remain idle. This asymmetry severely constrains overall system throughput. Yongtong Wu, Shaoyuan Chen, Rilin Huang, Yixuan Tan, Yinmin Zhong, Xin Jin 0008, Panpan Huang |
SIGCOMM | 7 |
| 2026 | Epiphron: Resource-Efficient Distributed Key-Value StorageabstractIn-memory key-value storage necessitates a substantial quantity of computation and storage resources for both performance and scalability, thereby diminishing the resources available for user applications. The emergence of programmable network hardware, including SmartNICs and programmable switches, provides the opportunity to offload operations from server CPUs. We present Epiphron, a novel distributed in-memory key-value store architecture that co-designs with off-path SmartNICs and programmable switches. Facing the limited performance of off-path SmartNICs, Epiphron successfully achieves high resource efficiency while keeping load balancing and fault tolerance by$(i)$hybridizing erasure coding with replication in storage management,$(ii)$accelerating read operations with a new data plane design (conflict detection and RDMA-compatible forwarding) on programmable switches,$(iii)$employing a network protocol extended from one-sided RDMA. We evaluate Epiphron on Barefoot Tofino switches, NVIDIA BlueField-2 SmartNICs, and commodity servers. The experimental results demonstrate that compared to existing solutions, Epiphron improves throughput by up to 2.2× and consumes 47% less memory while completely bypassing server CPUs. Ruidong Zhu, Bingyang Wu, Xin Yao 0008, Renhai Chen, Gong Zhang 0001, Xuanzhe Liu, Xin Jin 0008 |
IEEE Trans. Knowl. Data Eng. | 8 |
| 2026 | RAGCache: Efficient Knowledge Caching for Retrieval-Augmented GenerationabstractRetrieval-Augmented Generation (RAG) has demonstrated substantial advancements in various natural language processing tasks by integrating the strengths of large language models (LLMs) and external knowledge databases. However, the retrieval step introduces long sequence generation and extra data dependency, resulting in long end-to-end latency. Our analysis benchmarks current RAG systems and reveals that, while the retrieval step poses performance challenges, it also offers optimization opportunities through its retrieval pattern and streaming search behavior. We propose RAGCache, a latency-optimized serving system tailored for RAG. RAGCache leverages the retrieval pattern to organize and cache the intermediate states of retrieved knowledge in a knowledge tree across the GPU and host memory hierarchy, reducing LLM generation time. RAGCache employs dynamic speculative pipelining to exploit the streaming search behavior, overlapping retrieval with LLM generation to minimize end-to-end latency. We implement RAGCache based on vLLM and Faiss, and evaluate it on both open-source and production datasets. Experimental results demonstrate that RAGCache reduces the time to first token (TTFT) by up to 4× and improves the throughput by up to 2.1× compared to vLLM integrated with Faiss. Chao Jin 0007, Xuanlin Jiang, Fangyue Liu, Shufan Liu, Xuanzhe Liu, Xin Jin 0008 |
ACM Trans. Comput. Syst. | 7 |
| 2025 | CloudChurn: Optimizing Enterprise Customer Churn Prediction in Cloud Services for Huawei Cloud
Hengyu Ye, Yulong Song, Zhipeng Bian, Xiaofeng Gao 0001, Guihai Chen, Xin Jin 0008, Zhenli Sheng |
DASFAA (6) | 6 |
| 2025 | OmniBal: Towards Fast Instruction-Tuning for Vision-Language Models via Omniverse Computation BalanceabstractVision-language instruction-tuning models have recently achieved significant performance improvements. In this work, we discover that large-scale 3D parallel training on those models leads to an imbalanced computation load across different devices. The vision and language parts are inherently heterogeneous: their data distribution and model architecture differ significantly, which affects distributed training efficiency. To address this issue, we rebalance the computational load from data, model, and memory perspectives, achieving more balanced computation across devices. Specifically, for the data, instances are grouped into new balanced mini-batches within and across devices. A search-based method is employed for the model to achieve a more balanced partitioning. For memory optimization, we adaptively adjust the re-computation strategy for each partition to utilize the available memory fully. These three perspectives are not independent but are closely connected, forming an omniverse balanced training framework. Extensive experiments are conducted to validate the effectiveness of our method. Compared with the open-source training code of InternVL-Chat, training time is reduced greatly, achieving about 1.8$\times$ speed-up. Our method’s efficacy and generalizability are further validated across various models and datasets. Codes will be released at https://github.com/ModelTC/OmniBal. Yongqiang Yao, Jingru Tan, Feizhao Zhang, Yazhe Niu, Xin Jin 0008, Bo Li 0126, Pengfei Liu 0003, Ruihao Gong, Dahua Lin, Ningyi Xu |
ICML | 6 |
| 2025 | MeshTest: End-to-End Testing for Service Mesh Traffic Management
Naiqian Zheng, Tianshuo Qiao, Xuanzhe Liu, Xin Jin 0008 |
NSDI | 4 |
| 2025 | Optimizing RLHF Training for Large Language Models with Stage Fusion
Yinmin Zhong, Bingyang Wu, Changyi Wan, Hanpeng Hu, Ranchen Ming, Yibo Zhu 0001, Xin Jin 0008 |
NSDI | 11 |
| 2025 | Fornax: A Hardware-Centric Session Management in Large Public Cloud NetworkabstractSmartNIC is increasingly utilized to accelerate cloud network components. The effectiveness and correctness of hardware acceleration heavily rely on its management mechanism. Unfortunately, traditional management mechanisms adopt software-centric architecture, which treats flow as the basic management unit and completely relies on one-way commands to manage the flow table, making it challenging to support various cloud network scenarios while managing extremely large tables. In this paper, we advocate for a radical new mechanism to shift the management paradigm from software-centric architecture to hardware-centric architecture, which adopts session as the basic management unit and designs two-way protocols to facilitate the management process. We propose and implement a first-of-its-kind system, called Fornax, a novel management architecture for large public cloud networks. At the core of Fornax is leveraging a session-empowered hardware engine to provide various management capabilities. Besides, Fornax utilizes a light-weight software manager to enhance system scalability, and hardware-driven management protocols to improve resource efficiency. Our testbed evaluations demonstrate that Fornax can reduce the software storage usage by 80% and CPU usage by 77% with little hardware resource overhead. Our large-scale production results show that Fornax can manage up to 16M session entries while significantly reducing the resource overhead by over 79%. Heng Yu 0005, Jian Zhao 0006, Guozhi Lin, Baozeng Zhang, Yunpeng Guan, Jiajun Liang, Chao Pei, Yachen Wang, Xin Jin 0008, Jilong Wang 0001, Congcong Miao |
SIGCOMM | 14 |
| 2025 | DistTrain: Addressing Model and Data Heterogeneity with Disaggregated Training for Multimodal Large Language ModelsabstractMultimodal large language models (LLMs) empower LLMs to ingest inputs and generate outputs in multiple forms, such as text, image, and audio. However, the integration of multiple modalities introduces heterogeneity in both the model and training data, creating unique systems challenges. Yinmin Zhong, Hanpeng Hu, Jianjian Sun, Zheng Ge, Yibo Zhu 0001, Daxin Jiang, Xin Jin 0008 |
SIGCOMM | 9 |
| 2025 | MegaScale-Infer: Efficient Mixture-of-Experts Model Serving with Disaggregated Expert ParallelismabstractMixture-of-Experts (MoE) showcases tremendous potential to scale large language models (LLMs) with enhanced performance and reduced computational complexity. However, its sparsely activated architecture shifts feed-forward networks (FFNs) from being compute-intensive to memory-intensive during inference, leading to substantially lower GPU utilization and increased operational costs. Ruidong Zhu, Ziheng Jiang, Chao Jin 0007, Cesar A. Stuardo, Huaping Zhou, Jianzhe Xiao, Lingjun Liu, Haibin Lin, Li-Wen Chang, Jianxi Ye, Xuanzhe Liu, Xin Jin 0008, Xin Liu 0086 |
SIGCOMM | 19 |
| 2025 | Aegaeon: Effective GPU Pooling for Concurrent LLM Serving on the Market
Yuxing Xiang, Xue Li 0024, Kun Qian 0021, Diwen Zhu, Wenyuan Yu, Ennan Zhai, Xuanzhe Liu, Xin Jin 0008, Jingren Zhou 0001 |
SOSP | 9 |
| 2025 | EdgeLLM: Fast On-Device LLM Inference With Speculative DecodingabstractGenerative tasks, such as text generation and question answering, are essential for mobile applications. Given their inherent privacy sensitivity, executing them on devices is demanded. Nowadays, the execution of these generative tasks heavily relies on the Large Language Models (LLMs). However, the scarce device memory severely hinders the scalability of these models. We presentEdgeLLM, an efficient on-device LLM inference system for models whose sizes exceed the device's memory capacity.EdgeLLMis built atop speculative decoding, which delegates most tokens to a smaller, memory-resident (draft) LLM.EdgeLLMintegrates three novel techniques: (1) Instead of generating a fixed width and depth token tree,EdgeLLMproposes compute-efficient branch navigation and verification to pace the progress of different branches according to their accepted probability to prevent the wasteful allocation of computing resources to the wrong branch and to verify them all at once efficiently. (2) It uses a self-adaptive fallback strategy that promptly initiates the verification process when the smaller LLM generates an incorrect token. (3) To not block the generation,EdgeLLMproposes speculatively generating tokens during large LLM verification with the compute-IO pipeline. Through extensive experiments,EdgeLLMexhibits impressive token generation speed which is up to 9.3× faster than existing engines. Daliang Xu, Wangsong Yin, Hao Zhang 0108, Xin Jin 0008, Ying Zhang 0012, Shiyun Wei, Mengwei Xu 0001, Xuanzhe Liu |
IEEE Trans. Mob. Comput. | 4 |
| 2025 | Efficient Fault Tolerance for Stateful Serverless Computing with Asymmetric LoggingabstractServerless computing separates function execution from state management. Simple retry-based fault tolerance might corrupt the shared state with duplicate updates. Existing solutions employ log-based fault tolerance to achieve exactly-once semantics, where every single read or write to the external state is associated with a log for deterministic replay. However, logging is not a free lunch, which introduces considerable overhead to stateful serverless applications. We present Halfmoon, a serverless runtime system for fault-tolerant stateful serverless computing. Our key insight is that it is unnecessary to symmetrically log both reads and writes. Instead, it suffices to log either reads or writes, i.e., asymmetrically. We design two logging protocols that enforce exactly-once semantics while providing log-free reads and writes, which are suitable for read- and write-intensive workloads, respectively. We theoretically prove that the two protocols are log-optimal , i.e., no other protocols can achieve lower logging overhead than our protocols. We provide a criterion for choosing the right protocol for a given workload, and a pauseless switching mechanism to switch protocols for dynamic workloads. We implement a prototype of Halfmoon. Experiments show that Halfmoon achieves 20%–40% lower latency and 1.5–4.0× lower logging overhead than the state-of-the-art solution Boki. Haoyu Feng, Xuanzhe Liu, Xin Jin 0008 |
ACM Trans. Comput. Syst. | 4 |
| 2025 | FaaSPR: Latency-Oriented Placement and Routing Optimization for Serverless Workflow ProcessingabstractWorkflow processing enhances the applicability of serverless computing while retaining the characteristics of fine-grained resource management and elastic scalability. However, current serverless platforms lack targeted optimization of placement and routing strategies for workflow processing, leading to high overheads due to inter-server data transmission, instance cold starts, and function request queuing. We propose FaaSPR, a serverless scheduling system that exploits placement and routing optimizations to minimize workflow processing latency. FaaSPR groups instances with potential data transmission and proportionally distributes groups with heterogeneous instances across multiple servers, taking into account resource constraints and historical placement traces. This method addresses the issues of poor scalability and frequent instance migrations in existing solutions. Utilizing a routing algorithm based on multi-stage linear programming, FaaSPR minimizes cross-server data transmission within and between instance groups while ensuring load balancing among instances. Experiments show that, compared to the state-of-the-art solution FaaSFlow, FaaSPR decreases the average and 99th percentile tail latency by up to 68.03% and 93.46%, respectively. Additionally, reducing workflow processing latency leads to up to 46.18% decrease in resource consumption for FaaS users. Yunshan Jia, Chao Jin 0007, Qing Li 0028, Xuanzhe Liu, Xin Jin 0008 |
IEEE Trans. Netw. | 5 |
| 2025 | Efficient Far Memory-Aware Scheduling With FaMASabstractBenefiting from the low-latency RDMA network, far memory techniques that enable applications to leverage memory on remote machines have been proposed to improve job performance and resource utilization in datacenters. In a far memory-equipped datacenter, applications exhibit diverse characteristics when using far memory. However, existing datacenter schedulers are oblivious to the presence of far memory, resulting in suboptimal scheduling decisions that can compromise the performance guarantee of tasks. To this end, we present FaMAS, a far memory-aware datacenter scheduler that exploits the benefits of far memory to improve datacenter efficiency and avoid severe performance degradation of tasks. FaMAS tackles the far-memory scheduling problem for two objectives: reducing average job completion time and meeting deadlines. FaMAS introduces two novel far memory-aware algorithms: Job Completion Time First (JCTF) and Service Level Objective First (SLOF), tailored to the above two objectives respectively. JCTF formalizes the scheduling target with insights from the trade-off between the performance and the resource usage of individual tasks, while SLOF incorporates the growth rate of resource usage to represent tasks. The experimental results show that JCTF improves the average job completion time by up to$1.81\times $and SLOF achieves up to$3.95\times $lower deadline violations over the best traditional algorithm. Chiheng Lou, Xin Jin 0008 |
IEEE Trans. Netw. | 2 |
| 2025 | GreenFlow: A Carbon-Efficient Scheduler for Deep Learning WorkloadsabstractDeep learning (DL) has become a key component of modern software. Training DL models leads to huge carbon emissions. In data centers, it is important to reduce carbon emissions while completing DL training jobs early. In this article, we propose GreenFlow, a GPU cluster scheduler that reduces the average Job Completion Time (JCT) under a carbon emission budget. We first present performance models for DL training jobs to predict the throughput and energy consumption performance under different configurations. Based on the performance models and the carbon intensity of the grid, GreenFlow dynamically allocates GPUs, and adjusts the GPU-level and job-level configurations of DL training jobs. GreenFlow applies network packing and buddy allocation to job placement, thus avoiding extra carbon incurred by resource fragmentations. Evaluations on a real testbed show that when emitting the same amount of carbon, GreenFlow can improve the average JCT by up to 2.15×, compared to competitive baselines. Diandian Gu, Peng Sun 0006, Xin Jin 0008, Xuanzhe Liu |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2025 | Cannikin: No Lagger of SLO in Concurrent Multiple LoRA LLM ServingabstractLow-rank adaptation (LoRA) is widely used to efficiently fine-tune large language models (LLMs), leading to multiple models fine-tuned from the same pre-trained LLM. State-of-the-art LLM serving systems colocate these LoRA models on the same GPU instances for concurrent serving, which decreases memory usage and boosts efficiency. However, the unawareness of the SLO requirements of each LoRA service and the interference between requests from different LoRA services can cause significant SLO violations. This paper presents Cannikin, a multi-LoRA inference serving system that optimizes the minimum of the SLO attainments of all LoRA services in the serving system, denoted as lagger-SLO attainment. We obtain insights from the characterization of a real-world multi-LoRA serving trace, which reveals the stable input/output lengths of the most popular LoRA services. This motivates Cannikin to propose an SLO-aware scheduling algorithm that prioritizes requests based on efficient deadline estimation. Cannikin further detects the influence of interference between different LoRA services on SLO violations and eliminates the bias between these services. The evaluation using real-world traces demonstrates that compared to the state-of-the-art multi-LoRA serving systems, Cannikin can handle up to 3.6× higher rates or 2.8× more burstiness while maintaining the SLO attainment of each LoRA service$> $90% . Ruidong Zhu, Ziyue Jiang 0002, Zhi Zhang 0005, Xin Liu 0086, Xuanzhe Liu, Xin Jin 0008 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2025 | Niagara+: Scheduling Live ML Analytics Across Heterogeneous Device Processors and Edge ServersabstractIntelligent applications rely significantly on the live machine learning pipeline, a couple of deep neural network (DNN) inference services, executed on mobile devices to meet functional requirements while ensuring user data privacy. However, executing these DNN services on resource-constrained mobile devices presents a considerable challenge: low throughput and high energy consumption of inference tasks. To address this issue, we proposeNiagara+, a novel system designed to enhance throughput by jointly scheduling DNN inference services across heterogeneous processors on mobile devices and offloading services to powerful edge servers. To achieve this,Niagara+encounters two critical challenges: unpredictable workload dynamics and high scheduling complexity. To effectively tackle these challenges,Niagara+employs a predictive model to forecast incoming workload patterns and orchestrates service allocation across device heterogeneous processors and edge servers through a combination of two-step offline scheduling optimization and online service dispatching strategies. We implementedNiagara+and conducted comprehensive experiments, demonstrating its superiority over state-of-the-art approaches, reducing DNN service latency by up to 2.6× under high-bandwidth networks and 9.1× under low-bandwidth networks, while consistently meeting stringent inference latency requirements. Daliang Xu, Qing Li 0028, Mengwei Xu 0001, Gang Huang 0001, Shangguang Wang, Qun Wei, Xin Jin 0008, Yun Ma 0002, Xuanzhe Liu |
IEEE Trans. Serv. Comput. | 8 |
| 2024 | SoCFlow: Efficient and Scalable DNN Training on SoC-Clustered Edge ServersabstractSoC-Cluster, a novel server architecture composed of massive mobile system-on-chips (SoCs), is gaining popularity in industrial edge computing due to its energy efficiency and compatibility with existing mobile applications. However, we observe that the deployed SoC-Cluster servers are not fully utilized, because the hosted workloads are mostly user-triggered and have significant tidal phenomena. To harvest the free cycles, we propose to co-locate deep learning tasks on them. Daliang Xu, Mengwei Xu 0001, Chiheng Lou, Li Zhang 0133, Gang Huang 0001, Xin Jin 0008, Xuanzhe Liu |
ASPLOS (1) | 6 |
| 2024 | Unison: A Parallel-Efficient and User-Transparent Network Simulation KernelabstractDiscrete-event simulation (DES) is a prevalent tool for evaluating network designs. Although DES offers full fidelity and generality, its slow performance limits its application. To speed up DES, many network simulators employ parallel discrete-event simulation (PDES). However, adapting existing network simulation models to PDES requires complex reconfigurations and often yields limited performance improvement. In this paper, we address this gap by proposing a parallel-efficient and user-transparent network simulation kernel, Unison, that adopts fine-grained partition and load-adaptive scheduling optimized for network scenarios. We prototype Unison based on ns-3. Existing network simulation models of ns-3 can be seamlessly transitioned to Unison. Testbed experiments on commodity servers demonstrate that Unison can achieve a 40× speedup over DES using 24 CPU cores, and a 10× speedup compared with existing PDES algorithms under the same CPU cores. Songyuan Bai, Chen Tian 0001, Xiaoliang Wang 0001, Chang Liu 0001, Xin Jin 0008, Fu Xiao 0001, Qiao Xiang, Wan-Chun Dou, Guihai Chen |
EuroSys | 6 |
| 2024 | MegaScale: Scaling Large Language Model Training to More Than 10, 000 GPUs
Ziheng Jiang, Haibin Lin, Yinmin Zhong, Qi Huang 0001, Yangrui Chen, Zhi Zhang 0005, Yanghua Peng, Xiang Li 0067, Shibiao Nong, Yulu Jia, Sun He, Hongmin Chen, Zhihao Bai, Qi Hou, Shipeng Yan, Yiyao Sheng, Zhuo Jiang, Haohan Xu, Zhang Zhang 0003, Pengfei Nie, Leqi Zou, Sida Zhao, Zherui Liu, Xiaoying Jia 0001, Jianxi Ye, Xin Jin 0008, Xin Liu 0086 |
NSDI | 31 |
| 2024 | Jolteon: Unleashing the Promise of Serverless for Serverless Workflows
Chao Jin 0007, Xin Jin 0008 |
NSDI | 3 |
| 2024 | Fast Vector Query Processing for Large Datasets Beyond GPU Memory with Reordered Pipelining
Fangyue Liu, Gang Huang 0001, Xuanzhe Liu, Xin Jin 0008 |
NSDI | 5 |
| 2024 | Burstable Cloud Block Storage with Data Processing Units
Junyi Shu, Kun Qian 0021, Ennan Zhai, Xuanzhe Liu, Xin Jin 0008 |
OSDI | 5 |
| 2024 | dLoRA: Dynamically Orchestrating Requests and Adapters for LoRA LLM Serving
Bingyang Wu, Ruidong Zhu, Peng Sun 0006, Xuanzhe Liu, Xin Jin 0008 |
OSDI | 6 |
| 2024 | DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
Yinmin Zhong, Junda Chen, Jianbo Hu, Yibo Zhu 0001, Xuanzhe Liu, Xin Jin 0008, Hao Zhang 0025 |
OSDI | 7 |
| 2024 | LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence ParallelismabstractThe context window of large language models (LLMs) is rapidly increasing, leading to a huge variance in resource usage between different requests as well as between different phases of the same request. Restricted by static parallelism strategies, existing LLM serving systems cannot efficiently utilize the underlying resources to serve variable-length requests in different phases. To address this problem, we propose a new parallelism paradigm, elastic sequence parallelism (ESP), to elastically adapt to the variance across different requests and phases. Based on ESP, we design and build LoongServe, an LLM serving system that (1) improves computation efficiency by elastically adjusting the degree of parallelism in real-time, (2) improves communication efficiency by reducing key-value cache migration overhead and overlapping partial decoding communication with computation, and (3) improves GPU memory efficiency by reducing key-value cache fragmentation across instances. Our evaluation under diverse real-world datasets shows that LoongServe improves the throughput by up to 3.85× compared to chunked prefill and 5.81× compared to prefill-decoding disaggregation. Bingyang Wu, Yinmin Zhong, Peng Sun 0006, Xuanzhe Liu, Xin Jin 0008 |
SOSP | 6 |
| 2024 | SDCC: software-defined collective communication for distributed training
Xin Jin 0008, Zhen Zhang 0063, Yunshan Jia, Yun Ma 0002, Xuanzhe Liu |
Sci. China Inf. Sci. | 1 |
| 2024 | MuxFlow: efficient GPU sharing in production-level clusters with more than 10000 GPUs
Xuanzhe Liu, Shufan Liu, Xiang Li 0067, Yibo Zhu 0001, Xin Liu 0086, Xin Jin 0008 |
Sci. China Inf. Sci. | 7 |
| 2024 | Efficient, Scalable, and Sustainable DNN Training on SoC-Clustered Edge ServersabstractIn the realm of industrial edge computing, a novel server architecture known as SoC-Cluster, characterized by its aggregation of numerous mobile systems-on-chips (SoCs), has emerged as a promising solution owing to its enhanced energy efficiency and seamless integration with prevalent mobile applications. Despite its advantages, the utilization of SoC-Cluster servers remains unsatisfactory, primarily attributed to the tidal patterns of user-initiated workloads. To address such inefficiency, we introduceSoCFlow+, a pioneering framework designed to facilitate the co-location of deep learning training tasks on SoC-Cluster servers, thereby optimizing resource utilization.SoCFlow+incorporates three novel techniques tailored to mitigate the inherent limitations of commercial SoC-Cluster servers. First, it employs group-wise parallelism complemented by delayed aggregation, a strategy engineered to enhance the training efficiency and scalability of deep learning models, effectively circumventing network bottlenecks. Second, it integrates a data-parallel mixed-precision training algorithm, optimized to exploit the heterogeneous processing capabilities inherent to mobile SoCs fully. Third,SoCFlow+employs an underclocking-aware workload re-balanacing mechanism to tackle the training performance degradation caused by the thermal control of mobile SoCs. Through rigorous experimental validation,SoCFlow+achieves a convergence speedup ranging from 1.6× to 740× across 32 SoCs, compared to conventional benchmarks. Furthermore, when juxtaposed with commodity GPU servers (e.g., NVIDIA V100) under identical power constraints,SoCFlow+not only exhibits comparable training speed but also achieves a remarkable reduction in energy consumption by a factor of 2.31× to 10.23×, all while preserving convergence accuracy. Mengwei Xu 0001, Daliang Xu, Chiheng Lou, Li Zhang 0133, Gang Huang 0001, Xin Jin 0008, Xuanzhe Liu |
IEEE Trans. Mob. Comput. | 6 |
| 2024 | FLASH: Heterogeneity-Aware Federated Learning at ScaleabstractFederated learning (FL) becomes a promising machine learning paradigm. The impact of heterogeneous hardware specifications and dynamic states on the FL process has not yet been studied systematically. This paper presents the first large-scale study of this impact based on real-world data collected from 136k smartphones. We conducted extensive experiments on our proposed heterogeneity-aware FL platform namelyFLASH, to systematically explore the performance of state-of-the-art FL algorithms and key FL configurations in heterogeneity-aware and -unaware settings, finding the following. (1) Heterogeneity causes accuracy to drop by up to 9.2% and convergence time to increase by 2.32×. (2) Heterogeneity negatively impacts popular aggregation algorithms, e.g., the accuracy variance reduction brought byq-FedAvgdrops by 17.5%. (3) Heterogeneity does not worsen the accuracy loss caused by gradient-compression algorithms significantly, but it compromises the convergence time by up to 2.5×. (4) Heterogeneity hinders client-selection algorithms from selecting wanted clients, thus reducing effectiveness. e.g., the accuracy increase brought by the state-of-the-art client-selection algorithm drops by 73.9%. (5) Heterogeneity causes the optimal FL hyper-parameters to drift significantly. More specifically, the heterogeneity-unaware setting favors looser deadline and higher reporting fraction to achieve better training performance. (6) Heterogeneity results in non-trivial failed clients (more than 10%) and leads to participation bias (the top 30% of clients contribute 86% of computations). Our FLASH platform and data have been publicly open sourced. Chengxu Yang, Mengwei Xu 0001, Qipeng Wang 0001, Zhenpeng Chen 0001, Yun Ma 0002, Kaigui Bian, Gang Huang 0001, Yunxin Liu 0001, Xin Jin 0008, Xuanzhe Liu |
IEEE Trans. Mob. Comput. | 10 |
| 2024 | DistMind: Efficient Resource Disaggregation for Deep Learning WorkloadsabstractDeep learning (DL) systems suffer from low resource utilization due to 1) monolithic server model that tightly couples compute and memory; and 2) limited sharing between different inference applications, and across inference and training, because of strict service level objectives (SLOs). To address this problem, we present, a disaggregated DL system that enables efficient multiplexing of DL applications with near-optimal resource utilization. decouples compute from host memory, and exposes the abstractions of a GPU pool and a memory pool, each of which can be independently provisioned. The key challenge is to dynamically allocate GPU resources to different applications based on their real-time demands while meeting strict SLOs. We tackle this challenge by exploiting the power of high-speed 100 Gbps networks, and design three-stage pipelining, cache-aware load balancing, and DNN-aware sharding mechanisms based on the characteristics of DL workloads, to achieve millisecond-scale application loading overhead and improve system efficiency. We have implemented a prototype of and integrated it with PyTorch. Experimental results on AWS EC2 show that achieves near 100% resource utilization, and compared with NVIDIA MPS and Ray, improves the throughput by up to 279% and reduces the inference latency by up to 94%. Xin Jin 0008, Zhihao Bai, Zhen Zhang 0063, Yibo Zhu 0001, Yinmin Zhong, Xuanzhe Liu |
IEEE/ACM Trans. Netw. | 1 |
| 2024 | Pyxis: Scheduling Mixed Tasks in Disaggregated DatacentersabstractDisaggregating compute from storage is an emerging trend in cloud computing. Effectively utilizing resources in both compute and storage pool is the key to high performance. The state-of-the-art scheduler provides optimal scheduling decisions for workloads with homogeneous tasks. However, cloud applications often generate a mix of tasks with diverse compute and IO characteristics, resulting in sub-optimal performance for existing solutions. We present Pyxis, a system that provides optimal scheduling decisions for mixed workloads in disaggregated datacenters with theoretical guarantees. Pyxis is capable of maximizing overall throughput while meeting latency SLOs. Pyxis decouples the scheduling of different tasks. Our insight is that the optimal solution has an “all-or-nothing” structure that can be captured by a singleturning pointin the spectrum of tasks. Based on task characteristics, the turning point partitions the tasks either all to storage nodes or all to compute nodes (none to storage nodes). We theoretically prove that the optimal solution has such a structure, and design an online algorithm with sub-second convergence. We implement a prototype of Pyxis. Experiments on CloudLab with various synthetic and application workloads show that Pyxis improves the throughput by 3–21× over the state-of-the-art solution. Chao Jin 0007, Mosharaf Chowdhury, Zhenming Liu, Xuanzhe Liu, Xin Jin 0008 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2024 | Aquifer: Transparent Microsecond-Scale Scheduling for vRAN WorkloadsabstractVirtual Radio Access Network (vRAN) is an emerging approach offered by cloud providers to accelerate 5G services deployment. Despite significant microsecond-scale traffic variations, vRAN instances are provisioned based on peak load to meet strict latency requirements, leading to significant resource waste. Conceivably, vRAN can share CPUs with other applications to increase CPU utilization. Yet, existing sharing solutions require modifications to vRAN source code, hindering their deployment on public clouds. We present Aquifer, a microsecond-scale scheduler providing transparent CPU sharing for vRAN workloads. Our key observation is a common producer-consumer task execution pattern in mainstream vRAN implementations. We exploit this pattern to reclaim CPU cores from worker threads only at the boundary of processing different tasks. This guarantees run-to-completion task processing, which is critical for vRAN to achieve low latency and stability. Aquifer intercepts system calls invoked by vRAN at the OS layer to achieve transparent load monitoring and core reallocation. Aquifer employs a set of system-level optimizations on thread state detection, signal transmission and core selection, which reduces the scheduling cycle to 2$\mu s$. Experimental results show that Aquifer reclaims up to 88.31% of wasted CPU resources for two mainstream vRAN implementations, FlexRAN and OAI, without any source code modifications. Yunshan Jia, Yinmin Zhong, Meng Wang 0018, Xuanzhe Liu, Xin Jin 0008 |
IEEE Trans. Serv. Comput. | 7 |
| 2023 | ElasticFlow: An Elastic Serverless Training Platform for Distributed Deep LearningabstractThis paper proposes ElasticFlow, an elastic serverless training platform for distributed deep learning. ElasticFlow provides a serverless interface with two distinct features: (i) users specify only the deep neural network (DNN) model and hyperparameters for a job, but not the number of GPUs; (ii) users specify the deadline for a job, but not the amount of time to occupy GPUs. In contrast to existing server-centric platforms, ElasticFlow provides performance guarantees in terms of meeting deadlines while alleviating tedious, low-level, and manual resource management for deep learning developers. The characteristics of distributed training introduce two challenges. First, the training throughput scales non-linearly with the number of GPUs. Second, the scaling efficiency is affected by worker placement. To address these challenges, we propose Minimum Satisfactory Share to capture the resource usage of training jobs to meet deadlines, and ElasticFlow performs admission control based on it. We develop a greedy algorithm that dynamically allocates resources to admitted jobs based on diminishing returns. We apply buddy allocation to worker placement to eliminate the effect of topology. Evaluation results on a cluster of 128 GPUs show that ElasticFlow increases the number of jobs that can meet their deadlines by 1.46–7.65× compared to existing solutions. Diandian Gu, Yinmin Zhong, Yifan Xiong 0001, Zhenhua Han, Peng Cheng 0005, Fan Yang 0024, Gang Huang 0001, Xin Jin 0008, Xuanzhe Liu |
ASPLOS (2) | 9 |
| 2023 | Disaggregated RAID Storage in Modern DatacentersabstractRAID (Redundant Array of Independent Disks) has been widely adopted for decades, as it provides enhanced throughput and redundancy beyond what a single disk can offer. Today, enabled by fast datacenter networks, accessing remote block devices with acceptable overhead (i.e. disaggregated storage) becomes a reality (e.g., for serverless applications). Combining RAID with remote storage can provide the same benefits while creating better fault tolerance and flexibility than its monolithic counterparts. The key challenge of disaggregated RAID is to handle extra network traffic generated by RAID, which can consume a vast amount of NIC bandwidth. We present dRAID, a disaggregated RAID system that achieves near-optimal read and write throughput. dRAID exploits peer-to-peer disaggregated data access to reduce bandwidth consumption in both normal and degraded states. It employs non-blocking multi-stage writes to maximize inter-node parallelism, and applies pipelined I/O processing to maximize inter-device parallelism. We introduce bandwidth-aware reconstruction for better load balancing. We show that dRAID provides up to 3× bandwidth improvement. The results on a lightweight object store show that dRAID brings 1.5×-2.35× throughput improvement on various workloads. Junyi Shu, Ruidong Zhu, Yun Ma 0002, Gang Huang 0001, Hong Mei 0001, Xuanzhe Liu, Xin Jin 0008 |
ASPLOS (3) | 7 |
| 2023 | Niagara: Scheduling DNN Inference Services on Heterogeneous Edge Processors
Daliang Xu, Qing Li 0028, Mengwei Xu 0001, Gang Huang 0001, Shangguang Wang, Xin Jin 0008, Yun Ma 0002, Xuanzhe Liu |
ICSOC (1) | 7 |
| 2023 | Transparent GPU Sharing in Container Clouds for Deep Learning Workloads
Bingyang Wu, Zhihao Bai, Xuanzhe Liu, Xin Jin 0008 |
NSDI | 5 |
| 2023 | Fast, Approximate Vector Queries on Very Large Unstructured Datasets
Chao Jin 0007, Linpeng Tang, Xuanzhe Liu, Xin Jin 0008 |
NSDI | 5 |
| 2023 | AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning Serving
Zhuohan Li 0001, Lianmin Zheng, Yinmin Zhong, Vincent Liu 0001, Ying Sheng 0007, Xin Jin 0008, Yanping Huang, Hao Zhang 0025, Joseph Gonzalez 0001, Ion Stoica |
OSDI | 6 |
| 2023 | Ditto: Efficient Serverless Analytics with Elastic ParallelismabstractServerless computing provides fine-grained resource elasticity for data analytics---a job can flexibly scale its resources for each stage, instead of sticking to a fixed pool of resources throughout its lifetime. Due to different data dependencies and different shuffling overheads caused by intra- and inter-server communication, the best degree of parallelism (DoP) for each stage varies based on runtime conditions. Chao Jin 0007, Xingyu Xiang, Songyun Zou, Gang Huang 0001, Xuanzhe Liu, Xin Jin 0008 |
SIGCOMM | 7 |
| 2023 | XRON: A Hybrid Elastic Cloud Overlay Network for Video Conferencing at Planetary ScaleabstractQuality and cost are two key considerations for video conferencing services. Service providers face a dilemma when selecting network tiers to build their infrastructure---relying on Internet links has poor quality, while using premium links brings excessive cost. Bingyang Wu, Kun Qian 0021, Bo Li 0061, Dennis Cai, Ennan Zhai, Xuanzhe Liu, Xin Jin 0008 |
SIGCOMM | 11 |
| 2023 | Understanding the Micro-Behaviors of Hardware Offloaded Network Stacks with LuminaabstractHardware offloaded network stacks are widely adopted in modern datacenters to meet the demand for high throughput, ultra-low latency and low CPU overhead. To fully leverage their exceptional performance, users need to have a deep understanding of their behaviors. Despite many efforts on testing software network stacks, hardware network stacks impose unique challenges to testing tools due to their kernel bypass nature and high performance. Zhuolong Yu, Wei Bai 0001, Shachar Raindel, Vladimir Braverman, Xin Jin 0008 |
SIGCOMM | 6 |
| 2023 | Klotski: Efficient and Safe Network Migration of Large Production DatacentersabstractThis paper presents the design, implementation, evaluation, and deployment of Meta's production network migration system. We first introduce the network migration problem for large-scale production datacenter networks (DCNs). A network migration task at Meta touches as many as hundreds of switches and tens of thousands of circuits per datacenter (DC), and involves physical deployment work on site that can last months. We describe real-world migration challenges, covering complex and evolving DCN architectures and operational constraints. We mathematically formalize the problem of generating efficient and safe migration plans, and exploit the inherent symmetry and locality of DCN topologies to prune the search space. We design an ordering-agnostic compact topology representation to eliminate redundant satisfiability checking, and apply the A* algorithm with a domain-specific priority function to find the optimal plan. Evaluation results on a range of production migration cases show that Klotski reduces the time to find optimal migration plans by up to 381× compared to prior solutions. We hope by introducing the problem and sharing our deployment experience, this work can provide a useful context for network migration in the real world and inspire future research. Xiaoxiang Zhang, Ying Zhang 0022, Zhaodong Wang, Yuandong Tian, Alex Nikulkov, Joao Ferreira, Xuanzhe Liu, Xin Jin 0008 |
SIGCOMM | 10 |
| 2023 | Oobleck: Resilient Distributed Training of Large Models Using Pipeline TemplatesabstractOobleck enables resilient distributed training of large DNN models with guaranteed fault tolerance. It takes a planning-execution co-design approach, where it first generates a set of heterogeneous pipeline templates and instantiates at least f + 1 logically equivalent pipeline replicas to tolerate any f simultaneous failures. During execution, it relies on already-replicated model states across the replicas to provide fast recovery. Oobleck provably guarantees that some combination of the initially created pipeline templates can be used to cover all available resources after f or fewer simultaneous failures, thereby avoiding resource idling at all times. Evaluation on large DNN models with billions of parameters shows that Oobleck provides consistently high throughput, and it outperforms state-of-the-art fault tolerance solutions like Bamboo and Varuna by up to 13.9×. Insu Jang, Zhenning Yang, Zhen Zhang 0063, Xin Jin 0008, Mosharaf Chowdhury |
SOSP | 4 |
| 2023 | Halfmoon: Log-Optimal Fault-Tolerant Stateful Serverless ComputingabstractServerless computing separates function execution from state management. Simple retry-based fault tolerance might corrupt the shared state with duplicate updates. Existing solutions employ log-based fault tolerance to achieve exactlyonce semantics, where every single read or write to the external state is associated with a log for deterministic replay. However, logging is not a free lunch, which introduces considerable overhead to stateful serverless applications. Xuanzhe Liu, Xin Jin 0008 |
SOSP | 3 |
| 2023 | Automated Verification of an In-Production DNS Authoritative EngineabstractThis paper presents DNS-V, a verification framework for our in-production DNS authoritative engine, which is the core of our DNS service. The key idea for automated verification in general is based on the layered verification principle. However, we face the challenge that our in-production DNS authoritative engine lacks modularity, more specifically, as can be seen with unclean interfaces and poor data structure encapsulation. This makes the layered verification hard to apply. To address this challenge, we propose a summarization approach that performs full-path symbolic execution to accumulate all path conditions and computation effects, and then represents a module's behavior in an abstract form as a set of input-effect pairs. In addition, for portability to future iterated versions of our DNS authoritative engine, we identify common dependency library modules that remain stable across different versions, and carefully design their abstractions to make them amenable to automated reasoning. Our framework has been successful in identifying and preventing tens of critical bugs in different versions of our DNS authoritative engine from reaching production, with a porting effort of less than one person-week. Naiqian Zheng, Mengqi Liu 0001, Yuxing Xiang, Linjian Song, Nan Wang 0041, Zhuo Liang, Dennis Cai, Ennan Zhai, Xuanzhe Liu, Xin Jin 0008 |
SOSP | 13 |
| 2023 | Scalable and Efficient Full-Graph GNN Training for Large GraphsabstractGraph Neural Networks (GNNs) have emerged as powerful tools to capture structural information from graph-structured data, achieving state-of-the-art performance on applications such as recommendation, knowledge graph, and search. Graphs in these domains typically contain hundreds of millions of nodes and billions of edges. However, previous GNN systems demonstrate poor scalability because large and interleaved computation dependencies in GNN training cause significant overhead in current parallelization methods. We present G3, a distributed system that can efficiently train GNNs over billion-edge graphs at scale. G3 introduces GNN hybrid parallelism which synthesizes three dimensions of parallelism to scale out GNN training by sharing intermediate results peer-to-peer in fine granularity, eliminating layer-wise barriers for global collective communication or neighbor replications as seen in prior works. G3 leverages locality-aware iterative partitioning and multi-level pipeline scheduling to exploit acceleration opportunities by distributing balanced workload among workers and overlapping computation with communication in both inter-layer and intra-layer training processes. We show via a prototype implementation and comprehensive experiments that G3 can achieve as much as 2.24x speedup in a 16-node cluster, and better final accuracy over prior works. Xinchen Wan, Kaiqiang Xu, Xudong Liao, Yilun Jin, Kai Chen 0005, Xin Jin 0008 |
Proc. ACM Manag. Data | 6 |
| 2023 | Enabling Edge-Cloud Video Analytics for Robotics ApplicationsabstractEmerging deep learning-based video analytics tasks demand computation-intensive neural networks and powerful computing resources on the cloud to achieve high accuracy. Due to the latency requirement and limited network bandwidth, edge-cloud systems adaptively compress the data to strike a balance between overall analytics accuracy and bandwidth consumption. However, the degraded data leads to another issue of poortail accuracy, which means the extremely low accuracy of a few semantic classes and video frames. Autonomous robotics applications especially value the tail accuracy performance but suffer using the prior edge-cloud systems. We present Runespoor, an edge-cloud video analytics system to manage the tail accuracy and enable emerging robotics applications. We train and deploy a super-resolution model tailored for the tail accuracy of analytics tasks on the server to significantly improves the performance on hard-to-detect classes and sophisticated frames. During online operation, we use an adaptive data rate controller to further improve the tail performance by instantly adjusting the data rate policy according to the video content. Our evaluation shows that Runespoor improves class-wise tail accuracy by up to 300%, frame-wise 90%/99% tail accuracy by up to 22%/54%, and greatly improves the overall accuracy and bandwidth trade-off. Weiyan Wang, Duowen Liu, Xin Jin 0008, Junchen Jiang, Kai Chen 0005 |
IEEE Trans. Cloud Comput. | 4 |
| 2023 | Rise of Distributed Deep Learning Training in the Big Model Era: From a Software Engineering PerspectiveabstractDeep learning (DL) has become a key component of modern software. In the “ big model ” era, the rich features of DL-based software (i.e., DL software) substantially rely on powerful DL models, e.g., BERT, GPT-3, and the recently emerging GPT-4, which are trained on the powerful cloud with large datasets. Hence, training effective DL models has become a vital stage in the whole software lifecycle. When training deep learning models, especially those big models, developers need to parallelize and distribute the computation and memory resources amongst multiple devices (e.g., a cluster of GPUs) in the training process, which is known as distributed deep learning training , or distributed training for short. However, the unique challenges that developers encounter in distributed training process have not been studied in the software engineering community. Given the increasingly heavy dependence of current DL-based software on distributed training, this paper aims to fill in the knowledge gap and presents the first comprehensive study on developers’ issues in distributed training. To this end, we focus on popular DL frameworks that support distributed training (including TensorFlow, PyTorch, Keras, and Horovod) and analyze 1,131 real-world developers’ issues about using these frameworks reported on Stack Overflow and GitHub. We construct a fine-grained taxonomy consisting of 30 categories regarding the fault symptoms and summarize common fix patterns for different symptoms. We find that: (1) many distributed-specific faults and non-distributed-specific faults inherently share the same fault symptoms, making it challenging to debug; (2) most of the fault symptoms have frequent fix patterns; (3) about half of the faults are related to system-level configurations. Based on the results, we suggest actionable implications on research avenues that can potentially facilitate the distributed training to develop DL-based software, such as focusing on the frequent and common fix patterns when designing testing or debugging tools, developing efficient testing and debugging techniques for communication configuration along with the synthesis of network configuration analysis, designing new multi-device checkpoint-and-replay techniques to help reproduction, and designing serverless APIs for cloud platforms. Xuanzhe Liu, Diandian Gu, Zhenpeng Chen 0001, Jinfeng Wen, Yun Ma 0002, Haoyu Wang 0001, Xin Jin 0008 |
ACM Trans. Softw. Eng. Methodol. | 8 |
| 2023 | FaaSLight: General Application-level Cold-start Latency Optimization for Function-as-a-Service in Serverless ComputingabstractServerless computing is a popular cloud computing paradigm that frees developers from server management. Function-as-a-Service (FaaS) is the most popular implementation of serverless computing, representing applications as event-driven and stateless functions. However, existing studies report that functions of FaaS applications severely suffer from cold-start latency. In this article, we propose an approach, namely, FaaSLight , to accelerating the cold start for FaaS applications through application-level optimization. We first conduct a measurement study to investigate the possible root cause of the cold-start problem of FaaS. The result shows that application code loading latency is a significant overhead. Therefore, loading only indispensable code from FaaS applications can be an adequate solution. Based on this insight, we identify code related to application functionalities by constructing the function-level call graph and separate other code (i.e., optional code) from FaaS applications. The separated optional code can be loaded on demand to avoid the inaccurate identification of indispensable code causing application failure. In particular, a key principle guiding the design of FaaSLight is inherently general, i.e., platform - and language-agnostic . In practice, FaaSLight can be effectively applied to FaaS applications developed in different programming languages (Python and JavaScript), and can be seamlessly deployed on popular serverless platforms such as AWS Lambda and Google Cloud Functions, without having to modify the underlying OSes or hypervisors, nor introducing any additional manual engineering efforts to developers. The evaluation results on real-world FaaS applications show that FaaSLight can significantly reduce the code loading latency (up to 78.95%, 28.78% on average), thereby reducing the cold-start latency. As a result, the total response latency of functions can be decreased by up to 42.05% (19.21% on average). Compared with the state-of-the-art, FaaSLight achieves a 21.25× improvement in reducing the average total response latency. Xuanzhe Liu, Jinfeng Wen, Zhenpeng Chen 0001, Ding Li 0001, Junkai Chen, Yi Liu 0014, Haoyu Wang 0001, Xin Jin 0008 |
ACM Trans. Softw. Eng. Methodol. | 8 |
| 2023 | Rise of the Planet of Serverless Computing: A Systematic ReviewabstractServerless computing is an emerging cloud computing paradigm, being adopted to develop a wide range of software applications. It allows developers to focus on the application logic in the granularity of function, thereby freeing developers from tedious and error-prone infrastructure management. Meanwhile, its unique characteristic poses new challenges to the development and deployment of serverless-based applications. To tackle these challenges, enormous research efforts have been devoted. This article provides a comprehensive literature review to characterize the current research state of serverless computing. Specifically, this article covers 164 articles on 17 research directions of serverless computing, including performance optimization, programming framework, application migration, multi-cloud development, testing and debugging, and so on. It also derives research trends, focus, and commonly-used platforms for serverless computing, as well as promising research opportunities. Jinfeng Wen, Zhenpeng Chen 0001, Xin Jin 0008, Xuanzhe Liu |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2022 | Multi-objective congestion controlabstractDecades of research on Internet congestion control (CC) have produced a plethora of algorithms that optimize for different performance objectives. Applications face the challenge of choosing the most suitable algorithm based on their needs, and it takes tremendous efforts and expertise to customize CC algorithms when new demands emerge. In this paper, we explore a basic question: can we design a single CC algorithm to satisfy different objectives? Yiqing Ma, Han Tian, Xudong Liao, Junxue Zhang 0001, Weiyan Wang, Kai Chen 0005, Xin Jin 0008 |
EuroSys | 7 |
| 2022 | Mandheling: mixed-precision on-device DNN training with DSP offloadingabstractThis paper proposes Mandheling, the first system that enables highly resource-efficient on-device training by orchestrating mixed-precision training with on-chip Digital Signal Processor (DSP) offloading. Mandheling fully explores the advantages of DSP in integer-based numerical calculations using four novel techniques: (1) a CPU-DSP co-scheduling scheme to situationally mitigate the overhead from DSP-unfriendly operators; (2) a self-adaptive rescaling algorithm to reduce the overhead of dynamic rescaling in backward propagation; (3) a batch-splitting algorithm to improve DSP cache efficiency; (4) a DSP compute subgraph-reusing mechanism to eliminate the preparation overhead on DSP. We have fully implemented Mandheling and demonstrated its effectiveness through extensive experiments. The results show that, compared to the state-of-the-art DNN engines from TFLite and MNN, Mandheling reduces per-batch training time by 5.5X and energy consumption by 8.9X on average. In end-to-end training tasks, Mandheling reduces convergence time by up to 10.7X and energy consumption by 13.1X, with only 1.9%--2.7% accuracy loss compared to the FP32 precision setting. Daliang Xu, Mengwei Xu 0001, Qipeng Wang 0001, Shangguang Wang, Yun Ma 0002, Gang Huang 0001, Xin Jin 0008, Xuanzhe Liu |
MobiCom | 8 |
| 2022 | Melon: breaking the memory wall for resource-efficient on-device machine learningabstractOn-device learning is a promising technique for emerging privacy-preserving machine learning paradigms. However, through quantitative experiments, we find that commodity mobile devices cannot well support state-of-the-art DNN training with a large enough batch size, due to the limited local memory capacity. To fill the gap, we propose Melon, a memory-friendly on-device learning framework that enables the training tasks with large batch size beyond the physical memory capacity. Melon judiciously retrofits existing memory saving techniques to fit into resource-constrained mobile devices, i.e., recomputation and micro-batch. Melon further incorporates novel techniques to deal with the high memory fragmentation and memory adaptation. We implement and evaluate Melon with various typical DNN models on commodity mobile devices. The results show that Melon can achieve up to 4.33× larger batch size under the same memory budget. Given the same batch size, Melon achieves 1.89× on average (up to 4.01×) higher training throughput, and saves up to 49.43% energy compared to competitive alternatives. Furthermore, Melon reduces 78.59% computation on average in terms of memory budget adaptation. Qipeng Wang 0001, Mengwei Xu 0001, Chao Jin 0007, Xinran Dong, Jinliang Yuan, Xin Jin 0008, Gang Huang 0001, Yunxin Liu 0001, Xuanzhe Liu |
MobiSys | 6 |
| 2022 | NetVRM: Virtual Register Memory for Programmable Networks
Tao Wang 0088, Dan R. K. Ports, Anirudh Sivaraman, Xin Jin 0008 |
NSDI | 6 |
| 2022 | Multi-resource interleaving for deep learning trainingabstractTraining Deep Learning (DL) model requires multiple resource types, including CPUs, GPUs, storage IO, and network IO. Advancements in DL have produced a wide spectrum of models that have diverse usage patterns on different resource types. Existing DL schedulers focus on only GPU allocation, while missing the opportunity of packing jobs along multiple resource types. Yuanqiang Liu, Yanghua Peng, Yibo Zhu 0001, Xuanzhe Liu, Xin Jin 0008 |
SIGCOMM | 6 |
| 2022 | Meissa: scalable network testing for programmable data planesabstractEnsuring the correctness of programmable data planes is important. Testing offers comprehensive correctness checking, including detecting both code bugs and non-code bugs. However, scalability is a key challenge for testing production-scale data planes to achieve high coverage. This paper presents Meissa, a scalable network testing system for programmable data planes with full path coverage. The core of Meissa is a domain-specific code summary technique that simplifies the control flow graph of a data plane program for scalable testing without sacrificing coverage. Code summary decomposes a data plane program into individual pipelines, and summarizes each pipeline with a succinct representation. We formally prove that Meissa with code summary achieves 100% path coverage. We use both open-source and production-scale data plane programs to evaluate Meissa. The evaluation shows that (i) Meissa is able to test production-scale data plane programs that cannot be supported by state-of-the-art efforts, and (ii) besides P4 code bugs, Meissa is able to not only identify known non-code bugs, but also detect previously-unknown non-code bugs. We also share in this paper several real cases tested by Meissa in a production programmable data plane. Naiqian Zheng, Mengqi Liu 0001, Ennan Zhai, Hongqiang Harry Liu, Kaicheng Yang 0001, Xuanzhe Liu, Xin Jin 0008 |
SIGCOMM | 8 |
| 2022 | MiCS: Near-linear Scaling for Training Gigantic Model on Public CloudabstractExisting general purpose frameworks for gigantic model training, i.e., dense models with billions of parameters, cannot scale efficiently on cloud environment with various networking conditions due to large communication overheads. In this paper, we propose MiCS, which Minimizes the Communication Scale to bring down communication overhead. Specifically, by decreasing the number of participants in a communication collective, MiCS can utilize heterogeneous network bandwidth, reduce network traffic over slower links, reduce the latency of communications for maintaining high network bandwidth utilization, and amortize expensive global gradient synchronization overhead. Our evaluation on AWS shows that the system throughput of MiCS is up to 2.89× that of the state-of-the-art large model training systems. MiCS achieves near-linear scaling efficiency, which is up to 1.27× that of DeepSpeed. MiCS allows us to train a proprietary model with 100 billion parameters on 512 GPUs with 99.4% weak-scaling efficiency, and it is able to saturate over 54.5% theoretical computation power of each GPU on a public cloud with less GPU memory and more restricted networks than DGX-A100 clusters. Zhen Zhang 0063, Shuai Zheng 0004, Yida Wang 0003, Justin Chiu, George Karypis, Trishul Chilimbi, Mu Li 0003, Xin Jin 0008 |
Proc. VLDB Endow. | 8 |
| 2021 | FAST: FPGA-based Subgraph Matching on Massive GraphsabstractSubgraph matching is a basic operation widely used in many applications. However, due to its NP-hardness and the explosive growth of graph data, it is challenging to compute subgraph matching, especially in large graphs. In this paper, we aim at scaling up subgraph matching on a single machine using FPGAs. Specifically, we propose a CPU-FPGA co-designed framework. On the CPU side, we first develop a novel auxiliary data structure called candidate search tree (CST) which serves as a complete search space of subgraph matching. CST can be partitioned and fully loaded into FPGAs' on-chip memory. Then, a workload estimation technique is proposed to balance the load between the CPU and FPGA. On the FPGA side, we design and implement the first FPGA-based subgraph matching algorithm, called FAST. To take full advantage of the pipeline mechanism on FPGAs, task parallelism optimization and task generator separation strategy are proposed for FAST, achieving massive parallelism. Moreover, we carefully develop a BRAM-only matching process to fully utilize FPGA's on-chip memory, which avoids the expensive intermediate data transfer between FPGA's BRAM and DRAM. Comprehensive experiments show that FAST achieves up to 462.0x and 150.0x speedup compared with the state-of-the-art algorithm DAF and CECI, respectively. In addition, FAST is the only algorithm that can handle the billion-scale graph using one machine in our experiments. Xin Jin 0008, Zhengyi Yang 0001, Xuemin Lin 0001, Shiyu Yang 0002, Lu Qin 0001 |
ICDE | 1 |
| 2021 | Enabling Edge-Cloud Video Analytics for Robotics ApplicationsabstractEmerging deep learning-based video analytics tasks demand computation-intensive neural networks and powerful computing resources on the cloud to achieve high accuracy. Due to the latency requirement and limited network bandwidth, edge-cloud systems adaptively compress the data to strike a balance between overall analytics accuracy and bandwidth consumption. However, the degraded data leads to another issue of poor tail accuracy, which means the extremely low accuracy of a few semantic classes and video frames. Autonomous robotics applications especially value the tail accuracy performance but suffer using the prior edge-cloud systems.We present Runespoor, an edge-cloud video analytics system to manage the tail accuracy and enable emerging robotics applications. We train and deploy a super-resolution model tailored for the tail accuracy of analytics tasks on the server to significantly improves the performance on hard-to-detect classes and sophisticated frames. During online operation, we use an adaptive data rate controller to further improve the tail performance by instantly adjusting the data rate policy according to the video content. Our evaluation shows that Runespoor improves class-wise tail accuracy by up to 300%, frame-wise 90%/99% tail accuracy by up to 22%/54%, and greatly improves the overall accuracy and bandwidth trade-off. Weiyan Wang, Duowen Liu, Xin Jin 0008, Junchen Jiang, Kai Chen 0005 |
INFOCOM | 4 |
| 2021 | Ship Compute or Ship Data? Why Not Both?
Jingfeng Wu, Xin Jin 0008, Mosharaf Chowdhury |
NSDI | 3 |
| 2021 | Twenty Years After: Hierarchical Core-Stateless Fair Queueing
Zhuolong Yu, Jingfeng Wu, Vladimir Braverman, Ion Stoica, Xin Jin 0008 |
NSDI | 5 |
| 2021 | Programmable packet scheduling with a single queueabstractProgrammable packet scheduling enables scheduling algorithms to be programmed into the data plane without changing the hardware. Existing proposals either have no hardware implementations for switch ASICs or require multiple strict-priority queues. Zhuolong Yu, Chuheng Hu, Jingfeng Wu, Xiao Sun 0004, Vladimir Braverman, Mosharaf Chowdhury, Zhenhua Liu 0002, Xin Jin 0008 |
SIGCOMM | 8 |
| 2021 | Network planning with deep reinforcement learningabstractNetwork planning is critical to the performance, reliability and cost of web services. This problem is typically formulated as an Integer Linear Programming (ILP) problem. Today's practice relies on hand-tuned heuristics from human experts to address the scalability challenge of ILP solvers. Satyajeet Ahuja, Yuandong Tian, Ying Zhang 0022, Xin Jin 0008 |
SIGCOMM | 6 |
| 2021 | An empirical study on challenges of application development in serverless computingabstractServerless computing is an emerging paradigm for cloud computing, gaining traction in a wide range of applications such as video processing and machine learning. This new paradigm allows developers to focus on the development of the logic of serverless computing based applications (abbreviated as serverless-based applications) in the granularity of function, thereby freeing developers from tedious and error-prone infrastructure management. Meanwhile, it also introduces new challenges on the design, implementation, and deployment of serverless-based applications, and current serverless computing platforms are far away from satisfactory. However, to the best of our knowledge, these challenges have not been well studied. To fill this knowledge gap, this paper presents the first comprehensive study on understanding the challenges in developing serverless-based applications from the developers’ perspective. We mine and analyze 22,731 relevant questions from Stack Overflow (a popular Q&A website for developers), and show the increasing popularity trend and the high difficulty level of serverless computing for developers. Through manual inspection of 619 sampled questions, we construct a taxonomy of challenges that developers encounter, and report a series of findings and actionable implications. Stakeholders including application developers, researchers, and cloud providers can leverage these findings and implications to better understand and further explore the serverless computing paradigm. Jinfeng Wen, Zhenpeng Chen 0001, Yi Liu 0014, Yiling Lou, Yun Ma 0002, Gang Huang 0001, Xin Jin 0008, Xuanzhe Liu |
ESEC/SIGSOFT FSE | 7 |
| 2021 | Runtime Recovery of Web Applications under Zero-Day ReDoS AttacksabstractRegular expression denial of service (ReDoS)— which exploits the super-linear running time of matching regular expressions against carefully crafted inputs—is an emerging class of DoS attacks to web services. One challenging question for a victim web service under ReDoS attacks is how to quickly recover its normal operation after ReDoS attacks, especially these zero-day ones exploiting previously unknown vulnerabilities.In this paper, we present RegexNet, the first payload-based, automated, reactive ReDoS recovery system for web services. RegexNet adopts a learning model, which is updated constantly in a feedback loop during runtime, to classify payloads of upcoming requests including the request contents and database query responses. If detected as a cause leading to ReDoS, RegexNet migrates those requests to a sandbox and isolates their execution for a fast, first-measure recovery.We have implemented a RegexNet prototype and integrated it with HAProxy and Node.js. Evaluation results show that RegexNet is effective in recovering the performance of web services against zero-day ReDoS attacks, responsive on reacting to attacks in sub-minute, and resilient to different ReDoS attack types including adaptive ones that are designed to evade RegexNet on purpose. Zhihao Bai, Ke Wang 0040, Yinzhi Cao, Xin Jin 0008 |
SP | 5 |
| 2021 | Jaqen: A High-Performance Switch-Native Approach for Detecting and Mitigating Volumetric DDoS Attacks with Programmable Switches
Zaoxing Liu, Hun Namkung, Georgios Nikolaidis, Jeongkeun Lee, Changhoon Kim, Xin Jin 0008, Vladimir Braverman, Minlan Yu, Vyas Sekar |
USENIX Security Symposium | 6 |
| 2020 | On Efficient Constructions of CheckpointsabstractEfficient construction of checkpoints/snapshots is a critical tool for training and diagnosing deep learning models. In this paper, we propose a lossy compression scheme for checkpoint constructions (called LC-Checkpoint). LC-Checkpoint simultaneously maximizes the compression rate and optimizes the recovery speed, under the assumption that SGD is used to train the model. LC-Checkpoint uses quantization and priority promotion to store the most crucial information for SGD to recover, and then uses a Huffman coding to leverage the non-uniform distribution of the gradient scales. Our extensive experiments show that LC-Checkpoint achieves a compression rate up to 28{\texttimes} and recovery speedup up to 5.77{\texttimes} over a state-of-the-art algorithm (SCAR). Yu Chen 0036, Zhenming Liu, Bin Ren 0002, Xin Jin 0008 |
ICML | 4 |
| 2020 | PipeSwitch: Fast Pipelined Context Switching for Deep Learning Applications
Zhihao Bai, Zhen Zhang 0063, Yibo Zhu 0001, Xin Jin 0008 |
OSDI | 4 |
| 2020 | Pegasus: Tolerating Skewed Workloads in Distributed Storage with In-Network Coherence Directories
Jialin Li 0001, Jacob Nelson 0001, Ellis Michael, Xin Jin 0008, Dan R. K. Ports |
OSDI | 4 |
| 2020 | RackSched: A Microsecond-Scale Scheduler for Rack-Scale Computers
Kostis Kaffes, Zixu Chen, Zhenming Liu, Christoforos E. Kozyrakis, Ion Stoica, Xin Jin 0008 |
OSDI | 7 |
| 2020 | NetLock: Fast, Centralized Lock Management Using Programmable SwitchesabstractLock managers are widely used by distributed systems. Traditional centralized lock managers can easily support policies between multiple users using global knowledge, but they suffer from low performance. In contrast, emerging decentralized approaches are faster but cannot provide flexible policy support. Furthermore, performance in both cases is limited by the server capability. Zhuolong Yu, Yiwen Zhang 0008, Vladimir Braverman, Mosharaf Chowdhury, Xin Jin 0008 |
SIGCOMM | 5 |
| 2019 | PatMat: A Distributed Pattern Matching Engine with CypherabstractGraph pattern matching is one of the most fundamental problems in graph database and is associated with a wide spectrum of applications. Due to its computational intensiveness, researchers have primarily devoted their efforts to improving the performance of the algorithm while constraining the graphs to have singular labels on vertices (edges) or no label. Whereas in practice graphs are typically associated with rich properties, thus the main focus in the industry is instead on powerful query languages that can express a sufficient number of pattern matching scenarios. We demo PatMat in this work to glue together the academic efforts on performance and the industrial efforts on expressiveness. To do so, we leverage the state-of-the-art join-based algorithms in the distributed contexts and Cypher query language - the most widely-adopted declarative language for graph pattern matching. The experiments demonstrate how we are capable of turning complex Cypher semantics into a distributed solution with high performance. Kongzhang Hao, Zhengyi Yang 0001, Longbin Lai, Zhengmin Lai, Xin Jin 0008, Xuemin Lin 0001 |
CIKM | 5 |
| 2019 | QPipe: quantiles sketch fully in the data planeabstractEfficient network management requires collecting a variety of statistics over the packet flows. Monitoring the flows directly in the data plane allows the system to detect anomalies faster. However, monitoring algorithms have to handle a throughput of 109 packets per second and to maintain a very low memory footprint. Widely adopted sampling-based approaches suffer from low accuracy in estimations. Thus, it is natural to ask: "Is it possible to maintain important statistics in the data plane using small memory footprint?". In this paper, we answer this question in affirmative for an important case of quantiles. We introduce QPipe, the first quantiles sketching algorithm that can be implemented entirely in the data plane. Our main technical contribution is an on-the-plane implementation of a variant of SweepKLL [27] algorithm. Specifically, we give novel implementations of argmin(), the major building block of SweepKLL which are usually not supported in the data plane of the commodity switch. We prototype QPipe in P4 and compare its performance with a sampling-based baseline. Our evaluations demonstrate 10× memory reduction for a fixed approximation error and 90× error improvement for a fixed amount of memory. We conclude that QPipe can be an attractive alternative to sampling-based methods. Nikita Ivkin, Zhuolong Yu, Vladimir Braverman, Xin Jin 0008 |
CoNEXT | 4 |
| 2019 | Flash: efficient dynamic routing for offchain networksabstractOffchain networks emerge as a promising solution to address the scalability challenge of blockchain. Participants make payments through offchain networks instead of committing transactions on-chain. Routing is critical to the performance of offchain networks. Existing solutions use either static routing with poor performance or dynamic routing with high overhead to obtain the dynamic channel balance information. In this paper, we propose Flash, a new dynamic routing solution that leverages the unique transactions characteristics in offchain networks to strike a better tradeoff between path optimality and probing overhead. By studying the traces of real offchain networks, we find that the payment sizes are heavy-tailed, and most payments are highly recurrent. Flash thus differentiates the treatment of elephant payments from that of mice payments. It uses a modified max-flow algorithm for elephant payments to find paths with sufficient capacity, and strategically routes the payment across paths to minimize the transaction fees. Mice payments are sent directly by looking up a routing table with a few precomputed paths to reduce probing overhead. Testbed experiments and trace-driven simulations show that Flash improves the success volume of payments by up to 2.3x compared to the state-of-the-art routing algorithm. Peng Wang 0070, Hong Xu 0001, Xin Jin 0008, Tao Wang 0088 |
CoNEXT | 3 |
| 2019 | DistCache: Provable Load Balancing for Large-Scale Storage Systems with Distributed Caching
Zaoxing Liu, Zhihao Bai, Zhenming Liu, Changhoon Kim, Vladimir Braverman, Xin Jin 0008, Ion Stoica |
FAST | 7 |
| 2019 | Neural packet classificationabstractPacket classification is a fundamental problem in computer networking. This problem exposes a hard tradeoff between the computation and state complexity, which makes it particularly challenging. To navigate this tradeoff, existing solutions rely on complex hand-tuned heuristics, which are brittle and hard to optimize. Eric Liang, Xin Jin 0008, Ion Stoica |
SIGCOMM | 3 |
| 2019 | DistCache: Provable Load Balancing for Large-Scale Storage Systems with Distributed Caching
Zaoxing Liu, Zhihao Bai, Zhenming Liu, Changhoon Kim, Vladimir Braverman, Xin Jin 0008, Ion Stoica |
USENIX ATC | 7 |
| 2019 | Distributed Subgraph Matching on Timely DataflowabstractRecently there emerge many distributed algorithms that aim at solving subgraph matching at scale. Existing algorithm-level comparisons failed to provide a systematic view of distributed subgraph matching mainly due to the intertwining of strategy and optimization. In this paper, we identify four strategies and three general-purpose optimizations from representative state-of-the-art algorithms. We implement the four strategies with the optimizations based on the common Timely dataflow system for systematic strategy-level comparison. Our implementation covers all representative algorithms. We conduct extensive experiments for both unlabelled matching and labelled matching to analyze the performance of distributed subgraph matching under various settings, which is finally summarized as a practical guide. Longbin Lai, Zhengyi Yang 0001, Xin Jin 0008, Zhengmin Lai, Ran Wang 0008, Kongzhang Hao, Xuemin Lin 0001, Lu Qin 0001, Wenjie Zhang 0001, Ying Zhang 0001, Zhengping Qian, Jingren Zhou 0001 |
Proc. VLDB Endow. | 4 |
| 2019 | Harmonia: Near-Linear Scalability for Replicated Storage with In-Network Conflict DetectionabstractDistributed storage employs replication to mask failures and improve availability. However, these systems typically exhibit a hard tradeoff between consistency and performance. Ensuring consistency introduces coordination overhead, and as a result the system throughput does not scale with the number of replicas. We present Harmonia, a replicated storage architecture that exploits the capability of new-generation programmable switches to obviate this tradeoff by providing near-linear scalability without sacrificing consistency. To achieve this goal, Harmonia detects read-write conflicts in the network, which enables any replica to serve reads for objects with no pending writes. Harmonia implements this functionality at line rate, thus imposing no performance overhead. We have implemented a prototype of Harmonia on a cluster of commodity servers connected by a Barefoot Tofino switch, and have integrated it with Redis. We demonstrate the generality of our approach by supporting a variety of replication protocols, including primary-backup, chain replication, Viewstamped Replication, and NOPaxos. Experimental results show that Harmonia improves the throughput of these protocols by up to 10 x for a replication factor of 10, providing near-linear scalability up to the limit of our testbed. Zhihao Bai, Jialin Li 0001, Ellis Michael, Dan R. K. Ports, Ion Stoica, Xin Jin 0008 |
Proc. VLDB Endow. | 7 |
| 2018 | DumbNet: a smart data center network fabric with dumb switchesabstractToday's data center networks have already pushed many functions to hosts. A fundamental question is how to divide functions between network and software. We present DumbNet, a new data center network architecture with no state in switches. DumbNet switches have no forwarding tables, no state, and thus require no configurations. Almost all control plane functions are pushed to hosts: they determine the entire path of a packet and then write the path as tags in the packet header. Switches only need to examine the tags to forward packets and monitor the port state. We design a set of host-based mechanisms to make the new architecture viable, from network bootstrapping and topology maintenance to network routing and failure handling. We build a prototype with 7 switches and 27 servers, as well as an FPGA-based switch. Extensive evaluations show that DumbNet achieves performance comparable to traditional networks, supports application-specific extensions like flowlet-based traffic engineering, and stays extremely simple and easy-to-manage. Da Wei, Ziheng Song, Ruihan Wu, Xin Jin 0008, Wei Xu 0005 |
EuroSys | 7 |
| 2018 | Proactive Video Push for Optimizing Bandwidth Consumption in Hybrid CDN-P2P VoD SystemsabstractDecentralizing content delivery to edge devices has become a popular solution for saving the bandwidth consumption of CDN when the CDN bandwidth is expensive. One successful realization is the hybrid CDN-P2P VoD system, where a client is allowed to request video content from a number of seeds (seed clients) in the P2P network. However, the seed scarcity problem may arise for a video resource when there are an insufficient number of seeds to satisfy requests to the video. To alleviate this problem, many commercial VoD systems have employed a video push mechanism that directly sends the recent scarce video resources to randomly-chosen seeds to serve more requests. However, the current video push mechanism fails to consider which videos will become scarce in the future, or differentiate the uploading capability of different seeds. In this paper, we propose Proactive-Push, a video push mechanism that lowers the bandwidth consumption of CDN by predicting future scarce videos and proactively sending them to competent seeds with strong uploading capabilities. Proactive-Push trains neural network models to correctly predict 80% of future scarce video resources, and identify over 90% of competent seeds. We evaluate Proactive-Push using a trace-driven emulation and a real-world pilot deployment over a commercial VoD system. Results show that Proactive-Push can further reduce the proportion of direct download from CDN by 21%, and save the CDN bandwidth cost at peak time by 18%. Yuanxing Zhang, Chengliang Gao, Yangze Guo, Kaigui Bian, Xin Jin 0008, Zhi Yang 0001, Lingyang Song, Jiangang Cheng, Hu Tuo, Xiaoming Li 0001 |
INFOCOM | 5 |
| 2018 | NetChain: Scale-Free Sub-RTT Coordination
Xin Jin 0008, Nate Foster, Jeongkeun Lee, Robert Soulé, Changhoon Kim, Ion Stoica |
NSDI | 1 |
| 2018 | ASAP: Fast, Approximate Graph Pattern Mining at Scale
Anand Padmanabha Iyer, Zaoxing Liu, Xin Jin 0008, Shivaram Venkataraman, Vladimir Braverman, Ion Stoica |
OSDI | 3 |
| 2018 | AWStream: adaptive wide-area streaming analyticsabstractThe emerging class of wide-area streaming analytics faces the challenge of scarce and variable WAN bandwidth. Non-adaptive applications built with TCP or UDP suffer from increased latency or degraded accuracy. State-of-the-art approaches that adapt to network changes require developer writing sub-optimal manual policies or are limited to application-specific optimizations. Ben Zhang 0003, Xin Jin 0008, Sylvia Ratnasamy, John Wawrzynek, Edward A. Lee |
SIGCOMM | 2 |
| 2017 | Catalyst: Unlocking the Power of Choice to Speed up Network UpdatesabstractSpeeding up network updates is crucial to maintain high agility and to react quickly to network failures. In this paper, we present Ctalyst---a new design to reduce the network update time. We observe that networks offer a power of choice, where there are many equally-good alternative paths that traffic flows can be assigned to, which is facilitated by redundancy in networks. Catalyst exploits this power of choice to assign flows to alternative paths to merge stages in the dependency graph (that captures the update plan), which in turn reduces the total update time. Furthermore, we observe that because of the prevalence of switch stragglers---switches that unexpectedly take longer time to update, simply assigning a flow to a single (shortest) path is not an optimal design as even a single switch straggler can substantially increase the update time. Thus, the second principle in Catalyst is to compute multiple paths for individual flows offline, among which one would be selected at runtime based on temporal switch conditions, in order to enable a fast update. Our evaluation using a load-balancer setting in a data center network shows that Catalyst effectively reduces the total update time by 1.14--2.15x. Rohan Gandhi, Ori Rottenstreich, Xin Jin 0008 |
CoNEXT | 3 |
| 2017 | Competitive analysis for online scheduling in software-defined optical WANabstractModern planetary-scale online services have massive data to transfer over the wide area network (WAN). Due to the tremendous cost of building WANs and the stringent timing requirement of distributed applications, it is critical for network operators to make efficient use of network resources to optimize data transfers. By leveraging software-defined networking (SDN) and reconfigurable optical devices, recent solutions design centralized systems to jointly control the network layer and the optical layer. While these solutions show it is promising to significantly reduce data transfer times by centralized cross-layer control, they do not have any theoretical guarantees on the proposed algorithms. This paper presents approximation algorithms and theoretical analysis for the online transfer scheduling problem over optical WANs. The goal of the scheduling problem is to minimize the makespan (the time to finish all transfers) or the total sum of completion times. We design and analyze various greedy, online scheduling algorithms that can achieve 3-competitive ratio for makespan, 2-competitive ratio for minimum sum completion time for jobs of unit size, and 3α-competitive ratio for jobs of arbitrary transfer size and each node having degree constraint d, where α = 1 when d = 1 and α = 1.86 when d ≥ 2. We also evaluated the performance of these algorithms and compared the performance with prior heuristics. Su Jia, Xin Jin 0008, Golnaz Ghasemiesfeh, Jiaxin Ding 0001, Jie Gao 0001 |
INFOCOM | 2 |
| 2017 | SketchVisor: Robust Network Measurement for Software Packet ProcessingabstractNetwork measurement remains a missing piece in today's software packet processing platforms. Sketches provide a promising building block for filling this void by monitoring every packet with fixed-size memory and bounded errors. However, our analysis shows that existing sketch-based measurement solutions suffer from severe performance drops under high traffic load. Although sketches are efficiently designed, applying them in network measurement inevitably incurs heavy computational overhead. Qun Huang 0001, Xin Jin 0008, Patrick P. C. Lee, Runhui Li, Lu Tang 0004, Yi-Chao Chen 0001, Gong Zhang 0001 |
SIGCOMM | 2 |
| 2017 | NetCache: Balancing Key-Value Stores with Fast In-Network CachingabstractWe present NetCache, a new key-value store architecture that leverages the power and flexibility of new-generation programmable switches to handle queries on hot items and balance the load across storage nodes. NetCache provides high aggregate throughput and low latency even under highly-skewed and rapidly-changing workloads. The core of NetCache is a packet-processing pipeline that exploits the capabilities of modern programmable switch ASICs to efficiently detect, index, cache and serve hot key-value items in the switch data plane. Additionally, our solution guarantees cache coherence with minimal overhead. We implement a NetCache prototype on Barefoot Tofino switches and commodity servers and demonstrate that a single switch can process 2+ billion queries per second for 64K items with 16-byte keys and 128-byte values, while only consuming a small portion of its hardware resources. To the best of our knowledge, this is the first time that a sophisticated application-level functionality, such as in-network caching, has been shown to run at line rate on programmable switches. Furthermore, we show that NetCache improves the throughput by 3-10x and reduces the latency of up to 40% of queries by 50%, for high-performance, in-memory key-value stores. Xin Jin 0008, Robert Soulé, Jeongkeun Lee, Nate Foster, Changhoon Kim, Ion Stoica |
SOSP | 1 |
| 2016 | Increasing large-scale data center capacity by statistical power controlabstractGiven the high cost of large-scale data centers, an important design goal is to fully utilize available power resources to maximize the computing capacity. In this paper we present Ampere, a novel power management system for data centers to increase the computing capacity by over-provisioning the number of servers. Instead of doing power capping that degrades the performance of running jobs, we use a statistical control approach to implement dynamic power management by indirectly affecting the workload scheduling, which can enormously reduce the risk of power violations. Instead of being a part of the already over-complicated scheduler, Ampere only interacts with the scheduler with two basic APIs. Instead of power control on the rack level, we impose power constraint on the row level, which leads to more room for over provisioning. Guosai Wang, Shuhao Wang, Weisong Shi, Yinghang Zhu, Dianming Hu, Longbo Huang, Xin Jin 0008, Wei Xu 0005 |
EuroSys | 9 |
| 2016 | Optimizing Bulk Transfers with Software-Defined Optical WANabstractBulk transfer on the wide-area network (WAN) is a fundamental service to many globally-distributed applications. It is challenging to efficiently utilize expensive WAN bandwidth to achieve short transfer completion time and meet mission-critical deadlines. Advancements in software-defined networking (SDN) and optical hardware make it feasible and beneficial to quickly reconfigure optical devices in the optical layer, which brings a new opportunity for traffic management on the WAN. Xin Jin 0008, Da Wei, Siming Li, Jie Gao 0001, Guangzhi Li, Wei Xu 0005, Jennifer Rexford |
SIGCOMM | 1 |
| 2015 | CoVisor: A Compositional Hypervisor for Software-Defined Networks
Xin Jin 0008, Jennifer Gossels, Jennifer Rexford, David Walker 0001 |
NSDI | 1 |
| 2014 | Dynamic scheduling of network updatesabstractWe present Dionysus, a system for fast, consistent network updates in software-defined networks. Dionysus encodes as a graph the consistency-related dependencies among updates at individual switches, and it then dynamically schedules these updates based on runtime differences in the update speeds of different switches. This dynamic scheduling is the key to its speed; prior update methods are slow because they pre-determine a schedule, which does not adapt to runtime conditions. Testbed experiments and data-driven simulations show that Dionysus improves the median update speed by 53--88% in both wide area and data center networks compared to prior methods. Xin Jin 0008, Hongqiang Harry Liu, Rohan Gandhi, Srikanth Kandula, Ratul Mahajan, Ming Zhang 0005, Jennifer Rexford, Roger Wattenhofer |
SIGCOMM | 1 |
| 2013 | SoftCell: scalable and flexible cellular core network architectureabstractCellular core networks suffer from inflexible and expensive equipment, as well as from complex control-plane protocols. To address these challenges, we present SoftCell, a scalable architecture that supports fine-grained policies for mobile devices in cellular core networks, using commodity switches and servers. SoftCell enables operators to realize high-level service policies that direct traffic through sequences of middleboxes based on subscriber attributes and applications. To minimize the size of the forwarding tables, SoftCell aggregates traffic along multiple dimensions---the service policy, the base station, and the mobile device---at different switches in the network. Since most traffic originates from mobile devices, SoftCell performs fine-grained packet classification at the access switches, next to the base stations, where software switches can easily handle the state and bandwidth requirements. SoftCell guarantees that packets belonging to the same connection traverse the same sequence of middleboxes in both directions, even in the presence of mobility. We demonstrate that SoftCell improves the scalability and flexibility of cellular core networks by analyzing real LTE workloads, performing micro-benchmarks on our prototype controller as well as large-scale simulations. Xin Jin 0008, Li Erran Li, Laurent Vanbever, Jennifer Rexford |
CoNEXT | 1 |
| 2013 | Intra-data-center traffic engineering with ensemble routingabstractToday's data centers are shared among multiple tenants running a wide range of applications. These applications require a network with a scalable and robust layer-2 network management solution that enables load-balancing and QoS provisioning. Ensemble routing was proposed to achieve management scalability and robustness by using Virtual Local Area Networks (VLANs) and operating on the granularity of flow ensembles, i.e. group of flows. The key challenge of intra-data-center traffic engineering with ensemble routing is the combinatorial optimization of VLAN assignment, i.e., optimally assigning flow ensembles to VLANs to achieve load balancing and low network costs. Based on the Markov approximation framework, we solve the VLAN assignment problem with a general objective function and arbitrary network topologies by designing approximation algorithms with close-to-optimal performance guarantees. We study several properties of our algorithms, including performance optimality, perturbation bound, convergence of algorithms and impacts of algorithmic parameter choices. Then we extend these results to variants of VLAN assignment problem, including interaction with TCP congestion and QoS considerations. We validate our analytical results by conducting extensive numerical experiments. The results show that our algorithms can be tuned to meet different temporal constraints, incorporate fine-grained traffic management, overcome traffic measurement limitations, and tolerate imprecise and incomplete traffic matrices. Ziyu Shao, Xin Jin 0008, Wenjie Jiang 0001, Minghua Chen 0001, Mung Chiang |
INFOCOM | 2 |
| 2011 | A study of the VANET connectivity by percolation theoryabstractAs deploying Vehicular Ad Hoc NETworks (VANETs) costs large amounts of resources, it is crucial that governments and companies make a thorough estimation and comparison of the benefits and the costs. The network connectivity is an important factor we should take care of, because it can greatly affect the performance of VANETs and further affect how much we can benefit from VANETs. We use percolation theory to analyze the connectivity of VANETs. Through theoretical deduction, we discover the quantitative relationship among network connectivity, vehicle density and transmission range. We show that there is a jump of the network connectivity when vehicle density or transmission range is big enough. Simulations conducted in a large scenario validate our theoretical results. Our results have great meanings in the deployment of VANETs in real world. Given vehicle density, our theorem can be used to calculate the minimum transmission range to achieve good network connectivity. As a large transmission range can cause serious collisions in wireless links, it is a tradeoff to choose a proper transmission range. Our analysis can give a hand to this tradeoff and guide the deployment of VANETs in real world. Xin Jin 0008, Weijie J. Su |
CCNC | 1 |
| 2011 | Relative Link Quality Assessment and Hybrid Routing Scheme for Wireless Mesh NetworksabstractWe present a link capacity estimation model and a hybrid routing scheme for wireless mesh networks which aims at improving throughput. The multi-source interference between links and frame capture behaviors of nodes are analyzed. Our model provides estimated values which can indicate the relative quality of wireless links. And we present a routing metric which combines multiple factors based on the estimated values. The proposed routing scheme which is independent of special routing protocols adopts different strategies for the traffic of accessing the Internet and the end-to-end traffic. Besides, we introduce a new gateway-assisted routing technique to further improve the performance. Simulation results show that the estimation model provides accurate estimated values enough for route selection, and the routing scheme improves end-to-end throughput by 185% compared with legacy AODV[14] protocol. ChaoYi Bian, Xin Jin 0008, Xiaoming Li 0001, Wei Yan 0007 |
ICC | 2 |
| 2011 | Quantitative Analysis of the VANET Connectivity: Theory and ApplicationabstractNowadays, Vehicular Ad hoc NETworks (VANETs) attract more and more attentions both from academia and industry. Although it has achieved much success in the research field, large-scale deployments of VANETs are still lacking. One barrier facing us is that both governments and companies are doubtful about the performance of VANETs in large-scale deployment in real world. The connectivity of VANETs is a factor closely related to the performance of VANETs. We give a thorough theoretical analysis of the VANET connectivity using bond percolation model and Bollobas model in different scenarios. We discover the quantitative relationship among network connectivity, vehicle density and transmission range. Given the vehicle density, we can calculate the minimum transmission range to achieve good network connectivity. Simulations conducted in a large scenario validate our analysis. Our results not only give us insights about the properties of the network topology, but also have great meanings in real world, which can guide the deployment of VANETs. Xin Jin 0008, Weijie J. Su, Wei Yan 0007 |
VTC Spring | 1 |