Jialin Li 0001

dblp:75/4924-1 · DBLP profile ↗
← Back
42ranked-venue papers
6as first author
31since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 17 · 16 since 2021Systems, architecture and hardware · 11 · 1 first-author · 6 since 2021Software engineering, systems software and programming languages · 10 · 5 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 NutCracker: A Compilation Framework for Hybrid DPU Architectures
abstract
SoC-based SmartNICs, or data processing units (DPUs), are becoming a viable option for offloading infrastructure services. However, developers need deep hardware-level knowledge to fully utilize the available hardware accelerators on a target DPU. Hardware heterogeneity also makes porting across DPUs a formidable task. In this work, we propose a new compiler framework, NutCracker. Using NutCracker, programmers develop DPU applications using high-level target-independent languages. NutCracker applies a two-stage compilation process. It first performs progressive lowering to convert the source program to candidate intermediate representations (IRs) of the target DPU. Next, NutCracker applies cost-guided mapping optimization using equality saturation to select a final implementation on the target hardware with a configurable optimization goal. Evaluated on eight applications, NutCracker reduces developer effort by nearly 90% while delivering performance within 3% of handcrafted implementations for seven of the workloads. Moreover, its compilation time is comparable to standard toolchains such as GCC.
Haifeng Sun 0004, Antoine Kaufmann, Jialin Li 0001
EuroSys4
2026 Lemonshark: Asynchronous DAG-BFT With Early Finality
Michael Yiqing Hu, Alvin Hong Yao Yan, Xiang Liu 0017, Jialin Li 0001
NSDI5
2026 Cortex: Achieving Low-Latency, Cost-Efficient Remote Data Access For LLM via Semantic-Aware Knowledge Caching
Chaoyi Ruan, Chao Bi, Ziji Shi, Jialin Li 0001
NSDI6
2026 Libra: Flexible Request Partitioning and Scheduling for Serving Unbalanced and Dynamic LLM Workloads
Chaoyi Ruan, Yinhe Chen, Dongqi Tian, Yandong Shi, Jialin Li 0001
NSDI6
2026 PD3: Prefetching Data with DPUs for Disaggregated Memory
Sidharth Sankhe, Felix Zhang, Umayrah Chonee, Sherman Lim, Jiasheng Hu, Jialin Li 0001, Qizhen Zhang 0001
NSDI6
2026 HyperEdge: An Edge CDN Infrastructure for Cost Efficient Video Streaming
Dehui Wei, Jiao Zhang 0002, Zhichen Xue, Yajie Peng, Xiaofei Pang, Jialin Li 0001
NSDI9
2026 Brief Announcement: Amortized Asynchronous Byzantine Reliable Broadcast with Optimal Resilience
abstract
Byzantine Reliable Broadcast (BRB) is a fundamental primitive in distributed computing and cryptographic systems; reducing the communication cost of BRB thus remains an important research direction. However most existing work either focus strictly on the synchronous network model or forgo optimal resilience (f<⌊n3⌋) in the face of asynchrony.
Michael Yiqing Hu, Alvin Hong Yao Yan, Jialin Li 0001
PODC3
2026 Capybara: Dynamic Load Balancing with Microsecond-Scale TCP Migration
abstract
Layer-4 load balancers are a popular solution to high tail latencies but perform poorly under unpredictable skewed workloads because they statically assign connections to servers. We present Capybara, a new load balancer architecture that enables dynamic rebalancing of established connections. Capybara divides load balancing responsibility into a fast L4 load balancer, a host-switch co-designed connection migration protocol, and a transport interface for application-level connection state migration. Capybara leverages two trends - programmable switches and kernel-bypass - to efficiently implement connection migration without disruption, while maintaining transparency to clients. Under realistic workloads, Capybara achieves up to 149× lower tail latency and more than 2× higher throughput for scale-out services compared to state-of-the-art load balancing approaches.
Inho Choi, Nimish Wadekar, Guangda Sun, Raj Joshi, Joshua Fried, Omar S. Navarro Leija, Dan R. K. Ports, Irene Zhang, Jialin Li 0001
SIGCOMM9
2026 POSTER: VibeNIC: Toward LLM-driven Agile Development of FPGA SmartNICs
Jialin Li 0001
SIGCOMM2
2026 Horizon: A Hyper-Edge Observability Engine for Live Streaming Networks
abstract
Live streaming services power mainstream real-time interactions on top of dedicated live streaming networks (LiveNets). Yet making LiveNets reliable at scale is challenging: failures arise on the userfacing delivery path and within streaming protocol and application logic, so operators need both continuous runtime monitoring to detect and localize incidents quickly and proactive preflight testing to exercise changes under representative environments and sustained playback behavior. Meeting these goals hinges on the right vantage point: the observability workflow must traverse the same network paths and delivery stacks as users while remaining controllable and non-intrusive. We present Horizon, which leverages near-user, provider-managed hyper-edge devices and orchestrates them into a shared fleet that supports both always-on monitoring and customizable, scenario-driven validation. Horizon has been deployed in production for over three years; in 2025, it identified 2,000+ major network incidents using 100,000+ hyper-edge agents.
Daqian Ding, Shixian Guo, Zhendong Xie, Aifang Xu, Changqian Wang, Kefei Liu 0004, Jialin Li 0001, Yunming Xiao, Heming Cui, Yiming Qiu 0001
SIGCOMM9
2026 Turbo: Efficiently Serving Long-Context Large Language Models with In-Network Aggregation
abstract
LLM supporting long contexts faces a critical memory bottleneck due to the linear growth of KV cache. Distributing the storage across multiple GPUs alleviates this burden but introduces significant communication overhead or traffic incast, especially during the decoding phase. We propose Turbo, a first-of-its-kind in-network aggregation system that accelerates long-context inference by offloading query broadcast and attention aggregation to switches. We address three key challenges to map complex attention mechanisms onto restricted switch hardware: (i) To bypass the switch's inability to buffer global states or perform complex operations, we devise online table-based aggregation, which decomposes global reduction into pairwise operations and approximates nonlinear functions via lookup tables. (ii) To circumvent the restriction on retroactive state access in RMT pipelines, we introduce a rolling forward scheme that propagates states to enable cross-stage updates. (iii) To mitigate aggregation stragglers caused by topology-induced load imbalance, we construct a load-aware aggregation tree that optimizes workload distribution. Evaluations on a Tofino2-based testbed show that Turbo reduces end-to-end inference latency by up to 37%. Large-scale simulations on NS-3 demonstrate that Turbo significantly outperforms state-of-the-art baselines in both inference latency and network traffic reduction with negligible accuracy loss.
Ying Wan 0001, Yuchen Xu 0003, Chuwen Zhang, Yingsheng Huang, Wenquan Xu, Jialin Li 0001, Mingwei Xu 0001, Wenfei Wu, Congcong Miao
SIGCOMM7
2026 HyStream: A Hybrid System for Application Streaming via Predictive Delivery and Sequence-Linearized Caching
abstract
Traditional application delivery requires full local installation, introducing persistent security risks from outdated software and imposing significant download delays. While advances in network bandwidth and latency have made remote content delivery more viable, existing dynamic loading mechanisms, such as network filesystems, often remain constrained by performance bottlenecks. Worse still, these solutions degrade sharply under variable or weak connectivity, where untimely code delivery can stall execution altogether. We propose HyStream, a hybrid application streaming system that combines predictive remote delivery with local sequence-linearized caching to sustain responsive and robust execution without requiring installation. HyStream addresses three key challenges: (1) maintaining microsecond-level latency comparable to local storage; (2) bridging the semantic gap between stateless remote storage and stateful execution; and (3) mitigating the limitations of purely network-based solutions under degraded connectivity. To achieve this, HyStream integrates three core components: a dual-mode transmission mechanism that decouples synchronous demand-driven requests from asynchronous speculative prefetching; a thread-aware Markov-chain model that captures fine-grained, concurrent access patterns for accurate prediction; and a sequence-linearized cache that persists streamed blocks in predicted execution order to support deterministic fallback behavior. Together, these components transform irregular, latency-sensitive I/O into efficient structured access that masks network variability. Evaluation shows HyStream delivers near-native performance across diverse networks. On mobile devices, it achieves 16–29% better per-page access latency than local UFS3.1, even over variable Wi-Fi connectivity. On desktops, it typically sustains startup overheads below 30% relative to local NVMe. Under variable and degraded network conditions, the sequence-linearized cache increasingly serves execution-critical accesses, rendering application performance largely insensitive to network latency and jitter within intra-city and inter-city deployments.
Sheng Yue 0001, Xiang Liu 0017, Yongjian Fu 0004, Jialin Li 0001
IEEE Trans. Netw.6
2025 DPDPU: Data Processing with DPUs
Jason Hu, Philip A. Bernstein, Jialin Li 0001, Qizhen Zhang 0001
CIDR3
2025 Efficient Partitioning Vision Transformer on Edge Devices for Distributed Inference
abstract
Deep learning models are increasingly utilized on resource-constrained edge devices for real-time data analytics. Recently, Vision Transformer and their variants have shown exceptional performance in various computer vision tasks. However, their substantial computational requirements and low inference latency create significant challenges for deploying such models on resource-constrained edge devices. To address this issue, we propose a novel framework, ED-ViT, which is designed to efficiently split and execute complex Vision Transformers across multiple edge devices. Our approach involves partitioning Vision Transformer models into several sub-models, while each dedicated to handling a specific subset of data classes. To further reduce computational overhead and inference latency, we introduce a class-wise pruning technique that decreases the size of each sub-model. Through extensive experiments conducted on five datasets using three model architectures and actual implementation on edge devices, we demonstrate that our method significantly cuts down inference latency on edge devices and achieves a reduction in model size by up to 28.9 times and 34.1 times, respectively, while maintaining test accuracy comparable to the original Vision Transformer. Additionally, we compare ED-ViT with two state-of-the-art methods that deploy CNN and SNN models on edge devices, evaluating metrics such as accuracy, inference time, and overall model size. Our comprehensive evaluation underscores the effectiveness of the proposed ED-ViT framework.
Xiang Liu 0017, Yijun Song, Xia Li 0005, Huiying Lan, Linshan Jiang, Jialin Li 0001
ICDCS8
2025 TAPAS: Fast and Automatic Derivation of Tensor Parallel Strategies for Large Neural Networks
abstract
Tensor parallelism is an essential technique for distributed training of large neural networks. However, automatically determining an optimal tensor parallel strategy is challenging due to the gigantic search space, which grows exponentially with model size and tensor dimension. This prohibits the adoption of auto-parallel systems on larger models.
Ziji Shi, Ang Wang, Jie Zhang 0135, Chencan Wu, Yong Li 0045, Xiaokui Xiao, Wei Lin 0016, Jialin Li 0001
ICPP9
2025 One-shot Federated Learning Methods: A Practical Guide
abstract
One-shot Federated Learning (OFL) is a distributed machine learning paradigm that constrains client-server communication to a single round, addressing privacy and communication overhead issues associated with multiple rounds of data exchange in traditional Federated Learning (FL). OFL demonstrates the practical potential for integration with future approaches that require collaborative training models, such as large language models (LLMs). However, current OFL methods face two major challenges: data heterogeneity and model heterogeneity, which result in subpar performance compared to conventional FL methods. Worse still, despite numerous studies addressing these limitations, a comprehensive summary is still lacking. To address these gaps, this paper presents a systematic analysis of the challenges faced by OFL and thoroughly reviews the current methods. We also offer an innovative categorization method and analyze the trade-offs of various techniques. Additionally, we discuss the most promising future directions and the technologies that should be integrated into the OFL field. This work aims to provide guidance and insights for future research.
Xiang Liu 0017, Zhenheng Tang, Xia Li 0005, Yijun Song, Sijie Ji, Bo Han 0003, Linshan Jiang, Jialin Li 0001
IJCAI9
2025 QoE-Optimized MultiPath Scheduling for Video Services in Large-Scale Peer-to-Peer CDNs
abstract
Video content providers such as Douyin implement Peer-to-Peer Content Delivery Networks (PCDNs) to reduce the costs associated with Content Delivery Networks (CDNs) while still maintaining optimal user-perceived quality of experience (QoE). PCDNs rely on the remaining resources of edge devices, such as edge access devices and hosts, to store and distribute data with a Multiple-Server-to-One-Client (MS2OC) communication pattern. MS2OC parallel transmission pattern suffers from severe data out-of-order issues. PCDNs offer significant cost savings by using multiple low-cost edge devices. However, due to its unique characteristics, including pull-based streaming transmission, many heterogeneous paths, and large receiving buffers, directly applying existing schedulers designed for Multipath TCP (MPTCP) to PCDN fails to meet the two goals of high aggregate bandwidth and low end-to-end delivery latency. To tackle this issue, we provide a detailed overview of Douyin’s self-developed PCDN video transmission system and introduce the first QoE-enhanced packet-level scheduler for PCDN systems, named Pscheduler. Pscheduler evaluates path quality with a congestion-control-decoupled algorithm and employs our proposed path-pick-packet method for data distribution, ensuring a smooth video playback experience. Additionally, we propose a redundant transmission algorithm to enhance task download speeds for segmented video transmission. Our extensive online A/B tests, involving 100,000 Douyin users generating tens of millions of video data points, demonstrate that Pscheduler achieves an average improvement of 60% in goodput, a 20% reduction in data delivery waiting time, and a 30% reduction in rebuffering rates. Furthermore, we conducted simulation experiments that further validate the effectiveness of Pscheduler, confirming its improvements in performance metrics under various network conditions.
Dehui Wei, Jiao Zhang 0002, Xiang Liu 0017, Zhichen Xue, Tao Huang 0005, Linshan Jiang, Jialin Li 0001
IEEE J. Sel. Areas Commun.8
2025 Hare: A Systematic Framework for Efficient and Generally Automatic Hotspot Offloading on Programmable Switches
abstract
Switch-based hotspot offloading is a trendy solution for latency-sensitive applications to achieve high system throughput with an acceptable P99 query response latency. However, due to the varying object sizes, dynamic workloads, and complex query-processing functions of the latency-sensitive applications, existing switch-based dynamic hotspot offloading approaches struggle to handle these applications effectively. This is mainly because of their inefficient switch resource utilization and non-generalizable hotspot offloading designs. So we propose Hare, a systematic framework that consists of three techniques to address these issues. First, Hare uses a MAT-based cross-stage structure to store and perform hit-checks for large hotspots on the switch data plane. Second, Hare uses a switch-server co-offloading mechanism to support fast and precise offloading. Third, Hare is designed to enable generally automatic offloading by decoupling application-related query processing with hotspot offloading. Compared to the state-of-the-art approaches, Hare supports$8.86\times \sim 9.97\times $larger hotspot size, achieves$1.27 \times \sim 6.61 \times $higher system throughput, and can recover the system throughput and the P99 query response latency within 8s.
Xueying Zhu, Yingtao Li 0001, Xiang Li 0205, Jialin Li 0001, Zeke Wang
IEEE Trans. Netw.4
2024 ParaGAN: A Scalable Distributed Training Framework for Generative Adversarial Networks
abstract
Recent advances in Generative Artificial Intelligence have fueled numerous applications, particularly those involving Generative Adversarial Networks (GANs), which are essential for synthesizing realistic photos and videos. However, efficiently training GANs remains a critical challenge due to their computationally intensive and numerically unstable nature. Existing methods often require days or even weeks for training, posing significant resource and time constraints.
Ziji Shi, Jialin Li 0001, Yang You 0001
SoCC2
2024 FedLPA: One-shot Federated Learning with Layer-Wise Posterior Aggregation
abstract
Efficiently aggregating trained neural networks from local clients into a global model on a server is a widely researched topic in federated learning. Recently, motivated by diminishing privacy concerns, mitigating potential attacks, and reducing communication overhead, one-shot federated learning (i.e., limiting client-server communication into a single round) has gained popularity among researchers. However, the one-shot aggregation performances are sensitively affected by the non-identical training data distribution, which exhibits high statistical heterogeneity in some real-world scenarios. To address this issue, we propose a novel one-shot aggregation method with layer-wise posterior aggregation, named FedLPA. FedLPA aggregates local models to obtain a more accurate global model without requiring extra auxiliary datasets or exposing any private label information, e.g., label distributions. To effectively capture the statistics maintained in the biased local datasets in the practical non-IID scenario, we efficiently infer the posteriors of each layer in each local model using layer-wise Laplace approximation and aggregate them to train the global parameters. Extensive experimental results demonstrate that FedLPA significantly improves learning performance over state-of-the-art methods across several metrics.
Xiang Liu 0017, Liangxi Liu, Feiyang Ye 0004, Yunheng Shen, Xia Li 0005, Linshan Jiang, Jialin Li 0001
NeurIPS7
2023 Hydra: Serialization-Free Network Ordering for Strongly Consistent Distributed Applications
Inho Choi, Ellis Michael, Dan R. K. Ports, Jialin Li 0001
NSDI5
2023 Network Load Balancing with In-network Reordering Support for RDMA
abstract
Remote Direct Memory Access (RDMA) is widely used in high-performance computing (HPC) and data center networks. In this paper, we first show that RDMA does not work well with existing load balancing algorithms because of its traffic flow characteristics and assumption of in-order packet delivery. We then propose ConWeave, a load balancing framework designed for RDMA. The key idea of ConWeave is that with the right design, it is possible to perform fine granularity rerouting and mask the effect of out-of-order packet arrivals transparently in the network datapath using a programmable switch. We have implemented ConWeave on a Tofino2 switch. Evaluations show that ConWeave can achieve up to 42.3% and 66.8% improvement for average and 99-percentile FCT, respectively compared to the state-of-the-art load balancing algorithms.
Cha Hwan Song, Xin Zhe Khooi, Raj Joshi, Inho Choi, Jialin Li 0001, Mun Choon Chan
SIGCOMM5
2023 NeoBFT: Accelerating Byzantine Fault Tolerance Using Authenticated In-Network Ordering
abstract
Mission critical systems deployed in data centers today are facing more sophisticated failures. Byzantine fault-tolerant (BFT) protocols are capable of masking these types of failures, but are rarely deployed due to their performance cost and complexity. In this work, we propose a new approach to designing high performance BFT protocols in data centers. By re-examining the ordering responsibility between the network and the BFT protocol, we advocate a new abstraction offered by the data center network infrastructure. Concretely, we design a new authenticated ordered multicast primitive (aom) that provides transferable authentication and non-equivocation guarantees. Feasibility of the design is demonstrated by two hardware implementations of aom- one using HMAC and the other using public key cryptography for authentication - on new-generation programmable switches. We then co-design a new BFT protocol, NeoBFT, that leverages the guarantees of aom to eliminate cross-replica coordination and authentication in the common case. Evaluation results show that NeoBFT outperforms state-of-the-art protocols on both latency and throughput metrics by a wide margin, demonstrating the benefit of our new network ordering abstraction for BFT systems.
Guangda Sun, Mingliang Jiang, Xin Zhe Khooi, Jialin Li 0001
SIGCOMM5
2023 TEE-based General-purpose Computational Backend for Secure Delegated Data Processing
abstract
The increasing prevalence of data breaches necessitates robust data protection measures in computational tasks. Secure computation outsourcing (SCO) presents a viable solution by safeguarding the confidentiality of inputs and outputs in data processing without disclosure. Nonetheless, this approach assumes the existence of a trustworthy coordinator to orchestrate and oversee the process, typically implying that data owners must fulfill this role themselves. In this paper, we consider secure delegated data processing (SDDP), an expanded data processing scenario wherein data owners simply delegate their data to SDDP providers for subsequent value mining or other downstream applications, eliminating the necessary involvement of data owners or trusted entities to dive into data processing deeply. However, general-purpose SDDP poses significant challenges in permitting the discretionary execution of computational tasks by SDDP providers on sensitive data while ensuring confidentiality. Existing approaches are insufficient to support SDDP in either efficiency or universality. To tackle this issue, we propose TGCB, a TEE-based General-purpose Computational Backend, designed to endow general-purpose computation with SDDP capabilities from an engineering perspective, powered by TEE-based code integrity and data confidentiality. Central to TGCB is the Encryption Programming Language (EPL) that defines computational tasks in SDDP. Specifically, SDDP providers can express arbitrary computable functions as EPL scripts, processed by TGCB's interfaces, securely interpreted and executed in TEE, ensuring data confidentiality throughout the process. As a universal computational backend, TGCB extensively bolsters data security in existing general-purpose computational tasks, allowing data owners to leverage SDDP without privacy concerns.
Mo Sha 0002, Jialin Li 0001, Sheng Wang 0011, Feifei Li 0001, Kian-Lee Tan
Proc. ACM Manag. Data2
2023 P4SGD: Programmable Switch Enhanced Model-Parallel Training on Generalized Linear Models on Distributed FPGAs
abstract
Generalized linear models (GLMs) are a widely utilized family of machine learning models in real-world applications. As data size increases, it is essential to perform efficient distributed training for these models. However, existing systems for distributed training have a high cost for communication and often use large batch sizes to balance computation and communication, which negatively affects convergence. Therefore, we argue for an efficient distributed GLM training system that strives to achieve linear scalability, while keeping batch size reasonably low. As a start, we propose P4SGD, a distributed heterogeneous training system that efficiently trains GLMs through model parallelism between distributed FPGAs and through forward-communication-backward pipeline parallelism within an FPGA. Moreover, we propose a light-weight, latency-centric in-switch aggregation protocol to minimize the latency of the AllReduce operation between distributed FPGAs, powered by a programmable switch. As such, to our knowledge, P4SGD is the first solution that achieves almost linear scalability between distributed accelerators through model parallelism. We implement P4SGD on eight Xilinx U280 FPGAs and a Tofino P4 switch. Our experiments show P4SGD converges up to 6.5X faster than the state-of-the-art GPU counterpart.
Hongjing Huang, Yingtao Li 0001, Jie Sun 0017, Xueying Zhu, Jie Zhang 0081, Jialin Li 0001, Zeke Wang
IEEE Trans. Parallel Distributed Syst.7
2022 Linear-time Temporal Logic guided Greybox Fuzzing
abstract
Software model checking as well as runtime verification are verification techniques which are widely used for checking temporal properties of software systems. Even though they are property verification techniques, their common usage in practice is in "bug finding", that is, finding violations of temporal properties. Motivated by this observation and leveraging the recent progress in fuzzing, we build a greybox fuzzing framework to find violations of Linear-time Temporal Logic (LTL) properties.
Ruijie Meng, Jialin Li 0001, Ivan Beschastnikh, Abhik Roychoudhury
ICSE3
2022 SimBricks: end-to-end network system evaluation with modular simulation
abstract
Full system "end-to-end" measurements in physical testbeds are the gold standard for network systems evaluation but are often not feasible. When physical testbeds are not available we frequently turn to simulation for evaluation. Unfortunately, existing simulators are insufficient for end-to-end evaluation, as they either cannot simulate all components, or simulate them with inadequate detail.
Hejing Li, Jialin Li 0001, Antoine Kaufmann
SIGCOMM2
2022 Linear types for large-scale systems verification
abstract
Reasoning about memory aliasing and mutation in software verification is a hard problem. This is especially true for systems using SMT-based automated theorem provers. Memory reasoning in SMT verification typically requires a nontrivial amount of manual effort to specify heap invariants, as well as extensive alias reasoning from the SMT solver. In this paper, we present a hybrid approach that combines linear types with SMT-based verification for memory reasoning. We integrate linear types into Dafny, a verification language with an SMT backend, and show that the two approaches complement each other. By separating memory reasoning from verification conditions, linear types reduce the SMT solving time. At the same time, the expressiveness of SMT queries extends the flexibility of the linear type system. In particular, it allows our linear type system to easily and correctly mix linear and nonlinear data in novel ways, encapsulating linear data inside nonlinear data and vice-versa. We formalize the core of our extensions, prove soundness, and provide algorithms for linear type checking. We evaluate our approach by converting the implementation of a verified storage system (about 24K lines of code and proof) written in Dafny, to use our extended Dafny. The resulting system uses linear types for 91% of the code and SMT-based heap reasoning for the remaining 9%. We show that the converted system has 28% fewer lines of proofs and 30% shorter verification time overall. We discuss the development overhead in the original system due to SMT-based heap reasoning and highlight the improved developer experience when using linear types.
Jialin Li 0001, Andrea Lattuada 0001, Yi Zhou 0025, Jonathan Cameron, Jon Howell, Bryan Parno, Chris Hawblitzel
Proc. ACM Program. Lang.1
2021 An incremental path towards a safer OS kernel
abstract
Linux has become the de-facto operating system of our age, but its vulnerabilities are a constant threat to service availability, user privacy, and data integrity. While one might scrap Linux and start over, the cost of that would be prohibitive due to Linux's ubiquitous deployment. In this paper, we propose an alternative, incremental route to a safer Linux through proper modularization and gradual replacement module by module. We lay out the research challenges and potential solutions for this route, and discuss the open questions ahead.
Jialin Li 0001, Samantha Miller, Danyang Zhuo, Ang Chen 0001, Jon Howell, Thomas E. Anderson
HotOS1
2021 In-Network Applications: Beyond Single Switch Pipelines
abstract
The emergence of commodity programmable switches have spawned a series of innovations in the network data plane. By making the traditionally stateless network architectures to be stateful, we can realize a diverse set of applications, e.g., networking monitoring, load-balancing, firewalls, entirely in the data plane. On the other hand, many existing in-network applications assume that the underlying switch is single-pipelined, however, in reality, commodity programmable switches are designed with multiple pipelines in mind. While this approach enables high scalability, it has introduced a serious disadvantage: maintaining states across the pipelines is non-trivial. For instance, without involving the control plane it is infeasible to keep track of a request and its response in different pipelines, thereby rendering many in-network proposals impractical.In this paper, we highlight this fundamental limitation that holds back the practical widespread adoption of stateful applications in today’s multi-pipeline switches. By scrutinizing recent in-network approaches, we identify that majority of them cannot operate as they are proposed on multi-pipeline switches. After raising awareness of this inevitable consequence, we discuss a set of possible workarounds for in-network applications to overcome this issue on multi-pipeline switches.
Xin Zhe Khooi, Levente Csikor, Jialin Li 0001, Dinil Mon Divakaran
NetSoft3
2021 Revisiting Heavy-Hitter Detection on Commodity Programmable Switches
abstract
Existing in-network heavy-hitter detection algorithms suffer from several shortcomings. On the one hand, most of the algorithms perform monitoring in intervals and reset the data structures in between; consequently, a notable amount of heavy hitters (HH) spanning across the intervals go undetected. On the other hand, the algorithms consume substantial hardware resources, potentially hindering other data plane functionalities to be integrated on the same device.In this work, we revisit the state-of-the-art in-network approaches in this regard and identify that they fall short in over-coming the aforementioned issues. In particular, we investigate whether it is possible to design a heavy-hitter detection algorithm that provides high accuracy without consuming substantial re-sources, thereby making it feasible to integrate with concurrent applications. To this end, we propose dSketch, a time-decaying algorithm for in-network heavy-hitter detection. Trace-driven simulations and evaluations on the Intel Tofino-based commodity switches show that dSketch significantly improves the detection rate of HHs by 5–10% while being resource- and operation-efficient in contrast to state-of-the-art approaches. Moreover, we show that dSketch can be integrated with standard switch functionalities such as switch. p4 with additional resources spared, offering itself as a compelling solution for switch data plane designers.
Xin Zhe Khooi, Levente Csikor, Jialin Li 0001, Min Suk Kang, Dinil Mon Divakaran
NetSoft3
2020 Meerkat: multicore-scalable replicated transactions following the zero-coordination principle
abstract
Traditionally, the high cost of network communication between servers has hidden the impact of cross-core coordination in replicated systems. However, new technologies, like kernel-bypass networking and faster network links, have exposed hidden bottlenecks in distributed systems.
Adriana Szekeres, Michael J. Whittaker, Jialin Li 0001, Naveen Kr. Sharma, Arvind Krishnamurthy, Dan R. K. Ports, Irene Zhang
EuroSys3
2020 Pegasus: Tolerating Skewed Workloads in Distributed Storage with In-Network Coherence Directories
Jialin Li 0001, Jacob Nelson 0001, Ellis Michael, Xin Jin 0008, Dan R. K. Ports
OSDI1
2019 Harmonia: Near-Linear Scalability for Replicated Storage with In-Network Conflict Detection
abstract
Distributed storage employs replication to mask failures and improve availability. However, these systems typically exhibit a hard tradeoff between consistency and performance. Ensuring consistency introduces coordination overhead, and as a result the system throughput does not scale with the number of replicas. We present Harmonia, a replicated storage architecture that exploits the capability of new-generation programmable switches to obviate this tradeoff by providing near-linear scalability without sacrificing consistency. To achieve this goal, Harmonia detects read-write conflicts in the network, which enables any replica to serve reads for objects with no pending writes. Harmonia implements this functionality at line rate, thus imposing no performance overhead. We have implemented a prototype of Harmonia on a cluster of commodity servers connected by a Barefoot Tofino switch, and have integrated it with Redis. We demonstrate the generality of our approach by supporting a variety of replication protocols, including primary-backup, chain replication, Viewstamped Replication, and NOPaxos. Experimental results show that Harmonia improves the throughput of these protocols by up to 10 x for a replication factor of 10, providing near-linear scalability up to the limit of our testbed.
Zhihao Bai, Jialin Li 0001, Ellis Michael, Dan R. K. Ports, Ion Stoica, Xin Jin 0008
Proc. VLDB Endow.3
2017 Eris: Coordination-Free Consistent Transactions Using In-Network Concurrency Control
abstract
Distributed storage systems aim to provide strong consistency and isolation guarantees on an architecture that is partitioned across multiple shards for scalability and replicated for fault tolerance. Traditionally, achieving all of these goals has required an expensive combination of atomic commitment and replication protocols -- introducing extensive coordination overhead. Our system, Eris, takes a different approach. It moves a core piece of concurrency control functionality, which we term multi-sequencing, into the datacenter network itself. This network primitive takes on the responsibility for consistently ordering transactions, and a new lightweight transaction protocol ensures atomicity.
Jialin Li 0001, Ellis Michael, Dan R. K. Ports
SOSP1
2016 Specifying and Checking File System Crash-Consistency Models
abstract
Applications depend on persistent storage to recover state after system crashes. But the POSIX file system interfaces do not define the possible outcomes of a crash. As a result, it is difficult for application writers to correctly understand the ordering of and dependencies between file system operations, which can lead to corrupt application state and, in the worst case, catastrophic data loss. This paper presents crash-consistency models, analogous to memory consistency models, which describe the behavior of a file system across crashes. Crash-consistency models include both litmus tests, which demonstrate allowed and forbidden behaviors, and axiomatic and operational specifications. We present a formal framework for developing crash-consistency models, and a toolkit, called Ferrite, for validating those models against real file system implementations. We develop a crash-consistency model for ext4, and use Ferrite to demonstrate unintuitive crash behaviors of the ext4 implementation. To demonstrate the utility of crash-consistency models to application writers, we use our models to prototype proof-of-concept verification and synthesis tools, as well as new library interfaces for crash-safe applications.
James Bornholt, Antoine Kaufmann, Jialin Li 0001, Arvind Krishnamurthy, Emina Torlak, Xi Wang 0005
ASPLOS3
2016 Just Say NO to Paxos Overhead: Replacing Consensus with Network Ordering
Jialin Li 0001, Ellis Michael, Naveen Kr. Sharma, Adriana Szekeres, Dan R. K. Ports
OSDI1
2016 Arrakis: The Operating System Is the Control Plane
abstract
Recent device hardware trends enable a new approach to the design of network server operating systems. In a traditional operating system, the kernel mediates access to device hardware by server applications to enforce process isolation as well as network and disk security. We have designed and implemented a new operating system, Arrakis, that splits the traditional role of the kernel in two. Applications have direct access to virtualized I/O devices, allowing most I/O operations to skip the kernel entirely, while the kernel is re-engineered to provide network and disk protection without kernel mediation of every operation. We describe the hardware and software changes needed to take advantage of this new abstraction, and we illustrate its power by showing improvements of 2 to 5 × in latency and 9 × throughput for a popular persistent NoSQL store relative to a well-tuned Linux implementation.
Simon Peter 0001, Jialin Li 0001, Irene Zhang, Dan R. K. Ports, Doug Woos, Arvind Krishnamurthy, Thomas E. Anderson, Timothy Roscoe
ACM Trans. Comput. Syst.2
2015 Designing Distributed Systems Using Approximate Synchrony in Data Center Networks
Dan R. K. Ports, Jialin Li 0001, Vincent Liu 0001, Naveen Kr. Sharma, Arvind Krishnamurthy
NSDI2
2014 Tales of the Tail: Hardware, OS, and Application-level Sources of Tail Latency
abstract
Interactive services often have large-scale parallel implementations. To deliver fast responses, the median and tail latencies of a service's components must be low. In this paper, we explore the hardware, OS, and application-level sources of poor tail latency in high throughput servers executing on multi-core machines.
Jialin Li 0001, Naveen Kr. Sharma, Dan R. K. Ports, Steve D. Gribble
SoCC1
2014 Towards High-Performance Application-Level Storage Management
Simon Peter 0001, Jialin Li 0001, Irene Zhang, Dan R. K. Ports, Thomas E. Anderson, Arvind Krishnamurthy, Mark Zbikowski, Doug Woos
HotStorage2
2014 Arrakis: The Operating System is the Control Plane
Simon Peter 0001, Jialin Li 0001, Irene Zhang, Dan R. K. Ports, Doug Woos, Arvind Krishnamurthy, Thomas E. Anderson, Timothy Roscoe
OSDI2