Jiarong Xing

dblp:217/0857 · DBLP profile ↗
← Back
19ranked-venue papers
11as first author
16since 2021 · last 2026
0009-0006-6163-0569ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 5 first-author · 6 since 2021Systems, architecture and hardware · 5 · 2 first-author · 4 since 2021Security and privacy · 5 · 4 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 BlendServe: Optimizing Offline Inference with Resource-Aware Batching
abstract
Offline batch inference is gaining popularity as a cost-effective solution for latency-insensitive tasks, such as model evaluation and data curation. As the latency objective is highly relaxed, maximizing throughput is the primary goal in offline inference. Previous studies focused solely on throughput optimization within a batch. However, the diverse resource demands (compute-intensive vs. memory-intensive) across a wide range of applications make these approaches less effective, as imbalanced resource demands between batches restrict optimization opportunities.
Yilong Zhao 0002, Shuo Yang 0011, Kan Zhu, Lianmin Zheng, Baris Kasikci, Yifan Qiao 0002, Yang Zhou 0008, Jiarong Xing, Ion Stoica
ASPLOS (2)8
2026 SkyWalker: A Locality-Aware Cross-Region Load Balancer for LLM Inference
abstract
Serving Large Language Models (LLMs) efficiently in multi-region setups remains a challenge. Due to cost and GPU availability concerns, providers typically deploy LLMs in multiple regions using instance with long-term commitments, like reserved instances or on-premise clusters, which are often underutilized due to their region-local traffic handling and diurnal traffic variance. In this paper, we introduce SkyWalker, a multi-region load balancer for LLM inference that aggregates regional diurnal patterns through cross-region traffic handling. By doing so, SkyWalker enables providers to reserve instances based on expected global demand, rather than peak demand in each individual region. Meanwhile, SkyWalker preserves KV-Cache locality and load balancing, ensuring cost efficiency without sacrificing performance. SkyWalker achieves this with a cache-aware cross-region traffic handler and a selective pushing based load balancing mechanism. Our evaluation on real-world workloads shows that it achieves 1.12–2.06× higher throughput and 1.74–6.30× lower latency compared to existing load balancers, while reducing total serving cost by 25%.
Ziming Mao, Jamison Kerney, Ethan J. Jackson, Zhifei Li 0006, Jiarong Xing, Scott Shenker, Ion Stoica
EuroSys6
2025 Remote Direct Code Execution
abstract
We propose remote direct code execution (RDX), which elevates the power of RDMA from memory access to code execution. We target runtime extension frameworks such as Wasm filters, BPF programs, and UDF functions, where RDX enables an agentless architecture that unlocks capabilities such as fast extension injection, update consistency guarantees, and minimal resource contention. We outline the roadmap for RDX around a new CodeFlow abstraction, encompassing programming remote extensions, exposing management stubs, remotely validating and JIT compiling code, seamlessly linking code to local context, managing remote extension state, and synchronizing code to targets. The case studies and initial results demonstrate the feasibility of RDX and its potential to spark the next wave of RDMA innovations.
Yibo Huang 0005, Yiming Qiu 0001, Daqian Ding, Patrick Tser Jern Kon, Yiwen Zhang 0008, Yuzhou Mao, Archit Bhatnagar, Mosharaf Chowdhury, Srini Devadas, Jiarong Xing, Ang Chen 0001
HotNets10
2024 Occam: A Programming System for Reliable Network Management
abstract
The complexity of large networks makes their management a daunting task. State-of-the-art network management tools use workflow systems for automation, but they do not adequately address the substantial challenges in operation reliability. This paper presents Occam, a programming system that simplifies the development of reliable network management tasks. We leverage the fact that most modern network management systems are backed with a source-of-truth database, and thus customize database techniques to the context of network management. Occam exposes an easy-to-use programming model for network operators to express the key management logic, while shielding them from reliability concerns, such as operational conflicts and task atomicity. Instead, the Occam runtime provides these reliability guardrails automatically. Our evaluation demonstrates Occam's effectiveness in simplifying management tasks, minimizing network vulnerable time and assisting with failure recovery.
Jiarong Xing, Kuo-Feng Hsu, Yiting Xia, Yan Cai 0018, Ying Zhang 0022, Ang Chen 0001
EuroSys1
2024 On the Criticality of Integrity Protection in 5G Fronthaul Networks
Jiarong Xing, Sophia Yoo, Xenofon Foukas, Daehyeok Kim, Michael K. Reiter
USENIX Security Symposium1
2023 Simplifying Cloud Management with Cloudless Computing
abstract
Cloud computing has transformed the IT industry, but managing cloud infrastructures remains a difficult task. We make a case for putting today's management practices, known as "Infrastructure-as-Code," on a firmer ground via a principled design. We call this end goal Cloudless Computing: it aims to simplify cloud infrastructure management tasks by supporting them "as-a-service," analogous to serverless computing that relieves users of the burden of managing server instances. By assisting tenants with these tasks, cloud resources will be presented to their users more readily without the undue burden of complex control. We describe the research problems by examining the typical lifecycle of today's cloud infrastructure management, and identify places where a cloudless approach will advance the state of the art.
Yiming Qiu 0001, Patrick Tser Jern Kon, Jiarong Xing, Yibo Huang 0005, Xinyu Wang 0006, Peng Huang 0005, Mosharaf Chowdhury, Ang Chen 0001
HotNets3
2023 Enabling Resilience in Virtualized RANs with Atlas
abstract
Virtualized radio access networks (vRANs), which allow running RAN processing on commodity servers instead of proprietary hardware, are gaining adoption in cellular networks. Two properties of the vRAN's "Distributed Unit (DU)" that implements the lower RAN layers---its real-time deadlines and its black-box nature---make it challenging to provide resilience features such as upgrades and failover without long service disruptions. These properties preclude the use of existing resilience techniques like virtual machine migration or state replication that are used for typical workloads. This paper presents Atlas, the first system that provides resilience for the DU. The central insight in Atlas is to repurpose existing cellular mechanisms for wireless resilience, namely handovers and cell reselection, to provide software resilience for the DU. For planned resilience events like upgrades, we design a novel technique that simultaneously serves cells from both the old and new DUs via the same radio, and uses handovers between these cells to migrate user devices. For unplanned failures, we identify deficiencies in existing RAN protocols that disrupt cell reselection after DU failure, and show how we can eliminate these disruptions using a middlebox between the DU and higher layers. Our evaluation with a state-of-the-art 5G vRAN testbed shows that Atlas achieves minimal disruption to cellular connectivity during resilience events, while incurring low overhead.
Jiarong Xing, Junzhi Gong, Xenofon Foukas, Anuj Kalia, Daehyeok Kim, Manikanta Kotaru
MobiCom1
2023 Unleashing SmartNIC Packet Processing Performance in P4
abstract
SmartNICs are on the rise as a packet processing platform, with the trend towards a uniform P4 programming model. However, unleashing SmartNIC packet processing performance in P4 is a formidable task. Traditional SmartNIC optimizations rely on low-level program tuning, but P4 abstractions operate at one level above. At the same time, today's P4 optimizations primarily focus on resource packing rather than performance tuning. We develop Pipeleon, an automated performance optimization framework for P4 programmable SmartNICs. We introduce techniques that are tailored to the performance characteristics of SmartNICs, and further leverage dynamic workload patterns for profile-guided optimization. Pipeleon pinpoints program hotspots at the P4 level and computes runtime optimization plans to specialize the program layout based on the latest profile. We have prototyped Pipeleon and applied it to optimize two popular P4 SmartNICs---Nvidia BlueField2 and Netronome Agilio CX---as well as a software SmartNIC emulator extended based on BMv2. Our results show that Pipeleon significantly improves SmartNIC packet processing performance in realistic scenarios.
Jiarong Xing, Yiming Qiu 0001, Kuo-Feng Hsu, Songyuan Sui, Khalid Manaa, Omer Shabtai, Yonatan Piasetzky, Matty Kadosh, Arvind Krishnamurthy, T. S. Eugene Ng, Ang Chen 0001
SIGCOMM1
2023 Remote Direct Memory Introspection
Jiarong Xing, Yibo Huang 0005, Danyang Zhuo, Srini Devadas, Ang Chen 0001
USENIX Security Symposium2
2022 Symbolic Distillation for Learned TCP Congestion Control
abstract
Recent advances in TCP congestion control (CC) have achieved tremendous success with deep reinforcement learning (RL) approaches, which use feedforward neural networks (NN) to learn complex environment conditions and make better decisions. However, such ``black-box'' policies lack interpretability and reliability, and often, they need to operate outside the traditional TCP datapath due to the use of complex NNs. This paper proposes a novel two-stage solution to achieve the best of both worlds: first to train a deep RL agent, then distill its (over-)parameterized NN policy into white-box, light-weight rules in the form of symbolic expressions that are much easier to understand and to implement in constrained environments. At the core of our proposal is a novel symbolic branching algorithm that enables the rule to be aware of the context in terms of various network conditions, eventually converting the NN policy into a symbolic tree. The distilled symbolic rules preserve and often improve performance over state-of-the-art NN policies while being faster and simpler than a standard neural network. We validate the performance of our distilled symbolic rules on both simulation and emulation environments. Our code is available at https://github.com/VITA-Group/SymbolicPCC.
S. P. Sharan, Wenqing Zheng, Kuo-Feng Hsu, Jiarong Xing, Ang Chen 0001, Zhangyang Wang
NeurIPS4
2022 Runtime Programmable Switches
Jiarong Xing, Kuo-Feng Hsu, Matty Kadosh, Alan Lo, Yonatan Piasetzky, Arvind Krishnamurthy, Ang Chen 0001
NSDI1
2022 Bedrock: Programmable Network Support for Secure RDMA Systems
Jiarong Xing, Kuo-Feng Hsu, Yiming Qiu 0001, Ziyang Yang, Ang Chen 0001
USENIX Security Symposium1
2021 Probabilistic profiling of stateful data planes for adversarial testing
abstract
Recently, there is a flurry of projects that develop data plane systems in programmable switches, and these systems perform far more sophisticated processing than simply deciding a packet's next hop (i.e., traditional forwarding). This presents challenges to existing network program profilers, which are developed primarily to handle stateless forwarding programs.
Qiao Kang, Jiarong Xing, Yiming Qiu 0001, Ang Chen 0001
ASPLOS2
2021 A Vision for Runtime Programmable Networks
abstract
Our community has made significant progress in developing programmable network infrastructure, starting from the control plane and expanding to the data plane. As a latest trend, network devices are becoming runtime programmable while serving live traffic. This allows for reprogramming of individual device programs at fine-grained timescales to add or remove network functions. Many applications and services, however, need control over a combination of devices, including end host stacks, NICs, and switches, to accomplish their goals. We lay out our vision for runtime programmable networks, building upon device-level features to provide live, network-wide, runtime reprogramming. A whole-stack approach is needed with new programming models, compiler support, and network management abstractions. We outline a research agenda as a call to arms to the community.
Jiarong Xing, Yiming Qiu 0001, Kuo-Feng Hsu, Matty Kadosh, Alan Lo, Aditya Akella, Thomas E. Anderson, Arvind Krishnamurthy, T. S. Eugene Ng, Ang Chen 0001
HotNets1
2021 Automated SmartNIC Offloading Insights for Network Functions
abstract
The gap between CPU and networking speeds has motivated the development of SmartNICs for NF (network functions) offloading. However, offloading performance is predicated upon intricate knowledge about SmartNIC hardware and careful hand-tuning of the ported programs. Today, developers cannot easily reason about the offloading performance or the effectiveness of different porting strategies without resorting to a trial-and-error approach.
Yiming Qiu 0001, Jiarong Xing, Kuo-Feng Hsu, Qiao Kang, Ming Liu 0027, Srinivas Narayana, Ang Chen 0001
SOSP2
2021 Ripple: A Programmable, Decentralized Link-Flooding Defense Against Adaptive Adversaries
Jiarong Xing, Ang Chen 0001
USENIX Security Symposium1
2020 NetWarden: Mitigating Network Covert Channels while Preserving Performance
Jiarong Xing, Qiao Kang, Ang Chen 0001
USENIX Security Symposium1
2019 Architecting Programmable Data Plane Defenses into the Network with FastFlex
abstract
This paper is motivated by the ever increasing scale and diversity of attacks that are best handled by the network infrastructure. FastFlex builds upon recent progress, which has developed a variety of network defenses in programmable data planes, and takes this trend one step further: it aims to develop architectural support for these defenses as a first-class citizen. We envision that the network architecture would support these defenses as naturally as it does routing---as the network routes traffic end-to-end, it also turns the defenses on and off as needed for attack mitigation. We propose a key abstraction: the multimode data plane. Normally, it operates under optimal configurations computed by centralized control, but upon attacks, it performs distributed mode changes entirely in data plane for mitigation. Mixed-vector attacks would trigger co-existing modes at different regions of the network, and attacks that rapidly change would be met with equally fast mode adaptations. We sketch this vision, discuss the opportunities and challenges it involves, and present a use case on link-flooding defense.
Jiarong Xing, Ang Chen 0001
HotNets1
2018 A distributed multi-level model with dynamic replacement for the storage of smart edge computing
Jiarong Xing, Hongjun Dai, Zhilou Yu
J. Syst. Archit.1