VLDB 2026 Research / reviewers in the wild / expert
Junxue Zhang 0001
dblp:155/4434
· DBLP profile ↗
47ranked-venue papers
7as first author
37since 2021 · last 2026
0000-0001-6926-7801ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 23 · 4 first-author · 19 since 2021Systems, architecture and hardware · 15 · 2 first-author · 12 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Efficient Serving of Network-intensive LLM InferencesabstractPrefix caching has become a key technique for LLM serving, and nowadays the reusable KVCache contents are often hosted on distributed servers. For long-context LLM inferences with high cache hit ratio, cross-server KVCache transmission has become an emerging performance bottleneck; such network-intensive LLM inferences are increasingly prevalent in the coming era of agentic AI. However, existing LLM inference engines are essentially compute-centric; we find that they are highly inefficient when serving such workloads due to compute-stage service blocking and ignorance of KVCache-transfer cost. Chen Chen 0067, Junxue Zhang 0001, Zhusheng Wang, Zixuan Guan, Qizhen Weng 0001, Minyi Guo |
APNet | 3 |
| 2026 | Learn-to-Probe: Achieving Signal Distinguishability in Learning-based Congestion ControlabstractInternet congestion control remains a fundamental challenge, and recent learning-based congestion control algorithms (CCAs) have shown potential in optimizing network performance. However, their reliance on heuristically chosen input signals often leads to suboptimal behavior across diverse network conditions. In this paper, we identify the root cause as the lack of signal distinguishability-the ability of signals to reflect meaningful differences in network states. To address this, we propose Learn-to-Probe (LTP), a novel signal engineering paradigm that actively generates distinguishable network signals to improve the learning process. LTP (i) employs Bayesian filtering to accurately estimate network states from historical signals, and (ii) introduces an intrinsic reinforcement learning reward that encourages the flow to probe the network, inducing signal sequences that minimize the estimation uncertainty in (i). This probing behavior naturally enhances signal distinguishability, enabling the learning model to make more informed decisions. Extensive evaluations show that LTP consistently achieves high link utilization, low queuing delay, and stable convergence across diverse environments. Our results underscore the importance of signal distinguishability and offer a new direction for robust, adaptive congestion control. Han Tian, Junxue Zhang 0001, Xudong Liao, Decang Sun, Bin Huang 0024, Wenxue Li 0004, Yong Wang 0046, Kai Chen 0005 |
EuroSys | 3 |
| 2026 | PolicyCache: Intra-flow Learning in Congestion Control
Han Tian, Xudong Liao, Decang Sun, Wenxue Li 0004, Bin Huang 0024, Senbo Fu, Junxue Zhang 0001, Dian Shen, Kai Chen 0005 |
NSDI | 10 |
| 2026 | Reliable RDMA Over Lossy Fabrics via Data-Control Partitioning
Wenxue Li 0004, Xiangzhou Liu, Yunxuan Zhang, Gaoxiong Zeng, Shoushou Ren, Zhenghang Ren, Bowen Liu 0002, Junxue Zhang 0001, Bingyang Liu, Kai Chen 0005 |
IEEE Trans. Netw. | 12 |
| 2026 | Towards Fair and Efficient Congestion Control Through Multi-Agent Deep Reinforcement LearningabstractRecent years have witnessed a plethora of learning-based solutions for congestion control (CC) that demonstrate better performance over traditional TCP schemes. However, they fail to provide consistently good convergence properties, includingfairness, fast convergenceandstability, due to the mismatch between their objective functions and these properties. Despite being intuitive, integrating these properties into existing learning-based CC is challenging, because: 1) their training environments are designed for the performance optimization of single flow but incapable of cooperative multi-flow optimization, and 2) there is no directly measurable metric to represent these properties into the training objective function. We present Astraea, a new learning-based congestion control that ensures fast convergence to fairness with stability. At the heart of Astraea is a multi-agent deep reinforcement learning framework that explicitly optimizes these convergence properties during the training process by enabling the learning of interactive policy between multiple competing flows, while maintaining high performance. We further build a faithful multi-flow environment that emulates the competing behaviors of concurrent flows, explicitly expressing convergence properties to enable their optimization during training. We have fully implemented Astraea and our comprehensive experiments show that Astraea can quickly converge to fairness point and exhibit better stability than its counterparts. For example, Astraea achieves near-optimal bandwidth sharing (i.e., fairness) when multiple flows compete for the same bottleneck, delivers up to 8.4× faster convergence speed and 2.8× smaller throughput deviation, while achieving comparable or even better performance over prior solutions. Han Tian, Xudong Liao, Chaoliang Zeng, Xinchen Wan, Junxue Zhang 0001, Kai Chen 0005 |
IEEE Trans. Netw. | 5 |
| 2025 | Cache-Aware I/O Rate Control for RDMA
Qijing Li, Bowen Liu 0002, Junxue Zhang 0001, Kai Chen 0005 |
APNet | 5 |
| 2025 | Design and Operation of Shared Machine Learning Clusters on CampusabstractThe rapid advancement of large machine learning (ML) models has driven universities worldwide to invest heavily in GPU clusters. Effectively sharing these resources among multiple users is essential for maximizing both utilization and accessibility. However, managing shared GPU clusters presents significant challenges, ranging from system configuration to fair resource allocation among users. This paper introduces SING, a full-stack solution tailored to simplify shared GPU cluster management. Aimed at addressing the pressing need for efficient resource sharing with limited staffing, SING enhances operational efficiency by reducing maintenance costs and optimizing resource utilization. We provide a comprehensive overview of its four extensible architectural layers, explore the features of each layer, and share insights from real-world deployment, including usage patterns and incident management strategies. As part of our commitment to advancing shared ML cluster management, we open-source SING's resources to support the development and operation of similar systems. Kaiqiang Xu, Decang Sun, Hao Wang 0116, Zhenghang Ren, Xinchen Wan, Xudong Liao, Zilong Wang 0007, Junxue Zhang 0001, Kai Chen 0005 |
ASPLOS (1) | 8 |
| 2025 | Achieving Fairness Generalizability for Learning-based Congestion Control with JuryabstractInternet congestion control (CC) has long posed a challenging control problem in networking systems, with recent approaches increasingly incorporating deep reinforcement learning (DRL) to enhance adaptability and performance. Despite promising, DRL-based CC schemes often suffer from poor fairness, particularly when applied to network environments unseen during training. This paper introduces Jury, a novel DRL-based CC scheme designed to achieve fairness generalizability. At its heart, Jury decouples the fairness control from the principal DRL model with two design elements: i) By transforming network signals, it provides a universal view of network environments among competing flows, and ii) It adopts a post-processing phase to dynamically module the sending rate based on flow bandwidth occupancy estimation, ensuring large flows behave more conservatively and smaller flows more aggressively, thus achieving a fair and balanced bandwidth allocation. We have fully implemented Jury, and extensive evaluations demonstrate its robust convergence properties and high performance across a broad spectrum of both emulated and real-world network conditions. Han Tian, Xudong Liao, Decang Sun, Chaoliang Zeng, Yilun Jin, Junxue Zhang 0001, Xinchen Wan, Zilong Wang 0007, Yong Wang 0046, Kai Chen 0005 |
EuroSys | 6 |
| 2025 | eNetSTL: Towards an In-kernel Library for High-Performance eBPF-based Network FunctionsabstractUsing extended Berkeley Packet Filter (eBPF) to implement networking functions (NFs) has been a promising trend for modern network infrastructure. In this paper, we endeavor to implement 35 representative NFs with eBPF, but encounter inherent problems of either incomplete functionality or performance degradation of up to 49.2%. Conventional solutions like modifying the eBPF infrastructure or implementing functions directly in the kernel can lead to intrusive and unstable modifications. Bin Yang 0027, Dian Shen, Junxue Zhang 0001, Lunqi Zhao, Beilun Wang, Guyue Liu, Kai Chen 0005 |
EuroSys | 3 |
| 2025 | GREEN: Carbon-efficient Resource Scheduling for Machine Learning Clusters
Kaiqiang Xu, Decang Sun, Han Tian, Junxue Zhang 0001, Kai Chen 0005 |
NSDI | 4 |
| 2025 | Enabling Efficient GPU Communication over Multiple NICs with FuseLink
Zhenghang Ren, Zilong Wang 0007, Wenxue Li 0004, Kaiqiang Xu, Xudong Liao, Yijun Sun, Bowen Liu 0002, Han Tian, Junxue Zhang 0001, Mingfei Wang, Zhizhen Zhong, Guyue Liu, Ying Zhang 0022, Kai Chen 0005 |
OSDI | 11 |
| 2025 | CEIO: A Cache-Efficient Network I/O Architecture for NIC-CPU Data PathsabstractEfficient Input/Output (I/O) data path between NICs and CPUs/DRAMs is critical for supporting datacenter applications with high-performance network transmission, especially as link speed scales to 100Gbps and beyond. Traditional I/O acceleration strategies, such as Data Direct I/O (DDIO) and Remote Direct Memory Access (RDMA), perform suboptimally due to the inefficient utilization of the Last-Level Cache (LLC). This paper presents CEIO, a novel cache-efficient network I/O architecture that employs proactive rate control and elastic buffering to achieve zero LLC misses in the I/O data path while ensuring the effectiveness of DDIO and RDMA under various network conditions. We have implemented CEIO on commodity SmartNICs and incorporated it into widely-used DPDK and RDMA libraries. Experiments with well-optimized RPC framework and distributed file system under realistic workloads demonstrate that CEIO achieves up to 2.9× higher throughput and 1.9× lower P99.9 latency over prior work. Bowen Liu 0002, Qijing Li, Zhuobin Huang, Yijun Sun, Wenxue Li 0004, Junxue Zhang 0001, Ping Yin, Kai Chen 0005 |
SIGCOMM | 7 |
| 2025 | Revisiting RDMA Reliability for Lossy FabricsabstractDue to the high operational complexity and limited deployment scale of lossless RDMA networks, the community has been exploring efficient RDMA communication over lossy fabrics. State-of-the-art (SOTA) lossy RDMA solutions implement a simplified selective repeat mechanism in RDMA NICs (RNICs) to enhance loss recovery efficiency. However, these solutions still face performance challenges, such as unavoidable ECMP hash collisions and excessive retransmission timeouts (RTOs). In this paper, we revisit RDMA reliability with the goals of being independent of PFC, compatible with packet-level load balancing, free from RTO, and friendly to hardware offloading. To this end, we propose DCP, a transport architecture that co-designs both the switch and RNICs, fully meeting the design goals. At its core, DCP-Switch introduces a simple yet effective lossless control plane, which is leveraged by DCP-RNIC to enhance reliability support for high-speed lossy fabrics, primarily including header-only-based retransmission and bitmap-free packet tracking. We prototype DCP-Switch using P4 switch and DCP-RNIC using FPGA. Extensive experiments demonstrate that DCP achieves 1.6× and 2.1× performance improvements, compared to SOTA lossless and lossy RDMA solutions, respectively. Wenxue Li 0004, Xiangzhou Liu, Yunxuan Zhang, Gaoxiong Zeng, Shoushou Ren, Zhenghang Ren, Bowen Liu 0002, Junxue Zhang 0001, Kai Chen 0005, Bingyang Liu |
SIGCOMM | 12 |
| 2025 | Accelerating Distributed Graph Learning by Using Collaborative In-Network Multicast and Aggregation
Jiawei Huang 0001, Yijun Li 0002, Jingling Liu, Junxue Zhang 0001, Hui Li 0120, Shengwen Zhou, Xiaojuan Lu, Qichen Su, Jianxin Wang 0001, Chee-Wei Tan 0001, Yong Cui 0001, Kai Chen 0005 |
USENIX ATC | 5 |
| 2025 | Towards Optimal Rack-scale μs-level CPU Scheduling through In-Network Workload Shaping
Xudong Liao, Han Tian, Xinchen Wan, Chaoliang Zeng, Hao Wang 0116, Junxue Zhang 0001, Mengyu Ma, Guyue Liu, Kai Chen 0005 |
USENIX ATC | 6 |
| 2025 | Sequoia: An Accessible and Extensible Framework for Privacy-Preserving Machine Learning over Distributed DataabstractPrivacy-preserving machine learning (PPML) algorithms use secure computation protocols to allow multiple data parties to collaboratively train machine learning (ML) models while maintaining their data confidentiality. However, current PPML frameworks couple secure protocols with ML models in PPML algorithm implementations, making it challenging for non-experts to develop and optimize PPML applications, limiting their accessibility and performance. We propose Sequoia, a novel PPML framework that decouples ML models and secure protocols to optimize the development and execution of PPML applications across data parties. Sequoia offers JAX-compatible APIs for users to program their ML models, while using a compiler-executor architecture to automatically apply PPML algorithms and system optimizations for model execution over distributed data. The compiler in Sequoia incorporates cross-party PPML processes into user-defined ML models by transparently adding computation, encryption, and communication steps with extensible policies, and the executor efficiently schedules code execution across multiple data parties, considering data dependencies and device heterogeneity. Compared to existing PPML frameworks, Sequoia requires 64%-92% fewer lines of code for users to implement the same PPML algorithms, and achieves 88% speedup of training throughput in horizontal PPML. Kaiqiang Xu, Di Chai, Junxue Zhang 0001, Fan Lai 0001, Kai Chen 0005 |
Proc. ACM Manag. Data | 3 |
| 2024 | Accelerating Privacy-Preserving Machine Learning With GeniBatchabstractCross-silo privacy-preserving machine learning (PPML) adopt; Partial Homomorphic Encryption (PHE) for secure data combination and high-quality model training across multiple organizations (e.g., medical and financial). However, PHE introduces significant computation and communication overheads due to data inflation. Batch optimization is an encouraging direction to mitigate the problem by compressing multiple data into a single ciphertext. While promising, it is impractical for a large number of cross-silo PPML applications due to the limited vector operations support and severe data corruption. Junxue Zhang 0001, Xiaodian Cheng, Hong Zhang 0025, Yilun Jin, Shuihai Hu, Han Tian, Kai Chen 0005 |
EuroSys | 2 |
| 2024 | Fast, Scalable, and Accurate Rate Limiter for RDMA NICsabstractRDMA NICs desire a rate limiter that is accurate, scalable, and fast: to precisely enforce the policies such as congestion control and traffic isolation, to support a large number of flows, and to sustain high packet rates. Prior works such as SENIC and PIEO can achieve accuracy and scalability, but they are not fast enough, thus fail to fulfill the performance requirement of RNICs, due primarily to their monolithic design and one-packet-per-sorting transmission. We present Tassel, a hierarchical rate limiter for RDMA NICs that can deliver high packet rates by enabling multiple-packet-per-sorting transmission, while preserving accuracy and scalability. At its heart, Tassel renovates the workflow of the rate limiter hierarchically: by first applying scalable rate limiting to the flows to be scheduled, followed by accurate rate limiting to the packets to be transmitted, while leveraging adaptive batching and packet filtering to improve the performance of these two steps. We integrate Tassel into the RNIC architecture by replacing the original QP scheduler module and implement the prototype of Tassel using FPGA. Experimental results show that Tassel delivers 125 Mpps packet rate, outperforming SENIC and PIEO by 3.6×, while supporting 16 K flows with low resource usage, 7.5% - 25.6% as compared to SENIC and PIEO, and preserving high accuracy, precisely enforcing rate limits from 100 Kbps to 100 Gbps. Zilong Wang 0007, Xinchen Wan, Yijun Sun, Qingsong Ning, Junxue Zhang 0001, Kai Chen 0005 |
SIGCOMM | 8 |
| 2024 | Efficient Decentralized Federated Singular Vector Decomposition
Di Chai, Junxue Zhang 0001, Liu Yang 0008, Yilun Jin, Leye Wang, Kai Chen 0005, Qiang Yang 0001 |
USENIX ATC | 2 |
| 2024 | Accelerating Secure Collaborative Machine Learning with Protocol-Aware RDMA
Zhenghang Ren, Mingxuan Fan, Zilong Wang 0007, Junxue Zhang 0001, Chaoliang Zeng, Cheng Hong 0001, Kai Chen 0005 |
USENIX Security Symposium | 4 |
| 2024 | A Survey for Federated Learning Evaluations: Goals and MeasuresabstractEvaluation is a systematic approach to assessing how well a system achieves its intended purpose. Federated learning (FL) is a novel paradigm for privacy-preserving machine learning that allows multiple parties to collaboratively train models without sharing sensitive data. However, evaluating FL is challenging due to its interdisciplinary nature and diverse goals, such as utility, efficiency, and security. In this survey, we first review the major evaluation goals adopted in the existing studies and then explore the evaluation metrics used for each goal. We also introduceFedEval, an open-source platform that provides a standardized and comprehensive evaluation framework for FL algorithms in terms of their utility, efficiency, and security. Finally, we discuss several challenges and future research directions for FL evaluation. Di Chai, Leye Wang, Liu Yang 0008, Junxue Zhang 0001, Kai Chen 0005, Qiang Yang 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2024 | Load Balancing With Multi-Level Signals for Lossless Datacenter NetworksabstractVarious datacenter network (DCN) load balancing schemes have been proposed in the past decade. Unfortunately, most of these solutions designed for lossy DCNs do not work well for Priority Flow Control (PFC) enabled lossless DCNs, primarily due to the reason that the individual congestion signals used in these solutions, e.g., link load, queue length, Round Trip Time (RTT) and Explicit Congestion Notification (ECN), may not be able to correctly or timely reflect the hop-by-hop PFC pausing. This paper first reveals the above problems via extensive experiments, and then based on the insights learned, we present Proteus, a PFC-aware load balancing scheme that is resilient to PFC pausing by exploring a combination of multi-level congestion signals. At its heart, Proteus leverages RTT-level signals (i.e., RTT and link utilization) to detect path status for initial routing decision, and exploits sub-RTT level signal (i.e., cumulative sojourn time) to reflect instantaneous PFC pausing and make timely rerouting choices based on the idea of better-late-than-never. We have implemented Proteus in the hardware programmable switch. Our testbed experiments as well as large-scale simulations show that Proteus can effectively handle PFC pausing under realistic workloads and achieve up to 35%, 31%, 28%, 22% and 46%, 42%, 34%, 29% better average FCT and$99^{th}$percentile FCT than CONGA, DRILL, Hermes and MP-RDMA, respectively. Jinbin Hu 0001, Chaoliang Zeng, Zilong Wang 0007, Junxue Zhang 0001, Kun Guo 0003, Hong Xu 0001, Jiawei Huang 0001, Kai Chen 0005 |
IEEE/ACM Trans. Netw. | 4 |
| 2024 | eMPTCP: A Framework to Fully Extend Multipath TCPabstractMPTCP provides the basic multipath support for network applications to deliver high throughput and robust communication. However, the original MPTCP is designed with limited extensibility. Various research works have tried to extend MPTCP to attain better performance or richer functionalities. These existing approaches either modify the kernel implementation of MPTCP, which involve considerable engineering efforts and may accidentally introduce safety issues, or control MPTCP via userspace tools, which suffer from restricted functionality support. To address this issue, we propose eMPTCP, an easy-to-use framework to fully extend MPTCP without safety risks. Internally, eMPTCP has a modular and pluggable model which allows operators to specify a comprehensive MPTCP extension as a chain of sub-policies. eMPTCP further enforces the policies through packet header manipulations. To ensure safety, eMPTCP is implemented using eBPF. Despite the stringent constraints of eBPF, we show that it is possible to implement an elaborated framework for a fully extensible MPTCP. Through verifying MPTCP in a number of real-world cases and extensive experiments, we show that eMPTCP is able to support a wide range of MPTCP extensions, while the overhead of eMPTCP operations in the kernel is in the scale of nanosecond, and the extra processing time accounts for only about 0.63% of flows’ transmission time. Dian Shen, Bin Yang 0027, Junxue Zhang 0001, Fang Dong 0001, John C. S. Lui |
IEEE/ACM Trans. Netw. | 3 |
| 2024 | Efficient DRL-Based Congestion Control With Ultra-Low OverheadabstractPrevious congestion control (CC) algorithms based on deep reinforcement learning (DRL) directly adjust flow sending rate to respond to dynamic bandwidth change, resulting in high inference overhead. Such overhead may consume considerable CPU resources and hurt the datapath performance. In this paper, we present, a hierarchical congestion control algorithm that fully utilizes the performance gain from deep reinforcement learning but with ultra-low overhead. At its heart, decouples the congestion control task into two subtasks in different timescales and handles them with different components: 1) lightweight CC executor that performs fine-grained control responding to dynamic bandwidth changes; and 2) RL agent that works at a coarse-grained level that generates control sub-policies for the CC executor. Such two-level control architecture can provide fine-grained DRL-based control with a low model inference overhead. Real-world experiments and emulations show that achieves consistent high performance across various network conditions with an ultra-low control overhead reduced by at least 80% compared to its DRL-based counterparts, similar to classic CC schemes such as Cubic. Han Tian, Xudong Liao, Chaoliang Zeng, Decang Sun, Junxue Zhang 0001, Kai Chen 0005 |
IEEE/ACM Trans. Netw. | 5 |
| 2024 | LiteFlow: Toward High-Performance Adaptive Neural Networks for Kernel DatapathabstractAdaptive neural networks (NN) have been used to optimize OS kernel datapath functions because they can achieve superior performance under changing environments. However, how to deploy these NNs remains a challenge. One approach is to deploy these adaptive NNs in the userspace. However, such userspace deployments suffer from either high cross-space communication overhead or low responsiveness, significantly compromising the function performance. On the other hand, pure kernel-space deployments also incur a large performance degradation because the computation logic of model tuning algorithm is typically complex, interfering with the performance of normal datapath execution. This paper presents LiteFlow, a hybrid solution to build high-performance adaptive NNs for kernel datapath. At its core, LiteFlow decouples the control path of adaptive NNs into: 1) a kernel-space fast path for efficient model inference; and 2) a userspace slow path for effective model tuning. We have implemented LiteFlow with Linux kernel datapath and evaluated it with three popular datapath functions including congestion control, flow scheduling, and load balancing. Compared to prior works, LiteFlow achieves 44.4% better goodput for congestion control, and improves the completion time for long flows by 33.7% and 56.7% for flow scheduling and load balancing, respectively. Junxue Zhang 0001, Chaoliang Zeng, Hong Zhang 0025, Shuihai Hu, Kai Chen 0005 |
IEEE/ACM Trans. Netw. | 1 |
| 2024 | High-Performance Hardware Acceleration Architecture for Cross-Silo Federated LearningabstractCross-silo federated learning (FL) adopts various cryptographic operations to preserve data privacy, which introduces significant performance overhead. In this paper, we identify nine widely-used cryptographic operations and design an efficient hardware architecture to accelerate them. However, directly offloading them on hardware statically leads to (1) inadequate hardware acceleration due to the limited resources allocated to each operation; (2) insufficient resource utilization, since different operations are used at different times. To address these challenges, we propose FLASH, a high-performance hardware acceleration architecture for cross-silo FL systems. At its heart, FLASH extracts two basic operators—modular exponentiation and multiplication—behind the nine cryptographic operations and implements them as highly-performant engines to achieve adequate acceleration. Furthermore, it leverages a dataflow scheduling scheme to dynamically compose different cryptographic operations based on these basic engines to obtain sufficient resource utilization. We have implemented a fully-functional FLASH prototype with Xilinx VU13P FPGA and integrated it with FATE, the most widely-adopted cross-silo FL framework. Experimental results show that, for the nine cryptographic operations, FLASH achieves up to$14.0\times$and$3.4\times$acceleration over CPU and GPU, translating to up to$6.8\times$and$2.0\times$speedup for realistic FL applications, respectively. We finally evaluate the FLASH design as an ASIC, and it achieves$23.6\times$performance improvement upon the FPGA prototype. Junxue Zhang 0001, Xiaodian Cheng, Liu Yang 0008, Jinbin Hu 0001, Han Tian, Kai Chen 0005 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2023 | Enabling Load Balancing for Lossless DatacentersabstractVarious datacenter network (DCN) load balancing schemes have been proposed in the past decade. Unfortunately, most of these solutions designed for lossy DCNs do not work well for Priority Flow Control (PFC) enabled lossless DCNs, primarily due to the reason that the individual congestion signals used in these solutions, e.g., link load, queue length, Round Trip Time (RTT) and Explicit Congestion Notification (ECN), may not be able to correctly or timely reflect the hop-by-hop PFC pausing. This paper first reveals the above problems via extensive experiments, and then based on the insights learned, we present Proteus, a PFC-aware load balancing scheme that is resilient to PFC pausing by exploring a combination of multi-level congestion signals. At its heart, Proteus leverages RTT-Ievel signals (i.e., RTT and link utilization) to detect path status for initial routing decision, and exploits sub-RTT level signal (i.e., cumulative sojourn time) to reflect instantaneous PFC pausing and make timely rerouting choices based on the idea of better-late-than-never. We have implemented Proteus in the hardware programmable switch. Our testbed experiments as well as large-scale simulations show that Proteus can effectively handle PFC pausing under realistic workloads and achieve up to 35 %, 31 %, 28%, 22% and 46 %, 42 %, 34 %, 29 % better average FCT and 99thpercentile FCT than CONGA, DRILL, Hermes and MP-RDMA, respectively. Jinbin Hu 0001, Chaoliang Zeng, Zilong Wang 0007, Junxue Zhang 0001, Kun Guo 0003, Hong Xu 0001, Jiawei Huang 0001, Kai Chen 0005 |
ICNP | 4 |
| 2023 | Communication Efficient Secret Sharing with Dynamic Communication-Computation ConversionabstractSecret Sharing (SS) is widely adopted in secure Multi-Party Computation (MPC) with its simplicity and computational efficiency. However, SS-based MPC protocol introduces significant communication overhead due to interactive operations on secret sharings over the network. For instance, training a neural network model with SS-based MPC may incur tens of thousands of communication rounds among parties, making it extremely hard for real-world deployment.To reduce the communication overhead of SS, prior works statically convert interactive operations to equivalent non-interactive operations with extra computation cost. However, we show that such static conversion misses chances for optimization, and further present SOLAR, an SS-based MPC framework that aims to reduce the communication overhead through dynamic communication-computation conversion. At its heart, SOLAR converts interactive operations that involve communication among parties to equivalent non-interactive operations within each party with extra computations and introduces a speculative strategy to perform opportunistic conversion when CPU is idle for network transmission. We have implemented and evaluated SOLAR on several popular MPC applications, and achieved 1.6-8.1 times speedup in multi-thread setting compared to the basic SS and 1.2-8.6 times speedup over static conversion. Zhenghang Ren, Xiaodian Cheng, Mingxuan Fan, Junxue Zhang 0001, Cheng Hong 0001 |
INFOCOM | 4 |
| 2023 | FLASH: Towards a High-performance Hardware Acceleration Architecture for Cross-silo Federated Learning
Junxue Zhang 0001, Xiaodian Cheng, Liu Yang 0008, Jinbin Hu 0001, Kai Chen 0005 |
NSDI | 1 |
| 2023 | Enabling ECN for Datacenter Networks With RTT VariationsabstractECN has been widely employed in production datacenters to deliver high throughput low latency communications. Despite being successful, prior ECN-based transports have an important drawback: they adopt a fixed RTT value in calculating instantaneous ECN marking threshold while overlooking the RTT variations in practice. In this paper, we reveal that the current practice of using a fixed high-percentile RTT for ECN threshold calculation can lead to persistent queue buildups, significantly increasing packet latency. On the other hand, directly adopting lower percentile RTTs results in throughput degradation. To handle the problem, we introduce$\sf{ECN}^{\unicode{x266F}}$, a simple yet effective solution to enable ECN for RTT variations. At its heart,$\sf{ECN}^{\unicode{x266F}}$inherits the current instantaneous ECN marking (based on a high-percentile RTT) to achieve high throughput and burst tolerance, while further marking packets (conservatively) upon detecting long-term queue buildups to eliminate unnecessary queueing delay without degrading throughput. We implement$\sf{ECN}^{\unicode{x266F}}$on a Barefoot Tofino switch and evaluate it through extensive testbed experiments and large-scale simulations. Our evaluation confirms that$\sf{ECN}^{\unicode{x266F}}$can effectively reduce latency without hurting throughput. For example, compared to the current practice,$\sf{ECN}^{\unicode{x266F}}$achieves up to$23.4\%$($31.2\%$) lower average (99th percentile) flow completion time (FCT) for short flows while delivering similar FCT for large flows under production workloads. Junxue Zhang 0001, Wei Bai 0001, Kai Chen 0005 |
IEEE Trans. Cloud Comput. | 1 |
| 2022 | Spine: an efficient DRL-based congestion control with ultra-low overheadabstractPrevious congestion control (CC) algorithms based on deep reinforcement learning (DRL) directly adjust flow sending rate to respond to dynamic bandwidth change, resulting in high inference overhead. Such overhead may consume considerable CPU resources and hurt the datapath performance. In this paper, we present Spine, a hierarchical congestion control algorithm that fully utilizes the performance gain from deep reinforcement learning but with ultra-low overhead. At its heart, Spine decouples the congestion control task into two subtasks in different timescales and handles them with different components: i) a lightweight CC executor that performs fine-grained control responding to dynamic bandwidth changes, and ii) an RL agent that works at a coarse-grained level that generates control sub-policies for the CC executor. Such two-level control architecture can provide fine-grained DRL-based control with a low model inference overhead. Real-world experiments and emulations show that Spine achieves consistent high performance across various network conditions with an ultra-low control overhead reduced by at least 80% compared to its DRL-based counterparts, similar to classic CC schemes such as Cubic. Han Tian, Xudong Liao, Chaoliang Zeng, Junxue Zhang 0001, Kai Chen 0005 |
CoNEXT | 4 |
| 2022 | Multi-objective congestion controlabstractDecades of research on Internet congestion control (CC) have produced a plethora of algorithms that optimize for different performance objectives. Applications face the challenge of choosing the most suitable algorithm based on their needs, and it takes tremendous efforts and expertise to customize CC algorithms when new demands emerge. In this paper, we explore a basic question: can we design a single CC algorithm to satisfy different objectives? Yiqing Ma, Han Tian, Xudong Liao, Junxue Zhang 0001, Weiyan Wang, Kai Chen 0005, Xin Jin 0008 |
EuroSys | 4 |
| 2022 | Towards the Full Extensibility of Multipath TCP with eMPTCPabstractMPTCP provides the basic multipath support for network applications to deliver high throughput and robust communication. However, the original MPTCP is designed with limited extensibility. Various research works have tried to extend MPTCP to attain better performance or richer functionalities. These existing approaches either modify the kernel implementation of MPTCP, which involve considerable engineering efforts and may accidentally introduce security issues, or control MPTCP via user-space tools, which suffer from restricted functionality support. To address this issue, we propose eMPTCP, an easy-to-use framework to fully extend MPTCP without security risks. Internally, eMPTCP has a modular and pluggable model which allows operators to specify a comprehensive MPTCP extension as a chain of sub-policies. eMPTCP further enforces the policies through packet header manipulations. To ensure safety, eMPTCP is implemented using eBPF. Despite the stringent constraints of eBPF, we show that it is possible to implement an elaborated framework for a fully extensible MPTCP. Through verifying MPTCP in a number of real-world cases and extensive experiments, we show that eMPTCP is able to support a wide range of MPTCP extensions, while the overhead of eMPTCP operations in the kernel is in the scale of nanosecond, and the extra processing time accounts for only about 0.63% of flows' transmission time. Bin Yang 0027, Dian Shen, Junxue Zhang 0001, Fang Dong 0001, Junzhou Luo, John C. S. Lui |
ICNP | 3 |
| 2022 | Practical Lossless Federated Singular Vector Decomposition over Billion-Scale DataabstractWith the enactment of privacy-preserving regulations, e.g., GDPR, federated SVD is proposed to enable SVD-based applications over different data sources without revealing the original data. However, many SVD-based applications cannot be well supported by existing federated SVD solutions. The crux is that these solutions, adopting either differential privacy (DP) or homomorphic encryption (HE), suffer from accuracy loss caused by unremovable noise or degraded efficiency due to inflated data. Di Chai, Leye Wang, Junxue Zhang 0001, Liu Yang 0008, Shuowei Cai, Kai Chen 0005, Qiang Yang 0001 |
KDD | 3 |
| 2022 | LiteFlow: towards high-performance adaptive neural networks for kernel datapathabstractAdaptive neural networks (NN) have been used to optimize OS kernel datapath functions because they can achieve superior performance under changing environments. However, how to deploy these NNs remains a challenge. One approach is to deploy these adaptive NNs in the userspace. However, such userspace deployments suffer from either high cross-space communication overhead or low responsiveness, significantly compromising the function performance. On the other hand, pure kernel-space deployments also incur a large performance degradation because the computation logic of model tuning algorithm is typically complex, interfering with the performance of normal datapath execution. Junxue Zhang 0001, Chaoliang Zeng, Hong Zhang 0025, Shuihai Hu, Kai Chen 0005 |
SIGCOMM | 1 |
| 2022 | Sphinx: Enabling Privacy-Preserving Online Learning over the CloudabstractWith the growing complexity of deep learning applications, users have started to delegate their data and models to the cloud. Among these applications, online learning services, which involve both training and inference procedures, are widely deployed. To ensure privacy guarantee on the public cloud, researchers have proposed a plethora of privacy-preserving deep learning algorithms with different techniques, ranging from obfuscation mechanisms to cryptographic tools. However, none of them is applicable to online learning services. They either focus only on inference or training procedure while ignoring the other, or require non-colluding or trusted third parties. In this paper, we present Sphinx, an efficient and privacy-preserving online deep learning system without any trusted third parties. Sphinx strikes a balance between model performance, computational efficiency, and privacy preservation with systematical optimizations on both private inference and training protocols. At its core, Sphinx synthesizes homomorphic encryption and differential privacy reciprocally to maintain the model by keeping most of its parameters as plaintexts, enabling fast training and inference protocol designs. Meanwhile, by refining the homomorphic operation behaviors, Sphinx avoids most of the heavyweight homomorphic operations and minimizes the communication cost. As a result, Sphinx is able to reduce the training time significantly while achieving real-time inference without exposing user privacy. In our experiments, we find that compared to the pure homomorphic encryption solution, Sphinx is $35 \times$ faster for training and 4 orders of magnitude faster for inference, providing real-time inference response (0.05 seconds for MNIST and 0.08 seconds for CIFAR-10). Our experiments also demonstrate that Sphinx achieves promising model accuracy under a tight privacy budget (96% accuracy under $\epsilon=2, \delta=10^{-5}$ for MNIST) without a trusted data aggregator, and is more robust against practical reconstruction attacks. Han Tian, Chaoliang Zeng, Zhenghang Ren, Di Chai, Junxue Zhang 0001, Kai Chen 0005, Qiang Yang 0001 |
SP | 5 |
| 2021 | Enabling Low Latency Edge Intelligence based on Multi-exit DNNs in the WildabstractIn recent years, deep neural networks (DNNs) have witnessed a booming of artificial intelligence Internet of Things applications with stringent demands across high accuracy and low latency. A widely adopted solution is to process such computation-intensive DNNs inference tasks with edge computing. Nevertheless, existing edge-based DNN processing methods still cannot achieve acceptable performance due to the intensive transmission data and unnecessary computation. To address the above limitations, we take the advantage of Multi-exit DNNs (ME-DNNs) that allows the tasks to exit early at different depths of the DNN during inference, based on the input complexity. However, naively deploying ME-DNNs in edge still fails to deliver fast and consistent inference in the wild environment. Specifically, 1) at the model-level, unsuitable exit settings will increase additional computational overhead and will lead to excessive queuing delay; 2) at the computation-level, it is hard to sustain high performance consistently in the dynamic edge computing environment. In this paper, we present a Low Latency Edge Intelligence Scheme based on Multi-Exit DNNs (LEIME) to tackle the aforementioned problem. At the model-level, we propose an exit setting algorithm to automatically build optimal ME-DNNs with lower time complexity; At the computation-level, we present a distributed offloading mechanism to fine-tune the task dispatching at runtime to sustain high performance in the dynamic environment, which has the property of close-to-optimal performance guarantee. Finally, we implement a prototype system and extensively evaluate it through testbed and large-scale simulation experiments. Experimental results demonstrate that LEIME significantly improves applications' performance, achieving 1.1–18.7 × speedup in different situations. Zhaowu Huang, Fang Dong 0001, Dian Shen, Junxue Zhang 0001, Huitian Wang, Guangxing Cai, Qiang He 0001 |
ICDCS | 4 |
| 2020 | RAT - Resilient Allreduce Tree for Distributed Machine LearningabstractParameter/gradient exchange plays an important role in large-scale distributed machine learning (DML). However, prior solutions such as parameter server (PS) or ring-allreduce (Ring) fall short since they are not resilient to issues or uncertainties like oversubscription, congestion or failures that may occur in datacenter networks (DCN). Xinchen Wan, Hong Zhang 0025, Hao Wang 0116, Shuihai Hu, Junxue Zhang 0001, Kai Chen 0005 |
APNet | 5 |
| 2020 | Facilitating Application-Aware Bandwidth Allocation in the Cloud with One-Step-Ahead Traffic InformationabstractBandwidth allocation to virtual machines (VMs) has a significant impact on the performance of communication-intensive big data applications hosted in VMs. It is crucial to accurately determine how much bandwidth to be reserved for VMs and when to adjust it. Past approaches typically resort to predicting the long-term network demands of applications for bandwidth allocation. However, lacking of prediction accuracy, these methods lead to the unpredictable application performance. Recently, it is conceded that the network demands of applications can only be accurately derived right before each of their execution phases. Hence, it is challenging to timely allocate the bandwidth to VMs with limited information. In this paper, we design and implement AppBag, an Application-aware Bandwidth guarantee framework, which allocates the accurate bandwidth to VMs with one-step-ahead traffic information. We propose an algorithm to allocate the bandwidth to VMs and map them onto feasible hosts. To reduce the overhead when adjusting the allocation, an efficient Lazy Migration (LM) algorithm is proposed with bounded performance. We conduct extensive evaluations using real-world applications, showing that AppBag can handle the bandwidth requests at run-time, while reducing the execution time of applications by 47.3 percent and the global traffic by 36.7 percent, compared to the state-of-the-art methods. Dian Shen, Junzhou Luo, Fang Dong 0001, Jiahui Jin 0001, Junxue Zhang 0001, Jun Shen 0001 |
IEEE Trans. Serv. Comput. | 5 |
| 2019 | Rethinking Transport Layer Design for Distributed Machine LearningabstractMotivated by the increasing scale of data, we see a growing need of high performance distributed machine learning systems. Many research works are being proposed to improve distributed machine learning performance. Jiacheng Xia, Gaoxiong Zeng, Junxue Zhang 0001, Weiyan Wang, Wei Bai 0001, Junchen Jiang, Kai Chen 0005 |
APNet | 3 |
| 2019 | Enabling ECN for datacenter networks with RTT variationsabstractECN has been widely employed in production datacenters to deliver high throughput low latency communications. Despite being successful, prior ECN-based transports have an important drawback: they adopt a fixed RTT value in calculating instantaneous ECN marking threshold while overlooking the RTT variations in practice. Junxue Zhang 0001, Wei Bai 0001, Kai Chen 0005 |
CoNEXT | 1 |
| 2017 | Resilient Datacenter Load Balancing in the WildabstractProduction datacenters operate under various uncertainties such as traffic dynamics, topology asymmetry, and failures. Therefore, datacenter load balancing schemes must be resilient to these uncertainties; i.e., they should accurately sense path conditions and timely react to mitigate the fallouts. Despite significant efforts, prior solutions have important drawbacks. On the one hand, solutions such as Presto and DRB are oblivious to path conditions and blindly reroute at fixed granularity. On the other hand, solutions such as CONGA and CLOVE can sense congestion, but they can only reroute when flowlets emerge; thus, they cannot always react timely to uncertainties. To make things worse, these solutions fail to detect/handle failures such as blackholes and random packet drops, which greatly degrades their performance. Hong Zhang 0025, Junxue Zhang 0001, Wei Bai 0001, Kai Chen 0005, Mosharaf Chowdhury |
SIGCOMM | 2 |
| 2017 | Enabling application-aware flexible graph partition mechanism for parallel graph processing systemsabstractSummary With the emerging of the large‐scale graph data,Pregel‐like graph parallel processing systems have been an essential tool to efficiently process the graph data. The first step to use thePregel‐like systems is to partition the graph into multiple blocks and distribute them on multiple machines. The partition strategy plays a significant role in determining the performance because a good partition could both ensure load balance and optimize network communication overhead, and vice versa. However, existing partition strategies fail to meet the requirements because they suffer from the following drawbacks: (1) they ignore the application features and (2) they ignore the multi‐application feature in productive environment. To overcome those drawbacks, we proposed thesuperblockpartition strategy, which utilizes theatomic blocksgenerated by pre‐processing of the original graph and could be constructed and re‐constructed dynamically according to the submitted applications in real time. The hash‐based and clustering‐based pre‐partition methods are covered in details. The application feature extraction method and heuristicsuperblockpartition algorithm are proposed to construct the superblocks. Experimental results show that thesuperblockpartition strategy could boost the graph processing performance and its partition efficiency also outperforms the hash‐based and topology optimal partition strategy. Copyright © 2016 John Wiley & Sons, Ltd. Fang Dong 0001, Junxue Zhang 0001, Junzhou Luo, Dian Shen, Jiahui Jin 0001 |
Concurr. Comput. Pract. Exp. | 2 |
| 2016 | AppBag: Application-Aware Bandwidth Allocation for Virtual Machines in Cloud EnvironmentabstractIt is challenging to allocate the network bandwidth to virtual machines(VMs) hosting communication-intensive applications. Due to the temporal and spatial variability of the hosted applications, it is crucial how much bandwidth to be reserved for each VM and when to adjust it. Prior approaches typically resort to predicting the applications' network demands, according to which the VMs are placed once for all or periodically migrated. However, recent works conceded that the network demands of applications can only be accurately derived right before each execution phase. In this paper, we propose AppBag, an Application-aware Bandwidth guarantee framework which allocates the bandwidth to VMs using only one-stepahead information. An efficient VM migration algorithm is then proposed to adjust the bandwidth allocation and corresponding VM placement, subjected to the network demands variation in future execution phases. We further implement AppBag with OpenStack and deploy it on the testbed environment in our data center. Extensive evaluations using popular applications show that AppBag can handle the bandwidth requests at run-time while improving applications' performance and reducing the global traffic in the data center fabric. Dian Shen, Junzhou Luo, Fang Dong 0001, Junxue Zhang 0001 |
ICPP | 4 |
| 2016 | A client-side directory prefetching mechanism for GlusterFSabstractDistributed file system has the characteristics of large capacity, good scalability and high reliability, which make it widely used in many areas involving large-scale data storage. It offers simplified, highly-available services for users to access data. However, due to the non-metadata design, the performance of traversal operation on large directories in those non-metadata distributed file systems is poor. With the increasing amount of files, it severely affects the user experience. In this paper, we present a directory prefetching mechanism on the client side to reduce directory traversal operation latency in non-metadata distributed file system. The mechanism, combined with the client's cache, adopts the directory access history to predict future access pattern and fetches the content of the directory without user intervention. Our goal is to reduce the overall access latency in the non-metadata distributed file system in order to better satisfy the user experience. Fang Dong 0001, Junxue Zhang 0001, Zhuqing Xu, Junzhou Luo |
SMC | 3 |
| 2015 | Towards optimized scheduling for data-intensive scientific workflow in multiple datacenter environmentabstractSummary In the big data era, scientific workflow exhibits the characteristics of data intensity and becomes increasingly popular in scientific domains. Efficient scheduling of data‐intensive scientific workflow in a multiple datacenter (DC) environment has been a long‐standing challenge. Most of previous work on data‐intensive scientific workflow scheduling primarily focused on the optimization of reducing the volumes of data transfer between workflow tasks. In this paper, novel scheduling strategies for the execution of data‐intensive scientific workflow in multi‐DC environment are proposed aiming at the optimization of the overall data transfer time. A novel DC selection approach is proposed to minimize the number of DCs having enough storage capacity for the execution of scientific workflow as well as optimized inter‐DC network bandwidth for efficient data transfer between workflow tasks. A k‐means clustering‐based data placement strategy is adopted to intelligently place the initial data of scientific workflow thereby reducing the volume of initial data transfer between different DCs. A multilevel task replication scheduling strategy is invented to reduce the volumes of intermediate data transfer between DCs during the runtime of the scientific workflow. Simulations spanning a broad range of scientific workflow and multi‐DC settings are performed in order to verify the proposed approaches. The numerical results show that our combined scheduling strategy significantly reduces the overall data transfer time and data transfer volume when scientific workflow is scheduled in multi‐DC environment. Copyright © 2015 John Wiley & Sons, Ltd. Jinghui Zhang 0001, Junzhou Luo, Fang Dong 0001, Junxue Zhang 0001 |
Concurr. Comput. Pract. Exp. | 5 |
| 2014 | Game theory based dynamic resource allocation for hybrid environment with cloud and big data applicationabstractVirtualization based cloud and big data applications have been widely adopted in various fields. Because deploying the big data applications on the cloud will cause obvious performance degradation, the cloud and big data applications are provided with fixed resource separately. However, the traditional fixed resource allocation mechanism has two drawbacks: (1) low resource utility and (2) unresponsiveness to the performance degradation. To address these drawbacks, the cloud and big data hybrid environment is designed, where fair resource allocation is used to ensure fairness between cloud and big data applications while virtual machine migration is used to make each virtual machine in cloud application reach its own satisfactory. Herein, game theory is used to model the conflict and negotiation between cloud and big data applications. Firstly, the Nash Equilibrium is used to discover the best strategy for both applications. Secondly, as for virtual machine migration, we use Nash Bargaining game to present the situation where virtual machines compete for more resources allocation while their minimal demand is ensured. Finally, experiments are carried out to prove that the hybrid environment outperforms the traditional method both in resource utility and application performance. Junxue Zhang 0001, Fang Dong 0001, Dian Shen, Junzhou Luo |
SMC | 1 |