Xuya Jia

dblp:192/3465 · DBLP profile ↗
← Back
12ranked-venue papers
7as first author
4since 2021 · last 2026
0000-0001-8995-879XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 12 · 7 first-author · 4 since 2021
YearPublicationVenuePosition
2026 An Efficient Computing and Communication Framework for Large-Scale Data Processing Cluster
Xuya Jia, Zhiyi Yao, Edison Liu, Congcong Miao, Yuedong Xu 0001
IEEE Trans. Netw.1
2025 Holmes: Localizing Irregularities in LLM Training with Mega-scale GPU Clusters
Zhiyi Yao, Pengbo Hu, Congcong Miao, Xuya Jia, Zuning Liang, Yuedong Xu 0001, Chunzhi He, Mingzhuo Chen, Xiang Li 0010, Zekun He, Yachen Wang, Xianneng Zou, Junchen Jiang
NSDI4
2024 Turbo: Efficient Communication Framework for Large-scale Data Processing Cluster
abstract
Big data processing clusters are suffering from a long job completion time due to the inefficient utilization of the RDMA capability. Our production measurement results in a large-scale cluster with hundreds of server nodes to process large-scale jobs have shown that the existing deployment of RDMA technique results in a long-tail job completion time, with some jobs even taking up more than twice the average time to complete. In this paper, we present the design and implementation of Turbo, an efficient communication framework for the large-scale data processing cluster to achieve high performance and scalability. The core of Turbo's approach is to leverage a dynamic block-level flowlet transmission mechanism and a non-blocking communication middleware to improve the network throughput and enhance system's scalability. Furthermore, Turbo ensures high system reliability by utilizing an external shuffle service as well as TCP serving as a backup. We integrate Turbo into Apache Spark and evaluate Turbo in a small-scale testbed and a large-scale cluster consisting of hundreds of server nodes. The small-scale testbed evaluation results show that Turbo improves the network throughput by 15.1% while maintaining high system reliability. The large-scale production results have shown Turbo can reduce the job completion time by 23.9% and increase the job completion rate by 2.03× over the existing RDMA solutions.
Xuya Jia, Zhiyi Yao, Edison Liu, Xiang Li 0223, Zekun He, Yachen Wang, Xianneng Zou, Chongqing Zhao, Jinhui Chu, Jilong Wang 0001, Congcong Miao
SIGCOMM1
2021 Making Multi-String Pattern Matching Scalable and Cost-Efficient with Programmable Switching ASICs
abstract
Multi-string pattern matching is a crucial building block for many network security applications, and thus of great importance. Since every byte of a packet has to be inspected by a large set of patterns, it often becomes a bottleneck of these applications and dominates the performance of an entire system. Many existing works have been devoted to alleviate this performance bottleneck either by algorithm optimization or hardware acceleration. However, neither one provides the desired scalability and costs that keep pace with the dramatic increase of the network bandwidth and network traffic today. In this paper, we present BOLT, a scalable and cost-efficient multi-string pattern matching system leveraging the capability of emerging programmable switches. BOLT combines the following two techniques, a smart state encoding scheme to fit a large number of strings into the limited memory on the programmable switch, and a variable k-stride transition mechanism to increase the throughput significantly with the same level of memory costs. We implement a prototype of BOLT and make its source code publicly available. Extensive evaluations demonstrate that BOLT could provide orders of magnitude improvement in throughput which is scalable with pattern sets and workloads, and could also significantly decrease the number of entries and memory requirement.
Menghao Zhang 0001, Chang Liu 0021, Ying Liu 0024, Xuya Jia, Mingwei Xu 0001
INFOCOM6
2019 How Powerful Switches Should be Deployed: A Precise Estimation Based on Queuing Theory
abstract
Software-Defined Networking (SDN) provides a tractable and efficient architecture for operators to customize their network functions. Many traditional Data Center Networks (DCNs) are upgraded by SDN to improve link utilization and management flexibility, but they are lack of the instructions for selecting the substitutive SDN switches with the proper flow table space to achieve cost-effective and energy-saving networks. In this paper, we fill the gap of solving the flow table space estimation problem based on queuing theory. First, we divide the life process of a flow table entry into the packet-in process, the handling process and the serving process to establish a queuing system to estimate the least required number of the flow table entries of SDN switches. Second, we analyze the traffic distribution of DCNs to calculate the critical parameters in our model. Third, on the basis of the essence of the structured topologies in DCNs, we construct a probability model of routing strategies to quantize the influence of path selection. Comprehensive experiments show that the relative flow table space estimation error of our model can be less than 10%, which can give operators insights into the requirement of the SDN switches at specific positions.
Gengbiao Shen, Qing Li 0006, Shuo Ai, Yong Jiang 0001, Mingwei Xu 0001, Xuya Jia
INFOCOM6
2019 Metro: An Efficient Traffic Fast Rerouting Scheme With Low Overhead
abstract
Failure is common instead of exception in large-scale networks. To provide high service quality to upper-layer applications, it is desired that a converged backup path can be rapidly launched when failure occurs. In this paper, we design an IP based Fast ReRouting (FRR) scheme called Metro, which can solve the traffic rerouting convergence problem after arbitrary single link/node failure with low stretch for the backup path. When failure occurs in the network, Metro first indicates all the network areas that would be affected by the failure, and then finds out a few bridge links to drain the traffic in the affected network area to the network area that is not affected by the failure. In this way, Metro does not configure tunnels, encapsulate or modify data packets, and hence it is easy to be deployed in current networks. Extensive simulations show that Metro can solve arbitrary single link/node failure with backup paths shorter than the state-of-the-art solutions, and about 98% of the backup path stretch in Metro are the same as the optimal tunnel scheme.
Xuya Jia, Dan Li 0001, Jing Zhu 0007, Yong Jiang 0001
IEEE/ACM Trans. Netw.1
2018 Cross-Layer Self-Similar Coflow Scheduling for Machine Learning Clusters
abstract
In recent years, many companies have developed various distributed computation frameworks for processing machine learning (ML) jobs in clusters. Networking is a well-known bottleneck for ML systems and the cluster demands efficient scheduling for huge traffic (up to 1GB per flow) generated by ML jobs. Coflow has been proven an effective abstraction to schedule flows of such data-parallel applications. However, the implementation of coflow scheduling policy is constrained when coflow characteristics are unknown a prior, and when TCP congestion control misinterprets the congestion signal leading to low throughput. Fortunately, traffic patterns experienced by some ML jobs support to speculate the complete coflow characteristic with limited information. Hence this paper summarizes coflow from these ML jobs as self-similar coflow and proposes a decentralized self-similar coflow scheduler Cicada. Cicada assigns each coflow a probe flow to speculate its characteristics during the transportation and employs the Shortest Job First (SJF) to separate coflow into strict priority queues based on the speculation result. To achieve full bandwidth for throughput- sensitive ML jobs, and to guarantee the scheduling policy implementation, Cicada promotes the elastic transport-layer rate control that outperforms prior works. Large-scale simulations show that Cicada completes coflow 2.08x faster than the state-of-the-art schemes in the information-agnostic scenario.
Yong Jiang 0001, Qing Li 0006, Xuya Jia, Mingwei Xu 0001
ICCCN4
2018 Intelligent path control for energy-saving in hybrid SDN networks
Xuya Jia, Yong Jiang 0001, Zehua Guo 0001, Gengbiao Shen, Lei Wang 0071
Comput. Networks1
2018 Woodpecker: Detecting and mitigating link-flooding attacks via SDN
Lei Wang 0071, Qing Li 0006, Yong Jiang 0001, Xuya Jia
Comput. Networks4
2017 A low overhead flow-holding algorithm in software-defined networks
Xuya Jia, Qing Li 0006, Yong Jiang 0001, Zehua Guo 0001
Comput. Networks1
2016 Incremental Switch Deployment for Hybrid Software-Defined Networks
abstract
Software-Defined Networking (SDN) brings great opportunities to improve network performance. However, due to budget constraints and technique limitations, Internet Service Providers (ISPs) can upgrade only a limited number of conventional switches to SDN switches in real backbone networks at one time. In this paper, we propose one heuristic scheme for deploying SDN switches in hybrid SDNs. Our scheme works for two different cases: (1) maximizing the network control ability with a given upgrading budget constraint, and (2) minimizing the upgrading cost to achieve the best network control ability. We evaluate our scheme in real topologies. We evaluate our scheme in real topologies. The results show that our scheme can achieve 95% of flows controlled with only 10% upgrading cost.
Xuya Jia, Yong Jiang 0001, Zehua Guo 0001
LCN1
2016 Reducing and Balancing Flow Table Entries in Software-Defined Networks
abstract
Software-Defined Networking (SDN) allows flexible and efficient management of networks. However, the limited capacity of flow tables in SDN switches hinders the deployment of SDN. In this paper, we propose a novel routing scheme to improve the efficiency of flow tables in SDNs. To efficiently use the routing scheme, we formulate an optimization problem with the objective to maximize the number of flows in the network, constrained by the limited flow table space in SDN switches. The problem is NP-hard, and we propose the K Similar Greedy Tree (KSGT) algorithm to solve it. We evaluate the performance of KSGT against "traditional" SDN solutions with real-world topologies and traffic. The results show that, compared to the existing solutions, KSGT can reduce about 60% of flow entries when processing the same amount of flows, and improve about 25% of the successful installation and forwarding flows under the same flow table space.
Xuya Jia, Yong Jiang 0001, Zehua Guo 0001, Zhenwei Wu
LCN1