Xinchen Wan

dblp:272/9234 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
16since 2021 · last 2026
0000-0001-6503-5309ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 11 · 4 first-author · 10 since 2021Systems, architecture and hardware · 5 · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Towards Fair and Efficient Congestion Control Through Multi-Agent Deep Reinforcement Learning
abstract
Recent years have witnessed a plethora of learning-based solutions for congestion control (CC) that demonstrate better performance over traditional TCP schemes. However, they fail to provide consistently good convergence properties, includingfairness, fast convergenceandstability, due to the mismatch between their objective functions and these properties. Despite being intuitive, integrating these properties into existing learning-based CC is challenging, because: 1) their training environments are designed for the performance optimization of single flow but incapable of cooperative multi-flow optimization, and 2) there is no directly measurable metric to represent these properties into the training objective function. We present Astraea, a new learning-based congestion control that ensures fast convergence to fairness with stability. At the heart of Astraea is a multi-agent deep reinforcement learning framework that explicitly optimizes these convergence properties during the training process by enabling the learning of interactive policy between multiple competing flows, while maintaining high performance. We further build a faithful multi-flow environment that emulates the competing behaviors of concurrent flows, explicitly expressing convergence properties to enable their optimization during training. We have fully implemented Astraea and our comprehensive experiments show that Astraea can quickly converge to fairness point and exhibit better stability than its counterparts. For example, Astraea achieves near-optimal bandwidth sharing (i.e., fairness) when multiple flows compete for the same bottleneck, delivers up to 8.4× faster convergence speed and 2.8× smaller throughput deviation, while achieving comparable or even better performance over prior solutions.
Han Tian, Xudong Liao, Chaoliang Zeng, Xinchen Wan, Junxue Zhang 0001, Kai Chen 0005
IEEE Trans. Netw.4
2025 Harmonia: A Unified Framework for Heterogeneous FPGA Acceleration in the Cloud
abstract
FPGAs are gaining popularity in the cloud as accelerators for various applications. To make FPGAs more accessible for users and streamline system management, cloud providers have widely adopted the shell-role architecture on their homogeneous FPGA servers. However, the increasing heterogeneity of cloud FPGAs poses new challenges for this architecture. Previous studies either focus on homogeneous FPGAs or only partially address the portability issues for roles, while still requiring laborious shell development for providers and ad-hoc software modifications for users.
Xinchen Wan, Zilong Wang 0007, Qian Zhao 0001, Feng Ning, Qingsong Ning, Shideng Zhang, Zhenyu Li 0001, Layong Luo, Gaogang Xie
ASPLOS (2)3
2025 Design and Operation of Shared Machine Learning Clusters on Campus
abstract
The rapid advancement of large machine learning (ML) models has driven universities worldwide to invest heavily in GPU clusters. Effectively sharing these resources among multiple users is essential for maximizing both utilization and accessibility. However, managing shared GPU clusters presents significant challenges, ranging from system configuration to fair resource allocation among users. This paper introduces SING, a full-stack solution tailored to simplify shared GPU cluster management. Aimed at addressing the pressing need for efficient resource sharing with limited staffing, SING enhances operational efficiency by reducing maintenance costs and optimizing resource utilization. We provide a comprehensive overview of its four extensible architectural layers, explore the features of each layer, and share insights from real-world deployment, including usage patterns and incident management strategies. As part of our commitment to advancing shared ML cluster management, we open-source SING's resources to support the development and operation of similar systems.
Kaiqiang Xu, Decang Sun, Hao Wang 0116, Zhenghang Ren, Xinchen Wan, Xudong Liao, Zilong Wang 0007, Junxue Zhang 0001, Kai Chen 0005
ASPLOS (1)5
2025 Achieving Fairness Generalizability for Learning-based Congestion Control with Jury
abstract
Internet congestion control (CC) has long posed a challenging control problem in networking systems, with recent approaches increasingly incorporating deep reinforcement learning (DRL) to enhance adaptability and performance. Despite promising, DRL-based CC schemes often suffer from poor fairness, particularly when applied to network environments unseen during training. This paper introduces Jury, a novel DRL-based CC scheme designed to achieve fairness generalizability. At its heart, Jury decouples the fairness control from the principal DRL model with two design elements: i) By transforming network signals, it provides a universal view of network environments among competing flows, and ii) It adopts a post-processing phase to dynamically module the sending rate based on flow bandwidth occupancy estimation, ensuring large flows behave more conservatively and smaller flows more aggressively, thus achieving a fair and balanced bandwidth allocation. We have fully implemented Jury, and extensive evaluations demonstrate its robust convergence properties and high performance across a broad spectrum of both emulated and real-world network conditions.
Han Tian, Xudong Liao, Decang Sun, Chaoliang Zeng, Yilun Jin, Junxue Zhang 0001, Xinchen Wan, Zilong Wang 0007, Yong Wang 0046, Kai Chen 0005
EuroSys7
2025 A Generic and Efficient Communication Framework for Message-Level In-Network Computing
Xinchen Wan, Han Tian, Xudong Liao, Chaoliang Zeng, Zilong Wang 0007, Qingsong Ning, Guyue Liu, Layong Luo, Kai Chen 0005
INFOCOM1
2025 MixNet: A Runtime Reconfigurable Optical-Electrical Fabric for Distributed Mixture-of-Experts Training
abstract
Mixture-of-Expert (MoE) models outperform conventional models by selectively activating different subnets, named experts, on a per-token basis. This gated computation generates dynamic communications that cannot be determined beforehand, challenging the existing GPU interconnects that remain static during distributed training. In this paper, we advocate for a first-of-its-kind system, called MixNet, that unlocks topology reconfiguration during distributed MoE training. Towards this vision, we first perform a production measurement study and show that the MoE dynamic communication pattern has strong locality, alleviating the need for global reconfiguration. Based on this, we design and implement a regionally reconfigurable high-bandwidth domain that augments existing electrical interconnects using optical circuit switching (OCS), achieving scalability while maintaining rapid adaptability. We build a fully functional MixNet prototype with commodity hardware and a customized collective communication runtime. Our prototype trains state-of-the-art MoE models with in-training topology reconfiguration across 32 A100 GPUs. Large-scale packet-level simulations show that MixNet achieves performance comparable to a non-blocking fat-tree fabric while boosting the networking cost efficiency (e.g., performance per dollar) of four representative MoE models by 1.2×–1.5× and 1.9×–2.3× at 100 Gbps and 400 Gbps link bandwidths, respectively.
Xudong Liao, Yijun Sun, Han Tian, Xinchen Wan, Yilun Jin, Zilong Wang 0007, Zhenghang Ren, Wenxue Li 0004, Kin Fai Tse, Zhizhen Zhong, Guyue Liu, Ying Zhang 0022, Xiaofeng Ye, Yiming Zhang 0003, Kai Chen 0005
SIGCOMM4
2025 Coflow Scheduling for LLM Training
abstract
Training large language models (LLMs) generates diverse coflows within a cluster, requiring optimized scheduling to enhance communication-computation overlap and minimize training time. Existing schedulers inadequately handle contention both across and within coflows, resulting in suboptimal performance.
Xinchen Wan, Kaiqiang Xu, Xudong Liao, Yilun Jin, Yijun Sun, Zhenghang Ren, Han Tian, Kai Chen 0005
SIGCOMM1
2025 Towards Optimal Rack-scale μs-level CPU Scheduling through In-Network Workload Shaping
Xudong Liao, Han Tian, Xinchen Wan, Chaoliang Zeng, Hao Wang 0116, Junxue Zhang 0001, Mengyu Ma, Guyue Liu, Kai Chen 0005
USENIX ATC3
2024 Astraea: Towards Fair and Efficient Learning-based Congestion Control
abstract
Recent years have witnessed a plethora of learning-based solutions for congestion control (CC) that demonstrate better performance over traditional TCP schemes. However, they fail to provide consistently good convergence properties, including fairness, fast convergence and stability, due to the mismatch between their objective functions and these properties. Despite being intuitive, integrating these properties into existing learning-based CC is challenging, because: 1) their training environments are designed for the performance optimization of single flow but incapable of cooperative multi-flow optimization, and 2) there is no directly measurable metric to represent these properties into the training objective function.
Xudong Liao, Han Tian, Chaoliang Zeng, Xinchen Wan, Kai Chen 0005
EuroSys4
2024 Towards Domain-Specific Network Transport for Distributed DNN Training
Hao Wang 0116, Han Tian, Jingrong Chen 0004, Xinchen Wan, Jiacheng Xia, Gaoxiong Zeng, Wei Bai 0001, Junchen Jiang, Yong Wang 0046, Kai Chen 0005
NSDI4
2024 Accelerating Neural Recommendation Training with Embedding Scheduling
Chaoliang Zeng, Xudong Liao, Xiaodian Cheng, Han Tian, Xinchen Wan, Hao Wang 0116, Kai Chen 0005
NSDI5
2024 Fast, Scalable, and Accurate Rate Limiter for RDMA NICs
abstract
RDMA NICs desire a rate limiter that is accurate, scalable, and fast: to precisely enforce the policies such as congestion control and traffic isolation, to support a large number of flows, and to sustain high packet rates. Prior works such as SENIC and PIEO can achieve accuracy and scalability, but they are not fast enough, thus fail to fulfill the performance requirement of RNICs, due primarily to their monolithic design and one-packet-per-sorting transmission. We present Tassel, a hierarchical rate limiter for RDMA NICs that can deliver high packet rates by enabling multiple-packet-per-sorting transmission, while preserving accuracy and scalability. At its heart, Tassel renovates the workflow of the rate limiter hierarchically: by first applying scalable rate limiting to the flows to be scheduled, followed by accurate rate limiting to the packets to be transmitted, while leveraging adaptive batching and packet filtering to improve the performance of these two steps. We integrate Tassel into the RNIC architecture by replacing the original QP scheduler module and implement the prototype of Tassel using FPGA. Experimental results show that Tassel delivers 125 Mpps packet rate, outperforming SENIC and PIEO by 3.6×, while supporting 16 K flows with low resource usage, 7.5% - 25.6% as compared to SENIC and PIEO, and preserving high accuracy, precisely enforcing rate limits from 100 Kbps to 100 Gbps.
Zilong Wang 0007, Xinchen Wan, Yijun Sun, Qingsong Ning, Junxue Zhang 0001, Kai Chen 0005
SIGCOMM2
2023 Accurate and Scalable Rate Limiter for RDMA NICs
abstract
Rate limiter is required by RDMA NIC (RNIC) to enforce the rate limits calculated by congestion control. RNIC expects the rate limiter to be accurate and scalable: to precisely shape the traffic for numerous flows with minimized resource consumption, thereby mitigating the incasts and congestions and improving the network performance. Previous works, however, fail to meet the performance requirements of RNIC while achieving accuracy and scalability.
Zilong Wang 0007, Xinchen Wan, Chaoliang Zeng, Kai Chen 0005
APNet2
2023 SRNIC: A Scalable Architecture for RDMA NICs
Zilong Wang 0007, Layong Luo, Qingsong Ning, Chaoliang Zeng, Wenxue Li 0004, Xinchen Wan, Xiongfei Geng, Tianhao Wang 0025, Weicheng Ling, Kejia Huo, Pingbo An, Kui Ji, Shideng Zhang, Ruiqing Feng, Kai Chen 0005, Chuanxiong Guo
NSDI6
2023 Scalable and Efficient Full-Graph GNN Training for Large Graphs
abstract
Graph Neural Networks (GNNs) have emerged as powerful tools to capture structural information from graph-structured data, achieving state-of-the-art performance on applications such as recommendation, knowledge graph, and search. Graphs in these domains typically contain hundreds of millions of nodes and billions of edges. However, previous GNN systems demonstrate poor scalability because large and interleaved computation dependencies in GNN training cause significant overhead in current parallelization methods. We present G3, a distributed system that can efficiently train GNNs over billion-edge graphs at scale. G3 introduces GNN hybrid parallelism which synthesizes three dimensions of parallelism to scale out GNN training by sharing intermediate results peer-to-peer in fine granularity, eliminating layer-wise barriers for global collective communication or neighbor replications as seen in prior works. G3 leverages locality-aware iterative partitioning and multi-level pipeline scheduling to exploit acceleration opportunities by distributing balanced workload among workers and overlapping computation with communication in both inter-layer and intra-layer training processes. We show via a prototype implementation and comprehensive experiments that G3 can achieve as much as 2.24x speedup in a 16-node cluster, and better final accuracy over prior works.
Xinchen Wan, Kaiqiang Xu, Xudong Liao, Yilun Jin, Kai Chen 0005, Xin Jin 0008
Proc. ACM Manag. Data1
2022 DGS: Communication-Efficient Graph Sampling for Distributed GNN Training
abstract
Distributed GNN training tends to generate huge volumes of communication. To reduce communication cost, the state-of-the-art sampling-based techniques sample and retrieve only a subset of the nodes. However, our analysis shows that current sampling algorithms are still inefficient in network communication for distributed GNN training, which is mainly because of three problems: first, they overlook the locality of the sampled neighbor nodes in the cluster; second, they sample data only at the coarse-grained graph node level; and third, some mechanisms they adopted fall short in distributed scenarios. This paper presents a graph sampling framework (DGS) for distributed GNN training, which effectively reduces network communication cost while preserving the final GNN model accuracy. To achieve this, DGS samples neighborhood information based on the locality of the neighbor nodes in the cluster, and samples data at the levels of not only graph nodes but also node features based on explanation. Specifically, DGS constructs an explanation graph which preserves the relationship between the local graph and remote nodes, and leverages the recently-proposed model explanation technique to design an online explanation scheme that interprets the importance of nodes and features. Evaluation results show that DGS achieves up to 1.25× throughput speedup over the state-of-the-art FastGCN and reduces the communication cost by up to 28.3%, while preserving the final model accuracy almost the same as that of full-batch training.
Xinchen Wan, Kai Chen 0005, Yiming Zhang 0003
ICNP1
2020 RAT - Resilient Allreduce Tree for Distributed Machine Learning
abstract
Parameter/gradient exchange plays an important role in large-scale distributed machine learning (DML). However, prior solutions such as parameter server (PS) or ring-allreduce (Ring) fall short since they are not resilient to issues or uncertainties like oversubscription, congestion or failures that may occur in datacenter networks (DCN).
Xinchen Wan, Hong Zhang 0025, Hao Wang 0116, Shuihai Hu, Junxue Zhang 0001, Kai Chen 0005
APNet1