EDBT 2026 Demo / reviewers in the wild / expert
Shenzhi Yuan
dblp:286/5809
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2026
0009-0003-1790-2588ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 5 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SprayCast: Congestion-Adaptive Native Multicast for Dynamic Sparse All-to-All CommunicationabstractMixture-of-Experts (MoE) models outperform traditional dense models through sparse expert activation, where each token is dynamically routed to a small subset of experts. Across many tokens, these sparse Dispatch operations induce all-to-all traffic, making communication a major bottleneck for both training and inference: unicast replication wastes bandwidth, while table-driven multicast struggles with receiver-set churn and incast. In this paper, we propose SprayCast, a congestion-adaptive native RDMA multicast scheme for dynamic sparse token Dispatch. To avoid maintaining multicast forwarding tables in switches, SprayCast encodes each packet’s destination node set in its packet header using hierarchical bitmaps, enabling table-free in-network replication. It uses in-band network telemetry (INT) feedback to steer replication away from congested multicast branches and range-based negative acknowledgments (NACKs) for localized loss recovery, saving bandwidth and reducing tail latency in dynamic all-to-all communication. In htsim simulations on a 128-server fat-tree, SprayCast achieves better scalability as top-K dispatch fanout increases and reduces P99 dispatch tail latency by up to 6 × at K = 8 compared with representative baselines. Yingying Zeng, Xiaobin Tan, Shenzhi Yuan, Feng Yang 0013 |
APNet | 5 |
| 2025 | CacheMon: In-Network Cache Coordination for Massively Scalable Distributed Storage SystemsabstractThe exponential growth of data creates significant challenges for distributed storage systems. Conventional cache management architectures face limitations in performance and scalability, primarily due to uncoordinated resource utilization across front-end and back-end networks, skewed data access patterns, and the demands of ultra-high concurrency. In this paper, we propose CacheMon, an in-network cache coordination system based on hybrid topology for massively scalable distributed storage. CacheMon establishes a control plane in the programmable switch to integrate front-end and back-end resources, ending the inefficiency of traditionally isolated networks. We design two key mechanisms in CacheMon: a cache tracking mechanism for maintaining data consistency through location recording, and a unified load balancing mechanism that leverages in-network measurements to optimize resource allocation and adaptively counter workload skew. Experimental results confirm that CacheMon significantly improves throughput performance and system scalability under skewed and concurrent workloads, leveraging its hybrid topology to efficiently coordinate network resources. Kexin Ju 0003, Xiaobin Tan, Shenzhi Yuan, Shangwei Li, Chaoming Huang, Quan Zheng 0002 |
ICPADS | 3 |
| 2025 | Straggler Dynamic Management for Distributed DNN TrainingabstractStraggler nodes are a major bottleneck in large-scale distributed training, degrading efficiency and stability. However, current solutions, including In-Network Aggregation (INA), lack the adaptability to effectively manage these stragglers in dynamic environments. This paper proposes Straggler Dynamic Management (SDM), an adaptive method for large-scale distributed training that performs dynamic straggler management by coordinating the data and control planes to achieve accurate, time-based detection and efficient mitigation via a performanceaware redundancy strategy and semi-asynchronous aggregation. SDM manages stragglers through a coordinated architecture that decouples the data and control planes for efficient detection and response. It leverages the data plane to estimate each node's remaining completion time, ensuring accurate and low-overhead straggler identification. The control plane then mitigates their impact using two key strategies: a performance-aware redundancy scheme to reduce waiting delays, and a semi-asynchronous aggregation mechanism that dynamically adjusts synchronization to alleviate gradient staleness and improve model convergence. We implement and deploy SDM on a real-world hardware testbed and evaluate its performance under various straggler scenarios. Experimental results demonstrate that SDM significantly improves training efficiency and convergence stability in the presence of straggler nodes, particularly when multiple stragglers occur simultaneously, exhibiting greater robustness and adaptability than existing methods. Tiance Li, Bo Chai, Xiaobin Tan, Shenzhi Yuan, Kexin Ju 0003, Shiyin Zhu |
ICPADS | 4 |
| 2024 | dotPS: Disorder Tolerant Load Balancing Scheme for Datacenter NetworkabstractRecently, with the development of AI technology, load balancing methods are commonly integrated to enhance datacenter network throughput within a spine-leaf architecture. However, traditional load balancing inevitably leads to packet disorder, affecting system efficiency. Considering that network applications can actually tolerate a certain disorder degree, we think the transmission efficiency will be improved without requiring additional processing of packets within a certain disorder tolerance threshold. Thus, we propose a disorder tolerant load balancing scheme called Disorder Tolerant Packet Spray, dotPS. We design the architecture of dotPS, controlling the degree of disorder tolerance to improve transmission efficiency, and propose the concept of a group interval which can be dynamically adjusted as a measurement of disorder tolerance. Then we design a load balancing algorithm based on group interval, requiring strict ordering for inter-group packets while tolerating disorder in intra-group packets. Finally, simulation results dedicate that compared to the state-of-the-art load balancing methods under different network loads, the proposed method reduces the average flow completion time by 8% to 18%. Chenzhao Huang, Xiaobin Tan, Shenzhi Yuan, Shiyin Zhu |
HPCC | 3 |
| 2024 | A Two-phase Encrypted Traffic Classification Scheme in Programmable Data PlaneabstractThe importance of encrypted traffic classification for network management and security is self-evident. The emergence of programmable data plane (PDP) technology makes it possible to directly implement encrypted traffic classification in the data plane, which can classify network traffics in line-rate. In this paper, we propose a two-phase encrypted traffic classification (TP-ETC) scheme in programmable data plane. In TP-ETC, Convolutional Neural Network (CNN) is employed for classifying highly similar traffic with high accuracy in the first phase, and Long Short-Term Memory (LSTM) model is responsible for classifying all remaining traffic with low storage overhead in the second phase, achieving the best balance between accuracy and storage overhead. We also design a feature extraction method suitable for PDP, effectively reducing the overhead of feature storage. In addition, we design a table segmentation algorithm to reduce the growth rate of table entries to a linear level. The experimental results demonstrate the superiority of the proposed scheme TP-ETC. Xiaobin Tan, Shenzhi Yuan, Mengxiang Li, Jiansong Wu, Quan Zheng 0002 |
ISPA | 3 |
| 2024 | Adaptive Gradient Data Partition and Route Selection for Distributed DNN Training
Bo Chai, Xiaobin Tan, Shenzhi Yuan, Guangge Jia, Qiushi Meng, Shiyin Zhu |
NPC (2) | 3 |