EDBT 2026 Demo / reviewers in the wild / expert
Shiyin Zhu
dblp:342/9964
· DBLP profile ↗
8ranked-venue papers
0as first author
8since 2021 · last 2026
0009-0008-5219-0885ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 6 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reducing Cross-Pod Communication Overhead for MoE Model Training With Hybrid Parallelism in Multi-Tenant ClustersabstractThe massive parameter scale of sparsely-activated Mixture-of-Experts (MoE) models necessitates distributed training with hybrid parallelism. Placing such training tasks,i.e.mapping the logical partitions of an MoE model to available physical NPUs, is challenging. Due to the bandwidth and latency discrepancies between intra- and inter- Pods, the cross-Pod communication usually becomes a bottleneck. The high dispersion of NPUs in multi-tenant clusters exacerbates this issue further. However, a paucity of studies has considered the cross-Pod model placement problem. To address this challenge, we propose a novel model placement scheme tailored for MoE model training with hybrid parallelism in multi-tenant clusters. By quantifying the cross-Pod communication overhead incurred during MoE model training, the model placement is formulated as a 0-1 integer quadratic problem, which is NP hard. Motivated by the traffic difference between different parallelism, we decompose this problem into two subproblems. To solve the subproblems, we propose a lightweight two-stage algorithm based on Best-Fit strategy and neighborhood search. Experiments under different models and network topologies show that our model placement scheme can reduce cross-Pod traffic by 35.9% and cut communication time by 18.7% compared to state-of-the-art methods. Huihuang Qin, Shuangwu Chen, Tao Zhang 0170, Ziyang Zou, Xiaobin Tan, Shiyin Zhu, Jian Yang 0014 |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2025 | INTMCC: An In-Network Telemetry-Based Multipath Congestion Control Algorithm for Data Center Networks
Chenzhao Huang, Yingying Zeng, Shiyin Zhu, Xiaobin Tan |
ICA3PP (5) | 5 |
| 2025 | Straggler Dynamic Management for Distributed DNN TrainingabstractStraggler nodes are a major bottleneck in large-scale distributed training, degrading efficiency and stability. However, current solutions, including In-Network Aggregation (INA), lack the adaptability to effectively manage these stragglers in dynamic environments. This paper proposes Straggler Dynamic Management (SDM), an adaptive method for large-scale distributed training that performs dynamic straggler management by coordinating the data and control planes to achieve accurate, time-based detection and efficient mitigation via a performanceaware redundancy strategy and semi-asynchronous aggregation. SDM manages stragglers through a coordinated architecture that decouples the data and control planes for efficient detection and response. It leverages the data plane to estimate each node's remaining completion time, ensuring accurate and low-overhead straggler identification. The control plane then mitigates their impact using two key strategies: a performance-aware redundancy scheme to reduce waiting delays, and a semi-asynchronous aggregation mechanism that dynamically adjusts synchronization to alleviate gradient staleness and improve model convergence. We implement and deploy SDM on a real-world hardware testbed and evaluate its performance under various straggler scenarios. Experimental results demonstrate that SDM significantly improves training efficiency and convergence stability in the presence of straggler nodes, particularly when multiple stragglers occur simultaneously, exhibiting greater robustness and adaptability than existing methods. Tiance Li, Bo Chai, Xiaobin Tan, Shenzhi Yuan, Kexin Ju 0003, Shiyin Zhu |
ICPADS | 6 |
| 2025 | Beamforming Design for RIS-Aided ISCC in Internet of Vehicles SystemsabstractWith the development of communication technology, the Internet of Vehicles (IoV) is becoming increasingly important, enabling vehicle-to-everything communication for real-time information exchange and processing, thereby significantly enhancing traffic efficiency and safety. In this article, we consider a joint beamforming design problem in IoV, where the objective is to minimize transmission power, computation rate, and communication rate within the integrated sensing, communication, and computation (ISCC) framework. Moreover, reconfigurable intelligent surfaces (RISs) can provide additional spatial degrees of freedom to enhance the performance of ISCC systems in IoV within limited spectrum, energy resources, and complex interference management. To address the joint beamforming design problem, we present a cooperative beamforming algorithm called weight performance optimization (WPO), which explores three single-objective optimization problems in sensing, computation, and communication within the IoV context, using alternating optimization (AO) to simplify and solve these foundational elements of the WPO framework within limited resources and vehicle mobility, enhancing resource distribution while maintaining a balance between power efficiency and system performance. Numerical results demonstrate the efficiency and potential advantages of our proposed algorithms. Specifically, the results show that the sensing error of the WPO algorithm is reduced by up to 92.2% compared to existing popular algorithms, while the computation rate and communication rate are increased by more than 29.5% and 23.9%, respectively. Ruihang Yang, Dezhi Wang 0001, Shiyin Zhu, Jianrong Bao, Zhaohui Yang 0001, Chongwen Huang |
IEEE Internet Things J. | 3 |
| 2025 | Joint Dynamic Data and Model Parallelism for Distributed Training of DNNs Over Heterogeneous InfrastructureabstractDistributed training of deep neural networks (DNNs) suffers from efficiency declines in dynamic heterogeneous environments, due to the resource wastage brought by the straggler problem in data parallelism (DP) and pipeline bubbles in model parallelism (MP). Additionally, the limited resource availability requires a trade-off between training performance and long-term costs, particularly in online settings. To address these challenges, this article presents a novel online approach to maximize long-term training efficiency in heterogeneous environments through uneven data assignment and communication-aware model partitioning. A group-based hierarchical architecture combining DP and MP is developed to balance discrepant computation and communication capabilities, and offer a flexible parallel mechanism. In order to jointly optimize the performance and long-term cost of the online DL training process, we formulate this problem as a stochastic optimization with time-averaged constraints. By utilizing Lyapunov’s stochastic network optimization theory, we decompose it into several instantaneous sub-optimizations, and devise an effective online solution to address them based on tentative searching and linear solving. We have implemented a prototype system and evaluated the effectiveness of our solution based on realistic experiments, reducing batch training time by up to 68.59% over state-of-the-art methods. Xiaofeng Jiang, Xiaobin Tan, Huasen He, Shiyin Zhu, Jian Yang 0014 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2024 | dotPS: Disorder Tolerant Load Balancing Scheme for Datacenter NetworkabstractRecently, with the development of AI technology, load balancing methods are commonly integrated to enhance datacenter network throughput within a spine-leaf architecture. However, traditional load balancing inevitably leads to packet disorder, affecting system efficiency. Considering that network applications can actually tolerate a certain disorder degree, we think the transmission efficiency will be improved without requiring additional processing of packets within a certain disorder tolerance threshold. Thus, we propose a disorder tolerant load balancing scheme called Disorder Tolerant Packet Spray, dotPS. We design the architecture of dotPS, controlling the degree of disorder tolerance to improve transmission efficiency, and propose the concept of a group interval which can be dynamically adjusted as a measurement of disorder tolerance. Then we design a load balancing algorithm based on group interval, requiring strict ordering for inter-group packets while tolerating disorder in intra-group packets. Finally, simulation results dedicate that compared to the state-of-the-art load balancing methods under different network loads, the proposed method reduces the average flow completion time by 8% to 18%. Chenzhao Huang, Xiaobin Tan, Shenzhi Yuan, Shiyin Zhu |
HPCC | 6 |
| 2024 | Adaptive Gradient Data Partition and Route Selection for Distributed DNN Training
Bo Chai, Xiaobin Tan, Shenzhi Yuan, Guangge Jia, Qiushi Meng, Shiyin Zhu |
NPC (2) | 7 |
| 2023 | Poster: Programmable Cycle-Specified Queue for Deterministic NetworkingabstractThe emerging time-critical applications pose intense demands for enabling large-scale deterministic networks. In this paper, we propose a new Programmable Cycle-Specified Queue (PCSQ) for wide-area deterministic packet scheduling. We implement the first end-to-end high-precision rotation dequeuing, which enables microsecond-level time slot resource reservation (noted as T) and especially jitter control of up to 2T. We prototype the PCSQ scheduler on an FPGA. The PCSQ-enabled switches can guarantee bounded delay and jitter transmission on a realistic testbed. Yudong Huang, Shuo Wang 0006, Shiyin Zhu, Guoyu Peng, Xinyuan Zhang 0011, Tian Pan 0001, Tao Huang 0005, Zuopin Cheng, Daorong Guo, Lianqing Zhang, Juyan Lei, Liangzhang Xu, Wei Wang 0494, Xinmin Liu, Xuejun You, Yunjie Liu 0001 |
SIGCOMM | 3 |