Guoyong Jiang

dblp:140/1027 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 BCD-Megatron: A Cost-Effective Training System for Large Language Models
Yunquan Zhang, Guoyong Jiang, Daning Cheng
ICDCS5
2026 Pulse: Training Acceleration for Large Diffusion Models with Automatic Pipeline Parallelism
Boran Sun, Guoyong Jiang, Yuechen Tao, Zhishu Che, Jieling Yu, Shan Chang, Huaxi Gu, Fangming Liu
ICDCS2
2024 Deep Reinforcement Learning Based Dynamic Flowlet Switching for DCN
abstract
Flowlet switching has been proven to be an effective technology for fine-grained load balancing in data center networks. However, flowlet detection based on static flowlet timeout values, lacks accuracy and effectiveness in complex network environments. In this paper, we propose a new deep reinforcement learning approach, called DRLet, to dynamically detect flowlets. DRLet offers two advantages: first, it provides dynamic flowlet timeout values to detect bursts into fine-grained flowlets; second, flowlet timeout values are automatically configured by the deep reinforcement learning agent, which only requires simple and measurable network states, instead of any prior knowledge, to achieve the pre-defined goal. With our approach, the flowlet timeout value dynamically matches the network load scenario, ensuring the accuracy and effectiveness of flowlet detection while suppressing packet reordering. Our results show that DRLet achieves superior performance compared to existing schemes based on static flowlet timeout values in both baseline and asymmetric topologies.
Xinglong Diao, Huaxi Gu, Wenting Wei, Guoyong Jiang, Baochun Li
IEEE Trans. Cloud Comput.4
2023 DRL-TAL: Deep Reinforcement Learning-Based Traffic-Aware Load Balancing in Data Center Networks
abstract
Load balancing in data center networks is crucial to effectively utilize network resources and enhance Quality of Service (QoS). Especially, the flowlet-level load balancing has been proven efficient in reducing latency and increasing throughput simultaneously. However, most existing work relying on empirical static timeout encounters performance degradation in dynamic network scenarios, due to a mismatch between the static timeout and changing traffic conditions. To address this problem, we propose a Deep Reinforcement Learning-Based Traffic-Aware Load Balancing scheme (DRL-TAL), which uses deep reinforcement learning (DRL) to update the flowlet timeout adaptively. The agent using a deep deterministic policy gradient (DDPG) algorithm continuously senses network throughput and generates the timeout threshold dynamically for the next time slot. The flowlet granularity is deployed for elephant flows to achieve a balance between throughput and disorder, where the timeout value relies on the threshold generated by the agent. Furthermore, the mice flow gets forwarded under packet granularity by selecting the port with the smallest queue length to ensure a shorter flow completion time. The results demonstrate that DRL-TAL performs impressively well in the symmetric topology, with no packet loss and minimal disorder under high load compared to the state-of-the-art schemes. Moreover, it significantly reduces flow completion time by up to 45% compared to Conga in the asymmetric topology.
Guoyong Jiang, Wenting Wei, Kun Wang 0001, Chengding Pang, Yong Liu 0038
GLOBECOM1