Guojun Yuan

dblp:220/7047 · DBLP profile ↗
← Back
16ranked-venue papers
0as first author
11since 2021 · last 2026
0000-0002-4864-7677ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 9 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Credit-Guided Congestion Control on Wafer-Scale On-Chip Networks for Molecular Dynamics
abstract
Molecular dynamics (MD) is a cornerstone of scientific computing, but strong scaling often collapses at high parallelism because communication is bursty and highly sensitive to tail latency. MD advances by repeating a fixed timestep loop (one iteration of force computation and state update), and performance is largely determined by how quickly timesteps complete. A key reason is that each timestep contains short, synchronized communication phases, followed by a global dependency before the next timestep. Wafer-scale chips (WSCs) offer cycle-level latency and high on-chip bandwidth, yet their 2D mesh fabrics can still suffer burst-induced queue buildup; existing wavelet scheduling relies on a static stride that either over-injects (triggering credit backpressure) or over-throttles (wasting bandwidth) as conditions evolve.
Shixiong Qi, Zhan Wang 0003, Ning Kang 0007, Fan Yang 0096, Yuanzhe Wang, Guanglei Chen, Guangming Tan, Guojun Yuan
SIGCOMM12
2026 AFCC: ACK-Based Fast Congestion Control in Lossless Networks
abstract
Congestion control is vital for achieving ultra-low latency and ultra-high throughput in large-scale data centers. However, existing mechanisms like DCQCN and HPCC often respond slowly and inaccurately to congestion, leading to extended queuing delays and reduced throughput. The notification delay of nearly one round-trip time (RTT) exacerbates congestion issues. Furthermore, these algorithms fail to identify queue growth caused by priority flow control (PFC) pause frames from the downstream device in lossless networks, resulting in unnecessary rate reductions on innocent flows. In this paper, we introduce the ACK-based Fast Congestion Control (AFCC) mechanism, designed to achieve sub-RTT notification delays by leveraging ACK packets. To enhance responsiveness to last-hop congestion, AFCC incorporates the number of concurrent congested flows within each ACK. Furthermore, AFCC utilizes ACK’s In-band Network Telemetry (INT) to identify the paused phase, thereby facilitating accurate congestion detection. Experimental results show that AFCC effectively maintains high utilization and significantly reduces flow completion time by 31.2% compared to HPCC and 78.4% compared to DCQCN.
Guojun Yuan, Zhan Wang 0003, Ninghui Sun, Guangming Tan
IEEE Trans. Netw.2
2025 MD-pipe: A Strong Scaling Enhanced Pipeline Architecture for Ab Initio Accuracy Molecular Dynamics
abstract
Molecular Dynamics (MD) simulations with first-principles accuracy are widely applied in various fields, including materials science and molecular pharmacology.Current research focus on reducing the solution time of ab initio molecular dynamics (AIMD) from both
Ning Kang 0007, Guojun Yuan, Beining Zhang, Guanglei Chen, Jiayi Rao, Zhan Wang 0003, Weile Jia, Ninghui Sun, Guangming Tan
ISCA2
2024 Palos: Fair and Flexible Flow Scheduling on RNIC
abstract
In recent years, Remote Direct Memory Access (RDMA) has gained significant attraction within modern hyperscale data centers. However, RNIC fails to provide fine-grained performance isolation among network flows with different traffic patterns which co-exist in multi-tenant data centers and typically have various bandwidth, throughput and latency requirements.In this paper, we reveal that the drawbacks on isolation root in the packet-level flow scheduling mechanism implemented in the RNIC hardware. To solve this problem, we introduce Palos, a fair and flexible flow-scheduling mechanism. In the hardware layer, Palos adopts a data chunk based scheduling mechanism by reconstructing communication descriptors. The data chunk based scheduling diminishes the performance interference between large flows and small flows. Palos configures the scheduler in the software layer using a hierarchical weight setting to enable customized performance policy while preventing the configuration of users from interfering each other. Our experiments demonstrate that Palos provides better performance isolation and performance control flexibility compared with the commodity RDMA NIC and existing optimization framework.
Zhenlong Ma, Fan Yang 0096, Ning Kang 0007, Guojun Yuan, Zhan Wang 0003, Ninghui Sun
HPCC5
2024 FNCC: Fast Notification Congestion Control in Data Center Networks
abstract
Congestion control plays a pivotal role in large-scale data centers, facilitating ultra-low latency, high bandwidth, and optimal utilization. Even with the deployment of data center congestion control mechanisms such as DCQCN and HPCC, these algorithms often respond to congestion sluggishly. This sluggishness is primarily due to the slow notification of congestion. It takes almost one round-trip time (RTT) for the congestion information to reach the sender. In this paper, we introduce the Fast Notification Congestion Control (FNCC) mechanism, which achieves sub-RTT notification. FNCC leverages the acknowledgment packet (ACK) from the return path to carry in-network telemetry (INT) information of the request path, offering the sender more timely and accurate INT. To further accelerate the responsiveness of last-hop congestion control, we propose that the receiver notifies the sender of the number of concurrent congested flows, which can be used to adjust the congested flows to a fair rate quickly. Our experimental results demonstrate that FNCC reduces flow completion time by 27.4% and 88.9% compared to HPCC and DCQCN, respectively. Moreover, FNCC triggers minimal pause frames and maintains high utilization even at 400Gbps.
Zhan Wang 0003, Fan Yang 0096, Ning Kang 0007, Zhenlong Ma, Guojun Yuan, Guangming Tan, Ninghui Sun
ICPP6
2024 Scaling Molecular Dynamics with ab initio Accuracy to 149 Nanoseconds per Day
abstract
Physical phenomena such as chemical reactions, bond breaking, and phase transition require molecular dynamics (MD) simulation with ab initio accuracy ranging from milliseconds to microseconds. However, previous state-of-the-art neural network based MD packages such as DeePMD-kit can only reach 4.7 nanoseconds per day on the Fugaku supercomputer. In this paper, we present a novel node-based parallelization scheme to reduce communication by 81%, then optimize the computationally intensive kernels with sve-gemm and mixed precision. Finally, we implement intra-node load balance to further improve the scalability. Numerical results on the Fugaku supercomputer show that our work has significantly improved the time-to-solution of the DeePMD-kit by a factor of 31.7 x, reaching 149 nanoseconds per day on 12,000 computing nodes. This work has opened the door for millisecond simulation with ab initio accuracy within one week for the first time.
Zhuoqiang Guo, Mingzhen Li 0001, Enji Li, Guojun Yuan, Zhan Wang 0003, Guangming Tan, Weile Jia
SC7
2024 Towards connection-scalable RNIC architecture
Ning Kang 0007, Zhan Wang 0003, Fan Yang 0096, Xiaoxiao Ma 0004, Zhenlong Ma, Guojun Yuan, Guangming Tan
J. Supercomput.6
2023 A Scalable RDMA Network Interface Card with Efficient Cache Management
abstract
Remote Direct Memory Access (RDMA) has been applied in large-scale clusters due to its high bandwidth and low latency features in recent decades. However, the scalability issue is a known intractable problem that the current RDMA Network Interface Card (RNIC) cannot overcome as the number of connections increases to thousands. The key to the scalability bottleneck is managing the connection information cached on the RNIC. To solve the scalability problem, firstly, we test and analyze the behavior of the commercial RNIC when the network performance declines in the case of large-scale connections. Further, we give the performance model of RNIC and point out that the cache design is the key to scalability issues. Then, we propose a Scalable RDMA NIC (ScalaRNIC) architecture with a non-blocking and priority-programmable cache design. Besides, our cache design supports parametric configuration by the extended API. ScalaRNIC can maintain$\mathrm{a}\approx 100\%$cache hit ratio for higher priority connections and keep message rates nearly equal to peak performance. In contrast, commercial RNIC's performance drops by 47% when there are a few thousand connections.
Xiaoxiao Ma 0004, Fan Yang 0096, Zhan Wang 0003, Ning Kang 0007, Guojun Yuan, Xunjun An
ISCAS5
2023 Enhance the Strong Scaling of LAMMPS on Fugaku
abstract
Physical phenomenon such as protein folding requires simulation up to microseconds of physical time, which directly corresponds to the strong scaling of molecular dynamics(MD) on modern supercomputers. In this paper, we present a highly scalable implementation of the state-of-the-art MD code LAMMPS on Fugaku by exploiting the 6D mesh/torus topology of the TofuD network. Based on our detailed analysis of the MD communication pattern, we first adapt coarse-grained peer-to-peer ghost-region communication with uTofu interface, then further improve the scalability via fine-grained thread pool. Finally, Remote direct memory access (RDMA) primitives are utilized to avoid buffer overhead. Numerical results show that our optimized code can reduce 77% of the communication time, improving the performance of baseline LAMMPS by a factor of 2.9x and 2.2x for Lennard-Jones and embedded-atom method potentials when scaling to 36, 846 computing nodes. Our optimization techniques can also benefit other applications with stencil or domain decomposition methods.
Zhuoqiang Guo, Shunchen Shi, Guangming Tan, Weile Jia, Guojun Yuan, Zhan Wang 0003
SC8
2022 csRNA: Connection-Scalable RDMA NIC Architecture in Datacenter Environment
abstract
RDMA has been widely deployed in datacenter networking as an ideal optimization strategy in recent years. Due to its mechanisms such as kernel bypass and hardware offloading, RDMA is expected to offer better performance than traditional kernel-based TCP/IP networking. However, the hardware offloading in RDMA requires the RDMA Network Interface Card (RNIC) to manage the connection metadata, and the limited on-chip memory size in RNIC leads to its limited connection scalability. When the RNIC maintains a large number of connections, its performance drops dramatically.This paper first finds that the head-of-line blocking in connection metadata management is a major factor affecting RNIC scalability. Based on the findings, we propose csRNA, a connection-scalable RNIC architecture that maintains near-peak performance when connection scales. To achieve the non-blocking RNIC processing path, csRNA utilizes a non-blocking connection scheduler to schedule different connections when blocking. Furthermore, using a non-blocking connection management model, csRNA departs from the conventional RNIC design by returning the prepared connections first. csRNA effectively avoids the performance degradation caused by the head-of-line blocking of connection metadata management when the number of connections increases. We implement and evaluate csRNA and demonstrate that with less on-chip memory occupancy, csRNA could still maintain near-peak performance when scaling up to more than 15,000 connections.
Ning Kang 0007, Zhan Wang 0003, Fan Yang 0096, Xiaoxiao Ma 0004, Zhenlong Ma, Guojun Yuan, Guangming Tan
ICCD6
2021 A New Optoelectronic Hybrid Network Based on Scheduling Optimization of Optical Links
abstract
The emergence of exascale computers will represent a milestone in high-performance computing (HPC). Optoelectronic interconnections and configurable switches will change the traditional supercomputer architecture. However, new hardware is not easily adapted to dynamic running conditions. Based on scheduling optimization of optical links, we propose a new optoelectronic hybrid network, the software-defined network accelerator (sDNA), for an exascale computer. Our scheduling optimization contains an optical interconnection method and an adaptive routing method. The main contribution of our work is an extended edge forwarding index (E-EFI) optical interconnection method based on slow-switching optical devices. The optical link connections are established by evaluating the traffic offloading revenue for each optical link candidate. To support optical interconnection, sDNA selects a suitable routing strategy according to the job-schedule information and prior HPC application knowledge. We tested sDNA in a network simulator and a prototype exascale computer system using both the US Department of Energy (DOE) application and real-world communication benchmarks. The verification results for traffic offloading reveal that our optical interconnection method not only offloads traffic from electrical links to optical links but also avoids the congestion inherent to electrical links. sDNA maintains a throughput of more than 80 percent bandwidth and reduces the communication delay by 10 percent in our real prototype system and simulator. Thus, sDNA is an ideal candidate for accelerating the communication performance of exascale computers.
En Shao, Guangming Tan, Zhan Wang 0003, Guojun Yuan, Zheng Cao 0003, Ninghui Sun
IEEE Trans. Computers4
2019 SwitchAgg: A Further Step Towards In-Network Computation
abstract
Many distributed applications adopt a partition/aggregation pattern to achieve high performance and scalability. The aggregation process, which usually takes a large portion of the overall execution time, incurs large amount of network traffic and bottlenecks the system performance. To reduce network traffic, some researches take advantage of network devices to commit innetwork aggregation. However, these approaches use either special topology or middle-boxes, which cannot be easily deployed in current datacenters. The emerging programmable RMT switch brings us new opportunities to implement in-network computation task. However, we argue that the architecture of RMT switch is not suitable for in-network aggregation since it is designed primarily for implementing traditional network functions. In this paper, we first give a detailed analysis of in-network aggregation, and point out the key factor that affects the data reduction ratio. We then propose SwitchAgg, which is an innetwork aggregation system that is compatible with current datacenter infrastructures. We also evaluate the performance improvement we have gained from SwitchAgg. Our results show that, SwitchAgg can process data aggregation tasks at line rate and gives a high data reduction rate, which helps us to cut down network traffic and alleviate pressure on server CPU. In the system performance test, the job-completion-time can be reduced as much as 50%.
Fan Yang 0096, Zhan Wang 0003, Xiaoxiao Ma 0004, Guojun Yuan, Xuejun An
FPGA4
2019 A New Traffic Offloading Method with Slow Switching Optical Device in Exascale Computer
abstract
The expected exascale computer will comprise tens of thousands of computing nodes and nearly 5000 interconnected nodes in years to come. Such a large-scale system will represent a milestone in the progress of High-Performance Computing (HPC). The more efficient network hardware, like optoelectronic interconnection and configurable switches, is reforming the traditional architecture of supercomputers. However, the present architecture containing new hardware is not easy to adapt to the dynamically running condition, because the newly developed hardware is normally unable to effectively improve overall performance. Here, we propose a new accelerated system called Software Defined Network Accelerator (sDNA) for the exascale computer. Inspired by edge forwarding index (EFI), the main contribution of our work is that it presents an extended EFI-based optical interconnection method with slow switching optical device. The optical link is connected by the evaluation of each optical link candidate's traffic offloading revenue. As the supporting method for optical interconnection, sDNA selects the most suitable routing configuration according to the job-schedule information and the prior-knowledge of HPC applications. We tested sDNA in a network simulator and a prototype system for the exascale computer, using both DOE application benchmarks and a real-world communication benchmark. From the result of verification of traffic offloading, we found that our optical interconnection method based on our extended EFI evaluation is not only able to offload the traffic from an electrical link to an optical link but is also able to avoid congestion inherent to electrical link. Furthermore, our experimental results show that sDNA maintains the throughput of more than 80% bandwidth and reduced the communication delay by 10% in our real prototype system and simulator. Together, our sDNA is an ideal candidate for accelerating communication performance of the exascale computer.
En Shao, Guangming Tan, Zhan Wang 0003, Guojun Yuan, Ninghui Sun
ICCD4
2019 T2HT : Traffic-Driven Machine Learning Based Hierarchical Topology Generation Model
abstract
In high-performance computing (HPC) and distributed computing area, network performance greatly influenced application efficiency. However, due to the diversity of traffic patterns, the traditional network with fixed topology may achieve good performance under some applications, while performs poorly under other forms. Network reconfiguration technologies which can change the topology dynamically have been developed to obtain a balanced performance for different traffic patterns. Nonetheless, selecting an appropriate network topology from the wide variety of options remains difficult due to the complexity of analyzing traffic alongside topology performance characteristics. Traditional research focused on congestion estimation and specific parameter adjustment without reconfiguring the global topology. In this paper, we propose a generic Traffic to Hierarchical Topology (T2HT) method to analyze traffic patterns and choose an appropriate network configuration for the given traffic T2HT makes use of actual traffic data with a hierarchical model to predict network performance with a given topology and uses a machine learning (ML) algorithm to score the better options in order to determine the best topology. We performed 8000 simulations of dataset-topology combinations to verify the feasibility of the model. Our results show that T2HT achieved marked improvements with its recommendations, making it feasible for use in hierarchical network design. Under the DOE testbed, the throughput of the topology generated by T2HT can reach above 90% of theoretical limit(full connection), and the latency is improved by about 24.6% compared to typical topology 3D Torus with the same physical restrictions.
Hongrui Zhu, Guojun Yuan, Guangming Tan, Zhan Wang 0003, Xuejun An
ICPADS2
2019 Wormhole optical network: a new architecture to solve long diameter problem in exascale computer
En Shao, Zhan Wang 0003, Guojun Yuan, Guangming Tan, Ninghui Sun
CCF Trans. High Perform. Comput.3
2017 Regional Congestion Control in Datacenter Networks
abstract
The rapid deployment of cloud computing and online services poses great challenges for data center networks, and congestion control is one of the top concerns. Although numbers of proposals in different network layers have been put forward to alleviate the negative impact of congestion, the short-lived flows, which are latency-sensitive and constitute the majority of total traffic in data centers, still suffer severe performance degradation. Since the existing congestion control methods all rely on end hosts to perceive congestion and then adjust their network sending rate, the response time is relatively long when compared with the duration of short-lived flows, which increases latency significantly. In this paper, we propose RCC, a regional congestion control mechanism, which aims to respond to congestion more quickly and eliminate the mismatch mentioned above. Different from host-based mechanisms, RCC is implemented in the switch, which detects the congestion state and schedule the traffic around the congestion point locally, without sending feedback to the distal host. Evaluation has shown that, compared with the host-based mechanism, our method achieves better performance for short-lived flows and maintains stable buffer occupancy of the switch. In addition, mixed long- and short-lived flows which contend for the same bottleneck link can share the bandwidth more fairly.
Fan Yang 0096, Zhan Wang 0003, Xiaoli Liu 0002, Zheng Cao 0003, Guojun Yuan, Xuejun An
ICPADS5