VLDB 2026 Research / reviewers in the wild / expert
Zhan Wang 0003
dblp:91/6269-3
· DBLP profile ↗
29ranked-venue papers
1as first author
15since 2021 · last 2026
0009-0003-4274-7671ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 23 · 1 first-author · 11 since 2021Computer networks · 4 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CCL-D: A High-Precision Diagnostic System for Slow and Hang Anomalies in Large-Scale Model TrainingabstractAs training scales grow, collective communication libraries (CCL) increasingly face anomalies arising from complex interactions among hardware, software, and environmental factors. These anomalies typically manifest as slow/hang communication, the most frequent and time-consuming category to diagnose. However, traditional diagnostic methods remain inaccurate and inefficient, frequently requiring hours or even days for root cause analysis. To address this, we propose CCL-D, a high-precision diagnostic system designed to detect and locate slow/hang anomalies in large-scale distributed training. CCL-D integrates a rank-level real-time probe with an intelligent decision analyzer. The probe measures cross-layer anomaly metrics using a lightweight distributed tracing framework to monitor communication traffic. The analyzer performs automated anomaly detection and root-cause location, precisely identifying the faulty GPU rank. Deployed on a 4,000-GPU cluster over one year, CCL-D achieved near-complete coverage of known slow/hang anomalies and pinpointed affected ranks within 6 minutes—substantially outperforming existing solutions. Yida Gu, Fakang Wang, Jianhao Fu, Zhenhang Sun, Qianyu Zhang 0001, Hairui Zhao 0002, Wenjing Huang 0002, Jinwu Yang, Yueyuan Zhou, Qian Zhao 0021, Haoxu Li, Zhan Wang 0003, Guangming Tan, Dingwen Tao |
PPoPP | 18 |
| 2026 | COCCL: A Collective Communication Library Supporting Easy Integration and Configuration of Customized Compression for Scalable LLM TrainingabstractCollective communication is critical to scaling large language model (LLM) training across various parallelism strategies, including data, tensor, and pipeline parallelism on GPU clusters. However, as model sizes and training scales increase, communication overhead is emerging as a major performance bottleneck. While compression is a promising mitigation strategy, existing solutions often lack user-transparency, hinder deployment and extensibility, and are not co-designed with communication algorithms. To address these limitations, we present COCCL, a high-performance collective communication library built on top of NCCL. COCCL introduces a novel programming model that can easily integrate compression into communication workflows with flexible configurability. It features a suite of compression-aware collective algorithms and runtime overlap mechanisms that mitigate error propagation and reduce computational overhead. We integrate well-established compression techniques into COCCL and tune the compression configurations during 3D-parallel training on GPT and Qwen models with up to 7 billion parameters. Using the optimal configuration (COCCL-3D), we achieve 1.24× throughput improvement while maintaining training accuracy. Haoran Kong, Hairui Zhao 0002, Shengkai Lyu, Xingjian Tian, Liyang Zhao, Zhuohan Chen, Fakang Wang, Zizhong Chen, Zhan Wang 0003, Guangming Tan, Dingwen Tao |
PPoPP | 12 |
| 2026 | Credit-Guided Congestion Control on Wafer-Scale On-Chip Networks for Molecular DynamicsabstractMolecular dynamics (MD) is a cornerstone of scientific computing, but strong scaling often collapses at high parallelism because communication is bursty and highly sensitive to tail latency. MD advances by repeating a fixed timestep loop (one iteration of force computation and state update), and performance is largely determined by how quickly timesteps complete. A key reason is that each timestep contains short, synchronized communication phases, followed by a global dependency before the next timestep. Wafer-scale chips (WSCs) offer cycle-level latency and high on-chip bandwidth, yet their 2D mesh fabrics can still suffer burst-induced queue buildup; existing wavelet scheduling relies on a static stride that either over-injects (triggering credit backpressure) or over-throttles (wasting bandwidth) as conditions evolve. Shixiong Qi, Zhan Wang 0003, Ning Kang 0007, Fan Yang 0096, Yuanzhe Wang, Guanglei Chen, Guangming Tan, Guojun Yuan |
SIGCOMM | 3 |
| 2026 | EPIC: Abstraction and Polymorphism of In-Network Collectives on Ethernet
Yitao Yuan, Jianglong Nie, Tianyu Bai, Ruizhe Zhou, Siyuan Cao, Xujie Fan, Yuchen Xu 0003, Junkai Chen, Chenqi Zhao, Nengyuan Zhang, Shaoke Fang, Jiangyuan Chen, Yuanfeng Chen, Zhan Wang 0003, Yuchao Zhang 0004, Yang Liu 0038, Xiangrui Yang 0002, Xiaohe Hu, Limin Xiao 0001, Weifeng Zhang 0003, Yazhu Lan, Jianbo Dong, Binzhang Fu, Wenfei Wu |
SIGCOMM | 15 |
| 2026 | AFCC: ACK-Based Fast Congestion Control in Lossless NetworksabstractCongestion control is vital for achieving ultra-low latency and ultra-high throughput in large-scale data centers. However, existing mechanisms like DCQCN and HPCC often respond slowly and inaccurately to congestion, leading to extended queuing delays and reduced throughput. The notification delay of nearly one round-trip time (RTT) exacerbates congestion issues. Furthermore, these algorithms fail to identify queue growth caused by priority flow control (PFC) pause frames from the downstream device in lossless networks, resulting in unnecessary rate reductions on innocent flows. In this paper, we introduce the ACK-based Fast Congestion Control (AFCC) mechanism, designed to achieve sub-RTT notification delays by leveraging ACK packets. To enhance responsiveness to last-hop congestion, AFCC incorporates the number of concurrent congested flows within each ACK. Furthermore, AFCC utilizes ACK’s In-band Network Telemetry (INT) to identify the paused phase, thereby facilitating accurate congestion detection. Experimental results show that AFCC effectively maintains high utilization and significantly reduces flow completion time by 31.2% compared to HPCC and 78.4% compared to DCQCN. Guojun Yuan, Zhan Wang 0003, Ninghui Sun, Guangming Tan |
IEEE Trans. Netw. | 4 |
| 2025 | MD-pipe: A Strong Scaling Enhanced Pipeline Architecture for Ab Initio Accuracy Molecular DynamicsabstractMolecular Dynamics (MD) simulations with first-principles accuracy are widely applied in various fields, including materials science and molecular pharmacology.Current research focus on reducing the solution time of ab initio molecular dynamics (AIMD) from both Ning Kang 0007, Guojun Yuan, Beining Zhang, Guanglei Chen, Jiayi Rao, Zhan Wang 0003, Weile Jia, Ninghui Sun, Guangming Tan |
ISCA | 10 |
| 2024 | Palos: Fair and Flexible Flow Scheduling on RNICabstractIn recent years, Remote Direct Memory Access (RDMA) has gained significant attraction within modern hyperscale data centers. However, RNIC fails to provide fine-grained performance isolation among network flows with different traffic patterns which co-exist in multi-tenant data centers and typically have various bandwidth, throughput and latency requirements.In this paper, we reveal that the drawbacks on isolation root in the packet-level flow scheduling mechanism implemented in the RNIC hardware. To solve this problem, we introduce Palos, a fair and flexible flow-scheduling mechanism. In the hardware layer, Palos adopts a data chunk based scheduling mechanism by reconstructing communication descriptors. The data chunk based scheduling diminishes the performance interference between large flows and small flows. Palos configures the scheduler in the software layer using a hierarchical weight setting to enable customized performance policy while preventing the configuration of users from interfering each other. Our experiments demonstrate that Palos provides better performance isolation and performance control flexibility compared with the commodity RDMA NIC and existing optimization framework. Zhenlong Ma, Fan Yang 0096, Ning Kang 0007, Guojun Yuan, Zhan Wang 0003, Ninghui Sun |
HPCC | 6 |
| 2024 | FNCC: Fast Notification Congestion Control in Data Center NetworksabstractCongestion control plays a pivotal role in large-scale data centers, facilitating ultra-low latency, high bandwidth, and optimal utilization. Even with the deployment of data center congestion control mechanisms such as DCQCN and HPCC, these algorithms often respond to congestion sluggishly. This sluggishness is primarily due to the slow notification of congestion. It takes almost one round-trip time (RTT) for the congestion information to reach the sender. In this paper, we introduce the Fast Notification Congestion Control (FNCC) mechanism, which achieves sub-RTT notification. FNCC leverages the acknowledgment packet (ACK) from the return path to carry in-network telemetry (INT) information of the request path, offering the sender more timely and accurate INT. To further accelerate the responsiveness of last-hop congestion control, we propose that the receiver notifies the sender of the number of concurrent congested flows, which can be used to adjust the congested flows to a fair rate quickly. Our experimental results demonstrate that FNCC reduces flow completion time by 27.4% and 88.9% compared to HPCC and DCQCN, respectively. Moreover, FNCC triggers minimal pause frames and maintains high utilization even at 400Gbps. Zhan Wang 0003, Fan Yang 0096, Ning Kang 0007, Zhenlong Ma, Guojun Yuan, Guangming Tan, Ninghui Sun |
ICPP | 2 |
| 2024 | Scaling Molecular Dynamics with ab initio Accuracy to 149 Nanoseconds per DayabstractPhysical phenomena such as chemical reactions, bond breaking, and phase transition require molecular dynamics (MD) simulation with ab initio accuracy ranging from milliseconds to microseconds. However, previous state-of-the-art neural network based MD packages such as DeePMD-kit can only reach 4.7 nanoseconds per day on the Fugaku supercomputer. In this paper, we present a novel node-based parallelization scheme to reduce communication by 81%, then optimize the computationally intensive kernels with sve-gemm and mixed precision. Finally, we implement intra-node load balance to further improve the scalability. Numerical results on the Fugaku supercomputer show that our work has significantly improved the time-to-solution of the DeePMD-kit by a factor of 31.7 x, reaching 149 nanoseconds per day on 12,000 computing nodes. This work has opened the door for millisecond simulation with ab initio accuracy within one week for the first time. Zhuoqiang Guo, Mingzhen Li 0001, Enji Li, Guojun Yuan, Zhan Wang 0003, Guangming Tan, Weile Jia |
SC | 8 |
| 2024 | Towards connection-scalable RNIC architecture
Ning Kang 0007, Zhan Wang 0003, Fan Yang 0096, Xiaoxiao Ma 0004, Zhenlong Ma, Guojun Yuan, Guangming Tan |
J. Supercomput. | 2 |
| 2023 | Understanding the Scalability Problem of RNIC Cache at the Micro-architecture LevelabstractHigh-performance Remote Direct Memory Access (RDMA) has been applied in High-Performance Computing (HPC) clusters and large-scale data centers as InfiniBand(IB) or RDMA over Converged Ethernet (RoCE). As the number of connections and the data space accessed by the network increases, the scalability problem of the RDMA Network Interface Card (RNIC) is getting worse because of cache misses on the RNIC. However, the RNIC is a black box for users. To understand the scalability problem, we give a detailed analysis of RNIC communication behavior based on the commercial RNIC driver and PCIe logic analyzer. We design a set of Cache Tests (CacheT) and a universal test methodology to profile the micro-architecture of the cache on the RNIC. We get statistical data by software tests and PCIe logic analyzer and obtain the cache micro-architecture parameters of commercial RNIC. Finally, we present an accurate performance model with 98% accuracy. Xiaoxiao Ma 0004, Fan Yang 0096, Zhan Wang 0003, Ning Kang 0007, Xunjun An |
ICC | 3 |
| 2023 | A Scalable RDMA Network Interface Card with Efficient Cache ManagementabstractRemote Direct Memory Access (RDMA) has been applied in large-scale clusters due to its high bandwidth and low latency features in recent decades. However, the scalability issue is a known intractable problem that the current RDMA Network Interface Card (RNIC) cannot overcome as the number of connections increases to thousands. The key to the scalability bottleneck is managing the connection information cached on the RNIC. To solve the scalability problem, firstly, we test and analyze the behavior of the commercial RNIC when the network performance declines in the case of large-scale connections. Further, we give the performance model of RNIC and point out that the cache design is the key to scalability issues. Then, we propose a Scalable RDMA NIC (ScalaRNIC) architecture with a non-blocking and priority-programmable cache design. Besides, our cache design supports parametric configuration by the extended API. ScalaRNIC can maintain$\mathrm{a}\approx 100\%$cache hit ratio for higher priority connections and keep message rates nearly equal to peak performance. In contrast, commercial RNIC's performance drops by 47% when there are a few thousand connections. Xiaoxiao Ma 0004, Fan Yang 0096, Zhan Wang 0003, Ning Kang 0007, Guojun Yuan, Xunjun An |
ISCAS | 3 |
| 2023 | Enhance the Strong Scaling of LAMMPS on FugakuabstractPhysical phenomenon such as protein folding requires simulation up to microseconds of physical time, which directly corresponds to the strong scaling of molecular dynamics(MD) on modern supercomputers. In this paper, we present a highly scalable implementation of the state-of-the-art MD code LAMMPS on Fugaku by exploiting the 6D mesh/torus topology of the TofuD network. Based on our detailed analysis of the MD communication pattern, we first adapt coarse-grained peer-to-peer ghost-region communication with uTofu interface, then further improve the scalability via fine-grained thread pool. Finally, Remote direct memory access (RDMA) primitives are utilized to avoid buffer overhead. Numerical results show that our optimized code can reduce 77% of the communication time, improving the performance of baseline LAMMPS by a factor of 2.9x and 2.2x for Lennard-Jones and embedded-atom method potentials when scaling to 36, 846 computing nodes. Our optimization techniques can also benefit other applications with stencil or domain decomposition methods. Zhuoqiang Guo, Shunchen Shi, Guangming Tan, Weile Jia, Guojun Yuan, Zhan Wang 0003 |
SC | 9 |
| 2022 | csRNA: Connection-Scalable RDMA NIC Architecture in Datacenter EnvironmentabstractRDMA has been widely deployed in datacenter networking as an ideal optimization strategy in recent years. Due to its mechanisms such as kernel bypass and hardware offloading, RDMA is expected to offer better performance than traditional kernel-based TCP/IP networking. However, the hardware offloading in RDMA requires the RDMA Network Interface Card (RNIC) to manage the connection metadata, and the limited on-chip memory size in RNIC leads to its limited connection scalability. When the RNIC maintains a large number of connections, its performance drops dramatically.This paper first finds that the head-of-line blocking in connection metadata management is a major factor affecting RNIC scalability. Based on the findings, we propose csRNA, a connection-scalable RNIC architecture that maintains near-peak performance when connection scales. To achieve the non-blocking RNIC processing path, csRNA utilizes a non-blocking connection scheduler to schedule different connections when blocking. Furthermore, using a non-blocking connection management model, csRNA departs from the conventional RNIC design by returning the prepared connections first. csRNA effectively avoids the performance degradation caused by the head-of-line blocking of connection metadata management when the number of connections increases. We implement and evaluate csRNA and demonstrate that with less on-chip memory occupancy, csRNA could still maintain near-peak performance when scaling up to more than 15,000 connections. Ning Kang 0007, Zhan Wang 0003, Fan Yang 0096, Xiaoxiao Ma 0004, Zhenlong Ma, Guojun Yuan, Guangming Tan |
ICCD | 2 |
| 2021 | A New Optoelectronic Hybrid Network Based on Scheduling Optimization of Optical LinksabstractThe emergence of exascale computers will represent a milestone in high-performance computing (HPC). Optoelectronic interconnections and configurable switches will change the traditional supercomputer architecture. However, new hardware is not easily adapted to dynamic running conditions. Based on scheduling optimization of optical links, we propose a new optoelectronic hybrid network, the software-defined network accelerator (sDNA), for an exascale computer. Our scheduling optimization contains an optical interconnection method and an adaptive routing method. The main contribution of our work is an extended edge forwarding index (E-EFI) optical interconnection method based on slow-switching optical devices. The optical link connections are established by evaluating the traffic offloading revenue for each optical link candidate. To support optical interconnection, sDNA selects a suitable routing strategy according to the job-schedule information and prior HPC application knowledge. We tested sDNA in a network simulator and a prototype exascale computer system using both the US Department of Energy (DOE) application and real-world communication benchmarks. The verification results for traffic offloading reveal that our optical interconnection method not only offloads traffic from electrical links to optical links but also avoids the congestion inherent to electrical links. sDNA maintains a throughput of more than 80 percent bandwidth and reduces the communication delay by 10 percent in our real prototype system and simulator. Thus, sDNA is an ideal candidate for accelerating the communication performance of exascale computers. En Shao, Guangming Tan, Zhan Wang 0003, Guojun Yuan, Zheng Cao 0003, Ninghui Sun |
IEEE Trans. Computers | 3 |
| 2019 | SwitchAgg: A Further Step Towards In-Network ComputationabstractMany distributed applications adopt a partition/aggregation pattern to achieve high performance and scalability. The aggregation process, which usually takes a large portion of the overall execution time, incurs large amount of network traffic and bottlenecks the system performance. To reduce network traffic, some researches take advantage of network devices to commit innetwork aggregation. However, these approaches use either special topology or middle-boxes, which cannot be easily deployed in current datacenters. The emerging programmable RMT switch brings us new opportunities to implement in-network computation task. However, we argue that the architecture of RMT switch is not suitable for in-network aggregation since it is designed primarily for implementing traditional network functions. In this paper, we first give a detailed analysis of in-network aggregation, and point out the key factor that affects the data reduction ratio. We then propose SwitchAgg, which is an innetwork aggregation system that is compatible with current datacenter infrastructures. We also evaluate the performance improvement we have gained from SwitchAgg. Our results show that, SwitchAgg can process data aggregation tasks at line rate and gives a high data reduction rate, which helps us to cut down network traffic and alleviate pressure on server CPU. In the system performance test, the job-completion-time can be reduced as much as 50%. Fan Yang 0096, Zhan Wang 0003, Xiaoxiao Ma 0004, Guojun Yuan, Xuejun An |
FPGA | 2 |
| 2019 | A New Traffic Offloading Method with Slow Switching Optical Device in Exascale ComputerabstractThe expected exascale computer will comprise tens of thousands of computing nodes and nearly 5000 interconnected nodes in years to come. Such a large-scale system will represent a milestone in the progress of High-Performance Computing (HPC). The more efficient network hardware, like optoelectronic interconnection and configurable switches, is reforming the traditional architecture of supercomputers. However, the present architecture containing new hardware is not easy to adapt to the dynamically running condition, because the newly developed hardware is normally unable to effectively improve overall performance. Here, we propose a new accelerated system called Software Defined Network Accelerator (sDNA) for the exascale computer. Inspired by edge forwarding index (EFI), the main contribution of our work is that it presents an extended EFI-based optical interconnection method with slow switching optical device. The optical link is connected by the evaluation of each optical link candidate's traffic offloading revenue. As the supporting method for optical interconnection, sDNA selects the most suitable routing configuration according to the job-schedule information and the prior-knowledge of HPC applications. We tested sDNA in a network simulator and a prototype system for the exascale computer, using both DOE application benchmarks and a real-world communication benchmark. From the result of verification of traffic offloading, we found that our optical interconnection method based on our extended EFI evaluation is not only able to offload the traffic from an electrical link to an optical link but is also able to avoid congestion inherent to electrical link. Furthermore, our experimental results show that sDNA maintains the throughput of more than 80% bandwidth and reduced the communication delay by 10% in our real prototype system and simulator. Together, our sDNA is an ideal candidate for accelerating communication performance of the exascale computer. En Shao, Guangming Tan, Zhan Wang 0003, Guojun Yuan, Ninghui Sun |
ICCD | 3 |
| 2019 | T2HT : Traffic-Driven Machine Learning Based Hierarchical Topology Generation ModelabstractIn high-performance computing (HPC) and distributed computing area, network performance greatly influenced application efficiency. However, due to the diversity of traffic patterns, the traditional network with fixed topology may achieve good performance under some applications, while performs poorly under other forms. Network reconfiguration technologies which can change the topology dynamically have been developed to obtain a balanced performance for different traffic patterns. Nonetheless, selecting an appropriate network topology from the wide variety of options remains difficult due to the complexity of analyzing traffic alongside topology performance characteristics. Traditional research focused on congestion estimation and specific parameter adjustment without reconfiguring the global topology. In this paper, we propose a generic Traffic to Hierarchical Topology (T2HT) method to analyze traffic patterns and choose an appropriate network configuration for the given traffic T2HT makes use of actual traffic data with a hierarchical model to predict network performance with a given topology and uses a machine learning (ML) algorithm to score the better options in order to determine the best topology. We performed 8000 simulations of dataset-topology combinations to verify the feasibility of the model. Our results show that T2HT achieved marked improvements with its recommendations, making it feasible for use in hierarchical network design. Under the DOE testbed, the throughput of the topology generated by T2HT can reach above 90% of theoretical limit(full connection), and the latency is improved by about 24.6% compared to typical topology 3D Torus with the same physical restrictions. Hongrui Zhu, Guojun Yuan, Guangming Tan, Zhan Wang 0003, Xuejun An |
ICPADS | 4 |
| 2019 | Wormhole optical network: a new architecture to solve long diameter problem in exascale computer
En Shao, Zhan Wang 0003, Guojun Yuan, Guangming Tan, Ninghui Sun |
CCF Trans. High Perform. Comput. | 2 |
| 2017 | Optimizing the Datapath for Key-value Middleware with NVMe SSDs over RDMA InterconnectsabstractIn-memory key-value store is a crucial building block of large-scale web architecture. Given the growth of the data volume and the need for low-latency responses, cost-effective storage expansion and fast large-message processing are the major challenges. In this paper, we explore the design of key-value middleware that takes advantage of modern NVMe SSDs and RDMA interconnects to achieve high performance without excessive DRAM deployment. We propose an all-in-userland approach to improve the data plane efficiency. Both NVMe and RDMA are interfaced directly from the user-space for effective data access and tailored data management. We present a low-latency storage extension framework based on NVMe and a new design of JVM-aware Memcache protocol based on RDMA. To further accelerate large-message transfer, we provide a hybrid communication protocol fusing Eager and Rendezvous schemas, and a united I/O staging approach to achieve maximum latency hiding through pipelining. As the benchmarking results indicate, with the non-negligible JVM overhead taken into account, our solution obtains comparable communication performance with the RDMA-Memcached released by the OSU. For SSD-involved operations, the latency decreases by up to 31% compared to the kernel-based I/O processing. Zhongqi An, Qiang Li 0045, Zhan Wang 0003, Zhigang Huo |
CLUSTER | 6 |
| 2017 | Regional Congestion Control in Datacenter NetworksabstractThe rapid deployment of cloud computing and online services poses great challenges for data center networks, and congestion control is one of the top concerns. Although numbers of proposals in different network layers have been put forward to alleviate the negative impact of congestion, the short-lived flows, which are latency-sensitive and constitute the majority of total traffic in data centers, still suffer severe performance degradation. Since the existing congestion control methods all rely on end hosts to perceive congestion and then adjust their network sending rate, the response time is relatively long when compared with the duration of short-lived flows, which increases latency significantly. In this paper, we propose RCC, a regional congestion control mechanism, which aims to respond to congestion more quickly and eliminate the mismatch mentioned above. Different from host-based mechanisms, RCC is implemented in the switch, which detects the congestion state and schedule the traffic around the congestion point locally, without sending feedback to the distal host. Evaluation has shown that, compared with the host-based mechanism, our method achieves better performance for short-lived flows and maintains stable buffer occupancy of the switch. In addition, mixed long- and short-lived flows which contend for the same bottleneck link can share the bandwidth more fairly. Fan Yang 0096, Zhan Wang 0003, Xiaoli Liu 0002, Zheng Cao 0003, Guojun Yuan, Xuejun An |
ICPADS | 2 |
| 2017 | Regional Congestion Mitigation in Lossless Datacenter Networks
Xiaoli Liu 0002, Fan Yang 0096, Yanan Jin, Zhan Wang 0003, Zheng Cao 0003, Ninghui Sun |
NPC | 4 |
| 2016 | Modeling Traffic of Big Data Platform for Large Scale Datacenter NetworksabstractPrior to deployment, network designers often use simulators to pre-evaluate the performance of designed network with artificial network traffic. The traditional way of separating network design from real applications will not only result in over-designed network configurations, wasting money and energy, but also miss the real network demands of applications, degrading system performance. In this paper, we provide a method to model the network traffic of current popular big data platforms, which can observably improve the matching between network design and applications. The new method extracts communication behavior from the popular big data applications and replays the behavior instead of the packet traces. Experiments show that the traffic generated by the model is almost match the real traffic and the model can easily scale to thousands of nodes. Zheng Cao 0003, Zhan Wang 0003, Dawei Zang, En Shao, Ninghui Sun |
ICPADS | 3 |
| 2015 | PROP: Using PCIe-Based RDMA to Accelerate Rack-Scale Communications in Data CentersabstractIn order to reduce the demands on bandwidth of core layer network, data center operators usually assign tasks of the same job to servers that are located in the same rack, leading to the fact that 80% of the traffic originated from servers retains in the same rack. As a result, providing sufficient network capacity inside racks becomes critical to the Quality-of-Service of current data center applications. In this paper, we propose PROP, a novel hybrid network architecture which leverages PCIe-based RDMA to reinforce rack-scale connectivity in data centers. In our design, intra-rack bulk data transfers will be accelerated by a dedicated high-bandwidth PCIe-compliant network while complemented with the existing Ethernet network. In addition, we develop a proprietary PCIe-based RDMA hardware which can allow the servers in the same rack to exchange data in main memory without involving the operating system and the processors. We also implement a software stack to enable existing socket-based applications to transparently utilize the proposed dedicated network system. As the preliminary stage, this paper focuses on exploiting the unique design point and implements an FPGA-based prototype to validate the technical feasibility of the proposed architecture. Dawei Zang, Zheng Cao 0003, Xiaoli Liu 0002, Lin Wang 0015, Zhan Wang 0003, Ninghui Sun |
ICPADS | 5 |
| 2014 | Building a large-scale direct network with low-radix routersabstractCommunication locality is an important characteristic of parallel applications. A great deal of research shows that utilizing the characteristic will favor most applications. Aiming at communication locality, we present a hierarchical direct network topology to accelerate neighbor communication. Combining mesh topology and complete graph topology, it can be used to optimize local communication and build large-scale network with low radix routers. Analyzing the characteristic of hierarchical topology, we find the presented topology has high cost performance and excellent expandability. We also design two minimum path routing algorithms and compare them with Mesh, Dragonfly and PERCS topologies. The results show the saturated throughput of hierarchical topology is nearly 40% with uniform random trace and 70% with local communication model of 4K nodes. That indicates high scalability for applications with local communication and cost efficiency for uniform random trace. Zheng Cao 0003, Zhiguo Fan, Zhan Wang 0003, Xiaoli Liu 0002, Li Qiang, Xuejun An, Ninghui Sun |
ICPADS | 4 |
| 2014 | HiNetSim: A Parallel Simulator for Large-Scale Hierarchical Direct Networks
Zhiguo Fan, Zheng Cao 0003, Xiaoli Liu 0002, Zhan Wang 0003, Dawei Zang, Xuejun An |
NPC | 5 |
| 2014 | An Intra-Server Interconnect Fabric for Heterogeneous Computing
Zheng Cao 0003, Xiaoli Liu 0002, Qiang Li 0045, Zhan Wang 0003, Xuejun An |
J. Comput. Sci. Technol. | 5 |
| 2013 | cHPP controller: A High Performance Hyper-node Hardware AcceleratorabstractThe high-density blade server provides an attractive solution for the rapid increasing demand on computing. The degree of parallelism inside a blade enclosure nowadays has reach up to hundreds of cores. In such parallelism, it is necessary to accelerate communications inside a blade enclosure. However, commercial products seldom set foot in the optimization based on hardware. A hyper-node controller is proposed to provide a low overhead and high performance interconnection based on PCIe, which supports global address space, user-level communication, and efficient communication primitives. Furthermore, the efficient sharing of I/O resource is another goal of this design. The prototype of the hyper-node controller is implemented in FPGA. The testing results show the lowest latency is only 1.242us and the highest bandwidth is 3.19GB/s, which is almost 99.7% of the theoretic peak bandwidth. Zheng Cao 0003, Zhan Wang 0003, Xiaoli Liu 0002, Xuejun An, Ninghui Sun |
PDCAT | 4 |
| 2012 | Design of Hardware-Based Communication Performance Measurement ToolabstractWith the popularity and development of heterogeneous computing, proper communication performance measurement tools are needed to explore new communication patterns under heterogeneous computing systems and optimize program's performance. This paper proposes a hardware-based communication performance measurement tool, named as HCPM, which brings little influence on original program, and can collect communication traces generated by heterogeneous processors which implement PCIe or HT as their system bus. HCPM firstly provides basic communication primitives to set up a communication system. Then based on these primitives, it collects communication trace. Real-time collected traces are transmitted to a dedicated computer for further analysis. Evaluation shows that with the use of proper compression in hardware, HCPM can transmit at least five processors' communication traces with a single Gigabit Ethernet link. Zhan Wang 0003, Zheng Cao 0003, Xiaoli Liu 0002, Xuejun An |
CLUSTER | 1 |