Zheng Cao 0003

dblp:11/5229-3 · DBLP profile ↗
← Back
35ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0002-1565-3683ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 1 first-author · 5 since 2021Computer networks · 5 · 2 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Software engineering, systems software and programming languages · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 eGPU: Production-Scale Elastic Sharing Over 10,000 GPUs
abstract
As the cost of GPUs continues to rise, GPU-sharing solutions have become increasingly important for improving efficiency and maximizing resource utilization. At the same time, large-scale operational deployments of such solutions remain relatively less explored, especially in heterogeneous production environments where workload dynamics and orchestration complexity introduce new practical considerations. In this paper, we introduce eGPU, an elastic, efficient, and scalable GPU-sharing framework tailored for production-scale concurrent machine learning (ML) training and inference. eGPU enables fine-grained, runtime-adjustable sharing of GPUs across multiple jobs, while preserving high resource utilization and fault isolation. To address communication bottlenecks, eGPU supports native NVLink/NCCL-based communication between shared GPU instances, capabilities that are limited or unavailable in many existing designs. Built with production deployment in mind, eGPU integrates with Kubernetes (K8s) to support large-scale orchestration. It has been deployed and running stably in production clusters with over$\text{1 0, 0 0 0 ~ G P U s}$for five years. Our evaluation results show that eGPU achieves elastic and precise control over instance sizes, improves job efficiency by 21 % to 31% than SOTA sharing solutions, saves the number of GPUs required by up to$8 \times$, and improves cluster GPU utilization by more than$\mathrm{3} \times$.
Xiaochuan Tang, Hao Qi 0008, Jianbo Dong, Yinghao Yu, Zhennan Xue, Daocheng Ying, Zheng Cao 0003, Xiaoyi Lu 0001
HPCA8
2024 Kspeed: Beating I/O Bottlenecks of Data Provisioning for RDMA Training Clusters
abstract
The rapidly-increasing computing power of GPUs has rendered the I/O subsystem a bottleneck for distributed deep learning (DL) training. Currently, substantial data preprocessing work (e.g., decoding) has to be conducted on CPUs for a wide range of training scenarios such as computer vision (CV) and audio. Unfortunately, the involvement of training nodes' host memory and/or CPUs on the critical path of loading data to GPUs incurs significant GPU stalls in modern RDMA training clusters, because CPUs are much slower than GPUs and the connection from PCIe switches to host memory tends to suffer from incast problems. Moreover, this also incurs high CPU usage and resource contention, which consequently causes data loading performance variation and stragglers. This paper presents KSpeed, a novel data provisioning framework for large-scale RDMA training clusters. As many data preprocessing tasks need to be done by CPUs, KSpeed organizes host memory and CPU resources in the cluster to build a disaggregated memory/CPU pool, where the nodes can read raw input data from backend storage to their host memory, preprocess the data by their CPUs if necessary, and write cached/preprocessed data (on demand) directly to the training workers' GPU memory to minimize GPU stalls. KSpeed leverages the multi-rail RDMA network to eliminate unnecessary memory copies, interference, and congestion. Evaluation on a 96-GPU cluster shows that KSpeed delivers$5.4 \times \sim 100 \times$higher data loading performance over the state-of-the-art designs (DPP and Alluxio). KSpeed achieves near-linear scalability as the GPU number increases from 8 to 512.
Jianbo Dong, Hao Qi 0008, Tianjing Xu, Xiaoli Liu 0002, Rongyao Wang, Xiaoyi Lu 0001, Zheng Cao 0003, Binzhang Fu
ICNP8
2023 Fisc: A Large-scale Cloud-native-oriented File System
Qiang Li 0045, Lulu Chen, Xiaoliang Wang 0001, Qiao Xiang, Wenhui Yao, Minfei Huang, Puyuan Yang, Shanyang Liu, Zhaosheng Zhu, Huayong Wang, Haonan Qiu, Derui Liu, Shaozong Liu, Yaohui Wu, Zhiwu Wu, Zicheng Luo, Yuchao Shao, Gexiao Tian, Zhongjie Wu, Zheng Cao 0003, Jiwu Shu, Jie Wu 0003, Jiesheng Wu
FAST25
2023 More Than Capacity: Performance-oriented Evolution of Pangu in Alibaba
Qiang Li 0045, Qiao Xiang, Yuxin Wang 0003, Ridi Wen, Wenhui Yao, Shuqi Zhao, Zhaosheng Zhu, Huayong Wang, Shanyang Liu, Lulu Chen, Zhiwu Wu, Haonan Qiu, Derui Liu, Gexiao Tian, Shaozong Liu, Yaohui Wu, Zicheng Luo, Yuchao Shao, Junping Wu, Zheng Cao 0003, Zhongjie Wu, Jiaji Zhu, Jiwu Shu, Jiesheng Wu
FAST24
2023 Flor: An Open High Performance RDMA Framework Over Heterogeneous RNICs
Qiang Li 0045, Yixiao Gao, Xiaoliang Wang 0001, Haonan Qiu, Yanfang Le, Derui Liu, Qiao Xiang, Bo Li 0061, Jianbo Dong, Lingbo Tang, Hongqiang Harry Liu, Shaozong Liu, Rui Miao 0001, Yaohui Wu, Zhiwu Wu, Zheng Cao 0003, Zhongjie Wu, Chen Tian 0001, Guihai Chen, Dennis Cai, Jiaji Zhu, Jiesheng Wu, Jiwu Shu
OSDI21
2022 GALAXY: A Generative Pre-trained Model for Task-Oriented Dialog with Semi-supervised Learning and Explicit Policy Injection
abstract
Pre-trained models have proved to be powerful in enhancing task-oriented dialog systems. However, current pre-training methods mainly focus on enhancing dialog understanding and generation tasks while neglecting the exploitation of dialog policy. In this paper, we propose GALAXY, a novel pre-trained dialog model that explicitly learns dialog policy from limited labeled dialogs and large-scale unlabeled dialog corpora via semi-supervised learning. Specifically, we introduce a dialog act prediction task for policy optimization during pre-training and employ a consistency regularization term to refine the learned representation with the help of unlabeled dialogs. We also implement a gating mechanism to weigh suitable unlabeled dialog samples. Empirical results show that GALAXY substantially improves the performance of task-oriented dialog systems, and achieves new state-of-the-art results on benchmark datasets: In-Car, MultiWOZ2.0 and MultiWOZ2.1, improving their end-to-end combined scores by 2.5, 5.3 and 5.5 points, respectively. We also show that GALAXY has a stronger few-shot ability than existing models under various low-resource settings. For reproducibility, we release the code and data at https://github.com/siat-nlp/GALAXY.
Wanwei He, Yinpei Dai, Yinhe Zheng, Yuchuan Wu, Zheng Cao 0003, Dermot Liu, Min Yang 0007, Fei Huang 0002, Luo Si, Jian Sun 0021, Yongbin Li 0001
AAAI5
2022 SPACE-2: Tree-Structured Semi-Supervised Contrastive Pre-training for Task-Oriented Dialog Understanding
abstract
Pre-training methods with contrastive learning objectives have shown remarkable success in dialog understanding tasks. However, current contrastive learning solely considers the self-augmented dialog samples as positive samples and treats all other dialog samples as negative ones, which enforces dissimilar representations even for dialogs that are semantically related. In this paper, we propose SPACE-2, a tree-structured pre-trained conversation model, which learns dialog representations from limited labeled dialogs and large-scale unlabeled dialog corpora via semi-supervised contrastive pre-training. Concretely, we first define a general semantic tree structure (STS) to unify the inconsistent annotation schema across different dialog datasets, so that the rich structural information stored in all labeled data can be exploited. Then we propose a novel multi-view score function to increase the relevance of all possible dialogs that share similar STSs and only push away other completely different dialogs during supervised contrastive pre-training. To fully exploit unlabeled dialogs, a basic self-supervised contrastive loss is also added to refine the learned representations. Experiments show that our method can achieve new state-of-the-art results on the DialoGLUE benchmark consisting of seven datasets and four popular dialog understanding tasks.
Wanwei He, Yinpei Dai, Binyuan Hui, Min Yang 0007, Zheng Cao 0003, Jianbo Dong, Fei Huang 0002, Luo Si, Yongbin Li 0001
COLING5
2022 CGoDial: A Large-Scale Benchmark for Chinese Goal-oriented Dialog Evaluation
abstract
Practical dialog systems need to deal with various knowledge sources, noisy user expressions, and the shortage of annotated data.To better solve the above problems, we propose CGoDial 1 , a new challenging and comprehensive Chinese benchmark for multi-domain Goal-oriented Dialog evaluation.It contains 96,763 dialog sessions, and 574,949 dialog turns totally, covering three datasets with different knowledge sources: 1) a slot-based dialog (SBD) dataset with table-formed knowledge, 2) a flow-based dialog (FBD) dataset with treeformed knowledge, and a retrieval-based dialog (RBD) dataset with candidate-formed knowledge.To bridge the gap between academic benchmarks and spoken dialog scenarios, we either collect data from real conversations or add spoken features to existing datasets via crowdsourcing.The proposed experimental settings include the combinations of training with either the entire training set or a few-shot training set, and testing with either the standard test set or a hard test subset, which can assess model capabilities in terms of general prediction, fast adaptability and reliable robustness.
Yinpei Dai, Wanwei He, Bowen Li 0002, Yuchuan Wu, Zheng Cao 0003, Zhongqi An, Jian Sun 0021, Yongbin Li 0001
EMNLP5
2022 mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections
abstract
Chenliang Li, Haiyang Xu, Junfeng Tian, Wei Wang, Ming Yan, Bin Bi, Jiabo Ye, He Chen, Guohai Xu, Zheng Cao, Ji Zhang, Songfang Huang, Fei Huang, Jingren Zhou, Luo Si. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Chenliang Li 0003, Haiyang Xu 0001, Wei Wang 0225, Ming Yan 0008, Bin Bi, Jiabo Ye, Guohai Xu, Zheng Cao 0003, Ji Zhang 0011, Songfang Huang, Fei Huang 0002, Jingren Zhou 0001, Luo Si
EMNLP10
2021 AIBench Scenario: Scenario-Distilling AI Benchmarking
abstract
Modern real-world application scenarios like Internet services consist of a diversity of AI and non-AI modules with huge code sizes and long and complicated execution paths, which raises serious benchmarking or evaluating challenges. Using AI components or micro benchmarks alone can lead to error-prone conclusions. This paper presents a methodology to attack the above challenge. We formalize a real-world application scenario as a Directed Acyclic Graph-based model and propose the rules to distill it into a permutation of essential AI and non-AI tasks, which we call a scenario benchmark. Together with seventeen industry partners, we extract nine typical scenario benchmarks. We design and implement an extensible, configurable, and flexible benchmark framework. We implement two Internet service AI scenario benchmarks based on the framework as proxies to two real-world application scenarios. We consider scenario, component, and micro benchmarks as three indispensable parts for evaluating. Our evaluation shows the advantage of our methodology against using component or micro AI benchmarks alone. The specifications, source code11Zenodo: https://doi.org/10.5281/zenodo.5158715 GitHub: https://github.com/BenchCouncil/aibench_scenario, testbed, and results are publicly available from https://www.benchcouncil.org/aibench/scenario/.
Wanling Gao, Fei Tang 0003, Jianfeng Zhan, Lei Wang 0004, Zheng Cao 0003, Chuanxin Lan, Chunjie Luo, Xiaoli Liu 0002, Zihan Jiang 0006
PACT6
2021 AIBench Training: Balanced Industry-Standard AI Training Benchmarking
abstract
Earlier-stage evaluations of a new AI architecture/system need affordable AI benchmarks. Only using a few AI component benchmarks like MLPerf alone in the other stages may lead to misleading conclusions. Moreover, the learning dynamics are not well understood, and the benchmarks' shelf-life is short. This paper proposes a balanced benchmarking methodology. We use real-world benchmarks to cover the factors space that impacts the learning dynamics to the most considerable extent. After performing an exhaustive survey on Internet service AI domains, we identify and implement nineteen representative AI tasks with state-of-the-art models. For repeatable performance ranking (RPR subset) and workload characterization (WC subset), we keep two subsets to a minimum for affordability. We contribute by far the most comprehensive AI training benchmark suite. The evaluations show: (1) AIBench Training (v1.1) outperforms MLPerf Training (v0.7) in terms of diversity and representativeness of model complexity, computational cost, convergent rate, computation, and memory access patterns, and hotspot functions; (2) Against the AIBench full benchmarks, its RPR subset shortens the benchmarking cost by 64%, while maintaining the primary workload characteristics; (3) The performance ranking shows the single-purpose AI accelerator like TPU with the optimized TensorFlow framework performs better than that of GPUs while losing the latter's general support for various AI models. The specification, source code, and performance numbers are available from the AIBench homepage https://www.benchcouncil.org/aibench-training/index.html.
Fei Tang 0003, Wanling Gao, Jianfeng Zhan, Chuanxin Lan, Lei Wang 0004, Chunjie Luo, Zheng Cao 0003, Xingwang Xiong, Zihan Jiang 0006, Tianshu Hao, Fanda Fan, Fan Zhang 0047, Yunyou Huang, Jianan Chen 0003, Mengjia Du, Chen Zheng 0001, Daoyi Zheng, Haoning Tang, Kunlin Zhan, Defei Kong, Chongkang Tan, Xinhui Tian, Yatao Li, Junchao Shao, Xiaoyu Wang 0002, Jiahui Dai, Hainan Ye
ISPASS8
2021 Reducing BERT Computation by Padding Removal and Curriculum Learning
abstract
BERT is very computationally expensive, which is a hurdle for its training and deployment. This work focuses on removing the unnecessary computation due to input padding in BERT. The input of BERT consists of two concatenated sentences. If the length of the two concatenated sentences is shorter than the maximum sequence length, padding must be added to the end of the sentences to fill the empty slots in the input. Because the lengths of sentences vary greatly, there can be a large amount of padding in input. For the English Wikipedia & BooksCorpus dataset, the percentage of padding among all the input tokens is 17% and 48%, respectively, when the max sequence length is set to 128 and 512. For the Chinese Wikipedia dataset, this percentage is 35% and 79%, respectively, when the max sequence length is 128 and 512. For SQuAD-v1.1 [2], padding accounts for 54% of the total input tokens when the max sequence length is 384. Thus, there is a lot of wasted computation on padding, which significantly increases the training and inference time.
Wei Zhang 0127, Wei Wei 0021, Wen Wang 0001, Lingling Jin, Zheng Cao 0003
ISPASS5
2021 When Cloud Storage Meets RDMA
Yixiao Gao, Qiang Li 0045, Lingbo Tang, Yongqing Xi, Wenwen Peng, Bo Li 0061, Yaohui Wu, Shaozong Liu, Xingkui Liu, Zhongjie Wu, Junping Wu, Zheng Cao 0003, Chen Tian 0001, Jiaji Zhu, Haiyong Wang, Dennis Cai, Jiesheng Wu
NSDI18
2021 A New Optoelectronic Hybrid Network Based on Scheduling Optimization of Optical Links
abstract
The emergence of exascale computers will represent a milestone in high-performance computing (HPC). Optoelectronic interconnections and configurable switches will change the traditional supercomputer architecture. However, new hardware is not easily adapted to dynamic running conditions. Based on scheduling optimization of optical links, we propose a new optoelectronic hybrid network, the software-defined network accelerator (sDNA), for an exascale computer. Our scheduling optimization contains an optical interconnection method and an adaptive routing method. The main contribution of our work is an extended edge forwarding index (E-EFI) optical interconnection method based on slow-switching optical devices. The optical link connections are established by evaluating the traffic offloading revenue for each optical link candidate. To support optical interconnection, sDNA selects a suitable routing strategy according to the job-schedule information and prior HPC application knowledge. We tested sDNA in a network simulator and a prototype exascale computer system using both the US Department of Energy (DOE) application and real-world communication benchmarks. The verification results for traffic offloading reveal that our optical interconnection method not only offloads traffic from electrical links to optical links but also avoids the congestion inherent to electrical links. sDNA maintains a throughput of more than 80 percent bandwidth and reduces the communication delay by 10 percent in our real prototype system and simulator. Thus, sDNA is an ideal candidate for accelerating the communication performance of exascale computers.
En Shao, Guangming Tan, Zhan Wang 0003, Guojun Yuan, Zheng Cao 0003, Ninghui Sun
IEEE Trans. Computers5
2020 DWT: Decoupled Workload Tracing for Data Centers
abstract
Workload tracing is the foundational technology that many applications hinge upon. However, recent paradigm shift to-ward cloud computing has caused tremendous challenges to traditional workload tracing. Existing solutions either require a dedicated offline cluster or fail to capture the full-spectrum workload characteristics. This paper proposes DWT, a novel framework that leverages fast online instruction tracing, and uses synthetic data offline for memory access pattern reconstruction, thereby capturing the full workload characteristics while obviating the need of dedicated clusters. Experiment results show that the stack distance profiles generated from synthetic address traces match well with the original ones across all SPEC CPU 2017 programs and representative cloud applications, with correlation coefficient R^2 no less than 0.9. The page-level access frequencies also match well with those of the original programs. This decoupled tracing approach not only removes the roadblocks on workload characterization for data centers, but also enables new applications such as efficient online resource management.
Xiaowei Jiang, Zheng Cao 0003
HPCA5
2020 EFLOPS: Algorithm and System Co-Design for a High Performance Distributed Training Platform
abstract
Deep neural networks (DNNs) have gained tremendous attractions as compelling solutions for applications such as image classification, object detection, speech recognition, and so forth. Its great success comes with excessive trainings to make sure the model accuracy is good enough for those applications. Nowadays, it becomes challenging to train a DNN model because of 1) the model size and data size keep increasing, which usually needs more iterations to train; 2) DNN algorithms evolve rapidly, which requires the training phase to be short for a quick deployment. To address those challenges, distributed training platforms have been proposed to leverage massive server nodes for training, with the hope of significant training time reduction. Therefore, scalability is a critical performance metric to evaluate a distributed training platform. Nevertheless, our analysis reveals that traditional server clusters have poor scalability for training due to the traffic congestions within the server and beyond. The intra-server traffic on the I/O fabric can result in severe congestions and skewed quality of service as high performance devices are competing with each other. Moreover, the traffic congestions on the Ethernet for inter-server communication could also incur significant performance degradation. In this work, we devise a novel distributed training platform, EFLOPS, that adopts an algorithm and system co-design methodology to achieve good scalability. A new server architecture is proposed to alleviate the intra-server congestions. Moreover, a new network topology, BiGraph, is proposed to divide the network into two separate parts, so that there is always a direct connection between any nodes from different parts. Finally, accompany with BiGraph, a topology-aware allreduce algorithm is proposed to eliminate the traffic congestion on the direct connection. The experimental results show that eliminating the congestions on network interface can gain up to 11.3xcommunication speedup. The proposed algorithm and topology can provide further improvement up to 6.08x. The overall performance of ResNet-50 training achieves near-linear scalability, and is competitive to the top-rankings of MLPerf results.
Jianbo Dong, Zheng Cao 0003, Jianxi Ye, Shaochuang Wang, Liuyihan Song, Liwei Peng, Yiqun Guo, Xiaowei Jiang, Lingbo Tang, Yin Du, Yingya Zhang
HPCA2
2020 Dissecting the Communication Latency in Distributed Deep Sparse Learning
abstract
Distributed deep learning (DDL) uses a cluster of servers to train models in parallel. This has been applied to a multiplicity of problems, e.g. online advertisement, friend recommendations. However, the distribution of training means that the communication network becomes a key component in system performance. In this paper, we measure the Alibaba's DDL system, with a focus on understanding the bottlenecks introduced by the network. Our key finding is that the communications overhead has a surprisingly large impact on performance. To explore this, we analyse latency logs of 1.38M Remote Procedure Calls between servers during model training for two real applications of high-dimensional sparse data. We reveal the major contributors of the latency, including concurrent write/read operations of different connections and network connection management. We further observe a skewed distribution of update frequency for individual parameters, motivating us to propose using in-network computation capacity to offload server tasks.
Zhenyu Li 0001, Jianbo Dong, Zheng Cao 0003, Tao Lan, Gareth Tyson, Gaogang Xie
Internet Measurement Conference4
2020 CETUS: Towards Proportional Capacity Provisioning and Cost-Effectiveness in Frontend Servers
abstract
Hyper-scale data centers emerged in the last decade largely adopt a multi-tiered architecture with the frontend clusters dedicated to serve high-demand user facing web traffic. In order to mitigate the overhead caused by Transport Layer Security (TLS) that protects the communications between users and the data center, the frontend clusters usually apply hardware TLS acceleration. In this paper, we analyze the inefficiencies that lie in today's frontend clusters, and propose CETUS, an improved data center frontend system architecture. CETUS improves the cost effectiveness of frontend clusters through cluster consolidation that fully offloads the TLS and network stack to a CETUS SoC; it enables proportional capacity provisioning through pooling of resources and dynamic division of frontend tasks with a flow diverter. Compared to existing frontend clusters that are equipped with commercial TLS acceleration solutions, CETUS balances out the frontend resource utilization, and provides up to 86.2% in cost reduction while maintaining at the same level of throughput and latency of the frontend.
Xiaowei Jiang, Zheng Cao 0003
ISPASS6
2019 HPCC: high precision congestion control
abstract
Congestion control (CC) is the key to achieving ultra-low latency, high bandwidth and network stability in high-speed networks. From years of experience operating large-scale and high-speed RDMA networks, we find the existing high-speed CC schemes have inherent limitations for reaching these goals. In this paper, we present HPCC (High Precision Congestion Control), a new high-speed CC mechanism which achieves the three goals simultaneously. HPCC leverages in-network telemetry (INT) to obtain precise link load information and controls traffic precisely. By addressing challenges such as delayed INT information during congestion and overreac-tion to INT information, HPCC can quickly converge to utilize free bandwidth while avoiding congestion, and can maintain near-zero in-network queues for ultra-low latency. HPCC is also fair and easy to deploy in hardware. We implement HPCC with commodity programmable NICs and switches. In our evaluation, compared to DCQCN and TIMELY, HPCC shortens flow completion times by up to 95%, causing little congestion even under large-scale incasts.
Rui Miao 0001, Hongqiang Harry Liu, Lingbo Tang, Zheng Cao 0003, Ming Zhang 0005, Frank Kelly, Mohammad Alizadeh, Minlan Yu
SIGCOMM7
2017 Regional Congestion Control in Datacenter Networks
abstract
The rapid deployment of cloud computing and online services poses great challenges for data center networks, and congestion control is one of the top concerns. Although numbers of proposals in different network layers have been put forward to alleviate the negative impact of congestion, the short-lived flows, which are latency-sensitive and constitute the majority of total traffic in data centers, still suffer severe performance degradation. Since the existing congestion control methods all rely on end hosts to perceive congestion and then adjust their network sending rate, the response time is relatively long when compared with the duration of short-lived flows, which increases latency significantly. In this paper, we propose RCC, a regional congestion control mechanism, which aims to respond to congestion more quickly and eliminate the mismatch mentioned above. Different from host-based mechanisms, RCC is implemented in the switch, which detects the congestion state and schedule the traffic around the congestion point locally, without sending feedback to the distal host. Evaluation has shown that, compared with the host-based mechanism, our method achieves better performance for short-lived flows and maintains stable buffer occupancy of the switch. In addition, mixed long- and short-lived flows which contend for the same bottleneck link can share the bandwidth more fairly.
Fan Yang 0096, Zhan Wang 0003, Xiaoli Liu 0002, Zheng Cao 0003, Guojun Yuan, Xuejun An
ICPADS4
2017 Regional Congestion Mitigation in Lossless Datacenter Networks
Xiaoli Liu 0002, Fan Yang 0096, Yanan Jin, Zhan Wang 0003, Zheng Cao 0003, Ninghui Sun
NPC5
2016 Modeling Traffic of Big Data Platform for Large Scale Datacenter Networks
abstract
Prior to deployment, network designers often use simulators to pre-evaluate the performance of designed network with artificial network traffic. The traditional way of separating network design from real applications will not only result in over-designed network configurations, wasting money and energy, but also miss the real network demands of applications, degrading system performance. In this paper, we provide a method to model the network traffic of current popular big data platforms, which can observably improve the matching between network design and applications. The new method extracts communication behavior from the popular big data applications and replays the behavior instead of the packet traces. Experiments show that the traffic generated by the model is almost match the real traffic and the model can easily scale to thousands of nodes.
Zheng Cao 0003, Zhan Wang 0003, Dawei Zang, En Shao, Ninghui Sun
ICPADS2
2015 PROP: Using PCIe-Based RDMA to Accelerate Rack-Scale Communications in Data Centers
abstract
In order to reduce the demands on bandwidth of core layer network, data center operators usually assign tasks of the same job to servers that are located in the same rack, leading to the fact that 80% of the traffic originated from servers retains in the same rack. As a result, providing sufficient network capacity inside racks becomes critical to the Quality-of-Service of current data center applications. In this paper, we propose PROP, a novel hybrid network architecture which leverages PCIe-based RDMA to reinforce rack-scale connectivity in data centers. In our design, intra-rack bulk data transfers will be accelerated by a dedicated high-bandwidth PCIe-compliant network while complemented with the existing Ethernet network. In addition, we develop a proprietary PCIe-based RDMA hardware which can allow the servers in the same rack to exchange data in main memory without involving the operating system and the processors. We also implement a software stack to enable existing socket-based applications to transparently utilize the proposed dedicated network system. As the preliminary stage, this paper focuses on exploiting the unique design point and implements an FPGA-based prototype to validate the technical feasibility of the proposed architecture.
Dawei Zang, Zheng Cao 0003, Xiaoli Liu 0002, Lin Wang 0015, Zhan Wang 0003, Ninghui Sun
ICPADS2
2014 Building a large-scale direct network with low-radix routers
abstract
Communication locality is an important characteristic of parallel applications. A great deal of research shows that utilizing the characteristic will favor most applications. Aiming at communication locality, we present a hierarchical direct network topology to accelerate neighbor communication. Combining mesh topology and complete graph topology, it can be used to optimize local communication and build large-scale network with low radix routers. Analyzing the characteristic of hierarchical topology, we find the presented topology has high cost performance and excellent expandability. We also design two minimum path routing algorithms and compare them with Mesh, Dragonfly and PERCS topologies. The results show the saturated throughput of hierarchical topology is nearly 40% with uniform random trace and 70% with local communication model of 4K nodes. That indicates high scalability for applications with local communication and cost efficiency for uniform random trace.
Zheng Cao 0003, Zhiguo Fan, Zhan Wang 0003, Xiaoli Liu 0002, Li Qiang, Xuejun An, Ninghui Sun
ICPADS2
2014 HiNetSim: A Parallel Simulator for Large-Scale Hierarchical Direct Networks
Zhiguo Fan, Zheng Cao 0003, Xiaoli Liu 0002, Zhan Wang 0003, Dawei Zang, Xuejun An
NPC2
2014 An Intra-Server Interconnect Fabric for Heterogeneous Computing
Zheng Cao 0003, Xiaoli Liu 0002, Qiang Li 0045, Zhan Wang 0003, Xuejun An
J. Comput. Sci. Technol.1
2013 Accelerating Allreduce Operation: A Switch-Based Solution
abstract
Collective operations, such as all reduce, are widely treated as the critical limiting factors in achieving high performance in massively parallel applications. Conventional host-based implementations, which introduce a large amount of point-to-point communications, are less efficient in large-scale systems. To address this issue, we propose a design of switch chip to accelerate collective operations, especially the allreduce operation. The major advantage of the proposed solution is the high scalability since expensive point-to-point communications are avoided. Two kinds of allreduce operations, namely block-allreduce and burst-allreduce, are implemented for short and long messages, respectively. We evaluated the proposed design with both a cycle-accurate simulator and a FPGA prototype system. The experimental results prove that switch-based allreduce implementation is quite efficient and scalable, especially in large-scale systems. In the prototype, our switch-based implementation significantly outperforms the host-based one, with a 16 times improvement in MPI time on 16 nodes. Furthermore, the simulation shows that, upon scaling from 2 to 4096 nodes, the switch-based allreduce latency only increases slightly by less than 2 us.
Nongda Hu, Zheng Cao 0003, Xuejun An, Ninghui Sun
ICCCN3
2013 cHPP controller: A High Performance Hyper-node Hardware Accelerator
abstract
The high-density blade server provides an attractive solution for the rapid increasing demand on computing. The degree of parallelism inside a blade enclosure nowadays has reach up to hundreds of cores. In such parallelism, it is necessary to accelerate communications inside a blade enclosure. However, commercial products seldom set foot in the optimization based on hardware. A hyper-node controller is proposed to provide a low overhead and high performance interconnection based on PCIe, which supports global address space, user-level communication, and efficient communication primitives. Furthermore, the efficient sharing of I/O resource is another goal of this design. The prototype of the hyper-node controller is implemented in FPGA. The testing results show the lowest latency is only 1.242us and the highest bandwidth is 3.19GB/s, which is almost 99.7% of the theoretic peak bandwidth.
Zheng Cao 0003, Zhan Wang 0003, Xiaoli Liu 0002, Xuejun An, Ninghui Sun
PDCAT3
2012 Design of Hardware-Based Communication Performance Measurement Tool
abstract
With the popularity and development of heterogeneous computing, proper communication performance measurement tools are needed to explore new communication patterns under heterogeneous computing systems and optimize program's performance. This paper proposes a hardware-based communication performance measurement tool, named as HCPM, which brings little influence on original program, and can collect communication traces generated by heterogeneous processors which implement PCIe or HT as their system bus. HCPM firstly provides basic communication primitives to set up a communication system. Then based on these primitives, it collects communication trace. Real-time collected traces are transmitted to a dedicated computer for further analysis. Evaluation shows that with the use of proper compression in hardware, HCPM can transmit at least five processors' communication traces with a single Gigabit Ethernet link.
Zhan Wang 0003, Zheng Cao 0003, Xiaoli Liu 0002, Xuejun An
CLUSTER2
2011 Design of HPC Node with Heterogeneous Processors
abstract
Heterogeneous Computing is becoming an important technology trend in HPC, where more and more heterogeneous processors are used. However, in traditional node architecture, heterogeneous processors are always used as coprocessors. Such usage increases the communication latency between heterogeneous processors and prevents the node from achieving high density. With the purpose of improving communication efficiency between heterogeneous processors, this paper proposed a new node architecture named HeteNode. In HeteNode, general purpose processors and heterogeneous processors are interconnected by a system controller directly and play the same role in both process of communication and process of computation. The prototype of HeteNode which contains nine processors in 1U chassis is built. Evaluation carried out on the prototype shows that 580ns minimum intra-node latency and 1.78us minimum inter-node latency between heterogeneous processors are achieved. Besides, NPB benchmarks show good scalability in HeteNode.
Zheng Cao 0003, Hongwei Tang, Qiang Li 0045, Bo Li 0009, Xuejun An, Ninghui Sun
CLUSTER1
2010 Adding an Expressway to Accelerate the Neighborhood Communication
abstract
The blade system is very popular in high performance computing. In a blade system, the blade is a fundamental element in which are symmetric multi-processors (SMP). About ten blades constitute a blade box, several blade boxes constitute a cabinet and some cabinets constitute a blade system at last. The blades in a blade box are neighbors because they have relatively short distance. Programmers always try to place the tightly related processes into the same blade box. However, there's seldom any optimization made by hardware to accelerate the communication in a blade box. Thus, a single chip design called hyper-node controller is presented to provide ultra low latency and high bandwidth which resembles an expressway between neighbors. All the nodes in a blade box can act as a single hyper node by using the hyper-node controller. It is apparent that the additional controller is a useful supplement to efficiently enhance the communication in a blade box and finally enhance the entire blade system. A FPGA prototype of the hyper-node controller has been implemented and it can connect five blades simultaneously. In the preliminary performance evaluation, the latency for an 8-byte payload between two blades is less than 1us, 1.33GB/s which is nearly 94% of the peak effective bandwidth can be obtained by transferring messages with a payload of only 256 bytes.
Zheng Cao 0003, Xuejun An, Ninghui Sun
HPCC3
2010 HPP Controller: A System Controller Dedicated for Message Passing
abstract
The traditional system controller in symmetric multi-processors (SMP) controls the memory, so it is suitable for the shared memory programming model. With the emergence of the processors which integrate memory controllers, the system controller seems less important than before. However, since the system controller resides in the center of a computer system, it acts as an artery which directly connects to the processors and the high-speed IO devices. Thus making full use of its position advantage can no doubt gain performance enhancement. By now, the message passing programming model has dominated the high performance computing (HPC) field, however the system controller makes little contribution to it. Thus, a system controller called HPP controller which is dedicated for the message passing programming model is presented in this paper. The HPP controller is connected to several processors simultaneously, and the communication between these processors uses the message passing programming model. The HPP controller has powerful DMA engines embedded which can provide flexible and sufficient message passing capability. Two key techniques: supporting arbitrary byte alignment and virtualizing the DMA engine are introduced in detail. The preliminary result of the FPGA prototype shows that the HPP controller has ultra low hardware latency and relatively high bandwidth. Besides, the NPB result shows that it can provide high efficiency for the message passing programming model.
Zheng Cao 0003, Xuejun An, Ninghui Sun
PDCAT3
2010 HPP controller: a system controller for high performance computing
Zheng Cao 0003, Xuejun An, Ninghui Sun
Frontiers Comput. Sci. China2
2009 SimK: A Large-Scale Parallel Simulation Engine
Mingyu Chen 0001, Gui Zheng, Zheng Cao 0003, Huiwei Lv, Ninghui Sun
J. Comput. Sci. Technol.4
2005 A Reconfigurable Optical Interconnect System for DSAG
abstract
High performance computing research is facing challenges and innovation on architecture is urgent. DSAG architecture is proposed and delivers "Architecture on Demand" feature. In DSAG, components in different catalogs are parted, while the ones in same catalog are congregated. This architecture can be enabled by optical interconnect and reconfigurable computing technology. Using advanced optical devices and enhanced reconfigurable computing devices (FPGA), we build a prototype system for DSAG. Optical interconnect can reach 16Gbps bandwidth; DDRAM interface is selected as host communication interface to match the bandwidth of optical channel; Reconfigurable logic and embedded processors are employed for flexible reconfiguration. The system is featured by high bandwidth, owerful, flexible.
Lei Li 0005, Zheng Cao 0003, Mingyu Chen 0001, Jianping Fan 0002
PDCAT2