Keqiang He

dblp:95/1050 · DBLP profile ↗
← Back
26ranked-venue papers
6as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 19 · 5 first-author · 8 since 2021Systems, architecture and hardware · 6 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2026 OSCAR: O(1)-Step Convergence and Readily-deployable Congestion Control
Zhaochen Zhang, Feiyang Xue, Rui Ning, Keqiang He, Gianni Antichi, Zhimeng Yin 0001, Rui Li 0020, Zhengqi Cui, Zhehao Lin, Peirui Cao, Guihai Chen, Chen Tian 0001
NSDI4
2026 Virtual Slicing: Achieving Control Plane Availability and Traffic Engineering Efficiency in Data Centers
abstract
Many proposals have demonstrated the efficiency advantages of software-defined networking (SDN) in managing data center networks. Common practices employ centralized traffic engineering (TE) in the SDN control plane to optimize load balancing and throughput. Meanwhile, for high availability purposes, the control plane is partitioned to ensure the impact of a single faulty controller is contained. However, the interaction between these two aspects is often overlooked. In particular, we show that the current control plane partitioning approach leads to imbalanced link loads and degraded application performance. To address this issue, we proposevirtual slicing, a new control plane partitioning scheme. Virtual slicing achieves desirable traffic engineering performance while retaining the availability guarantees from the current approach. Virtual slicing is implemented and evaluated with real-world and synthetic traffic traces on production spine-free data center networks. Results show that virtual slicing reduces tail link utilizations by up to 28.4%, and improves flow completion times by up to 36%.
Brian Chang, Keqiang He, Shawn Shuoshuo Chen, Mingyang Zhang 0005, Wenfei Wu, Fan Wu 0006, Chen Tian 0001, Aditya Akella
IEEE Trans. Netw.2
2026 UDMP: Unified Delay-Driven Multipath Protocol for AI Clusters
abstract
Distributed AI model training generates bursty, low-entropy elephant flows that challenge existing single-path transport protocols in multi-stage Clos networks, leading to congestion and inefficiency. Multipath transport emerges as a promising solution, leveraging multiple paths to balance traffic and enhance resilience. However, current multipath RDMA solutions suffer from scalability, congestion control, and load-balancing inefficiencies. This paper introduces Unified Delay-driven Multipath Protocol (UDMP), a novel approach that co-designs congestion control and load balancing using network delay as a unified signal. UDMP employs delay-gradient-based congestion control to precisely resolve unavoidable congestion. Moreover, UDMP leverages delay-assisted load balancing to shift traffic across paths with minimal latency adaptively, maintaining throughput when encountering avoidable congestion. A novel Token Pool design integrates these components, eliminating per-path state overhead while achieving fine-grained traffic distribution. Implementations on DPDK and NS3 demonstrate that UDMP achieves up to 2x higher throughput and reduces flow completion times by up to 30% compared to state-of-the-art methods like MPRDMA and QP-Scaling. These results highlight UDMP’s effectiveness in meeting the stringent performance requirements of modern distributed AI training workloads.
Chengyuan Huang, Zhengqi Cui, Jun Xu 0037, Zhaochen Zhang, Li Wang 0110, Peirui Cao, Zhongming Ji, Jilei Chen, Shengju Zhang, Lingkun Meng, Ahmed M. Abdelmoniem, Fu Xiao 0001, Wan-Chun Dou, Guihai Chen, Keqiang He, Chen Tian 0001
IEEE Trans. Netw.16
2025 Marlin: Enabling High-Throughput Congestion Control Testing in Large-Scale Networks
abstract
Cloud providers require high-throughput traffic to test the effectiveness of congestion control (CC) configurations (i.e., CC algorithm selection and their parameter settings) in networks. A network tester capable of evaluating CC configurations needs to fulfill the following requirements: (R1) Capable of generating traffic with CC behaviors. (R2) Ability to customize CC algorithms. (R3) High throughput CC traffic generation. However, existing network testers fail to meet these requirements simultaneously. The paper presents Marlin, a novel high-throughput network tester designed for CC evaluation. Marlin leverages a high-throughput, low-programmability device to amplify the traffic generated by a low-throughput, high-programmability device. The low-throughput device is responsible for complex computational tasks, such as running CC and flow scheduling algorithms, and communicates with the high-throughput device at a high frequency using small packets to instruct it to generate high-throughput traffic with CC behaviors. This hybrid approach allows for customizable, high-throughput CC testing. Our experiments demonstrate that Marlin can accurately emulate CC behaviors and replicate real-world scenarios. Marlin can generate 1.2 Tbps of CC traffic using a single programmable switch pipeline and one 100 Gbps port of an FPGA NIC, supporting up to 65,536 concurrent flows.
Li Wang 0110, Jingzhi Wang, Songyue Liu, Keqiang He, Jian Wang 0025, Xiaoliang Wang 0001, Wan-Chun Dou, Guihai Chen, Chen Tian 0001
EuroSys5
2025 Enabling Virtual Priority in Data Center Congestion Control
abstract
In data center networks, various types of traffic with strict performance requirements operate simultaneously, necessitating effective isolation and scheduling through priority queues. However, most switches support only around ten priority queues. Virtual priority can address this limitation by emulating multi-priority queues on a single physical queue, but existing solutions often require complex switch-level scheduling and hardware changes. Our key insight is that virtual priority can be achieved by carefully managing bandwidth contention in a physical queue, which is traditionally handled by congestion control (CC) algorithms. Hence, the virtual priority mechanism needs to be tightly coupled with CC. In this paper, we propose PrioPlus, a CC enhancement algorithm that can be integrated with existing congestion control schemes to enable virtual priority transmission. PrioPlus assigns specific delay ranges to different priority levels, ensuring that flows transmit only when the delay is within the assigned range, effectively meeting virtual priority requirements. Compared to Swift CC with physical priority queues, PrioPlus provides strict priority for high-priority flows without impacting performance sensibly. Meanwhile, it benefits low-priority flows from 25% to 41% as its priority-aware design enhances CC's ability to fully utilize available bandwidth once higher-priority traffic completes. As a result, in coflow and model training scenarios, PrioPlus improves job completion times by 21% and 33%, respectively, compared to Swift with physical priority queues.
Zhaochen Zhang, Feiyang Xue, Keqiang He, Zhimeng Yin 0001, Gianni Antichi, Yizhi Wang 0004, Rui Ning, Haixin Nan, Xu Zhang 0006, Peirui Cao, Xiaoliang Wang 0001, Wan-Chun Dou, Guihai Chen, Chen Tian 0001
EuroSys3
2025 FlowGram: Resource Allocation for Network Flow Measurement in Clouds
abstract
Network measurement is essential for cloud tenants in their virtual network management. However, existing measurement systems exhibit suboptimal resource efficiency while accommodating an increasing number of tenants with limited switch resources. To address this issue, this paper proposes FlowGram, a measurement framework to provide flow frequency estimation services for cloud tenants. FlowGram provides a user-friendly interface for tenants and automatically adjusts the memory allocation to meet the tenants' demands. In the data plane, FlowGram employs sketches on programmable switches to perform the measurement. In the control plane, FlowGram utilizes a precise error model of sketches and an efficient memory allocation algorithm that ensures the measurement errors are within the specified error bounds. We prototype FlowGram on programmable switches and commodity servers and conduct comprehensive experiments. Experiment results demonstrate that FlowGram outperforms the state-of-the-art system, achieving a 14 % higher task satisfaction rate while utilizing the same amount of resources.
Chenqi Zhao, Wenfei Wu, Qun Huang 0001, Keqiang He
IWQoS4
2025 SGLB: Scalable and Robust Global Load Balancing in Commodity AI Clusters
abstract
Internet companies are constructing large-scale AI clusters with commodity Ethernet switches for AI model training to support their businesses. AI training workloads impose stringent network requirements, mandating that cluster networks deliver high peak throughput while maintaining robustness and resilience in the face of link failures. We present SGLB, a distributed, global congestion-aware load balancing system for AI clusters. SGLB operates a control-plane protocol, SyncMesh, to enable a new load balancing abstraction in modern commodity switches—Global Load Balancing (GLB) engine—which utilizes global congestion information to distribute traffic across all available paths. We address three key challenges in designing SGLB: fast routing convergence to minimize downtime in the event of link failures, scalable maintenance of congestion profiles within the constraints of limited switch hardware resources, and preventing GLB throughput suppression in scenarios where path bandwidths are asymmetric. We prototype SGLB and conduct extensive experiments to evaluate SGLB. SGLB ensures rapid routing convergence in the event of link failures, recovering in as little as 45 μs to guarantee network robustness for long-term, stable model training. Additionally, SGLB effectively load-balances traffic across paths, avoiding those with global congestion, which accelerates All-to-All collective communication by up to 60%.
Chenchen Qi, Wenfei Wu, Yongcan Wang, Keqiang He, Yu-Hsiang Kao, Zongying He, Chen-Yu Yen, Zhuo Jiang, Feng Luo 0006, Surendra Anubolu, Yanjin Gao, Bingfeng Lin, Wenda Ni, Donglin Wei, Shan Ding
SIGCOMM4
2025 Reunion: Receiver-driven network load balancing mechanism in AI training clusters
Mingyao Wang, Keqiang He, Peirui Cao, Jiong Duan, Dongliang Lv, Chengyuan Huang, Wan-Chun Dou, Guihai Chen, Chen Tian 0001
Comput. Networks2
2024 Balancing Sdn Control Plane Availability and Traffic Engineering Efficiency in Data Centers
abstract
Many proposals have demonstrated the efficiency advantages of software-defined networking (SDN) in managing data center networks. Common practices employ centralized traffic engineering (TE) in the SDN control plane to optimize load balancing and throughput. Meanwhile, for high availability purposes, the control plane is partitioned to ensure the impact of a single faulty controller is contained. However, the interaction between these two aspects is often overlooked. In particular, we show that the current control plane partitioning approach leads to imbalanced link loads and degraded application performance. To address this issue, we propose virtual slicing, a new control plane partitioning scheme. Virtual slicing achieves desirable traffic engineering performance while retaining the availability guarantees from the current approach. Virtual slicing is implemented and evaluated with real-world and synthetic traffic traces on production spine-free data center networks. Results show that virtual slicing reduces tail link utilizations by up to 28.4 %, and improves flow completion times by up to 36 %.
Brian Chang, Keqiang He, Shawn Shuoshuo Chen, Mingyang Zhang 0005, Wenfei Wu, Aditya Akella
ICNP2
2024 Precise Data Center Traffic Engineering with Constrained Hardware Resources
Shawn Shuoshuo Chen, Keqiang He, Rui Wang 0025, Srinivasan Seshan, Peter Steenkiste
NSDI2
2022 Hashing Design in Modern Networks: Challenges and Mitigation Techniques
Yunhong Xu, Keqiang He, Rui Wang 0025, Minlan Yu, Nick G. Duffield, Hassan M. G. Wassel, Shidong Zhang, Leonid B. Poutievski, Junlan Zhou, Amin Vahdat
USENIX ATC2
2017 Low Latency Software Rate Limiters for Cloud Networks
abstract
A lot of recent work has focused on reducing in network queueing latency in datacenter networks. In this paper, we focus on a less explored topic --- latency increases caused by queueing in rate limiters on the end-host. First, we show that latency can be increased by an order of magnitude by rate limiters in cloud networks. To solve this problem, we extend ECN marking into rate limiters and use a datacenter congestion control algorithm --- DCTCP. Unfortunately, while this reduces latency, it also leads to throughput oscillation. Thus, this solution is not sufficient. In this paper, we also analyze the specific reasons that ECN marking in software rate limiters leads to the throughput oscillation problem. Finally, we propose two potential solutions to design software rate limiters that can achieve stable high throughput and low latency.
Keqiang He, Weite Qin, Wenfei Wu, Tian Pan 0001, Chengchen Hu, Jiao Zhang 0002, Brent E. Stephens, Aditya Akella, Ying Zhang 0022
APNet1
2016 AC/DC TCP: Virtual Congestion Control Enforcement for Datacenter Networks
abstract
Multi-tenant datacenters are successful because tenants can seamlessly port their applications and services to the cloud. Virtual Machine (VM) technology plays an integral role in this success by enabling a diverse set of software to be run on a unified underlying framework. This flexibility, however, comes at the cost of dealing with out-dated, inefficient, or misconfigured TCP stacks implemented in the VMs. This paper investigates if administrators can take control of a VM's TCP congestion control algorithm without making changes to the VM or network hardware. We propose AC/DC TCP, a scheme that exerts fine-grained control over arbitrary tenant TCP stacks by enforcing per-flow congestion control in the virtual switch (vSwitch). Our scheme is light-weight, flexible, scalable and can police non-conforming flows. In our evaluation the computational overhead of AC/DC TCP is less than one percentage point and we show implementing an administrator-defined congestion control algorithm in the vSwitch (i.e., DCTCP) closely tracks its native performance, regardless of the VM's TCP stack.
Keqiang He, Eric Rozner, Kanak Agarwal 0001, Yu Gu 0001, Wes Felter, John B. Carter, Aditya Akella
SIGCOMM1
2015 PerfSight: Performance Diagnosis for Software Dataplanes
abstract
The advent of network functions virtualization (NFV) means that data planes are no longer simply composed of routers and switches. Instead they are very complex and involve a variety of sophisticated packet processing elements that reside on the OSes and software running on compute servers where network functions (NFs) are hosted. In this paper, we argue that these new "software data planes" are susceptible to at least three new classes of performance problems. To diagnose such problems, we design, implement and evaluate, PerfSight, a ground-up system that works by extracting comprehensive low-level information regarding packet processing and I/O performance of the various elements in the software data plane. Name then analyzes the information gathered in various dimensions (e.g., across all VMs on a machine, or all VMs deployed by a tenant). By looking across aggregates, we show that it becomes possible to detect and diagnose key performance problems. Experimental results show that our framework can result in accurate detection of the root causes of key performance problems in software data planes, and it imposes very little overhead.
Wenfei Wu, Keqiang He, Aditya Akella
Internet Measurement Conference2
2015 Presto: Edge-based Load Balancing for Fast Datacenter Networks
abstract
Datacenter networks deal with a variety of workloads, ranging from latency-sensitive small flows to bandwidth-hungry large flows. Load balancing schemes based on flow hashing, e.g., ECMP, cause congestion when hash collisions occur and can perform poorly in asymmetric topologies. Recent proposals to load balance the network require centralized traffic engineering, multipath-aware transport, or expensive specialized hardware. We propose a mechanism that avoids these limitations by (i) pushing load-balancing functionality into the soft network edge (e.g., virtual switches) such that no changes are required in the transport layer, customer VMs, or networking hardware, and (ii) load balancing on fine-grained, near-uniform units of data (flowcells) that fit within end-host segment offload optimizations used to support fast networking speeds. We design and implement such a soft-edge load balancing scheme, called Presto, and evaluate it on a 10 Gbps physical testbed. We demonstrate the computational impact of packet reordering on receivers and propose a mechanism to handle reordering in the TCP receive offload functionality. Presto's performance closely tracks that of a single, non-blocking switch over many workloads and is adaptive to failures and topology asymmetry.
Keqiang He, Eric Rozner, Kanak Agarwal 0001, Wes Felter, John B. Carter, Aditya Akella
SIGCOMM1
2015 Latency in Software Defined Networks: Measurements and Mitigation Techniques
abstract
We conduct a comprehensive measurement study of switch control plane latencies using four types of production SDN switches. Our measurements show that control actions, such as rule installation, have surprisingly high latency, due to both software implementation inefficiencies and fundamental traits of switch hardware. We also propose three measurement-driven latency mitigation techniques---optimizing route selection, spreading rules across switches, and reordering rule installations---to effectively tame the flow setup latencies in SDN.
Keqiang He, Junaid Khalid, Aaron Gember, Chaithan Prakash, Aditya Akella, Li Erran Li, Marina Thottan
SIGMETRICS1
2013 Next stop, the cloud: understanding modern web service deployment in EC2 and azure
abstract
An increasingly large fraction of Internet services are hosted on a cloud computing system such as Amazon EC2 or Windows Azure. But to date, no in-depth studies about cloud usage by Internet services has been performed. We provide a detailed measurement study to shed light on how modern web service deployments use the cloud and to identify ways in which cloud-using services might improve these deployments. Our results show that: 4% of the Alexa top million use EC2/Azure; there exist several common deployment patterns for cloud-using web service front ends; and services can significantly improve their wide-area performance and failure tolerance by making better use of existing regional diversity in EC2. Driving these analyses are several new datasets, including one with over 34 million DNS records for Alexa websites and a packet capture from a large university network.
Keqiang He, Alexis Fisher, Liang Wang 0023, Aaron Gember, Aditya Akella, Thomas Ristenpart
Internet Measurement Conference1
2012 Greening the Internet Using Multi-frequency Scaling Scheme
abstract
In this paper, we have designed a Multi-Frequency Scaling scheme for energy conservation of network devices, especially routers and switches. The frequency of components in a network device is scaled dynamically according to the real time workload. A Markov model is developed for performance analysis of this mechanism. We implement a prototype of this scheme in the data path of a general IPv4 router based on a real hardware platform - NetFPGA. Experimental results show excellent energy savings at the cost of a tolerable latency, under various ranges of traffic loads. Our work indicates the feasibility and possibility of deploying this mechanism into real network devices for energy saving.
Wei Meng 0001, Yi Wang 0004, Chengchen Hu, Keqiang He, Jun Li 0003, Bin Liu 0001
AINA4
2012 Scalable Name Lookup in NDN Using Effective Name Component Encoding
abstract
Name-based route lookup is a key function for Named Data Networking (NDN). The NDN names are hierarchical and have variable and unbounded lengths, which are much longer than IPv4/6 address, making fast name lookup a challenging issue. In this paper, we propose an effective Name Component Encoding (NCE) solution with the following two techniques: (1) A code allocation mechanism is developed to achieve memory-efficient encoding for name components, (2) We apply an improved State Transition Arrays to accelerate the longest name prefix matching and design a fast and incremental update mechanism which satisfies the special requirements of NDN forwarding process, namely to insert, modify, and delete name prefixes frequently. Furthermore, we analyze the memory consumption and time complexity of NCE. Experimental results on a name set containing 3,000,000 names demonstrate that compared with the character trie NCE reduces overall 30% memory. Besides, NCE performs a few millions lookups per second (on an Intel 2.8 GHz CPU), a speedup of over 7 times compared with the character trie. Our evaluation results also show that NCE can scale up to accommodate the potential future growth of the name sets.
Yi Wang 0004, Keqiang He, Huichen Dai, Wei Meng 0001, Junchen Jiang, Bin Liu 0001, Yan Chen 0004
ICDCS2
2012 Reducing power of traffic manager in routers via dynamic on/off-chip scheduling
abstract
Green networking in the Internet becomes increasingly important. In a high-performance router, the dominant power consumer on the Internet, half of its total power usage goes into the line-cards, where the traffic managers inside consume most of it. In this paper, we propose an energy-efficient design on the traffic manager architecture for packet buffering and storage. Unlike traditional routers where packets are always kept in off-chip memory, we propose a dynamic on-chip and off-chip scheduling mechanism, called Dynamic Packet Manager (DPM), to reduce both peak and average power consumption caused by the traffic manager. DPM buffers packets in a small on-chip memory in the light-traffic period, and activates the off-chip memory on when the on-chip memory is to overflow. In this design, when the traffic is light, the off-chip memory is put into power saving state by clock gating so that the average power consumption is reduced. With an on-chip flow based and off-chip class-based design, DPM can save one off-chip memory otherwise used for the per-flow index information storage, therefore further reduce the peak power usage. We present the theoretic analysis guiding the implementation of the DPM mechanism. Experiments on three prototypes implemented on different hardware show that the peak and average power consumptions can be reduced by 27.9% and 37.5% respectively, along with less on-chip memory cost. Besides, the traffic manger with DPM shows better performance on average packet scheduling delay than the one without DPM.
Jindou Fan, Chengchen Hu, Keqiang He, Junchen Jiang, Bin Liu 0001
INFOCOM3
2012 Measurements on movie distribution behaviour in peer-to-peer networks
abstract
Peer-to-Peer (P2P) mode dominates the way that files are shared over the Internet today. A measurement study on the user behaviour during the P2P file sharing is important and helpful to better understand and design P2P networks. In this study, the authors developed a method to collect information about peers and connections in movie sharing at the BitTorrent client side. Movie is selected as the investigation object since its immense popularity and large size among all the file types over P2P networks. The method proposed in this study can be easily applied to study the distribution behaviour of other types of files. Based on the collected data, the authors have derived 10 observations in three categories: (i) distributions of peers and connections over globe time and local time (after adjustment of time differences); (ii) distributions of peers and connections over geographic areas (at different levels of continents, countries, cities); and (iii) the influence to the above distributions by differences of population, gross domestic product (GDP) and life style.
Chengchen Hu, Xiaojun Wang 0001, Keqiang He, Bin Liu 0001
IET Commun.3
2011 Parallel Name Lookup for Named Data Networking
abstract
Name-based route lookup is a key function for Named Data Networking (NDN). The NDN names are hierarchical and have variable and unbounded lengths, which are much longer than IPv4/6 address, making fast name lookup a challenging issue. In this paper, we propose a parallel architecture for NDN name lookup called Parallel Name Lookup (PNL) which leverages hardware parallelism to achieve high lookup speedup while keeping a low and controllable memory redundancy. The core of PNL is an allocation algorithm that maps the logically tree-based structure to physically parallel modules, with low computational complexity. We evaluate the PNL's performance and show that PNL dramatically accelerates the name lookup process. Furthermore, with certain knowledge of prior probability, the speedup can be significantly improved.
Yi Wang 0004, Huichen Dai, Junchen Jiang, Keqiang He, Wei Meng 0001, Bin Liu 0001
GLOBECOM4
2011 Measurements on movie distribution behavior in Peer-to-Peer networks
abstract
Peer-to-Peer (P2P) mode dominates the way that files are shared over the Internet today. A measurement study on the user behavior during the P2P file sharing is important and helpful to better understand and design P2P networks. In this paper, we developed a method to collect information about peers and connections in movie sharing at the BitTorrent client side. Based on the collected data, we have derived 5 observations in the influence upon peers and connections distributions over geographic areas (at different levels of continents, countries, cities) by differences of population, GDP (Gross Domestic Product), time zone and life style.
Xiaofei Wang 0006, Xiaojun Wang 0001, Chengchen Hu, Keqiang He, Junchen Jiang, Bin Liu 0001
Integrated Network Management4
2010 A2C: Anti-Attack Counters for Traffic Measurement
abstract
Flow-level sampling methods have been widely studied and extensively employed in network traffic measurement systems. However, traffic anomalies are becoming more prevalent and severe in the Internet, which pose great challenges to the traffic measurement. Existing solutions targeted at such scenario have either low accuracy or high memory usage. In this paper, we propose a two-stage sampling approach-Anti Attack Counters (A2C) and an efficient parameter adapting method to solve the problem. The proposed sampling mechanism can adapt to the network condition automatically and collect more information even under severe traffic attacks. Theoretical analysis on accuracy and resource requirement is presented in our work. Furthermore, we validate our approach using both synthetic and real traces. The experimental results demonstrate that A2C is of high resilience while providing significantly improved measurement accuracy with reduced memory occupation comparing with other existing anti-attack countermeasures.
Keqiang He, Chengchen Hu, Junchen Jiang, Yachao Zhou, Bin Liu 0001
GLOBECOM1
2010 Parallel Architecture for High Throughput DFA-Based Deep Packet Inspection
abstract
Multi-pattern matching is a key technique for implementing network security applications such as Network Intrusion Detection/Protection Systems (NIDS/NIPSes) where every packet is inspected against predefined attack signatures written in regular expressions (regexes). To this end, Deterministic Finite Automaton (DFA) is widely used for multi-regex matching, but existing DFAbased researches have claimed high throughput at an expenses of extremely high memory cost. In this paper, we propose a parallel architecture of DFA called Parallel DFA (PDFA), using multiple flow aggregations to increase the throughput with nearly no extra memory cost. The basic idea is to selectively store the DFA in multiple memory modules which can be accessed in parallel and to explore the potential parallelism. The memory cost of our system in both the average cases and the worst cases is analyzed, optimized and evaluated by numerical results. The evaluation shows that we obtain an average speedup of about 0.5k to 0.7k where k is the number of parallel memory modules under our synthetic trace and compressed real trace in a statistical average case, compared with the traditional DFA-based matching approaches.
Junchen Jiang, Xiaofei Wang 0006, Keqiang He, Bin Liu 0001
ICC3
1994 Latency Metric: An Experimental Method for Measuring and Evaluating Parallel Program and Architecture Scalability
Xiaodong Zhang 0001, Yong Yan 0003, Keqiang He
J. Parallel Distributed Comput.3