EDBT 2026 Demo / reviewers in the wild / expert
Binzhang Fu
dblp:60/7673
· DBLP profile ↗
36ranked-venue papers
6as first author
17since 2021 · last 2026
0009-0008-1213-0554ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 23 · 1 first-author · 16 since 2021Systems, architecture and hardware · 10 · 4 first-author · 1 since 2021Software engineering, systems software and programming languages · 4 · 4 first-authorSecurity and privacy · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Balancing and Beyond: Communication-Centric Optimizations in Expert ParallelismabstractThe Mixture-of-Experts (MoE) architecture scales large language models (LLMs) to trillions of parameters by activating only a small subset of experts per token. In practice, MoE inference is commonly deployed with Expert Parallelism (EP), which places whole experts on different GPUs to preserve kernel efficiency. However, production EP deployments often suffer from two bottlenecks: (1) expert workload imbalance, which creates computation and communication stragglers, and (2) communication inefficiency, where inter-GPU transfers dominate latency even after balancing. We present EPIC, an experience-driven EP inference system that addresses these issues progressively for real deployments. EPIC mitigates imbalance via performance-aware expert migration and runtime expert activation, and then improves communication with topology-adaptive transport kernels and fine-grained computation-communication overlap. EPIC has been deployed at scale across O(10K) GPUs in our online inference service for both open-source models (e.g., Qwen3-Coder and DeepSeek-R1) and internal models, reducing communication time and per-token latency by up to 40% and 21%, respectively. Jiamin Cao, Qingxu Li, Yaozhong Liu, Shangfeng Shi, Kunling He, Ennan Zhai, Jianbo Dong, Binzhang Fu, Dennis Cai |
SIGCOMM | 13 |
| 2026 | EPIC: Abstraction and Polymorphism of In-Network Collectives on Ethernet
Yitao Yuan, Jianglong Nie, Tianyu Bai, Ruizhe Zhou, Siyuan Cao, Xujie Fan, Yuchen Xu 0003, Junkai Chen, Chenqi Zhao, Nengyuan Zhang, Shaoke Fang, Jiangyuan Chen, Yuanfeng Chen, Zhan Wang 0003, Yuchao Zhang 0004, Yang Liu 0038, Xiangrui Yang 0002, Xiaohe Hu, Limin Xiao 0001, Weifeng Zhang 0003, Yazhu Lan, Jianbo Dong, Binzhang Fu, Wenfei Wu |
SIGCOMM | 30 |
| 2026 | Anytest: Localizing the Root Cause of Hardware Transport Performance Anomalies
Zhaochen Zhang, Sheng Cheng 0002, Feiyang Xue, Chang Liu 0001, Boliang Liu, Rui Li 0020, Li Wang 0110, Peirui Cao, Qingkai Meng 0001, Guihai Chen, Shuguang Cheng, Yongqing Xi, Binzhang Fu, Dennis Cai, Chen Tian 0001 |
SIGCOMM | 20 |
| 2026 | Advancing RDMA Scalability With High PerformanceabstractDue to its superior performance, Remote Direct Memory Access (RDMA) has been widely deployed in data center networks. It provides applications with ultra-high throughput, ultra-low latency, and far lower CPU utilization than TCP/IP software network stack. However, the connection states that must be stored on the RDMA NIC (RNIC) and the small NIC memory result in poor scalability. The performance drops significantly when the RNIC needs to maintain a large number of concurrent connections. We propose StaR (Stateless RDMA), which solves the scalability problem of RDMA by transferring states to the other communication end in a trusted network. Leveraging the asymmetric communication pattern in data center applications, StaRlets the communication end with low NIC memory usage to save states for the other end with high NIC memory usage, thus making the RNIC on the bottleneck side stateless. We implemented StaR on an FPGA board with a 10Gbps network port and NS-3, evaluating its performance on a testbed with 9 machines, each equipped with StaR NICs, and verified its scalability stability by conducting a larger-scale simulation with 200 fully connected nodes using a 100Gbps link. The experimental results show that in high concurrency scenarios, the throughput of StaR can reach up to 4.13x and 1.35x of the original RNIC and the latest software-based solution, respectively. Xijin Yin, Guo Chen 0001, Xizheng Wang, Huichen Dai, Bojie Li, Binzhang Fu, Kun Tan 0002 |
IEEE Trans. Netw. | 7 |
| 2025 | Enhancing Large-Scale AI Training Efficiency: The C4 Solution for Real-Time Anomaly Detection and Communication OptimizationabstractThe emergence of Large Language Models (LLMs) has necessitated the adoption of distributed training techniques, involving the deployment of thousands of GPUs to train a single model. Unfortunately, the efficiency of large-scale distributed training systems is often suboptimal due to the increased likelihood of hardware errors in high-end GPU products and the heightened risk of network traffic collisions. Specifically, GPUs involved in the same job require periodic synchronization to exchange necessary data, such as gradients, parameters, or activations. As a result, any local hardware failure can disrupt training tasks, and the inability to swiftly identify faulty components leads to a significant waste of GPU resources. Moreover, prolonged communication due to traffic collisions can substantially increase GPU waiting times. To address these challenges, we propose a communication-driven solution, namely the C 4. The key insights of C 4 are twofold. First, the load in distributed training exhibits homogeneous characteristics and is divided into iterations through periodic synchronization, therefore hardware anomalies would incur certain syndrome in collective communication. By leveraging this feature, $\mathbf{C} 4$ can rapidly identify the faulty components, swiftly isolate the anomaly, and restart the task, thereby avoiding resource wastage caused by delays in anomaly detection. Second, the predictable communication model of collective communication, involving a limited number of long-lived flows, allows C 4 to efficiently execute traffic planning, substantially reducing bandwidth competition among these flows. The $\mathbf{C 4}$ has been extensively deployed across real-world production systems in a hyperscale cloud provider, yielding a significant improvement in system efficiency, from 30% to $\mathbf{4 5 \%}$. This enhancement is attributed to a $\mathbf{3 0 \%}$ reduction in error-induced overhead and a 15% reduction in communication costs. Jianbo Dong, Yikai Zhu, Hairong Jiao, Ennan Zhai, Wencong Xiao, Man Yuan, Siran Yang, Jiamang Wang, Rui Men, Dennis Cai, Binzhang Fu |
HPCA | 25 |
| 2025 | Evolution of Aegis: Fault Diagnosis for AI Model Training Service in Production
Jianbo Dong, Kun Qian 0021, Zhilong Zheng, Liang Chen 0001, Yichi Xu, Yikai Zhu, Xue Li 0024, Zhihui Ren, Yang Liu 0245, Yu Guan 0005, Chaojie Yang, Yang Zhang 0102, Man Yuan, Yong Li 0008, Xianlong Zeng, Zhiping Yao, Binzhang Fu, Ennan Zhai, Wei Lin 0016, Dennis Cai |
NSDI | 28 |
| 2025 | SimAI: Unifying Architecture Design and Performance Tuning for Large-Scale Large Language Model Training with Scalability and Precision
Xizheng Wang, Qingxu Li, Yichi Xu, Dan Li 0001, Li Chen 0008, Heyang Zhou, Linkang Zheng, Yikai Zhu, Yang Liu 0245, Kun Qian 0021, Kunling He, Ennan Zhai, Dennis Cai, Binzhang Fu |
NSDI | 18 |
| 2025 | GPU-Disaggregated Serving for Deep Learning Recommendation Models at Scale
Lingyun Yang, Yongchen Wang, Yinghao Yu, Qizhen Weng 0001, Jianbo Dong, Chi Zhang 0005, Yanyi Zi, Zechao Zhang, Menglei Zheng, Lanlan Xi, Binzhang Fu, Tao Lan, Liping Zhang 0013, Lin Qu, Wei Wang 0030 |
NSDI | 18 |
| 2025 | SyCCL: Exploiting Symmetry for Efficient Collective Communication SchedulingabstractThe performance of collective communication schedules is crucial for the efficiency of machine learning jobs and GPU cluster utilization. Existing open-source collective communication libraries (such as NCCL and RCCL) rely on fixed schedules and cannot adjust to varying topology and model requirements. State-of-the-art collective schedule synthesizers (such as TECCL and TACCL) utilize Mixed Integer Linear Program for modeling but encounter search space explosion and scalability challenges. In this paper, we propose SyCCL, a scalable collective schedule synthesizer that aims to synthesize near-optimal schedules in tens of minutes for production-scale machine-learning jobs. SyCCL leverages collective and topology symmetries to decompose the original collective communication demand into smaller sub-demands within smaller topology subsets. SyCCL proposes efficient search strategies to quickly explore potential sub-demands, synthesizes corresponding sub-schedules, and integrates these sub-schedules into complete schedules. Our 32-A100 testbed and production-scale simulation experiments show that SyCCL improves collective performance by up to 127% while reducing synthesis time by 2 to 4 orders of magnitude compared to state-of-the-art efforts. Jiamin Cao, Shangfeng Shi, Weisen Liu, Yifan Yang 0009, Yichi Xu, Zhilong Zheng, Yu Guan 0005, Kun Qian 0021, Ying Liu 0024, Mingwei Xu 0001, Ning Wang 0001, Jianbo Dong, Binzhang Fu, Dennis Cai, Ennan Zhai |
SIGCOMM | 15 |
| 2025 | Alibaba Stellar: A New Generation RDMA Network for Cloud AIabstractThe rapid adoption of Large Language Models (LLMs) in cloud environments has intensified the demand for high-performance AI training and inference, where Remote Direct Memory Access (RDMA) plays a critical role. However, existing RDMA virtualization solutions, such as Single-Root Input/Output Virtualization (SR-IOV), face significant limitations in scalability, performance, and stability. These issues include lengthy container initialization times, hardware resource constraints, and inefficient traffic steering. To address these challenges, we propose Stellar, a new generation RDMA network for cloud AI. Stellar introduces three key innovations: Para-Virtualized Direct Memory Access (PVDMA) for on-demand memory pinning, extended Memory Translation Table (eMTT) for optimized GPU Direct RDMA (GDR) performance, and RDMA Packet Spray for efficient multi-path utilization. Deployed in our large-scale AI clusters, Stellar spins up virtual devices in seconds, reduces container initialization time by 15 times, and improves LLM training speed by up to 14%. Our evaluations demonstrate that Stellar significantly outperforms existing solutions, offering a scalable, stable, and high-performance RDMA network for cloud AI. Menglei Zheng, Binbin Liao, Suwei Xu, Yongjia Mo, Qinghua Peng, Jilie Luo, Qingxu Li, Zishu Wang, Jianbo Dong, Kunling He, Sheng Cheng 0002, Jiamin Cao, Hairong Jiao, Lingjun Zhu, Yiquan Chen, Wei Wang 0030, Shuhong Zhu, Xingru Li, Qiang Wang 0022, Wei Lin 0016, Ennan Zhai, Jiesheng Wu, Qiang Liu 0036, Binzhang Fu, Dennis Cai |
SIGCOMM | 38 |
| 2024 | Kspeed: Beating I/O Bottlenecks of Data Provisioning for RDMA Training ClustersabstractThe rapidly-increasing computing power of GPUs has rendered the I/O subsystem a bottleneck for distributed deep learning (DL) training. Currently, substantial data preprocessing work (e.g., decoding) has to be conducted on CPUs for a wide range of training scenarios such as computer vision (CV) and audio. Unfortunately, the involvement of training nodes' host memory and/or CPUs on the critical path of loading data to GPUs incurs significant GPU stalls in modern RDMA training clusters, because CPUs are much slower than GPUs and the connection from PCIe switches to host memory tends to suffer from incast problems. Moreover, this also incurs high CPU usage and resource contention, which consequently causes data loading performance variation and stragglers. This paper presents KSpeed, a novel data provisioning framework for large-scale RDMA training clusters. As many data preprocessing tasks need to be done by CPUs, KSpeed organizes host memory and CPU resources in the cluster to build a disaggregated memory/CPU pool, where the nodes can read raw input data from backend storage to their host memory, preprocess the data by their CPUs if necessary, and write cached/preprocessed data (on demand) directly to the training workers' GPU memory to minimize GPU stalls. KSpeed leverages the multi-rail RDMA network to eliminate unnecessary memory copies, interference, and congestion. Evaluation on a 96-GPU cluster shows that KSpeed delivers$5.4 \times \sim 100 \times$higher data loading performance over the state-of-the-art designs (DPP and Alluxio). KSpeed achieves near-linear scalability as the GPU number increases from 8 to 512. Jianbo Dong, Hao Qi 0008, Tianjing Xu, Xiaoli Liu 0002, Rongyao Wang, Xiaoyi Lu 0001, Zheng Cao 0003, Binzhang Fu |
ICNP | 9 |
| 2024 | Crux: GPU-Efficient Communication Scheduling for Deep Learning TrainingabstractDeep learning training (DLT), e.g., large language model (LLM) training, has become one of the most important services in multitenant cloud computing. By deeply studying in-production DLT jobs, we observed that communication contention among different DLT jobs seriously influences the overall GPU computation utilization, resulting in the low efficiency of the training cluster. In this paper, we present Crux, a communication scheduler that aims to maximize GPU computation utilization by mitigating the communication contention among DLT jobs. Maximizing GPU computation utilization for DLT, nevertheless, is NP-Complete; thus, we formulate and prove a novel theorem to approach this goal by GPU intensity-aware communication scheduling. Then, we propose an approach that prioritizes the DLT flows with high GPU computation intensity, reducing potential communication contention. Our 96-GPU testbed experiments show that Crux improves 8.3% to 14.8% GPU computation utilization. The large-scale production trace-based simulation further shows that Crux increases GPU computation utilization by up to 23% compared with alternatives including Sincronia, TACCL, and CASSINI. Jiamin Cao, Yu Guan 0005, Kun Qian 0021, Wencong Xiao, Jianbo Dong, Binzhang Fu, Dennis Cai, Ennan Zhai |
SIGCOMM | 7 |
| 2024 | Alibaba HPN: A Data Center Network for Large Language Model TrainingabstractThis paper presents HPN, Alibaba Cloud's data center network for large language model (LLM) training. Due to the differences between LLMs and general cloud computing (e.g., in terms of traffic patterns and fault tolerance), traditional data center networks are not well-suited for LLM training. LLM training produces a small number of periodic, bursty flows (e.g., 400Gbps) on each host. This characteristic of LLM training predisposes Equal-Cost Multi-Path (ECMP) to hash polarization, causing issues such as uneven traffic distribution. HPN introduces a 2-tier, dual-plane architecture capable of interconnecting 15K GPUs within one Pod, typically accommodated by the traditional 3-tier Clos architecture. Such a new architecture design not only avoids hash polarization but also greatly reduces the search space for path selection. Another challenge in LLM training is that its requirement for GPUs to complete iterations in synchronization makes it more sensitive to singlepoint failure (typically occurring on ToR). HPN proposes a new dual-ToR design to replace the single-ToR in traditional data center networks. HPN has been deployed in our production for more than eight months. We share our experience in designing, and building HPN, as well as the operational lessons of HPN in production. Kun Qian 0021, Yongqing Xi, Jiamin Cao, Yichi Xu, Yu Guan 0005, Binzhang Fu, Xuemei Shi, Fangbo Zhu, Rui Miao 0001, Peng Wang 0185, Xianlong Zeng, Eddie Ruan, Zhiping Yao, Ennan Zhai, Dennis Cai |
SIGCOMM | 7 |
| 2023 | Dependable Virtualized Fabric on Programmable Data PlaneabstractIn modern multi-tenant data centers, each tenant desires reassuring dependability from the virtualized network fabric – bandwidth guarantee with work conservation, bounded tail latency and resilient reachability. However, the slow convergence of prior works under network dynamics and uncertainties can hardly provide the dependability for tenants. Further, state-of-the-art load balance schemes are guarantee-agnostic and bring great risks on breaking bandwidth guarantee, which is overlooked in prior works. In this paper, we propose vFab, a dependable virtualized fabric framework which can (1) quickly detect network failure in data plane, (2) explicitly select proper paths for all flows, and (3) converge to ideal bandwidth allocation at sub-millisecond. The core idea of vFab is to leverage the programmable data plane to build a fusion of an active edge (e.g., NIC) and an informative core (e.g., switch), where the core sends link status and tenant information to the edge via telemetry to help the latter make a timely and accurate decision on path selection and traffic admission. We fully implement vFab with commodity SmartNICs and programmable switches. Extensive evaluations show that vFab can keep bandwidth guarantee with high bandwidth utilization, low and bounded latency, and resilient reachability under various network scenarios with limited overhead. Application-level experiments show that vFab can improve QPS by$2.4\times $and cut tail latency by$10\times $compared to the alternatives. Kaihui Gao, Shuai Wang 0028, Kun Qian 0021, Dan Li 0001, Rui Miao 0001, Bo Li 0061, Yu Zhou 0008, Ennan Zhai, Chen Sun 0005, Binzhang Fu, Frank Kelly, Dennis Cai, Hongqiang Harry Liu, Tao Sun 0010 |
IEEE/ACM Trans. Netw. | 12 |
| 2022 | Predictable vFabric on informative data planeabstractIn multi-tenant data centers, each tenant desires reassuring predictability from the virtual network fabric - bandwidth guarantee, work conservation, and bounded tail latency. Achieving these goals simultaneously relies on rapid and precise traffic admission. However, the slow convergence (tens of milliseconds) of prior works can hardly satisfy the increasingly rigorous performance demand under dynamic traffic patterns. Further, state-of-the-art load balance schemes are all guarantee-agnostic and bring great risks on breaking bandwidth guarantee, which is overlooked in prior works. Shuai Wang 0028, Kaihui Gao, Kun Qian 0021, Dan Li 0001, Rui Miao 0001, Bo Li 0061, Yu Zhou 0008, Ennan Zhai, Chen Sun 0005, Binzhang Fu, Frank Kelly, Dennis Cai, Hongqiang Harry Liu, Ming Zhang 0005 |
SIGCOMM | 12 |
| 2022 | From luna to solar: the evolutions of the compute-to-storage networks in Alibaba cloudabstractThis paper presents the two generations of storage network stacks that reduced the average I/O latency of Alibaba Cloud's EBS service by 72% in the last five years: Luna, a user-space TCP stack that corresponds the latency of network to the speed of SSD; and Solar, a storage-oriented UDP stack that enables both storage and network hardware accelerations. Rui Miao 0001, Lingjun Zhu, Kun Qian 0021, Shujun Zhuang, Bo Li 0061, Shuguang Cheng, Binzhang Fu, Jiaji Zhu, Jiesheng Wu, Dennis Cai, Hongqiang Harry Liu |
SIGCOMM | 13 |
| 2021 | StaR: Breaking the Scalability Limit for RDMAabstractDue to its superior performance, Remote Direct Memory Access (RDMA) has been widely deployed in data center networks. It provides applications with ultra-high throughput, ultra-low latency, and far lower CPU utilization than TCP/IP software network stack. However, the connection states that must be stored on the RDMA NIC (RNIC) and the small NIC memory result in poor scalability. The performance drops significantly when the RNIC needs to maintain a large number of concurrent connections.We propose StaR (Stateless RDMA), which solves the scalability problem of RDMA by transferring states to the other communication end. Leveraging the asymmetric communication pattern in data center applications, StaR lets the communication end with low concurrency save states for the other end with high concurrency, thus making the RNIC on the bottleneck side to be stateless. We have implemented StaR on an FPGA board with 10Gbps network port and evaluated its performance on a testbed with 9 machines all equipped with StaR NICs. The experimental results show that in high concurrency scenarios, the throughput of StaR can reach up to 4.13x and 1.35x of the original RNIC and the latest software-based solution, respectively. Xizheng Wang, Guo Chen 0001, Xijin Yin, Huichen Dai, Bojie Li, Binzhang Fu, Kun Tan 0002 |
ICNP | 6 |
| 2020 | MasQ: RDMA for Virtual Private CloudabstractRDMA communication in virtual private cloud (VPC) networks is still a challenging job due to the difficulty in fulfilling all virtualization requirements without sacrificing RDMA communication performance. To address this problem, this paper proposes a software-defined solution, namely, MasQ, which is short for "queue masquerade". The core insight of MasQ is that all RDMA communications should associate with at least one queue pair (QP). Thus, the requirements of virtualization, such as network isolation and the application of security rules, can be easily fulfilled if QP's behavior is properly defined. In particular, MasQ exploits the virtio-based paravirtualization technique to realize the control path. Moreover, to avoid performance overhead, MasQ leaves all data path operations, such as sending and receiving, to the hardware. We have implemented MasQ in the OpenFabrics Enterprise Distribution (OFED) framework and proved its scalability and performance efficiency by evaluating it against typical applications. The results demonstrate that MasQ achieves almost the same performance as bare-metal RDMA for data communication. Binzhang Fu, Kun Tan 0002, Bei Hua, Zhi-Li Zhang, Kai Zheng 0003 |
SIGCOMM | 3 |
| 2020 | Optimizing TCP Loss Recovery Performance Over Mobile Data NetworksabstractRecent advances in high-speed mobile networks have revealed new bottlenecks in ubiquitous TCP protocol deployed in the Internet. In addition to differentiating non-congestive loss from congestive loss, our experiments revealed two significant performance bottlenecks during the loss recovery phase: flow control bottleneck and application stall, resulting in degradation in QoS performance. To tackle these two problems, we first develop a novel opportunistic retransmission algorithm to eliminate the flow control bottleneck, which enables TCP sender to transmit new packets even if receiver's receiving window is exhausted. Second, application stall can be significantly alleviated by carefully monitoring and tuning the TCP sending buffer growth mechanism. We implemented and modularized the proposed algorithms in the Linux kernel thus they can plug-and-play with the existing TCP loss recovery algorithms easily. We evaluated our proposed algorithms over emulated and real experiments and showed that, compared to existing TCP loss recovery algorithms, the proposed optimization algorithms improve the bandwidth efficiency by up to 133 percent and completely mitigate RTT spikes, i.e., over 50 percent RTT reduction, over the loss recovery phase. Ke Liu 0004, Zhongbin Zha, Wenkai Wan, Vaneet Aggarwal, Binzhang Fu, Mingyu Chen 0001 |
IEEE Trans. Mob. Comput. | 5 |
| 2019 | Towards Stateless RNIC for Data Center NetworksabstractBecause of small NIC on-chip memory, the massive connection states maintained on Remote Direct Memory Access (RDMA) NIC (RNIC) significantly limit its scalability. When the number of concurrent connections grows, RNICs have to frequently fetch connection states from host memory, leading to dramatic performance degradation. In this paper, we propose StaR, which fundamentally solves this scalability issue by making RNIC stateless. Leveraging the asymmetric communication pattern in data center applications, the StaR RNIC stores zero connection-related states by moving all the connection states to the other end. Through careful design, StaR RNICs can maintain unchanged RDMA semantics and avoid security issues even when processing traffic statelessly. Preliminary simulation results show that StaR can improve the aggregate throughput by more than 160x (stress test) and 4x (application) compared to original RNICs. Pulin Pan, Guo Chen 0001, Xizheng Wang, Huichen Dai, Bojie Li, Binzhang Fu, Kun Tan 0002 |
APNet | 6 |
| 2019 | vSocket: virtual socket interface for RDMA in public cloudsabstractRDMA has been widely adopted as a promising solution for high performance networks, but is still unavailable for a large number of socket-based applications running in public clouds due to the following reasons. There is no available virtualization technique of RDMA that can meet the cloud's requirements. Moreover, it is cost prohibitive to rewrite the socket-based applications with the Verbs API. To address the above problems, we present vSocket, a software-based RDMA virtualization framework for socket-based applications in public clouds. vSocket takes into account the demands of clouds such as security rules and network isolation, so it can be deployed in the current public clouds. Furthermore, vSocket provides native socket API so that socket-based applications can use it without any modifications. Finally, to validate the performance gains, we implemented a prototype and compared it with current virtual network solutions against 1) basic network benchmarks and 2) the Redis, a typical I/O intensive application. Experimental results show that the latency of basic benchmarks can be reduced by 88% and the throughput of Redis is improved by 4 times. Binzhang Fu, Bei Hua |
VEE | 2 |
| 2017 | Footprint: Regulating Routing Adaptiveness in Networks-on-Chip
Binzhang Fu, John Kim 0001 |
ISCA | 1 |
| 2017 | Stem: A Table-Based Congestion Control Framework for Virtualized Data Center Networks
Binzhang Fu, Mingyu Chen 0001 |
NPC | 2 |
| 2017 | Efficient Regional Congestion Awareness (ERCA) for Load Balance with Aggregated Congestion InformationabstractIn this paper, we propose Efficient Regional Congestion Awareness (ERCA), a novel adaptive routing technique which utilizes both local and non-local/aggregated congestion information to estimate network congestion. To reduce the interference caused by the noises of the congestion information, two optimizations are performed: first, to minimize the noises of the congestion information, ERCA only considers the congestion status of the links adjacent to the nodes defined from current to the corresponding boundary, second, to minimize the influence of the noises, ERCA exploits dynamic instead of static weights to evaluate port's congestion. Furthermore, ERCA uses one wire per dimension direction or quadrant to transmit congestion information between two adjacent nodes, so the wiring overhead of ERCA is minimal. Compared with the DBAR, ERCA can maximally improve the saturation throughput by 11.7% and averagely improve the saturation throughput by 6.02%. Binzhang Fu, Mingyu Chen 0001, Lixin Zhang 0002 |
PDP | 3 |
| 2016 | Isolating bandwidth guarantees from work conservation in the cloudabstractTo predict lower bounds on the performance of applications, the cloud should provide guarantees on bandwidth that each virtual machine can obtain. By competing for spare network bandwidth, current solutions that provide bandwidth guarantees can achieve work conservation as well. However, they usually fail to provide accurate bandwidth guarantees, for the interference between traffic for achieving the two objectives respectively. In order to eliminate the interference, they reserve sufficient bandwidth headroom for every link, which cannot be allocated to tenants as guarantees, incurring a decrease in the total of guarantees that each link can offer and thus a decline in the revenue of the cloud provider. To address the problem, this paper proposes a new mechanism, namely DFlow, which achieves bandwidth guarantees and work conservation simultaneously. Specifically, DFlow isolates its solutions for achieving bandwidth guarantees and work conservation from each other by splitting every flow into two subflows with distinct priorities. They are then used to achieve the two objectives respectively. Our evaluations show that DFlow can provide accurate bandwidth guarantees without reserving any bandwidth headroom while achieving work conservation to effectively utilize spare network bandwidth. Ke Liu 0004, Binzhang Fu, Mingyu Chen 0001, Lixin Zhang 0002 |
ISCC | 3 |
| 2015 | Optimizing TCP loss recovery performance over mobile data networksabstractRecent advances in high-speed mobile networks have revealed new bottlenecks in ubiquitous TCP protocol deployed in the Internet. In addition to differentiating non-congestive loss from congestive loss, our experiments revealed two significant performance bottlenecks during loss recovery phase: flow control bottleneck and application stall, resulting in degradation in QoS performance. To tackle these two problems we firstly develop a novel opportunistic retransmission algorithm to eliminate the flow control bottleneck, which enables TCP sender to transmit new packets even if receiver's receiving window is exhausted. Secondly, the application stall can be significantly alleviated by carefully monitoring and tuning the TCP send buffer growth mechanism. We implemented and modularized the proposed algorithms in the Linux kernel thus they can plug-and-play with the existing TCP loss recovery algorithms easily. Using emulated experiments we showed that with the proposed optimization techniques the existing loss recovery algorithms can at most achieve 98.3% bandwidth utilization during loss recovery phase, and reduce RTT by at most 80% after loss recovery phase. Zhongbin Zha, Ke Liu 0004, Binzhang Fu, Mingyu Chen 0001 |
SECON | 3 |
| 2015 | RISO: Enforce Noninterfered Performance With Relaxed Network-on-Chip Isolation in Many-Core Cloud ProcessorsabstractWorkload consolidation is widely used in modern cloud processors to reduce total cost of ownership. Performance isolation has to be enforced between consolidated workloads to achieve controllable quality of service. Networks-on-chip (NoCs), as a major shared resource, often incur traffic interference and violate performance isolation criteria. Previous work resorts to strict isolation strategy that partitions NoC into independent regions to isolate core-to-core communication traffic. However, strict isolation either results in low consolidation density or degrades network performance, and more importantly, cannot be applied to memory access traffic. To address these weaknesses, we propose a novel performance isolation strategy in NoC, called relaxed isolation (RISO). It permits underutilized routers and links to be shared by multiple applications, and, at the same time, it keeps the aggregated traffic in check to enforce performance isolation. Experimental results show that RISO could effectively improve consolidation density and network performance in synergy. Binzhang Fu, Ying Wang 0001, Yinhe Han 0001, Guihai Yan, Xiaowei Li 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2015 | Economizing TSV Resources in 3-D Network-on-Chip DesignabstractThe confluence of 3-D integration and network-on-chip (NoC) provides an effective solution to the scalability problem of on-chip interconnects. In 3-D integration, through-silicon via (TSV) is considered to be the most promising bonding technology. However, TSVs are also precious link resources because they consume significant chip area and possibly lead to routing congestion in the physical design stage. In addition, TSVs suffer from serious yield losses that shrink the effective TSV density. Thus, it is necessary to implement a TSV-economical 3-D NoC architecture in cost-effective design. For symmetric 3-D mesh NoCs, we observe that the TSVs bandwidth utilization is low and they rarely become the contention spots in networks as planar links. Based on this observation, we propose the TSV sharing (TS) scheme to save TSVs in 3-D NoC by enabling neighboring routers to share the vertical channels in a time division multiplexing way. We also investigate different TS implementation alternatives and show how TS improves TSV-effectiveness (TE) in multicore processors through a design space exploration. In experiments, we comprehensively evaluate TSs influence on all layers of system. It is shown that the proposed method significantly promotes TE with negligible performance overhead. Ying Wang 0001, Yinhe Han 0001, Lei Zhang 0008, Binzhang Fu, Cheng Liu 0008, Huawei Li 0001, Xiaowei Li 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2014 | Dandelion: A locally-high-performance and globally-high-scalability hierarchical data center networkabstractThe increasing customer demand is driving modern data centers to embrace the freely-expandable network architecture. Unfortunately, state-of-the-art freely-expandable networks suffer from either the large granularity of expansion or the prohibitive implementation cost. Furthermore, a recent research showed that data center traffic tends to be highly clustered. Based on above observations, this paper proposes a freely-expandable network architecture, namely the dandelion. Dandelion is a two-level hierarchical network, where the first level aims at “high performance” and the second level aims at “high scalability”. The resulting network has two distinct advantages. First, it could arbitrarily expand with a reasonable granularity. Second, the router architecture is efficient as well as highly scalable since 1) the routing table is significantly compressed and 2) a fixed number of virtual channels per physical channel are required regardless of the network size. Finally, the traffic characteristics of four typical cloud applications are analyzed, and the generated traffic patterns are used to evaluate the proposed network architecture. Simulation results prove that the dandelion is a promising network architecture for future data centers. Binzhang Fu, Wentao Bao, Guolong Jiang, Mingyu Chen 0001, Lixin Zhang 0002, Yidong Tao, Junfeng Zhao 0003 |
ICCCN | 1 |
| 2014 | QBLESS: A case for QoS-aware bufferless NoCsabstractDatacenters consolidate diverse applications to improve utilization. However when multiple applications are co-located on such platforms, contention for shared resources like Networks-on-Chip (NoCs) can degrade the performance of latency-critical online services (high-priority applications). Recently proposed bufferless NoCs have the advantages of requiring less area and power, but they pose challenges in quality-of-service (QoS) support, which usually relies on buffer-based virtual channels (VCs). We propose QBLESS, a QoS-aware bufferless NoC scheme for datacenters. QBLESS consists of two components: a routing mechanism (QBLESS-R) that can substantially reduce flit deflection for high-priority applications, and a congestion-control mechanism (QBLESS-CC) that guarantees performance for high-priority applications and improves overall system throughput. We use trace-driven simulation to model a 64-core system, finding that when compared to BLESS, a previous state-of-the-art bufferless NoC design, QBLESS improves performance of high-priority applications by an average of 33.2%. Zhicheng Yao, Xiufeng Sui, Tianni Xu, Jiuyue Ma, Sally A. McKee, Binzhang Fu, Yungang Bao |
IWQoS | 7 |
| 2014 | A High-Performance and Cost-Efficient Interconnection Network for High-Density Servers
Wentao Bao, Binzhang Fu, Mingyu Chen 0001, Lixin Zhang 0002 |
J. Comput. Sci. Technol. | 2 |
| 2014 | ZoneDefense: A Fault-Tolerant Routing for 2-D Meshes Without Virtual ChannelsabstractFault-tolerant routing is usually used to provide reliable on-chip communication for many-core processors. This paper focuses on a special class of algorithms that do not use virtual channels. One of the major challenges is to keep the network deadlock free in the presence of faults, especially those locating on network edges. State-of-the-art solutions address this problem by either disabling all nodes of the faulty network edges or including all faults into one faulty block. Therefore, a large number of fault-free nodes will be sacrificed. To address this problem, the proposed ZoneDefense routing not only includes faults into convex faulty blocks but also spreads the faulty blocks' position information in corresponding columns. The nodes, which know the position of faulty blocks, form the defense zones. Therefore, packets can find the faulty blocks and route around them in advance. Exploiting the defense zones, the proposed ZoneDefense routing could tolerate many more faults with significantly reduced sacrificed fault-free nodes compared with the state-of-the-art algorithms. Furthermore, the ZoneDefense routing does not degrade the network performance in the absence of faults, and could get similar performance as its counterparts in the presence of faults. Binzhang Fu, Yinhe Han 0001, Huawei Li 0001, Xiaowei Li 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2013 | RISO: relaxed network-on-chip isolation for cloud processorsabstractCloud service providers use workload consolidation technique in many-core cloud processors to optimize system utilization and augment performance for ever extending scale-out workloads. Performance isolation usually has to be enforced for the consolidated workloads sharing the same many-core resources. Networks-on-chip (NoC) serves as a major shared resource, also needs to be isolated to avoid violating performance isolation. Prior work uses strict network isolation to fulfill performance isolation. However, strict network isolation either results in low consolidation density, or complex routing mechanisms which indicates prohibitive high hardware cost and large latency. In view of this limitation, we propose a novel NoC isolation strategy for many-core cloud processors, called relaxed isolation (RISO). It permits underutilized links to be shared by multiple applications, at the same time keeps the aggregated traffic in check to enforce performance isolation. The experimental results show that the consolidation density is improved more than 12% in comparison with previous strict isolation scheme, meanwhile reducing network latency by 38.4% on average. Guihai Yan, Yinhe Han 0001, Binzhang Fu, Xiaowei Li 0001 |
DAC | 4 |
| 2011 | An abacus turn model for time/space-efficient reconfigurable routingabstractApplications' traffic tends to be bursty and the location of hot-spot nodes moves as time goes by. This will significantly aggregate the blocking problem of wormhole-routed Network-on-Chip (NoC). Most of state-of-the-art traffic balancing solutions are based on fully adaptive routing algorithms which may introduce large time/space overhead to routers. Partially adaptive routing algorithms, on the other hand, are time/space efficient, but lack of even or sufficient routing adaptiveness. Reconfigurable routing algorithms could provide on-demand routing adaptiveness for reducing blocking, but most of them are off-line solutions due to the lack of a practical model to dynamically generate deadlock-free routing algorithms. Binzhang Fu, Yinhe Han 0001, Huawei Li 0001, Xiaowei Li 0001 |
ISCA | 1 |
| 2010 | Accelerating Lightpath setup via broadcasting in binary-tree waveguide in Optical NoCsabstractIn this paper, we propose a binary-tree waveguide connected Optical-Network-on-Chip (ONoC) to accelerate the establishment of the lightpath. By broadcasting the control data in the proposed power-efficient binary-tree waveguide, the maximal hops for establishing lightpath is reduced to two. With extensive simulations and analysis, we demonstrate that the proposed ONoC significantly reduces the setup time, and then the packet latency. Binzhang Fu, Yinhe Han 0001, Huawei Li 0001, Xiaowei Li 0001 |
DATE | 1 |
| 2009 | A New Multiple-Round DOR Routing for 2D Network-on-Chip MeshesabstractThe Network-on-Chip (NoC) meshes are limited by the reliability constraint, which impels us to exploit the fault tolerant routing. Particularly, one of the main design issues is minimizing the loss of non-faulty routers at the presence of faults. To address that problem, we propose a new fault tolerant routing, which has the following two distinct advantages: First, it keeps a network deadlock-free by utilizing restricted intermediate nodes rather than adding virtual channels (VC). This characteristic leads to an area-efficient router. Second, in the proposed routing algorithm, the rounds of DOR are not limited by the number of VC's anymore. As a consequence, the number of sacrificed non-faulty routers is significantly reduced. We demonstrate above advantages through extensive simulations. The experimental results show that under the limitation of VC's, the proposed routing algorithm always sacrifices the minimal number of non-faulty routers compared to previous solutions. Binzhang Fu, Yinhe Han 0001, Huawei Li 0001, Xiaowei Li 0001 |
PRDC | 1 |