Guo Chen 0001

dblp:24/858-1 · DBLP profile ↗
← Back
43ranked-venue papers
8as first author
23since 2021 · last 2026
0000-0002-6069-6869ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 28 · 5 first-author · 12 since 2021Systems, architecture and hardware · 5 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Boosting Adversarial Transferability via Ensemble Non-Attention
abstract
Ensemble attacks integrate the outputs of surrogate models with diverse architectures, which can be combined with various gradient-based attacks to improve adversarial transferability. However, previous work shows unsatisfactory attack performance when transferring across heterogeneous model architectures. The main reason is that the gradient update directions of heterogeneous surrogate models differ widely, making it hard to reduce the gradient variance of ensemble models while making the best of individual model. To tackle this challenge, we design a novel ensemble attack, NAMEA, which for the first time integrates the gradients from the non-attention areas of ensemble models into the iterative gradient optimization process. Our design is inspired by the observation that the attention areas of heterogeneous models vary sharply, thus the non-attention areas of ViTs are likely to be the focus of CNNs and vice versa. Therefore, we merge the gradients respectively from the attention and non-attention areas of ensemble models so as to fuse the transfer information of CNNs and ViTs. Specifically, we pioneer a new way of decoupling the gradients of non-attention areas from those of attention areas, while merging gradients by meta-learning. Empirical evaluations on ImageNet dataset indicate that NAMEA outperforms AdaEA and SMER, the state-of-the-art ensemble attacks by an average of 15.0% and 9.6%, respectively. This work is the first attempt to explore the power of ensemble non-attention in boosting cross-architecture transferability, providing new insights into launching ensemble attacks.
Yipeng Zou, Qin Liu 0001, Jie Wu 0001, Yu Peng 0003, Guo Chen 0001, Hui Zhou 0014, Guanghui Ye
AAAI5
2026 Efficient Cloud-Edge Collaborative Approaches to Sparql Queries Over Large RDF Graphs
abstract
With the increasing use of RDF graphs, storing and querying such data using SPARQL remains a critical problem. Current mainstream solutions rely on cloud-based data management architectures, but often suffer from performance bottlenecks in environments with limited bandwidth or high system load. To address this issue, this paper explores for the first time the integration of edge computing to move graph data storage and processing to edge environments, thereby improving query performance. This approach requires offloading query processing to edge servers, which involves addressing two challenges: data localization and network scheduling. First, the data localization challenge lies in computing the subgraphs maintained on edge servers to quickly identify the servers that can handle specific queries. To address this challenge, we introduce a new concept of pattern-induced subgraphs. Second, the network scheduling challenge involves efficiently assigning queries to edge and cloud servers to optimize overall system performance. We tackle this by constructing a overall system model that jointly captures data distribution, query characteristics, network communication, and computational resources. Accordingly, we further propose a joint formulation of query assignment and computational resource allocation, modeling it as a Mixed Integer Nonlinear Programming (MINLP) problem and solve this problem using a modified branch-and-bound algorithm. Experimental results on real datasets under a real cloud platform demonstrate that our proposed method outperforms the state-of-the-art baseline methods in terms of efficiency. The codes are available on GitHub
Shidan Ma, Peng Peng 0001, Xu Zhou 0001, M. Tamer Özsu, Lei Zou 0001, Guo Chen 0001
ICDE6
2026 CAA: Toward Camouflaged and Transferable Adversarial Examples
abstract
Transferable adversarial examples (AEs) are visually indistinguishable from benign images, but can successfully mislead unknown deep neural networks. However, existing AEs normally vary considerably from benign images in the feature space, making them hard to pass label checking and adversarial detection. Therefore, how to make AEs camouflaged, disguising as benign images during detection is still an open problem. In this paper, we propose a novel camouflaged adversarial attack (CAA), which produces camouflaged adversarial examples (CAEs) for the first time. Our main idea is to make CAEs’ adversarial properties keep “dormant” state until the target model inadvertently triggers the “activated” state. To this end, we craftattackandcamouflageperturbations, so that CAEs are visually and feature/label-wise indistinguishable from benign images at first, but will implicitly turn into AEs once being triggered. Specifically, we exploit two common preprocessing operations, image scaling and JPEG compression, as the trigger, and propose a two-stage optimization strategy. As the preprocessing details of target models are unknown, the first stage trains a well-designed generative adversarial network under varying scaling/compression parameters to enhance the robustness of attack perturbations. The second stage uses feature (dis)similarities and contrastive distances to improve the transferability of camouflage perturbations. Extensive experiments on ImageNet dataset validate the effectiveness of CAA. Especially for robust models, the average fooling rate after preprocessing could reach 96.3% outperforming the state-of-the-art adversarial attack by 13.5%.
Yipeng Zou, Qin Liu 0001, Jie Wu 0001, Tian Wang 0001, Guo Chen 0001, Tao Peng 0011, Guojun Wang 0001
IEEE Trans. Inf. Forensics Secur.5
2026 A Multi-Granularity Self-Guiding Graph Diffusion Model for Predicting Private Car Activity
Bo Liu 0104, Tong Li 0013, Zhu Xiao, Yong Qiang Hei, Jingtao Ding, Guo Chen 0001
IEEE Trans. Intell. Transp. Syst.7
2026 Advancing RDMA Scalability With High Performance
abstract
Due to its superior performance, Remote Direct Memory Access (RDMA) has been widely deployed in data center networks. It provides applications with ultra-high throughput, ultra-low latency, and far lower CPU utilization than TCP/IP software network stack. However, the connection states that must be stored on the RDMA NIC (RNIC) and the small NIC memory result in poor scalability. The performance drops significantly when the RNIC needs to maintain a large number of concurrent connections. We propose StaR (Stateless RDMA), which solves the scalability problem of RDMA by transferring states to the other communication end in a trusted network. Leveraging the asymmetric communication pattern in data center applications, StaRlets the communication end with low NIC memory usage to save states for the other end with high NIC memory usage, thus making the RNIC on the bottleneck side stateless. We implemented StaR on an FPGA board with a 10Gbps network port and NS-3, evaluating its performance on a testbed with 9 machines, each equipped with StaR NICs, and verified its scalability stability by conducting a larger-scale simulation with 200 fully connected nodes using a 100Gbps link. The experimental results show that in high concurrency scenarios, the throughput of StaR can reach up to 4.13x and 1.35x of the original RNIC and the latest software-based solution, respectively.
Xijin Yin, Guo Chen 0001, Xizheng Wang, Huichen Dai, Bojie Li, Binzhang Fu, Kun Tan 0002
IEEE Trans. Netw.2
2026 Accelerating Hardware/Software Combined Traffic Processing With Fast and Efficient Asynchronous Flow Offloading
abstract
eHardware/software (hw/sw) combined systems are necessary to meet modern clouds’ requirements for processing huge amounts of network traffic by efficiently offloading large flows to hardware. However, existing hw/sw flow offloading systems typically perform traffic statistics collection and large flow selection within a time window—a time-window-based approach. Their offloading decision of large flows issynchronizedin the unit of a time window, which is mismatched to the asynchronous and dynamic nature of each flow’s sending rate. Additionally, the flow measurement and selection for large flows are decoupled in these solutions, leading to memory and CPU inefficiency. In this paper, we introduce TAO, a novel solution to the hw/sw combined flow offloading problem byasynchronouslyselecting and offloading flows based on flow table entries. TAO can reactfasterto the rapid dynamics of flows by taking actions at each table entry and ismore efficientby coupling flow measurement and selection into the entry. We have implemented a full-fledged TAO prototype based on the P4 switch and DPDK. Testbed results demonstrate that TAO can offload ∼16% more traffic to hardware, outperforming existing solutions by achieving 42× lower memory overhead. Meanwhile, it reduces software CPU utilization by 66.7% and cuts tail forwarding latency by 95.59% compared to state-of-the-art methods.
Xijin Yin, Yuanwei Lu, Xin Zhang 0117, Xingtong Lin, Shengli Zheng, Bangwen Deng, Xianneng Zou, Yachen Wang, Guo Chen 0001
IEEE Trans. Netw.9
2025 UCM: Fast and Maintainable User-space RDMA Connection Setup
Huijun Shen, Zelong Yue, Xingyu Guo, Xijin Yin, Lang An, Jianxi Ye, Guo Chen 0001
APNet12
2025 PIAR: Path-Improved Adaptive Routing for Dragonfly Networks
abstract
For the next-generation exascale supercomputing communication systems, Dragonfly topology offers strong scalability, low latency, and cost efficiency. Dragonfly networks have already been implemented in current supercomputers and will continue to expand in future systems. Adaptive routing in Dragonfly topologies is critical for network performance. The traditional UGAL routing algorithm, which uses the valiant mechanism to select non-minimal paths, does not adequately consider the impact of high hops in non-minimal paths, often unnecessarily increasing the average path length, thereby increasing network latency and load. Furthermore, UGAL inaccurately estimates the congestion of the entire routing path based on local information, leading to suboptimal routing decisions that limit the algorithm's performance. In this paper, we propose PIAR, a novel pathimproved adaptive routing algorithm. PIAR dynamically selects paths based on the status of local and global channels, prioritizing non-minimal paths with fewer hops to reduce network latency and load, thereby improving network performance. Additionally, we present the microarchitecture of the routing computation unit. Our evaluation results demonstrate that, compared with advanced algorithms such as PAR$_{\text {PH }}$, TPR, and UGAL LE, PIAR achieves an average throughput improvement of 19.2 % and reduces latency by up to$\mathbf{1 3. 4 \%}$under the single synthetic traffic. Under mixed traffic, PIAR achieves an average throughput improvement of$\mathbf{2 3. 6 \%}$and reduces the latency by up to$\mathbf{3 3. 8 \%}$. For application workloads, PIAR achieves an average reduction of 24.0 % in packet latency.
Qiang Wang 0006, Jinbo Xu, Guo Chen 0001
CLUSTER7
2025 Fast and Scalable Selective Retransmission for RDMA
Peihao Huang, Guo Chen 0001, Xin Zhang 0117, Huijun Shen, Ying Bian, Yuanwei Lu, Zhenyuan Ruan, Bojie Li, Jiansong Zhang 0001, Yongfeng Liu, Zhigang Chen 0001
INFOCOM2
2024 LEFT: LightwEight and FasT packet Reordering for RDMA
abstract
RDMA, as a cutting-edge networking technology, has gained extensive adoption in large-scale data centers due to its exceptional characteristics, such as low and stable latency, high throughput and low CPU utilization. However, due to the limited on-chip memory of the RDMA Network Interface Cards (RNIC), commercial RDMA usually only supports single-path transmission and cannot fully utilize the rich parallel paths within the DCN, resulting in insufficient bandwidth utilization. Multipath transmission can improve bandwidth utilization, but the out-of-order (OoO) packets it brings negatively impacts the performance of RNICs. Recent works have attempted to address this issue by using bitmaps to record OoO packets to better support multipath transmission. However, these approaches either consume excessive memory for maintaining bitmaps, leading to poor connection scalability, or introduce high latency in bitmap sharing. Consequently, implementing efficient packet reordering in RDMA remains a challenge.
Peihao Huang, Xin Zhang 0117, Zhigang Chen 0001, Guo Chen 0001
APNet5
2024 ActiveDNS: Is There Room for DNS Optimization Beyond CDNs?
abstract
Domain Name System (DNS) converts domain names into IP addresses. The specific IP it returns to a client has significant implications for optimizing user experiences on web services. The establishment of Content Distribution Networks (CDNs) facilitates the spread of content across diverse cache servers, thereby enabling quick responses to user requests. But not all distributed Internet services have sufficient budget or resources to deploy CDN servers. For these services, is there a cost-efficient way to enhance users’ access experience to them? In this paper, we build a data-driven DNS resolver system named ActiveDNS. We conduct a historical analysis of billions of DNS request-response logs from a public DNS resolver. Inspired by these findings, we probed dozens of public DNS servers to obtain available IPs behind carefully selected target domain names and measure their performance metrics. We built a domain ip performance database of 2.9 million records that can be easily incorporated into the open-source DNS framework BIND via DLZ technology. ActiveDNS has been deployed in real-world scenarios and the evaluations reveal promising results. Compared to currently used DNS resolver, the 75th service round-trip time (RTT) for the IP returned by ActiveDNS decreased from 67 ms to 43 ms.
Changhua Pei, Yingqiang Wang, Guo Chen 0001, Yuchao Zhang 0004, Gaogang Xie
LCN5
2024 Slim and Fast: Low-Overhead Container Overlay Network With Fast Connection Setup
abstract
Large-scale cloud applications today are often deployed using multiple containers, and a container overlay network is the de facto method to provide connectivity among these containers. However, the existing tunneling-based overlay network incurs significant performance overhead due to the need of transformation for every packet. Recent workSlim, through manipulating connection-level meta-data, allows containers to use host OS sockets directly thus they can achieve good performance without extra packet tunneling. Nevertheless, the connection setup is significantly slowed down, which requires an extra round-trip communication between both sides to pass the mapping information of the host OS socket and the container socket. This greatly hurts the performance of many cloud applications that must process short connections at high speed. We proposeSlimFast, a low-overhead container overlay network which provides a fast connection setup.SlimFastdirectly uses the host OS socket for container communication asSlim. However,SlimFastneeds no extra communication during connection setup. We reserve a dedicated host port for the container network and use socket mapping table to locally find the right container socket during connection setup. We implementSlimFastwhich is compatible with existing container applications. Experiments show that,SlimFastcan improve the connection setup time by about 2.1x compared withSlim, meanwhile maintaining low-overhead during data transmission asSlim. This brings significant performance improvement to real applications. Particularly, testbed results show thatSlimFastimproves the throughput of Nginx proxy and Memcached by about 0.9x and 2.2x, respectively.
Fusheng Lin, Xin Zhang 0117, Guo Chen 0001, Li Chen 0008, Kenli Li 0001, Hongbo Jiang 0001
IEEE Trans. Cloud Comput.3
2023 Beyond the Content: Considering the Network for Online Video Recommendation
abstract
Online recommendation systems play critical roles in enhancing user experience by helping them find the most interesting videos from a vast amount of content. However, the existing recommendation modules and video transmission modules in the industry often operate independently, resulting in the recommendation model providing some videos that cannot be transmitted within the specified deadlines successfully. This can lead to an inferior watching experience for users and resource waste for video providers. To address this, we propose a novel framework called NetRec, which for the first time optimizes the recommendation quality by jointly considering the network transmission. We accomplish this by re-ranking the top-N videos obtained from the recommendation system and selecting the top-M (M is approximately half of N) videos that provide the maximum overall revenue, e.g., video playing time while considering the network status. The entire system comprises network measurement, video quality estimation, and multi-objective optimization modules. Real-world Internet results show that our framework can increase users’ video playing time by 20% to 160%. Furthermore, we provide several promising directions for further improving the video recommendation quality under our NetRec framework, which jointly considers the network for the recommendation.
Lihui Lang, Meiqi Hu, Changhua Pei, Guo Chen 0001
APNet4
2023 An Attention-based Label Mapping and Multi-factor Domain Adaptation Approach for ACS Prediction
abstract
Acute Coronary Syndrome (ACS), an emergent medical condition, is intricately linked to environmental factors like air pollution and meteorological conditions. Harnessing regional environmental data, such as weather metrics, can promptly forecast ACS incidence rates, enabling optimised medical resource allocation and increased patient recovery rates. However, the prediction task is rendered complex due to disparities in data collection capabilities across institutions, yielding datasets with analogous features but significant label variations, impeding the application of universal models. Challenges abound due to the heterogeneity of multi-factor data, temporal alignment disparities, and the intricacies of sparse data. To address these challenges, this paper introduces the Domain Adaptation with Multi-factor Associative Structures (DAMAS), a time-series domain adaptation approach based on multi-factor sparse associative frameworks. Augmented by an isomorphic attention-driven variable label mapping scheme and combined with multi-layer perceptrons, our approach skilfully negotiates label imbalances. This results in refined prediction precision connecting environmental factors to regional ACS incidences.
Yutao Dou, Xiongjun Zhao, Kun Xie 0001, Guo Chen 0001, Shaoliang Peng
BIBM5
2023 LDCSF: Local depth convolution-based Swim framework for classifying multi-label histopathology images
abstract
Histopathological images are the gold standard for diagnosing liver cancer. However, the accuracy of fully digital diagnosis in computational pathology needs to be improved. In this paper, in order to solve the problem of multi-label and low classification accuracy of histopathology images, we propose a locally deep convolutional Swim framework (LDCSF) to classify multi-label histopathology images. In order to be able to provide local field of view diagnostic results, we propose the LDCSF model, which consists of a Swin transformer module, a local depth convolution (LDC) module, a feature reconstruction (FR) module, and a ResNet module. The Swin transformer module reduces the amount of computation generated by the attention mechanism by limiting the attention to each window. The LDC then reconstructs the attention map and performs convolution operations in multiple channels, passing the resulting feature map to the next layer. The FR module uses the corresponding weight coefficient vectors obtained from the channels to dot product with the original feature map vector matrix to generate representative feature maps. Finally, the residual network undertakes the final classification task. As a result, the classification accuracy of LDCSF for interstitial area, necrosis, non-tumor and tumor reached 0.9460, 0.9960, 0.9808, 0.9847, respectively.
Liangrui Pan, Guo Chen 0001, Wenjuan Liu, Xuan Liu 0001, Shaoliang Peng
BIBM2
2023 ESM-NBR: fast and accurate nucleic acid-binding residue prediction via protein language model feature representation and multi-task learning
abstract
Protein-nucleic acid interactions play a very important role in a variety of biological activities. Accurate identification of nucleic acid-binding residues is a critical step in understanding the interaction mechanisms. Although many computationally based methods have been developed to predict nucleic acid-binding residues, challenges remain. In this study, a fast and accurate sequence-based method, called ESM-NBR, is proposed. In ESM-NBR, we first use the large protein language model ESM2 to extract discriminative biological properties feature representation from protein primary sequences; then, a multi-task deep learning model composed of stacked bidirectional long short-term memory (BiLSTM) and multi-layer perceptron (MLP) networks is employed to explore common and private information of DNA- and RNA-binding residues with ESM2 feature as input. Experimental results on benchmark data sets demonstrate that the prediction performance of ESM2 feature representation comprehensively outperforms evolutionary information-based hidden Markov model (HMM) features. Meanwhile, the ESM-NBR obtains the MCC values for DNA-binding residues prediction of 0.427 and 0.391 on two independent test sets, which are 18.61 and 10.45% higher than those of the second-best methods, respectively. Moreover, by completely discarding the time-cost multiple sequence alignment process, the prediction speed of ESM-NBR far exceeds that of existing methods (5.52s for a protein sequence of length 500, which is about 16 times faster than the second-fastest method). A user-friendly standalone package and the data of ESM-NBR are freely available for academic use at: https://github.com/wwzll123/ESM-NBR.
Wenwu Zeng, Dafeng Lv, Xuan Liu 0001, Guo Chen 0001, Wenjuan Liu, Shaoliang Peng
BIBM4
2023 Modeling the Training Iteration Time for Heterogeneous Distributed Deep Learning Systems
abstract
Distributed deep learning systems effectively respond to the increasing demand for large‐scale data processing in recent years. However, the significant investment in building distributed learning systems with powerful computing nodes places a huge financial burden on developers and researchers. It will be good to predict the precise benefit, i.e., how many times of speedup it can get compared with training on single machine (or a few), before actually building such big learning systems. To address this problem, this paper presents a novel performance model on training iteration time for heterogeneous distributed deep learning systems based on the characteristics of the parameter server (PS) system with bulk synchronous parallel (BSP) synchronization style. The accuracy of our performance model is demonstrated by comparing real measurement results on TensorFlow when training different neural networks with various kinds of hardware testbeds: the prediction accuracy is higher than 90% in most cases.
Yifu Zeng, Pulin Pan, Kenli Li 0001, Guo Chen 0001
Int. J. Intell. Syst.5
2023 MA-STS-Based Social Intimacy Analysis Algorithm Using Real Campus Network Data
abstract
In recent years, the widespread availability of Wi‐Fi in various settings, including universities, enterprises, and large shopping centers, has become increasingly prevalent. The user’s time and location information embedded in wireless network systems can reveal individual and group social relationships, which indirectly reflect each person’s psychological well‐being. However, due to challenges in obtaining complete data, the high complexity of related data, and the absence of suitable data analysis models, few studies have analyzed student social behavior using data from university campus networks. This paper employs real‐world data from a renowned Chinese university’s wireless campus network for in‐depth analysis and introduces a novel multiangle semantic trajectory similarity (MA‐STS) algorithm to infer the intimacy and relationship types (such as teacher‐student, friends, classmates, or romantic partners) between users. The experiments demonstrate that the proposed algorithm achieves an accuracy of over 95%.
Yifu Zeng, Xiangshu Qi, Weiping Yang, Nian Pan, Guo Chen 0001
Int. J. Intell. Syst.6
2023 Fast, Scalable and Robust Centralized Routing for Data Center Networks
abstract
This paper presents a fast and robust centralized data center network (DCN) routing solution, called . For fast routing calculation, uses centralized controllers to collect/disseminate the network’s link-states (LS), and offload the actual routing calculation onto each switch. Observing that the routing changes can be classified into a few fixed patterns in DCNs which have regular topologies, we simplify each switch’s routing calculation into a table-lookup manner, i.e., comparing LS changes with pre-installed base topology and updating routing paths according to predefined rules. As such, the routing calculation time at each switch only needs 10s of us even in a large network topology containing 10K+ switches. For efficient controller fault-tolerance, purposely uses reporter switch to ensure the LS updates successfully delivered to all affected switches. As such, can use multiple stateless controllers and little redundant traffic to tolerate failures, which incurs little overhead under normal case, and keeps 10s of ms fast routing reaction time even under complex data-/control-plane failures. We design, implement and evaluate with extensive experiments on Linux-machine controllers and white-box switches. provides$\sim$1200x and$\sim$100x shorter convergence time than current distributed protocol BGP and the state-of-the-art centralized routing solution, respectively. Furthermore, Primus maintains good routing controllability/manageability thanks to its centralized architecture, which enables us to build several advanced routing features in our testbed, including routing failure visualization and weighted-cost-multi-path routing.
Fusheng Lin, Guo Chen 0001, Guihua Zhou, Dehui Wei, Li Chen 0008, Yuanwei Lu, Andrew Qu, Hongbo Jiang 0001
IEEE/ACM Trans. Netw.3
2021 StaR: Breaking the Scalability Limit for RDMA
abstract
Due to its superior performance, Remote Direct Memory Access (RDMA) has been widely deployed in data center networks. It provides applications with ultra-high throughput, ultra-low latency, and far lower CPU utilization than TCP/IP software network stack. However, the connection states that must be stored on the RDMA NIC (RNIC) and the small NIC memory result in poor scalability. The performance drops significantly when the RNIC needs to maintain a large number of concurrent connections.We propose StaR (Stateless RDMA), which solves the scalability problem of RDMA by transferring states to the other communication end. Leveraging the asymmetric communication pattern in data center applications, StaR lets the communication end with low concurrency save states for the other end with high concurrency, thus making the RNIC on the bottleneck side to be stateless. We have implemented StaR on an FPGA board with 10Gbps network port and evaluated its performance on a testbed with 9 machines all equipped with StaR NICs. The experimental results show that in high concurrency scenarios, the throughput of StaR can reach up to 4.13x and 1.35x of the original RNIC and the latest software-based solution, respectively.
Xizheng Wang, Guo Chen 0001, Xijin Yin, Huichen Dai, Bojie Li, Binzhang Fu, Kun Tan 0002
ICNP2
2021 NFD: Using Behavior Models to Develop Cross-Platform Network Functions
abstract
NFV ecosystem is flourishing and more and more NF platforms appear, but this makes NF vendors difficult to deliver NFs rapidly to diverse platforms. We propose an NF development framework named NFD for cross-platform NF development. NFD's main idea is to decouple the functional logic from the platform logic -it provides a platform-independent language to program NFs' behavior models, and a compiler with interfaces to develop platform-specific plugins. By enabling a plugin on the compiler, various NF models would be compiled to executables integrated with the target platform. We prototype NFD, build 14 NFs, and support 6 platforms (standard Linux, OpenNetVM, GPU, SGX, DPDK, OpenNF). Our evaluation shows that NFD can save development workload for cross-platform NFs and output valid and performant NFs.
Hongyi Huang, Wenfei Wu, Yongchao He, Bangwen Deng, Ying Zhang 0022, Yongqiang Xiong, Guo Chen 0001, Yong Cui 0001, Peng Cheng 0005
INFOCOM7
2021 Primus: Fast and Robust Centralized Routing for Large-scale Data Center Networks
abstract
This paper presents a fast and robust centralized data center network (DCN) routing solution called Primus. For fast routing calculation, Primus uses centralized controller to collect/disseminates the network's link-states (LS), and offload the actual routing calculation onto each switch. Observing that the routing changes can be classified into a few fixed patterns in DCNs which have regular topologies, we simplify each switch's routing calculation into a table-lookup manner, i.e., comparing LS changes with pre-installed base topology and updating routing paths according to predefined rules. As such, the routing calculation time at each switch only needs 10s of us even in a large network topology containing 10K+ switches. For efficient controller fault-tolerance, Primus purposely uses reporter switch to ensure the LS updates successfully delivered to all affected switches. As such, Primus can use multiple stateless controllers and little redundant traffic to tolerate failures, which incurs little overhead under normal case, and keeps 10s of ms fast routing reaction time even under complex data-/control-plane failures. We design, implement and evaluate Primus with extensive experiments on Linux-machine controllers and white-box switches. Primus provides ~1200x and ~100x shorter convergence time than current distributed protocol BGP and the state-of-the-art centralized routing solution, respectively.
Guihua Zhou, Guo Chen 0001, Fusheng Lin, Dehui Wei, Jianbing Wu, Li Chen 0008, Yuanwei Lu, Andrew Qu, Hongbo Jiang 0001
INFOCOM2
2021 CQPPS: A scalable multi-path switch fabric without back pressure
abstract
Abstract Generally, in order to guarantee a good throughput and decrease the complexity of reassembling, the majority of current commodity routers take elaborate closed‐loop flow control schemes such as back pressure to prevent cell loss in the switch fabrics. As commodity router's port number grows larger and link rate becomes faster, the huge I/O pins and memory consumption makes these closed‐loop flow controls very difficult to implement engineeringly. This paper approaches the problem of building ultra‐large‐capacity router from a different angle. Crosspoint‐queued‐based Parallel Packet Switch (CQPPS), a highly scalable switch architecture with no need of any closed‐loop flow control schemes, is proposed. And the authors propose the Padded Frame plus Round‐Robin scheduling scheme for CQPPS architecture. By allowing potential cell loss in the switch fabrics and slightly higher light‐load delay, CQPPS achieves loss rate orders of magnitudes lower than back pressure schemes, and high‐load delay 10 times less than back pressure schemes. It also greatly reduces the complexity of engineering implementation.
Boyan Pan, Guo Chen 0001
IET Commun.3
2019 Towards Stateless RNIC for Data Center Networks
abstract
Because of small NIC on-chip memory, the massive connection states maintained on Remote Direct Memory Access (RDMA) NIC (RNIC) significantly limit its scalability. When the number of concurrent connections grows, RNICs have to frequently fetch connection states from host memory, leading to dramatic performance degradation. In this paper, we propose StaR, which fundamentally solves this scalability issue by making RNIC stateless. Leveraging the asymmetric communication pattern in data center applications, the StaR RNIC stores zero connection-related states by moving all the connection states to the other end. Through careful design, StaR RNICs can maintain unchanged RDMA semantics and avoid security issues even when processing traffic statelessly. Preliminary simulation results show that StaR can improve the aggregate throughput by more than 160x (stress test) and 4x (application) compared to original RNICs.
Pulin Pan, Guo Chen 0001, Xizheng Wang, Huichen Dai, Bojie Li, Binzhang Fu, Kun Tan 0002
APNet2
2019 Direct Universal Access: Making Data Center Resources Available to FPGA
Ran Shu 0001, Peng Cheng 0005, Guo Chen 0001, Yongqiang Xiong, Derek Chiou, Thomas Moscibroda
NSDI3
2019 Dynamic TCP Initial Windows and Congestion Control Schemes Through Reinforcement Learning
abstract
Despite many years of improvements to it, TCP still suffers from an unsatisfactory performance. For services dominated by short flows (e.g., web search and e-commerce), TCP suffers from the flow startup problem and cannot fully utilize the available bandwidth in the modern Internet: TCP starts from a conservative and static initial window (IW, 2-4 or 10), while most of the web flows are too short to converge to the best sending rate before the session ends. For services dominated by long flows (e.g., video streaming and file downloading), the congestion control (CC) scheme manually and statically configured might not offer the best performance for the latest network conditions. To address these two challenges, we propose TCP-RL, which uses reinforcement learning (RL) techniques to dynamically configure IW and CC in order to improve the performance of TCP flow transmission. Basing on the latest network conditions observed at the server side of a web service, TCP-RL dynamically configures a suitable IW for short flows through group-based RL, and dynamically configures a suitable CC scheme for long flows through deep RL. Our extensive experiments show that for short flows, TCP-RL can reduce the average transmission time by about 23%; and for long flows, compared with the performance of 14 CC schemes, TCP-RL's performance ranks top 5 for about 85% of the 288 given static network conditions, whereas for about 90% of conditions, its performance drops by less than 12% compared with that of the best-performing CC schemes for the same network conditions.
Xiaohui Nie, Youjian Zhao, Zhihan Li 0002, Guo Chen 0001, Kaixin Sui, Zijie Ye, Dan Pei
IEEE J. Sel. Areas Commun.4
2019 M-Skyline: Taking sunk cost and alternative recommendation in consideration for skyline query on uncertain data
Yifu Zeng, Guo Chen 0001, Kenli Li 0001, Yantao Zhou, Xu Zhou 0001, Keqin Li 0001
Knowl. Based Syst.2
2019 MP-RDMA: Enabling RDMA With Multi-Path Transport in Datacenters
abstract
RDMA is becoming prevalent because of its low latency, high throughput and low CPU overhead. However, in current datacenters, RDMA remains a single path transport which is prone to failures and falls short to utilize the rich parallel network paths. Unlike previous multi-path approaches, which mainly focus on TCP, this paper presents a multi-path transport for RDMA, i.e. MP-RDMA, which efficiently utilizes the rich network paths in datacenters. MP-RDMA employs three novel techniques to address the challenge of limited RDMA NICs on-chip memory size: 1) a multi-path ACK-clocking mechanism to distribute traffic in a congestion-aware manner without incurring per-path states; 2) an out-of-order aware path selection mechanism to control the level of out-of-order delivered packets, thus minimizes the meta data required to them; 3) a synchronise mechanism to ensure in-order memory update whenever needed. With all these techniques, MP-RDMA only adds 66B to each connection state compared to single-path RDMA. Our evaluation with an FPGA-based prototype demonstrates that compared with single-path RDMA, MP-RDMA can significantly improve the robustness under failures ( $2\times \sim 4\times $ higher throughput under 0.5%~10% link loss ratio) and improve the overall network utilization by up to 47%.
Guo Chen 0001, Yuanwei Lu, Bojie Li, Kun Tan 0002, Yongqiang Xiong, Peng Cheng 0005, Jiansong Zhang 0001, Thomas Moscibroda
IEEE/ACM Trans. Netw.1
2018 Reducing Web Latency Through Dynamically Setting TCP Initial Window with Reinforcement Learning
abstract
Latency, which directly affects the user experience and revenue of web services, is far from ideal in reality, due to the well-known TCP flow startup problem. Specifically, since TCP starts from a conservative and static initial window (IW, 2~4 or 10), most of the web flows are too short to have enough time to find its best congestion window before the session ends. As a result, TCP cannot fully utilize the available bandwidth in the modern Internet. In this paper, we propose to use group-based reinforcement learning (RL) to enable a web server, through trial-and-error, to dynamically set a suitable IW for a web flow before its transmission starts. Our proposed system, SmartIW, collects TCP flow performance metrics (e.g., transmission time, loss rate, RTT) in real-time without any client assistance. Then these metrics are aggregated into groups with similar features (subnet, ISP, province, etc.) to satisfy RL's requirement. SmartIW has been deployed in one of the top global search engines for more than a year. Our online and testbed experiments show that, compared to the common practice of IW=10, SmartIW can reduce the average transmission time by 23% to 29%.
Xiaohui Nie, Youjian Zhao, Dan Pei, Guo Chen 0001, Kaixin Sui
IWQoS4
2018 Multi-Path Transport for RDMA in Datacenters
Yuanwei Lu, Guo Chen 0001, Bojie Li, Kun Tan 0002, Yongqiang Xiong, Peng Cheng 0005, Jiansong Zhang 0001, Enhong Chen, Thomas Moscibroda
NSDI2
2018 FUSO: Fast Multi-Path Loss Recovery for Data Center Networks
Guo Chen 0001, Yuanwei Lu, Yuan Meng 0002, Bojie Li, Kun Tan 0002, Dan Pei, Peng Cheng 0005, Layong Luo, Yongqiang Xiong, Xiaoliang Wang 0001, Youjian Zhao
IEEE/ACM Trans. Netw.1
2017 Memory Efficient Loss Recovery for Hardware-based Transport in Datacenter
abstract
Limited by the small on-chip memory, hardware-based transport typically implements go-back-N loss recovery mechanism, which costs very few memory but is well-known to perform inferior even under small packet loss ratio. We present MELO, an efficient selective retransmission mechanism for hardware-based transport, which consumes only a constant small memory regardless of the number of concurrent connections. Specifically, MELO employs an architectural separation between data and meta data storage and uses a shared bits pool allocation mechanism to reduce meta data on-chip memory footprint. By only adding in average 23B extra on-chip states for each connection, MELO achieves up to 14.02x throughput while reduces 99% tail FCT by 3.11x compared with go-back-N under certain loss ratio.
Yuanwei Lu, Guo Chen 0001, Zhenyuan Ruan, Wencong Xiao, Bojie Li, Jiansong Zhang 0001, Yongqiang Xiong, Peng Cheng 0005, Enhong Chen
APNet2
2017 Network Stack as a Service in the Cloud
abstract
The tenant network stack is implemented inside the virtual machines in today's public cloud. This legacy architecture presents a barrier to protocol stack innovation due to the tight coupling between the network stack and the guest OS. In particular, it causes many deployment troubles to tenants and management and efficiency problems to the cloud provider. To address these issues, we articulate a vision of providing the network stack as a service. The central idea is to decouple the network stack from the guest OS, and offer it as an independent entity implemented by the cloud provider. This re-architecting allows tenants to readily deploy any stack independent of its kernel, and the provider to offer meaningful SLAs to tenants by gaining control over the network stack. We sketch an initial design called NetKernel to accomplish this vision. Our preliminary testbed evaluation with a prototype shows the feasibility and benefits of our idea.
Zhixiong Niu, Hong Xu 0001, Dongsu Han, Peng Cheng 0005, Yongqiang Xiong, Guo Chen 0001, Keith Winstein
HotNets6
2017 How Much Are Your Neighbors Interfering with Your WiFi Delay?
abstract
Previous studies have shown the WiFi, as the dominant last hop access to Internet, has become the weakest link in the round-trip network delay. Therefore it is critical to understand and minimize the WiFi interference in order to reduce the WiFi hop delay. For the first time in the literature, this paper defines an intuitive and accurate metric to quantify the impact of interference on each actual packet. For each packet traveling through the access point, it measures the percentage of MAC layer delay wasted due to neighbor APs' interference. This metric is defined based on a packet's various (measured or inferred) timestamps and can be measured with a small kernel modification on a commodity AP with little overhead. Our 29-AP two- month measurement results in the wild show that this metric is a strong indicator of interference's impact on WiFi hop delay. Using this metric as input, distributed channel selection on individual APs reduces the median WiFi hop delay by up to 5X. Collaborative optimization on multiple APs reduces the overall WiFi hop delay by 5X compared to the default channel.
Changhua Pei, Youjian Zhao, Guo Chen 0001, Yuan Meng 0002, Yang Liu 0442, Ya Su, Ruming Tang, Dan Pei
ICCCN3
2017 One more queue is enough: Minimizing flow completion time with explicit priority notification
abstract
Ideally, minimizing the flow completion time (FCT) requires millions of priorities supported by the underlying network so that each flow has its unique priority. However, in production datacenters, the available switch priority queues for flow scheduling are very limited (merely 2 or 3). This practical constraint seriously degrades the performance of previous approaches. In this paper, we introduce Explicit Priority Notification (EPN), a novel scheduling mechanism which emulates fine-grained priorities (i.e., desired priorities or DP) using only two switch priority queues. EPN can support various flow scheduling disciplines with or without flow size information. We have implemented EPN on commodity switches and evaluated its performance with both testbed experiments and extensive simulations. Our results show that, with flow size information, EPN achieves comparable FCT as pFabric that requires clean-slate switch hardware. And EPN also outperforms TCP by up to 60.5% if it bins the traffic into two priority queues according to flow size. In information-agnostic setting, EPN outperforms PIAS with two priority queues by up to 37.7%. To the best of our knowledge, EPN is the first system that provides millions of priorities for flow scheduling with commodity switches.
Yuanwei Lu, Guo Chen 0001, Larry Luo, Kun Tan 0002, Yongqiang Xiong, Xiaoliang Wang 0001, Enhong Chen
INFOCOM2
2017 TCP WISE: One initial congestion window is not enough
abstract
Current TCP is very inefficient for web services. Web transactions are often very short-lived. TCP flow starts with a conservative initial congestion window (IW), which causes multiple round-trip times to finish the transmission even if the end-to-end bandwidth is sufficient for the transaction to be finished in one round-trip time. Previous research efforts have been focusing on finding the overall best IW for the entire Internet or a service company. However, we observe that one-IW-fits-all is suboptimal after one year of online measurement in Baidu, one of the top global search engine companies. To reduce the TCP latency, we propose TCP WISE, which dynamically assigns suitable IWs for different user cluster at different times on the server side. The values of users' IWs are proactively learned based on the historical experience on the server-side. Our testbed experiments show that our learning algorithm can handle the network changes and converge to the best IW. We have deployed TCP WISE in one of Baidu's production data center, and results shows that the 80thpercentile latency of the HTTP responses has been reduced by 10.4% compared with current TCP with a fixed IW of 10.
Xiaohui Nie, Youjian Zhao, Guo Chen 0001, Kaixin Sui, Yazheng Chen, Dan Pei
IPCCC3
2017 𝔽2 Tree: Rapid Failure Recovery for Routing in Production Data Center Networks
abstract
Failures are not uncommon in production data center networks (DCNs) nowadays. It takes long time for the DCN routing to recover from a failure and find new forwarding paths, significantly impacting realtime and interactive applications at the upper layer. In this paper, we present a fault-tolerant DCN solution, called F2Tree, which is readily deployed in existing DNCs. F2Tree can significantly improve the failure recovery time only through a small amount of link rewiring and switch configuration changes. Through testbed and emulation experiments, we show that F2Tree can greatly reduce the routing recovery time after failure (by 78%) and improve the performance of upper layer applications when routing failure happens (96% less deadline-missing requests).
Guo Chen 0001, Youjian Zhao, Hailiang Xu, Dan Pei, Dan Li 0001
IEEE/ACM Trans. Netw.1
2016 WiFi can be the weakest link of round trip network latency in the wild
abstract
As mobile Internet is now indispensable in our daily lives, WiFi's latency performance has become critical to mobile applications' quality of experience. Unfortunately, WiFi hop latency in the wild remains largely unknown. In this paper, we first propose an effective approach to break down the round trip network latency. Then we provide the first systematic study on WiFi hop latency in the wild based on the latency and WiFi factors collected from 47 APs on T university campus for two months. We observe that WiFi hop can be the weakest link in the round trip network latency: more than 50% (10%) of TCP packets suffer from WiFi hop latency larger than 20ms (100ms), and WiFi hop latency occupies more than 60% in more than half of the round trip network latency. To help understand, troubleshoot, and optimize WiFi hop latency for WiFi APs in general, we train a decision tree model. Based on the model's output, we are able to reduce the median latency by 80% from 50ms to 10ms in one real case, and reduce the maximum latency from 250ms to 50ms in another real case.
Changhua Pei, Youjian Zhao, Guo Chen 0001, Ruming Tang, Yuan Meng 0002, Minghua Ma, Ken Ling, Dan Pei
INFOCOM3
2016 Fast and Cautious: Leveraging Multi-path Diversity for Transport Loss Recovery in Data Centers
Guo Chen 0001, Yuanwei Lu, Yuan Meng 0002, Bojie Li, Kun Tan 0002, Dan Pei, Peng Cheng 0005, Layong Luo, Yongqiang Xiong, Xiaoliang Wang 0001, Youjian Zhao
USENIX ATC1
2015 Rewiring 2 Links Is Enough: Accelerating Failure Recovery in Production Data Center Networks
abstract
Failures are not uncommon in production data center networks (DCNs) nowadays, and it takes long time for the network to recover from a failure and find new forwarding paths, significantly impacting real time and interactive applications at the upper layer. The slow failure recovery is due to two primary reasons. First, there lacks immediate backup paths for downward links in DCN with multi-rooted tree topology. Second, distributed routing protocols in DCN take time to converge after failures. In this paper, we present a fault-tolerant DCN solution, called F2Tree, that can significantly improve the failure recovery time in current DCNs, only through a small amount of link rewiring and switch configuration changes. Because F2Tree does not change any existing software or hardware, it is readily deployed in production DCNs, where other existing proposals fail to achieve. Through testbed and emulation experiments, we show that F2Tree can greatly reduce the time of failure recovery by 78%. Our experimental results also show that, for partition-aggregate applications (popular in DCN) under various failure conditions, F2Tree reduces the ratio of deadline-missing requests by more than 96% compared to current DCNs.
Guo Chen 0001, Youjian Zhao, Dan Pei, Dan Li 0001
ICDCS1
2015 Alleviating flow interference in data center networks through fine-grained switch queue management
Guo Chen 0001, Youjian Zhao, Dan Pei
Comput. Networks1
2014 CQRD: A switch-based approach to flow interference in Data Center Networks
abstract
Modern data centers need to satisfy stringent low-latency for real-time interactive applications (e.g. search, web retail). However, short delay-sensitive flows often have to wait a long time for memory and link resource occupied by a few of long bandwidth-greedy flows because they share the same switch output queue (OQ). To address the above flow interference problem, this paper advocates more fine-grained flow separation in the switches than traditional OQ. We propose CQRD, a simple and cost-effective queue management scheme for data center switches, through only minor changes to the buffering and scheduling scheme. No change to the transport layer or coordination among switches is required. Simulation results show that CQRD can reduce the FCT of short flows by 20–44% in a single switch and 8–30% in a multi-stage data center switch network, only at the cost of a minor goodput decrease of large of flows.
Guo Chen 0001, Dan Pei, Youjian Zhao
LCN1
2014 Designing Buffer Capacity of Crosspoint-Queued Switch
Guo Chen 0001, Dan Pei, Youjian Zhao, Yongqian Sun
NPC1