VLDB 2026 Research / reviewers in the wild / expert
Guyue Liu
dblp:151/5134 · also Guyue (Grace) Liu
· DBLP profile ↗
47ranked-venue papers
5as first author
33since 2021 · last 2026
0009-0001-4933-0276ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 25 · 3 first-author · 21 since 2021Systems, architecture and hardware · 9 · 7 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 2 since 2021Security and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Practical and Scalable RDMA Connection Sharing for HPC WorkloadabstractRDMA is a fundamental communication infrastructure in high-performance computing (HPC). However, as the number of RDMA connections increases, system performance rapidly declines and memory consumption increases sharply. Previous research demonstrates that sharing RDMA connections among processes is necessary and effective to address the scalability problem. Unfortunately, previous work shares connections in software, thus incurring substantial overhead to each packet operation, and fails to comprehensively explore control policies to achieve superior sharing decisions. Yuejie Wang, Tuo Fang, Biyu Peng, Xin Sun 0027, Chengchao Xu, Yuxin Ren 0001, Ning Jia 0004, Xinwei Hu, Yunfei Du 0001, Guyue Liu |
EuroSys | 12 |
| 2026 | Limitless Scalability: A High-Throughput and Replica-Agnostic BFT Consensus
Chenyu Zhang 0008, Xiulong Liu 0001, Hao Xu 0025, Haochen Ren, Muhammad Shahzad 0001, Guyue Liu, Keqiu Li |
NDSS | 6 |
| 2026 | CascadeNet: Generating Network Traffic with High-Fidelity Temporal Patterns
Runwei Lu, Yanran Deng, Ruixuan Li 0016, Jinting Liu, Yuejie Wang, Deming Xu, Han Tian, Kai Chen 0005, Guyue Liu |
NSDI | 10 |
| 2026 | MirrorNet: High-fidelity and Scalable Network Emulation for Software-defined WAN
Congcong Miao, Yuejie Wang, Xuefeng Ji, Guozhi Shan, Pan Fang, Yanke Zhang, Xianneng Zou, Guyue Liu |
NSDI | 11 |
| 2026 | Achieving Network Efficiency Through Service Collaborative Capacity Sharing and EnforcementabstractMeta's rapid expansion in users, business operations, and AI workloads is straining our backbone network, while physical constraints—such as fiber, space, and power—limit the speed of capacity growth. To address these challenges, we present a service-aware network capacity planning suite that systematically improves network efficiency with a service collaboration approach. We propose the "safe capacity" abstraction which enables services to incorporate current and projected network conditions into their compute and storage allocation decisions. We introduce a hose-carving method that efficiently translates service-level traffic demands into detailed traffic matrices, allowing for more precise bandwidth allocation. To promote responsible network usage, we design a network rate card which attributes network consumption to individual services, incentivizing optimization and resource trade-offs. Additionally, new enforcement features at the end-host layer dynamically adjust resource allocations and traffic flows at runtime to maximize utilization. This paper is the first to detail a collaborative, service-aware approach to backbone network efficiency at Meta scale. Based on years of operational experience, we share practical insights and highlight new directions for research in network efficiency. Vinayak Dangui, Alaleh Razmjoo, Guanqing Yan, Mahesh Nayak, Mansi Babbar, Brian Bierig, Tejas Birajdar, Prabhakaran Ganesan, Lilian Liu, Matt Maia, Jerry Yang, Shrinivas Petale, Satyajeet Ahuja, Abhinav Triguna, Guyue Liu, Ying Zhang 0022 |
SIGCOMM | 15 |
| 2026 | Towards High-Performance Intrusion Detection with Robustness Guarantees on Programmable Switches at ISP ScaleabstractIn order to provide security connections to the enterprise campus sites, internet service providers are offering comprehensive intrusion detection services at the network layer. However, existing network intrusion detection systems (NIDS) are either ineffective or inefficient for high-speed network protection, especially for encrypted traffic analysis. In this paper, we design and implement SiteGuard, an inline network intrusion detection system with programmable switches specifically developed to protect enterprise campus sites connecting to ISP. SiteGuard proposes a dual-plane feature extraction model to extract extensive traffic features at near line-speed. SiteGuard also proposes a lightweight one-class classification model that trains the best parameters exclusively on benign traffic to identify malicious traffic. In addition, SiteGuard introduces an online update mechanism that aims to dynamically adjust the detection model in response to environmental changes. SiteGuard has been in production for more than three years. Our production and testbed evaluations demonstrate SiteGuard can detect malicious traffic with approximately 90% accuracy in minutes. Han Zhang 0009, Linqiang Qian, Guyue Liu, Kaiyang Zhao 0004, Yantu Tong, Zeji Xiao, Dongbiao He, Ke Ruan, Jilong Wang 0001, Xia Yin 0001 |
SIGCOMM | 4 |
| 2025 | Squeezing Operator Performance Potential for the Ascend ArchitectureabstractWith the rise of deep learning, many companies have developed domain-specific architectures (DSAs) optimized for AI workloads, with Ascend being a representative. To fully realize the operator performance on Ascend, effective analysis and optimization is urgently needed. Compared to GPU, Ascend requires users to manage operations manually, leading to complex performance issues that require precise analysis. However, existing roofline models face challenges of visualization complexity and inaccurate performance assessment. To address these needs, we introduce a component-based roofline model that abstracts components to capture operator performance, thereby effectively identifying bottleneck components. Furthermore, through practical operator optimization case studies, we illustrate a comprehensive process of optimization based on roofline analysis, summarizing common performance issues and optimization strategies. Finally, extensive end-to-end optimization experiments demonstrate significant model speed improvements, ranging from 1.07× to 2.15×, along with valuable insights from practice. Zhibin Wang 0002, Guyue Liu, Yongzhong Wang, Fuchun Wei, Zhiheng Hu, Yanlin Liu, Yaoyuan Wang, Wan-Chun Dou, Guihai Chen, Chen Tian 0001 |
ASPLOS (2) | 3 |
| 2025 | Orcas: A DAG-based Consensus Approach with Linear Communication OverheadabstractTo enable parallel transaction processing in blockchain systems, recent consensus protocols have adopted directed acyclic graph (DAG) structures where DAG is used to organize and parallelize the blocks. Unfortunately, these protocols suffer from high communication overhead. Our experiment on the state-of-the-art Graded DAG[12] reveals that dissemination of transaction and consensus vote messages account for the majority of network traffic. We analyze that the overall overhead is O (N2) per replica and O (N3) for the entire system, where N is the number of replicas, and note that existing approaches have not succeeded in reducing this overhead. Xiulong Liu 0001, Hao Xu 0025, Chenyu Zhang 0008, Gaowei Shi, Keqiu Li, Muhammad Shahzad 0001, Guyue Liu |
SoCC | 8 |
| 2025 | Fork: A Dual Congestion Control Loop for Small and Large Flows in DatacentersabstractMany existing transport designs aim to deliver ultra-low latency and high bandwidth for applications in high-speed datacenter networks. However, almost all of them intertwine the control of small and large flows using the same control entity (e.g., sender or receiver) and congestion feedback signal (e.g., ECN or credit), thus bringing significant performance impairments. By contrast, we seek to decouple the rate control of small flows from that of large ones. Wenxin Li 0001, Yulong Li 0001, Lide Suo, Xuan Gao 0001, Xin Xie 0001, Sheng Chen 0015, Ziqi Fan, Wenyu Qu, Guyue Liu |
EuroSys | 10 |
| 2025 | eNetSTL: Towards an In-kernel Library for High-Performance eBPF-based Network FunctionsabstractUsing extended Berkeley Packet Filter (eBPF) to implement networking functions (NFs) has been a promising trend for modern network infrastructure. In this paper, we endeavor to implement 35 representative NFs with eBPF, but encounter inherent problems of either incomplete functionality or performance degradation of up to 49.2%. Conventional solutions like modifying the eBPF infrastructure or implementing functions directly in the kernel can lead to intrusive and unstable modifications. Bin Yang 0027, Dian Shen, Junxue Zhang 0001, Lunqi Zhao, Beilun Wang, Guyue Liu, Kai Chen 0005 |
EuroSys | 7 |
| 2025 | A Generic and Efficient Communication Framework for Message-Level In-Network Computing
Xinchen Wan, Han Tian, Xudong Liao, Chaoliang Zeng, Zilong Wang 0007, Qingsong Ning, Guyue Liu, Layong Luo, Kai Chen 0005 |
INFOCOM | 11 |
| 2025 | Heimdall: Towards Risk-Aware Network Management Outsourcing
Yuejie Wang, Qiutong Men, Yongting Chen, Jiajin Liu, Gengyu Chen, Ying Zhang 0022, Guyue Liu, Vyas Sekar |
NDSS | 7 |
| 2025 | Ladder: A Convergence-based Structured DAG Blockchain for High Throughput and Low Latency
Dengcheng Hu, Jianrong Wang, Xiulong Liu 0001, Hao Xu 0025, Xujing Wu, Muhammad Shahzad 0001, Guyue Liu, Keqiu Li |
NSDI | 7 |
| 2025 | Enabling Efficient GPU Communication over Multiple NICs with FuseLink
Zhenghang Ren, Zilong Wang 0007, Wenxue Li 0004, Kaiqiang Xu, Xudong Liao, Yijun Sun, Bowen Liu 0002, Han Tian, Junxue Zhang 0001, Mingfei Wang, Zhizhen Zhong, Guyue Liu, Ying Zhang 0022, Kai Chen 0005 |
OSDI | 14 |
| 2025 | Achieving High-Speed and Robust Encrypted Traffic Anomaly Detection with Programmable SwitchesabstractAttacks against data centers are becoming more common as a result of the fast expansion of applications. In order to keep pace with the growing amount of data centers connected to their networks, internet service providers must offer comprehensive security services. However, existing network intrusion detection systems (NIDS) are either ineffective or inefficient for the high-speed encrypted network traffic. In this paper, we design and implement Mazu, an inline network intrusion detection system with programmable switches specifically developed to protect data centers connecting to the internet service provider. Mazu proposes a dual-plane feature extraction model to extract extensive traffic features at near line-speed. Mazu also proposes a lightweight one-class classification model that trains the best parameters exclusively on benign traffic to identify the malicious traffic. In addition, Mazu introduces an online update mechanism aimed at dynamically adjusting the detection model in response to environmental changes. Mazu has been in production for two years, during which time it has identified over 10 critical attack events and protect more than 10 million servers for two ISPs. Our production and testbed evaluations demonstrate that Mazu can detect malicious traffic entering the data center sites with approximately 90% accuracy within minutes. Han Zhang 0009, Guyue Liu, Xingang Shi, Dongbiao He, Jilong Wang 0001, Ke Ruan, Xia Yin 0001 |
SIGCOMM | 2 |
| 2025 | DNSLogzip: A Novel Approach to Fast and High-Ratio Compression for DNS LogsabstractDomain Name System (DNS) logs capture detailed records of the queries and responses exchanged between DNS servers and clients, playing a crucial role in applications such as cybersecurity monitoring and regulatory compliance, which often require long-term data retention. With the rapid growth of Internet traffic, the volume of DNS logs has surged, presenting significant storage challenges. Although many DNS operators use general-purpose compression algorithms to reduce storage costs, these solutions fail to fully exploit the unique characteristics of DNS data, leading to inefficiencies and rising storage demands. Yunwei Dai, Guyue Liu, Tao Huang 0005, Shuo Wang 0006, Xingli Wu, Heshun Li, Fanglong Hu |
SIGCOMM | 2 |
| 2025 | MixNet: A Runtime Reconfigurable Optical-Electrical Fabric for Distributed Mixture-of-Experts TrainingabstractMixture-of-Expert (MoE) models outperform conventional models by selectively activating different subnets, named experts, on a per-token basis. This gated computation generates dynamic communications that cannot be determined beforehand, challenging the existing GPU interconnects that remain static during distributed training. In this paper, we advocate for a first-of-its-kind system, called MixNet, that unlocks topology reconfiguration during distributed MoE training. Towards this vision, we first perform a production measurement study and show that the MoE dynamic communication pattern has strong locality, alleviating the need for global reconfiguration. Based on this, we design and implement a regionally reconfigurable high-bandwidth domain that augments existing electrical interconnects using optical circuit switching (OCS), achieving scalability while maintaining rapid adaptability. We build a fully functional MixNet prototype with commodity hardware and a customized collective communication runtime. Our prototype trains state-of-the-art MoE models with in-training topology reconfiguration across 32 A100 GPUs. Large-scale packet-level simulations show that MixNet achieves performance comparable to a non-blocking fat-tree fabric while boosting the networking cost efficiency (e.g., performance per dollar) of four representative MoE models by 1.2×–1.5× and 1.9×–2.3× at 100 Gbps and 400 Gbps link bandwidths, respectively. Xudong Liao, Yijun Sun, Han Tian, Xinchen Wan, Yilun Jin, Zilong Wang 0007, Zhenghang Ren, Wenxue Li 0004, Kin Fai Tse, Zhizhen Zhong, Guyue Liu, Ying Zhang 0022, Xiaofeng Ye, Yiming Zhang 0003, Kai Chen 0005 |
SIGCOMM | 12 |
| 2025 | InfiniteHBD: Building Datacenter-Scale High-Bandwidth Domain for LLM with Optical Circuit Switching TransceiversabstractScaling Large Language Model (LLM) training relies on multidimensional parallelism, where High-Bandwidth Domains (HBDs) are critical for communication-intensive parallelism like Tensor Parallelism. However, existing HBD architectures face fundamental limitations in scalability, cost, and fault resiliency: switch-centric HBDs (e.g., NVL-72) incur prohibitive scaling costs, while GPU-centric HBDs (e.g., TPUv3/Dojo) suffer from severe fault propagation. Switch-GPU hybrid HBDs (e.g., TPUv4) take a middle-ground approach, but the fault explosion radius remains large. Chenchen Shou, Guyue Liu, Hao Nie, Huaiyu Meng, Yu Zhou 0008, Wenqing Lv, Yelong Xu, Yuanwei Lu, Yanbo Yu, Yichen Shen 0001, Yibo Zhu 0001, Daxin Jiang |
SIGCOMM | 2 |
| 2025 | Towards LLM-Based Failure Localization in Production-Scale NetworksabstractRoot causing and failure localization are critical to maintain reliability in cloud network operations. When an incident is reported, network operators must review massive volumes of monitoring data and identify the root cause (i.e., error device) as fast as possible, making it extremely challenging even for experienced operators. Large language models (LLMs) have shown great potential in text understanding and reasoning. In this paper, we present BiAn, an LLM-based framework designed to assist operators in efficient incident investigation. BiAn processes monitoring data and generates error device rankings with detailed explanations. To date, BiAn has been deployed in our network infrastructure for 10 months and it has successfully assisted operators in identifying error devices more quickly, reducing time to root causing by 20.5% (55.2% for high-risk incidents). Extensive performance evaluations based on 17 months of real cases further demonstrate that BiAn achieves accurate and fast failure localization. It improves accuracy by 9.2% compared to the baseline approach. Chenxu Wang 0007, Xumiao Zhang, Runwei Lu, Xianshang Lin, Xuan Zeng 0002, Zhe An, Gongwei Wu, Chen Tian 0001, Guihai Chen, Guyue Liu, Yuhong Liao, Dennis Cai, Ennan Zhai |
SIGCOMM | 12 |
| 2025 | Towards Optimal Rack-scale μs-level CPU Scheduling through In-Network Workload Shaping
Xudong Liao, Han Tian, Xinchen Wan, Chaoliang Zeng, Hao Wang 0116, Junxue Zhang 0001, Mengyu Ma, Guyue Liu, Kai Chen 0005 |
USENIX ATC | 8 |
| 2025 | Accelerating Model Training on Ascend Chips: An Industrial System for Profiling, Analysis and Optimization
Zhibin Wang 0002, Ruyi Zhang 0005, Chen Tian 0001, Xiaoliang Wang 0001, Wan-Chun Dou, Guihai Chen, Bingqiang Wang, Yonghong Tian 0001, Yan Zhang 0002, Hui Wang 0030, Fuchun Wei, Boquan Sun, Bin She, Teng Su, Yaoyuan Wang, Guyue Liu |
USENIX ATC | 23 |
| 2025 | Using a multi-strain infectious disease model with physical information neural networks to study the time dependence of SARS-CoV-2 variants of concernabstractWith the ongoing evolution of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) and its increasing adaptation to humans, several variants of concern (VOCs) and variants of interest (VOIs) have been identified since late 2020. These include Alpha, Beta, Gamma, Delta, Omicron parent lineage, and other variants. These variants may show distinct levels of virulence, antigenicity, and infectivity, which require specific defense and control measures. In this study, we propose an [Formula: see text] infectious disease model to simulate the spread of SARS-CoV-2 variants among the human population. We combine the proposed epidemic model and reported infected data of variants with physical information neural networks (PINNs) to develop a novel mechanism called VOCs-informed neural network (VOCs-INN). In our experiments, we found that this algorithm can accurately fit the reported data of the British Columbia (BC) province and its five internal health agencies in Canada. Furthermore, it can simulate observed or unobserved dynamics, infer time-dependent parameters, and enable short-term predictions. The experimental results also reveal variations in the intensity of control strategies implemented across these regions. VOCs-INN performs well in fitting and forecasting when analyzing long-term or multi-wave data. Xu Chen 0049, Suli Liu, Guyue Liu |
PLoS Comput. Biol. | 5 |
| 2024 | Understanding Communication Characteristics of Distributed TrainingabstractCommunication is pivotal in distributed training and a thorough understanding of its characteristics is essential for future optimizations. However, prior works are limited, either focusing on customized optimizations or conducting incomplete explorations on communication characteristics. In this work, we systematically analyze the communication characteristics of distributed training, considering two key aspects of communication: pattern and overhead, and assessing a broad spectrum of determinant factors. In particular, we extensively investigate the features of communication patterns, such as predictability, and comprehensively evaluate the impact of various factors on communication overhead. Additionally, we develop and validate an analytical formulation to estimate communication overhead, providing a mathematical understanding of models with predictability. Wenxue Li 0004, Xiangzhou Liu, Yilun Jin, Han Tian, Zhizhen Zhong, Guyue Liu, Ying Zhang 0022, Kai Chen 0005 |
APNet | 7 |
| 2024 | Flow Scheduling with Imprecise Knowledge
Wenxin Li 0001, Xin He 0043, Keqiu Li, Kai Chen 0005, Zhao Ge, Zewei Guan, Heng Qi, Song Zhang 0008, Guyue Liu |
NSDI | 10 |
| 2024 | PPT: A Pragmatic Transport for DatacentersabstractThis paper introduces PPT, a pragmatic transport that achieves comparable performance to proactive transports while maintaining good deployability as reactive transports. Our key idea is to run a low-priority control loop to leverage the available bandwidth left by the reactive transports. The main challenge is to send just enough packets to improve performance without harming the primary control loop. We combine two unconventional techniques: an intermittent loop initialization and an exponential window decrease, enabling us to dynamically identify and fill the spare bandwidth. We further complement PPT's design with a buffer-aware flow scheduling scheme to optimize the average FCT of small flows without prior knowledge of flow size information. We have implemented a PPT prototype in the Linux kernel with ~400 lines of code and demonstrated that compared to Homa, it delivers up to 46.3% lower overall average FCT and even 25%/55.5% lower average/tail FCT of small flows in an Memcached workload. Lide Suo, Yiren Pang, Wenxin Li 0001, Renjie Pei, Keqiu Li, Xiulong Liu 0001, Xin He 0043, Yitao Hu, Guyue Liu |
SIGCOMM | 9 |
| 2024 | ConfMask: Enabling Privacy-Preserving Configuration Sharing via AnonymizationabstractReal-world network configurations play a critical role in network management and research tasks. While valuable, data holders often hesitate to share them due to business and privacy concerns. Existing methods are deficient in concealing the implicit information that can be inferred from configurations, such as topology and routing paths. To address this, we present ConfMask, a novel framework designed to systematically anonymize network topology and routing paths in configurations. Our approach tackles key privacy, utility, and scalability challenges, which arise from the strong dependency between different datasets and complex routing protocols. Our anonymization algorithm is scalable to large networks and effectively mitigates de-anonymization risk. Moreover, it maintains essential network properties such as reachability, waypointing and multi-path consistency, making it suitable for a wide range of downstream tasks. Compared to existing dataplane anonymization algorithm (i.e., NetHide), ConfMask reduces ~75% specification differences between the original and the anonymized networks. Yuejie Wang, Qiutong Men, Yongting Chen, Guyue Liu |
SIGCOMM | 5 |
| 2024 | Inversion impact of approximate PIFO to Start-Time Fair Queueing
Junda Song, Jiajin Liu, Peixuan Gao, Guyue Liu, H. Jonathan Chao |
Comput. Networks | 4 |
| 2023 | LemonNFV: Consolidating Heterogeneous Network Functions at Line Speed
Hao Li 0011, Yihan Dang, Guangda Sun, Guyue Liu, Danfeng Shan, Peng Zhang 0011 |
NSDI | 4 |
| 2023 | ClickINC: In-network Computing as a Service in Heterogeneous Programmable Data-center NetworksabstractIn-Network Computing (INC) has found many applications for performance boosts or cost reduction. However, given heterogeneous devices, diverse applications, and multi-path network typologies, it is cumbersome and error-prone for application developers to effectively utilize the available network resources and gain predictable benefits without impeding normal network functions. Previous work is oriented to network operators more than application developers. We develop ClickINC to streamline the INC programming and deployment using a unified and automated workflow. Click-INC provides INC developers a modular programming abstractions, without concerning to the states of the devices and the network topology. We describe the ClickINC framework, model, language, workflow, and corresponding algorithms. Experiments on both an emulator and a prototype system demonstrate its feasibility and benefits. Wenquan Xu, Haoyu Song 0001, Zhikang Chen, Wenfei Wu, Guyue Liu, Yinchao Zhang, Zerui Tian, Bin Liu 0001 |
SIGCOMM | 7 |
| 2022 | Network entitlement: contract-based network sharing with agility and SLO guaranteesabstractThis paper presents Meta's Production Wide Area Network (WAN) Entitlement solution used by thousands of Meta's services to share the network safely and efficiently. We first introduce the Network Entitlement problem, i.e., how to share WAN bandwidth across services with flexibility and SLO guarantees. We present a new abstraction entitlement contract, which is stable, simple, and operationally friendly. The contract defines services' network quota and is set up between the network team and services teams to govern their obligations. Our framework includes two key parts: (1) an entitlement granting system that establishes an agile contract while achieving network efficiency and meeting long-term SLO guarantees, and (2) a large-scale distributed run-time enforcement system that enforces the contract on the production traffic. We demonstrate its effectiveness through extensive simulations and real-world end-to-end tests. The system has been deployed and operated for over two years in production. We hope that our years of experience provide a new angle to viewing WAN network sharing in production and will inspire follow-up research. Satyajeet Ahuja, Vinayak Dangui, Kirtesh Patil, Manikandan Somasundaram, Mario A. Sánchez, Guanqing Yan, Mohammad Noormohammadpour, Alaleh Razmjoo, Grace Smith, Abhinav Triguna, Soshant Bali, Yuxiang Xiang, Prabhakaran Ganesan, Mikel Jimenez Fernandez, Petr Lapukhov, Guyue Liu, Ying Zhang 0022 |
SIGCOMM | 19 |
| 2021 | Watching the watchmen: Least privilege for managed network servicesabstractMany enterprises outsource network management (e.g., troubleshooting failures, monitoring performance) to third-party managed service providers (MSPs) to reduce cost. Unfortunately, recent incidents show that MSPs themselves have become an attractive launchpad to gain access to customer networks. In this work, we argue that such incidents arise due to a violation of the least privilege principle. We revisit the MSP outsourcing problem through this least-privilege view, identify key challenges in realizing this framework, and present initial ideas toward this goal. In particular, we propose providing the MSP provider an isolated "digital twin" environment to resolve problems and prevent providers from directly accessing the customer production network. Changes are verified before importing them into the production network, ensuring there are no privilege violations. Our preliminary experiments show that our approach can resolve practical problems (e.g., misconfigurations) and is effective in reducing the attack surfaces for MSP customers. Guyue Liu, Ao Li 0009, Christopher Canel, Vyas Sekar |
HotNets | 1 |
| 2021 | Don't Yank My Chain: Auditable NF Service Chaining
Guyue Liu, Hugo Sadok, Anne Kohlbrenner, Bryan Parno, Vyas Sekar, Justine Sherry |
NSDI | 1 |
| 2021 | Formalizing an Architectural Model of a Trustworthy Edge IoT Security Gateway‡abstractToday’s edge networks continue to see an increasing number of deployed IoT devices. These IoT devices aim to increase productivity and efficiency; however, they are plagued by a myriad of vulnerabilities. Industry and academia have proposed protecting these devices by deploying a “bolt-on” security gateway to these edge networks. The gateway applies security protections at the network level. While security gateways are an attractive solution, they raise a fundamental concern: Can the bolt-on security gateway be trusted? This paper identifies key challenges in realizing this goal and sketches a roadmap for providing trust in bolt-on edge IoT security gateways. Specifically, we show the promise of using a micro-hypervisor driven approach for delivering practical (deployable today) trust that is catered to both end-users and gateway vendors alike in terms of cost, generality, capabilities, and performance. We describe the challenges in establishing trust on today’s edge security gateways, formalize the adversary and trust properties, describe our system architecture, encode and prove our architecture trust properties using the Alloy formal modeling language. We foresee our trustworthy security gateway architecture becoming a practical and extensible formal foundation towards realizing robust trust properties on today’s edge security gateway implementations. Matt McCormack, Amit Vasudevan, Guyue Liu, Vyas Sekar |
RTCSA | 3 |
| 2020 | Fine-Grained Isolation for Scalable, Dynamic, Multi-tenant Edge Clouds
Yuxin Ren 0001, Guyue Liu, Vlad Nitu, Wenyuan Shao, Riley Kennedy, Gabriel Parmer, Timothy Wood 0001, Alain Tchana |
USENIX ATC | 2 |
| 2020 | REINFORCE: Achieving Efficient Failure Resiliency for Network Function Virtualization-Based ServicesabstractEnsuring high availability (HA) for software-based networks is a critical design feature that will help the adoption of software-based network functions (NFs) in production networks. It is important for NFs to avoid outages and maintain mission-critical operations. However, HA support for NFs on the critical data path can result in unacceptable performance degradation. We present REINFORCE, an integrated framework to support efficient resiliency for NF service chains. REINFORCE includes timely failure detection and consistent failover mechanisms. REINFORCE replicates state to standby NFs (local and remote) while enforcing correctness. It minimizes the number of state transfers by exploiting the concept of external synchrony, and leverages opportunistic batching and multi-buffering to optimize performance. Experimental results show that, even at line-rate packet processing (10 Gbps), REINFORCE achieves chain-level failover across servers in a LAN within 10ms, incurring less than 10% performance overhead, and adds average latency only ~400 μs, with a worst-case latency of less than 1ms. REINFORCE also recovers from software failures within the same node in less than 100 μs, incurring less than 1% performance overhead and adds less than 5 μs latency during normal operation. Sameer G. Kulkarni, Guyue Liu, K. K. Ramakrishnan, Mayutan Arumaithurai, Timothy Wood 0001, Xiaoming Fu 0001 |
IEEE/ACM Trans. Netw. | 2 |
| 2019 | Living on the Edge: Serverless Computing and the Cost of Failure ResiliencyabstractServerless computing platforms have gained popularity because they allow easy deployment of services in a highly scalable and cost-effective manner. By enabling just-in-time startup of container-based services, these platforms can achieve good multiplexing and automatically respond to traffic growth, making them particularly desirable for edge cloud data centers where resources are scarce. Edge cloud data centers are also gaining attention because of their promise to provide responsive, low-latency shared computing and storage resources. Bringing serverless capabilities to edge cloud data centers must continue to achieve the goals of low latency and reliability. The reliability guarantees provided by serverless computing however are weak, with node failures causing requests to be dropped or executed multiple times. Thus serverless computing only provides a best effort infrastructure, leaving application developers responsible for implementing stronger reliability guarantees at a higher level. Current approaches for providing stronger semantics such as “exactly once” guarantees could be integrated into serverless platforms, but they come at high cost in terms of both latency and resource consumption. As edge cloud services move towards applications such as autonomous vehicle control that require strong guarantees for both reliability and performance, these approaches may no longer be sufficient. In this paper we evaluate the latency, throughput, and resource costs of providing different reliability guarantees, with a focus on these emerging edge cloud platforms and applications. Sameer G. Kulkarni, Guyue Liu, K. K. Ramakrishnan, Timothy Wood 0001 |
LANMAN | 2 |
| 2019 | Advancing Network Function Virtualization Platforms with Programmable NICsabstractNetwork Function Virtualization seeks to run high performance middleboxes in a flexible, more configurable software environment. Even with advances such as kernel bypass and zero-copy IO, middlebox platforms still struggle to meet stringent throughput and latency requirements. To achieve line rates as network bandwidths rise, these platforms often must make tradeoffs such as inefficiently dedicating more CPU cores or weakening security and isolation properties. In this paper we explore how advances in programmable “smart NICs” can be leveraged by software middlebox platforms to improve performance, resource efficiency, and security. Our evaluation shows several use cases for smart NICs, which improve performance significantly while reducing resource consumption and providing strong isolation. Zhen Ni, Guyue Liu, Dennis Afanasev, Timothy Wood 0001, Jinho Hwang |
LANMAN | 2 |
| 2018 | REINFORCE: achieving efficient failure resiliency for network function virtualization based servicesabstractEnsuring high availability (HA) for software-based networks is a critical design feature that will help the adoption of software-based network functions (NFs) in production networks. It is important for NFs to avoid outages and maintain mission-critical operations. However, HA support for NFs on the critical data path can result in unacceptable performance degradation. We present REINFORCE, an integrated framework to support efficient resiliency for NFs and NF service chains. REINFORCE includes timely failure detection and consistent failover mechanisms. REINFORCE replicates state to standby NFs (local and remote) while enforcing correctness. It minimizes the number of state transfers by exploiting the concept of external synchrony, and leverages opportunistic batching and multi-buffering to optimize performance. Experimental results show that, even at line-rate packet processing (10 Gbps), REINFORCE achieves chain-level failover across servers in a LAN (or within the same node) within 10ms (100/μs), incurring less than 10% (1%) performance overhead, and adds average latency of only ~400/μs (5/μs), with a worst-case latency of less than 1ms (10/μs). Sameer G. Kulkarni, Guyue Liu, K. K. Ramakrishnan, Mayutan Arumaithurai, Timothy Wood 0001, Xiaoming Fu 0001 |
CoNEXT | 2 |
| 2018 | Scalable Memory Reclamation for Multi-Core, Real-Time SystemsabstractA core challenge in best utilizing an increasing number of cores in real-time systems is addressing the problem of efficient and predictable resource sharing. Traditional mechanisms for mutual exclusion, such as locks, limit parallelism due to serialized resource access. Relaxing mutual exclusion, reader-writer locks enable selective parallelism for a subset of accesses, but can suffer from increased implementation overheads. In all such implementations, the costs of cache-coherency alone can be prohibitive for an increasing number of cores. This paper investigates the use of techniques such as Read-Copy Update (RCU) to enable truly parallel access to data-structures. Such techniques optimize for data-structure read-paths, and can completely avoid stores to shared structures, thus avoiding cache-coherency overheads. We show that existing implementations of preemptive RCU aren't designed to provide real-time latencies, and require a potentially unbounded amount of dynamically allocated memory. Thus, we introduce two new implementations that are both predictable and efficient, and a matching analysis that establishes bounds on memory consumption. We additionally provide a schedulability analysis that demonstrates the effectiveness of scalable read-side operations, achieving consistently higher schedulability than existing techniques. We further apply the analysis to provide admission control for a soft real-time application to both achieve higher throughput than existing approaches (up to 40% higher) while limiting 99th percentile read-path latencies (4x lower than existing techniques). Yuxin Ren 0001, Guyue Liu, Gabriel Parmer, Björn B. Brandenburg |
RTAS | 2 |
| 2018 | Microboxes: high performance NFV with customizable, asynchronous TCP stacks and dynamic subscriptionsabstractExisting network service chaining frameworks are based on a "packet-centric" model where each NF in a chain is given every packet for processing. This approach becomes both inefficient and inconvenient for more complex network functions that operate at higher levels of the protocol stack. We propose Microboxes, a novel service chaining abstraction designed to support transport- and application-layer middle-boxes, or even end-system like services. Simply including a TCP stack in an NFV platform is insufficient because there is a wide spectrum of middlebox types-from NFs requiring only simple TCP bytestream reconstruction to full endpoint termination. By exposing a publish/subscribe-based API for NFs to access packets or protocol events as needed, Microboxes eliminates redundant processing across a chain and enables a modular design. Our implementation on a DPDK-based NFV framework can double throughput by consolidating stack operations and provide a 51% throughput gain by customizing TCP processing to the appropriate level. Guyue Liu, Yuxin Ren 0001, Mykola Yurchenko, K. K. Ramakrishnan, Timothy Wood 0001 |
SIGCOMM | 1 |
| 2016 | OpenNetVM: Flexible, high performance NFV (Demo)abstractNetwork Function Virtualization promises to enable dynamic management of software-based network functions. We envision a dynamic and flexible network that can support a smarter data plane than just simple switches that forward packets. This network architecture supports complex stateful routing of flows where processing by network functions (NFs) can transform packet data, customized on a per-flow basis, as it moves between end points. This demo will present OpenNetVM, a highly efficient packet processing framework that greatly simplifies the development of network functions, as well as their management and optimization. OpenNetVM runs network functions in lightweight Docker containers that start in less than a second. The OpenNetVM platform manager provides load balancing, flexible flow management, and service name abstractions. OpenNetVM uses DPDK for high performance I/O, and efficiently routes packets through dynamically created service chains. We will demonstrate how the research community can easily build new network functions and rapidly deploy them to see their effectiveness in high performance network environments. Wei Zhang 0052, Guyue Liu, Phil Lopreiato, Grégoire Todeschi, K. K. Ramakrishnan, Timothy Wood 0001 |
LANMAN | 2 |
| 2016 | NetAlytics: Cloud-Scale Application Performance Monitoring with SDN and NFV
Guyue Liu, Michael Trotter, Yuxin Ren 0001, Timothy Wood 0001 |
Middleware | 1 |
| 2016 | SDNFV: Flexible and Dynamic Software Defined Control of an Application- and Flow-Aware Data Plane
Wei Zhang 0052, Guyue Liu, Ali Mohammadkhan, Jinho Hwang, K. K. Ramakrishnan, Timothy Wood 0001 |
Middleware | 2 |
| 2015 | Cloud-Scale Application Performance Monitoring with SDN and NFVabstractIn cloud data centers, more and more services are deployed across multiple tiers to increase flexibility and scalability. However, this makes it difficult for the cloud provider to identify which tier of the application is the bottleneck and how to resolve performance problems. Existing solutions approach this problem by constantly monitoring either in end-hosts or physical switches. Host based monitoring usually needs instrumentation of application code, making it less practical, while network hardware based monitoring is expensive and requires special features in each physical switch. Instead, we believe network wide monitoring should be flexible and easy to deploy in a non-intrusive way by exploiting recent advances in software-based network services. Towards this end we are developing a distributed software-based network monitoring framework for cloud data centers. Our system leverages knowledge of topology and routing information to build relationships between each tier of the application, and detect and locate performance bottlenecks by monitoring the network inside software switches. Guyue Liu, Timothy Wood 0001 |
IC2E | 1 |
| 2015 | Virtual function placement and traffic steering in flexible and dynamic software defined networksabstractThe integration of network function virtualization (NFV) and software defined networks (SDN) seeks to create a more flexible and dynamic software-based network environment. The line between entities involved in forwarding and those involved in more complex middle box functionality in the network is blurred by the use of high-performance virtualized platforms capable of performing these functions. A key problem is how and where network functions should be placed in the network and how traffic is routed through them. An efficient placement and appropriate routing increases system capacity while also minimizing the delay seen by flows. In this paper, we formulate the problem of network function placement and routing as a mixed integer linear programming problem. This formulation not only determines the placement of services and routing of the flows, but also seeks to minimize the resource utilization. We develop heuristics to solve the problem incrementally, allowing us to support a large number of flows and to solve the problem for incoming flows without impacting existing flows. Ali Mohammadkhan, Sheida Ghapani, Guyue Liu, Wei Zhang 0052, K. K. Ramakrishnan, Timothy Wood 0001 |
LANMAN | 3 |
| 2015 | CloudNet: Dynamic Pooling of Cloud Resources by Live WAN Migration of Virtual MachinesabstractVirtualization technology and the ease with which virtual machines (VMs) can be migrated within the LAN have changed the scope of resource management from allocating resources on a single server to manipulating pools of resources within a data center. We expect WAN migration of virtual machines to likewise transform the scope of provisioning resources from a single data center to multiple data centers spread across the country or around the world. In this paper, we present the CloudNet architecture consisting of cloud computing platforms linked with a virtual private network (VPN)-based network infrastructure to provide seamless and secure connectivity between enterprise and cloud data center sites. To realize our vision of efficiently pooling geographically distributed data center resources, CloudNet provides optimized support for live WAN migration of virtual machines. Specifically, we present a set of optimizations that minimize the cost of transferring storage and virtual machine memory during migrations over low bandwidth and high-latency Internet links. We evaluate our system on an operational cloud platform distributed across the continental US. During simultaneous migrations of four VMs between data centers in Texas and Illinois, CloudNet's optimizations reduce memory migration time by 65% and lower bandwidth consumption for the storage and memory transfer by 19 GB, a 50% reduction. Timothy Wood 0001, K. K. Ramakrishnan, Prashant J. Shenoy, Jacobus E. van der Merwe, Jinho Hwang, Guyue Liu, Lucas Chaufournier |
IEEE/ACM Trans. Netw. | 6 |
| 2014 | Topology Discovery and Service Classification for Distributed-Aware CloudsabstractCloud data centers are difficult to manage because providers have no knowledge of what applications are being run by customers or how they interact. As a consequence, current clouds provide minimal automated management functionality, passing the problem on to users who have access to even fewer tools since they lack insight into the underlying infrastructure. Ideally, the cloud platform, not the customer, should be managing data center resources in order to both use them efficiently and provide strong application-level performance and reliability guarantees. To do this, we believe that clouds must become "distibuted-aware" so that they can deduce the overall structure and dependencies within a client's distributed applications and use that knowledge to better guide management services. Towards this end we are developing a light-weight topology detection system that maps distributed applications and a service classification algorithm that can determine not only overall application types, but individual VM roles as well. Jinho Hwang, Guyue Liu, Sai Zeng, Frederick Y. Wu, Timothy Wood 0001 |
IC2E | 2 |