EDBT 2026 Demo / reviewers in the wild / expert
Ennan Zhai
dblp:23/7560
· DBLP profile ↗
84ranked-venue papers
10as first author
46since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 56 · 5 first-author · 38 since 2021Software engineering, systems software and programming languages · 11 · 3 first-author · 5 since 2021Systems, architecture and hardware · 8 · 2 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-authorSecurity and privacy · 4 · 1 since 2021Theory of computation · 3Artificial intelligence and machine learning · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EROICA: Online Performance Troubleshooting for Large-scale Model Training
Yu Guan 0005, Zhiyu Yin, Sheng Cheng 0002, Chaojie Yang, Kun Qian 0004, Tianyin Xu, Yang Zhang 0102, Yong Li 0008, Dennis Cai, Ennan Zhai |
NSDI | 13 |
| 2026 | HeteCCL: Synthesizing Near-Optimal Collective Communication Schedules for Heterogeneous GPU Clusters
Chenyang Hei, Jiamin Cao, Chengxi Gao, Xiuzhu Sha, Tongrui Liu, Dengke Zhang, Ennan Zhai, Xingwei Wang 0001 |
NSDI | 8 |
| 2026 | Come Hell or Still Water: Alleviating Tail Latency in Cloud Block Store
Chaolei Hu, Kun Qian 0004, Erci Xu, Xue Li 0024, Yuesheng Gu, Lingjun Zhu, Fengyuan Ren, Ennan Zhai |
NSDI | 10 |
| 2026 | ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production
Yuxing Xiang, Xue Li 0024, Kun Qian 0004, Yan Zhang 0117, Wenyuan Yu, Ennan Zhai, Xin Jin 0008, Jingren Zhou 0001 |
NSDI | 6 |
| 2026 | Diagnosing and Repairing Distributed Routing Configurations Using Selective Symbolic Simulation
Rulan Yang, Gao Han, Hanyang Shao, Xiaoqiang Zheng, Lizhao You, Ruiting Zhou, Linghe Kong, Ennan Zhai, Qiao Xiang, Jiwu Shu |
NSDI | 10 |
| 2026 | Balancing and Beyond: Communication-Centric Optimizations in Expert ParallelismabstractThe Mixture-of-Experts (MoE) architecture scales large language models (LLMs) to trillions of parameters by activating only a small subset of experts per token. In practice, MoE inference is commonly deployed with Expert Parallelism (EP), which places whole experts on different GPUs to preserve kernel efficiency. However, production EP deployments often suffer from two bottlenecks: (1) expert workload imbalance, which creates computation and communication stragglers, and (2) communication inefficiency, where inter-GPU transfers dominate latency even after balancing. We present EPIC, an experience-driven EP inference system that addresses these issues progressively for real deployments. EPIC mitigates imbalance via performance-aware expert migration and runtime expert activation, and then improves communication with topology-adaptive transport kernels and fine-grained computation-communication overlap. EPIC has been deployed at scale across O(10K) GPUs in our online inference service for both open-source models (e.g., Qwen3-Coder and DeepSeek-R1) and internal models, reducing communication time and per-token latency by up to 40% and 21%, respectively. Jiamin Cao, Qingxu Li, Yaozhong Liu, Shangfeng Shi, Kunling He, Ennan Zhai, Jianbo Dong, Binzhang Fu, Dennis Cai |
SIGCOMM | 11 |
| 2026 | From Nimitz to NetPila: The Evolution of Production-Scale Container Network
Sheng Cheng 0002, Jiamin Cao, Shuhong Zhu, Ennan Zhai, Dennis Cai |
SIGCOMM | 8 |
| 2026 | Networked Agent Memory and Causality Representation: Experiences towards Interpretable Cloud-Scale Root-Causing
Yanyu Ren, Xianshang Lin, Chenxu Wang 0007, Li Chen 0008, Shuai Wang 0028, Kaihui Gao, Dan Li 0001, Chen Tian 0001, Yunguang Li, Ennan Zhai |
SIGCOMM | 13 |
| 2026 | AIDA: Accelerating Root Cause Analysis for Multi-Vendor Device Failures with LLM-Powered ReasoningabstractRoot cause analysis (RCA) of network device failures is critical to cloud reliability. While monitoring can identify which device has failed, diagnosing why remains a slow, manual process, increasing the risk of recurring failures and cascading service disruptions. Existing automated methods are inadequate: traditional methods lack precision, while prior machine learning (ML) and large language model (LLM) approaches are often too coarse-grained, require heavy manual configuration, or fail to produce verifiable reasoning essential for operator trust. This paper presents AIDA, the first system to deliver automated, fine-grained RCA of network device failures, deployed at scale in Alibaba Cloud's production network. AIDA's contributions include: (1) fine-tuning an LLM with reinforcement learning to distill expert logic into interpretable reasoning chains; (2) synthesizing these chains into an evolving knowledge graph (KG) to support retrieval-augmented generation (RAG); and (3) employing RAG-driven multi-step inference wherein the LLM is sequentially guided by the KG to construct robust, verifiable reasoning. Deployed for over a year, AIDA has achieved 95.4% precision with interpretable output and reduced the median RCA time from 72.6 hours to 1.6 minutes. Notably, it curtails the 90th-percentile diagnosis latency from 329.9 hours to 19.4 hours. Xuan Zeng 0002, Xumiao Zhang, Xiaoxi Zhang 0001, Deke Guo, Ennan Zhai |
SIGCOMM | 8 |
| 2026 | Retriever: A Distributed Intrusion Detection System for NOS-Enabled NetworksabstractNetwork Operating Systems (NOS) are being widely deployed on edge devices by cloud service providers to perform fast configurations and offer high availability for new network protocols. However, NOS-enabled networks open the door to intruders that can stealthily corrupt less-guarded programmable switches to launch attacks on the entire network. Traditional centralized intrusion detection systems may neglect anomalous events on NOS-equipped switches and fail to detect such attacks. In this paper, we make the first attempt towards intrusion detection for NOS-enabled networks by designingRetriever.Retrieverfeatures a lightweight local anomaly detection module on programmable switches and a central anomaly assessment module on the central server. The local anomaly detection module selectively traces both system and network events on switches, based on which a provenance graph of events is established. Upcoming events unmatched by the provenance graph are aggregated to construct a suspicious subgraph to report to the central server. The central anomaly assessment module extracts semantic representations from reported suspicious subgraphs and computes their anomaly scores. Large-scale experiments show thatRetrievercan achieve high intrusion detection accuracy (nearly 100%) with low overheads. Runmin Ou, Yijie Bai, Yanjiao Chen, Bingchuan Tian, Zhiming Ji, Ennan Zhai, Dennis Cai, Wenyuan Xu 0001 |
IEEE Trans. Dependable Secur. Comput. | 7 |
| 2025 | Enhancing Large-Scale AI Training Efficiency: The C4 Solution for Real-Time Anomaly Detection and Communication OptimizationabstractThe emergence of Large Language Models (LLMs) has necessitated the adoption of distributed training techniques, involving the deployment of thousands of GPUs to train a single model. Unfortunately, the efficiency of large-scale distributed training systems is often suboptimal due to the increased likelihood of hardware errors in high-end GPU products and the heightened risk of network traffic collisions. Specifically, GPUs involved in the same job require periodic synchronization to exchange necessary data, such as gradients, parameters, or activations. As a result, any local hardware failure can disrupt training tasks, and the inability to swiftly identify faulty components leads to a significant waste of GPU resources. Moreover, prolonged communication due to traffic collisions can substantially increase GPU waiting times. To address these challenges, we propose a communication-driven solution, namely the C 4. The key insights of C 4 are twofold. First, the load in distributed training exhibits homogeneous characteristics and is divided into iterations through periodic synchronization, therefore hardware anomalies would incur certain syndrome in collective communication. By leveraging this feature, $\mathbf{C} 4$ can rapidly identify the faulty components, swiftly isolate the anomaly, and restart the task, thereby avoiding resource wastage caused by delays in anomaly detection. Second, the predictable communication model of collective communication, involving a limited number of long-lived flows, allows C 4 to efficiently execute traffic planning, substantially reducing bandwidth competition among these flows. The $\mathbf{C 4}$ has been extensively deployed across real-world production systems in a hyperscale cloud provider, yielding a significant improvement in system efficiency, from 30% to $\mathbf{4 5 \%}$. This enhancement is attributed to a $\mathbf{3 0 \%}$ reduction in error-induced overhead and a 15% reduction in communication costs. Jianbo Dong, Yikai Zhu, Hairong Jiao, Ennan Zhai, Wencong Xiao, Man Yuan, Siran Yang, Jiamang Wang, Rui Men, Dennis Cai, Binzhang Fu |
HPCA | 13 |
| 2025 | Mitigating Scalability Walls of RDMA-based Container Networks
Wei Liu 0148, Kun Qian 0021, Zhenhua Li 0001, Feng Qian 0001, Tianyin Xu, Yunhao Liu 0001, Yu Guan 0005, Shuhong Zhu, Hongfei Xu, Lanlan Xi, Ennan Zhai |
NSDI | 12 |
| 2025 | Evolution of Aegis: Fault Diagnosis for AI Model Training Service in Production
Jianbo Dong, Kun Qian 0021, Zhilong Zheng, Liang Chen 0001, Yichi Xu, Yikai Zhu, Xue Li 0024, Zhihui Ren, Yang Liu 0245, Yu Guan 0005, Chaojie Yang, Yang Zhang 0102, Man Yuan, Yong Li 0008, Xianlong Zeng, Zhiping Yao, Binzhang Fu, Ennan Zhai, Wei Lin 0016, Dennis Cai |
NSDI | 29 |
| 2025 | SimAI: Unifying Architecture Design and Performance Tuning for Large-Scale Large Language Model Training with Scalability and Precision
Xizheng Wang, Qingxu Li, Yichi Xu, Dan Li 0001, Li Chen 0008, Heyang Zhou, Linkang Zheng, Yikai Zhu, Yang Liu 0245, Kun Qian 0021, Kunling He, Ennan Zhai, Dennis Cai, Binzhang Fu |
NSDI | 16 |
| 2025 | Learning Production-Optimized Congestion Control Selection for Alibaba Cloud CDN
Xuan Zeng 0002, Xumiao Zhang, Xiaoxi Zhang 0001, Xu Chen 0004, Guihai Chen, Yubing Qiu, Chong Hao, Ennan Zhai |
NSDI | 11 |
| 2025 | New Evolution of Hoyan: Enhancing Scalability, Usability, and Accuracy for Alibaba's Global WAN VerificationabstractThe network verification system Hoyan has been deployed for Alibaba Cloud's wide-area network (WAN) for years and achieved considerable success in preventing misconfiguration-caused network incidents. However, recent years have seen the emergence of new challenges in scalability, usability, and accuracy for Hoyan. This paper presents the new evolution of Hoyan to address these challenges. First, to support the large increase in the number of routers and prefixes on our WAN, Hoyan's simulation has evolved from a centralized fashion to a distributed framework, which improves the efficiency by 5 times and can scale to O(104) routers, millions of prefixes, and billions of flows. Second, to improve Hoyan's usability in checking route change intents, we developed a specification language RCL, which supports the easy specification and automatic verification of route change intents. Third, to ensure high accuracy we enhanced Hoyan's accuracy diagnosis framework, which helped us identify and fix dozens of implementation and modeling issues. Hoyan is used on a daily basis for our WAN. It supports O(100) verification requests each week, prevents O(10) incidents each year, and helps reduce the percentage of misconfiguration-caused network incidents from 56% to 5%. Yifei Yuan 0001, Fangdan Ye, Jingkai Zhang, Mengqi Liu 0001, Yuyang Sang, Ruizhen Yang, Duncheng She, Zhiqing Ye, Tianchen Guo, Xinji Tang, Zhongyu Guan, Lingpeng Su, Ci Wang, Ruiyang Feng, Zhonghui Xie, Xianlong Zeng, Dennis Cai, Ennan Zhai |
SIGCOMM | 25 |
| 2025 | SkeletonHunter: Diagnosing and Localizing Network Failures in Containerized Large Model TrainingabstractThe flexibility and portability characteristics have made containers a popular serverless environment for large model training in recent years. Unfortunately, these advantages render the network support for containerized large model training extremely challenging, due to the high dynamics of containers, the complex interplay between underlay and overlay networks, and the stringent requirements on failure detection and localization. Existing data center network debugging tools, which rely on comprehensive or opportunistic monitoring, are either inefficient or inaccurate in this setting. Wei Liu 0148, Kun Qian 0021, Zhenhua Li 0001, Tianyin Xu, Yunhao Liu 0001, Jiakang Li, Shuhong Zhu, Xue Li 0024, Hongfei Xu, Ennan Zhai |
SIGCOMM | 13 |
| 2025 | SyCCL: Exploiting Symmetry for Efficient Collective Communication SchedulingabstractThe performance of collective communication schedules is crucial for the efficiency of machine learning jobs and GPU cluster utilization. Existing open-source collective communication libraries (such as NCCL and RCCL) rely on fixed schedules and cannot adjust to varying topology and model requirements. State-of-the-art collective schedule synthesizers (such as TECCL and TACCL) utilize Mixed Integer Linear Program for modeling but encounter search space explosion and scalability challenges. In this paper, we propose SyCCL, a scalable collective schedule synthesizer that aims to synthesize near-optimal schedules in tens of minutes for production-scale machine-learning jobs. SyCCL leverages collective and topology symmetries to decompose the original collective communication demand into smaller sub-demands within smaller topology subsets. SyCCL proposes efficient search strategies to quickly explore potential sub-demands, synthesizes corresponding sub-schedules, and integrates these sub-schedules into complete schedules. Our 32-A100 testbed and production-scale simulation experiments show that SyCCL improves collective performance by up to 127% while reducing synthesis time by 2 to 4 orders of magnitude compared to state-of-the-art efforts. Jiamin Cao, Shangfeng Shi, Weisen Liu, Yifan Yang 0009, Yichi Xu, Zhilong Zheng, Yu Guan 0005, Kun Qian 0021, Ying Liu 0024, Mingwei Xu 0001, Ning Wang 0001, Jianbo Dong, Binzhang Fu, Dennis Cai, Ennan Zhai |
SIGCOMM | 17 |
| 2025 | ParserHawk: Hardware-aware parser generator using program synthesisabstractParser programs are becoming increasingly complex to accommodate intricate network packet formats and advanced protocols. Existing parser compilers incorporate predefined program rewrite rules to output the low-level parser implementation. Yet, these rules are often brittle and sensitive to how the input parser program is written. As a result, generated implementations could consume more hardware resources than necessary. In some cases, these compilers unnecessarily reject valid parser programs that could have fit within the target device parser's resource constraints. Karan Kumar G., Ennan Zhai, Bili Dong, Joseph Tassarotti, Srinivas Narayana, Anirudh Sivaraman |
SIGCOMM | 5 |
| 2025 | ResCCL: Resource-Efficient Scheduling for Collective CommunicationabstractAs distributed deep learning training (DLT) systems scale, collective communication has become a significant performance bottleneck. While current approaches optimize bandwidth utilization and task completion time, existing communication libraries (CCLs) backends fail to efficiently manage GPU resources during algorithm execution, limiting the performance of advanced algorithms. This paper proposes ResCCL, a novel CCL backend designed for Resource-Efficient Scheduling to address key limitations in current systems. ResCCL enhances execution efficiency by optimizing scheduling at the primitive level (e.g., send and recvReduceCopy), enabling flexible thread block (TB) allocation, and generating lightweight communication kernels to minimize runtime overhead. Our approach tackles the global scheduling problem, reduces idle TB resources, and enhances communication bandwidth. Evaluation results demonstrate that ResCCL achieves up to 2.5× improvement in bandwidth performance compared to both NCCL and MSCCL. It reduces SM resource overhead by 77.8% and increases TB utilization by 41.6% while running the same algorithms. In end-to-end DLT, ResCCL boosts Megatron's throughput by up to 39%. Tongrui Liu, Chenyang Hei, Fuliang Li, Chengxi Gao, Jiamin Cao, Ennan Zhai, Xingwei Wang 0001 |
SIGCOMM | 7 |
| 2025 | Alibaba Stellar: A New Generation RDMA Network for Cloud AIabstractThe rapid adoption of Large Language Models (LLMs) in cloud environments has intensified the demand for high-performance AI training and inference, where Remote Direct Memory Access (RDMA) plays a critical role. However, existing RDMA virtualization solutions, such as Single-Root Input/Output Virtualization (SR-IOV), face significant limitations in scalability, performance, and stability. These issues include lengthy container initialization times, hardware resource constraints, and inefficient traffic steering. To address these challenges, we propose Stellar, a new generation RDMA network for cloud AI. Stellar introduces three key innovations: Para-Virtualized Direct Memory Access (PVDMA) for on-demand memory pinning, extended Memory Translation Table (eMTT) for optimized GPU Direct RDMA (GDR) performance, and RDMA Packet Spray for efficient multi-path utilization. Deployed in our large-scale AI clusters, Stellar spins up virtual devices in seconds, reduces container initialization time by 15 times, and improves LLM training speed by up to 14%. Our evaluations demonstrate that Stellar significantly outperforms existing solutions, offering a scalable, stable, and high-performance RDMA network for cloud AI. Menglei Zheng, Binbin Liao, Suwei Xu, Yongjia Mo, Qinghua Peng, Jilie Luo, Qingxu Li, Zishu Wang, Jianbo Dong, Kunling He, Sheng Cheng 0002, Jiamin Cao, Hairong Jiao, Lingjun Zhu, Yiquan Chen, Wei Wang 0030, Shuhong Zhu, Xingru Li, Qiang Wang 0022, Wei Lin 0016, Ennan Zhai, Jiesheng Wu, Qiang Liu 0036, Binzhang Fu, Dennis Cai |
SIGCOMM | 35 |
| 2025 | Towards LLM-Based Failure Localization in Production-Scale NetworksabstractRoot causing and failure localization are critical to maintain reliability in cloud network operations. When an incident is reported, network operators must review massive volumes of monitoring data and identify the root cause (i.e., error device) as fast as possible, making it extremely challenging even for experienced operators. Large language models (LLMs) have shown great potential in text understanding and reasoning. In this paper, we present BiAn, an LLM-based framework designed to assist operators in efficient incident investigation. BiAn processes monitoring data and generates error device rankings with detailed explanations. To date, BiAn has been deployed in our network infrastructure for 10 months and it has successfully assisted operators in identifying error devices more quickly, reducing time to root causing by 20.5% (55.2% for high-risk incidents). Extensive performance evaluations based on 17 months of real cases further demonstrate that BiAn achieves accurate and fast failure localization. It improves accuracy by 9.2% compared to the baseline approach. Chenxu Wang 0007, Xumiao Zhang, Runwei Lu, Xianshang Lin, Xuan Zeng 0002, Zhe An, Gongwei Wu, Chen Tian 0001, Guihai Chen, Guyue Liu, Yuhong Liao, Dennis Cai, Ennan Zhai |
SIGCOMM | 16 |
| 2025 | SkyNet: Analyzing Alert Flooding from Severe Network Failures in Large Cloud InfrastructuresabstractFor providers operating large-scale global networks, the timeliness of network failure recovery significantly affects the reliability of network services. Ideally, a network monitoring system should have enough coverage to detect even minor issues, but high coverage means alert floods during severe network failures. In practice, there is a gap between the flooding raw alerts data collected by network monitoring tools and the readable information needed for failure diagnosis. Existing solutions using limited network monitoring data sources and heuristic diagnostic rules, lack comprehensive coverage and the capability to address severe failures, especially which network operators have never handled a similar one before. This paper presents SkyNet, a network analysis system to extract scope and severity information from alert floods. SkyNet ensures comprehensive coverage by integrating multiple monitoring data sources through a uniform input format, enhancing extensibility for new network monitoring tools. During alert floods, SkyNet groups alerts, assesses their severity, and filters out insignificant ones to aid network operators in mitigating network failures. To date, SkyNet has been running stably on our network for one and a half years without any false negatives and has successfully reduced the time-to-mitigation for over 80% of network failures since its deployment in production. Huanwu Hu, Yunguang Li, Xiangyu Tang, Bingchuan Tian, Gongwei Wu, Xumiao Zhang, Ennan Zhai, Yuhong Liao, Dennis Cai |
SIGCOMM | 12 |
| 2025 | Aegaeon: Effective GPU Pooling for Concurrent LLM Serving on the Market
Yuxing Xiang, Xue Li 0024, Kun Qian 0021, Diwen Zhu, Wenyuan Yu, Ennan Zhai, Xuanzhe Liu, Xin Jin 0008, Jingren Zhou 0001 |
SOSP | 7 |
| 2025 | Roaming Free in the VR World with MP2
Xumiao Zhang, Yuning Chen, Xuan Zeng 0002, Zhilong Zheng, Xianshang Lin, Yanmei Liu, Songwu Lu, Z. Morley Mao, Wan Du, Dennis Cai, Ennan Zhai |
USENIX ATC | 13 |
| 2024 | Cross-Platform Transpilation of Packet-Processing Programs using Program SynthesisabstractThe proliferation of programmable network devices offers a wide range of device options for developers of packet processing programs. However, there are several differences in programming language usage, hardware resource constraints, and hardware architecture across these devices. Programmers must understand multiple programming languages and hardware designs to write programs for various devices. Karan Kumar G., Ennan Zhai, Srinivas Narayana, Anirudh Sivaraman |
APNet | 4 |
| 2024 | LuoShen: A Hyper-Converged Programmable Gateway for Multi-Tenant Multi-Service Edge Clouds
Tian Pan 0001, Xionglie Wei, Yisong Qiao, Tiesheng Cheng, Wenqiang Su, Yuke Hong, Zhengzhong Wang, Chongjing Dai, Peiqiao Wang, Xuetao Jia, Jianyuan Lu, Enge Song, Biao Lyu, Ennan Zhai, Jiao Zhang 0002, Tao Huang 0005, Dennis Cai, Shunmin Zhu |
NSDI | 21 |
| 2024 | Sirius: Composing Network Function Chains into P4-Capable Edge Gateways
Jiamin Cao, Mengqi Liu 0001, Dennis Cai, Ennan Zhai |
NSDI | 7 |
| 2024 | Reasoning about Network Traffic Load Property at Production Scale
Fangdan Ye, Yifei Yuan 0001, Ruizhen Yang, Bingchuan Tian, Tianchen Guo, Zhongyu Guan, Xianlong Zeng, Chenren Xu, Dennis Cai, Ennan Zhai |
NSDI | 14 |
| 2024 | Burstable Cloud Block Storage with Data Processing Units
Junyi Shu, Kun Qian 0021, Ennan Zhai, Xuanzhe Liu, Xin Jin 0008 |
OSDI | 3 |
| 2024 | Crux: GPU-Efficient Communication Scheduling for Deep Learning TrainingabstractDeep learning training (DLT), e.g., large language model (LLM) training, has become one of the most important services in multitenant cloud computing. By deeply studying in-production DLT jobs, we observed that communication contention among different DLT jobs seriously influences the overall GPU computation utilization, resulting in the low efficiency of the training cluster. In this paper, we present Crux, a communication scheduler that aims to maximize GPU computation utilization by mitigating the communication contention among DLT jobs. Maximizing GPU computation utilization for DLT, nevertheless, is NP-Complete; thus, we formulate and prove a novel theorem to approach this goal by GPU intensity-aware communication scheduling. Then, we propose an approach that prioritizes the DLT flows with high GPU computation intensity, reducing potential communication contention. Our 96-GPU testbed experiments show that Crux improves 8.3% to 14.8% GPU computation utilization. The large-scale production trace-based simulation further shows that Crux increases GPU computation utilization by up to 23% compared with alternatives including Sincronia, TACCL, and CASSINI. Jiamin Cao, Yu Guan 0005, Kun Qian 0021, Wencong Xiao, Jianbo Dong, Binzhang Fu, Dennis Cai, Ennan Zhai |
SIGCOMM | 9 |
| 2024 | A General and Efficient Approach to Verifying Traffic Load Properties under Arbitrary k FailuresabstractThis paper presents YU, the first verification system for checking traffic load properties under arbitrary failure scenarios that can scale to production Wide Area Networks (WANs). Building a practical YU requires us to address two challenges in terms of generality and efficiency. The state-of-the-art efforts either assume shortest-path-based forwarding (e.g., QARC) or only target single-failure reasoning (e.g., Jingubang). As a result, the former inherently cannot generalize to widely used protocols (e.g., SR and iBGP) that are beyond shortest-path forwarding, while the latter cannot efficiently handle arbitrary failure scenarios. For the generality challenge, we propose an approach inspired by symbolic execution, called symbolic traffic execution, to model the forwarding behavior of a range of practically deployed protocols (e.g., eBGP, iBGP, iGP, and SR) under failure scenarios. For the efficiency challenge, we propose diverse equivalence classification techniques (i.e., k-failure-equivalence and link-local-equivalence reduction) to reduce the symbolic traffic execution overhead caused by both the large size of the production WAN and the huge number of traffic flows traversing it. YU has been used in the daily verification of our WAN for several months and has successfully identified potential failure scenarios that would lead to traffic load violations. Yifei Yuan 0001, Fangdan Ye, Mengqi Liu 0001, Ruizhen Yang, Tianchen Guo, Xianlong Zeng, Chenren Xu, Dennis Cai, Ennan Zhai |
SIGCOMM | 12 |
| 2024 | Alibaba HPN: A Data Center Network for Large Language Model TrainingabstractThis paper presents HPN, Alibaba Cloud's data center network for large language model (LLM) training. Due to the differences between LLMs and general cloud computing (e.g., in terms of traffic patterns and fault tolerance), traditional data center networks are not well-suited for LLM training. LLM training produces a small number of periodic, bursty flows (e.g., 400Gbps) on each host. This characteristic of LLM training predisposes Equal-Cost Multi-Path (ECMP) to hash polarization, causing issues such as uneven traffic distribution. HPN introduces a 2-tier, dual-plane architecture capable of interconnecting 15K GPUs within one Pod, typically accommodated by the traditional 3-tier Clos architecture. Such a new architecture design not only avoids hash polarization but also greatly reduces the search space for path selection. Another challenge in LLM training is that its requirement for GPUs to complete iterations in synchronization makes it more sensitive to singlepoint failure (typically occurring on ToR). HPN proposes a new dual-ToR design to replace the single-ToR in traditional data center networks. HPN has been deployed in our production for more than eight months. We share our experience in designing, and building HPN, as well as the operational lessons of HPN in production. Kun Qian 0021, Yongqing Xi, Jiamin Cao, Yichi Xu, Yu Guan 0005, Binzhang Fu, Xuemei Shi, Fangbo Zhu, Rui Miao 0001, Peng Wang 0185, Xianlong Zeng, Eddie Ruan, Zhiping Yao, Ennan Zhai, Dennis Cai |
SIGCOMM | 17 |
| 2024 | Relational Network VerificationabstractRelational network verification is a new approach for validating network changes. In contrast to traditional network verification, which analyzes specifications for a single network snapshot, it analyzes specifications that capture similarities and differences between two network snapshots (e.g., pre- and post-change snapshots). Relational specifications are compact and precise because they focus on the flows and paths that change between snapshots and then simply mandate that all other network behaviors "stay the same", without enumerating them. To achieve similar guarantees, single-snapshot specifications would need to enumerate all flow and path behaviors that are not expected to change in order to enable checking that nothing has accidentally changed. Such specifications are proportional to network size, which makes them impractical to generate for many real-world networks. Xieyang Xu, Yifei Yuan 0001, Zachary Kincaid, Arvind Krishnamurthy, Ratul Mahajan, David Walker 0001, Ennan Zhai |
SIGCOMM | 7 |
| 2023 | Norma: Towards Practical Network Load Testing
Bingchuan Tian, Chen Tian 0001, Yu Zhou 0008, Mengjing Ma, Zhewen Yang, Guihai Chen, Dennis Cai, Ennan Zhai |
NSDI | 12 |
| 2023 | CellFusion: Multipath Vehicle-to-Cloud Video Streaming with Network Coding in the WildabstractThis paper presents CellFusion, a system designed for high-quality, real-time video streaming from vehicles to the cloud. It leverages an innovative blend of multipath QUIC transport and network coding. Surpassing the limitations of individual cellular carriers, CellFusion uses a unique last-mile overlay that integrates multiple cellular networks into a single, unified cloud connection. This integration is made possible through the use of in-vehicle Customer Premises Equipment (CPEs) and edge-cloud proxy servers. Yunzhe Ni, Zhilong Zheng, Xianshang Lin, Fengyu Gao, Xuan Zeng 0002, Yirui Liu 0001, Senlang Du, Guang Yang 0006, Yuanchao Su, Dennis Cai, Hongqiang Harry Liu, Chenren Xu, Ennan Zhai |
SIGCOMM | 16 |
| 2023 | XRON: A Hybrid Elastic Cloud Overlay Network for Video Conferencing at Planetary ScaleabstractQuality and cost are two key considerations for video conferencing services. Service providers face a dilemma when selecting network tiers to build their infrastructure---relying on Internet links has poor quality, while using premium links brings excessive cost. Bingyang Wu, Kun Qian 0021, Bo Li 0061, Dennis Cai, Ennan Zhai, Xuanzhe Liu, Xin Jin 0008 |
SIGCOMM | 9 |
| 2023 | Automated Verification of an In-Production DNS Authoritative EngineabstractThis paper presents DNS-V, a verification framework for our in-production DNS authoritative engine, which is the core of our DNS service. The key idea for automated verification in general is based on the layered verification principle. However, we face the challenge that our in-production DNS authoritative engine lacks modularity, more specifically, as can be seen with unclean interfaces and poor data structure encapsulation. This makes the layered verification hard to apply. To address this challenge, we propose a summarization approach that performs full-path symbolic execution to accumulate all path conditions and computation effects, and then represents a module's behavior in an abstract form as a set of input-effect pairs. In addition, for portability to future iterated versions of our DNS authoritative engine, we identify common dependency library modules that remain stable across different versions, and carefully design their abstractions to make them amenable to automated reasoning. Our framework has been successful in identifying and preventing tens of critical bugs in different versions of our DNS authoritative engine from reaching production, with a porting effort of less than one person-week. Naiqian Zheng, Mengqi Liu 0001, Yuxing Xiang, Linjian Song, Nan Wang 0041, Zhuo Liang, Dennis Cai, Ennan Zhai, Xuanzhe Liu, Xin Jin 0008 |
SOSP | 11 |
| 2023 | Dependable Virtualized Fabric on Programmable Data PlaneabstractIn modern multi-tenant data centers, each tenant desires reassuring dependability from the virtualized network fabric – bandwidth guarantee with work conservation, bounded tail latency and resilient reachability. However, the slow convergence of prior works under network dynamics and uncertainties can hardly provide the dependability for tenants. Further, state-of-the-art load balance schemes are guarantee-agnostic and bring great risks on breaking bandwidth guarantee, which is overlooked in prior works. In this paper, we propose vFab, a dependable virtualized fabric framework which can (1) quickly detect network failure in data plane, (2) explicitly select proper paths for all flows, and (3) converge to ideal bandwidth allocation at sub-millisecond. The core idea of vFab is to leverage the programmable data plane to build a fusion of an active edge (e.g., NIC) and an informative core (e.g., switch), where the core sends link status and tenant information to the edge via telemetry to help the latter make a timely and accurate decision on path selection and traffic admission. We fully implement vFab with commodity SmartNICs and programmable switches. Extensive evaluations show that vFab can keep bandwidth guarantee with high bandwidth utilization, low and bounded latency, and resilient reachability under various network scenarios with limited overhead. Application-level experiments show that vFab can improve QPS by$2.4\times $and cut tail latency by$10\times $compared to the alternatives. Kaihui Gao, Shuai Wang 0028, Kun Qian 0021, Dan Li 0001, Rui Miao 0001, Bo Li 0061, Yu Zhou 0008, Ennan Zhai, Chen Sun 0005, Binzhang Fu, Frank Kelly, Dennis Cai, Hongqiang Harry Liu, Tao Sun 0010 |
IEEE/ACM Trans. Netw. | 8 |
| 2022 | Cetus: Releasing P4 Programmers from the Chore of Trial and Error Compiling
Ennan Zhai, Mengqi Liu 0001, Hongqiang Harry Liu |
NSDI | 3 |
| 2022 | Predictable vFabric on informative data planeabstractIn multi-tenant data centers, each tenant desires reassuring predictability from the virtual network fabric - bandwidth guarantee, work conservation, and bounded tail latency. Achieving these goals simultaneously relies on rapid and precise traffic admission. However, the slow convergence (tens of milliseconds) of prior works can hardly satisfy the increasingly rigorous performance demand under dynamic traffic patterns. Further, state-of-the-art load balance schemes are all guarantee-agnostic and bring great risks on breaking bandwidth guarantee, which is overlooked in prior works. Shuai Wang 0028, Kaihui Gao, Kun Qian 0021, Dan Li 0001, Rui Miao 0001, Bo Li 0061, Yu Zhou 0008, Ennan Zhai, Chen Sun 0005, Binzhang Fu, Frank Kelly, Dennis Cai, Hongqiang Harry Liu, Ming Zhang 0005 |
SIGCOMM | 8 |
| 2022 | Meissa: scalable network testing for programmable data planesabstractEnsuring the correctness of programmable data planes is important. Testing offers comprehensive correctness checking, including detecting both code bugs and non-code bugs. However, scalability is a key challenge for testing production-scale data planes to achieve high coverage. This paper presents Meissa, a scalable network testing system for programmable data planes with full path coverage. The core of Meissa is a domain-specific code summary technique that simplifies the control flow graph of a data plane program for scalable testing without sacrificing coverage. Code summary decomposes a data plane program into individual pipelines, and summarizes each pipeline with a succinct representation. We formally prove that Meissa with code summary achieves 100% path coverage. We use both open-source and production-scale data plane programs to evaluate Meissa. The evaluation shows that (i) Meissa is able to test production-scale data plane programs that cannot be supported by state-of-the-art efforts, and (ii) besides P4 code bugs, Meissa is able to not only identify known non-code bugs, but also detect previously-unknown non-code bugs. We also share in this paper several real cases tested by Meissa in a production programmable data plane. Naiqian Zheng, Mengqi Liu 0001, Ennan Zhai, Hongqiang Harry Liu, Kaicheng Yang 0001, Xuanzhe Liu, Xin Jin 0008 |
SIGCOMM | 3 |
| 2022 | Learning CI Configuration Correctness for Early Build FeedbackabstractContinuous Integration (CI) allows developers to check whether their code can build successfully and pass tests across various system environments with every commit. To use a CI platform, a developer must provide configuration files within a code repository to specify build conditions. Incorrect configuration settings lead to CI build failures, which can take hours to run, wasting valuable developer time and delaying product release dates. Debugging CI configurations is a slow and error-prone process. The only way to check the correctness of CI configurations is to push a commit and wait for the build result. We present VeriCI, the first system for localizing CI configuration errors at the code level. VeriCI runs as a static analysis tool, before the developer sends the build request to the CI server. Our key insight is that the commit history and the corresponding build histories available in CI environments can be used both for build error prediction and build error localization. We leverage the build history as a labeled dataset to automatically derive customized rules describing correct CI configurations, using supervised machine learning techniques. To more accurately identify root causes, we train a neural network that filters out constraints that are less likely to be connected to the root cause of build failure. We evaluate VeriCI on real world data from GitHub and achieve 91% accuracy of predicting a build failure and correctly identify the root cause in 75% of cases. We also conducted a between-subjects user study with 20 software developers, showing that VeriCI significantly helps users in identifying and fixing errors in CI. Mark Santolucito, Jialu Zhang 0002, Ennan Zhai, Jürgen Cito, Ruzica Piskac |
SANER | 3 |
| 2021 | Campion: debugging router configuration differencesabstractWe present a new approach for debugging two router configurations that are intended to be behaviorally equivalent. Existing router verification techniques cannot identify all differences or localize those differences to relevant configuration lines. Our approach addresses these limitations through a _modular_ analysis, which separately analyzes pairs of corresponding configuration components. It handles all router components that affect routing and forwarding, including configuration for BGP, OSPF, static routes, route maps and ACLs. Further, for many configuration components our modular approach enables simple _structural equivalence_ checks to be used without additional loss of precision versus modular semantic checks, aiding both efficiency and error localization. We implemented this approach in the tool Campion and applied it to debugging pairs of backup routers from different manufacturers and validating replacement of critical routers. Campion analyzed 30 proposed router replacements in a production cloud network and proactively detected four configuration bugs, including a route reflector bug that could have caused a severe outage. Campion also found multiple differences between backup routers from different vendors in a university network. These were undetected for three years, and depended on subtle semantic differences that the operators said they were "highly unlikely" to detect by "just eyeballing the configs." Alan Tang, Siva Kesava Reddy K., Ryan Beckett, Ennan Zhai, Matt Brown, Todd D. Millstein, Yuval Tamir, George Varghese |
SIGCOMM | 4 |
| 2021 | Aquila: a practically usable verification system for production-scale programmable data planesabstractThis paper presents Aquila, the first practically usable verification system for Alibaba's production-scale programmable data planes. Aquila addresses four challenges in building a practically usable verification: (1) specification complexity; (2) verification scalability; (3) bug localization; and (4) verifier self validation. Specifically, first, Aquila proposes a high-level language that facilitates easy expression of specifications, reducing lines of specification codes by tenfold compared to the state-of-the-art. Second, Aquila constructs a sequential encoding algorithm to circumvent the exponential growth of states associated with the upscaling of data plane programs to production level. Third, Aquila adopts an automatic and accurate bug localization approach that can narrow down suspects based on reported violations and pinpoint the culprit by simulating a fix for each suspect. Fourth and finally, Aquila can perform self validation based on refinement proof, which involves the construction of an alternative representation and subsequent equivalence checking. To this date, Aquila has been used in the verification of our production-scale programmable edge networks for over half a year, and it has successfully prevented many potential failures resulting from data plane bugs. Bingchuan Tian, Mengqi Liu 0001, Ennan Zhai, Yu Zhou 0008, Mengjing Ma, Xionglie Wei, Hongqiang Harry Liu, Ming Zhang 0005, Chen Tian 0001, Minlan Yu |
SIGCOMM | 4 |
| 2021 | Static detection of silent misconfigurations with deep interaction analysisabstractThe behavior of large systems is guided by their configurations: users set parameters in the configuration file to dictate which corresponding part of the system code is executed. However, it is often the case that, although some parameters are set in the configuration file, they do not influence the system runtime behavior, thus failing to meet the user’s intent. Moreover, such misconfigurations rarely lead to an error message or raising an exception. We introduce the notion of silent misconfigurations which are prohibitively hard to identify due to (1) lack of feedback and (2) complex interactions between configurations and code. This paper presents ConfigX, the first tool for the detection of silent misconfigurations. The main challenge is to understand the complex interactions between configurations and the code that they affected. Our goal is to derive a specification describing non-trivial interactions between the configuration parameters that lead to silent misconfigurations. To this end, ConfigX uses static analysis to determine which parts of the system code are associated with configuration parameters. ConfigX then infers the connections between configuration parameters by analyzing their associated code blocks. We design customized control- and data-flow analysis to derive a specification of configurations. Additionally, we conduct reachability analysis to eliminate spurious rules to reduce false positives. Upon evaluation on five real-world datasets across three widely-used systems, Apache, vsftpd, and PostgreSQL, ConfigX detected more than 2200 silent misconfigurations. We additionally conducted a user study where we ran ConfigX on misconfigurations reported on user forums by real-world users. ConfigX easily detected issues and suggested repairs for those misconfigurations. Our solutions were accepted and confirmed in the interaction with the users, who originally posted the problems. Jialu Zhang 0002, Ruzica Piskac, Ennan Zhai, Tianyin Xu |
Proc. ACM Program. Lang. | 3 |
| 2020 | Lock-Free Collaboration Support for Cloud Storage Services with Operation Inference and Transformation
Minghao Zhao 0001, Zhenhua Li 0001, Ennan Zhai, Feng Qian 0001, Yunhao Liu 0001, Tianyin Xu |
FAST | 4 |
| 2020 | Check before You Change: Preventing Correlated Failures in Service Updates
Ennan Zhai, Ang Chen 0001, Ruzica Piskac, Mahesh Balakrishnan 0001, Bingchuan Tian, Haoliang Zhang |
NSDI | 1 |
| 2020 | Lyra: A Cross-Platform Language and Compiler for Data Plane Programming on Heterogeneous ASICsabstractProgrammable data plane has been moving towards deployments in data centers as mainstream vendors of switching ASICs enable programmability in their newly launched products, such as Broadcom's Trident-4, Intel/Barefoot's Tofino, and Cisco's Silicon One. However, current data plane programs are written in low-level, chip-specific languages (e.g., P4 and NPL) and thus tightly coupled to the chip-specific architecture. As a result, it is arduous and error-prone to develop, maintain, and composite data plane programs in production networks. This paper presents Lyra, the first cross-platform, high-level language & compiler system that aids the programmers in programming data planes efficiently. Lyra offers a one-big-pipeline abstraction that allows programmers to use simple statements to express their intent, without laboriously taking care of the details in hardware; Lyra also proposes a set of synthesis and optimization techniques to automatically compile this "big-pipeline" program into multiple pieces of runnable chip-specific code that can be launched directly on the individual programmable switches of the target network. We built and evaluated Lyra. Lyra not only generates runnable real-world programs (in both P4 and NPL), but also uses up to 87.5% fewer hardware resources and up to 78% fewer lines of code than human-written programs. Ennan Zhai, Hongqiang Harry Liu, Rui Miao 0001, Yu Zhou 0008, Bingchuan Tian, Chen Sun 0005, Dennis Cai, Ming Zhang 0005, Minlan Yu |
SIGCOMM | 2 |
| 2020 | Accuracy, Scalability, Coverage: A Practical Configuration Verifier on a Global WANabstractThis paper presents Hoyan-- the first reported large scale deployment of configuration verification in a global-scale wide area network (WAN). Hoyan has been running in production for more than two years and is currently used for all critical configuration auditing and updates on the WAN. We highlight our innovative designs and real-life experience to make Hoyan accurate and scalable in practice. For accuracy under the inconsistencies of devices' vendor-specific behaviors (VSBs), Hoyan continuously discovers the flaws in device behavior models, thus aiding the operators in fixing the models. For scalability to verify our global WAN, Hoyan introduces a "global-simulation & local formal-modeling" strategy to model uncertainties in small scales and perform aggressive pruning of possibilities during the protocol simulations. Hoyan achieves near-100% verification accuracy after it detected and fixed O(10) VSBs on our WAN. Hoyan has prevented many potential service failures resulting from misconfiguration and reduced the failure rate of updates of our WAN by more than half in 2019. Fangdan Ye, Ennan Zhai, Hongqiang Harry Liu, Bingchuan Tian, Qiaobo Ye, Chunsheng Wang, Tianchen Guo, Duncheng She, Biao Cheng, Ming Zhang 0005, Rodrigo Fonseca |
SIGCOMM | 3 |
| 2020 | Automated repair by example for firewalls
William T. Hallahan, Ennan Zhai, Ruzica Piskac |
Formal Methods Syst. Des. | 2 |
| 2020 | PriFi: Low-Latency Anonymity for Organizational NetworksabstractOrganizational networks are vulnerable to trafficanalysis attacks that enable adversaries to infer sensitive information fromnetwork traffic—even if encryption is used. Typical anonymous communication networks are tailored to the Internet and are poorly suited for organizational networks.We present PriFi, an anonymous communication protocol for LANs, which protects users against eavesdroppers and provides high-performance traffic-analysis resistance. PriFi builds onDining Cryptographers networks (DC-nets), but reduces the high communication latency of prior designs via a new client/relay/server architecture, in which a client’s packets remain on their usual network path without additional hops, and in which a set of remote servers assist the anonymization process without adding latency. PriFi also solves the challenge of equivocation attacks, which are not addressed by related work, by encrypting traffic based on communication history. Our evaluation shows that PriFi introduces modest latency overhead (≈ 100ms for 100 clients) and is compatible with delay-sensitive applications such as Voice-over-IP. Ludovic Barman, Italo Dacosta, Mahdi Zamani, Ennan Zhai, Apostolos Pyrgelis, Bryan Ford, Joan Feigenbaum, Jean-Pierre Hubaux |
Proc. Priv. Enhancing Technol. | 4 |
| 2020 | HyCloud: Tweaking Hybrid Cloud Storage Services for Cost-Efficient Filesystem HostingabstractToday's cloud storage infrastructures typically provide two distinct types of services for hosting files: object storage like Amazon S3 and filesystem storage like Amazon EFS. In practice, a cloud storage user often desires the advantages of both-efficient filesystem operations with a low unit storage price. An intuitive approach to achieving this goal is to combine the two types of services, e.g., by hosting large files in S3 and small files together with directory structures in EFS. Unfortunately, our benchmark experiments indicate that the clients' download performance for large files becomes a severe system bottleneck. In this article, we attempt to address the bottleneck with little overhead by carefully tweaking the usages of S3 and EFS. Guided by two key observations, we design and implement an open-source system called HyCloud. It automatically invokes the data APIs of S3 and EFS on behalf of users, and intelligently schedules the data transfer among S3, EFS and the clients in a distributed manner. Real-world evaluations demonstrate that the unit storage price of HyCloud is close to that of S3, and the filesystem operations are executed as quickly as in EFS in most times (sometimes even more quickly than in EFS). Jinlong E, Yong Cui 0001, Zhenhua Li 0001, Mingkang Ruan, Ennan Zhai |
IEEE/ACM Trans. Netw. | 5 |
| 2019 | HyCloud: Tweaking Hybrid Cloud Storage Services for Cost-Efficient Filesystem HostingabstractToday's cloud storage infrastructures typically provide two distinct types of services for hosting files: object storage like Amazon S3 and filesystem storage like Amazon EFS. The former supports simple, flat object operations with a low unit storage price, while the latter supports complex, hierarchical filesystem operations with a high unit storage price. In practice, however, a cloud storage user often desires the advantages of both-efficient filesystem operations with a low unit storage price. An intuitive approach to achieving this goal is to combine the two types of services, e.g., by hosting large files in S3 and small files together with directory structures in EFS. Unfortunately, our benchmark experiments indicate that the clients' download performance for large files becomes a severe system bottleneck. In this paper, we attempt to address the bottleneck with little overhead by carefully tweaking the usages of S3 and EFS. This attempt is enabled by two key observations. First, since S3 and EFS have the same unit network-traffic price and the data transfer between S3 and EFS is free of charge, we can employ EFS as a relay for the clients' quickly downloading large files. Second, noticing that significant similarity exists between the files hosted at the cloud and its users, in most times we can convert large-size file downloads into small-size file synchronizations (through delta encoding and data compression). Guided by the observations, we design and implement an open-source system called HyCloud. It automatically invokes the data APIs of S3 and EFS on behalf of users, and handles the data transfer among S3, EFS and the clients. Real-world evaluations demonstrate that the unit storage price of HyCloud is close to that of S3, and the filesystem operations are executed as quickly as in EFS in most times (sometimes even more quickly than in EFS). Jinlong E, Yong Cui 0001, Mingkang Ruan, Zhenhua Li 0001, Ennan Zhai |
INFOCOM | 5 |
| 2019 | Mobile Gaming on Personal Computers with Direct Android EmulationabstractPlaying Android games on Windows x86 PCs has gained enormous popularity in recent years, and the de facto solution is to use mobile emulators built with the AOVB (Android-x86 On VirtualBox) architecture. When playing heavy 3D Android games with AOVB, however, users often suffer unsatisfactory smoothness due to the considerable overhead of full virtualization. This paper presents DAOW, a game-oriented Android emulator implementing the idea of direct Android emulation, which eliminates the overhead of full virtualization by directly executing Android app binaries on top of x86-based Windows. Based on pragmatic, efficient instruction rewriting and syscall emulation, DAOW offers foreign Android binaries direct access to the domestic PC hardware through Windows kernel interfaces, achieving nearly native hardware performance. Moreover, it leverages graphics and security techniques to enhance user experiences and prevent cheating in gaming. As of late 2018, DAOW has been adopted by over 50 million PC users to run thousands of heavy 3D Android games. Compared with AOVB, DAOW improves the smoothness by 21% on average, decreases the game startup time by 48%, and reduces the memory usage by 22%. Zhenhua Li 0001, Yunhao Liu 0001, Hai Long, Yuanchao Huang, Jiaming He, Tianyin Xu, Ennan Zhai |
MobiCom | 8 |
| 2019 | Demo: Mobile Gaming on Personal Computers with Direct Android EmulationabstractPlaying Android games with Windows x86 PCs is now popular, and the common solution is to use mobile emulators built with the AOVB (Android-x86 On VirtualBox) architecture. Nevertheless, running heavy 3D Android games on AOVB incurs considerable overhead of full virtualization, thus often leading to unsatisfactory smoothness. To tackle this issue, we present DAOW, a commercial game-oriented Android emulator implementing the idea of direct Android emulation, which eliminates the overhead of full virtualization by providing foreign Android binaries with direct access to the domestic PC hardware through Windows kernel interfaces. In this demo, we will demonstrate that DAOW essentially outperforms traditional AOVB-based emulators in terms of running smoothness, game startup time, and memory usage. Xinlei Yang, Zhenhua Li 0001, Yunhao Liu 0001, Guoyang Du, Ziwen Wu, Tianyin Xu, Ennan Zhai |
MobiCom | 9 |
| 2019 | Understanding Fileless Attacks on Linux-based IoT Devices with HoneyCloudabstractWith the wide adoption, Linux-based IoT devices have emerged as one primary target of today's cyber attacks. Traditional malware-based attacks can quickly spread across these devices, but they are well-understood threats with effective defense techniques such as malware fingerprinting and community-based fingerprint sharing. Recently, fileless attacks---attacks that do not rely on malware files---have been increasing on Linux-based IoT devices, and posing significant threats to the security and privacy of IoT systems. Little has been known in terms of their characteristics and attack vectors, which hinders research and development efforts to defend against them. In this paper, we present our endeavor in understanding fileless attacks on Linux-based IoT devices in the wild. Over a span of twelve months, we deploy 4 hardware IoT honeypots and 108 specially designed software IoT honeypots, and successfully attract a wide variety of real-world IoT attacks. We present our measurement study on these attacks, with a focus on fileless attacks, including the prevalence, exploits, environments, and impacts. Our study further leads to multi-fold insights towards actionable defense strategies that can be adopted by IoT vendors and end users. Fan Dang 0001, Zhenhua Li 0001, Yunhao Liu 0001, Ennan Zhai, Qi Alfred Chen, Tianyin Xu, Yan Chen 0004 |
MobiSys | 4 |
| 2019 | Understanding and Detecting Overlay-based Android Malware at Market ScalesabstractAs a key UI feature of Android, overlay enables one app to draw over other apps by creating an extra View layer on top of the host View. While greatly facilitating user interactions with multiple apps at the same time, it is often exploited by malicious apps (malware) to attack users. To combat this threat, prior countermeasures concentrate on restricting the capabilities of overlays at the OS level, while barely seeing adoption by Android due to the concern of sacrificing overlays' usability. To address this dilemma, a more pragmatic approach is to enable the early detection of overlay-based malware at the app market level during the app review process, so that all the capabilities of overlays can stay unchanged. Unfortunately, little has been known about the feasibility and effectiveness of this approach for lack of understanding of malicious overlays in the wild. To fill this gap, in this paper we perform the first large-scale comparative study of overlay characteristics in benign and malicious apps using static and dynamic analyses. Our results reveal a set of suspicious overlay properties strongly correlated with the malice of apps, including several novel features. Guided by the study insights, we build OverlayChecker, a system that is able to automatically detect overlay-based malware at market scales. OverlayChecker has been adopted by one of the world's largest Android app stores to check around 10K newly submitted apps per day. It can efficiently (within 2 minutes per app) detect nearly all (96%) overlay-based malware using a single commodity server. Yuxuan Yan, Zhenhua Li 0001, Qi Alfred Chen, Christo Wilson, Tianyin Xu, Ennan Zhai, Yong Li 0008, Yunhao Liu 0001 |
MobiSys | 6 |
| 2019 | Safely and automatically updating in-network ACL configurations with intent languageabstractIn-network Access Control List (ACL) is an important technique in ensuring network-wide connectivity and security. As cloud-scale WANs today constantly evolve in size and complexity, in-network ACL rules are becoming increasingly more complex. This presents a great challenge to the updating process of ACL configurations: network operators are frequently required to update "tangled" ACL rules across thousands of devices to meet diverse business requirements, and even a single ACL misconfiguration may lead to network disruptions. Such increasing challenges call for an automated system to improve the efficiency and correctness of ACL updates. This paper presents Jinjing, a system that aids Alibaba's network operators in automatically and correctly updating ACL configurations in Alibaba's global WAN. Jinjing allows the operators to express in a declarative language, named LAI, their update intent (e.g., ACL migration and traffic control). Then, Jinjing automatically synthesizes ACL update plans that satisfy their intent. At the heart of Jinjing, we develop a set of novel verification and synthesis techniques to rigorously guarantee the correctness of update plans. In Alibaba, our operators have used Jinjing to efficiently update their ACLs and have thus prevented significant service downtime. Bingchuan Tian, Xinyi Zhang 0003, Ennan Zhai, Hongqiang Harry Liu, Qiaobo Ye, Chunsheng Wang, Zhiming Ji, Yihong Sang, Ming Zhang 0005, Chen Tian 0001, Haitao Zheng 0001, Ben Y. Zhao |
SIGCOMM | 3 |
| 2019 | Pricing Data Tampering in Automated Fare Collection with NFC-Equipped SmartphonesabstractAutomated Fare Collection (AFC) systems have been globally deployed for decades, particularly in the public transportation network where the transit fee is calculated based on the length of the trip (a.k.a., distance-based pricing AFC systems). Although most messages of AFC systems are insecurely transferred in plaintext, system operators did not pay much attention to this vulnerability, since the AFC network is basically isolated from the public network (e.g., the Internet)-there is no way of exploiting such a vulnerability from the outside of the AFC network. Nevertheless, in recent years, the advent of Near Field Communication (NFC)-equipped smartphones has opened up a channel to invade into the AFC network from the mobile Internet, i.e., by Host-based Card Emulation (HCE) over NFC-equipped smartphones. In this paper, we identify a novel paradigm of attacks, called LessPay, against modern distance-based pricing AFC systems, enabling users to pay much less than what they are supposed to be charged. The identified attack has two important properties: 1) it is invisible to AFC system operators because the attack never causes any inconsistency in the back-end database of the operators; and 2) it can be scalable to affect a large number of users (e.g., 10,000) by only requiring a moderate-sized AFC card pool (e.g., containing 150 cards). To evaluate the efficacy of the attack, we developed an HCE app to launch the LessPay attack; and the real-world experiments demonstrate not only the feasibility of the LessPay attack (with 97.6 percent success rate) but also its low cost in terms of bandwidth and computation. Finally, we propose, implement and evaluate four types of countermeasures, and present security analysis and comparison of these countermeasures on defending against the LessPay attack. Fan Dang 0001, Ennan Zhai, Zhenhua Li 0001, David Mohaisen, Kaigui Bian, Qingfu Wen, Mo Li 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2018 | Towards Web-based Delta Synchronization for Cloud Storage Services
Zhenhua Li 0001, Ennan Zhai, Tianyin Xu, Yang Li 0092, Yunhao Liu 0001, Quanlu Zhang, Yao Liu 0001 |
FAST | 3 |
| 2018 | H2Cloud: Maintaining the Whole Filesystem in an Object Storage CloudabstractObject storage clouds (e.g., Amazon S3) have become extremely popular due to their highly usable interface and cost-effectiveness. They are, therefore, widely used by various applications (e.g., Dropbox) to host user data. However, because object storage clouds are flat and lack the concept of a directory, it becomes necessary to maintain file meta-data and directory structure in a separate index cloud. This paper investigates the possibility of using a single object storage cloud to efficiently host the whole filesystem for users, including both the file content and directories, while avoiding meta-data loss caused by index cloud failures. We design a novel data structure, Hierarchical Hash (or H2), to natively enable the efficient mapping from filesystem operations to object-level operations. Based on H2, we implement a prototype system, H2Cloud, that can maintain large filesystems of users in an object storage cloud and support fast directory operations. Both theoretical analysis and real-world experiments confirm the efficacy of our solution: H2Cloud achieves faster directory operations than OpenStack Swift by orders of magnitude, and has similar performance to Dropbox but yet does not need a separate index cloud. Minghao Zhao 0001, Zhenhua Li 0001, Ennan Zhai, Gareth Tyson, Chen Qian 0001, Zhenyu Li 0001, Leiyu Zhao |
ICPP | 3 |
| 2018 | Minimizing the Cask Effect of Multi-Source Content DeliveryabstractThis paper reveals the performance anomaly (i.e., the decline of delivery speed) when the client upgrades a task from single-source content delivery to multi-source content delivery. This anomaly is mainly caused by two aspects: (1) data sources with different types vary greatly in terms of acceleration reward (AR), and data sources with certain types are particularly easy to become inferior; (2) When the data sources remain fixed for a period of time, the large diversity of participant time (DPT) of data sources disturb the acceleration and the data sources with less participant time are inferior. Combing these insights, we figure out that the multi-source content delivery is limited by the so-called cask effect, i.e., the acceleration effect mainly depends on the inferior data sources. Zhenhua Li 0001, Zhenyu Li 0001, Tianyin Xu, Ennan Zhai, Yao Liu 0001, Minghao Zhao 0001, Yunhao Liu 0001 |
IWQoS | 5 |
| 2018 | On the Synchronization Bottleneck of OpenStack Swift-Like Cloud Storage SystemsabstractAs one type of the most popular cloud storage services, OpenStack Swift and its follow-up systems replicate each object across multiple storage nodes and leverageobject sync protocolsto achieve high reliability andeventual consistency. The performance of object sync protocols heavily relies on two key parameters:$r$(number of replicas for each object) and$n$(number of objects hosted by each storage node). In existing tutorials and demos, the configurations are usually$r=3$and$n<1,000$by default, and the sync process seems to perform well. However, we discover in data-intensive scenarios, e.g., when$r>3$and$n\gg 1,000$, the sync process is significantly delayed and produces massive network overhead, referred to as thesync bottleneck problem. By reviewing the source code of OpenStack Swift, we find that its object sync protocol utilizes a fairly simple and network-intensive approach to check the consistency among replicas of objects. Hence in a sync round, the number of exchanged hash values per node is$\Theta (n\times r)$. To tackle the problem, we propose a lightweight and practical object sync protocol,LightSync, which not only remarkably reduces the sync overhead, but also preserves high reliability and eventual consistency. LightSync derives this capability from three novel building blocks: 1)Hashing of Hashes, which aggregates all the$h$hash values of each data partition into a single but representative hash value with the Merkle tree; 2)Circular Hash Checking, which checks the consistency of different partition replicas by only sending the aggregated hash value to the clockwise neighbor; and 3)Failed Neighbor Handling, which properly detects and handles node failures with moderate overhead to effectively strengthen the robustness of LightSync. The design of LightSync offers provable guarantee on reducing the per-node network overhead from$\Theta (n\times r)$to$\Theta (\frac{n}{h})$. Furthermore, we have implemented LightSync as an open-source patch and adopted it to OpenStack Swift, thus reducing the sync delay by up to 879$\times$and the network overhead by up to 47.5$\times$. Mingkang Ruan, Thierry Titcheu Chekam, Ennan Zhai, Zhenhua Li 0001, Yao Liu 0001, Jinlong E, Yong Cui 0001, Hong Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2017 | Automated repair by example for firewallsabstractFirewalls are widely deployed to manage enterprise networks. Because enterprise-scale firewalls contain hundreds or thousands of rules, ensuring the correctness of firewalls - that the rules in the firewalls meet the specifications of their administrators - is an important but challenging problem. Although existing firewall diagnosis and verification techniques can identify potentially faulty rules, they offer administrators little or no help with automatically fixing faulty rules. This paper presents FireMason, the first effort that offers automated repair by example for firewalls. Once an administrator observes undesired behavior in a firewall, she may provide input/output examples that comply with the intended behaviors. Based on the examples, FireMason automatically synthesizes new firewall rules for the existing firewall. This new firewall correctly handles packets specified by the examples, while maintaining the rest of the behaviors of the original firewall. Through a conversion of the firewalls to SMT formulas, we offer formal guarantees that the change is correct. Our evaluation results from real-world case studies show that FireMason can efficiently find repairs. William T. Hallahan, Ennan Zhai, Ruzica Piskac |
FMCAD | 2 |
| 2017 | Practical Web-based Delta Synchronization for Cloud Storage Services
Zhenhua Li 0001, Ennan Zhai, Tianyin Xu |
HotStorage | 3 |
| 2017 | Large-scale invisible attack on AFC systems with NFC-equipped smartphonesabstractAutomated Fare Collection (AFC) systems have been globally deployed for decades, particularly in public transportation. Although the transaction messages of AFC systems are mostly transferred in plaintext, which is obviously insecure, system operators do not need to pay much attention to this issue, since the AFC network is well isolated from public network (e.g., the Internet). Nevertheless, in recent years, the advent of Near Field Communication (NFC)-equipped smartphones has bridged the gap between the AFC network and the Internet through Host-based Card Emulation (HCE). Motivated by this fact, we design and practice a novel paradigm of attack on modern distance-based pricing AFC systems, enabling users to pay much less than actually required. Our constructed attack has two important properties: 1) it is invisible to AFC system operators because the attack never causes any inconsistency in the backend database of the operators; and 2) it can be scalable to large number of users (e.g., 10,000) by maintaining a moderate-sized AFC card pool (e.g., containing 150 cards). Based upon this constructed attack, we developed an HCE app, named LessPay. Our real-world experiments on LessPay demonstrate not only the feasibility of our attack (with 97.6% success rate), but also its low-overhead in terms of bandwidth and computation. Fan Dang 0001, Zhenhua Li 0001, Ennan Zhai, David Mohaisen, Qingfu Wen, Mo Li 0001 |
INFOCOM | 4 |
| 2017 | Synthesizing configuration file specifications with association rule learningabstractSystem failures resulting from configuration errors are one of the major reasons for the compromised reliability of today's software systems. Although many techniques have been proposed for configuration error detection, these approaches can generally only be applied after an error has occurred. Proactively verifying configuration files is a challenging problem, because 1) software configurations are typically written in poorly structured and untyped “languages”, and 2) specifying rules for configuration verification is challenging in practice. This paper presents ConfigV, a verification framework for general software configurations. Our framework works as follows: in the pre-processing stage, we first automatically derive a specification. Once we have a specification, we check if a given configuration file adheres to that specification. The process of learning a specification works through three steps. First, ConfigV parses a training set of configuration files (not necessarily all correct) into a well-structured and probabilistically-typed intermediate representation. Second, based on the association rule learning algorithm, ConfigV learns rules from these intermediate representations. These rules establish relationships between the keywords appearing in the files. Finally, ConfigV employs rule graph analysis to refine the resulting rules. ConfigV is capable of detecting various configuration errors, including ordering errors, integer correlation errors, type errors, and missing entry errors. We evaluated ConfigV by verifying public configuration files on GitHub, and we show that ConfigV can detect known configuration errors in these files. Mark Santolucito, Ennan Zhai, Rahul Dhodapkar, Aaron Shim, Ruzica Piskac |
Proc. ACM Program. Lang. | 2 |
| 2017 | An auditing language for preventing correlated failures in the cloudabstractToday's cloud services extensively rely on replication techniques to ensure availability and reliability. In complex datacenter network architectures, however, seemingly independent replica servers may inadvertently share deep dependencies (e.g., aggregation switches). Such unexpected common dependencies may potentially result in correlated failures across the entire replication deployments, invalidating the efforts. Although existing cloud management and diagnosis tools have been able to offer post-failure forensics, they, nevertheless, typically lead to quite prolonged failure recovery time in the cloud-scale systems. In this paper, we propose a novel language framework, named RepAudit, that manages to prevent correlated failure risks before service outages occur, by allowing cloud administrators to proactively audit the replication deployments of interest. In particular, RepAudit consists of three new components: 1) a declarative domain-specific language, RAL, for cloud administrators to write auditing programs expressing diverse auditing tasks; 2) a high-performance RAL auditing engine that generates the auditing results by accurately and efficiently analyzing the underlying structures of the target replication deployments; and 3) an RAL-code generator that can automatically produce complex RAL programs based on easily written specifications. Our evaluation result shows that RepAudit uses 80x less lines of code than state-of-the-art efforts in expressing the auditing task of determining the top-20 critical correlated-failure root causes. To the best of our knowledge, RepAudit is the first effort capable of simultaneously offering expressive, accurate and efficient correlated failure auditing to the cloud-scale replication systems. Ennan Zhai, Ruzica Piskac, Ronghui Gu, Xun Lao |
Proc. ACM Program. Lang. | 1 |
| 2016 | Probabilistic Automated Language Learning for Configuration Files
Mark Santolucito, Ennan Zhai, Ruzica Piskac |
CAV (2) | 2 |
| 2016 | Building Privacy-Preserving Cryptographic Credentials from Federated Online IdentitiesabstractFederated identity providers, e.g., Facebook and PayPal, offer a convenient means for authenticating users to third-party applications. Unfortunately such cross-site authentications carry privacy and tracking risks. For example, federated identity providers can learn what applications users are accessing; meanwhile, the applications can know the users' identities in reality. John Maheswaran, Daniel Jackowitz, Ennan Zhai, David Wolinsky, Bryan Ford |
CODASPY | 3 |
| 2016 | On the synchronization bottleneck of OpenStack Swift-like cloud storage systemsabstractAs one type of the most popular cloud storage services, OpenStack Swift and its follow-up systems replicate each data object across multiple storage nodes and leverage object sync protocols to achieve high availability and eventual consistency. The performance of object sync protocols heavily relies on two key parameters: r (number of replicas for each object) and η (number of objects hosted by each storage node). In existing tutorials and demos, the configurations are usually r = 3 and n3 and n ≫ 1000, the object sync process is significantly delayed and produces massive network overhead. This phenomenon is referred to as the sync bottleneck problem. Then, to explore the root cause, we review the source code of OpenStack Swift and find that its object sync protocol utilizes a fairly simple and network-intensive approach to check the consistency among replicas of objects. In particular, each storage node is required to periodically multicast the hash values of all its hosted objects to all the other replica nodes. Thus in a sync round, the number of exchanged hash values per node is Θ(n×r). Further, to tackle the problem, we propose a lightweight object sync protocol called LightSync. It remarkably reduces the sync overhead by using two novel building blocks: 1) Hashing of Hashes, which aggregates all the h hash values of each data partition into a single but representative hash value with the Merkle tree; 2) Circular Hash Checking, which checks the consistency of different partition replicas by only sending the aggregated hash value to the clockwise neighbor. Its design provably reduces the per-node network overhead from Θ(n×r) to Θ(n/h). In addition, we have implemented LightSync as an open-source patch and adopted it to OpenStack Swift, thus reducing sync delay by up to 28.8× and network overhead by up to 14.2×. Thierry Titcheu Chekam, Ennan Zhai, Zhenhua Li 0001, Yong Cui 0001, Kui Ren 0001 |
INFOCOM | 2 |
| 2016 | AnonRep: Towards Tracking-Resistant Anonymous Reputation
Ennan Zhai, David Wolinsky, Ruichuan Chen, Ewa Syta, Chao Teng, Bryan Ford |
NSDI | 1 |
| 2016 | Resisting Tag Spam by Leveraging Implicit User BehaviorsabstractTagging systems are vulnerable to tag spam attacks. However, defending against tag spam has been challenging in practice, since adversaries can easily launch spam attacks in various ways and scales. To deeply understand users' tagging behaviors and explore more effective defense, this paper first conducts measurement experiments on public datasets of two representative tagging systems: Del.icio.us and CiteULike. Our key finding is that a significant fraction of correct tag-resource annotations are contributed by a small number of implicit similarity cliques, where users annotate common resources with similar tags. Guided by the above finding, we propose a new service, called Spam-Resistance-as-a-Service (or SRaaS), to effectively defend against heterogeneous tag spam attacks even at very large scales. At the heart of SRaaS is a novel reputation assessment protocol, whose design leverages the implicit similarity cliques coupled with the social networks inherent to typical tagging systems. With such a design, SRaaS manages to offer provable guarantees on diminishing the influence of tag spam attacks. We build an SRaaS prototype and evaluate it using a large-scale spam-oriented research dataset (which is much more polluted by tag spam than Del.icio.us and CiteULike datasets). Our evaluational results demonstrate that SRaaS outperforms existing tag spam defenses deployed in real-world systems, while introducing low overhead. Ennan Zhai, Zhenhua Li 0001, Zhenyu Li 0001, Fan Wu 0006, Guihai Chen |
Proc. VLDB Endow. | 1 |
| 2015 | A Risk-Evaluation Assisted System for Service SelectionabstractWith the rapid adoption of Service Oriented Architecture (SOA), increasingly more application-level services are developed through composing service components offered by different service providers. While such application development mode offers advantages in terms of cost-effectiveness and flexibility, application developers cannot understand or deal with risks potentially resulting from vulnerabilities within composed services due to non-transparency of the service providers. Furthermore, some of the vulnerabilities in practice are deeply hidden in dependency structures underlying composed services, thus making even the service providers fail to know the vulnerabilities. This paper proposes a risk-evaluation assisted service selection system, called Risk Evaluation-as-a-Service(or REaaS), which aims to assist application developers to understand vulnerability risks hidden within alternative services when the developers at first attempt to adopt their applications. In particular, for a given application developer's service selection requirement, REaaS produces a ranking list based upon vulnerability risks of alternative services to serve as a guideline regarding which service has the lowest potential risks (e.g., Bugs) for this application deployment. REaaS achieves this goal through the following three steps: 1) generating a package dependency graph for each alternative service, 2) assigning threat-degrees to packages in each dependency graph, and 3) analyzing each dependency graph and evaluating service-risk of each service. We have built a REaaS prototype and used real case study to demonstrate the practicality of REaaS. Ennan Zhai, Liang Gu, Yumei Hai |
ICWS | 1 |
| 2014 | Heading Off Correlated Failures through Independence-as-a-Service
Ennan Zhai, Ruichuan Chen, David Wolinsky, Bryan Ford |
OSDI | 1 |
| 2011 | SecGuard: Secure and Practical Integrity Protection Model for Operating Systems
Ennan Zhai, Qingni Shen, Tao Yang 0015, Liping Ding, Sihan Qing |
APWeb | 1 |
| 2011 | A Multi-compositional Enforcement on Information Flow Security
Cong Sun 0001, Ennan Zhai, Zhong Chen 0001, Jianfeng Ma 0001 |
ICICS | 2 |
| 2011 | Sorcery: Overcoming deceptive votes in P2P content sharing systems
Ennan Zhai, Huiping Sun, Sihan Qing, Zhong Chen 0001 |
Peer-to-Peer Netw. Appl. | 1 |
| 2010 | DSpam: Defending Against Spam in Tagging Systems via Users' ReliabilityabstractResisting spam in tagging system is very challenging. This paper presents DSpam, a novel spam-resistant tagging system which can significantly diminish spam in tag search results with users’ reliabilities. DSpam client groups other users into unfamiliar users and interacted users according to the fact whether the client has interacted with such users. For an unfamiliar user, the client computes his reliability by tagging behavior-based mechanism which reflects correlation of annotations between them. For an interacted user, the reliability includes two parts: feedback-based reliability, which indicates direct interactions between that user and the client, and recommendation reliability, which indicates the evaluation about that user from the client’s friends. The client ranks search result with the average reliabilities of himself with respect to annotators of each result. Experimental results show DSpam can effectively resist tag spam and work better than existing tag search schemes. Ennan Zhai, Cui Cao, Yongqiang Xie, Zhaojun Wang, Jian-bin Hu, Zhong Chen 0001 |
ICPADS | 2 |
| 2010 | SWORDS: Improving Sensor Networks Immunity under Worm Attacks
Nike Gui, Ennan Zhai, Jian-bin Hu, Zhong Chen 0001 |
WAIM | 2 |
| 2009 | Filtering Spam in Social Tagging System with Dynamic Behavior AnalysisabstractSpam in social tagging systems introduced by some malicious participants has become a serious problem for its global popularizing. Some studies which can be deduced to static user data analysis have been presented to combat tag spam, but either they do not give an exact evaluation or the algorithms' performances are not good enough. In this paper, we proposed a novel method based on analysis of dynamic user behavior data for the notion that users' behaviors in social tagging system can reflect the quality of tags more accurately. Through modeling the different categories of participants' behaviors, we extract tag-associated actions which can be used to estimate whether tag is spam, and then present our algorithm that can filter the tag spam in the results of social search. The experiment results show that our method indeed outperforms the existing methods based on static data and effectively defends against the tag spam in various spam attacks. Ennan Zhai, Huiping Sun, Yelu Chen, Zhong Chen 0001 |
ASONAM | 2 |
| 2009 | SpamResist: Making Peer-to-Peer Tagging Systems Robust to SpamabstractTagging systems are known to be particularly vulnerable to tag spam. Due to the self-organization and self-maintenance nature of Peer-to-Peer (P2P) overlay networks, users in the P2P tagging systems are more vulnerable to tag spam than the centralized ones. This paper proposes SpamResist, a novel social reliability-based mechanism. For each tag search, SpamResist client groups the search respondents into two categories, namely unfamiliar peers and interacted peers according to the fact whether the client has interacted with such respondents. For the two different categories of peers, the client computes their reliability degrees, and then utilizes these reliability degrees as weights to rank search results. To obtain higher quality search results, we propose a socially-enhanced mechanism, considering social friends can share their previous experience and help improve both the performance and convergence of SpamResist. Finally, the experimental results illustrate that SpamResist can effectively defend against tag spam and work better than the existing search models in P2P tagging systems. Ennan Zhai, Ruichuan Chen, Eng Keong Lua, Long Zhang 0003, Huiping Sun, Zhuhua Cai, Sihan Qing, Zhong Chen 0001 |
GLOBECOM | 1 |
| 2009 | Sorcery: Could We Make P2P Content Sharing Systems Robust to Deceivers?abstractDeceptive behaviors of peers in peer-to-peer (P2P) content sharing systems have become a serious problem due to the features of P2P overlay networks such as anonymity, self-organization, etc. This paper presents Sorcery, a novel active challenge-response mechanism based on the notion that one side of interaction with dominant information can detect whether the other side is telling a lie. To make each client obtain the dominant information, our approach introduces social network to the P2P content sharing system; thus, the client can establish friend-relationships with peers who are either acquaintances in reality or those reliable online friends. Using the confidential voting histories of friends as own dominant information, the client can challenge the content providers with the overlapping votes of both his friends and the content provider, thus detecting whether the content provider is a deceiver. Moreover, Sorcery provides the punishment mechanism which can reduce the impact brought by deceptive behaviors, and our work also discusses some key practical issues. The experimental results illustrate that Sorcery can effectively address the problem of deceptive behaviors, and work better than the existing reputation models. Ennan Zhai, Ruichuan Chen, Zhuhua Cai, Long Zhang 0003, Eng Keong Lua, Huiping Sun, Sihan Qing, Liyong Tang, Zhong Chen 0001 |
Peer-to-Peer Computing | 1 |