EDBT 2026 Demo / reviewers in the wild / expert
Shixiong Zhao
dblp:141/1483
· DBLP profile ↗
23ranked-venue papers
5as first author
16since 2021 · last 2026
0000-0002-1643-2583ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 4 first-author · 11 since 2021Security and privacy · 5 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 3 since 2021Computer networks · 3 · 1 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in ProductionabstractAs the foundational component of versatile AI applications, training an multimodal large language model (MLLM) relies on multimodal datasets with dynamic modality mixture proportions and sample length distributions. However, existing MLLM systems remain inefficient under dynamic workloads, due to statically coupled decisions of resource allocation and model parallelization between encoders and the LLM backbone. This paper presents MegaScale-Omni, an industrial-grade MLLM training system tailored for dynamic workload adaption and hyper-scale deployment. MegaScale-Omni is built upon the training scheme of encoder-LLM multiplexing with three key innovations: (1) Decoupled parallelism strategies with long-short sequence parallelism for encoders to process variable-length samples, and full-fledged 5D parallelism for the LLM backbone, both organized under a communication-efficient parallelization layout. (2) Unified encoder-LLM representations for flexible, extensible colocation, and a new paradigm of encoder-LLM joint pipeline with workload resilience. (3) Workload balancing techniques via decentralized grouped reordering in data loaders and adaptive resharding from encoder to LLM ranks. MegaScale-Omni is deployed as the foundation of our in-house large-scale MLLM training tasks with thousands of GPUs. Our experimental results demonstrate 1.27×–7.57× throughput improvement under production-grade dynamic workloads, as compared to four state-of-the-art systems. Chunyu Xue, Yangrui Chen, Jianyu Jiang, Ningxin Zheng, Junda Feng, Jingji Chen, Shixiong Zhao, Zanbo Wang, Lishu Luo, Faming Wu, Haibin Lin, Yanghua Peng, Xin Liu 0086, Quan Chen 0002 |
EuroSys | 7 |
| 2025 | PipeMesh: Achieving Memory-Efficient Computation-Communication Overlap for Training Large Language ModelsabstractEfficiently training large language models (LLMs) on commodity cloud resources remains challenging due to limitations in network bandwidth and accelerator memory capacity. Existing training systems can be categorized based on their pipeline schedules. Depth-first scheduling, employed by systems like Megatron, prioritizes memory efficiency but restricts the overlap between communication and computation, causing accelerators to remain idle for over 20% of the training time. Conversely, breadth-first scheduling maximizes communication overlap but generates excessive intermediate activations, exceeding memory capacity and slowing computation by more than 34%. To address these limitations, we propose a novel elastic pipeline schedule that enables fine-grained control over the trade-off between communication overlap and memory consumption. Our approach determines the number of micro-batches scheduled together according to the communication time and the memory available. Furthermore, we introduce a mixed sharding strategy and a pipeline-aware selective recomputation technique to reduce memory usage. Experimental results demonstrate that our system eliminates most of the 28% all-accelerator idle time caused by communication, with recomputation accounting for less than 1.9% of the training time. Compared to existing baselines,PIPEMESHimproves training throughput on commodity clouds by 20.1% to 33.8%. Fanxin Li, Shixiong Zhao, Yuhao Qing, Jianyu Jiang, Xusheng Chen, Heming Cui |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2024 | AGRNav: Efficient and Energy-Saving Autonomous Navigation for Air-Ground Robots in Occlusion-Prone EnvironmentsabstractThe exceptional mobility and long endurance of air-ground robots are raising interest in their usage to navigate complex environments (e.g., forests and large buildings). However, such environments often contain occluded and unknown regions, and without accurate prediction of unobserved obstacles, the movement of the air-ground robot often suffers a sub-optimal trajectory under existing mapping-based and learning-based navigation methods. In this work, we present AGRNav, a novel framework designed to search for safe and energy-saving air-ground hybrid paths. AGRNav contains a lightweight semantic scene completion network (SCONet) with self-attention to enable accurate obstacle predictions by capturing contextual information and occlusion area features. The framework subsequently employs a query-based method for low-latency updates of prediction results to the grid map. Finally, based on the updated map, the hierarchical path planner efficiently searches for energy-saving paths for navigation. We validate AGRNav’s performance through benchmarks in both simulated and real-world environments, demonstrating its superiority over classical and state-of-the-art methods. The open-source code is available at https://github.com/jmwang0117/AGRNav. Junming Wang 0001, Zekai Sun, Xiuxian Guan, Tianxiang Shen, Zongyuan Zhang, Tianyang Duan, Dong Huang 0005, Shixiong Zhao, Heming Cui |
ICRA | 8 |
| 2024 | Scene-wise Adaptive Network for Dynamic Cold-start Scenes Optimization in CTR PredictionabstractIn the realm of modern mobile E-commerce, providing users with nearby commercial service recommendations through location-based online services has become increasingly vital. While machine learning approaches have shown promise in multi-scene recommendation, existing methodologies often struggle to address cold-start problems in unprecedented scenes: the increasing diversity of commercial choices, along with the short online lifespan of scenes, give rise to the complexity of effective recommendations in online and dynamic scenes. In this work, we propose Scene-wise Adaptive Network (SwAN 1), a novel approach that emphasizes high-performance cold-start online recommendations for new scenes. Our approach introduces several crucial capabilities, including scene similarity learning, user-specific scene transition cognition, scene-specific information construction for the new scene, and enhancing the diverged logical information between scenes. We demonstrate SwAN’s potential to optimize dynamic multi-scene recommendation problems by effectively online handling cold-start recommendations for any newly arrived scenes. More encouragingly, SwAN has been successfully deployed in Meituan’s online catering recommendation service, which serves millions of customers per day, and SwAN has achieved a 5.64% CTR index improvement relative to the baselines and a 5.19% increase in daily order volume proportion. Jie Zhou 0029, Chuan Luo 0002, Shixiong Zhao |
RecSys | 6 |
| 2023 | Coorp: Satisfying Low-Latency and High-Throughput Requirements of Wireless Network for Coordinated Robotic LearningabstractIn coordinated robotic learning, multiple robots share the same wireless channel for communication, and bring together latency-sensitive (LS) network flows for control and bandwidth-hungry (BH) flows for distributed learning. Unfortunately, existing wireless network supporting systems cannot coordinate these two network flows to meet their own requirements: 1) prioritized contention systems (e.g., EDCA) prevent LS messages from timely acquiring the wireless channel because multiple wireless network interface cards (WNICs) with BH messages are contending for the channel 2) global planning systems (e.g., SchedWiFi) have to reserve a notable time window in the shared channel for each LS flow, suffering from severe bandwidth degradation (up to 42%). We present the coordinated preemption method to meet both requirements for LS flows and BH flows. Globally (among multiple robots), coordinated preemption eliminates unnecessary contention of BH flows by making them transmit in a round-robin manner, such that LS flows have the highest chance to win the contention against BH flows, without sacrificing overall bandwidth from the perspective of coordinated robotic learning applications. Locally (within the same robot), coordinated preemption in real time predicts the periodic transmission of LS flows from the upper application and conservatively limits packets of BH flows buffered in the WNIC only before LS packets arriving, reducing the bandwidth devoted to preemption. COORP, our implementation of coordinated preemption, reduced the violation of latency requirements from 53.9% (EDCA) to 8.8% (comparable to SchedWiFi). Regarding learning quality, COORP achieved a comparable (at times the same) learning reward with EDCA, which grew up to 76% faster than SchedWiFi. Shengliang Deng, Xiuxian Guan, Zekai Sun, Shixiong Zhao, Tianxiang Shen, Xusheng Chen, Tianyang Duan, Jia Pan 0001, Libo Zhang 0001, Heming Cui |
IEEE Internet Things J. | 4 |
| 2023 | Fold3D: Rethinking and Parallelizing Computational and Communicational Tasks in the Training of Large DNN ModelsabstractTraining a large DNN (e.g., GPT3) efficiently on commodity clouds is challenging even with the latest 3D parallel training systems (e.g., Megatron v3.0). In particular, along the pipeline parallelism dimension, computational tasks that produce a whole DNN's gradients with multiple input batches should be concurrently activated; along the data parallelism dimension, a set of heavy-weight communications (for aggregating the accumulated outputs of computational tasks) isinevitably serializedafter the pipelined tasks, undermining the training performance (e.g., in Megatron, data parallelism caused all GPUs idle for over 44% of the training time) over commodity cloud networks. To deserialize these communicational and computational tasks, we propose the AIAO scheduling (for 3D parallelism) which slices a DNN into multiple segments, so that the computational tasks processing the same DNN segment can be scheduled together, and the communicational tasks that synchronize this segment can be launched and overlapped (deserialized) with other segments’ computational tasks. We realized this idea in ourFold3Dtraining system. Extensive evaluation showsFold3Deliminated most of the all-GPU 44% idle time in Megatron (caused by data parallelism), leading to 25.2%–42.1% training throughput improvement compared to four notable baselines over various settings;Fold3D's high performance scaled to many GPUs. Fanxin Li, Shixiong Zhao, Yuhao Qing, Xusheng Chen, Xiuxian Guan, Sen Wang 0004, Gong Zhang 0001, Heming Cui |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2023 | A Geography-Based P2P Overlay Network for Fast and Robust Blockchain SystemsabstractNumerous blockchain systems with various consensus protocols have emerged to achieve high transaction rates (2$\sim$10K tps). However, their underlying P2P network primitives constrain further improvements due to two problems (i) high message redundancy and (ii) long broadcast convergence time. The first problem is caused by the excessive robustness of the dominant broadcast approach Gossip. All state-of-the-art blockchain systems only tolerate 20-50% node failure while Gossip can withstand up to 90%. The reason for (ii) is that existing broadcast topologies ignore geographical distances among nodes and incur paths with unnecessarily high latency. We presentFRing, a geography-based P2P overlay network for fast and robust broadcast in blockchain systems.FRinghas three main features: sufficient robustness, low message redundancy, and fast convergence. To reduce convergence time,FRingforms the network topology by considering geographical proximity. A novel broadcast algorithm based onFRingtopology is proposed to lower message redundancy while maintaining sufficient robustness. One major challenge is to eliminate the risk of topology inference by traffic pattern analysis.FRingleverages Intel SGX to guarantee nodes’ behavior integrity and incorporates pattern obfuscation to prevent traffic pattern analysis. The evaluation shows thatFRingimproved the throughput of EOS by 220% and Hyperledger Fabric by 210%. Haoran Qiu, Shixiong Zhao, Xusheng Chen, Ji Qi 0002, Heming Cui, Sen Wang 0004 |
IEEE Trans. Serv. Comput. | 3 |
| 2022 | NASPipe: high performance and reproducible pipeline parallel supernet training via causal synchronous parallelismabstractSupernet training, a prevalent and important paradigm in Neural Architecture Search, embeds the whole DNN architecture search space into one monolithic supernet, iteratively activates a subset of the supernet (i.e., a subnet) for fitting each batch of data, and searches a high-quality subnet which meets specific requirements. Although training subnets in parallel on multiple GPUs is desirable for acceleration, there inherently exists a race hazard that concurrent subnets may access the same DNN layers. Existing systems support neither efficiently parallelizing subnets’ training executions, nor resolving the race hazard deterministically, leading to unreproducible training procedures and potentiallly non-trivial accuracy loss. Shixiong Zhao, Fanxin Li, Xusheng Chen, Tianxiang Shen, Li Chen 0008, Sen Wang 0004, Nicholas Zhang, Cheng Li 0001, Heming Cui |
ASPLOS | 1 |
| 2022 | ROG: A High Performance and Robust Distributed Training System for Robotic IoTabstractCritical robotic tasks such as rescue and disaster response are more prevalently leveraging ML (Machine Learning) models deployed on a team of wireless robots, on which data parallel (DP) training over Internet of Things of these robots (robotic IoT) can harness the distributed hardware resources to adapt their models to changing environments as soon as possible. Unfortunately, due to the need for DP synchronization across all robots, the instability in wireless networks (i.e., fluctuating bandwidth due to occlusion and varying communication distance) often leads to severe stall of robots, which affects the training accuracy within a tight time budget and wastes energy stalling. Existing methods to cope with the instability of datacenter networks are incapable of handling such straggler effect. That is because they are conducting model-granulated transmission scheduling, which is much more coarse-grained than the granularity of transient network instability in real-world robotic IoT networks, making a previously reached schedule mismatch with the varying bandwidth during transmission. We present ROG, the first ROw-Granulated distributed training system optimized for ML training over unstable wireless networks. ROG confines the granularity of transmission and synchronization to each row of a layer’s parameters and schedules the transmission of each row adaptively to the fluctuating bandwidth. In this way the ML training process can update partial and the most important gradients of a stale robot to avoid triggering stalls, while provably guaranteeing convergence. The evaluation shows that, given the same training time, ROG achieved about 4.9%~6.5% training accuracy gain compared with the baselines and saved 20.4%~50.7% of the energy to achieve the same training accuracy. Xiuxian Guan, Zekai Sun, Shengliang Deng, Xusheng Chen, Shixiong Zhao, Zongyuan Zhang, Tianyang Duan, Chenshu Wu, Yong Cui 0001, Libo Zhang 0001, Rui Wang 0007, Heming Cui |
MICRO | 5 |
| 2022 | CRONUS: Fault-isolated, Secure and High-performance Heterogeneous Computing for Trusted Execution EnvironmentabstractWith the trend of processing a large volume of sensitive data on PaaS services (e.g., DNN training), a TEE architecture that supports general heterogeneous accelerators, enables spatial sharing on one accelerator, and enforces strong isolation across accelerators is highly desirable. However, none of the existing TEE solutions meet all three requirements. In this paper, we propose CRONUS, the first TEE architecture that achieves the three crucial requirements. The key idea of CRONUS is to partition heterogeneous computation into isolated TEE enclaves, where each enclave encapsulates only one kind of computation (e.g., GPU computation), and multiple enclaves can spatially share an accelerator. Then, CRONUS constructs heterogeneous computing using remote procedure calls (RPCs) among enclaves. With CRONUS, each accelerator’s hardware and its software stack are strongly isolated from others’, and each enclave trusts only its own hardware. To tackle the security challenge caused by inter-enclave interactions, we design a new streaming remote procedure call abstraction to enable secure RPCs with high performance. CRONUS is software-based, making it general to diverse accelerators. We implemented CRONUS on ARM TrustZone. Evaluation on diverse workloads with CPUs, GPUs and NPUs shows that, CRONUS achieves less than 7.1% extra computation time compared to native (unprotected) executions. Jianyu Jiang, Ji Qi 0002, Tianxiang Shen, Xusheng Chen, Shixiong Zhao, Sen Wang 0004, Li Chen 0008, Gong Zhang 0001, Xiapu Luo, Heming Cui |
MICRO | 5 |
| 2022 | SOTER: Guarding Black-box Inference for General Neural Networks at the Edge
Tianxiang Shen, Ji Qi 0002, Jianyu Jiang, Siyuan Wen, Xusheng Chen, Shixiong Zhao, Sen Wang 0004, Li Chen 0008, Xiapu Luo, Fengwei Zhang, Heming Cui |
USENIX ATC | 7 |
| 2022 | Efficient and DoS-resistant Consensus for Permissioned Blockchains
Xusheng Chen, Shixiong Zhao, Ji Qi 0002, Jianyu Jiang, Haoze Song, Cheng Wang 0021, Tsz On Li, T.-H. Hubert Chan, Fengwei Zhang, Xiapu Luo, Sen Wang 0004, Gong Zhang 0001, Heming Cui |
Perform. Evaluation | 2 |
| 2022 | DAENet: Making Strong Anonymity Scale in a Fully Decentralized NetworkabstractTraditional anonymous networks (e.g., Tor) are vulnerable to traffic analysis attacks that monitor the whole network traffic to determine which users are communicating. To preserve user anonymity against traffic analysis attacks, the emerging mix networks mess up the order of packets through a set of centralized and explicit shuffling nodes. However, this centralized design of mix networks is insecure against targeted DoS attacks that can completely block these shuffling nodes. In this article, we presentDAENet, an efficient mix network that resists both targeted DoS attacks and traffic analysis attacks with a new abstraction calledStealthy Peer-to-Peer (P2P) Network. Thestealthy P2P networkeffectively hides the shuffling nodes used in a routing path into the whole network, such that adversaries cannot distinguish specific shuffling nodes and conduct targeted DoS attacks to block these nodes. In addition, to handle traffic analysis attacks, we leverage the confidentiality and integrity protection of Intel SGX to ensure trustworthy packet shuffles at each distributed host and use multiple routing paths to prevent adversaries from tracking and revealing user identities. We show that our system is scalable with moderate latency (2.2s) when running in a cluster of 10,000 participants and is robust in the case of machine failures, making it an attractive new design for decentralized anonymous communication. DAENet ’s code is released onhttps://github.com/hku-systems/DAENet. Tianxiang Shen, Jianyu Jiang, Yunpeng Jiang, Xusheng Chen, Ji Qi 0002, Shixiong Zhao, Fengwei Zhang, Xiapu Luo, Heming Cui |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2022 | vPipe: A Virtualized Acceleration System for Achieving Efficient and Scalable Pipeline Parallel DNN TrainingabstractThe increasing computational complexity of DNNs achieved unprecedented successes in various areas such as machine vision and natural language processing (NLP), e.g., the recent advanced Transformer has billions of parameters. However, as large-scale DNNs significantly exceed GPU's physical memory limit, they cannot be trained by conventional methods such as data parallelism. Pipeline parallelism that partitions a large DNN into small subnets and trains them on different GPUs is a plausible solution. Unfortunately, the layer partitioning and memory management in existing pipeline parallel systems are fixed during training, making them easily impeded by out-of-memory errors and the GPU under-utilization. These drawbacks amplify when performing neural architecture search (NAS) such as the evolved Transformer, where different network architectures of Transformer needed to be trained repeatedly. vPipe is the first system that transparently provides dynamic layer partitioning and memory management for pipeline parallelism. vPipe has two unique contributions, including (1) an online algorithm for searching a near-optimal layer partitioning and memory management plan, and (2) a live layer migration protocol for re-balancing the layer distribution across a training pipeline. vPipe improved the training throughput of two notable baselines (Pipedream and GPipe) by 61.4-463.4 percent and 24.8-291.3 percent on various large DNNs and training settings. Shixiong Zhao, Fanxin Li, Xusheng Chen, Xiuxian Guan, Jianyu Jiang, Dong Huang 0005, Yuhao Qing, Sen Wang 0004, Peng Wang 0037, Gong Zhang 0001, Cheng Li 0001, Ping Luo 0002, Heming Cui |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | MVSAS: Semantic-Aware Scheduling for Low Latency and High Precision in Wireless Multi-View ApplicationabstractMulti-view models for various multi-view applications (e.g., pose recognition, facial recognition) achieve higher accuracy when more sensing data (views) from different sensors are flexibly collected via wireless networks and combined into inference input. However, when the view number scales up, the application suffers a long latency to collect all the latest views before inference (vanilla workflow). We observed that collecting all the latest views before inference is unnecessary, because different views are often not equally important and important views have major contribution to the output. In this paper, we present a Multi-View Semantic-Aware Scheduling (MVSAS) system that automatically prioritizes views according to their importance and schedules the early transmission of the important views. We tackled the challenge to infer view importance by analyzing the inference intermediates and extracting the semantics (e.g., number of persons) of each view. Once important views are collected, needless to wait for other less important views, the important views are combined with stale version of less important views as inference input, so as to retain high accuracy while reducing the latency to collect views. Evaluation shows that MVSAS achieved at most 36.9% latency reduction while retaining at most 98.7% accuracy compared to the vanilla workflow. Xiuxian Guan, Zekai Sun, Shengliang Deng, Shixiong Zhao, Tianxiang Shen, Tsz On Li, Rui Wang 0007, Heming Cui |
ICPADS | 4 |
| 2021 | Bidl: A High-throughput, Low-latency Permissioned Blockchain Framework for Datacenter NetworksabstractA permissioned blockchain framework typically runs an efficient Byzantine consensus protocol and is attractive to deploy fast trading applications among a large number of mutually untrusted participants (e.g., companies). Unfortunately, all existing permissioned blockchain frameworks adopt sequential workflows for invoking the consensus protocol and executing applications' transactions, making the performance of these applications much lower than deploying them in traditional systems (e.g., in-datacenter stock exchange). Ji Qi 0002, Xusheng Chen, Yunpeng Jiang, Jianyu Jiang, Tianxiang Shen, Shixiong Zhao, Sen Wang 0004, Gong Zhang 0001, Li Chen 0008, Man Ho Au, Heming Cui |
SOSP | 6 |
| 2020 | Uranus: Simple, Efficient SGX Programming and its ApplicationsabstractApplications written in Java have strengths to tackle diverse threats in public clouds, but these applications are still prone to privileged attacks when processing plaintext data. Intel SGX is powerful to tackle these attacks, and traditional SGX systems rewrite a Java application's sensitive functions, which process plaintext data, using C/C++ SGX API. Although this code-rewrite approach achieves good efficiency and a small TCB, it requires SGX expert knowledge and can be tedious and error-prone. To tackle the limitations of rewriting Java to C/C++, recent SGX systems propose a code-reuse approach, which runs a default JVM in an SGX enclave to execute the sensitive Java functions. However, both recent study and this paper find that running a default JVM in enclaves incurs two major vulnerabilities, Iago attacks, and control flow leakage of sensitive functions, due to the usage of OS features in JVM. In this paper, Uranus creates easy-to-use Java programming abstractions for application developers to annotate sensitive functions, and Uranus automatically runs these functions in SGX at runtime. Uranus effectively tackles the two major vulnerabilities in the code-reuse approach by presenting two new protocols: 1) a Java bytecode attestation protocol for dynamically loaded functions; and 2) an OS-decoupled, efficient GC protocol optimized for data-handling applications running in enclaves. We implemented Uranus in Linux and applied it to two diverse data-handling applications: Spark and ZooKeeper. Evaluation shows that: 1) Uranus achieves the same security guarantees as two relevant SGX systems for these two applications with only a few annotations; 2) Uranus has reasonable performance overhead compared to the native, insecure applications; and 3) Uranus defends against privileged attacks. Uranus source code and evaluation results are released on https://github.com/hku-systems/uranus. Jianyu Jiang, Xusheng Chen, Tsz On Li, Cheng Wang 0021, Tianxiang Shen, Shixiong Zhao, Heming Cui, Cho-Li Wang, Fengwei Zhang |
AsiaCCS | 6 |
| 2020 | HAMS: High Availability for Distributed Machine Learning Service GraphsabstractMission-critical services often deploy multiple Machine Learning (ML) models in a distributed graph manner, where each model can be deployed on a distinct physical host. Practical fault tolerance for such ML service graphs should meet three crucial requirements: high availability (fast failover), low normal case performance overhead, and global consistency under non-determinism (e.g., threads in a GPU can do floating point additions in random order). Unfortunately, despite much effort, existing fault tolerance systems, including those taking the primary-backup approach or the checkpoint-replay approach, cannot meet all these three requirements. To tackle this problem, we present HAMS, which starts from the primary-backup approach to replicate each stateful ML model, and we leverage the causal logging technique from the checkpoint-replay approach to eliminate the notorious stop-and-buffer delay in the primary-backup approach. Extensive evaluation on 25 ML models and six ML services shows that: (1) in normal case, HAMS achieved 0.5%-3.7% overhead on latency compared with bare metal; (2) HAMS took 116.12ms-254.19ms to recover one stateful model in all services, 155.1X-1067.9X faster than a relevant system Lineage Stash (LS); and (3) HAMS recovered these services with global consistency even when the GPU non-determinism exists, not supported by LS. HAMS's code is released ongithub.com/hku-systems/hams. Shixiong Zhao, Xusheng Chen, Cheng Wang 0021, Fanxin Li, Heming Cui, Cheng Li 0001, Sen Wang 0004 |
DSN | 1 |
| 2019 | NFVactor: A Resilient NFV System Using the Distributed Actor ModelabstractResilience functionality, including failure resilience and flow migration, is of pivotal importance in practical network function virtualization (NFV) systems. However, existing failure recovery procedures incur high packet processing delay due to heavyweight process checkpointing, while flow migration has poor performance due to centralized control. This paper proposes NFVactor, a novel NFV system that aims to provide lightweight failure resilience and high-performance flow migration. NFVactorenables these by using actor model to provide a per-flow execution environment, so that each flow can replicate and migrate itself with improved parallelism, while the efficiency of the actor model is guaranteed by a carefully designed runtime system. Moreover, NFVactorachieves transparent resilience: once a new network function (NF) is implemented for NFVactor, the NF automatically acquires resilience support. Our evaluation result shows that NFVactorachieves 10-Gbps packet processing, flow migration completion time that is 144 times faster than the existing system, and packet processing delay stabilized at around 20 μs during replication. Jingpu Duan, Xiaodong Yi 0001, Shixiong Zhao, Chuan Wu 0001, Heming Cui, Franck Le |
IEEE J. Sel. Areas Commun. | 3 |
| 2018 | OWL: Understanding and Detecting Concurrency AttacksabstractJust like bugs in single-threaded programs can lead to vulnerabilities, bugs in multithreaded programs can also lead to concurrency attacks. We studied 31 real-world concurrency attacks, including privilege escalations, hijacking code executions, and bypassing security checks. We found that compared to concurrency bugs' traditional consequences (e.g., program crashes), concurrency attacks' consequences are often implicit, extremely hard to be observed and diagnosed by program developers. Moreover, in addition to bug-inducing inputs, extra subtle inputs are often needed to trigger the attacks. These subtle features make existing tools ineffective to detect concurrency attacks. To tackle this problem, we present OWL, the first practical tool that models general concurrency attacks' implicit consequences and automatically detects them. We implemented OWL in Linux and successfully detected five new concurrency attacks, including three confirmed and fixed by developers, and two exploited from previously known and well-studied concurrency bugs. OWL has also detected seven known concurrency attacks. Our evaluation shows that OWL eliminates 94.1% of the reports generated by existing concurrency bug detectors as false positive, greatly reducing developers' efforts on diagnosis. All OWL source code, concurrency attack exploit scripts, and results are available on github.com/hku-systems/owl. Shixiong Zhao, Haoran Qiu, Tsz On Li, Heming Cui |
DSN | 1 |
| 2018 | PLOVER: Fast, Multi-core Scalable Virtual Machine Fault-tolerance
Cheng Wang 0021, Xusheng Chen, Weiwei Jia 0001, Boxuan Li, Haoran Qiu, Shixiong Zhao, Heming Cui |
NSDI | 6 |
| 2017 | Kakute: A Precise, Unified Information Flow Analysis System for Big-data SecurityabstractBig-data frameworks (e.g., Spark) enable computations on tremendous data records generated by third parties, causing various security and reliability problems such as information leakage and programming bugs. Existing systems for big-data security (e.g., Titian) track data transformations in a record level, so they are imprecise and too coarse-grained for these problems. For instance, when we ran Titian to drill down input records that produced a buggy output record, Titian reported 3 to 9 orders of magnitude more input records than the actual ones. Information Flow Tracking (IFT) is a conventional approach for precise information control. However, extant IFT systems are neither efficient nor complete for big-data frameworks, because theses frameworks are data-intensive, and data flowing across hosts is often ignored by IFT. Jianyu Jiang, Shixiong Zhao, Danish Alsayed, Heming Cui, Feng Liang 0004, Zhaoquan Gu |
ACSAC | 2 |
| 2015 | Towards Effective Developer Recommendation in Software CrowdsourcingabstractCrowdsourcing has attracted increasing attention from both industry and academia since it was proposed.Now a lot of work is finished by crowdsourcing, such as logo design, website promotion, industrial design, copywriting, software development, translation and image annotation.Although software crowdsourcing achieves positive results in practice, we still face a challenge of assigning suitable developers to specific tasks.In this paper, we propose a novel approach that recommends developers.In particular, our approach supports: comprehensively measuring the tasks and developers in software crowdsourcing, and recommending developers on the basis of the developer-task competence, task-task similarity, and soft power. Shixiong Zhao, Beijun Shen, Yuting Chen 0001, Hao Zhong 0001 |
SEKE | 1 |