Ming Wu 0007

dblp:32/2572-7 · DBLP profile ↗
← Back
32ranked-venue papers
2as first author
8since 2021 · last 2025
0009-0001-6663-4861ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 17 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 10 · 1 first-author · 4 since 2021Computer networks · 4 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 CCAgent: Coordinating Collaborative Data Scaling for Operating System Agents via Web3
Liang Chen 0024, Haozhe Zhao, Yinzhen Huang, Tsekai Lin, Weichu Xie, Peiyi Wang, Runxin Xu, Ming Wu 0007, Baobao Chang
CIKM10
2025 HEMVM: A Heterogeneous Blockchain Framework for Interoperable Virtual Machines
abstract
This paper introduces HEMVM, an innovative heterogeneous blockchain framework that seamlessly integrates diverse virtual machines (VMs), including the Ethereum Virtual Machine (EVM) and the Move Virtual Machine (MoveVM), into a unified system. This integration facilitates interoperability while retaining compatibility with existing Ethereum and Move toolchains by preserving high-level language constructs. HEMVM's unique cross-VM operations allow users to interact with contracts across various VMs using any wallet software, effectively resolving the fragmentation in user experience caused by differing VM designs. Our experimental results demonstrate that HEMVM is both fast and efficient, incurring minimal overhead (less than 4.4 %) for intra-VM transactions and achieving up to 9300 TPS for cross-VM transactions. Our results also show that the cross-VM operations in HEMVM are sufficiently expressive to support complex decentralized finance interactions across multiple VMs. Finally, the parallelized prototype of HEMVM shows performance improvements up to 44.8 % compared to the sequential version of HEMVM under workloads with mixed transaction types.
Vladyslav Nekriach, Sidi Mohamed Beillahi, Chenxing Li, Peilun Li, Ming Wu 0007, Andreas G. Veneris, Fan Long
Proc. ACM Program. Lang.5
2025 Fine-Grained Structured Sparse Computing for FPGA-Based AI Inference
abstract
With the explosive growth in the number of parameters in deep neural networks (DNNs), sparsity-centric algorithm and hardware designs have become critical for low-latency AI serving systems. However, the inherent randomness in pruning methods often leads to fragmented data access and irregular computation patterns in sparse matrices, resulting in significantly reduced hardware efficiency. Addressing the balance between the ‘randomness’ required to maintain model accuracy and the ‘regularity’ needed for efficient hardware design is crucial for realizing effective sparse computing in AI. This article proposes a fine-grained structured sparsity (FSS) paradigm. The pruned sparse matrices in this paradigm exhibit characteristics of ‘local randomness’ and ‘global regularity’. This dual-feature design allows AI accelerator hardware based on the FSS paradigm to maintain both high model accuracy and efficient hardware design. We implemented this novel accelerator on the Xilinx Alveo U280 and validated our concept across three different AI models, including CNN, RNN, and LLM, demonstrating performance that significantly outperforms prior methods.
Chen Zhang 0001, Shijie Cao, Guohao Dai 0001, Chenbo Geng, Zhuliang Yao, Wencong Xiao, Yunxin Liu 0001, Ming Wu 0007, Guangyu Sun 0003, Zhigang Ji, Runsheng Wang, Ru Huang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.8
2024 LVMT: An Efficient Authenticated Storage for Blockchain
abstract
Authenticated storage access is the performance bottleneck of a blockchain, because each access can be amplified to potentially O (log n ) disk I/O operations in the standard Merkle Patricia Trie (MPT) storage structure. In this article, we propose a multi-Layer Versioned Multipoint Trie (LVMT), a novel high-performance blockchain storage with significantly reduced I/O amplifications. LVMT uses the authenticated multipoint evaluation tree vector commitment protocol to update commitment proofs in constant time. LVMT adopts a multi-layer design to support unlimited key–value pairs and stores version numbers instead of value hashes to avoid costly elliptic curve multiplication operations. In our experiment, LVMT outperforms the MPT in real Ethereum traces, delivering read and write operations 6× faster. It also boosts blockchain system execution throughput by up to 2.7×.
Chenxing Li, Sidi Mohamed Beillahi, Guang Yang 0020, Ming Wu 0007, Wei Xu 0005, Fan Long
ACM Trans. Storage4
2023 Mercury: Fast Transaction Broadcast in High Performance Blockchain Systems
abstract
Blockchain systems must be secure and offer high performance. These systems rely on transaction broadcast mechanisms to provide both of these features. Unfortunately, in today’s systems, the broadcast mechanisms are highly inefficient.We present Mercury, a new transaction broadcast protocol designed for high performance blockchains. Mercury shortens the transaction propagation delay using two techniques: a virtual coordinate system and an early outburst strategy. Simulation results show that Mercury outperforms prior propagation schemes and decreases overall propagation latency by up to 44%. When implemented in Conflux, an open-source high-throughput blockchain system, Mercury reduces transaction propagation latency by over 50% with less than 5% bandwidth overhead.
Mingxun Zhou, Liyi Zeng, Peilun Li, Fan Long, Dong Zhou 0006, Ivan Beschastnikh, Ming Wu 0007
INFOCOM8
2023 LVMT: An Efficient Authenticated Storage for Blockchain
Chenxing Li, Sidi Mohamed Beillahi, Guang Yang 0020, Ming Wu 0007, Wei Xu 0005, Fan Long
OSDI4
2022 REFTY: Refinement Types for Valid Deep Learning Models
abstract
Deep learning has been increasingly adopted in many application areas. To construct valid deep learning models, developers must conform to certain computational constraints by carefully selecting appropriate neural architectures and hyperparameter values. For example, the kernel size hyperparameter of the 2D convolution operator cannot be overlarge to ensure that the height and width of the output tensor remain positive. Because model construction is largely manual and lacks necessary tooling support, it is possible to violate those constraints and raise type errors of deep learning models, causing either runtime exceptions or wrong output results. In this paper, we propose Refty, a refinement type-based tool for statically checking the validity of deep learning models ahead of job execution. Refty refines each type of deep learning operator with framework-independent logical formulae that describe the computational constraints on both tensors and hyperparameters. Given the neural architecture and hyperparameter domains of a model, Refty visits every operator, generates a set of constraints that the model should satisfy, and utilizes an SMT solver for solving the constraints. We have evaluated Refty on both individual operators and representative real-world models with various hyperparameter values under PyTorch and TensorFlow. We also compare it with an existing shape-checking tool. The experimental results show that Refty finds all the type errors and achieves 100% Precision and Recall, demonstrating its effectiveness.
Yanjie Gao, Zhengxian Li, Haoxiang Lin, Hongyu Zhang 0002, Ming Wu 0007, Mao Yang 0004
ICSE5
2022 Utilizing Parallelism in Smart Contracts on Decentralized Blockchains by Taming Application-Inherent Conflicts
abstract
Traditional public blockchain systems typically had very limited transaction throughput because of the bottleneck of the consensus protocol itself. With recent advances in consensus technology, the performance limit has been greatly lifted, typically to thousands of transactions per second. With this, transaction execution has become a new performance bottleneck. Exploiting parallelism in transaction execution is a clear and direct way to address this and to further increase transaction throughput. Although some recent literature introduced concurrency control mechanisms to execute smart contract transactions in parallel, the reported speedup that they can achieve is far from ideal. The main reason is that the proposed parallel execution mechanisms cannot effectively deal with the conflicts inherent in many blockchain applications.
Péter Garamvölgyi, Dong Zhou 0006, Fan Long, Ming Wu 0007
ICSE5
2020 Shrec: bandwidth-efficient transaction relay in high-throughput blockchain systems
abstract
The success of Bitcoin and Ethereum has attracted many efforts to build high-throughput blockchain systems. This paper focuses on transaction dissemination --- a rather overlooked issue in these systems. We argue that efficient transaction dissemination is the key for a blockchain system to sustain at high-throughput --- usually thousands of transactions per second --- and the existing solutions fell short at doing so.
Chenxing Li, Peilun Li, Ming Wu 0007, Dong Zhou 0006, Fan Long
SoCC4
2020 A Decentralized Blockchain with High Throughput and Fast Confirmation
Chenxing Li, Peilun Li, Dong Zhou 0006, Ming Wu 0007, Guang Yang 0020, Wei Xu 0005, Fan Long, Andrew Chi-Chih Yao
USENIX ATC5
2020 Distributed Graph Computation Meets Machine Learning
abstract
TuX2is a new distributed graph engine that bridges graph computation and distributed machine learning.TuX2inherits the benefits of elegant graph computation model, efficient graph layout, and balanced parallelism to scale to billion-edge graphs, while extended and optimized for distributed machine learning to support heterogeneity in data model, Stale Synchronous Parallel in scheduling, and a new Mini-batch, Exchange, GlobalSync, and Apply (MEGA) model for programming.TuX2further introduces a hybrid vertex-cut graph optimization and supports various consistency models in fault tolerance for machine learning. We have developed a set of representative distributed machine learning algorithms inTuX2, covering both supervised and unsupervised learning. Compared to the implementations on distributed machine learning platforms, writing those algorithms inTuX2takes only about 25 percent of the code: our graph computation model hides the detailed management of data layout, partitioning, and parallelism from developers. The extensive evaluation ofTuX2, using large datasets with up to 64 billion of edges, shows thatTuX2outperforms PowerGraph/PowerLyra, the state-of-the-art distributed graph engines, by an order of magnitude, while beating two state-of-the-art distributed machine learning systems by at least 60 percent.
Wencong Xiao, Jilong Xue, Youshan Miao, Ming Wu 0007, Wei Li 0022, Lidong Zhou
IEEE Trans. Parallel Distributed Syst.6
2019 Fast Distributed Deep Learning over RDMA
abstract
Deep learning emerges as an important new resource-intensive workload and has been successfully applied in computer vision, speech, natural language processing, and so on. Distributed deep learning is becoming a necessity to cope with growing data and model sizes. Its computation is typically characterized by a simple tensor data abstraction to model multi-dimensional matrices, a dataflow graph to model computation, and iterative executions with relatively frequent synchronizations, thereby making it substantially different from Map/Reduce style distributed big data computation.
Jilong Xue, Youshan Miao, Ming Wu 0007, Lidong Zhou
EuroSys4
2019 Efficient and Effective Sparse LSTM on FPGA with Bank-Balanced Sparsity
abstract
Neural networks based on Long Short-Term Memory (LSTM) are widely deployed in latency-sensitive language and speech applications. To speed up LSTM inference, previous research proposes weight pruning techniques to reduce computational cost. Unfortunately, irregular computation and memory accesses in unrestricted sparse LSTM limit the realizable parallelism, especially when implemented on FPGA. To address this issue, some researchers propose block-based sparsity patterns to increase the regularity of sparse weight matrices, but these approaches suffer from deteriorated prediction accuracy. This work presents Bank-Balanced Sparsity (BBS), a novel sparsity pattern that can maintain model accuracy at a high sparsity level while still enable an efficient FPGA implementation. BBS partitions each weight matrix row into banks for parallel computing, while adopts fine-grained pruning inside each bank to maintain model accuracy. We develop a 3-step software-hardware co-optimization approach to apply BBS in real FPGA hardware. First, we propose a bank-balanced pruning method to induce the BBS pattern on weight matrices. Then we introduce a decoding-free sparse matrix format, Compressed Sparse Banks (CSB), that transparently exposes inter-bank parallelism in BBS to hardware. Finally, we design an FPGA accelerator that takes advantage of BBS to eliminate irregular computation and memory accesses. Implemented on Intel Arria-10 FPGA, the BBS accelerator can achieve 750.9 GOPs on sparse LSTM networks with a batch size of 1. Compared to state-of-the-art FPGA accelerators for LSTM with different compression techniques, the BBS accelerator achieves 2.3 ~ 3.7x improvement on energy efficiency and 7.0 ~ 34.4x reduction on latency with negligible loss of model accuracy.
Shijie Cao, Chen Zhang 0001, Zhuliang Yao, Wencong Xiao, Lanshun Nie, Dechen Zhan, Yunxin Liu 0001, Ming Wu 0007
FPGA8
2019 NeuGraph: Parallel Deep Neural Network Computation on Large Graphs
Lingxiao Ma, Zhi Yang 0001, Youshan Miao, Jilong Xue, Ming Wu 0007, Lidong Zhou, Yafei Dai
USENIX ATC5
2019 Bandwidth and Locality Aware Task-stealing for Manycore Architectures with Bandwidth-Asymmetric Memory
abstract
Parallel computers now start to adopt Bandwidth-Asymmetric Memory architecture that consists of traditional DRAM memory and new High Bandwidth Memory (HBM) for high memory bandwidth. However, existing task schedulers suffer from low bandwidth usage and poor data locality problems in bandwidth-asymmetric memory architectures. To solve the two problems, we propose a Bandwidth and Locality Aware Task-stealing (BATS) system, which consists of an HBM-aware data allocator, a bandwidth-aware traffic balancer, and a hierarchical task-stealing scheduler. Leveraging compile-time code transformation and run-time data distribution, the data allocator enables HBM usage automatically without user interference. According to data access hotness, the traffic balancer migrates data to balance memory traffic across memory nodes proportional to their bandwidth. The hierarchical scheduler improves data locality at runtime without a priori program knowledge. Experiments on an Intel Knights Landing server that adopts bandwidth-asymmetric memory show that BATS reduces the execution time of memory-bound programs up to 83.5% compared with traditional task-stealing schedulers.
Han Zhao 0005, Quan Chen 0002, Yuxian Qiu, Ming Wu 0007, Jingwen Leng, Chao Li 0009, Minyi Guo
ACM Trans. Archit. Code Optim.4
2019 FlexSaaS: A Reconfigurable Accelerator for Web Search Selection
abstract
Web search engines deploy large-scale selection services on CPUs to identify a set of web pages that match user queries. An FPGA-based accelerator can exploit various levels of parallelism and provide a lower latency, higher throughput, more energy-efficient solution than commodity CPUs. However, maintaining such a customized accelerator in a commercial search engine is challenging because selection services are changed often. This article presents our design for FlexSaaS (Flexible Selection as a Service), an FPGA-based accelerator for web search selection. To address efficiency and flexibility challenges, FlexSaaS abstracts computing models and separates memory access from computation. Specifically, FlexSaaS (i) contains a reconfigurable number of matching processors that can handle various possible query plans, (ii) decouples index stream reading from matching computation to fetch and decode index files, and (iii) includes a universal memory accessor that hides the complex memory hierarchy and reduces host data access latency. Evaluated on FPGAs in the selection service of a commercial web search--the Bing web search engine—FlexSaaS can be evolved quickly to adapt to new updates. Compared to the software baseline, FlexSaaS on Arria 10 reduces average latency by 30% and increases throughput by 1.5×.
Shijie Cao, Lanshun Nie, Dechen Zhan, Ningyi Xu, Ramashis Das, Ming Wu 0007, Derek Chiou
ACM Trans. Reconfigurable Technol. Syst.7
2018 SDPaxos: Building Efficient Semi-Decentralized Geo-replicated State Machines
abstract
Existing state machine replication protocols are confronting two major challenges in geo-replication: (1) limited performance caused by load imbalance, and (2) severe performance degradation in heterogeneous environments or under high-contention workloads. This paper presents a new semi-decentralized approach to addressing both the challenges at the same time. Our protocol, SDPaxos, divides the task of a replication protocol into two parts: durably replicating each command across replicas without global order, and ordering all commands to enforce the consistency guarantee. We decentralize the process of replicating commands, which accounts for the largest proportion of load, to provide high performance. In contrast, we centralize the process of ordering commands, which is lightweight but needs a global view, for better performance stability against heterogeneity or contention. The key novelty lies in that SDPaxos achieves the optimal one-round-trip latency under realistic configurations, despite the two separated steps, replicating and ordering, which are both based on Paxos. We also design a recovery protocol to do rapid failover under failures, and a series of optimizations to boost performance. We show via a prototype implementation the significant advantage of SDPaxos on both throughput and latency, facing different environments and workloads.
Quanlu Zhang, Zhi Yang 0001, Ming Wu 0007, Yafei Dai
SoCC4
2017 Tux2: Distributed Graph Computation for Machine Learning
Wencong Xiao, Jilong Xue, Youshan Miao, Ming Wu 0007, Wei Li 0022, Lidong Zhou
NSDI6
2015 GraM: scaling graph computation to the trillions
abstract
GraM is an efficient and scalable graph engine for a large class of widely used graph algorithms. It is designed to scale up to multicores on a single server, as well as scale out to multiple servers in a cluster, offering significant, often over an order-of-magnitude, improvement over existing distributed graph engines on evaluated graph algorithms. GraM is also capable of processing graphs that are significantly larger than previously reported. In particular, using 64 servers (1,024 physical cores), it performs a PageRank iteration in 140 seconds on a synthetic graph with over one trillion edges, setting a new milestone for graph engines.
Ming Wu 0007, Fan Yang 0024, Jilong Xue, Wencong Xiao, Youshan Miao, Haoxiang Lin, Yafei Dai, Lidong Zhou
SoCC1
2015 ImmortalGraph: A System for Storage and Analysis of Temporal Graphs
abstract
Temporal graphs that capture graph changes over time are attracting increasing interest from research communities, for functions such as understanding temporal characteristics of social interactions on a time-evolving social graph. ImmortalGraph is a storage and execution engine designed and optimized specifically for temporal graphs. Locality is at the center of ImmortalGraph’s design: temporal graphs are carefully laid out in both persistent storage and memory, taking into account data locality in both time and graph-structure dimensions. ImmortalGraph introduces the notion of locality-aware batch scheduling in computation, so that common “bulk” operations on temporal graphs are scheduled to maximize the benefit of in-memory data locality. The design of ImmortalGraph explores an interesting interplay among locality, parallelism, and incremental computation in supporting common mining tasks on temporal graphs. The result is a high-performance temporal-graph system that is up to 5 times more efficient than existing database solutions for graph queries. The locality optimizations in ImmortalGraph offer up to an order of magnitude speedup for temporal iterative graph mining compared to a straightforward application of existing graph engines on a series of snapshots.
Youshan Miao, Ming Wu 0007, Fan Yang 0024, Lidong Zhou, Vijayan Prabhakaran, Enhong Chen
ACM Trans. Storage4
2014 Chronos: a graph engine for temporal graph analysis
abstract
Temporal graphs capture changes in graphs over time and are becoming a subject that attracts increasing interest from the research communities, for example, to understand temporal characteristics of social interactions on a time-evolving social graph. Chronos is a storage and execution engine designed and optimized specifically for running in-memory iterative graph computation on temporal graphs. Locality is at the center of the Chronos design, where the in-memory layout of temporal graphs and the scheduling of the iterative computation on temporal graphs are carefully designed, so that common "bulk" operations on temporal graphs are scheduled to maximize the benefit of in-memory data locality. The design of Chronos further explores the interesting interplay among locality, parallelism, and incremental computation in supporting common mining tasks on temporal graphs. The result is a high-performance temporal-graph system that offers up to an order of magnitude speedup for temporal iterative graph mining compared to a straightforward application of existing graph engines on a series of snapshots.
Youshan Miao, Ming Wu 0007, Fan Yang 0024, Lidong Zhou, Vijayan Prabhakaran, Enhong Chen
EuroSys4
2014 Apollo: Scalable and Coordinated Scheduling for Cloud-Scale Computing
Eric Boutin, Jaliya Ekanayake, Wei Lin 0016, Jingren Zhou 0001, Zhengping Qian, Ming Wu 0007, Lidong Zhou
OSDI7
2013 Tango: distributed data structures over a shared log
abstract
Distributed systems are easier to build than ever with the emergence of new, data-centric abstractions for storing and computing over massive datasets. However, similar abstractions do not exist for storing and accessing meta-data. To fill this gap, Tango provides developers with the abstraction of a replicated, in-memory data structure (such as a map or a tree) backed by a shared log. Tango objects are easy to build and use, replicating state via simple append and read operations on the shared log instead of complex distributed protocols; in the process, they obtain properties such as linearizability, persistence and high availability from the shared log. Tango also leverages the shared log to enable fast transactions across different objects, allowing applications to partition state across machines and scale to the limits of the underlying log without sacrificing consistency.
Mahesh Balakrishnan 0001, Dahlia Malkhi, Ted Wobber, Ming Wu 0007, Vijayan Prabhakaran, Michael Wei, John D. Davis, Sriram Rao, Tao Zou 0002, Aviad Zuck
SOSP4
2012 Kineograph: taking the pulse of a fast-changing and connected world
abstract
Kineograph is a distributed system that takes a stream of incoming data to construct a continuously changing graph, which captures the relationships that exist in the data feed. As a computing platform, Kineograph further supports graph-mining algorithms to extract timely insights from the fast-changing graph structure. To accommodate graph-mining algorithms that assume a static underlying graph, Kineograph creates a series of consistent snapshots, using a novel and efficient epoch commit protocol. To keep up with continuous updates on the graph, Kineograph includes an incremental graph-computation engine. We have developed three applications on top of Kineograph to analyze Twitter data: user ranking, approximate shortest paths, and controversial topic detection. For these applications, Kineograph takes a live Twitter data feed and maintains a graph of edges between all users and hashtags. Our evaluation shows that with 40 machines processing 100K tweets per second, Kineograph is able to continuously compute global properties, such as user ranks, with less than 2.5-minute timeliness guarantees. This rate of traffic is more than 10 times the reported peak rate of Twitter as of October 2011.
Raymond Cheng 0001, Aapo Kyrola, Youshan Miao, Xuetian Weng, Ming Wu 0007, Fan Yang 0024, Lidong Zhou, Feng Zhao 0001, Enhong Chen
EuroSys6
2012 Managing Large Graphs on Multi-Cores with Graph Awareness
Vijayan Prabhakaran, Ming Wu 0007, Xuetian Weng, Frank McSherry, Lidong Zhou, Maya Haradasan
USENIX ATC2
2011 Practical software model checking via dynamic interface reduction
abstract
Implementation-level software model checking explores the state space of a system implementation directly to find potential software defects without requiring any specification or modeling. Despite early successes, the effectiveness of this approach remains severely constrained due to poor scalability caused by state-space explosion. DeMeter makes software model checking more practical with the following contributions: (i) proposing dynamic interface reduction, a new state-space reduction technique, (ii) introducing a framework that enables dynamic interface reduction in an existing model checker with a reasonable amount of effort, and (iii) providing the framework with a distributed runtime engine that supports parallel distributed model checking.
Huayang Guo, Ming Wu 0007, Lidong Zhou
SOSP2
2010 Language-based replay via data flow cut
abstract
A replay tool aiming to reproduce a program's execution interposes itself at an appropriate replay interface between the program and the environment. During recording, it logs all non-deterministic side effects passing through the interface from the environment and feeds them back during replay. The replay interface is critical for correctness and recording overhead of replay tools.
Ming Wu 0007, Fan Long, Xi Wang 0005, Zhilei Xu, Haoxiang Lin, Xuezheng Liu, Huayang Guo, Lidong Zhou, Zheng Zhang 0001
SIGSOFT FSE1
2009 MODIST: Transparent Model Checking of Unmodified Distributed Systems
Tisheng Chen, Ming Wu 0007, Zhilei Xu, Xuezheng Liu, Haoxiang Lin, Mao Yang 0004, Fan Long, Lidong Zhou
NSDI3
2009 MPIWiz: subgroup reproducible replay of mpi applications
abstract
Message Passing Interface (MPI) is a widely used standard for managing coarse-grained concurrency on distributed computers. Debugging parallel MPI applications, however, has always been a particularly challenging task due to their high degree of concurrent execution and non-deterministic behavior. Deterministic replay is a potentially powerful technique for addressing these challenges, with existing MPI replay tools adopting either data-replay or order-replay approaches. Unfortunately, each approach has its tradeoffs. Data-replay generates substantial log sizes by recording every communication message. Order-replay generates small logs, but requires all processes to be replayed together. We believe that these drawbacks are the primary reasons that inhibit the wide adoption of deterministic replay as the critical enabler of cyclic debugging of MPI applications.
Ruini Xue, Xuezheng Liu, Ming Wu 0007, Zheng Zhang 0001, Geoffrey M. Voelker
PPoPP3
2008 D3S: Debugging Deployed Distributed Systems
Xuezheng Liu, Xi Wang 0005, Feibo Chen, Xiaochen Lian, Ming Wu 0007, M. Frans Kaashoek, Zheng Zhang 0001
NSDI7
2008 Corey: An Operating System for Many Cores
Silas Boyd-Wickizer, Haibo Chen 0001, Rong Chen 0001, Yandong Mao, M. Frans Kaashoek, Robert Morris 0005, Aleksey Pesterev, Lex Stein, Ming Wu 0007, Yue-hua Dai, Zheng Zhang 0001
OSDI9
2008 R2: An Application-Level Kernel for Record and Replay
Xi Wang 0005, Xuezheng Liu, Zhilei Xu, Ming Wu 0007, M. Frans Kaashoek, Zheng Zhang 0001
OSDI6