Michael Wei

dblp:132/0329 · DBLP profile ↗
← Back
23ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0003-0433-3745ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 6 · 2 since 2021Computer networks · 3 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 DeepErr: Automatic Root-Cause Analysis of System Call Failures
abstract
System call failures present significant challenges for operating system (OS) users, as the failures are often cryptic and difficult to diagnose due to limited error codes and missing documentation. As a result, software developers struggle to utilize system calls effectively, and power users encounter difficulties configuring the OS and resolving environment problems. Existing automatic root-cause analysis tools are inadequate, primarily due to dependence on comparative analysis, which requires similar successful executions that are often unavailable.
Nadav Amit, Michael Wei
SYSTOR2
2022 Graham: Synchronizing Clocks by Leveraging Local Clock Properties
Ali Najafi, Michael Wei
NSDI2
2022 Optimizing Storage Performance with Calibrated Interrupts
abstract
After request completion, an I/O device must decide whether to minimize latency by immediately firing an interrupt or to optimize for throughput by delaying the interrupt, anticipating that more requests will complete soon and help amortize the interrupt cost. Devices employ adaptive interrupt coalescing heuristics that try to balance between these opposing goals. Unfortunately, because devices lack the semantic information about which I/O requests are latency-sensitive, these heuristics can sometimes lead to disastrous results. Instead, we propose addressing the root cause of the heuristics problem by allowing software to explicitly specify to the device if submitted requests are latency-sensitive. The device then “calibrates” its interrupts to completions of latency-sensitive requests. We focus on NVMe storage devices and show that it is natural to express these semantics in the kernel and the application and only requires a modest two-bit change to the device interface. Calibrated interrupts increase throughput by up to 35%, reduce CPU consumption by as much as 30%, and achieve up to 37% lower latency when interrupts are coalesced.
Amy Tai, Igor Smolyar, Michael Wei, Dan Tsafrir
ACM Trans. Storage3
2021 Systems research is running out of time
abstract
Most sciences conduct experiments with a thorough understanding of the accuracy and precision of the instruments used for making measurements. Time is the most frequently used measurement in systems research, yet most of the literature does not consider the precision and accuracy of clocks. In this paper, we argue for the importance of understanding timekeeping and providing precise and accurate time for general systems research.
Ali Najafi, Amy Tai, Michael Wei
HotOS3
2021 Optimizing Storage Performance with Calibrated Interrupts
Amy Tai, Igor Smolyar, Michael Wei, Dan Tsafrir
OSDI3
2021 Dealing with (some of) the fallout from meltdown
abstract
The meltdown vulnerability allows users to read kernel memory by exploiting a hardware flaw in speculative execution. Processor vendors recommend "page table isolation" (PTI) as a software fix, but PTI can significantly degrade the performance of system-call-heavy programs. Leveraging the fact that 32-bit pointers cannot access 64-bit kernel memory, we propose "Shrink", a safe alternative to PTI, which is applicable to programs capable of running in 32-bit address spaces. We show that Shrink can restore the performance of some workloads, suggest additional potential alternatives, and argue that vendors must be more open about hardware flaws to allow developers to design protection schemes that are safe and performant.
Nadav Amit, Michael Wei, Dan Tsafrir
SYSTOR2
2021 RainBlock: Faster Transaction Processing in Public Blockchains
Soujanya Ponnapalli, Aashaka Shah, Souvik Banerjee, Dahlia Malkhi, Amy Tai, Vijay Chidambaram, Michael Wei
USENIX ATC7
2020 Don't shoot down TLB shootdowns!
abstract
Translation Lookaside Buffers (TLBs) are critical for building performant virtual memory systems. Because most processors do not provide coherence for TLB mappings, TLB shootdowns provide a software mechanism that invokes inter-processor interrupts (IPLs) to synchronize TLBs. TLB shootdowns are expensive, so recent work has aimed to avoid the frequency of shootdowns through techniques such as batching. We show that aggressive batching can cause correctness issues and addressing them can obviate the benefits of batching. Instead, our work takes a different approach which focuses on both improving the performance of TLB shootdowns and carefully selecting where to avoid shootdowns. We introduce four general techniques to improve shootdown performance: (1) concurrently flush initiator and remote TLBs, (2) early acknowledgement from remote cores, (3) cacheline consolidation of kernel data structures to reduce cacheline contention, and (4) in-context flushing of userspace entries to address the overheads introduced by Spectre and Meltdown mitigations. We also identify that TLB flushing can be avoiding when handling copy-on-write (CoW) faults and some TLB shootdowns can be batched in certain system calls. Overall, we show that our approach results in significant speedups without sacrificing safety and correctness in both microbenchmarks and real-world applications.
Nadav Amit, Amy Tai, Michael Wei
EuroSys3
2019 Just-In-Time Compilation for Verilog: A New Technique for Improving the FPGA Programming Experience
abstract
FPGAs offer compelling acceleration opportunities for modern applications. However compilation for FPGAs is painfully slow, potentially requiring hours or longer. We approach this problem with a solution from the software domain: the use of a JIT. Code is executed immediately in a software simulator, and compilation is performed in the background. When finished, the code is moved into hardware, and from the user's perspective it simply gets faster. We have embodied these ideas in Cascade: the first JIT compiler for Verilog. Cascade reduces the time between initiating compilation and running code to less than a second, and enables generic printf debugging from hardware. Cascade preserves program performance to within 3× in a debugging environment, and has minimal effect on a finalized design. Crucially, these properties hold even for programs that perform side effects on connected IO devices. A user study demonstrates the value to experts and non-experts alike: Cascade encourages more frequent compilation, and reduces the time to produce working hardware designs.
Eric Schkufza, Michael Wei, Christopher J. Rossbach
ASPLOS2
2019 Storm: a fast transactional dataplane for remote data structures
abstract
RDMA technology enables a host to access the memory of a remote host without involving the remote CPU, improving the performance of distributed in-memory storage systems. Previous studies argued that RDMA suffers from scalability issues, because the NIC's limited resources are unable to simultaneously cache the state of all the concurrent network streams. These concerns led to various software-based proposals to reduce the size of this state by trading off performance.
Stanko Novakovic, Yizhou Shan, Aasheesh Kolli, Michael Cui, Yiying Zhang 0005, Haggai Eran, Boris Pismenny, Liran Liss, Michael Wei, Dan Tsafrir, Marcos K. Aguilera
SYSTOR9
2019 JumpSwitches: Restoring the Performance of Indirect Branches In the Era of Spectre
Nadav Amit, Fred Jacobs, Michael Wei
USENIX ATC3
2018 Sharing, Protection, and Compatibility for Reconfigurable Fabric with AmorphOS
Ahmed Khawaja, Joshua Landgraf, Rohith Prakash, Michael Wei, Eric Schkufza, Christopher J. Rossbach
OSDI4
2018 Remote regions: a simple abstraction for remote memory
Marcos K. Aguilera, Nadav Amit, Irina Calciu, Xavier Deguillard, Jayneel Gandhi, Stanko Novakovic, Arun Ramanathan, Pratap Subrahmanyam, Lalith Suresh 0001, Kiran Tati, Rajesh Venkatasubramanian, Michael Wei
USENIX ATC12
2018 The Design and Implementation of Hyperupcalls
Nadav Amit, Michael Wei
USENIX ATC2
2017 Remote memory in the age of fast networks
abstract
As the latency of the network approaches that of memory, it becomes increasingly attractive for applications to use remote memory---random-access memory at another computer that is accessed using the virtual memory subsystem. This is an old idea whose time has come, in the age of fast networks. To work effectively, remote memory must address many technical challenges. In this paper, we enumerate these challenges, discuss their feasibility, explain how some of them are addressed by recent work, and indicate other promising ways to tackle them. Some challenges remain as open problems, while others deserve more study. In this paper, we hope to provide a broad research agenda around this topic, by proposing more problems than solutions.
Marcos K. Aguilera, Nadav Amit, Irina Calciu, Xavier Deguillard, Jayneel Gandhi, Pratap Subrahmanyam, Lalith Suresh 0001, Kiran Tati, Rajesh Venkatasubramanian, Michael Wei
SoCC10
2017 Hypercallbacks: Decoupling Policy Decisions and Execution
abstract
research-article Share on Hypercallbacks: Decoupling Policy Decisions and Execution Authors: Nadav Amit VMware Research Group VMware Research GroupView Profile , Michael Wei VMware Research Group VMware Research GroupView Profile , Cheng-Chun Tu VMware VMwareView Profile Authors Info & Claims HotOS '17: Proceedings of the 16th Workshop on Hot Topics in Operating SystemsMay 2017 Pages 37–41https://doi.org/10.1145/3102980.3102987Published:07 May 2017Publication History 2citation337DownloadsMetricsTotal Citations2Total Downloads337Last 12 Months21Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Nadav Amit, Michael Wei, Cheng-Chun Tu
HotOS2
2017 vCorfu: A Cloud-Scale Object Store on a Shared Log
Michael Wei, Amy Tai, Christopher J. Rossbach, Ittai Abraham, Maithem Munshed, Medhavi Dhawan, Jim Stabile, Udi Wieder, Scott Fritchie, Steven Swanson, Michael J. Freedman, Dahlia Malkhi
NSDI1
2016 Silver: A Scalable, Distributed, Multi-versioning, Always Growing (Ag) File System
Michael Wei, Christopher J. Rossbach, Ittai Abraham, Udi Wieder, Steven Swanson, Dahlia Malkhi, Amy Tai
HotStorage1
2016 Replex: A Scalable, Highly Available Multi-Index Data Store
Amy Tai, Michael Wei, Michael J. Freedman, Ittai Abraham, Dahlia Malkhi
USENIX ATC2
2013 Tango: distributed data structures over a shared log
abstract
Distributed systems are easier to build than ever with the emergence of new, data-centric abstractions for storing and computing over massive datasets. However, similar abstractions do not exist for storing and accessing meta-data. To fill this gap, Tango provides developers with the abstraction of a replicated, in-memory data structure (such as a map or a tree) backed by a shared log. Tango objects are easy to build and use, replicating state via simple append and read operations on the shared log instead of complex distributed protocols; in the process, they obtain properties such as linearizability, persistence and high availability from the shared log. Tango also leverages the shared log to enable fast transactions across different objects, allowing applications to partition state across machines and scale to the limits of the underlying log without sacrificing consistency.
Mahesh Balakrishnan 0001, Dahlia Malkhi, Ted Wobber, Ming Wu 0007, Vijayan Prabhakaran, Michael Wei, John D. Davis, Sriram Rao, Tao Zou 0002, Aviad Zuck
SOSP6
2013 Beyond block I/O: implementing a distributed shared log in hardware
abstract
The basic block I/O interface used for interacting with storage devices hasn't changed much in 30 years. With the advent of very fast I/O devices based on solid-state memory, it becomes increasingly attractive to make many devices directly and concurrently available to many clients. However, when multiple clients share media at fine grain, retaining data consistency is problematic: SCSI, IDE, and their descendants don't offer much help. We propose an interface to networked storage that reduces an existing software implementation of a distributed shared log to hardware. Our system achieves both scalable throughput and strong consistency, while obtaining significant benefits in cost and power over the software implementation.
Michael Wei, John D. Davis, Ted Wobber, Mahesh Balakrishnan 0001, Dahlia Malkhi
SYSTOR1
2013 CORFU: A distributed shared log
abstract
CORFU is a global log which clients can append-to and read-from over a network. Internally, CORFU is distributed over a cluster of machines in such a way that there is no single I/O bottleneck to either appends or reads. Data is fully replicated for fault tolerance, and a modest cluster of about 16--32 machines with SSD drives can sustain 1 million 4-KByte operations per second. The CORFU log enabled the construction of a variety of distributed applications that require strong consistency at high speeds, such as databases, transactional key-value stores, replicated state machines, and metadata services.
Mahesh Balakrishnan 0001, Dahlia Malkhi, John D. Davis, Vijayan Prabhakaran, Michael Wei, Ted Wobber
ACM Trans. Comput. Syst.5
2012 CORFU: A Shared Log Design for Flash Clusters
Mahesh Balakrishnan 0001, Dahlia Malkhi, Vijayan Prabhakaran, Ted Wobber, Michael Wei, John D. Davis
NSDI5