Ali José Mashtizadeh

dblp:135/6767 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0002-8672-5138ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 7 · 2 first-author · 4 since 2021Computer networks · 3 · 2 since 2021Security and privacy · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 HA/TCP: A Reliable and Scalable Framework for TCP Network Functions
Haoyu Gu, Ali José Mashtizadeh, Bernard Wong 0001
NSDI2
2024 MemSnap μCheckpoints: A Data Single Level Store for Fearless Persistence
abstract
Single level stores (SLSes) have recently resurfaced as a system for persisting application data. SLSes like EROS, Aurora, and TreeSLS use application checkpointing to replace file-based APIs. These systems checkpoint at a coarse granularity and must be combined with file persistence mechanisms like WALs, undermining the benefits of their SLS design.
Emil Tsalapatis, Ryan Hancock, Rakeeb Hossain, Ali José Mashtizadeh
ASPLOS (3)4
2024 Draconis: Network-Accelerated Scheduling for Microsecond-Scale Workloads
abstract
We present Draconis, a novel scheduler for workloads in the range of tens to hundreds of microseconds. Draconis challenges the popular belief that programmable switches cannot house the complex data structures, such as queues, needed to support an in-network scheduler. Using programmable switches, Draconis achieves the low scheduling tail latency and high throughput needed to support these microsecond-scale workloads on large clusters. Furthermore, Draconis supports a wide range of complex scheduling policies, including locality-aware scheduling, priority-based scheduling, and resource-based scheduling.
Sreeharsha Udayashankar, Ashraf Abdel-Hadi, Ali José Mashtizadeh, Samer Al-Kiswany
EuroSys3
2024 Simple, Fast and Widely Applicable Concurrent Memory Reclamation via Neutralization
abstract
Reclaiming memory in non-blocking dynamic data structures in unmanaged languages like C/C++ presents a unique challenge due to the risk of use-after-free errors caused by concurrent accesses. Existing safe memory reclamation (SMR) algorithms fall short of satisfying five key properties: high performance, bounded garbage, usability, consistency, and applicability. In particular, bounded garbage and high performance are quite difficult to achieve simultaneously. In this paper, we address this limitation by proposing a new, provably correct technique called neutralization based reclamation (NBR) that neutralizes threads using POSIX signals to provide the synchronization required for safe memory reclamation. NBR uses atomic reads and writes and achieves bounded garbage and high performance without imposing significant overhead on concurrent readers and writers. An extensive experimental evaluation serves to demonstrate the efficiency of our technique across various data structures, reclamation algorithms, and workloads. A detailed survey of popular concurrent data structures suggests NBR is applicable to a wide range of data structures, many of which could not be used with prior SMR algorithms that guarantee bounded garbage.
Ajay Singh 0002, Trevor Brown 0001, Ali José Mashtizadeh
IEEE Trans. Parallel Distributed Syst.3
2023 Metal: An Open Architecture for Developing Processor Features
abstract
In recent years, an increasing number of hardware devices started providing programming interfaces to developers such as smart NICs. Processor vendors use microcode to extend processors' features such as Intel SGX and VT-x. This enables processor architects to quickly evolve processor designs and features. However, modern processors still lack general programmability as microcode is inaccessible to system developers. Developers still cannot define custom processor features. We argue that processors should expose this capability to developers, which enables new operating system and application designs.
Siyao Zhao, Ali José Mashtizadeh
HotOS2
2022 OrcBench: A Representative Serverless Benchmark
abstract
Serverless computing is rapidly growing area of research. No standardized benchmark currently exists for evaluating orchestration level decisions or executing large serverless workloads because of the limited data provided by cloud providers. Current benchmarks focus on other aspects, such as the cost of running general types of functions and their runtimes.We introduce OrcBench, the first orchestration benchmark based on the recently published Microsoft Azure serverless data set. OrcBench categorizes 8622 serverless functions into 17 distinct models, which represent 5.6 million invocations from the original trace.OrcBench also incorporates a time-series analysis to identify function chains within the dataset. OrcBench can use these to create workloads that mimic complete serverless applications, which includes simulating CPU and memory usage. The modeling allows these workloads to be scaled according to the target hardware configuration.
Ryan Hancock, Sreeharsha Udayashankar, Ali José Mashtizadeh, Samer Al-Kiswany
CLOUD3
2022 Accelerating Reads With In-Network Consistency-Aware Load Balancing
abstract
We present FLAIR, a novel approach for accelerating read operations in leader-based consensus protocols. FLAIR leverages the capabilities of the new generation of programmable switches to serve reads from follower replicas without compromising consistency. The core of the new approach is a packet-processing pipeline that can track client requests and system replies, identify consistent replicas, and at line speed, forward read requests to replicas that can serve the read without sacrificing linearizability. An additional benefit of FLAIR is that it facilitates devising novel consistency-aware load balancing techniques. Following the new approach, we designed FlairKV, a key-value store atop Raft. FlairKV implements the processing pipeline using the P4 programming language. We evaluate the benefits of the proposed approach and compare it to previous approaches using a cluster with a Barefoot Tofino switch. Our evaluation indicates that, compared to state-of-the-art alternatives, the proposed approach can bring significant performance gains: up to 42% higher throughput and 35–97% lower latency for most workloads. Furthermore, our evaluation shows that our novel load balancing techniques can cope with heterogeneous load and hardware to achieve higher performance, and that FLAIR can scale to support large data sets and clusters.
Ibrahim Kettaneh, Ahmed Alquraan, Hatem Takruri, Ali José Mashtizadeh, Samer Al-Kiswany
IEEE/ACM Trans. Netw.4
2021 The Aurora operating system: revisiting the single level store
abstract
Applications on modern operating systems manage their ephemeral state in memory, and persistent state on disk. Ensuring consistency between them is a source of significant developer effort, yet still a source of significant bugs in mature applications. We present the Aurora single level store (SLS), an OS that simplifies persistence by automatically persisting all traditionally ephemeral application state. With recent storage hardware like NVMe SSDs and NVDIMMs, Aurora is able to continuously checkpoint entire applications with millisecond granularity.
Emil Tsalapatis, Ryan Hancock, Tavian Barnes, Ali José Mashtizadeh
HotOS4
2021 NBR: neutralization based reclamation
abstract
Safe memory reclamation (SMR) algorithms suffer from a trade-off between bounding unreclaimed memory and the speed of reclamation. Hazard pointer (HP) based algorithms bound unreclaimed memory at all times, but tend to be slower than other approaches. Epoch based reclamation (EBR) algorithms are faster, but do not bound memory reclamation. Other algorithms follow hybrid approaches, requiring special compiler or hardware support, changes to record layouts, and/or extensive code changes. Not all SMR algorithms can be used to reclaim memory for all data structures.
Ajay Singh 0002, Trevor Brown 0001, Ali José Mashtizadeh
PPoPP3
2021 The Aurora Single Level Store Operating System
abstract
Applications on modern operating systems manage their ephemeral state in memory and persistent state on disk. Ensuring consistency between them is a source of significant developer effort and application bugs. We present the Aurora single level store, an OS that eliminates the distinction between ephemeral and persistent application state.
Emil Tsalapatis, Ryan Hancock, Tavian Barnes, Ali José Mashtizadeh
SOSP4
2021 SKQ: Event Scheduling for Optimizing Tail Latency in a Traditional OS Kernel
Siyao Zhao, Haoyu Gu, Ali José Mashtizadeh
USENIX ATC3
2020 Fault Tolerant Service Function Chaining
abstract
Network traffic typically traverses a sequence of middleboxes forming a service function chain, or simply a chain. Tolerating failures when they occur along chains is imperative to the availability and reliability of enterprise applications. Making a chain fault-tolerant is challenging since, in the event of failures, the state of faulty middleboxes must be correctly and quickly recovered while providing high throughput and low latency.
Milad Ghaznavi, Elaheh Jalalpour, Bernard Wong 0001, Raouf Boutaba, Ali José Mashtizadeh
SIGCOMM5
2017 Towards Practical Default-On Multi-Core Record/Replay
abstract
We present Castor, a record/replay system for multi-core applications that provides consistently low and predictable overheads. With Castor, developers can leave record and replay on by default, making it practical to record and reproduce production bugs, or employ fault tolerance to recover from hardware failures.
Ali José Mashtizadeh, Tal Garfinkel, David Terei, David Mazières, Mendel Rosenblum
ASPLOS1
2015 CCFI: Cryptographically Enforced Control Flow Integrity
abstract
Control flow integrity (CFI) restricts jumps and branches within a program to prevent attackers from executing arbitrary code in vulnerable programs. However, traditional CFI still offers attackers too much freedom to chose between valid jump targets, as seen in recent attacks.
Ali José Mashtizadeh, Andrea Bittau, Dan Boneh, David Mazières
CCS1
2014 Hacking Blind
abstract
We show that it is possible to write remote stack buffer overflow exploits without possessing a copy of the target binary or source code, against services that restart after a crash. This makes it possible to hack proprietary closed-binary services, or open-source servers manually compiled and installed from source where the binary remains unknown to the attacker. Traditional techniques are usually paired against a particular binary and distribution where the hacker knows the location of useful gadgets for Return Oriented Programming (ROP). Our Blind ROP (BROP) attack instead remotely finds enough ROP gadgets to perform a write system call and transfers the vulnerable binary over the network, after which an exploit can be completed using known techniques. This is accomplished by leaking a single bit of information based on whether a process crashed or not when given a particular input string. BROP requires a stack vulnerability and a service that restarts after a crash. We implemented Braille, a fully automated exploit that yielded a shell in under 4,000 requests (20 minutes) against a contemporary nginx vulnerability, yaSSL + MySQL, and a toy proprietary server written by a colleague. The attack works against modern 64-bit Linux with address space layout randomization (ASLR), no-execute page protection (NX) and stack canaries.
Andrea Bittau, Adam Belay, Ali José Mashtizadeh, David Mazières, Dan Boneh
IEEE Symposium on Security and Privacy3
2014 XvMotion: Unified Virtual Machine Migration over Long Distance
Ali José Mashtizadeh, Min Cai, Gabriel Tarasuk-Levin, Ricardo Koller, Tal Garfinkel, Sreekanth Setty
USENIX ATC1
2013 Replication, history, and grafting in the Ori file system
abstract
Ori is a file system that manages user data in a modern setting where users have multiple devices and wish to access files everywhere, synchronize data, recover from disk failure, access old versions, and share data. The key to satisfying these needs is keeping and replicating file system history across devices, which is now practical as storage space has outpaced both wide-area network (WAN) bandwidth and the size of managed data. Replication provides access to files from multiple devices. History provides synchronization and offline access. Replication and history together subsume backup by providing snapshots and avoiding any single point of failure. In fact, Ori is fully peer-to-peer, offering opportunistic synchronization between user devices in close proximity and ensuring that the file system is usable so long as a single replica remains. Cross-file system data sharing with history is provided by a new mechanism called grafting. An evaluation shows that as a local file system, Ori has low overhead compared to a File system in User Space (FUSE) loopback driver; as a network file system, Ori over a WAN outperforms NFS over a LAN.
Ali José Mashtizadeh, Andrea Bittau, Yifeng Frank Huang, David Mazières
SOSP1
2012 Dune: Safe User-level Access to Privileged CPU Features
Adam Belay, Andrea Bittau, Ali José Mashtizadeh, David Terei, David Mazières, Christoforos E. Kozyrakis
OSDI3
2011 vIC: Interrupt Coalescing for Virtual Machine Storage Device IO
Irfan Ahmad 0005, Ajay Gulati, Ali José Mashtizadeh
USENIX ATC3
2011 The Design and Evolution of Live Storage Migration in VMware ESX
Ali José Mashtizadeh, Emré Celebi, Tal Garfinkel, Min Cai
USENIX ATC1