EDBT 2026 Demo / reviewers in the wild / expert
Mohammad Shahrad
dblp:166/3165
· DBLP profile ↗
23ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0002-8214-9583ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 7 first-author · 9 since 2021Computer networks · 3 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Demystifying Serverless Costs on Public Platforms: Bridging Billing, Architecture, and OS SchedulingabstractPublic cloud serverless platforms have attracted a large user base due to their high scalability, plug-and-play deployment model, and pay-per-use billing. However, compared to virtual machines and container hosting services, modern serverless offerings typically impose higher per-unit time and resource charges. Additionally, billing practices such as wall-clock time allocation-based billing, invocation fees, and usage rounding up can further increase costs. Changyuan Lin, Yuanzhi Ma, Mohammad Shahrad |
EuroSys | 3 |
| 2026 | In-Production Characterization of an Open Source Serverless Platform and New Scaling StrategiesabstractServerless computing has become more popular and evolved to support more complex tasks than the original Function as a Service (FaaS) model. The design of serverless systems has advanced to accommodate application demands and offer flexibility. Careful characterization of modern serverless systems and understanding of current gaps are warranted. Publicly available datasets on workloads in select production serverless systems do not fully represent all offerings or capture traces at the required time resolution to identify changes in application-level request-response patterns. Nima Nasiri, Nalin Munshi, Simon Moser, Marius Pirvu, Vijay Sundaresan, Daryl Maier, Thatta Premnath, Norman Böwing, Sathish Gopalakrishnan, Mohammad Shahrad |
EuroSys | 10 |
| 2026 | Hierarchical Integration of WebAssembly in Serverless for Efficiency and Interoperability
Mohammadamin Baqershahi, Changyuan Lin, Visal Saosuo, Paul Chen, Mohammad Shahrad |
NSDI | 5 |
| 2025 | Scheduling Job Streams on Uniprocessors with Cold Start DelaysabstractWe consider uniprocessor job scheduling where jobs have deadlines and each job belongs to a job family. Each job family has an associated setup time or cold start delay, and when a job is scheduled, if the predecessor job does not belong to the same job family then this setup time needs to be included. We examine the scheduling problem when the objective is to minimize the number of tardy jobs, and the challenge of including the setup time for switching between job families results in poor performance of well-known policies such as earliest deadline first (EDF). We propose a near-optimal online scheduling policy for jobs with deadlines, on uniprocessor platforms. This problem arises in a variety of contexts including serverless computing and MLaaS (machine learning as a service) where a job request may need a suitable resource container to be provisioned if an earlier request was not of the same type. The general offline problem of job scheduling with job families and setup costs has previously been studied and shown to be NP-Hard. In an effort to improve our understanding of the online problem with the objective of maximizing the number of jobs that meet their deadlines, we focus on the case where all jobs have the same execution time. We show that even this special case is NP-Hard in the offline setting. The policy we propose, which requires job buffering, is nearly 1-competitive when each job has a reasonably large slack. We also propose a heuristic that performs well in many situations despite a weak competitive ratio. Sathish Gopalakrishnan, Grady Thompson, Jonathan Cao, Mohammad Shahrad |
RTAS | 4 |
| 2025 | Growlithe: A Developer-Centric Compliance Tool for Serverless ApplicationsabstractServerless applications consist of functions written in heterogeneous programming languages, use diverse data stores and communication services, and evolve rapidly. Consequently, it is challenging for serverless tenants to protect their application data from inadvertent leaks due to bugs, misconfigurations, and human errors. Cloud security tools, such as Identity and Access Management (IAM), lack observability into a tenant's application, whereas the state-of-the-art dataflow tracking tools require support from the cloud platform and incur significant runtime overheads. We present Growlithe, a tool that integrates with the serverless application development toolchain and enables continuous compliance with data policies by design. Growlithe allows declarative specification of access and data flow control policies over a language- and platform-independent dataflow graph abstraction of a serverless application, and enforces these policies through a combination of static analysis and runtime enforcement. We used Growlithe with applications using Python and JavaScript functions that can be hosted on AWS Lambda and Google Cloud Functions platforms. We empirically demonstrate that Growlithe is cross-cutting, portable and efficient, and enables developers to easily adapt their application and policies to evolving requirements. Arshia Moghimi, Devam Sisodraker, Mohammad Shahrad, Aastha Mehta |
SP | 4 |
| 2024 | The Sunk Carbon Fallacy: Rethinking Carbon Footprint Metrics for Effective Carbon-Aware SchedulingabstractThe rapid increase in computing demand and corresponding energy consumption have focused attention on computing's impact on the climate and sustainability. Prior work proposes metrics that quantify computing's carbon footprint across several lifecycle phases, including its supply chain, operation, and end-of-life. Industry uses these metrics to optimize the carbon footprint of manufacturing hardware and running computing applications. Unfortunately, prior work on optimizing datacenters' carbon footprint often succumbs to the sunk cost fallacy by considering embodied carbon emissions (a sunk cost) when making operational decisions (i.e., job scheduling and placement), which leads to operational decisions that do not always reduce the total carbon footprint. Noman Bashir, Varun Gohil, Anagha Belavadi Subramanya, Mohammad Shahrad, David Irwin 0001, Elsa Olivetti, Christina Delimitrou |
SoCC | 4 |
| 2024 | Caribou: Fine-Grained Geospatial Shifting of Serverless Applications for SustainabilityabstractSustainability in computing is critical as environmental concerns rise. The cloud industry's carbon footprint is significant and rapidly growing. We show that dynamic geospatial shifting of cloud workloads to regions with lower carbon emission energy sources, particularly for more portable cloud workloads such as serverless applications, has a high potential to lower operational carbon emissions. To make the case, we build a comprehensive framework called Caribou that offloads serverless workflows across geo-distributed regions. Caribou requires no change in the application logic, nor on the provider side. It dynamically determines the best deployment plans, automatically (re-) deploys functions to appropriate regions, and redirects traffic to new endpoints. In reducing operational carbon through fine-grained, function-level offloading, Caribou does not undermine standard metrics such as performance and cost. We show how this approach can reduce the carbon footprint by an average of 22.9% to 66.6% across the North American continent. We demonstrate how a detailed specification of location constraints (e.g., to ensure compliance of one stage) can allow emission reductions for workflows (e.g., by offloading other stages). By showcasing the feasibility of carbon-aware geospatial application deployment, Caribou aims to push the boundaries of system techniques available to curtail cloud carbon emissions and provide a framework for future research. Viktor Gsteiger, Pin Hong (Daniel) Long, Yiran (Jerry) Sun, Parshan Javanrood, Mohammad Shahrad |
SOSP | 5 |
| 2023 | Parrotfish: Parametric Regression for Optimizing Serverless FunctionsabstractServerless computing is a new paradigm that aims to remove the burdens of cloud management from developers. Yet rightsizing serverless functions remains a pain point for developers. Choosing the right memory configuration is necessary to ensure cost and/or performance optimality for serverless workloads. In this work, we identify that using parametric regression can significantly simplify function rightsizing compared to black-box optimization techniques currently available. With this insight, we build a tool, called Parrotfish, which finds optimal configurations through an online learning process. It also allows users to communicate constraints on execution time, or to relax cost optimality to gain performance. Parrotfish achieves substantially lower exploration costs (1.81-9.96×) compared with the state-of-the-art tools, while delivering similar or better recommendations. Arshia Moghimi, Joe Hattori, Alexander Li, Mehdi Ben Chikha, Mohammad Shahrad |
SoCC | 5 |
| 2023 | VectorVisor: A Binary Translation Scheme for Throughput-Oriented GPU Acceleration
Samuel Ginzburg, Mohammad Shahrad, Michael J. Freedman |
USENIX ATC | 2 |
| 2023 | UnFaaSener: Latency and Cost Aware Offloading of Functions from Serverless Platforms
Ghazal Sadeghian, Mohamed Elsakhawy, Mohanna Shahrad, Joe Hattori, Mohammad Shahrad |
USENIX ATC | 5 |
| 2021 | On Merits and Viability of Multi-Cloud ServerlessabstractServerless computing is a rapidly growing paradigm in the cloud industry that envisions functions as the computational building blocks of an application. Instead of forcing the application developer to provision cloud resources for their application, the cloud provider provisions the required resources for each function "under the hood." In this work, we envision virtual serverless providers (VSPs) to aggregate serverless offerings. In doing so, VSPs allow developers (and businesses) to get rid of vendor lock-in problems and exploit pricing and performance variation across providers by adaptively utilizing the best provider at each time, forcing the providers to compete to offer cheaper and superior services. We discuss the merits of a VSP and show that serverless systems are well-suited to cross-provider aggregation, compared to virtual machines. We propose a VSP system architecture and implement an initial version. Using experimental evaluations, our preliminary results show that a VSP can improve maximum sustained throughput by 1.2x to 4.2x, reduces SLO violations by 98.8%, and improves the total invocations' costs by 54%. Ataollah Fatahi Baarzi, George Kesidis, Carlee Joe-Wong, Mohammad Shahrad |
SoCC | 4 |
| 2021 | Provisioning Differentiated Last-Level Cache Allocations to VMs in Public CloudsabstractPublic cloud providers offer access to hardware resources and users rent resources by choosing among many VM sizes. While users choose the CPU core count and main memory size per VM, they cannot specify last-level cache (LLC) requirements. LLC is typically shared among all cores of a modern CPU causing cache contention and performance interference among co-located VMs. Consequently, a user's only way to avoid this interference is purchasing a full-server VM to prevent co-tenants. Although researchers have studied LLC partitioning and despite its availability in commodity processors, LLC QoS has not been offered to public cloud users today. Existing techniques rely mostly on performance profiling, which is not feasible in public cloud settings with opaque VMs. Moreover, prior work does not address how to deliver differentiated LLC allocations at scale. Mohammad Shahrad, Sameh Elnikety, Ricardo Bianchini |
SoCC | 1 |
| 2020 | Serverless in the Wild: Characterizing and Optimizing the Serverless Workload at a Large Cloud Provider
Mohammad Shahrad, Rodrigo Fonseca, Íñigo Goiri, Gohar Irfan Chaudhry, Paul Batum, Jason Cooke, Eduardo Laureano, Colby Tresness, Mark Russinovich, Ricardo Bianchini |
USENIX ATC | 1 |
| 2020 | Burstable Instances for Clouds: Performance Modeling, Equilibrium Analysis, and Revenue MaximizationabstractLeading cloud providers recently introduced a new instance type named burstable instances to better match the time-varying workloads of tenants and further reduce their costs. In the research community, however, little has been done to understand burstable instances from a theoretical perspective. This paper presents the first unified framework to model, analyze, and optimize the operation of burstable instances. Specifically, we model the resource provisioning of burstable instances, identify key performance metrics, and derive the analytical performance given the resource provisioning decisions. We then characterize the equilibrium behind tenants' responses to the prices offered for different burstable instance service classes, taking into account the impact of tenants' actions on the performance achieved by each service class. In addition, we investigate how a cloud provider can leverage knowledge of this equilibrium to find the prices that maximize its total revenue. Finally, we validate our framework on real traces and demonstrate its usage to price burstable offerings in a public cloud. Yuxuan Jiang 0001, Mohammad Shahrad, David Wentzlaff, Danny H. K. Tsang, Carlee Joe-Wong |
IEEE/ACM Trans. Netw. | 2 |
| 2019 | Burstable Instances for Clouds: Performance Modeling, Equilibrium Analysis, and Revenue MaximizationabstractLeading cloud providers recently introduced a new instance type named burstable instances to better match the time-varying workloads of tenants and further reduce their costs. In the research community, however, little has been done to understand burstable instances from a theoretical perspective. This paper presents the first unified framework to model, analyze, and optimize the operation of burstable instances. Specifically, we model the resource provisioning of burstable instances in different service classes, identify key performance metrics, and derive the performance given the resource provisioning decisions. We then characterize the equilibrium behind tenants' responses to the prices offered for different burstable instance service classes, taking into account the impact of tenants' actions on the performance achieved by each service class. In addition, we investigate how a cloud provider can leverage the knowledge of this equilibrium to find the prices that maximize its total revenue. Finally, we validate our framework on real traces and demonstrate its usage to price a public cloud. Yuxuan Jiang 0001, Mohammad Shahrad, David Wentzlaff, Danny H. K. Tsang, Carlee Joe-Wong |
INFOCOM | 2 |
| 2019 | Architectural Implications of Function-as-a-Service ComputingabstractServerless computing is a rapidly growing cloud application model, popularized by Amazon's Lambda platform. Serverless cloud services provide fine-grained provisioning of resources, which scale automatically with user demand. Function-as-a-Service (FaaS) applications follow this serverless model, with the developer providing their application as a set of functions which are executed in response to a user- or system-generated event. Functions are designed to be short-lived and execute inside containers or virtual machines, introducing a range of system-level overheads. This paper studies the architectural implications of this emerging paradigm. Using the commercial-grade Apache OpenWhisk FaaS platform on real servers, this work investigates and identifies the architectural implications of FaaS serverless computing. The workloads, along with the way that FaaS inherently interleaves short functions from many tenants frustrates many of the locality-preserving architectural structures common in modern processors. In particular, we find that: FaaS containerization brings up to 20x slowdown compared to native execution, cold-start can be over 10x a short function's execution time, branch mispredictions per kilo-instruction are 20x higher for short functions, memory bandwidth increases by 6x due to the invocation pattern, and IPC decreases by as much as 35% due to inter-function interference. We open-source FaaSProfiler, the FaaS testing and profiling platform that we developed for this work. Mohammad Shahrad, Jonathan Balkind, David Wentzlaff |
MICRO | 1 |
| 2018 | Power and Energy Characterization of an Open Source 25-Core Manycore ProcessorabstractThe end of Dennard's scaling and the looming power wall have made power and energy primary design goals for modern processors. Further, new applications such as cloud computing and Internet of Things (IoT) continue to necessitate increased performance and energy efficiency. Manycore processors show potential in addressing some of these issues. However, there is little detailed power and energy data on manycore processors. In this work, we carefully study detailed power and energy characteristics of Piton, a 25-core modern open source academic processor, including voltage versus frequency scaling, energy per instruction (EPI), memory system energy, network-on-chip (NoC) energy, thermal characteristics, and application performance and power consumption. This is the first detailed power and energy characterization of an open source manycore design implemented in silicon. The open source nature of the processor provides increased value, enabling detailed characterization verified against simulation and the ability to correlate results with the design and register transfer level (RTL) model. Additionally, this enables other researchers to utilize this work to build new power models, devise new research directions, and perform accurate power and energy research using the open source processor. The characterization data reveals a number of interesting insights, including that operand values have a large impact on EPI, recomputing data can be more energy efficient than loading it from memory, on-chip data transmission (NoC) energy is low, and insights on energy efficient multithreaded core design. All data collected and the hardware infrastructure used is open source and available for download at http://www.openpiton.org. Michael McKeown, Alexey Lavrov, Mohammad Shahrad, Paul J. Jackson, Yaosheng Fu, Jonathan Balkind, Tri Minh Nguyen 0003, Katie Lim, Yanqi Zhou, David Wentzlaff |
HPCA | 3 |
| 2017 | Incentivizing self-capping to increase cloud utilizationabstractCloud Infrastructure as a Service (IaaS) providers continually seek higher resource utilization to better amortize capital costs. Higher utilization not only can enable higher profit for IaaS providers but also provides a mechanism to raise energy efficiency; therefore creating greener cloud services. Unfortunately, achieving high utilization is difficult mainly due to infrastructure providers needing to maintain spare capacity to service demand fluctuations. Mohammad Shahrad, Cristian Klein, Liang Zheng 0002, Mung Chiang, Erik Elmroth, David Wentzlaff |
SoCC | 1 |
| 2017 | Symmetric split-row LDPC decodersabstractLDPC codes are deployed in many modern wired and wireless communication systems. While fully-parallel LDPC decoders are very efficient, they typically suffer from routing complexity. The Split-Row method effectively reduces this complexity with a minor performance loss. This paper shows the importance of symmetry in Split-Row architectures and proves that the implementation of Split-Row decoders based on new proposed smart column-permuted versions of parity check matrices leads to a better error performance as well as a more efficient hardware. Moreover, in order to achieve optimized column-permuted parity check matrices, a heuristic approach is proposed. This method is then generalized to support QC-LDPC codes. Applied to IEEE 802.3an (10GBASE-T Ethernet) and IEEE 802.11n (Wi-Fi) LDPC decoders, the new technique improves the error performance, while leading to almost 3× speed-up in the synthesis compile time and about 10% reduction in the critical path. Mohammad Shahrad, Mahdi Shabany |
ISCAS | 1 |
| 2016 | OpenPiton: An Open Source Manycore Research FrameworkabstractIndustry is building larger, more complex, manycore processors on the back of strong institutional knowledge, but academic projects face difficulties in replicating that scale. To alleviate these difficulties and to develop and share knowledge, the community needs open architecture frameworks for simulation, synthesis, and software exploration which support extensibility, scalability, and configurability, alongside an established base of verification tools and supported software. In this paper we present OpenPiton, an open source framework for building scalable architecture research prototypes from 1 core to 500 million cores. OpenPiton is the world's first open source, general-purpose, multithreaded manycore processor and framework. OpenPiton leverages the industry hardened OpenSPARC T1 core with modifications and builds upon it with a scratch-built, scalable uncore creating a flexible, modern manycore design. In addition, OpenPiton provides synthesis and backend scripts for ASIC and FPGA to enable other researchers to bring their designs to implementation. OpenPiton provides a complete verification infrastructure of over 8000 tests, is supported by mature software tools, runs full-stack multiuser Debian Linux, and is written in industry standard Verilog. Multiple implementations of OpenPiton have been created including a taped-out 25-core implementation in IBM's 32nm process and multiple Xilinx FPGA prototypes. Jonathan Balkind, Michael McKeown, Yaosheng Fu, Tri Minh Nguyen 0003, Yanqi Zhou, Alexey Lavrov, Mohammad Shahrad, Adi Fuchs, Samuel Payne, Xiaohua Liang, Matthew Matl, David Wentzlaff |
ASPLOS | 7 |
| 2016 | Availability Knob: Flexible User-Defined Availability in the CloudabstractFailure is inevitable in cloud environments. Finding the root cause of a failure can be very complex or at times nearly impossible. Different cloud customers have varying availability demands as well as a diverse willingness to pay for availability. In contrast to existing solutions that try to provide higher and higher availability in the cloud, we propose the Availability Knob (AK). AK provides flexible, user-defined, availability in IaaS clouds, allowing the IaaS cloud customer to express their desire for availability to the cloud provider. Complementary to existing high-reliability solutions and not requiring hardware changes, AK enables more efficient markets. This leads to reduced provider costs, increased provider profit, and improved user satisfaction when compared to an IaaS cloud with no ability to convey availability needs. We leverage game theory to derive incentive compatible pricing, which not only enables AK to function with no knowledge of the root cause of failure but also function under adversarial situations where users deliberately cause downtime. We develop a high-level stochastic simulator to test AK in large-scale IaaS clouds over long time periods. We also prototype AK in OpenStack to explore availability-API tradeoffs and to provide a grounded, real-world, implementation. Our results show that deploying AK leads to more than 10% cost reduction for providers and improves user satisfaction. It also enables providers to set variable profit margins based on the risk of not meeting availability guarantees and the disparity in availability supply/demand. Variable profit margins enable cloud providers to improve their profit by as much as 20%. Mohammad Shahrad, David Wentzlaff |
SoCC | 1 |
| 2016 | Piton: A 25-core academic manycore research processorabstractPresents a collection of slides covering the following: many-core processors; cloud computing; data warehouses; and data centers. Michael McKeown, Yaosheng Fu, Tri Minh Nguyen 0003, Yanqi Zhou, Jonathan Balkind, Alexey Lavrov, Mohammad Shahrad, Samuel Payne, David Wentzlaff |
Hot Chips Symposium | 7 |
| 2015 | TTCN: A new approach for low-power split-row LDPC decodersabstractSplit-Row technique is proved to be one of the most effective methods to reduce the routing complexity of fully-parallel LDPC decoders. This technique is based on the idea of splitting each check node processor to multiple smaller processors. This paper introduces a new method, to increase the power-efficiency of Split-Row LDPC decoders. The proposed method is called trust to the truthful check node (TTCN), enabling the decoder to only depend on a portion of check node processors at specific decoding iterations. This leads to an average reduction of 30%-40% in the check node dynamic power consumption. This is achieved by means of trust to a minority of check node processors and gating the others. In fact, a great side effect of the proposed method is also a slight improvement in the error performance. Mohammad Shahrad, Mahdi Shabany |
ISCAS | 1 |