Cesar A. Stuardo

dblp:204/3580 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
2since 2021 · last 2025
0000-0002-5052-0077ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-authorSoftware engineering, systems software and programming languages · 2Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Distributed systems · 66% Embedded and real-time systems · 12% Cloud and datacenter computing · 11%
Artificial intelligence
1 paper
Efficient and distributed learning · 44% Deep learning architectures and training · 44% Language models and text generation · 13%
Software engineering, system software, and programming languages
3 papers
Operating systems · 54% Software testing · 46%
Databases, data mining, and information retrieval
1 paper
Transaction processing and concurrency control · 100%

Topics — the 12 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
inference serving
0.912025
MegaScale-Infer: Efficient Mixture-of-Experts Model Serving with Disaggregated Expert Parallelism · SIGCOMM 2025
Machine learning › Deep learning architectures and training › mixture of experts
mixture-of-experts inference
0.912025
MegaScale-Infer: Efficient Mixture-of-Experts Model Serving with Disaggregated Expert Parallelism · SIGCOMM 2025
Distributed systems › distributed machine learning
expert parallelism
0.912025
MegaScale-Infer: Efficient Mixture-of-Experts Model Serving with Disaggregated Expert Parallelism · SIGCOMM 2025
Distributed systems
fault tolerance
0.722019
FlyMC: Highly Scalable Testing of Complex Interleavings in Distributed Systems · EuroSys 2019
Transactuations: Where Transactions Meet the Physical World · ACM Trans. Comput. Syst. 2018
Software testing › system testing
distributed system testing
0.522019
FlyMC: Highly Scalable Testing of Complex Interleavings in Distributed Systems · EuroSys 2019
ScaleCheck: A Single-Machine Approach for Discovering Scalability Bugs in Large Distributed Systems · FAST 2019
Embedded and real-time systems
cyber-physical system platforms
0.422019
Transactuations: Where Transactions Meet the Physical World · ACM Trans. Comput. Syst. 2018
Transactuations: Where Transactions Meet the Physical World · USENIX ATC 2019
Distributed systems
distributed system testing
0.412019
FlyMC: Highly Scalable Testing of Complex Interleavings in Distributed Systems · EuroSys 2019
Operating systems › i/o › i/o subsystem
i/o scheduling
0.312017
MittOS: Supporting Millisecond Tail Tolerance with Fast Rejecting SLO-Aware OS Interface · SOSP 2017
Storage systems › i/o architecture › i/o subsystem
i/o stack
0.312017
MittOS: Supporting Millisecond Tail Tolerance with Fast Rejecting SLO-Aware OS Interface · SOSP 2017
Cloud and datacenter computing › quality of service
tail latency
0.312017
MittOS: Supporting Millisecond Tail Tolerance with Fast Rejecting SLO-Aware OS Interface · SOSP 2017
Natural language and speech › Language models and text generation
large language model inference
0.312025
MegaScale-Infer: Efficient Mixture-of-Experts Model Serving with Disaggregated Expert Parallelism · SIGCOMM 2025
Parallel and multicore computing › data parallelism
data-parallel applications
0.112017
MittOS: Supporting Millisecond Tail Tolerance with Fast Rejecting SLO-Aware OS Interface · SOSP 2017

Methods — techniques the papers use, named apart from their topics

expert parallelism · 1.7disaggregated serving · 1.7state symmetry · 0.8parallel flips · 0.8event independence · 0.8transaction abstraction · 0.7runtime system · 0.7fast rejection · 0.6SLO prediction · 0.6
YearPublicationVenuePosition
2025 MegaScale-Infer: Efficient Mixture-of-Experts Model Serving with Disaggregated Expert Parallelism
abstract
Mixture-of-Experts (MoE) showcases tremendous potential to scale large language models (LLMs) with enhanced performance and reduced computational complexity. However, its sparsely activated architecture shifts feed-forward networks (FFNs) from being compute-intensive to memory-intensive during inference, leading to substantially lower GPU utilization and increased operational costs.
Ruidong Zhu, Ziheng Jiang, Chao Jin 0007, Cesar A. Stuardo, Huaping Zhou, Jianzhe Xiao, Lingjun Liu, Haibin Lin, Li-Wen Chang, Jianxi Ye, Xuanzhe Liu, Xin Jin 0008, Xin Liu 0086
SIGCOMM5
2022 Layered Contention Mitigation for Cloud Storage
abstract
We introduce an ecosystem of contention mitigation supports within the operating system, runtime and library layers. This ecosystem provides an end-to-end request abstraction that enables a uniform type of contention mitigation capabilities, namely request cancellation and delay prediction, that can be stackable together across multiple resource layers. Our evaluation shows that in our ecosystem, multi-resource storage applications are faster by 5-70% starting at 90P (the 90thpercentile) compared to popular practices such as speculative execution and is only 3% slower on average compared to a best-case (no contention) scenario.
Meng Wang 0056, Cesar A. Stuardo, Daniar Heri Kurniawan, Ray A. O. Sinurat, Haryadi S. Gunawi
CLOUD2
2019 FlyMC: Highly Scalable Testing of Complex Interleavings in Distributed Systems
abstract
We present a fast and scalable testing approach for datacenter/cloud systems such as Cassandra, Hadoop, Spark, and ZooKeeper. The uniqueness of our approach is in its ability to overcome the path/state-space explosion problem in testing workloads with complex interleavings of messages and faults. We introduce three powerful algorithms: state symmetry, event independence, and parallel flips, which collectively makes our approach on average 16x (up to 78x) faster than other state-of-the-art solutions. We have integrated our techniques with 8 popular datacenter systems, successfully reproduced 12 old bugs, and found 10 new bugs --- all were done without random walks or manual checkpoints.
Jeffrey F. Lukman, Huan Ke, Cesar A. Stuardo, Riza O. Suminto, Daniar Heri Kurniawan, Dikaimin Simon, Satria Priambada, Chen Tian 0002, Tanakorn Leesatapornwongsa, Aarti Gupta, Shan Lu 0001, Haryadi S. Gunawi
EuroSys3
2019 ScaleCheck: A Single-Machine Approach for Discovering Scalability Bugs in Large Distributed Systems
Cesar A. Stuardo, Tanakorn Leesatapornwongsa, Riza O. Suminto, Huan Ke, Jeffrey F. Lukman, Wei-Chiu Chuang, Shan Lu 0001, Haryadi S. Gunawi
FAST1
2019 Transactuations: Where Transactions Meet the Physical World
Aritra Sengupta, Tanakorn Leesatapornwongsa, Masoud Saeida Ardekani, Cesar A. Stuardo
USENIX ATC4
2018 Transactuations: Where Transactions Meet the Physical World
abstract
A large class of IoT applications read sensors, execute application logic, and actuate actuators. However, the lack of high-level programming abstractions compromises correctness, especially in the presence of failures and unwanted interleaving between applications. A key problem arises when operations on IoT devices or the application itself fails, which leads to inconsistencies between the physical state and application state, breaking application semantics and causing undesired consequences. Transactions are a well-established abstraction for correctness, but assume properties that are absent in an IoT context. In this article, we study one such environment, smart home, and establish inconsistencies manifesting out of failures. We propose an abstraction called transactuation that empowers developers to build reliable applications. Our runtime, Relacs , implements the abstraction atop a real smart-home platform. We evaluate programmability, performance, and effectiveness of transactuations to demonstrate its potential as a powerful abstraction and execution model.
Tanakorn Leesatapornwongsa, Aritra Sengupta, Masoud Saeida Ardekani, Gustavo Petri, Cesar A. Stuardo
ACM Trans. Comput. Syst.5
2017 PBSE: a robust path-based speculative execution for degraded-network tail tolerance in data-parallel frameworks
abstract
We reveal loopholes of Speculative Execution (SE) implementations under a unique fault model: node-level network throughput degradation. This problem appears in many data-parallel frameworks such as Hadoop MapReduce and Spark. To address this, we present PBSE, a robust, path-based speculative execution that employs three key ingredients: path progress, path diversity, and path-straggler detection and speculation. We show how PBSE is superior to other approaches such as cloning and aggressive speculation under the aforementioned fault model. PBSE is a general solution, applicable to many data-parallel frameworks such as Hadoop/HDFS+QFS, Spark and Flume.
Riza O. Suminto, Cesar A. Stuardo, Alexandra Clark, Huan Ke, Tanakorn Leesatapornwongsa, Daniar Heri Kurniawan, Vincentius Martin, Maheswara Rao G. Uma, Haryadi S. Gunawi
SoCC2
2017 Scalability Bugs: When 100-Node Testing is Not Enough
abstract
We highlight the problem of scalability bugs, a new class of bugs that appear in "cloud-scale" distributed systems. Scalability bugs are latent bugs that are cluster-scale dependent, whose symptoms typically surface in large-scale deployments, but not in small or medium-scale deployments. The standard practice to test large distributed systems is to deploy them on a large number of machines ("real-scale testing"), which is difficult and expensive. New methods are needed to reduce developers' burdens in finding, reproducing, and debugging scalability bugs. We propose "scale check," an approach that helps developers find and replay scalability bugs at real scales, but do so only on one machine and still achieve a high accuracy (i.e., similar observed behaviors as if the nodes are deployed in real-scale testing).
Tanakorn Leesatapornwongsa, Cesar A. Stuardo, Riza O. Suminto, Huan Ke, Jeffrey F. Lukman, Haryadi S. Gunawi
HotOS2
2017 MittOS: Supporting Millisecond Tail Tolerance with Fast Rejecting SLO-Aware OS Interface
abstract
MittOS provides operating system support to cut millisecond-level tail latencies for data-parallel applications. In MittOS, we advocate a new principle that operating system should quickly reject IOs that cannot be promptly served. To achieve this, MittOS exposes a fast rejecting SLO-aware interface wherein applications can provide their SLOs (e.g., IO deadlines). If MittOS predicts that the IO SLOs cannot be met, MittOS will promptly return EBUSY signal, allowing the application to failover (retry) to another less-busy node without waiting. We build MittOS within the storage stack (disk, SSD, and OS cache managements), but the principle is extensible to CPU and runtime memory managements as well. MittOS' no-wait approach helps reduce IO completion time up to 35% compared to wait-then-speculate approaches.
Mingzhe Hao, Huaicheng Li, Michael Hao Tong, Chrisma Pakha, Riza O. Suminto, Cesar A. Stuardo, Andrew A. Chien, Haryadi S. Gunawi
SOSP6