Daniar Heri Kurniawan

dblp:206/4503 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0003-2646-3321ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 2 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Storage systems · 72% Distributed systems · 25% Cloud and datacenter computing · 4%
Computer networks
1 paper
Content delivery and video streaming · 77% Network measurement and analytics · 23%
Databases, data mining, and information retrieval
2 papers
Machine learning and data management · 57% Recommender systems · 43%
Software engineering, system software, and programming languages
1 paper
Software testing · 100%

Topics — the 7 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems › flash and SSD › flash memory
flash storage
0.912025
Heimdall: Optimizing Storage I/O Admission with Extensive Machine Learning Pipeline · EuroSys 2025
Storage systems › key-value storage
embedding table storage
0.712023
EVStore: Storage and Caching Capabilities for Scaling Embedding Tables in Deep Recommendation Systems · ASPLOS (2) 2023
Storage systems
key-value storage
0.712023
EVStore: Storage and Caching Capabilities for Scaling Embedding Tables in Deep Recommendation Systems · ASPLOS (2) 2023
Software testing › system testing
distributed system testing
0.412019
FlyMC: Highly Scalable Testing of Complex Interleavings in Distributed Systems · EuroSys 2019
Distributed systems
distributed system testing
0.412019
FlyMC: Highly Scalable Testing of Complex Interleavings in Distributed Systems · EuroSys 2019
Distributed systems
fault tolerance
0.412019
FlyMC: Highly Scalable Testing of Complex Interleavings in Distributed Systems · EuroSys 2019
Machine learning and data management
learned database components
0.312025
Heimdall: Optimizing Storage I/O Admission with Extensive Machine Learning Pipeline · EuroSys 2025

Methods — techniques the papers use, named apart from their topics

noise filtering · 1.7machine learning pipeline · 1.7feature engineering · 1.7domain-specific approximation · 1.3caching · 1.3state symmetry · 0.8parallel flips · 0.8event independence · 0.8
YearPublicationVenuePosition
2025 Heimdall: Optimizing Storage I/O Admission with Extensive Machine Learning Pipeline
abstract
This paper introduces Heimdall, a highly accurate and efficient machine learning-powered I/O admission policy for flash storage, designed to operate in a black-box manner. We make domain-specific innovations in various ML stages by introducing accurate period-based labeling, 3-stage noise filtering, in-depth feature engineering, and fine-grained tuning, which together improve the decision accuracy from 67% up to 93%. We perform various deployment optimizations to reach a sub-μs inference latency and a small, 28KB, memory overhead. With 500 unbiased random experiments derived from production traces, we show Heimdall delivers 15-35% lower average I/O latency compared to the state of the art and up to 2x faster to a baseline. Heimdall is ready for user-level, in-kernel, and distributed deployments.
Daniar Heri Kurniawan, Rani Ayu Putri, Peiran Qin, Kahfi S. Zulkifli, Ray A. O. Sinurat, Janki Bhimani, Sandeep Madireddy, Achmad I. Kistijantoro, Haryadi S. Gunawi
EuroSys1
2023 EVStore: Storage and Caching Capabilities for Scaling Embedding Tables in Deep Recommendation Systems
abstract
Modern recommendation systems, primarily driven by deep-learning models, depend on fast model inferences to be useful. To tackle the sparsity in the input space, particularly for categorical variables, such inferences are made by storing increasingly large embedding vector (EV) tables in memory. A core challenge is that the inference operation has an all-or-nothing property: each inference requires multiple EV table lookups, but if any memory access is slow, the whole inference request is slow. In our paper, we design, implement and evaluate EVStore, a 3-layer EV table lookup system that harnesses both structural regularity in inference operations and domain-specific approximations to provide optimized caching, yielding up to 23% and 27% reduction on the average and p90 latency while quadrupling throughput at 0.2% loss in accuracy. Finally, we show that at a minor cost of accuracy, EVStore can reduce the Deep Recommendation System (DRS) memory usage by up to 94%, yielding potentially enormous savings for these costly, pervasive systems.
Daniar Heri Kurniawan, Ruipu Wang, Kahfi S. Zulkifli, Fandi A. Wiranata, John Bent, Ymir Vigfusson, Haryadi S. Gunawi
ASPLOS (2)1
2022 Layered Contention Mitigation for Cloud Storage
abstract
We introduce an ecosystem of contention mitigation supports within the operating system, runtime and library layers. This ecosystem provides an end-to-end request abstraction that enables a uniform type of contention mitigation capabilities, namely request cancellation and delay prediction, that can be stackable together across multiple resource layers. Our evaluation shows that in our ecosystem, multi-resource storage applications are faster by 5-70% starting at 90P (the 90thpercentile) compared to popular practices such as speculative execution and is only 3% slower on average compared to a best-case (no contention) scenario.
Meng Wang 0056, Cesar A. Stuardo, Daniar Heri Kurniawan, Ray A. O. Sinurat, Haryadi S. Gunawi
CLOUD3
2019 FlyMC: Highly Scalable Testing of Complex Interleavings in Distributed Systems
abstract
We present a fast and scalable testing approach for datacenter/cloud systems such as Cassandra, Hadoop, Spark, and ZooKeeper. The uniqueness of our approach is in its ability to overcome the path/state-space explosion problem in testing workloads with complex interleavings of messages and faults. We introduce three powerful algorithms: state symmetry, event independence, and parallel flips, which collectively makes our approach on average 16x (up to 78x) faster than other state-of-the-art solutions. We have integrated our techniques with 8 popular datacenter systems, successfully reproduced 12 old bugs, and found 10 new bugs --- all were done without random walks or manual checkpoints.
Jeffrey F. Lukman, Huan Ke, Cesar A. Stuardo, Riza O. Suminto, Daniar Heri Kurniawan, Dikaimin Simon, Satria Priambada, Chen Tian 0002, Tanakorn Leesatapornwongsa, Aarti Gupta, Shan Lu 0001, Haryadi S. Gunawi
EuroSys5
2019 E2E: embracing user heterogeneity to improve quality of experience on the web
abstract
Conventional wisdom states that to improve quality of experience (QoE), web service providers should reduce the median or other percentiles of server-side delays. This work shows that doing so can be inefficient due to user heterogeneity in how the delays impact QoE. From the perspective of QoE, the sensitivity of a request to delays can vary greatly even among identical requests arriving at the service, because they differ in the wide-area network latency experienced prior to arriving at the service. In other words, saving 50ms of server-side delay affects different users differently.
Siddhartha Sen 0001, Daniar Heri Kurniawan, Haryadi S. Gunawi, Junchen Jiang
SIGCOMM3
2017 PBSE: a robust path-based speculative execution for degraded-network tail tolerance in data-parallel frameworks
abstract
We reveal loopholes of Speculative Execution (SE) implementations under a unique fault model: node-level network throughput degradation. This problem appears in many data-parallel frameworks such as Hadoop MapReduce and Spark. To address this, we present PBSE, a robust, path-based speculative execution that employs three key ingredients: path progress, path diversity, and path-straggler detection and speculation. We show how PBSE is superior to other approaches such as cloning and aggressive speculation under the aforementioned fault model. PBSE is a general solution, applicable to many data-parallel frameworks such as Hadoop/HDFS+QFS, Spark and Flume.
Riza O. Suminto, Cesar A. Stuardo, Alexandra Clark, Huan Ke, Tanakorn Leesatapornwongsa, Daniar Heri Kurniawan, Vincentius Martin, Maheswara Rao G. Uma, Haryadi S. Gunawi
SoCC7