EDBT 2026 Demo / reviewers in the wild / expert
Avani Wildani
dblp:69/8170
· DBLP profile ↗
31ranked-venue papers
8as first author
12since 2021 · last 2026
0000-0001-9457-8863ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 22 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Security and privacy · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Pontus: Identifying intrusions from massive logs via accurate provenance clustering and efficient graph serialization with minimum provenance lossabstractIdentifying intrusions from massive logs has long been a great challenge. To address this issue, this paper proposes Pontus, a novel host-based intrusion detection method via accurate provenance clustering, efficient graph serialization and classification with minimum provenance loss. Pontus first utilizes a novel multi-round label propagation algorithm (MLPA) based on overlapping community discovery to cluster the behavior instances that constitute user behavior accurately. In this way, Pontus can analyze behavior instances to extract behavior features effectively while reducing the analysis workload. Then, Pontus enables efficient graph serialization via neighbor node aggregation to convert the behavior instance into vectors while maximizing the retention of provenance information. Finally, Pontus uses a hybrid method that combines the convolutional autoencoder with Bisecting Kmeans clustering to accurately extract the provenance features of behavior instances to identify host-based intrusions. The experimental results show that compared with the state-of-the-art methods, Pontus’s accuracy increases by an average of 0.243, and F1-score increases by an average of 0.201, with small runtime overheads. Yulai Xie 0002, Heyu Zhang, Yafeng Wu, Dan Feng 0001, Pan Zhou 0001, Avani Wildani |
Expert Syst. Appl. | 8 |
| 2025 | Valet: Efficient Data Placement on Modern SSDsabstractThe increasing demand for ssds coupled with scaling difficulties has left manufacturers scrambling for newer ssd interfaces which promise better performance and durability. While these interfaces reduce the rigidity of traditional abstractions, they require application or system-level changes that can impact the stability, security, and portability of systems. To make matters worse, such changes are rendered futile with the introduction of next-generation interfaces. It is therefore no surprise that such interfaces have seen limited adoption, leaving behind a graveyard of experimental interfaces ranging from open-channel ssds to stream ssds. Devashish R. Purandare, Peter Alvaro, Avani Wildani, Darrell D. E. Long, Ethan L. Miller |
SoCC | 3 |
| 2025 | Rethinking Web Cache Design for the AI EraabstractWeb caches have long been effective at reducing latency and backend load by storing popular content close to users, exploiting the temporal and spatial locality of human-driven access patterns. However, the rise of AI-generated traffic is challenging this assumption. AI agents such as search crawlers and data scrapers issue large volumes of diverse, low-referrer requests with minimal reuse, which degrade cache effectiveness, interfere with human-relevant content, and increase pressure on backend systems. In this paper, we argue that caching infrastructure must evolve to address this shift. We analyze emerging AI traffic patterns and study their impact on caching performance using a CDN prototype based on Wikimedia's architecture. Our results show that even modest amounts of AI traffic lead to significant cache inefficiency. We envision a workload-aware caching paradigm that serves human and AI traffic through differentiated tiers and policies, preserving responsiveness for users while adapting to the diverse access patterns and requirements of AI workloads. Yazhuo Zhang, Jinqing Cai, Avani Wildani, Ana Klimovic |
SoCC | 3 |
| 2025 | Spatially Adaptive PM2.5 Estimation in Low-Sensor Regions Using Variational Gaussian ProcessesabstractAir pollution, particularly particulate matter 2.5 (PM2.5), poses a significant public health challenge in densely populated developing regions. Moreover, deploying an extensive ground sensor network to monitor PM2.5accurately is economically unfeasible in such regions. To address this problem, we utilize Sparse Variational Gaussian Process (SVGP) models to generate approximate data using the limited ground sensor data. Since SVGPs use computational approximators for Gaussian Process modeling, we hypothesize that their inducing points can be trained to adapt spatially, i.e., these points, when optimized, can spread over the region of interest. Hence, well-initialized inducing points allow SVGPs to model PM2.5data by capturing spatial variations of the region. We evaluate our hypothesis using PM2.5data from Lima, Peru, one of the most polluted cities in the Americas, and with very few PM2.5ground sensors. Our experiments qualitatively validate our hypothesis of spatial adaptation and provide a quantitative justification of improved performance over the baseline models. Shrey Gupta, Avani Wildani, Yang Liu 0037 |
DSAA | 2 |
| 2025 | Efficient intrusion detection via heterogeneous graph attention networks and parallel provenance analysis
Yulai Xie 0002, Shixun Zhao, Pan Zhou 0001, Dan Feng 0001, Avani Wildani, Yafeng Wu |
Comput. Networks | 6 |
| 2024 | Rethinking the Networking Stack for Serverless Environments: A Sidecar ApproachabstractServerless platforms rely on legacy networking stacks for communication and data movement. We quantitatively analyze the performance of these stacks and show their mismatch with highly consolidated, virtualized modern serverless environments, focusing on Firecracker, the most common serverless virtualization framework. As serverless applications grow in complexity and interaction, the resulting network bottleneck is a prime source of user-perceived, end-to-end latency. In this paper, we present a detailed vision of a new, sidecar-based networking stack for serverless environments. Our primary design goal is to provide low-overhead networking while maintaining existing security guarantees. We outline the research challenges in both the control and the data plane that the community needs to tackle before such a sidecar architecture can be used in practice. Vishwanath Seshagiri, Vahab Jabrayilov, Avani Wildani, Kostis Kaffes |
SoCC | 4 |
| 2024 | Accurate Generation of I/O Workloads Using Generative Adversarial NetworksabstractIt is essential to utilize a large number of I/O workloads to analyze commodity system performance or simulate scientific phenomena in high-performance scientific computing. I/O traces are often unavailable at scale due to trace storage overhead, privacy concerns, and the performance impact of trace instrumentation. We study how to generate sufficiently representative I/O workloads using Generative Adversarial Networks (GANs). The best GAN architecture can generate I/O workloads with maximum mean discrepancy (MMD) as low as 0.015-0.05, which implies the synthetic I/O workloads have successfully learned the potential distribution of real I/O traces. We demonstrate that the performance similarity between the original I/O trace and the generated I/O workload through trace replay can be 90.36%-97.32%. Heyu Zhang, Yulai Xie 0002, Yafeng Wu, Dan Feng 0001, Avani Wildani, Darrell D. E. Long |
NAS | 7 |
| 2024 | Spatial Transfer Learning for Estimating PM2.5 in Data-Poor Regions
Shrey Gupta, Yongbee Park, Jianzhao Bi, Suyash Gupta 0001, Andreas Züfle, Avani Wildani, Yang Liu 0037 |
ECML/PKDD (9) | 6 |
| 2024 | Accelerating multi-tier storage cache simulations using knee detection
Tyler Estro, Mário Antunes 0001, Pranav Bhandari, Anshul Gandhi, Geoffrey H. Kuenning, Carl A. Waldspurger, Avani Wildani, Erez Zadok |
Perform. Evaluation | 8 |
| 2023 | Guiding Simulations of Multi-Tier Storage Caches Using Knee DetectionabstractSimulating storage cache hierarchies enables efficient exploration of their configuration space, including diverse topologies, parameters and policies, and devices with varied performance characteristics, while avoiding expensive physical experiments. Miss Ratio Curves (MRCs) efficiently characterize the performance of a cache over a range of cache sizes. These useful tools reveal “key points” for cache simulation, such as knees in the curve that immediately follow sharp cliffs. Unfortunately, there are no automated techniques for efficiently finding key points in MRCs, and the cross-application of existing knee-detection algorithms yields inaccurate results. We present a multi-stage framework that identifies key points in any MRC, for both stack-based (e.g., LRU) and more sophis-ticated eviction algorithms (e.g., ARC). Our approach quickly locates candidates using efficient hash-based sampling, curve simplification, knee detection, and novel post-processing filters. We introduce Z-Method, a new multi-knee detection algorithm that employs statistical outlier detection to choose promising points robustly and efficiently. We evaluate our framework against seven other knee-detection algorithms, using both ARC and LRU MRCs from 106 diverse real-world workloads, and apply it to identify key points in multi-tier MRCs. Compared to naive approaches, our framework reduces the total number of points needed to accurately identify the best two-tier cache hierarchies by an average factor of approximately$5.5\times$for ARC and$7.7\times$for LRU. Tyler Estro, Mário Antunes 0001, Pranav Bhandari, Anshul Gandhi, Geoffrey H. Kuenning, Carl A. Waldspurger, Avani Wildani, Erez Zadok |
MASCOTS | 8 |
| 2023 | Portunus: Re-imagining Access Control in Distributed Systems
Watson Ladd, Tanya Verma, Marloes Venema, Armando Faz-Hernández, Brendan McMillion, Avani Wildani, Nick Sullivan |
USENIX ATC | 6 |
| 2023 | Paradise: Real-Time, Generalized, and Distributed Provenance-Based Intrusion DetectionabstractIdentifying intrusion from massive and multi-source logs accurately and in real-time presents challenges for today's users. This article presents Paradise, a real-time, generalized, and distributed provenance-based intrusion detection method. Paradise introduces a novel extract strategy to prune and extract process feature vectors from provenance dependencies at the system log level, and it stores them in high-efficiency memory databases. Using this strategy, Paradise does not depend on the specific operating system type or provenance collection framework. Provenance-based dependencies are calculated independently during the detection phase, thus, Paradise can negotiate all detection results from multiple detectors without extra communication overhead between detectors. Paradise also employs an efficient load-balanced distribution scheme that enhances the Kafka architecture to efficiently distribute provenance graph feature vectors to the detectors. The experimental results demonstrate that our method has a high detection accuracy with a low time overhead. Yafeng Wu, Yulai Xie 0002, Xuelong Liao, Pan Zhou 0001, Dan Feng 0001, Avani Wildani, Darrell D. E. Long |
IEEE Trans. Dependable Secur. Comput. | 8 |
| 2020 | Position: Can Microservices Drive a Renaissance in Workload-Aware Storage Management?
Pranav Bhandari, Avani Wildani, Dimitrios Skourtis, Vasily Tarasov, Deepavali Bhagwat, Lukas Rupprecht, Ali Anwar 0001 |
HotStorage | 2 |
| 2020 | Desperately Seeking ... Optimal Multi-Tier Cache Configurations
Tyler Estro, Pranav Bhandari, Avani Wildani, Erez Zadok |
HotStorage | 3 |
| 2020 | Autonomic Formation of Large-Scale Wireless Mesh NetworksabstractReal-world deployments of low-cost, peer-to-peer Wireless Mesh Networks (WMNs) for communication in under-served settings are hampered by low throughput capacity and high complexity of network control. We present a design of autonomic agents that manipulate the formation of WMN topologies by organizing a node placement into dynamic network partitions while enforcing inter-partition connectivity to promote the WMNs' capacity through density control, increased frequency diversity and multi-domain SDN-based control. We show that our competing Self-Organizing and Self-Healing agents achieve fast convergence to stable partition sets and global re-connectivity, relying on local information. Moreover, the design achieves global inter-partition connectivity with less than 20% of healing agents on nodes, converging under extreme node churn conditions. The design is robust to the average node placement density, producing partitions isolated at the physical and link layers with the properties of bounded diameter and node degree, and elected partition control node to act as an SDN domain controller. Sergio Gramacho, Felipe Gramacho, Avani Wildani |
NetSoft | 3 |
| 2019 | Autonomic Partitioning for the Smart Control of Wireless Mesh NetworksabstractReal-world deployments of low-cost, peer-to-peer Wireless Mesh Networks (WMNs) for communication in under-served settings are hampered by low throughput capacity and high complexity of network control. We present an autonomic smart-agent based WMN design that organizes a node placement (NP) into dynamic network partitions to promote the WMNs' capacity through increased frequency diversity and advanced network control. Through a custom-built simulation framework that supports agent decision making under real concurrency settings, we show that our smart partitioning technique achieves fast convergence to stable partition sets, relying on local information, under extreme node churn conditions. The design is robust to the average WMN NP density and produces partitions isolated at the physical and link layers with the properties of bounded diameter and node degree, and elected partition control node. Sergio Gramacho, Felipe Gramacho, Avani Wildani |
WiMob | 3 |
| 2017 | Mithril: mining sporadic associations for cache prefetchingabstractThe growing pressure on cloud application scalability has accentuated storage performance as a critical bottleneck. Although cache replacement algorithms have been extensively studied, cache prefetching - reducing latency by retrieving items before they are actually requested - remains an underexplored area. Existing approaches to history-based prefetching, in particular, provide too few benefits for real systems for the resources they cost. Juncheng Yang, Reza Karimi, Trausti Saemundsson, Avani Wildani, Ymir Vigfusson |
SoCC | 4 |
| 2017 | Snapshot Judgements: Obtaining Data Insights without Tracing
Ian F. Adams, Avani Wildani |
HotStorage | 3 |
| 2017 | Fighting for a Niche: An Evolutionary Model of Storage
Avani Wildani |
HotStorage | 1 |
| 2016 | Effects of prolonged media usage and long-term planning on archival systemsabstractIn archival systems, storage media are often replaced much earlier than their expected service life in exchange for other benefits of new media, such as higher capacity, bandwidth, and I/O operations per second, or lower costs. In an era of decreasing media density growth rates, retiring media early by considering only short-term benefits while discarding potential long-term cost benefits could have a negative long-term impact on an archival system's economics. To extend an archival system's life, at low cost, while limiting performance degradation, we suggest extending media lifetime past manufacturer recommendations as well as increasing the horizon for planning and provisioning future media purchases. We present a cost-benefit analysis of the impact of prolonged media usage and long-term planning. Through Monte Carlo simulation, we simulate the behavior of an archival system using tapes, hard disk drives (HDDs), solid state devices (SSDs), and Blu-ray discs. We show that leaving older media in the archival system makes economic sense for SSDs without significantly affecting reliability; we show cost improvements of approximately 10% for SSDs for a low annual media density growth rate, such as 5%, which would have been a loss of 35%, for a high annual media density rate, such as 20%. We show that, for SSDs and hard disks, the optimal planning time of an archival system is at least as long as the media service life. Combining prolonged media usage with an extended planning horizon reduced costs by 15% for a system using SSDs. Avani Wildani, Ethan L. Miller, David S. H. Rosenthal, Darrell D. E. Long |
MSST | 2 |
| 2016 | Can We Group Storage? Statistical Techniques to Identify Predictive Groupings in Storage System AccessesabstractStoring large amounts of data for different users has become the new normal in a modern distributed cloud storage environment. Storing data successfully requires a balance of availability, reliability, cost, and performance. Typically, systems design for this balance with minimal information about the data that will pass through them. We propose a series of methods to derive groupings from data that have predictive value, informing layout decisions for data on disk. Unlike previous grouping work, we focus on dynamically identifying groupings in data that can be gathered from active systems in real time with minimal impact using spatiotemporal locality. We outline several techniques we have developed and discuss how we select particular techniques for particular workloads and application domains. Our statistical and machine-learning-based grouping algorithms answer questions such as “What can a grouping be based on?” and “Is a given grouping meaningful for a given application?” We design our models to be flexible and require minimal domain information so that our results are as broadly applicable as possible. We intend for this work to provide a launchpad for future specialized system design using groupings in combination with caching policies and architectural distinctions such as tiered storage to create the next generation of scalable storage systems. Avani Wildani, Ethan L. Miller |
ACM Trans. Storage | 1 |
| 2015 | A Case for Rigorous Workload ClassificationabstractTraditional workload labels such as "archival" and "HPC" are poorly understood and inconsistently applied. As usage of systems has evolved, the language to describe this usage has stagnated. To better understand how workload type translates into system design requirements, we use a combination of longitudinal analysis and statistical feature extraction to categorize workload traces and study how the properties of classical workload types, such as the "write-once, read-maybe" assumption for archives, have evolved over time. Once this step is complete, we intend to move to active classification of workloads to replace these broad, poorly specified categories with quantitative metrics that can be used to improve metrics such as power, availability, and performance by mathematically relating storage algorithms with workload properties. Avani Wildani, Ian F. Adams |
MASCOTS | 1 |
| 2014 | An Economic Perspective of Disk vs. Flash Media in Archival StorageabstractFor three decades, Kryder's law correctly predicted an exponential increase in bit density on disk platters, leading to an exponential drop in cost per gigabyte, and thus to an entrenched expectation that if data could be stored for a few years the incremental cost of storing it forever would be minimal. However, disk now is over 7 times as expensive as Kryder's law would have predicted, and industry projections suggest that in 2020 the gap will reach 200 times, disrupting this expectation. Our model shows that archives based upon alternative media are surprisingly cost competitive with archives based upon traditional disk media over the long-term. We propose using Archival Flash for long-term data preservation, with the trade off between longer data retention period and lower write cycles. Avani Wildani, Ethan L. Miller, Daniel C. Rosenthal, Ian F. Adams, Christina E. Strong, Andy Hospodor |
MASCOTS | 2 |
| 2014 | PERSES: Data Layout for Low Impact FailuresabstractGrowth in disk capacity continues to outpace advances in read speed and device reliability. This has led to storage systems spending increasing amounts of time in a degraded state while failed disks reconstruct. Users and applications that do not use the data on the failed or degraded drives are negligibly impacted by the failure, increasing the perceived performance of the system. We leverage this observation with PERSES, a statistical data allocation scheme to reduce the performance impact of reconstruction after disk failure. PERSES reduces degradation from the perspective of the user by clustering data on disks such that data with high probability of co-access is placed on the same device as often as possible. Trace-driven simulations show that, by laying out data with PERSES, we can reduce the perceived time lost due to failure over three years by up to 80% compared to arbitrary allocation. Avani Wildani, Ethan L. Miller, Ian F. Adams, Darrell D. E. Long |
MASCOTS | 1 |
| 2013 | HANDS: A heuristically arranged non-backup in-line deduplication systemabstractDeduplicating in-line data on primary storage is hampered by the disk bottleneck problem, an issue which results from the need to keep an index mapping portions of data to hash values in memory in order to detect duplicate data without paying the performance penalty of disk paging. The index size is proportional to the volume of unique data, so placing the entire index into RAM is not cost effective with a deduplication ratio below 45%. HANDS reduces the amount of in-memory index storage required by up to 99% while still achieving between 30% and 90% of the deduplication a full memory-resident index provides, making primary deduplication cost effective in workloads with deduplication rates as low as 8%. HANDS is a framework that dynamically pre-fetches fingerprints from disk into memory cache according to working sets statistically derived from access patterns. We use a simple neighborhood grouping as our statistical technique to demonstrate the effectiveness of our approach. HANDS is modular and requires only spatio-temporal data, making it suitable for a wide range of storage systems without the need to modify host file systems. Avani Wildani, Ethan L. Miller, Ohad Rodeh |
ICDE | 1 |
| 2013 | Validating Storage System InstrumentationabstractThere is a large body of work-such as system administration and intrusion detection-that relies upon storage system logs and snapshots. These solutions rely on accurate system records, however, little effort has been made to verify the correctness of logging instrumentation and log reliability. We present a solution, called ExDiff, that uses expectation differencing to validate storage system logs. Our solution can identify development errors such as the omission of a logging point and runtime errors such as log crashes. ExDiff uses metadata snapshots and activity logs to predict the expected state of the system and compares that with the system's actual state. Mismatches between the expected and actual metadata states can then be used to highlight gaps in log coverage, as well as aid in identifying specific types of missing entries. We show that ExDiff provides valuable insight to system designers, administrators and researchers by accurately identifying gaps in log coverage, providing clues useful in isolating specific types of missing log entries, and highlighting potential misunderstandings in logged action. Ian F. Adams, Mark W. Storer, Avani Wildani, Ethan L. Miller, Brian A. Madden |
MASCOTS | 3 |
| 2013 | Single-Snapshot File System AnalysisabstractMetadata snapshots are a common method for gaining insight into file systems due to their small size and relative ease of acquisition. Since they are static, most researchers have used them for relatively simple analyses such as file size distributions and age of files. We hypothesize that it is possible to gain much richer insights into file system and user behavior by clustering features in metadata snapshots and comparing the entropy within clusters to the entropy within natural partitions such as directory hierarchies. We discuss several different methods for gaining deeper insights into metadata snapshots, and show a small proof of concept using data from Los Alamos National Laboratories. In our initial work, we see evidence that it is possible to identify user locality information, traditionally the purview of dynamic traces, using a single static snapshot. Avani Wildani, Ian F. Adams, Ethan L. Miller |
MASCOTS | 1 |
| 2013 | Examining extended and scientific metadata for scalable index designsabstractWhile file system metadata is well characterized by a variety of workload studies, scientific metadata is much less well understood. We characterize scientific metadata, in order to better understand the implications for index design. Based on our findings, existing solutions for either file system or scientific search will not suffice for indexing a large scientific file system. We describe the problems with existing solutions, and suggest column stores as an alternative approach. Aleatha Parker-Wood, Darrell D. E. Long, Brian A. Madden, Ian F. Adams, Michael McThrow, Avani Wildani |
SYSTOR | 6 |
| 2011 | Efficiently identifying working sets in block I/O streamsabstractIdentifying groups of blocks that tend to be read or written together in a given environment is the first step towards powerful techniques for device failure isolation and power management. For example, identified groups can be placed together on a single disk, avoiding excess drive activity across an exascale storage system. Unlike previous grouping work, we focus on identifying groupings in data that can be gathered from real, running systems with minimal impact. Using temporal, spatial, and access ordering information from an enterprise data set, we identified a set of groupings that consistently appear, indicating that these are working sets that are likely to be accessed together. We present several techniques to obtain groupings along with a discussion of what techniques best apply to particular types of real systems. We intend to use these preliminary results to inform our search for new types of workloads with a goal of identifying properties of easily separable workloads across different systems and dynamically moving groups in these workloads to reduce disk activity in large storage systems. Avani Wildani, Ethan L. Miller, Lee Ward |
SYSTOR | 1 |
| 2010 | GAUL: Gestalt Analysis of Unstructured Logs for Diagnosing Recurring Problems in Large Enterprise Storage SystemsabstractWe present GAUL, a system to automate the whole log comparison between a new problem and the ones diagnosed in the past to identify recurring problems. GAUL uses a fuzzy match algorithm based on the contextual overlap between log lines and efficiently implements this using scalable index/search. The accuracy and efficiency of the comparison is further improved by leveraging problem set information and noise tolerance techniques. We evaluate GAUL using 4339 customer problems that occurred in all field deployments of an enterprise storage system over the course of a year. Our results show that with human-filtered logs, GAUL can identify the correct problem set 66% of the time among the top10 matches, which is 15% more accurate than the VSM system that uses cosine similarity and 19% more accurate than the ERRCMP system that uses error codes for log comparison. With unfiltered logs, the top10 match accuracy of GAUL is 40%, which is 22% more accurate than VSM and 26% more accurate than ERRCMP. Pin Zhou, Binny S. Gill, Wendy Belluomini, Avani Wildani |
SRDS | 4 |
| 2009 | Protecting against rare event failures in archival systemsabstractDigital archives are growing rapidly, necessitating stronger reliability measures than RAID to avoid data loss from device failure. Mirroring, a popular solution, is too expensive over time. We present a compromise solution that uses multi-level redundancy coding to reduce the probability of data loss from multiple simultaneous device failures. This approach handles small-scale failures of one or two devices efficiently while still allowing the system to survive rare-event, larger-scale failures of four or more devices. In our approach, each disk is split into a set of fixed size disklets which are used to construct reliability stripes. To protect against rare event failures, reliability stripes are grouped into larger super-groups, each of which has a corresponding super-parity; super-parity is only used to recover data when disk failures overwhelm the redundancy in a single reliability stripe. Super-parity can be stored on a variety of devices such as NV-RAM and always-on disks to offset write bottlenecks while still keeping the number of active devices low. Our calculations of failure probabilities show that adding super-parity allows our system to absorb many more disk failures without data loss. Through discrete event simulation, we found that adding super-groups has a significant impact on mean time to data loss and that rebuilds are slow but not unmanageable. Finally, we showed that robustness against rare events can be achieved for a fraction of total system cost. Avani Wildani, Thomas J. E. Schwarz, Ethan L. Miller, Darrell D. E. Long |
MASCOTS | 1 |