EDBT 2026 Demo / reviewers in the wild / expert
Elizabeth Varki
dblp:44/6651
· DBLP profile ↗
9ranked-venue papers
4as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 4 first-authorSoftware engineering, systems software and programming languages · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Storage systems · 66% Performance modeling and evaluation · 23% Memory systems · 8% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 50% Computational science and engineering · 50% |
Topics — the 13 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
bioinformatics workflow |
0.5 | 1 | 2021 | RepeatFS: a file system providing reproducibility through provenance and automation · Bioinform. 2021 |
Computational science and engineering
computational reproducibility |
0.5 | 1 | 2021 | RepeatFS: a file system providing reproducibility through provenance and automation · Bioinform. 2021 |
Storage systems
file systems |
0.5 | 1 | 2021 | RepeatFS: a file system providing reproducibility through provenance and automation · Bioinform. 2021 |
Memory systems › cache
prefetching |
0.1 | 1 | 2008 | TaP: Table-based Prefetching for Storage Caches · FAST 2008 |
Performance modeling and evaluation › queueing models › parallel-server system
fork-join queue |
0.1 | 3 | 2001 | Response Time Analysis of Parallel Computer and Storage Systems · IEEE Trans. Parallel Distributed Syst. 2001 Mean Value Technique for Closed Fork-Join Networks · SIGMETRICS 1999 Analysis of Balanced Fork-Join Queueing Networks · SIGMETRICS 1996 |
Performance modeling and evaluation
queueing models |
0.1 | 3 | 2001 | Response Time Analysis of Parallel Computer and Storage Systems · IEEE Trans. Parallel Distributed Syst. 2001 Mean Value Technique for Closed Fork-Join Networks · SIGMETRICS 1999 Analysis of Balanced Fork-Join Queueing Networks · SIGMETRICS 1996 |
Storage systems
disk array |
0.0 | 1 | 2004 | Issues and Challenges in the Performance Analysis of Real Disk Arrays · IEEE Trans. Parallel Distributed Syst. 2004 |
Storage systems › disk array
disk array performance |
0.0 | 1 | 2004 | Issues and Challenges in the Performance Analysis of Real Disk Arrays · IEEE Trans. Parallel Distributed Syst. 2004 |
Performance modeling and evaluation
storage performance modeling |
0.0 | 1 | 2004 | Issues and Challenges in the Performance Analysis of Real Disk Arrays · IEEE Trans. Parallel Distributed Syst. 2004 |
Embedded and real-time systems › real-time scheduling › schedulability analysis
response time analysis |
0.0 | 1 | 2001 | Response Time Analysis of Parallel Computer and Storage Systems · IEEE Trans. Parallel Distributed Syst. 2001 |
Performance modeling and evaluation › queueing models
mean value analysis |
0.0 | 1 | 1999 | Mean Value Technique for Closed Fork-Join Networks · SIGMETRICS 1999 |
Performance modeling and evaluation › queueing models
queueing network analysis |
0.0 | 1 | 1996 | Analysis of Balanced Fork-Join Queueing Networks · SIGMETRICS 1996 |
Performance modeling and evaluation
measurement-based modeling |
0.0 | 1 | 2004 | Issues and Challenges in the Performance Analysis of Real Disk Arrays · IEEE Trans. Parallel Distributed Syst. 2004 |
Methods — techniques the papers use, named apart from their topics
task automation · 1.0provenance tracking · 1.0simulation · 0.1measurement data analysis · 0.0black-box modeling · 0.0response time approximation · 0.0mean value analysis · 0.0queueing analysis · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | RepeatFS: a file system providing reproducibility through provenance and automationabstractMOTIVATION: Reproducibility is of central importance to the scientific process. The difficulty of consistently replicating and verifying experimental results is magnified in the era of big data, in which bioinformatics analysis often involves complex multi-application pipelines operating on terabytes of data. These processes result in thousands of possible permutations of data preparation steps, software versions and command-line arguments. Existing reproducibility frameworks are cumbersome and involve redesigning computational methods. To address these issues, we developed RepeatFS, a file system that records, replicates and verifies informatics workflows with no alteration to the original methods. RepeatFS also provides several other features to help promote analytical transparency and reproducibility, including provenance visualization and task automation. RESULTS: We used RepeatFS to successfully visualize and replicate a variety of bioinformatics tasks consisting of over a million operations with no alteration to the original methods. RepeatFS correctly identified all software inconsistencies that resulted in replication differences. AVAILABILITYAND IMPLEMENTATION: RepeatFS is implemented in Python 3. Its source code and documentation are available at https://github.com/ToniWestbrook/repeatfs. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Anthony Westbrook, Elizabeth Varki, William Kelley Thomas |
Bioinform. | 2 |
| 2016 | RAIDX: RAID without StripingabstractEach disk of traditional RAID is logically divided into stripe units, and stripe units at the same location on each disk form a stripe. Thus, RAID striping forces homogeneity on its disks. Now, consider a heterogeneous array of hard disks, solid state disks, and RAM storage devices, with various access speeds and sizes. If this array is organized as a RAID system, then larger disks have wasted space and faster disks are under utilized. This paper proposes RAIDX, a new organization for a heterogeneous array. RAIDX disks are divided into chunks, larger disks have more chunks. Chunks from one or more disks are grouped into bundles, and RAIDX bundles chunks of data across its disks. The heterogeneity of disks causes unbalanced load distribution with some under-utilized disks and some bottleneck disks. To balance load across disks, RAIDX moves most frequently accessed chunks to under-utilized, faster disks and least frequently used chunks to larger, slower disks. Chunk remapping is done at the RAIDX level and does not impact file system to storage addressing. Experiments comparing local and networked RAIDX against local RAID show that RAIDX has faster throughput than RAID when the array is composed of heterogeneous disks: local RAIDX is 2.5x faster than RAID, networked RAIDX is 1.6x faster than local RAID. András Fekete, Elizabeth Varki |
MASCOTS | 2 |
| 2010 | Sequential Prefetch Cache Sizing for Maximal Hit RateabstractWe propose a prefetch cache sizing module for use with any sequential prefetching scheme and evaluate its impact on the hit rate. Disk array caches perform sequential prefetching by loading data contiguous to I/O request data into the array cache. If the I/O workload has sequential locality, then data prefetched in response to sequential accesses in the workload will receive hits. Different schemes prefetch different data, so the prefetch cache size requirement varies. Moreover, the proportion of sequential and random requests in the workload and their interleaving pattern affects the size requirement. If the cache is too small, then prefetched data would get evicted from the cache before a request for the data arrives, thus lowering the hit rate. If the cache is too large, then valuable cache space is wasted. We present a simple sizing module that can be added to any prefetching scheme to ensure that the prefetch cache size is adequately matched to the requirement of the prefetching scheme on a dynamic workload comprising multiple streams. We analytically compute the maximal hit rate achievable by popular prefetching schemes and through simulations, show that our sizing module maintains the prefetch cache at a size that nearly achieves this maximal hit rate. Swapnil Bhatia, Elizabeth Varki, Arif Merchant |
MASCOTS | 2 |
| 2008 | TaP: Table-based Prefetching for Storage Caches
Mingju Li, Elizabeth Varki, Swapnil Bhatia, Arif Merchant |
FAST | 2 |
| 2008 | Co-allocation in Data Grids: A Global, Multi-user Perspective
Adam H. Villa, Elizabeth Varki |
GPC | 2 |
| 2004 | Issues and Challenges in the Performance Analysis of Real Disk ArraysabstractThe performance modeling and analysis of disk arrays is challenging due to the presence of multiple disks, large array caches, and sophisticated array controllers. Moreover, storage manufacturers may not reveal the internal algorithms implemented in their devices, so real disk arrays are effectively black-boxes. We use standard performance techniques to develop an integrated performance model that incorporates some of the complexities of real disk arrays. We show how measurement data and baseline performance models can be used to extract information about the various features implemented in a disk array. In this process, we identify areas for future research in the performance analysis of real disk arrays. Elizabeth Varki, Arif Merchant, Jianzhang Xu, Xiaozhou Qiu |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2001 | Response Time Analysis of Parallel Computer and Storage SystemsabstractFork-join structures have gained increased importance in recent years as a means of modeling parallelism in computer and storage systems. The basic fork-join model is one in which a job arriving at a parallel system splits into K independent tasks that are assigned to K unique, homogeneous servers. In the paper, a simple response time approximation is derived for parallel systems with exponential service time distributions. The approximation holds for networks modeling several devices, both parallel and nonparallel. (In the case of closed networks containing a stand-alone parallel system, a mean response time bound is derived.) In addition, the response time approximation is extended to cover the more realistic case wherein a job splits into an arbitrary number of tasks upon arrival at a parallel system. Simulation results for closed networks with stand-alone parallel subsystems and exponential service time distributions indicate that the response time approximation is, on average, within 3 percent of the seeded response times. Similarly, simulation results with nonexponential distributions also indicate that the response time approximation is close to the seeded values. Potential applications of our results include the modeling of data placement in disk arrays and the execution of parallel programs in multiprocessor and distributed systems. Elizabeth Varki |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 1999 | Mean Value Technique for Closed Fork-Join NetworksabstractA simple technique for computing mean performance measures of closed single-class fork-join networks with exponential service time distribution is given here.This technique is similar to the mean value analysis technique for closed product-form networks and iterates on the number of customers in the network.Mean performance measures like the mean response times, queue lengths, and throughput of closed fork-join networks can be computed recursively without calculating the steady-state distribution of the network.The technique is based on the mean value equation for fork-join networks which relates the response time of a network to the mean service times at the service centers and the mean queue length of the system with one customer less.Unlike product-form networks, the mean value equation for fork-join networks is an approximation and the technique computes lower performance bound values for the fork-join network.However, it is a good approximation since the mean value equation is derived from an equation that exactly relates the response time of parallel systems to the degree of parallelism and the mean arrival queue length.Using simulation, it is shown that the relative error in the approximation is less than 5% in most cases.The error does not increase with each iteration.Permlsslon to make digital or hard copies of all or part of this work for personal or classroom use is granted without tee provided that copa are not made or distributed for pro10 01 commercial advantage and that copies bear this notice and the full cNa110n on the first Page.To copy otherwise, to republish.to post on servers Or t0 redwribute to lists.requires pnor specific Elizabeth Varki |
SIGMETRICS | 1 |
| 1996 | Analysis of Balanced Fork-Join Queueing Networks
Elizabeth Varki, Lawrence W. Dowdy |
SIGMETRICS | 1 |