Elizabeth Varki

dblp:44/6651 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 4 first-authorSoftware engineering, systems software and programming languages · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Storage systems · 66% Performance modeling and evaluation · 23% Memory systems · 8%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 50% Computational science and engineering · 50%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
bioinformatics workflow
0.512021
RepeatFS: a file system providing reproducibility through provenance and automation · Bioinform. 2021
Computational science and engineering
computational reproducibility
0.512021
RepeatFS: a file system providing reproducibility through provenance and automation · Bioinform. 2021
Storage systems
file systems
0.512021
RepeatFS: a file system providing reproducibility through provenance and automation · Bioinform. 2021
Memory systems › cache
prefetching
0.112008
TaP: Table-based Prefetching for Storage Caches · FAST 2008
Performance modeling and evaluation › queueing models › parallel-server system
fork-join queue
0.132001
Response Time Analysis of Parallel Computer and Storage Systems · IEEE Trans. Parallel Distributed Syst. 2001
Mean Value Technique for Closed Fork-Join Networks · SIGMETRICS 1999
Analysis of Balanced Fork-Join Queueing Networks · SIGMETRICS 1996
Performance modeling and evaluation
queueing models
0.132001
Response Time Analysis of Parallel Computer and Storage Systems · IEEE Trans. Parallel Distributed Syst. 2001
Mean Value Technique for Closed Fork-Join Networks · SIGMETRICS 1999
Analysis of Balanced Fork-Join Queueing Networks · SIGMETRICS 1996
Storage systems
disk array
0.012004
Issues and Challenges in the Performance Analysis of Real Disk Arrays · IEEE Trans. Parallel Distributed Syst. 2004
Storage systems › disk array
disk array performance
0.012004
Issues and Challenges in the Performance Analysis of Real Disk Arrays · IEEE Trans. Parallel Distributed Syst. 2004
Performance modeling and evaluation
storage performance modeling
0.012004
Issues and Challenges in the Performance Analysis of Real Disk Arrays · IEEE Trans. Parallel Distributed Syst. 2004
Embedded and real-time systems › real-time scheduling › schedulability analysis
response time analysis
0.012001
Response Time Analysis of Parallel Computer and Storage Systems · IEEE Trans. Parallel Distributed Syst. 2001
Performance modeling and evaluation › queueing models
mean value analysis
0.011999
Mean Value Technique for Closed Fork-Join Networks · SIGMETRICS 1999
Performance modeling and evaluation › queueing models
queueing network analysis
0.011996
Analysis of Balanced Fork-Join Queueing Networks · SIGMETRICS 1996
Performance modeling and evaluation
measurement-based modeling
0.012004
Issues and Challenges in the Performance Analysis of Real Disk Arrays · IEEE Trans. Parallel Distributed Syst. 2004

Methods — techniques the papers use, named apart from their topics

task automation · 1.0provenance tracking · 1.0simulation · 0.1measurement data analysis · 0.0black-box modeling · 0.0response time approximation · 0.0mean value analysis · 0.0queueing analysis · 0.0
YearPublicationVenuePosition
2021 RepeatFS: a file system providing reproducibility through provenance and automation
abstract
MOTIVATION: Reproducibility is of central importance to the scientific process. The difficulty of consistently replicating and verifying experimental results is magnified in the era of big data, in which bioinformatics analysis often involves complex multi-application pipelines operating on terabytes of data. These processes result in thousands of possible permutations of data preparation steps, software versions and command-line arguments. Existing reproducibility frameworks are cumbersome and involve redesigning computational methods. To address these issues, we developed RepeatFS, a file system that records, replicates and verifies informatics workflows with no alteration to the original methods. RepeatFS also provides several other features to help promote analytical transparency and reproducibility, including provenance visualization and task automation. RESULTS: We used RepeatFS to successfully visualize and replicate a variety of bioinformatics tasks consisting of over a million operations with no alteration to the original methods. RepeatFS correctly identified all software inconsistencies that resulted in replication differences. AVAILABILITYAND IMPLEMENTATION: RepeatFS is implemented in Python 3. Its source code and documentation are available at https://github.com/ToniWestbrook/repeatfs. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Anthony Westbrook, Elizabeth Varki, William Kelley Thomas
Bioinform.2
2016 RAIDX: RAID without Striping
abstract
Each disk of traditional RAID is logically divided into stripe units, and stripe units at the same location on each disk form a stripe. Thus, RAID striping forces homogeneity on its disks. Now, consider a heterogeneous array of hard disks, solid state disks, and RAM storage devices, with various access speeds and sizes. If this array is organized as a RAID system, then larger disks have wasted space and faster disks are under utilized. This paper proposes RAIDX, a new organization for a heterogeneous array. RAIDX disks are divided into chunks, larger disks have more chunks. Chunks from one or more disks are grouped into bundles, and RAIDX bundles chunks of data across its disks. The heterogeneity of disks causes unbalanced load distribution with some under-utilized disks and some bottleneck disks. To balance load across disks, RAIDX moves most frequently accessed chunks to under-utilized, faster disks and least frequently used chunks to larger, slower disks. Chunk remapping is done at the RAIDX level and does not impact file system to storage addressing. Experiments comparing local and networked RAIDX against local RAID show that RAIDX has faster throughput than RAID when the array is composed of heterogeneous disks: local RAIDX is 2.5x faster than RAID, networked RAIDX is 1.6x faster than local RAID.
András Fekete, Elizabeth Varki
MASCOTS2
2010 Sequential Prefetch Cache Sizing for Maximal Hit Rate
abstract
We propose a prefetch cache sizing module for use with any sequential prefetching scheme and evaluate its impact on the hit rate. Disk array caches perform sequential prefetching by loading data contiguous to I/O request data into the array cache. If the I/O workload has sequential locality, then data prefetched in response to sequential accesses in the workload will receive hits. Different schemes prefetch different data, so the prefetch cache size requirement varies. Moreover, the proportion of sequential and random requests in the workload and their interleaving pattern affects the size requirement. If the cache is too small, then prefetched data would get evicted from the cache before a request for the data arrives, thus lowering the hit rate. If the cache is too large, then valuable cache space is wasted. We present a simple sizing module that can be added to any prefetching scheme to ensure that the prefetch cache size is adequately matched to the requirement of the prefetching scheme on a dynamic workload comprising multiple streams. We analytically compute the maximal hit rate achievable by popular prefetching schemes and through simulations, show that our sizing module maintains the prefetch cache at a size that nearly achieves this maximal hit rate.
Swapnil Bhatia, Elizabeth Varki, Arif Merchant
MASCOTS2
2008 TaP: Table-based Prefetching for Storage Caches
Mingju Li, Elizabeth Varki, Swapnil Bhatia, Arif Merchant
FAST2
2008 Co-allocation in Data Grids: A Global, Multi-user Perspective
Adam H. Villa, Elizabeth Varki
GPC2
2004 Issues and Challenges in the Performance Analysis of Real Disk Arrays
abstract
The performance modeling and analysis of disk arrays is challenging due to the presence of multiple disks, large array caches, and sophisticated array controllers. Moreover, storage manufacturers may not reveal the internal algorithms implemented in their devices, so real disk arrays are effectively black-boxes. We use standard performance techniques to develop an integrated performance model that incorporates some of the complexities of real disk arrays. We show how measurement data and baseline performance models can be used to extract information about the various features implemented in a disk array. In this process, we identify areas for future research in the performance analysis of real disk arrays.
Elizabeth Varki, Arif Merchant, Jianzhang Xu, Xiaozhou Qiu
IEEE Trans. Parallel Distributed Syst.1
2001 Response Time Analysis of Parallel Computer and Storage Systems
abstract
Fork-join structures have gained increased importance in recent years as a means of modeling parallelism in computer and storage systems. The basic fork-join model is one in which a job arriving at a parallel system splits into K independent tasks that are assigned to K unique, homogeneous servers. In the paper, a simple response time approximation is derived for parallel systems with exponential service time distributions. The approximation holds for networks modeling several devices, both parallel and nonparallel. (In the case of closed networks containing a stand-alone parallel system, a mean response time bound is derived.) In addition, the response time approximation is extended to cover the more realistic case wherein a job splits into an arbitrary number of tasks upon arrival at a parallel system. Simulation results for closed networks with stand-alone parallel subsystems and exponential service time distributions indicate that the response time approximation is, on average, within 3 percent of the seeded response times. Similarly, simulation results with nonexponential distributions also indicate that the response time approximation is close to the seeded values. Potential applications of our results include the modeling of data placement in disk arrays and the execution of parallel programs in multiprocessor and distributed systems.
Elizabeth Varki
IEEE Trans. Parallel Distributed Syst.1
1999 Mean Value Technique for Closed Fork-Join Networks
abstract
A simple technique for computing mean performance measures of closed single-class fork-join networks with exponential service time distribution is given here.This technique is similar to the mean value analysis technique for closed product-form networks and iterates on the number of customers in the network.Mean performance measures like the mean response times, queue lengths, and throughput of closed fork-join networks can be computed recursively without calculating the steady-state distribution of the network.The technique is based on the mean value equation for fork-join networks which relates the response time of a network to the mean service times at the service centers and the mean queue length of the system with one customer less.Unlike product-form networks, the mean value equation for fork-join networks is an approximation and the technique computes lower performance bound values for the fork-join network.However, it is a good approximation since the mean value equation is derived from an equation that exactly relates the response time of parallel systems to the degree of parallelism and the mean arrival queue length.Using simulation, it is shown that the relative error in the approximation is less than 5% in most cases.The error does not increase with each iteration.Permlsslon to make digital or hard copies of all or part of this work for personal or classroom use is granted without tee provided that copa are not made or distributed for pro10 01 commercial advantage and that copies bear this notice and the full cNa110n on the first Page.To copy otherwise, to republish.to post on servers Or t0 redwribute to lists.requires pnor specific
Elizabeth Varki
SIGMETRICS1
1996 Analysis of Balanced Fork-Join Queueing Networks
Elizabeth Varki, Lawrence W. Dowdy
SIGMETRICS1