Sara Alspaugh

dblp:53/8397 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Performance modeling and evaluation · 48% Cloud and datacenter computing · 36% Energy-efficient computing · 16%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 100%
Human-computer interaction and pervasive computing
1 paper
Usability and user experience research · 100%

Topics — the 8 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visualization and visual analytics › visualization evaluation
user study
0.412019
Futzing and Moseying: Interviews with Professional Data Analysts on Exploration Practices · IEEE Trans. Vis. Comput. Graph. 2019
Cloud and datacenter computing
cluster resource management and scheduling
0.222012
Energy efficiency for large-scale MapReduce workloads with significant interactive analysis · EuroSys 2012
Interactive Analytical Processing in Big Data Systems: A Cross-Industry Study of MapReduce Workloads · Proc. VLDB Endow. 2012
Performance modeling and evaluation
benchmarking
0.112012
Interactive Analytical Processing in Big Data Systems: A Cross-Industry Study of MapReduce Workloads · Proc. VLDB Endow. 2012
Energy-efficient computing
datacenter energy efficiency
0.112012
Energy efficiency for large-scale MapReduce workloads with significant interactive analysis · EuroSys 2012
Performance modeling and evaluation › workload characterization › parallel workload analysis
mapreduce workload characterization
0.112012
Interactive Analytical Processing in Big Data Systems: A Cross-Industry Study of MapReduce Workloads · Proc. VLDB Endow. 2012
Cloud and datacenter computing › job scheduling
workload-aware scheduling
0.112012
Energy efficiency for large-scale MapReduce workloads with significant interactive analysis · EuroSys 2012
Performance modeling and evaluation
workload characterization
0.112012
Interactive Analytical Processing in Big Data Systems: A Cross-Industry Study of MapReduce Workloads · Proc. VLDB Endow. 2012
Query processing and optimization
interactive data exploration
0.012012
Energy efficiency for large-scale MapReduce workloads with significant interactive analysis · EuroSys 2012

Methods — techniques the papers use, named apart from their topics

interview study · 0.8trace analysis · 0.4workload characterization · 0.3empirical study · 0.1
YearPublicationVenuePosition
2019 Futzing and Moseying: Interviews with Professional Data Analysts on Exploration Practices
abstract
We report the results of interviewing thirty professional data analysts working in a range of industrial, academic, and regulatory environments. This study focuses on participants' descriptions of exploratory activities and tool usage in these activities. Highlights of the findings include: distinctions between exploration as a precursor to more directed analysis versus truly open-ended exploration; confirmation that some analysts see "finding something interesting" as a valid goal of data exploration while others explicitly disavow this goal; conflicting views about the role of intelligent tools in data exploration; and pervasive use of visualization for exploration, but with only a subset using direct manipulation interfaces. These findings provide guidelines for future tool development, as well as a better understanding of the meaning of the term "data exploration" based on the words of practitioners "in the wild."
Sara Alspaugh, Nava Zokaei, Andrea Liu, Cindy Jin, Marti A. Hearst
IEEE Trans. Vis. Comput. Graph.1
2014 Analyzing Log Analysis: An Empirical Study of User Log Mining
Sara Alspaugh, Bei Di Chen, Jessica Lin 0003, Archana Ganapathi, Marti A. Hearst, Randy H. Katz
LISA1
2012 Cake: enabling high-level SLOs on shared storage systems
abstract
Cake is a coordinated, multi-resource scheduler for shared distributed storage environments with the goal of achieving both high throughput and bounded latency. Cake uses a two-level scheduling scheme to enforce high-level service-level objectives (SLOs). First-level schedulers control consumption of resources such as disk and CPU. These schedulers (1) provide mechanisms for differentiated scheduling, (2) split large requests into smaller chunks, and (3) limit the number of outstanding device requests, which together allow for effective control over multi-resource consumption within the storage system. Cake's second-level scheduler coordinates the first-level schedulers to map high-level SLO requirements into actual scheduling parameters. These parameters are dynamically adjusted over time to enforce high-level performance specifications for changing workloads. We evaluate Cake using multiple workloads derived from real-world traces. Our results show that Cake allows application programmers to explore the latency vs. throughput trade-off by setting different high-level performance requirements on their workloads. Furthermore, we show that using Cake has concrete economic and business advantages, reducing provisioning costs by up to 50% for a consolidated workload and reducing the completion time of an analytics cycle by up to 40%.
Andrew Wang 0002, Shivaram Venkataraman, Sara Alspaugh, Randy H. Katz, Ion Stoica
SoCC3
2012 Energy efficiency for large-scale MapReduce workloads with significant interactive analysis
abstract
MapReduce workloads have evolved to include increasing amounts of time-sensitive, interactive data analysis; we refer to such workloads as MapReduce with Interactive Analysis (MIA). Such workloads run on large clusters, whose size and cost make energy efficiency a critical concern. Prior works on MapReduce energy efficiency have not yet considered this workload class. Increasing hardware utilization helps improve efficiency, but is challenging to achieve for MIA workloads. These concerns lead us to develop BEEMR (Berkeley Energy Efficient MapReduce), an energy efficient MapReduce workload manager motivated by empirical analysis of real-life MIA traces at Facebook. The key insight is that although MIA clusters host huge data volumes, the interactive jobs operate on a small fraction of the data, and thus can be served by a small pool of dedicated machines; the less time-sensitive jobs can run on the rest of the cluster in a batch fashion. BEEMR achieves 40-50% energy savings under tight design constraints, and represents a first step towards improving energy efficiency for an increasingly important class of datacenter workloads.
Yanpei Chen, Sara Alspaugh, Dhruba Borthakur, Randy H. Katz
EuroSys2
2012 Interactive Analytical Processing in Big Data Systems: A Cross-Industry Study of MapReduce Workloads
abstract
Within the past few years, organizations in diverse industries have adopted MapReduce-based systems for large-scale data processing. Along with these new users, important new workloads have emerged which feature many small, short, and increasingly interactive jobs in addition to the large, long-running batch jobs for which MapReduce was originally designed. As interactive, large-scale query processing is a strength of the RDBMS community, it is important that lessons from that field be carried over and applied where possible in this new domain. However, these new workloads have not yet been described in the literature. We fill this gap with an empirical analysis of MapReduce traces from six separate business-critical deployments inside Facebook and at Cloudera customers in e-commerce, telecommunications, media, and retail. Our key contribution is a characterization of new MapReduce workloads which are driven in part by interactive analysis, and which make heavy use of query-like programming frameworks on top of MapReduce. These workloads display diverse behaviors which invalidate prior assumptions about MapReduce such as uniform data access, regular diurnal patterns, and prevalence of large jobs. A secondary contribution is a first step towards creating a TPC-like data processing benchmark for MapReduce.
Yanpei Chen, Sara Alspaugh, Randy H. Katz
Proc. VLDB Endow.2