Evan Randall Sparks

dblp:136/5765 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-authorDatabases, data management, data science and information retrieval · 4 · 2 first-authorSystems, architecture and hardware · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
High-performance computing · 48% Performance modeling and evaluation · 23% Distributed systems · 13%
Databases, data mining, and information retrieval
2 papers
Machine learning and data management · 67% Data integration and cleaning · 33%
Artificial intelligence
3 papers
Efficient and distributed learning · 100%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data integration and cleaning
data provenance
0.312017
Diagnosing Machine Learning Pipelines with Fine-grained Lineage · HPDC 2017
Machine learning and data management › scalable machine learning
distributed learning
0.312017
KeystoneML: Optimizing Pipelines for Large-Scale Advanced Analytics · ICDE 2017
Machine learning and data management › automated machine learning
machine learning pipeline optimization
0.312017
KeystoneML: Optimizing Pipelines for Large-Scale Advanced Analytics · ICDE 2017
Machine learning › Efficient and distributed learning
machine learning libraries
0.212016
MLlib: Machine Learning in Apache Spark · J. Mach. Learn. Res. 2016
High-performance computing › numerical linear algebra
distributed matrix computation
0.212016
Matrix Computations and Optimization in Apache Spark · KDD 2016
High-performance computing › numerical linear algebra
singular value decomposition
0.212016
Matrix Computations and Optimization in Apache Spark · KDD 2016
Mathematical optimization › continuous optimization
convex optimization
0.212016
Matrix Computations and Optimization in Apache Spark · KDD 2016
Distributed systems
distributed machine learning
0.112017
Diagnosing Machine Learning Pipelines with Fine-grained Lineage · HPDC 2017
High-performance computing › data-intensive computing
large-scale data analytics
0.112017
KeystoneML: Optimizing Pipelines for Large-Scale Advanced Analytics · ICDE 2017
Cloud and datacenter computing
cluster computing framework
0.112016
Matrix Computations and Optimization in Apache Spark · KDD 2016
Cloud and datacenter computing
cluster resource management and scheduling
0.112016
MLlib: Machine Learning in Apache Spark · J. Mach. Learn. Res. 2016
Distributed systems
distributed data processing
0.112016
MLlib: Machine Learning in Apache Spark · J. Mach. Learn. Res. 2016
Memory systems
data-centric computing
0.012013
MLI: An API for Distributed Machine Learning · ICDM 2013

Methods — techniques the papers use, named apart from their topics

pipeline optimization · 0.6performance modeling · 0.6linear algebra primitives · 0.5distributed optimization · 0.5distributed matrix operations · 0.5convex optimization · 0.5application programming interface design · 0.3
YearPublicationVenuePosition
2020 Kira: Processing Astronomy Imagery Using Big Data Technology
abstract
Scientific analyses commonly compose multiple single-process programs into a dataflow. An end-to-end dataflow of single-process programs is known as a many-task application. Typically, HPC tools are used to parallelize these analyses. In this work, we investigate an alternate approach that uses Apache Spark-a modern platform for data intensive computing-to parallelize many-task applications. We implement Kira, a flexible and distributed astronomy image processing toolkit, and its Source Extractor (Kira SE) application. Using Kira SE as a case study, we examine the programming flexibility, dataflow richness, scheduling capacity and performance of Apache Spark running on the Amazon EC2 cloud. By exploiting data locality, Kira SE achieves a 4.1× speedup over an equivalent C program when analyzing a 1TB dataset using 512 cores on the Amazon EC2 cloud. Furthermore, Kira SE on the Amazon EC2 cloud achieves a 1.8× speedup over the C program on the NERSC Edison supercomputer. A 128-core Amazon EC2 cloud deployment of Kira SE using Spark Streaming can achieve a second-scale latency with a sustained throughput of 800 MB/s. Our experience with Kira demonstrates that data intensive computing platforms like Apache Spark are a performant alternative for many-task scientific applications.
Zhao Zhang 0007, Kyle Barbary, Frank A. Nothaft, Evan Randall Sparks, Oliver Zahn, Michael J. Franklin, David A. Patterson 0001, Saul Perlmutter
IEEE Trans. Big Data4
2017 Random projection design for scalable implicit smoothing of randomly observed stochastic processes
abstract
Sampling at random timestamps, long range dependencies, and scale hamper standard meth- ods for multivariate time series analysis. In this paper we present a novel estimator for cross-covariance of randomly observed time series which unravels the dynamics of an unobserved stochastic process. We analyze the statistical properties of our estimator without needing the assumption that observation timestamps are independent from the process of interest and show that our solution is not hindered by the issues affecting standard estimators for cross-covariance. We implement and evaluate our statistically sound and scalable approach in the distributed setting using Apache Spark and demonstrate its ability to unravel causal dynamics on both simulations and high-frequency financial trading data.
Francois Belletti, Evan Randall Sparks, Alexandre M. Bayen, Joseph Gonzalez 0001
AISTATS2
2017 Diagnosing Machine Learning Pipelines with Fine-grained Lineage
abstract
We present the Hippo system to enable the diagnosis of distributed machine learning (ML) pipelines by leveraging fine-grained data lineage. Hippo exposes a concise yet powerful API, derived from primitive lineage types, to capture fine-grained data lineage for each data transformation. It records the input datasets, the output datasets and the cell-level mapping between them. It also collects sufficient information that is needed to reproduce the computation. Hippo efficiently enables common ML diagnosis operations such as code debugging, result analysis, data anomaly removal, and computation replay. By exploiting the metadata separation and high-order function encoding strategies, we observe an O(10^3)x total improvement in lineage storage efficiency vs. the baseline of cell-wise mapping recording while maintaining the lineage integrity. Hippo can answer the real use case lineage queries within a few seconds, which is low enough to enable interactive diagnosis of ML pipelines.
Zhao Zhang 0007, Evan Randall Sparks, Michael J. Franklin
HPDC2
2017 KeystoneML: Optimizing Pipelines for Large-Scale Advanced Analytics
abstract
Modern advanced analytics applications make use of machine learning techniques and contain multiple steps of domain-specific and general-purpose processing with high resource requirements. We present KeystoneML, a system that captures and optimizes the end-to-end large-scale machine learning applications for high-throughput training in a distributed environment with a high-level API. This approach offers increased ease of use and higher performance over existing systems for large scale learning. We demonstrate the effectiveness of KeystoneML in achieving high quality statistical accuracy and scalable training using real world datasets in several domains.
Evan Randall Sparks, Shivaram Venkataraman, Tomer Kaftan, Michael J. Franklin, Benjamin Recht
ICDE1
2017 Paleo: A Performance Model for Deep Neural Networks
Evan Randall Sparks, Ameet Talwalkar
ICLR (Poster)2
2016 Matrix Computations and Optimization in Apache Spark
abstract
We describe matrix computations available in the cluster programming framework, Apache Spark. Out of the box, Spark provides abstractions and implementations for distributed matrices and optimization routines using these matrices. When translating single-node algorithms to run on a distributed cluster, we observe that often a simple idea is enough: separating matrix operations from vector operations and shipping the matrix operations to be ran on the cluster, while keeping vector operations local to the driver. In the case of the Singular Value Decomposition, by taking this idea to an extreme, we are able to exploit the computational power of a cluster, while running code written decades ago for a single core. Another example is our Spark port of the popular TFOCS optimization package, originally built for MATLAB, which allows for solving Linear programs as well as a variety of other convex programs. We conclude with a comprehensive set of benchmarks for hardware accelerated matrix computations from the JVM, which is interesting in its own right, as many cluster programming frameworks use the JVM. The contributions described in this paper are already merged into Apache Spark and available on Spark installations by default, and commercially supported by a slew of companies which provide further services.
Reza Bosagh Zadeh, Alexander Ulanov, Burak Yavuz, Li Pu, Shivaram Venkataraman, Evan Randall Sparks, Aaron Staple, Matei Zaharia
KDD7
2016 MLlib: Machine Learning in Apache Spark
abstract
Apache Spark is a popular open-source platform for large-scale data processing that is well-suited for iterative machine learning tasks. In this paper we present MLlib, Spark's open- source distributed machine learning library. MLlib provides efficient functionality for a wide range of learning settings and includes several underlying statistical, optimization, and linear algebra primitives. Shipped with Spark, MLlib supports several languages and provides a high-level API that leverages Spark's rich ecosystem to simplify the development of end-to-end machine learning pipelines. MLlib has experienced a rapid growth due to its vibrant open-source community of over 140 contributors, and includes extensive documentation to support further growth and to let users quickly get up to speed.
Joseph K. Bradley, Burak Yavuz, Evan Randall Sparks, Shivaram Venkataraman, Davies Liu, Jeremy Freeman, D. B. Tsai, Manish Amde, Sean Owen, Doris Xin, Reynold Xin, Michael J. Franklin, Reza Bosagh Zadeh, Matei Zaharia, Ameet Talwalkar
J. Mach. Learn. Res.4
2015 Scientific computing meets big data technology: An astronomy use case
abstract
Scientific analyses commonly compose multiple single-process programs into a dataflow. An end-to-end dataflow of single-process programs is known as a many-task application. Typically, tools from the HPC software stack are used to parallelize these analyses. In this work, we investigate an alternate approach that uses Apache Spark - a modern big data platform - to parallelize many-task applications. We present Kira, a flexible and distributed astronomy image processing toolkit using Apache Spark. We then use the Kira toolkit to implement a Source Extractor application for astronomy images, called Kira SE. With Kira SE as the use case, we study the programming flexibility, dataflow richness, scheduling capacity and performance of Apache Spark running on the EC2 cloud. By exploiting data locality, Kira SE achieves a 3.7 χ speedup over an equivalent C program when analyzing a 1TB dataset using 512 cores on the Amazon EC2 cloud. Furthermore, we show that by leveraging software originally designed for big data infrastructure, Kira SE achieves competitive performance to the C implementation running on the NERSC Edison supercomputer. Our experience with Kira indicates that emerging Big Data platforms such as Apache Spark are a performant alternative for many-task scientific applications.
Zhao Zhang 0007, Kyle Barbary, Frank A. Nothaft, Evan Randall Sparks, Oliver Zahn, Michael J. Franklin, David A. Patterson 0001, Saul Perlmutter
IEEE BigData4
2015 Automating model search for large scale machine learning
abstract
The proliferation of massive datasets combined with the development of sophisticated analytical techniques has enabled a wide variety of novel applications such as improved product recommendations, automatic image tagging, and improved speech-driven interfaces. A major obstacle to supporting these predictive applications is the challenging and expensive process of identifying and training an appropriate predictive model. Recent efforts aiming to automate this process have focused on single node implementations and have assumed that model training itself is a black box, limiting their usefulness for applications driven by large-scale datasets. In this work, we build upon these recent efforts and propose an architecture for automatic machine learning at scale comprised of a cost-based cluster resource allocation estimator, advanced hyper-parameter tuning techniques, bandit resource allocation via runtime algorithm introspection, and physical optimization via batching and optimal resource allocation. The result is TuPAQ, a component of the MLbase system that automatically finds and trains models for a user's predictive application with comparable quality to those found using exhaustive strategies, but an order of magnitude more efficiently than the standard baseline approach. TuPAQ scales to models trained on Terabytes of data across hundreds of machines.
Evan Randall Sparks, Ameet Talwalkar, Daniel Haas, Michael J. Franklin, Michael I. Jordan, Tim Kraska
SoCC1
2013 MLI: An API for Distributed Machine Learning
abstract
MLI is an Application Programming Interface designed to address the challenges of building Machine Learning algorithms in a distributed setting based on data-centric computing. Its primary goal is to simplify the development of high-performance, scalable, distributed algorithms. Our initial results show that, relative to existing systems, this interface can be used to build distributed implementations of a wide variety of common Machine Learning algorithms with minimal complexity and highly competitive performance and scalability.
Evan Randall Sparks, Ameet Talwalkar, Virginia Smith, Jey Kottalam, Xinghao Pan, Joseph Gonzalez 0001, Michael J. Franklin, Michael I. Jordan, Tim Kraska
ICDM1