John Robert Wernsing

dblp:08/7942 · also John Wernsing · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
0since 2021 · last 2017
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-authorArtificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Network and information security
2 papers
Cryptographic primitives and cryptanalysis · 48% Cryptographic protocols and secure computation · 30% Privacy and data protection · 22%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
High-performance computing · 32% Electronic design automation · 32% Performance modeling and evaluation · 17%
Databases, data mining, and information retrieval
1 paper
Query processing and optimization · 33% Data stream processing · 33% Database system architecture and tuning · 33%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 16 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cryptographic primitives and cryptanalysis
homomorphic encryption
0.522017
Manual for Using Homomorphic Encryption for Bioinformatics · Proc. IEEE 2017
CryptoNets: Applying Neural Networks to Encrypted Data with High Throughput and Accuracy · ICML 2016
Privacy and data protection
privacy-preserving machine learning
0.212016
CryptoNets: Applying Neural Networks to Encrypted Data with High Throughput and Accuracy · ICML 2016
Cryptographic protocols and secure computation › secure inference
secure neural network inference
0.212016
CryptoNets: Applying Neural Networks to Encrypted Data with High Throughput and Accuracy · ICML 2016
Data stream processing
continuous query processing
0.212014
Trill: A High-Performance Incremental Query Processor for Diverse Analytics · Proc. VLDB Endow. 2014
Query processing and optimization › incremental computation
incremental query processing
0.212014
Trill: A High-Performance Incremental Query Processor for Diverse Analytics · Proc. VLDB Endow. 2014
Performance modeling and evaluation › benchmarking › computer architecture benchmarking
accelerator benchmarking
0.212013
A performance and energy comparison of convolution on GPUs, FPGAs, and multicore processors · ACM Trans. Archit. Code Optim. 2013
High-performance computing › tensor computation
convolution
0.212013
A performance and energy comparison of convolution on GPUs, FPGAs, and multicore processors · ACM Trans. Archit. Code Optim. 2013
Electronic design automation
design space exploration
0.212013
A performance and energy comparison of convolution on GPUs, FPGAs, and multicore processors · ACM Trans. Archit. Code Optim. 2013
Electronic design automation › design space exploration
automated design space exploration
0.112012
RACECAR: a heuristic for automatic function specialization on multi-core heterogeneous systems · PPoPP 2012
High-performance computing
performance optimization
0.112012
RACECAR: a heuristic for automatic function specialization on multi-core heterogeneous systems · PPoPP 2012
Reconfigurable computing and FPGAs
FPGA accelerator
0.122013
A performance and energy comparison of convolution on GPUs, FPGAs, and multicore processors · ACM Trans. Archit. Code Optim. 2013
RACECAR: a heuristic for automatic function specialization on multi-core heterogeneous systems · PPoPP 2012
Bioinformatics and computational biology
genomic privacy
0.112017
Manual for Using Homomorphic Encryption for Bioinformatics · Proc. IEEE 2017
Cryptographic protocols and secure computation
secure computation on encrypted data
0.112017
Manual for Using Homomorphic Encryption for Bioinformatics · Proc. IEEE 2017
Machine learning › Deep learning architectures and training
neural network inference
0.112016
CryptoNets: Applying Neural Networks to Encrypted Data with High Throughput and Accuracy · ICML 2016
Processor architecture and microarchitecture
chip multiprocessor
0.012013
A performance and energy comparison of convolution on GPUs, FPGAs, and multicore processors · ACM Trans. Archit. Code Optim. 2013
Parallel and multicore computing
task partitioning
0.012012
RACECAR: a heuristic for automatic function specialization on multi-core heterogeneous systems · PPoPP 2012

Methods — techniques the papers use, named apart from their topics

homomorphic encryption · 1.1neural network conversion · 0.5encrypted inference · 0.5dynamic compilation · 0.2batched-columnar data representation · 0.2pareto-optimal trade-off analysis · 0.2heuristic search · 0.1
YearPublicationVenuePosition
2017 Manual for Using Homomorphic Encryption for Bioinformatics
abstract
Biological data science is an emerging field facing multiple challenges for hosting, sharing, computing on, and interacting with large data sets. Privacy regulations and concerns about the risks of leaking sensitive personal health and genomic data add another layer of complexity to the problem. Recent advances in cryptography over the last five years have yielded a tool, homomorphic encryption, which can be used to encrypt data in such a way that storage can be outsourced to an untrusted cloud, and the data can be computed on in a meaningful way in encrypted form, without access to decryption keys. This paper introduces homomorphic encryption to the bioinformatics community, and presents an informal “manual” for using the Simple Encrypted Arithmetic Library (SEAL), which we have made publicly available for bioinformatic, genomic, and other research purposes.
Nathan Dowlin, Ran Gilad-Bachrach, Kim Laine, Kristin E. Lauter, Michael Naehrig, John Robert Wernsing
Proc. IEEE6
2016 CryptoNets: Applying Neural Networks to Encrypted Data with High Throughput and Accuracy
abstract
Applying machine learning to a problem which involves medical, financial, or other types of sensitive data, not only requires accurate predictions but also careful attention to maintaining data privacy and security. Legal and ethical requirements may prevent the use of cloud-based machine learning solutions for such tasks. In this work, we will present a method to convert learned neural networks to CryptoNets, neural networks that can be applied to encrypted data. This allows a data owner to send their data in an encrypted form to a cloud service that hosts the network. The encryption ensures that the data remains confidential since the cloud does not have access to the keys needed to decrypt it. Nevertheless, we will show that the cloud service is capable of applying the neural network to the encrypted data to make encrypted predictions, and also return them in encrypted form. These encrypted predictions can be sent back to the owner of the secret key who can decrypt them. Therefore, the cloud service does not gain any information about the raw data nor about the prediction it made. We demonstrate CryptoNets on the MNIST optical character recognition tasks. CryptoNets achieve 99% accuracy and can make around 59000 predictions per hour on a single PC. Therefore, they allow high throughput, accurate, and private predictions.
Ran Gilad-Bachrach, Nathan Dowlin, Kim Laine, Kristin E. Lauter, Michael Naehrig, John Robert Wernsing
ICML6
2015 Tempe: Live scripting for live data
abstract
Data scientists are increasingly working with live streaming data, for example, business telemetry and signals from wearable devices and the Internet of Things. Unfortunately, current tools for exploratory data analysis provide poor support for streaming data. This paper presents Tempe, a data science environment for temporal and streaming data. Tempe's extensible scripting environment allows for live programming, displays interactive, continually updating visualizations, and provides a uniform query language for both stored and live data. We discuss the streaming features of Tempe and evaluate our design choices with a deployment study at Microsoft with a product team who used Tempe continuously for six months.
Robert DeLine, Danyel Fisher, Badrish Chandramouli, Jonathan Goldstein, Michael Barnett 0001, James F. Terwilliger, John Robert Wernsing
VL/HCC7
2014 Trill: A High-Performance Incremental Query Processor for Diverse Analytics
abstract
This paper introduces Trill -- a new query processor for analytics. Trill fulfills a combination of three requirements for a query processor to serve the diverse big data analytics space: (1) Query Model : Trill is based on a tempo-relational model that enables it to handle streaming and relational queries with early results, across the latency spectrum from real-time to offline; (2) Fabric and Language Integration : Trill is architected as a high-level language library that supports rich data-types and user libraries, and integrates well with existing distribution fabrics and applications; and (3) Performance : Trill's throughput is high across the latency spectrum. For streaming data, Trill's throughput is 2-4 orders of magnitude higher than comparable streaming engines. For offline relational queries, Trill's throughput is comparable to a major modern commercial columnar DBMS. Trill uses a streaming batched-columnar data representation with a new dynamic compilation-based system architecture that addresses all these requirements. In this paper, we describe Trill's new design and architecture, and report experimental results that demonstrate Trill's high performance across diverse analytics scenarios. We also describe how Trill's ability to support diverse analytics has resulted in its adoption across many usage scenarios at Microsoft.
Badrish Chandramouli, Jonathan Goldstein, Michael Barnett 0001, Robert DeLine, John C. Platt, James F. Terwilliger, John Robert Wernsing
Proc. VLDB Endow.7
2013 A performance and energy comparison of convolution on GPUs, FPGAs, and multicore processors
abstract
Recent architectural trends have focused on increased parallelism via multicore processors and increased heterogeneity via accelerator devices (e.g., graphics-processing units, field-programmable gate arrays). Although these architectures have significant performance and energy potential, application designers face many device-specific challenges when choosing an appropriate accelerator or when customizing an algorithm for an accelerator. To help address this problem, in this article we thoroughly evaluate convolution, one of the most common operations in digital-signal processing, on multicores, graphics-processing units, and field-programmable gate arrays. Whereas many previous application studies evaluate a specific usage of an application, this article assists designers with design space exploration for numerous use cases by analyzing effects of different input sizes, different algorithms, and different devices, while also determining Pareto-optimal trade-offs between performance and energy.
Jeremy Fowers, Greg Brown, John Robert Wernsing, Greg Stitt
ACM Trans. Archit. Code Optim.3
2012 The RACECAR heuristic for automatic function specialization on multi-core heterogeneous systems
abstract
Embedded systems increasingly combine multi-core processors and heterogeneous resources such as graphics-processing units and field-programmable gate arrays. However, significant application design complexity for such systems caused by parallel programming and device-specific challenges has often led to untapped performance potential. Application developers targeting such systems currently must determine how to parallelize computation, create different device-specialized implementations for each heterogeneous resource, and then determine how to apportion work to each resource. In this paper, we present the RACECAR heuristic to automate the optimization of applications for multi-core heterogeneous systems by automatically exploring implementation alternatives that include different algorithms, parallelization strategies, and work distributions. Experimental results show RACECAR-specialized implementations can effectively incorporate provided implementations and parallelize computation across multiple cores, graphics-processing units, and field-programmable gate arrays, improving performance by an average of 47x compared to a CPU, while the fastest provided implementations are only able to average 33x.
John Robert Wernsing, Greg Stitt, Jeremy Fowers
CASES1
2012 RACECAR: a heuristic for automatic function specialization on multi-core heterogeneous systems
abstract
High-performance computing systems increasingly combine multi-core processors and heterogeneous resources such as graphics-processing units and field-programmable gate arrays. However, significant application design complexity for such systems has often led to untapped performance potential. Application designers targeting such systems currently must determine how to parallelize computation, create device-specialized implementations for each heterogeneous resource, and determine how to partition work for each resource. In this paper, we present the RACECAR heuristic to automate the optimization of applications for multi-core heterogeneous systems by automatically exploring implementation alternatives that include different algorithms, parallelization strategies, and work distributions. Experimental results show RACECAR-specialized implementations achieve speedups up to 117x and average 11x compared to a single CPU thread when parallelizing computation across multiple cores, graphics-processing units, and field-programmable gate arrays.
John Robert Wernsing, Greg Stitt
PPoPP1
2012 Elastic computing: A portable optimization framework for hybrid computers
John Robert Wernsing, Greg Stitt
Parallel Comput.1
2010 Elastic computing: a framework for transparent, portable, and adaptive multi-core heterogeneous computing
abstract
Over the past decade, system architectures have started on a clear trend towards increased parallelism and heterogeneity, often resulting in speedups of 10x to 100x. Despite numerous compiler and high-level synthesis studies, usage of such systems has largely been limited to device experts, due to significantly increased application design complexity. To reduce application design complexity, we introduce elastic computing - a framework that separates functionality from implementation details by enabling designers to use specialized functions, called elastic functions, which enable an optimization framework to explore thousands of possible implementations, even ones using different algorithms. Elastic functions allow designers to execute the same application code efficiently on potentially any architecture and for different runtime parameters such as input size, battery life, etc. In this paper, we present an initial elastic computing framework that transparently optimizes application code onto diverse systems, achieving significant speedups ranging from 1.3x to 46x on a hyper-threaded Xeon system with an FPGA accelerator, a 16-CPU Opteron system, and a quad-core Xeon system.
John Robert Wernsing, Greg Stitt
LCTES1