Shoumik Palkar

dblp:168/9023 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
2since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 6 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 first-authorSystems, architecture and hardware · 1Computer networks · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Query processing and optimization · 58% Machine learning and data management · 26% Database system architecture and tuning · 10%
Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 63% Runtime systems and virtual machines · 37%
Computer architecture, parallel and distributed computing, and storage systems
4 papers
Parallel and multicore computing · 57% Cloud and datacenter computing · 24% Storage systems · 19%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%
Computer networks
1 paper
Software-defined and programmable networks · 77% Internet architecture and protocols · 23%

Topics — the 9 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization › memory optimization
data movement optimization
0.412019
Optimizing data-intensive computations in existing libraries with split annotations · SOSP 2019
Compilers and program optimization › loop transformation
loop fusion
0.412019
Optimizing data-intensive computations in existing libraries with split annotations · SOSP 2019
Parallel and multicore computing
parallel programming runtimes
0.412019
Optimizing data-intensive computations in existing libraries with split annotations · SOSP 2019
Query processing and optimization
query optimization
0.312018
Filter Before You Parse: Faster Analytics on Raw Data with Sparser · Proc. VLDB Endow. 2018
Software-defined and programmable networks
network function virtualization
0.212015
E2: a framework for NFV applications · SOSP 2015
Cloud and datacenter computing
cluster resource management and scheduling
0.212015
E2: a framework for NFV applications · SOSP 2015
Storage systems › data management › database storage
columnar storage
0.212022
Photon: A Fast Query Engine for Lakehouse Systems · SIGMOD Conference 2022
Parallel and multicore computing
parallel programming models
0.112020
Offload Annotations: Bringing Heterogeneous Computing to Existing Libraries and Workloads · USENIX ATC 2020
Internet architecture and protocols
packet processing
0.112015
E2: a framework for NFV applications · SOSP 2015

Methods — techniques the papers use, named apart from their topics

feature cascading · 0.9cost model · 0.9split annotations · 0.8runtime measurement · 0.7pipelining · 0.7adaptive optimization · 0.7selectivity estimation · 0.3dynamic optimization · 0.3SIMD-based filtering · 0.3
YearPublicationVenuePosition
2022 Photon: A High-Performance Query Engine for the Lakehouse
Alexander Behm, Shoumik Palkar
CIDR2
2022 Photon: A Fast Query Engine for Lakehouse Systems
abstract
Many organizations are shifting to a data management paradigm called the "Lakehouse," which implements the functionality of structured data warehouses on top of unstructured data lakes. This presents new challenges for query execution engines. The engine needs to provide good performance on the raw uncurated datasets that are ubiquitous in data lakes, and excellent performance on structured data stored in popular columnar file formats like Apache Parquet. Toward these goals, we present Photon, a vectorized query engine for Lakehouse environments that we developed at Databricks. Photon can outperform existing warehouses on SQL workloads and also supports the Apache Spark API. We discuss the design choices we made in Photon (e.g., vectorization vs. code generation) and describe its integration with our existing SQL and Apache Spark runtimes, its task model, and its memory manager. Photon has accelerated some customer workloads by over 10x and has recently allowed Databricks to set a new audited performance record for the official 100TB TPC-DS benchmark.
Alexander Behm, Shoumik Palkar, Utkarsh Agarwal, Timothy Armstrong, David Cashman, Ankur Dave, Todd Greenstein, Shant Hovsepian, Arvind Sai Krishnan, Paul Leventis, Ala Luszczak, Prashanth Menon, Mostafa Mokhtar, Gene Pang, Sameer Paranjpye, Greg Rahn, Bart Samwel, Tom van Bussel, Herman Van Hövell, Maryann Xue, Reynold Xin, Matei Zaharia
SIGMOD Conference2
2020 Offload Annotations: Bringing Heterogeneous Computing to Existing Libraries and Workloads
Gina Yuan, Shoumik Palkar, Deepak Narayanan, Matei Zaharia
USENIX ATC2
2020 A Demonstration of Willump: A Statistically-Aware End-to-end Optimizer for Machine Learning Inference
abstract
Systems for ML inference are widely deployed today, but they typically optimize ML inference workloads using techniques designed for conventional data serving workloads and miss critical opportunities to leverage the statistical nature of ML. In this demo, we present Willump, an optimizer for ML inference that introduces statistically-motivated optimizations targeting ML applications whose performance bottleneck is feature computation. Willump automatically cascades feature computation for classification queries: Willump classifies most data inputs using only high-value, low-cost features selected by a cost model, improving query performance by up to 5 x without statistically significant accuracy loss. In this demo, we use interactive and easily-downloadable Jupyter notebooks to show VLDB attendees which applications Willump can speed up, how to use Willump, and how Willump produces such large performance gains.
Peter Kraft, Daniel Kang 0001, Deepak Narayanan, Shoumik Palkar, Peter Bailis, Matei Zaharia
Proc. VLDB Endow.4
2019 Optimizing data-intensive computations in existing libraries with split annotations
abstract
Data movement between main memory and the CPU is a major bottleneck in parallel data-intensive applications. In response, researchers have proposed using compilers and intermediate representations (IRs) that apply optimizations such as loop fusion under existing high-level APIs such as NumPy and TensorFlow. Even though these techniques generally do not require changes to user applications, they require intrusive changes to the library itself: often, library developers must rewrite each function using a new IR. In this paper, we propose a new technique called split annotations (SAs) that enables key data movement optimizations over unmodified library functions. SAs only require developers to annotate functions and implement an API that specifies how to partition data in the library. The annotation and API describe how to enable cross-function data pipelining and parallelization, while respecting each function's correctness constraints. We implement a parallel runtime for SAs in a system called Mozart. We show that Mozart can accelerate workloads in libraries such as Intel MKL and Pandas by up to 15x, with no library modifications. Mozart also provides performance gains competitive with solutions that require rewriting libraries, and can sometimes outperform these systems by up to 2x by leveraging existing hand-optimized code.
Shoumik Palkar, Matei Zaharia
SOSP1
2018 Filter Before You Parse: Faster Analytics on Raw Data with Sparser
abstract
Exploratory big data applications often run on raw unstructured or semi-structured data formats, such as JSON files or text logs. These applications can spend 80--90% of their execution time parsing the data. In this paper, we propose a new approach for reducing this overhead: apply filters on the data's raw bytestream before parsing. This technique, which we call raw filtering, leverages the features of modern hardware and the high selectivity of queries found in many exploratory applications. With raw filtering, a user-specified query predicate is compiled into a set of filtering primitives called raw filters (RFs). RFs are fast, SIMD-based operators that occasionally yield false positives, but never false negatives. We combine multiple RFs into an RF cascade to decrease the false positive rate and maximize parsing throughput. Because the best RF cascade is data-dependent, we propose an optimizer that dynamically selects the combination of RFs with the best expected throughput, achieving within 10% of the global optimum cascade while adding less than 1.2% overhead. We implement these techniques in a system called Sparser, which automatically manages a parsing cascade given a data stream in a supported format (e.g., JSON, Avro, Parquet) and a user query. We show that many real-world applications are highly selective and benefit from Sparser. Across diverse workloads, Sparser accelerates state-of-the-art parsers such as Mison by up to 22 × and improves end-to-end application performance by up to 9 ×.
Shoumik Palkar, Firas Abuzaid, Peter Bailis, Matei Zaharia
Proc. VLDB Endow.1
2018 Evaluating End-to-End Optimization for Data Analytics Applications in Weld
abstract
Modern analytics applications use a diverse mix of libraries and functions. Unfortunately, there is no optimization across these libraries, resulting in performance penalties as high as an order of magnitude in many applications. To address this problem, we proposed Weld, a common runtime for existing data analytics libraries that performs key physical optimizations such as pipelining under existing, imperative library APIs. In this work, we further develop the Weld vision by designing an automatic adaptive optimizer for Weld applications, and evaluating its impact on realistic data science workloads. Our optimizer eliminates multiple forms of overhead that arise when composing imperative libraries like Pandas and NumPy, and uses lightweight measurements to make data-dependent decisions at run-time in ad-hoc workloads where no statistics are available, with sub-second overhead. We also evaluate which optimizations have the largest impact in practice and whether Weld can be integrated into libraries incrementally. Our results are promising: using our optimizer, Weld accelerates data science workloads by up to 23X on one thread and 80X on eight threads, and its adaptive optimizations provide up to a 3.75X speedup over rule-based optimization. Moreover, Weld provides benefits if even just 4--5 operators in a library are ported to use it. Our results show that common runtime designs like Weld may be a viable approach to accelerate analytics.
Shoumik Palkar, James Thomas 0003, Deepak Narayanan, Pratiksha Thaker, Rahul Palamuttam, Parimarjan Negi, Anil Shanbhag, Malte Schwarzkopf, Holger Pirk, Saman P. Amarasinghe, Samuel Madden 0001, Matei Zaharia
Proc. VLDB Endow.1
2017 A Common Runtime for High Performance Data Analysis
Shoumik Palkar, James Thomas 0003, Anil Shanbhag, Deepak Narayanan, Holger Pirk, Malte Schwarzkopf, Saman P. Amarasinghe, Matei Zaharia
CIDR1
2017 DIY Hosting for Online Privacy
abstract
Web users today rely on centralized services for applications such as email, file transfer and chat. Unfortunately, these services create a significant privacy risk: even with a benevolent provider, a single breach can put millions of users' data at risk. One alternative would be for users to host their own servers, but this would be highly expensive for most applications: a single VM deployed in a high-availability mode can cost many dollars per month. In this paper, we propose Deploy It Yourself (DIY), a new model for hosting applications based on serverless computing platforms such as Amazon Lambda. DIY allows users to run a highly available service with much stronger privacy guarantees than current centralized providers, and at a dramatically lower cost than traditional server hosting. DIY only relies on the security of container isolation and a key manager as opposed to the large codebase of a high-level application such as Gmail (and all the Google teams using Gmail data). With attestation technology such as SGX, DIY's execution could also be verified remotely. We show that a DIY email server that sends 500 messages/day costs $0.26/month, which is 50x cheaper than a highly available EC2 server. We also implement a DIY chat service and show that it performs well. Finally, we argue that DIY applications are simple enough to operate that cloud providers could offer a simple "app store" for using them.
Shoumik Palkar, Matei Zaharia
HotNets1
2015 E2: a framework for NFV applications
abstract
By moving network appliance functionality from proprietary hardware to software, Network Function Virtualization promises to bring the advantages of cloud computing to network packet processing. However, the evolution of cloud computing (particularly for data analytics) has greatly benefited from application-independent methods for scaling and placement that achieve high efficiency while relieving programmers of these burdens. NFV has no such general management solutions. In this paper, we present a scalable and application-agnostic scheduling framework for packet processing, and compare its performance to current approaches.
Shoumik Palkar, Chang Lan, Sangjin Han, Keon Jang, Aurojit Panda, Sylvia Ratnasamy, Luigi Rizzo, Scott Shenker
SOSP1