Tathagata Das

dblp:31/6214 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
1since 2021 · last 2023
0009-0007-1453-9591ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 1 first-authorDatabases, data management, data science and information retrieval · 4 · 1 since 2021Systems, architecture and hardware · 2 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 first-authorSecurity and privacy · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
7 papers
Storage systems · 41% Cloud and datacenter computing · 29% Distributed systems · 20%
Databases, data mining, and information retrieval
2 papers
Database system architecture and tuning · 36% Query processing and optimization · 36% Data stream processing · 28%
Computer networks
2 papers
Datacenter networks · 70% Network management and operations · 23% Internet of things and sensor networks · 7%
Network and information security
2 papers
Authentication and access control · 53% Systems and software security · 32% Privacy and data protection · 16%

Topics — the 24 heaviest of 29, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems › object storage
cloud object store
0.412020
Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores · Proc. VLDB Endow. 2020
Storage systems
key-value storage
0.412020
Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores · Proc. VLDB Endow. 2020
Storage systems
storage reliability
0.412020
Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores · Proc. VLDB Endow. 2020
Query processing and optimization › incremental computation
incremental query processing
0.312018
Structured Streaming: A Declarative API for Real-Time Applications in Apache Spark · SIGMOD Conference 2018
Cloud and datacenter computing
cluster resource management and scheduling
0.322015
Scaling Spark in the Real World: Performance and Usability · Proc. VLDB Endow. 2015
Discretized streams: fault-tolerant streaming computation at scale · SOSP 2013
Distributed systems
fault tolerance
0.232015
Resilient Distributed Datasets: A Fault-Tolerant Abstraction for In-Memory Cluster Computing · NSDI 2012
Scaling Spark in the Real World: Performance and Usability · Proc. VLDB Endow. 2015
BFT Protocols Under Fire · NSDI 2008
Cloud and datacenter computing
serverless computing
0.212015
Scaling Spark in the Real World: Performance and Usability · Proc. VLDB Endow. 2015
Distributed systems › fault tolerance › fault-tolerant distributed systems
fault-tolerant stream processing
0.212013
Discretized streams: fault-tolerant streaming computation at scale · SOSP 2013
Distributed systems
stream processing
0.212013
Discretized streams: fault-tolerant streaming computation at scale · SOSP 2013
Datacenter networks
datacenter transport
0.112012
DeTail: reducing the flow completion time tail in datacenter networks · SIGCOMM 2012
Datacenter networks › datacenter transport
flow completion time
0.112012
DeTail: reducing the flow completion time tail in datacenter networks · SIGCOMM 2012
High-performance computing
cluster computing
0.112012
Resilient Distributed Datasets: A Fault-Tolerant Abstraction for In-Memory Cluster Computing · NSDI 2012
Cloud and datacenter computing › cluster computing framework
in-memory cluster computing
0.112012
Resilient Distributed Datasets: A Fault-Tolerant Abstraction for In-Memory Cluster Computing · NSDI 2012
Ubiquitous computing and smart environments › mobile crowdsourcing › crowdsensing
mobile crowdsensing
0.112010
PRISM: platform for remote sensing using smartphones · MobiSys 2010
Cloud and datacenter computing
virtualization
0.112010
LiteGreen: Saving Energy in Networked Desktops Using Virtualization · USENIX ATC 2010
Query processing and optimization
query compilation
0.112018
Structured Streaming: A Declarative API for Real-Time Applications in Apache Spark · SIGMOD Conference 2018
Network management and operations › fault management
fault diagnosis
0.112009
NetPrints: Diagnosing Home Network Misconfigurations Using Shared Knowledge · NSDI 2009
Distributed systems › fault tolerance
byzantine fault tolerance
0.112008
BFT Protocols Under Fire · NSDI 2008
Cloud and datacenter computing
datacenter application performance
0.012012
DeTail: reducing the flow completion time tail in datacenter networks · SIGCOMM 2012
Parallel and multicore computing
parallel programming models
0.012012
Resilient Distributed Datasets: A Fault-Tolerant Abstraction for In-Memory Cluster Computing · NSDI 2012
Systems and software security
configuration security
0.012010
Baaz: A System for Detecting Access Control Misconfigurations · USENIX Security Symposium 2010
Systems and software security › operating system security
sandboxing
0.012010
PRISM: platform for remote sensing using smartphones · MobiSys 2010
Privacy and data protection › privacy-preserving sensing
sensor data privacy
0.012010
PRISM: platform for remote sensing using smartphones · MobiSys 2010
Internet of things and sensor networks
home network
0.012009
NetPrints: Diagnosing Home Network Misconfigurations Using Shared Knowledge · NSDI 2009

Methods — techniques the papers use, named apart from their topics

incremental view maintenance · 0.3code generation · 0.3lineage-based recovery · 0.3parallel recovery · 0.2
YearPublicationVenuePosition
2023 Analyzing and Comparing Lakehouse Storage Systems
Paras Jain 0001, Peter Kraft, Conor Power, Tathagata Das, Ion Stoica, Matei Zaharia
CIDR4
2020 Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores
abstract
Cloud object stores such as Amazon S3 are some of the largest and most cost-effective storage systems on the planet, making them an attractive target to store large data warehouses and data lakes. Unfortunately, their implementation as key-value stores makes it difficult to achieve ACID transactions and high performance: metadata operations such as listing objects are expensive, and consistency guarantees are limited. In this paper, we present Delta Lake, an open source ACID table storage layer over cloud object stores initially developed at Databricks. Delta Lake uses a transaction log that is compacted into Apache Parquet format to provide ACID properties, time travel, and significantly faster metadata operations for large tabular datasets (e.g., the ability to quickly search billions of table partitions for those relevant to a query). It also leverages this design to provide high-level features such as automatic data layout optimization, upserts, caching, and audit logs. Delta Lake tables can be accessed from Apache Spark, Hive, Presto, Redshift and other systems. Delta Lake is deployed at thousands of Databricks customers that process exabytes of data per day, with the largest instances managing exabyte-scale datasets and billions of objects.
Michael Armbrust, Tathagata Das, Sameer Paranjpye, Reynold Xin, Shixiong Zhu, Ali Ghodsi 0002, Burak Yavuz, Mukul Murthy, Joseph Torres, Liwen Sun, Peter Boncz, Mostafa Mokhtar, Herman Van Hövell, Adrian Ionescu, Alicja Luszczak, Michal Switakowski, Takuya Ueshin, Xiao Li 0087, Michal Szafranski, Pieter Senster, Matei Zaharia
Proc. VLDB Endow.2
2018 Structured Streaming: A Declarative API for Real-Time Applications in Apache Spark
abstract
With the ubiquity of real-time data, organizations need streaming systems that are scalable, easy to use, and easy to integrate into business applications. Structured Streaming is a new high-level streaming API in Apache Spark based on our experience with Spark Streaming. Structured Streaming differs from other recent streaming APIs, such as Google Dataflow, in two main ways. First, it is a purely declarative API based on automatically incrementalizing a static relational query (expressed using SQL or DataFrames), in contrast to APIs that ask the user to build a DAG of physical operators. Second, Structured Streaming aims to support end-to-end real-time applications that integrate streaming with batch and interactive analysis. We found that this integration was often a key challenge in practice. Structured Streaming achieves high performance via Spark SQL's code generation engine and can outperform Apache Flink by up to 2x and Apache Kafka Streams by 90x. It also offers rich operational features such as rollbacks, code updates, and mixed streaming/batch execution. We describe the system's design and use cases from several hundred production deployments on Databricks, the largest of which process over 1 PB of data per month.
Michael Armbrust, Tathagata Das, Joseph Torres, Burak Yavuz, Shixiong Zhu, Reynold Xin, Ali Ghodsi 0002, Ion Stoica, Matei Zaharia
SIGMOD Conference2
2015 Scaling Spark in the Real World: Performance and Usability
abstract
Apache Spark is one of the most widely used open source processing engines for big data, with rich language-integrated APIs and a wide range of libraries. Over the past two years, our group has worked to deploy Spark to a wide range of organizations through consulting relationships as well as our hosted service, Databricks. We describe the main challenges and requirements that appeared in taking Spark to a wide set of users, and usability and performance improvements we have made to the engine in response.
Michael Armbrust, Tathagata Das, Aaron Davidson, Ali Ghodsi 0002, Andrew Or, Josh Rosen, Ion Stoica, Patrick Wendell, Reynold Xin, Matei Zaharia
Proc. VLDB Endow.2
2014 Adaptive Stream Processing using Dynamic Batch Sizing
abstract
The need for real-time processing of "big data" has led to the development of frameworks for distributed stream processing in clusters. It is important for such frameworks to be robust against variable operating conditions such as server failures, changes in data ingestion rates, and workload characteristics. To provide fault tolerance and efficient stream processing at scale, recent stream processing frameworks have proposed to treat streaming workloads as a series of batch jobs on small batches of streaming data. However, the robustness of such frameworks against variable operating conditions has not been explored.
Tathagata Das, Yuan Zhong 0001, Ion Stoica, Scott Shenker
SoCC1
2013 Discretized streams: fault-tolerant streaming computation at scale
abstract
Many "big data" applications must act on data in real time. Running these applications at ever-larger scales requires parallel platforms that automatically handle faults and stragglers. Unfortunately, current distributed stream processing models provide fault recovery in an expensive manner, requiring hot replication or long recovery times, and do not handle stragglers. We propose a new processing model, discretized streams (D-Streams), that overcomes these challenges. D-Streams enable a parallel recovery mechanism that improves efficiency over traditional replication and backup schemes, and tolerates stragglers. We show that they support a rich set of operators while attaining high per-node throughput similar to single-node systems, linear scaling to 100 nodes, sub-second latency, and sub-second fault recovery. Finally, D-Streams can easily be composed with batch and interactive query models like MapReduce, enabling rich applications that combine these modes. We implement D-Streams in a system called Spark Streaming.
Matei Zaharia, Tathagata Das, Haoyuan Li 0001, Timothy Hunter, Scott Shenker, Ion Stoica
SOSP2
2013 Large-Scale Estimation in Cyberphysical Systems Using Streaming Data: A Case Study With Arterial Traffic Estimation
abstract
Controlling and analyzing cyberphysical and robotics systems is increasingly becoming a Big Data challenge. We study the case of predicting drivers' travel times in a large urban area from sparse GPS traces. We present a framework that can accommodate a wide variety of traffic distributions and spread all the computations on a cluster to achieve small latencies. Our framework is built on Discretized Streams, a recently proposed approach to stream processing at scale. We demonstrate the usefulness of Discretized Streams with a novel algorithm to estimate vehicular traffic in urban networks. Our online EM algorithm can estimate traffic on a very large city network (the San Francisco Bay Area) by processing tens of thousands of observations per second, with a latency of a few seconds.
Timothy Hunter, Tathagata Das, Matei Zaharia, Pieter Abbeel, Alexandre M. Bayen
IEEE Trans Autom. Sci. Eng.2
2012 Resilient Distributed Datasets: A Fault-Tolerant Abstraction for In-Memory Cluster Computing
Matei Zaharia, Mosharaf Chowdhury, Tathagata Das, Ankur Dave, Justin Ma, Murphy McCauly, Michael J. Franklin, Scott Shenker, Ion Stoica
NSDI3
2012 DeTail: reducing the flow completion time tail in datacenter networks
abstract
Web applications have now become so sophisticated that rendering a typical page may require hundreds of intra-datacenter flows. At the same time, web sites must meet strict page creation deadlines of 200-300ms to satisfy user demands for interactivity. Long-tailed flow completion times make it challenging for web sites to meet these constraints. They are forced to choose between rendering a subset of the complex page, or delay its rendering, thus missing deadlines and sacrificing either quality or responsiveness. Either option leads to potential financial loss.
David Zats, Tathagata Das, Prashanth Mohan, Dhruba Borthakur, Randy H. Katz
SIGCOMM2
2010 PRISM: platform for remote sensing using smartphones
abstract
To realize the potential of opportunistic and participatory sensing using mobile smartphones, a key challenge is ensuring the ease of developing and deploying such applications, without the need for the application writer to reinvent the wheel each time. To this end, we present a Platform for Remote Sensing using Smartphones (PRISM) that balances the interconnected goals of generality, security, and scalability. PRISM allows application writers to package their applications as executable binaries, which offers efficiency and also the flexibility of reusing existing code modules. PRISM then pushes the application out automatically to an appropriate set of phones based on a specified set of predicates. This push model enables timely and scalable application deployment while still ensuring a good degree of privacy. To safely execute untrusted applications on the smartphones, while allowing them controlled access to sensitive sensor data, we augment standard software sandboxing with several PRISM-specific elements like resource metering and forced amnesia.
Tathagata Das, Prashanth Mohan, Venkat N. Padmanabhan, Ramachandran Ramjee, Asankhaya Sharma
MobiSys1
2010 LiteGreen: Saving Energy in Networked Desktops Using Virtualization
Tathagata Das, Pradeep Padala, Venkat N. Padmanabhan, Ramachandran Ramjee, Kang G. Shin
USENIX ATC1
2010 Baaz: A System for Detecting Access Control Misconfigurations
Tathagata Das, Ranjita Bhagwan, Prasad Naldurg
USENIX Security Symposium1
2009 NetPrints: Diagnosing Home Network Misconfigurations Using Shared Knowledge
Bhavish Agarwal, Ranjita Bhagwan, Tathagata Das, Siddharth Eswaran, Venkat N. Padmanabhan, Geoffrey M. Voelker
NSDI3
2008 BFT Protocols Under Fire
Atul Singh, Tathagata Das, Petros Maniatis, Peter Druschel, Timothy Roscoe
NSDI2
2008 Bio-inspired Search and Distributed Memory Formation on Power-Law Networks
Tathagata Das, Subrata Nandi, Andreas Deutsch, Niloy Ganguly
PPSN1