VLDB 2026 Research / reviewers in the wild / expert
Tathagata Das
dblp:31/6214
· DBLP profile ↗
15ranked-venue papers
5as first author
1since 2021 · last 2023
0009-0007-1453-9591ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 5 · 1 first-authorDatabases, data management, data science and information retrieval · 4 · 1 since 2021Systems, architecture and hardware · 2 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 first-authorSecurity and privacy · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
7 papers |
Storage systems · 41% Cloud and datacenter computing · 29% Distributed systems · 20% | |
| Databases, data mining, and information retrieval
2 papers |
Database system architecture and tuning · 36% Query processing and optimization · 36% Data stream processing · 28% | |
| Computer networks
2 papers |
Datacenter networks · 70% Network management and operations · 23% Internet of things and sensor networks · 7% | |
| Network and information security
2 papers |
Authentication and access control · 53% Systems and software security · 32% Privacy and data protection · 16% |
Topics — the 24 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems › object storage
cloud object store |
0.4 | 1 | 2020 | Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores · Proc. VLDB Endow. 2020 |
Storage systems
key-value storage |
0.4 | 1 | 2020 | Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores · Proc. VLDB Endow. 2020 |
Storage systems
storage reliability |
0.4 | 1 | 2020 | Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores · Proc. VLDB Endow. 2020 |
Query processing and optimization › incremental computation
incremental query processing |
0.3 | 1 | 2018 | Structured Streaming: A Declarative API for Real-Time Applications in Apache Spark · SIGMOD Conference 2018 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.3 | 2 | 2015 | Scaling Spark in the Real World: Performance and Usability · Proc. VLDB Endow. 2015 Discretized streams: fault-tolerant streaming computation at scale · SOSP 2013 |
Distributed systems
fault tolerance |
0.2 | 3 | 2015 | Resilient Distributed Datasets: A Fault-Tolerant Abstraction for In-Memory Cluster Computing · NSDI 2012 Scaling Spark in the Real World: Performance and Usability · Proc. VLDB Endow. 2015 BFT Protocols Under Fire · NSDI 2008 |
Cloud and datacenter computing
serverless computing |
0.2 | 1 | 2015 | Scaling Spark in the Real World: Performance and Usability · Proc. VLDB Endow. 2015 |
Distributed systems › fault tolerance › fault-tolerant distributed systems
fault-tolerant stream processing |
0.2 | 1 | 2013 | Discretized streams: fault-tolerant streaming computation at scale · SOSP 2013 |
Distributed systems
stream processing |
0.2 | 1 | 2013 | Discretized streams: fault-tolerant streaming computation at scale · SOSP 2013 |
Datacenter networks
datacenter transport |
0.1 | 1 | 2012 | DeTail: reducing the flow completion time tail in datacenter networks · SIGCOMM 2012 |
Datacenter networks › datacenter transport
flow completion time |
0.1 | 1 | 2012 | DeTail: reducing the flow completion time tail in datacenter networks · SIGCOMM 2012 |
High-performance computing
cluster computing |
0.1 | 1 | 2012 | Resilient Distributed Datasets: A Fault-Tolerant Abstraction for In-Memory Cluster Computing · NSDI 2012 |
Cloud and datacenter computing › cluster computing framework
in-memory cluster computing |
0.1 | 1 | 2012 | Resilient Distributed Datasets: A Fault-Tolerant Abstraction for In-Memory Cluster Computing · NSDI 2012 |
Ubiquitous computing and smart environments › mobile crowdsourcing › crowdsensing
mobile crowdsensing |
0.1 | 1 | 2010 | PRISM: platform for remote sensing using smartphones · MobiSys 2010 |
Cloud and datacenter computing
virtualization |
0.1 | 1 | 2010 | LiteGreen: Saving Energy in Networked Desktops Using Virtualization · USENIX ATC 2010 |
Query processing and optimization
query compilation |
0.1 | 1 | 2018 | Structured Streaming: A Declarative API for Real-Time Applications in Apache Spark · SIGMOD Conference 2018 |
Network management and operations › fault management
fault diagnosis |
0.1 | 1 | 2009 | NetPrints: Diagnosing Home Network Misconfigurations Using Shared Knowledge · NSDI 2009 |
Distributed systems › fault tolerance
byzantine fault tolerance |
0.1 | 1 | 2008 | BFT Protocols Under Fire · NSDI 2008 |
Cloud and datacenter computing
datacenter application performance |
0.0 | 1 | 2012 | DeTail: reducing the flow completion time tail in datacenter networks · SIGCOMM 2012 |
Parallel and multicore computing
parallel programming models |
0.0 | 1 | 2012 | Resilient Distributed Datasets: A Fault-Tolerant Abstraction for In-Memory Cluster Computing · NSDI 2012 |
Systems and software security
configuration security |
0.0 | 1 | 2010 | Baaz: A System for Detecting Access Control Misconfigurations · USENIX Security Symposium 2010 |
Systems and software security › operating system security
sandboxing |
0.0 | 1 | 2010 | PRISM: platform for remote sensing using smartphones · MobiSys 2010 |
Privacy and data protection › privacy-preserving sensing
sensor data privacy |
0.0 | 1 | 2010 | PRISM: platform for remote sensing using smartphones · MobiSys 2010 |
Internet of things and sensor networks
home network |
0.0 | 1 | 2009 | NetPrints: Diagnosing Home Network Misconfigurations Using Shared Knowledge · NSDI 2009 |
Methods — techniques the papers use, named apart from their topics
incremental view maintenance · 0.3code generation · 0.3lineage-based recovery · 0.3parallel recovery · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Analyzing and Comparing Lakehouse Storage Systems
Paras Jain 0001, Peter Kraft, Conor Power, Tathagata Das, Ion Stoica, Matei Zaharia |
CIDR | 4 |
| 2020 | Delta Lake: High-Performance ACID Table Storage over Cloud Object StoresabstractCloud object stores such as Amazon S3 are some of the largest and most cost-effective storage systems on the planet, making them an attractive target to store large data warehouses and data lakes. Unfortunately, their implementation as key-value stores makes it difficult to achieve ACID transactions and high performance: metadata operations such as listing objects are expensive, and consistency guarantees are limited. In this paper, we present Delta Lake, an open source ACID table storage layer over cloud object stores initially developed at Databricks. Delta Lake uses a transaction log that is compacted into Apache Parquet format to provide ACID properties, time travel, and significantly faster metadata operations for large tabular datasets (e.g., the ability to quickly search billions of table partitions for those relevant to a query). It also leverages this design to provide high-level features such as automatic data layout optimization, upserts, caching, and audit logs. Delta Lake tables can be accessed from Apache Spark, Hive, Presto, Redshift and other systems. Delta Lake is deployed at thousands of Databricks customers that process exabytes of data per day, with the largest instances managing exabyte-scale datasets and billions of objects. Michael Armbrust, Tathagata Das, Sameer Paranjpye, Reynold Xin, Shixiong Zhu, Ali Ghodsi 0002, Burak Yavuz, Mukul Murthy, Joseph Torres, Liwen Sun, Peter Boncz, Mostafa Mokhtar, Herman Van Hövell, Adrian Ionescu, Alicja Luszczak, Michal Switakowski, Takuya Ueshin, Xiao Li 0087, Michal Szafranski, Pieter Senster, Matei Zaharia |
Proc. VLDB Endow. | 2 |
| 2018 | Structured Streaming: A Declarative API for Real-Time Applications in Apache SparkabstractWith the ubiquity of real-time data, organizations need streaming systems that are scalable, easy to use, and easy to integrate into business applications. Structured Streaming is a new high-level streaming API in Apache Spark based on our experience with Spark Streaming. Structured Streaming differs from other recent streaming APIs, such as Google Dataflow, in two main ways. First, it is a purely declarative API based on automatically incrementalizing a static relational query (expressed using SQL or DataFrames), in contrast to APIs that ask the user to build a DAG of physical operators. Second, Structured Streaming aims to support end-to-end real-time applications that integrate streaming with batch and interactive analysis. We found that this integration was often a key challenge in practice. Structured Streaming achieves high performance via Spark SQL's code generation engine and can outperform Apache Flink by up to 2x and Apache Kafka Streams by 90x. It also offers rich operational features such as rollbacks, code updates, and mixed streaming/batch execution. We describe the system's design and use cases from several hundred production deployments on Databricks, the largest of which process over 1 PB of data per month. Michael Armbrust, Tathagata Das, Joseph Torres, Burak Yavuz, Shixiong Zhu, Reynold Xin, Ali Ghodsi 0002, Ion Stoica, Matei Zaharia |
SIGMOD Conference | 2 |
| 2015 | Scaling Spark in the Real World: Performance and UsabilityabstractApache Spark is one of the most widely used open source processing engines for big data, with rich language-integrated APIs and a wide range of libraries. Over the past two years, our group has worked to deploy Spark to a wide range of organizations through consulting relationships as well as our hosted service, Databricks. We describe the main challenges and requirements that appeared in taking Spark to a wide set of users, and usability and performance improvements we have made to the engine in response. Michael Armbrust, Tathagata Das, Aaron Davidson, Ali Ghodsi 0002, Andrew Or, Josh Rosen, Ion Stoica, Patrick Wendell, Reynold Xin, Matei Zaharia |
Proc. VLDB Endow. | 2 |
| 2014 | Adaptive Stream Processing using Dynamic Batch SizingabstractThe need for real-time processing of "big data" has led to the development of frameworks for distributed stream processing in clusters. It is important for such frameworks to be robust against variable operating conditions such as server failures, changes in data ingestion rates, and workload characteristics. To provide fault tolerance and efficient stream processing at scale, recent stream processing frameworks have proposed to treat streaming workloads as a series of batch jobs on small batches of streaming data. However, the robustness of such frameworks against variable operating conditions has not been explored. Tathagata Das, Yuan Zhong 0001, Ion Stoica, Scott Shenker |
SoCC | 1 |
| 2013 | Discretized streams: fault-tolerant streaming computation at scaleabstractMany "big data" applications must act on data in real time. Running these applications at ever-larger scales requires parallel platforms that automatically handle faults and stragglers. Unfortunately, current distributed stream processing models provide fault recovery in an expensive manner, requiring hot replication or long recovery times, and do not handle stragglers. We propose a new processing model, discretized streams (D-Streams), that overcomes these challenges. D-Streams enable a parallel recovery mechanism that improves efficiency over traditional replication and backup schemes, and tolerates stragglers. We show that they support a rich set of operators while attaining high per-node throughput similar to single-node systems, linear scaling to 100 nodes, sub-second latency, and sub-second fault recovery. Finally, D-Streams can easily be composed with batch and interactive query models like MapReduce, enabling rich applications that combine these modes. We implement D-Streams in a system called Spark Streaming. Matei Zaharia, Tathagata Das, Haoyuan Li 0001, Timothy Hunter, Scott Shenker, Ion Stoica |
SOSP | 2 |
| 2013 | Large-Scale Estimation in Cyberphysical Systems Using Streaming Data: A Case Study With Arterial Traffic EstimationabstractControlling and analyzing cyberphysical and robotics systems is increasingly becoming a Big Data challenge. We study the case of predicting drivers' travel times in a large urban area from sparse GPS traces. We present a framework that can accommodate a wide variety of traffic distributions and spread all the computations on a cluster to achieve small latencies. Our framework is built on Discretized Streams, a recently proposed approach to stream processing at scale. We demonstrate the usefulness of Discretized Streams with a novel algorithm to estimate vehicular traffic in urban networks. Our online EM algorithm can estimate traffic on a very large city network (the San Francisco Bay Area) by processing tens of thousands of observations per second, with a latency of a few seconds. Timothy Hunter, Tathagata Das, Matei Zaharia, Pieter Abbeel, Alexandre M. Bayen |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2012 | Resilient Distributed Datasets: A Fault-Tolerant Abstraction for In-Memory Cluster Computing
Matei Zaharia, Mosharaf Chowdhury, Tathagata Das, Ankur Dave, Justin Ma, Murphy McCauly, Michael J. Franklin, Scott Shenker, Ion Stoica |
NSDI | 3 |
| 2012 | DeTail: reducing the flow completion time tail in datacenter networksabstractWeb applications have now become so sophisticated that rendering a typical page may require hundreds of intra-datacenter flows. At the same time, web sites must meet strict page creation deadlines of 200-300ms to satisfy user demands for interactivity. Long-tailed flow completion times make it challenging for web sites to meet these constraints. They are forced to choose between rendering a subset of the complex page, or delay its rendering, thus missing deadlines and sacrificing either quality or responsiveness. Either option leads to potential financial loss. David Zats, Tathagata Das, Prashanth Mohan, Dhruba Borthakur, Randy H. Katz |
SIGCOMM | 2 |
| 2010 | PRISM: platform for remote sensing using smartphonesabstractTo realize the potential of opportunistic and participatory sensing using mobile smartphones, a key challenge is ensuring the ease of developing and deploying such applications, without the need for the application writer to reinvent the wheel each time. To this end, we present a Platform for Remote Sensing using Smartphones (PRISM) that balances the interconnected goals of generality, security, and scalability. PRISM allows application writers to package their applications as executable binaries, which offers efficiency and also the flexibility of reusing existing code modules. PRISM then pushes the application out automatically to an appropriate set of phones based on a specified set of predicates. This push model enables timely and scalable application deployment while still ensuring a good degree of privacy. To safely execute untrusted applications on the smartphones, while allowing them controlled access to sensitive sensor data, we augment standard software sandboxing with several PRISM-specific elements like resource metering and forced amnesia. Tathagata Das, Prashanth Mohan, Venkat N. Padmanabhan, Ramachandran Ramjee, Asankhaya Sharma |
MobiSys | 1 |
| 2010 | LiteGreen: Saving Energy in Networked Desktops Using Virtualization
Tathagata Das, Pradeep Padala, Venkat N. Padmanabhan, Ramachandran Ramjee, Kang G. Shin |
USENIX ATC | 1 |
| 2010 | Baaz: A System for Detecting Access Control Misconfigurations
Tathagata Das, Ranjita Bhagwan, Prasad Naldurg |
USENIX Security Symposium | 1 |
| 2009 | NetPrints: Diagnosing Home Network Misconfigurations Using Shared Knowledge
Bhavish Agarwal, Ranjita Bhagwan, Tathagata Das, Siddharth Eswaran, Venkat N. Padmanabhan, Geoffrey M. Voelker |
NSDI | 3 |
| 2008 | BFT Protocols Under Fire
Atul Singh, Tathagata Das, Petros Maniatis, Peter Druschel, Timothy Roscoe |
NSDI | 2 |
| 2008 | Bio-inspired Search and Distributed Memory Formation on Power-Law Networks
Tathagata Das, Subrata Nandi, Andreas Deutsch, Niloy Ganguly |
PPSN | 1 |