EDBT 2026 Demo / reviewers in the wild / expert
Pratiksha Thaker
dblp:123/5171
· DBLP profile ↗
11ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0001-9977-5081ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Computer networks · 3 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Parallel and multicore computing · 44% Cloud and datacenter computing · 22% Storage systems · 17% | |
| Network and information security
2 papers |
Privacy and data protection · 100% | |
| Computer networks
3 papers |
Network optimization and economics · 48% Datacenter networks · 37% Transport protocols and congestion control · 11% | |
| Artificial intelligence
2 papers |
Trustworthy machine learning · 74% Language models and text generation · 26% | |
| Theoretical computer science
1 paper |
Algorithmic game theory and mechanism design · 100% |
Topics — the 15 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Privacy and data protection › differential privacy › privacy auditing
privacy violation detection |
1.0 | 1 | 2026 | Generate-then-Verify: Reconstructing Data from Limited Published Statistics · SP 2026 |
Privacy and data protection
statistical database privacy |
1.0 | 1 | 2026 | Generate-then-Verify: Reconstructing Data from Limited Published Statistics · SP 2026 |
Parallel and multicore computing › parallel query processing
intra-query parallelism |
0.9 | 1 | 2025 | PARALLELPROMPT: Extracting Parallelism from Large Language Model Queries · NeurIPS 2025 |
Cloud and datacenter computing › inference serving
LLM serving |
0.9 | 1 | 2025 | PARALLELPROMPT: Extracting Parallelism from Large Language Model Queries · NeurIPS 2025 |
Parallel and multicore computing
parallel query processing |
0.9 | 1 | 2025 | PARALLELPROMPT: Extracting Parallelism from Large Language Model Queries · NeurIPS 2025 |
Machine learning › Trustworthy machine learning › robustness
distribution shift |
0.8 | 1 | 2024 | On the Benefits of Public Representations for Private Transfer Learning under Distribution Shift · NeurIPS 2024 |
Privacy and data protection
differential privacy |
0.8 | 1 | 2024 | On the Benefits of Public Representations for Private Transfer Learning under Distribution Shift · NeurIPS 2024 |
Privacy and data protection › privacy-preserving machine learning
private training |
0.8 | 1 | 2024 | On the Benefits of Public Representations for Private Transfer Learning under Distribution Shift · NeurIPS 2024 |
Network optimization and economics
fairness |
0.7 | 1 | 2023 | Congestion Control Safety via Comparative Statics · INFOCOM 2023 |
Storage systems
key-value storage |
0.7 | 1 | 2023 | Cornflakes: Zero-Copy Serialization for Microsecond-Scale Networking · SOSP 2023 |
Distributed systems › concurrency control
serialization |
0.7 | 1 | 2023 | Cornflakes: Zero-Copy Serialization for Microsecond-Scale Networking · SOSP 2023 |
Algorithmic game theory and mechanism design
congestion games |
0.7 | 1 | 2023 | Congestion Control Safety via Comparative Statics · INFOCOM 2023 |
Transport protocols and congestion control › transport protocols
UDP |
0.2 | 1 | 2023 | Cornflakes: Zero-Copy Serialization for Microsecond-Scale Networking · SOSP 2023 |
Network optimization and economics
utility function |
0.2 | 1 | 2023 | Congestion Control Safety via Comparative Statics · INFOCOM 2023 |
Network performance modeling
protocol performance analysis |
0.1 | 1 | 2014 | An experimental study of the learnability of congestion control · SIGCOMM 2014 |
Methods — techniques the papers use, named apart from their topics
rule-based multilingual validation · 1.7large language model prompting · 1.7representation learning · 1.5public pretraining · 1.5game theory · 1.3comparative statics · 1.3integer programming · 1.0generate-then-verify · 1.0runtime measurement · 0.7pipelining · 0.7adaptive optimization · 0.7experimental study · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Generate-then-Verify: Reconstructing Data from Limited Published StatisticsabstractWe study the problem of reconstructing tabular data from aggregate statistics, in which the attacker aims to identify interesting claims about the sensitive data that can be verified with 100% certainty given the aggregates. Successful attempts in prior work have conducted studies in settings where the set of published statistics is rich enough that entire datasets can be reconstructed with certainty. In our work, we instead focus on the regime where many possible datasets match the published statistics, making it impossible to reconstruct the entire private dataset perfectly (i.e., when approaches in prior work fail). We propose the problem of partial data reconstruction, in which the goal of the adversary is to instead output a $\textit{subset}$ of rows and/or columns that are $\textit{guaranteed to be correct}$. We introduce a novel integer programming approach that first $\textbf{generates}$ a set of claims and then $\textbf{verifies}$ whether each claim holds for all possible datasets consistent with the published aggregates. We evaluate our approach on the housing-level microdata from the U.S. Decennial Census release, demonstrating that privacy violations can still persist even when information published about such data is relatively sparse. Terrance Liu, Eileen Xiao, Adam D. Smith 0001, Pratiksha Thaker, Steven Z. Wu |
SP | 4 |
| 2025 | PARALLELPROMPT: Extracting Parallelism from Large Language Model QueriesabstractLLM serving systems typically treat user prompts as monolithic inputs, optimizing inference through decoding tricks or inter-query batching. However, many real-world prompts contain *latent semantic parallelism*—decomposable structures where subtasks can be executed independently to reduce latency while preserving meaning.We introduce PARALLELPROMPT, the first benchmark for measuring intra-query parallelism in natural user prompts. Our dataset comprises over 37,000 real-world prompts from public LLM chat logs, each annotated with a structured schema capturing task templates, shared context, and iteration inputs. These schemas are extracted using LLM-assisted prompting with rule-based multilingual validation.To evaluate the benefits of decomposition, we provide an execution suite that benchmarks serial vs. parallel strategies, measuring latency, structural adherence, and semantic fidelity. Our results show that intra-query parallelism can be successfully parsed in over 75\% of curated datasets, unlocking up to *$5\times$ speedups* on tasks like translation, comprehension, and comparative analysis, with minimal quality degradation.By releasing this benchmark, curation pipeline, and evaluation suite, we provide the first standardized testbed for studying structure-aware execution in LLM serving pipelines. Steven Kolawole, Keshav Santhanam, Virginia Smith, Pratiksha Thaker |
NeurIPS | 4 |
| 2024 | On the Benefits of Public Representations for Private Transfer Learning under Distribution ShiftabstractPublic pretraining is a promising approach to improve differentially private model training. However, recent work has noted that many positive research results studying this paradigm only consider in-distribution tasks, and may not apply to settings where there is distribution shift between the pretraining and finetuning data---a scenario that is likely when finetuning private tasks due to the sensitive nature of the data. In this work, we show empirically across three tasks that even in settings with large distribution shift, where both zero-shot performance from public data and training from scratch with private data give unusably weak results, public features can in fact improve private training accuracy by up to 67\% over private training from scratch. We provide a theoretical explanation for this phenomenon, showing that if the public and private data share a low-dimensional representation, public representations can improve the sample complexity of private training even if it is \emph{impossible} to learn the private task from the public data alone. Altogether, our results provide evidence that public data can indeed make private training practical in realistic settings of extreme distribution shift. Pratiksha Thaker, Amrith Setlur, Steven Z. Wu, Virginia Smith |
NeurIPS | 1 |
| 2023 | Congestion Control Safety via Comparative StaticsabstractWhen congestion control algorithms compete on shared links, unfair outcomes can result, especially between algorithms that aim to prioritize different objectives. For example, a throughput-maximizing application could make the link completely unusable for a latency-sensitive application. In order to study these outcomes formally, we model the congestion control problem as a game in which applications have heterogeneous utility functions. We draw on the comparative statics literature in economics to derive simple and practically useful conditions under which all applications achieve at least ε utility at equilibrium, a minimal safety condition for the network to be useful for any application. Compared to prior analyses of similar games, we show that our framework supports a more realistic class of utility functions that includes highly latency-sensitive applications such as teleconferencing and online gaming. Pratiksha Thaker, Matei Zaharia, Tatsunori B. Hashimoto |
INFOCOM | 1 |
| 2023 | Cornflakes: Zero-Copy Serialization for Microsecond-Scale NetworkingabstractData serialization is critical for many datacenter applications, but the memory copies required to move application data into packets are costly. Recent zero-copy APIs expose NIC scatter-gather capabilities, raising the possibility of offloading this data movement to the NIC. However, as the memory coordination required for scatter-gather adds bookkeeping overhead, scatter-gather is not always useful. We describe Cornflakes, a hybrid serialization library stack that uses scatter-gather for serialization when it improves performance and falls back to memory copies otherwise. We have implemented Cornflakes within a UDP and TCP networking stack, across Mellanox and Intel NICs. On a Twitter cache trace, Cornflakes achieves 15.4% higher throughput than prior software approaches on a custom key-value store and 8.8% higher throughput than Redis serialization within Redis. Deepti Raghavan, Shreya Ravi, Gina Yuan, Pratiksha Thaker, Sanjari Srivastava, Micah Murray, Pedro Henrique de Mello Morado Penna, Amy Ousterhout, Philip Alexander Levis, Matei Zaharia, Irene Zhang |
SOSP | 4 |
| 2021 | Clamor: Extending Functional Cluster Computing Frameworks with Fine-Grained Remote Memory AccessabstractWe propose Clamor, a functional cluster computing framework that adds support for fine-grained, transparent access to global variables for distributed, data-parallel tasks. Clamor targets workloads that perform sparse accesses and updates within the bulk synchronous parallel execution model, a setting where the standard technique of broadcasting global variables is highly inefficient. Clamor implements a novel dynamic replication mechanism in order to enable efficient access to popular data regions on the fly, and tracks finegrained dependencies in order to retain the lineage-based fault tolerance model of systems like Spark. Clamor can integrate with existing Rust and C++ libraries to transparently distribute programs on the cluster. We show that Clamor is competitive with Spark in simple functional workloads and can improve performance significantly compared to custom systems on workloads that sparsely access large global variables: from 5x for sparse logistic regression to over 100x on distributed geospatial queries. Pratiksha Thaker, Hudson Ayers, Deepti Raghavan, Ning Niu, Philip Alexander Levis, Matei Zaharia |
SoCC | 1 |
| 2021 | Don't Hate the Player, Hate the Game: Safety and Utility in Multi-Agent Congestion ControlabstractWe posit that unfairness between congestion control algorithms in a network often results from actors optimizing, or attempting to optimize, incompatible utility functions. In order to mitigate this, we propose that algorithm designers should explicitly declare a utility function they hope to optimize, which enables theoretical analysis of the safety of the utility function with respect to other utilities in the network. When we can place utilities in a common game-theoretic framework, we can analytically determine the potential for an application with one of those utilities to be unsafe before it is deployed in a network, rather than determining safety properties ad-hoc from measurements after deployment. We give examples of the types of restrictions and guarantees that can arise from such a model in the context of rate-based congestion control protocols. Pratiksha Thaker, Matei Zaharia, Tatsunori B. Hashimoto |
HotNets | 1 |
| 2018 | Evaluating End-to-End Optimization for Data Analytics Applications in WeldabstractModern analytics applications use a diverse mix of libraries and functions. Unfortunately, there is no optimization across these libraries, resulting in performance penalties as high as an order of magnitude in many applications. To address this problem, we proposed Weld, a common runtime for existing data analytics libraries that performs key physical optimizations such as pipelining under existing, imperative library APIs. In this work, we further develop the Weld vision by designing an automatic adaptive optimizer for Weld applications, and evaluating its impact on realistic data science workloads. Our optimizer eliminates multiple forms of overhead that arise when composing imperative libraries like Pandas and NumPy, and uses lightweight measurements to make data-dependent decisions at run-time in ad-hoc workloads where no statistics are available, with sub-second overhead. We also evaluate which optimizations have the largest impact in practice and whether Weld can be integrated into libraries incrementally. Our results are promising: using our optimizer, Weld accelerates data science workloads by up to 23X on one thread and 80X on eight threads, and its adaptive optimizations provide up to a 3.75X speedup over rule-based optimization. Moreover, Weld provides benefits if even just 4--5 operators in a library are ported to use it. Our results show that common runtime designs like Weld may be a viable approach to accelerate analytics. Shoumik Palkar, James Thomas 0003, Deepak Narayanan, Pratiksha Thaker, Rahul Palamuttam, Parimarjan Negi, Anil Shanbhag, Malte Schwarzkopf, Holger Pirk, Saman P. Amarasinghe, Samuel Madden 0001, Matei Zaharia |
Proc. VLDB Endow. | 4 |
| 2014 | An experimental study of the learnability of congestion controlabstractWhen designing a distributed network protocol, typically it is infeasible to fully define the target network where the protocol is intended to be used. It is therefore natural to ask: How faithfully do protocol designers really need to understand the networks they design for? What are the important signals that endpoints should listen to? How can researchers gain confidence that systems that work well on well-characterized test networks during development will also perform adequately on real networks that are inevitably more complex, or future networks yet to be developed? Is there a tradeoff between the performance of a protocol and the breadth of its intended operating range of networks? What is the cost of playing fairly with cross-traffic that is governed by another protocol? Anirudh Sivaraman, Keith Winstein, Pratiksha Thaker, Hari Balakrishnan |
SIGCOMM | 3 |
| 2014 | Learning perceptually grounded word meanings from unaligned parallel data
Stefanie Tellex, Pratiksha Thaker, Joshua Mason Joseph, Nicholas Roy |
Mach. Learn. | 2 |
| 2013 | Clarifying commands with information-theoretic human-robot dialogabstractOur goal is to improve the efficiency and effectiveness of natural language communication between humans and robots. Human language is frequently ambiguous, and a robot's limited sensing makes complete understanding of a statement even more difficult. To address these challenges, we describe an approach for enabling a robot to engage in clarifying dialog with a human partner, just as a human might do in a similar situation. Given an unconstrained command from a human operator, the robot asks one or more questions and receives natural language answers from the human. We apply an information-theoretic approach to choosing questions for the robot to ask. Specifically, we choose the type and subject of questions in order to maximize the reduction in Shannon entropy of the robot's mapping between language and entities in the world. Within the framework of the G3 graphical model, we derive a method to estimate this entropy reduction, choose the optimal question to ask, and merge the information gained from the human operator's answer. We demonstrate that this improves the accuracy of command understanding over prior work while asking fewer questions as compared to baseline question-selection strategies. Robin Deits, Stefanie Tellex, Pratiksha Thaker, Dimitar Simeonov, Thomas Kollar, Nicholas Roy |
J. Hum. Robot Interact. | 3 |