VLDB 2026 Research / reviewers in the wild / expert
Jayshree Sarathy
dblp:270/0818
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0001-5852-5182ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | "Having Confidence in My Confidence Intervals": How Data Users Engage with Privacy-Protected Wikipedia DataabstractIn response to calls for open data and growing privacy threats, organizations are increasingly adopting privacy-preserving techniques that add noise to published datasets. These techniques seek to protect privacy of data subjects while enabling useful analyses. With expert feedback, we developed empirically-driven documentation explaining the noise characteristics of two Wikipedia pageview datasets: one using rounding (heuristic privacy) and another using differential privacy (DP, formal privacy). We then used these documents to conduct a task-based contextual inquiry (n=15) exploring how data users—largely unfamiliar with these methods—perceive, interact with, and interpret privacy-preserving noise during data analysis. Harold Triedman, Jayshree Sarathy, Priyanka Nanayakkara, Rachel Cummings, Gabriel Kaptchuk, Sean Kross, Elissa M. Redmiles |
CHI | 2 |
| 2025 | Analyzing the Differentially Private Theil-Sen Estimator for Simple Linear RegressionabstractIn this paper, we study differentially private point and confidence interval estimators for simple linear regression. Motivated by recent work that highlights the strong empirical performance of an algorithm based on robust statistics, DPTheilSen, we provide a rigorous, finite-sample analysis of its privacy and accuracy properties, offer guidance on setting hyperparameters, and show how to produce differentially private confidence intervals to accompany its point estimates. Jayshree Sarathy, Salil P. Vadhan |
Proc. Priv. Enhancing Technol. | 1 |
| 2024 | "I inherently just trust that it works": Investigating Mental Models of Open-Source Libraries for Differential PrivacyabstractDifferential privacy (DP) is a promising framework for privacy-preserving data science, but recent studies have exposed challenges in bringing this theoretical framework for privacy into practice. These tensions are particularly salient in the context of open-source software libraries for DP data analysis, which are emerging tools to help data stewards and analysts build privacy-preserving data pipelines for their applications. While there has been significant investment into such libraries, we need further inquiry into the role of these libraries in promoting understanding of and trust in DP, and in turn, the ways in which design of these open-source libraries can shed light on the challenges of creating trustworthy data infrastructures in practice. In this study, we use qualitative methods and mental models approaches to analyze the differences between conceptual models used to design open-source DP libraries and mental models of DP held by users. Through a two-stage study design involving formative interviews with 5 developers of open-source DP libraries and user studies with 17 data analysts, we find that DP libraries often struggle to bridge the gaps between developer and user mental models. In particular, we highlight the tension DP libraries face in maintaining rigorous DP implementations and facilitating user interaction. We conclude by offering practical recommendations for further development of DP libraries. Patrick Song, Jayshree Sarathy, Michael Shoemate, Salil P. Vadhan |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2023 | Don't Look at the Data! How Differential Privacy Reconfigures the Practices of Data ScienceabstractAcross academia, government, and industry, data stewards are facing increasing pressure to make datasets more openly accessible for researchers while also protecting the privacy of data subjects. Differential privacy (DP) is one promising way to offer privacy along with open access, but further inquiry is needed into the tensions between DP and data science. In this study, we conduct interviews with 19 data practitioners who are non-experts in DP as they use a DP data analysis prototype to release privacy-preserving statistics about sensitive data, in order to understand perceptions, challenges, and opportunities around using DP. We find that while DP is promising for providing wider access to sensitive datasets, it also introduces challenges into every stage of the data science workflow. We identify ethics and governance questions that arise when socializing data scientists around new privacy constraints and offer suggestions to better integrate DP and data science. Jayshree Sarathy, Sophia Song, Audrey Haque, Tania Schlatter, Salil P. Vadhan |
CHI | 1 |
| 2022 | Differentially Private Simple Linear RegressionabstractAbstract Economics and social science research often require analyzing datasets of sensitive personal information at fine granularity, with models fit to small subsets of the data. Unfortunately, such fine-grained analysis can easily reveal sensitive individual information. We study regression algorithms that satisfy differential privacy, a constraint which guarantees that an algorithm’s output reveals little about any individual input data record, even to an attacker with side information about the dataset. Motivated by the Opportunity Atlas, a high-profile, small-area analysis tool in economics research, we perform a thorough experimental evaluation of differentially private algorithms for simple linear regression on small datasets with tens to hundreds of records—a particularly challenging regime for differential privacy. In contrast, prior work on differentially private linear regression focused on multivariate linear regression on large datasets or asymptotic analysis. Through a range of experiments, we identify key factors that affect the relative performance of the algorithms. We find that algorithms based on robust estimators—in particular, the median-based estimator of Theil and Sen—perform best on small datasets (e.g., hundreds of datapoints), while algorithms based on Ordinary Least Squares or Gradient Descent perform better for large datasets. However, we also discuss regimes in which this general finding does not hold. Notably, the differentially private analogues of Theil–Sen (one of which was suggested in a theoretical work of Dwork and Lei) have not been studied in any prior experimental work on differentially private linear regression. Daniel Alabi, Audra McMillan, Jayshree Sarathy, Adam D. Smith 0001, Salil P. Vadhan |
Proc. Priv. Enhancing Technol. | 3 |