VLDB 2026 Research / reviewers in the wild / expert
Michael Shoemate
dblp:224/1003
· DBLP profile ↗
5ranked-venue papers
0as first author
4since 2021 · last 2024
0000-0003-1190-3363ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Artificial intelligence and machine learning · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | "I inherently just trust that it works": Investigating Mental Models of Open-Source Libraries for Differential PrivacyabstractDifferential privacy (DP) is a promising framework for privacy-preserving data science, but recent studies have exposed challenges in bringing this theoretical framework for privacy into practice. These tensions are particularly salient in the context of open-source software libraries for DP data analysis, which are emerging tools to help data stewards and analysts build privacy-preserving data pipelines for their applications. While there has been significant investment into such libraries, we need further inquiry into the role of these libraries in promoting understanding of and trust in DP, and in turn, the ways in which design of these open-source libraries can shed light on the challenges of creating trustworthy data infrastructures in practice. In this study, we use qualitative methods and mental models approaches to analyze the differences between conceptual models used to design open-source DP libraries and mental models of DP held by users. Through a two-stage study design involving formative interviews with 5 developers of open-source DP libraries and user studies with 17 data analysts, we find that DP libraries often struggle to bridge the gaps between developer and user mental models. In particular, we highlight the tension DP libraries face in maintaining rigorous DP implementations and facilitating user interaction. We conclude by offering practical recommendations for further development of DP libraries. Patrick Song, Jayshree Sarathy, Michael Shoemate, Salil P. Vadhan |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2023 | Concurrent Composition for Interactive Differential Privacy with Adaptive Privacy-Loss ParametersabstractIn this paper, we study the concurrent composition of interactive mechanisms with adaptively chosen privacy-loss parameters. In this setting, the adversary can interleave queries to existing interactive mechanisms, as well as create new ones. We prove that every valid privacy filter and odometer for noninteractive mechanisms extends to the concurrent composition of interactive mechanisms if privacy loss is measured using (ε, δ)-DP, ƒ-DP, or Rényi DP of fixed order. Our results offer strong theoretical foundations for enabling full adaptivity in composing differentially private interactive mechanisms, showing that concurrency does not affect the privacy guarantees. We also provide an implementation for users to deploy in practice. Samuel Haney, Michael Shoemate, Grace Tian, Salil P. Vadhan, Andrew Vyrros, Vicki Xu, Wanrong Zhang 0001 |
CCS | 2 |
| 2022 | Widespread Underestimation of Sensitivity in Differentially Private Libraries and How to Fix ItabstractWe identify a new class of vulnerabilities in implementations of differential privacy. Specifically, they arise when computing basic statistics such as sums, thanks to discrepancies between the implemented arithmetic using finite data types (namely, ints or floats) and idealized arithmetic over the reals or integers. These discrepancies cause the sensitivity of the implemented statistics (i.e., how much one individual's data can affect the result) to be much larger than the sensitivity we expect. Consequently, essentially all differential privacy libraries fail to introduce enough noise to hide individual-level information as required by differential privacy, and we show that this may be exploited in realistic attacks on differentially private query systems. In addition to presenting these vulnerabilities, we also provide a number of solutions, which modify or constrain the way in which the sum is implemented in order to recover the idealized or near-idealized bounds on sensitivity. Sílvia Casacuberta, Michael Shoemate, Salil P. Vadhan, Connor Wagaman |
CCS | 2 |
| 2021 | An Ecosystem of Applications for Modeling Political ViolenceabstractConflict researchers face many challenges, including (1) how to model conflicts, (2) how to measure them, (3) how to manage their spatio-temporal character, and (4) how to handle a potential abundance of information and explanation. In this paper, we describe an ecosystem of tools designed for use by subject matter experts that addresses these challenges. Three case studies show workflows that are facilitated by this ecosystem. Aline Bessa, Sonia Castelo Quispe, Rémi Rampin, Aécio S. R. Santos, Michael Shoemate, Vito D'Orazio, Juliana Freire |
SIGMOD Conference | 5 |
| 2019 | Modeling and Forecasting Armed Conflict: AutoML with Human-Guided Machine LearningabstractMachine learning has made slow inroads into quantitative social science due to both a mismatch of machine learning's strengths to the causal and inferential tasks domain researchers pursue [1] and also a lack of algorithmic training among many domain experts [2]. However, conflict research- the empirical examination of political unrest, violence and civil war-has seen a growing emphasis on prediction and forecasting models. We describe automated machine learning (AutoML) to identify models, and human-guided machine learning (HGML), and show how these can incorporate domain knowledge and research requirements into model selection and assessment, and provide high quality machine learning pipelines to domain experts comparable to state-of-the-literature solutions. We examine three peer-reviewed papers with predictive models of conflict [3, 4, 5] and run their data through our HGML system using multiple AutoML engines and find this system produces slightly elevated performance on each paper's model, without any ML expertise required of the user. Our research has three takeaways for computational social science. First, predictive models of conflict would benefit from even minimal applications of AutoML; Secondly, human-guided machine learning offers the attractive option of constraining AutoML systems to address the kinds of questions conflict researchers assess with predictive models; Finally, current existing AutoML implementations produce divergent solutions and so can be productively harnessed in parallel. Vito D'Orazio, James Honaker, Raman Prasady, Michael Shoemate |
IEEE BigData | 4 |