Stephen Meisenbacher

dblp:326/8319 · also Stephen Joseph Meisenbacher · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0001-9230-5001ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Security and privacy · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Spend Your Budget Wisely: Towards an Intelligent Distribution of the Privacy Budget in Differentially Private Text Rewriting
abstract
The task of Differentially Private Text Rewriting is a class of text privatization techniques in which (sensitive) input textual documents are rewritten under Differential Privacy (DP) guarantees. The motivation behind such methods is to hide both explicit and implicit identifiers that could be contained in text, while still retaining the semantic meaning of the original text, thus preserving utility. Recent years have seen an uptick in research output in this field, offering a diverse array of word-, sentence-, and document-level DP rewriting methods. Common to these methods is the selection of a privacy budget (i.e., the ε parameter), which governs the degree to which a text is privatized. One major limitation of previous works, stemming directly from the unique structure of language itself, is the lack of consideration of where the privacy budget should be allocated, as not all aspects of language, and therefore text, are equally sensitive or personal. In this work, we are the first to address this shortcoming, asking the question of how a given privacy budget can be intelligently and sensibly distributed amongst a target document. We construct and evaluate a toolkit of linguistics- and NLP-based methods used to allocate a privacy budget to constituent tokens in a text document. In a series of privacy and utility experiments, we empirically demonstrate that given the same privacy budget, intelligent distribution leads to higher privacy levels and more positive trade-offs than a naive distribution of epsilon. Our work highlights the intricacies of text privatization with DP, and furthermore, it calls for further work on finding more efficient ways to maximize the privatization benefits offered by DP in text rewriting.
Stephen Meisenbacher, Chaeeun Joy Lee, Florian Matthes
CODASPY1
2025 Leveraging Semantic Triples for Private Document Generation with Local Differential Privacy Guarantees
abstract
Many works at the intersection of Differential Privacy (DP) in Natural Language Processing aim to protect privacy by transforming texts under DP guarantees.This can be performed in a variety of ways, from word perturbations to full document rewriting, and most often under local DP.Here, an input text must be made indistinguishable from any other potential text, within some bound governed by the privacy parameter ε.Such a guarantee is quite demanding, and recent works show that privatizing texts under local DP can only be done reasonably under very high ε values.Addressing this challenge, we introduce DP-ST, which leverages semantic triples for neighborhoodaware private document generation under local DP guarantees.Through the evaluation of our method, we demonstrate the effectiveness of the divide-and-conquer paradigm, particularly when limiting the DP notion (and privacy guarantees) to that of a privatization neighborhood.When combined with LLM post-processing, our method allows for coherent text generation even at lower ε values, while still balancing privacy and utility.These findings highlight the importance of coherence in achieving balanced privatization outputs at reasonable ε levels.
Stephen Meisenbacher, Maulik Chevli, Florian Matthes
EMNLP1
2025 Lexical Substitution is not Synonym Substitution: On the Importance of Producing Contextually Relevant Word Substitutes
Juraj Vladika, Stephen Meisenbacher, Florian Matthes
ICAART (3)2
2025 "We are not Future-ready": Understanding AI Privacy Risks and Existing Mitigation Strategies from the Perspective of AI Developers in Europe
Alexandra Klymenko, Stephen Meisenbacher, Patrick Gage Kelley, Sai Teja Peddinti, Kurt Thomas, Florian Matthes
SOUPS2
2024 Just Rewrite It Again: A Post-Processing Method for Enhanced Semantic Similarity and Privacy Preservation of Differentially Private Rewritten Text
abstract
The study of Differential Privacy (DP) in Natural Language Processing often views the task of text privatization as a rewriting task, in which sensitive input texts are rewritten to hide explicit or implicit private information. In order to evaluate the privacy-preserving capabilities of a DP text rewriting mechanism, empirical privacy tests are frequently employed. In these tests, an adversary is modeled, who aims to infer sensitive information (e.g., gender) about the author behind a (privatized) text. Looking to improve the empirical protections provided by DP rewriting methods, we propose a simple post-processing method based on the goal of aligning rewritten texts with their original counterparts, where DP rewritten texts are rewritten again. Our results show that such an approach not only produces outputs that are more semantically reminiscent of the original inputs, but also texts which score on average better in empirical privacy evaluations. Therefore, our approach raises the bar for DP rewriting methods in their empirical privacy evaluations, providing an extra layer of protection against malicious adversaries.
Stephen Meisenbacher, Florian Matthes
ARES1
2024 A Comparative Analysis of Word-Level Metric Differential Privacy: Benchmarking the Privacy-Utility Trade-off
abstract
The application of Differential Privacy to Natural Language Processing techniques has emerged in relevance in recent years, with an increasing number of studies published in established NLP outlets. In particular, the adaptation of Differential Privacy for use in NLP tasks has first focused on the word-level, where calibrated noise is added to word embedding vectors to achieve “noisy” representations. To this end, several implementations have appeared in the literature, each presenting an alternative method of achieving word-level Differential Privacy. Although each of these includes its own evaluation, no comparative analysis has been performed to investigate the performance of such methods relative to each other. In this work, we conduct such an analysis, comparing seven different algorithms on two NLP tasks with varying hyperparameters, including the epsilon parameter, or privacy budget. In addition, we provide an in-depth analysis of the results with a focus on the privacy-utility trade-off, as well as open-source our implementation code for further reproduction. As a result of our analysis, we give insight into the benefits and challenges of word-level Differential Privacy, and accordingly, we suggest concrete steps forward for the research field.
Stephen Meisenbacher, Nihildev Nandakumar, Alexandra Klymenko, Florian Matthes
LREC/COLING1
2024 Thinking Outside of the Differential Privacy Box: A Case Study in Text Privatization with Language Model Prompting
abstract
The field of privacy-preserving Natural Language Processing has risen in popularity, particularly at a time when concerns about privacy grow with the proliferation of Large Language Models.One solution consistently appearing in recent literature has been the integration of Differential Privacy (DP) into NLP techniques.In this paper, we take these approaches into critical view, discussing the restrictions that DP integration imposes, as well as bring to light the challenges that such restrictions entail.To accomplish this, we focus on DP-PROMPT, a recent method for text privatization leveraging language models to rewrite texts.In particular, we explore this rewriting task in multiple scenarios, both with DP and without DP.To drive the discussion on the merits of DP in NLP, we conduct empirical utility and privacy experiments.Our results demonstrate the need for more discussion on the usability of DP in NLP and its benefits over non-DP approaches.
Stephen Meisenbacher, Florian Matthes
EMNLP1
2022 Understanding the Implementation of Technical Measures in the Process of Data Privacy Compliance: A Qualitative Study
abstract
Background: Modern privacy regulations, such as the General Data Protection Regulation (GDPR), address privacy in software systems in a technologically agnostic way by mentioning general ”technical measures” for data privacy compliance rather than dictating how these should be implemented. An understanding of the concept of technical measures and how exactly these can be handled in practice, however, is not trivial due to its interdisciplinary nature and the necessary technical-legal interactions.
Alexandra Klymenko, Oleksandr Kosenkov, Stephen Meisenbacher, Parisa Elahidoost, Daniel Méndez 0001, Florian Matthes
ESEM3