EDBT 2026 Demo / reviewers in the wild / expert
Yves-Alexandre de Montjoye
dblp:75/7560
· DBLP profile ↗
25ranked-venue papers
1as first author
17since 2021 · last 2026
0000-0002-2559-5616ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 15 · 14 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3Applied, interdisciplinary, general and emerging computing · 3Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Automated Privacy Risk Estimation of Limited Fixed Aggregate StatisticsabstractEmpirical inference attacks are a popular approach for evaluating the privacy risk of data release mechanisms in practice. While an active attack literature exists to evaluate machine learning models or synthetic data release, we currently lack comparable methods for fixed aggregate statistics, in particular when only a limited number of statistics are released. We here propose an inference attack framework against fixed aggregate statistics and an attribute inference attack called DeSIA (Deterministic-Stochastic Inference Attack). DeSIA consists of two modules: a deterministic module, which verifies whether the target user's attribute can be uniquely inferred, and a stochastic module, which estimates the most likely attribute value when it cannot be uniquely inferred. We instantiate DeSIA against the U.S. Census PPMF dataset and show it to identify risks missed by reconstruction-based attacks. In particular, we show DeSIA to be highly effective in identifying vulnerable users, achieving a true positive rate of 0.14 at a false positive rate of $10^{-3}$. We show DeSIA to outperform all reconstruction-based attacks even without access to a real-world auxiliary dataset. We show DeSIA to be robust to varying levels of noise addition and across varying numbers of released aggregate statistics. We perform an extensive ablation study of DeSIA and show how DeSIA can be successfully adapted to the membership inference task. Overall, our results show that (a) even a limited number of aggregate statistics can reveal sufficient information to expose users to risk from inference attacks, leaving some users disproportionately vulnerable, and (b) emphasize the need for formal privacy mechanisms and testing before aggregate statistics are released. Yifeng Mao, Bozhidar Stevanoski, Yves-Alexandre de Montjoye |
Proc. Priv. Enhancing Technol. | 3 |
| 2025 | The DCR Delusion: Measuring the Privacy Risk of Synthetic Data
Zexi Yao, Natasa Krco, Georgi Ganev, Yves-Alexandre de Montjoye |
ESORICS (1) | 4 |
| 2025 | Certification for Differentially Private Prediction in Gradient-Based TrainingabstractWe study private prediction where differential privacy is achieved by adding noise to the outputs of a non-private model. Existing methods rely on noise proportional to the global sensitivity of the model, often resulting in sub-optimal privacy-utility trade-offs compared to private training. We introduce a novel approach for computing dataset-specific upper bounds on prediction sensitivity by leveraging convex relaxation and bound propagation techniques. By combining these bounds with the smooth sensitivity mechanism, we significantly improve the privacy analysis of private prediction compared to global sensitivity-based approaches. Experimental results across real-world datasets in medical image classification and natural language processing demonstrate that our sensitivity bounds are can be orders of magnitude tighter than global sensitivity. Our approach provides a strong basis for the development of novel privacy preserving technologies. Matthew Wicker, Philip Sosnin, Igor Shilov, Adrianna Janik, Mark Niklas Müller, Yves-Alexandre de Montjoye, Adrian Weller, Calvin Tsay |
ICML | 6 |
| 2025 | Exploring the limits of strong membership inference attacks on large language modelsabstractState-of-the-art membership inference attacks (MIAs) typically require training many reference models, making it difficult to scale these attacks to large pre-trained language models (LLMs). As a result, prior research has either relied on weaker attacks that avoid training references (e.g., fine-tuning attacks), or on stronger attacks applied to small models and datasets. However, weaker attacks have been shown to be brittle and insights from strong attacks in simplified settings do not translate to today's LLMs. These challenges prompt an important question: are the limitations observed in prior work due to attack design choices, or are MIAs fundamentally ineffective on LLMs? We address this question by scaling LiRA--one of the strongest MIAs--to GPT-2 architectures ranging from 10M to 1B parameters, training references on over 20B tokens from the C4 dataset. Our results advance the understanding of MIAs on LLMs in four key ways. While (1) strong MIAs can succeed on pre-trained LLMs, (2) their effectiveness, remains limited (e.g., AUC<0.7) in practical settings. (3) Even when strong MIAs achieve better-than-random AUC, aggregate metrics can conceal substantial per-sample MIA decision instability: due to training randomness, many decisions are so unstable that they are statistically indistinguishable from a coin flip. Finally, (4) the relationship between MIA success and related LLM privacy metrics is not as straightforward as prior work has suggested. Jamie Hayes, Ilia Shumailov, Christopher A. Choquette-Choo, Matthew Jagielski, Georgios Kaissis, Milad Nasr, Meenatchi Sundaram Muthu Selva Annamalai, Niloofar Mireshghallah, Igor Shilov, Matthieu Meeus, Yves-Alexandre de Montjoye, Katherine Lee, Franziska Boenisch, Adam Dziedzic, A. Feder Cooper |
NeurIPS | 11 |
| 2025 | Free Record-Level Privacy Risk Evaluation Through Artifact-Based Methods
Joseph Pollock, Igor Shilov, Euodia Dodd, Yves-Alexandre de Montjoye |
USENIX Security Symposium | 4 |
| 2024 | QueryCheetah: Fast Automated Discovery of Attribute Inference Attacks Against Query-Based SystemsabstractQuery-based systems (QBSs) are one of the key approaches for sharing data. QBSs allow analysts to request aggregate information from a private protected dataset. Attacks are a crucial part of ensuring QBSs are truly privacy-preserving. The development and testing of attacks is however very labor-intensive and unable to cope with the increasing complexity of systems. Automated approaches have been shown to be promising but are currently extremely computationally intensive, limiting their applicability in practice. We here propose QueryCheetah, a fast and effective method for automated discovery of privacy attacks against QBSs. We instantiate QueryCheetah on attribute inference attacks and show it to discover stronger attacks than previous methods while being 18 times faster than the state-of-the-art automated approach. We then show how QueryCheetah allows system developers to thoroughly evaluate the privacy risk, including for various attacker strengths and target individuals. We finally show how QueryCheetah can be used out-of-the-box to find attacks in larger syntaxes and workarounds around ad-hoc defenses. Bozhidar Stevanoski, Ana-Maria Cretu 0002, Yves-Alexandre de Montjoye |
CCS | 3 |
| 2024 | Re-pseudonymization Strategies for Smart Meter Data Are Not Robust to Deep Learning Profiling AttacksabstractSmart meters, devices measuring the electricity and gas consumption of a household, are currently being deployed at a fast rate throughout the world. The data they collect are extremely useful, including in the fight against climate change. However, these data and the information that can be inferred from them are highly sensitive. Re-pseudonymization, i.e., the frequent replacement of random identifiers over time, is widely used to share smart meter data while mitigating the risk of re-identification. We here show how, in spite of re-pseudonymization, households' consumption records can be pieced together with high accuracy in large-scale datasets. More specifically, we propose the first deep learning-based profiling attack against re-pseudonymized smart meter data. Our attack combines neural network embeddings, which are used to extract features from weekly consumption records and are tailored to the smart meter identification task, with a nearest neighbor classifier. We evaluate six neural networks architectures as the embedding model. Our results suggest that the Transformer and CNN-LSTM architectures vastly outperform previous methods as well as other architectures, successfully identifying the correct household 73.4% of the time among 5139 households based on electricity and gas consumption records (54.5% for electricity only). We further show that the features extracted by the embedding model maintain their effectiveness when transferred to a set of users disjoint from the one used to train the model. Finally, we extensively evaluate the robustness of our results. In particular, we show that the accuracy of the attack only slowly decreases with the size of the dataset, how less frequent re-pseudonymization will further increase accuracy, and how an attacker can evaluate the likelihood of a match to be correct. Taken together, our results strongly suggest that even frequent re-pseudonymization strategies can be reversed, strongly limiting their ability to prevent re-identification in practice. Ana-Maria Cretu 0002, Miruna Rusu, Yves-Alexandre de Montjoye |
CODASPY | 3 |
| 2024 | Copyright Traps for Large Language ModelsabstractQuestions of fair use of copyright-protected content to train Large Language Models (LLMs) are being actively debated. Document-level inference has been proposed as a new task: inferring from black-box access to the trained model whether a piece of content has been seen during training. SOTA methods however rely on naturally occurring memorization of (part of) the content. While very effective against models that memorize significantly, we hypothesize - and later confirm - that they will not work against models that do not naturally memorize, e.g. medium-size 1B models. We here propose to use copyright traps, the inclusion of fictitious entries in original content, to detect the use of copyrighted materials in LLMs with a focus on models where memorization does not naturally occur. We carefully design a randomized controlled experimental setup, inserting traps into original content (books) and train a 1.3B LLM from scratch. We first validate that the use of content in our target model would be undetectable using existing methods. We then show, contrary to intuition, that even medium-length trap sentences repeated a significant number of times (100) are not detectable using existing methods. However, we show that longer sequences repeated a large number of times can be reliably detected (AUC=0.75) and used as copyright traps. Beyond copyright applications, our findings contribute to the study of LLM memorization: the randomized controlled setup enables us to draw causal relationships between memorization and certain sequence properties such as repetition in model training data and perplexity. Matthieu Meeus, Igor Shilov, Manuel Faysse, Yves-Alexandre de Montjoye |
ICML | 4 |
| 2024 | Did the Neurons Read your Book? Document-level Membership Inference for Large Language Models
Matthieu Meeus, Shubham Jain 0006, Marek Rei, Yves-Alexandre de Montjoye |
USENIX Security Symposium | 4 |
| 2024 | Investigating the Effect of Misalignment on Membership Privacy in the White-box SettingabstractMachine learning models have been shown to leak sensitive information about their training datasets. Models are increasingly deployed on devices, raising concerns that white-box access to the model parameters increases the attack surface compared to black-box access which only provides query access. Directly extending the shadow modelling technique from the black-box to the white-box setting has been shown, in general, not to perform better than black-box only attacks. A potential reason is misalignment, a known characteristic of deep neural networks. In the shadow modelling context, misalignment means that, while the shadow models learn similar features in each layer, the features are located in different positions. We here present the first systematic analysis of the causes of misalignment in shadow models and show the use of a different weight initialisation to be the main cause. We then extend several re-alignment techniques, previously developed in the model fusion literature, to the shadow modelling context, where the goal is to re-align the layers of a shadow model to those of the target model.We show re-alignment techniques to significantly reduce the measured misalignment between the target and shadow models. Finally, we perform a comprehensive evaluation of white-box membership inference attacks (MIA). Our analysis reveals that internal layer activation-based MIAs suffer strongly from shadow model misalignment, while gradient-based MIAs are only sometimes significantly affected. We show that re-aligning the shadow models strongly improves the former's performance and can also improve the latter's performance, although less frequently. On the CIFAR10 dataset with a false positive rate of 1%, white-box MIA using re-aligned shadow models improves the true positive rate by 4.5%.Taken together, our results highlight that on-device deployment increases the attack surface and that the newly available information can be used to build more powerful attacks. Ana-Maria Cretu 0002, Yves-Alexandre de Montjoye, Shruti Tople |
Proc. Priv. Enhancing Technol. | 3 |
| 2024 | A Zero Auxiliary Knowledge Membership Inference Attack on Aggregate Location DataabstractLocation data is frequently collected from populations and shared in aggregate form to guide policy and decision making. However, the prevalence of aggregated data also raises the privacy concern of membership inference attacks (MIAs). MIAs infer whether an individual's data contributed to the aggregate release. Although effective MIAs have been developed for aggregate location data, these require access to an extensive auxiliary dataset of individual traces over the same locations, which are collected from a similar population. This assumption is often impractical given common privacy practices surrounding location data. To measure the risk of an MIA performed by a realistic adversary, we develop the first Zero Auxiliary Knowledge (ZK) MIA on aggregate location data, which eliminates the need for an auxiliary dataset of real individual traces. Instead, we develop a novel synthetic approach, such that suitable synthetic traces are generated from the released aggregate. We also develop methods to correct for bias and noise, to show that our synthetic-based attack is still applicable when privacy mechanisms are applied prior to release. Using two large-scale location datasets, we demonstrate that our ZK MIA matches the state-of-the-art Knock-Knock (KK) MIA across a wide range of settings, including popular implementations of differential privacy (DP) and suppression of small counts. Furthermore, we show that ZK MIA remains highly effective even when the adversary only knows a small fraction (10%) of their target's location history. This demonstrates that effective MIAs can be performed by realistic adversaries, highlighting the need for strong DP protection. Vincent Guan, Florent Guépin, Ana-Maria Cretu 0002, Yves-Alexandre de Montjoye |
Proc. Priv. Enhancing Technol. | 4 |
| 2023 | Achilles' Heels: Vulnerable Record Identification in Synthetic Data Publishing
Matthieu Meeus, Florent Guépin, Ana-Maria Cretu 0002, Yves-Alexandre de Montjoye |
ESORICS (2) | 4 |
| 2023 | Deep perceptual hashing algorithms with hidden dual purpose: when client-side scanning does facial recognitionabstractEnd-to-end encryption (E2EE) provides strong technical protections to individuals from interferences. Governments and law enforcement agencies around the world have however raised concerns that E2EE also allows illegal content to be shared undetected. Client-side scanning (CSS), using perceptual hashing (PH) to detect known illegal content before it is shared, is seen as a promising solution to prevent the diffusion of illegal content while preserving encryption. While these proposals raise strong privacy concerns, proponents of the solutions have argued that the risk is limited as the technology has a limited scope: detecting known illegal content. In this paper, we show that modern perceptual hashing algorithms are actually fairly flexible pieces of technology and that this flexibility could be used by an adversary to add a secondary hidden feature to a client-side scanning system. More specifically, we show that an adversary providing the PH algorithm can "hide" a secondary purpose of face recognition of a target individual alongside its primary purpose of image copy detection. We first propose a procedure to train a dual-purpose deep perceptual hashing model by jointly optimizing for both the image copy detection and the targeted facial recognition task. Second, we extensively evaluate our dual-purpose model and show it to be able to reliably identify a target individual 67% of the time while not impacting its performance at detecting illegal content. We also show that our model is neither a general face detection nor a facial recognition model, allowing its secondary purpose to be hidden. Finally, we show that the secondary purpose can be enabled by adding a single illegal looking image to the database. Taken together, our results raise concerns that a deep perceptual hashing-based CSS system could turn billions of user devices into tools to locate targeted individuals. Shubham Jain 0006, Ana-Maria Cretu 0002, Antoine Cully, Yves-Alexandre de Montjoye |
SP | 4 |
| 2023 | Web Privacy: A Formal Adversarial Model for Query ObfuscationabstractThe queries we perform, the searches we make, and the websites we visit – this sensitive data is collected at scale by companies as part of the services they provide. Query obfuscation, intertwining the genuine queries of the user with artificial queries, has been proposed as a solution to protect the privacy of individuals on the web.We here present a formal model and formulate through attack models three privacy requirements for obfuscators: (1)indistinguishability, that the user query should be hard to identify; (2)coverage, that its topic should be hard to identify; and (3)imprecision, that the query should still be hard to identify for an attacker with additional auxiliary information. The latter is needed to make the former two guarantees “future-proof”. Using our framework, we derive two important results for obfuscators. First, we show that indistinguishability imposes strong bounds on the coverage and imprecision achievable by an obfuscator. Second, we prove an important tradeoff between coverage and imprecision, which inherently limits the strength and robustness of the privacy guarantees that an obfuscator can provide. We then introduce a family of obfuscators with provable indistinguishability guarantees, which we callk–ball obfuscators, and show, for a range of parameter values, the achievable coverage and imprecision. We show empirically that our theoretical tradeoff holds, and that its bound is not tight in practice: even in a simple idealized setting, there is a significant gap between practical coverage and imprecision guarantees, and the optimal bounds. While obfuscators have proven popular with the general public, all obfuscators currently available provide adhoc guarantees, and have been shown to be vulnerable to attacks, putting the data of users at risk. We hope this work to be a first step towards a robust evaluation of the properties of query obfuscators and the development of principled obfuscators. Florimond Houssiau, Thibaut Liénart, Julien M. Hendrickx, Yves-Alexandre de Montjoye |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2022 | QuerySnout: Automating the Discovery of Attribute Inference Attacks against Query-Based SystemsabstractAlthough query-based systems (QBS) have become one of the main solutions to share data anonymously, building QBSes that robustly protect the privacy of individuals contributing to the dataset is a hard problem. Theoretical solutions relying on differential privacy guarantees are difficult to implement correctly with reasonable accuracy, while ad-hoc solutions might contain unknown vulnerabilities. Evaluating the privacy provided by QBSes must thus be done by evaluating the accuracy of a wide range of privacy attacks. However, existing attacks against QBSes require time and expertise to develop, need to be manually tailored to the specific systems attacked, and are limited in scope. In this paper, we develop QuerySnout, the first method to automatically discover vulnerabilities in query-based systems. QuerySnout takes as input a target record and the QBS as a black box, analyzes its behavior on one or more datasets, and outputs a multiset of queries together with a rule to combine answers to them in order to reveal the sensitive attribute of the target record. QuerySnout uses evolutionary search techniques based on a novel mutation operator to find a multiset of queries susceptible to lead to an attack, and a machine learning classifier to infer the sensitive attribute from answers to the queries selected. We showcase the versatility of QuerySnout by applying it to two attack scenarios (assuming access to either the private dataset or to a different dataset from the same distribution), three real-world datasets, and a variety of protection mechanisms. We show the attacks found by QuerySnout to consistently equate or outperform, sometimes by a large margin, the best attacks from the literature. We finally show how QuerySnout can be extended to QBSes that require a budget, and apply QuerySnout to a simple QBS based on the Laplace mechanism. Taken together, our results show how powerful and accurate attacks against QBSes can already be found by an automated system, allowing for highly complex QBSes to be automatically tested "at the pressing of a button". We believe this line of research to be crucial to improve the robustness of systems providing privacy-preserving access to personal data in theory and in practice. Ana-Maria Cretu 0002, Florimond Houssiau, Antoine Cully, Yves-Alexandre de Montjoye |
CCS | 4 |
| 2022 | Pool Inference Attacks on Local Differential Privacy: Quantifying the Privacy Guarantees of Apple's Count Mean Sketch in Practice
Andrea Gadotti, Florimond Houssiau, Meenatchi Sundaram Muthu Selva Annamalai, Yves-Alexandre de Montjoye |
USENIX Security Symposium | 4 |
| 2022 | Adversarial Detection Avoidance Attacks: Evaluating the robustness of perceptual hashing-based client-side scanning
Shubham Jain 0006, Ana-Maria Cretu 0002, Yves-Alexandre de Montjoye |
USENIX Security Symposium | 3 |
| 2020 | Inference of node attributes from social network assortativity
Dounia Mulders, Cyril de Bodt, Johannes Bjelland, Alex Pentland, Michel Verleysen, Yves-Alexandre de Montjoye |
Neural Comput. Appl. | 6 |
| 2020 | Towards Matching User Mobility Traces in Large-Scale DatasetsabstractThe problem of unicity and reidentifiability of records in large-scale databases has been studied in different contexts and approaches, with focus on preserving privacy or matching records from different data sources. With an increasing number of service providers nowadays routinely collecting location traces of their users on unprecedented scales, there is a pronounced interest in the possibility of matching records and datasets based on spatial trajectories. Extending previous work on reidentifiability of spatial data and trajectory matching, we present the first large-scale analysis of user matchability in real mobility datasets on realistic scales, i.e. among two datasets that consist of several million people's mobility traces, coming from a mobile network operator and transportation smart card usage. We extract the relevant statistical properties which influence the matching process and analyze their impact on the matchability of users. We show that for individuals with typical activity in the transportation system (those making 3-4 trips per day on average), a matching algorithm based on the co-occurrence of their activities is expected to achieve a 16.8 percent success only after a one-week long observation of their mobility traces, and over 55 percent after four weeks. We show that the main determinant of matchability is the expected number of co-occurring records in the two datasets. Finally, we discuss different scenarios in terms of data collection frequency and give estimates of matchability over time. We show that with higher frequency data collection becoming more common, we can expect much higher success rates in even shorter intervals. Dániel Kondor, Behrooz Hashemian, Yves-Alexandre de Montjoye, Carlo Ratti |
IEEE Trans. Big Data | 3 |
| 2019 | OPAL: High performance platform for large-scale privacy-preserving location data analyticsabstractMobile phones and other ubiquitous technologies are generating vast amounts of high-resolution location data. This data has been shown to have a great potential for the public good, e.g. to monitor human migration during crises or to predict the spread of epidemic diseases. Location data is, however, considered one of the most sensitive types of data, and a large body of research has shown the limits of traditional data anonymization methods for big data. Privacy concerns have so far strongly limited the use of location data collected by telcos, especially in developing countries.In this paper, we introduce OPAL (for OPen ALgorithms), an open-source, scalable, and privacy-preserving platform for location data. At its core, OPAL relies on an open algorithm to extract key aggregated statistics from location data for a wide range of potential use cases. We first discuss how we designed the OPAL platform, building a modular and resilient framework for efficient location analytics. We then describe the layered mechanisms we have put in place to protect privacy and discuss the example of a population density algorithm. We finally evaluate the scalability and extensibility of the platform and discuss related work.The code will be open-sourced on GitHub upon publication. Axel Oehmichen, Shubham Jain 0006, Andrea Gadotti, Yves-Alexandre de Montjoye |
IEEE BigData | 4 |
| 2019 | Differentially Private Compressive K-meansabstractThis work addresses the problem of learning from large collections of data with privacy guarantees. The sketched learning framework proposes to deal with the large scale of datasets by compressing them into a single vector of generalized random moments, from which the learning task is then performed. We modify the standard sketching mechanism to provide differential privacy, using addition of Laplace noise combined with a subsampling mechanism (each moment is computed from a subset of the dataset). The data can be divided between several sensors, each applying the privacy-preserving mechanism locally, yielding a differentially-private sketch of the whole dataset when reunited. We apply this framework to the k-means clustering problem, for which a measure of utility of the mechanism in terms of a signal-to-noise ratio is provided, and discuss the obtained privacy-utility tradeoff. Vincent Schellekens, Antoine Chatalic, Florimond Houssiau, Yves-Alexandre de Montjoye, Laurent Jacques, Rémi Gribonval |
ICASSP | 4 |
| 2019 | When the Signal is in the Noise: Exploiting Diffix's Sticky Noise
Andrea Gadotti, Florimond Houssiau, Luc Rocher, Benjamin Livshits, Yves-Alexandre de Montjoye |
USENIX Security Symposium | 5 |
| 2019 | UNVEIL: Capture and Visualise WiFi Data LeakagesabstractIn the past few years, numerous privacy vulnerabilities have been discovered in the WiFi standards and their implementations for mobile devices. These vulnerabilities allow an attacker to collect large amounts of data on the device user, which could be used to infer sensitive information such as religion, gender, and sexual orientation. Solutions for these vulnerabilities are often hard to design and typically require many years to be widely adopted, leaving many devices at risk. Shubham Jain 0006, Eden Bensaid, Yves-Alexandre de Montjoye |
WWW | 3 |
| 2017 | Modeling the Temporal Nature of Human Behavior for Demographics Prediction
Bjarke Felbo, Pål Roe Sundsøy, Alex Pentland, Sune Lehmann, Yves-Alexandre de Montjoye |
ECML/PKDD (3) | 5 |
| 2016 | bandicoot: a Python Toolbox for Mobile Phone Metadataabstractbandicoot is an open-source Python toolbox to extract more than 1442 features from standard mobile phone metadata. bandicoot makes it easy for machine learning researchers and practitioners to load mobile phone data, to analyze and visualize them, and to extract robust features which can be used for various classification and clustering tasks. Emphasis is put on ease of use, consistency, and documentation. bandicoot has no dependencies and is distributed under MIT license. Yves-Alexandre de Montjoye, Luc Rocher, Alex Pentland |
J. Mach. Learn. Res. | 1 |