Antoine Boutet

dblp:27/4142 · DBLP profile ↗
← Back
11ranked-venue papers in the field
7as first author
6since 2021 · last 2025
0000-0002-4057-416XORCID · corroborated

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 7 (4 first)Big Data, Cloud & Distributed Data Systems · 2 (1 first)Database Systems & Data Management · 1 (1 first)Data Mining & Knowledge Discovery · 1 (1 first)
YearPublicationVenuePosition
2025 "I'm Not for Sale" - Perceptions and Limited Awareness of Privacy Risks by Digital Natives About Location Data
abstract
Although mobile devices benefit users in their daily lives in numerous ways, they also raise several privacy concerns. For instance, they can reveal sensitive information that can be inferred from location data. This location data is shared through service providers as well as mobile applications. Understanding how and with whom users share their location data as well as users’ perception of the underlying privacy risks, are important notions to grasp in order to design usable privacy-enhancing technologies. In this work, we perform a quantitative and qualitative analysis of smartphone users’ awareness, perception and self-reported behavior towards location data-sharing through a survey of n = 99 young adult participants (i.e., digital natives). We compare stated practices with actual behaviors to better understand their mental models, and survey participants’ understanding of privacy risks before and after the inspection of location traces and the information that can be inferred therefrom. Our empirical results show that participants have risky privacy practices: about 54% of participants underestimate the number of mobile applications to which they have granted access to their data, and 33% forget or do not think of revoking access to their data. Furthermore, most of the participants do not have a realistic perception of privacy risks and have generally heard little about privacy-related scandals. Also, by using a demonstrator to perform inferences from location data, we observe that slightly more than half of participants (57%) are surprised by the extent of potentially inferred information, and that 47% intend to reduce access to their data via permissions as a result of using the demonstrator. Last, a majority of participants have little knowledge of the tools to better protect themselves, but are nonetheless willing to follow suggestions to improve privacy (51%). Educating people, including digital natives, about privacy risks through transparency tools seems a promising approach.
Antoine Boutet, Victor Morel
ICWSM1
2024 On the Alignment of Group Fairness with Attribute Privacy
Jan Aalmoes, Vasisht Duddu, Antoine Boutet
WISE (2)3
2024 Synthetic Data: Generate Avatar Data on Demand
Thomas Lebrun, Louis Béziaud, Tristan Allard, Antoine Boutet, Sébastien Gambs, Mohamed Maouche
WISE (5)4
2023 Toward training NLP models to take into account privacy leakages
abstract
With the rise of machine learning and data-driven models especially in the field of Natural Language Processing (NLP), a strong demand for sharing data between organisations has emerged. However datasets are usually composed of personal data and thus subject to numerous regulations which require anonymization before disseminating the data. In the medical domain for instance, patient records are extremely sensitive and private, but the de-identification of medical documents is a complex task. Recent advances in NLP models have shown encouraging results in this field, but the question of whether deploying such models is safe remains.In this paper, we evaluate three privacy risks on NLP models trained on sensitive data. Specifically, we evaluate counterfactual memorization, which corresponds to rare and sensitive information which has too much influence on the model. We also evaluate membership inference as well as the ability to extract verbatim training data from the model. With this evaluation, we can cure data at risk from the training data and calibrate hyper parameters to provide a supplementary utility and privacy trade-off to the usual mitigation strategies such as using differential privacy. We exhaustively illustrate the privacy leakage of NLP models through a use-case using medical texts and discuss the impact of both the proposed methodology and mitigation schemes.
Gaspard Berthelier, Antoine Boutet
IEEE Big Data2
2023 Towards an evolution in the characterization of the risk of re-identification of medical images
abstract
As facial recognition technology proliferates, concerns emerge regarding its application to medical imaging, specifically Magnetic Resonance Imaging (MRI). This paper investigates privacy risks associated with MRI data, including re-identification through social network photographs and sensitive attribute inference. The exponential growth in MRI quality coincides with the increasing sophistication of facial recognition tools, raising the potential for re-identification using medical images. Our attack involves reconstructing faces and applying facial recognition techniques to extract identifying features that can be compared to photographs. Legal frameworks like GDPR mandate the assessment and protection of personal data, necessitating continuous risk evaluation. Beyond re-identification, we explore the inference of individual attributes from MRI images, such as age, gender, and ethnic group. This research assesses the privacy risks associated with MRI data by taking into account the evolution of facial recognition and reconstruction tools that have become increasingly accessible. We also show that facial hair removal technique on photographs increases the risk of re-identification. Overall, our results highlight vulnerabilities in sharing MRI data, emphasizing the need for enhanced privacy safeguards.
Antoine Boutet, Carole Frindel, Mohamed Maouche
IEEE Big Data1
2022 Inferring Sensitive Attributes from Model Explanations
abstract
Model explanations provide transparency into a trained machine learning model's blackbox behavior to a model builder. They indicate the influence of different input attributes to its corresponding model prediction. The dependency of explanations on input raises privacy concerns for sensitive user data. However, current literature has limited discussion on privacy risks of model explanations.
Vasisht Duddu, Antoine Boutet
CIKM2
2019 Inspect What Your Location History Reveals About You: Raising user awareness on privacy threats associated with disclosing his location data
abstract
Location is one of the most extensively collected personal data on mobile by applications and third-party services. However, how the location of users is actually processed in practice by the actors of targeted advertising ecosystem remains unclear. Nonetheless, these providers have a strong incentive to create very detailed profile of users to better monetize the collected data. End users are usually not aware about the strength and wide range of inference that can be performed from their mobility traces. In this demonstration, users interact with a web-based application to inspect their location history and to discover the inferential power of this kind of data. Moreover to better understand the possible countermeasures, users can apply a sanitization to protect their data and visualize the impact on both the mobility traces and the associated inferred information. The objective of this demonstration is to raise the user awareness on the profiling capabilities and the privacy threats associated with disclosing his location data as well as how sanitization mechanisms can be efficient to mitigate these privacy risks. In addition, by collecting users feedbacks on the personal information revealed and the usage of a geosanitization mechanism, we hope that this demonstration will also be useful to constitute a new and valuable dataset on users perceptions on these questions.
Antoine Boutet, Sébastien Gambs
CIKM1
2016 Being prepared in a sparse world: The case of KNN graph construction
abstract
K-Nearest-Neighbor (KNN) graphs have emerged as a fundamental building block of many on-line services providing recommendation, similarity search and classification. Constructing a KNN graph rapidly and accurately is, however, a computationally intensive task. As data volumes keep growing, speed and the ability to scale out are becoming critical factors when deploying a KNN algorithm. In this work, we present KIFF, a generic, fast and scalable KNN graph construction algorithm. KIFF directly exploits the bipartite nature of most datasets to which KNN algorithms are applied. This simple but powerful strategy drastically limits the computational cost required to rapidly converge to an accurate KNN solution, especially for sparse datasets. Our evaluation on a representative range of datasets show that KIFF provides, on average, a speed-up factor of 14 against recent state-of-the art solutions while improving the quality of the KNN approximation by 18%.
Antoine Boutet, Anne-Marie Kermarrec, Nupur Mittal, François Taïani
ICDE1
2015 C3PO: A Network and Application Framework for Spontaneous and Ephemeral Social Networks
Antoine Boutet, Stéphane Frénot, Frédérique Laforest, Pascale Launay, Nicolas Le Sommer, Yves Mahéo, Damien Reimert
WISE (2)1
2012 What's in Twitter: I Know What Parties are Popular and Who You are Supporting Now!
abstract
In modern politics, parties and individual candidates must have an online presence and usually have dedicated social media coordinators. In this context, we study the usefulness of analysing Twitter messages to identify both the characteristics of political parties and the political leaning of users. As a case study, we collected the main stream of Twitter related to the 2010 UK General Election during the associated period -- gathering around 1,150,000 messages from about 220,000 users. We examined the characteristics of the three main parties in the election and highlighted the main differences between parties. First, Lab our members were the most active and influential during the election while Conservative members were the most organized to promote their activities. Second, the websites and blogs that each political party's members supported are clearly different from those that all the other political parties' members supported. From these observations, we develop a simple and practical classification method which uses the number of Twitter messages referring to a particular political party. The experimental results showed that the proposed classification method achieved about 86% classification accuracy and outperforms other classification methods that require expensive costs for tuning classifier parameters and/or knowledge about network topology.
Antoine Boutet, Hyoungshick Kim, Eiko Yoneki
ASONAM1
2012 What's in Your Tweets? I Know Who You Supported in the UK 2010 General Election
Antoine Boutet, Hyoungshick Kim, Eiko Yoneki
ICWSM1