Benjamin Nguyen

dblp:69/6924 · DBLP profile ↗
← Back
32ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0003-0719-9684ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 21 · 1 first-author · 3 since 2021Security and privacy · 8 · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Verifiable Anomaly and Similarity Detection Using Matrix Profile in Private Time-series
Xavier Bultel, Charlène Jojon, Benjamin Nguyen, Haoying Zhang
ACISP (3)3
2026 Privacy Attacks on Matrix Profiles via Reconstruction Techniques
abstract
Matrix Profile (MP) is a data mining structure increasingly used for time series analysis in both academic and industrial contexts. Given its application to sensitive domains such as healthcare or energy monitoring, it is crucial to examine associated privacy risks, especially since MPs are often shared or processed in untrusted environments like the cloud. While recent studies suggest that MPs offer some privacy protection, this assumption remains largely untested. This paper analyzes the privacy risks of MP publication through the lens of EU data protection law, focusing on singling-out, linkability, and inference risks. We introduce a reconstruction technique based on constraint optimization, capable of recovering approximate original time series from their MPs, leading to severe privacy attacks. Experiments on real-world datasets reveal vulnerabilities to all attack types, with reconstructed series reaching up to 0.99 Pearson Correlation with the original.
Haoying Zhang, Nicolas Anciaux, Benjamin Nguyen, Fabien Girard, José María de Fuentes, Adrien Boiret
Proc. Priv. Enhancing Technol.3
2025 Cryptographic Commitments on Anonymizable Data
abstract
Local Differential Privacy (LDP) mechanisms consist of (locally) adding controlled noise to data in order to protect the privacy of their owner. In this paper, we introduce a new cryptographic primitive called LDP commitment. Usually, a commitment ensures that the committed value cannot be modified before it is revealed. In the case of an LDP commitment, however, the value is revealed after being perturbed by an LDP mechanism. Opening an LDP commitment therefore requires a proof that the mechanism has been correctly applied to the value, to ensure that the value is still usable for statistical purposes. We also present a security model for this primitive, in which we define the hiding and binding properties. Finally, we present a concrete scheme for an LDP staircase mechanism (generalizing the randomized response technique), based on classical cryptographic tools and standard assumptions. We provide an implementation in Rust that demonstrates its practical efficiency (the generation of a commitment requires just a few milliseconds).On the application side, we show how our primitive can be used to ensure simultaneously privacy, usability and traceability of medical data when it is used for statistical studies in an open science context. We consider a scenario where a hospital provides sensitive patients data signed by doctors to a research center after it has been anonymized, so that the research center can verify both the provenance of the data (i.e. verify the doctors’ signatures even though the data has been noised) and that the data has been correctly anonymized (i.e. is usable even though it has been anonymized).
Xavier Bultel, Céline Chevalier, Charlène Jojon, Diandian Liu, Benjamin Nguyen
EuroS&P5
2025 TELESAFE - Detecting Private/Work Boundary Crossings in Energy Consumption Trails in Telework
abstract
Teleworking has become a social gain following the COVID-19 lock-downs. In many professions, remote work is becoming a common practice, either at the employee's home or in a shared space nearby. However, this creates an implicit private/work-life tension as private activities may be carried out during work time and vice versa. Detecting boundary crossings is of outmost relevance - they serve as evidence of the workers' breaks and right to rest. However, this must be achieved without excessive surveillance. Existing activity recognition techniques either do not address the border crossing problem or require a priori training. To address this issue, this article proposes TELESAFE , a boundary crossing detector solution for teleworking. TELESAFE does not require any training nor instrumentation of the teleworker home and can be run locally in resource-constrained devices. To illustrate its suitability, it is applied on electric consumption trails so as to enable self and third-party assessment (e.g., work inspectors) on working conditions. Results on real-world datasets show a Fscore over 90% for identifying private activities involving one or more devices with usage patterns of varying lengths. Interestingly, TELESAFE outperforms Machine and Deep-Learning approaches in the most complex settings, without the burden of training.
Haoying Zhang, Mariem Brahem, Nicolas Anciaux, Benjamin Nguyen, José María de Fuentes
Proc. VLDB Endow.4
2024 A new PET for Data Collection via Forms with Data Minimization, Full Accuracy and Informed Consent
abstract
International audience
Nicolas Anciaux, Sabine Frittella, Baptiste Joffroy, Benjamin Nguyen, Guillaume Scerri
EDBT4
2024 Cohesive Database Neighborhoods for Differential Privacy: Mapping Relational Databases to RDF
Sara Taki, Adrien Boiret, Cédric Eichler, Benjamin Nguyen
WISE (5)4
2023 Demo: Data Minimization and Informed Consent in Administrative Forms
abstract
This article proposes a demonstration implementing the data minimization privacy principle, focusing on reducing data collected by government administrations through forms. Data minimization is defined in many privacy regulations worldwide, but has not seen extensive real-world application. We propose a model based on logic and game theory and show that it is possible to create a practical and efficient solution for a real French welfare benefit case.
Nicolas Anciaux, Sabine Frittella, Baptiste Joffroy, Benjamin Nguyen
CCS4
2022 Mouse-tracking meta-cognitive ratings of comprehension during garden-path sentences
Benjamin Nguyen, Michael J. Spivey
CogSci1
2022 Privacy Analysis with a Distributed Transition System and a Data-Wise Metric
Siva Anantharaman, Sabine Frittella, Benjamin Nguyen
PSD3
2019 Personal Data Management Systems: The security and functionality standpoint
Nicolas Anciaux, Philippe Bonnet, Luc Bouganim, Benjamin Nguyen, Philippe Pucheral, Iulian Sandu Popa, Guillaume Scerri
Inf. Syst.4
2017 Managing Distributed Queries under Personalized Anonymity Constraints
abstract
The benefit of performing Big data computations over individual's microdata is manifold, in the medical, energy or transportation fields to cite only a few, and this interest is growing with the emergence of smart disclosure initiatives around the world. However, these computations often expose microdata to privacy leakages , explaining the reluctance of individuals to participate in studies despite the privacy guarantees promised by statistical institutes. This paper proposes a novel approach to push personalized privacy guarantees in the processing of database queries so that individuals can disclose different amounts of information (i.e. data at different levels of accuracy) depending on their own perception of the risk. Moreover, we propose a decentralized computing infrastructure based on secure hardware enforcing these personalized privacy guarantees all along the query execution process. A performance analysis conducted on a real platform shows the effectiveness of the approach.
Axel Michel, Benjamin Nguyen, Philippe Pucheral
DATA2
2016 DatShA : A Data Sharing Algebra for access control plans
abstract
International audience
Luc Bouganim, Athanasia Katsouraki, Benjamin Nguyen
EDBT3
2016 Private and Scalable Execution of SQL Aggregates on a Secure Decentralized Architecture
abstract
Current applications, from complex sensor systems (e.g., quantified self) to online e-markets, acquire vast quantities of personal information that usually end up on central servers where they are exposed to prying eyes. Conversely, decentralized architectures that help individuals keep full control of their data complexify global treatments and queries, impeding the development of innovative services. This article aims precisely at reconciling individual's privacy on one side and global benefits for the community and business perspectives on the other. It promotes the idea of pushing the security to secure hardware devices controlling the data at the place of their acquisition. Thanks to these tangible physical elements of trust, secure distributed querying protocols can reestablish the capacity to perform global computations, such as Structured Query Language (SQL) aggregates, without revealing any sensitive information to central servers. This article studies how to secure the execution of such queries in the presence of honest-but-curious and malicious attackers. It also discusses how the resulting querying protocols can be integrated in a concrete decentralized architecture. Cost models and experiments on SQL/Asymmetric Architecture (AA), our distributed prototype running on real tamper-resistant hardware, demonstrate that this approach can scale to nationwide applications.
Quoc-Cuong To, Benjamin Nguyen, Philippe Pucheral
ACM Trans. Database Syst.2
2015 Limiting Data Exposure in Multi-Label Classification Processes
abstract
Administrative services such social care, tax reduction, and many others using complex decision processes, request individuals to provide large amounts of private data items, in order to calibrate their proposal to the specific situation of the applicant. This data is subsequently processed and stored by the organization. However, all the requested information is not needed to reach the same decision. We have recently proposed an approach, termed Minimum Exposure, to reduce the quantity of information provided by the users, in order to protect her privacy, reduce processing costs for the organization, and financial lost in the case of a data breach. In this paper, we address the case of decision making processes based on sets of classifiers, typically multi-label classifiers. We propose a practical implementation using state of the art multi-label classifiers, and analyze the effectiveness of our solution on several real multi-label data sets.
Nicolas Anciaux, Danae Boutara, Benjamin Nguyen, Michalis Vazirgiannis
Fundam. Informaticae3
2014 Tutorial: Managing Personal Data with Strong Privacy Guarantees
abstract
International audience
Nicolas Anciaux, Benjamin Nguyen, Iulian Sandu Popa
EDBT2
2014 Privacy-Preserving Query Execution using a Decentralized Architecture and Tamper Resistant Hardware
abstract
Current applications, from complex sensor systems (e.g. quantified self) to online e-markets acquire vast quantities of personal information which usually ends-up on central servers. Decentralized architectures, devised to help individuals keep full control of their data, hinder global treatments and queries, impeding the development of services of great interest. This paper promotes the idea of pushing the security to the edges of applications, through the use of secure hardware devices controlling the data at the place of their acquisition. To solve this problem, we propose secure distributed querying protocols based on the use of a tangible physical element of trust, reestablishing the capacity to perform global computations without revealing any sensitive information to central servers. There are two main problems when trying to support SQL in this context: perform joins and perform aggregations. In this paper, we study the subset of SQL queries without joins and show how to secure their execution in the presence of honest-but-curious attackers. Cost models and experiments demonstrate that this approach can scale to nationwide infrastructures.
Quoc-Cuong To, Benjamin Nguyen, Philippe Pucheral
EDBT2
2014 METAP: revisiting Privacy-Preserving Data Publishing using secure devices
Tristan Allard, Benjamin Nguyen, Philippe Pucheral
Distributed Parallel Databases2
2014 SQL/AA: Executing SQL on an Asymmetric Architecture
abstract
Current applications, from complex sensor systems (e.g. quantified self) to online e-markets acquire vast quantities of personal information which usually end-up on central servers. This information represents an unprecedented potential for user customized applications and business (e.g., car insurance billing, carbon tax, traffic decongestion, resource optimization in smart grids, healthcare surveillance, participatory sensing). However, the PRISM affair has shown that public opinion is starting to wonder whether these new services are not bringing us closer to science fiction dystopias. It has become clear that centralizing and processing all one's data on a single server is a major problem with regards to privacy concerns. Conversely, decentralized architectures, devised to help individuals keep full control of their data, complexify global treatments and queries, often impeding the development of innovative services and applications.
Quoc-Cuong To, Benjamin Nguyen, Philippe Pucheral
Proc. VLDB Endow.2
2013 I can do text analytics!: designing development tools for novice developers
abstract
Text analytics, an increasingly important application domain, is hampered by the high barrier to entry due to the many conceptual difficulties novice developers encounter. This work addresses the problem by developing a tool to guide novice developers to adopt the best practices employed by expert developers in text analytics and to quickly harness the full power of the underlying system. Taking a user centered task analytical approach, the tool development went through multiple design iterations and evaluation cycles. In the latest evaluation, we found that our tool enables novice developers to develop high quality extractors on par with the state of art within a few hours and with minimal training. Finally, we discuss our experience and lessons learned in the context of designing user interfaces to reduce the barriers to entry into complex domains of expertise.
Huahai Yang, Daina Pupons Wickham, Laura Chiticariu, Yunyao Li 0001, Benjamin Nguyen, Arnaldo Carreno-Fuentes
CHI5
2013 Trusted Cells: A Sea Change for Personal Data Services
Nicolas Anciaux, Philippe Bonnet, Luc Bouganim, Benjamin Nguyen, Iulian Sandu Popa, Philippe Pucheral
CIDR4
2013 MinExp-card: limiting data collection using a smart card
abstract
Online services such as social care, tax services, bank loans and many others, request individuals to fill in application forms with hundreds of private data items, in order to calibrate their offer. In practice, far too much data is requested, leading to over data disclosure. As shown in our previous works, avoiding this problem would (1) improve the privacy of the applicants and (2) decrease costs for service providers. We demonstrate here a prototype designed and implemented in partnership with the General Council of Yvelines District in France. The prototype targets forms used to calibrate social care for dependant people. To maintain the privacy of the decision process used to calibrate the social care, we propose a smartcard implementation. We will show that a 50% reduction of the items exposed in application forms can be achieved, explore the quality and scalability of our smartcard implementation, and demonstrate its scope.
Nicolas Anciaux, Walid Bezza, Benjamin Nguyen, Michalis Vazirgiannis
EDBT3
2013 Personal Data Management with Secure Hardware: How to Keep Your Data at Hand
abstract
We review existing solutions for personal data management, present a functional architecture for such decentralized alternatives, expose recent techniques dealing with embedded data management and global query processing in this architecture, and conclude by presenting existing and future implementations.
Nicolas Anciaux, Benjamin Nguyen, Iulian Sandu Popa
MDM (2)2
2012 Limiting data collection in application forms: A real-case application of a founding privacy principle
abstract
Application forms are often used by companies and administrations to collect personal data about applicants and tailor services to their specific situation. For example, taxes rates, social care, or personal loans, are usually calibrated based on a set of personal data collected through application forms. In the eyes of privacy laws and directives, the set of personal data collected to achieve a service must be restricted to the minimum necessary. This reduces the impact of data breaches both in the interest of service providers and applicants. In this article, we study the problem of limiting data collection in those application forms, used to collect data and subsequently feed decision making processes. In practice, the set of data collected is far excessive because application forms are filled in without any means to know what data will really impact the decision. To overcome this problem, we propose a reverse approach, where the set of strictly required data items to fill in the application form can be computed on the user's side. We formalize the underlying NP Hard optimization problem, propose algorithms to compute a solution, and validate them with experiments. Our proposal leads to a significant reduction of the quantity of personal data filled in application forms while still reaching the same decision.
Nicolas Anciaux, Benjamin Nguyen, Michalis Vazirgiannis
PST2
2011 Sanitizing Microdata without Leak: Combining Preventive and Curative Actions
Tristan Allard, Benjamin Nguyen, Philippe Pucheral
ISPEC2
2011 Towards a Safe Realization of Privacy-Preserving Data Publishing Mechanisms
abstract
This article addresses the issue of adapting the traditional model of Privacy-Preserving Data Publishing (PPDP) to an environment composed of a large number of tamper-resistant Secure Portable Tokens (SPTs) containing private personal data. Our model assumes that the SPTs seldom connect to a highly available but untrusted infrastructure. We illustrate the problem by studying the feasability of the simple generalization privacy mechanism to enforce k-anonymity.
Tristan Allard, Benjamin Nguyen, Philippe Pucheral
Mobile Data Management (2)2
2011 Safe realization of the Generalization privacy mechanism
abstract
An increasing number of surveys and articles high-light the failure of database servers to keep confidential data really private. Even without considering their vulnerability against external or internal attacks, mere negligences often lead to privacy disasters. The advent of powerful smart portable tokens, combining the security of smart card microcontrollers with the storage capacity of NAND Flash chips, introduces today credible alternatives to the systematic centralization of personal data on servers. Individuals can now store their personal data (e.g., their medical folder) in their own smart tokens, kept under their control, and never disclose in clear their private data to the outside untrusted world. However, this new opportunity of managing and protecting personal data conflicts with the objective of implementing knowledge-based decision making tools on top of centralized data. This paper precisely addresses this issue and proposes to adapt the traditional Generalization privacy mechanism to an environment composed of a large set of tamper-resistant smart portable tokens seldom connected to a highly available but untrusted infrastructure. This conjunction of hypothesis makes the problem fundamentally different from any previously studied privacy-preserving data publishing problem we are aware of.
Tristan Allard, Benjamin Nguyen, Philippe Pucheral
PST2
2010 Secure Personal Data Servers: a Vision Paper
abstract
An increasing amount of personal data is automatically gathered and stored on servers by administrations, hospitals, insurance companies, etc. Citizen themselves often count on internet companies to store their data and make them reliable and highly available through the internet. However, these benefits must be weighed against privacy risks incurred by centralization. This paper suggests a radically different way of considering the management of personal data. It builds upon the emergence of new portable and secure devices combining the security of smart cards and the storage capacity of NAND Flash chips. By embedding a full-fledged Personal Data Server in such devices, user control of how her sensitive data is shared by others (by whom, for how long, according to which rule, for which purpose) can be fully reestablished and convincingly enforced. To give sense to this vision, Personal Data Servers must be able to interoperate with external servers and must provide traditional database services like durability, availability, query facilities, transactions. This paper proposes an initial design for the Personal Data Server approach, identifies the main technical challenges associated with it and sketches preliminary solutions. We expect that this paper will open exciting perspectives for future database research.
Tristan Allard, Nicolas Anciaux, Luc Bouganim, Yanli Guo, Lionel Le Folgoc, Benjamin Nguyen, Philippe Pucheral, Indrajit Ray, Indrakshi Ray, Shaoyi Yin
Proc. VLDB Endow.6
2008 WebContent: efficient P2P Warehousing of web data
abstract
We present the WebContent platform for managing distributed repositories of XML and semantic Web data. The platform allows integrating various data processing building blocks (crawling, translation, semantic annotation, full-text search, structured XML querying, and semantic querying), presented as Web services, into a large-scale efficient platform. Calls to various services are combined inside ActiveXML [8] documents, which are XML documents including service calls. An ActiveXML optimizer is used to: ( i ) efficiently distribute computations among sites; ( ii ) perform XQuery-specific optimizations by leveraging an algebraic XQuery optimizer; and ( iii ) given an XML query, chose among several distributed indices the most appropriate in order to answer the query.
Serge Abiteboul, Tristan Allard, Philippe Chatalic, Georges Gardarin, A. Ghitescu, François Goasdoué, Ioana Manolescu, Benjamin Nguyen, M. Ouazara, A. Somani, Nicolas Travers, Gabriel Vasile, Spyros Zoupanos
Proc. VLDB Endow.8
2007 P2PTester: a tool for measuring P2P platform performance
abstract
The current abundance and complexity of P2P architectures makes it extremely difficult to assess their performance. P2PTester is the first tool devised to interface with, and measure the performance of, existing P2P data management platforms. We isolate basic components present in current P2P platforms, and insert "hooks" for P2PTester to capture, analyze and trace the interactions taking place in the underlying distributed system.
Bogdan Butnaru, Florin Dragan 0001, Georges Gardarin, Ioana Manolescu, Benjamin Nguyen, Radu Pop, Nicoleta Preda, Laurent Yeh
ICDE5
2004 THESUS, a Closer View on Web Content Management Enhanced with Link Semantics
abstract
With the unstoppable growth of the world wide Web, the great success of Web search engines, such as Google and AltaVista, users now turn to the Web whenever looking for information. However, many users are neophytes when it comes to computer science, yet they are often specialists of a certain domain. These users would like to add more semantics to guide their search through world wide Web material, whereas currently most search features are based on raw lexical content. We show how the use of the incoming links of a page can be used efficiently to classify a page in a concise manner. This enhances the browsing and querying of Web pages. We focus on the tools needed in order to manage the links and their semantics. We further process these links using a hierarchy of concepts, akin to an ontology, and a thesaurus. This work is demonstrated by an prototype system, called THESUS, that organizes thematic Web documents into semantic clusters. Our contributions are the following: 1) a model and language to exploit link semantics information, 2) the THESUS prototype system, 3) its innovative aspects and algorithms, more specifically, the novel similarity measure between Web documents applied to different clustering schemes (DB-Scan and COBWEB), and 4) a thorough experimental evaluation proving the value of our approach.
Iraklis Varlamis, Michalis Vazirgiannis, Maria Halkidi, Benjamin Nguyen
IEEE Trans. Knowl. Data Eng.4
2003 THESUS: Organizing Web document collections based on link semantics
Maria Halkidi, Benjamin Nguyen, Iraklis Varlamis, Michalis Vazirgiannis
VLDB J.2
2001 Monitoring XML Data on the Web
abstract
We consider the monitoring of a flow of incoming documents. More precisely, we present here the monitoring used in a very large warehouse built from XML documents found on the web. The flow of documents consists in XML pages (that are warehoused) and HTML pages (that are not). Our contributions are the following:
Benjamin Nguyen, Serge Abiteboul, Gregory Cobena, Mihai Preda
SIGMOD Conference1