Héber Hwang Arcolezi

dblp:248/5342 · DBLP profile ↗
← Back
28ranked-venue papers
17as first author
23since 2021 · last 2026
0000-0001-8059-7094ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 11 · 6 first-author · 10 since 2021Databases, data management, data science and information retrieval · 7 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 5 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author
YearPublicationVenuePosition
2026 Private Frequency Estimation via Residue Number Systems
abstract
We present Modular Subset Selection (MSS), a new algorithm for locally differentially private (LDP) frequency estimation. Given a universe of size k and n users, our ε-LDP mechanism encodes each input via a Residue Number System (RNS) over ℓ pairwise-coprime moduli m0, ..., m_{ℓ−1}, and reports a randomly chosen index j ∊ [ℓ] along with the perturbed residue using the statistically optimal Subset Selection (SS) mechanism. This design reduces the user communication cost from Θ(ω log₂(k/ω)) bits required by standard SS (with ω ≈ k/(e^ε+1)) down to ⌈ log₂ ℓ ⌉ + ⌈ log₂ m_j ⌉ bits, where m_j < k. Server-side decoding runs in Θ(n + r k ℓ) time, where r is the number of LSMR iterations. In practice, with well-conditioned moduli (i.e., constant r and ℓ = Θ(log k)), this becomes Θ(n + k log k). We prove that MSS achieves worst-case MSE within a constant factor of state-of-the-art protocols such as SS and Projective Geometry Response (PGR), while avoiding the algebraic prerequisites and dynamic-programming decoder required by PGR. Empirically, MSS matches the estimation accuracy of SS, PGR, and RAPPOR across realistic (k, ε) settings, while offering faster decoding than PGR and shorter user messages than SS. Lastly, by sampling from multiple moduli and reporting only a single perturbed residue, MSS achieves the lowest reconstruction-attack success rate among all evaluated LDP protocols.
Héber Hwang Arcolezi
AAAI1
2026 Estimating the True Distribution of Data Collected with Randomized Response
abstract
Randomized Response (RR) is a protocol designed to collect and analyze categorical data with local differential privacy guarantees. It has been used as a building block of mechanisms deployed by Big tech companies to collect app or web users' data. Each user reports an automatic random alteration of their true value to the analytics server, which then estimates the histogram of the true unseen values of all users using a debiasing rule to compensate for the added randomness. A known issue is that the standard debiasing rule can yield a vector with negative values (which can not be interpreted as a histogram), and there is no consensus on the best fix. An elegant but slow solution is the Iterative Bayesian Update algorithm (IBU), which converges to the Maximum Likelihood Estimate (MLE) as the number of iterations goes to infinity. This paper bypasses IBU by providing a simple formula for the exact MLE of RR and compares it with other estimation methods experimentally to help practitioners decide which one to use.
Carlos Antonio Pinzón, Ehab ElSalamouny, Lucas Massot, Alexis Miller, Héber Hwang Arcolezi, Catuscia Palamidessi
AAAI5
2026 How Tough Is Location Anonymization? Re-identifying 100K Real-User Trajectories in Japan
abstract
Mobility traces are among the most revealing forms of personal data, yet trajectory releases are often protected only by ad hoc transformations. We stress-test such practices on recently-released YJMob100K, an anonymized dataset of 100,000 user trajectories in Japan. First, we show that the applied protection leaves enough spatial and temporal structure to recover both the real-world geographic frame and the actual calendar timeline by exploiting density signatures, urban correlations, and temporal activity profiles. On top of this reconstruction, we quantify privacy risks through trajectory-level metrics that capture spatio-temporal k-anonymity, m-point unicity, home-work and multi-anchor uniqueness, and exposure to secluded and sensitive locations. These metrics reveal extensive re-identification surfaces: a small number of observations, anchors, or sensitive venues often suffices to uniquely pinpoint users or their social neighborhoods. Finally, we evaluate representative sanitization strategies: geo-indistinguishability, local differential privacy, and aggressive spatial de-structuring; and observe a consistent pattern: strong privacy parameters destroy downstream utility, while utility-preserving settings leave structural leakage largely intact. Overall, our findings show that current sanitization techniques are insufficient for large-scale mobility data, and they highlight the urgent need for trajectory-aware privacy mechanisms and stronger publication standards.
Abhishek Kumar Mishra 0001, Mathieu Cunche, Héber Hwang Arcolezi
AsiaCCS3
2026 Data Poisoning in Longitudinal Local Differential Privacy: Attacks and Analysis
Antonio A. Marreiras Neto, Héber Hwang Arcolezi, Javam C. Machado
DBSec2
2026 Revisiting Locally Differentially Private Protocols: Towards Better Trade-Offs in Privacy, Utility, and Attack Resistance
abstract
Local Differential Privacy (LDP) offers strong privacy protection, especially in settings in which the server collecting the data is untrusted. However, designing LDP mechanisms that achieve an optimal trade-off between privacy, utility and robustness to adversarial inference and integrity attacks remains challenging. In this work, we introduce a general multi-objective optimization framework for refining LDP protocols, enabling the joint optimization of privacy and utility under various adversarial settings. While our framework is flexible to accommodate multiple privacy and security attacks as well as utility metrics, in this paper, we specifically optimize for Attacker Success Rate (ASR) under \emph{data reconstruction attack} as a concrete measure of privacy leakage and Mean Squared Error (MSE) as a measure of utility. Complementarily, we evaluate integrity-oriented threats through data poisoning attacks, providing an additional adversarial perspective. More precisely, we systematically revisit these trade-offs by analyzing eight state-of-the-art LDP frequency estimation protocols and proposing refined counterparts that leverage tailored optimization techniques. Experimental results demonstrate that our proposed adaptive mechanisms consistently outperform their non-adaptive counterparts, achieving substantial reductions in ASR while preserving utility, and pushing closer to the ASR-MSE Pareto frontier. By bridging the gap between theoretical guarantees and real-world vulnerabilities, our framework enables modular and context-aware deployment of LDP mechanisms with tunable privacy-utility-attackability trade-offs.
Héber Hwang Arcolezi, Sébastien Gambs
ICDE1
2026 Where to Intervene? Benchmarking Fairness-Aware Learning on Differentially Private Synthetic Tabular Data
abstract
Machine learning models are increasingly deployed in high-stakes domains, raising concerns about both privacy and fairness. Differential Privacy (DP) has become a gold standard for privacy-preserving data analysis, while fairness-aware mechanisms aim to mitigate discrimination against underrepresented groups. However, these objectives can conflict: DP often amplifies disparities across demographic groups, and little is known about whether established fairness interventions remain effective under DP constraints. In this work, we present, to our knowledge, the first systematic evaluation of fairness interventions on differentially private synthetic tabular data. Our benchmark centers on the Adaptive Iterative Mechanism (AIM), identified as the state-of-the-art marginal-based DP synthesizer in recent KDD & VLDB 2025 tutorials by Cormode et al. We thus evaluate fairness interventions across four datasets, multiple group fairness metrics, and three categories of mitigation strategies (pre-processing, in-processing, and post-processing) under a wide range of privacy budgets. We compare four pipeline configurations: (Baseline) training on original data; (DP-only) training on DP synthetic data; (Fair-only) applying fairness mechanisms on original data; and (DP+Fair) combining fairness mechanisms with DP synthetic data. Our results demonstrate that while DP alone can degrade both utility and fairness, applying fairness interventions can partially restore equitable outcomes. Among them, post-processing methods tend to provide more stable fairness–utility trade-offs across privacy budgets and synthesizers, achieving strong fairness improvements while preserving competitive utility relative to other intervention stages. We release all code, data, and experimental artifacts in an open-source repository (https://github.com/vinicius-verona/dp-fair-intervention-benchmark) to ensure full reproducibility and to support future research on the privacy-fairness-utility trade-off.
Vinícius Gabriel Angelozzi, Héber Hwang Arcolezi
Proc. Priv. Enhancing Technol.2
2026 Understanding Disclosure Risk in Differential Privacy with Applications to Noise Calibration and Auditing
Patricia Guerra-Balboa, Annika Sauer, Héber Hwang Arcolezi, Thorsten Strufe
Proc. VLDB Endow.3
2025 Demo: Exploring Utility and Attackability Trade-offs in Local Differential Privacy
abstract
Local Differential Privacy (LDP) provides strong, formal privacy guarantees without requiring a trusted curator, making it a promising approach for privacy-preserving data collection and analysis. However, despite extensive research, practitioners may struggle to understand how to tune LDP parameters and anticipate the impact on data utility and attack risks for their specific scenarios. To address this gap, we demonstrate LDP-Toolbox, the first interactive, web-based toolbox (implemented in Python) that enables practical, analytical visualization of trade-offs between privacy loss (ε), utility loss, and vulnerability to attacks. The toolbox supports exploration of these trade-offs using real-world datasets from different domains; in this demonstration, we focus on discrete personal attributes and location-based scenarios. By providing intuitive, visual insights, LDP-Toolbox lowers the barrier to deploying LDP in real applications and helps bridge the gap between theoretical guarantees and practical adoption. The toolbox is open-source on PyPI (https://pypi.org/project/ldp-toolbox) and a video is available on our GitHub repository (https://github.com/hharcolezi/ldp-toolbox).
Haoying Zhang, Abhishek Kumar Mishra 0001, Héber Hwang Arcolezi
CCS3
2025 Fair Play for Individuals, Foul Play for Groups? Auditing Anonymization's Impact on ML Fairness
abstract
Machine learning (ML) algorithms are heavily based on the availability of training data, which, depending on the domain, often includes sensitive information about data providers. This raises critical privacy concerns. Anonymization techniques have emerged as a practical solution to address these issues by generalizing features or suppressing data to make it more difficult to accurately identify individuals. Although recent studies have shown that privacy-enhancing technologies can influence ML predictions across different subgroups, thus affecting fair decision-making, the specific effects of anonymization techniques, such as k-anonymity, ℓ-diversity, and t-closeness, on ML fairness remain largely unexplored. In this work, we systematically audit the impact of anonymization techniques on ML fairness, evaluating both individual and group fairness. Our quantitative study reveals that anonymization can degrade group fairness metrics by up to fourfold. Conversely, similarity-based individual fairness metrics tend to improve under stronger anonymization, largely as a result of increased input homogeneity. By analyzing varying levels of anonymization across diverse privacy settings and data distributions, this study provides critical insights into the trade-offs between privacy, fairness, and utility, offering actionable guidelines for responsible AI development. Our code is publicly available at: https://github.com/hharcolezi/anonymity-impact-fairness.
Héber Hwang Arcolezi, Mina Alishahi, Adda-Akram Bendoukha, Nesrine Kaaniche
ECAI1
2025 Group fairness under obfuscated sensitive information
abstract
In the era of Big Data, the development of artificial intelligence (AI) systems presents both opportunities and challenges, particularly concerning privacy and fairness. While differential privacy (DP) has emerged as a robust methodology for preserving privacy in real-world applications, its local variant (LDP) specifically addresses trust issues by removing the reliance on a centralized server. Equally critical, conducting fairness audits of AI systems helps identify and mitigate discriminatory outcomes in machine learning. Although the relationship between DP and fairness is inherently multifaceted, this paper offers a detailed empirical examination of how collecting multi-dimensional sensitive attributes under LDP affects fairness in binary classification tasks. Our findings reveal that LDP can slightly improve fairness without substantially degrading model performance—challenging the notion that DP necessarily exacerbates unfairness. We demonstrate these results by evaluating seven state-of-the-art LDP protocols on three benchmark datasets, using established group fairness metrics. Moreover, we propose a novel privacy budget allocation scheme that incorporates varying domain sizes of sensitive attributes, achieving a superior privacy–utility–fairness trade-off compared to existing solutions.
Héber Hwang Arcolezi, Karima Makhlouf, Catuscia Palamidessi
J. Comput. Secur.1
2024 A Systematic and Formal Study of the Impact of Local Differential Privacy on Fairness: Preliminary Results
abstract
Machine learning (ML) algorithms rely primarily on the availability of training data, and, depending on the domain, these data may include sensitive information about the data providers, thus leading to significant privacy issues. Differential privacy (DP) is the predominant solution for privacy-preserving ML, and the local model of DP is the preferred choice when the server or the data collector are not trusted. Recent experimental studies have shown that local DP can impact ML prediction for different subgroups of individuals, thus affecting fair decision-making. However, the results are conflicting in the sense that some studies show a positive impact of privacy on fairness while others show a negative one. In this work, we conduct a systematic and formal study of the effect of local DP on fairness. Specifically, we perform a quantitative study of how the fairness of the decisions made by the ML model changes under local DP for different levels of privacy and data distributions. In particular, we provide bounds in terms of the joint distributions and the privacy level, delimiting the extent to which local DP can impact the fairness of the model. We characterize the cases in which privacy reduces discrimination and those with the opposite effect. We validate our theoretical findings on synthetic and real-world datasets. Our results are preliminary in the sense that, for now, we study only the case of one sensitive attribute, and only statistical disparity, conditional statistical disparity, and equal opportunity difference.
Karima Makhlouf, Tamara Stefanovic, Héber Hwang Arcolezi, Catuscia Palamidessi
CSF3
2024 Nob-MIAs: Non-biased Membership Inference Attacks Assessment on Large Language Models with Ex-Post Dataset Construction
Cédric Eichler, Nathan Champeil, Nicolas Anciaux, Alexandra Bensamoun, Héber Hwang Arcolezi, José María de Fuentes
WISE (3)5
2024 On the impact of multi-dimensional local differential privacy on fairness
Karima Makhlouf, Héber Hwang Arcolezi, Sami Zhioua, Ghassen Ben Brahim, Catuscia Palamidessi
Data Min. Knowl. Discov.2
2024 Revealing the True Cost of Locally Differentially Private Protocols: An Auditing Perspective
abstract
While the existing literature on Differential Privacy (DP) auditing predominantly focuses on the centralized model (e.g., in auditing the DP-SGD algorithm), we advocate for extending this approach to audit Local DP (LDP). To achieve this, we introduce the LDP-Auditor framework for empirically estimating the privacy loss of locally differentially private mechanisms. This approach leverages recent advances in designing privacy attacks against LDP frequency estimation protocols. More precisely, through the analysis of numerous state-of-the-art LDP protocols, we extensively explore the factors influencing the privacy audit, such as the impact of different encoding and perturbation functions. Additionally, we investigate the influence of the domain size and the theoretical privacy loss parameters ϵ and δ on local privacy estimation. In-depth case studies are also conducted to explore specific aspects of LDP auditing, including distinguishability attacks on LDP protocols for longitudinal studies and multidimensional data. Finally, we present a notable achievement of our LDP-Auditor framework, which is the discovery of a bug in a state-of-the-art LDP Python package. Overall, our LDP-Auditor framework as well as our study offer valuable insights into the sources of randomness and information loss in LDP protocols. These contributions collectively provide a realistic understanding of the local privacy loss, which can help practitioners in selecting the LDP mechanism and privacy parameters that best align with their specific requirements. We open-sourced LDP-Auditor in [4].
Héber Hwang Arcolezi, Sébastien Gambs
Proc. Priv. Enhancing Technol.1
2023 On the Utility Gain of Iterative Bayesian Update for Locally Differentially Private Mechanisms
Héber Hwang Arcolezi, Selene Leya Cerna Ñahuis, Catuscia Palamidessi
DBSec1
2023 (Local) Differential Privacy has NO Disparate Impact on Fairness
Héber Hwang Arcolezi, Karima Makhlouf, Catuscia Palamidessi
DBSec1
2023 Frequency Estimation of Evolving Data Under Local Differential Privacy
abstract
International audience
Héber Hwang Arcolezi, Carlos Antonio Pinzón, Catuscia Palamidessi, Sébastien Gambs
EDBT1
2023 On the Risks of Collecting Multidimensional Data Under Local Differential Privacy
abstract
The private collection of multiple statistics from a population is a fundamental statistical problem. One possible approach to realize this is to rely on the local model of differential privacy (LDP). Numerous LDP protocols have been developed for the task of frequency estimation of single and multiple attributes. These studies mainly focused on improving the utility of the algorithms to ensure the server performs the estimations accurately. In this paper, we investigate privacy threats (re-identification and attribute inference attacks) against LDP protocols for multidimensional data following two state-of-the-art solutions for frequency estimation of multiple attributes. To broaden the scope of our study, we have also experimentally assessed five widely used LDP protocols, namely, generalized randomized response, optimal local hashing, subset selection, RAPPOR and optimal unary encoding. Finally, we also proposed a countermeasure that improves both utility and robustness against the identified threats. Our contributions can help practitioners aiming to collect users' statistics privately to decide which LDP mechanism best fits their needs.
Héber Hwang Arcolezi, Sébastien Gambs, Jean-François Couchot, Catuscia Palamidessi
Proc. VLDB Endow.1
2022 Multi-Freq-LDPy: Multiple Frequency Estimation Under Local Differential Privacy in Python
Héber Hwang Arcolezi, Jean-François Couchot, Sébastien Gambs, Catuscia Palamidessi, Majid Zolfaghari
ESORICS (3)1
2022 Differentially private multivariate time series forecasting of aggregated human mobility with deep learning: Input or gradient perturbation?
Héber Hwang Arcolezi, Jean-François Couchot, Denis Renaud, Bechara al Bouna, Xiaokui Xiao
Neural Comput. Appl.1
2022 Privacy-Preserving Prediction of Victim's Mortality and Their Need for Transportation to Health Facilities
abstract
Emergency medical services (EMS) provide crucial prehospital care, such as in the case of cardiac arrest, where the victim requires immediate first-aid. For this reason, it is vital to improving EMS response time. This article proposes a novel methodology based on machine learning (ML) techniques to predict both the victims’ mortality and their need for transportation to health facilities using data gathered from the start of the emergency call until the Departmental Fire and Rescue Service of the Doubs (SDIS25) is notified. We first analyzed SDIS25 calls to find out associations between the call processing times and victims’ mortality, and to measure the variables’ importance. Next, we validated our proposed ML-based methodology, where mortality could be predicted with accuracy and area under the receiver operating characteristic curve (AUC) scores of 96.44% and 96.04%, respectively, while the need for transportation achieved an accuracy and AUC scores of 73.62% and 78.91%, respectively. What is more, we found out that it was still possible to predict both targets perturbating the input data by applyingk-anonymity and differential privacy techniques. In conclusion, the results showed the potential of ML for EMS, which can be used as a decision-support tool to early identify mortality and the use of resources (transportation) and, thus, help EMS to save more lives and avoid service disruptions.
Héber Hwang Arcolezi, Selene Leya Cerna Ñahuis, Jean-François Couchot, Christophe Guyeux, Abdallah Makhoul
IEEE Trans. Ind. Informatics1
2021 Random Sampling Plus Fake Data: Multidimensional Frequency Estimates With Local Differential Privacy
abstract
With local differential privacy (LDP), users can privatize their data and thus guarantee privacy properties before transmitting it to the server (a.k.a. the aggregator). One primary objective of LDP is frequency (or histogram) estimation, in which the aggregator estimates the number of users for each possible value. In practice, when a study with rich content on a population is desired, the interest is in the multiple attributes of the population, that is to say, in multidimensional data (d ≥ 2). However, contrary to the problem of frequency estimation of a single attribute (the majority of the works), the multidimensional aspect imposes to pay particular attention to the privacy budget. This one can indeed grow extremely quickly due to the composition theorem. To the authors' knowledge, two solutions seem to stand out for this task: 1) splitting the privacy budget for each attribute, i.e., send each value with ε d ≥-LDP (Spl), and 2) random sampling a single attribute and spend all the privacy budget to send it with ε-LDP (Smp). AlthoughSmp adds additional sampling error, it has proven to provide higher data utility than the formerSpl solution. However, we argue that aggregators (who are also seen as attackers) are aware of the sampled attribute and its LDP value, which is protected by a "less strict" eε probability bound (rather than e^ε/d ). This way, we propose a solution named Random S ampling plus Fake Data (RS+FD), which allows creatinguncertainty over the sampled attribute by generating fake data for each non-sampled attribute; RS+FD further benefits from amplification by sampling. We theoretically and experimentally validate our proposed solution on both synthetic and real-world datasets to show that RS+FD achieves nearly the same or better utility than the state-of-the-artSmp solution.
Héber Hwang Arcolezi, Jean-François Couchot, Bechara al Bouna, Xiaokui Xiao
CIKM1
2021 RISE controller tuning and system identification through machine learning for human lower limb rehabilitation via neuromuscular electrical stimulation
Héber Hwang Arcolezi, Willian R. B. M. Nunes, Rafael A. de Araujo, Selene Leya Cerna Ñahuis, Marcelo A. A. Sanches, Marcelo C. M. Teixeira, Aparecido Augusto de Carvalho
Eng. Appl. Artif. Intell.1
2020 Mobility modeling through mobile data: generating an optimized and open dataset respecting privacy
abstract
Modeling and understanding people's mobility at a temporal and geographical space are very strict requirements for developing better strategies of urban public and private transportation systems as well as establishing improved business techniques. This work proposes a random-search based approach to instantiate statistical indicators through an improved mobility scenario which provides specific information about people attending one or several days for some events. Then, we recreate that scenario with virtual humans, proposing a synthetic and open dataset that matches the original statistical data. The results show the proposed approach is very efficient to model people's mobility, and the generated data has a low error rate compared to the original one.
Héber Hwang Arcolezi, Jean-François Couchot, Oumaya Baala, Jean-Michel Contet, Bechara al Bouna, Xiaokui Xiao
IWCMC1
2020 A Comparison of LSTM and XGBoost for Predicting Firemen Interventions
Selene Leya Cerna Ñahuis, Christophe Guyeux, Héber Hwang Arcolezi, Raphaël Couturier, Guillaume Royer
WorldCIST (2)3
2020 Forecasting the number of firefighter interventions per region with local-differential-privacy-based data
Héber Hwang Arcolezi, Jean-François Couchot, Selene Leya Cerna Ñahuis, Christophe Guyeux, Guillaume Royer, Bechara al Bouna, Xiaokui Xiao
Comput. Secur.1
2019 A RISE-based Controller Fine-tuned by an Improved Genetic Algorithm for Human Lower Limb Rehabilitation via Neuromuscular Electrical Stimulation
abstract
In the last few years, several studies have been carried out showing that Functional Electrical Stimulation (FES) and Neuromuscular Electrical Stimulation (NMES) produce good therapeutic results in patients with Spinal Cord Injury (SCI). This paper presents the proposal of a fine-tuning method based on an Improved Genetic Algorithm (IGA) to a continuous and robust control technique for uncertain nonlinear systems named Robust Integral of the Sign of the Error (RISE), for knee joint control. Simulation results are provided for three paraplegic and one healthy identified patients on ideal and nonideal conditions. Although in the literature this controller presents good results without any fine tuning method, we provide an approach to improve it, even more, believing on the minimization of fatigue and other problems that often occurs in SCI patients treated with FES/NMES, by selecting adequately the gain parameters of the RISE controller.
Héber Hwang Arcolezi, Willian R. B. M. Nunes, Selene Leya Cerna Ñahuis, Marcelo A. A. Sanches, Marcelo C. M. Teixeira, Aparecido Augusto de Carvalho
CoDIT1
2019 Long Short-Term Memory for Predicting Firemen Interventions
abstract
Many environmental, economic and societal factors are leading fire brigades to be increasingly solicited, and they, therefore, face an ever-increasing number of interventions, most of the time with constant resources. On the other hand, these interventions are directly related to human activity, which itself is predictable: swimming pool drownings occur in summer while road accidents due to ice storms occur in winter. One solution to improve the response of firefighters with constant resources is therefore to predict their workload, i.e., their number of interventions per hour, based on explanatory variables conditioning human activity. The purpose of this article is to show that these interventions can indeed be predicted, in a nonabsurd way, from state-of-the-art tools such as recurrent long short-term memory neural networks (LSTM). From the list of interventions in the Doubs (France), we show that it is possible to build, from scratch, a neural network capable of reasonably predicting the interventions of 2017 from those of 2012-2016. While the results could be improved, they are already promising and would allow the actions of firefighters with a constant resource to be optimized.
Selene Leya Cerna Ñahuis, Christophe Guyeux, Héber Hwang Arcolezi, Raphaël Couturier, Guillaume Royer, Anna Diva P. Lotufo
CoDIT3